diff --git a/README.md b/README.md index 7758052c..63e9d5fa 100644 --- a/README.md +++ b/README.md @@ -29,19 +29,38 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **序列处理** | Biopython `Bio.Seq` | 序列对象、**MutableSeq可变序列**、互补、转录、翻译、序列特征 | ✅ | | **序列 I/O** | Biopython `Bio.SeqIO` | FASTA/FASTQ/GenBank 解析与写入 | ✅ | | **序列比对** | Biopython / scikit-bio | Needleman-Wunsch、Smith-Waterman、多序列比对、替换矩阵(BLOSUM/PAM) | ✅ | -| **BLAST解析** | Biopython `Bio.Blast` | BLAST结果解析、tabular/xml格式、HSP过滤、最佳匹配 | ✅ | -| **SearchIO** | Biopython `Bio.SearchIO` | 统一搜索结果模型、HMMER3解析、BLAT PSL解析、BLAST转换 | ✅ | +| **BLAST基础解析** | BioSeqs compatibility API | 历史tabular/XML标签解析、HSP过滤与最佳匹配 | ✅ | +| **现代BLAST XML** | Biopython `Bio.Blast` | XML1/XML2严格解析与规范写回、多query/report、参数/统计、描述与taxonomy、八类程序的链向/translated坐标路径 | ✅ | +| **SearchIO** | Biopython `Bio.SearchIO` | 统一搜索结果模型、HMMER3/Infernal解析、BLAT PSL解析、BLAST转换 | ✅ | +| **HH-suite HHR** | Biopython `Bio.Align.hhr` | HHsearch/HHblits HHR解析、profile-profile比对、命中筛选、坐标映射、序列化往返 | ✅ | +| **共享参考比对合并** | Biopython `Bio.Align.Alignment` | 合并共享同一参考序列的PWA/MSA、同步insertion slots、保留局部坐标与metadata、双向坐标映射 | ✅ | +| **Alignment详细统计** | Biopython `Bio.Align.Alignment.counts` | 左/内部/右 insertion/deletion、gap open/extend、identity/mismatch/positive、wildcard、替换矩阵和十二类affine gap评分 | ✅ | +| **PSL/PSLX成对比对** | Biopython `Bio.Align.psl` | 21/23列严格读写、核酸与translated DNA-protein路径、正反链坐标、block/gap统计、sequence-aware recount及坐标映射 | ✅ | +| **Alignment-aware SAM** | Biopython `Bio.Align.sam` | SAM header与typed tag严格读写、CIGAR坐标路径、正反链、soft/hard clipping、PHRED、MD/NM及坐标映射 | ✅ | +| **A2M状态感知多序列比对** | Biopython `Bio.Align.a2m` | match/insertion列状态、大小写与点语义、严格读写、坐标映射、插入槽、统计、共识及match-only投影 | ✅ | +| **EMBOSS alignment输出** | Biopython `Bio.Align.emboss` | srspair/pair/simple报告、多alignment与多序列、局部/反向坐标、consensus统计、坐标路径及规范往返 | ✅ | +| **Exonerate alignment输出** | Biopython `Bio.Align.exonerate` | cigar/vulgar严格读写、完整operation path、正反链与protein strand、3:1 translated坐标、双向映射及规范往返 | ✅ | +| **Exonerate C4文本报告** | Biopython `Bio.SearchIO.ExonerateIO.exonerate_text` | C4层次结果、3/4/5行模型、wrapped blocks、intron/NER/split codon/frameshift、蛋白质翻译与链感知坐标 | ✅ | +| **GCG MSF多序列比对** | Biopython `Bio.Align.msf` | AA/NA/PileUp严格解析、interleaved rows、标准GCG checksum、gap规范化、坐标路径、统计及canonical writer | ✅ | +| **NEXUS多序列比对** | Biopython `Bio.Align.nexus` | DATA/CHARACTERS/TAXA、nested comments、quoted taxa、sequential/interleaved MATRIX、MATCHCHAR、坐标统计及canonical writer | ✅ | +| **Stockholm注释型多序列比对** | Biopython `Bio.Align.stockholm` | 多记录严格读写、GF/GS/GR/GC、reference与database reference、insertion/deletion列、all-gap压缩、坐标统计及canonical writer | ✅ | +| **UCSC Chain成对比对** | Biopython `Bio.Align.chain` | 12/13字段严格读写、连续多记录、正反双轴绝对坐标路径、size/dt/dq块、双向位置/区间映射、反转及canonical writer | ✅ | +| **现代MAF多基因组比对** | Biopython `Bio.Align.maf` | track/header与a/s/i/e/q严格读写、正负链绝对坐标路径、任意component映射、MafIndex半开区间查询、多外显子拼接及canonical writer | ✅ | +| **现代BED成对比对** | Biopython `Bio.Align.bed` | BED3-BED12严格读写、正负链target/query路径、exon block重建、双向residue映射、半开区间查询及分级writer | ✅ | | **系统发育树** | Biopython `Bio.Phylo` | 树结构、Newick 解析、距离计算、可视化 | ✅ | | **PDB 结构** | Biopython `Bio.PDB` | 原子/残基/链解析、结构操作 | ✅ | +| **BinaryCIF** | Biopython `Bio.PDB.binary_cif` | MessagePack解析、七类逆编码、三态缺失值、类别查询、PDB Structure转换 | ✅ | +| **CE 结构比对** | Biopython `Bio.PDB.cealign` | CA/C4'引导原子、AFP路径搜索、CE显著性、QCP刚体叠合、全原子变换 | ✅ | | **SAM/BAM/VCF** | pysam | 比对文件、变异检测、基因型查询 | ✅ | | **FASTA 索引** | pyfaidx | 快速随机访问、.fai 索引 | ✅ | | **机器学习特征** | scikit-learn | k-mer 频率、氨基酸组成、理化性质 | ✅ | | **Biostrings** | Bioconductor Biostrings | IUPAC 支持、RSCU、复杂度、Tm 计算、模式匹配(matchPattern/vmatchPattern)、错配和插入缺失检测、回文序列查找 | ✅ | -| **GenomicRanges** | Bioconductor GenomicRanges | GRanges、区间操作、集合运算、precede/follow、coverage计算、distance_to_nearest | ✅ | +| **GenomicRanges** | Bioconductor GenomicRanges | GRanges/GRangesList、复合特征、区间操作、集合运算、precede/follow、coverage计算、distance_to_nearest | ✅ | | **plyranges** | Bioconductor plyranges | dplyr-like tidy verbs for GRanges: filter/mutate/select/arrange/rename/group_by+summarise/join_by、metadata管理 | ✅ | | **pheatmap** | Bioconductor pheatmap | 增强型热图可视化:层次聚类(complete/average/ward)、距离矩阵(euclidean/manhattan/correlation)、行/列注释、颜色方案、聚类间隙 | ✅ | | **factoextra** | Bioconductor factoextra | PCA/因子分析工具:特征值计算、方差解释率、个体/变量坐标、cos2质量、贡献度评分、维度描述 | ✅ | | **DESeq2** | Bioconductor DESeq2 | 差异表达分析、size factors归一化、分散度估计、负二项GLM拟合、Wald检验、LFC收缩 | ✅ | +| **miloR** | Bioconductor miloR | 单细胞KNN邻域采样、样本计数、负二项差异丰度、graph spatial FDR、SCE接入 | ✅ | | **dplyr** | R dplyr | DataFrame 数据操作 | ✅ | | **enrichplot** | Bioconductor enrichplot | 富集分析结果可视化、dotplot/barplot/heatmap/cnetplot/enrichment map | ✅ | | **IsoformSwitchAnalyzeR** | Bioconductor IsoformSwitchAnalyzeR | 转录本异构体切换分析、PSI/DPSI/DIF值计算、功能后果预测 | ✅ | @@ -68,8 +87,13 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | ✅ | Bio.protein_analysis | 蛋白质序列高级分析: Kyte-Doolittle疏水性滑动窗口、GOR二级结构预测、Hopp-Woods抗原性、跨膜区段预测、氨基酸/二肽/三肽组成、Shannon熵保守性评分 | | ✅ | Bio.PCD | 质谱PCD格式解析: 质谱图谱(Scan/RT/PEPMASS)、峰列表提取、总离子流色谱图(TIC)、基峰色谱图(BPC)、m/z范围过滤、前体离子信息、序列化往返 | | ✅ | Bio.PDB.Dice + Selection | PDB结构切割(链/残基/原子/模型提取)、B因子过滤、几何选择、结构统计、序列提取 | -| ✅ | Bio.Align.MAF | MAF (Multiple Alignment Format) 多序列比对格式解析、块处理、百分比一致性计算、统计分析、格式转换 | -| ✅ | Bio.Align.Mauve | Mauve 基因组比对格式解析、LCB(共线性块)检测、倒位检测、断点检测、基因组覆盖率、BED导出 | +| ✅ | Bio.Align.MAF | 宽松MAF块解析、选择/过滤、百分比一致性、统计分析与格式转换 | +| ✅ | Bio.Align.maf | 现代MAF document/track、严格a/s/i/e/q、绝对坐标路径、component映射、参考区间索引、多外显子拼接与规范往返 | +| ✅ | Bio.Align.bed | BED3-BED12 pairwise alignment、双轴链向坐标、block投影、双向residue映射、区间搜索与分级写回 | +| ✅ | Bio.Align.Mauve (legacy) | 历史 MAF-like 块的 LCB 重排摘要、倒位/断点检测、覆盖率与 BED 导出 | +| ✅ | Bio.Align.mauve | 现代 XMFA header/LCB 严格读写、combined/separate source、正负链坐标路径、索引、跨序列投影、统计与序列重建 | +| ✅ | Bio.Align.clustal | 现代CLUSTAL metadata与interleaved block严格读写、累计残基数、consensus、坐标路径、跨行投影、统计与规范写回 | +| ✅ | Bio.Blast XML | XML1/XML2 document/record/hit/HSP模型、多描述与taxonomy、Karlin-Altschul统计、frame/strand及1:1/3:1坐标路径、规范往返 | | ✅ | Bio.Stockholm | Stockholm 格式解析 (Pfam/Rfam比对格式)、二级结构注释、百分比一致性、保守性分析、FASTA转换 | | ✅ | Bio.PopGen (advanced) | 高级群体遗传学统计: Tajima's D, Fu & Li's D/F, McDonald-Kreitman检验, 等位基因频率谱, 中性分析 | | ✅ | Bio.SeqUtils.CodonUsage (advanced) | 高级密码子分析: CAI密码子适应指数、RSCU相对同义密码子使用、ENC有效密码子数、GC3偏斜、最优/稀有密码子检测 | @@ -107,7 +131,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | ✅ | GENIE3 | Bioconductor GENIE3基因调控网络推断: 回归树特征重要性、方差缩减、加权邻接矩阵、对称化网络 | | ✅ | decoupleR | Bioconductor decoupleR功能活性推断: WSum/WMean/Norm/ULM/MLM方法、先验知识网络(PKN)、调控子活性评分 | | ✅ | BayesSpace | Bioconductor BayesSpace空间转录组聚类: t分布混合模型、马尔可夫随机场(MRF)先验、EM算法、六边形/方形网格邻居 | -| ✅ | muscat | Bioconductor muscat单细胞差异状态分析: 伪批量聚合(Sum/Mean/Median)、EdgeR/DESeq2/Limma DS检验、BH-FDR校正、样本QC指标 | +| ✅ | muscat | Bioconductor muscat 1.27.4多样本多亚群差异状态分析: gene × cell严格合同、五类cluster-sample伪批量、任意满秩设计/多contrast、负二项IRLS、经验贝叶斯dispersion、DS/DD与局部/全局BH-FDR、两阶段确认及不可变SCE写回;保留基础聚合/QC兼容API | | ✅ | infercnv | Bioconductor infercnv单细胞拷贝数变异推断: 染色体位置排序基因、参考细胞比较、log2FC有界计算、金字塔权重基因组平滑、每细胞中位数中心化+噪声过滤、CNV分数+肿瘤细胞预测 | | ✅ | SCENIC | Bioconductor SCENIC单细胞调控网络推断与聚类: TF-target共表达模块(GENIE3风格)、Regulon构建(权重剪枝/cisTarget motif排名剪枝)、AUCell活性评分(recovery curve AUC)、二值化阈值(MeanStd/KMeans2/Median)、细胞状态聚类+主控调控因子识别 | | ✅ | CIBERSORT | 免疫细胞去卷积: 非负最小二乘(NNLS)求解细胞类型分数、投影梯度下降、LM22风格特征矩阵(40标记基因×10免疫细胞类型)、Pearson拟合优度+RMSE、分数归一化(Σ=1.0) | @@ -121,6 +145,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | ✅ | Bioconductor fishpond (Swish) | 非参数差异表达分析: Mann-Whitney-Wilcoxon秩和统计、置换检验、BH-FDR校正、log2FC方向判定 | | ✅ | Bioconductor MatrixGenerics | 矩阵行/列汇总统计: rowMeans/colMeans、rowSums/colSums、rowVars/colVars、rowSds/colSds、rowMedians/colMedians、rowMins/colMins、rowMaxs/colMaxs、rowRanges/colRanges、rowMad/colMad、rowCounts、rowAnys/colAnys、rowAlls/colAlls、块处理 | | ✅ | Bioconductor beachmat | 矩阵访问API: 列/行块处理(bmat_apply_col_blocks/bmat_apply_row_blocks)、线性迭代器(BmatIterator)、子集/转置/绑定、逐元素操作、类型安全访问 | +| ✅ | Bioconductor SparseArray | N维稀疏数组: 规范化COO、重复坐标合并、R列主序转换、切片/置换/绑定、稀疏算术、统计与矩阵乘法 | | ✅ | Bioconductor glmGamPoi | Gamma-Poisson广义线性模型: size factors估计、伪批量聚合、单基因拟合(IWLCS迭代加权最小二乘)、Wald差异表达检验、BH-FDR校正、线性代数求解、正态CDF与p值计算 | | ✅ | Bioconductor survival | 生存分析: Kaplan-Meier估计器(Greenwood标准误)、log-rank检验(两组比较)、Cox比例风险模型(Newton-Raphson偏似然拟合、Breslow ties)、卡方p值、中位生存期 | | ✅ | Bioconductor methylKit | 亚硫酸氢盐测序甲基化分析: 甲基化胞嘧啶统计、覆盖率过滤/归一化、Fisher精确检验差异甲基化、BH-FDR校正、DMR识别、样本相关性/聚类、BED导出 | @@ -148,8 +173,29 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | ✅ | velociraptor | 单细胞RNA velocity分析: 稳态线性回归(gamma/beta比估计)、EM算法动力学模型(alpha/beta/gamma参数估计)、velocity向量计算、KNN加权嵌入投影、根细胞识别 | | ✅ | Bio.Compass | COMPASS profile-profile比对输出解析: 版本提取、多记录解析(SW分数/E值/百分比一致性/比对序列)、E值与一致性过滤、比对长度统计、摘要生成 | | ✅ | Bio.SearchIO.ExonerateIO | Exonerate比对输出解析: vulgar格式解析(比对块三元组)、cigar格式解析、格式自动检测、分数过滤、内含子统计、vulgar/cigar字符串重建 | +| ✅ | Bio.SearchIO.ExonerateIO.exonerate_text | Exonerate C4人类可读报告: Document/Query/Hit/HSP/Fragment层次、3/4/5行模型、wrapped block拼接、intron/NER/split codon/frameshift、翻译与坐标投影 | | ✅ | Bio.PDB.mmcifio | mmCIF文件写入: Structure对象序列化(data block/header/atom_site loop)、20列原子坐标格式化、HETATM支持、值转义、round-trip验证 | | ✅ | Bio.SearchIO.InterproscanIO | InterProScan输出解析: TSV 14列格式解析(蛋白质ID/分析数据库/签名/位置/分数/IPR/GO)、按数据库/蛋白质过滤、GO条目提取、按蛋白质分组 | +| ✅ | Bio.SearchIO.InfernalIO | Infernal cmscan/cmsearch解析: tabular格式1/2/3自动检测、non-verbose文本与--noali、CM/HMM-only、Query/Hit/HSP/Fragment层次、正负链坐标、local-end多片段、过滤与SearchIO转换 | +| ✅ | Bioconductor variancePartition | 多随机截距线性混合模型、ML/REML方差分量、固定/随机/残差方差占比、precision weights、BLUP、dream contrast、数值Satterthwaite检验、BH-FDR与SummarizedExperiment接入 | +| ✅ | Bio.Align.Alignment shared-reference merge | 共享参考PWA/MSA合并、reference-boundary insertion slot同步、局部reference/query坐标、metadata、统计、MSA与aligned FASTA转换 | +| ✅ | Bio.Align.Alignment map/mapall | 零起始半开区间alignment path组合、局部clipping、gap与正反链传播、双向坐标查询、PSL、批量map及protein MSA到codon-aware nucleotide MSA投影 | +| ✅ | Bio.Align.Alignment counts | pairwise/MSA详细gap分类、open/extend事件、identity/mismatch/positive、wildcard、替换矩阵与完整affine总分 | +| ✅ | Bio.Align.psl | alignment-aware PSL/PSLX 21/23列严格解析与写出、核酸和translated 3:1路径、双轴链向、match/repeat/N recount、block序列及坐标互映 | +| ✅ | Bio.Align.sam | alignment-aware SAM header/record严格读写、显式CIGAR path、反向链序列与PHRED语义、typed tags、MD/NM和双向坐标映射 | +| ✅ | Bio.Align.a2m | A2M match/insertion状态推导、大小写与点gap规范读写、插入槽、逐行坐标映射、pair counts、共识和match projection | +| ✅ | Bio.Align.emboss | EMBOSS srspair/pair/simple文件元数据与多alignment解析、任意序列数、纯gap block、局部/反向坐标、consensus统计、compact path和规范写回 | +| ✅ | Bio.Align.exonerate | Exonerate cigar/vulgar文件元数据与严格读写、完整M/5/I/3/C/G/N/S/F操作、正反链/protein strand、3:1 translated path、坐标互映和统计 | +| ✅ | Bio.Align.msf | GCG/PileUp MSF蛋白质与核酸MSA、Check/CompCheck、interleaved blocks、标准checksum、短行补齐、坐标映射、统计与规范写回 | +| ✅ | Bio.Align.nexus | NEXUS DATA/CHARACTERS与TAXA block、nested comments、quoted/duplicate taxa、sequential/interleaved MATRIX、datatype校验、MATCHCHAR、坐标映射、统计与规范写回 | +| ✅ | Bio.Align.stockholm | Stockholm多记录、GF/GS/GR/GC标准与自定义注释、reference/database/nested-domain、M/D/I列操作、all-gap压缩、坐标映射、统计与规范写回 | +| ✅ | Bio.Align.chain | UCSC Chain严格多记录读写、float score与可选ID、正反target/query绝对坐标路径、size/dt/dq重建、双向位置/区间映射、反转和查询 | +| ✅ | Bio.Align.maf | MAF track/header与a/s/i/e/q严格读写、plus/minus绝对路径、任意component映射、MafIndex半开查询、多外显子拼接与canonical writer | +| ✅ | Bio.Align.bed | BED3-BED12严格解析与写出、numeric/text score、正负链query坐标、exon blocks、双向residue mapping、search与summary | +| ✅ | Bio.Align.bigbed | BigBed v4二进制读写、BED3-BED12、AutoSQL扩展字段、多级chromosome B+ tree与R-tree、zlib/DEFLATE解码、区间/名称查询和BED导出 | +| ✅ | Bio.Align.bigmaf | 标准bedMaf bed3+1读写、完整MAF a/s/i/e/q语义、正负链坐标映射、BigBed压缩索引查询、MAF导出和严格损坏数据诊断 | +| ✅ | Bioconductor decontX | 单细胞ambient RNA去污染: cluster-native/contaminant多项式混合、Beta/Dirichlet先验EM、empty-droplet background、自动聚类、计数分解、诊断与SingleCellExperiment接入 | +| ✅ | Bioconductor celda | `celda_CG`细胞群与基因模块联合聚类: 分层Dirichlet-multinomial、collapsed likelihood、EM/Gibbs、多链、K/L模型选择、预测与SingleCellExperiment接入 | | ✅ | Bio.PDB.SASA | 溶剂可及表面积计算: Shrake-Rupley滚动球算法(Fibonacci球面采样)、范德华半径查表、逐原子/残基/链SASA、骨架/侧链拆分 | | ✅ | Bio.SeqIO.NibIO | nib 2-bit二进制序列格式: DNA 2-bit编码(T=0/C=1/A=2/G=3)、4碱基/字节打包、hex I/O、子序列提取、反向互补、GC含量、压缩比 | | ✅ | ChIPseeker | ChIP-seq峰注释: 峰-TSS距离计算、基因组特征分配(Promoter/5'UTR/3'UTR/Exon/Intron/Downstream/Distal Intergenic)、最近基因查找、注释摘要 | @@ -194,12 +240,49 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **群体遗传学** | Biopython `Bio.PopGen` | 等位基因频率、FST、哈迪-温伯格检验 | ✅ | | **edgeR** | Bioconductor edgeR | 差异表达分析、DGEList、精确检验、GLM拟合 | ✅ | | **limma** | Bioconductor limma | 差异表达分析、线性模型拟合、经验贝叶斯、voom变换、RPKM/CPM/quantile归一化、ComBat/removeBatchEffect批次校正、treat严格检验 | ✅ | +| **variancePartition** | Bioconductor variancePartition | 重复测量线性混合模型、方差分解、BLUP、precision weights、dream contrast与Satterthwaite检验 | ✅ | +| **dreamlet** | Bioconductor dreamlet | sample×cell-type pseudobulk、TMM、cell/sample/gene过滤、logCPM、Poisson/voom precision weights、分cell-type重复测量模型及study-wide FDR | ✅ | +| **nnSVG** | Bioconductor nnSVG | nearest-neighbor Gaussian process、空间变异基因检验、gene-specific length scale、协变量设计、空间方差占比、BH-FDR及SpatialExperiment接入 | ✅ | +| **Banksy** | Bioconductor Banksy | H0邻域均值与H1+方位harmonic、六类空间核、lambda联合特征、分组标准化、PCA、多起点k-means、标签平滑、参数扫描及SpatialExperiment接入 | ✅ | +| **Voyager** | Bioconductor Voyager | kNN/distance-band/inverse-distance空间权重(W/B/C/S编码)、全局Moran's I与Geary's c(Cliff-Ord随机化期望/方差/正态p)、局部Moran's I(LISA象限+置换推断)、局部Geary's c、Getis-Ord Gi/Gi*(Ord-Getis z)、Lee's L(全局+局部)、多元局部Geary、经验变差函数与spherical/exponential/gaussian拟合、Moran correlogram、确定性splitmix64置换、BH-FDR及SpatialExperiment不可变写回 | ✅ | +| **spicyR** | Bioconductor spicyR | 有序细胞类型对cross-L曲线、矩形窗口边界校正、图像级共定位统计、precision weights、重复受试者随机截距、条件对比、BH-FDR及SpatialExperiment接入 | ✅ | +| **lisaClust** | Bioconductor lisaClust | 每细胞多类型local-K/centered local-L曲线、Gaussian KDE强度校正、矩形/凸包窗口、圆盘边界修正、确定性多起点k-means、silhouette、区域富集及SpatialExperiment写回 | ✅ | +| **SpatialDecon** | Bioconductor SpatialDecon | 背景感知加权log-normal非负回归、两阶段异常点重拟合、Hessian不确定度、细胞丰度/比例/计数尺度、cell-type collapse、reverse deconvolution、负探针背景、单细胞profile构建及SpatialExperiment写回 | ✅ | +| **ALDEx2** | Bioconductor ALDEx2 | Dirichlet Monte Carlo组成型差异丰度、六类denominator、Welch/Wilcoxon与配对检验、effect/overlap、Aitchison距离及SummarizedExperiment接入 | ✅ | +| **DirichletMultinomial** | Bioconductor DirichletMultinomial | Dirichlet-multinomial概率、有限混合EM/BFGS聚类、Laplace/AIC/BIC选K、生成式分组分类、分层交叉验证、ROC及SummarizedExperiment接入 | ✅ | | **SummarizedExperiment** | Bioconductor SummarizedExperiment | 多维基因组数据容器、Assays、行/列操作 | ✅ | +| **RangedSummarizedExperiment** | Bioconductor SummarizedExperiment | GRanges/GRangesList行范围、复合特征精确重叠/最近邻、覆盖度、区间变换与协调子集 | ✅ | +| **TreeSummarizedExperiment** | Bioconductor TreeSummarizedExperiment | 行/列树、节点链接、树节点子集、祖先/后代查询、层级聚合 | ✅ | | **IRanges** | Bioconductor IRanges | 整数区间操作、集合运算、重叠检测、findOverlaps高级类型、nearest、coverage、距离矩阵计算 | ✅ | | **TxDb** | Bioconductor GenomicFeatures | 转录本数据库、GTF解析、基因/转录本/外显子/CDS提取、UTR/内含子计算、启动子提取 | ✅ | | **ExPASy** | Biopython `Bio.ExPASy` | 蛋白质分析工具接口、Swiss-Prot条目解析、酶数据库查询、蛋白质参数计算(分子量、等电点、GRAVY、不稳定指数) | ✅ | +| **Cellosaurus** | Biopython `Bio.ExPASy.cellosaurus` | Cellosaurus平面文本解析、类型化细胞系记录、数据库交叉引用、物种查询、序列化往返 | ✅ | +| **UniGene** | Biopython `Bio.UniGene` | NCBI UniGene固定宽度记录解析、类型化序列/蛋白相似性/STS/转录本映射、严格SCOUNT校验、序列化往返 | ✅ | +| **HH-suite HHR** | Biopython `Bio.Align.hhr` | HHR元数据、命中摘要、多块profile比对、consensus/二级结构/DSSP/confidence注释、概率和E-value查询 | ✅ | +| **共享参考比对合并** | Biopython `Bio.Align.Alignment.from_alignments_with_same_reference` | 混合PWA/MSA输入、首端/内部/末端insertion同步、多query投影、局部坐标与metadata保留、统计和格式转换 | ✅ | +| **Alignment坐标组合** | Biopython `Bio.Align.Alignment.map/mapall` | alignment path组合、局部overhang clipping、exon/intron与indel gap、正反链组合、坐标双向查询、PSL及1:1/1:3 MSA投影 | ✅ | +| **Alignment详细计数与评分** | Biopython `Bio.Align.Alignment.counts` | 十二类affine gap事件、identity/mismatch/positive、wildcard、BLOSUM/PAM评分、反向链和MSA全部序列对汇总 | ✅ | +| **SAM坐标比对读写** | Biopython `Bio.Align.sam` | header/reference模型、M/I/D/N/=/X显式路径、反向链与clipping、typed tags、PHRED、MD/NM和严格往返 | ✅ | +| **A2M状态感知比对** | Biopython `Bio.Align.a2m` | D/I列状态、大小写与点gap编码、wrapped/CRLF读写、坐标互映、插入槽、统计、共识与match投影 | ✅ | +| **EMBOSS alignment报告** | Biopython `Bio.Align.emboss` | srspair/pair/simple元数据、固定列block、多alignment/多序列、正反向绝对坐标、pair counts与canonical writer | ✅ | +| **Exonerate alignment报告** | Biopython `Bio.Align.exonerate` | header/footer与cigar/vulgar、operation normalization、链感知绝对路径、protein-DNA 3:1映射、统计与canonical writer | ✅ | +| **Exonerate C4 SearchIO报告** | Biopython `Bio.SearchIO.ExonerateIO.exonerate_text` | metadata与查询层次、固定列alignment body、剪接/NER/frameshift语义、phase/frame及双轴区间 | ✅ | +| **GCG MSF alignment** | Biopython `Bio.Align.msf` | AA/NA/PileUp header、Name metadata、interleaved blocks、GCG checksum、三类gap、coordinate path、pair counts与writer | ✅ | +| **NEXUS alignment** | Biopython `Bio.Align.nexus` | DATA/CHARACTERS/TAXA block、quote/comment词法、sequential/interleaved rows、DNA/RNA/protein/standard校验、MATCHCHAR、坐标与writer | ✅ | +| **Stockholm alignment** | Biopython `Bio.Align.stockholm` | 严格header/terminator与多记录、GF/GS/GR/GC映射、M/D/I操作、all-gap列压缩、reference、坐标、统计与writer | ✅ | +| **UCSC Chain alignment** | Biopython `Bio.Align.chain` | 12/13列header、连续记录、float score、双轴链向、绝对half-open路径、size/dt/dq、坐标查询、反转与规范写回 | ✅ | +| **MAF alignment/index** | Biopython `Bio.Align.maf` | track/header与a/s/i/e/q、正负链绝对坐标、任意component映射、MafIndex半开查询、多外显子拼接与规范写回 | ✅ | +| **BED pairwise alignment** | Biopython `Bio.Align.bed` | BED3-BED12、target/query双轴路径、反链转录本坐标、block/counts、双向位置映射、区间搜索与分级写回 | ✅ | +| **BigBed二进制区间索引** | Biopython `Bio.Align.bigbed` | BigBed v4读写、BED3-BED12与AutoSQL、多级B+ tree/R-tree、stored/fixed/dynamic DEFLATE、区间/名称查询、链感知exon坐标及BED导出 | ✅ | +| **BigMaf多物种比对索引** | Biopython `Bio.Align.bigmaf` | 标准bedMaf AutoSQL、MAF a/s/i/e/q块、score/pass/comment、正负链坐标映射、压缩BigBed索引查询及普通MAF导出 | ✅ | +| **BigPsl成对比对索引** | Biopython `Bio.Align.bigpsl` | 标准bed12+13 AutoSQL、核酸与translated protein坐标路径、正反链、match/repeat/N recount、压缩索引查询及PSL导出 | ✅ | +| **zinbwave** | Bioconductor zinbwave | 零膨胀负二项低维模型、cell/gene协变量与offset、确定性latent factors、dispersion shrinkage、observational weights、残差/归一化/插补及SCE集成 | ✅ | +| **apeglm** | Bioconductor apeglm | 负二项GLM、自适应经验贝叶斯Cauchy/Student-t先验、确定性多起点MAP、Laplace后验SD/区间、FSR/FSOS/s-value、DESeq2与SummarizedExperiment接入 | ✅ | | **Prosite** | Biopython `Bio.Prosite` | 蛋白质模体数据库搜索、Prosite模式解析、模体匹配算法、模体得分计算 | ✅ | -| **PAML** | Biopython `Bio.PAML` | 分子进化分析、dN/dS计算(Nei-Gojobori方法)、Jukes-Cantor校正、密码子使用分析 | ✅ | +| **PAML内置近似兼容层** | BioSeqs compatibility API | Nei-Gojobori风格dN/dS、Jukes-Cantor校正和密码子使用分析;不解析PAML控制文件或原生输出 | ✅ | +| **CODEML控制与结果分析** | Biopython `Bio.Phylo.PAML.codeml` | 严格control读写、CODONML/AAML、多NSsites/branch-site/clade/free-ratio、pairwise、距离矩阵、多基因、BEB/NEB及AIC/BIC/LRT | ✅ | +| **BASEML核苷酸模型分析** | Biopython `Bio.Phylo.PAML.baseml` | 严格control读写、JC69至UNRESTu、参数/SE、kappa、REV/UNREST Q矩阵、rate-class、auto-dGamma、nhomo节点及AIC/BIC/LRT | ✅ | +| **YN00成对密码子替换分析** | Biopython `Bio.Phylo.PAML.yn00` | 严格control读写、NG86/YN00/LWL85/LWL85m/LPB93、PAML 4.1-4.9i兼容、非有限值、对称矩阵及统计汇总 | ✅ | | **Graphics** | Biopython `Bio.Graphics` | 生物信息学可视化、序列Logo绘制、序列比对可视化、基因组特征绘图 | ✅ | | **BSgenome** | Bioconductor BSgenome | 基因组序列数据库、染色体序列检索、子序列提取、链特异性基因提取 | ✅ | | **biomaRt** | Bioconductor biomaRt | 基因ID映射、基因注释查询、批量查询、外部数据库映射 | ✅ | @@ -227,6 +310,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **Affy** | Biopython `Bio.Affy` | Affymetrix芯片数据分析、RMA标准化、背景校正、分位数归一化 | ✅ | | **SVDSuperimposer** | Biopython `Bio.PDB.SVDSuperimposer` | SVD蛋白质结构叠合、旋转矩阵、平移向量、RMSD计算 | ✅ | | **QCPSuperimposer** | Biopython `Bio.PDB.QCPSuperimposer` | 四元数特征多项式结构叠合、高精度旋转矩阵、平移向量、RMSD计算 | ✅ | +| **CEAligner** | Biopython `Bio.PDB.cealign` | 组合扩展结构比对、AFP单调路径、CE Z-score、局部索引优化、QCP最优叠合 | ✅ | | **ResidueDepth** | Biopython `Bio.PDB.ResidueDepth` | 残基深度计算、溶剂可及表面积(SASA)、表面/核心残基识别 | ✅ | | **StructureAlignment** | Biopython `Bio.PDB.StructureAlignment` | 多蛋白质结构比对、动态规划比对、RMSD/TM-score计算、渐进式多结构比对 | ✅ | | **KEGG** | Biopython `Bio.KEGG` | KEGG基因/通路/化合物/酶记录解析、通路分析 | ✅ | @@ -240,6 +324,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **UniProtIO** | Biopython `Bio.SeqIO.UniprotIO` | UniProt XML格式解析、蛋白质条目提取、基因名、物种、序列、功能注释、数据库交叉引用 | ✅ | | **chem_utils** | Biopython `Bio.PDB.chem_utils` | 化学计算工具:范德华半径、共价半径、键长、键角、二面角、经验式、分子式量、氢键长度 | ✅ | | **mmCIF** | Biopython `Bio.PDB.MMCIFParser` | mmCIF格式解析、数据块、类别、原子位点提取 | ✅ | +| **BinaryCIF** | Biopython `Bio.PDB.binary_cif` | 纯MoonBit MessagePack解析、ByteArray/FixedPoint/IntervalQuantization/RunLength/Delta/IntegerPacking/StringArray逆编码、mask与Structure转换 | ✅ | | **Nexus** | Biopython `Bio.Nexus` | NEXUS格式解析、数据矩阵、系统发育树、距离矩阵 | ✅ | | **EMBOSS** | EMBOSS suite | GC偏斜、AT偏斜、分子量、Tm值、ORF查找、距离计算、蛋白质参数 | ✅ | | **ChIPseeker** | Bioconductor ChIPseeker | ChIP-seq峰注释、基因距离计算、注释分类(启动子/外显子/内含子/UTR/基因间区)、BED格式读取、peak2gene关联分析、多峰值集重叠分析(peakOverlap)、Venn图可视化、饼图可视化、结果汇总与可视化 | ✅ | @@ -260,16 +345,23 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **microbiome** | Bioconductor microbiome | 微生物组分析、Alpha多样性(Shannon/Simpson/Chao1/ACE/Fisher/Pielou)、Beta多样性(Bray-Curtis/Jaccard/JSD/weighted/unweighted UniFrac)、PCoA主坐标分析、差异丰度分析(Welch t检验/Wilcoxon秩和检验/BH校正) | ✅ | | **BiocParallel** | Bioconductor BiocParallel | 并行计算框架、任务分块、并行求和、均值计算、进度追踪 | ✅ | | **ensembldb** | Bioconductor ensembldb | Ensembl注释数据库接口、基因/转录本/外显子/CDS检索、染色体过滤、biotype过滤、基因长度计算 | ✅ | -| **DropletUtils** | Bioconductor DropletUtils | 空液滴检测、barcode排序、knee点检测、emptyDrops算法、细胞过滤 | ✅ | +| **DropletUtils** | Bioconductor DropletUtils 1.33.0 | 严格feature × barcode计数模型、Simple Good-Turing ambient profile、曲线追踪knee/inflection、multinomial/Dirichlet-multinomial概率、alpha估计、确定性Monte Carlo、Phipson–Smyth p值、BH-FDR、高计数保留及SingleCellExperiment写回 | ✅ | | **rhdf5** | Bioconductor rhdf5 | HDF5文件格式支持、数据集读写、组管理、属性操作、文件列表查看 | ✅ | | **Matrix** | Bioconductor Matrix | 稀疏矩阵操作、CSC/CSR格式、矩阵运算(加法、乘法、转置)、行列统计、范数计算 | ✅ | | **BiocGenerics** | Bioconductor BiocGenerics | Bioconductor通用函数、NA处理、排序、集合运算、匹配、表统计、序列生成 | ✅ | | **scran** | Bioconductor scran | 单细胞归一化(sum_factors)、SNN图构建、Leiden聚类、差异标志物分析 | ✅ | +| **scrapper** | Bioconductor scrapper | 批次感知RNA QC、大小因子清洗与居中、log-normalization、LOWESS方差趋势、HVG选择、多因子pseudo-bulk、不可变SCE集成 | ✅ | +| **scuttle** | Bioconductor scuttle 1.23.1 | batch-aware median/MAD异常值、subset per-feature QC、重叠feature-set聚合、精确无放回count downsampling、batch/block coverage等化及不可变SCE集成 | ✅ | +| **bluster** | Bioconductor bluster 1.23.0 | 多起点K-means、精确KNN与三类加权SNN图、Louvain风格聚类、two-step量化聚类、Rand/ARI、silhouette、purity、RMSD、modularity、bootstrap稳定性及不可变SCE集成 | ✅ | +| **FlowSOM** | Bioconductor FlowSOM 2.21.0 | 规则网格SOM、KWSP/random/PCA初始化、四类距离、分阶段MST邻域、meta-clustering与自动elbow、节点MFI/CV/SD/MAD/阳性率、MAD outlier、新数据映射及FlowFrame/SCE集成 | ✅ | +| **decontX** | Bioconductor decontX | cluster-aware ambient RNA混合模型、每细胞污染率、Beta/Dirichlet先验EM、empty-droplet profile、自动k-means、native/contaminant计数分解与SCE集成 | ✅ | +| **celda** | Bioconductor celda | `celda_CG`细胞群/基因模块联合聚类、collapsed likelihood、EM/Gibbs、多链、K/L网格选择、新细胞预测与SCE集成 | ✅ | +| **miloR** | Bioconductor miloR | 精确KNN图、精炼重叠邻域、邻域×样本计数、NB-GLM/Wald检验、BH与四种graph spatial FDR、SingleCellExperiment接入 | ✅ | | **monocle3** | Bioconductor monocle3 | 单细胞轨迹分析、PCA/UMAP降维、主图学习、拟时间排序、差异表达分析、分支点检测、分支特异性差异表达 | ✅ | | **ShortRead** | Bioconductor ShortRead | 短读序列质量控制、QA统计、adapter修剪、质量修剪、读长过滤、FastQC报告生成 | ✅ | | **scater** | Bioconductor scater | 单细胞质量控制、QC指标计算、细胞/基因过滤、CPM/log-CPM标准化、HVG检测、PCA降维 | ✅ | -| **MAST** | Bioconductor MAST | 单细胞差异表达分析、Hurdle模型、离散/连续组分检验、BH-FDR校正、结果汇总 | ✅ | -| **SingleR** | Bioconductor SingleR | 细胞类型注释、Spearman/Pearson相关性、参考图谱匹配、精细调优(Fine-tuning)、Delta score置信度评估 | ✅ | +| **MAST** | Bioconductor MAST 1.39.0 | 严格feature × cell数据模型、任意设计矩阵与CDR、Bayesian logistic/Gaussian hurdle GLM、嵌套模型LRT、经验贝叶斯方差收缩、BH-FDR、边际logFC及SingleCellExperiment写回 | ✅ | +| **SingleR** | Bioconductor SingleR 2.15.2 | 严格gene × sample/cell模型、基因名对齐、成对classic markers、ties-aware Spearman、标签内相关分位数、迭代fine-tuning、delta/MAD剪枝、cluster注释、多参考重算及SingleCellExperiment写回 | ✅ | | **Cyclone** | Bioconductor cyclone | 细胞周期评分、基因对(Gene pairs)比较、G1/S/G2/M期相预测、相别分布统计、平均得分分析 | ✅ | | **dorothea** | Bioconductor dorothea | 转录因子活性预测、Regulon分析、VIPER算法、置换检验、Z-score评估、Top TF筛选 | ✅ | | **GenomicFiles** | Bioconductor GenomicFiles | 分布式基因组文件处理、按区间扫描BAM/BED/VCF、批量查询、归约、覆盖度计算 | ✅ | @@ -282,8 +374,9 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **GSVA** | Bioconductor GSVA | 基因集变异分析、单样本通路评分(ssGSEA/zscore/PLAGE)、富集分析、置换检验、富集图可视化(enrichment map)、表型相关性分析(phenotype correlation)、生存分析(survival analysis)、分数分布分析与可视化 | ✅ | | **ChromVAR** | Bioconductor chromVAR | 染色质变异分析、TF motif富集、GC偏差校正、细胞聚类、变异性分析 | ✅ | | **DelayedArray** | Bioconductor DelayedArray | 延迟计算数组、懒加载操作、分块处理、行/列聚合、子集操作 | ✅ | +| **SparseArray** | Bioconductor SparseArray | N维规范化COO稀疏数组、坐标/线性索引、稀疏切片、维度置换与绑定、不可变赋值、算术/统计、矩阵乘法 | ✅ | | **AnnotationFilter** | Bioconductor AnnotationFilter | 基因注释过滤、染色体筛选、生物类型过滤、区域重叠检测、符号模式匹配 | ✅ | -| **scDblFinder** | Bioconductor scDblFinder | 单细胞双细胞检测、Doublet评分计算、最近邻搜索、PCA降维、细胞过滤 | ✅ | +| **scDblFinder** | Bioconductor scDblFinder 1.27.6 | 严格cell × gene数据模型、top-variable特征选择、library normalization/PCA、随机或跨cluster人工doublet、精确kNN比例与来源推断、cxds互斥共表达分数、迭代正则化logistic分类、预期doublet rate阈值优化、capture分层、homotypic修正、pairwise来源富集及SingleCellExperiment写回 | ✅ | | **Batchelor** | Bioconductor batchelor | 单细胞批次校正、rescaleBatches缩放校正、mutual nearest neighbor、fastMNN多批次校正、批次混合评分 | ✅ | | **Seurat** | Bioconductor Seurat | 单细胞数据分析核心、LogNormalize标准化、高可变基因检测、PCA降维、图聚类、UMAP可视化、差异标志物分析、跨样本整合(FindIntegrationAnchors/IntegrateData) | ✅ | | **ChIPseeker** | Bioconductor ChIPseeker | ChIP-seq峰值注释、基因组区域分类(启动子/外显子/内含子/UTR/基因间区)、距离TSS分布、BED格式读取、peak2gene关联分析、注释可视化、统计分析 | ✅ | @@ -312,7 +405,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **bamsignals** | Bioconductor bamsignals | ChIP-seq信号提取(计数模式、RPM/RPKM归一化、基因组区域信号分析、染色质状态分析) | ✅ | | **nucleR** | Bioconductor nucleR | 核小体定位分析(信号平滑、峰值检测、核小体occupancy计算、位置比较、动态变化分析) | ✅ | | **csaw** | Bioconductor csaw | ChIP-seq窗口差异分析(滑动窗口计数、TMM归一化、窗口过滤、负二项GLM检验、差异区域检测) | ✅ | -| **slingshot** | Bioconductor slingshot | 单细胞轨迹推断(MST构建、主曲线拟合、拟时间计算、分支检测) | ✅ | +| **slingshot** | Bioconductor slingshot 2.21.0 | 保留基础MST/主曲线兼容API,并新增hard/soft cluster membership、协方差缩放距离、start/end约束、omega forest、root-to-leaf lineage、同时主曲线、rank重加权/重分配、Optional pseudotime、分支ID、新数据映射及不可变SCE写回 | ✅ | | **SCnorm** | Bioconductor SCnorm | 单细胞RNA-seq归一化(分位数回归、深度依赖偏差校正、基因特异性归一化) | ✅ | | **EDASeq** | Bioconductor EDASeq | RNA-seq探索性分析(GC含量归一化、基因长度校正(Loess)、样本间归一化、RPKM计算) | ✅ | | **Bio.phenotype** | Biopython `Bio.phenotype` | 表型微阵列分析(PlateRecord/WellRecord、logistic/Gompertz生长曲线拟合、CSV/JSON解析、控制减法) | ✅ | @@ -324,19 +417,41 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **destiny** | Bioconductor destiny | 单细胞扩散映射(Diffusion Maps)降维、距离矩阵计算、高斯核构建、特征分解、扩散分量计算 | ✅ | | **Rtsne** | Bioconductor Rtsne | t-SNE降维算法、成对距离计算、条件概率估计(perplexity优化)、联合概率矩阵构建、梯度下降优化(动量/early exaggeration)、Barnes-Hut近似 | ✅ | | **uwot** | Bioconductor uwot | UMAP降维算法、k近邻搜索、模糊单纯集构建、局部模糊集并集、低维嵌入优化(SGD/负采样)、min_dist/spread参数控制 | ✅ | -| **tradeSeq** | Bioconductor tradeSeq | 轨迹差异表达分析、GAM(广义可加模型)拟合、样条基函数、差异表达检验、BH-FDR校正 | ✅ | +| **tradeSeq** | Bioconductor tradeSeq 1.27.0 | 保留基础轨迹差异表达兼容API,并新增gene × cell计数合同、cell × lineage拟时间/权重、library offset、多lineage负二项GAM、惩罚B-spline、dispersion/AIC/协方差、五类Wald检验、log2FC阈值、平滑预测、knot评估、Slingshot直连及不可变SCE写回 | ✅ | +| **muscat advanced** | Bioconductor muscat 1.27.4 | cluster × sample伪批量、sum/mean/median/prop.detected/num.detected、任意设计与多contrast、NB-IRLS、dispersion收缩、CDR归一化DD、local/global FDR、DS/DD stagewise及SingleCellExperiment接入 | ✅ | | **PROGENy** | Bioconductor PROGENy | 通路活性推断、L2正则化线性回归(Ridge回归)、通路基因集权重矩阵、样本通路活性计算 | ✅ | | **AUCell** | Bioconductor AUCell | 单细胞基因集评分、AUC(曲线下面积)计算、基因排序、min-max归一化、细胞/基因集评分查询 | ✅ | | **ggtree** | Bioconductor ggtree | 系统发育树可视化布局算法、矩形布局(phylogram)、放射状布局、无根布局、节点坐标映射、边缘/标签数据生成 | ✅ | | ✅ | BiocSingular | SVD奇异值分解,支持Exact/IRLBA/Randomized三种算法,用于单细胞降维 | | ✅ | BiocNeighbors | KMKNN和Annoy最近邻搜索,支持欧几里得/曼哈顿/余弦距离 | | ✅ | mixOmics | 多组学整合方法,包括PLS回归、稀疏PLS (sPLS)、DIABLO多块整合 | -| **MAF格式解析** | Biopython `Bio.Align` | MAF多序列比对格式解析、块操作、百分比一致性、统计分析、选择/过滤/写回 | ✅ | +| **MAF宽松块分析** | Biopython `Bio.Align` | MAF块解析、百分比一致性、统计分析、选择/过滤/写回 | ✅ | +| **现代MAF alignment API** | Biopython `Bio.Align.maf` | track/header、严格a/s/i/e/q、正负链绝对坐标、component映射、MafIndex区间搜索、多外显子拼接及canonical写回 | ✅ | +| **现代BED alignment API** | Biopython `Bio.Align.bed` | 分级BED3-BED12读写、链感知双轴路径、exon block投影、residue mapping、overlap search与summary | ✅ | +| **HH-suite HHR格式** | Biopython `Bio.Align.hhr` | HHsearch/HHblits结果严格解析、0-based坐标、query-target映射、规范化写回 | ✅ | +| **共享参考比对同步** | Biopython `Bio.Align.Alignment` | 相同参考PWA/MSA合并、边界插入宽度归一化、query原始比对结构保留、reference/query/column坐标互映 | ✅ | +| **Alignment gap/composition统计** | Biopython `Bio.Align.Alignment.counts` | pairwise与MSA逐对统计、端部/内部gap分类、open/extend事件、替换和gap总分 | ✅ | +| **Alignment-aware tabular搜索结果** | Biopython `Bio.Align.tabular` | BLAST outfmt 7、FASTA 8CB/8CC、BTOP/aln_code路径、链向及translated坐标 | ✅ | +| **Alignment-aware PSL/PSLX** | Biopython `Bio.Align.psl` | 21/23列格式、header、block/gap一致性、正反链、translated 3:1坐标、序列重计数与严格诊断 | ✅ | +| **Alignment-aware SAM** | Biopython `Bio.Align.sam` | SAM 1.6 header/record、typed optional tags、CIGAR path、clipping、反向链、PHRED、MD/NM与规范写回 | ✅ | +| **A2M状态感知MSA** | Biopython `Bio.Align.a2m` | match/deletion与insertion列、canonical大小写/点编码、严格往返、坐标映射、插入槽、统计与共识 | ✅ | +| **EMBOSS alignment output** | Biopython `Bio.Align.emboss` | water/needle/stretcher/matcher/alignret输出、srspair/pair/simple、metadata、consensus、局部/反向坐标与严格诊断 | ✅ | +| **Exonerate alignment output** | Biopython `Bio.Align.exonerate` | cigar/vulgar报告、M/5/I/3/C/G/N/S/F路径、正反链与protein strand、translated coordinates、严格诊断与双格式写回 | ✅ | +| **Exonerate C4 text output** | Biopython `Bio.SearchIO.ExonerateIO.exonerate_text` | C4 alignment text、query/hit/HSP/fragment聚合、intron/NER/split codon/frameshift、translated coordinates与严格诊断 | ✅ | +| **GCG MSF alignment format** | Biopython `Bio.Align.msf` | protein/nucleotide MSF、PileUp与EMBOSS变体、interleaved MSA、GCG checksum、坐标和统计、严格读写 | ✅ | +| **NEXUS alignment format** | Biopython `Bio.Align.nexus` | NEXUS header与block、nested comments、quoted taxa、sequential/interleaved MATRIX、MATCHCHAR、all-gap压缩、坐标统计与严格读写 | ✅ | +| **Stockholm alignment format** | Biopython `Bio.Align.stockholm` | 多记录严格解析、GF/GS/GR/GC与自定义注释、reference与database reference、insertion/deletion列、all-gap压缩、坐标统计与严格写回 | ✅ | +| **UCSC Chain alignment format** | Biopython `Bio.Align.chain` | 严格header/block/span校验、连续多记录、正反双轴absolute path、canonical block重建、位置/区间映射、反转与ID/overlap查询 | ✅ | +| **MAF alignment format/index** | Biopython `Bio.Align.maf` | track/##maf metadata、a/s/i/e/q关联校验、链感知absolute path、任意序列坐标映射、reference interval index与spliced alignment | ✅ | +| **BED alignment format** | Biopython `Bio.Align.bed` | BED3-BED12字段层级、numeric/text score、正负链transcript path、block几何校验、双向mapping及canonical projection | ✅ | | **UCSC Chain文件/liftOver** | Bioconductor rtracklayer | Chain格式解析、基因组坐标liftOver转换、链段查找、染色体间坐标映射、位置/区间转换 | ✅ | | **Biostrings matchPDict** | Bioconductor Biostrings | 字典模式匹配(matchPDict/vmatchPattern)、多序列模式计数(vcountPattern)、错配容忍、最佳匹配查找 | ✅ | | **GenomicRanges gaps/reduce/disjoin** | Bioconductor GenomicRanges | gaps检测、reduce合并、disjoin拆分、setdiff/交集/并集集合运算、coverage计算、promoters提取、trim | ✅ | +| **GenomicRanges GRangesList** | Bioconductor GenomicRanges | 命名复合基因组特征、split/unlist/relist、逐组区间变换、并行集合运算、特征级重叠/最近邻/覆盖度 | ✅ | | **NGS质量修剪与接头去除** | Bioconductor ShortRead | 质量修剪(滑动窗口)、接头去除、poly-A修剪、长度/GC含量过滤、批量修剪、Fastq解析与序列化、统计计算 | ✅ | -| **Mauve基因组比对** | Biopython `Bio.Align` | Mauve基因组比对格式、LCB检测、倒位/断点检测、覆盖率分析、BED导出、基因组重排率 | ✅ | +| **Mauve重排分析(legacy)** | BioSeqs compatibility API | 历史 MAF-like 块、LCB检测、倒位/断点检测、覆盖率、BED导出与重排率 | ✅ | +| **XMFA多基因组比对** | Biopython `Bio.Align.mauve` | `#SequenceN*` metadata、LCB、正负链坐标、区间索引、跨序列位置/区间投影、统计、重建与规范写回 | ✅ | +| **BLAST XML1/XML2** | Biopython `Bio.Blast` | 严格XML树解析、多query与XML2多report、HitDescr/taxonomy、参数与统计、八类BLAST坐标语义、canonical writer与显式跨格式转换 | ✅ | | **Stockholm格式** | Biopython `Bio.Stockholm` | Stockholm/Pfam格式解析、二级结构注释、百分比一致性、保守性、FASTA/Stockholm互转 | ✅ | | **高级群体遗传学** | Biopython `Bio.PopGen` | Tajima's D中性检验、Fu & Li's D/F、McDonald-Kreitman检验、等位基因频率谱、综合中性分析 | ✅ | | **高级密码子分析** | Biopython `Bio.SeqUtils.CodonUsage` | CAI密码子适应指数、RSCU相对同义密码子使用、ENC有效密码子数、GC3偏斜、最优/稀有密码子、物种特异性参考表 | ✅ | @@ -344,6 +459,7 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **Storey q-value** | Bioconductor qvalue | π₀估计、Storey q-value计算、自助法π₀、FDR校正、显著性检验 | ✅ | | **独立假设加权(IHW)** | Bioconductor IHW | 协变量加权Bonferroni、局部/全局加权、Storey pi0加权、多协变量支持、迭代权重优化 | ✅ | | **DelayedMatrixStats** | Bioconductor DelayedMatrixStats | DelayedArray统计层、row/col统计(mean/var/sd/median/min/max/sum)、NA处理、子集操作 | ✅ | +| **SparseArray N维稀疏计算** | Bioconductor SparseArray | COO规范化、重复坐标合并与零消除、R列主序、稀疏子集/aperm/bind、Hadamard运算、crossprod/tcrossprod | ✅ | | **GC-RMA芯片分析** | Bioconductor gcrma | GC校正RMA、背景校正(IdealMM/Express)、GC查找表、分位数归一化、探针组汇总 | ✅ | | **ACE contig格式** | Biopython `Bio.Sequencing.Ace` | ACE组装格式解析、reads/contigs提取、共有序列生成、覆盖度分析、GC含量计算、格式化输出 | ✅ | | **蛋白质组学分析** | Biopython `Bio.SeqUtils.Proteomics` | 8种蛋白酶切(胰酶/糜酶/胃酶/LysC/ArgC/CNBr/GluC/AspN)、单同位素/平均质量计算、同位素分布、b/y碎片离子 | ✅ | @@ -365,6 +481,8 @@ BioSeqs 是一个基于 **MoonBit** 语言开发的生物信息学工具库, | **Wise2 DNA-蛋白比对** | Biopython `Bio.Wise` | GeneWise输出解析、外显子/内含子/比对列、剪接位点相位、比特分数、参数提取、蛋白质/DNA序列、基因预测结果 | ✅ | | **stageR 两阶段检验** | Bioconductor stageR | 两阶段假设检验(筛选+确认)、Simes聚合、BH-FDR校正、Holm步降程序、OFDR控制、Dte/Dtu方法、确认p值重缩放 | ✅ | | **EnrichedHeatmap 富集热图** | Bioconductor EnrichedHeatmap | 基因组信号归一化、目标区域窗口化、四种均值模式(absolute/weighted/w0/coverage)、行平滑、百分位裁剪、链方向处理 | ✅ | +| **高级密码子比对与选择压力检验** | Biopython `Bio.codonalign` | Z-test选择检验(Nei-Gojobori近似方差)、Fisher精确检验中性度、密码子比对构建器、滑窗dN/dS、BH-FDR多重校正、成对Ka/Ks表 | ✅ | +| **高级蛋白质序列预测** | Biopython `Bio.SeqUtils` | Chou-Fasman二级结构预测、IUPred无序区预测、COILS卷曲螺旋预测、Kolaskar-Tongaonkar抗原性、Emini表面可及性、Karplus-Schulz柔柔性 | ✅ | 项目致力于打造一个完整、高效的生物信息学工具库,覆盖从基础序列处理到高级序列组装的全流程。 @@ -385,6 +503,27 @@ IvanAXu/BioSeqs/ │ ├── fastq_io.mbt # FASTQ 格式解析 │ ├── genbank_io.mbt # GenBank 格式解析 │ ├── align.mbt # MultipleSeqAlignment 多序列比对 +│ ├── shared_reference_alignment.mbt # Bio.Align共享参考PWA/MSA合并、insertion slot同步与坐标映射 +│ ├── alignment_map.mbt # Bio.Align.Alignment map/mapall坐标路径组合与MSA投影 +│ ├── alignment_counts.mbt # Bio.Align.Alignment.counts详细gap/composition统计与评分 +│ ├── align_tabular.mbt # Bio.Align.tabular BLAST/FASTA traceback表格解析与坐标路径 +│ ├── align_psl.mbt # Bio.Align.psl PSL/PSLX严格读写、链向路径、统计与坐标映射 +│ ├── align_sam.mbt # Bio.Align.sam严格读写、CIGAR路径、typed tags、MD/NM与反向链 +│ ├── a2m.mbt # Bio.Align.a2m状态感知MSA读写、坐标、插入槽、统计与投影 +│ ├── align_emboss.mbt # Bio.Align.emboss srspair/pair/simple解析、坐标、统计与规范写回 +│ ├── align_exonerate.mbt # Bio.Align.exonerate cigar/vulgar路径、链向、translated映射与规范写回 +│ ├── msf.mbt # Bio.Align.msf GCG/PileUp MSA、checksum、坐标统计与规范写回 +│ ├── align_nexus.mbt # Bio.Align.nexus DATA/CHARACTERS、interleave、MATCHCHAR、坐标统计与写回 +│ ├── align_stockholm.mbt # Bio.Align.stockholm GF/GS/GR/GC、多记录、列操作、坐标统计与写回 +│ ├── align_chain.mbt # Bio.Align.chain严格读写、双轴链向绝对路径、block重建与坐标映射 +│ ├── align_maf.mbt # Bio.Align.maf严格文档读写、绝对坐标、MafIndex查询与多外显子拼接 +│ ├── align_mauve.mbt # Bio.Align.mauve现代XMFA文档、LCB、双向坐标、索引、投影与规范写回 +│ ├── align_clustal.mbt # Bio.Align.clustal现代MSA metadata、consensus、坐标、统计与规范写回 +│ ├── align_phylip.mbt # Bio.Align.phylip现代sequential/interleaved MSA、坐标、统计与规范写回 +│ ├── align_bed.mbt # Bio.Align.bed BED3-BED12、双轴路径、block投影、坐标映射与分级写回 +│ ├── bigbed.mbt # Bio.Align.bigbed v4、BED/AutoSQL、多级B+ tree/R-tree与DEFLATE +│ ├── bigmaf.mbt # Bio.Align.bigmaf bedMaf、MAF注释、链向坐标映射与索引查询 +│ ├── bigpsl.mbt # Bio.Align.bigpsl bed12+13、核酸/translated protein路径与PSL │ ├── alignio.mbt # 比对文件 I/O │ ├── clustal_io.mbt # Clustal 格式 │ ├── phylip_io.mbt # PHYLIP 格式 @@ -397,6 +536,7 @@ IvanAXu/BioSeqs/ │ ├── phylo.mbt # 系统发育树 (Clade/Tree) │ ├── tree_io.mbt # 进化树格式解析 (Newick、NHX格式解析、树操作) │ ├── blast.mbt # BLAST结果解析 (tabular/xml格式、HSP、Hit、Record、过滤) +│ ├── blast_xml_advanced.mbt # Bio.Blast XML1/XML2严格解析、坐标路径、查询分析与规范写回 │ ├── searchio.mbt # SearchIO 统一搜索结果模型 (HSPFragment、HSP、Hit、QueryResult、HMMER3/BLAT解析) │ ├── search_io.mbt # SearchIO 统一搜索结果模型 (HMMER3解析、BLAT PSL解析、BLAST转换) │ ├── subsmat.mbt # 替换矩阵 (BLOSUM62/45、PAM250/30、矩阵解析、分数查询) @@ -430,6 +570,7 @@ IvanAXu/BioSeqs/ │ ├── bioc_parallel.mbt # Bioconductor 并行计算框架 (任务分块、并行求和) │ ├── genomic_ranges.mbt # GenomicRanges 基因组区间操作 (GRanges、IRanges) │ ├── genomic_ranges_advanced.mbt # GenomicRanges tile/slidingWindows/区间运算 +│ ├── granges_list.mbt # GenomicRanges GRangesList 复合特征、重组、重叠与集合运算 │ ├── iranges.mbt # IRanges 整数区间操作 (集合运算、重叠检测) │ ├── genomic_alignments.mbt # GenomicAlignments 基因组比对分析 (GAlignments、coverage、summarizeOverlaps、pileup) │ ├── txdb.mbt # TxDb 转录本数据库 (GTF解析、基因/转录本/外显子/CDS提取、UTR/内含子计算) @@ -438,12 +579,18 @@ IvanAXu/BioSeqs/ │ ├── rhdf5.mbt # Bioconductor rhdf5 HDF5文件格式支持 │ ├── deseq2.mbt # DESeq2 差异表达分析 (size factors归一化、分散度估计、负二项GLM拟合、Wald检验、LFC收缩) │ ├── deseq2_advanced.mbt # DESeq2 VST方差稳定化变换、PCA可视化 +│ ├── apeglm.mbt # apeglm 自适应重尾LFC收缩、Laplace后验、FSR/FSOS与容器接入 +│ ├── aldex2.mbt # ALDEx2 Dirichlet Monte Carlo组成型推断、检验、effect与容器接入 +│ ├── dirichlet_multinomial.mbt # DirichletMultinomial有限混合EM、模型选择、分类、CV与ROC │ ├── edger.mbt # edgeR 差异表达分析 (DGEList、精确检验、GLM拟合) │ ├── edger_advanced.mbt # edgeR准似然F检验、camera/roast基因集检验 │ ├── limma.mbt # limma 差异表达、归一化、批次校正 (线性模型、经验贝叶斯、voom、RPKM/CPM/quantile、ComBat) │ ├── matrix.mbt # Bioconductor Matrix 稀疏矩阵操作 (CSC/CSR格式、矩阵运算) +│ ├── sparse_array.mbt # Bioconductor SparseArray N维规范化COO、切片/置换/绑定、稀疏算术与矩阵乘法 │ ├── bioc_neighbors.mbt # BiocNeighbors 最近邻搜索 (KMKNN/Annoy) │ ├── summarized_experiment.mbt # SummarizedExperiment 多维基因组数据容器 +│ ├── ranged_summarized_experiment.mbt # RangedSummarizedExperiment GRanges/GRangesList行范围与协调操作 +│ ├── tree_summarized_experiment.mbt # TreeSummarizedExperiment 树结构实验容器、节点子集与层级聚合 │ ├── dplyr.mbt # dplyr 数据操作 (DataFrame、filter、select、mutate、arrange、group_by、summarize、join) │ ├── plyranges.mbt # plyranges tidy基因组数据操作 (GRanges的filter/mutate/select/arrange/rename/summarise/join) │ ├── smith_waterman.mbt # Smith-Waterman 局部序列比对 (动态规划、自定义打分、回溯矩阵) @@ -453,7 +600,10 @@ IvanAXu/BioSeqs/ │ ├── de_bruijn.mbt # De Bruijn Graph (k-mer节点、欧拉路径、序列组装、图简化) │ ├── suffix_array_tree.mbt # Suffix Array & Suffix Tree (前缀倍增、LCP数组、模式匹配、最长重复子串) │ ├── olc.mbt # Overlap-Layout-Consensus (重叠检测、哈密顿路径、一致性序列生成) -│ ├── paml.mbt # Bio.PAML 分子进化分析 (dN/dS、Jukes-Cantor校正) +│ ├── paml.mbt # 历史内置近似API (Nei-Gojobori风格dN/dS、Jukes-Cantor校正) +│ ├── paml_codeml.mbt # Bio.Phylo.PAML.codeml control/CODONML/AAML解析与模型比较 +│ ├── paml_baseml.mbt # Bio.Phylo.PAML.baseml control/核苷酸模型结果与统计比较 +│ ├── paml_yn00.mbt # Bio.Phylo.PAML.yn00 control与五种成对密码子替换估计 │ ├── hmm.mbt # Hidden Markov Model (前向/后向算法、维特比算法、Baum-Welch训练、基因预测) │ ├── kmeans.mbt # K-means Clustering (距离计算、K-means++初始化、聚类、轮廓系数评估) │ ├── kmer.mbt # Bio.Kmer k-mer计数与频率分析 @@ -465,6 +615,8 @@ IvanAXu/BioSeqs/ │ ├── restriction.mbt # 限制性内切酶分析 (酶切位点查找、片段分析) │ ├── protparam.mbt # ProtParam 蛋白质参数分析 (不稳定指数、等电点、信号肽预测、二级结构倾向) │ ├── prosite.mbt # Bio.Prosite 蛋白质模体数据库搜索 +│ ├── cellosaurus.mbt # Bio.ExPASy.cellosaurus 细胞系数据库平面文件解析 +│ ├── unigene.mbt # Bio.UniGene NCBI UniGene固定宽度记录解析、查询与序列化 │ ├── affy.mbt # Affy Affymetrix芯片数据分析 (RMA标准化、背景校正、分位数归一化) │ ├── feature_extraction.mbt # 机器学习特征提取 │ ├── faidx.mbt # FASTA 快速索引访问 (pyfaidx) @@ -480,6 +632,7 @@ IvanAXu/BioSeqs/ │ ├── ballgown.mbt # ballgown 转录组水平差异表达分析 (FPKM计算、t检验、基因/转录本结构) │ ├── align_info.mbt # AlignInfo 比对统计 (一致性序列、保守位点、Shannon熵、成对序列同一性) │ ├── codon_align.mbt # CodonAlign 密码子比对 (密码子替换分类、dN/dS选择压力分析、密码子使用偏好) +│ ├── codon_align_advanced.mbt # CodonAlign 高级密码子比对 (Z-test选择检验、Fisher精确检验、密码子比对构建器、滑窗dN/dS、BH-FDR、成对Ka/Ks表) │ ├── entrez.mbt # Entrez NCBI数据库访问 (ESearch、EFetch、PubMed/Gene/Taxonomy解析) │ ├── genome_info_db.mbt # GenomeInfoDb 基因组信息管理 (染色体信息、着丝粒位置、基因组构建、染色体臂) │ ├── interaction_set.mbt # InteractionSet 染色质交互数据 (Hi-C交互、锚点对、交互矩阵、距离分布) @@ -492,6 +645,7 @@ IvanAXu/BioSeqs/ │ ├── chem_utils.mbt # 化学计算工具 (范德华半径、共价半径、键长、键角、二面角、分子式量) │ ├── jaspar.mbt # JASPAR PFM格式解析 (模体矩阵、PWM转换、共有序列、序列扫描) │ ├── mmcif.mbt # mmCIF格式解析 (Bio.PDB.MMCIFParser、数据块、类别、原子位点) +│ ├── binary_cif.mbt # Bio.PDB.binary_cif MessagePack解析、七类逆编码、mask与Structure转换 │ ├── nexus.mbt # Nexus格式解析 (Bio.Nexus、数据矩阵、系统发育树、距离矩阵) │ ├── emboss.mbt # EMBOSS工具接口 (GC偏斜、AT偏斜、分子量、Tm值、ORF查找、距离计算) │ ├── chipseeker.mbt # ChIPseeker ChIP-seq峰注释分析 (峰-基因距离计算、注释分类(启动子/外显子/内含子/UTR/基因间区)、BED格式读取、peak2gene关联分析、结果汇总与可视化) @@ -506,13 +660,31 @@ IvanAXu/BioSeqs/ │ ├── annotation_hub.mbt # AnnotationHub 中心化注释资源访问 (资源搜索、类型/提供者/基因组查询、资源管理) │ ├── genomic_features.mbt # GenomicFeatures 基因组注释功能 (Gene/Transcript/Exon数据结构、GTF解析、区域查询) │ ├── graph.mbt # graph 图数据结构 (有向/无向图、最短路径、连通分量、DOT输出) -│ ├── droplet_utils.mbt # DropletUtils 空液滴检测 (emptyDrops算法、knee点检测、细胞过滤) +│ ├── droplet_utils.mbt # DropletUtils 旧版空液滴检测兼容API +│ ├── droplet_utils_advanced.mbt # DropletUtils 1.33.0高级emptyDrops (Good-Turing、DM概率、Monte Carlo、knee/inflection与SCE接入) │ ├── scran.mbt # scran 单细胞归一化与聚类 (sum_factors、SNN图、Leiden聚类、标志物分析) +│ ├── scrapper.mbt # scrapper 单细胞预处理 (批次感知RNA QC、大小因子、LOWESS/HVG、pseudo-bulk、SCE集成) +│ ├── scuttle.mbt # scuttle 1.23.1 batch MAD、per-feature QC、feature聚合、精确downsampling与SCE集成 +│ ├── bluster.mbt # bluster 1.23.0 K-means、KNN/SNN图聚类、two-step、聚类诊断、稳定性与SCE集成 +│ ├── flowsom.mbt # FlowSOM 2.21.0拓扑SOM、MST、meta-clustering、节点统计、outlier与容器集成 +│ ├── decontx.mbt # decontX ambient RNA去污染 (Bayesian EM、background、自动聚类、计数分解、SCE集成) +│ ├── celda.mbt # celda_CG 细胞群与基因模块联合聚类 (collapsed likelihood、EM/Gibbs、多链、模型选择、SCE集成) +│ ├── milo.mbt # miloR KNN邻域差异丰度 (精炼采样、NB-GLM、graph spatial FDR、SCE接入) +│ ├── zinbwave.mbt # zinbwave 零膨胀NB低维模型 (EM/IRLS、latent factors、observational weights、SCE接入) +│ ├── variance_partition.mbt # variancePartition 混合模型方差分解、BLUP与dream重复测量检验 +│ ├── dreamlet.mbt # dreamlet pseudobulk、TMM/voom权重、分cell-type混合模型与study-wide FDR +│ ├── nnsvg.mbt # nnSVG nearest-neighbor GP、空间变异检验、length scale与SpatialExperiment接入 +│ ├── banksy.mbt # Banksy空间邻域harmonic、lambda联合特征、PCA、聚类、平滑与SpatialExperiment接入 +│ ├── voyager.mbt # Voyager空间自相关:kNN/distance-band/inverse-distance权重、Moran's I/Geary's c(全局+局部)、Getis-Ord Gi*、Lee's L、变差函数拟合、Moran correlogram、置换检验BH-FDR与SpatialExperiment接入 +│ ├── spicyr.mbt # spicyR cross-L共定位、边界校正、加权/随机截距模型与SpatialExperiment接入 +│ ├── lisaclust.mbt # lisaClust local-K/L曲线、KDE、窗口边界修正、区域聚类与SpatialExperiment接入 +│ ├── spatialdecon.mbt # SpatialDecon背景感知log-normal解卷积、异常点重拟合、不确定度与容器接入 │ ├── monocle3.mbt # monocle3 单细胞轨迹分析 (PCA/UMAP降维、主图学习、拟时间排序) │ ├── short_read.mbt # ShortRead 短读序列质量控制 (QA统计、adapter修剪、质量修剪、读长过滤、FastQC报告) │ ├── seq_quality_trim.mbt # NGS质量修剪与接头去除 (质量修剪、接头去除、poly-A修剪、长度/GC过滤、批量修剪) │ ├── scater.mbt # scater 单细胞质量控制 (QC指标计算、细胞/基因过滤、标准化、HVG检测、PCA) -│ ├── mast.mbt # MAST 单细胞差异表达分析 (Hurdle模型、离散/连续检验、BH-FDR校正) +│ ├── mast.mbt # MAST 兼容层 (检测率/Welch检验、BH-FDR与旧版结果API) +│ ├── mast_advanced.mbt # MAST 1.39.0 Bayesian hurdle GLM、嵌套LRT、eBayes与SCE接入 │ ├── genomic_files.mbt # GenomicFiles 分布式基因组文件处理 (BAM/BED/VCF扫描、区间查询、归约、覆盖度) │ ├── diffbind.mbt # DiffBind ChIP-seq差异结合分析 (峰值重叠、共识峰、TMM归一化、NB检验) │ ├── minfi.mbt # minfi DNA甲基化分析 (NOOB/Illumina/分位数/功能归一化、β/M值、DMP/DMR分析) @@ -539,6 +711,7 @@ IvanAXu/BioSeqs/ │ ├── delayed_array.mbt # DelayedArray 延迟计算数组 (懒加载操作、分块处理、行/列聚合、子集操作) │ ├── annotation_filter.mbt # AnnotationFilter 基因注释过滤 (染色体筛选、生物类型过滤、区域重叠检测、符号模式匹配) │ ├── sc_dbl_finder.mbt # scDblFinder 单细胞双细胞检测 (Doublet评分计算、最近邻搜索、PCA降维、细胞过滤) +│ ├── sc_dbl_finder_advanced.mbt # scDblFinder 1.27.6高级流程 (人工doublet、kNN/cxds特征、迭代分类、分层阈值、来源富集与SCE接入) │ ├── batchelor.mbt # Batchelor 单细胞批次校正 (rescaleBatches、mutual nearest neighbor、fastMNN、批次混合评分) │ ├── seurat.mbt # Seurat 单细胞数据分析核心 (标准化、高可变基因、PCA、聚类、UMAP、差异表达、跨样本整合) │ ├── variation.mbt # Bio.Variation 变异分析 (SNP分析、突变检测、氨基酸替换分析、BLOSUM62/Grantham矩阵) @@ -556,7 +729,8 @@ IvanAXu/BioSeqs/ │ ├── metagenomeseq.mbt # metagenomeSeq 零膨胀模型微生物组差异丰度分析 (归一化、零膨胀概率计算) │ ├── hilbertcurve.mbt # HilbertCurve Hilbert曲线坐标映射 (编码/解码、距离计算、基因组线性化) │ ├── taxonomy.mbt # Taxonomy 分类学分析 (Taxon/TaxonomyDatabase、谱系查询、共同祖先计算) -│ ├── single_r.mbt # SingleR 细胞类型注释 (参考图谱、Spearman/Pearson相关性、精细调优) +│ ├── single_r.mbt # SingleR 兼容层 (参考图谱、Spearman/Pearson相关性、旧版结果API) +│ ├── single_r_advanced.mbt # SingleR 2.15.2 marker训练、分位数分类、fine-tuning、剪枝、多参考与SCE接入 │ ├── cyclone.mbt # Cyclone 细胞周期评分 (基因对比较、G1/S/G2/M期相预测) │ ├── dorothea.mbt # dorothea 转录因子活性预测 (Regulon、VIPER、置换检验) │ ├── gff.mbt # GFF GFF3格式解析 (GFFFeature/GFFRecord、属性解析、特征提取) @@ -580,11 +754,13 @@ IvanAXu/BioSeqs/ │ ├── phenotype.mbt # Bio.phenotype 表型微阵列分析 (WellRecord/PlateRecord/PhenFitParams、logistic/Gompertz拟合、CSV/JSON解析) │ ├── blast_applications.mbt # Bio.Blast.Applications BLAST命令行工具包装 (8种BLAST变体、快速构建器、参数管理) │ ├── qcp_superimposer.mbt # QCP叠加 (四元数旋转、结构比对、RMSD计算、最优叠加) +│ ├── cealign.mbt # Bio.PDB.cealign CE组合扩展结构比对 (AFP路径、Z-score、QCP叠合、全原子变换) │ ├── psea.mbt # Bio.PDB.PSEA 二级结构预测 (PseaAtom/PseaResult、CA-CA距离、虚拟二面角、H/E/C分配、三态到八态转换) │ ├── sff_io.mbt # Bio.SeqIO.SffIO SFF二进制格式解析 (SffHeader/SffRead/SffFile、二进制编码/解码、质量修剪) │ ├── seq_complexity.mbt # 序列复杂度与组成分析 (Shannon熵、GC偏斜、混沌游戏表示) │ ├── csaw.mbt # csaw ChIP-seq窗口差异分析 (滑动窗口计数、TMM归一化、窗口过滤、负二项GLM检验、差异区域检测) │ ├── slingshot.mbt # slingshot 单细胞轨迹推断 (MST构建、主曲线拟合、拟时间计算、分支检测) +│ ├── slingshot_advanced.mbt # slingshot 2.21.0 soft membership、约束forest、同时主曲线、预测与SCE接入 │ ├── scnorm.mbt # SCnorm 单细胞RNA-seq归一化 (分位数回归、深度依赖偏差校正、基因特异性归一化) │ ├── edaseq.mbt # EDASeq RNA-seq探索性分析 (GC含量归一化、基因长度校正Loess、样本间归一化、RPKM计算) │ ├── searchio.mbt # Bio.SearchIO 统一搜索结果模型 (BLAST/HMMER解析、QueryResult/Hit/HSP层次结构、E-value过滤) @@ -598,13 +774,14 @@ IvanAXu/BioSeqs/ │ ├── rtsne.mbt # Rtsne t-SNE降维算法 (距离矩阵、条件概率、梯度下降、动量优化) │ ├── uwot.mbt # uwot UMAP降维算法 (k近邻、模糊单纯集、SGD优化、负采样) │ ├── tradeseq.mbt # tradeSeq 轨迹差异表达分析 (TrajectoryPoint、GAM拟合、样条基函数、差异检验) +│ ├── tradeseq_advanced.mbt # tradeSeq 1.27.0多lineage NB-GAM、Wald检验、预测、knot评估与容器接入 │ ├── progeny.mbt # PROGENy 通路活性推断 (L2正则化线性回归、Ridge回归、通路基因集权重矩阵) │ ├── aucell.mbt # AUCell 单细胞基因集评分 (AUC计算、基因排序、min-max归一化) │ ├── geoquery.mbt # GEO数据库查询 (GDS/GSE/GSM解析、数据下载、平台信息) │ ├── ggtree.mbt # ggtree 系统发育树可视化布局 (矩形/放射状/无根布局、节点坐标映射) │ ├── mix_omics.mbt # mixOmics 多组学整合 (PLS/sPLS/DIABLO) -│ ├── maf.mbt # MAF (Multiple Alignment Format) 多序列比对格式解析与分析 -│ ├── mauve.mbt # Mauve 基因组比对格式解析与重排分析 +│ ├── maf.mbt # 早期宽松MAF块解析、选择、过滤与统计分析 +│ ├── mauve.mbt # 历史MAF-like块兼容解析与LCB重排分析(非现代XMFA) │ ├── stockholm.mbt # Stockholm 格式解析 (Pfam/Rfam比对) 与二级结构分析 │ ├── popgen_advanced.mbt # 高级群体遗传学统计 (Tajima's D, Fu & Li's D/F, MK检验) │ ├── codon_advanced.mbt # 高级密码子分析 (CAI, RSCU, ENC, GC3) @@ -630,6 +807,7 @@ IvanAXu/BioSeqs/ │ ├── karyoploter.mbt # karyoploteR 核型可视化 (染色体轨道、数据点、ASCII渲染) │ ├── system_piper.mbt # SystemPipeR 流水线编排 (步骤管理、依赖关系、进度追踪) │ ├── muscat.mbt # muscat 单细胞差异状态分析 (伪批量聚合、DS检验) +│ ├── muscat_advanced.mbt # muscat 1.27.4 NB-IRLS DS/DD、stagewise检验与SCE接入 │ ├── infercnv.mbt # infercnv 单细胞CNV推断 (基因组位置平滑、参考细胞比较、CNV评分) │ ├── scenic.mbt # SCENIC 单细胞调控网络推断 (共表达模块、Regulon构建、AUCell活性评分) │ ├── cibersort.mbt # CIBERSORT 免疫细胞去卷积 (NNLS求解、LM22风格特征矩阵、分数归一化) @@ -685,6 +863,7 @@ IvanAXu/BioSeqs/ │ ├── velociraptor.mbt # velociraptor 单细胞RNA velocity (稳态回归、EM动力学模型、velocity向量、KNN嵌入投影) │ ├── compass.mbt # Bio.Compass COMPASS profile-profile比对输出解析 (版本提取、多记录解析、E值/一致性过滤、摘要) │ ├── exonerate.mbt # Bio.SearchIO.ExonerateIO Exonerate输出解析 (vulgar/cigar格式、比对块解析、内含子统计、字符串重建) +│ ├── exonerate_text.mbt # Bio.SearchIO.ExonerateIO C4文本报告、层次结果、剪接/翻译语义与坐标投影 │ ├── mmcifio.mbt # Bio.PDB.mmcifio mmCIF文件写入 (Structure序列化、20列atom_site loop、HETATM支持、值转义) │ ├── interproscan.mbt # Bio.SearchIO.InterproscanIO InterProScan输出解析 (TSV 14列、数据库/蛋白质过滤、GO提取、分组) │ ├── sasa.mbt # Bio.PDB.SASA 溶剂可及表面积 (Shrake-Rupley滚动球、Fibonacci球面采样、VDW半径、骨架/侧链拆分) @@ -713,6 +892,8 @@ IvanAXu/BioSeqs/ │ ├── transfac.mbt # Bio.Motifs.Transfac TRANSFAC转录因子结合谱解析 (PFM频率矩阵、AC/ID/DE/BF/CC字段、参考文献、共识序列) │ ├── hmmer_io.mbt # Bio.SearchIO.HmmerIO HMMER3输出解析 (domtblout域表、文本格式、Query/Hit/HSP/Domain聚合) │ ├── fasta_search_io.mbt # Bio.SearchIO.FastaIO FASTA搜索输出解析 (-m8紧凑表格、-m9带注释头、元数据提取) +│ ├── infernal_io.mbt # Bio.SearchIO.InfernalIO cmscan/cmsearch解析 (tabular 1/2/3、non-verbose文本、local-end片段) +│ ├── hhr.mbt # Bio.Align.hhr HH-suite HHR解析、命中查询、坐标映射与序列化 │ ├── gene_pop.mbt # Bio.PopGen.GenePop GenePop群体遗传学 (基因型解析、等位基因频率、杂合度、序列化往返) │ ├── stage_r.mbt # Bioconductor stageR 两阶段假设检验 (筛选+确认、Simes聚合、BH-FDR、Holm步降、OFDR控制) │ ├── enriched_heatmap.mbt # Bioconductor EnrichedHeatmap 基因组信号归一化 (窗口化、四种均值模式、行平滑、百分位裁剪) @@ -721,6 +902,7 @@ IvanAXu/BioSeqs/ │ ├── phylo_cdao.mbt # Bio.Phylo.CDAO CDAO本体RDF/XML格式 (Tree/Node/TU/Edge、Newick双向转换、命名空间处理) │ ├── smart.mbt # Bio.Smart SMART蛋白质结构域数据库解析 (结构域分类、E值过滤、GO注释、查询与摘要) │ ├── protein_analysis.mbt # Bio.protein_analysis 蛋白质序列高级分析 (疏水性、GOR二级结构、抗原性、跨膜预测、保守性) +│ ├── protein_analysis_advanced.mbt # 高级蛋白质序列预测 (Chou-Fasman二级结构、IUPred无序区、COILS卷曲螺旋、Kolaskar抗原性、Emini表面可及性、Karplus-Schulz柔柔性) │ ├── pcd.mbt # Bio.PCD 质谱PCD格式解析 (图谱解析、TIC/BPC色谱图、峰过滤、前体离子、序列化) │ └── utils.mbt # 通用工具函数 ├── examples/ # 示例程序 @@ -751,6 +933,7 @@ IvanAXu/BioSeqs/ │ ├── consensus_cluster_plus_demo/ # ConsensusClusterPlus 共识聚类示例 │ ├── cyclone_demo/ # Cyclone 细胞周期评分示例 (基因对比较、G1/S/G2/M期相预测) │ ├── codon_align_demo/ # CodonAlign 密码子比对示例 (密码子替换分类、dN/dS选择压力分析、密码子使用偏好) +│ ├── codon_align_advanced_demo/ # CodonAlign 高级密码子比对示例 (Z-test选择检验、Fisher精确检验、密码子比对构建器、滑窗dN/dS、成对Ka/Ks表) │ ├── codon_usage_demo/ # CodonUsage 密码子使用分析示例 (CAI、ENC、RSCU、GC3、CBI、Fop、最优密码子检测) │ ├── cram_demo/ # CRAM 格式解析示例 (压缩二进制序列比对格式、CRAM转BAM、参考序列管理) │ ├── de_bruijn_demo/ # De Bruijn Graph 序列组装示例 @@ -765,6 +948,8 @@ IvanAXu/BioSeqs/ │ ├── enrichplot_demo/ # enrichplot 富集分析结果可视化示例 │ ├── ensembldb_demo/ # ensembldb Ensembl注释数据库接口示例 │ ├── expasy_demo/ # ExPASy 蛋白质分析工具接口示例 +│ ├── cellosaurus_demo/ # Cellosaurus 记录解析、查询与序列化示例 +│ ├── unigene_demo/ # UniGene cluster解析、子记录查询与序列化往返示例 │ ├── entrez_demo/ # Entrez NCBI数据库访问示例 (ESearch、EFetch、PubMed/Gene/Taxonomy解析) │ ├── faidx_demo/ # FASTA 索引示例 │ ├── fgsea_demo/ # fgsea 快速基因集富集分析示例 (置换检验、NES/ES计算、Leading Edge基因) @@ -772,6 +957,7 @@ IvanAXu/BioSeqs/ │ ├── genomic_alignments_demo/ # GenomicAlignments 基因组比对分析示例 (GAlignments、coverage、summarizeOverlaps、pileup) │ ├── genomic_ranges_demo/ # GenomicRanges 基因组区间操作示例 │ ├── genomic_ranges_advanced_demo/ # GenomicRanges tile/slidingWindows/区间运算示例 +│ ├── granges_list_demo/ # GRangesList 复合转录本、精确重叠与分组实验容器示例 │ ├── geoquery_demo/ # GEOquery GEO数据库示例 (Series Matrix解析、SOFT格式解析、ExpressionSet转换、基因过滤) │ ├── go_enrichment_demo/ # GOEnrichment GO功能富集分析示例 (超几何检验、BH校正、富集结果过滤) │ ├── hmm_demo/ # Hidden Markov Model 基因预测示例 @@ -787,6 +973,7 @@ IvanAXu/BioSeqs/ │ ├── medline_demo/ # Medline/PubMed解析示例 (文献记录、APA引用、MeSH过滤) │ ├── ml_features/ # 机器学习特征提取示例 │ ├── mmcif_demo/ # mmCIF格式解析示例 (数据块解析、类别查询、原子位点提取) +│ ├── binary_cif_demo/ # BinaryCIF MessagePack解析、编码管线、mask与PDB Structure转换示例 │ ├── motifs_demo/ # 序列模体识别示例 │ ├── motifs_advanced_demo/ # 模体高级功能示例 (JASPAR/TRANSFAC解析、模体比对、KL/JS散度、模体聚类) │ ├── multi_assay_experiment_demo/ # MultiAssayExperiment 多组学数据协调示例 (实验协调、样本映射) @@ -794,7 +981,11 @@ IvanAXu/BioSeqs/ │ ├── neighbor_search_demo/ # NeighborSearch KD树近邻搜索示例 (半径搜索、最近邻、原子对搜索) │ ├── nexus_demo/ # Nexus格式解析示例 (数据矩阵、系统发育树、距离矩阵) │ ├── olc_demo/ # Overlap-Layout-Consensus 序列组装示例 -│ ├── paml_demo/ # Bio.PAML 分子进化分析示例 +│ ├── paml_demo/ # 历史内置近似dN/dS示例 +│ ├── paml_codeml_demo/ # CODEML control/result解析、BEB与模型比较示例 +│ ├── paml_baseml_demo/ # BASEML control、REV矩阵、离散gamma与LRT示例 +│ ├── paml_yn00_demo/ # YN00 control、五种估计、对称矩阵与汇总示例 +│ ├── blast_xml_advanced_demo/ # BLAST XML1/XML2、命中/HSP、translated坐标与规范往返示例 │ ├── pdb_analysis_demo/ # PDB 高级结构分析示例 (主链二面角、氢键检测、二级结构分配、Ramachandran图、SASA计算、疏水性分析) │ ├── pdb_demo/ # PDB 结构解析示例 │ ├── pdb_list_demo/ # Bio.PDB.PDBList PDB结构下载管理示例 @@ -818,12 +1009,14 @@ IvanAXu/BioSeqs/ │ ├── seqfeature_advanced_demo/ # Bio.SeqFeature CompoundLocation与LocationParser │ ├── rna_structure_demo/ # RNA二级结构预测示例 │ ├── single_cell_demo/ # SingleCell 单细胞数据分析示例 (QC指标、Log标准化、PCA降维、高变异基因) -│ ├── single_r_demo/ # SingleR 细胞类型注释示例 (参考图谱、Spearman/Pearson相关性、精细调优) +│ ├── single_r_demo/ # SingleR 2.15.2 markers、分位数分类、cluster、多参考与SCE写回示例 │ ├── smith_waterman_demo/ # Smith-Waterman 局部序列比对示例 │ ├── subsmat_demo/ # 替换矩阵示例 (BLOSUM62/45、PAM250/30矩阵查询、蛋白质比对打分) │ ├── substitution_matrices_demo/ # 现代替换矩阵示例 (矩阵注册表、频率矩阵计算、log-odds打分、Shannon熵、KL散度、NCBI解析) │ ├── suffix_array_tree_demo/ # Suffix Array & Suffix Tree 示例 │ ├── summarized_experiment_demo/ # SummarizedExperiment 数据容器示例 +│ ├── ranged_summarized_experiment_demo/ # RangedSummarizedExperiment GRanges重叠、最近邻、区间变换与排序示例 +│ ├── tree_summarized_experiment_demo/ # TreeSummarizedExperiment 行/列树链接、节点子集与聚合示例 │ ├── sva_demo/ # sva 替代变量分析与ComBat批次校正示例 (经验贝叶斯方法、PCA分析) │ ├── svd_superimposer_demo/ # SVDSuperimposer SVD蛋白质结构叠合示例 (旋转矩阵、平移向量、RMSD计算) │ ├── structure_alignment_demo/ # Bio.PDB.StructureAlignment 多蛋白质结构比对示例 @@ -840,12 +1033,31 @@ IvanAXu/BioSeqs/ │ ├── annotation_hub_demo/ # AnnotationHub 中心化注释资源访问示例 (资源搜索、类型/提供者/基因组查询) │ ├── genomic_features_demo/ # GenomicFeatures 基因组注释示例 (GTF解析、基因/转录本/外显子查询、区域查询) │ ├── graph_demo/ # graph 图数据结构示例 (有向/无向图构建、最短路径、连通分量、DOT输出) -│ ├── droplet_utils_demo/ # DropletUtils 空液滴检测示例 (emptyDrops算法、knee点检测、细胞过滤) +│ ├── droplet_utils_demo/ # DropletUtils ambient profile、alpha估计、barcode-rank、emptyDrops、过滤与SCE写回示例 │ ├── scran_demo/ # scran 单细胞归一化与聚类示例 (sum_factors、SNN图、Leiden聚类、标志物分析) +│ ├── scrapper_demo/ # scrapper 批次感知RNA QC、归一化、LOWESS/HVG、pseudo-bulk与SCE集成示例 +│ ├── scuttle_demo/ # scuttle batch MAD、subset QC、feature聚合、coverage等化与不可变SCE示例 +│ ├── bluster_demo/ # bluster K-means、SNN图、two-step、聚类诊断、bootstrap稳定性与不可变SCE示例 +│ ├── flowsom_demo/ # FlowSOM拓扑训练、MST、节点统计、outlier、新数据与FlowFrame/SCE示例 +│ ├── decontx_demo/ # decontX cluster/background去污染、marker校正、诊断与SCE输出示例 +│ ├── celda_demo/ # celda_CG联合聚类、module marker、细胞预测、模型选择与SCE输出示例 +│ ├── milo_demo/ # miloR KNN图、精炼邻域、NB差异丰度、spatial FDR与SCE接入示例 +│ ├── zinbwave_demo/ # zinbwave latent factors、dropout权重、残差/插补与SCE集成示例 +│ ├── apeglm_demo/ # apeglm MLE/MAP、重尾收缩、FSR/FSOS、TSV与SE接入示例 +│ ├── aldex2_demo/ # ALDEx2 IQLR、Dirichlet实例、effect/eBH、距离与SE接入示例 +│ ├── dirichlet_multinomial_demo/ # DMM聚类、选K、分组分类、交叉验证、ROC与SE接入示例 +│ ├── variance_partition_demo/ # variancePartition 方差分解、BLUP、precision weights、dream与SE接入示例 +│ ├── dreamlet_demo/ # dreamlet SCE pseudobulk、TMM/voom、donor随机截距与跨cell-type FDR示例 +│ ├── nnsvg_demo/ # nnSVG空间变异基因、length scale、过滤与SpatialExperiment接入示例 +│ ├── banksy_demo/ # Banksy H0/H1、lambda扫描、PCA聚类、平滑与SpatialExperiment接入示例 +│ ├── voyager_demo/ # Voyager空间权重、Moran/Geary/Getis-Ord/Lee's L、变差函数、correlogram与SpatialExperiment接入示例 +│ ├── spicyr_demo/ # spicyR cross-L、条件对比、重复受试者模型与SpatialExperiment接入示例 +│ ├── lisaclust_demo/ # lisaClust local-K/L、区域聚类、富集与SpatialExperiment写回示例 +│ ├── spatialdecon_demo/ # SpatialDecon背景校正、丰度/计数、collapse、reverse与SpatialExperiment示例 │ ├── monocle3_demo/ # monocle3 单细胞轨迹分析示例 (PCA/UMAP降维、主图学习、拟时间排序) │ ├── short_read_demo/ # ShortRead 短读序列质量控制示例 (QA统计、adapter修剪、质量修剪、FastQC报告) │ ├── scater_demo/ # scater 单细胞质量控制示例 (QC指标计算、细胞/基因过滤、标准化、HVG检测、PCA) -│ ├── mast_demo/ # MAST 单细胞差异表达分析示例 (Hurdle模型、离散/连续检验、BH-FDR校正) +│ ├── mast_demo/ # MAST 1.39.0双组分GLM、LRT/eBayes、边际效应与SCE写回示例 │ ├── genomic_files_demo/ # GenomicFiles 分布式基因组文件处理示例 (BAM/BED/VCF扫描、区间查询、归约、覆盖度) │ ├── diffbind_demo/ # DiffBind ChIP-seq差异结合分析示例 (峰值重叠、共识峰、TMM归一化、NB检验) │ ├── minfi_demo/ # minfi DNA甲基化分析示例 (NOOB/Illumina/分位数/功能归一化、β/M值、DMP/DMR分析) @@ -864,9 +1076,10 @@ IvanAXu/BioSeqs/ │ ├── gsva_demo/ # GSVA 基因集变异分析示例 (ssGSEA/zscore/PLAGE评分、富集分析、置换检验、富集图可视化、表型相关性分析、生存分析、分数分布分析) │ ├── chromvar_demo/ # ChromVAR 染色质变异分析示例 (TF motif富集、GC偏差校正、细胞聚类、变异性分析) │ ├── delayed_array_demo/ # DelayedArray 延迟计算数组示例 (懒加载操作、分块处理、行/列聚合、子集操作) +│ ├── sparse_array_demo/ # SparseArray N维稀疏张量、切片/置换、统计、算术与矩阵乘法示例 │ ├── annotation_filter_demo/ # AnnotationFilter 基因注释过滤示例 (染色体筛选、生物类型过滤、区域重叠检测、符号模式匹配) │ ├── sc3_demo/ # SC3 单细胞共识聚类示例 -│ ├── sc_dbl_finder_demo/ # scDblFinder 单细胞双细胞检测示例 (Doublet评分计算、最近邻搜索、PCA降维、细胞过滤) +│ ├── sc_dbl_finder_demo/ # scDblFinder高级流程示例 (capture分层、已知doublet、来源富集、自动聚类、过滤与SCE写回) │ ├── batchelor_demo/ # Batchelor 单细胞批次校正示例 (rescaleBatches、fastMNN、mutual nearest neighbor、批次混合评分) │ ├── seurat_demo/ # Seurat 单细胞数据分析示例 (标准化、高可变基因、PCA、聚类、UMAP、差异标志物分析、跨样本整合) │ ├── chipseeker_demo/ # ChIPseeker ChIP-seq峰值注释示例 (基因组区域分类(启动子/外显子/内含子/UTR/基因间区)、距离TSS分布、BED格式读取、peak2gene关联分析、注释可视化、统计分析) @@ -906,10 +1119,12 @@ IvanAXu/BioSeqs/ │ ├── sff_io_demo/ # Bio.SeqIO.SffIO SFF二进制解析示例 (二进制编码/解码往返、质量修剪、按名称查找) │ ├── csaw_demo/ # csaw ChIP-seq窗口差异分析示例 (滑动窗口、TMM归一化、差异区域检测) │ ├── slingshot_demo/ # slingshot 单细胞轨迹推断示例 (MST构建、主曲线、拟时间计算) +│ ├── slingshot_advanced_demo/ # slingshot约束MST、同时主曲线、分支、预测及不可变SCE示例 │ ├── scnorm_demo/ # SCnorm 单细胞RNA-seq归一化示例 (分位数回归、深度依赖校正) │ ├── edaseq_demo/ # EDASeq RNA-seq探索性分析示例 (GC归一化、Loess校正、RPKM计算) │ ├── pdb_vectors_demo/ # Bio.PDB.vectors 3D向量与旋转矩阵示例 (Vector3运算、Kabsch叠合、二面角计算) │ ├── qcp_superimposer_demo/ # Bio.PDB.QCPSuperimposer 四元数结构叠合示例 +│ ├── cealign_demo/ # Bio.PDB.cealign CE组合扩展结构比对、路径统计与不可变全原子变换示例 │ ├── circ_seq_demo/ # Bio.SeqUtils.CircSeq 环状DNA操作示例 (酶切分析、PCR引物设计、序列旋转) │ ├── align_abstract_demo/ # Bio.Align.AlignAbstract 抽象比对示例 (一致性序列、Shannon熵、同一性矩阵、简约信息位点) │ ├── maftools_demo/ # maftools 癌症基因组学示例 (MAF数据创建、突变分类、TMB计算、突变谱分析) @@ -919,12 +1134,13 @@ IvanAXu/BioSeqs/ │ ├── rtsne_demo/ # Rtsne t-SNE降维示例 (距离矩阵、条件概率、梯度下降、动量优化) │ ├── uwot_demo/ # uwot UMAP降维示例 (k近邻、模糊单纯集、SGD优化、负采样) │ ├── tradeseq_demo/ # tradeSeq 轨迹差异表达示例 (GAM拟合、基因平滑、差异检验、可视化) +│ ├── tradeseq_advanced_demo/ # tradeSeq NB-GAM、五类检验、预测、knot、Slingshot与SCE示例 │ ├── progeny_demo/ # PROGENy 通路活性推断示例 (L2正则化回归、通路基因集、样本活性计算) │ ├── aucell_demo/ # AUCell 单细胞基因集评分示例 (AUC计算、归一化、细胞/基因集评分) │ ├── ggtree_demo/ # ggtree 系统发育树可视化示例 (矩形/放射状/无根布局、节点坐标) │ ├── mix_omics_demo/ # mixOmics 多组学整合演示 │ ├── maf_demo/ # MAF 格式解析与分析示例 (解析、统计、选择、过滤、写回) -│ ├── mauve_demo/ # Mauve 基因组比对分析示例 (倒位检测、断点检测、覆盖率、BED导出) +│ ├── mauve_demo/ # 历史MAF-like块重排分析示例 (倒位检测、断点检测、覆盖率、BED导出) │ ├── stockholm_demo/ # Stockholm 格式解析示例 (Pfam/Rfam格式、二级结构、保守性分析) │ ├── popgen_advanced_demo/ # 高级群体遗传学示例 (Tajima's D、Fu & Li、MK检验、中性分析) │ ├── codon_advanced_demo/ # 高级密码子分析示例 (CAI、RSCU、ENC、最优/稀有密码子) @@ -950,6 +1166,7 @@ IvanAXu/BioSeqs/ │ ├── karyoploter_demo/ # karyoploteR 核型可视化示例 │ ├── system_piper_demo/ # SystemPipeR 流水线编排示例 │ ├── muscat_demo/ # muscat 单细胞差异状态分析示例 +│ ├── muscat_advanced_demo/ # muscat 1.27.4伪批量、NB-IRLS DS/DD、stagewise与SCE示例 │ ├── infercnv_demo/ # infercnv 单细胞拷贝数变异推断示例 │ ├── scenic_demo/ # SCENIC 单细胞调控网络推断示例 │ ├── cibersort_demo/ # CIBERSORT 免疫细胞去卷积示例 @@ -1011,6 +1228,7 @@ IvanAXu/BioSeqs/ │ ├── velociraptor_demo/ # velociraptor RNA velocity示例 (稳态/动力学模型、velocity向量、嵌入投影、根细胞识别) │ ├── compass_demo/ # Bio.Compass COMPASS比对输出解析示例 (profile-profile比对解析、E值过滤、摘要) │ ├── exonerate_demo/ # Bio.SearchIO.ExonerateIO Exonerate输出解析示例 (vulgar/cigar解析、内含子统计、字符串重建) +│ ├── exonerate_text_demo/ # Exonerate C4层次结果、fragment坐标、intron区间与统计的离线示例 │ ├── mmcifio_demo/ # Bio.PDB.mmcifio mmCIF写入示例 (Structure序列化、atom_site loop、round-trip验证) │ ├── interproscan_demo/ # Bio.SearchIO.InterproscanIO InterProScan解析示例 (TSV解析、数据库过滤、GO条目、分组) │ ├── sasa_demo/ # Bio.PDB.SASA 溶剂可及表面积示例 (Shrake-Rupley算法、逐原子/残基SASA、骨架/侧链拆分) @@ -1039,6 +1257,29 @@ IvanAXu/BioSeqs/ │ ├── transfac_demo/ # TRANSFAC转录因子结合谱解析示例 (PFM矩阵、共识序列、频率计算、序列化、参考文献) │ ├── hmmer_io_demo/ # HMMER3输出解析示例 (domtblout域表、文本格式、Query/Hit/HSP聚合、多域比对) │ ├── fasta_search_io_demo/ # FASTA搜索输出解析示例 (-m8表格、-m9注释头、元数据、Query/Hit/HSP聚合) +│ ├── infernal_io_demo/ # Infernal cmscan/cmsearch解析示例 (tabular 3、文本local-end、过滤、SearchIO转换) +│ ├── hhr_demo/ # HH-suite HHR解析、命中筛选、坐标映射与序列化往返示例 +│ ├── shared_reference_alignment_demo/ # 共享参考PWA/MSA合并、insertion同步、坐标映射与FASTA转换示例 +│ ├── alignment_map_demo/ # Alignment.map/mapall、反链、PSL与protein-to-codon MSA投影示例 +│ ├── alignment_counts_demo/ # Alignment.counts gap分类、affine/BLOSUM评分、反链与MSA汇总示例 +│ ├── align_tabular_demo/ # BLAST BTOP、FASTA aln_code与translated反链坐标解析示例 +│ ├── align_psl_demo/ # PSL/PSLX读写、反链映射、translated recount与文档摘要示例 +│ ├── align_sam_demo/ # SAM header/path、反链、splicing、typed tags、MD/NM与往返示例 +│ ├── a2m_demo/ # A2M状态解析、插入槽、共识、坐标、统计、投影与往返示例 +│ ├── align_emboss_demo/ # EMBOSS元数据、局部/反向坐标、path、统计与wrapped往返示例 +│ ├── align_exonerate_demo/ # Exonerate剪接/translated path、坐标映射、统计与双格式往返示例 +│ ├── msf_demo/ # GCG MSF解析、checksum、坐标映射、统计、宽度异常与规范往返示例 +│ ├── align_nexus_demo/ # NEXUS interleave解析、quoted taxa、坐标统计、all-gap压缩与规范往返示例 +│ ├── align_stockholm_demo/ # Stockholm GF/GS/GR/GC、insertion列、坐标统计、all-gap压缩与规范往返示例 +│ ├── align_chain_demo/ # Chain多记录、正反链路径、block/counts、坐标映射、反转与规范往返示例 +│ ├── align_maf_demo/ # MAF track/a/s/i/e/q、绝对路径、component映射、索引、拼接与规范往返示例 +│ ├── align_mauve_demo/ # XMFA metadata/LCB、正负链路径、统计、索引、投影、重建与规范往返示例 +│ ├── align_clustal_demo/ # CLUSTAL metadata、consensus、坐标路径、投影、统计与累计计数往返示例 +│ ├── align_phylip_demo/ # PHYLIP布局识别、名称规范化、坐标、统计及两种布局往返示例 +│ ├── align_bed_demo/ # BED3-BED12、正负链路径、block、坐标映射、搜索、统计与规范往返示例 +│ ├── bigbed_demo/ # BigBed写入/解析、索引查询、负链exon坐标、BED导出与损坏诊断 +│ ├── bigmaf_demo/ # BigMaf压缩写入、bedMaf schema、区间查询、负链映射与MAF导出 +│ ├── bigpsl_demo/ # BigPsl压缩写入、R-tree查询、反链/translated坐标与PSL导出 │ ├── gene_pop_demo/ # GenePop群体遗传学示例 (基因型解析、等位基因频率、杂合度统计、序列化往返) │ ├── stage_r_demo/ # stageR两阶段检验示例 (筛选+确认、Simes聚合、BH-FDR、Holm步降、OFDR控制) │ ├── enriched_heatmap_demo/ # EnrichedHeatmap富集热图示例 (信号归一化、四种均值模式、行平滑、链方向处理) @@ -1047,6 +1288,7 @@ IvanAXu/BioSeqs/ │ ├── phylo_cdao_demo/ # CDAO本体RDF/XML示例 (Tree/Node/TU构建、解析往返、Newick转换) │ ├── smart_demo/ # SMART结构域解析示例 (结构域检测、E值过滤、GO注释、摘要报告) │ ├── protein_analysis_demo/ # 蛋白质分析示例 (疏水性、GOR二级结构、抗原性、跨膜预测、保守性) +│ ├── protein_analysis_advanced_demo/ # 高级蛋白质序列预测示例 (Chou-Fasman二级结构、IUPred无序区、COILS卷曲螺旋、Kolaskar抗原性、Emini表面可及性、Karplus-Schulz柔柔性) │ ├── pcd_demo/ # 质谱PCD格式示例 (图谱解析、TIC/BPC色谱图、峰过滤、序列化往返) ├── test/ │ ├── moonbit/ # MoonBit 测试文件 @@ -1059,6 +1301,7 @@ IvanAXu/BioSeqs/ │ │ ├── bio_seq_wb_test.mbt │ │ ├── biostrings_test.mbt │ │ ├── blast_test.mbt +│ │ ├── blast_xml_advanced_test.mbt │ │ ├── bloom_filter_test.mbt │ │ ├── bwt_fm_test.mbt │ │ ├── cluster_test.mbt @@ -1076,6 +1319,7 @@ IvanAXu/BioSeqs/ │ │ ├── genomic_alignments_test.mbt │ │ ├── genomic_ranges_test.mbt │ │ ├── genomic_ranges_advanced_test.mbt +│ │ ├── granges_list_test.mbt │ │ ├── go_enrichment_test.mbt │ │ ├── hmm_test.mbt │ │ ├── hmm_wbtest.mbt @@ -1101,12 +1345,15 @@ IvanAXu/BioSeqs/ │ │ ├── sequtils_test.mbt │ │ ├── single_cell_test.mbt │ │ ├── single_r_test.mbt +│ │ ├── single_r_advanced_test.mbt │ │ ├── smith_waterman_test.mbt │ │ ├── subsmat_test.mbt │ │ ├── substitution_matrices_test.mbt │ │ ├── suffix_array_tree_test.mbt │ │ ├── suffix_array_tree_wbtest.mbt │ │ ├── summarized_experiment_test.mbt +│ │ ├── ranged_summarized_experiment_test.mbt +│ │ ├── tree_summarized_experiment_test.mbt │ │ ├── svd_superimposer_test.mbt │ │ ├── tree_io_test.mbt │ │ ├── txdb_test.mbt @@ -1145,6 +1392,8 @@ IvanAXu/BioSeqs/ │ │ ├── bioc_generics_test.mbt │ │ ├── bioc_parallel_test.mbt │ │ ├── bsseq_test.mbt +│ │ ├── cellosaurus_test.mbt +│ │ ├── unigene_test.mbt │ │ ├── checksum_test.mbt │ │ ├── chipseeker_test.mbt │ │ ├── chromvar_test.mbt @@ -1155,6 +1404,7 @@ IvanAXu/BioSeqs/ │ │ ├── consensus_cluster_plus_test.mbt │ │ ├── csaw_test.mbt │ │ ├── delayed_array_test.mbt +│ │ ├── sparse_array_test.mbt │ │ ├── destiny_test.mbt │ │ ├── rtsne_test.mbt │ │ ├── uwot_test.mbt @@ -1162,6 +1412,7 @@ IvanAXu/BioSeqs/ │ │ ├── diffbind_test.mbt │ │ ├── dose_test.mbt │ │ ├── droplet_utils_test.mbt +│ │ ├── droplet_utils_advanced_test.mbt │ │ ├── dss_test.mbt │ │ ├── dssp_test.mbt │ │ ├── edaseq_test.mbt @@ -1187,6 +1438,7 @@ IvanAXu/BioSeqs/ │ │ ├── kmer_test.mbt │ │ ├── maftools_test.mbt │ │ ├── mast_test.mbt +│ │ ├── mast_advanced_test.mbt │ │ ├── matrix_test.mbt │ │ ├── medline_test.mbt │ │ ├── melting_temp_test.mbt @@ -1198,6 +1450,9 @@ IvanAXu/BioSeqs/ │ │ ├── nexus_test.mbt │ │ ├── nucle_r_test.mbt │ │ ├── paml_test.mbt +│ │ ├── paml_codeml_test.mbt +│ │ ├── paml_baseml_test.mbt +│ │ ├── paml_yn00_test.mbt │ │ ├── pathway_test.mbt │ │ ├── pdb_analysis_test.mbt │ │ ├── pdb_dice_test.mbt @@ -1218,15 +1473,55 @@ IvanAXu/BioSeqs/ │ │ ├── prosite_test.mbt │ │ ├── psea_test.mbt │ │ ├── qcp_superimposer_test.mbt +│ │ ├── cealign_test.mbt │ │ ├── reactome_pa_test.mbt │ │ ├── residue_depth_test.mbt │ │ ├── rhdf5_test.mbt │ │ ├── s4vectors_test.mbt │ │ ├── sc3_test.mbt │ │ ├── sc_dbl_finder_test.mbt +│ │ ├── sc_dbl_finder_advanced_test.mbt │ │ ├── scater_test.mbt │ │ ├── scnorm_test.mbt │ │ ├── scran_test.mbt +│ │ ├── scrapper_test.mbt +│ │ ├── scuttle_test.mbt +│ │ ├── bluster_test.mbt +│ │ ├── flowsom_test.mbt +│ │ ├── decontx_test.mbt +│ │ ├── celda_test.mbt +│ │ ├── milo_test.mbt +│ │ ├── zinbwave_test.mbt +│ │ ├── apeglm_test.mbt +│ │ ├── aldex2_test.mbt +│ │ ├── dirichlet_multinomial_test.mbt +│ │ ├── variance_partition_test.mbt +│ │ ├── dreamlet_test.mbt +│ │ ├── nnsvg_test.mbt +│ │ ├── banksy_test.mbt +│ │ ├── voyager_test.mbt +│ │ ├── shared_reference_alignment_test.mbt +│ │ ├── alignment_map_test.mbt +│ │ ├── alignment_counts_test.mbt +│ │ ├── align_tabular_test.mbt +│ │ ├── align_psl_test.mbt +│ │ ├── align_sam_test.mbt +│ │ ├── a2m_test.mbt +│ │ ├── align_emboss_test.mbt +│ │ ├── align_exonerate_test.mbt +│ │ ├── exonerate_text_test.mbt +│ │ ├── msf_test.mbt +│ │ ├── align_nexus_test.mbt +│ │ ├── align_stockholm_test.mbt +│ │ ├── align_chain_test.mbt +│ │ ├── align_maf_test.mbt +│ │ ├── align_mauve_test.mbt +│ │ ├── align_clustal_test.mbt +│ │ ├── align_phylip_test.mbt +│ │ ├── align_bed_test.mbt +│ │ ├── bigbed_test.mbt +│ │ ├── bigmaf_test.mbt +│ │ ├── bigpsl_test.mbt │ │ ├── search_io_test.mbt │ │ ├── searchio_new_test.mbt │ │ ├── seq_complexity_test.mbt @@ -1236,12 +1531,14 @@ IvanAXu/BioSeqs/ │ │ ├── seq_quality_trim_test.mbt │ │ ├── single_cell_experiment_test.mbt │ │ ├── slingshot_test.mbt +│ │ ├── slingshot_advanced_test.mbt │ │ ├── spatial_experiment_test.mbt │ │ ├── statistics_test.mbt │ │ ├── structure_alignment_test.mbt │ │ ├── taxonomy_test.mbt │ │ ├── topgo_test.mbt │ │ ├── tradeseq_test.mbt +│ │ ├── tradeseq_advanced_test.mbt │ │ ├── tximport_test.mbt │ │ ├── universalmotif_test.mbt │ │ ├── variant_filtering_test.mbt @@ -1254,6 +1551,8 @@ IvanAXu/BioSeqs/ │ │ ├── stockholm_test.mbt │ │ ├── popgen_advanced_test.mbt │ │ ├── codon_advanced_test.mbt +│ │ ├── codon_align_advanced_test.mbt +│ │ ├── protein_analysis_advanced_test.mbt │ │ ├── pdb_packing_test.mbt │ │ ├── qvalue_test.mbt │ │ ├── ihw_test.mbt @@ -1276,6 +1575,7 @@ IvanAXu/BioSeqs/ │ │ ├── karyoploter_test.mbt │ │ ├── system_piper_test.mbt │ │ ├── muscat_test.mbt +│ │ ├── muscat_advanced_test.mbt │ │ ├── infercnv_test.mbt │ │ ├── scenic_test.mbt │ │ ├── cibersort_test.mbt @@ -1355,6 +1655,7 @@ IvanAXu/BioSeqs/ │ │ ├── gck_io_test.mbt │ │ ├── alignace_test.mbt │ │ ├── mmtf_test.mbt +│ │ ├── binary_cif_test.mbt │ │ ├── naccess_test.mbt │ │ ├── wise_test.mbt │ │ ├── dnashape_test.mbt @@ -1366,6 +1667,8 @@ IvanAXu/BioSeqs/ │ │ ├── transfac_full_test.mbt │ │ ├── hmmer_io_test.mbt │ │ ├── fasta_search_io_test.mbt +│ │ ├── infernal_io_test.mbt +│ │ ├── hhr_test.mbt │ │ ├── gene_pop_test.mbt │ │ ├── stage_r_test.mbt │ │ ├── enriched_heatmap_test.mbt @@ -1400,7 +1703,7 @@ IvanAXu/BioSeqs/ ### 样例测试 ``` moon build # ✅ 成功 -moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全部通过 +moon test # ✅ 12209 个测试全部通过 ``` ### 模块对照表 @@ -1437,13 +1740,39 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `alignio.mbt` | BioPython `Bio.AlignIO` | 比对文件 I/O | | `align_io.mbt` | BioPython `Bio.AlignIO` | ClustalW/FASTA/Stockholm 解析 | | `clustal_io.mbt` | BioPython `Bio.AlignIO.ClustalIO` | Clustal 格式 | +| `align_clustal.mbt` | BioPython `Bio.Align.clustal` | 现代CLUSTAL generator metadata、严格interleaved blocks、consensus、coordinate path、统计与canonical writer | | `phylip_io.mbt` | BioPython `Bio.AlignIO.PhylipIO` | PHYLIP 格式 | +| `align_phylip.mbt` | BioPython `Bio.Align.phylip` | 现代PHYLIP严格header与固定宽度名称、sequential/interleaved自动识别、coordinate path、统计及canonical writer | | `subsmat.mbt` | BioPython `Bio.SubsMat` | BLOSUM/PAM 替换矩阵 | | `substitution_matrices.mbt` | BioPython `Bio.Align.substitution_matrices` | 现代替换矩阵基础设施 (ArrayData、矩阵注册表、频率矩阵、log-odds、Shannon熵、KL散度) | | `align_info.mbt` | BioPython `Bio.Align.AlignInfo` | 比对统计与一致性序列 | | `align_abstract.mbt` | BioPython `Bio.Align.AlignAbstract` | 抽象比对类型、Shannon熵、同一性矩阵、简约信息位点 | | `codon_align.mbt` | BioPython `Bio.codonalign` | 密码子比对与 dN/dS 分析 | +| `codon_align_advanced.mbt` | BioPython `Bio.codonalign` | 高级密码子比对 (Z-test选择检验、Fisher精确检验、密码子比对构建器、滑窗dN/dS、BH-FDR、成对Ka/Ks表) | +| `protein_analysis_advanced.mbt` | Biopython `Bio.SeqUtils` | 高级蛋白质序列预测 (Chou-Fasman二级结构、IUPred无序区、COILS卷曲螺旋、Kolaskar抗原性、Emini表面可及性、Karplus-Schulz柔柔性) | | `searchio.mbt` | BioPython `Bio.SearchIO` | 统一搜索结果模型、BLAST/HMMER解析、E-value过滤 | +| `blast_xml_advanced.mbt` | BioPython `Bio.Blast` | XML1/XML2类型化文档、严格parser/writer、多query/report、parameters/statistics、description/taxonomy及链向/translated HSP坐标 | +| `exonerate_text.mbt` | BioPython `Bio.SearchIO.ExonerateIO.exonerate_text` | C4文本Document/Query/Hit/HSP/Fragment层次、3/4/5行模型、剪接/NER/frameshift和链感知坐标 | +| `hhr.mbt` | BioPython `Bio.Align.hhr` | HHsearch/HHblits HHR解析、profile比对注释、命中筛选、坐标映射与规范序列化 | +| `shared_reference_alignment.mbt` | BioPython `Bio.Align.Alignment.from_alignments_with_same_reference` | 共享参考PWA/MSA合并、insertion slot同步、局部坐标、双向映射、统计、MSA与FASTA转换 | +| `alignment_map.mbt` | BioPython `Bio.Align.Alignment.map/mapall` | alignment path组合、local clipping、gap与链向传播、双向坐标查询、PSL、map_many及codon-aware MSA投影 | +| `alignment_counts.mbt` | BioPython `Bio.Align.Alignment.counts` | left/internal/right insertion/deletion、open/extend、composition、wildcard、替换矩阵与十二类affine gap评分 | +| `align_tabular.mbt` | BioPython `Bio.Align.tabular` | BLAST outfmt 7与FASTA 8CB/8CC元数据、BTOP/CIGAR traceback、链向及translated坐标 | +| `align_psl.mbt` | BioPython `Bio.Align.psl` | PSL/PSLX header与21/23列严格读写、核酸/translated block路径、正反链、match分类、recount及坐标转换 | +| `align_sam.mbt` | BioPython `Bio.Align.sam` | SAM header/reference与record严格读写、typed tags、CIGAR path、反向链、clipping、PHRED、MD/NM及坐标互映 | +| `a2m.mbt` | BioPython `Bio.Align.a2m` | D/I列状态感知A2M读写、大小写/点编码、逐行坐标映射、插入槽、pair counts、共识及match projection | +| `align_emboss.mbt` | BioPython `Bio.Align.emboss` | srspair/pair/simple report解析、文件/比对元数据、多序列block、局部/反向坐标、pair counts、compact path与规范写回 | +| `align_exonerate.mbt` | BioPython `Bio.Align.exonerate` | alignment-aware cigar/vulgar严格读写、完整operation path、链向与protein strand、translated mapping、统计和规范写回 | +| `msf.mbt` | BioPython `Bio.Align.msf` | GCG/PileUp MSF严格解析、AA/NA metadata、interleaved rows、标准checksum、坐标映射、统计与canonical writer | +| `align_nexus.mbt` | BioPython `Bio.Align.nexus` | DATA/CHARACTERS/TAXA block、quote/comment-aware解析、sequential/interleaved MATRIX、datatype与MATCHCHAR、坐标统计和canonical writer | +| `align_stockholm.mbt` | BioPython `Bio.Align.stockholm` | 严格多记录Stockholm读写、GF/GS/GR/GC映射、reference与database reference、M/D/I列、all-gap压缩、坐标统计和canonical writer | +| `align_chain.mbt` | BioPython `Bio.Align.chain` | UCSC Chain 12/13字段header与size/dt/dq严格读写、双轴正反链absolute path、坐标映射、反转、ID/overlap查询与canonical writer | +| `align_maf.mbt` | BioPython `Bio.Align.maf` | MAF document/track与a/s/i/e/q严格读写、链感知absolute path、component映射、MafIndex半开查询、多外显子拼接与canonical writer | +| `align_mauve.mbt` | BioPython `Bio.Align.mauve` | 现代XMFA header与LCB严格读写、combined/separate source、正负链coordinate path、区间索引、位置/区间投影、统计及重建 | +| `align_bed.mbt` | BioPython `Bio.Align.bed` | BED3-BED12 pairwise alignment读写、numeric/text score、链感知target/query path、blocks、双向映射、区间查询与分级writer | +| `bigbed.mbt` | BioPython `Bio.Align.bigbed` | BigBed v4读写、BED3-BED12、AutoSQL、多级chromosome B+ tree/R-tree、zlib/DEFLATE及索引查询 | +| `bigmaf.mbt` | BioPython `Bio.Align.bigmaf` | bedMaf bed3+1读写、MAF a/s/i/e/q、正负链坐标映射、压缩索引查询、摘要和MAF导出 | +| `bigpsl.mbt` | BioPython `Bio.Align.bigpsl` | bed12+13读写、核酸/translated protein路径、双向坐标映射、recount、压缩查询、摘要和PSL导出 | #### 系统发育树 @@ -1453,6 +1782,10 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `tree_io.mbt` | BioPython `Bio.TreeIO` | Newick/NHX 格式解析 | | `tree_construction.mbt` | BioPython `Bio.Phylo.TreeConstruction` | UPGMA/WPGMA/NJ 建树算法 | | `phylo_xml.mbt` | BioPython `Bio.Phylo.PhyloXML` | PhyloXML格式解析、序列化、Newick双向转换、分类单元注释 | +| `paml.mbt` | BioSeqs compatibility API | 内置Nei-Gojobori风格dN/dS与Jukes-Cantor近似;不运行或解析PAML | +| `paml_codeml.mbt` | BioPython `Bio.Phylo.PAML.codeml` | CODEML control规范读写、CODONML/AAML结果、分支/位点/多基因模型与AIC/BIC/LRT | +| `paml_baseml.mbt` | BioPython `Bio.Phylo.PAML.baseml` | BASEML control规范读写、11类核苷酸模型、Q/率类别/非齐次节点结果与AIC/BIC/LRT | +| `paml_yn00.mbt` | BioPython `Bio.Phylo.PAML.yn00` | YN00 control规范读写、五种成对估计、跨版本解析、对称矩阵与均值 | #### 结构生物学 @@ -1461,8 +1794,10 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `pdb.mbt` | BioPython `Bio.PDB` | PDB 数据类型 | | `pdb_io.mbt` | BioPython `Bio.PDB.PDBIO` | PDB 文件 I/O | | `svd_superimposer.mbt` | BioPython `Bio.PDB.SVDSuperimposer` | SVD 蛋白质结构叠合 | +| `cealign.mbt` | BioPython `Bio.PDB.cealign` | CE组合扩展结构比对、AFP路径、CE显著性、QCP叠合与全原子变换 | | `neighbor_search.mbt` | BioPython `Bio.PDB.NeighborSearch` | KD 树近邻搜索 | | `mmcif.mbt` | BioPython `Bio.PDB.MMCIFParser` | mmCIF 格式解析 | +| `binary_cif.mbt` | BioPython `Bio.PDB.binary_cif` | BinaryCIF MessagePack解析、七类逆编码、三态mask、类别查询与PDB Structure转换 | | `pdb_vectors.mbt` | BioPython `Bio.PDB.vectors` | 3D向量/旋转矩阵、叉积、Kabsch叠合、二面角 | | `pdb_analysis.mbt` | BioPython `Bio.PDB.StructureAnalysis` | 二面角计算、距离矩阵、接触图、氢键检测、二级结构分配、Ramachandran图、SASA计算(Shrake-Rupley)、结构质量评估、疏水性分析 | | `pdb_header.mbt` | BioPython `Bio.PDB.ParsePDBHeader` | PDB头部元数据解析 (HEADER/TITLE/COMPOUND/SOURCE/REMARK/AUTH/DBREF) | @@ -1472,12 +1807,14 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | MoonBit 文件 | 对应 Python 库 | 核心功能 | | :--- | :--- | :--- | | `sam.mbt` | pysam | SAM 文件解析 | +| `align_sam.mbt` | Biopython `Bio.Align.sam` | alignment-aware SAM严格解析与写出、显式坐标路径、typed tags、MD/NM、PHRED及正反链映射 | | `bam.mbt` | pysam | BAM 文件解析 | | `bgzf.mbt` | pysam | BGZF 解压缩 | | `vcf.mbt` | pysam | VCF 文件解析 | | `cram_wbtest.mbt` | pysam | CRAM 格式解析 | | `genomic_ranges.mbt` | Bioconductor GenomicRanges | GRanges 区间操作 | | `genomic_ranges_advanced.mbt` | GenomicRanges Tile/Windows | tile分箱、sliding_windows滑窗、tile_genome基因组覆盖、coverage_by_window覆盖度计算、bin_genome分箱统计、promoters启动子、gaps间隙、subtract区间减法 | +| `granges_list.mbt` | Bioconductor GenomicRanges | GRangesList复合特征、split/unlist/relist、逐组变换、集合运算、重叠/最近邻/覆盖度 | | `iranges.mbt` | Bioconductor IRanges | 整数区间操作 | | `genomic_alignments.mbt` | Bioconductor GenomicAlignments | GAlignments 比对分析 | | `variant_annotation.mbt` | Bioconductor VariantAnnotation | 变异注释 | @@ -1496,13 +1833,26 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `edger.mbt` | Bioconductor edgeR | DGEList 差异表达 | | `edger_advanced.mbt` | edgeR QLF/Camera/Roast | 准似然F检验、QL分散度估计、camera竞争性基因集检验、roast自足基因集检验 | | `limma.mbt` | Bioconductor limma | 线性模型与 voom 变换 | +| `variance_partition.mbt` | Bioconductor variancePartition | 多随机截距LMM、ML/REML方差分量、固定/随机/残差占比、BLUP、precision weights、dream contrast、数值Satterthwaite与BH-FDR | +| `dreamlet.mbt` | Bioconductor dreamlet | sample×cell-type pseudobulk、完整TMM、CPM/logCPM、Poisson/voom权重、typed fixed/random design、分cell-type dream拟合及两级BH-FDR | +| `nnsvg.mbt` | Bioconductor nnSVG | 坐标缩放与前驱kNN、指数协方差NNGP、协变量GLS、profile ML、空间/非空间LR检验、gene-specific length scale、BH-FDR及SpatialExperiment接入 | +| `banksy.mbt` | Bioconductor Banksy | H0邻域均值、H1+方位Fourier/Gabor harmonic、六类空间核、lambda联合矩阵、分组标准化、PCA、多起点k-means、平滑、ARI与SpatialExperiment接入 | +| `voyager.mbt` | Bioconductor Voyager | kNN/distance-band/inverse-distance权重(W/B/C/S编码)、全局Moran's I与Geary's c(Cliff-Ord随机化期望/方差/正态p)、局部Moran's I(LISA象限分类+置换推断)、局部Geary's c、Getis-Ord Gi/Gi*(Ord-Getis z)、Lee's L(全局+局部)、多元局部Geary、经验变差函数与spherical/exponential/gaussian拟合、Moran correlogram、确定性splitmix64置换、BH-FDR与SpatialExperiment不可变写回 | +| `spicyr.mbt` | Bioconductor spicyR | 图像内有序细胞类型对cross-L、矩形窗口边界校正、图像级统计、cell-count precision weights、加权固定/随机截距模型、条件对比、BH-FDR及SpatialExperiment接入 | +| `lisaclust.mbt` | Bioconductor lisaClust | 每图像local-K/centered local-L、Gaussian KDE密度权重、矩形/凸包窗口、圆盘可见面积边界修正、确定性多起点k-means、silhouette、区域富集及SpatialExperiment写回 | +| `spatialdecon.mbt` | Bioconductor SpatialDecon | 背景感知加权log-normal非负回归、两阶段异常点重拟合、observed/expected Hessian协方差、细胞丰度尺度、cell-type collapse、reverse deconvolution、负探针背景、单细胞profile与SpatialExperiment接入 | +| `decontx.mbt` | Bioconductor decontX | cluster-native/contaminant多项式混合、Beta/Dirichlet先验EM、background、自动聚类、计数分解与SCE输出 | +| `celda.mbt` | Bioconductor celda | `celda_CG`分层Dirichlet-multinomial、细胞群/基因模块联合推断、collapsed likelihood、EM/Gibbs、多链、K/L选择、预测与SCE输出 | | `summarized_experiment.mbt` | Bioconductor SummarizedExperiment | 多维数据容器 | +| `ranged_summarized_experiment.mbt` | Bioconductor RangedSummarizedExperiment | GRanges/GRangesList行范围、链特异重叠/最近邻、覆盖度、区间变换与协调子集 | +| `tree_summarized_experiment.mbt` | Bioconductor TreeSummarizedExperiment | 行/列树与数据链接、按节点子集、层级聚合 | | `ballgown.mbt` | Bioconductor ballgown | 转录组水平差异表达 | | `ruvseq.mbt` | Bioconductor RUVSeq | RNA-seq 批次效应去除 | | `sva.mbt` | Bioconductor sva | 替代变量分析与 ComBat | | `single_cell.mbt` | Bioconductor SingleCellExperiment | 单细胞数据分析 | | `csaw.mbt` | Bioconductor csaw | ChIP-seq窗口差异分析、TMM归一化、负二项GLM检验 | | `slingshot.mbt` | Bioconductor slingshot | 单细胞轨迹推断、MST构建、主曲线拟合、拟时间计算 | +| `slingshot_advanced.mbt` | Bioconductor slingshot 2.21.0 | soft membership、协方差缩放距离、约束MST/omega forest、root-to-leaf lineage、同时主曲线、Optional pseudotime、预测与SCE写回 | | `scnorm.mbt` | Bioconductor SCnorm | 单细胞RNA-seq归一化、分位数回归、深度依赖偏差校正 | | `edaseq.mbt` | Bioconductor EDASeq | RNA-seq探索性分析、GC含量归一化、基因长度校正 | | `glm_gampoi.mbt` | Bioconductor glmGamPoi | Gamma-Poisson GLM、size factors、伪批量聚合、IWLCS拟合、Wald检验、BH-FDR校正 | @@ -1511,6 +1861,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | MoonBit 文件 | 对应 Python 库 | 核心功能 | | :--- | :--- | :--- | +| `sparse_array.mbt` | Bioconductor SparseArray | N维规范化COO、重复坐标合并与零消除、R列主序转换、稀疏切片/置换/绑定、算术/统计、矩阵乘法 | | `matrix_generics.mbt` | Bioconductor MatrixGenerics | rowMeans/colMeans、rowSums/colSums、rowVars/colVars、rowSds/colSds、rowMedians/colMedians、rowMins/colMins、rowMaxs/colMaxs、rowRanges/colRanges、rowMad/colMad、rowCounts、rowAnys/colAnys、rowAlls/colAlls、块处理 | | `beachmat.mbt` | Bioconductor beachmat | 列/行块处理、线性迭代器(BmatIterator)、子集/转置/绑定、逐元素操作、类型安全矩阵访问API | | `survival.mbt` | R/Bioconductor survival | Kaplan-Meier估计器(Greenwood标准误)、log-rank检验、Cox比例风险模型(Newton-Raphson偏似然拟合) | @@ -1533,12 +1884,15 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | MoonBit 文件 | 对应 Python 库 | 核心功能 | | :--- | :--- | :--- | -| `blast.mbt` | BioPython `Bio.Blast` | BLAST 结果解析 | +| `blast.mbt` | BioSeqs compatibility API | 历史BLAST tabular/XML基础解析 | +| `blast_xml_advanced.mbt` | BioPython `Bio.Blast` | BLAST XML1/XML2严格解析、八类程序坐标语义、多report与规范写回 | | `search_io.mbt` | BioPython `Bio.SearchIO` | 统一搜索结果模型 | | `kegg.mbt` | BioPython `Bio.KEGG` | KEGG 数据库解析 | | `medline.mbt` | BioPython `Bio.Medline` | Medline/PubMed 解析 | | `entrez.mbt` | BioPython `Bio.Entrez` | NCBI 数据库访问 | | `swissprot.mbt` | BioPython `Bio.SwissProt` | UniProt 记录解析 | +| `cellosaurus.mbt` | BioPython `Bio.ExPASy.cellosaurus` | Cellosaurus记录解析、交叉引用查询与平面文本序列化 | +| `unigene.mbt` | BioPython `Bio.UniGene` | NCBI UniGene固定宽度记录解析、类型化子记录查询、SCOUNT校验与序列化往返 | | `uniprot_io.mbt` | BioPython `Bio.SeqIO.UniprotIO` | UniProt XML 格式解析 | | `chem_utils.mbt` | BioPython `Bio.PDB.chem_utils` | 化学计算工具(键长、键角、二面角、分子式量) | | `jaspar.mbt` | BioPython `Bio.motifs.Jaspar` | JASPAR PFM 格式解析与模体分析 | @@ -1572,6 +1926,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `transfac.mbt` | BioPython `Bio.Motifs.Transfac` | TRANSFAC转录因子结合谱解析(两字母字段码AC/ID/DE/NA/OS/BF/CC/P0/XX//、位置频率矩阵PFM、参考文献RN/RA/RT/RL/RX、共识序列、频率计算、字母索引、序列化往返) | | `hmmer_io.mbt` | BioPython `Bio.SearchIO.HmmerIO` | HMMER3输出解析(domtblout域表23列格式、人类可读文本格式、Query/Hit/HSP/HSPFragment聚合、多域比对、i-Evalue/c-Evalue、bitscore、条件E值、Query/Domain/Alignment段标记) | | `fasta_search_io.mbt` | BioPython `Bio.SearchIO.FastaIO` | FASTA搜索输出解析(-m8紧凑表格12列、-m9带#注释头、程序/版本/数据库元数据提取、Query/Hit/HSP聚合、正负链判定、E-value/bitscore) | +| `infernal_io.mbt` | BioPython `Bio.SearchIO.InfernalIO` | Infernal cmscan/cmsearch解析(tabular格式1/2/3自动检测、non-verbose文本与--noali、CM/HMM-only、local-end多片段、正负链坐标、过滤与SearchIO转换) | | `gene_pop.mbt` | BioPython `Bio.PopGen.GenePop` | GenePop群体遗传学(基因型diploid/haploid解析、2/3位等位基因数字检测、Pop人口分隔、Locus名自动生成/#Loci注释、等位基因频率、观察/期望杂合度、序列化往返) | | `stage_r.mbt` | Bioconductor stageR | 两阶段假设检验(筛选阶段BH-FDR、确认阶段Holm步降、Simes聚合、OFDR控制、Dte/Dtu方法、确认p值G/R重缩放) | | `enriched_heatmap.mbt` | Bioconductor EnrichedHeatmap | 基因组信号归一化(目标区域窗口化、四种均值模式absolute/weighted/w0/coverage、行平滑、百分位裁剪、负链窗口反转) | @@ -1612,13 +1967,24 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `geoquery.mbt` | Bioconductor GEOquery | GEO数据库数据获取、Series Matrix解析 | | `tximport.mbt` | Bioconductor tximport | 转录本量化数据导入、基因级别汇总 | | `single_cell_experiment.mbt` | Bioconductor SingleCellExperiment | 单细胞核心容器 (多assay、PCA/tSNE/UMAP降维、size factors) | +| `scuttle.mbt` | Bioconductor scuttle 1.23.1 | batch-aware MAD异常值、subset per-feature QC、feature-set/ID聚合、精确count downsampling、batch/block coverage等化与不可变SCE包装 | +| `bluster.mbt` | Bioconductor bluster 1.23.0 | observation × variable聚类、K-means++多起点、精确KNN、rank/number/Jaccard SNN、Louvain风格优化、two-step、诊断与SingleCellExperiment接入 | +| `flowsom.mbt` | Bioconductor FlowSOM 2.21.0 | cell × marker拓扑SOM、KWSP/PCA、分阶段MST、meta-clustering、节点统计、outlier、FlowFrame与SingleCellExperiment接入 | +| `mast.mbt` | Bioconductor MAST兼容层 | 检测比例与阳性表达的轻量双组分检验、BH-FDR及旧版结果API | +| `mast_advanced.mbt` | Bioconductor MAST 1.39.0 | 任意设计矩阵、Cauchy稳定化logistic/Gaussian hurdle GLM、嵌套LRT、eBayes方差收缩、边际logFC与SingleCellExperiment接入 | +| `single_r.mbt` | Bioconductor SingleR兼容层 | 参考profile、Spearman/Pearson相关性、旧版fine-tuning和结果API | +| `single_r_advanced.mbt` | Bioconductor SingleR 2.15.2 | 成对classic markers、标签内相关分位数、迭代fine-tuning、MAD剪枝、cluster与多参考注释及SingleCellExperiment接入 | +| `droplet_utils.mbt` | Bioconductor DropletUtils兼容层 | 旧版液滴统计、barcode排序、简化emptyDrops与细胞过滤API | +| `droplet_utils_advanced.mbt` | Bioconductor DropletUtils 1.33.0 | Simple Good-Turing ambient profile、curve-tracing knee/inflection、multinomial/Dirichlet-multinomial、alpha MLE、Monte Carlo、BH-FDR与SingleCellExperiment接入 | | `complex_heatmap.mbt` | Bioconductor ComplexHeatmap | 复杂热图可视化 (行/列聚类、颜色映射、热图注释) | | `pheatmap.mbt` | Bioconductor pheatmap | 增强型热图 (层次聚类、距离矩阵、行/列注释、颜色方案) | | `gsva.mbt` | Bioconductor GSVA | 基因集变异分析 (ssGSEA/zscore/PLAGE评分) | | `chromvar.mbt` | Bioconductor chromVAR | 染色质变异分析 (TF motif富集、GC偏差校正) | | `delayed_array.mbt` | Bioconductor DelayedArray | 延迟计算数组 (懒加载操作、分块处理、行/列聚合) | +| `sparse_array.mbt` | Bioconductor SparseArray | N维稀疏数组、规范化COO存储、切片/置换/绑定、稀疏算术、统计与矩阵乘法 | | `annotation_filter.mbt` | Bioconductor AnnotationFilter | 基因注释过滤 (染色体筛选、生物类型过滤、区域重叠检测) | -| `sc_dbl_finder.mbt` | Bioconductor scDblFinder | 单细胞双细胞检测 (Doublet评分、最近邻搜索、PCA降维) | +| `sc_dbl_finder.mbt` | Bioconductor scDblFinder兼容层 | 旧版Doublet评分、最近邻搜索、PCA降维与细胞过滤API | +| `sc_dbl_finder_advanced.mbt` | Bioconductor scDblFinder 1.27.6 | 人工doublet、共同归一化/PCA、精确kNN与cxds特征、迭代分类、分层阈值、来源富集及SingleCellExperiment接入 | | `seurat.mbt` | Bioconductor Seurat | 单细胞数据分析核心 (标准化、高可变基因、PCA、聚类、UMAP、差异表达、跨样本整合) | | `chipseeker.mbt` | Bioconductor ChIPseeker | ChIP-seq峰值注释 (基因组区域分类(启动子/外显子/内含子/UTR/基因间区)、距离TSS分布、BED格式读取、peak2gene关联分析、注释可视化) | | `topgo.mbt` | Bioconductor topGO | 拓扑GO富集分析 (TopGOTerm/TopGOGraph/TopGOEnrichmentResult数据结构、elim算法、weight01算法、Fisher精确检验、GO图构建) | @@ -1641,6 +2007,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `sff_io.mbt` | Biopython `Bio.SeqIO.SffIO` | SFF二进制格式解析 (SffHeader/SffRead/SffFile数据结构、大端字节序u16/u32/u64读写、二进制编码/解码、质量修剪、均值质量、按名称查找) | | `csaw.mbt` | Bioconductor csaw | ChIP-seq窗口差异分析 (CswWindow/CswDataSet/CswNormResult/CswResult数据结构、滑动窗口计数、TMM归一化、窗口过滤、负二项GLM检验、BH-FDR校正、差异区域检测) | | `slingshot.mbt` | Bioconductor slingshot | 单细胞轨迹推断 (SlingshotNode/SlingshotEdge/SlingshotCurve/SlingshotResult数据结构、MST构建、主曲线拟合、拟时间计算、分支检测) | +| `slingshot_advanced.mbt` | Bioconductor slingshot 2.21.0 | hard/soft cluster输入、weighted center/covariance、三类cluster距离、start/end约束Kruskal forest、自动omega、component root与lineage、同时主曲线、rank reweight/reassign、分支/预测及SingleCellExperiment接入 | | `scnorm.mbt` | Bioconductor SCnorm | 单细胞RNA-seq归一化 (SCnormQuantFit/SCnormGeneNormResult/SCnormResult数据结构、分位数回归、深度依赖偏差校正、基因特异性归一化) | | `edaseq.mbt` | Bioconductor EDASeq | RNA-seq探索性分析 (EDASeqGeneAnno/EDASeqDataSet/EDASeqWithinResult/EDASeqBetweenResult数据结构、GC含量归一化、基因长度Loess校正、样本间归一化、RPKM计算) | | `maftools.mbt` | Bioconductor maftools | 癌症基因组学MAF分析 (MAFMutation/MAFData/MutationSpectrum/TMBResult数据结构、SNV/Indel分类、TMB计算、突变谱分析、共现分析、Oncoplot数据生成、MAF文件解析) | @@ -1650,8 +2017,16 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `rtsne.mbt` | Bioconductor Rtsne | t-SNE降维算法 (TsneConfig/TsneResult数据结构、成对距离计算、条件概率估计与perplexity优化、联合概率矩阵构建、梯度下降优化、动量/early exaggeration调度) | | `uwot.mbt` | Bioconductor uwot | UMAP降维算法 (UmapConfig/UmapResult数据结构、k近邻搜索、模糊单纯集构建、局部模糊集并集、SGD低维嵌入优化、负采样、min_dist/spread参数控制) | | `tradeseq.mbt` | Bioconductor tradeSeq | 轨迹差异表达分析 (TrajectoryPoint/GeneExpressionData/GAMFit/DifferentialExpressionResult数据结构、GAM广义可加模型拟合、样条基函数、条件效应检验、BH-FDR校正) | -| `maf.mbt` | BioPython `Bio.Align` | MAF多序列比对格式解析与分析 (块处理、百分比一致性、统计分析、选择/过滤) | -| `mauve.mbt` | BioPython `Bio.Align` | Mauve基因组比对格式解析 (LCB检测、倒位/断点、覆盖率、BED导出) | +| `maf.mbt` | BioPython `Bio.Align` | 早期宽松MAF块解析与分析 (块处理、百分比一致性、统计分析、选择/过滤) | +| `align_maf.mbt` | BioPython `Bio.Align.maf` | 现代MAF文档、严格a/s/i/e/q、绝对坐标路径、reference index、多外显子拼接与规范写回 | +| `align_bed.mbt` | BioPython `Bio.Align.bed` | 现代BED pairwise alignment、正负链transcript路径、exon block、坐标映射、搜索、统计与BED3-BED12写回 | +| `mauve.mbt` | BioSeqs compatibility API | 历史MAF-like块解析与LCB重排分析(倒位/断点、覆盖率、BED导出;不解析现代XMFA) | +| `align_mauve.mbt` | BioPython `Bio.Align.mauve` | 现代XMFA文档、严格header/LCB、双向坐标路径、索引、跨序列投影、统计、重建与规范写回 | +| `align_clustal.mbt` | BioPython `Bio.Align.clustal` | 现代CLUSTAL alignment、严格block顺序和累计计数、列注释、坐标映射、统计与规范写回 | +| `align_phylip.mbt` | BioPython `Bio.Align.phylip` | 现代PHYLIP alignment、wrapped sequential/interleaved blocks、10列名称规范化、坐标映射、统计与规范写回 | +| `exonerate_text.mbt` | BioPython `Bio.SearchIO.ExonerateIO.exonerate_text` | Exonerate C4固定列alignment text、层次聚合、wrapped blocks、intron/NER/split codon/frameshift及坐标投影 | +| `blast.mbt` | BioSeqs compatibility API | 历史BLAST tabular/XML标签解析、HSP过滤与最佳匹配 | +| `blast_xml_advanced.mbt` | BioPython `Bio.Blast` | XML1/XML2 document/record/hit/HSP模型、严格结构与数值校验、1:1/3:1坐标、查询分析及canonical writer | | `stockholm.mbt` | BioPython `Bio.Stockholm` | Stockholm格式解析与分析 (Pfam/Rfam格式、二级结构、保守性、FASTA转换) | | `popgen_advanced.mbt` | BioPython `Bio.PopGen` | 高级群体遗传学统计 (Tajima's D、Fu & Li's D/F、MK检验、等位基因频率谱) | | `codon_advanced.mbt` | BioPython `Bio.SeqUtils.CodonUsage` | 高级密码子分析 (CAI、RSCU、ENC、GC3、物种特异性参考表) | @@ -1691,6 +2066,55 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `karyoploter.mbt` | `karyoploteR` | 核型可视化(染色体轨道、数据点、ASCII 渲染) | | `system_piper.mbt` | `SystemPipeR` | 流水线编排(步骤管理、依赖关系、进度追踪) | | `muscat.mbt` | `muscat` | 单细胞差异状态分析(伪批量聚合、DS 检验、QC) | +| `muscat_advanced.mbt` | `muscat` 1.27.4 | gene × cell合同、五类cluster-sample伪批量、任意设计/contrast、NB-IRLS与dispersion收缩、DS/DD、stagewise及不可变SCE写回 | +| `droplet_utils_advanced.mbt` | `DropletUtils` | feature × barcode严格计数、Simple Good-Turing、barcode-rank曲线追踪、multinomial/Dirichlet-multinomial及alpha MLE、确定性Monte Carlo、Phipson–Smyth校正、BH-FDR和不可变SCE写回 | +| `scrapper.mbt` | `scrapper` | 批次感知RNA QC、大小因子清洗/居中、count与log归一化、LOWESS方差趋势、HVG选择、多因子pseudo-bulk、SingleCellExperiment不可变包装 | +| `scuttle.mbt` | `scuttle` | R风格batch median/MAD及阈值共享、cell-subset feature QC、任意重叠feature-set聚合、精确无放回downsampling、batch/block coverage等化与SingleCellExperiment不可变包装 | +| `bluster.mbt` | `bluster` | K-means++与多起点Lloyd、精确KNN、rank/number/Jaccard SNN、seed可复现的Louvain风格优化、two-step聚类、Rand/ARI与cluster diagnostics、bootstrap稳定性及SingleCellExperiment写回 | +| `flowsom.mbt` | `FlowSOM` | Manhattan/Euclidean/Chebyshev/cosine距离、random/KWSP/PCA码本、规则网格与MST拓扑训练、K-means meta-clustering/elbow、节点MFI/CV/SD/MAD、purity/F-measure、MAD outlier及不可变SCE写回 | +| `slingshot_advanced.mbt` | `slingshot` | soft cluster membership、weighted covariance/Mahalanobis距离、受约束Kruskal forest、固定/自动omega、同时主曲线、cosine共享前缀收缩、rank重加权/重分配、新数据投影及不可变SCE写回 | +| `decontx.mbt` | `decontX` | 每细胞native/contaminant Bayesian mixture、确定性EM、empty-droplet ambient profile、自动k-means、诊断与SingleCellExperiment不可变包装 | +| `celda.mbt` | `celda` | `celda_CG`细胞群/基因模块联合聚类、四层Dirichlet-multinomial、collapsed EM/Gibbs、多链诊断、K/L网格选择、预测与SingleCellExperiment不可变包装 | +| `milo.mbt` | `miloR` | 精确KNN图、median精炼重叠邻域、邻域计数/表达、固定效应NB-GLM/Wald检验、graph spatial FDR与SCE接入 | +| `zinbwave.mbt` | `zinbwave` | ZINB交替EM/IRLS、cell/gene design与offset、确定性低维因子、gene dispersion shrinkage、observational weights、deviance residual及SingleCellExperiment包装 | +| `apeglm.mbt` | `apeglm` | 负二项GLM MLE、自适应Cauchy/Student-t先验、阻尼Newton多起点MAP、Laplace后验SD/区间、FSR/FSOS/s-value、DESeq2与SummarizedExperiment包装 | +| `aldex2.mbt` | `ALDEx2` | count+prior Dirichlet Monte Carlo、all/median/IQLR/zero/LVHA/user分母、两组与配对检验、posterior expected BH、effect/overlap、距离和SummarizedExperiment包装 | +| `dirichlet_multinomial.mbt` | `DirichletMultinomial` | sample×taxon DMM概率、soft k-means初始化、log-alpha BFGS/EM、Gamma prior、Hessian区间、Laplace/AIC/BIC、dmngroup分类、分层CV、ROC与SummarizedExperiment入口 | +| `variance_partition.mbt` | `variancePartition` | typed fixed/random design、多随机截距LMM、ML/REML方差分解、BLUP、weighted dream contrast、Satterthwaite自由度与SummarizedExperiment入口 | +| `dreamlet.mbt` | `dreamlet` | SCE到sample×cluster pseudobulk、TMM/logCPM、两阶段voom precision weights、固定/随机效应筛选、逐cell-type dream和study-wide FDR | +| `nnsvg.mbt` | `nnSVG` | AMMD/坐标和排序前驱kNN、指数协方差NNGP、covariate GLS、空间方差比例与length scale优化、LR/p-value/BH-FDR、过滤和SpatialExperiment rowData输出 | +| `banksy.mbt` | `Banksy` | kNN/radius邻域核、H0/H1+空间harmonic、lambda加权BANKSY矩阵、global/group scaling、Gram-Jacobi PCA、确定性多起点聚类、平滑与SpatialExperiment输出 | +| `voyager.mbt` | `Voyager` | kNN/distance-band/inverse-distance权重(W/B/C/S)、全局Moran's I与Geary's c(Cliff-Ord随机化方差+正态p)、局部Moran's I(象限+置换推断)、局部Geary's c、Getis-Ord Gi/Gi*(Ord-Getis z)、Lee's L(全局+局部)、多元局部Geary、经验变差函数+spherical/exponential/gaussian拟合、Moran correlogram、splitmix64置换、BH-FDR与SpatialExperiment不可变写回 | +| `spicyr.mbt` | `spicyR` | ordered cell-type-pair cross-L、矩形窗口disc-intersection边界校正、图像级localization统计、precision weights、加权LMM、条件对比、BH-FDR与SpatialExperiment metadata输出 | +| `lisaclust.mbt` | `lisaClust` | 多细胞类型local-K/L特征、KDE intensity correction、矩形/凸包窗口、disc-window边界积分、确定性多起点k-means、regionMap observed/expected富集与SpatialExperiment region输出 | +| `spatialdecon.mbt` | `SpatialDecon` | background-aware weighted log-normal non-negative regression、algorithm2异常点重拟合、Hessian协方差、abundance/count scaling、cell-type collapse、reverseDecon、GeoMx background、profile构建与SpatialExperiment输出 | +| `shared_reference_alignment.mbt` | `Bio.Align.Alignment` | 同参考PWA/MSA的reference-boundary insertion同步、原query投影、局部坐标、metadata、统计与格式转换 | +| `alignment_map.mbt` | `Bio.Align.Alignment.map/mapall` | 两层alignment坐标组合、正反链与gap传播、PSL、批量映射及protein/nucleotide MSA投影 | +| `alignment_counts.mbt` | `Bio.Align.Alignment.counts` | pairwise/MSA gap事件和composition汇总、正反链、wildcard、替换矩阵及完整affine评分 | +| `align_tabular.mbt` | `Bio.Align.tabular` | BLAST/FASTA query block、完整字段词汇、BTOP/aln_code路径、translated轴换算、过滤与coordinate alignment转换 | +| `align_psl.mbt` | `Bio.Align.psl` | PSL/PSLX类型模型、严格字段与block一致性、核酸反向query、translated反向target、3:1 codon映射、序列重计数与往返 | +| `align_sam.mbt` | `Bio.Align.sam` | header/reference与alignment类型模型、CIGAR路径、反向链与clipping、typed tags、PHRED、MD/NM、规范读写和坐标查询 | +| `a2m.mbt` | `Bio.Align.a2m` | match/insertion状态模型、canonical字符编码、wrapped/CRLF解析、行/列坐标互映、插入槽、统计、共识、切片和match-only投影 | +| `align_emboss.mbt` | `Bio.Align.emboss` | EMBOSS文件与alignment类型模型、固定21列body、多alignment/多序列、纯gap block、链感知绝对坐标、consensus统计和canonical writer | +| `align_exonerate.mbt` | `Bio.Align.exonerate` | Exonerate document/alignment/operation模型、cigar/vulgar、正反链/protein坐标、3:1 translated mapping、双向查询、统计和canonical writer | +| `exonerate_text.mbt` | `Bio.SearchIO.ExonerateIO.exonerate_text` | C4 document/query/hit/HSP/fragment模型、固定列3/4/5行parser、剪接/翻译/phase/frame、区间查询与防御复制 | +| `msf.mbt` | `Bio.Align.msf` | GCG MSF metadata/sequence/alignment模型、标准checksum、interleaved parser、短行补齐、coordinate path、pair counts、consensus和writer | +| `align_nexus.mbt` | `Bio.Align.nexus` | 独立alignment metadata/sequence模型、nested comment与quoted token lexer、sequential/interleaved parser、MATCHCHAR、all-gap压缩、坐标查询、统计和writer | +| `align_stockholm.mbt` | `Bio.Align.stockholm` | 有序GF/GS/GR/GC与reference模型、多记录parser、M/D/I列语义、all-gap压缩、annotation-aware切片、坐标查询、统计和writer | +| `align_chain.mbt` | `Bio.Align.chain` | 不可变alignment/block/counts模型、严格多记录parser、链感知absolute coordinate path、block重建、双向位置/区间映射、pairs、invert和writer | +| `align_maf.mbt` | `Bio.Align.maf` | 不可变document/track/index/spliced模型、严格a/s/i/e/q parser、链感知absolute path、任意component映射、区间查询、拼接和writer | +| `align_mauve.mbt` | `Bio.Align.mauve` | XMFA document/source/block/row模型、严格metadata与wrapped-row parser、正负链boundary path、pair counts、half-open index、跨序列投影、重建和writer | +| `align_clustal.mbt` | `Bio.Align.clustal` | metadata/sequence/alignment/counts模型、六类generator header、严格interleaved parser、consensus、coordinate path、投影、统计和writer | +| `align_phylip.mbt` | `Bio.Align.phylip` | sequence/alignment/counts模型、严格sequential/interleaved parser、固定10列名称、coordinate path、投影、统计和writer | +| `align_bed.mbt` | `Bio.Align.bed` | 不可变alignment/document/score/block模型、BED3-BED12 parser、双轴路径校验、block投影、双向residue mapping、search、summary和writer | +| `bigbed.mbt` | `Bio.Align.bigbed` | BigBed v4二进制读写、BED/AutoSQL、平衡B+ tree/R-tree、完整DEFLATE块解码、区间/名称查询与BED导出 | +| `bigmaf.mbt` | `Bio.Align.bigmaf` | bedMaf AutoSQL与MAF块往返、a/s/i/e/q注释、链感知坐标映射、BigBed索引查询和普通MAF导出 | +| `bigpsl.mbt` | `Bio.Align.bigpsl` | 标准bigPsl AutoSQL、核酸与translated DNA-protein坐标、链感知映射、match分类、BigBed索引和PSL导出 | +| `paml.mbt` | BioSeqs compatibility API | 内置Nei-Gojobori风格dN/dS和Jukes-Cantor近似,保留历史调用兼容性 | +| `paml_codeml.mbt` | `Bio.Phylo.PAML.codeml` | 类型化control/CODONML/AAML模型、NSsites/branch-site/clade/free-ratio、pairwise/距离矩阵、多基因、BEB/NEB和模型比较 | +| `paml_baseml.mbt` | `Bio.Phylo.PAML.baseml` | 类型化control与BASEML 4.x结果、JC69-UNRESTu、参数/SE、kappa/Q矩阵、nparK/auto-dGamma、nhomo节点和模型比较 | +| `paml_yn00.mbt` | `Bio.Phylo.PAML.yn00` | 类型化control与PAML 4.1-4.9i结果、NG86/YN00/LWL85/LWL85m/LPB93、矩阵和统计汇总 | +| `blast_xml_advanced.mbt` | `Bio.Blast` | XML1/XML2类型化文档、多query/report、HitDescr/taxonomy、参数/统计、strand/frame/translated path、过滤与同/跨dialect写回 | | `infercnv.mbt` | `infercnv` | 单细胞拷贝数变异推断(基因组位置排序、参考细胞有界 LFC 计算、金字塔权重平滑、CNV 分数与恶性细胞预测) | | `scenic.mbt` | `SCENIC` | 单细胞调控网络推断与聚类(TF-target 共表达模块、Regulon 构建、AUCell 活性评分、二值化阈值、细胞状态与主控调控因子) | | `cibersort.mbt` | `CIBERSORT` | 免疫细胞去卷积(NNLS 求解、LM22 风格特征矩阵、Pearson 拟合优度、分数归一化) | @@ -1716,6 +2140,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `velociraptor.mbt` | `velociraptor` | 单细胞RNA velocity分析(稳态线性回归gamma/beta比、EM动力学模型alpha/beta/gamma估计、velocity向量计算、KNN加权嵌入投影、根细胞识别、转移矩阵) | | `compass.mbt` | `Bio.Compass` | COMPASS profile-profile比对输出解析(版本提取、多记录解析SW分数/E值/百分比一致性/比对序列/共识行、E值与一致性过滤、比对长度统计、摘要生成) | | `exonerate.mbt` | `Bio.SearchIO.ExonerateIO` | Exonerate比对输出解析(vulgar格式三元组比对块解析M/I/5/3/S/G/U/V、cigar格式解析、格式自动检测、分数过滤、内含子统计、vulgar/cigar字符串重建) | +| `exonerate_text.mbt` | `Bio.SearchIO.ExonerateIO.exonerate_text` | Exonerate C4人类可读文本严格解析(多query/hit/HSP聚合、wrapped body、intron/NER/split codon/frameshift、三字母氨基酸和链感知坐标) | | `mmcifio.mbt` | `Bio.PDB.mmcifio` | mmCIF文件写入(Structure对象序列化data block/header/atom_site loop、20列原子坐标格式化、HETATM支持、值转义、round-trip验证) | | `interproscan.mbt` | `Bio.SearchIO.InterproscanIO` | InterProScan输出解析(TSV 14列格式解析蛋白质ID/MD5/长度/分析数据库/签名/位置/分数/IPR/GO、按数据库/蛋白质过滤、GO条目提取、按蛋白质分组、摘要) | | `sasa.mbt` | `Bio.PDB.SASA` | 溶剂可及表面积计算(Shrake-Rupley滚动球算法、Fibonacci球面采样、范德华半径查表H/C/N/O/S/P、逐原子/残基SASA、骨架/侧链拆分、总量统计) | @@ -1738,6 +2163,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | `transfac.mbt` | Biopython `Bio.Motifs.Transfac` | TRANSFAC转录因子结合谱解析(两字母字段码AC/ID/DE/NA/OS/BF/CC/P0/XX//、位置频率矩阵PFM、参考文献RN/RA/RT/RL/RX、共识序列、频率计算、字母索引、序列化往返) | | `hmmer_io.mbt` | Biopython `Bio.SearchIO.HmmerIO` | HMMER3输出解析(domtblout域表23列格式、人类可读文本格式、Query/Hit/HSP/HSPFragment聚合、多域比对、i-Evalue/c-Evalue、bitscore、条件E值、Query/Domain/Alignment段标记) | | `fasta_search_io.mbt` | Biopython `Bio.SearchIO.FastaIO` | FASTA搜索输出解析(-m8紧凑表格12列、-m9带#注释头、程序/版本/数据库元数据提取、Query/Hit/HSP聚合、正负链判定、E-value/bitscore) | +| `infernal_io.mbt` | Biopython `Bio.SearchIO.InfernalIO` | Infernal cmscan/cmsearch解析(tabular格式1/2/3自动检测、non-verbose文本与--noali、CM/HMM-only、local-end多片段、正负链坐标、过滤与SearchIO转换) | | `gene_pop.mbt` | Biopython `Bio.PopGen.GenePop` | GenePop群体遗传学(基因型diploid/haploid解析、2/3位等位基因数字检测、Pop人口分隔、Locus名自动生成/#Loci注释、等位基因频率、观察/期望杂合度、序列化往返) | | `stage_r.mbt` | Bioconductor stageR | 两阶段假设检验(StageRMethod/StageRConfig/StageRResult数据结构、筛选阶段BH-FDR校正、确认阶段Holm步降程序、Simes聚合、OFDR控制、Dte/Dtu调整向量、确认p值G/R重缩放、显著性基因/假设提取) | | `enriched_heatmap.mbt` | Bioconductor EnrichedHeatmap | 基因组信号归一化(GenomicSignal/TargetRegion/MeanMode/EnrichedHeatmapConfig/NormalizedMatrix数据结构、目标区域窗口化、四种均值模式absolute/weighted/w0/coverage、行平滑、百分位裁剪、负链窗口反转、富集谱计算) | @@ -1799,7 +2225,7 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 ### 12. DESeq2 差异表达分析 (Bioconductor DESeq2) -实现完整的 RNA-seq 差异表达分析流程,支持从原始计数到差异表达基因筛选的全流程分析。可以创建 DESeqDataSet 对象管理计数矩阵、样本信息和设计矩阵。支持 size factors 估计(中位数比率法)进行测序深度校正,以及计数矩阵归一化和 log2 CPM 计算。支持分散度估计(parametric fit),结合经验贝叶斯收缩方法。支持负二项 GLM 拟合,通过迭代加权最小二乘法估计回归系数。支持 Wald 检验进行差异表达显著性检验,计算 log2 fold change、标准误、检验统计量和 p 值。支持 Benjamini-Hochberg 多重检验校正。支持 LFC 收缩(apeglm-like 方法),减小低表达基因的 fold change 估计偏差。支持显著基因筛选(按 adjusted p-value 和 LFC 阈值)和获取 top 差异表达基因。适用于 RNA-seq 差异表达分析。 +实现完整的 RNA-seq 差异表达分析流程,支持从原始计数到差异表达基因筛选的全流程分析。可以创建 DESeqDataSet 对象管理计数矩阵、样本信息和设计矩阵。支持 size factors 估计(中位数比率法)进行测序深度校正,以及计数矩阵归一化和 log2 CPM 计算。支持分散度估计(parametric fit),结合经验贝叶斯收缩方法。支持负二项 GLM 拟合,通过迭代加权最小二乘法估计回归系数。支持 Wald 检验进行差异表达显著性检验,计算 log2 fold change、标准误、检验统计量和 p 值。支持 Benjamini-Hochberg 多重检验校正。旧 `lfc_shrink` API 保留固定正态先验的轻量收缩行为;完整的自适应重尾 apeglm 实现在独立 `apeglm.mbt` 和第 254 节中。支持显著基因筛选(按 adjusted p-value 和 LFC 阈值)和获取 top 差异表达基因。适用于 RNA-seq 差异表达分析。 ### 13. Suffix Array & Suffix Tree (libdivsufsort) @@ -1833,9 +2259,9 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 提供统一的搜索结果模型,支持 HMMER3 tabular 格式和 BLAT PSL 格式的解析。可以获取查询 ID、命中数、top hits(按 E-value 排序)和 HSP 数量统计。支持 BLAST 结果转换为 SearchIO 模型,便于不同搜索工具结果的统一处理。 -### 21. BLAST 结果解析 (Bio.Blast) +### 21. BLAST 基础解析兼容层 -支持 BLAST tabular 和 XML 格式的解析,提供丰富的结果过滤和访问接口。可以按 E-value 和 identity 过滤 hits,获取最佳匹配和最佳 HSP。支持所有 HSPs 的获取和查询序列长度的访问。 +`blast.mbt`保留早期BLAST tabular和简单XML标签解析API,支持按E-value和identity过滤、最佳命中/HSP及基础汇总。完整的Biopython `Bio.Blast` XML1/XML2文档语义、严格校验、链向/translated坐标和规范写回由第288节的`blast_xml_advanced.mbt`提供。 ### 22. 替换矩阵 (Bio.SubsMat) @@ -2075,7 +2501,13 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 ### 81. DropletUtils 空液滴检测 (Bioconductor DropletUtils) -实现单细胞 RNA-seq 数据的空液滴检测功能,支持 barcode 排序、knee 点检测和 emptyDrops 算法。可以计算液滴统计指标(总计数、检测基因数),对液滴按总计数排序,找到 knee 点估计细胞数量。支持基于 Monte Carlo 模拟的空液滴检测,计算每个液滴为空的概率和 FDR 值,进行细胞过滤。 +保留 `droplet_utils.mbt` 的旧版液滴统计、barcode 排序和过滤 API;`droplet_utils_advanced.mbt` 对齐 Bioconductor `DropletUtils` 1.33.0 的核心 `emptyDrops` 工作流。`DropletCountMatrix` 固定采用 feature × barcode 方向,严格校验矩形、非负整数计数和唯一非空名称,并通过防御性复制保持输入不可变。 + +ambient profile 可由 `lower`、`by_rank` 或显式 known-empty mask 选择空液滴,聚合 feature counts 后使用 Simple Good-Turing count-of-counts 回归和平滑频率切换;对样本中存在但 ambient 计数为零的 feature,使用安全伪概率保护。barcode-rank 在 log10(rank)-log10(total) 曲线上按固定弧长窗口追踪,以平均 rank 处理 ties,并分别由最短上凸 chord 和最负 gradient 求 knee 与 inflection。 + +概率内核实现完整 multinomial 与 Dirichlet-multinomial 对数概率,后者可在 log-alpha 区间上通过 golden-section 最大似然估计 concentration。Monte Carlo 使用 Park–Miller RNG、Box–Muller 正态和 Marsaglia–Tsang Gamma 采样;相同 library size 共享递增抽样路径,尾部计数采用 Phipson–Smyth `(b+1)/(B+1)` 校正,并以 `Limited` 标记零极端次数。最终对普通 barcode 执行 BH-FDR,高计数 barcode 按显式阈值或自动 knee 无条件保留,未检验值通过 `Option` 保留上游 `NA` 语义。 + +`droplet_utils_empty_drops_sce` 从 `SingleCellExperiment` assay 读取计数,可解析 known-empty 列,并向副本写入 total、log-probability、p-value、Limited、FDR、class、retained、ambient profile 和运行元数据。可移植实现使用稠密矩阵和串行确定性 Monte Carlo;配置中 `alpha < 0` 表示普通 multinomial,`alpha > 0` 表示 Dirichlet-multinomial。上游的稀疏/延迟矩阵调度、BiocParallel 后端和磁盘支持不在当前范围内。 ### 82. scran 单细胞归一化与聚类 (Bioconductor scran) @@ -2095,7 +2527,19 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 ### 86. MAST 单细胞差异表达分析 (Bioconductor MAST) -实现单细胞差异表达分析功能,采用 Hurdle 模型(零膨胀模型)处理单细胞数据的零膨胀特性。模型包含两个组分:离散组分(Fisher 精确检验检测率差异)和连续组分(Welch t 检验表达水平差异)。支持使用卡叉分布合并两个 p 值得到联合检验结果,使用 Benjamini-Hochberg 方法进行 FDR 多重检验校正。支持计算 log2 倍数变化、检测率统计和结果汇总(差异基因计数、Top 基因列表)。适用于单细胞转录组差异表达分析。 +`mast.mbt` 保留原有检测比例、阳性表达 Welch 检验和结果汇总 API 作为兼容层。`mast_advanced.mbt` 参考 Bioconductor MAST 1.39.0 的 `zlm`、`bayesglm`、`lrTest`、`ebayes` 和 `logFC` 实现完整的可移植高级流程,输入统一采用 feature × cell 非负表达矩阵。`MastAdvancedData::create` 接受任意满秩设计矩阵,`mast_advanced_from_groups` 提供可指定 reference 的 treatment coding,并可自动追加每细胞检测率 `cngeneson`/CDR 协变量;构造阶段严格检查矩阵方向、矩形性、有限值、名称唯一性、设计维度和秩。 + +每个基因分别拟合检测事件的 logistic GLM 和仅使用阳性表达的 Gaussian GLM。离散组分使用 IRLS 与非截距系数 scale 2.5 的局部二次 Cauchy 稳定化,连续组分通过 Cholesky 正规方程求解。待检验设计列从完整模型中删除后重新拟合 reduced model,分别计算离散、连续 likelihood-ratio statistic,并将可检验组分的 statistic 与自由度相加形成 hurdle 检验;chi-square survival probability 由 regularized upper incomplete gamma 计算。不可拟合的全零、全阳性或阳性设计秩不足基因保留为 `Double?::None`,BH-FDR 只校正可检验条目。 + +经验贝叶斯层支持 MAST 默认 H0(按基因阳性表达中心化)和 H1(完整设计残差)两种 sufficient statistics,通过 inverse-gamma 边际似然估计先验方差和自由度,再对基因残差方差进行收缩。结果同时提供组分系数/标准误、收敛状态、LRT/FDR、检测率、moderated variance,以及 MAST 定义的离散概率 × 连续均值边际 logFC 和 delta-method 方差。`mast_advanced_zlm_sce` 从 `SingleCellExperiment` assay 和 group `col_data` 读取输入,在深复制的容器中写入检测率、组分/联合 p 值、FDR、logFC、分类、CDR 与运行元数据,不修改调用方对象。当前实现使用稠密数组和串行确定性求解,不包含上游并行后端、稀疏矩阵调度、混合效应模型或绘图接口。 + +### SingleR 2.15.2 参考驱动细胞类型注释 (Bioconductor SingleR) + +`single_r.mbt` 保留参考 profile、Spearman/Pearson 相关性和旧版结果 API 作为兼容层;`single_r_advanced.mbt` 对齐 Bioconductor SingleR 2.15.2 的训练、分类、fine-tuning、剪枝和多参考整合语义。参考矩阵固定为 gene × sample,测试矩阵固定为 gene × cell,构造阶段严格检查矩形性、有限值、名称唯一性和标签长度,并通过防御性复制保持输入不可变。训练按测试基因顺序对齐共享基因,保持标签首次出现顺序,对每个标签计算逐基因中位数,再按有向标签对的正中位数差选择 classic markers;自动 marker 数使用 `500 × (2/3)^log2(N)`。 + +分类对测试细胞和每个参考样本计算支持 ties 平均秩的 Spearman 相关,并以标签内相关分布的线性插值分位数作为 score,默认分位数为 0.8。fine-tuning 从距最高 score 不超过阈值的标签开始,重新合并候选标签间 markers、重算分位数并迭代收缩候选集合;结果同时保留初始 score、最终标签、`delta.next` 和相对所有标签中位数的 `delta.median`。剪枝按最终标签计算 `median(delta) - nmads × 1.4826 × MAD(delta)`,并支持硬 `delta.median`/`delta.next` 下限,未通过项使用 `Option::None` 保留上游 `NA` 语义。 + +高级入口还支持按 cluster 汇总基因表达后注释,以及对多个参考独立分类后,在所有参考和测试共享的 marker 空间中重算 score 并选择来源参考。`single_r_advanced_sce` 默认读取 `logcounts`,可选 cluster `col_data` 列,并向深复制的 `SingleCellExperiment` 写入 label、pruned label、score、delta、cluster 和运行元数据,不修改原容器。当前可移植实现使用稠密数组和串行相关计算,不包含上游 BiocNeighbors、DelayedArray、BiocParallel、HDF5 索引或 celldex 数据下载后端。 ### 87. GenomicFiles 分布式基因组文件处理 (Bioconductor GenomicFiles) @@ -2143,7 +2587,13 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 ### 98. scDblFinder 单细胞双细胞检测 (Bioconductor scDblFinder) -实现单细胞 RNA-seq 数据的双细胞(doublet)检测功能,支持 Doublet 评分计算和细胞过滤。可以创建 SingleCellData 对象(计数矩阵、细胞名称、基因名称)和 DoubletScore 对象(细胞名称、评分、是否为双细胞)。支持距离计算(scdf_compute_distance)、最近邻搜索(scdf_find_nearest_neighbors)、Doublet 评分计算(scdf_compute_doublet_score)、双细胞检测(scdf_detect_doublets)、结果汇总(scdf_doublet_summary)和细胞过滤(scdf_filter_doublets)。支持 PCA 降维(scdf_compute_pca)用于降维后距离计算。适用于单细胞数据的质量控制和双细胞去除。 +在保留 `sc_dbl_finder.mbt` 旧版 `SingleCellData`、`DoubletScore` 和 8 个兼容测试的基础上,`sc_dbl_finder_advanced.mbt` 实现 Bioconductor `scDblFinder` 1.27.6 的高级单细胞 RNA-seq doublet 检测流程。严格构造器要求矩形 cell × gene 非负有限计数、唯一非空标识和非零文库,并执行防御性复制;`ScDblFinderConfig` 对预期 doublet rate、人工样本数、特征数、维度、邻居数、迭代和分类器参数执行完整边界校验。 + +算法按 capture/sample 独立拟合。每批选择高方差基因,对真实和人工细胞共同进行 library normalization、log1p 转换和确定性 PCA;人工 doublet 可随机配对或按 cluster 跨群配对,并支持 half-size 缩放。精确 kNN 计算人工邻居比例、逆距离与 rank 加权比例、最近真实/人工距离、`nearestClass`、最可能来源和来源歧义度。额外计算 library size、检测基因数、`nAbove2` 及 cxds 风格互斥基因共表达 surprise 分数。 + +分类阶段使用确定性、类别平衡、L2 正则化 logistic gradient descent 代替上游 XGBoost,并在每轮重新训练前排除高疑似真实细胞和不可识别人工 doublet。阈值通过预期 doublet rate 偏差、假阳性率和人工 doublet 假阴性率的联合损失优化;cluster-aware 模式使用 `Σp_c²` 修正 homotypic 比例,并提供 Poisson 上尾与 BH-FDR 校正的 pairwise origin enrichment。结果支持按名称查询、top doublet 排序、摘要、singlet 过滤和自动快速聚类。 + +`sc_dbl_finder_single_cell_experiment` 接受 gene × cell assay,显式转置后拟合,并以不可变复制写回 `score`、`class`、邻居比例、加权比例、most likely origin、selected genes、PCA 和元数据。为控制 MoonBit 精确 kNN 的计算量,自动人工 doublet 数使用 `max(150, min(1500, 2 × nCells))`,不同于上游最少 1500 的默认策略;当前不复刻 XGBoost、BiocNeighbors 后端、Poisson resampling、meta-cell/triplet、scATAC Amulet、fragment overlap 和 feature aggregation。 ### 99. ChIPseeker ChIP-seq峰值注释 (Bioconductor ChIPseeker) @@ -2699,6 +3149,436 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 实现 CDAO(Comparative Data Analysis Ontology,比较数据分析本体)RDF/XML 格式的解析与序列化,参考 Biopython `Bio.Phylo.CDAO`。CDAO 是基于 RDF 的系统发育数据表示标准,使用 CDAO 本体术语将树结构编码为 RDF 三元组(subject-predicate-object),便于与语义网和本体推理系统互操作。核心 CDAO 本体术语:cdao:Tree(系统发育树)、cdao:Node(树节点)、cdao:Edge(树枝/边)、cdao:has_Root(树→根节点)、cdao:has_Child/has_Descendant(父→子节点)、cdao:has_Ancestor/has_Parent(子→父节点)、cdao:belongs_to_TU(节点→分类单元)、cdao:TU(分类单元/OTU/叶标签)、rdfs:label(标签文字)。核心数据结构:Cdaotree(id/rooted/root_node_id/name?);CdaoNode(id/children : Array[String]/parent_id?/tu_id?/branch_length?/label?,关键字段标记 mut 以便构建时修改);CdaoTU(id/label?);CdaoDocument(trees : Array[Cdaotree]/nodes : Map[String, CdaoNode]/tus : Map[String, CdaoTU] 完整 RDF 图)。核心函数:cdao_namespace()/cdao_rdf_namespace()/cdao_rdfs_namespace() 返回命名空间 URI;cdao_parse(xml) 主解析入口 → cdao_parse_rdf_xml 提取 CdaoTriple 三元组(手写 XML 解析器,处理标签/属性/rdf:about/rdf:resource/文本内容/自闭合/嵌套子元素)→ cdao_build_document 三元组分类填充 Document(rdf:type 创建节点/TU、has_Root 创建 Tree、has_Child 填 children、has_Ancestor 填 parent_id、belongs_to_TU 填 tu_id、rdfs:label 填 label、has_branch_length 填 branch_length);cdao_to_trees(doc) 递归 cdao_build_clade 将 CdaoDocument 转为 BioSeqs Tree 数组(TU 标签优先于节点标签);cdao_write(tree) 将 Tree 序列化为 RDF/XML 字符串(CdaoWriteState 管理 node_counter/tu_counter/tu_map,递归 cdao_write_clade 输出节点与边,末尾输出 TU 元素,cdao_escape_xml 处理 & < > 实体转义)。适用于系统发育数据语义网交换、本体推理、CDAO 兼容工具链互操作。 +### 236. TreeSummarizedExperiment 树结构实验容器 (Bioconductor TreeSummarizedExperiment) + +实现结合实验矩阵与层级树的 `TreeSummarizedExperiment` 容器,复用现有 `SummarizedExperiment`、`Tree` 和 `Clade` 类型。容器支持 `row_tree`/`col_tree`、`row_links`/`col_links` 和 `reference_sequences`,其中 `TseLink` 记录节点标签、稳定别名、一基节点编号、叶节点状态和树名称,对应 Bioconductor 的 `rowTree`、`rowLinks`、`colTree`、`colLinks` 与 `referenceSeq` 语义。`subset_rows`/`subset_cols` 同步裁剪 assay、链接和参考序列;`subset_by_row_nodes`/`subset_by_col_nodes` 可按内部节点或叶节点选择所有已链接后代,保留原树结构。`aggregate_rows`/`aggregate_cols` 对目标节点覆盖的数据执行 Sum、Mean、Min 或 Max 聚合,并为结果重建节点链接。`tse_find_descendants`、`tse_find_ancestors` 和 `tse_is_leaf` 提供树节点查询。`is_valid` 检查 assay 维度、链接长度、树存在性和参考序列长度。适用于微生物分类丰度、系统发育表达矩阵和具有样本层级的数据分析。 + +### 237. Cellosaurus 细胞系数据库解析 (Bio.ExPASy.cellosaurus) + +实现与 Biopython `Bio.ExPASy.cellosaurus` 对应的 Cellosaurus 平面文本解析。`cellosaurus_parse` 支持批量记录,`cellosaurus_read` 支持零或一条记录;解析器识别 `ID`、`AC`、`AS`、`SY`、`DR`、`RX`、`WW`、`CC`、`ST`、`DI`、`OX`、`HI`、`OI`、`SX`、`AG`、`CA` 和 `DT` 字段,兼容数据库头部、未知扩展字段与 CRLF。`CellosaurusRecord` 提供类型化记录,`CellosaurusCrossReference` 将 `DR` 字段拆分为数据库和登录号;辅助方法支持次级登录号/同义名拆分、按数据库筛选交叉引用和物种文本查询。`to_string` 可生成规范平面文本并支持解析-序列化往返。缺失 `//` 终止符、记录嵌套、非法 `DR` 或单记录读取到多条记录时抛出 `CellosaurusError`。 + +### 238. RangedSummarizedExperiment 基因组区间实验容器 (Bioconductor SummarizedExperiment) + +实现支持 `GRanges` 或 `GRangesList` assay 行范围的 `RangedSummarizedExperiment`,并复用现有 `SummarizedExperiment` 管理 assays、`col_data` 和 metadata。`new_with_range_groups` 和 `from_experiment_with_range_groups` 用复合范围表示转录本及其外显子等特征,外层元素与 assay 行严格平行;未显式提供 `row_data` 时会采用 `GRangesList.element_metadata`。容器内部使用每个复合特征的外接范围维持行维度,但 `find_overlaps`、`count_overlaps`、`overlaps_any`、`subset_by_overlaps`、`nearest`、`distance_to_nearest` 和 `coverage` 始终使用真实成员范围,因此不会将外显子之间的内含子误判为重叠。`subset_rows`/`subset_cols` 支持选择、重排和重复索引,并同步更新 assay、分组范围、行列注释与名称;`shift`、`narrow`、`resize`、`flank` 和 `promoters` 逐成员变换范围并保持实验数据不变。原有 `GRanges` 构造器与 API 保持兼容,`with_row_ranges` 可显式切回平面范围模式。 + +### 239. UniGene 基因聚类记录解析 (Bio.UniGene) + +实现与 Biopython `Bio.UniGene` 对应的 NCBI UniGene 固定宽度平面文件解析。`unigene_parse` 支持多记录输入,`unigene_read` 强制读取单条记录;解析器覆盖 `ID`、`TITLE`、`GENE`、`CYTOBAND`、`EXPRESS`、`RESTR_EXPR`、`GNM_TERMINUS`、`GENE_ID`、`LOCUSLINK`、`HOMOL`、`CHROMOSOME`、`PROTSIM`、`TXMAP`、`SCOUNT`、`SEQUENCE` 和 `STS` 标签。`UniGeneRecord` 以类型化数组保存序列、蛋白相似性、STS 和转录本映射子记录,支持按序列类型、登录号、相似物种和 IMAGE clone 查询,并保留未知子字段用于兼容扩展格式。解析器支持 CRLF,严格检查固定 12 列标签、记录终止符、布尔值和非负 `SCOUNT`,且要求声明数量与实际 `SEQUENCE` 数量一致。`to_string` 生成规范固定宽度文本并支持解析-序列化往返;格式错误抛出 `UniGeneError`。 + + +### 240. GRangesList 复合基因组特征 (Bioconductor GenomicRanges) + +实现 Bioconductor `GenomicRanges::GRangesList` 的复合特征语义,每个命名外层元素保存一组 `GRanges`,适合表示转录本-外显子、基因-调控区等一对多结构。构造器严格校验外层名称、元素元数据和内部范围维度;`granges_split_as_list` 按首次出现顺序分组,`granges_list_from_partition` 按元素长度分区,`unlist` 展平成单个 `GRanges`,`relist` 按原分区重建并保留名称和元数据。`subset`、`concat`、`parallel_concat` 支持外层选择和组合;`shift`、`narrow`、`resize`、`flank`、`promoters`、`reduce`、`disjoin` 和 `sort_ranges` 逐元素执行,`parallel_union`、`parallel_intersect` 和 `parallel_setdiff` 提供同位置元素间集合运算。`find_overlaps`、`count_overlaps`、`overlaps_any` 和分组对分组查询返回外层复合特征索引,并对同一特征的多个成员命中去重;`nearest`、`distance_to_nearest` 和 `coverage` 基于所有真实成员范围计算。`feature_bounds` 仅用于生成每个复合特征的协调外接范围,不替代精确区间计算。 + +### 241. HH-suite HHR profile-profile 比对解析 (Bio.Align.hhr) + +实现与 Biopython `Bio.Align.hhr` 对应的 HHsearch/HHblits HHR 文本解析。`hhr_parse` 读取查询元数据、命中摘要表和多块 profile-profile 比对,使用 Biopython 风格的 0-based、end-exclusive 坐标,并保留 query/target consensus、预测二级结构、DSSP、逐列分数和 confidence。`HhrRecord` 支持按 target 查询、probability/E-value 过滤和最佳命中选择;`HhrAlignment` 提供去 gap 序列、identity/coverage 统计、`query_to_target` 坐标映射和 `aligned_pairs`。解析器严格校验 rank、摘要与详情数量、坐标跨度、跨块连续性、比对宽度、`Aligned_cols` 和终止标记,兼容 CRLF、无空行的官方布局、零命中及 EOF 结束的完整末块。`to_string` 生成规范 HHR 文本并支持解析-序列化往返;格式错误抛出 `HhrError`。 + +### 242. N维稀疏数组基础设施 (Bioconductor SparseArray) + +实现与 Bioconductor `SparseArray` 核心语义对应的 N 维稀疏数组,作为现有二维 CSC/CSR `BiocMatrix` 的补充。`SparseArray::from_coo` 使用 0-based 坐标构建规范化 COO:构造时校验维度和坐标、深复制输入、按坐标排序、合并重复坐标,并删除合并后为零的条目;`from_flat`/`to_flat` 遵循 R 风格列主序(第一维变化最快),另提供矩阵和零数组构造器。查询 API 覆盖维度、长度、非零坐标/值、密度、坐标及线性随机访问和稠密转换,所有公开坐标访问器均返回防御性副本。 + +稀疏变换支持不可变单点/批量赋值、重复索引子集、0-based end-exclusive 切片、任意维度 `aperm`、二维转置和按指定维度绑定;加减、Hadamard 乘积、缩放和非零映射直接处理规范化非零条目。统计 API 包括全数组 sum/mean/min/max(极值正确纳入隐式零)以及二维 row/column sums、means 和非零计数;二维稀疏矩阵还支持 `matmul`、`crossprod` 和 `tcrossprod`。当前实现采用可移植的规范化 COO,并未宣称覆盖官方包的完整 SVT 存储后端。 + +### 243. CEAligner 组合扩展结构比对 (Biopython Bio.PDB.cealign) + +实现与 Biopython `Bio.PDB.cealign.CEAligner` 对应的组合扩展结构比对。`cealign_get_guide_atoms` 按模型、链和残基顺序提取引导原子,蛋白质优先使用 `CA`,缺失时回退到核酸 `C4'`;`CEAligner::set_reference` 保存不可变参考坐标,`align` 对移动结构建立分子内距离矩阵和 AFP 相似度矩阵,以 Biopython 阈值扩展严格单调的 CE 路径并保留最多 20 条候选。在最长候选路径中使用 `QCPSuperimposer` 选择最低 RMSD 叠合,默认窗口 8 时计算 CE 经验 Z-score,并可在显著路径上执行只接受 RMSD 降低的局部索引优化。 + +`CeAlignmentResult` 返回对齐索引、片段数、RMSD、Z-score、覆盖率、旋转矩阵和平移向量。`transform=true` 会将刚体变换应用到移动结构的全部原子,同时重建 `Structure` 以保证输入对象不被原地修改;`transform=false` 仅计算比对。公开辅助 API 还包括距离矩阵、片段相似度、路径搜索和独立结构变换,非法参数、引导原子缺失或结构长度不足时抛出 `CeAlignError`。 + +### 244. scrapper 单细胞预处理 (Bioconductor scrapper) + +实现 Bioconductor `scrapper` 核心单细胞 RNA-seq 预处理流程,矩阵统一采用 feature × cell 方向。RNA QC 计算每个细胞的文库总量、检测基因数和命名 feature subset 比例,并用 `log(value + 1)` 空间的 median/MAD 下限及 subset 比例上限执行批次感知过滤;block 顺序按首次出现保留。大小因子支持非法值清洗、文库大小估计、全局/逐批次居中,以及保留批次间尺度的最低批次居中。count scaling 和可配置底数、pseudo-count 的 log-normalization 均返回新矩阵,不修改输入。 + +基因建模提供均值、sample variance、quarter-root LOWESS 局部线性趋势、左侧向原点外推、残差方差和带 ties/bound 控制的 HVG 选择。`scrapper_aggregate_across_cells` 可按一个或多个分类因子的唯一组合生成 pseudo-bulk sums、detected counts、means 和 medians,并返回组组合及每个细胞的组索引。`scrapper_normalize_rna_counts_sce` 与 `scrapper_quick_rna_qc_sce` 深复制 assay 和主要注释后写入结果,避免修改原 `SingleCellExperiment`。当前实现是无需 libscran C++ 的可移植 MoonBit 版本,不保证 LOWESS 与上游后端位级一致,也不覆盖 `scrapper` 的全部导出接口。 + +### scuttle 1.23.1 单细胞基础工具 (Bioconductor scuttle) + +在避免复制 `scrapper.mbt` 已覆盖的细胞 QC、归一化、方差建模和 pseudo-bulk 功能的前提下,`scuttle.mbt` 对齐 Bioconductor `scuttle` 1.23.1 中仍缺失的复杂语义,所有矩阵统一采用 feature × cell 方向。`scuttle_is_outlier` 使用 R 的 `1.4826 × median absolute deviation`,支持 lower/higher/both tail、log2 空间、阈值估计 subset、逐 batch 估计、median/MAD 共享及缺失 batch 恢复;batch 名按因子式字典序输出,无法估计的阈值和状态分别由 validity 标记与 `Bool?::None` 保留。 + +per-feature QC 计算全体及命名 cell subset 的均值、严格检测百分比和 subset/global ratio。feature 聚合支持任意重叠集合、sum/average、检测计数及按字符串 ID 的字典序分组。downsampling 会先将数值 count 四舍五入并把负值截为零,再通过固定 seed 的 Park–Miller RNG 执行无放回抽样;逐列或全矩阵输出总量严格等于 `round(total × proportion)`。batch downsampling 可按 median、mean 或 `exp(mean(log(1+x)))` 汇总 coverage,并可在 block 内等化到最浅 batch。 + +`scuttle_per_feature_qc_sce`、`scuttle_aggregate_feature_sets_sce` 和 `scuttle_downsample_sce` 深复制 assay、row/column metadata、reduced dimensions、metadata 与 alternative experiments。QC 写入 row metadata,feature-set 聚合替换所选 assay 并清除失效的 gene-level metadata,downsampling 添加新 assay,均不修改输入容器。当前实现采用稠密数组和串行确定性抽样,不覆盖上游稀疏/延迟矩阵、BiocParallel 后端及已由 `scrapper` 提供的重叠接口。 + +### 245. BinaryCIF 二进制结构格式 (Bio.PDB.binary_cif) + +实现与 Biopython `Bio.PDB.binary_cif` 对应的 BinaryCIF 解析与结构转换。`binary_cif_parse` 使用纯 MoonBit MessagePack 读取器解析 data block、category 和 column,并支持 `ByteArray`、`FixedPoint`、`IntervalQuantization`、`RunLength`、`Delta`、`IntegerPacking` 和 `StringArray` 七类 BinaryCIF 逆编码;编码流水线按规范逆序执行。列 API 提供整数、浮点、文本和原始 CIF token 查询,并将 mask 的 `0/1/2` 分别表示为 present、`.` 和 `?`。解析器严格检查 UTF-8、字节范围、数组长度、整数打包值域、字符串 offsets、MessagePack 嵌套深度及尾随数据。 + +`binary_cif_to_structure` 和 `binary_cif_parse_structure` 将 `_atom_site` 转为现有 `Structure -> Model -> Chain -> Residue -> Atom` 层次,保留模型号、链、残基、坐标、occupancy、B-factor、元素、altloc、插入码、formal charge 和 ATOM/HETATM 语义。模块内置真实 MessagePack fixture,覆盖两模型、蛋白质、水分子和三态 mask。输入 API 接收原始 `Array[Int]` 字节;gzip 数据需由调用方预先解压。BinaryCIF 类别仍保留完整多字符 chain ID,但现有 PDB `Chain.id` 为 `Char`,转换时使用首字符,并拒绝首字符冲突的链 ID。 + +### 246. miloR 单细胞邻域差异丰度 (Bioconductor miloR) + +实现 Bioconductor `miloR` 的核心单细胞邻域差异丰度流程。`milo_build_graph` 在 cell × dimension 降维坐标上构建排除自身、稳定处理距离 ties 的精确 KNN,并将有向边对称化为无向图;`make_neighborhoods` 使用确定性无放回采样,以种子邻域的逐维 median profile 查找精炼代表细胞,去重后生成重叠邻域。`count_cells` 按样本首次出现顺序生成 neighborhood × sample 计数,另提供邻域平均表达、重叠矩阵、距离度量和 `SingleCellExperiment` reduced-dimension 构造入口。 + +`test_neighborhoods` 使用 library-size offset、method-of-moments 离散度及向全局 median 的收缩,为每个邻域拟合 log-link 负二项 GLM,并输出指定系数的 log2 fold change、Wald 统计量、p-value 和 BH FDR。`milo_graph_spatial_fdr` 实现频率加权 BH,支持 k-distance、neighbour-distance、max-distance 和 graph-overlap 四类 connectivity 权重,也可显式禁用 spatial correction。当前模块是无外部 edgeR 依赖的可移植 fixed-effect 实现,不覆盖 NB-GLMM、edgeR TMM/RLE 与 quasi-likelihood 后端,也不提供上游绘图接口。 + +### 247. Infernal cmscan/cmsearch 输出解析 (Bio.SearchIO.InfernalIO) + +实现 Biopython 1.86 `Bio.SearchIO.InfernalIO` 的 Infernal `cmscan`/`cmsearch` 结果读取。`infernal_parse_tabular` 自动识别或显式选择 tabular 格式 1、2、3,保留 clan、模型/序列长度、截断、pipeline pass、GC、bias、bit score、E-value、included、overlap 及格式 2 的重叠索引和比例。所有序列坐标从 Infernal 的 1-based inclusive 规范化为 0-based half-open,并统一处理正负链。结果采用 `InfernalQueryResult -> InfernalHit -> InfernalHSP -> InfernalFragment` 类型层次,按输入顺序聚合重复 query 和 hit。 + +`infernal_parse_text` 支持 non-verbose plain text、`--noali`、CM pipeline 和 HMM-only pipeline,解析 query metadata、hit score、模型/序列比对及 CS、NC、similarity、PP 注释。模型和序列中的 `*[NN]*` local-end 标记会同步拆分为多个 fragment,并记录两侧 omission 长度和链方向坐标。查询 API 提供 best-HSP、E-value/included 过滤、摘要及到通用 `QueryResult` 的转换。当前范围不包括 verbose text、writer 或完整 Infernal 命令行封装。 + +### 248. 重复测量方差分解与差异表达 (Bioconductor variancePartition) + +实现 Bioconductor `variancePartition`/`dream` 的可移植线性混合模型核心。`vp_numeric_effect`、`vp_categorical_effect`、`vp_random_effect` 和 `vp_design` 使用 typed design 表达连续变量、分类固定效应及稳定 level 编码的随机截距,避免解析 R formula 字符串。每个基因拟合 `y = Xβ + ΣZₖbₖ + ε`,以 Cholesky/GLS 计算固定效应,并在 log-variance 空间用 ML 或 REML 估计多个随机效应及异方差残差;`fit_extract_variance_partition` 按上游语义默认使用 ML。结果报告每个固定 term 的 `var(Xⱼβⱼ)`、随机方差、残差方差、归一化占比、拟合值、残差和各 level 的 BLUP。 + +observation-level precision weights 会先缩放到均值 1,并进入 `V = ΣτₖZₖZₖᵀ + σ²diag(1/w)`。`dream`/`dream_se` 对任意固定效应 contrast 计算估计值、标准误、数值 Satterthwaite 自由度、双侧 Student-t p 值和 BH-FDR;`fit_extract_variance_partition_se` 与 `dream_se` 可直接读取 `SummarizedExperiment` 的表达和 weights assay。当前范围支持随机截距,不包括随机斜率、Kenward-Roger、voom mean-variance trend、limma empirical Bayes、缺失值省略及上游绘图接口。 + +### 249. 共享参考序列比对合并 (Biopython Bio.Align) + +实现 Biopython 1.86 `Alignment.from_alignments_with_same_reference` 的共享参考比对合并语义。`shared_reference_input` 表示一条参考和一条或多条 query 的 PWA/MSA,`shared_reference_input_from_pairwise` 可直接接入现有 `PairwiseAlignment`;构造过程验证 raw/aligned sequence、0-based half-open 局部坐标、行宽以及所有输入的参考内容和覆盖区间。核心 `alignments_with_same_reference` 将参考覆盖区间表示为 `L + 1` 个 reference-boundary insertion slots,取各输入同一边界插入宽度的最大值,再将原 query 行投影到统一列空间,因此不会重新比对 query,并支持首端、内部、末端 insertion 及单个输入中的多条 query。 + +`SharedReferenceAlignment` 保留参考/query名称、描述和局部坐标,提供 row/column 查询、reference/query/column 双向坐标映射,以及 identity、mismatch、insertion、deletion 统计。合并结果可转换为现有 `MultipleSeqAlignment` 或 aligned FASTA。算法复杂度为 `O(total input columns + merged rows × merged columns)`;当前要求所有输入覆盖同一参考区间,使用 `-` 表示 gap,不自动执行反向互补或 query 间二次比对。 + +### 250. ambient RNA 去污染 (Bioconductor decontX) + +实现 Bioconductor `decontX` 的可移植单细胞 ambient RNA 去污染核心,输入统一为 gene × cell 非负计数矩阵。每个细胞由所属 cluster 的 native multinomial profile 与 contaminant multinomial profile 混合,污染率按细胞估计;无外部 background 时,污染 profile 由其他 cluster 的本征表达加权形成,提供 empty-droplet/background 矩阵时则使用固定的全局 ambient profile。确定性 EM 在 E-step 分解每个 gene × cell 的 native/contaminant 期望计数,在 M-step 更新 cluster profile 和污染率,并使用 Beta contamination prior、Dirichlet-style profile pseudocount、概率下限及 likelihood tolerance 控制收敛。 + +`decontx` 返回校正计数、污染计数、每细胞污染率、native/contaminant profiles、cluster 编码和完整 likelihood 诊断,且对每个观测保持 `corrected + contaminant = original`。查询 API 提供 cell/cluster estimates、最高污染细胞、名称索引和摘要;`decontx_auto` 对 library-size scaling 后的 `log(1 + count)` 表达执行确定性 k-means 初始化。`decontx_sce` 从指定 assay 和 cluster `colData` 读取输入,在复制的 `SingleCellExperiment` 中增加校正 assay、污染率、cluster 和迭代 metadata,不修改调用方对象。当前范围采用 cluster 标签或 k-means,而非上游 variational Bayes 聚类后端;不包含 GPU/稀疏矩阵专用求解器和绘图接口。 + +### 251. Alignment 坐标组合 (Biopython Bio.Align.Alignment.map/mapall) + +实现 Biopython `Alignment.map` 的 coordinate path composition:第一层表示 outer target 到 shared middle,第二层表示 shared middle 到 final query,组合过程只扫描并相交两条 alignment path,不重新执行序列比对,也不依赖序列内容。`CoordinatePairwiseAlignment` 使用零起始、半开区间坐标,支持局部 alignment 的左右 overhang clipping、exon/intron 与 insertion/deletion gap、正链/反链及双反链组合;可从现有 `PairwiseAlignment` 适配,也可仅提供序列长度和坐标。结果提供 aligned blocks/counts、target/query 双向坐标查询、可选反向互补的 gapped rows、PSL、摘要及 `map_many`。 + +`CoordinateMultipleAlignment.mapall` 将 MSA 每行通过对应 pairwise mapping 投影到统一列空间,支持 nucleotide 1:1 和 protein:nucleotide 1:3 两类一致比例,因此可将 protein MSA 转换为 codon-aware nucleotide MSA,并保留氨基酸 gap 对应的三碱基 gap。构造器会验证名称、序列、坐标边界、单调方向、step size、共享序列长度和跨行映射比例;当前不负责生成原始 pairwise alignment,也不支持混合比例、frameshift 或非整数缩放。 + +### 252. 零膨胀负二项低维表示 (Bioconductor zinbwave) + +实现 Bioconductor `zinbwave` 的可移植 ZINB-WaVE 核心,输入统一为 gene × cell 非负整数计数矩阵。均值子模型使用 log link,零膨胀子模型使用 logit link,两者共享已知 cell-level design、gene-level design、显式 mean/zero offset 和未知 cell latent factors;每个基因单独估计 inverse-dispersion 对应的 dispersion。latent factors 从 library-offset corrected `log(count + 1)` 的 cell Gram matrix确定性初始化,随后与 NB mean、zero-inflation 和 gene/cell effects 交替更新,并在每轮中心化和 RMS 缩放以控制可识别性。 + +E-step 对零计数计算来自 NB component 的后验 responsibility,正计数 responsibility 固定为 1;M-step 分别使用 log-link NB IRLS、logistic IRLS 和带 ridge 的线性求解更新参数。gene dispersion 使用加权矩估计并向跨基因 median 收缩,最终结果提供完整 ZINB log-likelihood 轨迹、AIC/BIC、fitted means、structural-zero probabilities、下游差异分析 observational weights、NB deviance residuals、library-normalized values、零值后验插补和逐基因诊断。 + +`zinbwave_sce` 从指定 `SingleCellExperiment` assay 拟合模型,在不可变副本中加入 `zinbwave_weights`、`zinbwave_residuals`、`zinbwave_normalized`、`zinbwave_imputed` assays、低维表示和模型 metadata,不修改输入对象。构造器会诊断 ragged/负数/非整数/非有限计数、空 cell library、design/offset 维度和标识符问题。当前实现使用 dense MoonBit arrays 和确定性交替求解,不包含上游 R 包的并行后端、稀疏矩阵专用优化、epsilon penalty 路径或绘图接口。 + +### 253. BigBed 二进制区间索引 (Biopython Bio.Align.bigbed) + +实现 Biopython 1.86 `Bio.Align.bigbed` 对应的 UCSC BigBed v4 二进制读写核心。`bigbed_write` 将按染色体和起点排序的 BED3-BED12 记录编码为 binary BED 数据块,写入 64 字节主头、AutoSQL schema、chromosome B+ tree、record count、R-tree interval index、total summary 和尾部 magic;`bigbed_parse` 同时支持小端/大端文件并严格验证版本、偏移、树深度、记录计数、染色体边界及索引一致性。writer 可通过 `items_per_slot` 和 `block_size` 构建多数据块及平衡多级索引,而不是仅支持单叶节点。 + +标准 BED 渐进字段和 custom AutoSQL scalar/array 字段均可往返保留。压缩 reader 实现 RFC 1950/1951 zlib 的 stored、fixed Huffman 和 dynamic Huffman DEFLATE,包括 canonical code、code-length repeat、LZ77重叠复制、32 KiB距离、FCHECK/FDICT/CINFO及Adler-32校验;writer 使用确定性 stored DEFLATE。`BigBedFile::search` 通过R-tree剪枝执行0-based half-open区间查询,另提供全量记录、名称查询、汇总和BED文本导出。BED12记录可恢复含intron jump的target/query坐标路径,负链query坐标按转录本长度反向。 + +构造器和解析器会诊断无效BED字段、未排序记录、未知染色体、损坏magic、截断数据、异常树节点、非法zlib流及校验和不一致。当前范围不写入zoom levels或extra string indices;offset/count虽按u64布局读取和写入,但内存模型仍拒绝高32位非零的超大文件。writer的stored DEFLATE保证互操作性和确定性,但不追求压缩率。 + +### 254. 自适应重尾效应量收缩 (Bioconductor apeglm) + +实现 Bioconductor `apeglm` 的 negative-binomial 路径,用于对 RNA-seq GLM 的目标系数执行自适应重尾后验收缩。`apeglm_fit` 接受 gene × sample 非负整数计数、sample × coefficient 设计矩阵、逐基因 dispersion,以及可选 log offset 和 observation weight;模型使用 `log(mu) = offset + X beta`,所有内部系数、阈值和先验尺度统一使用 natural-log scale。每个基因先拟合无先验 NB-GLM MLE,再以异方差 Efron-Morris 方程从跨基因 MLE/SE 自适应估计先验方差;目标系数使用可配置自由度的 Student-t 先验,默认 `df=1` 即 Cauchy,intercept 和其他 nuisance 系数使用宽 Normal no-shrink prior。 + +MAP 求解器实现阻尼 Newton、Cholesky 信息矩阵求解、逐级 ridge 正定化、最大步长、28 级回溯线搜索和严格参数边界;从 MLE、0 和先验尺度的正负倍数生成精确数量的确定性多起点并选取最高 posterior mode。MAP 处逆信息矩阵给出 Laplace posterior covariance、SD 和可配置可信区间;结果同时计算 local false sign rate、超过效应阈值时发生 false sign or small effect 的 FSOS probability,以及按 FSR 排序累计均值定义的 s-value。查询 API 支持名称/索引查找、MAP/SD 矩阵、排序、筛选、摘要和 natural-log/log2 TSV 导出。 + +`apeglm_from_deseq2` 读取 `DESeqDataSet` 的 counts、design、dispersion 和 size-factor log offset;`apeglm_summarized_experiment` 在不可变副本中增加 `apeglm_map`、`apeglm_sd`、`apeglm_fsr`、`apeglm_svalue` 和 `apeglm_fsos` assays 及模型 metadata。构造器会诊断 ragged、负数、非整数或非有限计数,无效 design/dispersion/offset/weight、全零权重和重复标识符。当前范围实现 dense negative-binomial backend 和 Laplace/Normal posterior approximation,不包含上游 beta-binomial backend、grid/HPD integration、稀疏矩阵专用优化或并行执行。 + +### 255. BigMaf 多物种比对索引 (Biopython Bio.Align.bigmaf) + +实现 Biopython 1.86 `Bio.Align.bigmaf` 对应的 UCSC BigMaf 格式。`bigmaf_write` 将多物种 `BigMafBlock` 编码为标准 `bedMaf` `bed3+1` BigBed v4 文件,第四个 AutoSQL `lstring mafBlock` 字段保存分号分隔的完整 MAF block;二进制头、chromosome B+ tree、R-tree 和 DEFLATE 复用经过验证的 BigBed 实现。writer 验证第一条 component 为 `.` 正链参考、source size 与目标长度一致,并按目标顺序和半开区间稳定排序后写入压缩或非压缩数据块。 + +严格 MAF 模型完整保留 `a` 行的 score/pass、`s` sequence component、`i` insertion context、`q` aligned quality、`e` empty component 和注释。component API 提供去 gap 序列、正向序列/区间、alignment column 与 source coordinate 双向转换,以及参考位置到任意物种的映射;负链坐标按 MAF forward-coordinate 规则转换。`BigMafFile::search` 通过底层 R-tree 查询目标区间,同时接受裸 chromosome 或 reference-qualified 名称;另提供全量 block、pairwise identity、摘要和标准 MAF 文本导出。 + +解析器要求 `definedFieldCount=3`、`fieldCount=4` 及标准 `bedMaf` schema,并交叉验证 BED 区间、嵌入 MAF 第一条 component、reference prefix 和 chromosome target。构造器会诊断非法状态字符、负坐标、source 越界、size 与非 gap 长度不符、quality/gap 不同步、重复注释及不一致列宽;二进制层继续诊断损坏 magic、树节点、DEFLATE 和校验和。当前范围不生成 BigMaf zoom levels 或 extra indices,也不实现远程 HTTP range reader。 + +### 256. cohort-scale 单细胞重复测量分析 (Bioconductor dreamlet) + +实现 Bioconductor `dreamlet` 的 sample×cell-type pseudobulk 混合模型流程,输入统一为 gene×cell 非负整数计数。`dreamlet_aggregate_to_pseudobulk` 按 sample 和 cluster 首次出现顺序求 raw count sum,为缺失的 sample×cluster 组合补零列和 `cell_counts=0`,并验证被选择 metadata 在同一样本内保持一致;`dreamlet_aggregate_sce` 可直接读取 `SingleCellExperiment` assay 与 `colData`。每个 cell type 独立过滤低细胞数或零文库样本,并按 total count、最小 count、CPM 达标样本比例过滤基因,避免不同 cell type 共享不适用的 retained set。 + +归一化实现 edgeR 风格 TMM:按归一化 count 的 75% 分位数选择参考样本,计算 M/A 值,执行 log-ratio 与 abundance 双裁剪,以 inverse asymptotic variance 加权,并将因子几何中心化到 1。effective library size 用于 normalized CPM 和带 prior count 的 log2 CPM。precision-weight 流程先从 count 均值和 library scale 构造可配置的 Poisson 初始权重,再拟合 residual variance 四次方根对平均表达的 LOWESS 趋势,最终使用预测方差倒数作为 voom-style observation weights。 + +`DreamletEffectSpec` 和 `DreamletModelSpec` 以 typed numeric、categorical、random effects 代替 R formula 解析,并转换到现有 `variancePartition` 设计与求解器;常量固定效应和无重复 level 的随机效应会按 cell type 删除并记录。`dreamlet_process_assays` 输出每个 cell type 的过滤、TMM、表达、权重、design 和趋势诊断,`dreamlet` 对指定 coefficient 运行 weighted fixed/random mixed model,报告数值 Satterthwaite 检验、cluster 内 BH-FDR 和跨全部 gene×cell-type hypotheses 的 study-wide BH-FDR。查询 API 提供 assay lookup、top table 和分阶段摘要。当前范围支持 dense arrays 和随机截距,不解析 R formula,不包含随机斜率、Kenward-Roger、limma empirical Bayes、稀疏/并行后端、绘图或上游 `aggr_means` 的变化 numeric cell metadata 聚合。 + +### 257. BigPsl 成对比对索引 (Biopython Bio.Align.bigpsl) + +实现 Biopython 1.86 `Bio.Align.bigpsl` 对应的 UCSC BigPsl 格式。`bigpsl_write` 将 `BigPslAlignment` 编码为标准 BigBed v4 `bed12+13` 文件,完整写入 25 字段 `bigPsl` AutoSQL、chromosome B+ tree、R-tree 和可选 DEFLATE 数据块;`bigpsl_parse` 复用 BigBed 二进制校验并额外交叉验证 target/query 区间、block count、`oChromStarts`、`chromSize`、match 分类和 `seqType`。writer 支持选择性保存 query sequence 与 NCBI CDS 字段,默认保持文件紧凑。 + +坐标模型使用 0-based half-open path。核酸 block 要求 target 递增且 target/query 长度 1:1,支持 query 正反链;translated DNA-to-protein block 要求 query 递增且 target/query 长度 3:1,支持 target 正反链。API 提供 block/gap 统计、target/query 双向映射、amino acid 到 codon interval 映射、区间和 query-name 查询、摘要及标准 PSL 导出,其中反向 translated alignment 使用 PSL `+-` strand。`recount` 可按实际序列重新计算 match、mismatch、repeat-match 和 wildcard/N,核酸支持 lower/upper repeat mask,translated alignment 使用标准遗传密码并在反向 target 上先做 reverse complement。 + +构造器和解析器诊断名称、坐标边界/方向、非 1:1 或 3:1 block、无 aligned block、非法 score/thick interval、序列长度、match 总数、schema 及 strand 组合。当前范围不生成 zoom levels 或额外 string index,不实现远程 HTTP range reader;二进制容器能力与 BigBed 保持一致。 + +### 258. Dirichlet Monte Carlo 组成型差异丰度 (Bioconductor ALDEx2) + +实现 Bioconductor `ALDEx2` 的可移植两组组成型推断核心,输入统一为 feature × sample 非负整数计数。`aldex2_clr` 先过滤跨全部样本均为零的 feature,以 `count + prior`(默认 prior 0.5)作为 Dirichlet shape,使用显式 seed、Park-Miller LCG、Box-Muller normal 和 Marsaglia-Tsang gamma sampler 生成确定性 Monte Carlo 实例。log2 ratio 变换支持 `all`、`median`、IQLR、condition-specific `zero`、LVHA 和调用方指定原始输入 feature index 的 `user` denominator,也可接收 sample × Monte Carlo instance 的正 scale matrix 直接构造 scale-aware log abundance。 + +`aldex2_ttest` 在每个实例内执行两组 Welch 与 Mann-Whitney/Wilcoxon 检验,配对模式改用 paired t 和 signed-rank;每个实例独立执行标准 Benjamini-Hochberg 校正,再跨 posterior 实例求 expected p/eBH。`aldex2_effect` 报告组内 relative abundance、between/within difference、标准化 effect、可信区间和 sign overlap;另提供 posterior expected Aitchison距离、名称查询、排序、阈值筛选、摘要及 ALDEx2 风格 TSV。`aldex2_summarized_experiment` 在不可变副本中加入 effect、overlap、Welch eBH 和 Wilcoxon eBH assays,并按原始行索引回填被过滤的全零 feature。 + +构造器会诊断 ragged、负数、非整数或非有限计数,空样本文库、名称/condition/配对维度、denominator 和 scale matrix 错误。当前范围不包含 `aldex.glm`、Kruskal-Wallis、相关性、绘图、BiocParallel 或自动 gamma scale uncertainty simulation;effect posterior 使用同一实例内的确定性 pairwise 组合,不保证与上游最多 10000 次随机重采样逐位一致。 + +### 259. Alignment 详细计数与评分 (Biopython Bio.Align.Alignment.counts) + +实现 Biopython 1.86 `Alignment.counts` / `AlignmentCounts` 的坐标路径统计模型,并保留原有轻量 `CoordinatePairwiseAlignment::counts()` API。新 `alignment_counts` API 将 gap 分为 left/internal/right insertion 和 deletion 十二类 open/extend 事件;同方向连续 gap 计为 extension,diagonal step 重置 gap path。结果同时提供各层级聚合 getter、aligned、identity、mismatch、positive、gap/substitution/total score 和摘要。 + +composition 支持 wildcard、match/mismatch score 或替换矩阵;矩阵模式计算正分 substitution 的 positives,并对未知 residue 给出明确诊断。gap score 可使用统一 affine、按 insertion/deletion 方向区分,或完整十二参数配置。坐标读取支持反向链的归一化与反向互补;无序列文本时仍可统计坐标和 gap。`CoordinateMultipleAlignment::alignment_counts` 对全部无序序列对求和,并忽略为其他行插入的双 gap 列而保持当前 gap path。 + +实现覆盖路径长度、坐标单调性、aligned step、有限 score 和矩阵 alphabet 校验;专项测试包含官方 BLOSUM62/BLOSUM45 示例、左右/内部 gap、连续 open/extend、wildcard、反向链、length-only alignment 和 MSA 汇总。同步修正既有标准 20×20 `BLOSUM45` 的 139 个错误分值,与 Biopython 1.86 官方矩阵逐项一致。当前 API 统计已有 coordinate alignment,不负责执行新的序列比对。 + +### 260. Dirichlet-multinomial 混合聚类与分类 (Bioconductor DirichletMultinomial) + +实现 Bioconductor `DirichletMultinomial` 的可移植有限混合模型核心,原生输入与上游 `dmn()` 一致,采用 sample × taxon 非负整数计数矩阵。基础 API 提供 Dirichlet-multinomial log-PMF、均值和含过度离散膨胀的协方差。混合拟合先在相对丰度空间执行确定性 soft k-means,再以 posterior responsibility 和 component weight 进行 EM;每个 component 的 alpha 使用 log 参数化、Gamma(shape=0.1, rate=0.1) prior、inverse-BFGS、Armijo 回溯线搜索和显式参数边界优化,避免正值约束被迭代破坏。 + +`DmnFit` 返回按 mixture weight 降序排列的 alpha、权重、sample responsibility、component proportions、浓度、Hessian 近似区间和 likelihood 轨迹,并提供新样本 evidence、posterior 与 assignment。goodness-of-fit 遵循上游参数计数 `P = K × taxa + K - 1`,报告 negative log evidence、log determinant、Laplace、AIC 和 BIC;`dirichlet_multinomial_select` 比较连续 K。`dirichlet_multinomial_group_fit` 为每个 phenotype 拟合独立 DMM 并结合经验 group prior 构造生成式分类器,可指定每组 K 或按 Laplace 自动选择;另提供确定性分层交叉验证、概率输出和二分类 ROC/AUC。 + +`dirichlet_multinomial_fit_se` 从 `SummarizedExperiment` 的 feature × sample assay 校验并转置为 sample × taxon。构造器会诊断空/ragged矩阵、负数、非整数或非有限 assay、零文库、组件数、名称、group、fold 和预测维度错误。当前实现使用 dense MoonBit arrays 和确定性单线程求解,不依赖上游 C/GSL,也不包含稀疏矩阵专用优化、并行多起点、绘图或完整 S4 方法分派。 + +### 261. Alignment-aware tabular 搜索结果解析 (Biopython Bio.Align.tabular) + +实现 Biopython 1.86 `Bio.Align.tabular` 的 alignment-aware 表格解析器,支持 NCBI BLAST `-outfmt 7`、FASTA `-m 8CB` BTOP 和 `-m 8CC` `aln_code`。解析结果按 query block 保留 program/version、command line、database、RID、query 描述与长度、完整上游字段词汇、声明命中数和零命中 query;`AlignTabularDocument` 提供扁平化、query 查询,`AlignTabularQueryResult` 提供 E-value 最佳命中与阈值过滤。 + +BTOP 和 FASTA CIGAR 会合并连续 operation 并重建显式 target/query coordinate path,区分 aligned、query gap 和 target gap。输入的 1-based inclusive 区间规范化为 0-based half-open,同时保留正反链方向;BLASTX、TBLASTN、TBLASTX、RPSTBLASTN、FASTX/FASTY 和 TFASTX/TFASTY 按核酸轴每个 residue 三个单位换算。绝对路径可转换为 `CoordinatePairwiseAlignment`,并提供 alignment columns、aligned residues、两类 gap residues、gap events 和摘要。 + +解析器严格诊断 header/字段重复或缺失、未知字段、列数、整数溢出、非法或非有限浮点、百分比和 E-value 范围、ID/长度冲突、坐标边界、traceback 截断与 operation、BTOP/CIGAR 同时出现、traceback span/序列消耗/alignment length 不一致,以及声明命中数和 processed query 数不匹配。当前范围聚焦文本解析和坐标重建,不执行 BLAST/FASTA 搜索,也不解析 XML、ASN.1 或普通无注释 outfmt 6 文档。 + +### 262. 最近邻高斯过程空间变异基因检测 (Bioconductor nnSVG) + +实现 Bioconductor `nnSVG` 的可移植 spatially variable gene 检测核心,输入统一为 gene × spot 表达矩阵与 spot × dimension 空间坐标。坐标按各维最大 range 统一缩放,支持确定性的 approximate maximum-minimum-distance (AMMD) 和坐标和排序;每个 spot 只连接处理顺序中的前驱近邻。指数协方差 `R(i,j)=exp(-distance/length_scale)` 通过 NNGP 条件分解生成局部回归系数和条件方差,避免构造完整高斯过程精度矩阵。 + +每个基因拟合 `y=X beta+w+epsilon`,以 NNGP generalized least squares profile maximum likelihood 联合搜索 gene-specific spatial length scale 和 spatial variance proportion,并通过确定性局部细化优化。结果报告 `sigma_sq`、`tau_sq`、`phi=1/length_scale`、空间方差占比、回归系数及收敛状态;空间模型与非空间线性模型使用 likelihood-ratio statistic 比较,以 chi-square df=2 tail 计算 p-value,再执行 Benjamini-Hochberg FDR 和稳定排名。满秩协变量设计受到严格校验,常量或达到方差下限的基因显式回退到非空间模型。 + +`nnsvg_filter_genes` 实现按最小计数、表达 spot 百分比和 `MT-`/`mt-` 前缀过滤;`nnsvg_spatial_experiment` 从指定 assay 与空间坐标运行模型,在不可变 `SpatialExperiment` 副本的 rowData 中写入 13 项 nnSVG 统计和 metadata。API 另提供基因查询、top/significant 结果、摘要与示例数据。当前实现使用 dense 小型前驱协方差矩阵和确定性单线程网格优化,不依赖 BRISC、R、BiocParallel 或稀疏矩阵后端。 + +### 263. Alignment-aware PSL/PSLX 读写与坐标模型 (Biopython Bio.Align.psl) + +实现 Biopython 1.86 `Bio.Align.psl` 对应的 alignment-aware UCSC PSL/PSLX 模型。`psl_parse` 接受带或不带 `psLayout version` header 的 21 列 PSL 和 23 列 PSLX,使用 0-based half-open 坐标重建显式 target/query path;`PslAlignment` 保留 match、mismatch、repeat match、N、query/target insertion 和 block sequence,不把全部对齐位置简化为 match。核酸比对支持 `+`/`-` query 方向,translated DNA-to-protein 比对支持 `++`/`+-` target 方向及严格 3:1 步长,PSLX translated target block 保存翻译后的氨基酸片段。 + +模块提供 block/gap 统计、identity/score、target-query 双向映射和 query residue 到 target codon interval 映射。`recount` 可从完整序列重新计算匹配分类,支持核酸反向互补、lower/upper repeat mask、自定义 wildcard,以及正向或反向 target DNA 翻译;writer 可选择 header、PSLX、自动 recount 和版本。`PslDocument` 支持 query/target 过滤、跨记录摘要和解析-写出往返,并可与 `CoordinatePairwiseAlignment` 互转;由于现有通用坐标模型要求 target 递增,`+-` translated alignment 在适配时给出明确错误。 + +解析器严格拒绝非十进制或溢出整数、非法 strand/列数、blockCount 与 CSV/PSLX 数量不一致、零长度/重叠/逆序/越界 block、首尾 gap、错误的核酸 1:1 或 translated 3:1 比例、match 分类总和、q/t insertion 统计及声明区间不一致。该模块负责普通文本 PSL/PSLX;`search_io.mbt` 继续提供简化的 BLAT 搜索结果适配,`bigpsl.mbt` 负责 bed12+13 BigBed 二进制索引。 + +### 264. 空间邻域增强聚类 (Bioconductor Banksy) + +实现 Bioconductor `Banksy` 1.9.1 的空间转录组邻域增强特征与聚类工作流。输入统一为 gene × spot 表达矩阵及 spot × dimension 坐标;`H0` 计算归一化加权邻域均值,`H1+` 在局部非加权均值中心化后计算方位 Fourier/Gabor harmonic 幅值。支持 `kNN_median`、inverse-distance、inverse-power、rank、uniform 和 radius-Gaussian 六类空间核、每阶独立邻域大小、确定性邻居采样,以及二维方位角和多维欧氏距离。 + +`banksy_get_matrix` 按 `sqrt(1-lambda)` 加权原始表达,并将 `lambda * 2^-m` 在 `H0..HM` 间归一化后开方加权;支持全局或按 section/sample 分组的 feature 标准化。下游提供 spots Gram 矩阵上的 Jacobi PCA、确定性 farthest-point 多起点 k-means、空 cluster 重播种、silhouette、空间近邻一致率、Adjusted Rand Index、同步标签平滑和多 lambda 参数扫描,适配 cell typing 与 tissue domain segmentation 两类用法。 + +`banksy_spatial_experiment` 从指定 assay、rowData、colData 和二维/三维 spatial coordinates 构建模型,在不可变容器副本中写回 `H0..HM` assays、原始/平滑 cluster 标签与 metadata。所有入口校验矩阵方向、矩形性、有限值、名称唯一性、邻域/采样边界、lambda、PCA/聚类维度及分组完整性。当前实现使用 dense MoonBit arrays、确定性单线程 Jacobi PCA 与 k-means,不依赖 R、BiocParallel、igraph、Leiden 或稀疏矩阵后端。 + +### 265. Alignment-aware SAM 严格读写与坐标模型 (Biopython Bio.Align.sam) + +实现 Biopython 1.86 `Bio.Align.sam` 的 alignment-aware SAM 模型,并与已有宽松记录级 `sam.mbt` 并存。`align_sam_parse` 解析 `@HD/@SQ/@RG/@PG/@CO` 及扩展 header,索引 reference metadata;每条 mapped record 将 1-based `POS/PNEXT` 规范化为 0-based 坐标,并从 `M/I/D/N/=/X` 构建显式 target/query path。soft/hard clipping 保留在 CIGAR 元数据中,`N` 与 deletion 分开统计;反向链 query sequence 和 PHRED 转为生物学正向语义,坐标递减,写回时恢复 SAM 存储方向。 + +optional tags 保留 `A/i/f/Z/H/B:c/C/s/S/i/I/f` 的类型和数组 subtype,提供类型化构造与防御性查询。模块支持 CIGAR 统计、flag/reference/mate访问、target-query 双向坐标映射、aligned row 重建、MD reference reconstruction、NM 计算及基于完整 reference 的 MD 生成。`align_sam_format` 和 `align_sam_write` 提供规范 record/document 往返,`align_sam_create` 从生物学方向的 sequence、qualities、CIGAR 和 typed tags 构造 mapped alignment。 + +解析器严格诊断 header 顺序与重复 reference、字段数和整数边界、mapped/unmapped 一致性、reference 越界、CIGAR clipping/P 操作、SEQ/QUAL/CIGAR 长度、PHRED 范围、重复或非法 tag、B-array subtype/range,以及 MD token 与 CIGAR `=/X/D` 的逐碱基结构冲突。当前范围聚焦 SAM 文本和 alignment coordinate semantics,不解码 BAM/CRAM;二进制格式继续由现有 `bam.mbt`、`cram_wbtest.mbt` 负责。 + +### 266. 细胞群与基因模块联合聚类 (Bioconductor celda) + +实现 Bioconductor `celda` 的 `celda_CG` 可移植核心,输入统一为 feature × cell 非负整数计数矩阵。模型联合推断 cell population 标签与 feature module 标签,并以 `Theta`(sample 内 population)、`Phi`(population 内 module)、`Psi`(module 内 feature)和 `Eta`(全局 module abundance)构成四层 Dirichlet-multinomial;collapsed log-likelihood 使用 `alpha/beta/delta/gamma` 超参数,并保留上游每个 module 一个 pseudogene 的平滑语义。 + +推断支持确定性 hard-EM 与 seeded Gibbs、多链 farthest-first/balanced 初始化、最佳状态保存、提前停止、非空 population/module 约束及稳定标签重排。结果提供 posterior 参数、fitted counts、population/module 成员与 top features、perplexity、AIC/BIC、独立 likelihood 计算和新细胞 posterior prediction;`celda_cg_grid_search` 可按 BIC、perplexity 或 likelihood 比较 K/L 候选。 + +`celda_cg_sce` 从指定 assay 和可选 sample `colData` 读取输入,在不可变 `SingleCellExperiment` 副本中写入 fitted assay、1-based population/module 标签及模型诊断 metadata。所有入口校验矩阵方向、矩形性、有限非负整数、名称唯一性、sample/初始标签完整性和超参数边界。当前实现采用 dense MoonBit arrays 和单线程完整 collapsed likelihood 重算,面向中小型矩阵及可验证工作流,不等同于上游 C++/OpenMP 大规模性能后端。 + +### 267. A2M 状态感知多序列比对 (Biopython Bio.Align.a2m) + +实现 Biopython 1.86 `Bio.Align.a2m` 的单 MSA 严格读写。解析器由首行逐列推导 `D`(match/deletion)或 `I`(insertion)状态,验证后续行在同列使用相同字符类别;内部将残基统一为大写、`.`/`-` 统一为 gap,同时独立保留状态,因此写回时可准确恢复 match 列大写/连字符和 insertion 列小写/点。支持 wrapped sequence、空行、CRLF、header 描述和 canonical line wrapping,并对空记录、非等宽行、非法字符、状态错位及非法构造参数给出类型化错误。 + +`A2mAlignment` 提供 sequence position、alignment column 与跨行 residue 的 0-based 映射,gap 返回 `None`;连续 insertion-state 列按 reference-boundary slot 汇总。分析 API 覆盖逐列坐标对、identity/mismatch/gap/double-gap、match/insertion aligned counts、占用率、阈值共识、列切片和 match-only 投影。独立状态模型避免传统 `MultipleSeqAlignment` 归一化后丢失 A2M 的模型列语义。 + +### 268. EMBOSS alignment 输出解析与坐标模型 (Biopython Bio.Align.emboss) + +实现 Biopython 1.86 `Bio.Align.emboss` 的 alignment-aware parser,读取 water、needle、stretcher、matcher、alignret 等工具产生的 `srspair`、`pair` 和 `simple` 报告。类型化 document 保留 Program、Rundate、Commandline、Align_format、Report_file;每个 alignment 保留任意数量的序列、matrix、gap/extend penalty、score、Identity/Similarity/Gaps 及 longest/shortest 注释。固定 21 列正文支持截断 identifier、多 block、多 alignment、全空格 consensus 和纯 gap block。 + +EMBOSS 的 1-based inclusive 坐标在内部规范为 0-based boundary/residue 坐标;正向、反向和局部区间均支持 column-to-position、position-to-column、跨行映射、aligned pairs 与 compact coordinate path。pair counts 区分 identity、mismatch、insertion/deletion、double-gap、gap-open 和 consensus positive。严格校验覆盖 header/annotation、数值范围、row顺序、block宽度、坐标连续性、declared Length 和报告统计;额外提供 wrapped canonical writer,以补足上游只读模块并保证严格往返。 + +### 269. Exonerate alignment 输出读写与坐标模型 (Biopython Bio.Align.exonerate) + +实现 Biopython 1.86 `Bio.Align.exonerate` 的 alignment-aware Exonerate 模型,与既有 `exonerate.mbt` 的 `Bio.SearchIO.ExonerateIO` 搜索结果聚合 API 并存。`align_exonerate_parse` 严格读取 `Command line`、`Hostname`、completion marker 和零个或多个 `cigar:`/`vulgar:` alignment;不可变 document、alignment 和 operation 类型保留 query/target identifier、0-based boundary、`+`/`-`/`.` strand、score 及每段双轴步长。 + +vulgar 的 `M/5/I/3/C/G/N/S/F` 操作规范为显式 `M/5/N/3/C/D/I/U/S/F` path,其中双轴 non-equivalenced region 拆成可查询的 target/query movement,并在写回时无损重组。模块支持正向、反向和 protein strand、DNA/protein 3:1 translated CIGAR、绝对 coordinate path、query-target 双向 residue/codon 映射、aligned pairs,以及 match、gap open、intron、non-equivalenced、split codon 和 frame shift 统计。vulgar writer 保留完整操作语义,cigar writer 将特殊操作规范投影为 `M/I/D` 且保持路径;严格诊断覆盖 header/footer、字段与数值、strand方向、operation合法性和 endpoint span。当前模块负责 alignment coordinate semantics,不替代搜索结果层的 `Bio.SearchIO.ExonerateIO`。 + +### 270. GCG MSF 多序列比对读写与坐标模型 (Biopython Bio.Align.msf) + +实现 Biopython 1.86 `Bio.Align.msf` 的 GCG/PileUp 多序列比对格式。`msf_parse` 支持 `!!AA_MULTIPLE_ALIGNMENT`、`!!NA_MULTIPLE_ALIGNMENT` 和 `PileUp` header,解析 `MSF:/Type:/Check:`、EMBOSS `CompCheck:`、自由 preamble/title/date、每行 `Name:/Len:/Check:/Weight:` metadata、可选 `oo`、数字坐标行和 interleaved sequence blocks。`.`、`~`、`-` gap 在内部统一为 `-`,蛋白质与核酸残基分别校验,CRLF 与小写输入被规范化。 + +模块实现位置权重 1..57 循环的标准 GCG checksum,并支持默认严格校验或显式关闭验证;第三方零 checksum 视为未提供。官方 W protein fixture 的 93-residue 行会补齐到 99 列;DOA fixture 一类 header width 与实际宽度不一致的文件保留 `declared_length`,由 `length_mismatch` 暴露而不中止解析。不可变 metadata、sequence、alignment 和 pair-count 模型提供 row/column 双向映射、跨行 residue 映射、aligned pairs、compact coordinate path、identity/mismatch/gap-open、occupancy、阈值 consensus 和 checksum 摘要。 + +`msf_write` 重新计算行级与文件级 checksum,按可配置 block/group width 输出 canonical interleaved MSF,并可选择 `.`, `~` 或 `-` gap。严格诊断覆盖 header/type 冲突、数值溢出、重复 ID、descriptor/body 损坏、声明 residue 数不一致、非法字符、all-gap column 以及 writer 参数。当前实现聚焦单个文本 MSF alignment,不负责流式多记录容器。 + +### 271. NEXUS 多序列比对读写与坐标模型 (Biopython Bio.Align.nexus) + +实现 Biopython 1.86 `Bio.Align.nexus` 的现代 alignment 语义,并与原有 `nexus.mbt` 的浅层 `Bio.Nexus` block API 分离。`align_nexus_parse` 严格检查 `#NEXUS`,以 quote/comment-aware lexer 处理任意嵌套 `[...]` comment、分号命令边界、单/双引号 taxon name 和 doubled apostrophe;支持 `DATA`/`CHARACTERS`、独立 `TAXA`/`TAXLABELS`、`DIMENSIONS NTAX/NCHAR`、`FORMAT DATATYPE/MISSING/GAP/MATCHCHAR/INTERLEAVE/RESPECTCASE/SYMBOLS`,以及带标签、无标签、wrapped sequential 和 interleaved MATRIX。后续 interleave block 可按 NEXUS 的空格/下划线等价规则关联 taxon,duplicate taxon 仍保留原始 ID。 + +DNA、RNA、protein 和 standard datatype 分别执行残基集合校验;custom missing/gap 被规范化,其中 missing 仍推进 residue coordinate,只有 gap 不推进。`MATCHCHAR` 以首行同列残基展开,source 中全 gap 列在最终 alignment 中移除,同时 metadata 保留原始宽度和删除列数。不可变 metadata、sequence、alignment 与 counts 模型提供 row/column 双向位置查询、跨序列 residue mapping、aligned pairs、compact coordinate path、逐对及全 MSA identity/mismatch/gap-open 统计、occupancy、阈值 consensus 和摘要。 + +`align_nexus_write` 输出 canonical DATA block,正确 quote taxon name 并将 apostrophe 写为 `''`,支持 custom gap、可配置 block width 和显式 sequential/interleaved;未指定模式时,alignment 宽度大于 1000 自动 interleave,与 Biopython writer 一致。严格诊断覆盖 header、comment/quote、block、dimensions、format、matrix、datatype、taxon关联、MATCHCHAR 和 writer 参数。官方 9×48 fixture 在移除两个全 gap 列后得到 9×46 alignment,全部序列对统计为 862 aligned、256 identities、606 mismatches 和 596 gap columns。 + +### 272. Stockholm 多序列比对读写、注释与坐标模型 (Biopython Bio.Align.stockholm) + +实现 Biopython 1.86 `Bio.Align.stockholm` 的现代 alignment 语义,并保留原有 `stockholm.mbt` 的宽松 block-oriented API。`align_stockholm_parse` 严格检查 `# STOCKHOLM 1.0` header 和 `//` terminator,`align_stockholm_parse_all` 支持连续多记录;sequence row 必须唯一且等宽。类型化不可变模型保留 GF alignment annotation、GS sequence annotation、GR residue annotation 和 GC column annotation,映射标准字段并保留自定义 GS/GR/GC;reference 的 RN/RM/RT/RA/RL/RC、database reference 的 DR/DC、nested domain 的 NE/NL 和重复 AU/WK/SM 均按输入顺序建模。 + +解析器区分 `-` deletion gap 与 `.` insertion gap,形成每列 `M/D/I` operation;同列禁止混用两类 gap。内部 aligned row 将 gap 规范为 `-`,GR 同时保存 residue-level 与 aligned value。source 中全 gap 列会与 GC、GR 和 operation 同步压缩,并保留 source width 与删除列数。查询 API 提供 row/column 双向位置映射、跨行 residue mapping、aligned pairs、compact coordinate path、annotation-aware column slicing、逐对及全 MSA identity/mismatch/gap-open 统计、occupancy、阈值 consensus 和摘要。 + +`align_stockholm_write` 输出 canonical GF/GS/GR/GC 顺序,根据 operation 将 insertion gap 恢复为 `.`,按 row gap 展开 residue-level GR,并对长 CC/RC/RT 文本折行;`align_stockholm_write_all` 支持多记录往返。严格诊断覆盖 header/terminator、annotation字段、SQ row count、数值溢出、orphan/duplicate reference字段、非法列操作、annotation宽度和writer不支持的未知GF。官方 HAT fixture 保留 3×33 alignment、完整注释与 insertion column;专项 fixture 验证 source 7列压缩为 retained 4列且 operation 为 `MIMM`。 + +### 273. UCSC Chain成对比对读写与坐标模型 (Biopython Bio.Align.chain) + +实现 Biopython 1.86 `Bio.Align.chain` 的现代 pairwise alignment 语义,并与现有 `chain_liftover.mbt` 的宽松 rtracklayer-style liftOver API 并存。`align_chain_parse` 严格读取 12/13 字段 header,支持有限浮点 score、可选 chain ID、连续多记录、CRLF/空白分隔及 target/query 任意 `+`/`-` 组合;每条 `size dt dq` 和末尾 `size` 记录被重建为 zero-based half-open 的绝对坐标路径,负链通过递减边界表示。 + +不可变 `AlignChainAlignment`、`AlignChainCoordinate`、`AlignChainBlock`、`AlignChainCounts` 和 range/pair 模型提供 canonical block 重建、`M/D/I` operation path、aligned/gap/open统计、双向 residue 映射、区间拆分映射、受上限保护的 aligned-pair 展开、target/query反转、chain ID查找和target overlap查询。`align_chain_write`/`align_chain_write_all` 从绝对路径恢复规范 header 与block,可稳定多记录往返。 + +严格诊断覆盖 identifier、strand、非负整数、32位溢出、NaN/极值score、header区间、block字段数、末尾block、零长度segment、单调性、aligned step等长及累计span一致性。官方风格 181-base reverse-query fixture 重建 `MDMDM`、181 aligned bases、1530 target gap bases和1711列;第二条fixture覆盖reverse target与双向gap。旧 `cl_*` API 保持兼容,继续面向宽松的基因组liftOver工作流。 + +### 274. MAF多基因组比对读写、坐标路径与参考索引 (Biopython Bio.Align.maf) + +实现 Biopython 1.86 `Bio.Align.maf` 的现代 alignment 语义,并保留 `maf.mbt` 的早期宽松块分析 API。新模块复用 `bigmaf.mbt` 已验证的 `BigMafComponent`、`BigMafEmptyComponent`、`BigMafInsertion` 和 `BigMafBlock` 作为严格数据层,在其上提供不可变 `AlignMafTrack`、`AlignMafDocument`、`AlignMafIndex`、`AlignMafSplicedAlignment` 和摘要模型。parser 支持 quoted/escaped UCSC `track` metadata、`##maf version/scoring/program`、文档与块注释,以及标准 `a/s/i/e/q` 记录;`i/q` 必须紧随同源 `s` 行并与其列宽、gap位置一致。 + +component 坐标使用 zero-based half-open forward genomic axis;负链 MAF start 被转换为递减的绝对边界。`align_maf_coordinate_path` 按列状态变化压缩多序列路径,`align_maf_map_position` 可在任意两个 component 间双向映射 residue,并在目标gap处返回 `None`。gap字符 `.`, `=`, `_` 统一规范为 `-`;数值解析拒绝负整数、32位溢出、NaN和极端score,source span 使用减法边界检查避免加法溢出。 + +`AlignMafIndex::create` 要求每块参考 component 唯一且source size/strand一致,并按forward interval构建内存索引。`search`/`search_ranges` 使用半开overlap、文件顺序返回和跨区间去重;`get_spliced` 验证递增不重叠外显子,保留reference gap对应插入列,未覆盖参考填 `N`、其他物种填 `-`,并支持完整结果反向互补。重叠reference block会被诊断为歧义;当前拼接明确要求plus-strand reference。`align_maf_write` 输出canonical track/header与a/s/i/q/e顺序,可稳定文档往返;BigMaf二进制容器和旧宽松MAF分析接口继续独立存在。 + +### 275. BED成对比对读写、双轴坐标与分级写回 (Biopython Bio.Align.bed) + +实现 Biopython 1.86 `Bio.Align.bed` 的现代 pairwise alignment 语义,并与 `bigbed.mbt` 的 BigBed v4二进制容器、AutoSQL和索引职责分离。不可变 `AlignBedAlignment`、`AlignBedDocument`、`AlignBedCoordinate`、`AlignBedBlock`、`AlignBedCounts` 和 `AlignBedSummary` 建模 target/query 双轴路径;`AlignBedScore` 同时保留有限数值 score 与无空白文本 score。parser 支持 BED3-BED12、CRLF和通用空白,其中 BED3-BED9按连续单块解释,BED12重建完整exon path;BED10仅接受单块,BED11仅在缺失的blockStarts可无损推断时接受。 + +坐标统一为 zero-based half-open boundary。target规范化为递增轴,plus query递增、minus query递减;BED12的query size由blockSizes求和,target-only segment表示intron。`blocks` 从一般路径投影aligned segments,`map_target_position`/`map_query_position` 仅映射aligned residue并在intron或不可表示gap返回 `None`。`counts` 汇总aligned bases、双轴skip bases/open和columns,document提供half-open `search`、有序target/query集合及跨记录summary。 + +严格校验覆盖3-12列、非负32位整数、有限score、strand、thick interval、block count/list长度、正size、首尾span、排序/重叠/越界、坐标单调性、零长度segment和aligned step等长。writer可从一般target/query path生成canonical blocks,并按 `bed_columns` 输出 BED3-BED12;query-only segment按BED可表达能力被投影跳过,反向target path先规范化。该层处理文本pairwise alignment,`bigbed.mbt` 继续独立处理压缩二进制存储和索引。 + +### 276. 差异空间细胞共定位分析 (Bioconductor spicyR) + +实现 Bioconductor `spicyR` 1.25.0 的图像级差异空间共定位核心。每张图像对有序细胞类型对计算 cross-K/cross-L 曲线,排除细胞自身匹配并保留上游同型对 `n²` intensity normalization;半径限制在最短窗口跨度约一半并折叠重复截断值。图像统计量按上游语义累加 `sum(L(r)-r)`,缺失类型对可省略或用 Poisson 基线补齐,单个同型细胞保留为可用基线。 + +坐标窗口由每张图像的范围和可选 padding 构造;边界校正使用以源细胞为圆心的可见圆面积倒数,因此有方向性。矩形与圆的交面积通过固定 96 分片 Simpson 积分确定性求解。图像 precision weights 由同型 `n(n-1)` 或异型 harmonic effective count 构造、归一化至均值一并应用 `weight_factor`;这是不依赖 R/scam 的可移植近似,不复刻上游 GAM 平滑。 + +无重复受试者时拟合加权线性模型,有重复受试者时复用 `variancePartition` 的 ML/REML 随机截距模型、数值 Satterthwaite 自由度和 Student-t 检验;支持数值协变量、多条件相对参考组对比及跨细胞类型对 BH-FDR。API 可直接分析细胞坐标、拟合预计算 association,或从 `SpatialExperiment` 的 `colData`/`spatialCoords` 提取输入并在不可变副本 metadata 中写回结果;另提供 pair/condition/image 查询、排序、显著性过滤、摘要和合成重复测量数据。 + +### 277. 局部空间关联曲线与组织区域聚类 (Bioconductor lisaClust) + +实现 Bioconductor `lisaClust` 1.21.0 的局部空间统计与区域发现核心。模块按图像独立处理每个源细胞,对所有目标细胞类型和半径累积排除自身匹配的局部邻居;目标类型强度为 `n/area`,期望值为 `πr² × visible_fraction × intensity`。标准化 local-K 输出 `(observed-expected)/sqrt(expected)`,centered local-L 输出 `sqrt(observed)-sqrt(expected)`,并通过 Gaussian KDE 的均值归一化逆密度权重校正不均匀采样,支持密度下限、有限值归零和按窗口短边约 `1/2.01` 截断半径。 + +空间窗口支持带 padding 的矩形和 monotonic-chain 凸包;圆盘位于凸窗口内的可见面积使用固定角度中点积分确定性计算。完整 cell × cell-type × radius 曲线矩阵进入确定性多起点 k-means,采用 farthest-first 初始化、空簇恢复和稳定标签规范化,同时输出 inertia、silhouette、cluster sizes、centroids 和收敛诊断。区域富集按 `observed / (cell_type_total × region_total / N)` 计算,并提供细胞区域查询、富集检索/排名、区域摘要及预计算曲线聚类入口。 + +`lisaclust_spatial_experiment` 从 `colData` 与 `spatialCoords` 提取图像、细胞类型和可选 cell ID,在深复制的 `SpatialExperiment` 中写入区域列及细胞数、特征数、区域数和 silhouette metadata,原容器保持不变。可移植实现不支持上游 concave window;KDE 在细胞位置直接计算而非复刻 `spatstat::density.ppp` 像素栅格;圆盘相交面积使用角度积分近似;不同图像截断后的 effective radii 单独记录,但保留用户请求半径对应的统一特征列。 + +### 278. 背景感知空间混合细胞解卷积 (Bioconductor SpatialDecon) + +实现 Bioconductor `SpatialDecon` 1.23.0 的核心混合细胞解卷积流程。输入为 gene × spot 表达矩阵、同维背景与 precision-weight 矩阵,以及 gene × cell-type profile;数据和 profile 按基因名对齐,并要求共享基因数不少于细胞类型数。profile 可按全矩阵指定分位数缩放至目标值,默认复刻上游 `2 / Q0.99(X)` 尺度。每个 spot 拟合非负丰度 `β`,最小化 `Σ wᵢ(log(yᵢ)-log(bᵢ+max(Xᵢβ,10⁻⁴)))²`,因此直接建模加性技术背景上的 log-normal 生物信号。 + +优化器使用确定性投影阻尼 Newton、partial-pivot Gaussian elimination 和 backtracking line search;observed Hessian 非正定、奇异或不产生下降方向时,回退到 expected-Hessian 缩放梯度。两阶段流程先拟合全部基因,再以 `log2(max(y,lower))-log2(max(fitted,lower))` 标记超过阈值的数据点并重拟合;保留数据不足以识别模型时自动取消剔除。协方差优先取带 ridge 的 observed Hessian 逆,在无效时回退 expected Hessian,由此输出标准误、t 统计量和正态近似双侧 p 值。 + +结果同时提供 abundance、spot 内 proportion、按最大 spot 总丰度归一化的 cells-per-100,以及结合 nuclei count 的细胞数尺度,并记录 fitted、log2 residual、outlier mask、RMSE、相关性、目标值和收敛诊断。`collapse_spatial_decon` 通过 `β'=Aβ` 与 `Σ'=AΣAᵀ` 合并细胞类型;`reverse_spatial_decon` 对每个 gene 拟合非负 intercept 与变化 cell scores;辅助 API 可按 probe pool 平均负探针推导 GeoMx 背景,也可从单细胞 count matrix 经细胞/基因过滤、可选 library normalization 和 cell-type 均值构建 profile。 + +`spatial_decon_spatial_experiment` 从 assay、rowData 和 colData 提取输入,在深复制容器的 colData 中写入 abundance/proportion,并保留原对象不变。MoonBit 版本以确定性投影 Newton 代替 R `optim(method="L-BFGS-B")`;不在线下载上游约 75 个 profile matrix,不内置 `safeTME` 数据,也不包含 pure-tumor profile 推断和 tumor clustering。不同 row 维度的细胞丰度不会作为 assay 写入,而以 spot 级 colData 字段保存。 + +### 279. 通用单细胞聚类与诊断 (Bioconductor bluster) + +实现 Bioconductor `bluster` 1.23.0 的通用聚类、图构建与诊断核心,矩阵统一采用 observation × variable 方向。精确近邻支持 Euclidean、Manhattan 和 cosine 距离,排除自身、按观察索引稳定处理距离 ties,并在 `k > n-1` 时截断。KNN 图可保持有向关系或对任一方向近邻边进行对称化;SNN 图把自身作为 rank 0 邻居,支持上游 rank 权重 `max(k-(rᵢ+rⱼ)/2,10⁻⁶)`、共享邻居数和 Jaccard 权重。 + +K-means 使用 K-means++ 初始化、多起点 Lloyd 迭代、空簇恢复和固定 seed 的 Park-Miller RNG,返回质心、逐簇/总 within-cluster sum of squares 与迭代诊断。图聚类使用无需 R、igraph 或 C++ 的 seed 可复现 Louvain 风格局部 modularity 优化,并支持 resolution;two-step 流程先以 K-means 向量量化观察,默认使用 `round(sqrt(n))` 个质心,再在质心 SNN 图上聚类并将标签映射回原观察。 + +诊断 API 覆盖 pairwise Rand 分解与 adjusted Rand index、RMS distance 的 approximate silhouette、cluster RMSD、以第 k 邻居距离中位数为半径的 neighbor purity、pairwise modularity、贪心 community merging、nested cluster mapping、minimum/maximum/union cluster correspondence,以及 bootstrap K-means stability。`bluster_cluster_sce` 从 `SingleCellExperiment.reduced_dims` 读取 cell × dimension 坐标,深复制 assays、row/column metadata、reduced dimensions、metadata 和 alternative experiments 后写入聚类标签,输入对象保持不变。当前实现使用稠密精确近邻和串行局部优化,不覆盖上游 BiocNeighbors 近似索引、igraph/cluster_leiden 后端或 BiocParallel 调度。 + +### 280. 拓扑自组织映射与meta-clustering (Bioconductor FlowSOM) + +实现 Bioconductor `FlowSOM` 2.21.0 的拓扑聚类核心,输入统一为 cell × marker。距离支持 Manhattan、Euclidean、Chebyshev 和 cosine,并定义零向量的稳定 cosine 语义;码本可由无放回 random、最远点迭代 KWSP 或 covariance/power-iteration PCA 网格初始化。训练从二维规则网格的 Chebyshev 邻域开始,按 stage 线性衰减学习率与半径,每个 stage 后以码本 Euclidean 距离构造确定性 Prim MST,并用无权最短路径重建下一阶段邻域。固定 seed 的 Park-Miller RNG 保证初始化、抽样和训练可复现。 + +训练结果包含 BMU 与距离、完整 MST、拓扑距离、quantization/topographic error,以及基于原始未加权 marker 值的节点 counts、percentages、median fluorescence intensity、sample SD、CV 和 R 风格 `1.4826 × MAD`。meta-clustering 复用 `bluster` 多起点 K-means;自动模式平滑不同 k 的 within-cluster SSE,并以双线性拟合残差选择 elbow。辅助 API 覆盖新数据投影、节点距离 MAD outlier、marker 级双侧 outlier、节点阳性比例、meta counts/medians、weighted/unweighted purity 和 F-measure。 + +`flowsom_train_flow_frame` 保持 flowCore event × marker 方向并支持 marker 子集;`flowsom_cluster_sce` 将 SingleCellExperiment 的 marker × cell assay 转置后训练,在递归深复制的容器中写入 1-based `FlowSOM.cluster` 和 `FlowSOM.metacluster`,不修改输入对象。实现不依赖 R、C/C++、igraph 或 ConsensusClusterPlus;当前采用稠密矩阵、串行在线更新和 K-means meta-clustering,不覆盖上游并行后端、共识聚类插件及可视化层。 + +### 281. 约束谱系与同时主曲线 (Bioconductor slingshot) + +实现 Bioconductor `slingshot` 2.21.0 的高级轨迹推断核心,坐标统一为 cell × dimension,并保留原 `slingshot.mbt` 兼容API。输入支持hard cluster label和逐行归一化的soft membership;cluster摘要支持weighted mean/median、有效样本分母协方差和ridge稳定化,cluster距离支持质心Euclidean、pooled diagonal缩放及full Mahalanobis,并在矩阵不可逆时回退到diagonal。 + +谱系推断使用稳定排序的确定性Kruskal,支持start cluster、end cluster叶节点约束、固定`omega`阈值和基于无限制MST边长中位数的自动`omega` forest。每个connected component独立选择root并枚举root-to-leaf lineage。曲线可使用none、line或endpoint cluster PC1扩展,通过arc-length重采样、折线投影和Gaussian-kernel local-linear smoother迭代拟合;cell × lineage距离秩用于`1-rank²`重加权和重新分配,共享cluster前缀使用cosine taper同时收缩。 + +结果提供cell × lineage Optional pseudotime、weights、平均pseudotime、1-based branch ID及收敛诊断。`slingshot_predict`将新cell投影到已拟合曲线并按训练距离90%分位数衰减权重;`slingshot_advanced_sce`递归深复制SingleCellExperiment及alternative experiments,再写入branch、pseudotime和weight字段,输入对象保持不变。当前实现采用稠密矩阵与串行平滑,不覆盖上游S4/PseudotimeOrdering容器、BiocParallel后端和可视化层。 + +### 282. 多谱系轨迹差异表达 (Bioconductor tradeSeq) + +实现 Bioconductor `tradeSeq` 1.27.0 的高级轨迹差异表达核心,并保留 `tradeseq.mbt` 的基础兼容API。输入遵循上游 gene × cell非负整数计数、cell × lineage pseudotime和weight方向;每个cell的lineage weight逐行归一化,正权重对应的拟时间必须有限且每条lineage必须覆盖非零范围。默认offset由文库大小的中心化log值生成,也可显式传入。上游随机cell-to-lineage分配被替换为确定性weighted expansion,使每个cell对所有活跃lineage的总贡献严格为1。 + +每条lineage使用独立open-uniform B-spline coefficient block,degree为`min(3, nKnots-1)`;联合负二项IRLS使用`Var(Y)=mu+phi*mu^2`,并加入二阶差分平滑惩罚与ridge稳定化。dispersion通过Pearson moment迭代更新并按配置截断;最终模型保留系数、penalized information逆矩阵协方差、gene × cell拟合值、NB log-likelihood、AIC、迭代次数及收敛状态。`tradeseq_evaluate_k_advanced`对候选knot数逐一重拟合,返回gene × candidate AIC及mean-AIC选择结果。 + +Wald contrast引擎通过Gram-Schmidt去除线性相关行,提供`associationTest`、`startVsEndTest`、`diffEndTest`、`patternTest`和`earlyDETest`,并统一支持log2 fold-change阈值、chi-square tail probability与gene-level BH-FDR。平滑预测返回各lineage原始拟时间网格、均值和delta-method标准误;`tradeseq_fit_from_slingshot`直接消费Slingshot Optional pseudotime/weight,`tradeseq_advanced_sce`递归深复制SingleCellExperiment及alternative experiments后写回拟合assay、检验p值/FDR、dispersion和metadata。当前实现采用稠密串行线性代数,不依赖R、mgcv、edgeR或BiocParallel,也不覆盖上游零膨胀模型和可视化层。 + +### 283. 多样本多亚群差异状态与检测 (Bioconductor muscat) + +实现 Bioconductor `muscat` 1.27.4 的高级多样本、多亚群分析核心,并保留 `muscat.mbt` 的基础兼容API。输入严格采用 gene × cell非负整数计数,验证sample、cluster和group元数据长度、非空ID、样本到组的一对一关系,以及gene/cell名称的数量和唯一性。cluster × sample伪批量支持sum、mean、median、proportion detected和number detected五类聚合;缺失组合补零,并同时保留cell count与library size。 + +设计层支持与伪批量样本顺序完全一致的任意有限满秩矩阵、默认reference-coded group design及多个contrast。差异状态(DS)模型使用cluster-sample library-size offset、负二项IRLS和gene-wise Pearson dispersion,并向cluster内median trend执行经验贝叶斯收缩。Wald检验支持log2 fold-change阈值;结果同时提供每个cluster/contrast内的local BH-FDR和每个contrast跨cluster的global BH-FDR。 + +差异检测(DD)遵循上游CDR归一化语义:先过滤median detection fraction达到阈值的普遍检测基因,再使用`log(nCells × mean detection fraction)` offset拟合检测计数。DS与DD结果可通过harmonic-mean screening和两假设confirmation组合为`DS`、`DD`、`both`、`screen_only`或`none`。`muscat_advanced_sce`从SingleCellExperiment读取assay和cell metadata,递归深复制assay、row/col data、reduced dimensions、metadata及alternative experiments后写回cluster级logFC、FDR和stagewise分类,不修改输入对象。 + +当前实现是无外部依赖的MoonBit NB-Wald流程,不调用或宣称复刻edgeR、DESeq2、limma、MAST及stageR的R后端;采用稠密串行线性代数,也不覆盖上游并行、随机效应模型和可视化层。 + +### 284. 现代 XMFA 多基因组比对读写与坐标投影 (Biopython Bio.Align.mauve) + +实现 Biopython 1.86 `Bio.Align.mauve` 的现代 XMFA 文档语义,并将历史 `mauve.mbt` 明确保留为 MAF-like 重排分析兼容层。严格 parser 支持 `#FormatVersion`、有序 `#SequenceNFile/Entry/Format` metadata、combined-file 与 separate-file identifier、`> N:start-end strand description` 行、wrapped alignment、CRLF、`=` 块终止符及多个 locally collinear block;同时校验 source 编号、entry、名称唯一性、比对宽度、IUPAC 字符和 `0-0` 全 gap 行。 + +文件中的 1-based inclusive 区间统一转换为内部 0-based half-open 区间。每行保留完整 boundary coordinate path:正链递增,负链递减;compact path 在任一行 residue/gap movement 状态变化时保留边界,与 Biopython printed-alignment 坐标一致。API 提供边界/残基/列查询、LCB 与文档级 pairwise identity/mismatch/gap/gap-open 统计、half-open interval index,以及跨序列位置和区间投影;区间映射会在 gap、列不连续或坐标步长变化处拆分,并保留目标链方向。 + +序列重建将各 LCB 的 source-oriented segment 写回 forward coordinates,未覆盖位置使用可配置字符填充,重叠冲突会返回类型化错误。canonical writer 支持固定宽度折行并可稳定 round-trip;公开数组访问器递归复制 block、row 与 coordinate path,避免调用方修改文档内部状态。专项测试采用 Biopython 官方 `combined.xmfa`、`separate.xmfa` 和 simple fixtures 的 metadata、负链坐标及 compact path 语义,并覆盖 wrapped/CRLF、统计、索引、投影、重建、写回和错误边界。 + +### 285. CODEML控制文件、原生输出与模型比较 (Biopython Bio.Phylo.PAML.codeml) + +实现 Biopython 1.86 `Bio.Phylo.PAML.codeml` 中不依赖外部可执行文件的完整离线工作流,并保留旧 `paml.mbt` 作为明确标注的内置近似兼容层。control parser 支持 `seqfile`、`outfile`、`treefile` 和 Biopython Codeml option集合,处理星号行尾注释、CRLF与多值`NSsites`;严格拒绝缺失路径、重复键、未知option、空值、错误数值类型及负NSsites,并提供稳定的canonical writer。 + +结果解析器识别CODONML和AAML header、PAML版本、模型与codon-frequency metadata、序列/位点数量,以及多个NSsites或gene结果。类型化模型覆盖lnL/np/参数/SE、主树与dN/dS/omega树、kappa/omega、branch表、site-class比例、branch-site A前景/背景类别、clade model C branch types、free-ratio、多基因relative rates和gene-wise参数;同时支持pairwise dN/dS、AAML raw/ML下三角距离矩阵以及BEB/NEB阳性位点。 + +分析层提供AIC、BIC、最小AIC选择和嵌套模型似然比检验。LRT使用Lanczos log-gamma与regularized incomplete gamma Q计算chi-square尾概率,并验证参数嵌套、似然单调性、显著性阈值及观测数边界。PAML输出中的`nan`分支量以`Double?`保留,公开数组访问器执行深复制。本模块只解析和分析已有artifact,不启动、捆绑或伪装外部`codeml`二进制;60项黑盒测试改编自Biopython PAML 4.1-4.9 fixtures,覆盖成功路径、数值信号、复制语义和严格错误边界。 + +### 286. BASEML核苷酸替换模型、非齐次频率与率类别 (Biopython Bio.Phylo.PAML.baseml) + +实现 Biopython 1.86 `Bio.Phylo.PAML.baseml` 中可确定性测试的离线工作流。control parser覆盖`seqfile`、`outfile`、`treefile`和全部Baseml option,支持星号注释、CRLF以及model 9/10附加rate-group定义;规范化整数和浮点表示,严格拒绝缺失路径、重复/未知键、空值、非法model、错误类型、无效`ncatG/nparK/nhomo`及越界二元开关。`baseml_model_name`提供JC69、K80、F81、F84、HKY85、T92、TN93、REV、UNREST、REVu和UNRESTu的完整编号映射。 + +结果解析器兼容PAML 4.1-4.7 header与base-frequency布局,提取unconstrained/fitted lnL、np、完整参数向量、SE、tree length和branch-length Newick。类型化结果进一步覆盖单值/多值/branch-specific kappa、REV/UNREST rate parameters与4x4 Q矩阵、平均Ts/Tv、离散gamma alpha/rates/frequencies、auto-dGamma rho与rate-category transition matrix,以及`nhomo`节点frequency parameters、可选realized T/C/A/G频率和root标记。矩阵维度、参数/SE长度、节点唯一性和必需估计均执行严格验证,所有公开数组和嵌套矩阵访问器均返回深复制。 + +分析层提供AIC、BIC和嵌套模型LRT,复用经过CODEML测试的Lanczos log-gamma与regularized incomplete gamma Q实现chi-square尾概率,同时校验参数自由度、似然单调性、观测数和alpha边界。本模块不启动或捆绑外部`baseml`二进制;64项黑盒测试改编自Biopython官方model、SE、alpha1rho1、nparK、nhomo及4.1-4.7 fixtures,覆盖control round-trip、数值信号、复制语义和错误边界。 + +### 287. YN00成对密码子替换估计与跨版本结果解析 (Biopython Bio.Phylo.PAML.yn00) + +实现 Biopython 1.86 `Bio.Phylo.PAML.yn00` 的离线control和结果工作流。control parser覆盖`seqfile`、`outfile`、`verbose`、`icode`、`weighting`、`commonf3x4`和`ndata`,支持星号注释及CRLF;严格拒绝缺失路径、重复或未知键、空值、多重等号、非整数option、越界开关和非法遗传密码编号,并提供稳定的canonical writer。 + +结果状态机兼容PAML 4.1-4.9i输出差异,解析NG86、Yang-Nielsen 2000、LWL85、modified LWL85和LPB93五种方法。类型化结果保留dN、dS、omega、kappa、时间、同义/非同义位点、标准误和rho;支持旧式无空格NG86矩阵、相邻负值、长名称及带点数字名称,并将`nan`、`inf`和Windows `-1.#IND`安全映射为未定义值。解析过程验证序列索引和名称一致性、名称唯一性、三角pair数量、方法完整性、重复记录及矩阵维度。 + +分析API提供无序序列对查询、任意方法/统计量的对称矩阵和跳过未定义值的pair均值;公开名称、pair和嵌套矩阵访问器均执行防御复制。本模块只解析和分析已有YN00 artifact,不启动或捆绑外部`yn00`二进制;80项黑盒测试覆盖Biopython官方跨版本fixture语义、数值信号、名称边界、非有限值、复制语义和严格错误路径。 + +### 288. 现代BLAST XML1/XML2解析、坐标路径与规范写回 (Biopython Bio.Blast) + +实现 Biopython 1.86 `Bio.Blast` 的现代离线XML工作流,并保留`blast.mbt`作为历史基础解析兼容层。递归XML parser支持declaration、comment、DOCTYPE、CDATA、自闭合元素、namespace和entity,严格检查标签配对、尾随数据、唯一header、query/report一致性及必需字段。类型化文档覆盖XML1和XML2、多query、普通XML2多report、PSI-BLAST iterations、多个HitDescr、taxonomy、Parameters及Karlin-Altschul Statistics。 + +HSP层将BLAST的1-based inclusive字段转换为0-based boundary coordinate path。`blastn`和`megablast`按strand/frame递增或递减;`blastp`、`rpsblast`和`psiblast`使用蛋白质1:1轴;`blastx`、`tblastn`和`tblastx`在翻译的核酸轴上按每个氨基酸3个碱基移动,并保留`coded_by`正向或`complement(...)`位置。路径在match、query gap或target gap状态切换处压缩,同时校验alignment长度、midline、gap/identity/positive计数、raw span、frame范围、方向和序列边界。 + +canonical writer支持同dialect写回和显式XML1/XML2转换;XML2普通多query按XSD生成多个`BlastOutput2` report,PSI-BLAST使用`Results/iterations/Iteration`。无法无损表达的转换会返回类型化错误,writer结果在返回前由严格parser复验。公开record/hit/HSP、描述、mask和坐标数组执行防御复制,并提供query查找、best hit/HSP、E-value过滤和文档汇总。61项黑盒测试覆盖八类BLAST程序、XML实体、多描述、反链和双翻译坐标、规范往返、跨dialect转换、`NaN`/Infinity及结构错误;模块不访问NCBI网络,也不启动本地BLAST可执行文件。 + +### 289. 现代CLUSTAL多序列比对、列注释与坐标模型 (Biopython Bio.Align.clustal) + +实现 Biopython 1.86 `Bio.Align.clustal` 的现代单MSA工作流,并保留`clustal_io.mbt`作为历史`Bio.AlignIO.ClustalIO`兼容层。parser识别CLUSTAL、PROBCONS、MUSCLE、MSAPROBS、Kalign和Biopython generator header,支持CRLF、interleaved blocks、可选累计ungapped residue count以及稀疏、全空格或缺失的consensus。第一块建立唯一identifier顺序,后续块必须严格复现该顺序和等宽片段;累计计数、字符集、列宽、全gap行与全gap列均执行类型化校验。 + +`AlignClustalAlignment`保留原始generator/version、ungapped和aligned row以及`*`、`:`、`.`列注释。坐标API提供零起始residue/column双向查询、跨行位置投影、逐列aligned pairs和按movement vector压缩的Biopython风格多行coordinate path;分析层提供pair/all-pairs identity与gap-open统计、occupancy、majority consensus,并按标准ClustalW strong/weak蛋白质保守组计算列符号,核酸则只标记无gap完全一致列。 + +canonical writer支持配置block/name width、累计残基数和consensus,拒绝静默截断长identifier,并在返回前用严格parser复验输出。构造器和数组结果执行防御复制。75项黑盒测试覆盖官方fixture风格header、block和计数语义、坐标/统计信号、规范往返、复制语义及失败边界;离线示例不启动CLUSTAL、MUSCLE或其他外部aligner。 + +### 290. 现代PHYLIP布局识别、固定宽度名称与坐标模型 (Biopython Bio.Align.phylip) + +实现 Biopython 1.86 `Bio.Align.phylip` 的现代单MSA工作流,并保留`phylip_io.mbt`作为历史`Bio.AlignIO.PhylipIO`兼容层。parser严格读取两个正整数header字段和首块固定10列identifier,支持CRLF、wrapped sequential、interleaved blocks及片段内分组空格,并按Biopython首块规则自动识别布局。每块行数、片段宽度、累计列数、尾部数据和声明维度都会校验;`.`、全gap列、不完整block和溢出维度返回类型化错误,同时保留空、内部空格或重复identifier的合法fixed-width语义。 + +`AlignPhylipAlignment`同时保存aligned与ungapped rows及来源布局。坐标API提供零起始residue/column双向查询、跨行位置投影、逐列aligned pairs和按movement vector压缩的多行coordinate path;统计层提供pair/all-pairs identities、mismatches、gap columns、double-gap columns、gap opens、occupancy和majority consensus。 + +writer默认生成与现代Biopython一致的单行sequential格式,也可配置wrapped sequential或interleaved blocks及分组宽度。identifier写出前执行trim、标点替换和10字符截断,允许上游兼容的规范化碰撞;输出返回前由严格parser复验。构造器及公开数组执行防御复制。72项黑盒测试覆盖官方fixture风格布局、名称边界、坐标/统计、稳定往返、复制语义和失败路径;离线示例不启动外部PHYLIP程序或aligner。 + +### 291. Exonerate C4文本结果、剪接语义与坐标投影 (Biopython Bio.SearchIO.ExonerateIO.exonerate_text) + +实现 Biopython 1.86 `Bio.SearchIO.ExonerateIO.exonerate_text` 的 C4 人类可读报告层,并与既有`exonerate.mbt` vulgar/cigar摘要兼容API及`align_exonerate.mbt` operation path模型并存。严格parser读取`Command line`、`Hostname`、`C4 Alignment` header和completion marker,将多query、hit及HSP聚合为不可变`Document -> QueryResult -> Hit -> HSP -> Fragment`层次;query/target描述、model、raw score、header boundary和alignment body均保留类型语义。 + +alignment body通过coordinate row的冒号位置确定固定列宽,支持跨physical block拼接及3、4、5行模型。translated模型解析protein/DNA annotation、三字母氨基酸、leading/trailing partial triplet和codon flip;特殊区域识别target/query/joint/reverse intron、NER、split codon与frameshift,并分别生成双轴fragment及inter-range。坐标统一为0-based half-open区间,保留forward/reverse/protein strand、phase、frame及每轴step,支持fragment residue位置投影、HSP区间/计数、query/hit查找和best-HSP选择。 + +parser严格检查metadata顺序、固定列行宽、坐标方向与终点、模型维度、splice/NER长度、split codon配对和frameshift归属;公开层次查询与数组结果执行深层防御复制。42项黑盒测试覆盖wrapped blocks、正反链intron、joint intron、NER、protein/DNA翻译、5行coding模型、partial phase、特殊氨基酸、聚合及畸形输入;示例完全离线,不调用外部`exonerate`程序。 + +### 292. 空间自相关统计与变差函数建模 (Bioconductor Voyager) + +实现 Bioconductor `Voyager` 的单变量、双变量与多变量空间自相关核心,矩阵统一采用 spot × feature 之外的独立向量视图。空间权重由 `voyager_weights_knn`(kNN,`k` 超过 `n-1` 时截断,距离 ties 按邻居索引稳定排序)、`voyager_weights_distance_band`(固定半径,含自身排除)或 `voyager_weights_inverse_distance`(`w_ij = 1/d^power`,可选带宽上限)从二维/三维坐标构建,并支持四种编码风格:`W` 行标准化、`B` 二值、`C` 全局标准化(总和为 1)和 `S` Caussinus–Mestre(`1/√(k_i·k_j)`)。权重矩阵以邻居索引数组的稀疏形式存储,S0/S1/S2 由 `voyager_weights_s0/s1/s2` 按定义确定性计算。 + +全局统计提供 Moran's I 和 Geary's c,均采用 Cliff–Ord (1981) 随机化期望与方差(`b2 = n·m4/m2²`),并以 Abramowitz–Stegun 正态 CDF 计算双侧 p 值;方差非正时回退为 0。局部统计覆盖 Anselin (1995) 局部 Moran's I(LISA)及 HH/HL/LH/LL 象限分类、局部 Geary's c(similar/dissimilar 分类)、Getis–Ord Gi/Gi*(`star=true` 时含自身 `w_ii=1`)及精确 Ord–Getis 随机化 z 分数与 hotspot/coldspot 分类。双变量 Lee's L 同时返回全局 L 与逐点局部 L;`voyager_multivariate_local_geary` 对多特征矩阵按每特征独立置换并汇总逐点统计。置换推断使用固定 seed 的 splitmix64 PRNG 驱动 Fisher–Yates 洗牌,p 值采用 `(extreme+1)/(perm+1)` 校正,并通过 Benjamini–Hochberg 步降控制 FDR。 + +经验变差函数将点对按等距 lag 分箱(默认上限为最大成对距离一半)并计算半方差 `γ = Σ(x_i-x_j)²/(2·n_pairs)`;`voyager_fit_variogram` 在有界 range 网格上搜索、对每个候选 range 用闭式线性最小二乘求解 (nugget, partial sill),最小化残差平方和,支持 spherical/exponential/gaussian 三种模型,`voyager_variogram_predict` 据此预测任意距离的半方差。`voyager_correlogram` 按距离分箱逐 bin 构建行标准化权重并计算 Moran's I,自动跳过无观测对的 bin。`voyager_run_univariate_sfe` 从 `SpatialExperiment` 的 assay、`spatialCoords`、`rowData`/`colData` 提取输入,在深复制容器中把全局统计写入 `rowData`、局部统计(local estimate/FDR/quadrant)写入逐 spot `colData`,并在 metadata 记录方法与特征数,原对象保持不变;基因名按 `gene_name`→`gene_id`→`gene_N` 回退解析,三维坐标在 z 非恒定时自动启用。当前实现不依赖 R、spdep 或 sf,采用稠密成对距离与串行计算,不覆盖上游 `listw`/`nb` S4 对象、并行后端、协变量残差化与可视化层;58 项黑盒测试覆盖手算链状格点、置换确定性、FDR 单调性与 SpatialExperiment 不可变性。 + +### 293. 高级密码子比对与选择压力检验 (Bio.codonalign advanced) + +实现 Biopython `Bio.codonalign` 模块中的高级选择压力分析功能,基于 Nei–Gojobori (1986) 框架。`codon_test_selection` 执行 Z-test 选择检验,使用 NG86 大样本近似方差(`V(dN) = pN(1−pN) / [Nn·(1−4pN/3)²]`,`V(dS)` 同理,协方差忽略),Z = (dN−dS)/√V(dN−dS),支持三种备择假设:正选择(H₁: dN > dS,单尾 p = 1−Φ(Z))、净化选择(H₁: dN < dS,单尾 p = Φ(Z))、中性(H₁: dN ≠ dS,双尾 p = 2(1−Φ(|Z|))),在 α=0.05 水平给出结论。`codon_test_neutrality` 执行 Fisher 精确检验,构建 2×2 列联表(行:非同义/同义;列:差异/相同),支持 two-sided/greater/less 三种检验方向,返回 p 值与优势比。 + +`build_codon_alignment` 从蛋白质比对和未比对的编码序列构建密码子比对:逐位扫描蛋白质比对,遇 gap (`-`/`.`) 插入 `---`,否则消费编码序列的下一个密码子并验证翻译与蛋白质残基一致。`sliding_window_dnds` 以可配置窗口大小和步长沿密码子比对滑窗,逐窗计算 NG86 dN/dS 以检测选择热点,自动跳过不足 3 个有效密码子对的窗口。`codon_align_advanced_bh_fdr` 实现 Benjamini–Hochberg FDR 校正(步降法,`q_i = min(q_{i+1}, p_i·m/rank)`)。`pairwise_kaks_table` 对多条序列执行成对 Z-test 正选择检验,对所有 p 值统一 BH-FDR 校正,按阈值给出显著性结论。32 项黑盒测试覆盖正/净化选择检测、Fisher 检验多方向、构建器 gap 处理与翻译验证、滑窗边界与退化窗口跳过、FDR 单调性、成对表完整性与错误输入拒绝。 + +### 294. 高级蛋白质序列预测 (Bio.SeqUtils advanced) + +实现六种经典经验蛋白质序列分析算法,覆盖二级结构、无序区、卷曲螺旋、抗原性、表面可及性和柔柔性预测。`chou_fasman_predict` 使用 Chou–Fasman (1974) 氨基酸倾向值表(Pα 螺旋、Pβ 折叠、Pt 转角)进行二级结构预测:先扫描螺旋成核位点(≥6 残基均值 Pα > 1.03 且 Pα > Pβ)和折叠成核位点(≥3 残基均值 Pβ > 1.05 且 Pβ > Pα),再向两侧延伸至倾向值低于 1.0,转角区域由 4 残基窗口 Pt > 1.0 且 Pα、Pβ < 1.0 识别,螺旋/折叠冲突按区域均值倾向值高者优先裁决,输出逐残基 H/E/T/C 预测及各结构区域起止。 + +`iupred` 使用 Dosztányi 等 (2005) 的成对相互作用能矩阵估计每残基在滑动窗口内的能量,通过 logistic 变换 `1/(1+exp(-(E+0.45)·4))` 映射到 [0,1] 无序分,阈值 0.5 以上判为无序,连续 ≥5 残基无序归为一个无序区段;`iupred_long` (window=100) 和 `iupred_short` (window=25) 分别提供全局和局部模式。`predict_coiled_coils` 使用 Lupas 等 (1991) 的七肽重复 (a–g) 评分矩阵,位置 a 和 d(疏水核心)权重 2.5×,在滑动窗口内尝试全部 7 种读框取最高分,归一化分超过阈值(默认 0.9)判为卷曲螺旋区。`kolaskar_tongaonkar_antigenicity` 以 7 残基滑窗计算 Kolaskar–Tongaonkar (1990) 抗原倾向均值,高于全序列均值的连续 ≥6 残基区域为抗原位点。`emini_surface_accessibility` 按 Emini 等 (1985) 公式计算滑窗表面概率(中心残基权重 2×),`karplus_schulz_flexibility` 以归一化 B 因子参数的滑窗均值估计链柔柔性,值 >1.0 表示高于平均柔柔性。26 项黑盒测试覆盖各算法的正常用例、边界条件(短序列、空序列、非标准残基拒绝)、输出范围验证与大小写不敏感处理。 ## 性能优化 @@ -2801,8 +3681,8 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 | 指标 | 数值 | | :--- | :---: | -| 总测试数 | 8105 | -| 通过数 | 8105 | +| 总测试数 | 12209 | +| 通过数 | 12209 | | 失败数 | 0 | | 通过率 | 100% | @@ -2813,10 +3693,10 @@ moon test --package IvanAXu/BioSeqs/test/moonbit # ✅ 8105 个测试全 moon build # 运行所有测试 -moon test --package IvanAXu/BioSeqs/test/moonbit +moon test # 运行单个模块测试 -moon test --package IvanAXu/BioSeqs/test/moonbit --test bio_seq_test +moon test test/moonbit/bio_seq_test.mbt # 更新快照测试 moon test --update @@ -2842,6 +3722,7 @@ moon test --update | 特征提取 | `feature_extraction_test.mbt` | 19 | | Biostrings | `biostrings_test.mbt` | 21 | | GenomicRanges | `genomic_ranges_test.mbt` | 22 | +| GRangesList | `granges_list_test.mbt` | 38 | | plyranges | `plyranges_test.mbt` | 15 | | DESeq2 | `deseq2_test.mbt` | 10 | | dplyr | `dplyr_test.mbt` | 9 | @@ -2866,6 +3747,8 @@ moon test --update | edgeR | `edger_test.mbt` | 7 | | limma | `limma_test.mbt` | 10 | | SummarizedExperiment | `summarized_experiment_test.mbt` | 7 | +| RangedSummarizedExperiment | `ranged_summarized_experiment_test.mbt` | 16 | +| TreeSummarizedExperiment | `tree_summarized_experiment_test.mbt` | 11 | | IRanges | `iranges_test.mbt` | 14 | | AlignIO | `align_io_test.mbt` | 12 | | Cluster | `cluster_test.mbt` | 12 | @@ -2898,7 +3781,37 @@ moon test --update | NeighborSearch | `neighbor_search_test.mbt` | 6 | | BiocNeighbors | `bioc_neighbors_test.mbt` | 61 | | SwissProt | `swissprot_test.mbt` | 8 | +| Cellosaurus | `cellosaurus_test.mbt` | 16 | +| UniGene | `unigene_test.mbt` | 23 | +| Bio.Align.hhr | `hhr_test.mbt` | 33 | +| Bio.Align shared-reference merge | `shared_reference_alignment_test.mbt` | 41 | +| Bio.Align Alignment.map/mapall | `alignment_map_test.mbt` | 44 | +| Bio.Align Alignment.counts | `alignment_counts_test.mbt` | 57 | +| Bio.Align.tabular | `align_tabular_test.mbt` | 90 | +| Bio.Align.psl | `align_psl_test.mbt` | 93 | +| Bio.Align.sam | `align_sam_test.mbt` | 94 | +| Bio.Align.a2m | `a2m_test.mbt` | 71 | +| Bio.Align.emboss | `align_emboss_test.mbt` | 81 | +| Bio.Align.exonerate | `align_exonerate_test.mbt` | 86 | +| Bio.Align.msf | `msf_test.mbt` | 95 | +| Bio.Align.nexus | `align_nexus_test.mbt` | 100 | +| Bio.Align.stockholm | `align_stockholm_test.mbt` | 148 | +| Bio.Align.chain | `align_chain_test.mbt` | 156 | +| Bio.Align.maf | `align_maf_test.mbt` | 155 | +| Bio.Align.mauve | `align_mauve_test.mbt` | 67 | +| Bio.Align.clustal | `align_clustal_test.mbt` | 75 | +| Bio.Align.phylip | `align_phylip_test.mbt` | 72 | +| Bio.SearchIO.ExonerateIO.exonerate_text | `exonerate_text_test.mbt` | 42 | +| Bio.Phylo.PAML.codeml | `paml_codeml_test.mbt` | 60 | +| Bio.Phylo.PAML.baseml | `paml_baseml_test.mbt` | 64 | +| Bio.Phylo.PAML.yn00 | `paml_yn00_test.mbt` | 80 | +| Bio.Blast XML1/XML2 | `blast_xml_advanced_test.mbt` | 61 | +| Bio.Align.bed | `align_bed_test.mbt` | 152 | +| Bio.Align.bigbed | `bigbed_test.mbt` | 64 | +| Bio.Align.bigmaf | `bigmaf_test.mbt` | 79 | +| SparseArray | `sparse_array_test.mbt` | 41 | | mmCIF | `mmcif_test.mbt` | 2 | +| BinaryCIF | `binary_cif_test.mbt` | 37 | | Nexus | `nexus_test.mbt` | 2 | | EMBOSS | `emboss_test.mbt` | 15 | | ChIPseeker | `chipseeker_test.mbt` | 14 | @@ -2913,12 +3826,27 @@ moon test --update | AnnotationHub | `annotation_hub_test.mbt` | 8 | | GenomicFeatures | `genomic_features_test.mbt` | 6 | | graph | `graph_test.mbt` | 8 | -| DropletUtils | `droplet_utils_test.mbt` | 6 | +| DropletUtils兼容层 | `droplet_utils_test.mbt` | 7 | +| Bioconductor DropletUtils高级流程 | `droplet_utils_advanced_test.mbt` | 58 | | scran | `scran_test.mbt` | 8 | +| scrapper | `scrapper_test.mbt` | 35 | +| Bioconductor scuttle | `scuttle_test.mbt` | 64 | +| Bioconductor bluster | `bluster_test.mbt` | 58 | +| Bioconductor FlowSOM | `flowsom_test.mbt` | 66 | +| decontX | `decontx_test.mbt` | 43 | +| celda_CG | `celda_test.mbt` | 75 | +| miloR | `milo_test.mbt` | 37 | +| zinbwave | `zinbwave_test.mbt` | 54 | +| apeglm | `apeglm_test.mbt` | 66 | +| ALDEx2 | `aldex2_test.mbt` | 80 | +| DirichletMultinomial | `dirichlet_multinomial_test.mbt` | 58 | | monocle3 | `monocle3_test.mbt` | 10 | | ShortRead | `short_read_test.mbt` | 15 | | scater | `scater_test.mbt` | 17 | -| MAST | `mast_test.mbt` | 12 | +| MAST兼容层 | `mast_test.mbt` | 12 | +| Bioconductor MAST高级流程 | `mast_advanced_test.mbt` | 42 | +| SingleR兼容层 | `single_r_test.mbt` | 26 | +| Bioconductor SingleR高级流程 | `single_r_advanced_test.mbt` | 54 | | GenomicFiles | `genomic_files_test.mbt` | 28 | | DiffBind | `diffbind_test.mbt` | 36 | | minfi | `minfi_test.mbt` | 40 | @@ -2931,7 +3859,8 @@ moon test --update | ChromVAR | `chromvar_test.mbt` | 12 | | DelayedArray | `delayed_array_test.mbt` | 10 | | AnnotationFilter | `annotation_filter_test.mbt` | 10 | -| scDblFinder | `sc_dbl_finder_test.mbt` | 8 | +| scDblFinder兼容层 | `sc_dbl_finder_test.mbt` | 8 | +| Bioconductor scDblFinder高级流程 | `sc_dbl_finder_advanced_test.mbt` | 65 | | ChIPseeker | `chipseeker_test.mbt` | 14 | | Taxonomy | `taxonomy_test.mbt` | 7 | | GFF | `gff_test.mbt` | 5 | @@ -2944,7 +3873,10 @@ moon test --update | uwot | `uwot_test.mbt` | 9 | | microbiome | `microbiome_test.mbt` | 33 | | tradeSeq | `tradeseq_test.mbt` | 12 | -| QCP叠加 | `qcp_superimposer_test.mbt` | 7 | +| Bioconductor tradeSeq高级流程 | `tradeseq_advanced_test.mbt` | 44 | +| Bioconductor muscat高级流程 | `muscat_advanced_test.mbt` | 62 | +| QCP叠加 | `qcp_superimposer_test.mbt` | 8 | +| CEAligner | `cealign_test.mbt` | 35 | | 残基深度 | `residue_depth_test.mbt` | 10 | | 结构比对 | `structure_alignment_test.mbt` | 8 | | PDB向量 | `pdb_vectors_test.mbt` | 55 | @@ -2964,6 +3896,7 @@ moon test --update | SFF_IO | `sff_io_test.mbt` | 16 | | csaw | `csaw_test.mbt` | 9 | | slingshot | `slingshot_test.mbt` | 10 | +| slingshot高级流程 | `slingshot_advanced_test.mbt` | 57 | | SCnorm | `scnorm_test.mbt` | 8 | | EDASeq | `edaseq_test.mbt` | 10 | | SearchIO新 | `searchio_new_test.mbt` | 30 | @@ -3091,6 +4024,17 @@ moon test --update | Bio.Motifs.Transfac | `transfac_full_test.mbt` | 18 | | Bio.SearchIO.HmmerIO | `hmmer_io_test.mbt` | 19 | | Bio.SearchIO.FastaIO | `fasta_search_io_test.mbt` | 19 | +| Bio.SearchIO.InfernalIO | `infernal_io_test.mbt` | 37 | +| Bioconductor variancePartition | `variance_partition_test.mbt` | 38 | +| Bioconductor dreamlet | `dreamlet_test.mbt` | 58 | +| Bioconductor nnSVG | `nnsvg_test.mbt` | 67 | +| Bioconductor Banksy | `banksy_test.mbt` | 78 | +| Bioconductor spicyR | `spicyr_test.mbt` | 39 | +| Bioconductor lisaClust | `lisaclust_test.mbt` | 51 | +| Bioconductor SpatialDecon | `spatialdecon_test.mbt` | 60 | +| Bioconductor Voyager | `voyager_test.mbt` | 58 | +| CodonAlign Advanced | `codon_align_advanced_test.mbt` | 32 | +| Protein Analysis Advanced | `protein_analysis_advanced_test.mbt` | 26 | | Bio.PopGen.GenePop | `gene_pop_test.mbt` | 34 | | Bioconductor stageR | `stage_r_test.mbt` | 25 | | Bioconductor EnrichedHeatmap | `enriched_heatmap_test.mbt` | 20 | @@ -3156,7 +4100,7 @@ bash test/python/compare_seqio.sh moon build # 运行所有测试 -moon test --package IvanAXu/BioSeqs/test/moonbit +moon test # 更新接口文件 moon info @@ -3186,7 +4130,7 @@ moon run cmd/bench/main.mbt ### 示例程序 -项目提供 336 个示例程序,展示各模块的典型用法: +项目提供 408 个示例程序,展示各模块的典型用法: | 示例 | 说明 | 运行命令 | |------|------|----------| @@ -3194,6 +4138,13 @@ moon run cmd/bench/main.mbt | seqio_demo | 序列 I/O(FASTA/FASTQ/GenBank 解析与写入、FASTA 索引) | `moon run examples/seqio_demo/main.mbt` | | alignment_demo | 序列比对(Needleman-Wunsch、Smith-Waterman、Clustal/Phylip) | `moon run examples/alignment_demo/main.mbt` | | phylo_demo | 系统发育树(Newick 解析、遍历、距离计算、ASCII 可视化) | `moon run examples/phylo_demo/main.mbt` | +| paml_codeml_demo | CODEML control规范往返、CODONML多模型、branch/site class、BEB、AIC/BIC与LRT | `moon run examples/paml_codeml_demo` | +| paml_baseml_demo | BASEML control规范往返、参数/SE、REV Q矩阵、离散gamma及AIC/BIC/LRT | `moon run examples/paml_baseml_demo` | +| paml_yn00_demo | YN00 control规范往返、五种成对估计、未定义值、对称矩阵与均值 | `moon run examples/paml_yn00_demo` | +| blast_xml_advanced_demo | XML1/XML2解析、多描述/taxonomy、最佳命中/HSP、负translated frame、坐标路径及canonical/cross-dialect往返 | `moon run examples/blast_xml_advanced_demo` | +| align_clustal_demo | CLUSTAL generator metadata、interleaved rows、consensus、coordinate path、跨行投影、pair counts及累计计数canonical往返 | `moon run examples/align_clustal_demo` | +| align_phylip_demo | PHYLIP sequential/interleaved识别、名称规范化、coordinate path、跨行投影、pair counts及canonical往返 | `moon run examples/align_phylip_demo` | +| exonerate_text_demo | C4 metadata与层次结果、fragment/内含子区间、phase、统计及residue位置投影 | `moon run examples/exonerate_text_demo` | | pdb_demo | PDB 结构解析(原子/残基/链访问、距离计算) | `moon run examples/pdb_demo/main.mbt` | | sam_vcf_demo | SAM/VCF 解析(比对记录、变异检测、基因型查询) | `moon run examples/sam_vcf_demo/main.mbt` | | faidx_demo | FASTA 索引(pyfaidx 风格随机访问、.fai 序列化) | `moon run examples/faidx_demo/main.mbt` | @@ -3202,6 +4153,7 @@ moon run cmd/bench/main.mbt | cram_demo | CRAM 格式解析(压缩二进制序列比对格式、CRAM转BAM、参考序列管理) | `moon run examples/cram_demo/main.mbt` | | biostrings_demo | Biostrings 序列分析(IUPAC、RSCU、复杂度、Tm) | `moon run examples/biostrings_demo/main.mbt` | | genomic_ranges_demo | GenomicRanges 基因组区间操作(GRanges、区间运算、集合操作) | `moon run examples/genomic_ranges_demo/main.mbt` | +| granges_list_demo | GRangesList复合转录本、unlist/relist、精确外显子重叠与分组RangedSummarizedExperiment | `moon run examples/granges_list_demo/main.mbt` | | deseq2_demo | DESeq2 差异表达分析(size factors归一化、分散度估计、负二项GLM拟合、Wald检验、LFC收缩) | `moon run examples/deseq2_demo/main.mbt` | | dplyr_demo | dplyr 数据操作(filter、select、mutate、arrange、group_by、summarize、join) | `moon run examples/dplyr_demo/main.mbt` | | plyranges_demo | plyranges tidy基因组数据操作(GRanges的filter/mutate/select/arrange/rename/summarise) | `moon run examples/plyranges_demo/main.mbt` | @@ -3220,6 +4172,8 @@ moon run cmd/bench/main.mbt | edger_demo | edgeR 差异表达分析(DGEList创建、归一化因子、分散度估计、精确检验、GLM拟合) | `moon run examples/edger_demo/main.mbt` | | limma_demo | limma 差异表达分析(voom变换、线性模型拟合、经验贝叶斯、topTable、对比矩阵) | `moon run examples/limma_demo/main.mbt` | | summarized_experiment_demo | SummarizedExperiment 多维数据容器(Assays、行/列操作、合并) | `moon run examples/summarized_experiment_demo/main.mbt` | +| ranged_summarized_experiment_demo | RangedSummarizedExperiment GRanges行范围、链特异重叠、最近邻、promoter变换与协调排序 | `moon run examples/ranged_summarized_experiment_demo/main.mbt` | +| tree_summarized_experiment_demo | TreeSummarizedExperiment 行/列树链接、节点查询、树节点子集与层级聚合 | `moon run examples/tree_summarized_experiment_demo/main.mbt` | | iranges_demo | IRanges 整数区间操作(shift、resize、reduce、集合运算、重叠检测) | `moon run examples/iranges_demo/main.mbt` | | align_io_demo | 比对格式解析(ClustalW、FASTA、Stockholm格式解析与写入) | `moon run examples/align_io_demo/main.mbt` | | cluster_demo | 序列聚类分析(距离矩阵、层次聚类、Newick输出、轮廓系数) | `moon run examples/cluster_demo/main.mbt` | @@ -3248,6 +4202,8 @@ moon run cmd/bench/main.mbt | seq_complexity_demo | 序列复杂度与组成分析(Shannon熵、语言学复杂度、DUST评分、CGR、序列相似度) | `moon run examples/seq_complexity_demo/main.mbt` | | align_info_demo | AlignInfo 比对统计(一致性序列、保守位点、Shannon熵、成对序列同一性) | `moon run examples/align_info_demo/main.mbt` | | codon_align_demo | CodonAlign 密码子比对(密码子替换分类、dN/dS选择压力分析、密码子使用偏好、ENC) | `moon run examples/codon_align_demo/main.mbt` | +| codon_align_advanced_demo | CodonAlign 高级密码子比对(Z-test选择检验、Fisher精确检验、密码子比对构建器、滑窗dN/dS、BH-FDR、成对Ka/Ks表) | `moon run examples/codon_align_advanced_demo/main.mbt` | +| protein_analysis_advanced_demo | 高级蛋白质序列预测(Chou-Fasman二级结构、IUPred无序区、COILS卷曲螺旋、Kolaskar抗原性、Emini表面可及性、Karplus-Schulz柔柔性) | `moon run examples/protein_analysis_advanced_demo/main.mbt` | | entrez_demo | Entrez NCBI数据库访问(ESearch、EFetch、PubMed/Gene/Taxonomy解析) | `moon run examples/entrez_demo/main.mbt` | | genome_info_db_demo | GenomeInfoDb 基因组信息管理(染色体信息、着丝粒位置、染色体臂、基因组构建) | `moon run examples/genome_info_db_demo/main.mbt` | | interaction_set_demo | InteractionSet 染色质交互(Hi-C交互、锚点对、交互矩阵、距离分布、Top交互) | `moon run examples/interaction_set_demo/main.mbt` | @@ -3255,6 +4211,11 @@ moon run cmd/bench/main.mbt | tree_construction_demo | TreeConstruction 系统发育树构建(UPGMA/WPGMA/NJ算法、替换模型、距离矩阵) | `moon run examples/tree_construction_demo/main.mbt` | | neighbor_search_demo | NeighborSearch KD树近邻搜索(半径搜索、最近邻、原子对搜索) | `moon run examples/neighbor_search_demo/main.mbt` | | swissprot_demo | SwissProt 蛋白数据库解析(记录解析、特征提取、参考文献) | `moon run examples/swissprot_demo/main.mbt` | +| cellosaurus_demo | Cellosaurus细胞系记录解析、物种/同义名/交叉引用查询和序列化往返 | `moon run examples/cellosaurus_demo/main.mbt` | +| unigene_demo | NCBI UniGene cluster解析、序列/蛋白相似性/STS/转录本映射查询和序列化往返 | `moon run examples/unigene_demo/main.mbt` | +| hhr_demo | HH-suite HHR元数据与profile比对解析、命中筛选、query-target坐标映射和序列化往返 | `moon run examples/hhr_demo` | +| sparse_array_demo | SparseArray N维稀疏张量、切片/aperm、行列统计、稀疏算术和矩阵乘法 | `moon run examples/sparse_array_demo` | +| cealign_demo | CE组合扩展结构比对、AFP路径统计、CE显著性、QCP叠合与不可变全原子变换 | `moon run examples/cealign_demo` | | uniprot_io_demo | UniProt XML格式解析(蛋白质条目解析、功能注释提取、序列转换) | `moon run examples/uniprot_io_demo/main.mbt` | | chem_utils_demo | 化学计算工具(键长、键角、二面角、分子式量、氢键长度) | `moon run examples/chem_utils_demo/main.mbt` | | jaspar_demo | JASPAR PFM格式解析(模体矩阵解析、共有序列、PWM转换、序列扫描) | `moon run examples/jaspar_demo/main.mbt` | @@ -3262,12 +4223,29 @@ moon run cmd/bench/main.mbt | sva_demo | SVA 替代变量分析与ComBat批次校正(经验贝叶斯方法、PCA分析、批次效应去除) | `moon run examples/sva_demo/main.mbt` | | ballgown_demo | Ballgown 转录组水平差异表达分析(FPKM计算、t检验、转录本/基因水平DE分析) | `moon run examples/ballgown_demo/main.mbt` | | mmcif_demo | mmCIF格式解析(数据块解析、类别查询、原子位点提取、结构信息) | `moon run examples/mmcif_demo/main.mbt` | +| binary_cif_demo | BinaryCIF MessagePack解析、七类逆编码、三态mask与PDB Structure转换 | `moon run examples/binary_cif_demo` | | nexus_demo | Nexus格式解析(数据矩阵、系统发育树、距离矩阵、分类单元) | `moon run examples/nexus_demo/main.mbt` | | emboss_demo | EMBOSS工具接口(GC偏斜、AT偏斜、分子量、Tm值、ORF查找、距离计算) | `moon run examples/emboss_demo/main.mbt` | | bioconductor_demo | Bioconductor模块综合示例(ChIPseeker峰注释(外显子/内含子/UTR分类、peak2gene关联)、DOSE疾病富集、ReactomePA通路分析、AnnotationDbi注释数据库、clusterProfiler富集框架、WGCNA共表达网络、Batchelor单细胞批次校正、Seurat单细胞分析) | `moon run examples/bioconductor_demo/main.mbt` | | short_read_demo | ShortRead 短读序列质量控制(QA统计、adapter修剪、质量修剪、读长过滤、FastQC报告生成) | `moon run examples/short_read_demo/main.mbt` | | scater_demo | scater 单细胞质量控制(QC指标计算、细胞/基因过滤、CPM/log-CPM标准化、HVG检测、PCA降维) | `moon run examples/scater_demo/main.mbt` | -| mast_demo | MAST 单细胞差异表达分析(Hurdle模型、离散/连续检验、BH-FDR校正、结果汇总) | `moon run examples/mast_demo/main.mbt` | +| scrapper_demo | 批次感知RNA QC、大小因子归一化、LOWESS/HVG、多因子pseudo-bulk和不可变SCE集成 | `moon run examples/scrapper_demo` | +| scuttle_demo | batch-aware MAD异常值、subset per-feature QC、重叠feature-set聚合、精确batch coverage等化和不可变SCE写回 | `moon run examples/scuttle_demo` | +| bluster_demo | 多起点K-means、SNN图、Louvain风格聚类、two-step、ARI/silhouette/purity/RMSD、bootstrap稳定性和不可变SCE写回 | `moon run examples/bluster_demo` | +| flowsom_demo | 规则网格与MST拓扑SOM、meta-clustering、节点MFI/CV/阳性率、MAD outlier、新数据映射及FlowFrame/SCE写回 | `moon run examples/flowsom_demo` | +| slingshot_advanced_demo | Y型轨迹、start/end约束MST、同时主曲线、pseudotime/branch weights、新数据投影及不可变SCE写回 | `moon run examples/slingshot_advanced_demo` | +| tradeseq_advanced_demo | 多lineage负二项GAM、五类Wald检验、平滑预测、AIC knot评估、Slingshot直连及不可变SCE写回 | `moon run examples/tradeseq_advanced_demo` | +| muscat_advanced_demo | cluster-sample伪批量、任意设计、NB-IRLS DS、CDR归一化DD、两阶段确认及不可变SCE写回 | `moon run examples/muscat_advanced_demo` | +| droplet_utils_demo | Simple Good-Turing ambient profile、alpha估计、barcode-rank knee/inflection、Dirichlet-multinomial emptyDrops、细胞过滤及不可变SCE写回 | `moon run examples/droplet_utils_demo` | +| decontx_demo | cluster/background ambient RNA去污染、每细胞污染率、marker校正、cluster诊断和不可变SCE输出 | `moon run examples/decontx_demo` | +| celda_demo | `celda_CG`细胞群/基因模块联合聚类、top markers、新细胞预测、BIC模型选择和不可变SCE输出 | `moon run examples/celda_demo` | +| milo_demo | 精确KNN图、精炼重叠邻域、样本计数、NB-GLM差异丰度、graph spatial FDR和SCE接入 | `moon run examples/milo_demo` | +| zinbwave_demo | ZINB latent-factor拟合、dropout后验权重、归一化/插补/deviance residual和不可变SCE输出 | `moon run examples/zinbwave_demo` | +| apeglm_demo | NB-GLM MLE与自适应重尾MAP、FSR/s-value/FSOS、log2 TSV及不可变SummarizedExperiment输出 | `moon run examples/apeglm_demo` | +| aldex2_demo | Dirichlet Monte Carlo、IQLR、posterior expected eBH、effect/overlap、Aitchison距离及不可变SummarizedExperiment输出 | `moon run examples/aldex2_demo` | +| dirichlet_multinomial_demo | Dirichlet-multinomial混合聚类、Laplace选K、分组分类、分层交叉验证、ROC与SummarizedExperiment转置入口 | `moon run examples/dirichlet_multinomial_demo` | +| mast_demo | MAST 1.39.0任意设计/CDR、Bayesian logistic与Gaussian hurdle GLM、嵌套LRT、eBayes、边际logFC和不可变SCE写回 | `moon run examples/mast_demo` | +| single_r_demo | SingleR 2.15.2成对classic markers、标签内相关分位数、fine-tuning/MAD剪枝、cluster、多参考重算和不可变SCE写回 | `moon run examples/single_r_demo` | | genomic_files_demo | GenomicFiles 分布式基因组文件处理(BAM/BED/VCF扫描、区间查询、归约、覆盖度计算) | `moon run examples/genomic_files_demo/main.mbt` | | diffbind_demo | DiffBind ChIP-seq差异结合分析(峰值重叠、共识峰识别、TMM归一化、负二项分布检验) | `moon run examples/diffbind_demo/main.mbt` | | minfi_demo | minfi DNA甲基化分析(NOOB/Illumina/分位数/功能归一化、β/M值计算、DMP/DMR分析) | `moon run examples/minfi_demo/main.mbt` | @@ -3281,7 +4259,7 @@ moon run cmd/bench/main.mbt | chromvar_demo | ChromVAR 染色质变异分析(TF motif富集、GC偏差校正、细胞聚类、变异性分析、偏差图) | `moon run examples/chromvar_demo/main.mbt` | | delayed_array_demo | DelayedArray 延迟计算数组(懒加载操作、分块处理、行/列聚合、转置、子集操作) | `moon run examples/delayed_array_demo/main.mbt` | | annotation_filter_demo | AnnotationFilter 基因注释过滤(染色体筛选、生物类型过滤、链过滤、区域重叠检测、符号模式匹配) | `moon run examples/annotation_filter_demo/main.mbt` | -| sc_dbl_finder_demo | scDblFinder 单细胞双细胞检测(Doublet评分计算、最近邻搜索、双细胞检测、PCA降维、细胞过滤) | `moon run examples/sc_dbl_finder_demo/main.mbt` | +| sc_dbl_finder_demo | scDblFinder 1.27.6高级流程(capture分层、已知doublet训练、人工doublet、kNN/cxds特征、迭代分类、来源富集、自动聚类、singlet过滤与SingleCellExperiment写回) | `moon run examples/sc_dbl_finder_demo/main.mbt` | | chipseeker_demo | ChIPseeker ChIP-seq峰值注释(基因组区域分类(启动子/外显子/内含子/UTR/基因间区)、距离TSS分布、BED格式读取、peak2gene关联分析、多峰值集重叠分析、Venn图、饼图可视化、统计分析) | `moon run examples/chipseeker_demo/main.mbt` | | taxonomy_demo | Taxonomy 分类学分析(分类数据库创建、谱系查询、共同祖先计算、分类单元管理) | `moon run examples/taxonomy_demo/main.mbt` | | gff_demo | GFF GFF3格式解析(基因注释特征提取、属性解析、基因/转录本/CDS/外显子结构分析) | `moon run examples/gff_demo/main.mbt` | @@ -3359,6 +4337,34 @@ moon run cmd/bench/main.mbt | transfac_demo | TRANSFAC转录因子结合谱解析(PFM矩阵、共识序列、频率计算、序列化、参考文献) | `moon run examples/transfac_demo/main.mbt` | | hmmer_io_demo | HMMER3输出解析(domtblout域表、文本格式、Query/Hit/HSP聚合、多域比对) | `moon run examples/hmmer_io_demo/main.mbt` | | fasta_search_io_demo | FASTA搜索输出解析(-m8表格、-m9注释头、元数据、Query/Hit/HSP聚合) | `moon run examples/fasta_search_io_demo/main.mbt` | +| infernal_io_demo | Infernal cmscan/cmsearch解析(tabular 3、non-verbose文本、local-end片段、过滤与SearchIO转换) | `moon run examples/infernal_io_demo` | +| variance_partition_demo | typed固定/随机设计、ML方差分解、BLUP、precision weights、dream contrast与SummarizedExperiment接入 | `moon run examples/variance_partition_demo` | +| dreamlet_demo | SingleCellExperiment pseudobulk、TMM/logCPM、Poisson/voom权重、donor随机截距和跨cell-type FDR | `moon run examples/dreamlet_demo` | +| nnsvg_demo | nearest-neighbor GP空间变异基因检验、length scale、空间方差占比、基因过滤与SpatialExperiment接入 | `moon run examples/nnsvg_demo` | +| banksy_demo | H0/H1空间邻域特征、cell-typing/domain lambda、PCA聚类、标签平滑、参数扫描与SpatialExperiment接入 | `moon run examples/banksy_demo` | +| voyager_demo | kNN/distance-band/inverse-distance权重、全局Moran/Geary、局部Moran LISA、Getis-Ord Gi*、Lee's L、变差函数拟合、correlogram与SpatialExperiment接入 | `moon run examples/voyager_demo` | +| spicyr_demo | 有序细胞类型对cross-L、矩形窗口边界校正、precision weights、重复受试者模型、条件对比与SpatialExperiment接入 | `moon run examples/spicyr_demo` | +| lisaclust_demo | 每细胞local-K/L曲线、KDE与边界校正、确定性区域聚类、silhouette、observed/expected富集及SpatialExperiment写回 | `moon run examples/lisaclust_demo` | +| spatialdecon_demo | 背景感知log-normal解卷积、丰度/比例/计数、cell-type collapse、reverse deconvolution、负探针背景、单细胞profile与SpatialExperiment写回 | `moon run examples/spatialdecon_demo` | +| shared_reference_alignment_demo | 共享参考PWA/MSA合并、reference insertion同步、坐标映射、统计、MSA与aligned FASTA转换 | `moon run examples/shared_reference_alignment_demo` | +| alignment_map_demo | chromosome→transcript→read坐标组合、intron gap、反链、坐标查询、PSL与protein-to-codon MSA投影 | `moon run examples/alignment_map_demo` | +| alignment_counts_demo | left/internal/right gap与open/extend、affine/BLOSUM45评分、反向链和MSA逐对汇总 | `moon run examples/alignment_counts_demo` | +| align_tabular_demo | BLAST outfmt 7 BTOP、FASTA 8CC aln_code、最佳命中和TBLASTX反链translated坐标 | `moon run examples/align_tabular_demo` | +| align_psl_demo | PSL/PSLX多记录读写、反向query映射、translated target recount、block sequence和摘要 | `moon run examples/align_psl_demo` | +| align_sam_demo | SAM header/reference、正反链CIGAR路径、splicing、typed tags、PHRED、MD/NM、aligned rows与规范往返 | `moon run examples/align_sam_demo` | +| a2m_demo | A2M D/I列状态、插入槽、match共识、跨行坐标映射、pair counts、match projection与wrapped往返 | `moon run examples/a2m_demo` | +| align_emboss_demo | EMBOSS srspair元数据、局部/反向坐标、consensus、compact path、pair counts与wrapped规范往返 | `moon run examples/align_emboss_demo` | +| align_exonerate_demo | Exonerate spliced vulgar path、反向protein-to-DNA 3:1映射、operation统计及vulgar/cigar规范往返 | `moon run examples/align_exonerate_demo` | +| msf_demo | GCG/PileUp MSF metadata与checksum、interleaved rows、coordinate path、pair counts、consensus、宽度异常及规范往返 | `moon run examples/msf_demo` | +| align_nexus_demo | NEXUS interleaved DATA、quoted taxa、MATCHCHAR、coordinate path、residue mapping、统计、all-gap压缩与canonical往返 | `moon run examples/align_nexus_demo` | +| align_stockholm_demo | Stockholm GF/GS/GR/GC与reference、insertion operation、coordinate path、residue mapping、统计、all-gap压缩与canonical往返 | `moon run examples/align_stockholm_demo` | +| align_chain_demo | UCSC Chain连续记录、正反双轴绝对路径、block/counts、双向位置与区间映射、反转及canonical往返 | `moon run examples/align_chain_demo` | +| align_maf_demo | MAF track/header、a/s/i/e/q、plus/minus绝对路径、component映射、MafIndex查询、多外显子拼接与canonical往返 | `moon run examples/align_maf_demo` | +| align_mauve_demo | XMFA source metadata与LCB、正负链coordinate path、pair counts、区间索引、跨序列投影、forward重建与wrapped canonical往返 | `moon run examples/align_mauve_demo` | +| align_bed_demo | BED3/BED12、numeric/text score、正负链双轴路径、exon blocks、双向mapping、搜索、统计、BED3-BED12分级写回与canonical往返 | `moon run examples/align_bed_demo` | +| bigbed_demo | BigBed v4压缩写入/解析、AutoSQL、多级索引查询、负链exon坐标、BED导出及损坏文件诊断 | `moon run examples/bigbed_demo` | +| bigmaf_demo | BigMaf压缩bed3+1写入、bedMaf schema、R-tree区间查询、负链物种坐标映射和MAF a/s/i/e/q导出 | `moon run examples/bigmaf_demo` | +| bigpsl_demo | BigPsl bed12+13压缩写入、AutoSQL、R-tree查询、反向核酸/translated protein坐标和PSL导出 | `moon run examples/bigpsl_demo` | | gene_pop_demo | GenePop群体遗传学(基因型解析、等位基因频率、杂合度统计、序列化往返) | `moon run examples/gene_pop_demo/main.mbt` | | stage_r_demo | stageR两阶段假设检验(筛选+确认、Simes聚合、BH-FDR、Holm步降、OFDR控制) | `moon run examples/stage_r_demo/main.mbt` | | enriched_heatmap_demo | EnrichedHeatmap富集热图(信号归一化、四种均值模式、行平滑、链方向处理) | `moon run examples/enriched_heatmap_demo/main.mbt` | @@ -3403,6 +4409,7 @@ moon run cmd/bench/main.mbt - ✅ 实现 Biostrings 序列分析(IUPAC、RSCU、复杂度、Tm) - ✅ 实现 DESeq2 差异表达分析(数据集创建、结果分析、显著基因筛选) - ✅ 实现 GenomicRanges 基因组区间操作(GRanges、区间运算、集合操作) +- ✅ 实现 GenomicRanges GRangesList 复合特征(split/unlist/relist、逐组变换、集合运算、特征级重叠/最近邻/覆盖度) - ✅ 实现 dplyr 数据操作 - ✅ 实现 Smith-Waterman 局部序列比对(DNA/蛋白质比对、自定义打分、得分矩阵) - ✅ 实现 Needleman-Wunsch 全局序列比对(DNA/蛋白质比对、自定义打分、得分矩阵) @@ -3418,6 +4425,8 @@ moon run cmd/bench/main.mbt - ✅ 实现 群体遗传学分析(等位基因频率、基因型频率、哈迪-温伯格检验、FST统计、Watterson's theta) - ✅ 实现 edgeR 差异表达分析(DGEList创建、归一化因子、分散度估计、精确检验、GLM拟合) - ✅ 实现 SummarizedExperiment 多维数据容器(Assays、行/列操作、合并) +- ✅ 实现 RangedSummarizedExperiment 基因组区间实验容器(GRanges/GRangesList行范围、复合特征精确重叠/最近邻、覆盖度、区间变换与协调子集) +- ✅ 实现 TreeSummarizedExperiment 树结构实验容器(行/列树链接、节点子集、层级聚合) - ✅ 实现 IRanges 整数区间操作(shift、resize、reduce、集合运算、重叠检测) - ✅ 实现 比对格式解析(ClustalW、FASTA、Stockholm格式解析与写入) - ✅ 实现 序列聚类分析(距离矩阵、层次聚类、Newick输出、轮廓系数) @@ -3443,6 +4452,8 @@ moon run cmd/bench/main.mbt - ✅ 实现 seq_complexity 序列复杂度与组成分析(Shannon熵、语言学复杂度、DUST评分、CGR、序列相似度) - ✅ 实现 AlignInfo 比对统计(一致性序列、保守位点、Shannon熵、成对序列同一性) - ✅ 实现 CodonAlign 密码子比对(密码子替换分类、dN/dS选择压力分析、密码子使用偏好、ENC) +- ✅ 实现 CodonAlign 高级密码子比对(Z-test选择检验、Fisher精确检验、密码子比对构建器、滑窗dN/dS、BH-FDR多重校正、成对Ka/Ks表) +- ✅ 实现高级蛋白质序列预测(Chou-Fasman二级结构、IUPred无序区、COILS卷曲螺旋、Kolaskar-Tongaonkar抗原性、Emini表面可及性、Karplus-Schulz柔柔性) - ✅ 实现 Entrez NCBI数据库访问(ESearch、EFetch、PubMed/Gene/Taxonomy解析) - ✅ 实现 GenomeInfoDb 基因组信息管理(染色体信息、着丝粒位置、染色体臂、基因组构建) - ✅ 实现 InteractionSet 染色质交互(Hi-C交互、锚点对、交互矩阵、距离分布、Top交互) @@ -3450,6 +4461,61 @@ moon run cmd/bench/main.mbt - ✅ 实现 TreeConstruction 系统发育树构建(UPGMA/WPGMA/NJ算法、替换模型、距离矩阵) - ✅ 实现 NeighborSearch KD树近邻搜索(半径搜索、最近邻、原子对搜索) - ✅ 实现 SwissProt 蛋白数据库解析(记录解析、特征提取、参考文献) +- ✅ 实现 Cellosaurus 细胞系数据库解析(多记录读取、交叉引用/物种查询、平面文本序列化) +- ✅ 实现 UniGene 基因聚类记录解析(固定宽度多记录读取、类型化子记录查询、SCOUNT校验、平面文本序列化) +- ✅ 实现 Bio.Align.hhr HH-suite HHR解析(元数据、命中摘要、多块profile比对、注释保留、过滤、坐标映射与序列化) +- ✅ 实现 Bioconductor SparseArray N维稀疏数组(规范化COO、R列主序、切片/置换/绑定、稀疏算术、统计与矩阵乘法) +- ✅ 实现 Bio.PDB.cealign CE组合扩展结构比对(CA/C4'引导原子、AFP路径、CE Z-score、QCP叠合、局部优化与全原子变换) +- ✅ 实现 Bioconductor scrapper 单细胞预处理(批次感知RNA QC、大小因子与log-normalization、LOWESS/HVG、多因子pseudo-bulk、不可变SCE集成) +- ✅ 实现 Bioconductor scuttle 1.23.1基础工具(batch-aware MAD异常值、subset per-feature QC、重叠feature-set聚合、精确count/batch downsampling及不可变SCE集成) +- ✅ 实现 Bioconductor bluster 1.23.0通用聚类(K-means++、KNN/SNN图、Louvain风格优化、two-step、聚类诊断、bootstrap稳定性与SingleCellExperiment接入) +- ✅ 实现 Bioconductor FlowSOM 2.21.0拓扑自组织映射(KWSP/PCA、规则网格、多阶段MST、meta-clustering、节点统计、MAD outlier及FlowFrame/SingleCellExperiment接入) +- ✅ 实现 Bioconductor slingshot 2.21.0高级轨迹推断(soft membership、约束MST/omega forest、同时主曲线、重加权/重分配、预测及SingleCellExperiment接入) +- ✅ 实现 Bioconductor tradeSeq 1.27.0高级轨迹差异表达(确定性lineage权重、多lineage惩罚NB-GAM、dispersion/AIC/协方差、五类Wald检验、平滑预测、knot选择、Slingshot及SingleCellExperiment接入) +- ✅ 实现 Bioconductor muscat 1.27.4高级多样本多亚群分析(五类cluster-sample伪批量、任意设计/多contrast、NB-IRLS DS、CDR归一化DD、local/global FDR、两阶段确认及SingleCellExperiment接入) +- ✅ 实现 Bio.PDB.binary_cif BinaryCIF解析(MessagePack、七类逆编码、三态mask、类别查询与PDB Structure转换) +- ✅ 实现 Bioconductor miloR 单细胞邻域差异丰度(精确KNN图、median精炼采样、邻域计数、固定效应NB-GLM、graph spatial FDR与SCE接入) +- ✅ 实现 Bio.SearchIO.InfernalIO Infernal cmscan/cmsearch输出解析(tabular 1/2/3、non-verbose文本、--noali、CM/HMM-only、local-end多片段与SearchIO转换) +- ✅ 实现 Bioconductor variancePartition 重复测量混合模型(ML/REML方差分解、BLUP、precision weights、dream contrast、数值Satterthwaite、BH-FDR与SummarizedExperiment接入) +- ✅ 实现 Bioconductor dreamlet cohort-scale单细胞重复测量分析(sample×cell-type pseudobulk、TMM、过滤、logCPM、Poisson/voom权重、typed混合模型与study-wide FDR) +- ✅ 实现 Bio.Align共享参考序列比对合并(混合PWA/MSA、reference-boundary insertion同步、局部坐标、双向映射、统计与格式转换) +- ✅ 实现 Bio.Align.Alignment map/mapall(alignment path组合、local clipping、gap与正反链传播、坐标查询、PSL及protein-to-codon MSA投影) +- ✅ 实现 Bio.Align.Alignment counts(left/internal/right insertion/deletion、gap open/extend、identity/mismatch/positive、wildcard、替换矩阵、affine评分及MSA逐对汇总) +- ✅ 实现 Bio.Align.tabular alignment-aware搜索结果解析(BLAST outfmt 7、FASTA 8CB/8CC、BTOP/aln_code、正反链与translated坐标、零命中query) +- ✅ 实现 Bio.Align.psl alignment-aware PSL/PSLX(21/23列严格读写、核酸/translated路径、双轴链向、block/gap统计、match recount、坐标映射与往返) +- ✅ 实现 Bio.Align.sam alignment-aware SAM(header/reference与typed tags、CIGAR坐标路径、正反链与clipping、PHRED、MD/NM、严格校验及规范往返) +- ✅ 实现 Bio.Align.a2m 状态感知多序列比对(D/I列状态、大小写与点gap编码、wrapped/CRLF严格读写、坐标映射、插入槽、统计、共识与match-only投影) +- ✅ 实现 Bio.Align.emboss alignment输出(srspair/pair/simple、多alignment/多序列、局部/反向坐标、纯gap block、consensus统计、compact path与规范往返) +- ✅ 实现 Bio.Align.exonerate alignment输出(cigar/vulgar、完整operation path、正反链/protein strand、3:1 translated坐标、双向映射、统计与规范往返) +- ✅ 实现 Bio.Align.msf GCG/PileUp多序列比对(AA/NA header、interleaved rows、标准checksum、gap规范化、坐标映射、统计与canonical writer) +- ✅ 实现 Bio.Align.nexus NEXUS多序列比对(DATA/CHARACTERS/TAXA、nested comments、quoted/duplicate taxa、sequential/interleaved MATRIX、datatype/MATCHCHAR、坐标统计与canonical writer) +- ✅ 实现 Bio.Align.stockholm Stockholm多序列比对(严格多记录、GF/GS/GR/GC、reference/database/nested-domain、M/D/I列、all-gap压缩、坐标统计与canonical writer) +- ✅ 实现 Bio.Align.chain UCSC Chain成对比对(12/13字段严格多记录、float score、正反双轴absolute path、size/dt/dq、坐标映射、反转、查询与canonical writer) +- ✅ 实现 Bio.Align.maf MAF多基因组比对(track/header、严格a/s/i/e/q、正负链absolute path、component映射、MafIndex半开查询、多外显子拼接与canonical writer) +- ✅ 实现 Bio.Align.mauve 现代XMFA多基因组比对(严格metadata/LCB、combined/separate source、正负链coordinate path、区间索引、跨序列投影、统计、重建与canonical writer) +- ✅ 实现 Bio.Align.clustal 现代CLUSTAL多序列比对(六类generator header、严格interleaved blocks、累计残基数、consensus、坐标投影、统计与canonical writer) +- ✅ 实现 Bio.Align.phylip 现代PHYLIP多序列比对(固定10列名称、sequential/interleaved自动识别、wrapped/grouped blocks、坐标投影、统计与canonical writer) +- ✅ 实现 Bio.SearchIO.ExonerateIO.exonerate_text C4人类可读报告(层次聚合、3/4/5行模型、wrapped blocks、intron/NER/split codon/frameshift、翻译与链感知坐标) +- ✅ 实现 Bio.Phylo.PAML.codeml 离线工作流(严格control读写、CODONML/AAML、NSsites/branch-site/clade/free-ratio、pairwise/距离矩阵、多基因、BEB/NEB与AIC/BIC/LRT) +- ✅ 实现 Bio.Phylo.PAML.baseml 离线工作流(严格control读写、JC69-UNRESTu、参数/SE、kappa/Q矩阵、nparK/auto-dGamma、nhomo节点与AIC/BIC/LRT) +- ✅ 实现 Bio.Phylo.PAML.yn00 离线工作流(严格control读写、NG86/YN00/LWL85/LWL85m/LPB93、PAML 4.1-4.9i、非有限值、对称矩阵与均值) +- ✅ 实现 Bio.Blast现代XML1/XML2离线工作流(严格XML树、多query/report、HitDescr/taxonomy、参数/统计、八类程序链向/translated坐标及canonical writer) +- ✅ 实现 Bio.Align.bed BED成对比对(BED3-BED12严格读写、numeric/text score、正负链双轴路径、block投影、双向residue映射、区间查询、统计与分级writer) +- ✅ 实现 Bio.Align.bigbed BigBed v4二进制区间格式(BED3-BED12、AutoSQL、多级B+ tree/R-tree、DEFLATE、区间/名称查询与BED导出) +- ✅ 实现 Bio.Align.bigmaf BigMaf多物种比对索引(标准bedMaf、MAF a/s/i/e/q、正负链坐标映射、压缩BigBed查询与MAF导出) +- ✅ 实现 Bio.Align.bigpsl BigPsl成对比对索引(标准bed12+13、核酸/translated protein坐标、match recount、压缩BigBed查询与PSL导出) +- ✅ 实现 Bioconductor decontX ambient RNA去污染(cluster-native/contaminant混合、Beta/Dirichlet先验EM、empty-droplet background、自动聚类、计数分解与SCE集成) +- ✅ 实现 Bioconductor celda `celda_CG`细胞群与基因模块联合聚类(分层Dirichlet-multinomial、collapsed EM/Gibbs、多链、K/L选择、预测与SCE集成) +- ✅ 实现 Bioconductor zinbwave 零膨胀负二项低维模型(cell/gene协变量、offset、latent factors、dispersion shrinkage、observational weights、残差/插补与SCE集成) +- ✅ 实现 Bioconductor apeglm 自适应重尾效应量收缩(NB-GLM MLE、经验贝叶斯Cauchy/Student-t先验、多起点MAP、Laplace后验、FSR/FSOS/s-value与容器接入) +- ✅ 实现 Bioconductor ALDEx2 组成型差异丰度(Dirichlet Monte Carlo、六类denominator、两组/配对检验、posterior expected BH、effect/overlap、距离与SummarizedExperiment接入) +- ✅ 实现 Bioconductor DirichletMultinomial 混合聚类与分类(DMM概率、soft k-means、log-alpha BFGS/EM、Gamma prior、Laplace/AIC/BIC、dmngroup分类、分层CV、ROC与SummarizedExperiment接入) +- ✅ 实现 Bioconductor nnSVG 空间变异基因检测(坐标缩放与前驱kNN、指数协方差NNGP、covariate GLS、gene-specific length scale、空间/非空间LR检验、BH-FDR与SpatialExperiment接入) +- ✅ 实现 Bioconductor Banksy 空间邻域增强聚类(六类空间核、H0/H1+、lambda联合矩阵、分组标准化、PCA、多起点k-means、标签平滑、参数扫描与SpatialExperiment接入) +- ✅ 实现 Bioconductor spicyR 差异空间细胞共定位分析(有序细胞类型对cross-L、矩形边界校正、图像级统计、precision weights、加权/随机截距模型、多条件对比、BH-FDR与SpatialExperiment接入) +- ✅ 实现 Bioconductor lisaClust 局部空间区域发现(多细胞类型local-K/L、Gaussian KDE强度校正、矩形/凸包窗口、圆盘边界积分、确定性多起点k-means、silhouette、区域富集与SpatialExperiment写回) +- ✅ 实现 Bioconductor SpatialDecon 背景感知空间解卷积(加权log-normal非负回归、异常点重拟合、Hessian不确定度、丰度/计数尺度、cell-type collapse、reverse deconvolution、负探针背景、单细胞profile与SpatialExperiment写回) +- ✅ 实现 Bioconductor DropletUtils 1.33.0高级emptyDrops(Simple Good-Turing ambient profile、barcode-rank knee/inflection、multinomial/Dirichlet-multinomial、alpha MLE、确定性Monte Carlo、Phipson–Smyth p值、BH-FDR、高计数保留与SingleCellExperiment写回) - ✅ 实现 FGSEA 快速基因集富集分析(基因排名、富集分数、NES、p值、Leading Edge基因、BH校正) - ✅ 实现 SVA 替代变量分析与ComBat批次校正(经验贝叶斯方法、PCA分析、批次效应去除) - ✅ 实现 Ballgown 转录组水平差异表达分析(FPKM计算、t检验、转录本/基因水平DE分析) @@ -3462,7 +4528,8 @@ moon run cmd/bench/main.mbt - ✅ 实现 AnnotationDbi注释数据库、clusterProfiler富集框架、WGCNA共表达网络 - ✅ 实现 ShortRead 短读序列质量控制(QA统计、adapter修剪、质量修剪、读长过滤、FastQC报告生成) - ✅ 实现 scater 单细胞质量控制(QC指标计算、细胞/基因过滤、CPM/log-CPM标准化、HVG检测、PCA降维) -- ✅ 实现 MAST 单细胞差异表达分析(Hurdle模型、离散/连续检验、BH-FDR校正、结果汇总) +- ✅ 实现 Bioconductor MAST 1.39.0高级单细胞差异表达(任意设计/CDR、Bayesian logistic与Gaussian hurdle GLM、嵌套LRT、H0/H1 eBayes、NA/FDR、边际logFC及SingleCellExperiment写回) +- ✅ 实现 Bioconductor SingleR 2.15.2高级参考注释(严格基因对齐、成对classic markers、标签内相关分位数、迭代fine-tuning、delta/MAD剪枝、cluster、多参考重算及SingleCellExperiment写回) - ✅ 实现 GenomicFiles 分布式基因组文件处理(BAM/BED/VCF扫描、区间查询、归约、覆盖度计算) - ✅ 实现 DiffBind ChIP-seq差异结合分析(峰值重叠、共识峰识别、TMM归一化、负二项分布检验) - ✅ 实现 minfi DNA甲基化分析(NOOB/Illumina/分位数/功能归一化、β/M值计算、DMP/DMR分析) @@ -3474,7 +4541,7 @@ moon run cmd/bench/main.mbt - ✅ 实现 ChromVAR 染色质变异分析(TF motif富集、GC偏差校正、细胞聚类、变异性分析、偏差图) - ✅ 实现 DelayedArray 延迟计算数组(懒加载操作、分块处理、行/列聚合、转置、子集操作) - ✅ 实现 AnnotationFilter 基因注释过滤(染色体筛选、生物类型过滤、链过滤、区域重叠检测、符号模式匹配) -- ✅ 实现 scDblFinder 单细胞双细胞检测(Doublet评分计算、最近邻搜索、双细胞检测、PCA降维、细胞过滤) +- ✅ 实现 Bioconductor scDblFinder 1.27.6高级双细胞检测(人工doublet、共同归一化/PCA、精确kNN与cxds特征、迭代正则化分类、capture分层阈值、homotypic修正、来源富集及SingleCellExperiment写回) - ✅ 实现 ChIPseeker ChIP-seq峰值注释(基因组区域分类、距离TSS分布、注释可视化、统计分析) - ✅ 实现 DESeq2 差异表达分析(size factors归一化、分散度估计、负二项GLM拟合、Wald检验、LFC收缩) - ✅ 实现 ChIPseeker峰注释、DOSE疾病富集、ReactomePA通路分析 diff --git a/examples/a2m_demo/main.mbt b/examples/a2m_demo/main.mbt new file mode 100644 index 00000000..e9e11e35 --- /dev/null +++ b/examples/a2m_demo/main.mbt @@ -0,0 +1,77 @@ +///| +fn main { + println("=== Biopython Bio.Align.a2m Demo ===") + + let alignment = @src.a2m_parse(@src.a2m_example_text()) catch { + _ => abort("failed to parse A2M sample") + } + + println("\n1. State-aware A2M parsing") + println(" " + alignment.summary()) + println(" Column states: " + alignment.states) + println(" Reference sequence: " + alignment.sequences[0].sequence) + + println("\n2. Insertion slots and match consensus") + for insertion in alignment.insertion_runs() { + println( + " Slot " + + insertion.slot.to_string() + + ": columns [" + + insertion.start_column.to_string() + + ", " + + insertion.end_column.to_string() + + "), width " + + insertion.width.to_string(), + ) + } + let consensus = alignment.consensus(include_insertions=false) catch { + _ => abort("failed to calculate A2M consensus") + } + println(" Match-state consensus: " + consensus) + + println("\n3. Coordinate mapping") + match alignment.map_position(0, 1, 5) { + Some(position) => + println( + " Reference residue 5 maps to query_one residue " + + position.to_string() + + " (zero-based)", + ) + None => println(" Reference residue 5 maps to a query_one gap") + } + + println("\n4. Pairwise alignment statistics") + let counts = alignment.pair_counts(0, 1) catch { + _ => abort("failed to calculate A2M pair counts") + } + println( + " Aligned/identity/mismatch: " + + counts.aligned.to_string() + + "/" + + counts.identities.to_string() + + "/" + + counts.mismatches.to_string(), + ) + println(" Identity fraction: " + counts.identity().to_string()) + + println("\n5. Match-only projection") + let projected = alignment.match_projection() catch { + _ => abort("failed to project A2M match columns") + } + println(" " + projected.summary()) + println(" Projected states: " + projected.states) + + println("\n6. Wrapped canonical serialization") + let serialized = @src.a2m_write(alignment, line_width=6) catch { + _ => abort("failed to write A2M alignment") + } + let reparsed = @src.a2m_parse(serialized) catch { + _ => abort("failed to reparse A2M alignment") + } + println(" Serialized bytes: " + serialized.length().to_string()) + println( + " Round trip preserved alignment: " + (alignment == reparsed).to_string(), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/a2m_demo/moon.pkg b/examples/a2m_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/a2m_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/aldex2_demo/main.mbt b/examples/aldex2_demo/main.mbt new file mode 100644 index 00000000..bf815237 --- /dev/null +++ b/examples/aldex2_demo/main.mbt @@ -0,0 +1,145 @@ +///| +fn print_feature(result : @src.Aldex2Result, name : String) -> Unit { + let effect = match result.effects.find_feature(name) { + Some(value) => value + None => abort("ALDEx2 example feature not found: " + name) + } + let tests = match result.tests.find_feature(name) { + Some(value) => value + None => abort("ALDEx2 example test result not found: " + name) + } + println( + " " + + name + + ": effect=" + + effect.effect.to_string() + + ", overlap=" + + effect.overlap.to_string() + + ", Welch eBH=" + + tests.welch_adjusted.to_string() + + ", Wilcoxon eBH=" + + tests.wilcoxon_adjusted.to_string(), + ) +} + +///| +fn main { + println("=== Bioconductor ALDEx2 Demo ===") + let (counts, conditions, features, samples) = @src.aldex2_example_data() + let config = @src.Aldex2Config::create( + mc_samples=64, + denominator=@src.Aldex2Iqlr, + seed=2026, + ) catch { + _ => abort("failed to create ALDEx2 configuration") + } + + println("\n1. Draw deterministic Dirichlet instances and apply IQLR") + let clr = @src.aldex2_clr( + counts, + conditions, + config~, + feature_names=features, + sample_names=samples, + ) catch { + Aldex2Error(message) => abort("failed to construct ALDEx2 CLR: " + message) + } + println( + " retained features=" + + clr.feature_count().to_string() + + ", removed all-zero features=" + + clr.removed_feature_count().to_string() + + ", samples=" + + clr.sample_count().to_string() + + ", Monte Carlo instances=" + + clr.mc_sample_count().to_string(), + ) + + println("\n2. Estimate posterior tests and standardized effects") + let result = @src.aldex2( + counts, + conditions, + config~, + feature_names=features, + sample_names=samples, + ) catch { + Aldex2Error(message) => abort("failed to run ALDEx2: " + message) + } + print_feature(result, "increased") + print_feature(result, "decreased") + print_feature(result, "stable_high") + print_feature(result, "rare_increased") + + println("\n3. Rank, select, and summarize compositional effects") + let ranked = result.ranked() + let summary = result.summary(maximum_adjusted=0.1, minimum_effect=1.0) + println( + " largest absolute effect=" + + ranked[0].feature_name + + " (" + + ranked[0].effect.to_string() + + ")", + ) + println( + " Welch eBH <= 0.10=" + + summary.significant_welch_count.to_string() + + ", |effect| >= 1=" + + summary.large_effect_count.to_string() + + ", jointly selected=" + + result.select(maximum_adjusted=0.1, minimum_effect=1.0).length().to_string(), + ) + + println("\n4. Compute posterior expected Aitchison distances") + let distances = result.clr.expected_distance() + println( + " distance(C1, T1)=" + + distances[0][4].to_string() + + ", symmetric=" + + (distances[0][4] == distances[4][0]).to_string(), + ) + + println("\n5. Add posterior summaries to an immutable container copy") + let assays : Map[String, Array[Array[Double]]] = Map([("counts", counts)]) + let metadata : Map[String, String] = Map([("source", "aldex2_demo")]) + let experiment = @src.summarized_experiment(assays, [], [], metadata) + let output = @src.aldex2_summarized_experiment( + experiment, + conditions, + config~, + feature_names=features, + sample_names=samples, + ) catch { + Aldex2Error(message) => + abort("failed to integrate ALDEx2 with SummarizedExperiment: " + message) + } + let effect_assay = match @src.se_assay(output.experiment, "aldex2_effect") { + Some(value) => value + None => abort("ALDEx2 effect assay missing") + } + println( + " effect rows=" + + effect_assay.length().to_string() + + ", source object unchanged=" + + (@src.se_assay(experiment, "aldex2_effect") is None).to_string(), + ) + + println("\n6. Export an ALDEx2-compatible tabular summary") + println(result.to_tsv()) + + println("7. Invalid count matrices produce explicit diagnostics") + let rejected = try { + ignore( + @src.aldex2([[2.0, -1.0, 3.0, 4.0], [5.0, 6.0, 7.0, 8.0]], [ + "A", "A", "B", "B", + ]), + ) + false + } catch { + Aldex2Error(message) => { + println(" " + message) + true + } + } + println(" malformed counts rejected=" + rejected.to_string()) + println("\n=== Demo Complete ===") +} diff --git a/examples/aldex2_demo/moon.pkg b/examples/aldex2_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/aldex2_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/align_bed_demo/main.mbt b/examples/align_bed_demo/main.mbt new file mode 100644 index 00000000..47424924 --- /dev/null +++ b/examples/align_bed_demo/main.mbt @@ -0,0 +1,157 @@ +///| +fn bed_coordinate_path_text( + coordinates : Array[@src.AlignBedCoordinate], +) -> String { + let output = StringBuilder::new() + for index = 0; index < coordinates.length(); index = index + 1 { + if index > 0 { + output.write_string(" -> ") + } + output.write_string( + "(" + + coordinates[index].target.to_string() + + ", " + + coordinates[index].query.to_string() + + ")", + ) + } + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.bed Demo ===") + + let document = @src.align_bed_parse(@src.align_bed_example_text()) catch { + AlignBedError(message) => abort("BED parsing failed: " + message) + } + let plus = document.alignments[0] + let minus = document.alignments[1] + + println("\n1. Parse mixed BED3 and BED12 records") + println( + " records=" + + document.alignments.length().to_string() + + ", targets=" + + document.targets().length().to_string() + + ", queries=" + + document.queries().length().to_string(), + ) + println( + " numeric score=" + + plus.score.unwrap().number().unwrap_or(-1.0).to_string(), + ) + let text_score_document = @src.align_bed_parse( + "chr1\t10\t20\tcurated-hit\thigh-confidence\t+", + ) catch { + AlignBedError(message) => abort("text BED score parsing failed: " + message) + } + println( + " text score=" + + text_score_document.alignments[0].score.unwrap().text_value().unwrap_or( + "missing", + ), + ) + + println("\n2. Inspect plus/minus target-query coordinate paths") + println(" plus: " + bed_coordinate_path_text(plus.coordinates)) + println(" minus: " + bed_coordinate_path_text(minus.coordinates)) + println( + " plus strand=" + + plus.strand() + + ", minus strand=" + + minus.strand(), + ) + + println("\n3. Reconstruct exon blocks and transcript sizes") + for index = 0; index < plus.blocks().length(); index = index + 1 { + let block = plus.blocks()[index] + println( + " exon " + + (index + 1).to_string() + + ": target[" + + block.target_start.to_string() + + ", " + + block.target_end.to_string() + + ") query[" + + block.query_start.to_string() + + ", " + + block.query_end.to_string() + + ")", + ) + } + println( + " plus query size=" + + plus.query_size().to_string() + + ", minus query size=" + + minus.query_size().to_string(), + ) + + println("\n4. Map aligned residues in both directions") + println( + " plus target 1000 -> query " + + plus.map_target_position(1000).unwrap_or(-1).to_string(), + ) + println( + " plus query 1054 -> target " + + plus.map_query_position(1054).unwrap_or(-1).to_string(), + ) + println( + " minus target 2000 -> query " + + minus.map_target_position(2000).unwrap_or(-1).to_string(), + ) + println( + " intron target 3000 -> query " + + plus.map_target_position(3000).unwrap_or(-1).to_string(), + ) + + println("\n5. Count path operations and query half-open intervals") + let counts = plus.counts() + let summary = document.summary() + let hits = document.search("chr22", 4900, 5100) catch { + AlignBedError(message) => abort("BED interval search failed: " + message) + } + println( + " aligned=" + + counts.aligned.to_string() + + ", intron bases=" + + counts.target_skip_bases.to_string() + + ", blocks=" + + counts.blocks.to_string(), + ) + println( + " document aligned=" + + summary.aligned_bases.to_string() + + ", plus/minus=" + + summary.plus_strand_count.to_string() + + "/" + + summary.minus_strand_count.to_string(), + ) + println(" chr22 overlaps [4900, 5100): " + hits.length().to_string()) + + println("\n6. Project one alignment to BED3 through BED12") + for columns = 3; columns <= 12; columns = columns + 1 { + let line = @src.align_bed_format(plus, bed_columns=columns) catch { + AlignBedError(message) => abort("BED formatting failed: " + message) + } + println( + " BED" + + columns.to_string() + + ": " + + line[0:line.length() - 1].to_owned(), + ) + } + + println("\n7. Write canonical BED12 and parse it again") + let single = @src.AlignBedDocument::create([plus]) catch { + AlignBedError(message) => abort("BED document creation failed: " + message) + } + let canonical = @src.align_bed_write(single) catch { + AlignBedError(message) => abort("BED serialization failed: " + message) + } + let round_trip = @src.align_bed_parse(canonical) catch { + AlignBedError(message) => abort("BED round-trip parsing failed: " + message) + } + println(" " + canonical[0:canonical.length() - 1].to_owned()) + println(" round-trip preserved: " + (round_trip == single).to_string()) +} diff --git a/examples/align_bed_demo/moon.pkg b/examples/align_bed_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_bed_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_chain_demo/main.mbt b/examples/align_chain_demo/main.mbt new file mode 100644 index 00000000..91c30ec5 --- /dev/null +++ b/examples/align_chain_demo/main.mbt @@ -0,0 +1,114 @@ +///| +fn chain_coordinate_path_text( + coordinates : Array[@src.AlignChainCoordinate], +) -> String { + let output = StringBuilder::new() + for index = 0; index < coordinates.length(); index = index + 1 { + if index > 0 { + output.write_string(" -> ") + } + output.write_string( + "(" + + coordinates[index].target.to_string() + + ", " + + coordinates[index].query.to_string() + + ")", + ) + } + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.chain Demo ===") + + let alignments = @src.align_chain_parse_all(@src.align_chain_example_text()) catch { + AlignChainError(message) => abort("Chain parsing failed: " + message) + } + let alignment = alignments[0] + + println("\n1. Parse consecutive UCSC Chain records") + println(" records: " + alignments.length().to_string()) + for record in alignments { + println(" " + record.summary()) + } + + println("\n2. Inspect absolute coordinates and canonical blocks") + println(" operations: " + alignment.operation_path()) + println(" coordinates: " + chain_coordinate_path_text(alignment.coordinates)) + let blocks = alignment.blocks() catch { + AlignChainError(message) => + abort("Chain block reconstruction failed: " + message) + } + for index = 0; index < blocks.length(); index = index + 1 { + let block = blocks[index] + println( + " block " + + (index + 1).to_string() + + ": size=" + + block.size.to_string() + + ", dt=" + + block.target_gap.to_string() + + ", dq=" + + block.query_gap.to_string(), + ) + } + + println("\n3. Calculate coordinate-only counts") + let counts = alignment.counts() + println( + " aligned=" + + counts.aligned.to_string() + + ", target-gap bases=" + + counts.target_gap_bases.to_string() + + ", query-gap bases=" + + counts.query_gap_bases.to_string(), + ) + println( + " alignment length=" + + alignment.alignment_length().to_string() + + ", aligned blocks=" + + counts.aligned_blocks.to_string(), + ) + + println("\n4. Map residues and intervals in both directions") + println( + " target 42530895 -> query " + + alignment.map_target_position(42530895).unwrap_or(-1).to_string(), + ) + println( + " query 180 -> target " + + alignment.map_query_position(180).unwrap_or(-1).to_string(), + ) + let pieces = alignment.map_target_range(42530950, 42532030) catch { + AlignChainError(message) => abort("Chain range mapping failed: " + message) + } + for piece in pieces { + println( + " target[" + + piece.target_start.to_string() + + ", " + + piece.target_end.to_string() + + ") -> query[" + + piece.query_start.to_string() + + ", " + + piece.query_end.to_string() + + ")", + ) + } + + println("\n5. Invert and round-trip canonical Chain output") + let inverted = alignment.invert() catch { + AlignChainError(message) => abort("Chain inversion failed: " + message) + } + println(" inverted: " + inverted.summary()) + let canonical = @src.align_chain_write_all(alignments) catch { + AlignChainError(message) => abort("Chain serialization failed: " + message) + } + let round_trip = @src.align_chain_parse_all(canonical) catch { + AlignChainError(message) => + abort("Chain round-trip parsing failed: " + message) + } + println(" output bytes: " + canonical.length().to_string()) + println(" round-trip preserved: " + (round_trip == alignments).to_string()) +} diff --git a/examples/align_chain_demo/moon.pkg b/examples/align_chain_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_chain_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_clustal_demo/main.mbt b/examples/align_clustal_demo/main.mbt new file mode 100644 index 00000000..24c9acc8 --- /dev/null +++ b/examples/align_clustal_demo/main.mbt @@ -0,0 +1,104 @@ +///| +fn clustal_demo_parse(text : String) -> @bio.AlignClustalAlignment { + @bio.align_clustal_parse(text) catch { + AlignClustalError(message) => abort("CLUSTAL parse failed: " + message) + } +} + +///| +fn clustal_demo_write(alignment : @bio.AlignClustalAlignment) -> String { + @bio.align_clustal_write( + alignment, + block_width=5, + name_width=16, + include_counts=true, + include_consensus=true, + ) catch { + AlignClustalError(message) => abort("CLUSTAL write failed: " + message) + } +} + +///| +fn clustal_demo_path(path : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < path.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(path[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.clustal Offline Demo ===") + let alignment = clustal_demo_parse(@bio.align_clustal_example_text()) + + println("\n1. Generator metadata and interleaved rows") + println(" " + alignment.summary()) + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": aligned=" + + sequence.aligned_sequence + + ", ungapped=" + + sequence.sequence, + ) + } + println( + " stored consensus=" + alignment.consensus.unwrap_or(""), + ) + println(" calculated consensus=" + alignment.calculated_consensus()) + + println("\n2. Compressed coordinate path") + let paths = alignment.coordinate_path() + for row = 0; row < paths.length(); row = row + 1 { + println( + " " + alignment.sequences[row].id + ": " + clustal_demo_path(paths[row]), + ) + } + println( + " reference residue 3 -> query_one residue " + + alignment.map_position(0, 1, 3).unwrap_or(-1).to_string(), + ) + + println("\n3. Pairwise statistics and majority residues") + let counts = alignment.pair_counts(0, 1) catch { + AlignClustalError(message) => abort("CLUSTAL counting failed: " + message) + } + let majority = alignment.majority_consensus(minimum_fraction=0.5) catch { + AlignClustalError(message) => abort("CLUSTAL consensus failed: " + message) + } + println( + " aligned=" + + counts.aligned.to_string() + + ", identities=" + + counts.identities.to_string() + + ", gap opens=" + + counts.gap_opens.to_string() + + ", identity=" + + counts.identity().to_string(), + ) + println(" majority=" + majority) + + println("\n4. Canonical writing with cumulative residue counts") + let canonical = clustal_demo_write(alignment) + let reparsed = clustal_demo_parse(canonical) + println( + " canonical generator=" + + reparsed.metadata.program + + " " + + reparsed.metadata.version, + ) + println( + " stable round-trip=" + + (clustal_demo_write(reparsed) == canonical).to_string(), + ) + println( + "\nThe demo parses and writes existing artifacts; it does not launch an aligner.", + ) +} diff --git a/examples/align_clustal_demo/moon.pkg b/examples/align_clustal_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/align_clustal_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/align_emboss_demo/main.mbt b/examples/align_emboss_demo/main.mbt new file mode 100644 index 00000000..b69ad51f --- /dev/null +++ b/examples/align_emboss_demo/main.mbt @@ -0,0 +1,130 @@ +///| +fn coordinate_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index in 0.. 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.emboss Demo ===") + + let document = @src.align_emboss_parse(@src.align_emboss_example_text()) catch { + AlignEmbossError(message) => abort("EMBOSS parsing failed: " + message) + } + let alignment = document.alignments[0] + + println("\n1. Parse report and alignment metadata") + println(" " + document.summary()) + println(" " + alignment.summary()) + println( + " matrix=" + + alignment.annotations.matrix + + ", score=" + + alignment.annotations.score.unwrap_or(0.0).to_string(), + ) + + println("\n2. Preserve local coordinates and consensus") + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": boundaries [" + + sequence.start.to_string() + + ", " + + sequence.end.to_string() + + "), residues=" + + sequence.sequence.length().to_string(), + ) + } + println(" consensus: " + alignment.consensus) + + println("\n3. Map absolute residue coordinates") + match alignment.map_position(0, 1, 87) { + Some(position) => + println( + " reference residue 87 maps to query residue " + + position.to_string() + + " (zero-based absolute coordinates)", + ) + None => println(" reference residue 87 maps to a query gap") + } + let path = alignment.coordinate_path() + println(" reference path: " + coordinate_text(path[0])) + println(" query path: " + coordinate_text(path[1])) + + println("\n4. Calculate pairwise EMBOSS statistics") + let counts = alignment.pair_counts(0, 1) catch { + AlignEmbossError(message) => abort("pair counting failed: " + message) + } + println( + " identities/mismatches/gaps: " + + counts.identities.to_string() + + "/" + + counts.mismatches.to_string() + + "/" + + (counts.insertions + counts.deletions).to_string(), + ) + println( + " positives=" + + counts.positives.to_string() + + ", identity=" + + counts.identity().to_string(), + ) + + println("\n5. Handle reverse-strand boundaries") + let annotations = @src.AlignEmbossAnnotations::create(4) catch { + AlignEmbossError(message) => + abort("annotation construction failed: " + message) + } + let forward = @src.AlignEmbossSequence::create("target", "ACGT", 2, 6) catch { + AlignEmbossError(message) => + abort("forward row construction failed: " + message) + } + let reverse = @src.AlignEmbossSequence::create("reverse", "ACGT", 20, 16) catch { + AlignEmbossError(message) => + abort("reverse row construction failed: " + message) + } + let reverse_alignment = @src.AlignEmbossAlignment::create( + [forward, reverse], + annotations, + consensus="||||", + ) catch { + AlignEmbossError(message) => + abort("reverse alignment construction failed: " + message) + } + println( + " reverse columns 0..3 map to residues " + + reverse_alignment + .column_to_sequence_position(1, 0) + .unwrap_or(-1) + .to_string() + + ".." + + reverse_alignment + .column_to_sequence_position(1, 3) + .unwrap_or(-1) + .to_string(), + ) + + println("\n6. Wrapped canonical serialization") + let serialized = @src.align_emboss_write(document, line_width=7) catch { + AlignEmbossError(message) => abort("EMBOSS writing failed: " + message) + } + let reparsed = @src.align_emboss_parse(serialized) catch { + AlignEmbossError(message) => + abort("serialized EMBOSS parsing failed: " + message) + } + println(" serialized bytes: " + serialized.length().to_string()) + println( + " round trip preserved report: " + (document == reparsed).to_string(), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/align_emboss_demo/moon.pkg b/examples/align_emboss_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_emboss_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_exonerate_demo/main.mbt b/examples/align_exonerate_demo/main.mbt new file mode 100644 index 00000000..70e666b5 --- /dev/null +++ b/examples/align_exonerate_demo/main.mbt @@ -0,0 +1,101 @@ +///| +fn coordinate_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index in 0.. 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.exonerate Demo ===") + + let document = @src.align_exonerate_parse(@src.align_exonerate_example_text()) catch { + AlignExonerateError(message) => + abort("Exonerate parsing failed: " + message) + } + let spliced = document.alignments[0] + let translated = document.alignments[1] + + println("\n1. Parse report metadata and alignments") + println(" " + document.summary()) + println(" command: " + document.metadata.command_line) + for alignment in document.alignments { + println(" " + alignment.summary()) + } + + println("\n2. Preserve spliced vulgar operations") + let spliced_path = spliced.coordinate_path() + println(" operations: " + spliced.operation_codes()) + println(" target path: " + coordinate_text(spliced_path[0])) + println(" query path: " + coordinate_text(spliced_path[1])) + let spliced_counts = spliced.counts() + println( + " introns=" + + spliced_counts.introns.to_string() + + ", intron units=" + + spliced_counts.intron_units.to_string() + + ", matching operations=" + + spliced_counts.matching_operations.to_string(), + ) + + println("\n3. Map residues across an intron") + println( + " query residue 6 -> target residue " + + spliced.query_position_to_target(6).unwrap_or(-1).to_string(), + ) + println( + " target residue 108 -> query residue " + + spliced.target_position_to_query(108).unwrap_or(-1).to_string() + + " (-1 denotes an intron)", + ) + + println("\n4. Handle protein-to-DNA reverse translation") + let translated_path = translated.coordinate_path() + println(" target path: " + coordinate_text(translated_path[0])) + println(" query path: " + coordinate_text(translated_path[1])) + for residue in 0.. reverse-strand nucleotide " + + translated.query_position_to_target(residue).unwrap_or(-1).to_string(), + ) + } + + println("\n5. Write canonical vulgar and cigar reports") + let vulgar = @src.align_exonerate_write(document, format="vulgar") catch { + AlignExonerateError(message) => + abort("vulgar serialization failed: " + message) + } + let cigar = @src.align_exonerate_write(document, format="cigar") catch { + AlignExonerateError(message) => + abort("cigar serialization failed: " + message) + } + let vulgar_round_trip = @src.align_exonerate_parse(vulgar) catch { + AlignExonerateError(message) => + abort("vulgar round trip failed: " + message) + } + let cigar_round_trip = @src.align_exonerate_parse(cigar) catch { + AlignExonerateError(message) => abort("cigar round trip failed: " + message) + } + println( + " vulgar preserves document: " + + (vulgar_round_trip == document).to_string(), + ) + println( + " cigar preserves first coordinate path: " + + (cigar_round_trip.alignments[0].coordinate_path() == + spliced.coordinate_path()).to_string(), + ) + println(" vulgar bytes: " + vulgar.length().to_string()) + println(" cigar bytes: " + cigar.length().to_string()) + + println("\n=== Demo Complete ===") +} diff --git a/examples/align_exonerate_demo/moon.pkg b/examples/align_exonerate_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_exonerate_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_maf_demo/main.mbt b/examples/align_maf_demo/main.mbt new file mode 100644 index 00000000..8ab22b35 --- /dev/null +++ b/examples/align_maf_demo/main.mbt @@ -0,0 +1,155 @@ +///| +fn maf_integer_array_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < values.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.maf Demo ===") + + let document = @src.align_maf_parse(@src.align_maf_example_text()) catch { + AlignMafError(message) => abort("MAF parsing failed: " + message) + } + let track = match document.track { + Some(value) => value + None => abort("example MAF track metadata is missing") + } + + println("\n1. Parse MAF header, track metadata, and a/s/i/e/q records") + println( + " version=" + + document.version + + ", scoring=" + + document.scoring.unwrap_or("none") + + ", program=" + + document.program.unwrap_or("none"), + ) + println( + " track=" + + track.name.unwrap_or("unnamed") + + ", visibility=" + + track.visibility.unwrap_or("default") + + ", species=" + + track.species_order.length().to_string(), + ) + println( + " blocks=" + + document.blocks.length().to_string() + + ", comments=" + + document.comments.length().to_string(), + ) + let first = document.blocks[0] + let reference = first.components[0] + let insertion = match reference.insertion { + Some(value) => + value.left_status + + ":" + + value.left_count.to_string() + + "/" + + value.right_status + + ":" + + value.right_count.to_string() + None => "none" + } + println( + " reference=" + + reference.source + + ", sequence=" + + reference.sequence() + + ", insertion=" + + insertion + + ", quality=" + + reference.quality.unwrap_or("none"), + ) + println( + " empty component=" + + first.empty_components[0].source + + ", status=" + + first.empty_components[0].status, + ) + + println("\n2. Build absolute coordinate paths across plus/minus strands") + let path = @src.align_maf_coordinate_path(first) + for point in path { + println( + " column " + + point.column.to_string() + + " -> " + + maf_integer_array_text(point.positions), + ) + } + let to_mouse = @src.align_maf_map_position( + first, "hg38.chr7", 100, "mm39.chr5", + ) catch { + AlignMafError(message) => abort("MAF position mapping failed: " + message) + } + let to_human = @src.align_maf_map_position( + first, "mm39.chr5", 449, "hg38.chr7", + ) catch { + AlignMafError(message) => abort("MAF reverse mapping failed: " + message) + } + println( + " hg38:100 -> mm39:" + + to_mouse.unwrap_or(-1).to_string() + + "; mm39:449 -> hg38:" + + to_human.unwrap_or(-1).to_string(), + ) + + println("\n3. Summarize the multi-block document") + let summary = document.summary() + println( + " components=" + + summary.component_count.to_string() + + ", empty=" + + summary.empty_component_count.to_string() + + ", aligned columns=" + + summary.aligned_columns.to_string(), + ) + println( + " reference bases=" + + summary.reference_bases.to_string() + + ", distinct sources=" + + summary.source_count.to_string(), + ) + + println("\n4. Query a MafIndex-style half-open reference interval") + let index = @src.AlignMafIndex::create(document, "hg38.chr7") catch { + AlignMafError(message) => abort("MAF index creation failed: " + message) + } + let hits = index.search(103, 111) catch { + AlignMafError(message) => abort("MAF interval search failed: " + message) + } + println( + " entries=" + + index.entry_count().to_string() + + ", overlaps [103, 111)=" + + hits.length().to_string(), + ) + + println("\n5. Splice multiple exons with missing-data filling") + let spliced = index.get_spliced([102, 110], [106, 113]) catch { + AlignMafError(message) => abort("MAF splicing failed: " + message) + } + println(" hg38.chr7: " + spliced.sequence("hg38.chr7").unwrap_or("missing")) + println(" mm39.chr5: " + spliced.sequence("mm39.chr5").unwrap_or("missing")) + println(" output columns: " + spliced.columns.to_string()) + + println("\n6. Write canonical MAF and parse it again") + let canonical = @src.align_maf_write(document) catch { + AlignMafError(message) => abort("MAF serialization failed: " + message) + } + let round_trip = @src.align_maf_parse(canonical) catch { + AlignMafError(message) => abort("MAF round-trip parsing failed: " + message) + } + println(" output bytes: " + canonical.length().to_string()) + println(" round-trip preserved: " + (round_trip == document).to_string()) +} diff --git a/examples/align_maf_demo/moon.pkg b/examples/align_maf_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_maf_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_mauve_demo/main.mbt b/examples/align_mauve_demo/main.mbt new file mode 100644 index 00000000..e1d7cb73 --- /dev/null +++ b/examples/align_mauve_demo/main.mbt @@ -0,0 +1,106 @@ +///| +fn align_mauve_demo_document() -> @bio.AlignMauveDocument { + @bio.align_mauve_parse(@bio.align_mauve_example_text()) catch { + AlignMauveError(message) => + abort("failed to parse XMFA example: " + message) + } +} + +///| +fn align_mauve_demo_path(values : Array[Int]) -> String { + "[" + values.map(fn(value) { value.to_string() }).join(", ") + "]" +} + +///| +fn main { + println("=== Biopython Bio.Align.mauve XMFA Demo ===") + let document = align_mauve_demo_document() + let summary = @bio.align_mauve_summary(document) + println("\n1. Strict document and locally collinear blocks") + println( + " format=" + + document.format_version() + + ", identifiers=" + + document.identifiers().join(", "), + ) + println( + " sources=" + + summary.sources().to_string() + + ", blocks=" + + summary.blocks().to_string() + + ", rows=" + + summary.rows().to_string(), + ) + + let first = document.blocks()[0] + let compact = first.compact_coordinates() + println("\n2. Forward and reverse coordinate paths") + println( + " first block: " + + first.row_count().to_string() + + " rows x " + + first.width().to_string() + + " columns", + ) + println(" reverse row compact path: " + align_mauve_demo_path(compact[0])) + println(" forward row compact path: " + align_mauve_demo_path(compact[2])) + + let counts = first.counts() + println("\n3. Pairwise alignment statistics") + println( + " aligned=" + + counts.aligned().to_string() + + ", identities=" + + counts.identities().to_string() + + ", mismatches=" + + counts.mismatches().to_string() + + ", gaps=" + + counts.gaps().to_string() + + ", gap opens=" + + counts.gap_opens().to_string(), + ) + + let index = @bio.align_mauve_index(document) + let overlaps = index.query("0", 1, 10) catch { + AlignMauveError(message) => abort("interval query failed: " + message) + } + let positions = @bio.align_mauve_map_position(document, "0", "2", 48) catch { + AlignMauveError(message) => abort("position projection failed: " + message) + } + println("\n4. Interval index and cross-sequence projection") + println(" source 0 [1,10) overlaps: " + overlaps.length().to_string()) + if positions.length() > 0 { + println( + " source 0 position 48 -> source 2 position " + + positions[0].target_position().to_string() + + " in block " + + positions[0].block_index().to_string(), + ) + } + + let reconstructed = @bio.align_mauve_reconstruct_sequence(document, "0") catch { + AlignMauveError(message) => + abort("sequence reconstruction failed: " + message) + } + println("\n5. Forward source reconstruction") + println( + " source 0 length=" + + reconstructed.length().to_string() + + ", sequence=" + + reconstructed, + ) + + let written = @bio.align_mauve_write(document, line_width=20) catch { + AlignMauveError(message) => abort("XMFA write failed: " + message) + } + let reparsed = @bio.align_mauve_parse(written) catch { + AlignMauveError(message) => abort("XMFA round-trip failed: " + message) + } + println("\n6. Canonical wrapped writer") + println( + " round-trip blocks=" + + reparsed.block_count().to_string() + + ", equivalent=" + + (reparsed == document).to_string(), + ) +} diff --git a/examples/align_mauve_demo/moon.pkg b/examples/align_mauve_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/align_mauve_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/align_nexus_demo/main.mbt b/examples/align_nexus_demo/main.mbt new file mode 100644 index 00000000..db5cccc9 --- /dev/null +++ b/examples/align_nexus_demo/main.mbt @@ -0,0 +1,134 @@ +///| +fn nexus_integer_array_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < values.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.nexus Demo ===") + + let alignment = @src.align_nexus_parse(@src.align_nexus_example_text()) catch { + AlignNexusError(message) => abort("NEXUS parsing failed: " + message) + } + + println("\n1. Parse an interleaved DATA matrix") + println(" " + alignment.summary()) + println( + " datatype: " + + alignment.metadata.data_type.code() + + " (" + + alignment.metadata.data_type.molecule_type() + + ")", + ) + let match_character = match alignment.metadata.match_character { + Some(value) => value + None => "none" + } + println( + " missing=" + + alignment.metadata.missing_character + + ", gap=" + + alignment.metadata.gap_character + + ", matchchar=" + + match_character, + ) + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": " + + sequence.aligned_sequence + + " -> " + + sequence.sequence, + ) + } + + println("\n2. Inspect coordinate paths and residue mapping") + let path = alignment.coordinate_path() + for row = 0; row < alignment.num_sequences(); row = row + 1 { + println( + " " + + alignment.sequences[row].id + + ": " + + nexus_integer_array_text(path[row]), + ) + } + println( + " reference residue 5 -> query one residue " + + alignment.map_position(0, 1, 5).unwrap_or(-1).to_string(), + ) + println( + " reference residue 7 -> query one residue " + + alignment.map_position(0, 1, 7).unwrap_or(-1).to_string() + + " (-1 denotes a gap)", + ) + + println("\n3. Calculate all-pairs statistics") + let counts = alignment.counts() + println( + " pairs=" + + counts.pairs.to_string() + + ", aligned=" + + counts.aligned.to_string() + + ", identities=" + + counts.identities.to_string() + + ", gaps=" + + counts.gap_columns.to_string(), + ) + let consensus = alignment.consensus(minimum_fraction=0.67) catch { + AlignNexusError(message) => abort("consensus failed: " + message) + } + println(" consensus: " + consensus) + println( + " occupancy at source gap column: " + alignment.occupancy()[4].to_string(), + ) + + println("\n4. Write canonical interleaved NEXUS and parse it again") + let canonical = @src.align_nexus_write( + alignment, + interleave=Some(true), + block_width=4, + ) catch { + AlignNexusError(message) => abort("NEXUS serialization failed: " + message) + } + let round_trip = @src.align_nexus_parse(canonical) catch { + AlignNexusError(message) => abort("round-trip parsing failed: " + message) + } + println(" output bytes: " + canonical.length().to_string()) + println( + " quoted IDs and rows preserved: " + + (round_trip.sequences == alignment.sequences).to_string(), + ) + println( + " coordinate path preserved: " + + (round_trip.coordinate_path() == alignment.coordinate_path()).to_string(), + ) + + println("\n5. Remove all-gap source columns during construction") + let compact = @src.align_nexus_from_aligned( + ["reference", "query"], + ["A-C--GT", "ATC--GT"], + @src.AlignNexusDna, + ) catch { + AlignNexusError(message) => + abort("alignment construction failed: " + message) + } + println( + " source columns=" + + compact.source_alignment_length().to_string() + + ", alignment columns=" + + compact.alignment_length().to_string() + + ", removed=" + + compact.metadata.removed_all_gap_columns.to_string(), + ) + println(" compact reference: " + compact.sequences[0].aligned_sequence) +} diff --git a/examples/align_nexus_demo/moon.pkg b/examples/align_nexus_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_nexus_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_phylip_demo/main.mbt b/examples/align_phylip_demo/main.mbt new file mode 100644 index 00000000..0cf00789 --- /dev/null +++ b/examples/align_phylip_demo/main.mbt @@ -0,0 +1,101 @@ +///| +fn phylip_demo_parse(text : String) -> @bio.AlignPhylipAlignment { + @bio.align_phylip_parse(text) catch { + AlignPhylipError(message) => abort("PHYLIP parse failed: " + message) + } +} + +///| +fn phylip_demo_write( + alignment : @bio.AlignPhylipAlignment, + layout : @bio.AlignPhylipLayout, +) -> String { + @bio.align_phylip_write(alignment, layout~, block_width=4, group_width=0) catch { + AlignPhylipError(message) => abort("PHYLIP write failed: " + message) + } +} + +///| +fn phylip_demo_path(path : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < path.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(path[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.phylip Offline Demo ===") + let alignment = phylip_demo_parse(@bio.align_phylip_example_text()) + + println("\n1. Automatic layout detection and fixed-width rows") + println(" " + alignment.summary()) + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": aligned=" + + sequence.aligned_sequence + + ", ungapped=" + + sequence.sequence, + ) + } + + println("\n2. Biopython-compatible identifier normalization") + let normalized = @bio.align_phylip_normalize_id("Taxon[1]:extra") catch { + AlignPhylipError(message) => + abort("PHYLIP name normalization failed: " + message) + } + println(" Taxon[1]:extra -> " + normalized) + + println("\n3. Coordinate path and cross-row projection") + let paths = alignment.coordinate_path() + for row = 0; row < paths.length(); row = row + 1 { + println( + " " + alignment.sequences[row].id + ": " + phylip_demo_path(paths[row]), + ) + } + println( + " reference residue 3 -> query_one residue " + + alignment.map_position(0, 1, 3).unwrap_or(-1).to_string(), + ) + + println("\n4. Pairwise statistics and majority consensus") + let counts = alignment.pair_counts(0, 1) catch { + AlignPhylipError(message) => abort("PHYLIP counting failed: " + message) + } + let consensus = alignment.majority_consensus(minimum_fraction=0.5) catch { + AlignPhylipError(message) => abort("PHYLIP consensus failed: " + message) + } + println( + " aligned=" + + counts.aligned.to_string() + + ", identities=" + + counts.identities.to_string() + + ", gap opens=" + + counts.gap_opens.to_string() + + ", identity=" + + counts.identity().to_string(), + ) + println(" majority=" + consensus) + + println("\n5. Canonical sequential and interleaved writing") + let sequential = phylip_demo_write(alignment, @bio.PhylipSequential) + let interleaved = phylip_demo_write(alignment, @bio.PhylipInterleaved) + let reparsed = phylip_demo_parse(interleaved) + let sequential_lines = sequential.split("\n").to_array() + println(" sequential first row=" + sequential_lines[1].to_owned()) + println( + " stable interleaved round-trip=" + + (phylip_demo_write(reparsed, @bio.PhylipInterleaved) == interleaved).to_string(), + ) + println( + "\nThe demo parses and writes existing artifacts; it does not launch an aligner.", + ) +} diff --git a/examples/align_phylip_demo/moon.pkg b/examples/align_phylip_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/align_phylip_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/align_psl_demo/main.mbt b/examples/align_psl_demo/main.mbt new file mode 100644 index 00000000..1830867c --- /dev/null +++ b/examples/align_psl_demo/main.mbt @@ -0,0 +1,112 @@ +///| +fn parse_psl(text : String) -> @src.PslDocument { + @src.psl_parse(text) catch { + PslError(message) => abort("PSL parsing failed: " + message) + } +} + +///| +fn no_header(pslx : Bool) -> @src.PslWriteConfig { + @src.PslWriteConfig::create(header=false, pslx~) catch { + PslError(message) => abort("PSL writer configuration failed: " + message) + } +} + +///| +fn main { + println("=== Biopython Bio.Align.psl Demo ===") + + println("\n1. Write and parse a multi-record PSL document") + let alignments = @src.psl_example_data() catch { + PslError(message) => abort("PSL example construction failed: " + message) + } + let text = @src.psl_write(alignments) catch { + PslError(message) => abort("PSL writing failed: " + message) + } + let document = parse_psl(text) + let summary = document.summary() + println( + " records=" + + summary.alignment_count.to_string() + + ", nucleotide=" + + summary.nucleotide_count.to_string() + + ", translated=" + + summary.protein_count.to_string(), + ) + println( + " aligned query units=" + + summary.aligned_query_units.to_string() + + ", query insert bases=" + + summary.query_insert_bases.to_string() + + ", target insert bases=" + + summary.target_insert_bases.to_string(), + ) + + println("\n2. Map a reverse-strand nucleotide query") + let reverse = document.query("readReverse")[0] + println(" strand reverse=" + reverse.is_reverse().to_string()) + match reverse.target_to_query(4) { + Some(position) => + println(" target base 4 maps to query base " + position.to_string()) + None => println(" target base 4 is outside aligned blocks") + } + match reverse.query_to_target_interval(7) { + Some((start, end)) => + println( + " query base 7 maps to target interval [" + + start.to_string() + + ", " + + end.to_string() + + ")", + ) + None => println(" query base 7 is outside aligned blocks") + } + + println("\n3. Emit and parse PSLX block sequences") + let protein = alignments[2].recount() catch { + PslError(message) => abort("translated recount failed: " + message) + } + let pslx = protein.format(config=no_header(true)) catch { + PslError(message) => abort("PSLX formatting failed: " + message) + } + let pslx_alignment = parse_psl(pslx).alignments[0] + println( + " PSLX blocks=" + + pslx_alignment.blocks().length().to_string() + + ", query fragment=" + + pslx_alignment.query_block_sequences[0] + + ", translated target fragment=" + + pslx_alignment.target_block_sequences[0], + ) + + println("\n4. Recount a reverse-target translated alignment") + let translated_reverse = @src.PslAlignment::create( + "codingReverse", + 18, + "peptide", + 4, + [18, 12, 9, 3], + [0, 2, 2, 4], + sequence_type=@src.PslProtein, + target_sequence="AAATCCAAAAAAAGCCAT", + query_sequence="MAFG", + ) catch { + PslError(message) => + abort("reverse translated alignment failed: " + message) + } + let recounted = translated_reverse.recount() catch { + PslError(message) => abort("reverse translated recount failed: " + message) + } + let translated_line = recounted.format(config=no_header(false)) catch { + PslError(message) => abort("translated PSL formatting failed: " + message) + } + println( + " strand=+-, matches=" + + recounted.matches.to_string() + + ", mismatches=" + + recounted.mismatches.to_string(), + ) + println(" " + translated_line) + + println("=== Demo Complete ===") +} diff --git a/examples/align_psl_demo/moon.pkg b/examples/align_psl_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_psl_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_sam_demo/main.mbt b/examples/align_sam_demo/main.mbt new file mode 100644 index 00000000..9e7a59ca --- /dev/null +++ b/examples/align_sam_demo/main.mbt @@ -0,0 +1,153 @@ +///| +fn parse_sam(text : String) -> @src.AlignSamDocument { + @src.align_sam_parse(text) catch { + AlignSamError(message) => abort("SAM parsing failed: " + message) + } +} + +///| +fn demo_reference() -> String { + "N".repeat(4) + "ACGTTCGGAA" + "N".repeat(26) +} + +///| +fn optional_int(value : Int?) -> String { + match value { + Some(item) => item.to_string() + None => "unknown" + } +} + +///| +fn coordinate_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index in 0.. 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.sam Demo ===") + + println("\n1. Parse headers, references, and records") + let document = parse_sam(@src.align_sam_example_data()) + println( + " headers=" + + document.headers.length().to_string() + + ", references=" + + document.references.length().to_string() + + ", alignments=" + + document.alignments.length().to_string(), + ) + match document.reference("chr1") { + Some(reference) => + println( + " reference " + + reference.name + + ", length=" + + reference.length.to_string(), + ) + None => abort("chr1 was not indexed") + } + + println("\n2. Inspect a clipped and gapped forward alignment") + let forward = document.alignments[0] + let stats = forward.stats() + println(" " + forward.summary()) + println( + " CIGAR=" + + @src.align_sam_cigar_string(forward.cigar) + + ", query bases=" + + stats.query_consumed.to_string() + + ", reference bases=" + + stats.reference_consumed.to_string(), + ) + match forward.target_to_query(8) { + Some(position) => + println(" target base 8 maps to query base " + position.to_string()) + None => println(" target base 8 lies in a gap") + } + let md = match forward.string_tag("MD") { + Some(value) => value + None => "unknown" + } + println(" NM=" + optional_int(forward.computed_nm()) + ", MD=" + md) + + println("\n3. Reconstruct aligned rows and calculate MD") + let rows = forward.aligned_rows(reference_sequence=demo_reference()) catch { + AlignSamError(message) => abort("row reconstruction failed: " + message) + } + let calculated_md = @src.align_sam_calculate_md(forward, demo_reference()) catch { + AlignSamError(message) => abort("MD calculation failed: " + message) + } + println(" target " + rows.0) + println(" query " + rows.1) + println(" calculated MD=" + calculated_md) + + println("\n4. Inspect reverse-strand and spliced coordinates") + let reverse = document.alignments[1] + println( + " reverse=" + + reverse.is_reverse().to_string() + + ", biological query=" + + reverse.query_sequence, + ) + println( + " query path " + + coordinate_text(reverse.query_coordinates) + + ", skipped reference bases=" + + reverse.stats().skipped.to_string(), + ) + match reverse.target_to_query(27) { + Some(position) => + println( + " target base 27 maps to reverse query base " + position.to_string(), + ) + None => println(" target base 27 lies in the skipped interval") + } + + println("\n5. Construct and round-trip a typed SAM record") + let tags = [ + @src.align_sam_integer_tag("NM", 0), + @src.align_sam_string_tag("RG", "demo"), + @src.align_sam_float_array_tag("BF", [0.25, 0.75]), + ] catch { + AlignSamError(message) => abort("tag construction failed: " + message) + } + let created = @src.align_sam_create( + "created_read", + "chr1", + 40, + 15, + "AACGTT", + "1S4M1S", + reverse=true, + mapq=50, + qualities=[30, 31, 32, 33, 34, 35], + tags~, + ) catch { + AlignSamError(message) => abort("alignment construction failed: " + message) + } + let record = @src.align_sam_format(created) catch { + AlignSamError(message) => abort("SAM formatting failed: " + message) + } + let reparsed = parse_sam("@SQ\tSN:chr1\tLN:40\n" + record).alignments[0] + println(" " + record) + println( + " round-trip query=" + + reparsed.query_sequence + + ", mapq=" + + optional_int(reparsed.mapq) + + ", tags=" + + reparsed.tags.length().to_string(), + ) + + println("=== Demo Complete ===") +} diff --git a/examples/align_sam_demo/moon.pkg b/examples/align_sam_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_sam_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_stockholm_demo/main.mbt b/examples/align_stockholm_demo/main.mbt new file mode 100644 index 00000000..30d4676d --- /dev/null +++ b/examples/align_stockholm_demo/main.mbt @@ -0,0 +1,119 @@ +///| +fn stockholm_integer_array_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < values.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.stockholm Demo ===") + + let alignment = @src.align_stockholm_parse( + @src.align_stockholm_example_text(), + ) catch { + AlignStockholmError(message) => + abort("Stockholm parsing failed: " + message) + } + + println("\n1. Parse the official Pfam HAT record") + println(" " + alignment.summary()) + println( + " family: " + + alignment.annotation("identifier").unwrap_or("unknown") + + " (" + + alignment.annotation("accession").unwrap_or("unknown") + + ")", + ) + println( + " references=" + + alignment.references.length().to_string() + + ", database references=" + + alignment.database_references.length().to_string(), + ) + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": " + + sequence.aligned_sequence + + " -> " + + sequence.sequence, + ) + } + + println("\n2. Inspect insertion operations and coordinates") + println(" operations: " + alignment.operations) + let path = alignment.coordinate_path() + for row = 0; row < alignment.num_sequences(); row = row + 1 { + println( + " " + + alignment.sequences[row].id + + ": " + + stockholm_integer_array_text(path[row]), + ) + } + println( + " row 3 residue 17 -> row 1 residue " + + alignment.map_position(2, 0, 17).unwrap_or(-1).to_string() + + " (-1 denotes a gap)", + ) + + println("\n3. Calculate alignment statistics") + let counts = alignment.counts() + println( + " pairs=" + + counts.pairs.to_string() + + ", aligned=" + + counts.aligned.to_string() + + ", identities=" + + counts.identities.to_string() + + ", gap columns=" + + counts.gap_columns.to_string(), + ) + let consensus = alignment.consensus() catch { + AlignStockholmError(message) => abort("consensus failed: " + message) + } + println(" consensus: " + consensus) + println( + " insertion-column occupancy: " + alignment.occupancy()[17].to_string(), + ) + + println("\n4. Write canonical Stockholm and parse it again") + let canonical = @src.align_stockholm_write(alignment) catch { + AlignStockholmError(message) => + abort("Stockholm serialization failed: " + message) + } + let round_trip = @src.align_stockholm_parse(canonical) catch { + AlignStockholmError(message) => + abort("Stockholm round-trip parsing failed: " + message) + } + println(" output bytes: " + canonical.length().to_string()) + println( + " rows preserved: " + + (round_trip.sequences == alignment.sequences).to_string(), + ) + println( + " operations preserved: " + + (round_trip.operations == alignment.operations).to_string(), + ) + + println("\n5. Remove all-gap source columns during construction") + let compact = @src.align_stockholm_from_aligned(["alpha", "beta"], [ + "A.-C--G", "AT-C--G", + ]) catch { + AlignStockholmError(message) => + abort("Stockholm construction failed: " + message) + } + println(" " + compact.summary()) + println(" operations: " + compact.operations) + println(" alpha: " + compact.sequences[0].aligned_sequence) + println(" beta: " + compact.sequences[1].aligned_sequence) +} diff --git a/examples/align_stockholm_demo/moon.pkg b/examples/align_stockholm_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_stockholm_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/align_tabular_demo/main.mbt b/examples/align_tabular_demo/main.mbt new file mode 100644 index 00000000..3e35a1e3 --- /dev/null +++ b/examples/align_tabular_demo/main.mbt @@ -0,0 +1,119 @@ +///| +fn parse_document(text : String) -> @src.AlignTabularDocument { + @src.align_tabular_parse(text) catch { + AlignTabularError(message) => abort("tabular parsing failed: " + message) + } +} + +///| +fn show_path(values : Array[Int]) -> String { + let mut result = "[" + for index in 0.. 0 { + result = result + ", " + } + result = result + values[index].to_string() + } + result + "]" +} + +///| +fn main { + println("=== Biopython Bio.Align.tabular Demo ===") + + println("\n1. Reconstruct BLAST outfmt 7 BTOP coordinate paths") + let blast = "# BLASTP 2.14.1+\n" + + "# Query: q1 protein query\n" + + "# Database: proteins\n" + + "# Fields: query id, subject id, % identity, alignment length, mismatches, gap opens, q. start, q. end, s. start, s. end, evalue, bit score, BTOP, query length, subject length, query seq, subject seq\n" + + "# 2 hits found\n" + + "q1\ts1\t75.0\t8\t2\t2\t3\t9\t10\t16\t1e-5\t42.0\t2AC1-GT-2\t20\t30\tAAAT-TGG\tAACTG-GG\n" + + "q1\ts2\t75.0\t8\t2\t2\t3\t9\t4\t10\t0.02\t31.0\t2AC1-GT-2\t20\t25\tAAAT-TGG\tAACTG-GG\n" + + "# BLAST processed 1 queries\n" + let blast_document = parse_document(blast) + let query = blast_document.queries[0] + println( + " program=" + + query.metadata.program + + ", query=" + + query.query_id + + ", hits=" + + query.alignments.length().to_string(), + ) + for alignment in query.alignments { + println(" " + alignment.summary()) + println( + " target path=" + + show_path(alignment.target_coordinates) + + ", query path=" + + show_path(alignment.query_coordinates), + ) + } + let best = query.best_by_evalue() catch { + AlignTabularError(message) => abort("best-hit selection failed: " + message) + } + match best { + Some(alignment) => println(" best E-value target=" + alignment.target_id) + None => println(" no E-value-bearing hit") + } + let coordinate = query.alignments[0].coordinate_alignment() catch { + AlignTabularError(message) => + abort("coordinate conversion failed: " + message) + } + println( + " coordinate alignment columns=" + + show_path(coordinate.target_coordinates) + + " / " + + show_path(coordinate.query_coordinates), + ) + + println("\n2. Decode FASTA -m 8CC aln_code") + let fasta = "# fasta36 -q -m 8CC query.aa database.aa\n" + + "# FASTA 36.3.8h May, 2020\n" + + "# Query: qf protein query - 218 aa\n" + + "# Database: database.aa\n" + + "# Fields: query id, subject id, % identity, alignment length, mismatches, gap opens, q. start, q. end, s. start, s. end, evalue, bit score, aln_code\n" + + "# 1 hits found\n" + + "qf\tsf\t50.0\t9\t4\t2\t4\t10\t6\t13\t1e-4\t30.0\t3M2I1D3M\n" + + "# FASTA processed 1 queries\n" + let fasta_document = parse_document(fasta) + let fasta_alignment = fasta_document.queries[0].alignments[0] + println(" command=" + fasta_alignment.metadata.command_line) + println( + " " + + fasta_alignment.summary() + + ", target gaps=" + + fasta_alignment.target_gap_residues().to_string() + + ", query gaps=" + + fasta_alignment.query_gap_residues().to_string(), + ) + + println("\n3. Preserve reverse translated TBLASTX coordinates") + let translated = "# TBLASTX 2.14.1+\n" + + "# Query: qn translated query\n" + + "# Database: nucleotides\n" + + "# Fields: query id, subject id, query length, subject length, q. start, q. end, s. start, s. end, alignment length, BTOP\n" + + "# 1 hits found\n" + + "qn\tsn\t60\t90\t30\t19\t45\t34\t4\t4\n" + let translated_alignment = parse_document(translated).queries[0].alignments[0] + println( + " query units/residue=" + + translated_alignment.query_units_per_residue.to_string() + + ", target units/residue=" + + translated_alignment.target_units_per_residue.to_string(), + ) + println( + " reverse query=" + + translated_alignment.query_is_reverse().to_string() + + ", reverse target=" + + translated_alignment.target_is_reverse().to_string(), + ) + println( + " target path=" + + show_path(translated_alignment.target_coordinates) + + ", query path=" + + show_path(translated_alignment.query_coordinates), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/align_tabular_demo/moon.pkg b/examples/align_tabular_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/align_tabular_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/alignment_counts_demo/main.mbt b/examples/alignment_counts_demo/main.mbt new file mode 100644 index 00000000..78f02f79 --- /dev/null +++ b/examples/alignment_counts_demo/main.mbt @@ -0,0 +1,131 @@ +///| +fn show_optional_int(value : Int?) -> String { + match value { + Some(number) => number.to_string() + None => "unknown" + } +} + +///| +fn show_optional_double(value : Double?) -> String { + match value { + Some(number) => number.to_string() + None => "unknown" + } +} + +///| +fn main { + println("=== Biopython Alignment.counts Demo ===") + + println("\n1. Classify and score an internal deletion") + let gapped = @src.coordinate_pairwise_alignment( + "target", + "GAACT", + "query", + "GAT", + [0, 2, 4, 5], + [0, 2, 2, 3], + ) catch { + _ => abort("failed to build gapped alignment") + } + let affine = @src.AlignmentGapScores::affine(-5.0, -1.0) catch { + _ => abort("failed to configure affine gap scores") + } + let scored = @src.AlignmentCountsConfig::scored( + match_score=2.0, + mismatch_score=-1.0, + gap_scores=affine, + ) catch { + _ => abort("failed to configure alignment scoring") + } + let counts = gapped.alignment_counts(config=scored) catch { + _ => abort("failed to calculate detailed counts") + } + println(" " + counts.summary()) + println( + " internal deletions=" + + counts.internal_deletions().to_string() + + " (open=" + + counts.open_internal_deletions.to_string() + + ", extend=" + + counts.extend_internal_deletions.to_string() + + ")", + ) + println( + " substitution=" + + show_optional_double(counts.substitution_score) + + ", gap=" + + show_optional_double(counts.gap_score) + + ", total=" + + show_optional_double(counts.score), + ) + + println("\n2. Count positives with the official BLOSUM45 matrix") + let protein = @src.coordinate_pairwise_alignment( + "protein_1", + "EPQSDPSVEPPLSQETFSDLWKLLPE", + "protein_2", + "EPSSETGMDPPLSQETFEDLWSLLPD", + [0, 26], + [0, 26], + ) catch { + _ => abort("failed to build protein alignment") + } + let protein_counts = protein.alignment_counts( + config=@src.AlignmentCountsConfig::with_matrix(@src.blosum45()), + ) catch { + _ => abort("failed to score protein alignment") + } + println( + " identities=" + + protein_counts.identities.to_string() + + ", mismatches=" + + protein_counts.mismatches.to_string() + + ", positives=" + + show_optional_int(protein_counts.positives) + + ", substitution score=" + + show_optional_double(protein_counts.substitution_score), + ) + + println("\n3. Read residues through a reverse-strand coordinate path") + let reverse = @src.coordinate_pairwise_alignment( + "reference", + "GAACT", + "reverse_read", + "ATC", + [0, 2, 4, 5], + [3, 1, 1, 0], + ) catch { + _ => abort("failed to build reverse-strand alignment") + } + let reverse_counts = reverse.alignment_counts() catch { + _ => abort("failed to count reverse-strand alignment") + } + println(" " + reverse_counts.summary()) + println( + " reverse internal deletions=" + + reverse_counts.internal_deletions().to_string(), + ) + + println("\n4. Sum counts and scores over every unordered MSA row pair") + let msa = @src.coordinate_multiple_alignment(["a", "b", "c"], [ + "A-AA", "AAAA", "A--A", + ]) catch { + _ => abort("failed to build multiple alignment") + } + let msa_counts = msa.alignment_counts(config=scored) catch { + _ => abort("failed to count multiple alignment") + } + println(" " + msa_counts.summary()) + println( + " gap opens=" + + msa_counts.open_gaps().to_string() + + ", gap extensions=" + + msa_counts.extend_gaps().to_string() + + ", total score=" + + show_optional_double(msa_counts.score), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/alignment_counts_demo/moon.pkg b/examples/alignment_counts_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/alignment_counts_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/alignment_map_demo/main.mbt b/examples/alignment_map_demo/main.mbt new file mode 100644 index 00000000..b5e0833f --- /dev/null +++ b/examples/alignment_map_demo/main.mbt @@ -0,0 +1,135 @@ +///| +fn main { + println("=== Biopython Alignment.map/mapall Demo ===") + + println("\n1. Compose chromosome -> transcript -> read") + let chromosome_to_transcript = @src.coordinate_pairwise_alignment( + "chromosome", + "AAAAAAAACCCCCCCAAAAAAAAAAAGGGGGGAAAAAAAA", + "transcript", + "CCCCCCCGGGGGG", + [8, 15, 26, 32], + [0, 7, 7, 13], + ) catch { + _ => abort("failed to build chromosome-to-transcript alignment") + } + let transcript_to_read = @src.coordinate_pairwise_alignment( + "transcript", + "CCCCCCCGGGGGG", + "read", + "CCCCGGGG", + [3, 11], + [0, 8], + ) catch { + _ => abort("failed to build transcript-to-read alignment") + } + let chromosome_to_read = chromosome_to_transcript.map(transcript_to_read) catch { + _ => abort("failed to compose coordinate alignments") + } + println(" " + chromosome_to_read.summary()) + println( + " target coordinates: " + chromosome_to_read.target_coordinates.to_string(), + ) + println( + " query coordinates: " + chromosome_to_read.query_coordinates.to_string(), + ) + + println("\n2. Render the exon blocks and intron gap") + match chromosome_to_read.aligned_rows() { + Some((target, query)) => { + println(" chromosome: " + target) + println(" read: " + query) + } + None => abort("concrete alignment rows should render") + } + let counts = chromosome_to_read.counts() + println( + " blocks=" + + counts.blocks.to_string() + + ", aligned=" + + counts.aligned.to_string() + + ", query-gap bases=" + + counts.query_gap_bases.to_string(), + ) + + println("\n3. Query coordinates and serialize PSL") + println( + " chromosome position 14 -> read: " + + chromosome_to_read.target_to_query(14).to_string(), + ) + println( + " read position 4 -> chromosome: " + + chromosome_to_read.query_to_target(4).to_string(), + ) + let psl = chromosome_to_read.to_psl() catch { + _ => abort("failed to serialize PSL") + } + println(" PSL: " + psl) + + println("4. Preserve strand through composition") + let chromosome_to_reverse_transcript = @src.coordinate_pairwise_alignment( + "chromosome", + "AAAAAAAAAAAAGGGGGGGCCCCCGGGGGGAAAAAAAAAA", + "reverse_transcript", + "TCCCCCCGGGGGCCCCCCC", + [12, 31], + [19, 0], + ) catch { + _ => abort("failed to build reverse transcript alignment") + } + let reverse_transcript_to_read = @src.coordinate_pairwise_alignment( + "reverse_transcript", + "TCCCCCCGGGGGCCCCCCC", + "reverse_read", + "CCCGGGGGCC", + [4, 14], + [0, 10], + ) catch { + _ => abort("failed to build reverse read alignment") + } + let reverse_projection = chromosome_to_reverse_transcript.map( + reverse_transcript_to_read, + ) catch { + _ => abort("failed to compose reverse-strand alignment") + } + println(" " + reverse_projection.summary()) + println( + " projected query coordinates: " + + reverse_projection.query_coordinates.to_string(), + ) + + println("\n5. Project a protein MSA to a codon-aware nucleotide MSA") + let protein_msa = @src.coordinate_multiple_alignment( + ["protein_1", "protein_2"], + ["M-K", "MTK"], + ) catch { + _ => abort("failed to build protein MSA") + } + let protein_1_to_dna = @src.coordinate_pairwise_alignment( + "protein_1", + "MK", + "dna_1", + "ATGAAA", + [0, 2], + [0, 6], + ) catch { + _ => abort("failed to build first protein-to-DNA mapping") + } + let protein_2_to_dna = @src.coordinate_pairwise_alignment( + "protein_2", + "MTK", + "dna_2", + "ATGACCAAA", + [0, 3], + [0, 9], + ) catch { + _ => abort("failed to build second protein-to-DNA mapping") + } + let nucleotide_msa = protein_msa.mapall([protein_1_to_dna, protein_2_to_dna]) catch { + _ => abort("failed to project protein MSA") + } + println(" " + nucleotide_msa.summary()) + println(nucleotide_msa.to_fasta()) + + println("=== Demo Complete ===") +} diff --git a/examples/alignment_map_demo/moon.pkg b/examples/alignment_map_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/alignment_map_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/apeglm_demo/main.mbt b/examples/apeglm_demo/main.mbt new file mode 100644 index 00000000..c09f3945 --- /dev/null +++ b/examples/apeglm_demo/main.mbt @@ -0,0 +1,126 @@ +///| +fn print_gene(result : @src.ApeglmResult, name : String) -> Unit { + match result.find_gene(name) { + Some(gene) => + println( + " " + + name + + ": MLE=" + + (gene.mle[result.coefficient] / @math.ln(2.0)).to_string() + + ", MAP=" + + gene.effect(result.coefficient, log2_scale=true).to_string() + + ", posterior SD=" + + gene.standard_error(result.coefficient, log2_scale=true).to_string() + + ", FSR=" + + gene.fsr.to_string() + + ", s-value=" + + gene.s_value.to_string() + + ", FSOS=" + + gene.threshold_probability.to_string(), + ) + None => abort("example gene not found: " + name) + } +} + +///| +fn main { + println("=== Bioconductor apeglm Demo ===") + let (counts, design, dispersions, genes, coefficients) = @src.apeglm_example_data() + let config = @src.ApeglmConfig::create(coefficient=2, threshold=@math.ln(2.0)) catch { + _ => abort("failed to create apeglm configuration") + } + + println("\n1. Fit a batch-adjusted negative-binomial GLM") + let result = @src.apeglm_fit( + counts, + design, + dispersions, + config~, + gene_names=genes, + coefficient_names=coefficients, + ) catch { + _ => abort("failed to fit apeglm model") + } + let summary = result.summary(maximum_s_value=0.1) + println( + " coefficient=" + + summary.coefficient_name + + ", genes=" + + summary.gene_count.to_string() + + ", converged=" + + summary.converged_count.to_string() + + ", selected=" + + summary.selected_count.to_string(), + ) + println( + " adaptive Cauchy scale=" + + result.prior_control.prior_scale.to_string() + + " (natural-log scale)", + ) + + println("\n2. Compare unshrunk MLEs with posterior MAP estimates") + print_gene(result, "strong_up") + print_gene(result, "strong_down") + print_gene(result, "low_noisy") + print_gene(result, "stable") + + println("\n3. Rank and select effects by directional error control") + let ranked = result.ranked() + println( + " top-ranked gene=" + + ranked[0].gene_name + + ", absolute MAP=" + + ranked[0].map[result.coefficient].abs().to_string(), + ) + println( + " genes with s-value <= 0.10: " + + result.select(maximum_s_value=0.1).length().to_string(), + ) + + println("\n4. Add posterior summaries to an immutable container copy") + let assays : Map[String, Array[Array[Double]]] = Map([]) + assays["counts"] = counts + let metadata : Map[String, String] = Map([]) + metadata["source"] = "apeglm_demo" + let experiment = @src.summarized_experiment(assays, [], [], metadata) + let output = @src.apeglm_summarized_experiment( + experiment, + design, + dispersions, + config~, + gene_names=genes, + coefficient_names=coefficients, + ) catch { + _ => abort("failed to integrate apeglm with SummarizedExperiment") + } + let posterior_map = match @src.se_assay(output.experiment, "apeglm_map") { + Some(value) => value + None => abort("posterior MAP assay missing") + } + println( + " posterior MAP rows=" + + posterior_map.length().to_string() + + ", source object unchanged=" + + (@src.se_assay(experiment, "apeglm_map") is None).to_string(), + ) + + println("\n5. Export a DESeq2-style log2 table") + println(result.to_tsv(log2_scale=true)) + + println("6. Invalid model inputs produce explicit diagnostics") + let rejected = try { + ignore( + @src.apeglm_fit([[1.0, -1.0], [2.0, 3.0]], [[1.0, 0.0], [1.0, 1.0]], [ + 0.1, 0.1, + ]), + ) + false + } catch { + ApeglmError(message) => { + println(" " + message) + true + } + } + println(" malformed counts rejected=" + rejected.to_string()) + println("\n=== Demo Complete ===") +} diff --git a/examples/apeglm_demo/moon.pkg b/examples/apeglm_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/apeglm_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/banksy_demo/main.mbt b/examples/banksy_demo/main.mbt new file mode 100644 index 00000000..48d9e83f --- /dev/null +++ b/examples/banksy_demo/main.mbt @@ -0,0 +1,183 @@ +// Bioconductor Banksy-inspired spatial clustering workflow. + +///| +fn banksy_demo_config() -> @src.BanksyComputeConfig { + @src.BanksyComputeConfig::create( + max_harmonic=1, + k_geom=[4, 4], + spatial_mode=@src.banksy_knn_median(), + seed=7, + ) catch { + BanksyError(message) => abort("invalid BANKSY configuration: " + message) + } +} + +///| +fn banksy_demo_cluster_config() -> @src.BanksyClusterConfig { + @src.BanksyClusterConfig::create( + n_clusters=2, + n_starts=8, + max_iterations=100, + tolerance=1.0e-8, + seed=3, + ) catch { + BanksyError(message) => + abort("invalid BANKSY cluster configuration: " + message) + } +} + +///| +fn main { + println("=== Bioconductor Banksy Demo ===") + let (expression, coordinates, gene_names, spot_names) = @src.banksy_example_data() + let config = banksy_demo_config() + let model = @src.banksy_compute_with_config( + expression, + coordinates, + config, + gene_names~, + spot_names~, + ) catch { + BanksyError(message) => + abort("BANKSY feature computation failed: " + message) + } + + println("\n1. Compute spatial neighborhood harmonics") + println(" " + model.summary()) + println( + " first spot left_marker: expression=" + + model.expression[0][0].to_string() + + ", H0=" + + model.harmonics[0][0][0].to_string() + + ", H1=" + + model.harmonics[1][0][0].to_string(), + ) + let mut weight_sum = 0.0 + for neighbor in model.neighbors[0][0] { + weight_sum = weight_sum + neighbor.weight + } + println( + " H0 neighbors=" + + model.neighbors[0][0].length().to_string() + + ", normalized weight sum=" + + weight_sum.to_string(), + ) + + println("\n2. Compare cell-typing and spatial-domain lambda values") + let typing = @src.banksy_run( + model, + lambda=0.2, + n_components=3, + cluster_config=banksy_demo_cluster_config(), + max_harmonic=1, + scale=true, + smooth=true, + smoothing_k=3, + smoothing_threshold=0.5, + smoothing_iterations=10, + ) catch { + BanksyError(message) => abort("BANKSY cell-typing run failed: " + message) + } + let domains = @src.banksy_run( + model, + lambda=0.8, + n_components=3, + cluster_config=banksy_demo_cluster_config(), + max_harmonic=1, + scale=true, + smooth=true, + smoothing_k=3, + smoothing_threshold=0.5, + smoothing_iterations=10, + ) catch { + BanksyError(message) => abort("BANKSY domain run failed: " + message) + } + println(" lambda=0.2: " + typing.summary()) + println(" lambda=0.8: " + domains.summary()) + println( + " PCA dimensions=" + + typing.pca.n_spots().to_string() + + "x" + + typing.pca.n_components.to_string() + + ", cluster sizes=" + + typing.clustering.cluster_sizes.to_string(), + ) + let ari = @src.banksy_adjusted_rand_index( + typing.final_labels(), + domains.final_labels(), + ) catch { + BanksyError(message) => abort("BANKSY ARI failed: " + message) + } + println(" adjusted Rand index between lambda runs=" + ari.to_string()) + + println("\n3. Sweep spatial weights") + let sweep = @src.banksy_parameter_sweep( + model, + [0.0, 0.2, 0.5, 0.8], + n_components=3, + cluster_config=banksy_demo_cluster_config(), + max_harmonic=1, + scale=true, + ) catch { + BanksyError(message) => abort("BANKSY parameter sweep failed: " + message) + } + for run in sweep { + println( + " lambda=" + + run.lambda.to_string() + + ", silhouette=" + + run.clustering.silhouette.to_string() + + ", spatial agreement=" + + run.spatial_agreement.to_string(), + ) + } + + println("\n4. Enrich an immutable SpatialExperiment") + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "logcounts", expression)) + for gene_name in gene_names { + ignore(@src.se_add_row(experiment, Map([("gene_name", gene_name)]))) + } + for spot in 0.. + abort("BANKSY SpatialExperiment integration failed: " + message) + } + println( + " assays include H0=" + + integrated.experiment.assay.contains("H0").to_string() + + ", H1=" + + integrated.experiment.assay.contains("H1").to_string(), + ) + println( + " first cluster=" + + integrated.experiment.col_data[0]["banksy_cluster"] + + ", input remains unchanged=" + + (!experiment.assay.contains("H0")).to_string(), + ) + println("=== Demo Complete ===") +} diff --git a/examples/banksy_demo/moon.pkg b/examples/banksy_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/banksy_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/bigbed_demo/main.mbt b/examples/bigbed_demo/main.mbt new file mode 100644 index 00000000..00676ad2 --- /dev/null +++ b/examples/bigbed_demo/main.mbt @@ -0,0 +1,166 @@ +///| +fn format_coordinate_path(path : Array[Array[Int]]) -> String { + let output = StringBuilder::new() + output.write_string("[") + for i in 0.. 0 { + output.write_string(", ") + } + output.write_string("[") + for j in 0.. 0 { + output.write_string(", ") + } + output.write_string(path[i][j].to_string()) + } + output.write_string("]") + } + output.write_string("]") + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.bigbed Demo ===") + let targets = [ + @src.BigBedTarget::create("chr1", 1000000) catch { + _ => abort("failed to create chr1") + }, + @src.BigBedTarget::create("chr2", 500000) catch { + _ => abort("failed to create chr2") + }, + ] + let records = [ + @src.BigBedRecord::create( + "chr1", + 100, + 280, + name="transcript-A", + score=960, + strand="+", + thick_start=120, + thick_end=260, + item_rgb="40,120,220", + block_count=3, + block_sizes=[50, 40, 30], + block_starts=[0, 80, 150], + extra_fields=["GENE1"], + ) catch { + _ => abort("failed to create transcript-A") + }, + @src.BigBedRecord::create( + "chr1", + 250, + 410, + name="transcript-B", + score=870, + strand="-", + thick_start=270, + thick_end=390, + item_rgb="220,80,80", + block_count=2, + block_sizes=[60, 50], + block_starts=[0, 110], + extra_fields=["GENE1"], + ) catch { + _ => abort("failed to create transcript-B") + }, + @src.BigBedRecord::create( + "chr2", + 1000, + 1120, + name="transcript-C", + score=700, + strand="+", + block_count=2, + block_sizes=[45, 50], + block_starts=[0, 70], + extra_fields=["GENE2"], + ) catch { + _ => abort("failed to create transcript-C") + }, + ] + let gene_field = @src.BigBedField::create( + "string", + "gene", + comment="Gene identifier", + ) catch { + _ => abort("failed to create custom AutoSQL field") + } + let config = @src.BigBedWriteConfig::create( + bed_columns=12, + items_per_slot=2, + block_size=2, + compress=true, + schema_name="transcripts", + schema_comment="Indexed transcript annotations", + custom_fields=[gene_field], + ) catch { + _ => abort("failed to create BigBed configuration") + } + + println("\n1. Write and parse a compressed BigBed v4 stream") + let bytes = @src.bigbed_write(targets, records, config~) catch { + _ => abort("failed to write BigBed") + } + let file = @src.bigbed_parse(bytes) catch { + _ => abort("failed to parse BigBed") + } + println(" bytes=" + bytes.length().to_string()) + let summary = file.summary() catch { + _ => abort("failed to summarize BigBed") + } + println( + " records=\{summary.record_count}, targets=\{summary.target_count}, blocks=\{summary.data_block_count}, covered bases=\{summary.covered_bases}, compressed=\{summary.compressed}", + ) + + println("\n2. Inspect the embedded AutoSQL schema") + println(file.schema.to_auto_sql()) + + println("3. Query an overlapping chromosome interval through the R-tree") + let hits = file.search("chr1", start=240, end=300) catch { + _ => abort("failed to search BigBed") + } + for hit in hits { + let gene = match + hit.annotation(file.schema, "gene", file.header.defined_field_count) { + Some(value) => value + None => "missing" + } + println( + " " + + hit.name + + " " + + hit.chrom + + ":" + + hit.chrom_start.to_string() + + "-" + + hit.chrom_end.to_string() + + " gene=" + + gene, + ) + } + + println("\n4. Recover exon-aware reverse-strand coordinates") + let reverse = file.find_by_name("transcript-B") catch { + _ => abort("failed to query transcript name") + } + println( + " target/query path=" + format_coordinate_path(reverse[0].coordinates()), + ) + + println("\n5. Export records as BED text") + println(file.to_bed() catch { _ => abort("failed to export BED") }) + + println("\n6. Diagnose a damaged BigBed header") + let damaged = bytes.copy() + damaged[0] = 0 + try { + ignore(@src.bigbed_parse(damaged)) + abort("damaged BigBed unexpectedly parsed") + } catch { + BigBedError(message) => println(" rejected: " + message) + } + + println("\n=== Demo Complete ===") +} diff --git a/examples/bigbed_demo/moon.pkg b/examples/bigbed_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/bigbed_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/bigmaf_demo/main.mbt b/examples/bigmaf_demo/main.mbt new file mode 100644 index 00000000..f4c6f70d --- /dev/null +++ b/examples/bigmaf_demo/main.mbt @@ -0,0 +1,72 @@ +///| +fn main { + println("=== Biopython Bio.Align.bigmaf Demo ===") + let (reference, targets, blocks) = @src.bigmaf_example_data() catch { + BigMafError(message) => abort("failed to create BigMaf data: " + message) + } + let config = @src.BigMafWriteConfig::create( + compress=true, + block_size=3, + items_per_slot=2, + ) catch { + BigMafError(message) => abort("failed to configure BigMaf: " + message) + } + + println("\n1. Write and parse a compressed bed3+1 BigMaf stream") + let bytes = @src.bigmaf_write(reference, targets, blocks, config~) catch { + BigMafError(message) => abort("failed to write BigMaf: " + message) + } + let file = @src.bigmaf_parse(bytes) catch { + BigMafError(message) => abort("failed to parse BigMaf: " + message) + } + let summary = file.summary() catch { + BigMafError(message) => abort("failed to summarize BigMaf: " + message) + } + println( + " bytes=\{bytes.length()}, alignments=\{summary.alignment_count}, targets=\{summary.target_count}, components=\{summary.component_count}, compressed=\{summary.compressed}", + ) + + println("\n2. Inspect the standard bedMaf AutoSQL schema") + println(@src.bigmaf_schema().to_auto_sql()) + + println("3. Query chr7 through the BigBed R-tree") + let hits = file.search("hg38.chr7", start=95, end=110) catch { + BigMafError(message) => abort("failed to query BigMaf: " + message) + } + let block = hits[0] + let reference_component = block.reference_component() + println( + " hits=\{hits.length()}, reference=\{reference_component.source}:\{reference_component.start}-\{reference_component.start + reference_component.size}, columns=\{block.aligned_columns()}", + ) + println( + " score=\{block.score}, pass=\{block.pass_number}, identity=\{block.pairwise_identity()}", + ) + + println("\n4. Map a reference base to a reverse-strand component") + match block.map_reference_to(104, "mm39.chr5") { + Some(position) => + println(" hg38.chr7:104 -> mm39.chr5:" + position.to_string()) + None => println(" reference base maps to a gap") + } + match block.component("mm39.chr5") { + Some(component) => { + let (start, end) = component.forward_interval() + println( + " mouse forward interval=\{start}-\{end}, sequence=\{component.forward_sequence()}", + ) + } + None => abort("example mouse component is missing") + } + + println("\n5. Preserve MAF annotations and export standard MAF") + println( + " empty components=\{block.empty_components.length()}, comments=\{block.comments.length()}", + ) + println( + file.to_maf() catch { + BigMafError(message) => abort("failed to export MAF: " + message) + }, + ) + + println("=== Demo Complete ===") +} diff --git a/examples/bigmaf_demo/moon.pkg b/examples/bigmaf_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/bigmaf_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/bigpsl_demo/main.mbt b/examples/bigpsl_demo/main.mbt new file mode 100644 index 00000000..87590755 --- /dev/null +++ b/examples/bigpsl_demo/main.mbt @@ -0,0 +1,71 @@ +///| +fn main { + println("=== Biopython Bio.Align.bigpsl Demo ===") + let (targets, alignments) = @src.bigpsl_example_data() catch { + BigPslError(message) => abort("failed to create BigPsl data: " + message) + } + let config = @src.BigPslWriteConfig::create( + compress=true, + block_size=3, + items_per_slot=1, + ) catch { + BigPslError(message) => abort("failed to configure BigPsl: " + message) + } + + println("\n1. Write and parse a compressed bed12+13 BigPsl stream") + let bytes = @src.bigpsl_write(targets, alignments, config~) catch { + BigPslError(message) => abort("failed to write BigPsl: " + message) + } + let file = @src.bigpsl_parse(bytes) catch { + BigPslError(message) => abort("failed to parse BigPsl: " + message) + } + let summary = file.summary() catch { + BigPslError(message) => abort("failed to summarize BigPsl: " + message) + } + println( + " bytes=\{bytes.length()}, alignments=\{summary.alignment_count}, targets=\{summary.target_count}, nucleotide=\{summary.nucleotide_count}, protein=\{summary.protein_count}, compressed=\{summary.compressed}", + ) + + println("\n2. Inspect the standard bigPsl AutoSQL schema") + println(@src.bigpsl_schema().to_auto_sql()) + + println("3. Query chr1 through the BigBed R-tree") + let hits = file.search("chr1", start=105, end=106) catch { + BigPslError(message) => abort("failed to query BigPsl: " + message) + } + println(" hits=" + hits.length().to_string()) + for alignment in hits { + println(" " + alignment.summary()) + println( + " target chr1:105 -> query " + + alignment.target_to_query(105).to_string(), + ) + } + + println("\n4. Recover reverse-query nucleotide coordinates") + let reverse = file.find_by_query("rna-reverse") catch { + BigPslError(message) => abort("failed to find reverse query: " + message) + } + println(" " + reverse[0].summary()) + println( + " target chr1:200 -> query " + reverse[0].target_to_query(200).to_string(), + ) + + println("\n5. Map amino acids to reverse-target codon intervals") + let protein = file.find_by_query("protein-reverse") catch { + BigPslError(message) => abort("failed to find protein query: " + message) + } + println(" " + protein[0].summary()) + println( + " amino acid 0 -> target codon " + + protein[0].query_to_target_interval(0).to_string(), + ) + + println("\n6. Export standard PSL text") + println( + file.to_psl() catch { + BigPslError(message) => abort("failed to export PSL: " + message) + }, + ) + println("=== Demo Complete ===") +} diff --git a/examples/bigpsl_demo/moon.pkg b/examples/bigpsl_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/bigpsl_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/binary_cif_demo/main.mbt b/examples/binary_cif_demo/main.mbt new file mode 100644 index 00000000..b4836f31 --- /dev/null +++ b/examples/binary_cif_demo/main.mbt @@ -0,0 +1,98 @@ +///| +fn main { + println("=== Biopython Bio.PDB.binary_cif Demo ===") + + let bytes = @src.binary_cif_sample_bytes() + let file = @src.binary_cif_parse(bytes) catch { + _ => abort("failed to parse BinaryCIF sample") + } + + println("\n1. MessagePack BinaryCIF document") + println(" Encoded bytes: " + bytes.length().to_string()) + println(" Encoder: " + file.encoder) + println(" " + file.summary()) + + let block = file.get_block(0).unwrap() + let atom_site = block.get_category("atom_site").unwrap() + println("\n2. Category and column access") + println(" Block header: " + block.header) + println(" _atom_site rows: " + atom_site.row_count.to_string()) + let first_component = atom_site + .get_column("label_comp_id") + .unwrap() + .string_at(0) catch { + _ => abort("failed to read component column") + } + println( + " First component: " + first_component.unwrap(), + ) + let first_x = atom_site + .get_column("Cartn_x") + .unwrap() + .double_at(0) catch { + _ => abort("failed to read coordinate column") + } + println( + " First x coordinate: " + first_x.unwrap().to_string(), + ) + + println("\n3. CIF missing-value masks") + let altloc = atom_site.get_column("label_alt_id").unwrap() + let token0 = altloc.cif_token_at(0) catch { + _ => abort("failed to read mask row 0") + } + let token3 = altloc.cif_token_at(3) catch { + _ => abort("failed to read mask row 3") + } + let token6 = altloc.cif_token_at(6) catch { + _ => abort("failed to read mask row 6") + } + println(" Row 0 token: " + token0) + println(" Row 3 token: " + token3) + println(" Row 6 token: " + token6) + + println("\n4. PDB structure construction") + let structure = @src.binary_cif_to_structure(file) catch { + _ => abort("failed to construct PDB structure") + } + println(" Structure ID: " + structure.get_id()) + println(" Models: " + structure.get_num_models().to_string()) + println(" Chains: " + structure.get_num_chains().to_string()) + println(" Residues: " + structure.get_num_residues().to_string()) + println(" Atoms: " + structure.get_num_atoms().to_string()) + for model in structure.get_models() { + println( + " Model " + + model.get_id().to_string() + + ": " + + model.get_num_chains().to_string() + + " chain(s), " + + model.get_atoms().length().to_string() + + " atom(s)", + ) + } + + println("\n5. Atom annotations") + let atoms = structure.get_atoms() + let first = atoms[0].get_coord() + println( + " First atom " + + atoms[0].name + + ": (" + + first.x.to_string() + + ", " + + first.y.to_string() + + ", " + + first.z.to_string() + + ")", + ) + println( + " Alternate location on atom 4: " + atoms[3].altloc.to_string(), + ) + println(" Water residue is hetero: " + structure.get_residues()[2].is_het().to_string()) + + println("\nNotes") + println(" Gzip-compressed .bcif input must be decompressed before parsing.") + println(" PDB Structure chain IDs are one character; category data keeps full IDs.") + println("\n=== Demo Complete ===") +} diff --git a/examples/binary_cif_demo/moon.pkg b/examples/binary_cif_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/binary_cif_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/blast_xml_advanced_demo/main.mbt b/examples/blast_xml_advanced_demo/main.mbt new file mode 100644 index 00000000..01b7aeb7 --- /dev/null +++ b/examples/blast_xml_advanced_demo/main.mbt @@ -0,0 +1,167 @@ +///| +fn blast_demo_parse(text : String) -> @bio.BlastXmlDocument { + @bio.blast_xml_parse(text) catch { + BlastXmlError(message) => abort("BLAST XML parse failed: " + message) + } +} + +///| +fn blast_demo_write(document : @bio.BlastXmlDocument) -> String { + @bio.blast_xml_write(document) catch { + BlastXmlError(message) => abort("BLAST XML write failed: " + message) + } +} + +///| +fn blast_demo_write_xml2(document : @bio.BlastXmlDocument) -> String { + @bio.blast_xml_write_as(document, @bio.BlastXmlFormat::Xml2) catch { + BlastXmlError(message) => abort("BLAST XML conversion failed: " + message) + } +} + +///| +fn blast_demo_xml1() -> String { + "\n" + + "\n" + + "" + + "blastp" + + "BLASTP 2.15.0+" + + "offline demo" + + "swissprot" + + "Query_1" + + "protein query" + + "12" + + "" + + "BLOSUM62" + + "0.001" + + "11" + + "1" + + "" + + "" + + "1" + + "Query_1" + + "protein query" + + "12" + + "" + + "1sp|DEMO.1|" + + "demonstration targetDEMO" + + "20" + + "142.5" + + "1002e-12" + + "17" + + "18" + + "00" + + "77" + + "18" + + "MKKL-AAAMKKLTAAA" + + "MKKL AAA" + + "482816" + + "183558113" + + "72" + + "4627826878" + + "0.041" + + "0.267" + + "0.14" + + "" + + "\n" +} + +///| +fn blast_demo_xml2() -> String { + "\n" + + "" + + "blastxBLASTX 2.15.0+" + + "offline demoswissprot" + + "BLOSUM62" + + "0.001111" + + "2" + + "translated_querynegative frame query" + + "301" + + "sp|PRIMARY.1|PRIMARY" + + "primary target83333" + + "sp|SECONDARY.1|SECONDARY" + + "secondary target224308" + + "Bacillus subtilis10" + + "13060" + + "1e-83" + + "157-1" + + "243" + + "0MKTMKT" + + "" + + "\n" +} + +///| +fn main { + println("=== Biopython Bio.Blast XML1/XML2 Offline Demo ===") + let xml1 = blast_demo_parse(blast_demo_xml1()) + let record = xml1.records()[0] + let best = record.best_hit().unwrap() + let hsp = best.best_hsp().unwrap() + let statistics = record.statistics().unwrap() + println("\n1. XML1 records, best hit, and statistics") + println( + " program=" + + xml1.program() + + ", records=" + + xml1.records().length().to_string() + + ", hits=" + + xml1.total_hits().to_string(), + ) + println( + " best=" + + best.primary_description().accession() + + ", E-value=" + + hsp.evalue().to_string() + + ", identity=" + + hsp.identity_fraction().to_string(), + ) + println( + " database sequences=" + + statistics.database_sequences().to_string() + + ", letters=" + + statistics.database_letters().to_string(), + ) + + println("\n2. Compressed alignment path") + for coordinate in hsp.coordinates() { + println( + " target=" + + coordinate.target().to_string() + + ", query=" + + coordinate.query().to_string(), + ) + } + + let xml2 = blast_demo_parse(blast_demo_xml2()) + let translated_hit = xml2.records()[0].hits()[0] + let translated_hsp = translated_hit.hsps()[0] + println("\n3. XML2 descriptions and translated coordinates") + println( + " descriptions=" + + translated_hit.descriptions().length().to_string() + + ", secondary taxon=" + + translated_hit.descriptions()[1].scientific_name().unwrap_or("unknown"), + ) + println( + " query frame=" + + translated_hsp.query_frame().unwrap().to_string() + + ", coded_by=" + + translated_hsp.query_coded_by().unwrap_or("undefined"), + ) + + let canonical = blast_demo_write(xml2) + let converted = blast_demo_write_xml2(xml1) + println("\n4. Canonical writing and dialect conversion") + println( + " XML2 round-trip stable: " + + (blast_demo_write(blast_demo_parse(canonical)) == canonical).to_string(), + ) + println( + " XML1 -> XML2 records: " + + blast_demo_parse(converted).records().length().to_string(), + ) + println( + "\nThe demo parses and writes existing artifacts; it does not use NCBI or launch BLAST.", + ) +} diff --git a/examples/blast_xml_advanced_demo/moon.pkg b/examples/blast_xml_advanced_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/blast_xml_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/bluster_demo/main.mbt b/examples/bluster_demo/main.mbt new file mode 100644 index 00000000..86d16393 --- /dev/null +++ b/examples/bluster_demo/main.mbt @@ -0,0 +1,160 @@ +///| +fn bluster_demo_ints(values : Array[Int]) -> String { + values.map(fn(value) { value.to_string() }).join(", ") +} + +///| +fn bluster_demo_doubles(values : Array[Double]) -> String { + values.map(fn(value) { value.to_string() }).join(", ") +} + +///| +fn bluster_demo_optional(value : Double?) -> String { + match value { + Some(actual) => actual.to_string() + None => "NA" + } +} + +///| +fn main { + println("=== Bioconductor bluster 1.23.0 Demo ===") + + let coordinates = [ + [0.0, 0.0], + [0.1, 0.0], + [0.2, 0.1], + [0.3, 0.1], + [10.0, 10.0], + [10.1, 10.0], + [10.2, 10.1], + [10.3, 10.1], + ] + + let kmeans_config = @bio.BlusterKmeansConfig::create(2, starts=8, seed=101) catch { + _ => abort("failed to create the k-means configuration") + } + let kmeans = @bio.bluster_cluster_rows_kmeans(coordinates, kmeans_config) catch { + _ => abort("k-means clustering failed") + } + println("\n1. Deterministic multi-start K-means") + println(" clusters: " + bluster_demo_ints(kmeans.clusters)) + println( + " total within sum of squares: " + + kmeans.total_within_sum_squares.to_string(), + ) + + let graph = @bio.bluster_make_snn_graph( + coordinates, + k=2, + weighting=@bio.bluster_rank_weight(), + ) catch { + _ => abort("SNN graph construction failed") + } + let graph_config = @bio.BlusterGraphConfig::create( + k=2, + resolution=0.5, + seed=103, + ) catch { + _ => abort("failed to create the graph configuration") + } + let graph_clusters = @bio.bluster_cluster_rows_graph( + coordinates, graph_config, + ) catch { + _ => abort("graph clustering failed") + } + println("\n2. Exact SNN graph and Louvain-style optimization") + println(" graph edges: " + graph.edge_count().to_string()) + println(" clusters: " + bluster_demo_ints(graph_clusters.clusters)) + println(" modularity: " + graph_clusters.modularity.to_string()) + + let two_step_config = @bio.BlusterTwoStepConfig::create( + first_centers=4, + first_starts=5, + second_k=1, + resolution=0.5, + seed=107, + ) catch { + _ => abort("failed to create the two-step configuration") + } + let two_step = @bio.bluster_cluster_rows_two_step( + coordinates, two_step_config, + ) catch { + _ => abort("two-step clustering failed") + } + println("\n3. Two-step vector quantization") + println(" centroids: " + two_step.centroids.length().to_string()) + println(" clusters: " + bluster_demo_ints(two_step.clusters)) + + let rand = @bio.bluster_pairwise_rand( + kmeans.clusters, + graph_clusters.clusters, + ) catch { + _ => abort("Rand comparison failed") + } + let silhouette = @bio.bluster_approx_silhouette( + coordinates, + graph_clusters.clusters, + ) catch { + _ => abort("silhouette approximation failed") + } + let purity = @bio.bluster_neighbor_purity( + coordinates, + graph_clusters.clusters, + k=2, + ) catch { + _ => abort("neighbor purity failed") + } + let rmsd = @bio.bluster_cluster_rmsd(coordinates, graph_clusters.clusters) catch { + _ => abort("cluster RMSD failed") + } + println("\n4. Cluster diagnostics") + println(" adjusted Rand index: " + bluster_demo_optional(rand.index)) + println(" silhouette widths: " + bluster_demo_doubles(silhouette.widths)) + println(" neighbor purities: " + bluster_demo_doubles(purity.purity)) + println(" first cluster RMSD: " + bluster_demo_optional(rmsd.values[0])) + + let stability = @bio.bluster_bootstrap_kmeans_stability( + coordinates, + kmeans_config, + iterations=8, + seed=109, + ) catch { + _ => abort("bootstrap stability failed") + } + println("\n5. Bootstrap stability") + println( + " first cluster coherence: " + + bluster_demo_optional(stability.ratios[0][0]), + ) + println( + " cluster separation: " + bluster_demo_optional(stability.ratios[0][1]), + ) + + let experiment = @bio.SingleCellExperiment::new( + [ + [8.0, 7.0, 9.0, 8.0, 1.0, 2.0, 1.0, 2.0], + [1.0, 2.0, 1.0, 2.0, 8.0, 7.0, 9.0, 8.0], + ], + ["MarkerA", "MarkerB"], + ["C1", "C2", "C3", "C4", "C5", "C6", "C7", "C8"], + ) + experiment.reduced_dims["PCA"] = coordinates + experiment.metadata["source"] = "bluster_demo" + let clustered = @bio.bluster_cluster_sce( + experiment, + graph_config, + output_column="community", + ) catch { + _ => abort("SCE clustering failed") + } + println("\n6. Immutable SingleCellExperiment integration") + println( + " written labels: " + clustered.experiment.col_data["community"].join(", "), + ) + println( + " original unchanged: " + + (!experiment.col_data.contains("community")).to_string(), + ) + println("\nbluster demo completed.") +} diff --git a/examples/bluster_demo/moon.pkg b/examples/bluster_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/bluster_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/cealign_demo/main.mbt b/examples/cealign_demo/main.mbt new file mode 100644 index 00000000..a43e1551 --- /dev/null +++ b/examples/cealign_demo/main.mbt @@ -0,0 +1,87 @@ +///| +fn main { + println("=== Biopython Bio.PDB.cealign Demo ===") + let (reference, mobile) = @src.cealign_sample_structures() + let original_mobile_coordinate = mobile.get_atoms()[0].coord + + println("\n1. Configure a CE aligner") + let aligner = @src.CEAligner::new() catch { + _ => abort("failed to create CEAligner") + } + let configured = aligner.set_reference(reference) catch { + _ => abort("failed to set CE reference structure") + } + println( + " Window size / maximum gap: " + + configured.window_size.to_string() + + " / " + + configured.max_gap.to_string(), + ) + println( + " Reference guide atoms: " + configured.reference_length().to_string(), + ) + + println("\n2. Search CE fragment paths and superimpose structures") + let result = configured.align(mobile, transform=true, final_optimization=true) catch { + _ => abort("failed to align sample structures") + } + println(" " + result.summary()) + println( + " Aligned guide atoms / fragments: " + + result.path.aligned_length().to_string() + + " / " + + result.path.fragment_count.to_string(), + ) + println( + " RMSD / CE Z-score: " + + result.rmsd.to_string() + + " / " + + result.path.z_score.to_string(), + ) + println( + " Reference / mobile coverage: " + + result.reference_coverage().to_string() + + " / " + + result.mobile_coverage().to_string(), + ) + println( + " Final local optimization applied: " + + result.optimization_applied.to_string(), + ) + + println("\n3. Inspect the selected monotonic path") + let pairs = result.path.aligned_pairs() + let first = pairs[0] + let last = pairs[pairs.length() - 1] + println( + " First reference/mobile index: " + + first.0.to_string() + + " / " + + first.1.to_string(), + ) + println( + " Last reference/mobile index: " + + last.0.to_string() + + " / " + + last.1.to_string(), + ) + + println("\n4. Verify immutable all-atom transformation") + let unchanged_mobile_coordinate = mobile.get_atoms()[0].coord + let transformed_coordinate = result.structure.get_atoms()[0].coord + let reference_coordinate = reference.get_atoms()[0].coord + println( + " Input mobile coordinate unchanged: " + + (unchanged_mobile_coordinate.distance(original_mobile_coordinate) < 1.0e-12).to_string(), + ) + println( + " Transformed first guide distance to reference: " + + transformed_coordinate.distance(reference_coordinate).to_string(), + ) + println( + " Transformed atom count: " + + result.structure.get_atoms().length().to_string(), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/cealign_demo/moon.pkg b/examples/cealign_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/cealign_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/celda_demo/main.mbt b/examples/celda_demo/main.mbt new file mode 100644 index 00000000..ccab72a9 --- /dev/null +++ b/examples/celda_demo/main.mbt @@ -0,0 +1,147 @@ +// Bioconductor celda_CG-inspired joint cell and feature clustering workflow. + +///| +fn celda_demo_config() -> @src.CeldaCGConfig { + @src.CeldaCGConfig::create( + 3, + 4, + max_iterations=20, + stop_iterations=4, + chains=3, + seed=2026, + ) catch { + _ => abort("invalid celda_CG configuration") + } +} + +///| +fn main { + println("=== Bioconductor celda_CG Demo ===") + let (counts, features, cells, samples) = @src.celda_cg_example_data() + + println("\n1. Jointly cluster cells and feature modules") + let result = @src.celda_cg_fit( + counts, + samples, + celda_demo_config(), + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("celda_CG model fitting failed") + } + println(" " + result.summary()) + println( + " chains=" + + result.diagnostics.chain_scores.length().to_string() + + ", selected chain=" + + (result.diagnostics.best_chain + 1).to_string(), + ) + for population in 0.. abort("population query failed") + } + println( + " population " + + (population + 1).to_string() + + ": " + + members.length().to_string() + + " cells", + ) + } + + println("\n2. Inspect feature modules and their strongest markers") + for module_ in 0.. abort("feature module query failed") + } + let mut marker_text = "" + for index in 0.. 0 { + marker_text = marker_text + ", " + } + marker_text = marker_text + + top[index].feature_name + + " (" + + top[index].probability.to_string() + + ")" + } + println(" module " + (module_ + 1).to_string() + ": " + marker_text) + } + + println("\n3. Predict populations for new cells") + let new_counts : Array[Array[Double]] = [] + for feature in 0.. abort("celda_CG prediction failed") + } + for cell in 0.. population " + + (prediction.assignments[cell] + 1).to_string(), + ) + } + + println("\n4. Compare candidate K values by BIC") + let grid_config = @src.CeldaCGConfig::create( + 1, + 1, + max_iterations=8, + stop_iterations=2, + chains=1, + seed=31, + ) catch { + _ => abort("invalid grid-search configuration") + } + let grid = @src.celda_cg_grid_search( + counts, + samples, + [2, 3], + [4], + grid_config, + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("celda_CG grid search failed") + } + for entry in grid.entries { + println( + " K=" + + entry.cell_populations.to_string() + + ", L=" + + entry.feature_modules.to_string() + + ", BIC=" + + entry.score.to_string(), + ) + } + println( + " selected K=" + grid.best_result().config.cell_populations.to_string(), + ) + + println("\n5. Add fitted counts and labels to an SCE copy") + let sce = @src.SingleCellExperiment::new(counts, features, cells) + sce.col_data["donor"] = samples + let output = @src.celda_cg_sce( + sce, + celda_demo_config(), + sample_column="donor", + ) catch { + _ => abort("celda_CG SCE integration failed") + } + println( + " fitted assay rows=" + + @src.sce_get_assay(output.experiment, "celda_fitted").length().to_string(), + ) + println( + " cell labels=" + + @src.sce_get_col_data(output.experiment, "celda_cell_population") + .length() + .to_string() + + ", source unchanged=" + + (@src.sce_get_assay(sce, "celda_fitted").length() == 0).to_string(), + ) + println("\n=== Demo Complete ===") +} diff --git a/examples/celda_demo/moon.pkg b/examples/celda_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/celda_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/cellosaurus_demo/main.mbt b/examples/cellosaurus_demo/main.mbt new file mode 100644 index 00000000..bce3a9e6 --- /dev/null +++ b/examples/cellosaurus_demo/main.mbt @@ -0,0 +1,35 @@ +///| +/// Cellosaurus example: parse, query, and serialize a cell-line record. +/// +/// Run: moon run examples/cellosaurus_demo/main.mbt +fn main { + println("=== Cellosaurus Record Demo ===") + try { + let record = match @bio.cellosaurus_read(@bio.cellosaurus_sample_text()) { + Some(record) => record + None => abort("expected one Cellosaurus record") + } + + println("Record: \{record.repr()}") + println("Accession: \{record.accession}") + println("Human origin: \{record.has_species("Homo sapiens")}") + + println("Synonyms:") + for synonym in record.synonym_list() { + println(" - \{synonym}") + } + + println("Database cross-references:") + for reference in record.cross_references { + println(" - \{reference.database}: \{reference.accession}") + } + let ecacc = record.cross_references_for("ECACC") + println("ECACC accessions: \{ecacc.length()}") + + let serialized = record.to_string() + let reparsed = @bio.cellosaurus_parse(serialized) + println("Round-trip records: \{reparsed.length()}") + } catch { + _ => println("Failed to parse the Cellosaurus sample") + } +} diff --git a/examples/cellosaurus_demo/moon.pkg b/examples/cellosaurus_demo/moon.pkg new file mode 100644 index 00000000..f3a37d11 --- /dev/null +++ b/examples/cellosaurus_demo/moon.pkg @@ -0,0 +1,7 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +options( + "is-main": true, +) diff --git a/examples/codon_align_advanced_demo/main.mbt b/examples/codon_align_advanced_demo/main.mbt new file mode 100644 index 00000000..95f9f58d --- /dev/null +++ b/examples/codon_align_advanced_demo/main.mbt @@ -0,0 +1,186 @@ +// Bio.codonalign-inspired advanced codon alignment and selection pressure +// analysis workflow. +// +// Exercises the full selection-pressure pipeline on synthetic codon +// alignments: +// 1. Z-test for selection (positive / purifying / neutrality). +// 2. Fisher's exact test for neutrality (2×2 contingency table). +// 3. Codon alignment builder from protein alignment + coding sequences. +// 4. Sliding-window dN/dS scan for selection hotspots. +// 5. Benjamini–Hochberg FDR correction of multiple p-values. +// 6. Pairwise Ka/Ks table with Z-test and FDR across a small panel. + +///| +fn caa_demo_round(value : Double, digits : Int) -> Double { + let scale = @math.pow(10.0, digits.to_double()) + (value * scale).round() / scale +} + +///| +fn caa_demo_print_d(label : String, value : Double) -> Unit { + println(" " + label + ": " + caa_demo_round(value, 6).to_string()) +} + +///| +fn caa_demo_print_s(label : String, value : String) -> Unit { + println(" " + label + ": " + value) +} + +///| +fn main { + println("=== Bio.codonalign Advanced Demo ===") + + // ----------------------------------------------------------------------- + // 1. Z-test for selection + // ----------------------------------------------------------------------- + println("\n1. Z-test for selection (Nei–Gojobori approximate variance)") + let pos_seqs = @src.create_positive_selection_alignment() + let pur_seqs = @src.create_purifying_selection_alignment() + let s1_pos = pos_seqs[0] + let s2_pos = pos_seqs[1] + let s1_pur = pur_seqs[0] + let s2_pur = pur_seqs[1] + let pos_result = @src.codon_test_selection(s1_pos, s2_pos, test_type="positive") catch { + CodonAlignAdvancedError(msg) => abort("Z-test positive failed: " + msg) + } + println(" Positive-selection pair:") + caa_demo_print_d("dN", pos_result.dn()) + caa_demo_print_d("dS", pos_result.ds()) + caa_demo_print_d("dN/dS", pos_result.dnds()) + caa_demo_print_d("Z-score", pos_result.z_score()) + caa_demo_print_d("p-value", pos_result.p_value()) + caa_demo_print_s("conclusion", pos_result.conclusion()) + let pur_result = @src.codon_test_selection(s1_pur, s2_pur, test_type="purifying") catch { + CodonAlignAdvancedError(msg) => abort("Z-test purifying failed: " + msg) + } + println(" Purifying-selection pair:") + caa_demo_print_d("dN", pur_result.dn()) + caa_demo_print_d("dS", pur_result.ds()) + caa_demo_print_d("dN/dS", pur_result.dnds()) + caa_demo_print_d("Z-score", pur_result.z_score()) + caa_demo_print_d("p-value", pur_result.p_value()) + caa_demo_print_s("conclusion", pur_result.conclusion()) + + // ----------------------------------------------------------------------- + // 2. Fisher's exact test for neutrality + // ----------------------------------------------------------------------- + println("\n2. Fisher's exact test for neutrality") + let fisher_pos = @src.codon_test_neutrality(s1_pos, s2_pos, test_type="greater") catch { + CodonAlignAdvancedError(msg) => abort("Fisher test failed: " + msg) + } + println(" Positive-selection pair (greater):") + caa_demo_print_d("p-value", fisher_pos.p_value()) + caa_demo_print_d("odds ratio", fisher_pos.odds_ratio()) + println( + " Nd=" + + fisher_pos.n_diff().to_string() + + ", Sd=" + + fisher_pos.s_diff().to_string(), + ) + caa_demo_print_s("conclusion", fisher_pos.conclusion()) + + // ----------------------------------------------------------------------- + // 3. Codon alignment builder + // ----------------------------------------------------------------------- + println("\n3. Codon alignment builder (protein alignment + CDS)") + let (protein_aln, coding_seqs) = @src.create_demo_protein_alignment() + let codon_aln = @src.build_codon_alignment( + protein_aln, + coding_seqs, + names=["geneA", "geneB", "geneC"], + ) catch { + CodonAlignAdvancedError(msg) => abort("Codon alignment builder failed: " + msg) + } + println(" n_codons (columns): " + codon_aln.n_codons().to_string()) + let aln_seqs = codon_aln.sequences() + let aln_names = codon_aln.names() + for i in 0.. abort("Sliding window failed: " + msg) + } + let windows = sw_result.windows() + println(" " + windows.length().to_string() + " windows (size=9, step=6):") + for w in windows { + println( + " codons " + + w.start_codon().to_string() + + "-" + + w.end_codon().to_string() + + " (n=" + + w.n_codons().to_string() + + "): dN=" + + caa_demo_round(w.dn(), 6).to_string() + + ", dS=" + + caa_demo_round(w.ds(), 6).to_string() + + ", dN/dS=" + + caa_demo_round(w.dnds(), 4).to_string(), + ) + } + + // ----------------------------------------------------------------------- + // 5. Benjamini–Hochberg FDR correction + // ----------------------------------------------------------------------- + println("\n5. Benjamini–Hochberg FDR correction") + let raw_pvalues = [0.001, 0.04, 0.03, 0.5, 0.01, 0.2] + let adjusted = @src.codon_align_advanced_bh_fdr(raw_pvalues) + println(" raw → adjusted:") + for i in 0.. abort("Pairwise Ka/Ks table failed: " + msg) + } + println(" " + kaks_table.length().to_string() + " pairwise comparisons:") + for row in kaks_table { + println( + " " + + row.seq1_name() + + " vs " + + row.seq2_name() + + ": dN/dS=" + + caa_demo_round(row.dnds(), 4).to_string() + + ", p=" + + caa_demo_round(row.p_value(), 4).to_string() + + ", FDR=" + + caa_demo_round(row.fdr(), 4).to_string() + + " — " + + row.conclusion(), + ) + } + + println("\n=== Demo complete ===") +} diff --git a/examples/codon_align_advanced_demo/moon.pkg b/examples/codon_align_advanced_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/codon_align_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/decontx_demo/main.mbt b/examples/decontx_demo/main.mbt new file mode 100644 index 00000000..6cc663dd --- /dev/null +++ b/examples/decontx_demo/main.mbt @@ -0,0 +1,109 @@ +///| +fn format_doubles(values : Array[Double]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index].to_string() + index = index + 1 + } + result + "]" +} + +///| +fn main { + println("=== Bioconductor decontX Demo ===") + let (counts, genes, cells, clusters, background) = @src.decontx_example_data() + + println("\n1. Infer cluster-aware ambient RNA contamination") + let result = @src.decontx(counts, genes, cells, clusters) catch { + _ => abort("failed to run cluster-aware decontX") + } + println(" " + result.summary()) + println(" per-cell contamination: " + format_doubles(result.contamination)) + + println("\n2. Inspect cross-population marker correction") + println( + " B marker in T1: observed=" + + counts[2][0].to_string() + + ", native=" + + result.corrected_counts[2][0].to_string() + + ", contaminant=" + + result.contamination_counts[2][0].to_string(), + ) + println( + " T marker in B1: observed=" + + counts[0][4].to_string() + + ", native=" + + result.corrected_counts[0][4].to_string() + + ", contaminant=" + + result.contamination_counts[0][4].to_string(), + ) + + println("\n3. Summarize contamination by cluster") + for estimate in result.cluster_estimates() { + println( + " " + + estimate.cluster + + ": cells=" + + estimate.n_cells.to_string() + + ", mean contamination=" + + estimate.mean_contamination.to_string() + + ", contaminant counts=" + + estimate.contaminant_counts.to_string(), + ) + } + let most_contaminated = result.most_contaminated_cells(2) + for estimate in most_contaminated { + println( + " high-contamination cell " + + estimate.cell_name + + ": " + + estimate.contamination.to_string(), + ) + } + + println("\n4. Use empty droplets as an explicit ambient profile") + let background_result = @src.decontx( + counts, + genes, + cells, + clusters, + background~, + ) catch { + _ => abort("failed to run background-aware decontX") + } + println(" " + background_result.summary()) + println( + " shared ambient profile: " + + format_doubles(background_result.contaminant_profiles[0]), + ) + + println("\n5. Add corrected counts and diagnostics to an SCE copy") + let sce = @src.SingleCellExperiment::new(counts, genes, cells) + sce.col_data["cluster"] = clusters + let output = @src.decontx_sce(sce, background~) catch { + _ => abort("failed to integrate decontX with SingleCellExperiment") + } + println( + " corrected assay dimensions: " + + @src.sce_get_assay(output.experiment, "decontXcounts").length().to_string() + + " genes x " + + output.result.n_cells().to_string() + + " cells", + ) + println( + " contamination metadata entries: " + + @src.sce_get_col_data(output.experiment, "decontX_contamination") + .length() + .to_string(), + ) + println( + " source object unchanged: " + + (@src.sce_get_assay(sce, "decontXcounts").length() == 0).to_string(), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/decontx_demo/moon.pkg b/examples/decontx_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/decontx_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/dirichlet_multinomial_demo/main.mbt b/examples/dirichlet_multinomial_demo/main.mbt new file mode 100644 index 00000000..b0626c54 --- /dev/null +++ b/examples/dirichlet_multinomial_demo/main.mbt @@ -0,0 +1,162 @@ +///| +fn example_counts() -> Array[Array[Int]] { + [ + [40, 3, 2], + [35, 4, 1], + [42, 2, 3], + [38, 5, 2], + [45, 3, 1], + [36, 2, 4], + [2, 40, 3], + [4, 35, 2], + [3, 43, 1], + [5, 38, 2], + [1, 44, 3], + [3, 36, 4], + ] +} + +///| +fn example_groups() -> Array[String] { + ["A", "A", "A", "A", "A", "A", "B", "B", "B", "B", "B", "B"] +} + +///| +fn main { + println("=== Bioconductor DirichletMultinomial Demo ===") + let counts = example_counts() + let groups = example_groups() + let taxa = ["taxon_a", "taxon_b", "taxon_c"] + let config = @src.DmnConfig::create( + components=2, + max_iterations=50, + optimizer_iterations=120, + soft_kmeans_iterations=60, + seed=2026, + ) catch { + DirichletMultinomialError(message) => + abort("failed to create DirichletMultinomial configuration: " + message) + } + + println("\n1. Fit a finite Dirichlet-multinomial mixture") + let fit = @src.dirichlet_multinomial_fit(counts, taxon_names=taxa, config~) catch { + DirichletMultinomialError(message) => + abort("failed to fit Dirichlet-multinomial mixture: " + message) + } + println(" " + fit.summary()) + println(" mixture weights=" + fit.mixture_weights().to_string()) + println(" component proportions=" + fit.component_proportions().to_string()) + println(" assignments=" + fit.assignments().to_string()) + + println("\n2. Compare K=1 and K=2 with the upstream Laplace criterion") + let selection = @src.dirichlet_multinomial_select( + counts, + 2, + criterion=@src.DmnLaplace, + taxon_names=taxa, + config=config.with_components(1), + ) catch { + DirichletMultinomialError(message) => + abort("failed to select Dirichlet-multinomial model: " + message) + } + println(" Laplace scores=" + selection.scores.to_string()) + println( + " selected components=" + selection.best_component_count().to_string(), + ) + + println("\n3. Train a dmngroup-style generative classifier") + let classifier = @src.dirichlet_multinomial_group_fit( + counts, + groups, + components_by_group=[1, 1], + taxon_names=taxa, + config=config.with_components(1), + ) catch { + DirichletMultinomialError(message) => + abort("failed to fit Dirichlet-multinomial classifier: " + message) + } + let novel = [[50, 1, 1], [1, 50, 1]] + let probabilities = classifier.predict(novel) catch { + DirichletMultinomialError(message) => + abort("failed to classify novel samples: " + message) + } + let predictions = classifier.predict_assignments(novel) catch { + DirichletMultinomialError(message) => + abort("failed to assign novel samples: " + message) + } + println(" class probabilities=" + probabilities.to_string()) + println(" predictions=" + predictions.to_string()) + + println("\n4. Run deterministic stratified cross-validation and ROC") + let cross_validation = @src.dirichlet_multinomial_cross_validate( + counts, + groups, + 3, + components_by_group=[1, 1], + taxon_names=taxa, + config=config.with_components(1), + ) catch { + DirichletMultinomialError(message) => + abort("failed to cross-validate classifier: " + message) + } + let truth : Array[Bool] = [] + let scores : Array[Double] = [] + for sample in 0.. + abort("failed to compute ROC: " + message) + } + println( + " accuracy=" + + cross_validation.accuracy.to_string() + + ", AUC=" + + roc.auc.to_string() + + ", best threshold=" + + roc.best_threshold.to_string(), + ) + + println("\n5. Fit a feature x sample SummarizedExperiment assay") + let assays : Map[String, Array[Array[Double]]] = Map([ + ( + "counts", + [ + [40.0, 35.0, 42.0, 2.0, 4.0, 3.0], + [3.0, 4.0, 2.0, 40.0, 35.0, 43.0], + [2.0, 1.0, 3.0, 3.0, 2.0, 1.0], + ], + ), + ]) + let experiment = @src.summarized_experiment(assays, [], [], Map([])) + let se_fit = @src.dirichlet_multinomial_fit_se( + experiment, + "counts", + taxon_names=taxa, + config~, + ) catch { + DirichletMultinomialError(message) => + abort("failed to fit SummarizedExperiment assay: " + message) + } + println( + " transposed dimensions=" + + se_fit.sample_count().to_string() + + " samples x " + + se_fit.taxon_count().to_string() + + " taxa", + ) + + println("\n6. Invalid count matrices produce explicit diagnostics") + let rejected = try { + ignore(@src.dirichlet_multinomial_fit([[3, -1], [2, 4]])) + false + } catch { + DirichletMultinomialError(message) => { + println(" " + message) + true + } + } + println(" malformed counts rejected=" + rejected.to_string()) + println("\n=== Demo Complete ===") +} diff --git a/examples/dirichlet_multinomial_demo/moon.pkg b/examples/dirichlet_multinomial_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/dirichlet_multinomial_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/dreamlet_demo/main.mbt b/examples/dreamlet_demo/main.mbt new file mode 100644 index 00000000..f79b1660 --- /dev/null +++ b/examples/dreamlet_demo/main.mbt @@ -0,0 +1,149 @@ +///| +fn demo_data() -> ( + Array[Array[Double]], + Array[String], + Array[String], + Array[String], + Array[String], + Array[String], + Array[String], +) { + let genes = [ + "response_up", "stable", "T_cell_marker", "donor_signal", "response_down", + ] + let counts : Array[Array[Double]] = [[], [], [], [], []] + let cells : Array[String] = [] + let samples : Array[String] = [] + let clusters : Array[String] = [] + let treatments : Array[String] = [] + let donors : Array[String] = [] + for donor in 0..<4 { + for treatment in 0..<2 { + let sample = "D" + + (donor + 1).to_string() + + (if treatment == 0 { "_C" } else { "_T" }) + for cluster in 0..<2 { + let cluster_name = if cluster == 0 { "T_cell" } else { "Monocyte" } + for replicate in 0..<3 { + cells.push( + sample + "_" + cluster_name + "_" + (replicate + 1).to_string(), + ) + samples.push(sample) + clusters.push(cluster_name) + treatments.push(if treatment == 0 { "Control" } else { "Treated" }) + donors.push("D" + (donor + 1).to_string()) + counts[0].push( + (10 + treatment * 10 + donor + (if replicate == 2 { 1 } else { 0 })).to_double(), + ) + counts[1].push((8 + replicate % 2).to_double()) + counts[2].push( + (if cluster == 0 { 18 + treatment } else { 4 + treatment }).to_double(), + ) + counts[3].push((6 + donor * 2 + replicate).to_double()) + counts[4].push( + (18 - treatment * 8 + cluster + (if replicate == 1 { 1 } else { 0 })).to_double(), + ) + } + } + } + } + (counts, genes, cells, samples, clusters, treatments, donors) +} + +///| +fn main { + println("=== Bioconductor dreamlet Demo ===") + let (counts, genes, cells, samples, clusters, treatments, donors) = demo_data() + + println("\n1. Build a SingleCellExperiment with cohort metadata") + let sce = @src.SingleCellExperiment::new(counts, genes, cells) + sce.col_data["sample_id"] = samples + sce.col_data["cluster_id"] = clusters + sce.col_data["Treatment"] = treatments + sce.col_data["Donor"] = donors + println( + " " + + @src.sce_get_n_genes(sce).to_string() + + " genes x " + + @src.sce_get_n_cells(sce).to_string() + + " cells", + ) + + println("\n2. Aggregate raw counts by sample and cell type") + let pseudobulk = @src.dreamlet_aggregate_sce(sce, "sample_id", "cluster_id", metadata_fields=[ + "Treatment", "Donor", + ]) catch { + _ => abort("failed to aggregate the SingleCellExperiment") + } + println(pseudobulk.summary()) + + println("3. Apply filtering, TMM normalization, and voom weights") + let model = @src.dreamlet_model([ + @src.dreamlet_categorical_effect("Treatment", reference="Control"), + @src.dreamlet_random_effect("Donor"), + ]) catch { + _ => abort("failed to build the repeated-measures model") + } + let config = @src.DreamletProcessConfig::create( + min_cells=3, + min_count=2.0, + min_samples=4, + min_prop=0.5, + min_total_count=10.0, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("failed to build dreamlet processing controls") + } + let processed = @src.dreamlet_process_assays(pseudobulk, model, config~) catch { + _ => abort("failed to process pseudobulk assays") + } + println(processed.summary()) + for assay in processed.assays { + println( + " " + + assay.cluster_id + + ": " + + assay.gene_names.length().to_string() + + " genes, TMM reference=" + + assay.tmm_reference_sample, + ) + } + + println("4. Fit the donor random-intercept treatment contrast") + let result = @src.dreamlet( + processed, + "Treatment:Treated", + contrast_name="Treated-Control", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("failed to run dreamlet differential expression") + } + println(result.summary()) + + println("5. Inspect the study-wide top table") + let top = result.top_table() catch { + _ => abort("failed to create the dreamlet top table") + } + let limit = if top.length() < 6 { top.length() } else { 6 } + for index in 0.. abort("failed to create the DropletUtils example") + } + println( + "Droplet matrix: " + + data.n_features().to_string() + + " features x " + + data.n_barcodes().to_string() + + " barcodes", + ) - println("\n2. Barcode ranking:") - let ranked = @src.barcode_ranking(stats.total_counts) - println(" Top 5 barcodes by UMI count:") - let mut i = 0 - while i < 5 && i < ranked.length() { - println(" " + barcodes[ranked[i].1] + ": " + ranked[i].0.to_string()) - i = i + 1 + let ambient = @src.droplet_utils_ambient_profile(data, lower=15) catch { + _ => abort("failed to estimate the ambient profile") } + let alpha = @src.droplet_utils_estimate_alpha(data, ambient) catch { + _ => abort("failed to estimate Dirichlet-multinomial alpha") + } + println( + "Ambient droplets: " + + ambient.n_empty().to_string() + + ", ambient molecules: " + + ambient.total.to_string() + + ", estimated alpha: " + + alpha.to_string(), + ) - println("\n3. Find knee point:") - let knee = @src.find_knee_point(ranked) - println(" Knee point at index: " + knee.to_string()) + let ranks = @src.droplet_utils_barcode_ranks( + data, + lower=15, + exclude_from=0, + window=0.2, + ) catch { + _ => abort("failed to trace the barcode-rank curve") + } println( - " Estimated cell number: " + - @src.estimate_cell_number(stats.total_counts).to_string(), + "Barcode-rank knee: " + + ranks.knee.to_string() + + ", inflection: " + + ranks.inflection.to_string(), ) - println("\n4. Empty droplet detection (emptyDrops):") - let result = @src.empty_drops(counts, barcodes, 10, 100) - println(" Tested barcodes: " + result.barcode.length().to_string()) + let config = @src.DropletUtilsConfig::create( + lower=15, + iterations=499, + fdr_threshold=0.05, + retain=80.0, + alpha~, + rank_exclude=0, + rank_window=0.2, + seed=31, + ) catch { + _ => abort("failed to create the DropletUtils configuration") + } + let result = @src.droplet_utils_empty_drops(data, config~) catch { + _ => abort("emptyDrops failed") + } + println(result.summary()) - let mut n_cells = 0 - i = 0 - while i < result.is_cell.length() { - if result.is_cell[i] { - n_cells = n_cells + 1 + println("\nCalled barcodes") + for barcode in 0.. value.to_string() + None => "NA" + } + let fdr = match result.fdr[barcode] { + Some(value) => value.to_string() + None => "NA" + } + println( + result.barcodes[barcode] + + "\ttotal=" + + result.totals[barcode].to_string() + + "\tp=" + + p_value + + "\tFDR=" + + fdr + + "\tretained=" + + result.always_retained[barcode].to_string(), + ) } - i = i + 1 } - println(" Detected cells: " + n_cells.to_string()) - println("\n5. Filter cells by emptyDrops results:") - let (filtered_counts, filtered_barcodes) = @src.filter_cells_by_empty_drops( - counts, barcodes, result, + let filtered = result.filter(data) catch { + _ => abort("failed to filter called barcodes") + } + println( + "\nFiltered matrix: " + + filtered.n_features().to_string() + + " features x " + + filtered.n_barcodes().to_string() + + " called barcodes", ) - println(" Filtered barcodes: " + filtered_barcodes.length().to_string()) - println("\nDropletUtils demo completed!") + let assay : Array[Array[Double]] = [] + for feature in 0.. abort("DropletUtils SingleCellExperiment integration failed") + } + println( + "SingleCellExperiment write-back: class=" + + output.experiment.col_data["droplet.class"][40] + + ", ambient=" + + output.experiment.row_data["droplet.ambient"][0] + + ", original_unchanged=" + + (!experiment.col_data.contains("droplet.class")).to_string(), + ) } diff --git a/examples/exonerate_text_demo/main.mbt b/examples/exonerate_text_demo/main.mbt new file mode 100644 index 00000000..9721d5a9 --- /dev/null +++ b/examples/exonerate_text_demo/main.mbt @@ -0,0 +1,91 @@ +///| +fn exonerate_text_demo_parse(text : String) -> @bio.ExonerateTextDocument { + @bio.exonerate_text_parse(text) catch { + ExonerateTextError(message) => { + println("Exonerate text parse failed: " + message) + abort(message) + } + } +} + +///| +fn exonerate_text_demo_range(range : @bio.ExonerateTextRange) -> String { + "[" + range.start.to_string() + ", " + range.end.to_string() + ")" +} + +///| +fn main { + println("=== Biopython Bio.SearchIO.ExonerateIO Offline Demo ===") + let document = exonerate_text_demo_parse(@bio.exonerate_text_example()) + println("\n1. Report metadata and SearchIO hierarchy") + println(" command=" + document.metadata.command_line) + println(" host=" + document.metadata.hostname) + println(" " + document.summary()) + + let query = document.queries[0] + let hit = query.hits[0] + let hsp = hit.hsps[0] + println( + " query=" + + query.id + + ", target=" + + hit.id + + ", model=" + + hsp.model + + ", score=" + + hsp.score.to_string(), + ) + + println("\n2. Spliced fragments and normalized coordinates") + let query_ranges = hsp.query_ranges() + let hit_ranges = hsp.hit_ranges() + for index = 0; index < hsp.fragments.length(); index = index + 1 { + let fragment = hsp.fragments[index] + println( + " fragment " + + (index + 1).to_string() + + ": query=" + + exonerate_text_demo_range(query_ranges[index]) + + ", target=" + + exonerate_text_demo_range(hit_ranges[index]) + + ", phase=" + + fragment.phase.to_string(), + ) + println( + " " + + fragment.query_sequence + + "\n " + + fragment.similarity + + "\n " + + fragment.hit_sequence, + ) + } + + println("\n3. Inter-fragment intron ranges") + println( + " query=" + + exonerate_text_demo_range(hsp.query_inter_ranges()[0]) + + ", target=" + + exonerate_text_demo_range(hsp.hit_inter_ranges()[0]), + ) + + let counts = hsp.counts() + println("\n4. Alignment statistics and position projection") + println( + " columns=" + + counts.alignment_columns.to_string() + + ", identities=" + + counts.identities.to_string() + + ", gaps=" + + counts.gap_columns.to_string() + + ", intervening gaps=" + + counts.introns_or_ner_gaps.to_string(), + ) + println( + " first aligned query base -> " + + hsp.fragments[0].query_position(0).unwrap_or(-1).to_string(), + ) + println( + "\nThe demo parses an existing report artifact; it does not launch Exonerate.", + ) +} diff --git a/examples/exonerate_text_demo/moon.pkg b/examples/exonerate_text_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/exonerate_text_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/flowsom_demo/main.mbt b/examples/flowsom_demo/main.mbt new file mode 100644 index 00000000..2c625bd7 --- /dev/null +++ b/examples/flowsom_demo/main.mbt @@ -0,0 +1,189 @@ +///| +fn flowsom_demo_ints(values : Array[Int]) -> String { + values.map(fn(value) { value.to_string() }).join(", ") +} + +///| +fn flowsom_demo_optional(value : Double?) -> String { + match value { + Some(actual) => actual.to_string() + None => "NA" + } +} + +///| +fn flowsom_demo_bool_count(values : Array[Bool]) -> Int { + let mut total = 0 + for value in values { + if value { + total = total + 1 + } + } + total +} + +///| +fn main { + println("=== Bioconductor FlowSOM 2.21.0 Demo ===") + + // FlowSOM uses cell x marker matrices. + let cells = [ + [8.0, 1.0, 5.0], + [8.2, 1.1, 5.1], + [7.8, 0.9, 4.9], + [8.1, 1.2, 5.0], + [1.0, 8.0, 5.0], + [1.1, 8.2, 5.2], + [0.9, 7.8, 4.8], + [1.2, 8.1, 5.1], + [4.0, 4.0, 9.0], + [4.2, 4.1, 9.2], + [3.8, 3.9, 8.8], + [4.1, 4.2, 9.1], + ] + let config = @bio.FlowSomConfig::create( + xdim=3, + ydim=2, + rlen=5, + mst_runs=2, + alpha_start=0.08, + alpha_end=0.01, + radius_start=2.0, + radius_end=0.0, + importance=[1.0, 1.0, 0.75], + metaclusters=3, + meta_starts=6, + outlier_mad=3.0, + seed=101, + ) catch { + _ => abort("failed to create the FlowSOM configuration") + } + let model = @bio.flowsom_train(cells, config, marker_names=[ + "CD3", "CD19", "CD45", + ]) catch { + _ => abort("FlowSOM training failed") + } + + println("\n1. Topology-aware SOM training") + println(@bio.flowsom_summary(model)) + println(" node counts: " + flowsom_demo_ints(model.node_counts)) + println( + " cell meta-clusters: " + + flowsom_demo_ints(model.cell_meta_clusters.map(fn(value) { value + 1 })), + ) + + println("\n2. Minimum spanning tree") + for edge in model.mst_edges { + println( + " node " + + (edge.from + 1).to_string() + + " -> " + + (edge.to + 1).to_string() + + ", weight=" + + edge.weight.to_string(), + ) + } + + let mut occupied = 0 + while occupied < model.node_counts.length() && + model.node_counts[occupied] == 0 { + occupied = occupied + 1 + } + println("\n3. Node statistics on unweighted marker values") + println(" first occupied node: " + (occupied + 1).to_string()) + println( + " CD3 median fluorescence: " + + flowsom_demo_optional(model.node_medians[occupied][0]), + ) + println( + " CD3 coefficient of variation: " + + flowsom_demo_optional(model.node_cvs[occupied][0]), + ) + let positive = @bio.flowsom_node_positive_percentages(model, [5.0, 5.0, 7.0]) catch { + _ => abort("node positivity calculation failed") + } + println( + " CD3-positive fraction: " + flowsom_demo_optional(positive[occupied][0]), + ) + println( + " training outliers: " + + flowsom_demo_bool_count(model.outliers.per_cell).to_string(), + ) + + let projected = @bio.flowsom_map_new(model, [ + [model.codes[0][0], model.codes[0][1], model.codes[0][2] / 0.75], + [20.0, 20.0, 20.0], + ]) catch { + _ => abort("new data projection failed") + } + println("\n4. New-data projection and MAD outliers") + println( + " mapped nodes: " + + flowsom_demo_ints(projected.mapping.clusters.map(fn(value) { value + 1 })), + ) + println( + " mapped meta-clusters: " + + flowsom_demo_ints(projected.meta_clusters.map(fn(value) { value + 1 })), + ) + println( + " outlier flags: " + + projected.outliers.map(fn(value) { value.to_string() }).join(", "), + ) + + let frame = @bio.FlowFrame::new(cells, [ + @bio.ParameterDescription::new("CD3", 25.0, 0.0, 25.0), + @bio.ParameterDescription::new("CD19", 25.0, 0.0, 25.0), + @bio.ParameterDescription::new("CD45", 25.0, 0.0, 25.0), + ]) + let frame_config = @bio.FlowSomConfig::create( + xdim=3, + ydim=2, + rlen=5, + mst_runs=2, + importance=[1.0, 0.75], + metaclusters=3, + meta_starts=6, + seed=103, + ) catch { + _ => abort("failed to create the FlowFrame configuration") + } + let frame_model = @bio.flowsom_train_flow_frame(frame, frame_config, marker_indices=[ + 0, 2, + ]) catch { + _ => abort("FlowFrame integration failed") + } + println("\n5. flowCore FlowFrame integration") + println(" selected markers: " + frame_model.marker_names.join(", ")) + println( + " mapped events: " + frame_model.mapping.clusters.length().to_string(), + ) + + // SingleCellExperiment assays use marker x cell orientation. + let experiment = @bio.SingleCellExperiment::new( + [ + [8.0, 8.2, 7.8, 8.1, 1.0, 1.1, 0.9, 1.2, 4.0, 4.2, 3.8, 4.1], + [1.0, 1.1, 0.9, 1.2, 8.0, 8.2, 7.8, 8.1, 4.0, 4.1, 3.9, 4.2], + [5.0, 5.1, 4.9, 5.0, 5.0, 5.2, 4.8, 5.1, 9.0, 9.2, 8.8, 9.1], + ], + ["CD3", "CD19", "CD45"], + ["C1", "C2", "C3", "C4", "C5", "C6", "C7", "C8", "C9", "C10", "C11", "C12"], + ) + experiment.metadata["source"] = "flowsom_demo" + let clustered = @bio.flowsom_cluster_sce(experiment, config) catch { + _ => abort("SingleCellExperiment integration failed") + } + println("\n6. Immutable SingleCellExperiment integration") + println( + " SOM labels: " + + clustered.experiment.col_data["FlowSOM.cluster"].join(", "), + ) + println( + " meta-cluster labels: " + + clustered.experiment.col_data["FlowSOM.metacluster"].join(", "), + ) + println( + " original unchanged: " + + (!experiment.col_data.contains("FlowSOM.cluster")).to_string(), + ) + println("\nFlowSOM demo completed.") +} diff --git a/examples/flowsom_demo/moon.pkg b/examples/flowsom_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/flowsom_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/granges_list_demo/main.mbt b/examples/granges_list_demo/main.mbt new file mode 100644 index 00000000..8312fcd9 --- /dev/null +++ b/examples/granges_list_demo/main.mbt @@ -0,0 +1,110 @@ +///| +fn print_element(name : String, ranges : @src.GRanges) -> Unit { + println(" \{name} (\{@src.granges_length(ranges)} ranges)") + for index = 0; index < @src.granges_length(ranges); index = index + 1 { + println( + " \{ranges.seqnames[index]}:\{ranges.starts[index]}-\{ranges.ends[index]}", + ) + } +} + +///| +fn main { + println("=== Bioconductor GRangesList Demo ===") + + let transcripts = @src.granges_list( + [ + @src.granges(["chr1", "chr1"], [(100, 109), (200, 209)], [ + @src.strand_plus(), + @src.strand_plus(), + ]), + @src.granges(["chr1", "chr1"], [(150, 159), (300, 309)], [ + @src.strand_minus(), + @src.strand_minus(), + ]), + @src.granges_single("chr2", 50, 59, @src.strand_star()), + ], + names=["txA", "txB", "txC"], + element_metadata=[ + Map([("gene", "geneA")]), + Map([("gene", "geneB")]), + Map([("gene", "geneC")]), + ], + ) catch { + _ => abort("failed to construct transcript ranges") + } + + println("\n1. Compound genomic features") + println(" " + transcripts.summary()) + let names = transcripts.names() + for index = 0; index < transcripts.length(); index = index + 1 { + match transcripts.get(index) { + Some(ranges) => print_element(names[index], ranges) + None => abort("missing transcript ranges") + } + } + + println("\n2. Unlist and relist") + let flat = transcripts.unlist() + println(" Flattened range count: \{@src.granges_length(flat)}") + let rebuilt = transcripts.relist(flat) catch { + _ => abort("failed to relist flattened ranges") + } + println( + " Rebuilt element sizes: " + + rebuilt.element_lengths().map(fn(size) { size.to_string() }).join(", "), + ) + + println("\n3. Feature-level overlaps use member exons") + let queries = @src.granges(["chr1", "chr1"], [(120, 130), (205, 206)], [ + @src.strand_plus(), + @src.strand_plus(), + ]) + println(" Query 0 lies inside txA bounds but only in its intron.") + for hit in transcripts.find_overlaps(queries) { + println(" \{names[hit.0]} overlaps query \{hit.1}") + } + + println("\n4. Grouped RangedSummarizedExperiment") + let experiment = @src.RangedSummarizedExperiment::new_with_range_groups( + assays=Map([("counts", [[12.0, 18.0], [25.0, 30.0], [8.0, 11.0]])]), + row_range_groups=transcripts, + col_data=[ + Map([("sample", "S1"), ("condition", "control")]), + Map([("sample", "S2"), ("condition", "treated")]), + ], + row_names=names, + metadata=Map([("organism", "human")]), + ) catch { + _ => abort("failed to construct grouped experiment") + } + println(" " + experiment.summary()) + let overlapping = experiment.subset_by_overlaps(queries) catch { + _ => abort("failed to subset grouped experiment") + } + println(" Overlapping assay rows: " + overlapping.row_names().join(", ")) + + let shifted = experiment.shift(10) catch { + _ => abort("failed to shift grouped experiment") + } + match shifted.row_range_groups() { + Some(groups) => + match groups.get(0) { + Some(ranges) => + println( + " Shifted txA exon starts: " + + ranges.starts.map(fn(start) { start.to_string() }).join(", "), + ) + None => abort("missing shifted transcript") + } + None => abort("grouping was not preserved") + } + + println("\n5. Coverage uses all member ranges") + let coverage = experiment.coverage(Map([("chr1", 320), ("chr2", 80)])) + println( + " chr1 coverage at exon/intron positions 100/125: \{coverage["chr1"][99]}/\{coverage["chr1"][124]}", + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/granges_list_demo/moon.pkg b/examples/granges_list_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/granges_list_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/hhr_demo/main.mbt b/examples/hhr_demo/main.mbt new file mode 100644 index 00000000..f9a5eef2 --- /dev/null +++ b/examples/hhr_demo/main.mbt @@ -0,0 +1,66 @@ +///| +fn main { + println("=== Biopython Bio.Align.hhr Demo ===") + + let record = @src.hhr_parse(@src.hhr_sample_text()) catch { + _ => abort("failed to parse HHR sample") + } + + println("\n1. HH-suite run metadata") + println(" " + record.summary()) + println(" Effective sequences: " + record.metadata.neff.to_string()) + println(" Searched HMMs: " + record.metadata.searched_hmms.to_string()) + + println("\n2. Ranked profile hits") + for alignment in record.alignments { + println(" " + alignment.summary()) + println( + " query/target coverage: " + + alignment.query_coverage().to_string() + + " / " + + alignment.target_coverage().to_string(), + ) + } + + println("\n3. Filtering and target lookup") + let confident = record.filter_by_probability(90.0) + println(" Hits with probability >= 90%: " + confident.length().to_string()) + match record.find_target("target_B") { + Some(alignment) => + println( + " target_B aligned query: " + + alignment.query_sequence + + "\n target_B aligned target: " + + alignment.target_sequence, + ) + None => abort("missing target_B") + } + + println("\n4. Coordinate mapping") + match record.best_alignment() { + Some(alignment) => { + println(" Best hit: " + alignment.target_name) + match alignment.query_to_target(3) { + Some(target_position) => + println( + " Query position 3 maps to target position " + + target_position.to_string() + + " (zero-based)", + ) + None => println(" Query position 3 maps to a target gap") + } + println(" Preserved confidence columns: " + alignment.confidence) + } + None => abort("missing best alignment") + } + + println("\n5. Canonical serialization") + let serialized = record.to_string() + let reparsed = @src.hhr_parse(serialized) catch { + _ => abort("failed to reparse canonical HHR") + } + println(" Serialized bytes: " + serialized.length().to_string()) + println(" Round trip preserved record: " + (record == reparsed).to_string()) + + println("\n=== Demo Complete ===") +} diff --git a/examples/hhr_demo/moon.pkg b/examples/hhr_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/hhr_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/infernal_io_demo/main.mbt b/examples/infernal_io_demo/main.mbt new file mode 100644 index 00000000..6914b337 --- /dev/null +++ b/examples/infernal_io_demo/main.mbt @@ -0,0 +1,101 @@ +///| +fn main { + println("=== Biopython Bio.SearchIO.InfernalIO Demo ===") + + println("\n1. Auto-detect and parse Infernal tabular output") + let tabular = @src.infernal_tabular_sample() + let format = @src.infernal_tabular_format(tabular) catch { + _ => abort("failed to detect Infernal tabular format") + } + let format_name = match format { + Format1 => "1" + Format2 => "2" + Format3 => "3" + } + println(" detected format: " + format_name) + let queries = @src.infernal_parse_tabular(tabular) catch { + _ => abort("failed to parse Infernal tabular sample") + } + for query in queries { + println(" " + query.summary()) + for hit in query.hits { + println( + " hit=" + + hit.id + + ", description=" + + hit.description + + ", HSPs=" + + hit.hsps.length().to_string(), + ) + } + } + + println("\n2. Reverse-strand normalized coordinates") + let reverse = queries[0].hits[0].hsps[1].fragments[0] + println( + " Infernal 480..412 -> zero-based half-open [" + + reverse.hit_start.to_string() + + ", " + + reverse.hit_end.to_string() + + "), strand=" + + (if reverse.hit_strand == @src.strand_minus() { "-" } else { "+" }), + ) + + println("\n3. Parse plain text and split local-end alignment") + let text_queries = @src.infernal_parse_text(@src.infernal_text_sample()) catch { + _ => abort("failed to parse Infernal plain-text sample") + } + let text_query = text_queries[0] + let local_hsp = text_query.hits[0].hsps[0] + println(" " + text_query.summary()) + println(" local-end fragments: " + local_hsp.fragments.length().to_string()) + for index in 0.. abort("failed to filter Infernal results") + } + println(" retained HSPs: " + filtered.count_hsps().to_string()) + match queries[0].hits[0].best_hsp() { + Some(best) => + println( + " best HSP: score=" + + best.bitscore.to_string() + + ", E-value=" + + best.evalue.to_string(), + ) + None => abort("expected a best Infernal HSP") + } + + println("\n5. Convert to the generic SearchIO hierarchy") + let generic = text_query.to_searchio() + println( + " QueryResult id=" + + generic.id + + ", hits=" + + generic.hits.length().to_string() + + ", fragments=" + + generic.hits[0].hsps[0].fragments.length().to_string(), + ) + + println( + "\nScope: Infernal 1.0+ non-verbose cmscan/cmsearch tabular formats 1/2/3, plain text, --noali, CM and HMM-only output.", + ) + println("\n=== Demo Complete ===") +} diff --git a/examples/infernal_io_demo/moon.pkg b/examples/infernal_io_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/infernal_io_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/lisaclust_demo/main.mbt b/examples/lisaclust_demo/main.mbt new file mode 100644 index 00000000..b159c24a --- /dev/null +++ b/examples/lisaclust_demo/main.mbt @@ -0,0 +1,142 @@ +// Bioconductor lisaClust-inspired local spatial association and region discovery. + +///| +fn main { + let cells = @src.lisaclust_example_data() + let config = @src.LisaConfig::create( + radii=[2.5, 5.0], + curve_kind=@src.lisa_k_curve(), + window_kind=@src.lisa_rectangle_window(), + bandwidth=4.0, + min_density=0.05, + edge_correction=true, + window_padding=0.1, + edge_samples=96, + n_clusters=2, + n_starts=8, + seed=11, + ) catch { + LisaError(message) => abort("invalid lisaClust configuration: " + message) + } + let result = @src.lisaclust(cells, config~) catch { + LisaError(message) => abort("lisaClust analysis failed: " + message) + } + + println("=== Direct lisaClust analysis ===") + println(result.summary()) + println( + "iterations=" + + result.iterations.to_string() + + ", converged=" + + result.converged.to_string(), + ) + for image_radii in result.curves.effective_radii { + println( + image_radii.image_id + + " effective radii: " + + image_radii.radii.to_string(), + ) + } + + println("\n=== Region summaries ===") + println("region\tsize\tdominant cell type\tmaximum enrichment") + for summary in result.region_summaries { + println( + summary.region + + "\t" + + summary.size.to_string() + + "\t" + + summary.dominant_cell_type + + "\t" + + summary.maximum_enrichment.to_string(), + ) + } + + let top = result.top_enrichments(limit=6, minimum_relative_frequency=1.0) catch { + LisaError(message) => abort("enrichment ranking failed: " + message) + } + println("\n=== Top observed/expected enrichments ===") + println("region\tcell type\tobserved\texpected\tratio") + for entry in top { + println( + entry.region + + "\t" + + entry.cell_type + + "\t" + + entry.observed.to_string() + + "\t" + + entry.expected.to_string() + + "\t" + + entry.relative_frequency.to_string(), + ) + } + + let l_config = @src.LisaConfig::create( + radii=[2.5, 5.0], + curve_kind=@src.lisa_l_curve(), + window_kind=@src.lisa_rectangle_window(), + bandwidth=4.0, + min_density=0.05, + edge_correction=true, + window_padding=0.1, + edge_samples=96, + ) catch { + LisaError(message) => abort("invalid local-L configuration: " + message) + } + let l_curves = @src.lisa_curves(cells, config=l_config) catch { + LisaError(message) => abort("local-L curve computation failed: " + message) + } + println("\n=== Centered local-L curve for the first cell ===") + println("cell=" + l_curves.cell_ids[0]) + for feature in 0.. + abort("lisaClust SpatialExperiment integration failed: " + message) + } + println("\n=== SpatialExperiment integration ===") + println( + "cells=" + + integrated.experiment.metadata["lisaclust_cells"] + + ", features=" + + integrated.experiment.metadata["lisaclust_features"] + + ", regions=" + + integrated.experiment.metadata["lisaclust_regions"], + ) + println( + "first assigned tissueRegion=" + + integrated.experiment.col_data[0]["tissueRegion"] + + ", original unchanged=" + + (!experiment.col_data[0].contains("tissueRegion")).to_string(), + ) +} diff --git a/examples/lisaclust_demo/moon.pkg b/examples/lisaclust_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/lisaclust_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/mast_demo/main.mbt b/examples/mast_demo/main.mbt index f2af296a..a5951ae0 100644 --- a/examples/mast_demo/main.mbt +++ b/examples/mast_demo/main.mbt @@ -1,145 +1,115 @@ ///| -fn main { - println("=== MAST Demo ===") +fn mast_demo_option(value : Double?) -> String { + match value { + Some(number) => number.to_string() + None => "NA" + } +} - let n_genes = 10 - let n_group1 = 6 - let n_group2 = 6 - let n_cells = n_group1 + n_group2 +///| +fn main { + println("=== Bioconductor MAST 1.39.0 Advanced Demo ===") - let expression : Array[Array[Double]] = Array::make( - n_cells, - Array::make(n_genes, 0.0), + let (data, tested) = @src.mast_advanced_example() catch { + _ => abort("failed to construct the MAST example") + } + println("\n1. Feature x cell model and treatment-coded design") + println( + " genes=" + + data.n_genes.to_string() + + ", cells=" + + data.n_cells.to_string() + + ", coefficients=" + + data.coefficient_names.join(", "), + ) + println( + " tested coefficient=" + + data.coefficient_names[tested[0]] + + ", first-cell CDR=" + + data.cdr[0].to_string(), ) - let gene_names : Array[String] = Array::make(n_genes, "") - let cell_names : Array[String] = Array::make(n_cells, "") - let groups : Array[Int] = Array::make(n_cells, 0) - let mut g = 0 - while g < n_genes { - gene_names[g] = "Gene" + g.to_string() - g = g + 1 + let config = @src.MastAdvancedConfig::create( + empirical_bayes=true, + ebayes_use_full_model=false, + fdr_threshold=0.05, + ) catch { + _ => abort("failed to construct the MAST configuration") } - - let mut i = 0 - while i < n_cells { - cell_names[i] = "Cell" + i.to_string() - groups[i] = if i < n_group1 { 0 } else { 1 } - - let row : Array[Double] = Array::make(n_genes, 0.0) - let mut j = 0 - while j < n_genes { - let base_expr = if j < 5 { - if i < n_group1 { - 10.0 - } else { - 50.0 - } - } else if i < n_group1 { - 30.0 - } else { - 25.0 - } - let noise = (i * 7 + j * 13).to_double() % 5.0 - let zero_prob = if j < 3 { - if i < n_group1 { - 0.4 - } else { - 0.1 - } - } else { - 0.2 - } - let is_zero = ((i * 31 + j * 17) % 100).to_double() / 100.0 < zero_prob - row[j] = if is_zero { 0.0 } else { base_expr + noise } - j = j + 1 - } - expression[i] = row - i = i + 1 + let result = @src.mast_advanced_zlm(data, tested, config~) catch { + _ => abort("advanced MAST zlm fit failed") } + println("\n2. Bayesian logistic + positive Gaussian hurdle fit") + println(" " + result.summary()) + println( + " eBayes prior variance=" + + result.prior_variance.to_string() + + ", prior df=" + + result.prior_df.to_string(), + ) - println("\n1. Create MAST Data:") - let data = @src.MastData::new(expression, gene_names, cell_names, groups) - println(" Number of cells: " + data.n_cells.to_string()) - println(" Number of genes: " + data.n_genes.to_string()) - println(" Number of groups: " + data.n_groups.to_string()) - - println("\n2. Detection rate (cngeneson) per cell:") - let mut c = 0 - while c < 5 && c < data.n_cells { + println("\n3. Component LRT, hurdle LRT, FDR and marginal logFC") + for gene in 0.. abort("advanced MAST SingleCellExperiment fit failed") } + println( + " original has mast39.hurdleFdr: " + + experiment.row_data.contains("mast39.hurdleFdr").to_string(), + ) + println( + " enriched row columns include hurdle FDR: " + + output.experiment.row_data.contains("mast39.hurdleFdr").to_string(), + ) + println( + " contrast metadata: " + + output.experiment.metadata["mast39.contrast"] + + ", tested genes=" + + output.experiment.metadata["mast39.tested"], + ) - println("\nMAST demo completed!") + println("\nMAST advanced demo completed.") } diff --git a/examples/milo_demo/main.mbt b/examples/milo_demo/main.mbt new file mode 100644 index 00000000..4ab931c0 --- /dev/null +++ b/examples/milo_demo/main.mbt @@ -0,0 +1,148 @@ +///| +fn format_ints(values : Array[Int]) -> String { + let mut output = "[" + for index in 0.. 0 { + output = output + ", " + } + output = output + values[index].to_string() + } + output + "]" +} + +///| +fn format_doubles(values : Array[Double]) -> String { + let mut output = "[" + for index in 0.. 0 { + output = output + ", " + } + output = output + values[index].to_string() + } + output + "]" +} + +///| +fn main { + println("=== Bioconductor miloR Demo ===") + let data = @src.milo_sample_data() + + println("\n1. Exact KNN graph from reduced dimensions") + let graph = @src.milo_build_graph( + data.coordinates, + cell_names=data.cell_names, + k=8, + dimensions=2, + ) catch { + _ => abort("failed to build the Milo KNN graph") + } + println(" " + graph.summary()) + println(" first directed KNN: " + format_ints(graph.knn_indices[0])) + println( + " first undirected neighbors: " + format_ints(graph.graph_neighbors[0]), + ) + + println("\n2. Refined, overlapping neighborhoods") + let neighborhoods = graph.make_neighborhoods( + proportion=0.5, + refined=true, + seed=19, + ) catch { + _ => abort("failed to sample Milo neighborhoods") + } + println( + " refined representatives: " + + neighborhoods.neighborhood_indices.length().to_string(), + ) + println( + " first neighborhood: " + format_ints(neighborhoods.neighborhoods[0]), + ) + + println("\n3. Neighborhood-by-sample counts and mean expression") + let counted = neighborhoods.count_cells(data.sample_ids) catch { + _ => abort("failed to count cells by sample") + } + let feature_x : Array[Double] = [] + let feature_y : Array[Double] = [] + for coordinate in data.coordinates { + feature_x.push(coordinate[0]) + feature_y.push(coordinate[1]) + } + let means = counted.neighborhood_expression([feature_x, feature_y]) catch { + _ => abort("failed to aggregate neighborhood expression") + } + println(" sample order: " + counted.sample_names.to_string()) + println(" first count row: " + format_ints(counted.neighborhood_counts[0])) + println(" mean x by neighborhood: " + format_doubles(means[0])) + + println("\n4. Fixed-effect negative-binomial differential abundance") + let design = @src.milo_binary_design(data.sample_conditions, reference="A") catch { + _ => abort("failed to construct the Milo design matrix") + } + let results = counted.test_neighborhoods(design.matrix, 1, cell_sizes=[ + 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, + ]) catch { + _ => abort("failed to test Milo neighborhoods") + } + println(" coefficient: " + design.coefficient_name) + let mut left_depleted = 0 + let mut right_enriched = 0 + for result in results { + let x = counted.coordinates[result.index_cell][0] + if x < 2.5 && result.log_fc < 0.0 { + left_depleted = left_depleted + 1 + } else if x >= 2.5 && result.log_fc > 0.0 { + right_enriched = right_enriched + 1 + } + println( + " nhood " + + (result.neighborhood + 1).to_string() + + ": log2FC=" + + result.log_fc.to_string() + + ", FDR=" + + result.fdr.to_string() + + ", spatial FDR=" + + result.spatial_fdr.to_string(), + ) + } + println( + " state directions: left depleted=" + + left_depleted.to_string() + + ", right enriched=" + + right_enriched.to_string(), + ) + + println("\n5. Alternative graph-aware FDR and SCE integration") + let p_values : Array[Double] = [] + for result in results { + p_values.push(result.p_value) + } + let overlap_fdr = @src.milo_graph_spatial_fdr( + counted, + p_values, + weighting=@src.milo_graph_overlap_weighting(), + ) catch { + _ => abort("failed to compute graph-overlap FDR") + } + println(" graph-overlap FDR: " + format_doubles(overlap_fdr)) + + let sce = @src.SingleCellExperiment::new( + [Array::make(data.cell_names.length(), 1.0)], + ["marker"], + data.cell_names, + ) + let with_pca = @src.sce_set_reduced_dim(sce, "PCA", data.coordinates) + let sce_graph = @src.milo_from_single_cell_experiment( + with_pca, + reduced_dim="PCA", + k=8, + ) catch { + _ => abort("failed to construct Milo from SingleCellExperiment") + } + println(" SCE-backed graph: " + sce_graph.summary()) + + println( + "\nScope: portable fixed-effect NB-GLM; GLMM, edgeR TMM/RLE/QL backends and plotting are not included.", + ) + println("\n=== Demo Complete ===") +} diff --git a/examples/milo_demo/moon.pkg b/examples/milo_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/milo_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/msf_demo/main.mbt b/examples/msf_demo/main.mbt new file mode 100644 index 00000000..e89bee88 --- /dev/null +++ b/examples/msf_demo/main.mbt @@ -0,0 +1,113 @@ +///| +fn integer_array_text(values : Array[Int]) -> String { + let output = StringBuilder::new() + output.write_char('[') + for index = 0; index < values.length(); index = index + 1 { + if index > 0 { + output.write_string(", ") + } + output.write_string(values[index].to_string()) + } + output.write_char(']') + output.to_string() +} + +///| +fn main { + println("=== Biopython Bio.Align.msf Demo ===") + + let alignment = @src.msf_parse(@src.msf_example_text()) catch { + MsfError(message) => abort("MSF parsing failed: " + message) + } + + println("\n1. Parse interleaved MSF metadata and rows") + println(" " + alignment.summary()) + println(" title: " + alignment.metadata.title) + for sequence in alignment.sequences { + println( + " " + + sequence.id + + ": residues=" + + sequence.length.to_string() + + ", weight=" + + sequence.weight.to_string() + + ", checksum=" + + sequence.checksum.to_string(), + ) + } + + println("\n2. Build a Biopython-style coordinate path") + let path = alignment.coordinate_path() + for row = 0; row < alignment.num_sequences(); row = row + 1 { + println( + " " + alignment.sequences[row].id + ": " + integer_array_text(path[row]), + ) + } + println( + " reference residue 4 -> query_one residue " + + alignment.map_position(0, 1, 4).unwrap_or(-1).to_string(), + ) + println( + " reference residue 3 -> query_one residue " + + alignment.map_position(0, 1, 3).unwrap_or(-1).to_string() + + " (-1 denotes a gap)", + ) + + println("\n3. Calculate pair statistics and consensus") + let counts = alignment.pair_counts(0, 2) catch { + MsfError(message) => abort("pair counting failed: " + message) + } + println( + " aligned=" + + counts.aligned.to_string() + + ", identities=" + + counts.identities.to_string() + + ", mismatches=" + + counts.mismatches.to_string() + + ", gap opens=" + + counts.gap_opens.to_string(), + ) + let consensus = alignment.consensus(minimum_fraction=0.67) catch { + MsfError(message) => abort("consensus failed: " + message) + } + println(" consensus: " + consensus) + println(" occupancy at column 1: " + alignment.occupancy()[1].to_string()) + + println("\n4. Write canonical MSF and parse it again") + let canonical = @src.msf_write( + alignment, + block_width=12, + group_width=4, + gap_character="~", + ) catch { + MsfError(message) => abort("MSF serialization failed: " + message) + } + let round_trip = @src.msf_parse(canonical) catch { + MsfError(message) => abort("round-trip parsing failed: " + message) + } + println(" output bytes: " + canonical.length().to_string()) + println(" checksums valid: " + round_trip.checksums_valid().to_string()) + println( + " aligned rows preserved: " + + (round_trip.sequences[1].aligned_sequence == + alignment.sequences[1].aligned_sequence).to_string(), + ) + + println("\n5. Preserve third-party declared-width mismatches") + let mismatch_text = "!!AA_MULTIPLE_ALIGNMENT\n" + + "MSF: 2 Type: P Check: 0 ..\n" + + "Name: full Len: 4 Check: 0 Weight: 1.0\n" + + "Name: short Len: 2 Check: 0 Weight: 1.0\n" + + "//\n\nfull ACDE\nshort AC\n" + let mismatch = @src.msf_parse(mismatch_text) catch { + MsfError(message) => abort("length-mismatch parsing failed: " + message) + } + println( + " declared=" + + mismatch.metadata.declared_length.to_string() + + ", actual=" + + mismatch.alignment_length().to_string() + + ", mismatch=" + + mismatch.length_mismatch().to_string(), + ) +} diff --git a/examples/msf_demo/moon.pkg b/examples/msf_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/msf_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/muscat_advanced_demo/main.mbt b/examples/muscat_advanced_demo/main.mbt new file mode 100644 index 00000000..0eaca8c8 --- /dev/null +++ b/examples/muscat_advanced_demo/main.mbt @@ -0,0 +1,169 @@ +///| +fn muscat_demo_config() -> @bio.MuscatAdvancedConfig { + @bio.MuscatAdvancedConfig::create( + min_cells=3, + min_count=1.0, + min_samples=2, + max_iterations=100, + tolerance=1.0e-8, + dispersion_prior_df=10.0, + ridge=1.0e-6, + fdr_threshold=0.1, + detection_filter=0.9, + ) catch { + _ => abort("failed to create muscat advanced configuration") + } +} + +///| +fn muscat_demo_result( + result : @bio.MuscatAdvancedResults, + genes : Array[String], + clusters : Array[String], + contrast : String, +) -> Unit { + println("\n" + result.summary()) + for cluster in clusters { + println(" cluster " + cluster) + for gene in genes { + match result.get(gene, cluster, contrast) { + Some(value) => + println( + " " + + gene + + ": log2FC=" + + value.log_fc.to_string() + + ", Wald=" + + value.statistic.to_string() + + ", p=" + + value.p_value.to_string() + + ", local FDR=" + + value.local_fdr.to_string() + + ", global FDR=" + + value.global_fdr.to_string(), + ) + None => () + } + } + } +} + +///| +fn muscat_demo_sce(data : @bio.MuscatAdvancedData) -> @bio.SingleCellExperiment { + let experiment = @bio.SingleCellExperiment::new( + data.counts, + data.gene_names, + data.cell_names, + ) + experiment.col_data["sample"] = data.sample_ids.copy() + experiment.col_data["cluster"] = data.cluster_ids.copy() + experiment.col_data["group"] = data.group_ids.copy() + experiment.row_data["symbol"] = data.gene_names.copy() + experiment.metadata["source"] = "muscat advanced demo" + experiment +} + +///| +fn main { + println("=== Bioconductor muscat 1.27.4 Advanced Demo ===") + let data = @bio.muscat_advanced_example() catch { + _ => abort("failed to construct muscat example data") + } + let config = muscat_demo_config() + + println("\n1. Strict gene x cell data contract") + println( + " " + + data.gene_names.length().to_string() + + " genes, " + + data.cell_names.length().to_string() + + " cells, " + + data.sample_names.length().to_string() + + " samples, " + + data.cluster_names.length().to_string() + + " clusters", + ) + + let sums = @bio.muscat_aggregate_advanced( + data, + aggregation=@bio.muscat_sum_counts(), + ) catch { + _ => abort("sum-count aggregation failed") + } + let detections = @bio.muscat_aggregate_advanced( + data, + aggregation=@bio.muscat_number_detected(), + ) catch { + _ => abort("number-detected aggregation failed") + } + println("\n2. Cluster-sample pseudobulk aggregation") + for cluster in 0.. abort("group design construction failed") + } + let contrasts = @bio.muscat_default_contrasts_advanced(design) catch { + _ => abort("default contrast construction failed") + } + let contrast = contrasts[0].name + println("\n3. Replicated design and contrast") + println(" coefficients: " + design.coefficient_names.join(", ")) + println(" contrast: " + contrast) + + let ds = @bio.muscat_pbds_advanced(sums, design, contrasts, config~) catch { + _ => abort("differential-state model failed") + } + let dd = @bio.muscat_pbdd_advanced(detections, design, contrasts, config~) catch { + _ => abort("differential-detection model failed") + } + muscat_demo_result(ds, data.gene_names, data.cluster_names, contrast) + muscat_demo_result(dd, data.gene_names, data.cluster_names, contrast) + + let stagewise = @bio.muscat_stagewise_ds_dd(ds, dd, alpha=0.1) catch { + _ => abort("stagewise DS/DD testing failed") + } + println("\n4. Harmonic-mean screening and two-stage confirmation") + for result in stagewise.results { + if result.classification != "none" { + println( + " " + + result.cluster_name + + "/" + + result.gene_name + + ": " + + result.classification + + ", screen FDR=" + + result.screen_fdr.to_string(), + ) + } + } + + let sce_output = @bio.muscat_advanced_sce( + muscat_demo_sce(data), + "sample", + "cluster", + "group", + reference="ctrl", + output_prefix="muscat", + config~, + ) catch { + _ => abort("SingleCellExperiment integration failed") + } + println("\n5. Immutable SingleCellExperiment write-back") + println( + " metadata version: " + sce_output.experiment.metadata["muscat.version"], + ) + println( + " row annotations: muscat.A.dsLogFC, muscat.A.dsFdr, " + + "muscat.A.ddLogFC, muscat.A.ddFdr, muscat.A.stageClass", + ) +} diff --git a/examples/muscat_advanced_demo/moon.pkg b/examples/muscat_advanced_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/muscat_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/nnsvg_demo/main.mbt b/examples/nnsvg_demo/main.mbt new file mode 100644 index 00000000..9d33a1be --- /dev/null +++ b/examples/nnsvg_demo/main.mbt @@ -0,0 +1,99 @@ +// Bioconductor nnSVG-inspired spatially variable gene analysis. + +///| +fn main { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + let config = @src.NnsvgConfig::create( + n_neighbors=4, + length_scale_grid=8, + proportion_grid=8, + refinement_steps=2, + fdr_threshold=0.1, + ) catch { + NnsvgError(message) => abort("invalid nnSVG configuration: " + message) + } + + let result = @src.nnsvg_with_config( + expression, + coordinates, + config, + gene_names~, + ) catch { + NnsvgError(message) => abort("nnSVG fit failed: " + message) + } + + println("=== Direct nnSVG analysis ===") + println(result.summary()) + println("rank\tgene\tLR\tp-value\tFDR\tprop_sv\tlength_scale") + for gene in result.top(result.n_genes()) { + println( + gene.rank.to_string() + + "\t" + + gene.gene_name + + "\t" + + gene.likelihood_ratio.to_string() + + "\t" + + gene.p_value.to_string() + + "\t" + + gene.adjusted_p_value.to_string() + + "\t" + + gene.proportion_spatial_variance.to_string() + + "\t" + + gene.length_scale.to_string(), + ) + } + + let counts = [ + [0.0, 0.0, 4.0, 5.0, 0.0, 0.0], + [8.0, 9.0, 7.0, 8.0, 9.0, 8.0], + [5.0, 5.0, 5.0, 5.0, 5.0, 5.0], + ] + let filter = @src.nnsvg_filter_genes( + counts, + ["low_expression", "spatial_candidate", "MT-ND1"], + minimum_count=3.0, + minimum_spot_percentage=50.0, + ) catch { + NnsvgError(message) => abort("nnSVG filtering failed: " + message) + } + println("\n=== Upstream-compatible gene filtering ===") + println( + "kept=" + + filter.kept_indices.length().to_string() + + ", low-expression=" + + filter.removed_low_expression.length().to_string() + + ", mitochondrial=" + + filter.removed_mitochondrial.length().to_string(), + ) + + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "logcounts", expression)) + for gene_name in gene_names { + ignore(@src.se_add_row(experiment, Map([("gene_name", gene_name)]))) + } + for spot in 0.. + abort("nnSVG SpatialExperiment integration failed: " + message) + } + println("\n=== SpatialExperiment integration ===") + println( + "rowData fields written for " + + integrated.experiment.row_data.length().to_string() + + " genes; first-gene padj=" + + integrated.experiment.row_data[0]["nnsvg_padj"], + ) +} diff --git a/examples/nnsvg_demo/moon.pkg b/examples/nnsvg_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/nnsvg_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/paml_baseml_demo/main.mbt b/examples/paml_baseml_demo/main.mbt new file mode 100644 index 00000000..31452251 --- /dev/null +++ b/examples/paml_baseml_demo/main.mbt @@ -0,0 +1,127 @@ +///| +fn baseml_demo_control() -> @bio.BasemlControl { + @bio.baseml_parse_control(@bio.baseml_example_control_text()) catch { + BasemlError(message) => abort("failed to parse BASEML control: " + message) + } +} + +///| +fn baseml_demo_results() -> @bio.BasemlResults { + @bio.baseml_parse_results(@bio.baseml_example_results_text()) catch { + BasemlError(message) => abort("failed to parse BASEML results: " + message) + } +} + +///| +fn baseml_demo_null_model() -> @bio.BasemlResults { + @bio.baseml_parse_results( + "BASEML (in paml version 4.7, January 2013) alignment.phylip JC69 dGamma\n" + + "lnL(ntime: 7 np: 8): -330.000000 +0.000000\n" + + "0.01 0.01 0.01 0.01 0.01 0.01 0.01 1.0\n" + + "tree length = 0.018\n" + + "((A:0.004,B:0.004):0.002,C:0.008);\n", + ) catch { + BasemlError(message) => abort("failed to parse null model: " + message) + } +} + +///| +fn main { + println("=== Biopython Bio.Phylo.PAML.baseml Offline Demo ===") + let control = baseml_demo_control() + let model = match control.model_number() { + Some(value) => value + None => abort("control has no model") + } + println("\n1. Strict BASEML control parsing") + println(" sequence file: " + control.sequence_file()) + println(" tree file: " + control.tree_file()) + println( + " model: " + + model.to_string() + + " (" + + (@bio.baseml_model_name(model) catch { + BasemlError(message) => abort(message) + }) + + ")", + ) + let canonical = @bio.baseml_write_control(control) + let reparsed = @bio.baseml_parse_control(canonical) catch { + BasemlError(message) => abort("control round-trip failed: " + message) + } + println(" canonical round-trip: " + (reparsed == control).to_string()) + + let result = baseml_demo_results() + println("\n2. PAML result, parameters, and standard errors") + println( + " PAML=" + + result.version() + + ", model=" + + result.model_description() + + ", lnL=" + + result.ln_likelihood().to_string(), + ) + println( + " parameters=" + + result.parameter_count().to_string() + + ", SE values=" + + result.standard_errors().length().to_string() + + ", tree length=" + + result.tree_length().to_string(), + ) + + println("\n3. REV rate matrix and discrete gamma") + match result.q_matrix() { + Some(matrix) => { + let average = match matrix.average_ts_tv() { + Some(value) => value.to_string() + None => "not reported" + } + println( + " Q dimensions=" + + matrix.rows().length().to_string() + + "x" + + matrix.rows()[0].length().to_string() + + ", average Ts/Tv=" + + average, + ) + } + None => abort("result has no Q matrix") + } + let alpha_text = match result.alpha() { + Some(value) => value.to_string() + None => "not reported" + } + println( + " rate categories=" + + result.rates().length().to_string() + + ", alpha=" + + alpha_text, + ) + + let comparison = @bio.baseml_likelihood_ratio( + baseml_demo_null_model(), + result, + ) catch { + BasemlError(message) => abort("likelihood-ratio test failed: " + message) + } + let bic = result.bic(222) catch { + BasemlError(message) => abort("BIC calculation failed: " + message) + } + println("\n4. Information criteria and nested-model LRT") + println(" REV AIC=" + result.aic().to_string() + ", BIC=" + bic.to_string()) + println( + " statistic=" + + comparison.statistic().to_string() + + ", df=" + + comparison.degrees_of_freedom().to_string() + + ", p=" + + comparison.p_value().to_string() + + ", significant=" + + comparison.significant().to_string(), + ) + + println( + "\nThe demo parses existing BASEML artifacts; it does not launch PAML.", + ) +} diff --git a/examples/paml_baseml_demo/moon.pkg b/examples/paml_baseml_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/paml_baseml_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/paml_codeml_demo/main.mbt b/examples/paml_codeml_demo/main.mbt new file mode 100644 index 00000000..3343b774 --- /dev/null +++ b/examples/paml_codeml_demo/main.mbt @@ -0,0 +1,114 @@ +///| +fn codeml_demo_control() -> @bio.CodemlControl { + @bio.codeml_parse_control(@bio.codeml_example_control_text()) catch { + CodemlError(message) => abort("failed to parse CODEML control: " + message) + } +} + +///| +fn codeml_demo_results() -> @bio.CodemlResults { + @bio.codeml_parse_results(@bio.codeml_example_results_text()) catch { + CodemlError(message) => abort("failed to parse CODEML results: " + message) + } +} + +///| +fn codeml_demo_ints(values : Array[Int]) -> String { + "[" + values.map(fn(value) { value.to_string() }).join(", ") + "]" +} + +///| +fn main { + println("=== Biopython Bio.Phylo.PAML.codeml Offline Demo ===") + let control = codeml_demo_control() + println("\n1. Strict CODEML control parsing") + println(" sequence file: " + control.sequence_file()) + println(" tree file: " + control.tree_file()) + println(" NSsites: " + codeml_demo_ints(control.ns_sites())) + let written = @bio.codeml_write_control(control) + let reparsed = @bio.codeml_parse_control(written) catch { + CodemlError(message) => abort("control round-trip failed: " + message) + } + println( + " canonical round-trip: " + + (reparsed.ns_sites() == control.ns_sites()).to_string(), + ) + + let results = codeml_demo_results() + let models = results.models() + let null_model = models[0] + let selection_model = models[1] + println("\n2. CODONML metadata and multiple NSsites models") + println( + " program=" + + results.program() + + ", PAML=" + + results.version() + + ", sequences=" + + results.sequence_count().to_string() + + ", sites=" + + results.site_count().to_string(), + ) + println( + " M0 lnL=" + + null_model.ln_likelihood().to_string() + + ", parameters=" + + null_model.parameter_count().to_string(), + ) + println( + " M2 lnL=" + + selection_model.ln_likelihood().to_string() + + ", parameters=" + + selection_model.parameter_count().to_string(), + ) + + let branches = null_model.branches() + let site_classes = selection_model.site_classes() + let positive_sites = selection_model.positive_sites() + println("\n3. Branch, site-class, and BEB estimates") + println( + " branches=" + + branches.length().to_string() + + ", first branch=" + + branches[0].branch(), + ) + println( + " site classes=" + + site_classes.length().to_string() + + ", selected class proportion=" + + site_classes[2].proportion().to_string(), + ) + println( + " BEB site " + + positive_sites[0].position().to_string() + + positive_sites[0].amino_acid() + + ", probability=" + + positive_sites[0].probability().to_string() + + positive_sites[0].significance(), + ) + + let bic = selection_model.bic(results.site_count()) catch { + CodemlError(message) => abort("BIC calculation failed: " + message) + } + let comparison = @bio.codeml_likelihood_ratio(null_model, selection_model) catch { + CodemlError(message) => abort("likelihood-ratio test failed: " + message) + } + println("\n4. Information criteria and nested-model LRT") + println( + " M2 AIC=" + selection_model.aic().to_string() + ", BIC=" + bic.to_string(), + ) + println( + " statistic=" + + comparison.statistic().to_string() + + ", df=" + + comparison.degrees_of_freedom().to_string() + + ", p=" + + comparison.p_value().to_string() + + ", significant=" + + comparison.significant().to_string(), + ) + + println( + "\nThe demo parses existing CODEML artifacts; it does not launch PAML.", + ) +} diff --git a/examples/paml_codeml_demo/moon.pkg b/examples/paml_codeml_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/paml_codeml_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/paml_yn00_demo/main.mbt b/examples/paml_yn00_demo/main.mbt new file mode 100644 index 00000000..73da807c --- /dev/null +++ b/examples/paml_yn00_demo/main.mbt @@ -0,0 +1,122 @@ +///| +fn optional_value(value : Double?) -> String { + match value { + Some(number) => number.to_string() + None => "undefined" + } +} + +///| +fn yn00_demo_control(text : String) -> @bio.Yn00Control { + @bio.yn00_parse_control(text) catch { + Yn00Error(message) => abort("failed to parse YN00 control: " + message) + } +} + +///| +fn yn00_demo_results() -> @bio.Yn00Results { + @bio.yn00_parse_results(@bio.yn00_example_results_text()) catch { + Yn00Error(message) => abort("failed to parse YN00 results: " + message) + } +} + +///| +fn yn00_demo_matrix(results : @bio.Yn00Results) -> @bio.Yn00DistanceMatrix { + results.matrix("YN00", "dS") catch { + Yn00Error(message) => abort("failed to build YN00 matrix: " + message) + } +} + +///| +fn yn00_demo_mean(results : @bio.Yn00Results) -> Double { + results.mean("YN00", "dS") catch { + Yn00Error(message) => abort("failed to calculate YN00 mean: " + message) + } +} + +///| +fn main { + println("=== Biopython Bio.Phylo.PAML.yn00 Offline Demo ===") + let control = yn00_demo_control(@bio.yn00_example_control_text()) + let canonical = @bio.yn00_write_control(control) + println("\n1. Strict YN00 control parsing") + println(" sequence file: " + control.sequence_file()) + println(" output file: " + control.output_file()) + println(" genetic code: " + control.option("icode").unwrap_or("default")) + println( + " canonical round-trip: " + + (@bio.yn00_write_control(yn00_demo_control(canonical)) == canonical).to_string(), + ) + + let results = yn00_demo_results() + println("\n2. Complete pairwise result") + println( + " alignment=" + + results.alignment_file() + + ", sequences=" + + results.sequence_count().to_string() + + ", codons=" + + results.codon_count().to_string() + + ", pairs=" + + results.pairs().length().to_string(), + ) + println(" names: " + results.sequence_names().join(", ")) + + let pair = results.pair("Homo_sapie", "Pan_troglo").unwrap() + let yn = pair.yang_nielsen() + println("\n3. Five methods for Homo_sapie vs Pan_troglo") + println( + " NG86: omega=" + + pair.ng86().omega().to_string() + + ", dN=" + + pair.ng86().dn().to_string() + + ", dS=" + + pair.ng86().ds().to_string(), + ) + println( + " YN00: kappa=" + + yn.kappa().to_string() + + ", omega=" + + yn.omega().to_string() + + ", dN=" + + yn.dn().to_string() + + " +- " + + yn.dn_standard_error().to_string() + + ", dS=" + + yn.ds().to_string() + + " +- " + + yn.ds_standard_error().to_string(), + ) + println( + " LWL85: omega=" + + optional_value(pair.lwl85().omega()) + + ", dS=" + + optional_value(pair.lwl85().ds()), + ) + println( + " LWL85m: omega=" + + optional_value(pair.lwl85_modified().omega()) + + ", rho=" + + optional_value(pair.lwl85_modified().rho()), + ) + println( + " LPB93: omega=" + + optional_value(pair.lpb93().omega()) + + ", dS=" + + optional_value(pair.lpb93().ds()), + ) + + let matrix = yn00_demo_matrix(results) + println("\n4. Symmetric Yang-Nielsen dS matrix") + let names = matrix.names() + for first in names { + let row : Array[String] = [] + for second in names { + row.push(optional_value(matrix.value(first, second))) + } + println(" " + first + ": " + row.join(", ")) + } + println(" mean pairwise dS: " + yn00_demo_mean(results).to_string()) + + println("\nThe demo parses existing YN00 artifacts; it does not launch PAML.") +} diff --git a/examples/paml_yn00_demo/moon.pkg b/examples/paml_yn00_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/paml_yn00_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/protein_analysis_advanced_demo/main.mbt b/examples/protein_analysis_advanced_demo/main.mbt new file mode 100644 index 00000000..30b4b7dc --- /dev/null +++ b/examples/protein_analysis_advanced_demo/main.mbt @@ -0,0 +1,183 @@ +// Advanced protein-sequence prediction demo. +// +// Showcases six empirical protein-analysis algorithms: +// 1. Chou–Fasman secondary-structure prediction. +// 2. IUPred intrinsically-disordered-region prediction. +// 3. COILS coiled-coil prediction. +// 4. Kolaskar–Tongaonkar antigenicity prediction. +// 5. Emini surface-accessibility prediction. +// 6. Karplus–Schulz flexibility prediction. + +///| +fn paa_demo_round(value : Double, digits : Int) -> Double { + let scale = @math.pow(10.0, digits.to_double()) + (value * scale).round() / scale +} + +///| +fn paa_demo_format_ss(prediction : Array[String]) -> String { + let mut result = "" + for p in prediction { + result = result + p + } + result +} + +///| +fn main { + println("=== Protein Analysis Advanced Demo ===") + + // ----------------------------------------------------------------------- + // 1. Chou-Fasman secondary-structure prediction + // ----------------------------------------------------------------------- + println("\n1. Chou-Fasman secondary-structure prediction") + let helical = @src.demo_helical_sequence() + let cf = @src.chou_fasman_predict(helical) catch { + ProteinAdvancedError(msg) => abort("Chou-Fasman failed: " + msg) + } + println(" Sequence: " + helical) + println(" SS: " + paa_demo_format_ss(cf.prediction())) + println( + " Helices: " + + cf.helices().length().to_string() + + ", Sheets: " + + cf.sheets().length().to_string() + + ", Turns: " + + cf.turns().length().to_string(), + ) + for region in cf.helices() { + let (start, end) = region + println(" helix " + start.to_string() + "-" + end.to_string()) + } + + // ----------------------------------------------------------------------- + // 2. IUPred disorder prediction + // ----------------------------------------------------------------------- + println("\n2. IUPred disorder prediction") + let disorder_seq = @src.demo_disorder_sequence() + let iu = @src.iupred(disorder_seq, window_size=25) catch { + ProteinAdvancedError(msg) => abort("IUPred failed: " + msg) + } + let iu_scores = iu.scores() + let mut max_score = 0.0 + for s in iu_scores { + if s > max_score { + max_score = s + } + } + println(" Sequence: " + disorder_seq) + println(" Max IUPred score: " + paa_demo_round(max_score, 4).to_string()) + let regions = iu.disordered_regions() + if regions.length() > 0 { + for region in regions { + let (start, end) = region + println( + " Disordered region: " + + start.to_string() + + "-" + + end.to_string(), + ) + } + } else { + println(" No disordered regions detected") + } + + // ----------------------------------------------------------------------- + // 3. COILS coiled-coil prediction + // ----------------------------------------------------------------------- + println("\n3. COILS coiled-coil prediction") + let cc_seq = @src.demo_coiled_coil_sequence() + let cc = @src.predict_coiled_coils(cc_seq, window=14, threshold=0.5) catch { + ProteinAdvancedError(msg) => abort("COILS failed: " + msg) + } + println(" Sequence: " + cc_seq) + let cc_scores = cc.scores() + let mut cc_max = 0.0 + for s in cc_scores { + if s > cc_max { + cc_max = s + } + } + println(" Max COILS score: " + paa_demo_round(cc_max, 4).to_string()) + let cc_regions = cc.coiled_coil_regions() + if cc_regions.length() > 0 { + for region in cc_regions { + let (start, end) = region + println( + " Coiled-coil region: " + start.to_string() + "-" + end.to_string() + ) + } + } else { + println(" No coiled-coil regions detected") + } + + // ----------------------------------------------------------------------- + // 4. Kolaskar-Tongaonkar antigenicity prediction + // ----------------------------------------------------------------------- + println("\n4. Kolaskar-Tongaonkar antigenicity prediction") + let ag_seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let kt = @src.kolaskar_tongaonkar_antigenicity(ag_seq, window=7) catch { + ProteinAdvancedError(msg) => abort("Kolaskar failed: " + msg) + } + println(" Sequence: " + ag_seq) + let kt_scores = kt.scores() + let mut kt_max = 0.0 + for s in kt_scores { + if s > kt_max { + kt_max = s + } + } + println(" Max antigenicity score: " + paa_demo_round(kt_max, 4).to_string()) + let sites = kt.antigenic_sites() + println(" Antigenic sites: " + sites.length().to_string()) + for site in sites { + let (start, end) = site + println(" site " + start.to_string() + "-" + end.to_string()) + } + + // ----------------------------------------------------------------------- + // 5. Emini surface accessibility prediction + // ----------------------------------------------------------------------- + println("\n5. Emini surface accessibility prediction") + let em_scores = @src.emini_surface_accessibility(ag_seq, window=6) catch { + ProteinAdvancedError(msg) => abort("Emini failed: " + msg) + } + let mut em_max = 0.0 + let mut em_sum = 0.0 + for s in em_scores { + em_sum = em_sum + s + if s > em_max { + em_max = s + } + } + println( + " Mean: " + + paa_demo_round(em_sum / em_scores.length().to_double(), 4).to_string() + + ", Max: " + + paa_demo_round(em_max, 4).to_string(), + ) + + // ----------------------------------------------------------------------- + // 6. Karplus-Schulz flexibility prediction + // ----------------------------------------------------------------------- + println("\n6. Karplus-Schulz flexibility prediction") + let ks_scores = @src.karplus_schulz_flexibility(ag_seq, window=5) catch { + ProteinAdvancedError(msg) => abort("Karplus-Schulz failed: " + msg) + } + let mut ks_sum = 0.0 + let mut ks_max = 0.0 + for s in ks_scores { + ks_sum = ks_sum + s + if s > ks_max { + ks_max = s + } + } + println( + " Mean flexibility: " + + paa_demo_round(ks_sum / ks_scores.length().to_double(), 4).to_string() + + ", Max: " + + paa_demo_round(ks_max, 4).to_string(), + ) + + println("\n=== Demo complete ===") +} diff --git a/examples/protein_analysis_advanced_demo/moon.pkg b/examples/protein_analysis_advanced_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/protein_analysis_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/ranged_summarized_experiment_demo/main.mbt b/examples/ranged_summarized_experiment_demo/main.mbt new file mode 100644 index 00000000..9034770d --- /dev/null +++ b/examples/ranged_summarized_experiment_demo/main.mbt @@ -0,0 +1,112 @@ +///| +fn print_ranges(names : Array[String], ranges : @src.GRanges) -> Unit { + for index = 0; index < @src.granges_length(ranges); index = index + 1 { + println( + " \{names[index]}: \{ranges.seqnames[index]}:\{ranges.starts[index]}-\{ranges.ends[index]}", + ) + } +} + +///| +fn main { + println("=== RangedSummarizedExperiment Demo ===") + + let experiment = @src.RangedSummarizedExperiment::new( + assays=Map([ + ( + "counts", + [ + [30.0, 31.0, 32.0], + [10.0, 11.0, 12.0], + [20.0, 21.0, 22.0], + [40.0, 41.0, 42.0], + ], + ), + ( + "logcounts", + [ + [3.40, 3.43, 3.47], + [2.40, 2.48, 2.56], + [3.04, 3.09, 3.14], + [3.71, 3.74, 3.76], + ], + ), + ]), + row_ranges=@src.granges( + ["chr2", "chr1", "chr1", "chr3"], + [(300, 349), (100, 149), (200, 249), (50, 99)], + [ + @src.strand_plus(), + @src.strand_plus(), + @src.strand_minus(), + @src.strand_star(), + ], + ), + col_data=[ + Map([("sample", "S1"), ("condition", "control")]), + Map([("sample", "S2"), ("condition", "treated")]), + Map([("sample", "S3"), ("condition", "treated")]), + ], + row_data=[ + Map([("biotype", "protein_coding")]), + Map([("biotype", "lncRNA")]), + Map([("biotype", "protein_coding")]), + Map([("biotype", "protein_coding")]), + ], + row_names=["geneC", "geneA", "geneB", "geneD"], + metadata=Map([("study", "range-aware-expression")]), + ) catch { + _ => abort("failed to construct RangedSummarizedExperiment") + } + + println("\n1. Container and row ranges") + println(" " + experiment.summary()) + print_ranges(experiment.row_names(), experiment.row_ranges()) + + let regions = @src.granges(["chr1", "chr1"], [(120, 170), (220, 230)], [ + @src.strand_plus(), + @src.strand_minus(), + ]) + println("\n2. Strand-aware overlaps") + for hit in experiment.find_overlaps(regions) { + println(" \{experiment.row_names()[hit.0]} overlaps query \{hit.1}") + } + let overlapping = experiment.subset_by_overlaps(regions) catch { + _ => abort("failed to subset overlapping rows") + } + println(" Retained rows: " + overlapping.row_names().join(", ")) + + println("\n3. Nearest ranges and interval distances") + let nearest = experiment.nearest(regions) + let distances = experiment.distance_to_nearest(regions) + for index = 0; index < experiment.nrow(); index = index + 1 { + println( + " \{experiment.row_names()[index]}: subject=\{nearest[index]}, distance=\{distances[index]}", + ) + } + + println("\n4. Strand-aware promoter ranges") + let promoter_experiment = experiment.promoters(upstream=100, downstream=20) catch { + _ => abort("failed to create promoter ranges") + } + print_ranges( + promoter_experiment.row_names(), + promoter_experiment.row_ranges(), + ) + + println("\n5. Sort ranges and assays together") + let sorted = experiment.sort() catch { + _ => abort("failed to sort experiment") + } + println(" Row order: " + sorted.row_names().join(", ")) + match sorted.assay("counts") { + Some(counts) => + println( + " First-sample counts: " + + counts.map(fn(row) { row[0].to_string() }).join(", "), + ) + None => println(" counts assay is unavailable") + } + + println("\n=== Demo Complete ===") +} diff --git a/examples/ranged_summarized_experiment_demo/moon.pkg b/examples/ranged_summarized_experiment_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/ranged_summarized_experiment_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/sc_dbl_finder_demo/main.mbt b/examples/sc_dbl_finder_demo/main.mbt index 3b727dd7..f6a035bb 100644 --- a/examples/sc_dbl_finder_demo/main.mbt +++ b/examples/sc_dbl_finder_demo/main.mbt @@ -1,31 +1,134 @@ ///| fn main { - let data = @bio.scdf_create_example_data() + let (data, clusters, samples, known_doublets) = @bio.scdf_create_advanced_example() catch { + _ => abort("failed to create the scDblFinder example") + } + let config = @bio.ScDblFinderConfig::create( + expected_doublet_rate=0.15, + artificial_doublets=48, + selected_features=18, + dimensions=5, + neighbors=8, + iterations=2, + classifier_steps=120, + seed=17, + ) catch { + _ => abort("failed to create the scDblFinder configuration") + } + let result = @bio.sc_dbl_finder( + data, + clusters~, + samples~, + known_doublets~, + config~, + ) catch { + _ => abort("scDblFinder fitting failed") + } - println( - "Single cell data: " + - data.cell_names.length().to_string() + - " cells, " + - data.gene_names.length().to_string() + - " genes", - ) + println(result.summary()) + println("\nCapture thresholds") + for capture in 0.. abort("failed to rank doublets") + } + for cell in top { + println( + cell.cell_name + + "\tscore=" + + cell.score.to_string() + + "\tclass=" + + (if cell.is_doublet { "doublet" } else { "singlet" }) + + "\torigin=" + + cell.origin, + ) + } + + println("\nPairwise origin enrichment") + let enrichment = @bio.scdf_pairwise_enrichment(result) catch { + _ => abort("failed to compute origin enrichment") + } + for row in enrichment { + println( + row.combination + + "\tobserved=" + + row.observed.to_string() + + "\texpected=" + + row.expected.to_string() + + "\tadjusted_p=" + + row.adjusted_p_value.to_string(), + ) + } - let scores = @bio.scdf_compute_doublet_score(data, 10) + let singlets = result.filter_singlets(data) catch { + _ => abort("failed to filter singlets") + } println( - "Doublet scores computed for " + scores.length().to_string() + " cells", + "\nRetained " + + singlets.cell_names.length().to_string() + + " singlets from " + + data.cell_names.length().to_string() + + " cells", ) - let doublets = @bio.scdf_detect_doublets(scores, 0.3) - println("Detected doublets: " + doublets.length().to_string()) - - let summary = @bio.scdf_doublet_summary(scores) - println(summary) + let automatic_clusters = @bio.scdf_fast_cluster( + data, + 3, + dimensions=5, + selected_features=18, + ) catch { + _ => abort("automatic clustering failed") + } + println("Automatic clusters for first five cells") + for cell in 0..<5 { + println(data.cell_names[cell] + "\t" + automatic_clusters[cell]) + } - let filtered = @bio.scdf_filter_doublets(data, scores) + let assay : Array[Array[Double]] = [] + for gene in 0.. abort("SingleCellExperiment integration failed") + } println( - "After filtering: " + filtered.cell_names.length().to_string() + " cells", + "\nSingleCellExperiment write-back: score=" + + output.experiment.col_data["scDblFinder.score"][0] + + ", original_unchanged=" + + (!experiment.col_data.contains("scDblFinder.score")).to_string(), ) - - let pca = @bio.scdf_compute_pca(data, 5) - println("PCA computed: " + pca.length().to_string() + " cells, 5 components") } diff --git a/examples/scrapper_demo/main.mbt b/examples/scrapper_demo/main.mbt new file mode 100644 index 00000000..f3d8e329 --- /dev/null +++ b/examples/scrapper_demo/main.mbt @@ -0,0 +1,145 @@ +///| +fn format_doubles(values : Array[Double]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index].to_string() + index = index + 1 + } + result + "]" +} + +///| +fn format_ints(values : Array[Int]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index].to_string() + index = index + 1 + } + result + "]" +} + +///| +fn format_strings(values : Array[String]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index] + index = index + 1 + } + result + "]" +} + +///| +fn main { + println("=== Bioconductor scrapper Demo ===") + let counts = @src.scrapper_sample_counts() + let blocks = ["batch1", "batch1", "batch1", "batch2", "batch2", "batch2"] + let mitochondrial = @src.ScrapperNamedSubset::new("mito", [2, 4]) + + println("\n1. Batch-aware RNA quality control") + let qc = @src.scrapper_quick_rna_qc(counts, subsets=[mitochondrial], blocks~) catch { + _ => abort("failed to compute scrapper RNA QC") + } + println(" " + qc.summary()) + println(" library sums: " + format_doubles(qc.metrics.sums)) + println(" detected genes: " + format_ints(qc.metrics.detected)) + println( + " mitochondrial proportions: " + + format_doubles(qc.metrics.subset_proportions[0]), + ) + + println("\n2. Library-size factors and log-normalization") + let size_factors = @src.scrapper_library_size_factors( + counts, + blocks~, + mode=@src.scrapper_center_per_block(), + ) catch { + _ => abort("failed to compute library size factors") + } + let logcounts = @src.scrapper_normalize_counts(counts, size_factors) catch { + _ => abort("failed to normalize counts") + } + println(" size factors: " + format_doubles(size_factors)) + println(" first normalized feature: " + format_doubles(logcounts[0])) + + println("\n3. Mean-variance trend and highly variable genes") + let model = @src.scrapper_model_gene_variances( + logcounts, + mean_filter=false, + min_window_count=3, + ) catch { + _ => abort("failed to model gene variances") + } + let highly_variable = model.highly_variable_genes(top=3) catch { + _ => abort("failed to choose highly variable genes") + } + println(" means: " + format_doubles(model.means)) + println(" residual variances: " + format_doubles(model.residuals)) + println(" selected 0-based genes: " + format_ints(highly_variable)) + + println("\n4. Multi-factor pseudo-bulk aggregation") + let aggregate = @src.scrapper_aggregate_across_cells( + counts, + [ + @src.ScrapperFactor::new("cluster", ["T", "T", "B", "T", "B", "B"]), + @src.ScrapperFactor::new("batch", blocks), + ], + compute_median=true, + ) catch { + _ => abort("failed to aggregate cells") + } + let pseudo_bulk_means = aggregate.means() catch { + _ => abort("failed to compute pseudo-bulk means") + } + println(" " + aggregate.summary()) + println(" groups: " + format_strings(aggregate.group_names)) + println(" cells per group: " + format_ints(aggregate.counts)) + println(" first feature means: " + format_doubles(pseudo_bulk_means[0])) + + println("\n5. Immutable SingleCellExperiment integration") + let sce = @src.SingleCellExperiment::new( + counts, + ["Gene1", "Gene2", "Gene3", "Gene4", "Gene5", "Gene6"], + ["Cell1", "Cell2", "Cell3", "Cell4", "Cell5", "Cell6"], + ) + let normalized_sce = @src.scrapper_normalize_rna_counts_sce( + sce, + blocks~, + mode=@src.scrapper_center_per_block(), + ) catch { + _ => abort("failed to normalize SingleCellExperiment") + } + let (annotated_sce, sce_qc) = @src.scrapper_quick_rna_qc_sce( + normalized_sce, + subsets=[mitochondrial], + blocks~, + ) catch { + _ => abort("failed to annotate SingleCellExperiment QC") + } + println( + " original logcounts rows: " + + @src.sce_get_assay(sce, "logcounts").length().to_string(), + ) + println( + " copied logcounts rows: " + + @src.sce_get_assay(annotated_sce, "logcounts").length().to_string(), + ) + println( + " copied QC metadata columns populated: " + + @src.sce_get_col_data(annotated_sce, "scrapper_keep").length().to_string(), + ) + println(" SCE QC retained cells: " + sce_qc.retained_count().to_string()) + + println("\n=== Demo Complete ===") +} diff --git a/examples/scrapper_demo/moon.pkg b/examples/scrapper_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/scrapper_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/scuttle_demo/main.mbt b/examples/scuttle_demo/main.mbt new file mode 100644 index 00000000..fe33a1e2 --- /dev/null +++ b/examples/scuttle_demo/main.mbt @@ -0,0 +1,124 @@ +///| +fn scuttle_demo_doubles(values : Array[Double]) -> String { + values.map(fn(value) { value.to_string() }).join(", ") +} + +///| +fn main { + println("=== Bioconductor scuttle 1.23.1 Demo ===") + + let outlier_config = @bio.ScuttleOutlierConfig::create( + nmads=2.0, + direction=@bio.scuttle_outlier_higher(), + batches=["A", "A", "A", "B", "B", "B"], + ) catch { + _ => abort("failed to create the outlier configuration") + } + let outliers = @bio.scuttle_is_outlier( + [9.0, 10.0, 11.0, 98.0, 100.0, 150.0], + config=outlier_config, + ) catch { + _ => abort("batch-aware outlier detection failed") + } + println("\n1. Batch-aware MAD outlier detection") + println(" batches: " + outliers.batch_names.join(", ")) + println(" upper thresholds: " + scuttle_demo_doubles(outliers.higher)) + println(" discarded observations: " + outliers.discarded_count().to_string()) + + let counts = [ + [8.0, 4.0, 16.0, 8.0], + [2.0, 6.0, 4.0, 12.0], + [0.0, 2.0, 0.0, 4.0], + [10.0, 8.0, 20.0, 16.0], + ] + let controls = @bio.ScuttleNamedCellSubset::create("controls", [0, 1]) catch { + _ => abort("failed to create the control-cell subset") + } + let qc = @bio.scuttle_per_feature_qc(counts, subsets=[controls]) catch { + _ => abort("per-feature QC failed") + } + println("\n2. Per-feature QC with a named cell subset") + println(" means: " + scuttle_demo_doubles(qc.means)) + println( + " detected percentages: " + scuttle_demo_doubles(qc.detected_percent), + ) + println( + " control/global ratios: " + scuttle_demo_doubles(qc.subset_ratios[0]), + ) + + let pathway_a = @bio.ScuttleFeatureSet::create("PathwayA", [0, 1]) catch { + _ => abort("failed to create PathwayA") + } + let pathway_b = @bio.ScuttleFeatureSet::create("PathwayB", [1, 2, 3]) catch { + _ => abort("failed to create PathwayB") + } + let aggregated = @bio.scuttle_aggregate_feature_sets(counts, [ + pathway_a, pathway_b, + ]) catch { + _ => abort("feature-set aggregation failed") + } + println("\n3. Overlapping feature-set aggregation") + for index in 0.. abort("batch coverage equalization failed") + } + println("\n4. Exact batch coverage equalization") + println(" batch summaries: " + scuttle_demo_doubles(balanced.summaries)) + println(" proportions: " + scuttle_demo_doubles(balanced.proportions)) + + let experiment = @bio.SingleCellExperiment::new( + counts, + ["G1", "G2", "G3", "G4"], + ["C1", "C2", "C3", "C4"], + ) + experiment.col_data["batch"] = ["A", "A", "B", "B"] + experiment.metadata["source"] = "scuttle_demo" + let qc_output = @bio.scuttle_per_feature_qc_sce(experiment, subsets=[controls]) catch { + _ => abort("SCE per-feature QC failed") + } + let downsampled = @bio.scuttle_downsample_sce( + qc_output.experiment, + 0.5, + output_assay="half", + seed=42, + ) catch { + _ => abort("SCE downsampling failed") + } + let pathway_sce = @bio.scuttle_aggregate_feature_sets_sce( + downsampled, + [pathway_a, pathway_b], + assay_names=["counts", "half"], + ) catch { + _ => abort("SCE feature aggregation failed") + } + println("\n5. Immutable SingleCellExperiment integration") + println( + " QC rows written: " + + qc_output.experiment.row_data["scuttle.mean"].length().to_string(), + ) + println( + " downsampled assay rows: " + + downsampled.assays["half"].length().to_string(), + ) + println(" aggregated rows: " + pathway_sce.row_names.join(", ")) + println( + " original unchanged: " + + (!experiment.row_data.contains("scuttle.mean") && + !experiment.assays.contains("half")).to_string(), + ) + println("\nscuttle demo completed.") +} diff --git a/examples/scuttle_demo/moon.pkg b/examples/scuttle_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/scuttle_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/shared_reference_alignment_demo/main.mbt b/examples/shared_reference_alignment_demo/main.mbt new file mode 100644 index 00000000..125bd418 --- /dev/null +++ b/examples/shared_reference_alignment_demo/main.mbt @@ -0,0 +1,103 @@ +///| +fn main { + println("=== Biopython Shared-reference Alignment Demo ===") + + println("\n1. Adapt an existing pairwise alignment") + let pairwise = @src.pairaligner_align("ACGT", "ACT") + let pairwise_input = @src.shared_reference_input_from_pairwise( + pairwise, + reference_name="reference", + reference_description="shared DNA template", + query_name="pairwise_read", + query_description="one deletion", + ) catch { + _ => abort("failed to adapt pairwise alignment") + } + println(" reference: " + pairwise_input.aligned_reference) + println(" query: " + pairwise_input.queries[0].aligned_sequence) + + println("\n2. Build a multi-query alignment with an insertion") + let inserted = @src.shared_reference_sequence( + "inserted_read", + "ACGGT", + description="one insertion", + ) catch { + _ => abort("failed to build inserted query") + } + let deleted = @src.shared_reference_sequence( + "deleted_read", + "A---T", + description="three deletions", + ) catch { + _ => abort("failed to build deleted query") + } + let multiple_input = @src.shared_reference_input( + "ACGT", + "ACG-T", + [inserted, deleted], + reference_name="reference", + reference_description="shared DNA template", + ) catch { + _ => abort("failed to build multi-query input") + } + + println("\n3. Synchronize reference-boundary insertion slots") + let alignment = @src.alignments_with_same_reference([ + pairwise_input, multiple_input, + ]) catch { + _ => abort("failed to merge shared-reference alignments") + } + println(" " + alignment.summary()) + println(" insertion widths: " + alignment.insertion_widths.to_string()) + for index in 0.. reference: " + + alignment.query_to_reference(1, 3).to_string(), + ) + println( + " inserted_read query position 4 -> reference: " + + alignment.query_to_reference(1, 4).to_string(), + ) + println( + " reference position 3 -> inserted_read query: " + + alignment.reference_to_query(1, 3).to_string(), + ) + + println("\n5. Inspect alignment statistics") + for index in 0.. abort("failed to convert merged alignment") + } + println( + " MSA dimensions: " + + msa.num_records().to_string() + + " x " + + msa.get_alignment_length().to_string(), + ) + println(alignment.to_fasta()) + + println("=== Demo Complete ===") +} diff --git a/examples/shared_reference_alignment_demo/moon.pkg b/examples/shared_reference_alignment_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/shared_reference_alignment_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/single_r_demo/main.mbt b/examples/single_r_demo/main.mbt index f179fd02..c40140d4 100644 --- a/examples/single_r_demo/main.mbt +++ b/examples/single_r_demo/main.mbt @@ -1,81 +1,148 @@ ///| -/// SingleR Demo - Cell type annotation by reference expression profiles. -/// Inspired by Bioconductor SingleR package. +fn single_r_demo_optional_label(value : String?) -> String { + match value { + Some(label) => label + None => "NA" + } +} ///| fn main { - println("=== SingleR Cell Type Annotation Demo ===") - println("Inspired by Bioconductor SingleR package") - println("") - - println("1. Creating reference dataset with known cell type profiles...") - let ref_data = @bio.single_r_create_reference_data() - println(" Reference profiles: \{ref_data.n_profiles.to_string()}") - println(" Cell types:") - for ct in ref_data.cell_types { - println(" - \{ct}") + println("=== Bioconductor SingleR 2.15.2 Advanced Demo ===") + let (reference, data) = @bio.single_r_advanced_example() catch { + _ => abort("failed to construct the SingleR example") + } + println( + "\n1. Gene x sample/cell inputs: reference=" + + reference.n_genes.to_string() + + "x" + + reference.n_samples.to_string() + + ", test=" + + data.n_genes.to_string() + + "x" + + data.n_cells.to_string(), + ) + + let training = @bio.single_r_advanced_train(reference, data.gene_names) catch { + _ => abort("SingleR marker training failed") + } + let t_vs_b = training.markers_between("T cell", "B cell") catch { + _ => abort("failed to query pairwise markers") + } + println("\n2. Directed classic marker training") + println( + " shared genes=" + + training.common_gene_names.length().to_string() + + ", selected markers=" + + training.marker_gene_names().length().to_string(), + ) + println(" T cell > B cell markers: " + t_vs_b.join(", ")) + + let config = @bio.SingleRAdvancedConfig::create( + quantile=0.8, + fine_tune=true, + prune=true, + nmads=3.0, + ) catch { + _ => abort("failed to construct the SingleR configuration") + } + let result = @bio.single_r_advanced_classify(data, training, config~) catch { + _ => abort("SingleR cell annotation failed") } - println(" Genes per profile: \{ref_data.n_genes.to_string()}") - println("") - - println("2. Creating synthetic single-cell data to annotate...") - let test_data = @bio.single_r_create_test_data() - println(" Cells: \{test_data.n_cells.to_string()}") - println(" Genes: \{test_data.n_genes.to_string()}") - println("") - - println("3. Running SingleR annotation (Spearman correlation)...") - let params = @bio.SingleRParams::new() - let result = @bio.single_r_annotate_cells(test_data, ref_data, params) - println(" Annotation completed for \{result.cell_ids.length().to_string()} cells") - println("") - - println("4. Annotation results summary:") - println(" First 10 annotations:") - let mut i = 0 - while i < 10 && i < result.cell_ids.length() { - println(" Cell \{result.cell_ids[i]}: \{result.labels[i]} (score: \{result.scores[i].to_string()})") - i = i + 1 + println("\n3. Quantile scores, iterative fine-tuning and pruning") + for cell in 0.. abort("SingleR cluster annotation failed") + } + println("\n4. Cluster-level annotation from summed profiles") + for cluster in 0.. abort("failed to construct the second reference") } - println("") - - println("6. Running SingleR with Pearson correlation...") - let params_pearson = @bio.SingleRParams::with_method("pearson") - let result_pearson = @bio.single_r_annotate_cells(test_data, ref_data, params_pearson) - println(" Pearson annotation completed") - println(" First 5 annotations (Pearson):") - let mut j = 0 - while j < 5 && j < result_pearson.cell_ids.length() { - println(" Cell \{result_pearson.cell_ids[j]}: \{result_pearson.labels[j]} (score: \{result_pearson.scores[j].to_string()})") - j = j + 1 + let combined = @bio.single_r_advanced_combine( + data, + [reference, alternate], + config~, + ) catch { + _ => abort("SingleR multi-reference integration failed") + } + println("\n5. Multi-reference marker-space recomputation") + for cell in 0.. abort("SingleR SingleCellExperiment integration failed") } - println("") - - println("7. Correlation analysis example:") - let t_cell_expr = ref_data.profiles[0].expression - let b_cell_expr = ref_data.profiles[1].expression - let t_vs_b_spearman = @bio.single_r_spearman_correlation(t_cell_expr, b_cell_expr) - let t_vs_self_spearman = @bio.single_r_spearman_correlation(t_cell_expr, t_cell_expr) - let t_vs_b_pearson = @bio.single_r_pearson_correlation(t_cell_expr, b_cell_expr) - println(" T cell vs B cell (Spearman): \{t_vs_b_spearman.to_string()}") - println(" T cell vs T cell (Spearman): \{t_vs_self_spearman.to_string()}") - println(" T cell vs B cell (Pearson): \{t_vs_b_pearson.to_string()}") - println("") - - println("=== Demo completed successfully! ===") - println("") - println("Key takeaways:") - println("- SingleR annotates cells by comparing expression profiles to reference datasets") - println("- Spearman correlation is robust to outliers; Pearson is sensitive to magnitude") - println("- Fine-tuning can resolve ambiguous assignments by comparing against related cell types") - println("- The delta score (difference between top and 2nd best) indicates annotation confidence") - println("- SingleR can be used with any reference dataset (bulk RNA-seq, sorted populations)") -} \ No newline at end of file + println("\n6. Immutable SingleCellExperiment write-back") + println( + " mode=" + + output.experiment.metadata["SingleR215.mode"] + + ", labels=" + + output.experiment.col_data["SingleR215.labels"].join(", "), + ) + println( + " original unchanged=" + + (!experiment.col_data.contains("SingleR215.labels")).to_string(), + ) + println("\nSingleR advanced demo completed.") +} diff --git a/examples/slingshot_advanced_demo/main.mbt b/examples/slingshot_advanced_demo/main.mbt new file mode 100644 index 00000000..af7837b8 --- /dev/null +++ b/examples/slingshot_advanced_demo/main.mbt @@ -0,0 +1,163 @@ +///| +fn slingshot_demo_coordinates() -> Array[Array[Double]] { + [ + [-0.1, 0.0], + [0.0, 0.1], + [0.1, -0.1], + [0.9, 0.0], + [1.0, 0.1], + [1.1, -0.1], + [1.9, 0.0], + [2.0, 0.1], + [2.1, -0.1], + [2.9, 0.9], + [3.0, 1.0], + [3.1, 1.1], + [2.9, -0.9], + [3.0, -1.0], + [3.1, -1.1], + ] +} + +///| +fn slingshot_demo_labels() -> Array[String] { + ["A", "A", "A", "B", "B", "B", "C", "C", "C", "D", "D", "D", "E", "E", "E"] +} + +///| +fn slingshot_demo_lineage( + lineage : Array[Int], + cluster_names : Array[String], +) -> String { + lineage.map(fn(index) { cluster_names[index] }).join(" -> ") +} + +///| +fn slingshot_demo_doubles(values : Array[Double]) -> String { + values.map(fn(value) { value.to_string() }).join(", ") +} + +///| +fn main { + println("=== Bioconductor slingshot 2.21.0 Advanced Demo ===") + let coordinates = slingshot_demo_coordinates() + let labels = slingshot_demo_labels() + let config = @bio.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@bio.slingshot_scaled_full(), + extension=@bio.slingshot_extend_line(), + shrink=1.0, + reweight=true, + reassign=true, + max_iterations=8, + tolerance=1.0e-4, + smoother_span=0.4, + curve_points=30, + ) catch { + _ => abort("failed to create the advanced slingshot configuration") + } + let result = @bio.slingshot_advanced(coordinates, labels, config) catch { + _ => abort("advanced slingshot fitting failed") + } + + println("\n1. Constrained cluster minimum spanning tree") + for edge in result.lineage_model.edges { + println( + " " + + result.lineage_model.cluster_names[edge.from] + + " -> " + + result.lineage_model.cluster_names[edge.to] + + ", distance=" + + edge.distance.to_string(), + ) + } + + println("\n2. Root-to-leaf lineages and simultaneous curves") + for index in 0.. abort("new-cell trajectory projection failed") + } + println("\n4. New-cell projection") + for cell in 0.. abort("SingleCellExperiment trajectory integration failed") + } + println("\n5. Immutable SingleCellExperiment integration") + println(" output columns: slingshot.branch, slingshot.pseudotime") + println( + " pseudotime matrix: " + + sce_output.experiment.reduced_dims["slingshot.pseudotime"] + .length() + .to_string() + + " cells x " + + sce_output.experiment.reduced_dims["slingshot.pseudotime"][0] + .length() + .to_string() + + " lineages", + ) + println( + " original unchanged: " + + (!experiment.col_data.contains("slingshot.branch")).to_string(), + ) + println("\n" + @bio.slingshot_advanced_summary(result)) +} diff --git a/examples/slingshot_advanced_demo/moon.pkg b/examples/slingshot_advanced_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/slingshot_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/sparse_array_demo/main.mbt b/examples/sparse_array_demo/main.mbt new file mode 100644 index 00000000..bde89909 --- /dev/null +++ b/examples/sparse_array_demo/main.mbt @@ -0,0 +1,138 @@ +///| +fn format_doubles(values : Array[Double]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index].to_string() + index = index + 1 + } + result + "]" +} + +///| +fn format_ints(values : Array[Int]) -> String { + let mut result = "[" + let mut index = 0 + while index < values.length() { + if index > 0 { + result = result + ", " + } + result = result + values[index].to_string() + index = index + 1 + } + result + "]" +} + +///| +fn format_matrix(values : Array[Array[Double]]) -> String { + let mut result = "[" + let mut row = 0 + while row < values.length() { + if row > 0 { + result = result + ", " + } + result = result + format_doubles(values[row]) + row = row + 1 + } + result + "]" +} + +///| +fn main { + println("=== Bioconductor SparseArray Demo ===\n") + + let tensor = @src.sparse_array_sample() catch { + _ => abort("failed to create sparse tensor") + } + let tensor_summary = tensor.summary() catch { + _ => abort("failed to summarize sparse tensor") + } + let tensor_value = tensor.get([0, 4, 1]) catch { + _ => abort("failed to access sparse tensor") + } + println("1. Multidimensional sparse tensor") + println(" " + tensor_summary) + println(" value[0, 4, 1] = " + tensor_value.to_string()) + println( + " stored coordinates = " + tensor.nzcoordinates().length().to_string(), + ) + + let permuted = tensor.aperm([2, 0, 1]) catch { + _ => abort("failed to permute sparse tensor") + } + let permuted_summary = permuted.summary() catch { + _ => abort("failed to summarize permuted tensor") + } + let permuted_value = permuted.get([1, 0, 4]) catch { + _ => abort("failed to access permuted tensor") + } + println("\n2. Dimension permutation") + println(" " + permuted_summary) + println(" old [0, 4, 1] -> new [1, 0, 4] = " + permuted_value.to_string()) + + let sliced = tensor.slice([0, 0, 1], [4, 5, 2]) catch { + _ => abort("failed to slice sparse tensor") + } + let sliced_summary = sliced.summary() catch { + _ => abort("failed to summarize sparse slice") + } + println("\n3. Sparse slice without dense materialization") + println(" " + sliced_summary) + + let counts = @src.SparseArray::from_matrix([ + [1.0, 0.0, 2.0], + [0.0, 3.0, 0.0], + [4.0, 0.0, 5.0], + ]) catch { + _ => abort("failed to create sparse count matrix") + } + let counts_summary = counts.summary() catch { + _ => abort("failed to summarize sparse count matrix") + } + let row_sums = counts.row_sums() catch { + _ => abort("failed to compute row sums") + } + let column_sums = counts.column_sums() catch { + _ => abort("failed to compute column sums") + } + let row_nonzero_counts = counts.row_nonzero_counts() catch { + _ => abort("failed to count row nonzero values") + } + println("\n4. Matrix summaries") + println(" " + counts_summary) + println(" row sums: " + format_doubles(row_sums)) + println(" column sums: " + format_doubles(column_sums)) + println(" row nonzero counts: " + format_ints(row_nonzero_counts)) + + let design = @src.SparseArray::from_matrix([ + [1.0, 0.0], + [0.0, 1.0], + [1.0, 1.0], + ]) catch { + _ => abort("failed to create sparse design matrix") + } + let product = counts.matmul(design) catch { + _ => abort("failed to multiply sparse matrices") + } + println("\n5. Sparse matrix multiplication") + println(" result: " + format_matrix(product)) + + let normalized = counts.scale(0.5) catch { + _ => abort("failed to scale sparse matrix") + } + let combined = counts.add(normalized) catch { + _ => abort("failed to add sparse matrices") + } + let combined_summary = combined.summary() catch { + _ => abort("failed to summarize sparse arithmetic result") + } + let dense_combined = combined.to_dense_matrix() catch { + _ => abort("failed to materialize sparse arithmetic result") + } + println("\n6. Sparse arithmetic") + println(" counts + counts * 0.5: " + combined_summary) + println(" dense view: " + format_matrix(dense_combined)) +} diff --git a/examples/sparse_array_demo/moon.pkg b/examples/sparse_array_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/sparse_array_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/spatialdecon_demo/main.mbt b/examples/spatialdecon_demo/main.mbt new file mode 100644 index 00000000..400d8f23 --- /dev/null +++ b/examples/spatialdecon_demo/main.mbt @@ -0,0 +1,171 @@ +// Bioconductor SpatialDecon-inspired background-aware spatial deconvolution. + +///| +fn main { + let (data, profile, nuclei_counts) = @src.spatial_decon_example_data() catch { + SpatialDeconError(message) => abort("example data failed: " + message) + } + let config = @src.SpatialDeconConfig::create( + rescale_profile=false, + tolerance=1.0e-9, + ) catch { + SpatialDeconError(message) => abort("invalid configuration: " + message) + } + let result = @src.spatial_decon(data, profile, nuclei_counts~, config~) catch { + SpatialDeconError(message) => abort("deconvolution failed: " + message) + } + + println("=== Background-aware spatial deconvolution ===") + println(result.summary()) + println("spot\tcell type\tabundance\tproportion\tcells") + for spot in 0.. abort("ranking failed: " + message) + } + for entry in ranked { + let cell_type = entry.cell_type + let mut cell_index = 0 + while result.cell_types[cell_index] != cell_type { + cell_index = cell_index + 1 + } + println( + entry.spot_id + + "\t" + + cell_type + + "\t" + + entry.abundance.to_string() + + "\t" + + entry.proportion.to_string() + + "\t" + + result.cell_counts[cell_index][spot].to_string(), + ) + } + } + println("fit diagnostics:") + for spot in 0.. abort("invalid cell merge: " + message) + } + let myeloid = @src.SpatialDeconCellMerge::create("Myeloid", ["Myeloid"]) catch { + SpatialDeconError(message) => abort("invalid cell merge: " + message) + } + let collapsed = @src.collapse_spatial_decon(result, [lymphoid, myeloid]) catch { + SpatialDeconError(message) => abort("cell-type collapse failed: " + message) + } + println("\n=== Collapsed cell types ===") + for group in 0.. + abort("reverse deconvolution failed: " + message) + } + println("\n=== Reverse deconvolution ===") + println("first-gene intercept=" + reverse.coefficients[0][0].to_string()) + for cell_type in 0.. + abort("background estimation failed: " + message) + } + println("\n=== Probe-pool background ===") + for row in background { + println(row[0].to_string() + "\t" + row[1].to_string()) + } + + let learned_profile = @src.create_spatial_decon_profile( + ["g1", "g2", "g3"], + ["a1", "a2", "b1", "b2"], + ["A", "A", "B", "B"], + [[10.0, 8.0, 1.0, 2.0], [1.0, 2.0, 9.0, 11.0], [2.0, 2.0, 2.0, 2.0]], + normalize=true, + scaling_factor=5.0, + min_cells=1, + min_genes=0, + ) catch { + SpatialDeconError(message) => + abort("profile construction failed: " + message) + } + println("\n=== Single-cell-derived profile ===") + println( + "genes=" + + learned_profile.n_genes().to_string() + + ", cell types=" + + learned_profile.n_cell_types().to_string(), + ) + for cell_type in learned_profile.cell_types { + println("profile cell type: " + cell_type) + } + + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "counts", data.values)) + ignore(@src.se_add_assay(experiment, "background", data.background)) + for gene in data.gene_names { + ignore(@src.se_add_row(experiment, Map([("gene_id", gene)]))) + } + for spot in 0.. + abort("SpatialExperiment integration failed: " + message) + } + println("\n=== SpatialExperiment integration ===") + println( + "genes=" + + integrated.experiment.metadata["spatialdecon_genes"] + + ", spots=" + + integrated.experiment.metadata["spatialdecon_spots"] + + ", original unchanged=" + + (!experiment.col_data[0].contains("SpatialDecon:T_cell")).to_string(), + ) +} diff --git a/examples/spatialdecon_demo/moon.pkg b/examples/spatialdecon_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/spatialdecon_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/spicyr_demo/main.mbt b/examples/spicyr_demo/main.mbt new file mode 100644 index 00000000..85d34925 --- /dev/null +++ b/examples/spicyr_demo/main.mbt @@ -0,0 +1,141 @@ +// Bioconductor spicyR-inspired differential spatial colocalization. + +///| +fn main { + let (cells, metadata, pairs) = @src.spicyr_example_data() + let config = @src.SpicyConfig::create( + radii=[0.5, 1.0, 2.0], + edge_correction=true, + use_weights=true, + reference_condition="control", + fdr_threshold=0.1, + tolerance=0.01, + ) catch { + SpicyError(message) => abort("invalid spicyR configuration: " + message) + } + + let result = @src.spicyr( + cells, + metadata, + pairs~, + covariate_names=["age"], + config~, + ) catch { + SpicyError(message) => abort("spicyR analysis failed: " + message) + } + + println("=== Direct spicyR analysis ===") + println(result.summary()) + println("pair\tcondition\testimate\tSE\tp-value\tFDR\tmodel") + for pair_result in result.pairs { + for condition in pair_result.conditions { + let model = if pair_result.model + is @src.SpicyModelKind::WeightedRandomIntercept { + "random-intercept" + } else { + "linear" + } + println( + pair_result.pair.label() + + "\t" + + condition.condition_level + + "\t" + + condition.estimate.to_string() + + "\t" + + condition.standard_error.to_string() + + "\t" + + condition.p_value.to_string() + + "\t" + + condition.adjusted_p_value.to_string() + + "\t" + + model, + ) + } + } + + let first_pair = result.pairs[0] + let first_image = first_pair.associations[0] + println("\n=== Per-image cross-L curve ===") + println( + "image=" + + first_image.image_id + + ", pair=" + + first_pair.pair.label() + + ", summary=" + + first_image.statistic.to_string() + + ", precision_weight=" + + first_image.weight.to_string(), + ) + for index in 0.. + abort("precomputed spicyR association fit failed: " + message) + } + println("\n=== Precomputed association input ===") + println(refitted.summary()) + + let experiment = @src.SpatialExperiment::new() + for cell in cells { + let mut image_index = -1 + for index in 0.. + abort("spicyR SpatialExperiment integration failed: " + message) + } + println("\n=== SpatialExperiment integration ===") + println( + "images=" + + integrated.experiment.metadata["spicyr_images"] + + ", pairs=" + + integrated.experiment.metadata["spicyr_pairs"] + + ", reference=" + + integrated.experiment.metadata["spicyr_reference"], + ) +} diff --git a/examples/spicyr_demo/moon.pkg b/examples/spicyr_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/spicyr_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/tradeseq_advanced_demo/main.mbt b/examples/tradeseq_advanced_demo/main.mbt new file mode 100644 index 00000000..986a6a27 --- /dev/null +++ b/examples/tradeseq_advanced_demo/main.mbt @@ -0,0 +1,290 @@ +///| +fn tradeseq_demo_counts() -> Array[Array[Double]] { + let flat : Array[Double] = [] + let association : Array[Double] = [] + let endpoint : Array[Double] = [] + let early_branch : Array[Double] = [] + let zero : Array[Double] = [] + for cell in 0..<24 { + flat.push((10 + cell % 3).to_double()) + if cell < 8 { + association.push((3 + cell).to_double()) + endpoint.push((6 + cell % 2).to_double()) + early_branch.push((8 + cell % 2).to_double()) + } else if cell < 16 { + let step = cell - 8 + association.push((12 + 3 * step).to_double()) + endpoint.push((10 + 4 * step).to_double()) + early_branch.push( + if step < 4 { + (26 + step).to_double() + } else { + (14 + step % 2).to_double() + }, + ) + } else { + let step = cell - 16 + association.push((12 + 3 * step).to_double()) + endpoint.push((10 + step / 3).to_double()) + early_branch.push( + if step < 4 { + (4 + step).to_double() + } else { + (14 + step % 2).to_double() + }, + ) + } + zero.push(0.0) + } + [flat, association, endpoint, early_branch, zero] +} + +///| +fn tradeseq_demo_trajectory() -> (Array[Array[Double]], Array[Array[Double]]) { + let pseudotime : Array[Array[Double]] = [] + let weights : Array[Array[Double]] = [] + for cell in 0..<24 { + if cell < 8 { + let time = cell.to_double() * 0.05 + pseudotime.push([time, time]) + weights.push([0.5, 0.5]) + } else if cell < 16 { + let time = 0.4 + (cell - 8).to_double() * 0.08 + pseudotime.push([time, 0.0]) + weights.push([1.0, 0.0]) + } else { + let time = 0.4 + (cell - 16).to_double() * 0.08 + pseudotime.push([0.0, time]) + weights.push([0.0, 1.0]) + } + } + (pseudotime, weights) +} + +///| +fn tradeseq_demo_config() -> @bio.TradeSeqAdvancedConfig { + @bio.TradeSeqAdvancedConfig::create( + n_knots=4, + smoothing_penalty=1.0, + max_iterations=80, + tolerance=1.0e-5, + ridge=1.0e-6, + fdr_threshold=0.1, + test_points=6, + ) catch { + _ => abort("failed to create tradeSeq configuration") + } +} + +///| +fn tradeseq_demo_print_test(table : @bio.TradeSeqAdvancedTestTable) -> Unit { + println("\n" + table.test_name) + for result in table.results { + println( + " " + + result.gene_id + + ": Wald=" + + result.wald_statistic.to_string() + + ", df=" + + result.degrees_freedom.to_string() + + ", p=" + + result.p_value.to_string() + + ", FDR=" + + result.adjusted_p_value.to_string() + + ", max log2FC=" + + result.log2_fold_change.to_string(), + ) + } +} + +///| +fn tradeseq_demo_slingshot() -> @bio.SlingshotAdvancedResult { + let coordinates = [ + [-0.1, 0.0], + [0.0, 0.1], + [0.1, -0.1], + [0.9, 0.0], + [1.0, 0.1], + [1.1, -0.1], + [1.9, 0.0], + [2.0, 0.1], + [2.1, -0.1], + [2.9, 0.9], + [3.0, 1.0], + [3.1, 1.1], + [2.9, -0.9], + [3.0, -1.0], + [3.1, -1.1], + ] + let labels = [ + "A", "A", "A", "B", "B", "B", "C", "C", "C", "D", "D", "D", "E", "E", "E", + ] + let config = @bio.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@bio.slingshot_center_euclidean(), + extension=@bio.slingshot_extend_none(), + max_iterations=6, + curve_points=24, + ) catch { + _ => abort("failed to create Slingshot configuration") + } + @bio.slingshot_advanced(coordinates, labels, config) catch { + _ => abort("Slingshot fitting failed") + } +} + +///| +fn main { + println("=== Bioconductor tradeSeq 1.27.0 Advanced Demo ===") + let counts = tradeseq_demo_counts() + let (pseudotime, weights) = tradeseq_demo_trajectory() + let gene_names = ["flat", "association", "endpoint", "early", "zero"] + let config = tradeseq_demo_config() + let fit = @bio.tradeseq_fit_advanced( + counts, + pseudotime, + weights, + gene_names~, + lineage_names=["left", "right"], + offsets=Array::make(24, 0.0), + config~, + ) catch { + _ => abort("tradeSeq NB-GAM fitting failed") + } + + println("\n1. Multi-lineage negative-binomial GAM") + println(@bio.tradeseq_advanced_summary(fit)) + for model in fit.models { + println( + " " + + model.gene_id + + ": dispersion=" + + model.dispersion.to_string() + + ", AIC=" + + model.aic.to_string() + + ", converged=" + + model.converged.to_string(), + ) + } + + println("\n2. Trajectory Wald tests with BH-FDR") + let association = @bio.tradeseq_association_test_advanced(fit) catch { + _ => abort("associationTest failed") + } + let start_end = @bio.tradeseq_start_vs_end_test_advanced(fit) catch { + _ => abort("startVsEndTest failed") + } + let diff_end = @bio.tradeseq_diff_end_test_advanced(fit) catch { + _ => abort("diffEndTest failed") + } + let pattern = @bio.tradeseq_pattern_test_advanced(fit, n_points=5) catch { + _ => abort("patternTest failed") + } + let early = @bio.tradeseq_early_de_test_advanced(fit, 0.35, 0.75, n_points=5) catch { + _ => abort("earlyDETest failed") + } + tradeseq_demo_print_test(association) + tradeseq_demo_print_test(start_end) + tradeseq_demo_print_test(diff_end) + tradeseq_demo_print_test(pattern) + tradeseq_demo_print_test(early) + + let prediction = @bio.tradeseq_predict_smooth_advanced( + fit, + "endpoint", + n_points=8, + ) catch { + _ => abort("smooth prediction failed") + } + println("\n3. Smooth prediction with delta-method uncertainty") + for lineage in 0.. abort("knot evaluation failed") + } + println("\n4. AIC knot evaluation") + for index in 0.. abort("Slingshot to tradeSeq integration failed") + } + println("\n5. Direct Slingshot integration") + println( + " " + + slingshot_fit.n_cells.to_string() + + " cells, lineages=" + + slingshot_fit.lineage_names.join(","), + ) + + let cell_names : Array[String] = [] + for cell in 0..<24 { + cell_names.push("cell" + (cell + 1).to_string()) + } + let experiment = @bio.SingleCellExperiment::new( + counts, gene_names, cell_names, + ) + experiment.reduced_dims["slingshot.pseudotime"] = pseudotime + experiment.reduced_dims["slingshot.weights"] = weights + let sce_output = @bio.tradeseq_advanced_sce(experiment, config~) catch { + _ => abort("SingleCellExperiment tradeSeq integration failed") + } + println("\n6. Immutable SingleCellExperiment write-back") + println( + " fitted assay: " + + sce_output.experiment.assays["tradeSeq.fitted"].length().to_string() + + " genes x " + + sce_output.experiment.assays["tradeSeq.fitted"][0].length().to_string() + + " cells", + ) + println( + " original unchanged: " + + (!experiment.assays.contains("tradeSeq.fitted")).to_string(), + ) +} diff --git a/examples/tradeseq_advanced_demo/moon.pkg b/examples/tradeseq_advanced_demo/moon.pkg new file mode 100644 index 00000000..4ecfd216 --- /dev/null +++ b/examples/tradeseq_advanced_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src" @bio, +} + +pkgtype(kind: "executable") diff --git a/examples/tree_summarized_experiment_demo/main.mbt b/examples/tree_summarized_experiment_demo/main.mbt new file mode 100644 index 00000000..84fac27d --- /dev/null +++ b/examples/tree_summarized_experiment_demo/main.mbt @@ -0,0 +1,133 @@ +///| +fn make_taxa_tree() -> @src.Tree { + let taxon_a = @src.Clade::new(name=Some("TaxonA")) + let taxon_b = @src.Clade::new(name=Some("TaxonB")) + let taxon_c = @src.Clade::new(name=Some("TaxonC")) + let taxon_d = @src.Clade::new(name=Some("TaxonD")) + let firmicutes = @src.Clade::new(name=Some("Firmicutes"), clades=[ + taxon_a, taxon_b, + ]) + let bacteroidota = @src.Clade::new(name=Some("Bacteroidota"), clades=[ + taxon_c, taxon_d, + ]) + let bacteria = @src.Clade::new(name=Some("Bacteria"), clades=[ + firmicutes, bacteroidota, + ]) + @src.Tree::new(bacteria, rooted=true, name=Some("taxonomy")) +} + +///| +fn make_sample_tree() -> @src.Tree { + let control_1 = @src.Clade::new(name=Some("Control1")) + let control_2 = @src.Clade::new(name=Some("Control2")) + let treated_1 = @src.Clade::new(name=Some("Treated1")) + let treated_2 = @src.Clade::new(name=Some("Treated2")) + let controls = @src.Clade::new(name=Some("Controls"), clades=[ + control_1, control_2, + ]) + let treated = @src.Clade::new(name=Some("Treated"), clades=[ + treated_1, treated_2, + ]) + let samples = @src.Clade::new(name=Some("Samples"), clades=[controls, treated]) + @src.Tree::new(samples, rooted=true, name=Some("sample_groups")) +} + +///| +fn print_assay(data : Array[Array[Double]]) -> Unit { + for row in data { + println(" [" + row.map(fn(value) { value.to_string() }).join(", ") + "]") + } +} + +///| +fn main { + println("=== TreeSummarizedExperiment Demo ===") + + let experiment = @src.summarized_experiment( + Map([ + ( + "counts", + [ + [10.0, 12.0, 20.0, 24.0], + [5.0, 7.0, 15.0, 17.0], + [30.0, 32.0, 18.0, 20.0], + [8.0, 10.0, 14.0, 16.0], + ], + ), + ]), + [("TaxonA", 0, 0), ("TaxonB", 0, 0), ("TaxonC", 0, 0), ("TaxonD", 0, 0)], + [ + Map([("sample", "Control1")]), + Map([("sample", "Control2")]), + Map([("sample", "Treated1")]), + Map([("sample", "Treated2")]), + ], + Map([("study", "microbiome-treatment")]), + ) + + let tse = @src.TreeSummarizedExperiment::new( + experiment~, + row_tree=Some(make_taxa_tree()), + row_node_labels=["TaxonA", "TaxonB", "TaxonC", "TaxonD"], + col_tree=Some(make_sample_tree()), + col_node_labels=["Control1", "Control2", "Treated1", "Treated2"], + reference_sequences=["ACGT", "AAGT", "CCGT", "TTGT"], + ) + + println("\n1. Container and links") + println(" Valid: \{tse.is_valid()}") + println(" Dimensions: \{tse.nrow()} rows x \{tse.ncol()} columns") + for link in tse.row_links() { + println( + " \{link.node_alias()} -> \{link.node_label()}, leaf=\{link.is_leaf()}", + ) + } + + println("\n2. Tree queries") + let descendants = @src.tse_find_descendants( + make_taxa_tree(), + "Firmicutes", + leaves_only=true, + ) + println(" Firmicutes descendants: " + descendants.join(", ")) + println( + " TaxonC ancestors: " + + @src.tse_find_ancestors(make_taxa_tree(), "TaxonC").join(" -> "), + ) + + println("\n3. Subset by an internal tree node") + match tse.subset_by_row_nodes(["Firmicutes"]) { + Some(subset) => + match subset.assay("counts") { + Some(data) => print_assay(data) + None => () + } + None => println(" Row tree is unavailable") + } + + println("\n4. Aggregate taxa to phylum level (sum)") + match tse.aggregate_rows(["Firmicutes", "Bacteroidota"]) { + Some(aggregated) => + match aggregated.assay("counts") { + Some(data) => print_assay(data) + None => () + } + None => println(" No matching row-tree nodes") + } + + println("\n5. Aggregate samples to treatment groups (mean)") + match + tse.aggregate_cols( + ["Controls", "Treated"], + aggregation=@src.tse_aggregation_mean(), + ) { + Some(aggregated) => + match aggregated.assay("counts") { + Some(data) => print_assay(data) + None => () + } + None => println(" No matching column-tree nodes") + } + + println("\n=== Demo Complete ===") +} diff --git a/examples/tree_summarized_experiment_demo/moon.pkg b/examples/tree_summarized_experiment_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/tree_summarized_experiment_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/unigene_demo/main.mbt b/examples/unigene_demo/main.mbt new file mode 100644 index 00000000..1971282c --- /dev/null +++ b/examples/unigene_demo/main.mbt @@ -0,0 +1,68 @@ +///| +fn main { + println("=== Bio.UniGene Demo ===") + + let record = @src.unigene_read(@src.unigene_sample_text()) catch { + _ => abort("failed to parse UniGene sample") + } + + println("\n1. Cluster") + println(" " + record.summary()) + println(" Species prefix: " + record.species) + println(" Cytoband: " + record.cytoband) + println(" Expression: " + record.expression.join(", ")) + println(" Restricted expression: " + record.restricted_expression.join(", ")) + + println("\n2. Protein similarities") + for similarity in record.protein_similarities { + println( + " " + + similarity.protein_id + + " (taxon " + + similarity.organism + + "): " + + similarity.percent + + "% over " + + similarity.alignment_length + + " aa", + ) + } + + println("\n3. Sequence entries") + println(" mRNA: " + record.sequences_of_type("mRNA").length().to_string()) + println(" EST: " + record.sequences_of_type("EST").length().to_string()) + for sequence in record.image_sequences() { + println( + " IMAGE clone " + + sequence.image_id + + ": " + + sequence.accession + + " (" + + sequence.read_end + + " end)", + ) + } + + println("\n4. Mapping information") + for transcript_map in record.transcript_maps { + println( + " Marker " + + transcript_map.marker + + " on panel " + + transcript_map.radiation_hybrid_panel, + ) + } + for site in record.sts { + println(" STS " + site.accession + ", UniSTS " + site.unists) + } + + println("\n5. Canonical serialization") + let serialized = record.to_string() + let reparsed = @src.unigene_read(serialized) catch { + _ => abort("failed to reparse serialized UniGene record") + } + println(" Serialized bytes: " + serialized.length().to_string()) + println(" Round trip preserved record: " + (record == reparsed).to_string()) + + println("\n=== Demo Complete ===") +} diff --git a/examples/unigene_demo/moon.pkg b/examples/unigene_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/unigene_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/variance_partition_demo/main.mbt b/examples/variance_partition_demo/main.mbt new file mode 100644 index 00000000..0fd31a6b --- /dev/null +++ b/examples/variance_partition_demo/main.mbt @@ -0,0 +1,186 @@ +///| +fn demo_samples() -> Array[String] { + let samples : Array[String] = [] + for subject in 0..<4 { + for replicate in 0..<4 { + samples.push( + "S" + (subject + 1).to_string() + "_R" + (replicate + 1).to_string(), + ) + } + } + samples +} + +///| +fn demo_subjects() -> Array[String] { + let subjects : Array[String] = [] + for subject in 0..<4 { + for _ in 0..<4 { + subjects.push("S" + (subject + 1).to_string()) + } + } + subjects +} + +///| +fn demo_treatments() -> Array[String] { + let treatments : Array[String] = [] + for _ in 0..<4 { + treatments.push("Control") + treatments.push("Treated") + treatments.push("Control") + treatments.push("Treated") + } + treatments +} + +///| +fn demo_subject_gene() -> Array[Double] { + let values : Array[Double] = [] + let noise = [0.08, 0.04, -0.08, -0.04] + for subject in 0..<4 { + let subject_effect = (subject.to_double() - 1.5) * 2.0 + for replicate in 0..<4 { + values.push(10.0 + subject_effect + noise[replicate]) + } + } + values +} + +///| +fn demo_treatment_gene() -> Array[Double] { + let values : Array[Double] = [] + let noise = [0.05, -0.03, -0.05, 0.03] + for subject in 0..<4 { + let subject_effect = (subject.to_double() - 1.5) * 0.08 + for replicate in 0..<4 { + let treatment_effect = if replicate % 2 == 1 { 4.0 } else { 0.0 } + values.push(6.0 + subject_effect + treatment_effect + noise[replicate]) + } + } + values +} + +///| +fn main { + println("=== Bioconductor variancePartition Demo ===") + let treatment = @src.vp_categorical_effect( + "Treatment", + demo_treatments(), + reference="Control", + ) catch { + _ => abort("failed to encode treatment") + } + let subject = @src.vp_random_effect("Subject", demo_subjects()) catch { + _ => abort("failed to encode subject") + } + let design = @src.vp_design(demo_samples(), [treatment], [subject]) catch { + _ => abort("failed to build mixed-model design") + } + let expression = [demo_subject_gene(), demo_treatment_gene()] + let weights = [ + Array::make(16, 1.0), + [ + 0.5, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, + 1.0, + ], + ] + + println("\n1. Typed fixed and random effects") + println(" coefficients: " + design.coefficient_names.to_string()) + println(" random effect: " + design.random_effects[0].name) + println(" levels: " + design.random_effects[0].levels.to_string()) + + println("\n2. ML variance fractions with observation-level weights") + let partition = @src.fit_extract_variance_partition( + expression, + design, + gene_names=["subject_gene", "treatment_gene"], + weights~, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("failed to partition variance") + } + println(partition.summary()) + for fit in partition.fits { + println(" " + fit.gene_name + ":") + for component in 0.. abort("failed to build treatment contrast") + } + let differential = @src.dream( + expression, + design, + contrast, + gene_names=["subject_gene", "treatment_gene"], + weights~, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("failed to run dream") + } + for gene in differential.genes { + println( + " " + + gene.gene_name + + ": estimate=" + + gene.estimate.to_string() + + ", t=" + + gene.statistic.to_string() + + ", df=" + + gene.degrees_of_freedom.to_string() + + ", FDR=" + + gene.adjusted_p_value.to_string(), + ) + } + + println("\n5. SummarizedExperiment assay integration") + let experiment = @src.summarized_experiment( + Map([("logcounts", expression), ("weights", weights)]), + [], + [], + Map([("study", "repeated-measures-demo")]), + ) + let assay_result = @src.dream_se( + experiment, + "logcounts", + design, + contrast, + gene_names=["subject_gene", "treatment_gene"], + weights_assay="weights", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("failed to run assay-backed dream") + } + println(" " + assay_result.summary()) + + println( + "\nScope: random intercepts and numerical Satterthwaite tests; random slopes, Kenward-Roger, voom/eBayes and plotting are not included.", + ) + println("\n=== Demo Complete ===") +} diff --git a/examples/variance_partition_demo/moon.pkg b/examples/variance_partition_demo/moon.pkg new file mode 100644 index 00000000..8824b1ac --- /dev/null +++ b/examples/variance_partition_demo/moon.pkg @@ -0,0 +1,6 @@ +import { + "IvanAXu/BioSeqs/src", + "moonbitlang/core/hashmap", +} + +pkgtype(kind: "executable") diff --git a/examples/voyager_demo/main.mbt b/examples/voyager_demo/main.mbt new file mode 100644 index 00000000..a2d79b3c --- /dev/null +++ b/examples/voyager_demo/main.mbt @@ -0,0 +1,266 @@ +// Bioconductor Voyager-inspired spatial autocorrelation workflow. +// +// Exercises the full univariate pipeline on a synthetic SpatialExperiment: +// 1. Spatial weight construction (kNN, distance-band, inverse-distance). +// 2. Global Moran's I and Geary's c with analytical inference. +// 3. Local Moran's I (LISA) with permutation inference and quadrants. +// 4. Local Getis–Ord Gi* hotspot detection. +// 5. Bivariate Lee's L spatial association. +// 6. Empirical variogram and spherical model fitting. +// 7. Moran correlogram over distance bins. +// 8. Immutable SpatialExperiment write-back of local statistics. + +///| +fn voyager_demo_coords(se : @src.SpatialExperiment) -> Array[Array[Double]] { + let coords : Array[Array[Double]] = [] + for sc in se.spatial_coords { + coords.push([sc.x, sc.y]) + } + coords +} + +///| +fn voyager_demo_round(value : Double, digits : Int) -> Double { + let scale = @math.pow(10.0, digits.to_double()) + (value * scale).round() / scale +} + +///| +fn voyager_demo_print_row(label : String, value : Double) -> Unit { + println(" " + label + ": " + voyager_demo_round(value, 4).to_string()) +} + +///| +fn main { + let se = @src.voyager_example_spatial_experiment() + let coords = voyager_demo_coords(se) + // assay[feature][spot]; the example ships 4 features on a 6x6 grid. + let assay = se.assay["logcounts"] + let gradient = assay[0] + let hotspot = assay[1] + let anti = assay[3] + println("=== Bioconductor Voyager Demo ===") + println( + "SpatialExperiment: " + + se.col_data.length().to_string() + + " spots, " + + assay.length().to_string() + + " features, platform=" + + se.metadata["platform"], + ) + + // ----------------------------------------------------------------------- + // 1. Spatial weight construction + // ----------------------------------------------------------------------- + println("\n1. Spatial weight construction") + let knn = @src.voyager_weights_knn(coords, 4) catch { + VoyagerError(message) => abort("kNN weights failed: " + message) + } + println( + " kNN k=4: style=" + + knn.style + + ", S0=" + + voyager_demo_round(@src.voyager_weights_s0(knn), 4).to_string(), + ) + let dband = @src.voyager_weights_distance_band(coords, 1.0) catch { + VoyagerError(message) => abort("distance-band weights failed: " + message) + } + println( + " distance-band 1.0: centre spot neighbours=" + + dband.neighbors[18].length().to_string(), + ) + let idw = @src.voyager_weights_inverse_distance(coords, 1.0) catch { + VoyagerError(message) => + abort("inverse-distance weights failed: " + message) + } + println( + " inverse-distance power=1: centre spot neighbours=" + + idw.neighbors[18].length().to_string(), + ) + + // ----------------------------------------------------------------------- + // 2. Global Moran's I and Geary's c + // ----------------------------------------------------------------------- + println("\n2. Global spatial autocorrelation (gene_gradient)") + let moran = @src.voyager_global_morans_i(gradient, knn) catch { + VoyagerError(message) => abort("global Moran's I failed: " + message) + } + voyager_demo_print_row("Moran's I estimate", moran.estimate) + voyager_demo_print_row("Moran's I expectation", moran.expectation) + voyager_demo_print_row("Moran's I z-score", moran.z_score) + voyager_demo_print_row("Moran's I p-value", moran.p_value) + let geary = @src.voyager_global_gearys_c(gradient, knn) catch { + VoyagerError(message) => abort("global Geary's c failed: " + message) + } + voyager_demo_print_row("Geary's c estimate", geary.estimate) + voyager_demo_print_row("Geary's c p-value", geary.p_value) + + // ----------------------------------------------------------------------- + // 3. Local Moran's I (LISA) with permutation inference + // ----------------------------------------------------------------------- + println("\n3. Local Moran's I (LISA) with permutation inference") + let lisa = @src.voyager_local_morans_i( + gradient, + knn, + permutations=199, + seed=20240501, + fdr_threshold=0.1, + ) catch { + VoyagerError(message) => abort("local Moran's I failed: " + message) + } + let mut hh = 0 + let mut ll = 0 + let mut sig = 0 + for i in 0.. abort("Getis–Ord failed: " + message) + } + let mut max_z = -1.0e300 + let mut max_spot = 0 + for i in 0.. max_z { + max_z = getis.z_scores[i] + max_spot = i + } + } + println( + " peak Gi* z-score=" + + voyager_demo_round(max_z, 4).to_string() + + " at spot " + + max_spot.to_string() + + " (" + + getis.classifications[max_spot] + + ")", + ) + + // ----------------------------------------------------------------------- + // 5. Bivariate Lee's L + // ----------------------------------------------------------------------- + println("\n5. Bivariate Lee's L (gene_gradient vs gene_anti)") + let lees = @src.voyager_lees_l( + gradient, + anti, + knn, + permutations=199, + seed=20240501, + fdr_threshold=0.1, + ) catch { + VoyagerError(message) => abort("Lee's L failed: " + message) + } + voyager_demo_print_row("global Lee's L", lees.global_l) + + // ----------------------------------------------------------------------- + // 6. Empirical variogram and spherical model fit + // ----------------------------------------------------------------------- + println("\n6. Empirical variogram and spherical model fit (gene_hotspot)") + let empirical = @src.voyager_empirical_variogram(coords, hotspot, n_lags=8) catch { + VoyagerError(message) => abort("empirical variogram failed: " + message) + } + println(" empirical points=" + empirical.length().to_string()) + let model = @src.voyager_fit_variogram(empirical, model_type="spherical") catch { + VoyagerError(message) => abort("variogram fit failed: " + message) + } + println( + " spherical: nugget=" + + voyager_demo_round(model.nugget, 4).to_string() + + ", sill=" + + voyager_demo_round(model.sill, 4).to_string() + + ", range=" + + voyager_demo_round(model.range, 4).to_string() + + ", SSE=" + + voyager_demo_round(model.fitted_sse, 4).to_string(), + ) + voyager_demo_print_row( + "variogram predicted semivariance at range", + @src.voyager_variogram_predict(model, model.range), + ) + + // ----------------------------------------------------------------------- + // 7. Moran correlogram + // ----------------------------------------------------------------------- + println("\n7. Moran correlogram over distance bins") + let corr = @src.voyager_correlogram(coords, gradient, n_lags=5) catch { + VoyagerError(message) => abort("correlogram failed: " + message) + } + println(" bins=" + corr.length().to_string()) + for point in corr { + println( + " lag=" + + voyager_demo_round(point.lag, 3).to_string() + + ", Moran's I=" + + voyager_demo_round(point.morans_i, 4).to_string() + + ", npairs=" + + point.npairs.to_string(), + ) + } + + // ----------------------------------------------------------------------- + // 8. Immutable SpatialExperiment write-back + // ----------------------------------------------------------------------- + println("\n8. Immutable SpatialExperiment write-back") + let output = @src.voyager_run_univariate_sfe( + se, + [0, 1], + assay_name="logcounts", + stat_method="moran", + permutations=99, + seed=2024, + fdr_threshold=0.1, + output_prefix="voyager", + ) catch { + VoyagerError(message) => + abort("Voyager SpatialExperiment integration failed: " + message) + } + println( + " features analyzed=" + + output.results.length().to_string() + + ", method=" + + output.results[0].stat_method, + ) + println( + " input unchanged: " + + (!se.col_data[0].contains("voyager.moran.local.gene_gradient")).to_string(), + ) + println( + " output colData carries local Moran: " + + output.experiment.col_data[0] + .contains("voyager.moran.local.gene_gradient") + .to_string(), + ) + println( + " metadata method=" + + output.experiment.metadata["voyager.method"] + + ", n_features=" + + output.experiment.metadata["voyager.n_features"], + ) + println("=== Demo Complete ===") +} diff --git a/examples/voyager_demo/moon.pkg b/examples/voyager_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/voyager_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/examples/zinbwave_demo/main.mbt b/examples/zinbwave_demo/main.mbt new file mode 100644 index 00000000..1e04a6c1 --- /dev/null +++ b/examples/zinbwave_demo/main.mbt @@ -0,0 +1,95 @@ +///| +fn main { + println("=== Bioconductor zinbwave Demo ===") + let (counts, genes, cells, batch) = @src.zinbwave_example_data() + let config = @src.ZinbWaveConfig::create( + n_factors=2, + max_iterations=35, + inner_iterations=2, + tolerance=1.0e-5, + ridge=0.01, + ) catch { + _ => abort("failed to create zinbwave configuration") + } + + println("\n1. Fit a ZINB latent-factor model with a batch covariate") + let model = @src.zinbwave_fit( + counts, + config~, + cell_covariates=batch, + gene_names=genes, + cell_names=cells, + ) catch { + _ => abort("failed to fit zinbwave model") + } + println(" " + model.summary()) + println( + " log-likelihood: " + + model.diagnostics.initial_log_likelihood.to_string() + + " -> " + + model.diagnostics.final_log_likelihood.to_string(), + ) + + println("\n2. Inspect the inferred cell representation") + for cell in 0.. + println( + " " + + summary.gene_name + + ": observed zero fraction=" + + summary.observed_zero_fraction.to_string() + + ", fitted=" + + summary.fitted_zero_fraction.to_string() + + ", dispersion=" + + summary.dispersion.to_string(), + ) + None => abort("dropout gene summary should exist") + } + println( + " structural-zero posterior weight at dropout_gene/A1: " + + (1.0 - model.observational_weights[5][0]).to_string(), + ) + + println("\n4. Generate normalized, imputed, and residual matrices") + let normalized = model.normalized_values() + let imputed = model.impute_zeros(counts) catch { + _ => abort("failed to impute model zeros") + } + println(" normalized A_marker_1/A1=" + normalized[0][0].to_string()) + println(" imputed dropout_gene/A1=" + imputed[5][0].to_string()) + println( + " deviance residual A_marker_1/B1=" + + model.deviance_residuals[0][4].to_string(), + ) + + println("\n5. Add outputs to an immutable SingleCellExperiment copy") + let sce = @src.SingleCellExperiment::new(counts, genes, cells) + let output = @src.zinbwave_sce(sce, config~, cell_covariates=batch) catch { + _ => abort("failed to integrate zinbwave with SingleCellExperiment") + } + println( + " reduced dimensions: " + + @src.sce_get_reduced_dim(output.experiment, "zinbwave").length().to_string() + + " cells", + ) + println( + " observation-weight assay: " + + @src.sce_get_assay(output.experiment, "zinbwave_weights") + .length() + .to_string() + + " genes", + ) + println( + " source object unchanged: " + + (@src.sce_get_assay(sce, "zinbwave_weights").length() == 0).to_string(), + ) + + println("\n=== Demo Complete ===") +} diff --git a/examples/zinbwave_demo/moon.pkg b/examples/zinbwave_demo/moon.pkg new file mode 100644 index 00000000..5363c0f5 --- /dev/null +++ b/examples/zinbwave_demo/moon.pkg @@ -0,0 +1,5 @@ +import { + "IvanAXu/BioSeqs/src", +} + +pkgtype(kind: "executable") diff --git a/moon.mod b/moon.mod index 1049a3c3..9c192bde 100644 --- a/moon.mod +++ b/moon.mod @@ -11,7 +11,7 @@ name = "IvanAXu/BioSeqs" -version = "0.1.6" +version = "0.1.8" readme = "README.mbt.md" diff --git a/src/a2m.mbt b/src/a2m.mbt new file mode 100644 index 00000000..2876f818 --- /dev/null +++ b/src/a2m.mbt @@ -0,0 +1,930 @@ +// Biopython-compatible A2M multiple sequence alignment support. +// +// A2M uses upper-case residues and '-' in model match columns, and lower-case +// residues and '.' in insertion columns. The first row defines the state of +// every column. Internally rows are normalized to upper-case residues and '-'; +// the separate state string preserves the information required for round-trip +// serialization. + +///| +/// Error raised for malformed A2M data or invalid alignment operations. +pub suberror A2mError { + A2mError(String) +} + +///| +/// State assigned to an A2M alignment column. +pub(all) enum A2mColumnState { + A2mMatch + A2mInsertion +} derive(Eq, Debug) + +///| +/// One normalized sequence row in an A2M alignment. +/// +/// `aligned_sequence` contains upper-case residues and '-' gaps. `sequence` +/// contains the ungapped residues. +pub struct A2mSequence { + id : String + description : String + sequence : String + aligned_sequence : String +} derive(Eq, Debug) + +///| +/// A state-aware A2M multiple sequence alignment. +/// +/// `states` contains one `D` (model match/deletion) or `I` (insertion) for +/// each alignment column, following Biopython's column annotation convention. +pub struct A2mAlignment { + sequences : Array[A2mSequence] + states : String +} derive(Eq, Debug) + +///| +/// A contiguous run of insertion-state columns. +/// +/// `slot` is the number of match columns preceding the run. Therefore slot 0 +/// is before the first match column and slot `match_columns()` is after the +/// final match column. +pub struct A2mInsertionRun { + slot : Int + start_column : Int + end_column : Int + width : Int +} derive(Eq, Debug) + +///| +/// Pairwise statistics for two rows in an A2M alignment. +pub struct A2mPairCounts { + columns : Int + aligned : Int + identities : Int + mismatches : Int + gap_columns : Int + double_gap_columns : Int + match_aligned : Int + insertion_aligned : Int +} derive(Eq, Debug) + +///| +/// Construct a normalized A2M sequence row. +pub fn A2mSequence::create( + id : String, + aligned_sequence : String, + description? : String = "", +) -> A2mSequence raise A2mError { + a2m_validate_id(id) + a2m_validate_description(description) + if aligned_sequence.length() == 0 { + raise A2mError("A2M aligned sequence must not be empty") + } + let normalized = a2m_normalize_aligned_sequence(aligned_sequence) + A2mSequence::{ + id, + description, + sequence: a2m_remove_gaps(normalized), + aligned_sequence: normalized, + } +} + +///| +/// Construct an A2M alignment from normalized rows and a `D`/`I` state string. +pub fn A2mAlignment::create( + sequences : Array[A2mSequence], + states : String, +) -> A2mAlignment raise A2mError { + let copied : Array[A2mSequence] = [] + for sequence in sequences { + copied.push(sequence) + } + let alignment = A2mAlignment::{ sequences: copied, states } + a2m_validate_alignment(alignment) + alignment +} + +///| +/// Construct an A2M alignment directly from row metadata and normalized rows. +pub fn a2m_from_aligned( + ids : Array[String], + aligned_sequences : Array[String], + states : String, + descriptions? : Array[String] = [], +) -> A2mAlignment raise A2mError { + if ids.length() == 0 { + raise A2mError("A2M alignment must contain at least one sequence") + } + if ids.length() != aligned_sequences.length() { + raise A2mError( + "A2M identifiers and aligned sequences must have equal lengths", + ) + } + if descriptions.length() != 0 && descriptions.length() != ids.length() { + raise A2mError("A2M descriptions must be empty or match the sequence count") + } + let sequences : Array[A2mSequence] = [] + for index = 0; index < ids.length(); index = index + 1 { + let description = if descriptions.length() == 0 { + "" + } else { + descriptions[index] + } + sequences.push( + A2mSequence::create(ids[index], aligned_sequences[index], description~), + ) + } + A2mAlignment::create(sequences, states) +} + +///| +/// Parse one A2M multiple sequence alignment. +/// +/// The first row defines the state of each column. All following rows must use +/// the same character class in each column. Wrapped sequence lines and CRLF +/// input are accepted; blank lines are ignored. +pub fn a2m_parse(text : String) -> A2mAlignment raise A2mError { + let lines = a2m_normalize_lines(text) + let ids : Array[String] = [] + let descriptions : Array[String] = [] + let encoded_rows : Array[String] = [] + let mut current = -1 + for raw_line in lines { + let line = raw_line.trim().to_owned() + if line.length() == 0 { + continue + } + if line.unsafe_get(0).to_int() == '>'.to_int() { + let header = line[1:].trim().to_owned() + if header.length() == 0 { + raise A2mError("A2M header must contain a sequence identifier") + } + let (id, description) = a2m_parse_header(header) + a2m_validate_id(id) + a2m_validate_description(description) + ids.push(id) + descriptions.push(description) + encoded_rows.push("") + current = encoded_rows.length() - 1 + } else { + if current < 0 { + raise A2mError("A2M sequence data appears before the first header") + } + a2m_validate_encoded_fragment(line) + encoded_rows[current] = encoded_rows[current] + line + } + } + if encoded_rows.length() == 0 { + raise A2mError("Empty A2M input") + } + let width = encoded_rows[0].length() + if width == 0 { + raise A2mError("A2M sequences must not be empty") + } + let states_builder = StringBuilder::new(size_hint=width) + for column = 0; column < width; column = column + 1 { + let code = encoded_rows[0].unsafe_get(column).to_int() + if a2m_is_upper(code) || code == '-'.to_int() { + states_builder.write_char('D') + } else if a2m_is_lower(code) || code == '.'.to_int() { + states_builder.write_char('I') + } else { + raise A2mError( + "Invalid A2M character in first sequence at column " + + column.to_string(), + ) + } + } + let states = states_builder.to_string() + let sequences : Array[A2mSequence] = [] + for row = 0; row < encoded_rows.length(); row = row + 1 { + let encoded = encoded_rows[row] + if encoded.length() != width { + raise A2mError( + "A2M row '" + + ids[row] + + "' has width " + + encoded.length().to_string() + + "; expected " + + width.to_string(), + ) + } + let normalized = StringBuilder::new(size_hint=width) + for column = 0; column < width; column = column + 1 { + let code = encoded.unsafe_get(column).to_int() + let state = states.unsafe_get(column).to_int() + if state == 'D'.to_int() { + if code == '-'.to_int() { + normalized.write_char('-') + } else if a2m_is_upper(code) { + normalized.write_char(code.unsafe_to_char()) + } else { + raise A2mError( + "A2M row '" + + ids[row] + + "' uses an insertion character in match column " + + column.to_string(), + ) + } + } else if code == '.'.to_int() { + normalized.write_char('-') + } else if a2m_is_lower(code) { + normalized.write_char(a2m_upper_code(code).unsafe_to_char()) + } else { + raise A2mError( + "A2M row '" + + ids[row] + + "' uses a match character in insertion column " + + column.to_string(), + ) + } + } + sequences.push(A2mSequence::{ + id: ids[row], + description: descriptions[row], + sequence: a2m_remove_gaps(normalized.to_string()), + aligned_sequence: normalized.to_string(), + }) + } + A2mAlignment::create(sequences, states) +} + +///| +/// Serialize one A2M alignment. +/// +/// A `line_width` of zero emits one sequence line per row. Positive values +/// wrap rows without changing column states. +pub fn a2m_write( + alignment : A2mAlignment, + line_width? : Int = 0, +) -> String raise A2mError { + a2m_validate_alignment(alignment) + if line_width < 0 { + raise A2mError("A2M line width must be non-negative") + } + let output = StringBuilder::new() + for row = 0; row < alignment.sequences.length(); row = row + 1 { + let sequence = alignment.sequences[row] + output.write_char('>') + output.write_string(sequence.id) + if sequence.description.length() > 0 { + output.write_char(' ') + output.write_string(sequence.description) + } + output.write_char('\n') + let encoded = alignment.encoded_sequence(row).unwrap() + if line_width == 0 { + output.write_string(encoded) + output.write_char('\n') + } else { + let mut start = 0 + while start < encoded.length() { + let end = if start + line_width < encoded.length() { + start + line_width + } else { + encoded.length() + } + output.write_string(encoded[start:end].to_owned()) + output.write_char('\n') + start = end + } + } + } + output.to_string() +} + +///| +/// Return the number of rows. +pub fn A2mAlignment::num_sequences(self : A2mAlignment) -> Int { + self.sequences.length() +} + +///| +/// Return the number of alignment columns. +pub fn A2mAlignment::alignment_length(self : A2mAlignment) -> Int { + self.states.length() +} + +///| +/// Return the number of model match/deletion columns. +pub fn A2mAlignment::match_columns(self : A2mAlignment) -> Int { + let mut count = 0 + for index = 0; index < self.states.length(); index = index + 1 { + if self.states.unsafe_get(index).to_int() == 'D'.to_int() { + count = count + 1 + } + } + count +} + +///| +/// Return the number of insertion-state columns. +pub fn A2mAlignment::insertion_columns(self : A2mAlignment) -> Int { + self.alignment_length() - self.match_columns() +} + +///| +/// Return a column state by zero-based alignment coordinate. +pub fn A2mAlignment::column_state( + self : A2mAlignment, + column : Int, +) -> A2mColumnState? { + if column < 0 || column >= self.states.length() { + None + } else if self.states.unsafe_get(column).to_int() == 'D'.to_int() { + Some(A2mMatch) + } else { + Some(A2mInsertion) + } +} + +///| +/// Return all row characters at an alignment column. +pub fn A2mAlignment::column(self : A2mAlignment, column : Int) -> String? { + if column < 0 || column >= self.states.length() { + return None + } + let result = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + result.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(result.to_string()) +} + +///| +/// Find the first row with an exact sequence identifier. +pub fn A2mAlignment::find_sequence(self : A2mAlignment, id : String) -> Int? { + for index = 0; index < self.sequences.length(); index = index + 1 { + if self.sequences[index].id == id { + return Some(index) + } + } + None +} + +///| +/// Return a row in canonical A2M character form. +pub fn A2mAlignment::encoded_sequence( + self : A2mAlignment, + row : Int, +) -> String? { + if row < 0 || row >= self.sequences.length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + let result = StringBuilder::new(size_hint=self.states.length()) + for column = 0; column < self.states.length(); column = column + 1 { + let code = aligned.unsafe_get(column).to_int() + if self.states.unsafe_get(column).to_int() == 'D'.to_int() { + result.write_char(code.unsafe_to_char()) + } else if code == '-'.to_int() { + result.write_char('.') + } else { + result.write_char(a2m_lower_code(code).unsafe_to_char()) + } + } + Some(result.to_string()) +} + +///| +/// Map an ungapped row coordinate to its alignment column. +pub fn A2mAlignment::sequence_position_to_column( + self : A2mAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || row >= self.sequences.length() || position < 0 { + return None + } + let aligned = self.sequences[row].aligned_sequence + let mut coordinate = 0 + for column = 0; column < aligned.length(); column = column + 1 { + if aligned.unsafe_get(column).to_int() != '-'.to_int() { + if coordinate == position { + return Some(column) + } + coordinate = coordinate + 1 + } + } + None +} + +///| +/// Map an alignment column to an ungapped row coordinate. +/// +/// Gap columns return `None`. +pub fn A2mAlignment::column_to_sequence_position( + self : A2mAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.states.length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + if aligned.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut coordinate = 0 + for index = 0; index < column; index = index + 1 { + if aligned.unsafe_get(index).to_int() != '-'.to_int() { + coordinate = coordinate + 1 + } + } + Some(coordinate) +} + +///| +/// Map a residue coordinate from one row through the alignment to another. +/// +/// Returns `None` for out-of-range coordinates and target gaps. +pub fn A2mAlignment::map_position( + self : A2mAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column coordinates for two rows. +/// +/// Residues are represented by zero-based coordinates and gaps by `None`. +pub fn A2mAlignment::aligned_pairs( + self : A2mAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise A2mError { + a2m_validate_row(self, first_row) + a2m_validate_row(self, second_row) + let first = self.sequences[first_row].aligned_sequence + let second = self.sequences[second_row].aligned_sequence + let result : Array[(Int?, Int?)] = [] + let mut first_coordinate = 0 + let mut second_coordinate = 0 + for column = 0; column < self.states.length(); column = column + 1 { + let first_gap = first.unsafe_get(column).to_int() == '-'.to_int() + let second_gap = second.unsafe_get(column).to_int() == '-'.to_int() + let first_value : Int? = if first_gap { + None + } else { + Some(first_coordinate) + } + let second_value : Int? = if second_gap { + None + } else { + Some(second_coordinate) + } + result.push((first_value, second_value)) + if !first_gap { + first_coordinate = first_coordinate + 1 + } + if !second_gap { + second_coordinate = second_coordinate + 1 + } + } + result +} + +///| +/// Compute pairwise identity, mismatch, and gap counts for two rows. +pub fn A2mAlignment::pair_counts( + self : A2mAlignment, + first_row : Int, + second_row : Int, +) -> A2mPairCounts raise A2mError { + a2m_validate_row(self, first_row) + a2m_validate_row(self, second_row) + let first = self.sequences[first_row].aligned_sequence + let second = self.sequences[second_row].aligned_sequence + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut match_aligned = 0 + let mut insertion_aligned = 0 + for column = 0; column < self.states.length(); column = column + 1 { + let first_code = first.unsafe_get(column).to_int() + let second_code = second.unsafe_get(column).to_int() + let first_gap = first_code == '-'.to_int() + let second_gap = second_code == '-'.to_int() + if first_gap && second_gap { + double_gap_columns = double_gap_columns + 1 + } else if first_gap || second_gap { + gap_columns = gap_columns + 1 + } else { + aligned = aligned + 1 + if first_code == second_code { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + if self.states.unsafe_get(column).to_int() == 'D'.to_int() { + match_aligned = match_aligned + 1 + } else { + insertion_aligned = insertion_aligned + 1 + } + } + } + A2mPairCounts::{ + columns: self.states.length(), + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + match_aligned, + insertion_aligned, + } +} + +///| +/// Return identity fraction over columns containing residues in both rows. +pub fn A2mPairCounts::identity(self : A2mPairCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Locate contiguous insertion-state runs and their model slots. +pub fn A2mAlignment::insertion_runs( + self : A2mAlignment, +) -> Array[A2mInsertionRun] { + let result : Array[A2mInsertionRun] = [] + let mut column = 0 + let mut slot = 0 + while column < self.states.length() { + if self.states.unsafe_get(column).to_int() == 'D'.to_int() { + slot = slot + 1 + column = column + 1 + } else { + let start = column + while column < self.states.length() && + self.states.unsafe_get(column).to_int() == 'I'.to_int() { + column = column + 1 + } + result.push(A2mInsertionRun::{ + slot, + start_column: start, + end_column: column, + width: column - start, + }) + } + } + result +} + +///| +/// Calculate a majority consensus. +/// +/// Gaps do not vote. Columns without residues produce `-`; columns below +/// `minimum_fraction` produce `X`. Setting `include_insertions` to false +/// returns a model match-state consensus. +pub fn A2mAlignment::consensus( + self : A2mAlignment, + include_insertions? : Bool = true, + minimum_fraction? : Double = 0.0, +) -> String raise A2mError { + if minimum_fraction != minimum_fraction || + minimum_fraction < 0.0 || + minimum_fraction > 1.0 { + raise A2mError("A2M consensus minimum fraction must be between 0 and 1") + } + let result = StringBuilder::new() + for column = 0; column < self.states.length(); column = column + 1 { + if !include_insertions && + self.states.unsafe_get(column).to_int() == 'I'.to_int() { + continue + } + let counts = Array::make(26, 0) + let mut residues = 0 + for sequence in self.sequences { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if a2m_is_upper(code) { + let index = code - 'A'.to_int() + counts[index] = counts[index] + 1 + residues = residues + 1 + } + } + if residues == 0 { + result.write_char('-') + continue + } + let mut best_index = 0 + let mut best_count = counts[0] + for index = 1; index < counts.length(); index = index + 1 { + if counts[index] > best_count { + best_index = index + best_count = counts[index] + } + } + let fraction = best_count.to_double() / residues.to_double() + if fraction < minimum_fraction { + result.write_char('X') + } else { + result.write_char((best_index + 'A'.to_int()).unsafe_to_char()) + } + } + result.to_string() +} + +///| +/// Return per-column residue occupancy as a fraction of row count. +pub fn A2mAlignment::occupancy(self : A2mAlignment) -> Array[Double] { + let result : Array[Double] = [] + for column = 0; column < self.states.length(); column = column + 1 { + let mut residues = 0 + for sequence in self.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + result.push(residues.to_double() / self.sequences.length().to_double()) + } + result +} + +///| +/// Return a new alignment containing only model match/deletion columns. +pub fn A2mAlignment::match_projection( + self : A2mAlignment, +) -> A2mAlignment raise A2mError { + let match_count = self.match_columns() + if match_count == 0 { + raise A2mError("A2M alignment has no match columns to project") + } + let states = StringBuilder::new(size_hint=match_count) + let mut emitted = 0 + while emitted < match_count { + states.write_char('D') + emitted = emitted + 1 + } + let sequences : Array[A2mSequence] = [] + for sequence in self.sequences { + let row = StringBuilder::new(size_hint=match_count) + for column = 0; column < self.states.length(); column = column + 1 { + if self.states.unsafe_get(column).to_int() == 'D'.to_int() { + row.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + } + sequences.push( + A2mSequence::create( + sequence.id, + row.to_string(), + description=sequence.description, + ), + ) + } + A2mAlignment::create(sequences, states.to_string()) +} + +///| +/// Extract a non-empty half-open range of alignment columns. +pub fn A2mAlignment::slice_columns( + self : A2mAlignment, + start : Int, + end : Int, +) -> A2mAlignment raise A2mError { + if start < 0 || end <= start || end > self.states.length() { + raise A2mError("A2M column slice is out of bounds or empty") + } + let sequences : Array[A2mSequence] = [] + for sequence in self.sequences { + sequences.push( + A2mSequence::create( + sequence.id, + sequence.aligned_sequence[start:end].to_owned(), + description=sequence.description, + ), + ) + } + A2mAlignment::create(sequences, self.states[start:end].to_owned()) +} + +///| +/// Return a compact alignment summary. +pub fn A2mAlignment::summary(self : A2mAlignment) -> String { + "A2mAlignment(sequences=" + + self.sequences.length().to_string() + + ", columns=" + + self.states.length().to_string() + + ", match=" + + self.match_columns().to_string() + + ", insertion=" + + self.insertion_columns().to_string() + + ")" +} + +///| +/// Return a small state-aware A2M example. +pub fn a2m_example_text() -> String { + ">reference SAM model reference\n" + + "ACDefG-HIkLM\n" + + ">query_one insertion and deletion\n" + + "ACD..GTHI.LM\n" + + ">query_two divergent homolog\n" + + "ATDqrG-HI.L-\n" +} + +///| +fn a2m_validate_alignment(alignment : A2mAlignment) -> Unit raise A2mError { + if alignment.sequences.length() == 0 { + raise A2mError("A2M alignment must contain at least one sequence") + } + let width = alignment.states.length() + if width == 0 { + raise A2mError("A2M state string must not be empty") + } + for column = 0; column < width; column = column + 1 { + let state = alignment.states.unsafe_get(column).to_int() + if state != 'D'.to_int() && state != 'I'.to_int() { + raise A2mError( + "A2M state string contains a character other than D or I at column " + + column.to_string(), + ) + } + } + for row = 0; row < alignment.sequences.length(); row = row + 1 { + let sequence = alignment.sequences[row] + a2m_validate_id(sequence.id) + a2m_validate_description(sequence.description) + if sequence.aligned_sequence.length() != width { + raise A2mError( + "A2M row '" + + sequence.id + + "' has width " + + sequence.aligned_sequence.length().to_string() + + "; expected " + + width.to_string(), + ) + } + for column = 0; column < width; column = column + 1 { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if code != '-'.to_int() && !a2m_is_upper(code) { + raise A2mError( + "Normalized A2M row '" + + sequence.id + + "' contains an invalid character at column " + + column.to_string(), + ) + } + } + if a2m_remove_gaps(sequence.aligned_sequence) != sequence.sequence { + raise A2mError( + "A2M row '" + sequence.id + "' has inconsistent ungapped sequence", + ) + } + } +} + +///| +fn a2m_validate_row(alignment : A2mAlignment, row : Int) -> Unit raise A2mError { + if row < 0 || row >= alignment.sequences.length() { + raise A2mError("A2M row index is out of bounds") + } +} + +///| +fn a2m_validate_id(id : String) -> Unit raise A2mError { + if id.length() == 0 { + raise A2mError("A2M sequence identifier must not be empty") + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if code == ' '.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() { + raise A2mError("A2M sequence identifier must not contain whitespace") + } + } +} + +///| +fn a2m_validate_description(description : String) -> Unit raise A2mError { + for index = 0; index < description.length(); index = index + 1 { + let code = description.unsafe_get(index).to_int() + if code == '\n'.to_int() || code == '\r'.to_int() { + raise A2mError("A2M description must not contain line breaks") + } + } +} + +///| +fn a2m_validate_encoded_fragment(fragment : String) -> Unit raise A2mError { + for index = 0; index < fragment.length(); index = index + 1 { + let code = fragment.unsafe_get(index).to_int() + if !a2m_is_upper(code) && + !a2m_is_lower(code) && + code != '-'.to_int() && + code != '.'.to_int() { + raise A2mError( + "Invalid A2M sequence character at wrapped-line offset " + + index.to_string(), + ) + } + } +} + +///| +fn a2m_normalize_aligned_sequence(sequence : String) -> String raise A2mError { + let result = StringBuilder::new(size_hint=sequence.length()) + for index = 0; index < sequence.length(); index = index + 1 { + let code = sequence.unsafe_get(index).to_int() + if code == '-'.to_int() { + result.write_char('-') + } else if a2m_is_upper(code) { + result.write_char(code.unsafe_to_char()) + } else if a2m_is_lower(code) { + result.write_char(a2m_upper_code(code).unsafe_to_char()) + } else { + raise A2mError( + "Normalized A2M sequence contains an invalid character at column " + + index.to_string(), + ) + } + } + result.to_string() +} + +///| +fn a2m_remove_gaps(sequence : String) -> String { + let result = StringBuilder::new(size_hint=sequence.length()) + for index = 0; index < sequence.length(); index = index + 1 { + let code = sequence.unsafe_get(index).to_int() + if code != '-'.to_int() { + result.write_char(code.unsafe_to_char()) + } + } + result.to_string() +} + +///| +fn a2m_parse_header(header : String) -> (String, String) { + for index = 0; index < header.length(); index = index + 1 { + let code = header.unsafe_get(index).to_int() + if code == ' '.to_int() || code == '\t'.to_int() { + return (header[0:index].to_owned(), header[index:].trim().to_owned()) + } + } + (header, "") +} + +///| +fn a2m_normalize_lines(text : String) -> Array[String] { + let result : Array[String] = [] + for view in text.split("\n") { + let mut line = view.to_owned() + if line.length() > 0 && + line.unsafe_get(line.length() - 1).to_int() == '\r'.to_int() { + line = line[0:line.length() - 1].to_owned() + } + result.push(line) + } + result +} + +///| +fn a2m_is_upper(code : Int) -> Bool { + code >= 'A'.to_int() && code <= 'Z'.to_int() +} + +///| +fn a2m_is_lower(code : Int) -> Bool { + code >= 'a'.to_int() && code <= 'z'.to_int() +} + +///| +fn a2m_upper_code(code : Int) -> Int { + if a2m_is_lower(code) { + code - ('a'.to_int() - 'A'.to_int()) + } else { + code + } +} + +///| +fn a2m_lower_code(code : Int) -> Int { + if a2m_is_upper(code) { + code + ('a'.to_int() - 'A'.to_int()) + } else { + code + } +} diff --git a/src/abi.mbt b/src/abi.mbt index 2ca30206..0addc721 100644 --- a/src/abi.mbt +++ b/src/abi.mbt @@ -245,15 +245,27 @@ pub fn AbiTrace::quality_at(self : AbiTrace, index : Int) -> Int { ///| /// Format an ABIF date as "YYYY-MM-DD". pub fn AbiDate::to_string(self : AbiDate) -> String { - let m = if self.month < 10 { "0" + self.month.to_string() } else { self.month.to_string() } - let d = if self.day < 10 { "0" + self.day.to_string() } else { self.day.to_string() } + let m = if self.month < 10 { + "0" + self.month.to_string() + } else { + self.month.to_string() + } + let d = if self.day < 10 { + "0" + self.day.to_string() + } else { + self.day.to_string() + } self.year.to_string() + "-" + m + "-" + d } ///| /// Format an ABIF time as "HH:MM:SS". pub fn AbiTime::to_string(self : AbiTime) -> String { - let h = if self.hours < 10 { "0" + self.hours.to_string() } else { self.hours.to_string() } + let h = if self.hours < 10 { + "0" + self.hours.to_string() + } else { + self.hours.to_string() + } let mi = if self.minutes < 10 { "0" + self.minutes.to_string() } else { @@ -275,7 +287,10 @@ fn abi_read_u32(bytes : Array[Int], pos : Int) -> Int { if pos + 3 >= bytes.length() { return 0 } - bytes[pos] * 16777216 + bytes[pos + 1] * 65536 + bytes[pos + 2] * 256 + bytes[pos + 3] + bytes[pos] * 16777216 + + bytes[pos + 1] * 65536 + + bytes[pos + 2] * 256 + + bytes[pos + 3] } ///| @@ -428,10 +443,7 @@ fn abi_find_entry( ///| /// Read raw data bytes for a directory entry from the file bytes. -fn abi_read_entry_data( - bytes : Array[Int], - entry : AbiDirEntry, -) -> Array[Int] { +fn abi_read_entry_data(bytes : Array[Int], entry : AbiDirEntry) -> Array[Int] { let data : Array[Int] = Array::new() // If data fits inline (data_size <= 4), use inline_data if entry.data_size <= 4 { @@ -459,10 +471,7 @@ fn abi_read_entry_data( ///| /// Read trace data (array of Int) for a directory entry. /// Trace data is typically 2-byte (word) values. -fn abi_read_trace_data( - bytes : Array[Int], - entry : AbiDirEntry, -) -> Array[Int] { +fn abi_read_trace_data(bytes : Array[Int], entry : AbiDirEntry) -> Array[Int] { let data : Array[Int] = Array::new() let raw = abi_read_entry_data(bytes, entry) // If element type is word (3) or short (4), read as 2-byte values @@ -492,10 +501,7 @@ fn abi_read_trace_data( ///| /// Read integer array data for a directory entry (e.g., base positions). -fn abi_read_int_array( - bytes : Array[Int], - entry : AbiDirEntry, -) -> Array[Int] { +fn abi_read_int_array(bytes : Array[Int], entry : AbiDirEntry) -> Array[Int] { let data : Array[Int] = Array::new() let raw = abi_read_entry_data(bytes, entry) // For PLOC (base positions), element type is short (4), size 2 diff --git a/src/ace.mbt b/src/ace.mbt index 2f0e716d..7cb075c8 100644 --- a/src/ace.mbt +++ b/src/ace.mbt @@ -260,9 +260,7 @@ pub fn ace_parse(content : String) -> AceData { reads_map[r.read_id] = r i = ace_skip_read_block(lines, i) } - None => { - i = i + 1 - } + None => i = i + 1 } continue } @@ -274,9 +272,7 @@ pub fn ace_parse(content : String) -> AceData { contigs.push(contig) i = new_idx } - None => { - i = i + 1 - } + None => i = i + 1 } continue } @@ -301,7 +297,7 @@ pub fn ace_parse_reads(content : String) -> Array[AceRead] { let read = ace_parse_read_block(lines, i) match read { Some(r) => reads.push(r) - None => { () } + None => () } i = ace_skip_read_block(lines, i) } else { @@ -330,9 +326,7 @@ pub fn ace_parse_contigs(content : String) -> Array[AceContig] { reads_map[r.read_id] = r i = ace_skip_read_block(lines, i) } - None => { - i = i + 1 - } + None => i = i + 1 } } else if trimmed.has_prefix("CT ") { let contig_and_new_idx = ace_parse_contig_block(lines, i, reads_map) @@ -341,9 +335,7 @@ pub fn ace_parse_contigs(content : String) -> Array[AceContig] { contigs.push(contig) i = new_idx } - None => { - i = i + 1 - } + None => i = i + 1 } } else { i = i + 1 @@ -354,10 +346,7 @@ pub fn ace_parse_contigs(content : String) -> Array[AceContig] { ///| /// Parse a single RD block starting at the given line index. -fn ace_parse_read_block( - lines : Array[String], - start_idx : Int, -) -> AceRead? { +fn ace_parse_read_block(lines : Array[String], start_idx : Int) -> AceRead? { if start_idx >= lines.length() { return None } @@ -394,7 +383,8 @@ fn ace_parse_read_block( } if (first_char >= 'A'.to_int() && first_char <= 'Z'.to_int()) || (first_char >= 'a'.to_int() && first_char <= 'z'.to_int()) || - first_char == '-'.to_int() || first_char == '.'.to_int() { + first_char == '-'.to_int() || + first_char == '.'.to_int() { let remaining = read_len - sequence.length() let chunk = if trimmed.length() > remaining { trimmed[0:remaining].to_owned() @@ -412,17 +402,11 @@ fn ace_parse_read_block( let (quality, new_idx) = ace_parse_quality_lines(lines, idx, read_len) idx = new_idx - Some(AceRead::new( - read_id, - sequence, - quality, - 1, - read_len, - "+", - "", - chemistry, - dye, - )) + Some( + AceRead::new( + read_id, sequence, quality, 1, read_len, "+", "", chemistry, dye, + ), + ) } ///| @@ -454,12 +438,13 @@ fn ace_skip_read_block(lines : Array[String], start_idx : Int) -> Int { idx = idx + 1 break } - if (first_char >= '0'.to_int() && first_char <= '9'.to_int()) { + if first_char >= '0'.to_int() && first_char <= '9'.to_int() { break } if (first_char >= 'A'.to_int() && first_char <= 'Z'.to_int()) || (first_char >= 'a'.to_int() && first_char <= 'z'.to_int()) || - first_char == '-'.to_int() || first_char == '.'.to_int() { + first_char == '-'.to_int() || + first_char == '.'.to_int() { seq_chars = seq_chars + trimmed.length() idx = idx + 1 } else { @@ -528,7 +513,8 @@ fn ace_parse_contig_block( } if (first_char >= 'A'.to_int() && first_char <= 'Z'.to_int()) || (first_char >= 'a'.to_int() && first_char <= 'z'.to_int()) || - first_char == '-'.to_int() || first_char == '.'.to_int() { + first_char == '-'.to_int() || + first_char == '.'.to_int() { let remaining = contig_len - sequence.length() let chunk = if trimmed.length() > remaining { trimmed[0:remaining].to_owned() @@ -543,12 +529,17 @@ fn ace_parse_contig_block( } // Parse base qualities - let (base_qualities, new_idx) = ace_parse_quality_lines(lines, idx, contig_len) + let (base_qualities, new_idx) = ace_parse_quality_lines( + lines, idx, contig_len, + ) idx = new_idx // Parse AF alignment lines let reads : Array[AceRead] = Array::new() - let alignment_positions : Map[String, (Int, Int, Int, Int, Int, Int, String)] = Map([], capacity=32) + let alignment_positions : Map[String, (Int, Int, Int, Int, Int, Int, String)] = Map( + [], + capacity=32, + ) while idx < lines.length() { let line = lines[idx].to_string() @@ -571,13 +562,7 @@ fn ace_parse_contig_block( let strand = if af_tokens.length() >= 9 { af_tokens[8] } else { "+" } alignment_positions[read_name] = ( - contig_start, - contig_end, - read_start, - read_end, - qual_start, - qual_end, - strand, + contig_start, contig_end, read_start, read_end, qual_start, qual_end, strand, ) } idx = idx + 1 @@ -640,17 +625,12 @@ fn ace_parse_contig_block( ri = ri + 1 } - Some(( - AceContig::new( - contig_name, - sequence, - reads, - base_qualities, - "", - false, + Some( + ( + AceContig::new(contig_name, sequence, reads, base_qualities, "", false), + idx, ), - idx, - )) + ) } // ===== Serialization ===== @@ -784,7 +764,11 @@ pub fn ace_read_coverage(contig : AceContig) -> Map[Int, Int] { let read = contig.reads[ri] let start = if read.clip_start > 0 { read.clip_start } else { 1 } let end = if read.clip_end > 0 { - if read.clip_end > contig.length { contig.length } else { read.clip_end } + if read.clip_end > contig.length { + contig.length + } else { + read.clip_end + } } else { contig.length } @@ -839,19 +823,20 @@ pub fn ace_contig_gc_content(contig : AceContig) -> Double { /// Create sample ACE content for testing. pub fn sample_ace_content() -> String { "AF contig1 60\n" + - "RD read1 60 chemistry1 dye1\n" + - "ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG\n" + - "q 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30\n" + - "RD read2 60 chemistry2 dye2\n" + - "GCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTA\n" + - "q 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25\n" + - "CT contig1 60 0 60\n" + - "ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG\n" + - "q 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40\n" + - "AF read1 1 60 1 60 1 60 +\n" + - "AF read2 1 60 1 60 1 60 -\n" + "RD read1 60 chemistry1 dye1\n" + + "ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG\n" + + "q 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30 30\n" + + "RD read2 60 chemistry2 dye2\n" + + "GCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTAGCTA\n" + + "q 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25 25\n" + + "CT contig1 60 0 60\n" + + "ATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCGATCG\n" + + "q 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40 40\n" + + "AF read1 1 60 1 60 1 60 +\n" + + "AF read2 1 60 1 60 1 60 -\n" } +///| pub fn sample_ace_data() -> AceData { ace_parse(sample_ace_content()) -} \ No newline at end of file +} diff --git a/src/aldex2.mbt b/src/aldex2.mbt new file mode 100644 index 00000000..c95c961e --- /dev/null +++ b/src/aldex2.mbt @@ -0,0 +1,1538 @@ +// Compositional differential-abundance inference inspired by Bioconductor +// ALDEx2. Count matrices use feature x sample orientation throughout. + +///| +pub suberror Aldex2Error { + Aldex2Error(String) +} + +///| +pub(all) enum Aldex2Denominator { + Aldex2All + Aldex2Median + Aldex2Iqlr + Aldex2Zero + Aldex2Lvha + Aldex2User +} derive(Eq, Debug) + +///| +pub struct Aldex2Config { + mc_samples : Int + prior : Double + denominator : Aldex2Denominator + denominator_indices : Array[Int] + seed : Int + paired : Bool + interval_level : Double +} derive(Eq, Debug) + +///| +pub fn Aldex2Config::create( + mc_samples? : Int = 128, + prior? : Double = 0.5, + denominator? : Aldex2Denominator = Aldex2All, + denominator_indices? : Array[Int] = [], + seed? : Int = 42, + paired? : Bool = false, + interval_level? : Double = 0.95, +) -> Aldex2Config raise Aldex2Error { + if mc_samples < 1 { + raise Aldex2Error("ALDEx2 Monte Carlo sample count must be positive") + } + if !aldex2_is_finite(prior) || prior <= 0.0 { + raise Aldex2Error("ALDEx2 Dirichlet prior must be finite and positive") + } + if seed <= 0 { + raise Aldex2Error("ALDEx2 random seed must be positive") + } + if !aldex2_is_finite(interval_level) || + interval_level <= 0.0 || + interval_level >= 1.0 { + raise Aldex2Error("ALDEx2 interval level must be finite and in (0, 1)") + } + if denominator == Aldex2User && denominator_indices.length() == 0 { + raise Aldex2Error("ALDEx2 user denominator must contain feature indices") + } + if denominator != Aldex2User && denominator_indices.length() > 0 { + raise Aldex2Error( + "ALDEx2 denominator indices require the user denominator mode", + ) + } + Aldex2Config::{ + mc_samples, + prior, + denominator, + denominator_indices: denominator_indices.copy(), + seed, + paired, + interval_level, + } +} + +///| +pub fn Aldex2Config::default() -> Aldex2Config { + Aldex2Config::{ + mc_samples: 128, + prior: 0.5, + denominator: Aldex2All, + denominator_indices: [], + seed: 42, + paired: false, + interval_level: 0.95, + } +} + +///| +pub struct Aldex2Clr { + counts : Array[Array[Double]] + feature_names : Array[String] + sample_names : Array[String] + conditions : Array[String] + group_names : Array[String] + group_indices : Array[Array[Int]] + original_feature_indices : Array[Int] + denominator_by_sample : Array[Array[Int]] + dirichlet : Array[Array[Array[Double]]] + analysis : Array[Array[Array[Double]]] + scale_samples : Array[Array[Double]] + original_feature_count : Int + config : Aldex2Config +} + +///| +pub struct Aldex2TestFeature { + feature_name : String + welch_p_value : Double + welch_adjusted : Double + wilcoxon_p_value : Double + wilcoxon_adjusted : Double +} derive(Eq, Debug) + +///| +pub struct Aldex2TestResult { + features : Array[Aldex2TestFeature] + group_a : String + group_b : String + paired : Bool + mc_samples : Int +} + +///| +pub struct Aldex2EffectFeature { + feature_name : String + rab_all : Double + rab_group_a : Double + rab_group_b : Double + diff_between : Double + diff_within : Double + effect : Double + effect_low : Double + effect_high : Double + overlap : Double +} derive(Eq, Debug) + +///| +pub struct Aldex2EffectResult { + features : Array[Aldex2EffectFeature] + group_a : String + group_b : String + paired : Bool + interval_level : Double +} + +///| +pub struct Aldex2Result { + clr : Aldex2Clr + tests : Aldex2TestResult + effects : Aldex2EffectResult +} + +///| +pub struct Aldex2Summary { + feature_count : Int + sample_count : Int + removed_feature_count : Int + significant_welch_count : Int + significant_wilcoxon_count : Int + large_effect_count : Int + group_a : String + group_b : String + mc_samples : Int + denominator : String + scale_model : Bool +} derive(Eq, Debug) + +///| +pub struct Aldex2SummarizedExperimentOutput { + experiment : SummarizedExperiment + result : Aldex2Result +} + +///| +fn aldex2_is_finite(value : Double) -> Bool { + value == value && value.abs() <= 1.0e300 +} + +///| +fn aldex2_log2(value : Double) -> Double { + @math.ln(value) / @math.ln(2.0) +} + +///| +fn aldex2_mean(values : Array[Double]) -> Double { + if values.length() == 0 { + return 0.0 + } + let mut total = 0.0 + for value in values { + total = total + value + } + total / values.length().to_double() +} + +///| +fn aldex2_variance(values : Array[Double]) -> Double { + if values.length() < 2 { + return 0.0 + } + let center = aldex2_mean(values) + let mut total = 0.0 + for value in values { + let delta = value - center + total = total + delta * delta + } + total / (values.length() - 1).to_double() +} + +///| +fn aldex2_quantile(values : Array[Double], probability : Double) -> Double { + if values.length() == 0 { + return 0.0 + } + let sorted = values.copy() + sorted.sort_by(fn(left : Double, right : Double) -> Int { + if left < right { + -1 + } else if left > right { + 1 + } else { + 0 + } + }) + if sorted.length() == 1 { + return sorted[0] + } + let position = probability.max(0.0).min(1.0) * + (sorted.length() - 1).to_double() + let lower = @math.floor(position).to_int() + let upper = (lower + 1).min(sorted.length() - 1) + let fraction = position - lower.to_double() + sorted[lower] * (1.0 - fraction) + sorted[upper] * fraction +} + +///| +fn aldex2_median(values : Array[Double]) -> Double { + aldex2_quantile(values, 0.5) +} + +///| +fn aldex2_default_names(prefix : String, count : Int) -> Array[String] { + let names : Array[String] = [] + for index in 0.. Unit raise Aldex2Error { + if names.length() != expected { + raise Aldex2Error("ALDEx2 " + label + " name count does not match data") + } + for index in 0.. (Array[String], Array[Array[Int]]) raise Aldex2Error { + let names : Array[String] = [] + let indices : Array[Array[Int]] = [] + for sample in 0.. Array[Int] { + let output : Array[Int] = [] + for index in 0.. Bool { + for value in values { + if value == target { + return true + } + } + false +} + +///| +fn aldex2_intersection(left : Array[Int], right : Array[Int]) -> Array[Int] { + let output : Array[Int] = [] + for value in left { + if aldex2_contains(right, value) { + output.push(value) + } + } + output +} + +///| +fn aldex2_lcg(state : Int) -> (Int, Double) { + let normalized = (state - 1) % 2147483646 + 1 + let quotient = normalized / 44488 + let remainder = normalized % 44488 + let candidate = 48271 * remainder - 3399 * quotient + let next = if candidate > 0 { candidate } else { candidate + 2147483647 } + (next, next.to_double() / 2147483647.0) +} + +///| +fn aldex2_normal(state : Int) -> (Int, Double) { + let (state1, raw1) = aldex2_lcg(state) + let (state2, raw2) = aldex2_lcg(state1) + let first = raw1.max(1.0e-15) + let value = (-2.0 * @math.ln(first)).sqrt() * + @math.cos(6.283185307179586 * raw2) + (state2, value) +} + +///| +fn aldex2_gamma(shape : Double, state : Int) -> (Int, Double) { + if shape < 1.0 { + let (state1, uniform) = aldex2_lcg(state) + let (state2, sampled) = aldex2_gamma(shape + 1.0, state1) + return (state2, sampled * @math.exp(@math.ln(uniform.max(1.0e-15)) / shape)) + } + let d = shape - 1.0 / 3.0 + let c = 1.0 / (9.0 * d).sqrt() + let mut current = state + while true { + let (state1, normal) = aldex2_normal(current) + current = state1 + let base = 1.0 + c * normal + if base > 0.0 { + let volume = base * base * base + let (state2, uniform) = aldex2_lcg(current) + current = state2 + if uniform < 1.0 - 0.0331 * normal * normal * normal * normal { + return (current, d * volume) + } + if @math.ln(uniform.max(1.0e-15)) < + 0.5 * normal * normal + d * (1.0 - volume + @math.ln(volume)) { + return (current, d * volume) + } + } + } + (current, shape) +} + +///| +fn aldex2_pseudoclr( + counts : Array[Array[Double]], + prior : Double, +) -> Array[Array[Double]] { + let feature_count = counts.length() + let sample_count = counts[0].length() + let output : Array[Array[Double]] = [] + for sample in 0.. Array[Double] { + let feature_count = transformed[0].length() + let output : Array[Double] = [] + for feature in 0.. Array[Int] raise Aldex2Error { + let transformed = aldex2_pseudoclr(counts, prior) + let sets : Array[Array[Int]] = [] + for samples in group_indices { + let variances = aldex2_variances_for_samples(transformed, samples) + let lower = aldex2_quantile(variances, 0.25) + let upper = aldex2_quantile(variances, 0.75) + let selected : Array[Int] = [] + for feature in 0.. lower && variances[feature] < upper { + selected.push(feature) + } + } + sets.push(selected) + } + let all_samples = aldex2_all_indices(transformed.length()) + let all_variances = aldex2_variances_for_samples(transformed, all_samples) + let all_lower = aldex2_quantile(all_variances, 0.25) + let all_upper = aldex2_quantile(all_variances, 0.75) + let all_selected : Array[Int] = [] + for feature in 0.. all_lower && all_variances[feature] < all_upper { + all_selected.push(feature) + } + } + sets.push(all_selected) + let mut intersection = sets[0].copy() + for index in 1.. Array[Int] raise Aldex2Error { + let transformed = aldex2_pseudoclr(counts, prior) + let sets : Array[Array[Int]] = [] + for samples in group_indices { + let variances = aldex2_variances_for_samples(transformed, samples) + let abundances : Array[Double] = [] + for feature in 0..= abundance_cutoff { + selected.push(feature) + } + } + sets.push(selected) + } + let mut intersection = sets[0].copy() + for index in 1.. Array[Array[Int]] raise Aldex2Error { + let feature_count = counts.length() + let shared = match config.denominator { + Aldex2Iqlr => aldex2_iqlr_indices(counts, groups, config.prior) + Aldex2Lvha => aldex2_lvha_indices(counts, groups, config.prior) + Aldex2User => { + let selected : Array[Int] = [] + for index in config.denominator_indices { + if index < 0 || index >= feature_count { + raise Aldex2Error( + "ALDEx2 user denominator feature index is out of bounds", + ) + } + if aldex2_contains(selected, index) { + raise Aldex2Error( + "ALDEx2 user denominator contains duplicate feature indices", + ) + } + selected.push(index) + } + selected + } + _ => aldex2_all_indices(feature_count) + } + let by_group : Array[Array[Int]] = [] + if config.denominator == Aldex2Zero { + for samples in groups { + let selected : Array[Int] = [] + for feature in 0.. 0.0 { + selected.push(feature) + } + } + if selected.length() == 0 { + raise Aldex2Error( + "ALDEx2 zero denominator has no observed features in a condition", + ) + } + by_group.push(selected) + } + } else { + for _ in groups { + by_group.push(shared.copy()) + } + } + let by_sample : Array[Array[Int]] = [] + for sample in 0.. Unit raise Aldex2Error { + if scale_samples.length() == 0 { + return + } + if scale_samples.length() != sample_count { + raise Aldex2Error("ALDEx2 scale sample rows must match samples") + } + for sample in 0.. Aldex2Clr raise Aldex2Error { + if counts.length() == 0 { + raise Aldex2Error("ALDEx2 count matrix must contain features") + } + let sample_count = counts[0].length() + if sample_count == 0 { + raise Aldex2Error("ALDEx2 count matrix must contain samples") + } + if conditions.length() != sample_count { + raise Aldex2Error("ALDEx2 condition count must match samples") + } + let supplied_features = if feature_names.length() == 0 { + aldex2_default_names("feature_", counts.length()) + } else { + feature_names.copy() + } + let supplied_samples = if sample_names.length() == 0 { + aldex2_default_names("sample_", sample_count) + } else { + sample_names.copy() + } + aldex2_validate_names(supplied_features, counts.length(), "feature") + aldex2_validate_names(supplied_samples, sample_count, "sample") + let library_sizes = Array::make(sample_count, 0.0) + let filtered_counts : Array[Array[Double]] = [] + let filtered_names : Array[String] = [] + let original_indices : Array[Int] = [] + for feature in 0.. 1.0e-9 { + raise Aldex2Error("ALDEx2 counts must be integers") + } + copied.push(value) + row_total = row_total + value + library_sizes[sample] = library_sizes[sample] + value + } + if row_total > 0.0 { + filtered_counts.push(copied) + filtered_names.push(supplied_features[feature]) + original_indices.push(feature) + } + } + if filtered_counts.length() == 0 { + raise Aldex2Error("ALDEx2 count matrix contains only zero features") + } + for sample in 0..= counts.length() { + raise Aldex2Error( + "ALDEx2 user denominator feature index is out of bounds", + ) + } + let mut filtered_index = -1 + for index in 0.. 0 { + let scale = aldex2_log2(scale_samples[sample][instance]) + for feature in 0.. Int { + self.feature_names.length() +} + +///| +pub fn Aldex2Clr::sample_count(self : Aldex2Clr) -> Int { + self.sample_names.length() +} + +///| +pub fn Aldex2Clr::mc_sample_count(self : Aldex2Clr) -> Int { + self.config.mc_samples +} + +///| +pub fn Aldex2Clr::removed_feature_count(self : Aldex2Clr) -> Int { + self.original_feature_count - self.feature_count() +} + +///| +pub fn Aldex2Clr::sample_instance( + self : Aldex2Clr, + sample : Int, + instance : Int, +) -> Array[Double]? { + if sample < 0 || + sample >= self.sample_count() || + instance < 0 || + instance >= self.config.mc_samples { + return None + } + let output : Array[Double] = [] + for feature in 0.. Array[Double]? { + if feature < 0 || + feature >= self.feature_count() || + instance < 0 || + instance >= self.config.mc_samples { + return None + } + let output : Array[Double] = [] + for sample in 0.. Double? { + if feature < 0 || + feature >= self.feature_count() || + sample < 0 || + sample >= self.sample_count() { + return None + } + Some(aldex2_median(self.analysis[sample][feature])) +} + +///| +pub fn Aldex2Clr::expected_proportion( + self : Aldex2Clr, + feature : Int, + sample : Int, +) -> Double? { + if feature < 0 || + feature >= self.feature_count() || + sample < 0 || + sample >= self.sample_count() { + return None + } + Some(aldex2_median(self.dirichlet[sample][feature])) +} + +///| +fn aldex2_welch( + group_a : Array[Double], + group_b : Array[Double], +) -> (Double, Double) { + if group_a.length() < 2 || group_b.length() < 2 { + return (0.0, 1.0) + } + let mean_a = aldex2_mean(group_a) + let mean_b = aldex2_mean(group_b) + let variance_a = aldex2_variance(group_a) + let variance_b = aldex2_variance(group_b) + let component_a = variance_a / group_a.length().to_double() + let component_b = variance_b / group_b.length().to_double() + let standard_error = (component_a + component_b).sqrt() + if standard_error <= 1.0e-15 { + if (mean_a - mean_b).abs() <= 1.0e-15 { + return (0.0, 1.0) + } + return (if mean_b > mean_a { 1.0e12 } else { -1.0e12 }, 0.0) + } + let statistic = (mean_b - mean_a) / standard_error + let denominator = component_a * + component_a / + (group_a.length() - 1).to_double() + + component_b * component_b / (group_b.length() - 1).to_double() + let degrees = if denominator > 0.0 { + (component_a + component_b) * (component_a + component_b) / denominator + } else { + (group_a.length() + group_b.length() - 2).to_double() + } + let p_value = (2.0 * (1.0 - stat_t_cdf(statistic.abs(), degrees))) + .max(0.0) + .min(1.0) + (statistic, p_value) +} + +///| +fn aldex2_paired_t( + group_a : Array[Double], + group_b : Array[Double], +) -> (Double, Double) { + if group_a.length() != group_b.length() || group_a.length() < 2 { + return (0.0, 1.0) + } + let differences : Array[Double] = [] + for index in 0.. 0.0 { 1.0e12 } else { -1.0e12 }, 0.0) + } + let statistic = center / standard_error + let p_value = (2.0 * + (1.0 - stat_t_cdf(statistic.abs(), (differences.length() - 1).to_double()))) + .max(0.0) + .min(1.0) + (statistic, p_value) +} + +///| +fn aldex2_bh_adjust(p_values : Array[Double]) -> Array[Double] { + let count = p_values.length() + if count == 0 { + return [] + } + let sorted : Array[(Double, Int)] = [] + for index in 0.. Int { + if left.0 < right.0 { + -1 + } else if left.0 > right.0 { + 1 + } else { + left.1 - right.1 + } + }) + let adjusted = Array::make(count, 1.0) + let mut running_minimum = 1.0 + let mut rank = count - 1 + while rank >= 0 { + let candidate = (sorted[rank].0 * count.to_double() / (rank + 1).to_double()).min( + 1.0, + ) + running_minimum = running_minimum.min(candidate) + adjusted[sorted[rank].1] = running_minimum + rank = rank - 1 + } + adjusted +} + +///| +pub fn aldex2_ttest(clr : Aldex2Clr) -> Aldex2TestResult raise Aldex2Error { + if clr.group_names.length() != 2 { + raise Aldex2Error("ALDEx2 t-test requires exactly two conditions") + } + if clr.group_indices[0].length() < 2 || clr.group_indices[1].length() < 2 { + raise Aldex2Error( + "ALDEx2 t-test requires at least two replicates per condition", + ) + } + let feature_count = clr.feature_count() + let mc_samples = clr.config.mc_samples + let welch : Array[Array[Double]] = [] + let wilcoxon : Array[Array[Double]] = [] + let welch_adjusted : Array[Array[Double]] = [] + let wilcoxon_adjusted : Array[Array[Double]] = [] + for _ in 0.. Array[Double] { + let output : Array[Double] = [] + for sample in samples { + for instance in 0.. (Array[Double], Array[Double]) { + let between : Array[Double] = [] + let within : Array[Double] = [] + for instance in 0.. (Array[Double], Array[Double]) { + let between : Array[Double] = [] + for instance in 0.. Aldex2EffectResult raise Aldex2Error { + if clr.group_names.length() != 2 { + raise Aldex2Error( + "ALDEx2 effect estimation requires exactly two conditions", + ) + } + if clr.group_indices[0].length() < 2 || clr.group_indices[1].length() < 2 { + raise Aldex2Error( + "ALDEx2 effect estimation requires at least two replicates per condition", + ) + } + let all_samples = aldex2_all_indices(clr.sample_count()) + let alpha = (1.0 - clr.config.interval_level) / 2.0 + let features : Array[Aldex2EffectFeature] = [] + for feature in 0.. 1.0e-12 { + between[index] / spread + } else if between[index].abs() <= 1.0e-12 { + 0.0 + } else { + between[index] / 1.0e-12 + } + effect_distribution.push(effect) + if effect < 0.0 { + negative = negative + 1 + } else if effect > 0.0 { + positive = positive + 1 + } + } + let overlap_denominator = (negative + positive).to_double() + 1.0 + let overlap = if overlap_denominator > 0.0 { + (negative.min(positive).to_double() + 0.5) / overlap_denominator + } else { + 0.5 + } + features.push(Aldex2EffectFeature::{ + feature_name: clr.feature_names[feature], + rab_all: aldex2_median(pooled_all), + rab_group_a: aldex2_median(pooled_a), + rab_group_b: aldex2_median(pooled_b), + diff_between: aldex2_median(between), + diff_within: aldex2_median(within), + effect: aldex2_median(effect_distribution), + effect_low: aldex2_quantile(effect_distribution, alpha), + effect_high: aldex2_quantile(effect_distribution, 1.0 - alpha), + overlap, + }) + } + Aldex2EffectResult::{ + features, + group_a: clr.group_names[0], + group_b: clr.group_names[1], + paired: clr.config.paired, + interval_level: clr.config.interval_level, + } +} + +///| +pub fn aldex2( + counts : Array[Array[Double]], + conditions : Array[String], + config? : Aldex2Config = Aldex2Config::default(), + feature_names? : Array[String] = [], + sample_names? : Array[String] = [], + scale_samples? : Array[Array[Double]] = [], +) -> Aldex2Result raise Aldex2Error { + let clr = aldex2_clr( + counts, + conditions, + config~, + feature_names~, + sample_names~, + scale_samples~, + ) + let tests = aldex2_ttest(clr) + let effects = aldex2_effect(clr) + Aldex2Result::{ clr, tests, effects } +} + +///| +pub fn Aldex2TestResult::find_feature( + self : Aldex2TestResult, + name : String, +) -> Aldex2TestFeature? { + for feature in self.features { + if feature.feature_name == name { + return Some(feature) + } + } + None +} + +///| +pub fn Aldex2EffectResult::find_feature( + self : Aldex2EffectResult, + name : String, +) -> Aldex2EffectFeature? { + for feature in self.features { + if feature.feature_name == name { + return Some(feature) + } + } + None +} + +///| +pub fn Aldex2Result::select( + self : Aldex2Result, + maximum_adjusted? : Double = 0.05, + minimum_effect? : Double = 1.0, +) -> Array[Aldex2EffectFeature] { + let selected : Array[Aldex2EffectFeature] = [] + for index in 0..= minimum_effect { + selected.push(self.effects.features[index]) + } + } + selected +} + +///| +pub fn Aldex2Result::ranked(self : Aldex2Result) -> Array[Aldex2EffectFeature] { + let ranked = self.effects.features.copy() + ranked.sort_by(fn( + left : Aldex2EffectFeature, + right : Aldex2EffectFeature, + ) -> Int { + if left.effect.abs() > right.effect.abs() { + -1 + } else if left.effect.abs() < right.effect.abs() { + 1 + } else if left.feature_name < right.feature_name { + -1 + } else if left.feature_name > right.feature_name { + 1 + } else { + 0 + } + }) + ranked +} + +///| +fn aldex2_denominator_name(value : Aldex2Denominator) -> String { + match value { + Aldex2All => "all" + Aldex2Median => "median" + Aldex2Iqlr => "iqlr" + Aldex2Zero => "zero" + Aldex2Lvha => "lvha" + Aldex2User => "user" + } +} + +///| +pub fn Aldex2Result::summary( + self : Aldex2Result, + maximum_adjusted? : Double = 0.05, + minimum_effect? : Double = 1.0, +) -> Aldex2Summary { + let mut welch_count = 0 + let mut wilcoxon_count = 0 + let mut effect_count = 0 + for index in 0..= minimum_effect { + effect_count = effect_count + 1 + } + } + Aldex2Summary::{ + feature_count: self.clr.feature_count(), + sample_count: self.clr.sample_count(), + removed_feature_count: self.clr.removed_feature_count(), + significant_welch_count: welch_count, + significant_wilcoxon_count: wilcoxon_count, + large_effect_count: effect_count, + group_a: self.tests.group_a, + group_b: self.tests.group_b, + mc_samples: self.clr.config.mc_samples, + denominator: aldex2_denominator_name(self.clr.config.denominator), + scale_model: self.clr.scale_samples.length() > 0, + } +} + +///| +pub fn Aldex2Clr::expected_distance(self : Aldex2Clr) -> Array[Array[Double]] { + let sample_count = self.sample_count() + let distances : Array[Array[Double]] = [] + for _ in 0.. String { + let buffer = StringBuilder::new() + buffer.write_string( + "feature\trab.all\trab.group.a\trab.group.b\tdiff.btw\tdiff.win\teffect\teffect.low\teffect.high\toverlap\twe.ep\twe.eBH\twi.ep\twi.eBH\n", + ) + for index in 0.. SummarizedExperiment { + let assays : Map[String, Array[Array[Double]]] = Map([]) + for key in experiment.assays.keys() { + let matrix : Array[Array[Double]] = [] + for row in experiment.assays[key] { + matrix.push(row.copy()) + } + assays[key] = matrix + } + let metadata : Map[String, String] = Map([]) + for key in experiment.metadata.keys() { + metadata[key] = experiment.metadata[key] + } + SummarizedExperiment::{ + assays, + row_ranges: experiment.row_ranges.copy(), + col_data: experiment.col_data.copy(), + metadata, + } +} + +///| +pub fn aldex2_summarized_experiment( + experiment : SummarizedExperiment, + conditions : Array[String], + config? : Aldex2Config = Aldex2Config::default(), + assay_name? : String = "counts", + feature_names? : Array[String] = [], + sample_names? : Array[String] = [], + scale_samples? : Array[Array[Double]] = [], +) -> Aldex2SummarizedExperimentOutput raise Aldex2Error { + let counts = match se_assay(experiment, assay_name) { + Some(value) => value + None => + raise Aldex2Error( + "ALDEx2 SummarizedExperiment assay not found: " + assay_name, + ) + } + let result = aldex2( + counts, + conditions, + config~, + feature_names~, + sample_names~, + scale_samples~, + ) + let enriched = aldex2_copy_experiment(experiment) + let effect = Array::make(result.clr.original_feature_count, [0.0]) + let overlap = Array::make(result.clr.original_feature_count, [1.0]) + let welch = Array::make(result.clr.original_feature_count, [1.0]) + let wilcoxon = Array::make(result.clr.original_feature_count, [1.0]) + for index in 0.. ( + Array[Array[Double]], + Array[String], + Array[String], + Array[String], +) { + let counts = [ + [120.0, 132.0, 115.0, 126.0, 480.0, 510.0, 495.0, 525.0], + [360.0, 340.0, 375.0, 355.0, 88.0, 95.0, 82.0, 90.0], + [210.0, 220.0, 205.0, 215.0, 225.0, 218.0, 230.0, 222.0], + [75.0, 82.0, 78.0, 80.0, 92.0, 86.0, 89.0, 94.0], + [42.0, 38.0, 45.0, 40.0, 44.0, 41.0, 43.0, 39.0], + [18.0, 24.0, 20.0, 22.0, 19.0, 21.0, 23.0, 20.0], + [6.0, 8.0, 5.0, 7.0, 26.0, 30.0, 28.0, 32.0], + [14.0, 12.0, 16.0, 13.0, 15.0, 14.0, 13.0, 16.0], + [90.0, 96.0, 88.0, 93.0, 97.0, 91.0, 95.0, 92.0], + [33.0, 36.0, 31.0, 35.0, 34.0, 32.0, 37.0, 33.0], + [55.0, 51.0, 58.0, 54.0, 57.0, 53.0, 56.0, 52.0], + [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + ] + let conditions = [ + "control", "control", "control", "control", "treated", "treated", "treated", + "treated", + ] + let features = [ + "increased", "decreased", "stable_high", "mild", "stable_1", "stable_2", "rare_increased", + "stable_3", "stable_4", "stable_5", "stable_6", "all_zero", + ] + let samples = ["C1", "C2", "C3", "C4", "T1", "T2", "T3", "T4"] + (counts, conditions, features, samples) +} diff --git a/src/align_abstract.mbt b/src/align_abstract.mbt index 3f1160dc..7491f610 100644 --- a/src/align_abstract.mbt +++ b/src/align_abstract.mbt @@ -54,8 +54,12 @@ pub fn AbstractAlignment::new( let n_seqs = sequences.length() let alignment_length = if n_seqs > 0 { sequences[0].length() } else { 0 } AbstractAlignment::{ - sequences, identifiers, alignment_type, - validated: false, n_seqs, alignment_length + sequences, + identifiers, + alignment_type, + validated: false, + n_seqs, + alignment_length, } } @@ -88,7 +92,9 @@ pub fn AbstractAlignment::validate(self : AbstractAlignment) -> (Bool, String) { ///| /// Validate characters in the alignment. -pub fn AbstractAlignment::validate_characters(self : AbstractAlignment) -> (Bool, String) { +pub fn AbstractAlignment::validate_characters( + self : AbstractAlignment, +) -> (Bool, String) { let valid_chars = get_valid_chars(self.alignment_type) let mut seq_idx = 0 @@ -101,7 +107,12 @@ pub fn AbstractAlignment::validate_characters(self : AbstractAlignment) -> (Bool let ch = c.to_string() let pos = i.to_string() let id = self.identifiers[seq_idx] - let msg = "Invalid character " + ch + " at position " + pos + " in sequence " + id + let msg = "Invalid character " + + ch + + " at position " + + pos + + " in sequence " + + id return (false, msg) } i = i + 1 @@ -111,25 +122,20 @@ pub fn AbstractAlignment::validate_characters(self : AbstractAlignment) -> (Bool (true, "") } +///| fn get_valid_chars(align_type : AlignAbstractType) -> Array[UInt16] { let nucleotide_chars : Array[UInt16] = [ - 65, 84, 71, 67, 85, 78, 45, 46, - 82, 89, 83, 87, 75, 77, 66, 86, 68, 72 + 65, 84, 71, 67, 85, 78, 45, 46, 82, 89, 83, 87, 75, 77, 66, 86, 68, 72, ] let protein_chars : Array[UInt16] = [ - 65, 82, 78, 68, 67, 81, 69, 71, 72, 73, - 76, 75, 77, 70, 80, 83, 84, 87, 89, 86, - 66, 90, 88, 45, 46, 85, 79 + 65, 82, 78, 68, 67, 81, 69, 71, 72, 73, 76, 75, 77, 70, 80, 83, 84, 87, 89, 86, + 66, 90, 88, 45, 46, 85, 79, ] let generic_chars : Array[UInt16] = [ - 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, - 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, - 117, 118, 119, 120, 121, 122, - 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, - 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, - 85, 86, 87, 88, 89, 90, - 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, - 45, 46, 42 + 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, + 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 65, 66, 67, 68, 69, 70, 71, + 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 48, + 49, 50, 51, 52, 53, 54, 55, 56, 57, 45, 46, 42, ] match align_type { @@ -139,14 +145,25 @@ fn get_valid_chars(align_type : AlignAbstractType) -> Array[UInt16] { } } +///| fn char_to_upper(c : UInt16) -> UInt16 { - if c >= 97 && c <= 122 { c - 32 } else { c } + if c >= 97 && c <= 122 { + c - 32 + } else { + c + } } +///| fn char_to_lower(c : UInt16) -> UInt16 { - if c >= 65 && c <= 90 { c + 32 } else { c } + if c >= 65 && c <= 90 { + c + 32 + } else { + c + } } +///| /// Convert a UInt16 character code to its 1-character String representation. fn u16_to_str(c : UInt16) -> String { match c { @@ -221,6 +238,7 @@ fn u16_to_str(c : UInt16) -> String { } } +///| fn char_in_array(c : UInt16, arr : Array[UInt16]) -> Bool { for ch in arr { if ch == c { @@ -238,13 +256,18 @@ pub fn AbstractAlignment::abstract_n_seqs(self : AbstractAlignment) -> Int { ///| /// Get alignment length. -pub fn AbstractAlignment::abstract_alignment_length(self : AbstractAlignment) -> Int { +pub fn AbstractAlignment::abstract_alignment_length( + self : AbstractAlignment, +) -> Int { self.alignment_length } ///| /// Get a sequence by index. -pub fn AbstractAlignment::abstract_get_seq(self : AbstractAlignment, idx : Int) -> String { +pub fn AbstractAlignment::abstract_get_seq( + self : AbstractAlignment, + idx : Int, +) -> String { if idx >= 0 && idx < self.n_seqs { self.sequences[idx] } else { @@ -254,7 +277,10 @@ pub fn AbstractAlignment::abstract_get_seq(self : AbstractAlignment, idx : Int) ///| /// Get sequence identifier by index. -pub fn AbstractAlignment::abstract_get_id(self : AbstractAlignment, idx : Int) -> String { +pub fn AbstractAlignment::abstract_get_id( + self : AbstractAlignment, + idx : Int, +) -> String { if idx >= 0 && idx < self.n_seqs { self.identifiers[idx] } else { @@ -264,7 +290,10 @@ pub fn AbstractAlignment::abstract_get_id(self : AbstractAlignment, idx : Int) - ///| /// Get a column from the alignment (array of characters). -pub fn AbstractAlignment::abstract_get_column(self : AbstractAlignment, col_idx : Int) -> Array[UInt16] { +pub fn AbstractAlignment::abstract_get_column( + self : AbstractAlignment, + col_idx : Int, +) -> Array[UInt16] { let column : Array[UInt16] = Array::new() if col_idx >= 0 && col_idx < self.alignment_length { for seq in self.sequences { @@ -276,7 +305,9 @@ pub fn AbstractAlignment::abstract_get_column(self : AbstractAlignment, col_idx ///| /// Get the alignment type. -pub fn AbstractAlignment::abstract_type(self : AbstractAlignment) -> AlignAbstractType { +pub fn AbstractAlignment::abstract_type( + self : AbstractAlignment, +) -> AlignAbstractType { self.alignment_type } @@ -295,13 +326,20 @@ pub struct AlignAbstractColumnStats { ///| /// Compute column statistics for a given alignment column. -pub fn abstract_column_stats(alignment : AbstractAlignment, col_idx : Int) -> AlignAbstractColumnStats { +pub fn abstract_column_stats( + alignment : AbstractAlignment, + col_idx : Int, +) -> AlignAbstractColumnStats { let column = alignment.abstract_get_column(col_idx) let n = column.length() if n == 0 { return AlignAbstractColumnStats::{ - column_index: col_idx, conservation: 0.0, diversity: 0.0, - gap_fraction: 0.0, n_unique_chars: 0, consensus_char: 45 + column_index: col_idx, + conservation: 0.0, + diversity: 0.0, + gap_fraction: 0.0, + n_unique_chars: 0, + consensus_char: 45, } } @@ -357,13 +395,15 @@ pub fn abstract_column_stats(alignment : AbstractAlignment, col_idx : Int) -> Al diversity: entropy, gap_fraction, n_unique_chars: counts.length(), - consensus_char: consensus + consensus_char: consensus, } } ///| /// Get an array of all column statistics for the alignment. -pub fn abstract_all_column_stats(alignment : AbstractAlignment) -> Array[AlignAbstractColumnStats] { +pub fn abstract_all_column_stats( + alignment : AbstractAlignment, +) -> Array[AlignAbstractColumnStats] { let stats : Array[AlignAbstractColumnStats] = Array::new() let mut i = 0 while i < alignment.alignment_length { @@ -375,7 +415,10 @@ pub fn abstract_all_column_stats(alignment : AbstractAlignment) -> Array[AlignAb ///| /// Get the consensus sequence of the alignment (using most common character per column). -pub fn abstract_consensus_sequence(alignment : AbstractAlignment, threshold? : Double = 0.5) -> String { +pub fn abstract_consensus_sequence( + alignment : AbstractAlignment, + threshold? : Double = 0.5, +) -> String { let mut consensus = "" let mut i = 0 while i < alignment.alignment_length { @@ -392,7 +435,9 @@ pub fn abstract_consensus_sequence(alignment : AbstractAlignment, threshold? : D ///| /// Calculate sequence identity matrix (pairwise). -pub fn abstract_identity_matrix(alignment : AbstractAlignment) -> Array[Array[Double]] { +pub fn abstract_identity_matrix( + alignment : AbstractAlignment, +) -> Array[Array[Double]] { let n = alignment.n_seqs let matrix : Array[Array[Double]] = Array::new() let mut i = 0 @@ -400,7 +445,10 @@ pub fn abstract_identity_matrix(alignment : AbstractAlignment) -> Array[Array[Do let row : Array[Double] = Array::new() let mut j = 0 while j < n { - let identity = seq_identity(alignment.sequences[i], alignment.sequences[j]) + let identity = seq_identity( + alignment.sequences[i], + alignment.sequences[j], + ) row.push(identity) j = j + 1 } @@ -410,6 +458,7 @@ pub fn abstract_identity_matrix(alignment : AbstractAlignment) -> Array[Array[Do matrix } +///| fn seq_identity(seq1 : String, seq2 : String) -> Double { if seq1.length() != seq2.length() { return 0.0 @@ -493,6 +542,7 @@ pub fn abstract_parsimony_sites(alignment : AbstractAlignment) -> Int { count } +///| fn is_parsimony_informative(alignment : AbstractAlignment, col : Int) -> Bool { let column = alignment.abstract_get_column(col) let counts : Map[UInt16, Int] = Map([], capacity=20) @@ -529,6 +579,7 @@ pub fn abstract_variable_sites(alignment : AbstractAlignment) -> Int { count } +///| fn is_variable_site(alignment : AbstractAlignment, col : Int) -> Bool { let column = alignment.abstract_get_column(col) if column.length() == 0 { @@ -557,6 +608,7 @@ pub fn abstract_singleton_sites(alignment : AbstractAlignment) -> Int { count } +///| fn is_singleton_site(alignment : AbstractAlignment, col : Int) -> Bool { let column = alignment.abstract_get_column(col) let counts : Map[UInt16, Int] = Map([], capacity=20) @@ -588,7 +640,9 @@ fn is_singleton_site(alignment : AbstractAlignment, col : Int) -> Bool { ///| /// Calculate alignment distance matrix (using simple p-distance). -pub fn abstract_distance_matrix(alignment : AbstractAlignment) -> Array[Array[Double]] { +pub fn abstract_distance_matrix( + alignment : AbstractAlignment, +) -> Array[Array[Double]] { let n = alignment.n_seqs let matrix : Array[Array[Double]] = Array::new() let mut i = 0 @@ -596,7 +650,10 @@ pub fn abstract_distance_matrix(alignment : AbstractAlignment) -> Array[Array[Do let row : Array[Double] = Array::new() let mut j = 0 while j < n { - let dist = pairwise_distance(alignment.sequences[i], alignment.sequences[j]) + let dist = pairwise_distance( + alignment.sequences[i], + alignment.sequences[j], + ) row.push(dist) j = j + 1 } @@ -606,6 +663,7 @@ pub fn abstract_distance_matrix(alignment : AbstractAlignment) -> Array[Array[Do matrix } +///| fn pairwise_distance(seq1 : String, seq2 : String) -> Double { if seq1.length() != seq2.length() { return 1.0 @@ -629,7 +687,10 @@ fn pairwise_distance(seq1 : String, seq2 : String) -> Double { ///| /// Filter alignment columns by gap fraction. -pub fn abstract_filter_gaps(alignment : AbstractAlignment, max_gap_fraction : Double) -> AbstractAlignment { +pub fn abstract_filter_gaps( + alignment : AbstractAlignment, + max_gap_fraction : Double, +) -> AbstractAlignment { let keep_columns : Array[Int] = Array::new() let mut col = 0 while col < alignment.alignment_length { @@ -655,12 +716,19 @@ pub fn abstract_filter_gaps(alignment : AbstractAlignment, max_gap_fraction : Do seq_idx = seq_idx + 1 } - AbstractAlignment::new(sequences=new_seqs, identifiers=alignment.identifiers.copy(), alignment_type=alignment.alignment_type) + AbstractAlignment::new( + sequences=new_seqs, + identifiers=alignment.identifiers.copy(), + alignment_type=alignment.alignment_type, + ) } ///| /// Filter alignment sequences by minimum coverage. -pub fn abstract_filter_coverage(alignment : AbstractAlignment, min_coverage : Double) -> AbstractAlignment { +pub fn abstract_filter_coverage( + alignment : AbstractAlignment, + min_coverage : Double, +) -> AbstractAlignment { let keep_seqs : Array[Int] = Array::new() let mut idx = 0 while idx < alignment.n_seqs { @@ -688,14 +756,26 @@ pub fn abstract_filter_coverage(alignment : AbstractAlignment, min_coverage : Do new_ids.push(alignment.identifiers[idx2]) } - AbstractAlignment::new(sequences=new_seqs, identifiers=new_ids, alignment_type=alignment.alignment_type) + AbstractAlignment::new( + sequences=new_seqs, + identifiers=new_ids, + alignment_type=alignment.alignment_type, + ) } ///| /// Trim alignment to include only the region between start and end columns. -pub fn abstract_trim(alignment : AbstractAlignment, start : Int, end : Int) -> AbstractAlignment { +pub fn abstract_trim( + alignment : AbstractAlignment, + start : Int, + end : Int, +) -> AbstractAlignment { let s = if start > 0 { start } else { 0 } - let e = if end < alignment.alignment_length - 1 { end } else { alignment.alignment_length - 1 } + let e = if end < alignment.alignment_length - 1 { + end + } else { + alignment.alignment_length - 1 + } let new_seqs : Array[String] = Array::new() for seq in alignment.sequences { @@ -703,7 +783,11 @@ pub fn abstract_trim(alignment : AbstractAlignment, start : Int, end : Int) -> A new_seqs.push(trimmed) } - AbstractAlignment::new(sequences=new_seqs, identifiers=alignment.identifiers.copy(), alignment_type=alignment.alignment_type) + AbstractAlignment::new( + sequences=new_seqs, + identifiers=alignment.identifiers.copy(), + alignment_type=alignment.alignment_type, + ) } ///| @@ -720,16 +804,33 @@ pub fn abstract_summary(alignment : AbstractAlignment) -> String { let tstr = alignment_type_str(alignment.alignment_type) "Alignment Summary:\n" + - " Sequences: " + n_seqs + "\n" + - " Length: " + alen + "\n" + - " Type: " + tstr + "\n" + - " Valid: " + vstr + "\n" + - " Coverage: " + coverage + "\n" + - " Overall identity: " + identity + "\n" + - " Variable sites: " + variable + "\n" + - " Parsimony-informative sites: " + parsimony + "\n" + " Sequences: " + + n_seqs + + "\n" + + " Length: " + + alen + + "\n" + + " Type: " + + tstr + + "\n" + + " Valid: " + + vstr + + "\n" + + " Coverage: " + + coverage + + "\n" + + " Overall identity: " + + identity + + "\n" + + " Variable sites: " + + variable + + "\n" + + " Parsimony-informative sites: " + + parsimony + + "\n" } +///| pub fn alignment_type_str(t : AlignAbstractType) -> String { match t { AlignAbstractType::Nucleotide => "Nucleotide" diff --git a/src/align_analysis.mbt b/src/align_analysis.mbt index 1ca5b877..8cc56540 100644 --- a/src/align_analysis.mbt +++ b/src/align_analysis.mbt @@ -22,56 +22,67 @@ pub fn AlnAnalysisResult::new( dn_ds_ratio : Double, dn_ds_ratio_sem : Double, n_synonymous : Int, - n_nonsynonymous : Int + n_nonsynonymous : Int, ) -> AlnAnalysisResult { - AlnAnalysisResult::{ dn, ds, dn_ds_ratio, dn_ds_ratio_sem, n_synonymous, n_nonsynonymous } + AlnAnalysisResult::{ + dn, + ds, + dn_ds_ratio, + dn_ds_ratio_sem, + n_synonymous, + n_nonsynonymous, + } } ///| /// Get dn (non-synonymous substitution rate). -pub fn aln_get_dn(self : AlnAnalysisResult) -> Double { +pub fn AlnAnalysisResult::aln_get_dn(self : AlnAnalysisResult) -> Double { self.dn } ///| /// Get ds (synonymous substitution rate). -pub fn aln_get_ds(self : AlnAnalysisResult) -> Double { +pub fn AlnAnalysisResult::aln_get_ds(self : AlnAnalysisResult) -> Double { self.ds } ///| /// Get dn/ds ratio (omega). -pub fn aln_get_ratio(self : AlnAnalysisResult) -> Double { +pub fn AlnAnalysisResult::aln_get_ratio(self : AlnAnalysisResult) -> Double { self.dn_ds_ratio } ///| /// Get standard error of mean for dn/ds ratio. -pub fn aln_get_ratio_sem(self : AlnAnalysisResult) -> Double { +pub fn AlnAnalysisResult::aln_get_ratio_sem(self : AlnAnalysisResult) -> Double { self.dn_ds_ratio_sem } ///| /// Get number of synonymous substitutions. -pub fn aln_get_n_syn(self : AlnAnalysisResult) -> Int { +pub fn AlnAnalysisResult::aln_get_n_syn(self : AlnAnalysisResult) -> Int { self.n_synonymous } ///| /// Get number of non-synonymous substitutions. -pub fn aln_get_n_nonsyn(self : AlnAnalysisResult) -> Int { +pub fn AlnAnalysisResult::aln_get_n_nonsyn(self : AlnAnalysisResult) -> Int { self.n_nonsynonymous } ///| /// Check if the result indicates positive selection (omega > 1). -pub fn aln_has_positive_selection(self : AlnAnalysisResult) -> Bool { +pub fn AlnAnalysisResult::aln_has_positive_selection( + self : AlnAnalysisResult, +) -> Bool { self.dn_ds_ratio > 1.0 } ///| /// Check if the result indicates purifying selection (omega < 1). -pub fn aln_has_purifying_selection(self : AlnAnalysisResult) -> Bool { +pub fn AlnAnalysisResult::aln_has_purifying_selection( + self : AlnAnalysisResult, +) -> Bool { self.dn_ds_ratio < 1.0 } @@ -79,31 +90,57 @@ pub fn aln_has_purifying_selection(self : AlnAnalysisResult) -> Bool { /// Simple genetic code table (standard code 1). /// Returns amino acid for a codon. pub fn aln_aa_from_codon(codon : String) -> String { - if codon == "TTT" || codon == "TTC" { "F" } - else if codon == "TTA" || codon == "TTG" { "L" } - else if codon == "CTT" || codon == "CTC" || codon == "CTA" || codon == "CTG" { "L" } - else if codon == "ATT" || codon == "ATC" || codon == "ATA" { "I" } - else if codon == "ATG" { "M" } - else if codon == "GTT" || codon == "GTC" || codon == "GTA" || codon == "GTG" { "V" } - else if codon == "TCT" || codon == "TCC" || codon == "TCA" || codon == "TCG" { "S" } - else if codon == "CCT" || codon == "CCC" || codon == "CCA" || codon == "CCG" { "P" } - else if codon == "ACT" || codon == "ACC" || codon == "ACA" || codon == "ACG" { "T" } - else if codon == "GCT" || codon == "GCC" || codon == "GCA" || codon == "GCG" { "A" } - else if codon == "TAT" || codon == "TAC" { "Y" } - else if codon == "TAA" || codon == "TAG" || codon == "TGA" { "*" } - else if codon == "CAT" || codon == "CAC" { "H" } - else if codon == "CAA" || codon == "CAG" { "Q" } - else if codon == "AAT" || codon == "AAC" { "N" } - else if codon == "AAA" || codon == "AAG" { "K" } - else if codon == "GAT" || codon == "GAC" { "D" } - else if codon == "GAA" || codon == "GAG" { "E" } - else if codon == "TGT" || codon == "TGC" { "C" } - else if codon == "TGG" { "W" } - else if codon == "CGT" || codon == "CGC" || codon == "CGA" || codon == "CGG" { "R" } - else if codon == "AGT" || codon == "AGC" { "S" } - else if codon == "AGA" || codon == "AGG" { "R" } - else if codon == "GGT" || codon == "GGC" || codon == "GGA" || codon == "GGG" { "G" } - else { "?" } + if codon == "TTT" || codon == "TTC" { + "F" + } else if codon == "TTA" || codon == "TTG" { + "L" + } else if codon == "CTT" || codon == "CTC" || codon == "CTA" || codon == "CTG" { + "L" + } else if codon == "ATT" || codon == "ATC" || codon == "ATA" { + "I" + } else if codon == "ATG" { + "M" + } else if codon == "GTT" || codon == "GTC" || codon == "GTA" || codon == "GTG" { + "V" + } else if codon == "TCT" || codon == "TCC" || codon == "TCA" || codon == "TCG" { + "S" + } else if codon == "CCT" || codon == "CCC" || codon == "CCA" || codon == "CCG" { + "P" + } else if codon == "ACT" || codon == "ACC" || codon == "ACA" || codon == "ACG" { + "T" + } else if codon == "GCT" || codon == "GCC" || codon == "GCA" || codon == "GCG" { + "A" + } else if codon == "TAT" || codon == "TAC" { + "Y" + } else if codon == "TAA" || codon == "TAG" || codon == "TGA" { + "*" + } else if codon == "CAT" || codon == "CAC" { + "H" + } else if codon == "CAA" || codon == "CAG" { + "Q" + } else if codon == "AAT" || codon == "AAC" { + "N" + } else if codon == "AAA" || codon == "AAG" { + "K" + } else if codon == "GAT" || codon == "GAC" { + "D" + } else if codon == "GAA" || codon == "GAG" { + "E" + } else if codon == "TGT" || codon == "TGC" { + "C" + } else if codon == "TGG" { + "W" + } else if codon == "CGT" || codon == "CGC" || codon == "CGA" || codon == "CGG" { + "R" + } else if codon == "AGT" || codon == "AGC" { + "S" + } else if codon == "AGA" || codon == "AGG" { + "R" + } else if codon == "GGT" || codon == "GGC" || codon == "GGA" || codon == "GGG" { + "G" + } else { + "?" + } } ///| @@ -111,7 +148,11 @@ pub fn aln_aa_from_codon(codon : String) -> String { /// Returns true if the substitution is synonymous (doesn't change amino acid). pub fn aln_is_synonymous(codon : String, pos : Int, new_base : String) -> Bool { let first_part = if pos > 0 { substring(codon, 0, pos) } else { "" } - let second_part = if pos < 2 { substring(codon, pos + 1, 3 - pos - 1) } else { "" } + let second_part = if pos < 2 { + substring(codon, pos + 1, 3 - pos - 1) + } else { + "" + } let new_codon = first_part + new_base + second_part let old_aa = aln_aa_from_codon(codon) let new_aa = aln_aa_from_codon(new_codon) @@ -123,17 +164,17 @@ pub fn aln_is_synonymous(codon : String, pos : Int, new_base : String) -> Bool { pub fn aln_count_sites(sequence : String) -> Array[Int] { let seq_len = sequence.length() let n_codons = seq_len / 3 - + let mut n_syn = 0 let mut n_nonsyn = 0 - + let bases = ["A", "C", "G", "T"] - + let mut codon_idx = 0 while codon_idx < n_codons { let codon_start = codon_idx * 3 let codon = substring(sequence, codon_start, 3) - + // Check each position in the codon let mut pos = 0 while pos < 3 { @@ -154,7 +195,7 @@ pub fn aln_count_sites(sequence : String) -> Array[Int] { } codon_idx = codon_idx + 1 } - + // Each site has 3 possible changes (excluding the original), so divide by 3 let sites = [n_syn / 3, n_nonsyn / 3] sites @@ -164,32 +205,36 @@ pub fn aln_count_sites(sequence : String) -> Array[Int] { /// Calculate dn/ds ratio between two sequences using the NG86 method. /// Nei-Gojobori method for estimating dn and ds. pub fn aln_analyze_dn_ds(seq1 : String, seq2 : String) -> AlnAnalysisResult { - let len = if seq1.length() < seq2.length() { seq1.length() } else { seq2.length() } + let len = if seq1.length() < seq2.length() { + seq1.length() + } else { + seq2.length() + } let n_codons = len / 3 - + let sites = aln_count_sites(seq1) let total_syn_sites = sites[0].to_double() let total_nonsyn_sites = sites[1].to_double() - + if total_syn_sites == 0.0 || total_nonsyn_sites == 0.0 { return AlnAnalysisResult::new(0.0, 0.0, 0.0, 0.0, 0, 0) } - + let mut n_syn_changes = 0.0 let mut n_nonsyn_changes = 0.0 - + let mut codon_idx = 0 while codon_idx < n_codons { let codon_start = codon_idx * 3 let codon1 = substring(seq1, codon_start, 3) let codon2 = substring(seq2, codon_start, 3) - + // Check each position in the codon let mut pos = 0 while pos < 3 { let base1 = substring(codon1, pos, 1) let base2 = substring(codon2, pos, 1) - + if base1 != base2 { if aln_is_synonymous(codon1, pos, base2) { n_syn_changes = n_syn_changes + 1.0 @@ -201,47 +246,49 @@ pub fn aln_analyze_dn_ds(seq1 : String, seq2 : String) -> AlnAnalysisResult { } codon_idx = codon_idx + 1 } - + // Calculate rates let ds = if total_syn_sites > 0.0 { n_syn_changes / total_syn_sites } else { 0.0 } - + let dn = if total_nonsyn_sites > 0.0 { n_nonsyn_changes / total_nonsyn_sites } else { 0.0 } - + // Calculate ratio let ratio = if ds > 0.0 { dn / ds } else { 0.0 } - + // Calculate SEM (simplified) let variance = if total_syn_sites > 0.0 && total_nonsyn_sites > 0.0 { - (1.0 / total_syn_sites) + (1.0 / total_nonsyn_sites) + 1.0 / total_syn_sites + 1.0 / total_nonsyn_sites } else { 1.0 } let sem = ratio * variance.sqrt() - + AlnAnalysisResult::new( dn, ds, ratio, sem, n_syn_changes.to_int(), - n_nonsyn_changes.to_int() + n_nonsyn_changes.to_int(), ) } ///| /// Calculate dn/ds matrix for multiple sequence comparisons. -pub fn aln_calculate_dn_ds_matrix(sequences : Array[String]) -> Array[Array[AlnAnalysisResult]] { +pub fn aln_calculate_dn_ds_matrix( + sequences : Array[String], +) -> Array[Array[AlnAnalysisResult]] { let n = sequences.length() let matrix : Array[Array[AlnAnalysisResult]] = Array::new() - + let mut i = 0 while i < n { let row : Array[AlnAnalysisResult] = Array::new() @@ -262,7 +309,7 @@ pub fn aln_calculate_dn_ds_matrix(sequences : Array[String]) -> Array[Array[AlnA matrix.push(row) i = i + 1 } - + // Fill symmetric entries let mut x = 0 while x < n { @@ -273,15 +320,19 @@ pub fn aln_calculate_dn_ds_matrix(sequences : Array[String]) -> Array[Array[AlnA } x = x + 1 } - + matrix } ///| /// Calculate Jukes-Cantor distance between two sequences. pub fn aln_jukes_cantor_distance(seq1 : String, seq2 : String) -> Double { - let len = if seq1.length() < seq2.length() { seq1.length() } else { seq2.length() } - + let len = if seq1.length() < seq2.length() { + seq1.length() + } else { + seq2.length() + } + let mut diffs = 0 let mut i = 0 while i < len { @@ -290,13 +341,13 @@ pub fn aln_jukes_cantor_distance(seq1 : String, seq2 : String) -> Double { } i = i + 1 } - + let p = diffs.to_double() / len.to_double() if p >= 0.75 { // Distance is too large for JC correction 10.0 } else { - let arg = 1.0 - (4.0 / 3.0) * p + let arg = 1.0 - 4.0 / 3.0 * p if arg > 0.0 { -0.75 * @math.ln(arg) } else { @@ -308,20 +359,32 @@ pub fn aln_jukes_cantor_distance(seq1 : String, seq2 : String) -> Double { ///| /// Calculate Kimura 2-parameter distance. pub fn aln_kimura_2p_distance(seq1 : String, seq2 : String) -> Double { - let len = if seq1.length() < seq2.length() { seq1.length() } else { seq2.length() } - + let len = if seq1.length() < seq2.length() { + seq1.length() + } else { + seq2.length() + } + let mut transitions = 0 let mut transversions = 0 - + let mut i = 0 while i < len { let base1 = substring(seq1, i, 1) let base2 = substring(seq2, i, 1) - + if base1 != base2 { - let is_base1_purine = if base1 == "A" || base1 == "G" { true } else { false } - let is_base2_purine = if base2 == "A" || base2 == "G" { true } else { false } - + let is_base1_purine = if base1 == "A" || base1 == "G" { + true + } else { + false + } + let is_base2_purine = if base2 == "A" || base2 == "G" { + true + } else { + false + } + if is_base1_purine == is_base2_purine { // Same type (both purine or both pyrimidine) = transition transitions = transitions + 1 @@ -332,13 +395,13 @@ pub fn aln_kimura_2p_distance(seq1 : String, seq2 : String) -> Double { } i = i + 1 } - + let p = transitions.to_double() / len.to_double() let q = transversions.to_double() / len.to_double() - + let arg1 = 1.0 - 2.0 * p - q let arg2 = 1.0 - 2.0 * q - + if arg1 <= 0.0 || arg2 <= 0.0 { 10.0 } else { @@ -348,10 +411,12 @@ pub fn aln_kimura_2p_distance(seq1 : String, seq2 : String) -> Double { ///| /// Calculate distance matrix using Jukes-Cantor method. -pub fn aln_jukes_cantor_matrix(sequences : Array[String]) -> Array[Array[Double]] { +pub fn aln_jukes_cantor_matrix( + sequences : Array[String], +) -> Array[Array[Double]] { let n = sequences.length() let matrix : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < n { let row : Array[Double] = Array::make(n, 0.0) @@ -364,7 +429,7 @@ pub fn aln_jukes_cantor_matrix(sequences : Array[String]) -> Array[Array[Double] matrix.push(row) i = i + 1 } - + // Fill symmetric part let mut x = 0 while x < n { @@ -375,7 +440,7 @@ pub fn aln_jukes_cantor_matrix(sequences : Array[String]) -> Array[Array[Double] } x = x + 1 } - + matrix } @@ -384,7 +449,7 @@ pub fn aln_jukes_cantor_matrix(sequences : Array[String]) -> Array[Array[Double] pub fn aln_kimura_2p_matrix(sequences : Array[String]) -> Array[Array[Double]] { let n = sequences.length() let matrix : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < n { let row : Array[Double] = Array::make(n, 0.0) @@ -397,7 +462,7 @@ pub fn aln_kimura_2p_matrix(sequences : Array[String]) -> Array[Array[Double]] { matrix.push(row) i = i + 1 } - + // Fill symmetric part let mut x = 0 while x < n { @@ -408,15 +473,19 @@ pub fn aln_kimura_2p_matrix(sequences : Array[String]) -> Array[Array[Double]] { } x = x + 1 } - + matrix } ///| /// Calculate number of substitutions per site (simple p-distance). pub fn aln_p_distance(seq1 : String, seq2 : String) -> Double { - let len = if seq1.length() < seq2.length() { seq1.length() } else { seq2.length() } - + let len = if seq1.length() < seq2.length() { + seq1.length() + } else { + seq2.length() + } + let mut diffs = 0 let mut i = 0 while i < len { @@ -425,6 +494,6 @@ pub fn aln_p_distance(seq1 : String, seq2 : String) -> Double { } i = i + 1 } - + diffs.to_double() / len.to_double() } diff --git a/src/align_applications.mbt b/src/align_applications.mbt index e45734fc..65b1fc66 100644 --- a/src/align_applications.mbt +++ b/src/align_applications.mbt @@ -12,33 +12,63 @@ pub fn ClustalwCommandline::new(executable : String) -> ClustalwCommandline { } ///| -pub fn ClustalwCommandline::set_input(self : ClustalwCommandline, infile : String) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-infile", Some(infile)) } +pub fn ClustalwCommandline::set_input( + self : ClustalwCommandline, + infile : String, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-infile", Some(infile)), + } } ///| -pub fn ClustalwCommandline::set_output(self : ClustalwCommandline, outfile : String) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-outfile", Some(outfile)) } +pub fn ClustalwCommandline::set_output( + self : ClustalwCommandline, + outfile : String, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-outfile", Some(outfile)), + } } ///| -pub fn ClustalwCommandline::set_output_format(self : ClustalwCommandline, format : String) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-output", Some(format)) } +pub fn ClustalwCommandline::set_output_format( + self : ClustalwCommandline, + format : String, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-output", Some(format)), + } } ///| -pub fn ClustalwCommandline::set_matrix(self : ClustalwCommandline, matrix : String) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-matrix", Some(matrix)) } +pub fn ClustalwCommandline::set_matrix( + self : ClustalwCommandline, + matrix : String, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-matrix", Some(matrix)), + } } ///| -pub fn ClustalwCommandline::set_gap_open(self : ClustalwCommandline, penalty : Double) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-gapopen", Some(penalty.to_string())) } +pub fn ClustalwCommandline::set_gap_open( + self : ClustalwCommandline, + penalty : Double, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-gapopen", Some(penalty.to_string())), + } } ///| -pub fn ClustalwCommandline::set_gap_extend(self : ClustalwCommandline, penalty : Double) -> ClustalwCommandline { - ClustalwCommandline::{ commandline: self.commandline.add_arg("-gapext", Some(penalty.to_string())) } +pub fn ClustalwCommandline::set_gap_extend( + self : ClustalwCommandline, + penalty : Double, +) -> ClustalwCommandline { + ClustalwCommandline::{ + commandline: self.commandline.add_arg("-gapext", Some(penalty.to_string())), + } } ///| @@ -52,33 +82,59 @@ pub struct ClustalOmegaCommandline { } ///| -pub fn ClustalOmegaCommandline::new(executable : String) -> ClustalOmegaCommandline { +pub fn ClustalOmegaCommandline::new( + executable : String, +) -> ClustalOmegaCommandline { ClustalOmegaCommandline::{ commandline: AbstractCommandline::new(executable) } } ///| -pub fn ClustalOmegaCommandline::set_input(self : ClustalOmegaCommandline, infile : String) -> ClustalOmegaCommandline { - ClustalOmegaCommandline::{ commandline: self.commandline.add_arg("-i", Some(infile)) } +pub fn ClustalOmegaCommandline::set_input( + self : ClustalOmegaCommandline, + infile : String, +) -> ClustalOmegaCommandline { + ClustalOmegaCommandline::{ + commandline: self.commandline.add_arg("-i", Some(infile)), + } } ///| -pub fn ClustalOmegaCommandline::set_output(self : ClustalOmegaCommandline, outfile : String) -> ClustalOmegaCommandline { - ClustalOmegaCommandline::{ commandline: self.commandline.add_arg("-o", Some(outfile)) } +pub fn ClustalOmegaCommandline::set_output( + self : ClustalOmegaCommandline, + outfile : String, +) -> ClustalOmegaCommandline { + ClustalOmegaCommandline::{ + commandline: self.commandline.add_arg("-o", Some(outfile)), + } } ///| -pub fn ClustalOmegaCommandline::set_output_format(self : ClustalOmegaCommandline, format : String) -> ClustalOmegaCommandline { - ClustalOmegaCommandline::{ commandline: self.commandline.add_arg("--outfmt", Some(format)) } +pub fn ClustalOmegaCommandline::set_output_format( + self : ClustalOmegaCommandline, + format : String, +) -> ClustalOmegaCommandline { + ClustalOmegaCommandline::{ + commandline: self.commandline.add_arg("--outfmt", Some(format)), + } } ///| -pub fn ClustalOmegaCommandline::set_iterations(self : ClustalOmegaCommandline, n : Int) -> ClustalOmegaCommandline { - ClustalOmegaCommandline::{ commandline: self.commandline.add_arg("--iterations", Some(n.to_string())) } +pub fn ClustalOmegaCommandline::set_iterations( + self : ClustalOmegaCommandline, + n : Int, +) -> ClustalOmegaCommandline { + ClustalOmegaCommandline::{ + commandline: self.commandline.add_arg("--iterations", Some(n.to_string())), + } } ///| -pub fn ClustalOmegaCommandline::set_full_matrix(self : ClustalOmegaCommandline) -> ClustalOmegaCommandline { - ClustalOmegaCommandline::{ commandline: self.commandline.add_arg("--full", None) } +pub fn ClustalOmegaCommandline::set_full_matrix( + self : ClustalOmegaCommandline, +) -> ClustalOmegaCommandline { + ClustalOmegaCommandline::{ + commandline: self.commandline.add_arg("--full", None), + } } ///| @@ -97,23 +153,40 @@ pub fn MuscleCommandline::new(executable : String) -> MuscleCommandline { } ///| -pub fn MuscleCommandline::set_input(self : MuscleCommandline, infile : String) -> MuscleCommandline { - MuscleCommandline::{ commandline: self.commandline.add_arg("-in", Some(infile)) } +pub fn MuscleCommandline::set_input( + self : MuscleCommandline, + infile : String, +) -> MuscleCommandline { + MuscleCommandline::{ + commandline: self.commandline.add_arg("-in", Some(infile)), + } } ///| -pub fn MuscleCommandline::set_output(self : MuscleCommandline, outfile : String) -> MuscleCommandline { - MuscleCommandline::{ commandline: self.commandline.add_arg("-out", Some(outfile)) } +pub fn MuscleCommandline::set_output( + self : MuscleCommandline, + outfile : String, +) -> MuscleCommandline { + MuscleCommandline::{ + commandline: self.commandline.add_arg("-out", Some(outfile)), + } } ///| -pub fn MuscleCommandline::set_diags(self : MuscleCommandline) -> MuscleCommandline { +pub fn MuscleCommandline::set_diags( + self : MuscleCommandline, +) -> MuscleCommandline { MuscleCommandline::{ commandline: self.commandline.add_arg("-diags", None) } } ///| -pub fn MuscleCommandline::set_max_iterations(self : MuscleCommandline, n : Int) -> MuscleCommandline { - MuscleCommandline::{ commandline: self.commandline.add_arg("-maxiters", Some(n.to_string())) } +pub fn MuscleCommandline::set_max_iterations( + self : MuscleCommandline, + n : Int, +) -> MuscleCommandline { + MuscleCommandline::{ + commandline: self.commandline.add_arg("-maxiters", Some(n.to_string())), + } } ///| @@ -132,13 +205,23 @@ pub fn MAFFTCommandline::new(executable : String) -> MAFFTCommandline { } ///| -pub fn MAFFTCommandline::set_input(self : MAFFTCommandline, infile : String) -> MAFFTCommandline { - MAFFTCommandline::{ commandline: self.commandline.add_arg("--input", Some(infile)) } +pub fn MAFFTCommandline::set_input( + self : MAFFTCommandline, + infile : String, +) -> MAFFTCommandline { + MAFFTCommandline::{ + commandline: self.commandline.add_arg("--input", Some(infile)), + } } ///| -pub fn MAFFTCommandline::set_output(self : MAFFTCommandline, outfile : String) -> MAFFTCommandline { - MAFFTCommandline::{ commandline: self.commandline.add_arg("--output", Some(outfile)) } +pub fn MAFFTCommandline::set_output( + self : MAFFTCommandline, + outfile : String, +) -> MAFFTCommandline { + MAFFTCommandline::{ + commandline: self.commandline.add_arg("--output", Some(outfile)), + } } ///| @@ -147,8 +230,13 @@ pub fn MAFFTCommandline::set_auto(self : MAFFTCommandline) -> MAFFTCommandline { } ///| -pub fn MAFFTCommandline::set_threads(self : MAFFTCommandline, n : Int) -> MAFFTCommandline { - MAFFTCommandline::{ commandline: self.commandline.add_arg("--thread", Some(n.to_string())) } +pub fn MAFFTCommandline::set_threads( + self : MAFFTCommandline, + n : Int, +) -> MAFFTCommandline { + MAFFTCommandline::{ + commandline: self.commandline.add_arg("--thread", Some(n.to_string())), + } } ///| @@ -194,4 +282,4 @@ pub fn create_example_mafft() -> MAFFTCommandline { let cmd = cmd.set_auto() let cmd = cmd.set_threads(4) cmd -} \ No newline at end of file +} diff --git a/src/align_bed.mbt b/src/align_bed.mbt new file mode 100644 index 00000000..805d4fb7 --- /dev/null +++ b/src/align_bed.mbt @@ -0,0 +1,1256 @@ +// Biopython-compatible Bio.Align.bed support. +// +// This module models standalone BED records as pairwise alignment coordinate +// paths. BigBed binary storage remains implemented separately in bigbed.mbt. + +///| +/// Error raised for malformed BED data or invalid coordinate operations. +pub suberror AlignBedError { + AlignBedError(String) +} + +///| +/// A BED score. Biopython preserves non-numeric score tokens as text. +pub enum AlignBedScore { + AlignBedNumeric(Double) + AlignBedText(String) +} derive(Eq, Debug) + +///| +/// One point in the target/query absolute coordinate path. +pub struct AlignBedCoordinate { + target : Int + query : Int +} derive(Eq, Debug) + +///| +/// One aligned BED block in absolute target and query coordinates. +pub struct AlignBedBlock { + target_start : Int + target_end : Int + query_start : Int + query_end : Int +} derive(Eq, Debug) + +///| +/// Coordinate-only alignment counts. +pub struct AlignBedCounts { + aligned : Int + target_skip_bases : Int + query_skip_bases : Int + target_skip_opens : Int + query_skip_opens : Int + blocks : Int + columns : Int +} derive(Eq, Debug) + +///| +/// One BED pairwise alignment. +/// +/// Target coordinates may be stored in either orientation by callers; BED +/// serialization normalizes them to increasing genomic coordinates. Query +/// coordinates decrease for reverse-strand alignments. +pub struct AlignBedAlignment { + target_id : String + query_id : String? + coordinates : Array[AlignBedCoordinate] + score : AlignBedScore? + thick_start : Int? + thick_end : Int? + item_rgb : String? + source_columns : Int +} derive(Eq, Debug) + +///| +/// Summary across all records in a BED document. +pub struct AlignBedSummary { + alignment_count : Int + target_count : Int + query_count : Int + plus_strand_count : Int + minus_strand_count : Int + aligned_bases : Int + target_skip_bases : Int +} derive(Eq, Debug) + +///| +/// A standalone BED document containing zero or more pairwise alignments. +pub struct AlignBedDocument { + alignments : Array[AlignBedAlignment] +} derive(Eq, Debug) + +///| +fn align_bed_fail(message : String) -> Unit raise AlignBedError { + raise AlignBedError(message) +} + +///| +fn align_bed_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn align_bed_min(left : Int, right : Int) -> Int { + if left < right { + left + } else { + right + } +} + +///| +fn align_bed_max(left : Int, right : Int) -> Int { + if left > right { + left + } else { + right + } +} + +///| +fn align_bed_sign(value : Int) -> Int { + if value < 0 { + -1 + } else if value > 0 { + 1 + } else { + 0 + } +} + +///| +fn align_bed_copy_coordinates( + coordinates : Array[AlignBedCoordinate], +) -> Array[AlignBedCoordinate] { + let copied : Array[AlignBedCoordinate] = [] + for coordinate in coordinates { + copied.push(coordinate) + } + copied +} + +///| +fn align_bed_copy_alignments( + alignments : Array[AlignBedAlignment], +) -> Array[AlignBedAlignment] { + let copied : Array[AlignBedAlignment] = [] + for alignment in alignments { + copied.push(alignment) + } + copied +} + +///| +fn align_bed_validate_token( + value : String, + label : String, + allow_empty : Bool, +) -> Unit raise AlignBedError { + if !allow_empty && value.length() == 0 { + align_bed_fail(label + " must not be empty") + } + for index in 0.. Unit raise AlignBedError { + match score { + AlignBedNumeric(value) => + if value.is_nan() || value.abs() > 1.0e300 { + align_bed_fail("BED score must be finite") + } + AlignBedText(value) => align_bed_validate_token(value, "BED score", false) + } +} + +///| +fn align_bed_reversed_coordinates( + coordinates : Array[AlignBedCoordinate], +) -> Array[AlignBedCoordinate] { + let reversed : Array[AlignBedCoordinate] = [] + for index in 0.. Array[AlignBedCoordinate] { + if coordinates.length() > 1 && + coordinates[0].target > coordinates[coordinates.length() - 1].target { + align_bed_reversed_coordinates(coordinates) + } else { + align_bed_copy_coordinates(coordinates) + } +} + +///| +fn align_bed_blocks_from_coordinates( + coordinates : Array[AlignBedCoordinate], +) -> Array[AlignBedBlock] { + let forward = align_bed_forward_coordinates(coordinates) + let blocks : Array[AlignBedBlock] = [] + if forward.length() < 2 { + return blocks + } + let mut target_start = forward[0].target + let mut query_start = forward[0].query + for index in 1.. Unit raise AlignBedError { + align_bed_validate_token(alignment.target_id, "BED target identifier", false) + match alignment.query_id { + Some(value) => + align_bed_validate_token(value, "BED query identifier", false) + None => () + } + if alignment.source_columns < 3 || alignment.source_columns > 12 { + align_bed_fail("BED column count must be between 3 and 12") + } + match alignment.score { + Some(value) => align_bed_validate_score(value) + None => () + } + match alignment.item_rgb { + Some(value) => align_bed_validate_token(value, "BED itemRgb", false) + None => () + } + if alignment.coordinates.length() < 2 { + align_bed_fail("BED coordinate path must contain at least two points") + } + let mut target_direction = 0 + let mut query_direction = 0 + let mut aligned_segments = 0 + for index in 0.. + if value < interval_start || value > interval_end { + align_bed_fail("BED thickStart must lie inside the target interval") + } + None => () + } + match alignment.thick_end { + Some(value) => + if value < interval_start || value > interval_end { + align_bed_fail("BED thickEnd must lie inside the target interval") + } + None => () + } + match (alignment.thick_start, alignment.thick_end) { + (Some(start), Some(end)) => + if start > end { + align_bed_fail("BED thickStart must not exceed thickEnd") + } + _ => () + } +} + +///| +/// Construct one coordinate point. +pub fn AlignBedCoordinate::create( + target : Int, + query : Int, +) -> AlignBedCoordinate raise AlignBedError { + if target < 0 || query < 0 { + align_bed_fail("BED coordinates must be non-negative") + } + AlignBedCoordinate::{ target, query } +} + +///| +/// Construct a finite numeric BED score. +pub fn AlignBedScore::numeric( + value : Double, +) -> AlignBedScore raise AlignBedError { + let score = AlignBedNumeric(value) + align_bed_validate_score(score) + score +} + +///| +/// Construct a non-empty textual BED score token. +pub fn AlignBedScore::text(value : String) -> AlignBedScore raise AlignBedError { + let score = AlignBedText(value) + align_bed_validate_score(score) + score +} + +///| +/// Return the numeric score, if this score was parsed as a number. +pub fn AlignBedScore::number(self : AlignBedScore) -> Double? { + match self { + AlignBedNumeric(value) => Some(value) + AlignBedText(_) => None + } +} + +///| +/// Return the original textual score, if this is a non-numeric token. +pub fn AlignBedScore::text_value(self : AlignBedScore) -> String? { + match self { + AlignBedNumeric(_) => None + AlignBedText(value) => Some(value) + } +} + +///| +/// Construct and validate a pairwise BED alignment. +pub fn AlignBedAlignment::create( + target_id : String, + query_id : String?, + coordinates : Array[AlignBedCoordinate], + score? : AlignBedScore? = None, + thick_start? : Int? = None, + thick_end? : Int? = None, + item_rgb? : String? = None, + source_columns? : Int = 12, +) -> AlignBedAlignment raise AlignBedError { + let alignment = AlignBedAlignment::{ + target_id, + query_id, + coordinates: align_bed_copy_coordinates(coordinates), + score, + thick_start, + thick_end, + item_rgb, + source_columns, + } + align_bed_validate_alignment(alignment) + alignment +} + +///| +/// Construct a BED document and copy the input alignment array. +pub fn AlignBedDocument::create( + alignments : Array[AlignBedAlignment], +) -> AlignBedDocument raise AlignBedError { + let copied = align_bed_copy_alignments(alignments) + for alignment in copied { + align_bed_validate_alignment(alignment) + } + AlignBedDocument::{ alignments: copied } +} + +///| +fn align_bed_strip_cr(value : String) -> String { + if value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '\r'.to_int() { + value[0:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn align_bed_split_lines(value : String) -> Array[String] { + let lines : Array[String] = [] + let mut start = 0 + for index in 0.. 0 { + lines.push("") + } + lines +} + +///| +fn align_bed_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} + +///| +fn align_bed_trim(value : String) -> String { + let mut start = 0 + while start < value.length() && + align_bed_is_whitespace(value.unsafe_get(start).to_int()) { + start = start + 1 + } + let mut end = value.length() + while end > start && + align_bed_is_whitespace(value.unsafe_get(end - 1).to_int()) { + end = end - 1 + } + value[start:end].to_owned() +} + +///| +fn align_bed_split_whitespace(value : String) -> Array[String] { + let words : Array[String] = [] + let mut index = 0 + while index < value.length() { + while index < value.length() && + align_bed_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= value.length() { + break + } + let start = index + while index < value.length() && + !align_bed_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + words.push(value[start:index].to_owned()) + } + words +} + +///| +fn align_bed_parse_uint( + value : String, + label : String, +) -> Int raise AlignBedError { + if value.length() == 0 { + align_bed_fail(label + " must be an unsigned integer") + } + let mut result = 0 + for index in 0.. '9'.to_int() { + align_bed_fail(label + " must be an unsigned integer") + } + let digit = code - '0'.to_int() + if result > (2147483647 - digit) / 10 { + align_bed_fail(label + " exceeds the supported integer range") + } + result = result * 10 + digit + } + result +} + +///| +fn align_bed_is_number(value : String) -> Bool { + if value.length() == 0 { + return false + } + let mut index = 0 + if value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() { + index = index + 1 + } + let mut digits = 0 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + if index < value.length() && value.unsafe_get(index).to_int() == '.'.to_int() { + index = index + 1 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + } + if digits == 0 { + return false + } + if index < value.length() && + ( + value.unsafe_get(index).to_int() == 'e'.to_int() || + value.unsafe_get(index).to_int() == 'E'.to_int() + ) { + index = index + 1 + if index < value.length() && + ( + value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() + ) { + index = index + 1 + } + let exponent_start = index + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + index = index + 1 + } + if index == exponent_start { + return false + } + } + index == value.length() +} + +///| +fn align_bed_parse_score(value : String) -> AlignBedScore raise AlignBedError { + if !align_bed_is_number(value) { + return AlignBedScore::text(value) + } + let parsed_value = if value.unsafe_get(0).to_int() == '+'.to_int() { + value[1:value.length()].to_owned() + } else { + value + } + let number = match parse_double(parsed_value) { + Some(result) => result + None => { + align_bed_fail("BED numeric score could not be parsed") + 0.0 + } + } + AlignBedScore::numeric(number) +} + +///| +fn align_bed_parse_uint_list( + value : String, + label : String, +) -> Array[Int] raise AlignBedError { + let values : Array[Int] = [] + let mut start = 0 + for index in 0.. Unit raise AlignBedError { + if count <= 0 { + align_bed_fail("BED blockCount must be positive") + } + if sizes.length() != count { + align_bed_fail("BED blockCount does not match blockSizes") + } + if starts.length() != count { + align_bed_fail("BED blockCount does not match blockStarts") + } + if starts[0] != 0 { + align_bed_fail("the first BED block must start at chromStart") + } + let mut previous_end = 0 + for index in 0.. span || size > span - start { + align_bed_fail("BED block extends beyond chromEnd") + } + previous_end = start + size + } + if previous_end != span { + align_bed_fail("the last BED block must end at chromEnd") + } +} + +///| +fn align_bed_infer_starts( + span : Int, + sizes : Array[Int], +) -> Array[Int] raise AlignBedError { + let starts : Array[Int] = [] + let mut position = 0 + for size in sizes { + if size <= 0 || position > span || size > span - position { + align_bed_fail("BED11 blockSizes cannot be placed inside chromEnd") + } + starts.push(position) + position = position + size + } + if position != span { + align_bed_fail( + "BED11 blockStarts are ambiguous unless blockSizes cover the full span", + ) + } + starts +} + +///| +fn align_bed_coordinates_from_blocks( + chrom_start : Int, + strand : String, + sizes : Array[Int], + starts : Array[Int], +) -> Array[AlignBedCoordinate] { + let coordinates : Array[AlignBedCoordinate] = [] + let mut query_position = 0 + let mut target_position = chrom_start + coordinates.push(AlignBedCoordinate::{ target: chrom_start, query: 0 }) + for index in 0.. AlignBedAlignment raise AlignBedError { + let words = align_bed_split_whitespace(line) + let bed_columns = words.length() + if bed_columns < 3 || bed_columns > 12 { + align_bed_fail( + "expected between 3 and 12 BED columns at line " + + line_number.to_string() + + ", found " + + bed_columns.to_string(), + ) + } + let target_id = words[0] + align_bed_validate_token(target_id, "BED target identifier", false) + let chrom_start = align_bed_parse_uint(words[1], "BED chromStart") + let chrom_end = align_bed_parse_uint(words[2], "BED chromEnd") + if chrom_end <= chrom_start { + align_bed_fail("BED coordinates must define a positive half-open interval") + } + let span = chrom_end - chrom_start + let query_id = if bed_columns >= 4 { + align_bed_validate_token(words[3], "BED query identifier", false) + Some(words[3]) + } else { + None + } + let score = if bed_columns >= 5 { + Some(align_bed_parse_score(words[4])) + } else { + None + } + let strand = if bed_columns >= 6 { + if words[5] != "+" && words[5] != "-" && words[5] != "." { + align_bed_fail("BED strand must be '+', '-', or '.'") + } + if words[5] == "-" { + "-" + } else { + "+" + } + } else { + "+" + } + let thick_start = if bed_columns >= 7 { + Some(align_bed_parse_uint(words[6], "BED thickStart")) + } else { + None + } + let thick_end = if bed_columns >= 8 { + Some(align_bed_parse_uint(words[7], "BED thickEnd")) + } else { + None + } + let item_rgb = if bed_columns >= 9 { + align_bed_validate_token(words[8], "BED itemRgb", false) + Some(words[8]) + } else { + None + } + let sizes : Array[Int] = [] + let starts : Array[Int] = [] + if bed_columns <= 9 { + sizes.push(span) + starts.push(0) + } else { + let block_count = align_bed_parse_uint(words[9], "BED blockCount") + if bed_columns == 10 { + if block_count != 1 { + align_bed_fail( + "BED10 records with multiple blocks lack blockSizes and blockStarts", + ) + } + sizes.push(span) + starts.push(0) + } else { + let parsed_sizes = align_bed_parse_uint_list(words[10], "BED blockSizes") + if parsed_sizes.length() != block_count { + align_bed_fail("BED blockCount does not match blockSizes") + } + for size in parsed_sizes { + sizes.push(size) + } + if bed_columns == 11 { + let inferred = align_bed_infer_starts(span, sizes) + for start in inferred { + starts.push(start) + } + } else { + let parsed_starts = align_bed_parse_uint_list( + words[11], + "BED blockStarts", + ) + for start in parsed_starts { + starts.push(start) + } + } + align_bed_validate_blocks(span, block_count, sizes, starts) + } + } + let coordinates = align_bed_coordinates_from_blocks( + chrom_start, strand, sizes, starts, + ) + AlignBedAlignment::create( + target_id, + query_id, + coordinates, + score~, + thick_start~, + thick_end~, + item_rgb~, + source_columns=bed_columns, + ) +} + +///| +/// Parse all BED records from a string. +pub fn align_bed_parse(text : String) -> AlignBedDocument raise AlignBedError { + if text.length() == 0 { + return AlignBedDocument::create([]) + } + let lines = align_bed_split_lines(text) + let alignments : Array[AlignBedAlignment] = [] + for index in 0.. Array[AlignBedBlock] { + align_bed_blocks_from_coordinates(self.coordinates) +} + +///| +/// Return the relative query strand encoded by this coordinate path. +pub fn AlignBedAlignment::strand(self : AlignBedAlignment) -> String { + let forward = align_bed_forward_coordinates(self.coordinates) + if forward[0].query > forward[forward.length() - 1].query { + "-" + } else { + "+" + } +} + +///| +/// Return the first aligned target coordinate. +pub fn AlignBedAlignment::target_start(self : AlignBedAlignment) -> Int { + self.blocks()[0].target_start +} + +///| +/// Return the exclusive final aligned target coordinate. +pub fn AlignBedAlignment::target_end(self : AlignBedAlignment) -> Int { + let blocks = self.blocks() + blocks[blocks.length() - 1].target_end +} + +///| +/// Return the query length represented by BED blocks. +pub fn AlignBedAlignment::query_size(self : AlignBedAlignment) -> Int { + let mut size = 0 + for block in self.blocks() { + size = size + (block.target_end - block.target_start) + } + size +} + +///| +/// Return the number of aligned target/query residues. +pub fn AlignBedAlignment::aligned_bases(self : AlignBedAlignment) -> Int { + self.query_size() +} + +///| +/// Return the target span between the first and final aligned blocks. +pub fn AlignBedAlignment::target_span(self : AlignBedAlignment) -> Int { + self.target_end() - self.target_start() +} + +///| +/// Return coordinate-only alignment statistics. +pub fn AlignBedAlignment::counts(self : AlignBedAlignment) -> AlignBedCounts { + let mut aligned = 0 + let mut target_skips = 0 + let mut query_skips = 0 + let mut target_opens = 0 + let mut query_opens = 0 + let mut columns = 0 + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = align_bed_abs(next.target - current.target) + let query_step = align_bed_abs(next.query - current.query) + if target_step > 0 && query_step > 0 { + aligned = aligned + target_step + } else if target_step > 0 { + target_skips = target_skips + target_step + target_opens = target_opens + 1 + } else { + query_skips = query_skips + query_step + query_opens = query_opens + 1 + } + columns = columns + align_bed_max(target_step, query_step) + } + AlignBedCounts::{ + aligned, + target_skip_bases: target_skips, + query_skip_bases: query_skips, + target_skip_opens: target_opens, + query_skip_opens: query_opens, + blocks: self.blocks().length(), + columns, + } +} + +///| +fn align_bed_map_on_segment( + source : Int, + source_start : Int, + source_end : Int, + destination_start : Int, + destination_end : Int, +) -> Int? { + let source_step = source_end - source_start + let destination_step = destination_end - destination_start + if source_step == 0 || + destination_step == 0 || + align_bed_abs(source_step) != align_bed_abs(destination_step) { + return None + } + let offset = if source_step > 0 { + if source < source_start || source >= source_end { + return None + } + source - source_start + } else { + if source < source_end || source >= source_start { + return None + } + source_start - 1 - source + } + if destination_step > 0 { + Some(destination_start + offset) + } else { + Some(destination_start - 1 - offset) + } +} + +///| +/// Map one target residue to the query, returning `None` in an intron/gap. +pub fn AlignBedAlignment::map_target_position( + self : AlignBedAlignment, + position : Int, +) -> Int? { + if position < 0 { + return None + } + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + match + align_bed_map_on_segment( + position, + current.target, + next.target, + current.query, + next.query, + ) { + Some(mapped) => return Some(mapped) + None => () + } + } + None +} + +///| +/// Map one query residue to the target, returning `None` in an unaligned gap. +pub fn AlignBedAlignment::map_query_position( + self : AlignBedAlignment, + position : Int, +) -> Int? { + if position < 0 { + return None + } + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + match + align_bed_map_on_segment( + position, + current.query, + next.query, + current.target, + next.target, + ) { + Some(mapped) => return Some(mapped) + None => () + } + } + None +} + +///| +/// Return true if the aligned target interval overlaps a half-open query. +pub fn AlignBedAlignment::overlaps( + self : AlignBedAlignment, + target_id : String, + start : Int, + end : Int, +) -> Bool { + end > start && + self.target_id == target_id && + self.target_start() < end && + start < self.target_end() +} + +///| +fn align_bed_join_tabs(fields : Array[String]) -> String { + let output = StringBuilder::new() + for index in 0.. 0 { + output.write_char('\t') + } + output.write_string(fields[index]) + } + output.to_string() +} + +///| +fn align_bed_integer_list(values : Array[Int]) -> String { + let output = StringBuilder::new() + for value in values { + output.write_string(value.to_string()) + output.write_char(',') + } + output.to_string() +} + +///| +fn align_bed_score_text(score : AlignBedScore) -> String { + match score { + AlignBedNumeric(value) => value.to_string() + AlignBedText(value) => value + } +} + +///| +/// Format one alignment using BED3 through BED12. +pub fn align_bed_format( + alignment : AlignBedAlignment, + bed_columns? : Int = 12, +) -> String raise AlignBedError { + if bed_columns < 3 || bed_columns > 12 { + align_bed_fail("BED column count must be between 3 and 12") + } + align_bed_validate_alignment(alignment) + let blocks = alignment.blocks() + if blocks.length() == 0 { + return "" + } + let chrom_start = blocks[0].target_start + let chrom_end = blocks[blocks.length() - 1].target_end + let fields : Array[String] = [ + alignment.target_id, + chrom_start.to_string(), + chrom_end.to_string(), + ] + if bed_columns == 3 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(alignment.query_id.unwrap_or("query")) + if bed_columns == 4 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push( + match alignment.score { + Some(score) => align_bed_score_text(score) + None => "0" + }, + ) + if bed_columns == 5 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(alignment.strand()) + if bed_columns == 6 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(alignment.thick_start.unwrap_or(chrom_start).to_string()) + if bed_columns == 7 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(alignment.thick_end.unwrap_or(chrom_end).to_string()) + if bed_columns == 8 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(alignment.item_rgb.unwrap_or("0")) + if bed_columns == 9 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(blocks.length().to_string()) + if bed_columns == 10 { + return align_bed_join_tabs(fields) + "\n" + } + let block_sizes : Array[Int] = [] + let block_starts : Array[Int] = [] + for block in blocks { + block_sizes.push(block.target_end - block.target_start) + block_starts.push(block.target_start - chrom_start) + } + fields.push(align_bed_integer_list(block_sizes)) + if bed_columns == 11 { + return align_bed_join_tabs(fields) + "\n" + } + fields.push(align_bed_integer_list(block_starts)) + align_bed_join_tabs(fields) + "\n" +} + +///| +/// Serialize all records in a BED document. +pub fn align_bed_write( + document : AlignBedDocument, + bed_columns? : Int = 12, +) -> String raise AlignBedError { + let validated = AlignBedDocument::create(document.alignments) + let output = StringBuilder::new() + for alignment in validated.alignments { + output.write_string(align_bed_format(alignment, bed_columns~)) + } + output.to_string() +} + +///| +/// Find all records overlapping one target half-open interval. +pub fn AlignBedDocument::search( + self : AlignBedDocument, + target_id : String, + start : Int, + end : Int, +) -> Array[AlignBedAlignment] raise AlignBedError { + if start < 0 || end <= start { + align_bed_fail("BED search interval must satisfy 0 <= start < end") + } + let results : Array[AlignBedAlignment] = [] + for alignment in self.alignments { + if alignment.overlaps(target_id, start, end) { + results.push(alignment) + } + } + results +} + +///| +/// Return distinct target identifiers in first-seen order. +pub fn AlignBedDocument::targets(self : AlignBedDocument) -> Array[String] { + let targets : Array[String] = [] + let seen : Map[String, Bool] = Map([]) + for alignment in self.alignments { + if !seen.contains(alignment.target_id) { + seen[alignment.target_id] = true + targets.push(alignment.target_id) + } + } + targets +} + +///| +/// Return distinct non-missing query identifiers in first-seen order. +pub fn AlignBedDocument::queries(self : AlignBedDocument) -> Array[String] { + let queries : Array[String] = [] + let seen : Map[String, Bool] = Map([]) + for alignment in self.alignments { + match alignment.query_id { + Some(query) => + if !seen.contains(query) { + seen[query] = true + queries.push(query) + } + None => () + } + } + queries +} + +///| +/// Summarize all records in the document. +pub fn AlignBedDocument::summary(self : AlignBedDocument) -> AlignBedSummary { + let mut plus = 0 + let mut minus = 0 + let mut aligned = 0 + let mut target_skips = 0 + for alignment in self.alignments { + if alignment.strand() == "-" { + minus = minus + 1 + } else { + plus = plus + 1 + } + let counts = alignment.counts() + aligned = aligned + counts.aligned + target_skips = target_skips + counts.target_skip_bases + } + AlignBedSummary::{ + alignment_count: self.alignments.length(), + target_count: self.targets().length(), + query_count: self.queries().length(), + plus_strand_count: plus, + minus_strand_count: minus, + aligned_bases: aligned, + target_skip_bases: target_skips, + } +} + +///| +/// A compact Biopython-style BED3/BED12 fixture. +pub fn align_bed_example_text() -> String { + "chr22\t1000\t5000\tmRNA1\t960\t+\t1200\t4900\t255,0,0\t2\t567,488,\t0,3512,\n" + + "chr22\t2000\t6000\tmRNA2\t900\t-\t2300\t5960\t0,255,0\t2\t433,399,\t0,3601,\n" + + "chr7\t100\t180" +} diff --git a/src/align_chain.mbt b/src/align_chain.mbt new file mode 100644 index 00000000..cdaac195 --- /dev/null +++ b/src/align_chain.mbt @@ -0,0 +1,1403 @@ +// Biopython-compatible Bio.Align.chain support. +// +// This module is intentionally separate from chain_liftover.mbt. The older +// module provides a permissive rtracklayer-style liftOver API, while this file +// models modern Bio.Align pairwise alignments as absolute coordinate paths. + +///| +/// Error raised for malformed chain data or invalid coordinate operations. +pub suberror AlignChainError { + AlignChainError(String) +} + +///| +/// One point in a two-row absolute coordinate path. +pub struct AlignChainCoordinate { + target : Int + query : Int +} derive(Eq, Debug) + +///| +/// One canonical UCSC chain block. +/// +/// A non-final row contains all three fields. The final row uses `size` only, +/// and therefore has zero target and query gaps. +pub struct AlignChainBlock { + size : Int + target_gap : Int + query_gap : Int +} derive(Eq, Debug) + +///| +/// One aligned part of a requested genomic range. +/// +/// Coordinates are always zero-based, forward-axis, half-open intervals. +pub struct AlignChainRangeMapping { + target_start : Int + target_end : Int + query_start : Int + query_end : Int + target_strand : String + query_strand : String +} derive(Eq, Debug) + +///| +/// One aligned residue pair in forward genomic coordinates. +pub struct AlignChainPair { + target : Int + query : Int +} derive(Eq, Debug) + +///| +/// Coordinate-only chain alignment statistics. +pub struct AlignChainCounts { + aligned : Int + target_gap_bases : Int + query_gap_bases : Int + gap_columns : Int + target_gap_opens : Int + query_gap_opens : Int + aligned_blocks : Int + path_segments : Int +} derive(Eq, Debug) + +///| +/// A coordinate-aware pairwise chain alignment. +/// +/// `coordinates` follows Biopython's `Alignment.coordinates` convention. +/// Reverse strands are represented by decreasing coordinates. Chain files do +/// not store sequence letters, only sequence lengths and the coordinate path. +pub struct AlignChainAlignment { + target_id : String + target_size : Int + query_id : String + query_size : Int + score : Double + chain_id : String + coordinates : Array[AlignChainCoordinate] +} derive(Eq, Debug) + +///| +fn align_chain_fail(message : String) -> Unit raise AlignChainError { + raise AlignChainError(message) +} + +///| +fn align_chain_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn align_chain_min(left : Int, right : Int) -> Int { + if left < right { + left + } else { + right + } +} + +///| +fn align_chain_max(left : Int, right : Int) -> Int { + if left > right { + left + } else { + right + } +} + +///| +fn align_chain_sign(value : Int) -> Int { + if value < 0 { + -1 + } else if value > 0 { + 1 + } else { + 0 + } +} + +///| +fn align_chain_copy_coordinates( + coordinates : Array[AlignChainCoordinate], +) -> Array[AlignChainCoordinate] { + let copied : Array[AlignChainCoordinate] = [] + for coordinate in coordinates { + copied.push(coordinate) + } + copied +} + +///| +fn align_chain_copy_alignments( + alignments : Array[AlignChainAlignment], +) -> Array[AlignChainAlignment] { + let copied : Array[AlignChainAlignment] = [] + for alignment in alignments { + copied.push(alignment) + } + copied +} + +///| +fn align_chain_validate_token( + value : String, + label : String, + allow_empty : Bool, +) -> Unit raise AlignChainError { + if !allow_empty && value.length() == 0 { + align_chain_fail(label + " must not be empty") + } + for index in 0.. Unit raise AlignChainError { + if strand != "+" && strand != "-" { + align_chain_fail(label + " strand must be '+' or '-'") + } +} + +///| +fn align_chain_validate_interval( + start : Int, + end : Int, + size : Int, + label : String, +) -> Unit raise AlignChainError { + if size < 0 { + align_chain_fail(label + " size must be non-negative") + } + if start < 0 || end < start || end > size { + align_chain_fail(label + " interval must satisfy 0 <= start <= end <= size") + } +} + +///| +fn align_chain_validate_alignment( + alignment : AlignChainAlignment, +) -> Unit raise AlignChainError { + align_chain_validate_token(alignment.target_id, "target identifier", false) + align_chain_validate_token(alignment.query_id, "query identifier", false) + align_chain_validate_token(alignment.chain_id, "chain identifier", true) + if alignment.target_size < 0 { + align_chain_fail("target size must be non-negative") + } + if alignment.query_size < 0 { + align_chain_fail("query size must be non-negative") + } + if alignment.score.is_nan() || alignment.score.abs() > 1.0e300 { + align_chain_fail("chain score must be finite") + } + if alignment.coordinates.length() < 2 { + align_chain_fail("chain coordinate path must contain at least two points") + } + let mut target_direction = 0 + let mut query_direction = 0 + let mut aligned = 0 + for index in 0.. alignment.target_size { + align_chain_fail("target coordinate is outside the declared sequence") + } + if point.query < 0 || point.query > alignment.query_size { + align_chain_fail("query coordinate is outside the declared sequence") + } + if index + 1 < alignment.coordinates.length() { + let next = alignment.coordinates[index + 1] + let target_step = next.target - point.target + let query_step = next.query - point.query + if target_step == 0 && query_step == 0 { + align_chain_fail("chain coordinate path contains a zero-length segment") + } + if target_step != 0 { + let direction = align_chain_sign(target_step) + if target_direction == 0 { + target_direction = direction + } else if target_direction != direction { + align_chain_fail("target coordinates must be monotonic") + } + } + if query_step != 0 { + let direction = align_chain_sign(query_step) + if query_direction == 0 { + query_direction = direction + } else if query_direction != direction { + align_chain_fail("query coordinates must be monotonic") + } + } + if target_step != 0 && query_step != 0 { + if align_chain_abs(target_step) != align_chain_abs(query_step) { + align_chain_fail( + "aligned target and query steps must have equal lengths", + ) + } + aligned = aligned + align_chain_abs(target_step) + } + } + } + if target_direction == 0 || query_direction == 0 || aligned == 0 { + align_chain_fail("chain path must contain at least one aligned segment") + } +} + +///| +/// Construct one absolute coordinate. +pub fn AlignChainCoordinate::create( + target : Int, + query : Int, +) -> AlignChainCoordinate raise AlignChainError { + if target < 0 || query < 0 { + align_chain_fail("chain coordinates must be non-negative") + } + AlignChainCoordinate::{ target, query } +} + +///| +/// Construct one chain block. +pub fn AlignChainBlock::create( + size : Int, + target_gap? : Int = 0, + query_gap? : Int = 0, +) -> AlignChainBlock raise AlignChainError { + if size < 0 || target_gap < 0 || query_gap < 0 { + align_chain_fail("chain block values must be non-negative") + } + AlignChainBlock::{ size, target_gap, query_gap } +} + +///| +/// Construct and validate a coordinate-aware chain alignment. +pub fn AlignChainAlignment::create( + target_id : String, + target_size : Int, + query_id : String, + query_size : Int, + coordinates : Array[AlignChainCoordinate], + score? : Double = 0.0, + chain_id? : String = "", +) -> AlignChainAlignment raise AlignChainError { + let alignment = AlignChainAlignment::{ + target_id, + target_size, + query_id, + query_size, + score, + chain_id, + coordinates: align_chain_copy_coordinates(coordinates), + } + align_chain_validate_alignment(alignment) + alignment +} + +///| +fn align_chain_oriented_coordinate( + coordinate : Int, + size : Int, + strand : String, +) -> Int { + if strand == "+" { + coordinate + } else { + size - coordinate + } +} + +///| +/// Build an alignment from normalized UCSC header coordinates and blocks. +pub fn align_chain_from_blocks( + target_id : String, + target_size : Int, + target_strand : String, + target_start : Int, + target_end : Int, + query_id : String, + query_size : Int, + query_strand : String, + query_start : Int, + query_end : Int, + blocks : Array[AlignChainBlock], + score? : Double = 0.0, + chain_id? : String = "", +) -> AlignChainAlignment raise AlignChainError { + align_chain_validate_token(target_id, "target identifier", false) + align_chain_validate_token(query_id, "query identifier", false) + align_chain_validate_token(chain_id, "chain identifier", true) + align_chain_validate_strand(target_strand, "target") + align_chain_validate_strand(query_strand, "query") + align_chain_validate_interval(target_start, target_end, target_size, "target") + align_chain_validate_interval(query_start, query_end, query_size, "query") + if score.is_nan() || score.abs() > 1.0e300 { + align_chain_fail("chain score must be finite") + } + if blocks.length() == 0 { + align_chain_fail("chain record must contain at least one block row") + } + let relative : Array[AlignChainCoordinate] = [ + AlignChainCoordinate::{ target: 0, query: 0 }, + ] + let mut target_position = 0 + let mut query_position = 0 + let mut aligned = 0 + for index in 0.. 0 { + target_position = target_position + block.size + query_position = query_position + block.size + relative.push(AlignChainCoordinate::{ + target: target_position, + query: query_position, + }) + aligned = aligned + block.size + } + if block.target_gap > 0 { + target_position = target_position + block.target_gap + relative.push(AlignChainCoordinate::{ + target: target_position, + query: query_position, + }) + } + if block.query_gap > 0 { + query_position = query_position + block.query_gap + relative.push(AlignChainCoordinate::{ + target: target_position, + query: query_position, + }) + } + if target_position > target_end - target_start { + align_chain_fail("chain blocks exceed the declared target span") + } + if query_position > query_end - query_start { + align_chain_fail("chain blocks exceed the declared query span") + } + } + if aligned == 0 { + align_chain_fail("chain record must contain aligned bases") + } + if target_position != target_end - target_start { + align_chain_fail("chain blocks do not match the declared target span") + } + if query_position != query_end - query_start { + align_chain_fail("chain blocks do not match the declared query span") + } + let coordinates : Array[AlignChainCoordinate] = [] + for point in relative { + let target_oriented = target_start + point.target + let query_oriented = query_start + point.query + coordinates.push(AlignChainCoordinate::{ + target: align_chain_oriented_coordinate( + target_oriented, target_size, target_strand, + ), + query: align_chain_oriented_coordinate( + query_oriented, query_size, query_strand, + ), + }) + } + AlignChainAlignment::create( + target_id, + target_size, + query_id, + query_size, + coordinates, + score~, + chain_id~, + ) +} + +///| +fn align_chain_strip_cr(value : String) -> String { + if value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '\r'.to_int() { + value[0:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn align_chain_split_lines(value : String) -> Array[String] { + let lines : Array[String] = [] + let mut start = 0 + for index in 0.. 0 { + lines.push("") + } + lines +} + +///| +fn align_chain_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} + +///| +fn align_chain_trim(value : String) -> String { + let mut start = 0 + while start < value.length() && + align_chain_is_whitespace(value.unsafe_get(start).to_int()) { + start = start + 1 + } + let mut end = value.length() + while end > start && + align_chain_is_whitespace(value.unsafe_get(end - 1).to_int()) { + end = end - 1 + } + value[start:end].to_owned() +} + +///| +fn align_chain_split_whitespace(value : String) -> Array[String] { + let words : Array[String] = [] + let mut index = 0 + while index < value.length() { + while index < value.length() && + align_chain_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= value.length() { + break + } + let start = index + while index < value.length() && + !align_chain_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + words.push(value[start:index].to_owned()) + } + words +} + +///| +fn align_chain_parse_uint( + value : String, + label : String, +) -> Int raise AlignChainError { + if value.length() == 0 { + align_chain_fail(label + " must be an unsigned integer") + } + let mut result = 0 + for index in 0.. '9'.to_int() { + align_chain_fail(label + " must be an unsigned integer") + } + let digit = code - '0'.to_int() + if result > (2147483647 - digit) / 10 { + align_chain_fail(label + " exceeds the supported integer range") + } + result = result * 10 + digit + } + result +} + +///| +fn align_chain_is_score(value : String) -> Bool { + if value.length() == 0 { + return false + } + let mut index = 0 + if value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() { + index = index + 1 + } + let mut digits = 0 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + if index < value.length() && value.unsafe_get(index).to_int() == '.'.to_int() { + index = index + 1 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + } + if digits == 0 { + return false + } + if index < value.length() && + ( + value.unsafe_get(index).to_int() == 'e'.to_int() || + value.unsafe_get(index).to_int() == 'E'.to_int() + ) { + index = index + 1 + if index < value.length() && + ( + value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() + ) { + index = index + 1 + } + let exponent_start = index + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + index = index + 1 + } + if index == exponent_start { + return false + } + } + index == value.length() +} + +///| +fn align_chain_parse_score(value : String) -> Double raise AlignChainError { + if !align_chain_is_score(value) { + align_chain_fail("chain score must be a number") + } + let parse_value = if value.length() > 0 && + value.unsafe_get(0).to_int() == '+'.to_int() { + value[1:value.length()].to_owned() + } else { + value + } + let number = match parse_double(parse_value) { + Some(number) => number + None => { + align_chain_fail("chain score must be a number") + 0.0 + } + } + if number.is_nan() || number.abs() > 1.0e300 { + align_chain_fail("chain score must be finite") + } + number +} + +///| +fn align_chain_parse_record( + lines : Array[String], + start : Int, +) -> (AlignChainAlignment, Int) raise AlignChainError { + let header = align_chain_split_whitespace(align_chain_trim(lines[start])) + if header.length() != 12 && header.length() != 13 { + align_chain_fail( + "chain header must contain exactly 12 or 13 whitespace-delimited fields", + ) + } + if header[0] != "chain" { + align_chain_fail("chain record must start with the 'chain' keyword") + } + let score = align_chain_parse_score(header[1]) + let target_id = header[2] + let target_size = align_chain_parse_uint(header[3], "target size") + let target_strand = header[4] + let target_start = align_chain_parse_uint(header[5], "target start") + let target_end = align_chain_parse_uint(header[6], "target end") + let query_id = header[7] + let query_size = align_chain_parse_uint(header[8], "query size") + let query_strand = header[9] + let query_start = align_chain_parse_uint(header[10], "query start") + let query_end = align_chain_parse_uint(header[11], "query end") + let chain_id = if header.length() == 13 { header[12] } else { "" } + align_chain_validate_strand(target_strand, "target") + align_chain_validate_strand(query_strand, "query") + align_chain_validate_interval(target_start, target_end, target_size, "target") + align_chain_validate_interval(query_start, query_end, query_size, "query") + let blocks : Array[AlignChainBlock] = [] + let mut index = start + 1 + let mut terminated = false + while index < lines.length() { + let line = align_chain_trim(lines[index]) + if line.length() == 0 { + align_chain_fail("blank line found before the final chain block") + } + let words = align_chain_split_whitespace(line) + if words.length() > 0 && words[0] == "chain" { + align_chain_fail("chain record is missing its final size-only block") + } + if words.length() != 1 && words.length() != 3 { + align_chain_fail("chain block must contain one or three integer fields") + } + let size = align_chain_parse_uint(words[0], "chain block size") + if words.length() == 1 { + blocks.push(AlignChainBlock::{ size, target_gap: 0, query_gap: 0 }) + index = index + 1 + terminated = true + break + } + let target_gap = align_chain_parse_uint(words[1], "target gap") + let query_gap = align_chain_parse_uint(words[2], "query gap") + if target_gap == 0 && query_gap == 0 { + align_chain_fail("non-final chain block must contain a non-zero gap") + } + blocks.push(AlignChainBlock::{ size, target_gap, query_gap }) + index = index + 1 + } + if !terminated { + align_chain_fail("chain record is missing its final size-only block") + } + let alignment = align_chain_from_blocks( + target_id, + target_size, + target_strand, + target_start, + target_end, + query_id, + query_size, + query_strand, + query_start, + query_end, + blocks, + score~, + chain_id~, + ) + (alignment, index) +} + +///| +/// Parse all chain records in a string. +pub fn align_chain_parse_all( + text : String, +) -> Array[AlignChainAlignment] raise AlignChainError { + let lines = align_chain_split_lines(text) + let alignments : Array[AlignChainAlignment] = [] + let mut index = 0 + while index < lines.length() { + let line = align_chain_trim(lines[index]) + if line.length() == 0 { + index = index + 1 + continue + } + if align_chain_split_whitespace(line)[0] != "chain" { + align_chain_fail( + "unexpected content before chain record at line " + + (index + 1).to_string(), + ) + } + let (alignment, next) = align_chain_parse_record(lines, index) + alignments.push(alignment) + index = next + } + alignments +} + +///| +/// Parse exactly one chain record. +pub fn align_chain_parse( + text : String, +) -> AlignChainAlignment raise AlignChainError { + let alignments = align_chain_parse_all(text) + if alignments.length() == 0 { + align_chain_fail("chain input contains no records") + } + if alignments.length() != 1 { + align_chain_fail("expected exactly one chain record") + } + alignments[0] +} + +///| +/// Return the target strand encoded by the coordinate path. +pub fn AlignChainAlignment::target_strand(self : AlignChainAlignment) -> String { + if self.coordinates[self.coordinates.length() - 1].target > + self.coordinates[0].target { + "+" + } else { + "-" + } +} + +///| +/// Return the query strand encoded by the coordinate path. +pub fn AlignChainAlignment::query_strand(self : AlignChainAlignment) -> String { + if self.coordinates[self.coordinates.length() - 1].query > + self.coordinates[0].query { + "+" + } else { + "-" + } +} + +///| +/// Return the normalized target start. +pub fn AlignChainAlignment::target_start(self : AlignChainAlignment) -> Int { + align_chain_min( + self.coordinates[0].target, + self.coordinates[self.coordinates.length() - 1].target, + ) +} + +///| +/// Return the normalized target end. +pub fn AlignChainAlignment::target_end(self : AlignChainAlignment) -> Int { + align_chain_max( + self.coordinates[0].target, + self.coordinates[self.coordinates.length() - 1].target, + ) +} + +///| +/// Return the normalized query start. +pub fn AlignChainAlignment::query_start(self : AlignChainAlignment) -> Int { + align_chain_min( + self.coordinates[0].query, + self.coordinates[self.coordinates.length() - 1].query, + ) +} + +///| +/// Return the normalized query end. +pub fn AlignChainAlignment::query_end(self : AlignChainAlignment) -> Int { + align_chain_max( + self.coordinates[0].query, + self.coordinates[self.coordinates.length() - 1].query, + ) +} + +///| +fn align_chain_target_oriented( + alignment : AlignChainAlignment, + coordinate : Int, +) -> Int { + if alignment.target_strand() == "+" { + coordinate + } else { + alignment.target_size - coordinate + } +} + +///| +fn align_chain_query_oriented( + alignment : AlignChainAlignment, + coordinate : Int, +) -> Int { + if alignment.query_strand() == "+" { + coordinate + } else { + alignment.query_size - coordinate + } +} + +///| +/// Convert the coordinate path to canonical chain block rows. +pub fn AlignChainAlignment::blocks( + self : AlignChainAlignment, +) -> Array[AlignChainBlock] raise AlignChainError { + align_chain_validate_alignment(self) + let blocks : Array[AlignChainBlock] = [] + let mut size = 0 + let mut target_gap = 0 + let mut query_gap = 0 + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = align_chain_target_oriented(self, next.target) - + align_chain_target_oriented(self, current.target) + let query_step = align_chain_query_oriented(self, next.query) - + align_chain_query_oriented(self, current.query) + if target_step < 0 || query_step < 0 { + align_chain_fail("chain coordinate path is not monotonic on its strand") + } + if target_step > 0 && query_step > 0 { + if target_step != query_step { + align_chain_fail( + "aligned target and query steps must have equal lengths", + ) + } + if target_gap > 0 || query_gap > 0 { + blocks.push(AlignChainBlock::{ size, target_gap, query_gap }) + size = target_step + target_gap = 0 + query_gap = 0 + } else { + size = size + target_step + } + } else if target_step > 0 { + target_gap = target_gap + target_step + } else if query_step > 0 { + query_gap = query_gap + query_step + } + } + if target_gap > 0 || query_gap > 0 { + blocks.push(AlignChainBlock::{ size, target_gap, query_gap }) + size = 0 + } + blocks.push(AlignChainBlock::{ size, target_gap: 0, query_gap: 0 }) + blocks +} + +///| +fn align_chain_header_target_start(alignment : AlignChainAlignment) -> Int { + align_chain_target_oriented(alignment, alignment.coordinates[0].target) +} + +///| +fn align_chain_header_target_end(alignment : AlignChainAlignment) -> Int { + align_chain_target_oriented( + alignment, + alignment.coordinates[alignment.coordinates.length() - 1].target, + ) +} + +///| +fn align_chain_header_query_start(alignment : AlignChainAlignment) -> Int { + align_chain_query_oriented(alignment, alignment.coordinates[0].query) +} + +///| +fn align_chain_header_query_end(alignment : AlignChainAlignment) -> Int { + align_chain_query_oriented( + alignment, + alignment.coordinates[alignment.coordinates.length() - 1].query, + ) +} + +///| +/// Serialize one alignment as a canonical UCSC chain record. +pub fn align_chain_write( + alignment : AlignChainAlignment, +) -> String raise AlignChainError { + align_chain_validate_alignment(alignment) + let output = StringBuilder::new() + output.write_string("chain ") + output.write_string(alignment.score.to_string()) + output.write_char(' ') + output.write_string(alignment.target_id) + output.write_char(' ') + output.write_string(alignment.target_size.to_string()) + output.write_char(' ') + output.write_string(alignment.target_strand()) + output.write_char(' ') + output.write_string(align_chain_header_target_start(alignment).to_string()) + output.write_char(' ') + output.write_string(align_chain_header_target_end(alignment).to_string()) + output.write_char(' ') + output.write_string(alignment.query_id) + output.write_char(' ') + output.write_string(alignment.query_size.to_string()) + output.write_char(' ') + output.write_string(alignment.query_strand()) + output.write_char(' ') + output.write_string(align_chain_header_query_start(alignment).to_string()) + output.write_char(' ') + output.write_string(align_chain_header_query_end(alignment).to_string()) + if alignment.chain_id.length() > 0 { + output.write_char(' ') + output.write_string(alignment.chain_id) + } + output.write_char('\n') + let blocks = alignment.blocks() + for index in 0.. String raise AlignChainError { + let output = StringBuilder::new() + for alignment in alignments { + output.write_string(align_chain_write(alignment)) + } + output.to_string() +} + +///| +/// Return one operation per coordinate-path segment. +/// +/// `M` consumes both rows, `D` consumes target only, and `I` consumes query +/// only. +pub fn AlignChainAlignment::operation_path( + self : AlignChainAlignment, +) -> String { + let output = StringBuilder::new() + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = next.target - current.target + let query_step = next.query - current.query + if target_step != 0 && query_step != 0 { + output.write_char('M') + } else if target_step != 0 { + output.write_char('D') + } else { + output.write_char('I') + } + } + output.to_string() +} + +///| +/// Count the number of alignment columns represented by the path. +pub fn AlignChainAlignment::alignment_length(self : AlignChainAlignment) -> Int { + let mut length = 0 + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + length = length + + align_chain_max( + align_chain_abs(next.target - current.target), + align_chain_abs(next.query - current.query), + ) + } + length +} + +///| +/// Count aligned bases. +pub fn AlignChainAlignment::aligned_bases(self : AlignChainAlignment) -> Int { + let mut aligned = 0 + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = align_chain_abs(next.target - current.target) + let query_step = align_chain_abs(next.query - current.query) + if target_step > 0 && query_step > 0 { + aligned = aligned + target_step + } + } + aligned +} + +///| +/// Compute coordinate-only alignment counts. +pub fn AlignChainAlignment::counts( + self : AlignChainAlignment, +) -> AlignChainCounts { + let mut aligned = 0 + let mut target_gap_bases = 0 + let mut query_gap_bases = 0 + let mut target_gap_opens = 0 + let mut query_gap_opens = 0 + let mut aligned_blocks = 0 + let mut previous = 'X' + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = align_chain_abs(next.target - current.target) + let query_step = align_chain_abs(next.query - current.query) + if target_step > 0 && query_step > 0 { + aligned = aligned + target_step + if previous != 'M' { + aligned_blocks = aligned_blocks + 1 + } + previous = 'M' + } else if target_step > 0 { + target_gap_bases = target_gap_bases + target_step + if previous != 'D' { + target_gap_opens = target_gap_opens + 1 + } + previous = 'D' + } else { + query_gap_bases = query_gap_bases + query_step + if previous != 'I' { + query_gap_opens = query_gap_opens + 1 + } + previous = 'I' + } + } + AlignChainCounts::{ + aligned, + target_gap_bases, + query_gap_bases, + gap_columns: target_gap_bases + query_gap_bases, + target_gap_opens, + query_gap_opens, + aligned_blocks, + path_segments: self.coordinates.length() - 1, + } +} + +///| +fn align_chain_map_on_segment( + source : Int, + source_start : Int, + source_end : Int, + destination_start : Int, + destination_end : Int, +) -> Int? { + let source_step = source_end - source_start + let destination_step = destination_end - destination_start + if source_step == 0 || + destination_step == 0 || + align_chain_abs(source_step) != align_chain_abs(destination_step) { + return None + } + let offset = if source_step > 0 { + if source < source_start || source >= source_end { + return None + } + source - source_start + } else { + if source < source_end || source >= source_start { + return None + } + source_start - 1 - source + } + if destination_step > 0 { + Some(destination_start + offset) + } else { + Some(destination_start - 1 - offset) + } +} + +///| +/// Map one target residue to the query, returning `None` in an unaligned gap. +pub fn AlignChainAlignment::map_target_position( + self : AlignChainAlignment, + position : Int, +) -> Int? { + if position < 0 || position >= self.target_size { + return None + } + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + match + align_chain_map_on_segment( + position, + current.target, + next.target, + current.query, + next.query, + ) { + Some(mapped) => return Some(mapped) + None => () + } + } + None +} + +///| +/// Map one query residue to the target, returning `None` in an unaligned gap. +pub fn AlignChainAlignment::map_query_position( + self : AlignChainAlignment, + position : Int, +) -> Int? { + if position < 0 || position >= self.query_size { + return None + } + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + match + align_chain_map_on_segment( + position, + current.query, + next.query, + current.target, + next.target, + ) { + Some(mapped) => return Some(mapped) + None => () + } + } + None +} + +///| +fn align_chain_range_mapping( + alignment : AlignChainAlignment, + target_start : Int, + target_end : Int, + query_start : Int, + query_end : Int, +) -> AlignChainRangeMapping { + AlignChainRangeMapping::{ + target_start, + target_end, + query_start, + query_end, + target_strand: alignment.target_strand(), + query_strand: alignment.query_strand(), + } +} + +///| +/// Map a target interval to all aligned query pieces. +/// +/// Target-only gap portions are omitted, so one input interval may yield +/// multiple output pieces. +pub fn AlignChainAlignment::map_target_range( + self : AlignChainAlignment, + start : Int, + end : Int, +) -> Array[AlignChainRangeMapping] raise AlignChainError { + if start < 0 || end <= start || end > self.target_size { + align_chain_fail("target range must satisfy 0 <= start < end <= size") + } + let mappings : Array[AlignChainRangeMapping] = [] + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_low = align_chain_min(current.target, next.target) + let target_high = align_chain_max(current.target, next.target) + if current.target != next.target && current.query != next.query { + let piece_start = align_chain_max(start, target_low) + let piece_end = align_chain_min(end, target_high) + if piece_start < piece_end { + let query_first = match self.map_target_position(piece_start) { + Some(value) => value + None => { + align_chain_fail("internal target range mapping failure") + 0 + } + } + let query_last = match self.map_target_position(piece_end - 1) { + Some(value) => value + None => { + align_chain_fail("internal target range mapping failure") + 0 + } + } + mappings.push( + align_chain_range_mapping( + self, + piece_start, + piece_end, + align_chain_min(query_first, query_last), + align_chain_max(query_first, query_last) + 1, + ), + ) + } + } + } + mappings +} + +///| +/// Map a query interval to all aligned target pieces. +pub fn AlignChainAlignment::map_query_range( + self : AlignChainAlignment, + start : Int, + end : Int, +) -> Array[AlignChainRangeMapping] raise AlignChainError { + if start < 0 || end <= start || end > self.query_size { + align_chain_fail("query range must satisfy 0 <= start < end <= size") + } + let mappings : Array[AlignChainRangeMapping] = [] + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let query_low = align_chain_min(current.query, next.query) + let query_high = align_chain_max(current.query, next.query) + if current.target != next.target && current.query != next.query { + let piece_start = align_chain_max(start, query_low) + let piece_end = align_chain_min(end, query_high) + if piece_start < piece_end { + let target_first = match self.map_query_position(piece_start) { + Some(value) => value + None => { + align_chain_fail("internal query range mapping failure") + 0 + } + } + let target_last = match self.map_query_position(piece_end - 1) { + Some(value) => value + None => { + align_chain_fail("internal query range mapping failure") + 0 + } + } + mappings.push( + align_chain_range_mapping( + self, + align_chain_min(target_first, target_last), + align_chain_max(target_first, target_last) + 1, + piece_start, + piece_end, + ), + ) + } + } + } + mappings +} + +///| +/// Materialize aligned residue pairs, subject to a caller-controlled limit. +pub fn AlignChainAlignment::aligned_pairs( + self : AlignChainAlignment, + limit? : Int = 1000000, +) -> Array[AlignChainPair] raise AlignChainError { + if limit < 0 { + align_chain_fail("aligned-pair limit must be non-negative") + } + let aligned = self.aligned_bases() + if aligned > limit { + align_chain_fail("aligned-pair count exceeds the requested limit") + } + let pairs : Array[AlignChainPair] = [] + for index in 0..<(self.coordinates.length() - 1) { + let current = self.coordinates[index] + let next = self.coordinates[index + 1] + let target_step = next.target - current.target + let query_step = next.query - current.query + if target_step != 0 && query_step != 0 { + let length = align_chain_abs(target_step) + let target_direction = align_chain_sign(target_step) + let query_direction = align_chain_sign(query_step) + for offset in 0.. 0 { + current.target + offset + } else { + current.target - 1 - offset + }, + query: if query_direction > 0 { + current.query + offset + } else { + current.query - 1 - offset + }, + }) + } + } + } + pairs +} + +///| +/// Swap target and query while preserving score, ID, and path orientation. +pub fn AlignChainAlignment::invert( + self : AlignChainAlignment, +) -> AlignChainAlignment raise AlignChainError { + let coordinates : Array[AlignChainCoordinate] = [] + for point in self.coordinates { + coordinates.push(AlignChainCoordinate::{ + target: point.query, + query: point.target, + }) + } + AlignChainAlignment::create( + self.query_id, + self.query_size, + self.target_id, + self.target_size, + coordinates, + score=self.score, + chain_id=self.chain_id, + ) +} + +///| +/// Find the first alignment with the requested chain ID. +pub fn align_chain_find_by_id( + alignments : Array[AlignChainAlignment], + chain_id : String, +) -> AlignChainAlignment? { + for alignment in alignments { + if alignment.chain_id == chain_id { + return Some(alignment) + } + } + None +} + +///| +/// Return alignments whose target span overlaps a half-open interval. +pub fn align_chain_target_overlaps( + alignments : Array[AlignChainAlignment], + target_id : String, + start : Int, + end : Int, +) -> Array[AlignChainAlignment] raise AlignChainError { + if start < 0 || end <= start { + align_chain_fail("overlap query must satisfy 0 <= start < end") + } + let matches : Array[AlignChainAlignment] = [] + for alignment in alignments { + if alignment.target_id == target_id && + alignment.target_start() < end && + alignment.target_end() > start { + matches.push(alignment) + } + } + align_chain_copy_alignments(matches) +} + +///| +/// Return a concise alignment summary. +pub fn AlignChainAlignment::summary(self : AlignChainAlignment) -> String { + let counts = self.counts() + "AlignChainAlignment(target=" + + self.target_id + + ":" + + self.target_start().to_string() + + "-" + + self.target_end().to_string() + + self.target_strand() + + ", query=" + + self.query_id + + ":" + + self.query_start().to_string() + + "-" + + self.query_end().to_string() + + self.query_strand() + + ", aligned=" + + counts.aligned.to_string() + + ", gaps=" + + counts.gap_columns.to_string() + + ")" +} + +///| +/// Return a compact two-record fixture based on Biopython's chain examples. +pub fn align_chain_example_text() -> String { + "chain 176 chr3 198295559 + 42530895 42532606 NR_046654.1 181 - 0 181 1\n" + + "63\t1062\t0\n" + + "75\t468\t0\n" + + "43\n\n" + + "chain 500 chr1 1000 - 100 200 queryA 500 + 20 130 reverse-target\n" + + "40\t10\t20\n" + + "50\n" +} diff --git a/src/align_clustal.mbt b/src/align_clustal.mbt new file mode 100644 index 00000000..0bccc917 --- /dev/null +++ b/src/align_clustal.mbt @@ -0,0 +1,1363 @@ +// Biopython-compatible Bio.Align.clustal support. +// +// This module is separate from clustal_io.mbt. The older module exposes the +// historical AlignIO-style MultipleSeqAlignment API, while this file retains +// modern file metadata, column annotations, and coordinate-aware rows. + +///| +/// Error raised for malformed CLUSTAL data or invalid alignment operations. +pub suberror AlignClustalError { + AlignClustalError(String) +} + +///| +/// File-level metadata from a CLUSTAL alignment. +/// +/// Known generators follow Biopython 1.86: CLUSTAL, PROBCONS, MUSCLE, +/// MSAPROBS, Kalign, and Biopython. +pub struct AlignClustalMetadata { + program : String + version : String + header : String +} derive(Eq, Debug) + +///| +/// One coordinate-bearing row in a CLUSTAL alignment. +pub struct AlignClustalSequence { + id : String + sequence : String + aligned_sequence : String +} derive(Eq, Debug) + +///| +/// A modern CLUSTAL multiple sequence alignment. +/// +/// `consensus` retains the `*`, `:`, `.`, and space column annotation when it +/// is present. Sparse source annotations are padded with spaces. +pub struct AlignClustalAlignment { + metadata : AlignClustalMetadata + sequences : Array[AlignClustalSequence] + consensus : String? +} derive(Eq, Debug) + +///| +/// Pairwise or all-pairs alignment statistics. +pub struct AlignClustalCounts { + pairs : Int + aligned : Int + identities : Int + mismatches : Int + gap_columns : Int + double_gap_columns : Int + gap_opens : Int +} derive(Eq, Debug) + +///| +/// One parsed interleaved block. +priv struct AlignClustalBlock { + ids : Array[String] + segments : Array[String] + residue_counts : Array[Int?] + consensus : String? + width : Int +} + +///| +/// Construct validated CLUSTAL file metadata. +pub fn AlignClustalMetadata::create( + program : String, + version? : String = "", + header? : String = "", +) -> AlignClustalMetadata raise AlignClustalError { + align_clustal_validate_program(program) + align_clustal_validate_version(version) + let canonical_header = if header.length() == 0 { + align_clustal_canonical_header(program, version) + } else { + align_clustal_validate_plain_text(header, "CLUSTAL header") + header + } + AlignClustalMetadata::{ program, version, header: canonical_header } +} + +///| +/// Construct one validated CLUSTAL row. +pub fn AlignClustalSequence::create( + id : String, + aligned_sequence : String, +) -> AlignClustalSequence raise AlignClustalError { + align_clustal_validate_id(id) + if aligned_sequence.length() == 0 { + raise AlignClustalError("CLUSTAL aligned sequence must not be empty") + } + align_clustal_validate_segment(aligned_sequence) + let sequence = align_clustal_remove_gaps(aligned_sequence) + if sequence.length() == 0 { + raise AlignClustalError( + "CLUSTAL sequence row must contain at least one residue", + ) + } + AlignClustalSequence::{ id, sequence, aligned_sequence } +} + +///| +/// Construct a validated modern CLUSTAL alignment. +pub fn AlignClustalAlignment::create( + metadata : AlignClustalMetadata, + sequences : Array[AlignClustalSequence], + consensus? : String? = None, +) -> AlignClustalAlignment raise AlignClustalError { + let copied : Array[AlignClustalSequence] = [] + for sequence in sequences { + copied.push(sequence) + } + let alignment = AlignClustalAlignment::{ + metadata, + sequences: copied, + consensus, + } + align_clustal_validate_alignment(alignment) + alignment +} + +///| +/// Construct a coordinate-aware alignment from printed rows. +pub fn align_clustal_from_aligned( + ids : Array[String], + aligned_sequences : Array[String], + program? : String = "Biopython", + version? : String = "", + consensus? : String? = None, +) -> AlignClustalAlignment raise AlignClustalError { + if ids.length() == 0 { + raise AlignClustalError( + "CLUSTAL alignment must contain at least one sequence", + ) + } + if ids.length() != aligned_sequences.length() { + raise AlignClustalError( + "CLUSTAL identifiers and aligned rows must have equal lengths", + ) + } + let sequences : Array[AlignClustalSequence] = [] + for index = 0; index < ids.length(); index = index + 1 { + sequences.push( + AlignClustalSequence::create(ids[index], aligned_sequences[index]), + ) + } + let metadata = AlignClustalMetadata::create(program, version~) + AlignClustalAlignment::create(metadata, sequences, consensus~) +} + +///| +/// Parse one CLUSTAL alignment. +/// +/// Blocks must repeat the first block's unique identifiers in the same order. +/// Optional right-hand residue counts are validated as cumulative ungapped +/// counts. Consensus lines may be absent from some blocks; missing portions +/// are represented by spaces in the retained column annotation. +pub fn align_clustal_parse( + text : String, +) -> AlignClustalAlignment raise AlignClustalError { + let lines = align_clustal_lines(text) + if lines.length() == 0 || lines[0].trim().length() == 0 { + raise AlignClustalError("Empty CLUSTAL input") + } + let metadata = align_clustal_parse_header(lines[0]) + let blocks : Array[AlignClustalBlock] = [] + let mut block_lines : Array[String] = [] + for line_index = 1; line_index < lines.length(); line_index = line_index + 1 { + let line = lines[line_index] + if line.trim().length() == 0 { + if line.length() >= 10 && block_lines.length() > 0 { + block_lines.push(line) + } else if block_lines.length() > 0 { + blocks.push(align_clustal_parse_block(block_lines)) + block_lines = [] + } + } else { + block_lines.push(line) + } + } + if block_lines.length() > 0 { + blocks.push(align_clustal_parse_block(block_lines)) + } + if blocks.length() == 0 { + raise AlignClustalError("CLUSTAL input contains no alignment rows") + } + + let ids = blocks[0].ids + let builders : Array[StringBuilder] = [] + let cumulative = Array::make(ids.length(), 0) + let seen_ids : Array[String] = [] + for id in ids { + align_clustal_validate_id(id) + if align_clustal_find_string(seen_ids, id) >= 0 { + raise AlignClustalError( + "Duplicated CLUSTAL sequence identifier '" + id + "'", + ) + } + seen_ids.push(id) + builders.push(StringBuilder::new()) + } + + let consensus_builder = StringBuilder::new() + let mut has_consensus = false + for block_index = 0 + block_index < blocks.length() + block_index = block_index + 1 { + let block = blocks[block_index] + if block.ids.length() != ids.length() { + raise AlignClustalError( + "CLUSTAL block sequence count differs from the first block", + ) + } + for row = 0; row < ids.length(); row = row + 1 { + if block.ids[row] != ids[row] { + raise AlignClustalError( + "CLUSTAL block identifiers must repeat in the same order", + ) + } + builders[row].write_string(block.segments[row]) + cumulative[row] = cumulative[row] + + align_clustal_residue_count(block.segments[row]) + match block.residue_counts[row] { + Some(declared) => + if declared != cumulative[row] { + raise AlignClustalError( + "CLUSTAL cumulative residue count mismatch for '" + ids[row] + "'", + ) + } + None => () + } + } + match block.consensus { + Some(annotation) => { + has_consensus = true + consensus_builder.write_string(annotation) + } + None => consensus_builder.write_string(" ".repeat(block.width)) + } + } + + let sequences : Array[AlignClustalSequence] = [] + for row = 0; row < ids.length(); row = row + 1 { + sequences.push( + AlignClustalSequence::create(ids[row], builders[row].to_string()), + ) + } + let consensus : String? = if has_consensus { + Some(consensus_builder.to_string()) + } else { + None + } + AlignClustalAlignment::create(metadata, sequences, consensus~) +} + +///| +/// Write canonical CLUSTAL text. +/// +/// Biopython-compatible defaults use 50 alignment columns and a 36-character +/// name field. Identifiers are never silently truncated; names longer than 30 +/// characters require a larger `name_width`. +pub fn align_clustal_write( + alignment : AlignClustalAlignment, + block_width? : Int = 50, + name_width? : Int = 36, + include_counts? : Bool = false, + include_consensus? : Bool = true, +) -> String raise AlignClustalError { + align_clustal_validate_alignment(alignment) + if block_width <= 0 { + raise AlignClustalError("CLUSTAL block width must be positive") + } + if name_width <= 1 { + raise AlignClustalError("CLUSTAL name width must be greater than one") + } + for sequence in alignment.sequences { + if sequence.id.length() >= name_width { + raise AlignClustalError( + "CLUSTAL name width must exceed every identifier length", + ) + } + if name_width == 36 && sequence.id.length() > 30 { + raise AlignClustalError( + "CLUSTAL canonical writer does not silently truncate identifiers " + + "longer than 30 characters", + ) + } + } + let output = StringBuilder::new() + output.write_string( + align_clustal_canonical_header( + alignment.metadata.program, + alignment.metadata.version, + ), + ) + output.write_string("\n\n\n") + let cumulative = Array::make(alignment.num_sequences(), 0) + let width = alignment.alignment_length() + let mut start = 0 + while start < width { + let stop = if start + block_width < width { + start + block_width + } else { + width + } + for row = 0; row < alignment.num_sequences(); row = row + 1 { + let sequence = alignment.sequences[row] + let segment = sequence.aligned_sequence[start:stop].to_owned() + output.write_string(sequence.id) + output.write_string(" ".repeat(name_width - sequence.id.length())) + output.write_string(segment) + cumulative[row] = cumulative[row] + align_clustal_residue_count(segment) + if include_counts { + output.write_char(' ') + output.write_string(cumulative[row].to_string()) + } + output.write_char('\n') + } + if include_consensus { + match alignment.consensus { + Some(consensus) => { + output.write_string(" ".repeat(name_width)) + output.write_string(consensus[start:stop].to_owned()) + output.write_char('\n') + } + None => () + } + } + output.write_char('\n') + start = stop + } + output.write_char('\n') + let result = output.to_string() + ignore(align_clustal_parse(result)) + result +} + +///| +/// Return the number of sequence rows. +pub fn AlignClustalAlignment::num_sequences( + self : AlignClustalAlignment, +) -> Int { + self.sequences.length() +} + +///| +/// Return the alignment width including gaps. +pub fn AlignClustalAlignment::alignment_length( + self : AlignClustalAlignment, +) -> Int { + if self.sequences.length() == 0 { + 0 + } else { + self.sequences[0].aligned_sequence.length() + } +} + +///| +/// Locate a row by exact identifier. +pub fn AlignClustalAlignment::find_sequence( + self : AlignClustalAlignment, + id : String, +) -> Int? { + for index = 0; index < self.sequences.length(); index = index + 1 { + if self.sequences[index].id == id { + return Some(index) + } + } + None +} + +///| +/// Return one printed alignment column. +pub fn AlignClustalAlignment::column( + self : AlignClustalAlignment, + column : Int, +) -> String? { + if column < 0 || column >= self.alignment_length() { + return None + } + let result = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + result.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(result.to_string()) +} + +///| +/// Map a zero-based ungapped row position to an alignment column. +pub fn AlignClustalAlignment::sequence_position_to_column( + self : AlignClustalAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + position < 0 || + position >= self.sequences[row].sequence.length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + let mut coordinate = 0 + for column = 0; column < aligned.length(); column = column + 1 { + if aligned.unsafe_get(column).to_int() != '-'.to_int() { + if coordinate == position { + return Some(column) + } + coordinate = coordinate + 1 + } + } + None +} + +///| +/// Map one alignment column to a zero-based ungapped row position. +pub fn AlignClustalAlignment::column_to_sequence_position( + self : AlignClustalAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.alignment_length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + if aligned.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut coordinate = 0 + for index = 0; index < column; index = index + 1 { + if aligned.unsafe_get(index).to_int() != '-'.to_int() { + coordinate = coordinate + 1 + } + } + Some(coordinate) +} + +///| +/// Project one row position through the alignment to another row. +pub fn AlignClustalAlignment::map_position( + self : AlignClustalAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column residue positions for two rows. +pub fn AlignClustalAlignment::aligned_pairs( + self : AlignClustalAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise AlignClustalError { + align_clustal_validate_row(self, first_row) + align_clustal_validate_row(self, second_row) + let result : Array[(Int?, Int?)] = [] + let mut first_position = 0 + let mut second_position = 0 + for column = 0; column < self.alignment_length(); column = column + 1 { + let first_gap = self.sequences[first_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + let second_gap = self.sequences[second_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + result.push( + ( + if first_gap { + None + } else { + Some(first_position) + }, + if second_gap { + None + } else { + Some(second_position) + }, + ), + ) + if !first_gap { + first_position = first_position + 1 + } + if !second_gap { + second_position = second_position + 1 + } + } + result +} + +///| +/// Return a compact Biopython-style coordinate path for all rows. +pub fn AlignClustalAlignment::coordinate_path( + self : AlignClustalAlignment, +) -> Array[Array[Int]] { + let paths : Array[Array[Int]] = [] + let coordinates = Array::make(self.sequences.length(), 0) + for _row = 0; _row < self.sequences.length(); _row = _row + 1 { + paths.push([0]) + } + let width = self.alignment_length() + for column = 0; column < width; column = column + 1 { + for row = 0; row < self.sequences.length(); row = row + 1 { + if self.sequences[row].aligned_sequence.unsafe_get(column).to_int() != + '-'.to_int() { + coordinates[row] = coordinates[row] + 1 + } + } + let boundary = if column + 1 == width { + true + } else { + align_clustal_movement_changes(self, column, column + 1) + } + if boundary { + for row = 0; row < self.sequences.length(); row = row + 1 { + paths[row].push(coordinates[row]) + } + } + } + paths +} + +///| +/// Compute statistics for one pair of rows. +pub fn AlignClustalAlignment::pair_counts( + self : AlignClustalAlignment, + first_row : Int, + second_row : Int, +) -> AlignClustalCounts raise AlignClustalError { + align_clustal_validate_row(self, first_row) + align_clustal_validate_row(self, second_row) + align_clustal_count_pair(self, first_row, second_row) +} + +///| +/// Aggregate statistics across every unordered pair. +pub fn AlignClustalAlignment::counts( + self : AlignClustalAlignment, +) -> AlignClustalCounts { + let mut pairs = 0 + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + for first = 0; first < self.sequences.length(); first = first + 1 { + for second = first + 1 + second < self.sequences.length() + second = second + 1 { + let counts = align_clustal_count_pair(self, first, second) + pairs = pairs + 1 + aligned = aligned + counts.aligned + identities = identities + counts.identities + mismatches = mismatches + counts.mismatches + gap_columns = gap_columns + counts.gap_columns + double_gap_columns = double_gap_columns + counts.double_gap_columns + gap_opens = gap_opens + counts.gap_opens + } + } + AlignClustalCounts::{ + pairs, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +/// Return exact identity over columns containing residues in both rows. +pub fn AlignClustalCounts::identity(self : AlignClustalCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Return per-column non-gap occupancy. +pub fn AlignClustalAlignment::occupancy( + self : AlignClustalAlignment, +) -> Array[Double] { + let result : Array[Double] = [] + for column = 0; column < self.alignment_length(); column = column + 1 { + let mut residues = 0 + for sequence in self.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + result.push(residues.to_double() / self.sequences.length().to_double()) + } + result +} + +///| +/// Calculate a majority residue consensus. Gaps do not vote. +pub fn AlignClustalAlignment::majority_consensus( + self : AlignClustalAlignment, + minimum_fraction? : Double = 0.0, +) -> String raise AlignClustalError { + if minimum_fraction != minimum_fraction || + minimum_fraction < 0.0 || + minimum_fraction > 1.0 { + raise AlignClustalError( + "CLUSTAL consensus minimum fraction must be between 0 and 1", + ) + } + let output = StringBuilder::new(size_hint=self.alignment_length()) + for column = 0; column < self.alignment_length(); column = column + 1 { + let frequencies = Array::make(128, 0) + let mut residues = 0 + for sequence in self.sequences { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if code != '-'.to_int() { + let upper = align_clustal_upper_code(code) + if upper >= 0 && upper < frequencies.length() { + frequencies[upper] = frequencies[upper] + 1 + } + residues = residues + 1 + } + } + if residues == 0 { + output.write_char('-') + continue + } + let mut best_code = 0 + let mut best_count = -1 + for code = 0; code < frequencies.length(); code = code + 1 { + if frequencies[code] > best_count { + best_code = code + best_count = frequencies[code] + } + } + if best_count.to_double() / residues.to_double() < minimum_fraction { + output.write_char('X') + } else { + output.write_char(best_code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +/// Calculate CLUSTAL `*`, `:`, `.`, and space conservation symbols. +/// +/// Nucleotide alignments receive `*` only for exact fully occupied columns. +/// Protein alignments additionally use the standard strong and weak ClustalW +/// residue groups. +pub fn AlignClustalAlignment::calculated_consensus( + self : AlignClustalAlignment, +) -> String { + let output = StringBuilder::new(size_hint=self.alignment_length()) + let nucleotide = align_clustal_is_nucleotide(self) + let strong_groups = [ + "STA", "NEQK", "NHQK", "NDEQ", "QHRK", "MILV", "MILF", "HY", "FYW", + ] + let weak_groups = [ + "CSA", "ATV", "SAG", "STNK", "STPA", "SGND", "SNDEQK", "NDEQHK", "NEQHRK", "FVLIM", + "HFY", + ] + for column = 0; column < self.alignment_length(); column = column + 1 { + if align_clustal_column_has_gap(self, column) { + output.write_char(' ') + } else if align_clustal_column_identical(self, column) { + output.write_char('*') + } else if !nucleotide && + align_clustal_column_in_any_group(self, column, strong_groups) { + output.write_char(':') + } else if !nucleotide && + align_clustal_column_in_any_group(self, column, weak_groups) { + output.write_char('.') + } else { + output.write_char(' ') + } + } + output.to_string() +} + +///| +/// Return a compact alignment summary. +pub fn AlignClustalAlignment::summary(self : AlignClustalAlignment) -> String { + "AlignClustalAlignment(program=" + + self.metadata.program + + ", version=" + + self.metadata.version + + ", sequences=" + + self.num_sequences().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ", consensus=" + + (match self.consensus { + Some(_) => "yes" + None => "no" + }) + + ")" +} + +///| +/// Return a compact interleaved CLUSTAL example with counts and consensus. +pub fn align_clustal_example_text() -> String { + "CLUSTAL W (2.1) multiple sequence alignment\n" + + "\n" + + "\n" + + "reference MKT--AI 5\n" + + "query_one M-TQQAI 6\n" + + "query_two MKT--AV 5\n" + + " * * *:\n" + + "\n" + + "reference LGH 8\n" + + "query_one LG- 8\n" + + "query_two LGH 8\n" + + " ** \n" +} + +///| +fn align_clustal_validate_alignment( + alignment : AlignClustalAlignment, +) -> Unit raise AlignClustalError { + align_clustal_validate_program(alignment.metadata.program) + align_clustal_validate_version(alignment.metadata.version) + align_clustal_validate_plain_text(alignment.metadata.header, "CLUSTAL header") + if alignment.sequences.length() == 0 { + raise AlignClustalError( + "CLUSTAL alignment must contain at least one sequence", + ) + } + let width = alignment.sequences[0].aligned_sequence.length() + if width == 0 { + raise AlignClustalError( + "CLUSTAL alignment must contain at least one column", + ) + } + let ids : Array[String] = [] + for sequence in alignment.sequences { + align_clustal_validate_id(sequence.id) + if align_clustal_find_string(ids, sequence.id) >= 0 { + raise AlignClustalError( + "Duplicated CLUSTAL sequence identifier '" + sequence.id + "'", + ) + } + ids.push(sequence.id) + if sequence.aligned_sequence.length() != width { + raise AlignClustalError("CLUSTAL aligned rows must have equal widths") + } + align_clustal_validate_segment(sequence.aligned_sequence) + if align_clustal_remove_gaps(sequence.aligned_sequence) != sequence.sequence { + raise AlignClustalError( + "CLUSTAL sequence row has inconsistent derived fields", + ) + } + } + for column = 0; column < width; column = column + 1 { + let mut residues = 0 + for sequence in alignment.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + if residues == 0 { + raise AlignClustalError( + "CLUSTAL alignment contains an all-gap column at " + column.to_string(), + ) + } + } + match alignment.consensus { + Some(consensus) => { + if consensus.length() != width { + raise AlignClustalError( + "CLUSTAL consensus length must equal the alignment width", + ) + } + align_clustal_validate_consensus(consensus) + } + None => () + } +} + +///| +fn align_clustal_parse_header( + line : String, +) -> AlignClustalMetadata raise AlignClustalError { + let words = align_clustal_split_whitespace(line) + if words.length() == 0 { + raise AlignClustalError("Empty CLUSTAL header") + } + let program = words[0] + align_clustal_validate_program(program) + let mut version = "" + for index = 1; index < words.length(); index = index + 1 { + let word = align_clustal_strip_parentheses(words[index]) + if word.length() > 0 { + let first = word.unsafe_get(0).to_int() + if first >= '0'.to_int() && first <= '9'.to_int() { + version = word + break + } + } + } + AlignClustalMetadata::create(program, version~, header=line) +} + +///| +fn align_clustal_parse_block( + lines : Array[String], +) -> AlignClustalBlock raise AlignClustalError { + let ids : Array[String] = [] + let segments : Array[String] = [] + let residue_counts : Array[Int?] = [] + let mut segment_start = -1 + let mut width = -1 + let mut consensus : String? = None + let mut saw_consensus = false + for line in lines { + if line.length() == 0 { + continue + } + let first = line.unsafe_get(0).to_int() + if align_clustal_is_whitespace(first) { + if ids.length() == 0 { + raise AlignClustalError( + "CLUSTAL consensus line appears before sequence rows", + ) + } + if saw_consensus { + raise AlignClustalError( + "CLUSTAL block contains multiple consensus lines", + ) + } + saw_consensus = true + let annotation = align_clustal_extract_consensus( + line, segment_start, width, + ) + consensus = Some(annotation) + } else { + if saw_consensus { + raise AlignClustalError( + "CLUSTAL sequence row appears after a consensus line", + ) + } + let fields = align_clustal_split_whitespace(line) + if fields.length() < 2 || fields.length() > 3 { + raise AlignClustalError( + "CLUSTAL sequence line must contain two or three fields", + ) + } + let id = fields[0] + let segment = fields[1] + align_clustal_validate_id(id) + align_clustal_validate_segment(segment) + if segment_start < 0 { + segment_start = align_clustal_segment_start(line) + } + if width < 0 { + width = segment.length() + } else if segment.length() != width { + raise AlignClustalError( + "CLUSTAL rows in one block must have equal widths", + ) + } + ids.push(id) + segments.push(segment) + if fields.length() == 3 { + residue_counts.push( + Some( + align_clustal_parse_nonnegative_int( + fields[2], + "CLUSTAL residue count", + ), + ), + ) + } else { + residue_counts.push(None) + } + } + } + if ids.length() == 0 || width <= 0 { + raise AlignClustalError("CLUSTAL block contains no sequence rows") + } + AlignClustalBlock::{ ids, segments, residue_counts, consensus, width } +} + +///| +fn align_clustal_extract_consensus( + line : String, + start : Int, + width : Int, +) -> String raise AlignClustalError { + if start < 0 || width <= 0 { + raise AlignClustalError("Invalid CLUSTAL consensus position") + } + for index = 0; index < start && index < line.length(); index = index + 1 { + if !align_clustal_is_whitespace(line.unsafe_get(index).to_int()) { + raise AlignClustalError( + "CLUSTAL consensus line is not aligned with sequence data", + ) + } + } + let output = StringBuilder::new(size_hint=width) + for offset = 0; offset < width; offset = offset + 1 { + let index = start + offset + if index >= line.length() { + output.write_char(' ') + } else { + let code = line.unsafe_get(index).to_int() + if code == ' '.to_int() { + output.write_char(' ') + } else if code == '*'.to_int() || + code == ':'.to_int() || + code == '.'.to_int() { + output.write_char(code.unsafe_to_char()) + } else { + raise AlignClustalError("CLUSTAL consensus contains an invalid symbol") + } + } + } + for index = start + width; index < line.length(); index = index + 1 { + if !align_clustal_is_whitespace(line.unsafe_get(index).to_int()) { + raise AlignClustalError( + "CLUSTAL consensus has non-whitespace trailing data", + ) + } + } + output.to_string() +} + +///| +fn align_clustal_count_pair( + alignment : AlignClustalAlignment, + first_row : Int, + second_row : Int, +) -> AlignClustalCounts { + let first = alignment.sequences[first_row].aligned_sequence + let second = alignment.sequences[second_row].aligned_sequence + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + let mut in_gap = false + for column = 0; column < alignment.alignment_length(); column = column + 1 { + let first_code = first.unsafe_get(column).to_int() + let second_code = second.unsafe_get(column).to_int() + let first_gap = first_code == '-'.to_int() + let second_gap = second_code == '-'.to_int() + if first_gap && second_gap { + double_gap_columns = double_gap_columns + 1 + in_gap = false + } else if first_gap || second_gap { + gap_columns = gap_columns + 1 + if !in_gap { + gap_opens = gap_opens + 1 + } + in_gap = true + } else { + aligned = aligned + 1 + if first_code == second_code { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + in_gap = false + } + } + AlignClustalCounts::{ + pairs: 1, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +fn align_clustal_movement_changes( + alignment : AlignClustalAlignment, + first_column : Int, + second_column : Int, +) -> Bool { + for sequence in alignment.sequences { + let first_moves = sequence.aligned_sequence + .unsafe_get(first_column) + .to_int() != + '-'.to_int() + let second_moves = sequence.aligned_sequence + .unsafe_get(second_column) + .to_int() != + '-'.to_int() + if first_moves != second_moves { + return true + } + } + false +} + +///| +fn align_clustal_column_has_gap( + alignment : AlignClustalAlignment, + column : Int, +) -> Bool { + for sequence in alignment.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + return true + } + } + false +} + +///| +fn align_clustal_column_identical( + alignment : AlignClustalAlignment, + column : Int, +) -> Bool { + let first = align_clustal_upper_code( + alignment.sequences[0].aligned_sequence.unsafe_get(column).to_int(), + ) + for index = 1; index < alignment.sequences.length(); index = index + 1 { + let code = align_clustal_upper_code( + alignment.sequences[index].aligned_sequence.unsafe_get(column).to_int(), + ) + if code != first { + return false + } + } + true +} + +///| +fn align_clustal_column_in_any_group( + alignment : AlignClustalAlignment, + column : Int, + groups : Array[String], +) -> Bool { + for group in groups { + let mut all_present = true + for sequence in alignment.sequences { + let code = align_clustal_upper_code( + sequence.aligned_sequence.unsafe_get(column).to_int(), + ) + if !align_clustal_string_has_code(group, code) { + all_present = false + break + } + } + if all_present { + return true + } + } + false +} + +///| +fn align_clustal_is_nucleotide(alignment : AlignClustalAlignment) -> Bool { + let symbols = "ACGTUNRYKMSWBDHV" + for sequence in alignment.sequences { + for index = 0; index < sequence.sequence.length(); index = index + 1 { + let code = align_clustal_upper_code( + sequence.sequence.unsafe_get(index).to_int(), + ) + if code != '?'.to_int() && !align_clustal_string_has_code(symbols, code) { + return false + } + } + } + true +} + +///| +fn align_clustal_validate_program( + program : String, +) -> Unit raise AlignClustalError { + if program != "CLUSTAL" && + program != "PROBCONS" && + program != "MUSCLE" && + program != "MSAPROBS" && + program != "Kalign" && + program != "Biopython" { + raise AlignClustalError( + "Unknown CLUSTAL generator '" + + program + + "'; expected CLUSTAL, PROBCONS, MUSCLE, MSAPROBS, Kalign, or Biopython", + ) + } +} + +///| +fn align_clustal_validate_version( + version : String, +) -> Unit raise AlignClustalError { + for index = 0; index < version.length(); index = index + 1 { + let code = version.unsafe_get(index).to_int() + if align_clustal_is_whitespace(code) || code < 33 || code > 126 { + raise AlignClustalError("CLUSTAL version must be one printable token") + } + } +} + +///| +fn align_clustal_validate_id(id : String) -> Unit raise AlignClustalError { + if id.length() == 0 { + raise AlignClustalError("CLUSTAL sequence identifier must not be empty") + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if align_clustal_is_whitespace(code) || code < 33 || code > 126 { + raise AlignClustalError( + "CLUSTAL sequence identifier contains whitespace or a control character", + ) + } + } +} + +///| +fn align_clustal_validate_segment( + segment : String, +) -> Unit raise AlignClustalError { + if segment.length() == 0 { + raise AlignClustalError("CLUSTAL sequence segment must not be empty") + } + for index = 0; index < segment.length(); index = index + 1 { + let code = segment.unsafe_get(index).to_int() + let letter = (code >= 'A'.to_int() && code <= 'Z'.to_int()) || + (code >= 'a'.to_int() && code <= 'z'.to_int()) + if !letter && + code != '-'.to_int() && + code != '?'.to_int() && + code != '*'.to_int() { + raise AlignClustalError( + "CLUSTAL sequence contains an invalid residue or gap symbol", + ) + } + } +} + +///| +fn align_clustal_validate_consensus( + consensus : String, +) -> Unit raise AlignClustalError { + for index = 0; index < consensus.length(); index = index + 1 { + let code = consensus.unsafe_get(index).to_int() + if code != ' '.to_int() && + code != '*'.to_int() && + code != ':'.to_int() && + code != '.'.to_int() { + raise AlignClustalError("CLUSTAL consensus contains an invalid symbol") + } + } +} + +///| +fn align_clustal_validate_plain_text( + value : String, + label : String, +) -> Unit raise AlignClustalError { + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code == '\n'.to_int() || code == '\r'.to_int() || code < 32 { + raise AlignClustalError( + label + " contains a line break or control character", + ) + } + } +} + +///| +fn align_clustal_validate_row( + alignment : AlignClustalAlignment, + row : Int, +) -> Unit raise AlignClustalError { + if row < 0 || row >= alignment.sequences.length() { + raise AlignClustalError("CLUSTAL row index is out of bounds") + } +} + +///| +fn align_clustal_canonical_header(program : String, version : String) -> String { + if version.length() == 0 { + program + " multiple sequence alignment" + } else { + program + " " + version + " multiple sequence alignment" + } +} + +///| +fn align_clustal_lines(text : String) -> Array[String] { + let raw = text.split("\n").to_array() + let lines : Array[String] = [] + for line in raw { + if line.length() > 0 && + line.unsafe_get(line.length() - 1).to_int() == '\r'.to_int() { + lines.push(line[0:line.length() - 1].to_owned()) + } else { + lines.push(line.to_owned()) + } + } + lines +} + +///| +fn align_clustal_split_whitespace(value : String) -> Array[String] { + let fields : Array[String] = [] + let mut start = 0 + let mut in_field = false + for index = 0; index < value.length(); index = index + 1 { + let whitespace = align_clustal_is_whitespace( + value.unsafe_get(index).to_int(), + ) + if whitespace { + if in_field { + fields.push(value[start:index].to_owned()) + in_field = false + } + } else if !in_field { + start = index + in_field = true + } + } + if in_field { + fields.push(value[start:value.length()].to_owned()) + } + fields +} + +///| +fn align_clustal_segment_start(line : String) -> Int raise AlignClustalError { + let mut index = 0 + while index < line.length() && + !align_clustal_is_whitespace(line.unsafe_get(index).to_int()) { + index = index + 1 + } + while index < line.length() && + align_clustal_is_whitespace(line.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= line.length() { + raise AlignClustalError("CLUSTAL sequence line has no sequence segment") + } + index +} + +///| +fn align_clustal_parse_nonnegative_int( + text : String, + label : String, +) -> Int raise AlignClustalError { + if text.length() == 0 { + raise AlignClustalError(label + " is empty") + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignClustalError(label + " must be a non-negative integer") + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + raise AlignClustalError(label + " is outside the supported integer range") + } + value = value * 10 + digit + } + value +} + +///| +fn align_clustal_residue_count(segment : String) -> Int { + let mut count = 0 + for index = 0; index < segment.length(); index = index + 1 { + if segment.unsafe_get(index).to_int() != '-'.to_int() { + count = count + 1 + } + } + count +} + +///| +fn align_clustal_remove_gaps(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code != '-'.to_int() { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_clustal_strip_parentheses(value : String) -> String { + if value.length() >= 2 && + value.unsafe_get(0).to_int() == '('.to_int() && + value.unsafe_get(value.length() - 1).to_int() == ')'.to_int() { + value[1:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn align_clustal_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 'a'.to_int() + 'A'.to_int() + } else { + code + } +} + +///| +fn align_clustal_string_has_code(value : String, code : Int) -> Bool { + for index = 0; index < value.length(); index = index + 1 { + if value.unsafe_get(index).to_int() == code { + return true + } + } + false +} + +///| +fn align_clustal_find_string(values : Array[String], target : String) -> Int { + align_clustal_find_string_from(values, target, 0) +} + +///| +fn align_clustal_find_string_from( + values : Array[String], + target : String, + start : Int, +) -> Int { + for index = start; index < values.length(); index = index + 1 { + if values[index] == target { + return index + } + } + -1 +} + +///| +fn align_clustal_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} diff --git a/src/align_cluster.mbt b/src/align_cluster.mbt index b92a5060..6340b954 100644 --- a/src/align_cluster.mbt +++ b/src/align_cluster.mbt @@ -106,19 +106,27 @@ pub fn MSADistanceMatrix::new( ///| /// Get the distance matrix as a 2D array. -pub fn MSADistanceMatrix::get_matrix(self : MSADistanceMatrix) -> Array[Array[Double]] { +pub fn MSADistanceMatrix::get_matrix( + self : MSADistanceMatrix, +) -> Array[Array[Double]] { self.matrix } ///| /// Get the sequence IDs. -pub fn MSADistanceMatrix::get_sequence_ids(self : MSADistanceMatrix) -> Array[String] { +pub fn MSADistanceMatrix::get_sequence_ids( + self : MSADistanceMatrix, +) -> Array[String] { self.sequence_ids } ///| /// Get the distance between two sequences by index. -pub fn MSADistanceMatrix::get(self : MSADistanceMatrix, i : Int, j : Int) -> Double { +pub fn MSADistanceMatrix::get( + self : MSADistanceMatrix, + i : Int, + j : Int, +) -> Double { self.matrix[i][j] } @@ -169,7 +177,7 @@ pub struct GuideTree { ///| /// Create a new GuideTree from an array of nodes. pub fn GuideTree::new(nodes : Array[GuideTreeNode]) -> GuideTree { - GuideTree::{ nodes } + GuideTree::{ nodes, } } ///| @@ -308,11 +316,7 @@ pub fn upgma_tree(distance_matrix : MSADistanceMatrix) -> GuideTree { let size_i = count_leaves_in_node(nodes, i_idx) let size_j = count_leaves_in_node(nodes, j_idx) - let new_node = guide_tree_internal( - next_id, - [i_idx, j_idx], - height, - ) + let new_node = guide_tree_internal(next_id, [i_idx, j_idx], height) nodes.push(new_node) let new_node_idx = next_id next_id = next_id + 1 @@ -321,9 +325,10 @@ pub fn upgma_tree(distance_matrix : MSADistanceMatrix) -> GuideTree { for k = 0; k < m; k = k + 1 { if k != min_i && k != min_j { let d = ( - size_i.to_double() * dist[min_i][k] + - size_j.to_double() * dist[min_j][k] - ) / (size_i.to_double() + size_j.to_double()) + size_i.to_double() * dist[min_i][k] + + size_j.to_double() * dist[min_j][k] + ) / + (size_i.to_double() + size_j.to_double()) new_dists_row.push(d) } } @@ -356,15 +361,12 @@ pub fn upgma_tree(distance_matrix : MSADistanceMatrix) -> GuideTree { dist = new_dist } - GuideTree::{ nodes } + GuideTree::{ nodes, } } ///| /// Count the number of leaf nodes under a given node. -fn count_leaves_in_node( - nodes : Array[GuideTreeNode], - node_idx : Int, -) -> Int { +fn count_leaves_in_node(nodes : Array[GuideTreeNode], node_idx : Int) -> Int { let node = nodes[node_idx] if node.is_leaf { 1 @@ -379,9 +381,7 @@ fn count_leaves_in_node( ///| /// Copy a distance matrix. -fn copy_distance_matrix( - matrix : Array[Array[Double]], -) -> Array[Array[Double]] { +fn copy_distance_matrix(matrix : Array[Array[Double]]) -> Array[Array[Double]] { let result : Array[Array[Double]] = Array::new() for row in matrix { let new_row : Array[Double] = Array::new() @@ -400,18 +400,8 @@ fn copy_distance_matrix( /// `seq2` — second sequence string. /// /// Returns a tuple of (aligned_seq1, aligned_seq2, alignment_score). -pub fn align_pairwise( - seq1 : String, - seq2 : String, -) -> (String, String, Double) { - align_pairwise_with_scoring( - seq1, - seq2, - 2.0, - -1.0, - -2.0, - -0.5, - ) +pub fn align_pairwise(seq1 : String, seq2 : String) -> (String, String, Double) { + align_pairwise_with_scoring(seq1, seq2, 2.0, -1.0, -2.0, -0.5) } ///| @@ -469,11 +459,7 @@ pub fn align_pairwise_with_scoring( for j = 1; j <= m; j = j + 1 { let c1 = seq1.unsafe_get(i - 1) let c2 = seq2.unsafe_get(j - 1) - let sub_score = if c1 == c2 { - match_score - } else { - mismatch_score - } + let sub_score = if c1 == c2 { match_score } else { mismatch_score } let diag = dp[i - 1][j - 1] + sub_score let up = dp[i - 1][j] + @@ -784,12 +770,11 @@ pub fn progressive_alignment( let mut all_processed = true for child_idx in node.children { match processed.get(child_idx) { - Some(true) => { + Some(true) => match alignment_map.get(child_idx) { Some(profile) => child_results.push(profile) None => all_processed = false } - } _ => all_processed = false } } @@ -797,8 +782,7 @@ pub fn progressive_alignment( let left_profile = child_results[0] let right_profile = child_results[1] let (aligned_left, aligned_right, _) = align_profiles( - left_profile, - right_profile, + left_profile, right_profile, ) let left_aligned_profile : Array[String] = Array::new() let right_aligned_profile : Array[String] = Array::new() @@ -867,7 +851,7 @@ pub fn progressive_alignment( } else { "seq_" + seq_id_idx.to_string() } - records.push(SeqRecord::new(Seq::new(s), id=id)) + records.push(SeqRecord::new(Seq::new(s), id~)) seq_id_idx = seq_id_idx + 1 } create_msa_safe(records) @@ -881,9 +865,7 @@ fn create_msa_safe(records : Array[SeqRecord]) -> MultipleSeqAlignment { annotations: Map([], capacity=0), column_annotations: Map([], capacity=0), } - try { - MultipleSeqAlignment::new(records) - } catch { + MultipleSeqAlignment::new(records) catch { _ => empty } } @@ -988,4 +970,4 @@ fn repeat_char(c : Char, n : Int) -> String { buf.write_char(c) } buf.to_string() -} \ No newline at end of file +} diff --git a/src/align_emboss.mbt b/src/align_emboss.mbt new file mode 100644 index 00000000..f837e7f4 --- /dev/null +++ b/src/align_emboss.mbt @@ -0,0 +1,1792 @@ +// Alignment-aware EMBOSS srspair/pair/simple support. +// +// This module follows Biopython 1.86 Bio.Align.emboss semantics. EMBOSS +// sequence lines use a fixed 21-character prefix and one-based inclusive +// printed coordinates. Public coordinates are normalized to zero-based +// residue positions and boundary coordinates. + +///| +/// Error raised for malformed EMBOSS alignment output or invalid operations. +pub suberror AlignEmbossError { + AlignEmbossError(String) +} + +///| +/// Orientation of a sequence row in an EMBOSS alignment. +pub(all) enum AlignEmbossStrand { + AlignEmbossForward + AlignEmbossReverse +} derive(Eq, Debug) + +///| +/// File-level metadata from the EMBOSS report header. +pub struct AlignEmbossMetadata { + program : String + rundate : String + report_file : String + align_format : String + command_line : String +} derive(Eq, Debug) + +///| +/// Per-alignment annotations reported by EMBOSS. +pub struct AlignEmbossAnnotations { + matrix : String + gap_penalty : Double? + extend_penalty : Double? + length : Int + identity : Int? + similarity : Int? + gaps : Int? + score : Double? + longest_identity : String + longest_similarity : String + shortest_identity : String + shortest_similarity : String +} derive(Eq, Debug) + +///| +/// One sequence row in an EMBOSS alignment. +/// +/// `start` and `end` are zero-based boundary coordinates. Forward rows have +/// `start <= end`; reverse rows have `start > end`. `aligned_sequence` is in +/// displayed alignment orientation and `sequence` is the same row without +/// gap characters. +pub struct AlignEmbossSequence { + id : String + sequence : String + aligned_sequence : String + start : Int + end : Int + strand : AlignEmbossStrand +} derive(Eq, Debug) + +///| +/// One parsed EMBOSS pairwise or multiple sequence alignment. +pub struct AlignEmbossAlignment { + sequences : Array[AlignEmbossSequence] + consensus : String + annotations : AlignEmbossAnnotations +} derive(Eq, Debug) + +///| +/// A complete EMBOSS report containing one or more alignments. +pub struct AlignEmbossDocument { + metadata : AlignEmbossMetadata + alignments : Array[AlignEmbossAlignment] +} derive(Eq, Debug) + +///| +/// Pairwise statistics for two rows in an EMBOSS alignment. +pub struct AlignEmbossPairCounts { + columns : Int + aligned : Int + identities : Int + mismatches : Int + insertions : Int + deletions : Int + double_gap_columns : Int + insertion_opens : Int + deletion_opens : Int + positives : Int +} derive(Eq, Debug) + +///| +/// Construct file-level EMBOSS metadata. +pub fn AlignEmbossMetadata::create( + program? : String = "", + rundate? : String = "", + report_file? : String = "", + align_format? : String = "srspair", + command_line? : String = "", +) -> AlignEmbossMetadata raise AlignEmbossError { + align_emboss_validate_header_value(program, "Program") + align_emboss_validate_header_value(rundate, "Rundate") + align_emboss_validate_header_value(report_file, "Report_file") + align_emboss_validate_header_value(command_line, "Command line") + if align_format != "srspair" && + align_format != "pair" && + align_format != "simple" { + raise AlignEmbossError( + "EMBOSS Align_format must be srspair, pair, or simple", + ) + } + AlignEmbossMetadata::{ + program, + rundate, + report_file, + align_format, + command_line, + } +} + +///| +/// Construct per-alignment EMBOSS annotations. +pub fn AlignEmbossAnnotations::create( + length : Int, + matrix? : String = "", + gap_penalty? : Double? = None, + extend_penalty? : Double? = None, + identity? : Int? = None, + similarity? : Int? = None, + gaps? : Int? = None, + score? : Double? = None, + longest_identity? : String = "", + longest_similarity? : String = "", + shortest_identity? : String = "", + shortest_similarity? : String = "", +) -> AlignEmbossAnnotations raise AlignEmbossError { + if length <= 0 { + raise AlignEmbossError("EMBOSS alignment length must be positive") + } + align_emboss_validate_header_value(matrix, "Matrix") + align_emboss_validate_optional_nonnegative_double(gap_penalty, "Gap_penalty") + align_emboss_validate_optional_nonnegative_double( + extend_penalty, "Extend_penalty", + ) + align_emboss_validate_optional_count(identity, length, "Identity") + align_emboss_validate_optional_count(similarity, length, "Similarity") + align_emboss_validate_optional_count(gaps, length, "Gaps") + match score { + Some(value) => + if value.abs() > 1.0e300 { + raise AlignEmbossError("EMBOSS Score must be finite") + } + None => () + } + AlignEmbossAnnotations::{ + matrix, + gap_penalty, + extend_penalty, + length, + identity, + similarity, + gaps, + score, + longest_identity, + longest_similarity, + shortest_identity, + shortest_similarity, + } +} + +///| +/// Construct one normalized EMBOSS sequence row. +pub fn AlignEmbossSequence::create( + id : String, + aligned_sequence : String, + start : Int, + end : Int, +) -> AlignEmbossSequence raise AlignEmbossError { + align_emboss_validate_id(id) + if aligned_sequence.length() == 0 { + raise AlignEmbossError("EMBOSS aligned sequence must not be empty") + } + align_emboss_validate_aligned_fragment(aligned_sequence, "sequence row") + if start < 0 || end < 0 { + raise AlignEmbossError("EMBOSS sequence coordinates must be non-negative") + } + let sequence = align_emboss_remove_gaps(aligned_sequence) + let span = if start >= end { start - end } else { end - start } + if sequence.length() != span { + raise AlignEmbossError( + "EMBOSS ungapped sequence length does not match its coordinate span", + ) + } + let strand = if start <= end { + AlignEmbossForward + } else { + AlignEmbossReverse + } + AlignEmbossSequence::{ id, sequence, aligned_sequence, start, end, strand } +} + +///| +/// Construct one EMBOSS alignment. +pub fn AlignEmbossAlignment::create( + sequences : Array[AlignEmbossSequence], + annotations : AlignEmbossAnnotations, + consensus? : String = "", +) -> AlignEmbossAlignment raise AlignEmbossError { + let copied : Array[AlignEmbossSequence] = [] + for sequence in sequences { + copied.push(sequence) + } + let alignment = AlignEmbossAlignment::{ + sequences: copied, + consensus, + annotations, + } + align_emboss_validate_alignment(alignment) + alignment +} + +///| +/// Construct a complete EMBOSS report. +pub fn AlignEmbossDocument::create( + metadata : AlignEmbossMetadata, + alignments : Array[AlignEmbossAlignment], +) -> AlignEmbossDocument raise AlignEmbossError { + if alignments.length() == 0 { + raise AlignEmbossError("EMBOSS document must contain an alignment") + } + let copied : Array[AlignEmbossAlignment] = [] + for alignment in alignments { + align_emboss_validate_alignment(alignment) + copied.push(alignment) + } + AlignEmbossDocument::{ metadata, alignments: copied } +} + +///| +/// Parse an EMBOSS srspair, pair, or simple alignment report. +/// +/// The parser accepts LF or CRLF input and validates file/alignment headers, +/// declared row order, fixed sequence-line fields, block widths, coordinate +/// continuity, final alignment width, consensus width, and reported counts. +pub fn align_emboss_parse( + text : String, +) -> AlignEmbossDocument raise AlignEmbossError { + let lines = align_emboss_lines(text) + if lines.length() == 0 || (lines.length() == 1 && lines[0].length() == 0) { + raise AlignEmbossError("Empty EMBOSS input") + } + let header_divider = "########################################" + if lines[0].trim().to_owned() != header_divider { + raise AlignEmbossError("EMBOSS file is missing its header divider") + } + let mut program = "" + let mut rundate = "" + let mut report_file = "" + let mut align_format = "srspair" + let mut command_line = "" + let mut index = 1 + let mut header_closed = false + while index < lines.length() { + let line = lines[index] + if line.trim().to_owned() == header_divider { + header_closed = true + index = index + 1 + break + } + if align_emboss_has_prefix(line, "# ") { + if command_line.length() == 0 { + raise AlignEmbossError( + "EMBOSS command continuation appears before Commandline", + ) + } + command_line = command_line + " " + line[1:].trim().to_owned() + index = index + 1 + continue + } + if !align_emboss_has_prefix(line, "# ") { + raise AlignEmbossError( + "Unexpected EMBOSS header line " + (index + 1).to_string(), + ) + } + let content = line[2:].to_owned() + let (key, value) = align_emboss_split_key_value( + content, + "header", + index + 1, + ) + match key { + "Program" => program = value + "Rundate" => rundate = value + "Report_file" => report_file = value + "Align_format" => align_format = value + "Commandline" => command_line = value + _ => () + } + index = index + 1 + } + if !header_closed { + raise AlignEmbossError("Truncated EMBOSS file header") + } + let metadata = AlignEmbossMetadata::create( + program~, + rundate~, + report_file~, + align_format~, + command_line~, + ) + let alignments : Array[AlignEmbossAlignment] = [] + while index < lines.length() { + let line = lines[index].trim().to_owned() + if line.length() == 0 || align_emboss_is_separator(line) { + index = index + 1 + continue + } + if line != "#=======================================" { + raise AlignEmbossError( + "Unexpected content before EMBOSS alignment at line " + + (index + 1).to_string(), + ) + } + let (alignment, next_index) = align_emboss_parse_alignment( + lines, index, metadata, + ) + alignments.push(alignment) + index = next_index + } + AlignEmbossDocument::create(metadata, alignments) +} + +///| +/// Serialize an EMBOSS document in a canonical parseable form. +/// +/// The upstream Biopython module is read-only; this writer is a MoonBit +/// extension for reproducible round trips. Sequence blocks default to the +/// EMBOSS width of 50 columns. +pub fn align_emboss_write( + document : AlignEmbossDocument, + line_width? : Int = 50, +) -> String raise AlignEmbossError { + if line_width <= 0 || line_width > 50 { + raise AlignEmbossError("EMBOSS line width must be between 1 and 50") + } + if document.alignments.length() == 0 { + raise AlignEmbossError("EMBOSS document must contain an alignment") + } + let output = StringBuilder::new() + output.write_string("########################################\n") + if document.metadata.program.length() > 0 { + output.write_string("# Program: " + document.metadata.program + "\n") + } + if document.metadata.rundate.length() > 0 { + output.write_string("# Rundate: " + document.metadata.rundate + "\n") + } + if document.metadata.command_line.length() > 0 { + output.write_string( + "# Commandline: " + document.metadata.command_line + "\n", + ) + } + output.write_string( + "# Align_format: " + document.metadata.align_format + "\n", + ) + if document.metadata.report_file.length() > 0 { + output.write_string( + "# Report_file: " + document.metadata.report_file + "\n", + ) + } + output.write_string("########################################\n") + for alignment in document.alignments { + align_emboss_validate_alignment(alignment) + output.write_string("#=======================================\n#\n") + output.write_string( + "# Aligned_sequences: " + alignment.sequences.length().to_string() + "\n", + ) + for row = 0; row < alignment.sequences.length(); row = row + 1 { + output.write_string( + "# " + (row + 1).to_string() + ": " + alignment.sequences[row].id + "\n", + ) + } + if alignment.annotations.matrix.length() > 0 { + output.write_string("# Matrix: " + alignment.annotations.matrix + "\n") + } + align_emboss_write_optional_double( + output, + "Gap_penalty", + alignment.annotations.gap_penalty, + ) + align_emboss_write_optional_double( + output, + "Extend_penalty", + alignment.annotations.extend_penalty, + ) + output.write_string("#\n") + output.write_string( + "# Length: " + alignment.annotations.length.to_string() + "\n", + ) + align_emboss_write_optional_count( + output, + "Identity", + alignment.annotations.identity, + alignment.annotations.length, + ) + align_emboss_write_optional_count( + output, + "Similarity", + alignment.annotations.similarity, + alignment.annotations.length, + ) + align_emboss_write_optional_count( + output, + "Gaps", + alignment.annotations.gaps, + alignment.annotations.length, + ) + match alignment.annotations.score { + Some(value) => output.write_string("# Score: " + value.to_string() + "\n") + None => () + } + align_emboss_write_long_annotation( + output, + "Longest_Identity", + alignment.annotations.longest_identity, + ) + align_emboss_write_long_annotation( + output, + "Longest_Similarity", + alignment.annotations.longest_similarity, + ) + align_emboss_write_long_annotation( + output, + "Shortest_Identity", + alignment.annotations.shortest_identity, + ) + align_emboss_write_long_annotation( + output, + "Shortest_Similarity", + alignment.annotations.shortest_similarity, + ) + output.write_string("#\n#=======================================\n") + let width = alignment.alignment_length() + let mut column = 0 + while column < width { + let block_end = if column + line_width < width { + column + line_width + } else { + width + } + for row = 0; row < alignment.sequences.length(); row = row + 1 { + let sequence = alignment.sequences[row] + let fragment = sequence.aligned_sequence[column:block_end].to_owned() + let boundary = align_emboss_boundary_at(sequence, column) + let residues = align_emboss_count_residues(fragment) + let (printed_start, printed_end) = align_emboss_printed_coordinates( + sequence, boundary, residues, + ) + output.write_string( + align_emboss_sequence_prefix(sequence.id, printed_start) + + fragment + + " " + + align_emboss_pad_left(printed_end.to_string(), 7) + + "\n", + ) + } + if alignment.consensus.length() > 0 { + output.write_string( + " ".repeat(21) + + alignment.consensus[column:block_end].to_owned() + + "\n", + ) + } + output.write_char('\n') + column = block_end + } + output.write_string( + "#---------------------------------------\n" + + "#---------------------------------------\n", + ) + } + output.to_string() +} + +///| +/// Return the number of alignments in a report. +pub fn AlignEmbossDocument::num_alignments(self : AlignEmbossDocument) -> Int { + self.alignments.length() +} + +///| +/// Return an alignment by zero-based index. +pub fn AlignEmbossDocument::get( + self : AlignEmbossDocument, + index : Int, +) -> AlignEmbossAlignment? { + if index < 0 || index >= self.alignments.length() { + None + } else { + Some(self.alignments[index]) + } +} + +///| +/// Return the number of sequence rows. +pub fn AlignEmbossAlignment::num_sequences(self : AlignEmbossAlignment) -> Int { + self.sequences.length() +} + +///| +/// Return the alignment width including gap columns. +pub fn AlignEmbossAlignment::alignment_length( + self : AlignEmbossAlignment, +) -> Int { + self.annotations.length +} + +///| +/// Find the first row with an exact declared identifier. +pub fn AlignEmbossAlignment::find_sequence( + self : AlignEmbossAlignment, + id : String, +) -> Int? { + for row = 0; row < self.sequences.length(); row = row + 1 { + if self.sequences[row].id == id { + return Some(row) + } + } + None +} + +///| +/// Return all row characters at an alignment column. +pub fn AlignEmbossAlignment::column( + self : AlignEmbossAlignment, + column : Int, +) -> String? { + if column < 0 || column >= self.alignment_length() { + return None + } + let result = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + result.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(result.to_string()) +} + +///| +/// Return the zero-based absolute residue position at an alignment column. +/// +/// Gap columns and invalid indices return `None`. +pub fn AlignEmbossAlignment::column_to_sequence_position( + self : AlignEmbossAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.alignment_length() { + return None + } + let sequence = self.sequences[row] + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let boundary = align_emboss_boundary_at(sequence, column) + if sequence.strand == AlignEmbossForward { + Some(boundary) + } else { + Some(boundary - 1) + } +} + +///| +/// Map a zero-based absolute residue position to its alignment column. +pub fn AlignEmbossAlignment::sequence_position_to_column( + self : AlignEmbossAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || row >= self.sequences.length() || position < 0 { + return None + } + for column = 0; column < self.alignment_length(); column = column + 1 { + match self.column_to_sequence_position(row, column) { + Some(value) => if value == position { return Some(column) } + None => () + } + } + None +} + +///| +/// Map a residue position from one row through the alignment to another row. +pub fn AlignEmbossAlignment::map_position( + self : AlignEmbossAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column absolute coordinates for two rows. +/// +/// Gaps are represented by `None`. +pub fn AlignEmbossAlignment::aligned_pairs( + self : AlignEmbossAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise AlignEmbossError { + align_emboss_validate_row(self, first_row) + align_emboss_validate_row(self, second_row) + let pairs : Array[(Int?, Int?)] = [] + for column = 0; column < self.alignment_length(); column = column + 1 { + pairs.push( + ( + self.column_to_sequence_position(first_row, column), + self.column_to_sequence_position(second_row, column), + ), + ) + } + pairs +} + +///| +/// Build a compact Biopython-style boundary coordinate path. +/// +/// The outer array contains one coordinate array per row. Breakpoints are +/// emitted whenever the per-column movement vector changes. +pub fn AlignEmbossAlignment::coordinate_path( + self : AlignEmbossAlignment, +) -> Array[Array[Int]] { + let result : Array[Array[Int]] = [] + for sequence in self.sequences { + result.push([sequence.start]) + } + if self.alignment_length() == 0 { + return result + } + for column = 1; column < self.alignment_length(); column = column + 1 { + let mut changed = false + for row = 0; row < self.sequences.length(); row = row + 1 { + let sequence = self.sequences[row] + let previous = align_emboss_column_step(sequence, column - 1) + let current = align_emboss_column_step(sequence, column) + if previous != current { + changed = true + break + } + } + if changed { + for row = 0; row < self.sequences.length(); row = row + 1 { + result[row].push(align_emboss_boundary_at(self.sequences[row], column)) + } + } + } + for row = 0; row < self.sequences.length(); row = row + 1 { + result[row].push(self.sequences[row].end) + } + result +} + +///| +/// Compute pairwise residue, gap, gap-open, and positive-column counts. +/// +/// A gap in the first row is an insertion; a gap in the second row is a +/// deletion. `positives` uses `|` and `:` from the EMBOSS consensus when +/// available, otherwise it equals exact identities. +pub fn AlignEmbossAlignment::pair_counts( + self : AlignEmbossAlignment, + first_row : Int, + second_row : Int, +) -> AlignEmbossPairCounts raise AlignEmbossError { + align_emboss_validate_row(self, first_row) + align_emboss_validate_row(self, second_row) + let first = self.sequences[first_row].aligned_sequence + let second = self.sequences[second_row].aligned_sequence + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut insertions = 0 + let mut deletions = 0 + let mut double_gap_columns = 0 + let mut insertion_opens = 0 + let mut deletion_opens = 0 + let mut positives = 0 + let mut in_insertion = false + let mut in_deletion = false + for column = 0; column < self.alignment_length(); column = column + 1 { + let first_code = first.unsafe_get(column).to_int() + let second_code = second.unsafe_get(column).to_int() + let first_gap = first_code == '-'.to_int() + let second_gap = second_code == '-'.to_int() + if first_gap && second_gap { + double_gap_columns = double_gap_columns + 1 + in_insertion = false + in_deletion = false + } else if first_gap { + insertions = insertions + 1 + if !in_insertion { + insertion_opens = insertion_opens + 1 + } + in_insertion = true + in_deletion = false + } else if second_gap { + deletions = deletions + 1 + if !in_deletion { + deletion_opens = deletion_opens + 1 + } + in_deletion = true + in_insertion = false + } else { + aligned = aligned + 1 + if first_code == second_code { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + if self.consensus.length() == self.alignment_length() { + let marker = self.consensus.unsafe_get(column).to_int() + if marker == '|'.to_int() || marker == ':'.to_int() { + positives = positives + 1 + } + } else if first_code == second_code { + positives = positives + 1 + } + in_insertion = false + in_deletion = false + } + } + AlignEmbossPairCounts::{ + columns: self.alignment_length(), + aligned, + identities, + mismatches, + insertions, + deletions, + double_gap_columns, + insertion_opens, + deletion_opens, + positives, + } +} + +///| +/// Return exact identity divided by columns where both rows have residues. +pub fn AlignEmbossPairCounts::identity(self : AlignEmbossPairCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Return a compact alignment description. +pub fn AlignEmbossAlignment::summary(self : AlignEmbossAlignment) -> String { + "EMBOSS alignment(rows=" + + self.sequences.length().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ", matrix=" + + self.annotations.matrix + + ")" +} + +///| +/// Return a compact report description. +pub fn AlignEmbossDocument::summary(self : AlignEmbossDocument) -> String { + "EMBOSS report(program=" + + self.metadata.program + + ", format=" + + self.metadata.align_format + + ", alignments=" + + self.alignments.length().to_string() + + ")" +} + +///| +/// Small EMBOSS srspair document used by tests and examples. +pub fn align_emboss_example_text() -> String { + "########################################\n" + + "# Program: water\n" + + "# Rundate: Wed Jan 16 17:23:19 2002\n" + + "# Commandline: water\n" + + "# -asequence reference.fa\n" + + "# -bsequence query.fa\n" + + "# Align_format: srspair\n" + + "# Report_file: stdout\n" + + "########################################\n" + + "#=======================================\n" + + "#\n" + + "# Aligned_sequences: 2\n" + + "# 1: reference_sequence\n" + + "# 2: query_sequence\n" + + "# Matrix: EDNAFULL\n" + + "# Gap_penalty: 10.0\n" + + "# Extend_penalty: 0.5\n" + + "#\n" + + "# Length: 18\n" + + "# Identity: 16/18 (88.9%)\n" + + "# Similarity: 17/18 (94.4%)\n" + + "# Gaps: 1/18 ( 5.6%)\n" + + "# Score: 72.5\n" + + "#\n" + + "#=======================================\n" + + "reference_seq 79 ACGTTGAGT-CTGGGATG 95\n" + + " ||||||||| ||||:|||\n" + + "query_sequence 1 ACGTTGAGTACTGGAATG 18\n" + + "#---------------------------------------\n" + + "#---------------------------------------\n" +} + +///| +fn align_emboss_parse_alignment( + lines : Array[String], + start_index : Int, + metadata : AlignEmbossMetadata, +) -> (AlignEmbossAlignment, Int) raise AlignEmbossError { + let mut index = start_index + 1 + let identifiers : Array[String] = [] + let mut number_of_sequences = -1 + let mut length = -1 + let mut matrix = "" + let mut gap_penalty : Double? = None + let mut extend_penalty : Double? = None + let mut identity : Int? = None + let mut similarity : Int? = None + let mut gaps : Int? = None + let mut score : Double? = None + let mut longest_identity = "" + let mut longest_similarity = "" + let mut shortest_identity = "" + let mut shortest_similarity = "" + let mut metadata_closed = false + while index < lines.length() { + let line = lines[index] + let trimmed = line.trim().to_owned() + if trimmed == "#=======================================" { + metadata_closed = true + index = index + 1 + break + } + if trimmed == "#" || trimmed.length() == 0 { + index = index + 1 + continue + } + if !align_emboss_has_prefix(line, "# ") { + raise AlignEmbossError( + "Unexpected EMBOSS alignment header line " + (index + 1).to_string(), + ) + } + let content = line[2:].to_owned() + let (key, value) = align_emboss_split_annotation(content, index + 1) + if key == "Aligned_sequences" { + if number_of_sequences >= 0 { + raise AlignEmbossError("Duplicate EMBOSS Aligned_sequences") + } + number_of_sequences = align_emboss_parse_positive_int( + value, + "Aligned_sequences", + index + 1, + ) + index = index + 1 + for expected = 1; expected <= number_of_sequences; expected = expected + 1 { + if index >= lines.length() || + !align_emboss_has_prefix(lines[index], "# ") { + raise AlignEmbossError("Truncated EMBOSS sequence identifier list") + } + let (number_text, identifier) = align_emboss_split_key_value( + lines[index][2:].to_owned(), + "sequence identifier", + index + 1, + ) + let number = align_emboss_parse_positive_int( + number_text, + "sequence number", + index + 1, + ) + if number != expected { + raise AlignEmbossError( + "Non-contiguous EMBOSS sequence identifier numbering", + ) + } + align_emboss_validate_id(identifier) + identifiers.push(identifier) + index = index + 1 + } + continue + } + match key { + "Matrix" => matrix = value + "Gap_penalty" => + gap_penalty = Some( + align_emboss_parse_number(value, "Gap_penalty", index + 1), + ) + "Extend_penalty" => + extend_penalty = Some( + align_emboss_parse_number(value, "Extend_penalty", index + 1), + ) + "Length" => + length = align_emboss_parse_positive_int(value, "Length", index + 1) + "Identity" => + identity = Some( + align_emboss_parse_fraction_count(value, "Identity", index + 1), + ) + "Similarity" => + similarity = Some( + align_emboss_parse_fraction_count(value, "Similarity", index + 1), + ) + "Gaps" => + gaps = Some(align_emboss_parse_fraction_count(value, "Gaps", index + 1)) + "Score" => + score = Some(align_emboss_parse_number(value, "Score", index + 1)) + "Longest_Identity" => longest_identity = value + "Longest_Similarity" => longest_similarity = value + "Shortest_Identity" => shortest_identity = value + "Shortest_Similarity" => shortest_similarity = value + _ => + raise AlignEmbossError( + "Unknown EMBOSS alignment annotation '" + key + "'", + ) + } + index = index + 1 + } + if !metadata_closed { + raise AlignEmbossError("Truncated EMBOSS alignment header") + } + if number_of_sequences <= 0 || identifiers.length() != number_of_sequences { + raise AlignEmbossError("Number of EMBOSS sequences is missing") + } + if length <= 0 { + raise AlignEmbossError("Length of EMBOSS alignment is missing") + } + let annotations = AlignEmbossAnnotations::create( + length, + matrix~, + gap_penalty~, + extend_penalty~, + identity~, + similarity~, + gaps~, + score~, + longest_identity~, + longest_similarity~, + shortest_identity~, + shortest_similarity~, + ) + let aligned_rows : Array[String] = [] + let ungapped_rows : Array[String] = [] + let starts : Array[Int] = [] + let ends : Array[Int] = [] + let directions : Array[Int] = [] + for _row = 0; _row < number_of_sequences; _row = _row + 1 { + aligned_rows.push("") + ungapped_rows.push("") + starts.push(0) + ends.push(0) + directions.push(0) + } + let consensus_builder = StringBuilder::new(size_hint=length) + let mut consensus_seen = false + let mut row_index = 0 + let mut block_width = 0 + while index < lines.length() { + let line = lines[index] + let trimmed = line.trim().to_owned() + if align_emboss_is_separator(trimmed) { + index = index + 1 + break + } + if trimmed.length() == 0 && line.length() < 21 { + if row_index == number_of_sequences { + row_index = 0 + block_width = 0 + } + index = index + 1 + continue + } + if line.length() < 21 { + raise AlignEmbossError( + "Malformed EMBOSS alignment body at line " + (index + 1).to_string(), + ) + } + let prefix = line[0:21].to_owned().trim().to_owned() + if prefix.length() == 0 { + if block_width <= 0 { + raise AlignEmbossError( + "EMBOSS consensus appears before a sequence block", + ) + } + let raw = line[21:].to_owned() + let chunk = align_emboss_consensus_chunk(raw, block_width, index + 1) + consensus_builder.write_string(chunk) + consensus_seen = true + index = index + 1 + continue + } + if row_index == number_of_sequences { + row_index = 0 + block_width = 0 + } + let prefix_words = align_emboss_split_whitespace(prefix) + if prefix_words.length() != 2 { + raise AlignEmbossError( + "Malformed EMBOSS sequence prefix at line " + (index + 1).to_string(), + ) + } + let printed_id = prefix_words[0] + if !align_emboss_has_prefix(identifiers[row_index], printed_id) { + raise AlignEmbossError( + "Unexpected EMBOSS sequence identifier at line " + + (index + 1).to_string(), + ) + } + let printed_start = align_emboss_parse_nonnegative_int( + prefix_words[1], + "sequence start", + index + 1, + ) + let suffix_words = align_emboss_split_whitespace(line[21:].to_owned()) + if suffix_words.length() != 2 { + raise AlignEmbossError( + "Malformed EMBOSS sequence fields at line " + (index + 1).to_string(), + ) + } + let fragment = suffix_words[0] + align_emboss_validate_aligned_fragment( + fragment, + "line " + (index + 1).to_string(), + ) + let printed_end = align_emboss_parse_nonnegative_int( + suffix_words[1], + "sequence end", + index + 1, + ) + if block_width == 0 { + block_width = fragment.length() + } else if fragment.length() != block_width { + raise AlignEmbossError( + "Inconsistent EMBOSS sequence block width at line " + + (index + 1).to_string(), + ) + } + let ungapped = align_emboss_remove_gaps(fragment) + let consumed = ungapped.length() + let previous_total = ungapped_rows[row_index].length() + if previous_total == 0 && consumed > 0 { + if printed_start < printed_end { + let start_boundary = printed_start - 1 + if printed_start <= 0 || printed_end != start_boundary + consumed { + raise AlignEmbossError( + "Invalid forward EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + starts[row_index] = start_boundary + ends[row_index] = printed_end + directions[row_index] = 1 + } else if printed_start > printed_end { + let end_boundary = printed_end - 1 + if printed_end <= 0 || end_boundary != printed_start - consumed { + raise AlignEmbossError( + "Invalid reverse EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + starts[row_index] = printed_start + ends[row_index] = end_boundary + directions[row_index] = -1 + } else { + if consumed != 1 || printed_start <= 0 { + raise AlignEmbossError( + "Ambiguous EMBOSS coordinates at line " + (index + 1).to_string(), + ) + } + // A one-residue block prints the same coordinate in either direction. + // Keep the one-based residue coordinate until another block resolves it. + starts[row_index] = printed_start + ends[row_index] = printed_end + directions[row_index] = 0 + } + } else if consumed == 0 { + if directions[row_index] >= 0 { + if printed_start != ends[row_index] || printed_end != printed_start { + raise AlignEmbossError( + "Invalid forward gap-only EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + } else if printed_start - 1 != ends[row_index] || + printed_end != printed_start { + raise AlignEmbossError( + "Invalid reverse gap-only EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + } else if directions[row_index] > 0 { + let start_boundary = printed_start - 1 + if printed_start <= 0 || + start_boundary != ends[row_index] || + printed_end != start_boundary + consumed { + raise AlignEmbossError( + "Discontinuous forward EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + ends[row_index] = printed_end + } else if directions[row_index] < 0 { + let end_boundary = printed_end - 1 + if printed_end <= 0 || + printed_start != ends[row_index] || + end_boundary != printed_start - consumed { + raise AlignEmbossError( + "Discontinuous reverse EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + ends[row_index] = end_boundary + } else if printed_start == ends[row_index] + 1 { + let start_boundary = printed_start - 1 + if printed_end != start_boundary + consumed { + raise AlignEmbossError( + "Discontinuous forward EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + starts[row_index] = starts[row_index] - 1 + ends[row_index] = printed_end + directions[row_index] = 1 + } else if printed_start + 1 == ends[row_index] { + let end_boundary = printed_end - 1 + if printed_end <= 0 || end_boundary != printed_start - consumed { + raise AlignEmbossError( + "Discontinuous reverse EMBOSS coordinates at line " + + (index + 1).to_string(), + ) + } + ends[row_index] = end_boundary + directions[row_index] = -1 + } else { + raise AlignEmbossError( + "Cannot determine EMBOSS row orientation at line " + + (index + 1).to_string(), + ) + } + aligned_rows[row_index] = aligned_rows[row_index] + fragment + ungapped_rows[row_index] = ungapped_rows[row_index] + ungapped + row_index = row_index + 1 + index = index + 1 + } + let sequences : Array[AlignEmbossSequence] = [] + for row = 0; row < number_of_sequences; row = row + 1 { + if directions[row] == 0 && ungapped_rows[row].length() == 1 { + starts[row] = starts[row] - 1 + directions[row] = 1 + } + if aligned_rows[row].length() != length { + raise AlignEmbossError( + "EMBOSS row '" + + identifiers[row] + + "' has width " + + aligned_rows[row].length().to_string() + + "; expected " + + length.to_string(), + ) + } + sequences.push( + AlignEmbossSequence::create( + identifiers[row], + aligned_rows[row], + starts[row], + ends[row], + ), + ) + } + let consensus = if consensus_seen { + let value = consensus_builder.to_string() + if value.length() != length { + raise AlignEmbossError( + "EMBOSS consensus width does not match declared Length", + ) + } + value + } else { + "" + } + let alignment = AlignEmbossAlignment::create( + sequences, + annotations, + consensus~, + ) + align_emboss_validate_reported_counts(alignment) + ignore(metadata) + (alignment, index) +} + +///| +fn align_emboss_validate_alignment( + alignment : AlignEmbossAlignment, +) -> Unit raise AlignEmbossError { + if alignment.sequences.length() == 0 { + raise AlignEmbossError("EMBOSS alignment must contain a sequence") + } + if alignment.annotations.length <= 0 { + raise AlignEmbossError("EMBOSS alignment length must be positive") + } + for sequence in alignment.sequences { + align_emboss_validate_id(sequence.id) + if sequence.aligned_sequence.length() != alignment.annotations.length { + raise AlignEmbossError( + "EMBOSS sequence width does not match alignment length", + ) + } + align_emboss_validate_aligned_fragment( + sequence.aligned_sequence, + "sequence row", + ) + if align_emboss_remove_gaps(sequence.aligned_sequence) != sequence.sequence { + raise AlignEmbossError("EMBOSS ungapped sequence is inconsistent") + } + let span = if sequence.start >= sequence.end { + sequence.start - sequence.end + } else { + sequence.end - sequence.start + } + if span != sequence.sequence.length() { + raise AlignEmbossError("EMBOSS sequence coordinate span is inconsistent") + } + if sequence.start <= sequence.end && sequence.strand != AlignEmbossForward { + raise AlignEmbossError("EMBOSS sequence strand is inconsistent") + } + if sequence.start > sequence.end && sequence.strand != AlignEmbossReverse { + raise AlignEmbossError("EMBOSS sequence strand is inconsistent") + } + } + if alignment.consensus.length() != 0 && + alignment.consensus.length() != alignment.annotations.length { + raise AlignEmbossError("EMBOSS consensus width must match alignment length") + } + for index = 0; index < alignment.consensus.length(); index = index + 1 { + let code = alignment.consensus.unsafe_get(index).to_int() + if code < 32 || code > 126 { + raise AlignEmbossError("EMBOSS consensus contains a control character") + } + } +} + +///| +fn align_emboss_validate_reported_counts( + alignment : AlignEmbossAlignment, +) -> Unit raise AlignEmbossError { + if alignment.sequences.length() != 2 { + return + } + let counts = alignment.pair_counts(0, 1) + match alignment.annotations.identity { + Some(value) => + if value != counts.identities { + raise AlignEmbossError( + "Reported EMBOSS Identity does not match sequence rows", + ) + } + None => () + } + match alignment.annotations.gaps { + Some(value) => + if value != + counts.insertions + counts.deletions + counts.double_gap_columns { + raise AlignEmbossError( + "Reported EMBOSS Gaps does not match sequence rows", + ) + } + None => () + } + match alignment.annotations.similarity { + Some(value) => + if alignment.consensus.length() == alignment.alignment_length() && + value != counts.positives { + raise AlignEmbossError( + "Reported EMBOSS Similarity does not match consensus", + ) + } + None => () + } +} + +///| +fn align_emboss_validate_row( + alignment : AlignEmbossAlignment, + row : Int, +) -> Unit raise AlignEmbossError { + if row < 0 || row >= alignment.sequences.length() { + raise AlignEmbossError("EMBOSS row index is out of range") + } +} + +///| +fn align_emboss_boundary_at( + sequence : AlignEmbossSequence, + column : Int, +) -> Int { + let mut consumed = 0 + for index = 0; index < column; index = index + 1 { + if sequence.aligned_sequence.unsafe_get(index).to_int() != '-'.to_int() { + consumed = consumed + 1 + } + } + if sequence.strand == AlignEmbossForward { + sequence.start + consumed + } else { + sequence.start - consumed + } +} + +///| +fn align_emboss_column_step( + sequence : AlignEmbossSequence, + column : Int, +) -> Int { + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + 0 + } else if sequence.strand == AlignEmbossForward { + 1 + } else { + -1 + } +} + +///| +fn align_emboss_printed_coordinates( + sequence : AlignEmbossSequence, + boundary : Int, + residues : Int, +) -> (Int, Int) { + if sequence.strand == AlignEmbossForward { + if residues == 0 { + (boundary, boundary) + } else { + (boundary + 1, boundary + residues) + } + } else if residues == 0 { + (boundary + 1, boundary + 1) + } else { + (boundary, boundary - residues + 1) + } +} + +///| +fn align_emboss_sequence_prefix(id : String, start : Int) -> String { + let displayed = if id.length() > 13 { id[0:13].to_owned() } else { id } + align_emboss_pad_right(displayed, 13) + + " " + + align_emboss_pad_left(start.to_string(), 6) + + " " +} + +///| +fn align_emboss_pad_left(text : String, width : Int) -> String { + if text.length() >= width { + text + } else { + " ".repeat(width - text.length()) + text + } +} + +///| +fn align_emboss_pad_right(text : String, width : Int) -> String { + if text.length() >= width { + text + } else { + text + " ".repeat(width - text.length()) + } +} + +///| +fn align_emboss_consensus_chunk( + raw : String, + width : Int, + line_number : Int, +) -> String raise AlignEmbossError { + if raw.length() <= width { + raw + " ".repeat(width - raw.length()) + } else { + let extra = raw[width:].to_owned() + if extra.trim().length() > 0 { + raise AlignEmbossError( + "EMBOSS consensus exceeds block width at line " + + line_number.to_string(), + ) + } + raw[0:width].to_owned() + } +} + +///| +fn align_emboss_split_annotation( + content : String, + line_number : Int, +) -> (String, String) raise AlignEmbossError { + let colon = align_emboss_find_char(content, ':'.to_int()) + if colon >= 0 { + return ( + content[0:colon].trim().to_owned(), + content[colon + 1:].trim().to_owned(), + ) + } + let marker = " = " + let equal = align_emboss_find_substring(content, marker) + if equal >= 0 { + return ( + content[0:equal].trim().to_owned(), + content[equal + marker.length():].trim().to_owned(), + ) + } + raise AlignEmbossError( + "Malformed EMBOSS annotation at line " + line_number.to_string(), + ) +} + +///| +fn align_emboss_split_key_value( + content : String, + context : String, + line_number : Int, +) -> (String, String) raise AlignEmbossError { + let separator = align_emboss_find_char(content, ':'.to_int()) + if separator <= 0 { + raise AlignEmbossError( + "Malformed EMBOSS " + context + " at line " + line_number.to_string(), + ) + } + let key = content[0:separator].trim().to_owned() + let value = content[separator + 1:].trim().to_owned() + if key.length() == 0 || value.length() == 0 { + raise AlignEmbossError( + "Empty EMBOSS " + context + " field at line " + line_number.to_string(), + ) + } + (key, value) +} + +///| +fn align_emboss_parse_fraction_count( + text : String, + field : String, + line_number : Int, +) -> Int raise AlignEmbossError { + let slash = align_emboss_find_char(text, '/'.to_int()) + let count_text = if slash < 0 { + text + } else { + text[0:slash].trim().to_owned() + } + align_emboss_parse_nonnegative_int(count_text, field, line_number) +} + +///| +fn align_emboss_parse_positive_int( + text : String, + field : String, + line_number : Int, +) -> Int raise AlignEmbossError { + let value = align_emboss_parse_nonnegative_int(text, field, line_number) + if value <= 0 { + raise AlignEmbossError( + "EMBOSS " + field + " must be positive at line " + line_number.to_string(), + ) + } + value +} + +///| +fn align_emboss_parse_nonnegative_int( + text : String, + field : String, + line_number : Int, +) -> Int raise AlignEmbossError { + if text.length() == 0 { + raise AlignEmbossError( + "Empty EMBOSS " + field + " at line " + line_number.to_string(), + ) + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignEmbossError( + "Invalid EMBOSS " + field + " at line " + line_number.to_string(), + ) + } + value = value * 10 + code - '0'.to_int() + } + value +} + +///| +fn align_emboss_parse_number( + text : String, + field : String, + line_number : Int, +) -> Double raise AlignEmbossError { + if text.length() == 0 { + raise AlignEmbossError( + "Empty EMBOSS " + field + " at line " + line_number.to_string(), + ) + } + let mut has_digit = false + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + has_digit = true + } else if code != '-'.to_int() && + code != '+'.to_int() && + code != '.'.to_int() && + code != 'e'.to_int() && + code != 'E'.to_int() { + raise AlignEmbossError( + "Invalid EMBOSS " + field + " at line " + line_number.to_string(), + ) + } + } + if !has_digit { + raise AlignEmbossError( + "Invalid EMBOSS " + field + " at line " + line_number.to_string(), + ) + } + match parse_double(text) { + Some(value) => + if value.abs() > 1.0e300 { + raise AlignEmbossError( + "Non-finite EMBOSS " + field + " at line " + line_number.to_string(), + ) + } else { + value + } + None => + raise AlignEmbossError( + "Invalid EMBOSS " + field + " at line " + line_number.to_string(), + ) + } +} + +///| +fn align_emboss_validate_id(id : String) -> Unit raise AlignEmbossError { + if id.length() == 0 { + raise AlignEmbossError("EMBOSS sequence identifier must not be empty") + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if code <= 32 || code == 127 { + raise AlignEmbossError( + "EMBOSS sequence identifier must not contain whitespace", + ) + } + } +} + +///| +fn align_emboss_validate_header_value( + value : String, + field : String, +) -> Unit raise AlignEmbossError { + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code == '\n'.to_int() || code == '\r'.to_int() { + raise AlignEmbossError("EMBOSS " + field + " must not contain a newline") + } + } +} + +///| +fn align_emboss_validate_aligned_fragment( + fragment : String, + context : String, +) -> Unit raise AlignEmbossError { + if fragment.length() == 0 { + raise AlignEmbossError("Empty EMBOSS aligned fragment in " + context) + } + for index = 0; index < fragment.length(); index = index + 1 { + let code = fragment.unsafe_get(index).to_int() + if code <= 32 || code > 126 { + raise AlignEmbossError("Invalid EMBOSS aligned character in " + context) + } + } +} + +///| +fn align_emboss_validate_optional_count( + value : Int?, + length : Int, + field : String, +) -> Unit raise AlignEmbossError { + match value { + Some(count) => + if count < 0 || count > length { + raise AlignEmbossError( + "EMBOSS " + field + " must be between zero and Length", + ) + } + None => () + } +} + +///| +fn align_emboss_validate_optional_nonnegative_double( + value : Double?, + field : String, +) -> Unit raise AlignEmbossError { + match value { + Some(number) => + if number < 0.0 || number.abs() > 1.0e300 { + raise AlignEmbossError( + "EMBOSS " + field + " must be finite and non-negative", + ) + } + None => () + } +} + +///| +fn align_emboss_remove_gaps(sequence : String) -> String { + let result = StringBuilder::new(size_hint=sequence.length()) + for index = 0; index < sequence.length(); index = index + 1 { + let code = sequence.unsafe_get(index).to_int() + if code != '-'.to_int() { + result.write_char(code.unsafe_to_char()) + } + } + result.to_string() +} + +///| +fn align_emboss_count_residues(sequence : String) -> Int { + let mut count = 0 + for index = 0; index < sequence.length(); index = index + 1 { + if sequence.unsafe_get(index).to_int() != '-'.to_int() { + count = count + 1 + } + } + count +} + +///| +fn align_emboss_lines(text : String) -> Array[String] { + let raw = text.split("\n") + let lines : Array[String] = [] + for line in raw { + let owned = line.to_owned() + if owned.length() > 0 && + owned.unsafe_get(owned.length() - 1).to_int() == '\r'.to_int() { + lines.push(owned[0:owned.length() - 1].to_owned()) + } else { + lines.push(owned) + } + } + lines +} + +///| +fn align_emboss_split_whitespace(text : String) -> Array[String] { + let result : Array[String] = [] + let mut index = 0 + while index < text.length() { + while index < text.length() && + align_emboss_is_whitespace(text.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= text.length() { + break + } + let start = index + while index < text.length() && + !align_emboss_is_whitespace(text.unsafe_get(index).to_int()) { + index = index + 1 + } + result.push(text[start:index].to_owned()) + } + result +} + +///| +fn align_emboss_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} + +///| +fn align_emboss_has_prefix(text : String, prefix : String) -> Bool { + if prefix.length() > text.length() { + return false + } + for index = 0; index < prefix.length(); index = index + 1 { + if text.unsafe_get(index).to_int() != prefix.unsafe_get(index).to_int() { + return false + } + } + true +} + +///| +fn align_emboss_is_separator(text : String) -> Bool { + text == "#---------------------------------------" +} + +///| +fn align_emboss_find_char(text : String, target : Int) -> Int { + for index = 0; index < text.length(); index = index + 1 { + if text.unsafe_get(index).to_int() == target { + return index + } + } + -1 +} + +///| +fn align_emboss_find_substring(text : String, target : String) -> Int { + if target.length() == 0 { + return 0 + } + if target.length() > text.length() { + return -1 + } + for start = 0; start <= text.length() - target.length(); start = start + 1 { + let mut equal = true + for offset = 0; offset < target.length(); offset = offset + 1 { + if text.unsafe_get(start + offset).to_int() != + target.unsafe_get(offset).to_int() { + equal = false + break + } + } + if equal { + return start + } + } + -1 +} + +///| +fn align_emboss_write_optional_double( + output : StringBuilder, + key : String, + value : Double?, +) -> Unit { + match value { + Some(number) => + output.write_string("# " + key + ": " + number.to_string() + "\n") + None => () + } +} + +///| +fn align_emboss_write_optional_count( + output : StringBuilder, + key : String, + value : Int?, + length : Int, +) -> Unit { + match value { + Some(count) => + output.write_string( + "# " + key + ": " + count.to_string() + "/" + length.to_string() + "\n", + ) + None => () + } +} + +///| +fn align_emboss_write_long_annotation( + output : StringBuilder, + key : String, + value : String, +) -> Unit { + if value.length() > 0 { + output.write_string("# " + key + " = " + value + "\n") + } +} diff --git a/src/align_exonerate.mbt b/src/align_exonerate.mbt new file mode 100644 index 00000000..7f78318b --- /dev/null +++ b/src/align_exonerate.mbt @@ -0,0 +1,1319 @@ +// Alignment-aware Exonerate cigar/vulgar support. +// +// This module follows Biopython 1.86 Bio.Align.exonerate semantics. Exonerate +// coordinates are zero-based boundaries. Operations keep query and target +// steps separately so DNA/protein 3:1 alignments do not lose information. + +///| +/// Error raised for malformed alignment-aware Exonerate data. +pub suberror AlignExonerateError { + AlignExonerateError(String) +} + +///| +/// Orientation or molecule type reported for an Exonerate sequence. +pub(all) enum AlignExonerateStrand { + AlignExonerateForward + AlignExonerateReverse + AlignExonerateProtein +} derive(Eq, Debug) + +///| +/// File-level Exonerate metadata. +pub struct AlignExonerateMetadata { + program : String + command_line : String + hostname : String +} derive(Eq, Debug) + +///| +/// One normalized Exonerate path operation. +/// +/// Codes use alignment-aware semantics: +/// `M` match, `5`/`3` splice sites, `N` intron, `C` codon, `D` deletion, +/// `I` insertion, `U` non-equivalenced region, `S` split codon, and `F` +/// frame shift. Steps are non-negative molecular coordinates. +pub struct AlignExonerateOperation { + code : String + query_step : Int + target_step : Int +} derive(Eq, Debug) + +///| +/// One pairwise alignment from an Exonerate report. +pub struct AlignExonerateAlignment { + query_id : String + query_start : Int + query_end : Int + query_strand : AlignExonerateStrand + target_id : String + target_start : Int + target_end : Int + target_strand : AlignExonerateStrand + score : Double + operations : Array[AlignExonerateOperation] +} derive(Eq, Debug) + +///| +/// A complete Exonerate report. +pub struct AlignExonerateDocument { + metadata : AlignExonerateMetadata + alignments : Array[AlignExonerateAlignment] +} derive(Eq, Debug) + +///| +/// Operation and coordinate statistics for one Exonerate alignment. +pub struct AlignExonerateCounts { + operations : Int + matching_operations : Int + aligned_query_units : Int + aligned_target_units : Int + insertions : Int + deletions : Int + insertion_opens : Int + deletion_opens : Int + introns : Int + intron_units : Int + non_equivalenced_units : Int + split_codons : Int + frameshifts : Int +} derive(Eq, Debug) + +///| +/// Construct Exonerate file metadata. +pub fn AlignExonerateMetadata::create( + command_line? : String = "", + hostname? : String = "", +) -> AlignExonerateMetadata raise AlignExonerateError { + align_exonerate_validate_header_value(command_line, "Command line") + align_exonerate_validate_header_value(hostname, "Hostname") + AlignExonerateMetadata::{ program: "exonerate", command_line, hostname } +} + +///| +/// Construct and validate one normalized Exonerate operation. +pub fn AlignExonerateOperation::create( + code : String, + query_step : Int, + target_step : Int, +) -> AlignExonerateOperation raise AlignExonerateError { + align_exonerate_validate_operation(code, query_step, target_step) + AlignExonerateOperation::{ code, query_step, target_step } +} + +///| +/// Construct one alignment and validate operation totals against its bounds. +pub fn AlignExonerateAlignment::create( + query_id : String, + query_start : Int, + query_end : Int, + query_strand : AlignExonerateStrand, + target_id : String, + target_start : Int, + target_end : Int, + target_strand : AlignExonerateStrand, + score : Double, + operations : Array[AlignExonerateOperation], +) -> AlignExonerateAlignment raise AlignExonerateError { + align_exonerate_validate_id(query_id, "query") + align_exonerate_validate_id(target_id, "target") + align_exonerate_validate_axis(query_start, query_end, query_strand, "query") + align_exonerate_validate_axis( + target_start, target_end, target_strand, "target", + ) + if score.abs() > 1.0e300 { + raise AlignExonerateError("Exonerate score must be finite") + } + if operations.length() == 0 { + raise AlignExonerateError("Exonerate alignment must contain an operation") + } + let copied : Array[AlignExonerateOperation] = [] + let mut query_span = 0 + let mut target_span = 0 + for operation in operations { + align_exonerate_validate_operation( + operation.code, + operation.query_step, + operation.target_step, + ) + query_span = query_span + operation.query_step + target_span = target_span + operation.target_step + copied.push(operation) + } + if query_span != align_exonerate_abs(query_end - query_start) { + raise AlignExonerateError( + "Exonerate query operation span does not match its coordinates", + ) + } + if target_span != align_exonerate_abs(target_end - target_start) { + raise AlignExonerateError( + "Exonerate target operation span does not match its coordinates", + ) + } + AlignExonerateAlignment::{ + query_id, + query_start, + query_end, + query_strand, + target_id, + target_start, + target_end, + target_strand, + score, + operations: copied, + } +} + +///| +/// Construct a complete report. Exonerate reports with no hits are valid. +pub fn AlignExonerateDocument::create( + metadata : AlignExonerateMetadata, + alignments : Array[AlignExonerateAlignment], +) -> AlignExonerateDocument raise AlignExonerateError { + if metadata.program != "exonerate" { + raise AlignExonerateError("Exonerate metadata program must be exonerate") + } + let copied : Array[AlignExonerateAlignment] = [] + for alignment in alignments { + let validated = AlignExonerateAlignment::create( + alignment.query_id, + alignment.query_start, + alignment.query_end, + alignment.query_strand, + alignment.target_id, + alignment.target_start, + alignment.target_end, + alignment.target_strand, + alignment.score, + alignment.operations, + ) + copied.push(validated) + } + AlignExonerateDocument::{ metadata, alignments: copied } +} + +///| +pub fn AlignExonerateDocument::num_alignments( + self : AlignExonerateDocument, +) -> Int { + self.alignments.length() +} + +///| +pub fn AlignExonerateDocument::get( + self : AlignExonerateDocument, + index : Int, +) -> AlignExonerateAlignment? { + if index < 0 || index >= self.alignments.length() { + None + } else { + Some(self.alignments[index]) + } +} + +///| +pub fn AlignExonerateDocument::find_by_query( + self : AlignExonerateDocument, + query_id : String, +) -> Array[AlignExonerateAlignment] { + let result : Array[AlignExonerateAlignment] = [] + for alignment in self.alignments { + if alignment.query_id == query_id { + result.push(alignment) + } + } + result +} + +///| +pub fn AlignExonerateDocument::best_alignment( + self : AlignExonerateDocument, +) -> AlignExonerateAlignment? { + if self.alignments.length() == 0 { + return None + } + let mut best = self.alignments[0] + for index = 1; index < self.alignments.length(); index = index + 1 { + if self.alignments[index].score > best.score { + best = self.alignments[index] + } + } + Some(best) +} + +///| +pub fn AlignExonerateDocument::summary(self : AlignExonerateDocument) -> String { + "Exonerate report(host=" + + self.metadata.hostname + + ", alignments=" + + self.alignments.length().to_string() + + ")" +} + +///| +pub fn AlignExonerateAlignment::query_length( + self : AlignExonerateAlignment, +) -> Int { + if self.query_strand == AlignExonerateReverse { + self.query_start + } else { + self.query_end + } +} + +///| +pub fn AlignExonerateAlignment::target_length( + self : AlignExonerateAlignment, +) -> Int { + if self.target_strand == AlignExonerateReverse { + self.target_start + } else { + self.target_end + } +} + +///| +pub fn AlignExonerateAlignment::num_operations( + self : AlignExonerateAlignment, +) -> Int { + self.operations.length() +} + +///| +pub fn AlignExonerateAlignment::operation_codes( + self : AlignExonerateAlignment, +) -> String { + let output = StringBuilder::new(size_hint=self.operations.length()) + for operation in self.operations { + output.write_string(operation.code) + } + output.to_string() +} + +///| +/// Return `[target_coordinates, query_coordinates]` boundary paths. +pub fn AlignExonerateAlignment::coordinate_path( + self : AlignExonerateAlignment, +) -> Array[Array[Int]] { + let target = [self.target_start] + let query = [self.query_start] + let target_sign = align_exonerate_strand_sign(self.target_strand) + let query_sign = align_exonerate_strand_sign(self.query_strand) + let mut target_position = self.target_start + let mut query_position = self.query_start + for operation in self.operations { + target_position = target_position + target_sign * operation.target_step + query_position = query_position + query_sign * operation.query_step + target.push(target_position) + query.push(query_position) + } + [target, query] +} + +///| +/// Map an absolute query residue to the corresponding target residue. +/// +/// A target gap returns `None`. For translated alignments, a protein residue +/// maps to the first nucleotide of its codon and each nucleotide maps to its +/// containing protein residue. +pub fn AlignExonerateAlignment::query_position_to_target( + self : AlignExonerateAlignment, + position : Int, +) -> Int? { + align_exonerate_map_position(self, position, true) +} + +///| +/// Map an absolute target residue to the corresponding query residue. +pub fn AlignExonerateAlignment::target_position_to_query( + self : AlignExonerateAlignment, + position : Int, +) -> Int? { + align_exonerate_map_position(self, position, false) +} + +///| +/// Return target/query residue pairs for all query units in aligned segments. +pub fn AlignExonerateAlignment::aligned_pairs( + self : AlignExonerateAlignment, +) -> Array[(Int, Int)] { + let pairs : Array[(Int, Int)] = [] + let query_sign = align_exonerate_strand_sign(self.query_strand) + let target_sign = align_exonerate_strand_sign(self.target_strand) + let mut query_boundary = self.query_start + let mut target_boundary = self.target_start + for operation in self.operations { + if operation.query_step > 0 && operation.target_step > 0 { + for offset = 0; offset < operation.query_step; offset = offset + 1 { + let query_position = if query_sign > 0 { + query_boundary + offset + } else { + query_boundary - 1 - offset + } + let target_offset = offset * + operation.target_step / + operation.query_step + let target_position = if target_sign > 0 { + target_boundary + target_offset + } else { + target_boundary - 1 - target_offset + } + pairs.push((target_position, query_position)) + } + } + query_boundary = query_boundary + query_sign * operation.query_step + target_boundary = target_boundary + target_sign * operation.target_step + } + pairs +} + +///| +/// Calculate operation-aware alignment statistics. +pub fn AlignExonerateAlignment::counts( + self : AlignExonerateAlignment, +) -> AlignExonerateCounts { + let mut matching_operations = 0 + let mut aligned_query_units = 0 + let mut aligned_target_units = 0 + let mut insertions = 0 + let mut deletions = 0 + let mut insertion_opens = 0 + let mut deletion_opens = 0 + let mut introns = 0 + let mut intron_units = 0 + let mut non_equivalenced_units = 0 + let mut split_codons = 0 + let mut frameshifts = 0 + let mut previous_insertion = false + let mut previous_deletion = false + for operation in self.operations { + if operation.query_step > 0 && operation.target_step > 0 { + aligned_query_units = aligned_query_units + operation.query_step + aligned_target_units = aligned_target_units + operation.target_step + if operation.code == "M" || operation.code == "C" { + matching_operations = matching_operations + 1 + } + } + let insertion = operation.target_step == 0 && operation.query_step > 0 + let deletion = operation.query_step == 0 && operation.target_step > 0 + if insertion { + insertions = insertions + operation.query_step + if !previous_insertion { + insertion_opens = insertion_opens + 1 + } + } + if deletion { + deletions = deletions + operation.target_step + if !previous_deletion { + deletion_opens = deletion_opens + 1 + } + } + if operation.code == "N" { + introns = introns + 1 + intron_units = intron_units + operation.query_step + operation.target_step + } else if operation.code == "U" { + non_equivalenced_units = non_equivalenced_units + + operation.query_step + + operation.target_step + } else if operation.code == "S" { + split_codons = split_codons + 1 + } else if operation.code == "F" { + frameshifts = frameshifts + 1 + } + previous_insertion = insertion + previous_deletion = deletion + } + AlignExonerateCounts::{ + operations: self.operations.length(), + matching_operations, + aligned_query_units, + aligned_target_units, + insertions, + deletions, + insertion_opens, + deletion_opens, + introns, + intron_units, + non_equivalenced_units, + split_codons, + frameshifts, + } +} + +///| +pub fn AlignExonerateAlignment::summary( + self : AlignExonerateAlignment, +) -> String { + "Exonerate alignment(query=" + + self.query_id + + ", target=" + + self.target_id + + ", score=" + + self.score.to_string() + + ", operations=" + + self.operations.length().to_string() + + ")" +} + +///| +/// Parse an alignment-aware Exonerate cigar/vulgar report. +pub fn align_exonerate_parse( + text : String, +) -> AlignExonerateDocument raise AlignExonerateError { + let lines = align_exonerate_lines(text) + if lines.length() < 3 { + raise AlignExonerateError("Truncated Exonerate report header") + } + let command_line = align_exonerate_parse_header_line( + lines[0], + "Command line: ", + "Command line", + 1, + ) + let hostname = align_exonerate_parse_header_line( + lines[1], + "Hostname: ", + "Hostname", + 2, + ) + let metadata = AlignExonerateMetadata::create(command_line~, hostname~) + let alignments : Array[AlignExonerateAlignment] = [] + let mut footer_seen = false + for index = 2; index < lines.length(); index = index + 1 { + let trimmed = lines[index].trim().to_owned() + if footer_seen { + if trimmed.length() > 0 { + raise AlignExonerateError( + "Additional data after Exonerate completion marker at line " + + (index + 1).to_string(), + ) + } + continue + } + if trimmed.length() == 0 { + continue + } + if trimmed == "-- completed exonerate analysis" { + footer_seen = true + } else if trimmed.has_prefix("vulgar: ") { + alignments.push( + align_exonerate_parse_vulgar_line(trimmed[8:].to_owned(), index + 1), + ) + } else if trimmed.has_prefix("cigar: ") { + alignments.push( + align_exonerate_parse_cigar_line(trimmed[7:].to_owned(), index + 1), + ) + } else { + raise AlignExonerateError( + "Unexpected Exonerate body line at line " + (index + 1).to_string(), + ) + } + } + if !footer_seen { + raise AlignExonerateError( + "Failed to find completed Exonerate analysis marker", + ) + } + AlignExonerateDocument::create(metadata, alignments) +} + +///| +/// Serialize a complete report as canonical `vulgar` or `cigar` output. +pub fn align_exonerate_write( + document : AlignExonerateDocument, + format? : String = "vulgar", +) -> String raise AlignExonerateError { + if format != "vulgar" && format != "cigar" { + raise AlignExonerateError("Exonerate output format must be vulgar or cigar") + } + let validated = AlignExonerateDocument::create( + document.metadata, + document.alignments, + ) + let output = StringBuilder::new() + output.write_string( + "Command line: [" + validated.metadata.command_line + "]\n", + ) + output.write_string("Hostname: [" + validated.metadata.hostname + "]\n") + for alignment in validated.alignments { + if format == "vulgar" { + output.write_string(align_exonerate_format_vulgar(alignment)) + } else { + output.write_string(align_exonerate_format_cigar(alignment)) + } + } + output.write_string("-- completed exonerate analysis\n") + output.to_string() +} + +///| +fn align_exonerate_parse_vulgar_line( + content : String, + line_number : Int, +) -> AlignExonerateAlignment raise AlignExonerateError { + let words = align_exonerate_split_whitespace(content) + if words.length() < 12 || (words.length() - 9) % 3 != 0 { + raise AlignExonerateError( + "Malformed Exonerate vulgar field count at line " + + line_number.to_string(), + ) + } + let query_id = words[0] + let query_start = align_exonerate_parse_int( + words[1], + "query start", + line_number, + ) + let query_end = align_exonerate_parse_int(words[2], "query end", line_number) + let query_strand = align_exonerate_parse_strand( + words[3], + "query", + line_number, + ) + let target_id = words[4] + let target_start = align_exonerate_parse_int( + words[5], + "target start", + line_number, + ) + let target_end = align_exonerate_parse_int( + words[6], + "target end", + line_number, + ) + let target_strand = align_exonerate_parse_strand( + words[7], + "target", + line_number, + ) + let score = align_exonerate_parse_score(words[8], line_number) + let operations : Array[AlignExonerateOperation] = [] + let mut index = 9 + while index < words.length() { + let raw_code = words[index] + let query_step = align_exonerate_parse_int( + words[index + 1], + "query step", + line_number, + ) + let target_step = align_exonerate_parse_int( + words[index + 2], + "target step", + line_number, + ) + if raw_code == "G" { + if query_step == 0 && target_step > 0 { + operations.push(AlignExonerateOperation::create("D", 0, target_step)) + } else if target_step == 0 && query_step > 0 { + operations.push(AlignExonerateOperation::create("I", query_step, 0)) + } else { + raise AlignExonerateError( + "Exonerate vulgar gap must advance exactly one sequence at line " + + line_number.to_string(), + ) + } + } else if raw_code == "I" { + operations.push( + AlignExonerateOperation::create("N", query_step, target_step), + ) + } else if raw_code == "N" { + if target_step > 0 { + operations.push(AlignExonerateOperation::create("U", 0, target_step)) + } + if query_step > 0 { + operations.push(AlignExonerateOperation::create("U", query_step, 0)) + } + if query_step == 0 && target_step == 0 { + raise AlignExonerateError( + "Empty Exonerate non-equivalenced operation at line " + + line_number.to_string(), + ) + } + } else { + operations.push( + AlignExonerateOperation::create(raw_code, query_step, target_step), + ) + } + index = index + 3 + } + AlignExonerateAlignment::create( + query_id, query_start, query_end, query_strand, target_id, target_start, target_end, + target_strand, score, operations, + ) +} + +///| +fn align_exonerate_parse_cigar_line( + content : String, + line_number : Int, +) -> AlignExonerateAlignment raise AlignExonerateError { + let words = align_exonerate_split_whitespace(content) + if words.length() < 11 || (words.length() - 9) % 2 != 0 { + raise AlignExonerateError( + "Malformed Exonerate cigar field count at line " + line_number.to_string(), + ) + } + let query_id = words[0] + let query_start = align_exonerate_parse_int( + words[1], + "query start", + line_number, + ) + let query_end = align_exonerate_parse_int(words[2], "query end", line_number) + let query_strand = align_exonerate_parse_strand( + words[3], + "query", + line_number, + ) + let target_id = words[4] + let target_start = align_exonerate_parse_int( + words[5], + "target start", + line_number, + ) + let target_end = align_exonerate_parse_int( + words[6], + "target end", + line_number, + ) + let target_strand = align_exonerate_parse_strand( + words[7], + "target", + line_number, + ) + let score = align_exonerate_parse_score(words[8], line_number) + let operations : Array[AlignExonerateOperation] = [] + let mut raw_query = 0 + let mut raw_target = 0 + let mut normalized_query = 0 + let mut normalized_target = 0 + let mut index = 9 + while index < words.length() { + let code = words[index] + let step = align_exonerate_parse_positive_int( + words[index + 1], + "cigar step", + line_number, + ) + if code == "M" { + raw_query = raw_query + step + raw_target = raw_target + step + } else if code == "I" { + if query_strand == AlignExonerateProtein && + target_strand != AlignExonerateProtein { + raw_query = raw_query + step * 3 + } else { + raw_query = raw_query + step + } + } else if code == "D" { + if target_strand == AlignExonerateProtein && + query_strand != AlignExonerateProtein { + raw_target = raw_target + step * 3 + } else { + raw_target = raw_target + step + } + } else { + raise AlignExonerateError( + "Unknown Exonerate cigar operation " + + code + + " at line " + + line_number.to_string(), + ) + } + let next_query = align_exonerate_normalize_cigar_offset( + raw_query, query_strand, target_strand, + ) + let next_target = align_exonerate_normalize_cigar_offset( + raw_target, target_strand, query_strand, + ) + let query_step = next_query - normalized_query + let target_step = next_target - normalized_target + operations.push( + AlignExonerateOperation::create(code, query_step, target_step), + ) + normalized_query = next_query + normalized_target = next_target + index = index + 2 + } + AlignExonerateAlignment::create( + query_id, query_start, query_end, query_strand, target_id, target_start, target_end, + target_strand, score, operations, + ) +} + +///| +fn align_exonerate_format_header( + alignment : AlignExonerateAlignment, +) -> Array[String] { + [ + alignment.query_id, + alignment.query_start.to_string(), + alignment.query_end.to_string(), + align_exonerate_strand_symbol(alignment.query_strand), + alignment.target_id, + alignment.target_start.to_string(), + alignment.target_end.to_string(), + align_exonerate_strand_symbol(alignment.target_strand), + alignment.score.to_string(), + ] +} + +///| +fn align_exonerate_format_vulgar( + alignment : AlignExonerateAlignment, +) -> String raise AlignExonerateError { + let words = ["vulgar:"] + for word in align_exonerate_format_header(alignment) { + words.push(word) + } + let mut index = 0 + while index < alignment.operations.length() { + let operation = alignment.operations[index] + if operation.code == "U" { + let mut query_step = 0 + let mut target_step = 0 + while index < alignment.operations.length() && + alignment.operations[index].code == "U" { + query_step = query_step + alignment.operations[index].query_step + target_step = target_step + alignment.operations[index].target_step + index = index + 1 + } + align_exonerate_append_vulgar(words, "N", query_step, target_step) + continue + } + let code = if operation.code == "N" { + "I" + } else if operation.code == "D" || operation.code == "I" { + "G" + } else { + operation.code + } + align_exonerate_append_vulgar( + words, + code, + operation.query_step, + operation.target_step, + ) + index = index + 1 + } + words.join(" ") + "\n" +} + +///| +fn align_exonerate_append_vulgar( + words : Array[String], + code : String, + query_step : Int, + target_step : Int, +) -> Unit { + words.push(code) + words.push(query_step.to_string()) + words.push(target_step.to_string()) +} + +///| +fn align_exonerate_format_cigar( + alignment : AlignExonerateAlignment, +) -> String raise AlignExonerateError { + let words = ["cigar:"] + for word in align_exonerate_format_header(alignment) { + words.push(word) + } + for operation in alignment.operations { + if operation.code == "M" { + let step = if alignment.query_strand == AlignExonerateProtein && + alignment.target_strand != AlignExonerateProtein { + operation.target_step + } else if alignment.target_strand == AlignExonerateProtein && + alignment.query_strand != AlignExonerateProtein { + operation.query_step + } else { + if operation.query_step != operation.target_step { + raise AlignExonerateError( + "Cannot encode unequal non-translated match steps as cigar", + ) + } + operation.query_step + } + align_exonerate_append_cigar(words, "M", step) + } else if operation.code == "5" || operation.code == "3" { + align_exonerate_append_cigar_movement(words, operation) + } else if operation.code == "N" { + align_exonerate_append_cigar_movement(words, operation) + } else if operation.code == "C" { + if operation.query_step != operation.target_step { + raise AlignExonerateError("Cannot encode unequal codon steps as cigar") + } + align_exonerate_append_cigar(words, "M", operation.query_step) + } else if operation.code == "D" { + align_exonerate_append_cigar(words, "D", operation.target_step) + } else if operation.code == "I" { + align_exonerate_append_cigar(words, "I", operation.query_step) + } else if operation.code == "U" || + operation.code == "S" || + operation.code == "F" { + if operation.target_step > 0 { + align_exonerate_append_cigar(words, "D", operation.target_step) + } + if operation.query_step > 0 { + align_exonerate_append_cigar(words, "I", operation.query_step) + } + } else { + raise AlignExonerateError( + "Cannot encode Exonerate operation " + operation.code + " as cigar", + ) + } + } + words.join(" ") + "\n" +} + +///| +fn align_exonerate_append_cigar_movement( + words : Array[String], + operation : AlignExonerateOperation, +) -> Unit raise AlignExonerateError { + if operation.query_step == 0 { + align_exonerate_append_cigar(words, "D", operation.target_step) + } else if operation.target_step == 0 { + align_exonerate_append_cigar(words, "I", operation.query_step) + } else if operation.query_step == operation.target_step { + align_exonerate_append_cigar(words, "M", operation.query_step) + } else { + raise AlignExonerateError( + "Cannot encode two-axis splice/intron movement as cigar", + ) + } +} + +///| +fn align_exonerate_append_cigar( + words : Array[String], + code : String, + step : Int, +) -> Unit raise AlignExonerateError { + if step <= 0 { + raise AlignExonerateError("Exonerate cigar step must be positive") + } + words.push(code) + words.push(step.to_string()) +} + +///| +fn align_exonerate_map_position( + alignment : AlignExonerateAlignment, + position : Int, + from_query : Bool, +) -> Int? { + let query_sign = align_exonerate_strand_sign(alignment.query_strand) + let target_sign = align_exonerate_strand_sign(alignment.target_strand) + let mut query_boundary = alignment.query_start + let mut target_boundary = alignment.target_start + for operation in alignment.operations { + let source_boundary = if from_query { + query_boundary + } else { + target_boundary + } + let source_sign = if from_query { query_sign } else { target_sign } + let source_step = if from_query { + operation.query_step + } else { + operation.target_step + } + let destination_boundary = if from_query { + target_boundary + } else { + query_boundary + } + let destination_sign = if from_query { target_sign } else { query_sign } + let destination_step = if from_query { + operation.target_step + } else { + operation.query_step + } + let offset = if source_sign > 0 { + position - source_boundary + } else { + source_boundary - 1 - position + } + if offset >= 0 && offset < source_step { + if destination_step == 0 { + return None + } + let destination_offset = offset * destination_step / source_step + if destination_sign > 0 { + return Some(destination_boundary + destination_offset) + } + return Some(destination_boundary - 1 - destination_offset) + } + query_boundary = query_boundary + query_sign * operation.query_step + target_boundary = target_boundary + target_sign * operation.target_step + } + None +} + +///| +fn align_exonerate_validate_operation( + code : String, + query_step : Int, + target_step : Int, +) -> Unit raise AlignExonerateError { + if query_step < 0 || target_step < 0 { + raise AlignExonerateError("Exonerate operation steps must be non-negative") + } + if query_step == 0 && target_step == 0 { + raise AlignExonerateError("Exonerate operation must advance a sequence") + } + if code == "M" { + if query_step == 0 || target_step == 0 { + raise AlignExonerateError("Exonerate match must advance both sequences") + } + } else if code == "5" || code == "3" { + if query_step != 2 && target_step != 2 { + raise AlignExonerateError( + "Exonerate splice-site operation must contain a two-unit step", + ) + } + } else if code == "N" { + if (query_step == 0) == (target_step == 0) { + raise AlignExonerateError( + "Exonerate intron must advance exactly one sequence", + ) + } + } else if code == "C" { + if query_step == 0 || + target_step == 0 || + query_step % 3 != 0 || + target_step % 3 != 0 { + raise AlignExonerateError( + "Exonerate codon operation must use positive multiples of three", + ) + } + } else if code == "D" { + if query_step != 0 || target_step <= 0 { + raise AlignExonerateError( + "Exonerate deletion must advance only the target", + ) + } + } else if code == "I" { + if target_step != 0 || query_step <= 0 { + raise AlignExonerateError( + "Exonerate insertion must advance only the query", + ) + } + } else if code == "U" { + if (query_step == 0) == (target_step == 0) { + raise AlignExonerateError( + "Normalized Exonerate U must advance exactly one sequence", + ) + } + } else if code == "S" { + () + } else if code == "F" { + if (query_step == 0) == (target_step == 0) { + raise AlignExonerateError( + "Exonerate frame shift must advance exactly one sequence", + ) + } + } else { + raise AlignExonerateError("Unknown Exonerate operation " + code) + } +} + +///| +fn align_exonerate_validate_axis( + start : Int, + end : Int, + strand : AlignExonerateStrand, + field : String, +) -> Unit raise AlignExonerateError { + if start < 0 || end < 0 { + raise AlignExonerateError( + "Exonerate " + field + " coordinates must be non-negative", + ) + } + if strand == AlignExonerateReverse { + if end >= start { + raise AlignExonerateError( + "Reverse Exonerate " + field + " coordinates must decrease", + ) + } + } else if end <= start { + raise AlignExonerateError( + "Forward/protein Exonerate " + field + " coordinates must increase", + ) + } +} + +///| +fn align_exonerate_validate_id( + id : String, + field : String, +) -> Unit raise AlignExonerateError { + if id.length() == 0 { + raise AlignExonerateError( + "Exonerate " + field + " identifier must not be empty", + ) + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if align_exonerate_is_whitespace(code) || code == 127 { + raise AlignExonerateError( + "Exonerate " + field + " identifier must not contain whitespace", + ) + } + } +} + +///| +fn align_exonerate_validate_header_value( + value : String, + field : String, +) -> Unit raise AlignExonerateError { + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code == '\n'.to_int() || code == '\r'.to_int() { + raise AlignExonerateError( + "Exonerate " + field + " must not contain a newline", + ) + } + } +} + +///| +fn align_exonerate_parse_header_line( + line : String, + prefix : String, + field : String, + line_number : Int, +) -> String raise AlignExonerateError { + if !line.has_prefix(prefix) { + raise AlignExonerateError( + "Missing Exonerate " + field + " at line " + line_number.to_string(), + ) + } + let value = line[prefix.length():].trim().to_owned() + if value.length() < 2 || + value.unsafe_get(0).to_int() != '['.to_int() || + value.unsafe_get(value.length() - 1).to_int() != ']'.to_int() { + raise AlignExonerateError( + "Malformed bracketed Exonerate " + + field + + " at line " + + line_number.to_string(), + ) + } + value[1:value.length() - 1].to_owned() +} + +///| +fn align_exonerate_parse_strand( + text : String, + field : String, + line_number : Int, +) -> AlignExonerateStrand raise AlignExonerateError { + if text == "+" { + AlignExonerateForward + } else if text == "-" { + AlignExonerateReverse + } else if text == "." { + AlignExonerateProtein + } else { + raise AlignExonerateError( + "Invalid Exonerate " + + field + + " strand at line " + + line_number.to_string(), + ) + } +} + +///| +fn align_exonerate_strand_symbol(strand : AlignExonerateStrand) -> String { + if strand == AlignExonerateForward { + "+" + } else if strand == AlignExonerateReverse { + "-" + } else { + "." + } +} + +///| +fn align_exonerate_strand_sign(strand : AlignExonerateStrand) -> Int { + if strand == AlignExonerateReverse { + -1 + } else { + 1 + } +} + +///| +fn align_exonerate_parse_positive_int( + text : String, + field : String, + line_number : Int, +) -> Int raise AlignExonerateError { + let value = align_exonerate_parse_int(text, field, line_number) + if value <= 0 { + raise AlignExonerateError( + "Exonerate " + + field + + " must be positive at line " + + line_number.to_string(), + ) + } + value +} + +///| +fn align_exonerate_parse_int( + text : String, + field : String, + line_number : Int, +) -> Int raise AlignExonerateError { + if text.length() == 0 { + raise AlignExonerateError( + "Empty Exonerate " + field + " at line " + line_number.to_string(), + ) + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignExonerateError( + "Invalid Exonerate " + field + " at line " + line_number.to_string(), + ) + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + raise AlignExonerateError( + "Exonerate " + + field + + " is too large at line " + + line_number.to_string(), + ) + } + value = value * 10 + digit + } + value +} + +///| +fn align_exonerate_parse_score( + text : String, + line_number : Int, +) -> Double raise AlignExonerateError { + let mut has_digit = false + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + has_digit = true + } else if code != '+'.to_int() && + code != '-'.to_int() && + code != '.'.to_int() && + code != 'e'.to_int() && + code != 'E'.to_int() { + raise AlignExonerateError( + "Invalid Exonerate score at line " + line_number.to_string(), + ) + } + } + if !has_digit { + raise AlignExonerateError( + "Invalid Exonerate score at line " + line_number.to_string(), + ) + } + match parse_double(text) { + Some(value) => + if value.abs() > 1.0e300 { + raise AlignExonerateError( + "Non-finite Exonerate score at line " + line_number.to_string(), + ) + } else { + value + } + None => + raise AlignExonerateError( + "Invalid Exonerate score at line " + line_number.to_string(), + ) + } +} + +///| +fn align_exonerate_normalize_cigar_offset( + raw : Int, + strand : AlignExonerateStrand, + other : AlignExonerateStrand, +) -> Int { + if strand == AlignExonerateProtein && other != AlignExonerateProtein { + (raw + 2) / 3 + } else { + raw + } +} + +///| +fn align_exonerate_lines(text : String) -> Array[String] { + let raw = text.split("\n") + let lines : Array[String] = [] + for line in raw { + let owned = line.to_owned() + if owned.length() > 0 && + owned.unsafe_get(owned.length() - 1).to_int() == '\r'.to_int() { + lines.push(owned[0:owned.length() - 1].to_owned()) + } else { + lines.push(owned) + } + } + lines +} + +///| +fn align_exonerate_split_whitespace(text : String) -> Array[String] { + let result : Array[String] = [] + let mut index = 0 + while index < text.length() { + while index < text.length() && + align_exonerate_is_whitespace(text.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= text.length() { + break + } + let start = index + while index < text.length() && + !align_exonerate_is_whitespace(text.unsafe_get(index).to_int()) { + index = index + 1 + } + result.push(text[start:index].to_owned()) + } + result +} + +///| +fn align_exonerate_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() +} + +///| +fn align_exonerate_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +/// Small report covering splicing and translated reverse-strand coordinates. +pub fn align_exonerate_example_text() -> String { + "Command line: [exonerate -m est2genome query.fa target.fa --showvulgar yes]\n" + + "Hostname: [moonbit]\n" + + "vulgar: transcript 0 18 + chromosome 100 218 + 250 M 6 6 5 0 2 I 0 96 3 0 2 M 12 12\n" + + "vulgar: protein 0 4 . genome 500 488 - 80 M 4 12\n" + + "-- completed exonerate analysis\n" +} diff --git a/src/align_maf.mbt b/src/align_maf.mbt new file mode 100644 index 00000000..f05cb40d --- /dev/null +++ b/src/align_maf.mbt @@ -0,0 +1,1418 @@ +// Biopython-compatible Bio.Align.maf support. +// +// The legacy maf.mbt API remains available for permissive block analysis. +// This module models modern MAF documents, strict a/s/i/e/q records, +// absolute coordinate paths, and in-memory MafIndex-style interval queries. + +///| +pub suberror AlignMafError { + AlignMafError(String) +} + +///| +pub struct AlignMafTrack { + name : String? + description : String? + frames : String? + maf_dot : String? + visibility : String? + species_order : Array[String] +} derive(Eq, Debug) + +///| +pub struct AlignMafDocument { + version : String + scoring : String? + program : String? + comments : Array[String] + track : AlignMafTrack? + blocks : Array[BigMafBlock] +} derive(Eq, Debug) + +///| +pub struct AlignMafCoordinate { + column : Int + positions : Array[Int] +} derive(Eq, Debug) + +///| +pub struct AlignMafIndexEntry { + block_index : Int + start : Int + end : Int +} derive(Eq, Debug) + +///| +pub struct AlignMafIndex { + reference : String + source_size : Int + reference_strand : String + entries : Array[AlignMafIndexEntry] + blocks : Array[BigMafBlock] +} derive(Eq, Debug) + +///| +pub struct AlignMafSplicedSequence { + source : String + text : String +} derive(Eq, Debug) + +///| +pub struct AlignMafSplicedAlignment { + reference : String + strand : String + starts : Array[Int] + ends : Array[Int] + sequences : Array[AlignMafSplicedSequence] + columns : Int +} derive(Eq, Debug) + +///| +pub struct AlignMafSummary { + block_count : Int + component_count : Int + empty_component_count : Int + aligned_columns : Int + reference_bases : Int + source_count : Int +} derive(Eq, Debug) + +///| +fn align_maf_fail(message : String) -> Unit raise AlignMafError { + raise AlignMafError(message) +} + +///| +fn align_maf_is_space(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} + +///| +fn align_maf_trim(value : String) -> String { + let mut start = 0 + while start < value.length() && + align_maf_is_space(value.unsafe_get(start).to_int()) { + start = start + 1 + } + let mut end = value.length() + while end > start && align_maf_is_space(value.unsafe_get(end - 1).to_int()) { + end = end - 1 + } + value[start:end].to_owned() +} + +///| +fn align_maf_strip_cr(value : String) -> String { + if value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '\r'.to_int() { + value[0:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn align_maf_lines(value : String) -> Array[String] { + let raw = value.split("\n").to_array() + let lines : Array[String] = [] + for line in raw { + lines.push(align_maf_strip_cr(line.to_owned())) + } + lines +} + +///| +fn align_maf_words(value : String) -> Array[String] { + let words : Array[String] = [] + let mut index = 0 + while index < value.length() { + while index < value.length() && + align_maf_is_space(value.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= value.length() { + break + } + let start = index + while index < value.length() && + !align_maf_is_space(value.unsafe_get(index).to_int()) { + index = index + 1 + } + words.push(value[start:index].to_owned()) + } + words +} + +///| +fn align_maf_shell_words(value : String) -> Array[String] raise AlignMafError { + let words : Array[String] = [] + let mut current = "" + let mut quote = 0 + let mut escaped = false + let mut has_value = false + for index in 0.. (String, String) raise AlignMafError { + let mut equal = -1 + for index in 0.. Unit raise AlignMafError { + if value.length() == 0 { + align_maf_fail(label + " must not be empty") + } + for index in 0.. Unit raise AlignMafError { + align_maf_validate_scalar(value, label) + for index in 0.. Int raise AlignMafError { + if value.length() == 0 { + align_maf_fail(label + " must be an unsigned integer") + } + let mut result = 0 + for index in 0.. '9'.to_int() { + align_maf_fail(label + " must be an unsigned integer") + } + let digit = code - '0'.to_int() + if result > (2147483647 - digit) / 10 { + align_maf_fail(label + " exceeds the supported integer range") + } + result = result * 10 + digit + } + result +} + +///| +fn align_maf_is_number(value : String) -> Bool { + if value.length() == 0 { + return false + } + let mut index = 0 + if value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() { + index = index + 1 + } + let mut digits = 0 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + if index < value.length() && value.unsafe_get(index).to_int() == '.'.to_int() { + index = index + 1 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + digits = digits + 1 + index = index + 1 + } + } + if digits == 0 { + return false + } + if index < value.length() && + ( + value.unsafe_get(index).to_int() == 'e'.to_int() || + value.unsafe_get(index).to_int() == 'E'.to_int() + ) { + index = index + 1 + if index < value.length() && + ( + value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() + ) { + index = index + 1 + } + let exponent_start = index + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + index = index + 1 + } + if index == exponent_start { + return false + } + } + index == value.length() +} + +///| +fn align_maf_parse_double( + value : String, + label : String, +) -> Double raise AlignMafError { + if !align_maf_is_number(value) { + align_maf_fail(label + " must be a number") + } + let parsed_value = if value.unsafe_get(0).to_int() == '+'.to_int() { + value[1:value.length()].to_owned() + } else { + value + } + let number = match parse_double(parsed_value) { + Some(number) => number + None => { + align_maf_fail(label + " must be a number") + 0.0 + } + } + if number.is_nan() || number.abs() > 1.0e300 { + align_maf_fail(label + " must be finite") + } + number +} + +///| +fn align_maf_copy_strings(values : Array[String]) -> Array[String] { + let copied : Array[String] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_maf_copy_blocks(values : Array[BigMafBlock]) -> Array[BigMafBlock] { + let copied : Array[BigMafBlock] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_maf_copy_ints(values : Array[Int]) -> Array[Int] { + let copied : Array[Int] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +pub fn AlignMafTrack::create( + name? : String? = None, + description? : String? = None, + frames? : String? = None, + maf_dot? : String? = None, + visibility? : String? = None, + species_order? : Array[String] = [], +) -> AlignMafTrack raise AlignMafError { + match name { + Some(value) => align_maf_validate_scalar(value, "MAF track name") + None => () + } + match description { + Some(value) => align_maf_validate_scalar(value, "MAF track description") + None => () + } + match frames { + Some(value) => align_maf_validate_scalar(value, "MAF track frames") + None => () + } + match maf_dot { + Some(value) => + if value != "on" && value != "off" { + align_maf_fail("MAF track mafDot must be 'on' or 'off'") + } + None => () + } + match visibility { + Some(value) => + if value != "dense" && value != "pack" && value != "full" { + align_maf_fail( + "MAF track visibility must be 'dense', 'pack', or 'full'", + ) + } + None => () + } + for source in species_order { + align_maf_validate_token(source, "MAF track species") + } + AlignMafTrack::{ + name, + description, + frames, + maf_dot, + visibility, + species_order: align_maf_copy_strings(species_order), + } +} + +///| +pub fn AlignMafDocument::create( + blocks : Array[BigMafBlock], + version? : String = "1", + scoring? : String? = None, + program? : String? = None, + comments? : Array[String] = [], + track? : AlignMafTrack? = None, +) -> AlignMafDocument raise AlignMafError { + if version != "1" { + align_maf_fail("MAF version must be 1") + } + match scoring { + Some(value) => align_maf_validate_token(value, "MAF scoring metadata") + None => () + } + match program { + Some(value) => align_maf_validate_token(value, "MAF program metadata") + None => () + } + for comment in comments { + align_maf_validate_scalar(comment, "MAF comment") + } + AlignMafDocument::{ + version, + scoring, + program, + comments: align_maf_copy_strings(comments), + track, + blocks: align_maf_copy_blocks(blocks), + } +} + +///| +fn align_maf_parse_track(line : String) -> AlignMafTrack raise AlignMafError { + let words = align_maf_shell_words(line) + if words.length() < 2 || words[0] != "track" { + align_maf_fail("Malformed MAF track line") + } + let mut name : String? = None + let mut description : String? = None + let mut frames : String? = None + let mut maf_dot : String? = None + let mut visibility : String? = None + let species_order : Array[String] = [] + for index in 1.. 0 { + align_maf_fail("Duplicate MAF track speciesOrder") + } + for source in align_maf_words(value) { + species_order.push(source) + } + if species_order.length() == 0 { + align_maf_fail("MAF track speciesOrder must not be empty") + } + } else { + align_maf_fail("Unexpected MAF track variable '" + key + "'") + } + } + AlignMafTrack::create( + name~, + description~, + frames~, + maf_dot~, + visibility~, + species_order~, + ) +} + +///| +fn align_maf_normalize_text(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index in 0.. BigMafBlock raise AlignMafError { + if lines.length() == 0 { + align_maf_fail("MAF alignment block must not be empty") + } + let header = align_maf_words(align_maf_trim(lines[0])) + if header.length() == 0 || header[0] != "a" { + align_maf_fail("MAF alignment block must start with an 'a' line") + } + let mut score : Double? = None + let mut pass_number : Int? = None + for index in 1.. source_size || + component_size > source_size - component_start { + align_maf_fail("MAF component coordinates exceed the source sequence") + } + let component = BigMafComponent::create( + words[1], + component_start, + component_size, + words[4], + source_size, + align_maf_normalize_text(words[6]), + ) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } + components.push(component) + last_component = components.length() - 1 + } else if words[0] == "i" { + if last_component < 0 || words.length() != 6 { + align_maf_fail( + "MAF insertion lines must follow an s line and have six fields", + ) + } + if words[1] != components[last_component].source { + align_maf_fail( + "MAF insertion source does not match the preceding s line", + ) + } + if components[last_component].insertion is Some(_) { + align_maf_fail("Duplicate MAF insertion line for one component") + } + let left_count = align_maf_parse_uint( + words[3], + "MAF left insertion count", + ) + let right_count = align_maf_parse_uint( + words[5], + "MAF right insertion count", + ) + let insertion = BigMafInsertion::create( + words[2], + left_count, + words[4], + right_count, + ) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } + components[last_component] = components[last_component].with_insertion( + insertion, + ) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } + } else if words[0] == "q" { + if last_component < 0 || words.length() != 3 { + align_maf_fail( + "MAF quality lines must follow an s line and have three fields", + ) + } + if words[1] != components[last_component].source { + align_maf_fail("MAF quality source does not match the preceding s line") + } + if components[last_component].quality is Some(_) { + align_maf_fail("Duplicate MAF quality line for one component") + } + components[last_component] = components[last_component].with_quality( + words[2], + ) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } + } else if words[0] == "e" { + if words.length() != 7 { + align_maf_fail("MAF empty component lines must contain seven fields") + } + let empty_start = align_maf_parse_uint( + words[2], + "MAF empty component start", + ) + let empty_size = align_maf_parse_uint( + words[3], + "MAF empty component size", + ) + let empty_source_size = align_maf_parse_uint( + words[5], + "MAF empty source size", + ) + if empty_source_size <= 0 || + empty_start > empty_source_size || + empty_size > empty_source_size - empty_start { + align_maf_fail( + "MAF empty component coordinates exceed the source sequence", + ) + } + let empty = BigMafEmptyComponent::create( + words[1], + empty_start, + empty_size, + words[4], + empty_source_size, + words[6], + ) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } + empty_components.push(empty) + last_component = -1 + } else { + align_maf_fail("Unexpected MAF line type '" + words[0] + "'") + } + } + BigMafBlock::create(components, score~, pass_number~, empty_components~) catch { + BigMafError(message) => { + align_maf_fail(message) + abort("") + } + } +} + +///| +pub fn align_maf_parse(text : String) -> AlignMafDocument raise AlignMafError { + let lines = align_maf_lines(text) + let mut index = 0 + while index < lines.length() && align_maf_trim(lines[index]).length() == 0 { + index = index + 1 + } + if index >= lines.length() { + align_maf_fail("MAF input is empty") + } + let mut track : AlignMafTrack? = None + let first = align_maf_trim(lines[index]) + if first == "track" || first.has_prefix("track ") { + track = Some(align_maf_parse_track(first)) + index = index + 1 + } + while index < lines.length() && align_maf_trim(lines[index]).length() == 0 { + index = index + 1 + } + if index >= lines.length() { + align_maf_fail("MAF input is missing its header") + } + let header = align_maf_words(align_maf_trim(lines[index])) + if header.length() == 0 || header[0] != "##maf" { + align_maf_fail("MAF header line must start with ##maf") + } + let mut version : String? = None + let mut scoring : String? = None + let mut program : String? = None + for field_index in 1.. value + None => { + align_maf_fail("MAF header must declare version=1") + "" + } + } + if found_version != "1" { + align_maf_fail("MAF version must be 1") + } + index = index + 1 + let comments : Array[String] = [] + let blocks : Array[BigMafBlock] = [] + let mut current : Array[String] = [] + let mut saw_block = false + while index < lines.length() { + let line = align_maf_trim(lines[index]) + if line.length() == 0 { + index = index + 1 + continue + } + if line.unsafe_get(0).to_int() == '#'.to_int() && + current.length() == 0 && + !saw_block { + comments.push(align_maf_trim(line[1:line.length()].to_owned())) + } else if line == "a" || line.has_prefix("a ") || line.has_prefix("a\t") { + if current.length() > 0 { + blocks.push(align_maf_parse_block(current)) + } + current = [line] + saw_block = true + } else { + if current.length() == 0 { + align_maf_fail("MAF data line appears before an alignment line") + } + current.push(line) + } + index = index + 1 + } + if current.length() > 0 { + blocks.push(align_maf_parse_block(current)) + } + AlignMafDocument::create( + blocks, + version=found_version, + scoring~, + program~, + comments~, + track~, + ) +} + +///| +fn align_maf_quote_track(value : String) -> String { + let mut needs_quotes = false + let output = StringBuilder::new(size_hint=value.length()) + for index in 0.. Unit { + output.write_string("track") + match track.name { + Some(value) => output.write_string(" name=" + align_maf_quote_track(value)) + None => () + } + match track.description { + Some(value) => + output.write_string(" description=" + align_maf_quote_track(value)) + None => () + } + match track.frames { + Some(value) => + output.write_string(" frames=" + align_maf_quote_track(value)) + None => () + } + match track.maf_dot { + Some(value) => output.write_string(" mafDot=" + value) + None => () + } + match track.visibility { + Some(value) => output.write_string(" visibility=" + value) + None => () + } + if track.species_order.length() > 0 { + let species = StringBuilder::new() + for index in 0.. 0 { + species.write_char(' ') + } + species.write_string(track.species_order[index]) + } + output.write_string( + " speciesOrder=" + align_maf_quote_track(species.to_string()), + ) + } + output.write_char('\n') +} + +///| +pub fn align_maf_write( + document : AlignMafDocument, +) -> String raise AlignMafError { + let validated = AlignMafDocument::create( + document.blocks, + version=document.version, + scoring=document.scoring, + program=document.program, + comments=document.comments, + track=document.track, + ) + let output = StringBuilder::new() + match validated.track { + Some(track) => align_maf_write_track(output, track) + None => () + } + output.write_string("##maf version=1") + match validated.scoring { + Some(value) => output.write_string(" scoring=" + value) + None => () + } + match validated.program { + Some(value) => output.write_string(" program=" + value) + None => () + } + output.write_char('\n') + for comment in validated.comments { + output.write_string("# " + comment + "\n") + } + output.write_char('\n') + for block in validated.blocks { + output.write_string(block.to_maf()) + } + output.to_string() +} + +///| +fn align_maf_pattern_equal(left : Array[Bool], right : Array[Bool]) -> Bool { + if left.length() != right.length() { + return false + } + for index in 0.. Array[Bool] { + let pattern : Array[Bool] = [] + for component in block.components { + pattern.push(component.text.unsafe_get(column).to_int() != '-'.to_int()) + } + pattern +} + +///| +fn align_maf_copy_positions(positions : Array[Int]) -> Array[Int] { + let copied : Array[Int] = [] + for position in positions { + copied.push(position) + } + copied +} + +///| +pub fn align_maf_coordinate_path( + block : BigMafBlock, +) -> Array[AlignMafCoordinate] { + let width = block.aligned_columns() + let positions : Array[Int] = [] + for component in block.components { + positions.push( + if component.strand == "+" { + component.start + } else { + component.source_size - component.start + }, + ) + } + let path : Array[AlignMafCoordinate] = [ + AlignMafCoordinate::{ + column: 0, + positions: align_maf_copy_positions(positions), + }, + ] + if width == 0 { + return path + } + let mut previous = align_maf_column_pattern(block, 0) + for column in 0.. 0 && !align_maf_pattern_equal(previous, pattern) { + path.push(AlignMafCoordinate::{ + column, + positions: align_maf_copy_positions(positions), + }) + } + for row in 0.. Int? raise AlignMafError { + let from = match block.component(from_source) { + Some(component) => component + None => return None + } + let to = match block.component(to_source) { + Some(component) => component + None => return None + } + match from.source_to_column(position) { + Some(column) => + to.column_to_source(column) catch { + BigMafError(message) => { + align_maf_fail(message) + None + } + } + None => None + } +} + +///| +pub fn AlignMafDocument::sources(self : AlignMafDocument) -> Array[String] { + let sources : Array[String] = [] + for block in self.blocks { + for component in block.components { + if !sources.contains(component.source) { + sources.push(component.source) + } + } + for empty in block.empty_components { + if !sources.contains(empty.source) { + sources.push(empty.source) + } + } + } + sources +} + +///| +pub fn AlignMafDocument::summary(self : AlignMafDocument) -> AlignMafSummary { + let mut component_count = 0 + let mut empty_component_count = 0 + let mut aligned_columns = 0 + let mut reference_bases = 0 + for block in self.blocks { + component_count = component_count + block.components.length() + empty_component_count = empty_component_count + + block.empty_components.length() + aligned_columns = aligned_columns + block.aligned_columns() + if block.components.length() > 0 { + reference_bases = reference_bases + block.components[0].size + } + } + AlignMafSummary::{ + block_count: self.blocks.length(), + component_count, + empty_component_count, + aligned_columns, + reference_bases, + source_count: self.sources().length(), + } +} + +///| +fn align_maf_entry_sort(entries : Array[AlignMafIndexEntry]) -> Unit { + for index in 1.. 0 && + ( + entries[position - 1].start > value.start || + ( + entries[position - 1].start == value.start && + entries[position - 1].end > value.end + ) + ) { + entries[position] = entries[position - 1] + position = position - 1 + } + entries[position] = value + } +} + +///| +pub fn AlignMafIndex::create( + document : AlignMafDocument, + reference : String, +) -> AlignMafIndex raise AlignMafError { + align_maf_validate_token(reference, "MAF index reference") + if document.blocks.length() == 0 { + align_maf_fail("MAF index requires at least one alignment block") + } + let entries : Array[AlignMafIndexEntry] = [] + let mut source_size = -1 + let mut reference_strand = "" + for block_index in 0.. value + None => { + align_maf_fail("MAF index reference is missing from an alignment block") + block.components[0] + } + } + if source_size < 0 { + source_size = component.source_size + reference_strand = component.strand + } else { + if component.source_size != source_size { + align_maf_fail("MAF index reference source sizes are inconsistent") + } + if component.strand != reference_strand { + align_maf_fail("MAF index reference strands are inconsistent") + } + } + let (start, end) = component.forward_interval() + entries.push(AlignMafIndexEntry::{ block_index, start, end }) + } + align_maf_entry_sort(entries) + AlignMafIndex::{ + reference, + source_size, + reference_strand, + entries, + blocks: align_maf_copy_blocks(document.blocks), + } +} + +///| +fn align_maf_validate_range( + start : Int, + end : Int, + source_size : Int, +) -> Unit raise AlignMafError { + if start < 0 || end <= start || end > source_size { + align_maf_fail( + "MAF query interval must satisfy 0 <= start < end <= source size", + ) + } +} + +///| +pub fn AlignMafIndex::search( + self : AlignMafIndex, + start : Int, + end : Int, +) -> Array[BigMafBlock] raise AlignMafError { + align_maf_validate_range(start, end, self.source_size) + let blocks : Array[BigMafBlock] = [] + for entry in self.entries { + if entry.start < end && entry.end > start { + blocks.push(self.blocks[entry.block_index]) + } + } + blocks +} + +///| +pub fn AlignMafIndex::search_ranges( + self : AlignMafIndex, + starts : Array[Int], + ends : Array[Int], +) -> Array[BigMafBlock] raise AlignMafError { + if starts.length() == 0 || starts.length() != ends.length() { + align_maf_fail("MAF range arrays must be non-empty and have equal lengths") + } + let selected = Array::make(self.blocks.length(), false) + for index in 0.. starts[index] { + selected[entry.block_index] = true + } + } + } + let blocks : Array[BigMafBlock] = [] + for index in 0.. Int { + self.entries.length() +} + +///| +fn align_maf_repeat(value : String, count : Int) -> String { + let output = StringBuilder::new(size_hint=count) + for _ in 0.. Array[String] { + let sources : Array[String] = [index.reference] + for block in index.blocks { + for component in block.components { + if component.source != index.reference && + !sources.contains(component.source) { + sources.push(component.source) + } + } + } + sources +} + +///| +fn align_maf_append_missing(texts : Array[String], count : Int) -> Unit { + if count <= 0 { + return + } + texts[0] = texts[0] + align_maf_repeat("N", count) + for index in 1.. AlignMafSplicedAlignment raise AlignMafError { + if starts.length() == 0 || starts.length() != ends.length() { + align_maf_fail("MAF exon arrays must be non-empty and have equal lengths") + } + if strand != "+" && strand != "-" { + align_maf_fail("MAF spliced strand must be '+' or '-'") + } + if self.reference_strand != "+" { + align_maf_fail( + "MAF splicing currently requires a plus-strand reference component", + ) + } + for index in 0.. 0 && starts[index] < ends[index - 1] { + align_maf_fail("MAF exon intervals must be sorted and non-overlapping") + } + } + let sources = align_maf_splice_sources(self) + let texts = Array::make(sources.length(), "") + for exon_index in 0..= exon_end { + continue + } + let overlap_start = if entry.start > exon_start { + entry.start + } else { + exon_start + } + let overlap_end = if entry.end < exon_end { entry.end } else { exon_end } + if overlap_start < current { + align_maf_fail( + "Overlapping MAF reference blocks make splicing ambiguous", + ) + } + align_maf_append_missing(texts, overlap_start - current) + let block = self.blocks[entry.block_index] + let reference = match block.component(self.reference) { + Some(component) => component + None => { + align_maf_fail("Indexed MAF block lost its reference component") + block.components[0] + } + } + let first_column = match reference.source_to_column(overlap_start) { + Some(column) => column + None => { + align_maf_fail("MAF reference interval does not map to a column") + 0 + } + } + let last_column = match reference.source_to_column(overlap_end - 1) { + Some(column) => column + None => { + align_maf_fail("MAF reference interval does not map to a column") + 0 + } + } + let column_start = if first_column < last_column { + first_column + } else { + last_column + } + let column_end = if first_column > last_column { + first_column + 1 + } else { + last_column + 1 + } + let width = column_end - column_start + for source_index in 0.. + texts[source_index] = texts[source_index] + + component.text[column_start:column_end].to_owned() + None => + texts[source_index] = texts[source_index] + + align_maf_repeat("-", width) + } + } + current = overlap_end + } + align_maf_append_missing(texts, exon_end - current) + } + if strand == "-" { + for index in 0.. String? { + for sequence in self.sequences { + if sequence.source == source { + return Some(sequence.text) + } + } + None +} + +///| +pub fn AlignMafSplicedAlignment::ungapped_length( + self : AlignMafSplicedAlignment, + source : String, +) -> Int? { + match self.sequence(source) { + Some(text) => { + let mut count = 0 + for index in 0.. None + } +} + +///| +pub fn align_maf_example_text() -> String { + "track name=\"Multiz demo\" description=\"MAF coordinate example\" " + + "mafDot=on visibility=pack speciesOrder=\"hg38 mm39 rn7\"\n" + + "##maf version=1 scoring=blastz program=multiz\n" + + "# generated fixture\n\n" + + "a score=23262.0 pass=1\n" + + "s hg38.chr7 100 8 + 1000 AC-TGCAAT\n" + + "i hg38.chr7 C 0 I 2\n" + + "q hg38.chr7 98-765432\n" + + "s mm39.chr5 50 8 - 500 ACCTG-AAT\n" + + "e rn7.chr1 80 6 + 700 I\n\n" + + "a score=125.5\n" + + "s hg38.chr7 110 4 + 1000 GGTT\n" + + "s mm39.chr5 70 3 - 500 G-TT\n\n" + + "a score=80\n" + + "s hg38.chr7 120 4 + 1000 AAAA\n" + + "s rn7.chr1 90 4 + 700 AATA\n" +} diff --git a/src/align_mauve.mbt b/src/align_mauve.mbt new file mode 100644 index 00000000..7f1d2eb7 --- /dev/null +++ b/src/align_mauve.mbt @@ -0,0 +1,2107 @@ +// Biopython-compatible Bio.Align.mauve support. +// +// This module is intentionally separate from mauve.mbt. The older module +// provides MAF-like rearrangement summaries, while this file models genuine +// Mauve/ProgressiveMauve XMFA documents, aligned rows, and coordinate paths. + +///| +/// Error raised for malformed XMFA data or invalid coordinate operations. +pub suberror AlignMauveError { + AlignMauveError(String) +} + +///| +/// One non-sequence file-level XMFA metadata entry. +pub struct AlignMauveMetadataEntry { + key : String + value : String +} derive(Eq, Debug) + +///| +/// One sequence declared by `#SequenceN*` header fields. +/// +/// `number` and `entry` retain XMFA's one-based values. `entry == None` +/// denotes separate-file input without a `SequenceNEntry` line. +pub struct AlignMauveSource { + number : Int + file : String + entry : Int? + format : String + identifier : String +} derive(Eq, Debug) + +///| +/// One aligned row in a locally collinear XMFA block. +/// +/// `start` and `end` are zero-based forward-axis half-open coordinates. +/// Reverse rows have a decreasing `coordinates` path from `end` to `start`. +/// `source_sequence` is always stored in the forward source orientation. +pub struct AlignMauveRow { + source_number : Int + identifier : String + start : Int + end : Int + strand : String + description : String + aligned_sequence : String + source_sequence : String + coordinates : Array[Int] +} derive(Eq, Debug) + +///| +/// One Mauve locally collinear block. +pub struct AlignMauveBlock { + rows : Array[AlignMauveRow] + width : Int +} derive(Eq, Debug) + +///| +/// A complete XMFA document. +pub struct AlignMauveDocument { + format_version : String + sources : Array[AlignMauveSource] + metadata : Array[AlignMauveMetadataEntry] + blocks : Array[AlignMauveBlock] +} derive(Eq, Debug) + +///| +/// Pairwise statistics summed over all unordered row pairs in one or more +/// blocks. Columns where both rows contain gaps are reported separately and +/// do not contribute to gap counts. +pub struct AlignMauveCounts { + pair_count : Int + columns : Int + aligned : Int + identities : Int + mismatches : Int + gaps : Int + gap_opens : Int + double_gap_columns : Int +} derive(Eq, Debug) + +///| +/// One interval-index entry for an aligned source segment. +pub struct AlignMauveIndexEntry { + identifier : String + start : Int + end : Int + block_index : Int + row_index : Int + strand : String +} derive(Eq, Debug) + +///| +/// In-memory interval index over all non-empty XMFA rows. +pub struct AlignMauveIndex { + identifiers : Array[String] + entries : Array[AlignMauveIndexEntry] +} derive(Eq, Debug) + +///| +/// One exact residue-to-residue coordinate projection. +pub struct AlignMauvePositionMapping { + block_index : Int + column : Int + source_position : Int + target_position : Int +} derive(Eq, Debug) + +///| +/// One contiguous aligned part of a projected source interval. +/// +/// Source and target intervals are zero-based forward-axis half-open ranges. +/// `column_start` and `column_end` identify the shared alignment columns. +pub struct AlignMauveRangeMapping { + block_index : Int + column_start : Int + column_end : Int + source_start : Int + source_end : Int + target_start : Int + target_end : Int + source_strand : String + target_strand : String +} derive(Eq, Debug) + +///| +/// Aggregate document summary. +pub struct AlignMauveSummary { + sources : Int + blocks : Int + rows : Int + alignment_columns : Int + aligned_pairs : Int + identities : Int + mismatches : Int + gaps : Int +} derive(Eq, Debug) + +///| +priv struct AlignMauvePendingRow { + source_number : Int + start : Int + end : Int + strand : String + description : String + mut aligned_sequence : String +} + +///| +fn align_mauve_fail(message : String) -> Unit raise AlignMauveError { + raise AlignMauveError(message) +} + +///| +fn align_mauve_copy_sources( + values : Array[AlignMauveSource], +) -> Array[AlignMauveSource] { + let copied : Array[AlignMauveSource] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_mauve_clone_row(value : AlignMauveRow) -> AlignMauveRow { + let coordinates : Array[Int] = [] + for coordinate in value.coordinates { + coordinates.push(coordinate) + } + AlignMauveRow::{ + source_number: value.source_number, + identifier: value.identifier, + start: value.start, + end: value.end, + strand: value.strand, + description: value.description, + aligned_sequence: value.aligned_sequence, + source_sequence: value.source_sequence, + coordinates, + } +} + +///| +fn align_mauve_copy_rows(values : Array[AlignMauveRow]) -> Array[AlignMauveRow] { + let copied : Array[AlignMauveRow] = [] + for value in values { + copied.push(align_mauve_clone_row(value)) + } + copied +} + +///| +fn align_mauve_clone_block(value : AlignMauveBlock) -> AlignMauveBlock { + AlignMauveBlock::{ + rows: align_mauve_copy_rows(value.rows), + width: value.width, + } +} + +///| +fn align_mauve_copy_blocks( + values : Array[AlignMauveBlock], +) -> Array[AlignMauveBlock] { + let copied : Array[AlignMauveBlock] = [] + for value in values { + copied.push(align_mauve_clone_block(value)) + } + copied +} + +///| +fn align_mauve_copy_metadata( + values : Array[AlignMauveMetadataEntry], +) -> Array[AlignMauveMetadataEntry] { + let copied : Array[AlignMauveMetadataEntry] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_mauve_copy_ints(values : Array[Int]) -> Array[Int] { + let copied : Array[Int] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_mauve_copy_index_entries( + values : Array[AlignMauveIndexEntry], +) -> Array[AlignMauveIndexEntry] { + let copied : Array[AlignMauveIndexEntry] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn align_mauve_starts_with(value : String, prefix : String) -> Bool { + if prefix.length() > value.length() { + return false + } + for index in 0.. Bool { + if suffix.length() > value.length() { + return false + } + let offset = value.length() - suffix.length() + for index in 0.. Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} + +///| +fn align_mauve_trim_line(value : String) -> String { + let mut start = 0 + let mut end = value.length() + while start < end && align_mauve_is_space(value.unsafe_get(start).to_int()) { + start = start + 1 + } + while end > start && align_mauve_is_space(value.unsafe_get(end - 1).to_int()) { + end = end - 1 + } + value[start:end].to_owned() +} + +///| +fn align_mauve_validate_plain_text( + value : String, + label : String, + allow_empty : Bool, +) -> Unit raise AlignMauveError { + if !allow_empty && value.length() == 0 { + align_mauve_fail(label + " must not be empty") + } + for index in 0.. Int raise AlignMauveError { + if text.length() == 0 { + align_mauve_fail(label + " must not be empty") + } + let mut value = 0 + for index in 0.. '9'.to_int() { + align_mauve_fail(label + " must be a decimal integer") + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + align_mauve_fail(label + " exceeds the supported integer range") + } + value = value * 10 + digit + } + value +} + +///| +fn align_mauve_index_of_char(value : String, wanted : Int, start : Int) -> Int { + let mut index = start + while index < value.length() { + if value.unsafe_get(index).to_int() == wanted { + return index + } + index = index + 1 + } + -1 +} + +///| +fn align_mauve_header_parts( + line : String, + line_number : Int, +) -> (String, String) raise AlignMauveError { + let body = align_mauve_trim_line(line[1:line.length()].to_owned()) + let mut split = 0 + while split < body.length() && + !align_mauve_is_space(body.unsafe_get(split).to_int()) { + split = split + 1 + } + if split == 0 || split == body.length() { + align_mauve_fail( + "XMFA header line " + + line_number.to_string() + + " must contain a key and value", + ) + } + let key = body[0:split].to_owned() + let value = align_mauve_trim_line(body[split:body.length()].to_owned()) + if value.length() == 0 { + align_mauve_fail( + "XMFA header line " + line_number.to_string() + " has an empty value", + ) + } + (key, value) +} + +///| +fn align_mauve_find_metadata( + values : Array[AlignMauveMetadataEntry], + key : String, +) -> String? { + for value in values { + if value.key == key { + return Some(value.value) + } + } + None +} + +///| +fn align_mauve_find_source_property( + values : Array[AlignMauveMetadataEntry], + number : Int, + suffix : String, +) -> String? { + align_mauve_find_metadata(values, "Sequence" + number.to_string() + suffix) +} + +///| +fn align_mauve_parse_sequence_key( + key : String, +) -> (Int, String)? raise AlignMauveError { + if !align_mauve_starts_with(key, "Sequence") { + return None + } + let suffix = if align_mauve_ends_with(key, "File") { + "File" + } else if align_mauve_ends_with(key, "Entry") { + "Entry" + } else if align_mauve_ends_with(key, "Format") { + "Format" + } else { + align_mauve_fail("Unexpected XMFA sequence header keyword '" + key + "'") + "" + } + let start = "Sequence".length() + let end = key.length() - suffix.length() + if end <= start { + align_mauve_fail("Malformed XMFA sequence header keyword '" + key + "'") + } + let number = align_mauve_parse_positive_int( + key[start:end].to_owned(), + "XMFA sequence number", + ) + if number <= 0 { + align_mauve_fail("XMFA sequence numbers are one-based and must be positive") + } + Some((number, suffix)) +} + +///| +fn align_mauve_find_source( + sources : Array[AlignMauveSource], + number : Int, +) -> AlignMauveSource? { + for source in sources { + if source.number == number { + return Some(source) + } + } + None +} + +///| +fn align_mauve_find_source_by_identifier( + sources : Array[AlignMauveSource], + identifier : String, +) -> AlignMauveSource? { + for source in sources { + if source.identifier == identifier { + return Some(source) + } + } + None +} + +///| +fn align_mauve_validate_strand(strand : String) -> Unit raise AlignMauveError { + if strand != "+" && strand != "-" { + align_mauve_fail("XMFA strand must be '+' or '-'") + } +} + +///| +fn align_mauve_complement(base : Char) -> Char raise AlignMauveError { + match base { + 'A' => 'T' + 'C' => 'G' + 'G' => 'C' + 'T' => 'A' + 'U' => 'A' + 'R' => 'Y' + 'Y' => 'R' + 'S' => 'S' + 'W' => 'W' + 'K' => 'M' + 'M' => 'K' + 'B' => 'V' + 'D' => 'H' + 'H' => 'D' + 'V' => 'B' + 'N' => 'N' + 'X' => 'X' + 'a' => 't' + 'c' => 'g' + 'g' => 'c' + 't' => 'a' + 'u' => 'a' + 'r' => 'y' + 'y' => 'r' + 's' => 's' + 'w' => 'w' + 'k' => 'm' + 'm' => 'k' + 'b' => 'v' + 'd' => 'h' + 'h' => 'd' + 'v' => 'b' + 'n' => 'n' + 'x' => 'x' + _ => { + align_mauve_fail( + "XMFA sequence contains non-IUPAC residue '" + base.to_string() + "'", + ) + 'N' + } + } +} + +///| +fn align_mauve_validate_aligned_sequence( + sequence : String, +) -> Unit raise AlignMauveError { + if sequence.length() == 0 { + align_mauve_fail("XMFA aligned sequence must not be empty") + } + for index in 0.. String { + let output = StringBuilder::new(size_hint=sequence.length()) + for index in 0.. String raise AlignMauveError { + let output = StringBuilder::new(size_hint=sequence.length()) + let mut index = sequence.length() - 1 + while index >= 0 { + output.write_char( + align_mauve_complement(sequence.unsafe_get(index).unsafe_to_char()), + ) + index = index - 1 + } + output.to_string() +} + +///| +fn align_mauve_make_coordinates( + aligned_sequence : String, + start : Int, + end : Int, + strand : String, +) -> Array[Int] raise AlignMauveError { + let coordinates : Array[Int] = [] + let mut current = if strand == "+" { start } else { end } + coordinates.push(current) + for column in 0.. AlignMauvePendingRow raise AlignMauveError { + let body = align_mauve_trim_line(line[1:line.length()].to_owned()) + let mut first_end = 0 + while first_end < body.length() && + !align_mauve_is_space(body.unsafe_get(first_end).to_int()) { + first_end = first_end + 1 + } + if first_end == 0 || first_end == body.length() { + align_mauve_fail( + "Malformed XMFA row description at line " + line_number.to_string(), + ) + } + let locus = body[0:first_end].to_owned() + let mut strand_start = first_end + while strand_start < body.length() && + align_mauve_is_space(body.unsafe_get(strand_start).to_int()) { + strand_start = strand_start + 1 + } + let mut strand_end = strand_start + while strand_end < body.length() && + !align_mauve_is_space(body.unsafe_get(strand_end).to_int()) { + strand_end = strand_end + 1 + } + if strand_start == strand_end || strand_end == body.length() { + align_mauve_fail( + "XMFA row description at line " + + line_number.to_string() + + " must include strand and description", + ) + } + let strand = body[strand_start:strand_end].to_owned() + align_mauve_validate_strand(strand) + let description = align_mauve_trim_line( + body[strand_end:body.length()].to_owned(), + ) + align_mauve_validate_plain_text(description, "XMFA row description", false) + let colon = align_mauve_index_of_char(locus, ':'.to_int(), 0) + if colon <= 0 || + colon + 1 >= locus.length() || + align_mauve_index_of_char(locus, ':'.to_int(), colon + 1) >= 0 { + align_mauve_fail( + "Malformed XMFA sequence locus at line " + line_number.to_string(), + ) + } + let dash = align_mauve_index_of_char(locus, '-'.to_int(), colon + 1) + if dash <= colon + 1 || + dash + 1 >= locus.length() || + align_mauve_index_of_char(locus, '-'.to_int(), dash + 1) >= 0 { + align_mauve_fail( + "Malformed XMFA coordinate range at line " + line_number.to_string(), + ) + } + let source_number = align_mauve_parse_positive_int( + locus[0:colon].to_owned(), + "XMFA row sequence number", + ) + if source_number <= 0 { + align_mauve_fail("XMFA row sequence number must be positive") + } + let file_start = align_mauve_parse_positive_int( + locus[colon + 1:dash].to_owned(), + "XMFA row start", + ) + let file_end = align_mauve_parse_positive_int( + locus[dash + 1:locus.length()].to_owned(), + "XMFA row end", + ) + let (start, end) = if file_start == 0 { + if file_end != 0 { + align_mauve_fail("XMFA start coordinate zero is only valid for 0-0 rows") + } + (0, 0) + } else { + if file_end < file_start { + align_mauve_fail("XMFA row end must not precede its start") + } + (file_start - 1, file_end) + } + AlignMauvePendingRow::{ + source_number, + start, + end, + strand, + description, + aligned_sequence: "", + } +} + +///| +fn align_mauve_finalize_block( + pending : Array[AlignMauvePendingRow], + sources : Array[AlignMauveSource], +) -> AlignMauveBlock raise AlignMauveError { + if pending.length() == 0 { + align_mauve_fail("XMFA alignment block must contain at least one row") + } + let rows : Array[AlignMauveRow] = [] + let mut width = -1 + let mut has_residue = false + for item in pending { + for row in rows { + if row.source_number == item.source_number { + align_mauve_fail( + "XMFA alignment block contains a duplicated sequence number", + ) + } + } + let source = match align_mauve_find_source(sources, item.source_number) { + Some(value) => value + None => { + align_mauve_fail( + "XMFA row refers to undeclared sequence " + + item.source_number.to_string(), + ) + sources[0] + } + } + let row = AlignMauveRow::create( + source, + item.start, + item.end, + item.strand, + item.description, + item.aligned_sequence, + ) + if width < 0 { + width = row.aligned_sequence.length() + } else if row.aligned_sequence.length() != width { + align_mauve_fail("XMFA rows in one block must have equal widths") + } + if row.end > row.start { + has_residue = true + } + rows.push(row) + } + if width <= 0 { + align_mauve_fail("XMFA alignment block width must be positive") + } + if !has_residue { + align_mauve_fail("XMFA alignment block must contain at least one residue") + } + AlignMauveBlock::{ rows, width } +} + +///| +/// Construct one validated metadata entry. +pub fn AlignMauveMetadataEntry::create( + key : String, + value : String, +) -> AlignMauveMetadataEntry raise AlignMauveError { + align_mauve_validate_plain_text(key, "XMFA metadata key", false) + align_mauve_validate_plain_text(value, "XMFA metadata value", false) + for index in 0.. AlignMauveSource raise AlignMauveError { + if number <= 0 { + align_mauve_fail("XMFA source number must be positive") + } + align_mauve_validate_plain_text(file, "XMFA source file", false) + align_mauve_validate_plain_text(format, "XMFA source format", false) + if entry < 0 { + align_mauve_fail("XMFA source entry must be non-negative") + } + let resolved_identifier = if identifier.length() > 0 { + identifier + } else if entry > 0 { + (entry - 1).to_string() + } else { + file + } + align_mauve_validate_plain_text( + resolved_identifier, "XMFA source identifier", false, + ) + AlignMauveSource::{ + number, + file, + entry: if entry == 0 { + None + } else { + Some(entry) + }, + format, + identifier: resolved_identifier, + } +} + +///| +/// Construct one validated aligned XMFA row. +pub fn AlignMauveRow::create( + source : AlignMauveSource, + start : Int, + end : Int, + strand : String, + description : String, + aligned_sequence : String, +) -> AlignMauveRow raise AlignMauveError { + if start < 0 || end < start { + align_mauve_fail("XMFA row interval must satisfy 0 <= start <= end") + } + align_mauve_validate_strand(strand) + align_mauve_validate_plain_text(description, "XMFA row description", false) + align_mauve_validate_aligned_sequence(aligned_sequence) + let displayed_sequence = align_mauve_remove_gaps(aligned_sequence) + if displayed_sequence.length() != end - start { + align_mauve_fail( + "XMFA row coordinate span must equal its ungapped sequence length", + ) + } + if start == end && displayed_sequence.length() != 0 { + align_mauve_fail("XMFA 0-0 row must contain gaps only") + } + let source_sequence = if strand == "+" { + displayed_sequence + } else { + align_mauve_reverse_complement(displayed_sequence) + } + let coordinates = align_mauve_make_coordinates( + aligned_sequence, start, end, strand, + ) + AlignMauveRow::{ + source_number: source.number, + identifier: source.identifier, + start, + end, + strand, + description, + aligned_sequence, + source_sequence, + coordinates, + } +} + +///| +/// Construct one validated XMFA block. +pub fn AlignMauveBlock::create( + rows : Array[AlignMauveRow], +) -> AlignMauveBlock raise AlignMauveError { + if rows.length() == 0 { + align_mauve_fail("XMFA block must contain at least one row") + } + let copied = align_mauve_copy_rows(rows) + let width = copied[0].aligned_sequence.length() + let mut has_residue = false + for index in 0.. row.start { + has_residue = true + } + for other in 0.. AlignMauveDocument raise AlignMauveError { + if text.length() == 0 { + align_mauve_fail("Empty XMFA input") + } + let raw_lines = text.split("\n").to_array() + let lines : Array[String] = [] + for raw in raw_lines { + let owned = raw.to_owned() + let line = if owned.length() > 0 && + owned.unsafe_get(owned.length() - 1).to_int() == '\r'.to_int() { + owned[0:owned.length() - 1].to_owned() + } else { + owned + } + lines.push(line) + } + let all_metadata : Array[AlignMauveMetadataEntry] = [] + let mut index = 0 + while index < lines.length() { + let trimmed = align_mauve_trim_line(lines[index]) + if trimmed.length() == 0 { + index = index + 1 + continue + } + if !align_mauve_starts_with(trimmed, "#") { + break + } + let (key, value) = align_mauve_header_parts(trimmed, index + 1) + if align_mauve_find_metadata(all_metadata, key) is Some(_) { + align_mauve_fail("Duplicated XMFA header keyword '" + key + "'") + } + all_metadata.push(AlignMauveMetadataEntry::create(key, value)) + index = index + 1 + } + if all_metadata.length() == 0 { + align_mauve_fail("XMFA input must begin with metadata header lines") + } + let format_version = match + align_mauve_find_metadata(all_metadata, "FormatVersion") { + Some(value) => value + None => { + align_mauve_fail("XMFA header is missing FormatVersion") + "" + } + } + let mut maximum_source = 0 + for entry in all_metadata { + match align_mauve_parse_sequence_key(entry.key) { + Some((number, _)) => + if number > maximum_source { + maximum_source = number + } + None => () + } + } + if maximum_source == 0 { + align_mauve_fail("XMFA header declares no source sequences") + } + let source_files : Array[String] = [] + let source_formats : Array[String] = [] + let source_entries : Array[Int] = [] + for number = 1; number <= maximum_source; number = number + 1 { + let file = match + align_mauve_find_source_property(all_metadata, number, "File") { + Some(value) => value + None => { + align_mauve_fail( + "XMFA header is missing Sequence" + number.to_string() + "File", + ) + "" + } + } + let format = match + align_mauve_find_source_property(all_metadata, number, "Format") { + Some(value) => value + None => { + align_mauve_fail( + "XMFA header is missing Sequence" + number.to_string() + "Format", + ) + "" + } + } + let source_entry = match + align_mauve_find_source_property(all_metadata, number, "Entry") { + Some(value) => { + let parsed = align_mauve_parse_positive_int(value, "XMFA source entry") + if parsed <= 0 { + align_mauve_fail("XMFA source entry must be positive") + } + parsed + } + None => 0 + } + source_files.push(file) + source_formats.push(format) + source_entries.push(source_entry) + } + let mut combined = true + for source_index in 1..") { + pending.push(align_mauve_parse_description(trimmed, line_number)) + continue + } + if pending.length() == 0 { + align_mauve_fail( + "XMFA sequence data at line " + + line_number.to_string() + + " appears before a row description", + ) + } + for char_index in 0.. 0 { + align_mauve_fail("XMFA final alignment block is missing its '=' terminator") + } + AlignMauveDocument::{ format_version, sources, metadata, blocks } +} + +///| +/// Serialize an XMFA document in canonical Mauve form. +/// +/// `line_width == 0` writes one physical line per aligned row, matching +/// Biopython's writer. Positive widths wrap sequence rows deterministically. +pub fn align_mauve_write( + document : AlignMauveDocument, + line_width? : Int = 0, +) -> String raise AlignMauveError { + if line_width < 0 { + align_mauve_fail("XMFA line width must be non-negative") + } + if document.sources.length() == 0 { + align_mauve_fail("XMFA document must declare at least one source") + } + let output = StringBuilder::new() + output.write_string("#FormatVersion ") + output.write_string(document.format_version) + output.write_char('\n') + for source in document.sources { + output.write_string("#Sequence") + output.write_string(source.number.to_string()) + output.write_string("File\t") + output.write_string(source.file) + output.write_char('\n') + match source.entry { + Some(entry) => { + output.write_string("#Sequence") + output.write_string(source.number.to_string()) + output.write_string("Entry\t") + output.write_string(entry.to_string()) + output.write_char('\n') + } + None => () + } + output.write_string("#Sequence") + output.write_string(source.number.to_string()) + output.write_string("Format\t") + output.write_string(source.format) + output.write_char('\n') + } + for entry in document.metadata { + if entry.key != "FormatVersion" { + output.write_char('#') + output.write_string(entry.key) + output.write_char('\t') + output.write_string(entry.value) + output.write_char('\n') + } + } + for block in document.blocks { + for row in block.rows { + output.write_string("> ") + output.write_string(row.source_number.to_string()) + output.write_char(':') + if row.start == 0 && row.end == 0 { + output.write_string("0-0") + } else { + output.write_string((row.start + 1).to_string()) + output.write_char('-') + output.write_string(row.end.to_string()) + } + output.write_char(' ') + output.write_string(row.strand) + output.write_char(' ') + output.write_string(row.description) + output.write_char('\n') + if line_width == 0 { + output.write_string(row.aligned_sequence) + output.write_char('\n') + } else { + let mut offset = 0 + while offset < row.aligned_sequence.length() { + let end = if offset + line_width < row.aligned_sequence.length() { + offset + line_width + } else { + row.aligned_sequence.length() + } + output.write_string(row.aligned_sequence[offset:end].to_owned()) + output.write_char('\n') + offset = end + } + } + } + output.write_string("=\n") + } + output.to_string() +} + +///| +pub fn AlignMauveMetadataEntry::key(self : AlignMauveMetadataEntry) -> String { + self.key +} + +///| +pub fn AlignMauveMetadataEntry::value(self : AlignMauveMetadataEntry) -> String { + self.value +} + +///| +pub fn AlignMauveSource::number(self : AlignMauveSource) -> Int { + self.number +} + +///| +pub fn AlignMauveSource::file(self : AlignMauveSource) -> String { + self.file +} + +///| +pub fn AlignMauveSource::entry(self : AlignMauveSource) -> Int? { + self.entry +} + +///| +pub fn AlignMauveSource::format(self : AlignMauveSource) -> String { + self.format +} + +///| +pub fn AlignMauveSource::identifier(self : AlignMauveSource) -> String { + self.identifier +} + +///| +pub fn AlignMauveRow::source_number(self : AlignMauveRow) -> Int { + self.source_number +} + +///| +pub fn AlignMauveRow::identifier(self : AlignMauveRow) -> String { + self.identifier +} + +///| +pub fn AlignMauveRow::start(self : AlignMauveRow) -> Int { + self.start +} + +///| +pub fn AlignMauveRow::end(self : AlignMauveRow) -> Int { + self.end +} + +///| +pub fn AlignMauveRow::strand(self : AlignMauveRow) -> String { + self.strand +} + +///| +pub fn AlignMauveRow::description(self : AlignMauveRow) -> String { + self.description +} + +///| +pub fn AlignMauveRow::aligned_sequence(self : AlignMauveRow) -> String { + self.aligned_sequence +} + +///| +pub fn AlignMauveRow::source_sequence(self : AlignMauveRow) -> String { + self.source_sequence +} + +///| +pub fn AlignMauveRow::coordinates(self : AlignMauveRow) -> Array[Int] { + align_mauve_copy_ints(self.coordinates) +} + +///| +/// Return the source boundary at one alignment-column boundary. +pub fn AlignMauveRow::coordinate_at_boundary( + self : AlignMauveRow, + column : Int, +) -> Int raise AlignMauveError { + if column < 0 || column >= self.coordinates.length() { + align_mauve_fail("XMFA alignment-column boundary is out of range") + } + self.coordinates[column] +} + +///| +/// Return the zero-based forward source coordinate represented by one residue +/// column, or `None` when this row has a gap in that column. +pub fn AlignMauveRow::residue_coordinate( + self : AlignMauveRow, + column : Int, +) -> Int? raise AlignMauveError { + if column < 0 || column >= self.aligned_sequence.length() { + align_mauve_fail("XMFA alignment column is out of range") + } + if self.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + if self.strand == "+" { + Some(self.coordinates[column]) + } else { + Some(self.coordinates[column] - 1) + } +} + +///| +/// Locate one forward-axis source residue in this aligned row. +pub fn AlignMauveRow::column_for_coordinate( + self : AlignMauveRow, + coordinate : Int, +) -> Int? raise AlignMauveError { + if coordinate < 0 { + align_mauve_fail("XMFA source coordinate must be non-negative") + } + for column in 0.. if value == coordinate { return Some(column) } + None => () + } + } + None +} + +///| +pub fn AlignMauveBlock::rows(self : AlignMauveBlock) -> Array[AlignMauveRow] { + align_mauve_copy_rows(self.rows) +} + +///| +pub fn AlignMauveBlock::row_count(self : AlignMauveBlock) -> Int { + self.rows.length() +} + +///| +pub fn AlignMauveBlock::width(self : AlignMauveBlock) -> Int { + self.width +} + +///| +pub fn AlignMauveBlock::row_by_identifier( + self : AlignMauveBlock, + identifier : String, +) -> AlignMauveRow? { + for row in self.rows { + if row.identifier == identifier { + return Some(row) + } + } + None +} + +///| +pub fn AlignMauveBlock::row_by_source_number( + self : AlignMauveBlock, + source_number : Int, +) -> AlignMauveRow? { + for row in self.rows { + if row.source_number == source_number { + return Some(row) + } + } + None +} + +///| +/// Return Biopython-style compact coordinate rows. +/// +/// A boundary is retained whenever any row changes between residue movement +/// and gap movement, plus the first and final boundaries. +pub fn AlignMauveBlock::compact_coordinates( + self : AlignMauveBlock, +) -> Array[Array[Int]] { + let boundaries : Array[Int] = [0] + for column = 1; column < self.width; column = column + 1 { + let mut changed = false + for row in self.rows { + let previous = row.coordinates[column] - row.coordinates[column - 1] + let next = row.coordinates[column + 1] - row.coordinates[column] + if previous != next { + changed = true + } + } + if changed { + boundaries.push(column) + } + } + boundaries.push(self.width) + let compact : Array[Array[Int]] = [] + for row in self.rows { + let coordinates : Array[Int] = [] + for boundary in boundaries { + coordinates.push(row.coordinates[boundary]) + } + compact.push(coordinates) + } + compact +} + +///| +/// Project one source residue to a target residue through a shared block +/// column. A gap in either row yields `None`. +pub fn AlignMauveBlock::map_position( + self : AlignMauveBlock, + source_identifier : String, + target_identifier : String, + source_position : Int, +) -> Int? raise AlignMauveError { + if source_position < 0 { + align_mauve_fail("XMFA source position must be non-negative") + } + let source = match self.row_by_identifier(source_identifier) { + Some(row) => row + None => return None + } + let target = match self.row_by_identifier(target_identifier) { + Some(row) => row + None => return None + } + let column = match source.column_for_coordinate(source_position) { + Some(value) => value + None => return None + } + target.residue_coordinate(column) +} + +///| +/// Calculate pairwise alignment counts over all unordered row pairs. +pub fn AlignMauveBlock::counts(self : AlignMauveBlock) -> AlignMauveCounts { + let mut pair_count = 0 + let mut columns = 0 + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gaps = 0 + let mut gap_opens = 0 + let mut double_gap_columns = 0 + for left = 0; left < self.rows.length(); left = left + 1 { + for right = left + 1; right < self.rows.length(); right = right + 1 { + pair_count = pair_count + 1 + columns = columns + self.width + let mut active_gap = 0 + for column in 0.. String { + self.format_version +} + +///| +pub fn AlignMauveDocument::sources( + self : AlignMauveDocument, +) -> Array[AlignMauveSource] { + align_mauve_copy_sources(self.sources) +} + +///| +pub fn AlignMauveDocument::blocks( + self : AlignMauveDocument, +) -> Array[AlignMauveBlock] { + align_mauve_copy_blocks(self.blocks) +} + +///| +pub fn AlignMauveDocument::metadata_entries( + self : AlignMauveDocument, +) -> Array[AlignMauveMetadataEntry] { + align_mauve_copy_metadata(self.metadata) +} + +///| +pub fn AlignMauveDocument::metadata_value( + self : AlignMauveDocument, + key : String, +) -> String? { + align_mauve_find_metadata(self.metadata, key) +} + +///| +pub fn AlignMauveDocument::source_count(self : AlignMauveDocument) -> Int { + self.sources.length() +} + +///| +pub fn AlignMauveDocument::block_count(self : AlignMauveDocument) -> Int { + self.blocks.length() +} + +///| +pub fn AlignMauveDocument::identifiers( + self : AlignMauveDocument, +) -> Array[String] { + let identifiers : Array[String] = [] + for source in self.sources { + identifiers.push(source.identifier) + } + identifiers +} + +///| +pub fn AlignMauveDocument::source( + self : AlignMauveDocument, + identifier : String, +) -> AlignMauveSource? { + align_mauve_find_source_by_identifier(self.sources, identifier) +} + +///| +/// Sum pairwise counts over all XMFA blocks. +pub fn AlignMauveDocument::counts( + self : AlignMauveDocument, +) -> AlignMauveCounts { + let mut pair_count = 0 + let mut columns = 0 + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gaps = 0 + let mut gap_opens = 0 + let mut double_gap_columns = 0 + for block in self.blocks { + let counts = block.counts() + pair_count = pair_count + counts.pair_count + columns = columns + counts.columns + aligned = aligned + counts.aligned + identities = identities + counts.identities + mismatches = mismatches + counts.mismatches + gaps = gaps + counts.gaps + gap_opens = gap_opens + counts.gap_opens + double_gap_columns = double_gap_columns + counts.double_gap_columns + } + AlignMauveCounts::{ + pair_count, + columns, + aligned, + identities, + mismatches, + gaps, + gap_opens, + double_gap_columns, + } +} + +///| +pub fn AlignMauveCounts::pair_count(self : AlignMauveCounts) -> Int { + self.pair_count +} + +///| +pub fn AlignMauveCounts::columns(self : AlignMauveCounts) -> Int { + self.columns +} + +///| +pub fn AlignMauveCounts::aligned(self : AlignMauveCounts) -> Int { + self.aligned +} + +///| +pub fn AlignMauveCounts::identities(self : AlignMauveCounts) -> Int { + self.identities +} + +///| +pub fn AlignMauveCounts::mismatches(self : AlignMauveCounts) -> Int { + self.mismatches +} + +///| +pub fn AlignMauveCounts::gaps(self : AlignMauveCounts) -> Int { + self.gaps +} + +///| +pub fn AlignMauveCounts::gap_opens(self : AlignMauveCounts) -> Int { + self.gap_opens +} + +///| +pub fn AlignMauveCounts::double_gap_columns(self : AlignMauveCounts) -> Int { + self.double_gap_columns +} + +///| +/// Build an in-memory interval index over all non-empty rows. +pub fn align_mauve_index(document : AlignMauveDocument) -> AlignMauveIndex { + let entries : Array[AlignMauveIndexEntry] = [] + for block_index in 0.. row.start { + entries.push(AlignMauveIndexEntry::{ + identifier: row.identifier, + start: row.start, + end: row.end, + block_index, + row_index, + strand: row.strand, + }) + } + } + } + AlignMauveIndex::{ identifiers: document.identifiers(), entries } +} + +///| +pub fn AlignMauveIndex::entries( + self : AlignMauveIndex, +) -> Array[AlignMauveIndexEntry] { + align_mauve_copy_index_entries(self.entries) +} + +///| +/// Find rows overlapping one zero-based half-open source interval. +pub fn AlignMauveIndex::query( + self : AlignMauveIndex, + identifier : String, + start : Int, + end : Int, +) -> Array[AlignMauveIndexEntry] raise AlignMauveError { + if start < 0 || end <= start { + align_mauve_fail("XMFA index query must satisfy 0 <= start < end") + } + if !self.identifiers.contains(identifier) { + align_mauve_fail("Unknown XMFA source identifier '" + identifier + "'") + } + let matches : Array[AlignMauveIndexEntry] = [] + for entry in self.entries { + if entry.identifier == identifier && entry.start < end && entry.end > start { + matches.push(entry) + } + } + matches +} + +///| +pub fn AlignMauveIndexEntry::identifier(self : AlignMauveIndexEntry) -> String { + self.identifier +} + +///| +pub fn AlignMauveIndexEntry::start(self : AlignMauveIndexEntry) -> Int { + self.start +} + +///| +pub fn AlignMauveIndexEntry::end(self : AlignMauveIndexEntry) -> Int { + self.end +} + +///| +pub fn AlignMauveIndexEntry::block_index(self : AlignMauveIndexEntry) -> Int { + self.block_index +} + +///| +pub fn AlignMauveIndexEntry::row_index(self : AlignMauveIndexEntry) -> Int { + self.row_index +} + +///| +pub fn AlignMauveIndexEntry::strand(self : AlignMauveIndexEntry) -> String { + self.strand +} + +///| +/// Project one source residue through every block containing both sequences. +pub fn align_mauve_map_position( + document : AlignMauveDocument, + source_identifier : String, + target_identifier : String, + source_position : Int, +) -> Array[AlignMauvePositionMapping] raise AlignMauveError { + if align_mauve_find_source_by_identifier(document.sources, source_identifier) + is None { + align_mauve_fail( + "Unknown XMFA source identifier '" + source_identifier + "'", + ) + } + if align_mauve_find_source_by_identifier(document.sources, target_identifier) + is None { + align_mauve_fail( + "Unknown XMFA target identifier '" + target_identifier + "'", + ) + } + if source_position < 0 { + align_mauve_fail("XMFA source position must be non-negative") + } + let mappings : Array[AlignMauvePositionMapping] = [] + for block_index in 0.. row + None => continue + } + let target = match block.row_by_identifier(target_identifier) { + Some(row) => row + None => continue + } + let column = match source.column_for_coordinate(source_position) { + Some(value) => value + None => continue + } + match target.residue_coordinate(column) { + Some(target_position) => + mappings.push(AlignMauvePositionMapping::{ + block_index, + column, + source_position, + target_position, + }) + None => () + } + } + mappings +} + +///| +pub fn AlignMauvePositionMapping::block_index( + self : AlignMauvePositionMapping, +) -> Int { + self.block_index +} + +///| +pub fn AlignMauvePositionMapping::column( + self : AlignMauvePositionMapping, +) -> Int { + self.column +} + +///| +pub fn AlignMauvePositionMapping::source_position( + self : AlignMauvePositionMapping, +) -> Int { + self.source_position +} + +///| +pub fn AlignMauvePositionMapping::target_position( + self : AlignMauvePositionMapping, +) -> Int { + self.target_position +} + +///| +fn align_mauve_append_range_mapping( + mappings : Array[AlignMauveRangeMapping], + block_index : Int, + column_start : Int, + column_end : Int, + first_source : Int, + last_source : Int, + first_target : Int, + last_target : Int, + source_strand : String, + target_strand : String, +) -> Unit { + let source_start = if first_source < last_source { + first_source + } else { + last_source + } + let source_end = if first_source > last_source { + first_source + 1 + } else { + last_source + 1 + } + let target_start = if first_target < last_target { + first_target + } else { + last_target + } + let target_end = if first_target > last_target { + first_target + 1 + } else { + last_target + 1 + } + mappings.push(AlignMauveRangeMapping::{ + block_index, + column_start, + column_end, + source_start, + source_end, + target_start, + target_end, + source_strand, + target_strand, + }) +} + +///| +/// Project a source interval through all shared, residue-aligned block runs. +/// +/// Gaps split the result into separate mappings. Reverse rows retain their +/// strand labels while returned coordinate intervals remain forward-axis. +pub fn align_mauve_map_range( + document : AlignMauveDocument, + source_identifier : String, + target_identifier : String, + start : Int, + end : Int, +) -> Array[AlignMauveRangeMapping] raise AlignMauveError { + if start < 0 || end <= start { + align_mauve_fail("XMFA source interval must satisfy 0 <= start < end") + } + if align_mauve_find_source_by_identifier(document.sources, source_identifier) + is None { + align_mauve_fail( + "Unknown XMFA source identifier '" + source_identifier + "'", + ) + } + if align_mauve_find_source_by_identifier(document.sources, target_identifier) + is None { + align_mauve_fail( + "Unknown XMFA target identifier '" + target_identifier + "'", + ) + } + let mappings : Array[AlignMauveRangeMapping] = [] + for block_index in 0.. row + None => continue + } + let target = match block.row_by_identifier(target_identifier) { + Some(row) => row + None => continue + } + let source_step = if source.strand == "+" { 1 } else { -1 } + let target_step = if target.strand == "+" { 1 } else { -1 } + let mut active = false + let mut column_start = 0 + let mut first_source = 0 + let mut last_source = 0 + let mut first_target = 0 + let mut last_target = 0 + let mut previous_column = -2 + for column in 0.. + source_value >= start && source_value < end + _ => false + } + if usable { + let source_value = source_position.unwrap() + let target_value = target_position.unwrap() + let contiguous = active && + column == previous_column + 1 && + source_value == last_source + source_step && + target_value == last_target + target_step + if !contiguous { + if active { + align_mauve_append_range_mapping( + mappings, + block_index, + column_start, + previous_column + 1, + first_source, + last_source, + first_target, + last_target, + source.strand, + target.strand, + ) + } + active = true + column_start = column + first_source = source_value + first_target = target_value + } + last_source = source_value + last_target = target_value + previous_column = column + } else if active { + align_mauve_append_range_mapping( + mappings, + block_index, + column_start, + previous_column + 1, + first_source, + last_source, + first_target, + last_target, + source.strand, + target.strand, + ) + active = false + } + } + if active { + align_mauve_append_range_mapping( + mappings, + block_index, + column_start, + previous_column + 1, + first_source, + last_source, + first_target, + last_target, + source.strand, + target.strand, + ) + } + } + mappings +} + +///| +pub fn AlignMauveRangeMapping::block_index( + self : AlignMauveRangeMapping, +) -> Int { + self.block_index +} + +///| +pub fn AlignMauveRangeMapping::column_start( + self : AlignMauveRangeMapping, +) -> Int { + self.column_start +} + +///| +pub fn AlignMauveRangeMapping::column_end(self : AlignMauveRangeMapping) -> Int { + self.column_end +} + +///| +pub fn AlignMauveRangeMapping::source_start( + self : AlignMauveRangeMapping, +) -> Int { + self.source_start +} + +///| +pub fn AlignMauveRangeMapping::source_end(self : AlignMauveRangeMapping) -> Int { + self.source_end +} + +///| +pub fn AlignMauveRangeMapping::target_start( + self : AlignMauveRangeMapping, +) -> Int { + self.target_start +} + +///| +pub fn AlignMauveRangeMapping::target_end(self : AlignMauveRangeMapping) -> Int { + self.target_end +} + +///| +pub fn AlignMauveRangeMapping::source_strand( + self : AlignMauveRangeMapping, +) -> String { + self.source_strand +} + +///| +pub fn AlignMauveRangeMapping::target_strand( + self : AlignMauveRangeMapping, +) -> String { + self.target_strand +} + +///| +/// Reconstruct the known forward source sequence from all XMFA blocks. +/// +/// Unknown positions up to the largest observed endpoint use `fill`, which +/// must be exactly one non-gap IUPAC character. Conflicting overlapping rows +/// are rejected. +pub fn align_mauve_reconstruct_sequence( + document : AlignMauveDocument, + identifier : String, + fill? : String = "N", +) -> String raise AlignMauveError { + if align_mauve_find_source_by_identifier(document.sources, identifier) is None { + align_mauve_fail("Unknown XMFA source identifier '" + identifier + "'") + } + if fill.length() != 1 || fill == "-" { + align_mauve_fail( + "XMFA reconstruction fill must be one non-gap IUPAC character", + ) + } + let fill_char = fill.unsafe_get(0).unsafe_to_char() + ignore(align_mauve_complement(fill_char)) + let mut length = 0 + for block in document.blocks { + for row in block.rows { + if row.identifier == identifier && row.end > length { + length = row.end + } + } + } + let sequence : Array[Char] = [] + for _ in 0.. AlignMauveSummary { + let counts = document.counts() + let mut rows = 0 + let mut alignment_columns = 0 + for block in document.blocks { + rows = rows + block.rows.length() + alignment_columns = alignment_columns + block.width + } + AlignMauveSummary::{ + sources: document.sources.length(), + blocks: document.blocks.length(), + rows, + alignment_columns, + aligned_pairs: counts.aligned, + identities: counts.identities, + mismatches: counts.mismatches, + gaps: counts.gaps, + } +} + +///| +pub fn AlignMauveSummary::sources(self : AlignMauveSummary) -> Int { + self.sources +} + +///| +pub fn AlignMauveSummary::blocks(self : AlignMauveSummary) -> Int { + self.blocks +} + +///| +pub fn AlignMauveSummary::rows(self : AlignMauveSummary) -> Int { + self.rows +} + +///| +pub fn AlignMauveSummary::alignment_columns(self : AlignMauveSummary) -> Int { + self.alignment_columns +} + +///| +pub fn AlignMauveSummary::aligned_pairs(self : AlignMauveSummary) -> Int { + self.aligned_pairs +} + +///| +pub fn AlignMauveSummary::identities(self : AlignMauveSummary) -> Int { + self.identities +} + +///| +pub fn AlignMauveSummary::mismatches(self : AlignMauveSummary) -> Int { + self.mismatches +} + +///| +pub fn AlignMauveSummary::gaps(self : AlignMauveSummary) -> Int { + self.gaps +} + +///| +/// Small deterministic XMFA fixture adapted from Biopython's official Mauve +/// combined-file test data. +pub fn align_mauve_example_text() -> String { + "#FormatVersion Mauve1\n" + + "#Sequence1File\tcombined.fa\n" + + "#Sequence1Entry\t1\n" + + "#Sequence1Format\tFastA\n" + + "#Sequence2File\tcombined.fa\n" + + "#Sequence2Entry\t2\n" + + "#Sequence2Format\tFastA\n" + + "#Sequence3File\tcombined.fa\n" + + "#Sequence3Entry\t3\n" + + "#Sequence3Format\tFastA\n" + + "#BackboneFile\tcombined.xmfa.bbcols\n" + + "> 1:2-49 - combined.fa\n" + + "AAGCCCTCCTAGCACACACCCGGAGTGG-CCGGGCCGTACTTTCCTTTT\n" + + "> 2:0-0 + combined.fa\n" + + "-------------------------------------------------\n" + + "> 3:2-48 + combined.fa\n" + + "AAGCCCTGC--GCGCTCAGCCGGAGTGTCCCGGGCCCTGCTTTCCTTTT\n" + + "=\n" + + "> 1:1-1 + combined.fa\n" + + "G\n" + + "=\n" + + "> 1:50-50 + combined.fa\n" + + "A\n" + + "=\n" + + "> 2:1-41 + combined.fa\n" + + "GAAGAGGAAAAGTAGATCCCTGGCGTCCGGAGCTGGGACGT\n" + + "=\n" + + "> 3:1-1 + combined.fa\n" + + "C\n" + + "=\n" + + "> 3:49-49 + combined.fa\n" + + "C\n" + + "=\n" +} diff --git a/src/align_nexus.mbt b/src/align_nexus.mbt new file mode 100644 index 00000000..631f335b --- /dev/null +++ b/src/align_nexus.mbt @@ -0,0 +1,2111 @@ +// Biopython-compatible Bio.Align.nexus support. +// +// This module is intentionally separate from nexus.mbt. The older module +// exposes Bio.Nexus-style block access, while this file models one modern, +// coordinate-aware alignment matrix. + +///| +/// Error raised for malformed NEXUS alignment data or invalid operations. +pub suberror AlignNexusError { + AlignNexusError(String) +} + +///| +/// Sequence datatype declared by a NEXUS FORMAT command. +pub(all) enum AlignNexusDataType { + AlignNexusDna + AlignNexusRna + AlignNexusProtein + AlignNexusStandard +} derive(Eq, Debug) + +///| +/// File-level metadata retained from a DATA or CHARACTERS block. +/// +/// `declared_characters` is the NCHAR value in the source file. +/// `removed_all_gap_columns` records columns omitted from the alignment, +/// matching Biopython's `Bio.Align.nexus` behavior. +pub struct AlignNexusMetadata { + block_name : String + declared_taxa : Int + declared_characters : Int + data_type : AlignNexusDataType + missing_character : String + gap_character : String + match_character : String? + interleaved : Bool + respect_case : Bool + symbols : String + removed_all_gap_columns : Int +} derive(Eq, Debug) + +///| +/// One NEXUS matrix row. +/// +/// `aligned_sequence` always uses `-` for gaps. `sequence` is the same row +/// with gaps removed; missing-data symbols remain coordinate-bearing. +pub struct AlignNexusSequence { + id : String + sequence : String + aligned_sequence : String +} derive(Eq, Debug) + +///| +/// A coordinate-aware NEXUS multiple sequence alignment. +pub struct AlignNexusAlignment { + metadata : AlignNexusMetadata + sequences : Array[AlignNexusSequence] +} derive(Eq, Debug) + +///| +/// Pairwise or all-pairs alignment statistics. +pub struct AlignNexusCounts { + pairs : Int + aligned : Int + identities : Int + mismatches : Int + gap_columns : Int + double_gap_columns : Int + gap_opens : Int +} derive(Eq, Debug) + +///| +/// Internal FORMAT state used while parsing commands. +priv struct AlignNexusFormat { + data_type : AlignNexusDataType + missing_character : String + gap_character : String + match_character : String? + interleaved : Bool + respect_case : Bool + symbols : String + labels : Bool +} + +///| +/// Construct validated NEXUS alignment metadata. +pub fn AlignNexusMetadata::create( + declared_taxa : Int, + declared_characters : Int, + data_type : AlignNexusDataType, + block_name? : String = "data", + missing_character? : String = "?", + gap_character? : String = "-", + match_character? : String? = None, + interleaved? : Bool = false, + respect_case? : Bool = false, + symbols? : String = "", + removed_all_gap_columns? : Int = 0, +) -> AlignNexusMetadata raise AlignNexusError { + if declared_taxa <= 0 { + raise AlignNexusError("NEXUS NTAX must be positive") + } + if declared_characters <= 0 { + raise AlignNexusError("NEXUS NCHAR must be positive") + } + let normalized_block = block_name.to_lower().trim().to_owned() + if normalized_block != "data" && normalized_block != "characters" { + raise AlignNexusError("NEXUS alignment block must be DATA or CHARACTERS") + } + align_nexus_validate_format_char( + missing_character, "NEXUS missing-data character", + ) + align_nexus_validate_format_char(gap_character, "NEXUS gap character") + if missing_character == gap_character { + raise AlignNexusError( + "NEXUS missing-data and gap characters must be different", + ) + } + match match_character { + Some(value) => { + align_nexus_validate_format_char(value, "NEXUS match character") + if value == missing_character || value == gap_character { + raise AlignNexusError( + "NEXUS match character must differ from missing and gap characters", + ) + } + } + None => () + } + align_nexus_validate_symbols(data_type, symbols, respect_case) + if removed_all_gap_columns < 0 || + removed_all_gap_columns >= declared_characters { + raise AlignNexusError("NEXUS removed all-gap column count is outside NCHAR") + } + AlignNexusMetadata::{ + block_name: normalized_block, + declared_taxa, + declared_characters, + data_type, + missing_character, + gap_character, + match_character, + interleaved, + respect_case, + symbols, + removed_all_gap_columns, + } +} + +///| +/// Construct and normalize one standalone NEXUS row. +pub fn AlignNexusSequence::create( + id : String, + aligned_sequence : String, + data_type : AlignNexusDataType, + missing_character? : String = "?", + gap_character? : String = "-", + symbols? : String = "", + respect_case? : Bool = false, +) -> AlignNexusSequence raise AlignNexusError { + align_nexus_validate_id(id) + if aligned_sequence.length() == 0 { + raise AlignNexusError("NEXUS aligned sequence must not be empty") + } + align_nexus_validate_format_char( + missing_character, "NEXUS missing-data character", + ) + align_nexus_validate_format_char(gap_character, "NEXUS gap character") + if missing_character == gap_character { + raise AlignNexusError( + "NEXUS missing-data and gap characters must be different", + ) + } + align_nexus_validate_symbols(data_type, symbols, respect_case) + let normalized = align_nexus_normalize_row( + aligned_sequence, data_type, missing_character, gap_character, symbols, respect_case, + ) + AlignNexusSequence::{ + id, + sequence: align_nexus_remove_gaps(normalized), + aligned_sequence: normalized, + } +} + +///| +/// Construct a validated coordinate-aware NEXUS alignment. +pub fn AlignNexusAlignment::create( + metadata : AlignNexusMetadata, + sequences : Array[AlignNexusSequence], +) -> AlignNexusAlignment raise AlignNexusError { + let copied : Array[AlignNexusSequence] = [] + for sequence in sequences { + copied.push(sequence) + } + let alignment = AlignNexusAlignment::{ metadata, sequences: copied } + align_nexus_validate_alignment(alignment) + alignment +} + +///| +/// Construct an alignment directly from identifiers and printed rows. +/// +/// Source gap symbols are normalized to `-`, and columns containing only gaps +/// are removed exactly as they are when parsing a NEXUS file. +pub fn align_nexus_from_aligned( + ids : Array[String], + aligned_sequences : Array[String], + data_type : AlignNexusDataType, + missing_character? : String = "?", + gap_character? : String = "-", + symbols? : String = "", + respect_case? : Bool = false, +) -> AlignNexusAlignment raise AlignNexusError { + if ids.length() == 0 { + raise AlignNexusError("NEXUS alignment must contain at least one sequence") + } + if ids.length() != aligned_sequences.length() { + raise AlignNexusError( + "NEXUS identifiers and aligned rows must have equal lengths", + ) + } + let width = aligned_sequences[0].length() + if width == 0 { + raise AlignNexusError("NEXUS alignment must contain at least one column") + } + for row in aligned_sequences { + if row.length() != width { + raise AlignNexusError("NEXUS aligned rows must have equal widths") + } + } + let normalized : Array[String] = [] + for row in aligned_sequences { + normalized.push( + align_nexus_normalize_row( + row, data_type, missing_character, gap_character, symbols, respect_case, + ), + ) + } + let prepared = align_nexus_remove_all_gap_columns(normalized) + if prepared.0.length() == 0 || prepared.0[0].length() == 0 { + raise AlignNexusError( + "NEXUS alignment must contain a non-gap alignment column", + ) + } + let metadata = AlignNexusMetadata::create( + ids.length(), + width, + data_type, + missing_character~, + gap_character~, + respect_case~, + symbols~, + removed_all_gap_columns=prepared.1, + ) + let sequences : Array[AlignNexusSequence] = [] + for index = 0; index < ids.length(); index = index + 1 { + align_nexus_validate_id(ids[index]) + sequences.push(AlignNexusSequence::{ + id: ids[index], + sequence: align_nexus_remove_gaps(prepared.0[index]), + aligned_sequence: prepared.0[index], + }) + } + AlignNexusAlignment::create(metadata, sequences) +} + +///| +/// Parse one Biopython-compatible NEXUS alignment. +/// +/// The parser recognizes DATA and CHARACTERS matrices, TAXA/TAXLABELS, +/// sequential and interleaved layouts, nested comments, quoted identifiers, +/// doubled quote escaping, custom missing/gap/match symbols, and the DNA, RNA, +/// protein, nucleotide, and standard datatypes. +pub fn align_nexus_parse( + text : String, +) -> AlignNexusAlignment raise AlignNexusError { + let lines = align_nexus_normalize_lines(text) + if lines.length() == 0 || + (lines.length() == 1 && lines[0].trim().length() == 0) { + raise AlignNexusError("Empty NEXUS file") + } + if lines[0].trim() != "#NEXUS" { + raise AlignNexusError("File does not start with NEXUS header") + } + let body_builder = StringBuilder::new() + for index = 1; index < lines.length(); index = index + 1 { + if index > 1 { + body_builder.write_char('\n') + } + body_builder.write_string(lines[index]) + } + let commands = align_nexus_scan_commands(body_builder.to_string()) + let mut active_block = "" + let mut tax_ntax = -1 + let mut data_ntax = -1 + let mut data_nchar = -1 + let taxlabels : Array[String] = [] + let mut block_name = "" + let mut matrix_body : String? = None + let mut format = AlignNexusFormat::{ + data_type: AlignNexusDna, + missing_character: "?", + gap_character: "-", + match_character: None, + interleaved: false, + respect_case: false, + symbols: "", + labels: true, + } + for command in commands { + let split = align_nexus_split_keyword(command) + let keyword = split.0 + let options = split.1 + if keyword == "begin" { + if active_block.length() != 0 { + raise AlignNexusError("NEXUS blocks must not be nested") + } + let words = align_nexus_tokenize(options) + if words.length() != 1 { + raise AlignNexusError("Malformed NEXUS BEGIN command") + } + active_block = words[0].to_lower() + continue + } + if keyword == "end" || keyword == "endblock" { + if active_block.length() == 0 { + raise AlignNexusError("Unmatched NEXUS END command") + } + active_block = "" + continue + } + if active_block.length() == 0 { + raise AlignNexusError( + "NEXUS command '" + keyword + "' appears outside a block", + ) + } + if active_block == "taxa" { + if keyword == "dimensions" { + let dimensions = align_nexus_parse_dimensions(options) + if dimensions.0 >= 0 { + tax_ntax = dimensions.0 + } + if dimensions.1 >= 0 { + raise AlignNexusError("NCHAR is not valid in a TAXA block") + } + } else if keyword == "taxlabels" { + let parsed_labels = align_nexus_parse_taxlabels(options) + taxlabels.clear() + for label in parsed_labels { + taxlabels.push(label) + } + } else if keyword != "title" && keyword != "link" { + raise AlignNexusError( + "Unsupported command '" + keyword + "' in NEXUS TAXA block", + ) + } + continue + } + if active_block == "data" || active_block == "characters" { + if keyword == "dimensions" { + let dimensions = align_nexus_parse_dimensions(options) + if dimensions.0 >= 0 { + data_ntax = dimensions.0 + } + if dimensions.1 >= 0 { + data_nchar = dimensions.1 + } + } else if keyword == "format" { + format = align_nexus_parse_format(options, format) + } else if keyword == "taxlabels" { + let parsed_labels = align_nexus_parse_taxlabels(options) + taxlabels.clear() + for label in parsed_labels { + taxlabels.push(label) + } + } else if keyword == "matrix" { + match matrix_body { + Some(_) => + raise AlignNexusError( + "A NEXUS file may contain only one alignment matrix", + ) + None => { + matrix_body = Some(options) + block_name = active_block + } + } + } else if keyword != "title" && + keyword != "link" && + keyword != "options" && + keyword != "charlabels" && + keyword != "charstatelabels" && + keyword != "statelabels" && + keyword != "eliminate" { + raise AlignNexusError( + "Unsupported command '" + + keyword + + "' in NEXUS " + + active_block.to_upper() + + " block", + ) + } + continue + } + // Alignment parsing deliberately ignores unrelated SETS, TREES, CODONS, + // and vendor-specific blocks after their command boundaries are scanned. + } + if active_block.length() != 0 { + raise AlignNexusError("NEXUS block is missing its END command") + } + let body = match matrix_body { + Some(value) => value + None => + raise AlignNexusError("NEXUS file does not contain an alignment matrix") + } + let ntax = if data_ntax >= 0 { data_ntax } else { tax_ntax } + if ntax <= 0 { + raise AlignNexusError("NEXUS NTAX must be specified before MATRIX") + } + if data_nchar <= 0 { + raise AlignNexusError("NEXUS NCHAR must be specified before MATRIX") + } + if tax_ntax >= 0 && data_ntax >= 0 && tax_ntax != data_ntax { + raise AlignNexusError("NEXUS TAXA and alignment NTAX values disagree") + } + if taxlabels.length() != 0 && taxlabels.length() != ntax { + raise AlignNexusError("NEXUS TAXLABELS count does not match NTAX") + } + align_nexus_validate_format(format) + let parsed = align_nexus_parse_matrix( + body, ntax, data_nchar, format, taxlabels, + ) + let resolved = align_nexus_resolve_and_normalize_rows(parsed.1, format) + let prepared = align_nexus_remove_all_gap_columns(resolved) + if prepared.0.length() == 0 || prepared.0[0].length() == 0 { + raise AlignNexusError("NEXUS matrix contains no non-gap alignment column") + } + let metadata = AlignNexusMetadata::create( + ntax, + data_nchar, + format.data_type, + block_name~, + missing_character=format.missing_character, + gap_character=format.gap_character, + match_character=format.match_character, + interleaved=format.interleaved, + respect_case=format.respect_case, + symbols=format.symbols, + removed_all_gap_columns=prepared.1, + ) + let sequences : Array[AlignNexusSequence] = [] + for index = 0; index < ntax; index = index + 1 { + sequences.push(AlignNexusSequence::{ + id: parsed.0[index], + sequence: align_nexus_remove_gaps(prepared.0[index]), + aligned_sequence: prepared.0[index], + }) + } + AlignNexusAlignment::create(metadata, sequences) +} + +///| +/// Write one canonical NEXUS DATA alignment. +/// +/// When `interleave` is omitted, alignments wider than 1000 columns are +/// interleaved, matching Biopython 1.86. Interleaved blocks default to 70 +/// columns. +pub fn align_nexus_write( + alignment : AlignNexusAlignment, + interleave? : Bool? = None, + block_width? : Int = 70, +) -> String raise AlignNexusError { + align_nexus_validate_alignment(alignment) + let width = alignment.alignment_length() + if width == 0 { + raise AlignNexusError("Non-empty NEXUS sequences are required") + } + if block_width <= 0 { + raise AlignNexusError("NEXUS writer block width must be positive") + } + let use_interleave = match interleave { + Some(value) => value + None => width > 1000 + } + let names : Array[String] = [] + let mut name_width = 0 + for sequence in alignment.sequences { + let safe = align_nexus_safe_name(sequence.id) + names.push(safe) + if safe.length() > name_width { + name_width = safe.length() + } + } + let output = StringBuilder::new() + output.write_string("#NEXUS\n") + output.write_string("begin data;\n") + output.write_string( + "dimensions ntax=" + + alignment.num_sequences().to_string() + + " nchar=" + + width.to_string() + + ";\n", + ) + output.write_string( + "format datatype=" + + alignment.metadata.data_type.code() + + " missing=" + + alignment.metadata.missing_character + + " gap=" + + alignment.metadata.gap_character, + ) + if alignment.metadata.data_type == AlignNexusStandard { + output.write_string( + " symbols=\"" + + align_nexus_escape_double_quotes(alignment.metadata.symbols) + + "\"", + ) + } + if alignment.metadata.respect_case { + output.write_string(" respectcase") + } + if use_interleave { + output.write_string(" interleave") + } + output.write_string(";\n") + output.write_string("matrix\n") + if use_interleave { + let mut start = 0 + while start < width { + let end = align_nexus_min(start + block_width, width) + for row = 0; row < alignment.sequences.length(); row = row + 1 { + output.write_string(align_nexus_pad_right(names[row], name_width + 1)) + output.write_string( + align_nexus_output_fragment( + alignment.sequences[row].aligned_sequence, + start, + end, + alignment.metadata.gap_character, + ), + ) + output.write_char('\n') + } + output.write_char('\n') + start = end + } + } else { + for row = 0; row < alignment.sequences.length(); row = row + 1 { + output.write_string(align_nexus_pad_right(names[row], name_width + 1)) + output.write_string( + align_nexus_output_fragment( + alignment.sequences[row].aligned_sequence, + 0, + width, + alignment.metadata.gap_character, + ), + ) + output.write_char('\n') + } + } + output.write_string(";\n") + output.write_string("end;\n") + output.to_string() +} + +///| +/// Return the canonical lowercase FORMAT datatype name. +pub fn AlignNexusDataType::code(self : AlignNexusDataType) -> String { + match self { + AlignNexusDna => "dna" + AlignNexusRna => "rna" + AlignNexusProtein => "protein" + AlignNexusStandard => "standard" + } +} + +///| +/// Return the Biopython-style molecule type annotation. +pub fn AlignNexusDataType::molecule_type(self : AlignNexusDataType) -> String { + match self { + AlignNexusDna => "DNA" + AlignNexusRna => "RNA" + AlignNexusProtein => "protein" + AlignNexusStandard => "" + } +} + +///| +/// Return the number of sequence rows. +pub fn AlignNexusAlignment::num_sequences(self : AlignNexusAlignment) -> Int { + self.sequences.length() +} + +///| +/// Return the alignment width after removal of all-gap source columns. +pub fn AlignNexusAlignment::alignment_length(self : AlignNexusAlignment) -> Int { + if self.sequences.length() == 0 { + 0 + } else { + self.sequences[0].aligned_sequence.length() + } +} + +///| +/// Return the source width before all-gap columns were removed. +pub fn AlignNexusAlignment::source_alignment_length( + self : AlignNexusAlignment, +) -> Int { + self.metadata.declared_characters +} + +///| +/// Locate the first row with an exact identifier. +pub fn AlignNexusAlignment::find_sequence( + self : AlignNexusAlignment, + id : String, +) -> Int? { + for index = 0; index < self.sequences.length(); index = index + 1 { + if self.sequences[index].id == id { + return Some(index) + } + } + None +} + +///| +/// Return one printed alignment column. +pub fn AlignNexusAlignment::column( + self : AlignNexusAlignment, + column : Int, +) -> String? { + if column < 0 || column >= self.alignment_length() { + return None + } + let output = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + output.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(output.to_string()) +} + +///| +/// Map a zero-based ungapped row position to an alignment column. +pub fn AlignNexusAlignment::sequence_position_to_column( + self : AlignNexusAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || row >= self.sequences.length() || position < 0 { + return None + } + let aligned = self.sequences[row].aligned_sequence + let mut coordinate = 0 + for column = 0; column < aligned.length(); column = column + 1 { + if aligned.unsafe_get(column).to_int() != '-'.to_int() { + if coordinate == position { + return Some(column) + } + coordinate = coordinate + 1 + } + } + None +} + +///| +/// Map a printed column to a zero-based ungapped row position. +pub fn AlignNexusAlignment::column_to_sequence_position( + self : AlignNexusAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.alignment_length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + if aligned.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut coordinate = 0 + for index = 0; index < column; index = index + 1 { + if aligned.unsafe_get(index).to_int() != '-'.to_int() { + coordinate = coordinate + 1 + } + } + Some(coordinate) +} + +///| +/// Map one residue through the alignment to another row. +pub fn AlignNexusAlignment::map_position( + self : AlignNexusAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column residue coordinates for two rows. +pub fn AlignNexusAlignment::aligned_pairs( + self : AlignNexusAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise AlignNexusError { + align_nexus_validate_row(self, first_row) + align_nexus_validate_row(self, second_row) + let result : Array[(Int?, Int?)] = [] + let mut first_position = 0 + let mut second_position = 0 + for column = 0; column < self.alignment_length(); column = column + 1 { + let first_gap = self.sequences[first_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + let second_gap = self.sequences[second_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + result.push( + ( + if first_gap { + None + } else { + Some(first_position) + }, + if second_gap { + None + } else { + Some(second_position) + }, + ), + ) + if !first_gap { + first_position = first_position + 1 + } + if !second_gap { + second_position = second_position + 1 + } + } + result +} + +///| +/// Return a compact Biopython-style coordinate path for all rows. +pub fn AlignNexusAlignment::coordinate_path( + self : AlignNexusAlignment, +) -> Array[Array[Int]] { + let paths : Array[Array[Int]] = [] + let coordinates = Array::make(self.sequences.length(), 0) + for _row = 0; _row < self.sequences.length(); _row = _row + 1 { + paths.push([0]) + } + let width = self.alignment_length() + for column = 0; column < width; column = column + 1 { + for row = 0; row < self.sequences.length(); row = row + 1 { + if self.sequences[row].aligned_sequence.unsafe_get(column).to_int() != + '-'.to_int() { + coordinates[row] = coordinates[row] + 1 + } + } + let boundary = if column + 1 == width { + true + } else { + align_nexus_movement_changes(self, column, column + 1) + } + if boundary { + for row = 0; row < self.sequences.length(); row = row + 1 { + paths[row].push(coordinates[row]) + } + } + } + paths +} + +///| +/// Compute statistics for one pair of rows. +pub fn AlignNexusAlignment::pair_counts( + self : AlignNexusAlignment, + first_row : Int, + second_row : Int, +) -> AlignNexusCounts raise AlignNexusError { + align_nexus_validate_row(self, first_row) + align_nexus_validate_row(self, second_row) + align_nexus_count_pair(self, first_row, second_row) +} + +///| +/// Aggregate statistics across every unordered pair of rows. +pub fn AlignNexusAlignment::counts( + self : AlignNexusAlignment, +) -> AlignNexusCounts { + let mut pairs = 0 + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + for first = 0; first < self.sequences.length(); first = first + 1 { + for second = first + 1 + second < self.sequences.length() + second = second + 1 { + let counts = align_nexus_count_pair(self, first, second) + pairs = pairs + 1 + aligned = aligned + counts.aligned + identities = identities + counts.identities + mismatches = mismatches + counts.mismatches + gap_columns = gap_columns + counts.gap_columns + double_gap_columns = double_gap_columns + counts.double_gap_columns + gap_opens = gap_opens + counts.gap_opens + } + } + AlignNexusCounts::{ + pairs, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +/// Return exact identity over columns containing residues in both rows. +pub fn AlignNexusCounts::identity(self : AlignNexusCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Calculate a majority consensus. Gaps do not vote. +pub fn AlignNexusAlignment::consensus( + self : AlignNexusAlignment, + minimum_fraction? : Double = 0.0, +) -> String raise AlignNexusError { + if minimum_fraction != minimum_fraction || + minimum_fraction < 0.0 || + minimum_fraction > 1.0 { + raise AlignNexusError( + "NEXUS consensus minimum fraction must be between 0 and 1", + ) + } + let output = StringBuilder::new(size_hint=self.alignment_length()) + for column = 0; column < self.alignment_length(); column = column + 1 { + let counts = Array::make(128, 0) + let mut residues = 0 + for sequence in self.sequences { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if code != '-'.to_int() { + if code >= 0 && code < counts.length() { + counts[code] = counts[code] + 1 + } + residues = residues + 1 + } + } + if residues == 0 { + output.write_char('-') + continue + } + let mut best_code = 0 + let mut best_count = -1 + for code = 0; code < counts.length(); code = code + 1 { + if counts[code] > best_count { + best_code = code + best_count = counts[code] + } + } + if best_count.to_double() / residues.to_double() < minimum_fraction { + output.write_char('X') + } else { + output.write_char(best_code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +/// Return per-column non-gap occupancy. +pub fn AlignNexusAlignment::occupancy( + self : AlignNexusAlignment, +) -> Array[Double] { + let result : Array[Double] = [] + for column = 0; column < self.alignment_length(); column = column + 1 { + let mut residues = 0 + for sequence in self.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + result.push(residues.to_double() / self.sequences.length().to_double()) + } + result +} + +///| +/// Return a compact alignment summary. +pub fn AlignNexusAlignment::summary(self : AlignNexusAlignment) -> String { + "AlignNexusAlignment(type=" + + self.metadata.data_type.code() + + ", sequences=" + + self.num_sequences().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ", source_columns=" + + self.source_alignment_length().to_string() + + ", removed_all_gap=" + + self.metadata.removed_all_gap_columns.to_string() + + ")" +} + +///| +/// Return a compact interleaved NEXUS example with quoting and MATCHCHAR. +pub fn align_nexus_example_text() -> String { + "#NEXUS\n" + + "[MoonBit NEXUS example with a [nested] comment]\n" + + "begin data;\n" + + "dimensions ntax=3 nchar=12;\n" + + "format datatype=dna missing=? gap=- matchchar=. interleave=yes;\n" + + "matrix\n" + + "reference ACGT--\n" + + "'query one' A.GT??\n" + + "'isn''t_query' .C-T--\n" + + "\n" + + "reference ACGTAC\n" + + "query_one ..G-AC\n" + + "'isn''t_query' A-GTAC\n" + + ";\n" + + "end;\n" +} + +///| +fn align_nexus_validate_alignment( + alignment : AlignNexusAlignment, +) -> Unit raise AlignNexusError { + if alignment.sequences.length() == 0 { + raise AlignNexusError("NEXUS alignment must contain at least one sequence") + } + if alignment.sequences.length() != alignment.metadata.declared_taxa { + raise AlignNexusError("NEXUS sequence count does not match declared NTAX") + } + let width = alignment.sequences[0].aligned_sequence.length() + if width + alignment.metadata.removed_all_gap_columns != + alignment.metadata.declared_characters { + raise AlignNexusError( + "NEXUS alignment width and removed columns do not match NCHAR", + ) + } + align_nexus_validate_format_char( + alignment.metadata.missing_character, + "NEXUS missing-data character", + ) + align_nexus_validate_format_char( + alignment.metadata.gap_character, + "NEXUS gap character", + ) + align_nexus_validate_symbols( + alignment.metadata.data_type, + alignment.metadata.symbols, + alignment.metadata.respect_case, + ) + for sequence in alignment.sequences { + align_nexus_validate_id(sequence.id) + if sequence.aligned_sequence.length() != width { + raise AlignNexusError("NEXUS aligned rows must have equal widths") + } + for index = 0; index < sequence.aligned_sequence.length(); index = index + 1 { + let code = sequence.aligned_sequence.unsafe_get(index).to_int() + if code != '-'.to_int() && + code != alignment.metadata.missing_character.unsafe_get(0).to_int() && + !align_nexus_valid_residue( + code, + alignment.metadata.data_type, + alignment.metadata.symbols, + alignment.metadata.respect_case, + ) { + raise AlignNexusError( + "NEXUS sequence contains a residue inconsistent with FORMAT", + ) + } + } + if align_nexus_remove_gaps(sequence.aligned_sequence) != sequence.sequence { + raise AlignNexusError( + "NEXUS sequence row has inconsistent derived fields", + ) + } + } + for column = 0; column < width; column = column + 1 { + let mut has_residue = false + for sequence in alignment.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + has_residue = true + } + } + if !has_residue { + raise AlignNexusError( + "NEXUS alignment contains an unremoved all-gap column", + ) + } + } +} + +///| +fn align_nexus_validate_row( + alignment : AlignNexusAlignment, + row : Int, +) -> Unit raise AlignNexusError { + if row < 0 || row >= alignment.sequences.length() { + raise AlignNexusError("NEXUS row index is out of bounds") + } +} + +///| +fn align_nexus_count_pair( + alignment : AlignNexusAlignment, + first_row : Int, + second_row : Int, +) -> AlignNexusCounts { + let first = alignment.sequences[first_row].aligned_sequence + let second = alignment.sequences[second_row].aligned_sequence + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + let mut in_gap = false + for column = 0; column < alignment.alignment_length(); column = column + 1 { + let first_code = first.unsafe_get(column).to_int() + let second_code = second.unsafe_get(column).to_int() + let first_gap = first_code == '-'.to_int() + let second_gap = second_code == '-'.to_int() + if first_gap && second_gap { + double_gap_columns = double_gap_columns + 1 + in_gap = false + } else if first_gap || second_gap { + gap_columns = gap_columns + 1 + if !in_gap { + gap_opens = gap_opens + 1 + } + in_gap = true + } else { + aligned = aligned + 1 + if first_code == second_code { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + in_gap = false + } + } + AlignNexusCounts::{ + pairs: 1, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +fn align_nexus_scan_commands( + body : String, +) -> Array[String] raise AlignNexusError { + let commands : Array[String] = [] + let mut current = StringBuilder::new() + let mut quote = 0 + let mut comment_depth = 0 + let mut index = 0 + while index < body.length() { + let code = body.unsafe_get(index).to_int() + if comment_depth > 0 { + if code == '['.to_int() { + comment_depth = comment_depth + 1 + } else if code == ']'.to_int() { + comment_depth = comment_depth - 1 + } else if code == '\n'.to_int() { + current.write_char('\n') + } + index = index + 1 + continue + } + if quote != 0 { + current.write_char(code.unsafe_to_char()) + if code == quote { + if index + 1 < body.length() && + body.unsafe_get(index + 1).to_int() == quote { + current.write_char(quote.unsafe_to_char()) + index = index + 2 + continue + } + quote = 0 + } + index = index + 1 + continue + } + if code == '['.to_int() { + comment_depth = 1 + } else if code == ']'.to_int() { + raise AlignNexusError("Unmatched closing NEXUS comment bracket") + } else if code == '\''.to_int() || code == '"'.to_int() { + quote = code + current.write_char(code.unsafe_to_char()) + } else if code == ';'.to_int() { + let command = current.to_string().trim().to_owned() + if command.length() > 0 { + commands.push(command) + } + current = StringBuilder::new() + } else { + current.write_char(code.unsafe_to_char()) + } + index = index + 1 + } + if comment_depth != 0 { + raise AlignNexusError("Unterminated NEXUS comment") + } + if quote != 0 { + raise AlignNexusError("Unterminated quoted NEXUS token") + } + if current.to_string().trim().length() != 0 { + raise AlignNexusError("NEXUS command is missing a semicolon") + } + commands +} + +///| +fn align_nexus_split_keyword(command : String) -> (String, String) { + let mut index = 0 + while index < command.length() && + align_nexus_is_whitespace(command.unsafe_get(index).to_int()) { + index = index + 1 + } + let start = index + while index < command.length() && + !align_nexus_is_whitespace(command.unsafe_get(index).to_int()) { + index = index + 1 + } + ( + command[start:index].to_owned().to_lower(), + command[index:].to_owned().trim().to_owned(), + ) +} + +///| +fn align_nexus_tokenize(text : String) -> Array[String] raise AlignNexusError { + let tokens : Array[String] = [] + let mut index = 0 + while index < text.length() { + while index < text.length() && + align_nexus_is_whitespace(text.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= text.length() { + break + } + let code = text.unsafe_get(index).to_int() + if code == '='.to_int() || code == ','.to_int() { + tokens.push(code.unsafe_to_char().to_string()) + index = index + 1 + continue + } + if code == '\''.to_int() || code == '"'.to_int() { + let quote = code + let output = StringBuilder::new() + index = index + 1 + let mut closed = false + while index < text.length() { + let current = text.unsafe_get(index).to_int() + if current == quote { + if index + 1 < text.length() && + text.unsafe_get(index + 1).to_int() == quote { + output.write_char(quote.unsafe_to_char()) + index = index + 2 + } else { + index = index + 1 + closed = true + break + } + } else { + output.write_char(current.unsafe_to_char()) + index = index + 1 + } + } + if !closed { + raise AlignNexusError("Unterminated quoted NEXUS token") + } + tokens.push(output.to_string()) + continue + } + let start = index + while index < text.length() { + let current = text.unsafe_get(index).to_int() + if align_nexus_is_whitespace(current) || + current == '='.to_int() || + current == ','.to_int() { + break + } + index = index + 1 + } + tokens.push(text[start:index].to_owned()) + } + tokens +} + +///| +fn align_nexus_parse_dimensions( + text : String, +) -> (Int, Int) raise AlignNexusError { + let tokens = align_nexus_tokenize(text) + let mut ntax = -1 + let mut nchar = -1 + let mut index = 0 + while index < tokens.length() { + let key = tokens[index].to_lower() + if key == "newtaxa" { + index = index + 1 + continue + } + if key != "ntax" && key != "nchar" { + raise AlignNexusError( + "Unknown NEXUS DIMENSIONS option '" + tokens[index] + "'", + ) + } + if index + 2 >= tokens.length() || tokens[index + 1] != "=" { + raise AlignNexusError("Malformed NEXUS DIMENSIONS assignment") + } + let value = align_nexus_parse_positive_int( + tokens[index + 2], + "NEXUS " + key.to_upper(), + ) + if key == "ntax" { + if ntax >= 0 { + raise AlignNexusError("Duplicated NEXUS NTAX assignment") + } + ntax = value + } else { + if nchar >= 0 { + raise AlignNexusError("Duplicated NEXUS NCHAR assignment") + } + nchar = value + } + index = index + 3 + } + (ntax, nchar) +} + +///| +fn align_nexus_parse_format( + text : String, + initial : AlignNexusFormat, +) -> AlignNexusFormat raise AlignNexusError { + let tokens = align_nexus_tokenize(text) + let mut data_type = initial.data_type + let mut missing = initial.missing_character + let mut gap = initial.gap_character + let mut match_character = initial.match_character + let mut interleaved = initial.interleaved + let mut respect_case = initial.respect_case + let mut symbols = initial.symbols + let mut labels = initial.labels + let mut index = 0 + while index < tokens.length() { + let key = tokens[index].to_lower() + if key == "interleave" { + if index + 1 < tokens.length() && tokens[index + 1] == "=" { + if index + 2 >= tokens.length() { + raise AlignNexusError("Missing NEXUS INTERLEAVE value") + } + interleaved = align_nexus_parse_yes_no( + tokens[index + 2], + "NEXUS INTERLEAVE", + ) + index = index + 3 + } else { + interleaved = true + index = index + 1 + } + continue + } + if key == "nointerleave" { + interleaved = false + index = index + 1 + continue + } + if key == "respectcase" { + respect_case = true + index = index + 1 + continue + } + if key == "labels" { + if index + 1 < tokens.length() && tokens[index + 1] == "=" { + if index + 2 >= tokens.length() { + raise AlignNexusError("Missing NEXUS LABELS value") + } + let value = tokens[index + 2].to_lower() + if value == "no" || value == "none" { + labels = false + } else if value == "yes" || value == "left" { + labels = true + } else { + raise AlignNexusError("Unsupported NEXUS LABELS value") + } + index = index + 3 + } else { + labels = true + index = index + 1 + } + continue + } + if key == "nolabels" { + labels = false + index = index + 1 + continue + } + if key == "transpose" || + key == "tokens" || + key == "notokens" || + key == "equate" { + raise AlignNexusError( + "NEXUS FORMAT option '" + key + "' is not supported for alignments", + ) + } + if key != "datatype" && + key != "missing" && + key != "gap" && + key != "matchchar" && + key != "symbols" { + raise AlignNexusError( + "Unknown NEXUS FORMAT option '" + tokens[index] + "'", + ) + } + if index + 2 >= tokens.length() || tokens[index + 1] != "=" { + raise AlignNexusError( + "Malformed NEXUS FORMAT assignment for " + key.to_upper(), + ) + } + let value = tokens[index + 2] + if key == "datatype" { + data_type = match value.to_lower() { + "dna" | "nucleotide" => AlignNexusDna + "rna" => AlignNexusRna + "protein" => AlignNexusProtein + "standard" => AlignNexusStandard + other => + raise AlignNexusError("Unsupported NEXUS datatype '" + other + "'") + } + } else if key == "missing" { + missing = value + } else if key == "gap" { + gap = value + } else if key == "matchchar" { + match_character = Some(value) + } else { + symbols = value + } + index = index + 3 + } + let result = AlignNexusFormat::{ + data_type, + missing_character: missing, + gap_character: gap, + match_character, + interleaved, + respect_case, + symbols, + labels, + } + align_nexus_validate_format(result) + result +} + +///| +fn align_nexus_validate_format( + format : AlignNexusFormat, +) -> Unit raise AlignNexusError { + align_nexus_validate_format_char( + format.missing_character, + "NEXUS missing-data character", + ) + align_nexus_validate_format_char(format.gap_character, "NEXUS gap character") + if format.missing_character == format.gap_character { + raise AlignNexusError( + "NEXUS missing-data and gap characters must be different", + ) + } + match format.match_character { + Some(value) => { + align_nexus_validate_format_char(value, "NEXUS match character") + if value == format.missing_character || value == format.gap_character { + raise AlignNexusError( + "NEXUS match character must differ from missing and gap characters", + ) + } + } + None => () + } + align_nexus_validate_symbols( + format.data_type, + format.symbols, + format.respect_case, + ) +} + +///| +fn align_nexus_parse_taxlabels( + text : String, +) -> Array[String] raise AlignNexusError { + let tokens = align_nexus_tokenize(text) + let labels : Array[String] = [] + for token in tokens { + if token != "," { + align_nexus_validate_id(token) + labels.push(token) + } + } + if labels.length() == 0 { + raise AlignNexusError("NEXUS TAXLABELS must not be empty") + } + labels +} + +///| +fn align_nexus_parse_matrix( + body : String, + ntax : Int, + nchar : Int, + format : AlignNexusFormat, + taxlabels : Array[String], +) -> (Array[String], Array[String]) raise AlignNexusError { + let raw_lines = align_nexus_normalize_lines(body) + let lines : Array[String] = [] + for raw in raw_lines { + let trimmed = raw.trim().to_owned() + if trimmed.length() > 0 { + lines.push(trimmed) + } + } + if lines.length() == 0 { + raise AlignNexusError("NEXUS MATRIX must not be empty") + } + if format.interleaved { + align_nexus_parse_interleaved_matrix( + lines, + ntax, + nchar, + format.labels, + taxlabels, + ) + } else { + align_nexus_parse_sequential_matrix( + lines, + ntax, + nchar, + format.labels, + taxlabels, + ) + } +} + +///| +fn align_nexus_parse_interleaved_matrix( + lines : Array[String], + ntax : Int, + nchar : Int, + labels : Bool, + taxlabels : Array[String], +) -> (Array[String], Array[String]) raise AlignNexusError { + if !labels && taxlabels.length() != ntax { + raise AlignNexusError("NEXUS LABELS=NO matrices require NTAX TAXLABELS") + } + let entry_ids : Array[String] = [] + let entry_chunks : Array[String] = [] + let mut line_index = 0 + while line_index < lines.length() { + if labels { + let parsed = align_nexus_parse_labeled_line(lines[line_index]) + let mut chunk = align_nexus_remove_whitespace(parsed.1) + line_index = line_index + 1 + if chunk.length() == 0 { + if line_index >= lines.length() { + raise AlignNexusError("Missing sequence after NEXUS matrix label") + } + chunk = align_nexus_remove_whitespace(lines[line_index]) + line_index = line_index + 1 + } + align_nexus_validate_id(parsed.0) + entry_ids.push(parsed.0) + entry_chunks.push(chunk) + } else { + entry_ids.push(taxlabels[entry_ids.length() % ntax]) + entry_chunks.push(align_nexus_remove_whitespace(lines[line_index])) + line_index = line_index + 1 + } + } + if entry_ids.length() == 0 || entry_ids.length() % ntax != 0 { + raise AlignNexusError( + "Interleaved NEXUS MATRIX does not contain complete NTAX blocks", + ) + } + let ids : Array[String] = [] + let rows = Array::make(ntax, "") + let block_count = entry_ids.length() / ntax + for block = 0; block < block_count; block = block + 1 { + let first_chunk_length = entry_chunks[block * ntax].length() + if first_chunk_length == 0 { + raise AlignNexusError("NEXUS interleaved matrix chunk must not be empty") + } + for row = 0; row < ntax; row = row + 1 { + let entry = block * ntax + row + if entry_chunks[entry].length() != first_chunk_length { + raise AlignNexusError( + "Rows in one NEXUS interleaved block must have equal widths", + ) + } + if block == 0 { + ids.push(entry_ids[entry]) + } else if !align_nexus_labels_equivalent(ids[row], entry_ids[entry]) { + raise AlignNexusError( + "NEXUS interleaved taxon order or label changed between blocks", + ) + } + rows[row] = rows[row] + entry_chunks[entry] + } + } + for row = 0; row < ntax; row = row + 1 { + if rows[row].length() != nchar { + raise AlignNexusError( + "NEXUS NCHAR does not match interleaved row length for '" + + ids[row] + + "'", + ) + } + } + (ids, rows) +} + +///| +fn align_nexus_parse_sequential_matrix( + lines : Array[String], + ntax : Int, + nchar : Int, + labels : Bool, + taxlabels : Array[String], +) -> (Array[String], Array[String]) raise AlignNexusError { + if !labels && taxlabels.length() != ntax { + raise AlignNexusError("NEXUS LABELS=NO matrices require NTAX TAXLABELS") + } + let ids : Array[String] = [] + let rows : Array[String] = [] + let mut line_index = 0 + for row = 0; row < ntax; row = row + 1 { + if line_index >= lines.length() { + raise AlignNexusError("Not enough taxa in NEXUS MATRIX") + } + let mut sequence = "" + if labels { + let parsed = align_nexus_parse_labeled_line(lines[line_index]) + align_nexus_validate_id(parsed.0) + ids.push(parsed.0) + sequence = align_nexus_remove_whitespace(parsed.1) + } else { + ids.push(taxlabels[row]) + sequence = align_nexus_remove_whitespace(lines[line_index]) + } + line_index = line_index + 1 + while sequence.length() < nchar { + if line_index >= lines.length() { + raise AlignNexusError("NEXUS MATRIX row is shorter than declared NCHAR") + } + sequence = sequence + align_nexus_remove_whitespace(lines[line_index]) + line_index = line_index + 1 + } + if sequence.length() != nchar { + raise AlignNexusError("NEXUS MATRIX row is longer than declared NCHAR") + } + rows.push(sequence) + } + if line_index != lines.length() { + raise AlignNexusError("Too many taxa or sequence lines in NEXUS MATRIX") + } + (ids, rows) +} + +///| +fn align_nexus_parse_labeled_line( + line : String, +) -> (String, String) raise AlignNexusError { + let mut index = 0 + while index < line.length() && + align_nexus_is_whitespace(line.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= line.length() { + raise AlignNexusError("Empty NEXUS MATRIX row") + } + let code = line.unsafe_get(index).to_int() + if code == '\''.to_int() || code == '"'.to_int() { + let quote = code + let label = StringBuilder::new() + index = index + 1 + let mut closed = false + while index < line.length() { + let current = line.unsafe_get(index).to_int() + if current == quote { + if index + 1 < line.length() && + line.unsafe_get(index + 1).to_int() == quote { + label.write_char(quote.unsafe_to_char()) + index = index + 2 + } else { + index = index + 1 + closed = true + break + } + } else { + label.write_char(current.unsafe_to_char()) + index = index + 1 + } + } + if !closed { + raise AlignNexusError("Unterminated quoted NEXUS taxon name") + } + if index < line.length() && + !align_nexus_is_whitespace(line.unsafe_get(index).to_int()) { + raise AlignNexusError( + "Quoted NEXUS taxon name must be followed by whitespace", + ) + } + (label.to_string(), line[index:].to_owned().trim().to_owned()) + } else { + let start = index + while index < line.length() && + !align_nexus_is_whitespace(line.unsafe_get(index).to_int()) { + index = index + 1 + } + (line[start:index].to_owned(), line[index:].to_owned().trim().to_owned()) + } +} + +///| +fn align_nexus_resolve_and_normalize_rows( + rows : Array[String], + format : AlignNexusFormat, +) -> Array[String] raise AlignNexusError { + let resolved : Array[String] = [] + for row_index = 0; row_index < rows.length(); row_index = row_index + 1 { + let output = StringBuilder::new(size_hint=rows[row_index].length()) + for column = 0; column < rows[row_index].length(); column = column + 1 { + let mut code = rows[row_index].unsafe_get(column).to_int() + match format.match_character { + Some(character) => + if code == character.unsafe_get(0).to_int() { + if row_index == 0 { + raise AlignNexusError( + "NEXUS match character is not allowed in the first row", + ) + } + code = resolved[0].unsafe_get(column).to_int() + } + None => () + } + output.write_char(code.unsafe_to_char()) + } + resolved.push( + align_nexus_normalize_row( + output.to_string(), + format.data_type, + format.missing_character, + format.gap_character, + format.symbols, + format.respect_case, + ), + ) + } + resolved +} + +///| +fn align_nexus_normalize_row( + row : String, + data_type : AlignNexusDataType, + missing_character : String, + gap_character : String, + symbols : String, + respect_case : Bool, +) -> String raise AlignNexusError { + let output = StringBuilder::new(size_hint=row.length()) + let missing_code = missing_character.unsafe_get(0).to_int() + let gap_code = gap_character.unsafe_get(0).to_int() + for index = 0; index < row.length(); index = index + 1 { + let code = row.unsafe_get(index).to_int() + if code == gap_code { + output.write_char('-') + } else if code == missing_code || + align_nexus_valid_residue(code, data_type, symbols, respect_case) { + output.write_char(code.unsafe_to_char()) + } else { + raise AlignNexusError( + "Illegal " + + data_type.code() + + " character '" + + code.unsafe_to_char().to_string() + + "' in NEXUS MATRIX", + ) + } + } + output.to_string() +} + +///| +fn align_nexus_valid_residue( + code : Int, + data_type : AlignNexusDataType, + symbols : String, + respect_case : Bool, +) -> Bool { + match data_type { + AlignNexusDna => + align_nexus_contains_code("ACGTRYSWKMBDHVNacgtryswkmbdhvn", code) + AlignNexusRna => + align_nexus_contains_code("ACGURYSWKMBDHVNacguryswkmbdhvn", code) + AlignNexusProtein => + align_nexus_contains_code( + "ACDEFGHIKLMNPQRSTVWYBZX*acdefghiklmnpqrstvywbzx", code, + ) + AlignNexusStandard => + if align_nexus_contains_code(symbols, code) { + true + } else if !respect_case { + align_nexus_contains_code(symbols, align_nexus_swap_case(code)) + } else { + false + } + } +} + +///| +fn align_nexus_validate_symbols( + data_type : AlignNexusDataType, + symbols : String, + respect_case : Bool, +) -> Unit raise AlignNexusError { + if data_type == AlignNexusStandard && symbols.length() == 0 { + raise AlignNexusError( + "NEXUS standard datatype requires a non-empty SYMBOLS value", + ) + } + let seen : Array[Int] = [] + for index = 0; index < symbols.length(); index = index + 1 { + let code = symbols.unsafe_get(index).to_int() + if code < 33 || + code > 126 || + align_nexus_is_whitespace(code) || + code == '\''.to_int() || + code == '"'.to_int() || + code == '['.to_int() || + code == ']'.to_int() || + code == '('.to_int() || + code == ')'.to_int() || + code == ','.to_int() || + code == ';'.to_int() || + code == '='.to_int() { + raise AlignNexusError( + "NEXUS SYMBOLS must contain distinct printable state characters", + ) + } + let comparable = if respect_case { + code + } else { + align_nexus_upper_code(code) + } + for previous in seen { + if previous == comparable { + raise AlignNexusError("NEXUS SYMBOLS contains a duplicate state") + } + } + seen.push(comparable) + } +} + +///| +fn align_nexus_validate_format_char( + value : String, + label : String, +) -> Unit raise AlignNexusError { + if value.length() != 1 { + raise AlignNexusError(label + " must be one ASCII character") + } + let code = value.unsafe_get(0).to_int() + if code < 33 || + code > 126 || + code == '['.to_int() || + code == ']'.to_int() || + code == '\''.to_int() || + code == '"'.to_int() || + code == ';'.to_int() { + raise AlignNexusError( + label + " must be one printable non-reserved character", + ) + } +} + +///| +fn align_nexus_validate_id(id : String) -> Unit raise AlignNexusError { + if id.length() == 0 { + raise AlignNexusError("NEXUS taxon identifier must not be empty") + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if code < 32 || + code == 127 || + code == '\n'.to_int() || + code == '\r'.to_int() { + raise AlignNexusError( + "NEXUS taxon identifier contains a control character", + ) + } + } +} + +///| +fn align_nexus_remove_all_gap_columns( + rows : Array[String], +) -> (Array[String], Int) { + if rows.length() == 0 { + return ([], 0) + } + let width = rows[0].length() + let keep = Array::make(width, false) + let mut removed = 0 + for column = 0; column < width; column = column + 1 { + for row in rows { + if row.unsafe_get(column).to_int() != '-'.to_int() { + keep[column] = true + } + } + if !keep[column] { + removed = removed + 1 + } + } + let result : Array[String] = [] + for row in rows { + let output = StringBuilder::new(size_hint=width - removed) + for column = 0; column < width; column = column + 1 { + if keep[column] { + output.write_char(row.unsafe_get(column).unsafe_to_char()) + } + } + result.push(output.to_string()) + } + (result, removed) +} + +///| +fn align_nexus_movement_changes( + alignment : AlignNexusAlignment, + first_column : Int, + second_column : Int, +) -> Bool { + for row = 0; row < alignment.sequences.length(); row = row + 1 { + let first = alignment.sequences[row].aligned_sequence + .unsafe_get(first_column) + .to_int() != + '-'.to_int() + let second = alignment.sequences[row].aligned_sequence + .unsafe_get(second_column) + .to_int() != + '-'.to_int() + if first != second { + return true + } + } + false +} + +///| +fn align_nexus_labels_equivalent(first : String, second : String) -> Bool { + if first == second { + return true + } + let left = StringBuilder::new(size_hint=first.length()) + let right = StringBuilder::new(size_hint=second.length()) + for index = 0; index < first.length(); index = index + 1 { + let code = first.unsafe_get(index).to_int() + left.write_char( + (if code == ' '.to_int() { '_'.to_int() } else { code }).unsafe_to_char(), + ) + } + for index = 0; index < second.length(); index = index + 1 { + let code = second.unsafe_get(index).to_int() + right.write_char( + (if code == ' '.to_int() { '_'.to_int() } else { code }).unsafe_to_char(), + ) + } + left.to_string() == right.to_string() +} + +///| +fn align_nexus_safe_name(id : String) -> String { + let mut quote = false + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + let safe = (code >= 'A'.to_int() && code <= 'Z'.to_int()) || + (code >= 'a'.to_int() && code <= 'z'.to_int()) || + (code >= '0'.to_int() && code <= '9'.to_int()) || + code == '_'.to_int() || + code == '-'.to_int() || + code == '.'.to_int() + if !safe { + quote = true + } + } + if !quote { + return id + } + let output = StringBuilder::new(size_hint=id.length() + 2) + output.write_char('\'') + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + output.write_char(code.unsafe_to_char()) + if code == '\''.to_int() { + output.write_char('\'') + } + } + output.write_char('\'') + output.to_string() +} + +///| +fn align_nexus_output_fragment( + row : String, + start : Int, + end : Int, + gap_character : String, +) -> String { + let output = StringBuilder::new(size_hint=end - start) + let gap_code = gap_character.unsafe_get(0).to_int() + for index = start; index < end; index = index + 1 { + let code = row.unsafe_get(index).to_int() + output.write_char( + (if code == '-'.to_int() { gap_code } else { code }).unsafe_to_char(), + ) + } + output.to_string() +} + +///| +fn align_nexus_escape_double_quotes(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + output.write_char(code.unsafe_to_char()) + if code == '"'.to_int() { + output.write_char('"') + } + } + output.to_string() +} + +///| +fn align_nexus_remove_gaps(sequence : String) -> String { + let output = StringBuilder::new(size_hint=sequence.length()) + for index = 0; index < sequence.length(); index = index + 1 { + let code = sequence.unsafe_get(index).to_int() + if code != '-'.to_int() { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_nexus_remove_whitespace(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if !align_nexus_is_whitespace(code) { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_nexus_normalize_lines(text : String) -> Array[String] { + let raw_lines = text.split("\n") + let lines : Array[String] = [] + for raw in raw_lines { + let owned = raw.to_owned() + if owned.length() > 0 && + owned.unsafe_get(owned.length() - 1).to_int() == '\r'.to_int() { + lines.push(owned[0:owned.length() - 1].to_owned()) + } else { + lines.push(owned) + } + } + lines +} + +///| +fn align_nexus_parse_positive_int( + text : String, + label : String, +) -> Int raise AlignNexusError { + if text.length() == 0 { + raise AlignNexusError(label + " is empty") + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignNexusError(label + " must be a positive integer") + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + raise AlignNexusError(label + " is outside the supported integer range") + } + value = value * 10 + digit + } + if value <= 0 { + raise AlignNexusError(label + " must be positive") + } + value +} + +///| +fn align_nexus_parse_yes_no( + value : String, + label : String, +) -> Bool raise AlignNexusError { + match value.to_lower() { + "yes" | "true" => true + "no" | "false" => false + _ => raise AlignNexusError(label + " must be YES or NO") + } +} + +///| +fn align_nexus_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() +} + +///| +fn align_nexus_contains_code(values : String, target : Int) -> Bool { + for index = 0; index < values.length(); index = index + 1 { + if values.unsafe_get(index).to_int() == target { + return true + } + } + false +} + +///| +fn align_nexus_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 'a'.to_int() + 'A'.to_int() + } else { + code + } +} + +///| +fn align_nexus_swap_case(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 'a'.to_int() + 'A'.to_int() + } else if code >= 'A'.to_int() && code <= 'Z'.to_int() { + code - 'A'.to_int() + 'a'.to_int() + } else { + code + } +} + +///| +fn align_nexus_pad_right(value : String, width : Int) -> String { + if value.length() >= width { + return value + } + let output = StringBuilder::new(size_hint=width) + output.write_string(value) + for _index = value.length(); _index < width; _index = _index + 1 { + output.write_char(' ') + } + output.to_string() +} + +///| +fn align_nexus_min(first : Int, second : Int) -> Int { + if first < second { + first + } else { + second + } +} diff --git a/src/align_phylip.mbt b/src/align_phylip.mbt new file mode 100644 index 00000000..061a3c77 --- /dev/null +++ b/src/align_phylip.mbt @@ -0,0 +1,1161 @@ +// Biopython-compatible Bio.Align.phylip support. +// +// This module is separate from phylip_io.mbt. The older module exposes the +// historical AlignIO-style MultipleSeqAlignment API, while this file models +// the modern coordinate-bearing Alignment representation and strict PHYLIP +// sequential/interleaved parsing semantics. + +///| +/// Error raised for malformed PHYLIP data or invalid alignment operations. +pub suberror AlignPhylipError { + AlignPhylipError(String) +} + +///| +/// Physical layout detected in or requested for a PHYLIP document. +pub(all) enum AlignPhylipLayout { + PhylipSequential + PhylipInterleaved +} derive(Eq, Debug) + +///| +/// One coordinate-bearing PHYLIP row. +pub struct AlignPhylipSequence { + id : String + sequence : String + aligned_sequence : String +} derive(Eq, Debug) + +///| +/// A modern PHYLIP multiple sequence alignment. +/// +/// PHYLIP identifiers read from files are the trimmed fixed-width 10-column +/// fields. Programmatically constructed rows may retain longer identifiers; +/// the writer applies Biopython's PHYLIP normalization when serializing them. +pub struct AlignPhylipAlignment { + sequences : Array[AlignPhylipSequence] + source_layout : AlignPhylipLayout +} derive(Eq, Debug) + +///| +/// Pairwise or all-pairs alignment statistics. +pub struct AlignPhylipCounts { + pairs : Int + aligned : Int + identities : Int + mismatches : Int + gap_columns : Int + double_gap_columns : Int + gap_opens : Int +} derive(Eq, Debug) + +///| +/// Construct one validated PHYLIP row from its printed aligned sequence. +pub fn AlignPhylipSequence::create( + id : String, + aligned_sequence : String, +) -> AlignPhylipSequence raise AlignPhylipError { + align_phylip_validate_id(id) + if aligned_sequence.length() == 0 { + raise AlignPhylipError("PHYLIP aligned sequence must not be empty") + } + align_phylip_validate_segment(aligned_sequence) + AlignPhylipSequence::{ + id, + sequence: align_phylip_remove_gaps(aligned_sequence), + aligned_sequence, + } +} + +///| +/// Construct a validated modern PHYLIP alignment. +pub fn AlignPhylipAlignment::create( + sequences : Array[AlignPhylipSequence], + source_layout? : AlignPhylipLayout = PhylipSequential, +) -> AlignPhylipAlignment raise AlignPhylipError { + let copied : Array[AlignPhylipSequence] = [] + for sequence in sequences { + copied.push(sequence) + } + let alignment = AlignPhylipAlignment::{ sequences: copied, source_layout } + align_phylip_validate_alignment(alignment) + alignment +} + +///| +/// Construct a coordinate-aware PHYLIP alignment from printed rows. +pub fn align_phylip_from_aligned( + ids : Array[String], + aligned_sequences : Array[String], + source_layout? : AlignPhylipLayout = PhylipSequential, +) -> AlignPhylipAlignment raise AlignPhylipError { + if ids.length() == 0 { + raise AlignPhylipError( + "PHYLIP alignment must contain at least one sequence", + ) + } + if ids.length() != aligned_sequences.length() { + raise AlignPhylipError( + "PHYLIP identifiers and aligned rows must have equal lengths", + ) + } + let sequences : Array[AlignPhylipSequence] = [] + for index = 0; index < ids.length(); index = index + 1 { + sequences.push( + AlignPhylipSequence::create(ids[index], aligned_sequences[index]), + ) + } + AlignPhylipAlignment::create(sequences, source_layout~) +} + +///| +/// Parse one strict PHYLIP alignment. +/// +/// The first line must contain the declared row and column counts. Identifiers +/// occupy exactly the first 10 columns of each first-block or first-row line. +/// As in Biopython 1.86, a blank line after the first `n` physical rows selects +/// interleaved parsing; otherwise wrapped sequential parsing is used. A file +/// containing exactly `n` complete rows is treated as sequential. +pub fn align_phylip_parse( + text : String, +) -> AlignPhylipAlignment raise AlignPhylipError { + let lines = align_phylip_lines(text) + if lines.length() == 0 || lines[0].trim().length() == 0 { + raise AlignPhylipError("Empty PHYLIP input") + } + let header = align_phylip_split_whitespace(lines[0]) + if header.length() != 2 { + raise AlignPhylipError("PHYLIP header must contain exactly two integers") + } + let number_of_sequences = align_phylip_parse_positive_int( + header[0], + "PHYLIP sequence count", + ) + let number_of_columns = align_phylip_parse_positive_int( + header[1], + "PHYLIP column count", + ) + let mut last = lines.length() - 1 + while last > 0 && lines[last].trim().length() == 0 { + last = last - 1 + } + if last < number_of_sequences { + raise AlignPhylipError( + "PHYLIP input has fewer rows than declared in its header", + ) + } + for index = 1; index <= number_of_sequences; index = index + 1 { + if lines[index].trim().length() == 0 { + raise AlignPhylipError( + "PHYLIP first block contains an empty sequence row", + ) + } + } + let marker = number_of_sequences + 1 + let layout = if marker > last { + PhylipSequential + } else if lines[marker].trim().length() == 0 { + PhylipInterleaved + } else { + PhylipSequential + } + let sequences = match layout { + PhylipSequential => + align_phylip_parse_sequential( + lines, last, number_of_sequences, number_of_columns, + ) + PhylipInterleaved => + align_phylip_parse_interleaved( + lines, last, number_of_sequences, number_of_columns, + ) + } + AlignPhylipAlignment::create(sequences, source_layout=layout) +} + +///| +/// Serialize a PHYLIP alignment. +/// +/// Defaults match Biopython's modern writer: sequential rows, no wrapping, +/// fixed 10-column normalized identifiers, and no grouping whitespace. +/// `block_width` and `group_width` enable canonical wrapped or interleaved +/// presentation without changing alignment coordinates. +pub fn align_phylip_write( + alignment : AlignPhylipAlignment, + layout? : AlignPhylipLayout = PhylipSequential, + block_width? : Int = 0, + group_width? : Int = 0, +) -> String raise AlignPhylipError { + align_phylip_validate_alignment(alignment) + if block_width < 0 { + raise AlignPhylipError("PHYLIP block width must be non-negative") + } + if group_width < 0 { + raise AlignPhylipError("PHYLIP group width must be non-negative") + } + let width = alignment.alignment_length() + let effective_block_width = if block_width == 0 { + if layout == PhylipInterleaved { + 60 + } else { + width + } + } else { + block_width + } + if effective_block_width <= 0 { + raise AlignPhylipError("PHYLIP block width must be positive") + } + let names : Array[String] = [] + for sequence in alignment.sequences { + names.push(align_phylip_normalize_id(sequence.id)) + } + let output = StringBuilder::new() + output.write_string(alignment.num_sequences().to_string()) + output.write_char(' ') + output.write_string(width.to_string()) + output.write_char('\n') + match layout { + PhylipSequential => + for row = 0; row < alignment.num_sequences(); row = row + 1 { + let aligned = alignment.sequences[row].aligned_sequence + let mut start = 0 + while start < width { + let stop = if start + effective_block_width < width { + start + effective_block_width + } else { + width + } + if start == 0 { + align_phylip_write_name(output, names[row]) + } else { + output.write_string(" ".repeat(10)) + } + align_phylip_write_grouped( + output, + aligned[start:stop].to_owned(), + group_width, + ) + output.write_char('\n') + start = stop + } + } + PhylipInterleaved => { + let mut start = 0 + while start < width { + let stop = if start + effective_block_width < width { + start + effective_block_width + } else { + width + } + if start > 0 { + output.write_char('\n') + } + for row = 0; row < alignment.num_sequences(); row = row + 1 { + if start == 0 { + align_phylip_write_name(output, names[row]) + } else { + output.write_string(" ".repeat(10)) + } + align_phylip_write_grouped( + output, + alignment.sequences[row].aligned_sequence[start:stop].to_owned(), + group_width, + ) + output.write_char('\n') + } + start = stop + } + } + } + let result = output.to_string() + let reparsed = align_phylip_parse(result) + if reparsed.num_sequences() != alignment.num_sequences() || + reparsed.alignment_length() != alignment.alignment_length() { + raise AlignPhylipError("PHYLIP writer self-validation failed") + } + for row = 0; row < alignment.num_sequences(); row = row + 1 { + if reparsed.sequences[row].id != names[row] || + reparsed.sequences[row].aligned_sequence != + alignment.sequences[row].aligned_sequence { + raise AlignPhylipError("PHYLIP writer self-validation failed") + } + } + result +} + +///| +/// Apply Biopython's strict PHYLIP identifier normalization. +/// +/// Outer whitespace is removed, `[](),` are deleted, `:;` become `|`, and +/// the result is truncated to the fixed 10-column identifier field. +pub fn align_phylip_normalize_id(id : String) -> String raise AlignPhylipError { + align_phylip_validate_id(id) + let trimmed = id.trim().to_owned() + let output = StringBuilder::new(size_hint=trimmed.length()) + for index = 0; index < trimmed.length(); index = index + 1 { + let code = trimmed.unsafe_get(index).to_int() + if code == '['.to_int() || + code == ']'.to_int() || + code == '('.to_int() || + code == ')'.to_int() || + code == ','.to_int() { + continue + } + if code == ':'.to_int() || code == ';'.to_int() { + output.write_char('|') + } else { + output.write_char(code.unsafe_to_char()) + } + } + let normalized = output.to_string() + if normalized.length() > 10 { + normalized[0:10].to_owned() + } else { + normalized + } +} + +///| +/// Return serialized identifiers in row order. +pub fn AlignPhylipAlignment::normalized_ids( + self : AlignPhylipAlignment, +) -> Array[String] { + let result : Array[String] = [] + for sequence in self.sequences { + result.push( + align_phylip_normalize_id(sequence.id) catch { + AlignPhylipError(_) => "" + }, + ) + } + result +} + +///| +/// Return whether PHYLIP normalization maps two rows to the same identifier. +pub fn AlignPhylipAlignment::has_normalized_id_collisions( + self : AlignPhylipAlignment, +) -> Bool { + let names = self.normalized_ids() + for first = 0; first < names.length(); first = first + 1 { + for second = first + 1; second < names.length(); second = second + 1 { + if names[first] == names[second] { + return true + } + } + } + false +} + +///| +/// Return the number of sequence rows. +pub fn AlignPhylipAlignment::num_sequences(self : AlignPhylipAlignment) -> Int { + self.sequences.length() +} + +///| +/// Return the alignment width including gaps. +pub fn AlignPhylipAlignment::alignment_length( + self : AlignPhylipAlignment, +) -> Int { + if self.sequences.length() == 0 { + 0 + } else { + self.sequences[0].aligned_sequence.length() + } +} + +///| +/// Locate every row with an exact identifier. Duplicate strict identifiers +/// are legal in PHYLIP and therefore all matches are returned. +pub fn AlignPhylipAlignment::find_sequences( + self : AlignPhylipAlignment, + id : String, +) -> Array[Int] { + let result : Array[Int] = [] + for index = 0; index < self.sequences.length(); index = index + 1 { + if self.sequences[index].id == id { + result.push(index) + } + } + result +} + +///| +/// Return one printed alignment column. +pub fn AlignPhylipAlignment::column( + self : AlignPhylipAlignment, + column : Int, +) -> String? { + if column < 0 || column >= self.alignment_length() { + return None + } + let result = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + result.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(result.to_string()) +} + +///| +/// Map a zero-based ungapped row position to an alignment column. +pub fn AlignPhylipAlignment::sequence_position_to_column( + self : AlignPhylipAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + position < 0 || + position >= self.sequences[row].sequence.length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + let mut coordinate = 0 + for column = 0; column < aligned.length(); column = column + 1 { + if aligned.unsafe_get(column).to_int() != '-'.to_int() { + if coordinate == position { + return Some(column) + } + coordinate = coordinate + 1 + } + } + None +} + +///| +/// Map one alignment column to a zero-based ungapped row position. +pub fn AlignPhylipAlignment::column_to_sequence_position( + self : AlignPhylipAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.alignment_length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + if aligned.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut coordinate = 0 + for index = 0; index < column; index = index + 1 { + if aligned.unsafe_get(index).to_int() != '-'.to_int() { + coordinate = coordinate + 1 + } + } + Some(coordinate) +} + +///| +/// Project one row position through the alignment to another row. +pub fn AlignPhylipAlignment::map_position( + self : AlignPhylipAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column residue positions for two rows. +pub fn AlignPhylipAlignment::aligned_pairs( + self : AlignPhylipAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise AlignPhylipError { + align_phylip_validate_row(self, first_row) + align_phylip_validate_row(self, second_row) + let result : Array[(Int?, Int?)] = [] + let mut first_position = 0 + let mut second_position = 0 + for column = 0; column < self.alignment_length(); column = column + 1 { + let first_gap = self.sequences[first_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + let second_gap = self.sequences[second_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + result.push( + ( + if first_gap { + None + } else { + Some(first_position) + }, + if second_gap { + None + } else { + Some(second_position) + }, + ), + ) + if !first_gap { + first_position = first_position + 1 + } + if !second_gap { + second_position = second_position + 1 + } + } + result +} + +///| +/// Return a compact Biopython-style coordinate path for all rows. +pub fn AlignPhylipAlignment::coordinate_path( + self : AlignPhylipAlignment, +) -> Array[Array[Int]] { + let paths : Array[Array[Int]] = [] + let coordinates = Array::make(self.sequences.length(), 0) + for _row = 0; _row < self.sequences.length(); _row = _row + 1 { + paths.push([0]) + } + let width = self.alignment_length() + for column = 0; column < width; column = column + 1 { + for row = 0; row < self.sequences.length(); row = row + 1 { + if self.sequences[row].aligned_sequence.unsafe_get(column).to_int() != + '-'.to_int() { + coordinates[row] = coordinates[row] + 1 + } + } + let boundary = if column + 1 == width { + true + } else { + align_phylip_movement_changes(self, column, column + 1) + } + if boundary { + for row = 0; row < self.sequences.length(); row = row + 1 { + paths[row].push(coordinates[row]) + } + } + } + paths +} + +///| +/// Compute statistics for one pair of rows. +pub fn AlignPhylipAlignment::pair_counts( + self : AlignPhylipAlignment, + first_row : Int, + second_row : Int, +) -> AlignPhylipCounts raise AlignPhylipError { + align_phylip_validate_row(self, first_row) + align_phylip_validate_row(self, second_row) + align_phylip_count_pair(self, first_row, second_row) +} + +///| +/// Aggregate statistics across every unordered pair. +pub fn AlignPhylipAlignment::counts( + self : AlignPhylipAlignment, +) -> AlignPhylipCounts { + let mut pairs = 0 + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + for first = 0; first < self.sequences.length(); first = first + 1 { + for second = first + 1 + second < self.sequences.length() + second = second + 1 { + let counts = align_phylip_count_pair(self, first, second) + pairs = pairs + 1 + aligned = aligned + counts.aligned + identities = identities + counts.identities + mismatches = mismatches + counts.mismatches + gap_columns = gap_columns + counts.gap_columns + double_gap_columns = double_gap_columns + counts.double_gap_columns + gap_opens = gap_opens + counts.gap_opens + } + } + AlignPhylipCounts::{ + pairs, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +/// Return exact identity over columns containing residues in both rows. +pub fn AlignPhylipCounts::identity(self : AlignPhylipCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Return per-column non-gap occupancy. +pub fn AlignPhylipAlignment::occupancy( + self : AlignPhylipAlignment, +) -> Array[Double] { + let result : Array[Double] = [] + for column = 0; column < self.alignment_length(); column = column + 1 { + let mut residues = 0 + for sequence in self.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + result.push(residues.to_double() / self.sequences.length().to_double()) + } + result +} + +///| +/// Calculate a majority residue consensus. Gaps do not vote. +pub fn AlignPhylipAlignment::majority_consensus( + self : AlignPhylipAlignment, + minimum_fraction? : Double = 0.0, +) -> String raise AlignPhylipError { + if minimum_fraction != minimum_fraction || + minimum_fraction < 0.0 || + minimum_fraction > 1.0 { + raise AlignPhylipError( + "PHYLIP consensus minimum fraction must be between 0 and 1", + ) + } + let output = StringBuilder::new(size_hint=self.alignment_length()) + for column = 0; column < self.alignment_length(); column = column + 1 { + let frequencies = Array::make(128, 0) + let mut residues = 0 + for sequence in self.sequences { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if code != '-'.to_int() { + let upper = align_phylip_upper_code(code) + if upper >= 0 && upper < frequencies.length() { + frequencies[upper] = frequencies[upper] + 1 + } + residues = residues + 1 + } + } + if residues == 0 { + output.write_char('-') + continue + } + let mut best_code = 0 + let mut best_count = -1 + for code = 0; code < frequencies.length(); code = code + 1 { + if frequencies[code] > best_count { + best_code = code + best_count = frequencies[code] + } + } + if best_count.to_double() / residues.to_double() < minimum_fraction { + output.write_char('X') + } else { + output.write_char(best_code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +/// Return a compact alignment summary. +pub fn AlignPhylipAlignment::summary(self : AlignPhylipAlignment) -> String { + "AlignPhylipAlignment(layout=" + + (match self.source_layout { + PhylipSequential => "sequential" + PhylipInterleaved => "interleaved" + }) + + ", sequences=" + + self.num_sequences().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ", normalized_collisions=" + + self.has_normalized_id_collisions().to_string() + + ")" +} + +///| +/// Return a compact official-style interleaved PHYLIP fixture. +pub fn align_phylip_example_text() -> String { + "3 12\n" + + "Reference MKT--AIL\n" + + "Query_one M-TQQAIL\n" + + "Query_two MKT--AVL\n" + + "\n" + + " GHAA\n" + + " G-AA\n" + + " GHAA\n" +} + +///| +fn align_phylip_parse_sequential( + lines : Array[String], + last : Int, + number_of_sequences : Int, + number_of_columns : Int, +) -> Array[AlignPhylipSequence] raise AlignPhylipError { + let result : Array[AlignPhylipSequence] = [] + let mut line_index = 1 + for _row = 0; _row < number_of_sequences; _row = _row + 1 { + if line_index > last || lines[line_index].trim().length() == 0 { + raise AlignPhylipError( + "PHYLIP sequential data ended before all rows were read", + ) + } + let (id, first_segment) = align_phylip_parse_named_row(lines[line_index]) + line_index = line_index + 1 + let output = StringBuilder::new(size_hint=number_of_columns) + output.write_string(first_segment) + let mut length = first_segment.length() + if length > number_of_columns { + raise AlignPhylipError( + "PHYLIP sequential row exceeds the declared column count", + ) + } + while length < number_of_columns { + if line_index > last || lines[line_index].trim().length() == 0 { + raise AlignPhylipError( + "PHYLIP sequential row ended before its declared length", + ) + } + let segment = align_phylip_compact_segment(lines[line_index]) + align_phylip_validate_segment(segment) + if length + segment.length() > number_of_columns { + raise AlignPhylipError( + "PHYLIP sequential row exceeds the declared column count", + ) + } + output.write_string(segment) + length = length + segment.length() + line_index = line_index + 1 + } + result.push(AlignPhylipSequence::create(id, output.to_string())) + } + while line_index <= last { + if lines[line_index].trim().length() != 0 { + raise AlignPhylipError( + "PHYLIP input contains trailing non-whitespace data", + ) + } + line_index = line_index + 1 + } + result +} + +///| +fn align_phylip_parse_interleaved( + lines : Array[String], + last : Int, + number_of_sequences : Int, + number_of_columns : Int, +) -> Array[AlignPhylipSequence] raise AlignPhylipError { + let ids : Array[String] = [] + let builders : Array[StringBuilder] = [] + let lengths = Array::make(number_of_sequences, 0) + let mut first_width = -1 + for row = 0; row < number_of_sequences; row = row + 1 { + let (id, segment) = align_phylip_parse_named_row(lines[row + 1]) + if first_width < 0 { + first_width = segment.length() + } else if segment.length() != first_width { + raise AlignPhylipError( + "PHYLIP rows in an interleaved block must have equal widths", + ) + } + if segment.length() > number_of_columns { + raise AlignPhylipError( + "PHYLIP interleaved row exceeds the declared column count", + ) + } + ids.push(id) + let output = StringBuilder::new(size_hint=number_of_columns) + output.write_string(segment) + builders.push(output) + lengths[row] = segment.length() + } + let mut line_index = number_of_sequences + 1 + while lengths[0] < number_of_columns { + while line_index <= last && lines[line_index].trim().length() == 0 { + line_index = line_index + 1 + } + if line_index > last { + raise AlignPhylipError( + "PHYLIP interleaved data ended before its declared length", + ) + } + let segments : Array[String] = [] + let mut block_width = -1 + for _row = 0; _row < number_of_sequences; _row = _row + 1 { + if line_index > last || lines[line_index].trim().length() == 0 { + raise AlignPhylipError( + "PHYLIP interleaved block has fewer rows than declared", + ) + } + let segment = align_phylip_compact_segment(lines[line_index]) + align_phylip_validate_segment(segment) + if block_width < 0 { + block_width = segment.length() + } else if segment.length() != block_width { + raise AlignPhylipError( + "PHYLIP rows in an interleaved block must have equal widths", + ) + } + segments.push(segment) + line_index = line_index + 1 + } + for row = 0; row < number_of_sequences; row = row + 1 { + if lengths[row] + segments[row].length() > number_of_columns { + raise AlignPhylipError( + "PHYLIP interleaved row exceeds the declared column count", + ) + } + builders[row].write_string(segments[row]) + lengths[row] = lengths[row] + segments[row].length() + } + } + for row = 0; row < number_of_sequences; row = row + 1 { + if lengths[row] != number_of_columns { + raise AlignPhylipError( + "PHYLIP row length differs from the declared column count", + ) + } + } + while line_index <= last { + if lines[line_index].trim().length() != 0 { + raise AlignPhylipError( + "PHYLIP input contains trailing non-whitespace data", + ) + } + line_index = line_index + 1 + } + let result : Array[AlignPhylipSequence] = [] + for row = 0; row < number_of_sequences; row = row + 1 { + result.push( + AlignPhylipSequence::create(ids[row], builders[row].to_string()), + ) + } + result +} + +///| +fn align_phylip_parse_named_row( + line : String, +) -> (String, String) raise AlignPhylipError { + if line.length() < 10 { + raise AlignPhylipError( + "PHYLIP sequence row is shorter than the 10-column identifier field", + ) + } + let id = line[0:10].trim().to_owned() + align_phylip_validate_id(id) + let segment = align_phylip_compact_segment(line[10:line.length()].to_owned()) + align_phylip_validate_segment(segment) + (id, segment) +} + +///| +fn align_phylip_validate_alignment( + alignment : AlignPhylipAlignment, +) -> Unit raise AlignPhylipError { + if alignment.sequences.length() == 0 { + raise AlignPhylipError( + "PHYLIP alignment must contain at least one sequence", + ) + } + let width = alignment.sequences[0].aligned_sequence.length() + if width == 0 { + raise AlignPhylipError("PHYLIP alignment must contain at least one column") + } + for sequence in alignment.sequences { + align_phylip_validate_id(sequence.id) + if sequence.aligned_sequence.length() != width { + raise AlignPhylipError("PHYLIP aligned rows must have equal widths") + } + align_phylip_validate_segment(sequence.aligned_sequence) + if align_phylip_remove_gaps(sequence.aligned_sequence) != sequence.sequence { + raise AlignPhylipError( + "PHYLIP sequence row has inconsistent derived fields", + ) + } + } + for column = 0; column < width; column = column + 1 { + let mut residues = 0 + for sequence in alignment.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + if residues == 0 { + raise AlignPhylipError( + "PHYLIP alignment contains an all-gap column at " + column.to_string(), + ) + } + } +} + +///| +fn align_phylip_count_pair( + alignment : AlignPhylipAlignment, + first_row : Int, + second_row : Int, +) -> AlignPhylipCounts { + let first = alignment.sequences[first_row].aligned_sequence + let second = alignment.sequences[second_row].aligned_sequence + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + let mut in_gap = false + for column = 0; column < alignment.alignment_length(); column = column + 1 { + let first_code = first.unsafe_get(column).to_int() + let second_code = second.unsafe_get(column).to_int() + let first_gap = first_code == '-'.to_int() + let second_gap = second_code == '-'.to_int() + if first_gap && second_gap { + double_gap_columns = double_gap_columns + 1 + in_gap = false + } else if first_gap || second_gap { + gap_columns = gap_columns + 1 + if !in_gap { + gap_opens = gap_opens + 1 + } + in_gap = true + } else { + aligned = aligned + 1 + if first_code == second_code { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + in_gap = false + } + } + AlignPhylipCounts::{ + pairs: 1, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +fn align_phylip_movement_changes( + alignment : AlignPhylipAlignment, + first_column : Int, + second_column : Int, +) -> Bool { + for sequence in alignment.sequences { + let first_moves = sequence.aligned_sequence + .unsafe_get(first_column) + .to_int() != + '-'.to_int() + let second_moves = sequence.aligned_sequence + .unsafe_get(second_column) + .to_int() != + '-'.to_int() + if first_moves != second_moves { + return true + } + } + false +} + +///| +fn align_phylip_write_name(output : StringBuilder, name : String) -> Unit { + output.write_string(name) + if name.length() < 10 { + output.write_string(" ".repeat(10 - name.length())) + } +} + +///| +fn align_phylip_write_grouped( + output : StringBuilder, + segment : String, + group_width : Int, +) -> Unit { + if group_width == 0 || group_width >= segment.length() { + output.write_string(segment) + return + } + let mut start = 0 + while start < segment.length() { + let stop = if start + group_width < segment.length() { + start + group_width + } else { + segment.length() + } + if start > 0 { + output.write_char(' ') + } + output.write_string(segment[start:stop].to_owned()) + start = stop + } +} + +///| +fn align_phylip_validate_id(id : String) -> Unit raise AlignPhylipError { + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if code < 32 || code > 126 || code == '\n'.to_int() || code == '\r'.to_int() { + raise AlignPhylipError( + "PHYLIP sequence identifier contains a control or non-ASCII character", + ) + } + } +} + +///| +fn align_phylip_validate_segment( + segment : String, +) -> Unit raise AlignPhylipError { + if segment.length() == 0 { + raise AlignPhylipError("PHYLIP sequence segment must not be empty") + } + for index = 0; index < segment.length(); index = index + 1 { + let code = segment.unsafe_get(index).to_int() + if code == '.'.to_int() { + raise AlignPhylipError("PHYLIP format no longer allows dots in sequences") + } + if code <= 32 || code > 126 { + raise AlignPhylipError( + "PHYLIP sequence contains whitespace or a non-ASCII character", + ) + } + } +} + +///| +fn align_phylip_validate_row( + alignment : AlignPhylipAlignment, + row : Int, +) -> Unit raise AlignPhylipError { + if row < 0 || row >= alignment.sequences.length() { + raise AlignPhylipError("PHYLIP row index is out of bounds") + } +} + +///| +fn align_phylip_compact_segment(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code != ' '.to_int() && code != '\t'.to_int() { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_phylip_remove_gaps(value : String) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code != '-'.to_int() { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_phylip_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 'a'.to_int() + 'A'.to_int() + } else { + code + } +} + +///| +fn align_phylip_lines(text : String) -> Array[String] { + let raw = text.split("\n").to_array() + let lines : Array[String] = [] + for line in raw { + if line.length() > 0 && + line.unsafe_get(line.length() - 1).to_int() == '\r'.to_int() { + lines.push(line[0:line.length() - 1].to_owned()) + } else { + lines.push(line.to_owned()) + } + } + lines +} + +///| +fn align_phylip_split_whitespace(value : String) -> Array[String] { + let fields : Array[String] = [] + let mut start = 0 + let mut in_field = false + for index = 0; index < value.length(); index = index + 1 { + let whitespace = align_phylip_is_whitespace( + value.unsafe_get(index).to_int(), + ) + if whitespace { + if in_field { + fields.push(value[start:index].to_owned()) + in_field = false + } + } else if !in_field { + start = index + in_field = true + } + } + if in_field { + fields.push(value[start:value.length()].to_owned()) + } + fields +} + +///| +fn align_phylip_parse_positive_int( + text : String, + label : String, +) -> Int raise AlignPhylipError { + if text.length() == 0 { + raise AlignPhylipError(label + " is empty") + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignPhylipError(label + " must be a positive integer") + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + raise AlignPhylipError(label + " is outside the supported integer range") + } + value = value * 10 + digit + } + if value <= 0 { + raise AlignPhylipError(label + " must be positive") + } + value +} + +///| +fn align_phylip_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\r'.to_int() || + code == '\n'.to_int() +} diff --git a/src/align_psl.mbt b/src/align_psl.mbt new file mode 100644 index 00000000..20978e8e --- /dev/null +++ b/src/align_psl.mbt @@ -0,0 +1,1509 @@ +// Alignment-aware UCSC PSL/PSLX support. +// +// This module follows Biopython Bio.Align.psl coordinate semantics. PSL uses +// zero-based half-open intervals. Nucleotide alignments may reverse the query +// axis; translated DNA-to-protein alignments may reverse the target axis. + +///| +pub suberror PslError { + PslError(String) +} + +///| +pub(all) enum PslSequenceType { + PslNucleotide + PslProtein +} derive(Eq, Debug) + +///| +pub(all) enum PslMaskMode { + PslNoMask + PslMaskLower + PslMaskUpper +} derive(Eq, Debug) + +///| +pub struct PslBlock { + target_start : Int + target_end : Int + query_start : Int + query_end : Int + target_size : Int + query_size : Int + query_sequence : String + target_sequence : String +} derive(Eq, Debug) + +///| +pub struct PslCounts { + aligned_query_units : Int + query_insert_count : Int + query_insert_bases : Int + target_insert_count : Int + target_insert_bases : Int + block_count : Int +} derive(Eq, Debug) + +///| +pub struct PslAlignment { + target_name : String + target_size : Int + query_name : String + query_size : Int + target_coordinates : Array[Int] + query_coordinates : Array[Int] + sequence_type : PslSequenceType + matches : Int + mismatches : Int + repeat_matches : Int + n_count : Int + target_sequence : String + query_sequence : String + query_block_sequences : Array[String] + target_block_sequences : Array[String] +} + +///| +pub struct PslDocument { + version : String + has_header : Bool + alignments : Array[PslAlignment] +} + +///| +pub struct PslWriteConfig { + header : Bool + pslx : Bool + recount : Bool + mask : PslMaskMode + wildcard : Char +} + +///| +pub struct PslSummary { + alignment_count : Int + pslx_count : Int + nucleotide_count : Int + protein_count : Int + aligned_query_units : Int + query_insert_bases : Int + target_insert_bases : Int +} + +///| +priv struct PslStorage { + strand : String + query_start : Int + query_end : Int + target_start : Int + target_end : Int + block_sizes : Array[Int] + query_starts : Array[Int] + target_starts : Array[Int] +} + +///| +fn psl_fail(message : String) -> Unit raise PslError { + raise PslError(message) +} + +///| +fn psl_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn psl_min(left : Int, right : Int) -> Int { + if left < right { + left + } else { + right + } +} + +///| +fn psl_max(left : Int, right : Int) -> Int { + if left > right { + left + } else { + right + } +} + +///| +fn psl_copy_ints(values : Array[Int]) -> Array[Int] { + let copy : Array[Int] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn psl_copy_strings(values : Array[String]) -> Array[String] { + let copy : Array[String] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn psl_reverse_ints(values : Array[Int]) -> Array[Int] { + let reversed : Array[Int] = [] + let mut index = values.length() + while index > 0 { + index = index - 1 + reversed.push(values[index]) + } + reversed +} + +///| +fn psl_strip_cr(value : String) -> String { + if value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '\r'.to_int() { + value[0:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn psl_validate_text( + value : String, + label : String, + allow_empty : Bool, +) -> Unit raise PslError { + if !allow_empty && value.length() == 0 { + psl_fail(label + " must not be empty") + } + for index in 0.. Unit raise PslError { + if sequence.length() == 0 { + return + } + if sequence.length() != expected_size { + psl_fail(label + " length does not match its declaration") + } + for index in 0.. Unit raise PslError { + if sequence.length() != expected_size { + psl_fail(label + " length does not match its block size") + } + for index in 0.. Int raise PslError { + let mut direction = 0 + for index in 0..<(coordinates.length() - 1) { + let step = coordinates[index + 1] - coordinates[index] + if step > 0 { + if direction < 0 { + psl_fail(label + " coordinates change direction") + } + direction = 1 + } else if step < 0 { + if direction > 0 { + psl_fail(label + " coordinates change direction") + } + direction = -1 + } + } + direction +} + +///| +fn psl_raw_blocks( + target_coordinates : Array[Int], + query_coordinates : Array[Int], + query_sequences : Array[String], + target_sequences : Array[String], +) -> Array[PslBlock] { + let blocks : Array[PslBlock] = [] + let mut block_index = 0 + for index in 0..<(target_coordinates.length() - 1) { + let target_start = target_coordinates[index] + let target_end = target_coordinates[index + 1] + let query_start = query_coordinates[index] + let query_end = query_coordinates[index + 1] + let target_size = psl_abs(target_end - target_start) + let query_size = psl_abs(query_end - query_start) + if target_size > 0 && query_size > 0 { + blocks.push(PslBlock::{ + target_start, + target_end, + query_start, + query_end, + target_size, + query_size, + query_sequence: if query_sequences.length() > 0 { + query_sequences[block_index] + } else { + "" + }, + target_sequence: if target_sequences.length() > 0 { + target_sequences[block_index] + } else { + "" + }, + }) + block_index = block_index + 1 + } + } + blocks +} + +///| +fn psl_path_counts( + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> PslCounts { + let mut aligned_query_units = 0 + let mut query_insert_count = 0 + let mut query_insert_bases = 0 + let mut target_insert_count = 0 + let mut target_insert_bases = 0 + let mut block_count = 0 + for index in 0..<(target_coordinates.length() - 1) { + let target_step = psl_abs( + target_coordinates[index + 1] - target_coordinates[index], + ) + let query_step = psl_abs( + query_coordinates[index + 1] - query_coordinates[index], + ) + if target_step > 0 && query_step > 0 { + aligned_query_units = aligned_query_units + query_step + block_count = block_count + 1 + } else if target_step == 0 && query_step > 0 { + query_insert_count = query_insert_count + 1 + query_insert_bases = query_insert_bases + query_step + } else if target_step > 0 && query_step == 0 { + target_insert_count = target_insert_count + 1 + target_insert_bases = target_insert_bases + target_step + } + } + PslCounts::{ + aligned_query_units, + query_insert_count, + query_insert_bases, + target_insert_count, + target_insert_bases, + block_count, + } +} + +///| +pub fn PslAlignment::create( + target_name : String, + target_size : Int, + query_name : String, + query_size : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], + sequence_type? : PslSequenceType = PslNucleotide, + matches? : Int = -1, + mismatches? : Int = -1, + repeat_matches? : Int = -1, + n_count? : Int = -1, + target_sequence? : String = "", + query_sequence? : String = "", + query_block_sequences? : Array[String] = [], + target_block_sequences? : Array[String] = [], +) -> PslAlignment raise PslError { + psl_validate_text(target_name, "PSL target name", false) + psl_validate_text(query_name, "PSL query name", false) + if target_size <= 0 || query_size <= 0 { + psl_fail("PSL target and query sizes must be positive") + } + if target_coordinates.length() != query_coordinates.length() { + psl_fail("PSL target and query coordinate arrays must have equal length") + } + if target_coordinates.length() < 2 { + psl_fail("PSL alignment path must contain at least two points") + } + let mut targets = psl_copy_ints(target_coordinates) + let mut queries = psl_copy_ints(query_coordinates) + let initial_target_direction = psl_axis_direction(targets, "PSL target") + let initial_query_direction = psl_axis_direction(queries, "PSL query") + match sequence_type { + PslNucleotide => + if initial_target_direction < 0 { + targets = psl_reverse_ints(targets) + queries = psl_reverse_ints(queries) + } + PslProtein => + if initial_query_direction < 0 { + targets = psl_reverse_ints(targets) + queries = psl_reverse_ints(queries) + } + } + let target_direction = psl_axis_direction(targets, "PSL target") + let query_direction = psl_axis_direction(queries, "PSL query") + match sequence_type { + PslNucleotide => + if target_direction <= 0 || query_direction == 0 { + psl_fail( + "nucleotide PSL requires an increasing target and an oriented query", + ) + } + PslProtein => + if query_direction <= 0 || target_direction == 0 { + psl_fail( + "translated PSL requires an increasing query and an oriented target", + ) + } + } + for index in 0.. target_size { + psl_fail("PSL target coordinate is out of bounds") + } + if queries[index] < 0 || queries[index] > query_size { + psl_fail("PSL query coordinate is out of bounds") + } + } + let mut block_count = 0 + for index in 0..<(targets.length() - 1) { + let target_step = psl_abs(targets[index + 1] - targets[index]) + let query_step = psl_abs(queries[index + 1] - queries[index]) + if target_step == 0 && query_step == 0 { + psl_fail("PSL path contains a zero-length step") + } + if target_step > 0 && query_step > 0 { + match sequence_type { + PslNucleotide => + if target_step != query_step { + psl_fail("nucleotide PSL aligned steps must have equal lengths") + } + PslProtein => + if target_step != 3 * query_step { + psl_fail("translated PSL aligned steps must have a 3:1 ratio") + } + } + block_count = block_count + 1 + } + } + if block_count == 0 { + psl_fail("PSL alignment must contain at least one aligned block") + } + let first_target_step = psl_abs(targets[1] - targets[0]) + let first_query_step = psl_abs(queries[1] - queries[0]) + let last = targets.length() - 1 + let last_target_step = psl_abs(targets[last] - targets[last - 1]) + let last_query_step = psl_abs(queries[last] - queries[last - 1]) + if first_target_step == 0 || + first_query_step == 0 || + last_target_step == 0 || + last_query_step == 0 { + psl_fail("PSL path cannot begin or end with an unaligned gap") + } + psl_validate_sequence(target_sequence, target_size, "PSL target sequence") + psl_validate_sequence(query_sequence, query_size, "PSL query sequence") + let has_query_blocks = query_block_sequences.length() > 0 + let has_target_blocks = target_block_sequences.length() > 0 + if has_query_blocks != has_target_blocks { + psl_fail("PSLX requires both query and target block sequences") + } + if has_query_blocks { + if query_block_sequences.length() != block_count || + target_block_sequences.length() != block_count { + psl_fail("PSLX block sequence count does not match blockCount") + } + let blocks = psl_raw_blocks(targets, queries, [], []) + for index in 0.. PslWriteConfig raise PslError { + let code = wildcard.to_int() + if code == 0 || + code == ','.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() { + psl_fail("PSL wildcard must be a printable non-delimiter character") + } + PslWriteConfig::{ header, pslx, recount, mask, wildcard } +} + +///| +pub fn PslWriteConfig::default() -> PslWriteConfig { + PslWriteConfig::{ + header: true, + pslx: false, + recount: false, + mask: PslNoMask, + wildcard: 'N', + } +} + +///| +pub fn PslAlignment::blocks(self : PslAlignment) -> Array[PslBlock] { + psl_raw_blocks( + self.target_coordinates, + self.query_coordinates, + self.query_block_sequences, + self.target_block_sequences, + ) +} + +///| +pub fn PslAlignment::counts(self : PslAlignment) -> PslCounts { + psl_path_counts(self.target_coordinates, self.query_coordinates) +} + +///| +pub fn PslAlignment::is_reverse(self : PslAlignment) -> Bool { + match self.sequence_type { + PslNucleotide => + self.query_coordinates[0] > + self.query_coordinates[self.query_coordinates.length() - 1] + PslProtein => + self.target_coordinates[0] > + self.target_coordinates[self.target_coordinates.length() - 1] + } +} + +///| +pub fn PslAlignment::is_pslx(self : PslAlignment) -> Bool { + self.query_block_sequences.length() > 0 +} + +///| +pub fn PslAlignment::identity(self : PslAlignment) -> Double { + let total = self.matches + + self.mismatches + + self.repeat_matches + + self.n_count + if total == 0 { + 0.0 + } else { + (self.matches + self.repeat_matches).to_double() / total.to_double() + } +} + +///| +pub fn PslAlignment::score(self : PslAlignment) -> Int { + let counts = self.counts() + self.matches + + self.repeat_matches / 2 - + self.mismatches - + counts.query_insert_count - + counts.target_insert_count +} + +///| +pub fn PslAlignment::target_to_query( + self : PslAlignment, + position : Int, +) -> Int? { + if position < 0 || position >= self.target_size { + return None + } + for block in self.blocks() { + let target_minimum = psl_min(block.target_start, block.target_end) + let target_maximum = psl_max(block.target_start, block.target_end) + if position >= target_minimum && position < target_maximum { + let target_offset = if block.target_end > block.target_start { + position - block.target_start + } else { + block.target_start - 1 - position + } + let query_offset = match self.sequence_type { + PslNucleotide => target_offset + PslProtein => target_offset / 3 + } + return Some( + if block.query_end > block.query_start { + block.query_start + query_offset + } else { + block.query_start - 1 - query_offset + }, + ) + } + } + None +} + +///| +pub fn PslAlignment::query_to_target_interval( + self : PslAlignment, + position : Int, +) -> (Int, Int)? { + if position < 0 || position >= self.query_size { + return None + } + for block in self.blocks() { + let query_minimum = psl_min(block.query_start, block.query_end) + let query_maximum = psl_max(block.query_start, block.query_end) + if position >= query_minimum && position < query_maximum { + let query_offset = if block.query_end > block.query_start { + position - block.query_start + } else { + block.query_start - 1 - position + } + let scale = match self.sequence_type { + PslNucleotide => 1 + PslProtein => 3 + } + let oriented = query_offset * scale + let first = if block.target_end > block.target_start { + block.target_start + oriented + } else { + block.target_start - oriented - scale + } + return Some( + (psl_min(first, first + scale), psl_max(first, first + scale)), + ) + } + } + None +} + +///| +pub fn PslAlignment::query_to_target( + self : PslAlignment, + position : Int, +) -> Int? { + match self.query_to_target_interval(position) { + Some((start, _)) => Some(start) + None => None + } +} + +///| +fn psl_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 32 + } else { + code + } +} + +///| +fn psl_is_lower(code : Int) -> Bool { + code >= 'a'.to_int() && code <= 'z'.to_int() +} + +///| +fn psl_is_upper(code : Int) -> Bool { + code >= 'A'.to_int() && code <= 'Z'.to_int() +} + +///| +fn psl_complement_code(code : Int) -> Int { + match psl_upper_code(code) { + 65 => 84 + 67 => 71 + 71 => 67 + 84 => 65 + 85 => 65 + 82 => 89 + 89 => 82 + 77 => 75 + 75 => 77 + 66 => 86 + 86 => 66 + 68 => 72 + 72 => 68 + value => value + } +} + +///| +fn psl_translated_block( + sequence : String, + start : Int, + end : Int, +) -> String raise PslError { + let minimum = psl_min(start, end) + let maximum = psl_max(start, end) + let dna = sequence[minimum:maximum].to_owned() + let oriented = if end < start { + Seq::new(dna).reverse_complement() + } else { + Seq::new(dna) + } + let protein = oriented.translate() catch { + SeqError(message) => { + psl_fail("cannot translate PSL target block: " + message) + Seq::new("") + } + } + protein.to_string() +} + +///| +pub fn PslAlignment::recount( + self : PslAlignment, + target_sequence? : String = "", + query_sequence? : String = "", + mask? : PslMaskMode = PslNoMask, + wildcard? : Char = 'N', +) -> PslAlignment raise PslError { + let target = if target_sequence.length() > 0 { + target_sequence + } else { + self.target_sequence + } + let query = if query_sequence.length() > 0 { + query_sequence + } else { + self.query_sequence + } + psl_validate_sequence(target, self.target_size, "PSL target sequence") + psl_validate_sequence(query, self.query_size, "PSL query sequence") + if target.length() == 0 || query.length() == 0 { + psl_fail("PSL recount requires concrete target and query sequences") + } + if self.sequence_type == PslProtein && mask != PslNoMask { + psl_fail("repeat masking is only defined for nucleotide PSL") + } + let wildcard_code = psl_upper_code(wildcard.to_int()) + let mut matches = 0 + let mut mismatches = 0 + let mut repeat_matches = 0 + let mut n_count = 0 + for block in self.blocks() { + match self.sequence_type { + PslProtein => { + let translated = psl_translated_block( + target, + block.target_start, + block.target_end, + ) + if translated.length() != block.query_size { + psl_fail("translated PSL target block has an unexpected length") + } + for offset in 0.. + for offset in 0.. block.query_start { + block.query_start + offset + } else { + block.query_start - 1 - offset + } + let raw_target = target.unsafe_get(target_index).to_int() + let target_code = psl_upper_code(raw_target) + let raw_query = query.unsafe_get(query_index).to_int() + let query_code = if block.query_end < block.query_start { + psl_complement_code(raw_query) + } else { + psl_upper_code(raw_query) + } + if target_code == wildcard_code || query_code == wildcard_code { + n_count = n_count + 1 + } else if target_code != query_code { + mismatches = mismatches + 1 + } else { + let masked = match mask { + PslMaskLower => psl_is_lower(raw_target) + PslMaskUpper => psl_is_upper(raw_target) + PslNoMask => false + } + if masked { + repeat_matches = repeat_matches + 1 + } else { + matches = matches + 1 + } + } + } + } + } + PslAlignment::create( + self.target_name, + self.target_size, + self.query_name, + self.query_size, + self.target_coordinates, + self.query_coordinates, + sequence_type=self.sequence_type, + matches~, + mismatches~, + repeat_matches~, + n_count~, + target_sequence=target, + query_sequence=query, + ) +} + +///| +fn psl_storage(alignment : PslAlignment) -> PslStorage { + let blocks = alignment.blocks() + let reverse = alignment.is_reverse() + let block_sizes : Array[Int] = [] + let query_starts : Array[Int] = [] + let target_starts : Array[Int] = [] + let mut query_start = alignment.query_size + let mut query_end = 0 + let mut target_start = alignment.target_size + let mut target_end = 0 + for block in blocks { + let query_minimum = psl_min(block.query_start, block.query_end) + let query_maximum = psl_max(block.query_start, block.query_end) + let target_minimum = psl_min(block.target_start, block.target_end) + let target_maximum = psl_max(block.target_start, block.target_end) + query_start = psl_min(query_start, query_minimum) + query_end = psl_max(query_end, query_maximum) + target_start = psl_min(target_start, target_minimum) + target_end = psl_max(target_end, target_maximum) + block_sizes.push(block.query_size) + query_starts.push( + if alignment.sequence_type == PslNucleotide && reverse { + alignment.query_size - query_maximum + } else { + query_minimum + }, + ) + target_starts.push( + if alignment.sequence_type == PslProtein && reverse { + alignment.target_size - target_maximum + } else { + target_minimum + }, + ) + } + let strand = match alignment.sequence_type { + PslNucleotide => if reverse { "-" } else { "+" } + PslProtein => if reverse { "+-" } else { "++" } + } + PslStorage::{ + strand, + query_start, + query_end, + target_start, + target_end, + block_sizes, + query_starts, + target_starts, + } +} + +///| +fn psl_join_ints(values : Array[Int]) -> String { + let output = StringBuilder::new() + for index in 0.. 0 { + output.write_char(',') + } + output.write_string(values[index].to_string()) + } + if values.length() > 0 { + output.write_char(',') + } + output.to_string() +} + +///| +fn psl_join_strings(values : Array[String]) -> String { + let output = StringBuilder::new() + for index in 0.. 0 { + output.write_char(',') + } + output.write_string(values[index]) + } + if values.length() > 0 { + output.write_char(',') + } + output.to_string() +} + +///| +fn psl_block_sequences( + alignment : PslAlignment, +) -> (Array[String], Array[String]) raise PslError { + if alignment.query_block_sequences.length() > 0 { + return ( + psl_copy_strings(alignment.query_block_sequences), + psl_copy_strings(alignment.target_block_sequences), + ) + } + if alignment.target_sequence.length() == 0 || + alignment.query_sequence.length() == 0 { + psl_fail( + "PSLX output requires block sequences or concrete target/query sequences", + ) + } + let query_sequences : Array[String] = [] + let target_sequences : Array[String] = [] + for block in alignment.blocks() { + match alignment.sequence_type { + PslNucleotide => { + let target = alignment.target_sequence[block.target_start:block.target_end].to_owned() + let query_minimum = psl_min(block.query_start, block.query_end) + let query_maximum = psl_max(block.query_start, block.query_end) + let raw_query = alignment.query_sequence[query_minimum:query_maximum].to_owned() + let query = if block.query_end < block.query_start { + Seq::new(raw_query).reverse_complement().to_string() + } else { + raw_query + } + target_sequences.push(target) + query_sequences.push(query) + } + PslProtein => { + let target = psl_translated_block( + alignment.target_sequence, + block.target_start, + block.target_end, + ) + let query = alignment.query_sequence[block.query_start:block.query_end].to_owned() + target_sequences.push(target) + query_sequences.push(query) + } + } + } + (query_sequences, target_sequences) +} + +///| +pub fn psl_header(version? : String = "3") -> String raise PslError { + psl_validate_text(version, "PSL version", false) + for index in 0.. String raise PslError { + let alignment = if config.recount { + self.recount(mask=config.mask, wildcard=config.wildcard) + } else { + self + } + let storage = psl_storage(alignment) + let counts = alignment.counts() + let output = StringBuilder::new() + let fields = [ + alignment.matches.to_string(), + alignment.mismatches.to_string(), + alignment.repeat_matches.to_string(), + alignment.n_count.to_string(), + counts.query_insert_count.to_string(), + counts.query_insert_bases.to_string(), + counts.target_insert_count.to_string(), + counts.target_insert_bases.to_string(), + storage.strand, + alignment.query_name, + alignment.query_size.to_string(), + storage.query_start.to_string(), + storage.query_end.to_string(), + alignment.target_name, + alignment.target_size.to_string(), + storage.target_start.to_string(), + storage.target_end.to_string(), + counts.block_count.to_string(), + psl_join_ints(storage.block_sizes), + psl_join_ints(storage.query_starts), + psl_join_ints(storage.target_starts), + ] + for index in 0.. 0 { + output.write_char('\t') + } + output.write_string(fields[index]) + } + if config.pslx { + let (query_sequences, target_sequences) = psl_block_sequences(alignment) + output.write_char('\t') + output.write_string(psl_join_strings(query_sequences)) + output.write_char('\t') + output.write_string(psl_join_strings(target_sequences)) + } + output.write_char('\n') + output.to_string() +} + +///| +pub fn psl_write( + alignments : Array[PslAlignment], + config? : PslWriteConfig = PslWriteConfig::default(), + version? : String = "3", +) -> String raise PslError { + let output = StringBuilder::new() + if config.header { + output.write_string(psl_header(version~)) + } + for alignment in alignments { + output.write_string(alignment.format(config~)) + } + output.to_string() +} + +///| +fn psl_parse_int(value : String, label : String) -> Int raise PslError { + if value.length() == 0 { + psl_fail(label + " must not be empty") + } + let mut result = 0 + for index in 0.. '9'.to_int() { + psl_fail(label + " must be a non-negative integer") + } + let digit = code - '0'.to_int() + if result > (2147483647 - digit) / 10 { + psl_fail(label + " exceeds the supported integer range") + } + result = result * 10 + digit + } + result +} + +///| +fn psl_parse_csv_ints( + value : String, + label : String, +) -> Array[Int] raise PslError { + let parts = split_by_char(value, ','.to_int()) + let values : Array[Int] = [] + for index in 0.. Array[String] raise PslError { + let parts = split_by_char(value, ','.to_int()) + let values : Array[String] = [] + for index in 0.. PslAlignment raise PslError { + let words = split_by_char(line, '\t'.to_int()) + if words.length() != 21 && words.length() != 23 { + psl_fail( + "PSL line " + + line_number.to_string() + + " has " + + words.length().to_string() + + " columns; expected 21 or 23", + ) + } + let matches = psl_parse_int(words[0], "PSL matches") + let mismatches = psl_parse_int(words[1], "PSL misMatches") + let repeat_matches = psl_parse_int(words[2], "PSL repMatches") + let n_count = psl_parse_int(words[3], "PSL nCount") + let declared_query_insert_count = psl_parse_int(words[4], "PSL qNumInsert") + let declared_query_insert_bases = psl_parse_int(words[5], "PSL qBaseInsert") + let declared_target_insert_count = psl_parse_int(words[6], "PSL tNumInsert") + let declared_target_insert_bases = psl_parse_int(words[7], "PSL tBaseInsert") + let strand = words[8] + let sequence_type = match strand { + "+" | "-" => PslNucleotide + "++" | "+-" => PslProtein + _ => { + psl_fail("PSL strand must be '+', '-', '++', or '+-'") + PslNucleotide + } + } + let query_name = words[9] + let query_size = psl_parse_int(words[10], "PSL qSize") + let declared_query_start = psl_parse_int(words[11], "PSL qStart") + let declared_query_end = psl_parse_int(words[12], "PSL qEnd") + let target_name = words[13] + let target_size = psl_parse_int(words[14], "PSL tSize") + let declared_target_start = psl_parse_int(words[15], "PSL tStart") + let declared_target_end = psl_parse_int(words[16], "PSL tEnd") + let block_count = psl_parse_int(words[17], "PSL blockCount") + if block_count <= 0 { + psl_fail("PSL blockCount must be positive") + } + let block_sizes = psl_parse_csv_ints(words[18], "PSL blockSizes") + let query_starts = psl_parse_csv_ints(words[19], "PSL qStarts") + let target_starts = psl_parse_csv_ints(words[20], "PSL tStarts") + if block_sizes.length() != block_count || + query_starts.length() != block_count || + target_starts.length() != block_count { + psl_fail("PSL block arrays do not match blockCount") + } + let query_block_sequences : Array[String] = if words.length() == 23 { + psl_parse_csv_strings(words[21], "PSLX query sequences") + } else { + [] + } + let target_block_sequences : Array[String] = if words.length() == 23 { + psl_parse_csv_strings(words[22], "PSLX target sequences") + } else { + [] + } + if words.length() == 23 && + ( + query_block_sequences.length() != block_count || + target_block_sequences.length() != block_count + ) { + psl_fail("PSLX sequence arrays do not match blockCount") + } + let target_coordinates : Array[Int] = [target_starts[0]] + let query_coordinates : Array[Int] = [query_starts[0]] + let mut target_position = target_starts[0] + let mut query_position = query_starts[0] + for index in 0.. block_size + PslProtein => 3 * block_size + } + let target_start = target_starts[index] + let query_start = query_starts[index] + if target_start < target_position || query_start < query_position { + psl_fail("PSL blocks overlap or are not sorted") + } + if target_start + target_block_size > target_size || + query_start + block_size > query_size { + psl_fail("PSL block extends beyond a declared sequence size") + } + if target_start != target_position { + target_coordinates.push(target_start) + query_coordinates.push(query_position) + target_position = target_start + } + if query_start != query_position { + target_coordinates.push(target_position) + query_coordinates.push(query_start) + query_position = query_start + } + target_position = target_position + target_block_size + query_position = query_position + block_size + target_coordinates.push(target_position) + query_coordinates.push(query_position) + } + if strand == "-" { + for index in 0.. Bool { + if line.length() == 0 { + return false + } + for index in 0.. PslDocument raise PslError { + let raw_lines = split_by_char(content, '\n'.to_int()) + let lines : Array[String] = [] + for line in raw_lines { + lines.push(psl_strip_cr(line)) + } + let mut index = 0 + while index < lines.length() && lines[index].length() == 0 { + index = index + 1 + } + if index >= lines.length() { + psl_fail("PSL input is empty") + } + let mut version = "" + let mut has_header = false + if starts_with(lines[index], 0, "psLayout ") { + let words = split_by_whitespace(lines[index]) + if words.length() != 3 || words[1] != "version" { + psl_fail("malformed PSL header") + } + version = words[2] + psl_validate_text(version, "PSL version", false) + has_header = true + index = index + 1 + let mut found_separator = false + while index < lines.length() { + if psl_is_separator(lines[index]) { + found_separator = true + index = index + 1 + break + } + index = index + 1 + } + if !found_separator { + psl_fail("PSL header separator was not found") + } + } + let alignments : Array[PslAlignment] = [] + while index < lines.length() { + if lines[index].length() > 0 { + alignments.push(psl_parse_record(lines[index], index + 1)) + } + index = index + 1 + } + PslDocument::{ version, has_header, alignments } +} + +///| +pub fn PslDocument::query( + self : PslDocument, + name : String, +) -> Array[PslAlignment] { + let results : Array[PslAlignment] = [] + for alignment in self.alignments { + if alignment.query_name == name { + results.push(alignment) + } + } + results +} + +///| +pub fn PslDocument::target( + self : PslDocument, + name : String, +) -> Array[PslAlignment] { + let results : Array[PslAlignment] = [] + for alignment in self.alignments { + if alignment.target_name == name { + results.push(alignment) + } + } + results +} + +///| +pub fn PslDocument::summary(self : PslDocument) -> PslSummary { + let mut pslx_count = 0 + let mut nucleotide_count = 0 + let mut protein_count = 0 + let mut aligned_query_units = 0 + let mut query_insert_bases = 0 + let mut target_insert_bases = 0 + for alignment in self.alignments { + if alignment.is_pslx() { + pslx_count = pslx_count + 1 + } + match alignment.sequence_type { + PslNucleotide => nucleotide_count = nucleotide_count + 1 + PslProtein => protein_count = protein_count + 1 + } + let counts = alignment.counts() + aligned_query_units = aligned_query_units + counts.aligned_query_units + query_insert_bases = query_insert_bases + counts.query_insert_bases + target_insert_bases = target_insert_bases + counts.target_insert_bases + } + PslSummary::{ + alignment_count: self.alignments.length(), + pslx_count, + nucleotide_count, + protein_count, + aligned_query_units, + query_insert_bases, + target_insert_bases, + } +} + +///| +pub fn PslDocument::write( + self : PslDocument, + config? : PslWriteConfig = PslWriteConfig::default(), +) -> String raise PslError { + let version = if self.version.length() > 0 { self.version } else { "3" } + psl_write(self.alignments, config~, version~) +} + +///| +pub fn PslAlignment::to_coordinate_alignment( + self : PslAlignment, +) -> CoordinatePairwiseAlignment raise PslError { + if self.sequence_type == PslProtein && self.is_reverse() { + psl_fail( + "reverse-target translated PSL cannot use the target-increasing coordinate adapter", + ) + } + coordinate_pairwise_alignment_with_lengths( + self.target_name, + self.target_size, + self.query_name, + self.query_size, + self.target_coordinates, + self.query_coordinates, + target_sequence=self.target_sequence, + query_sequence=self.query_sequence, + ) catch { + AlignmentMapError(message) => { + psl_fail("cannot convert PSL coordinate path: " + message) + coordinate_pairwise_alignment_with_lengths( + "target", + 1, + "query", + 1, + [0, 1], + [0, 1], + ) catch { + _ => abort("unreachable PSL coordinate fallback") + } + } + } +} + +///| +pub fn psl_from_coordinate_alignment( + alignment : CoordinatePairwiseAlignment, + sequence_type? : PslSequenceType = PslNucleotide, +) -> PslAlignment raise PslError { + PslAlignment::create( + alignment.target_name, + alignment.target_length, + alignment.query_name, + alignment.query_length, + alignment.target_coordinates, + alignment.query_coordinates, + sequence_type~, + target_sequence=alignment.target_sequence, + query_sequence=alignment.query_sequence, + ) +} + +///| +pub fn psl_example_data() -> Array[PslAlignment] raise PslError { + let nucleotide = PslAlignment::create( + "chrExample", + 24, + "readForward", + 12, + [2, 6, 9, 9, 13], + [0, 4, 4, 6, 10], + target_sequence="TTACGTGGGACGTCCCAAATTTGG", + query_sequence="ACGTGGACGTAA", + matches=6, + mismatches=1, + repeat_matches=0, + n_count=1, + ) + let reverse = PslAlignment::create( + "chrExample", + 24, + "readReverse", + 8, + [4, 8, 10, 14], + [8, 4, 4, 0], + ) + let protein = PslAlignment::create( + "codingDna", + 12, + "peptide", + 2, + [0, 6], + [0, 2], + sequence_type=PslProtein, + target_sequence="ATGGCTTAATAG", + query_sequence="MA", + ).recount() + [nucleotide, reverse, protein] +} diff --git a/src/align_sam.mbt b/src/align_sam.mbt new file mode 100644 index 00000000..d8064c83 --- /dev/null +++ b/src/align_sam.mbt @@ -0,0 +1,2072 @@ +// Alignment-aware SAM parsing, coordinate paths, typed tags, and writing. +// +// This module follows Biopython 1.86 Bio.Align.sam semantics. SAM POS and +// PNEXT fields are converted from one-based coordinates to zero-based values. + +///| +pub suberror AlignSamError { + AlignSamError(String) +} + +///| +pub(all) enum AlignSamOperation { + SamAligned + SamInsertion + SamDeletion + SamSkipped + SamEqual + SamMismatch +} derive(Eq, Debug) + +///| +pub(all) enum AlignSamTagValue { + SamTagInteger(Int) + SamTagFloat(Double) + SamTagCharacter(String) + SamTagString(String) + SamTagHex(String) + SamTagIntegerArray(String, Array[Int]) + SamTagFloatArray(Array[Double]) +} derive(Debug) + +///| +pub struct AlignSamHeaderField { + key : String + value : String +} derive(Eq, Debug) + +///| +pub struct AlignSamHeader { + record_type : String + fields : Array[AlignSamHeaderField] + comment : String +} derive(Debug) + +///| +pub struct AlignSamReference { + name : String + length : Int + annotations : Array[AlignSamHeaderField] +} derive(Debug) + +///| +pub struct AlignSamCigarElement { + operation : String + length : Int +} derive(Eq, Debug) + +///| +pub struct AlignSamTag { + name : String + value : AlignSamTagValue +} derive(Debug) + +///| +pub struct AlignSamStats { + reference_consumed : Int + query_consumed : Int + aligned : Int + inserted : Int + deleted : Int + skipped : Int + soft_clipped : Int + hard_clipped : Int +} derive(Eq, Debug) + +///| +pub struct AlignSamAlignment { + query_name : String + flag : Int + reference_name : String + reference_length : Int? + target_coordinates : Array[Int] + query_coordinates : Array[Int] + operations : Array[AlignSamOperation] + cigar : Array[AlignSamCigarElement] + mapq : Int? + mate_reference : String? + mate_position : Int? + template_length : Int? + query_sequence : String + sequence_known : Bool + qualities : Array[Int] + qualities_known : Bool + hard_clip_left : Int? + hard_clip_right : Int? + tags : Array[AlignSamTag] +} + +///| +pub struct AlignSamDocument { + headers : Array[AlignSamHeader] + references : Array[AlignSamReference] + alignments : Array[AlignSamAlignment] +} + +///| +fn align_sam_fail(message : String) -> Unit raise AlignSamError { + raise AlignSamError(message) +} + +///| +fn align_sam_copy_ints(values : Array[Int]) -> Array[Int] { + let copy : Array[Int] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn align_sam_copy_doubles(values : Array[Double]) -> Array[Double] { + let copy : Array[Double] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn align_sam_copy_tag_value(value : AlignSamTagValue) -> AlignSamTagValue { + match value { + SamTagInteger(item) => SamTagInteger(item) + SamTagFloat(item) => SamTagFloat(item) + SamTagCharacter(item) => SamTagCharacter(item) + SamTagString(item) => SamTagString(item) + SamTagHex(item) => SamTagHex(item) + SamTagIntegerArray(kind, items) => + SamTagIntegerArray(kind, align_sam_copy_ints(items)) + SamTagFloatArray(items) => SamTagFloatArray(align_sam_copy_doubles(items)) + } +} + +///| +fn align_sam_copy_cigar( + values : Array[AlignSamCigarElement], +) -> Array[AlignSamCigarElement] { + let copy : Array[AlignSamCigarElement] = [] + for value in values { + copy.push(AlignSamCigarElement::{ + operation: value.operation, + length: value.length, + }) + } + copy +} + +///| +fn align_sam_copy_fields( + values : Array[AlignSamHeaderField], +) -> Array[AlignSamHeaderField] { + let copy : Array[AlignSamHeaderField] = [] + for value in values { + copy.push(AlignSamHeaderField::{ key: value.key, value: value.value }) + } + copy +} + +///| +fn align_sam_strip_cr(value : String) -> String { + if value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '\r'.to_int() { + value[0:value.length() - 1].to_owned() + } else { + value + } +} + +///| +fn align_sam_split_lines(value : String) -> Array[String] { + let lines : Array[String] = [] + let mut start = 0 + for index in 0.. Array[String] { + let parts : Array[String] = [] + let mut start = 0 + for index in 0.. Int? { + for index in start.. Int raise AlignSamError { + if value.length() == 0 { + align_sam_fail(label + " is missing") + } + let mut index = 0 + let negative = value.unsafe_get(0).to_int() == '-'.to_int() + if negative || value.unsafe_get(0).to_int() == '+'.to_int() { + index = 1 + } + if index == value.length() { + align_sam_fail(label + " is not an integer") + } + if value == "-2147483648" { + return -2147483647 - 1 + } + let mut result = 0 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + align_sam_fail(label + " is not an integer: " + value) + } + let digit = code - '0'.to_int() + if result > (2147483647 - digit) / 10 { + align_sam_fail(label + " exceeds the MoonBit Int range") + } + result = result * 10 + digit + index = index + 1 + } + if negative { + -result + } else { + result + } +} + +///| +fn align_sam_parse_double( + value : String, + label : String, +) -> Double raise AlignSamError { + match parse_double(value) { + Some(result) => { + if result.abs() > 1.0e300 { + align_sam_fail(label + " must be finite") + } + result + } + None => { + align_sam_fail(label + " is not a floating-point number: " + value) + 0.0 + } + } +} + +///| +fn align_sam_validate_token( + value : String, + label : String, + allow_star : Bool, +) -> Unit raise AlignSamError { + if value.length() == 0 { + align_sam_fail(label + " must not be empty") + } + if value == "*" && allow_star { + return + } + for index in 0.. Unit raise AlignSamError { + if name.length() != 2 { + align_sam_fail("SAM tag names must contain exactly two characters") + } + let first = name.unsafe_get(0).to_int() + let second = name.unsafe_get(1).to_int() + let first_ok = (first >= 'A'.to_int() && first <= 'Z'.to_int()) || + (first >= 'a'.to_int() && first <= 'z'.to_int()) + let second_ok = (second >= 'A'.to_int() && second <= 'Z'.to_int()) || + (second >= 'a'.to_int() && second <= 'z'.to_int()) || + (second >= '0'.to_int() && second <= '9'.to_int()) + if !first_ok || !second_ok { + align_sam_fail("invalid SAM tag name " + name) + } +} + +///| +fn align_sam_is_hex(code : Int) -> Bool { + (code >= '0'.to_int() && code <= '9'.to_int()) || + (code >= 'A'.to_int() && code <= 'F'.to_int()) || + (code >= 'a'.to_int() && code <= 'f'.to_int()) +} + +///| +fn align_sam_is_md_base(code : Int) -> Bool { + code == 'A'.to_int() || + code == 'C'.to_int() || + code == 'G'.to_int() || + code == 'T'.to_int() || + code == 'N'.to_int() || + code == 'a'.to_int() || + code == 'c'.to_int() || + code == 'g'.to_int() || + code == 't'.to_int() || + code == 'n'.to_int() +} + +///| +fn align_sam_operation_from_text( + operation : String, +) -> AlignSamOperation raise AlignSamError { + if operation == "M" { + SamAligned + } else if operation == "I" { + SamInsertion + } else if operation == "D" { + SamDeletion + } else if operation == "N" { + SamSkipped + } else if operation == "=" { + SamEqual + } else if operation == "X" { + SamMismatch + } else { + align_sam_fail("CIGAR operation " + operation + " has no alignment path") + SamAligned + } +} + +///| +fn align_sam_is_clip(operation : String) -> Bool { + operation == "S" || operation == "H" +} + +///| +fn align_sam_is_core(operation : String) -> Bool { + operation == "M" || + operation == "I" || + operation == "D" || + operation == "N" || + operation == "=" || + operation == "X" +} + +///| +pub fn align_sam_parse_cigar( + text : String, +) -> Array[AlignSamCigarElement] raise AlignSamError { + if text == "*" { + return [] + } + if text.length() == 0 { + align_sam_fail("CIGAR must not be empty") + } + let elements : Array[AlignSamCigarElement] = [] + let mut start = 0 + let mut index = 0 + while index < text.length() { + let code = text.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + index = index + 1 + continue + } + if index == start { + align_sam_fail("CIGAR operation is missing a length") + } + let length = align_sam_parse_int( + text[start:index].to_owned(), + "CIGAR length", + ) + if length <= 0 { + align_sam_fail("CIGAR lengths must be positive") + } + let operation = text[index:index + 1].to_owned() + if operation == "P" { + align_sam_fail( + "CIGAR padding operation P is not supported by Bio.Align.sam", + ) + } + if !align_sam_is_clip(operation) && !align_sam_is_core(operation) { + align_sam_fail("unknown CIGAR operation " + operation) + } + elements.push(AlignSamCigarElement::{ operation, length }) + index = index + 1 + start = index + } + if start != text.length() { + align_sam_fail("CIGAR ends with a length but no operation") + } + let mut first_core = elements.length() + let mut last_core = -1 + for element_index in 0.. last_core + 1 { + element_index = element_index - 1 + let operation = elements[element_index].operation + if operation == "S" { + seen_soft = true + } else if operation == "H" && !seen_soft { + // Right hard clipping is outermost and therefore visited first. + } else if operation == "H" { + align_sam_fail("right hard clipping must follow right soft clipping") + } + } + align_sam_copy_cigar(elements) +} + +///| +pub fn align_sam_cigar_string(elements : Array[AlignSamCigarElement]) -> String { + if elements.length() == 0 { + return "*" + } + let output = StringBuilder::new() + for element in elements { + output.write_string(element.length.to_string()) + output.write_string(element.operation) + } + output.to_string() +} + +///| +fn align_sam_cigar_query_consumed( + elements : Array[AlignSamCigarElement], +) -> Int { + let mut total = 0 + for element in elements { + if element.operation == "M" || + element.operation == "I" || + element.operation == "S" || + element.operation == "=" || + element.operation == "X" { + total = total + element.length + } + } + total +} + +///| +fn align_sam_left_soft(elements : Array[AlignSamCigarElement]) -> Int { + let mut total = 0 + for element in elements { + if align_sam_is_core(element.operation) { + break + } + if element.operation == "S" { + total = total + element.length + } + } + total +} + +///| +fn align_sam_raw_hard_clips( + elements : Array[AlignSamCigarElement], +) -> (Int?, Int?) { + let mut left = 0 + let mut right = 0 + for element in elements { + if align_sam_is_core(element.operation) { + break + } + if element.operation == "H" { + left = left + element.length + } + } + let mut index = elements.length() + while index > 0 { + index = index - 1 + let element = elements[index] + if align_sam_is_core(element.operation) { + break + } + if element.operation == "H" { + right = right + element.length + } + } + ( + if left == 0 { + None + } else { + Some(left) + }, + if right == 0 { + None + } else { + Some(right) + }, + ) +} + +///| +fn align_sam_build_path( + target_start : Int, + cigar : Array[AlignSamCigarElement], + query_length : Int, + reverse : Bool, +) -> (Array[Int], Array[Int], Array[AlignSamOperation]) raise AlignSamError { + let target_coordinates : Array[Int] = [target_start] + let query_coordinates : Array[Int] = [align_sam_left_soft(cigar)] + let operations : Array[AlignSamOperation] = [] + let mut target = target_start + let mut query = query_coordinates[0] + for element in cigar { + let operation = element.operation + if operation == "S" || operation == "H" { + continue + } + if operation == "M" || operation == "=" || operation == "X" { + target = target + element.length + query = query + element.length + } else if operation == "I" { + query = query + element.length + } else if operation == "D" || operation == "N" { + target = target + element.length + } + target_coordinates.push(target) + query_coordinates.push(query) + operations.push(align_sam_operation_from_text(operation)) + } + if query_length != align_sam_cigar_query_consumed(cigar) { + align_sam_fail( + "query length does not match query-consuming CIGAR operations", + ) + } + if reverse { + for index in 0.. AlignSamTag raise AlignSamError { + let first = match align_sam_find_char(field, ':', 0) { + Some(value) => value + None => { + align_sam_fail("SAM optional tag is missing its datatype") + 0 + } + } + let second = match align_sam_find_char(field, ':', first + 1) { + Some(value) => value + None => { + align_sam_fail("SAM optional tag is missing its value") + 0 + } + } + let name = field[0:first].to_owned() + let datatype = field[first + 1:second].to_owned() + let text = field[second + 1:field.length()].to_owned() + align_sam_validate_tag_name(name) + let value = if datatype == "i" { + SamTagInteger(align_sam_parse_int(text, "SAM integer tag " + name)) + } else if datatype == "f" { + SamTagFloat(align_sam_parse_double(text, "SAM float tag " + name)) + } else if datatype == "A" { + if text.length() != 1 { + align_sam_fail("SAM A tag " + name + " must contain one character") + } + let code = text.unsafe_get(0).to_int() + if code < 33 || code > 126 { + align_sam_fail("SAM A tag " + name + " must be printable ASCII") + } + SamTagCharacter(text) + } else if datatype == "Z" { + for index in 0.. 126 { + align_sam_fail("SAM Z tag " + name + " contains invalid text") + } + } + SamTagString(text) + } else if datatype == "H" { + if text.length() % 2 != 0 { + align_sam_fail("SAM H tag " + name + " must contain byte pairs") + } + for index in 0.. 127) { + align_sam_fail("SAM B:c value is outside the signed byte range") + } + if subtype == "C" && (item < 0 || item > 255) { + align_sam_fail("SAM B:C value is outside the unsigned byte range") + } + if subtype == "s" && (item < -32768 || item > 32767) { + align_sam_fail("SAM B:s value is outside the signed short range") + } + if subtype == "S" && (item < 0 || item > 65535) { + align_sam_fail("SAM B:S value is outside the unsigned short range") + } + if subtype == "I" && item < 0 { + align_sam_fail("SAM B:I values must be non-negative") + } + items.push(item) + } + SamTagIntegerArray(subtype, items) + } else { + align_sam_fail("unknown SAM B tag subtype " + subtype) + SamTagIntegerArray("i", []) + } + } else { + align_sam_fail("unknown SAM optional tag datatype " + datatype) + SamTagString("") + } + AlignSamTag::{ name, value } +} + +///| +fn align_sam_tag_text(tag : AlignSamTag) -> String { + let prefix = tag.name + ":" + match tag.value { + SamTagInteger(value) => prefix + "i:" + value.to_string() + SamTagFloat(value) => prefix + "f:" + value.to_string() + SamTagCharacter(value) => prefix + "A:" + value + SamTagString(value) => prefix + "Z:" + value + SamTagHex(value) => prefix + "H:" + value + SamTagIntegerArray(kind, values) => { + let output = StringBuilder::new() + output.write_string(prefix) + output.write_string("B:") + output.write_string(kind) + for value in values { + output.write_char(',') + output.write_string(value.to_string()) + } + output.to_string() + } + SamTagFloatArray(values) => { + let output = StringBuilder::new() + output.write_string(prefix) + output.write_string("B:f") + for value in values { + output.write_char(',') + output.write_string(value.to_string()) + } + output.to_string() + } + } +} + +///| +pub fn align_sam_integer_tag( + name : String, + value : Int, +) -> AlignSamTag raise AlignSamError { + align_sam_validate_tag_name(name) + AlignSamTag::{ name, value: SamTagInteger(value) } +} + +///| +pub fn align_sam_float_tag( + name : String, + value : Double, +) -> AlignSamTag raise AlignSamError { + align_sam_validate_tag_name(name) + if value.abs() > 1.0e300 { + align_sam_fail("SAM floating-point tag value must be finite") + } + AlignSamTag::{ name, value: SamTagFloat(value) } +} + +///| +pub fn align_sam_string_tag( + name : String, + value : String, +) -> AlignSamTag raise AlignSamError { + align_sam_parse_tag(name + ":Z:" + value) +} + +///| +pub fn align_sam_character_tag( + name : String, + value : String, +) -> AlignSamTag raise AlignSamError { + align_sam_parse_tag(name + ":A:" + value) +} + +///| +pub fn align_sam_hex_tag( + name : String, + value : String, +) -> AlignSamTag raise AlignSamError { + align_sam_parse_tag(name + ":H:" + value) +} + +///| +pub fn align_sam_integer_array_tag( + name : String, + subtype : String, + values : Array[Int], +) -> AlignSamTag raise AlignSamError { + let output = StringBuilder::new() + output.write_string(name) + output.write_string(":B:") + output.write_string(subtype) + for value in values { + output.write_char(',') + output.write_string(value.to_string()) + } + align_sam_parse_tag(output.to_string()) +} + +///| +pub fn align_sam_float_array_tag( + name : String, + values : Array[Double], +) -> AlignSamTag raise AlignSamError { + let output = StringBuilder::new() + output.write_string(name) + output.write_string(":B:f") + for value in values { + output.write_char(',') + output.write_string(value.to_string()) + } + align_sam_parse_tag(output.to_string()) +} + +///| +fn align_sam_parse_header(line : String) -> AlignSamHeader raise AlignSamError { + if line.length() < 3 || line.unsafe_get(0).to_int() != '@'.to_int() { + align_sam_fail("invalid SAM header line") + } + let parts = align_sam_split_char(line, '\t') + let record_type = parts[0][1:parts[0].length()].to_owned() + if record_type.length() != 2 { + align_sam_fail("SAM header record type must contain two characters") + } + if record_type == "CO" { + let mut comment = "" + for index in 1.. 1 { + comment = comment + "\t" + } + comment = comment + parts[index] + } + return AlignSamHeader::{ record_type, fields: [], comment } + } + let fields : Array[AlignSamHeaderField] = [] + let seen : Map[String, Bool] = Map([]) + for index in 1.. value + None => { + align_sam_fail("SAM header field is missing ':'") + 0 + } + } + let key = part[0:separator].to_owned() + let value = part[separator + 1:part.length()].to_owned() + if key.length() != 2 || value.length() == 0 { + align_sam_fail("invalid SAM header field " + part) + } + if seen.contains(key) { + align_sam_fail("duplicate SAM header field " + key) + } + seen[key] = true + fields.push(AlignSamHeaderField::{ key, value }) + } + AlignSamHeader::{ record_type, fields, comment: "" } +} + +///| +fn align_sam_header_value(header : AlignSamHeader, key : String) -> String? { + for field in header.fields { + if field.key == key { + return Some(field.value) + } + } + None +} + +///| +fn align_sam_parse_reference( + header : AlignSamHeader, +) -> AlignSamReference raise AlignSamError { + let name = match align_sam_header_value(header, "SN") { + Some(value) => value + None => { + align_sam_fail("@SQ header is missing SN") + "" + } + } + align_sam_validate_token(name, "@SQ SN", false) + let length = match align_sam_header_value(header, "LN") { + Some(value) => align_sam_parse_int(value, "@SQ LN") + None => { + align_sam_fail("@SQ header is missing LN") + 0 + } + } + if length <= 0 { + align_sam_fail("@SQ LN must be positive") + } + let annotations : Array[AlignSamHeaderField] = [] + for field in header.fields { + if field.key != "SN" && field.key != "LN" { + if field.key == "TP" && + field.value != "linear" && + field.value != "circular" { + align_sam_fail("@SQ TP must be linear or circular") + } + annotations.push(AlignSamHeaderField::{ + key: field.key, + value: field.value, + }) + } + } + AlignSamReference::{ name, length, annotations } +} + +///| +fn align_sam_validate_sequence(sequence : String) -> Unit raise AlignSamError { + for index in 0..= 'A'.to_int() && code <= 'Z'.to_int()) || + (code >= 'a'.to_int() && code <= 'z'.to_int()) || + code == '='.to_int() || + code == '.'.to_int() + if !valid { + align_sam_fail("SAM SEQ contains an invalid character") + } + } +} + +///| +fn align_sam_reverse_ints(values : Array[Int]) -> Array[Int] { + let reversed : Array[Int] = [] + let mut index = values.length() + while index > 0 { + index = index - 1 + reversed.push(values[index]) + } + reversed +} + +///| +fn align_sam_reverse_complement(sequence : String) -> String { + Seq::new(sequence).reverse_complement().to_string() +} + +///| +fn align_sam_reference_length( + references : Array[AlignSamReference], + name : String, +) -> Int? { + for reference in references { + if reference.name == name { + return Some(reference.length) + } + } + None +} + +///| +fn align_sam_validate_md( + md : String, + cigar : Array[AlignSamCigarElement], +) -> Unit raise AlignSamError { + if md.length() == 0 { + align_sam_fail("MD tag must not be empty") + } + let expected : Array[String] = [] + for element in cigar { + if element.operation == "M" || + element.operation == "=" || + element.operation == "X" || + element.operation == "D" { + for _ in 0.. '9'.to_int() { + align_sam_fail( + "MD tag must start with and alternate through match counts", + ) + } + let start = index + while index < md.length() { + let digit = md.unsafe_get(index).to_int() + if digit < '0'.to_int() || digit > '9'.to_int() { + break + } + index = index + 1 + } + let matches = align_sam_parse_int( + md[start:index].to_owned(), + "MD match count", + ) + for _ in 0..= expected.length() { + align_sam_fail("MD match count exceeds the CIGAR reference path") + } + if expected[pointer] == "D" || expected[pointer] == "X" { + align_sam_fail("MD match count conflicts with the CIGAR operation") + } + pointer = pointer + 1 + } + expect_count = false + continue + } + if code == '^'.to_int() { + index = index + 1 + let start = index + while index < md.length() && + align_sam_is_md_base(md.unsafe_get(index).to_int()) { + if pointer >= expected.length() || expected[pointer] != "D" { + align_sam_fail("MD deletion does not match a CIGAR D operation") + } + pointer = pointer + 1 + index = index + 1 + } + if index == start { + align_sam_fail("MD deletion marker must be followed by bases") + } + expect_count = true + continue + } + if align_sam_is_md_base(code) { + if pointer >= expected.length() || expected[pointer] == "D" { + align_sam_fail("MD mismatch does not match an aligned CIGAR operation") + } + if expected[pointer] == "=" { + align_sam_fail("MD mismatch conflicts with a CIGAR = operation") + } + pointer = pointer + 1 + index = index + 1 + expect_count = true + continue + } + align_sam_fail("MD tag contains an invalid character") + } + if expect_count || pointer != expected.length() { + align_sam_fail( + "MD tag consumes " + + pointer.to_string() + + " reference bases but CIGAR requires " + + expected.length().to_string(), + ) + } +} + +///| +fn align_sam_parse_alignment( + line : String, + references : Array[AlignSamReference], +) -> AlignSamAlignment raise AlignSamError { + let fields = align_sam_split_char(line, '\t') + if fields.length() < 11 { + align_sam_fail( + "SAM alignment has " + + fields.length().to_string() + + " fields; expected at least 11", + ) + } + let query_name = fields[0] + align_sam_validate_token(query_name, "QNAME", false) + let flag = align_sam_parse_int(fields[1], "FLAG") + if flag < 0 || flag > 65535 { + align_sam_fail("FLAG must be between 0 and 65535") + } + let unmapped = (flag & 0x4) != 0 + let reverse = (flag & 0x10) != 0 + let reference_name = fields[2] + align_sam_validate_token(reference_name, "RNAME", true) + let one_based_position = align_sam_parse_int(fields[3], "POS") + let raw_mapq = align_sam_parse_int(fields[4], "MAPQ") + if raw_mapq < 0 || raw_mapq > 255 { + align_sam_fail("MAPQ must be between 0 and 255") + } + let cigar = align_sam_parse_cigar(fields[5]) catch { + AlignSamError(message) => + if unmapped && fields[5] == "*" { + [] + } else { + raise AlignSamError(message) + } + } + let raw_mate_reference = fields[6] + align_sam_validate_token(raw_mate_reference, "RNEXT", true) + let raw_mate_position = align_sam_parse_int(fields[7], "PNEXT") + let raw_template_length = align_sam_parse_int(fields[8], "TLEN") + let raw_sequence = fields[9] + let sequence_known = raw_sequence != "*" + if sequence_known { + align_sam_validate_sequence(raw_sequence) + } + let quality_text = fields[10] + let qualities_known = quality_text != "*" + let raw_qualities : Array[Int] = [] + if qualities_known { + if !sequence_known || quality_text.length() != raw_sequence.length() { + align_sam_fail("QUAL length must equal SEQ length") + } + for index in 0.. 126 { + align_sam_fail("QUAL contains a character outside PHRED+33 range") + } + raw_qualities.push(code - 33) + } + } + let tags : Array[AlignSamTag] = [] + let seen_tags : Map[String, Bool] = Map([]) + let mut md : String? = None + for index in 11.. md = Some(value) + _ => align_sam_fail("MD tag must use Z datatype") + } + } + if tag.name == "AS" { + match tag.value { + SamTagInteger(_) => () + _ => align_sam_fail("AS tag must use i datatype") + } + } + tags.push(tag) + } + if unmapped { + if reference_name != "*" || one_based_position != 0 || fields[5] != "*" { + align_sam_fail( + "unmapped SAM records must use RNAME *, POS 0, and CIGAR *", + ) + } + } else { + if reference_name == "*" { + align_sam_fail("mapped SAM records require a reference name") + } + if one_based_position <= 0 { + align_sam_fail("mapped SAM POS must be positive") + } + if fields[5] == "*" { + align_sam_fail("mapped SAM records require a CIGAR") + } + } + if raw_mate_position < 0 { + align_sam_fail("PNEXT must be non-negative") + } + let query_length = if sequence_known { + raw_sequence.length() + } else if unmapped { + 0 + } else { + align_sam_cigar_query_consumed(cigar) + } + let reference_length : Int? = if unmapped { + None + } else { + let length = align_sam_reference_length(references, reference_name) + if references.length() > 0 && length is None { + align_sam_fail( + "mapped reference " + reference_name + " is absent from @SQ", + ) + } + length + } + let (target_coordinates, query_coordinates, operations) = if unmapped { + ([], [], []) + } else { + let target_start = one_based_position - 1 + let path = align_sam_build_path(target_start, cigar, query_length, reverse) + match reference_length { + Some(length) => + if path.0[path.0.length() - 1] > length { + align_sam_fail("alignment extends beyond its @SQ reference length") + } + None => () + } + path + } + match md { + Some(value) => { + if unmapped { + align_sam_fail("unmapped SAM records cannot carry an MD tag") + } + align_sam_validate_md(value, cigar) + } + None => () + } + let raw_hard_clips : (Int?, Int?) = if unmapped { + (None, None) + } else { + align_sam_raw_hard_clips(cigar) + } + let raw_hard_left = raw_hard_clips.0 + let raw_hard_right = raw_hard_clips.1 + let hard_clip_left = if reverse { raw_hard_right } else { raw_hard_left } + let hard_clip_right = if reverse { raw_hard_left } else { raw_hard_right } + let query_sequence = if !sequence_known { + "" + } else if reverse { + align_sam_reverse_complement(raw_sequence) + } else { + raw_sequence + } + let qualities = if !qualities_known { + [] + } else if reverse { + align_sam_reverse_ints(raw_qualities) + } else { + raw_qualities + } + let mate_reference : String? = if raw_mate_reference == "*" { + None + } else if raw_mate_reference == "=" { + if reference_name == "*" { + align_sam_fail("RNEXT = requires a mapped reference") + } + Some(reference_name) + } else { + Some(raw_mate_reference) + } + AlignSamAlignment::{ + query_name, + flag, + reference_name, + reference_length, + target_coordinates, + query_coordinates, + operations, + cigar, + mapq: if raw_mapq == 255 { + None + } else { + Some(raw_mapq) + }, + mate_reference, + mate_position: if raw_mate_position == 0 { + None + } else { + Some(raw_mate_position - 1) + }, + template_length: if raw_template_length == 0 { + None + } else { + Some(raw_template_length) + }, + query_sequence, + sequence_known, + qualities, + qualities_known, + hard_clip_left, + hard_clip_right, + tags, + } +} + +///| +pub fn align_sam_parse(text : String) -> AlignSamDocument raise AlignSamError { + let headers : Array[AlignSamHeader] = [] + let references : Array[AlignSamReference] = [] + let alignments : Array[AlignSamAlignment] = [] + let reference_names : Map[String, Bool] = Map([]) + let mut saw_alignment = false + let mut saw_hd = false + for line in align_sam_split_lines(text) { + if line.length() == 0 { + continue + } + if line.unsafe_get(0).to_int() == '@'.to_int() { + if saw_alignment { + align_sam_fail("SAM headers must precede alignment records") + } + let header = align_sam_parse_header(line) + if header.record_type == "HD" { + if saw_hd || headers.length() > 0 { + align_sam_fail("@HD must be the first and only @HD record") + } + saw_hd = true + if align_sam_header_value(header, "VN") is None { + align_sam_fail("@HD header is missing VN") + } + } + if header.record_type == "SQ" { + let reference = align_sam_parse_reference(header) + if reference_names.contains(reference.name) { + align_sam_fail("duplicate @SQ reference " + reference.name) + } + reference_names[reference.name] = true + references.push(reference) + } + headers.push(header) + } else { + saw_alignment = true + alignments.push(align_sam_parse_alignment(line, references)) + } + } + AlignSamDocument::{ headers, references, alignments } +} + +///| +pub fn AlignSamDocument::reference( + self : AlignSamDocument, + name : String, +) -> AlignSamReference? { + for reference in self.references { + if reference.name == name { + return Some(AlignSamReference::{ + name: reference.name, + length: reference.length, + annotations: align_sam_copy_fields(reference.annotations), + }) + } + } + None +} + +///| +pub fn AlignSamAlignment::is_unmapped(self : AlignSamAlignment) -> Bool { + (self.flag & 0x4) != 0 +} + +///| +pub fn AlignSamAlignment::is_reverse(self : AlignSamAlignment) -> Bool { + (self.flag & 0x10) != 0 +} + +///| +pub fn AlignSamAlignment::is_paired(self : AlignSamAlignment) -> Bool { + (self.flag & 0x1) != 0 +} + +///| +pub fn AlignSamAlignment::is_secondary(self : AlignSamAlignment) -> Bool { + (self.flag & 0x100) != 0 +} + +///| +pub fn AlignSamAlignment::is_supplementary(self : AlignSamAlignment) -> Bool { + (self.flag & 0x800) != 0 +} + +///| +pub fn AlignSamAlignment::query_length(self : AlignSamAlignment) -> Int { + if self.sequence_known { + self.query_sequence.length() + } else { + align_sam_cigar_query_consumed(self.cigar) + } +} + +///| +pub fn AlignSamAlignment::reference_start(self : AlignSamAlignment) -> Int? { + if self.target_coordinates.length() == 0 { + None + } else { + Some(self.target_coordinates[0]) + } +} + +///| +pub fn AlignSamAlignment::reference_end(self : AlignSamAlignment) -> Int? { + if self.target_coordinates.length() == 0 { + None + } else { + Some(self.target_coordinates[self.target_coordinates.length() - 1]) + } +} + +///| +pub fn AlignSamAlignment::tag( + self : AlignSamAlignment, + name : String, +) -> AlignSamTag? { + for tag in self.tags { + if tag.name == name { + return Some(AlignSamTag::{ + name: tag.name, + value: align_sam_copy_tag_value(tag.value), + }) + } + } + None +} + +///| +pub fn AlignSamAlignment::integer_tag( + self : AlignSamAlignment, + name : String, +) -> Int? { + match self.tag(name) { + Some(tag) => + match tag.value { + SamTagInteger(value) => Some(value) + _ => None + } + None => None + } +} + +///| +pub fn AlignSamAlignment::string_tag( + self : AlignSamAlignment, + name : String, +) -> String? { + match self.tag(name) { + Some(tag) => + match tag.value { + SamTagCharacter(value) => Some(value) + SamTagString(value) => Some(value) + SamTagHex(value) => Some(value) + _ => None + } + None => None + } +} + +///| +pub fn AlignSamAlignment::stats(self : AlignSamAlignment) -> AlignSamStats { + let mut reference_consumed = 0 + let mut query_consumed = 0 + let mut aligned = 0 + let mut inserted = 0 + let mut deleted = 0 + let mut skipped = 0 + let mut soft_clipped = 0 + let mut hard_clipped = 0 + for element in self.cigar { + let length = element.length + if element.operation == "M" || + element.operation == "=" || + element.operation == "X" { + reference_consumed = reference_consumed + length + query_consumed = query_consumed + length + aligned = aligned + length + } else if element.operation == "I" { + query_consumed = query_consumed + length + inserted = inserted + length + } else if element.operation == "D" { + reference_consumed = reference_consumed + length + deleted = deleted + length + } else if element.operation == "N" { + reference_consumed = reference_consumed + length + skipped = skipped + length + } else if element.operation == "S" { + query_consumed = query_consumed + length + soft_clipped = soft_clipped + length + } else if element.operation == "H" { + hard_clipped = hard_clipped + length + } + } + AlignSamStats::{ + reference_consumed, + query_consumed, + aligned, + inserted, + deleted, + skipped, + soft_clipped, + hard_clipped, + } +} + +///| +fn align_sam_is_aligned_operation(operation : AlignSamOperation) -> Bool { + operation == SamAligned || operation == SamEqual || operation == SamMismatch +} + +///| +pub fn AlignSamAlignment::target_to_query( + self : AlignSamAlignment, + position : Int, +) -> Int? { + for index in 0..= target_start && position < target_end { + let query_start = self.query_coordinates[index] + let query_end = self.query_coordinates[index + 1] + let offset = position - target_start + if query_end >= query_start { + return Some(query_start + offset) + } else { + return Some(query_start - 1 - offset) + } + } + } + None +} + +///| +pub fn AlignSamAlignment::query_to_target( + self : AlignSamAlignment, + position : Int, +) -> Int? { + for index in 0..= query_start { + if position >= query_start && position < query_end { + return Some(self.target_coordinates[index] + position - query_start) + } + } else if position >= query_end && position < query_start { + return Some(self.target_coordinates[index] + query_start - 1 - position) + } + } + None +} + +///| +fn align_sam_raw_sequence(alignment : AlignSamAlignment) -> String { + if !alignment.sequence_known { + "*" + } else if alignment.is_reverse() { + align_sam_reverse_complement(alignment.query_sequence) + } else { + alignment.query_sequence + } +} + +///| +fn align_sam_md_positions( + alignment : AlignSamAlignment, +) -> (Array[Int], Array[String]) { + let positions : Array[Int] = [] + let defaults : Array[String] = [] + let raw_sequence = align_sam_raw_sequence(alignment) + let mut target = match alignment.reference_start() { + Some(value) => value + None => 0 + } + let mut query = 0 + for element in alignment.cigar { + let operation = element.operation + if operation == "S" { + query = query + element.length + } else if operation == "M" || operation == "=" || operation == "X" { + for offset in 0.. (Array[Int], Array[String]) raise AlignSamError { + let (positions, values) = align_sam_md_positions(alignment) + let md = match alignment.string_tag("MD") { + Some(value) => value + None => return (positions, values) + } + let mut pointer = 0 + let mut index = 0 + while index < md.length() { + let code = md.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + let start = index + while index < md.length() { + let digit = md.unsafe_get(index).to_int() + if digit < '0'.to_int() || digit > '9'.to_int() { + break + } + index = index + 1 + } + pointer = pointer + + align_sam_parse_int(md[start:index].to_owned(), "MD match count") + } else if code == '^'.to_int() { + index = index + 1 + while index < md.length() && + align_sam_is_md_base(md.unsafe_get(index).to_int()) { + if pointer >= values.length() { + align_sam_fail("MD deletion exceeds the CIGAR path") + } + values[pointer] = md[index:index + 1].to_owned() + pointer = pointer + 1 + index = index + 1 + } + } else { + if pointer >= values.length() { + align_sam_fail("MD mismatch exceeds the CIGAR path") + } + values[pointer] = md[index:index + 1].to_owned() + pointer = pointer + 1 + index = index + 1 + } + } + if pointer != values.length() { + align_sam_fail("MD tag does not cover the complete CIGAR reference path") + } + (positions, values) +} + +///| +fn align_sam_reference_char( + positions : Array[Int], + values : Array[String], + position : Int, +) -> String { + for index in 0.. (String, String) raise AlignSamError { + if self.is_unmapped() { + align_sam_fail("unmapped SAM records do not have aligned rows") + } + if reference_sequence.length() > 0 { + match self.reference_length { + Some(length) => + if reference_sequence.length() != length { + align_sam_fail("reference sequence length does not match @SQ LN") + } + None => () + } + } + let (md_positions, md_values) = align_sam_reference_from_md(self) + let raw_sequence = align_sam_raw_sequence(self) + let target_row = StringBuilder::new() + let query_row = StringBuilder::new() + let mut target = match self.reference_start() { + Some(value) => value + None => 0 + } + let mut query = 0 + for element in self.cigar { + let operation = element.operation + if operation == "S" { + query = query + element.length + } else if operation == "M" || operation == "=" || operation == "X" { + for offset in 0.. 0 { + target_row.write_string( + reference_sequence[target_position:target_position + 1].to_owned(), + ) + } else { + target_row.write_string( + align_sam_reference_char(md_positions, md_values, target_position), + ) + } + if self.sequence_known { + query_row.write_string( + raw_sequence[query + offset:query + offset + 1].to_owned(), + ) + } else { + query_row.write_char('?') + } + } + target = target + element.length + query = query + element.length + } else if operation == "I" { + target_row.write_string("-".repeat(element.length)) + if self.sequence_known { + query_row.write_string( + raw_sequence[query:query + element.length].to_owned(), + ) + } else { + query_row.write_string("?".repeat(element.length)) + } + query = query + element.length + } else if operation == "D" || operation == "N" { + for offset in 0.. 0 { + target_row.write_string( + reference_sequence[target_position:target_position + 1].to_owned(), + ) + } else { + target_row.write_string( + align_sam_reference_char(md_positions, md_values, target_position), + ) + } + } + query_row.write_string("-".repeat(element.length)) + target = target + element.length + } + } + (target_row.to_string(), query_row.to_string()) +} + +///| +pub fn AlignSamAlignment::computed_nm(self : AlignSamAlignment) -> Int? { + if self.is_unmapped() { + return None + } + let mut count = 0 + for element in self.cigar { + if element.operation == "I" || + element.operation == "D" || + element.operation == "X" { + count = count + element.length + } + } + match self.string_tag("MD") { + Some(md) => { + let mut index = 0 + let mut in_deletion = false + while index < md.length() { + let code = md.unsafe_get(index).to_int() + if code == '^'.to_int() { + in_deletion = true + } else if code >= '0'.to_int() && code <= '9'.to_int() { + in_deletion = false + } else if align_sam_is_md_base(code) && !in_deletion { + count = count + 1 + } + index = index + 1 + } + // X operations and MD mismatch letters describe the same substitutions. + let mut explicit_mismatches = 0 + for element in self.cigar { + if element.operation == "X" { + explicit_mismatches = explicit_mismatches + element.length + } + } + Some(count - explicit_mismatches) + } + None => { + for element in self.cigar { + if element.operation == "M" { + return None + } + } + Some(count) + } + } +} + +///| +pub fn align_sam_calculate_md( + alignment : AlignSamAlignment, + reference_sequence : String, +) -> String raise AlignSamError { + if alignment.is_unmapped() { + align_sam_fail("cannot calculate MD for an unmapped record") + } + match alignment.reference_length { + Some(length) => + if reference_sequence.length() != length { + align_sam_fail("reference sequence length does not match @SQ LN") + } + None => () + } + let raw_sequence = align_sam_raw_sequence(alignment) + if !alignment.sequence_known { + align_sam_fail("cannot calculate MD without a query sequence") + } + let mut target = match alignment.reference_start() { + Some(value) => value + None => 0 + } + let mut query = 0 + let mut matches = 0 + let output = StringBuilder::new() + for element in alignment.cigar { + let operation = element.operation + if operation == "S" { + query = query + element.length + } else if operation == "M" || operation == "=" || operation == "X" { + for offset in 0.. String raise AlignSamError { + if !alignment.qualities_known { + return "*" + } + let values = if alignment.is_reverse() { + align_sam_reverse_ints(alignment.qualities) + } else { + align_sam_copy_ints(alignment.qualities) + } + let output = StringBuilder::new() + for value in values { + if value < 0 || value > 93 { + align_sam_fail("PHRED quality must be between 0 and 93") + } + output.write_char((value + 33).unsafe_to_char()) + } + output.to_string() +} + +///| +pub fn align_sam_format( + alignment : AlignSamAlignment, + calculate_md? : Bool = false, + reference_sequence? : String = "", +) -> String raise AlignSamError { + let output = StringBuilder::new() + output.write_string(alignment.query_name) + output.write_char('\t') + output.write_string(alignment.flag.to_string()) + output.write_char('\t') + output.write_string(alignment.reference_name) + output.write_char('\t') + let position = match alignment.reference_start() { + Some(value) => value + 1 + None => 0 + } + output.write_string(position.to_string()) + output.write_char('\t') + output.write_string( + match alignment.mapq { + Some(value) => value.to_string() + None => "255" + }, + ) + output.write_char('\t') + output.write_string(align_sam_cigar_string(alignment.cigar)) + output.write_char('\t') + let mate_reference = match alignment.mate_reference { + Some(value) => + if value == alignment.reference_name && alignment.reference_name != "*" { + "=" + } else { + value + } + None => "*" + } + output.write_string(mate_reference) + output.write_char('\t') + output.write_string( + match alignment.mate_position { + Some(value) => (value + 1).to_string() + None => "0" + }, + ) + output.write_char('\t') + output.write_string( + match alignment.template_length { + Some(value) => value.to_string() + None => "0" + }, + ) + output.write_char('\t') + output.write_string(align_sam_raw_sequence(alignment)) + output.write_char('\t') + output.write_string(align_sam_quality_text(alignment)) + for tag in alignment.tags { + if calculate_md && tag.name == "MD" { + continue + } + output.write_char('\t') + output.write_string(align_sam_tag_text(tag)) + } + if calculate_md { + if reference_sequence.length() == 0 { + align_sam_fail("reference_sequence is required when calculate_md is true") + } + output.write_string("\tMD:Z:") + output.write_string(align_sam_calculate_md(alignment, reference_sequence)) + } + output.write_char('\n') + output.to_string() +} + +///| +fn align_sam_header_text(header : AlignSamHeader) -> String { + let output = StringBuilder::new() + output.write_char('@') + output.write_string(header.record_type) + if header.record_type == "CO" { + if header.comment.length() > 0 { + output.write_char('\t') + output.write_string(header.comment) + } + } else { + for field in header.fields { + output.write_char('\t') + output.write_string(field.key) + output.write_char(':') + output.write_string(field.value) + } + } + output.write_char('\n') + output.to_string() +} + +///| +pub fn align_sam_write( + document : AlignSamDocument, +) -> String raise AlignSamError { + let output = StringBuilder::new() + for header in document.headers { + output.write_string(align_sam_header_text(header)) + } + for alignment in document.alignments { + output.write_string(align_sam_format(alignment)) + } + output.to_string() +} + +///| +fn align_sam_format_tags(tags : Array[AlignSamTag]) -> String { + let output = StringBuilder::new() + for tag in tags { + output.write_char('\t') + output.write_string(align_sam_tag_text(tag)) + } + output.to_string() +} + +///| +fn align_sam_quality_from_values( + qualities : Array[Int], +) -> String raise AlignSamError { + let output = StringBuilder::new() + for value in qualities { + if value < 0 || value > 93 { + align_sam_fail("PHRED quality must be between 0 and 93") + } + output.write_char((value + 33).unsafe_to_char()) + } + output.to_string() +} + +///| +pub fn align_sam_create( + query_name : String, + reference_name : String, + reference_length : Int, + target_start : Int, + query_sequence : String, + cigar : String, + reverse? : Bool = false, + mapq? : Int = 255, + qualities? : Array[Int] = [], + tags? : Array[AlignSamTag] = [], + flag? : Int = 0, +) -> AlignSamAlignment raise AlignSamError { + if reference_length <= 0 { + align_sam_fail("reference_length must be positive") + } + if target_start < 0 { + align_sam_fail("target_start must be non-negative") + } + if (flag & 0x4) != 0 { + align_sam_fail("align_sam_create constructs mapped alignments") + } + let final_flag = if reverse { flag | 0x10 } else { flag & 0xffef } + let stored_sequence = if reverse { + align_sam_reverse_complement(query_sequence) + } else { + query_sequence + } + let stored_qualities = if reverse { + align_sam_reverse_ints(qualities) + } else { + align_sam_copy_ints(qualities) + } + if qualities.length() > 0 && qualities.length() != query_sequence.length() { + align_sam_fail("qualities length must equal query_sequence length") + } + let quality_text = if qualities.length() == 0 { + "*" + } else { + align_sam_quality_from_values(stored_qualities) + } + let text = "@SQ\tSN:" + + reference_name + + "\tLN:" + + reference_length.to_string() + + "\n" + + query_name + + "\t" + + final_flag.to_string() + + "\t" + + reference_name + + "\t" + + (target_start + 1).to_string() + + "\t" + + mapq.to_string() + + "\t" + + cigar + + "\t*\t0\t0\t" + + stored_sequence + + "\t" + + quality_text + + align_sam_format_tags(tags) + + "\n" + let document = align_sam_parse(text) + document.alignments[0] +} + +///| +pub fn AlignSamAlignment::summary(self : AlignSamAlignment) -> String { + if self.is_unmapped() { + return "SAM alignment " + self.query_name + ": unmapped" + } + let stats = self.stats() + let start = match self.reference_start() { + Some(value) => value.to_string() + None => "?" + } + let end = match self.reference_end() { + Some(value) => value.to_string() + None => "?" + } + "SAM alignment " + + self.query_name + + " -> " + + self.reference_name + + ":" + + start + + "-" + + end + + ", aligned=" + + stats.aligned.to_string() + + ", inserted=" + + stats.inserted.to_string() + + ", deleted=" + + stats.deleted.to_string() + + ", skipped=" + + stats.skipped.to_string() +} + +///| +pub fn align_sam_example_data() -> String { + "@HD\tVN:1.6\tSO:coordinate\n" + + "@SQ\tSN:chr1\tLN:40\tAS:demo\n" + + "@RG\tID:rg1\tSM:sample1\n" + + "read_forward\t0\tchr1\t5\t60\t3S4M1I2M1D3M2S\t*\t0\t0\tGGGACGTATCGTAAA\tIIIIIIIIIIIIIII\tNM:i:3\tMD:Z:6^G1A1\tAS:i:8\tRG:Z:rg1\n" + + "read_reverse\t16\tchr1\t20\t40\t2H2S5M3N3M1S1H\t*\t0\t0\tTACGTACGTAA\tJJJJJJJJJJJ\tNM:i:0\tMD:Z:8\tAS:i:8\n" + + "read_unmapped\t4\t*\t0\t0\t*\t*\t0\t0\tACGT\t!!!!\tRG:Z:rg1\n" +} diff --git a/src/align_stockholm.mbt b/src/align_stockholm.mbt new file mode 100644 index 00000000..9252ad95 --- /dev/null +++ b/src/align_stockholm.mbt @@ -0,0 +1,2426 @@ +// Biopython-compatible Bio.Align.stockholm support. +// +// This module is intentionally separate from stockholm.mbt. The older module +// exposes a permissive block-oriented API, while this file models modern, +// coordinate-aware Stockholm alignments and their typed annotations. + +///| +/// Error raised for malformed Stockholm data or invalid operations. +pub suberror AlignStockholmError { + AlignStockholmError(String) +} + +///| +/// A mapped GF or GS annotation. +/// +/// `code` is the Stockholm tag and `feature` is Biopython's descriptive name. +pub struct AlignStockholmAnnotation { + code : String + feature : String + value : String +} derive(Eq, Debug) + +///| +/// A mapped GC annotation with one character per retained alignment column. +pub struct AlignStockholmColumnAnnotation { + code : String + feature : String + value : String +} derive(Eq, Debug) + +///| +/// A mapped GR annotation. +/// +/// `value` contains one character per ungapped residue. `aligned_value` +/// contains one character per retained alignment column. +pub struct AlignStockholmLetterAnnotation { + code : String + feature : String + value : String + aligned_value : String +} derive(Eq, Debug) + +///| +/// One structured Stockholm reference. +pub struct AlignStockholmReference { + number : Int + medline : String + title : String + author : String + location : String + comment : String +} derive(Eq, Debug) + +///| +/// One structured database reference. +pub struct AlignStockholmDatabaseReference { + reference : String + comment : String +} derive(Eq, Debug) + +///| +/// One nested-domain annotation. +pub struct AlignStockholmNestedDomain { + accession : String + location : String +} derive(Eq, Debug) + +///| +/// One sequence row and its GS/GR annotations. +pub struct AlignStockholmSequence { + id : String + description : String + sequence : String + aligned_sequence : String + annotations : Array[AlignStockholmAnnotation] + dbxrefs : Array[String] + letter_annotations : Array[AlignStockholmLetterAnnotation] +} derive(Eq, Debug) + +///| +/// A coordinate-aware Stockholm 1.0 alignment. +/// +/// `operations` contains one `M`, `D`, or `I` per retained column. A source +/// `-` gap marks a deletion column and a source `.` gap marks an insertion +/// column, matching Biopython's `Bio.Align.stockholm` behavior. +pub struct AlignStockholmAlignment { + version : String + sequences : Array[AlignStockholmSequence] + annotations : Array[AlignStockholmAnnotation] + column_annotations : Array[AlignStockholmColumnAnnotation] + references : Array[AlignStockholmReference] + database_references : Array[AlignStockholmDatabaseReference] + nested_domains : Array[AlignStockholmNestedDomain] + operations : String + source_columns : Int + removed_all_gap_columns : Int +} derive(Eq, Debug) + +///| +/// Pairwise or all-pairs Stockholm alignment statistics. +pub struct AlignStockholmCounts { + pairs : Int + aligned : Int + identities : Int + mismatches : Int + gap_columns : Int + double_gap_columns : Int + gap_opens : Int +} derive(Eq, Debug) + +///| +priv struct AlignStockholmRawAnnotation { + code : String + value : String +} + +///| +priv struct AlignStockholmRawSequenceAnnotation { + sequence_id : String + code : String + value : String +} + +///| +priv struct AlignStockholmReferenceBuilder { + number : Int + mut medline : String + mut title : String + mut author : String + mut location : String + comment : String +} + +///| +priv struct AlignStockholmDatabaseReferenceBuilder { + reference : String + mut comment : String +} + +///| +priv struct AlignStockholmNestedDomainBuilder { + accession : String + mut location : String +} + +///| +/// Construct a generic GF or GS annotation. +pub fn AlignStockholmAnnotation::create( + code : String, + value : String, + feature? : String? = None, +) -> AlignStockholmAnnotation raise AlignStockholmError { + align_stockholm_validate_annotation_code(code) + AlignStockholmAnnotation::{ + code, + feature: match feature { + Some(name) => name + None => code + }, + value, + } +} + +///| +/// Construct a generic GC annotation. +pub fn AlignStockholmColumnAnnotation::create( + code : String, + value : String, + feature? : String? = None, +) -> AlignStockholmColumnAnnotation raise AlignStockholmError { + align_stockholm_validate_annotation_code(code) + AlignStockholmColumnAnnotation::{ + code, + feature: match feature { + Some(name) => name + None => align_stockholm_gc_feature(code) + }, + value, + } +} + +///| +/// Construct a generic GR annotation. +pub fn AlignStockholmLetterAnnotation::create( + code : String, + value : String, + aligned_value : String, + feature? : String? = None, +) -> AlignStockholmLetterAnnotation raise AlignStockholmError { + align_stockholm_validate_annotation_code(code) + AlignStockholmLetterAnnotation::{ + code, + feature: match feature { + Some(name) => name + None => align_stockholm_gr_feature(code) + }, + value, + aligned_value, + } +} + +///| +/// Construct one standalone Stockholm row. +pub fn AlignStockholmSequence::create( + id : String, + aligned_sequence : String, + description? : String = "", + annotations? : Array[AlignStockholmAnnotation] = [], + dbxrefs? : Array[String] = [], + letter_annotations? : Array[AlignStockholmLetterAnnotation] = [], +) -> AlignStockholmSequence raise AlignStockholmError { + align_stockholm_validate_id(id) + if aligned_sequence.length() == 0 { + raise AlignStockholmError("Stockholm aligned sequence must not be empty") + } + let normalized = align_stockholm_normalize_row(aligned_sequence) + let copied_annotations : Array[AlignStockholmAnnotation] = [] + for annotation in annotations { + copied_annotations.push(annotation) + } + let copied_dbxrefs : Array[String] = [] + for dbxref in dbxrefs { + copied_dbxrefs.push(dbxref) + } + let copied_letter_annotations : Array[AlignStockholmLetterAnnotation] = [] + for annotation in letter_annotations { + copied_letter_annotations.push(annotation) + } + let sequence = AlignStockholmSequence::{ + id, + description, + sequence: align_stockholm_remove_gaps(normalized), + aligned_sequence: normalized, + annotations: copied_annotations, + dbxrefs: copied_dbxrefs, + letter_annotations: copied_letter_annotations, + } + align_stockholm_validate_sequence_annotations(sequence) + sequence +} + +///| +/// Construct and validate one coordinate-aware Stockholm alignment. +pub fn AlignStockholmAlignment::create( + sequences : Array[AlignStockholmSequence], + annotations? : Array[AlignStockholmAnnotation] = [], + column_annotations? : Array[AlignStockholmColumnAnnotation] = [], + references? : Array[AlignStockholmReference] = [], + database_references? : Array[AlignStockholmDatabaseReference] = [], + nested_domains? : Array[AlignStockholmNestedDomain] = [], + operations? : String = "", + source_columns? : Int = 0, + removed_all_gap_columns? : Int = 0, +) -> AlignStockholmAlignment raise AlignStockholmError { + let copied_sequences : Array[AlignStockholmSequence] = [] + for sequence in sequences { + copied_sequences.push(sequence) + } + let copied_annotations : Array[AlignStockholmAnnotation] = [] + for annotation in annotations { + copied_annotations.push(annotation) + } + let copied_column_annotations : Array[AlignStockholmColumnAnnotation] = [] + for annotation in column_annotations { + copied_column_annotations.push(annotation) + } + let copied_references : Array[AlignStockholmReference] = [] + for reference in references { + copied_references.push(reference) + } + let copied_database_references : Array[AlignStockholmDatabaseReference] = [] + for reference in database_references { + copied_database_references.push(reference) + } + let copied_nested_domains : Array[AlignStockholmNestedDomain] = [] + for domain in nested_domains { + copied_nested_domains.push(domain) + } + let width = if copied_sequences.length() == 0 { + 0 + } else { + copied_sequences[0].aligned_sequence.length() + } + let actual_operations = if operations.length() == 0 && width > 0 { + align_stockholm_infer_operations(copied_sequences) + } else { + operations + } + let actual_source_columns = if source_columns == 0 { + width + removed_all_gap_columns + } else { + source_columns + } + let alignment = AlignStockholmAlignment::{ + version: "1.0", + sequences: copied_sequences, + annotations: copied_annotations, + column_annotations: copied_column_annotations, + references: copied_references, + database_references: copied_database_references, + nested_domains: copied_nested_domains, + operations: actual_operations, + source_columns: actual_source_columns, + removed_all_gap_columns, + } + align_stockholm_validate_alignment(alignment) + alignment +} + +///| +/// Construct an alignment from identifiers and printed rows. +/// +/// Source `.` and `-` gaps retain their insertion/deletion operation class. +/// Columns containing only gaps are removed. +pub fn align_stockholm_from_aligned( + ids : Array[String], + aligned_sequences : Array[String], +) -> AlignStockholmAlignment raise AlignStockholmError { + if ids.length() == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one sequence", + ) + } + if ids.length() != aligned_sequences.length() { + raise AlignStockholmError( + "Stockholm identifiers and aligned rows must have equal lengths", + ) + } + let width = aligned_sequences[0].length() + if width == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one column", + ) + } + for row in aligned_sequences { + if row.length() != width { + raise AlignStockholmError("Stockholm aligned rows must have equal widths") + } + } + let prepared = align_stockholm_prepare_rows(aligned_sequences) + if prepared.0.length() == 0 || prepared.0[0].length() == 0 { + raise AlignStockholmError( + "Stockholm alignment contains no non-gap alignment column", + ) + } + let sequences : Array[AlignStockholmSequence] = [] + for index = 0; index < ids.length(); index = index + 1 { + sequences.push( + AlignStockholmSequence::create(ids[index], prepared.0[index]), + ) + } + AlignStockholmAlignment::create( + sequences, + operations=prepared.1, + source_columns=width, + removed_all_gap_columns=align_stockholm_count_true(prepared.2), + ) +} + +///| +/// Parse exactly one Stockholm 1.0 alignment. +pub fn align_stockholm_parse( + text : String, +) -> AlignStockholmAlignment raise AlignStockholmError { + let alignments = align_stockholm_parse_all(text) + if alignments.length() == 0 { + raise AlignStockholmError("Stockholm input contains no alignment") + } + if alignments.length() != 1 { + raise AlignStockholmError( + "Expected one Stockholm alignment, found " + + alignments.length().to_string(), + ) + } + alignments[0] +} + +///| +/// Parse all concatenated Stockholm 1.0 alignments. +pub fn align_stockholm_parse_all( + text : String, +) -> Array[AlignStockholmAlignment] raise AlignStockholmError { + let lines = align_stockholm_normalize_lines(text) + let alignments : Array[AlignStockholmAlignment] = [] + let mut index = 0 + while index < lines.length() { + while index < lines.length() && lines[index].trim().length() == 0 { + index = index + 1 + } + if index >= lines.length() { + break + } + if lines[index].trim() != "# STOCKHOLM 1.0" { + raise AlignStockholmError( + "Expected '# STOCKHOLM 1.0' at line " + (index + 1).to_string(), + ) + } + let start = index + 1 + index = start + while index < lines.length() && lines[index].trim() != "//" { + if lines[index].trim() == "# STOCKHOLM 1.0" { + raise AlignStockholmError( + "Stockholm alignment is missing its // terminator", + ) + } + index = index + 1 + } + if index >= lines.length() { + raise AlignStockholmError( + "Stockholm alignment is missing its // terminator", + ) + } + alignments.push(align_stockholm_parse_record(lines, start, index)) + index = index + 1 + } + alignments +} + +///| +/// Write one canonical Stockholm 1.0 alignment. +pub fn align_stockholm_write( + alignment : AlignStockholmAlignment, +) -> String raise AlignStockholmError { + align_stockholm_validate_alignment(alignment) + let rows = alignment.num_sequences() + let columns = alignment.alignment_length() + if rows == 0 { + raise AlignStockholmError("Must have at least one Stockholm sequence") + } + if columns == 0 { + raise AlignStockholmError("Non-empty Stockholm sequences are required") + } + let output = StringBuilder::new() + output.write_string("# STOCKHOLM 1.0\n") + for annotation in alignment.annotations { + if !align_stockholm_is_known_gf(annotation.code) { + raise AlignStockholmError( + "Unknown Stockholm GF annotation " + annotation.code, + ) + } + if annotation.code == "CC" { + output.write_string( + align_stockholm_wrap_text("#=GF CC ", annotation.value), + ) + } else { + output.write_string( + "#=GF " + annotation.code + " " + annotation.value + "\n", + ) + } + } + for domain in alignment.nested_domains { + if domain.accession.length() > 0 { + output.write_string("#=GF NE " + domain.accession + "\n") + } + if domain.location.length() > 0 { + output.write_string("#=GF NL " + domain.location + "\n") + } + } + for reference in alignment.references { + if reference.comment.length() > 0 { + output.write_string( + align_stockholm_wrap_text("#=GF RC ", reference.comment), + ) + } + output.write_string("#=GF RN [" + reference.number.to_string() + "]\n") + if reference.medline.length() > 0 { + output.write_string("#=GF RM " + reference.medline + "\n") + } + if reference.title.length() > 0 { + output.write_string( + align_stockholm_wrap_text("#=GF RT ", reference.title), + ) + } + if reference.author.length() > 0 { + output.write_string("#=GF RA " + reference.author + "\n") + } + if reference.location.length() > 0 { + output.write_string("#=GF RL " + reference.location + "\n") + } + } + for reference in alignment.database_references { + output.write_string("#=GF DR " + reference.reference + "\n") + if reference.comment.length() > 0 { + output.write_string("#=GF DC " + reference.comment + "\n") + } + } + output.write_string("#=GF SQ " + rows.to_string() + "\n") + let mut name_width = 0 + for sequence in alignment.sequences { + if sequence.id.length() > name_width { + name_width = sequence.id.length() + } + } + let start = align_stockholm_max(name_width, 20) + 12 + for sequence in alignment.sequences { + let padded_name = align_stockholm_pad_right(sequence.id, name_width) + for annotation in sequence.annotations { + output.write_string( + "#=GS " + + padded_name + + " " + + annotation.code + + " " + + annotation.value + + "\n", + ) + } + if sequence.description.length() > 0 { + output.write_string( + "#=GS " + padded_name + " DE " + sequence.description + "\n", + ) + } + for dbxref in sequence.dbxrefs { + output.write_string("#=GS " + padded_name + " DR " + dbxref + "\n") + } + } + for row = 0; row < alignment.sequences.length(); row = row + 1 { + let sequence = alignment.sequences[row] + output.write_string(align_stockholm_pad_right(sequence.id, start)) + output.write_string( + align_stockholm_render_row( + sequence.aligned_sequence, + alignment.operations, + ), + ) + output.write_char('\n') + let padded_name = align_stockholm_pad_right(sequence.id, name_width) + for annotation in sequence.letter_annotations { + output.write_string( + align_stockholm_pad_right( + "#=GR " + padded_name + " " + annotation.code + " ", + start, + ), + ) + output.write_string( + align_stockholm_render_letter_annotation(sequence, annotation.value), + ) + output.write_char('\n') + } + } + for annotation in alignment.column_annotations { + output.write_string( + align_stockholm_pad_right("#=GC " + annotation.code + " ", start), + ) + output.write_string(annotation.value) + output.write_char('\n') + } + output.write_string("//\n") + output.to_string() +} + +///| +/// Write concatenated Stockholm records. +pub fn align_stockholm_write_all( + alignments : Array[AlignStockholmAlignment], +) -> String raise AlignStockholmError { + let output = StringBuilder::new() + for alignment in alignments { + output.write_string(align_stockholm_write(alignment)) + } + output.to_string() +} + +///| +/// Return the number of rows. +pub fn AlignStockholmAlignment::num_sequences( + self : AlignStockholmAlignment, +) -> Int { + self.sequences.length() +} + +///| +/// Return the retained alignment width. +pub fn AlignStockholmAlignment::alignment_length( + self : AlignStockholmAlignment, +) -> Int { + if self.sequences.length() == 0 { + 0 + } else { + self.sequences[0].aligned_sequence.length() + } +} + +///| +/// Return the source width before all-gap columns were removed. +pub fn AlignStockholmAlignment::source_alignment_length( + self : AlignStockholmAlignment, +) -> Int { + self.source_columns +} + +///| +/// Find a sequence by exact identifier. +pub fn AlignStockholmAlignment::find_sequence( + self : AlignStockholmAlignment, + id : String, +) -> Int? { + for index = 0; index < self.sequences.length(); index = index + 1 { + if self.sequences[index].id == id { + return Some(index) + } + } + None +} + +///| +/// Return all GF values matching a Stockholm code or mapped feature. +pub fn AlignStockholmAlignment::annotation_values( + self : AlignStockholmAlignment, + name : String, +) -> Array[String] { + let values : Array[String] = [] + for annotation in self.annotations { + if annotation.code == name || annotation.feature == name { + values.push(annotation.value) + } + } + values +} + +///| +/// Return the first GF value matching a Stockholm code or mapped feature. +pub fn AlignStockholmAlignment::annotation( + self : AlignStockholmAlignment, + name : String, +) -> String? { + for annotation in self.annotations { + if annotation.code == name || annotation.feature == name { + return Some(annotation.value) + } + } + None +} + +///| +/// Return a GC value matching a Stockholm code or mapped feature. +pub fn AlignStockholmAlignment::column_annotation( + self : AlignStockholmAlignment, + name : String, +) -> String? { + for annotation in self.column_annotations { + if annotation.code == name || annotation.feature == name { + return Some(annotation.value) + } + } + None +} + +///| +/// Return the first GS value matching a Stockholm code or mapped feature. +pub fn AlignStockholmSequence::annotation( + self : AlignStockholmSequence, + name : String, +) -> String? { + for annotation in self.annotations { + if annotation.code == name || annotation.feature == name { + return Some(annotation.value) + } + } + None +} + +///| +/// Return a residue-level GR value matching a code or mapped feature. +pub fn AlignStockholmSequence::letter_annotation( + self : AlignStockholmSequence, + name : String, +) -> String? { + for annotation in self.letter_annotations { + if annotation.code == name || annotation.feature == name { + return Some(annotation.value) + } + } + None +} + +///| +/// Return one printed alignment column. +pub fn AlignStockholmAlignment::column( + self : AlignStockholmAlignment, + column : Int, +) -> String? { + if column < 0 || column >= self.alignment_length() { + return None + } + let output = StringBuilder::new(size_hint=self.sequences.length()) + for sequence in self.sequences { + output.write_char( + sequence.aligned_sequence.unsafe_get(column).unsafe_to_char(), + ) + } + Some(output.to_string()) +} + +///| +/// Map a zero-based ungapped sequence position to an alignment column. +pub fn AlignStockholmAlignment::sequence_position_to_column( + self : AlignStockholmAlignment, + row : Int, + position : Int, +) -> Int? { + if row < 0 || row >= self.sequences.length() || position < 0 { + return None + } + let aligned = self.sequences[row].aligned_sequence + let mut coordinate = 0 + for column = 0; column < aligned.length(); column = column + 1 { + if aligned.unsafe_get(column).to_int() != '-'.to_int() { + if coordinate == position { + return Some(column) + } + coordinate = coordinate + 1 + } + } + None +} + +///| +/// Map a retained column to a zero-based ungapped sequence position. +pub fn AlignStockholmAlignment::column_to_sequence_position( + self : AlignStockholmAlignment, + row : Int, + column : Int, +) -> Int? { + if row < 0 || + row >= self.sequences.length() || + column < 0 || + column >= self.alignment_length() { + return None + } + let aligned = self.sequences[row].aligned_sequence + if aligned.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut coordinate = 0 + for index = 0; index < column; index = index + 1 { + if aligned.unsafe_get(index).to_int() != '-'.to_int() { + coordinate = coordinate + 1 + } + } + Some(coordinate) +} + +///| +/// Map one residue through the alignment to another row. +pub fn AlignStockholmAlignment::map_position( + self : AlignStockholmAlignment, + source_row : Int, + target_row : Int, + source_position : Int, +) -> Int? { + if target_row < 0 || target_row >= self.sequences.length() { + return None + } + match self.sequence_position_to_column(source_row, source_position) { + Some(column) => self.column_to_sequence_position(target_row, column) + None => None + } +} + +///| +/// Return per-column residue coordinates for two rows. +pub fn AlignStockholmAlignment::aligned_pairs( + self : AlignStockholmAlignment, + first_row : Int, + second_row : Int, +) -> Array[(Int?, Int?)] raise AlignStockholmError { + align_stockholm_validate_row(self, first_row) + align_stockholm_validate_row(self, second_row) + let pairs : Array[(Int?, Int?)] = [] + let mut first_position = 0 + let mut second_position = 0 + for column = 0; column < self.alignment_length(); column = column + 1 { + let first_gap = self.sequences[first_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + let second_gap = self.sequences[second_row].aligned_sequence + .unsafe_get(column) + .to_int() == + '-'.to_int() + pairs.push( + ( + if first_gap { + None + } else { + Some(first_position) + }, + if second_gap { + None + } else { + Some(second_position) + }, + ), + ) + if !first_gap { + first_position = first_position + 1 + } + if !second_gap { + second_position = second_position + 1 + } + } + pairs +} + +///| +/// Return a compact Biopython-style coordinate path for all rows. +pub fn AlignStockholmAlignment::coordinate_path( + self : AlignStockholmAlignment, +) -> Array[Array[Int]] { + let paths : Array[Array[Int]] = [] + let coordinates = Array::make(self.sequences.length(), 0) + for _row = 0; _row < self.sequences.length(); _row = _row + 1 { + paths.push([0]) + } + let width = self.alignment_length() + for column = 0; column < width; column = column + 1 { + for row = 0; row < self.sequences.length(); row = row + 1 { + if self.sequences[row].aligned_sequence.unsafe_get(column).to_int() != + '-'.to_int() { + coordinates[row] = coordinates[row] + 1 + } + } + let boundary = if column + 1 == width { + true + } else { + align_stockholm_movement_changes(self, column, column + 1) + } + if boundary { + for row = 0; row < self.sequences.length(); row = row + 1 { + paths[row].push(coordinates[row]) + } + } + } + paths +} + +///| +/// Slice retained alignment columns and all GC/GR annotations together. +pub fn AlignStockholmAlignment::slice_columns( + self : AlignStockholmAlignment, + start : Int, + end : Int, +) -> AlignStockholmAlignment raise AlignStockholmError { + if start < 0 || end <= start || end > self.alignment_length() { + raise AlignStockholmError("Invalid Stockholm column slice") + } + let sequences : Array[AlignStockholmSequence] = [] + for sequence in self.sequences { + let aligned = sequence.aligned_sequence[start:end].to_owned() + let letter_annotations : Array[AlignStockholmLetterAnnotation] = [] + for annotation in sequence.letter_annotations { + let aligned_value = annotation.aligned_value[start:end].to_owned() + let value = align_stockholm_residue_annotation_value( + aligned, aligned_value, + ) + letter_annotations.push(AlignStockholmLetterAnnotation::{ + code: annotation.code, + feature: annotation.feature, + value, + aligned_value, + }) + } + sequences.push(AlignStockholmSequence::{ + id: sequence.id, + description: sequence.description, + sequence: align_stockholm_remove_gaps(aligned), + aligned_sequence: aligned, + annotations: sequence.annotations, + dbxrefs: sequence.dbxrefs, + letter_annotations, + }) + } + let column_annotations : Array[AlignStockholmColumnAnnotation] = [] + for annotation in self.column_annotations { + column_annotations.push(AlignStockholmColumnAnnotation::{ + code: annotation.code, + feature: annotation.feature, + value: annotation.value[start:end].to_owned(), + }) + } + AlignStockholmAlignment::create( + sequences, + annotations=self.annotations, + column_annotations~, + references=self.references, + database_references=self.database_references, + nested_domains=self.nested_domains, + operations=self.operations[start:end].to_owned(), + source_columns=end - start, + ) +} + +///| +/// Compute statistics for one pair of rows. +pub fn AlignStockholmAlignment::pair_counts( + self : AlignStockholmAlignment, + first_row : Int, + second_row : Int, +) -> AlignStockholmCounts raise AlignStockholmError { + align_stockholm_validate_row(self, first_row) + align_stockholm_validate_row(self, second_row) + align_stockholm_count_pair(self, first_row, second_row) +} + +///| +/// Aggregate statistics across every unordered row pair. +pub fn AlignStockholmAlignment::counts( + self : AlignStockholmAlignment, +) -> AlignStockholmCounts { + let mut result = AlignStockholmCounts::{ + pairs: 0, + aligned: 0, + identities: 0, + mismatches: 0, + gap_columns: 0, + double_gap_columns: 0, + gap_opens: 0, + } + for first = 0; first < self.sequences.length(); first = first + 1 { + for second = first + 1 + second < self.sequences.length() + second = second + 1 { + let counts = align_stockholm_count_pair(self, first, second) + result = AlignStockholmCounts::{ + pairs: result.pairs + counts.pairs, + aligned: result.aligned + counts.aligned, + identities: result.identities + counts.identities, + mismatches: result.mismatches + counts.mismatches, + gap_columns: result.gap_columns + counts.gap_columns, + double_gap_columns: result.double_gap_columns + + counts.double_gap_columns, + gap_opens: result.gap_opens + counts.gap_opens, + } + } + } + result +} + +///| +/// Return identities divided by aligned residue pairs. +pub fn AlignStockholmCounts::identity(self : AlignStockholmCounts) -> Double { + if self.aligned == 0 { + 0.0 + } else { + self.identities.to_double() / self.aligned.to_double() + } +} + +///| +/// Calculate a majority consensus; gaps do not vote. +pub fn AlignStockholmAlignment::consensus( + self : AlignStockholmAlignment, + minimum_fraction? : Double = 0.5, + ambiguous? : String = "X", +) -> String raise AlignStockholmError { + if minimum_fraction < 0.0 || minimum_fraction > 1.0 { + raise AlignStockholmError( + "Stockholm consensus minimum fraction must be between 0 and 1", + ) + } + if ambiguous.length() != 1 { + raise AlignStockholmError( + "Stockholm consensus ambiguous symbol must be one character", + ) + } + let output = StringBuilder::new(size_hint=self.alignment_length()) + for column = 0; column < self.alignment_length(); column = column + 1 { + let symbols : Array[Int] = [] + let counts : Array[Int] = [] + let mut residues = 0 + for sequence in self.sequences { + let code = sequence.aligned_sequence.unsafe_get(column).to_int() + if code != '-'.to_int() { + residues = residues + 1 + let normalized = align_stockholm_upper_code(code) + let mut found = -1 + for index = 0; index < symbols.length(); index = index + 1 { + if symbols[index] == normalized { + found = index + break + } + } + if found < 0 { + symbols.push(normalized) + counts.push(1) + } else { + counts[found] = counts[found] + 1 + } + } + } + if residues == 0 { + output.write_string(ambiguous) + continue + } + let mut best = 0 + for index = 1; index < counts.length(); index = index + 1 { + if counts[index] > counts[best] || + (counts[index] == counts[best] && symbols[index] < symbols[best]) { + best = index + } + } + if counts[best].to_double() / residues.to_double() >= minimum_fraction { + output.write_char(symbols[best].unsafe_to_char()) + } else { + output.write_string(ambiguous) + } + } + output.to_string() +} + +///| +/// Return per-column non-gap occupancy. +pub fn AlignStockholmAlignment::occupancy( + self : AlignStockholmAlignment, +) -> Array[Double] { + let values : Array[Double] = [] + for column = 0; column < self.alignment_length(); column = column + 1 { + let mut residues = 0 + for sequence in self.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + residues = residues + 1 + } + } + values.push( + if self.sequences.length() == 0 { + 0.0 + } else { + residues.to_double() / self.sequences.length().to_double() + }, + ) + } + values +} + +///| +/// Return a compact alignment summary. +pub fn AlignStockholmAlignment::summary( + self : AlignStockholmAlignment, +) -> String { + "AlignStockholmAlignment(version=" + + self.version + + ", sequences=" + + self.num_sequences().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ", source_columns=" + + self.source_columns.to_string() + + ", removed_all_gap=" + + self.removed_all_gap_columns.to_string() + + ")" +} + +///| +/// Official Biopython HAT fixture used by the example and black-box tests. +pub fn align_stockholm_example_text() -> String { + "# STOCKHOLM 1.0\n" + + "#=GF ID HAT\n" + + "#=GF AC PF02184.18\n" + + "#=GF DE HAT (Half-A-TPR) repeat\n" + + "#=GF AU SMART;\n" + + "#=GF SE Alignment kindly provided by SMART\n" + + "#=GF GA 21.00 21.00;\n" + + "#=GF TC 21.00 21.00;\n" + + "#=GF NC 20.90 20.90;\n" + + "#=GF BM hmmbuild HMM.ann SEED.ann\n" + + "#=GF SM hmmsearch -Z 57096847 -E 1000 --cpu 4 HMM pfamseq\n" + + "#=GF TP Repeat\n" + + "#=GF CL CL0020\n" + + "#=GF RN [1]\n" + + "#=GF RM 9478129\n" + + "#=GF RT The HAT helix, a repetitive motif implicated in RNA processing.\n" + + "#=GF RA Preker PJ, Keller W;\n" + + "#=GF RL Trends Biochem Sci 1998;23:15-16.\n" + + "#=GF DR INTERPRO; IPR003107;\n" + + "#=GF DR SMART; HAT;\n" + + "#=GF DR SO; 0001068; polypeptide_repeat;\n" + + "#=GF CC The HAT (Half A TPR) repeat is found in several RNA processing\n" + + "#=GF CC proteins [1].\n" + + "#=GF SQ 3\n" + + "#=GS CRN_DROME/191-222 AC P17886.2\n" + + "#=GS CLF1_SCHPO/185-216 AC P87312.1\n" + + "#=GS CLF1_SCHPO/185-216 DR PDB; 3JB9 R; 185-216;\n" + + "#=GS O16376_CAEEL/201-233 AC O16376.2\n" + + "CRN_DROME/191-222 KEIDRAREIYERFVYVH.PDVKNWIKFARFEES\n" + + "CLF1_SCHPO/185-216 HENERARGIYERFVVVH.PEVTNWLRWARFEEE\n" + + "#=GR CLF1_SCHPO/185-216 SS --HHHHHHHHHHHHHHS.--HHHHHHHHHHHHH\n" + + "O16376_CAEEL/201-233 KEIDRARSVYQRFLHVHGINVQNWIKYAKFEER\n" + + "#=GC SS_cons --HHHHHHHHHHHHHHS.--HHHHHHHHHHHHH\n" + + "#=GC seq_cons KEIDRARuIYERFVaVH.P-VpNWIKaARFEEc\n" + + "//\n" +} + +///| +fn align_stockholm_parse_record( + lines : Array[String], + start : Int, + end : Int, +) -> AlignStockholmAlignment raise AlignStockholmError { + let ids : Array[String] = [] + let rows : Array[String] = [] + let gf : Array[AlignStockholmRawAnnotation] = [] + let gs : Array[AlignStockholmRawSequenceAnnotation] = [] + let gc : Array[AlignStockholmRawAnnotation] = [] + let gr : Array[AlignStockholmRawSequenceAnnotation] = [] + let mut last_sequence_id = "" + for line_index = start; line_index < end; line_index = line_index + 1 { + let line = lines[line_index].trim().to_owned() + if line.length() == 0 { + continue + } + if align_stockholm_starts_with(line, "#=GF") { + let rest = align_stockholm_markup_rest(line, "#=GF", line_index) + let fields = align_stockholm_take_fields(rest, 1) + if fields.0.length() != 1 || fields.1.length() == 0 { + raise AlignStockholmError( + "Malformed GF annotation at line " + (line_index + 1).to_string(), + ) + } + gf.push(AlignStockholmRawAnnotation::{ + code: fields.0[0], + value: fields.1, + }) + continue + } + if align_stockholm_starts_with(line, "#=GS") { + let rest = align_stockholm_markup_rest(line, "#=GS", line_index) + let fields = align_stockholm_take_fields(rest, 2) + if fields.0.length() != 2 { + raise AlignStockholmError( + "Malformed GS annotation at line " + (line_index + 1).to_string(), + ) + } + gs.push(AlignStockholmRawSequenceAnnotation::{ + sequence_id: fields.0[0], + code: fields.0[1], + value: fields.1, + }) + continue + } + if align_stockholm_starts_with(line, "#=GC") { + let rest = align_stockholm_markup_rest(line, "#=GC", line_index) + let fields = align_stockholm_take_fields(rest, 1) + if fields.0.length() != 1 || fields.1.length() == 0 { + raise AlignStockholmError( + "Malformed GC annotation at line " + (line_index + 1).to_string(), + ) + } + align_stockholm_append_raw(gc, fields.0[0], fields.1) + continue + } + if align_stockholm_starts_with(line, "#=GR") { + let rest = align_stockholm_markup_rest(line, "#=GR", line_index) + let fields = align_stockholm_take_fields(rest, 2) + if fields.0.length() != 2 || fields.1.length() == 0 { + raise AlignStockholmError( + "Malformed GR annotation at line " + (line_index + 1).to_string(), + ) + } + if last_sequence_id.length() == 0 || fields.0[0] != last_sequence_id { + raise AlignStockholmError("GR annotation must follow its sequence row") + } + for annotation in gr { + if annotation.sequence_id == fields.0[0] && + annotation.code == fields.0[1] { + raise AlignStockholmError( + "Duplicate GR annotation " + fields.0[1] + " for " + fields.0[0], + ) + } + } + gr.push(AlignStockholmRawSequenceAnnotation::{ + sequence_id: fields.0[0], + code: fields.0[1], + value: fields.1, + }) + continue + } + if line.unsafe_get(0).to_int() == '#'.to_int() { + continue + } + let tokens = align_stockholm_split_whitespace(line) + if tokens.length() != 2 { + raise AlignStockholmError( + "Could not split Stockholm sequence line " + + (line_index + 1).to_string() + + " into identifier and aligned sequence", + ) + } + align_stockholm_validate_id(tokens[0]) + for existing in ids { + if existing == tokens[0] { + raise AlignStockholmError( + "Duplicate Stockholm sequence identifier " + tokens[0], + ) + } + } + if rows.length() > 0 && tokens[1].length() != rows[0].length() { + raise AlignStockholmError( + "Aligned sequence " + + tokens[0] + + " consists of " + + tokens[1].length().to_string() + + " letters, expected " + + rows[0].length().to_string() + + " letters", + ) + } + if tokens[1].length() == 0 { + raise AlignStockholmError("Stockholm sequence row must not be empty") + } + align_stockholm_validate_printed_row(tokens[1]) + ids.push(tokens[0]) + rows.push(tokens[1]) + last_sequence_id = tokens[0] + } + if ids.length() == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one sequence", + ) + } + let prepared = align_stockholm_prepare_rows(rows) + if prepared.0[0].length() == 0 { + raise AlignStockholmError( + "Stockholm alignment contains no non-gap alignment column", + ) + } + let annotations = align_stockholm_build_gf(gf, ids.length()) + let references = align_stockholm_build_references(gf) + let database_references = align_stockholm_build_database_references(gf) + let nested_domains = align_stockholm_build_nested_domains(gf) + let column_annotations = align_stockholm_build_gc( + gc, + rows[0].length(), + prepared.2, + ) + let sequences : Array[AlignStockholmSequence] = [] + for row_index = 0; row_index < ids.length(); row_index = row_index + 1 { + let sequence_annotations : Array[AlignStockholmAnnotation] = [] + let dbxrefs : Array[String] = [] + let mut description = "" + for annotation in gs { + if annotation.sequence_id != ids[row_index] { + continue + } + if annotation.code == "DR" { + dbxrefs.push(annotation.value) + continue + } + if annotation.code == "DE" { + if description.length() > 0 { + raise AlignStockholmError( + "Duplicate GS DE annotation for " + ids[row_index], + ) + } + description = annotation.value + continue + } + for existing in sequence_annotations { + if existing.code == annotation.code { + raise AlignStockholmError( + "Duplicate GS annotation " + + annotation.code + + " for " + + ids[row_index], + ) + } + } + sequence_annotations.push(AlignStockholmAnnotation::{ + code: annotation.code, + feature: align_stockholm_gs_feature(annotation.code), + value: annotation.value, + }) + } + let letter_annotations : Array[AlignStockholmLetterAnnotation] = [] + for annotation in gr { + if annotation.sequence_id != ids[row_index] { + continue + } + if annotation.value.length() != rows[0].length() { + raise AlignStockholmError( + "GR " + + annotation.code + + " length is " + + annotation.value.length().to_string() + + ", expected " + + rows[0].length().to_string(), + ) + } + let aligned_value = align_stockholm_remove_columns( + annotation.value, + prepared.2, + ) + let value = if annotation.code == "CSA" { + align_stockholm_remove_character(annotation.value, '-'.to_int()) + } else { + align_stockholm_remove_character(annotation.value, '.'.to_int()) + } + let ungapped = align_stockholm_remove_gaps(prepared.0[row_index]) + if value.length() != ungapped.length() { + raise AlignStockholmError( + "GR " + + annotation.code + + " does not contain one value per residue for " + + ids[row_index], + ) + } + letter_annotations.push(AlignStockholmLetterAnnotation::{ + code: annotation.code, + feature: align_stockholm_gr_feature(annotation.code), + value, + aligned_value, + }) + } + sequences.push(AlignStockholmSequence::{ + id: ids[row_index], + description, + sequence: align_stockholm_remove_gaps(prepared.0[row_index]), + aligned_sequence: prepared.0[row_index], + annotations: sequence_annotations, + dbxrefs, + letter_annotations, + }) + } + for annotation in gs { + if align_stockholm_find_string(ids, annotation.sequence_id) < 0 { + raise AlignStockholmError( + "Failed to find GS sequence " + annotation.sequence_id, + ) + } + } + for annotation in gr { + if align_stockholm_find_string(ids, annotation.sequence_id) < 0 { + raise AlignStockholmError( + "Failed to find GR sequence " + annotation.sequence_id, + ) + } + } + AlignStockholmAlignment::create( + sequences, + annotations~, + column_annotations~, + references~, + database_references~, + nested_domains~, + operations=prepared.1, + source_columns=rows[0].length(), + removed_all_gap_columns=align_stockholm_count_true(prepared.2), + ) +} + +///| +fn align_stockholm_build_gf( + raw : Array[AlignStockholmRawAnnotation], + rows : Int, +) -> Array[AlignStockholmAnnotation] raise AlignStockholmError { + let annotations : Array[AlignStockholmAnnotation] = [] + let codes = [ + "ID", "AC", "DE", "AU", "SE", "SS", "GA", "TC", "NC", "BM", "SM", "TP", "PI", + "CL", "WK", "CB", "**", "CC", + ] + for code in codes { + let values : Array[String] = [] + for annotation in raw { + if annotation.code == code { + values.push(annotation.value) + } + } + if values.length() == 0 { + continue + } + if code == "AU" { + for value in values { + annotations.push(AlignStockholmAnnotation::{ + code, + feature: align_stockholm_gf_feature(code), + value, + }) + } + } else if code == "WK" { + let merged = align_stockholm_merge_wikipedia(values) + for value in merged { + annotations.push(AlignStockholmAnnotation::{ + code, + feature: align_stockholm_gf_feature(code), + value, + }) + } + } else if code == "SM" || code == "CC" || code == "**" { + annotations.push(AlignStockholmAnnotation::{ + code, + feature: align_stockholm_gf_feature(code), + value: align_stockholm_join(values, " "), + }) + } else { + if values.length() != 1 { + raise AlignStockholmError("GF " + code + " must occur at most once") + } + annotations.push(AlignStockholmAnnotation::{ + code, + feature: align_stockholm_gf_feature(code), + value: values[0], + }) + } + } + let sq_values : Array[String] = [] + for annotation in raw { + if annotation.code == "SQ" { + sq_values.push(annotation.value) + } + } + if sq_values.length() > 1 { + raise AlignStockholmError("GF SQ must occur at most once") + } + if sq_values.length() == 1 { + let declared = align_stockholm_parse_positive_int( + sq_values[0], + "Stockholm GF SQ", + ) + if declared != rows { + raise AlignStockholmError( + "Inconsistent number of sequences in Stockholm alignment", + ) + } + } + annotations +} + +///| +fn align_stockholm_build_references( + raw : Array[AlignStockholmRawAnnotation], +) -> Array[AlignStockholmReference] raise AlignStockholmError { + let builders : Array[AlignStockholmReferenceBuilder] = [] + let pending_comments : Array[String] = [] + for annotation in raw { + if annotation.code == "RC" { + pending_comments.push(annotation.value) + } else if annotation.code == "RN" { + let value = annotation.value + if value.length() < 3 || + value.unsafe_get(0).to_int() != '['.to_int() || + value.unsafe_get(value.length() - 1).to_int() != ']'.to_int() { + raise AlignStockholmError("Malformed Stockholm GF RN annotation") + } + let number = align_stockholm_parse_positive_int( + value[1:value.length() - 1].to_owned(), + "Stockholm reference number", + ) + builders.push(AlignStockholmReferenceBuilder::{ + number, + medline: "", + title: "", + author: "", + location: "", + comment: align_stockholm_join(pending_comments, " "), + }) + pending_comments.clear() + } else if annotation.code == "RM" || + annotation.code == "RT" || + annotation.code == "RA" || + annotation.code == "RL" { + if builders.length() == 0 { + raise AlignStockholmError( + "Stockholm GF " + annotation.code + " appears before GF RN", + ) + } + let last = builders.length() - 1 + if annotation.code == "RM" { + if builders[last].medline.length() > 0 { + raise AlignStockholmError("Duplicate Stockholm GF RM annotation") + } + builders[last].medline = annotation.value + } else if annotation.code == "RT" { + builders[last].title = align_stockholm_append_text( + builders[last].title, + annotation.value, + ) + } else if annotation.code == "RA" { + builders[last].author = align_stockholm_append_text( + builders[last].author, + annotation.value, + ) + } else { + builders[last].location = align_stockholm_append_text( + builders[last].location, + annotation.value, + ) + } + } + } + if pending_comments.length() > 0 { + raise AlignStockholmError( + "Stockholm GF RC annotation is not followed by GF RN", + ) + } + let references : Array[AlignStockholmReference] = [] + for builder in builders { + references.push(AlignStockholmReference::{ + number: builder.number, + medline: builder.medline, + title: builder.title, + author: builder.author, + location: builder.location, + comment: builder.comment, + }) + } + references +} + +///| +fn align_stockholm_build_database_references( + raw : Array[AlignStockholmRawAnnotation], +) -> Array[AlignStockholmDatabaseReference] raise AlignStockholmError { + let builders : Array[AlignStockholmDatabaseReferenceBuilder] = [] + for annotation in raw { + if annotation.code == "DR" { + builders.push(AlignStockholmDatabaseReferenceBuilder::{ + reference: annotation.value, + comment: "", + }) + } else if annotation.code == "DC" { + if builders.length() == 0 { + raise AlignStockholmError( + "Stockholm GF DC appears before a database reference", + ) + } + let last = builders.length() - 1 + if builders[last].comment.length() > 0 { + raise AlignStockholmError("Duplicate Stockholm GF DC annotation") + } + builders[last].comment = annotation.value + } + } + let references : Array[AlignStockholmDatabaseReference] = [] + for builder in builders { + references.push(AlignStockholmDatabaseReference::{ + reference: builder.reference, + comment: builder.comment, + }) + } + references +} + +///| +fn align_stockholm_build_nested_domains( + raw : Array[AlignStockholmRawAnnotation], +) -> Array[AlignStockholmNestedDomain] raise AlignStockholmError { + let builders : Array[AlignStockholmNestedDomainBuilder] = [] + for annotation in raw { + if annotation.code == "NE" { + builders.push(AlignStockholmNestedDomainBuilder::{ + accession: annotation.value, + location: "", + }) + } else if annotation.code == "NL" { + if builders.length() == 0 { + raise AlignStockholmError( + "Stockholm GF NL appears before a nested domain", + ) + } + let last = builders.length() - 1 + if builders[last].location.length() > 0 { + raise AlignStockholmError("Duplicate Stockholm GF NL annotation") + } + builders[last].location = annotation.value + } + } + let domains : Array[AlignStockholmNestedDomain] = [] + for builder in builders { + domains.push(AlignStockholmNestedDomain::{ + accession: builder.accession, + location: builder.location, + }) + } + domains +} + +///| +fn align_stockholm_build_gc( + raw : Array[AlignStockholmRawAnnotation], + source_width : Int, + removed : Array[Bool], +) -> Array[AlignStockholmColumnAnnotation] raise AlignStockholmError { + let annotations : Array[AlignStockholmColumnAnnotation] = [] + for annotation in raw { + if annotation.value.length() != source_width { + raise AlignStockholmError( + "GC " + + annotation.code + + " length is " + + annotation.value.length().to_string() + + ", expected " + + source_width.to_string(), + ) + } + annotations.push(AlignStockholmColumnAnnotation::{ + code: annotation.code, + feature: align_stockholm_gc_feature(annotation.code), + value: align_stockholm_remove_columns(annotation.value, removed), + }) + } + annotations +} + +///| +fn align_stockholm_prepare_rows( + rows : Array[String], +) -> (Array[String], String, Array[Bool]) raise AlignStockholmError { + if rows.length() == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one sequence", + ) + } + let width = rows[0].length() + let operations = Array::make(width, 'M'.to_int()) + let all_gap = Array::make(width, true) + for row in rows { + if row.length() != width { + raise AlignStockholmError("Stockholm aligned rows must have equal widths") + } + align_stockholm_validate_printed_row(row) + for column = 0; column < width; column = column + 1 { + let code = row.unsafe_get(column).to_int() + if code == '-'.to_int() { + if operations[column] == 'I'.to_int() { + raise AlignStockholmError( + "Stockholm column mixes insertion and deletion gap symbols", + ) + } + operations[column] = 'D'.to_int() + } else if code == '.'.to_int() { + if operations[column] == 'D'.to_int() { + raise AlignStockholmError( + "Stockholm column mixes insertion and deletion gap symbols", + ) + } + operations[column] = 'I'.to_int() + } else { + all_gap[column] = false + } + } + } + let normalized : Array[String] = [] + for row in rows { + let output = StringBuilder::new(size_hint=width) + for column = 0; column < width; column = column + 1 { + if !all_gap[column] { + let code = row.unsafe_get(column).to_int() + output.write_char( + (if code == '.'.to_int() { '-'.to_int() } else { code }).unsafe_to_char(), + ) + } + } + normalized.push(output.to_string()) + } + let operation_output = StringBuilder::new(size_hint=width) + for column = 0; column < width; column = column + 1 { + if !all_gap[column] { + operation_output.write_char(operations[column].unsafe_to_char()) + } + } + (normalized, operation_output.to_string(), all_gap) +} + +///| +fn align_stockholm_validate_alignment( + alignment : AlignStockholmAlignment, +) -> Unit raise AlignStockholmError { + if alignment.version != "1.0" { + raise AlignStockholmError("Only Stockholm version 1.0 is supported") + } + if alignment.sequences.length() == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one sequence", + ) + } + let width = alignment.sequences[0].aligned_sequence.length() + if width == 0 { + raise AlignStockholmError( + "Stockholm alignment must contain at least one column", + ) + } + let ids : Array[String] = [] + for sequence in alignment.sequences { + align_stockholm_validate_id(sequence.id) + if align_stockholm_find_string(ids, sequence.id) >= 0 { + raise AlignStockholmError( + "Duplicate Stockholm sequence identifier " + sequence.id, + ) + } + ids.push(sequence.id) + if sequence.aligned_sequence.length() != width { + raise AlignStockholmError("Stockholm aligned rows must have equal widths") + } + if sequence.sequence != + align_stockholm_remove_gaps(sequence.aligned_sequence) { + raise AlignStockholmError( + "Stockholm ungapped sequence does not match aligned row", + ) + } + align_stockholm_validate_sequence_annotations(sequence) + } + if alignment.operations.length() != width { + raise AlignStockholmError( + "Stockholm operations length must match alignment width", + ) + } + for column = 0; column < width; column = column + 1 { + let operation = alignment.operations.unsafe_get(column).to_int() + if operation != 'M'.to_int() && + operation != 'D'.to_int() && + operation != 'I'.to_int() { + raise AlignStockholmError( + "Stockholm operations may contain only M, D, and I", + ) + } + let mut all_gap = true + let mut has_gap = false + for sequence in alignment.sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + has_gap = true + } else { + all_gap = false + } + } + if all_gap { + raise AlignStockholmError( + "Stockholm alignment must not retain all-gap columns", + ) + } + if operation == 'M'.to_int() && has_gap { + raise AlignStockholmError( + "Stockholm match operation cannot contain a gap", + ) + } + } + for annotation in alignment.column_annotations { + if annotation.code.length() == 0 { + raise AlignStockholmError("Stockholm GC code must not be empty") + } + if annotation.value.length() != width { + raise AlignStockholmError( + "Stockholm GC annotation length must match alignment width", + ) + } + } + if alignment.removed_all_gap_columns < 0 || + alignment.source_columns != width + alignment.removed_all_gap_columns { + raise AlignStockholmError("Invalid Stockholm source-column metadata") + } +} + +///| +fn align_stockholm_validate_sequence_annotations( + sequence : AlignStockholmSequence, +) -> Unit raise AlignStockholmError { + for annotation in sequence.annotations { + if annotation.code.length() == 0 { + raise AlignStockholmError("Stockholm GS code must not be empty") + } + } + for annotation in sequence.letter_annotations { + if annotation.code.length() == 0 { + raise AlignStockholmError("Stockholm GR code must not be empty") + } + if annotation.value.length() != sequence.sequence.length() { + raise AlignStockholmError( + "Stockholm GR residue annotation length must match sequence length", + ) + } + if annotation.aligned_value.length() != sequence.aligned_sequence.length() { + raise AlignStockholmError( + "Stockholm aligned GR annotation length must match alignment width", + ) + } + } +} + +///| +fn align_stockholm_count_pair( + alignment : AlignStockholmAlignment, + first_row : Int, + second_row : Int, +) -> AlignStockholmCounts { + let mut aligned = 0 + let mut identities = 0 + let mut mismatches = 0 + let mut gap_columns = 0 + let mut double_gap_columns = 0 + let mut gap_opens = 0 + let mut previous_gap_state = 0 + for column = 0; column < alignment.alignment_length(); column = column + 1 { + let first = alignment.sequences[first_row].aligned_sequence + .unsafe_get(column) + .to_int() + let second = alignment.sequences[second_row].aligned_sequence + .unsafe_get(column) + .to_int() + let first_gap = first == '-'.to_int() + let second_gap = second == '-'.to_int() + let gap_state = if first_gap && second_gap { + 3 + } else if first_gap { + 1 + } else if second_gap { + 2 + } else { + 0 + } + if gap_state == 0 { + aligned = aligned + 1 + if align_stockholm_upper_code(first) == align_stockholm_upper_code(second) { + identities = identities + 1 + } else { + mismatches = mismatches + 1 + } + } else { + if gap_state == 3 { + double_gap_columns = double_gap_columns + 1 + } else { + gap_columns = gap_columns + 1 + } + if gap_state != previous_gap_state { + gap_opens = gap_opens + 1 + } + } + previous_gap_state = gap_state + } + AlignStockholmCounts::{ + pairs: 1, + aligned, + identities, + mismatches, + gap_columns, + double_gap_columns, + gap_opens, + } +} + +///| +fn align_stockholm_movement_changes( + alignment : AlignStockholmAlignment, + first_column : Int, + second_column : Int, +) -> Bool { + for sequence in alignment.sequences { + let first_gap = sequence.aligned_sequence.unsafe_get(first_column).to_int() == + '-'.to_int() + let second_gap = sequence.aligned_sequence + .unsafe_get(second_column) + .to_int() == + '-'.to_int() + if first_gap != second_gap { + return true + } + } + false +} + +///| +fn align_stockholm_gf_feature(code : String) -> String { + match code { + "ID" => "identifier" + "AC" => "accession" + "DE" => "definition" + "AU" => "author" + "SE" => "source of seed" + "SS" => "source of structure" + "GA" => "gathering method" + "TC" => "trusted cutoff" + "NC" => "noise cutoff" + "BM" => "build method" + "SM" => "search method" + "TP" => "type" + "PI" => "previous identifier" + "CC" => "comment" + "CL" => "clan" + "WK" => "wikipedia" + "CB" => "calibration method" + "**" => "**" + _ => code + } +} + +///| +fn align_stockholm_gs_feature(code : String) -> String { + match code { + "AC" => "accession" + "OS" => "organism" + "OC" => "organism classification" + "LO" => "look" + _ => code + } +} + +///| +fn align_stockholm_gr_feature(code : String) -> String { + match code { + "SS" => "secondary structure" + "PP" => "posterior probability" + "CSA" => "Catalytic Site Atlas" + "SA" => "surface accessibility" + "TM" => "transmembrane" + "LI" => "ligand binding" + "AS" => "active site" + "pAS" => "active site - Pfam predicted" + "sAS" => "active site - from SwissProt" + "IN" => "intron" + _ => code + } +} + +///| +fn align_stockholm_gc_feature(code : String) -> String { + match code { + "RF" => "reference coordinate annotation" + "seq_cons" => "consensus sequence" + "scorecons" => "consensus score" + "scorecons_70" => "consensus score 70" + "scorecons_80" => "consensus score 80" + "scorecons_90" => "consensus score 90" + "MM" => "model mask" + "SS_cons" => "consensus secondary structure" + "PP_cons" => "consensus posterior probability" + "CSA_cons" => "consensus Catalytic Site Atlas" + "SA_cons" => "consensus surface accessibility" + "TM_cons" => "consensus transmembrane" + "LI_cons" => "consensus ligand binding" + "AS_cons" => "consensus active site" + "pAS_cons" => "consensus active site - Pfam predicted" + "sAS_cons" => "consensus active site - from SwissProt" + "IN_cons" => "consensus intron" + "RNA_elements" => "RNA elements" + "RNA_structural_element" => "RNA structural element" + "RNA_structural_elements" => "RNA structural elements" + "RNA_ligand_AdoCbl" => "RNA ligand AdoCbl" + "RNA_ligand_AqCbl" => "RNA ligand AqCbl" + "RNA_ligand_FMN" => "RNA ligand FMN" + "RNA_ligand_Guanidinium" => "RNA ligand Guanidinium" + "RNA_ligand_SAM" => "RNA ligand SAM" + "RNA_ligand_THF_1" => "RNA ligand THF 1" + "RNA_ligand_THF_2" => "RNA ligand THF 2" + "RNA_ligand_TPP" => "RNA ligand TPP" + "RNA_ligand_preQ1" => "RNA ligand preQ1" + "RNA_motif_k_turn" => "RNA motif k turn" + "Repeat_unit" => "Repeat unit" + "2L3J_B_SS" => "2L3J B SS" + "CORE" => "CORE" + "PK" => "PK" + "PK_SS" => "PK SS" + "cons" => "cons" + _ => code + } +} + +///| +fn align_stockholm_is_known_gf(code : String) -> Bool { + code == "ID" || + code == "AC" || + code == "DE" || + code == "AU" || + code == "SE" || + code == "SS" || + code == "GA" || + code == "TC" || + code == "NC" || + code == "BM" || + code == "SM" || + code == "TP" || + code == "PI" || + code == "CC" || + code == "CL" || + code == "WK" || + code == "CB" || + code == "**" +} + +///| +fn align_stockholm_infer_operations( + sequences : Array[AlignStockholmSequence], +) -> String { + let width = if sequences.length() == 0 { + 0 + } else { + sequences[0].aligned_sequence.length() + } + let output = StringBuilder::new(size_hint=width) + for column = 0; column < width; column = column + 1 { + let mut has_gap = false + for sequence in sequences { + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + has_gap = true + break + } + } + output.write_char(if has_gap { 'D' } else { 'M' }) + } + output.to_string() +} + +///| +fn align_stockholm_render_row(row : String, operations : String) -> String { + let output = StringBuilder::new(size_hint=row.length()) + for column = 0; column < row.length(); column = column + 1 { + let code = row.unsafe_get(column).to_int() + if code == '-'.to_int() && + operations.unsafe_get(column).to_int() == 'I'.to_int() { + output.write_char('.') + } else { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_stockholm_render_letter_annotation( + sequence : AlignStockholmSequence, + value : String, +) -> String { + let output = StringBuilder::new(size_hint=sequence.aligned_sequence.length()) + let mut residue = 0 + for column = 0 + column < sequence.aligned_sequence.length() + column = column + 1 { + if sequence.aligned_sequence.unsafe_get(column).to_int() == '-'.to_int() { + output.write_char('.') + } else { + output.write_char(value.unsafe_get(residue).unsafe_to_char()) + residue = residue + 1 + } + } + output.to_string() +} + +///| +fn align_stockholm_residue_annotation_value( + aligned_sequence : String, + aligned_value : String, +) -> String { + let output = StringBuilder::new() + for column = 0; column < aligned_sequence.length(); column = column + 1 { + if aligned_sequence.unsafe_get(column).to_int() != '-'.to_int() { + output.write_char(aligned_value.unsafe_get(column).unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_stockholm_remove_columns( + value : String, + removed : Array[Bool], +) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + if !removed[index] { + output.write_char(value.unsafe_get(index).unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_stockholm_count_true(values : Array[Bool]) -> Int { + let mut count = 0 + for value in values { + if value { + count = count + 1 + } + } + count +} + +///| +fn align_stockholm_remove_gaps(sequence : String) -> String { + let output = StringBuilder::new(size_hint=sequence.length()) + for index = 0; index < sequence.length(); index = index + 1 { + let code = sequence.unsafe_get(index).to_int() + if code != '-'.to_int() && code != '.'.to_int() { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_stockholm_remove_character(value : String, target : Int) -> String { + let output = StringBuilder::new(size_hint=value.length()) + for index = 0; index < value.length(); index = index + 1 { + let code = value.unsafe_get(index).to_int() + if code != target { + output.write_char(code.unsafe_to_char()) + } + } + output.to_string() +} + +///| +fn align_stockholm_normalize_row( + row : String, +) -> String raise AlignStockholmError { + align_stockholm_validate_printed_row(row) + let output = StringBuilder::new(size_hint=row.length()) + for index = 0; index < row.length(); index = index + 1 { + let code = row.unsafe_get(index).to_int() + output.write_char( + (if code == '.'.to_int() { '-'.to_int() } else { code }).unsafe_to_char(), + ) + } + output.to_string() +} + +///| +fn align_stockholm_validate_printed_row( + row : String, +) -> Unit raise AlignStockholmError { + for index = 0; index < row.length(); index = index + 1 { + let code = row.unsafe_get(index).to_int() + if align_stockholm_is_whitespace(code) || code < 33 || code > 126 { + raise AlignStockholmError( + "Stockholm sequence rows may contain only printable non-space ASCII", + ) + } + } +} + +///| +fn align_stockholm_validate_id(id : String) -> Unit raise AlignStockholmError { + if id.length() == 0 { + raise AlignStockholmError("Stockholm sequence identifier must not be empty") + } + for index = 0; index < id.length(); index = index + 1 { + let code = id.unsafe_get(index).to_int() + if align_stockholm_is_whitespace(code) || code < 33 || code > 126 { + raise AlignStockholmError( + "Stockholm sequence identifiers may not contain whitespace", + ) + } + } +} + +///| +fn align_stockholm_validate_annotation_code( + code : String, +) -> Unit raise AlignStockholmError { + if code.length() == 0 { + raise AlignStockholmError("Stockholm annotation code must not be empty") + } + for index = 0; index < code.length(); index = index + 1 { + let value = code.unsafe_get(index).to_int() + if align_stockholm_is_whitespace(value) || value < 33 || value > 126 { + raise AlignStockholmError( + "Stockholm annotation codes may not contain whitespace", + ) + } + } +} + +///| +fn align_stockholm_validate_row( + alignment : AlignStockholmAlignment, + row : Int, +) -> Unit raise AlignStockholmError { + if row < 0 || row >= alignment.sequences.length() { + raise AlignStockholmError( + "Stockholm row index " + row.to_string() + " is out of range", + ) + } +} + +///| +fn align_stockholm_append_raw( + values : Array[AlignStockholmRawAnnotation], + code : String, + value : String, +) -> Unit { + for index = 0; index < values.length(); index = index + 1 { + if values[index].code == code { + values[index] = AlignStockholmRawAnnotation::{ + code, + value: values[index].value + value, + } + return + } + } + values.push(AlignStockholmRawAnnotation::{ code, value }) +} + +///| +fn align_stockholm_markup_rest( + line : String, + prefix : String, + line_index : Int, +) -> String raise AlignStockholmError { + if line.length() == prefix.length() || + !align_stockholm_is_whitespace(line.unsafe_get(prefix.length()).to_int()) { + raise AlignStockholmError( + "Malformed Stockholm markup at line " + (line_index + 1).to_string(), + ) + } + line[prefix.length():].to_owned().trim().to_owned() +} + +///| +fn align_stockholm_take_fields( + value : String, + count : Int, +) -> (Array[String], String) { + let fields : Array[String] = [] + let mut index = 0 + while fields.length() < count { + while index < value.length() && + align_stockholm_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= value.length() { + break + } + let start = index + while index < value.length() && + !align_stockholm_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + fields.push(value[start:index].to_owned()) + } + while index < value.length() && + align_stockholm_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + (fields, if index < value.length() { value[index:].to_owned() } else { "" }) +} + +///| +fn align_stockholm_split_whitespace(value : String) -> Array[String] { + let fields : Array[String] = [] + let mut index = 0 + while index < value.length() { + while index < value.length() && + align_stockholm_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + if index >= value.length() { + break + } + let start = index + while index < value.length() && + !align_stockholm_is_whitespace(value.unsafe_get(index).to_int()) { + index = index + 1 + } + fields.push(value[start:index].to_owned()) + } + fields +} + +///| +fn align_stockholm_merge_wikipedia(values : Array[String]) -> Array[String] { + let merged : Array[String] = [] + let mut index = 0 + while index < values.length() { + let output = StringBuilder::new() + let mut value = values[index] + while value.length() > 0 && + value.unsafe_get(value.length() - 1).to_int() == '/'.to_int() && + index + 1 < values.length() { + output.write_string(value[0:value.length() - 1].to_owned()) + index = index + 1 + value = values[index] + } + output.write_string(value) + merged.push(output.to_string()) + index = index + 1 + } + merged +} + +///| +fn align_stockholm_wrap_text(prefix : String, text : String) -> String { + if text.length() == 0 { + return prefix + "\n" + } + let words = align_stockholm_split_whitespace(text) + if words.length() == 0 { + return prefix + "\n" + } + let output = StringBuilder::new() + let mut line = prefix + for word in words { + let separator = if line.length() == prefix.length() { "" } else { " " } + if line.length() > prefix.length() && line.length() + 1 + word.length() > 79 { + output.write_string(line) + output.write_char('\n') + line = prefix + word + } else { + line = line + separator + word + } + } + output.write_string(line) + output.write_char('\n') + output.to_string() +} + +///| +fn align_stockholm_parse_positive_int( + text : String, + label : String, +) -> Int raise AlignStockholmError { + if text.length() == 0 { + raise AlignStockholmError(label + " is empty") + } + let mut value = 0 + for index = 0; index < text.length(); index = index + 1 { + let code = text.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + raise AlignStockholmError(label + " must be a positive integer") + } + let digit = code - '0'.to_int() + if value > 214748364 || (value == 214748364 && digit > 7) { + raise AlignStockholmError( + label + " is outside the supported integer range", + ) + } + value = value * 10 + digit + } + if value <= 0 { + raise AlignStockholmError(label + " must be positive") + } + value +} + +///| +fn align_stockholm_normalize_lines(text : String) -> Array[String] { + let lines : Array[String] = [] + for raw in text.split("\n") { + let owned = raw.to_owned() + if owned.length() > 0 && + owned.unsafe_get(owned.length() - 1).to_int() == '\r'.to_int() { + lines.push(owned[0:owned.length() - 1].to_owned()) + } else { + lines.push(owned) + } + } + lines +} + +///| +fn align_stockholm_find_string(values : Array[String], target : String) -> Int { + for index = 0; index < values.length(); index = index + 1 { + if values[index] == target { + return index + } + } + -1 +} + +///| +fn align_stockholm_join(values : Array[String], separator : String) -> String { + let output = StringBuilder::new() + for index = 0; index < values.length(); index = index + 1 { + if index > 0 { + output.write_string(separator) + } + output.write_string(values[index]) + } + output.to_string() +} + +///| +fn align_stockholm_append_text(existing : String, value : String) -> String { + if existing.length() == 0 { + value + } else { + existing + " " + value + } +} + +///| +fn align_stockholm_pad_right(value : String, width : Int) -> String { + if value.length() >= width { + return value + } + let output = StringBuilder::new(size_hint=width) + output.write_string(value) + for _index = value.length(); _index < width; _index = _index + 1 { + output.write_char(' ') + } + output.to_string() +} + +///| +fn align_stockholm_starts_with(value : String, prefix : String) -> Bool { + if prefix.length() > value.length() { + return false + } + for index = 0; index < prefix.length(); index = index + 1 { + if value.unsafe_get(index).to_int() != prefix.unsafe_get(index).to_int() { + return false + } + } + true +} + +///| +fn align_stockholm_is_whitespace(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() +} + +///| +fn align_stockholm_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 'a'.to_int() + 'A'.to_int() + } else { + code + } +} + +///| +fn align_stockholm_max(first : Int, second : Int) -> Int { + if first > second { + first + } else { + second + } +} diff --git a/src/align_tabular.mbt b/src/align_tabular.mbt new file mode 100644 index 00000000..51b56935 --- /dev/null +++ b/src/align_tabular.mbt @@ -0,0 +1,1788 @@ +// Parse alignment-aware BLAST outfmt 7 and FASTA -m 8CB/-m 8CC output. +// +// This follows the field and traceback semantics of Biopython 1.86 +// Bio.Align.tabular. Unlike hit-table parsers, BTOP and aln_code values are +// expanded into explicit pairwise coordinate paths. + +///| +/// Error raised by alignment-aware tabular parsing and conversion. +pub suberror AlignTabularError { + AlignTabularError(String) +} + +///| +/// Source used to construct an alignment path. +pub enum AlignTabularTrace { + NoTrace + Btop + Cigar +} derive(Eq, Debug) + +///| +/// One run in a tabular alignment path. +/// +/// `Aligned` consumes both rows, `QueryGap` consumes only the target row, and +/// `TargetGap` consumes only the query row. +pub enum AlignTabularOperation { + Aligned + QueryGap + TargetGap +} derive(Eq, Debug) + +///| +/// One original field/value pair from a tabular row. +pub struct AlignTabularField { + name : String + value : String +} derive(Eq, Debug) + +///| +/// Metadata attached to one BLAST or FASTA query block. +pub struct AlignTabularMetadata { + program : String + version : String + command_line : String + database : String + rid : String +} derive(Eq, Debug) + +///| +/// One pairwise alignment parsed from a tabular row. +/// +/// Intervals are zero-based and half-open. Coordinate paths preserve the +/// reported strand direction. For translated searches, nucleotide axes use +/// nucleotide coordinates and have `units_per_residue == 3`. +pub struct AlignTabularAlignment { + metadata : AlignTabularMetadata + query_id : String + query_description : String + target_id : String + query_length : Int? + target_length : Int? + query_start : Int? + query_end : Int? + target_start : Int? + target_end : Int? + query_strand : Strand + target_strand : Strand + query_sequence : String + target_sequence : String + target_coordinates : Array[Int] + query_coordinates : Array[Int] + operations : Array[AlignTabularOperation] + trace : AlignTabularTrace + trace_text : String + path_is_absolute : Bool + target_units_per_residue : Int + query_units_per_residue : Int + fields : Array[AlignTabularField] +} derive(Debug) + +///| +/// One query block, including zero-hit queries. +pub struct AlignTabularQueryResult { + metadata : AlignTabularMetadata + query_id : String + query_description : String + query_length : Int? + query_unit : String + field_names : Array[String] + declared_hits : Int + alignments : Array[AlignTabularAlignment] +} derive(Debug) + +///| +/// Complete parsed document. +pub struct AlignTabularDocument { + queries : Array[AlignTabularQueryResult] + processed_queries : Int? +} derive(Debug) + +///| +priv struct AlignTabularRelativePath { + target_coordinates : Array[Int] + query_coordinates : Array[Int] + operations : Array[AlignTabularOperation] + columns : Int +} + +///| +fn align_tabular_starts_with(value : String, prefix : String) -> Bool { + if prefix.length() > value.length() { + return false + } + for index in 0.. Bool { + if suffix.length() > value.length() { + return false + } + let offset = value.length() - suffix.length() + for index in 0.. String { + value.trim().to_owned() +} + +///| +fn align_tabular_split_char(value : String, separator : Char) -> Array[String] { + let result : Array[String] = [] + let mut start = 0 + for index in 0.. Array[String] { + let result : Array[String] = [] + let mut start = 0 + let mut in_token = false + for index in 0.. String { + let mut result = "" + for index in start.. start { + result = result + " " + } + result = result + values[index] + } + result +} + +///| +fn align_tabular_last_delimiter(value : String, delimiter : String) -> Int? { + if delimiter.length() == 0 || delimiter.length() > value.length() { + return None + } + let mut found : Int? = None + for start in 0..<=(value.length() - delimiter.length()) { + let mut matches = true + for offset in 0.. Int raise AlignTabularError { + if value.length() == 0 { + raise AlignTabularError("missing integer field " + label) + } + let mut start = 0 + let negative = value.unsafe_get(0).to_int() == '-'.to_int() + if negative || value.unsafe_get(0).to_int() == '+'.to_int() { + start = 1 + } + if start == value.length() { + raise AlignTabularError("invalid integer in " + label + ": " + value) + } + for index in start.. '9'.to_int() { + raise AlignTabularError("invalid integer in " + label + ": " + value) + } + } + let mut first_digit = start + while first_digit < value.length() && + value.unsafe_get(first_digit).to_int() == '0'.to_int() { + first_digit = first_digit + 1 + } + if first_digit == value.length() { + return 0 + } + let significant_length = value.length() - first_digit + let limit = if negative { "2147483648" } else { "2147483647" } + let mut exceeds_limit = significant_length > limit.length() + let mut equals_limit = significant_length == limit.length() + if significant_length == limit.length() { + for offset in 0.. bound { + exceeds_limit = true + equals_limit = false + break + } else if digit < bound { + equals_limit = false + break + } + } + } + if exceeds_limit { + raise AlignTabularError("integer overflow in " + label + ": " + value) + } + if negative && equals_limit { + return -2147483647 - 1 + } + let mut result = 0 + for index in first_digit.. Double raise AlignTabularError { + if value.length() == 0 { + raise AlignTabularError("missing number field " + label) + } + let mut index = 0 + if value.unsafe_get(0).to_int() == '-'.to_int() || + value.unsafe_get(0).to_int() == '+'.to_int() { + index = 1 + } + let mut mantissa_digits = 0 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + mantissa_digits = mantissa_digits + 1 + index = index + 1 + } + if index < value.length() && value.unsafe_get(index).to_int() == '.'.to_int() { + index = index + 1 + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + mantissa_digits = mantissa_digits + 1 + index = index + 1 + } + } + if mantissa_digits == 0 { + raise AlignTabularError("invalid number in " + label + ": " + value) + } + if index < value.length() && + ( + value.unsafe_get(index).to_int() == 'e'.to_int() || + value.unsafe_get(index).to_int() == 'E'.to_int() + ) { + index = index + 1 + let exponent_value_start = index + if index < value.length() && + ( + value.unsafe_get(index).to_int() == '-'.to_int() || + value.unsafe_get(index).to_int() == '+'.to_int() + ) { + index = index + 1 + } + let exponent_start = index + while index < value.length() { + let code = value.unsafe_get(index).to_int() + if code < '0'.to_int() || code > '9'.to_int() { + break + } + index = index + 1 + } + if index == exponent_start { + raise AlignTabularError("invalid number in " + label + ": " + value) + } + ignore( + align_tabular_parse_int( + value[exponent_value_start:index].to_owned(), + label + " exponent", + ), + ) + } + if index != value.length() { + raise AlignTabularError("invalid number in " + label + ": " + value) + } + let parse_value = if value.unsafe_get(0).to_int() == '+'.to_int() { + value[1:value.length()].to_owned() + } else { + value + } + match parse_double(parse_value) { + Some(number) => + if number.is_nan() || number.abs() > 1.0e300 { + raise AlignTabularError("non-finite number in " + label + ": " + value) + } else { + number + } + None => raise AlignTabularError("invalid number in " + label + ": " + value) + } +} + +///| +fn align_tabular_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn align_tabular_copy_ints(values : Array[Int]) -> Array[Int] { + let result : Array[Int] = [] + for value in values { + result.push(value) + } + result +} + +///| +fn align_tabular_copy_strings(values : Array[String]) -> Array[String] { + let result : Array[String] = [] + for value in values { + result.push(value) + } + result +} + +///| +fn align_tabular_copy_operations( + values : Array[AlignTabularOperation], +) -> Array[AlignTabularOperation] { + let result : Array[AlignTabularOperation] = [] + for value in values { + result.push(value) + } + result +} + +///| +fn align_tabular_known_blast_program(program : String) -> Bool { + match program { + "BLASTN" + | "BLASTP" + | "BLASTX" + | "TBLASTN" + | "TBLASTX" + | "DELTABLAST" + | "PSIBLAST" + | "RPSBLAST" + | "RPSTBLASTN" => true + _ => false + } +} + +///| +fn align_tabular_known_fasta_program(program : String) -> Bool { + match program { + "FASTA" + | "SSEARCH" + | "GGSEARCH" + | "GLSEARCH" + | "FASTX" + | "FASTY" + | "TFASTX" + | "TFASTY" => true + _ => false + } +} + +///| +fn align_tabular_program_header(value : String) -> (String, String)? { + let parts = align_tabular_split_whitespace(value) + if parts.length() < 2 { + return None + } + let program = parts[0].to_upper() + if !align_tabular_known_blast_program(program) && + !align_tabular_known_fasta_program(program) { + return None + } + Some((program, align_tabular_join(parts, 1))) +} + +///| +fn align_tabular_field_value( + fields : Array[AlignTabularField], + name : String, +) -> String? { + for field in fields { + if field.name == name { + return Some(field.value) + } + } + None +} + +///| +fn align_tabular_optional_int( + fields : Array[AlignTabularField], + name : String, +) -> Int? raise AlignTabularError { + match align_tabular_field_value(fields, name) { + Some(value) => Some(align_tabular_parse_int(value, name)) + None => None + } +} + +///| +fn align_tabular_optional_double( + fields : Array[AlignTabularField], + name : String, +) -> Double? raise AlignTabularError { + match align_tabular_field_value(fields, name) { + Some(value) => Some(align_tabular_parse_double(value, name)) + None => None + } +} + +///| +fn align_tabular_is_known_field(name : String) -> Bool { + match name { + "query id" + | "subject id" + | "% identity" + | "alignment length" + | "mismatches" + | "gap opens" + | "q. start" + | "q. end" + | "s. start" + | "s. end" + | "evalue" + | "bit score" + | "BTOP" + | "aln_code" + | "query gi" + | "query acc." + | "query acc.ver" + | "query length" + | "subject ids" + | "subject gi" + | "subject gis" + | "subject acc." + | "subject acc.ver" + | "subject accs." + | "subject length" + | "query seq" + | "subject seq" + | "score" + | "identical" + | "positives" + | "gaps" + | "% positives" + | "% hsp coverage" + | "query/sbjct frames" + | "query frame" + | "sbjct frame" + | "subject tax ids" + | "subject sci names" + | "subject com names" + | "subject blast names" + | "subject super kingdoms" + | "subject title" + | "subject titles" + | "subject strand" + | "% subject coverage" => true + _ => false + } +} + +///| +fn align_tabular_validate_field_names( + names : Array[String], +) -> Unit raise AlignTabularError { + if names.length() == 0 { + raise AlignTabularError("Fields header must not be empty") + } + let seen : Map[String, Bool] = Map([], capacity=names.length()) + for name in names { + if name.length() == 0 { + raise AlignTabularError("Fields header contains an empty field") + } + if !align_tabular_is_known_field(name) { + raise AlignTabularError("unexpected tabular field: " + name) + } + if seen.contains(name) { + raise AlignTabularError("duplicate tabular field: " + name) + } + seen.set(name, true) + } +} + +///| +fn align_tabular_push_step( + target_coordinates : Array[Int], + query_coordinates : Array[Int], + operations : Array[AlignTabularOperation], + operation : AlignTabularOperation, + target_count : Int, + query_count : Int, +) -> Unit { + let last = operations.length() - 1 + if last >= 0 && operations[last] == operation { + let point = target_coordinates.length() - 1 + target_coordinates[point] = target_coordinates[point] + target_count + query_coordinates[point] = query_coordinates[point] + query_count + } else { + let point = target_coordinates.length() - 1 + target_coordinates.push(target_coordinates[point] + target_count) + query_coordinates.push(query_coordinates[point] + query_count) + operations.push(operation) + } +} + +///| +/// Parse a BLAST traceback operations (BTOP) string into a relative path. +pub fn align_tabular_parse_btop( + btop : String, +) -> (Array[Int], Array[Int], Array[AlignTabularOperation]) raise AlignTabularError { + if btop.length() == 0 { + raise AlignTabularError("BTOP value must not be empty") + } + let target_coordinates = [0] + let query_coordinates = [0] + let operations : Array[AlignTabularOperation] = [] + let mut index = 0 + while index < btop.length() { + let code = btop.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + let start = index + while index < btop.length() { + let digit = btop.unsafe_get(index).to_int() + if digit >= '0'.to_int() && digit <= '9'.to_int() { + index = index + 1 + } else { + break + } + } + let count = align_tabular_parse_int( + btop[start:index].to_owned(), + "BTOP match length", + ) + if count <= 0 { + raise AlignTabularError("BTOP match lengths must be positive") + } + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::Aligned, + count, + count, + ) + } else { + if index + 1 >= btop.length() { + raise AlignTabularError("truncated BTOP residue pair") + } + let first = btop.unsafe_get(index).unsafe_to_char() + let second = btop.unsafe_get(index + 1).unsafe_to_char() + if first == '-' && second == '-' { + raise AlignTabularError("BTOP residue pair cannot contain two gaps") + } + if first == '-' { + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::QueryGap, + 1, + 0, + ) + } else if second == '-' { + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::TargetGap, + 0, + 1, + ) + } else { + let first_valid = first.is_ascii_alphabetic() || first == '*' + let second_valid = second.is_ascii_alphabetic() || second == '*' + if !first_valid || !second_valid { + raise AlignTabularError("invalid BTOP residue pair") + } + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::Aligned, + 1, + 1, + ) + } + index = index + 2 + } + } + (target_coordinates, query_coordinates, operations) +} + +///| +/// Parse FASTA `aln_code` CIGAR syntax into a relative path. +/// +/// FASTA uses `I` for residues present only in the subject/target and `D` for +/// residues present only in the query, matching Biopython's parser. +pub fn align_tabular_parse_cigar( + cigar : String, +) -> (Array[Int], Array[Int], Array[AlignTabularOperation]) raise AlignTabularError { + if cigar.length() == 0 { + raise AlignTabularError("aln_code value must not be empty") + } + let target_coordinates = [0] + let query_coordinates = [0] + let operations : Array[AlignTabularOperation] = [] + let mut index = 0 + while index < cigar.length() { + let start = index + while index < cigar.length() { + let code = cigar.unsafe_get(index).to_int() + if code >= '0'.to_int() && code <= '9'.to_int() { + index = index + 1 + } else { + break + } + } + if start == index || index >= cigar.length() { + raise AlignTabularError("invalid aln_code operation") + } + let count = align_tabular_parse_int( + cigar[start:index].to_owned(), + "aln_code length", + ) + if count <= 0 { + raise AlignTabularError("aln_code lengths must be positive") + } + let operation = cigar.unsafe_get(index).unsafe_to_char() + match operation { + 'M' => + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::Aligned, + count, + count, + ) + 'I' => + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::QueryGap, + count, + 0, + ) + 'D' => + align_tabular_push_step( + target_coordinates, + query_coordinates, + operations, + AlignTabularOperation::TargetGap, + 0, + count, + ) + _ => raise AlignTabularError("unsupported aln_code operation") + } + index = index + 1 + } + (target_coordinates, query_coordinates, operations) +} + +///| +fn align_tabular_path( + fields : Array[AlignTabularField], +) -> AlignTabularRelativePath? raise AlignTabularError { + let btop = align_tabular_field_value(fields, "BTOP") + let cigar = align_tabular_field_value(fields, "aln_code") + if btop is Some(_) && cigar is Some(_) { + raise AlignTabularError("a row cannot contain both BTOP and aln_code") + } + let parsed = match (btop, cigar) { + (Some(value), None) => { + let (target, query, operations) = align_tabular_parse_btop(value) + Some((target, query, operations)) + } + (None, Some(value)) => { + let (target, query, operations) = align_tabular_parse_cigar(value) + Some((target, query, operations)) + } + _ => None + } + match parsed { + Some((target, query, operations)) => { + let mut columns = 0 + for index in 0.. query_count { target_count } else { query_count }) + } + Some(AlignTabularRelativePath::{ + target_coordinates: target, + query_coordinates: query, + operations, + columns, + }) + } + None => None + } +} + +///| +fn align_tabular_axis_factor(program : String, query : Bool) -> Int { + if query { + match program { + "BLASTX" | "TBLASTX" | "FASTX" | "FASTY" => 3 + _ => 1 + } + } else { + match program { + "TBLASTN" | "TBLASTX" | "RPSTBLASTN" | "TFASTX" | "TFASTY" => 3 + _ => 1 + } + } +} + +///| +fn align_tabular_interval( + reported_start : Int?, + reported_end : Int?, + label : String, +) -> (Int?, Int?, Strand) raise AlignTabularError { + match (reported_start, reported_end) { + (None, None) => (None, None, Strand::Star) + (Some(first), Some(last)) => { + if first <= 0 || last <= 0 { + raise AlignTabularError(label + " coordinates must be positive") + } + if first <= last { + (Some(first - 1), Some(last), Strand::Plus) + } else { + (Some(last - 1), Some(first), Strand::Minus) + } + } + _ => raise AlignTabularError(label + " start and end must appear together") + } +} + +///| +fn align_tabular_absolute_axis( + relative : Array[Int], + reported_start : Int, + reported_end : Int, + factor : Int, + label : String, +) -> Array[Int] raise AlignTabularError { + let consumed = relative[relative.length() - 1] - relative[0] + let expected = align_tabular_abs(reported_end - reported_start) + 1 + if consumed * factor != expected { + raise AlignTabularError( + label + + " traceback span " + + (consumed * factor).to_string() + + " does not match reported span " + + expected.to_string(), + ) + } + let direction = if reported_start <= reported_end { 1 } else { -1 } + let origin = if direction > 0 { reported_start - 1 } else { reported_start } + let result : Array[Int] = [] + for coordinate in relative { + result.push(origin + direction * coordinate * factor) + } + result +} + +///| +fn align_tabular_scale_axis(relative : Array[Int], factor : Int) -> Array[Int] { + let result : Array[Int] = [] + for coordinate in relative { + result.push(coordinate * factor) + } + result +} + +///| +fn align_tabular_ungapped_length(sequence : String) -> Int { + let mut length = 0 + for index in 0.. Unit raise AlignTabularError { + for index in 0.. (String, String, Int?, String) raise AlignTabularError { + let mut query_text = align_tabular_trim(value) + let mut length : Int? = None + let mut unit = "" + if fasta_style { + match align_tabular_last_delimiter(query_text, " - ") { + Some(delimiter) => { + let suffix = align_tabular_split_whitespace( + query_text[delimiter + 3:query_text.length()].to_owned(), + ) + if suffix.length() == 2 && (suffix[1] == "nt" || suffix[1] == "aa") { + let parsed_length = align_tabular_parse_int(suffix[0], "query size") + if parsed_length <= 0 { + raise AlignTabularError("query size must be positive") + } + length = Some(parsed_length) + unit = suffix[1] + query_text = align_tabular_trim(query_text[0:delimiter].to_owned()) + } + } + None => () + } + } + let parts = align_tabular_split_whitespace(query_text) + if parts.length() == 0 { + raise AlignTabularError("Query header must contain an identifier") + } + let identifier = parts[0] + let description = align_tabular_join(parts, 1) + (identifier, description, length, unit) +} + +///| +fn align_tabular_fields_header(value : String) -> Array[String] { + let raw = align_tabular_split_char(value, ',') + let result : Array[String] = [] + for field in raw { + result.push(align_tabular_trim(field)) + } + result +} + +///| +fn align_tabular_row_columns(line : String, expected : Int) -> Array[String] { + let tabular = align_tabular_split_char(line, '\t') + if tabular.length() == expected { + let result : Array[String] = [] + for value in tabular { + result.push(align_tabular_trim(value)) + } + return result + } + align_tabular_split_whitespace(line) +} + +///| +fn align_tabular_fields( + names : Array[String], + columns : Array[String], +) -> Array[AlignTabularField] raise AlignTabularError { + if columns.length() != names.length() { + raise AlignTabularError( + "tabular row has " + + columns.length().to_string() + + " columns but Fields declares " + + names.length().to_string(), + ) + } + let result : Array[AlignTabularField] = [] + for index in 0.. Unit raise AlignTabularError { + for field in fields { + match field.name { + "alignment length" + | "mismatches" + | "gap opens" + | "q. start" + | "q. end" + | "s. start" + | "s. end" + | "query length" + | "subject length" + | "score" + | "identical" + | "positives" + | "gaps" => { + let value = align_tabular_parse_int(field.value, field.name) + if value < 0 { + raise AlignTabularError(field.name + " must be non-negative") + } + if ( + field.name == "alignment length" || + field.name == "q. start" || + field.name == "q. end" || + field.name == "s. start" || + field.name == "s. end" || + field.name == "query length" || + field.name == "subject length" + ) && + value == 0 { + raise AlignTabularError(field.name + " must be positive") + } + } + "% identity" + | "evalue" + | "bit score" + | "% positives" + | "% hsp coverage" + | "% subject coverage" => { + let value = align_tabular_parse_double(field.value, field.name) + if value < 0.0 { + raise AlignTabularError(field.name + " must be non-negative") + } + if ( + field.name == "% identity" || + field.name == "% positives" || + field.name == "% hsp coverage" || + field.name == "% subject coverage" + ) && + value > 100.0 { + raise AlignTabularError(field.name + " must not exceed 100") + } + } + _ => () + } + } +} + +///| +fn align_tabular_alignment_from_row( + metadata : AlignTabularMetadata, + header_query_id : String, + header_description : String, + header_query_length : Int?, + names : Array[String], + line : String, +) -> AlignTabularAlignment raise AlignTabularError { + let columns = align_tabular_row_columns(line, names.length()) + let fields = align_tabular_fields(names, columns) + align_tabular_validate_numeric_fields(fields) + let query_id = match align_tabular_field_value(fields, "query id") { + Some(value) => value + None => + match align_tabular_field_value(fields, "query acc.ver") { + Some(value) => value + None => header_query_id + } + } + if query_id.length() == 0 { + raise AlignTabularError("tabular row has no query identifier") + } + if header_query_id.length() > 0 && query_id != header_query_id { + raise AlignTabularError( + "row query identifier " + + query_id + + " does not match Query header " + + header_query_id, + ) + } + let target_id = match align_tabular_field_value(fields, "subject id") { + Some(value) => value + None => + match align_tabular_field_value(fields, "subject acc.ver") { + Some(value) => value + None => "" + } + } + if target_id.length() == 0 { + raise AlignTabularError("tabular row has no subject identifier") + } + let row_query_length = align_tabular_optional_int(fields, "query length") + let query_length = match (header_query_length, row_query_length) { + (Some(header), Some(row)) => + if header != row { + raise AlignTabularError("query length disagrees with the Query header") + } else { + Some(header) + } + (Some(header), None) => Some(header) + (None, Some(row)) => Some(row) + (None, None) => None + } + match query_length { + Some(value) => + if value <= 0 { + raise AlignTabularError("query length must be positive") + } + None => () + } + let target_length = align_tabular_optional_int(fields, "subject length") + match target_length { + Some(value) => + if value <= 0 { + raise AlignTabularError("subject length must be positive") + } + None => () + } + let query_reported_start = align_tabular_optional_int(fields, "q. start") + let query_reported_end = align_tabular_optional_int(fields, "q. end") + let target_reported_start = align_tabular_optional_int(fields, "s. start") + let target_reported_end = align_tabular_optional_int(fields, "s. end") + let (query_start, query_end, query_strand) = align_tabular_interval( + query_reported_start, query_reported_end, "query", + ) + let (target_start, target_end, target_strand) = align_tabular_interval( + target_reported_start, target_reported_end, "subject", + ) + match (query_length, query_end) { + (Some(length), Some(end)) => + if end > length { + raise AlignTabularError("query coordinates exceed query length") + } + _ => () + } + match (target_length, target_end) { + (Some(length), Some(end)) => + if end > length { + raise AlignTabularError("subject coordinates exceed subject length") + } + _ => () + } + let query_sequence = align_tabular_field_value(fields, "query seq").unwrap_or( + "", + ) + let target_sequence = align_tabular_field_value(fields, "subject seq").unwrap_or( + "", + ) + align_tabular_validate_sequence(query_sequence, "query sequence") + align_tabular_validate_sequence(target_sequence, "subject sequence") + if query_sequence.length() > 0 && + target_sequence.length() > 0 && + query_sequence.length() != target_sequence.length() { + raise AlignTabularError( + "query and subject aligned sequences have different lengths", + ) + } + let path = align_tabular_path(fields) + let query_factor = align_tabular_axis_factor(metadata.program, true) + let target_factor = align_tabular_axis_factor(metadata.program, false) + let mut target_coordinates : Array[Int] = [] + let mut query_coordinates : Array[Int] = [] + let mut operations : Array[AlignTabularOperation] = [] + let mut path_is_absolute = false + let trace = if align_tabular_field_value(fields, "BTOP") is Some(_) { + AlignTabularTrace::Btop + } else if align_tabular_field_value(fields, "aln_code") is Some(_) { + AlignTabularTrace::Cigar + } else { + AlignTabularTrace::NoTrace + } + let trace_text = match trace { + Btop => align_tabular_field_value(fields, "BTOP").unwrap() + Cigar => align_tabular_field_value(fields, "aln_code").unwrap() + NoTrace => "" + } + match path { + Some(relative) => { + let alignment_length = align_tabular_optional_int( + fields, "alignment length", + ) + match alignment_length { + Some(length) => + if length != relative.columns { + raise AlignTabularError( + "traceback columns do not match alignment length", + ) + } + None => () + } + if query_sequence.length() > 0 { + if query_sequence.length() != relative.columns { + raise AlignTabularError( + "query sequence length does not match traceback columns", + ) + } + let consumed = relative.query_coordinates[relative.query_coordinates.length() - + 1] + if align_tabular_ungapped_length(query_sequence) != consumed { + raise AlignTabularError( + "query sequence consumption does not match traceback", + ) + } + } + if target_sequence.length() > 0 { + if target_sequence.length() != relative.columns { + raise AlignTabularError( + "subject sequence length does not match traceback columns", + ) + } + let consumed = relative.target_coordinates[relative.target_coordinates.length() - + 1] + if align_tabular_ungapped_length(target_sequence) != consumed { + raise AlignTabularError( + "subject sequence consumption does not match traceback", + ) + } + } + match + ( + query_reported_start, query_reported_end, target_reported_start, target_reported_end, + ) { + (Some(qs), Some(qe), Some(ts), Some(te)) => { + query_coordinates = align_tabular_absolute_axis( + relative.query_coordinates, + qs, + qe, + query_factor, + "query", + ) + target_coordinates = align_tabular_absolute_axis( + relative.target_coordinates, + ts, + te, + target_factor, + "subject", + ) + path_is_absolute = true + } + (None, None, None, None) => { + query_coordinates = align_tabular_scale_axis( + relative.query_coordinates, + query_factor, + ) + target_coordinates = align_tabular_scale_axis( + relative.target_coordinates, + target_factor, + ) + } + _ => + raise AlignTabularError( + "traceback requires all four query and subject coordinates", + ) + } + operations = align_tabular_copy_operations(relative.operations) + } + None => () + } + AlignTabularAlignment::{ + metadata, + query_id, + query_description: header_description, + target_id, + query_length, + target_length, + query_start, + query_end, + target_start, + target_end, + query_strand, + target_strand, + query_sequence, + target_sequence, + target_coordinates, + query_coordinates, + operations, + trace, + trace_text, + path_is_absolute, + target_units_per_residue: target_factor, + query_units_per_residue: query_factor, + fields, + } +} + +///| +fn align_tabular_parse_processed( + value : String, +) -> Int? raise AlignTabularError { + let suffix = " queries" + if !align_tabular_ends_with(value, suffix) { + return None + } + let prefixes = ["BLAST processed ", "FASTA processed "] + for prefix in prefixes { + if align_tabular_starts_with(value, prefix) { + let number = value[prefix.length():value.length() - suffix.length()].to_owned() + return Some(align_tabular_parse_int(number, "processed query count")) + } + } + None +} + +///| +fn align_tabular_parse_hits(value : String) -> Int? raise AlignTabularError { + let suffix = " hits found" + if !align_tabular_ends_with(value, suffix) { + return None + } + let number = align_tabular_trim( + value[0:value.length() - suffix.length()].to_owned(), + ) + Some(align_tabular_parse_int(number, "hit count")) +} + +///| +/// Parse BLAST outfmt 7 or FASTA `-m 8CB`/`-m 8CC` text. +pub fn align_tabular_parse( + content : String, +) -> AlignTabularDocument raise AlignTabularError { + if align_tabular_trim(content).length() == 0 { + raise AlignTabularError("empty alignment tabular document") + } + let lines = align_tabular_split_char(content, '\n') + let queries : Array[AlignTabularQueryResult] = [] + let mut processed_queries : Int? = None + let mut index = 0 + while index < lines.length() { + let line = align_tabular_trim(lines[index]) + if line.length() == 0 { + index = index + 1 + continue + } + if !align_tabular_starts_with(line, "# ") { + raise AlignTabularError("missing alignment tabular header") + } + let first_value = align_tabular_trim(line[2:line.length()].to_owned()) + match align_tabular_parse_processed(first_value) { + Some(count) => { + if count < 0 { + raise AlignTabularError("processed query count must be non-negative") + } + if processed_queries is Some(_) { + raise AlignTabularError("duplicate processed query summary") + } + processed_queries = Some(count) + index = index + 1 + continue + } + None => () + } + let mut command_line = "" + let mut program = "" + let mut version = "" + let mut fasta_style = false + match align_tabular_program_header(first_value) { + Some((parsed_program, parsed_version)) => { + program = parsed_program + version = parsed_version + fasta_style = align_tabular_known_fasta_program(program) + index = index + 1 + } + None => { + command_line = first_value + fasta_style = true + index = index + 1 + if index >= lines.length() { + raise AlignTabularError("missing FASTA program header") + } + let program_line = align_tabular_trim(lines[index]) + if !align_tabular_starts_with(program_line, "# ") { + raise AlignTabularError("missing FASTA program header") + } + let program_value = align_tabular_trim( + program_line[2:program_line.length()].to_owned(), + ) + match align_tabular_program_header(program_value) { + Some((parsed_program, parsed_version)) => + if !align_tabular_known_fasta_program(parsed_program) { + raise AlignTabularError( + "command line must be followed by a FASTA-suite program header", + ) + } else { + program = parsed_program + version = parsed_version + } + None => raise AlignTabularError("invalid FASTA program header") + } + index = index + 1 + } + } + let mut query_id = "" + let mut query_description = "" + let mut query_length : Int? = None + let mut query_unit = "" + let mut database = "" + let mut rid = "" + let mut field_names : Array[String] = [] + let mut declared_hits : Int? = None + let mut has_query = false + let mut has_database = false + let mut has_rid = false + let mut has_fields = false + while index < lines.length() { + let header = align_tabular_trim(lines[index]) + if !align_tabular_starts_with(header, "# ") { + break + } + let value = align_tabular_trim(header[2:header.length()].to_owned()) + match align_tabular_parse_hits(value) { + Some(count) => { + if count < 0 { + raise AlignTabularError("hit count must be non-negative") + } + declared_hits = Some(count) + index = index + 1 + break + } + None => () + } + if align_tabular_starts_with(value, "Query:") { + if has_query { + raise AlignTabularError("duplicate Query header") + } + has_query = true + let parsed = align_tabular_parse_query_header( + value[6:value.length()].to_owned(), + fasta_style, + ) + query_id = parsed.0 + query_description = parsed.1 + query_length = parsed.2 + query_unit = parsed.3 + } else if align_tabular_starts_with(value, "Database:") { + if has_database { + raise AlignTabularError("duplicate Database header") + } + has_database = true + database = align_tabular_trim(value[9:value.length()].to_owned()) + } else if align_tabular_starts_with(value, "RID:") { + if has_rid { + raise AlignTabularError("duplicate RID header") + } + has_rid = true + rid = align_tabular_trim(value[4:value.length()].to_owned()) + } else if align_tabular_starts_with(value, "Fields:") { + if has_fields { + raise AlignTabularError("duplicate Fields header") + } + has_fields = true + field_names = align_tabular_fields_header( + value[7:value.length()].to_owned(), + ) + align_tabular_validate_field_names(field_names) + } else { + raise AlignTabularError("unexpected header line: " + value) + } + index = index + 1 + } + let hit_count = match declared_hits { + Some(value) => value + None => raise AlignTabularError("query block is missing a hit count") + } + if query_id.length() == 0 { + raise AlignTabularError("query block is missing a Query header") + } + if database.length() == 0 { + raise AlignTabularError("query block is missing a Database header") + } + if hit_count > 0 && field_names.length() == 0 { + raise AlignTabularError("non-empty query block is missing Fields") + } + let metadata = AlignTabularMetadata::{ + program, + version, + command_line, + database, + rid, + } + let alignments : Array[AlignTabularAlignment] = [] + while index < lines.length() { + let row = align_tabular_trim(lines[index]) + if row.length() == 0 { + index = index + 1 + continue + } + if align_tabular_starts_with(row, "# ") { + break + } + alignments.push( + align_tabular_alignment_from_row( + metadata, query_id, query_description, query_length, field_names, row, + ), + ) + index = index + 1 + } + if alignments.length() != hit_count { + raise AlignTabularError( + "declared " + + hit_count.to_string() + + " hits but parsed " + + alignments.length().to_string() + + " rows", + ) + } + let mut inferred_query_length : Int? = None + for alignment in alignments { + if inferred_query_length is None && alignment.query_length is Some(_) { + inferred_query_length = alignment.query_length + } + } + let final_query_length = if query_length is Some(_) { + query_length + } else { + inferred_query_length + } + for alignment in alignments { + if final_query_length is Some(_) && + alignment.query_length is Some(_) && + final_query_length != alignment.query_length { + raise AlignTabularError("query length changes within a query block") + } + } + queries.push(AlignTabularQueryResult::{ + metadata, + query_id, + query_description, + query_length: final_query_length, + query_unit, + field_names: align_tabular_copy_strings(field_names), + declared_hits: hit_count, + alignments, + }) + } + match processed_queries { + Some(count) => + if count != queries.length() { + raise AlignTabularError( + "processed query count does not match parsed query blocks", + ) + } + None => () + } + AlignTabularDocument::{ queries, processed_queries } +} + +///| +/// Return an original row field. +pub fn AlignTabularAlignment::field( + self : AlignTabularAlignment, + name : String, +) -> String? { + align_tabular_field_value(self.fields, name) +} + +///| +/// Return an integer row field. +pub fn AlignTabularAlignment::integer_field( + self : AlignTabularAlignment, + name : String, +) -> Int? raise AlignTabularError { + align_tabular_optional_int(self.fields, name) +} + +///| +/// Return a floating-point row field. +pub fn AlignTabularAlignment::number_field( + self : AlignTabularAlignment, + name : String, +) -> Double? raise AlignTabularError { + align_tabular_optional_double(self.fields, name) +} + +///| +/// Number of printed alignment columns represented by the path. +pub fn AlignTabularAlignment::alignment_columns( + self : AlignTabularAlignment, +) -> Int { + let mut result = 0 + for index in 0.. query { target } else { query }) + } + result +} + +///| +/// Number of aligned residue pairs represented by the path. +pub fn AlignTabularAlignment::aligned_residues( + self : AlignTabularAlignment, +) -> Int { + let mut result = 0 + for index in 0.. Int { + let mut result = 0 + for index in 0.. Int { + let mut result = 0 + for index in 0.. Int { + let mut result = 0 + for operation in self.operations { + if operation != AlignTabularOperation::Aligned { + result = result + 1 + } + } + result +} + +///| +/// True if the query was reported on the reverse strand. +pub fn AlignTabularAlignment::query_is_reverse( + self : AlignTabularAlignment, +) -> Bool { + self.query_strand == Strand::Minus +} + +///| +/// True if the subject was reported on the reverse strand. +pub fn AlignTabularAlignment::target_is_reverse( + self : AlignTabularAlignment, +) -> Bool { + self.target_strand == Strand::Minus +} + +///| +/// Convert the path to the repository's coordinate alignment abstraction. +/// +/// The target axis is normalized to increasing order; the query axis then +/// represents relative strand orientation. +pub fn AlignTabularAlignment::coordinate_alignment( + self : AlignTabularAlignment, +) -> CoordinatePairwiseAlignment raise AlignTabularError { + if self.target_coordinates.length() == 0 || !self.path_is_absolute { + raise AlignTabularError( + "an absolute traceback path is required for coordinate conversion", + ) + } + let mut target = align_tabular_copy_ints(self.target_coordinates) + let mut query = align_tabular_copy_ints(self.query_coordinates) + if target[0] > target[target.length() - 1] { + let reversed_target : Array[Int] = [] + let reversed_query : Array[Int] = [] + let mut index = target.length() + while index > 0 { + index = index - 1 + reversed_target.push(target[index]) + reversed_query.push(query[index]) + } + target = reversed_target + query = reversed_query + } + let mut inferred_target_length = 0 + let mut inferred_query_length = 0 + for coordinate in target { + if coordinate > inferred_target_length { + inferred_target_length = coordinate + } + } + for coordinate in query { + if coordinate > inferred_query_length { + inferred_query_length = coordinate + } + } + let target_length = self.target_length.unwrap_or(inferred_target_length) + let query_length = self.query_length.unwrap_or(inferred_query_length) + coordinate_pairwise_alignment_with_lengths( + self.target_id, + target_length, + self.query_id, + query_length, + target, + query, + ) catch { + AlignmentMapError(message) => + raise AlignTabularError("coordinate conversion failed: " + message) + } +} + +///| +/// Compact human-readable alignment summary. +pub fn AlignTabularAlignment::summary(self : AlignTabularAlignment) -> String { + let mut result = self.query_id + " -> " + self.target_id + match self.field("evalue") { + Some(value) => result = result + ", E=" + value + None => () + } + match self.field("bit score") { + Some(value) => result = result + ", bits=" + value + None => () + } + if self.operations.length() > 0 { + result = result + + ", columns=" + + self.alignment_columns().to_string() + + ", gaps=" + + (self.query_gap_residues() + self.target_gap_residues()).to_string() + } + result +} + +///| +/// Flatten all query blocks into input-order alignments. +pub fn AlignTabularDocument::alignments( + self : AlignTabularDocument, +) -> Array[AlignTabularAlignment] { + let result : Array[AlignTabularAlignment] = [] + for query in self.queries { + for alignment in query.alignments { + result.push(alignment) + } + } + result +} + +///| +/// Find a query block by identifier. +pub fn AlignTabularDocument::find_query( + self : AlignTabularDocument, + query_id : String, +) -> AlignTabularQueryResult? { + for query in self.queries { + if query.query_id == query_id { + return Some(query) + } + } + None +} + +///| +/// Return the lowest-E-value alignment in a query block. +pub fn AlignTabularQueryResult::best_by_evalue( + self : AlignTabularQueryResult, +) -> AlignTabularAlignment? raise AlignTabularError { + let mut best : AlignTabularAlignment? = None + let mut best_value = 0.0 + for alignment in self.alignments { + match alignment.number_field("evalue") { + Some(value) => + if best is None || value < best_value { + best = Some(alignment) + best_value = value + } + None => () + } + } + best +} + +///| +/// Return alignments with E-value at or below a threshold. +pub fn AlignTabularQueryResult::filter_evalue( + self : AlignTabularQueryResult, + threshold : Double, +) -> Array[AlignTabularAlignment] raise AlignTabularError { + if threshold.is_nan() || threshold < 0.0 || threshold.abs() > 1.0e300 { + raise AlignTabularError("E-value threshold must be finite and non-negative") + } + let result : Array[AlignTabularAlignment] = [] + for alignment in self.alignments { + match alignment.number_field("evalue") { + Some(value) => if value <= threshold { result.push(alignment) } + None => () + } + } + result +} diff --git a/src/alignace.mbt b/src/alignace.mbt index c8225bd2..7b1ec010 100644 --- a/src/alignace.mbt +++ b/src/alignace.mbt @@ -45,7 +45,11 @@ pub struct AlignAceMotif { ///| /// Construct an AlignAceMotif from a count matrix. pub fn AlignAceMotif::new(count_matrix : Array[Array[Int]]) -> AlignAceMotif { - let width = if count_matrix.length() > 0 { count_matrix[0].length() } else { 0 } + let width = if count_matrix.length() > 0 { + count_matrix[0].length() + } else { + 0 + } let mut total = 0 if width > 0 { for base_idx in 0..<4 { @@ -124,7 +128,10 @@ pub fn AlignAceMotif::add_site(self : AlignAceMotif, s : AlignAceSite) -> Unit { ///| /// Add a motif to the record. -pub fn AlignAceRecord::add_motif(self : AlignAceRecord, m : AlignAceMotif) -> Unit { +pub fn AlignAceRecord::add_motif( + self : AlignAceRecord, + m : AlignAceMotif, +) -> Unit { self.motifs.push(m) } @@ -218,10 +225,7 @@ pub fn alignace_parse(text : String) -> AlignAceRecord { ///| /// Parse a motif matrix starting at the given line index. /// Returns Some(motif) if successful, None otherwise. -fn alignace_parse_motif( - lines : Array[String], - start : Int, -) -> AlignAceMotif? { +fn alignace_parse_motif(lines : Array[String], start : Int) -> AlignAceMotif? { // Line 0: "i T G A C T C G A T" (consensus letters) // Line 1: " 0 1 2 3 4 5 6 7 8 9" (column indices) // Line 2: "A 1 0 8 0 0 0 0 0 0 0" @@ -269,23 +273,13 @@ fn alignace_parse_site(text : String) -> AlignAceSite? { let strand = parts[2] let sequence = parts[3] return Some( - AlignAceSite::new( - sequence_id=seq_id, - position=pos, - strand=strand, - sequence=sequence, - ), + AlignAceSite::new(sequence_id=seq_id, position=pos, strand~, sequence~), ) } // No strand field, assume "+" let sequence = parts[2] Some( - AlignAceSite::new( - sequence_id=seq_id, - position=pos, - strand="+", - sequence=sequence, - ), + AlignAceSite::new(sequence_id=seq_id, position=pos, strand="+", sequence~), ) } @@ -360,7 +354,8 @@ pub fn alignace_to_pwm(motif : AlignAceMotif) -> Array[Array[Double]] { motif.count_matrix[2][col] + motif.count_matrix[3][col] if total > 0 { - freq_row[col] = motif.count_matrix[row][col].to_double() / total.to_double() + freq_row[col] = motif.count_matrix[row][col].to_double() / + total.to_double() } else { freq_row[col] = 0.25 } @@ -474,8 +469,15 @@ fn alignace_write_motif(motif : AlignAceMotif) -> String { sb.write_string("# Sites:\n") for site in motif.sites { sb.write_string( - "# " + site.sequence_id + "\t" + site.position.to_string() + "\t" + - site.strand + "\t" + site.sequence + "\n", + "# " + + site.sequence_id + + "\t" + + site.position.to_string() + + "\t" + + site.strand + + "\t" + + site.sequence + + "\n", ) } sb.write_string("#\n") @@ -495,7 +497,10 @@ pub fn alignace_num_motifs(record : AlignAceRecord) -> Int { ///| /// Get a motif by index (0-based). Returns None if out of range. -pub fn alignace_get_motif(record : AlignAceRecord, index : Int) -> AlignAceMotif? { +pub fn alignace_get_motif( + record : AlignAceRecord, + index : Int, +) -> AlignAceMotif? { if index >= 0 && index < record.motifs.length() { Some(record.motifs[index]) } else { @@ -510,7 +515,9 @@ pub fn alignace_summary(record : AlignAceRecord) -> String { sb.write_string("AlignACE Record Summary:\n") sb.write_string(" Version: " + record.version + "\n") sb.write_string(" Command: " + record.command + "\n") - sb.write_string(" Parameters: " + record.parameters.size().to_string() + "\n") + sb.write_string( + " Parameters: " + record.parameters.size().to_string() + "\n", + ) let keys = record.parameters.keys() for k in keys { sb.write_string(" " + k + " = " + record.parameters[k] + "\n") @@ -519,9 +526,15 @@ pub fn alignace_summary(record : AlignAceRecord) -> String { for i in 0.. AlignAceRecord { record.parameters["numcols"] = "10" record.parameters["expect"] = "10" // Motif 1: TGACTCGAT - let m1 = AlignAceMotif::new( - [ - [1, 0, 8, 0, 0, 0, 0, 0, 0, 0], - [0, 0, 0, 0, 1, 0, 8, 0, 0, 1], - [0, 8, 0, 0, 0, 0, 0, 8, 0, 0], - [7, 0, 0, 8, 7, 8, 0, 0, 8, 7], - ], - ) + let m1 = AlignAceMotif::new([ + [1, 0, 8, 0, 0, 0, 0, 0, 0, 0], + [0, 0, 0, 0, 1, 0, 8, 0, 0, 1], + [0, 8, 0, 0, 0, 0, 0, 8, 0, 0], + [7, 0, 0, 8, 7, 8, 0, 0, 8, 7], + ]) m1.sites.push( AlignAceSite::new( sequence_id="seq1", @@ -563,14 +574,12 @@ pub fn alignace_sample() -> AlignAceRecord { ) record.motifs.push(m1) // Motif 2: AATAAACAAA - let m2 = AlignAceMotif::new( - [ - [8, 8, 1, 8, 8, 8, 0, 8, 8, 8], - [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], - [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], - [1, 1, 8, 1, 1, 1, 9, 1, 1, 1], - ], - ) + let m2 = AlignAceMotif::new([ + [8, 8, 1, 8, 8, 8, 0, 8, 8, 8], + [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], + [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], + [1, 1, 8, 1, 1, 1, 9, 1, 1, 1], + ]) record.motifs.push(m2) record } diff --git a/src/alignment_counts.mbt b/src/alignment_counts.mbt new file mode 100644 index 00000000..a40a8f58 --- /dev/null +++ b/src/alignment_counts.mbt @@ -0,0 +1,849 @@ +// Detailed coordinate alignment statistics compatible with Biopython +// Bio.Align.Alignment.counts and AlignmentCounts. + +///| +/// Error raised while configuring or calculating detailed alignment counts. +pub suberror AlignmentCountsError { + AlignmentCountsError(String) +} + +///| +/// Affine gap scores for each gap side and direction. +/// +/// An insertion is a gap in the target row; a deletion is a gap in the query +/// row. A gap of length n contributes one open score and n - 1 extend scores. +pub struct AlignmentGapScores { + open_left_insertion_score : Double + extend_left_insertion_score : Double + open_left_deletion_score : Double + extend_left_deletion_score : Double + open_internal_insertion_score : Double + extend_internal_insertion_score : Double + open_internal_deletion_score : Double + extend_internal_deletion_score : Double + open_right_insertion_score : Double + extend_right_insertion_score : Double + open_right_deletion_score : Double + extend_right_deletion_score : Double +} derive(Eq, Debug) + +///| +fn alignment_counts_is_finite(value : Double) -> Bool { + value == value && value.abs() <= 1.0e300 +} + +///| +fn alignment_counts_validate_score( + value : Double, + label : String, +) -> Unit raise AlignmentCountsError { + if !alignment_counts_is_finite(value) { + raise AlignmentCountsError(label + " must be finite") + } +} + +///| +/// Construct fully position- and direction-specific affine gap scores. +pub fn AlignmentGapScores::create( + open_left_insertion_score~ : Double, + extend_left_insertion_score~ : Double, + open_left_deletion_score~ : Double, + extend_left_deletion_score~ : Double, + open_internal_insertion_score~ : Double, + extend_internal_insertion_score~ : Double, + open_internal_deletion_score~ : Double, + extend_internal_deletion_score~ : Double, + open_right_insertion_score~ : Double, + extend_right_insertion_score~ : Double, + open_right_deletion_score~ : Double, + extend_right_deletion_score~ : Double, +) -> AlignmentGapScores raise AlignmentCountsError { + let values = [ + open_left_insertion_score, extend_left_insertion_score, open_left_deletion_score, + extend_left_deletion_score, open_internal_insertion_score, extend_internal_insertion_score, + open_internal_deletion_score, extend_internal_deletion_score, open_right_insertion_score, + extend_right_insertion_score, open_right_deletion_score, extend_right_deletion_score, + ] + let labels = [ + "open left insertion score", "extend left insertion score", "open left deletion score", + "extend left deletion score", "open internal insertion score", "extend internal insertion score", + "open internal deletion score", "extend internal deletion score", "open right insertion score", + "extend right insertion score", "open right deletion score", "extend right deletion score", + ] + let mut index = 0 + while index < values.length() { + alignment_counts_validate_score(values[index], labels[index]) + index = index + 1 + } + AlignmentGapScores::{ + open_left_insertion_score, + extend_left_insertion_score, + open_left_deletion_score, + extend_left_deletion_score, + open_internal_insertion_score, + extend_internal_insertion_score, + open_internal_deletion_score, + extend_internal_deletion_score, + open_right_insertion_score, + extend_right_insertion_score, + open_right_deletion_score, + extend_right_deletion_score, + } +} + +///| +/// Use one open score and one extend score for every gap class. +pub fn AlignmentGapScores::affine( + open_score : Double, + extend_score : Double, +) -> AlignmentGapScores raise AlignmentCountsError { + alignment_counts_validate_score(open_score, "gap open score") + alignment_counts_validate_score(extend_score, "gap extend score") + AlignmentGapScores::{ + open_left_insertion_score: open_score, + extend_left_insertion_score: extend_score, + open_left_deletion_score: open_score, + extend_left_deletion_score: extend_score, + open_internal_insertion_score: open_score, + extend_internal_insertion_score: extend_score, + open_internal_deletion_score: open_score, + extend_internal_deletion_score: extend_score, + open_right_insertion_score: open_score, + extend_right_insertion_score: extend_score, + open_right_deletion_score: open_score, + extend_right_deletion_score: extend_score, + } +} + +///| +/// Use direction-specific affine scores independent of gap position. +pub fn AlignmentGapScores::directional( + insertion_open_score~ : Double, + insertion_extend_score~ : Double, + deletion_open_score~ : Double, + deletion_extend_score~ : Double, +) -> AlignmentGapScores raise AlignmentCountsError { + alignment_counts_validate_score(insertion_open_score, "insertion open score") + alignment_counts_validate_score( + insertion_extend_score, "insertion extend score", + ) + alignment_counts_validate_score(deletion_open_score, "deletion open score") + alignment_counts_validate_score( + deletion_extend_score, "deletion extend score", + ) + AlignmentGapScores::{ + open_left_insertion_score: insertion_open_score, + extend_left_insertion_score: insertion_extend_score, + open_left_deletion_score: deletion_open_score, + extend_left_deletion_score: deletion_extend_score, + open_internal_insertion_score: insertion_open_score, + extend_internal_insertion_score: insertion_extend_score, + open_internal_deletion_score: deletion_open_score, + extend_internal_deletion_score: deletion_extend_score, + open_right_insertion_score: insertion_open_score, + extend_right_insertion_score: insertion_extend_score, + open_right_deletion_score: deletion_open_score, + extend_right_deletion_score: deletion_extend_score, + } +} + +///| +/// Optional composition and gap scoring used by `alignment_counts`. +/// +/// A substitution matrix takes precedence over match/mismatch scores. +/// Substitution and gap scores remain unknown unless their corresponding +/// configuration is present. +pub struct AlignmentCountsConfig { + wildcard : Char? + substitution_matrix : SubstitutionMatrix? + match_score : Double? + mismatch_score : Double? + gap_scores : AlignmentGapScores? +} + +///| +pub fn AlignmentCountsConfig::create( + wildcard? : Char? = None, + substitution_matrix? : SubstitutionMatrix? = None, + match_score? : Double? = None, + mismatch_score? : Double? = None, + gap_scores? : AlignmentGapScores? = None, +) -> AlignmentCountsConfig raise AlignmentCountsError { + match (match_score, mismatch_score) { + (Some(match_value), Some(mismatch_value)) => { + alignment_counts_validate_score(match_value, "match score") + alignment_counts_validate_score(mismatch_value, "mismatch score") + } + (None, None) => () + _ => + raise AlignmentCountsError( + "match and mismatch scores must be provided together", + ) + } + AlignmentCountsConfig::{ + wildcard, + substitution_matrix, + match_score, + mismatch_score, + gap_scores, + } +} + +///| +pub fn AlignmentCountsConfig::default() -> AlignmentCountsConfig { + AlignmentCountsConfig::{ + wildcard: None, + substitution_matrix: None, + match_score: None, + mismatch_score: None, + gap_scores: None, + } +} + +///| +/// Configure wildcard-aware counting without calculating scores. +pub fn AlignmentCountsConfig::with_wildcard( + wildcard : Char, +) -> AlignmentCountsConfig { + AlignmentCountsConfig::{ + wildcard: Some(wildcard), + substitution_matrix: None, + match_score: None, + mismatch_score: None, + gap_scores: None, + } +} + +///| +/// Configure substitution scoring and positive-match counting. +pub fn AlignmentCountsConfig::with_matrix( + matrix : SubstitutionMatrix, + wildcard? : Char? = None, +) -> AlignmentCountsConfig { + AlignmentCountsConfig::{ + wildcard, + substitution_matrix: Some(matrix), + match_score: None, + mismatch_score: None, + gap_scores: None, + } +} + +///| +/// Configure complete match/mismatch and affine gap scoring. +pub fn AlignmentCountsConfig::scored( + match_score~ : Double, + mismatch_score~ : Double, + gap_scores~ : AlignmentGapScores, + wildcard? : Char? = None, +) -> AlignmentCountsConfig raise AlignmentCountsError { + alignment_counts_validate_score(match_score, "match score") + alignment_counts_validate_score(mismatch_score, "mismatch score") + AlignmentCountsConfig::{ + wildcard, + substitution_matrix: None, + match_score: Some(match_score), + mismatch_score: Some(mismatch_score), + gap_scores: Some(gap_scores), + } +} + +///| +/// Configure complete substitution-matrix and affine gap scoring. +pub fn AlignmentCountsConfig::scored_matrix( + matrix~ : SubstitutionMatrix, + gap_scores~ : AlignmentGapScores, + wildcard? : Char? = None, +) -> AlignmentCountsConfig { + AlignmentCountsConfig::{ + wildcard, + substitution_matrix: Some(matrix), + match_score: None, + mismatch_score: None, + gap_scores: Some(gap_scores), + } +} + +///| +/// Detailed counts and optional scores for a coordinate alignment. +pub struct AlignmentCounts { + open_left_insertions : Int + extend_left_insertions : Int + open_left_deletions : Int + extend_left_deletions : Int + open_internal_insertions : Int + extend_internal_insertions : Int + open_internal_deletions : Int + extend_internal_deletions : Int + open_right_insertions : Int + extend_right_insertions : Int + open_right_deletions : Int + extend_right_deletions : Int + aligned : Int + identities : Int + mismatches : Int + positives : Int? + gap_score : Double? + substitution_score : Double? + score : Double? +} derive(Debug) + +///| +pub fn AlignmentCounts::left_insertions(self : AlignmentCounts) -> Int { + self.open_left_insertions + self.extend_left_insertions +} + +///| +pub fn AlignmentCounts::internal_insertions(self : AlignmentCounts) -> Int { + self.open_internal_insertions + self.extend_internal_insertions +} + +///| +pub fn AlignmentCounts::right_insertions(self : AlignmentCounts) -> Int { + self.open_right_insertions + self.extend_right_insertions +} + +///| +pub fn AlignmentCounts::left_deletions(self : AlignmentCounts) -> Int { + self.open_left_deletions + self.extend_left_deletions +} + +///| +pub fn AlignmentCounts::internal_deletions(self : AlignmentCounts) -> Int { + self.open_internal_deletions + self.extend_internal_deletions +} + +///| +pub fn AlignmentCounts::right_deletions(self : AlignmentCounts) -> Int { + self.open_right_deletions + self.extend_right_deletions +} + +///| +pub fn AlignmentCounts::open_insertions(self : AlignmentCounts) -> Int { + self.open_left_insertions + + self.open_internal_insertions + + self.open_right_insertions +} + +///| +pub fn AlignmentCounts::extend_insertions(self : AlignmentCounts) -> Int { + self.extend_left_insertions + + self.extend_internal_insertions + + self.extend_right_insertions +} + +///| +pub fn AlignmentCounts::insertions(self : AlignmentCounts) -> Int { + self.open_insertions() + self.extend_insertions() +} + +///| +pub fn AlignmentCounts::open_deletions(self : AlignmentCounts) -> Int { + self.open_left_deletions + + self.open_internal_deletions + + self.open_right_deletions +} + +///| +pub fn AlignmentCounts::extend_deletions(self : AlignmentCounts) -> Int { + self.extend_left_deletions + + self.extend_internal_deletions + + self.extend_right_deletions +} + +///| +pub fn AlignmentCounts::deletions(self : AlignmentCounts) -> Int { + self.open_deletions() + self.extend_deletions() +} + +///| +pub fn AlignmentCounts::open_left_gaps(self : AlignmentCounts) -> Int { + self.open_left_insertions + self.open_left_deletions +} + +///| +pub fn AlignmentCounts::extend_left_gaps(self : AlignmentCounts) -> Int { + self.extend_left_insertions + self.extend_left_deletions +} + +///| +pub fn AlignmentCounts::left_gaps(self : AlignmentCounts) -> Int { + self.open_left_gaps() + self.extend_left_gaps() +} + +///| +pub fn AlignmentCounts::open_internal_gaps(self : AlignmentCounts) -> Int { + self.open_internal_insertions + self.open_internal_deletions +} + +///| +pub fn AlignmentCounts::extend_internal_gaps(self : AlignmentCounts) -> Int { + self.extend_internal_insertions + self.extend_internal_deletions +} + +///| +pub fn AlignmentCounts::internal_gaps(self : AlignmentCounts) -> Int { + self.open_internal_gaps() + self.extend_internal_gaps() +} + +///| +pub fn AlignmentCounts::open_right_gaps(self : AlignmentCounts) -> Int { + self.open_right_insertions + self.open_right_deletions +} + +///| +pub fn AlignmentCounts::extend_right_gaps(self : AlignmentCounts) -> Int { + self.extend_right_insertions + self.extend_right_deletions +} + +///| +pub fn AlignmentCounts::right_gaps(self : AlignmentCounts) -> Int { + self.open_right_gaps() + self.extend_right_gaps() +} + +///| +pub fn AlignmentCounts::open_gaps(self : AlignmentCounts) -> Int { + self.open_insertions() + self.open_deletions() +} + +///| +pub fn AlignmentCounts::extend_gaps(self : AlignmentCounts) -> Int { + self.extend_insertions() + self.extend_deletions() +} + +///| +pub fn AlignmentCounts::gaps(self : AlignmentCounts) -> Int { + self.open_gaps() + self.extend_gaps() +} + +///| +pub fn AlignmentCounts::summary(self : AlignmentCounts) -> String { + "AlignmentCounts(aligned=" + + self.aligned.to_string() + + ", identities=" + + self.identities.to_string() + + ", mismatches=" + + self.mismatches.to_string() + + ", gaps=" + + self.gaps().to_string() + + ")" +} + +///| +struct AlignmentCountsAccumulator { + mut open_left_insertions : Int + mut extend_left_insertions : Int + mut open_left_deletions : Int + mut extend_left_deletions : Int + mut open_internal_insertions : Int + mut extend_internal_insertions : Int + mut open_internal_deletions : Int + mut extend_internal_deletions : Int + mut open_right_insertions : Int + mut extend_right_insertions : Int + mut open_right_deletions : Int + mut extend_right_deletions : Int + mut aligned : Int + mut identities : Int + mut mismatches : Int + mut positives : Int + mut matrix_score : Double +} + +///| +fn alignment_counts_accumulator() -> AlignmentCountsAccumulator { + AlignmentCountsAccumulator::{ + open_left_insertions: 0, + extend_left_insertions: 0, + open_left_deletions: 0, + extend_left_deletions: 0, + open_internal_insertions: 0, + extend_internal_insertions: 0, + open_internal_deletions: 0, + extend_internal_deletions: 0, + open_right_insertions: 0, + extend_right_insertions: 0, + open_right_deletions: 0, + extend_right_deletions: 0, + aligned: 0, + identities: 0, + mismatches: 0, + positives: 0, + matrix_score: 0.0, + } +} + +///| +fn alignment_counts_normalize_coordinates( + coordinates : Array[Int], + sequence_length : Int, + reverse : Bool, +) -> Array[Int] { + let normalized : Array[Int] = [] + for coordinate in coordinates { + normalized.push( + if reverse { + sequence_length - coordinate + } else { + coordinate + }, + ) + } + normalized +} + +///| +fn alignment_counts_complement(value : Char) -> Char { + match value { + 'A' => 'T' + 'a' => 't' + 'C' => 'G' + 'c' => 'g' + 'G' => 'C' + 'g' => 'c' + 'T' => 'A' + 't' => 'a' + 'U' => 'A' + 'u' => 'a' + _ => value + } +} + +///| +fn alignment_counts_residue( + sequence : String, + sequence_length : Int, + position : Int, + reverse : Bool, +) -> Char { + if reverse { + alignment_counts_complement( + sequence.unsafe_get(sequence_length - position - 1).unsafe_to_char(), + ) + } else { + sequence.unsafe_get(position).unsafe_to_char() + } +} + +///| +fn alignment_counts_add_gap( + counts : AlignmentCountsAccumulator, + insertion : Bool, + continuation : Bool, + start : Int, + end : Int, + left : Int, + right : Int, + size : Int, +) -> Unit { + let side = if start == left { 0 } else if end == right { 2 } else { 1 } + let open = if continuation { 0 } else { 1 } + let extend = if continuation { size } else { size - 1 } + if insertion { + if side == 0 { + counts.open_left_insertions = counts.open_left_insertions + open + counts.extend_left_insertions = counts.extend_left_insertions + extend + } else if side == 1 { + counts.open_internal_insertions = counts.open_internal_insertions + open + counts.extend_internal_insertions = counts.extend_internal_insertions + + extend + } else { + counts.open_right_insertions = counts.open_right_insertions + open + counts.extend_right_insertions = counts.extend_right_insertions + extend + } + } else if side == 0 { + counts.open_left_deletions = counts.open_left_deletions + open + counts.extend_left_deletions = counts.extend_left_deletions + extend + } else if side == 1 { + counts.open_internal_deletions = counts.open_internal_deletions + open + counts.extend_internal_deletions = counts.extend_internal_deletions + extend + } else { + counts.open_right_deletions = counts.open_right_deletions + open + counts.extend_right_deletions = counts.extend_right_deletions + extend + } +} + +///| +fn alignment_counts_add_residues( + counts : AlignmentCountsAccumulator, + first : Char, + second : Char, + config : AlignmentCountsConfig, +) -> Unit raise AlignmentCountsError { + match config.wildcard { + Some(wildcard) => if first == wildcard || second == wildcard { return } + None => () + } + if first == second { + counts.identities = counts.identities + 1 + } else { + counts.mismatches = counts.mismatches + 1 + } + match config.substitution_matrix { + Some(matrix) => { + let first_text = first.to_string() + let second_text = second.to_string() + if !matrix.contains(first_text) { + raise AlignmentCountsError( + "target residue '" + first_text + "' is not in substitution matrix", + ) + } + if !matrix.contains(second_text) { + raise AlignmentCountsError( + "query residue '" + second_text + "' is not in substitution matrix", + ) + } + let value = matrix.get_score_case_insensitive(first_text, second_text) + counts.matrix_score = counts.matrix_score + value.to_double() + if value > 0 { + counts.positives = counts.positives + 1 + } + } + None => () + } +} + +///| +fn alignment_counts_add_pair( + counts : AlignmentCountsAccumulator, + first_sequence : String, + first_length : Int, + first_coordinates : Array[Int], + first_reverse : Bool, + second_sequence : String, + second_length : Int, + second_coordinates : Array[Int], + second_reverse : Bool, + config : AlignmentCountsConfig, +) -> Unit raise AlignmentCountsError { + if first_coordinates.length() != second_coordinates.length() { + raise AlignmentCountsError( + "alignment coordinate rows must have equal length", + ) + } + if first_coordinates.length() == 1 { + raise AlignmentCountsError( + "an alignment path must be empty or contain at least two points", + ) + } + let first = alignment_counts_normalize_coordinates( + first_coordinates, first_length, first_reverse, + ) + let second = alignment_counts_normalize_coordinates( + second_coordinates, second_length, second_reverse, + ) + if first.length() == 0 { + return + } + let first_left = first[0] + let first_right = first[first.length() - 1] + let second_left = second[0] + let second_right = second[second.length() - 1] + let mut path = 0 + let mut index = 0 + while index + 1 < first.length() { + let first_start = first[index] + let first_end = first[index + 1] + let second_start = second[index] + let second_end = second[index + 1] + let first_step = first_end - first_start + let second_step = second_end - second_start + if first_step < 0 || second_step < 0 { + raise AlignmentCountsError( + "normalized alignment coordinates must be non-decreasing", + ) + } + if first_step == 0 && second_step == 0 { + // Multiple-alignment row pairs may share a gap inserted for other rows. + } else if first_step == 0 { + alignment_counts_add_gap( + counts, + true, + path == 1, + first_start, + first_end, + first_left, + first_right, + second_step, + ) + path = 1 + } else if second_step == 0 { + alignment_counts_add_gap( + counts, + false, + path == 2, + second_start, + second_end, + second_left, + second_right, + first_step, + ) + path = 2 + } else { + if first_step != second_step { + raise AlignmentCountsError( + "aligned coordinate steps must have equal length", + ) + } + path = 0 + counts.aligned = counts.aligned + first_step + if first_sequence.length() > 0 && second_sequence.length() > 0 { + let mut offset = 0 + while offset < first_step { + let first_residue = alignment_counts_residue( + first_sequence, + first_length, + first_start + offset, + first_reverse, + ) + let second_residue = alignment_counts_residue( + second_sequence, + second_length, + second_start + offset, + second_reverse, + ) + alignment_counts_add_residues( + counts, first_residue, second_residue, config, + ) + offset = offset + 1 + } + } + } + index = index + 1 + } +} + +///| +fn alignment_counts_gap_score( + counts : AlignmentCountsAccumulator, + scores : AlignmentGapScores, +) -> Double { + counts.open_left_insertions.to_double() * scores.open_left_insertion_score + + counts.extend_left_insertions.to_double() * scores.extend_left_insertion_score + + counts.open_left_deletions.to_double() * scores.open_left_deletion_score + + counts.extend_left_deletions.to_double() * scores.extend_left_deletion_score + + counts.open_internal_insertions.to_double() * + scores.open_internal_insertion_score + + counts.extend_internal_insertions.to_double() * + scores.extend_internal_insertion_score + + counts.open_internal_deletions.to_double() * + scores.open_internal_deletion_score + + counts.extend_internal_deletions.to_double() * + scores.extend_internal_deletion_score + + counts.open_right_insertions.to_double() * scores.open_right_insertion_score + + counts.extend_right_insertions.to_double() * + scores.extend_right_insertion_score + + counts.open_right_deletions.to_double() * scores.open_right_deletion_score + + counts.extend_right_deletions.to_double() * scores.extend_right_deletion_score +} + +///| +fn alignment_counts_finish( + counts : AlignmentCountsAccumulator, + config : AlignmentCountsConfig, +) -> AlignmentCounts { + let positives = match config.substitution_matrix { + Some(_) => Some(counts.positives) + None => None + } + let substitution_score = match config.substitution_matrix { + Some(_) => Some(counts.matrix_score) + None => + match (config.match_score, config.mismatch_score) { + (Some(match_value), Some(mismatch_value)) => + Some( + counts.identities.to_double() * match_value + + counts.mismatches.to_double() * mismatch_value, + ) + _ => None + } + } + let gap_score = match config.gap_scores { + Some(scores) => Some(alignment_counts_gap_score(counts, scores)) + None => None + } + let score = match (substitution_score, gap_score) { + (Some(substitution_value), Some(gap_value)) => + Some(substitution_value + gap_value) + _ => None + } + AlignmentCounts::{ + open_left_insertions: counts.open_left_insertions, + extend_left_insertions: counts.extend_left_insertions, + open_left_deletions: counts.open_left_deletions, + extend_left_deletions: counts.extend_left_deletions, + open_internal_insertions: counts.open_internal_insertions, + extend_internal_insertions: counts.extend_internal_insertions, + open_internal_deletions: counts.open_internal_deletions, + extend_internal_deletions: counts.extend_internal_deletions, + open_right_insertions: counts.open_right_insertions, + extend_right_insertions: counts.extend_right_insertions, + open_right_deletions: counts.open_right_deletions, + extend_right_deletions: counts.extend_right_deletions, + aligned: counts.aligned, + identities: counts.identities, + mismatches: counts.mismatches, + positives, + gap_score, + substitution_score, + score, + } +} + +///| +/// Calculate Biopython-style detailed statistics for a pairwise alignment. +pub fn CoordinatePairwiseAlignment::alignment_counts( + self : CoordinatePairwiseAlignment, + config? : AlignmentCountsConfig = AlignmentCountsConfig::default(), +) -> AlignmentCounts raise AlignmentCountsError { + let counts = alignment_counts_accumulator() + alignment_counts_add_pair( + counts, + self.target_sequence, + self.target_length, + self.target_coordinates, + false, + self.query_sequence, + self.query_length, + self.query_coordinates, + self.is_reverse(), + config, + ) + alignment_counts_finish(counts, config) +} + +///| +/// Sum detailed statistics over every unordered pair of MSA rows. +pub fn CoordinateMultipleAlignment::alignment_counts( + self : CoordinateMultipleAlignment, + config? : AlignmentCountsConfig = AlignmentCountsConfig::default(), +) -> AlignmentCounts raise AlignmentCountsError { + if self.names.length() != self.sequences.length() || + self.names.length() != self.coordinates.length() { + raise AlignmentCountsError( + "MSA names, sequences, and coordinates must have equal row counts", + ) + } + let counts = alignment_counts_accumulator() + let mut first = 0 + while first < self.sequences.length() { + let mut second = first + 1 + while second < self.sequences.length() { + alignment_counts_add_pair( + counts, + self.sequences[first], + self.sequences[first].length(), + self.coordinates[first], + false, + self.sequences[second], + self.sequences[second].length(), + self.coordinates[second], + false, + config, + ) + second = second + 1 + } + first = first + 1 + } + alignment_counts_finish(counts, config) +} diff --git a/src/alignment_map.mbt b/src/alignment_map.mbt new file mode 100644 index 00000000..88f6219e --- /dev/null +++ b/src/alignment_map.mbt @@ -0,0 +1,1369 @@ +// Compose coordinate-based alignments without realigning their sequences. +// +// This implements the core semantics of Biopython Bio.Align.Alignment.map +// and mapall. Coordinates are zero-based, half-open sequence boundaries. + +///| +/// Error raised by coordinate alignment construction or projection. +pub suberror AlignmentMapError { + AlignmentMapError(String) +} + +///| +/// One aligned block in a coordinate-based pairwise alignment. +pub struct CoordinateAlignmentBlock { + target_start : Int + target_end : Int + query_start : Int + query_end : Int + size : Int +} derive(Eq, Debug) + +///| +/// Coordinate representation of a pairwise alignment. +/// +/// Sequence text may be empty when only lengths and coordinates are known. +/// Target coordinates are non-decreasing. Query coordinates may increase or +/// decrease, allowing reverse-strand alignments. +pub struct CoordinatePairwiseAlignment { + target_name : String + query_name : String + target_sequence : String + query_sequence : String + target_length : Int + query_length : Int + target_coordinates : Array[Int] + query_coordinates : Array[Int] +} derive(Eq, Debug) + +///| +/// Coordinate representation of a multiple sequence alignment. +pub struct CoordinateMultipleAlignment { + names : Array[String] + sequences : Array[String] + coordinates : Array[Array[Int]] +} derive(Eq, Debug) + +///| +/// Basic alignment statistics derived from a coordinate path. +pub struct CoordinateAlignmentCounts { + aligned : Int + target_gap_bases : Int + query_gap_bases : Int + target_gap_events : Int + query_gap_events : Int + blocks : Int +} derive(Eq, Debug) + +///| +fn alignment_map_copy_ints(values : Array[Int]) -> Array[Int] { + let copy : Array[Int] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn alignment_map_copy_strings(values : Array[String]) -> Array[String] { + let copy : Array[String] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn alignment_map_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn alignment_map_min(left : Int, right : Int) -> Int { + if left < right { + left + } else { + right + } +} + +///| +fn alignment_map_max(left : Int, right : Int) -> Int { + if left > right { + left + } else { + right + } +} + +///| +fn alignment_map_validate_name( + name : String, + label : String, +) -> Unit raise AlignmentMapError { + if name.length() == 0 { + raise AlignmentMapError(label + " must not be empty") + } +} + +///| +fn alignment_map_validate_sequence( + sequence : String, + expected_length : Int, + label : String, +) -> Unit raise AlignmentMapError { + if expected_length < 0 { + raise AlignmentMapError(label + " length must be non-negative") + } + if sequence.length() > 0 && sequence.length() != expected_length { + raise AlignmentMapError(label + " length does not match its declaration") + } + let mut index = 0 + while index < sequence.length() { + let residue = sequence.unsafe_get(index).unsafe_to_char() + if residue == '-' { + raise AlignmentMapError(label + " must not contain gap characters") + } + if residue == ' ' || residue == '\t' || residue == '\n' || residue == '\r' { + raise AlignmentMapError(label + " must not contain whitespace") + } + index = index + 1 + } +} + +///| +fn alignment_map_validate_coordinates( + target_length : Int, + query_length : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> Unit raise AlignmentMapError { + if target_coordinates.length() != query_coordinates.length() { + raise AlignmentMapError( + "target and query coordinates must have equal length", + ) + } + if target_coordinates.length() == 1 { + raise AlignmentMapError( + "an alignment path must be empty or contain at least two points", + ) + } + let mut index = 0 + while index < target_coordinates.length() { + let target = target_coordinates[index] + let query = query_coordinates[index] + if target < 0 || target > target_length { + raise AlignmentMapError("target coordinate is out of bounds") + } + if query < 0 || query > query_length { + raise AlignmentMapError("query coordinate is out of bounds") + } + index = index + 1 + } + let mut query_direction = 0 + index = 0 + while index + 1 < target_coordinates.length() { + let target_step = target_coordinates[index + 1] - target_coordinates[index] + let query_step = query_coordinates[index + 1] - query_coordinates[index] + if target_step < 0 { + raise AlignmentMapError("target coordinates must be non-decreasing") + } + if target_step == 0 && query_step == 0 { + raise AlignmentMapError("consecutive alignment coordinates must differ") + } + if query_step > 0 { + if query_direction < 0 { + raise AlignmentMapError("query coordinates change direction") + } + query_direction = 1 + } else if query_step < 0 { + if query_direction > 0 { + raise AlignmentMapError("query coordinates change direction") + } + query_direction = -1 + } + index = index + 1 + } +} + +///| +fn alignment_map_create( + target_name : String, + target_sequence : String, + target_length : Int, + query_name : String, + query_sequence : String, + query_length : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> CoordinatePairwiseAlignment raise AlignmentMapError { + alignment_map_validate_name(target_name, "target name") + alignment_map_validate_name(query_name, "query name") + alignment_map_validate_sequence( + target_sequence, target_length, "target sequence", + ) + alignment_map_validate_sequence( + query_sequence, query_length, "query sequence", + ) + alignment_map_validate_coordinates( + target_length, query_length, target_coordinates, query_coordinates, + ) + CoordinatePairwiseAlignment::{ + target_name, + query_name, + target_sequence, + query_sequence, + target_length, + query_length, + target_coordinates: alignment_map_copy_ints(target_coordinates), + query_coordinates: alignment_map_copy_ints(query_coordinates), + } +} + +///| +/// Construct a coordinate pairwise alignment from concrete sequences. +pub fn coordinate_pairwise_alignment( + target_name : String, + target_sequence : String, + query_name : String, + query_sequence : String, + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> CoordinatePairwiseAlignment raise AlignmentMapError { + alignment_map_create( + target_name, + target_sequence, + target_sequence.length(), + query_name, + query_sequence, + query_sequence.length(), + target_coordinates, + query_coordinates, + ) +} + +///| +/// Construct an alignment when sequence text is unavailable. +pub fn coordinate_pairwise_alignment_with_lengths( + target_name : String, + target_length : Int, + query_name : String, + query_length : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], + target_sequence? : String = "", + query_sequence? : String = "", +) -> CoordinatePairwiseAlignment raise AlignmentMapError { + alignment_map_create( + target_name, target_sequence, target_length, query_name, query_sequence, query_length, + target_coordinates, query_coordinates, + ) +} + +///| +/// Adapt the repository's existing PairwiseAlignment representation. +pub fn coordinate_alignment_from_pairwise( + alignment : PairwiseAlignment, + target_name? : String = "target", + query_name? : String = "query", +) -> CoordinatePairwiseAlignment raise AlignmentMapError { + if alignment.aligned_target.length() != alignment.aligned_query.length() { + raise AlignmentMapError("aligned pairwise rows must have equal length") + } + let target_coordinates = [alignment.target_start] + let query_coordinates = [alignment.query_start] + let mut target = alignment.target_start + let mut query = alignment.query_start + let mut column = 0 + while column < alignment.aligned_target.length() { + if alignment.aligned_target.unsafe_get(column).unsafe_to_char() != '-' { + target = target + 1 + } + if alignment.aligned_query.unsafe_get(column).unsafe_to_char() != '-' { + query = query + 1 + } + target_coordinates.push(target) + query_coordinates.push(query) + column = column + 1 + } + if target != alignment.target_end || query != alignment.query_end { + raise AlignmentMapError("pairwise coordinates disagree with aligned rows") + } + coordinate_pairwise_alignment( + target_name, + alignment.target, + query_name, + alignment.query, + target_coordinates, + query_coordinates, + ) +} + +///| +fn alignment_map_reverse(values : Array[Int]) -> Array[Int] { + let result : Array[Int] = [] + let mut index = values.length() + while index > 0 { + index = index - 1 + result.push(values[index]) + } + result +} + +///| +fn alignment_map_relationship( + target_coordinates : Array[Int], + query_coordinates : Array[Int], + label : String, +) -> Int raise AlignmentMapError { + let mut relationship = 0 + let mut index = 0 + while index + 1 < target_coordinates.length() { + let target_step = target_coordinates[index + 1] - target_coordinates[index] + let query_step = query_coordinates[index + 1] - query_coordinates[index] + if target_step != 0 && query_step != 0 { + let current = if target_step * query_step > 0 { 1 } else { -1 } + if relationship != 0 && relationship != current { + raise AlignmentMapError("inconsistent steps in " + label) + } + relationship = current + } + index = index + 1 + } + if relationship == 0 { + 1 + } else { + relationship + } +} + +///| +fn alignment_map_validate_equal_steps( + target_coordinates : Array[Int], + query_coordinates : Array[Int], + label : String, +) -> Unit raise AlignmentMapError { + let mut index = 0 + while index + 1 < target_coordinates.length() { + let target_step = target_coordinates[index + 1] - target_coordinates[index] + let query_step = query_coordinates[index + 1] - query_coordinates[index] + if target_step < 0 || query_step < 0 { + raise AlignmentMapError("normalized coordinates decrease in " + label) + } + if target_step > 0 && query_step > 0 && target_step != query_step { + raise AlignmentMapError("unequal aligned step sizes in " + label) + } + index = index + 1 + } +} + +///| +fn alignment_map_transform_against_reverse_middle( + alignment : CoordinatePairwiseAlignment, +) -> (Array[Int], Array[Int]) { + let target_reversed = alignment_map_reverse(alignment.target_coordinates) + let query_reversed = alignment_map_reverse(alignment.query_coordinates) + let targets : Array[Int] = [] + let queries : Array[Int] = [] + for value in target_reversed { + targets.push(alignment.target_length - value) + } + for value in query_reversed { + queries.push(alignment.query_length - value) + } + (targets, queries) +} + +///| +/// Compose `self` (outer target -> shared middle) with `next` +/// (shared middle -> final query). +pub fn CoordinatePairwiseAlignment::map( + self : CoordinatePairwiseAlignment, + next : CoordinatePairwiseAlignment, +) -> CoordinatePairwiseAlignment raise AlignmentMapError { + if self.query_length != next.target_length { + raise AlignmentMapError( + "first query length must equal second target length", + ) + } + if self.target_coordinates.length() == 0 || + next.target_coordinates.length() == 0 { + return alignment_map_create( + self.target_name, + self.target_sequence, + self.target_length, + next.query_name, + next.query_sequence, + next.query_length, + [], + [], + ) + } + let relationship1 = alignment_map_relationship( + self.target_coordinates, + self.query_coordinates, + "first alignment", + ) + let relationship2 = alignment_map_relationship( + next.target_coordinates, + next.query_coordinates, + "second alignment", + ) + let coordinates1_target = alignment_map_copy_ints(self.target_coordinates) + let coordinates1_query = alignment_map_copy_ints(self.query_coordinates) + let mut coordinates2_target = alignment_map_copy_ints(next.target_coordinates) + let mut coordinates2_query = alignment_map_copy_ints(next.query_coordinates) + if relationship1 > 0 { + if relationship2 < 0 { + let transformed : Array[Int] = [] + for value in coordinates2_query { + transformed.push(next.query_length - value) + } + coordinates2_query = transformed + } + } else { + let transformed_middle : Array[Int] = [] + for value in coordinates1_query { + transformed_middle.push(self.query_length - value) + } + coordinates1_query.clear() + for value in transformed_middle { + coordinates1_query.push(value) + } + let (reversed_target, reversed_query) = if relationship2 > 0 { + alignment_map_transform_against_reverse_middle(next) + } else { + let reversed_middle = alignment_map_reverse(next.target_coordinates) + let normalized_middle : Array[Int] = [] + for value in reversed_middle { + normalized_middle.push(next.target_length - value) + } + (normalized_middle, alignment_map_reverse(next.query_coordinates)) + } + coordinates2_target = reversed_target + coordinates2_query = reversed_query + } + alignment_map_validate_equal_steps( + coordinates1_target, coordinates1_query, "first alignment", + ) + alignment_map_validate_equal_steps( + coordinates2_target, coordinates2_query, "second alignment", + ) + let mut first_segment = 0 + while first_segment + 1 < coordinates1_target.length() { + if coordinates1_target[first_segment] < + coordinates1_target[first_segment + 1] && + coordinates1_query[first_segment] < coordinates1_query[first_segment + 1] { + break + } + first_segment = first_segment + 1 + } + if first_segment + 1 >= coordinates1_target.length() { + return alignment_map_create( + self.target_name, + self.target_sequence, + self.target_length, + next.query_name, + next.query_sequence, + next.query_length, + [], + [], + ) + } + let mut outer_start = coordinates1_target[first_segment] + let mut middle_start = coordinates1_query[first_segment] + let mut outer_end = coordinates1_target[first_segment + 1] + let mut middle_end = coordinates1_query[first_segment + 1] + let path_target : Array[Int] = [] + let path_query : Array[Int] = [] + let mut output_target_end = 2147483647 + let mut output_query_end = 2147483647 + let mut second_target_start = 2147483647 + let mut second_query_start = 2147483647 + let mut point2 = 0 + while point2 < coordinates2_target.length() { + let second_target_end = coordinates2_target[point2] + let second_query_end = coordinates2_query[point2] + while second_query_start < second_query_end && + second_target_start < second_target_end { + let mut size = 0 + let mut handled = false + while !handled { + if second_target_start < middle_start { + size = alignment_map_min(second_target_end, middle_start) - + second_target_start + handled = true + } else if second_target_start < middle_end { + let offset = second_target_start - middle_start + size = alignment_map_min(second_target_end, middle_end) - + second_target_start + let query_start = second_query_start + let target_start = outer_start + offset + if target_start != output_target_end || + query_start != output_query_end { + if target_start > output_target_end && + query_start > output_query_end { + path_target.push(target_start) + path_query.push(output_query_end) + } + path_target.push(target_start) + path_query.push(query_start) + } + output_query_end = query_start + size + output_target_end = target_start + size + path_target.push(output_target_end) + path_query.push(output_query_end) + handled = true + } else { + first_segment = first_segment + 1 + let mut found = false + while first_segment + 1 < coordinates1_target.length() { + outer_start = coordinates1_target[first_segment] + middle_start = coordinates1_query[first_segment] + outer_end = coordinates1_target[first_segment + 1] + middle_end = coordinates1_query[first_segment + 1] + if outer_start < outer_end && middle_start < middle_end { + found = true + break + } + first_segment = first_segment + 1 + } + if !found { + size = second_query_end - second_query_start + handled = true + } + } + } + if size <= 0 { + break + } + second_query_start = second_query_start + size + second_target_start = second_target_start + size + } + second_target_start = second_target_end + second_query_start = second_query_end + point2 = point2 + 1 + } + if relationship1 != relationship2 { + let transformed : Array[Int] = [] + for value in path_query { + transformed.push(next.query_length - value) + } + path_query.clear() + for value in transformed { + path_query.push(value) + } + } + alignment_map_create( + self.target_name, + self.target_sequence, + self.target_length, + next.query_name, + next.query_sequence, + next.query_length, + path_target, + path_query, + ) +} + +///| +/// Map several alignments through the same outer alignment. +pub fn CoordinatePairwiseAlignment::map_many( + self : CoordinatePairwiseAlignment, + alignments : Array[CoordinatePairwiseAlignment], +) -> Array[CoordinatePairwiseAlignment] raise AlignmentMapError { + let results : Array[CoordinatePairwiseAlignment] = [] + for alignment in alignments { + results.push(self.map(alignment)) + } + results +} + +///| +pub fn CoordinatePairwiseAlignment::is_empty( + self : CoordinatePairwiseAlignment, +) -> Bool { + self.target_coordinates.length() == 0 +} + +///| +pub fn CoordinatePairwiseAlignment::is_reverse( + self : CoordinatePairwiseAlignment, +) -> Bool { + if self.query_coordinates.length() < 2 { + return false + } + self.query_coordinates[self.query_coordinates.length() - 1] < + self.query_coordinates[0] +} + +///| +pub fn CoordinatePairwiseAlignment::blocks( + self : CoordinatePairwiseAlignment, +) -> Array[CoordinateAlignmentBlock] { + let blocks : Array[CoordinateAlignmentBlock] = [] + let mut index = 0 + while index + 1 < self.target_coordinates.length() { + let target_start = self.target_coordinates[index] + let target_end = self.target_coordinates[index + 1] + let query_start = self.query_coordinates[index] + let query_end = self.query_coordinates[index + 1] + let target_size = target_end - target_start + let query_size = alignment_map_abs(query_end - query_start) + if target_size > 0 && query_size > 0 { + blocks.push(CoordinateAlignmentBlock::{ + target_start, + target_end, + query_start, + query_end, + size: alignment_map_min(target_size, query_size), + }) + } + index = index + 1 + } + blocks +} + +///| +pub fn CoordinatePairwiseAlignment::counts( + self : CoordinatePairwiseAlignment, +) -> CoordinateAlignmentCounts { + let mut aligned = 0 + let mut target_gap_bases = 0 + let mut query_gap_bases = 0 + let mut target_gap_events = 0 + let mut query_gap_events = 0 + let mut blocks = 0 + let mut index = 0 + while index + 1 < self.target_coordinates.length() { + let target_step = self.target_coordinates[index + 1] - + self.target_coordinates[index] + let query_step = alignment_map_abs( + self.query_coordinates[index + 1] - self.query_coordinates[index], + ) + if target_step > 0 && query_step > 0 { + aligned = aligned + alignment_map_min(target_step, query_step) + blocks = blocks + 1 + } else if target_step == 0 && query_step > 0 { + target_gap_bases = target_gap_bases + query_step + target_gap_events = target_gap_events + 1 + } else if target_step > 0 && query_step == 0 { + query_gap_bases = query_gap_bases + target_step + query_gap_events = query_gap_events + 1 + } + index = index + 1 + } + CoordinateAlignmentCounts::{ + aligned, + target_gap_bases, + query_gap_bases, + target_gap_events, + query_gap_events, + blocks, + } +} + +///| +pub fn CoordinatePairwiseAlignment::target_to_query( + self : CoordinatePairwiseAlignment, + position : Int, +) -> Int? { + let mut index = 0 + while index + 1 < self.target_coordinates.length() { + let target_start = self.target_coordinates[index] + let target_end = self.target_coordinates[index + 1] + let query_start = self.query_coordinates[index] + let query_end = self.query_coordinates[index + 1] + if target_start <= position && + position < target_end && + query_start != query_end { + let offset = position - target_start + if query_end > query_start { + return Some(query_start + offset) + } + return Some(query_start - offset - 1) + } + index = index + 1 + } + None +} + +///| +pub fn CoordinatePairwiseAlignment::query_to_target( + self : CoordinatePairwiseAlignment, + position : Int, +) -> Int? { + let mut index = 0 + while index + 1 < self.query_coordinates.length() { + let target_start = self.target_coordinates[index] + let target_end = self.target_coordinates[index + 1] + let query_start = self.query_coordinates[index] + let query_end = self.query_coordinates[index + 1] + if target_start != target_end { + if query_end > query_start && + query_start <= position && + position < query_end { + return Some(target_start + position - query_start) + } + if query_end < query_start && + query_end <= position && + position < query_start { + return Some(target_start + query_start - position - 1) + } + } + index = index + 1 + } + None +} + +///| +fn alignment_map_repeat( + builder : StringBuilder, + value : Char, + count : Int, +) -> Unit { + let mut index = 0 + while index < count { + builder.write_char(value) + index = index + 1 + } +} + +///| +fn alignment_map_complement(value : Char) -> Char { + match value { + 'A' => 'T' + 'a' => 't' + 'C' => 'G' + 'c' => 'g' + 'G' => 'C' + 'g' => 'c' + 'T' => 'A' + 't' => 'a' + 'U' => 'A' + 'u' => 'a' + _ => value + } +} + +///| +fn alignment_map_write_segment( + builder : StringBuilder, + sequence : String, + start : Int, + end : Int, + complement_reverse : Bool, +) -> Unit { + if end >= start { + let mut index = start + while index < end { + builder.write_char(sequence.unsafe_get(index).unsafe_to_char()) + index = index + 1 + } + } else { + let mut index = start + while index > end { + index = index - 1 + let residue = sequence.unsafe_get(index).unsafe_to_char() + builder.write_char( + if complement_reverse { + alignment_map_complement(residue) + } else { + residue + }, + ) + } + } +} + +///| +/// Render gapped rows when both sequence contents are available. +pub fn CoordinatePairwiseAlignment::aligned_rows( + self : CoordinatePairwiseAlignment, + complement_reverse? : Bool = true, +) -> (String, String)? { + if self.target_sequence.length() == 0 || self.query_sequence.length() == 0 { + return None + } + let target = StringBuilder::new() + let query = StringBuilder::new() + let mut index = 0 + while index + 1 < self.target_coordinates.length() { + let target_start = self.target_coordinates[index] + let target_end = self.target_coordinates[index + 1] + let query_start = self.query_coordinates[index] + let query_end = self.query_coordinates[index + 1] + let target_step = target_end - target_start + let query_step = alignment_map_abs(query_end - query_start) + if target_step > 0 && query_step > 0 { + alignment_map_write_segment( + target, + self.target_sequence, + target_start, + target_end, + false, + ) + alignment_map_write_segment( + query, + self.query_sequence, + query_start, + query_end, + complement_reverse, + ) + } else if target_step > 0 { + alignment_map_write_segment( + target, + self.target_sequence, + target_start, + target_end, + false, + ) + alignment_map_repeat(query, '-', target_step) + } else if query_step > 0 { + alignment_map_repeat(target, '-', query_step) + alignment_map_write_segment( + query, + self.query_sequence, + query_start, + query_end, + complement_reverse, + ) + } + index = index + 1 + } + Some((target.to_string(), query.to_string())) +} + +///| +fn alignment_map_join_ints(values : Array[Int]) -> String { + let builder = StringBuilder::new() + let mut index = 0 + while index < values.length() { + if index > 0 { + builder.write_char(',') + } + builder.write_string(values[index].to_string()) + index = index + 1 + } + if values.length() > 0 { + builder.write_char(',') + } + builder.to_string() +} + +///| +/// Serialize the coordinate blocks as a PSL record. +/// +/// Aligned positions are reported as matches because composition does not +/// require sequence contents; gap event/base counts remain exact. +pub fn CoordinatePairwiseAlignment::to_psl( + self : CoordinatePairwiseAlignment, +) -> String raise AlignmentMapError { + let blocks = self.blocks() + if blocks.length() == 0 { + raise AlignmentMapError("cannot serialize an empty alignment as PSL") + } + let counts = self.counts() + let block_sizes : Array[Int] = [] + let query_starts : Array[Int] = [] + let target_starts : Array[Int] = [] + let reverse = self.is_reverse() + let mut query_min = self.query_length + let mut query_max = 0 + let mut target_min = self.target_length + let mut target_max = 0 + for block in blocks { + block_sizes.push(block.size) + target_starts.push(block.target_start) + let block_query_min = alignment_map_min(block.query_start, block.query_end) + let block_query_max = alignment_map_max(block.query_start, block.query_end) + query_min = alignment_map_min(query_min, block_query_min) + query_max = alignment_map_max(query_max, block_query_max) + target_min = alignment_map_min(target_min, block.target_start) + target_max = alignment_map_max(target_max, block.target_end) + query_starts.push( + if reverse { + self.query_length - block_query_max + } else { + block_query_min + }, + ) + } + let strand = if reverse { "-" } else { "+" } + counts.aligned.to_string() + + "\t0\t0\t0\t" + + counts.target_gap_events.to_string() + + "\t" + + counts.target_gap_bases.to_string() + + "\t" + + counts.query_gap_events.to_string() + + "\t" + + counts.query_gap_bases.to_string() + + "\t" + + strand + + "\t" + + self.query_name + + "\t" + + self.query_length.to_string() + + "\t" + + query_min.to_string() + + "\t" + + query_max.to_string() + + "\t" + + self.target_name + + "\t" + + self.target_length.to_string() + + "\t" + + target_min.to_string() + + "\t" + + target_max.to_string() + + "\t" + + blocks.length().to_string() + + "\t" + + alignment_map_join_ints(block_sizes) + + "\t" + + alignment_map_join_ints(query_starts) + + "\t" + + alignment_map_join_ints(target_starts) + + "\n" +} + +///| +pub fn CoordinatePairwiseAlignment::summary( + self : CoordinatePairwiseAlignment, +) -> String { + let counts = self.counts() + "CoordinatePairwiseAlignment(" + + self.target_name + + " <- " + + self.query_name + + ", blocks=" + + counts.blocks.to_string() + + ", aligned=" + + counts.aligned.to_string() + + ", strand=" + + (if self.is_reverse() { "-" } else { "+" }) + + ")" +} + +///| +fn alignment_map_ungap(sequence : String) -> String { + let builder = StringBuilder::new() + let mut index = 0 + while index < sequence.length() { + let residue = sequence.unsafe_get(index).unsafe_to_char() + if residue != '-' { + builder.write_char(residue) + } + index = index + 1 + } + builder.to_string() +} + +///| +/// Build a coordinate MSA from equal-width gapped rows. +pub fn coordinate_multiple_alignment( + names : Array[String], + aligned_sequences : Array[String], +) -> CoordinateMultipleAlignment raise AlignmentMapError { + if names.length() == 0 { + raise AlignmentMapError( + "a multiple alignment must contain at least one row", + ) + } + if names.length() != aligned_sequences.length() { + raise AlignmentMapError("MSA names and rows must have equal length") + } + let width = aligned_sequences[0].length() + let sequences : Array[String] = [] + let coordinates : Array[Array[Int]] = [] + let mut row_index = 0 + while row_index < names.length() { + alignment_map_validate_name(names[row_index], "MSA row name") + if aligned_sequences[row_index].length() != width { + raise AlignmentMapError("all MSA rows must have equal width") + } + let raw = alignment_map_ungap(aligned_sequences[row_index]) + alignment_map_validate_sequence(raw, raw.length(), "MSA sequence") + sequences.push(raw) + let row = [0] + let mut coordinate = 0 + let mut column = 0 + while column < width { + if aligned_sequences[row_index].unsafe_get(column).unsafe_to_char() != '-' { + coordinate = coordinate + 1 + } + row.push(coordinate) + column = column + 1 + } + coordinates.push(row) + row_index = row_index + 1 + } + CoordinateMultipleAlignment::{ + names: alignment_map_copy_strings(names), + sequences, + coordinates, + } +} + +///| +fn alignment_map_factor( + alignment : CoordinatePairwiseAlignment, +) -> Int raise AlignmentMapError { + let mut factor = 0 + let mut index = 0 + while index + 1 < alignment.target_coordinates.length() { + let target_step = alignment_map_abs( + alignment.target_coordinates[index + 1] - + alignment.target_coordinates[index], + ) + let query_step = alignment_map_abs( + alignment.query_coordinates[index + 1] - + alignment.query_coordinates[index], + ) + if target_step > 0 && query_step > 0 { + let current = if query_step == target_step { + 1 + } else if query_step == 3 * target_step { + 3 + } else { + raise AlignmentMapError( + "mapall supports only 1:1 or protein-to-codon mappings", + ) + } + if factor != 0 && factor != current { + raise AlignmentMapError("mapping contains inconsistent step factors") + } + factor = current + } + index = index + 1 + } + if factor == 0 { + raise AlignmentMapError("mapping has no aligned blocks") + } + factor +} + +///| +fn alignment_map_validate_msa( + alignment : CoordinateMultipleAlignment, +) -> Unit raise AlignmentMapError { + if alignment.names.length() == 0 || + alignment.names.length() != alignment.sequences.length() || + alignment.names.length() != alignment.coordinates.length() { + raise AlignmentMapError("invalid multiple alignment dimensions") + } + let points = alignment.coordinates[0].length() + if points < 2 { + raise AlignmentMapError( + "multiple alignment must contain at least one column", + ) + } + let mut row_index = 0 + while row_index < alignment.names.length() { + alignment_map_validate_name(alignment.names[row_index], "MSA row name") + if alignment.coordinates[row_index].length() != points { + raise AlignmentMapError("MSA coordinate rows must have equal length") + } + let mut point = 0 + while point < points { + let coordinate = alignment.coordinates[row_index][point] + if coordinate < 0 || coordinate > alignment.sequences[row_index].length() { + raise AlignmentMapError("MSA coordinate is out of bounds") + } + if point > 0 && coordinate < alignment.coordinates[row_index][point - 1] { + raise AlignmentMapError("MSA coordinates must be non-decreasing") + } + point = point + 1 + } + row_index = row_index + 1 + } +} + +///| +/// Project every row of an MSA through its corresponding sequence mapping. +/// +/// A 1:1 mapping produces a mapped sequence MSA. A 1:3 protein-to-nucleotide +/// mapping produces a codon-aware nucleotide MSA without translating or +/// realigning the nucleotide sequences. +pub fn CoordinateMultipleAlignment::mapall( + self : CoordinateMultipleAlignment, + mappings : Array[CoordinatePairwiseAlignment], +) -> CoordinateMultipleAlignment raise AlignmentMapError { + alignment_map_validate_msa(self) + if mappings.length() != self.sequences.length() { + raise AlignmentMapError("mapall requires one mapping per MSA row") + } + let mut factor = 0 + let mut row_index = 0 + while row_index < mappings.length() { + if mappings[row_index].target_length != self.sequences[row_index].length() { + raise AlignmentMapError( + "mapping target length does not match its MSA row", + ) + } + let current_factor = alignment_map_factor(mappings[row_index]) + if factor == 0 { + factor = current_factor + } else if factor != current_factor { + raise AlignmentMapError("mapall mappings use inconsistent step factors") + } + row_index = row_index + 1 + } + let points = self.coordinates[0].length() + let master_coordinates = [0] + let mut master = 0 + let mut point = 0 + while point + 1 < points { + let mut step = 0 + row_index = 0 + while row_index < self.coordinates.length() { + let current = alignment_map_abs( + self.coordinates[row_index][point + 1] - + self.coordinates[row_index][point], + ) + step = alignment_map_max(step, current) + row_index = row_index + 1 + } + master = master + factor * step + master_coordinates.push(master) + point = point + 1 + } + let projected : Array[CoordinatePairwiseAlignment] = [] + row_index = 0 + while row_index < mappings.length() { + let scaled_row : Array[Int] = [] + for coordinate in self.coordinates[row_index] { + scaled_row.push(factor * coordinate) + } + let outer = alignment_map_create( + "alignment", + "", + master, + self.names[row_index], + "", + factor * self.sequences[row_index].length(), + master_coordinates, + scaled_row, + ) + let scaled_target : Array[Int] = [] + for coordinate in mappings[row_index].target_coordinates { + scaled_target.push(factor * coordinate) + } + let inner = alignment_map_create( + mappings[row_index].target_name, + "", + factor * mappings[row_index].target_length, + mappings[row_index].query_name, + mappings[row_index].query_sequence, + mappings[row_index].query_length, + scaled_target, + mappings[row_index].query_coordinates, + ) + projected.push(outer.map(inner)) + row_index = row_index + 1 + } + let output_coordinates : Array[Array[Int]] = [] + let indices = Array::make(projected.length(), 0) + row_index = 0 + while row_index < projected.length() { + output_coordinates.push([]) + row_index = row_index + 1 + } + let mut previous = 0 + while true { + let mut found = false + let mut position = 2147483647 + row_index = 0 + while row_index < projected.length() { + let index = indices[row_index] + if index < projected[row_index].target_coordinates.length() { + position = alignment_map_min( + position, + projected[row_index].target_coordinates[index], + ) + found = true + } + row_index = row_index + 1 + } + if !found { + break + } + row_index = 0 + while row_index < projected.length() { + let index = indices[row_index] + let output = output_coordinates[row_index] + if index >= projected[row_index].target_coordinates.length() { + output.push( + if output.length() == 0 { + 0 + } else { + output[output.length() - 1] + }, + ) + } else { + let current_target = projected[row_index].target_coordinates[index] + let current_query = projected[row_index].query_coordinates[index] + if current_target == position { + output.push(current_query) + indices[row_index] = index + 1 + } else if current_target > position { + if output.length() == 0 { + output.push(current_query) + } else { + let last = output[output.length() - 1] + let step = if current_query > last { + position - previous + } else { + 0 + } + output.push(last + step) + } + } else { + raise AlignmentMapError("mapall coordinate merge became inconsistent") + } + } + row_index = row_index + 1 + } + previous = position + } + let names : Array[String] = [] + let sequences : Array[String] = [] + for alignment in projected { + names.push(alignment.query_name) + sequences.push(alignment.query_sequence) + } + CoordinateMultipleAlignment::{ + names, + sequences, + coordinates: output_coordinates, + } +} + +///| +pub fn CoordinateMultipleAlignment::num_sequences( + self : CoordinateMultipleAlignment, +) -> Int { + self.sequences.length() +} + +///| +pub fn CoordinateMultipleAlignment::alignment_length( + self : CoordinateMultipleAlignment, +) -> Int { + if self.coordinates.length() == 0 || self.coordinates[0].length() < 2 { + return 0 + } + let mut length = 0 + let mut point = 0 + while point + 1 < self.coordinates[0].length() { + let mut step = 0 + for row in self.coordinates { + step = alignment_map_max( + step, + alignment_map_abs(row[point + 1] - row[point]), + ) + } + length = length + step + point = point + 1 + } + length +} + +///| +pub fn CoordinateMultipleAlignment::row( + self : CoordinateMultipleAlignment, + index : Int, + complement_reverse? : Bool = true, +) -> String? { + if index < 0 || + index >= self.sequences.length() || + self.coordinates.length() == 0 { + return None + } + let sequence = self.sequences[index] + if sequence.length() == 0 { + return None + } + let builder = StringBuilder::new() + let mut point = 0 + while point + 1 < self.coordinates[index].length() { + let start = self.coordinates[index][point] + let end = self.coordinates[index][point + 1] + let row_step = alignment_map_abs(end - start) + let mut width = 0 + for row_coordinates in self.coordinates { + width = alignment_map_max( + width, + alignment_map_abs(row_coordinates[point + 1] - row_coordinates[point]), + ) + } + if row_step == 0 { + alignment_map_repeat(builder, '-', width) + } else { + alignment_map_write_segment( + builder, sequence, start, end, complement_reverse, + ) + alignment_map_repeat(builder, '-', width - row_step) + } + point = point + 1 + } + Some(builder.to_string()) +} + +///| +pub fn CoordinateMultipleAlignment::to_fasta( + self : CoordinateMultipleAlignment, +) -> String { + let builder = StringBuilder::new() + let mut index = 0 + while index < self.names.length() { + builder.write_char('>') + builder.write_string(self.names[index]) + builder.write_char('\n') + match self.row(index) { + Some(row) => builder.write_string(row) + None => () + } + builder.write_char('\n') + index = index + 1 + } + builder.to_string() +} + +///| +pub fn CoordinateMultipleAlignment::summary( + self : CoordinateMultipleAlignment, +) -> String { + "CoordinateMultipleAlignment(rows=" + + self.num_sequences().to_string() + + ", columns=" + + self.alignment_length().to_string() + + ")" +} + +///| +/// Small chromosome/transcript/read example from the Biopython tutorial. +pub fn alignment_map_example() -> CoordinatePairwiseAlignment raise AlignmentMapError { + let chromosome = coordinate_pairwise_alignment( + "chromosome", + "AAAAAAAACCCCCCCAAAAAAAAAAAGGGGGGAAAAAAAA", + "transcript", + "CCCCCCCGGGGGG", + [8, 15, 26, 32], + [0, 7, 7, 13], + ) + let read = coordinate_pairwise_alignment( + "transcript", + "CCCCCCCGGGGGG", + "read", + "CCCCGGGG", + [3, 11], + [0, 8], + ) + chromosome.map(read) +} diff --git a/src/alphabet.mbt b/src/alphabet.mbt index a9fde8f0..a6b26035 100644 --- a/src/alphabet.mbt +++ b/src/alphabet.mbt @@ -13,31 +13,35 @@ pub struct Alphabet { ///| /// Create a new alphabet. -pub fn Alphabet::new(name : String, letters : Array[String], is_gapped : Bool) -> Alphabet { +pub fn Alphabet::new( + name : String, + letters : Array[String], + is_gapped : Bool, +) -> Alphabet { Alphabet::{ name, letters, is_gapped } } ///| /// Get alphabet name. -pub fn get_name(self : Alphabet) -> String { +pub fn Alphabet::get_name(self : Alphabet) -> String { self.name } ///| /// Get alphabet letters. -pub fn get_letters(self : Alphabet) -> Array[String] { +pub fn Alphabet::get_letters(self : Alphabet) -> Array[String] { self.letters } ///| /// Check if alphabet is gapped. -pub fn is_gapped(self : Alphabet) -> Bool { +pub fn Alphabet::is_gapped(self : Alphabet) -> Bool { self.is_gapped } ///| /// Check if a character is valid in this alphabet. -pub fn is_valid(self : Alphabet, c : String) -> Bool { +pub fn Alphabet::is_valid(self : Alphabet, c : String) -> Bool { let mut i = 0 while i < self.letters.length() { if self.letters[i] == c { @@ -51,18 +55,14 @@ pub fn is_valid(self : Alphabet, c : String) -> Bool { ///| /// IUPAC unambiguous DNA alphabet (A, C, G, T). pub fn iupac_unambiguous_dna() -> Alphabet { - let letters = [ - "A", "C", "G", "T" - ] + let letters = ["A", "C", "G", "T"] Alphabet::new("IUPACUnambiguousDNA", letters, false) } ///| /// IUPAC unambiguous RNA alphabet (A, C, G, U). pub fn iupac_unambiguous_rna() -> Alphabet { - let letters = [ - "A", "C", "G", "U" - ] + let letters = ["A", "C", "G", "U"] Alphabet::new("IUPACUnambiguousRNA", letters, false) } @@ -70,9 +70,7 @@ pub fn iupac_unambiguous_rna() -> Alphabet { /// IUPAC ambiguous DNA alphabet (A, C, G, T, R, Y, S, W, K, M, B, D, H, V, N). pub fn iupac_ambiguous_dna() -> Alphabet { let letters = [ - "A", "C", "G", "T", - "R", "Y", "S", "W", "K", "M", - "B", "D", "H", "V", "N" + "A", "C", "G", "T", "R", "Y", "S", "W", "K", "M", "B", "D", "H", "V", "N", ] Alphabet::new("IUPACAmbiguousDNA", letters, false) } @@ -81,9 +79,7 @@ pub fn iupac_ambiguous_dna() -> Alphabet { /// IUPAC ambiguous RNA alphabet (A, C, G, U, R, Y, S, W, K, M, B, D, H, V, N). pub fn iupac_ambiguous_rna() -> Alphabet { let letters = [ - "A", "C", "G", "U", - "R", "Y", "S", "W", "K", "M", - "B", "D", "H", "V", "N" + "A", "C", "G", "U", "R", "Y", "S", "W", "K", "M", "B", "D", "H", "V", "N", ] Alphabet::new("IUPACAmbiguousRNA", letters, false) } @@ -92,9 +88,8 @@ pub fn iupac_ambiguous_rna() -> Alphabet { /// IUPAC protein alphabet (20 standard amino acids + B, Z, X). pub fn iupac_protein() -> Alphabet { let letters = [ - "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", - "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y", - "B", "Z", "X" + "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", + "T", "V", "W", "Y", "B", "Z", "X", ] Alphabet::new("IUPACProtein", letters, false) } @@ -103,9 +98,8 @@ pub fn iupac_protein() -> Alphabet { /// Extended IUPAC protein alphabet (20 standard + B, Z, X, U, O). pub fn iupac_extended_protein() -> Alphabet { let letters = [ - "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", - "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y", - "B", "Z", "X", "U", "O" + "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", + "T", "V", "W", "Y", "B", "Z", "X", "U", "O", ] Alphabet::new("IUPACExtendedProtein", letters, false) } @@ -113,18 +107,14 @@ pub fn iupac_extended_protein() -> Alphabet { ///| /// Gapped DNA alphabet (includes gap character '-'). pub fn gapped_dna() -> Alphabet { - let letters = [ - "A", "C", "G", "T", "-" - ] + let letters = ["A", "C", "G", "T", "-"] Alphabet::new("GappedDNA", letters, true) } ///| /// Gapped RNA alphabet (includes gap character '-'). pub fn gapped_rna() -> Alphabet { - let letters = [ - "A", "C", "G", "U", "-" - ] + let letters = ["A", "C", "G", "U", "-"] Alphabet::new("GappedRNA", letters, true) } @@ -132,9 +122,8 @@ pub fn gapped_rna() -> Alphabet { /// Gapped protein alphabet (includes gap character '-'). pub fn gapped_protein() -> Alphabet { let letters = [ - "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", - "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y", - "-" + "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", + "T", "V", "W", "Y", "-", ] Alphabet::new("GappedProtein", letters, true) } @@ -144,8 +133,8 @@ pub fn gapped_protein() -> Alphabet { /// Grouped by chemical properties. pub fn reduced_protein() -> Alphabet { let letters = [ - "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", - "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y" + "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", + "T", "V", "W", "Y", ] Alphabet::new("ReducedProtein", letters, false) } diff --git a/src/ancombc.mbt b/src/ancombc.mbt index be2433a2..77a57deb 100644 --- a/src/ancombc.mbt +++ b/src/ancombc.mbt @@ -62,7 +62,7 @@ fn abc_variance(arr : Array[Double]) -> Double { /// Welch-Satterthwaite degrees of freedom. fn abc_welch_t( group1 : Array[Double], - group2 : Array[Double] + group2 : Array[Double], ) -> (Double, Double) { let n1 = group1.length() let n2 = group2.length() @@ -79,9 +79,13 @@ fn abc_welch_t( } let t = (m1 - m2) / se // Welch-Satterthwaite df. - let num = (v1 / n1.to_double() + v2 / n2.to_double()) - let denom1 = v1 * v1 / ((n1.to_double() - 1.0) * n1.to_double() * n1.to_double()) - let denom2 = v2 * v2 / ((n2.to_double() - 1.0) * n2.to_double() * n2.to_double()) + let num = v1 / n1.to_double() + v2 / n2.to_double() + let denom1 = v1 * + v1 / + ((n1.to_double() - 1.0) * n1.to_double() * n1.to_double()) + let denom2 = v2 * + v2 / + ((n2.to_double() - 1.0) * n2.to_double() * n2.to_double()) let df = if denom1 + denom2 > 0.0 { num * num / (denom1 + denom2) } else { @@ -120,11 +124,12 @@ fn abc_normal_cdf(z : Double) -> Double { let t = 1.0 / (1.0 + 0.3275911 * x) let exp_val = @math.exp(0.0 - x * x / 2.0) let y2 = 1.0 - - (((((1.061405429 * t - 1.453152027) * t + 1.421413741) * t - - 0.284496736) * - t + - 0.254829592) * - t) * + ( + (((1.061405429 * t - 1.453152027) * t + 1.421413741) * t - 0.284496736) * + t + + 0.254829592 + ) * + t * exp_val 0.5 * (1.0 + sign * y2) } @@ -143,7 +148,13 @@ fn abc_bh_fdr(pvalues : Array[Double]) -> Array[Double] { indexed.push((i, pvalues[i])) } indexed.sort_by(fn(a : (Int, Double), b : (Int, Double)) -> Int { - if a.1 < b.1 { -1 } else if a.1 > b.1 { 1 } else { 0 } + if a.1 < b.1 { + -1 + } else if a.1 > b.1 { + 1 + } else { + 0 + } }) // Compute adjusted p-values from largest to smallest. let adjusted : Array[Double] = Array::make(m, 0.0) @@ -174,7 +185,10 @@ pub struct AncombcFeature { ///| /// Construct an AncombcFeature. -pub fn AncombcFeature::new(name : String, counts : Array[Int]) -> AncombcFeature { +pub fn AncombcFeature::new( + name : String, + counts : Array[Int], +) -> AncombcFeature { AncombcFeature::{ name, counts } } @@ -237,7 +251,7 @@ pub fn AncombcData::new( sample_names : Array[String], sample_groups : Array[String], feature_names : Array[String], - counts : Array[Array[Int]] + counts : Array[Array[Int]], ) -> AncombcData { AncombcData::{ sample_names, sample_groups, feature_names, counts } } @@ -277,7 +291,7 @@ pub fn AncombcData::counts(self : AncombcData) -> Array[Array[Int]] { pub fn AncombcData::get_count( self : AncombcData, feature_idx : Int, - sample_idx : Int + sample_idx : Int, ) -> Int { self.counts[feature_idx][sample_idx] } @@ -344,9 +358,7 @@ pub fn ancombc_sampling_fractions(data : AncombcData) -> Array[Double] { /// y[i,j] = log(count[i,j] + 0.5) - log(s_j * N + 0.5) /// where s_j is the sampling fraction and N is the number of features. /// This adjusts for differences in sequencing depth (library size). -pub fn ancombc_bias_corrected_log( - data : AncombcData -) -> Array[Array[Double]] { +pub fn ancombc_bias_corrected_log(data : AncombcData) -> Array[Array[Double]] { let n_features = data.n_features() let n_samples = data.n_samples() let fracs = ancombc_sampling_fractions(data) @@ -412,9 +424,16 @@ pub fn AncombcResult::new( w_stat : Int, p_value : Double, q_value : Double, - significant : Bool + significant : Bool, ) -> AncombcResult { - AncombcResult::{ feature, log_fold_change, w_stat, p_value, q_value, significant } + AncombcResult::{ + feature, + log_fold_change, + w_stat, + p_value, + q_value, + significant, + } } ///| @@ -461,14 +480,14 @@ pub struct AncombcTestResult { } ///| -pub fn AncombcTestResult::results(self : AncombcTestResult) -> Array[AncombcResult] { +pub fn AncombcTestResult::results( + self : AncombcTestResult, +) -> Array[AncombcResult] { self.results } ///| -pub fn AncombcTestResult::reference_feature( - self : AncombcTestResult -) -> String { +pub fn AncombcTestResult::reference_feature(self : AncombcTestResult) -> String { self.reference_feature } @@ -517,7 +536,7 @@ pub fn ancombc_test( data : AncombcData, group1 : String, group2 : String, - alpha : Double + alpha : Double, ) -> AncombcTestResult { let n_features = data.n_features() let n_samples = data.n_samples() @@ -535,11 +554,7 @@ pub fn ancombc_test( let log_data = ancombc_bias_corrected_log(data) // Select reference feature. let ref_idx = ancombc_reference_feature(data) - let ref_name = if ref_idx >= 0 { - data.feature_names[ref_idx] - } else { - "" - } + let ref_name = if ref_idx >= 0 { data.feature_names[ref_idx] } else { "" } // Compute log-ratios relative to reference: y_i - y_ref for each sample. let log_ratios : Array[Array[Double]] = Array::new() for i in 0.. String { sb.write_string(result.n_significant().to_string()) sb.write_char('\n') sb.write_string("--- Per-feature results ---\n") - sb.write_string( - "feature\tlogFC\tW\tp_value\tq_value\tsignificant\n", - ) + sb.write_string("feature\tlogFC\tW\tp_value\tq_value\tsignificant\n") for r in result.results { sb.write_string(r.feature) sb.write_char('\t') @@ -689,10 +702,13 @@ pub fn ancombc_result_summary(result : AncombcTestResult) -> String { pub fn ancombc_sample_data() -> AncombcData { let sample_names = ["S1", "S2", "S3", "S4", "S5", "S6", "S7", "S8"] let sample_groups = [ - "control", "control", "control", "control", "treatment", "treatment", - "treatment", "treatment", + "control", "control", "control", "control", "treatment", "treatment", "treatment", + "treatment", + ] + let feature_names = [ + "Bacteroides", "Prevotella", "Faecalibacterium", "Roseburia", "Eubacterium", + "Ruminococcus", ] - let feature_names = ["Bacteroides", "Prevotella", "Faecalibacterium", "Roseburia", "Eubacterium", "Ruminococcus"] // Counts: [feature][sample] // Features 0,1 are elevated in treatment; features 2-5 are similar. let counts = [ diff --git a/src/apeglm.mbt b/src/apeglm.mbt new file mode 100644 index 00000000..a0f4d891 --- /dev/null +++ b/src/apeglm.mbt @@ -0,0 +1,1548 @@ +// Approximate posterior estimation for negative-binomial GLM coefficients. +// +// This module follows Bioconductor apeglm's nbinom workflow. Coefficients and +// prior scales use the natural-log scale; convenience output methods can +// convert effects to log2 scale. + +///| +pub suberror ApeglmError { + ApeglmError(String) +} + +///| +pub struct ApeglmConfig { + coefficient : Int + threshold : Double + interval_level : Double + prior_scale : Double + prior_df : Double + prior_no_shrink_scale : Double + adaptive_prior : Bool + multiplier : Double + max_iterations : Int + tolerance : Double + max_step : Double + random_starts : Int +} derive(Eq, Debug) + +///| +pub fn ApeglmConfig::create( + coefficient? : Int = -1, + threshold? : Double = 0.0, + interval_level? : Double = 0.95, + prior_scale? : Double = 1.0, + prior_df? : Double = 1.0, + prior_no_shrink_scale? : Double = 15.0, + adaptive_prior? : Bool = true, + multiplier? : Double = 1.0, + max_iterations? : Int = 100, + tolerance? : Double = 1.0e-7, + max_step? : Double = 2.0, + random_starts? : Int = 5, +) -> ApeglmConfig raise ApeglmError { + if coefficient < -1 { + raise ApeglmError("apeglm coefficient must be -1 or a zero-based index") + } + if !apeglm_is_finite(threshold) || threshold < 0.0 { + raise ApeglmError("apeglm threshold must be finite and non-negative") + } + if !apeglm_is_finite(interval_level) || + interval_level <= 0.0 || + interval_level >= 1.0 { + raise ApeglmError("apeglm interval level must be finite and in (0, 1)") + } + if !apeglm_is_finite(prior_scale) || prior_scale <= 0.0 { + raise ApeglmError("apeglm prior scale must be finite and positive") + } + if !apeglm_is_finite(prior_df) || prior_df <= 0.0 { + raise ApeglmError( + "apeglm prior degrees of freedom must be finite and positive", + ) + } + if !apeglm_is_finite(prior_no_shrink_scale) || prior_no_shrink_scale <= 0.0 { + raise ApeglmError( + "apeglm no-shrink prior scale must be finite and positive", + ) + } + if !apeglm_is_finite(multiplier) || multiplier <= 0.0 { + raise ApeglmError("apeglm prior multiplier must be finite and positive") + } + if max_iterations < 1 { + raise ApeglmError("apeglm maximum iterations must be positive") + } + if !apeglm_is_finite(tolerance) || tolerance <= 0.0 { + raise ApeglmError("apeglm tolerance must be finite and positive") + } + if !apeglm_is_finite(max_step) || max_step <= 0.0 { + raise ApeglmError("apeglm maximum Newton step must be finite and positive") + } + if random_starts < 1 { + raise ApeglmError("apeglm start count must be positive") + } + ApeglmConfig::{ + coefficient, + threshold, + interval_level, + prior_scale, + prior_df, + prior_no_shrink_scale, + adaptive_prior, + multiplier, + max_iterations, + tolerance, + max_step, + random_starts, + } +} + +///| +pub fn ApeglmConfig::default() -> ApeglmConfig { + ApeglmConfig::{ + coefficient: -1, + threshold: 0.0, + interval_level: 0.95, + prior_scale: 1.0, + prior_df: 1.0, + prior_no_shrink_scale: 15.0, + adaptive_prior: true, + multiplier: 1.0, + max_iterations: 100, + tolerance: 1.0e-7, + max_step: 2.0, + random_starts: 5, + } +} + +///| +pub struct ApeglmPriorControl { + no_shrink : Array[Int] + prior_mean : Double + prior_scale : Double + prior_variance : Double + prior_df : Double + prior_no_shrink_scale : Double + adaptive : Bool +} derive(Debug) + +///| +pub struct ApeglmGeneResult { + gene_name : String + base_mean : Double + mle : Array[Double] + mle_sd : Array[Double] + map : Array[Double] + posterior_sd : Array[Double] + fsr : Double + s_value : Double + threshold_probability : Double + interval_lower : Double + interval_upper : Double + log_posterior : Double + converged : Bool + iterations : Int +} derive(Debug) + +///| +pub struct ApeglmResult { + gene_names : Array[String] + coefficient_names : Array[String] + coefficient : Int + genes : Array[ApeglmGeneResult] + prior_control : ApeglmPriorControl + threshold : Double + interval_level : Double + config : ApeglmConfig +} derive(Debug) + +///| +pub struct ApeglmSummary { + gene_count : Int + converged_count : Int + selected_count : Int + low_fsr_count : Int + prior_scale : Double + median_absolute_effect : Double + coefficient_name : String +} derive(Eq, Debug) + +///| +pub struct ApeglmSummarizedExperimentOutput { + experiment : SummarizedExperiment + result : ApeglmResult +} + +///| +priv struct ApeglmEvaluation { + value : Double + gradient : Array[Double] + information : Array[Array[Double]] +} + +///| +priv struct ApeglmOptimization { + beta : Array[Double] + covariance : Array[Array[Double]] + value : Double + converged : Bool + iterations : Int +} + +///| +fn apeglm_is_finite(value : Double) -> Bool { + value == value && value.abs() <= 1.0e300 +} + +///| +fn apeglm_clamp(value : Double, lower : Double, upper : Double) -> Double { + value.max(lower).min(upper) +} + +///| +fn apeglm_zero_matrix(rows : Int, columns : Int) -> Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for _ in 0.. Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for row in matrix { + output.push(row.copy()) + } + output +} + +///| +fn apeglm_mean(values : Array[Double]) -> Double { + if values.length() == 0 { + return 0.0 + } + let mut total = 0.0 + for value in values { + total = total + value + } + total / values.length().to_double() +} + +///| +fn apeglm_normal_cdf(value : Double) -> Double { + let absolute = value.abs() + let t = 1.0 / (1.0 + 0.2316419 * absolute) + let density = 0.3989422804014327 * @math.exp(-0.5 * absolute * absolute) + let tail = density * + t * + ( + 0.319381530 + + t * + (-0.356563782 + t * (1.781477937 + t * (-1.821255978 + t * 1.330274429))) + ) + if value >= 0.0 { + apeglm_clamp(1.0 - tail, 0.0, 1.0) + } else { + apeglm_clamp(tail, 0.0, 1.0) + } +} + +///| +fn apeglm_normal_quantile(probability : Double) -> Double { + let target = apeglm_clamp(probability, 1.0e-12, 1.0 - 1.0e-12) + let mut lower = -8.0 + let mut upper = 8.0 + for _ in 0..<80 { + let middle = 0.5 * (lower + upper) + if apeglm_normal_cdf(middle) < target { + lower = middle + } else { + upper = middle + } + } + 0.5 * (lower + upper) +} + +///| +fn apeglm_cholesky(matrix : Array[Array[Double]]) -> Array[Array[Double]]? { + let size = matrix.length() + if size == 0 { + return None + } + let lower = apeglm_zero_matrix(size, size) + for row in 0.. Array[Array[Double]]? { + let mut ridge = 0.0 + for attempt in 0..<14 { + let candidate = apeglm_copy_matrix(matrix) + for index in 0.. return Some(lower) + None => ridge = if attempt == 0 { 1.0e-10 } else { ridge * 10.0 } + } + } + None +} + +///| +fn apeglm_cholesky_solve( + lower : Array[Array[Double]], + right_hand_side : Array[Double], +) -> Array[Double] { + let size = lower.length() + let forward = Array::make(size, 0.0) + for row in 0..= 0 { + let mut value = forward[row] + for column in (row + 1).. Array[Array[Double]] { + let size = lower.length() + let inverse = apeglm_zero_matrix(size, size) + for column in 0.. Array[Double] { + match apeglm_regularized_cholesky(information) { + Some(lower) => apeglm_cholesky_solve(lower, gradient) + None => { + let output = Array::make(gradient.length(), 0.0) + for index in 0.. Array[Array[Double]] { + match apeglm_regularized_cholesky(information) { + Some(lower) => apeglm_cholesky_inverse(lower) + None => { + let output = apeglm_zero_matrix( + information.length(), + information.length(), + ) + for index in 0.. Unit raise ApeglmError { + if matrix.length() != rows { + raise ApeglmError(label + " row count does not match the expected shape") + } + for row in matrix { + if row.length() != columns { + raise ApeglmError(label + " must be rectangular with the expected shape") + } + for value in row { + if !apeglm_is_finite(value) || (non_negative && value < 0.0) { + raise ApeglmError(label + " contains an invalid value") + } + } + } +} + +///| +fn apeglm_validate_inputs( + counts : Array[Array[Double]], + design : Array[Array[Double]], + dispersions : Array[Double], + offsets : Array[Array[Double]], + weights : Array[Array[Double]], + gene_names : Array[String], + coefficient_names : Array[String], + config : ApeglmConfig, +) -> (Int, Int, Int, Int) raise ApeglmError { + if counts.length() == 0 { + raise ApeglmError("apeglm counts must contain at least one feature") + } + let samples = counts[0].length() + if samples < 2 { + raise ApeglmError("apeglm counts must contain at least two samples") + } + apeglm_validate_matrix( + counts, + counts.length(), + samples, + "apeglm counts", + true, + ) + for row in counts { + for value in row { + if (value - value.round()).abs() > 1.0e-8 { + raise ApeglmError("apeglm requires integer-valued counts") + } + } + } + if design.length() != samples { + raise ApeglmError("apeglm design rows must match the sample count") + } + if design.length() == 0 || design[0].length() < 2 { + raise ApeglmError( + "apeglm negative-binomial design requires an intercept and a coefficient", + ) + } + let coefficients = design[0].length() + apeglm_validate_matrix(design, samples, coefficients, "apeglm design", false) + let coefficient = if config.coefficient < 0 { + coefficients - 1 + } else { + config.coefficient + } + if coefficient <= 0 || coefficient >= coefficients { + raise ApeglmError( + "apeglm coefficient must select a non-intercept design column", + ) + } + if dispersions.length() != counts.length() { + raise ApeglmError("apeglm dispersions must match the feature count") + } + for dispersion in dispersions { + if !apeglm_is_finite(dispersion) || dispersion <= 0.0 { + raise ApeglmError("apeglm dispersions must be finite and positive") + } + } + if offsets.length() > 0 { + apeglm_validate_matrix( + offsets, + counts.length(), + samples, + "apeglm offsets", + false, + ) + } + if weights.length() > 0 { + apeglm_validate_matrix( + weights, + counts.length(), + samples, + "apeglm weights", + true, + ) + for row in weights { + let mut positive = false + for value in row { + if value > 0.0 { + positive = true + } + } + if !positive { + raise ApeglmError( + "apeglm each feature must have at least one positive weight", + ) + } + } + } + if gene_names.length() > 0 && gene_names.length() != counts.length() { + raise ApeglmError("apeglm gene names must match the feature count") + } + if coefficient_names.length() > 0 && + coefficient_names.length() != coefficients { + raise ApeglmError( + "apeglm coefficient names must match the design column count", + ) + } + (counts.length(), samples, coefficients, coefficient) +} + +///| +fn apeglm_names( + names : Array[String], + count : Int, + prefix : String, +) -> Array[String] raise ApeglmError { + if names.length() == 0 { + let generated : Array[String] = [] + for index in 0.. Array[Array[Double]] { + if offsets.length() == 0 { + apeglm_zero_matrix(genes, samples) + } else { + apeglm_copy_matrix(offsets) + } +} + +///| +fn apeglm_weights( + weights : Array[Array[Double]], + genes : Int, + samples : Int, +) -> Array[Array[Double]] { + if weights.length() == 0 { + let output : Array[Array[Double]] = [] + for _ in 0.. Double raise ApeglmError { + if counts.length() == 0 || design.length() != counts.length() { + raise ApeglmError( + "apeglm log-likelihood counts and design must have matching rows", + ) + } + let coefficients = beta.length() + for row in design { + if row.length() != coefficients { + raise ApeglmError("apeglm log-likelihood design columns must match beta") + } + } + if !apeglm_is_finite(dispersion) || dispersion <= 0.0 { + raise ApeglmError("apeglm log-likelihood dispersion must be positive") + } + if offsets.length() > 0 && offsets.length() != counts.length() { + raise ApeglmError("apeglm log-likelihood offsets have invalid length") + } + if weights.length() > 0 && weights.length() != counts.length() { + raise ApeglmError("apeglm log-likelihood weights have invalid length") + } + let size = 1.0 / dispersion + let mut result = 0.0 + for sample in 0.. ApeglmEvaluation { + let coefficients = beta.length() + let gradient = Array::make(coefficients, 0.0) + let information = apeglm_zero_matrix(coefficients, coefficients) + let size = 1.0 / dispersion + let mut value = 0.0 + for sample in 0.. Array[Double] { + let coefficients = design[0].length() + let beta = Array::make(coefficients, 0.0) + let mut count_total = 0.0 + let mut exposure_total = 0.0 + for sample in 0.. ApeglmOptimization { + let beta = initial.copy() + let mut evaluation = apeglm_evaluate( + counts, + design, + dispersion, + offsets, + weights, + beta, + coefficient, + with_prior, + prior_scale, + config.prior_df, + config.prior_no_shrink_scale, + ) + let mut converged = false + let mut iterations = 0 + for iteration in 0.. config.max_step { + config.max_step / maximum + } else { + 1.0 + } + let mut line_scale = step_scale + let mut accepted = false + let mut accepted_change = 0.0 + let mut accepted_beta = beta.copy() + let mut accepted_evaluation = evaluation + for _ in 0..<28 { + let candidate = beta.copy() + let mut change = 0.0 + for index in 0..= evaluation.value - 1.0e-10 { + accepted = true + accepted_change = change + accepted_beta = candidate + accepted_evaluation = next + break + } + line_scale = line_scale * 0.5 + } + if !accepted { + let mut maximum_gradient = 0.0 + for value in evaluation.gradient { + maximum_gradient = maximum_gradient.max(value.abs()) + } + converged = maximum_gradient <= config.tolerance * 10.0 + break + } + for index in 0.. Array[Array[Double]] { + let starts : Array[Array[Double]] = [mle.copy()] + if starts.length() < count { + let center = mle.copy() + center[coefficient] = 0.0 + starts.push(center) + } + let locations = [ + 2.0 * prior_scale, + -2.0 * prior_scale, + 4.0 * prior_scale, + -4.0 * prior_scale, + prior_scale, + -prior_scale, + ] + let mut index = 0 + while starts.length() < count && index < locations.length() { + let candidate = mle.copy() + candidate[coefficient] = locations[index] + starts.push(candidate) + index = index + 1 + } + let mut shell = 8.0 + while starts.length() < count { + let candidate = mle.copy() + let sign = if starts.length() % 2 == 0 { 1.0 } else { -1.0 } + candidate[coefficient] = apeglm_clamp( + sign * shell * prior_scale, + -30.0, + 30.0, + ) + starts.push(candidate) + if sign < 0.0 { + shell = shell * 2.0 + } + } + starts +} + +///| +fn apeglm_fit_map( + counts : Array[Double], + design : Array[Array[Double]], + dispersion : Double, + offsets : Array[Double], + weights : Array[Double], + mle : Array[Double], + coefficient : Int, + prior_scale : Double, + config : ApeglmConfig, +) -> ApeglmOptimization { + let starts = apeglm_target_starts( + mle, + coefficient, + prior_scale, + config.random_starts, + ) + let mut best = apeglm_optimize( + counts, + design, + dispersion, + offsets, + weights, + starts[0], + coefficient, + true, + prior_scale, + config, + ) + for index in 1.. best.value { + best = candidate + } + } + best +} + +///| +fn apeglm_standard_errors(covariance : Array[Array[Double]]) -> Array[Double] { + let output = Array::make(covariance.length(), 0.0) + for index in 0.. Double { + let mut numerator = 0.0 + let mut denominator = 0.0 + for index in 0.. Double raise ApeglmError { + if estimates.length() == 0 || estimates.length() != standard_errors.length() { + raise ApeglmError( + "apeglm prior adaptation requires matching non-empty MLE and SE arrays", + ) + } + if !apeglm_is_finite(minimum) || + !apeglm_is_finite(maximum) || + minimum <= 0.0 || + maximum <= minimum { + raise ApeglmError("apeglm prior variance bounds are invalid") + } + for index in 0..= 0.0 { + return maximum + } + let mut lower = minimum + let mut upper = maximum + for _ in 0..<100 { + let middle = 0.5 * (lower + upper) + if apeglm_prior_objective(middle, estimates, standard_errors) > 0.0 { + lower = middle + } else { + upper = middle + } + } + 0.5 * (lower + upper) +} + +///| +pub fn apeglm_s_values(local_fsr : Array[Double]) -> Array[Double] { + let count = local_fsr.length() + let order : Array[Int] = [] + for index in 0.. second || (first == second && order[right] > order[right + 1]) { + let temporary = order[right] + order[right] = order[right + 1] + order[right + 1] = temporary + } + } + } + let output = Array::make(count, 0.0) + let mut cumulative = 0.0 + for rank in 0.. Double { + let normalized : Array[Double] = [] + for sample in 0.. ApeglmGeneResult { + ApeglmGeneResult::{ + gene_name: gene.gene_name, + base_mean: gene.base_mean, + mle: gene.mle, + mle_sd: gene.mle_sd, + map: gene.map, + posterior_sd: gene.posterior_sd, + fsr: gene.fsr, + s_value, + threshold_probability: gene.threshold_probability, + interval_lower: gene.interval_lower, + interval_upper: gene.interval_upper, + log_posterior: gene.log_posterior, + converged: gene.converged, + iterations: gene.iterations, + } +} + +///| +pub fn apeglm_fit( + counts : Array[Array[Double]], + design : Array[Array[Double]], + dispersions : Array[Double], + config? : ApeglmConfig = ApeglmConfig::default(), + offsets? : Array[Array[Double]] = [], + weights? : Array[Array[Double]] = [], + gene_names? : Array[String] = [], + coefficient_names? : Array[String] = [], +) -> ApeglmResult raise ApeglmError { + let (gene_count, sample_count, coefficient_count, coefficient) = apeglm_validate_inputs( + counts, design, dispersions, offsets, weights, gene_names, coefficient_names, + config, + ) + let names = apeglm_names(gene_names, gene_count, "feature_") + let coefficients = apeglm_names( + coefficient_names, coefficient_count, "coefficient_", + ) + let model_offsets = apeglm_offsets(offsets, gene_count, sample_count) + let model_weights = apeglm_weights(weights, gene_count, sample_count) + let mle_fits : Array[ApeglmOptimization] = [] + let target_mle : Array[Double] = [] + let target_se : Array[Double] = [] + for gene in 0..= 0.0 { + apeglm_normal_cdf((config.threshold - estimate) / standard_error) + } else { + 1.0 - apeglm_normal_cdf((-config.threshold - estimate) / standard_error) + } + local_fsr.push(fsr) + raw_genes.push(ApeglmGeneResult::{ + gene_name: names[gene], + base_mean: apeglm_base_mean(counts[gene], model_offsets[gene]), + mle: mle_fit.beta.copy(), + mle_sd, + map: map_fit.beta.copy(), + posterior_sd, + fsr, + s_value: 0.0, + threshold_probability: apeglm_clamp(threshold_probability, 0.0, 1.0), + interval_lower: estimate - quantile * standard_error, + interval_upper: estimate + quantile * standard_error, + log_posterior: map_fit.value, + converged: map_fit.converged, + iterations: map_fit.iterations, + }) + } + let s_values = apeglm_s_values(local_fsr) + let genes : Array[ApeglmGeneResult] = [] + for index in 0.. Double { + let value = self.map[coefficient] + if log2_scale { + value / @math.ln(2.0) + } else { + value + } +} + +///| +pub fn ApeglmGeneResult::standard_error( + self : ApeglmGeneResult, + coefficient : Int, + log2_scale? : Bool = false, +) -> Double { + let value = self.posterior_sd[coefficient] + if log2_scale { + value / @math.ln(2.0) + } else { + value + } +} + +///| +pub fn ApeglmResult::map_matrix(self : ApeglmResult) -> Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for gene in self.genes { + output.push(gene.map.copy()) + } + output +} + +///| +pub fn ApeglmResult::sd_matrix(self : ApeglmResult) -> Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for gene in self.genes { + output.push(gene.posterior_sd.copy()) + } + output +} + +///| +pub fn ApeglmResult::gene( + self : ApeglmResult, + index : Int, +) -> ApeglmGeneResult? { + if index < 0 || index >= self.genes.length() { + None + } else { + Some(self.genes[index]) + } +} + +///| +pub fn ApeglmResult::find_gene( + self : ApeglmResult, + name : String, +) -> ApeglmGeneResult? { + for gene in self.genes { + if gene.gene_name == name { + return Some(gene) + } + } + None +} + +///| +pub fn ApeglmResult::ranked(self : ApeglmResult) -> Array[ApeglmGeneResult] { + let output = self.genes.copy() + for left in 0.. second.s_value || + (first.s_value == second.s_value && first_effect < second_effect) { + output[right] = second + output[right + 1] = first + } + } + } + output +} + +///| +pub fn ApeglmResult::select( + self : ApeglmResult, + maximum_s_value? : Double = 0.05, + minimum_absolute_effect? : Double = 0.0, + use_threshold_probability? : Bool = false, +) -> Array[ApeglmGeneResult] { + let output : Array[ApeglmGeneResult] = [] + for gene in self.genes { + let probability = if use_threshold_probability { + gene.threshold_probability + } else { + gene.s_value + } + if probability <= maximum_s_value && + gene.map[self.coefficient].abs() >= minimum_absolute_effect { + output.push(gene) + } + } + output +} + +///| +pub fn ApeglmResult::summary( + self : ApeglmResult, + maximum_s_value? : Double = 0.05, +) -> ApeglmSummary { + let mut converged_count = 0 + let mut low_fsr_count = 0 + let effects : Array[Double] = [] + for gene in self.genes { + if gene.converged { + converged_count = converged_count + 1 + } + if gene.fsr <= maximum_s_value { + low_fsr_count = low_fsr_count + 1 + } + effects.push(gene.map[self.coefficient].abs()) + } + effects.sort() + let median_absolute_effect = if effects.length() == 0 { + 0.0 + } else if effects.length() % 2 == 1 { + effects[effects.length() / 2] + } else { + let middle = effects.length() / 2 + 0.5 * (effects[middle - 1] + effects[middle]) + } + ApeglmSummary::{ + gene_count: self.genes.length(), + converged_count, + selected_count: self.select(maximum_s_value~).length(), + low_fsr_count, + prior_scale: self.prior_control.prior_scale, + median_absolute_effect, + coefficient_name: self.coefficient_names[self.coefficient], + } +} + +///| +pub fn ApeglmResult::to_tsv( + self : ApeglmResult, + log2_scale? : Bool = true, +) -> String { + let output = StringBuilder::new() + let scale_label = if log2_scale { "log2" } else { "ln" } + output.write_string( + "gene\tbaseMean\t" + + scale_label + + "MLE\t" + + scale_label + + "MAP\tposteriorSD\tFSR\tsvalue\tFSOS\tintervalLower\tintervalUpper\tconverged\n", + ) + let divisor = if log2_scale { @math.ln(2.0) } else { 1.0 } + for gene in self.genes { + output.write_string(gene.gene_name) + output.write_string("\t") + output.write_string(gene.base_mean.to_string()) + output.write_string("\t") + output.write_string((gene.mle[self.coefficient] / divisor).to_string()) + output.write_string("\t") + output.write_string((gene.map[self.coefficient] / divisor).to_string()) + output.write_string("\t") + output.write_string( + (gene.posterior_sd[self.coefficient] / divisor).to_string(), + ) + output.write_string("\t") + output.write_string(gene.fsr.to_string()) + output.write_string("\t") + output.write_string(gene.s_value.to_string()) + output.write_string("\t") + output.write_string(gene.threshold_probability.to_string()) + output.write_string("\t") + output.write_string((gene.interval_lower / divisor).to_string()) + output.write_string("\t") + output.write_string((gene.interval_upper / divisor).to_string()) + output.write_string("\t") + output.write_string(gene.converged.to_string()) + output.write_string("\n") + } + output.to_string() +} + +///| +pub fn apeglm_from_deseq2( + dataset : DESeqDataSet, + config? : ApeglmConfig = ApeglmConfig::default(), + coefficient_names? : Array[String] = [], +) -> ApeglmResult raise ApeglmError { + let counts : Array[Array[Double]] = [] + for row in dataset.counts { + let converted : Array[Double] = [] + for value in row { + converted.push(value.to_double()) + } + counts.push(converted) + } + let samples = if counts.length() == 0 { 0 } else { counts[0].length() } + if dataset.size_factors.length() != samples { + raise ApeglmError("apeglm DESeq2 size factors must match the sample count") + } + let offsets : Array[Array[Double]] = [] + for _ in 0.. SummarizedExperiment { + let assays : Map[String, Array[Array[Double]]] = Map([]) + for key in experiment.assays.keys() { + assays[key] = apeglm_copy_matrix(experiment.assays[key]) + } + let metadata : Map[String, String] = Map([]) + for key in experiment.metadata.keys() { + metadata[key] = experiment.metadata[key] + } + SummarizedExperiment::{ + assays, + row_ranges: experiment.row_ranges.copy(), + col_data: experiment.col_data.copy(), + metadata, + } +} + +///| +pub fn apeglm_summarized_experiment( + experiment : SummarizedExperiment, + design : Array[Array[Double]], + dispersions : Array[Double], + config? : ApeglmConfig = ApeglmConfig::default(), + assay_name? : String = "counts", + offsets? : Array[Array[Double]] = [], + weights? : Array[Array[Double]] = [], + gene_names? : Array[String] = [], + coefficient_names? : Array[String] = [], +) -> ApeglmSummarizedExperimentOutput raise ApeglmError { + let counts = match se_assay(experiment, assay_name) { + Some(value) => value + None => + raise ApeglmError( + "apeglm SummarizedExperiment assay not found: " + assay_name, + ) + } + let result = apeglm_fit( + counts, + design, + dispersions, + config~, + offsets~, + weights~, + gene_names~, + coefficient_names~, + ) + let enriched = apeglm_copy_summarized_experiment(experiment) + let map : Array[Array[Double]] = [] + let standard_error : Array[Array[Double]] = [] + let fsr : Array[Array[Double]] = [] + let s_value : Array[Array[Double]] = [] + let threshold_probability : Array[Array[Double]] = [] + for gene in result.genes { + map.push([gene.map[result.coefficient]]) + standard_error.push([gene.posterior_sd[result.coefficient]]) + fsr.push([gene.fsr]) + s_value.push([gene.s_value]) + threshold_probability.push([gene.threshold_probability]) + } + enriched.assays["apeglm_map"] = map + enriched.assays["apeglm_sd"] = standard_error + enriched.assays["apeglm_fsr"] = fsr + enriched.assays["apeglm_svalue"] = s_value + enriched.assays["apeglm_fsos"] = threshold_probability + enriched.metadata["apeglm_coefficient"] = result.coefficient_names[result.coefficient] + enriched.metadata["apeglm_prior_scale"] = result.prior_control.prior_scale.to_string() + enriched.metadata["apeglm_scale"] = "natural-log" + ApeglmSummarizedExperimentOutput::{ experiment: enriched, result } +} + +///| +pub fn apeglm_example_data() -> ( + Array[Array[Double]], + Array[Array[Double]], + Array[Double], + Array[String], + Array[String], +) { + let counts = [ + [48.0, 52.0, 45.0, 54.0, 198.0, 220.0, 205.0, 230.0], + [160.0, 148.0, 171.0, 155.0, 39.0, 44.0, 35.0, 41.0], + [72.0, 68.0, 75.0, 70.0, 79.0, 74.0, 77.0, 73.0], + [0.0, 1.0, 0.0, 0.0, 4.0, 0.0, 5.0, 0.0], + [12.0, 20.0, 15.0, 18.0, 27.0, 19.0, 31.0, 23.0], + [310.0, 290.0, 325.0, 305.0, 640.0, 615.0, 670.0, 650.0], + [9.0, 8.0, 11.0, 10.0, 8.0, 12.0, 9.0, 11.0], + [3.0, 0.0, 2.0, 1.0, 0.0, 1.0, 0.0, 2.0], + ] + let design = [ + [1.0, 0.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 0.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 0.0, 1.0], + [1.0, 1.0, 1.0], + [1.0, 0.0, 1.0], + [1.0, 1.0, 1.0], + ] + let dispersions = [0.08, 0.08, 0.1, 0.8, 0.25, 0.05, 0.2, 0.9] + let genes = [ + "strong_up", "strong_down", "stable", "low_noisy", "moderate_up", "abundant_up", + "null_low", "sparse_null", + ] + let coefficient_names = ["intercept", "batch", "condition"] + (counts, design, dispersions, genes, coefficient_names) +} diff --git a/src/application.mbt b/src/application.mbt index d4287a84..7c604cf2 100644 --- a/src/application.mbt +++ b/src/application.mbt @@ -19,12 +19,16 @@ pub fn AbstractCommandline::new(executable : String) -> AbstractCommandline { stdin: "", stdout: "", stderr: "", - env: Map([], capacity=0) + env: Map([], capacity=0), } } ///| -pub fn AbstractCommandline::add_arg(self : AbstractCommandline, arg : String, value : String?) -> AbstractCommandline { +pub fn AbstractCommandline::add_arg( + self : AbstractCommandline, + arg : String, + value : String?, +) -> AbstractCommandline { let args = self.arguments.copy() args.push(arg) if value is Some(_) { @@ -36,12 +40,16 @@ pub fn AbstractCommandline::add_arg(self : AbstractCommandline, arg : String, va stdin: self.stdin, stdout: self.stdout, stderr: self.stderr, - env: self.env + env: self.env, } } ///| -pub fn AbstractCommandline::set_parameter(self : AbstractCommandline, param : String, value : String) -> AbstractCommandline { +pub fn AbstractCommandline::set_parameter( + self : AbstractCommandline, + param : String, + value : String, +) -> AbstractCommandline { let args = self.arguments.copy() args.push(param) args.push(value) @@ -51,43 +59,52 @@ pub fn AbstractCommandline::set_parameter(self : AbstractCommandline, param : St stdin: self.stdin, stdout: self.stdout, stderr: self.stderr, - env: self.env + env: self.env, } } ///| -pub fn AbstractCommandline::set_stdout(self : AbstractCommandline, path : String) -> AbstractCommandline { +pub fn AbstractCommandline::set_stdout( + self : AbstractCommandline, + path : String, +) -> AbstractCommandline { AbstractCommandline::{ executable: self.executable, arguments: self.arguments, stdin: self.stdin, stdout: path, stderr: self.stderr, - env: self.env + env: self.env, } } ///| -pub fn AbstractCommandline::set_stderr(self : AbstractCommandline, path : String) -> AbstractCommandline { +pub fn AbstractCommandline::set_stderr( + self : AbstractCommandline, + path : String, +) -> AbstractCommandline { AbstractCommandline::{ executable: self.executable, arguments: self.arguments, stdin: self.stdin, stdout: self.stdout, stderr: path, - env: self.env + env: self.env, } } ///| -pub fn AbstractCommandline::set_stdin(self : AbstractCommandline, path : String) -> AbstractCommandline { +pub fn AbstractCommandline::set_stdin( + self : AbstractCommandline, + path : String, +) -> AbstractCommandline { AbstractCommandline::{ executable: self.executable, arguments: self.arguments, stdin: path, stdout: self.stdout, stderr: self.stderr, - env: self.env + env: self.env, } } @@ -115,7 +132,11 @@ pub struct CommandlineError { } ///| -pub fn CommandlineError::new(message : String, exit_code : Int, command : String) -> CommandlineError { +pub fn CommandlineError::new( + message : String, + exit_code : Int, + command : String, +) -> CommandlineError { CommandlineError::{ message, exit_code, command } } @@ -142,4 +163,4 @@ pub fn create_example_commandline() -> AbstractCommandline { cmd = cmd.add_arg("-out", Some("results.txt")) cmd = cmd.add_arg("-evalue", Some("1e-5")) cmd -} \ No newline at end of file +} diff --git a/src/bamsignals.mbt b/src/bamsignals.mbt index ae702a29..c71b5678 100644 --- a/src/bamsignals.mbt +++ b/src/bamsignals.mbt @@ -3,6 +3,7 @@ /// Provides functions for extracting and analyzing signals from BAM files. /// Supports signal counting and normalization for ChIP-seq and other sequencing data. +///| /// Count mode for signal extraction pub enum BamsigCountMode { /// Count all reads @@ -15,6 +16,7 @@ pub enum BamsigCountMode { PairedEnd } derive(Eq) +///| /// Normalization method for signal pub enum BamsigNormMethod { /// No normalization @@ -27,6 +29,7 @@ pub enum BamsigNormMethod { CPM } +///| /// Signal extraction parameters pub struct BamsigParams { /// Count mode @@ -41,31 +44,79 @@ pub struct BamsigParams { extend_len : Int } +///| /// Create default parameters pub fn BamsigParams::new() -> BamsigParams { - BamsigParams::{ count_mode: BamsigCountMode::All, filter_dup: false, min_mapq: 10, single_end: false, extend_len: 150 } + BamsigParams::{ + count_mode: BamsigCountMode::All, + filter_dup: false, + min_mapq: 10, + single_end: false, + extend_len: 150, + } } +///| /// Set count mode -pub fn BamsigParams::bamsig_set_count_mode(self : BamsigParams, mode : BamsigCountMode) -> BamsigParams { - BamsigParams::{ count_mode: mode, filter_dup: self.filter_dup, min_mapq: self.min_mapq, single_end: self.single_end, extend_len: self.extend_len } +pub fn BamsigParams::bamsig_set_count_mode( + self : BamsigParams, + mode : BamsigCountMode, +) -> BamsigParams { + BamsigParams::{ + count_mode: mode, + filter_dup: self.filter_dup, + min_mapq: self.min_mapq, + single_end: self.single_end, + extend_len: self.extend_len, + } } +///| /// Set filter duplicates -pub fn BamsigParams::bamsig_set_filter_dup(self : BamsigParams, filter : Bool) -> BamsigParams { - BamsigParams::{ count_mode: self.count_mode, filter_dup: filter, min_mapq: self.min_mapq, single_end: self.single_end, extend_len: self.extend_len } +pub fn BamsigParams::bamsig_set_filter_dup( + self : BamsigParams, + filter : Bool, +) -> BamsigParams { + BamsigParams::{ + count_mode: self.count_mode, + filter_dup: filter, + min_mapq: self.min_mapq, + single_end: self.single_end, + extend_len: self.extend_len, + } } +///| /// Set minimum mapping quality -pub fn BamsigParams::bamsig_set_min_mapq(self : BamsigParams, mapq : Int) -> BamsigParams { - BamsigParams::{ count_mode: self.count_mode, filter_dup: self.filter_dup, min_mapq: mapq, single_end: self.single_end, extend_len: self.extend_len } +pub fn BamsigParams::bamsig_set_min_mapq( + self : BamsigParams, + mapq : Int, +) -> BamsigParams { + BamsigParams::{ + count_mode: self.count_mode, + filter_dup: self.filter_dup, + min_mapq: mapq, + single_end: self.single_end, + extend_len: self.extend_len, + } } +///| /// Set extend length -pub fn BamsigParams::bamsig_set_extend(self : BamsigParams, extend_len : Int) -> BamsigParams { - BamsigParams::{ count_mode: self.count_mode, filter_dup: self.filter_dup, min_mapq: self.min_mapq, single_end: self.single_end, extend_len: extend_len } +pub fn BamsigParams::bamsig_set_extend( + self : BamsigParams, + extend_len : Int, +) -> BamsigParams { + BamsigParams::{ + count_mode: self.count_mode, + filter_dup: self.filter_dup, + min_mapq: self.min_mapq, + single_end: self.single_end, + extend_len, + } } +///| /// Genomic region for signal extraction pub struct BamsigRegion { /// Chromosome @@ -78,36 +129,48 @@ pub struct BamsigRegion { id : String } +///| /// Create new region -pub fn BamsigRegion::new(chrom : String, start : Int, end : Int, id : String) -> BamsigRegion { +pub fn BamsigRegion::new( + chrom : String, + start : Int, + end : Int, + id : String, +) -> BamsigRegion { BamsigRegion::{ chrom, start, end, id } } +///| /// Get chromosome pub fn BamsigRegion::bamsig_chrom(self : BamsigRegion) -> String { self.chrom } +///| /// Get start position pub fn BamsigRegion::bamsig_start(self : BamsigRegion) -> Int { self.start } +///| /// Get end position pub fn BamsigRegion::bamsig_end(self : BamsigRegion) -> Int { self.end } +///| /// Get width pub fn BamsigRegion::bamsig_width(self : BamsigRegion) -> Int { self.end - self.start } +///| /// Get ID pub fn BamsigRegion::bamsig_id(self : BamsigRegion) -> String { self.id } +///| /// BAM record for signal extraction pub struct BamsigRecord { /// Chromosome @@ -128,23 +191,39 @@ pub struct BamsigRecord { is_first : Bool } +///| /// Create new record -pub fn BamsigRecord::new(chrom : String, pos : Int, cigar : String, strand : Bool, mapq : Int, is_dup : Bool, is_paired : Bool, is_first : Bool) -> BamsigRecord { +pub fn BamsigRecord::new( + chrom : String, + pos : Int, + cigar : String, + strand : Bool, + mapq : Int, + is_dup : Bool, + is_paired : Bool, + is_first : Bool, +) -> BamsigRecord { BamsigRecord::{ chrom, pos, cigar, strand, mapq, is_dup, is_paired, is_first } } +///| /// Get alignment length from CIGAR pub fn BamsigRecord::bamsig_align_length(self : BamsigRecord) -> Int { parse_cigar_length(self.cigar) } +///| /// Get end position pub fn BamsigRecord::bamsig_end(self : BamsigRecord) -> Int { self.pos + self.bamsig_align_length() } +///| /// Check if record passes filter -pub fn BamsigRecord::bamsig_passes_filter(self : BamsigRecord, params : BamsigParams) -> Bool { +pub fn BamsigRecord::bamsig_passes_filter( + self : BamsigRecord, + params : BamsigParams, +) -> Bool { if self.is_dup && params.filter_dup { return false } @@ -157,6 +236,7 @@ pub fn BamsigRecord::bamsig_passes_filter(self : BamsigRecord, params : BamsigPa true } +///| /// Signal counts for regions pub struct BamsigSignal { /// Region IDs @@ -169,39 +249,68 @@ pub struct BamsigSignal { total_reads : Array[Double] } +///| /// Create new signal -pub fn BamsigSignal::new(region_ids : Array[String], counts : Array[Array[Double]], - norm_counts : Array[Array[Double]], total_reads : Array[Double]) -> BamsigSignal { +pub fn BamsigSignal::new( + region_ids : Array[String], + counts : Array[Array[Double]], + norm_counts : Array[Array[Double]], + total_reads : Array[Double], +) -> BamsigSignal { BamsigSignal::{ region_ids, counts, norm_counts, total_reads } } +///| /// Get number of regions pub fn BamsigSignal::bamsig_n_regions(self : BamsigSignal) -> Int { self.region_ids.length() } +///| /// Get number of samples pub fn BamsigSignal::bamsig_n_samples(self : BamsigSignal) -> Int { - if self.counts.length() > 0 { self.counts[0].length() } else { 0 } + if self.counts.length() > 0 { + self.counts[0].length() + } else { + 0 + } } +///| /// Get count for region and sample -pub fn BamsigSignal::bamsig_get_count(self : BamsigSignal, region_idx : Int, sample_idx : Int) -> Double { +pub fn BamsigSignal::bamsig_get_count( + self : BamsigSignal, + region_idx : Int, + sample_idx : Int, +) -> Double { self.counts[region_idx][sample_idx] } +///| /// Get normalized count -pub fn BamsigSignal::bamsig_get_norm_count(self : BamsigSignal, region_idx : Int, sample_idx : Int) -> Double { +pub fn BamsigSignal::bamsig_get_norm_count( + self : BamsigSignal, + region_idx : Int, + sample_idx : Int, +) -> Double { self.norm_counts[region_idx][sample_idx] } +///| /// Get region ID -pub fn BamsigSignal::bamsig_get_region_id(self : BamsigSignal, idx : Int) -> String { +pub fn BamsigSignal::bamsig_get_region_id( + self : BamsigSignal, + idx : Int, +) -> String { self.region_ids[idx] } +///| /// Sum counts across samples for a region -pub fn BamsigSignal::bamsig_row_sum(self : BamsigSignal, region_idx : Int) -> Double { +pub fn BamsigSignal::bamsig_row_sum( + self : BamsigSignal, + region_idx : Int, +) -> Double { let row = self.counts[region_idx] let mut sum = 0.0 let mut i = 0 @@ -212,8 +321,12 @@ pub fn BamsigSignal::bamsig_row_sum(self : BamsigSignal, region_idx : Int) -> Do sum } +///| /// Sum counts across regions for a sample -pub fn BamsigSignal::bamsig_col_sum(self : BamsigSignal, sample_idx : Int) -> Double { +pub fn BamsigSignal::bamsig_col_sum( + self : BamsigSignal, + sample_idx : Int, +) -> Double { let mut sum = 0.0 let mut i = 0 while i < self.counts.length() { @@ -223,17 +336,22 @@ pub fn BamsigSignal::bamsig_col_sum(self : BamsigSignal, sample_idx : Int) -> Do sum } +///| /// Filter regions by count threshold -pub fn BamsigSignal::bamsig_filter(self : BamsigSignal, min_count : Double, min_samples : Int) -> BamsigSignal { +pub fn BamsigSignal::bamsig_filter( + self : BamsigSignal, + min_count : Double, + min_samples : Int, +) -> BamsigSignal { let new_region_ids : Array[String] = Array::new() let new_counts : Array[Array[Double]] = Array::new() let new_norm_counts : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < self.region_ids.length() { let count_row = self.counts[i] let mut passing_samples = 0 - + let mut j = 0 while j < count_row.length() { if count_row[j] >= min_count { @@ -241,19 +359,25 @@ pub fn BamsigSignal::bamsig_filter(self : BamsigSignal, min_count : Double, min_ } j = j + 1 } - + if passing_samples >= min_samples { new_region_ids.push(self.region_ids[i]) new_counts.push(count_row) new_norm_counts.push(self.norm_counts[i]) } - + i = i + 1 } - - BamsigSignal::new(new_region_ids, new_counts, new_norm_counts, self.total_reads) + + BamsigSignal::new( + new_region_ids, + new_counts, + new_norm_counts, + self.total_reads, + ) } +///| /// Chromatin state analysis result pub struct BamsigChromState { /// State labels @@ -270,43 +394,73 @@ pub struct BamsigChromState { repressed_pct : Double } +///| /// Create new chromatin state result -pub fn BamsigChromState::new(labels : Array[String], colors : Array[String], frequencies : Array[Double], - active_promoter_pct : Double, active_enhancer_pct : Double, - repressed_pct : Double) -> BamsigChromState { - BamsigChromState::{ labels, colors, frequencies, active_promoter_pct, active_enhancer_pct, repressed_pct } +pub fn BamsigChromState::new( + labels : Array[String], + colors : Array[String], + frequencies : Array[Double], + active_promoter_pct : Double, + active_enhancer_pct : Double, + repressed_pct : Double, +) -> BamsigChromState { + BamsigChromState::{ + labels, + colors, + frequencies, + active_promoter_pct, + active_enhancer_pct, + repressed_pct, + } } +///| /// Get number of states pub fn BamsigChromState::bamsig_n_states(self : BamsigChromState) -> Int { self.labels.length() } +///| /// Get state label -pub fn BamsigChromState::bamsig_get_label(self : BamsigChromState, idx : Int) -> String { +pub fn BamsigChromState::bamsig_get_label( + self : BamsigChromState, + idx : Int, +) -> String { self.labels[idx] } +///| /// Get state frequency -pub fn BamsigChromState::bamsig_get_frequency(self : BamsigChromState, idx : Int) -> Double { +pub fn BamsigChromState::bamsig_get_frequency( + self : BamsigChromState, + idx : Int, +) -> Double { self.frequencies[idx] } +///| /// Get active promoter percentage -pub fn BamsigChromState::bamsig_active_promoter(self : BamsigChromState) -> Double { +pub fn BamsigChromState::bamsig_active_promoter( + self : BamsigChromState, +) -> Double { self.active_promoter_pct } +///| /// Get active enhancer percentage -pub fn BamsigChromState::bamsig_active_enhancer(self : BamsigChromState) -> Double { +pub fn BamsigChromState::bamsig_active_enhancer( + self : BamsigChromState, +) -> Double { self.active_enhancer_pct } +///| /// Get repressed percentage pub fn BamsigChromState::bamsig_repressed(self : BamsigChromState) -> Double { self.repressed_pct } +///| /// Parse CIGAR string to get alignment length fn parse_cigar_length(cigar : String) -> Int { let mut total = 0 @@ -335,6 +489,7 @@ fn parse_cigar_length(cigar : String) -> Int { total } +///| /// Create example records for testing pub fn bamsig_create_example_records() -> Array[BamsigRecord] { [ @@ -347,6 +502,7 @@ pub fn bamsig_create_example_records() -> Array[BamsigRecord] { ] } +///| /// Create example regions pub fn bamsig_create_example_regions() -> Array[BamsigRegion] { [ @@ -357,18 +513,21 @@ pub fn bamsig_create_example_regions() -> Array[BamsigRegion] { ] } +///| /// Count signals in regions -pub fn bamsig_count_signals(regions : Array[BamsigRegion], - records_by_chrom : Map[String, Array[BamsigRecord]], - params : BamsigParams) -> Array[Array[Double]] { +pub fn bamsig_count_signals( + regions : Array[BamsigRegion], + records_by_chrom : Map[String, Array[BamsigRecord]], + params : BamsigParams, +) -> Array[Array[Double]] { let n_regions = regions.length() let counts : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < n_regions { let region = regions[i] let mut region_count = 0.0 - + let chrom_records = records_by_chrom.get(region.chrom) if chrom_records is Some(records) { let recs = records @@ -378,15 +537,28 @@ pub fn bamsig_count_signals(regions : Array[BamsigRegion], if rec.bamsig_passes_filter(params) { let rec_end = rec.pos + rec.bamsig_align_length() if rec.pos < region.end && rec_end > region.start { - let overlap_start = if rec.pos > region.start { rec.pos } else { region.start } - let overlap_end = if rec_end < region.end { rec_end } else { region.end } - let _overlap = if overlap_end > overlap_start { overlap_end - overlap_start } else { 0 } - + let overlap_start = if rec.pos > region.start { + rec.pos + } else { + region.start + } + let overlap_end = if rec_end < region.end { + rec_end + } else { + region.end + } + let _overlap = if overlap_end > overlap_start { + overlap_end - overlap_start + } else { + 0 + } + if params.count_mode == BamsigCountMode::All { region_count = region_count + 1.0 } else if params.count_mode == BamsigCountMode::Sense && rec.strand { region_count = region_count + 1.0 - } else if params.count_mode == BamsigCountMode::Antisense && !rec.strand { + } else if params.count_mode == BamsigCountMode::Antisense && + !rec.strand { region_count = region_count + 1.0 } } @@ -394,77 +566,98 @@ pub fn bamsig_count_signals(regions : Array[BamsigRegion], j = j + 1 } } - + counts.push([region_count]) i = i + 1 } - + counts } +///| /// Normalize signals -pub fn bamsig_normalize_signals(counts : Array[Array[Double]], - total_reads : Array[Double], - norm_method : BamsigNormMethod) -> Array[Array[Double]] { +pub fn bamsig_normalize_signals( + counts : Array[Array[Double]], + total_reads : Array[Double], + norm_method : BamsigNormMethod, +) -> Array[Array[Double]] { let norm_counts : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < counts.length() { let row = counts[i] let norm_row : Array[Double] = Array::new() - + let mut j = 0 while j < row.length() { let c = row[j] let total = if j < total_reads.length() { total_reads[j] } else { 1.0 } - + let norm_val = match norm_method { BamsigNormMethod::None => c - BamsigNormMethod::RPM => if total > 0.0 { c / total * 1000000.0 } else { 0.0 } - BamsigNormMethod::CPM => if total > 0.0 { c / total * 1000000.0 } else { 0.0 } - BamsigNormMethod::RPKM => if total > 0.0 { c / total * 1000000.0 / 1000.0 } else { 0.0 } + BamsigNormMethod::RPM => + if total > 0.0 { + c / total * 1000000.0 + } else { + 0.0 + } + BamsigNormMethod::CPM => + if total > 0.0 { + c / total * 1000000.0 + } else { + 0.0 + } + BamsigNormMethod::RPKM => + if total > 0.0 { + c / total * 1000000.0 / 1000.0 + } else { + 0.0 + } } - + norm_row.push(norm_val) j = j + 1 } - + norm_counts.push(norm_row) i = i + 1 } - + norm_counts } +///| /// Analyze chromatin states -pub fn bamsig_analyze_chromatin_states(signal : BamsigSignal, - promoter_regions : Array[Int], - enhancer_regions : Array[Int], - repressed_regions : Array[Int]) -> BamsigChromState { +pub fn bamsig_analyze_chromatin_states( + signal : BamsigSignal, + promoter_regions : Array[Int], + enhancer_regions : Array[Int], + repressed_regions : Array[Int], +) -> BamsigChromState { let n_regions = signal.bamsig_n_regions() let n_states = 4 - + let labels = ["Active Promoter", "Active Enhancer", "Repressed", "Inactive"] let colors = ["#FF0000", "#00FF00", "#0000FF", "#CCCCCC"] let frequencies = Array::make(n_states, 0.0) - + let n_promoter = promoter_regions.length().to_double() let n_enhancer = enhancer_regions.length().to_double() let n_repressed = repressed_regions.length().to_double() - + let mut active_promoter = 0.0 let mut active_enhancer = 0.0 let mut repressed = 0.0 - + let mut i = 0 while i < n_regions { let total_count = signal.bamsig_row_sum(i) let is_active = total_count > 0.0 - + let is_promoter = promoter_regions.contains(i) let is_enhancer = enhancer_regions.contains(i) let is_repressed = repressed_regions.contains(i) - + if is_active { if is_promoter { active_promoter = active_promoter + 1.0 @@ -472,37 +665,51 @@ pub fn bamsig_analyze_chromatin_states(signal : BamsigSignal, active_enhancer = active_enhancer + 1.0 } } - + if is_repressed { repressed = repressed + 1.0 } - + i = i + 1 } - + let _total_annotated = n_promoter + n_enhancer + n_repressed - let total_promoter_pct = if n_promoter > 0.0 { active_promoter / n_promoter * 100.0 } else { 0.0 } - let enhancer_pct = if n_enhancer > 0.0 { active_enhancer / n_enhancer * 100.0 } else { 0.0 } - let repressed_pct = if n_repressed > 0.0 { repressed / n_repressed * 100.0 } else { 0.0 } - + let total_promoter_pct = if n_promoter > 0.0 { + active_promoter / n_promoter * 100.0 + } else { + 0.0 + } + let enhancer_pct = if n_enhancer > 0.0 { + active_enhancer / n_enhancer * 100.0 + } else { + 0.0 + } + let repressed_pct = if n_repressed > 0.0 { + repressed / n_repressed * 100.0 + } else { + 0.0 + } + frequencies[0] = total_promoter_pct frequencies[1] = enhancer_pct frequencies[2] = repressed_pct frequencies[3] = 100.0 - total_promoter_pct - enhancer_pct - repressed_pct - - BamsigChromState::new(labels, colors, frequencies, - total_promoter_pct, enhancer_pct, repressed_pct) + + BamsigChromState::new( + labels, colors, frequencies, total_promoter_pct, enhancer_pct, repressed_pct, + ) } +///| /// Create example signal data pub fn bamsig_create_example_signal() -> BamsigSignal { let regions = bamsig_create_example_regions() let n_regions = regions.length() let _n_samples = 2 - + let region_ids : Array[String] = Array::new() let counts : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < n_regions { region_ids.push(regions[i].id) @@ -512,9 +719,13 @@ pub fn bamsig_create_example_signal() -> BamsigSignal { counts.push(row) i = i + 1 } - + let total_reads = [10000.0, 9500.0] - let norm_counts = bamsig_normalize_signals(counts, total_reads, BamsigNormMethod::RPM) - + let norm_counts = bamsig_normalize_signals( + counts, + total_reads, + BamsigNormMethod::RPM, + ) + BamsigSignal::new(region_ids, counts, norm_counts, total_reads) } diff --git a/src/banksy.mbt b/src/banksy.mbt new file mode 100644 index 00000000..59ec00d4 --- /dev/null +++ b/src/banksy.mbt @@ -0,0 +1,2070 @@ +// Spatially aware clustering inspired by the Bioconductor Banksy package. +// +// Expression matrices use genes in rows and spots in columns. Neighborhood +// harmonics follow Banksy's H0, H1, ... convention: +// H0 = weighted neighborhood mean +// Hm = |sum_j weight_ij * centered(x_j) * exp(i * m * phi_ij)| + +///| +pub suberror BanksyError { + BanksyError(String) +} + +///| +pub(all) enum BanksySpatialMode { + BanksyKnnMedian + BanksyKnnInverseDistance + BanksyKnnInversePower + BanksyKnnRank + BanksyKnnUniform + BanksyRadiusGaussian +} derive(Eq, Debug) + +///| +pub fn banksy_knn_median() -> BanksySpatialMode { + BanksyKnnMedian +} + +///| +pub fn banksy_knn_inverse_distance() -> BanksySpatialMode { + BanksyKnnInverseDistance +} + +///| +pub fn banksy_knn_inverse_power() -> BanksySpatialMode { + BanksyKnnInversePower +} + +///| +pub fn banksy_knn_rank() -> BanksySpatialMode { + BanksyKnnRank +} + +///| +pub fn banksy_knn_uniform() -> BanksySpatialMode { + BanksyKnnUniform +} + +///| +pub fn banksy_radius_gaussian() -> BanksySpatialMode { + BanksyRadiusGaussian +} + +///| +pub struct BanksyComputeConfig { + max_harmonic : Int + k_geom : Array[Int] + spatial_mode : BanksySpatialMode + distance_power : Double + sigma : Double + alpha : Double + k_spatial : Int + sample_size : Int + sample_renormalize : Bool + seed : Int + center_harmonics : Bool +} derive(Debug) + +///| +pub fn BanksyComputeConfig::create( + max_harmonic? : Int = 0, + k_geom? : Array[Int] = [15], + spatial_mode? : BanksySpatialMode = BanksyKnnMedian, + distance_power? : Double = 2.0, + sigma? : Double = 1.5, + alpha? : Double = 0.05, + k_spatial? : Int = 100, + sample_size? : Int = 0, + sample_renormalize? : Bool = true, + seed? : Int = 1, + center_harmonics? : Bool = true, +) -> BanksyComputeConfig raise BanksyError { + if max_harmonic < 0 { + raise BanksyError("BANKSY maximum harmonic must be non-negative") + } + if k_geom.length() != 1 && k_geom.length() != max_harmonic + 1 { + raise BanksyError( + "BANKSY k_geom must contain one value or one value per harmonic", + ) + } + for value in k_geom { + if value < 1 { + raise BanksyError("BANKSY k_geom values must be positive") + } + } + if !banksy_is_finite(distance_power) || distance_power <= 0.0 { + raise BanksyError("BANKSY distance power must be finite and positive") + } + if !banksy_is_finite(sigma) || sigma <= 0.0 { + raise BanksyError("BANKSY Gaussian sigma must be finite and positive") + } + if !banksy_is_finite(alpha) || alpha <= 0.0 || alpha >= 1.0 { + raise BanksyError("BANKSY radial alpha must be finite and in (0, 1)") + } + if k_spatial < 1 { + raise BanksyError("BANKSY radial neighbor count must be positive") + } + if sample_size < 0 { + raise BanksyError("BANKSY sample size must be non-negative") + } + if seed < 0 { + raise BanksyError("BANKSY seed must be non-negative") + } + BanksyComputeConfig::{ + max_harmonic, + k_geom: k_geom.copy(), + spatial_mode, + distance_power, + sigma, + alpha, + k_spatial, + sample_size, + sample_renormalize, + seed, + center_harmonics, + } +} + +///| +pub fn BanksyComputeConfig::default() -> BanksyComputeConfig { + BanksyComputeConfig::{ + max_harmonic: 0, + k_geom: [15], + spatial_mode: BanksyKnnMedian, + distance_power: 2.0, + sigma: 1.5, + alpha: 0.05, + k_spatial: 100, + sample_size: 0, + sample_renormalize: true, + seed: 1, + center_harmonics: true, + } +} + +///| +pub struct BanksyNeighbor { + index : Int + distance : Double + weight : Double + angle : Double +} derive(Eq, Debug) + +///| +pub struct BanksyModel { + expression : Array[Array[Double]] + coordinates : Array[Array[Double]] + gene_names : Array[String] + spot_names : Array[String] + harmonics : Array[Array[Array[Double]]] + neighbors : Array[Array[Array[BanksyNeighbor]]] + config : BanksyComputeConfig +} derive(Debug) + +///| +pub struct BanksyMatrix { + values : Array[Array[Double]] + feature_names : Array[String] + component_weights : Array[Double] + lambda : Double + max_harmonic : Int + scaled : Bool +} derive(Debug) + +///| +pub struct BanksyPcaResult { + scores : Array[Array[Double]] + loadings : Array[Array[Double]] + eigenvalues : Array[Double] + variance_explained : Array[Double] + n_components : Int +} derive(Debug) + +///| +pub struct BanksyClusterConfig { + n_clusters : Int + n_starts : Int + max_iterations : Int + tolerance : Double + seed : Int +} derive(Eq, Debug) + +///| +pub fn BanksyClusterConfig::create( + n_clusters? : Int = 5, + n_starts? : Int = 10, + max_iterations? : Int = 100, + tolerance? : Double = 1.0e-6, + seed? : Int = 1, +) -> BanksyClusterConfig raise BanksyError { + if n_clusters < 1 { + raise BanksyError("BANKSY cluster count must be positive") + } + if n_starts < 1 { + raise BanksyError("BANKSY k-means start count must be positive") + } + if max_iterations < 1 { + raise BanksyError("BANKSY k-means iteration limit must be positive") + } + if !banksy_is_finite(tolerance) || tolerance <= 0.0 { + raise BanksyError("BANKSY k-means tolerance must be finite and positive") + } + if seed < 0 { + raise BanksyError("BANKSY k-means seed must be non-negative") + } + BanksyClusterConfig::{ n_clusters, n_starts, max_iterations, tolerance, seed } +} + +///| +pub fn BanksyClusterConfig::default() -> BanksyClusterConfig { + BanksyClusterConfig::{ + n_clusters: 5, + n_starts: 10, + max_iterations: 100, + tolerance: 1.0e-6, + seed: 1, + } +} + +///| +pub struct BanksyClusterResult { + labels : Array[Int] + centroids : Array[Array[Double]] + cluster_sizes : Array[Int] + inertia : Double + iterations : Int + converged : Bool + silhouette : Double +} derive(Debug) + +///| +pub struct BanksySmoothingResult { + labels : Array[Int] + iterations : Int + changed_labels : Int + converged : Bool +} derive(Eq, Debug) + +///| +pub struct BanksyRun { + lambda : Double + matrix : BanksyMatrix + pca : BanksyPcaResult + clustering : BanksyClusterResult + smoothing : BanksySmoothingResult? + spatial_agreement : Double +} derive(Debug) + +///| +pub struct BanksySpatialExperimentOutput { + experiment : SpatialExperiment + model : BanksyModel + run : BanksyRun +} + +///| +priv struct BanksyDistanceIndex { + index : Int + distance : Double +} + +///| +priv struct BanksyKmeansFit { + labels : Array[Int] + centroids : Array[Array[Double]] + inertia : Double + iterations : Int + converged : Bool +} + +///| +fn banksy_is_finite(value : Double) -> Bool { + value == value && value.abs() <= 1.0e300 +} + +///| +fn banksy_zero_matrix(rows : Int, columns : Int) -> Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for _ in 0.. Array[Array[Double]] { + let output : Array[Array[Double]] = [] + for row in matrix { + output.push(row.copy()) + } + output +} + +///| +fn banksy_validate_names( + names : Array[String], + expected : Int, + prefix : String, + label : String, +) -> Array[String] raise BanksyError { + if names.length() == 0 { + let generated : Array[String] = [] + for index in 0.. (Int, Int, Array[String], Array[String]) raise BanksyError { + if expression.length() == 0 { + raise BanksyError("BANKSY expression must contain at least one gene") + } + let spots = expression[0].length() + if spots < 2 { + raise BanksyError("BANKSY expression must contain at least two spots") + } + for row in expression { + if row.length() != spots { + raise BanksyError("BANKSY expression must be rectangular") + } + for value in row { + if !banksy_is_finite(value) { + raise BanksyError("BANKSY expression values must be finite") + } + } + } + if coordinates.length() != spots { + raise BanksyError("BANKSY coordinates must contain one row per spot") + } + let dimensions = coordinates[0].length() + if dimensions < 2 { + raise BanksyError("BANKSY coordinates require at least two dimensions") + } + for coordinate in coordinates { + if coordinate.length() != dimensions { + raise BanksyError("BANKSY coordinates must be rectangular") + } + for value in coordinate { + if !banksy_is_finite(value) { + raise BanksyError("BANKSY coordinates must be finite") + } + } + } + let genes = expression.length() + let checked_genes = banksy_validate_names(gene_names, genes, "gene_", "gene") + let checked_spots = banksy_validate_names(spot_names, spots, "spot_", "spot") + (genes, spots, checked_genes, checked_spots) +} + +///| +fn banksy_distance( + coordinates : Array[Array[Double]], + first : Int, + second : Int, +) -> Double { + let mut squared = 0.0 + for dimension in 0.. Double { + let dx = coordinates[to][0] - coordinates[from][0] + let dy = coordinates[to][1] - coordinates[from][1] + let angle = @math.atan2(dy, dx) + if angle < 0.0 { + angle + 2.0 * 3.141592653589793 + } else { + angle + } +} + +///| +fn banksy_sorted_distances( + coordinates : Array[Array[Double]], + source : Int, +) -> Array[BanksyDistanceIndex] { + let distances : Array[BanksyDistanceIndex] = [] + for target in 0.. Int { + if left.distance < right.distance { + -1 + } else if left.distance > right.distance { + 1 + } else if left.index < right.index { + -1 + } else if left.index > right.index { + 1 + } else { + 0 + } + }) + distances +} + +///| +fn banksy_median(values : Array[Double]) -> Double { + if values.length() == 0 { + return 0.0 + } + let sorted = values.copy() + sorted.sort() + let middle = sorted.length() / 2 + if sorted.length() % 2 == 1 { + sorted[middle] + } else { + (sorted[middle - 1] + sorted[middle]) / 2.0 + } +} + +///| +fn banksy_sample_order( + length : Int, + sample_size : Int, + seed : Int, + source : Int, +) -> Array[Int] { + let order : Array[Int] = [] + for index in 0..= length { + return order + } + let scores : Array[Int] = Array::make(length, 0) + for index in 0.. Int { + if scores[left] < scores[right] { + -1 + } else if scores[left] > scores[right] { + 1 + } else if left < right { + -1 + } else if left > right { + 1 + } else { + 0 + } + }) + let selected : Array[Int] = [] + for index in 0.. Array[BanksyNeighbor] { + let mut total = 0.0 + for neighbor in neighbors { + total = total + neighbor.weight + } + if total <= 0.0 { + return neighbors + } + let normalized : Array[BanksyNeighbor] = [] + for neighbor in neighbors { + normalized.push(BanksyNeighbor::{ + ..neighbor, + weight: neighbor.weight / total, + }) + } + normalized +} + +///| +fn banksy_global_nearest_median(coordinates : Array[Array[Double]]) -> Double { + let nearest : Array[Double] = [] + for source in 0.. 0 { + nearest.push(sorted[0].distance) + } + } + banksy_median(nearest).max(1.0e-12) +} + +///| +fn banksy_neighbor_graph( + coordinates : Array[Array[Double]], + config : BanksyComputeConfig, + harmonic : Int, +) -> Array[Array[BanksyNeighbor]] raise BanksyError { + let spots = coordinates.length() + let k_geom = if config.k_geom.length() == 1 { + config.k_geom[0] + } else { + config.k_geom[harmonic] + } + let requested = match config.spatial_mode { + BanksyRadiusGaussian => config.k_spatial + _ => k_geom + } + if requested >= spots { + raise BanksyError( + "BANKSY neighbor count must be smaller than the number of spots", + ) + } + if config.sample_size > requested { + raise BanksyError( + "BANKSY sample size must not exceed the available neighbor count", + ) + } + let nearest_median = if config.spatial_mode is BanksyRadiusGaussian { + banksy_global_nearest_median(coordinates) + } else { + 1.0 + } + let radius = config.sigma * + (-coordinates[0].length().to_double() * @math.ln(config.alpha)).sqrt() * + nearest_median + let graph : Array[Array[BanksyNeighbor]] = [] + for source in 0..= radius { + continue + } + let safe_distance = candidate.distance.max(1.0e-12) + let weight = match config.spatial_mode { + BanksyKnnMedian => + @math.exp( + -(candidate.distance * candidate.distance) / + (local_median * local_median), + ) + BanksyKnnInverseDistance => 1.0 / safe_distance + BanksyKnnInversePower => + 1.0 / @math.pow(safe_distance, config.distance_power) + BanksyKnnRank => { + let rank_value = (rank + 1).to_double() + let width = requested.to_double() / 1.5 + @math.exp(-(rank_value * rank_value) / (2.0 * width * width)) + } + BanksyKnnUniform => 1.0 + BanksyRadiusGaussian => + @math.exp( + -0.5 * + (candidate.distance / (nearest_median * config.sigma)) * + (candidate.distance / (nearest_median * config.sigma)), + ) + } + raw.push(BanksyNeighbor::{ + index: candidate.index, + distance: candidate.distance, + weight, + angle: banksy_angle(coordinates, source, candidate.index), + }) + } + if raw.length() == 0 { + let candidate = sorted[0] + raw.push(BanksyNeighbor::{ + index: candidate.index, + distance: candidate.distance, + weight: 1.0, + angle: banksy_angle(coordinates, source, candidate.index), + }) + } + let normalized = banksy_normalize_neighbors(raw) + if config.sample_size > 0 && config.sample_size < normalized.length() { + let order = banksy_sample_order( + normalized.length(), + config.sample_size, + config.seed + harmonic * 7919, + source, + ) + let sampled : Array[BanksyNeighbor] = [] + for index in order { + sampled.push(normalized[index]) + } + if config.sample_renormalize { + graph.push(banksy_normalize_neighbors(sampled)) + } else { + graph.push(sampled) + } + } else { + graph.push(normalized) + } + } + graph +} + +///| +fn banksy_compute_harmonic( + expression : Array[Array[Double]], + graph : Array[Array[BanksyNeighbor]], + harmonic : Int, + center : Bool, +) -> Array[Array[Double]] { + let genes = expression.length() + let spots = graph.length() + let output = banksy_zero_matrix(genes, spots) + for gene in 0.. 0 && center { + for neighbor in neighborhood { + local_mean = local_mean + expression[gene][neighbor.index] + } + local_mean = local_mean / neighborhood.length().to_double() + } + if harmonic == 0 { + let mut value = 0.0 + for neighbor in neighborhood { + value = value + neighbor.weight * expression[gene][neighbor.index] + } + output[gene][spot] = value + } else { + let mut real = 0.0 + let mut imaginary = 0.0 + for neighbor in neighborhood { + let value = expression[gene][neighbor.index] - local_mean + let phase = harmonic.to_double() * neighbor.angle + real = real + neighbor.weight * value * @math.cos(phase) + imaginary = imaginary + neighbor.weight * value * @math.sin(phase) + } + output[gene][spot] = (real * real + imaginary * imaginary).sqrt() + } + } + } + output +} + +///| +pub fn banksy_compute_with_config( + expression : Array[Array[Double]], + coordinates : Array[Array[Double]], + config : BanksyComputeConfig, + gene_names? : Array[String] = [], + spot_names? : Array[String] = [], +) -> BanksyModel raise BanksyError { + let (_, _, checked_genes, checked_spots) = banksy_validate_input( + expression, coordinates, gene_names, spot_names, + ) + let harmonics : Array[Array[Array[Double]]] = [] + let graphs : Array[Array[Array[BanksyNeighbor]]] = [] + for harmonic in 0..<=config.max_harmonic { + let graph = banksy_neighbor_graph(coordinates, config, harmonic) + let values = banksy_compute_harmonic( + expression, + graph, + harmonic, + harmonic > 0 && config.center_harmonics, + ) + graphs.push(graph) + harmonics.push(values) + } + BanksyModel::{ + expression: banksy_copy_matrix(expression), + coordinates: banksy_copy_matrix(coordinates), + gene_names: checked_genes, + spot_names: checked_spots, + harmonics, + neighbors: graphs, + config, + } +} + +///| +pub fn banksy_compute( + expression : Array[Array[Double]], + coordinates : Array[Array[Double]], + k_geom? : Int = 15, + compute_agf? : Bool = false, + gene_names? : Array[String] = [], + spot_names? : Array[String] = [], +) -> BanksyModel raise BanksyError { + let config = BanksyComputeConfig::create( + max_harmonic=if compute_agf { 1 } else { 0 }, + k_geom=[k_geom], + ) + banksy_compute_with_config( + expression, + coordinates, + config, + gene_names~, + spot_names~, + ) +} + +///| +pub fn BanksyModel::n_genes(self : BanksyModel) -> Int { + self.expression.length() +} + +///| +pub fn BanksyModel::n_spots(self : BanksyModel) -> Int { + if self.expression.length() == 0 { + 0 + } else { + self.expression[0].length() + } +} + +///| +pub fn BanksyModel::max_harmonic(self : BanksyModel) -> Int { + self.harmonics.length() - 1 +} + +///| +pub fn BanksyModel::harmonic( + self : BanksyModel, + order : Int, +) -> Array[Array[Double]]? { + if order < 0 || order >= self.harmonics.length() { + None + } else { + Some(banksy_copy_matrix(self.harmonics[order])) + } +} + +///| +pub fn BanksyModel::summary(self : BanksyModel) -> String { + "BANKSY model: " + + self.n_genes().to_string() + + " genes, " + + self.n_spots().to_string() + + " spots, H0-H" + + self.max_harmonic().to_string() +} + +///| +fn banksy_validate_groups( + groups : Array[String], + spots : Int, +) -> Array[String] raise BanksyError { + if groups.length() == 0 { + return [] + } + if groups.length() != spots { + raise BanksyError("BANKSY scaling groups must contain one label per spot") + } + for group in groups { + if group.length() == 0 { + raise BanksyError("BANKSY scaling group labels must not be empty") + } + } + groups.copy() +} + +///| +fn banksy_scale_component( + matrix : Array[Array[Double]], + groups : Array[String], +) -> Array[Array[Double]] { + let output = banksy_copy_matrix(matrix) + if matrix.length() == 0 { + return output + } + let spots = matrix[0].length() + if groups.length() == 0 { + for row in 0.. 1 { + (sum_square / (spots - 1).to_double()).sqrt() + } else { + 0.0 + } + for spot in 0.. 1.0e-14 { + (matrix[row][spot] - mean) / standard_deviation + } else { + 0.0 + } + } + } + return output + } + let unique : Array[String] = [] + let seen : Map[String, Bool] = Map([]) + for group in groups { + if !seen.contains(group) { + seen[group] = true + unique.push(group) + } + } + for row in 0.. 1 { + (sum_square / (indices.length() - 1).to_double()).sqrt() + } else { + 0.0 + } + for spot in indices { + output[row][spot] = if standard_deviation > 1.0e-14 { + (matrix[row][spot] - mean) / standard_deviation + } else { + 0.0 + } + } + } + } + output +} + +///| +pub fn banksy_component_weights( + lambda : Double, + max_harmonic : Int, +) -> Array[Double] raise BanksyError { + if !banksy_is_finite(lambda) || lambda < 0.0 || lambda > 1.0 { + raise BanksyError("BANKSY lambda must be finite and in [0, 1]") + } + if max_harmonic < 0 { + raise BanksyError("BANKSY maximum harmonic must be non-negative") + } + let raw : Array[Double] = [] + let mut total = 0.0 + for harmonic in 0..<=max_harmonic { + let value = @math.pow(2.0, -harmonic.to_double()) + raw.push(value) + total = total + value + } + let weights : Array[Double] = [(1.0 - lambda).sqrt()] + for value in raw { + weights.push((lambda * value / total).sqrt()) + } + weights +} + +///| +pub fn banksy_get_matrix( + model : BanksyModel, + lambda? : Double = 0.2, + max_harmonic? : Int = -1, + scale? : Bool = false, + groups? : Array[String] = [], +) -> BanksyMatrix raise BanksyError { + let maximum = if max_harmonic < 0 { + model.max_harmonic() + } else { + max_harmonic + } + if maximum < 0 || maximum > model.max_harmonic() { + raise BanksyError("BANKSY requested harmonic was not computed") + } + let checked_groups = banksy_validate_groups(groups, model.n_spots()) + let weights = banksy_component_weights(lambda, maximum) + let values : Array[Array[Double]] = [] + let feature_names : Array[String] = [] + let own = if scale { + banksy_scale_component(model.expression, checked_groups) + } else { + banksy_copy_matrix(model.expression) + } + for gene in 0.. Int { + self.values.length() +} + +///| +pub fn BanksyMatrix::n_spots(self : BanksyMatrix) -> Int { + if self.values.length() == 0 { + 0 + } else { + self.values[0].length() + } +} + +///| +pub fn BanksyMatrix::spot_matrix(self : BanksyMatrix) -> Array[Array[Double]] { + let output = banksy_zero_matrix(self.n_spots(), self.n_features()) + for feature in 0.. Array[Array[Double]] { + if data.length() == 0 { + return [] + } + let output = banksy_copy_matrix(data) + for column in 0.. Array[Array[Double]] { + let matrix = banksy_zero_matrix(size, size) + for index in 0.. (Array[Double], Array[Array[Double]]) { + let size = input.length() + let matrix = banksy_copy_matrix(input) + let vectors = banksy_identity(size) + let limit = (size * size * 30).max(60) + for _ in 0.. 1 { 1 } else { 0 } + let mut maximum = 0.0 + for row in 0.. maximum { + maximum = value + p = row + q = column + } + } + } + if maximum < 1.0e-12 || size < 2 { + break + } + let app = matrix[p][p] + let aqq = matrix[q][q] + let apq = matrix[p][q] + let tau = (aqq - app) / (2.0 * apq) + let t = if tau >= 0.0 { + 1.0 / (tau + (1.0 + tau * tau).sqrt()) + } else { + -1.0 / (-tau + (1.0 + tau * tau).sqrt()) + } + let cosine = 1.0 / (1.0 + t * t).sqrt() + let sine = t * cosine + for index in 0.. Int { + if values[left] > values[right] { + -1 + } else if values[left] < values[right] { + 1 + } else if left < right { + -1 + } else if left > right { + 1 + } else { + 0 + } + }) + let sorted_values : Array[Double] = [] + let sorted_vectors = banksy_zero_matrix(size, size) + for column in 0.. BanksyPcaResult raise BanksyError { + let spots = matrix.n_spots() + let features = matrix.n_features() + if spots < 2 || features < 1 { + raise BanksyError("BANKSY PCA requires at least two spots and one feature") + } + let maximum = (spots - 1).min(features) + if n_components < 1 || n_components > maximum { + raise BanksyError( + "BANKSY PCA component count must be in [1, min(spots - 1, features)]", + ) + } + let centered = banksy_center_columns(matrix.spot_matrix()) + let gram = banksy_zero_matrix(spots, spots) + for first in 0.. 0.0 { + total_variance = total_variance + value + } + } + if total_variance <= 1.0e-14 { + raise BanksyError("BANKSY PCA input has no non-constant variation") + } + let component_count = n_components + let scores = banksy_zero_matrix(spots, component_count) + let loadings = banksy_zero_matrix(features, component_count) + let eigenvalues : Array[Double] = [] + let variance_explained : Array[Double] = [] + for component in 0.. 1.0e-14 { + for feature in 0.. loadings[largest][component].abs() { + largest = feature + } + } + if loadings[largest][component] < 0.0 { + for spot in 0.. Int { + self.scores.length() +} + +///| +pub fn BanksyPcaResult::summary(self : BanksyPcaResult) -> String { + let mut cumulative = 0.0 + for value in self.variance_explained { + cumulative = cumulative + value + } + "BANKSY PCA: " + + self.n_components.to_string() + + " components, explained variance=" + + cumulative.to_string() +} + +///| +fn banksy_squared_distance( + left : Array[Double], + right : Array[Double], +) -> Double { + let mut value = 0.0 + for index in 0.. Array[Array[Double]] { + let centroids : Array[Array[Double]] = [data[first_index].copy()] + let selected = Array::make(data.length(), false) + selected[first_index] = true + while centroids.length() < clusters { + let mut best_index = -1 + let mut best_distance = -1.0 + for point in 0.. best_distance { + best_distance = minimum + best_index = point + } + } + if best_index < 0 { + break + } + selected[best_index] = true + centroids.push(data[best_index].copy()) + } + centroids +} + +///| +fn banksy_assign_clusters( + data : Array[Array[Double]], + centroids : Array[Array[Double]], +) -> (Array[Int], Double) { + let labels = Array::make(data.length(), 0) + let mut inertia = 0.0 + for point in 0.. Array[Array[Double]] { + let clusters = old_centroids.length() + let dimensions = data[0].length() + let centroids = banksy_zero_matrix(clusters, dimensions) + let counts = Array::make(clusters, 0) + for point in 0.. 0 { + for dimension in 0.. farthest_distance { + farthest = point + farthest_distance = distance + } + } + centroids[cluster] = data[farthest].copy() + } + } + centroids +} + +///| +fn banksy_kmeans_once( + data : Array[Array[Double]], + config : BanksyClusterConfig, + start : Int, +) -> BanksyKmeansFit { + let first = (config.seed % data.length() + start) % data.length() + let mut centroids = banksy_initialize_centroids( + data, + config.n_clusters, + first, + ) + let mut labels = Array::make(data.length(), -1) + let mut iterations = 0 + let mut converged = false + for iteration in 0.. Double raise BanksyError { + if data.length() != labels.length() { + raise BanksyError("BANKSY silhouette labels must match data rows") + } + if data.length() < 2 { + return 0.0 + } + let mut maximum_label = -1 + for label in labels { + if label < 0 { + raise BanksyError("BANKSY cluster labels must be non-negative") + } + maximum_label = maximum_label.max(label) + } + let clusters = maximum_label + 1 + let mut total = 0.0 + for point in 0.. 0 { + b = b.min(other_sums[cluster] / other_counts[cluster].to_double()) + } + } + if b < 1.0e299 { + let denominator = a.max(b) + if denominator > 0.0 { + total = total + (b - a) / denominator + } + } + } + total / data.length().to_double() +} + +///| +pub fn banksy_cluster( + pca : BanksyPcaResult, + config? : BanksyClusterConfig = BanksyClusterConfig::default(), +) -> BanksyClusterResult raise BanksyError { + if pca.scores.length() == 0 { + raise BanksyError("BANKSY clustering requires non-empty PCA scores") + } + if config.n_clusters > pca.scores.length() { + raise BanksyError("BANKSY cluster count must not exceed the spot count") + } + let mut best : BanksyKmeansFit? = None + for start in 0.. Some(fit) + Some(current) => + if fit.inertia < current.inertia { + Some(fit) + } else { + Some(current) + } + } + } + let fit = match best { + Some(value) => value + None => raise BanksyError("BANKSY k-means did not produce a fit") + } + let sizes = Array::make(config.n_clusters, 0) + for label in fit.labels { + sizes[label] = sizes[label] + 1 + } + let silhouette = banksy_silhouette(pca.scores, fit.labels) + BanksyClusterResult::{ + labels: fit.labels, + centroids: fit.centroids, + cluster_sizes: sizes, + inertia: fit.inertia, + iterations: fit.iterations, + converged: fit.converged, + silhouette, + } +} + +///| +pub fn BanksyClusterResult::n_clusters(self : BanksyClusterResult) -> Int { + self.cluster_sizes.length() +} + +///| +pub fn BanksyClusterResult::summary(self : BanksyClusterResult) -> String { + "BANKSY clustering: " + + self.n_clusters().to_string() + + " clusters, inertia=" + + self.inertia.to_string() + + ", silhouette=" + + self.silhouette.to_string() +} + +///| +fn banksy_plain_knn( + coordinates : Array[Array[Double]], + k : Int, +) -> Array[Array[Int]] raise BanksyError { + if k < 1 || k >= coordinates.length() { + raise BanksyError( + "BANKSY smoothing neighbor count must be in [1, spots - 1]", + ) + } + let output : Array[Array[Int]] = [] + for source in 0.. BanksySmoothingResult raise BanksyError { + if labels.length() != coordinates.length() || labels.length() < 2 { + raise BanksyError( + "BANKSY smoothing labels and coordinates must contain the same spots", + ) + } + if !banksy_is_finite(proportion_threshold) || + proportion_threshold < 0.0 || + proportion_threshold > 1.0 { + raise BanksyError("BANKSY smoothing proportion threshold must be in [0, 1]") + } + if max_iterations == 0 || max_iterations < -1 { + raise BanksyError( + "BANKSY smoothing iterations must be positive or -1 for convergence", + ) + } + let mut maximum_label = -1 + for label in labels { + if label < 0 { + raise BanksyError("BANKSY smoothing labels must be non-negative") + } + maximum_label = maximum_label.max(label) + } + let neighbors = banksy_plain_knn(coordinates, k) + let raw = labels.copy() + let mut current = labels.copy() + let limit = if max_iterations == -1 { 10000 } else { max_iterations } + let mut iterations = 0 + let mut converged = false + for iteration in 0.. counts[best] { + best = label + } + } + let proportion = counts[best].to_double() / (k + 1).to_double() + if proportion > proportion_threshold { + next[spot] = best + } + if next[spot] != current[spot] { + changes = changes + 1 + } + } + current = next + iterations = iteration + 1 + if changes == 0 { + converged = true + break + } + } + let mut total_changed = 0 + for spot in 0.. Double raise BanksyError { + if labels.length() != coordinates.length() || labels.length() < 2 { + raise BanksyError("BANKSY spatial agreement labels must match coordinates") + } + let neighbors = banksy_plain_knn(coordinates, k) + let mut matches = 0 + let mut total = 0 + for spot in 0.. Double raise BanksyError { + if left.length() != right.length() || left.length() < 2 { + raise BanksyError( + "BANKSY ARI label vectors must have equal length of at least two", + ) + } + let left_counts : Map[Int, Int] = Map([]) + let right_counts : Map[Int, Int] = Map([]) + let cells : Map[String, Int] = Map([]) + for index in 0.. value + 1 + None => 1 + } + right_counts[right[index]] = match right_counts.get(right[index]) { + Some(value) => value + 1 + None => 1 + } + let key = left[index].to_string() + ":" + right[index].to_string() + cells[key] = match cells.get(key) { + Some(value) => value + 1 + None => 1 + } + } + let choose_two = fn(value : Int) -> Double { + (value * (value - 1) / 2).to_double() + } + let mut cell_sum = 0.0 + for key in cells.keys() { + cell_sum = cell_sum + choose_two(cells[key]) + } + let mut left_sum = 0.0 + for key in left_counts.keys() { + left_sum = left_sum + choose_two(left_counts[key]) + } + let mut right_sum = 0.0 + for key in right_counts.keys() { + right_sum = right_sum + choose_two(right_counts[key]) + } + let total_pairs = choose_two(left.length()) + let expected = left_sum * right_sum / total_pairs + let maximum = (left_sum + right_sum) / 2.0 + if (maximum - expected).abs() <= 1.0e-14 { + if (cell_sum - expected).abs() <= 1.0e-14 { + 1.0 + } else { + 0.0 + } + } else { + (cell_sum - expected) / (maximum - expected) + } +} + +///| +pub fn banksy_run( + model : BanksyModel, + lambda? : Double = 0.2, + n_components? : Int = 20, + cluster_config? : BanksyClusterConfig = BanksyClusterConfig::default(), + max_harmonic? : Int = -1, + scale? : Bool = true, + groups? : Array[String] = [], + smooth? : Bool = false, + smoothing_k? : Int = 15, + smoothing_threshold? : Double = 0.5, + smoothing_iterations? : Int = 10, +) -> BanksyRun raise BanksyError { + let matrix = banksy_get_matrix(model, lambda~, max_harmonic~, scale~, groups~) + let pca = banksy_run_pca(matrix, n_components~) + let clustering = banksy_cluster(pca, config=cluster_config) + let smoothing = if smooth { + Some( + banksy_smooth_labels( + clustering.labels, + model.coordinates, + k=smoothing_k, + proportion_threshold=smoothing_threshold, + max_iterations=smoothing_iterations, + ), + ) + } else { + None + } + let final_labels = match smoothing { + Some(result) => result.labels + None => clustering.labels + } + let agreement_k = smoothing_k.min(model.n_spots() - 1).max(1) + let agreement = banksy_spatial_agreement( + final_labels, + model.coordinates, + k=agreement_k, + ) + BanksyRun::{ + lambda, + matrix, + pca, + clustering, + smoothing, + spatial_agreement: agreement, + } +} + +///| +pub fn banksy_parameter_sweep( + model : BanksyModel, + lambdas : Array[Double], + n_components? : Int = 20, + cluster_config? : BanksyClusterConfig = BanksyClusterConfig::default(), + max_harmonic? : Int = -1, + scale? : Bool = true, + groups? : Array[String] = [], +) -> Array[BanksyRun] raise BanksyError { + if lambdas.length() == 0 { + raise BanksyError("BANKSY parameter sweep requires at least one lambda") + } + let runs : Array[BanksyRun] = [] + for lambda in lambdas { + runs.push( + banksy_run( + model, + lambda~, + n_components~, + cluster_config~, + max_harmonic~, + scale~, + groups~, + ), + ) + } + runs +} + +///| +pub fn BanksyRun::final_labels(self : BanksyRun) -> Array[Int] { + match self.smoothing { + Some(result) => result.labels.copy() + None => self.clustering.labels.copy() + } +} + +///| +pub fn BanksyRun::summary(self : BanksyRun) -> String { + "BANKSY run: lambda=" + + self.lambda.to_string() + + ", " + + self.clustering.summary() + + ", spatial agreement=" + + self.spatial_agreement.to_string() +} + +///| +fn banksy_copy_string_map(source : Map[String, String]) -> Map[String, String] { + let copied : Map[String, String] = Map([]) + for key in source.keys() { + copied[key] = source[key] + } + copied +} + +///| +fn banksy_copy_spatial_experiment( + experiment : SpatialExperiment, +) -> SpatialExperiment { + let copied = SpatialExperiment::new() + for key in experiment.assay.keys() { + copied.assay[key] = banksy_copy_matrix(experiment.assay[key]) + } + for row in experiment.row_data { + copied.row_data.push(banksy_copy_string_map(row)) + } + for column in experiment.col_data { + copied.col_data.push(banksy_copy_string_map(column)) + } + for coordinate in experiment.spatial_coords { + copied.spatial_coords.push(coordinate) + } + for image in experiment.images { + copied.images.push( + SpatialImage::new( + image.id, + banksy_copy_matrix(image.data), + image.scale_factor, + ), + ) + } + for key in experiment.metadata.keys() { + copied.metadata[key] = experiment.metadata[key] + } + copied +} + +///| +fn banksy_experiment_gene_names( + experiment : SpatialExperiment, + genes : Int, +) -> Array[String] raise BanksyError { + if experiment.row_data.length() != 0 && experiment.row_data.length() != genes { + raise BanksyError("BANKSY SpatialExperiment rowData must match assay rows") + } + let names : Array[String] = [] + for gene in 0.. + if value.length() > 0 { + value + } else { + match row.get("gene_id") { + Some(identifier) => + if identifier.length() > 0 { + identifier + } else { + generated + } + None => generated + } + } + None => + match row.get("gene_id") { + Some(identifier) => + if identifier.length() > 0 { + identifier + } else { + generated + } + None => generated + } + } + names.push(name) + } + } + names +} + +///| +fn banksy_experiment_spot_names( + experiment : SpatialExperiment, + spots : Int, +) -> Array[String] raise BanksyError { + if experiment.col_data.length() != 0 && experiment.col_data.length() != spots { + raise BanksyError( + "BANKSY SpatialExperiment colData must match assay columns", + ) + } + let names : Array[String] = [] + for spot in 0.. if value.length() > 0 { value } else { generated } + None => generated + } + names.push(name) + } + } + names +} + +///| +fn banksy_experiment_coordinates( + experiment : SpatialExperiment, + spots : Int, +) -> Array[Array[Double]] raise BanksyError { + if experiment.spatial_coords.length() != spots { + raise BanksyError( + "BANKSY SpatialExperiment requires one spatial coordinate per spot", + ) + } + let mut use_z = false + if spots > 0 { + let first = experiment.spatial_coords[0].z + for coordinate in experiment.spatial_coords { + if (coordinate.z - first).abs() > 1.0e-14 { + use_z = true + } + } + } + let coordinates : Array[Array[Double]] = [] + for coordinate in experiment.spatial_coords { + if use_z { + coordinates.push([coordinate.x, coordinate.y, coordinate.z]) + } else { + coordinates.push([coordinate.x, coordinate.y]) + } + } + coordinates +} + +///| +pub fn banksy_spatial_experiment( + experiment : SpatialExperiment, + assay_name? : String = "logcounts", + config? : BanksyComputeConfig = BanksyComputeConfig::default(), + lambda? : Double = 0.2, + n_components? : Int = 20, + cluster_config? : BanksyClusterConfig = BanksyClusterConfig::default(), + max_harmonic? : Int = -1, + scale? : Bool = true, + group_key? : String = "", + smooth? : Bool = true, + smoothing_k? : Int = 15, + smoothing_threshold? : Double = 0.5, + smoothing_iterations? : Int = 10, +) -> BanksySpatialExperimentOutput raise BanksyError { + let expression = match experiment.assay.get(assay_name) { + Some(value) => value + None => + raise BanksyError( + "BANKSY SpatialExperiment assay not found: " + assay_name, + ) + } + if expression.length() == 0 { + raise BanksyError("BANKSY SpatialExperiment assay must not be empty") + } + let spots = expression[0].length() + let gene_names = banksy_experiment_gene_names(experiment, expression.length()) + let spot_names = banksy_experiment_spot_names(experiment, spots) + let coordinates = banksy_experiment_coordinates(experiment, spots) + let groups : Array[String] = [] + if group_key.length() > 0 { + if experiment.col_data.length() != spots { + raise BanksyError( + "BANKSY grouping requires complete SpatialExperiment colData", + ) + } + for column in experiment.col_data { + match column.get(group_key) { + Some(value) => + if value.length() == 0 { + raise BanksyError( + "BANKSY SpatialExperiment grouping values must not be empty", + ) + } else { + groups.push(value) + } + None => + raise BanksyError( + "BANKSY SpatialExperiment group field not found: " + group_key, + ) + } + } + } + let model = banksy_compute_with_config( + expression, + coordinates, + config, + gene_names~, + spot_names~, + ) + let run = banksy_run( + model, + lambda~, + n_components~, + cluster_config~, + max_harmonic~, + scale~, + groups~, + smooth~, + smoothing_k~, + smoothing_threshold~, + smoothing_iterations~, + ) + let enriched = banksy_copy_spatial_experiment(experiment) + if enriched.row_data.length() == 0 { + for gene_name in gene_names { + enriched.row_data.push(Map([("gene_name", gene_name)])) + } + } + if enriched.col_data.length() == 0 { + for spot_name in spot_names { + enriched.col_data.push(Map([("spot_id", spot_name)])) + } + } + for harmonic in 0..<=model.max_harmonic() { + enriched.assay["H" + harmonic.to_string()] = banksy_copy_matrix( + model.harmonics[harmonic], + ) + } + let labels = run.final_labels() + for spot in 0.. ( + Array[Array[Double]], + Array[Array[Double]], + Array[String], + Array[String], +) { + let width = 6 + let height = 4 + let spots = width * height + let coordinates : Array[Array[Double]] = [] + let spot_names : Array[String] = [] + let expression = banksy_zero_matrix(5, spots) + for y in 0..= width / 2 { 8.0 } else { 1.0 } + let upper = if y < height / 2 { 6.0 } else { 2.0 } + let boundary = if x == 2 || x == 3 { 7.0 } else { 1.0 } + let variation = ((x * 3 + y * 5) % 4).to_double() * 0.15 + expression[0][spot] = left + variation + expression[1][spot] = right + variation + expression[2][spot] = upper + variation + expression[3][spot] = boundary + variation + expression[4][spot] = 3.0 + variation + } + } + ( + expression, + coordinates, + ["left_marker", "right_marker", "upper_marker", "boundary", "baseline"], + spot_names, + ) +} diff --git a/src/batchelor.mbt b/src/batchelor.mbt index e8dc839b..a13a61be 100644 --- a/src/batchelor.mbt +++ b/src/batchelor.mbt @@ -18,12 +18,7 @@ pub fn BatchCorrectionResult::new( mnn_pairs : Array[(Int, Int)], var_explained : Array[Double], ) -> BatchCorrectionResult { - BatchCorrectionResult::{ - corrected, - batch_indices, - mnn_pairs, - var_explained, - } + BatchCorrectionResult::{ corrected, batch_indices, mnn_pairs, var_explained } } ///| @@ -307,7 +302,10 @@ pub fn fast_mnn( nn = nn + 1 } - let cov_matrix : Array[Array[Double]] = Array::make(n_genes, Array::make(n_genes, 0.0)) + let cov_matrix : Array[Array[Double]] = Array::make( + n_genes, + Array::make(n_genes, 0.0), + ) let mut pp = 0 while pp < n_cells { let cell = centered[pp] @@ -333,7 +331,10 @@ pub fn fast_mnn( } let eigenvalues : Array[Double] = Array::make(n_genes, 0.0) - let eigenvectors : Array[Array[Double]] = Array::make(n_genes, Array::make(n_genes, 0.0)) + let eigenvectors : Array[Array[Double]] = Array::make( + n_genes, + Array::make(n_genes, 0.0), + ) let mut uu = 0 while uu < n_genes { eigenvalues[uu] = cov_matrix[uu][uu] @@ -350,11 +351,7 @@ pub fn fast_mnn( while xx < n_genes { let p = cov_matrix[ww][xx] let d_val = eigenvalues[xx] - eigenvalues[ww] - let c = if d_val.abs() < 0.000001 { - 0.000001 - } else { - d_val - } + let c = if d_val.abs() < 0.000001 { 0.000001 } else { d_val } let t = p / c let cos_val = 1.0 / (1.0 + t * t).sqrt() let sin_val = t * cos_val @@ -480,7 +477,10 @@ pub fn fast_mnn( kk = kk + 1 } - let correction_vectors : Array[Array[Double]] = Array::make(n_cells, Array::make(top_components, 0.0)) + let correction_vectors : Array[Array[Double]] = Array::make( + n_cells, + Array::make(top_components, 0.0), + ) let mut mm = 0 while mm < n_batches - 1 { @@ -542,7 +542,9 @@ pub fn fast_mnn( let mut dist = 0.0 let mut uu = 0 while uu < top_components { - dist = dist + (pca_result[ss][uu] - pca_result[mnn_idx][uu]) * (pca_result[ss][uu] - pca_result[mnn_idx][uu]) + dist = dist + + (pca_result[ss][uu] - pca_result[mnn_idx][uu]) * + (pca_result[ss][uu] - pca_result[mnn_idx][uu]) uu = uu + 1 } let weight = @math.exp(-dist / (2.0 * sigma * sigma)) @@ -559,7 +561,8 @@ pub fn fast_mnn( if weight_sum > 0.0 { let mut ww = 0 while ww < top_components { - correction_vectors[ss][ww] = correction_vectors[ss][ww] + weighted_delta[ww] / weight_sum + correction_vectors[ss][ww] = correction_vectors[ss][ww] + + weighted_delta[ww] / weight_sum ww = ww + 1 } } @@ -583,7 +586,9 @@ pub fn fast_mnn( xx = xx + 1 } - BatchCorrectionResult::new(corrected, batch_indices, all_mnn_pairs, var_explained) + BatchCorrectionResult::new( + corrected, batch_indices, all_mnn_pairs, var_explained, + ) } ///| @@ -619,7 +624,11 @@ pub fn batchelor_create_example_data( x = x - 6.0 let base_expression = @math.exp(x) - let batch_factor = 1.0 + i.to_double() * batch_effect * (if k < n_genes / 2 { 1.0 } else { -1.0 }) * 0.1 + let batch_factor = 1.0 + + i.to_double() * + batch_effect * + (if k < n_genes / 2 { 1.0 } else { -1.0 }) * + 0.1 cell.push(base_expression * batch_factor) k = k + 1 @@ -637,7 +646,10 @@ pub fn batchelor_create_example_data( } ///| -pub fn compute_batch_mixing_score(corrected : Array[Array[Double]], batch_indices : Array[Int]) -> Double { +pub fn compute_batch_mixing_score( + corrected : Array[Array[Double]], + batch_indices : Array[Int], +) -> Double { let n_cells = corrected.length() let n_dims = if n_cells > 0 { corrected[0].length() } else { 0 } @@ -670,7 +682,9 @@ pub fn compute_batch_mixing_score(corrected : Array[Array[Double]], batch_indice let mut dist = 0.0 let mut k = 0 while k < n_dims { - dist = dist + (corrected[i][k] - corrected[j][k]) * (corrected[i][k] - corrected[j][k]) + dist = dist + + (corrected[i][k] - corrected[j][k]) * + (corrected[i][k] - corrected[j][k]) k = k + 1 } distances.push((dist.sqrt(), j)) @@ -703,7 +717,8 @@ pub fn compute_batch_mixing_score(corrected : Array[Array[Double]], batch_indice n = n + 1 } - score = score + (1.0 - same_batch_count.to_double() / n_neighbors.to_double()) + score = score + + (1.0 - same_batch_count.to_double() / n_neighbors.to_double()) i = i + 1 } diff --git a/src/bayes_space.mbt b/src/bayes_space.mbt index 4a042441..37f3e3d6 100644 --- a/src/bayes_space.mbt +++ b/src/bayes_space.mbt @@ -21,10 +21,10 @@ /// Spatial coordinates for a single spot. pub struct SpotCoord { spot_id : String - row : Int // array row index - col : Int // array column index - x : Double // continuous x coordinate (e.g. pixel) - y : Double // continuous y coordinate + row : Int // array row index + col : Int // array column index + x : Double // continuous x coordinate (e.g. pixel) + y : Double // continuous y coordinate } ///| @@ -43,10 +43,10 @@ pub fn spot_coord( /// BayesSpace clustering result. pub struct BayesSpaceResult { spot_ids : Array[String] - clusters : Array[Int] // cluster assignment per spot (0..q-1) - q : Int // number of clusters - cluster_centers : Array[Array[Double]] // q × n_features matrix - responsibilities : Array[Array[Double]] // n_spots × q soft assignments + clusters : Array[Int] // cluster assignment per spot (0..q-1) + q : Int // number of clusters + cluster_centers : Array[Array[Double]] // q × n_features matrix + responsibilities : Array[Array[Double]] // n_spots × q soft assignments n_iterations : Int log_likelihood : Double converged : Bool @@ -59,9 +59,7 @@ pub struct BayesSpaceResult { /// For square-grid compatibility we also include (r±1, c±1) as neighbors. /// /// Returns neighbor_indices[i] = list of spot indices that are neighbors of spot i. -pub fn bayes_space_hex_neighbors( - spots : Array[SpotCoord], -) -> Array[Array[Int]] { +pub fn bayes_space_hex_neighbors(spots : Array[SpotCoord]) -> Array[Array[Int]] { let n = spots.length() // Build a map from (row, col) to spot index let coord_map : Map[String, Int] = Map::new() @@ -353,7 +351,12 @@ pub fn bayes_space_run( while i < n_spots { let mut k = 0 while k < q { - let log_t = bayes_space_log_t_density(expression[i], centers[k], scales[k], d_df) + let log_t = bayes_space_log_t_density( + expression[i], + centers[k], + scales[k], + d_df, + ) let mut spatial = 0.0 let mut ni = 0 while ni < neighbors[i].length() { @@ -454,7 +457,12 @@ pub fn bayes_space_run( let mut max_lp = -1.0e18 let mut k = 0 while k < q { - let log_t = bayes_space_log_t_density(expression[i], centers[k], scales[k], d_df) + let log_t = bayes_space_log_t_density( + expression[i], + centers[k], + scales[k], + d_df, + ) let lp = @math.ln(pis[k]) + log_t if lp > max_lp { max_lp = lp @@ -464,7 +472,12 @@ pub fn bayes_space_run( let mut sum_exp = 0.0 k = 0 while k < q { - let log_t = bayes_space_log_t_density(expression[i], centers[k], scales[k], d_df) + let log_t = bayes_space_log_t_density( + expression[i], + centers[k], + scales[k], + d_df, + ) sum_exp = sum_exp + @math.exp(@math.ln(pis[k]) + log_t - max_lp) k = k + 1 } @@ -575,10 +588,18 @@ pub fn bayes_space_render_clusters( let mut max_c = spots[0].col let mut i = 1 while i < spots.length() { - if spots[i].row < min_r { min_r = spots[i].row } - if spots[i].row > max_r { max_r = spots[i].row } - if spots[i].col < min_c { min_c = spots[i].col } - if spots[i].col > max_c { max_c = spots[i].col } + if spots[i].row < min_r { + min_r = spots[i].row + } + if spots[i].row > max_r { + max_r = spots[i].row + } + if spots[i].col < min_c { + min_c = spots[i].col + } + if spots[i].col > max_c { + max_c = spots[i].col + } i = i + 1 } let label_chars = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ" @@ -638,8 +659,20 @@ pub fn bayes_space_sample_data() -> (Array[Array[Double]], Array[SpotCoord]) { 2.0 } // 3-feature expression vector per spot - expression.push([base + noise, base * 0.8 + noise * 0.5, base * 1.2 + noise * 0.3]) - spots.push(spot_coord("spot_" + r.to_string() + "_" + c.to_string(), r, c, c.to_double(), r.to_double())) + expression.push([ + base + noise, + base * 0.8 + noise * 0.5, + base * 1.2 + noise * 0.3, + ]) + spots.push( + spot_coord( + "spot_" + r.to_string() + "_" + c.to_string(), + r, + c, + c.to_double(), + r.to_double(), + ), + ) c = c + 1 } r = r + 1 diff --git a/src/bayseq.mbt b/src/bayseq.mbt index 673a0ba0..bbd7bcf2 100644 --- a/src/bayseq.mbt +++ b/src/bayseq.mbt @@ -33,7 +33,10 @@ pub struct BayesResult { /// Estimate the dispersion parameter for a single gene using the method of moments. /// For Negative Binomial: Var = mean + mean^2 / dispersion /// => dispersion = mean^2 / (var - mean) -pub fn bayseq_estimate_dispersion_gene(counts_a : Array[Double], counts_b : Array[Double]) -> Double { +pub fn bayseq_estimate_dispersion_gene( + counts_a : Array[Double], + counts_b : Array[Double], +) -> Double { let n_a = counts_a.length() let n_b = counts_b.length() if n_a < 2 || n_b < 2 { @@ -86,12 +89,21 @@ pub fn bayseq_estimate_dispersion_gene(counts_a : Array[Double], counts_b : Arra } // Bound dispersion to reasonable range - if disp < 0.01 { 0.01 } else if disp > 100.0 { 100.0 } else { disp } + if disp < 0.01 { + 0.01 + } else if disp > 100.0 { + 100.0 + } else { + disp + } } ///| /// Estimate dispersion prior parameters (shape, rate) from all genes. -pub fn bayseq_estimate_prior(counts : Array[Array[Double]], groups : Array[Int]) -> (Double, Double) { +pub fn bayseq_estimate_prior( + counts : Array[Array[Double]], + groups : Array[Int], +) -> (Double, Double) { let n_genes = counts.length() if n_genes == 0 { return (2.0, 0.5) @@ -103,8 +115,11 @@ pub fn bayseq_estimate_prior(counts : Array[Array[Double]], groups : Array[Int]) let mut n_group_1 = 0 let mut gi = 0 while gi < groups.length() { - if groups[gi] == 0 { n_group_0 = n_group_0 + 1 } - else if groups[gi] == 1 { n_group_1 = n_group_1 + 1 } + if groups[gi] == 0 { + n_group_0 = n_group_0 + 1 + } else if groups[gi] == 1 { + n_group_1 = n_group_1 + 1 + } gi = gi + 1 } @@ -164,7 +179,9 @@ pub fn hts_count_group(groups : Array[Int], group_id : Int) -> Int { let mut count = 0 let mut i = 0 while i < groups.length() { - if groups[i] == group_id { count = count + 1 } + if groups[i] == group_id { + count = count + 1 + } i = i + 1 } count @@ -217,7 +234,11 @@ pub fn bayseq_log_likelihood_ratio( ///| /// Negative Binomial log-likelihood for a set of counts given mean and dispersion. -pub fn bayseq_nb_log_likelihood(counts : Array[Double], mean : Double, disp : Double) -> Double { +pub fn bayseq_nb_log_likelihood( + counts : Array[Double], + mean : Double, + disp : Double, +) -> Double { if mean <= 0.0 || disp <= 0.0 { return 0.0 } @@ -234,7 +255,11 @@ pub fn bayseq_nb_log_likelihood(counts : Array[Double], mean : Double, disp : Do let lgamma_x1 = bayseq_lgamma(x + 1.0) let lgamma_r = bayseq_lgamma(r) let ln_p = if p > 0.0 { @math.ln(p) } else { -10000000000.0 } - let ln_1mp = if (1.0 - p) > 0.0 { @math.ln(1.0 - p) } else { -10000000000.0 } + let ln_1mp = if 1.0 - p > 0.0 { + @math.ln(1.0 - p) + } else { + -10000000000.0 + } ll = ll + lgamma_xr - lgamma_x1 - lgamma_r + r * ln_p + x * ln_1mp } else { let lgamma_r = bayseq_lgamma(r) @@ -268,12 +293,8 @@ pub fn bayseq_lgamma(x : Double) -> Double { // Lanczos approximation for x >= 2 // Coefficients for g=5 let c = [ - 76.18009172947146, - -86.50532032941677, - 24.01409824083091, - -1.231739572450155, - 0.1208650973866179e-2, - -0.5395239384953e-5 + 76.18009172947146, -86.50532032941677, 24.01409824083091, -1.231739572450155, + 0.1208650973866179e-2, -0.5395239384953e-5, ] let y = x let mut tmp = x + 5.5 @@ -326,9 +347,15 @@ pub fn bayseq_test( let mut sum_a = 0.0 let mut sum_b = 0.0 let mut i = 0 - while i < c_a.length() { sum_a = sum_a + c_a[i]; i = i + 1 } + while i < c_a.length() { + sum_a = sum_a + c_a[i] + i = i + 1 + } i = 0 - while i < c_b.length() { sum_b = sum_b + c_b[i]; i = i + 1 } + while i < c_b.length() { + sum_b = sum_b + c_b[i] + i = i + 1 + } let mean_a = sum_a / n_group_0.to_double() let mean_b = sum_b / n_group_1.to_double() @@ -349,9 +376,15 @@ pub fn bayseq_test( let prob_de = 1.0 / (1.0 + @math.exp(-llr)) let is_de = prob_de >= alpha - if is_de { de_count = de_count + 1 } + if is_de { + de_count = de_count + 1 + } - let name = if g < gene_names.length() { gene_names[g] } else { "Gene" + (g + 1).to_string() } + let name = if g < gene_names.length() { + gene_names[g] + } else { + "Gene" + (g + 1).to_string() + } results.push(BayesGeneResult::{ gene_index: g, @@ -361,7 +394,7 @@ pub fn bayseq_test( posterior_prob_de: prob_de, map_expression_a: mean_a, map_expression_b: mean_b, - is_de: is_de, + is_de, }) g = g + 1 } @@ -391,15 +424,26 @@ pub fn bayseq_get_de_genes(result : BayesResult) -> Array[BayesGeneResult] { ///| /// Get top DE genes sorted by absolute log fold change. -pub fn bayseq_get_top_de(result : BayesResult, n : Int) -> Array[BayesGeneResult] { +pub fn bayseq_get_top_de( + result : BayesResult, + n : Int, +) -> Array[BayesGeneResult] { let all = result.gene_results // Sort by absolute log_fold_change descending (bubble sort) let mut i = 0 while i < all.length() { let mut j = i + 1 while j < all.length() { - let abs_i = if all[i].log_fold_change >= 0.0 { all[i].log_fold_change } else { -all[i].log_fold_change } - let abs_j = if all[j].log_fold_change >= 0.0 { all[j].log_fold_change } else { -all[j].log_fold_change } + let abs_i = if all[i].log_fold_change >= 0.0 { + all[i].log_fold_change + } else { + -all[i].log_fold_change + } + let abs_j = if all[j].log_fold_change >= 0.0 { + all[j].log_fold_change + } else { + -all[j].log_fold_change + } if abs_j > abs_i { let tmp = all[i] all[i] = all[j] @@ -437,7 +481,7 @@ pub fn bayseq_sample_data() -> (Array[Array[Double]], Array[Int], Array[String]) let base_a = 500.0 let base_b = if g < 8 { 1500.0 } else if g < 16 { 200.0 } else { 500.0 } let base = if groups[s] == 0 { base_a } else { base_b } - let noise = (((g * 7 + s * 17) % 100).to_double() / 100.0) * base * 0.3 + let noise = ((g * 7 + s * 17) % 100).to_double() / 100.0 * base * 0.3 let val = base + noise row.push(val) s = s + 1 @@ -453,10 +497,17 @@ pub fn bayseq_sample_data() -> (Array[Array[Double]], Array[Int], Array[String]) /// Summarize the baySeq result. pub fn bayseq_summary(result : BayesResult) -> String { "baySeq Analysis Summary:\n" + - " Total genes: " + result.gene_results.length().to_string() + "\n" + - " DE genes (posterior >= " + result.alpha_level.to_string() + "): " + - result.n_de_genes.to_string() + "\n" + - " Dispersion prior: Gamma(shape=" + - result.dispersion_prior_shape.to_string() + - ", rate=" + result.dispersion_prior_rate.to_string() + ")" + " Total genes: " + + result.gene_results.length().to_string() + + "\n" + + " DE genes (posterior >= " + + result.alpha_level.to_string() + + "): " + + result.n_de_genes.to_string() + + "\n" + + " Dispersion prior: Gamma(shape=" + + result.dispersion_prior_shape.to_string() + + ", rate=" + + result.dispersion_prior_rate.to_string() + + ")" } diff --git a/src/beachmat.mbt b/src/beachmat.mbt index 7f52927a..11dba5b8 100644 --- a/src/beachmat.mbt +++ b/src/beachmat.mbt @@ -37,25 +37,45 @@ pub struct BmatParam { ///| /// Create default parameters. pub fn BmatParam::new() -> BmatParam { - { row_block_size: 0, col_block_size: 0, cache_data: true, storage_mode: BmatStorageMode::Dense } + { + row_block_size: 0, + col_block_size: 0, + cache_data: true, + storage_mode: BmatStorageMode::Dense, + } } ///| /// Create parameters with specific block sizes. pub fn BmatParam::with_blocks(row_block : Int, col_block : Int) -> BmatParam { - { row_block_size: row_block, col_block_size: col_block, cache_data: true, storage_mode: BmatStorageMode::Dense } + { + row_block_size: row_block, + col_block_size: col_block, + cache_data: true, + storage_mode: BmatStorageMode::Dense, + } } ///| /// Create parameters for column-wise access. pub fn BmatParam::column_param(block_size : Int) -> BmatParam { - { row_block_size: 0, col_block_size: block_size, cache_data: true, storage_mode: BmatStorageMode::Dense } + { + row_block_size: 0, + col_block_size: block_size, + cache_data: true, + storage_mode: BmatStorageMode::Dense, + } } ///| /// Create parameters for row-wise access. pub fn BmatParam::row_param(block_size : Int) -> BmatParam { - { row_block_size: block_size, col_block_size: 0, cache_data: true, storage_mode: BmatStorageMode::Dense } + { + row_block_size: block_size, + col_block_size: 0, + cache_data: true, + storage_mode: BmatStorageMode::Dense, + } } ///| @@ -217,7 +237,11 @@ pub fn bmat_apply_col_blocks( let n_cols = bmat.dim_cols let mut col = 0 while col < n_cols { - let end = if col + col_block_size < n_cols { col + col_block_size } else { n_cols } + let end = if col + col_block_size < n_cols { + col + col_block_size + } else { + n_cols + } let block_data : Array[Array[Double]] = [] let mut i = 0 while i < bmat.dim_rows { @@ -230,7 +254,13 @@ pub fn bmat_apply_col_blocks( block_data.push(row_data) i = i + 1 } - let block = BmatBlock::{ row_start: 0, row_end: bmat.dim_rows, col_start: col, col_end: end, block_data } + let block = BmatBlock::{ + row_start: 0, + row_end: bmat.dim_rows, + col_start: col, + col_end: end, + block_data, + } apply_fn(block) col = end } @@ -246,7 +276,11 @@ pub fn bmat_apply_row_blocks( let n_rows = bmat.dim_rows let mut row = 0 while row < n_rows { - let end = if row + row_block_size < n_rows { row + row_block_size } else { n_rows } + let end = if row + row_block_size < n_rows { + row + row_block_size + } else { + n_rows + } let block_data : Array[Array[Double]] = [] let mut i = row while i < end { @@ -259,7 +293,13 @@ pub fn bmat_apply_row_blocks( block_data.push(row_data) i = i + 1 } - let block = BmatBlock::{ row_start: row, row_end: end, col_start: 0, col_end: bmat.dim_cols, block_data } + let block = BmatBlock::{ + row_start: row, + row_end: end, + col_start: 0, + col_end: bmat.dim_cols, + block_data, + } apply_fn(block) row = end } @@ -267,10 +307,7 @@ pub fn bmat_apply_row_blocks( ///| /// Iterate over the full matrix and apply a function. -pub fn bmat_foreach( - bmat : Bmat, - iter_fn : (Int, Int, Double) -> Unit, -) -> Unit { +pub fn bmat_foreach(bmat : Bmat, iter_fn : (Int, Int, Double) -> Unit) -> Unit { let mut i = 0 while i < bmat.dim_rows { let mut j = 0 @@ -298,7 +335,12 @@ pub struct BmatIterator { ///| /// Create a new iterator. pub fn BmatIterator::new(bmat : Bmat) -> BmatIterator { - { data: bmat.data, total_cols: bmat.dim_cols, total: bmat.dim_rows * bmat.dim_cols, pos: 0 } + { + data: bmat.data, + total_cols: bmat.dim_cols, + total: bmat.dim_rows * bmat.dim_cols, + pos: 0, + } } ///| @@ -321,13 +363,21 @@ pub fn BmatIterator::next(self : BmatIterator) -> Double? { ///| /// Get current row index. pub fn BmatIterator::cur_row(self : BmatIterator) -> Int { - if self.total_cols == 0 { 0 } else { self.pos / self.total_cols } + if self.total_cols == 0 { + 0 + } else { + self.pos / self.total_cols + } } ///| /// Get current column index. pub fn BmatIterator::cur_col(self : BmatIterator) -> Int { - if self.total_cols == 0 { 0 } else { self.pos % self.total_cols } + if self.total_cols == 0 { + 0 + } else { + self.pos % self.total_cols + } } // ============================================================================ @@ -454,10 +504,7 @@ pub fn bmat_bind_rows(top : Bmat, bottom : Bmat) -> Bmat { ///| /// Apply a function to each element. -pub fn bmat_apply_elementwise( - bmat : Bmat, - map_fn : (Double) -> Double, -) -> Bmat { +pub fn bmat_apply_elementwise(bmat : Bmat, map_fn : (Double) -> Double) -> Bmat { let new_data : Array[Array[Double]] = [] let mut i = 0 while i < bmat.dim_rows { @@ -476,7 +523,11 @@ pub fn bmat_apply_elementwise( ///| /// Pretty-print a Bmat. pub fn Bmat::to_string(self : Bmat) -> String { - let mut s = "Bmat (" + self.dim_rows.to_string() + " x " + self.dim_cols.to_string() + "):\n" + let mut s = "Bmat (" + + self.dim_rows.to_string() + + " x " + + self.dim_cols.to_string() + + "):\n" let mut i = 0 while i < self.dim_rows { let mut j = 0 diff --git a/src/bigbed.mbt b/src/bigbed.mbt new file mode 100644 index 00000000..eb373646 --- /dev/null +++ b/src/bigbed.mbt @@ -0,0 +1,2475 @@ +// Bio.Align.bigbed - indexed UCSC BigBed records and interval queries. +// +// The implementation follows the BigBed version 4 binary layout used by +// Biopython: a chromosome B+ tree, binary BED data blocks, and an R-tree +// interval index. Writers use standards-compliant zlib stored blocks so the +// output remains deterministic without requiring a native compression library. + +///| +pub suberror BigBedError { + BigBedError(String) +} + +///| +pub struct BigBedField { + as_type : String + name : String + comment : String +} derive(Eq, Debug) + +///| +pub struct BigBedSchema { + name : String + comment : String + fields : Array[BigBedField] +} derive(Eq, Debug) + +///| +pub struct BigBedTarget { + name : String + length : Int +} derive(Eq, Debug) + +///| +pub struct BigBedRecord { + chrom : String + chrom_start : Int + chrom_end : Int + name : String + score : Int + strand : String + thick_start : Int + thick_end : Int + item_rgb : String + declared_block_count : Int + block_sizes : Array[Int] + block_starts : Array[Int] + extra_fields : Array[String] +} derive(Eq, Debug) + +///| +pub struct BigBedWriteConfig { + bed_columns : Int + items_per_slot : Int + block_size : Int + compress : Bool + schema_name : String + schema_comment : String + custom_fields : Array[BigBedField] +} derive(Eq, Debug) + +///| +pub struct BigBedHeader { + little_endian : Bool + version : Int + zoom_levels : Int + chromosome_tree_offset : Int + full_data_offset : Int + full_index_offset : Int + field_count : Int + defined_field_count : Int + auto_sql_offset : Int + total_summary_offset : Int + uncompress_buffer_size : Int + extra_indices_offset : Int +} derive(Eq, Debug) + +///| +pub struct BigBedBlockIndex { + start_chrom : Int + start_base : Int + end_chrom : Int + end_base : Int + data_offset : Int + data_size : Int +} derive(Eq, Debug) + +///| +pub struct BigBedSummary { + record_count : Int + target_count : Int + data_block_count : Int + covered_bases : Int + minimum_span : Int + maximum_span : Int + compressed : Bool +} derive(Eq, Debug) + +///| +pub struct BigBedFile { + header : BigBedHeader + schema : BigBedSchema + targets : Array[BigBedTarget] + source : Array[Int] + blocks : Array[BigBedBlockIndex] + record_count : Int +} + +///| +enum BigBedEndian { + BigBedLittleEndian + BigBedBigEndian +} derive(Eq) + +///| +let bigbed_magic : Int = 0x8789F2EB + +///| +let bigbed_version : Int = 4 + +///| +let bigbed_chrom_tree_magic : Int = 0x78CA8C91 + +///| +let bigbed_rtree_magic : Int = 0x2468ACE0 + +///| +let bigbed_standard_names : Array[String] = [ + "chrom", "chromStart", "chromEnd", "name", "score", "strand", "thickStart", "thickEnd", + "reserved", "blockCount", "blockSizes", "chromStarts", +] + +///| +let bigbed_standard_types : Array[String] = [ + "string", "uint", "uint", "string", "uint", "char[1]", "uint", "uint", "uint", + "int", "int[blockCount]", "int[blockCount]", +] + +///| +let bigbed_standard_comments : Array[String] = [ + "Reference sequence chromosome or scaffold", "Start position in chromosome", "End position in chromosome", + "Name of item.", "Score (0-1000)", "+ or - for strand", "Start of where display should be thick", + "End of where display should be thick", "Used as itemRgb", "Number of blocks", + "Comma separated list of block sizes", "Start positions relative to chromStart", +] + +///| +fn bigbed_fail(message : String) -> Unit raise BigBedError { + raise BigBedError(message) +} + +///| +fn bigbed_is_supported_type(as_type : String) -> Bool { + let base = if contains(as_type, "[") { + let mut end = 0 + while end < as_type.length() && + as_type.unsafe_get(end).to_int() != '['.to_int() { + end = end + 1 + } + as_type[0:end].to_owned() + } else { + as_type + } + base == "int" || + base == "uint" || + base == "short" || + base == "ushort" || + base == "byte" || + base == "ubyte" || + base == "float" || + base == "char" || + base == "string" || + base == "lstring" +} + +///| +fn bigbed_validate_text( + value : String, + label : String, +) -> Unit raise BigBedError { + for i in 0.. BigBedField raise BigBedError { + if !bigbed_is_supported_type(as_type) { + bigbed_fail("Unsupported AutoSQL field type '" + as_type + "'") + } + if name.length() == 0 { + bigbed_fail("AutoSQL field name must not be empty") + } + bigbed_validate_text(name, "AutoSQL field name") + BigBedField::{ as_type, name, comment } +} + +///| +fn bigbed_standard_fields(count : Int) -> Array[BigBedField] { + let fields : Array[BigBedField] = [] + for i in 0.. BigBedSchema raise BigBedError { + if name.length() == 0 { + bigbed_fail("AutoSQL table name must not be empty") + } + if fields.length() < 3 { + bigbed_fail("AutoSQL schema must contain at least three fields") + } + let seen : Map[String, Bool] = Map([]) + for field in fields { + if seen.contains(field.name) { + bigbed_fail("Duplicate AutoSQL field '" + field.name + "'") + } + seen[field.name] = true + } + BigBedSchema::{ name, comment, fields: fields.copy() } +} + +///| +pub fn BigBedSchema::default( + bed_columns? : Int = 12, +) -> BigBedSchema raise BigBedError { + if bed_columns < 3 || bed_columns > 12 { + bigbed_fail("BED column count must be between 3 and 12") + } + BigBedSchema::{ + name: "bed", + comment: "Browser Extensible Data", + fields: bigbed_standard_fields(bed_columns), + } +} + +///| +pub fn BigBedTarget::create( + name : String, + length : Int, +) -> BigBedTarget raise BigBedError { + if name.length() == 0 { + bigbed_fail("BigBed target name must not be empty") + } + bigbed_validate_text(name, "BigBed target name") + if length <= 0 { + bigbed_fail("BigBed target length must be positive") + } + BigBedTarget::{ name, length } +} + +///| +fn bigbed_validate_blocks( + span : Int, + sizes : Array[Int], + starts : Array[Int], +) -> Unit raise BigBedError { + if sizes.length() != starts.length() { + bigbed_fail("BED block sizes and starts must have equal lengths") + } + if sizes.length() == 0 { + return + } + if starts[0] != 0 { + bigbed_fail("The first BED block must start at chromStart") + } + let mut previous_end = 0 + for i in 0.. span { + bigbed_fail("BED block extends beyond chromEnd") + } + previous_end = block_end + } + if previous_end != span { + bigbed_fail("The last BED block must end at chromEnd") + } +} + +///| +pub fn BigBedRecord::create( + chrom : String, + chrom_start : Int, + chrom_end : Int, + name? : String = ".", + score? : Int = 0, + strand? : String = ".", + thick_start? : Int = -1, + thick_end? : Int = -1, + item_rgb? : String = "0", + block_count? : Int = -1, + block_sizes? : Array[Int] = [], + block_starts? : Array[Int] = [], + extra_fields? : Array[String] = [], +) -> BigBedRecord raise BigBedError { + if chrom.length() == 0 { + bigbed_fail("BED chromosome must not be empty") + } + bigbed_validate_text(chrom, "BED chromosome") + if chrom_start < 0 || chrom_end <= chrom_start { + bigbed_fail("BED coordinates must define a positive half-open interval") + } + if score < 0 || score > 1000 { + bigbed_fail("BED score must be between 0 and 1000") + } + if strand != "+" && strand != "-" && strand != "." { + bigbed_fail("BED strand must be '+', '-', or '.'") + } + bigbed_validate_text(name, "BED name") + bigbed_validate_text(item_rgb, "BED itemRgb") + for value in extra_fields { + bigbed_validate_text(value, "BigBed custom field") + } + let normalized_thick_start = if thick_start < 0 { + chrom_start + } else { + thick_start + } + let normalized_thick_end = if thick_end < 0 { chrom_end } else { thick_end } + if normalized_thick_start < chrom_start || + normalized_thick_start > normalized_thick_end || + normalized_thick_end > chrom_end { + bigbed_fail("BED thick interval must lie inside the record interval") + } + let normalized_block_count = if block_count < 0 { + block_sizes.length() + } else { + block_count + } + if normalized_block_count < 0 { + bigbed_fail("BED block count must not be negative") + } + if block_sizes.length() > 0 && block_sizes.length() != normalized_block_count { + bigbed_fail("BED block count does not match block sizes") + } + if block_starts.length() > 0 && + block_starts.length() != normalized_block_count { + bigbed_fail("BED block count does not match block starts") + } + if block_starts.length() > 0 && block_sizes.length() == 0 { + bigbed_fail("BED block starts require corresponding block sizes") + } + if block_sizes.length() > 0 && block_starts.length() > 0 { + bigbed_validate_blocks(chrom_end - chrom_start, block_sizes, block_starts) + } + BigBedRecord::{ + chrom, + chrom_start, + chrom_end, + name, + score, + strand, + thick_start: normalized_thick_start, + thick_end: normalized_thick_end, + item_rgb, + declared_block_count: normalized_block_count, + block_sizes: block_sizes.copy(), + block_starts: block_starts.copy(), + extra_fields: extra_fields.copy(), + } +} + +///| +pub fn BigBedWriteConfig::create( + bed_columns? : Int = 12, + items_per_slot? : Int = 512, + block_size? : Int = 256, + compress? : Bool = true, + schema_name? : String = "bed", + schema_comment? : String = "Browser Extensible Data", + custom_fields? : Array[BigBedField] = [], +) -> BigBedWriteConfig raise BigBedError { + if bed_columns < 3 || bed_columns > 12 { + bigbed_fail("BED column count must be between 3 and 12") + } + if items_per_slot <= 0 || items_per_slot > 65535 { + bigbed_fail("BigBed itemsPerSlot must be between 1 and 65535") + } + if block_size < 2 || block_size > 65535 { + bigbed_fail("BigBed blockSize must be between 2 and 65535") + } + let fields = bigbed_standard_fields(bed_columns) + for custom in custom_fields { + for standard in fields { + if standard.name == custom.name { + bigbed_fail( + "Custom AutoSQL field duplicates standard field '" + custom.name + "'", + ) + } + } + } + ignore( + BigBedSchema::create( + schema_name, + schema_comment, + { + let all = fields.copy() + for field in custom_fields { + all.push(field) + } + all + }, + ), + ) + BigBedWriteConfig::{ + bed_columns, + items_per_slot, + block_size, + compress, + schema_name, + schema_comment, + custom_fields: custom_fields.copy(), + } +} + +///| +pub fn BigBedWriteConfig::default() -> BigBedWriteConfig { + BigBedWriteConfig::{ + bed_columns: 12, + items_per_slot: 512, + block_size: 256, + compress: true, + schema_name: "bed", + schema_comment: "Browser Extensible Data", + custom_fields: [], + } +} + +///| +pub fn BigBedRecord::span(self : BigBedRecord) -> Int { + self.chrom_end - self.chrom_start +} + +///| +pub fn BigBedRecord::block_count(self : BigBedRecord) -> Int { + self.declared_block_count +} + +///| +pub fn BigBedRecord::query_size(self : BigBedRecord) -> Int { + if self.block_sizes.length() == 0 || + self.block_sizes.length() != self.block_starts.length() { + self.span() + } else { + let mut total = 0 + for size in self.block_sizes { + total = total + size + } + total + } +} + +///| +pub fn BigBedRecord::overlaps( + self : BigBedRecord, + chromosome : String, + start : Int, + end : Int, +) -> Bool { + self.chrom == chromosome && self.chrom_start < end && start < self.chrom_end +} + +///| +pub fn BigBedRecord::coordinates(self : BigBedRecord) -> Array[Array[Int]] { + let target : Array[Int] = [] + let query : Array[Int] = [] + let query_size = self.query_size() + if self.block_sizes.length() == 0 || + self.block_sizes.length() != self.block_starts.length() { + target.push(self.chrom_start) + target.push(self.chrom_end) + query.push(0) + query.push(query_size) + } else { + let mut query_position = 0 + for i in 0.. String? { + for i in 0.. String { + let builder = StringBuilder::new() + for i in 0.. String { + let output = StringBuilder::new() + output.write_string("table " + self.name + "\n") + output.write_string("\"" + bigbed_escape_auto_sql(self.comment) + "\"\n") + output.write_string("(\n") + for field in self.fields { + output.write_string( + " " + + field.as_type + + " " + + field.name + + "; \"" + + bigbed_escape_auto_sql(field.comment) + + "\"\n", + ) + } + output.write_string(")\n") + output.to_string() +} + +///| +fn bigbed_find_byte(text : String, code : Int, start : Int) -> Int { + for i in start.. String raise BigBedError { + let first = bigbed_find_byte(text, '"'.to_int(), 0) + if first < 0 { + bigbed_fail("Expected quoted AutoSQL text") + } + let second = bigbed_find_byte(text, '"'.to_int(), first + 1) + if second < 0 { + bigbed_fail("Unterminated quoted AutoSQL text") + } + text[first + 1:second].to_owned() +} + +///| +pub fn bigbed_parse_auto_sql(text : String) -> BigBedSchema raise BigBedError { + let lines = text.split("\n").to_array() + let mut table_name = "" + let mut table_comment = "" + let fields : Array[BigBedField] = [] + let mut saw_table = false + let mut in_fields = false + for raw in lines { + let line = trim(raw.to_owned()) + if line.length() == 0 { + continue + } + if !saw_table { + let words = split_by_char(line, ' '.to_int()) + if words.length() < 2 || words[0] != "table" { + bigbed_fail("AutoSQL declaration must start with 'table'") + } + table_name = words[1] + saw_table = true + continue + } + if table_comment.length() == 0 && !in_fields { + if line == "(" { + in_fields = true + } else { + table_comment = bigbed_quoted_value(line) + } + continue + } + if line == "(" { + in_fields = true + continue + } + if line == ")" { + in_fields = false + continue + } + if in_fields { + let semicolon = bigbed_find_byte(line, ';'.to_int(), 0) + if semicolon < 0 { + bigbed_fail("AutoSQL field declaration is missing ';'") + } + let definition = trim(line[0:semicolon].to_owned()) + let words = split_by_char(definition, ' '.to_int()) + let compact : Array[String] = [] + for word in words { + if word.length() > 0 { + compact.push(word) + } + } + if compact.length() != 2 { + bigbed_fail("AutoSQL field requires a type and name") + } + fields.push( + BigBedField::create( + compact[0], + compact[1], + comment=bigbed_quoted_value(line), + ), + ) + } + } + if !saw_table || in_fields || fields.length() == 0 { + bigbed_fail("Incomplete AutoSQL declaration") + } + BigBedSchema::create(table_name, table_comment, fields) +} + +///| +fn bigbed_check_range( + data : Array[Int], + position : Int, + size : Int, + label : String, +) -> Unit raise BigBedError { + if position < 0 || size < 0 || position > data.length() - size { + bigbed_fail("Truncated BigBed " + label) + } +} + +///| +fn bigbed_read_u16( + data : Array[Int], + position : Int, + endian : BigBedEndian, +) -> Int raise BigBedError { + bigbed_check_range(data, position, 2, "16-bit integer") + if endian == BigBedLittleEndian { + data[position] | (data[position + 1] << 8) + } else { + (data[position] << 8) | data[position + 1] + } +} + +///| +fn bigbed_read_u32( + data : Array[Int], + position : Int, + endian : BigBedEndian, +) -> Int raise BigBedError { + bigbed_check_range(data, position, 4, "32-bit integer") + if endian == BigBedLittleEndian { + data[position] | + (data[position + 1] << 8) | + (data[position + 2] << 16) | + (data[position + 3] << 24) + } else { + (data[position] << 24) | + (data[position + 1] << 16) | + (data[position + 2] << 8) | + data[position + 3] + } +} + +///| +fn bigbed_read_u64_local( + data : Array[Int], + position : Int, + endian : BigBedEndian, +) -> Int raise BigBedError { + let low_position = if endian == BigBedLittleEndian { + position + } else { + position + 4 + } + let high_position = if endian == BigBedLittleEndian { + position + 4 + } else { + position + } + let low = bigbed_read_u32(data, low_position, endian) + let high = bigbed_read_u32(data, high_position, endian) + if high != 0 || low < 0 { + bigbed_fail("BigBed offset or count exceeds the local Int range") + } + low +} + +///| +fn bigbed_push_u16(output : Array[Int], value : Int) -> Unit { + output.push(value & 0xFF) + output.push((value >> 8) & 0xFF) +} + +///| +fn bigbed_push_u32(output : Array[Int], value : Int) -> Unit { + output.push(value & 0xFF) + output.push((value >> 8) & 0xFF) + output.push((value >> 16) & 0xFF) + output.push((value >> 24) & 0xFF) +} + +///| +fn bigbed_push_u64_local(output : Array[Int], value : Int) -> Unit { + bigbed_push_u32(output, value) + bigbed_push_u32(output, 0) +} + +///| +fn bigbed_patch_u16(output : Array[Int], position : Int, value : Int) -> Unit { + output[position] = value & 0xFF + output[position + 1] = (value >> 8) & 0xFF +} + +///| +fn bigbed_patch_u32(output : Array[Int], position : Int, value : Int) -> Unit { + output[position] = value & 0xFF + output[position + 1] = (value >> 8) & 0xFF + output[position + 2] = (value >> 16) & 0xFF + output[position + 3] = (value >> 24) & 0xFF +} + +///| +fn bigbed_patch_u64_local( + output : Array[Int], + position : Int, + value : Int, +) -> Unit { + bigbed_patch_u32(output, position, value) + bigbed_patch_u32(output, position + 4, 0) +} + +///| +fn bigbed_push_text(output : Array[Int], value : String) -> Unit { + for i in 0.. String raise BigBedError { + bigbed_check_range(data, start, end - start, "text") + let builder = StringBuilder::new() + for i in start.. (String, Int) raise BigBedError { + if start < 0 || limit > data.length() || start >= limit { + bigbed_fail("Invalid BigBed string bounds") + } + let mut end = start + while end < limit && data[end] != 0 { + end = end + 1 + } + if end >= limit { + bigbed_fail("Unterminated BigBed string") + } + (bigbed_read_text(data, start, end), end + 1) +} + +///| +fn bigbed_parse_nonnegative_int( + value : String, + label : String, +) -> Int raise BigBedError { + if value.length() == 0 { + bigbed_fail("Missing integer for " + label) + } + let mut result = 0 + for i in 0.. '9'.to_int() { + bigbed_fail("Invalid integer '" + value + "' for " + label) + } + let digit = code - '0'.to_int() + if result > 214748364 || (result == 214748364 && digit > 7) { + bigbed_fail("Integer for " + label + " exceeds the local Int range") + } + result = result * 10 + digit + } + result +} + +///| +fn bigbed_parse_csv_ints( + value : String, + label : String, +) -> Array[Int] raise BigBedError { + let values : Array[Int] = [] + for word in split_by_char(value, ','.to_int()) { + if word.length() > 0 { + values.push(bigbed_parse_nonnegative_int(word, label)) + } + } + values +} + +///| +fn bigbed_csv(values : Array[Int]) -> String { + let output = StringBuilder::new() + for value in values { + output.write_string(value.to_string()) + output.write_string(",") + } + output.to_string() +} + +///| +fn bigbed_schema_for_config( + config : BigBedWriteConfig, +) -> BigBedSchema raise BigBedError { + let fields = bigbed_standard_fields(config.bed_columns) + for field in config.custom_fields { + fields.push(field) + } + BigBedSchema::create(config.schema_name, config.schema_comment, fields) +} + +///| +fn bigbed_validate_schema( + schema : BigBedSchema, + defined_fields : Int, +) -> Unit raise BigBedError { + if schema.fields.length() < defined_fields { + bigbed_fail("AutoSQL schema has fewer fields than the BigBed header") + } + for i in 0.. Int { + for i in 0.. Unit raise BigBedError { + if targets.length() == 0 { + bigbed_fail("BigBed requires at least one target sequence") + } + let seen : Map[String, Bool] = Map([]) + for target in targets { + ignore(BigBedTarget::create(target.name, target.length)) + if seen.contains(target.name) { + bigbed_fail("Duplicate BigBed target '" + target.name + "'") + } + seen[target.name] = true + } + let mut previous_chrom = -1 + let mut previous_start = -1 + for i in 0.. targets[chrom_id].length { + bigbed_fail( + "Record " + i.to_string() + " extends beyond its target sequence", + ) + } + if chrom_id < previous_chrom || + (chrom_id == previous_chrom && record.chrom_start < previous_start) { + bigbed_fail("BigBed records must be sorted by target and start") + } + if config.bed_columns >= 10 && record.declared_block_count <= 0 { + bigbed_fail("BED10-BED12 records require at least one block") + } + if config.bed_columns >= 11 && + record.block_sizes.length() != record.declared_block_count { + bigbed_fail("BED11-BED12 block count must match blockSizes") + } + if config.bed_columns >= 12 && + record.block_starts.length() != record.declared_block_count { + bigbed_fail("BED12 block count must match chromStarts") + } + if record.extra_fields.length() != config.custom_fields.length() { + bigbed_fail( + "Record " + i.to_string() + " has an incorrect number of custom fields", + ) + } + previous_chrom = chrom_id + previous_start = record.chrom_start + } +} + +///| +fn bigbed_rest_fields( + record : BigBedRecord, + bed_columns : Int, +) -> Array[String] { + let fields : Array[String] = [] + if bed_columns >= 4 { + fields.push(record.name) + } + if bed_columns >= 5 { + fields.push(record.score.to_string()) + } + if bed_columns >= 6 { + fields.push(record.strand) + } + if bed_columns >= 7 { + fields.push(record.thick_start.to_string()) + } + if bed_columns >= 8 { + fields.push(record.thick_end.to_string()) + } + if bed_columns >= 9 { + fields.push(record.item_rgb) + } + if bed_columns >= 10 { + fields.push(record.declared_block_count.to_string()) + } + if bed_columns >= 11 { + fields.push(bigbed_csv(record.block_sizes)) + } + if bed_columns >= 12 { + fields.push(bigbed_csv(record.block_starts)) + } + for value in record.extra_fields { + fields.push(value) + } + fields +} + +///| +fn bigbed_join_tabs(fields : Array[String]) -> String { + let output = StringBuilder::new() + for i in 0.. 0 { + output.write_string("\t") + } + output.write_string(fields[i]) + } + output.to_string() +} + +///| +fn bigbed_record_bytes( + record : BigBedRecord, + chrom_id : Int, + bed_columns : Int, +) -> Array[Int] { + let output : Array[Int] = [] + bigbed_push_u32(output, chrom_id) + bigbed_push_u32(output, record.chrom_start) + bigbed_push_u32(output, record.chrom_end) + bigbed_push_text( + output, + bigbed_join_tabs(bigbed_rest_fields(record, bed_columns)), + ) + output.push(0) + output +} + +///| +fn bigbed_adler32(data : Array[Int]) -> Int { + let mut a = 1 + let mut b = 0 + for byte in data { + a = (a + (byte & 0xFF)) % 65521 + b = (b + a) % 65521 + } + (b << 16) | a +} + +///| +struct BigBedBitReader { + data : Array[Int] + bit_position : Int + bit_limit : Int +} + +///| +struct BigBedHuffman { + counts : Array[Int] + symbols : Array[Int] +} + +///| +let bigbed_length_base : Array[Int] = [ + 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 15, 17, 19, 23, 27, 31, 35, 43, 51, 59, 67, 83, + 99, 115, 131, 163, 195, 227, 258, +] + +///| +let bigbed_length_extra : Array[Int] = [ + 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4, 5, 5, 5, + 5, 0, +] + +///| +let bigbed_distance_base : Array[Int] = [ + 1, 2, 3, 4, 5, 7, 9, 13, 17, 25, 33, 49, 65, 97, 129, 193, 257, 385, 513, 769, + 1025, 1537, 2049, 3073, 4097, 6145, 8193, 12289, 16385, 24577, +] + +///| +let bigbed_distance_extra : Array[Int] = [ + 0, 0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 8, 9, 9, 10, 10, 11, 11, + 12, 12, 13, 13, +] + +///| +fn bigbed_read_deflate_bits( + reader : BigBedBitReader, + count : Int, +) -> (Int, BigBedBitReader) raise BigBedError { + if count < 0 || count > 24 { + bigbed_fail("Invalid DEFLATE bit count") + } + if reader.bit_position + count > reader.bit_limit { + bigbed_fail("Truncated DEFLATE stream in BigBed data block") + } + let mut value = 0 + for i in 0..> (position % 8)) & 1 + value = value | (bit << i) + } + ( + value, + BigBedBitReader::{ + data: reader.data, + bit_position: reader.bit_position + count, + bit_limit: reader.bit_limit, + }, + ) +} + +///| +fn bigbed_align_deflate_reader( + reader : BigBedBitReader, +) -> BigBedBitReader raise BigBedError { + let aligned = (reader.bit_position + 7) / 8 * 8 + if aligned > reader.bit_limit { + bigbed_fail("Truncated DEFLATE byte alignment") + } + BigBedBitReader::{ + data: reader.data, + bit_position: aligned, + bit_limit: reader.bit_limit, + } +} + +///| +fn bigbed_build_huffman( + lengths : Array[Int], + label : String, +) -> BigBedHuffman raise BigBedError { + let counts : Array[Int] = Array::make(16, 0) + for length in lengths { + if length < 0 || length > 15 { + bigbed_fail("Invalid " + label + " Huffman code length") + } + counts[length] = counts[length] + 1 + } + if counts[0] == lengths.length() { + bigbed_fail("Empty " + label + " Huffman tree") + } + let mut remaining = 1 + for bits in 1..<16 { + remaining = (remaining << 1) - counts[bits] + if remaining < 0 { + bigbed_fail("Oversubscribed " + label + " Huffman tree") + } + } + let offsets : Array[Int] = Array::make(16, 0) + for bits in 1..<15 { + offsets[bits + 1] = offsets[bits] + counts[bits] + } + let next = offsets.copy() + let symbols : Array[Int] = Array::make(lengths.length() - counts[0], 0) + for symbol in 0.. 0 { + symbols[next[length]] = symbol + next[length] = next[length] + 1 + } + } + BigBedHuffman::{ counts, symbols } +} + +///| +fn bigbed_decode_huffman_symbol( + tree : BigBedHuffman, + input : BigBedBitReader, +) -> (Int, BigBedBitReader) raise BigBedError { + let mut reader = input + let mut code = 0 + let mut first = 0 + let mut index = 0 + for length in 1..<16 { + let (bit, next) = bigbed_read_deflate_bits(reader, 1) + reader = next + code = code | bit + let count = tree.counts[length] + if code < first + count { + let symbol_index = index + code - first + if symbol_index < 0 || symbol_index >= tree.symbols.length() { + bigbed_fail("Invalid canonical Huffman symbol index") + } + return (tree.symbols[symbol_index], reader) + } + index = index + count + first = (first + count) << 1 + code = code << 1 + } + bigbed_fail("Invalid Huffman code in BigBed DEFLATE stream") + (0, reader) +} + +///| +fn bigbed_fixed_huffman() -> (BigBedHuffman, BigBedHuffman) raise BigBedError { + let literal_lengths : Array[Int] = Array::make(288, 0) + for i in 0..<144 { + literal_lengths[i] = 8 + } + for i in 144..<256 { + literal_lengths[i] = 9 + } + for i in 256..<280 { + literal_lengths[i] = 7 + } + for i in 280..<288 { + literal_lengths[i] = 8 + } + let distance_lengths : Array[Int] = Array::make(32, 5) + ( + bigbed_build_huffman(literal_lengths, "fixed literal/length"), + bigbed_build_huffman(distance_lengths, "fixed distance"), + ) +} + +///| +fn bigbed_dynamic_huffman( + input : BigBedBitReader, +) -> (BigBedHuffman, BigBedHuffman, BigBedBitReader) raise BigBedError { + let mut reader = input + let (literal_bits, next) = bigbed_read_deflate_bits(reader, 5) + reader = next + let literal_count = literal_bits + 257 + let (distance_bits, next) = bigbed_read_deflate_bits(reader, 5) + reader = next + let distance_count = distance_bits + 1 + let (code_bits, next) = bigbed_read_deflate_bits(reader, 4) + reader = next + let code_count = code_bits + 4 + if literal_count > 286 || distance_count > 32 { + bigbed_fail("Invalid dynamic DEFLATE tree dimensions") + } + let code_order = [ + 16, 17, 18, 0, 8, 7, 9, 6, 10, 5, 11, 4, 12, 3, 13, 2, 14, 1, 15, + ] + let code_lengths : Array[Int] = Array::make(19, 0) + for i in 0.. total { + bigbed_fail("DEFLATE code-length repeat exceeds its tree") + } + let previous = lengths[position - 1] + for _ in 0.. total { + bigbed_fail("DEFLATE zero repeat exceeds its tree") + } + position = position + repeat + } else if symbol == 18 { + let (extra, next) = bigbed_read_deflate_bits(reader, 7) + reader = next + let repeat = extra + 11 + if position + repeat > total { + bigbed_fail("DEFLATE long zero repeat exceeds its tree") + } + position = position + repeat + } else { + bigbed_fail("Invalid DEFLATE code-length symbol") + } + } + let literal_lengths : Array[Int] = [] + for i in 0.. BigBedBitReader raise BigBedError { + let mut reader = input + let mut done = false + while !done { + let (symbol, next) = bigbed_decode_huffman_symbol(literal_tree, reader) + reader = next + if symbol < 256 { + if output.length() >= maximum_output { + bigbed_fail("BigBed data block exceeds the declared buffer size") + } + output.push(symbol) + } else if symbol == 256 { + done = true + } else { + if symbol < 257 || symbol > 285 { + bigbed_fail("Invalid DEFLATE length symbol") + } + let length_index = symbol - 257 + let (length_extra, next) = bigbed_read_deflate_bits( + reader, + bigbed_length_extra[length_index], + ) + reader = next + let length = bigbed_length_base[length_index] + length_extra + let (distance_symbol, next) = bigbed_decode_huffman_symbol( + distance_tree, reader, + ) + reader = next + if distance_symbol < 0 || distance_symbol >= 30 { + bigbed_fail("Invalid DEFLATE distance symbol") + } + let (distance_extra, next) = bigbed_read_deflate_bits( + reader, + bigbed_distance_extra[distance_symbol], + ) + reader = next + let distance = bigbed_distance_base[distance_symbol] + distance_extra + if distance <= 0 || distance > output.length() || distance > 32768 { + bigbed_fail("Invalid LZ77 distance in BigBed DEFLATE stream") + } + if output.length() + length > maximum_output { + bigbed_fail("BigBed data block exceeds the declared buffer size") + } + for _ in 0.. Array[Int] { + let output : Array[Int] = [0x78, 0x01] + let mut position = 0 + if data.length() == 0 { + output.push(1) + bigbed_push_u16(output, 0) + bigbed_push_u16(output, 0xFFFF) + } + while position < data.length() { + let remaining = data.length() - position + let size = if remaining > 65535 { 65535 } else { remaining } + let final_block = position + size == data.length() + output.push(if final_block { 1 } else { 0 }) + bigbed_push_u16(output, size) + bigbed_push_u16(output, size ^ 0xFFFF) + for i in position..<(position + size) { + output.push(data[i]) + } + position = position + size + } + let checksum = bigbed_adler32(data) + output.push((checksum >> 24) & 0xFF) + output.push((checksum >> 16) & 0xFF) + output.push((checksum >> 8) & 0xFF) + output.push(checksum & 0xFF) + output +} + +///| +fn bigbed_zlib_unstore( + data : Array[Int], + maximum_output : Int, +) -> Array[Int] raise BigBedError { + if data.length() < 8 { + bigbed_fail("Truncated zlib data block") + } + let cmf = data[0] + let flags = data[1] + if (cmf & 0x0F) != 8 || cmf >> 4 > 7 || ((cmf << 8) | flags) % 31 != 0 { + bigbed_fail("Invalid zlib header in BigBed data block") + } + if (flags & 0x20) != 0 { + bigbed_fail("Preset zlib dictionaries are not supported") + } + let payload_end = data.length() - 4 + let output : Array[Int] = [] + let mut reader = BigBedBitReader::{ + data, + bit_position: 16, + bit_limit: payload_end * 8, + } + let mut done = false + while !done { + let (final_block, next) = bigbed_read_deflate_bits(reader, 1) + reader = next + let (block_type, next) = bigbed_read_deflate_bits(reader, 2) + reader = next + done = final_block == 1 + if block_type == 0 { + reader = bigbed_align_deflate_reader(reader) + let (size, next) = bigbed_read_deflate_bits(reader, 16) + reader = next + let (inverse, next) = bigbed_read_deflate_bits(reader, 16) + reader = next + if (size ^ 0xFFFF) != inverse { + bigbed_fail("Invalid stored DEFLATE block length") + } + if output.length() + size > maximum_output { + bigbed_fail("BigBed data block exceeds the declared buffer size") + } + for _ in 0.. Unit { + for byte in source { + output.push(byte) + } +} + +///| +fn bigbed_patch_header( + output : Array[Int], + chromosome_tree_offset : Int, + full_data_offset : Int, + full_index_offset : Int, + field_count : Int, + defined_field_count : Int, + auto_sql_offset : Int, + total_summary_offset : Int, + uncompress_buffer_size : Int, +) -> Unit { + bigbed_patch_u32(output, 0, bigbed_magic) + bigbed_patch_u16(output, 4, bigbed_version) + bigbed_patch_u16(output, 6, 0) + bigbed_patch_u64_local(output, 8, chromosome_tree_offset) + bigbed_patch_u64_local(output, 16, full_data_offset) + bigbed_patch_u64_local(output, 24, full_index_offset) + bigbed_patch_u16(output, 32, field_count) + bigbed_patch_u16(output, 34, defined_field_count) + bigbed_patch_u64_local(output, 36, auto_sql_offset) + bigbed_patch_u64_local(output, 44, total_summary_offset) + bigbed_patch_u32(output, 52, uncompress_buffer_size) + bigbed_patch_u64_local(output, 56, 0) +} + +///| +fn bigbed_write_tree_key( + output : Array[Int], + key : String, + key_size : Int, +) -> Unit { + for i in 0.. Int { + let mut height = 1 + let mut capacity = block_size + while capacity < item_count { + capacity = if capacity > item_count / block_size { + item_count + } else { + capacity * block_size + } + height = height + 1 + } + height +} + +///| +fn bigbed_tree_capacity( + item_count : Int, + block_size : Int, + height : Int, +) -> Int { + let mut capacity = 1 + for _ in 0.. item_count / block_size { + item_count + } else { + capacity * block_size + } + } + capacity +} + +///| +fn bigbed_write_chrom_node( + output : Array[Int], + targets : Array[BigBedTarget], + order : Array[Int], + start : Int, + end : Int, + key_size : Int, + block_size : Int, + height : Int, +) -> Unit { + let count = end - start + if height == 1 { + output.push(1) + output.push(0) + bigbed_push_u16(output, count) + for position in start.. Unit { + let mut key_size = 1 + for target in targets { + if target.name.length() > key_size { + key_size = target.name.length() + } + } + let order : Array[Int] = [] + for i in 0.. Int { + if targets[left].name < targets[right].name { + -1 + } else if targets[left].name > targets[right].name { + 1 + } else { + 0 + } + }) + bigbed_push_u32(output, bigbed_chrom_tree_magic) + bigbed_push_u32(output, configured_block_size) + bigbed_push_u32(output, key_size) + bigbed_push_u32(output, 8) + bigbed_push_u64_local(output, targets.length()) + bigbed_push_u64_local(output, 0) + bigbed_write_chrom_node( + output, + targets, + order, + 0, + targets.length(), + key_size, + configured_block_size, + bigbed_tree_height(targets.length(), configured_block_size), + ) +} + +///| +fn bigbed_block_overlaps( + block : BigBedBlockIndex, + chrom_id : Int, + start : Int, + end : Int, +) -> Bool { + let starts_before_query_end = block.start_chrom < chrom_id || + (block.start_chrom == chrom_id && block.start_base < end) + let query_starts_before_block_end = chrom_id < block.end_chrom || + (chrom_id == block.end_chrom && start < block.end_base) + starts_before_query_end && query_starts_before_block_end +} + +///| +fn bigbed_rtree_end( + blocks : Array[BigBedBlockIndex], + start : Int, + end : Int, +) -> (Int, Int) { + let mut end_chrom = blocks[start].end_chrom + let mut end_base = blocks[start].end_base + for i in (start + 1).. end_chrom || + (block.end_chrom == end_chrom && block.end_base > end_base) { + end_chrom = block.end_chrom + end_base = block.end_base + } + } + (end_chrom, end_base) +} + +///| +fn bigbed_write_rtree_node( + output : Array[Int], + blocks : Array[BigBedBlockIndex], + start : Int, + end : Int, + block_size : Int, + height : Int, +) -> Unit { + let count = end - start + if height == 1 { + output.push(1) + output.push(0) + bigbed_push_u16(output, count) + for i in start.. Unit { + bigbed_push_u32(output, bigbed_rtree_magic) + bigbed_push_u32(output, configured_block_size) + bigbed_push_u64_local(output, item_count) + if blocks.length() == 0 { + for _ in 0..<4 { + bigbed_push_u32(output, 0) + } + } else { + bigbed_push_u32(output, blocks[0].start_chrom) + bigbed_push_u32(output, blocks[0].start_base) + let (end_chrom, end_base) = bigbed_rtree_end(blocks, 0, blocks.length()) + bigbed_push_u32(output, end_chrom) + bigbed_push_u32(output, end_base) + } + bigbed_push_u64_local(output, end_file_offset) + bigbed_push_u32(output, items_per_slot) + bigbed_push_u32(output, 0) + if blocks.length() == 0 { + output.push(1) + output.push(0) + bigbed_push_u16(output, 0) + } else { + bigbed_write_rtree_node( + output, + blocks, + 0, + blocks.length(), + configured_block_size, + bigbed_tree_height(blocks.length(), configured_block_size), + ) + } +} + +///| +pub fn bigbed_write( + targets : Array[BigBedTarget], + records : Array[BigBedRecord], + config? : BigBedWriteConfig = BigBedWriteConfig::default(), +) -> Array[Int] raise BigBedError { + bigbed_validate_records(targets, records, config) + let schema = bigbed_schema_for_config(config) + let output : Array[Int] = Array::make(64, 0) + let auto_sql_offset = output.length() + bigbed_push_text(output, schema.to_auto_sql()) + output.push(0) + let total_summary_offset = output.length() + for _ in 0..<40 { + output.push(0) + } + let chromosome_tree_offset = output.length() + bigbed_write_chrom_tree(output, targets, config.block_size) + let full_data_offset = output.length() + bigbed_push_u64_local(output, records.length()) + let blocks : Array[BigBedBlockIndex] = [] + let mut maximum_uncompressed = 0 + let mut record_index = 0 + while record_index < records.length() { + let block_start_index = record_index + let block_end_index = if record_index + config.items_per_slot < + records.length() { + record_index + config.items_per_slot + } else { + records.length() + } + let raw : Array[Int] = [] + while record_index < block_end_index { + let record = records[record_index] + let chrom_id = bigbed_target_index(targets, record.chrom) + bigbed_append( + raw, + bigbed_record_bytes(record, chrom_id, config.bed_columns), + ) + record_index = record_index + 1 + } + if raw.length() > maximum_uncompressed { + maximum_uncompressed = raw.length() + } + let payload = if config.compress { bigbed_zlib_store(raw) } else { raw } + let data_offset = output.length() + bigbed_append(output, payload) + let first = records[block_start_index] + let mut end_chrom = bigbed_target_index(targets, first.chrom) + let mut end_base = first.chrom_end + for i in (block_start_index + 1).. end_chrom || + (chrom_id == end_chrom && records[i].chrom_end > end_base) { + end_chrom = chrom_id + end_base = records[i].chrom_end + } + } + blocks.push(BigBedBlockIndex::{ + start_chrom: bigbed_target_index(targets, first.chrom), + start_base: first.chrom_start, + end_chrom, + end_base, + data_offset, + data_size: payload.length(), + }) + } + let full_index_offset = output.length() + bigbed_write_rtree( + output, + blocks, + config.block_size, + records.length(), + config.items_per_slot, + full_index_offset, + ) + bigbed_push_u32(output, bigbed_magic) + bigbed_patch_header( + output, + chromosome_tree_offset, + full_data_offset, + full_index_offset, + schema.fields.length(), + config.bed_columns, + auto_sql_offset, + total_summary_offset, + if config.compress { + maximum_uncompressed + } else { + 0 + }, + ) + output +} + +///| +fn bigbed_detect_endian(data : Array[Int]) -> BigBedEndian raise BigBedError { + if data.length() < 4 { + bigbed_fail("BigBed input is shorter than its magic number") + } + let little = data[0] | (data[1] << 8) | (data[2] << 16) | (data[3] << 24) + if little == bigbed_magic { + BigBedLittleEndian + } else { + let big = (data[0] << 24) | (data[1] << 16) | (data[2] << 8) | data[3] + if big == bigbed_magic { + BigBedBigEndian + } else { + bigbed_fail("Input does not contain the BigBed magic number") + BigBedLittleEndian + } + } +} + +///| +fn bigbed_read_header( + data : Array[Int], +) -> (BigBedHeader, BigBedEndian) raise BigBedError { + bigbed_check_range(data, 0, 64, "header") + let endian = bigbed_detect_endian(data) + let version = bigbed_read_u16(data, 4, endian) + if version != bigbed_version { + bigbed_fail( + "Unsupported BigBed version " + + version.to_string() + + "; expected version 4", + ) + } + let defined = bigbed_read_u16(data, 34, endian) + if defined < 3 || defined > 12 { + bigbed_fail("BigBed defined field count must be between 3 and 12") + } + let field_count = bigbed_read_u16(data, 32, endian) + if field_count < defined { + bigbed_fail("BigBed field count is smaller than its BED field count") + } + let header = BigBedHeader::{ + little_endian: endian == BigBedLittleEndian, + version, + zoom_levels: bigbed_read_u16(data, 6, endian), + chromosome_tree_offset: bigbed_read_u64_local(data, 8, endian), + full_data_offset: bigbed_read_u64_local(data, 16, endian), + full_index_offset: bigbed_read_u64_local(data, 24, endian), + field_count, + defined_field_count: defined, + auto_sql_offset: bigbed_read_u64_local(data, 36, endian), + total_summary_offset: bigbed_read_u64_local(data, 44, endian), + uncompress_buffer_size: bigbed_read_u32(data, 52, endian), + extra_indices_offset: bigbed_read_u64_local(data, 56, endian), + } + if header.chromosome_tree_offset < 64 || + header.full_data_offset <= header.chromosome_tree_offset || + header.full_index_offset <= header.full_data_offset || + header.full_index_offset >= data.length() { + bigbed_fail("BigBed header contains invalid section offsets") + } + (header, endian) +} + +///| +fn bigbed_read_key( + data : Array[Int], + position : Int, + key_size : Int, +) -> String raise BigBedError { + bigbed_check_range(data, position, key_size, "B+ tree key") + let mut end = position + while end < position + key_size && data[end] != 0 { + end = end + 1 + } + bigbed_read_text(data, position, end) +} + +///| +fn bigbed_read_chrom_node( + data : Array[Int], + position : Int, + key_size : Int, + endian : BigBedEndian, + targets : Array[BigBedTarget?], + depth : Int, +) -> Unit raise BigBedError { + if depth > 64 { + bigbed_fail("BigBed chromosome B+ tree exceeds 64 levels") + } + bigbed_check_range(data, position, 4, "B+ tree node") + let is_leaf = data[position] != 0 + let count = bigbed_read_u16(data, position + 2, endian) + let mut cursor = position + 4 + if is_leaf { + for _ in 0..= targets.length() { + bigbed_fail("Chromosome B+ tree contains an invalid target ID") + } + if targets[chrom_id] is Some(_) { + bigbed_fail("Chromosome B+ tree contains a duplicate target ID") + } + targets[chrom_id] = Some(BigBedTarget::create(name, chrom_size)) + } + } else { + let children : Array[Int] = [] + for _ in 0.. Array[BigBedTarget] raise BigBedError { + bigbed_check_range(data, offset, 32, "chromosome B+ tree header") + if bigbed_read_u32(data, offset, endian) != bigbed_chrom_tree_magic { + bigbed_fail("Invalid BigBed chromosome B+ tree magic number") + } + let key_size = bigbed_read_u32(data, offset + 8, endian) + let value_size = bigbed_read_u32(data, offset + 12, endian) + let item_count = bigbed_read_u64_local(data, offset + 16, endian) + if key_size <= 0 || value_size != 8 || item_count <= 0 { + bigbed_fail("Invalid BigBed chromosome B+ tree dimensions") + } + let slots : Array[BigBedTarget?] = Array::make(item_count, None) + bigbed_read_chrom_node(data, offset + 32, key_size, endian, slots, 0) + let targets : Array[BigBedTarget] = [] + for slot in slots { + match slot { + Some(target) => targets.push(target) + None => bigbed_fail("Chromosome B+ tree is missing a target ID") + } + } + targets +} + +///| +fn bigbed_read_rtree_node( + data : Array[Int], + position : Int, + endian : BigBedEndian, + blocks : Array[BigBedBlockIndex], + depth : Int, +) -> Unit raise BigBedError { + if depth > 64 { + bigbed_fail("BigBed R-tree exceeds 64 levels") + } + bigbed_check_range(data, position, 4, "R-tree node") + let is_leaf = data[position] != 0 + let count = bigbed_read_u16(data, position + 2, endian) + let mut cursor = position + 4 + if is_leaf { + for _ in 0.. (Array[BigBedBlockIndex], Int) raise BigBedError { + bigbed_check_range(data, offset, 48, "R-tree header") + if bigbed_read_u32(data, offset, endian) != bigbed_rtree_magic { + bigbed_fail("Invalid BigBed R-tree magic number") + } + let indexed_count = bigbed_read_u64_local(data, offset + 8, endian) + let blocks : Array[BigBedBlockIndex] = [] + bigbed_read_rtree_node(data, offset + 48, endian, blocks, 0) + for block in blocks { + bigbed_check_range( + data, + block.data_offset, + block.data_size, + "indexed data block", + ) + } + (blocks, indexed_count) +} + +///| +fn bigbed_read_schema( + data : Array[Int], + header : BigBedHeader, +) -> BigBedSchema raise BigBedError { + if header.auto_sql_offset == 0 { + return BigBedSchema::default(bed_columns=header.defined_field_count) + } + if header.auto_sql_offset < 64 || + header.auto_sql_offset >= header.chromosome_tree_offset { + bigbed_fail("BigBed AutoSQL offset is outside the metadata section") + } + let (text, _) = bigbed_read_c_string( + data, + header.auto_sql_offset, + header.chromosome_tree_offset, + ) + let schema = bigbed_parse_auto_sql(text) + if schema.fields.length() != header.field_count { + bigbed_fail("AutoSQL field count does not match the BigBed header") + } + bigbed_validate_schema(schema, header.defined_field_count) + schema +} + +///| +pub fn bigbed_parse(data : Array[Int]) -> BigBedFile raise BigBedError { + let (header, endian) = bigbed_read_header(data) + let trailing = bigbed_read_u32(data, data.length() - 4, endian) + if trailing != bigbed_magic { + bigbed_fail("BigBed trailing magic number is missing") + } + let schema = bigbed_read_schema(data, header) + let targets = bigbed_read_chrom_tree( + data, + header.chromosome_tree_offset, + endian, + ) + let record_count = bigbed_read_u64_local( + data, + header.full_data_offset, + endian, + ) + let (blocks, indexed_count) = bigbed_read_rtree( + data, + header.full_index_offset, + endian, + ) + if indexed_count != record_count { + bigbed_fail("BigBed data and R-tree record counts disagree") + } + BigBedFile::{ + header, + schema, + targets, + source: data.copy(), + blocks, + record_count, + } +} + +///| +fn bigbed_decode_block( + file : BigBedFile, + block : BigBedBlockIndex, +) -> Array[Int] raise BigBedError { + let payload : Array[Int] = [] + for i in block.data_offset..<(block.data_offset + block.data_size) { + payload.push(file.source[i]) + } + if file.header.uncompress_buffer_size > 0 { + bigbed_zlib_unstore(payload, file.header.uncompress_buffer_size) + } else { + payload + } +} + +///| +fn bigbed_record_from_words( + file : BigBedFile, + chrom_id : Int, + chrom_start : Int, + chrom_end : Int, + words : Array[String], +) -> BigBedRecord raise BigBedError { + if chrom_id < 0 || chrom_id >= file.targets.length() { + bigbed_fail("Binary BED record contains an invalid chromosome ID") + } + let expected_words = file.header.field_count - 3 + if words.length() != expected_words { + bigbed_fail( + "Binary BED record contains " + + words.length().to_string() + + " fields after its coordinates; expected " + + expected_words.to_string(), + ) + } + let bed_columns = file.header.defined_field_count + let name = if bed_columns >= 4 { words[0] } else { "." } + let score = if bed_columns >= 5 { + bigbed_parse_nonnegative_int(words[1], "BED score") + } else { + 0 + } + let strand = if bed_columns >= 6 { words[2] } else { "." } + let thick_start = if bed_columns >= 7 { + bigbed_parse_nonnegative_int(words[3], "BED thickStart") + } else { + chrom_start + } + let thick_end = if bed_columns >= 8 { + bigbed_parse_nonnegative_int(words[4], "BED thickEnd") + } else { + chrom_end + } + let item_rgb = if bed_columns >= 9 { words[5] } else { "0" } + let declared_blocks = if bed_columns >= 10 { + bigbed_parse_nonnegative_int(words[6], "BED blockCount") + } else { + 0 + } + let block_sizes = if bed_columns >= 11 { + bigbed_parse_csv_ints(words[7], "BED blockSizes") + } else { + [] + } + let block_starts = if bed_columns >= 12 { + bigbed_parse_csv_ints(words[8], "BED chromStarts") + } else { + [] + } + if bed_columns >= 11 && block_sizes.length() != declared_blocks { + bigbed_fail("BED blockCount does not match blockSizes") + } + if bed_columns >= 12 && block_starts.length() != declared_blocks { + bigbed_fail("BED blockCount does not match chromStarts") + } + let extra : Array[String] = [] + for i in (bed_columns - 3).. Array[BigBedRecord] raise BigBedError { + let data = bigbed_decode_block(file, block) + let endian = if file.header.little_endian { + BigBedLittleEndian + } else { + BigBedBigEndian + } + let records : Array[BigBedRecord] = [] + let mut position = 0 + while position < data.length() { + bigbed_check_range(data, position, 13, "binary BED record") + let chrom_id = bigbed_read_u32(data, position, endian) + let chrom_start = bigbed_read_u32(data, position + 4, endian) + let chrom_end = bigbed_read_u32(data, position + 8, endian) + let (rest, next) = bigbed_read_c_string(data, position + 12, data.length()) + let words = if rest.length() == 0 { + [] + } else { + split_by_char(rest, '\t'.to_int()) + } + records.push( + bigbed_record_from_words(file, chrom_id, chrom_start, chrom_end, words), + ) + position = next + } + records +} + +///| +pub fn BigBedFile::records( + self : BigBedFile, +) -> Array[BigBedRecord] raise BigBedError { + let records : Array[BigBedRecord] = [] + for block in self.blocks { + for record in bigbed_decode_records(self, block) { + records.push(record) + } + } + if records.length() != self.record_count { + bigbed_fail("Decoded BigBed record count does not match its header") + } + records +} + +///| +pub fn BigBedFile::search( + self : BigBedFile, + chromosome : String, + start? : Int = 0, + end? : Int = -1, +) -> Array[BigBedRecord] raise BigBedError { + let chrom_id = bigbed_target_index(self.targets, chromosome) + if chrom_id < 0 { + bigbed_fail("Unknown BigBed target '" + chromosome + "'") + } + let resolved_end = if end < 0 { self.targets[chrom_id].length } else { end } + if start < 0 || + resolved_end <= start || + resolved_end > self.targets[chrom_id].length { + bigbed_fail("BigBed search requires a valid positive half-open interval") + } + let results : Array[BigBedRecord] = [] + for block in self.blocks { + if bigbed_block_overlaps(block, chrom_id, start, resolved_end) { + for record in bigbed_decode_records(self, block) { + if record.overlaps(chromosome, start, resolved_end) { + results.push(record) + } + } + } + } + results +} + +///| +pub fn BigBedFile::find_by_name( + self : BigBedFile, + name : String, +) -> Array[BigBedRecord] raise BigBedError { + let results : Array[BigBedRecord] = [] + for record in self.records() { + if record.name == name { + results.push(record) + } + } + results +} + +///| +pub fn BigBedFile::target(self : BigBedFile, name : String) -> BigBedTarget? { + let index = bigbed_target_index(self.targets, name) + if index < 0 { + None + } else { + Some(self.targets[index]) + } +} + +///| +pub fn BigBedFile::is_compressed(self : BigBedFile) -> Bool { + self.header.uncompress_buffer_size > 0 +} + +///| +pub fn BigBedFile::summary( + self : BigBedFile, +) -> BigBedSummary raise BigBedError { + let mut covered_bases = 0 + let mut minimum_span = 0 + let mut maximum_span = 0 + let records = self.records() + for i in 0.. maximum_span { + maximum_span = span + } + } + BigBedSummary::{ + record_count: self.record_count, + target_count: self.targets.length(), + data_block_count: self.blocks.length(), + covered_bases, + minimum_span, + maximum_span, + compressed: self.is_compressed(), + } +} + +///| +pub fn BigBedFile::to_bed(self : BigBedFile) -> String raise BigBedError { + let output = StringBuilder::new() + let bed_columns = self.header.defined_field_count + for record in self.records() { + output.write_string(record.chrom) + output.write_string("\t") + output.write_string(record.chrom_start.to_string()) + output.write_string("\t") + output.write_string(record.chrom_end.to_string()) + let rest = bigbed_rest_fields(record, bed_columns) + for field in rest { + output.write_string("\t") + output.write_string(field) + } + output.write_string("\n") + } + output.to_string() +} + +///| +pub fn bigbed_example_data() -> (Array[BigBedTarget], Array[BigBedRecord]) { + let targets = [ + BigBedTarget::{ name: "chr1", length: 1000000 }, + BigBedTarget::{ name: "chr2", length: 500000 }, + ] + let records = [ + BigBedRecord::{ + chrom: "chr1", + chrom_start: 100, + chrom_end: 280, + name: "transcript-A", + score: 960, + strand: "+", + thick_start: 120, + thick_end: 260, + item_rgb: "40,120,220", + declared_block_count: 3, + block_sizes: [50, 40, 30], + block_starts: [0, 80, 150], + extra_fields: [], + }, + BigBedRecord::{ + chrom: "chr1", + chrom_start: 250, + chrom_end: 410, + name: "transcript-B", + score: 870, + strand: "-", + thick_start: 270, + thick_end: 390, + item_rgb: "220,80,80", + declared_block_count: 2, + block_sizes: [60, 50], + block_starts: [0, 110], + extra_fields: [], + }, + BigBedRecord::{ + chrom: "chr2", + chrom_start: 1000, + chrom_end: 1120, + name: "transcript-C", + score: 700, + strand: "+", + thick_start: 1000, + thick_end: 1120, + item_rgb: "0", + declared_block_count: 2, + block_sizes: [45, 50], + block_starts: [0, 70], + extra_fields: [], + }, + ] + (targets, records) +} diff --git a/src/bigmaf.mbt b/src/bigmaf.mbt new file mode 100644 index 00000000..73a4b5c7 --- /dev/null +++ b/src/bigmaf.mbt @@ -0,0 +1,1344 @@ +// Bio.Align.bigmaf - indexed UCSC BigMaf multiple alignments. +// +// BigMaf is a bed3+1 BigBed file whose custom lstring field contains a +// semicolon-delimited MAF block. The binary container is implemented by +// bigbed.mbt; this module owns strict MAF block validation and conversion. + +///| +pub suberror BigMafError { + BigMafError(String) +} + +///| +pub struct BigMafInsertion { + left_status : String + left_count : Int + right_status : String + right_count : Int +} derive(Eq, Debug) + +///| +pub struct BigMafComponent { + source : String + start : Int + size : Int + strand : String + source_size : Int + text : String + insertion : BigMafInsertion? + quality : String? +} derive(Eq, Debug) + +///| +pub struct BigMafEmptyComponent { + source : String + start : Int + size : Int + strand : String + source_size : Int + status : String +} derive(Eq, Debug) + +///| +pub struct BigMafBlock { + score : Double? + pass_number : Int? + components : Array[BigMafComponent] + empty_components : Array[BigMafEmptyComponent] + comments : Array[String] +} derive(Eq, Debug) + +///| +pub struct BigMafWriteConfig { + compress : Bool + block_size : Int + items_per_slot : Int +} derive(Eq, Debug) + +///| +pub struct BigMafSummary { + alignment_count : Int + target_count : Int + component_count : Int + empty_component_count : Int + aligned_columns : Int + covered_reference_bases : Int + compressed : Bool +} derive(Eq, Debug) + +///| +pub struct BigMafFile { + reference : String + bed : BigBedFile +} + +///| +fn bigmaf_fail(message : String) -> Unit raise BigMafError { + raise BigMafError(message) +} + +///| +fn bigmaf_is_space(code : Int) -> Bool { + code == ' '.to_int() || + code == '\t'.to_int() || + code == '\n'.to_int() || + code == '\r'.to_int() +} + +///| +fn bigmaf_words(text : String) -> Array[String] { + let words : Array[String] = [] + let mut start = -1 + for index in 0..= 0 { + words.push(text[start:index].to_owned()) + start = -1 + } + } else if start < 0 { + start = index + } + } + if start >= 0 { + words.push(text[start:text.length()].to_owned()) + } + words +} + +///| +fn bigmaf_validate_token( + value : String, + label : String, +) -> Unit raise BigMafError { + if value.length() == 0 { + bigmaf_fail(label + " must not be empty") + } + for index in 0.. Unit raise BigMafError { + for index in 0.. Int raise BigMafError { + if value.length() == 0 { + bigmaf_fail("Missing integer for " + label) + } + let mut first = 0 + if value.unsafe_get(0).to_int() == '-'.to_int() { + if value.length() == 1 { + bigmaf_fail("Invalid integer for " + label + ": " + value) + } + first = 1 + } + for index in first.. '9'.to_int() { + bigmaf_fail("Invalid integer for " + label + ": " + value) + } + } + parse_int(value) +} + +///| +fn bigmaf_parse_double( + value : String, + label : String, +) -> Double raise BigMafError { + match parse_double(value) { + Some(number) => + if number.is_nan() || number.abs() > 1.0e300 { + raise BigMafError("Non-finite number for " + label) + } else { + number + } + None => raise BigMafError("Invalid number for " + label + ": " + value) + } +} + +///| +fn bigmaf_non_gap_count(text : String) -> Int { + let mut count = 0 + for index in 0.. Unit raise BigMafError { + bigmaf_validate_token(text, "MAF aligned sequence") + for index in 0.. 126 { + bigmaf_fail("MAF aligned sequence contains a non-printable character") + } + } +} + +///| +fn bigmaf_insertion_status(status : String) -> Bool { + status == "C" || + status == "I" || + status == "N" || + status == "n" || + status == "M" || + status == "T" +} + +///| +fn bigmaf_empty_status(status : String) -> Bool { + status == "C" || status == "I" || status == "M" || status == "n" +} + +///| +fn bigmaf_source_parts(source : String) -> (String, String) raise BigMafError { + let mut dot = -1 + for index in 0..= source.length() { + bigmaf_fail( + "BigMaf reference component source must be '.'", + ) + } + (source[0:dot].to_owned(), source[dot + 1:source.length()].to_owned()) +} + +///| +pub fn BigMafInsertion::create( + left_status : String, + left_count : Int, + right_status : String, + right_count : Int, +) -> BigMafInsertion raise BigMafError { + if !bigmaf_insertion_status(left_status) || + !bigmaf_insertion_status(right_status) { + bigmaf_fail("MAF insertion status must be C, I, N, n, M, or T") + } + if left_count < 0 || right_count < 0 { + bigmaf_fail("MAF insertion counts must not be negative") + } + BigMafInsertion::{ left_status, left_count, right_status, right_count } +} + +///| +pub fn BigMafComponent::create( + source : String, + start : Int, + size : Int, + strand : String, + source_size : Int, + text : String, + insertion? : BigMafInsertion? = None, + quality? : String? = None, +) -> BigMafComponent raise BigMafError { + bigmaf_validate_token(source, "MAF source") + if start < 0 || size <= 0 || source_size <= 0 || start + size > source_size { + bigmaf_fail("MAF component coordinates exceed the source sequence") + } + if strand != "+" && strand != "-" { + bigmaf_fail("MAF component strand must be '+' or '-'") + } + bigmaf_validate_alignment_text(text) + if bigmaf_non_gap_count(text) != size { + bigmaf_fail("MAF component size does not match its non-gap sequence length") + } + match quality { + Some(value) => { + if value.length() != text.length() { + bigmaf_fail("MAF quality string must match the aligned sequence width") + } + for index in 0..= '0'.to_int() && quality_code <= '9'.to_int() + ) || + quality_code == 'F'.to_int() + if sequence_gap != quality_gap || (!quality_gap && !valid_score) { + bigmaf_fail( + "MAF quality must use 0-9/F and preserve aligned gap columns", + ) + } + } + } + None => () + } + match insertion { + Some(value) => + ignore( + BigMafInsertion::create( + value.left_status, + value.left_count, + value.right_status, + value.right_count, + ), + ) + None => () + } + BigMafComponent::{ + source, + start, + size, + strand, + source_size, + text, + insertion, + quality, + } +} + +///| +pub fn BigMafComponent::with_insertion( + self : BigMafComponent, + insertion : BigMafInsertion, +) -> BigMafComponent raise BigMafError { + BigMafComponent::create( + self.source, + self.start, + self.size, + self.strand, + self.source_size, + self.text, + insertion=Some(insertion), + quality=self.quality, + ) +} + +///| +pub fn BigMafComponent::with_quality( + self : BigMafComponent, + quality : String, +) -> BigMafComponent raise BigMafError { + BigMafComponent::create( + self.source, + self.start, + self.size, + self.strand, + self.source_size, + self.text, + insertion=self.insertion, + quality=Some(quality), + ) +} + +///| +pub fn BigMafEmptyComponent::create( + source : String, + start : Int, + size : Int, + strand : String, + source_size : Int, + status : String, +) -> BigMafEmptyComponent raise BigMafError { + bigmaf_validate_token(source, "MAF empty component source") + if start < 0 || size < 0 || source_size <= 0 || start + size > source_size { + bigmaf_fail("MAF empty component coordinates exceed the source sequence") + } + if strand != "+" && strand != "-" { + bigmaf_fail("MAF empty component strand must be '+' or '-'") + } + if !bigmaf_empty_status(status) { + bigmaf_fail("MAF empty component status must be C, I, M, or n") + } + BigMafEmptyComponent::{ source, start, size, strand, source_size, status } +} + +///| +pub fn BigMafBlock::create( + components : Array[BigMafComponent], + score? : Double? = None, + pass_number? : Int? = None, + empty_components? : Array[BigMafEmptyComponent] = [], + comments? : Array[String] = [], +) -> BigMafBlock raise BigMafError { + if components.length() == 0 { + bigmaf_fail("BigMaf blocks require at least one aligned component") + } + match score { + Some(value) => + if value.is_nan() || value.abs() > 1.0e300 { + bigmaf_fail("BigMaf block score must be finite") + } + None => () + } + match pass_number { + Some(value) => + if value <= 0 { + bigmaf_fail("MAF pass annotation must be positive") + } + None => () + } + let width = components[0].text.length() + for component in components { + ignore( + BigMafComponent::create( + component.source, + component.start, + component.size, + component.strand, + component.source_size, + component.text, + insertion=component.insertion, + quality=component.quality, + ), + ) + if component.text.length() != width { + bigmaf_fail("All MAF components in a block must have equal width") + } + } + for empty in empty_components { + ignore( + BigMafEmptyComponent::create( + empty.source, + empty.start, + empty.size, + empty.strand, + empty.source_size, + empty.status, + ), + ) + } + for comment in comments { + bigmaf_validate_comment(comment) + } + BigMafBlock::{ + score, + pass_number, + components: components.copy(), + empty_components: empty_components.copy(), + comments: comments.copy(), + } +} + +///| +pub fn BigMafWriteConfig::create( + compress? : Bool = true, + block_size? : Int = 256, + items_per_slot? : Int = 512, +) -> BigMafWriteConfig raise BigMafError { + if block_size < 2 || block_size > 65535 { + bigmaf_fail("BigMaf blockSize must be between 2 and 65535") + } + if items_per_slot <= 0 || items_per_slot > 65535 { + bigmaf_fail("BigMaf itemsPerSlot must be between 1 and 65535") + } + BigMafWriteConfig::{ compress, block_size, items_per_slot } +} + +///| +pub fn BigMafWriteConfig::default() -> BigMafWriteConfig { + BigMafWriteConfig::{ compress: true, block_size: 256, items_per_slot: 512 } +} + +///| +pub fn BigMafComponent::sequence(self : BigMafComponent) -> String { + let output = StringBuilder::new() + for index in 0.. String { + let sequence = self.sequence() + if self.strand == "+" { + sequence + } else { + reverse_complement(sequence) + } +} + +///| +pub fn BigMafComponent::forward_interval(self : BigMafComponent) -> (Int, Int) { + if self.strand == "+" { + (self.start, self.start + self.size) + } else { + (self.source_size - self.start - self.size, self.source_size - self.start) + } +} + +///| +pub fn BigMafEmptyComponent::forward_interval( + self : BigMafEmptyComponent, +) -> (Int, Int) { + if self.strand == "+" { + (self.start, self.start + self.size) + } else { + (self.source_size - self.start - self.size, self.source_size - self.start) + } +} + +///| +pub fn BigMafComponent::column_to_source( + self : BigMafComponent, + column : Int, +) -> Int? raise BigMafError { + if column < 0 || column >= self.text.length() { + bigmaf_fail("MAF alignment column is out of range") + } + if self.text.unsafe_get(column).to_int() == '-'.to_int() { + return None + } + let mut offset = 0 + for index in 0.. Int? { + let (forward_start, forward_end) = self.forward_interval() + if position < forward_start || position >= forward_end { + return None + } + let wanted = if self.strand == "+" { + position - self.start + } else { + self.source_size - self.start - position - 1 + } + let mut offset = 0 + for column in 0.. Int { + self.components[0].text.length() +} + +///| +pub fn BigMafBlock::reference_component(self : BigMafBlock) -> BigMafComponent { + self.components[0] +} + +///| +pub fn BigMafBlock::component( + self : BigMafBlock, + source : String, +) -> BigMafComponent? { + for component in self.components { + if component.source == source { + return Some(component) + } + } + None +} + +///| +pub fn BigMafBlock::sources(self : BigMafBlock) -> Array[String] { + let sources : Array[String] = [] + for component in self.components { + sources.push(component.source) + } + sources +} + +///| +pub fn BigMafBlock::map_reference_to( + self : BigMafBlock, + reference_position : Int, + source : String, +) -> Int? { + match self.components[0].source_to_column(reference_position) { + Some(column) => + match self.component(source) { + Some(component) => + component.column_to_source(column) catch { + _ => None + } + None => None + } + None => None + } +} + +///| +pub fn BigMafBlock::pairwise_identity(self : BigMafBlock) -> Double { + let mut matches = 0 + let mut comparisons = 0 + for column in 0.. Unit { + output.write_string( + "s " + + component.source + + " " + + component.start.to_string() + + " " + + component.size.to_string() + + " " + + component.strand + + " " + + component.source_size.to_string() + + " " + + component.text + + delimiter, + ) + match component.insertion { + Some(insertion) => + output.write_string( + "i " + + component.source + + " " + + insertion.left_status + + " " + + insertion.left_count.to_string() + + " " + + insertion.right_status + + " " + + insertion.right_count.to_string() + + delimiter, + ) + None => () + } + match component.quality { + Some(quality) => + output.write_string("q " + component.source + " " + quality + delimiter) + None => () + } +} + +///| +fn bigmaf_format_block(block : BigMafBlock, delimiter : String) -> String { + let output = StringBuilder::new() + for comment in block.comments { + output.write_string("# " + comment + delimiter) + } + output.write_string("a") + match block.score { + Some(score) => output.write_string(" score=" + score.to_string()) + None => () + } + match block.pass_number { + Some(pass_number) => output.write_string(" pass=" + pass_number.to_string()) + None => () + } + output.write_string(delimiter) + for component in block.components { + bigmaf_write_component(output, component, delimiter) + } + for empty in block.empty_components { + output.write_string( + "e " + + empty.source + + " " + + empty.start.to_string() + + " " + + empty.size.to_string() + + " " + + empty.strand + + " " + + empty.source_size.to_string() + + " " + + empty.status + + delimiter, + ) + } + output.to_string() +} + +///| +pub fn BigMafBlock::to_embedded_maf(self : BigMafBlock) -> String { + bigmaf_format_block(self, ";") +} + +///| +pub fn BigMafBlock::to_maf(self : BigMafBlock) -> String { + bigmaf_format_block(self, "\n") + "\n" +} + +///| +fn bigmaf_parse_score_line( + words : Array[String], +) -> (Double?, Int?) raise BigMafError { + let mut score : Double? = None + let mut pass_number : Int? = None + for index in 1.. BigMafBlock raise BigMafError { + if text.length() == 0 { + bigmaf_fail("BigMaf block text must not be empty") + } + let lines = text.split(";").to_array() + let components : Array[BigMafComponent] = [] + let empty_components : Array[BigMafEmptyComponent] = [] + let comments : Array[String] = [] + let mut score : Double? = None + let mut pass_number : Int? = None + let mut saw_alignment = false + let mut last_component = -1 + for raw in lines { + let line = trim(raw.to_owned()) + if line.length() == 0 { + continue + } + if line.unsafe_get(0).to_int() == '#'.to_int() { + comments.push(trim(line[1:line.length()].to_owned())) + continue + } + let words = bigmaf_words(line) + if words.length() == 0 { + continue + } + if words[0] == "a" { + if saw_alignment { + bigmaf_fail("BigMaf payload contains more than one MAF block") + } + let (found_score, found_pass) = bigmaf_parse_score_line(words) + score = found_score + pass_number = found_pass + saw_alignment = true + last_component = -1 + } else if words[0] == "s" { + if !saw_alignment { + bigmaf_fail("MAF sequence line appears before the alignment line") + } + if words.length() != 7 { + bigmaf_fail("MAF sequence lines must contain seven fields") + } + let component = BigMafComponent::create( + words[1], + bigmaf_parse_int(words[2], "MAF component start"), + bigmaf_parse_int(words[3], "MAF component size"), + words[4], + bigmaf_parse_int(words[5], "MAF source size"), + words[6], + ) + components.push(component) + last_component = components.length() - 1 + } else if words[0] == "i" { + if last_component < 0 || words.length() != 6 { + bigmaf_fail( + "MAF insertion lines must follow an s line and have six fields", + ) + } + if words[1] != components[last_component].source { + bigmaf_fail("MAF insertion source does not match the preceding s line") + } + if components[last_component].insertion is Some(_) { + bigmaf_fail("Duplicate MAF insertion line for one component") + } + let insertion = BigMafInsertion::create( + words[2], + bigmaf_parse_int(words[3], "MAF left insertion count"), + words[4], + bigmaf_parse_int(words[5], "MAF right insertion count"), + ) + components[last_component] = components[last_component].with_insertion( + insertion, + ) + } else if words[0] == "q" { + if last_component < 0 || words.length() != 3 { + bigmaf_fail( + "MAF quality lines must follow an s line and have three fields", + ) + } + if words[1] != components[last_component].source { + bigmaf_fail("MAF quality source does not match the preceding s line") + } + if components[last_component].quality is Some(_) { + bigmaf_fail("Duplicate MAF quality line for one component") + } + components[last_component] = components[last_component].with_quality( + words[2], + ) + } else if words[0] == "e" { + if !saw_alignment || words.length() != 7 { + bigmaf_fail("MAF empty component lines must contain seven fields") + } + empty_components.push( + BigMafEmptyComponent::create( + words[1], + bigmaf_parse_int(words[2], "MAF empty component start"), + bigmaf_parse_int(words[3], "MAF empty component size"), + words[4], + bigmaf_parse_int(words[5], "MAF empty source size"), + words[6], + ), + ) + last_component = -1 + } else { + bigmaf_fail("Unexpected MAF line type '" + words[0] + "'") + } + } + if !saw_alignment { + bigmaf_fail("BigMaf payload is missing its MAF alignment line") + } + BigMafBlock::create( + components, + score~, + pass_number~, + empty_components~, + comments~, + ) +} + +///| +pub fn bigmaf_schema() -> BigBedSchema { + BigBedSchema::{ + name: "bedMaf", + comment: "Bed3 with MAF block", + fields: [ + BigBedField::{ + as_type: "string", + name: "chrom", + comment: "Reference sequence chromosome or scaffold", + }, + BigBedField::{ + as_type: "uint", + name: "chromStart", + comment: "Start position in chromosome", + }, + BigBedField::{ + as_type: "uint", + name: "chromEnd", + comment: "End position in chromosome", + }, + BigBedField::{ + as_type: "lstring", + name: "mafBlock", + comment: "MAF block", + }, + ], + } +} + +///| +fn bigmaf_target_index( + targets : Array[BigBedTarget], + chromosome : String, +) -> Int { + for index in 0.. Array[Int] raise BigBedError, +) -> Array[Int] raise BigMafError { + operation() catch { + BigBedError(message) => raise BigMafError("BigBed: " + message) + } +} + +///| +pub fn bigmaf_write( + reference : String, + targets : Array[BigBedTarget], + blocks : Array[BigMafBlock], + config? : BigMafWriteConfig = BigMafWriteConfig::default(), +) -> Array[Int] raise BigMafError { + bigmaf_validate_token(reference, "BigMaf reference") + for index in 0.. bigmaf_fail("BigBed: " + message) + } + for previous in 0.. { + bigmaf_fail("BigBed: " + message) + BigBedRecord::{ + chrom: "", + chrom_start: 0, + chrom_end: 1, + name: ".", + score: 0, + strand: ".", + thick_start: 0, + thick_end: 1, + item_rgb: "0", + declared_block_count: 0, + block_sizes: [], + block_starts: [], + extra_fields: [], + } + } + } + records.push(record) + target_order.push(target_index) + } + let order : Array[Int] = [] + for index in 0.. Int { + if target_order[left] < target_order[right] { + -1 + } else if target_order[left] > target_order[right] { + 1 + } else if records[left].chrom_start < records[right].chrom_start { + -1 + } else if records[left].chrom_start > records[right].chrom_start { + 1 + } else if records[left].chrom_end < records[right].chrom_end { + -1 + } else if records[left].chrom_end > records[right].chrom_end { + 1 + } else { + 0 + } + }) + let sorted : Array[BigBedRecord] = [] + for index in order { + sorted.push(records[index]) + } + let field = BigBedField::{ + as_type: "lstring", + name: "mafBlock", + comment: "MAF block", + } + let bed_config = BigBedWriteConfig::create( + bed_columns=3, + items_per_slot=config.items_per_slot, + block_size=config.block_size, + compress=config.compress, + schema_name="bedMaf", + schema_comment="Bed3 with MAF block", + custom_fields=[field], + ) catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + BigBedWriteConfig::default() + } + } + bigmaf_wrap_bigbed(fn() -> Array[Int] raise BigBedError { + bigbed_write(targets, sorted, config=bed_config) + }) +} + +///| +fn bigmaf_validate_schema(bed : BigBedFile) -> Unit raise BigMafError { + if bed.header.defined_field_count != 3 || + bed.header.field_count != 4 || + bed.schema.name != "bedMaf" || + bed.schema.fields.length() != 4 || + bed.schema.fields[3].name != "mafBlock" || + bed.schema.fields[3].as_type != "lstring" { + bigmaf_fail( + "BigMaf requires the standard bedMaf bed3+1 AutoSQL declaration", + ) + } +} + +///| +fn bigmaf_decode_record( + bed : BigBedFile, + record : BigBedRecord, + expected_reference : String, +) -> (String, BigMafBlock) raise BigMafError { + if record.extra_fields.length() != 1 { + bigmaf_fail("BigMaf binary records require one mafBlock field") + } + let block = bigmaf_parse_block(record.extra_fields[0]) + let component = block.components[0] + let (reference, chromosome) = bigmaf_source_parts(component.source) + if expected_reference.length() > 0 && reference != expected_reference { + bigmaf_fail("BigMaf records use inconsistent reference prefixes") + } + if component.strand != "+" { + bigmaf_fail("BigMaf reference components must use the forward strand") + } + if chromosome != record.chrom || + component.start != record.chrom_start || + component.start + component.size != record.chrom_end { + bigmaf_fail("BigMaf BED coordinates disagree with the embedded MAF block") + } + match bed.target(chromosome) { + Some(target) => + if component.source_size != target.length { + bigmaf_fail( + "BigMaf reference component source size disagrees with its target", + ) + } + None => bigmaf_fail("BigMaf record refers to an unknown target") + } + (reference, block) +} + +///| +pub fn bigmaf_parse(data : Array[Int]) -> BigMafFile raise BigMafError { + let bed = bigbed_parse(data) catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + BigBedFile::{ + header: BigBedHeader::{ + little_endian: true, + version: 4, + zoom_levels: 0, + chromosome_tree_offset: 0, + full_data_offset: 0, + full_index_offset: 0, + field_count: 0, + defined_field_count: 0, + auto_sql_offset: 0, + total_summary_offset: 0, + uncompress_buffer_size: 0, + extra_indices_offset: 0, + }, + schema: bigmaf_schema(), + targets: [], + source: [], + blocks: [], + record_count: 0, + } + } + } + bigmaf_validate_schema(bed) + let records = bed.records() catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + [] + } + } + let mut reference = "" + for record in records { + let (found, _) = bigmaf_decode_record(bed, record, reference) + if reference.length() == 0 { + reference = found + } + } + BigMafFile::{ reference, bed } +} + +///| +pub fn BigMafFile::blocks( + self : BigMafFile, +) -> Array[BigMafBlock] raise BigMafError { + let records = self.bed.records() catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + [] + } + } + let blocks : Array[BigMafBlock] = [] + for record in records { + let (_, block) = bigmaf_decode_record(self.bed, record, self.reference) + blocks.push(block) + } + blocks +} + +///| +fn bigmaf_query_chromosome( + file : BigMafFile, + chromosome : String, +) -> String raise BigMafError { + if file.reference.length() == 0 { + return chromosome + } + let prefix = file.reference + "." + if chromosome.has_prefix(prefix) { + let suffix = chromosome[prefix.length():chromosome.length()].to_owned() + if suffix.length() == 0 { + bigmaf_fail("BigMaf query chromosome must not be empty") + } + suffix + } else { + chromosome + } +} + +///| +pub fn BigMafFile::search( + self : BigMafFile, + chromosome : String, + start? : Int = 0, + end? : Int = -1, +) -> Array[BigMafBlock] raise BigMafError { + let resolved = bigmaf_query_chromosome(self, chromosome) + let records = self.bed.search(resolved, start~, end~) catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + [] + } + } + let blocks : Array[BigMafBlock] = [] + for record in records { + let (_, block) = bigmaf_decode_record(self.bed, record, self.reference) + blocks.push(block) + } + blocks +} + +///| +pub fn BigMafFile::target( + self : BigMafFile, + chromosome : String, +) -> BigBedTarget? { + let resolved = bigmaf_query_chromosome(self, chromosome) catch { + _ => return None + } + self.bed.target(resolved) +} + +///| +pub fn BigMafFile::is_compressed(self : BigMafFile) -> Bool { + self.bed.is_compressed() +} + +///| +pub fn BigMafFile::to_maf(self : BigMafFile) -> String raise BigMafError { + let output = StringBuilder::new() + output.write_string("##maf version=1\n\n") + for block in self.blocks() { + output.write_string(block.to_maf()) + } + output.to_string() +} + +///| +pub fn BigMafFile::summary( + self : BigMafFile, +) -> BigMafSummary raise BigMafError { + let blocks = self.blocks() + let mut component_count = 0 + let mut empty_component_count = 0 + let mut aligned_columns = 0 + let mut covered_reference_bases = 0 + for block in blocks { + component_count = component_count + block.components.length() + empty_component_count = empty_component_count + + block.empty_components.length() + aligned_columns = aligned_columns + block.aligned_columns() + covered_reference_bases = covered_reference_bases + block.components[0].size + } + BigMafSummary::{ + alignment_count: blocks.length(), + target_count: self.bed.targets.length(), + component_count, + empty_component_count, + aligned_columns, + covered_reference_bases, + compressed: self.is_compressed(), + } +} + +///| +pub fn bigmaf_example_data() -> ( + String, + Array[BigBedTarget], + Array[BigMafBlock], +) raise BigMafError { + let reference = "hg38" + let targets = [ + BigBedTarget::create("chr7", 1000) catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + BigBedTarget::{ name: "chr7", length: 1000 } + } + }, + BigBedTarget::create("chr12", 800) catch { + BigBedError(message) => { + bigmaf_fail("BigBed: " + message) + BigBedTarget::{ name: "chr12", length: 800 } + } + }, + ] + let reference_one = BigMafComponent::create( + "hg38.chr7", + 100, + 8, + "+", + 1000, + "ACGT-ACGT", + quality=Some("9999-FFFF"), + ) + let chimp_one = BigMafComponent::create( + "panTro6.chr6", 200, 8, "+", 900, "ACGT-ACGT", + ) + let mouse_one = BigMafComponent::create( + "mm39.chr5", + 50, + 8, + "-", + 700, + "AC-TTACGT", + insertion=Some(BigMafInsertion::create("C", 0, "I", 12)), + ) + let empty = BigMafEmptyComponent::create( + "canFam6.chr6", 300, 20, "+", 850, "M", + ) + let first = BigMafBlock::create( + [reference_one, chimp_one, mouse_one], + score=Some(23262.0), + pass_number=Some(1), + empty_components=[empty], + comments=["conserved reference block"], + ) + let second = BigMafBlock::create( + [ + BigMafComponent::create("hg38.chr7", 300, 6, "+", 1000, "TTAAGA"), + BigMafComponent::create("panTro6.chr6", 410, 6, "+", 900, "TTAAGA"), + BigMafComponent::create("mm39.chr5", 120, 5, "+", 700, "TT-AGA"), + ], + score=Some(5062.0), + ) + let third = BigMafBlock::create( + [ + BigMafComponent::create("hg38.chr12", 40, 7, "+", 800, "GCA-GCTG"), + BigMafComponent::create("panTro6.chr10", 70, 7, "+", 600, "GCA-GCTG"), + ], + score=Some(6636.0), + ) + (reference, targets, [first, second, third]) +} diff --git a/src/bigpsl.mbt b/src/bigpsl.mbt new file mode 100644 index 00000000..8c31b1bd --- /dev/null +++ b/src/bigpsl.mbt @@ -0,0 +1,1672 @@ +// Bio.Align.bigpsl - indexed UCSC BigPsl pairwise alignments. +// +// BigPsl is a bed12+13 BigBed specialization. The BigBed container owns the +// binary layout and indices; this module owns pairwise-coordinate semantics, +// translated DNA-to-protein blocks, match categories, and strict schema +// validation. + +///| +pub suberror BigPslError { + BigPslError(String) +} + +///| +pub(all) enum BigPslSequenceType { + BigPslUnknown + BigPslNucleotide + BigPslProtein +} derive(Eq, Debug) + +///| +pub(all) enum BigPslMaskMode { + BigPslNoMask + BigPslMaskLower + BigPslMaskUpper +} derive(Eq, Debug) + +///| +pub struct BigPslBlock { + target_start : Int + target_end : Int + query_start : Int + query_end : Int + target_size : Int + query_size : Int +} derive(Eq, Debug) + +///| +pub struct BigPslAlignment { + target_name : String + target_size : Int + query_name : String + query_size : Int + target_coordinates : Array[Int] + query_coordinates : Array[Int] + query_sequence : String + cds : String + score : Int + thick_start : Int + thick_end : Int + item_rgb : String + matches : Int + mismatches : Int + repeat_matches : Int + n_count : Int + sequence_type : BigPslSequenceType +} derive(Eq, Debug) + +///| +pub struct BigPslWriteConfig { + compress : Bool + block_size : Int + items_per_slot : Int + store_query_sequence : Bool + store_cds : Bool +} derive(Eq, Debug) + +///| +pub struct BigPslAlignmentCounts { + aligned_query_units : Int + query_insert_units : Int + target_insert_bases : Int + query_insert_events : Int + target_insert_events : Int + block_count : Int +} derive(Eq, Debug) + +///| +pub struct BigPslSummary { + alignment_count : Int + target_count : Int + nucleotide_count : Int + protein_count : Int + unknown_count : Int + aligned_query_units : Int + covered_target_bases : Int + compressed : Bool +} derive(Eq, Debug) + +///| +pub struct BigPslFile { + bed : BigBedFile +} + +///| +priv struct BigPslStorage { + chrom_start : Int + chrom_end : Int + strand : String + other_start : Int + other_end : Int + other_strand : String + target_block_sizes : Array[Int] + target_block_starts : Array[Int] + other_block_starts : Array[Int] +} + +///| +fn bigpsl_fail(message : String) -> Unit raise BigPslError { + raise BigPslError(message) +} + +///| +fn bigpsl_abs(value : Int) -> Int { + if value < 0 { + -value + } else { + value + } +} + +///| +fn bigpsl_min(left : Int, right : Int) -> Int { + if left < right { + left + } else { + right + } +} + +///| +fn bigpsl_max(left : Int, right : Int) -> Int { + if left > right { + left + } else { + right + } +} + +///| +fn bigpsl_copy_ints(values : Array[Int]) -> Array[Int] { + let copy : Array[Int] = [] + for value in values { + copy.push(value) + } + copy +} + +///| +fn bigpsl_reverse_ints(values : Array[Int]) -> Array[Int] { + let reversed : Array[Int] = [] + let mut index = values.length() + while index > 0 { + index = index - 1 + reversed.push(values[index]) + } + reversed +} + +///| +fn bigpsl_validate_text( + value : String, + label : String, + allow_empty : Bool, +) -> Unit raise BigPslError { + if !allow_empty && value.length() == 0 { + bigpsl_fail(label + " must not be empty") + } + for index in 0.. Unit raise BigPslError { + if sequence.length() == 0 { + return + } + if sequence.length() != expected_size { + bigpsl_fail(label + " length does not match its declaration") + } + for index in 0.. Int { + match sequence_type { + BigPslUnknown => 0 + BigPslNucleotide => 1 + BigPslProtein => 2 + } +} + +///| +fn bigpsl_type_from_code(code : Int) -> BigPslSequenceType raise BigPslError { + match code { + 0 => BigPslUnknown + 1 => BigPslNucleotide + 2 => BigPslProtein + _ => { + bigpsl_fail("BigPsl seqType must be 0, 1, or 2") + BigPslUnknown + } + } +} + +///| +fn bigpsl_axis_direction( + coordinates : Array[Int], + label : String, +) -> Int raise BigPslError { + let mut direction = 0 + for index in 0..<(coordinates.length() - 1) { + let step = coordinates[index + 1] - coordinates[index] + if step > 0 { + if direction < 0 { + bigpsl_fail(label + " coordinates change direction") + } + direction = 1 + } else if step < 0 { + if direction > 0 { + bigpsl_fail(label + " coordinates change direction") + } + direction = -1 + } + } + direction +} + +///| +fn bigpsl_validate_coordinates( + target_size : Int, + query_size : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], + sequence_type : BigPslSequenceType, +) -> Unit raise BigPslError { + if target_size <= 0 || query_size <= 0 { + bigpsl_fail("BigPsl target and query sizes must be positive") + } + if target_coordinates.length() != query_coordinates.length() { + bigpsl_fail( + "BigPsl target and query coordinate arrays must have equal length", + ) + } + if target_coordinates.length() < 2 { + bigpsl_fail("BigPsl alignment paths require at least two coordinate points") + } + for index in 0.. target_size { + bigpsl_fail("BigPsl target coordinate is out of bounds") + } + if query_coordinates[index] < 0 || query_coordinates[index] > query_size { + bigpsl_fail("BigPsl query coordinate is out of bounds") + } + } + let target_direction = bigpsl_axis_direction( + target_coordinates, "BigPsl target", + ) + let query_direction = bigpsl_axis_direction(query_coordinates, "BigPsl query") + let mut aligned_blocks = 0 + for index in 0..<(target_coordinates.length() - 1) { + let target_step = target_coordinates[index + 1] - target_coordinates[index] + let query_step = query_coordinates[index + 1] - query_coordinates[index] + if target_step == 0 && query_step == 0 { + bigpsl_fail("Consecutive BigPsl coordinate points must differ") + } + if target_step != 0 && query_step != 0 { + let target_count = bigpsl_abs(target_step) + let query_count = bigpsl_abs(query_step) + match sequence_type { + BigPslProtein => + if target_count != 3 * query_count { + bigpsl_fail( + "Protein BigPsl aligned target blocks must be three times their query blocks", + ) + } + _ => + if target_count != query_count { + bigpsl_fail( + "Nucleotide BigPsl aligned target and query blocks must have equal size", + ) + } + } + aligned_blocks = aligned_blocks + 1 + } + } + if aligned_blocks == 0 { + bigpsl_fail("BigPsl alignment must contain at least one aligned block") + } + match sequence_type { + BigPslProtein => { + if query_direction != 1 { + bigpsl_fail("Protein BigPsl query coordinates must increase") + } + if target_direction == 0 { + bigpsl_fail("Protein BigPsl target coordinates must have a direction") + } + } + _ => { + if target_direction != 1 { + bigpsl_fail("Nucleotide BigPsl target coordinates must increase") + } + if query_direction == 0 { + bigpsl_fail("Nucleotide BigPsl query coordinates must have a direction") + } + } + } +} + +///| +fn bigpsl_raw_blocks( + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> Array[BigPslBlock] { + let blocks : Array[BigPslBlock] = [] + for index in 0..<(target_coordinates.length() - 1) { + let target_start = target_coordinates[index] + let target_end = target_coordinates[index + 1] + let query_start = query_coordinates[index] + let query_end = query_coordinates[index + 1] + let target_count = bigpsl_abs(target_end - target_start) + let query_count = bigpsl_abs(query_end - query_start) + if target_count > 0 && query_count > 0 { + blocks.push(BigPslBlock::{ + target_start, + target_end, + query_start, + query_end, + target_size: target_count, + query_size: query_count, + }) + } + } + blocks +} + +///| +fn bigpsl_aligned_query_units( + target_coordinates : Array[Int], + query_coordinates : Array[Int], +) -> Int { + let mut total = 0 + for block in bigpsl_raw_blocks(target_coordinates, query_coordinates) { + total = total + block.query_size + } + total +} + +///| +fn bigpsl_coordinate_interval(coordinates : Array[Int]) -> (Int, Int) { + let mut start = coordinates[0] + let mut end = coordinates[0] + for coordinate in coordinates { + start = bigpsl_min(start, coordinate) + end = bigpsl_max(end, coordinate) + } + (start, end) +} + +///| +pub fn BigPslAlignment::create( + target_name : String, + target_size : Int, + query_name : String, + query_size : Int, + target_coordinates : Array[Int], + query_coordinates : Array[Int], + sequence_type? : BigPslSequenceType = BigPslUnknown, + query_sequence? : String = "", + cds? : String = "", + score? : Int = 0, + thick_start? : Int = -1, + thick_end? : Int = -1, + item_rgb? : String = "0", + matches? : Int = -1, + mismatches? : Int = -1, + repeat_matches? : Int = -1, + n_count? : Int = -1, +) -> BigPslAlignment raise BigPslError { + bigpsl_validate_text(target_name, "BigPsl target name", false) + bigpsl_validate_text(query_name, "BigPsl query name", false) + bigpsl_validate_text(item_rgb, "BigPsl itemRgb", false) + bigpsl_validate_text(cds, "BigPsl CDS", true) + bigpsl_validate_sequence(query_sequence, query_size, "BigPsl query sequence") + bigpsl_validate_coordinates( + target_size, query_size, target_coordinates, query_coordinates, sequence_type, + ) + if score < 0 || score > 1000 { + bigpsl_fail("BigPsl score must be between 0 and 1000") + } + let (target_minimum, target_maximum) = bigpsl_coordinate_interval( + target_coordinates, + ) + let normalized_thick_start = if thick_start < 0 { + target_minimum + } else { + thick_start + } + let normalized_thick_end = if thick_end < 0 { + target_maximum + } else { + thick_end + } + if normalized_thick_start < target_minimum || + normalized_thick_start > normalized_thick_end || + normalized_thick_end > target_maximum { + bigpsl_fail("BigPsl thick interval must lie inside its target interval") + } + let aligned = bigpsl_aligned_query_units( + target_coordinates, query_coordinates, + ) + let all_unspecified = matches < 0 && + mismatches < 0 && + repeat_matches < 0 && + n_count < 0 + let normalized_matches = if all_unspecified { aligned } else { matches } + let normalized_mismatches = if all_unspecified { 0 } else { mismatches } + let normalized_repeat = if all_unspecified { 0 } else { repeat_matches } + let normalized_n = if all_unspecified { 0 } else { n_count } + if normalized_matches < 0 || + normalized_mismatches < 0 || + normalized_repeat < 0 || + normalized_n < 0 { + bigpsl_fail( + "BigPsl match categories must all be supplied or all be omitted", + ) + } + if normalized_matches + + normalized_mismatches + + normalized_repeat + + normalized_n != + aligned { + bigpsl_fail("BigPsl match categories must sum to the aligned query length") + } + BigPslAlignment::{ + target_name, + target_size, + query_name, + query_size, + target_coordinates: bigpsl_copy_ints(target_coordinates), + query_coordinates: bigpsl_copy_ints(query_coordinates), + query_sequence, + cds, + score, + thick_start: normalized_thick_start, + thick_end: normalized_thick_end, + item_rgb, + matches: normalized_matches, + mismatches: normalized_mismatches, + repeat_matches: normalized_repeat, + n_count: normalized_n, + sequence_type, + } +} + +///| +pub fn BigPslWriteConfig::create( + compress? : Bool = true, + block_size? : Int = 256, + items_per_slot? : Int = 512, + store_query_sequence? : Bool = false, + store_cds? : Bool = false, +) -> BigPslWriteConfig raise BigPslError { + if block_size < 2 || block_size > 65535 { + bigpsl_fail("BigPsl blockSize must be between 2 and 65535") + } + if items_per_slot <= 0 || items_per_slot > 65535 { + bigpsl_fail("BigPsl itemsPerSlot must be between 1 and 65535") + } + BigPslWriteConfig::{ + compress, + block_size, + items_per_slot, + store_query_sequence, + store_cds, + } +} + +///| +pub fn BigPslWriteConfig::default() -> BigPslWriteConfig { + BigPslWriteConfig::{ + compress: true, + block_size: 256, + items_per_slot: 512, + store_query_sequence: false, + store_cds: false, + } +} + +///| +pub fn BigPslAlignment::blocks(self : BigPslAlignment) -> Array[BigPslBlock] { + bigpsl_raw_blocks(self.target_coordinates, self.query_coordinates) +} + +///| +pub fn BigPslAlignment::target_interval(self : BigPslAlignment) -> (Int, Int) { + bigpsl_coordinate_interval(self.target_coordinates) +} + +///| +pub fn BigPslAlignment::query_interval(self : BigPslAlignment) -> (Int, Int) { + bigpsl_coordinate_interval(self.query_coordinates) +} + +///| +pub fn BigPslAlignment::is_reverse(self : BigPslAlignment) -> Bool { + let last = self.target_coordinates.length() - 1 + self.target_coordinates[0] > self.target_coordinates[last] || + self.query_coordinates[0] > self.query_coordinates[last] +} + +///| +pub fn BigPslAlignment::counts(self : BigPslAlignment) -> BigPslAlignmentCounts { + let mut aligned_query_units = 0 + let mut query_insert_units = 0 + let mut target_insert_bases = 0 + let mut query_insert_events = 0 + let mut target_insert_events = 0 + let mut block_count = 0 + for index in 0..<(self.target_coordinates.length() - 1) { + let target_step = bigpsl_abs( + self.target_coordinates[index + 1] - self.target_coordinates[index], + ) + let query_step = bigpsl_abs( + self.query_coordinates[index + 1] - self.query_coordinates[index], + ) + if target_step > 0 && query_step > 0 { + aligned_query_units = aligned_query_units + query_step + block_count = block_count + 1 + } else if target_step == 0 { + query_insert_units = query_insert_units + query_step + query_insert_events = query_insert_events + 1 + } else { + target_insert_bases = target_insert_bases + target_step + target_insert_events = target_insert_events + 1 + } + } + BigPslAlignmentCounts::{ + aligned_query_units, + query_insert_units, + target_insert_bases, + query_insert_events, + target_insert_events, + block_count, + } +} + +///| +pub fn BigPslAlignment::target_to_query( + self : BigPslAlignment, + position : Int, +) -> Int? { + if position < 0 || position >= self.target_size { + return None + } + for block in self.blocks() { + let minimum = bigpsl_min(block.target_start, block.target_end) + let maximum = bigpsl_max(block.target_start, block.target_end) + if position >= minimum && position < maximum { + let offset = if block.target_end > block.target_start { + position - block.target_start + } else { + block.target_start - 1 - position + } + let query_offset = match self.sequence_type { + BigPslProtein => offset / 3 + _ => offset + } + return Some( + if block.query_end > block.query_start { + block.query_start + query_offset + } else { + block.query_start - 1 - query_offset + }, + ) + } + } + None +} + +///| +pub fn BigPslAlignment::query_to_target_interval( + self : BigPslAlignment, + position : Int, +) -> (Int, Int)? { + if position < 0 || position >= self.query_size { + return None + } + for block in self.blocks() { + let minimum = bigpsl_min(block.query_start, block.query_end) + let maximum = bigpsl_max(block.query_start, block.query_end) + if position >= minimum && position < maximum { + let offset = if block.query_end > block.query_start { + position - block.query_start + } else { + block.query_start - 1 - position + } + let scale = match self.sequence_type { + BigPslProtein => 3 + _ => 1 + } + let oriented = offset * scale + let first = if block.target_end > block.target_start { + block.target_start + oriented + } else { + block.target_start - oriented - scale + } + return Some( + (bigpsl_min(first, first + scale), bigpsl_max(first, first + scale)), + ) + } + } + None +} + +///| +fn bigpsl_upper_code(code : Int) -> Int { + if code >= 'a'.to_int() && code <= 'z'.to_int() { + code - 32 + } else { + code + } +} + +///| +fn bigpsl_is_lower(code : Int) -> Bool { + code >= 'a'.to_int() && code <= 'z'.to_int() +} + +///| +fn bigpsl_is_upper(code : Int) -> Bool { + code >= 'A'.to_int() && code <= 'Z'.to_int() +} + +///| +fn bigpsl_complement_code(code : Int) -> Int { + match bigpsl_upper_code(code) { + 65 => 84 // A -> T + 67 => 71 // C -> G + 71 => 67 // G -> C + 84 => 65 // T -> A + 85 => 65 // U -> A + 82 => 89 // R -> Y + 89 => 82 // Y -> R + 77 => 75 // M -> K + 75 => 77 // K -> M + 66 => 86 // B -> V + 86 => 66 // V -> B + 68 => 72 // D -> H + 72 => 68 // H -> D + value => value + } +} + +///| +pub fn BigPslAlignment::recount( + self : BigPslAlignment, + target_sequence : String, + query_sequence : String, + mask? : BigPslMaskMode = BigPslNoMask, + wildcard? : Char = 'N', +) -> BigPslAlignment raise BigPslError { + bigpsl_validate_sequence( + target_sequence, + self.target_size, + "BigPsl target sequence", + ) + bigpsl_validate_sequence( + query_sequence, + self.query_size, + "BigPsl query sequence", + ) + if target_sequence.length() == 0 || query_sequence.length() == 0 { + bigpsl_fail("BigPsl recount requires concrete target and query sequences") + } + if self.sequence_type == BigPslProtein && mask != BigPslNoMask { + bigpsl_fail("Repeat masking is only defined for nucleotide BigPsl records") + } + let wildcard_code = bigpsl_upper_code(wildcard.to_int()) + let mut matches = 0 + let mut mismatches = 0 + let mut repeat_matches = 0 + let mut n_count = 0 + for block in self.blocks() { + match self.sequence_type { + BigPslProtein => { + let start = bigpsl_min(block.target_start, block.target_end) + let end = bigpsl_max(block.target_start, block.target_end) + let dna = target_sequence[start:end].to_owned() + let oriented = if block.target_end < block.target_start { + Seq::new(dna).reverse_complement() + } else { + Seq::new(dna) + } + let translated = oriented.translate() catch { + SeqError(message) => { + bigpsl_fail("Cannot translate BigPsl target block: " + message) + Seq::new("") + } + } + let amino = translated.to_string() + if amino.length() != block.query_size { + bigpsl_fail("Translated BigPsl target block has an unexpected length") + } + for offset in 0.. + for offset in 0.. block.query_start { + block.query_start + offset + } else { + block.query_start - 1 - offset + } + let raw_target = target_sequence.unsafe_get(target_index).to_int() + let target_code = bigpsl_upper_code(raw_target) + let raw_query = query_sequence.unsafe_get(query_index).to_int() + let query_code = if block.query_end < block.query_start { + bigpsl_complement_code(raw_query) + } else { + bigpsl_upper_code(raw_query) + } + if target_code == wildcard_code || query_code == wildcard_code { + n_count = n_count + 1 + } else if target_code != query_code { + mismatches = mismatches + 1 + } else { + let masked = match mask { + BigPslMaskLower => bigpsl_is_lower(raw_target) + BigPslMaskUpper => bigpsl_is_upper(raw_target) + BigPslNoMask => false + } + if masked { + repeat_matches = repeat_matches + 1 + } else { + matches = matches + 1 + } + } + } + } + } + BigPslAlignment::create( + self.target_name, + self.target_size, + self.query_name, + self.query_size, + self.target_coordinates, + self.query_coordinates, + sequence_type=self.sequence_type, + query_sequence~, + cds=self.cds, + score=self.score, + thick_start=self.thick_start, + thick_end=self.thick_end, + item_rgb=self.item_rgb, + matches~, + mismatches~, + repeat_matches~, + n_count~, + ) +} + +///| +fn bigpsl_storage(alignment : BigPslAlignment) -> BigPslStorage { + let blocks = alignment.blocks() + let target_reverse = alignment.target_coordinates[0] > + alignment.target_coordinates[alignment.target_coordinates.length() - 1] + let query_reverse = alignment.query_coordinates[0] > + alignment.query_coordinates[alignment.query_coordinates.length() - 1] + let strand = if target_reverse || query_reverse { "-" } else { "+" } + let other_strand = if target_reverse { "-" } else { "+" } + let ordered : Array[BigPslBlock] = [] + if target_reverse { + let mut index = blocks.length() + while index > 0 { + index = index - 1 + ordered.push(blocks[index]) + } + } else { + for block in blocks { + ordered.push(block) + } + } + let mut chrom_start = alignment.target_size + let mut chrom_end = 0 + let mut other_start = alignment.query_size + let mut other_end = 0 + for block in ordered { + chrom_start = bigpsl_min( + chrom_start, + bigpsl_min(block.target_start, block.target_end), + ) + chrom_end = bigpsl_max( + chrom_end, + bigpsl_max(block.target_start, block.target_end), + ) + other_start = bigpsl_min( + other_start, + bigpsl_min(block.query_start, block.query_end), + ) + other_end = bigpsl_max( + other_end, + bigpsl_max(block.query_start, block.query_end), + ) + } + let target_block_sizes : Array[Int] = [] + let target_block_starts : Array[Int] = [] + let other_block_starts : Array[Int] = [] + for block in ordered { + let target_minimum = bigpsl_min(block.target_start, block.target_end) + let query_minimum = bigpsl_min(block.query_start, block.query_end) + let query_maximum = bigpsl_max(block.query_start, block.query_end) + target_block_sizes.push(block.target_size) + target_block_starts.push(target_minimum - chrom_start) + other_block_starts.push( + if strand == "-" { + alignment.query_size - query_maximum + } else { + query_minimum + }, + ) + } + BigPslStorage::{ + chrom_start, + chrom_end, + strand, + other_start, + other_end, + other_strand, + target_block_sizes, + target_block_starts, + other_block_starts, + } +} + +///| +fn bigpsl_join_ints(values : Array[Int], trailing : Bool) -> String { + let output = StringBuilder::new() + for index in 0.. 0 { + output.write_char(',') + } + output.write_string(values[index].to_string()) + } + if trailing && values.length() > 0 { + output.write_char(',') + } + output.to_string() +} + +///| +pub fn bigpsl_schema() -> BigBedSchema { + let standard_names = [ + "chrom", "chromStart", "chromEnd", "name", "score", "strand", "thickStart", "thickEnd", + "reserved", "blockCount", "blockSizes", "chromStarts", + ] + let standard_types = [ + "string", "uint", "uint", "string", "uint", "char[1]", "uint", "uint", "uint", + "int", "int[blockCount]", "int[blockCount]", + ] + let standard_comments = [ + "Reference sequence chromosome or scaffold", "Start position in chromosome", + "End position in chromosome", "Name or ID of item, ideally both human readable and unique", + "Score (0-1000)", "+ or - indicates whether the query aligns to the + or - strand on the reference", + "Start of where display should be thick (start codon)", "End of where display should be thick (stop codon)", + "RGB value (use R,G,B string in input file)", "Number of blocks", "Comma separated list of block sizes", + "Start positions relative to chromStart", + ] + let custom_names = [ + "oChromStart", "oChromEnd", "oStrand", "oChromSize", "oChromStarts", "oSequence", + "oCDS", "chromSize", "match", "misMatch", "repMatch", "nCount", "seqType", + ] + let custom_types = [ + "uint", "uint", "char[1]", "uint", "int[blockCount]", "lstring", "string", "uint", + "uint", "uint", "uint", "uint", "uint", + ] + let custom_comments = [ + "Start position in other chromosome", "End position in other chromosome", "+ or -, - means that psl was reversed into BED-compatible coordinates", + "Size of other chromosome.", "Start positions relative to oChromStart or from oChromStart+oChromSize depending on strand", + "Sequence on other chrom (or edit list, or empty)", "CDS in NCBI format", "Size of target chromosome", + "Number of bases matched.", "Number of bases that don't match", "Number of bases that match but are part of repeats", + "Number of 'N' bases", "0=empty, 1=nucleotide, 2=amino_acid", + ] + let fields : Array[BigBedField] = [] + for index in 0.. Array[BigBedField] { + let schema = bigpsl_schema() + let fields : Array[BigBedField] = [] + for index in 12.. Int { + for index in 0.. BigPslAlignment raise BigPslError { + BigPslAlignment::create( + alignment.target_name, + alignment.target_size, + alignment.query_name, + alignment.query_size, + alignment.target_coordinates, + alignment.query_coordinates, + sequence_type=alignment.sequence_type, + query_sequence=alignment.query_sequence, + cds=alignment.cds, + score=alignment.score, + thick_start=alignment.thick_start, + thick_end=alignment.thick_end, + item_rgb=alignment.item_rgb, + matches=alignment.matches, + mismatches=alignment.mismatches, + repeat_matches=alignment.repeat_matches, + n_count=alignment.n_count, + ) +} + +///| +pub fn bigpsl_write( + targets : Array[BigBedTarget], + alignments : Array[BigPslAlignment], + config? : BigPslWriteConfig = BigPslWriteConfig::default(), +) -> Array[Int] raise BigPslError { + if targets.length() == 0 { + bigpsl_fail("BigPsl requires at least one target") + } + for index in 0.. bigpsl_fail("BigBed: " + message) + } + for previous in 0.. { + bigpsl_fail("BigBed: " + message) + BigBedRecord::{ + chrom: "", + chrom_start: 0, + chrom_end: 1, + name: ".", + score: 0, + strand: ".", + thick_start: 0, + thick_end: 1, + item_rgb: "0", + declared_block_count: 0, + block_sizes: [], + block_starts: [], + extra_fields: [], + } + } + } + records.push(record) + target_order.push(target_index) + } + let order : Array[Int] = [] + for index in 0.. Int { + if target_order[left] < target_order[right] { + -1 + } else if target_order[left] > target_order[right] { + 1 + } else if records[left].chrom_start < records[right].chrom_start { + -1 + } else if records[left].chrom_start > records[right].chrom_start { + 1 + } else if records[left].chrom_end < records[right].chrom_end { + -1 + } else if records[left].chrom_end > records[right].chrom_end { + 1 + } else { + 0 + } + }) + let sorted : Array[BigBedRecord] = [] + for index in order { + sorted.push(records[index]) + } + let bed_config = BigBedWriteConfig::create( + bed_columns=12, + items_per_slot=config.items_per_slot, + block_size=config.block_size, + compress=config.compress, + schema_name="bigPsl", + schema_comment="bigPsl pairwise alignment", + custom_fields=bigpsl_custom_fields(), + ) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + BigBedWriteConfig::default() + } + } + bigbed_write(targets, sorted, config=bed_config) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + [] + } + } +} + +///| +fn bigpsl_parse_int(value : String, label : String) -> Int raise BigPslError { + if value.length() == 0 { + bigpsl_fail("Missing integer for " + label) + } + for index in 0.. '9'.to_int() { + bigpsl_fail("Invalid integer for " + label + ": " + value) + } + } + parse_int(value) +} + +///| +fn bigpsl_split_csv( + value : String, + label : String, +) -> Array[Int] raise BigPslError { + if value.length() == 0 { + return [] + } + let values : Array[Int] = [] + let mut start = 0 + for index in 0..<=value.length() { + if index == value.length() || + value.unsafe_get(index).to_int() == ','.to_int() { + if index == start { + if index == value.length() && values.length() > 0 { + return values + } + bigpsl_fail("Empty value in " + label) + } + values.push(bigpsl_parse_int(value[start:index].to_owned(), label)) + start = index + 1 + } + } + values +} + +///| +fn bigpsl_validate_schema(bed : BigBedFile) -> Unit raise BigPslError { + let expected = bigpsl_schema() + if bed.header.defined_field_count != 12 || + bed.header.field_count != 25 || + bed.schema.name != "bigPsl" || + bed.schema.fields.length() != 25 { + bigpsl_fail("BigPsl requires the standard bigPsl bed12+13 declaration") + } + for index in 0.. BigPslAlignment raise BigPslError { + if record.extra_fields.length() != 13 { + bigpsl_fail("BigPsl binary records require thirteen custom fields") + } + if record.strand != "+" && record.strand != "-" { + bigpsl_fail("BigPsl BED strand must be '+' or '-'") + } + let other_start = bigpsl_parse_int( + record.extra_fields[0], + "BigPsl oChromStart", + ) + let other_end = bigpsl_parse_int(record.extra_fields[1], "BigPsl oChromEnd") + let other_strand = record.extra_fields[2] + if other_strand != "+" && other_strand != "-" { + bigpsl_fail("BigPsl oStrand must be '+' or '-'") + } + if other_strand == "-" && record.strand != "-" { + bigpsl_fail("BigPsl oStrand '-' requires BED strand '-'") + } + let query_size = bigpsl_parse_int(record.extra_fields[3], "BigPsl oChromSize") + let stored_query_starts = bigpsl_split_csv( + record.extra_fields[4], + "BigPsl oChromStarts", + ) + let query_sequence = record.extra_fields[5] + let cds = if record.extra_fields[6] == "n/a" { + "" + } else { + record.extra_fields[6] + } + let target_size = bigpsl_parse_int(record.extra_fields[7], "BigPsl chromSize") + let matches = bigpsl_parse_int(record.extra_fields[8], "BigPsl match") + let mismatches = bigpsl_parse_int(record.extra_fields[9], "BigPsl misMatch") + let repeat_matches = bigpsl_parse_int( + record.extra_fields[10], + "BigPsl repMatch", + ) + let n_count = bigpsl_parse_int(record.extra_fields[11], "BigPsl nCount") + let sequence_type = bigpsl_type_from_code( + bigpsl_parse_int(record.extra_fields[12], "BigPsl seqType"), + ) + match bed.target(record.chrom) { + Some(target) => + if target.length != target_size { + bigpsl_fail("BigPsl chromSize disagrees with its target declaration") + } + None => bigpsl_fail("BigPsl record refers to an unknown target") + } + let block_count = record.block_sizes.length() + if block_count == 0 || + record.block_starts.length() != block_count || + stored_query_starts.length() != block_count { + bigpsl_fail("BigPsl block arrays disagree with blockCount") + } + let target_starts : Array[Int] = [] + let target_block_sizes : Array[Int] = [] + let query_block_sizes : Array[Int] = [] + let query_starts : Array[Int] = [] + for index in 0.. { + if target_block_size % 3 != 0 { + bigpsl_fail( + "Protein BigPsl target block size must be divisible by three", + ) + } + target_block_size / 3 + } + _ => target_block_size + } + target_starts.push(record.chrom_start + record.block_starts[index]) + target_block_sizes.push(target_block_size) + query_block_sizes.push(query_block_size) + query_starts.push(stored_query_starts[index]) + if stored_query_starts[index] < 0 || + stored_query_starts[index] + query_block_size > query_size { + bigpsl_fail("BigPsl query block extends beyond oChromSize") + } + } + let ( + normalized_target_starts, + normalized_query_starts, + normalized_target_sizes, + normalized_query_sizes, + ) = if other_strand == "-" && record.strand == "-" { + let transformed_targets : Array[Int] = [] + let transformed_queries : Array[Int] = [] + for index in 0.. BigPslFile raise BigPslError { + let bed = bigbed_parse(data) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + BigBedFile::{ + header: BigBedHeader::{ + little_endian: true, + version: 4, + zoom_levels: 0, + chromosome_tree_offset: 0, + full_data_offset: 0, + full_index_offset: 0, + field_count: 0, + defined_field_count: 0, + auto_sql_offset: 0, + total_summary_offset: 0, + uncompress_buffer_size: 0, + extra_indices_offset: 0, + }, + schema: bigpsl_schema(), + targets: [], + source: [], + blocks: [], + record_count: 0, + } + } + } + bigpsl_validate_schema(bed) + let records = bed.records() catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + [] + } + } + for record in records { + ignore(bigpsl_decode_record(bed, record)) + } + BigPslFile::{ bed, } +} + +///| +pub fn BigPslFile::alignments( + self : BigPslFile, +) -> Array[BigPslAlignment] raise BigPslError { + let records = self.bed.records() catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + [] + } + } + let alignments : Array[BigPslAlignment] = [] + for record in records { + alignments.push(bigpsl_decode_record(self.bed, record)) + } + alignments +} + +///| +pub fn BigPslFile::search( + self : BigPslFile, + target : String, + start? : Int = 0, + end? : Int = -1, +) -> Array[BigPslAlignment] raise BigPslError { + let records = self.bed.search(target, start~, end~) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + [] + } + } + let alignments : Array[BigPslAlignment] = [] + for record in records { + alignments.push(bigpsl_decode_record(self.bed, record)) + } + alignments +} + +///| +pub fn BigPslFile::find_by_query( + self : BigPslFile, + query_name : String, +) -> Array[BigPslAlignment] raise BigPslError { + let records = self.bed.find_by_name(query_name) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + [] + } + } + let alignments : Array[BigPslAlignment] = [] + for record in records { + alignments.push(bigpsl_decode_record(self.bed, record)) + } + alignments +} + +///| +pub fn BigPslFile::target(self : BigPslFile, name : String) -> BigBedTarget? { + self.bed.target(name) +} + +///| +pub fn BigPslFile::is_compressed(self : BigPslFile) -> Bool { + self.bed.is_compressed() +} + +///| +pub fn BigPslAlignment::to_psl(self : BigPslAlignment) -> String { + let blocks = self.blocks() + let counts = self.counts() + let target_reverse = self.target_coordinates[0] > + self.target_coordinates[self.target_coordinates.length() - 1] + let query_reverse = self.query_coordinates[0] > + self.query_coordinates[self.query_coordinates.length() - 1] + let strand = if target_reverse { + "+-" + } else if query_reverse { + "-" + } else { + "+" + } + let block_sizes : Array[Int] = [] + let query_starts : Array[Int] = [] + let target_starts : Array[Int] = [] + for block in blocks { + block_sizes.push(block.query_size) + let query_minimum = bigpsl_min(block.query_start, block.query_end) + let query_maximum = bigpsl_max(block.query_start, block.query_end) + let target_minimum = bigpsl_min(block.target_start, block.target_end) + let target_maximum = bigpsl_max(block.target_start, block.target_end) + query_starts.push( + if query_reverse { + self.query_size - query_maximum + } else { + query_minimum + }, + ) + target_starts.push( + if target_reverse { + self.target_size - target_maximum + } else { + target_minimum + }, + ) + } + let (query_start, query_end) = self.query_interval() + let (target_start, target_end) = self.target_interval() + self.matches.to_string() + + "\t" + + self.mismatches.to_string() + + "\t" + + self.repeat_matches.to_string() + + "\t" + + self.n_count.to_string() + + "\t" + + counts.query_insert_events.to_string() + + "\t" + + counts.query_insert_units.to_string() + + "\t" + + counts.target_insert_events.to_string() + + "\t" + + counts.target_insert_bases.to_string() + + "\t" + + strand + + "\t" + + self.query_name + + "\t" + + self.query_size.to_string() + + "\t" + + query_start.to_string() + + "\t" + + query_end.to_string() + + "\t" + + self.target_name + + "\t" + + self.target_size.to_string() + + "\t" + + target_start.to_string() + + "\t" + + target_end.to_string() + + "\t" + + blocks.length().to_string() + + "\t" + + bigpsl_join_ints(block_sizes, true) + + "\t" + + bigpsl_join_ints(query_starts, true) + + "\t" + + bigpsl_join_ints(target_starts, true) + + "\n" +} + +///| +pub fn BigPslFile::to_psl(self : BigPslFile) -> String raise BigPslError { + let output = StringBuilder::new() + for alignment in self.alignments() { + output.write_string(alignment.to_psl()) + } + output.to_string() +} + +///| +pub fn BigPslAlignment::summary(self : BigPslAlignment) -> String { + let (start, end) = self.target_interval() + "BigPslAlignment(" + + self.target_name + + ":" + + start.to_string() + + "-" + + end.to_string() + + " <- " + + self.query_name + + ", blocks=" + + self.blocks().length().to_string() + + ", type=" + + (match self.sequence_type { + BigPslUnknown => "unknown" + BigPslNucleotide => "nucleotide" + BigPslProtein => "protein" + }) + + ")" +} + +///| +pub fn BigPslFile::summary( + self : BigPslFile, +) -> BigPslSummary raise BigPslError { + let alignments = self.alignments() + let mut nucleotide_count = 0 + let mut protein_count = 0 + let mut unknown_count = 0 + let mut aligned_query_units = 0 + let mut covered_target_bases = 0 + for alignment in alignments { + match alignment.sequence_type { + BigPslNucleotide => nucleotide_count = nucleotide_count + 1 + BigPslProtein => protein_count = protein_count + 1 + BigPslUnknown => unknown_count = unknown_count + 1 + } + aligned_query_units = aligned_query_units + + alignment.counts().aligned_query_units + let (start, end) = alignment.target_interval() + covered_target_bases = covered_target_bases + end - start + } + BigPslSummary::{ + alignment_count: alignments.length(), + target_count: self.bed.targets.length(), + nucleotide_count, + protein_count, + unknown_count, + aligned_query_units, + covered_target_bases, + compressed: self.is_compressed(), + } +} + +///| +pub fn bigpsl_example_data() -> (Array[BigBedTarget], Array[BigPslAlignment]) raise BigPslError { + let targets = [ + BigBedTarget::create("chr1", 1000) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + BigBedTarget::{ name: "chr1", length: 1000 } + } + }, + BigBedTarget::create("chr2", 800) catch { + BigBedError(message) => { + bigpsl_fail("BigBed: " + message) + BigBedTarget::{ name: "chr2", length: 800 } + } + }, + ] + let forward = BigPslAlignment::create( + "chr1", + 1000, + "rna-forward", + 30, + [100, 110, 120, 130], + [0, 10, 10, 20], + sequence_type=BigPslNucleotide, + score=900, + thick_start=103, + thick_end=127, + ) + let reverse = BigPslAlignment::create( + "chr1", + 1000, + "rna-reverse", + 40, + [200, 208, 208, 215], + [30, 22, 20, 13], + sequence_type=BigPslNucleotide, + score=850, + ) + let protein = BigPslAlignment::create( + "chr2", + 800, + "protein-reverse", + 20, + [500, 488, 470, 470, 455], + [0, 4, 4, 6, 11], + sequence_type=BigPslProtein, + score=1000, + ) + (targets, [forward, reverse, protein]) +} diff --git a/src/binary_cif.mbt b/src/binary_cif.mbt new file mode 100644 index 00000000..6b897458 --- /dev/null +++ b/src/binary_cif.mbt @@ -0,0 +1,2052 @@ +// Bio.PDB.binary_cif - BinaryCIF decoding and PDB structure construction. +// +// BinaryCIF stores CIF data blocks in MessagePack and applies a reversible +// encoding pipeline to each column. The encoding list is decoded in reverse. + +///| +pub suberror BinaryCifError { + BinaryCifError(String) +} + +///| +pub enum BinaryCifMask { + BinaryCifPresent + BinaryCifNotPresent + BinaryCifUnknown +} derive(Eq, Debug) + +///| +pub enum BinaryCifColumnKind { + BinaryCifInteger + BinaryCifFloat + BinaryCifText +} derive(Eq, Debug) + +///| +pub enum BinaryCifColumnData { + BinaryCifIntegers(Array[Int]) + BinaryCifFloats(Array[Double]) + BinaryCifStrings(Array[String]) +} + +///| +pub struct BinaryCifColumn { + name : String + data : BinaryCifColumnData + mask : Array[Int] +} + +///| +pub struct BinaryCifCategory { + name : String + row_count : Int + columns : Array[BinaryCifColumn] +} + +///| +pub struct BinaryCifDataBlock { + header : String + categories : Array[BinaryCifCategory] +} + +///| +pub struct BinaryCifFile { + version : String + encoder : String + data_blocks : Array[BinaryCifDataBlock] +} + +///| +pub fn BinaryCifColumn::kind(self : BinaryCifColumn) -> BinaryCifColumnKind { + match self.data { + BinaryCifIntegers(_) => BinaryCifInteger + BinaryCifFloats(_) => BinaryCifFloat + BinaryCifStrings(_) => BinaryCifText + } +} + +///| +pub fn BinaryCifColumn::length(self : BinaryCifColumn) -> Int { + match self.data { + BinaryCifIntegers(values) => values.length() + BinaryCifFloats(values) => values.length() + BinaryCifStrings(values) => values.length() + } +} + +///| +pub fn BinaryCifColumn::mask_at( + self : BinaryCifColumn, + index : Int, +) -> BinaryCifMask raise BinaryCifError { + if index < 0 || index >= self.length() { + raise BinaryCifError("BinaryCIF column index is out of bounds") + } + let value = if self.mask.length() == 0 { 0 } else { self.mask[index] } + match value { + 0 => BinaryCifPresent + 1 => BinaryCifNotPresent + 2 => BinaryCifUnknown + _ => raise BinaryCifError("BinaryCIF mask value must be 0, 1, or 2") + } +} + +///| +pub fn BinaryCifColumn::int_at( + self : BinaryCifColumn, + index : Int, +) -> Int? raise BinaryCifError { + match self.mask_at(index) { + BinaryCifPresent => + match self.data { + BinaryCifIntegers(values) => Some(values[index]) + _ => raise BinaryCifError("BinaryCIF column is not an integer column") + } + _ => None + } +} + +///| +pub fn BinaryCifColumn::double_at( + self : BinaryCifColumn, + index : Int, +) -> Double? raise BinaryCifError { + match self.mask_at(index) { + BinaryCifPresent => + match self.data { + BinaryCifIntegers(values) => Some(values[index].to_double()) + BinaryCifFloats(values) => Some(values[index]) + _ => raise BinaryCifError("BinaryCIF column is not a numeric column") + } + _ => None + } +} + +///| +pub fn BinaryCifColumn::string_at( + self : BinaryCifColumn, + index : Int, +) -> String? raise BinaryCifError { + match self.mask_at(index) { + BinaryCifPresent => + match self.data { + BinaryCifStrings(values) => Some(values[index]) + _ => raise BinaryCifError("BinaryCIF column is not a string column") + } + _ => None + } +} + +///| +pub fn BinaryCifColumn::cif_token_at( + self : BinaryCifColumn, + index : Int, +) -> String raise BinaryCifError { + match self.mask_at(index) { + BinaryCifNotPresent => "." + BinaryCifUnknown => "?" + BinaryCifPresent => + match self.data { + BinaryCifIntegers(values) => values[index].to_string() + BinaryCifFloats(values) => values[index].to_string() + BinaryCifStrings(values) => values[index] + } + } +} + +///| +pub fn BinaryCifCategory::get_column( + self : BinaryCifCategory, + name : String, +) -> BinaryCifColumn? { + for column in self.columns { + if column.name == name || self.name + "." + column.name == name { + return Some(column) + } + } + None +} + +///| +pub fn BinaryCifDataBlock::get_category( + self : BinaryCifDataBlock, + name : String, +) -> BinaryCifCategory? { + for category in self.categories { + if category.name == name || + ( + name.length() > 0 && + name.unsafe_get(0).to_int() != '_'.to_int() && + category.name == "_" + name + ) { + return Some(category) + } + } + None +} + +///| +pub fn BinaryCifFile::get_block( + self : BinaryCifFile, + index : Int, +) -> BinaryCifDataBlock? { + if index < 0 || index >= self.data_blocks.length() { + None + } else { + Some(self.data_blocks[index]) + } +} + +///| +pub fn BinaryCifFile::summary(self : BinaryCifFile) -> String { + let mut category_count = 0 + let mut row_count = 0 + for block in self.data_blocks { + category_count = category_count + block.categories.length() + for category in block.categories { + row_count = row_count + category.row_count + } + } + "BinaryCIF(version=" + + self.version + + ", blocks=" + + self.data_blocks.length().to_string() + + ", categories=" + + category_count.to_string() + + ", rows=" + + row_count.to_string() + + ")" +} + +// MessagePack model and decoder. + +///| +priv enum BcifMessage { + BcifNil + BcifBoolean(Bool) + BcifIntegerValue(Int) + BcifFloatValue(Double) + BcifTextValue(String) + BcifBinaryValue(Array[Int]) + BcifArrayValue(Array[BcifMessage]) + BcifMapValue(Map[String, BcifMessage]) +} + +///| +priv struct BcifReader { + data : Array[Int] + mut position : Int +} + +///| +fn bcif_read_byte(reader : BcifReader) -> Int raise BinaryCifError { + if reader.position >= reader.data.length() { + raise BinaryCifError("Truncated MessagePack input") + } + let value = reader.data[reader.position] + reader.position = reader.position + 1 + value +} + +///| +fn bcif_read_unsigned( + reader : BcifReader, + byte_count : Int, +) -> Int raise BinaryCifError { + if byte_count <= 0 || byte_count > 4 { + raise BinaryCifError("Unsupported MessagePack unsigned integer width") + } + if reader.position + byte_count > reader.data.length() { + raise BinaryCifError("Truncated MessagePack integer") + } + let first = reader.data[reader.position] + if byte_count == 4 && first >= 128 { + raise BinaryCifError("MessagePack unsigned integer exceeds MoonBit Int") + } + let mut value = 0 + for _ in 0.. Int raise BinaryCifError { + if byte_count <= 0 || byte_count > 4 { + raise BinaryCifError("Unsupported MessagePack signed integer width") + } + if reader.position + byte_count > reader.data.length() { + raise BinaryCifError("Truncated MessagePack integer") + } + let first = bcif_read_byte(reader) + let mut value = if first >= 128 { first - 256 } else { first } + for _ in 1.. Int raise BinaryCifError { + for _ in 0..<4 { + if bcif_read_byte(reader) != 0 { + raise BinaryCifError("MessagePack uint64 exceeds MoonBit Int") + } + } + bcif_read_unsigned(reader, 4) +} + +///| +fn bcif_read_i64(reader : BcifReader) -> Int raise BinaryCifError { + if reader.position + 8 > reader.data.length() { + raise BinaryCifError("Truncated MessagePack int64") + } + let sign_byte = reader.data[reader.position] + let expected = if sign_byte >= 128 { 255 } else { 0 } + for _ in 0..<4 { + if bcif_read_byte(reader) != expected { + raise BinaryCifError("MessagePack int64 exceeds MoonBit Int") + } + } + let next = reader.data[reader.position] + if (expected == 0 && next >= 128) || (expected == 255 && next < 128) { + raise BinaryCifError("MessagePack int64 exceeds MoonBit Int") + } + bcif_read_signed(reader, 4) +} + +///| +fn bcif_pow2(exponent : Int) -> Double { + let mut power = if exponent < 0 { -exponent } else { exponent } + let mut base = 2.0 + let mut result = 1.0 + while power > 0 { + if power % 2 == 1 { + result = result * base + } + base = base * base + power = power / 2 + } + if exponent < 0 { + 1.0 / result + } else { + result + } +} + +///| +fn bcif_float32_from_bytes( + b0 : Int, + b1 : Int, + b2 : Int, + b3 : Int, +) -> Double raise BinaryCifError { + let sign = if b0 >= 128 { -1.0 } else { 1.0 } + let exponent = (b0 & 0x7F) * 2 + (b1 >> 7) + let mantissa = (b1 & 0x7F) * 65536 + b2 * 256 + b3 + if exponent == 255 { + raise BinaryCifError("BinaryCIF does not accept non-finite floats") + } + if exponent == 0 { + sign * mantissa.to_double() * bcif_pow2(-149) + } else { + sign * (1.0 + mantissa.to_double() / 8388608.0) * bcif_pow2(exponent - 127) + } +} + +///| +fn bcif_float64_from_bytes( + bytes : Array[Int], + offset : Int, + little_endian : Bool, +) -> Double raise BinaryCifError { + let ordered = Array::make(8, 0) + for i in 0..<8 { + ordered[i] = if little_endian { + bytes[offset + 7 - i] + } else { + bytes[offset + i] + } + } + let sign = if ordered[0] >= 128 { -1.0 } else { 1.0 } + let exponent = (ordered[0] & 0x7F) * 16 + (ordered[1] >> 4) + if exponent == 2047 { + raise BinaryCifError("BinaryCIF does not accept non-finite doubles") + } + let mut fraction = (ordered[1] & 0x0F).to_double() / 16.0 + let mut scale = 1.0 / 4096.0 + for i in 2..<8 { + fraction = fraction + ordered[i].to_double() * scale + scale = scale / 256.0 + } + if exponent == 0 { + sign * fraction * bcif_pow2(-1022) + } else { + sign * (1.0 + fraction) * bcif_pow2(exponent - 1023) + } +} + +///| +fn bcif_read_float32(reader : BcifReader) -> Double raise BinaryCifError { + let b0 = bcif_read_byte(reader) + let b1 = bcif_read_byte(reader) + let b2 = bcif_read_byte(reader) + let b3 = bcif_read_byte(reader) + bcif_float32_from_bytes(b0, b1, b2, b3) +} + +///| +fn bcif_read_float64(reader : BcifReader) -> Double raise BinaryCifError { + if reader.position + 8 > reader.data.length() { + raise BinaryCifError("Truncated MessagePack float64") + } + let value = bcif_float64_from_bytes(reader.data, reader.position, false) + reader.position = reader.position + 8 + value +} + +///| +fn bcif_read_utf8( + reader : BcifReader, + length : Int, +) -> String raise BinaryCifError { + if length < 0 || reader.position + length > reader.data.length() { + raise BinaryCifError("Truncated MessagePack string") + } + let end = reader.position + length + let output = StringBuilder::new() + while reader.position < end { + let first = bcif_read_byte(reader) + if first < 0x80 { + output.write_char(first.unsafe_to_char()) + } else if first >= 0xC2 && first <= 0xDF { + if reader.position >= end { + raise BinaryCifError("Truncated UTF-8 sequence") + } + let second = bcif_read_byte(reader) + if second < 0x80 || second > 0xBF { + raise BinaryCifError("Invalid UTF-8 continuation byte") + } + output.write_char( + ((first & 0x1F) * 64 + (second & 0x3F)).unsafe_to_char(), + ) + } else if first >= 0xE0 && first <= 0xEF { + if reader.position + 2 > end { + raise BinaryCifError("Truncated UTF-8 sequence") + } + let second = bcif_read_byte(reader) + let third = bcif_read_byte(reader) + if second < 0x80 || + second > 0xBF || + third < 0x80 || + third > 0xBF || + (first == 0xE0 && second < 0xA0) || + (first == 0xED && second >= 0xA0) { + raise BinaryCifError("Invalid UTF-8 sequence") + } + output.write_char( + ((first & 0x0F) * 4096 + (second & 0x3F) * 64 + (third & 0x3F)).unsafe_to_char(), + ) + } else if first >= 0xF0 && first <= 0xF4 { + if reader.position + 3 > end { + raise BinaryCifError("Truncated UTF-8 sequence") + } + let second = bcif_read_byte(reader) + let third = bcif_read_byte(reader) + let fourth = bcif_read_byte(reader) + if second < 0x80 || + second > 0xBF || + third < 0x80 || + third > 0xBF || + fourth < 0x80 || + fourth > 0xBF || + (first == 0xF0 && second < 0x90) || + (first == 0xF4 && second >= 0x90) { + raise BinaryCifError("Invalid UTF-8 sequence") + } + output.write_char( + ((first & 0x07) * 262144 + + (second & 0x3F) * 4096 + + (third & 0x3F) * 64 + + (fourth & 0x3F)).unsafe_to_char(), + ) + } else { + raise BinaryCifError("Invalid UTF-8 leading byte") + } + } + output.to_string() +} + +///| +fn bcif_read_binary( + reader : BcifReader, + length : Int, +) -> Array[Int] raise BinaryCifError { + if length < 0 || reader.position + length > reader.data.length() { + raise BinaryCifError("Truncated MessagePack binary value") + } + let result = Array::make(length, 0) + for i in 0.. BcifMessage raise BinaryCifError { + if depth > 128 { + raise BinaryCifError("MessagePack nesting is too deep") + } + let marker = bcif_read_byte(reader) + if marker <= 0x7F { + return BcifIntegerValue(marker) + } + if marker >= 0xE0 { + return BcifIntegerValue(marker - 256) + } + if marker >= 0xA0 && marker <= 0xBF { + return BcifTextValue(bcif_read_utf8(reader, marker & 0x1F)) + } + if marker >= 0x90 && marker <= 0x9F { + let length = marker & 0x0F + let values = Array::new(capacity=length) + for _ in 0..= 0x80 && marker <= 0x8F { + return bcif_parse_map(reader, marker & 0x0F, depth + 1) + } + match marker { + 0xC0 => BcifNil + 0xC2 => BcifBoolean(false) + 0xC3 => BcifBoolean(true) + 0xC4 => + BcifBinaryValue(bcif_read_binary(reader, bcif_read_unsigned(reader, 1))) + 0xC5 => + BcifBinaryValue(bcif_read_binary(reader, bcif_read_unsigned(reader, 2))) + 0xC6 => + BcifBinaryValue(bcif_read_binary(reader, bcif_read_unsigned(reader, 4))) + 0xCA => BcifFloatValue(bcif_read_float32(reader)) + 0xCB => BcifFloatValue(bcif_read_float64(reader)) + 0xCC => BcifIntegerValue(bcif_read_unsigned(reader, 1)) + 0xCD => BcifIntegerValue(bcif_read_unsigned(reader, 2)) + 0xCE => BcifIntegerValue(bcif_read_unsigned(reader, 4)) + 0xCF => BcifIntegerValue(bcif_read_u64(reader)) + 0xD0 => BcifIntegerValue(bcif_read_signed(reader, 1)) + 0xD1 => BcifIntegerValue(bcif_read_signed(reader, 2)) + 0xD2 => BcifIntegerValue(bcif_read_signed(reader, 4)) + 0xD3 => BcifIntegerValue(bcif_read_i64(reader)) + 0xD9 => BcifTextValue(bcif_read_utf8(reader, bcif_read_unsigned(reader, 1))) + 0xDA => BcifTextValue(bcif_read_utf8(reader, bcif_read_unsigned(reader, 2))) + 0xDB => BcifTextValue(bcif_read_utf8(reader, bcif_read_unsigned(reader, 4))) + 0xDC => { + let length = bcif_read_unsigned(reader, 2) + let values = Array::new(capacity=length) + for _ in 0.. { + let length = bcif_read_unsigned(reader, 4) + let values = Array::new(capacity=length) + for _ in 0.. bcif_parse_map(reader, bcif_read_unsigned(reader, 2), depth + 1) + 0xDF => bcif_parse_map(reader, bcif_read_unsigned(reader, 4), depth + 1) + _ => + raise BinaryCifError( + "Unsupported MessagePack marker " + marker.to_string(), + ) + } +} + +///| +fn bcif_parse_map( + reader : BcifReader, + length : Int, + depth : Int, +) -> BcifMessage raise BinaryCifError { + let values : Map[String, BcifMessage] = Map([], capacity=length) + for _ in 0.. value + _ => raise BinaryCifError("MessagePack map keys must be strings") + } + values[key] = bcif_parse_message(reader, depth) + } + BcifMapValue(values) +} + +///| +fn bcif_as_map( + value : BcifMessage, + context : String, +) -> Map[String, BcifMessage] raise BinaryCifError { + match value { + BcifMapValue(result) => result + _ => raise BinaryCifError(context + " must be a MessagePack map") + } +} + +///| +fn bcif_as_array( + value : BcifMessage, + context : String, +) -> Array[BcifMessage] raise BinaryCifError { + match value { + BcifArrayValue(result) => result + _ => raise BinaryCifError(context + " must be a MessagePack array") + } +} + +///| +fn bcif_as_string( + value : BcifMessage, + context : String, +) -> String raise BinaryCifError { + match value { + BcifTextValue(result) => result + _ => raise BinaryCifError(context + " must be a string") + } +} + +///| +fn bcif_as_int( + value : BcifMessage, + context : String, +) -> Int raise BinaryCifError { + match value { + BcifIntegerValue(result) => result + _ => raise BinaryCifError(context + " must be an integer") + } +} + +///| +fn bcif_as_double( + value : BcifMessage, + context : String, +) -> Double raise BinaryCifError { + match value { + BcifIntegerValue(result) => result.to_double() + BcifFloatValue(result) => result + _ => raise BinaryCifError(context + " must be numeric") + } +} + +///| +fn bcif_as_bool( + value : BcifMessage, + context : String, +) -> Bool raise BinaryCifError { + match value { + BcifBoolean(result) => result + _ => raise BinaryCifError(context + " must be a boolean") + } +} + +///| +fn bcif_as_binary( + value : BcifMessage, + context : String, +) -> Array[Int] raise BinaryCifError { + match value { + BcifBinaryValue(result) => result + _ => raise BinaryCifError(context + " must be binary data") + } +} + +///| +fn bcif_required( + values : Map[String, BcifMessage], + key : String, + context : String, +) -> BcifMessage raise BinaryCifError { + match values.get(key) { + Some(value) => value + None => raise BinaryCifError(context + " is missing '" + key + "'") + } +} + +///| +fn bcif_encoding_int( + values : Map[String, BcifMessage], + key : String, + context : String, +) -> Int raise BinaryCifError { + bcif_as_int(bcif_required(values, key, context), context + "." + key) +} + +///| +fn bcif_encoding_double( + values : Map[String, BcifMessage], + key : String, + context : String, +) -> Double raise BinaryCifError { + bcif_as_double(bcif_required(values, key, context), context + "." + key) +} + +// BinaryCIF encoding decoders. + +///| +fn bcif_read_i16_le(data : Array[Int], offset : Int) -> Int { + let high = data[offset + 1] + (if high >= 128 { high - 256 } else { high }) * 256 + data[offset] +} + +///| +fn bcif_read_u16_le(data : Array[Int], offset : Int) -> Int { + data[offset] + data[offset + 1] * 256 +} + +///| +fn bcif_read_i32_le(data : Array[Int], offset : Int) -> Int { + let high = data[offset + 3] + (if high >= 128 { high - 256 } else { high }) * 16777216 + + data[offset + 2] * 65536 + + data[offset + 1] * 256 + + data[offset] +} + +///| +fn bcif_read_u32_le( + data : Array[Int], + offset : Int, +) -> Int raise BinaryCifError { + if data[offset + 3] >= 128 { + raise BinaryCifError("BinaryCIF UInt32 exceeds MoonBit Int") + } + data[offset + 3] * 16777216 + + data[offset + 2] * 65536 + + data[offset + 1] * 256 + + data[offset] +} + +///| +fn bcif_validate_bytes(data : Array[Int]) -> Unit raise BinaryCifError { + for value in data { + if value < 0 || value > 255 { + raise BinaryCifError("BinaryCIF byte array contains a non-byte value") + } + } +} + +///| +pub fn binary_cif_decode_int_bytes( + data : Array[Int], + type_code : Int, +) -> Array[Int] raise BinaryCifError { + bcif_validate_bytes(data) + let width = match type_code { + 1 | 4 => 1 + 2 | 5 => 2 + 3 | 6 => 4 + _ => raise BinaryCifError("BinaryCIF type is not an integer byte array") + } + if data.length() % width != 0 { + raise BinaryCifError("BinaryCIF byte array has an invalid length") + } + let result = Array::new(capacity=data.length() / width) + let mut offset = 0 + while offset < data.length() { + let value = match type_code { + 1 => { + let raw = data[offset] + if raw >= 128 { + raw - 256 + } else { + raw + } + } + 2 => bcif_read_i16_le(data, offset) + 3 => bcif_read_i32_le(data, offset) + 4 => data[offset] + 5 => bcif_read_u16_le(data, offset) + 6 => bcif_read_u32_le(data, offset) + _ => 0 + } + result.push(value) + offset = offset + width + } + result +} + +///| +pub fn binary_cif_decode_float_bytes( + data : Array[Int], + type_code : Int, +) -> Array[Double] raise BinaryCifError { + bcif_validate_bytes(data) + let width = match type_code { + 32 => 4 + 33 => 8 + _ => + raise BinaryCifError("BinaryCIF type is not a floating-point byte array") + } + if data.length() % width != 0 { + raise BinaryCifError( + "BinaryCIF floating-point byte array has an invalid length", + ) + } + let result = Array::new(capacity=data.length() / width) + let mut offset = 0 + while offset < data.length() { + if type_code == 32 { + result.push( + bcif_float32_from_bytes( + data[offset + 3], + data[offset + 2], + data[offset + 1], + data[offset], + ), + ) + } else { + result.push(bcif_float64_from_bytes(data, offset, true)) + } + offset = offset + width + } + result +} + +///| +pub fn binary_cif_decode_integer_packing( + data : Array[Int], + byte_count : Int, + is_unsigned : Bool, + source_size : Int, +) -> Array[Int] raise BinaryCifError { + if byte_count != 1 && byte_count != 2 { + raise BinaryCifError("BinaryCIF integer packing uses 1- or 2-byte values") + } + if source_size < 0 { + raise BinaryCifError("BinaryCIF integer packing source size is negative") + } + let upper = if is_unsigned { + if byte_count == 1 { + 255 + } else { + 65535 + } + } else if byte_count == 1 { + 127 + } else { + 32767 + } + let lower = if is_unsigned { + 0 + } else if byte_count == 1 { + -128 + } else { + -32768 + } + let result = Array::new(capacity=source_size) + let mut accumulator = 0 + for value in data { + if value < lower || value > upper { + raise BinaryCifError( + "BinaryCIF integer packing value is outside its byte range", + ) + } + accumulator = accumulator + value + let continuation = value == upper || (!is_unsigned && value == lower) + if !continuation { + result.push(accumulator) + accumulator = 0 + } + } + if accumulator != 0 || result.length() != source_size { + raise BinaryCifError("BinaryCIF integer packing source size does not match") + } + result +} + +///| +pub fn binary_cif_decode_run_length( + data : Array[Int], + source_size : Int, +) -> Array[Int] raise BinaryCifError { + if source_size < 0 || data.length() % 2 != 0 { + raise BinaryCifError("Invalid BinaryCIF run-length encoding") + } + let result = Array::new(capacity=source_size) + let mut index = 0 + while index < data.length() { + let value = data[index] + let count = data[index + 1] + if count < 0 || result.length() + count > source_size { + raise BinaryCifError("Invalid BinaryCIF run length") + } + for _ in 0.. Array[Int] { + let result = Array::new(capacity=data.length()) + let mut current = origin + for value in data { + current = current + value + result.push(current) + } + result +} + +///| +pub fn binary_cif_decode_fixed_point( + data : Array[Int], + factor : Double, +) -> Array[Double] raise BinaryCifError { + if factor == 0.0 || factor.abs() > 1.0e300 { + raise BinaryCifError( + "BinaryCIF fixed-point factor must be finite and non-zero", + ) + } + let result = Array::new(capacity=data.length()) + for value in data { + result.push(value.to_double() / factor) + } + result +} + +///| +pub fn binary_cif_decode_interval_quantization( + data : Array[Int], + minimum : Double, + maximum : Double, + number_of_steps : Int, +) -> Array[Double] raise BinaryCifError { + if number_of_steps < 2 || + minimum.abs() > 1.0e300 || + maximum.abs() > 1.0e300 || + maximum < minimum { + raise BinaryCifError("Invalid BinaryCIF interval quantization parameters") + } + let step = (maximum - minimum) / (number_of_steps - 1).to_double() + let result = Array::new(capacity=data.length()) + for value in data { + if value < 0 || value >= number_of_steps { + raise BinaryCifError("BinaryCIF quantized value is outside its interval") + } + result.push(minimum + value.to_double() * step) + } + result +} + +///| +pub fn binary_cif_decode_string_array( + indices : Array[Int], + offsets : Array[Int], + string_data : String, +) -> Array[String] raise BinaryCifError { + if offsets.length() == 0 || offsets[0] != 0 { + raise BinaryCifError("BinaryCIF string offsets must start at zero") + } + for i in 1.. string_data.length() { + raise BinaryCifError("BinaryCIF string offsets are invalid") + } + } + if offsets[offsets.length() - 1] != string_data.length() { + raise BinaryCifError( + "BinaryCIF string offsets do not cover the string data", + ) + } + let dictionary = Array::new(capacity=offsets.length() - 1) + for i in 0..<(offsets.length() - 1) { + dictionary.push(string_data[offsets[i]:offsets[i + 1]].to_owned()) + } + let result = Array::new(capacity=indices.length()) + for index in indices { + if index < 0 || index >= dictionary.length() { + raise BinaryCifError("BinaryCIF string lookup index is out of bounds") + } + result.push(dictionary[index]) + } + result +} + +///| +fn bcif_decoded_length(data : BinaryCifColumnData) -> Int { + match data { + BinaryCifIntegers(values) => values.length() + BinaryCifFloats(values) => values.length() + BinaryCifStrings(values) => values.length() + } +} + +///| +fn bcif_decode_raw( + bytes : Array[Int], + encodings : Array[BcifMessage], +) -> BinaryCifColumnData raise BinaryCifError { + let mut raw_bytes : Array[Int]? = Some(bytes) + let mut decoded : BinaryCifColumnData? = None + let mut index = encodings.length() - 1 + while index >= 0 { + let encoding = bcif_as_map(encodings[index], "BinaryCIF encoding") + let kind = bcif_as_string( + bcif_required(encoding, "kind", "BinaryCIF encoding"), + "BinaryCIF encoding.kind", + ) + match kind { + "ByteArray" => { + let source = match raw_bytes { + Some(value) => value + None => + raise BinaryCifError( + "ByteArray must be the final BinaryCIF encoding", + ) + } + let type_code = bcif_encoding_int(encoding, "type", "ByteArray") + decoded = Some( + if type_code == 32 || type_code == 33 { + BinaryCifFloats(binary_cif_decode_float_bytes(source, type_code)) + } else { + BinaryCifIntegers(binary_cif_decode_int_bytes(source, type_code)) + }, + ) + raw_bytes = None + } + "IntegerPacking" => { + let values = match decoded { + Some(BinaryCifIntegers(value)) => value + _ => raise BinaryCifError("IntegerPacking requires integer input") + } + let byte_count = bcif_encoding_int( + encoding, "byteCount", "IntegerPacking", + ) + let source_size = bcif_encoding_int( + encoding, "srcSize", "IntegerPacking", + ) + let is_unsigned = bcif_as_bool( + bcif_required(encoding, "isUnsigned", "IntegerPacking"), + "IntegerPacking.isUnsigned", + ) + decoded = Some( + BinaryCifIntegers( + binary_cif_decode_integer_packing( + values, byte_count, is_unsigned, source_size, + ), + ), + ) + } + "RunLength" => { + let values = match decoded { + Some(BinaryCifIntegers(value)) => value + _ => raise BinaryCifError("RunLength requires integer input") + } + decoded = Some( + BinaryCifIntegers( + binary_cif_decode_run_length( + values, + bcif_encoding_int(encoding, "srcSize", "RunLength"), + ), + ), + ) + } + "Delta" => { + let values = match decoded { + Some(BinaryCifIntegers(value)) => value + _ => raise BinaryCifError("Delta requires integer input") + } + decoded = Some( + BinaryCifIntegers( + binary_cif_decode_delta( + values, + bcif_encoding_int(encoding, "origin", "Delta"), + ), + ), + ) + } + "FixedPoint" => { + let values = match decoded { + Some(BinaryCifIntegers(value)) => value + _ => raise BinaryCifError("FixedPoint requires integer input") + } + decoded = Some( + BinaryCifFloats( + binary_cif_decode_fixed_point( + values, + bcif_encoding_double(encoding, "factor", "FixedPoint"), + ), + ), + ) + } + "IntervalQuantization" => { + let values = match decoded { + Some(BinaryCifIntegers(value)) => value + _ => + raise BinaryCifError("IntervalQuantization requires integer input") + } + let steps = match encoding.get("numSteps") { + Some(value) => bcif_as_int(value, "IntervalQuantization.numSteps") + None => + bcif_as_int( + bcif_required(encoding, "num_steps", "IntervalQuantization"), + "IntervalQuantization.num_steps", + ) + } + decoded = Some( + BinaryCifFloats( + binary_cif_decode_interval_quantization( + values, + bcif_encoding_double(encoding, "min", "IntervalQuantization"), + bcif_encoding_double(encoding, "max", "IntervalQuantization"), + steps, + ), + ), + ) + } + "StringArray" => { + let source = match raw_bytes { + Some(value) => value + None => raise BinaryCifError("StringArray requires raw byte input") + } + let data_encodings = bcif_as_array( + bcif_required(encoding, "dataEncoding", "StringArray"), + "StringArray.dataEncoding", + ) + let offset_encodings = bcif_as_array( + bcif_required(encoding, "offsetEncoding", "StringArray"), + "StringArray.offsetEncoding", + ) + let offsets_bytes = bcif_as_binary( + bcif_required(encoding, "offsets", "StringArray"), + "StringArray.offsets", + ) + let indices = match bcif_decode_raw(source, data_encodings) { + BinaryCifIntegers(value) => value + _ => + raise BinaryCifError( + "StringArray lookup data must decode to integers", + ) + } + let offsets = match bcif_decode_raw(offsets_bytes, offset_encodings) { + BinaryCifIntegers(value) => value + _ => + raise BinaryCifError("StringArray offsets must decode to integers") + } + decoded = Some( + BinaryCifStrings( + binary_cif_decode_string_array( + indices, + offsets, + bcif_as_string( + bcif_required(encoding, "stringData", "StringArray"), + "StringArray.stringData", + ), + ), + ), + ) + raw_bytes = None + } + _ => raise BinaryCifError("Unsupported BinaryCIF encoding '" + kind + "'") + } + index = index - 1 + } + match decoded { + Some(value) => value + None => raise BinaryCifError("BinaryCIF data has no encoding") + } +} + +///| +fn bcif_decode_data( + value : BcifMessage, + context : String, +) -> BinaryCifColumnData raise BinaryCifError { + let data = bcif_as_map(value, context) + let bytes = bcif_as_binary( + bcif_required(data, "data", context), + context + ".data", + ) + let encodings = bcif_as_array( + bcif_required(data, "encoding", context), + context + ".encoding", + ) + bcif_decode_raw(bytes, encodings) +} + +///| +fn bcif_parse_column( + value : BcifMessage, + row_count : Int, +) -> BinaryCifColumn raise BinaryCifError { + let column = bcif_as_map(value, "BinaryCIF column") + let name = bcif_as_string( + bcif_required(column, "name", "BinaryCIF column"), + "BinaryCIF column.name", + ) + let data = bcif_decode_data( + bcif_required(column, "data", "BinaryCIF column"), + "BinaryCIF column '" + name + "' data", + ) + if bcif_decoded_length(data) != row_count { + raise BinaryCifError( + "BinaryCIF column '" + name + "' length does not match rowCount", + ) + } + let mask = match column.get("mask") { + None | Some(BcifNil) => Array::make(row_count, 0) + Some(mask_value) => + match bcif_decode_data(mask_value, "BinaryCIF column mask") { + BinaryCifIntegers(values) => { + if values.length() != row_count { + raise BinaryCifError( + "BinaryCIF mask length does not match rowCount", + ) + } + for item in values { + if item < 0 || item > 2 { + raise BinaryCifError("BinaryCIF mask value must be 0, 1, or 2") + } + } + values + } + _ => raise BinaryCifError("BinaryCIF mask must decode to integers") + } + } + BinaryCifColumn::{ name, data, mask } +} + +///| +fn bcif_parse_category( + value : BcifMessage, +) -> BinaryCifCategory raise BinaryCifError { + let category = bcif_as_map(value, "BinaryCIF category") + let name = bcif_as_string( + bcif_required(category, "name", "BinaryCIF category"), + "BinaryCIF category.name", + ) + let row_count = bcif_as_int( + bcif_required(category, "rowCount", "BinaryCIF category"), + "BinaryCIF category.rowCount", + ) + if row_count < 0 { + raise BinaryCifError("BinaryCIF category rowCount is negative") + } + let raw_columns = bcif_as_array( + bcif_required(category, "columns", "BinaryCIF category"), + "BinaryCIF category.columns", + ) + let columns = Array::new(capacity=raw_columns.length()) + let seen : Map[String, Bool] = Map([], capacity=raw_columns.length()) + for raw_column in raw_columns { + let column = bcif_parse_column(raw_column, row_count) + if seen.contains(column.name) { + raise BinaryCifError("Duplicate BinaryCIF column '" + column.name + "'") + } + seen[column.name] = true + columns.push(column) + } + BinaryCifCategory::{ name, row_count, columns } +} + +///| +fn bcif_parse_block( + value : BcifMessage, +) -> BinaryCifDataBlock raise BinaryCifError { + let block = bcif_as_map(value, "BinaryCIF data block") + let header = bcif_as_string( + bcif_required(block, "header", "BinaryCIF data block"), + "BinaryCIF data block.header", + ) + let raw_categories = bcif_as_array( + bcif_required(block, "categories", "BinaryCIF data block"), + "BinaryCIF data block.categories", + ) + let categories = Array::new(capacity=raw_categories.length()) + let seen : Map[String, Bool] = Map([], capacity=raw_categories.length()) + for raw_category in raw_categories { + let category = bcif_parse_category(raw_category) + if seen.contains(category.name) { + raise BinaryCifError( + "Duplicate BinaryCIF category '" + category.name + "'", + ) + } + seen[category.name] = true + categories.push(category) + } + BinaryCifDataBlock::{ header, categories } +} + +///| +pub fn binary_cif_parse( + input : Array[Int], +) -> BinaryCifFile raise BinaryCifError { + if input.length() == 0 { + raise BinaryCifError("Empty BinaryCIF input") + } + for byte in input { + if byte < 0 || byte > 255 { + raise BinaryCifError("BinaryCIF input contains a non-byte value") + } + } + if input.length() >= 2 && input[0] == 0x1F && input[1] == 0x8B { + raise BinaryCifError( + "Gzip-compressed BinaryCIF input must be decompressed first", + ) + } + let reader = BcifReader::{ data: input, position: 0 } + let root = bcif_as_map(bcif_parse_message(reader, 0), "BinaryCIF document") + if reader.position != input.length() { + raise BinaryCifError("Additional bytes follow the BinaryCIF document") + } + let version = bcif_as_string( + bcif_required(root, "version", "BinaryCIF document"), + "BinaryCIF version", + ) + let encoder = bcif_as_string( + bcif_required(root, "encoder", "BinaryCIF document"), + "BinaryCIF encoder", + ) + let raw_blocks = bcif_as_array( + bcif_required(root, "dataBlocks", "BinaryCIF document"), + "BinaryCIF dataBlocks", + ) + if raw_blocks.length() == 0 { + raise BinaryCifError("BinaryCIF document has no data blocks") + } + let data_blocks = Array::new(capacity=raw_blocks.length()) + for raw_block in raw_blocks { + data_blocks.push(bcif_parse_block(raw_block)) + } + BinaryCifFile::{ version, encoder, data_blocks } +} + +// PDB Structure conversion. + +///| +priv struct BcifResidueBuilder { + resname : String + chain_id : Char + sequence_id : Int + insertion_code : Char + hetero_field : String + atoms : Array[Atom] +} + +///| +priv struct BcifChainBuilder { + full_id : String + chain_id : Char + residues : Array[BcifResidueBuilder] +} + +///| +priv struct BcifModelBuilder { + model_number : Int + chains : Array[BcifChainBuilder] +} + +///| +fn bcif_require_column( + category : BinaryCifCategory, + name : String, +) -> BinaryCifColumn raise BinaryCifError { + match category.get_column(name) { + Some(column) => column + None => + raise BinaryCifError( + "BinaryCIF category '" + + category.name + + "' is missing column '" + + name + + "'", + ) + } +} + +///| +fn bcif_required_string( + column : BinaryCifColumn, + row : Int, +) -> String raise BinaryCifError { + match column.string_at(row) { + Some(value) => value + None => + raise BinaryCifError( + "Required BinaryCIF value '" + column.name + "' is missing", + ) + } +} + +///| +fn bcif_required_int( + column : BinaryCifColumn, + row : Int, +) -> Int raise BinaryCifError { + match column.int_at(row) { + Some(value) => value + None => + raise BinaryCifError( + "Required BinaryCIF value '" + column.name + "' is missing", + ) + } +} + +///| +fn bcif_required_double( + column : BinaryCifColumn, + row : Int, +) -> Double raise BinaryCifError { + match column.double_at(row) { + Some(value) => value + None => + raise BinaryCifError( + "Required BinaryCIF value '" + column.name + "' is missing", + ) + } +} + +///| +fn bcif_optional_string( + category : BinaryCifCategory, + name : String, + row : Int, + fallback : String, +) -> String raise BinaryCifError { + match category.get_column(name) { + None => fallback + Some(column) => + match column.string_at(row) { + Some(value) => value + None => fallback + } + } +} + +///| +fn bcif_optional_int( + category : BinaryCifCategory, + name : String, + row : Int, + fallback : Int, +) -> Int raise BinaryCifError { + match category.get_column(name) { + None => fallback + Some(column) => + match column.int_at(row) { + Some(value) => value + None => fallback + } + } +} + +///| +fn bcif_optional_double( + category : BinaryCifCategory, + name : String, + row : Int, + fallback : Double, +) -> Double raise BinaryCifError { + match category.get_column(name) { + None => fallback + Some(column) => + match column.double_at(row) { + Some(value) => value + None => fallback + } + } +} + +///| +fn bcif_first_char_or_space(value : String) -> Char { + if value.length() == 0 || value == "." || value == "?" { + ' ' + } else { + value.unsafe_get(0).unsafe_to_char() + } +} + +///| +fn bcif_hetero_field(group : String, component : String) -> String { + if group == "HETATM" { + if component == "HOH" || component == "WAT" { + "W" + } else { + "H" + } + } else { + " " + } +} + +///| +fn bcif_entry_id(file : BinaryCifFile) -> String raise BinaryCifError { + let block = file.data_blocks[0] + match block.get_category("_entry") { + Some(category) => + match category.get_column("id") { + Some(column) => + match column.string_at(0) { + Some(value) => value + None => block.header + } + None => block.header + } + None => block.header + } +} + +///| +pub fn binary_cif_to_structure( + file : BinaryCifFile, + structure_id? : String = "", +) -> Structure raise BinaryCifError { + if file.data_blocks.length() == 0 { + raise BinaryCifError("BinaryCIF file has no data blocks") + } + let block = file.data_blocks[0] + let atom_site = match block.get_category("_atom_site") { + Some(category) => category + None => raise BinaryCifError("BinaryCIF file has no _atom_site category") + } + let names = bcif_require_column(atom_site, "label_atom_id") + let components = bcif_require_column(atom_site, "label_comp_id") + let chains = bcif_require_column(atom_site, "label_asym_id") + let sequence_ids = bcif_require_column(atom_site, "auth_seq_id") + let xs = bcif_require_column(atom_site, "Cartn_x") + let ys = bcif_require_column(atom_site, "Cartn_y") + let zs = bcif_require_column(atom_site, "Cartn_z") + let builders : Array[BcifModelBuilder] = [] + for row in 0.. 0 { + structure_id + } else { + bcif_entry_id(file) + }, + models~, + ) +} + +///| +pub fn binary_cif_parse_structure( + input : Array[Int], + structure_id? : String = "", +) -> Structure raise BinaryCifError { + binary_cif_to_structure(binary_cif_parse(input), structure_id~) +} + +// A deterministic BinaryCIF fixture used by tests and examples. + +///| +fn bcif_message_map(entries : Array[(String, BcifMessage)]) -> BcifMessage { + let values : Map[String, BcifMessage] = Map([], capacity=entries.length()) + for entry in entries { + let (key, value) = entry + values[key] = value + } + BcifMapValue(values) +} + +///| +fn bcif_message_array(values : Array[BcifMessage]) -> BcifMessage { + BcifArrayValue(values) +} + +///| +fn bcif_write_u16_be(output : Array[Int], value : Int) -> Unit { + output.push((value >> 8) & 0xFF) + output.push(value & 0xFF) +} + +///| +fn bcif_write_u32_be(output : Array[Int], value : Int) -> Unit { + output.push((value >> 24) & 0xFF) + output.push((value >> 16) & 0xFF) + output.push((value >> 8) & 0xFF) + output.push(value & 0xFF) +} + +///| +fn bcif_write_i32_le(output : Array[Int], value : Int) -> Unit { + output.push(value & 0xFF) + output.push((value >> 8) & 0xFF) + output.push((value >> 16) & 0xFF) + output.push((value >> 24) & 0xFF) +} + +///| +fn bcif_pack_length( + output : Array[Int], + small_base : Int, + marker16 : Int, + marker32 : Int, + length : Int, +) -> Unit { + if length < 16 { + output.push(small_base + length) + } else if length <= 65535 { + output.push(marker16) + bcif_write_u16_be(output, length) + } else { + output.push(marker32) + bcif_write_u32_be(output, length) + } +} + +///| +fn bcif_pack_string(output : Array[Int], value : String) -> Unit { + let length = value.length() + if length < 32 { + output.push(0xA0 + length) + } else if length <= 255 { + output.push(0xD9) + output.push(length) + } else if length <= 65535 { + output.push(0xDA) + bcif_write_u16_be(output, length) + } else { + output.push(0xDB) + bcif_write_u32_be(output, length) + } + for i in 0.. Unit { + output.push(0xCB) + if value == 0.0 { + for _ in 0..<8 { + output.push(0) + } + return + } + let sign = if value < 0.0 { 0x80 } else { 0 } + let mut normalized = value.abs() + let mut exponent = 0 + while normalized >= 2.0 { + normalized = normalized / 2.0 + exponent = exponent + 1 + } + while normalized < 1.0 { + normalized = normalized * 2.0 + exponent = exponent - 1 + } + let biased = exponent + 1023 + let mut fraction = normalized - 1.0 + fraction = fraction * 16.0 + let first_nibble = fraction.to_int() + fraction = fraction - first_nibble.to_double() + output.push(sign + (biased >> 4)) + output.push(((biased & 0x0F) << 4) + first_nibble) + for _ in 0..<6 { + fraction = fraction * 256.0 + let byte = fraction.to_int() + output.push(byte) + fraction = fraction - byte.to_double() + } +} + +///| +fn bcif_pack_message(output : Array[Int], value : BcifMessage) -> Unit { + match value { + BcifNil => output.push(0xC0) + BcifBoolean(false) => output.push(0xC2) + BcifBoolean(true) => output.push(0xC3) + BcifIntegerValue(number) => + if number >= 0 && number <= 127 { + output.push(number) + } else if number >= -32 && number < 0 { + output.push(number & 0xFF) + } else { + output.push(0xD2) + bcif_write_u32_be(output, number) + } + BcifFloatValue(number) => bcif_pack_float64(output, number) + BcifTextValue(text) => bcif_pack_string(output, text) + BcifBinaryValue(bytes) => { + if bytes.length() <= 255 { + output.push(0xC4) + output.push(bytes.length()) + } else if bytes.length() <= 65535 { + output.push(0xC5) + bcif_write_u16_be(output, bytes.length()) + } else { + output.push(0xC6) + bcif_write_u32_be(output, bytes.length()) + } + for byte in bytes { + output.push(byte) + } + } + BcifArrayValue(values) => { + bcif_pack_length(output, 0x90, 0xDC, 0xDD, values.length()) + for item in values { + bcif_pack_message(output, item) + } + } + BcifMapValue(values) => { + let keys = values.keys().collect() + bcif_pack_length(output, 0x80, 0xDE, 0xDF, keys.length()) + for key in keys { + bcif_pack_string(output, key) + bcif_pack_message(output, values[key]) + } + } + } +} + +///| +fn bcif_byte_array_encoding(type_code : Int) -> BcifMessage { + bcif_message_map([ + ("kind", BcifTextValue("ByteArray")), + ("type", BcifIntegerValue(type_code)), + ]) +} + +///| +fn bcif_sample_int_column(name : String, values : Array[Int]) -> BcifMessage { + let bytes = Array::new(capacity=values.length() * 4) + for value in values { + bcif_write_i32_le(bytes, value) + } + bcif_message_map([ + ("name", BcifTextValue(name)), + ( + "data", + bcif_message_map([ + ("data", BcifBinaryValue(bytes)), + ("encoding", bcif_message_array([bcif_byte_array_encoding(3)])), + ]), + ), + ]) +} + +///| +fn bcif_sample_fixed_column( + name : String, + scaled_values : Array[Int], + factor : Int, +) -> BcifMessage { + let deltas = Array::new(capacity=scaled_values.length()) + let mut previous = 0 + for value in scaled_values { + deltas.push(value - previous) + previous = value + } + let bytes = Array::new(capacity=deltas.length() * 4) + for value in deltas { + bcif_write_i32_le(bytes, value) + } + bcif_message_map([ + ("name", BcifTextValue(name)), + ( + "data", + bcif_message_map([ + ("data", BcifBinaryValue(bytes)), + ( + "encoding", + bcif_message_array([ + bcif_message_map([ + ("kind", BcifTextValue("FixedPoint")), + ("factor", BcifIntegerValue(factor)), + ("srcType", BcifIntegerValue(33)), + ]), + bcif_message_map([ + ("kind", BcifTextValue("Delta")), + ("origin", BcifIntegerValue(0)), + ("srcType", BcifIntegerValue(3)), + ]), + bcif_byte_array_encoding(3), + ]), + ), + ]), + ), + ]) +} + +///| +fn bcif_sample_string_column( + name : String, + values : Array[String], + mask? : Array[Int] = [], +) -> BcifMessage { + let dictionary : Array[String] = [] + let indices : Array[Int] = [] + let lookup : Map[String, Int] = Map([], capacity=values.length()) + for value in values { + let index = match lookup.get(value) { + Some(existing) => existing + None => { + let created = dictionary.length() + dictionary.push(value) + lookup[value] = created + created + } + } + indices.push(index) + } + let mut string_data = "" + let offsets = [0] + for value in dictionary { + string_data = string_data + value + offsets.push(string_data.length()) + } + let offset_bytes = Array::new(capacity=offsets.length() * 4) + for value in offsets { + bcif_write_i32_le(offset_bytes, value) + } + let index_bytes = Array::new(capacity=indices.length()) + for value in indices { + index_bytes.push(value) + } + let entries : Array[(String, BcifMessage)] = [ + ("name", BcifTextValue(name)), + ( + "data", + bcif_message_map([ + ("data", BcifBinaryValue(index_bytes)), + ( + "encoding", + bcif_message_array([ + bcif_message_map([ + ("kind", BcifTextValue("StringArray")), + ( + "dataEncoding", + bcif_message_array([bcif_byte_array_encoding(4)]), + ), + ("stringData", BcifTextValue(string_data)), + ( + "offsetEncoding", + bcif_message_array([bcif_byte_array_encoding(3)]), + ), + ("offsets", BcifBinaryValue(offset_bytes)), + ]), + ]), + ), + ]), + ), + ] + if mask.length() > 0 { + entries.push( + ( + "mask", + bcif_message_map([ + ("data", BcifBinaryValue(mask)), + ("encoding", bcif_message_array([bcif_byte_array_encoding(4)])), + ]), + ), + ) + } + bcif_message_map(entries) +} + +///| +pub fn binary_cif_sample_bytes() -> Array[Int] { + let atom_count = 8 + let atom_columns = [ + bcif_sample_string_column("group_PDB", [ + "ATOM", "ATOM", "ATOM", "ATOM", "ATOM", "HETATM", "ATOM", "ATOM", + ]), + bcif_sample_int_column("id", [1, 2, 3, 4, 5, 6, 7, 8]), + bcif_sample_string_column("type_symbol", [ + "N", "C", "C", "N", "C", "O", "N", "C", + ]), + bcif_sample_string_column("label_atom_id", [ + "N", "CA", "C", "N", "CA", "O", "N", "CA", + ]), + bcif_sample_string_column( + "label_alt_id", + ["", "", "", "A", "A", "", "", ""], + mask=[1, 1, 1, 0, 0, 1, 2, 1], + ), + bcif_sample_string_column("label_comp_id", [ + "GLY", "GLY", "GLY", "ALA", "ALA", "HOH", "SER", "SER", + ]), + bcif_sample_string_column("label_asym_id", [ + "A", "A", "A", "A", "A", "B", "A", "A", + ]), + bcif_sample_int_column("auth_seq_id", [1, 1, 1, 2, 2, 10, 1, 1]), + bcif_sample_string_column( + "pdbx_PDB_ins_code", + ["", "", "", "A", "A", "", "", ""], + mask=[1, 1, 1, 0, 0, 1, 1, 1], + ), + bcif_sample_fixed_column( + "Cartn_x", + [100, 220, 340, 460, 580, 700, 820, 940], + 100, + ), + bcif_sample_fixed_column( + "Cartn_y", + [200, 300, 400, 500, 600, 700, 800, 900], + 100, + ), + bcif_sample_fixed_column( + "Cartn_z", + [-100, 0, 100, 200, 300, 400, 500, 600], + 100, + ), + bcif_sample_fixed_column( + "occupancy", + [100, 100, 100, 50, 50, 100, 100, 100], + 100, + ), + bcif_sample_fixed_column( + "B_iso_or_equiv", + [1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700], + 100, + ), + bcif_sample_int_column("pdbx_PDB_model_num", [1, 1, 1, 1, 1, 1, 2, 2]), + ] + let document = bcif_message_map([ + ("version", BcifTextValue("0.3.0")), + ("encoder", BcifTextValue("BioSeqs MoonBit")), + ( + "dataBlocks", + bcif_message_array([ + bcif_message_map([ + ("header", BcifTextValue("BCIF")), + ( + "categories", + bcif_message_array([ + bcif_message_map([ + ("name", BcifTextValue("_entry")), + ("rowCount", BcifIntegerValue(1)), + ( + "columns", + bcif_message_array([bcif_sample_string_column("id", ["BCIF"])]), + ), + ]), + bcif_message_map([ + ("name", BcifTextValue("_atom_site")), + ("rowCount", BcifIntegerValue(atom_count)), + ("columns", bcif_message_array(atom_columns)), + ]), + ]), + ), + ]), + ]), + ), + ]) + let output : Array[Int] = [] + bcif_pack_message(output, document) + output +} diff --git a/src/bioc_generics.mbt b/src/bioc_generics.mbt index 6e010b0a..970297e7 100644 --- a/src/bioc_generics.mbt +++ b/src/bioc_generics.mbt @@ -164,7 +164,7 @@ pub fn order_int(arr : Array[Int]) -> Array[Int] { indices.push(i) i = i + 1 } - + let mut sorted = false while !sorted { sorted = true @@ -180,7 +180,7 @@ pub fn order_int(arr : Array[Int]) -> Array[Int] { i = i + 1 } } - + indices } @@ -193,7 +193,7 @@ pub fn order_double(arr : Array[Double]) -> Array[Int] { indices.push(i) i = i + 1 } - + let mut sorted = false while !sorted { sorted = true @@ -209,7 +209,7 @@ pub fn order_double(arr : Array[Double]) -> Array[Int] { i = i + 1 } } - + indices } @@ -241,7 +241,7 @@ pub fn sort_double(arr : Array[Double]) -> Array[Double] { pub fn unique_int(arr : Array[Int]) -> Array[Int] { let result : Array[Int] = Array::new() let seen = Map([], capacity=arr.length()) - + let mut i = 0 while i < arr.length() { if !seen.contains(arr[i].to_string()) { @@ -250,7 +250,7 @@ pub fn unique_int(arr : Array[Int]) -> Array[Int] { } i = i + 1 } - + result } @@ -258,7 +258,7 @@ pub fn unique_int(arr : Array[Int]) -> Array[Int] { pub fn unique_double(arr : Array[Double]) -> Array[Double] { let result : Array[Double] = Array::new() let seen = Map([], capacity=arr.length()) - + let mut i = 0 while i < arr.length() { if !seen.contains(arr[i].to_string()) { @@ -267,7 +267,7 @@ pub fn unique_double(arr : Array[Double]) -> Array[Double] { } i = i + 1 } - + result } @@ -275,7 +275,7 @@ pub fn unique_double(arr : Array[Double]) -> Array[Double] { pub fn unique_string(arr : Array[String]) -> Array[String] { let result : Array[String] = Array::new() let seen = Map([], capacity=arr.length()) - + let mut i = 0 while i < arr.length() { if !seen.contains(arr[i]) { @@ -284,7 +284,7 @@ pub fn unique_string(arr : Array[String]) -> Array[String] { } i = i + 1 } - + result } @@ -292,13 +292,13 @@ pub fn unique_string(arr : Array[String]) -> Array[String] { pub fn match_int(x : Array[Int], table : Array[Int]) -> Array[Int] { let result : Array[Int] = Array::new() let table_map = Map([], capacity=table.length()) - + let mut i = 0 while i < table.length() { table_map.set(table[i].to_string(), i) i = i + 1 } - + i = 0 while i < x.length() { if table_map.contains(x[i].to_string()) { @@ -308,7 +308,7 @@ pub fn match_int(x : Array[Int], table : Array[Int]) -> Array[Int] { } i = i + 1 } - + result } @@ -316,13 +316,13 @@ pub fn match_int(x : Array[Int], table : Array[Int]) -> Array[Int] { pub fn match_string(x : Array[String], table : Array[String]) -> Array[Int] { let result : Array[Int] = Array::new() let table_map = Map([], capacity=table.length()) - + let mut i = 0 while i < table.length() { table_map.set(table[i], i) i = i + 1 } - + i = 0 while i < x.length() { if table_map.contains(x[i]) { @@ -332,7 +332,7 @@ pub fn match_string(x : Array[String], table : Array[String]) -> Array[Int] { } i = i + 1 } - + result } @@ -340,13 +340,13 @@ pub fn match_string(x : Array[String], table : Array[String]) -> Array[Int] { pub fn intersect_int(x : Array[Int], y : Array[Int]) -> Array[Int] { let result : Array[Int] = Array::new() let y_set = Map([], capacity=y.length()) - + let mut i = 0 while i < y.length() { y_set.set(y[i].to_string(), true) i = i + 1 } - + i = 0 while i < x.length() { if y_set.contains(x[i].to_string()) { @@ -354,7 +354,7 @@ pub fn intersect_int(x : Array[Int], y : Array[Int]) -> Array[Int] { } i = i + 1 } - + result } @@ -362,13 +362,13 @@ pub fn intersect_int(x : Array[Int], y : Array[Int]) -> Array[Int] { pub fn intersect_string(x : Array[String], y : Array[String]) -> Array[String] { let result : Array[String] = Array::new() let y_set = Map([], capacity=y.length()) - + let mut i = 0 while i < y.length() { y_set.set(y[i], true) i = i + 1 } - + i = 0 while i < x.length() { if y_set.contains(x[i]) { @@ -376,7 +376,7 @@ pub fn intersect_string(x : Array[String], y : Array[String]) -> Array[String] { } i = i + 1 } - + result } @@ -384,7 +384,7 @@ pub fn intersect_string(x : Array[String], y : Array[String]) -> Array[String] { pub fn union_int(x : Array[Int], y : Array[Int]) -> Array[Int] { let result : Array[Int] = Array::new() let seen = Map([], capacity=x.length() + y.length()) - + let mut i = 0 while i < x.length() { if !seen.contains(x[i].to_string()) { @@ -393,7 +393,7 @@ pub fn union_int(x : Array[Int], y : Array[Int]) -> Array[Int] { } i = i + 1 } - + i = 0 while i < y.length() { if !seen.contains(y[i].to_string()) { @@ -402,7 +402,7 @@ pub fn union_int(x : Array[Int], y : Array[Int]) -> Array[Int] { } i = i + 1 } - + result } @@ -410,7 +410,7 @@ pub fn union_int(x : Array[Int], y : Array[Int]) -> Array[Int] { pub fn union_string(x : Array[String], y : Array[String]) -> Array[String] { let result : Array[String] = Array::new() let seen = Map([], capacity=x.length() + y.length()) - + let mut i = 0 while i < x.length() { if !seen.contains(x[i]) { @@ -419,7 +419,7 @@ pub fn union_string(x : Array[String], y : Array[String]) -> Array[String] { } i = i + 1 } - + i = 0 while i < y.length() { if !seen.contains(y[i]) { @@ -428,7 +428,7 @@ pub fn union_string(x : Array[String], y : Array[String]) -> Array[String] { } i = i + 1 } - + result } @@ -436,13 +436,13 @@ pub fn union_string(x : Array[String], y : Array[String]) -> Array[String] { pub fn setdiff_int(x : Array[Int], y : Array[Int]) -> Array[Int] { let result : Array[Int] = Array::new() let y_set = Map([], capacity=y.length()) - + let mut i = 0 while i < y.length() { y_set.set(y[i].to_string(), true) i = i + 1 } - + i = 0 while i < x.length() { if !y_set.contains(x[i].to_string()) { @@ -450,7 +450,7 @@ pub fn setdiff_int(x : Array[Int], y : Array[Int]) -> Array[Int] { } i = i + 1 } - + result } @@ -458,13 +458,13 @@ pub fn setdiff_int(x : Array[Int], y : Array[Int]) -> Array[Int] { pub fn setdiff_string(x : Array[String], y : Array[String]) -> Array[String] { let result : Array[String] = Array::new() let y_set = Map([], capacity=y.length()) - + let mut i = 0 while i < y.length() { y_set.set(y[i], true) i = i + 1 } - + i = 0 while i < x.length() { if !y_set.contains(x[i]) { @@ -472,14 +472,14 @@ pub fn setdiff_string(x : Array[String], y : Array[String]) -> Array[String] { } i = i + 1 } - + result } ///| pub fn table_int(arr : Array[Int]) -> Map[String, Int] { let result = Map([], capacity=arr.length()) - + let mut i = 0 while i < arr.length() { let key = arr[i].to_string() @@ -490,14 +490,14 @@ pub fn table_int(arr : Array[Int]) -> Map[String, Int] { } i = i + 1 } - + result } ///| pub fn table_string(arr : Array[String]) -> Map[String, Int] { let result = Map([], capacity=arr.length()) - + let mut i = 0 while i < arr.length() { let key = arr[i] @@ -508,7 +508,7 @@ pub fn table_string(arr : Array[String]) -> Map[String, Int] { } i = i + 1 } - + result } @@ -582,7 +582,7 @@ pub fn rep_string(x : String, times : Int) -> Array[String] { pub fn seq_int(from : Int, to : Int, by? : Int) -> Array[Int] { let step = if by is Some(_) { by.unwrap() } else { 1 } let result : Array[Int] = Array::new() - + if step > 0 { let mut i = from while i <= to { @@ -596,7 +596,7 @@ pub fn seq_int(from : Int, to : Int, by? : Int) -> Array[Int] { i = i + step } } - + result } @@ -604,7 +604,7 @@ pub fn seq_int(from : Int, to : Int, by? : Int) -> Array[Int] { pub fn seq_double(from : Double, to : Double, by? : Double) -> Array[Double] { let step = if by is Some(_) { by.unwrap() } else { 1.0 } let result : Array[Double] = Array::new() - + if step > 0.0 { let mut i = from while i <= to { @@ -618,7 +618,7 @@ pub fn seq_double(from : Double, to : Double, by? : Double) -> Array[Double] { i = i + step } } - + result } @@ -740,10 +740,10 @@ pub fn cbind(arrays : Array[Array[Double]]) -> Array[Array[Double]] { if arrays.length() == 0 { return [] } - + let nrow = arrays[0].length() let result : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < nrow { let row : Array[Double] = Array::new() @@ -759,19 +759,19 @@ pub fn cbind(arrays : Array[Array[Double]]) -> Array[Array[Double]] { result.push(row) i = i + 1 } - + result } ///| pub fn rbind(arrays : Array[Array[Double]]) -> Array[Array[Double]] { let result : Array[Array[Double]] = Array::new() - + let mut i = 0 while i < arrays.length() { result.push(arrays[i].copy()) i = i + 1 } - + result } diff --git a/src/bioc_neighbors.mbt b/src/bioc_neighbors.mbt index a0ce946d..91b07ca2 100644 --- a/src/bioc_neighbors.mbt +++ b/src/bioc_neighbors.mbt @@ -166,7 +166,13 @@ fn km_init_centroids( let (new_lcg, rand_val) = lcg_next_double(lcg) lcg = new_lcg let idx = (rand_val * n_points.to_double()).to_int() - let actual_idx = if idx >= n_points { n_points - 1 } else { if idx < 0 { 0 } else { idx } } + let actual_idx = if idx >= n_points { + n_points - 1 + } else if idx < 0 { + 0 + } else { + idx + } if !used_indices[actual_idx] { centroids.push(data[actual_idx].copy()) used_indices[actual_idx] = true @@ -249,12 +255,20 @@ fn km_kmeans( if n_points == 0 { return ([], []) } - let actual_clusters = if n_clusters > n_points { n_points } else { n_clusters } + let actual_clusters = if n_clusters > n_points { + n_points + } else { + n_clusters + } let mut centroids = km_init_centroids(data, n_points, n_dim, actual_clusters) let mut assignments = Array::make(n_points, 0) for iter = 0; iter < max_iter; iter = iter + 1 { - let new_assignments = km_assign_points(data, centroids, n_points, n_dim, actual_clusters) - let new_centroids = km_update_centroids(data, new_assignments, n_points, n_dim, actual_clusters) + let new_assignments = km_assign_points( + data, centroids, n_points, n_dim, actual_clusters, + ) + let new_centroids = km_update_centroids( + data, new_assignments, n_points, n_dim, actual_clusters, + ) let mut max_shift = 0.0 for c = 0; c < actual_clusters; c = c + 1 { let shift = euclidean_distance(centroids[c], new_centroids[c]) @@ -282,7 +296,7 @@ pub fn build_knn_index( match options.method { KMKNN => build_kmknn_index(data, n_points, n_dim, options) Annoy => build_annoy_index(data, n_points, n_dim, options) - BruteForce => { + BruteForce => IndexResult::{ method: BruteForce, n_points, @@ -291,7 +305,6 @@ pub fn build_knn_index( centroids: [], tree_nodes: [], } - } } } @@ -317,8 +330,12 @@ pub fn build_kmknn_index( let n_clusters = if n_points <= 10 { 1 } else { - let sqrt_val = ((n_points.to_double()).sqrt()).to_int() - if sqrt_val < 2 { 2 } else { sqrt_val } + let sqrt_val = n_points.to_double().sqrt().to_int() + if sqrt_val < 2 { + 2 + } else { + sqrt_val + } } let (centroids, _) = km_kmeans(data, n_points, n_dim, n_clusters, 20) IndexResult::{ @@ -368,7 +385,13 @@ fn build_single_annoy_tree( let (new_lcg, rand_dim) = lcg_next_double(current_lcg) current_lcg = new_lcg let split_dim = (rand_dim * n_dim.to_double()).to_int() - let actual_dim = if split_dim >= n_dim { n_dim - 1 } else { if split_dim < 0 { 0 } else { split_dim } } + let actual_dim = if split_dim >= n_dim { + n_dim - 1 + } else if split_dim < 0 { + 0 + } else { + split_dim + } let (new_lcg2, rand_val) = lcg_next_double(current_lcg) current_lcg = new_lcg2 @@ -377,8 +400,12 @@ fn build_single_annoy_tree( let mut max_val = data[point_indices[0]][actual_dim] for i = 1; i < n; i = i + 1 { let v = data[point_indices[i]][actual_dim] - if v < min_val { min_val = v } - if v > max_val { max_val = v } + if v < min_val { + min_val = v + } + if v > max_val { + max_val = v + } } let split_val = min_val + rand_val * (max_val - min_val) @@ -446,7 +473,9 @@ fn build_single_annoy_tree( point_indices: [], } - let idx = nodes_copy.length() - (nodes_after_right.length() - nodes_after_left.length()) - 1 + let idx = nodes_copy.length() - + (nodes_after_right.length() - nodes_after_left.length()) - + 1 nodes_copy[idx] = node (nodes_copy, node_id, current_lcg) @@ -476,7 +505,13 @@ pub fn build_annoy_index( 5 } else { let d = (@math.log2(n_points.to_double()) + 1.0).to_int() - if d > 20 { 20 } else { if d < 3 { 3 } else { d } } + if d > 20 { + 20 + } else if d < 3 { + 3 + } else { + d + } } let mut all_nodes : Array[AnnoyNode] = [] @@ -496,7 +531,13 @@ pub fn build_annoy_index( let (new_lcg2, rand_val) = lcg_next_double(shuf_lcg) shuf_lcg = new_lcg2 let idx = (rand_val * indices_left.length().to_double()).to_int() - let actual_idx = if idx >= indices_left.length() { indices_left.length() - 1 } else { if idx < 0 { 0 } else { idx } } + let actual_idx = if idx >= indices_left.length() { + indices_left.length() - 1 + } else if idx < 0 { + 0 + } else { + idx + } shuffled.push(indices_left[actual_idx]) let new_left : Array[Int] = [] for i = 0; i < indices_left.length(); i = i + 1 { @@ -509,14 +550,7 @@ pub fn build_annoy_index( let tree_start_id = all_nodes.length() let (tree_nodes, _, final_lcg) = build_single_annoy_tree( - data, - shuffled, - n_dim, - all_nodes, - tree_start_id, - 0, - max_depth, - lcg, + data, shuffled, n_dim, all_nodes, tree_start_id, 0, max_depth, lcg, ) all_nodes = tree_nodes lcg = final_lcg @@ -549,8 +583,14 @@ pub fn knn_quick_sort( let pivot = sorted_distances[pivot_idx] let pivot_index = sorted_indices[pivot_idx] while i <= j { - while sorted_distances[i] < pivot || (sorted_distances[i] == pivot && sorted_indices[i] < pivot_index) { i = i + 1 } - while sorted_distances[j] > pivot || (sorted_distances[j] == pivot && sorted_indices[j] > pivot_index) { j = j - 1 } + while sorted_distances[i] < pivot || + (sorted_distances[i] == pivot && sorted_indices[i] < pivot_index) { + i = i + 1 + } + while sorted_distances[j] > pivot || + (sorted_distances[j] == pivot && sorted_indices[j] > pivot_index) { + j = j - 1 + } if i <= j { let tmp_idx = sorted_indices[i] sorted_indices[i] = sorted_indices[j] @@ -574,14 +614,13 @@ pub fn knn_quick_sort( ///| /// Find indices of the k smallest distances. -pub fn knn_find_k_smallest( - distances : Array[Double], - k : Int, -) -> Array[Int] { +pub fn knn_find_k_smallest(distances : Array[Double], k : Int) -> Array[Int] { let n = distances.length() if k >= n { let result = Array::make(n, 0) - for i = 0; i < n; i = i + 1 { result[i] = i } + for i = 0; i < n; i = i + 1 { + result[i] = i + } return result } let indices : Array[Int] = [] @@ -610,12 +649,7 @@ pub fn knn_brute_force( ) -> KNNResult { if n_query == 0 || n_points == 0 || k == 0 { let actual_k = if k > n_points { n_points } else { k } - return KNNResult::{ - indices: [], - distances: [], - n_query, - k: actual_k, - } + return KNNResult::{ indices: [], distances: [], n_query, k: actual_k } } let actual_k = if k > n_points { n_points } else { k } let all_indices : Array[Array[Int]] = [] @@ -628,7 +662,12 @@ pub fn knn_brute_force( dists[i] = knn_compute_distance(query_point, data[i], distance) indices[i] = i } - let (sorted_indices, sorted_distances) = knn_quick_sort(dists, indices, 0, n_points - 1) + let (sorted_indices, sorted_distances) = knn_quick_sort( + dists, + indices, + 0, + n_points - 1, + ) let top_indices = Array::make(actual_k, 0) let top_distances = Array::make(actual_k, 0.0) for i = 0; i < actual_k; i = i + 1 { @@ -638,7 +677,12 @@ pub fn knn_brute_force( all_indices.push(top_indices) all_distances.push(top_distances) } - KNNResult::{ indices: all_indices, distances: all_distances, n_query, k: actual_k } + KNNResult::{ + indices: all_indices, + distances: all_distances, + n_query, + k: actual_k, + } } ///| @@ -652,15 +696,16 @@ pub fn run_knn( match index.method { KMKNN => knn_kmknn(index, query_data, n_query, options) Annoy => knn_annoy(index, query_data, n_query, options) - BruteForce => knn_brute_force( - index.data, - index.n_points, - index.n_dim, - query_data, - n_query, - options.k, - options.distance, - ) + BruteForce => + knn_brute_force( + index.data, + index.n_points, + index.n_dim, + query_data, + n_query, + options.k, + options.distance, + ) } } @@ -674,15 +719,18 @@ pub fn knn_kmknn( options : KNNOptions, ) -> KNNResult { if n_query == 0 || index.n_points == 0 || options.k == 0 { - let actual_k = if options.k > index.n_points { index.n_points } else { options.k } - return KNNResult::{ - indices: [], - distances: [], - n_query, - k: actual_k, + let actual_k = if options.k > index.n_points { + index.n_points + } else { + options.k } + return KNNResult::{ indices: [], distances: [], n_query, k: actual_k } + } + let actual_k = if options.k > index.n_points { + index.n_points + } else { + options.k } - let actual_k = if options.k > index.n_points { index.n_points } else { options.k } let n_clusters = index.centroids.length() let all_indices : Array[Array[Int]] = [] let all_distances : Array[Array[Double]] = [] @@ -716,8 +764,15 @@ pub fn knn_kmknn( centroid_dists[c] = euclidean_distance(query, index.centroids[c]) } let sorted_centroid_indices = Array::make(n_clusters, 0) - for c = 0; c < n_clusters; c = c + 1 { sorted_centroid_indices[c] = c } - let (sorted_c_idx, _) = knn_quick_sort(centroid_dists, sorted_centroid_indices, 0, n_clusters - 1) + for c = 0; c < n_clusters; c = c + 1 { + sorted_centroid_indices[c] = c + } + let (sorted_c_idx, _) = knn_quick_sort( + centroid_dists, + sorted_centroid_indices, + 0, + n_clusters - 1, + ) let candidates : Array[Int] = [] let mut cluster_idx = 0 @@ -733,7 +788,10 @@ pub fn knn_kmknn( for i = 0; i < index.n_points; i = i + 1 { let mut found = false for c = 0; c < candidates.length(); c = c + 1 { - if candidates[c] == i { found = true; break } + if candidates[c] == i { + found = true + break + } } if !found { candidates.push(i) @@ -743,9 +801,18 @@ pub fn knn_kmknn( let cand_dists = Array::make(candidates.length(), 0.0) for i = 0; i < candidates.length(); i = i + 1 { - cand_dists[i] = knn_compute_distance(query, index.data[candidates[i]], options.distance) + cand_dists[i] = knn_compute_distance( + query, + index.data[candidates[i]], + options.distance, + ) } - let (sorted_idx, sorted_dist) = knn_quick_sort(cand_dists, candidates, 0, candidates.length() - 1) + let (sorted_idx, sorted_dist) = knn_quick_sort( + cand_dists, + candidates, + 0, + candidates.length() - 1, + ) let top_indices = Array::make(actual_k, 0) let top_distances = Array::make(actual_k, 0.0) @@ -757,7 +824,12 @@ pub fn knn_kmknn( all_distances.push(top_distances) } - KNNResult::{ indices: all_indices, distances: all_distances, n_query, k: actual_k } + KNNResult::{ + indices: all_indices, + distances: all_distances, + n_query, + k: actual_k, + } } ///| @@ -770,17 +842,24 @@ pub fn knn_annoy( options : KNNOptions, ) -> KNNResult { if n_query == 0 || index.n_points == 0 || options.k == 0 { - let actual_k = if options.k > index.n_points { index.n_points } else { options.k } - return KNNResult::{ - indices: [], - distances: [], - n_query, - k: actual_k, + let actual_k = if options.k > index.n_points { + index.n_points + } else { + options.k } + return KNNResult::{ indices: [], distances: [], n_query, k: actual_k } + } + let actual_k = if options.k > index.n_points { + index.n_points + } else { + options.k } - let actual_k = if options.k > index.n_points { index.n_points } else { options.k } let n_trees = options.n_trees - let nodes_per_tree = if n_trees > 0 { index.tree_nodes.length() / n_trees } else { 0 } + let nodes_per_tree = if n_trees > 0 { + index.tree_nodes.length() / n_trees + } else { + 0 + } let all_indices : Array[Array[Int]] = [] let all_distances : Array[Array[Double]] = [] @@ -822,7 +901,10 @@ pub fn knn_annoy( let point_idx = node.point_indices[pi] let mut already = false for c = 0; c < candidates.length(); c = c + 1 { - if candidates[c] == point_idx { already = true; break } + if candidates[c] == point_idx { + already = true + break + } } if !already { candidates.push(point_idx) @@ -879,9 +961,18 @@ pub fn knn_annoy( if candidates.length() > actual_k { let cand_dists = Array::make(candidates.length(), 0.0) for i = 0; i < candidates.length(); i = i + 1 { - cand_dists[i] = knn_compute_distance(query, index.data[candidates[i]], options.distance) + cand_dists[i] = knn_compute_distance( + query, + index.data[candidates[i]], + options.distance, + ) } - let (sorted_idx, sorted_dist) = knn_quick_sort(cand_dists, candidates, 0, candidates.length() - 1) + let (sorted_idx, sorted_dist) = knn_quick_sort( + cand_dists, + candidates, + 0, + candidates.length() - 1, + ) let top_indices = Array::make(actual_k, 0) let top_distances = Array::make(actual_k, 0.0) for i = 0; i < actual_k; i = i + 1 { @@ -893,9 +984,18 @@ pub fn knn_annoy( } else if candidates.length() > 0 { let cand_dists = Array::make(candidates.length(), 0.0) for i = 0; i < candidates.length(); i = i + 1 { - cand_dists[i] = knn_compute_distance(query, index.data[candidates[i]], options.distance) + cand_dists[i] = knn_compute_distance( + query, + index.data[candidates[i]], + options.distance, + ) } - let (sorted_idx, sorted_dist) = knn_quick_sort(cand_dists, candidates, 0, candidates.length() - 1) + let (sorted_idx, sorted_dist) = knn_quick_sort( + cand_dists, + candidates, + 0, + candidates.length() - 1, + ) let top_indices = Array::make(sorted_idx.length(), 0) let top_distances = Array::make(sorted_dist.length(), 0.0) for i = 0; i < sorted_idx.length(); i = i + 1 { @@ -910,5 +1010,10 @@ pub fn knn_annoy( } } - KNNResult::{ indices: all_indices, distances: all_distances, n_query, k: actual_k } -} \ No newline at end of file + KNNResult::{ + indices: all_indices, + distances: all_distances, + n_query, + k: actual_k, + } +} diff --git a/src/bioc_parallel.mbt b/src/bioc_parallel.mbt index d9704c60..d184e113 100644 --- a/src/bioc_parallel.mbt +++ b/src/bioc_parallel.mbt @@ -10,12 +10,21 @@ pub struct BPPARAM { } ///| -pub fn BPPARAM::new(workers : Int, progressbar : Bool, timeout : Int) -> BPPARAM { +pub fn BPPARAM::new( + workers : Int, + progressbar : Bool, + timeout : Int, +) -> BPPARAM { BPPARAM::{ workers, progressbar, timeout, log_file: "" } } ///| -pub fn BPPARAM::new_with_log(workers : Int, progressbar : Bool, timeout : Int, log_file : String) -> BPPARAM { +pub fn BPPARAM::new_with_log( + workers : Int, + progressbar : Bool, + timeout : Int, + log_file : String, +) -> BPPARAM { BPPARAM::{ workers, progressbar, timeout, log_file } } @@ -70,13 +79,13 @@ pub fn bp_add_task(job : BPJob, task : Task) -> BPJob { ///| pub fn bp_sum(data : Array[Array[Double]], params : BPPARAM) -> Double { let mut total = 0.0 - + let chunk_size = if data.length() % params.workers == 0 { data.length() / params.workers } else { data.length() / params.workers + 1 } - + let mut start = 0 while start < data.length() { let end = if start + chunk_size < data.length() { @@ -84,7 +93,7 @@ pub fn bp_sum(data : Array[Array[Double]], params : BPPARAM) -> Double { } else { data.length() } - + let mut i = start while i < end { let mut j = 0 @@ -94,10 +103,10 @@ pub fn bp_sum(data : Array[Array[Double]], params : BPPARAM) -> Double { } i = i + 1 } - + start = end } - + total } @@ -105,14 +114,18 @@ pub fn bp_sum(data : Array[Array[Double]], params : BPPARAM) -> Double { pub fn bp_mean(data : Array[Array[Double]], params : BPPARAM) -> Double { let total = bp_sum(data, params) let mut count = 0.0 - + let mut i = 0 while i < data.length() { count = count + data[i].length().to_double() i = i + 1 } - - if count > 0.0 { total / count } else { 0.0 } + + if count > 0.0 { + total / count + } else { + 0.0 + } } ///| @@ -123,14 +136,17 @@ pub fn bp_mean_simple(data : Array[Double], workers : Int) -> Double { } ///| -pub fn bp_colsum(data : Array[Array[Double]], params : BPPARAM) -> Array[Double] { +pub fn bp_colsum( + data : Array[Array[Double]], + params : BPPARAM, +) -> Array[Double] { if data.length() == 0 { return Array::new() } - + let col_count = data[0].length() let results : Array[Double] = Array::new() - + let mut j = 0 while j < col_count { let mut sum = 0.0 @@ -144,14 +160,17 @@ pub fn bp_colsum(data : Array[Array[Double]], params : BPPARAM) -> Array[Double] results.push(sum) j = j + 1 } - + results } ///| -pub fn bp_rowsum(data : Array[Array[Double]], params : BPPARAM) -> Array[Double] { +pub fn bp_rowsum( + data : Array[Array[Double]], + params : BPPARAM, +) -> Array[Double] { let results : Array[Double] = Array::new() - + let mut i = 0 while i < data.length() { let mut sum = 0.0 @@ -163,20 +182,23 @@ pub fn bp_rowsum(data : Array[Array[Double]], params : BPPARAM) -> Array[Double] results.push(sum) i = i + 1 } - + results } ///| -pub fn bp_parallelize(data : Array[Double], workers : Int) -> Array[Array[Double]] { +pub fn bp_parallelize( + data : Array[Double], + workers : Int, +) -> Array[Array[Double]] { let chunks : Array[Array[Double]] = Array::new() - + let chunk_size = if data.length() % workers == 0 { data.length() / workers } else { data.length() / workers + 1 } - + let mut start = 0 while start < data.length() { let chunk : Array[Double] = Array::new() @@ -185,24 +207,24 @@ pub fn bp_parallelize(data : Array[Double], workers : Int) -> Array[Array[Double } else { data.length() } - + let mut i = start while i < end { chunk.push(data[i]) i = i + 1 } - + chunks.push(chunk) start = end } - + chunks } ///| pub fn bp_run(job : BPJob) -> Array[BPResult] { let results : Array[BPResult] = Array::new() - + let mut i = 0 while i < job.tasks.length() { let task = job.tasks[i] @@ -215,30 +237,33 @@ pub fn bp_run(job : BPJob) -> Array[BPResult] { results.push(BPResult::success(sum, i % job.params.workers)) i = i + 1 } - + results } ///| -pub fn bp_map_sum(data : Array[Array[Double]], params : BPPARAM) -> Array[BPResult] { +pub fn bp_map_sum( + data : Array[Array[Double]], + params : BPPARAM, +) -> Array[BPResult] { let results : Array[BPResult] = Array::new() - + let chunk_size = if data.length() % params.workers == 0 { data.length() / params.workers } else { data.length() / params.workers + 1 } - + let mut worker_id = 0 let mut start = 0 - + while start < data.length() { let end = if start + chunk_size < data.length() { start + chunk_size } else { data.length() } - + let mut i = start while i < end { let mut sum = 0.0 @@ -250,34 +275,37 @@ pub fn bp_map_sum(data : Array[Array[Double]], params : BPPARAM) -> Array[BPResu results.push(BPResult::success(sum, worker_id)) i = i + 1 } - + worker_id = worker_id + 1 start = end } - + results } ///| -pub fn bp_map_mean(data : Array[Array[Double]], params : BPPARAM) -> Array[BPResult] { +pub fn bp_map_mean( + data : Array[Array[Double]], + params : BPPARAM, +) -> Array[BPResult] { let results : Array[BPResult] = Array::new() - + let chunk_size = if data.length() % params.workers == 0 { data.length() / params.workers } else { data.length() / params.workers + 1 } - + let mut worker_id = 0 let mut start = 0 - + while start < data.length() { let end = if start + chunk_size < data.length() { start + chunk_size } else { data.length() } - + let mut i = start while i < end { let mut sum = 0.0 @@ -286,15 +314,19 @@ pub fn bp_map_mean(data : Array[Array[Double]], params : BPPARAM) -> Array[BPRes sum = sum + data[i][j] j = j + 1 } - let mean = if data[i].length() > 0 { sum / data[i].length().to_double() } else { 0.0 } + let mean = if data[i].length() > 0 { + sum / data[i].length().to_double() + } else { + 0.0 + } results.push(BPResult::success(mean, worker_id)) i = i + 1 } - + worker_id = worker_id + 1 start = end } - + results } @@ -307,7 +339,7 @@ pub fn create_example_bpparam() -> BPPARAM { pub fn create_example_bpjob() -> BPJob { let params = BPPARAM::new(2, false, 120) let job = BPJob::new("example_job", params) - + let task1_data : Array[Double] = Array::new() task1_data.push(1.0) task1_data.push(2.0) @@ -315,7 +347,7 @@ pub fn create_example_bpjob() -> BPJob { task1_data.push(4.0) task1_data.push(5.0) job.tasks.push(Task::new("task1", task1_data, 1)) - + let task2_data : Array[Double] = Array::new() task2_data.push(6.0) task2_data.push(7.0) @@ -323,7 +355,7 @@ pub fn create_example_bpjob() -> BPJob { task2_data.push(9.0) task2_data.push(10.0) job.tasks.push(Task::new("task2", task2_data, 1)) - + let task3_data : Array[Double] = Array::new() task3_data.push(11.0) task3_data.push(12.0) @@ -331,7 +363,7 @@ pub fn create_example_bpjob() -> BPJob { task3_data.push(14.0) task3_data.push(15.0) job.tasks.push(Task::new("task3", task3_data, 2)) - + let task4_data : Array[Double] = Array::new() task4_data.push(16.0) task4_data.push(17.0) @@ -339,6 +371,6 @@ pub fn create_example_bpjob() -> BPJob { task4_data.push(19.0) task4_data.push(20.0) job.tasks.push(Task::new("task4", task4_data, 2)) - + job -} \ No newline at end of file +} diff --git a/src/bioc_singular.mbt b/src/bioc_singular.mbt index 62682826..719ec424 100644 --- a/src/bioc_singular.mbt +++ b/src/bioc_singular.mbt @@ -25,15 +25,21 @@ pub enum SVDMethod { ///| /// Create ExactSVD method variant. -pub fn exact_svd_method() -> SVDMethod { ExactSVD } +pub fn exact_svd_method() -> SVDMethod { + ExactSVD +} ///| /// Create IRLBA method variant. -pub fn irlba_method() -> SVDMethod { IRLBA } +pub fn irlba_method() -> SVDMethod { + IRLBA +} ///| /// Create Randomized method variant. -pub fn randomized_method() -> SVDMethod { Randomized } +pub fn randomized_method() -> SVDMethod { + Randomized +} ///| /// Options for SVD computation. @@ -558,13 +564,7 @@ pub fn svd_options_full( n_oversamples : Int, method : SVDMethod, ) -> SVDOptions { - SVDOptions::{ - method, - rank, - tol, - max_iter, - n_oversamples, - } + SVDOptions::{ method, rank, tol, max_iter, n_oversamples } } // ============================================================================ @@ -598,7 +598,9 @@ pub fn run_exact_svd( if ncol <= nrow { let ata = svd_ata(matrix, nrow, ncol) let (eigenvalues, eigenvectors) = eigh_symmetric(ata, ncol) - let (sorted_evals, sorted_evecs) = sort_eigen(eigenvalues, eigenvectors, ncol) + let (sorted_evals, sorted_evecs) = sort_eigen( + eigenvalues, eigenvectors, ncol, + ) let d = Array::make(effective_rank, 0.0) let v : Array[Array[Double]] = Array::new() @@ -634,7 +636,9 @@ pub fn run_exact_svd( } else { let aat = svd_aat(matrix, nrow, ncol) let (eigenvalues, eigenvectors) = eigh_symmetric(aat, nrow) - let (sorted_evals, sorted_evecs) = sort_eigen(eigenvalues, eigenvectors, nrow) + let (sorted_evals, sorted_evecs) = sort_eigen( + eigenvalues, eigenvectors, nrow, + ) let d = Array::make(effective_rank, 0.0) let u : Array[Array[Double]] = Array::new() @@ -706,11 +710,7 @@ fn svd_bidiagonal_small( let u : Array[Array[Double]] = Array::new() for i = 0; i < n; i = i + 1 { - let sigma = if sorted_evals[i] > 0.0 { - sorted_evals[i].sqrt() - } else { - 0.0 - } + let sigma = if sorted_evals[i] > 0.0 { sorted_evals[i].sqrt() } else { 0.0 } d[i] = sigma let v_col = Array::make(n, 0.0) @@ -818,7 +818,8 @@ pub fn run_irlba( for j = 0; j < l; j = j + 1 { let u_new = svd_mat_vec(matrix, nrow, ncol, v_prev) for k = 0; k < nrow; k = k + 1 { - u_new[k] = u_new[k] - beta_prev * u_vecs[if j > 0 { j - 1 } else { 0 }][k] + u_new[k] = u_new[k] - + beta_prev * u_vecs[if j > 0 { j - 1 } else { 0 }][k] } let mut alpha = 0.0 @@ -1092,15 +1093,9 @@ pub fn run_svd( options : SVDOptions, ) -> SVResult { match options.method { - ExactSVD => { - run_exact_svd(matrix, nrow, ncol, options.rank) - } - IRLBA => { - run_irlba(matrix, nrow, ncol, options.rank, options) - } - Randomized => { - run_randomized_svd(matrix, nrow, ncol, options.rank, options) - } + ExactSVD => run_exact_svd(matrix, nrow, ncol, options.rank) + IRLBA => run_irlba(matrix, nrow, ncol, options.rank, options) + Randomized => run_randomized_svd(matrix, nrow, ncol, options.rank, options) } } @@ -1159,10 +1154,7 @@ pub fn create_test_matrix() -> (Array[Double], Int, Int) { let nrow = 5 let ncol = 4 let data = [ - 1.0, 2.0, 3.0, 4.0, - 2.0, 3.0, 4.0, 5.0, - 3.0, 4.0, 5.0, 6.0, - 4.0, 5.0, 6.0, 7.0, + 1.0, 2.0, 3.0, 4.0, 2.0, 3.0, 4.0, 5.0, 3.0, 4.0, 5.0, 6.0, 4.0, 5.0, 6.0, 7.0, 5.0, 6.0, 7.0, 8.0, ] (data, nrow, ncol) @@ -1178,4 +1170,4 @@ pub fn run_svd_truncated( ) -> SVResult { let opts = svd_options_full(k, 1.0e-7, 1000, 10, ExactSVD) run_svd(matrix, nrow, ncol, opts) -} \ No newline at end of file +} diff --git a/src/biostrings.mbt b/src/biostrings.mbt index 2eb195cb..e22de889 100644 --- a/src/biostrings.mbt +++ b/src/biostrings.mbt @@ -703,7 +703,9 @@ pub fn MatchPatternResult::widths(self : MatchPatternResult) -> Array[Int] { ///| /// Get all mismatch counts. -pub fn MatchPatternResult::mismatch_counts(self : MatchPatternResult) -> Array[Int] { +pub fn MatchPatternResult::mismatch_counts( + self : MatchPatternResult, +) -> Array[Int] { self.hits.map(fn(h) { h.mismatches }) } @@ -723,7 +725,13 @@ pub fn match_pattern( let hits : Array[MatchHit] = Array::new() if n == 0 || m == 0 { - return MatchPatternResult::{ hits, pattern, subject, max_mismatches, with_indels } + return MatchPatternResult::{ + hits, + pattern, + subject, + max_mismatches, + with_indels, + } } let max_errors = if with_indels { max_mismatches } else { max_mismatches } @@ -772,7 +780,7 @@ pub fn vmatch_pattern( with_indels? : Bool = false, ) -> Array[MatchPatternResult] { subjects.map(fn(subj) { - match_pattern(pattern=pattern, subject=subj, max_mismatches=max_mismatches, with_indels=with_indels) + match_pattern(pattern~, subject=subj, max_mismatches~, with_indels~) }) } @@ -870,10 +878,7 @@ pub fn find_inverted_repeats( ///| /// Find all occurrences of a motif allowing IUPAC ambiguity codes. -pub fn find_motif_iupac( - pattern : String, - subject : String, -) -> Array[Int] { +pub fn find_motif_iupac(pattern : String, subject : String) -> Array[Int] { let n = pattern.length() let m = subject.length() let positions : Array[Int] = Array::new() @@ -967,7 +972,9 @@ pub fn letter_frequency_matrix(seq : String) -> Map[String, Array[Int]] { let matrix : Map[String, Array[Int]] = Map([], capacity=8) // Initialize for DNA alphabet - let bases = ["A", "C", "G", "T", "U", "N", "R", "Y", "S", "W", "K", "M", "B", "D", "H", "V"] + let bases = [ + "A", "C", "G", "T", "U", "N", "R", "Y", "S", "W", "K", "M", "B", "D", "H", "V", + ] for base in bases { matrix.set(base, Array::make(n, 0)) } @@ -1007,7 +1014,9 @@ pub fn letter_frequency_matrix(seq : String) -> Map[String, Array[Int]] { /// Compute consensus sequence from a multiple sequence alignment. /// Returns the most frequent nucleotide at each position. pub fn biostrings_consensus_sequence(alignment : Array[String]) -> String { - if alignment.length() == 0 { return "" } + if alignment.length() == 0 { + return "" + } let n_seqs = alignment.length() let seq_len = alignment[0].length() @@ -1048,24 +1057,75 @@ pub fn biostrings_translate(seq : String, frame? : Int = 0) -> String { let aa_count = (n - offset) / 3 let result : FixedArray[UInt16] = FixedArray::make(aa_count, 0) - let codon_table : Map[String, UInt16] = Map([ - ("TTT", 70), ("TTC", 70), ("TTA", 76), ("TTG", 76), - ("CTT", 76), ("CTC", 76), ("CTA", 76), ("CTG", 76), - ("ATT", 73), ("ATC", 73), ("ATA", 73), ("ATG", 77), - ("GTT", 86), ("GTC", 86), ("GTA", 86), ("GTG", 86), - ("TCT", 83), ("TCC", 83), ("TCA", 83), ("TCG", 83), - ("CCT", 80), ("CCC", 80), ("CCA", 80), ("CCG", 80), - ("ACT", 65), ("ACC", 65), ("ACA", 65), ("ACG", 65), - ("GCT", 65), ("GCC", 65), ("GCA", 65), ("GCG", 65), - ("TAT", 89), ("TAC", 89), ("TAA", 42), ("TAG", 42), - ("CAT", 72), ("CAC", 72), ("CAA", 81), ("CAG", 81), - ("AAT", 78), ("AAC", 78), ("AAA", 75), ("AAG", 75), - ("GAT", 68), ("GAC", 68), ("GAA", 69), ("GAG", 69), - ("TGT", 67), ("TGC", 67), ("TGA", 42), ("TGG", 87), - ("CGT", 82), ("CGC", 82), ("CGA", 82), ("CGG", 82), - ("AGT", 83), ("AGC", 83), ("AGA", 82), ("AGG", 82), - ("GGT", 71), ("GGC", 71), ("GGA", 71), ("GGG", 71), - ], capacity=64) + let codon_table : Map[String, UInt16] = Map( + [ + ("TTT", 70), + ("TTC", 70), + ("TTA", 76), + ("TTG", 76), + ("CTT", 76), + ("CTC", 76), + ("CTA", 76), + ("CTG", 76), + ("ATT", 73), + ("ATC", 73), + ("ATA", 73), + ("ATG", 77), + ("GTT", 86), + ("GTC", 86), + ("GTA", 86), + ("GTG", 86), + ("TCT", 83), + ("TCC", 83), + ("TCA", 83), + ("TCG", 83), + ("CCT", 80), + ("CCC", 80), + ("CCA", 80), + ("CCG", 80), + ("ACT", 65), + ("ACC", 65), + ("ACA", 65), + ("ACG", 65), + ("GCT", 65), + ("GCC", 65), + ("GCA", 65), + ("GCG", 65), + ("TAT", 89), + ("TAC", 89), + ("TAA", 42), + ("TAG", 42), + ("CAT", 72), + ("CAC", 72), + ("CAA", 81), + ("CAG", 81), + ("AAT", 78), + ("AAC", 78), + ("AAA", 75), + ("AAG", 75), + ("GAT", 68), + ("GAC", 68), + ("GAA", 69), + ("GAG", 69), + ("TGT", 67), + ("TGC", 67), + ("TGA", 42), + ("TGG", 87), + ("CGT", 82), + ("CGC", 82), + ("CGA", 82), + ("CGG", 82), + ("AGT", 83), + ("AGC", 83), + ("AGA", 82), + ("AGG", 82), + ("GGT", 71), + ("GGC", 71), + ("GGA", 71), + ("GGG", 71), + ], + capacity=64, + ) let mut i = offset let mut out_idx = 0 @@ -1160,7 +1220,11 @@ pub fn expected_matches( // Expected number: (seq_length - pattern_length + 1) * probability let n_positions = (seq_length - n + 1).to_double() - if prob.is_nan() || prob < 0.0 { 0.0 } else { n_positions * prob } + if prob.is_nan() || prob < 0.0 { + 0.0 + } else { + n_positions * prob + } } ///| @@ -1203,7 +1267,8 @@ pub fn sequence_complexity(seq : String, word_size : Int) -> Double { entropy } -///| Test match_pattern with exact matching +///| +/// Test match_pattern with exact matching test "match_pattern_exact" { let result = match_pattern(pattern="ATG", subject="ATGATGATG") assert_eq(result.hits.length(), 3) @@ -1212,29 +1277,38 @@ test "match_pattern_exact" { assert_eq(result.hits[0].mismatches, 0) } -///| Test match_pattern with mismatches +///| +/// Test match_pattern with mismatches test "match_pattern_mismatch" { let result = match_pattern(pattern="ATG", subject="AAGATG", max_mismatches=1) assert_true(result.hits.length() >= 1) } -///| Test match_pattern with indels +///| +/// Test match_pattern with indels test "match_pattern_indels" { - let result = match_pattern(pattern="ATG", subject="ATGC", max_mismatches=1, with_indels=true) + let result = match_pattern( + pattern="ATG", + subject="ATGC", + max_mismatches=1, + with_indels=true, + ) assert_true(result.hits.length() >= 1) } -///| Test vmatch_pattern +///| +/// Test vmatch_pattern test "vmatch_pattern" { let subjects = ["ATGATG", "ATGCAT", "CCCCCC"] - let result = vmatch_pattern(pattern="ATG", subjects=subjects) + let result = vmatch_pattern(pattern="ATG", subjects~) assert_eq(result.length(), 3) assert_eq(result[0].hits.length(), 2) assert_eq(result[1].hits.length(), 1) assert_eq(result[2].hits.length(), 0) } -///| Test find_palindromes +///| +/// Test find_palindromes test "find_palindromes" { let result = find_palindromes(seq="ATAT", min_length=4) assert_eq(result.length(), 1) @@ -1242,76 +1316,93 @@ test "find_palindromes" { assert_eq(result[0].1, 4) } -///| Test find_palindromes_empty +///| +/// Test find_palindromes_empty test "find_palindromes_empty" { let result = find_palindromes(seq="ATGC", min_length=4) assert_eq(result.length(), 0) } -///| Test find_direct_repeats +///| +/// Test find_direct_repeats test "find_direct_repeats" { - let result = find_direct_repeats(seq="ATGATGATG", min_unit_length=3, max_unit_length=3, min_copies=2) + let result = find_direct_repeats( + seq="ATGATGATG", + min_unit_length=3, + max_unit_length=3, + min_copies=2, + ) assert_true(result.length() >= 1) } -///| Test find_inverted_repeats +///| +/// Test find_inverted_repeats test "find_inverted_repeats" { let result = find_inverted_repeats(seq="ATGCAT", min_length=3) assert_true(result.length() >= 0) } -///| Test biostrings_translate with standard genetic code +///| +/// Test biostrings_translate with standard genetic code test "biostrings_translate_basic" { // ATG = M, GCT = A, TAA = * let protein = biostrings_translate("ATGGCTTAA") assert_eq(protein, "MA*") } -///| Test biostrings_translate with different reading frame +///| +/// Test biostrings_translate with different reading frame test "biostrings_translate_frame1" { // Frame 1: TGG = W, CTT = L let protein = biostrings_translate("ATGGCTTAA", frame=1) assert_eq(protein, "WL") } -///| Test biostrings_translate with lowercase input +///| +/// Test biostrings_translate with lowercase input test "biostrings_translate_lowercase" { let protein = biostrings_translate("atggcttaa") assert_eq(protein, "MA*") } -///| Test biostrings_translate empty sequence +///| +/// Test biostrings_translate empty sequence test "biostrings_translate_empty" { let protein = biostrings_translate("") assert_eq(protein, "") } -///| Test biostrings_reverse_complement basic +///| +/// Test biostrings_reverse_complement basic test "biostrings_reverse_complement_basic" { let rc = biostrings_reverse_complement("ATGC") assert_eq(rc, "GCAT") } -///| Test biostrings_reverse_complement with IUPAC +///| +/// Test biostrings_reverse_complement with IUPAC test "biostrings_reverse_complement_iupac" { let rc = biostrings_reverse_complement("AR") assert_eq(rc, "YT") } -///| Test biostrings_reverse_complement with empty string +///| +/// Test biostrings_reverse_complement with empty string test "biostrings_reverse_complement_empty" { let rc = biostrings_reverse_complement("") assert_eq(rc, "") } -///| Test biostrings_reverse_complement palindrome +///| +/// Test biostrings_reverse_complement palindrome test "biostrings_reverse_complement_palindrome" { // ACGT's reverse complement is also ACGT let rc = biostrings_reverse_complement("ACGT") assert_eq(rc, "ACGT") } -///| Test letter_frequency_matrix basic +///| +/// Test letter_frequency_matrix basic test "letter_frequency_matrix_basic" { let matrix = letter_frequency_matrix("ACGT") let a_counts = matrix.get_or_default("A", Array::make(4, 0)) @@ -1324,7 +1415,8 @@ test "letter_frequency_matrix_basic" { assert_eq(t_counts[3], 1) } -///| Test biostrings_consensus_sequence +///| +/// Test biostrings_consensus_sequence test "biostrings_consensus_sequence_basic" { let alignment = ["ACGT", "ACGT", "TCGA"] let consensus = biostrings_consensus_sequence(alignment) @@ -1335,26 +1427,30 @@ test "biostrings_consensus_sequence_basic" { assert_eq(consensus, "ACGT") } -///| Test biostrings_consensus_sequence empty +///| +/// Test biostrings_consensus_sequence empty test "biostrings_consensus_sequence_empty" { let alignment : Array[String] = [] let consensus = biostrings_consensus_sequence(alignment) assert_eq(consensus, "") } -///| Test expected_matches +///| +/// Test expected_matches test "expected_matches_basic" { let n = expected_matches("ACGT", 12, 0.5) assert_true(n > 0.0) } -///| Test sequence_complexity +///| +/// Test sequence_complexity test "sequence_complexity_basic" { let entropy = sequence_complexity("ACGTACGTACGT", 2) assert_true(entropy > 0.0) } -///| Test sequence_complexity with degenerate sequence +///| +/// Test sequence_complexity with degenerate sequence test "sequence_complexity_degenerate" { let entropy = sequence_complexity("AAAA", 2) assert_true(entropy >= 0.0) diff --git a/src/biostrings_matchdict.mbt b/src/biostrings_matchdict.mbt index 828235f6..fa3dfeb2 100644 --- a/src/biostrings_matchdict.mbt +++ b/src/biostrings_matchdict.mbt @@ -60,11 +60,7 @@ pub fn bmd_create_pdict( max_mismatches? : Int = 0, with_indels? : Bool = false, ) -> PDict { - PDict::{ - patterns, - max_mismatches, - with_indels, - } + PDict::{ patterns, max_mismatches, with_indels } } ///| @@ -80,11 +76,7 @@ pub fn bmd_match_pdict(pdict~ : PDict, subject~ : String) -> MatchPDictResult { let hits : Array[MatchPDictHit] = Array::new() if subject.length() == 0 || pdict.patterns.length() == 0 { - return MatchPDictResult::{ - hits, - pdict, - subject_length: subject.length(), - } + return MatchPDictResult::{ hits, pdict, subject_length: subject.length() } } let mut p_idx = 0 @@ -129,11 +121,7 @@ pub fn bmd_match_pdict(pdict~ : PDict, subject~ : String) -> MatchPDictResult { p_idx = p_idx + 1 } - MatchPDictResult::{ - hits, - pdict, - subject_length: subject.length(), - } + MatchPDictResult::{ hits, pdict, subject_length: subject.length() } } ///| @@ -161,11 +149,7 @@ pub fn bmd_vcount_pattern( counts.push(0) i = i + 1 } - return CountPatternResult::{ - pattern, - counts, - total: 0, - } + return CountPatternResult::{ pattern, counts, total: 0 } } let mut i = 0 @@ -203,11 +187,7 @@ pub fn bmd_vcount_pattern( i = i + 1 } - CountPatternResult::{ - pattern, - counts, - total, - } + CountPatternResult::{ pattern, counts, total } } ///| @@ -231,8 +211,8 @@ pub fn bmd_vmatch_pattern( while i < patterns.length() { let result = bmd_vcount_pattern( pattern=patterns[i], - subjects=subjects, - max_mismatches=max_mismatches, + subjects~, + max_mismatches~, ) results.push(result) i = i + 1 @@ -303,7 +283,7 @@ pub fn bmd_which(pdict~ : PDict, subjects~ : Array[String]) -> Array[Bool] { /// @param subjects Array of subject sequences to test. /// @return An array of Int indices indicating matching subjects. pub fn bmd_which_index(pdict~ : PDict, subjects~ : Array[String]) -> Array[Int] { - let which_result = bmd_which(pdict=pdict, subjects=subjects) + let which_result = bmd_which(pdict~, subjects~) let indices : Array[Int] = Array::new() let mut i = 0 @@ -327,7 +307,7 @@ pub fn bmd_which_index(pdict~ : PDict, subjects~ : Array[String]) -> Array[Int] /// @param subject The subject sequence string to search. /// @return Total number of pattern occurrences found. pub fn bmd_count_occurrences(pdict~ : PDict, subject~ : String) -> Int { - let result = bmd_match_pdict(pdict=pdict, subject=subject) + let result = bmd_match_pdict(pdict~, subject~) result.hits.length() } @@ -341,7 +321,7 @@ pub fn bmd_count_occurrences(pdict~ : PDict, subject~ : String) -> Int { /// @param subject The subject sequence string to search. /// @return The best MatchPDictHit, or None if no matches found. pub fn bmd_find_best_match(pdict~ : PDict, subject~ : String) -> MatchPDictHit? { - let result = bmd_match_pdict(pdict=pdict, subject=subject) + let result = bmd_match_pdict(pdict~, subject~) if result.hits.length() == 0 { return None @@ -389,7 +369,11 @@ test "bmd_create_pdict_empty" { ///| test "bmd_create_pdict_with_indels" { - let pdict = bmd_create_pdict(patterns=["ATG"], max_mismatches=1, with_indels=true) + let pdict = bmd_create_pdict( + patterns=["ATG"], + max_mismatches=1, + with_indels=true, + ) assert_eq(pdict.with_indels, true) assert_eq(pdict.max_mismatches, 1) } @@ -397,7 +381,7 @@ test "bmd_create_pdict_with_indels" { ///| test "bmd_match_pdict_single_pattern" { let pdict = bmd_create_pdict(patterns=["ATG"]) - let result = bmd_match_pdict(pdict=pdict, subject="AATGCTAG") + let result = bmd_match_pdict(pdict~, subject="AATGCTAG") assert_eq(result.hits.length(), 1) assert_eq(result.hits[0].pattern, "ATG") assert_eq(result.hits[0].start, 2) @@ -408,7 +392,7 @@ test "bmd_match_pdict_single_pattern" { ///| test "bmd_match_pdict_multiple_patterns" { let pdict = bmd_create_pdict(patterns=["ATG", "GCT"]) - let result = bmd_match_pdict(pdict=pdict, subject="ATGGCT") + let result = bmd_match_pdict(pdict~, subject="ATGGCT") assert_eq(result.hits.length(), 2) assert_eq(result.hits[0].pattern, "ATG") assert_eq(result.hits[1].pattern, "GCT") @@ -417,21 +401,21 @@ test "bmd_match_pdict_multiple_patterns" { ///| test "bmd_match_pdict_overlapping_hits" { let pdict = bmd_create_pdict(patterns=["AAA"]) - let result = bmd_match_pdict(pdict=pdict, subject="AAAA") + let result = bmd_match_pdict(pdict~, subject="AAAA") assert_eq(result.hits.length(), 2) } ///| test "bmd_match_pdict_no_match" { let pdict = bmd_create_pdict(patterns=["XYZ"]) - let result = bmd_match_pdict(pdict=pdict, subject="ATGC") + let result = bmd_match_pdict(pdict~, subject="ATGC") assert_eq(result.hits.length(), 0) } ///| test "bmd_match_pdict_empty_subject" { let pdict = bmd_create_pdict(patterns=["ATG"]) - let result = bmd_match_pdict(pdict=pdict, subject="") + let result = bmd_match_pdict(pdict~, subject="") assert_eq(result.hits.length(), 0) assert_eq(result.subject_length, 0) } @@ -439,14 +423,14 @@ test "bmd_match_pdict_empty_subject" { ///| test "bmd_match_pdict_empty_patterns" { let pdict = bmd_create_pdict(patterns=[]) - let result = bmd_match_pdict(pdict=pdict, subject="ATGC") + let result = bmd_match_pdict(pdict~, subject="ATGC") assert_eq(result.hits.length(), 0) } ///| test "bmd_match_pdict_with_mismatches" { let pdict = bmd_create_pdict(patterns=["ATG"], max_mismatches=1) - let result = bmd_match_pdict(pdict=pdict, subject="AXG") + let result = bmd_match_pdict(pdict~, subject="AXG") assert_eq(result.hits.length(), 1) assert_eq(result.hits[0].mismatches, 1) } @@ -454,21 +438,21 @@ test "bmd_match_pdict_with_mismatches" { ///| test "bmd_match_pdict_mismatch_too_many" { let pdict = bmd_create_pdict(patterns=["ATG"], max_mismatches=1) - let result = bmd_match_pdict(pdict=pdict, subject="XYZ") + let result = bmd_match_pdict(pdict~, subject="XYZ") assert_eq(result.hits.length(), 0) } ///| test "bmd_match_pdict_pattern_longer_than_subject" { let pdict = bmd_create_pdict(patterns=["ATGCT"]) - let result = bmd_match_pdict(pdict=pdict, subject="AT") + let result = bmd_match_pdict(pdict~, subject="AT") assert_eq(result.hits.length(), 0) } ///| test "bmd_vcount_pattern_basic" { let subjects = ["ATGATG", "ATGCAT", "CCCCCC"] - let result = bmd_vcount_pattern(pattern="ATG", subjects=subjects) + let result = bmd_vcount_pattern(pattern="ATG", subjects~) assert_eq(result.pattern, "ATG") assert_eq(result.counts.length(), 3) assert_eq(result.counts[0], 2) @@ -480,7 +464,7 @@ test "bmd_vcount_pattern_basic" { ///| test "bmd_vcount_pattern_no_match" { let subjects = ["CCCC", "GGGG", "TTTT"] - let result = bmd_vcount_pattern(pattern="AAAA", subjects=subjects) + let result = bmd_vcount_pattern(pattern="AAAA", subjects~) assert_eq(result.total, 0) assert_eq(result.counts[0], 0) assert_eq(result.counts[1], 0) @@ -490,7 +474,7 @@ test "bmd_vcount_pattern_no_match" { ///| test "bmd_vcount_pattern_with_mismatches" { let subjects = ["AXG", "AYG", "AZG", "ATG"] - let result = bmd_vcount_pattern(pattern="ATG", subjects=subjects, max_mismatches=1) + let result = bmd_vcount_pattern(pattern="ATG", subjects~, max_mismatches=1) assert_eq(result.counts[0], 1) assert_eq(result.counts[1], 1) assert_eq(result.counts[2], 1) @@ -501,7 +485,7 @@ test "bmd_vcount_pattern_with_mismatches" { ///| test "bmd_vcount_pattern_empty_pattern" { let subjects = ["ATGC", "GCAT"] - let result = bmd_vcount_pattern(pattern="", subjects=subjects) + let result = bmd_vcount_pattern(pattern="", subjects~) assert_eq(result.total, 0) assert_eq(result.counts.length(), 2) } @@ -509,7 +493,7 @@ test "bmd_vcount_pattern_empty_pattern" { ///| test "bmd_vcount_pattern_empty_subjects" { let subjects : Array[String] = [] - let result = bmd_vcount_pattern(pattern="ATG", subjects=subjects) + let result = bmd_vcount_pattern(pattern="ATG", subjects~) assert_eq(result.total, 0) assert_eq(result.counts.length(), 0) } @@ -518,7 +502,7 @@ test "bmd_vcount_pattern_empty_subjects" { test "bmd_vmatch_pattern_basic" { let patterns = ["ATG", "CCC"] let subjects = ["ATGATG", "CCCCCC"] - let results = bmd_vmatch_pattern(patterns=patterns, subjects=subjects) + let results = bmd_vmatch_pattern(patterns~, subjects~) assert_eq(results.length(), 2) assert_eq(results[0].pattern, "ATG") assert_eq(results[0].total, 2) @@ -530,7 +514,7 @@ test "bmd_vmatch_pattern_basic" { test "bmd_vmatch_pattern_multiple_subjects" { let patterns = ["ATG"] let subjects = ["ATGATG", "ATXATG", "TTTTTT"] - let results = bmd_vmatch_pattern(patterns=patterns, subjects=subjects) + let results = bmd_vmatch_pattern(patterns~, subjects~) assert_eq(results.length(), 1) assert_eq(results[0].counts[0], 2) assert_eq(results[0].counts[2], 0) @@ -546,7 +530,7 @@ test "bmd_vmatch_pattern_empty" { test "bmd_which_basic" { let pdict = bmd_create_pdict(patterns=["ATG"]) let subjects = ["ATGCTA", "TTTTTT", "ATGCAT"] - let result = bmd_which(pdict=pdict, subjects=subjects) + let result = bmd_which(pdict~, subjects~) assert_eq(result.length(), 3) assert_eq(result[0], true) assert_eq(result[1], false) @@ -557,7 +541,7 @@ test "bmd_which_basic" { test "bmd_which_no_match" { let pdict = bmd_create_pdict(patterns=["ZZZZZ"]) let subjects = ["ATGC", "GCAT"] - let result = bmd_which(pdict=pdict, subjects=subjects) + let result = bmd_which(pdict~, subjects~) assert_eq(result[0], false) assert_eq(result[1], false) } @@ -566,7 +550,7 @@ test "bmd_which_no_match" { test "bmd_which_empty_pdict" { let pdict = bmd_create_pdict(patterns=[]) let subjects = ["ATGC", "GCAT"] - let result = bmd_which(pdict=pdict, subjects=subjects) + let result = bmd_which(pdict~, subjects~) assert_eq(result[0], false) assert_eq(result[1], false) } @@ -575,7 +559,7 @@ test "bmd_which_empty_pdict" { test "bmd_which_index_basic" { let pdict = bmd_create_pdict(patterns=["ATG"]) let subjects = ["TTTT", "ATGC", "GGGG", "ATGG"] - let indices = bmd_which_index(pdict=pdict, subjects=subjects) + let indices = bmd_which_index(pdict~, subjects~) assert_eq(indices.length(), 2) assert_eq(indices[0], 1) assert_eq(indices[1], 3) @@ -585,42 +569,42 @@ test "bmd_which_index_basic" { test "bmd_which_index_no_match" { let pdict = bmd_create_pdict(patterns=["ZZZZ"]) let subjects = ["ATGC", "GCAT"] - let indices = bmd_which_index(pdict=pdict, subjects=subjects) + let indices = bmd_which_index(pdict~, subjects~) assert_eq(indices.length(), 0) } ///| test "bmd_count_occurrences_basic" { let pdict = bmd_create_pdict(patterns=["ATG"]) - let count = bmd_count_occurrences(pdict=pdict, subject="ATGATGATG") + let count = bmd_count_occurrences(pdict~, subject="ATGATGATG") assert_eq(count, 3) } ///| test "bmd_count_occurrences_multiple_patterns" { let pdict = bmd_create_pdict(patterns=["AT", "TG"]) - let count = bmd_count_occurrences(pdict=pdict, subject="ATGATG") + let count = bmd_count_occurrences(pdict~, subject="ATGATG") assert_eq(count, 4) } ///| test "bmd_count_occurrences_no_match" { let pdict = bmd_create_pdict(patterns=["ZZZZ"]) - let count = bmd_count_occurrences(pdict=pdict, subject="ATGC") + let count = bmd_count_occurrences(pdict~, subject="ATGC") assert_eq(count, 0) } ///| test "bmd_count_occurrences_empty" { let pdict = bmd_create_pdict(patterns=[""]) - let count = bmd_count_occurrences(pdict=pdict, subject="ATGC") + let count = bmd_count_occurrences(pdict~, subject="ATGC") assert_eq(count, 0) } ///| test "bmd_find_best_match_exact" { let pdict = bmd_create_pdict(patterns=["ATG"]) - let best = bmd_find_best_match(pdict=pdict, subject="AATGCT") + let best = bmd_find_best_match(pdict~, subject="AATGCT") assert_true(best is Some(_)) assert_eq(best.unwrap().pattern, "ATG") assert_eq(best.unwrap().start, 2) @@ -629,7 +613,7 @@ test "bmd_find_best_match_exact" { ///| test "bmd_find_best_match_with_mismatches" { let pdict = bmd_create_pdict(patterns=["ATG", "AXG", "AYG"], max_mismatches=1) - let best = bmd_find_best_match(pdict=pdict, subject="AAAG") + let best = bmd_find_best_match(pdict~, subject="AAAG") assert_true(best is Some(_)) assert_eq(best.unwrap().mismatches, 1) assert_eq(best.unwrap().pattern, "ATG") @@ -638,14 +622,14 @@ test "bmd_find_best_match_with_mismatches" { ///| test "bmd_find_best_match_no_hit" { let pdict = bmd_create_pdict(patterns=["ZZZZZ"]) - let best = bmd_find_best_match(pdict=pdict, subject="ATGC") + let best = bmd_find_best_match(pdict~, subject="ATGC") assert_true(best is None) } ///| test "bmd_find_best_match_multiple_hits" { let pdict = bmd_create_pdict(patterns=["ATG"], max_mismatches=2) - let best = bmd_find_best_match(pdict=pdict, subject="AATGCTXG") + let best = bmd_find_best_match(pdict~, subject="AATGCTXG") assert_true(best is Some(_)) assert_eq(best.unwrap().pattern, "ATG") assert_eq(best.unwrap().mismatches, 0) @@ -655,7 +639,7 @@ test "bmd_find_best_match_multiple_hits" { test "bmd_which_with_mismatches" { let pdict = bmd_create_pdict(patterns=["ATG"], max_mismatches=1) let subjects = ["AXGCTA", "TTTTTT", "AYGCAT"] - let result = bmd_which(pdict=pdict, subjects=subjects) + let result = bmd_which(pdict~, subjects~) assert_eq(result[0], true) assert_eq(result[1], false) assert_eq(result[2], true) @@ -664,7 +648,7 @@ test "bmd_which_with_mismatches" { ///| test "bmd_vcount_pattern_single_subject" { let subjects = ["AAAA"] - let result = bmd_vcount_pattern(pattern="AA", subjects=subjects) + let result = bmd_vcount_pattern(pattern="AA", subjects~) assert_eq(result.counts.length(), 1) assert_eq(result.counts[0], 3) assert_eq(result.total, 3) @@ -673,7 +657,7 @@ test "bmd_vcount_pattern_single_subject" { ///| test "bmd_match_pdict_hit_properties" { let pdict = bmd_create_pdict(patterns=["ATG", "TGC"]) - let result = bmd_match_pdict(pdict=pdict, subject="ATGC") + let result = bmd_match_pdict(pdict~, subject="ATGC") assert_eq(result.hits.length(), 2) assert_eq(result.hits[0].pattern_idx, 0) assert_eq(result.hits[0].width, 3) @@ -686,7 +670,7 @@ test "bmd_match_pdict_hit_properties" { ///| test "bmd_count_occurrences_overlapping" { let pdict = bmd_create_pdict(patterns=["AA"]) - let count = bmd_count_occurrences(pdict=pdict, subject="AAAA") + let count = bmd_count_occurrences(pdict~, subject="AAAA") assert_eq(count, 3) } @@ -694,7 +678,7 @@ test "bmd_count_occurrences_overlapping" { test "bmd_vmatch_pattern_multiple_mismatches" { let patterns = ["ATG", "CCC"] let subjects = ["AXGATG", "CXCCCX"] - let results = bmd_vmatch_pattern(patterns=patterns, subjects=subjects, max_mismatches=1) + let results = bmd_vmatch_pattern(patterns~, subjects~, max_mismatches=1) assert_eq(results.length(), 2) assert_eq(results[0].counts[0], 2) assert_eq(results[1].counts[1], 4) diff --git a/src/blast_applications.mbt b/src/blast_applications.mbt index b67b476c..920cfad1 100644 --- a/src/blast_applications.mbt +++ b/src/blast_applications.mbt @@ -23,7 +23,13 @@ pub struct BlastParamSpec { ///| /// Create a parameter specification. -pub fn BlastParamSpec::new(name : String, description : String, takes_value : Bool, default_value : String, required : Bool) -> BlastParamSpec { +pub fn BlastParamSpec::new( + name : String, + description : String, + takes_value : Bool, + default_value : String, + required : Bool, +) -> BlastParamSpec { BlastParamSpec::{ name, description, takes_value, default_value, required } } @@ -32,25 +38,45 @@ pub fn BlastParamSpec::new(name : String, description : String, takes_value : Bo fn blastapp_common_params() -> Array[BlastParamSpec] { [ BlastParamSpec::new("-query", "Query sequence file", true, "", false), - BlastParamSpec::new("-query_loc", "Query location (start-stop)", true, "", false), + BlastParamSpec::new( + "-query_loc", "Query location (start-stop)", true, "", false, + ), BlastParamSpec::new("-db", "BLAST database name", true, "", false), BlastParamSpec::new("-out", "Output file name", true, "", false), - BlastParamSpec::new("-evalue", "Expectation value threshold", true, "10.0", false), - BlastParamSpec::new("-word_size", "Word size for initial match", true, "0", false), + BlastParamSpec::new( + "-evalue", "Expectation value threshold", true, "10.0", false, + ), + BlastParamSpec::new( + "-word_size", "Word size for initial match", true, "0", false, + ), BlastParamSpec::new("-gapopen", "Cost to open a gap", true, "0", false), BlastParamSpec::new("-gapextend", "Cost to extend a gap", true, "0", false), BlastParamSpec::new("-matrix", "Scoring matrix", true, "", false), BlastParamSpec::new("-threshold", "Minimum word score", true, "0", false), - BlastParamSpec::new("-comp_based_stats", "Composition-based stats", true, "0", false), - BlastParamSpec::new("-num_descriptions", "Number of descriptions", true, "500", false), - BlastParamSpec::new("-num_alignments", "Number of alignments", true, "250", false), - BlastParamSpec::new("-num_threads", "Number of CPU threads", true, "1", false), - BlastParamSpec::new("-max_target_seqs", "Max target sequences", true, "500", false), + BlastParamSpec::new( + "-comp_based_stats", "Composition-based stats", true, "0", false, + ), + BlastParamSpec::new( + "-num_descriptions", "Number of descriptions", true, "500", false, + ), + BlastParamSpec::new( + "-num_alignments", "Number of alignments", true, "250", false, + ), + BlastParamSpec::new( + "-num_threads", "Number of CPU threads", true, "1", false, + ), + BlastParamSpec::new( + "-max_target_seqs", "Max target sequences", true, "500", false, + ), BlastParamSpec::new("-dust", "DUST filter setting", true, "", false), BlastParamSpec::new("-seg", "SEG filter setting", true, "", false), BlastParamSpec::new("-soft_masking", "Soft masking", true, "false", false), - BlastParamSpec::new("-lcase_masking", "Use lowercase masking", false, "", false), - BlastParamSpec::new("-show_gis", "Show NCBI GIs in output", false, "", false), + BlastParamSpec::new( + "-lcase_masking", "Use lowercase masking", false, "", false, + ), + BlastParamSpec::new( + "-show_gis", "Show NCBI GIs in output", false, "", false, + ), BlastParamSpec::new("-html", "Produce HTML output", false, "", false), ] } @@ -61,7 +87,9 @@ fn blastapp_output_params() -> Array[BlastParamSpec] { [ BlastParamSpec::new("-outfmt", "Output format (0-18)", true, "0", false), BlastParamSpec::new("-max_hsps", "Max HSPs per subject", true, "0", false), - BlastParamSpec::new("-max_intron_length", "Max intron length", true, "0", false), + BlastParamSpec::new( + "-max_intron_length", "Max intron length", true, "0", false, + ), ] } @@ -82,18 +110,25 @@ pub struct BlastCommandline { ///| /// Create a new BLAST commandline wrapper. -pub fn BlastCommandline::new(executable : String, param_specs : Array[BlastParamSpec]) -> BlastCommandline { +pub fn BlastCommandline::new( + executable : String, + param_specs : Array[BlastParamSpec], +) -> BlastCommandline { BlastCommandline::{ executable, parameters: Map([], capacity=32), flags: Map([], capacity=16), - param_specs + param_specs, } } ///| /// Set a parameter value. -pub fn BlastCommandline::blastapp_set_parameter(self : BlastCommandline, name : String, value : String) -> BlastCommandline { +pub fn BlastCommandline::blastapp_set_parameter( + self : BlastCommandline, + name : String, + value : String, +) -> BlastCommandline { let new_params = Map([], capacity=32) let keys = self.parameters.keys().collect() let mut i = 0 @@ -106,13 +141,17 @@ pub fn BlastCommandline::blastapp_set_parameter(self : BlastCommandline, name : executable: self.executable, parameters: new_params, flags: self.flags, - param_specs: self.param_specs + param_specs: self.param_specs, } } ///| /// Set a flag (boolean parameter). -pub fn BlastCommandline::blastapp_set_flag(self : BlastCommandline, name : String, on : Bool) -> BlastCommandline { +pub fn BlastCommandline::blastapp_set_flag( + self : BlastCommandline, + name : String, + on : Bool, +) -> BlastCommandline { let new_flags = Map([], capacity=16) let keys = self.flags.keys().collect() let mut i = 0 @@ -125,25 +164,33 @@ pub fn BlastCommandline::blastapp_set_flag(self : BlastCommandline, name : Strin executable: self.executable, parameters: self.parameters, flags: new_flags, - param_specs: self.param_specs + param_specs: self.param_specs, } } ///| /// Get a parameter value (returns Option). -pub fn BlastCommandline::blastapp_get_parameter(self : BlastCommandline, name : String) -> String? { +pub fn BlastCommandline::blastapp_get_parameter( + self : BlastCommandline, + name : String, +) -> String? { self.parameters.get(name) } ///| /// Get a flag value (returns Option). -pub fn BlastCommandline::blastapp_get_flag(self : BlastCommandline, name : String) -> Bool? { +pub fn BlastCommandline::blastapp_get_flag( + self : BlastCommandline, + name : String, +) -> Bool? { self.flags.get(name) } ///| /// Build the full command-line string. -pub fn BlastCommandline::blastapp_build_command(self : BlastCommandline) -> String { +pub fn BlastCommandline::blastapp_build_command( + self : BlastCommandline, +) -> String { let mut cmd = self.executable // Add value parameters let keys = self.parameters.keys().collect() @@ -166,7 +213,9 @@ pub fn BlastCommandline::blastapp_build_command(self : BlastCommandline) -> Stri ///| /// Validate that all required parameters are set. -pub fn BlastCommandline::blastapp_validate(self : BlastCommandline) -> Array[String] { +pub fn BlastCommandline::blastapp_validate( + self : BlastCommandline, +) -> Array[String] { let errors : Array[String] = Array::new() let mut i = 0 while i < self.param_specs.length() { @@ -184,7 +233,9 @@ pub fn BlastCommandline::blastapp_validate(self : BlastCommandline) -> Array[Str ///| /// List all available parameter specifications. -pub fn BlastCommandline::blastapp_available_params(self : BlastCommandline) -> Array[BlastParamSpec] { +pub fn BlastCommandline::blastapp_available_params( + self : BlastCommandline, +) -> Array[BlastParamSpec] { self.param_specs.copy() } @@ -201,9 +252,15 @@ pub fn ncbi_blastp_commandline() -> BlastCommandline { i = i + 1 } // blastp-specific params - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) - specs.push(BlastParamSpec::new("-subject", "Subject sequence file", true, "", false)) - specs.push(BlastParamSpec::new("-subject_loc", "Subject location", true, "", false)) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) + specs.push( + BlastParamSpec::new("-subject", "Subject sequence file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-subject_loc", "Subject location", true, "", false), + ) BlastCommandline::new("blastp", specs) } @@ -218,11 +275,26 @@ pub fn ncbi_blastn_commandline() -> BlastCommandline { i = i + 1 } // blastn-specific params - specs.push(BlastParamSpec::new("-strand", "Query strand (both/minus/plus)", true, "both", false)) - specs.push(BlastParamSpec::new("-task", "Task (blastn/blastn-short/dc-megablast/etc.)", true, "megablast", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) - specs.push(BlastParamSpec::new("-subject", "Subject sequence file", true, "", false)) - specs.push(BlastParamSpec::new("-subject_loc", "Subject location", true, "", false)) + specs.push( + BlastParamSpec::new( + "-strand", "Query strand (both/minus/plus)", true, "both", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-task", "Task (blastn/blastn-short/dc-megablast/etc.)", true, "megablast", + false, + ), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) + specs.push( + BlastParamSpec::new("-subject", "Subject sequence file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-subject_loc", "Subject location", true, "", false), + ) BlastCommandline::new("blastn", specs) } @@ -236,12 +308,30 @@ pub fn ncbi_blastx_commandline() -> BlastCommandline { specs.push(out_specs[i]) i = i + 1 } - specs.push(BlastParamSpec::new("-strand", "Query strand (both/minus/plus)", true, "both", false)) - specs.push(BlastParamSpec::new("-query_gencode", "Query genetic code", true, "1", false)) - specs.push(BlastParamSpec::new("-max_intron_length", "Max intron length", true, "0", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) - specs.push(BlastParamSpec::new("-subject", "Subject sequence file", true, "", false)) - specs.push(BlastParamSpec::new("-subject_loc", "Subject location", true, "", false)) + specs.push( + BlastParamSpec::new( + "-strand", "Query strand (both/minus/plus)", true, "both", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-query_gencode", "Query genetic code", true, "1", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-max_intron_length", "Max intron length", true, "0", false, + ), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) + specs.push( + BlastParamSpec::new("-subject", "Subject sequence file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-subject_loc", "Subject location", true, "", false), + ) BlastCommandline::new("blastx", specs) } @@ -255,11 +345,25 @@ pub fn ncbi_tblastn_commandline() -> BlastCommandline { specs.push(out_specs[i]) i = i + 1 } - specs.push(BlastParamSpec::new("-db_gencode", "Database genetic code", true, "1", false)) - specs.push(BlastParamSpec::new("-frame_shift_penalty", "Frame shift penalty", true, "0", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) - specs.push(BlastParamSpec::new("-subject", "Subject sequence file", true, "", false)) - specs.push(BlastParamSpec::new("-subject_loc", "Subject location", true, "", false)) + specs.push( + BlastParamSpec::new( + "-db_gencode", "Database genetic code", true, "1", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-frame_shift_penalty", "Frame shift penalty", true, "0", false, + ), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) + specs.push( + BlastParamSpec::new("-subject", "Subject sequence file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-subject_loc", "Subject location", true, "", false), + ) BlastCommandline::new("tblastn", specs) } @@ -273,12 +377,30 @@ pub fn ncbi_tblastx_commandline() -> BlastCommandline { specs.push(out_specs[i]) i = i + 1 } - specs.push(BlastParamSpec::new("-strand", "Query strand (both/minus/plus)", true, "both", false)) - specs.push(BlastParamSpec::new("-query_gencode", "Query genetic code", true, "1", false)) - specs.push(BlastParamSpec::new("-db_gencode", "Database genetic code", true, "1", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) - specs.push(BlastParamSpec::new("-subject", "Subject sequence file", true, "", false)) - specs.push(BlastParamSpec::new("-subject_loc", "Subject location", true, "", false)) + specs.push( + BlastParamSpec::new( + "-strand", "Query strand (both/minus/plus)", true, "both", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-query_gencode", "Query genetic code", true, "1", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-db_gencode", "Database genetic code", true, "1", false, + ), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) + specs.push( + BlastParamSpec::new("-subject", "Subject sequence file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-subject_loc", "Subject location", true, "", false), + ) BlastCommandline::new("tblastx", specs) } @@ -292,15 +414,44 @@ pub fn ncbi_psiblast_commandline() -> BlastCommandline { specs.push(out_specs[i]) i = i + 1 } - specs.push(BlastParamSpec::new("-num_iterations", "Number of iterations", true, "1", false)) - specs.push(BlastParamSpec::new("-in_pssm", "Input PSSM file", true, "", false)) - specs.push(BlastParamSpec::new("-out_pssm", "Output PSSM file", true, "", false)) - specs.push(BlastParamSpec::new("-out_ascii_pssm", "Output ASCII PSSM file", true, "", false)) - specs.push(BlastParamSpec::new("-save_pssm_after_last_round", "Save PSSM after last iteration", false, "", false)) - specs.push(BlastParamSpec::new("-save_each_pssm", "Save PSSM after each iteration", false, "", false)) - specs.push(BlastParamSpec::new("-pseudocount", "Pseudocount", true, "0", false)) - specs.push(BlastParamSpec::new("-inclusion_ethresh", "Inclusion e-value threshold", true, "0.002", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) + specs.push( + BlastParamSpec::new( + "-num_iterations", "Number of iterations", true, "1", false, + ), + ) + specs.push( + BlastParamSpec::new("-in_pssm", "Input PSSM file", true, "", false), + ) + specs.push( + BlastParamSpec::new("-out_pssm", "Output PSSM file", true, "", false), + ) + specs.push( + BlastParamSpec::new( + "-out_ascii_pssm", "Output ASCII PSSM file", true, "", false, + ), + ) + specs.push( + BlastParamSpec::new( + "-save_pssm_after_last_round", "Save PSSM after last iteration", false, "", + false, + ), + ) + specs.push( + BlastParamSpec::new( + "-save_each_pssm", "Save PSSM after each iteration", false, "", false, + ), + ) + specs.push( + BlastParamSpec::new("-pseudocount", "Pseudocount", true, "0", false), + ) + specs.push( + BlastParamSpec::new( + "-inclusion_ethresh", "Inclusion e-value threshold", true, "0.002", false, + ), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) BlastCommandline::new("psiblast", specs) } @@ -308,9 +459,15 @@ pub fn ncbi_psiblast_commandline() -> BlastCommandline { /// NCBI rpsblast command line (reverse position-specific BLAST). pub fn ncbi_rpsblast_commandline() -> BlastCommandline { let specs = blastapp_common_params() - specs.push(BlastParamSpec::new("-outfmt", "Output format (0-18)", true, "0", false)) - specs.push(BlastParamSpec::new("-max_hsps", "Max HSPs per subject", true, "0", false)) - specs.push(BlastParamSpec::new("-remote", "Execute search remotely", false, "", false)) + specs.push( + BlastParamSpec::new("-outfmt", "Output format (0-18)", true, "0", false), + ) + specs.push( + BlastParamSpec::new("-max_hsps", "Max HSPs per subject", true, "0", false), + ) + specs.push( + BlastParamSpec::new("-remote", "Execute search remotely", false, "", false), + ) BlastCommandline::new("rpsblast", specs) } @@ -319,21 +476,34 @@ pub fn ncbi_rpsblast_commandline() -> BlastCommandline { pub fn ncbi_makeblastdb_commandline() -> BlastCommandline { let specs : Array[BlastParamSpec] = [ BlastParamSpec::new("-in", "Input FASTA file", true, "", false), - BlastParamSpec::new("-input_type", "Input file type (fasta/asn1_bin/asn1_txt)", true, "fasta", false), - BlastParamSpec::new("-dbtype", "Database type (nucl/prot)", true, "nucl", false), + BlastParamSpec::new( + "-input_type", "Input file type (fasta/asn1_bin/asn1_txt)", true, "fasta", + false, + ), + BlastParamSpec::new( + "-dbtype", "Database type (nucl/prot)", true, "nucl", false, + ), BlastParamSpec::new("-title", "Database title", true, "", false), BlastParamSpec::new("-parse_seqids", "Parse sequence IDs", false, "", false), BlastParamSpec::new("-hash_index", "Create hash index", false, "", false), - BlastParamSpec::new("-mask_data", "Masking algorithm data file", true, "", false), + BlastParamSpec::new( + "-mask_data", "Masking algorithm data file", true, "", false, + ), BlastParamSpec::new("-mask_algo", "Masking algorithm ID", true, "", false), BlastParamSpec::new("-gilist", "GI list file", true, "", false), BlastParamSpec::new("-seqidlist", "Sequence ID list file", true, "", false), - BlastParamSpec::new("-negative_gilist", "Negative GI list file", true, "", false), + BlastParamSpec::new( + "-negative_gilist", "Negative GI list file", true, "", false, + ), BlastParamSpec::new("-taxid", "Taxonomy ID", true, "", false), BlastParamSpec::new("-taxid_map", "Taxonomy ID map file", true, "", false), BlastParamSpec::new("-out", "Output database name", true, "", false), - BlastParamSpec::new("-blastdb_version", "BLAST database version (4 or 5)", true, "5", false), - BlastParamSpec::new("-max_file_sz", "Maximum file size", true, "1000000000", false), + BlastParamSpec::new( + "-blastdb_version", "BLAST database version (4 or 5)", true, "5", false, + ), + BlastParamSpec::new( + "-max_file_sz", "Maximum file size", true, "1000000000", false, + ), ] BlastCommandline::new("makeblastdb", specs) } @@ -342,47 +512,67 @@ pub fn ncbi_makeblastdb_commandline() -> BlastCommandline { ///| /// Build a typical blastp query against a local database. -pub fn blastapp_quick_blastp(query_file : String, db_name : String, evalue : String) -> BlastCommandline { +pub fn blastapp_quick_blastp( + query_file : String, + db_name : String, + evalue : String, +) -> BlastCommandline { ncbi_blastp_commandline() - .blastapp_set_parameter("-query", query_file) - .blastapp_set_parameter("-db", db_name) - .blastapp_set_parameter("-evalue", evalue) + .blastapp_set_parameter("-query", query_file) + .blastapp_set_parameter("-db", db_name) + .blastapp_set_parameter("-evalue", evalue) } ///| /// Build a typical blastn query against a local database. -pub fn blastapp_quick_blastn(query_file : String, db_name : String, evalue : String) -> BlastCommandline { +pub fn blastapp_quick_blastn( + query_file : String, + db_name : String, + evalue : String, +) -> BlastCommandline { ncbi_blastn_commandline() - .blastapp_set_parameter("-query", query_file) - .blastapp_set_parameter("-db", db_name) - .blastapp_set_parameter("-evalue", evalue) + .blastapp_set_parameter("-query", query_file) + .blastapp_set_parameter("-db", db_name) + .blastapp_set_parameter("-evalue", evalue) } ///| /// Build a typical blastx query against a local protein database. -pub fn blastapp_quick_blastx(query_file : String, db_name : String, evalue : String) -> BlastCommandline { +pub fn blastapp_quick_blastx( + query_file : String, + db_name : String, + evalue : String, +) -> BlastCommandline { ncbi_blastx_commandline() - .blastapp_set_parameter("-query", query_file) - .blastapp_set_parameter("-db", db_name) - .blastapp_set_parameter("-evalue", evalue) + .blastapp_set_parameter("-query", query_file) + .blastapp_set_parameter("-db", db_name) + .blastapp_set_parameter("-evalue", evalue) } ///| /// Build a typical tblastn query against a local nucleotide database. -pub fn blastapp_quick_tblastn(query_file : String, db_name : String, evalue : String) -> BlastCommandline { +pub fn blastapp_quick_tblastn( + query_file : String, + db_name : String, + evalue : String, +) -> BlastCommandline { ncbi_tblastn_commandline() - .blastapp_set_parameter("-query", query_file) - .blastapp_set_parameter("-db", db_name) - .blastapp_set_parameter("-evalue", evalue) + .blastapp_set_parameter("-query", query_file) + .blastapp_set_parameter("-db", db_name) + .blastapp_set_parameter("-evalue", evalue) } ///| /// Build a makeblastdb command for a given input file. -pub fn blastapp_quick_makeblastdb(input_file : String, dbtype : String, db_title : String) -> BlastCommandline { +pub fn blastapp_quick_makeblastdb( + input_file : String, + dbtype : String, + db_title : String, +) -> BlastCommandline { ncbi_makeblastdb_commandline() - .blastapp_set_parameter("-in", input_file) - .blastapp_set_parameter("-dbtype", dbtype) - .blastapp_set_parameter("-title", db_title) + .blastapp_set_parameter("-in", input_file) + .blastapp_set_parameter("-dbtype", dbtype) + .blastapp_set_parameter("-title", db_title) } // ===== Example data generators ===== @@ -391,37 +581,37 @@ pub fn blastapp_quick_makeblastdb(input_file : String, dbtype : String, db_title /// Create a sample blastp commandline for demonstration. pub fn blastapp_create_example_blastp() -> BlastCommandline { ncbi_blastp_commandline() - .blastapp_set_parameter("-query", "query.fasta") - .blastapp_set_parameter("-db", "nr") - .blastapp_set_parameter("-evalue", "0.001") - .blastapp_set_parameter("-out", "results.txt") - .blastapp_set_parameter("-outfmt", "7") - .blastapp_set_parameter("-num_threads", "4") - .blastapp_set_parameter("-max_target_seqs", "10") - .blastapp_set_flag("-show_gis", true) + .blastapp_set_parameter("-query", "query.fasta") + .blastapp_set_parameter("-db", "nr") + .blastapp_set_parameter("-evalue", "0.001") + .blastapp_set_parameter("-out", "results.txt") + .blastapp_set_parameter("-outfmt", "7") + .blastapp_set_parameter("-num_threads", "4") + .blastapp_set_parameter("-max_target_seqs", "10") + .blastapp_set_flag("-show_gis", true) } ///| /// Create a sample blastn commandline for demonstration. pub fn blastapp_create_example_blastn() -> BlastCommandline { ncbi_blastn_commandline() - .blastapp_set_parameter("-query", "gene.fasta") - .blastapp_set_parameter("-db", "nt") - .blastapp_set_parameter("-evalue", "1e-50") - .blastapp_set_parameter("-task", "megablast") - .blastapp_set_parameter("-out", "blastn_results.xml") - .blastapp_set_parameter("-outfmt", "5") - .blastapp_set_parameter("-num_descriptions", "100") - .blastapp_set_flag("-html", false) + .blastapp_set_parameter("-query", "gene.fasta") + .blastapp_set_parameter("-db", "nt") + .blastapp_set_parameter("-evalue", "1e-50") + .blastapp_set_parameter("-task", "megablast") + .blastapp_set_parameter("-out", "blastn_results.xml") + .blastapp_set_parameter("-outfmt", "5") + .blastapp_set_parameter("-num_descriptions", "100") + .blastapp_set_flag("-html", false) } ///| /// Create a sample makeblastdb commandline for demonstration. pub fn blastapp_create_example_makeblastdb() -> BlastCommandline { ncbi_makeblastdb_commandline() - .blastapp_set_parameter("-in", "genome.fa") - .blastapp_set_parameter("-dbtype", "nucl") - .blastapp_set_parameter("-title", "ExampleGenome") - .blastapp_set_flag("-parse_seqids", true) - .blastapp_set_flag("-hash_index", true) + .blastapp_set_parameter("-in", "genome.fa") + .blastapp_set_parameter("-dbtype", "nucl") + .blastapp_set_parameter("-title", "ExampleGenome") + .blastapp_set_flag("-parse_seqids", true) + .blastapp_set_flag("-hash_index", true) } diff --git a/src/blast_xml_advanced.mbt b/src/blast_xml_advanced.mbt new file mode 100644 index 00000000..0caacf52 --- /dev/null +++ b/src/blast_xml_advanced.mbt @@ -0,0 +1,3204 @@ +// Biopython-compatible modern Bio.Blast XML support. +// +// This module parses and writes NCBI BLAST XML1 and XML2 artifacts. It does +// not submit remote BLAST jobs or launch local BLAST executables. + +///| +/// Error raised for malformed or inconsistent BLAST XML. +pub suberror BlastXmlError { + BlastXmlError(String) +} + +///| +/// NCBI BLAST XML dialect. +pub(all) enum BlastXmlFormat { + Xml1 + Xml2 +} derive(Eq, Debug) + +///| +/// One masked interval on the query, using zero-based half-open coordinates. +pub struct BlastXmlMask { + start : Int + end : Int +} derive(Eq, Debug) + +///| +/// Query metadata for one BLAST record. +pub struct BlastXmlQuery { + id : String + description : String + length : Int + sequence : String? + masks : Array[BlastXmlMask] +} derive(Eq, Debug) + +///| +/// Search parameters stored in a BLAST XML header. +pub struct BlastXmlParameters { + matrix : String? + expect : Double + inclusion_threshold : Double? + score_match : Int? + score_mismatch : Int? + gap_open : Int? + gap_extend : Int? + filter : String? + pattern : String? + entrez_query : String? + composition_based_statistics : Int? + query_genetic_code : Int? + database_genetic_code : Int? + bl2seq_mode : Int? +} derive(Eq, Debug) + +///| +/// Database and Karlin-Altschul statistics for one query. +/// +/// `database_letters` is a `Double` because BLAST databases can exceed the +/// platform `Int` range. Integer values below 2^53 remain exactly represented. +pub struct BlastXmlStatistics { + database_sequences : Int + database_letters : Double + effective_hsp_length : Int + effective_search_space : Double + kappa : Double + lambda : Double + entropy : Double +} derive(Eq, Debug) + +///| +/// One target description. XML2 can attach several descriptions to one hit. +pub struct BlastXmlDescription { + id : String + accession : String + title : String + taxid : Int? + scientific_name : String? +} derive(Eq, Debug) + +///| +/// One breakpoint in an HSP coordinate path. +/// +/// Coordinates are zero-based boundaries. Reverse-strand axes decrease. +/// Translated nucleotide axes move in steps of three. +pub struct BlastXmlCoordinate { + target : Int + query : Int +} derive(Eq, Debug) + +///| +/// One high-scoring segment pair. +pub struct BlastXmlHsp { + number : Int + bit_score : Double + score : Double + evalue : Double + query_from : Int + query_to : Int + target_from : Int + target_to : Int + query_frame : Int? + target_frame : Int? + query_strand : String? + target_strand : String? + identity : Int + positive : Int? + gaps : Int? + alignment_length : Int? + density : Int? + pattern_from : Int? + pattern_to : Int? + query_sequence : String + target_sequence : String + midline : String? + coordinates : Array[BlastXmlCoordinate] + query_coded_by : String? + target_coded_by : String? +} derive(Eq, Debug) + +///| +/// One BLAST database hit and its HSPs. +pub struct BlastXmlHit { + number : Int + descriptions : Array[BlastXmlDescription] + length : Int + hsps : Array[BlastXmlHsp] +} derive(Eq, Debug) + +///| +/// Results for one query. +pub struct BlastXmlRecord { + number : Int + query : BlastXmlQuery + hits : Array[BlastXmlHit] + statistics : BlastXmlStatistics? + message : String? +} derive(Eq, Debug) + +///| +/// A complete BLAST XML1 or XML2 document. +pub struct BlastXmlDocument { + format : BlastXmlFormat + program : String + version : String + reference : String + database : String + global_query : BlastXmlQuery? + parameters : BlastXmlParameters + records : Array[BlastXmlRecord] + megablast_statistics : BlastXmlStatistics? +} derive(Eq, Debug) + +///| +priv struct BlastXmlNode { + name : String + text : String + children : Array[BlastXmlNode] +} + +///| +fn blast_xml_fail(message : String) -> Unit raise BlastXmlError { + raise BlastXmlError(message) +} + +///| +fn blast_xml_copy_masks(values : Array[BlastXmlMask]) -> Array[BlastXmlMask] { + let copied : Array[BlastXmlMask] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_descriptions( + values : Array[BlastXmlDescription], +) -> Array[BlastXmlDescription] { + let copied : Array[BlastXmlDescription] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_coordinates( + values : Array[BlastXmlCoordinate], +) -> Array[BlastXmlCoordinate] { + let copied : Array[BlastXmlCoordinate] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_hsps(values : Array[BlastXmlHsp]) -> Array[BlastXmlHsp] { + let copied : Array[BlastXmlHsp] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_hits(values : Array[BlastXmlHit]) -> Array[BlastXmlHit] { + let copied : Array[BlastXmlHit] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_records( + values : Array[BlastXmlRecord], +) -> Array[BlastXmlRecord] { + let copied : Array[BlastXmlRecord] = [] + for value in values { + copied.push(value) + } + copied +} + +///| +fn blast_xml_copy_query(query : BlastXmlQuery) -> BlastXmlQuery { + BlastXmlQuery::{ + id: query.id, + description: query.description, + length: query.length, + sequence: query.sequence, + masks: blast_xml_copy_masks(query.masks), + } +} + +///| +pub fn BlastXmlMask::start(self : BlastXmlMask) -> Int { + self.start +} + +///| +pub fn BlastXmlMask::end(self : BlastXmlMask) -> Int { + self.end +} + +///| +pub fn BlastXmlQuery::id(self : BlastXmlQuery) -> String { + self.id +} + +///| +pub fn BlastXmlQuery::description(self : BlastXmlQuery) -> String { + self.description +} + +///| +pub fn BlastXmlQuery::length(self : BlastXmlQuery) -> Int { + self.length +} + +///| +pub fn BlastXmlQuery::sequence(self : BlastXmlQuery) -> String? { + self.sequence +} + +///| +pub fn BlastXmlQuery::masks(self : BlastXmlQuery) -> Array[BlastXmlMask] { + blast_xml_copy_masks(self.masks) +} + +///| +pub fn BlastXmlParameters::matrix(self : BlastXmlParameters) -> String? { + self.matrix +} + +///| +pub fn BlastXmlParameters::expect(self : BlastXmlParameters) -> Double { + self.expect +} + +///| +pub fn BlastXmlParameters::inclusion_threshold( + self : BlastXmlParameters, +) -> Double? { + self.inclusion_threshold +} + +///| +pub fn BlastXmlParameters::score_match(self : BlastXmlParameters) -> Int? { + self.score_match +} + +///| +pub fn BlastXmlParameters::score_mismatch(self : BlastXmlParameters) -> Int? { + self.score_mismatch +} + +///| +pub fn BlastXmlParameters::gap_open(self : BlastXmlParameters) -> Int? { + self.gap_open +} + +///| +pub fn BlastXmlParameters::gap_extend(self : BlastXmlParameters) -> Int? { + self.gap_extend +} + +///| +pub fn BlastXmlParameters::filter(self : BlastXmlParameters) -> String? { + self.filter +} + +///| +pub fn BlastXmlParameters::pattern(self : BlastXmlParameters) -> String? { + self.pattern +} + +///| +pub fn BlastXmlParameters::entrez_query(self : BlastXmlParameters) -> String? { + self.entrez_query +} + +///| +pub fn BlastXmlParameters::composition_based_statistics( + self : BlastXmlParameters, +) -> Int? { + self.composition_based_statistics +} + +///| +pub fn BlastXmlParameters::query_genetic_code( + self : BlastXmlParameters, +) -> Int? { + self.query_genetic_code +} + +///| +pub fn BlastXmlParameters::database_genetic_code( + self : BlastXmlParameters, +) -> Int? { + self.database_genetic_code +} + +///| +pub fn BlastXmlParameters::bl2seq_mode(self : BlastXmlParameters) -> Int? { + self.bl2seq_mode +} + +///| +pub fn BlastXmlStatistics::database_sequences(self : BlastXmlStatistics) -> Int { + self.database_sequences +} + +///| +pub fn BlastXmlStatistics::database_letters( + self : BlastXmlStatistics, +) -> Double { + self.database_letters +} + +///| +pub fn BlastXmlStatistics::effective_hsp_length( + self : BlastXmlStatistics, +) -> Int { + self.effective_hsp_length +} + +///| +pub fn BlastXmlStatistics::effective_search_space( + self : BlastXmlStatistics, +) -> Double { + self.effective_search_space +} + +///| +pub fn BlastXmlStatistics::kappa(self : BlastXmlStatistics) -> Double { + self.kappa +} + +///| +pub fn BlastXmlStatistics::lambda(self : BlastXmlStatistics) -> Double { + self.lambda +} + +///| +pub fn BlastXmlStatistics::entropy(self : BlastXmlStatistics) -> Double { + self.entropy +} + +///| +pub fn BlastXmlDescription::id(self : BlastXmlDescription) -> String { + self.id +} + +///| +pub fn BlastXmlDescription::accession(self : BlastXmlDescription) -> String { + self.accession +} + +///| +pub fn BlastXmlDescription::title(self : BlastXmlDescription) -> String { + self.title +} + +///| +pub fn BlastXmlDescription::taxid(self : BlastXmlDescription) -> Int? { + self.taxid +} + +///| +pub fn BlastXmlDescription::scientific_name( + self : BlastXmlDescription, +) -> String? { + self.scientific_name +} + +///| +pub fn BlastXmlCoordinate::target(self : BlastXmlCoordinate) -> Int { + self.target +} + +///| +pub fn BlastXmlCoordinate::query(self : BlastXmlCoordinate) -> Int { + self.query +} + +///| +pub fn BlastXmlHsp::number(self : BlastXmlHsp) -> Int { + self.number +} + +///| +pub fn BlastXmlHsp::bit_score(self : BlastXmlHsp) -> Double { + self.bit_score +} + +///| +pub fn BlastXmlHsp::score(self : BlastXmlHsp) -> Double { + self.score +} + +///| +pub fn BlastXmlHsp::evalue(self : BlastXmlHsp) -> Double { + self.evalue +} + +///| +pub fn BlastXmlHsp::query_from(self : BlastXmlHsp) -> Int { + self.query_from +} + +///| +pub fn BlastXmlHsp::query_to(self : BlastXmlHsp) -> Int { + self.query_to +} + +///| +pub fn BlastXmlHsp::target_from(self : BlastXmlHsp) -> Int { + self.target_from +} + +///| +pub fn BlastXmlHsp::target_to(self : BlastXmlHsp) -> Int { + self.target_to +} + +///| +pub fn BlastXmlHsp::query_frame(self : BlastXmlHsp) -> Int? { + self.query_frame +} + +///| +pub fn BlastXmlHsp::target_frame(self : BlastXmlHsp) -> Int? { + self.target_frame +} + +///| +pub fn BlastXmlHsp::query_strand(self : BlastXmlHsp) -> String? { + self.query_strand +} + +///| +pub fn BlastXmlHsp::target_strand(self : BlastXmlHsp) -> String? { + self.target_strand +} + +///| +pub fn BlastXmlHsp::identity(self : BlastXmlHsp) -> Int { + self.identity +} + +///| +pub fn BlastXmlHsp::positive(self : BlastXmlHsp) -> Int? { + self.positive +} + +///| +pub fn BlastXmlHsp::gaps(self : BlastXmlHsp) -> Int? { + self.gaps +} + +///| +pub fn BlastXmlHsp::alignment_length(self : BlastXmlHsp) -> Int? { + self.alignment_length +} + +///| +pub fn BlastXmlHsp::density(self : BlastXmlHsp) -> Int? { + self.density +} + +///| +pub fn BlastXmlHsp::query_sequence(self : BlastXmlHsp) -> String { + self.query_sequence +} + +///| +pub fn BlastXmlHsp::target_sequence(self : BlastXmlHsp) -> String { + self.target_sequence +} + +///| +pub fn BlastXmlHsp::midline(self : BlastXmlHsp) -> String? { + self.midline +} + +///| +pub fn BlastXmlHsp::coordinates( + self : BlastXmlHsp, +) -> Array[BlastXmlCoordinate] { + blast_xml_copy_coordinates(self.coordinates) +} + +///| +pub fn BlastXmlHsp::query_coded_by(self : BlastXmlHsp) -> String? { + self.query_coded_by +} + +///| +pub fn BlastXmlHsp::target_coded_by(self : BlastXmlHsp) -> String? { + self.target_coded_by +} + +///| +pub fn BlastXmlHsp::identity_fraction(self : BlastXmlHsp) -> Double { + match self.alignment_length { + Some(length) => + if length == 0 { + 0.0 + } else { + self.identity.to_double() / length.to_double() + } + None => 0.0 + } +} + +///| +pub fn BlastXmlHit::number(self : BlastXmlHit) -> Int { + self.number +} + +///| +pub fn BlastXmlHit::descriptions( + self : BlastXmlHit, +) -> Array[BlastXmlDescription] { + blast_xml_copy_descriptions(self.descriptions) +} + +///| +pub fn BlastXmlHit::primary_description( + self : BlastXmlHit, +) -> BlastXmlDescription { + self.descriptions[0] +} + +///| +pub fn BlastXmlHit::length(self : BlastXmlHit) -> Int { + self.length +} + +///| +pub fn BlastXmlHit::hsps(self : BlastXmlHit) -> Array[BlastXmlHsp] { + blast_xml_copy_hsps(self.hsps) +} + +///| +pub fn BlastXmlHit::best_hsp(self : BlastXmlHit) -> BlastXmlHsp? { + if self.hsps.length() == 0 { + return None + } + let mut best = self.hsps[0] + for i = 1; i < self.hsps.length(); i = i + 1 { + let candidate = self.hsps[i] + if candidate.evalue < best.evalue || + (candidate.evalue == best.evalue && candidate.bit_score > best.bit_score) { + best = candidate + } + } + Some(best) +} + +///| +pub fn BlastXmlRecord::number(self : BlastXmlRecord) -> Int { + self.number +} + +///| +pub fn BlastXmlRecord::query(self : BlastXmlRecord) -> BlastXmlQuery { + blast_xml_copy_query(self.query) +} + +///| +pub fn BlastXmlRecord::hits(self : BlastXmlRecord) -> Array[BlastXmlHit] { + blast_xml_copy_hits(self.hits) +} + +///| +pub fn BlastXmlRecord::statistics(self : BlastXmlRecord) -> BlastXmlStatistics? { + self.statistics +} + +///| +pub fn BlastXmlRecord::message(self : BlastXmlRecord) -> String? { + self.message +} + +///| +pub fn BlastXmlRecord::hit( + self : BlastXmlRecord, + identifier : String, +) -> BlastXmlHit? { + for hit in self.hits { + for description in hit.descriptions { + if description.id == identifier || description.accession == identifier { + return Some(hit) + } + } + } + None +} + +///| +pub fn BlastXmlRecord::best_hit(self : BlastXmlRecord) -> BlastXmlHit? { + if self.hits.length() == 0 { + return None + } + let mut best_hit = self.hits[0] + let mut best_hsp = best_hit.best_hsp() + for i = 1; i < self.hits.length(); i = i + 1 { + let candidate_hsp = self.hits[i].best_hsp() + match (best_hsp, candidate_hsp) { + (None, Some(value)) => { + best_hit = self.hits[i] + best_hsp = Some(value) + } + (Some(current), Some(candidate)) => + if candidate.evalue < current.evalue || + ( + candidate.evalue == current.evalue && + candidate.bit_score > current.bit_score + ) { + best_hit = self.hits[i] + best_hsp = Some(candidate) + } + _ => () + } + } + Some(best_hit) +} + +///| +pub fn BlastXmlRecord::filter_evalue( + self : BlastXmlRecord, + maximum : Double, +) -> BlastXmlRecord raise BlastXmlError { + if maximum < 0.0 || maximum.abs() > 1.0e300 { + blast_xml_fail("BLAST XML E-value threshold must be finite and nonnegative") + } + let hits : Array[BlastXmlHit] = [] + for hit in self.hits { + let hsps : Array[BlastXmlHsp] = [] + for hsp in hit.hsps { + if hsp.evalue <= maximum { + hsps.push(hsp) + } + } + if hsps.length() > 0 { + hits.push(BlastXmlHit::{ + number: hit.number, + descriptions: blast_xml_copy_descriptions(hit.descriptions), + length: hit.length, + hsps, + }) + } + } + BlastXmlRecord::{ + number: self.number, + query: blast_xml_copy_query(self.query), + hits, + statistics: self.statistics, + message: self.message, + } +} + +///| +pub fn BlastXmlDocument::format(self : BlastXmlDocument) -> BlastXmlFormat { + self.format +} + +///| +pub fn BlastXmlDocument::program(self : BlastXmlDocument) -> String { + self.program +} + +///| +pub fn BlastXmlDocument::version(self : BlastXmlDocument) -> String { + self.version +} + +///| +pub fn BlastXmlDocument::reference(self : BlastXmlDocument) -> String { + self.reference +} + +///| +pub fn BlastXmlDocument::database(self : BlastXmlDocument) -> String { + self.database +} + +///| +pub fn BlastXmlDocument::global_query( + self : BlastXmlDocument, +) -> BlastXmlQuery? { + match self.global_query { + Some(query) => Some(blast_xml_copy_query(query)) + None => None + } +} + +///| +pub fn BlastXmlDocument::parameters( + self : BlastXmlDocument, +) -> BlastXmlParameters { + self.parameters +} + +///| +pub fn BlastXmlDocument::records( + self : BlastXmlDocument, +) -> Array[BlastXmlRecord] { + blast_xml_copy_records(self.records) +} + +///| +pub fn BlastXmlDocument::megablast_statistics( + self : BlastXmlDocument, +) -> BlastXmlStatistics? { + self.megablast_statistics +} + +///| +pub fn BlastXmlDocument::record( + self : BlastXmlDocument, + query_id : String, +) -> BlastXmlRecord? { + for record in self.records { + if record.query.id == query_id { + return Some(record) + } + } + None +} + +///| +pub fn BlastXmlDocument::total_hits(self : BlastXmlDocument) -> Int { + let mut total = 0 + for record in self.records { + total = total + record.hits.length() + } + total +} + +///| +pub fn BlastXmlDocument::total_hsps(self : BlastXmlDocument) -> Int { + let mut total = 0 + for record in self.records { + for hit in record.hits { + total = total + hit.hsps.length() + } + } + total +} + +///| +fn blast_xml_is_space(code : Int) -> Bool { + code == 32 || code == 9 || code == 10 || code == 13 +} + +///| +fn blast_xml_trim(value : String) -> String { + let mut start = 0 + let mut end = value.length() + while start < end && blast_xml_is_space(value.unsafe_get(start).to_int()) { + start = start + 1 + } + while end > start && blast_xml_is_space(value.unsafe_get(end - 1).to_int()) { + end = end - 1 + } + value[start:end].to_owned() +} + +///| +fn blast_xml_starts_at(text : String, index : Int, token : String) -> Bool { + if index < 0 || index + token.length() > text.length() { + return false + } + for i = 0; i < token.length(); i = i + 1 { + if text.unsafe_get(index + i) != token.unsafe_get(i) { + return false + } + } + true +} + +///| +fn blast_xml_find(text : String, token : String, start : Int) -> Int { + if token.length() == 0 { + return start + } + let mut index = start + while index + token.length() <= text.length() { + if blast_xml_starts_at(text, index, token) { + return index + } + index = index + 1 + } + -1 +} + +///| +fn blast_xml_parse_radix_digits(value : String, radix : Int) -> Int? { + if value.length() == 0 { + return None + } + let mut result = 0 + for i = 0; i < value.length(); i = i + 1 { + let code = value.unsafe_get(i).to_int() + let digit = if code >= 48 && code <= 57 { + code - 48 + } else if code >= 65 && code <= 70 { + code - 55 + } else if code >= 97 && code <= 102 { + code - 87 + } else { + return None + } + if digit >= radix { + return None + } + result = result * radix + digit + } + Some(result) +} + +///| +fn blast_xml_character(code : Int) -> Char { + code.unsafe_to_char() +} + +///| +fn blast_xml_entity_character(entity : String) -> Char? { + match entity { + "lt" => Some('<') + "gt" => Some('>') + "amp" => Some('&') + "quot" => Some('"') + "apos" => Some('\'') + "nbsp" => Some(blast_xml_character(160)) + "Auml" => Some(blast_xml_character(196)) + "auml" => Some(blast_xml_character(228)) + "Ouml" => Some(blast_xml_character(214)) + "ouml" => Some(blast_xml_character(246)) + "Uuml" => Some(blast_xml_character(220)) + "uuml" => Some(blast_xml_character(252)) + "szlig" => Some(blast_xml_character(223)) + "eacute" => Some(blast_xml_character(233)) + _ => + if entity.length() > 1 && entity.unsafe_get(0).to_int() == 35 { + let digits = entity[1:].to_owned() + let parsed = if digits.length() > 1 && + ( + digits.unsafe_get(0).to_int() == 120 || + digits.unsafe_get(0).to_int() == 88 + ) { + blast_xml_parse_radix_digits(digits[1:].to_owned(), 16) + } else { + blast_xml_parse_radix_digits(digits, 10) + } + match parsed { + Some(code) => + if code > 0 && code <= 0x10ffff { + Some(code.unsafe_to_char()) + } else { + None + } + None => None + } + } else { + None + } + } +} + +///| +fn blast_xml_unescape_once(value : String) -> String { + let output = StringBuilder::new() + let mut index = 0 + while index < value.length() { + if value.unsafe_get(index).to_int() == 38 { + let semicolon = blast_xml_find(value, ";", index + 1) + if semicolon >= 0 { + let entity = value[index + 1:semicolon].to_owned() + match blast_xml_entity_character(entity) { + Some(character) => { + output.write_char(character) + index = semicolon + 1 + continue + } + None => () + } + } + } + output.write_char(value.unsafe_get(index).unsafe_to_char()) + index = index + 1 + } + output.to_string() +} + +///| +fn blast_xml_unescape(value : String) -> String { + blast_xml_unescape_once(blast_xml_unescape_once(value)) +} + +///| +fn blast_xml_normalize_tag(raw : String) -> String { + let mut value = raw + let colon = blast_xml_find(value, ":", 0) + if colon >= 0 { + value = value[colon + 1:].to_owned() + } + let mut last_underscore = -1 + for i = 0; i < value.length(); i = i + 1 { + if value.unsafe_get(i).to_int() == 95 { + last_underscore = i + } + } + if last_underscore >= 0 { + value = value[last_underscore + 1:].to_owned() + } + value.to_lower() +} + +///| +fn blast_xml_open_end(text : String, start : Int) -> Int raise BlastXmlError { + let mut quote = 0 + let mut index = start + while index < text.length() { + let code = text.unsafe_get(index).to_int() + if quote == 0 && (code == 34 || code == 39) { + quote = code + } else if quote == code { + quote = 0 + } else if quote == 0 && code == 62 { + return index + } + index = index + 1 + } + blast_xml_fail("Unterminated BLAST XML opening tag") + 0 +} + +///| +fn blast_xml_tag_name(opening : String) -> String raise BlastXmlError { + let trimmed = blast_xml_trim(opening) + let mut end = 0 + while end < trimmed.length() && + !blast_xml_is_space(trimmed.unsafe_get(end).to_int()) && + trimmed.unsafe_get(end).to_int() != 47 { + end = end + 1 + } + if end == 0 { + blast_xml_fail("BLAST XML contains an empty tag name") + } + blast_xml_normalize_tag(trimmed[0:end].to_owned()) +} + +///| +fn blast_xml_skip_declaration( + text : String, + start : Int, +) -> Int raise BlastXmlError { + if blast_xml_starts_at(text, start, "", start + 4) + if end < 0 { + blast_xml_fail("Unterminated BLAST XML comment") + } + return end + 3 + } + if blast_xml_starts_at(text, start, "", start + 2) + if end < 0 { + blast_xml_fail("Unterminated BLAST XML processing instruction") + } + return end + 2 + } + if blast_xml_starts_at(text, start, " (BlastXmlNode, Int) raise BlastXmlError { + if start >= text.length() || text.unsafe_get(start).to_int() != 60 { + blast_xml_fail("Expected a BLAST XML opening tag") + } + if blast_xml_starts_at(text, start, " 0 && + trimmed_opening.unsafe_get(trimmed_opening.length() - 1).to_int() == 47 + if self_closing { + return (BlastXmlNode::{ name, text: "", children: [] }, open_end + 1) + } + let children : Array[BlastXmlNode] = [] + let content = StringBuilder::new() + let mut index = open_end + 1 + while index < text.length() { + if blast_xml_starts_at(text, index, "", index + 9) + if end < 0 { + blast_xml_fail("Unterminated BLAST XML CDATA section") + } + content.write_string(text[index + 9:end].to_owned()) + index = end + 3 + continue + } + if blast_xml_starts_at(text, index, "", + ) + assert_eq(blast_xml_test_parse(text).version(), "BLASTP 2.15.0+") +} diff --git a/test/moonbit/bluster_test.mbt b/test/moonbit/bluster_test.mbt new file mode 100644 index 00000000..8b906c56 --- /dev/null +++ b/test/moonbit/bluster_test.mbt @@ -0,0 +1,1039 @@ +///| +fn bluster_test_close( + left : Double, + right : Double, + tolerance : Double, +) -> Bool { + (left - right).abs() <= tolerance +} + +///| +fn bluster_test_some(value : Double?) -> Double { + match value { + Some(actual) => actual + None => abort("expected a numeric value") + } +} + +///| +fn bluster_test_unique_count(values : Array[Int]) -> Int { + let unique : Array[Int] = [] + for value in values { + let mut found = false + for existing in unique { + if existing == value { + found = true + break + } + } + if !found { + unique.push(value) + } + } + unique.length() +} + +///| +fn bluster_test_two_groups() -> Array[Array[Double]] { + [[0.0, 0.0], [0.1, 0.0], [0.2, 0.1], [10.0, 10.0], [10.1, 10.0], [10.2, 10.1]] +} + +///| +fn bluster_test_sce() -> @src.SingleCellExperiment { + let experiment = @src.SingleCellExperiment::new( + [[1.0, 2.0, 3.0, 4.0], [4.0, 3.0, 2.0, 1.0]], + ["G1", "G2"], + ["C1", "C2", "C3", "C4"], + ) + experiment.row_data["symbol"] = ["A", "B"] + experiment.col_data["batch"] = ["A", "A", "B", "B"] + experiment.reduced_dims["PCA"] = [ + [0.0, 0.0], + [0.1, 0.0], + [10.0, 10.0], + [10.1, 10.0], + ] + experiment.metadata["project"] = "bluster-test" + experiment.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[5.0, 6.0, 7.0, 8.0]], + ["Spike1"], + ["C1", "C2", "C3", "C4"], + ) + experiment +} + +///| +test "bluster: distance and weighting accessors expose enum values" { + assert_true( + @src.bluster_euclidean_distance() != @src.bluster_manhattan_distance(), + ) + assert_true( + @src.bluster_manhattan_distance() != @src.bluster_cosine_distance(), + ) + assert_true(@src.bluster_rank_weight() != @src.bluster_number_weight()) + assert_true(@src.bluster_number_weight() != @src.bluster_jaccard_weight()) +} + +///| +test "bluster: link denominator accessors expose enum values" { + assert_true(@src.bluster_link_minimum() != @src.bluster_link_maximum()) + assert_true(@src.bluster_link_maximum() != @src.bluster_link_union()) +} + +///| +test "bluster: k-means configuration preserves explicit settings" { + let config = @src.BlusterKmeansConfig::create( + 3, + max_iterations=25, + starts=4, + tolerance=1.0e-6, + seed=17, + ) catch { + _ => abort("valid k-means configuration should build") + } + assert_eq(config.centers, 3) + assert_eq(config.max_iterations, 25) + assert_eq(config.starts, 4) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.seed, 17) +} + +///| +test "bluster: k-means configuration rejects invalid controls" { + let mut failures = 0 + ignore(@src.BlusterKmeansConfig::create(0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterKmeansConfig::create(2, starts=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterKmeansConfig::create(2, max_iterations=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterKmeansConfig::create(2, tolerance=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterKmeansConfig::create(2, seed=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 5) +} + +///| +test "bluster: graph configuration preserves explicit settings" { + let config = @src.BlusterGraphConfig::create( + k=7, + shared=false, + snn_weight=@src.bluster_jaccard_weight(), + distance=@src.bluster_manhattan_distance(), + resolution=0.5, + max_iterations=30, + seed=19, + ) catch { + _ => abort("valid graph configuration should build") + } + assert_eq(config.k, 7) + assert_false(config.shared) + assert_true(config.snn_weight == @src.bluster_jaccard_weight()) + assert_true(config.distance == @src.bluster_manhattan_distance()) + assert_eq(config.resolution, 0.5) + assert_eq(config.max_iterations, 30) + assert_eq(config.seed, 19) +} + +///| +test "bluster: graph configuration rejects invalid controls" { + let mut failures = 0 + ignore(@src.BlusterGraphConfig::create(k=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterGraphConfig::create(resolution=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterGraphConfig::create(max_iterations=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterGraphConfig::create(seed=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "bluster: two-step configuration validates both stages" { + let valid = @src.BlusterTwoStepConfig::create( + first_centers=4, + first_starts=2, + second_k=3, + seed=23, + ) catch { + _ => abort("valid two-step configuration should build") + } + assert_eq(valid.first_centers, 4) + assert_eq(valid.first_starts, 2) + assert_eq(valid.second_k, 3) + let mut failures = 0 + ignore(@src.BlusterTwoStepConfig::create(first_centers=-1)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterTwoStepConfig::create(first_starts=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.BlusterTwoStepConfig::create(second_k=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "bluster: directed KNN uses exact nearest neighbors" { + let graph = @src.bluster_make_knn_graph( + [[0.0], [1.0], [3.0], [10.0]], + k=1, + directed=true, + ) catch { + _ => abort("valid directed KNN graph should build") + } + assert_true(graph.directed) + assert_eq(graph.k, 1) + assert_eq(graph.graph_kind, "KNN") + assert_eq(graph.node_count(), 4) + assert_eq(graph.edge_count(), 4) + assert_eq(graph.edge_weight(0, 1), 1.0) + assert_eq(graph.edge_weight(1, 0), 1.0) + assert_eq(graph.edge_weight(2, 1), 1.0) + assert_eq(graph.edge_weight(3, 2), 1.0) + assert_eq(graph.edge_weight(1, 2), 0.0) +} + +///| +test "bluster: KNN tie ordering is deterministic by observation index" { + let graph = @src.bluster_make_knn_graph( + [[0.0], [2.0], [-2.0]], + k=1, + directed=true, + ) catch { + _ => abort("tied KNN graph should build") + } + assert_eq(graph.edge_weight(0, 1), 1.0) + assert_eq(graph.edge_weight(0, 2), 0.0) +} + +///| +test "bluster: undirected KNN symmetrizes one-way relationships" { + let graph = @src.bluster_make_knn_graph([[0.0], [1.0], [3.0], [10.0]], k=1) catch { + _ => abort("valid undirected KNN graph should build") + } + assert_false(graph.directed) + assert_eq(graph.edge_count(), 3) + assert_eq(graph.edge_weight(1, 2), 1.0) + assert_eq(graph.edge_weight(2, 1), 1.0) +} + +///| +test "bluster: KNN caps k at the available observation count" { + let graph = @src.bluster_make_knn_graph([[0.0], [1.0], [2.0]], k=20) catch { + _ => abort("oversized k should be capped") + } + assert_eq(graph.k, 2) + assert_eq(graph.edge_count(), 3) +} + +///| +test "bluster: rank SNN weighting includes self at rank zero" { + let graph = @src.bluster_make_snn_graph( + [[0.0], [1.0], [10.0]], + k=1, + weighting=@src.bluster_rank_weight(), + ) catch { + _ => abort("rank SNN graph should build") + } + assert_eq(graph.graph_kind, "SNN") + assert_eq(graph.edge_count(), 3) + assert_true(bluster_test_close(graph.edge_weight(0, 1), 0.5, 1.0e-12)) + assert_true(bluster_test_close(graph.edge_weight(0, 2), 1.0e-6, 1.0e-15)) + assert_true(bluster_test_close(graph.edge_weight(1, 2), 0.5, 1.0e-12)) +} + +///| +test "bluster: number SNN weighting counts shared neighborhoods" { + let graph = @src.bluster_make_snn_graph( + [[0.0], [1.0], [10.0]], + k=1, + weighting=@src.bluster_number_weight(), + ) catch { + _ => abort("number SNN graph should build") + } + assert_eq(graph.edge_weight(0, 1), 2.0) + assert_eq(graph.edge_weight(0, 2), 1.0) + assert_eq(graph.edge_weight(1, 2), 1.0) +} + +///| +test "bluster: Jaccard SNN weighting uses neighborhood unions" { + let graph = @src.bluster_make_snn_graph( + [[0.0], [1.0], [10.0]], + k=1, + weighting=@src.bluster_jaccard_weight(), + ) catch { + _ => abort("Jaccard SNN graph should build") + } + assert_eq(graph.edge_weight(0, 1), 1.0) + assert_true(bluster_test_close(graph.edge_weight(0, 2), 1.0 / 3.0, 1.0e-12)) + assert_true(bluster_test_close(graph.edge_weight(1, 2), 1.0 / 3.0, 1.0e-12)) +} + +///| +test "bluster: SNN caps k and forms complete full-neighborhood graph" { + let graph = @src.bluster_make_snn_graph( + [[0.0], [1.0], [2.0]], + k=20, + weighting=@src.bluster_number_weight(), + ) catch { + _ => abort("oversized SNN k should be capped") + } + assert_eq(graph.k, 2) + assert_eq(graph.edge_count(), 3) + assert_eq(graph.edge_weight(0, 2), 3.0) +} + +///| +test "bluster: graph constructors reject invalid matrices and k" { + let mut failures = 0 + ignore(@src.bluster_make_snn_graph([], k=1)) catch { + _ => failures = failures + 1 + } + ignore(@src.bluster_make_knn_graph([[1.0], [1.0, 2.0]], k=1)) catch { + _ => failures = failures + 1 + } + ignore(@src.bluster_make_snn_graph([[1.0e301]], k=1)) catch { + _ => failures = failures + 1 + } + ignore(@src.bluster_make_knn_graph([[1.0]], k=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "bluster: k-means separates clear groups and reports WSS" { + let config = @src.BlusterKmeansConfig::create(2, starts=5, seed=31) catch { + _ => abort("valid k-means configuration should build") + } + let result = @src.bluster_cluster_rows_kmeans( + [[0.0, 0.0], [0.0, 1.0], [10.0, 10.0], [10.0, 11.0]], + config, + ) catch { + _ => abort("k-means should cluster valid data") + } + assert_eq(result.clusters.length(), 4) + assert_eq(result.centers.length(), 2) + assert_true(result.clusters[0] == result.clusters[1]) + assert_true(result.clusters[2] == result.clusters[3]) + assert_true(result.clusters[0] != result.clusters[2]) + assert_true(bluster_test_close(result.total_within_sum_squares, 1.0, 1.0e-12)) + assert_true( + bluster_test_close( + result.within_sum_squares[0] + result.within_sum_squares[1], + result.total_within_sum_squares, + 1.0e-12, + ), + ) +} + +///| +test "bluster: k-means is reproducible for a fixed seed" { + let config = @src.BlusterKmeansConfig::create(2, starts=4, seed=37) catch { + _ => abort("valid k-means configuration should build") + } + let data = bluster_test_two_groups() + let first = @src.bluster_cluster_rows_kmeans(data, config) catch { + _ => abort("first k-means run should work") + } + let second = @src.bluster_cluster_rows_kmeans(data, config) catch { + _ => abort("second k-means run should work") + } + assert_eq(first.clusters, second.clusters) + assert_eq(first.centers, second.centers) + assert_eq(first.total_within_sum_squares, second.total_within_sum_squares) +} + +///| +test "bluster: multiple starts cannot worsen the selected WSS" { + let one = @src.BlusterKmeansConfig::create(3, starts=1, seed=41) catch { + _ => abort("single-start configuration should build") + } + let many = @src.BlusterKmeansConfig::create(3, starts=8, seed=41) catch { + _ => abort("multi-start configuration should build") + } + let data = [ + [0.0, 0.0], + [0.2, 0.0], + [5.0, 5.0], + [5.2, 5.0], + [10.0, 0.0], + [10.2, 0.0], + ] + let first = @src.bluster_cluster_rows_kmeans(data, one) catch { + _ => abort("single-start k-means should work") + } + let second = @src.bluster_cluster_rows_kmeans(data, many) catch { + _ => abort("multi-start k-means should work") + } + assert_true( + second.total_within_sum_squares <= first.total_within_sum_squares + 1.0e-12, + ) +} + +///| +test "bluster: k-means rejects more centers than observations" { + let config = @src.BlusterKmeansConfig::create(3) catch { + _ => abort("configuration validation is independent of data") + } + let mut raised = false + ignore(@src.bluster_cluster_rows_kmeans([[0.0], [1.0]], config)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "bluster: graph clustering separates disconnected populations" { + let config = @src.BlusterGraphConfig::create(k=1, resolution=0.5, seed=43) catch { + _ => abort("valid graph configuration should build") + } + let result = @src.bluster_cluster_rows_graph( + bluster_test_two_groups(), + config, + ) catch { + _ => abort("graph clustering should work") + } + assert_eq(result.clusters.length(), 6) + assert_true(result.clusters[0] == result.clusters[1]) + assert_true(result.clusters[1] == result.clusters[2]) + assert_true(result.clusters[3] == result.clusters[4]) + assert_true(result.clusters[4] == result.clusters[5]) + assert_true(result.clusters[0] != result.clusters[3]) + assert_eq(bluster_test_unique_count(result.clusters), 2) +} + +///| +test "bluster: KNN graph clustering uses the non-shared path" { + let config = @src.BlusterGraphConfig::create( + k=1, + shared=false, + resolution=0.5, + seed=47, + ) catch { + _ => abort("valid KNN graph configuration should build") + } + let result = @src.bluster_cluster_rows_graph( + bluster_test_two_groups(), + config, + ) catch { + _ => abort("KNN graph clustering should work") + } + assert_eq(result.graph.graph_kind, "KNN") + assert_true(result.clusters[0] != result.clusters[3]) +} + +///| +test "bluster: graph clustering is reproducible for a fixed seed" { + let config = @src.BlusterGraphConfig::create(k=1, seed=53) catch { + _ => abort("valid graph configuration should build") + } + let first = @src.bluster_cluster_rows_graph(bluster_test_two_groups(), config) catch { + _ => abort("first graph clustering should work") + } + let second = @src.bluster_cluster_rows_graph( + bluster_test_two_groups(), + config, + ) catch { + _ => abort("second graph clustering should work") + } + assert_eq(first.clusters, second.clusters) + assert_eq(first.modularity, second.modularity) +} + +///| +test "bluster: higher resolution does not coarsen this graph" { + let low = @src.BlusterGraphConfig::create(k=1, resolution=0.2, seed=59) catch { + _ => abort("low-resolution configuration should build") + } + let high = @src.BlusterGraphConfig::create(k=1, resolution=3.0, seed=59) catch { + _ => abort("high-resolution configuration should build") + } + let data = bluster_test_two_groups() + let low_result = @src.bluster_cluster_rows_graph(data, low) catch { + _ => abort("low-resolution clustering should work") + } + let high_result = @src.bluster_cluster_rows_graph(data, high) catch { + _ => abort("high-resolution clustering should work") + } + assert_true( + bluster_test_unique_count(high_result.clusters) >= + bluster_test_unique_count(low_result.clusters), + ) +} + +///| +test "bluster: a one-node graph remains one cluster" { + let config = @src.BlusterGraphConfig::create(k=10) catch { + _ => abort("valid graph configuration should build") + } + let result = @src.bluster_cluster_rows_graph([[1.0, 2.0]], config) catch { + _ => abort("one-node graph clustering should work") + } + assert_eq(result.clusters, [0]) + assert_eq(result.graph.k, 0) + assert_eq(result.modularity, 0.0) +} + +///| +test "bluster: two-step clustering maps centroid labels to observations" { + let config = @src.BlusterTwoStepConfig::create( + first_centers=4, + first_starts=4, + second_k=1, + resolution=0.5, + seed=61, + ) catch { + _ => abort("valid two-step configuration should build") + } + let data = [ + [0.0, 0.0], + [0.1, 0.0], + [0.2, 0.1], + [0.3, 0.1], + [10.0, 10.0], + [10.1, 10.0], + [10.2, 10.1], + [10.3, 10.1], + ] + let result = @src.bluster_cluster_rows_two_step(data, config) catch { + _ => abort("two-step clustering should work") + } + assert_eq(result.clusters.length(), data.length()) + assert_eq(result.centroids.length(), 4) + assert_eq(result.centroids, result.first.centers) + assert_eq(result.centroid_clusters.length(), 4) + for observation in 0.. abort("automatic two-step configuration should build") + } + let result = @src.bluster_cluster_rows_two_step( + [[0.0], [1.0], [2.0], [10.0], [11.0], [12.0], [20.0], [21.0], [22.0]], + config, + ) catch { + _ => abort("automatic two-step clustering should work") + } + assert_eq(result.centroids.length(), 3) +} + +///| +test "bluster: unadjusted pairwise Rand decomposes correct pairs" { + let result = @src.bluster_pairwise_rand( + [0, 0, 1, 1], + [0, 1, 0, 1], + adjusted=false, + ) catch { + _ => abort("pairwise Rand should work") + } + assert_eq(result.cluster_names, [0, 1]) + assert_eq(bluster_test_some(result.correct[0][0]), 0.0) + assert_eq(bluster_test_some(result.totals[0][0]), 1.0) + assert_eq(bluster_test_some(result.correct[0][1]), 2.0) + assert_eq(bluster_test_some(result.totals[0][1]), 4.0) + assert_true( + bluster_test_close(bluster_test_some(result.index), 1.0 / 3.0, 1.0e-12), + ) + assert_true(result.ratios[1][0] is None) +} + +///| +test "bluster: adjusted pairwise Rand matches the standard ARI" { + let result = @src.bluster_pairwise_rand([0, 0, 1, 1], [0, 1, 0, 1]) catch { + _ => abort("adjusted pairwise Rand should work") + } + assert_true( + bluster_test_close(bluster_test_some(result.index), -0.5, 1.0e-12), + ) + assert_true( + bluster_test_close(bluster_test_some(result.ratios[0][0]), -0.5, 1.0e-12), + ) + assert_true( + bluster_test_close(bluster_test_some(result.ratios[0][1]), -0.5, 1.0e-12), + ) +} + +///| +test "bluster: identical clusterings have Rand index one" { + let result = @src.bluster_pairwise_rand([2, 2, 5, 5, 9, 9], [2, 2, 5, 5, 9, 9]) catch { + _ => abort("identical pairwise Rand should work") + } + assert_true(bluster_test_close(bluster_test_some(result.index), 1.0, 1.0e-12)) + for right in 0.. abort("clustering comparison should work") + } + assert_eq(bluster_test_some(compared[0][0]), 1.0) + assert_eq(bluster_test_some(compared[2][2]), 1.0) + assert_eq(compared[0][1], compared[1][0]) + assert_eq(compared[0][2], compared[2][0]) + assert_eq(bluster_test_some(compared[0][2]), 1.0) +} + +///| +test "bluster: Rand and comparison validate equal lengths" { + let mut failures = 0 + ignore(@src.bluster_pairwise_rand([0], [0, 1])) catch { + _ => failures = failures + 1 + } + ignore(@src.bluster_compare_clusterings([[0], [0, 1]])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "bluster: approximate silhouette is one for pure duplicates" { + let result = @src.bluster_approx_silhouette([[0.0], [0.0], [10.0], [10.0]], [ + 0, 0, 1, 1, + ]) catch { + _ => abort("silhouette approximation should work") + } + assert_eq(result.other_clusters, [1, 1, 0, 0]) + for width in result.widths { + assert_eq(width, 1.0) + } +} + +///| +test "bluster: approximate silhouette is zero for identical distributions" { + let result = @src.bluster_approx_silhouette([[0.0], [10.0], [0.0], [10.0]], [ + 0, 0, 1, 1, + ]) catch { + _ => abort("silhouette approximation should work") + } + for width in result.widths { + assert_true(bluster_test_close(width, 0.0, 1.0e-12)) + } +} + +///| +test "bluster: approximate silhouette handles one cluster" { + let result = @src.bluster_approx_silhouette([[0.0], [1.0], [2.0]], [7, 7, 7]) catch { + _ => abort("one-cluster silhouette should work") + } + assert_eq(result.other_clusters, [7, 7, 7]) + assert_eq(result.widths, [0.0, 0.0, 0.0]) +} + +///| +test "bluster: cluster RMSD uses sample variances" { + let result = @src.bluster_cluster_rmsd([[0.0], [2.0], [10.0]], [0, 0, 1]) catch { + _ => abort("cluster RMSD should work") + } + assert_eq(result.cluster_names, [0, 1]) + assert_true( + bluster_test_close(bluster_test_some(result.values[0]), 2.0.sqrt(), 1.0e-12), + ) + assert_true(result.values[1] is None) +} + +///| +test "bluster: cluster RMSD sum mode returns within-cluster sums" { + let result = @src.bluster_cluster_rmsd( + [[0.0], [2.0], [10.0], [14.0]], + [0, 0, 1, 1], + sum_squares=true, + ) catch { + _ => abort("cluster RMSD sum mode should work") + } + assert_eq(bluster_test_some(result.values[0]), 2.0) + assert_eq(bluster_test_some(result.values[1]), 8.0) +} + +///| +test "bluster: neighbor purity is one for separated duplicates" { + let result = @src.bluster_neighbor_purity( + [[0.0], [0.0], [10.0], [10.0]], + [0, 0, 1, 1], + k=1, + ) catch { + _ => abort("neighbor purity should work") + } + assert_eq(result.radius, 0.0) + assert_eq(result.purity, [1.0, 1.0, 1.0, 1.0]) + assert_eq(result.maximum_clusters, [0, 0, 1, 1]) +} + +///| +test "bluster: balanced purity removes cluster frequency effects" { + let result = @src.bluster_neighbor_purity( + [[0.0], [0.0], [0.0]], + [0, 0, 1], + k=1, + ) catch { + _ => abort("balanced neighbor purity should work") + } + assert_eq(result.purity, [0.5, 0.5, 0.5]) + assert_eq(result.maximum_clusters, [0, 0, 0]) +} + +///| +test "bluster: unbalanced purity retains cluster frequency effects" { + let result = @src.bluster_neighbor_purity( + [[0.0], [0.0], [0.0]], + [0, 0, 1], + k=1, + balanced=false, + ) catch { + _ => abort("unbalanced neighbor purity should work") + } + assert_true(bluster_test_close(result.purity[0], 2.0 / 3.0, 1.0e-12)) + assert_true(bluster_test_close(result.purity[2], 1.0 / 3.0, 1.0e-12)) +} + +///| +test "bluster: custom purity weights override frequency balancing" { + let result = @src.bluster_neighbor_purity( + [[0.0], [0.0], [0.0]], + [0, 0, 1], + k=1, + custom_weights=Some([1.0, 1.0, 2.0]), + ) catch { + _ => abort("custom-weight neighbor purity should work") + } + assert_eq(result.purity, [0.5, 0.5, 0.5]) +} + +///| +test "bluster: neighbor purity validates labels and custom weights" { + let mut failures = 0 + ignore(@src.bluster_neighbor_purity([[0.0]], [], k=1)) catch { + _ => failures = failures + 1 + } + ignore( + @src.bluster_neighbor_purity( + [[0.0], [1.0]], + [0, 1], + custom_weights=Some([1.0]), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.bluster_neighbor_purity( + [[0.0], [1.0]], + [0, 1], + custom_weights=Some([1.0, -1.0]), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "bluster: pairwise modularity decomposes observed and expected weights" { + let graph = @src.bluster_make_snn_graph([[0.0], [1.0], [10.0], [11.0]], k=1) catch { + _ => abort("SNN graph should build") + } + let result = @src.bluster_pairwise_modularity(graph, [0, 0, 1, 1]) catch { + _ => abort("pairwise modularity should work") + } + assert_true(bluster_test_close(result.total_weight, 1.0, 1.0e-12)) + assert_true( + bluster_test_close(bluster_test_some(result.observed[0][0]), 0.5, 1.0e-12), + ) + assert_eq(bluster_test_some(result.observed[0][1]), 0.0) + assert_true( + bluster_test_close(bluster_test_some(result.expected[0][0]), 0.25, 1.0e-12), + ) + assert_true( + bluster_test_close(bluster_test_some(result.expected[0][1]), 0.5, 1.0e-12), + ) + assert_true( + bluster_test_close( + bluster_test_some(result.modularity[0][0]), + 0.25, + 1.0e-12, + ), + ) + assert_eq(bluster_test_some(result.ratios[0][0]), 2.0) + assert_eq(bluster_test_some(result.ratios[0][1]), 0.0) +} + +///| +test "bluster: pairwise modularity upper triangle sums to zero" { + let graph = @src.bluster_make_knn_graph([[0.0], [1.0], [3.0], [10.0]], k=1) catch { + _ => abort("KNN graph should build") + } + let result = @src.bluster_pairwise_modularity(graph, [0, 0, 1, 1]) catch { + _ => abort("pairwise modularity should work") + } + let mut total = 0.0 + for right in 0.. abort("SNN graph should build") + } + let merged = @src.bluster_merge_communities(graph, [10, 11, 20, 21], 2) catch { + _ => abort("community merging should work") + } + assert_eq(bluster_test_unique_count(merged), 2) + assert_true(merged[0] == merged[1]) + assert_true(merged[2] == merged[3]) + assert_true(merged[0] != merged[2]) +} + +///| +test "bluster: community merging preserves labels when no merge is needed" { + let graph = @src.bluster_make_knn_graph([[0.0], [1.0], [10.0], [11.0]], k=1) catch { + _ => abort("KNN graph should build") + } + let labels = [10, 10, 20, 20] + let merged = @src.bluster_merge_communities(graph, labels, 5) catch { + _ => abort("no-op community merging should work") + } + assert_eq(merged, labels) +} + +///| +test "bluster: nested clusters identify perfect one-to-many mappings" { + let result = @src.bluster_nested_clusters([0, 0, 1, 1], [10, 11, 20, 20]) catch { + _ => abort("nested cluster mapping should work") + } + assert_eq(result.reference_names, [0, 1]) + assert_eq(result.alternative_names, [10, 11, 20]) + assert_eq(result.proportions[0], [1.0, 0.0]) + assert_eq(result.proportions[1], [1.0, 0.0]) + assert_eq(result.proportions[2], [0.0, 1.0]) + assert_eq(result.alternative_mapping, [0, 0, 1]) + assert_eq(result.alternative_maximum, [1.0, 1.0, 1.0]) + assert_eq(result.reference_scores, [1.0, 1.0]) +} + +///| +test "bluster: nested cluster mapping reports imperfect mixtures" { + let result = @src.bluster_nested_clusters([0, 0, 1, 1], [10, 10, 10, 10]) catch { + _ => abort("mixed nested cluster mapping should work") + } + assert_eq(result.proportions[0], [0.5, 0.5]) + assert_eq(result.alternative_mapping, [0]) + assert_eq(result.reference_scores, [0.5, 0.5]) +} + +///| +test "bluster: minimum link denominator favors nested clusters" { + let result = @src.bluster_link_clusters_matrix( + [0, 0, 1, 1], + [0, 1, 1, 1], + denominator=@src.bluster_link_minimum(), + ) catch { + _ => abort("minimum-denominator linking should work") + } + assert_eq(result, [[1.0, 0.5], [0.0, 1.0]]) +} + +///| +test "bluster: maximum link denominator requires similar sizes" { + let result = @src.bluster_link_clusters_matrix( + [0, 0, 1, 1], + [0, 1, 1, 1], + denominator=@src.bluster_link_maximum(), + ) catch { + _ => abort("maximum-denominator linking should work") + } + assert_true(bluster_test_close(result[0][0], 0.5, 1.0e-12)) + assert_true(bluster_test_close(result[0][1], 1.0 / 3.0, 1.0e-12)) + assert_true(bluster_test_close(result[1][1], 2.0 / 3.0, 1.0e-12)) +} + +///| +test "bluster: union link denominator computes Jaccard correspondence" { + let result = @src.bluster_link_clusters_matrix([0, 0, 1, 1], [0, 1, 1, 1]) catch { + _ => abort("union-denominator linking should work") + } + assert_true(bluster_test_close(result[0][0], 0.5, 1.0e-12)) + assert_true(bluster_test_close(result[0][1], 0.25, 1.0e-12)) + assert_true(bluster_test_close(result[1][1], 2.0 / 3.0, 1.0e-12)) +} + +///| +test "bluster: nested and linked mappings validate equal lengths" { + let mut failures = 0 + ignore(@src.bluster_nested_clusters([0], [0, 1])) catch { + _ => failures = failures + 1 + } + ignore(@src.bluster_link_clusters_matrix([0], [0, 1])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "bluster: bootstrap k-means stability is reproducible" { + let config = @src.BlusterKmeansConfig::create(2, starts=4, seed=71) catch { + _ => abort("bootstrap k-means configuration should build") + } + let data = [[0.0], [0.1], [0.2], [0.3], [10.0], [10.1], [10.2], [10.3]] + let first = @src.bluster_bootstrap_kmeans_stability( + data, + config, + iterations=8, + seed=73, + ) catch { + _ => abort("first bootstrap stability run should work") + } + let second = @src.bluster_bootstrap_kmeans_stability( + data, + config, + iterations=8, + seed=73, + ) catch { + _ => abort("second bootstrap stability run should work") + } + assert_eq(first.cluster_names, [0, 1]) + assert_eq(first.ratios, second.ratios) + assert_eq(first.adjusted_indices, second.adjusted_indices) + assert_eq(first.adjusted_indices.length(), 8) + assert_true(bluster_test_some(first.ratios[0][0]) > 0.5) + assert_true(bluster_test_some(first.ratios[0][1]) > 0.5) + assert_true(bluster_test_some(first.ratios[1][1]) > 0.5) +} + +///| +test "bluster: bootstrap mean aggregation returns the full matrix shape" { + let config = @src.BlusterKmeansConfig::create(2, starts=3, seed=79) catch { + _ => abort("bootstrap k-means configuration should build") + } + let result = @src.bluster_bootstrap_kmeans_stability( + [[0.0], [0.1], [0.2], [10.0], [10.1], [10.2]], + config, + iterations=5, + use_mean=true, + seed=83, + ) catch { + _ => abort("mean bootstrap stability should work") + } + assert_eq(result.ratios.length(), 2) + assert_eq(result.ratios[0].length(), 2) + assert_true(result.ratios[1][0] is None) +} + +///| +test "bluster: bootstrap stability rejects non-positive iterations" { + let config = @src.BlusterKmeansConfig::create(2) catch { + _ => abort("valid k-means configuration should build") + } + let mut raised = false + ignore( + @src.bluster_bootstrap_kmeans_stability( + [[0.0], [1.0]], + config, + iterations=0, + ), + ) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "bluster: SCE clustering writes labels without mutating the input" { + let experiment = bluster_test_sce() + let config = @src.BlusterGraphConfig::create(k=1, resolution=0.5, seed=89) catch { + _ => abort("valid SCE graph configuration should build") + } + let result = @src.bluster_cluster_sce( + experiment, + config, + output_column="community", + ) catch { + _ => abort("SCE graph clustering should work") + } + assert_false(experiment.col_data.contains("community")) + assert_true(result.experiment.col_data.contains("community")) + assert_eq(result.experiment.col_data["community"].length(), 4) + assert_eq(result.experiment.metadata["bluster.reduced_dim"], "PCA") + assert_eq(result.experiment.metadata["bluster.graph"], "SNN") + assert_eq(result.clustering.clusters.length(), 4) +} + +///| +test "bluster: SCE clustering deep-copies nested containers" { + let experiment = bluster_test_sce() + let config = @src.BlusterGraphConfig::create(k=1, seed=97) catch { + _ => abort("valid SCE graph configuration should build") + } + let result = @src.bluster_cluster_sce(experiment, config) catch { + _ => abort("SCE graph clustering should work") + } + result.experiment.assays["counts"][0][0] = 999.0 + result.experiment.row_data["symbol"][0] = "changed" + result.experiment.reduced_dims["PCA"][0][0] = 999.0 + result.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 999.0 + assert_eq(experiment.assays["counts"][0][0], 1.0) + assert_eq(experiment.row_data["symbol"][0], "A") + assert_eq(experiment.reduced_dims["PCA"][0][0], 0.0) + assert_eq( + experiment.alternative_experiments["spike"].assays["counts"][0][0], + 5.0, + ) +} + +///| +test "bluster: SCE clustering validates reduced dimensions" { + let config = @src.BlusterGraphConfig::create(k=1) catch { + _ => abort("valid SCE graph configuration should build") + } + let missing = @src.SingleCellExperiment::new([[1.0, 2.0]], ["G1"], [ + "C1", "C2", + ]) + let mut failures = 0 + ignore(@src.bluster_cluster_sce(missing, config)) catch { + _ => failures = failures + 1 + } + missing.reduced_dims["PCA"] = [[0.0]] + ignore(@src.bluster_cluster_sce(missing, config)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} diff --git a/test/moonbit/bsseq_test.mbt b/test/moonbit/bsseq_test.mbt index a7280c6b..e3775313 100644 --- a/test/moonbit/bsseq_test.mbt +++ b/test/moonbit/bsseq_test.mbt @@ -260,7 +260,10 @@ test "bsseq_find_hypo_dmr" { ///| test "bsseq_compute_methylation_diff" { assert_eq(@src.bsseq_compute_methylation_diff(0.8, 0.3), 0.5) - assert_true(@src.bsseq_compute_methylation_diff(0.1, 0.4) > -0.31 && @src.bsseq_compute_methylation_diff(0.1, 0.4) < -0.29) + assert_true( + @src.bsseq_compute_methylation_diff(0.1, 0.4) > -0.31 && + @src.bsseq_compute_methylation_diff(0.1, 0.4) < -0.29, + ) assert_eq(@src.bsseq_compute_methylation_diff(0.5, 0.5), 0.0) } diff --git a/test/moonbit/bumphunter_test.mbt b/test/moonbit/bumphunter_test.mbt index fc1fc622..8195361f 100644 --- a/test/moonbit/bumphunter_test.mbt +++ b/test/moonbit/bumphunter_test.mbt @@ -13,6 +13,7 @@ test "bump_position_creation" { assert_eq(p.values()[0], 0.5) } +///| test "bump_region_creation" { let r = @src.BumpRegion::new("chr1", 1000, 2000, 3.5, 10.0, 0, 5) assert_eq(r.chrom(), "chr1") @@ -25,10 +26,9 @@ test "bump_region_creation" { assert_eq(r.length(), 1000) } +///| test "bump_result_creation" { - let r = @src.BumpResult::new( - "chr1", 1000, 2000, 3.5, 10.0, 0.01, 0.05, 0, 5, - ) + let r = @src.BumpResult::new("chr1", 1000, 2000, 3.5, 10.0, 0.01, 0.05, 0, 5) assert_eq(r.chrom(), "chr1") assert_eq(r.start(), 1000) assert_eq(r.end_(), 2000) @@ -44,20 +44,18 @@ test "bump_result_creation" { // t-statistic computation // --------------------------------------------------------------------------- +///| test "bump_t_statistics_clear_difference" { // Position with clear difference between groups - let positions = [ - @src.BumpPosition::new("chr1", 100, [1.0, 1.0, 5.0, 5.0]), - ] + let positions = [@src.BumpPosition::new("chr1", 100, [1.0, 1.0, 5.0, 5.0])] let stats = @src.bump_compute_t_statistics_test(positions, 2) // Group1 mean=1, Group2 mean=5 => negative t-stat assert_true(stats[0] < 0.0) } +///| test "bump_t_statistics_no_difference" { - let positions = [ - @src.BumpPosition::new("chr1", 100, [3.0, 3.0, 3.0, 3.0]), - ] + let positions = [@src.BumpPosition::new("chr1", 100, [3.0, 3.0, 3.0, 3.0])] let stats = @src.bump_compute_t_statistics_test(positions, 2) assert_eq(stats[0], 0.0) } @@ -66,6 +64,7 @@ test "bump_t_statistics_no_difference" { // Smoothing // --------------------------------------------------------------------------- +///| test "bump_smooth_basic" { let stats = [1.0, 2.0, 3.0, 4.0, 5.0] let smoothed = @src.bump_smooth_test(stats, 1) @@ -77,6 +76,7 @@ test "bump_smooth_basic" { assert_eq(smoothed[4], 4.5) } +///| test "bump_smooth_preserves_constant" { let stats = [5.0, 5.0, 5.0, 5.0] let smoothed = @src.bump_smooth_test(stats, 1) @@ -89,6 +89,7 @@ test "bump_smooth_preserves_constant" { // Candidate bump finding // --------------------------------------------------------------------------- +///| test "bump_find_candidates_basic" { let positions = [ @src.BumpPosition::new("chr1", 100, [0.0]), @@ -104,6 +105,7 @@ test "bump_find_candidates_basic" { assert_eq(bumps[0].value(), 4.0) } +///| test "bump_find_candidates_none" { let positions = [ @src.BumpPosition::new("chr1", 100, [0.0]), @@ -114,6 +116,7 @@ test "bump_find_candidates_none" { assert_eq(bumps.length(), 0) } +///| test "bump_find_candidates_negative_values" { let positions = [ @src.BumpPosition::new("chr1", 100, [0.0]), @@ -127,6 +130,7 @@ test "bump_find_candidates_negative_values" { assert_true(bumps[0].value() < 0.0) } +///| test "bump_find_candidates_chromosome_boundary" { let positions = [ @src.BumpPosition::new("chr1", 100, [3.0]), @@ -144,6 +148,7 @@ test "bump_find_candidates_chromosome_boundary" { // Full analysis pipeline // --------------------------------------------------------------------------- +///| test "bump_hunt_detects_known_bump" { let data = @src.bump_sample_data() let results = @src.bump_hunt(data, 4, 2.0, 2, 20) @@ -153,17 +158,20 @@ test "bump_hunt_detects_known_bump" { assert_eq(results[0].chrom(), "chr1") } +///| test "bump_hunt_empty_input" { let results = @src.bump_hunt([], 4, 2.0, 2, 10) assert_eq(results.length(), 0) } +///| test "bump_hunt_zero_group_size" { let data = @src.bump_sample_data() let results = @src.bump_hunt(data, 0, 2.0, 2, 10) assert_eq(results.length(), 0) } +///| test "bump_hunt_high_cutoff_finds_nothing" { let data = @src.bump_sample_data() // Very high cutoff should find no bumps @@ -171,6 +179,7 @@ test "bump_hunt_high_cutoff_finds_nothing" { assert_eq(results.length(), 0) } +///| test "bump_hunt_result_has_valid_pvalue" { let data = @src.bump_sample_data() let results = @src.bump_hunt(data, 4, 2.0, 2, 20) @@ -180,6 +189,7 @@ test "bump_hunt_result_has_valid_pvalue" { } } +///| test "bump_hunt_result_has_valid_fdr" { let data = @src.bump_sample_data() let results = @src.bump_hunt(data, 4, 2.0, 2, 20) @@ -193,6 +203,7 @@ test "bump_hunt_result_has_valid_fdr" { // Sample data // --------------------------------------------------------------------------- +///| test "bump_sample_data_structure" { let data = @src.bump_sample_data() assert_eq(data.length(), 30) @@ -204,11 +215,24 @@ test "bump_sample_data_structure" { assert_eq(data[29].pos(), 3900) } +///| test "bump_sample_data_has_bump_signal" { let data = @src.bump_sample_data() // Cases (first 4 samples) should have higher values in positions 10-20 - let case_mean_at_5 = (data[5].values()[0] + data[5].values()[1] + data[5].values()[2] + data[5].values()[3]) / 4.0 - let case_mean_at_15 = (data[15].values()[0] + data[15].values()[1] + data[15].values()[2] + data[15].values()[3]) / 4.0 + let case_mean_at_5 = ( + data[5].values()[0] + + data[5].values()[1] + + data[5].values()[2] + + data[5].values()[3] + ) / + 4.0 + let case_mean_at_15 = ( + data[15].values()[0] + + data[15].values()[1] + + data[15].values()[2] + + data[15].values()[3] + ) / + 4.0 // Position 15 (in bump region) should have higher case mean than position 5 (outside) assert_true(case_mean_at_15 > case_mean_at_5) } @@ -217,16 +241,16 @@ test "bump_sample_data_has_bump_signal" { // Output formatting // --------------------------------------------------------------------------- +///| test "bump_result_to_string" { - let r = @src.BumpResult::new( - "chr1", 1000, 2000, 3.5, 10.0, 0.01, 0.05, 0, 5, - ) + let r = @src.BumpResult::new("chr1", 1000, 2000, 3.5, 10.0, 0.01, 0.05, 0, 5) let s = r.to_string() assert_true(s.contains("chr1")) assert_true(s.contains("1000")) assert_true(s.contains("2000")) } +///| test "bump_results_to_string" { let results = [ @src.BumpResult::new("chr1", 1000, 2000, 3.5, 10.0, 0.01, 0.05, 0, 5), @@ -240,6 +264,7 @@ test "bump_results_to_string" { // Area computation // --------------------------------------------------------------------------- +///| test "bump_area_is_sum_of_smoothed" { let positions = [ @src.BumpPosition::new("chr1", 100, [0.0]), @@ -258,6 +283,7 @@ test "bump_area_is_sum_of_smoothed" { // Bump length // --------------------------------------------------------------------------- +///| test "bump_length_calculation" { let r = @src.BumpRegion::new("chr1", 1000, 2500, 3.0, 9.0, 0, 3) assert_eq(r.length(), 1500) diff --git a/test/moonbit/caps_test.mbt b/test/moonbit/caps_test.mbt index 8d680a80..1226c2e6 100644 --- a/test/moonbit/caps_test.mbt +++ b/test/moonbit/caps_test.mbt @@ -6,9 +6,7 @@ // --------------------------------------------------------------------------- test "caps_differential_cutsite_creation" { - let dc = @src.CapsDifferentialCutsite::new( - 11, "EcoRI", [0], [1], - ) + let dc = @src.CapsDifferentialCutsite::new(11, "EcoRI", [0], [1]) assert_eq(dc.start(), 11) assert_eq(dc.enzyme_name(), "EcoRI") assert_eq(dc.cuts_in().length(), 1) @@ -17,10 +15,9 @@ test "caps_differential_cutsite_creation" { assert_eq(dc.blocked_in()[0], 1) } +///| test "caps_differential_cutsite_to_string" { - let dc = @src.CapsDifferentialCutsite::new( - 11, "EcoRI", [0], [1], - ) + let dc = @src.CapsDifferentialCutsite::new(11, "EcoRI", [0], [1]) let s = dc.to_string() assert_true(s.contains("DifferentialCutsite")) assert_true(s.contains("pos=11")) @@ -33,6 +30,7 @@ test "caps_differential_cutsite_to_string" { // CAPS map construction // --------------------------------------------------------------------------- +///| test "caps_map_basic" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -42,6 +40,7 @@ test "caps_map_basic" { assert_true(m.enzymes().length() >= 1) } +///| test "caps_map_named" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -51,6 +50,7 @@ test "caps_map_named" { assert_eq(m.sequence_names()[1], "strainB") } +///| test "caps_map_empty_sequences" { let enzymes = @src.caps_sample_enzymes() let m = @src.caps_map([], enzymes) @@ -59,6 +59,7 @@ test "caps_map_empty_sequences" { assert_eq(m.dcut_count(), 0) } +///| test "caps_map_single_sequence" { // Single sequence: no differential cutting possible let enzymes = @src.caps_sample_enzymes() @@ -71,6 +72,7 @@ test "caps_map_single_sequence" { // Differential cutsite detection // --------------------------------------------------------------------------- +///| test "caps_detects_ecori_differential" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -86,6 +88,7 @@ test "caps_detects_ecori_differential" { assert_eq(dc.blocked_in()[0], 1) } +///| test "caps_no_differential_for_identical_sequences" { // Identical sequences: no differential cutting let enzymes = @src.caps_sample_enzymes() @@ -97,6 +100,7 @@ test "caps_no_differential_for_identical_sequences" { assert_eq(m.dcut_count(), 0) } +///| test "caps_no_differential_without_enzyme_site" { // No enzyme recognition site in either sequence let enzymes = @src.caps_sample_enzymes() @@ -107,16 +111,14 @@ test "caps_no_differential_without_enzyme_site" { assert_eq(m.dcut_count(), 0) } +///| test "caps_multiple_differential_cutsites" { // Two sequences with multiple differential sites // seq0: GAATTC...AAGCTT (EcoRI + HindIII sites) // seq1: GAATTT...AAGCTT (only HindIII site) let enzymes = @src.caps_sample_enzymes() let m = @src.caps_map( - [ - "AAAAAAAAGAGAATTCAAAAAGCTTAAAAAA", - "AAAAAAAAGAGAATTTAAAAAGCTTAAAAAA", - ], + ["AAAAAAAAGAGAATTCAAAAAGCTTAAAAAA", "AAAAAAAAGAGAATTTAAAAAGCTTAAAAAA"], enzymes, ) // EcoRI should produce a differential cutsite @@ -131,6 +133,7 @@ test "caps_multiple_differential_cutsites" { // Query methods // --------------------------------------------------------------------------- +///| test "caps_get_dcuts_at_position" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -140,6 +143,7 @@ test "caps_get_dcuts_at_position" { assert_true(dcuts.length() >= 1) } +///| test "caps_has_dcuts_for_enzyme" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -149,6 +153,7 @@ test "caps_has_dcuts_for_enzyme" { assert_true(!m.has_dcuts_for_enzyme("HindIII")) } +///| test "caps_dcut_count" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -161,6 +166,7 @@ test "caps_dcut_count" { // Formatting // --------------------------------------------------------------------------- +///| test "caps_map_to_string" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -171,6 +177,7 @@ test "caps_map_to_string" { assert_true(s.contains("length=30")) } +///| test "caps_map_report" { let seqs = @src.caps_sample_sequences() let enzymes = @src.caps_sample_enzymes() @@ -181,6 +188,7 @@ test "caps_map_report" { assert_true(r.contains("EcoRI")) } +///| test "caps_map_report_no_dcuts" { let enzymes = @src.caps_sample_enzymes() let m = @src.caps_map( @@ -195,6 +203,7 @@ test "caps_map_report_no_dcuts" { // Sample data // --------------------------------------------------------------------------- +///| test "caps_sample_sequences_length" { let seqs = @src.caps_sample_sequences() assert_eq(seqs.length(), 2) @@ -203,6 +212,7 @@ test "caps_sample_sequences_length" { assert_eq(seqs[0].length(), 30) } +///| test "caps_sample_sequences_ecori_site" { let seqs = @src.caps_sample_sequences() // seq0 should contain GAATTC (EcoRI site) @@ -212,6 +222,7 @@ test "caps_sample_sequences_ecori_site" { assert_true(!seqs[1].contains("GAATTC")) } +///| test "caps_sample_enzymes" { let enzymes = @src.caps_sample_enzymes() // Should contain EcoRI and HindIII @@ -219,8 +230,12 @@ test "caps_sample_enzymes" { let mut has_ecori = false let mut has_hindiii = false for e in enzymes { - if e.name == "EcoRI" { has_ecori = true } - if e.name == "HindIII" { has_hindiii = true } + if e.name == "EcoRI" { + has_ecori = true + } + if e.name == "HindIII" { + has_hindiii = true + } } assert_true(has_ecori) assert_true(has_hindiii) @@ -230,14 +245,15 @@ test "caps_sample_enzymes" { // Edge cases // --------------------------------------------------------------------------- +///| test "caps_three_sequences_differential" { // Three sequences: seq0 cut, seq1 not cut, seq2 cut let enzymes = @src.caps_sample_enzymes() let m = @src.caps_map( [ "AAAAAAAAGAGAATTCAAAAAAAAAAAAAA", // has GAATTC - "AAAAAAAAGAGAATTTAAAAAAAAAAAAAA", // has GAATTT (no cut) - "AAAAAAAAGAGAATTCAAAAAAAAAAAAAA", // has GAATTC + "AAAAAAAAGAGAATTTAAAAAAAAAAAAAA", // has GAATTT (no cut) + "AAAAAAAAGAGAATTCAAAAAAAAAAAAAA", // has GAATTC ], enzymes, ) @@ -250,6 +266,7 @@ test "caps_three_sequences_differential" { assert_eq(dc.blocked_in()[0], 1) } +///| test "caps_no_enzymes" { let seqs = @src.caps_sample_sequences() let m = @src.caps_map(seqs, []) diff --git a/test/moonbit/cealign_test.mbt b/test/moonbit/cealign_test.mbt new file mode 100644 index 00000000..ba066304 --- /dev/null +++ b/test/moonbit/cealign_test.mbt @@ -0,0 +1,725 @@ +///| +/// Tests for Biopython Bio.PDB.cealign-compatible structure alignment. + +///| +fn ce_test_aligner(window_size : Int, max_gap : Int) -> @src.CEAligner { + @src.CEAligner::new(window_size~, max_gap~) catch { + _ => abort("valid CEAligner parameters should be accepted") + } +} + +///| +fn ce_test_curve(count : Int) -> Array[@src.Vector3] { + let coordinates : Array[@src.Vector3] = [] + let mut index = 0 + while index < count { + let value = index.to_double() + coordinates.push( + @src.Vector3::new( + value * 1.3, + value * value * 0.17 + (index % 3).to_double() * 0.4, + value * 0.9 + (index % 4).to_double() * 0.6, + ), + ) + index = index + 1 + } + coordinates +} + +///| +fn ce_test_transform_coordinates( + coordinates : Array[@src.Vector3], + cosine : Double, + sine : Double, + translation : @src.Vector3, +) -> Array[@src.Vector3] { + let result : Array[@src.Vector3] = [] + for coordinate in coordinates { + result.push( + @src.Vector3::new( + cosine * coordinate.x - sine * coordinate.y + translation.x, + sine * coordinate.x + cosine * coordinate.y + translation.y, + coordinate.z + translation.z, + ), + ) + } + result +} + +///| +fn ce_test_structure( + id : String, + coordinates : Array[@src.Vector3], + guide_name : String, +) -> @src.Structure { + let residues : Array[@src.Residue] = [] + let mut index = 0 + while index < coordinates.length() { + let guide = @src.Atom::new( + name=guide_name, + coord=coordinates[index], + resname="GLY", + chainid='A', + resseq=index + 1, + element=if guide_name == "CA" { "C" } else { "C" }, + ) + let companion = @src.Atom::new( + name="N", + coord=@src.Vector3::new( + coordinates[index].x + 0.3, + coordinates[index].y - 0.2, + coordinates[index].z + 0.1, + ), + resname="GLY", + chainid='A', + resseq=index + 1, + element="N", + ) + residues.push( + @src.Residue::new(resname="GLY", chainid='A', resseq=index + 1, atoms=[ + guide, companion, + ]), + ) + index = index + 1 + } + @src.Structure::new(id~, models=[ + @src.Model::new(id=0, chains=[@src.Chain::new(id='A', residues~)]), + ]) +} + +///| +fn ce_test_assert_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort("values differ beyond tolerance") + } +} + +///| +test "cealign: creates default and custom aligners" { + let default_aligner = @src.CEAligner::new() catch { + _ => abort("default CEAligner should be valid") + } + let custom = ce_test_aligner(4, 7) + assert_eq(default_aligner.window_size, 8) + assert_eq(default_aligner.max_gap, 30) + assert_eq(default_aligner.reference_length(), 0) + assert_eq(custom.window_size, 4) + assert_eq(custom.max_gap, 7) +} + +///| +test "cealign: rejects invalid parameters" { + let invalid_window = try { + ignore(@src.CEAligner::new(window_size=0)) + false + } catch { + CeAlignError(_) => true + } + let invalid_gap = try { + ignore(@src.CEAligner::new(max_gap=-1)) + false + } catch { + CeAlignError(_) => true + } + assert_true(invalid_window) + assert_true(invalid_gap) +} + +///| +test "cealign: extracts protein CA guide atoms" { + let structure = ce_test_structure("protein", ce_test_curve(4), "CA") + let atoms = @src.cealign_get_guide_atoms(structure) catch { + _ => abort("protein structure should provide guide atoms") + } + assert_eq(atoms.length(), 4) + assert_eq(atoms[0].atom_name, "CA") + assert_eq(atoms[0].residue_number, 1) +} + +///| +test "cealign: extracts nucleic acid C4 prime guide atoms" { + let structure = ce_test_structure("rna", ce_test_curve(4), "C4'") + let atoms = @src.cealign_get_guide_atoms(structure) catch { + _ => abort("nucleic acid structure should provide guide atoms") + } + assert_eq(atoms.length(), 4) + assert_eq(atoms[0].atom_name, "C4'") +} + +///| +test "cealign: prefers CA when a residue contains both guide atom types" { + let ca = @src.Atom::new(name="CA", coord=@src.Vector3::new(1.0, 2.0, 3.0)) + let c4 = @src.Atom::new(name="C4'", coord=@src.Vector3::new(9.0, 9.0, 9.0)) + let residue = @src.Residue::new(resname="MIX", chainid='A', resseq=1, atoms=[ + c4, ca, + ]) + let structure = @src.Structure::new(id="mixed", models=[ + @src.Model::new(id=0, chains=[@src.Chain::new(id='A', residues=[residue])]), + ]) + let atoms = @src.cealign_get_guide_atoms(structure) catch { + _ => abort("mixed residue should provide one guide atom") + } + assert_eq(atoms.length(), 1) + assert_eq(atoms[0].atom_name, "CA") + assert_eq(atoms[0].coordinate.x, 1.0) +} + +///| +test "cealign: sorts guide atoms by model chain and residue" { + let atom_a2 = @src.Atom::new( + name="CA", + coord=@src.Vector3::new(2.0, 0.0, 0.0), + ) + let atom_a1 = @src.Atom::new( + name="CA", + coord=@src.Vector3::new(1.0, 0.0, 0.0), + ) + let atom_b1 = @src.Atom::new( + name="CA", + coord=@src.Vector3::new(3.0, 0.0, 0.0), + ) + let chain_b = @src.Chain::new(id='B', residues=[ + @src.Residue::new(resname="GLY", chainid='B', resseq=1, atoms=[atom_b1]), + ]) + let chain_a = @src.Chain::new(id='A', residues=[ + @src.Residue::new(resname="GLY", chainid='A', resseq=2, atoms=[atom_a2]), + @src.Residue::new(resname="GLY", chainid='A', resseq=1, atoms=[atom_a1]), + ]) + let structure = @src.Structure::new(id="unordered", models=[ + @src.Model::new(id=0, chains=[chain_b, chain_a]), + ]) + let atoms = @src.cealign_get_guide_atoms(structure) catch { + _ => abort("guide atoms should be sortable") + } + assert_eq(atoms[0].chain_id, 'A') + assert_eq(atoms[0].residue_number, 1) + assert_eq(atoms[1].chain_id, 'A') + assert_eq(atoms[1].residue_number, 2) + assert_eq(atoms[2].chain_id, 'B') +} + +///| +test "cealign: rejects structures without guide atoms" { + let atom = @src.Atom::new(name="N", coord=@src.Vector3::new(0.0, 0.0, 0.0)) + let residue = @src.Residue::new(resname="GLY", chainid='A', resseq=1, atoms=[ + atom, + ]) + let structure = @src.Structure::new(id="no-guide", models=[ + @src.Model::new(id=0, chains=[@src.Chain::new(id='A', residues=[residue])]), + ]) + let raised = try { + ignore(@src.cealign_get_guide_atoms(structure)) + false + } catch { + CeAlignError(_) => true + } + assert_true(raised) +} + +///| +test "cealign: builds symmetric distance matrices" { + let matrix = @src.cealign_distance_matrix([ + @src.Vector3::new(0.0, 0.0, 0.0), + @src.Vector3::new(3.0, 4.0, 0.0), + @src.Vector3::new(3.0, 4.0, 12.0), + ]) + assert_eq(matrix.length(), 3) + assert_eq(matrix[0][0], 0.0) + assert_eq(matrix[0][1], 5.0) + assert_eq(matrix[1][0], 5.0) + assert_eq(matrix[1][2], 12.0) +} + +///| +test "cealign: identical fragments have zero dissimilarity" { + let coordinates = ce_test_curve(5) + let distances = @src.cealign_distance_matrix(coordinates) + let similarity = @src.cealign_fragment_similarity( + distances, distances, 0, 0, 4, + ) + assert_eq(similarity, 0.0) +} + +///| +test "cealign: distorted fragments have negative similarity" { + let first = ce_test_curve(5) + let second = first.copy() + second[3] = @src.Vector3::new(20.0, -10.0, 5.0) + let similarity = @src.cealign_fragment_similarity( + @src.cealign_distance_matrix(first), + @src.cealign_distance_matrix(second), + 0, + 0, + 4, + ) + assert_true(similarity < 0.0) +} + +///| +test "cealign: similarity matrix has all fragment starts" { + let similarity = @src.cealign_similarity_matrix( + ce_test_curve(7), + ce_test_curve(6), + 3, + ) catch { + _ => abort("valid coordinate sets should produce a matrix") + } + assert_eq(similarity.length(), 5) + assert_eq(similarity[0].length(), 4) + assert_eq(similarity[0][0], 0.0) +} + +///| +test "cealign: similarity matrix validates full windows" { + let raised = try { + ignore( + @src.cealign_similarity_matrix(ce_test_curve(2), ce_test_curve(4), 3), + ) + false + } catch { + CeAlignError(_) => true + } + assert_true(raised) +} + +///| +test "cealign: identical coordinates produce a full CE path" { + let coordinates = ce_test_curve(8) + let paths = @src.cealign_find_paths(coordinates, coordinates, 2, 0) catch { + _ => abort("identical coordinates should align") + } + assert_true(paths.length() > 0) + assert_eq(paths[0].fragment_count, 4) + assert_eq(paths[0].aligned_length(), 8) + assert_eq(paths[0].reference_path(), [0, 1, 2, 3, 4, 5, 6, 7]) + assert_eq(paths[0].mobile_path(), [0, 1, 2, 3, 4, 5, 6, 7]) +} + +///| +test "cealign: path search returns at most twenty ranked candidates" { + let coordinates = ce_test_curve(12) + let paths = @src.cealign_find_paths(coordinates, coordinates, 2, 2) catch { + _ => abort("valid coordinates should produce CE paths") + } + assert_true(paths.length() <= 20) + let mut index = 1 + while index < paths.length() { + assert_true(paths[index - 1].fragment_count >= paths[index].fragment_count) + index = index + 1 + } +} + +///| +test "cealign: CE paths are strictly monotonic" { + let coordinates = ce_test_curve(10) + let path = (@src.cealign_find_paths(coordinates, coordinates, 2, 1) catch { + _ => abort("valid coordinates should produce CE paths") + })[0] + let mut index = 1 + while index < path.aligned_length() { + assert_true( + path.reference_indices[index] > path.reference_indices[index - 1], + ) + assert_true(path.mobile_indices[index] > path.mobile_indices[index - 1]) + index = index + 1 + } +} + +///| +test "cealign: path accessors return defensive copies" { + let coordinates = ce_test_curve(8) + let path = (@src.cealign_find_paths(coordinates, coordinates, 2, 0) catch { + _ => abort("valid coordinates should produce CE paths") + })[0] + let reference_copy = path.reference_path() + reference_copy[0] = 99 + let mobile_copy = path.mobile_path() + mobile_copy[0] = 99 + assert_eq(path.reference_path()[0], 0) + assert_eq(path.mobile_path()[0], 0) +} + +///| +test "cealign: paths expose aligned index pairs" { + let coordinates = ce_test_curve(6) + let path = (@src.cealign_find_paths(coordinates, coordinates, 2, 0) catch { + _ => abort("valid coordinates should produce CE paths") + })[0] + let pairs = path.aligned_pairs() + assert_eq(pairs.length(), 6) + assert_eq(pairs[0], (0, 0)) + assert_eq(pairs[5], (5, 5)) +} + +///| +test "cealign: path search validates parameters" { + let invalid_gap = try { + ignore(@src.cealign_find_paths(ce_test_curve(4), ce_test_curve(4), 2, -1)) + false + } catch { + CeAlignError(_) => true + } + let invalid_window = try { + ignore(@src.cealign_find_paths(ce_test_curve(4), ce_test_curve(4), 0, 0)) + false + } catch { + CeAlignError(_) => true + } + assert_true(invalid_gap) + assert_true(invalid_window) +} + +///| +test "cealign: align requires a reference structure" { + let aligner = ce_test_aligner(2, 2) + let mobile = ce_test_structure("mobile", ce_test_curve(8), "CA") + let raised = try { + ignore(aligner.align(mobile, transform=false, final_optimization=false)) + false + } catch { + CeAlignError(_) => true + } + assert_true(raised) +} + +///| +test "cealign: reference must contain two complete windows" { + let aligner = ce_test_aligner(3, 2) + let reference = ce_test_structure("short", ce_test_curve(5), "CA") + let raised = try { + ignore(aligner.set_reference(reference)) + false + } catch { + CeAlignError(_) => true + } + assert_true(raised) +} + +///| +test "cealign: mobile structure must contain two complete windows" { + let reference = ce_test_structure("reference", ce_test_curve(8), "CA") + let mobile = ce_test_structure("short-mobile", ce_test_curve(3), "CA") + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should contain enough guide atoms") + } + let raised = try { + ignore(aligner.align(mobile, transform=false, final_optimization=false)) + false + } catch { + CeAlignError(_) => true + } + assert_true(raised) +} + +///| +test "cealign: set_reference retains ordered coordinates" { + let coordinates = ce_test_curve(8) + let reference = ce_test_structure("reference", coordinates, "CA") + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let stored = aligner.reference_coordinates() + assert_eq(aligner.reference_length(), 8) + ce_test_assert_close(stored[0].x, coordinates[0].x, 1.0e-12) + ce_test_assert_close(stored[7].z, coordinates[7].z, 1.0e-12) +} + +///| +test "cealign: translated structures superimpose with low RMSD" { + let reference_coordinates = ce_test_curve(12) + let mobile_coordinates = ce_test_transform_coordinates( + reference_coordinates, + 1.0, + 0.0, + @src.Vector3::new(7.0, -4.0, 2.0), + ) + let reference = ce_test_structure("reference", reference_coordinates, "CA") + let mobile = ce_test_structure("mobile", mobile_coordinates, "CA") + let aligner = ce_test_aligner(3, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = aligner.align(mobile, transform=false, final_optimization=false) catch { + _ => abort("translated structure should align") + } + assert_true(result.rmsd < 1.0e-5) + assert_eq(result.path.aligned_length(), 12) +} + +///| +test "cealign: rotated and translated sample structures align" { + let (reference, mobile) = @src.cealign_sample_structures() + let aligner = @src.CEAligner::new() catch { + _ => abort("default aligner should be valid") + } + let configured = aligner.set_reference(reference) catch { + _ => abort("sample reference should be accepted") + } + let result = configured.align( + mobile, + transform=true, + final_optimization=false, + ) catch { + _ => abort("sample structures should align") + } + assert_true(result.rmsd < 1.0e-4) + assert_eq(result.path.aligned_length(), 32) + assert_true(result.transformed) +} + +///| +test "cealign: transform false preserves returned coordinates" { + let coordinates = ce_test_curve(8) + let mobile_coordinates = ce_test_transform_coordinates( + coordinates, + 0.8, + 0.6, + @src.Vector3::new(2.0, 3.0, -1.0), + ) + let reference = ce_test_structure("reference", coordinates, "CA") + let mobile = ce_test_structure("mobile", mobile_coordinates, "CA") + let original = mobile.get_atoms()[0].coord + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = aligner.align(mobile, transform=false, final_optimization=false) catch { + _ => abort("mobile structure should align") + } + let returned = result.structure.get_atoms()[0].coord + assert_false(result.transformed) + ce_test_assert_close(returned.x, original.x, 1.0e-12) + ce_test_assert_close(returned.y, original.y, 1.0e-12) + ce_test_assert_close(returned.z, original.z, 1.0e-12) +} + +///| +test "cealign: transform true moves guide and companion atoms" { + let coordinates = ce_test_curve(10) + let mobile_coordinates = ce_test_transform_coordinates( + coordinates, + 0.8, + 0.6, + @src.Vector3::new(4.0, -2.0, 3.0), + ) + let reference = ce_test_structure("reference", coordinates, "CA") + let mobile = ce_test_structure("mobile", mobile_coordinates, "CA") + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = aligner.align(mobile, transform=true, final_optimization=false) catch { + _ => abort("mobile structure should align") + } + let transformed_atoms = result.structure.get_atoms() + assert_eq(transformed_atoms.length(), 20) + assert_true(transformed_atoms[0].coord.distance(coordinates[0]) < 1.0e-4) + assert_true( + transformed_atoms[1].coord.distance(mobile.get_atoms()[1].coord) > 0.1, + ) +} + +///| +test "cealign: alignment never mutates the input structure" { + let (reference, mobile) = @src.cealign_sample_structures() + let original = mobile.get_atoms()[0].coord + let aligner = @src.CEAligner::new() catch { + _ => abort("default aligner should be valid") + } + let configured = aligner.set_reference(reference) catch { + _ => abort("sample reference should be accepted") + } + ignore( + configured.align(mobile, transform=true, final_optimization=false) catch { + _ => abort("sample structures should align") + }, + ) + let after = mobile.get_atoms()[0].coord + assert_eq(after.x, original.x) + assert_eq(after.y, original.y) + assert_eq(after.z, original.z) +} + +///| +test "cealign: aligns nucleic acid structures through C4 prime atoms" { + let coordinates = ce_test_curve(8) + let mobile_coordinates = ce_test_transform_coordinates( + coordinates, + 0.8, + 0.6, + @src.Vector3::new(-2.0, 5.0, 1.0), + ) + let reference = ce_test_structure("rna-ref", coordinates, "C4'") + let mobile = ce_test_structure("rna-mobile", mobile_coordinates, "C4'") + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("nucleic reference should be accepted") + } + let result = aligner.align(mobile, transform=false, final_optimization=false) catch { + _ => abort("nucleic structures should align") + } + assert_true(result.rmsd < 1.0e-4) + assert_eq(result.path.aligned_length(), 8) +} + +///| +test "cealign: reports reference and mobile coverage" { + let coordinates = ce_test_curve(10) + let reference = ce_test_structure("reference", coordinates, "CA") + let mobile = ce_test_structure("mobile", coordinates, "CA") + let aligner = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = aligner.align(mobile, transform=false, final_optimization=false) catch { + _ => abort("identical structures should align") + } + assert_eq(result.reference_coverage(), 1.0) + assert_eq(result.mobile_coverage(), 1.0) +} + +///| +test "cealign: local path can bridge an insertion within max_gap" { + let reference_coordinates = ce_test_curve(10) + let mobile_coordinates : Array[@src.Vector3] = [] + let mut index = 0 + while index < reference_coordinates.length() { + if index == 4 { + mobile_coordinates.push(@src.Vector3::new(30.0, -20.0, 15.0)) + } + mobile_coordinates.push(reference_coordinates[index]) + index = index + 1 + } + let paths = @src.cealign_find_paths( + reference_coordinates, mobile_coordinates, 2, 2, + ) catch { + _ => abort("CE should search paths across a short insertion") + } + assert_true(paths.length() > 0) + assert_true(paths[0].aligned_length() >= 8) + let pairs = paths[0].aligned_pairs() + let mut bridges_insertion = false + for pair in pairs { + if pair.0 >= 4 && pair.1 == pair.0 + 1 { + bridges_insertion = true + } + } + assert_true(bridges_insertion) +} + +///| +test "cealign: final optimization never increases RMSD" { + let (reference, mobile) = @src.cealign_sample_structures() + let configured = (@src.CEAligner::new() catch { + _ => abort("default aligner should be valid") + }).set_reference(reference) catch { + _ => abort("sample reference should be accepted") + } + let baseline = configured.align( + mobile, + transform=false, + final_optimization=false, + ) catch { + _ => abort("baseline alignment should succeed") + } + let optimized = configured.align( + mobile, + transform=false, + final_optimization=true, + ) catch { + _ => abort("optimized alignment should succeed") + } + assert_true(optimized.rmsd <= baseline.rmsd + 1.0e-12) + assert_true(optimized.path.z_score >= 3.5) + assert_true(optimized.optimization_applied) +} + +///| +test "cealign: result transform accessors are defensive" { + let coordinates = ce_test_curve(8) + let reference = ce_test_structure("reference", coordinates, "CA") + let mobile = ce_test_structure("mobile", coordinates, "CA") + let configured = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = configured.align( + mobile, + transform=false, + final_optimization=false, + ) catch { + _ => abort("identical structures should align") + } + let rotation = result.rotation_matrix() + let translation = result.translation_vector() + rotation[0][0] = 99.0 + translation[0] = 99.0 + assert_true(result.rotation_matrix()[0][0] < 2.0) + assert_true(result.translation_vector()[0].abs() < 1.0e-6) +} + +///| +test "cealign: validates explicit rigid transforms" { + let structure = ce_test_structure("protein", ce_test_curve(4), "CA") + let bad_rotation = try { + ignore( + @src.cealign_transform_structure(structure, [[1.0, 0.0], [0.0, 1.0]], [ + 0.0, 0.0, 0.0, + ]), + ) + false + } catch { + CeAlignError(_) => true + } + let bad_translation = try { + ignore( + @src.cealign_transform_structure( + structure, + [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]], + [0.0, 0.0], + ), + ) + false + } catch { + CeAlignError(_) => true + } + assert_true(bad_rotation) + assert_true(bad_translation) +} + +///| +test "cealign: identity transform preserves atom coordinates" { + let structure = ce_test_structure("protein", ce_test_curve(4), "CA") + let transformed = @src.cealign_transform_structure( + structure, + [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]], + [0.0, 0.0, 0.0], + ) catch { + _ => abort("identity transform should be valid") + } + let original_coordinate = structure.get_atoms()[3].coord + let transformed_coordinate = transformed.get_atoms()[3].coord + assert_eq(transformed_coordinate.x, original_coordinate.x) + assert_eq(transformed_coordinate.y, original_coordinate.y) + assert_eq(transformed_coordinate.z, original_coordinate.z) +} + +///| +test "cealign: summary includes core alignment metrics" { + let coordinates = ce_test_curve(8) + let reference = ce_test_structure("reference", coordinates, "CA") + let mobile = ce_test_structure("mobile", coordinates, "CA") + let configured = ce_test_aligner(2, 2).set_reference(reference) catch { + _ => abort("reference should be accepted") + } + let result = configured.align( + mobile, + transform=false, + final_optimization=false, + ) catch { + _ => abort("identical structures should align") + } + let summary = result.summary() + assert_true(summary.has_prefix("CEAlignment(")) + assert_true(summary.contains("aligned=8")) + assert_true(summary.contains("rmsd=")) + assert_true(summary.contains("reference_coverage=1")) +} diff --git a/test/moonbit/celda_test.mbt b/test/moonbit/celda_test.mbt new file mode 100644 index 00000000..8066e127 --- /dev/null +++ b/test/moonbit/celda_test.mbt @@ -0,0 +1,1293 @@ +// Tests for Bioconductor celda_CG-inspired joint cell and feature clustering. + +///| +fn celda_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "celda_CG value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn celda_test_sum(values : Array[Double]) -> Double { + let mut total = 0.0 + for value in values { + total = total + value + } + total +} + +///| +fn celda_test_config() -> @src.CeldaCGConfig { + @src.CeldaCGConfig::create( + 3, + 4, + max_iterations=15, + stop_iterations=3, + chains=2, + seed=2026, + ) catch { + _ => abort("valid celda_CG configuration should build") + } +} + +///| +fn celda_test_result() -> @src.CeldaCGResult { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + @src.celda_cg_fit( + counts, + samples, + celda_test_config(), + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("example celda_CG analysis should succeed") + } +} + +///| +fn celda_test_sce() -> @src.SingleCellExperiment { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + let sce = @src.SingleCellExperiment::new(counts, features, cells) + sce.col_data["donor"] = samples + sce.metadata["source"] = "celda_test" + sce +} + +///| +test "celda_CG: default configuration preserves upstream hyperparameters" { + let config = @src.CeldaCGConfig::create(3, 4) catch { + _ => abort("default configuration should build") + } + assert_eq(config.cell_populations, 3) + assert_eq(config.feature_modules, 4) + assert_eq(config.alpha, 1.0) + assert_eq(config.beta, 1.0) + assert_eq(config.delta, 1.0) + assert_eq(config.gamma, 1.0) + assert_eq(config.algorithm, @src.CeldaEm) + assert_eq(config.max_iterations, 200) + assert_eq(config.stop_iterations, 10) + assert_eq(config.chains, 3) + assert_eq(config.tolerance, 1.0e-8) + assert_eq(config.seed, 12345) +} + +///| +test "celda_CG: custom configuration preserves every option" { + let config = @src.CeldaCGConfig::create( + 2, + 3, + alpha=0.5, + beta=0.25, + delta=0.75, + gamma=2.0, + algorithm=@src.CeldaGibbs, + max_iterations=20, + stop_iterations=4, + chains=5, + tolerance=1.0e-6, + seed=99, + ) catch { + _ => abort("custom configuration should build") + } + assert_eq(config.alpha, 0.5) + assert_eq(config.beta, 0.25) + assert_eq(config.delta, 0.75) + assert_eq(config.gamma, 2.0) + assert_eq(config.algorithm, @src.CeldaGibbs) + assert_eq(config.max_iterations, 20) + assert_eq(config.stop_iterations, 4) + assert_eq(config.chains, 5) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.seed, 99) +} + +///| +test "celda_CG: with_dimensions retains inference controls" { + let original = celda_test_config() + let updated = original.with_dimensions(2, 5, seed=77) catch { + _ => abort("valid dimensions should build") + } + assert_eq(updated.cell_populations, 2) + assert_eq(updated.feature_modules, 5) + assert_eq(updated.max_iterations, original.max_iterations) + assert_eq(updated.stop_iterations, original.stop_iterations) + assert_eq(updated.chains, original.chains) + assert_eq(updated.seed, 77) +} + +///| +test "celda_CG: configuration rejects non-positive dimensions" { + let bad_k = try { + ignore(@src.CeldaCGConfig::create(0, 2)) + false + } catch { + CeldaError(_) => true + } + let bad_l = try { + ignore(@src.CeldaCGConfig::create(2, 0)) + false + } catch { + CeldaError(_) => true + } + assert_true(bad_k) + assert_true(bad_l) +} + +///| +test "celda_CG: configuration rejects invalid Dirichlet priors" { + let alpha = try { + ignore(@src.CeldaCGConfig::create(2, 2, alpha=0.0)) + false + } catch { + CeldaError(_) => true + } + let beta = try { + ignore(@src.CeldaCGConfig::create(2, 2, beta=-1.0)) + false + } catch { + CeldaError(_) => true + } + let delta = try { + ignore(@src.CeldaCGConfig::create(2, 2, delta=Double::nan())) + false + } catch { + CeldaError(_) => true + } + let gamma = try { + ignore(@src.CeldaCGConfig::create(2, 2, gamma=1.0e301)) + false + } catch { + CeldaError(_) => true + } + assert_true(alpha) + assert_true(beta) + assert_true(delta) + assert_true(gamma) +} + +///| +test "celda_CG: configuration rejects invalid iteration limits" { + let maximum = try { + ignore(@src.CeldaCGConfig::create(2, 2, max_iterations=0)) + false + } catch { + CeldaError(_) => true + } + let stopping = try { + ignore(@src.CeldaCGConfig::create(2, 2, stop_iterations=0)) + false + } catch { + CeldaError(_) => true + } + assert_true(maximum) + assert_true(stopping) +} + +///| +test "celda_CG: configuration rejects invalid chains and tolerance" { + let chains = try { + ignore(@src.CeldaCGConfig::create(2, 2, chains=0)) + false + } catch { + CeldaError(_) => true + } + let tolerance = try { + ignore(@src.CeldaCGConfig::create(2, 2, tolerance=0.0)) + false + } catch { + CeldaError(_) => true + } + assert_true(chains) + assert_true(tolerance) +} + +///| +test "celda_CG: configuration rejects non-positive random seeds" { + let rejected = try { + ignore(@src.CeldaCGConfig::create(2, 2, seed=0)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: example data uses feature by cell orientation" { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + assert_eq(counts.length(), 12) + assert_eq(counts[0].length(), 12) + assert_eq(features.length(), 12) + assert_eq(cells.length(), 12) + assert_eq(samples.length(), 12) + assert_eq(features[0], "T_marker_1") + assert_eq(cells[8], "M1") +} + +///| +test "celda_CG: fit rejects an empty count matrix" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects a matrix without cells" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects a non-rectangular matrix" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0], [3.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects negative counts" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, -1.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects fractional counts" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 1.5]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects NaN counts" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, Double::nan()]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects effectively infinite counts" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 1.0e301]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects all-zero features" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0], [0.0, 0.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects all-zero cells" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 0.0], [2.0, 0.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects more populations than cells" { + let config = @src.CeldaCGConfig::create(3, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: fit rejects more modules than features" { + let config = @src.CeldaCGConfig::create(1, 2) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0]], [], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: sample labels must match the number of cells" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0]], ["S1"], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: sample labels cannot be blank" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_fit([[1.0, 2.0]], ["S1", " "], config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: sample order follows first appearance" { + let config = @src.CeldaCGConfig::create( + 1, + 1, + max_iterations=1, + stop_iterations=1, + chains=1, + ) catch { + _ => abort("configuration should build") + } + let result = @src.celda_cg_fit([[2.0, 3.0, 4.0]], ["B", "A", "B"], config) catch { + _ => abort("valid sample labels should fit") + } + assert_eq(result.sample_names, ["B", "A"]) + assert_eq(result.sample_indices, [0, 1, 0]) +} + +///| +test "celda_CG: omitted identifiers are generated deterministically" { + let config = @src.CeldaCGConfig::create( + 1, + 1, + max_iterations=1, + stop_iterations=1, + chains=1, + ) catch { + _ => abort("configuration should build") + } + let result = @src.celda_cg_fit([[2.0, 3.0], [3.0, 2.0]], [], config) catch { + _ => abort("generated names should be accepted") + } + assert_eq(result.feature_names, ["Feature1", "Feature2"]) + assert_eq(result.cell_names, ["Cell1", "Cell2"]) + assert_eq(result.sample_names, ["Sample1"]) +} + +///| +test "celda_CG: feature identifiers must match the matrix" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[1.0, 2.0], [2.0, 1.0]], [], config, feature_names=[ + "G1", + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: cell identifiers must match the matrix" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[1.0, 2.0], [2.0, 1.0]], [], config, cell_names=["C1"]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: identifiers cannot be blank" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[1.0, 2.0], [2.0, 1.0]], [], config, feature_names=[ + "G1", " ", + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: identifiers must be unique after trimming" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[1.0, 2.0], [2.0, 1.0]], [], config, feature_names=[ + "G1", " G1 ", + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: predefined cell labels must match and use all populations" { + let config = @src.CeldaCGConfig::create(2, 1) catch { + _ => abort("configuration should build") + } + let wrong_length = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0, 4.0]], [], config, initial_cell_clusters=[ + 0, 1, + ]), + ) + false + } catch { + CeldaError(_) => true + } + let missing = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0, 4.0]], [], config, initial_cell_clusters=[ + 0, 0, 0, + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(wrong_length) + assert_true(missing) +} + +///| +test "celda_CG: predefined cell labels reject out-of-range values" { + let config = @src.CeldaCGConfig::create(2, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0, 4.0]], [], config, initial_cell_clusters=[ + 0, 1, 2, + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: predefined feature labels must match and use all modules" { + let config = @src.CeldaCGConfig::create(1, 2) catch { + _ => abort("configuration should build") + } + let wrong_length = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0], [3.0, 2.0], [1.0, 1.0]], [], config, initial_feature_modules=[ + 0, 1, + ]), + ) + false + } catch { + CeldaError(_) => true + } + let missing = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0], [3.0, 2.0], [1.0, 1.0]], [], config, initial_feature_modules=[ + 0, 0, 0, + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(wrong_length) + assert_true(missing) +} + +///| +test "celda_CG: predefined feature labels reject out-of-range values" { + let config = @src.CeldaCGConfig::create(1, 2) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_fit([[2.0, 3.0], [3.0, 2.0], [1.0, 1.0]], [], config, initial_feature_modules=[ + 0, 1, -1, + ]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: one-population one-module likelihood matches closed form" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let value = @src.celda_cg_log_likelihood( + [[2.0, 1.0], [1.0, 2.0]], + [], + [0, 0], + [0, 0], + config, + ) catch { + _ => abort("valid labels should produce a likelihood") + } + celda_test_close(value, -6.327936783729195, 1.0e-9) +} + +///| +test "celda_CG: omitted and explicit default samples have equal likelihood" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let counts = [[2.0, 1.0], [1.0, 2.0]] + let implicit = @src.celda_cg_log_likelihood( + counts, + [], + [0, 0], + [0, 0], + config, + ) catch { + _ => abort("implicit sample should work") + } + let explicit = @src.celda_cg_log_likelihood( + counts, + ["Sample1", "Sample1"], + [0, 0], + [0, 0], + config, + ) catch { + _ => abort("explicit sample should work") + } + celda_test_close(implicit, explicit, 1.0e-12) +} + +///| +test "celda_CG: likelihood validates assignment occupancy" { + let config = @src.CeldaCGConfig::create(2, 2) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore( + @src.celda_cg_log_likelihood( + [[2.0, 1.0], [1.0, 2.0]], + [], + [0, 0], + [0, 1], + config, + ), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: result dimensions match the example matrix" { + let result = celda_test_result() + assert_eq(result.n_features(), 12) + assert_eq(result.n_cells(), 12) + assert_eq(result.n_samples(), 2) + assert_eq(result.cell_clusters.length(), 12) + assert_eq(result.feature_modules.length(), 12) +} + +///| +test "celda_CG: model selection metrics are finite" { + let result = celda_test_result() + assert_true(result.log_likelihood == result.log_likelihood) + assert_true(result.log_likelihood.abs() < 1.0e300) + assert_true(result.perplexity > 1.0) + assert_true(result.perplexity < 1.0e300) + assert_true(result.aic > 0.0) + assert_true(result.bic > 0.0) +} + +///| +test "celda_CG: diagnostics record bounded optimization work" { + let result = celda_test_result() + assert_true(result.diagnostics.iterations >= 1) + assert_true(result.diagnostics.iterations <= result.config.max_iterations) + assert_eq( + result.diagnostics.log_likelihoods.length(), + result.diagnostics.iterations + 1, + ) + assert_true(result.diagnostics.best_chain >= 0) + assert_true(result.diagnostics.best_chain < result.config.chains) +} + +///| +test "celda_CG: diagnostics include one finite score per chain" { + let result = celda_test_result() + assert_eq(result.diagnostics.chain_scores.length(), result.config.chains) + for score in result.diagnostics.chain_scores { + assert_true(score == score) + assert_true(score.abs() < 1.0e300) + } +} + +///| +test "celda_CG: hard EM is reproducible for a fixed seed" { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + let first = @src.celda_cg_fit( + counts, + samples, + celda_test_config(), + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("first fit should succeed") + } + let second = @src.celda_cg_fit( + counts, + samples, + celda_test_config(), + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("second fit should succeed") + } + assert_eq(first.cell_clusters, second.cell_clusters) + assert_eq(first.feature_modules, second.feature_modules) + assert_eq(first.log_likelihood, second.log_likelihood) +} + +///| +test "celda_CG: seeded Gibbs inference is reproducible" { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + let config = @src.CeldaCGConfig::create( + 3, + 4, + algorithm=@src.CeldaGibbs, + max_iterations=6, + stop_iterations=3, + chains=2, + seed=456, + ) catch { + _ => abort("Gibbs configuration should build") + } + let first = @src.celda_cg_fit( + counts, + samples, + config, + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("first Gibbs fit should succeed") + } + let second = @src.celda_cg_fit( + counts, + samples, + config, + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("second Gibbs fit should succeed") + } + assert_eq(first.cell_clusters, second.cell_clusters) + assert_eq(first.feature_modules, second.feature_modules) + assert_eq(first.log_likelihood, second.log_likelihood) +} + +///| +test "celda_CG: every inferred population remains occupied" { + let result = celda_test_result() + let sizes = result.population_sizes() + assert_eq(sizes.length(), 3) + for size in sizes { + assert_true(size > 0) + } + assert_eq(celda_test_sum(sizes.map(fn(value) { value.to_double() })), 12.0) +} + +///| +test "celda_CG: every inferred feature module remains occupied" { + let result = celda_test_result() + let sizes = result.module_sizes() + assert_eq(sizes.length(), 4) + let mut total = 0 + for size in sizes { + assert_true(size > 0) + total = total + size + } + assert_eq(total, 12) +} + +///| +test "celda_CG: example cell types form three distinct populations" { + let result = celda_test_result() + for index in 1..<4 { + assert_eq(result.cell_clusters[index], result.cell_clusters[0]) + } + for index in 5..<8 { + assert_eq(result.cell_clusters[index], result.cell_clusters[4]) + } + for index in 9..<12 { + assert_eq(result.cell_clusters[index], result.cell_clusters[8]) + } + assert_true(result.cell_clusters[0] != result.cell_clusters[4]) + assert_true(result.cell_clusters[0] != result.cell_clusters[8]) + assert_true(result.cell_clusters[4] != result.cell_clusters[8]) +} + +///| +test "celda_CG: example marker families form distinct modules" { + let result = celda_test_result() + assert_eq(result.feature_modules[0], result.feature_modules[1]) + assert_eq(result.feature_modules[1], result.feature_modules[2]) + assert_eq(result.feature_modules[3], result.feature_modules[4]) + assert_eq(result.feature_modules[4], result.feature_modules[5]) + assert_eq(result.feature_modules[6], result.feature_modules[7]) + assert_eq(result.feature_modules[7], result.feature_modules[8]) + assert_eq(result.feature_modules[9], result.feature_modules[10]) + assert_eq(result.feature_modules[10], result.feature_modules[11]) +} + +///| +test "celda_CG: theta is population by sample and column-normalized" { + let result = celda_test_result() + assert_eq(result.theta.length(), 3) + assert_eq(result.theta[0].length(), 2) + for sample in 0.. 0.0) + } else { + assert_eq(result.psi[feature][module_], 0.0) + } + } + } +} + +///| +test "celda_CG: psi reserves pseudogene probability in every module" { + let result = celda_test_result() + for module_ in 0.. 0.0) + assert_true(total < 1.0) + } +} + +///| +test "celda_CG: eta is a normalized feature-module abundance" { + let result = celda_test_result() + assert_eq(result.eta.length(), 4) + for probability in result.eta { + assert_true(probability > 0.0) + } + celda_test_close(celda_test_sum(result.eta), 1.0, 1.0e-12) +} + +///| +test "celda_CG: fitted counts match source dimensions" { + let result = celda_test_result() + assert_eq(result.fitted_counts.length(), 12) + assert_eq(result.fitted_counts[0].length(), 12) + for row in result.fitted_counts { + for value in row { + assert_true(value >= 0.0) + assert_true(value.abs() < 1.0e300) + } + } +} + +///| +test "celda_CG: fitted counts preserve every cell library size" { + let result = celda_test_result() + for cell in 0.. true + } + let module_ = try { + ignore(result.features_in_module(-1)) + false + } catch { + CeldaError(_) => true + } + assert_true(population) + assert_true(module_) +} + +///| +test "celda_CG: top features are sorted by posterior probability" { + let result = celda_test_result() + let module_ = result.feature_modules[0] + let scores = result.top_features(module_, limit=3) catch { + _ => abort("valid module should have top features") + } + assert_true(scores.length() >= 1) + for index in 1..= scores[index].probability) + } + for score in scores { + assert_eq(result.feature_modules[score.feature_index], module_) + assert_eq(result.feature_names[score.feature_index], score.feature_name) + } +} + +///| +test "celda_CG: top feature query validates module and limit" { + let result = celda_test_result() + let module_ = try { + ignore(result.top_features(4)) + false + } catch { + CeldaError(_) => true + } + let limit = try { + ignore(result.top_features(0, limit=0)) + false + } catch { + CeldaError(_) => true + } + assert_true(module_) + assert_true(limit) +} + +///| +test "celda_CG: summary reports dimensions and model family" { + let summary = celda_test_result().summary() + assert_true(summary.contains("celda_CG")) + assert_true(summary.contains("12 features x 12 cells")) + assert_true(summary.contains("K=3")) + assert_true(summary.contains("L=4")) +} + +///| +test "celda_CG: prediction returns cell by population probabilities" { + let (counts, _, _, _) = @src.celda_cg_example_data() + let result = celda_test_result() + let prediction = @src.celda_cg_predict_cells(result, counts) catch { + _ => abort("training matrix should be predictable") + } + assert_eq(prediction.probabilities.length(), 12) + assert_eq(prediction.probabilities[0].length(), 3) + assert_eq(prediction.assignments.length(), 12) +} + +///| +test "celda_CG: prediction probabilities are normalized" { + let (counts, _, _, _) = @src.celda_cg_example_data() + let prediction = @src.celda_cg_predict_cells(celda_test_result(), counts) catch { + _ => abort("training matrix should be predictable") + } + for probabilities in prediction.probabilities { + celda_test_close(celda_test_sum(probabilities), 1.0, 1.0e-12) + for probability in probabilities { + assert_true(probability >= 0.0) + assert_true(probability <= 1.0) + } + } +} + +///| +test "celda_CG: training cells predict their inferred populations" { + let (counts, _, _, _) = @src.celda_cg_example_data() + let result = celda_test_result() + let prediction = @src.celda_cg_predict_cells(result, counts) catch { + _ => abort("training matrix should be predictable") + } + assert_eq(prediction.assignments, result.cell_clusters) +} + +///| +test "celda_CG: prediction rejects a mismatched feature space" { + let rejected = try { + ignore( + @src.celda_cg_predict_cells(celda_test_result(), [[1.0, 2.0], [2.0, 1.0]]), + ) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: grid search evaluates the Cartesian K by L grid" { + let (counts, features, cells, samples) = @src.celda_cg_example_data() + let base = @src.CeldaCGConfig::create( + 1, + 1, + max_iterations=5, + stop_iterations=2, + chains=1, + seed=41, + ) catch { + _ => abort("base grid configuration should build") + } + let grid = @src.celda_cg_grid_search( + counts, + samples, + [2, 3], + [3, 4], + base, + feature_names=features, + cell_names=cells, + ) catch { + _ => abort("valid grid should fit") + } + assert_eq(grid.entries.length(), 4) + assert_true(grid.best_index >= 0) + assert_true(grid.best_index < 4) + assert_eq(grid.criterion, @src.CeldaBic) +} + +///| +test "celda_CG: grid best_result matches the selected entry" { + let (counts, _, _, samples) = @src.celda_cg_example_data() + let base = @src.CeldaCGConfig::create( + 1, + 1, + max_iterations=3, + stop_iterations=1, + chains=1, + ) catch { + _ => abort("base grid configuration should build") + } + let grid = @src.celda_cg_grid_search( + counts, + samples, + [2, 3], + [4], + base, + criterion=@src.CeldaLogLikelihood, + ) catch { + _ => abort("valid likelihood grid should fit") + } + let best = grid.best_result() + assert_eq(best.log_likelihood, grid.entries[grid.best_index].score) +} + +///| +test "celda_CG: grid search rejects empty dimensions" { + let (counts, _, _, samples) = @src.celda_cg_example_data() + let base = celda_test_config() + let rejected = try { + ignore(@src.celda_cg_grid_search(counts, samples, [], [2], base)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: grid search rejects duplicate K-L combinations" { + let (counts, _, _, samples) = @src.celda_cg_example_data() + let base = celda_test_config() + let rejected = try { + ignore(@src.celda_cg_grid_search(counts, samples, [2, 2], [3], base)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: SCE integration adds fitted assay and assignments" { + let output = @src.celda_cg_sce( + celda_test_sce(), + celda_test_config(), + sample_column="donor", + ) catch { + _ => abort("valid SCE integration should succeed") + } + assert_eq(@src.sce_get_assay(output.experiment, "celda_fitted").length(), 12) + assert_eq( + @src.sce_get_row_data(output.experiment, "celda_feature_module").length(), + 12, + ) + assert_eq( + @src.sce_get_col_data(output.experiment, "celda_cell_population").length(), + 12, + ) +} + +///| +test "celda_CG: SCE integration writes sample and model metadata" { + let output = @src.celda_cg_sce( + celda_test_sce(), + celda_test_config(), + sample_column="donor", + ) catch { + _ => abort("valid SCE integration should succeed") + } + assert_eq(@src.sce_get_col_data(output.experiment, "celda_sample_label"), [ + "Donor1", "Donor1", "Donor2", "Donor2", "Donor1", "Donor1", "Donor2", "Donor2", + "Donor1", "Donor1", "Donor2", "Donor2", + ]) + assert_eq(output.experiment.metadata["celda_model"], "celda_CG") + assert_eq(output.experiment.metadata["celda_K"], "3") + assert_eq(output.experiment.metadata["celda_L"], "4") + assert_eq(output.experiment.metadata["source"], "celda_test") +} + +///| +test "celda_CG: SCE integration does not mutate the source object" { + let source = celda_test_sce() + ignore( + @src.celda_cg_sce(source, celda_test_config(), sample_column="donor") catch { + _ => abort("valid SCE integration should succeed") + }, + ) + assert_eq(@src.sce_get_assay(source, "celda_fitted").length(), 0) + assert_eq(@src.sce_get_row_data(source, "celda_feature_module").length(), 0) + assert_eq(@src.sce_get_col_data(source, "celda_cell_population").length(), 0) +} + +///| +test "celda_CG: SCE integration supports custom assay names" { + let output = @src.celda_cg_sce( + celda_test_sce(), + celda_test_config(), + sample_column="donor", + fitted_assay="celda_reconstructed", + ) catch { + _ => abort("custom fitted assay should work") + } + assert_eq( + @src.sce_get_assay(output.experiment, "celda_reconstructed").length(), + 12, + ) + assert_eq(@src.sce_get_assay(output.experiment, "celda_fitted").length(), 0) +} + +///| +test "celda_CG: SCE integration validates the input assay" { + let source = @src.SCEBuilder::new() + |> @src.SCEBuilder::add_assay("other", [[1.0, 2.0], [2.0, 1.0]]) + |> @src.SCEBuilder::build + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_sce(source, config)) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} + +///| +test "celda_CG: SCE integration validates sample column metadata" { + let config = @src.CeldaCGConfig::create(1, 1) catch { + _ => abort("configuration should build") + } + let rejected = try { + ignore(@src.celda_cg_sce(celda_test_sce(), config, sample_column="missing")) + false + } catch { + CeldaError(_) => true + } + assert_true(rejected) +} diff --git a/test/moonbit/cellchat_test.mbt b/test/moonbit/cellchat_test.mbt index 95bfd817..cefcca2d 100644 --- a/test/moonbit/cellchat_test.mbt +++ b/test/moonbit/cellchat_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Bioconductor CellChat module - Cell-cell communication analysis. - test "lr_pair_create" { let pair = @src.lr_pair("TNF", "TNFR1") assert_eq(pair.ligand, "TNF") @@ -8,12 +7,14 @@ test "lr_pair_create" { assert_eq(pair.key, "TNF_TNFR1") } +///| test "lr_database" { let db = @src.cellchat_lr_database() assert_true(db.length() > 0) assert_eq(db[0].ligand, "TNF") } +///| test "mean_expr_basic" { let expr = [[10.0, 5.0], [20.0, 8.0], [30.0, 12.0]] let cell_types = ["TypeA", "TypeA", "TypeB"] @@ -21,6 +22,7 @@ test "mean_expr_basic" { assert_true((mean - 15.0).abs() < 0.01) } +///| test "mean_expr_no_cells" { let expr = [[10.0, 5.0]] let cell_types = ["TypeA"] @@ -28,40 +30,82 @@ test "mean_expr_no_cells" { assert_eq(mean, 0.0) } +///| test "analyze_basic" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=10, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=10, + seed=42, + ) assert_true(result.scores.length() > 0) assert_eq(result.cell_types.length(), 3) assert_true(result.lr_pairs.length() > 0) } +///| test "analyze_permutation_effect" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result_small = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) - let result_large = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=50, seed=42) + let result_small = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) + let result_large = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=50, + seed=42, + ) assert_true(result_small.scores.length() > 0) assert_true(result_large.scores.length() > 0) } +///| test "analyze_score_non_negative" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) let mut i = 0 while i < result.scores.length() { assert_true(result.scores[i].score >= 0.0) - assert_true(result.scores[i].p_value >= 0.0 && result.scores[i].p_value <= 1.0) + assert_true( + result.scores[i].p_value >= 0.0 && result.scores[i].p_value <= 1.0, + ) i = i + 1 } } +///| test "get_significant" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42, fdr=0.5) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + fdr=0.5, + ) let sig = @src.cellchat_get_significant(result) assert_true(sig.length() >= 0) // All significant scores should have p_value < fdr @@ -72,14 +116,23 @@ test "get_significant" { } } +///| test "aggregate" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) let agg = @src.cellchat_aggregate(result) assert_true(agg.size() > 0) } +///| test "sample_data_dimensions" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() assert_eq(expr.length(), 30) @@ -88,6 +141,7 @@ test "sample_data_dimensions" { assert_eq(gene_names.length(), 10) } +///| test "sample_cell_types" { let (_, cell_types, _) = @src.cellchat_sample_data() let mut type_a = 0 @@ -95,9 +149,13 @@ test "sample_cell_types" { let mut type_c = 0 let mut i = 0 while i < cell_types.length() { - if cell_types[i] == "TypeA" { type_a = type_a + 1 } - else if cell_types[i] == "TypeB" { type_b = type_b + 1 } - else if cell_types[i] == "TypeC" { type_c = type_c + 1 } + if cell_types[i] == "TypeA" { + type_a = type_a + 1 + } else if cell_types[i] == "TypeB" { + type_b = type_b + 1 + } else if cell_types[i] == "TypeC" { + type_c = type_c + 1 + } i = i + 1 } assert_eq(type_a, 10) @@ -105,35 +163,74 @@ test "sample_cell_types" { assert_eq(type_c, 10) } +///| test "get_top" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) let top = @src.cellchat_get_top(result, 5) assert_eq(top.length(), 5) } +///| test "get_top_fewer_than_n" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) let top = @src.cellchat_get_top(result, 100) assert_true(top.length() <= result.scores.length()) } +///| test "summary" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=42) + let result = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=42, + ) let summary = @src.cellchat_summary(result) assert_true(summary.length() > 0) } +///| test "analyze_seed_reproducible" { let (expr, cell_types, gene_names) = @src.cellchat_sample_data() let pairs = @src.cellchat_lr_database() - let result1 = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=123) - let result2 = @src.cellchat_analyze(expr, cell_types, gene_names, pairs, n_permutations=5, seed=123) + let result1 = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=123, + ) + let result2 = @src.cellchat_analyze( + expr, + cell_types, + gene_names, + pairs, + n_permutations=5, + seed=123, + ) // Same seed should give same results assert_eq(result1.scores.length(), result2.scores.length()) } diff --git a/test/moonbit/cellosaurus_test.mbt b/test/moonbit/cellosaurus_test.mbt new file mode 100644 index 00000000..cae72a98 --- /dev/null +++ b/test/moonbit/cellosaurus_test.mbt @@ -0,0 +1,194 @@ +///| +/// Tests for Bio.ExPASy.cellosaurus-compatible parsing. + +///| +test "cellosaurus_parse_official_sample_fields" { + let records = @src.cellosaurus_parse(@src.cellosaurus_sample_text()) + assert_eq(records.length(), 1) + let record = records[0] + assert_eq(record.identifier, "#15310-LN") + assert_eq(record.accession, "CVCL_E548") + assert_eq(record.sex, "Female") + assert_eq(record.age, "Age unspecified") + assert_eq(record.category, "Transformed cell line") + assert_eq(record.comments.length(), 2) + assert_eq(record.species.length(), 1) +} + +///| +test "cellosaurus_parse_cross_references" { + let record = @src.cellosaurus_parse(@src.cellosaurus_sample_text())[0] + assert_eq(record.cross_references.length(), 3) + assert_eq(record.cross_references[0].database, "dbMHC") + assert_eq(record.cross_references[0].accession, "48439") + assert_eq(record.cross_references[2].database, "Wikidata") + assert_eq(record.cross_references[2].accession, "Q54398957") +} + +///| +test "cellosaurus_list_helpers" { + let record = @src.CellosaurusRecord::new( + identifier="XP3OS", + secondary_accessions="CVCL_F511; CVCL_TEST", + synonyms="Xeroderma Pigmentosum 3 OSaka; GM04314; GM4314", + ) + assert_eq(record.secondary_accession_list(), ["CVCL_F511", "CVCL_TEST"]) + assert_eq(record.synonym_list(), [ + "Xeroderma Pigmentosum 3 OSaka", "GM04314", "GM4314", + ]) +} + +///| +test "cellosaurus_cross_reference_filter" { + let text = "ID Example\n" + + "AC CVCL_0001\n" + + "DR JCRB; JCRB0303\n" + + "DR JCRB; KURB1002\n" + + "DR Wikidata; Q1\n" + + "//\n" + let record = @src.cellosaurus_parse(text)[0] + assert_eq(record.cross_references_for("JCRB"), ["JCRB0303", "KURB1002"]) + assert_eq(record.cross_references_for("ATCC"), []) +} + +///| +test "cellosaurus_species_and_repr" { + let record = @src.cellosaurus_parse(@src.cellosaurus_sample_text())[0] + assert_true(record.has_species("Homo sapiens")) + assert_false(record.has_species("Mus musculus")) + assert_eq(record.repr(), "CellosaurusRecord (#15310-LN, CVCL_E548)") +} + +///| +test "cellosaurus_parse_multiple_records" { + let text = "ID XP3OS\n" + + "AC CVCL_3245\n" + + "AS CVCL_F511\n" + + "RX PubMed=832273;\n" + + "ST Amelogenin: X\n" + + "DI ORDO; Orphanet_910; Xeroderma pigmentosum\n" + + "OX NCBI_TaxID=9606; ! Homo sapiens (Human)\n" + + "//\n" + + "ID 1-5c-4\n" + + "AC CVCL_2260\n" + + "HI CVCL_0030 ! HeLa\n" + + "OI CVCL_0002 ! Same donor\n" + + "CA Cancer cell line\n" + + "//\n" + let records = @src.cellosaurus_parse(text) + assert_eq(records.length(), 2) + assert_eq(records[0].secondary_accessions, "CVCL_F511") + assert_eq(records[0].reference_identifiers, ["PubMed=832273;"]) + assert_eq(records[0].str_profiles, ["Amelogenin: X"]) + assert_eq(records[0].diseases.length(), 1) + assert_eq(records[1].identifier, "1-5c-4") + assert_eq(records[1].hierarchy, ["CVCL_0030 ! HeLa"]) + assert_eq(records[1].same_individual, ["CVCL_0002 ! Same donor"]) +} + +///| +test "cellosaurus_read_single" { + match @src.cellosaurus_read(@src.cellosaurus_sample_text()) { + Some(record) => assert_eq(record.accession, "CVCL_E548") + None => fail("expected one Cellosaurus record") + } +} + +///| +test "cellosaurus_read_empty" { + assert_true(@src.cellosaurus_read("") is None) + assert_true(@src.cellosaurus_read("//\n") is None) +} + +///| +test "cellosaurus_read_multiple_error" { + let text = "ID First\nAC CVCL_0001\n//\n" + + "ID Second\nAC CVCL_0002\n//\n" + let raised = try { + ignore(@src.cellosaurus_read(text)) + false + } catch { + CellosaurusError(_) => true + } + assert_true(raised) +} + +///| +test "cellosaurus_unterminated_record_error" { + let raised = try { + ignore(@src.cellosaurus_parse("ID Incomplete\nAC CVCL_0001\n")) + false + } catch { + CellosaurusError(_) => true + } + assert_true(raised) +} + +///| +test "cellosaurus_new_id_before_terminator_error" { + let text = "ID First\nAC CVCL_0001\nID Second\nAC CVCL_0002\n//\n" + let raised = try { + ignore(@src.cellosaurus_parse(text)) + false + } catch { + CellosaurusError(_) => true + } + assert_true(raised) +} + +///| +test "cellosaurus_invalid_cross_reference_error" { + let raised = try { + ignore(@src.cellosaurus_parse("ID Bad\nDR missing delimiter\n//\n")) + false + } catch { + CellosaurusError(_) => true + } + assert_true(raised) +} + +///| +test "cellosaurus_header_unknown_and_crlf" { + let text = "Cellosaurus release header\r\n" + + "__________\r\n" + + "//\r\n" + + "ID CRLF example\r\n" + + "AC CVCL_TEST\r\n" + + "ZZ ignored extension\r\n" + + "OX NCBI_TaxID=10090; ! Mus musculus (Mouse)\r\n" + + "//\r\n" + let records = @src.cellosaurus_parse(text) + assert_eq(records.length(), 1) + assert_eq(records[0].identifier, "CRLF example") + assert_eq(records[0].accession, "CVCL_TEST") + assert_true(records[0].has_species("Mus musculus")) +} + +///| +test "cellosaurus_repeated_single_value_fields_concatenate" { + let text = "ID Split\n" + + "AC CVCL_\n" + + "AC 1234\n" + + "SY First; \n" + + "SY Second\n" + + "//\n" + let record = @src.cellosaurus_parse(text)[0] + assert_eq(record.accession, "CVCL_1234") + assert_eq(record.synonyms, "First;Second") +} + +///| +test "cellosaurus_serialization_round_trip" { + let original = @src.cellosaurus_parse(@src.cellosaurus_sample_text())[0] + let serialized = original.to_string() + let reparsed = @src.cellosaurus_parse(serialized) + assert_eq(reparsed.length(), 1) + assert_true(reparsed[0] == original) +} + +///| +test "cellosaurus_empty_constructor_repr" { + let record = @src.CellosaurusRecord::new(identifier="") + assert_eq(record.repr(), "CellosaurusRecord ( )") + assert_eq(record.to_string(), "//\n") +} diff --git a/test/moonbit/chain_liftover_test.mbt b/test/moonbit/chain_liftover_test.mbt index 8dad6874..4c0939a2 100644 --- a/test/moonbit/chain_liftover_test.mbt +++ b/test/moonbit/chain_liftover_test.mbt @@ -154,7 +154,9 @@ test "cl_parse_chain_single_nucleotide_block" { ///| test "cl_parse_chain_id_preserved" { - let cf = @src.cl_parse_chain("chain 100 chr1 1000 + 100 200 chr2 500 + 50 150 99\n 100 0 0") + let cf = @src.cl_parse_chain( + "chain 100 chr1 1000 + 100 200 chr2 500 + 50 150 99\n 100 0 0", + ) assert_eq(cf.alignments[0].id, 99) } @@ -351,7 +353,9 @@ test "cl_get_chain_summary_basic" { ///| test "cl_get_chain_summary_empty_blocks" { - let cf = @src.cl_parse_chain("chain 5000 chr1 249250621 + 1 500 chr2 10000 + 1 500 43") + let cf = @src.cl_parse_chain( + "chain 5000 chr1 249250621 + 1 500 chr2 10000 + 1 500 43", + ) let summary = @src.cl_get_chain_summary(cf.alignments[0]) assert_eq(summary["total_block_size"], 0) assert_eq(summary["num_blocks"], 0) @@ -406,7 +410,9 @@ test "cl_chain_to_string_format" { ///| test "cl_chain_to_string_preserves_id" { - let cf = @src.cl_parse_chain("chain 7777 chr1 1000 + 100 200 chr2 500 + 50 150 99\n 100 0 0") + let cf = @src.cl_parse_chain( + "chain 7777 chr1 1000 + 100 200 chr2 500 + 50 150 99\n 100 0 0", + ) let s = @src.cl_chain_to_string(cf.alignments[0]) assert_true(s.contains("chain 7777")) assert_true(s.contains("99")) @@ -530,7 +536,9 @@ test "cl_chain_to_string_reverse_strand" { ///| test "cl_parse_chain_large_id" { - let cf = @src.cl_parse_chain("chain 100 chr1 1000 + 100 200 chr2 500 + 50 150 99999\n 100 0 0") + let cf = @src.cl_parse_chain( + "chain 100 chr1 1000 + 100 200 chr2 500 + 50 150 99999\n 100 0 0", + ) assert_eq(cf.alignments[0].id, 99999) } @@ -564,4 +572,4 @@ test "cl_find_chain_for_pos_at_boundary" { assert_true(r2 is None) let r3 = @src.cl_find_chain_for_pos(cf, "chr1", 6000) assert_true(r3 is Some(_)) -} \ No newline at end of file +} diff --git a/test/moonbit/checksum_test.mbt b/test/moonbit/checksum_test.mbt index 23af903a..df999342 100644 --- a/test/moonbit/checksum_test.mbt +++ b/test/moonbit/checksum_test.mbt @@ -50,4 +50,4 @@ test "verify_checksum_seguid" { let result = @src.checksum_seguid(seq) let verified = @src.verify_checksum(seq, result.checksum, "seguid") assert_true(verified) -} \ No newline at end of file +} diff --git a/test/moonbit/chromosome_visualization_test.mbt b/test/moonbit/chromosome_visualization_test.mbt index bdc64f5e..8c8197c2 100644 --- a/test/moonbit/chromosome_visualization_test.mbt +++ b/test/moonbit/chromosome_visualization_test.mbt @@ -20,11 +20,7 @@ test "chr_feature_type_values" { ///| test "chr_feature_creation" { - let feature = @src.ChrFeature::new( - name="Gene1", - start=100.0, - end=200.0, - ) + let feature = @src.ChrFeature::new(name="Gene1", start=100.0, end=200.0) assert_true((feature.start - 100.0).abs() < 1.0e-10) assert_true((feature.end - 200.0).abs() < 1.0e-10) assert_eq(feature.label, "Gene1") @@ -32,12 +28,8 @@ test "chr_feature_creation" { ///| test "chr_region_creation" { - let region = @src.ChrRegion::new( - name="Region1", - start=0.0, - end=1000.0, - ) - assert_true((region.start).abs() < 1.0e-10) + let region = @src.ChrRegion::new(name="Region1", start=0.0, end=1000.0) + assert_true(region.start.abs() < 1.0e-10) assert_true((region.end - 1000.0).abs() < 1.0e-10) assert_eq(region.label, "Region1") } @@ -57,50 +49,30 @@ test "chr_chromosome_creation" { ///| test "chr_chromosome_add_feature" { - let chr = @src.Chromosome::new( - name="chr1", - length=5000.0, - ) - let feature = @src.ChrFeature::new( - name="Gene1", - start=100.0, - end=200.0, - ) + let chr = @src.Chromosome::new(name="chr1", length=5000.0) + let feature = @src.ChrFeature::new(name="Gene1", start=100.0, end=200.0) let chr_with_feature = chr.add_feature(feature) assert_eq(chr_with_feature.features.length(), 1) } ///| test "chr_chromosome_add_region" { - let chr = @src.Chromosome::new( - name="chr1", - length=5000.0, - ) - let region = @src.ChrRegion::new( - name="Region1", - start=300.0, - end=500.0, - ) + let chr = @src.Chromosome::new(name="chr1", length=5000.0) + let region = @src.ChrRegion::new(name="Region1", start=300.0, end=500.0) let chr_with_region = chr.add_region(region) assert_eq(chr_with_region.regions.length(), 1) } ///| test "chr_chromosome_add_band" { - let chr = @src.Chromosome::new( - name="chr1", - length=5000.0, - ) + let chr = @src.Chromosome::new(name="chr1", length=5000.0) let chr_with_band = chr.add_band("Band1", 100.0, 200.0, "gpos50") assert_eq(chr_with_band.bands.length(), 1) } ///| test "chr_features_by_type" { - let chr = @src.Chromosome::new( - name="chr1", - length=5000.0, - ) + let chr = @src.Chromosome::new(name="chr1", length=5000.0) let exon1 = @src.ChrFeature::new( name="E1", start=100.0, @@ -128,20 +100,9 @@ test "chr_features_by_type" { ///| test "chr_features_in_region" { - let chr = @src.Chromosome::new( - name="chr1", - length=5000.0, - ) - let f1 = @src.ChrFeature::new( - name="F1", - start=100.0, - end=200.0, - ) - let f2 = @src.ChrFeature::new( - name="F2", - start=300.0, - end=400.0, - ) + let chr = @src.Chromosome::new(name="chr1", length=5000.0) + let f1 = @src.ChrFeature::new(name="F1", start=100.0, end=200.0) + let f2 = @src.ChrFeature::new(name="F2", start=300.0, end=400.0) let chr1 = chr.add_feature(f1).add_feature(f2) let in_region = chr1.features_in_region(150.0, 350.0) assert_eq(in_region.length(), 2) @@ -168,10 +129,7 @@ test "chr_diagram_creation" { length=5000.0, centromere_pos=2000.0, ) - let diagram = @src.ChrDiagram::new( - chromosomes=[chr], - title="Test Diagram", - ) + let diagram = @src.ChrDiagram::new(chromosomes=[chr], title="Test Diagram") assert_eq(diagram.chromosomes.length(), 1) assert_eq(diagram.title, "Test Diagram") } @@ -183,10 +141,7 @@ test "chr_to_svg_returns_string" { length=1000.0, centromere_pos=400.0, ) - let diagram = @src.ChrDiagram::new( - chromosomes=[chr], - title="Chromosome 1", - ) + let diagram = @src.ChrDiagram::new(chromosomes=[chr], title="Chromosome 1") let svg = diagram.to_svg() assert_true(svg.length() > 0) assert_true(svg.contains(" 0) assert_true(svg.contains(" 0) -} \ No newline at end of file +} diff --git a/test/moonbit/cibersort_test.mbt b/test/moonbit/cibersort_test.mbt index 9ac3372b..1168b02f 100644 --- a/test/moonbit/cibersort_test.mbt +++ b/test/moonbit/cibersort_test.mbt @@ -20,18 +20,19 @@ test "cib_signature_matrix_creation" { assert_eq(sig.matrix()[2][1], 1.0) } +///| test "cib_mixture_matrix_creation" { - let mix = @src.CibMixtureMatrix::new( - ["G1", "G2"], - ["S1", "S2"], - [[1.0, 2.0], [3.0, 4.0]], - ) + let mix = @src.CibMixtureMatrix::new(["G1", "G2"], ["S1", "S2"], [ + [1.0, 2.0], + [3.0, 4.0], + ]) assert_eq(mix.gene_names().length(), 2) assert_eq(mix.sample_names().length(), 2) assert_eq(mix.sample_names()[0], "S1") assert_eq(mix.matrix()[1][0], 3.0) } +///| test "cib_result_accessors" { // Use cib_run to get a real result and test accessors let sig = @src.cib_default_signature() @@ -49,6 +50,7 @@ test "cib_result_accessors" { // Built-in signature and marker genes // --------------------------------------------------------------------------- +///| test "cib_cell_type_names_count" { let names = @src.cib_cell_type_names() assert_eq(names.length(), 10) @@ -56,6 +58,7 @@ test "cib_cell_type_names_count" { assert_eq(names[9], "Mast cells") } +///| test "cib_marker_genes_count" { let genes = @src.cib_marker_genes() // 10 cell types × 4 markers each = 40 genes @@ -64,6 +67,7 @@ test "cib_marker_genes_count" { assert_eq(genes[4], "CD8A") // T cell CD8 marker } +///| test "cib_default_signature_dimensions" { let sig = @src.cib_default_signature() assert_eq(sig.gene_names().length(), 40) @@ -78,6 +82,7 @@ test "cib_default_signature_dimensions" { // NNLS deconvolution correctness // --------------------------------------------------------------------------- +///| test "cib_run_single_cell_type_recovery" { // Build a simple mixture: pure B cells (column 0 of signature) let sig = @src.cib_default_signature() @@ -100,6 +105,7 @@ test "cib_run_single_cell_type_recovery" { assert_true(r0.pearson_r() > 0.9) } +///| test "cib_run_mixed_composition_recovery" { // Use the built-in sample mixture generator let sig = @src.cib_default_signature() @@ -120,6 +126,7 @@ test "cib_run_mixed_composition_recovery" { } } +///| test "cib_run_sample1_dominant_t_cell_cd8" { // Sample1 was constructed as 60% T cells CD8 + 30% B cells + 10% NK cells let sig = @src.cib_default_signature() @@ -134,6 +141,7 @@ test "cib_run_sample1_dominant_t_cell_cd8" { assert_true(t_cd8 > nk) } +///| test "cib_run_sample2_dominant_monocytes" { // Sample2: 50% Monocytes + 30% Macrophages M1 + 20% Macrophages M2 let sig = @src.cib_default_signature() @@ -158,6 +166,7 @@ test "cib_run_sample2_dominant_monocytes" { ) } +///| test "cib_run_pearson_high_for_clean_data" { // For synthetic mixtures constructed from the signature, fit should be good let sig = @src.cib_default_signature() @@ -168,6 +177,7 @@ test "cib_run_pearson_high_for_clean_data" { } } +///| test "cib_run_handles_partial_gene_overlap" { // Mixture has only a subset of signature genes let sig = @src.cib_default_signature() @@ -186,14 +196,14 @@ test "cib_run_handles_partial_gene_overlap" { assert_true(r0.get_fraction("B cells") > 0.3) } +///| test "cib_run_empty_mixture_genes" { // Mixture has no genes in common with the signature -> all zeros let sig = @src.cib_default_signature() - let mix = @src.CibMixtureMatrix::new( - ["NONEXIST1", "NONEXIST2"], - ["Empty"], - [[1.0], [2.0]], - ) + let mix = @src.CibMixtureMatrix::new(["NONEXIST1", "NONEXIST2"], ["Empty"], [ + [1.0], + [2.0], + ]) let result = @src.cib_run(sig, mix) assert_eq(result.sample_names().length(), 1) let r0 = result.results()[0] @@ -206,6 +216,7 @@ test "cib_run_empty_mixture_genes" { // to_string formatting // --------------------------------------------------------------------------- +///| test "cib_deconvolution_to_string" { let sig = @src.cib_default_signature() let mix = @src.cib_sample_mixture() @@ -220,6 +231,7 @@ test "cib_deconvolution_to_string" { // NNLS edge cases // --------------------------------------------------------------------------- +///| test "cib_get_fraction_unknown_type" { // Use cib_run to get a real result and test get_fraction let sig = @src.cib_default_signature() @@ -233,6 +245,7 @@ test "cib_get_fraction_unknown_type" { assert_true(b_frac >= 0.0) } +///| test "cib_run_with_custom_tolerance" { let sig = @src.cib_default_signature() let mix = @src.cib_sample_mixture() diff --git a/test/moonbit/circ_seq_test.mbt b/test/moonbit/circ_seq_test.mbt index 8453b306..0bd88584 100644 --- a/test/moonbit/circ_seq_test.mbt +++ b/test/moonbit/circ_seq_test.mbt @@ -24,7 +24,7 @@ test "CircSeq::circ_gc_content" { test "CircSeq::circ_gc_content AT rich" { let circ = @src.CircSeq::new(sequence="ATATATAT") let gc = circ.circ_gc_content() - assert_true((gc).abs() < 1.0e-10) + assert_true(gc.abs() < 1.0e-10) } ///| diff --git a/test/moonbit/cluster_experiment_test.mbt b/test/moonbit/cluster_experiment_test.mbt index 7b2f652c..d4a9e983 100644 --- a/test/moonbit/cluster_experiment_test.mbt +++ b/test/moonbit/cluster_experiment_test.mbt @@ -128,8 +128,8 @@ test "ce_center_columns_basic" { // col 1: mean = 15 -> centered [-5, 5] let m = [[1.0, 10.0], [3.0, 20.0]] let c = @src.ce_center_columns(m) - assert_true((c[0][0] - (-1.0)).abs() < 0.001) - assert_true((c[0][1] - (-5.0)).abs() < 0.001) + assert_true((c[0][0] - -1.0).abs() < 0.001) + assert_true((c[0][1] - -5.0).abs() < 0.001) assert_true((c[1][0] - 1.0).abs() < 0.001) assert_true((c[1][1] - 5.0).abs() < 0.001) } @@ -376,11 +376,7 @@ test "ce_hclust_cut_k2" { ///| test "ce_hclust_cut_k_equals_n" { - let dist = [ - [0.0, 0.5, 0.5], - [0.5, 0.0, 0.5], - [0.5, 0.5, 0.0], - ] + let dist = [[0.0, 0.5, 0.5], [0.5, 0.0, 0.5], [0.5, 0.5, 0.0]] let labels = @src.ce_hclust_cut(dist, 3) assert_eq(labels.length(), 3) // each sample in its own cluster @@ -527,12 +523,7 @@ test "ce_sequential_cluster_isolates_small" { test "ce_rsec_pipeline" { // 4 samples, 2 clear clusters; small enough that sequential splits are // not triggered (each consensus cluster has < 4 points). - let data = [ - [0.0, 0.0], - [0.1, 0.0], - [10.0, 10.0], - [10.1, 10.0], - ] + let data = [[0.0, 0.0], [0.1, 0.0], [10.0, 10.0], [10.1, 10.0]] let params = @src.ce_default_params() let result = @src.ce_rsec(data, params) assert_eq(result.labels.length(), 4) @@ -719,12 +710,7 @@ test "ce_default_params" { ///| test "ce_default_params_used_in_rsec" { // Smoke test: default params produce a valid ClusterExperiment - let data = [ - [0.0, 0.0], - [0.1, 0.1], - [5.0, 5.0], - [5.1, 5.1], - ] + let data = [[0.0, 0.0], [0.1, 0.1], [5.0, 5.0], [5.1, 5.1]] let params = @src.ce_default_params() let result = @src.ce_rsec(data, params) assert_true(result.labels.length() == 4) diff --git a/test/moonbit/cnvkit_test.mbt b/test/moonbit/cnvkit_test.mbt index 0df6e9db..3b09e281 100644 --- a/test/moonbit/cnvkit_test.mbt +++ b/test/moonbit/cnvkit_test.mbt @@ -1,32 +1,22 @@ ///| test "cnvkit_create_probe" { let probe = @src.CNVProbe::new( - "probe_001", - "chr1", - 1000000, - 1001000, - -0.5, - 1.0 + "probe_001", "chr1", 1000000, 1001000, -0.5, 1.0, ) - + assert_eq(probe.probe_id, "probe_001") assert_eq(probe.chromosome, "chr1") assert_eq(probe.start, 1000000) assert_eq(probe.end, 1001000) - assert_true((probe.log2_ratio - (-0.5)).abs() < 0.001) + assert_true((probe.log2_ratio - -0.5).abs() < 0.001) } ///| test "cnvkit_create_segment" { let segment = @src.CNVSegment::new( - "chr1", - 1000000, - 2000000, - 50, - -0.8, - "deletion" + "chr1", 1000000, 2000000, 50, -0.8, "deletion", ) - + assert_eq(segment.chromosome, "chr1") assert_eq(segment.start, 1000000) assert_eq(segment.end, 2000000) @@ -37,20 +27,15 @@ test "cnvkit_create_segment" { ///| test "cnvkit_dataset_operations" { let mut dataset = @src.CNVDataset::new() - + assert_eq(dataset.count_probes(), 0) - + let probe = @src.CNVProbe::new( - "probe_001", - "chr1", - 1000000, - 1001000, - -0.5, - 1.0 + "probe_001", "chr1", 1000000, 1001000, -0.5, 1.0, ) - + dataset = dataset.add_probe(probe) - + assert_eq(dataset.count_probes(), 1) assert_eq(dataset.chromosomes.length(), 1) assert_eq(dataset.chromosomes[0], "chr1") @@ -59,15 +44,19 @@ test "cnvkit_dataset_operations" { ///| test "cnvkit_filter_chromosome" { let mut dataset = @src.CNVDataset::new() - - let probe1 = @src.CNVProbe::new("probe_001", "chr1", 1000000, 1001000, -0.5, 1.0) - let probe2 = @src.CNVProbe::new("probe_002", "chr2", 2000000, 2001000, 0.5, 1.0) - + + let probe1 = @src.CNVProbe::new( + "probe_001", "chr1", 1000000, 1001000, -0.5, 1.0, + ) + let probe2 = @src.CNVProbe::new( + "probe_002", "chr2", 2000000, 2001000, 0.5, 1.0, + ) + dataset = dataset.add_probe(probe1) dataset = dataset.add_probe(probe2) - + let filtered = dataset.filter_chromosome("chr1") - + assert_eq(filtered.count_probes(), 1) assert_eq(filtered.chromosomes.length(), 1) } @@ -75,24 +64,26 @@ test "cnvkit_filter_chromosome" { ///| test "cnvkit_cbs_segmentation" { let probes : Array[@src.CNVProbe] = Array::new() - + // Create probes with a clear change point let mut i = 0 while i < 20 { let ratio = if i < 10 { -0.8 } else { 0.1 } - probes.push(@src.CNVProbe::new( - "probe_" + i.to_string(), - "chr1", - 1000000 + i * 1000, - 1000000 + i * 1000 + 100, - ratio, - 1.0 - )) + probes.push( + @src.CNVProbe::new( + "probe_" + i.to_string(), + "chr1", + 1000000 + i * 1000, + 1000000 + i * 1000 + 100, + ratio, + 1.0, + ), + ) i = i + 1 } - + let result = @src.cbs_segment(probes, 0.05) - + assert_true(result.n_segments > 0) assert_true(result.breakpoints.length() > 0) } @@ -100,24 +91,26 @@ test "cnvkit_cbs_segmentation" { ///| test "cnvkit_smooth_log2_ratios" { let probes : Array[@src.CNVProbe] = Array::new() - + let mut i = 0 while i < 10 { - probes.push(@src.CNVProbe::new( - "probe_" + i.to_string(), - "chr1", - 1000000 + i * 1000, - 1000000 + i * 1000 + 100, - (i.to_double() - 5.0) * 0.1, - 1.0 - )) + probes.push( + @src.CNVProbe::new( + "probe_" + i.to_string(), + "chr1", + 1000000 + i * 1000, + 1000000 + i * 1000 + 100, + (i.to_double() - 5.0) * 0.1, + 1.0, + ), + ) i = i + 1 } - + let smoothed = @src.smooth_log2_ratios(probes, 3) - + assert_eq(smoothed.length(), 10) - + // Check that smoothed values are reasonable let val = smoothed[0].log2_ratio assert_true(val.abs() < 10.0) @@ -135,11 +128,11 @@ test "cnvkit_detect_breakpoints" { let segments : Array[@src.CNVSegment] = [ @src.CNVSegment::new("chr1", 1000000, 2000000, 50, -0.8, "deletion"), @src.CNVSegment::new("chr1", 2000000, 3000000, 50, 0.1, "neutral"), - @src.CNVSegment::new("chr1", 3000000, 4000000, 50, 0.9, "amplification") + @src.CNVSegment::new("chr1", 3000000, 4000000, 50, 0.9, "amplification"), ] - + let breakpoints = @src.detect_breakpoints(segments, 0.3) - + assert_true(breakpoints.length() > 0) } @@ -148,11 +141,11 @@ test "cnvkit_call_copy_numbers" { let segments : Array[@src.CNVSegment] = [ @src.CNVSegment::new("chr1", 1000000, 2000000, 50, -0.8, "deletion"), @src.CNVSegment::new("chr1", 2000000, 3000000, 50, 0.0, "neutral"), - @src.CNVSegment::new("chr1", 3000000, 4000000, 50, 0.8, "amplification") + @src.CNVSegment::new("chr1", 3000000, 4000000, 50, 0.8, "amplification"), ] - + let calls = @src.call_copy_numbers(segments, 2) - + assert_eq(calls.length(), 3) assert_eq(calls[0].state, "deletion") assert_eq(calls[1].state, "neutral") @@ -163,7 +156,7 @@ test "cnvkit_call_copy_numbers" { test "cnvkit_summarize_dataset" { let dataset = @src.create_example_cnv_dataset() let summary = @src.summarize_cnv(dataset) - + assert_true(summary.total_probes > 0) assert_true(summary.total_segments >= 0) } @@ -171,7 +164,7 @@ test "cnvkit_summarize_dataset" { ///| test "cnvkit_create_example" { let dataset = @src.create_example_cnv_dataset() - + assert_true(dataset.count_probes() > 0) assert_true(dataset.chromosomes.length() >= 2) } @@ -180,7 +173,7 @@ test "cnvkit_create_example" { test "cnvkit_summarize_string" { let dataset = @src.create_example_cnv_dataset() let summary_str = @src.cbs_summarize_dataset(dataset) - + assert_true(summary_str.contains("CNV Dataset Summary")) assert_true(summary_str.contains("Total probes:")) } @@ -189,7 +182,7 @@ test "cnvkit_summarize_string" { test "cnvkit_empty_segmentation" { let empty_probes : Array[@src.CNVProbe] = Array::new() let result = @src.cbs_segment(empty_probes, 0.05) - + assert_eq(result.n_segments, 0) assert_eq(result.breakpoints.length(), 0) } @@ -198,9 +191,9 @@ test "cnvkit_empty_segmentation" { test "cnvkit_single_probe" { let probe = @src.CNVProbe::new("p1", "chr1", 100, 200, 0.5, 1.0) let probes : Array[@src.CNVProbe] = [probe] - + let result = @src.cbs_segment(probes, 0.05) - + assert_eq(result.n_segments, 1) } @@ -208,6 +201,6 @@ test "cnvkit_single_probe" { test "cnvkit_smooth_edge_cases" { let empty_probes : Array[@src.CNVProbe] = Array::new() let smoothed = @src.smooth_log2_ratios(empty_probes, 3) - + assert_eq(smoothed.length(), 0) } diff --git a/test/moonbit/codon_advanced_test.mbt b/test/moonbit/codon_advanced_test.mbt index 6a20cc4d..d04efc7b 100644 --- a/test/moonbit/codon_advanced_test.mbt +++ b/test/moonbit/codon_advanced_test.mbt @@ -179,18 +179,29 @@ test "codon_advanced_calculate_enc" { ///| test "codon_advanced_calculate_enc_low_bias" { let counts = Map([], capacity=24) - counts.set("GCA", 1); counts.set("GCC", 1) - counts.set("GCG", 1); counts.set("GCT", 1) - counts.set("AAA", 1); counts.set("AAG", 1) - counts.set("GAA", 1); counts.set("GAG", 1) - counts.set("CTT", 1); counts.set("CTC", 1) - counts.set("CTA", 1); counts.set("CTG", 1) - counts.set("TTA", 1); counts.set("TTG", 1) - counts.set("ATT", 1); counts.set("ATC", 1) + counts.set("GCA", 1) + counts.set("GCC", 1) + counts.set("GCG", 1) + counts.set("GCT", 1) + counts.set("AAA", 1) + counts.set("AAG", 1) + counts.set("GAA", 1) + counts.set("GAG", 1) + counts.set("CTT", 1) + counts.set("CTC", 1) + counts.set("CTA", 1) + counts.set("CTG", 1) + counts.set("TTA", 1) + counts.set("TTG", 1) + counts.set("ATT", 1) + counts.set("ATC", 1) counts.set("ATA", 1) - counts.set("CGT", 1); counts.set("CGC", 1) - counts.set("CGA", 1); counts.set("CGG", 1) - counts.set("AGA", 1); counts.set("AGG", 1) + counts.set("CGT", 1) + counts.set("CGC", 1) + counts.set("CGA", 1) + counts.set("CGG", 1) + counts.set("AGA", 1) + counts.set("AGG", 1) let usage = @src.CodonUsageTable::new(counts) let enc_val = @src.calculate_enc(usage) @@ -395,4 +406,4 @@ test "codon_advanced_all_stop_codons" { assert_eq(@src.get_amino_acid("TAA"), "*") assert_eq(@src.get_amino_acid("TAG"), "*") assert_eq(@src.get_amino_acid("TGA"), "*") -} \ No newline at end of file +} diff --git a/test/moonbit/codon_align_advanced_test.mbt b/test/moonbit/codon_align_advanced_test.mbt new file mode 100644 index 00000000..450063c8 --- /dev/null +++ b/test/moonbit/codon_align_advanced_test.mbt @@ -0,0 +1,481 @@ +// Tests for the advanced codon alignment selection-pressure statistics module. + +///| +fn caa_test_close(actual : Double, expected : Double, tolerance : Double) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "CodonAlign advanced value " + + actual.to_string() + + " differs from " + + expected.to_string() + + " (tolerance " + + tolerance.to_string() + + ")", + ) + } +} + +///| +// Sequences with many nonsynonymous changes → positive selection pressure. +// 9 codons (27 nt): Seq1 M-A-L-K-W-Q-Q-P-W, Seq2 M-G-L-E-R-H-P-S-M (7 nonsyn, 2 syn) +fn caa_test_positive_selection_pair() -> (String, String) { + ( + "ATGGCCCTGAAATGGCAGCAGCCATGG", + "ATGGGTCTGGAGAGACACCCCAGCATG", + ) +} + +///| +// Sequences with mostly synonymous changes → purifying selection pressure. +// 9 codons (27 nt): Seq1 M-A-L-K-W-Q-Q-P-W, Seq2 has only synonymous codon switches. +fn caa_test_purifying_selection_pair() -> (String, String) { + ( + "ATGGCCCTGAAATGGCAGCAGCCATGG", + "ATGGCGCTCAAATGGCAGCAGCGATGG", + ) +} + +///| +// Identical sequences → no selection, dN = dS = 0. +fn caa_test_identical_pair() -> (String, String) { + ("ATGGCCCTGAAATGGCAGCAGCCATGG", "ATGGCCCTGAAATGGCAGCAGCCATGG") +} + +// =========================================================================== +// Z-test for selection +// =========================================================================== + +///| +test "CodonAlign advanced Z-test positive selection detects dN > dS" { + let (s1, s2) = caa_test_positive_selection_pair() + let result = @src.codon_test_selection(s1, s2, test_type="positive") catch { + _ => abort("Z-test should run for positive selection pair") + } + assert_eq(result.test_type(), "positive") + assert_true(result.dn() >= 0.0) + assert_true(result.ds() >= 0.0) + // With many nonsynonymous changes and no synonymous changes, dN should + // be greater than dS, and the Z-score should be positive. + assert_true(result.z_score() > 0.0 || result.variance() == 0.0) + assert_true(result.p_value() >= 0.0 && result.p_value() <= 1.0) +} + +///| +test "CodonAlign advanced Z-test purifying selection detects dN < dS" { + let (s1, s2) = caa_test_purifying_selection_pair() + let result = @src.codon_test_selection(s1, s2, test_type="purifying") catch { + _ => abort("Z-test should run for purifying selection pair") + } + assert_eq(result.test_type(), "purifying") + // With mostly synonymous changes, dS should be >= dN. + assert_true(result.ds() >= result.dn() || result.variance() == 0.0) + assert_true(result.p_value() >= 0.0 && result.p_value() <= 1.0) +} + +///| +test "CodonAlign advanced Z-test neutrality is two-tailed" { + let (s1, s2) = caa_test_positive_selection_pair() + let result = @src.codon_test_selection(s1, s2, test_type="neutrality") catch { + _ => abort("Z-test should run for neutrality test") + } + assert_eq(result.test_type(), "neutrality") + assert_true(result.p_value() >= 0.0 && result.p_value() <= 1.0) +} + +///| +test "CodonAlign advanced Z-test identical sequences has zero variance" { + let (s1, s2) = caa_test_identical_pair() + let result = @src.codon_test_selection(s1, s2) catch { + _ => abort("Z-test should run for identical sequences") + } + caa_test_close(result.dn(), 0.0, 1.0e-12) + caa_test_close(result.ds(), 0.0, 1.0e-12) + caa_test_close(result.variance(), 0.0, 1.0e-12) + caa_test_close(result.z_score(), 0.0, 1.0e-12) + caa_test_close(result.p_value(), 1.0, 1.0e-12) + assert_true(result.conclusion().contains("inconclusive")) +} + +///| +test "CodonAlign advanced Z-test rejects short sequences" { + let failed = try { + ignore(@src.codon_test_selection("AT", "AT")) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced Z-test rejects mismatched lengths" { + let failed = try { + ignore(@src.codon_test_selection("ATGGCC", "ATGGCCAAG")) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced Z-test rejects non-codon-aligned lengths" { + let failed = try { + ignore(@src.codon_test_selection("ATGGCCAT", "ATGGCCAT")) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Fisher's exact test for neutrality +// =========================================================================== + +///| +test "CodonAlign advanced Fisher test two-sided runs and returns valid p-value" { + let (s1, s2) = caa_test_positive_selection_pair() + let result = @src.codon_test_neutrality(s1, s2) catch { + _ => abort("Fisher exact test should run") + } + assert_eq(result.test_type(), "two-sided") + assert_true(result.p_value() >= 0.0 && result.p_value() <= 1.0) + assert_true(result.odds_ratio() >= 0.0) + assert_true(result.n_diff() >= 0) + assert_true(result.s_diff() >= 0) + assert_true(result.n_sites() >= 0.0) + assert_true(result.s_sites() >= 0.0) +} + +///| +test "CodonAlign advanced Fisher test greater detects excess nonsynonymous" { + let (s1, s2) = caa_test_positive_selection_pair() + let result = @src.codon_test_neutrality(s1, s2, test_type="greater") catch { + _ => abort("Fisher exact test greater should run") + } + assert_eq(result.test_type(), "greater") + // With more nonsynonymous differences, the one-sided test for greater + // should give a smaller p-value than two-sided. + let two_sided = @src.codon_test_neutrality(s1, s2) catch { + _ => abort("Fisher exact test two-sided should run") + } + assert_true(result.p_value() <= two_sided.p_value() + 1.0e-12) +} + +///| +test "CodonAlign advanced Fisher test identical sequences has no differences" { + let (s1, s2) = caa_test_identical_pair() + let result = @src.codon_test_neutrality(s1, s2) catch { + _ => abort("Fisher exact test should run for identical sequences") + } + assert_eq(result.n_diff(), 0) + assert_eq(result.s_diff(), 0) + // With no differences, p-value should be 1 (can't reject neutrality). + caa_test_close(result.p_value(), 1.0, 1.0e-10) + assert_true(result.conclusion().contains("not significant")) +} + +///| +test "CodonAlign advanced Fisher test purifying pair has more syn differences" { + let (s1, s2) = caa_test_purifying_selection_pair() + let result = @src.codon_test_neutrality(s1, s2, test_type="less") catch { + _ => abort("Fisher exact test less should run") + } + assert_eq(result.test_type(), "less") + // With mostly synonymous changes, s_diff >= n_diff. + assert_true(result.s_diff() >= result.n_diff()) +} + +// =========================================================================== +// Codon alignment builder +// =========================================================================== + +///| +test "CodonAlign advanced builder constructs codon alignment with gaps" { + let (protein_aln, coding_seqs) = @src.create_demo_protein_alignment() + let alignment = @src.build_codon_alignment(protein_aln, coding_seqs) catch { + _ => abort("codon alignment builder should succeed") + } + assert_eq(alignment.n_codons(), 5) + let seqs = alignment.sequences() + // Each sequence should have 5 codon columns × 3 nt = 15 characters. + for seq in seqs { + assert_eq(seq.length(), 15) + } + // Sequence 1 (M-A-K with gaps at pos 2,4): ATG(M) ---(gap) GCC(A) ---(gap) AAG(K) + assert_eq(seqs[0], "ATG---GCC---AAG") + // Sequence 2 (MTASK, no gaps): ATG(M) ACC(T) GCC(A) AGC(S) AAG(K) + assert_eq(seqs[1], "ATGACCGCCAGCAAG") +} + +///| +test "CodonAlign advanced builder with names" { + let protein_aln = ["MAK", "MAK"] + let coding_seqs = ["ATGGCCAAG", "ATGGCCAAG"] + let alignment = @src.build_codon_alignment( + protein_aln, + coding_seqs, + names=["gene1", "gene2"], + ) catch { + _ => abort("codon alignment builder should succeed") + } + let names = alignment.names() + assert_eq(names[0], "gene1") + assert_eq(names[1], "gene2") + assert_eq(alignment.n_codons(), 3) +} + +///| +test "CodonAlign advanced builder rejects translation mismatch" { + // Protein says M-A-K but coding sequence has T(ACC) instead of A(GCC). + let protein_aln = ["MAK"] + let coding_seqs = ["ATGACCAAG"] + let failed = try { + ignore(@src.build_codon_alignment(protein_aln, coding_seqs)) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced builder rejects exhausted coding sequence" { + // Protein alignment is longer than the coding sequence can support. + let protein_aln = ["MAKV"] + let coding_seqs = ["ATGGCC"] + let failed = try { + ignore(@src.build_codon_alignment(protein_aln, coding_seqs)) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced builder rejects leftover nucleotides" { + let protein_aln = ["MA"] + let coding_seqs = ["ATGGCCAAG"] + let failed = try { + ignore(@src.build_codon_alignment(protein_aln, coding_seqs)) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced builder rejects mismatched row counts" { + let failed = try { + ignore(@src.build_codon_alignment(["MAK"], ["ATGGCC", "ATGGCC"])) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced builder rejects unequal protein lengths" { + let failed = try { + ignore(@src.build_codon_alignment(["MAK", "MAKV"], ["ATGGCC", "ATGGCCAAG"])) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Sliding-window dN/dS +// =========================================================================== + +///| +test "CodonAlign advanced sliding window produces multiple windows" { + // 21-codon (63 nt) sequences. + let s1 = "ATGGCCCTGAAATGGCAGCAGCCATGGGCCCTGAAATGGCAGCAGCCATGGGCCCTGAAATGG" + let s2 = "ATGGGTCTGGAGAGACACCCCAGCATGGGGTCTGGAGAGACACCCCAGCATGGGGTCTGGAGA" + let result = @src.sliding_window_dnds(s1, s2, window_size=5, step_size=3) catch { + _ => abort("sliding window should run") + } + let windows = result.windows() + assert_true(windows.length() >= 2) + for w in windows { + assert_true(w.end_codon() > w.start_codon()) + assert_true(w.n_codons() >= 3) + assert_true(w.dn() >= 0.0) + assert_true(w.ds() >= 0.0) + } +} + +///| +test "CodonAlign advanced sliding window skips degenerate windows" { + // Identical sequences → all windows have dN = dS = 0. + let s1 = "ATGGCCCTGAAATGGCAGCAGCCATGGGCCCTGAAATGGCAGCAGCCATGGGCCCTGAAATGG" + let result = @src.sliding_window_dnds(s1, s1, window_size=5, step_size=3) catch { + _ => abort("sliding window should run for identical sequences") + } + let windows = result.windows() + for w in windows { + caa_test_close(w.dn(), 0.0, 1.0e-12) + caa_test_close(w.ds(), 0.0, 1.0e-12) + } +} + +///| +test "CodonAlign advanced sliding window rejects small window" { + let failed = try { + ignore(@src.sliding_window_dnds("ATGGCCATGGCC", "ATGGCCATGGCC", window_size=2)) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced sliding window rejects invalid step" { + let failed = try { + ignore(@src.sliding_window_dnds("ATGGCCATGGCC", "ATGGCCATGGCC", step_size=0)) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// BH-FDR +// =========================================================================== + +///| +test "CodonAlign advanced BH-FDR is monotonic non-decreasing in rank" { + let p_values = [0.001, 0.04, 0.03, 0.5, 0.01] + let fdr = @src.codon_align_advanced_bh_fdr(p_values) + assert_eq(fdr.length(), 5) + // Sort p-values and verify FDR is non-decreasing in sorted order. + let indexed : Array[(Double, Int)] = [] + for i in 0.. Int { + if a.0 < b.0 { + -1 + } else if a.0 > b.0 { + 1 + } else { + 0 + } + }) + let mut previous = 0.0 + for entry in indexed { + let adjusted = fdr[entry.1] + assert_true(adjusted >= previous - 1.0e-12) + previous = adjusted + } +} + +///| +test "CodonAlign advanced BH-FDR clamps p-values to [0,1]" { + let p_values = [-0.5, 0.5, 1.5] + let fdr = @src.codon_align_advanced_bh_fdr(p_values) + for v in fdr { + assert_true(v >= 0.0 && v <= 1.0) + } +} + +///| +test "CodonAlign advanced BH-FDR empty input returns empty" { + let fdr = @src.codon_align_advanced_bh_fdr([]) + assert_eq(fdr.length(), 0) +} + +///| +test "CodonAlign advanced BH-FDR single p-value equals itself" { + let fdr = @src.codon_align_advanced_bh_fdr([0.03]) + caa_test_close(fdr[0], 0.03, 1.0e-12) +} + +// =========================================================================== +// Pairwise Ka/Ks table +// =========================================================================== + +///| +test "CodonAlign advanced pairwise table computes all pairs with FDR" { + let seqs = @src.create_positive_selection_alignment() + let table = @src.pairwise_kaks_table( + seqs, + names=["g1", "g2", "g3", "g4", "g5", "g6", "g7", "g8", "g9", "g10"], + fdr_threshold=0.05, + ) catch { + _ => abort("pairwise Ka/Ks table should run") + } + // C(10, 2) = 45 pairs. + assert_eq(table.length(), 45) + for row in table { + assert_true(row.seq1_name().length() > 0) + assert_true(row.seq2_name().length() > 0) + assert_true(row.dn() >= 0.0) + assert_true(row.ds() >= 0.0) + assert_true(row.fdr() >= 0.0 && row.fdr() <= 1.0) + assert_true(row.p_value() >= 0.0 && row.p_value() <= 1.0) + } +} + +///| +test "CodonAlign advanced pairwise table rejects single sequence" { + let failed = try { + ignore(@src.pairwise_kaks_table(["ATGGCC"])) + false + } catch { + CodonAlignAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "CodonAlign advanced pairwise table uses default names when none provided" { + let table = @src.pairwise_kaks_table(["ATGGCCATG", "ATGGCCATG", "ATGGCCATG"]) catch { + _ => abort("pairwise Ka/Ks table should run") + } + // 3 sequences → 3 pairs. + assert_eq(table.length(), 3) + assert_eq(table[0].seq1_name(), "seq0") + assert_eq(table[0].seq2_name(), "seq1") +} + +// =========================================================================== +// Demo data +// =========================================================================== + +///| +test "CodonAlign advanced purifying selection demo has valid codons" { + let seqs = @src.create_purifying_selection_alignment() + assert_true(seqs.length() >= 2) + for seq in seqs { + assert_eq(seq.length() % 3, 0) + } +} + +///| +test "CodonAlign advanced positive selection demo has valid codons" { + let seqs = @src.create_positive_selection_alignment() + assert_true(seqs.length() >= 2) + for seq in seqs { + assert_eq(seq.length() % 3, 0) + } +} + +///| +test "CodonAlign advanced demo protein alignment is consistent" { + let (protein_aln, coding_seqs) = @src.create_demo_protein_alignment() + assert_eq(protein_aln.length(), coding_seqs.length()) + // All protein alignments should have the same length. + let aln_len = protein_aln[0].length() + for pa in protein_aln { + assert_eq(pa.length(), aln_len) + } +} diff --git a/test/moonbit/compass_test.mbt b/test/moonbit/compass_test.mbt index 7bcf69b9..f7206c55 100644 --- a/test/moonbit/compass_test.mbt +++ b/test/moonbit/compass_test.mbt @@ -7,22 +7,8 @@ test "compass_record_creation" { let r = @src.CompassRecord::new( - "query1.msa", - "template1.msa", - 5, - 4, - 120, - 110, - 185.5, - 3.5e-10, - 28.5, - 1, - 120, - 1, - 110, - "MALKSLVRLFG", - "MGVKSAVKT", - ": :: :::", + "query1.msa", "template1.msa", 5, 4, 120, 110, 185.5, 3.5e-10, 28.5, 1, 120, + 1, 110, "MALKSLVRLFG", "MGVKSAVKT", ": :: :::", ) assert_eq(r.query_name(), "query1.msa") assert_eq(r.template_name(), "template1.msa") @@ -42,10 +28,10 @@ test "compass_record_creation" { assert_eq(r.consensus_line(), ": :: :::") } +///| test "compass_record_accessors_individual" { let r = @src.CompassRecord::new( - "q", "t", 1, 2, 10, 20, 50.0, 0.001, 35.0, - 5, 50, 10, 60, "ACGT", "ACGT", ":::", + "q", "t", 1, 2, 10, 20, 50.0, 0.001, 35.0, 5, 50, 10, 60, "ACGT", "ACGT", ":::", ) assert_eq(r.query_name(), "q") assert_eq(r.template_name(), "t") @@ -65,13 +51,11 @@ test "compass_record_accessors_individual" { assert_eq(r.consensus_line(), ":::") } +///| test "compass_record_to_string" { let r = @src.CompassRecord::new( - "query1.msa", "template1.msa", - 5, 4, 120, 110, - 185.5, 3.5e-10, 28.5, - 1, 120, 1, 110, - "MALKSLVRLFG", "MGVKSAVKT", ": :: :::", + "query1.msa", "template1.msa", 5, 4, 120, 110, 185.5, 3.5e-10, 28.5, 1, 120, + 1, 110, "MALKSLVRLFG", "MGVKSAVKT", ": :: :::", ) let s = r.to_string() assert_true(s.contains("query1.msa")) @@ -83,56 +67,66 @@ test "compass_record_to_string" { // Parsing: sample data // --------------------------------------------------------------------------- +///| test "compass_parse_sample_data_count" { let text = @src.compass_sample_data() let records = @src.parse_compass(text) assert_eq(records.length(), 2) } +///| test "compass_parse_first_record_query_name" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].query_name(), "query1.msa") } +///| test "compass_parse_first_record_template_name" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].template_name(), "template1.msa") } +///| test "compass_parse_first_record_n_seqs" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].query_n_seqs(), 5) assert_eq(records[0].template_n_seqs(), 4) } +///| test "compass_parse_first_record_n_cols" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].query_n_cols(), 120) assert_eq(records[0].template_n_cols(), 110) } +///| test "compass_parse_first_record_sw_score" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].sw_score(), 185.5) } +///| test "compass_parse_first_record_e_value" { let records = @src.parse_compass(@src.compass_sample_data()) // Use tolerance comparison due to floating-point precision assert_true((records[0].e_value() - 3.5e-10).abs() < 1.0e-20) } +///| test "compass_parse_first_record_identity" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].percentage_identity(), 28.5) } +///| test "compass_parse_first_record_alignment" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].aligned_query(), "MALKSLVRLFG") assert_eq(records[0].aligned_template(), "MGVKSAVKT") } +///| test "compass_parse_first_record_positions" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].query_start(), 1) @@ -141,11 +135,13 @@ test "compass_parse_first_record_positions" { assert_eq(records[0].template_end(), 110) } +///| test "compass_parse_first_record_consensus" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[0].consensus_line(), ": :: :::") } +///| test "compass_parse_second_record" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(records[1].query_name(), "query2.msa") @@ -164,16 +160,19 @@ test "compass_parse_second_record" { // Version extraction // --------------------------------------------------------------------------- +///| test "compass_version_extraction" { let text = @src.compass_sample_data() let version = @src.compass_version(text) assert_eq(version, "2.4.2") } +///| test "compass_version_empty_input" { assert_eq(@src.compass_version(""), "") } +///| test "compass_version_no_version_line" { let text = "Some other content\nwithout version\n" assert_eq(@src.compass_version(text), "") @@ -183,6 +182,7 @@ test "compass_version_no_version_line" { // Filtering // --------------------------------------------------------------------------- +///| test "compass_filter_by_evalue" { let records = @src.parse_compass(@src.compass_sample_data()) // E-values: 3.5e-10 and 1.2e-05 @@ -191,18 +191,21 @@ test "compass_filter_by_evalue" { assert_eq(filtered[0].query_name(), "query1.msa") } +///| test "compass_filter_by_evalue_all_pass" { let records = @src.parse_compass(@src.compass_sample_data()) let filtered = @src.compass_filter_by_evalue(records, 1.0) assert_eq(filtered.length(), 2) } +///| test "compass_filter_by_evalue_none_pass" { let records = @src.parse_compass(@src.compass_sample_data()) let filtered = @src.compass_filter_by_evalue(records, 1.0e-20) assert_eq(filtered.length(), 0) } +///| test "compass_filter_by_identity" { let records = @src.parse_compass(@src.compass_sample_data()) // Identities: 28.5 and 15.3 @@ -211,6 +214,7 @@ test "compass_filter_by_identity" { assert_eq(filtered[0].query_name(), "query1.msa") } +///| test "compass_filter_by_identity_all_pass" { let records = @src.parse_compass(@src.compass_sample_data()) let filtered = @src.compass_filter_by_identity(records, 10.0) @@ -221,16 +225,17 @@ test "compass_filter_by_identity_all_pass" { // Alignment length // --------------------------------------------------------------------------- +///| test "compass_alignment_length" { let records = @src.parse_compass(@src.compass_sample_data()) assert_eq(@src.compass_alignment_length(records[0]), 11) assert_eq(@src.compass_alignment_length(records[1]), 10) } +///| test "compass_alignment_length_empty" { let r = @src.CompassRecord::new( - "", "", 0, 0, 0, 0, 0.0, 0.0, 0.0, - 0, 0, 0, 0, "", "", "", + "", "", 0, 0, 0, 0, 0.0, 0.0, 0.0, 0, 0, 0, 0, "", "", "", ) assert_eq(@src.compass_alignment_length(r), 0) } @@ -239,6 +244,7 @@ test "compass_alignment_length_empty" { // Summary // --------------------------------------------------------------------------- +///| test "compass_summary_content" { let records = @src.parse_compass(@src.compass_sample_data()) let summary = @src.compass_summary(records) @@ -250,6 +256,7 @@ test "compass_summary_content" { assert_true(summary.contains("28.5")) } +///| test "compass_summary_empty" { let records : Array[@src.CompassRecord] = [] let summary = @src.compass_summary(records) @@ -260,16 +267,19 @@ test "compass_summary_empty" { // Edge cases // --------------------------------------------------------------------------- +///| test "compass_parse_empty_input" { let records = @src.parse_compass("") assert_eq(records.length(), 0) } +///| test "compass_parse_whitespace_only" { let records = @src.parse_compass(" \n \n \n") assert_eq(records.length(), 0) } +///| test "compass_parse_single_record" { let text = "COMPASS version 2.4.2\n" + "Query alignment: single.msa\n" + @@ -296,10 +306,10 @@ test "compass_parse_single_record" { assert_eq(records[0].consensus_line(), ": : :") } +///| test "compass_parse_missing_fields" { // Record with only version and query name (missing other fields) - let text = "COMPASS version 1.0.0\n" + - "Query alignment: partial.msa\n" + let text = "COMPASS version 1.0.0\n" + "Query alignment: partial.msa\n" let records = @src.parse_compass(text) assert_eq(records.length(), 1) assert_eq(records[0].query_name(), "partial.msa") @@ -316,6 +326,7 @@ test "compass_parse_missing_fields" { assert_eq(records[0].consensus_line(), "") } +///| test "compass_parse_no_consensus_line" { // Alignment block without a consensus line between Query and Template let text = "COMPASS version 2.4.2\n" + @@ -337,6 +348,7 @@ test "compass_parse_no_consensus_line" { assert_eq(records[0].consensus_line(), "") } +///| test "compass_parse_three_records" { let mut text = "" text = text + "COMPASS version 2.4.2\n" @@ -360,6 +372,7 @@ test "compass_parse_three_records" { assert_eq(records[2].query_name(), "q3.msa") } +///| test "compass_parse_version_only" { let text = "COMPASS version 3.0.0\n" let records = @src.parse_compass(text) @@ -367,17 +380,20 @@ test "compass_parse_version_only" { assert_eq(records[0].query_name(), "") } +///| test "compass_sample_data_has_version" { let text = @src.compass_sample_data() assert_true(text.contains("COMPASS version 2.4.2")) } +///| test "compass_sample_data_has_evalue" { let text = @src.compass_sample_data() assert_true(text.contains("3.5e-10")) assert_true(text.contains("1.2e-05")) } +///| test "compass_filter_empty_records" { let records : Array[@src.CompassRecord] = [] let by_evalue = @src.compass_filter_by_evalue(records, 1.0) diff --git a/test/moonbit/compound_test.mbt b/test/moonbit/compound_test.mbt index 99123022..07ef7572 100644 --- a/test/moonbit/compound_test.mbt +++ b/test/moonbit/compound_test.mbt @@ -9,7 +9,9 @@ test "compound_new" { ///| test "compound_with_chemical" { - let c = @src.Compound::with_chemical("C0001", "Glucose", "C6H12O6", 0, "C(C1C(C(C(C(O1)O)O)O)O)O") + let c = @src.Compound::with_chemical( + "C0001", "Glucose", "C6H12O6", 0, "C(C1C(C(C(C(O1)O)O)O)O)O", + ) assert_eq(c.get_id(), "C0001") assert_eq(c.get_formula(), "C6H12O6") assert_eq(c.get_charge(), 0) diff --git a/test/moonbit/consensus_cluster_plus_test.mbt b/test/moonbit/consensus_cluster_plus_test.mbt index 9a67f008..6efe489a 100644 --- a/test/moonbit/consensus_cluster_plus_test.mbt +++ b/test/moonbit/consensus_cluster_plus_test.mbt @@ -9,24 +9,19 @@ test "consensus_cluster_plus_consensus_matrix" { [10.0, 11.0, 12.0], [11.0, 12.0, 13.0], ] - + let consensus = @src.ccp_calculate_consensus_matrix(data, 2, 10, 0.8) - + assert_true(consensus.length() == 4) assert_true(consensus[0].length() == 4) } ///| test "consensus_cluster_plus_kmeans" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - [11.0, 12.0], - ] - + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0], [11.0, 12.0]] + let labels = @src.ccp_kmeans_cluster(data, 2) - + assert_true(labels.length() == 4) } @@ -38,9 +33,9 @@ test "consensus_cluster_plus_consensus_score" { [0.8, 0.8, 1.0, 0.9], [0.7, 0.7, 0.9, 1.0], ] - + let score = @src.ccp_calculate_consensus_score(consensus) - + assert_true(score > 0.5) } @@ -54,9 +49,9 @@ test "consensus_cluster_plus_find_optimal_k" { [20.0, 21.0, 22.0], [21.0, 22.0, 23.0], ] - + let (best_k, scores) = @src.ccp_find_optimal_k(data, 2, 4) - + assert_true(best_k >= 2 && best_k <= 4) assert_true(scores.length() == 3) } @@ -69,9 +64,9 @@ test "consensus_cluster_plus_main" { [10.0, 11.0, 12.0], [11.0, 12.0, 13.0], ] - + let result = @src.bio_consensus_cluster(data, 2) - + assert_true(result.k == 2) assert_true(result.cluster_labels.length() == 4) assert_true(result.consensus_matrix.length() == 4) @@ -79,17 +74,13 @@ test "consensus_cluster_plus_main" { ///| test "consensus_cluster_plus_bio_api" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - ] - + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0]] + let (best_k, scores) = @src.bio_consensus_cluster_find_optimal_k(data, 2, 3) let consensus = @src.ccp_calculate_consensus_matrix(data, 2, 10, 0.8) let score = @src.bio_consensus_cluster_consensus_score(consensus) - + assert_true(best_k >= 2) assert_true(scores.length() > 0) assert_true(score >= 0.0) -} \ No newline at end of file +} diff --git a/test/moonbit/crystal_test.mbt b/test/moonbit/crystal_test.mbt index a8ecb034..7f4d43f3 100644 --- a/test/moonbit/crystal_test.mbt +++ b/test/moonbit/crystal_test.mbt @@ -51,7 +51,7 @@ test "unit_cell_volume_hexagonal" { gamma=120.0, ) let v = @src.unit_cell_volume(cell) - let expected = 5.0 * 5.0 * 7.0 * (3.0).sqrt() / 2.0 + let expected = 5.0 * 5.0 * 7.0 * 3.0.sqrt() / 2.0 assert_true((v - expected).abs() < 0.01) } @@ -142,13 +142,7 @@ test "lookup_space_group_unknown" { ///| test "crystal_atom_construction" { - let a = @src.CrystalAtom::new( - label="C1", - element="C", - x=0.5, - y=0.5, - z=0.5, - ) + let a = @src.CrystalAtom::new(label="C1", element="C", x=0.5, y=0.5, z=0.5) assert_eq(a.label, "C1") assert_eq(a.element, "C") assert_true((a.x - 0.5).abs() < 0.001) @@ -216,7 +210,7 @@ test "fractional_to_cartesian_orthorhombic" { beta=90.0, gamma=90.0, ) - let(x, y, z) = @src.fractional_to_cartesian(cell, 0.5, 0.25, 0.1) + let (x, y, z) = @src.fractional_to_cartesian(cell, 0.5, 0.25, 0.1) assert_true((x - 5.0).abs() < 0.001) assert_true((y - 5.0).abs() < 0.001) assert_true((z - 3.0).abs() < 0.001) @@ -232,7 +226,7 @@ test "cartesian_to_fractional_orthorhombic" { beta=90.0, gamma=90.0, ) - let(xf, yf, zf) = @src.cartesian_to_fractional(cell, 5.0, 5.0, 3.0) + let (xf, yf, zf) = @src.cartesian_to_fractional(cell, 5.0, 5.0, 3.0) assert_true((xf - 0.5).abs() < 0.001) assert_true((yf - 0.25).abs() < 0.001) assert_true((zf - 0.1).abs() < 0.001) @@ -251,8 +245,8 @@ test "coord_conversion_roundtrip_orthorhombic" { let xf0 = 0.3 let yf0 = 0.6 let zf0 = 0.9 - let(x, y, z) = @src.fractional_to_cartesian(cell, xf0, yf0, zf0) - let(xf, yf, zf) = @src.cartesian_to_fractional(cell, x, y, z) + let (x, y, z) = @src.fractional_to_cartesian(cell, xf0, yf0, zf0) + let (xf, yf, zf) = @src.cartesian_to_fractional(cell, x, y, z) assert_true((xf - xf0).abs() < 0.01) assert_true((yf - yf0).abs() < 0.01) assert_true((zf - zf0).abs() < 0.01) @@ -300,7 +294,7 @@ test "crystal_atom_distance_missing_atom" { ///| test "center_of_mass_basic" { let s = @src.crystal_sample_structure() - let(cx, cy, cz) = @src.center_of_mass(s) + let (cx, cy, cz) = @src.center_of_mass(s) // Center of mass should be a 3-tuple of finite doubles. assert_true(cx >= 0.0 || cx < 0.0) // just check it's a number assert_true(cy >= 0.0 || cy < 0.0) @@ -319,14 +313,14 @@ test "center_of_mass_empty_structure" { ) let s = @src.CrystalStructure::new( name="empty", - cell=cell, + cell~, space_group=@src.SpaceGroup::new(number=1, symbol="P 1"), atoms=[], bonds=[], z=1, molecular_weight=0.0, ) - let(cx, cy, cz) = @src.center_of_mass(s) + let (cx, cy, cz) = @src.center_of_mass(s) assert_true((cx - 0.0).abs() < 0.001) assert_true((cy - 0.0).abs() < 0.001) assert_true((cz - 0.0).abs() < 0.001) @@ -363,7 +357,7 @@ test "cif_block_loop_set_get" { let headers = ["_atom_site_label", "_atom_site_x"] let rows = [["C1", "0.5"], ["C2", "0.6"]] block.set_loop("_atom_site_label", headers, rows) - let(h, r) = block.get_loop("_atom_site_label") + let (h, r) = block.get_loop("_atom_site_label") assert_eq(h.length(), 2) assert_eq(r.length(), 2) assert_eq(r[0][0], "C1") @@ -372,7 +366,7 @@ test "cif_block_loop_set_get" { ///| test "cif_block_loop_missing_returns_empty" { let block = @src.CifBlock::new(name="test") - let(h, r) = block.get_loop("_nonexistent") + let (h, r) = block.get_loop("_nonexistent") assert_eq(h.length(), 0) assert_eq(r.length(), 0) } @@ -403,7 +397,7 @@ test "parse_cif_with_loop" { #|C3 0.7 let blocks = @src.parse_cif(text) assert_eq(blocks.length(), 1) - let(h, r) = blocks[0].get_loop("_atom_site_label") + let (h, r) = blocks[0].get_loop("_atom_site_label") assert_eq(h.length(), 2) assert_eq(r.length(), 3) assert_eq(r[0][0], "C1") @@ -521,7 +515,7 @@ test "crystal_summary_density_zero_when_mw_zero" { ) let s = @src.CrystalStructure::new( name="test", - cell=cell, + cell~, space_group=@src.SpaceGroup::new(number=1, symbol="P 1"), atoms=[], bonds=[], diff --git a/test/moonbit/csaw_test.mbt b/test/moonbit/csaw_test.mbt index 7ae988a4..094488bf 100644 --- a/test/moonbit/csaw_test.mbt +++ b/test/moonbit/csaw_test.mbt @@ -30,7 +30,7 @@ test "csaw_dataset_create" { let counts = [10.0, 20.0, 30.0] let windows = [ @src.CswWindow::new("w1", "chr1", 1000, 2000, counts), - @src.CswWindow::new("w2", "chr1", 3000, 4000, [50.0, 60.0, 70.0]) + @src.CswWindow::new("w2", "chr1", 3000, 4000, [50.0, 60.0, 70.0]), ] let samples = ["s1", "s2", "s3"] let lib_sizes = [1000.0, 2000.0, 3000.0] @@ -52,7 +52,7 @@ test "csaw_norm_library" { test "csaw_norm_tmm" { let windows = [ @src.CswWindow::new("w1", "chr1", 1000, 2000, [100.0, 200.0, 3000.0]), - @src.CswWindow::new("w2", "chr1", 3000, 4000, [150.0, 250.0, 350.0]) + @src.CswWindow::new("w2", "chr1", 3000, 4000, [150.0, 250.0, 350.0]), ] let samples = ["s1", "s2", "s3"] let lib_sizes = [1000.0, 2000.0, 3000.0] @@ -65,7 +65,7 @@ test "csaw_norm_tmm" { test "csaw_filter_abundance" { let windows = [ @src.CswWindow::new("w1", "chr1", 1000, 2000, [100.0, 200.0]), - @src.CswWindow::new("w2", "chr1", 3000, 4000, [5.0, 3.0]) + @src.CswWindow::new("w2", "chr1", 3000, 4000, [5.0, 3.0]), ] let samples = ["s1", "s2"] let lib_sizes = [1000.0, 2000.0] @@ -78,7 +78,7 @@ test "csaw_filter_abundance" { test "csaw_test_differential" { let windows = [ @src.CswWindow::new("w1", "chr1", 1000, 2000, [50.0, 100.0, 200.0, 250.0]), - @src.CswWindow::new("w2", "chr1", 3000, 4000, [80.0, 120.0, 180.0, 220.0]) + @src.CswWindow::new("w2", "chr1", 3000, 4000, [80.0, 120.0, 180.0, 220.0]), ] let samples = ["s1", "s2", "s3", "s4"] let lib_sizes = [1000.0, 2000.0, 3000.0, 4000.0] @@ -94,7 +94,7 @@ test "csaw_test_differential" { test "csaw_find_regions" { let windows = [ @src.CswWindow::new("w1", "chr1", 1000, 2000, [500.0, 100.0, 200.0, 250.0]), - @src.CswWindow::new("w2", "chr1", 3000, 4000, [80.0, 120.0, 180.0, 220.0]) + @src.CswWindow::new("w2", "chr1", 3000, 4000, [80.0, 120.0, 180.0, 220.0]), ] let samples = ["s1", "s2", "s3", "s4"] let lib_sizes = [1000.0, 2000.0, 3000.0, 4000.0] diff --git a/test/moonbit/cyclone_test.mbt b/test/moonbit/cyclone_test.mbt index 9f0b9570..5495bb8d 100644 --- a/test/moonbit/cyclone_test.mbt +++ b/test/moonbit/cyclone_test.mbt @@ -37,7 +37,7 @@ test "cyclone_score_phase" { let pairs = @src.cyclone_get_gene_pairs() let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + // Score the first cell (should be G1 phase) let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() @@ -46,7 +46,7 @@ test "cyclone_score_phase" { cell_expr.push(counts[g][0]) g = g + 1 } - + let g1_score = @src.cyclone_score_phase(cell_expr, gene_names, pairs, "G1") assert_true(g1_score >= 0.0) assert_true(g1_score <= 1.0) @@ -57,7 +57,7 @@ test "cyclone_score_cell" { let pairs = @src.cyclone_get_gene_pairs() let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -65,7 +65,7 @@ test "cyclone_score_cell" { cell_expr.push(counts[g][0]) g = g + 1 } - + let scores = @src.cyclone_score_cell(cell_expr, gene_names, pairs) assert_true(scores.size() > 0) assert_true(scores.contains("G1")) @@ -93,7 +93,7 @@ test "cyclone_score_single_cell" { let pairs = @src.cyclone_get_gene_pairs() let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -101,8 +101,10 @@ test "cyclone_score_single_cell" { cell_expr.push(counts[g][0]) g = g + 1 } - - let result = @src.cyclone_score_single_cell(cell_expr, "test_cell", gene_names, pairs) + + let result = @src.cyclone_score_single_cell( + cell_expr, "test_cell", gene_names, pairs, + ) assert_eq(result.cell_id, "test_cell") assert_true(result.scores.size() == 4) assert_true(result.phases.length() == 4) @@ -117,10 +119,10 @@ test "cyclone_score_single_cell" { test "cyclone_score_cells" { let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let results = @src.cyclone_score_cells(counts, cell_ids, gene_names) assert_eq(results.length(), 20) - + // Check first cell assert_eq(results[0].cell_id, "cell_1") assert_true(results[0].assigned_phase.length() > 0) @@ -130,12 +132,12 @@ test "cyclone_score_cells" { test "cyclone_phase_distribution" { let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let results = @src.cyclone_score_cells(counts, cell_ids, gene_names) let dist = @src.cyclone_phase_distribution(results) - + assert_true(dist.size() > 0) - + // Sum of proportions should be approximately 1.0 let keys = dist.keys() let mut sum = 0.0 @@ -149,7 +151,7 @@ test "cyclone_phase_distribution" { test "cyclone_create_test_data" { let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + assert_eq(cell_ids.length(), 20) assert_true(gene_names.length() > 0) assert_true(counts.length() > 0) @@ -175,7 +177,7 @@ test "cyclone_score_empty_cells" { let gene_names : Array[String] = Array::new() let cell_ids : Array[String] = Array::new() let counts : Array[Array[Double]] = Array::new() - + let results = @src.cyclone_score_cells(counts, cell_ids, gene_names) assert_eq(results.length(), 0) } @@ -191,10 +193,10 @@ test "cyclone_phase_distribution_empty" { test "cyclone_average_scores" { let test_data = @src.cyclone_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let results = @src.cyclone_score_cells(counts, cell_ids, gene_names) let avg = @src.cyclone_average_scores(results) - + assert_true(avg.size() == 4) assert_true(avg.contains("G1")) assert_true(avg.contains("S")) @@ -207,4 +209,4 @@ test "cyclone_average_scores_empty" { let results : Array[@src.CycloneResult] = Array::new() let avg = @src.cyclone_average_scores(results) assert_eq(avg.size(), 0) -} \ No newline at end of file +} diff --git a/test/moonbit/decontx_test.mbt b/test/moonbit/decontx_test.mbt new file mode 100644 index 00000000..378ef477 --- /dev/null +++ b/test/moonbit/decontx_test.mbt @@ -0,0 +1,656 @@ +// Tests for Bioconductor decontX-inspired ambient RNA decontamination. + +///| +fn decontx_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "decontX value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn decontx_test_result() -> @src.DecontXResult { + let (counts, genes, cells, clusters, _) = @src.decontx_example_data() + @src.decontx(counts, genes, cells, clusters) catch { + _ => abort("example decontX analysis should succeed") + } +} + +///| +fn decontx_test_sce() -> @src.SingleCellExperiment { + let (counts, _, _, clusters, _) = @src.decontx_example_data() + @src.SCEBuilder::new() + |> @src.SCEBuilder::add_assay("counts", counts) + |> @src.SCEBuilder::add_row_data("symbol", [ + "TR1", "TR2", "BR1", "BR2", "HK", "LOW", + ]) + |> @src.SCEBuilder::add_col_data("cluster", clusters) + |> @src.SCEBuilder::build +} + +///| +test "decontX: default configuration is scientifically conservative" { + let config = @src.DecontXConfig::default() + assert_eq(config.max_iterations, 100) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.contamination_prior_alpha, 1.0) + assert_eq(config.contamination_prior_beta, 9.0) + assert_eq(config.initial_contamination, 0.1) +} + +///| +test "decontX: custom configuration preserves all parameters" { + let config = @src.DecontXConfig::create( + max_iterations=25, + tolerance=1.0e-5, + contamination_prior_alpha=2.0, + contamination_prior_beta=8.0, + profile_pseudocount=0.25, + initial_contamination=0.2, + minimum_probability=1.0e-10, + ) catch { + _ => abort("valid custom configuration should build") + } + assert_eq(config.max_iterations, 25) + assert_eq(config.profile_pseudocount, 0.25) + assert_eq(config.minimum_probability, 1.0e-10) +} + +///| +test "decontX: configuration rejects invalid iteration and tolerance" { + let iterations = try { + ignore(@src.DecontXConfig::create(max_iterations=0)) + false + } catch { + DecontXError(_) => true + } + let tolerance = try { + ignore(@src.DecontXConfig::create(tolerance=0.0)) + false + } catch { + DecontXError(_) => true + } + assert_true(iterations) + assert_true(tolerance) +} + +///| +test "decontX: configuration rejects invalid priors" { + let beta = try { + ignore(@src.DecontXConfig::create(contamination_prior_beta=0.0)) + false + } catch { + DecontXError(_) => true + } + let profile = try { + ignore(@src.DecontXConfig::create(profile_pseudocount=-1.0)) + false + } catch { + DecontXError(_) => true + } + assert_true(beta) + assert_true(profile) +} + +///| +test "decontX: configuration rejects invalid probabilities" { + let initial = try { + ignore(@src.DecontXConfig::create(initial_contamination=1.0)) + false + } catch { + DecontXError(_) => true + } + let minimum = try { + ignore(@src.DecontXConfig::create(minimum_probability=0.0)) + false + } catch { + DecontXError(_) => true + } + assert_true(initial) + assert_true(minimum) +} + +///| +test "decontX: example data uses gene by cell orientation" { + let (counts, genes, cells, clusters, background) = @src.decontx_example_data() + assert_eq(counts.length(), 6) + assert_eq(counts[0].length(), 8) + assert_eq(genes.length(), 6) + assert_eq(cells.length(), 8) + assert_eq(clusters.length(), 8) + assert_eq(background.length(), 6) +} + +///| +test "decontX: result dimensions match the source matrix" { + let result = decontx_test_result() + assert_eq(result.n_genes(), 6) + assert_eq(result.n_cells(), 8) + assert_eq(result.n_clusters(), 2) + assert_eq(result.contamination_counts.length(), 6) + assert_eq(result.contamination_counts[0].length(), 8) +} + +///| +test "decontX: omitted identifiers are generated deterministically" { + let result = @src.decontx([[10.0, 1.0], [1.0, 10.0]], [], [], ["A", "B"]) catch { + _ => abort("generated names should be accepted") + } + assert_eq(result.gene_names, ["Gene1", "Gene2"]) + assert_eq(result.cell_names, ["Cell1", "Cell2"]) +} + +///| +test "decontX: cluster order follows first appearance" { + let result = @src.decontx([[10.0, 1.0, 5.0], [1.0, 10.0, 4.0]], [], [], [ + "B", "A", "B", + ]) catch { + _ => abort("valid labels should be accepted") + } + assert_eq(result.cluster_names, ["B", "A"]) + assert_eq(result.cluster_indices, [0, 1, 0]) +} + +///| +test "decontX: native and contaminant counts reconstruct observations" { + let (counts, _, _, _, _) = @src.decontx_example_data() + let result = decontx_test_result() + for gene in 0.. 0.0 && value < 1.0) + } +} + +///| +test "decontX: native profiles are normalized" { + let result = decontx_test_result() + for profile in result.native_profiles { + let mut total = 0.0 + for value in profile { + total = total + value + } + decontx_test_close(total, 1.0, 1.0e-10) + } +} + +///| +test "decontX: contaminant profiles are normalized" { + let result = decontx_test_result() + for profile in result.contaminant_profiles { + let mut total = 0.0 + for value in profile { + total = total + value + } + decontx_test_close(total, 1.0, 1.0e-10) + } +} + +///| +test "decontX: B-cell marker contamination is removed from T cells" { + let result = decontx_test_result() + assert_true(result.corrected_counts[2][0] < 8.0) + assert_true(result.contamination_counts[2][0] > 0.0) +} + +///| +test "decontX: T-cell marker contamination is removed from B cells" { + let result = decontx_test_result() + assert_true(result.corrected_counts[0][4] < 8.0) + assert_true(result.contamination_counts[0][4] > 0.0) +} + +///| +test "decontX: diagnostics record bounded inference work" { + let result = decontx_test_result() + assert_true(result.diagnostics.iterations >= 1) + assert_true(result.diagnostics.iterations <= 100) + assert_eq( + result.diagnostics.log_likelihoods.length(), + result.diagnostics.iterations, + ) +} + +///| +test "decontX: diagnostic likelihoods are finite" { + let result = decontx_test_result() + assert_true(result.diagnostics.initial_log_likelihood.abs() < 1.0e300) + assert_true(result.diagnostics.final_log_likelihood.abs() < 1.0e300) + for value in result.diagnostics.log_likelihoods { + assert_true(value == value) + assert_true(value.abs() < 1.0e300) + } +} + +///| +test "decontX: cell estimate reports decomposed library size" { + let result = decontx_test_result() + match result.cell_estimate(0) { + Some(estimate) => { + assert_eq(estimate.cell_name, "T1") + assert_eq(estimate.cluster, "T") + decontx_test_close( + estimate.native_counts + estimate.contaminant_counts, + estimate.library_size, + 1.0e-10, + ) + decontx_test_close(estimate.library_size, 214.0, 1.0e-10) + } + None => abort("valid cell estimate should exist") + } +} + +///| +test "decontX: cell estimate validates indices" { + let result = decontx_test_result() + assert_true(result.cell_estimate(-1) is None) + assert_true(result.cell_estimate(8) is None) +} + +///| +test "decontX: gene and cell lookup use preserved identifiers" { + let result = decontx_test_result() + assert_eq(result.gene_index("B_marker_1"), Some(2)) + assert_eq(result.cell_index("B3"), Some(6)) + assert_true(result.gene_index("missing") is None) + assert_true(result.cell_index("missing") is None) +} + +///| +test "decontX: cluster estimates partition all cells" { + let result = decontx_test_result() + let estimates = result.cluster_estimates() + assert_eq(estimates.length(), 2) + assert_eq(estimates[0].cluster, "T") + assert_eq(estimates[0].n_cells, 4) + assert_eq(estimates[1].cluster, "B") + assert_eq(estimates[1].n_cells, 4) + decontx_test_close( + estimates[0].native_counts + estimates[0].contaminant_counts, + estimates[0].total_counts, + 1.0e-10, + ) +} + +///| +test "decontX: most contaminated cells are sorted descending" { + let result = decontx_test_result() + let top = result.most_contaminated_cells(4) + assert_eq(top.length(), 4) + for index in 1..= top[index].contamination) + } +} + +///| +test "decontX: top-cell query handles zero and oversized limits" { + let result = decontx_test_result() + assert_eq(result.most_contaminated_cells(0).length(), 0) + assert_eq(result.most_contaminated_cells(100).length(), 8) +} + +///| +test "decontX: summary exposes model dimensions and convergence" { + let summary = decontx_test_result().summary() + assert_true(summary.contains("6 genes x 8 cells")) + assert_true(summary.contains("2 clusters")) + assert_true(summary.contains("mean contamination=")) + assert_true(summary.contains("converged=")) +} + +///| +test "decontX: explicit background fixes a common ambient profile" { + let (counts, genes, cells, clusters, background) = @src.decontx_example_data() + let result = @src.decontx(counts, genes, cells, clusters, background~) catch { + _ => abort("valid background should be accepted") + } + assert_true(result.used_background) + assert_eq(result.contaminant_profiles[0], result.contaminant_profiles[1]) +} + +///| +test "decontX: background changes the inferred contaminant distribution" { + let (counts, genes, cells, clusters, background) = @src.decontx_example_data() + let without = @src.decontx(counts, genes, cells, clusters) catch { + _ => abort("analysis without background should succeed") + } + let with_background = @src.decontx( + counts, + genes, + cells, + clusters, + background~, + ) catch { + _ => abort("analysis with background should succeed") + } + assert_true( + (without.contaminant_profiles[0][0] - + with_background.contaminant_profiles[0][0]).abs() > + 1.0e-6, + ) +} + +///| +test "decontX: one cluster is identifiable with external background" { + let result = @src.decontx( + [[10.0, 12.0], [2.0, 3.0]], + ["A", "B"], + ["C1", "C2"], + ["one", "one"], + background=[[1.0], [9.0]], + ) catch { + _ => abort("background should identify one-cluster contamination") + } + assert_eq(result.n_clusters(), 1) + assert_true(result.used_background) +} + +///| +test "decontX: one cluster without background is rejected" { + let raised = try { + ignore(@src.decontx([[10.0, 12.0], [2.0, 3.0]], [], [], ["one", "one"])) + false + } catch { + DecontXError(_) => true + } + assert_true(raised) +} + +///| +test "decontX: empty and ragged matrices are rejected" { + let empty = try { + ignore(@src.decontx([], [], [], [])) + false + } catch { + DecontXError(_) => true + } + let ragged = try { + ignore(@src.decontx([[1.0], [2.0, 3.0]], [], [], ["A"])) + false + } catch { + DecontXError(_) => true + } + assert_true(empty) + assert_true(ragged) +} + +///| +test "decontX: negative and non-finite counts are rejected" { + let negative = try { + ignore(@src.decontx([[1.0, -1.0]], [], [], ["A", "B"])) + false + } catch { + DecontXError(_) => true + } + let non_finite = try { + ignore(@src.decontx([[1.0, @double.not_a_number]], [], [], ["A", "B"])) + false + } catch { + DecontXError(_) => true + } + assert_true(negative) + assert_true(non_finite) +} + +///| +test "decontX: identifier dimensions are validated" { + let genes = try { + ignore(@src.decontx([[1.0, 2.0], [2.0, 1.0]], ["only_one"], [], ["A", "B"])) + false + } catch { + DecontXError(_) => true + } + let cells = try { + ignore(@src.decontx([[1.0, 2.0], [2.0, 1.0]], [], ["only_one"], ["A", "B"])) + false + } catch { + DecontXError(_) => true + } + assert_true(genes) + assert_true(cells) +} + +///| +test "decontX: duplicate and blank identifiers are rejected" { + let duplicate = try { + ignore( + @src.decontx([[1.0, 2.0], [2.0, 1.0]], ["same", "same"], [], ["A", "B"]), + ) + false + } catch { + DecontXError(_) => true + } + let blank = try { + ignore(@src.decontx([[1.0, 2.0], [2.0, 1.0]], [], ["C1", " "], ["A", "B"])) + false + } catch { + DecontXError(_) => true + } + assert_true(duplicate) + assert_true(blank) +} + +///| +test "decontX: cluster labels are validated" { + let length = try { + ignore(@src.decontx([[1.0, 2.0], [2.0, 1.0]], [], [], ["A"])) + false + } catch { + DecontXError(_) => true + } + let blank = try { + ignore(@src.decontx([[1.0, 2.0], [2.0, 1.0]], [], [], ["A", ""])) + false + } catch { + DecontXError(_) => true + } + assert_true(length) + assert_true(blank) +} + +///| +test "decontX: background dimensions and values are validated" { + let dimensions = try { + ignore( + @src.decontx([[1.0, 2.0], [2.0, 1.0]], [], [], ["A", "B"], background=[ + [1.0], + ]), + ) + false + } catch { + DecontXError(_) => true + } + let negative = try { + ignore( + @src.decontx([[1.0, 2.0], [2.0, 1.0]], [], [], ["A", "B"], background=[ + [1.0], + [-1.0], + ]), + ) + false + } catch { + DecontXError(_) => true + } + assert_true(dimensions) + assert_true(negative) +} + +///| +test "decontX: zero-library cells remain zero" { + let result = @src.decontx([[0.0, 10.0, 1.0], [0.0, 1.0, 10.0]], [], [], [ + "A", "A", "B", + ]) catch { + _ => abort("zero-library cells should be supported") + } + assert_eq(result.corrected_counts[0][0], 0.0) + assert_eq(result.corrected_counts[1][0], 0.0) + assert_eq(result.contamination_counts[0][0], 0.0) + assert_eq(result.contamination_counts[1][0], 0.0) +} + +///| +test "decontX: configured maximum iteration is respected" { + let (counts, genes, cells, clusters, _) = @src.decontx_example_data() + let config = @src.DecontXConfig::create(max_iterations=1, tolerance=1.0e-30) catch { + _ => abort("valid iteration limit should build") + } + let result = @src.decontx(counts, genes, cells, clusters, config~) catch { + _ => abort("bounded analysis should succeed") + } + assert_eq(result.diagnostics.iterations, 1) + assert_true(!result.diagnostics.converged) +} + +///| +test "decontX: stronger low-contamination prior shrinks estimates" { + let (counts, genes, cells, clusters, _) = @src.decontx_example_data() + let weak = @src.DecontXConfig::create( + max_iterations=1, + contamination_prior_alpha=1.0, + contamination_prior_beta=1.0, + ) catch { + _ => abort("weak prior should build") + } + let strong = @src.DecontXConfig::create( + max_iterations=1, + contamination_prior_alpha=1.0, + contamination_prior_beta=99.0, + ) catch { + _ => abort("strong prior should build") + } + let weak_result = @src.decontx(counts, genes, cells, clusters, config=weak) catch { + _ => abort("weak-prior analysis should succeed") + } + let strong_result = @src.decontx( + counts, + genes, + cells, + clusters, + config=strong, + ) catch { + _ => abort("strong-prior analysis should succeed") + } + assert_true( + strong_result.mean_contamination() < weak_result.mean_contamination(), + ) +} + +///| +test "decontX: automatic clustering finds both example populations" { + let (counts, genes, cells, _, _) = @src.decontx_example_data() + let result = @src.decontx_auto(counts, genes, cells, 2) catch { + _ => abort("automatic clustering should succeed") + } + assert_eq(result.n_cells(), 8) + assert_eq(result.n_clusters(), 2) + assert_true(result.cluster_names[0].has_prefix("cluster_")) +} + +///| +test "decontX: automatic clustering validates requested cluster count" { + let (counts, genes, cells, _, _) = @src.decontx_example_data() + let one = try { + ignore(@src.decontx_auto(counts, genes, cells, 1)) + false + } catch { + DecontXError(_) => true + } + let too_many = try { + ignore(@src.decontx_auto(counts, genes, cells, 9)) + false + } catch { + DecontXError(_) => true + } + assert_true(one) + assert_true(too_many) +} + +///| +test "decontX: SCE integration adds assay and diagnostics" { + let output = @src.decontx_sce(decontx_test_sce()) catch { + _ => abort("valid SCE integration should succeed") + } + let corrected = @src.sce_get_assay(output.experiment, "decontXcounts") + assert_eq(corrected.length(), 6) + assert_eq(corrected[0].length(), 8) + assert_eq( + @src.sce_get_col_data(output.experiment, "decontX_contamination").length(), + 8, + ) + assert_eq(@src.sce_get_col_data(output.experiment, "decontX_cluster"), [ + "T", "T", "T", "T", "B", "B", "B", "B", + ]) +} + +///| +test "decontX: SCE integration does not mutate the source container" { + let source = decontx_test_sce() + let _ = @src.decontx_sce(source) catch { + _ => abort("valid SCE integration should succeed") + } + assert_eq(@src.sce_get_assay(source, "decontXcounts").length(), 0) + assert_eq(@src.sce_get_col_data(source, "decontX_contamination").length(), 0) +} + +///| +test "decontX: SCE integration validates assay and cluster metadata" { + let missing_cluster = @src.SingleCellExperiment::new( + [[1.0, 2.0], [2.0, 1.0]], + ["G1", "G2"], + ["C1", "C2"], + ) + let cluster_error = try { + ignore(@src.decontx_sce(missing_cluster)) + false + } catch { + DecontXError(_) => true + } + let missing_assay = @src.SCEBuilder::new() + |> @src.SCEBuilder::add_assay("other", [[1.0, 2.0], [2.0, 1.0]]) + |> @src.SCEBuilder::add_col_data("cluster", ["A", "B"]) + |> @src.SCEBuilder::build + let assay_error = try { + ignore(@src.decontx_sce(missing_assay)) + false + } catch { + DecontXError(_) => true + } + assert_true(cluster_error) + assert_true(assay_error) +} + +///| +test "decontX: SCE integration supports a custom output assay" { + let output = @src.decontx_sce( + decontx_test_sce(), + output_assay="ambient_corrected", + ) catch { + _ => abort("custom output assay should be accepted") + } + assert_eq( + @src.sce_get_assay(output.experiment, "ambient_corrected").length(), + 6, + ) + assert_eq(@src.sce_get_assay(output.experiment, "decontXcounts").length(), 0) +} diff --git a/test/moonbit/decoupler_test.mbt b/test/moonbit/decoupler_test.mbt index 37251805..bb6e03a2 100644 --- a/test/moonbit/decoupler_test.mbt +++ b/test/moonbit/decoupler_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Bioconductor decoupleR module - Functional activity inference. - test "pkn_edge_create" { let e = @src.pkn_edge("TF1", "G1", 1.0) assert_eq(e.source, "TF1") @@ -8,6 +7,7 @@ test "pkn_edge_create" { assert_true((e.weight - 1.0).abs() < 1.0e-9) } +///| test "pkn_from_edges" { let edges = [ @src.pkn_edge("TF1", "G1", 1.0), @@ -20,6 +20,7 @@ test "pkn_from_edges" { assert_eq(pkn.targets.length(), 3) } +///| test "decoupler_sample_data_shape" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() assert_eq(samples.length(), 5) @@ -29,9 +30,16 @@ test "decoupler_sample_data_shape" { assert_true(pkn.edges.length() > 0) } +///| test "decoupler_wsum_basic" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) assert_eq(result.method, "wsum") assert_eq(result.samples.length(), 5) // 3 TFs in the PKN @@ -50,28 +58,50 @@ test "decoupler_wsum_basic" { assert_true(result.matrix[tf1_idx][0] > result.matrix[tf1_idx][3]) } +///| test "decoupler_wmean_basic" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wmean_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wmean_method(), + ) assert_eq(result.method, "wmean") assert_eq(result.regulators.length(), 3) // wmean normalizes by sum of absolute weights, so magnitude is smaller than wsum - let wsum_result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let wsum_result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) let mut i = 0 while i < result.regulators.length() { let mut j = 0 while j < result.samples.length() { // |wmean| should be <= |wsum| (since abs_w_sum >= 1.0) - assert_true(result.matrix[i][j].abs() <= wsum_result.matrix[i][j].abs() + 1.0e-9) + assert_true( + result.matrix[i][j].abs() <= wsum_result.matrix[i][j].abs() + 1.0e-9, + ) j = j + 1 } i = i + 1 } } +///| test "decoupler_norm_basic" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_norm_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_norm_method(), + ) assert_eq(result.method, "norm") // Normalized scores should have mean ~0 per regulator across samples let mut i = 0 @@ -88,9 +118,16 @@ test "decoupler_norm_basic" { } } +///| test "decoupler_ulm_basic" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_ulm_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_ulm_method(), + ) assert_eq(result.method, "ulm") assert_eq(result.regulators.length(), 3) // ULM scores are standardized; mean across samples should be ~0 @@ -108,9 +145,16 @@ test "decoupler_ulm_basic" { } } +///| test "decoupler_mlm_basic" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_mlm_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_mlm_method(), + ) assert_eq(result.method, "mlm") assert_eq(result.regulators.length(), 3) // MLM slope for TF1 should be positive (TF1 activates its targets) @@ -127,16 +171,33 @@ test "decoupler_mlm_basic" { assert_true(result.matrix[tf1_idx][0] > 0.0) } +///| test "decoupler_scores_per_cell" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) // Should have one score per regulator × sample combination - assert_eq(result.scores.length(), result.regulators.length() * result.samples.length()) + assert_eq( + result.scores.length(), + result.regulators.length() * result.samples.length(), + ) } +///| test "decoupler_top_regulators" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) let top = @src.decoupler_top_regulators(result, "S1", 2) assert_true(top.length() <= 2) // Top regulators for S1 should include TF1 (since S1 is TF1-high) @@ -151,16 +212,30 @@ test "decoupler_top_regulators" { assert_true(found_tf1) } +///| test "decoupler_top_regulators_unknown_sample" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) let top = @src.decoupler_top_regulators(result, "UNKNOWN", 3) assert_eq(top.length(), 0) } +///| test "decoupler_filter_scores" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) let filtered = @src.decoupler_filter_scores(result, 0.5) // All filtered scores should have |score| >= 0.5 let mut i = 0 @@ -170,6 +245,7 @@ test "decoupler_filter_scores" { } } +///| test "decoupler_method_to_string" { assert_eq(@src.decoupler_wsum_method().to_string(), "wsum") assert_eq(@src.decoupler_wmean_method().to_string(), "wmean") @@ -178,6 +254,7 @@ test "decoupler_method_to_string" { assert_eq(@src.decoupler_mlm_method().to_string(), "mlm") } +///| test "decoupler_default_method_is_ulm" { let (expr, samples, genes, pkn) = @src.decoupler_sample_data() // Default method argument should be ULM @@ -185,22 +262,28 @@ test "decoupler_default_method_is_ulm" { assert_eq(result.method, "ulm") } +///| test "decoupler_wsum_symmetric_in_sign" { // A PKN with negative weight should produce negative wsum - let edges = [ - @src.pkn_edge("REP", "G1", -1.0), - ] + let edges = [@src.pkn_edge("REP", "G1", -1.0)] let pkn = @src.pkn_from_edges(edges) let expr = [[0.0, 5.0, 0.0], [0.0, 0.0, 0.0]] let samples = ["S1", "S2"] let genes = ["REP", "G1", "G2"] - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) // REP's only target is G1 with weight -1.0; in S1, G1=5 → wsum=-5; in S2, G1=0 → wsum=0 assert_eq(result.regulators.length(), 1) assert_true((result.matrix[0][0] - -5.0).abs() < 1.0e-9) assert_true((result.matrix[0][1] - 0.0).abs() < 1.0e-9) } +///| test "decoupler_pkn_with_unknown_genes" { // PKN with regulator/target not in expression matrix should be filtered out let edges = [ @@ -212,18 +295,31 @@ test "decoupler_pkn_with_unknown_genes" { let expr = [[1.0, 2.0], [3.0, 4.0]] let samples = ["S1", "S2"] let genes = ["TF1", "G1"] - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) // Only TF1 (which appears as a gene) should be in regulators assert_eq(result.regulators.length(), 1) assert_eq(result.regulators[0], "TF1") } +///| test "decoupler_empty_pkn" { let pkn = @src.pkn_from_edges([]) let expr = [[1.0, 2.0], [3.0, 4.0]] let samples = ["S1", "S2"] let genes = ["TF1", "G1"] - let result = @src.decoupler_run(expr, samples, genes, pkn, method=@src.decoupler_wsum_method()) + let result = @src.decoupler_run( + expr, + samples, + genes, + pkn, + method=@src.decoupler_wsum_method(), + ) assert_eq(result.regulators.length(), 0) assert_eq(result.scores.length(), 0) } diff --git a/test/moonbit/delayed_matrix_stats_test.mbt b/test/moonbit/delayed_matrix_stats_test.mbt index c31880be..c090f620 100644 --- a/test/moonbit/delayed_matrix_stats_test.mbt +++ b/test/moonbit/delayed_matrix_stats_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for DelayedMatrixStats module. - test "create_delayed_matrix_basic" { let mat = @src.create_delayed_matrix( [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], @@ -22,11 +21,7 @@ test "create_delayed_matrix_empty" { ///| test "delayed_matrix_get_element" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0], [3.0, 4.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0], [3.0, 4.0]], [], []) assert_eq(mat.get(0, 0), 1.0) assert_eq(mat.get(0, 1), 2.0) assert_eq(mat.get(1, 0), 3.0) @@ -35,11 +30,7 @@ test "delayed_matrix_get_element" { ///| test "delayed_matrix_dim" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let d = mat.dim() assert_eq(d[0], 2) assert_eq(d[1], 3) @@ -63,11 +54,9 @@ test "delayed_matrix_subset" { ///| test "delayed_matrix_subset_names" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0], [3.0, 4.0]], - ["r1", "r2"], - ["c1", "c2"], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0], [3.0, 4.0]], ["r1", "r2"], [ + "c1", "c2", + ]) let sub = @src.delayed_matrix_subset(mat, [1], [0]) assert_eq(sub.row_names.length(), 1) assert_eq(sub.col_names.length(), 1) @@ -75,11 +64,7 @@ test "delayed_matrix_subset_names" { ///| test "row_stats_mean" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let means = @src.row_stats(mat, "mean") assert_eq(means.length(), 2) assert_eq(means[0], 2.0) @@ -88,11 +73,7 @@ test "row_stats_mean" { ///| test "col_stats_mean" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let means = @src.col_stats(mat, "mean") assert_eq(means.length(), 3) assert_eq(means[0], 2.5) @@ -102,11 +83,7 @@ test "col_stats_mean" { ///| test "row_medians" { - let mat = @src.create_delayed_matrix( - [[1.0, 3.0, 2.0], [4.0, 6.0, 5.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 3.0, 2.0], [4.0, 6.0, 5.0]], [], []) let medians = @src.row_medians(mat) assert_eq(medians.length(), 2) assert_eq(medians[0], 2.0) @@ -128,11 +105,7 @@ test "col_medians" { ///| test "row_medians_even" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0, 4.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0, 4.0]], [], []) let medians = @src.row_medians(mat) assert_eq(medians[0], 2.5) } @@ -165,11 +138,7 @@ test "dms_col_means" { ///| test "dms_row_vars" { - let mat = @src.create_delayed_matrix( - [[2.0, 4.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[2.0, 4.0, 6.0]], [], []) let vars = @src.dms_row_vars(mat) assert_eq(vars.length(), 1) assert_eq(vars[0], 4.0) @@ -177,11 +146,7 @@ test "dms_row_vars" { ///| test "col_vars" { - let mat = @src.create_delayed_matrix( - [[2.0], [4.0], [6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[2.0], [4.0], [6.0]], [], []) let vars = @src.col_vars(mat) assert_eq(vars.length(), 1) assert_eq(vars[0], 4.0) @@ -189,11 +154,7 @@ test "col_vars" { ///| test "row_allsums" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let sums = @src.row_allsums(mat) assert_eq(sums.length(), 2) assert_eq(sums[0], 6.0) @@ -202,11 +163,7 @@ test "row_allsums" { ///| test "row_n_basic" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let n = @src.row_n(mat) assert_eq(n.length(), 2) assert_eq(n[0], 3) @@ -243,11 +200,7 @@ test "row_stats_min_max" { ///| test "col_stats_min_max" { - let mat = @src.create_delayed_matrix( - [[3.0, 1.0, 4.0], [9.0, 2.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[3.0, 1.0, 4.0], [9.0, 2.0, 6.0]], [], []) let mins = @src.col_stats(mat, "min") let maxs = @src.col_stats(mat, "max") assert_eq(mins[0], 3.0) @@ -338,7 +291,10 @@ test "row_medians_with_na" { ///| test "col_stats_all_na" { let mat = @src.create_delayed_matrix( - [[@double.not_a_number, @double.not_a_number], [@double.not_a_number, @double.not_a_number]], + [ + [@double.not_a_number, @double.not_a_number], + [@double.not_a_number, @double.not_a_number], + ], [], [], ) @@ -376,11 +332,7 @@ test "row_stats_nna_stat" { ///| test "row_stats_nn_stat" { - let mat = @src.create_delayed_matrix( - [[1.0, @double.not_a_number, 3.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, @double.not_a_number, 3.0]], [], []) let nn_vals = @src.row_stats(mat, "nn") assert_eq(nn_vals.length(), 1) assert_eq(nn_vals[0], 2.0) @@ -424,11 +376,7 @@ test "row_vars_single_element" { ///| test "col_medians_odd" { - let mat = @src.create_delayed_matrix( - [[1.0], [5.0], [3.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0], [5.0], [3.0]], [], []) let medians = @src.col_medians(mat) assert_eq(medians[0], 3.0) } @@ -470,14 +418,10 @@ test "matrix_stats_result_creation" { ///| test "row_stats_sum_equals_allsums" { - let mat = @src.create_delayed_matrix( - [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], - [], - [], - ) + let mat = @src.create_delayed_matrix([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [], []) let sums1 = @src.row_allsums(mat) let sums2 = @src.row_stats(mat, "sum") assert_eq(sums1.length(), sums2.length()) assert_eq(sums1[0], sums2[0]) assert_eq(sums1[1], sums2[1]) -} \ No newline at end of file +} diff --git a/test/moonbit/deseq2_test.mbt b/test/moonbit/deseq2_test.mbt index a8f1c0fd..11f69c32 100644 --- a/test/moonbit/deseq2_test.mbt +++ b/test/moonbit/deseq2_test.mbt @@ -4,7 +4,7 @@ ///| test "test_sample_deseq_dataset" { let dds = @src.sample_deseq_dataset() - + assert_eq(dds.counts.length(), 20) assert_eq(dds.col_names.length(), 6) assert_eq(dds.row_names.length(), 20) @@ -18,15 +18,15 @@ test "test_sample_deseq_dataset" { test "test_estimate_size_factors" { let dds = @src.sample_deseq_dataset() let dds_sf = @src.estimate_size_factors(dds) - + assert_eq(dds_sf.size_factors.length(), 6) - + let mut sum_sf = 0.0 for sf in dds_sf.size_factors { assert_true(sf > 0.0) sum_sf = sum_sf + sf } - + let mean_sf = sum_sf / 6.0 assert_true(mean_sf > 0.9 && mean_sf < 1.1) } @@ -36,7 +36,7 @@ test "test_normalize_counts" { let dds = @src.sample_deseq_dataset() let dds_sf = @src.estimate_size_factors(dds) let normalized = @src.normalize_counts(dds_sf) - + assert_eq(normalized.length(), 20) assert_eq(normalized[0].length(), 6) } @@ -46,7 +46,7 @@ test "test_log2_cpm" { let dds = @src.sample_deseq_dataset() let dds_sf = @src.estimate_size_factors(dds) let log2_cpm = @src.log2_cpm(dds_sf) - + assert_eq(log2_cpm.length(), 20) assert_eq(log2_cpm[0].length(), 6) } @@ -56,9 +56,9 @@ test "test_estimate_dispersions" { let dds = @src.sample_deseq_dataset() let dds_sf = @src.estimate_size_factors(dds) let dds_disp = @src.estimate_dispersions(dds_sf) - + assert_eq(dds_disp.dispersions.length(), 20) - + for disp in dds_disp.dispersions { assert_true(disp >= 0.001) } @@ -70,7 +70,7 @@ test "test_deseq" { let dds_sf = @src.estimate_size_factors(dds) let dds_disp = @src.estimate_dispersions(dds_sf) let res = @src.deseq(dds_disp) - + assert_eq(res.row_names.length(), 20) assert_eq(res.base_mean.length(), 20) assert_eq(res.log2_fold_change.length(), 20) @@ -86,7 +86,7 @@ test "test_results" { let dds_sf = @src.estimate_size_factors(dds) let dds_disp = @src.estimate_dispersions(dds_sf) let res = @src.results(dds_disp) - + assert_eq(res.row_names.length(), 20) assert_eq(res.base_mean.length(), 20) assert_eq(res.log2_fold_change.length(), 20) @@ -94,11 +94,11 @@ test "test_results" { assert_eq(res.stat.length(), 20) assert_eq(res.p_value.length(), 20) assert_eq(res.padj.length(), 20) - + for p in res.p_value { assert_true(p >= 0.0 && p <= 1.0) } - + for p in res.padj { assert_true(p >= 0.0 && p <= 1.0) } @@ -111,10 +111,10 @@ test "test_lfc_shrink" { let dds_disp = @src.estimate_dispersions(dds_sf) let res = @src.results(dds_disp) let res_shrunk = @src.lfc_shrink(dds_disp, res) - + assert_eq(res_shrunk.row_names.length(), 20) assert_eq(res_shrunk.log2_fold_change.length(), 20) - + for i = 0; i < 20; i = i + 1 { assert_true(res_shrunk.lfc_se[i] <= res.lfc_se[i]) } @@ -126,12 +126,14 @@ test "test_significant_genes" { let dds_sf = @src.estimate_size_factors(dds) let dds_disp = @src.estimate_dispersions(dds_sf) let res = @src.results(dds_disp) - + let sig_genes = @src.significant_genes(res) assert_true(sig_genes.length() >= 0 && sig_genes.length() <= 20) - + let sig_genes_lfc = @src.significant_genes(res, alpha=0.05, lfc_threshold=0.5) - assert_true(sig_genes_lfc.length() >= 0 && sig_genes_lfc.length() <= sig_genes.length()) + assert_true( + sig_genes_lfc.length() >= 0 && sig_genes_lfc.length() <= sig_genes.length(), + ) } ///| @@ -140,11 +142,11 @@ test "test_top_genes" { let dds_sf = @src.estimate_size_factors(dds) let dds_disp = @src.estimate_dispersions(dds_sf) let res = @src.results(dds_disp) - + let top = @src.top_genes(res, n=5) assert_eq(top.length(), 5) - + for i = 0; i < 4; i = i + 1 { assert_true(top[i].2 <= top[i + 1].2) } -} \ No newline at end of file +} diff --git a/test/moonbit/destiny_test.mbt b/test/moonbit/destiny_test.mbt index f420a8f5..d0f8b74f 100644 --- a/test/moonbit/destiny_test.mbt +++ b/test/moonbit/destiny_test.mbt @@ -1,11 +1,7 @@ ///| test "destiny_create_cell_data" { - let cell = @src.CellData::new( - "cell_001", - [1.0, 2.0, 3.0], - "cluster_1" - ) - + let cell = @src.CellData::new("cell_001", [1.0, 2.0, 3.0], "cluster_1") + assert_eq(cell.cell_id, "cell_001") assert_eq(cell.expression.length(), 3) assert_eq(cell.cluster, "cluster_1") @@ -15,13 +11,15 @@ test "destiny_create_cell_data" { test "destiny_distance_matrix" { let cells = @src.create_example_sc_data(5, 3) let dist_matrix = @src.compute_distance_matrix(cells, "euclidean") - + assert_eq(dist_matrix.n_cells, 5) assert_eq(dist_matrix.cells.length(), 5) - + // Check symmetry - assert_true((dist_matrix.distances[0][1] - dist_matrix.distances[1][0]).abs() < 1.0e-10) - + assert_true( + (dist_matrix.distances[0][1] - dist_matrix.distances[1][0]).abs() < 1.0e-10, + ) + // Check diagonal is zero assert_true(dist_matrix.distances[0][0].abs() < 1.0e-10) } @@ -30,7 +28,7 @@ test "destiny_distance_matrix" { test "destiny_manhattan_distance" { let cells = @src.create_example_sc_data(3, 2) let dist_matrix = @src.compute_distance_matrix(cells, "manhattan") - + assert_eq(dist_matrix.n_cells, 3) assert_true(dist_matrix.distances[0][1] >= 0.0) } @@ -40,14 +38,14 @@ test "destiny_gaussian_kernel" { let cells = @src.create_example_sc_data(5, 3) let dist_matrix = @src.compute_distance_matrix(cells, "euclidean") let kernel = @src.compute_gaussian_kernel(dist_matrix, 1.0) - + assert_eq(kernel.cells.length(), 5) assert_eq(kernel.bandwidth, 1.0) - + // Check kernel values are between 0 and 1 assert_true(kernel.kernel[0][0] <= 1.0) assert_true(kernel.kernel[0][0] >= 0.0) - + // Diagonal should be 1 (exp(0)) assert_true((kernel.kernel[0][0] - 1.0).abs() < 1.0e-10) } @@ -57,7 +55,7 @@ test "destiny_find_sigma" { let cells = @src.create_example_sc_data(10, 3) let dist_matrix = @src.compute_distance_matrix(cells, "euclidean") let sigma = @src.find_sigma_automatic(dist_matrix) - + assert_true(sigma > 0.0) } @@ -65,12 +63,12 @@ test "destiny_find_sigma" { test "destiny_diffusion_map" { let cells = @src.create_example_sc_data(20, 5) let result = @src.compute_diffusion_map(cells, 3, 1.0) - + assert_eq(result.cell_ids.length(), 20) assert_eq(result.eigenvalues.length(), 3) assert_eq(result.embedding.length(), 20) assert_eq(result.explained_variance.length(), 3) - + // Check eigenvalues are non-negative let mut i = 0 while i < 3 { @@ -83,7 +81,7 @@ test "destiny_diffusion_map" { test "destiny_auto_embedding" { let cells = @src.create_example_sc_data(15, 4) let result = @src.destiny_create_embedding(cells, 2) - + assert_eq(result.cell_ids.length(), 15) assert_eq(result.embedding[0].length(), 2) } @@ -93,7 +91,7 @@ test "destiny_plot_coordinates" { let cells = @src.create_example_sc_data(10, 3) let result = @src.destiny_create_embedding(cells, 2) let coords = @src.destiny_plot_coordinates(result, 0, 1) - + assert_eq(coords.length(), 10) } @@ -102,7 +100,7 @@ test "destiny_summary" { let cells = @src.create_example_sc_data(8, 3) let result = @src.destiny_create_embedding(cells, 2) let summary = @src.destiny_summary(result) - + assert_true(summary.contains("Diffusion Map Summary")) assert_true(summary.contains("Number of cells:")) } @@ -110,17 +108,19 @@ test "destiny_summary" { ///| test "destiny_create_sc_data" { let cells = @src.create_example_sc_data(10, 3) - + assert_eq(cells.length(), 10) assert_eq(cells[0].expression.length(), 3) - assert_true(cells[0].cluster == "cluster_1" || cells[0].cluster == "cluster_2") + assert_true( + cells[0].cluster == "cluster_1" || cells[0].cluster == "cluster_2", + ) } ///| test "destiny_euclidean_distance_zero" { let cells = @src.create_example_sc_data(3, 2) let dist_matrix = @src.compute_distance_matrix(cells, "euclidean") - + // Same point should have zero distance assert_true(dist_matrix.distances[0][0].abs() < 1.0e-10) assert_true(dist_matrix.distances[1][1].abs() < 1.0e-10) @@ -130,7 +130,7 @@ test "destiny_euclidean_distance_zero" { test "destiny_small_dataset" { let cells = @src.create_example_sc_data(2, 2) let result = @src.compute_diffusion_map(cells, 1, 0.5) - + assert_eq(result.cell_ids.length(), 2) assert_eq(result.embedding.length(), 2) } diff --git a/test/moonbit/dexseq_test.mbt b/test/moonbit/dexseq_test.mbt index 04cbc660..d4e94b37 100644 --- a/test/moonbit/dexseq_test.mbt +++ b/test/moonbit/dexseq_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for DEXSeq module. - test "ExonCount creation" { let ec = @src.ExonCount::new("gene1", "exon1", [10, 20, 30]) assert_eq(ec.gene_id, "gene1") @@ -8,6 +7,7 @@ test "ExonCount creation" { assert_eq(ec.counts.length(), 3) } +///| test "DEXSeqDataSet creation" { let exon_counts : Array[@src.ExonCount] = Array::new() exon_counts.push(@src.ExonCount::new("gene1", "exon1", [10, 12, 15])) @@ -18,18 +18,21 @@ test "DEXSeqDataSet creation" { assert_eq(ds.gene_ids[0], "gene1") } +///| test "dexseq_normalize_counts" { let ds = @src.create_example_dexseq_dataset() let normalized = @src.dexseq_normalize_counts(ds) assert_eq(normalized.exon_counts.length(), ds.exon_counts.length()) } +///| test "dexseq_test_for_exon_usage" { let ds = @src.create_example_dexseq_dataset() let results = @src.dexseq_test_for_exon_usage(ds) assert_true(results.length() > 0) } +///| test "dexseq_filter_results" { let ds = @src.create_example_dexseq_dataset() let results = @src.dexseq_test_for_exon_usage(ds) diff --git a/test/moonbit/diffcyt_test.mbt b/test/moonbit/diffcyt_test.mbt index 3a9b7389..20852a5c 100644 --- a/test/moonbit/diffcyt_test.mbt +++ b/test/moonbit/diffcyt_test.mbt @@ -10,9 +10,7 @@ ///| test "dc_cell_creation" { - let cell = @src.CytometryCell::new( - "cell_1", "S1", "ctrl", [1.0, 2.0, 3.0] - ) + let cell = @src.CytometryCell::new("cell_1", "S1", "ctrl", [1.0, 2.0, 3.0]) assert_eq(cell.cell_id(), "cell_1") assert_eq(cell.sample_id(), "S1") assert_eq(cell.condition(), "ctrl") @@ -85,7 +83,9 @@ test "dc_assign_clusters_new" { let cells = @src.diffcyt_sample_data() let codebooks = @src.diffcyt_cluster_cells(cells, 4, 10) // Re-assign using existing codebooks. - let new_cells = [@src.CytometryCell::new("nc1", "S1", "ctrl", cells[0].marker_values())] + let new_cells = [ + @src.CytometryCell::new("nc1", "S1", "ctrl", cells[0].marker_values()), + ] @src.diffcyt_assign_clusters(new_cells, codebooks) assert_true(new_cells[0].cluster_id() >= 0) assert_true(new_cells[0].cluster_id() < 4) @@ -121,7 +121,9 @@ test "dc_calc_medians" { let cells = @src.diffcyt_sample_data() let _ = @src.diffcyt_cluster_cells(cells, 3, 10) let samples = @src.diffcyt_unique_samples(cells) - let medians = @src.diffcyt_calc_medians_by_cluster_marker(cells, 3, 3, samples) + let medians = @src.diffcyt_calc_medians_by_cluster_marker( + cells, 3, 3, samples, + ) assert_eq(medians.length(), 4) // 4 samples // Each row has 3 clusters * 3 markers = 9 entries. for row in medians { @@ -219,7 +221,9 @@ test "dc_testDS_returns_results" { conditions.push(cond_map.get(sid).unwrap_or("")) } let marker_names = ["marker_0", "marker_1", "marker_2"] - let ds_results = @src.diffcyt_testDS(cells, samples, conditions, 3, 3, marker_names) + let ds_results = @src.diffcyt_testDS( + cells, samples, conditions, 3, 3, marker_names, + ) // 3 clusters * 3 markers = 9 results (if 2 conditions). assert_true(ds_results.length() > 0) for r in ds_results { @@ -244,7 +248,9 @@ test "dc_testDS_marker_names" { conditions.push(cond_map.get(sid).unwrap_or("")) } let marker_names = ["CD4", "CD8", "CD3"] - let ds_results = @src.diffcyt_testDS(cells, samples, conditions, 2, 3, marker_names) + let ds_results = @src.diffcyt_testDS( + cells, samples, conditions, 2, 3, marker_names, + ) for r in ds_results { assert_true(marker_names.contains(r.marker_name())) } @@ -311,9 +317,9 @@ test "dc_top_table_ds" { for sid in samples { conditions.push(cond_map.get(sid).unwrap_or("")) } - let ds_results = @src.diffcyt_testDS( - cells, samples, conditions, 3, 3, ["m0", "m1", "m2"] - ) + let ds_results = @src.diffcyt_testDS(cells, samples, conditions, 3, 3, [ + "m0", "m1", "m2", + ]) let top5 = @src.diffcyt_top_table_ds(ds_results, 5) assert_true(top5.length() <= 5) if top5.length() >= 2 { diff --git a/test/moonbit/dirichlet_multinomial_test.mbt b/test/moonbit/dirichlet_multinomial_test.mbt new file mode 100644 index 00000000..84b4d502 --- /dev/null +++ b/test/moonbit/dirichlet_multinomial_test.mbt @@ -0,0 +1,994 @@ +// Tests for the Bioconductor DirichletMultinomial-inspired module. + +///| +fn dmn_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "DirichletMultinomial value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn dmn_test_row_sum(values : Array[Double]) -> Double { + let mut total = 0.0 + for value in values { + total = total + value + } + total +} + +///| +fn dmn_test_counts() -> Array[Array[Int]] { + [ + [40, 3, 2], + [35, 4, 1], + [42, 2, 3], + [38, 5, 2], + [45, 3, 1], + [36, 2, 4], + [2, 40, 3], + [4, 35, 2], + [3, 43, 1], + [5, 38, 2], + [1, 44, 3], + [3, 36, 4], + ] +} + +///| +fn dmn_test_groups() -> Array[String] { + ["A", "A", "A", "A", "A", "A", "B", "B", "B", "B", "B", "B"] +} + +///| +fn dmn_test_config(components : Int) -> @src.DmnConfig { + @src.DmnConfig::create( + components~, + max_iterations=40, + optimizer_iterations=100, + soft_kmeans_iterations=50, + tolerance=1.0e-5, + optimizer_tolerance=1.0e-5, + seed=2026, + ) catch { + DirichletMultinomialError(message) => + abort("valid DirichletMultinomial configuration failed: " + message) + } +} + +///| +fn dmn_test_fit(components : Int) -> @src.DmnFit { + @src.dirichlet_multinomial_fit( + dmn_test_counts(), + taxon_names=["taxon_a", "taxon_b", "taxon_c"], + config=dmn_test_config(components), + ) catch { + DirichletMultinomialError(message) => + abort("DirichletMultinomial test fit failed: " + message) + } +} + +///| +test "dirichlet_multinomial: default configuration matches upstream controls" { + let config = @src.DmnConfig::default() + assert_eq(config.components, 1) + assert_eq(config.max_iterations, 100) + assert_eq(config.optimizer_iterations, 250) + assert_eq(config.soft_kmeans_iterations, 200) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.optimizer_tolerance, 1.0e-5) + assert_eq(config.soft_beta, 50.0) + assert_eq(config.seed, 42) + assert_eq(config.prior_shape, 0.1) + assert_eq(config.prior_rate, 0.1) + assert_eq(config.min_alpha, 1.0e-8) +} + +///| +test "dirichlet_multinomial: custom configuration preserves controls" { + let config = @src.DmnConfig::create( + components=3, + max_iterations=9, + optimizer_iterations=11, + soft_kmeans_iterations=13, + tolerance=0.01, + optimizer_tolerance=0.02, + soft_beta=25.0, + seed=99, + prior_shape=0.2, + prior_rate=0.3, + min_alpha=1.0e-7, + ) catch { + _ => abort("custom DirichletMultinomial configuration should be valid") + } + assert_eq(config.components, 3) + assert_eq(config.max_iterations, 9) + assert_eq(config.optimizer_iterations, 11) + assert_eq(config.soft_kmeans_iterations, 13) + assert_eq(config.soft_beta, 25.0) + assert_eq(config.seed, 99) + assert_eq(config.prior_shape, 0.2) + assert_eq(config.prior_rate, 0.3) +} + +///| +test "dirichlet_multinomial: with_components preserves optimization controls" { + let original = dmn_test_config(1) + let changed = original.with_components(2, seed=77) catch { + _ => abort("with_components should accept positive component counts") + } + assert_eq(changed.components, 2) + assert_eq(changed.seed, 77) + assert_eq(changed.max_iterations, original.max_iterations) + assert_eq(changed.optimizer_iterations, original.optimizer_iterations) + assert_eq(changed.soft_beta, original.soft_beta) +} + +///| +test "dirichlet_multinomial: configuration rejects nonpositive components" { + let raised = try { + ignore(@src.DmnConfig::create(components=0)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: configuration rejects invalid iteration limits" { + let em = try { + ignore(@src.DmnConfig::create(max_iterations=0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let optimizer = try { + ignore(@src.DmnConfig::create(optimizer_iterations=0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let kmeans = try { + ignore(@src.DmnConfig::create(soft_kmeans_iterations=0)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(em) + assert_true(optimizer) + assert_true(kmeans) +} + +///| +test "dirichlet_multinomial: configuration rejects invalid tolerances" { + let em = try { + ignore(@src.DmnConfig::create(tolerance=0.0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let optimizer = try { + ignore(@src.DmnConfig::create(optimizer_tolerance=-0.1)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(em) + assert_true(optimizer) +} + +///| +test "dirichlet_multinomial: configuration rejects invalid soft beta and seed" { + let beta = try { + ignore(@src.DmnConfig::create(soft_beta=0.0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let seed = try { + ignore(@src.DmnConfig::create(seed=0)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(beta) + assert_true(seed) +} + +///| +test "dirichlet_multinomial: configuration rejects invalid Gamma priors" { + let shape = try { + ignore(@src.DmnConfig::create(prior_shape=0.0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let rate = try { + ignore(@src.DmnConfig::create(prior_rate=-0.1)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(shape) + assert_true(rate) +} + +///| +test "dirichlet_multinomial: configuration rejects invalid alpha floor" { + let zero = try { + ignore(@src.DmnConfig::create(min_alpha=0.0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let one = try { + ignore(@src.DmnConfig::create(min_alpha=1.0)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(zero) + assert_true(one) +} + +///| +test "dirichlet_multinomial: known beta-binomial probability is exact" { + let value = @src.dirichlet_multinomial_log_pmf([2, 1], [1.0, 1.0]) catch { + _ => abort("valid Dirichlet-multinomial PMF should succeed") + } + dmn_test_close(value, -1.3862943611198906, 1.0e-10) +} + +///| +test "dirichlet_multinomial: PMF normalizes over two-count compositions" { + let mut total = 0.0 + for left in 0..<=2 { + let log_probability = @src.dirichlet_multinomial_log_pmf([left, 2 - left], [ + 1.0, 1.0, + ]) catch { + _ => abort("valid Dirichlet-multinomial PMF should succeed") + } + total = total + @math.exp(log_probability) + } + dmn_test_close(total, 1.0, 1.0e-10) +} + +///| +test "dirichlet_multinomial: PMF is symmetric under paired permutation" { + let left = @src.dirichlet_multinomial_log_pmf([4, 1], [2.0, 3.0]) catch { + _ => abort("valid Dirichlet-multinomial PMF should succeed") + } + let right = @src.dirichlet_multinomial_log_pmf([1, 4], [3.0, 2.0]) catch { + _ => abort("valid Dirichlet-multinomial PMF should succeed") + } + dmn_test_close(left, right, 1.0e-12) +} + +///| +test "dirichlet_multinomial: all-zero count vector has unit probability" { + let value = @src.dirichlet_multinomial_log_pmf([0, 0, 0], [0.5, 1.5, 2.5]) catch { + _ => abort("zero-count Dirichlet-multinomial PMF should succeed") + } + dmn_test_close(value, 0.0, 1.0e-12) +} + +///| +test "dirichlet_multinomial: PMF rejects dimension mismatch" { + let raised = try { + ignore(@src.dirichlet_multinomial_log_pmf([1, 2], [1.0])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: PMF rejects negative counts" { + let raised = try { + ignore(@src.dirichlet_multinomial_log_pmf([1, -1], [1.0, 1.0])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: PMF rejects nonpositive alpha" { + let zero = try { + ignore(@src.dirichlet_multinomial_log_pmf([1, 1], [1.0, 0.0])) + false + } catch { + DirichletMultinomialError(_) => true + } + let huge = try { + ignore(@src.dirichlet_multinomial_log_pmf([1, 1], [1.0, 1.0e301])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(zero) + assert_true(huge) +} + +///| +test "dirichlet_multinomial: mean follows alpha proportions" { + let mean = @src.dirichlet_multinomial_mean(10, [2.0, 3.0]) catch { + _ => abort("valid Dirichlet-multinomial mean should succeed") + } + assert_eq(mean.length(), 2) + dmn_test_close(mean[0], 4.0, 1.0e-12) + dmn_test_close(mean[1], 6.0, 1.0e-12) +} + +///| +test "dirichlet_multinomial: zero-total mean is zero" { + let mean = @src.dirichlet_multinomial_mean(0, [2.0, 3.0]) catch { + _ => abort("zero-total Dirichlet-multinomial mean should succeed") + } + assert_eq(mean, [0.0, 0.0]) +} + +///| +test "dirichlet_multinomial: mean validates total and alpha" { + let total = try { + ignore(@src.dirichlet_multinomial_mean(-1, [1.0])) + false + } catch { + DirichletMultinomialError(_) => true + } + let alpha = try { + ignore(@src.dirichlet_multinomial_mean(2, [1.0, -1.0])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(total) + assert_true(alpha) +} + +///| +test "dirichlet_multinomial: covariance includes overdispersion inflation" { + let covariance = @src.dirichlet_multinomial_covariance(10, [2.0, 3.0]) catch { + _ => abort("valid Dirichlet-multinomial covariance should succeed") + } + dmn_test_close(covariance[0][0], 6.0, 1.0e-12) + dmn_test_close(covariance[1][1], 6.0, 1.0e-12) + dmn_test_close(covariance[0][1], -6.0, 1.0e-12) + dmn_test_close(covariance[1][0], -6.0, 1.0e-12) +} + +///| +test "dirichlet_multinomial: zero-total covariance is zero" { + let covariance = @src.dirichlet_multinomial_covariance(0, [1.0, 2.0, 3.0]) catch { + _ => abort("zero-total covariance should succeed") + } + for row in covariance { + assert_eq(row, [0.0, 0.0, 0.0]) + } +} + +///| +test "dirichlet_multinomial: single-component fit exposes dimensions and names" { + let fit = dmn_test_fit(1) + assert_eq(fit.component_count(), 1) + assert_eq(fit.sample_count(), 12) + assert_eq(fit.taxon_count(), 3) + assert_eq(fit.sample_names[0], "sample1") + assert_eq(fit.sample_names[11], "sample12") + assert_eq(fit.taxon_names, ["taxon_a", "taxon_b", "taxon_c"]) +} + +///| +test "dirichlet_multinomial: single-component posterior and weight are one" { + let fit = dmn_test_fit(1) + assert_eq(fit.mixture_weights(), [1.0]) + for row in fit.mixture() { + dmn_test_close(row[0], 1.0, 1.0e-12) + } + assert_eq(fit.assignments(), Array::make(12, 0)) +} + +///| +test "dirichlet_multinomial: estimates and intervals remain positive" { + let fit = dmn_test_fit(1) + for taxon in 0.. 0.0) + assert_true(fit.lower[0][taxon] > 0.0) + assert_true(fit.upper[0][taxon] > 0.0) + assert_true(fit.lower[0][taxon] <= fit.upper[0][taxon]) + } +} + +///| +test "dirichlet_multinomial: component proportions and fitted scale normalize" { + let fit = dmn_test_fit(1) + let proportions = fit.component_proportions() + dmn_test_close(dmn_test_row_sum(proportions[0]), 1.0, 1.0e-12) + let scaled = fit.fitted(scale=true) + assert_eq(scaled.length(), 3) + assert_eq(scaled[0].length(), 1) + dmn_test_close(scaled[0][0] + scaled[1][0] + scaled[2][0], 1.0, 1.0e-12) +} + +///| +test "dirichlet_multinomial: fitted matrix transposes component alpha" { + let fit = dmn_test_fit(1) + let fitted = fit.fitted() + for taxon in 0..= weights[1]) + dmn_test_close(weights[0] + weights[1], 1.0, 1.0e-10) +} + +///| +test "dirichlet_multinomial: two-component posterior rows normalize" { + let fit = dmn_test_fit(2) + for row in fit.mixture() { + assert_eq(row.length(), 2) + dmn_test_close(dmn_test_row_sum(row), 1.0, 1.0e-10) + assert_true(row[0] >= 0.0 && row[0] <= 1.0) + assert_true(row[1] >= 0.0 && row[1] <= 1.0) + } +} + +///| +test "dirichlet_multinomial: synthetic clusters receive distinct assignments" { + let assignments = dmn_test_fit(2).assignments() + let first = assignments[0] + let second = assignments[6] + assert_true(first != second) + for sample in 0..<6 { + assert_eq(assignments[sample], first) + } + for sample in 6..<12 { + assert_eq(assignments[sample], second) + } +} + +///| +test "dirichlet_multinomial: synthetic components recover dominant taxa" { + let fit = dmn_test_fit(2) + let assignments = fit.assignments() + let proportions = fit.component_proportions() + let first_component = assignments[0] + let second_component = assignments[6] + assert_true(proportions[first_component][0] > proportions[first_component][1]) + assert_true( + proportions[second_component][1] > proportions[second_component][0], + ) +} + +///| +test "dirichlet_multinomial: prediction normalizes and classifies novel samples" { + let fit = dmn_test_fit(2) + let posterior = fit.predict([[50, 1, 1], [1, 50, 1]]) catch { + _ => abort("valid DirichletMultinomial prediction should succeed") + } + for row in posterior { + dmn_test_close(dmn_test_row_sum(row), 1.0, 1.0e-10) + } + let assignments = fit.predict_assignments([[50, 1, 1], [1, 50, 1]]) catch { + _ => abort("valid DirichletMultinomial assignment should succeed") + } + assert_true(assignments[0] != assignments[1]) +} + +///| +test "dirichlet_multinomial: component evidence has sample by component shape" { + let fit = dmn_test_fit(2) + let evidence = fit.negative_log_evidence([[50, 1, 1], [1, 50, 1]]) catch { + _ => abort("valid DirichletMultinomial evidence should succeed") + } + assert_eq(evidence.length(), 2) + assert_eq(evidence[0].length(), 2) + for row in evidence { + for value in row { + assert_true(value.abs() < 1.0e300) + } + } +} + +///| +test "dirichlet_multinomial: concentrations equal alpha row sums" { + let fit = dmn_test_fit(2) + let concentrations = fit.concentrations() + for component in 0.. true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: model selection fits every requested K" { + let selection = @src.dirichlet_multinomial_select( + dmn_test_counts(), + 2, + taxon_names=["taxon_a", "taxon_b", "taxon_c"], + config=dmn_test_config(1), + ) catch { + _ => abort("DirichletMultinomial model selection should succeed") + } + assert_eq(selection.fits.length(), 2) + assert_eq(selection.scores.length(), 2) + assert_eq(selection.fits[0].component_count(), 1) + assert_eq(selection.fits[1].component_count(), 2) + assert_eq(selection.criterion, @src.DmnLaplace) +} + +///| +test "dirichlet_multinomial: model selection best index minimizes score" { + let selection = @src.dirichlet_multinomial_select( + dmn_test_counts(), + 2, + criterion=@src.DmnBic, + config=dmn_test_config(1), + ) catch { + _ => abort("DirichletMultinomial BIC selection should succeed") + } + assert_eq(selection.criterion, @src.DmnBic) + assert_eq( + selection.best_fit().component_count(), + selection.best_component_count(), + ) + assert_true(selection.scores[selection.best_index] <= selection.scores[0]) + assert_true(selection.scores[selection.best_index] <= selection.scores[1]) +} + +///| +test "dirichlet_multinomial: model selection validates component range" { + let zero = try { + ignore(@src.dirichlet_multinomial_select(dmn_test_counts(), 0)) + false + } catch { + DirichletMultinomialError(_) => true + } + let excessive = try { + ignore(@src.dirichlet_multinomial_select(dmn_test_counts(), 13)) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(zero) + assert_true(excessive) +} + +///| +test "dirichlet_multinomial: fit rejects empty and ragged matrices" { + let empty = try { + ignore(@src.dirichlet_multinomial_fit([])) + false + } catch { + DirichletMultinomialError(_) => true + } + let ragged = try { + ignore(@src.dirichlet_multinomial_fit([[1, 2], [3]])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(empty) + assert_true(ragged) +} + +///| +test "dirichlet_multinomial: fit rejects negative and zero-library samples" { + let negative = try { + ignore(@src.dirichlet_multinomial_fit([[1, -1], [2, 3]])) + false + } catch { + DirichletMultinomialError(_) => true + } + let zero = try { + ignore(@src.dirichlet_multinomial_fit([[1, 2], [0, 0]])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(negative) + assert_true(zero) +} + +///| +test "dirichlet_multinomial: fit rejects too many components" { + let raised = try { + ignore( + @src.dirichlet_multinomial_fit( + [[1, 2], [2, 1]], + config=dmn_test_config(3), + ), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: fit validates sample names" { + let count = try { + ignore( + @src.dirichlet_multinomial_fit([[1, 2], [2, 1]], sample_names=["one"]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + let duplicate = try { + ignore( + @src.dirichlet_multinomial_fit([[1, 2], [2, 1]], sample_names=[ + "same", "same", + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(count) + assert_true(duplicate) +} + +///| +test "dirichlet_multinomial: fit validates taxon names" { + let empty = try { + ignore( + @src.dirichlet_multinomial_fit([[1, 2], [2, 1]], taxon_names=[ + "taxon", " ", + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + let duplicate = try { + ignore( + @src.dirichlet_multinomial_fit([[1, 2], [2, 1]], taxon_names=[ + "same", "same", + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(empty) + assert_true(duplicate) +} + +///| +test "dirichlet_multinomial: returned mixture arrays are defensive copies" { + let fit = dmn_test_fit(1) + let weights = fit.mixture_weights() + weights[0] = 0.0 + let mixture = fit.mixture() + mixture[0][0] = 0.0 + assert_eq(fit.mixture_weights(), [1.0]) + assert_eq(fit.mixture()[0], [1.0]) +} + +///| +test "dirichlet_multinomial: group fit preserves labels and empirical priors" { + let classifier = @src.dirichlet_multinomial_group_fit( + dmn_test_counts(), + dmn_test_groups(), + components_by_group=[1, 1], + taxon_names=["taxon_a", "taxon_b", "taxon_c"], + config=dmn_test_config(1), + ) catch { + _ => abort("DirichletMultinomial group fit should succeed") + } + assert_eq(classifier.group_names, ["A", "B"]) + assert_eq(classifier.models.length(), 2) + assert_eq(classifier.priors, [0.5, 0.5]) + assert_eq(classifier.taxon_names, ["taxon_a", "taxon_b", "taxon_c"]) +} + +///| +test "dirichlet_multinomial: group prediction normalizes and separates classes" { + let classifier = @src.dirichlet_multinomial_group_fit( + dmn_test_counts(), + dmn_test_groups(), + components_by_group=[1, 1], + config=dmn_test_config(1), + ) catch { + _ => abort("DirichletMultinomial group fit should succeed") + } + let probabilities = classifier.predict([[50, 1, 1], [1, 50, 1]]) catch { + _ => abort("DirichletMultinomial group prediction should succeed") + } + for row in probabilities { + dmn_test_close(dmn_test_row_sum(row), 1.0, 1.0e-10) + } + let predictions = classifier.predict_assignments([[50, 1, 1], [1, 50, 1]]) catch { + _ => abort("DirichletMultinomial group assignment should succeed") + } + assert_eq(predictions, ["A", "B"]) +} + +///| +test "dirichlet_multinomial: group fitting validates labels and component counts" { + let labels = try { + ignore( + @src.dirichlet_multinomial_group_fit([[1, 2], [2, 1]], ["A"], components_by_group=[ + 1, 1, + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + let one_group = try { + ignore( + @src.dirichlet_multinomial_group_fit([[1, 2], [2, 1]], ["A", "A"], components_by_group=[ + 1, + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + let components = try { + ignore( + @src.dirichlet_multinomial_group_fit([[1, 2], [2, 1]], ["A", "B"], components_by_group=[ + 1, + ]), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(labels) + assert_true(one_group) + assert_true(components) +} + +///| +test "dirichlet_multinomial: stratified cross-validation returns held-out results" { + let result = @src.dirichlet_multinomial_cross_validate( + dmn_test_counts(), + dmn_test_groups(), + 3, + components_by_group=[1, 1], + taxon_names=["taxon_a", "taxon_b", "taxon_c"], + config=dmn_test_config(1), + ) catch { + _ => abort("DirichletMultinomial cross-validation should succeed") + } + assert_eq(result.probabilities.length(), 12) + assert_eq(result.predictions.length(), 12) + assert_eq(result.truth, dmn_test_groups()) + assert_eq(result.group_names, ["A", "B"]) + assert_eq(result.fold_ids, [0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2]) + for row in result.probabilities { + dmn_test_close(dmn_test_row_sum(row), 1.0, 1.0e-10) + } + assert_true(result.accuracy >= 0.9) + assert_eq(result.correct.to_double() / 12.0, result.accuracy) +} + +///| +test "dirichlet_multinomial: cross-validation validates folds and group sizes" { + let folds = try { + ignore( + @src.dirichlet_multinomial_cross_validate( + dmn_test_counts(), + dmn_test_groups(), + 1, + ), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + let sizes = try { + ignore( + @src.dirichlet_multinomial_cross_validate( + [[5, 1], [4, 1], [1, 5], [1, 4]], + ["A", "A", "B", "B"], + 3, + ), + ) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(folds) + assert_true(sizes) +} + +///| +test "dirichlet_multinomial: perfect ROC has unit AUC" { + let roc = @src.dirichlet_multinomial_roc([true, true, false, false], [ + 0.9, 0.8, 0.2, 0.1, + ]) catch { + _ => abort("valid DirichletMultinomial ROC should succeed") + } + assert_eq(roc.auc, 1.0) + assert_eq(roc.best_threshold, 0.8) + assert_eq(roc.sensitivity[0], 0.0) + assert_eq(roc.specificity[0], 1.0) + assert_eq(roc.sensitivity[roc.sensitivity.length() - 1], 1.0) + assert_eq(roc.specificity[roc.specificity.length() - 1], 0.0) +} + +///| +test "dirichlet_multinomial: tied ROC scores contribute half concordance" { + let roc = @src.dirichlet_multinomial_roc([true, false], [0.5, 0.5]) catch { + _ => abort("tied DirichletMultinomial ROC should succeed") + } + assert_eq(roc.auc, 0.5) + assert_eq(roc.thresholds.length(), 2) +} + +///| +test "dirichlet_multinomial: ROC rejects invalid labels and scores" { + let lengths = try { + ignore(@src.dirichlet_multinomial_roc([true], [0.5, 0.2])) + false + } catch { + DirichletMultinomialError(_) => true + } + let classes = try { + ignore(@src.dirichlet_multinomial_roc([true, true], [0.5, 0.2])) + false + } catch { + DirichletMultinomialError(_) => true + } + let score = try { + ignore(@src.dirichlet_multinomial_roc([true, false], [1.0e301, 0.2])) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(lengths) + assert_true(classes) + assert_true(score) +} + +///| +test "dirichlet_multinomial: SummarizedExperiment entry transposes feature by sample assay" { + let assays : Map[String, Array[Array[Double]]] = Map([ + ( + "counts", + [ + [40.0, 35.0, 42.0, 2.0, 4.0, 3.0], + [3.0, 4.0, 2.0, 40.0, 35.0, 43.0], + [2.0, 1.0, 3.0, 3.0, 2.0, 1.0], + ], + ), + ]) + let experiment = @src.summarized_experiment(assays, [], [], Map([])) + let fit = @src.dirichlet_multinomial_fit_se( + experiment, + "counts", + sample_names=["s1", "s2", "s3", "s4", "s5", "s6"], + taxon_names=["taxon_a", "taxon_b", "taxon_c"], + config=dmn_test_config(2), + ) catch { + _ => abort("DirichletMultinomial SummarizedExperiment fit should succeed") + } + assert_eq(fit.sample_count(), 6) + assert_eq(fit.taxon_count(), 3) + assert_eq(fit.sample_names[0], "s1") + assert_eq(fit.taxon_names[2], "taxon_c") + assert_true(fit.assignments()[0] != fit.assignments()[3]) +} + +///| +test "dirichlet_multinomial: SummarizedExperiment entry rejects missing assay" { + let experiment = @src.summarized_experiment(Map([]), [], [], Map([])) + let raised = try { + ignore(@src.dirichlet_multinomial_fit_se(experiment, "counts")) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: SummarizedExperiment entry rejects noninteger assay" { + let assays : Map[String, Array[Array[Double]]] = Map([ + ("counts", [[1.5, 2.0], [3.0, 4.0]]), + ]) + let experiment = @src.summarized_experiment(assays, [], [], Map([])) + let raised = try { + ignore(@src.dirichlet_multinomial_fit_se(experiment, "counts")) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: SummarizedExperiment entry rejects negative assay" { + let assays : Map[String, Array[Array[Double]]] = Map([ + ("counts", [[1.0, -2.0], [3.0, 4.0]]), + ]) + let experiment = @src.summarized_experiment(assays, [], [], Map([])) + let raised = try { + ignore(@src.dirichlet_multinomial_fit_se(experiment, "counts")) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} + +///| +test "dirichlet_multinomial: SummarizedExperiment entry rejects ragged assay" { + let assays : Map[String, Array[Array[Double]]] = Map([ + ("counts", [[1.0, 2.0], [3.0]]), + ]) + let experiment = @src.summarized_experiment(assays, [], [], Map([])) + let raised = try { + ignore(@src.dirichlet_multinomial_fit_se(experiment, "counts")) + false + } catch { + DirichletMultinomialError(_) => true + } + assert_true(raised) +} diff --git a/test/moonbit/dnashape_test.mbt b/test/moonbit/dnashape_test.mbt index 0629d99b..6f2259ff 100644 --- a/test/moonbit/dnashape_test.mbt +++ b/test/moonbit/dnashape_test.mbt @@ -113,10 +113,10 @@ test "ds_roll_table" { assert_eq(t.length(), 16) // AA (index 0) = -6.0 let v0 = t.get(0).unwrap_or(999.0) - assert_true((v0 - (-6.0)).abs() < 0.001) + assert_true((v0 - -6.0).abs() < 0.001) // TT (index 15) = -6.0 let v15 = t.get(15).unwrap_or(999.0) - assert_true((v15 - (-6.0)).abs() < 0.001) + assert_true((v15 - -6.0).abs() < 0.001) // CG (index 6) = 6.0 let v6 = t.get(6).unwrap_or(999.0) assert_true((v6 - 6.0).abs() < 0.001) @@ -131,7 +131,7 @@ test "ds_prot_table" { assert_true((v0 - 15.0).abs() < 0.001) // TA = -2.0 let v12 = t.get(12).unwrap_or(0.0) - assert_true((v12 - (-2.0)).abs() < 0.001) + assert_true((v12 - -2.0).abs() < 0.001) } ///| @@ -164,10 +164,10 @@ test "ds_ep_table" { assert_eq(t.length(), 16) // AA = -1.5 let v0 = t.get(0).unwrap_or(0.0) - assert_true((v0 - (-1.5)).abs() < 0.001) + assert_true((v0 - -1.5).abs() < 0.001) // CG = -0.6 let v6 = t.get(6).unwrap_or(0.0) - assert_true((v6 - (-0.6)).abs() < 0.001) + assert_true((v6 - -0.6).abs() < 0.001) } // =========================================================================== diff --git a/test/moonbit/dorothea_test.mbt b/test/moonbit/dorothea_test.mbt index 06e55874..7b770a63 100644 --- a/test/moonbit/dorothea_test.mbt +++ b/test/moonbit/dorothea_test.mbt @@ -5,7 +5,7 @@ test "dorothea_get_regulons" { let regulons = @src.dorothea_get_regulons() assert_true(regulons.length() > 0) - + // Check that key TFs are present let mut has_tp53 = false let mut has_myc = false @@ -28,7 +28,9 @@ test "dorothea_regulon_targets" { assert_true(regulon.targets.length() > 0) // Check that all targets have valid directions for target in regulon.targets { - assert_true(target.direction == "activation" || target.direction == "repression") + assert_true( + target.direction == "activation" || target.direction == "repression", + ) } } } @@ -38,7 +40,7 @@ test "dorothea_compute_activity_simple" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -46,14 +48,16 @@ test "dorothea_compute_activity_simple" { cell_expr.push(counts[g][0]) g = g + 1 } - + // Find a regulon with targets in our gene set let mut tf_activity = 0.0 for regulon in regulons { - let activity = @src.dorothea_compute_activity_simple(cell_expr, gene_names, regulon) + let activity = @src.dorothea_compute_activity_simple( + cell_expr, gene_names, regulon, + ) tf_activity = tf_activity + activity } - + assert_true(tf_activity >= 0.0) } @@ -62,7 +66,7 @@ test "dorothea_compute_viper_activity" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -70,15 +74,17 @@ test "dorothea_compute_viper_activity" { cell_expr.push(counts[g][0]) g = g + 1 } - + let mut has_activity = false for regulon in regulons { - let activity = @src.dorothea_compute_viper_activity(cell_expr, gene_names, regulon) + let activity = @src.dorothea_compute_viper_activity( + cell_expr, gene_names, regulon, + ) if activity.abs() > 0.0 { has_activity = true } } - + // VIPER activities should generally be non-zero if targets are found assert_true(has_activity || regulons.length() > 0) } @@ -88,7 +94,7 @@ test "dorothea_permutation_test" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -96,7 +102,7 @@ test "dorothea_permutation_test" { cell_expr.push(counts[g][0]) g = g + 1 } - + // Use a regulon with many targets for better test let mut tp53_regulon = regulons[0] // Default for regulon in regulons { @@ -105,8 +111,10 @@ test "dorothea_permutation_test" { break } } - - let (p_value, z_score) = @src.dorothea_permutation_test(cell_expr, gene_names, tp53_regulon, 100) + + let (p_value, z_score) = @src.dorothea_permutation_test( + cell_expr, gene_names, tp53_regulon, 100, + ) assert_true(p_value >= 0.0 && p_value <= 1.0) } @@ -115,7 +123,7 @@ test "dorothea_analyze_cell" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -123,10 +131,12 @@ test "dorothea_analyze_cell" { cell_expr.push(counts[g][0]) g = g + 1 } - + let params = @src.DorotheaParams::with_params(50, 3, 0.05) - let results = @src.dorothea_analyze_cell(cell_expr, gene_names, regulons, params) - + let results = @src.dorothea_analyze_cell( + cell_expr, gene_names, regulons, params, + ) + assert_true(results.length() > 0) for result in results { assert_true(result.n_targets >= 3) @@ -139,10 +149,12 @@ test "dorothea_analyze_cells" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let params = @src.DorotheaParams::with_params(50, 3, 0.05) - let results = @src.dorothea_analyze_cells(counts, cell_ids, gene_names, regulons, params) - + let results = @src.dorothea_analyze_cells( + counts, cell_ids, gene_names, regulons, params, + ) + assert_true(results.length() > 0) assert_true(results.length() <= regulons.length()) } @@ -151,7 +163,7 @@ test "dorothea_analyze_cells" { test "dorothea_create_test_data" { let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + assert_eq(cell_ids.length(), 10) assert_true(gene_names.length() > 0) assert_true(counts.length() > 0) @@ -188,9 +200,9 @@ test "dorothea_sort_by_activity" { [], [], @src.dorothea_get_regulons(), - @src.DorotheaParams::with_params(10, 3, 0.05) + @src.DorotheaParams::with_params(10, 3, 0.05), ) - + // Sort empty or small results let sorted = @src.dorothea_sort_by_activity(results) assert_true(sorted.length() <= results.length()) @@ -201,10 +213,12 @@ test "dorothea_get_top_tfs" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let params = @src.DorotheaParams::with_params(20, 3, 0.05) - let results = @src.dorothea_analyze_cells(counts, cell_ids, gene_names, regulons, params) - + let results = @src.dorothea_analyze_cells( + counts, cell_ids, gene_names, regulons, params, + ) + let top3 = @src.dorothea_get_top_tfs(results, 3) assert_true(top3.length() <= 3) } @@ -228,7 +242,7 @@ test "dorothea_get_target_expression" { let regulons = @src.dorothea_get_regulons() let test_data = @src.dorothea_create_test_data() let (counts, cell_ids, gene_names) = test_data - + let cell_expr : Array[Double] = Array::new() let n_genes = counts.length() let mut g = 0 @@ -236,8 +250,12 @@ test "dorothea_get_target_expression" { cell_expr.push(counts[g][0]) g = g + 1 } - + // Test with first regulon - let targets = @src.dorothea_get_target_expression(cell_expr, gene_names, regulons[0]) + let targets = @src.dorothea_get_target_expression( + cell_expr, + gene_names, + regulons[0], + ) assert_true(targets.length() >= 0) -} \ No newline at end of file +} diff --git a/test/moonbit/dreamlet_test.mbt b/test/moonbit/dreamlet_test.mbt new file mode 100644 index 00000000..a4bd4a9a --- /dev/null +++ b/test/moonbit/dreamlet_test.mbt @@ -0,0 +1,932 @@ +///| +/// Tests for the Bioconductor dreamlet-inspired pseudobulk mixed-model +/// workflow. + +///| +fn dreamlet_test_close( + actual : Double, + expected : Double, + tolerance? : Double = 1.0e-8, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "dreamlet value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn dreamlet_test_data() -> ( + Array[Array[Double]], + Array[String], + Array[String], + Array[String], + Array[String], + Map[String, Array[String]], +) { + let genes = [ + "response_up", "stable", "cell_type", "rare", "donor", "response_down", + ] + let counts : Array[Array[Double]] = [] + for _ in genes { + counts.push([]) + } + let cells : Array[String] = [] + let samples : Array[String] = [] + let clusters : Array[String] = [] + let treatments : Array[String] = [] + let donors : Array[String] = [] + let ages : Array[String] = [] + for donor in 0..<4 { + for treatment in 0..<2 { + let sample = "D" + + (donor + 1).to_string() + + (if treatment == 0 { "_C" } else { "_T" }) + for cluster in 0..<2 { + let cluster_name = if cluster == 0 { "T_cell" } else { "Monocyte" } + for replicate in 0..<3 { + cells.push( + sample + "_" + cluster_name + "_" + (replicate + 1).to_string(), + ) + samples.push(sample) + clusters.push(cluster_name) + treatments.push(if treatment == 0 { "Control" } else { "Treated" }) + donors.push("D" + (donor + 1).to_string()) + ages.push((30 + donor * 5).to_string()) + counts[0].push( + (10 + treatment * 10 + donor + (if replicate == 2 { 1 } else { 0 })).to_double(), + ) + counts[1].push((8 + replicate % 2).to_double()) + counts[2].push( + (if cluster == 0 { 18 + treatment } else { 4 + treatment }).to_double(), + ) + counts[3].push( + if donor == 0 && treatment == 0 && cluster == 0 && replicate == 0 { + 1.0 + } else { + 0.0 + }, + ) + counts[4].push((6 + donor * 2 + replicate).to_double()) + counts[5].push( + (18 - treatment * 8 + cluster + (if replicate == 1 { 1 } else { 0 })).to_double(), + ) + } + } + } + } + let metadata : Map[String, Array[String]] = Map([ + ("Treatment", treatments), + ("Donor", donors), + ("Age", ages), + ]) + (counts, genes, cells, samples, clusters, metadata) +} + +///| +fn dreamlet_test_pseudobulk() -> @src.DreamletPseudobulk { + let (counts, genes, cells, samples, clusters, metadata) = dreamlet_test_data() + @src.dreamlet_aggregate_to_pseudobulk( + counts, + genes, + cells, + samples, + clusters, + cell_metadata=metadata, + ) catch { + _ => abort("valid dreamlet pseudobulk aggregation should succeed") + } +} + +///| +fn dreamlet_test_config() -> @src.DreamletProcessConfig { + @src.DreamletProcessConfig::create( + min_cells=2, + min_count=1.0, + min_samples=6, + min_prop=0.5, + min_total_count=10.0, + span=0.7, + max_iterations=30, + tolerance=0.02, + ) catch { + _ => abort("valid dreamlet test configuration should build") + } +} + +///| +fn dreamlet_test_fixed_model() -> @src.DreamletModelSpec { + @src.dreamlet_model([ + @src.dreamlet_categorical_effect("Treatment", reference="Control"), + ]) catch { + _ => abort("valid dreamlet fixed model should build") + } +} + +///| +fn dreamlet_test_mixed_model() -> @src.DreamletModelSpec { + @src.dreamlet_model([ + @src.dreamlet_categorical_effect("Treatment", reference="Control"), + @src.dreamlet_random_effect("Donor"), + ]) catch { + _ => abort("valid dreamlet mixed model should build") + } +} + +///| +fn dreamlet_test_processed(mixed? : Bool = false) -> @src.DreamletProcessedData { + let model = if mixed { + dreamlet_test_mixed_model() + } else { + dreamlet_test_fixed_model() + } + @src.dreamlet_process_assays( + dreamlet_test_pseudobulk(), + model, + config=dreamlet_test_config(), + ) catch { + _ => abort("valid dreamlet processing should succeed") + } +} + +///| +fn dreamlet_test_result() -> @src.DreamletResult { + @src.dreamlet( + dreamlet_test_processed(), + "Treatment:Treated", + contrast_name="Treated-Control", + max_iterations=30, + tolerance=0.02, + ) catch { + _ => abort("valid dreamlet differential expression should succeed") + } +} + +///| +test "dreamlet: effect factories preserve kinds and references" { + let numeric = @src.dreamlet_numeric_effect("Age") + let categorical = @src.dreamlet_categorical_effect( + "Treatment", + reference="Control", + ) + let random = @src.dreamlet_random_effect("Donor") + assert_true(numeric.kind is Numeric) + assert_true(categorical.kind is Categorical) + assert_true(random.kind is Random) + assert_eq(categorical.reference, "Control") +} + +///| +test "dreamlet: typed model preserves effect order" { + let model = dreamlet_test_mixed_model() + assert_eq(model.effects.length(), 2) + assert_eq(model.effects[0].name, "Treatment") + assert_eq(model.effects[1].name, "Donor") + assert_true(model.intercept) +} + +///| +test "dreamlet: model rejects duplicate effect names" { + let raised = try { + ignore( + @src.dreamlet_model([ + @src.dreamlet_numeric_effect("Age"), + @src.dreamlet_random_effect("Age"), + ]), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: model rejects empty effect names" { + let raised = try { + ignore(@src.dreamlet_model([@src.dreamlet_numeric_effect("")])) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: model rejects a no-intercept design without fixed effects" { + let raised = try { + ignore( + @src.dreamlet_model( + [@src.dreamlet_random_effect("Donor")], + intercept=false, + ), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: default process configuration matches upstream defaults" { + let config = @src.DreamletProcessConfig::default() + assert_eq(config.min_cells, 5) + assert_eq(config.min_count, 5.0) + assert_eq(config.min_samples, 4) + assert_eq(config.min_prop, 0.4) + assert_eq(config.prior_count, 0.5) + assert_eq(config.logratio_trim, 0.3) + assert_eq(config.sum_trim, 0.05) +} + +///| +test "dreamlet: custom process configuration preserves controls" { + let config = dreamlet_test_config() + assert_eq(config.min_cells, 2) + assert_eq(config.min_samples, 6) + assert_eq(config.span, 0.7) + assert_true(config.use_poisson_initial_weights) + assert_false(config.rescale_initial_weights) +} + +///| +test "dreamlet: process configuration rejects invalid sample filters" { + let cells = try { + ignore(@src.DreamletProcessConfig::create(min_cells=0)) + false + } catch { + DreamletError(_) => true + } + let samples = try { + ignore(@src.DreamletProcessConfig::create(min_samples=2)) + false + } catch { + DreamletError(_) => true + } + assert_true(cells) + assert_true(samples) +} + +///| +test "dreamlet: process configuration rejects invalid proportions" { + let low = try { + ignore(@src.DreamletProcessConfig::create(min_prop=0.0)) + false + } catch { + DreamletError(_) => true + } + let high = try { + ignore(@src.DreamletProcessConfig::create(min_prop=1.1)) + false + } catch { + DreamletError(_) => true + } + assert_true(low) + assert_true(high) +} + +///| +test "dreamlet: process configuration rejects invalid TMM trims" { + let ratio = try { + ignore(@src.DreamletProcessConfig::create(logratio_trim=0.5)) + false + } catch { + DreamletError(_) => true + } + let abundance = try { + ignore(@src.DreamletProcessConfig::create(sum_trim=-0.1)) + false + } catch { + DreamletError(_) => true + } + assert_true(ratio) + assert_true(abundance) +} + +///| +test "dreamlet: pseudobulk uses gene by sample orientation" { + let pseudobulk = dreamlet_test_pseudobulk() + assert_eq(pseudobulk.gene_names.length(), 6) + assert_eq(pseudobulk.sample_names.length(), 8) + assert_eq(pseudobulk.assays.length(), 2) + assert_eq(pseudobulk.assays[0].counts.length(), 6) + assert_eq(pseudobulk.assays[0].counts[0].length(), 8) +} + +///| +test "dreamlet: pseudobulk preserves first-observed cell type order" { + let pseudobulk = dreamlet_test_pseudobulk() + assert_eq(pseudobulk.assays[0].cluster_id, "T_cell") + assert_eq(pseudobulk.assays[1].cluster_id, "Monocyte") +} + +///| +test "dreamlet: pseudobulk preserves first-observed sample order" { + let pseudobulk = dreamlet_test_pseudobulk() + assert_eq(pseudobulk.sample_names[0], "D1_C") + assert_eq(pseudobulk.sample_names[1], "D1_T") + assert_eq(pseudobulk.sample_names[7], "D4_T") +} + +///| +test "dreamlet: pseudobulk sums raw counts within sample and cell type" { + let pseudobulk = dreamlet_test_pseudobulk() + let assay = match pseudobulk.assay("T_cell") { + Some(value) => value + None => abort("T_cell assay should exist") + } + dreamlet_test_close(assay.counts[0][0], 31.0) + dreamlet_test_close(assay.counts[0][1], 61.0) +} + +///| +test "dreamlet: pseudobulk records observed cell counts" { + let pseudobulk = dreamlet_test_pseudobulk() + for assay in pseudobulk.assays { + for count in assay.cell_counts { + assert_eq(count, 3) + } + } +} + +///| +test "dreamlet: pseudobulk library sizes equal column sums" { + let assay = dreamlet_test_pseudobulk().assays[0] + for sample in 0.. abort("sparse combinations should aggregate") + } + let b = match result.assay("B") { + Some(value) => value + None => abort("B assay should exist") + } + assert_eq(b.cell_counts, [0, 1]) + assert_eq(b.counts[0], [0.0, 5.0]) +} + +///| +test "dreamlet: pseudobulk summary reports dimensions" { + let summary = dreamlet_test_pseudobulk().summary() + assert_true(summary.contains("Genes: 6")) + assert_true(summary.contains("Samples: 8")) + assert_true(summary.contains("Cell types: 2")) +} + +///| +test "dreamlet: aggregation rejects negative counts" { + let raised = try { + ignore( + @src.dreamlet_aggregate_to_pseudobulk([[-1.0]], ["g"], ["c"], ["s"], ["A"]), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: aggregation rejects non-integer raw counts" { + let raised = try { + ignore( + @src.dreamlet_aggregate_to_pseudobulk([[1.5]], ["g"], ["c"], ["s"], ["A"]), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: aggregation rejects duplicate gene names" { + let raised = try { + ignore( + @src.dreamlet_aggregate_to_pseudobulk( + [[1.0], [2.0]], + ["g", "g"], + ["c"], + ["s"], + ["A"], + ), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: aggregation rejects metadata varying within sample" { + let metadata : Map[String, Array[String]] = Map([ + ("Treatment", ["Control", "Treated"]), + ]) + let raised = try { + ignore( + @src.dreamlet_aggregate_to_pseudobulk( + [[1.0, 2.0]], + ["g"], + ["c1", "c2"], + ["s", "s"], + ["A", "A"], + cell_metadata=metadata, + ), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: aggregation rejects metadata dimension mismatch" { + let metadata : Map[String, Array[String]] = Map([("Group", ["A"])]) + let raised = try { + ignore( + @src.dreamlet_aggregate_to_pseudobulk( + [[1.0, 2.0]], + ["g"], + ["c1", "c2"], + ["s1", "s2"], + ["A", "A"], + cell_metadata=metadata, + ), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: TMM factors have geometric mean one" { + let counts = [ + [100.0, 200.0, 90.0], + [50.0, 100.0, 55.0], + [30.0, 60.0, 25.0], + [10.0, 20.0, 12.0], + ] + let result = @src.dreamlet_tmm(counts) catch { + _ => abort("valid TMM normalization should succeed") + } + let mut product = 1.0 + for factor in result.factors { + product = product * factor + } + dreamlet_test_close(product, 1.0, tolerance=1.0e-10) +} + +///| +test "dreamlet: TMM identical compositions produce unit factors" { + let result = @src.dreamlet_tmm([ + [10.0, 20.0, 30.0], + [5.0, 10.0, 15.0], + [2.0, 4.0, 6.0], + ]) catch { + _ => abort("proportional libraries should normalize") + } + for factor in result.factors { + dreamlet_test_close(factor, 1.0, tolerance=1.0e-10) + } +} + +///| +test "dreamlet: TMM reference index is valid" { + let result = @src.dreamlet_tmm([ + [10.0, 12.0, 9.0], + [3.0, 7.0, 5.0], + [8.0, 4.0, 6.0], + ]) catch { + _ => abort("valid TMM normalization should succeed") + } + assert_true(result.reference_sample >= 0) + assert_true(result.reference_sample < 3) +} + +///| +test "dreamlet: TMM is invariant to a common count multiplier" { + let first = @src.dreamlet_tmm([ + [10.0, 12.0, 9.0], + [3.0, 7.0, 5.0], + [8.0, 4.0, 6.0], + ]) catch { + _ => abort("first TMM normalization should succeed") + } + let second = @src.dreamlet_tmm([ + [20.0, 24.0, 18.0], + [6.0, 14.0, 10.0], + [16.0, 8.0, 12.0], + ]) catch { + _ => abort("second TMM normalization should succeed") + } + for sample in 0..<3 { + dreamlet_test_close( + first.factors[sample], + second.factors[sample], + tolerance=1.0e-10, + ) + } +} + +///| +test "dreamlet: TMM rejects zero-size libraries" { + let raised = try { + ignore(@src.dreamlet_tmm([[1.0, 0.0], [2.0, 0.0]])) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: normalized counts are counts per million" { + let normalized = @src.dreamlet_compute_norm_counts( + [[10.0, 20.0], [90.0, 180.0]], + [100.0, 200.0], + ) catch { + _ => abort("valid normalized counts should succeed") + } + assert_eq(normalized[0], [100000.0, 100000.0]) + assert_eq(normalized[1], [900000.0, 900000.0]) +} + +///| +test "dreamlet: log CPM uses prior counts for zeros" { + let transformed = @src.dreamlet_compute_log_cpm( + [[0.0, 0.0]], + [100.0, 200.0], + prior_count=0.5, + ) catch { + _ => abort("valid log CPM should succeed") + } + assert_true(transformed[0][0] == transformed[0][0]) + assert_true(transformed[0][1] == transformed[0][1]) + assert_true(transformed[0][0] > transformed[0][1]) +} + +///| +test "dreamlet: count transformations reject invalid library sizes" { + let normalized = try { + ignore(@src.dreamlet_compute_norm_counts([[1.0]], [0.0])) + false + } catch { + DreamletError(_) => true + } + let logged = try { + ignore(@src.dreamlet_compute_log_cpm([[1.0]], [-1.0])) + false + } catch { + DreamletError(_) => true + } + assert_true(normalized) + assert_true(logged) +} + +///| +test "dreamlet: processAssays retains both well-powered cell types" { + let processed = dreamlet_test_processed() + assert_eq(processed.assays.length(), 2) + assert_eq(processed.exclusions.length(), 0) +} + +///| +test "dreamlet: processAssays filters rare genes" { + let assay = dreamlet_test_processed().assays[0] + assert_false(assay.gene_names.contains("rare")) + assert_true(assay.dropped_gene_names.contains("rare")) + assert_eq(assay.gene_names.length(), 5) +} + +///| +test "dreamlet: processAssays builds treatment design" { + let assay = dreamlet_test_processed().assays[0] + assert_eq(assay.design.coefficient_names, ["(Intercept)", "Treatment:Treated"]) + assert_eq(assay.design.sample_names, assay.sample_names) +} + +///| +test "dreamlet: mixed processAssays reuses random-effect solver" { + let assay = dreamlet_test_processed(mixed=true).assays[0] + assert_eq(assay.design.random_effects.length(), 1) + assert_eq(assay.design.random_effects[0].name, "Donor") + assert_eq(assay.design.random_effects[0].levels.length(), 4) +} + +///| +test "dreamlet: processAssays stores TMM effective libraries" { + let assay = dreamlet_test_processed().assays[0] + assert_eq(assay.norm_factors.length(), 8) + assert_eq(assay.effective_library_sizes.length(), 8) + for sample in 0..<8 { + dreamlet_test_close( + assay.effective_library_sizes[sample], + assay.library_sizes[sample] * assay.norm_factors[sample], + ) + } +} + +///| +test "dreamlet: processAssays stores normalized expression dimensions" { + let assay = dreamlet_test_processed().assays[0] + assert_eq(assay.normalized_cpm.length(), assay.gene_names.length()) + assert_eq(assay.log_cpm.length(), assay.gene_names.length()) + assert_eq(assay.log_cpm[0].length(), assay.sample_names.length()) +} + +///| +test "dreamlet: Poisson initial weights are row-centered" { + let assay = dreamlet_test_processed().assays[0] + for row in assay.initial_weights { + let mut total = 0.0 + for weight in row { + assert_true(weight > 0.0) + total = total + weight + } + dreamlet_test_close(total / row.length().to_double(), 1.0) + } +} + +///| +test "dreamlet: voom precision weights are finite and positive" { + let assay = dreamlet_test_processed().assays[0] + assert_eq(assay.precision_weights.length(), assay.gene_names.length()) + for row in assay.precision_weights { + assert_eq(row.length(), assay.sample_names.length()) + for weight in row { + assert_true(weight > 0.0) + assert_true(weight <= 1.0e8) + assert_true(weight == weight) + } + } +} + +///| +test "dreamlet: voom trend coordinates are sorted" { + let assay = dreamlet_test_processed().assays[0] + assert_eq(assay.voom_x.length(), assay.gene_names.length()) + assert_eq(assay.voom_y.length(), assay.gene_names.length()) + assert_eq(assay.voom_fitted.length(), assay.gene_names.length()) + for index in 1.. abort("constant-effect model should build") + } + let pseudobulk = dreamlet_test_pseudobulk() + for assay in pseudobulk.assays { + assay.metadata["Constant"] = Array::make(8, "1") + } + let processed = @src.dreamlet_process_assays( + pseudobulk, + model, + config=dreamlet_test_config(), + ) catch { + _ => abort("constant effects should be dropped") + } + assert_true(processed.assays[0].dropped_effects[0].contains("Constant")) +} + +///| +test "dreamlet: processAssays reports missing model metadata" { + let model = @src.dreamlet_model([@src.dreamlet_numeric_effect("Missing")]) catch { + _ => abort("model itself should build") + } + let raised = try { + ignore( + @src.dreamlet_process_assays( + dreamlet_test_pseudobulk(), + model, + config=dreamlet_test_config(), + ), + ) + false + } catch { + DreamletError(message) => message.contains("retained no assays") + } + assert_true(raised) +} + +///| +test "dreamlet: processed summary reports model count" { + let summary = dreamlet_test_processed().summary() + assert_true(summary.contains("Retained cell types: 2")) + assert_true(summary.contains("Excluded cell types: 0")) + assert_true(summary.contains("Gene-cell type models: 10")) +} + +///| +test "dreamlet: differential expression fits every retained cell type" { + let result = dreamlet_test_result() + assert_eq(result.assays.length(), 2) + assert_eq(result.genes.length(), 10) + assert_eq(result.contrast_name, "Treated-Control") +} + +///| +test "dreamlet: response-up effect is positive in each cell type" { + let result = dreamlet_test_result() + let mut found = 0 + for gene in result.genes { + if gene.gene_name == "response_up" { + assert_true(gene.estimate > 0.5) + found = found + 1 + } + } + assert_eq(found, 2) +} + +///| +test "dreamlet: response-down effect is negative in each cell type" { + let result = dreamlet_test_result() + let mut found = 0 + for gene in result.genes { + if gene.gene_name == "response_down" { + assert_true(gene.estimate < -0.3) + found = found + 1 + } + } + assert_eq(found, 2) +} + +///| +test "dreamlet: result stores within-cluster and study-wide FDR" { + let result = dreamlet_test_result() + for gene in result.genes { + assert_true(gene.within_cluster_fdr >= gene.p_value) + assert_true(gene.study_wide_fdr >= gene.p_value) + assert_true(gene.study_wide_fdr <= 1.0) + } +} + +///| +test "dreamlet: top table sorts all cell types by p-value" { + let table = dreamlet_test_result().top_table() catch { + _ => abort("valid dreamlet top table should succeed") + } + assert_eq(table.length(), 10) + for index in 1.. abort("cell-type dreamlet top table should succeed") + } + assert_eq(table.length(), 5) + for gene in table { + assert_eq(gene.cluster_id, "Monocyte") + } +} + +///| +test "dreamlet: result assay lookup exposes underlying dream result" { + let result = dreamlet_test_result() + match result.assay("T_cell") { + Some(value) => { + assert_eq(value.contrast_name, "Treated-Control") + assert_eq(value.genes.length(), 5) + } + None => abort("T_cell dream result should exist") + } + assert_true(result.assay("missing") is None) +} + +///| +test "dreamlet: result summary reports study-wide analysis" { + let summary = dreamlet_test_result().summary() + assert_true(summary.contains("dreamlet Differential Expression Summary")) + assert_true(summary.contains("Contrast: Treated-Control")) + assert_true(summary.contains("Cell types: 2")) + assert_true(summary.contains("Tests: 10")) +} + +///| +test "dreamlet: differential expression rejects absent coefficient" { + let raised = try { + ignore( + @src.dreamlet( + dreamlet_test_processed(), + "Missing", + max_iterations=10, + tolerance=0.05, + ), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: top table validates FDR threshold" { + let raised = try { + ignore(dreamlet_test_result().top_table(maximum_fdr=1.1)) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: SCE integration preserves selected metadata" { + let (counts, genes, cells, samples, clusters, metadata) = dreamlet_test_data() + let sce = @src.SingleCellExperiment::new(counts, genes, cells) + sce.col_data["sample_id"] = samples + sce.col_data["cluster_id"] = clusters + sce.col_data["Treatment"] = metadata["Treatment"] + sce.col_data["Donor"] = metadata["Donor"] + let pseudobulk = @src.dreamlet_aggregate_sce(sce, "sample_id", "cluster_id", metadata_fields=[ + "Treatment", "Donor", + ]) catch { + _ => abort("valid SCE dreamlet aggregation should succeed") + } + assert_eq(pseudobulk.assays.length(), 2) + assert_eq(pseudobulk.assays[0].metadata["Treatment"][1], "Treated") +} + +///| +test "dreamlet: SCE integration rejects absent assay" { + let sce = @src.SingleCellExperiment::new([[1.0]], ["g"], ["c"]) + sce.col_data["sample"] = ["s"] + sce.col_data["cluster"] = ["A"] + let raised = try { + ignore( + @src.dreamlet_aggregate_sce(sce, "sample", "cluster", assay="missing"), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} + +///| +test "dreamlet: SCE integration rejects absent metadata field" { + let sce = @src.SingleCellExperiment::new([[1.0]], ["g"], ["c"]) + sce.col_data["sample"] = ["s"] + sce.col_data["cluster"] = ["A"] + let raised = try { + ignore( + @src.dreamlet_aggregate_sce(sce, "sample", "cluster", metadata_fields=[ + "Missing", + ]), + ) + false + } catch { + DreamletError(_) => true + } + assert_true(raised) +} diff --git a/test/moonbit/drimseq_test.mbt b/test/moonbit/drimseq_test.mbt index e372d7e3..e99dae90 100644 --- a/test/moonbit/drimseq_test.mbt +++ b/test/moonbit/drimseq_test.mbt @@ -1,11 +1,13 @@ ///| /// Test file for DRIMSeq module. - test "drimseq_transcript_count_creation" { let tc = @src.TranscriptCount::new( - transcript_id="tx1", gene_id="GeneA", - sample_id="s1", condition="ctrl", - count=100.0, gene_count=150.0, + transcript_id="tx1", + gene_id="GeneA", + sample_id="s1", + condition="ctrl", + count=100.0, + gene_count=150.0, ) assert_eq(tc.transcript_id, "tx1") assert_eq(tc.gene_id, "GeneA") @@ -13,47 +15,77 @@ test "drimseq_transcript_count_creation" { assert_true((tc.proportion() - 0.666666).abs() < 0.01) } +///| test "drimseq_transcript_count_proportion_zero" { let tc = @src.TranscriptCount::new( - transcript_id="tx1", gene_id="GeneA", - sample_id="s1", condition="ctrl", + transcript_id="tx1", + gene_id="GeneA", + sample_id="s1", + condition="ctrl", count=100.0, ) assert_eq(tc.proportion(), 0.0) } +///| test "drimseq_filter_counts" { let counts : Array[@src.TranscriptCount] = Array::new() - counts.push(@src.TranscriptCount::new( - transcript_id="tx1", gene_id="G1", - sample_id="s1", condition="c1", - count=50.0, gene_count=100.0, - )) - counts.push(@src.TranscriptCount::new( - transcript_id="tx2", gene_id="G2", - sample_id="s1", condition="c1", - count=5.0, gene_count=100.0, - )) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx1", + gene_id="G1", + sample_id="s1", + condition="c1", + count=50.0, + gene_count=100.0, + ), + ) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx2", + gene_id="G2", + sample_id="s1", + condition="c1", + count=5.0, + gene_count=100.0, + ), + ) let filtered = @src.drimseq_filter_counts(counts) assert_eq(filtered.length(), 1) assert_eq(filtered[0].transcript_id, "tx1") } +///| test "drimseq_aggregate_by_gene" { let counts : Array[@src.TranscriptCount] = Array::new() - counts.push(@src.TranscriptCount::new( - transcript_id="tx1", gene_id="G1", - sample_id="s1", condition="c1", count=50.0, - )) - counts.push(@src.TranscriptCount::new( - transcript_id="tx2", gene_id="G1", - sample_id="s1", condition="c1", count=30.0, - )) - counts.push(@src.TranscriptCount::new( - transcript_id="tx1", gene_id="G1", - sample_id="s2", condition="c1", count=40.0, - )) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx1", + gene_id="G1", + sample_id="s1", + condition="c1", + count=50.0, + ), + ) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx2", + gene_id="G1", + sample_id="s1", + condition="c1", + count=30.0, + ), + ) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx1", + gene_id="G1", + sample_id="s2", + condition="c1", + count=40.0, + ), + ) let aggregated = @src.drimseq_aggregate_by_gene(counts) assert_eq(aggregated.length(), 2) // G1-s1 and G1-s2 @@ -61,49 +93,72 @@ test "drimseq_aggregate_by_gene" { assert_eq(aggregated[1].count, 40.0) } +///| test "drimseq_compute_proportions" { let counts : Array[@src.TranscriptCount] = Array::new() - counts.push(@src.TranscriptCount::new( - transcript_id="tx1", gene_id="G1", - sample_id="s1", condition="c1", count=50.0, - )) - counts.push(@src.TranscriptCount::new( - transcript_id="tx2", gene_id="G1", - sample_id="s1", condition="c1", count=50.0, - )) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx1", + gene_id="G1", + sample_id="s1", + condition="c1", + count=50.0, + ), + ) + counts.push( + @src.TranscriptCount::new( + transcript_id="tx2", + gene_id="G1", + sample_id="s1", + condition="c1", + count=50.0, + ), + ) let gene_counts : Array[@src.TranscriptCount] = Array::new() - gene_counts.push(@src.TranscriptCount::new( - transcript_id="__gene__", gene_id="G1", - sample_id="s1", condition="c1", count=100.0, - )) + gene_counts.push( + @src.TranscriptCount::new( + transcript_id="__gene__", + gene_id="G1", + sample_id="s1", + condition="c1", + count=100.0, + ), + ) let props = @src.drimseq_compute_proportions(counts, gene_counts) assert_eq(props.length(), 2) assert_true((props[0].proportion() - 0.5).abs() < 0.01) } +///| test "drimseq_wald_test_same" { let a : Array[Double] = Array::new() - a.push(0.5); a.push(0.5) + a.push(0.5) + a.push(0.5) let b : Array[Double] = Array::new() - b.push(0.5); b.push(0.5) + b.push(0.5) + b.push(0.5) let (stat, pval, _) = @src.drimseq_wald_test(a, b) assert_true(pval >= 0.99) // Should not be significant } +///| test "drimseq_wald_test_different" { let a : Array[Double] = Array::new() - a.push(0.8); a.push(0.2) + a.push(0.8) + a.push(0.2) let b : Array[Double] = Array::new() - b.push(0.2); b.push(0.8) + b.push(0.2) + b.push(0.8) let (_, pval, delta) = @src.drimseq_wald_test(a, b) assert_true(pval < 0.05) // Should be significant assert_true(delta.length() == 2) } +///| test "drimseq_result_summary" { let counts = @src.drimseq_sample_data() let result = @src.drimseq_test_differential(counts) @@ -111,6 +166,7 @@ test "drimseq_result_summary" { assert_true(summary.length() > 0) } +///| test "drimseq_result_get_significant" { let counts = @src.drimseq_sample_data() let result = @src.drimseq_test_differential(counts) @@ -118,6 +174,7 @@ test "drimseq_result_get_significant" { assert_true(sig.length() >= 0) } +///| test "drimseq_result_get_top_genes" { let counts = @src.drimseq_sample_data() let result = @src.drimseq_test_differential(counts) @@ -125,20 +182,21 @@ test "drimseq_result_get_top_genes" { assert_true(top.length() <= 2) } +///| test "drimseq_config_creation" { let cfg = @src.DRIMSeqConfig::new() assert_true(cfg.min_count > 0.0) assert_true(cfg.alpha > 0.0) } +///| test "drimseq_config_setters" { - let cfg = @src.DRIMSeqConfig::new() - .set_min_count(val=5.0) - .set_alpha(val=0.01) + let cfg = @src.DRIMSeqConfig::new().set_min_count(val=5.0).set_alpha(val=0.01) assert_eq(cfg.min_count, 5.0) assert_eq(cfg.alpha, 0.01) } +///| test "drimseq_norm_creation" { let n1 = @src.drimseq_norm_none() let n2 = @src.drimseq_norm_tmm() @@ -148,6 +206,7 @@ test "drimseq_norm_creation" { assert_true(n3 == n3) } +///| test "drimseq_full_pipeline" { let counts = @src.drimseq_sample_data() let result = @src.drimseq_test_differential(counts) @@ -156,6 +215,7 @@ test "drimseq_full_pipeline" { assert_true(result.converged) } +///| test "drimseq_empty_data" { let empty : Array[@src.TranscriptCount] = Array::new() let result = @src.drimseq_test_differential(empty) diff --git a/test/moonbit/droplet_utils_advanced_test.mbt b/test/moonbit/droplet_utils_advanced_test.mbt new file mode 100644 index 00000000..dad9f1ad --- /dev/null +++ b/test/moonbit/droplet_utils_advanced_test.mbt @@ -0,0 +1,1017 @@ +///| +fn dua_test_small_data() -> @src.DropletCountMatrix { + @src.DropletCountMatrix::create( + [ + [3, 2, 4, 3, 2, 3, 20, 18], + [2, 3, 1, 2, 3, 4, 0, 0], + [0, 0, 0, 0, 0, 0, 15, 0], + [0, 0, 0, 0, 0, 0, 0, 15], + ], + feature_names=["ambient_a", "ambient_b", "marker_a", "marker_b"], + barcodes=[ + "empty_1", "empty_2", "empty_3", "empty_4", "empty_5", "empty_6", "cell_a", + "cell_b", + ], + ) catch { + _ => abort("valid DropletUtils test data should build") + } +} + +///| +fn dua_test_config() -> @src.DropletUtilsConfig { + @src.DropletUtilsConfig::create( + lower=5, + iterations=199, + fdr_threshold=0.1, + retain=30.0, + rank_exclude=0, + rank_window=0.25, + seed=23, + ) catch { + _ => abort("valid DropletUtils configuration should build") + } +} + +///| +fn dua_test_result() -> @src.DropletEmptyDropsResult { + @src.droplet_utils_empty_drops( + dua_test_small_data(), + config=dua_test_config(), + ) catch { + _ => abort("valid DropletUtils workflow should run") + } +} + +///| +fn dua_test_sce() -> @src.SingleCellExperiment { + let data = dua_test_small_data() + let assay : Array[Array[Double]] = [] + for feature in 0.. abort("valid matrix should build") + } + assert_eq(data.feature_names, ["feature1", "feature2"]) + assert_eq(data.barcodes, ["barcode1", "barcode2"]) + assert_eq(data.n_features(), 2) + assert_eq(data.n_barcodes(), 2) +} + +///| +test "DropletUtils advanced constructor copies matrix and names" { + let counts = [[1, 2], [3, 4]] + let features = ["g1", "g2"] + let barcodes = ["b1", "b2"] + let data = @src.DropletCountMatrix::create( + counts, + feature_names=features, + barcodes~, + ) catch { + _ => abort("valid matrix should build") + } + counts[0][0] = 99 + features[0] = "changed" + barcodes[0] = "changed" + assert_eq(data.counts[0][0], 1) + assert_eq(data.feature_names[0], "g1") + assert_eq(data.barcodes[0], "b1") +} + +///| +test "DropletUtils advanced copy_counts is defensive" { + let data = dua_test_small_data() + let copied = data.copy_counts() + copied[0][0] = 999 + assert_eq(data.counts[0][0], 3) +} + +///| +test "DropletUtils advanced computes feature and barcode totals" { + let data = dua_test_small_data() + assert_eq(data.barcode_totals(), [5, 5, 5, 5, 5, 7, 35, 33]) + assert_eq(data.feature_totals(), [55, 15, 15, 15]) +} + +///| +test "DropletUtils advanced rejects an empty matrix" { + let raised = try { + ignore(@src.DropletCountMatrix::create([])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced rejects an empty barcode dimension" { + let raised = try { + ignore(@src.DropletCountMatrix::create([[]])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced rejects ragged matrices" { + let raised = try { + ignore(@src.DropletCountMatrix::create([[1, 2], [3]])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced rejects negative counts" { + let raised = try { + ignore(@src.DropletCountMatrix::create([[1, -1], [3, 4]])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced validates identifier lengths" { + let raised = try { + ignore( + @src.DropletCountMatrix::create([[1, 2], [3, 4]], feature_names=["g1"], barcodes=[ + "b1", "b2", + ]), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced rejects empty and duplicate identifiers" { + let empty = try { + ignore( + @src.DropletCountMatrix::create([[1, 2], [3, 4]], feature_names=["g1", ""]), + ) + false + } catch { + _ => true + } + let duplicate = try { + ignore( + @src.DropletCountMatrix::create([[1, 2], [3, 4]], barcodes=["b1", "b1"]), + ) + false + } catch { + _ => true + } + assert_true(empty && duplicate) +} + +///| +test "DropletUtils advanced default configuration matches portable workflow" { + let config = @src.DropletUtilsConfig::default() + assert_eq(config.lower, 100) + assert_eq(config.iterations, 10000) + assert_eq(config.fdr_threshold, 0.001) + assert_eq(config.retain, -1.0) + assert_eq(config.by_rank, 0) + assert_eq(config.ignore, -1) + assert_eq(config.alpha, -1.0) + assert_eq(config.estimate_alpha, false) + assert_eq(config.seed, 1) +} + +///| +test "DropletUtils advanced configuration preserves explicit options" { + let config = @src.DropletUtilsConfig::create( + lower=12, + iterations=75, + fdr_threshold=0.2, + retain=50.0, + by_rank=3, + ignore=2, + test_ambient=true, + alpha=4.0, + estimate_alpha=true, + rank_exclude=2, + rank_window=0.5, + gradient_threshold=-0.5, + round_counts=false, + seed=19, + ) catch { + _ => abort("valid explicit configuration should build") + } + assert_eq(config.lower, 12) + assert_eq(config.alpha, 4.0) + assert_true(config.test_ambient && config.estimate_alpha) + assert_eq(config.round_counts, false) + assert_eq(config.seed, 19) +} + +///| +test "DropletUtils advanced rejects invalid count thresholds" { + let lower = try { + ignore(@src.DropletUtilsConfig::create(lower=-1)) + false + } catch { + _ => true + } + let iterations = try { + ignore(@src.DropletUtilsConfig::create(iterations=0)) + false + } catch { + _ => true + } + let by_rank = try { + ignore(@src.DropletUtilsConfig::create(by_rank=-1)) + false + } catch { + _ => true + } + let ignore_value = try { + ignore(@src.DropletUtilsConfig::create(ignore=-2)) + false + } catch { + _ => true + } + assert_true(lower && iterations && by_rank && ignore_value) +} + +///| +test "DropletUtils advanced rejects invalid probability configuration" { + let fdr = try { + ignore(@src.DropletUtilsConfig::create(fdr_threshold=1.1)) + false + } catch { + _ => true + } + let retain = try { + ignore(@src.DropletUtilsConfig::create(retain=0.0)) + false + } catch { + _ => true + } + let alpha = try { + ignore(@src.DropletUtilsConfig::create(alpha=0.0)) + false + } catch { + _ => true + } + let seed = try { + ignore(@src.DropletUtilsConfig::create(seed=0)) + false + } catch { + _ => true + } + assert_true(fdr && retain && alpha && seed) +} + +///| +test "DropletUtils advanced Good-Turing profile is normalized and positive" { + let profile = @src.droplet_utils_good_turing([1, 1, 2, 3, 0, 0]) catch { + _ => abort("Good-Turing profile should fit") + } + let mut total = 0.0 + for value in profile { + assert_true(value > 0.0) + total = total + value + } + assert_true((total - 1.0).abs() < 1.0e-10) + assert_true((profile[0] - profile[1]).abs() < 1.0e-12) + assert_true((profile[4] - profile[5]).abs() < 1.0e-12) +} + +///| +test "DropletUtils advanced Good-Turing protects unobserved active features" { + let profile = @src.droplet_utils_good_turing([17, 15, 0, 0]) catch { + _ => abort("safe Good-Turing profile should fit") + } + assert_true(profile[2] > 0.0) + assert_true(profile[3] > 0.0) + assert_true((profile[2] - profile[3]).abs() < 1.0e-12) +} + +///| +test "DropletUtils advanced Good-Turing rejects invalid inputs" { + let empty = try { + ignore(@src.droplet_utils_good_turing([])) + false + } catch { + _ => true + } + let all_zero = try { + ignore(@src.droplet_utils_good_turing([0, 0])) + false + } catch { + _ => true + } + let negative = try { + ignore(@src.droplet_utils_good_turing([1, -1])) + false + } catch { + _ => true + } + assert_true(empty && all_zero && negative) +} + +///| +test "DropletUtils advanced ambient profile uses lower threshold" { + let ambient = @src.droplet_utils_ambient_profile( + dua_test_small_data(), + lower=5, + ) catch { + _ => abort("ambient profile should fit") + } + assert_eq(ambient.assumed_empty, [ + true, true, true, true, true, false, false, false, + ]) + assert_eq(ambient.n_empty(), 5) + assert_eq(ambient.total, 25) + assert_eq(ambient.raw_counts, [14, 11, 0, 0]) +} + +///| +test "DropletUtils advanced ambient profile honors known-empty mask" { + let mask = [true, false, true, false, true, false, false, false] + let ambient = @src.droplet_utils_ambient_profile( + dua_test_small_data(), + lower=0, + known_empty=mask, + ) catch { + _ => abort("known-empty ambient profile should fit") + } + assert_eq(ambient.assumed_empty, mask) + assert_eq(ambient.n_empty(), 3) + assert_eq(ambient.total, 15) +} + +///| +test "DropletUtils advanced ambient profile supports by-rank selection" { + let data = @src.DropletCountMatrix::create([[50, 40, 30, 20, 10]], barcodes=[ + "b1", "b2", "b3", "b4", "b5", + ]) catch { + _ => abort("valid rank data should build") + } + let ambient = @src.droplet_utils_ambient_profile(data, by_rank=2) catch { + _ => abort("by-rank ambient profile should fit") + } + assert_eq(ambient.assumed_empty, [false, false, true, true, true]) + assert_eq(ambient.lower, 30) + assert_eq(ambient.total, 60) +} + +///| +test "DropletUtils advanced ambient profile rejects invalid masks" { + let wrong_length = try { + ignore( + @src.droplet_utils_ambient_profile(dua_test_small_data(), known_empty=[ + true, + ]), + ) + false + } catch { + _ => true + } + let none_selected = try { + ignore( + @src.droplet_utils_ambient_profile( + dua_test_small_data(), + known_empty=Array::make(8, false), + ), + ) + false + } catch { + _ => true + } + assert_true(wrong_length && none_selected) +} + +///| +test "DropletUtils advanced ambient profile rejects empty ambient counts" { + let raised = try { + ignore(@src.droplet_utils_ambient_profile(dua_test_small_data(), lower=0)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced barcode ranks average ties" { + let data = @src.DropletCountMatrix::create([[100, 100, 60, 60, 30, 15, 7, 3]]) catch { + _ => abort("valid rank data should build") + } + let ranks = @src.droplet_utils_barcode_ranks( + data, + lower=0, + exclude_from=0, + window=0.2, + ) catch { + _ => abort("barcode ranks should fit") + } + assert_eq(ranks.ranks[0], 1.5) + assert_eq(ranks.ranks[1], 1.5) + assert_eq(ranks.ranks[2], 3.5) + assert_eq(ranks.ranks[3], 3.5) +} + +///| +test "DropletUtils advanced barcode ranks return finite knees" { + let data = @src.DropletCountMatrix::create([ + [1000, 800, 600, 350, 180, 90, 45, 20, 8, 3], + ]) catch { + _ => abort("valid rank curve should build") + } + let ranks = @src.droplet_utils_barcode_ranks( + data, + lower=0, + exclude_from=0, + window=0.25, + ) catch { + _ => abort("barcode-rank curve should fit") + } + assert_true(ranks.knee > 0.0 && ranks.knee <= 1000.0) + assert_true(ranks.inflection > 0.0 && ranks.inflection <= 1000.0) + assert_eq(ranks.totals[0], 1000) +} + +///| +test "DropletUtils advanced barcode ranks reject insufficient curves" { + let data = @src.DropletCountMatrix::create([[10, 10, 10]]) catch { + _ => abort("valid tied rank data should build") + } + let raised = try { + ignore( + @src.droplet_utils_barcode_ranks( + data, + lower=0, + exclude_from=0, + window=0.2, + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced multinomial log probability is exact" { + let value = @src.droplet_utils_log_probability([2, 1], [0.25, 0.75]) catch { + _ => abort("multinomial probability should compute") + } + assert_true((@math.exp(value) - 0.140625).abs() < 1.0e-10) +} + +///| +test "DropletUtils advanced Dirichlet-multinomial probability is exact" { + let value = @src.droplet_utils_log_probability([1, 1], [0.5, 0.5], alpha=2.0) catch { + _ => abort("Dirichlet-multinomial probability should compute") + } + assert_true((@math.exp(value) - 1.0 / 3.0).abs() < 1.0e-10) +} + +///| +test "DropletUtils advanced multinomial and overdispersed models differ" { + let multinomial = @src.droplet_utils_log_probability([5, 0], [0.5, 0.5]) catch { + _ => abort("multinomial probability should compute") + } + let overdispersed = @src.droplet_utils_log_probability( + [5, 0], + [0.5, 0.5], + alpha=2.0, + ) catch { + _ => abort("Dirichlet-multinomial probability should compute") + } + assert_true(overdispersed > multinomial) +} + +///| +test "DropletUtils advanced probability validates dimensions and counts" { + let dimensions = try { + ignore(@src.droplet_utils_log_probability([1], [0.5, 0.5])) + false + } catch { + _ => true + } + let negative = try { + ignore(@src.droplet_utils_log_probability([-1, 2], [0.5, 0.5])) + false + } catch { + _ => true + } + assert_true(dimensions && negative) +} + +///| +test "DropletUtils advanced probability validates ambient proportions" { + let sum = try { + ignore(@src.droplet_utils_log_probability([1, 1], [0.4, 0.4])) + false + } catch { + _ => true + } + let negative = try { + ignore(@src.droplet_utils_log_probability([1, 1], [-0.1, 1.1])) + false + } catch { + _ => true + } + let impossible = try { + ignore(@src.droplet_utils_log_probability([1, 0], [0.0, 1.0])) + false + } catch { + _ => true + } + assert_true(sum && negative && impossible) +} + +///| +test "DropletUtils advanced Monte Carlo p-value is reproducible" { + let first = @src.droplet_utils_monte_carlo_p_value( + [8, 2], + [0.5, 0.5], + iterations=199, + seed=17, + ) catch { + _ => abort("Monte Carlo test should run") + } + let second = @src.droplet_utils_monte_carlo_p_value( + [8, 2], + [0.5, 0.5], + iterations=199, + seed=17, + ) catch { + _ => abort("Monte Carlo test should repeat") + } + assert_eq(first, second) + assert_true(first.0 >= 1.0 / 200.0 && first.0 <= 1.0) +} + +///| +test "DropletUtils advanced Monte Carlo uses Phipson-Smyth correction" { + let result = @src.droplet_utils_monte_carlo_p_value( + [30, 0], + [0.5, 0.5], + iterations=99, + seed=5, + ) catch { + _ => abort("extreme Monte Carlo test should run") + } + assert_eq(result.1, true) + assert_true((result.0 - 0.01).abs() < 1.0e-12) +} + +///| +test "DropletUtils advanced Monte Carlo supports Dirichlet-multinomial null" { + let result = @src.droplet_utils_monte_carlo_p_value( + [7, 3], + [0.5, 0.5], + iterations=99, + alpha=3.0, + seed=7, + ) catch { + _ => abort("Dirichlet-multinomial Monte Carlo test should run") + } + assert_true(result.0 >= 0.01 && result.0 <= 1.0) +} + +///| +test "DropletUtils advanced Monte Carlo rejects invalid controls" { + let iterations = try { + ignore( + @src.droplet_utils_monte_carlo_p_value([1, 1], [0.5, 0.5], iterations=0), + ) + false + } catch { + _ => true + } + let zero_total = try { + ignore(@src.droplet_utils_monte_carlo_p_value([0, 0], [0.5, 0.5])) + false + } catch { + _ => true + } + assert_true(iterations && zero_total) +} + +///| +test "DropletUtils advanced alpha estimation is finite and positive" { + let data = dua_test_small_data() + let ambient = @src.droplet_utils_ambient_profile(data, lower=5) catch { + _ => abort("ambient profile should fit") + } + let alpha = @src.droplet_utils_estimate_alpha(data, ambient) catch { + _ => abort("alpha should fit") + } + assert_true(alpha >= 0.01 && alpha <= 10000.0) + assert_true(alpha == alpha && alpha.abs() < 1.0e300) +} + +///| +test "DropletUtils advanced alpha estimation validates dimensions" { + let ambient = @src.droplet_utils_ambient_profile( + dua_test_small_data(), + lower=5, + ) catch { + _ => abort("ambient profile should fit") + } + let mismatched = @src.DropletCountMatrix::create([[1, 2]]) catch { + _ => abort("mismatched matrix should build") + } + let raised = try { + ignore(@src.droplet_utils_estimate_alpha(mismatched, ambient)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced emptyDrops preserves upstream NA semantics" { + let result = dua_test_result() + for barcode in 0..<5 { + assert_true(result.p_values[barcode] is None) + assert_true(result.log_probabilities[barcode] is None) + assert_true(result.limited[barcode] is None) + assert_true(result.fdr[barcode] is None) + } + assert_true(result.p_values[5] is Some(_)) + assert_true(result.p_values[6] is Some(_)) + assert_true(result.p_values[7] is Some(_)) +} + +///| +test "DropletUtils advanced emptyDrops retains high-count barcodes" { + let result = dua_test_result() + assert_eq(result.always_retained, [ + false, false, false, false, false, false, true, true, + ]) + assert_eq(result.calls[6], true) + assert_eq(result.calls[7], true) + match result.fdr[6] { + Some(value) => assert_eq(value, 0.0) + None => abort("retained barcode should have FDR") + } +} + +///| +test "DropletUtils advanced emptyDrops records Limited flags" { + let result = dua_test_result() + match result.limited[6] { + Some(value) => assert_eq(value, true) + None => abort("tested barcode should have Limited flag") + } + match result.p_values[6] { + Some(value) => assert_true(value >= 1.0 / 200.0 && value <= 1.0) + None => abort("tested barcode should have p-value") + } +} + +///| +test "DropletUtils advanced emptyDrops reports model diagnostics" { + let result = dua_test_result() + assert_eq(result.lower, 5) + assert_eq(result.retain, 30.0) + assert_eq(result.alpha, -1.0) + assert_eq(result.alpha_estimated, false) + assert_eq(result.iterations, 199) + assert_eq(result.fdr_threshold, 0.1) + assert_eq(result.ambient.n_empty(), 5) +} + +///| +test "DropletUtils advanced emptyDrops supports ambient diagnostics" { + let config = @src.DropletUtilsConfig::create( + lower=5, + iterations=49, + fdr_threshold=0.2, + retain=30.0, + test_ambient=true, + seed=11, + ) catch { + _ => abort("diagnostic configuration should build") + } + let result = @src.droplet_utils_empty_drops(dua_test_small_data(), config~) catch { + _ => abort("diagnostic workflow should run") + } + assert_eq(result.n_tested(), 8) + assert_true(result.p_values[0] is Some(_)) + assert_true(result.fdr[0] is None) + assert_eq(result.calls[0], false) +} + +///| +test "DropletUtils advanced emptyDrops supports ignore threshold" { + let config = @src.DropletUtilsConfig::create( + lower=5, + iterations=49, + fdr_threshold=0.2, + retain=100.0, + ignore=33, + seed=11, + ) catch { + _ => abort("ignore configuration should build") + } + let result = @src.droplet_utils_empty_drops(dua_test_small_data(), config~) catch { + _ => abort("ignore workflow should run") + } + assert_eq(result.n_tested(), 1) + assert_true(result.p_values[6] is Some(_)) + assert_true(result.p_values[7] is None) +} + +///| +test "DropletUtils advanced emptyDrops supports alpha estimation" { + let config = @src.DropletUtilsConfig::create( + lower=5, + iterations=29, + fdr_threshold=0.2, + retain=30.0, + estimate_alpha=true, + seed=13, + ) catch { + _ => abort("estimated-alpha configuration should build") + } + let result = @src.droplet_utils_empty_drops(dua_test_small_data(), config~) catch { + _ => abort("estimated-alpha workflow should run") + } + assert_eq(result.alpha_estimated, true) + assert_true(result.alpha > 0.0) +} + +///| +test "DropletUtils advanced emptyDrops supports known-empty droplets" { + let mask = [true, true, true, true, true, false, false, false] + let result = @src.droplet_utils_empty_drops( + dua_test_small_data(), + known_empty=mask, + config=dua_test_config(), + ) catch { + _ => abort("known-empty workflow should run") + } + assert_eq(result.ambient.assumed_empty, mask) + assert_eq(result.ambient.n_empty(), 5) +} + +///| +test "DropletUtils advanced emptyDrops computes automatic knee retention" { + let data = @src.droplet_utils_advanced_example() catch { + _ => abort("advanced example should build") + } + let config = @src.DropletUtilsConfig::create( + lower=15, + iterations=19, + fdr_threshold=0.2, + rank_exclude=0, + rank_window=0.2, + seed=3, + ) catch { + _ => abort("automatic-retain configuration should build") + } + let result = @src.droplet_utils_empty_drops(data, config~) catch { + _ => abort("automatic-retain workflow should run") + } + assert_true(result.retain > 0.0) + let mut retained = 0 + for value in result.always_retained { + if value { + retained = retained + 1 + } + } + assert_true(retained > 0 && retained < data.n_barcodes()) +} + +///| +test "DropletUtils advanced result query helpers are consistent" { + let result = dua_test_result() + assert_eq(result.n_called(), 2) + assert_eq(result.n_tested(), 3) + assert_eq(result.index_of("cell_a"), Some(6)) + assert_eq(result.index_of("missing"), None) + assert_eq(result.called_barcodes(), ["cell_a", "cell_b"]) + assert_true(result.summary().has_prefix("DropletUtils emptyDrops:")) +} + +///| +test "DropletUtils advanced result filters called barcodes" { + let filtered = dua_test_result().filter(dua_test_small_data()) catch { + _ => abort("called barcodes should filter") + } + assert_eq(filtered.n_features(), 4) + assert_eq(filtered.n_barcodes(), 2) + assert_eq(filtered.barcodes, ["cell_a", "cell_b"]) + assert_eq(filtered.barcode_totals(), [35, 33]) +} + +///| +test "DropletUtils advanced result rejects mismatched filtering data" { + let other = @src.DropletCountMatrix::create( + [[1, 2], [3, 4]], + feature_names=["g1", "g2"], + barcodes=["x", "y"], + ) catch { + _ => abort("other matrix should build") + } + let raised = try { + ignore(dua_test_result().filter(other)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced example has ambient and cell populations" { + let data = @src.droplet_utils_advanced_example() catch { + _ => abort("advanced example should build") + } + assert_eq(data.n_features(), 12) + assert_eq(data.n_barcodes(), 64) + let totals = data.barcode_totals() + assert_true(totals[0] < 20) + assert_true(totals[40] > 80) + assert_true(totals[52] > 80) +} + +///| +test "DropletUtils advanced SingleCellExperiment writes result columns" { + let output = @src.droplet_utils_empty_drops_sce( + dua_test_sce(), + output_prefix="droplet", + config=dua_test_config(), + ) catch { + _ => abort("SingleCellExperiment workflow should run") + } + assert_true(output.experiment.col_data.contains("droplet.total")) + assert_true(output.experiment.col_data.contains("droplet.pValue")) + assert_true(output.experiment.col_data.contains("droplet.limited")) + assert_true(output.experiment.col_data.contains("droplet.fdr")) + assert_true(output.experiment.col_data.contains("droplet.class")) + assert_true(output.experiment.col_data.contains("droplet.retained")) + assert_eq(output.experiment.col_data["droplet.class"][6], "cell") +} + +///| +test "DropletUtils advanced SingleCellExperiment writes ambient rows and metadata" { + let output = @src.droplet_utils_empty_drops_sce( + dua_test_sce(), + output_prefix="droplet", + config=dua_test_config(), + ) catch { + _ => abort("SingleCellExperiment workflow should run") + } + assert_eq(output.experiment.row_data["droplet.ambient"].length(), 4) + assert_eq(output.experiment.row_data["droplet.ambientCount"].length(), 4) + assert_eq(output.experiment.metadata["droplet.lower"], "5") + assert_eq(output.experiment.metadata["droplet.called"], "2") +} + +///| +test "DropletUtils advanced SingleCellExperiment does not mutate input" { + let input = dua_test_sce() + let output = @src.droplet_utils_empty_drops_sce( + input, + output_prefix="droplet", + config=dua_test_config(), + ) catch { + _ => abort("SingleCellExperiment workflow should run") + } + assert_eq(input.col_data.contains("droplet.class"), false) + output.experiment.assays["counts"][0][0] = 999.0 + assert_eq(input.assays["counts"][0][0], 3.0) +} + +///| +test "DropletUtils advanced SingleCellExperiment parses known-empty column" { + let input = dua_test_sce() + input.col_data["known"] = [ + "empty", "empty", "empty", "empty", "empty", "cell", "cell", "cell", + ] + let output = @src.droplet_utils_empty_drops_sce( + input, + known_empty_column="known", + config=dua_test_config(), + ) catch { + _ => abort("known-empty SingleCellExperiment workflow should run") + } + assert_eq(output.result.ambient.n_empty(), 5) + assert_eq(output.result.ambient.assumed_empty[5], false) +} + +///| +test "DropletUtils advanced SingleCellExperiment rounds counts by default" { + let input = dua_test_sce() + input.assays["counts"][0][0] = 3.4 + let output = @src.droplet_utils_empty_drops_sce( + input, + config=dua_test_config(), + ) catch { + _ => abort("rounded SingleCellExperiment workflow should run") + } + assert_eq(output.result.totals[0], 5) +} + +///| +test "DropletUtils advanced SingleCellExperiment can require integer counts" { + let input = dua_test_sce() + input.assays["counts"][0][0] = 3.4 + let config = @src.DropletUtilsConfig::create( + lower=5, + iterations=19, + fdr_threshold=0.2, + retain=30.0, + round_counts=false, + ) catch { + _ => abort("strict integer configuration should build") + } + let raised = try { + ignore(@src.droplet_utils_empty_drops_sce(input, config~)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced SingleCellExperiment rejects missing assay" { + let raised = try { + ignore( + @src.droplet_utils_empty_drops_sce( + dua_test_sce(), + assay_name="missing", + config=dua_test_config(), + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced SingleCellExperiment validates known-empty values" { + let input = dua_test_sce() + input.col_data["known"] = [ + "empty", "empty", "empty", "empty", "empty", "unknown", "cell", "cell", + ] + let raised = try { + ignore( + @src.droplet_utils_empty_drops_sce( + input, + known_empty_column="known", + config=dua_test_config(), + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "DropletUtils advanced SingleCellExperiment validates assay values" { + let negative = dua_test_sce() + negative.assays["counts"][0][0] = -1.0 + let negative_raised = try { + ignore( + @src.droplet_utils_empty_drops_sce(negative, config=dua_test_config()), + ) + false + } catch { + _ => true + } + let non_finite = dua_test_sce() + non_finite.assays["counts"][0][0] = Double::nan() + let finite_raised = try { + ignore( + @src.droplet_utils_empty_drops_sce(non_finite, config=dua_test_config()), + ) + false + } catch { + _ => true + } + assert_true(negative_raised && finite_raised) +} diff --git a/test/moonbit/dss_test.mbt b/test/moonbit/dss_test.mbt index e13ccb3b..0306cbc3 100644 --- a/test/moonbit/dss_test.mbt +++ b/test/moonbit/dss_test.mbt @@ -194,4 +194,4 @@ test "dss_disp_result_methods" { for sd in shrunken { assert_true(sd >= 0.0) } -} \ No newline at end of file +} diff --git a/test/moonbit/dssp_test.mbt b/test/moonbit/dssp_test.mbt index 86a41ec3..192558ac 100644 --- a/test/moonbit/dssp_test.mbt +++ b/test/moonbit/dssp_test.mbt @@ -31,4 +31,4 @@ test "dssp_analyze_structure_composition" { test "dssp_create_example_data" { let data = @src.create_example_dssp_data() assert_eq(data.records.length(), 8) -} \ No newline at end of file +} diff --git a/test/moonbit/edaseq_test.mbt b/test/moonbit/edaseq_test.mbt index bce3cb4d..d7b8b7db 100644 --- a/test/moonbit/edaseq_test.mbt +++ b/test/moonbit/edaseq_test.mbt @@ -26,13 +26,10 @@ test "edaseq_gene_anno_min_length" { test "edaseq_dataset" { let gene_ids = ["gene1", "gene2"] let sample_ids = ["sample1", "sample2", "sample3"] - let counts : Array[Array[Double]] = [ - [10.0, 20.0, 30.0], - [50.0, 60.0, 70.0] - ] + let counts : Array[Array[Double]] = [[10.0, 20.0, 30.0], [50.0, 60.0, 70.0]] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.45, 1000.0), - @src.EDASeqGeneAnno::new("gene2", 0.55, 2000.0) + @src.EDASeqGeneAnno::new("gene2", 0.55, 2000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) assert_eq(data.eda_n_genes(), 2) @@ -58,7 +55,7 @@ test "edaseq_sample_counts" { let counts : Array[Array[Double]] = [[10.0], [20.0]] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.5, 500.0), - @src.EDASeqGeneAnno::new("gene2", 0.6, 1000.0) + @src.EDASeqGeneAnno::new("gene2", 0.6, 1000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) let sample_counts = data.eda_sample_counts(0) @@ -74,12 +71,12 @@ test "edaseq_within_lane" { let counts : Array[Array[Double]] = [ [100.0, 200.0], [300.0, 400.0], - [500.0, 600.0] + [500.0, 600.0], ] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.4, 500.0), @src.EDASeqGeneAnno::new("gene2", 0.5, 1000.0), - @src.EDASeqGeneAnno::new("gene3", 0.6, 2000.0) + @src.EDASeqGeneAnno::new("gene3", 0.6, 2000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) let params = @src.EDASeqParams::new() @@ -93,11 +90,11 @@ test "edaseq_between_lane" { let sample_ids = ["s1", "s2", "s3"] let counts : Array[Array[Double]] = [ [100.0, 200.0, 300.0], - [500.0, 600.0, 700.0] + [500.0, 600.0, 700.0], ] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.5, 500.0), - @src.EDASeqGeneAnno::new("gene2", 0.55, 1000.0) + @src.EDASeqGeneAnno::new("gene2", 0.55, 1000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) let params = @src.EDASeqParams::new() @@ -112,12 +109,12 @@ test "edaseq_full_run" { let counts : Array[Array[Double]] = [ [100.0, 200.0], [300.0, 400.0], - [500.0, 600.0] + [500.0, 600.0], ] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.4, 500.0), @src.EDASeqGeneAnno::new("gene2", 0.5, 1000.0), - @src.EDASeqGeneAnno::new("gene3", 0.6, 2000.0) + @src.EDASeqGeneAnno::new("gene3", 0.6, 2000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) let params = @src.EDASeqParams::new() @@ -130,13 +127,10 @@ test "edaseq_full_run" { test "edaseq_result_accessors" { let gene_ids = ["gene1", "gene2"] let sample_ids = ["s1", "s2"] - let counts : Array[Array[Double]] = [ - [100.0, 200.0], - [300.0, 400.0] - ] + let counts : Array[Array[Double]] = [[100.0, 200.0], [300.0, 400.0]] let annotations = [ @src.EDASeqGeneAnno::new("gene1", 0.5, 500.0), - @src.EDASeqGeneAnno::new("gene2", 0.55, 1000.0) + @src.EDASeqGeneAnno::new("gene2", 0.55, 1000.0), ] let data = @src.EDASeqDataSet::new(gene_ids, sample_ids, counts, annotations) let params = @src.EDASeqParams::new() diff --git a/test/moonbit/edger_advanced_test.mbt b/test/moonbit/edger_advanced_test.mbt index 71e473ca..60fa5008 100644 --- a/test/moonbit/edger_advanced_test.mbt +++ b/test/moonbit/edger_advanced_test.mbt @@ -180,9 +180,9 @@ test "camera_with_all_genes_in_set" { ) let qlf = @src.glm_qlf_fit(dge) let result = @src.glm_qlf_test(qlf) - let (stat, pval, dir) = @src.camera( - result, ["GeneA", "GeneB", "GeneC", "GeneD"], - ) + let (stat, pval, dir) = @src.camera(result, [ + "GeneA", "GeneB", "GeneC", "GeneD", + ]) assert_true(pval >= 0.0 && pval <= 1.0) assert_true(dir == "up" || dir == "down" || dir == "none") } @@ -221,11 +221,7 @@ test "roast_with_empty_gene_set" { ///| test "qlf_test_handles_single_sample_per_group" { - let dge = @src.dge_list( - [[10, 20], [30, 40]], - ["A", "B"], - ["G1", "G2"], - ) + let dge = @src.dge_list([[10, 20], [30, 40]], ["A", "B"], ["G1", "G2"]) let qlf = @src.glm_qlf_fit(dge) let result = @src.glm_qlf_test(qlf) assert_eq(result.genes.length(), 2) @@ -238,4 +234,4 @@ test "qlf_test_handles_single_sample_per_group" { for f in result.fdr { assert_true(f >= 0.0 && f <= 1.0) } -} \ No newline at end of file +} diff --git a/test/moonbit/enhanced_volcano_test.mbt b/test/moonbit/enhanced_volcano_test.mbt index b60a0993..16410965 100644 --- a/test/moonbit/enhanced_volcano_test.mbt +++ b/test/moonbit/enhanced_volcano_test.mbt @@ -1,6 +1,5 @@ ///| /// Test file for EnhancedVolcano module. - test "volcano_plot_basic" { let ids = ["TP53", "BRCA1", "MYC", "KRAS", "EGFR"] let lfcs = [2.5, -3.0, 1.2, -0.5, 0.1] @@ -9,101 +8,194 @@ test "volcano_plot_basic" { assert_eq(result.get_n_genes(), 5) } +///| test "volcano_classification_up" { let ids = ["Gene1", "Gene2", "Gene3"] let lfcs = [2.0, -2.0, 0.5] let pvals = [0.001, 0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="Test", x_label="x", y_label="y") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="Test", + x_label="x", + y_label="y", + ) assert_eq(result.get_n_up(), 1) assert_eq(result.get_n_down(), 1) assert_eq(result.get_n_nonsig(), 1) } +///| test "volcano_classification_all_sig" { let ids = ["A", "B", "C", "D"] let lfcs = [2.5, -2.5, 1.5, -1.5] let pvals = [0.001, 0.001, 0.01, 0.01] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) assert_eq(result.get_n_up(), 2) assert_eq(result.get_n_down(), 2) assert_eq(result.get_n_nonsig(), 0) } +///| test "volcano_classification_none_sig" { let ids = ["A", "B", "C"] let lfcs = [0.1, -0.1, 0.5] let pvals = [0.1, 0.2, 0.1] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) assert_eq(result.get_n_up(), 0) assert_eq(result.get_n_down(), 0) assert_eq(result.get_n_nonsig(), 3) } +///| test "volcano_get_genes" { let ids = ["A", "B", "C"] let lfcs = [2.0, -2.0, 0.5] let pvals = [0.001, 0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) let genes = result.get_genes() assert_eq(genes.length(), 3) } +///| test "volcano_get_significant" { let ids = ["A", "B", "C"] let lfcs = [2.0, -2.0, 0.5] let pvals = [0.001, 0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) let sig = result.get_significant_genes() assert_eq(sig.length(), 2) // A (Up) and B (Down) } +///| test "volcano_get_up_genes" { let ids = ["A", "B", "C"] let lfcs = [2.0, -2.0, 1.5] let pvals = [0.001, 0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) let up = result.get_up_genes() assert_eq(up.length(), 2) } +///| test "volcano_get_down_genes" { let ids = ["A", "B", "C"] let lfcs = [2.0, -2.0, -1.5] let pvals = [0.001, 0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) let down = result.get_down_genes() assert_eq(down.length(), 2) } +///| test "volcano_get_top_genes" { let ids = ["A", "B", "C", "D", "E"] let lfcs = [2.5, -3.0, 1.5, -1.5, 0.3] let pvals = [0.01, 0.001, 0.02, 0.005, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="", + x_label="", + y_label="", + ) let top = result.get_top_genes(3) assert_eq(top.length(), 3) } +///| test "volcano_neg_log10_p" { let nlp = @src.neg_log10_p(0.05) assert_true(nlp > 1.29) // -log10(0.05) = 1.301 assert_true(nlp < 1.31) } +///| test "volcano_neg_log10_p_zero" { let nlp = @src.neg_log10_p(0.0) assert_true(nlp > 0.0) // should handle p=0 gracefully } +///| test "volcano_cutoffs" { let ids = ["A", "B"] let lfcs = [2.0, -2.0] let pvals = [0.001, 0.001] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.01, fc_cutoff=2.0, title="", x_label="", y_label="") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.01, + fc_cutoff=2.0, + title="", + x_label="", + y_label="", + ) assert_eq(result.get_p_cutoff(), 0.01) assert_eq(result.get_fc_cutoff(), 2.0) } +///| test "volcano_to_ascii" { let result = @src.volcano_sample(20) let ascii = result.to_ascii(width=60, height=20) @@ -111,6 +203,7 @@ test "volcano_to_ascii" { assert_true(ascii.contains("Volcano Plot")) } +///| test "volcano_summary" { let result = @src.volcano_sample(10) let summary = result.summary() @@ -118,6 +211,7 @@ test "volcano_summary" { assert_true(summary.contains("Total genes")) } +///| test "volcano_sample" { let result = @src.volcano_sample(30) assert_eq(result.get_n_genes(), 30) @@ -126,21 +220,32 @@ test "volcano_sample" { assert_true(result.get_n_nonsig() >= 0) } +///| test "volcano_classification_to_string" { // Test classification through the result of volcano_plot let ids = ["Gene1", "Gene2", "Gene3"] let lfcs = [2.0, -2.0, 0.5] let pvals = [0.001, 0.001, 0.5] - let result = @src.volcano_plot(ids, lfcs, pvals, p_cutoff=0.05, fc_cutoff=1.0, title="Test", x_label="log2FC", y_label="-log10(p)") + let result = @src.volcano_plot( + ids, + lfcs, + pvals, + p_cutoff=0.05, + fc_cutoff=1.0, + title="Test", + x_label="log2FC", + y_label="-log10(p)", + ) let genes = result.get_genes() assert_eq(genes.length(), 3) - + // Check that classifications work by examining result counts - assert_true(result.get_n_up() >= 1) // Gene1 should be Up - assert_true(result.get_n_down() >= 1) // Gene2 should be Down - assert_true(result.get_n_nonsig() >= 1) // Gene3 should be NonSig + assert_true(result.get_n_up() >= 1) // Gene1 should be Up + assert_true(result.get_n_down() >= 1) // Gene2 should be Down + assert_true(result.get_n_nonsig() >= 1) // Gene3 should be NonSig } +///| test "volcano_empty_inputs" { let ids : Array[String] = Array::new() let lfcs : Array[Double] = Array::new() @@ -149,9 +254,10 @@ test "volcano_empty_inputs" { assert_eq(result.get_n_genes(), 0) } +///| test "volcano_ascii_visual_elements" { let result = @src.volcano_sample(50) let ascii = result.to_ascii(width=60, height=20) assert_true(ascii.contains("+")) // up-regulated assert_true(ascii.contains("-")) // down-regulated -} \ No newline at end of file +} diff --git a/test/moonbit/enriched_heatmap_test.mbt b/test/moonbit/enriched_heatmap_test.mbt index b17afaa8..70985adc 100644 --- a/test/moonbit/enriched_heatmap_test.mbt +++ b/test/moonbit/enriched_heatmap_test.mbt @@ -247,7 +247,7 @@ test "eh_normalize_background" { // All values should be background (-1.0) since signal is on different chr. let mut j = 0 while j < mat.nCols { - assert_true((mat.matrix[0][j] - (-1.0)).abs() < 1.0e-10) + assert_true((mat.matrix[0][j] - -1.0).abs() < 1.0e-10) j = j + 1 } } @@ -286,7 +286,10 @@ test "eh_row_means" { ///| test "eh_minus_strand" { // Test that minus strand reverses window order. - let signals = [@src.GenomicSignal::new("chr1", 9800, 9900, 0.5), @src.GenomicSignal::new("chr1", 10100, 10200, 1.0)] + let signals = [ + @src.GenomicSignal::new("chr1", 9800, 9900, 0.5), + @src.GenomicSignal::new("chr1", 10100, 10200, 1.0), + ] let targets = [@src.TargetRegion::new("chr1", 10000, 10000, "-")] let config = @src.EnrichedHeatmapConfig::new( extendUp=500, diff --git a/test/moonbit/enrichplot_test.mbt b/test/moonbit/enrichplot_test.mbt index d932f441..f7a49a73 100644 --- a/test/moonbit/enrichplot_test.mbt +++ b/test/moonbit/enrichplot_test.mbt @@ -1,11 +1,16 @@ ///| - test "enrichplot_create_enrich_term" { let term = @src.EnrichTerm::new( - "GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, - ["gene1", "gene2", "gene3"] + "GO:0005623", + "cell", + 0.001, + 0.01, + 1.5, + 1.8, + 15, + ["gene1", "gene2", "gene3"], ) - + assert_eq(term.term_id, "GO:0005623") assert_eq(term.term_name, "cell") assert_eq(term.pvalue, 0.001) @@ -16,82 +21,124 @@ test "enrichplot_create_enrich_term" { assert_eq(term.genes.length(), 3) } +///| test "enrichplot_dotplot" { let terms = [ - @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, ["g1", "g2", "g3"]), - @src.EnrichTerm::new("GO:0005886", "plasma membrane", 0.002, 0.02, 1.2, 1.5, 10, ["g4", "g5"]), - @src.EnrichTerm::new("GO:0003674", "molecular_function", 0.003, 0.03, 0.8, 1.0, 8, ["g6"]), + @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, [ + "g1", "g2", "g3", + ]), + @src.EnrichTerm::new( + "GO:0005886", + "plasma membrane", + 0.002, + 0.02, + 1.2, + 1.5, + 10, + ["g4", "g5"], + ), + @src.EnrichTerm::new( + "GO:0003674", + "molecular_function", + 0.003, + 0.03, + 0.8, + 1.0, + 8, + ["g6"], + ), ] - + let result = @src.EnrichResult::new(terms, "GO", []) let plot = @src.bio_enrichplot_dotplot(result, 10, "Test Dotplot") - + assert_true(plot.contains("cell")) assert_true(plot.contains("plasma membrane")) assert_true(plot.contains("molecular_function")) } +///| test "enrichplot_barplot" { let terms = [ @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, ["g1"]), - @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, ["g2"]), + @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, [ + "g2", + ]), ] - + let result = @src.EnrichResult::new(terms, "GO", []) let plot = @src.bio_enrichplot_barplot(result, 5, "padj", "Test Barplot") - + assert_true(plot.contains("cell")) assert_true(plot.contains("membrane")) } +///| test "enrichplot_heatmap" { let terms = [ - @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, ["g1", "g2", "g3"]), - @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, ["g2", "g4"]), + @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, [ + "g1", "g2", "g3", + ]), + @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, [ + "g2", "g4", + ]), ] - + let result = @src.EnrichResult::new(terms, "GO", []) let plot = @src.bio_enrichplot_heatmap(result, 5, 5, "Test Heatmap") - + assert_true(plot.contains("cell")) assert_true(plot.contains("membrane")) } +///| test "enrichplot_cnetplot" { let terms = [ - @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, ["g1", "g2"]), - @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, ["g2", "g3"]), + @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, [ + "g1", "g2", + ]), + @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, [ + "g2", "g3", + ]), ] - + let result = @src.EnrichResult::new(terms, "GO", []) let plot = @src.bio_enrichplot_cnetplot(result, 5, "Test Cnetplot") - + assert_true(plot.contains("cell")) assert_true(plot.contains("membrane")) assert_true(plot.contains("g1")) assert_true(plot.contains("g2")) } +///| test "enrichplot_emapplot" { let terms = [ - @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, ["g1", "g2", "g3"]), - @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, ["g2", "g3", "g4"]), - @src.EnrichTerm::new("GO:0003674", "function", 0.003, 0.03, 0.8, 1.0, 8, ["g5"]), + @src.EnrichTerm::new("GO:0005623", "cell", 0.001, 0.01, 1.5, 1.8, 15, [ + "g1", "g2", "g3", + ]), + @src.EnrichTerm::new("GO:0005886", "membrane", 0.002, 0.02, 1.2, 1.5, 10, [ + "g2", "g3", "g4", + ]), + @src.EnrichTerm::new("GO:0003674", "function", 0.003, 0.03, 0.8, 1.0, 8, [ + "g5", + ]), ] - + let result = @src.EnrichResult::new(terms, "GO", []) let plot = @src.bio_enrichplot_emapplot(result, 0.3, "Test Emapplot") - + assert_true(plot.contains("cell")) assert_true(plot.contains("membrane")) } +///| test "enrichplot_empty_result" { let result = @src.EnrichResult::new([], "GO", []) - + let dotplot = @src.bio_enrichplot_dotplot(result, 10, "Test") assert_true(dotplot.contains("Test")) - + let barplot = @src.bio_enrichplot_barplot(result, 10, "padj", "Test") assert_true(barplot.contains("Test")) -} \ No newline at end of file +} diff --git a/test/moonbit/ensembldb_test.mbt b/test/moonbit/ensembldb_test.mbt index a6abe537..0aa49b11 100644 --- a/test/moonbit/ensembldb_test.mbt +++ b/test/moonbit/ensembldb_test.mbt @@ -1,38 +1,42 @@ ///| /// Tests for ensembldb module. - test "EnsDb creation" { let db = @src.create_example_ensdb() assert_eq(@src.edb_num_genes(db), 3) assert_eq(@src.edb_num_transcripts(db), 4) } +///| test "edb_get_gene_by_id" { let db = @src.create_example_ensdb() let gene = @src.edb_get_gene_by_id(db, "ENSG000001") assert_true(gene is Some(_)) } +///| test "edb_get_gene_by_name" { let db = @src.create_example_ensdb() let genes = @src.edb_get_gene_by_name(db, "ACTB") assert_eq(genes.length(), 1) } +///| test "edb_get_transcripts_by_gene" { let db = @src.create_example_ensdb() let txs = @src.edb_get_transcripts_by_gene(db, "ENSG000001") assert_eq(txs.length(), 2) } +///| test "edb_filter_by_chromosome" { let db = @src.create_example_ensdb() let filtered = @src.edb_filter_by_chromosome(db, "17") assert_eq(@src.edb_num_genes(filtered), 2) } +///| test "edb_get_gene_length" { let db = @src.create_example_ensdb() let length = @src.edb_get_gene_length(db, "ENSG000001") assert_eq(length, 6257) -} \ No newline at end of file +} diff --git a/test/moonbit/estimate_score_test.mbt b/test/moonbit/estimate_score_test.mbt index b59c15be..8e0b0581 100644 --- a/test/moonbit/estimate_score_test.mbt +++ b/test/moonbit/estimate_score_test.mbt @@ -7,27 +7,27 @@ // --------------------------------------------------------------------------- test "est_expression_creation" { - let expr = @src.EstExpression::new( - ["G1", "G2"], - ["S1", "S2"], - [[1.0, 2.0], [3.0, 4.0]], - ) + let expr = @src.EstExpression::new(["G1", "G2"], ["S1", "S2"], [ + [1.0, 2.0], + [3.0, 4.0], + ]) assert_eq(expr.gene_names().length(), 2) assert_eq(expr.sample_names().length(), 2) assert_eq(expr.matrix()[1][0], 3.0) } +///| test "est_signature_creation" { - let sig = @src.EstSignature::new( - ["STROMA1", "STROMA2"], - ["IMMUNE1", "IMMUNE2"], - ) + let sig = @src.EstSignature::new(["STROMA1", "STROMA2"], [ + "IMMUNE1", "IMMUNE2", + ]) assert_eq(sig.stromal_genes().length(), 2) assert_eq(sig.immune_genes().length(), 2) assert_eq(sig.stromal_genes()[0], "STROMA1") assert_eq(sig.immune_genes()[1], "IMMUNE2") } +///| test "est_sample_score_accessors" { // Use est_run to get real scores and test accessors let expr = @src.est_sample_data() @@ -43,6 +43,7 @@ test "est_sample_score_accessors" { ) } +///| test "est_result_accessors" { // Use est_run to get a real result and test accessors let expr = @src.est_sample_data() @@ -56,6 +57,7 @@ test "est_result_accessors" { // Built-in gene signatures // --------------------------------------------------------------------------- +///| test "est_stromal_genes_nonempty" { let genes = @src.est_stromal_genes() assert_true(genes.length() >= 30) @@ -63,24 +65,32 @@ test "est_stromal_genes_nonempty" { let mut has_col1a1 = false let mut has_dcn = false for g in genes { - if g == "COL1A1" { has_col1a1 = true } - if g == "DCN" { has_dcn = true } + if g == "COL1A1" { + has_col1a1 = true + } + if g == "DCN" { + has_dcn = true + } } assert_true(has_col1a1) assert_true(has_dcn) } +///| test "est_immune_genes_nonempty" { let genes = @src.est_immune_genes() assert_true(genes.length() >= 20) // Should contain typical immune markers like CD3D, CD3E let mut has_cd3d = false for g in genes { - if g == "CD3D" { has_cd3d = true } + if g == "CD3D" { + has_cd3d = true + } } assert_true(has_cd3d) } +///| test "est_default_signature" { let sig = @src.est_default_signature() assert_true(sig.stromal_genes().length() >= 30) @@ -91,6 +101,7 @@ test "est_default_signature" { // ECDF normalization behavior // --------------------------------------------------------------------------- +///| test "est_run_normalizes_across_samples" { // With 4 samples, ECDF values are 0.25, 0.5, 0.75, 1.0 // Scores should be in [-100, 100] range @@ -112,6 +123,7 @@ test "est_run_normalizes_across_samples" { // Sample data correctness // --------------------------------------------------------------------------- +///| test "est_sample_data_dimensions" { let expr = @src.est_sample_data() assert_eq(expr.sample_names().length(), 4) @@ -123,6 +135,7 @@ test "est_sample_data_dimensions" { } } +///| test "est_sample_data_sample_names" { let expr = @src.est_sample_data() assert_eq(expr.sample_names()[0], "HighStromal") @@ -135,6 +148,7 @@ test "est_sample_data_sample_names" { // Score semantics // --------------------------------------------------------------------------- +///| test "est_run_high_stromal_higher_than_low_stromal" { // HighStromal sample should have higher stromal score than LowBoth let expr = @src.est_sample_data() @@ -144,6 +158,7 @@ test "est_run_high_stromal_higher_than_low_stromal" { assert_true(high_stromal.stromal_score() > low_both.stromal_score()) } +///| test "est_run_high_immune_higher_than_low_immune" { // HighImmune sample should have higher immune score than LowBoth let expr = @src.est_sample_data() @@ -153,6 +168,7 @@ test "est_run_high_immune_higher_than_low_immune" { assert_true(high_immune.immune_score() > low_both.immune_score()) } +///| test "est_run_high_both_highest_estimate_score" { // HighBoth should have the highest combined ESTIMATE score let expr = @src.est_sample_data() @@ -162,6 +178,7 @@ test "est_run_high_both_highest_estimate_score" { assert_true(high_both.estimate_score() > low_both.estimate_score()) } +///| test "est_run_estimate_score_equals_stromal_plus_immune" { // ESTIMATEScore = StromalScore + ImmuneScore let expr = @src.est_sample_data() @@ -172,6 +189,7 @@ test "est_run_estimate_score_equals_stromal_plus_immune" { } } +///| test "est_run_low_both_has_low_scores" { // LowBoth (low infiltration) should have negative scores for both let expr = @src.est_sample_data() @@ -185,6 +203,7 @@ test "est_run_low_both_has_low_scores" { // Tumor purity // --------------------------------------------------------------------------- +///| test "est_run_tumor_purity_in_valid_range" { // Tumor purity should be in [0, 1] or NaN let expr = @src.est_sample_data() @@ -201,6 +220,7 @@ test "est_run_tumor_purity_in_valid_range" { } } +///| test "est_run_low_infiltration_higher_purity" { // LowBoth (low infiltration -> higher tumor content) should have higher // tumor purity than HighBoth (high infiltration -> lower tumor content) @@ -208,8 +228,7 @@ test "est_run_low_infiltration_higher_purity" { let result = @src.est_run(expr) let low_both = result.get_score("LowBoth") let high_both = result.get_score("HighBoth") - if !low_both.tumor_purity().is_nan() && - !high_both.tumor_purity().is_nan() { + if !low_both.tumor_purity().is_nan() && !high_both.tumor_purity().is_nan() { assert_true(low_both.tumor_purity() >= high_both.tumor_purity()) } } @@ -218,19 +237,16 @@ test "est_run_low_infiltration_higher_purity" { // Custom signature // --------------------------------------------------------------------------- +///| test "est_run_custom_signature" { // Build expression matrix with 4 genes, 2 samples // G1, G2 are stromal markers; G3, G4 are immune markers - let expr = @src.EstExpression::new( - ["G1", "G2", "G3", "G4"], - ["S1", "S2"], - [ - [10.0, 1.0], // G1 - stromal, high in S1 - [8.0, 1.0], // G2 - stromal, high in S1 - [1.0, 10.0], // G3 - immune, high in S2 - [1.0, 8.0], // G4 - immune, high in S2 - ], - ) + let expr = @src.EstExpression::new(["G1", "G2", "G3", "G4"], ["S1", "S2"], [ + [10.0, 1.0], // G1 - stromal, high in S1 + [8.0, 1.0], // G2 - stromal, high in S1 + [1.0, 10.0], // G3 - immune, high in S2 + [1.0, 8.0], // G4 - immune, high in S2 + ]) let sig = @src.EstSignature::new(["G1", "G2"], ["G3", "G4"]) let result = @src.est_run(expr, signature=sig) assert_eq(result.scores().length(), 2) @@ -241,13 +257,13 @@ test "est_run_custom_signature" { assert_true(s2.immune_score() > s1.immune_score()) } +///| test "est_run_empty_signature_returns_zeros" { // With empty gene sets, ssGSEA returns 0; ECDF of all-zeros is degenerate - let expr = @src.EstExpression::new( - ["G1", "G2"], - ["S1", "S2"], - [[1.0, 2.0], [3.0, 4.0]], - ) + let expr = @src.EstExpression::new(["G1", "G2"], ["S1", "S2"], [ + [1.0, 2.0], + [3.0, 4.0], + ]) let sig = @src.EstSignature::new([], []) let result = @src.est_run(expr, signature=sig) for s in result.scores() { @@ -261,6 +277,7 @@ test "est_run_empty_signature_returns_zeros" { // get_score and to_string // --------------------------------------------------------------------------- +///| test "est_result_get_score_known_sample" { let expr = @src.est_sample_data() let result = @src.est_run(expr) @@ -268,6 +285,7 @@ test "est_result_get_score_known_sample" { assert_eq(s.sample_id(), "HighStromal") } +///| test "est_result_get_score_unknown_sample" { let expr = @src.est_sample_data() let result = @src.est_run(expr) @@ -278,6 +296,7 @@ test "est_result_get_score_unknown_sample" { assert_true(s.estimate_score().abs() < 1.0e-9) } +///| test "est_result_to_string" { let expr = @src.est_sample_data() let result = @src.est_run(expr) @@ -293,13 +312,15 @@ test "est_result_to_string" { // Edge cases // --------------------------------------------------------------------------- +///| test "est_run_single_sample" { // Single sample: ECDF is degenerate (trivially all-equal), scores are 0.0 - let expr = @src.EstExpression::new( - ["G1", "G2", "G3", "G4"], - ["Only"], - [[10.0], [8.0], [5.0], [3.0]], - ) + let expr = @src.EstExpression::new(["G1", "G2", "G3", "G4"], ["Only"], [ + [10.0], + [8.0], + [5.0], + [3.0], + ]) let sig = @src.EstSignature::new(["G1", "G2"], ["G3", "G4"]) let result = @src.est_run(expr, signature=sig) assert_eq(result.scores().length(), 1) @@ -309,13 +330,14 @@ test "est_run_single_sample" { assert_true(s.immune_score().abs() < 1.0e-9) } +///| test "est_run_no_signature_genes_present" { // None of the signature genes are in the expression matrix - let expr = @src.EstExpression::new( - ["X1", "X2", "X3"], - ["S1", "S2"], - [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]], - ) + let expr = @src.EstExpression::new(["X1", "X2", "X3"], ["S1", "S2"], [ + [1.0, 2.0], + [3.0, 4.0], + [5.0, 6.0], + ]) let sig = @src.EstSignature::new(["G1", "G2"], ["G3", "G4"]) let result = @src.est_run(expr, signature=sig) // All scores should be 0 (or near 0) since no genes match @@ -325,6 +347,7 @@ test "est_run_no_signature_genes_present" { } } +///| test "est_run_all_samples_identical" { // Identical samples should have identical scores let expr = @src.EstExpression::new( diff --git a/test/moonbit/exonerate_test.mbt b/test/moonbit/exonerate_test.mbt index 6339760b..4686769e 100644 --- a/test/moonbit/exonerate_test.mbt +++ b/test/moonbit/exonerate_test.mbt @@ -38,9 +38,7 @@ test "exonerate_record_new" { @src.AlignmentBlock::new("I", 0, 0, 1000.0), ] let rec = @src.ExonerateRecord::new( - "query1", 0, 100, "+", - "target1", 0, 95, "+", - 1850.0, blocks, + "query1", 0, 100, "+", "target1", 0, 95, "+", 1850.0, blocks, ) assert_eq(rec.query_name(), "query1") assert_eq(rec.query_start(), 0) diff --git a/test/moonbit/exonerate_text_test.mbt b/test/moonbit/exonerate_text_test.mbt new file mode 100644 index 00000000..d0590ddc --- /dev/null +++ b/test/moonbit/exonerate_text_test.mbt @@ -0,0 +1,861 @@ +// Black-box tests for Bio.SearchIO.ExonerateIO.exonerate_text support. + +///| +fn exonerate_text_test_parse(text : String) -> @src.ExonerateTextDocument { + @src.exonerate_text_parse(text) catch { + ExonerateTextError(message) => + abort("valid Exonerate text failed: " + message) + } +} + +///| +fn exonerate_text_test_rejects(text : String) -> Bool { + try { + ignore(@src.exonerate_text_parse(text)) + false + } catch { + ExonerateTextError(_) => true + } +} + +///| +fn exonerate_text_test_wrap( + query : String, + query_description : String, + target : String, + target_description : String, + model : String, + score : Int, + query_start : Int, + query_end : Int, + target_start : Int, + target_end : Int, + body : String, +) -> String { + "Command line: [exonerate -m " + + model + + " query.fa target.fa]\n" + + "Hostname: [test-host]\n\n" + + exonerate_text_test_alignment( + query, query_description, target, target_description, model, score, query_start, + query_end, target_start, target_end, body, + ) + + "-- completed exonerate analysis\n" +} + +///| +fn exonerate_text_test_alignment( + query : String, + query_description : String, + target : String, + target_description : String, + model : String, + score : Int, + query_start : Int, + query_end : Int, + target_start : Int, + target_end : Int, + body : String, +) -> String { + "C4 Alignment:\n" + + "------------\n" + + " Query: " + + query + + (if query_description.length() > 0 { " " + query_description } else { "" }) + + "\n" + + " Target: " + + target + + (if target_description.length() > 0 { " " + target_description } else { "" }) + + "\n" + + " Model: " + + model + + "\n" + + " Raw score: " + + score.to_string() + + "\n" + + " Query range: " + + query_start.to_string() + + " -> " + + query_end.to_string() + + "\n" + + " Target range: " + + target_start.to_string() + + " -> " + + target_end.to_string() + + "\n\n" + + body + + "\n" +} + +///| +fn exonerate_text_test_dna_report() -> String { + exonerate_text_test_wrap( + "query1", + "demo query", + "target1", + "demo target", + "affine:local:dna2dna", + 31, + 0, + 6, + 100, + 107, + " 1 : ACGT-AC : 6\n" + " |||| ||\n" + " 101 : ACGTTAC : 107\n", + ) +} + +///| +fn exonerate_text_test_reverse_report() -> String { + exonerate_text_test_wrap( + "query1", + "", + "target1", + "reverse target:[revcomp]", + "ungapped:dna2dna", + 18, + 0, + 6, + 200, + 193, + " 1 : ACGTAC- : 6\n" + " |||||| \n" + " 200 : ACGTACT : 194\n", + ) +} + +///| +fn exonerate_text_test_center(value : String, width : Int) -> String { + let left = (width - value.length()) / 2 + " ".repeat(left) + value + " ".repeat(width - value.length() - left) +} + +///| +fn exonerate_text_test_intron_report() -> String { + let marker = " >>>> Target Intron 1 >>>> " + let splice = "gt" + ".".repeat(marker.length() - 4) + "ag" + let query_row = "ACG" + marker + "TAC" + let hit_row = "ACG" + splice + "TAC" + let similarity = "|||" + + exonerate_text_test_center("7 bp", marker.length()) + + "|||" + exonerate_text_test_wrap( + "transcript", + "", + "chromosome", + "", + "est2genome", + 42, + 0, + 6, + 100, + 113, + " 1 : " + + query_row + + " : 6\n" + + " " + + similarity + + "\n" + + " 101 : " + + hit_row + + " : 113\n", + ) +} + +///| +fn exonerate_text_test_ner_report() -> String { + let query_marker = "--< 4 >--" + let hit_marker = "--< 6 >--" + exonerate_text_test_wrap( + "query", + "", + "target", + "", + "NER:affine:local:dna2dna", + 25, + 0, + 10, + 100, + 112, + " 1 : AAA" + + query_marker + + "CCC : 10\n" + + " |||--< NER 1 >--|||\n" + + " 101 : AAA" + + hit_marker + + "CCC : 112\n", + ) +} + +///| +fn exonerate_text_test_protein_report() -> String { + exonerate_text_test_wrap( + "protein_query", + "", + "dna_target", + "", + "protein2dna:local", + 55, + 0, + 3, + 100, + 109, + " 1 : MetGly<->Lys : 3\n" + + " |||||| |||\n" + + " MetGly<->Lys\n" + + " 101 : ATGGGC---AAA : 109\n", + ) +} + +///| +fn exonerate_text_test_dna_to_protein_report() -> String { + exonerate_text_test_wrap( + "dna_query", + "", + "protein_target", + "", + "ungapped:dna2protein", + 44, + 0, + 9, + 10, + 13, + " 1 : ATGGGCAAA : 9\n" + + " MetGlyLys\n" + + " |||||||||\n" + + " 11 : MetGlyLys : 13\n", + ) +} + +///| +fn exonerate_text_test_frameshift_report() -> String { + exonerate_text_test_wrap( + "protein_query", + "", + "dna_target", + "", + "protein2dna:local", + 27, + 0, + 2, + 100, + 108, + " 1 : Asp--Ile : 2\n" + + " |||##|||\n" + + " Asp##Ile\n" + + " 101 : GATCCATT : 108\n", + ) +} + +///| +fn exonerate_text_test_wrapped_report() -> String { + exonerate_text_test_wrap( + "wrapped_query", + "", + "wrapped_target", + "", + "affine:local:dna2dna", + 36, + 0, + 8, + 100, + 108, + " 1 : ACGT : 4\n" + + " ||||\n" + + " 101 : ACGT : 104\n\n" + + " 5 : TGCA : 8\n" + + " ||||\n" + + " 105 : TGCA : 108\n", + ) +} + +///| +fn exonerate_text_test_joint_intron_report() -> String { + let marker = " >>>> Joint Intron 1 >>>> " + let similarity = "|||" + + exonerate_text_test_center("5 bp // 7 bp", marker.length()) + + "|||" + exonerate_text_test_wrap( + "joint_query", + "", + "joint_target", + "", + "est2genome", + 48, + 0, + 11, + 100, + 113, + " 1 : AAA" + + marker + + "CCC : 11\n" + + " " + + similarity + + "\n" + + " 101 : AAA" + + marker + + "CCC : 113\n", + ) +} + +///| +fn exonerate_text_test_reverse_intron_report() -> String { + let marker = " >>>> Target Intron 1 >>>> " + let splice = "gt" + ".".repeat(marker.length() - 4) + "ag" + let similarity = "|||" + + exonerate_text_test_center("7 bp", marker.length()) + + "|||" + exonerate_text_test_wrap( + "reverse_query", + "query strand:[revcomp]", + "reverse_target", + "target strand:[revcomp]", + "est2genome", + 41, + 10, + 4, + 200, + 187, + " 10 : ACG" + + marker + + "TAC : 5\n" + + " " + + similarity + + "\n" + + " 200 : ACG" + + splice + + "TAC : 188\n", + ) +} + +///| +fn exonerate_text_test_split_codon_report() -> String { + let marker = " >>>> Target Intron 1 >>>> " + let splice = "gt" + ".".repeat(marker.length() - 4) + "ag" + let similarity = "|||{||}" + + exonerate_text_test_center("7 bp", marker.length()) + + "{|}|||" + exonerate_text_test_wrap( + "protein_query", + "", + "genome_target", + "", + "protein2genome:local", + 73, + 0, + 3, + 100, + 116, + " 1 : Gly{Th}" + + marker + + "{r}Ala : 3\n" + + " " + + similarity + + "\n" + + " Gly{Th}" + + " ".repeat(marker.length()) + + "{r}Ala\n" + + " 101 : GGT{AC}" + + splice + + "{G}GCT : 116\n", + ) +} + +///| +fn exonerate_text_test_coding_report() -> String { + exonerate_text_test_wrap( + "coding_query", + "", + "coding_target", + "", + "coding2coding", + 52, + 0, + 6, + 100, + 106, + " 1 : ATGGAA : 6\n" + + " MetGlu\n" + + " ||||||\n" + + " MetGlu\n" + + " 101 : ATGGAA : 106\n", + ) +} + +///| +fn exonerate_text_test_coding_frameshift_report() -> String { + exonerate_text_test_wrap( + "coding_query", + "", + "coding_target", + "", + "coding2coding", + 39, + 0, + 6, + 100, + 106, + " 1 : ATGAC-A : 6\n" + + " Met#Thr\n" + + " |||#|||\n" + + " Met-Thr\n" + + " 101 : ATG-CTA : 106\n", + ) +} + +///| +fn exonerate_text_test_special_protein_report() -> String { + exonerate_text_test_wrap( + "special_protein", + "", + "special_dna", + "", + "protein2dna:local", + 29, + 0, + 3, + 100, + 109, + " 1 : ***UnkSec : 3\n" + + " |||||||||\n" + + " ***UnkSec\n" + + " 101 : TAANNNUGA : 109\n", + ) +} + +///| +fn exonerate_text_test_multi_report() -> String { + "Command line: [exonerate -m affine:local:dna2dna queries.fa targets.fa]\n" + + "Hostname: [aggregate-host]\n\n" + + exonerate_text_test_alignment( + "query1", "first query", "target1", "first target", "affine:local:dna2dna", 10, + 0, 3, 100, 103, " 1 : AAA : 3\n |||\n 101 : AAA : 103\n", + ) + + exonerate_text_test_alignment( + "query1", "first query", "target1", "first target", "affine:local:dna2dna", 20, + 3, 6, 200, 203, " 4 : CCC : 6\n |||\n 201 : CCC : 203\n", + ) + + exonerate_text_test_alignment( + "query1", "first query", "target2", "second target", "affine:local:dna2dna", + 15, 0, 3, 300, 303, " 1 : GGG : 3\n |||\n 301 : GGG : 303\n", + ) + + exonerate_text_test_alignment( + "query2", "second query", "target3", "third target", "affine:local:dna2dna", + 30, 0, 3, 400, 403, " 1 : TTT : 3\n |||\n 401 : TTT : 403\n", + ) + + "-- completed exonerate analysis\n" +} + +///| +test "Bio.SearchIO.ExonerateIO text parses metadata" { + let document = exonerate_text_test_parse(exonerate_text_test_dna_report()) + assert_eq(document.metadata.program, "exonerate") + assert_eq(document.metadata.hostname, "test-host") + assert_true(document.metadata.command_line.contains("affine:local:dna2dna")) +} + +///| +test "Bio.SearchIO.ExonerateIO text aggregates document shape" { + let document = exonerate_text_test_parse(exonerate_text_test_dna_report()) + assert_eq(document.num_queries(), 1) + assert_eq(document.num_hits(), 1) + assert_eq(document.num_hsps(), 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text preserves identifiers and descriptions" { + let query = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0] + assert_eq(query.id, "query1") + assert_eq(query.description, "demo query") + assert_eq(query.hits[0].id, "target1") + assert_eq(query.hits[0].description, "demo target") +} + +///| +test "Bio.SearchIO.ExonerateIO text parses DNA alignment rows" { + let fragment = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_sequence, "ACGT-AC") + assert_eq(fragment.similarity, "|||| ||") + assert_eq(fragment.hit_sequence, "ACGTTAC") +} + +///| +test "Bio.SearchIO.ExonerateIO text computes DNA coordinates" { + let hsp = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_start, 0) + assert_eq(hsp.query_end, 6) + assert_eq(hsp.hit_start, 100) + assert_eq(hsp.hit_end, 107) + assert_eq(hsp.fragments[0].query_start, 0) + assert_eq(hsp.fragments[0].query_end, 6) +} + +///| +test "Bio.SearchIO.ExonerateIO text reports strands" { + let hsp = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_strand, @src.ExonerateTextForward) + assert_eq(hsp.hit_strand, @src.ExonerateTextForward) +} + +///| +test "Bio.SearchIO.ExonerateIO text recognizes reverse complement suffix" { + let query = exonerate_text_test_parse(exonerate_text_test_reverse_report()).queries[0] + let hsp = query.hits[0].hsps[0] + assert_eq(query.hits[0].description, "reverse target") + assert_eq(hsp.hit_strand, @src.ExonerateTextReverse) + assert_eq(hsp.hit_start, 193) + assert_eq(hsp.hit_end, 200) +} + +///| +test "Bio.SearchIO.ExonerateIO text maps forward positions" { + let fragment = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_position(0), Some(0)) + assert_eq(fragment.query_position(4), None) + assert_eq(fragment.query_position(5), Some(4)) + assert_eq(fragment.hit_position(6), Some(106)) +} + +///| +test "Bio.SearchIO.ExonerateIO text maps reverse positions" { + let fragment = exonerate_text_test_parse(exonerate_text_test_reverse_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.hit_position(0), Some(199)) + assert_eq(fragment.hit_position(5), Some(194)) +} + +///| +test "Bio.SearchIO.ExonerateIO text splits target introns" { + let hsp = exonerate_text_test_parse(exonerate_text_test_intron_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.fragments[0].query_sequence, "ACG") + assert_eq(hsp.fragments[1].query_sequence, "TAC") +} + +///| +test "Bio.SearchIO.ExonerateIO text computes intron ranges" { + let hsp = exonerate_text_test_parse(exonerate_text_test_intron_report()).queries[0].hits[0].hsps[0] + let query_ranges = hsp.query_ranges() + let hit_ranges = hsp.hit_ranges() + assert_eq(query_ranges[0], @src.ExonerateTextRange::create(0, 3)) + assert_eq(query_ranges[1], @src.ExonerateTextRange::create(3, 6)) + assert_eq(hit_ranges[0], @src.ExonerateTextRange::create(100, 103)) + assert_eq(hit_ranges[1], @src.ExonerateTextRange::create(110, 113)) +} + +///| +test "Bio.SearchIO.ExonerateIO text exposes intron inter-ranges" { + let hsp = exonerate_text_test_parse(exonerate_text_test_intron_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_inter_ranges()[0], @src.ExonerateTextRange::create(3, 3)) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(103, 110), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text splits NER blocks" { + let hsp = exonerate_text_test_parse(exonerate_text_test_ner_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.fragments[0].query_sequence, "AAA") + assert_eq(hsp.fragments[1].hit_sequence, "CCC") +} + +///| +test "Bio.SearchIO.ExonerateIO text applies independent NER lengths" { + let hsp = exonerate_text_test_parse(exonerate_text_test_ner_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_inter_ranges()[0], @src.ExonerateTextRange::create(3, 7)) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(103, 109), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text converts protein2dna triplets" { + let fragment = exonerate_text_test_parse(exonerate_text_test_protein_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_sequence, "MG-K") + assert_eq(fragment.hit_sequence, "MG-K") + assert_eq(fragment.query_strand, @src.ExonerateTextProtein) + assert_eq(fragment.hit_strand, @src.ExonerateTextForward) +} + +///| +test "Bio.SearchIO.ExonerateIO text retains translated DNA annotation" { + let fragment = exonerate_text_test_parse(exonerate_text_test_protein_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.hit_annotation, Some("ATGGGC---AAA")) + assert_eq(fragment.query_step, 1) + assert_eq(fragment.hit_step, 3) +} + +///| +test "Bio.SearchIO.ExonerateIO text converts dna2protein triplets" { + let fragment = exonerate_text_test_parse( + exonerate_text_test_dna_to_protein_report(), + ).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_sequence, "MGK") + assert_eq(fragment.hit_sequence, "MGK") + assert_eq(fragment.query_annotation, Some("ATGGGCAAA")) + assert_eq(fragment.hit_strand, @src.ExonerateTextProtein) +} + +///| +test "Bio.SearchIO.ExonerateIO text splits frameshifts" { + let hsp = exonerate_text_test_parse(exonerate_text_test_frameshift_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.fragments[0].query_sequence, "D") + assert_eq(hsp.fragments[1].query_sequence, "I") + assert_eq(hsp.fragments[0].hit_sequence, "D") + assert_eq(hsp.fragments[1].hit_sequence, "I") +} + +///| +test "Bio.SearchIO.ExonerateIO text records frameshift coordinate gap" { + let hsp = exonerate_text_test_parse(exonerate_text_test_frameshift_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_inter_ranges()[0], @src.ExonerateTextRange::create(1, 1)) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(103, 105), + ) + assert_eq(hsp.fragments[0].hit_frame, 2) + assert_eq(hsp.fragments[1].hit_frame, 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text computes summary counts" { + let hsp = exonerate_text_test_parse(exonerate_text_test_dna_report()).queries[0].hits[0].hsps[0] + let counts = hsp.counts() + assert_eq(counts.fragments, 1) + assert_eq(counts.alignment_columns, 7) + assert_eq(counts.identities, 6) + assert_eq(counts.gap_columns, 1) + assert_eq(counts.query_gap_opens, 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text stitches wrapped physical blocks" { + let hsp = exonerate_text_test_parse(exonerate_text_test_wrapped_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 1) + assert_eq(hsp.fragments[0].query_sequence, "ACGTTGCA") + assert_eq(hsp.fragments[0].hit_sequence, "ACGTTGCA") + assert_eq(hsp.counts().alignment_columns, 8) +} + +///| +test "Bio.SearchIO.ExonerateIO text parses joint intron lengths" { + let hsp = exonerate_text_test_parse(exonerate_text_test_joint_intron_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.query_inter_ranges()[0], @src.ExonerateTextRange::create(3, 8)) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(103, 110), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text computes reverse intron coordinates" { + let hsp = exonerate_text_test_parse( + exonerate_text_test_reverse_intron_report(), + ).queries[0].hits[0].hsps[0] + assert_eq(hsp.query_strand, @src.ExonerateTextReverse) + assert_eq(hsp.hit_strand, @src.ExonerateTextReverse) + assert_eq(hsp.query_ranges(), [ + @src.ExonerateTextRange::create(7, 10), + @src.ExonerateTextRange::create(4, 7), + ]) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(190, 197), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text handles split protein codons" { + let hsp = exonerate_text_test_parse(exonerate_text_test_split_codon_report()).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.fragments[0].query_sequence, "GX") + assert_eq(hsp.fragments[1].query_sequence, "XA") + assert_eq(hsp.fragments[0].phase, 0) + assert_eq(hsp.fragments[1].phase, 1) + assert_eq(hsp.query_ranges(), [ + @src.ExonerateTextRange::create(0, 1), + @src.ExonerateTextRange::create(2, 3), + ]) + assert_eq(hsp.hit_ranges(), [ + @src.ExonerateTextRange::create(100, 105), + @src.ExonerateTextRange::create(112, 116), + ]) + assert_eq(hsp.hit_split_codons, [ + @src.ExonerateTextRange::create(103, 105), + @src.ExonerateTextRange::create(112, 113), + ]) +} + +///| +test "Bio.SearchIO.ExonerateIO text flips five-row coding alignments" { + let fragment = exonerate_text_test_parse(exonerate_text_test_coding_report()).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_sequence, "ATGGAA") + assert_eq(fragment.hit_sequence, "ATGGAA") + assert_eq(fragment.query_annotation, Some("MetGlu")) + assert_eq(fragment.hit_annotation, Some("MetGlu")) +} + +///| +test "Bio.SearchIO.ExonerateIO text assigns five-row frameshifts" { + let hsp = exonerate_text_test_parse( + exonerate_text_test_coding_frameshift_report(), + ).queries[0].hits[0].hsps[0] + assert_eq(hsp.num_fragments(), 2) + assert_eq(hsp.fragments[0].query_sequence, "ATG") + assert_eq(hsp.fragments[1].query_sequence, "C-A") + assert_eq(hsp.query_inter_ranges()[0], @src.ExonerateTextRange::create(3, 4)) + assert_eq( + hsp.hit_inter_ranges()[0], + @src.ExonerateTextRange::create(103, 103), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text maps special amino acids" { + let fragment = exonerate_text_test_parse( + exonerate_text_test_special_protein_report(), + ).queries[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_sequence, "*XU") + assert_eq(fragment.hit_sequence, "*XU") +} + +///| +test "Bio.SearchIO.ExonerateIO text aggregates query hit and HSP levels" { + let document = exonerate_text_test_parse(exonerate_text_test_multi_report()) + assert_eq(document.num_queries(), 2) + assert_eq(document.num_hits(), 3) + assert_eq(document.num_hsps(), 4) + assert_eq(document.queries[0].hits.length(), 2) + assert_eq(document.queries[0].hits[0].hsps.length(), 2) + assert_eq(document.best_hsp().unwrap().score, 30) +} + +///| +test "Bio.SearchIO.ExonerateIO text returns defensive hit copies" { + let document = exonerate_text_test_parse(exonerate_text_test_multi_report()) + let hit = document.queries[0].find_hit("target1").unwrap() + hit.hsps.clear() + assert_eq(document.queries[0].hits[0].hsps.length(), 2) +} + +///| +test "Bio.SearchIO.ExonerateIO text bundled example parses offline" { + let document = exonerate_text_test_parse(@src.exonerate_text_example()) + assert_eq(document.num_queries(), 1) + assert_eq(document.num_hsps(), 1) + assert_eq(document.queries[0].hits[0].hsps[0].num_fragments(), 2) +} + +///| +test "Bio.SearchIO.ExonerateIO text supports no-result reports" { + let document = exonerate_text_test_parse( + "Command line: [exonerate query.fa target.fa]\n" + + "Hostname: [none]\n\n" + + "-- completed exonerate analysis\n", + ) + assert_eq(document.num_queries(), 0) + assert_eq(document.num_hits(), 0) + assert_eq(document.best_hsp(), None) +} + +///| +test "Bio.SearchIO.ExonerateIO text supports CRLF" { + let text = exonerate_text_test_dna_report().replace_all(old="\n", new="\r\n") + assert_eq(exonerate_text_test_parse(text).num_hsps(), 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text returns defensive query copies" { + let document = exonerate_text_test_parse(exonerate_text_test_dna_report()) + let copy = document.find_query("query1").unwrap() + copy.hits.clear() + assert_eq(document.queries[0].hits.length(), 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text returns defensive best-HSP copies" { + let document = exonerate_text_test_parse(exonerate_text_test_dna_report()) + let copy = document.best_hsp().unwrap() + copy.fragments.clear() + assert_eq(document.queries[0].hits[0].hsps[0].fragments.length(), 1) +} + +///| +test "Bio.SearchIO.ExonerateIO text summarizes report shape" { + let summary = exonerate_text_test_parse(exonerate_text_test_dna_report()).summary() + assert_true(summary.contains("queries=1")) + assert_true(summary.contains("hits=1")) + assert_true(summary.contains("hsps=1")) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects missing metadata" { + assert_true(exonerate_text_test_rejects("-- completed exonerate analysis\n")) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects unfinished reports" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_dna_report().replace_all( + old="-- completed exonerate analysis", + new="", + ), + ), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects incomplete C4 headers" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_dna_report().replace_all( + old=" Model: affine:local:dna2dna\n", + new="", + ), + ), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects malformed ranges" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_dna_report().replace_all(old="0 -> 6", new="0 to 6"), + ), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects inconsistent range spans" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_dna_report().replace_all(old="0 -> 6", new="0 -> 7"), + ), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects reverse orientation mismatch" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_reverse_report().replace_all( + old="200 -> 193", + new="193 -> 200", + ), + ), + ) +} + +///| +test "Bio.SearchIO.ExonerateIO text rejects truncated physical rows" { + assert_true( + exonerate_text_test_rejects( + exonerate_text_test_dna_report().replace_all( + old=" |||| ||\n", + new=" |||\n", + ), + ), + ) +} diff --git a/test/moonbit/expasy_test.mbt b/test/moonbit/expasy_test.mbt index 46954e8e..ed587009 100644 --- a/test/moonbit/expasy_test.mbt +++ b/test/moonbit/expasy_test.mbt @@ -1,41 +1,45 @@ ///| /// Tests for ExPASy module. - test "ExPASyRecord creation" { let record = @src.ExPASyRecord::new("P04637", "Swiss-Prot") assert_eq(record.id, "P04637") assert_eq(record.database, "Swiss-Prot") } +///| test "ExPASyRecord add field" { let record = @src.ExPASyRecord::new("P04637", "Swiss-Prot") let record2 = record.add_field("description", "Tumor protein p53") - + let desc = record2.get_field("description") assert_true(desc is Some(_)) assert_eq(desc.unwrap(), "Tumor protein p53") } +///| test "ExPASyEntry creation" { let entry = @src.ExPASyEntry::new("P04637", "TP53") assert_eq(entry.accession, "P04637") assert_eq(entry.name, "TP53") } +///| test "ExPASyEntry add keyword" { let entry = @src.ExPASyEntry::new("P04637", "TP53") let entry2 = entry.add_keyword("Tumor suppressor") - + assert_eq(entry2.keywords.length(), 1) assert_eq(entry2.keywords[0], "Tumor suppressor") } +///| test "EnzymeEntry creation" { let enzyme = @src.EnzymeEntry::new("1.1.1.1", "Alcohol dehydrogenase") assert_eq(enzyme.ec_number, "1.1.1.1") assert_eq(enzyme.name, "Alcohol dehydrogenase") } +///| test "enzyme_parse_ec" { let (class, subclass, subsubclass, serial) = @src.enzyme_parse_ec("1.1.1.1") assert_eq(class, "1") @@ -44,35 +48,41 @@ test "enzyme_parse_ec" { assert_eq(serial, "1") } +///| test "expasy_get_prosite_ids" { let ids = @src.expasy_get_prosite_ids("PCNG") assert_true(ids.length() > 0) } +///| test "expasy_get_swissprot_entry" { let entry = @src.expasy_get_swissprot_entry("P04637") assert_true(entry is Some(_)) } +///| test "expasy_get_enzyme" { let enzyme = @src.expasy_get_enzyme("1.1.1.1") assert_true(enzyme is Some(_)) } +///| test "expasy_analyze_protein" { let results = @src.expasy_analyze_protein("AVG") assert_true(results.contains("molecular_weight")) assert_true(results.contains("gravy")) } +///| test "create_example_expasy_entry" { let entry = @src.create_example_expasy_entry() assert_eq(entry.accession, "P04637") assert_eq(entry.name, "TP53") } +///| test "create_example_enzyme_entry" { let enzyme = @src.create_example_enzyme_entry() assert_eq(enzyme.ec_number, "1.1.1.1") assert_eq(enzyme.name, "Alcohol dehydrogenase") -} \ No newline at end of file +} diff --git a/test/moonbit/factoextra_test.mbt b/test/moonbit/factoextra_test.mbt index 58d23123..a810ea59 100644 --- a/test/moonbit/factoextra_test.mbt +++ b/test/moonbit/factoextra_test.mbt @@ -17,11 +17,7 @@ test "factoextra_pca_basic" { ///| test "factoextra_pca_eigenvalues" { - let data = [ - [2.0, 3.0], - [5.0, 6.0], - [8.0, 9.0], - ] + let data = [[2.0, 3.0], [5.0, 6.0], [8.0, 9.0]] let result = @src.facto_pca(data) let eigenvalues = @src.facto_get_eigenvalue(result) assert_eq(eigenvalues.length(), 2) @@ -30,28 +26,22 @@ test "factoextra_pca_eigenvalues" { ///| test "factoextra_pca_variance" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - [10.0, 11.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0], [10.0, 11.0]] let result = @src.facto_pca(data) let eigenvalues = @src.facto_get_eigenvalue(result) let mut i = 1 while i < eigenvalues.length() { - assert_true(eigenvalues[i].cumulative_variance >= eigenvalues[i - 1].cumulative_variance) + assert_true( + eigenvalues[i].cumulative_variance >= + eigenvalues[i - 1].cumulative_variance, + ) i = i + 1 } } ///| test "factoextra_pca_individual_coords" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] let result = @src.facto_pca(data) let ind = @src.facto_get_pca_ind(result) assert_eq(ind.coord.length(), 3) @@ -60,11 +50,7 @@ test "factoextra_pca_individual_coords" { ///| test "factoextra_pca_variable_coords" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] let result = @src.facto_pca(data) let pca_var = @src.facto_get_pca_var(result) assert_eq(pca_var.coord.length(), 3) @@ -78,18 +64,14 @@ test "factoextra_pca_ncp" { [5.0, 6.0, 7.0, 8.0], [9.0, 10.0, 11.0, 12.0], ] - let result = @src.facto_pca(data, ncp = 2) + let result = @src.facto_pca(data, ncp=2) assert_eq(result.n_dims, 2) assert_eq(result.eigenvalues.length(), 2) } ///| test "factoextra_pca_cos2_individuals" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let ind = @src.facto_get_pca_ind(result) assert_eq(ind.cos2.length(), 3) @@ -107,11 +89,7 @@ test "factoextra_pca_cos2_individuals" { ///| test "factoextra_pca_contrib_variables" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let pca_var = @src.facto_get_pca_var(result) assert_eq(pca_var.contrib.length(), 2) @@ -120,11 +98,7 @@ test "factoextra_pca_contrib_variables" { ///| test "factoextra_pca_total_inertia" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let total = @src.facto_total_inertia(result) assert_true(total > 0.0) @@ -132,12 +106,7 @@ test "factoextra_pca_total_inertia" { ///| test "factoextra_pca_nb_dim" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - [10.0, 11.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0], [10.0, 11.0]] let result = @src.facto_pca(data) let nb = @src.facto_nb_dim(result, 80.0) assert_true(nb >= 1) @@ -146,10 +115,7 @@ test "factoextra_pca_nb_dim" { ///| test "factoextra_pca_summary" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let result = @src.facto_pca(data) let summary = @src.facto_summary(result) assert_eq(summary.n_individuals, 2) @@ -168,11 +134,7 @@ test "factoextra_pca_empty" { ///| test "factoextra_pca_contrib_dim" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let contribs = @src.facto_var_contrib_dim(result, 1) assert_eq(contribs.length(), 2) @@ -181,11 +143,7 @@ test "factoextra_pca_contrib_dim" { ///| test "factoextra_pca_cos2_dim" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let cos2s = @src.facto_var_cos2_dim(result, 1) assert_eq(cos2s.length(), 2) @@ -193,11 +151,7 @@ test "factoextra_pca_cos2_dim" { ///| test "factoextra_pca_dimdesc" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let desc = @src.facto_dimdesc(data, result, 1) assert_eq(desc.length(), 2) @@ -205,11 +159,7 @@ test "factoextra_pca_dimdesc" { ///| test "factoextra_pca_loadings" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - [7.0, 8.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]] let result = @src.facto_pca(data) let loadings = @src.facto_loadings(result) assert_eq(loadings.length(), 2) @@ -217,10 +167,7 @@ test "factoextra_pca_loadings" { ///| test "factoextra_pca_print_summary" { - let data = [ - [1.0, 2.0], - [4.0, 5.0], - ] + let data = [[1.0, 2.0], [4.0, 5.0]] let result = @src.facto_pca(data) let summary = @src.facto_print_summary(result) assert_true(summary.length() > 0) diff --git a/test/moonbit/feature_counts_test.mbt b/test/moonbit/feature_counts_test.mbt index 3f57568d..18cd06cc 100644 --- a/test/moonbit/feature_counts_test.mbt +++ b/test/moonbit/feature_counts_test.mbt @@ -1,9 +1,11 @@ ///| /// Test file for featureCounts module. - test "feature_annotation_creation" { let fa = @src.FeatureAnnotation::new( - chr="chr1", start=1000, end_=2000, strand="+", + chr="chr1", + start=1000, + end_=2000, + strand="+", gene_id="GeneA", ) assert_eq(fa.chr, "chr1") @@ -14,9 +16,13 @@ test "feature_annotation_creation" { assert_eq(fa.length(), 1001) } +///| test "feature_annotation_overlap" { let fa = @src.FeatureAnnotation::new( - chr="chr1", start=1000, end_=2000, strand="+", + chr="chr1", + start=1000, + end_=2000, + strand="+", gene_id="GeneA", ) assert_true(fa.overlaps(chr="chr1", start=1500, end_=1600)) @@ -26,9 +32,13 @@ test "feature_annotation_overlap" { assert_false(fa.overlaps(chr="chr1", start=2100, end_=2200)) } +///| test "feature_annotation_overlap_length" { let fa = @src.FeatureAnnotation::new( - chr="chr1", start=1000, end_=2000, strand="+", + chr="chr1", + start=1000, + end_=2000, + strand="+", gene_id="GeneA", ) assert_eq(fa.overlap_length(chr="chr1", start=1500, end_=1600), 101) @@ -37,9 +47,13 @@ test "feature_annotation_overlap_length" { assert_eq(fa.overlap_length(chr="chr2", start=1500, end_=1600), 0) } +///| test "read_alignment_creation" { let r = @src.ReadAlignment::new( - read_id="read1", chr="chr1", start=1500, end_=1600, + read_id="read1", + chr="chr1", + start=1500, + end_=1600, ) assert_eq(r.read_id, "read1") assert_eq(r.chr, "chr1") @@ -49,9 +63,13 @@ test "read_alignment_creation" { assert_eq(r.n_alignments, 1) } +///| test "read_alignment_paired" { let r = @src.ReadAlignment::new_paired( - read_id="read1", chr="chr1", start=1500, end_=1600, + read_id="read1", + chr="chr1", + start=1500, + end_=1600, mate_start=1700, ) assert_eq(r.is_paired, true) @@ -59,6 +77,7 @@ test "read_alignment_paired" { assert_true(r.fragment_length > 0) } +///| test "feature_counts_config_default" { let cfg = @src.FeatureCountsConfig::new() assert_eq(cfg.min_overlap, 1) @@ -66,6 +85,7 @@ test "feature_counts_config_default" { assert_eq(cfg.ignore_strand, false) } +///| test "feature_counts_config_setters" { let cfg = @src.FeatureCountsConfig::new() .set_min_overlap(val=5) @@ -76,16 +96,23 @@ test "feature_counts_config_setters" { assert_eq(cfg.count_fragments, true) } +///| test "feature_counts_count_basic" { let (features, reads, samples) = @src.feature_counts_sample_data() let cfg = @src.FeatureCountsConfig::new().set_min_mapq(val=40) - let result = @src.feature_counts_count(features, reads, sample_names=samples, config=cfg) + let result = @src.feature_counts_count( + features, + reads, + sample_names=samples, + config=cfg, + ) assert_eq(result.n_features(), 4) assert_eq(result.n_samples(), 1) assert_eq(result.total_reads[0], 7) assert_eq(result.assigned_reads[0], 5) } +///| test "feature_counts_get_count" { let (features, reads, samples) = @src.feature_counts_sample_data() let result = @src.feature_counts_count(features, reads, sample_names=samples) @@ -93,35 +120,56 @@ test "feature_counts_get_count" { assert_true(count >= 2.0) } +///| test "feature_counts_unassigned" { let (features, reads, samples) = @src.feature_counts_sample_data() let cfg = @src.FeatureCountsConfig::new().set_min_mapq(val=40) - let result = @src.feature_counts_count(features, reads, sample_names=samples, config=cfg) + let result = @src.feature_counts_count( + features, + reads, + sample_names=samples, + config=cfg, + ) assert_eq(result.unassigned_low_quality[0], 1) assert_eq(result.unassigned_no_features[0], 1) } +///| test "feature_counts_with_strand" { let (features, reads, samples) = @src.feature_counts_sample_data() - let cfg = @src.FeatureCountsConfig::new() - .set_strand_mode(mode=@src.strand_stranded()) - let result = @src.feature_counts_count(features, reads, sample_names=samples, config=cfg) + let cfg = @src.FeatureCountsConfig::new().set_strand_mode( + mode=@src.strand_stranded(), + ) + let result = @src.feature_counts_count( + features, + reads, + sample_names=samples, + config=cfg, + ) let count_b1 = result.get_count(feature_id="exon_B1", sample_idx=0) // read3 is on chr1:6500-6600, strand -, feature is on strand -, should match assert_true(count_b1 >= 1.0) } +///| test "feature_counts_reversed_strand" { let (features, reads, samples) = @src.feature_counts_sample_data() - let cfg = @src.FeatureCountsConfig::new() - .set_strand_mode(mode=@src.strand_reversed()) - let result = @src.feature_counts_count(features, reads, sample_names=samples, config=cfg) + let cfg = @src.FeatureCountsConfig::new().set_strand_mode( + mode=@src.strand_reversed(), + ) + let result = @src.feature_counts_count( + features, + reads, + sample_names=samples, + config=cfg, + ) // With reversed strand, read on + should NOT match feature on + // read1 (+) on chr1:1500 should NOT match exon_A1 (+) with reversed let count = result.get_count(feature_id="exon_A1", sample_idx=0) assert_true(count <= 1.0) } +///| test "feature_counts_library_size" { let (features, reads, samples) = @src.feature_counts_sample_data() let result = @src.feature_counts_count(features, reads, sample_names=samples) @@ -129,6 +177,7 @@ test "feature_counts_library_size" { assert_true(lib_size > 0.0) } +///| test "feature_counts_cpm" { let (features, reads, samples) = @src.feature_counts_sample_data() let result = @src.feature_counts_count(features, reads, sample_names=samples) @@ -136,6 +185,7 @@ test "feature_counts_cpm" { assert_true(cpm_val >= 0.0) } +///| test "feature_counts_summary" { let (features, reads, samples) = @src.feature_counts_sample_data() let result = @src.feature_counts_count(features, reads, sample_names=samples) @@ -143,26 +193,43 @@ test "feature_counts_summary" { assert_true(summary.length() > 0) } +///| test "feature_counts_multi_sample" { let features : Array[@src.FeatureAnnotation] = Array::new() - features.push(@src.FeatureAnnotation::new( - chr="chr1", start=1000, end_=2000, strand="+", - gene_id="GeneA", feature_id="exon_A1", - )) + features.push( + @src.FeatureAnnotation::new( + chr="chr1", + start=1000, + end_=2000, + strand="+", + gene_id="GeneA", + feature_id="exon_A1", + ), + ) let reads : Array[@src.ReadAlignment] = Array::new() - reads.push(@src.ReadAlignment::new( - read_id="read_s1_1", chr="chr1", start=1500, end_=1600, - )) - reads.push(@src.ReadAlignment::new( - read_id="read_s2_1", chr="chr1", start=1500, end_=1600, - )) + reads.push( + @src.ReadAlignment::new( + read_id="read_s1_1", + chr="chr1", + start=1500, + end_=1600, + ), + ) + reads.push( + @src.ReadAlignment::new( + read_id="read_s2_1", + chr="chr1", + start=1500, + end_=1600, + ), + ) let sample_names : Array[String] = Array::new() sample_names.push("sample1") sample_names.push("sample2") - let result = @src.feature_counts_count(features, reads, sample_names=sample_names) + let result = @src.feature_counts_count(features, reads, sample_names~) assert_eq(result.n_samples(), 2) assert_eq(result.total_reads[0], 1) assert_eq(result.total_reads[1], 1) @@ -170,6 +237,7 @@ test "feature_counts_multi_sample" { assert_eq(result.get_count(feature_id="exon_A1", sample_idx=1), 1.0) } +///| test "strand_mode_creation" { let s1 = @src.strand_unstranded() let s2 = @src.strand_stranded() @@ -180,6 +248,7 @@ test "strand_mode_creation" { assert_true(s3 == s3) } +///| test "feature_counts_feature_total" { let (features, reads, samples) = @src.feature_counts_sample_data() let result = @src.feature_counts_count(features, reads, sample_names=samples) diff --git a/test/moonbit/file_test.mbt b/test/moonbit/file_test.mbt index dcbe1f11..bf7d3d08 100644 --- a/test/moonbit/file_test.mbt +++ b/test/moonbit/file_test.mbt @@ -51,7 +51,11 @@ test "smart_file_creation_write" { ///| test "smart_file_with_format" { - let sf = @src.SmartFile::with_format("data.gz", @src.compression_gzip(), mode="r") + let sf = @src.SmartFile::with_format( + "data.gz", + @src.compression_gzip(), + mode="r", + ) assert_eq(sf.get_format().to_string(), "gzip") } diff --git a/test/moonbit/fishpond_test.mbt b/test/moonbit/fishpond_test.mbt index 97acc4e6..9120bd34 100644 --- a/test/moonbit/fishpond_test.mbt +++ b/test/moonbit/fishpond_test.mbt @@ -16,10 +16,9 @@ test "fish_counts_creation" { assert_eq(c.condition()[8], "control") } +///| test "fish_result_creation" { - let r = @src.FishResult::new( - "TX1", 15.0, 1.5, 0.01, 0.05, "up", - ) + let r = @src.FishResult::new("TX1", 15.0, 1.5, 0.01, 0.05, "up") assert_eq(r.transcript(), "TX1") assert_eq(r.statistic(), 15.0) assert_eq(r.log2_fold_change(), 1.5) @@ -32,35 +31,30 @@ test "fish_result_creation" { // Mann-Whitney-Wilcoxon statistic // --------------------------------------------------------------------------- +///| test "fish_mw_statistic_clear_separation" { // Cases all higher than controls - let stat = @src.fish_mann_whitney_test( - [10.0, 20.0, 30.0], - [1.0, 2.0, 3.0], - ) + let stat = @src.fish_mann_whitney_test([10.0, 20.0, 30.0], [1.0, 2.0, 3.0]) // With complete separation, W = sum of ranks 4,5,6 = 15 assert_eq(stat, 15.0) } +///| test "fish_mw_statistic_identical_groups" { - let stat = @src.fish_mann_whitney_test( - [5.0, 5.0, 5.0], - [5.0, 5.0, 5.0], - ) + let stat = @src.fish_mann_whitney_test([5.0, 5.0, 5.0], [5.0, 5.0, 5.0]) // All tied: average rank = 3.5 for each, sum for case = 3*3.5 = 10.5 assert_eq(stat, 10.5) } +///| test "fish_mw_statistic_overlap" { - let stat = @src.fish_mann_whitney_test( - [1.0, 3.0, 5.0], - [2.0, 4.0, 6.0], - ) + let stat = @src.fish_mann_whitney_test([1.0, 3.0, 5.0], [2.0, 4.0, 6.0]) // Ranks: 1->1(case), 2->2(ctrl), 3->3(case), 4->4(ctrl), 5->5(case), 6->6(ctrl) // W_case = 1+3+5 = 9 assert_eq(stat, 9.0) } +///| test "fish_mw_statistic_empty_group" { let stat = @src.fish_mann_whitney_test([], [1.0, 2.0]) assert_eq(stat, 0.0) @@ -70,17 +64,20 @@ test "fish_mw_statistic_empty_group" { // Log2 fold change // --------------------------------------------------------------------------- +///| test "fish_log2fc_positive" { let lfc = @src.fish_log2fc_test([8.0, 8.0], [2.0, 2.0]) // log2(9) - log2(3) = log2(3) ≈ 1.585 assert_true(lfc > 0.0) } +///| test "fish_log2fc_negative" { let lfc = @src.fish_log2fc_test([2.0, 2.0], [8.0, 8.0]) assert_true(lfc < 0.0) } +///| test "fish_log2fc_zero_means" { let lfc = @src.fish_log2fc_test([0.0, 0.0], [0.0, 0.0]) assert_eq(lfc, 0.0) @@ -90,6 +87,7 @@ test "fish_log2fc_zero_means" { // Full Swish analysis // --------------------------------------------------------------------------- +///| test "fish_swish_detects_upregulated" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -102,6 +100,7 @@ test "fish_swish_detects_upregulated" { assert_eq(results[2].direction(), "up") } +///| test "fish_swish_detects_downregulated" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -114,6 +113,7 @@ test "fish_swish_detects_downregulated" { assert_eq(results[5].direction(), "down") } +///| test "fish_swish_nonsignificant_stable" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -123,6 +123,7 @@ test "fish_swish_nonsignificant_stable" { } } +///| test "fish_swish_pvalues_in_range" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -132,6 +133,7 @@ test "fish_swish_pvalues_in_range" { } } +///| test "fish_swish_fdr_in_range" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -141,23 +143,21 @@ test "fish_swish_fdr_in_range" { } } +///| test "fish_swish_empty_input" { let counts = @src.FishCounts::new([], [], [], []) let results = @src.fish_swish(counts, "case", 10) assert_eq(results.length(), 0) } +///| test "fish_swish_too_few_samples" { - let counts = @src.FishCounts::new( - ["TX1"], - ["s1"], - ["case"], - [[5.0]], - ) + let counts = @src.FishCounts::new(["TX1"], ["s1"], ["case"], [[5.0]]) let results = @src.fish_swish(counts, "case", 10) assert_eq(results.length(), 0) } +///| test "fish_swish_log2fc_sign_correct" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -175,6 +175,7 @@ test "fish_swish_log2fc_sign_correct" { // Significant filtering // --------------------------------------------------------------------------- +///| test "fish_significant_filter" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -183,6 +184,7 @@ test "fish_significant_filter" { assert_true(sig.length() >= 6) } +///| test "fish_significant_strict_threshold" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -195,15 +197,15 @@ test "fish_significant_strict_threshold" { // Output formatting // --------------------------------------------------------------------------- +///| test "fish_result_to_string" { - let r = @src.FishResult::new( - "TX1", 15.0, 1.5, 0.01, 0.05, "up", - ) + let r = @src.FishResult::new("TX1", 15.0, 1.5, 0.01, 0.05, "up") let s = r.to_string() assert_true(s.contains("TX1")) assert_true(s.contains("up")) } +///| test "fish_results_to_string" { let counts = @src.fish_sample_counts() let results = @src.fish_swish(counts, "case", 500) @@ -216,28 +218,95 @@ test "fish_results_to_string" { // Sample data verification // --------------------------------------------------------------------------- +///| test "fish_sample_data_up_pattern" { let c = @src.fish_sample_counts() // TX1 (transcript 0): case values (indices 0-7) should be higher than control (indices 8-15) - let case_mean = (c.counts()[0][0] + c.counts()[0][1] + c.counts()[0][2] + c.counts()[0][3] + c.counts()[0][4] + c.counts()[0][5] + c.counts()[0][6] + c.counts()[0][7]) / 8.0 - let ctrl_mean = (c.counts()[0][8] + c.counts()[0][9] + c.counts()[0][10] + c.counts()[0][11] + c.counts()[0][12] + c.counts()[0][13] + c.counts()[0][14] + c.counts()[0][15]) / 8.0 + let case_mean = ( + c.counts()[0][0] + + c.counts()[0][1] + + c.counts()[0][2] + + c.counts()[0][3] + + c.counts()[0][4] + + c.counts()[0][5] + + c.counts()[0][6] + + c.counts()[0][7] + ) / + 8.0 + let ctrl_mean = ( + c.counts()[0][8] + + c.counts()[0][9] + + c.counts()[0][10] + + c.counts()[0][11] + + c.counts()[0][12] + + c.counts()[0][13] + + c.counts()[0][14] + + c.counts()[0][15] + ) / + 8.0 assert_true(case_mean > ctrl_mean) } +///| test "fish_sample_data_down_pattern" { let c = @src.fish_sample_counts() // TX4 (transcript 3): case values should be lower than control - let case_mean = (c.counts()[3][0] + c.counts()[3][1] + c.counts()[3][2] + c.counts()[3][3] + c.counts()[3][4] + c.counts()[3][5] + c.counts()[3][6] + c.counts()[3][7]) / 8.0 - let ctrl_mean = (c.counts()[3][8] + c.counts()[3][9] + c.counts()[3][10] + c.counts()[3][11] + c.counts()[3][12] + c.counts()[3][13] + c.counts()[3][14] + c.counts()[3][15]) / 8.0 + let case_mean = ( + c.counts()[3][0] + + c.counts()[3][1] + + c.counts()[3][2] + + c.counts()[3][3] + + c.counts()[3][4] + + c.counts()[3][5] + + c.counts()[3][6] + + c.counts()[3][7] + ) / + 8.0 + let ctrl_mean = ( + c.counts()[3][8] + + c.counts()[3][9] + + c.counts()[3][10] + + c.counts()[3][11] + + c.counts()[3][12] + + c.counts()[3][13] + + c.counts()[3][14] + + c.counts()[3][15] + ) / + 8.0 assert_true(case_mean < ctrl_mean) } +///| test "fish_sample_data_ns_pattern" { let c = @src.fish_sample_counts() // TX7 (transcript 6): similar values in both groups - let case_mean = (c.counts()[6][0] + c.counts()[6][1] + c.counts()[6][2] + c.counts()[6][3] + c.counts()[6][4] + c.counts()[6][5] + c.counts()[6][6] + c.counts()[6][7]) / 8.0 - let ctrl_mean = (c.counts()[6][8] + c.counts()[6][9] + c.counts()[6][10] + c.counts()[6][11] + c.counts()[6][12] + c.counts()[6][13] + c.counts()[6][14] + c.counts()[6][15]) / 8.0 + let case_mean = ( + c.counts()[6][0] + + c.counts()[6][1] + + c.counts()[6][2] + + c.counts()[6][3] + + c.counts()[6][4] + + c.counts()[6][5] + + c.counts()[6][6] + + c.counts()[6][7] + ) / + 8.0 + let ctrl_mean = ( + c.counts()[6][8] + + c.counts()[6][9] + + c.counts()[6][10] + + c.counts()[6][11] + + c.counts()[6][12] + + c.counts()[6][13] + + c.counts()[6][14] + + c.counts()[6][15] + ) / + 8.0 // Difference should be small - let diff = if case_mean > ctrl_mean { case_mean - ctrl_mean } else { ctrl_mean - case_mean } + let diff = if case_mean > ctrl_mean { + case_mean - ctrl_mean + } else { + ctrl_mean - case_mean + } assert_true(diff < 10.0) } diff --git a/test/moonbit/flowsom_test.mbt b/test/moonbit/flowsom_test.mbt new file mode 100644 index 00000000..b7689a57 --- /dev/null +++ b/test/moonbit/flowsom_test.mbt @@ -0,0 +1,1188 @@ +///| +fn flowsom_test_close( + left : Double, + right : Double, + tolerance : Double, +) -> Bool { + (left - right).abs() <= tolerance +} + +///| +fn flowsom_test_some(value : Double?) -> Double { + match value { + Some(actual) => actual + None => abort("expected a numeric value") + } +} + +///| +fn flowsom_test_unique_count(values : Array[Int]) -> Int { + let unique : Array[Int] = [] + for value in values { + if !unique.contains(value) { + unique.push(value) + } + } + unique.length() +} + +///| +fn flowsom_test_sum(values : Array[Int]) -> Int { + let mut total = 0 + for value in values { + total = total + value + } + total +} + +///| +fn flowsom_test_data() -> Array[Array[Double]] { + [ + [0.0, 0.0], + [0.1, 0.2], + [0.2, 0.1], + [0.3, 0.2], + [9.8, 10.0], + [10.0, 9.9], + [10.1, 10.2], + [10.3, 10.1], + ] +} + +///| +fn flowsom_test_config( + metaclusters? : Int = 0, + seed? : Int = 31, +) -> @src.FlowSomConfig { + @src.FlowSomConfig::create( + xdim=2, + ydim=2, + rlen=3, + mst_runs=2, + alpha_start=0.1, + alpha_end=0.02, + radius_start=1.0, + radius_end=0.0, + metaclusters~, + meta_starts=3, + seed~, + ) catch { + _ => abort("valid FlowSOM test configuration should build") + } +} + +///| +fn flowsom_test_model( + metaclusters? : Int = 0, + seed? : Int = 31, +) -> @src.FlowSomModel { + @src.flowsom_train( + flowsom_test_data(), + flowsom_test_config(metaclusters~, seed~), + marker_names=["CD3", "CD19"], + ) catch { + _ => abort("valid FlowSOM test model should train") + } +} + +///| +fn flowsom_test_one_node_model(metaclusters? : Int = 0) -> @src.FlowSomModel { + let config = @src.FlowSomConfig::create( + xdim=1, + ydim=1, + rlen=2, + alpha_start=0.1, + alpha_end=0.02, + radius_start=0.0, + radius_end=0.0, + metaclusters~, + seed=41, + ) catch { + _ => abort("valid one-node FlowSOM configuration should build") + } + @src.flowsom_train([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]], config, marker_names=[ + "A", "B", + ]) catch { + _ => abort("valid one-node FlowSOM model should train") + } +} + +///| +fn flowsom_test_sce() -> @src.SingleCellExperiment { + let experiment = @src.SingleCellExperiment::new( + [ + [0.0, 0.1, 0.2, 10.0, 10.1, 10.2], + [1.0, 1.1, 1.2, 11.0, 11.1, 11.2], + [2.0, 2.1, 2.2, 12.0, 12.1, 12.2], + ], + ["CD3", "CD19", "CD45"], + ["C1", "C2", "C3", "C4", "C5", "C6"], + ) + experiment.row_data["symbol"] = ["T", "B", "pan"] + experiment.col_data["batch"] = ["A", "A", "A", "B", "B", "B"] + experiment.reduced_dims["PCA"] = [ + [0.0, 0.0], + [0.1, 0.0], + [0.2, 0.1], + [10.0, 10.0], + [10.1, 10.0], + [10.2, 10.1], + ] + experiment.metadata["project"] = "flowsom-test" + experiment.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[5.0, 6.0, 7.0, 8.0, 9.0, 10.0]], + ["Spike1"], + ["C1", "C2", "C3", "C4", "C5", "C6"], + ) + experiment +} + +///| +test "FlowSOM: distance and initialization accessors expose enum values" { + assert_true( + @src.flowsom_manhattan_distance() != @src.flowsom_euclidean_distance(), + ) + assert_true( + @src.flowsom_euclidean_distance() != @src.flowsom_chebyshev_distance(), + ) + assert_true( + @src.flowsom_chebyshev_distance() != @src.flowsom_cosine_distance(), + ) + assert_true( + @src.flowsom_random_initialization() != @src.flowsom_kwsp_initialization(), + ) + assert_true( + @src.flowsom_kwsp_initialization() != @src.flowsom_pca_initialization(), + ) +} + +///| +test "FlowSOM: configuration defaults match the upstream workflow" { + let config = @src.FlowSomConfig::create() catch { + _ => abort("default FlowSOM configuration should build") + } + assert_eq(config.xdim, 10) + assert_eq(config.ydim, 10) + assert_eq(config.rlen, 10) + assert_eq(config.mst_runs, 1) + assert_eq(config.alpha_start, 0.05) + assert_eq(config.alpha_end, 0.01) + assert_eq(config.radius_start, -1.0) + assert_eq(config.radius_end, 0.0) + assert_true(config.distance == @src.flowsom_euclidean_distance()) + assert_true(config.initialization == @src.flowsom_kwsp_initialization()) + assert_eq(config.metaclusters, 0) + assert_eq(config.seed, 1) +} + +///| +test "FlowSOM: configuration preserves explicit controls" { + let config = @src.FlowSomConfig::create( + xdim=3, + ydim=2, + rlen=7, + mst_runs=3, + alpha_start=0.2, + alpha_end=0.03, + radius_start=2.0, + radius_end=0.5, + distance=@src.flowsom_cosine_distance(), + initialization=@src.flowsom_pca_initialization(), + importance=[2.0, 0.5], + metaclusters=3, + meta_starts=4, + outlier_mad=2.5, + seed=17, + ) catch { + _ => abort("explicit FlowSOM configuration should build") + } + assert_eq(config.xdim, 3) + assert_eq(config.ydim, 2) + assert_eq(config.rlen, 7) + assert_eq(config.mst_runs, 3) + assert_eq(config.importance, [2.0, 0.5]) + assert_eq(config.metaclusters, 3) + assert_eq(config.meta_starts, 4) + assert_eq(config.outlier_mad, 2.5) + assert_eq(config.seed, 17) +} + +///| +test "FlowSOM: configuration rejects invalid grid and run controls" { + let mut failures = 0 + ignore(@src.FlowSomConfig::create(xdim=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(ydim=-1)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(rlen=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(mst_runs=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "FlowSOM: configuration rejects invalid learning schedules" { + let mut failures = 0 + ignore(@src.FlowSomConfig::create(alpha_start=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(alpha_start=0.01, alpha_end=0.02)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(alpha_end=-0.01)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(alpha_start=1.0e301)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "FlowSOM: configuration rejects invalid radius schedules" { + let mut failures = 0 + ignore(@src.FlowSomConfig::create(radius_start=-2.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(radius_end=-0.1)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(radius_start=1.0, radius_end=2.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: configuration rejects invalid meta, outlier and seed controls" { + let mut failures = 0 + ignore(@src.FlowSomConfig::create(xdim=2, ydim=2, metaclusters=5)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(meta_starts=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(outlier_mad=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(seed=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.FlowSomConfig::create(importance=[1.0, 0.0])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 5) +} + +///| +test "FlowSOM: four distances follow their definitions" { + let left = [1.0, -2.0, 3.0] + let right = [4.0, 2.0, 3.0] + assert_eq( + @src.flowsom_distance(left, right, @src.flowsom_manhattan_distance()), + 7.0, + ) + assert_eq( + @src.flowsom_distance(left, right, @src.flowsom_euclidean_distance()), + 5.0, + ) + assert_eq( + @src.flowsom_distance(left, right, @src.flowsom_chebyshev_distance()), + 4.0, + ) + let cosine = @src.flowsom_distance( + [1.0, 0.0], + [0.0, 1.0], + @src.flowsom_cosine_distance(), + ) + assert_true(flowsom_test_close(cosine, 1.0, 1.0e-12)) +} + +///| +test "FlowSOM: cosine distance handles zero vectors" { + let cosine = @src.flowsom_cosine_distance() + assert_eq(@src.flowsom_distance([0.0, 0.0], [0.0, 0.0], cosine), 0.0) + assert_eq(@src.flowsom_distance([0.0, 0.0], [1.0, 0.0], cosine), 1.0) + assert_true( + flowsom_test_close( + @src.flowsom_distance([1.0, 1.0], [2.0, 2.0], cosine), + 0.0, + 1.0e-12, + ), + ) +} + +///| +test "FlowSOM: distance validates vector dimensions and values" { + let mut failures = 0 + ignore(@src.flowsom_distance([], [], @src.flowsom_euclidean_distance())) catch { + _ => failures = failures + 1 + } + ignore( + @src.flowsom_distance([1.0], [1.0, 2.0], @src.flowsom_euclidean_distance()), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.flowsom_distance([1.0e301], [0.0], @src.flowsom_euclidean_distance()), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: KWSP initialization repeatedly selects the farthest cell" { + let codes = @src.flowsom_initialize_kwsp([[0.0], [1.0], [10.0]], 3, seed=1) catch { + _ => abort("KWSP initialization should work") + } + assert_eq(codes, [[0.0], [10.0], [1.0]]) +} + +///| +test "FlowSOM: KWSP initialization is reproducible and copies rows" { + let data = [[0.0, 0.0], [1.0, 1.0], [5.0, 5.0], [10.0, 10.0]] + let first = @src.flowsom_initialize_kwsp(data, 3, seed=23) catch { + _ => abort("first KWSP initialization should work") + } + let second = @src.flowsom_initialize_kwsp(data, 3, seed=23) catch { + _ => abort("second KWSP initialization should work") + } + assert_eq(first, second) + first[0][0] = 999.0 + assert_true(data[0][0] != 999.0) + assert_true(data[1][0] != 999.0) + assert_true(data[2][0] != 999.0) + assert_true(data[3][0] != 999.0) +} + +///| +test "FlowSOM: KWSP initialization validates nodes and seed" { + let data = [[0.0], [1.0]] + let mut failures = 0 + ignore(@src.flowsom_initialize_kwsp(data, 0)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_initialize_kwsp(data, 3)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_initialize_kwsp(data, 1, seed=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: one-node PCA initialization returns the column means" { + let codes = @src.flowsom_initialize_pca( + [[1.0, 2.0], [3.0, 6.0], [5.0, 10.0]], + 1, + 1, + ) catch { + _ => abort("one-node PCA initialization should work") + } + assert_eq(codes.length(), 1) + assert_true(flowsom_test_close(codes[0][0], 3.0, 1.0e-12)) + assert_true(flowsom_test_close(codes[0][1], 6.0, 1.0e-12)) +} + +///| +test "FlowSOM: PCA grid is symmetric around the data mean" { + let codes = @src.flowsom_initialize_pca( + [[0.0, 0.0], [2.0, 1.0], [4.0, 2.0]], + 2, + 1, + ) catch { + _ => abort("PCA grid initialization should work") + } + assert_eq(codes.length(), 2) + assert_eq(codes[0].length(), 2) + assert_true( + flowsom_test_close((codes[0][0] + codes[1][0]) / 2.0, 2.0, 1.0e-9), + ) + assert_true( + flowsom_test_close((codes[0][1] + codes[1][1]) / 2.0, 1.0, 1.0e-9), + ) +} + +///| +test "FlowSOM: PCA initialization validates grid dimensions" { + let mut failures = 0 + ignore(@src.flowsom_initialize_pca([[1.0]], 0, 1)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_initialize_pca([[1.0]], 1, 0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: MST connects a one-dimensional chain with minimum weight" { + let edges = @src.flowsom_build_mst([[0.0], [1.0], [3.0], [7.0]]) catch { + _ => abort("MST construction should work") + } + assert_eq(edges.length(), 3) + assert_eq(edges[0].from, 0) + assert_eq(edges[0].to, 1) + assert_eq(edges[0].weight, 1.0) + assert_eq(edges[1].from, 1) + assert_eq(edges[1].to, 2) + assert_eq(edges[1].weight, 2.0) + assert_eq(edges[2].from, 2) + assert_eq(edges[2].to, 3) + assert_eq(edges[2].weight, 4.0) +} + +///| +test "FlowSOM: MST tie breaking is stable by node index" { + let edges = @src.flowsom_build_mst([[0.0], [1.0], [-1.0]]) catch { + _ => abort("tied MST construction should work") + } + assert_eq(edges.length(), 2) + assert_eq(edges[0].from, 0) + assert_eq(edges[0].to, 1) + assert_eq(edges[1].from, 0) + assert_eq(edges[1].to, 2) +} + +///| +test "FlowSOM: one-node MST has no edges" { + let edges = @src.flowsom_build_mst([[2.0, 3.0]]) catch { + _ => abort("one-node MST construction should work") + } + assert_eq(edges.length(), 0) +} + +///| +test "FlowSOM: MST topology reports unweighted path lengths" { + let distances = @src.flowsom_mst_distances([[0.0], [1.0], [3.0], [7.0]]) catch { + _ => abort("MST topology should work") + } + assert_eq(distances.length(), 4) + assert_eq(distances[0], [0, 1, 2, 3]) + assert_eq(distances[3], [3, 2, 1, 0]) +} + +///| +test "FlowSOM: MST validates rectangular finite codebooks" { + let mut failures = 0 + ignore(@src.flowsom_build_mst([])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_build_mst([[1.0], [1.0, 2.0]])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_build_mst([[1.0e301]])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: data mapping returns nearest nodes and distances" { + let mapping = @src.flowsom_map_data([[0.0, 0.0], [10.0, 10.0]], [ + [1.0, 0.0], + [9.0, 10.0], + ]) catch { + _ => abort("FlowSOM mapping should work") + } + assert_eq(mapping.clusters, [0, 1]) + assert_eq(mapping.distances, [1.0, 1.0]) +} + +///| +test "FlowSOM: data mapping resolves nearest-node ties by index" { + let mapping = @src.flowsom_map_data([[0.0], [2.0]], [[1.0]]) catch { + _ => abort("tied FlowSOM mapping should work") + } + assert_eq(mapping.clusters, [0]) + assert_eq(mapping.distances, [1.0]) +} + +///| +test "FlowSOM: data mapping supports non-Euclidean distance" { + let mapping = @src.flowsom_map_data( + [[0.0, 0.0], [4.0, 4.0]], + [[3.0, 0.0]], + distance=@src.flowsom_manhattan_distance(), + ) catch { + _ => abort("Manhattan FlowSOM mapping should work") + } + assert_eq(mapping.clusters, [0]) + assert_eq(mapping.distances, [3.0]) +} + +///| +test "FlowSOM: data mapping validates marker dimensions" { + let mut raised = false + ignore(@src.flowsom_map_data([[0.0, 1.0]], [[0.0]])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "FlowSOM: training returns a complete grid, MST and mapping" { + let model = flowsom_test_model() + assert_eq(model.data.length(), 8) + assert_eq(model.marker_names, ["CD3", "CD19"]) + assert_eq(model.codes.length(), 4) + assert_eq(model.grid.length(), 4) + assert_eq(model.grid[0], (0, 0)) + assert_eq(model.grid[3], (1, 1)) + assert_eq(model.mapping.clusters.length(), 8) + assert_eq(model.mapping.distances.length(), 8) + assert_eq(model.mst_edges.length(), 3) + assert_eq(model.topology_distances.length(), 4) + assert_eq(flowsom_test_sum(model.node_counts), 8) +} + +///| +test "FlowSOM: fixed seeds make training reproducible" { + let first = flowsom_test_model(seed=37) + let second = flowsom_test_model(seed=37) + assert_eq(first.codes, second.codes) + assert_eq(first.mapping.clusters, second.mapping.clusters) + assert_eq(first.mapping.distances, second.mapping.distances) + assert_eq(first.iterations, second.iterations) +} + +///| +test "FlowSOM: random initialization training path is available" { + let config = @src.FlowSomConfig::create( + xdim=2, + ydim=2, + rlen=2, + initialization=@src.flowsom_random_initialization(), + seed=43, + ) catch { + _ => abort("random FlowSOM configuration should build") + } + let model = @src.flowsom_train(flowsom_test_data(), config) catch { + _ => abort("random-initialized FlowSOM should train") + } + assert_eq(model.codes.length(), 4) + assert_eq(model.mapping.clusters.length(), 8) +} + +///| +test "FlowSOM: PCA training permits more nodes than cells" { + let config = @src.FlowSomConfig::create( + xdim=2, + ydim=2, + rlen=1, + initialization=@src.flowsom_pca_initialization(), + seed=47, + ) catch { + _ => abort("PCA FlowSOM configuration should build") + } + let model = @src.flowsom_train([[0.0, 0.0], [1.0, 1.0]], config) catch { + _ => abort("PCA FlowSOM with extra nodes should train") + } + assert_eq(model.codes.length(), 4) + assert_eq(model.mapping.clusters.length(), 2) +} + +///| +test "FlowSOM: multi-stage training tracks iterations and final topology" { + let model = flowsom_test_model() + assert_true(model.iterations > 0) + assert_true(model.iterations <= 3 * 8 * 2) + for node in 0.. abort("FlowSOM with generated marker names should train") + } + assert_eq(model.marker_names, ["marker.1", "marker.2"]) +} + +///| +test "FlowSOM: training validates marker names" { + let config = flowsom_test_config() + let mut failures = 0 + ignore(@src.flowsom_train(flowsom_test_data(), config, marker_names=["only"])) catch { + _ => failures = failures + 1 + } + ignore( + @src.flowsom_train(flowsom_test_data(), config, marker_names=["A", "A"]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.flowsom_train(flowsom_test_data(), config, marker_names=["A", ""]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: training validates importance dimensions" { + let config = @src.FlowSomConfig::create(xdim=1, ydim=1, importance=[1.0]) catch { + _ => abort("importance configuration should build before data validation") + } + let mut raised = false + ignore(@src.flowsom_train([[1.0, 2.0]], config)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "FlowSOM: training rejects empty, ragged and non-finite data" { + let config = flowsom_test_config() + let mut failures = 0 + ignore(@src.flowsom_train([], config)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_train([[1.0], [1.0, 2.0]], config)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_train([[1.0e301, 0.0]], config)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "FlowSOM: non-PCA initialization requires one cell per node" { + let config = @src.FlowSomConfig::create(xdim=2, ydim=2) catch { + _ => abort("valid FlowSOM configuration should build") + } + let mut raised = false + ignore(@src.flowsom_train([[0.0], [1.0]], config)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "FlowSOM: node summaries use medians, sample SD, CV and R-style MAD" { + let model = flowsom_test_one_node_model() + assert_eq(model.node_counts, [3]) + assert_eq(model.node_percentages, [1.0]) + assert_eq(flowsom_test_some(model.node_medians[0][0]), 3.0) + assert_eq(flowsom_test_some(model.node_medians[0][1]), 4.0) + assert_true( + flowsom_test_close(flowsom_test_some(model.node_sds[0][0]), 2.0, 1.0e-12), + ) + assert_true( + flowsom_test_close( + flowsom_test_some(model.node_cvs[0][0]), + 2.0 / 3.0, + 1.0e-12, + ), + ) + assert_true( + flowsom_test_close(flowsom_test_some(model.node_cvs[0][1]), 0.5, 1.0e-12), + ) + assert_true( + flowsom_test_close( + flowsom_test_some(model.node_mads[0][0]), + 2.9652, + 1.0e-12, + ), + ) +} + +///| +test "FlowSOM: empty nodes retain missing optional statistics" { + let config = @src.FlowSomConfig::create( + xdim=3, + ydim=1, + rlen=1, + initialization=@src.flowsom_pca_initialization(), + seed=53, + ) catch { + _ => abort("PCA FlowSOM configuration should build") + } + let model = @src.flowsom_train([[0.0], [10.0]], config) catch { + _ => abort("PCA FlowSOM should train") + } + let mut empty = -1 + for node in 0..= 0) + assert_true(model.node_medians[empty][0] is None) + assert_true(model.node_sds[empty][0] is None) + assert_true(model.node_mads[empty][0] is None) +} + +///| +test "FlowSOM: outlier report is aligned to nodes and cells" { + let model = flowsom_test_model() + assert_eq(model.outliers.median_distances.length(), 4) + assert_eq(model.outliers.mad_distances.length(), 4) + assert_eq(model.outliers.thresholds.length(), 4) + assert_eq(model.outliers.counts.length(), 4) + assert_eq(model.outliers.maximum_distances.length(), 4) + assert_eq(model.outliers.per_cell.length(), 8) + let mut flagged = 0 + for value in model.outliers.per_cell { + if value { + flagged = flagged + 1 + } + } + assert_eq(flowsom_test_sum(model.outliers.counts), flagged) +} + +///| +test "FlowSOM: quantization and topographic errors are bounded" { + let model = flowsom_test_model() + assert_true(model.quantization_error >= 0.0) + assert_true(model.topographic_error >= 0.0) + assert_true(model.topographic_error <= 1.0) +} + +///| +test "FlowSOM: node positive percentages use strict cutoffs" { + let model = flowsom_test_one_node_model() + let percentages = @src.flowsom_node_positive_percentages(model, [2.0, 4.0]) catch { + _ => abort("node positive percentages should work") + } + assert_true( + flowsom_test_close(flowsom_test_some(percentages[0][0]), 2.0 / 3.0, 1.0e-12), + ) + assert_true( + flowsom_test_close(flowsom_test_some(percentages[0][1]), 1.0 / 3.0, 1.0e-12), + ) +} + +///| +test "FlowSOM: node positive percentages validate cutoffs" { + let model = flowsom_test_one_node_model() + let mut failures = 0 + ignore(@src.flowsom_node_positive_percentages(model, [1.0])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_node_positive_percentages(model, [1.0, 1.0e301])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: one meta-cluster assigns every code to zero" { + let labels = @src.flowsom_meta_cluster_codes([[0.0], [1.0], [10.0]], 1) catch { + _ => abort("one FlowSOM meta-cluster should work") + } + assert_eq(labels, [0, 0, 0]) +} + +///| +test "FlowSOM: k-means meta-clustering separates distant code groups" { + let labels = @src.flowsom_meta_cluster_codes( + [[0.0], [0.1], [0.2], [10.0], [10.1], [10.2]], + 2, + starts=4, + seed=59, + ) catch { + _ => abort("FlowSOM meta-clustering should work") + } + assert_eq(flowsom_test_unique_count(labels), 2) + assert_true(labels[0] == labels[1] && labels[1] == labels[2]) + assert_true(labels[3] == labels[4] && labels[4] == labels[5]) + assert_true(labels[0] != labels[3]) +} + +///| +test "FlowSOM: meta-clustering validates cluster controls" { + let codes = [[0.0], [1.0]] + let mut failures = 0 + ignore(@src.flowsom_meta_cluster_codes(codes, 0)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_meta_cluster_codes(codes, 3)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_meta_cluster_codes(codes, 2, starts=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_meta_cluster_codes(codes, 2, seed=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "FlowSOM: automatic meta-clustering returns small maxima directly" { + assert_eq(@src.flowsom_determine_meta_clusters([[0.0]], 1), 1) + assert_eq(@src.flowsom_determine_meta_clusters([[0.0], [1.0]], 2), 2) + let mut failures = 0 + ignore(@src.flowsom_determine_meta_clusters([[0.0], [1.0]], 2, starts=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_determine_meta_clusters([[0.0], [1.0]], 2, seed=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: automatic meta-clustering selects an internal elbow" { + let selected = @src.flowsom_determine_meta_clusters( + [[0.0], [0.1], [0.2], [10.0], [10.1], [10.2]], + 5, + starts=3, + seed=61, + ) catch { + _ => abort("automatic FlowSOM meta-clustering should work") + } + assert_true(selected >= 2) + assert_true(selected < 5) +} + +///| +test "FlowSOM: automatic meta-clustering validates its maximum" { + let mut failures = 0 + ignore(@src.flowsom_determine_meta_clusters([[0.0]], 0)) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_determine_meta_clusters([[0.0]], 2)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: configured meta-clusters propagate to cells" { + let model = flowsom_test_model(metaclusters=2) + assert_eq(model.meta_clusters.length(), 4) + assert_eq(model.cell_meta_clusters.length(), 8) + assert_eq(flowsom_test_unique_count(model.meta_clusters), 2) + assert_eq( + @src.flowsom_meta_medians(model).length(), + @src.flowsom_meta_counts(model).length(), + ) + for cell in 0.. failures = failures + 1 + } + ignore(@src.flowsom_meta_medians(model)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: post-hoc meta-clustering leaves the original model unchanged" { + let model = flowsom_test_model() + let updated = @src.flowsom_with_metaclusters(model, 2, starts=3, seed=67) catch { + _ => abort("post-hoc FlowSOM meta-clustering should work") + } + assert_eq(model.meta_clusters.length(), 0) + assert_eq(updated.meta_clusters.length(), 4) + assert_eq(updated.cell_meta_clusters.length(), 8) + assert_eq(updated.config.metaclusters, 2) + updated.codes[0][0] = 999.0 + assert_true(model.codes[0][0] != 999.0) +} + +///| +test "FlowSOM: automatic post-hoc meta-clustering updates the model" { + let model = flowsom_test_model() + let updated = @src.flowsom_with_automatic_metaclusters( + model, + 3, + starts=3, + seed=71, + ) catch { + _ => abort("automatic post-hoc meta-clustering should work") + } + assert_true(updated.config.metaclusters >= 2) + assert_true(updated.config.metaclusters <= 3) + assert_eq(updated.meta_clusters.length(), 4) + assert_eq(updated.cell_meta_clusters.length(), 8) +} + +///| +test "FlowSOM: new data mapping applies node and meta-cluster labels" { + let model = flowsom_test_model(metaclusters=2) + let mapped = @src.flowsom_map_new(model, [[0.05, 0.1], [10.2, 10.0]]) catch { + _ => abort("new FlowSOM data should map") + } + assert_eq(mapped.mapping.clusters.length(), 2) + assert_eq(mapped.mapping.distances.length(), 2) + assert_eq(mapped.meta_clusters.length(), 2) + assert_eq(mapped.outliers.length(), 2) + for cell in 0..<2 { + assert_eq( + mapped.meta_clusters[cell], + model.meta_clusters[mapped.mapping.clusters[cell]], + ) + } +} + +///| +test "FlowSOM: new data mapping validates marker dimensions" { + let model = flowsom_test_model() + let mut raised = false + ignore(@src.flowsom_map_new(model, [[1.0]])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "FlowSOM: marker outliers classify low and high observations" { + let model = flowsom_test_one_node_model() + let outliers = @src.flowsom_marker_outliers( + model, + [[0.0, 4.0], [3.0, 4.0], [7.0, 8.0]], + mad_allowed=0.0, + ) catch { + _ => abort("marker-level FlowSOM outliers should work") + } + assert_eq(outliers, [[-1, 0], [0, 0], [1, 1]]) +} + +///| +test "FlowSOM: marker outliers validate dimensions and MAD control" { + let model = flowsom_test_one_node_model() + let mut failures = 0 + ignore(@src.flowsom_marker_outliers(model, [[1.0]])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_marker_outliers(model, [[1.0, 2.0]], mad_allowed=-1.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: weighted and unweighted purity follow contingency counts" { + let real = [0, 0, 1, 1] + let predicted = [0, 0, 0, 1] + let weighted = @src.flowsom_purity(real, predicted) catch { + _ => abort("weighted FlowSOM purity should work") + } + let unweighted = @src.flowsom_purity(real, predicted, weighted=false) catch { + _ => abort("unweighted FlowSOM purity should work") + } + assert_true(flowsom_test_close(weighted.mean, 0.75, 1.0e-12)) + assert_true(flowsom_test_close(weighted.worst, 2.0 / 3.0, 1.0e-12)) + assert_eq(weighted.below_075, 1) + assert_true(flowsom_test_close(unweighted.mean, 5.0 / 6.0, 1.0e-12)) +} + +///| +test "FlowSOM: F-measure weights best matches by real-cluster size" { + let score = @src.flowsom_f_measure([0, 0, 1, 1], [0, 0, 0, 1]) catch { + _ => abort("FlowSOM F-measure should work") + } + assert_true(flowsom_test_close(score, 11.0 / 15.0, 1.0e-12)) +} + +///| +test "FlowSOM: purity and F-measure validate label vectors" { + let mut failures = 0 + ignore(@src.flowsom_purity([], [])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_purity([0], [0, 1])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_f_measure([], [])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_f_measure([0], [0, 1])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "FlowSOM: summary reports model dimensions and quality" { + let summary = @src.flowsom_summary(flowsom_test_model(metaclusters=2)) + assert_true(summary.contains("FlowSOM model")) + assert_true(summary.contains("cells: 8")) + assert_true(summary.contains("markers: 2")) + assert_true(summary.contains("nodes: 4")) + assert_true(summary.contains("metaclusters: 2")) + assert_true(summary.contains("quantization error:")) + assert_true(summary.contains("topographic error:")) +} + +///| +test "FlowSOM: FlowFrame adapter selects markers in requested order" { + let frame = @src.FlowFrame::new( + [[0.0, 1.0, 2.0], [0.1, 1.1, 2.1], [10.0, 11.0, 12.0], [10.1, 11.1, 12.1]], + [ + @src.ParameterDescription::new("CD3", 20.0, 0.0, 20.0), + @src.ParameterDescription::new("CD19", 20.0, 0.0, 20.0), + @src.ParameterDescription::new("CD45", 20.0, 0.0, 20.0), + ], + ) + let config = @src.FlowSomConfig::create(xdim=2, ydim=1, rlen=2, seed=73) catch { + _ => abort("FlowFrame FlowSOM configuration should build") + } + let model = @src.flowsom_train_flow_frame(frame, config, marker_indices=[2, 0]) catch { + _ => abort("FlowFrame FlowSOM training should work") + } + assert_eq(model.marker_names, ["CD45", "CD3"]) + assert_eq(model.data[0], [2.0, 0.0]) + assert_eq(model.data[3], [12.1, 10.1]) +} + +///| +test "FlowSOM: FlowFrame adapter validates marker selection" { + let frame = @src.FlowFrame::new([[0.0, 1.0], [2.0, 3.0]], [ + @src.ParameterDescription::new("A", 10.0, 0.0, 10.0), + @src.ParameterDescription::new("B", 10.0, 0.0, 10.0), + ]) + let config = @src.FlowSomConfig::create(xdim=1, ydim=1) catch { + _ => abort("FlowFrame FlowSOM configuration should build") + } + let mut failures = 0 + ignore(@src.flowsom_train_flow_frame(frame, config, marker_indices=[0, 0])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_train_flow_frame(frame, config, marker_indices=[2])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "FlowSOM: SCE adapter writes one-based labels without mutating input" { + let experiment = flowsom_test_sce() + let config = @src.FlowSomConfig::create( + xdim=2, + ydim=1, + rlen=2, + metaclusters=2, + meta_starts=3, + seed=79, + ) catch { + _ => abort("SCE FlowSOM configuration should build") + } + let result = @src.flowsom_cluster_sce(experiment, config) catch { + _ => abort("SCE FlowSOM clustering should work") + } + assert_false(experiment.col_data.contains("FlowSOM.cluster")) + assert_false(experiment.col_data.contains("FlowSOM.metacluster")) + assert_true(result.experiment.col_data.contains("FlowSOM.cluster")) + assert_true(result.experiment.col_data.contains("FlowSOM.metacluster")) + assert_eq(result.experiment.col_data["FlowSOM.cluster"].length(), 6) + assert_eq(result.experiment.col_data["FlowSOM.metacluster"].length(), 6) + for label in result.experiment.col_data["FlowSOM.cluster"] { + assert_true(label == "1" || label == "2") + } + assert_eq(result.experiment.metadata["FlowSOM.assay"], "counts") + assert_eq(result.experiment.metadata["FlowSOM.grid"], "2x1") +} + +///| +test "FlowSOM: SCE adapter transposes assays and honors marker subsets" { + let experiment = flowsom_test_sce() + let config = @src.FlowSomConfig::create(xdim=2, ydim=1, rlen=2, seed=83) catch { + _ => abort("SCE FlowSOM configuration should build") + } + let result = @src.flowsom_cluster_sce( + experiment, + config, + marker_indices=[2, 0], + cluster_column="som", + meta_column="meta", + ) catch { + _ => abort("subset SCE FlowSOM clustering should work") + } + assert_eq(result.model.marker_names, ["CD45", "CD3"]) + assert_eq(result.model.data[0], [2.0, 0.0]) + assert_eq(result.model.data[5], [12.2, 10.2]) + assert_true(result.experiment.col_data.contains("som")) + assert_false(result.experiment.col_data.contains("meta")) + assert_eq(result.experiment.assays["counts"].length(), 3) + assert_eq(result.experiment.assays["counts"][0].length(), 6) +} + +///| +test "FlowSOM: SCE adapter recursively deep-copies nested containers" { + let experiment = flowsom_test_sce() + let config = @src.FlowSomConfig::create(xdim=2, ydim=1, rlen=2, seed=89) catch { + _ => abort("SCE FlowSOM configuration should build") + } + let result = @src.flowsom_cluster_sce(experiment, config) catch { + _ => abort("SCE FlowSOM clustering should work") + } + result.experiment.assays["counts"][0][0] = 999.0 + result.experiment.row_data["symbol"][0] = "changed" + result.experiment.col_data["batch"][0] = "changed" + result.experiment.reduced_dims["PCA"][0][0] = 999.0 + result.experiment.metadata["project"] = "changed" + result.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 999.0 + assert_eq(experiment.assays["counts"][0][0], 0.0) + assert_eq(experiment.row_data["symbol"][0], "T") + assert_eq(experiment.col_data["batch"][0], "A") + assert_eq(experiment.reduced_dims["PCA"][0][0], 0.0) + assert_eq(experiment.metadata["project"], "flowsom-test") + assert_eq( + experiment.alternative_experiments["spike"].assays["counts"][0][0], + 5.0, + ) +} + +///| +test "FlowSOM: SCE adapter validates assay, marker and output names" { + let experiment = flowsom_test_sce() + let config = @src.FlowSomConfig::create(xdim=2, ydim=1) catch { + _ => abort("SCE FlowSOM configuration should build") + } + let mut failures = 0 + ignore(@src.flowsom_cluster_sce(experiment, config, assay_name="missing")) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_cluster_sce(experiment, config, marker_indices=[0, 0])) catch { + _ => failures = failures + 1 + } + ignore(@src.flowsom_cluster_sce(experiment, config, cluster_column="")) catch { + _ => failures = failures + 1 + } + ignore( + @src.flowsom_cluster_sce( + experiment, + config, + cluster_column="same", + meta_column="same", + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} diff --git a/test/moonbit/fragment_mapper_test.mbt b/test/moonbit/fragment_mapper_test.mbt index 40863cac..83e05f59 100644 --- a/test/moonbit/fragment_mapper_test.mbt +++ b/test/moonbit/fragment_mapper_test.mbt @@ -145,12 +145,14 @@ test "fm_coverage_basic" { ///| test "fm_merge_fragments_adjacent_same_type" { - let frag1 = @src.Fragment::new("f1", 1, 5, - fragment_type=@src.ft_helix(), - residues=[@src.FragmentResidue::new("A", 1), @src.FragmentResidue::new("L", 2)]) - let frag2 = @src.Fragment::new("f2", 6, 10, - fragment_type=@src.ft_helix(), - residues=[@src.FragmentResidue::new("P", 6), @src.FragmentResidue::new("H", 7)]) + let frag1 = @src.Fragment::new("f1", 1, 5, fragment_type=@src.ft_helix(), residues=[ + @src.FragmentResidue::new("A", 1), + @src.FragmentResidue::new("L", 2), + ]) + let frag2 = @src.Fragment::new("f2", 6, 10, fragment_type=@src.ft_helix(), residues=[ + @src.FragmentResidue::new("P", 6), + @src.FragmentResidue::new("H", 7), + ]) let merged = @src.fm_merge_fragments([frag1, frag2]) assert_eq(merged.length(), 1) assert_eq(merged[0].start_residue, 1) @@ -367,4 +369,4 @@ test "fm_parse_fragments_sequence_mapping" { "END\n" let result = @src.fm_parse_fragments(content) assert_eq(result.protein_sequence, "AGF") -} \ No newline at end of file +} diff --git a/test/moonbit/freq_analysis_test.mbt b/test/moonbit/freq_analysis_test.mbt index cab60cdd..45824167 100644 --- a/test/moonbit/freq_analysis_test.mbt +++ b/test/moonbit/freq_analysis_test.mbt @@ -1,57 +1,72 @@ // Tests for Bio.FreqAnalysis module +///| test "fa_count_kmers single nucleotide" { let result = @src.fa_count_kmers("ACGT", 1) assert_true(result.fa_get_total_count() == 4) - assert_true(result.fa_get_frequency("A") > 0.24 && result.fa_get_frequency("A") < 0.26) + assert_true( + result.fa_get_frequency("A") > 0.24 && result.fa_get_frequency("A") < 0.26, + ) } +///| test "fa_count_kmers dinucleotide" { let result = @src.fa_count_kmers("ACGTAC", 2) assert_true(result.fa_get_total_count() == 5) } +///| test "fa_count_kmers empty" { let result = @src.fa_count_kmers("", 3) assert_true(result.fa_get_total_count() == 0) } +///| test "fa_nucleotide_frequency" { let result = @src.fa_nucleotide_frequency("AATTCCGG") assert_true(result.fa_get_total_count() == 8) - assert_true(result.fa_get_frequency("A") > 0.24 && result.fa_get_frequency("A") < 0.26) + assert_true( + result.fa_get_frequency("A") > 0.24 && result.fa_get_frequency("A") < 0.26, + ) } +///| test "fa_count_pattern" { let count = @src.fa_count_pattern("ATATAT", "AT") assert_true(count == 3) } +///| test "fa_count_pattern no match" { let count = @src.fa_count_pattern("ACGT", "TT") assert_true(count == 0) } +///| test "fa_codon_usage" { let result = @src.fa_codon_usage("ATGAAACCC") assert_true(result.fa_get_total_count() == 3) } +///| test "fa_gc_content" { let result = @src.fa_gc_content("AAGGCC") assert_true(result > 0.66 && result < 0.68) } +///| test "fa_gc_content all GC" { let result = @src.fa_gc_content("GCGC") assert_true(result == 1.0) } +///| test "fa_at_content" { let result = @src.fa_at_content("AATTCC") assert_true(result > 0.66 && result < 0.68) } +///| test "fa_find_motif" { let positions = @src.fa_find_motif("ATCGATCGATCG", "ATC") assert_true(positions.length() == 3) @@ -60,11 +75,13 @@ test "fa_find_motif" { assert_true(positions[2] == 8) } +///| test "fa_find_motif not found" { let positions = @src.fa_find_motif("ACGT", "TT") assert_true(positions.length() == 0) } +///| test "fa_get_patterns" { let result = @src.fa_count_kmers("AAACCC", 1) let patterns = result.fa_get_patterns() diff --git a/test/moonbit/freq_table_test.mbt b/test/moonbit/freq_table_test.mbt index cd24d654..fa8f8ec9 100644 --- a/test/moonbit/freq_table_test.mbt +++ b/test/moonbit/freq_table_test.mbt @@ -305,7 +305,10 @@ test "freq_table_from_sequence_basic" { ///| test "freq_table_from_sequence_with_alphabet_restricts" { - let ft = @src.freq_table_from_sequence("AATTGGCCNN", Some(["A", "T", "G", "C"])) + let ft = @src.freq_table_from_sequence( + "AATTGGCCNN", + Some(["A", "T", "G", "C"]), + ) assert_eq(ft.size(), 4) assert_eq(ft.count_of("A"), 2) // N is not in alphabet, so it should not be counted. diff --git a/test/moonbit/fssp_test.mbt b/test/moonbit/fssp_test.mbt index c79df6bb..b36eb564 100644 --- a/test/moonbit/fssp_test.mbt +++ b/test/moonbit/fssp_test.mbt @@ -7,8 +7,7 @@ test "fssp_header_creation" { let h = @src.FsspHeader::new( - "1dfa_A", "30-jul-1998", "crystal structure", - "17beta-hsd", "human", "holm", + "1dfa_A", "30-jul-1998", "crystal structure", "17beta-hsd", "human", "holm", 25, 3, 2.0, ) assert_eq(h.pdbid(), "1dfa_A") @@ -22,10 +21,10 @@ test "fssp_header_creation" { assert_eq(h.threshold(), "2") } +///| test "fssp_alignment_creation" { let a = @src.FsspAlignment::new( - "1csa_A", "1csa_A", 1, "MNIFVHEKDLFRTIVS", - 12.1, 2.0, 16, 68.8, + "1csa_A", "1csa_A", 1, "MNIFVHEKDLFRTIVS", 12.1, 2.0, 16, 68.8, ) assert_eq(a.pdbid(), "1csa_A") assert_eq(a.alignment_id(), "1csa_A") @@ -37,10 +36,10 @@ test "fssp_alignment_creation" { assert_eq(a.pid(), 68.8) } +///| test "fssp_data_creation" { let h = @src.FsspHeader::new( - "1dfa_A", "30-jul-1998", "", "", "", "", - 25, 2, 2.0, + "1dfa_A", "30-jul-1998", "", "", "", "", 25, 2, 2.0, ) let a1 = @src.FsspAlignment::new( "1dfa_A", "1dfa_A", 1, "ACGT", 15.0, 1.5, 4, 100.0, @@ -53,10 +52,9 @@ test "fssp_data_creation" { assert_true(data.reference().unwrap().pdbid() == "1dfa_A") } +///| test "fssp_filter_by_zscore" { - let h = @src.FsspHeader::new( - "1dfa_A", "", "", "", "", "", 10, 3, 2.0, - ) + let h = @src.FsspHeader::new("1dfa_A", "", "", "", "", "", 10, 3, 2.0) let a1 = @src.FsspAlignment::new("a1", "a1", 1, "AAA", 15.0, 1.0, 3, 100.0) let a2 = @src.FsspAlignment::new("a2", "a2", 1, "AAA", 10.0, 2.0, 3, 80.0) let a3 = @src.FsspAlignment::new("a3", "a3", 1, "AAA", 5.0, 3.0, 3, 60.0) @@ -67,10 +65,9 @@ test "fssp_filter_by_zscore" { assert_eq(filtered[1].pdbid(), "a2") } +///| test "fssp_empty_alignments" { - let h = @src.FsspHeader::new( - "1dfa_A", "", "", "", "", "", 0, 0, 0.0, - ) + let h = @src.FsspHeader::new("1dfa_A", "", "", "", "", "", 0, 0, 0.0) let data = @src.FsspData::new(h, []) assert_eq(data.n_alignments(), 0) assert_true(data.reference() is None) @@ -80,12 +77,14 @@ test "fssp_empty_alignments" { // Sample data // --------------------------------------------------------------------------- +///| test "fssp_sample_text_has_header" { let text = @src.fssp_sample_text() assert_true(text.contains("HEADER")) assert_true(text.contains("1dfa_A")) } +///| test "fssp_sample_text_has_alignments" { let text = @src.fssp_sample_text() assert_true(text.contains("## ALIGNMENTS")) @@ -93,6 +92,7 @@ test "fssp_sample_text_has_alignments" { assert_true(text.contains("1hsd_A")) } +///| test "fssp_sample_text_has_threshold" { let text = @src.fssp_sample_text() assert_true(text.contains("THRESHOLD")) @@ -103,6 +103,7 @@ test "fssp_sample_text_has_threshold" { // Parsing // --------------------------------------------------------------------------- +///| test "fssp_parse_header" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -113,6 +114,7 @@ test "fssp_parse_header" { assert_eq(h.threshold(), "2") } +///| test "fssp_parse_header_title" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -123,12 +125,14 @@ test "fssp_parse_header_title" { assert_eq(h.author(), "holm") } +///| test "fssp_parse_alignments_count" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) assert_eq(data.n_alignments(), 3) } +///| test "fssp_parse_alignment_details" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -144,6 +148,7 @@ test "fssp_parse_alignment_details" { assert_eq(ref_aln.pid(), 75.0) } +///| test "fssp_parse_second_alignment" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -153,6 +158,7 @@ test "fssp_parse_second_alignment" { assert_eq(a2.pid(), 68.8) } +///| test "fssp_parse_reference" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -160,11 +166,13 @@ test "fssp_parse_reference" { assert_eq(ref.pdbid(), "1dfa_A") } +///| test "fssp_parse_empty_input" { let data = @src.fssp_parse("") assert_eq(data.n_alignments(), 0) } +///| test "fssp_parse_header_only" { let text = "HEADER \\_1abc_A 1 01-jan-2000\nSEQLENGTH 10\n" let data = @src.fssp_parse(text) @@ -174,6 +182,7 @@ test "fssp_parse_header_only" { assert_eq(data.n_alignments(), 0) } +///| test "fssp_parse_filter_after_parse" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) @@ -187,6 +196,7 @@ test "fssp_parse_filter_after_parse" { // Formatting // --------------------------------------------------------------------------- +///| test "fssp_header_to_string" { let h = @src.FsspHeader::new( "1dfa_A", "30-jul-1998", "", "", "", "", 25, 3, 2.0, @@ -196,6 +206,7 @@ test "fssp_header_to_string" { assert_true(s.contains("25")) } +///| test "fssp_alignment_to_string" { let a = @src.FsspAlignment::new( "1csa_A", "1csa_A", 1, "ACGT", 12.1, 2.0, 4, 75.0, @@ -206,6 +217,7 @@ test "fssp_alignment_to_string" { assert_true(s.contains("12.1")) } +///| test "fssp_data_to_string" { let text = @src.fssp_sample_text() let data = @src.fssp_parse(text) diff --git a/test/moonbit/ga_test.mbt b/test/moonbit/ga_test.mbt index d1af22e9..32b86b63 100644 --- a/test/moonbit/ga_test.mbt +++ b/test/moonbit/ga_test.mbt @@ -17,9 +17,7 @@ test "ga_individual_creation" { ///| test "ga_individual_default" { - let ind = @src.GAIndividual::new( - sequence="ACGT", - ) + let ind = @src.GAIndividual::new(sequence="ACGT") assert_eq(ind.sequence, "ACGT") assert_true((ind.fitness - 0.0).abs() < 1.0e-10) assert_eq(ind.generation, 0) @@ -32,7 +30,7 @@ test "ga_population_creation" { @src.GAIndividual::new(sequence="CCCC", fitness=0.5, generation=0, id=1), @src.GAIndividual::new(sequence="GGGG", fitness=0.9, generation=0, id=2), ] - let pop = @src.GAPopulation::new(individuals=individuals) + let pop = @src.GAPopulation::new(individuals~) assert_eq(pop.individuals.length(), 3) assert_true((pop.best_fitness - 0.9).abs() < 1.0e-10) assert_true((pop.avg_fitness - 0.7333333333).abs() < 1.0e-10) @@ -110,7 +108,11 @@ test "ga_fitness_gc_content_zero" { test "ga_single_point_crossover" { let parent1 = @src.GAIndividual::new(sequence="AAAA") let parent2 = @src.GAIndividual::new(sequence="CCCC") - let (child1, child2) = @src.ga_single_point_crossover(parent1, parent2, point=2) + let (child1, child2) = @src.ga_single_point_crossover( + parent1, + parent2, + point=2, + ) assert_eq(child1.sequence.length(), 4) assert_eq(child2.sequence.length(), 4) } @@ -179,14 +181,24 @@ test "ga_evolution_loop_improves_fitness" { let fitness_fn = @src.ga_fitness_match(target) let result = @src.ga_evolve(config, fitness_fn) let initial_best = result.best_fitness_history[0] - let final_best = result.best_fitness_history[result.best_fitness_history.length() - 1] + let final_best = result.best_fitness_history[result.best_fitness_history.length() - + 1] assert_true(final_best >= initial_best) } ///| test "ga_result_creation" { - let best = @src.GAIndividual::new(sequence="ACGT", fitness=1.0, generation=10, id=5) - let stats = @src.GAGenerationStats::new(generation=10, best_fitness=1.0, avg_fitness=0.8) + let best = @src.GAIndividual::new( + sequence="ACGT", + fitness=1.0, + generation=10, + id=5, + ) + let stats = @src.GAGenerationStats::new( + generation=10, + best_fitness=1.0, + avg_fitness=0.8, + ) assert_eq(stats.generation, 10) assert_true((stats.best_fitness - 1.0).abs() < 1.0e-10) } @@ -260,4 +272,4 @@ test "ga_termination_threshold" { let fitness_fn = @src.ga_fitness_match(target) let result = @src.ga_evolve(config, fitness_fn) assert_true(result.generation_reached <= 200) -} \ No newline at end of file +} diff --git a/test/moonbit/gck_io_test.mbt b/test/moonbit/gck_io_test.mbt index e0d969b4..6bbf7007 100644 --- a/test/moonbit/gck_io_test.mbt +++ b/test/moonbit/gck_io_test.mbt @@ -33,8 +33,12 @@ test "gck_seq_type_from_int_round_trip" { let dna = @src.GckSeqType::from_int(0) let rna = @src.GckSeqType::from_int(1) let protein = @src.GckSeqType::from_int(2) - assert_true(@src.GckSeqType::from_int(dna.to_int()) is @src.GckSeqType::GckDna) - assert_true(@src.GckSeqType::from_int(rna.to_int()) is @src.GckSeqType::GckRna) + assert_true( + @src.GckSeqType::from_int(dna.to_int()) is @src.GckSeqType::GckDna, + ) + assert_true( + @src.GckSeqType::from_int(rna.to_int()) is @src.GckSeqType::GckRna, + ) assert_true( @src.GckSeqType::from_int(protein.to_int()) is @src.GckSeqType::GckProtein, ) @@ -74,12 +78,9 @@ test "gck_file_new_defaults" { ///| test "gck_feature_new" { - let f = @src.GckFeature::new( - name="AmpR", - type_="CDS", - direction=1, - segments=[(0, 19)], - ) + let f = @src.GckFeature::new(name="AmpR", type_="CDS", direction=1, segments=[ + (0, 19), + ]) assert_eq(f.name, "AmpR") assert_eq(f.type_, "CDS") assert_eq(f.direction, 1) @@ -91,12 +92,11 @@ test "gck_feature_new" { ///| test "gck_feature_new_multiple_segments" { - let f = @src.GckFeature::new( - name="gene1", - type_="CDS", - direction=2, - segments=[(0, 10), (20, 30), (40, 50)], - ) + let f = @src.GckFeature::new(name="gene1", type_="CDS", direction=2, segments=[ + (0, 10), + (20, 30), + (40, 50), + ]) assert_eq(f.name, "gene1") assert_eq(f.direction, 2) assert_eq(f.segments.length(), 3) @@ -140,12 +140,9 @@ test "gck_count_features_by_type_multiple" { @src.GckFeature::new(name="b", type_="CDS", direction=1, segments=[(20, 30)]), ) file.add_feature( - @src.GckFeature::new( - name="c", - type_="promoter", - direction=1, - segments=[(40, 50)], - ), + @src.GckFeature::new(name="c", type_="promoter", direction=1, segments=[ + (40, 50), + ]), ) assert_eq(@src.gck_count_features_by_type(file, "CDS"), 2) assert_eq(@src.gck_count_features_by_type(file, "promoter"), 1) @@ -320,7 +317,9 @@ test "gck_custom_file_round_trip" { file.set_sequence("ATGAAATAG") file.set_circular(false) file.add_feature( - @src.GckFeature::new(name="gene1", type_="CDS", direction=1, segments=[(0, 8)]), + @src.GckFeature::new(name="gene1", type_="CDS", direction=1, segments=[ + (0, 8), + ]), ) let hex = @src.gck_write(file) let parsed = @src.gck_parse(hex) @@ -339,20 +338,14 @@ test "gck_multiple_features_round_trip" { @src.GckFeature::new(name="f1", type_="CDS", direction=1, segments=[(0, 3)]), ) file.add_feature( - @src.GckFeature::new( - name="f2", - type_="promoter", - direction=2, - segments=[(4, 7)], - ), + @src.GckFeature::new(name="f2", type_="promoter", direction=2, segments=[ + (4, 7), + ]), ) file.add_feature( - @src.GckFeature::new( - name="f3", - type_="terminator", - direction=0, - segments=[(8, 11)], - ), + @src.GckFeature::new(name="f3", type_="terminator", direction=0, segments=[ + (8, 11), + ]), ) let hex = @src.gck_write(file) let parsed = @src.gck_parse(hex) @@ -370,12 +363,11 @@ test "gck_multi_segment_feature_round_trip" { let file = @src.GckFile::new() file.set_sequence("ATGCATGCATGC") file.add_feature( - @src.GckFeature::new( - name="gene", - type_="CDS", - direction=1, - segments=[(0, 2), (4, 6), (8, 10)], - ), + @src.GckFeature::new(name="gene", type_="CDS", direction=1, segments=[ + (0, 2), + (4, 6), + (8, 10), + ]), ) let hex = @src.gck_write(file) let parsed = @src.gck_parse(hex) @@ -398,24 +390,15 @@ test "gck_multi_segment_feature_round_trip" { ///| test "gck_feature_direction_values" { - let f1 = @src.GckFeature::new( - name="f1", - type_="CDS", - direction=0, - segments=[(0, 10)], - ) - let f2 = @src.GckFeature::new( - name="f2", - type_="CDS", - direction=1, - segments=[(0, 10)], - ) - let f3 = @src.GckFeature::new( - name="f3", - type_="CDS", - direction=2, - segments=[(0, 10)], - ) + let f1 = @src.GckFeature::new(name="f1", type_="CDS", direction=0, segments=[ + (0, 10), + ]) + let f2 = @src.GckFeature::new(name="f2", type_="CDS", direction=1, segments=[ + (0, 10), + ]) + let f3 = @src.GckFeature::new(name="f3", type_="CDS", direction=2, segments=[ + (0, 10), + ]) assert_eq(f1.direction, 0) assert_eq(f2.direction, 1) assert_eq(f3.direction, 2) diff --git a/test/moonbit/gcrma_test.mbt b/test/moonbit/gcrma_test.mbt index c4bebdac..6e70abd4 100644 --- a/test/moonbit/gcrma_test.mbt +++ b/test/moonbit/gcrma_test.mbt @@ -18,9 +18,7 @@ test "gcrma_probe_info_new" { ///| test "gcrma_probe_info_with_values" { - let probe = @src.ProbeInfo::with_values( - "probe_test", 8, "ACGTACGT", 2.5, - ) + let probe = @src.ProbeInfo::with_values("probe_test", 8, "ACGTACGT", 2.5) assert_eq(probe.probe_id, "probe_test") assert_eq(probe.gc_count, 8) assert_eq(probe.affinity, 2.5) @@ -113,9 +111,7 @@ test "gcrma_compute_gc_lookup_table_basic" { test "gcrma_background_correction_express" { let pm = [100.0, 200.0, 150.0, 300.0] let mm = [50.0, 80.0, 70.0, 120.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.5, "Express", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.5, "Express") assert_eq(result.length(), 4) assert_eq(result[0], 50.0) assert_eq(result[1], 120.0) @@ -125,9 +121,7 @@ test "gcrma_background_correction_express" { test "gcrma_background_correction_express_negative" { let pm = [30.0, 50.0] let mm = [80.0, 100.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.5, "Express", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.5, "Express") assert_eq(result.length(), 2) assert_eq(result[0], 1.0) } @@ -136,18 +130,14 @@ test "gcrma_background_correction_express_negative" { test "gcrma_background_correction_idealm" { let pm = [100.0, 200.0, 150.0, 300.0] let mm = [50.0, 80.0, 70.0, 120.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.5, "IdealMM", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.5, "IdealMM") assert_eq(result.length(), 4) assert_eq(result[0] > 1.0, true) } ///| test "gcrma_background_correction_empty" { - let result = @src.gcrma_background_correction( - [], [], 0.5, "IdealMM", - ) + let result = @src.gcrma_background_correction([], [], 0.5, "IdealMM") assert_eq(result.length(), 0) } @@ -180,11 +170,7 @@ test "gcrma_normalize_single_row" { ///| test "gcrma_normalize_multiple_rows" { - let data = [ - [100.0, 200.0], - [150.0, 250.0], - [120.0, 180.0], - ] + let data = [[100.0, 200.0], [150.0, 250.0], [120.0, 180.0]] let result = @src.gcrma_normalize(data) assert_eq(result.length(), 3) assert_eq(result[0].length(), 2) @@ -269,10 +255,7 @@ test "gcrma_process_empty" { ///| test "gcrma_process_with_config" { - let cel_data = [ - [100.0, 200.0], - [150.0, 250.0], - ] + let cel_data = [[100.0, 200.0], [150.0, 250.0]] let probe_info = [ @src.ProbeInfo::new("p1", "ACGTACGTACGT"), @src.ProbeInfo::new("p2", "GCGCGCGCGCGC"), @@ -285,10 +268,7 @@ test "gcrma_process_with_config" { ///| test "gcrma_process_no_normalize" { - let cel_data = [ - [100.0, 200.0], - [150.0, 250.0], - ] + let cel_data = [[100.0, 200.0], [150.0, 250.0]] let probe_info = [ @src.ProbeInfo::new("p1", "ACGTACGTACGT"), @src.ProbeInfo::new("p2", "GCGCGCGCGCGC"), @@ -302,9 +282,7 @@ test "gcrma_process_no_normalize" { test "gcrma_background_correction_mm_longer_than_pm" { let pm = [100.0, 200.0] let mm = [50.0, 80.0, 90.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.5, "Express", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.5, "Express") assert_eq(result.length(), 2) assert_eq(result[0], 50.0) assert_eq(result[1], 120.0) @@ -314,9 +292,7 @@ test "gcrma_background_correction_mm_longer_than_pm" { test "gcrma_background_correction_mm_shorter_than_pm" { let pm = [100.0, 200.0, 150.0] let mm = [50.0, 80.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.5, "Express", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.5, "Express") assert_eq(result.length(), 3) assert_eq(result[0], 50.0) assert_eq(result[2], 150.0) @@ -326,9 +302,7 @@ test "gcrma_background_correction_mm_shorter_than_pm" { test "gcrma_background_correction_idealm_clamps_to_minimum" { let pm = [10.0, 20.0] let mm = [8.0, 15.0] - let result = @src.gcrma_background_correction( - pm, mm, 0.9, "IdealMM", - ) + let result = @src.gcrma_background_correction(pm, mm, 0.9, "IdealMM") assert_eq(result.length(), 2) assert_eq(result[0] >= 1.0, true) assert_eq(result[1] >= 1.0, true) @@ -386,7 +360,10 @@ test "gcrma_full_pipeline_express" { @src.ProbeInfo::new("p5", "TGTGTGTGTGTG"), @src.ProbeInfo::new("p6", "CGCGCGCGCGCG"), ] - let config = @src.GCRMAConfig::new(background_method="Express", normalize=true) + let config = @src.GCRMAConfig::new( + background_method="Express", + normalize=true, + ) let result = @src.gcrma_process_with_config(cel_data, probe_info, config) assert_eq(result.expression_matrix.length(), 6) assert_eq(result.expression_matrix[0].length(), 2) @@ -412,12 +389,8 @@ test "gcrma_estimate_affinity_dinucleotide" { test "gcrma_background_correction_high_gc" { let pm = [100.0, 200.0, 150.0] let mm = [60.0, 100.0, 80.0] - let result_low = @src.gcrma_background_correction( - pm, mm, 0.2, "IdealMM", - ) - let result_high = @src.gcrma_background_correction( - pm, mm, 0.8, "IdealMM", - ) + let result_low = @src.gcrma_background_correction(pm, mm, 0.2, "IdealMM") + let result_high = @src.gcrma_background_correction(pm, mm, 0.8, "IdealMM") assert_eq(result_low.length(), 3) assert_eq(result_high.length(), 3) } @@ -435,4 +408,4 @@ test "gcrma_gc_correction_preserves_length" { ] let result = @src.gcrma_gc_correction(pm_values, gc_counts, probes) assert_eq(result.length(), 5) -} \ No newline at end of file +} diff --git a/test/moonbit/genefilter_test.mbt b/test/moonbit/genefilter_test.mbt index a0cfc966..9720c5ab 100644 --- a/test/moonbit/genefilter_test.mbt +++ b/test/moonbit/genefilter_test.mbt @@ -6,11 +6,13 @@ test "genefilter_row_ttest" { [5.0, 6.0, 5.5, 7.0, 6.5, 7.5], [100.0, 98.0, 102.0, 50.0, 48.0, 52.0], ] - let groups = ["control", "control", "control", "treatment", "treatment", "treatment"] - + let groups = [ + "control", "control", "control", "treatment", "treatment", "treatment", + ] + let ge = @src.GeneExpression::new(gene_ids, expression, groups) let result = @src.row_ttest(ge, "control", "treatment") - + assert_eq(result.passing_genes.length(), 3) } @@ -22,56 +24,45 @@ test "genefilter_row_wilcoxon" { [10.0, 11.0, 12.0, 1.0, 2.0, 3.0], ] let groups = ["groupA", "groupA", "groupA", "groupB", "groupB", "groupB"] - + let ge = @src.GeneExpression::new(gene_ids, expression, groups) let result = @src.row_wilcoxon(ge, "groupA", "groupB") - + assert_eq(result.passing_genes.length(), 2) } ///| test "genefilter_variance_filter" { let gene_ids = ["gene1", "gene2", "gene3"] - let expression = [ - [1.0, 1.0, 1.0], - [1.0, 2.0, 3.0], - [10.0, 20.0, 30.0], - ] + let expression = [[1.0, 1.0, 1.0], [1.0, 2.0, 3.0], [10.0, 20.0, 30.0]] let groups = ["g1", "g1", "g1"] - + let ge = @src.GeneExpression::new(gene_ids, expression, groups) let passing = @src.variance_filter(ge, 1.0) - + assert_eq(passing.length() >= 2, true) } ///| test "genefilter_cv_filter" { let gene_ids = ["gene1", "gene2"] - let expression = [ - [1.0, 1.0, 1.0], - [1.0, 2.0, 4.0], - ] + let expression = [[1.0, 1.0, 1.0], [1.0, 2.0, 4.0]] let groups = ["g1", "g1", "g1"] - + let ge = @src.GeneExpression::new(gene_ids, expression, groups) let passing = @src.cv_filter(ge, 0.5) - + assert_eq(passing.length() >= 1, true) } ///| test "genefilter_row_quantile_filter" { let gene_ids = ["gene1", "gene2", "gene3"] - let expression = [ - [5.0, 6.0, 7.0], - [1.0, 2.0, 3.0], - [9.0, 10.0, 11.0], - ] + let expression = [[5.0, 6.0, 7.0], [1.0, 2.0, 3.0], [9.0, 10.0, 11.0]] let groups = ["g1", "g1", "g1"] - + let ge = @src.GeneExpression::new(gene_ids, expression, groups) let passing = @src.row_quantile_filter(ge, 0.25, 0.75) - + assert_eq(passing.length() >= 1, true) -} \ No newline at end of file +} diff --git a/test/moonbit/genesis_test.mbt b/test/moonbit/genesis_test.mbt index c69d265d..d135caa5 100644 --- a/test/moonbit/genesis_test.mbt +++ b/test/moonbit/genesis_test.mbt @@ -9,9 +9,9 @@ test "genesis_estimate_kinship" { [1.0, 2.0, 0.0], [2.0, 0.0, 1.0], ] - + let result = @src.bio_estimate_kinship(genotypes) - + assert_true(result.n_samples == 4) assert_true(result.kinship_matrix.length() == 4) assert_true(result.sample_ids.length() == 4) @@ -25,9 +25,9 @@ test "genesis_pca" { [1.0, 2.0, 0.0], [2.0, 0.0, 1.0], ] - + let result = @src.bio_pca(genotypes, 2) - + assert_true(result.eigenvalues.length() == 2) assert_true(result.eigenvectors.length() > 0) assert_true(result.var_explained.length() == 2) @@ -35,14 +35,10 @@ test "genesis_pca" { ///| test "genesis_genetic_distance" { - let genotypes = [ - [0.0, 1.0, 2.0], - [0.0, 1.0, 2.0], - [1.0, 2.0, 0.0], - ] - + let genotypes = [[0.0, 1.0, 2.0], [0.0, 1.0, 2.0], [1.0, 2.0, 0.0]] + let distance = @src.bio_genetic_distance(genotypes) - + assert_true(distance.distance_matrix.length() == 3) assert_true(distance.sample_ids.length() == 3) } @@ -51,9 +47,9 @@ test "genesis_genetic_distance" { test "genesis_euclidean_distance" { let v1 = [0.0, 1.0, 2.0] let v2 = [0.0, 1.0, 2.0] - + let dist = @src.gs_euclidean_distance(v1, v2) - + assert_true(dist == 0.0) } @@ -61,9 +57,9 @@ test "genesis_euclidean_distance" { test "genesis_manhattan_distance" { let v1 = [0.0, 1.0, 2.0] let v2 = [0.0, 1.0, 2.0] - + let dist = @src.gs_manhattan_distance(v1, v2) - + assert_true(dist == 0.0) } @@ -71,8 +67,8 @@ test "genesis_manhattan_distance" { test "genesis_ibs_distance" { let v1 = [0.0, 1.0, 2.0] let v2 = [0.0, 1.0, 2.0] - + let dist = @src.gs_ibs_distance(v1, v2) - + assert_true(dist >= 0.0) } diff --git a/test/moonbit/genie3_test.mbt b/test/moonbit/genie3_test.mbt index bf49e02a..c04cda48 100644 --- a/test/moonbit/genie3_test.mbt +++ b/test/moonbit/genie3_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Bioconductor GENIE3 module - Gene regulatory network inference. - test "genie3_tree_leaf_basic" { let leaf = @src.genie3_tree_leaf(3.5) assert_true(leaf.is_leaf) @@ -10,6 +9,7 @@ test "genie3_tree_leaf_basic" { assert_eq(leaf.right_idx, -1) } +///| test "genie3_tree_split_basic" { let left = @src.genie3_tree_leaf(1.0) let right = @src.genie3_tree_leaf(2.0) @@ -23,6 +23,7 @@ test "genie3_tree_split_basic" { assert_true((node.importance_gain - 10.0).abs() < 1.0e-9) } +///| test "genie3_build_tree_simple" { // Simple dataset: y = 2 * x + noise let features = [[1.0], [2.0], [3.0], [4.0], [5.0], [6.0]] @@ -36,6 +37,7 @@ test "genie3_build_tree_simple" { assert_true(p1 < p2) } +///| test "genie3_build_tree_constant_target" { // Constant target should produce a single leaf let features = [[1.0], [2.0], [3.0]] @@ -45,6 +47,7 @@ test "genie3_build_tree_constant_target" { assert_true(tree.nodes[0].is_leaf) } +///| test "genie3_feature_importance" { // x0 strongly predicts y, x1 doesn't let features = [ @@ -62,6 +65,7 @@ test "genie3_feature_importance" { assert_true(importance[0] > 0.0) } +///| test "genie3_run_basic" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -75,6 +79,7 @@ test "genie3_run_basic" { assert_true(result.edges.length() > 0) } +///| test "genie3_run_symmetrize" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5, symmetrize=true) @@ -91,6 +96,7 @@ test "genie3_run_symmetrize" { } } +///| test "genie3_column_normalization" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -109,6 +115,7 @@ test "genie3_column_normalization" { } } +///| test "genie3_edges_sorted_descending" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -119,6 +126,7 @@ test "genie3_edges_sorted_descending" { } } +///| test "genie3_top_edges" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -133,6 +141,7 @@ test "genie3_top_edges" { } } +///| test "genie3_top_edges_more_than_available" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -140,6 +149,7 @@ test "genie3_top_edges_more_than_available" { assert_eq(top100.length(), result.edges.length()) } +///| test "genie3_regulators_of" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -158,6 +168,7 @@ test "genie3_regulators_of" { assert_true(found_g1) } +///| test "genie3_targets_of" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -171,6 +182,7 @@ test "genie3_targets_of" { } } +///| test "genie3_regulators_of_unknown_gene" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -178,6 +190,7 @@ test "genie3_regulators_of_unknown_gene" { assert_eq(regs.length(), 0) } +///| test "genie3_sample_data_shape" { let (expr, names) = @src.genie3_sample_data() assert_eq(names.length(), 5) @@ -188,6 +201,7 @@ test "genie3_sample_data_shape" { assert_true((expr[0][0] - expr2[0][0]).abs() < 1.0e-9) } +///| test "genie3_run_no_edges_when_uniform" { // All identical samples → no informative splits → no edges let expr = [[1.0, 2.0, 3.0], [1.0, 2.0, 3.0], [1.0, 2.0, 3.0]] @@ -196,6 +210,7 @@ test "genie3_run_no_edges_when_uniform" { assert_eq(result.edges.length(), 0) } +///| test "genie3_self_loops_excluded" { let (expr, names) = @src.genie3_sample_data() let result = @src.genie3_run(expr, names, max_depth=5) @@ -207,6 +222,7 @@ test "genie3_self_loops_excluded" { } } +///| test "genie3_predict_after_build" { // Build a tree and verify prediction lies within target range let features = [ diff --git a/test/moonbit/genome_diagram_test.mbt b/test/moonbit/genome_diagram_test.mbt index 3af6e3e3..d0ebbd24 100644 --- a/test/moonbit/genome_diagram_test.mbt +++ b/test/moonbit/genome_diagram_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for GenomeDiagram module. - test "gd_create_diagram - basic creation" { let d = @src.gd_create_diagram("test", 0, 1000) assert_eq(d.name(), "test") @@ -9,6 +8,7 @@ test "gd_create_diagram - basic creation" { assert_eq(@src.gd_diagram_length(d), 1000) } +///| test "gd_create_diagram - negative coordinates" { let d = @src.gd_create_diagram("neg", -100, 200) assert_eq(d.start(), -100) @@ -16,6 +16,7 @@ test "gd_create_diagram - negative coordinates" { assert_eq(@src.gd_diagram_length(d), 300) } +///| test "gd_add_track - single track" { let d = @src.gd_create_diagram("test", 0, 500) let d2 = @src.gd_add_track(d, "genes") @@ -24,6 +25,7 @@ test "gd_add_track - single track" { assert_eq(tracks[0].name(), "genes") } +///| test "gd_add_track - multiple tracks" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_track(d, "genes") @@ -32,9 +34,17 @@ test "gd_add_track - multiple tracks" { assert_eq(@src.gd_track_count(d4), 3) } +///| test "gd_add_feature - basic feature" { let d = @src.gd_create_diagram("test", 0, 1000) - let d2 = @src.gd_add_feature(d, 100, 300, "+", "GeneA", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 300, + "+", + "GeneA", + @src.gd_rectangle_shape(), + ) assert_eq(@src.gd_feature_count(d2), 1) let feats = d2.features() assert_eq(feats[0].start(), 100) @@ -43,14 +53,30 @@ test "gd_add_feature - basic feature" { assert_eq(feats[0].label(), "GeneA") } +///| test "gd_add_feature - multiple features" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_feature(d, 50, 150, "+", "Gene1", @src.gd_arrow_shape()) - let d3 = @src.gd_add_feature(d2, 200, 400, "-", "Gene2", @src.gd_diamond_shape()) - let d4 = @src.gd_add_feature(d3, 500, 800, "+", "Gene3", @src.gd_rectangle_shape()) + let d3 = @src.gd_add_feature( + d2, + 200, + 400, + "-", + "Gene2", + @src.gd_diamond_shape(), + ) + let d4 = @src.gd_add_feature( + d3, + 500, + 800, + "+", + "Gene3", + @src.gd_rectangle_shape(), + ) assert_eq(@src.gd_feature_count(d4), 3) } +///| test "gd_add_track_feature - basic track feature" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_track(d, "genes") @@ -62,6 +88,7 @@ test "gd_add_track_feature - basic track feature" { assert_eq(feats[0].label(), "GeneA") } +///| test "gd_add_track_feature - multiple features in same track" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_track(d, "genes") @@ -72,6 +99,7 @@ test "gd_add_track_feature - multiple features in same track" { assert_eq(feats.length(), 3) } +///| test "gd_add_track_feature - multiple tracks" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_track(d, "genes") @@ -86,36 +114,83 @@ test "gd_add_track_feature - multiple tracks" { assert_eq(exon_feats.length(), 1) } +///| test "gd_get_features - nonexistent track returns empty" { let d = @src.gd_create_diagram("test", 0, 1000) let feats = @src.gd_get_features(d, "nonexistent") assert_eq(feats.length(), 0) } +///| test "gd_find_overlapping_features - basic overlap" { let d = @src.gd_create_diagram("test", 0, 1000) - let d2 = @src.gd_add_feature(d, 100, 300, "+", "GeneA", @src.gd_rectangle_shape()) - let d3 = @src.gd_add_feature(d2, 400, 600, "+", "GeneB", @src.gd_rectangle_shape()) - let d4 = @src.gd_add_feature(d3, 250, 500, "+", "GeneC", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 300, + "+", + "GeneA", + @src.gd_rectangle_shape(), + ) + let d3 = @src.gd_add_feature( + d2, + 400, + 600, + "+", + "GeneB", + @src.gd_rectangle_shape(), + ) + let d4 = @src.gd_add_feature( + d3, + 250, + 500, + "+", + "GeneC", + @src.gd_rectangle_shape(), + ) let overlaps = @src.gd_find_overlapping_features(d4, 200, 400) assert_eq(overlaps.length(), 2) } +///| test "gd_find_overlapping_features - no overlap" { let d = @src.gd_create_diagram("test", 0, 1000) - let d2 = @src.gd_add_feature(d, 100, 200, "+", "GeneA", @src.gd_rectangle_shape()) - let d3 = @src.gd_add_feature(d2, 500, 600, "+", "GeneB", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 200, + "+", + "GeneA", + @src.gd_rectangle_shape(), + ) + let d3 = @src.gd_add_feature( + d2, + 500, + 600, + "+", + "GeneB", + @src.gd_rectangle_shape(), + ) let overlaps = @src.gd_find_overlapping_features(d3, 300, 400) assert_eq(overlaps.length(), 0) } +///| test "gd_find_overlapping_features - edge touching" { let d = @src.gd_create_diagram("test", 0, 1000) - let d2 = @src.gd_add_feature(d, 100, 200, "+", "GeneA", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 200, + "+", + "GeneA", + @src.gd_rectangle_shape(), + ) let overlaps = @src.gd_find_overlapping_features(d2, 200, 300) assert_eq(overlaps.length(), 0) } +///| test "gd_set_style - basic style" { let d = @src.gd_create_diagram("test", 0, 1000) let style = @src.DiagramStyle::new() @@ -125,6 +200,7 @@ test "gd_set_style - basic style" { assert_true(d2.style().border()) } +///| test "gd_set_style - circular mode" { let d = @src.gd_create_diagram("test", 0, 1000) let style = @src.DiagramStyle::new(circular=true, linear=false) @@ -133,6 +209,7 @@ test "gd_set_style - circular mode" { assert_false(d2.style().linear()) } +///| test "gd_diagram_length - different ranges" { let d1 = @src.gd_create_diagram("a", 0, 500) assert_eq(@src.gd_diagram_length(d1), 500) @@ -142,24 +219,41 @@ test "gd_diagram_length - different ranges" { assert_eq(@src.gd_diagram_length(d3), 100) } +///| test "gd_track_count - empty diagram" { let d = @src.gd_create_diagram("test", 0, 100) assert_eq(@src.gd_track_count(d), 0) } +///| test "gd_feature_count - mixed features" { let d = @src.gd_create_diagram("test", 0, 500) let d2 = @src.gd_add_track(d, "genes") - let d3 = @src.gd_add_feature(d2, 50, 150, "+", "G1", @src.gd_rectangle_shape()) + let d3 = @src.gd_add_feature( + d2, + 50, + 150, + "+", + "G1", + @src.gd_rectangle_shape(), + ) let d4 = @src.gd_add_track_feature(d3, "genes", 100, 200, "+", "G2") let d5 = @src.gd_add_feature(d4, 300, 400, "-", "G3", @src.gd_arrow_shape()) assert_eq(@src.gd_feature_count(d5), 3) } +///| test "gd_to_svg_string - generates valid SVG" { let d = @src.gd_create_diagram("test", 0, 500) let d2 = @src.gd_add_track(d, "genes") - let d3 = @src.gd_add_feature(d2, 50, 150, "+", "GeneA", @src.gd_rectangle_shape()) + let d3 = @src.gd_add_feature( + d2, + 50, + 150, + "+", + "GeneA", + @src.gd_rectangle_shape(), + ) let d4 = @src.gd_add_track_feature(d3, "genes", 200, 300, "+", "Exon1") let svg = @src.gd_to_svg_string(d4, 800, 400) assert_true(svg.length() > 0) @@ -167,19 +261,29 @@ test "gd_to_svg_string - generates valid SVG" { assert_true(svg.contains("")) } +///| test "gd_to_svg_string - empty diagram" { let d = @src.gd_create_diagram("empty", 0, 0) let svg = @src.gd_to_svg_string(d, 400, 200) assert_eq(svg, "") } +///| test "gd_to_svg_string - contains feature labels" { let d = @src.gd_create_diagram("test", 0, 500) - let d2 = @src.gd_add_feature(d, 100, 200, "+", "MyGene", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 200, + "+", + "MyGene", + @src.gd_rectangle_shape(), + ) let svg = @src.gd_to_svg_string(d2, 800, 400) assert_true(svg.contains("MyGene")) } +///| test "gd_set_feature_color - changes color" { let feat = @src.DiagramFeature::new(0, 100, color="#FF0000") let updated = @src.gd_set_feature_color(feat, "#00FF00") @@ -187,12 +291,19 @@ test "gd_set_feature_color - changes color" { assert_eq(feat.color(), "#FF0000") } +///| test "gd_set_feature_color - shape preserved" { - let feat = @src.DiagramFeature::new(0, 100, shape=@src.gd_arrow_shape(), color="#FF0000") + let feat = @src.DiagramFeature::new( + 0, + 100, + shape=@src.gd_arrow_shape(), + color="#FF0000", + ) let updated = @src.gd_set_feature_color(feat, "#00FF00") assert_eq(updated.shape().to_string(), "arrow") } +///| test "gd_label_features - auto labels large features" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_feature(d, 100, 300, "+", "", @src.gd_rectangle_shape()) @@ -203,14 +314,23 @@ test "gd_label_features - auto labels large features" { assert_eq(feats[1].label(), "") } +///| test "gd_label_features - existing labels preserved" { let d = @src.gd_create_diagram("test", 0, 1000) - let d2 = @src.gd_add_feature(d, 100, 300, "+", "Existing", @src.gd_rectangle_shape()) + let d2 = @src.gd_add_feature( + d, + 100, + 300, + "+", + "Existing", + @src.gd_rectangle_shape(), + ) let d3 = @src.gd_label_features(d2, 50) let feats = d3.features() assert_eq(feats[0].label(), "Existing") } +///| test "FeatureShape - to_string conversions" { let r = @src.gd_rectangle_shape() assert_eq(r.to_string(), "rectangle") @@ -224,6 +344,7 @@ test "FeatureShape - to_string conversions" { assert_eq(t.to_string(), "terminators") } +///| test "FeatureShape - from_string conversions" { let r = @src.FeatureShape::from_string("rectangle") assert_eq(r.to_string(), "rectangle") @@ -235,6 +356,7 @@ test "FeatureShape - from_string conversions" { assert_eq(u.to_string(), "rectangle") } +///| test "DiagramStyle - default values" { let style = @src.DiagramStyle::new() assert_eq(style.scale(), 1.0) @@ -244,6 +366,7 @@ test "DiagramStyle - default values" { assert_eq(style.color_scheme(), "default") } +///| test "DiagramStyle - setters" { let style = @src.DiagramStyle::new() let s2 = style.set_scale(2.5) @@ -254,6 +377,7 @@ test "DiagramStyle - setters" { assert_eq(s4.color_scheme(), "grayscale") } +///| test "DiagramStyle - circular and linear are mutually exclusive" { let style = @src.DiagramStyle::new() let circ = style.set_circular(true) @@ -264,6 +388,7 @@ test "DiagramStyle - circular and linear are mutually exclusive" { assert_true(lin.linear()) } +///| test "DiagramFeature - overlaps detection" { let feat = @src.DiagramFeature::new(100, 300) assert_true(feat.overlaps(50, 150)) @@ -273,16 +398,19 @@ test "DiagramFeature - overlaps detection" { assert_false(feat.overlaps(0, 100)) } +///| test "DiagramFeature - length" { let feat = @src.DiagramFeature::new(100, 300) assert_eq(feat.length(), 200) } +///| test "TrackFeature - length" { let feat = @src.TrackFeature::new(50, 150) assert_eq(feat.length(), 100) } +///| test "Track - feature operations" { let t = @src.Track::new("test_track") assert_eq(t.name(), "test_track") @@ -295,13 +423,22 @@ test "Track - feature operations" { assert_eq(t3.feature_count(), 2) } +///| test "gd_feature_count - diagram features only" { let d = @src.gd_create_diagram("test", 0, 1000) let d2 = @src.gd_add_feature(d, 50, 100, "+", "F1", @src.gd_rectangle_shape()) - let d3 = @src.gd_add_feature(d2, 200, 300, "+", "F2", @src.gd_rectangle_shape()) + let d3 = @src.gd_add_feature( + d2, + 200, + 300, + "+", + "F2", + @src.gd_rectangle_shape(), + ) assert_eq(@src.gd_feature_count(d3), 2) } +///| test "gd_to_svg_string - contains track names" { let d = @src.gd_create_diagram("test", 0, 500) let d2 = @src.gd_add_track(d, "MyTrack") @@ -310,6 +447,7 @@ test "gd_to_svg_string - contains track names" { assert_true(svg.contains("MyTrack")) } +///| test "gd_to_svg_string - border toggle" { let d = @src.gd_create_diagram("test", 0, 500) let style = @src.DiagramStyle::new(border=false) @@ -318,6 +456,7 @@ test "gd_to_svg_string - border toggle" { assert_true(svg.contains(" { @@ -240,6 +241,7 @@ test "gfa_parse_tag valid LN:i:100" { } } +///| test "gfa_parse_tag valid VN:Z:1.0" { match @src.gfa_parse_tag("VN:Z:1.0") { Some(t) => { @@ -251,6 +253,7 @@ test "gfa_parse_tag valid VN:Z:1.0" { } } +///| test "gfa_parse_tag invalid bad (no colons)" { match @src.gfa_parse_tag("bad") { Some(_) => assert_true(false) @@ -258,6 +261,7 @@ test "gfa_parse_tag invalid bad (no colons)" { } } +///| test "gfa_parse_tag missing type (one colon)" { match @src.gfa_parse_tag("LN:100") { Some(_) => assert_true(false) @@ -265,6 +269,7 @@ test "gfa_parse_tag missing type (one colon)" { } } +///| test "gfa_parse_tag empty string" { match @src.gfa_parse_tag("") { Some(_) => assert_true(false) @@ -272,6 +277,7 @@ test "gfa_parse_tag empty string" { } } +///| test "gfa_parse_tag with value containing colon" { // The value itself may legally contain ':' (e.g. JSON-style). The parser // splits on the first two colons only. @@ -289,6 +295,7 @@ test "gfa_parse_tag with value containing colon" { // 8. gfa_parse - full GFA document with H, S, L, P lines // ============================================================================ +///| test "gfa_parse full document" { let content = "H\tVN:Z:1.0\nS\tseq1\tACGTACGT\tLN:i:8\nS\tseq2\tTTTTGGGG\tLN:i:8\nL\tseq1\t+\tseq2\t-\t8M\nP\tpath1\tseq1+,seq2-\t8M,4M\n" let graph = @src.gfa_parse(content) @@ -315,6 +322,7 @@ test "gfa_parse full document" { assert_eq(graph.paths()[0].overlaps().length(), 2) } +///| test "gfa_parse handles CRLF line endings" { let content = "H\tVN:Z:1.0\r\nS\tseq1\tACGT\r\n" let graph = @src.gfa_parse(content) @@ -325,6 +333,7 @@ test "gfa_parse handles CRLF line endings" { assert_eq(graph.segments()[0].sequence(), "ACGT") } +///| test "gfa_parse skips blank and unknown lines" { let content = "\nH\tVN:Z:1.0\nX\tunknown\n\nS\tseq1\tACGT\n" let graph = @src.gfa_parse(content) @@ -333,6 +342,7 @@ test "gfa_parse skips blank and unknown lines" { assert_eq(graph.segments()[0].name(), "seq1") } +///| test "gfa_parse_containment" { let fields = ["C", "seq1", "+", "seq2", "+", "5", "8M"] let cont = @src.gfa_parse_containment(fields) @@ -346,6 +356,7 @@ test "gfa_parse_containment" { // 9. gfa_to_string round-trip // ============================================================================ +///| test "gfa_to_string round-trip preserves data" { let content = "H\tVN:Z:1.0\nS\tseq1\tACGTACGT\tLN:i:8\nS\tseq2\tTTTTGGGG\tLN:i:8\nL\tseq1\t+\tseq2\t-\t8M\nP\tpath1\tseq1+,seq2-\t8M,4M\n" let graph1 = @src.gfa_parse(content) @@ -366,6 +377,7 @@ test "gfa_to_string round-trip preserves data" { assert_eq(graph2.paths()[0].overlaps()[1], "4M") } +///| test "gfa_to_string matches expected lines" { let content = "H\tVN:Z:1.0\nS\tseq1\tACGT\tLN:i:4\nL\tseq1\t+\tseq2\t-\t8M\nP\tp1\tseq1+,seq2-\t8M,4M\n" let graph = @src.gfa_parse(content) @@ -380,71 +392,60 @@ test "gfa_to_string matches expected lines" { // Per-record serializers // ============================================================================ +///| test "gfa_header_to_line" { - let header = @src.GfaHeader::new( - "1.0", - [@src.GfaTag::new("VN", "Z", "1.0")], - ) + let header = @src.GfaHeader::new("1.0", [@src.GfaTag::new("VN", "Z", "1.0")]) assert_eq(@src.gfa_header_to_line(header), "H\tVN:Z:1.0") } +///| test "gfa_header_to_line multiple tags" { - let header = @src.GfaHeader::new( - "1.0", - [ - @src.GfaTag::new("VN", "Z", "1.0"), - @src.GfaTag::new("OR", "Z", "sample"), - ], - ) + let header = @src.GfaHeader::new("1.0", [ + @src.GfaTag::new("VN", "Z", "1.0"), + @src.GfaTag::new("OR", "Z", "sample"), + ]) assert_eq(@src.gfa_header_to_line(header), "H\tVN:Z:1.0\tOR:Z:sample") } +///| test "gfa_segment_to_line" { - let seg = @src.GfaSegment::new( - "seq1", - "ACGT", - [@src.GfaTag::new("LN", "i", "4")], - ) + let seg = @src.GfaSegment::new("seq1", "ACGT", [ + @src.GfaTag::new("LN", "i", "4"), + ]) assert_eq(@src.gfa_segment_to_line(seg), "S\tseq1\tACGT\tLN:i:4") } +///| test "gfa_segment_to_line no tags" { let seg = @src.GfaSegment::new("seq2", "TTTT", []) assert_eq(@src.gfa_segment_to_line(seg), "S\tseq2\tTTTT") } +///| test "gfa_link_to_line" { let link = @src.GfaLink::new("seq1", "+", "seq2", "-", "8M", []) assert_eq(@src.gfa_link_to_line(link), "L\tseq1\t+\tseq2\t-\t8M") } +///| test "gfa_link_to_line with tag" { - let link = @src.GfaLink::new( - "a", - "+", - "b", - "+", - "4M", - [@src.GfaTag::new("MQ", "i", "60")], - ) + let link = @src.GfaLink::new("a", "+", "b", "+", "4M", [ + @src.GfaTag::new("MQ", "i", "60"), + ]) assert_eq(@src.gfa_link_to_line(link), "L\ta\t+\tb\t+\t4M\tMQ:i:60") } +///| test "gfa_path_to_line" { let path = @src.GfaPath::new("p1", ["seq1+", "seq2-"], ["8M", "4M"], []) assert_eq(@src.gfa_path_to_line(path), "P\tp1\tseq1+,seq2-\t8M,4M") } +///| test "gfa_containment_to_line" { - let cont = @src.GfaContainment::new( - "seq1", - "+", - "seq2", - "+", - 5, - "8M", - [@src.GfaTag::new("MQ", "i", "60")], - ) + let cont = @src.GfaContainment::new("seq1", "+", "seq2", "+", 5, "8M", [ + @src.GfaTag::new("MQ", "i", "60"), + ]) assert_eq( @src.gfa_containment_to_line(cont), "C\tseq1\t+\tseq2\t+\t5\t8M\tMQ:i:60", @@ -455,6 +456,7 @@ test "gfa_containment_to_line" { // Per-record parsers // ============================================================================ +///| test "gfa_parse_header extracts version" { let fields = ["H", "VN:Z:1.0"] let header = @src.gfa_parse_header(fields) @@ -462,6 +464,7 @@ test "gfa_parse_header extracts version" { assert_eq(header.tags().length(), 1) } +///| test "gfa_parse_segment fields" { let fields = ["S", "seq1", "ACGTACGT", "LN:i:8"] let seg = @src.gfa_parse_segment(fields) @@ -472,6 +475,7 @@ test "gfa_parse_segment fields" { assert_eq(seg.tags()[0].value(), "8") } +///| test "gfa_parse_segment defaults sequence to *" { let fields = ["S", "seq1"] let seg = @src.gfa_parse_segment(fields) @@ -480,6 +484,7 @@ test "gfa_parse_segment defaults sequence to *" { assert_eq(seg.tags().length(), 0) } +///| test "gfa_parse_link fields" { let fields = ["L", "seq1", "+", "seq2", "-", "8M"] let link = @src.gfa_parse_link(fields) @@ -490,6 +495,7 @@ test "gfa_parse_link fields" { assert_eq(link.overlap(), "8M") } +///| test "gfa_parse_path fields" { let fields = ["P", "path1", "seq1+,seq2-", "8M,4M"] let path = @src.gfa_parse_path(fields) @@ -506,6 +512,7 @@ test "gfa_parse_path fields" { // 10. gfa_graph_n_segments / n_links / n_paths // ============================================================================ +///| test "gfa_graph_n_segments/links/paths on sample" { let graph = @src.gfa_sample_graph() assert_eq(@src.gfa_graph_n_segments(graph), 3) @@ -513,6 +520,7 @@ test "gfa_graph_n_segments/links/paths on sample" { assert_eq(@src.gfa_graph_n_paths(graph), 1) } +///| test "gfa_graph_n_segments/links/paths on empty" { let graph = @src.GfaGraph::new() assert_eq(@src.gfa_graph_n_segments(graph), 0) @@ -524,6 +532,7 @@ test "gfa_graph_n_segments/links/paths on empty" { // 11. gfa_get_segment - found and not found // ============================================================================ +///| test "gfa_get_segment found" { let graph = @src.gfa_sample_graph() match @src.gfa_get_segment(graph, "seq2") { @@ -535,6 +544,7 @@ test "gfa_get_segment found" { } } +///| test "gfa_get_segment found first" { let graph = @src.gfa_sample_graph() match @src.gfa_get_segment(graph, "seq1") { @@ -543,6 +553,7 @@ test "gfa_get_segment found first" { } } +///| test "gfa_get_segment not found" { let graph = @src.gfa_sample_graph() match @src.gfa_get_segment(graph, "missing") { @@ -551,6 +562,7 @@ test "gfa_get_segment not found" { } } +///| test "gfa_get_segment on empty graph" { let graph = @src.GfaGraph::new() match @src.gfa_get_segment(graph, "anything") { @@ -563,6 +575,7 @@ test "gfa_get_segment on empty graph" { // 12. gfa_get_segments_as_records - convert to SeqRecord array // ============================================================================ +///| test "gfa_get_segments_as_records sample graph" { let graph = @src.gfa_sample_graph() let records = @src.gfa_get_segments_as_records(graph) @@ -577,12 +590,14 @@ test "gfa_get_segments_as_records sample graph" { assert_eq(records[2].seq.to_string(), "CCCCAAAA") } +///| test "gfa_get_segments_as_records empty graph" { let graph = @src.GfaGraph::new() let records = @src.gfa_get_segments_as_records(graph) assert_eq(records.length(), 0) } +///| test "gfa_get_segments_as_records star sequence becomes empty" { let graph = @src.GfaGraph::new() graph.add_segment(@src.GfaSegment::new("masked", "*", [])) @@ -596,15 +611,11 @@ test "gfa_get_segments_as_records star sequence becomes empty" { // 13. gfa_reverse_link - from/to swapped, orientations swapped // ============================================================================ +///| test "gfa_reverse_link swaps segments and orientations" { - let link = @src.GfaLink::new( - "seq1", - "+", - "seq2", - "-", - "8M", - [@src.GfaTag::new("MQ", "i", "60")], - ) + let link = @src.GfaLink::new("seq1", "+", "seq2", "-", "8M", [ + @src.GfaTag::new("MQ", "i", "60"), + ]) let rev = @src.gfa_reverse_link(link) assert_eq(rev.from_segment(), "seq2") assert_eq(rev.from_orient(), "-") @@ -615,6 +626,7 @@ test "gfa_reverse_link swaps segments and orientations" { assert_eq(rev.tags()[0].name(), "MQ") } +///| test "gfa_reverse_link double-reverse is identity" { let link = @src.GfaLink::new("a", "+", "b", "-", "4M", []) let rev2 = @src.gfa_reverse_link(@src.gfa_reverse_link(link)) @@ -625,18 +637,12 @@ test "gfa_reverse_link double-reverse is identity" { assert_eq(rev2.overlap(), "4M") } +///| test "gfa_reverse_link preserves tags" { - let link = @src.GfaLink::new( - "x", - "+", - "y", - "+", - "0M", - [ - @src.GfaTag::new("MQ", "i", "40"), - @src.GfaTag::new("NM", "i", "0"), - ], - ) + let link = @src.GfaLink::new("x", "+", "y", "+", "0M", [ + @src.GfaTag::new("MQ", "i", "40"), + @src.GfaTag::new("NM", "i", "0"), + ]) let rev = @src.gfa_reverse_link(link) assert_eq(rev.tags().length(), 2) assert_eq(rev.tags()[0].name(), "MQ") @@ -647,6 +653,7 @@ test "gfa_reverse_link preserves tags" { // 14. gfa_graph_summary - non-empty string // ============================================================================ +///| test "gfa_graph_summary non-empty on sample" { let graph = @src.gfa_sample_graph() let summary = @src.gfa_graph_summary(graph) @@ -658,6 +665,7 @@ test "gfa_graph_summary non-empty on sample" { assert_true(summary.contains("total_sequence_length=24")) } +///| test "gfa_graph_summary on empty graph" { let graph = @src.GfaGraph::new() let summary = @src.gfa_graph_summary(graph) @@ -668,6 +676,7 @@ test "gfa_graph_summary on empty graph" { assert_true(summary.contains("total_sequence_length=0")) } +///| test "gfa_graph_summary ignores star sequences" { let graph = @src.GfaGraph::new() graph.add_segment(@src.GfaSegment::new("a", "ACGT", [])) @@ -681,6 +690,7 @@ test "gfa_graph_summary ignores star sequences" { // 15. Sample graph validation // ============================================================================ +///| test "gfa_sample_graph structure" { let graph = @src.gfa_sample_graph() assert_eq(graph.headers().length(), 1) @@ -691,6 +701,7 @@ test "gfa_sample_graph structure" { assert_eq(graph.containments().length(), 0) } +///| test "gfa_sample_graph segments" { let graph = @src.gfa_sample_graph() assert_eq(graph.segments()[0].name(), "seq1") @@ -704,6 +715,7 @@ test "gfa_sample_graph segments" { assert_eq(graph.segments()[0].tags()[0].value(), "8") } +///| test "gfa_sample_graph links" { let graph = @src.gfa_sample_graph() let l0 = graph.links()[0] @@ -716,6 +728,7 @@ test "gfa_sample_graph links" { assert_eq(l1.overlap(), "4M") } +///| test "gfa_sample_graph path" { let graph = @src.gfa_sample_graph() let p = graph.paths()[0] @@ -727,6 +740,7 @@ test "gfa_sample_graph path" { assert_eq(p.overlaps().length(), 2) } +///| test "gfa_sample_graph round-trips through gfa_to_string" { let graph = @src.gfa_sample_graph() let text = @src.gfa_to_string(graph) @@ -741,12 +755,14 @@ test "gfa_sample_graph round-trips through gfa_to_string" { // 16. Edge cases - empty graph, single segment, no links // ============================================================================ +///| test "edge case: empty graph serializes to empty string" { let graph = @src.GfaGraph::new() let s = @src.gfa_to_string(graph) assert_eq(s, "") } +///| test "edge case: parse empty content" { let graph = @src.gfa_parse("") assert_eq(@src.gfa_graph_n_segments(graph), 0) @@ -754,6 +770,7 @@ test "edge case: parse empty content" { assert_eq(@src.gfa_graph_n_paths(graph), 0) } +///| test "edge case: single segment, no links" { let graph = @src.GfaGraph::new() graph.add_segment(@src.GfaSegment::new("only", "ACGT", [])) @@ -764,6 +781,7 @@ test "edge case: single segment, no links" { assert_eq(out, "S\tonly\tACGT\n") } +///| test "edge case: single segment round-trip" { let graph = @src.GfaGraph::new() graph.add_segment(@src.GfaSegment::new("only", "ACGT", [])) @@ -774,6 +792,7 @@ test "edge case: single segment round-trip" { assert_eq(reparsed.segments()[0].sequence(), "ACGT") } +///| test "edge case: segment with star sequence" { let content = "S\tseq1\t*\tLN:i:100\n" let graph = @src.gfa_parse(content) @@ -782,6 +801,7 @@ test "edge case: segment with star sequence" { assert_eq(graph.segments()[0].tags()[0].name(), "LN") } +///| test "edge case: only headers" { let content = "H\tVN:Z:1.0\nH\tOR:Z:test\n" let graph = @src.gfa_parse(content) @@ -789,6 +809,7 @@ test "edge case: only headers" { assert_eq(@src.gfa_graph_n_segments(graph), 0) } +///| test "edge case: unknown record type ignored" { let content = "H\tVN:Z:1.0\nZ\tunknown\tfield\nS\tseq1\tACGT\n" let graph = @src.gfa_parse(content) diff --git a/test/moonbit/gff_test.mbt b/test/moonbit/gff_test.mbt index 4def9b1f..bf869e37 100644 --- a/test/moonbit/gff_test.mbt +++ b/test/moonbit/gff_test.mbt @@ -1,6 +1,5 @@ ///| /// GFF module tests - test "GFFFeature::new" { let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") assert_eq(feature.seqid, "chr1") @@ -9,6 +8,7 @@ test "GFFFeature::new" { assert_eq(feature.end, 2000) } +///| test "GFFFeature::add_attribute" { let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") .add_attribute("ID", "gene001") @@ -19,20 +19,24 @@ test "GFFFeature::add_attribute" { } } +///| test "GFFFeature::get_id" { - let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") - .add_attribute("ID", "gene001") + let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+").add_attribute( + "ID", "gene001", + ) match feature.get_id() { Some(id) => assert_eq(id, "gene001") None => assert_true(false) } } +///| test "GFFFeature::length" { let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") assert_eq(feature.length(), 1001) } +///| test "GFFFeature::is_coding" { let gene = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") assert_true(gene.is_coding()) @@ -40,11 +44,13 @@ test "GFFFeature::is_coding" { assert_true(cds.is_coding()) } +///| test "GFFRecord::new" { let record = @src.GFFRecord::new() assert_eq(record.count_features(), 0) } +///| test "GFFRecord::add_feature" { let record = @src.GFFRecord::new() let feature = @src.GFFFeature::new("chr1", "Ensembl", "gene", 1000, 2000, "+") @@ -52,18 +58,21 @@ test "GFFRecord::add_feature" { assert_eq(record.count_features(), 1) } +///| test "GFFRecord::get_features_by_type" { let record = @src.create_example_gff() let genes = record.get_genes() assert_true(genes.length() >= 1) } +///| test "GFFRecord::get_features_by_seqid" { let record = @src.create_example_gff() let features = record.get_features_by_seqid("chr1") assert_true(features.length() > 0) } +///| test "parse_gff" { let content = "##gff-version 3\nchr1\tEnsembl\tgene\t1000\t2000\t.\t+\t.\tID=gene001;Name=TP53\n" let record = @src.bio_parse_gff(content) @@ -71,8 +80,11 @@ test "parse_gff" { assert_eq(record.version, "3") } +///| test "parse_attributes" { - let attrs = @src.parse_attributes("ID=gene001;Name=TP53;biotype=protein_coding") + let attrs = @src.parse_attributes( + "ID=gene001;Name=TP53;biotype=protein_coding", + ) match attrs.get("ID") { Some(id) => assert_eq(id, "gene001") None => assert_true(false) @@ -83,6 +95,7 @@ test "parse_attributes" { } } +///| test "parse_attributes_unescape" { let attrs = @src.parse_attributes("ID=gene001;Name=value%3Bwith%3Bsemicolons") match attrs.get("Name") { @@ -91,25 +104,29 @@ test "parse_attributes_unescape" { } } +///| test "create_example_gff" { let record = @src.create_example_gff() assert_true(record.count_features() > 5) } +///| test "GFFRecord::get_child_features" { let record = @src.create_example_gff() let mrnas = record.get_child_features("gene:ENSG00000130203") assert_true(mrnas.length() >= 1) } +///| test "GFFRecord::get_unique_seqids" { let record = @src.create_example_gff() let seqids = record.get_unique_seqids() assert_true(seqids.length() >= 1) } +///| test "GFFRecord::get_features_in_range" { let record = @src.create_example_gff() let features = record.get_features_in_range("chr1", 10000, 11000) assert_true(features.length() > 0) -} \ No newline at end of file +} diff --git a/test/moonbit/ggtree_test.mbt b/test/moonbit/ggtree_test.mbt index 55834325..deabca30 100644 --- a/test/moonbit/ggtree_test.mbt +++ b/test/moonbit/ggtree_test.mbt @@ -32,8 +32,12 @@ test "ggtree_count_leaves" { test "ggtree_y_positions" { let nodes = @src.create_test_tree() let root = nodes[0] - let n_nodes = @src.ggtree_count_leaves(root) + @src.ggtree_collect_internal_ids(root).length() - let positions : Map[String, (Double, Double, Double)] = Map([], capacity=n_nodes) + let n_nodes = @src.ggtree_count_leaves(root) + + @src.ggtree_collect_internal_ids(root).length() + let positions : Map[String, (Double, Double, Double)] = Map( + [], + capacity=n_nodes, + ) let _ = @src.ggtree_compute_y_positions(root, 0.0, positions) // Root should have a y coordinate @@ -52,8 +56,12 @@ test "ggtree_y_positions" { test "ggtree_x_positions" { let nodes = @src.create_test_tree() let root = nodes[0] - let n_nodes = @src.ggtree_count_leaves(root) + @src.ggtree_collect_internal_ids(root).length() - let positions : Map[String, (Double, Double, Double)] = Map([], capacity=n_nodes) + let n_nodes = @src.ggtree_count_leaves(root) + + @src.ggtree_collect_internal_ids(root).length() + let positions : Map[String, (Double, Double, Double)] = Map( + [], + capacity=n_nodes, + ) let _ = @src.ggtree_compute_y_positions(root, 0.0, positions) @src.ggtree_compute_x_positions(root, 0.0, positions) diff --git a/test/moonbit/glm_gampoi_test.mbt b/test/moonbit/glm_gampoi_test.mbt index 4dc2f6dd..ff6ee84d 100644 --- a/test/moonbit/glm_gampoi_test.mbt +++ b/test/moonbit/glm_gampoi_test.mbt @@ -6,11 +6,11 @@ // --------------------------------------------------------------------------- test "glm_design_creation" { - let d = @src.GlmDesign::new( - ["s1", "s2", "s3"], - ["Intercept", "Treatment"], - [[1.0, 0.0], [1.0, 0.0], [1.0, 1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3"], ["Intercept", "Treatment"], [ + [1.0, 0.0], + [1.0, 0.0], + [1.0, 1.0], + ]) assert_eq(d.n_samples(), 3) assert_eq(d.n_coefs(), 2) assert_eq(d.get(0, 0), 1.0) @@ -19,27 +19,23 @@ test "glm_design_creation" { assert_eq(d.get(2, 1), 1.0) } +///| test "glm_design_intercept_only" { - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) assert_eq(d.n_samples(), 4) assert_eq(d.n_coefs(), 1) assert_eq(d.get(0, 0), 1.0) assert_eq(d.get(3, 0), 1.0) } +///| test "glm_fit_creation" { - let f = @src.GlmFit::new( - "GeneA", - [2.5, 1.2], - [0.3, 0.4], - 0.15, - 10, - 8.5, - ) + let f = @src.GlmFit::new("GeneA", [2.5, 1.2], [0.3, 0.4], 0.15, 10, 8.5) assert_eq(f.gene, "GeneA") assert_eq(f.coefficients.length(), 2) assert_eq(f.coefficients[0], 2.5) @@ -51,57 +47,33 @@ test "glm_fit_creation" { assert_eq(f.df_dispersion, 8.5) } +///| test "glm_fit_single_coef" { - let f = @src.GlmFit::new( - "GeneX", - [3.0], - [0.5], - 0.2, - 5, - 5.0, - ) + let f = @src.GlmFit::new("GeneX", [3.0], [0.5], 0.2, 5, 5.0) assert_eq(f.coefficients.length(), 1) assert_eq(f.coefficients[0], 3.0) assert_eq(f.std_errors[0], 0.5) } +///| test "glm_test_result_creation" { - let r = @src.GlmTestResult::new( - "GeneA", - "Treatment", - 1.5, - 0.3, - 5.0, - 0.001, - ) + let r = @src.GlmTestResult::new("GeneA", "Treatment", 1.5, 0.3, 5.0, 0.001) assert_eq(r.estimate(), 1.5) assert_eq(r.p_value(), 0.001) assert_eq(r.adj_p_value(), 0.001) } +///| test "glm_test_result_not_significant" { - let r = @src.GlmTestResult::new( - "GeneB", - "Treatment", - 0.1, - 0.5, - 0.2, - 0.8, - ) + let r = @src.GlmTestResult::new("GeneB", "Treatment", 0.1, 0.5, 0.2, 0.8) assert_false(r.is_significant) assert_eq(r.estimate(), 0.1) assert_eq(r.p_value(), 0.8) } +///| test "glm_test_result_to_string" { - let r = @src.GlmTestResult::new( - "GeneA", - "Treatment", - 1.5, - 0.3, - 5.0, - 0.001, - ) + let r = @src.GlmTestResult::new("GeneA", "Treatment", 1.5, 0.3, 5.0, 0.001) let s = r.to_string() assert_true(s.contains("GeneA")) assert_true(s.contains("Treatment")) @@ -109,19 +81,14 @@ test "glm_test_result_to_string" { assert_true(s.contains("p=")) } +///| test "glm_test_result_significant_marker" { - let r = @src.GlmTestResult::new( - "GeneC", - "Treatment", - 3.0, - 0.2, - 15.0, - 1.0e-10, - ) + let r = @src.GlmTestResult::new("GeneC", "Treatment", 3.0, 0.2, 15.0, 1.0e-10) let s = r.to_string() assert_true(s.contains("*")) } +///| test "glm_pseudobulk_creation" { let pb = @src.GlmPseudobulk::new( ["GeneA", "GeneB"], @@ -133,13 +100,9 @@ test "glm_pseudobulk_creation" { assert_eq(pb.n_groups(), 2) } +///| test "glm_pseudobulk_single_gene" { - let pb = @src.GlmPseudobulk::new( - ["GeneX"], - ["GroupA"], - [["s1"]], - [[42.0]], - ) + let pb = @src.GlmPseudobulk::new(["GeneX"], ["GroupA"], [["s1"]], [[42.0]]) assert_eq(pb.n_genes(), 1) assert_eq(pb.n_groups(), 1) } @@ -148,83 +111,101 @@ test "glm_pseudobulk_single_gene" { // Utility functions // --------------------------------------------------------------------------- +///| test "glm_sum_basic" { let s = @src.glm_sum([1.0, 2.0, 3.0, 4.0]) assert_eq(s, 10.0) } +///| test "glm_sum_single" { assert_eq(@src.glm_sum([5.0]), 5.0) } +///| test "glm_sum_empty" { assert_eq(@src.glm_sum([]), 0.0) } +///| test "glm_sum_negative" { assert_eq(@src.glm_sum([-1.0, -2.0, 3.0]), 0.0) } +///| test "glm_mean_basic" { let m = @src.glm_mean([2.0, 4.0, 6.0]) assert_eq(m, 4.0) } +///| test "glm_mean_single" { assert_eq(@src.glm_mean([7.0]), 7.0) } +///| test "glm_mean_empty" { assert_eq(@src.glm_mean([]), 0.0) } +///| test "glm_mean_zeros" { assert_eq(@src.glm_mean([0.0, 0.0, 0.0]), 0.0) } +///| test "glm_median_odd" { let m = @src.glm_median([3.0, 1.0, 2.0]) assert_eq(m, 2.0) } +///| test "glm_median_even" { let m = @src.glm_median([1.0, 3.0, 2.0, 4.0]) assert_eq(m, 2.5) } +///| test "glm_median_single" { assert_eq(@src.glm_median([5.0]), 5.0) } +///| test "glm_median_empty" { assert_eq(@src.glm_median([]), 0.0) } +///| test "glm_median_already_sorted" { let m = @src.glm_median([1.0, 2.0, 3.0, 4.0, 5.0]) assert_eq(m, 3.0) } +///| test "glm_min_positive_basic" { let v = @src.glm_min_positive([0.0, 3.0, 2.0, 5.0]) assert_eq(v, 2.0) } +///| test "glm_min_positive_no_positive" { let v = @src.glm_min_positive([0.0, -1.0, -5.0]) assert_eq(v, 0.0) } +///| test "glm_min_positive_all_positive" { let v = @src.glm_min_positive([10.0, 3.0, 7.0]) assert_eq(v, 3.0) } +///| test "glm_min_positive_single" { let v = @src.glm_min_positive([42.0]) assert_eq(v, 42.0) } +///| test "glm_min_positive_empty" { let v = @src.glm_min_positive([]) assert_eq(v, 0.0) @@ -234,6 +215,7 @@ test "glm_min_positive_empty" { // Linear algebra: solve linear system // --------------------------------------------------------------------------- +///| test "glm_solve_linear_system_2x2" { // 2x + 3y = 7 // x + y = 3 @@ -245,6 +227,7 @@ test "glm_solve_linear_system_2x2" { assert_true((x[1] - 1.0).abs() < 1.0e-6) } +///| test "glm_solve_linear_system_3x3" { // x + y + z = 6 // 2x + y - z = 1 @@ -258,15 +241,17 @@ test "glm_solve_linear_system_3x3" { assert_true((x[2] - 3.0).abs() < 1.0e-6) } +///| test "glm_solve_linear_system_identity" { // Identity * x = b => x = b let a = [[1.0, 0.0], [0.0, 1.0]] let b = [5.0, -3.0] let x = @src.glm_solve_linear_system(a, b, 2) assert_true((x[0] - 5.0).abs() < 1.0e-6) - assert_true((x[1] - (-3.0)).abs() < 1.0e-6) + assert_true((x[1] - -3.0).abs() < 1.0e-6) } +///| test "glm_solve_linear_system_1x1" { let a = [[3.0]] let b = [12.0] @@ -274,6 +259,7 @@ test "glm_solve_linear_system_1x1" { assert_true((x[0] - 4.0).abs() < 1.0e-6) } +///| test "glm_solve_linear_system_diagonal" { // Diagonal 3x3 system // 3x = 9, 2y = 8, 4z = 16 @@ -289,17 +275,19 @@ test "glm_solve_linear_system_diagonal" { // Linear algebra: matrix inverse // --------------------------------------------------------------------------- +///| test "glm_invert_matrix_2x2" { // A = [[4, 7], [2, 6]] // A^{-1} = [[0.6, -0.7], [-0.2, 0.4]] let a = [[4.0, 7.0], [2.0, 6.0]] let inv = @src.glm_invert_matrix(a, 2) assert_true((inv[0][0] - 0.6).abs() < 1.0e-6) - assert_true((inv[0][1] - (-0.7)).abs() < 1.0e-6) - assert_true((inv[1][0] - (-0.2)).abs() < 1.0e-6) + assert_true((inv[0][1] - -0.7).abs() < 1.0e-6) + assert_true((inv[1][0] - -0.2).abs() < 1.0e-6) assert_true((inv[1][1] - 0.4).abs() < 1.0e-6) } +///| test "glm_invert_matrix_identity_3x3" { let a = [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]] let inv = @src.glm_invert_matrix(a, 3) @@ -311,12 +299,14 @@ test "glm_invert_matrix_identity_3x3" { } } +///| test "glm_invert_matrix_1x1" { let a = [[5.0]] let inv = @src.glm_invert_matrix(a, 1) assert_true((inv[0][0] - 0.2).abs() < 1.0e-6) } +///| test "glm_invert_matrix_verify_product" { // A * A^{-1} should be identity let a = [[2.0, 1.0], [5.0, 3.0]] @@ -337,6 +327,7 @@ test "glm_invert_matrix_verify_product" { // Statistical utilities // --------------------------------------------------------------------------- +///| test "glm_phi_standard_normal" { // phi(0) ≈ 0.5 let p0 = @src.glm_phi(0.0) @@ -349,6 +340,7 @@ test "glm_phi_standard_normal" { assert_true((p2 - 0.025).abs() < 0.01) } +///| test "glm_phi_symmetry" { // phi(x) + phi(-x) = 1 let x = 1.5 @@ -356,6 +348,7 @@ test "glm_phi_symmetry" { assert_true((p_sum - 1.0).abs() < 1.0e-10) } +///| test "glm_phi_non_decreasing" { // phi should be non-decreasing let p1 = @src.glm_phi(-2.0) @@ -369,6 +362,7 @@ test "glm_phi_non_decreasing" { assert_true(p4 <= p5) } +///| test "glm_phi_clamped" { // phi should be in [0, 1] assert_true(@src.glm_phi(-10.0) >= 0.0) @@ -377,18 +371,21 @@ test "glm_phi_clamped" { assert_true(@src.glm_phi(10.0) <= 1.0) } +///| test "glm_normal_pvalue_two_sided_zero" { // z=0 => p=1.0 let p = @src.glm_normal_pvalue_two_sided(0.0) assert_true((p - 1.0).abs() < 0.01) } +///| test "glm_normal_pvalue_two_sided_196" { // z=1.96 => p≈0.05 let p = @src.glm_normal_pvalue_two_sided(1.96) assert_true((p - 0.05).abs() < 0.01) } +///| test "glm_normal_pvalue_two_sided_symmetry" { // z and -z should give same p-value let p1 = @src.glm_normal_pvalue_two_sided(2.5) @@ -396,6 +393,7 @@ test "glm_normal_pvalue_two_sided_symmetry" { assert_true((p1 - p2).abs() < 1.0e-10) } +///| test "glm_normal_pvalue_two_sided_in_range" { // All p-values should be in [0, 1] for z in [-5.0, -2.0, -1.0, 0.0, 1.0, 2.0, 5.0] { @@ -405,6 +403,7 @@ test "glm_normal_pvalue_two_sided_in_range" { } } +///| test "glm_normal_pvalue_two_sided_large_z" { // Large |z| should give very small p-value (clamped to 1e-15) let p = @src.glm_normal_pvalue_two_sided(10.0) @@ -415,6 +414,7 @@ test "glm_normal_pvalue_two_sided_large_z" { // Sorting // --------------------------------------------------------------------------- +///| test "glm_sort_pairs_basic" { let arr = [(3.0, 2), (1.0, 0), (2.0, 1)] @src.glm_sort_pairs(arr) @@ -426,6 +426,7 @@ test "glm_sort_pairs_basic" { assert_eq(arr[2].1, 2) } +///| test "glm_sort_pairs_already_sorted" { let arr = [(1.0, 0), (2.0, 1), (3.0, 2)] @src.glm_sort_pairs(arr) @@ -434,6 +435,7 @@ test "glm_sort_pairs_already_sorted" { assert_eq(arr[2].0, 3.0) } +///| test "glm_sort_pairs_single" { let arr = [(5.0, 0)] @src.glm_sort_pairs(arr) @@ -441,6 +443,7 @@ test "glm_sort_pairs_single" { assert_eq(arr[0].0, 5.0) } +///| test "glm_sort_pairs_empty" { let arr : Array[(Double, Int)] = [] @src.glm_sort_pairs(arr) @@ -451,12 +454,14 @@ test "glm_sort_pairs_empty" { // BH-FDR correction // --------------------------------------------------------------------------- +///| test "glm_bh_correct_single" { // Single p-value correction returns same value let p = @src.glm_bh_correct(0.05) assert_eq(p, 0.05) } +///| test "glm_bh_correct_array_known_values" { // p-values: [0.01, 0.04, 0.03, 0.005] // Sorted: [0.005, 0.01, 0.03, 0.04] @@ -473,6 +478,7 @@ test "glm_bh_correct_array_known_values" { assert_true((adj[3] - 0.02).abs() < 1.0e-6) } +///| test "glm_bh_correct_array_all_same" { let pvals = [0.05, 0.05, 0.05] let adj = @src.glm_bh_correct_array(pvals) @@ -481,17 +487,20 @@ test "glm_bh_correct_array_all_same" { assert_true((adj[1] - adj[2]).abs() < 1.0e-10) } +///| test "glm_bh_correct_array_single" { let adj = @src.glm_bh_correct_array([0.02]) assert_eq(adj.length(), 1) assert_true((adj[0] - 0.02).abs() < 1.0e-6) } +///| test "glm_bh_correct_array_empty" { let adj = @src.glm_bh_correct_array([]) assert_eq(adj.length(), 0) } +///| test "glm_bh_correct_array_values_in_range" { let pvals = [0.001, 0.02, 0.05, 0.1, 0.5] let adj = @src.glm_bh_correct_array(pvals) @@ -501,15 +510,11 @@ test "glm_bh_correct_array_values_in_range" { } } +///| test "glm_bh_correct_array_monotonic" { let pvals = [0.001, 0.01, 0.03, 0.04] let adj = @src.glm_bh_correct_array(pvals) - let indexed = [ - (pvals[0], 0), - (pvals[1], 1), - (pvals[2], 2), - (pvals[3], 3), - ] + let indexed = [(pvals[0], 0), (pvals[1], 1), (pvals[2], 2), (pvals[3], 3)] @src.glm_sort_pairs(indexed) let adj0 = adj[indexed[0].1] let adj1 = adj[indexed[1].1] @@ -524,15 +529,12 @@ test "glm_bh_correct_array_monotonic" { // Size factor calculation // --------------------------------------------------------------------------- +///| test "glm_calculate_sf_basic" { // 3 genes, 2 samples // Sample 0: [10, 20, 30], Sample 1: [20, 40, 60] // Sample 1 has twice the counts, so sf[1] ≈ 2 * sf[0] - let counts = [ - [10.0, 20.0], - [20.0, 40.0], - [30.0, 60.0], - ] + let counts = [[10.0, 20.0], [20.0, 40.0], [30.0, 60.0]] let sf = @src.glm_calculate_sf(counts) assert_eq(sf.length(), 2) assert_true(sf[0] > 0.0) @@ -541,11 +543,9 @@ test "glm_calculate_sf_basic" { assert_true((sf[1] / sf[0] - 2.0).abs() < 0.5) } +///| test "glm_calculate_sf_equal_samples" { - let counts = [ - [10.0, 10.0, 10.0], - [20.0, 20.0, 20.0], - ] + let counts = [[10.0, 10.0, 10.0], [20.0, 20.0, 20.0]] let sf = @src.glm_calculate_sf(counts) assert_eq(sf.length(), 3) // All samples have same counts, so sf should be equal @@ -553,6 +553,7 @@ test "glm_calculate_sf_equal_samples" { assert_true((sf[1] - sf[2]).abs() < 1.0e-6) } +///| test "glm_calculate_sf_single_sample" { let counts = [[5.0], [10.0], [15.0]] let sf = @src.glm_calculate_sf(counts) @@ -560,16 +561,15 @@ test "glm_calculate_sf_single_sample" { assert_true(sf[0] > 0.0) } +///| test "glm_calculate_sf_empty" { let sf = @src.glm_calculate_sf([]) assert_eq(sf.length(), 0) } +///| test "glm_calculate_sf_all_zeros" { - let counts = [ - [0.0, 0.0], - [0.0, 0.0], - ] + let counts = [[0.0, 0.0], [0.0, 0.0]] let sf = @src.glm_calculate_sf(counts) assert_eq(sf.length(), 2) // Should fallback to 1.0 for zero-count samples @@ -577,12 +577,9 @@ test "glm_calculate_sf_all_zeros" { assert_true(sf[1] > 0.0) } +///| test "glm_calculate_sf_genes_with_zeros" { - let counts = [ - [10.0, 0.0], - [0.0, 20.0], - [30.0, 30.0], - ] + let counts = [[10.0, 0.0], [0.0, 20.0], [30.0, 30.0]] let sf = @src.glm_calculate_sf(counts) assert_eq(sf.length(), 2) assert_true(sf[0] > 0.0) @@ -593,12 +590,10 @@ test "glm_calculate_sf_genes_with_zeros" { // Pseudobulk aggregation // --------------------------------------------------------------------------- +///| test "glm_pseudobulk_basic" { // 2 genes, 4 samples (2 control, 2 treated) - let counts = [ - [10.0, 12.0, 50.0, 48.0], - [20.0, 22.0, 18.0, 20.0], - ] + let counts = [[10.0, 12.0, 50.0, 48.0], [20.0, 22.0, 18.0, 20.0]] let gene_names = ["GeneA", "GeneB"] let group_labels = ["control", "control", "treated", "treated"] let pb = @src.glm_pseudobulk(counts, gene_names, group_labels) @@ -606,6 +601,7 @@ test "glm_pseudobulk_basic" { assert_eq(pb.n_groups(), 2) } +///| test "glm_pseudobulk_single_group" { let counts = [[10.0, 20.0, 30.0]] let gene_names = ["GeneX"] @@ -615,10 +611,9 @@ test "glm_pseudobulk_single_group" { assert_eq(pb.n_groups(), 1) } +///| test "glm_pseudobulk_single_sample_per_group" { - let counts = [ - [5.0, 10.0, 15.0], - ] + let counts = [[5.0, 10.0, 15.0]] let gene_names = ["GeneX"] let group_labels = ["A", "B", "C"] let pb = @src.glm_pseudobulk(counts, gene_names, group_labels) @@ -626,10 +621,11 @@ test "glm_pseudobulk_single_sample_per_group" { assert_eq(pb.n_groups(), 3) } +///| test "glm_pseudobulk_empty" { - let counts: Array[Array[Double]] = [] - let gene_names: Array[String] = [] - let group_labels: Array[String] = [] + let counts : Array[Array[Double]] = [] + let gene_names : Array[String] = [] + let group_labels : Array[String] = [] let pb = @src.glm_pseudobulk(counts, gene_names, group_labels) assert_eq(pb.n_genes(), 0) assert_eq(pb.n_groups(), 0) @@ -639,13 +635,15 @@ test "glm_pseudobulk_empty" { // Initial estimates and dispersion // --------------------------------------------------------------------------- +///| test "glm_initial_estimates_constant_counts" { let counts = [10.0, 10.0, 10.0, 10.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let coefs = @src.glm_initial_estimates(counts, d) assert_eq(coefs.length(), 1) // With constant counts and intercept-only design, @@ -654,6 +652,7 @@ test "glm_initial_estimates_constant_counts" { assert_true((coefs[0] - expected_ln).abs() < 0.5) } +///| test "glm_initial_estimates_two_groups" { let counts = [10.0, 12.0, 50.0, 48.0] let d = @src.GlmDesign::new( @@ -667,25 +666,29 @@ test "glm_initial_estimates_two_groups" { assert_true(coefs[1] > 0.0) } +///| test "glm_initial_estimates_with_zeros" { let counts = [0.0, 10.0, 20.0, 30.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let coefs = @src.glm_initial_estimates(counts, d) assert_eq(coefs.length(), 1) assert_true(coefs[0] > 0.0) } +///| test "glm_estimate_dispersion_low" { let counts = [10.0, 11.0, 10.5, 9.5] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let coefs = @src.glm_initial_estimates(counts, d) let disp = @src.glm_estimate_dispersion(counts, d, coefs) // Low dispersion for tightly clustered counts @@ -693,26 +696,30 @@ test "glm_estimate_dispersion_low" { assert_true(disp < 5.0) } +///| test "glm_estimate_dispersion_high" { let counts = [1.0, 50.0, 2.0, 48.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let coefs = @src.glm_initial_estimates(counts, d) let disp = @src.glm_estimate_dispersion(counts, d, coefs) // High dispersion for highly variable counts assert_true(disp >= 0.0) } +///| test "glm_estimate_dispersion_all_same" { let counts = [20.0, 20.0, 20.0, 20.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let coefs = @src.glm_initial_estimates(counts, d) let disp = @src.glm_estimate_dispersion(counts, d, coefs) // Zero variance => dispersion should be very small or 0 @@ -723,6 +730,7 @@ test "glm_estimate_dispersion_all_same" { // IWLCS fit // --------------------------------------------------------------------------- +///| test "glm_iwlcs_fit_produces_coefs" { let counts = [10.0, 12.0, 50.0, 48.0] let d = @src.GlmDesign::new( @@ -737,13 +745,15 @@ test "glm_iwlcs_fit_produces_coefs" { assert_true(coefs[1] > 0.0) } +///| test "glm_iwlcs_fit_converges" { let counts = [10.0, 10.0, 10.0, 10.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let initial = @src.glm_initial_estimates(counts, d) let coefs = @src.glm_iwlcs_fit(counts, d, 0.1, initial) assert_eq(coefs.length(), 1) @@ -756,6 +766,7 @@ test "glm_iwlcs_fit_converges" { // Standard error computation // --------------------------------------------------------------------------- +///| test "glm_compute_se_positive" { let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -770,12 +781,13 @@ test "glm_compute_se_positive" { assert_true(se[1] > 0.0) } +///| test "glm_compute_se_single_coef" { - let d = @src.GlmDesign::new( - ["s1", "s2", "s3"], - ["Intercept"], - [[1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + ]) let coefs = [2.5] let sf = [1.0, 1.0, 1.0] let se = @src.glm_compute_se(d, 0.2, coefs, sf) @@ -787,6 +799,7 @@ test "glm_compute_se_single_coef" { // Single-gene fitting // --------------------------------------------------------------------------- +///| test "glm_fit_one_gene_basic" { let counts = [10.0, 12.0, 50.0, 48.0] let d = @src.GlmDesign::new( @@ -804,6 +817,7 @@ test "glm_fit_one_gene_basic" { assert_true(fit.df_dispersion > 0.0) } +///| test "glm_fit_one_gene_treatment_positive" { let counts = [10.0, 12.0, 50.0, 48.0] let d = @src.GlmDesign::new( @@ -817,6 +831,7 @@ test "glm_fit_one_gene_treatment_positive" { assert_true(fit.coefficients[1] > 0.0) } +///| test "glm_fit_one_gene_no_treatment_effect" { let counts = [20.0, 22.0, 18.0, 20.0] let d = @src.GlmDesign::new( @@ -834,12 +849,10 @@ test "glm_fit_one_gene_no_treatment_effect" { // Differential expression testing // --------------------------------------------------------------------------- +///| test "glm_test_de_basic" { // 2 genes, 4 samples - let counts = [ - [10.0, 12.0, 50.0, 48.0], - [20.0, 22.0, 18.0, 20.0], - ] + let counts = [[10.0, 12.0, 50.0, 48.0], [20.0, 22.0, 18.0, 20.0]] let gene_names = ["GeneA", "GeneB"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -853,11 +866,9 @@ test "glm_test_de_basic" { assert_eq(results[1].gene, "GeneB") } +///| test "glm_test_de_pvalues_in_range" { - let counts = [ - [10.0, 12.0, 50.0, 48.0], - [20.0, 22.0, 18.0, 20.0], - ] + let counts = [[10.0, 12.0, 50.0, 48.0], [20.0, 22.0, 18.0, 20.0]] let gene_names = ["GeneA", "GeneB"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -874,12 +885,10 @@ test "glm_test_de_pvalues_in_range" { } } +///| test "glm_test_de_detects_differential" { // GeneA has strong treatment effect, GeneB does not - let counts = [ - [5.0, 6.0, 80.0, 75.0], - [20.0, 22.0, 18.0, 20.0], - ] + let counts = [[5.0, 6.0, 80.0, 75.0], [20.0, 22.0, 18.0, 20.0]] let gene_names = ["GeneA", "GeneB"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -894,11 +903,9 @@ test "glm_test_de_detects_differential" { assert_true(z_a > z_b) } +///| test "glm_test_de_estimate_sign" { - let counts = [ - [10.0, 12.0, 50.0, 48.0], - [50.0, 48.0, 10.0, 12.0], - ] + let counts = [[10.0, 12.0, 50.0, 48.0], [50.0, 48.0, 10.0, 12.0]] let gene_names = ["GeneUp", "GeneDown"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -913,6 +920,7 @@ test "glm_test_de_estimate_sign" { assert_true(results[1].estimate() < 0.0) } +///| test "glm_test_de_single_gene" { let counts = [[10.0, 12.0, 50.0, 48.0]] let gene_names = ["GeneA"] @@ -931,6 +939,7 @@ test "glm_test_de_single_gene" { // Significant filtering // --------------------------------------------------------------------------- +///| test "glm_significant_filters_correctly" { let r1 = @src.GlmTestResult::new("GeneA", "Treatment", 3.0, 0.5, 6.0, 1.0e-10) let r2 = @src.GlmTestResult::new("GeneB", "Treatment", 0.1, 0.3, 0.3, 0.8) @@ -947,11 +956,13 @@ test "glm_significant_filters_correctly" { } } +///| test "glm_significant_empty_input" { let sig = @src.glm_significant([], 0.05) assert_eq(sig.length(), 0) } +///| test "glm_significant_no_passing" { let r1 = @src.GlmTestResult::new("GeneA", "Treatment", 0.0, 0.5, 0.0, 0.5) let r2 = @src.GlmTestResult::new("GeneB", "Treatment", 0.1, 0.3, 0.3, 0.8) @@ -960,8 +971,11 @@ test "glm_significant_no_passing" { assert_eq(sig.length(), 0) } +///| test "glm_significant_all_pass" { - let r1 = @src.GlmTestResult::new("GeneA", "Treatment", 5.0, 0.5, 10.0, 1.0e-10) + let r1 = @src.GlmTestResult::new( + "GeneA", "Treatment", 5.0, 0.5, 10.0, 1.0e-10, + ) let r2 = @src.GlmTestResult::new("GeneB", "Treatment", 4.0, 0.3, 13.0, 1.0e-8) let results = [r1, r2] let sig = @src.glm_significant(results, 1.0) @@ -972,36 +986,35 @@ test "glm_significant_all_pass" { // Edge cases and integration // --------------------------------------------------------------------------- +///| test "glm_design_single_sample" { - let d = @src.GlmDesign::new( - ["s1"], - ["Intercept"], - [[1.0]], - ) + let d = @src.GlmDesign::new(["s1"], ["Intercept"], [[1.0]]) assert_eq(d.n_samples(), 1) assert_eq(d.n_coefs(), 1) assert_eq(d.get(0, 0), 1.0) } +///| test "glm_design_many_coefs" { - let d = @src.GlmDesign::new( - ["s1", "s2"], - ["A", "B", "C", "D"], - [[1.0, 0.0, 0.0, 0.0], [0.0, 1.0, 0.0, 0.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2"], ["A", "B", "C", "D"], [ + [1.0, 0.0, 0.0, 0.0], + [0.0, 1.0, 0.0, 0.0], + ]) assert_eq(d.n_samples(), 2) assert_eq(d.n_coefs(), 4) assert_eq(d.get(0, 0), 1.0) assert_eq(d.get(1, 1), 1.0) } +///| test "glm_fit_constant_counts" { let counts = [20.0, 20.0, 20.0, 20.0] - let d = @src.GlmDesign::new( - ["s1", "s2", "s3", "s4"], - ["Intercept"], - [[1.0], [1.0], [1.0], [1.0]], - ) + let d = @src.GlmDesign::new(["s1", "s2", "s3", "s4"], ["Intercept"], [ + [1.0], + [1.0], + [1.0], + [1.0], + ]) let sf = [1.0, 1.0, 1.0, 1.0] let fit = @src.glm_fit_one_gene(counts, d, sf, "GeneConst") assert_eq(fit.gene, "GeneConst") @@ -1009,11 +1022,9 @@ test "glm_fit_constant_counts" { assert_true(fit.dispersion >= 0.0) } +///| test "glm_test_de_with_size_factors" { - let counts = [ - [10.0, 20.0, 50.0, 100.0], - [20.0, 40.0, 18.0, 36.0], - ] + let counts = [[10.0, 20.0, 50.0, 100.0], [20.0, 40.0, 18.0, 36.0]] let gene_names = ["GeneA", "GeneB"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -1025,6 +1036,7 @@ test "glm_test_de_with_size_factors" { assert_eq(results.length(), 2) } +///| test "glm_full_pipeline" { // Simulate a full differential expression analysis pipeline let counts = [ @@ -1066,44 +1078,48 @@ test "glm_full_pipeline" { } } +///| test "glm_phi_05" { // phi(0.5) should be approximately 0.6915 let p = @src.glm_phi(0.5) assert_true((p - 0.6915).abs() < 0.01) } +///| test "glm_phi_2" { // phi(2.0) should be approximately 0.9772 let p = @src.glm_phi(2.0) assert_true((p - 0.9772).abs() < 0.01) } +///| test "glm_phi_negative" { // phi(-1.0) should be approximately 0.1587 let p = @src.glm_phi(-1.0) assert_true((p - 0.1587).abs() < 0.01) } +///| test "glm_median_two_elements" { let m = @src.glm_median([3.0, 7.0]) assert_eq(m, 5.0) } +///| test "glm_median_with_duplicates" { let m = @src.glm_median([1.0, 2.0, 2.0, 3.0, 3.0]) assert_eq(m, 2.0) } +///| test "glm_min_positive_all_zero" { let v = @src.glm_min_positive([0.0, 0.0, 0.0]) assert_eq(v, 0.0) } +///| test "glm_calculate_sf_two_genes" { - let counts = [ - [100.0, 10.0], - [200.0, 20.0], - ] + let counts = [[100.0, 10.0], [200.0, 20.0]] let sf = @src.glm_calculate_sf(counts) assert_eq(sf.length(), 2) // Sample 0 has 10x more counts than sample 1 @@ -1111,11 +1127,9 @@ test "glm_calculate_sf_two_genes" { assert_true((sf[0] / sf[1] - 10.0).abs() < 2.0) } +///| test "glm_test_de_all_zeros" { - let counts = [ - [0.0, 0.0, 0.0, 0.0], - [0.0, 0.0, 0.0, 0.0], - ] + let counts = [[0.0, 0.0, 0.0, 0.0], [0.0, 0.0, 0.0, 0.0]] let gene_names = ["GeneA", "GeneB"] let d = @src.GlmDesign::new( ["s1", "s2", "s3", "s4"], @@ -1132,6 +1146,7 @@ test "glm_test_de_all_zeros" { } } +///| test "glm_bh_correct_array_already_corrected" { // When all p-values are very small, BH should keep them small let pvals = [1.0e-10, 1.0e-8, 1.0e-6, 1.0e-4] @@ -1142,4 +1157,4 @@ test "glm_bh_correct_array_already_corrected" { } // The smallest p-value should remain small assert_true(adj[0] <= 1.0e-4) -} \ No newline at end of file +} diff --git a/test/moonbit/goa_test.mbt b/test/moonbit/goa_test.mbt index 72a7d575..7b6861e1 100644 --- a/test/moonbit/goa_test.mbt +++ b/test/moonbit/goa_test.mbt @@ -11,18 +11,21 @@ test "goa_aspect_from_string_F" { assert_eq(@src.GafAspect::description(a), "Molecular Function") } +///| test "goa_aspect_from_string_P" { let a = @src.GafAspect::from_string("P") assert_eq(@src.GafAspect::to_string(a), "P") assert_eq(@src.GafAspect::description(a), "Biological Process") } +///| test "goa_aspect_from_string_C" { let a = @src.GafAspect::from_string("C") assert_eq(@src.GafAspect::to_string(a), "C") assert_eq(@src.GafAspect::description(a), "Cellular Component") } +///| test "goa_aspect_from_string_unknown" { // Unknown inputs (e.g. "X", "", "foo") map to Unknown. let a1 = @src.GafAspect::from_string("X") @@ -36,6 +39,7 @@ test "goa_aspect_from_string_unknown" { assert_eq(@src.GafAspect::description(a3), "Unknown") } +///| test "goa_aspect_to_string_roundtrip" { // from_string -> to_string should round-trip for the three valid codes. assert_eq(@src.GafAspect::to_string(@src.GafAspect::from_string("F")), "F") @@ -47,6 +51,7 @@ test "goa_aspect_to_string_roundtrip" { // GafRecord creation and accessors // --------------------------------------------------------------------------- +///| test "goa_record_creation" { let r = @src.GafRecord::new( db="UniProtKB", @@ -86,6 +91,7 @@ test "goa_record_creation" { assert_eq(r.gene_product_form_id(), "") } +///| test "goa_record_creation_all_fields_populated" { // Record with every field populated, including the optional extension/form id. let r = @src.GafRecord::new( @@ -118,6 +124,7 @@ test "goa_record_creation_all_fields_populated" { // GafRecord utility methods // --------------------------------------------------------------------------- +///| test "goa_record_go_id_short" { let r = @src.GafRecord::new( db="UniProtKB", @@ -141,6 +148,7 @@ test "goa_record_go_id_short" { assert_eq(r.go_id_short(), "0003674") } +///| test "goa_record_go_id_short_no_prefix" { // A GO ID without "GO:" prefix should be returned unchanged. let r = @src.GafRecord::new( @@ -165,6 +173,7 @@ test "goa_record_go_id_short_no_prefix" { assert_eq(r.go_id_short(), "0003674") } +///| test "goa_record_qualifiers_split" { let r = @src.GafRecord::new( db="UniProtKB", @@ -191,6 +200,7 @@ test "goa_record_qualifiers_split" { assert_eq(qs[1], "enables") } +///| test "goa_record_qualifiers_single" { let r = @src.GafRecord::new( db="UniProtKB", @@ -216,6 +226,7 @@ test "goa_record_qualifiers_single" { assert_eq(qs[0], "enables") } +///| test "goa_record_qualifiers_empty" { let r = @src.GafRecord::new( db="UniProtKB", @@ -239,6 +250,7 @@ test "goa_record_qualifiers_empty" { assert_eq(r.qualifiers().length(), 0) } +///| test "goa_record_synonyms_split" { let r = @src.GafRecord::new( db="UniProtKB", @@ -265,6 +277,7 @@ test "goa_record_synonyms_split" { assert_eq(syns[1], "LFS1") } +///| test "goa_record_references_split" { let r = @src.GafRecord::new( db="UniProtKB", @@ -292,6 +305,7 @@ test "goa_record_references_split" { assert_eq(refs[2], "GO_REF:000001") } +///| test "goa_record_taxon_ids_with_prefix" { // "taxon:9606|taxon:9606" yields ["9606", "9606"] (duplicates preserved). let r = @src.GafRecord::new( @@ -319,6 +333,7 @@ test "goa_record_taxon_ids_with_prefix" { assert_eq(ids[1], "9606") } +///| test "goa_record_taxon_ids_bare_numeric" { // Numeric taxon IDs without the "taxon:" prefix should be returned as-is. let r = @src.GafRecord::new( @@ -345,6 +360,7 @@ test "goa_record_taxon_ids_bare_numeric" { assert_eq(ids[0], "9606") } +///| test "goa_record_taxon_ids_multiple_distinct" { // Two distinct taxon IDs (interactor case). let r = @src.GafRecord::new( @@ -372,6 +388,7 @@ test "goa_record_taxon_ids_multiple_distinct" { assert_eq(ids[1], "4932") } +///| test "goa_record_to_string" { let r = @src.GafRecord::new( db="UniProtKB", @@ -405,6 +422,7 @@ test "goa_record_to_string" { // Header line parsing // --------------------------------------------------------------------------- +///| test "goa_parse_header_line_valid" { let h = @src.goa_parse_header_line("!gaf-version: 2.2") assert_true(h is Some(_)) @@ -417,6 +435,7 @@ test "goa_parse_header_line_valid" { } } +///| test "goa_parse_header_line_generated_by" { let h = @src.goa_parse_header_line("!generated-by: UniProt") assert_true(h is Some(_)) @@ -429,6 +448,7 @@ test "goa_parse_header_line_generated_by" { } } +///| test "goa_parse_header_line_with_extra_spaces" { let h = @src.goa_parse_header_line("! gaf-version : 2.2 ") assert_true(h is Some(_)) @@ -441,23 +461,27 @@ test "goa_parse_header_line_with_extra_spaces" { } } +///| test "goa_parse_header_line_not_header" { // A line that does not start with "!" is not a header. let h = @src.goa_parse_header_line("UniProtKB\tQ12345") assert_true(h is None) } +///| test "goa_parse_header_line_no_colon" { // A header line with no colon is invalid. let h = @src.goa_parse_header_line("!this-has-no-colon") assert_true(h is None) } +///| test "goa_parse_header_line_empty" { let h = @src.goa_parse_header_line("") assert_true(h is None) } +///| test "goa_parse_header_line_empty_value" { // A header line with a colon but empty value is valid (value is ""). let h = @src.goa_parse_header_line("!gaf-version:") @@ -475,6 +499,7 @@ test "goa_parse_header_line_empty_value" { // Data line parsing // --------------------------------------------------------------------------- +///| test "goa_parse_line_full_17_columns" { let line = "UniProtKB\tQ12345\tPROT1\tenables\tGO:0003674\tPMID:12345\tIDA\tGO:0005515\tF\tProtein 1\tP1|PROT-1\tprotein\ttaxon:9606|taxon:9606\t20210115\tUniProt\t\t" let r = @src.goa_parse_line(line) @@ -503,6 +528,7 @@ test "goa_parse_line_full_17_columns" { } } +///| test "goa_parse_line_short_padded" { // A line with fewer than 17 columns should be right-padded with empty strings. let line = "UniProtKB\tQ12345\tPROT1\tenables\tGO:0003674" @@ -525,22 +551,26 @@ test "goa_parse_line_short_padded" { } } +///| test "goa_parse_line_empty" { let r = @src.goa_parse_line("") assert_true(r is None) } +///| test "goa_parse_line_whitespace_only" { let r = @src.goa_parse_line(" ") assert_true(r is None) } +///| test "goa_parse_line_header_returns_none" { // Header lines (starting with "!") should not be parsed as data records. let r = @src.goa_parse_line("!gaf-version: 2.2") assert_true(r is None) } +///| test "goa_parse_line_comment_returns_none" { // Any line starting with "!" is treated as a header/comment, not data. let r = @src.goa_parse_line("!some comment line") @@ -551,6 +581,7 @@ test "goa_parse_line_comment_returns_none" { // Full content parsing // --------------------------------------------------------------------------- +///| test "goa_parse_with_header_and_data" { let content = "!gaf-version: 2.2\n" + "!generated-by: UniProt\n" + @@ -565,6 +596,7 @@ test "goa_parse_with_header_and_data" { assert_eq(recs[1].db_object_symbol(), "TP53") } +///| test "goa_parse_empty_input" { let db = @src.goa_parse("") assert_eq(db.n_records(), 0) @@ -572,6 +604,7 @@ test "goa_parse_empty_input" { assert_eq(db.created_by(), "") } +///| test "goa_parse_only_headers" { let content = "!gaf-version: 2.2\n" + "!generated-by: SGD\n" let db = @src.goa_parse(content) @@ -580,6 +613,7 @@ test "goa_parse_only_headers" { assert_eq(db.created_by(), "SGD") } +///| test "goa_parse_skips_blank_lines" { let content = "!gaf-version: 2.2\n" + "\n" + @@ -590,6 +624,7 @@ test "goa_parse_skips_blank_lines" { assert_eq(db.n_records(), 2) } +///| test "goa_parse_ignores_unknown_header_keys" { // Header keys other than gaf-version / generated-by should not overwrite metadata. let content = "!gaf-version: 2.2\n" + @@ -606,6 +641,7 @@ test "goa_parse_ignores_unknown_header_keys" { // GoaDatabase properties // --------------------------------------------------------------------------- +///| test "goa_database_new_empty" { let db = @src.GoaDatabase::new() assert_eq(db.n_records(), 0) @@ -614,6 +650,7 @@ test "goa_database_new_empty" { assert_eq(db.created_by(), "") } +///| test "goa_database_records_accessor" { let db = @src.goa_sample_database() let recs = db.records() @@ -624,6 +661,7 @@ test "goa_database_records_accessor" { // Filter functions // --------------------------------------------------------------------------- +///| test "goa_filter_by_go_id" { let db = @src.goa_sample_database() // GO:0003674 (molecular_function) appears once in the sample database. @@ -633,12 +671,14 @@ test "goa_filter_by_go_id" { assert_eq(hits[0].go_id(), "GO:0003674") } +///| test "goa_filter_by_go_id_no_match" { let db = @src.goa_sample_database() let hits = @src.goa_filter_by_go_id(db, "GO:9999999") assert_eq(hits.length(), 0) } +///| test "goa_filter_by_aspect_molecular_function" { let db = @src.goa_sample_database() // F records: PROT1, TP53, YFG1, RPL5 (4 total). @@ -649,6 +689,7 @@ test "goa_filter_by_aspect_molecular_function" { } } +///| test "goa_filter_by_aspect_biological_process" { let db = @src.goa_sample_database() // P records: PROT1, TP53, Hsp70 (3 total). @@ -659,6 +700,7 @@ test "goa_filter_by_aspect_biological_process" { } } +///| test "goa_filter_by_aspect_cellular_component" { let db = @src.goa_sample_database() // C records: PROT1, YFG1, Hsp70 (3 total). @@ -669,6 +711,7 @@ test "goa_filter_by_aspect_cellular_component" { } } +///| test "goa_filter_by_aspect_unknown" { let db = @src.goa_sample_database() // No Unknown aspect records in the sample database. @@ -676,6 +719,7 @@ test "goa_filter_by_aspect_unknown" { assert_eq(hits.length(), 0) } +///| test "goa_filter_by_evidence_ida" { let db = @src.goa_sample_database() // IDA records: PROT1 (x3), YFG1 (C), Hsp70 (C) = 5 total. @@ -686,6 +730,7 @@ test "goa_filter_by_evidence_ida" { } } +///| test "goa_filter_by_evidence_iea" { let db = @src.goa_sample_database() // IEA records: TP53 (P), Hsp70 (P) = 2 total. @@ -693,12 +738,14 @@ test "goa_filter_by_evidence_iea" { assert_eq(hits.length(), 2) } +///| test "goa_filter_by_evidence_no_match" { let db = @src.goa_sample_database() let hits = @src.goa_filter_by_evidence(db, "NONEXISTENT") assert_eq(hits.length(), 0) } +///| test "goa_filter_by_taxon_human" { let db = @src.goa_sample_database() // taxon:9606 records: PROT1 (x3), TP53 (x2), RPL5 = 6 total. @@ -706,6 +753,7 @@ test "goa_filter_by_taxon_human" { assert_eq(hits.length(), 6) } +///| test "goa_filter_by_taxon_yeast" { let db = @src.goa_sample_database() // taxon:4932 records: YFG1 (x2) = 2 total. @@ -713,6 +761,7 @@ test "goa_filter_by_taxon_yeast" { assert_eq(hits.length(), 2) } +///| test "goa_filter_by_taxon_fly" { let db = @src.goa_sample_database() // taxon:7227 records: Hsp70 (x2) = 2 total. @@ -720,12 +769,14 @@ test "goa_filter_by_taxon_fly" { assert_eq(hits.length(), 2) } +///| test "goa_filter_by_taxon_no_match" { let db = @src.goa_sample_database() let hits = @src.goa_filter_by_taxon(db, "0000000") assert_eq(hits.length(), 0) } +///| test "goa_filter_by_db_object_id" { let db = @src.goa_sample_database() // Q12345 (PROT1) has 3 records (F, P, C). @@ -737,6 +788,7 @@ test "goa_filter_by_db_object_id" { } } +///| test "goa_filter_by_db_object_id_multiple" { let db = @src.goa_sample_database() // P04637 (TP53) has 2 records (F, P). @@ -747,6 +799,7 @@ test "goa_filter_by_db_object_id_multiple" { } } +///| test "goa_filter_by_db_object_id_no_match" { let db = @src.goa_sample_database() let hits = @src.goa_filter_by_db_object_id(db, "ZZZZZZ") @@ -757,6 +810,7 @@ test "goa_filter_by_db_object_id_no_match" { // Unique value extraction // --------------------------------------------------------------------------- +///| test "goa_unique_go_ids" { let db = @src.goa_sample_database() // Every record in the sample database has a distinct GO ID. @@ -766,6 +820,7 @@ test "goa_unique_go_ids" { assert_eq(ids[0], "GO:0003674") } +///| test "goa_unique_go_ids_dedup" { // Build a database with duplicate GO IDs to test deduplication. let rec = @src.GafRecord::new( @@ -787,14 +842,17 @@ test "goa_unique_go_ids_dedup" { annotation_extension="", gene_product_form_id="", ) - let content = @src.goa_record_to_gaf_line(rec) + "\n" + - @src.goa_record_to_gaf_line(rec) + "\n" + let content = @src.goa_record_to_gaf_line(rec) + + "\n" + + @src.goa_record_to_gaf_line(rec) + + "\n" let db = @src.goa_parse(content) let ids = @src.goa_unique_go_ids(db) assert_eq(ids.length(), 1) assert_eq(ids[0], "GO:0003674") } +///| test "goa_unique_evidence_codes" { let db = @src.goa_sample_database() // Sample DB uses IDA, EXP, IEA, ISS, IPI = 5 distinct evidence codes. @@ -802,6 +860,7 @@ test "goa_unique_evidence_codes" { assert_eq(codes.length(), 5) } +///| test "goa_unique_taxon_ids" { let db = @src.goa_sample_database() // Sample DB uses 9606, 4932, 7227 = 3 distinct taxon IDs. @@ -816,6 +875,7 @@ test "goa_unique_taxon_ids" { // Count functions // --------------------------------------------------------------------------- +///| test "goa_count_by_aspect" { let db = @src.goa_sample_database() let counts = @src.goa_count_by_aspect(db) @@ -827,6 +887,7 @@ test "goa_count_by_aspect" { assert_eq(counts.get("?"), None) } +///| test "goa_count_by_aspect_empty_db" { let db = @src.GoaDatabase::new() let counts = @src.goa_count_by_aspect(db) @@ -835,6 +896,7 @@ test "goa_count_by_aspect_empty_db" { assert_eq(counts.get("C"), None) } +///| test "goa_count_by_evidence" { let db = @src.goa_sample_database() let counts = @src.goa_count_by_evidence(db) @@ -848,6 +910,7 @@ test "goa_count_by_evidence" { assert_eq(counts.get("TAS"), None) } +///| test "goa_count_by_evidence_empty_db" { let db = @src.GoaDatabase::new() let counts = @src.goa_count_by_evidence(db) @@ -858,6 +921,7 @@ test "goa_count_by_evidence_empty_db" { // Summary generation // --------------------------------------------------------------------------- +///| test "goa_to_summary_has_basic_fields" { let db = @src.goa_sample_database() let s = @src.goa_to_summary(db) @@ -867,6 +931,7 @@ test "goa_to_summary_has_basic_fields" { assert_true(s.contains("Created by: UniProt")) } +///| test "goa_to_summary_has_aspect_section" { let db = @src.goa_sample_database() let s = @src.goa_to_summary(db) @@ -876,6 +941,7 @@ test "goa_to_summary_has_aspect_section" { assert_true(s.contains("Cellular Component")) } +///| test "goa_to_summary_has_evidence_section" { let db = @src.goa_sample_database() let s = @src.goa_to_summary(db) @@ -887,6 +953,7 @@ test "goa_to_summary_has_evidence_section" { assert_true(s.contains("IPI")) } +///| test "goa_to_summary_has_unique_counts" { let db = @src.goa_sample_database() let s = @src.goa_to_summary(db) @@ -895,6 +962,7 @@ test "goa_to_summary_has_unique_counts" { assert_true(s.contains("Unique taxon IDs: 3")) } +///| test "goa_to_summary_empty_db" { let db = @src.GoaDatabase::new() let s = @src.goa_to_summary(db) @@ -906,6 +974,7 @@ test "goa_to_summary_empty_db" { // Writing functions // --------------------------------------------------------------------------- +///| test "goa_record_to_gaf_line_roundtrip" { let r = @src.GafRecord::new( db="UniProtKB", @@ -949,6 +1018,7 @@ test "goa_record_to_gaf_line_roundtrip" { assert_eq(fields[16].to_string(), "") } +///| test "goa_record_to_gaf_line_parse_roundtrip" { // Serializing and re-parsing should produce an equivalent record. let r = @src.GafRecord::new( @@ -989,6 +1059,7 @@ test "goa_record_to_gaf_line_parse_roundtrip" { } } +///| test "goa_to_gaf_includes_header" { let db = @src.goa_sample_database() let text = @src.goa_to_gaf(db) @@ -996,6 +1067,7 @@ test "goa_to_gaf_includes_header" { assert_true(text.contains("!generated-by: UniProt")) } +///| test "goa_to_gaf_includes_all_records" { let db = @src.goa_sample_database() let text = @src.goa_to_gaf(db) @@ -1012,6 +1084,7 @@ test "goa_to_gaf_includes_all_records" { assert_true(text.contains("GO:0003735")) } +///| test "goa_to_gaf_roundtrip" { // Serializing the sample database and re-parsing should preserve its // record count and metadata. @@ -1023,6 +1096,7 @@ test "goa_to_gaf_roundtrip" { assert_eq(db2.created_by(), db.created_by()) } +///| test "goa_to_gaf_empty_db" { let db = @src.GoaDatabase::new() let text = @src.goa_to_gaf(db) @@ -1036,17 +1110,20 @@ test "goa_to_gaf_empty_db" { // Sample database validation // --------------------------------------------------------------------------- +///| test "goa_sample_database_has_10_records" { let db = @src.goa_sample_database() assert_eq(db.n_records(), 10) } +///| test "goa_sample_database_metadata" { let db = @src.goa_sample_database() assert_eq(db.version(), "2.2") assert_eq(db.created_by(), "UniProt") } +///| test "goa_sample_database_has_3_aspects" { let db = @src.goa_sample_database() let counts = @src.goa_count_by_aspect(db) @@ -1055,6 +1132,7 @@ test "goa_sample_database_has_3_aspects" { assert_eq(counts.get("C"), Some(3)) } +///| test "goa_sample_database_has_3_taxa" { let db = @src.goa_sample_database() let taxa = @src.goa_unique_taxon_ids(db) @@ -1064,6 +1142,7 @@ test "goa_sample_database_has_3_taxa" { assert_true(taxa.contains("7227")) } +///| test "goa_sample_database_has_5_evidence_codes" { let db = @src.goa_sample_database() let codes = @src.goa_unique_evidence_codes(db) @@ -1075,6 +1154,7 @@ test "goa_sample_database_has_5_evidence_codes" { assert_true(codes.contains("IPI")) } +///| test "goa_sample_database_first_record" { let db = @src.goa_sample_database() let rec = db.records()[0] @@ -1087,6 +1167,7 @@ test "goa_sample_database_first_record" { assert_eq(rec.taxon(), "taxon:9606|taxon:9606") } +///| test "goa_sample_database_last_record" { let db = @src.goa_sample_database() let rec = db.records()[9] @@ -1097,6 +1178,7 @@ test "goa_sample_database_last_record" { assert_eq(rec.aspect(), "F") } +///| test "goa_sample_database_all_uniprot_db" { let db = @src.goa_sample_database() for r in db.records() { @@ -1104,6 +1186,7 @@ test "goa_sample_database_all_uniprot_db" { } } +///| test "goa_sample_database_unique_go_ids_count" { let db = @src.goa_sample_database() // Each of the 10 records has a distinct GO ID. @@ -1115,6 +1198,7 @@ test "goa_sample_database_unique_go_ids_count" { // Edge cases // --------------------------------------------------------------------------- +///| test "goa_parse_line_single_column" { // A single-column line should be padded to 17 empty fields and parsed. let r = @src.goa_parse_line("UniProtKB") @@ -1130,6 +1214,7 @@ test "goa_parse_line_single_column" { } } +///| test "goa_parse_line_extra_columns_kept_in_field" { // A line with more than 17 columns: the parser splits on all tabs and uses // only the first 17 fields; extra columns are silently ignored. @@ -1146,6 +1231,7 @@ test "goa_parse_line_extra_columns_kept_in_field" { } } +///| test "goa_parse_line_with_trailing_newline" { // Trailing whitespace/newlines are trimmed before parsing. let line = "UniProtKB\tQ12345\tPROT1\tenables\tGO:0003674\tPMID:12345\tIDA\tGO:0005515\tF\tProtein 1\tP1\tprotein\ttaxon:9606\t20210115\tUniProt\t\t\n" @@ -1160,6 +1246,7 @@ test "goa_parse_line_with_trailing_newline" { } } +///| test "goa_parse_only_comment_lines" { // A line starting with "!" is treated as a header/comment regardless of // whether it has a colon, and never produces a data record. @@ -1172,6 +1259,7 @@ test "goa_parse_only_comment_lines" { assert_eq(db.created_by(), "UniProt") } +///| test "goa_parse_mixed_valid_and_invalid_lines" { let content = "!gaf-version: 2.2\n" + "!generated-by: UniProt\n" + @@ -1188,15 +1276,20 @@ test "goa_parse_mixed_valid_and_invalid_lines" { assert_eq(db.created_by(), "UniProt") } +///| test "goa_filter_on_empty_database" { let db = @src.GoaDatabase::new() assert_eq(@src.goa_filter_by_go_id(db, "GO:0003674").length(), 0) - assert_eq(@src.goa_filter_by_aspect(db, @src.GafAspect::from_string("F")).length(), 0) + assert_eq( + @src.goa_filter_by_aspect(db, @src.GafAspect::from_string("F")).length(), + 0, + ) assert_eq(@src.goa_filter_by_evidence(db, "IDA").length(), 0) assert_eq(@src.goa_filter_by_taxon(db, "9606").length(), 0) assert_eq(@src.goa_filter_by_db_object_id(db, "Q12345").length(), 0) } +///| test "goa_unique_on_empty_database" { let db = @src.GoaDatabase::new() assert_eq(@src.goa_unique_go_ids(db).length(), 0) diff --git a/test/moonbit/gosemsim_test.mbt b/test/moonbit/gosemsim_test.mbt index 3487d49a..92b611d5 100644 --- a/test/moonbit/gosemsim_test.mbt +++ b/test/moonbit/gosemsim_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for GOSemSim GO semantic similarity module. - test "gosemsim_example_graph" { let g = @src.gosemsim_example_graph() assert_true(g.root_id == "GO:0008150") @@ -10,81 +9,152 @@ test "gosemsim_example_graph" { } } +///| test "gosemsim_term_self_similarity" { let g = @src.gosemsim_example_graph() // A term compared to itself should have maximal similarity. - let resnik = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0006810", @src.resnik_measure()) + let resnik = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0006810", + @src.resnik_measure(), + ) assert_true(resnik > 0.0) - let lin = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0006810", @src.lin_measure()) + let lin = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0006810", + @src.lin_measure(), + ) assert_true(lin > 0.99 && lin <= 1.0 + 1.0e-6) } +///| test "gosemsim_resnik_similarity" { let g = @src.gosemsim_example_graph() // Parent-child should have high similarity - let sim = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0009987", @src.resnik_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0009987", + @src.resnik_measure(), + ) assert_true(sim > 0.0) } +///| test "gosemsim_lin_similarity" { let g = @src.gosemsim_example_graph() - let sim = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0007165", @src.lin_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0007165", + @src.lin_measure(), + ) assert_true(sim >= 0.0 && sim <= 1.0) } +///| test "gosemsim_rel_similarity" { let g = @src.gosemsim_example_graph() - let sim = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0007165", @src.rel_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0007165", + @src.rel_measure(), + ) assert_true(sim >= 0.0 && sim <= 1.0) } +///| test "gosemsim_jiang_similarity" { let g = @src.gosemsim_example_graph() - let sim = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0007165", @src.jiang_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0007165", + @src.jiang_measure(), + ) assert_true(sim >= 0.0) } +///| test "gosemsim_wang_similarity" { let g = @src.gosemsim_example_graph() - let sim = @src.gosemsim_term_sim(g, "GO:0006810", "GO:0007165", @src.wang_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:0006810", + "GO:0007165", + @src.wang_measure(), + ) assert_true(sim >= 0.0 && sim <= 1.0) } +///| test "gosemsim_missing_term" { let g = @src.gosemsim_example_graph() - let sim = @src.gosemsim_term_sim(g, "GO:9999999", "GO:0006810", @src.resnik_measure()) + let sim = @src.gosemsim_term_sim( + g, + "GO:9999999", + "GO:0006810", + @src.resnik_measure(), + ) assert_true(sim == 0.0) } +///| test "gosemsim_gen_sim_max" { let g = @src.gosemsim_example_graph() let go1 = ["GO:0006810", "GO:0007165"] let go2 = ["GO:0006810", "GO:0009987"] - let sim = @src.gosemsim_gen_sim(g, go1, go2, @src.resnik_measure(), combine = "max") + let sim = @src.gosemsim_gen_sim( + g, + go1, + go2, + @src.resnik_measure(), + combine="max", + ) assert_true(sim > 0.0) } +///| test "gosemsim_gen_sim_avg" { let g = @src.gosemsim_example_graph() let go1 = ["GO:0006810", "GO:0007165"] let go2 = ["GO:0006810", "GO:0009987"] - let sim = @src.gosemsim_gen_sim(g, go1, go2, @src.lin_measure(), combine = "avg") + let sim = @src.gosemsim_gen_sim( + g, + go1, + go2, + @src.lin_measure(), + combine="avg", + ) assert_true(sim >= 0.0 && sim <= 1.0) } +///| test "gosemsim_gen_sim_rcmax" { let g = @src.gosemsim_example_graph() let go1 = ["GO:0006810", "GO:0007165"] let go2 = ["GO:0006810", "GO:0009987"] - let sim = @src.gosemsim_gen_sim(g, go1, go2, @src.rel_measure(), combine = "rcmax") + let sim = @src.gosemsim_gen_sim( + g, + go1, + go2, + @src.rel_measure(), + combine="rcmax", + ) assert_true(sim >= 0.0 && sim <= 1.0) } +///| test "gosemsim_graph_operations" { let g = @src.GOGraph::new() let g1 = g.add_root("TEST:001") - let g2 = g1.add_term(@src.GOTermNode::new("TEST:001", ic = 0.0)) - let g3 = g2.add_term(@src.GOTermNode::new("TEST:002", ic = 0.5, parents = ["TEST:001"])) + let g2 = g1.add_term(@src.GOTermNode::new("TEST:001", ic=0.0)) + let g3 = g2.add_term( + @src.GOTermNode::new("TEST:002", ic=0.5, parents=["TEST:001"]), + ) match g3.get_term("TEST:002") { Some(term) => assert_true(term.ic == 0.5) None => assert_true(false) diff --git a/test/moonbit/granges_list_test.mbt b/test/moonbit/granges_list_test.mbt new file mode 100644 index 00000000..ea6be3ca --- /dev/null +++ b/test/moonbit/granges_list_test.mbt @@ -0,0 +1,681 @@ +///| +fn grl_test_object() -> @src.GRangesList { + @src.granges_list( + [ + @src.granges(["chr1", "chr1"], [(100, 109), (200, 209)], [ + @src.strand_plus(), + @src.strand_plus(), + ]), + @src.granges(["chr1", "chr1"], [(150, 159), (300, 309)], [ + @src.strand_minus(), + @src.strand_minus(), + ]), + @src.granges_single("chr2", 50, 59, @src.strand_star()), + @src.granges([], [], []), + ], + names=["txA", "txB", "txC", "empty"], + element_metadata=[ + Map([("gene", "geneA")]), + Map([("gene", "geneB")]), + Map([("gene", "geneC")]), + Map([("gene", "none")]), + ], + ) catch { + _ => abort("failed to construct GRangesList fixture") + } +} + +///| +fn grl_test_subject() -> @src.GRanges { + @src.granges( + ["chr1", "chr1", "chr1", "chr2", "chr3"], + [(105, 205), (155, 156), (205, 206), (55, 55), (1, 10)], + [ + @src.strand_plus(), + @src.strand_minus(), + @src.strand_plus(), + @src.strand_star(), + @src.strand_plus(), + ], + ) +} + +///| +fn grl_test_experiment() -> @src.RangedSummarizedExperiment { + @src.RangedSummarizedExperiment::new_with_range_groups( + assays=Map([ + ("counts", [[10.0, 11.0], [20.0, 21.0], [30.0, 31.0], [40.0, 41.0]]), + ]), + row_range_groups=grl_test_object(), + col_data=[Map([("sample", "S1")]), Map([("sample", "S2")])], + row_names=["txA", "txB", "txC", "empty"], + metadata=Map([("organism", "human")]), + ) catch { + _ => abort("failed to construct grouped RangedSummarizedExperiment") + } +} + +///| +test "granges_list: construct and access" { + let ranges = grl_test_object() + assert_true(ranges.is_valid()) + assert_eq(ranges.length(), 4) + assert_eq(ranges.names(), ["txA", "txB", "txC", "empty"]) + assert_eq(ranges.element_lengths(), [2, 2, 1, 0]) + assert_eq(ranges.element_metadata()[1]["gene"], "geneB") + assert_false(ranges.is_empty()) + assert_eq(ranges.summary(), "GRangesList(4 elements, 5 ranges)") +} + +///| +test "granges_list: get by index and name returns ranges" { + let ranges = grl_test_object() + match ranges.get(1) { + Some(element) => { + assert_eq(element.starts, [150, 300]) + assert_true(element.strands == [@src.strand_minus(), @src.strand_minus()]) + } + None => assert_true(false) + } + match ranges.get_by_name("txC") { + Some(element) => assert_eq(element.seqnames, ["chr2"]) + None => assert_true(false) + } + assert_true(ranges.get(-1) is None) + assert_true(ranges.get_by_name("missing") is None) +} + +///| +test "granges_list: empty semantics include all-empty elements" { + let no_elements = @src.granges_list([]) catch { + _ => abort("empty list should be valid") + } + let empty_elements = @src.granges_list( + [@src.granges([], [], []), @src.granges([], [], [])], + names=["a", "b"], + ) catch { + _ => abort("all-empty list should be valid") + } + assert_true(no_elements.is_empty()) + assert_true(empty_elements.is_empty()) +} + +///| +test "granges_list: rejects outer dimension mismatch" { + let names_raised = try { + ignore( + @src.granges_list( + [@src.granges_single("chr1", 1, 2, @src.strand_plus())], + names=["a", "b"], + ), + ) + false + } catch { + GRangesListError(_) => true + } + let metadata_raised = try { + ignore( + @src.granges_list( + [@src.granges_single("chr1", 1, 2, @src.strand_plus())], + element_metadata=[Map([]), Map([])], + ), + ) + false + } catch { + GRangesListError(_) => true + } + assert_true(names_raised) + assert_true(metadata_raised) +} + +///| +test "granges_list: rejects malformed inner GRanges" { + let raised = try { + ignore( + @src.granges_list([ + @src.granges(["chr1", "chr2"], [(1, 10), (20, 30)], [@src.strand_plus()]), + ]), + ) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: split uses first-appearance group order" { + let flat = @src.granges( + ["chr1", "chr2", "chr1", "chr2", "chr1"], + [(1, 5), (10, 15), (20, 25), (30, 35), (40, 45)], + [ + @src.strand_plus(), + @src.strand_minus(), + @src.strand_plus(), + @src.strand_minus(), + @src.strand_plus(), + ], + ) + let split = @src.granges_split_as_list(flat, [ + "txB", "txA", "txB", "txA", "txC", + ]) + assert_eq(split.names(), ["txB", "txA", "txC"]) + assert_eq(split.element_lengths(), [2, 2, 1]) + match split.get_by_name("txB") { + Some(element) => assert_eq(element.starts, [1, 20]) + None => assert_true(false) + } +} + +///| +test "granges_list: split rejects group length mismatch" { + let raised = try { + ignore( + @src.granges_split_as_list( + @src.granges_single("chr1", 1, 10, @src.strand_plus()), + [], + ), + ) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: partition supports empty features" { + let flat = @src.granges( + ["chr1", "chr1", "chr2"], + [(1, 5), (10, 15), (20, 25)], + [@src.strand_plus(), @src.strand_plus(), @src.strand_minus()], + ) + let partitioned = @src.granges_list_from_partition(flat, [2, 0, 1], names=[ + "tx1", "empty", "tx2", + ]) + assert_eq(partitioned.element_lengths(), [2, 0, 1]) + match partitioned.get(1) { + Some(element) => assert_eq(@src.granges_length(element), 0) + None => assert_true(false) + } +} + +///| +test "granges_list: partition validates lengths" { + let flat = @src.granges_single("chr1", 1, 10, @src.strand_plus()) + let sum_raised = try { + ignore(@src.granges_list_from_partition(flat, [2])) + false + } catch { + GRangesListError(_) => true + } + let negative_raised = try { + ignore(@src.granges_list_from_partition(flat, [-1, 2])) + false + } catch { + GRangesListError(_) => true + } + assert_true(sum_raised) + assert_true(negative_raised) +} + +///| +test "granges_list: unlist preserves member order" { + let flat = grl_test_object().unlist() + assert_eq(flat.seqnames, ["chr1", "chr1", "chr1", "chr1", "chr2"]) + assert_eq(flat.starts, [100, 200, 150, 300, 50]) + assert_eq(flat.ends, [109, 209, 159, 309, 59]) +} + +///| +test "granges_list: unlist and relist round trip" { + let original = grl_test_object() + let rebuilt = original.relist(original.unlist()) + assert_eq(rebuilt.names(), original.names()) + assert_eq(rebuilt.element_lengths(), original.element_lengths()) + assert_eq(rebuilt.element_metadata()[2]["gene"], "geneC") + for index = 0; index < original.length(); index = index + 1 { + match (original.get(index), rebuilt.get(index)) { + (Some(left), Some(right)) => { + assert_eq(left.seqnames, right.seqnames) + assert_eq(left.starts, right.starts) + assert_eq(left.ends, right.ends) + assert_true(left.strands == right.strands) + } + _ => assert_true(false) + } + } +} + +///| +test "granges_list: relist rejects incompatible flat length" { + let raised = try { + ignore( + grl_test_object().relist( + @src.granges_single("chr1", 1, 10, @src.strand_plus()), + ), + ) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: subset coordinates names and metadata" { + let subset = grl_test_object().subset([2, 0, 2]) + assert_eq(subset.names(), ["txC", "txA", "txC"]) + assert_eq(subset.element_lengths(), [1, 2, 1]) + assert_eq(subset.element_metadata()[0]["gene"], "geneC") + assert_eq(subset.element_metadata()[1]["gene"], "geneA") +} + +///| +test "granges_list: invalid subset index raises" { + let raised = try { + ignore(grl_test_object().subset([4])) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: feature bounds summarize compound features" { + let bounds = grl_test_object().feature_bounds() + assert_eq(bounds.seqnames, ["chr1", "chr1", "chr2", ""]) + assert_eq(bounds.starts, [100, 150, 50, 0]) + assert_eq(bounds.ends, [209, 309, 59, -1]) + assert_eq(bounds.widths, [110, 160, 10, 0]) + assert_true(bounds.strands[0] == @src.strand_plus()) + assert_true(bounds.strands[1] == @src.strand_minus()) +} + +///| +test "granges_list: mixed feature bounds use wildcard identifiers" { + let mixed = @src.granges_list([ + @src.granges(["chr1", "chr2"], [(10, 20), (30, 40)], [ + @src.strand_plus(), + @src.strand_minus(), + ]), + ]) catch { + _ => abort("mixed feature should be valid") + } + let bounds = mixed.feature_bounds() + assert_eq(bounds.seqnames, ["*"]) + assert_true(bounds.strands == [@src.strand_star()]) + assert_eq(bounds.starts, [10]) + assert_eq(bounds.ends, [40]) +} + +///| +test "granges_list: concat appends outer elements" { + let left = grl_test_object().subset([0]) + let right = grl_test_object().subset([1, 2]) + let combined = left.concat(right) + assert_eq(combined.names(), ["txA", "txB", "txC"]) + assert_eq(combined.element_lengths(), [2, 2, 1]) + assert_eq(combined.element_metadata()[2]["gene"], "geneC") +} + +///| +test "granges_list: parallel concat combines corresponding features" { + let shifted = grl_test_object().shift(1000) + let combined = grl_test_object().parallel_concat(shifted) + assert_eq(combined.element_lengths(), [4, 4, 2, 0]) + match combined.get(0) { + Some(element) => assert_eq(element.starts, [100, 200, 1100, 1200]) + None => assert_true(false) + } +} + +///| +test "granges_list: parallel operations require equal element counts" { + let raised = try { + ignore(grl_test_object().parallel_concat(grl_test_object().subset([0, 1]))) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: intra-range operations preserve groups" { + let shifted = grl_test_object().shift(10) + let narrowed = grl_test_object().narrow(2, 5) + let resized = grl_test_object().resize(5, "start") + assert_eq(shifted.element_lengths(), [2, 2, 1, 0]) + match shifted.get(0) { + Some(element) => assert_eq(element.starts, [110, 210]) + None => assert_true(false) + } + match narrowed.get(0) { + Some(element) => { + assert_eq(element.starts, [101, 201]) + assert_eq(element.ends, [104, 204]) + } + None => assert_true(false) + } + match resized.get(1) { + Some(element) => assert_eq(element.ends, [154, 304]) + None => assert_true(false) + } +} + +///| +test "granges_list: flank and promoters preserve outer metadata" { + let flanked = grl_test_object().flank(5, true, false) + let promoted = grl_test_object().promoters(upstream=10, downstream=3) + assert_eq(flanked.names(), grl_test_object().names()) + assert_eq(promoted.element_metadata()[0]["gene"], "geneA") + match flanked.get(0) { + Some(element) => assert_eq(element.starts, [95, 195]) + None => assert_true(false) + } + match promoted.get(1) { + Some(element) => { + assert_eq(element.starts, [157, 307]) + assert_eq(element.ends, [169, 319]) + } + None => assert_true(false) + } +} + +///| +test "granges_list: promoters reject negative widths" { + let raised = try { + ignore(grl_test_object().promoters(upstream=-1)) + false + } catch { + GRangesListError(_) => true + } + assert_true(raised) +} + +///| +test "granges_list: reduce disjoin and sort operate per feature" { + let ranges = @src.granges_list( + [ + @src.granges(["chr1", "chr1", "chr1"], [(20, 30), (1, 10), (8, 15)], [ + @src.strand_plus(), + @src.strand_plus(), + @src.strand_plus(), + ]), + ], + names=["tx"], + ) catch { + _ => abort("failed to construct interval fixture") + } + match ranges.sort_ranges().get(0) { + Some(element) => assert_eq(element.starts, [1, 8, 20]) + None => assert_true(false) + } + match ranges.reduce().get(0) { + Some(element) => { + assert_eq(element.starts, [1, 20]) + assert_eq(element.ends, [15, 30]) + } + None => assert_true(false) + } + match ranges.disjoin().get(0) { + Some(element) => assert_eq(element.starts, [1, 8, 11, 16, 20]) + None => assert_true(false) + } +} + +///| +test "granges_list: width summaries exclude introns" { + assert_eq(grl_test_object().total_widths(), [20, 20, 10, 0]) +} + +///| +test "granges_list: parallel set operations preserve grouping" { + let left = @src.granges_list([ + @src.granges_single("chr1", 1, 10, @src.strand_star()), + @src.granges_single("chr2", 20, 30, @src.strand_star()), + ]) catch { + _ => abort("failed to build left set") + } + let right = @src.granges_list([ + @src.granges_single("chr1", 5, 15, @src.strand_star()), + @src.granges_single("chr2", 25, 35, @src.strand_star()), + ]) catch { + _ => abort("failed to build right set") + } + match left.parallel_union(right).get(0) { + Some(element) => { + assert_eq(element.starts, [1]) + assert_eq(element.ends, [15]) + } + None => assert_true(false) + } + match left.parallel_intersect(right).get(1) { + Some(element) => { + assert_eq(element.starts, [25]) + assert_eq(element.ends, [30]) + } + None => assert_true(false) + } + match left.parallel_setdiff(right).get(0) { + Some(element) => { + assert_eq(element.starts, [1]) + assert_eq(element.ends, [4]) + } + None => assert_true(false) + } +} + +///| +test "granges_list: feature overlaps are deduplicated and strand-aware" { + let ranges = grl_test_object() + let subject = grl_test_subject() + assert_eq(ranges.find_overlaps(subject), [(0, 0), (0, 2), (1, 1), (2, 3)]) + assert_eq(ranges.count_overlaps(subject), [2, 1, 1, 0]) + assert_eq(ranges.overlaps_any(subject), [true, true, true, false]) + assert_eq(ranges.find_overlaps(subject, ignore_strand=true), [ + (0, 0), + (0, 2), + (1, 0), + (1, 1), + (2, 3), + ]) +} + +///| +test "granges_list: subset by overlap keeps each feature once" { + let subset = grl_test_object().subset_by_overlaps(grl_test_subject()) + assert_eq(subset.names(), ["txA", "txB", "txC"]) + assert_eq(subset.element_lengths(), [2, 2, 1]) +} + +///| +test "granges_list: grouped subjects use feature indices" { + let subject = @src.granges_list( + [ + @src.granges(["chr1", "chr1"], [(105, 106), (205, 206)], [ + @src.strand_plus(), + @src.strand_plus(), + ]), + @src.granges_single("chr2", 55, 56, @src.strand_plus()), + ], + names=["feature1", "feature2"], + ) catch { + _ => abort("failed to construct grouped subject") + } + let query = grl_test_object() + assert_eq(query.find_group_overlaps(subject), [(0, 0), (2, 1)]) + assert_eq(query.count_group_overlaps(subject), [1, 0, 1, 0]) + assert_eq(@src.granges_find_overlaps_list(query.unlist(), subject), [ + (0, 0), + (1, 0), + (4, 1), + ]) +} + +///| +test "granges_list: nearest uses minimum member distance" { + let subject = @src.granges( + ["chr1", "chr1", "chr2"], + [(120, 125), (170, 175), (70, 75)], + [@src.strand_plus(), @src.strand_minus(), @src.strand_plus()], + ) + let ranges = grl_test_object() + assert_eq(ranges.nearest(subject), [0, 1, 2, -1]) + assert_eq(ranges.distance_to_nearest(subject), [10, 10, 10, -1]) +} + +///| +test "granges_list: coverage uses all member ranges" { + let coverage = grl_test_object().coverage(Map([("chr1", 320), ("chr2", 80)])) + assert_eq(coverage["chr1"][99], 1) + assert_eq(coverage["chr1"][149], 1) + assert_eq(coverage["chr1"][199], 1) + assert_eq(coverage["chr1"][299], 1) + assert_eq(coverage["chr1"][250], 0) + assert_eq(coverage["chr2"][49], 1) +} + +///| +test "grouped RangedSummarizedExperiment: construct and access" { + let experiment = grl_test_experiment() + assert_true(experiment.is_valid()) + assert_true(experiment.has_grouped_row_ranges()) + assert_eq(experiment.nrow(), 4) + assert_eq(experiment.ncol(), 2) + assert_eq(experiment.row_ranges().starts, [100, 150, 50, 0]) + assert_eq(experiment.row_ranges().ends, [209, 309, 59, -1]) + assert_eq(experiment.row_data()[0]["gene"], "geneA") + assert_eq(experiment.metadata()["organism"], "human") + assert_eq( + experiment.summary(), + "RangedSummarizedExperiment(4 range groups x 2 samples, assays=[counts])", + ) + match experiment.row_range_groups() { + Some(groups) => assert_eq(groups.element_lengths(), [2, 2, 1, 0]) + None => assert_true(false) + } +} + +///| +test "grouped RangedSummarizedExperiment: rejects assay mismatch" { + let raised = try { + ignore( + @src.RangedSummarizedExperiment::new_with_range_groups( + assays=Map([("counts", [[1.0], [2.0]])]), + row_range_groups=grl_test_object(), + col_data=[Map([])], + ), + ) + false + } catch { + RangedSummarizedExperimentError(_) => true + } + assert_true(raised) +} + +///| +test "grouped RangedSummarizedExperiment: exact overlaps avoid introns" { + let query = @src.granges( + ["chr1", "chr1", "chr1"], + [(120, 130), (205, 206), (250, 260)], + [@src.strand_plus(), @src.strand_plus(), @src.strand_minus()], + ) + let experiment = grl_test_experiment() + assert_eq(experiment.find_overlaps(query), [(0, 1)]) + assert_eq(experiment.count_overlaps(query), [1, 0, 0, 0]) +} + +///| +test "grouped RangedSummarizedExperiment: row subset stays coordinated" { + let subset = grl_test_experiment().subset_rows([2, 0, 2]) + assert_true(subset.is_valid()) + assert_eq(subset.row_names(), ["txC", "txA", "txC"]) + assert_eq(subset.row_data()[0]["gene"], "geneC") + match subset.row_range_groups() { + Some(groups) => { + assert_eq(groups.names(), ["txC", "txA", "txC"]) + assert_eq(groups.element_lengths(), [1, 2, 1]) + } + None => assert_true(false) + } + match subset.assay("counts") { + Some(assay) => assert_eq(assay, [[30.0, 31.0], [10.0, 11.0], [30.0, 31.0]]) + None => assert_true(false) + } +} + +///| +test "grouped RangedSummarizedExperiment: column subset preserves groups" { + let subset = grl_test_experiment().subset_cols([1]) + assert_true(subset.has_grouped_row_ranges()) + assert_eq(subset.ncol(), 1) + assert_eq(subset.col_data()[0]["sample"], "S2") + match subset.row_range_groups() { + Some(groups) => assert_eq(groups.element_lengths(), [2, 2, 1, 0]) + None => assert_true(false) + } +} + +///| +test "grouped RangedSummarizedExperiment: transformations preserve groups" { + let shifted = grl_test_experiment().shift(10) + let promoted = grl_test_experiment().promoters(upstream=10, downstream=3) + assert_true(shifted.has_grouped_row_ranges()) + assert_eq(shifted.row_ranges().starts, [110, 160, 60, 0]) + match shifted.row_range_groups() { + Some(groups) => + match groups.get(0) { + Some(element) => assert_eq(element.starts, [110, 210]) + None => assert_true(false) + } + None => assert_true(false) + } + match promoted.row_range_groups() { + Some(groups) => + match groups.get(1) { + Some(element) => assert_eq(element.starts, [157, 307]) + None => assert_true(false) + } + None => assert_true(false) + } +} + +///| +test "grouped RangedSummarizedExperiment: nearest and coverage use members" { + let subject = @src.granges( + ["chr1", "chr1", "chr2"], + [(120, 125), (170, 175), (70, 75)], + [@src.strand_plus(), @src.strand_minus(), @src.strand_plus()], + ) + let experiment = grl_test_experiment() + assert_eq(experiment.nearest(subject), [0, 1, 2, -1]) + assert_eq(experiment.distance_to_nearest(subject), [10, 10, 10, -1]) + let coverage = experiment.coverage(Map([("chr1", 320), ("chr2", 80)])) + assert_eq(coverage["chr1"][120], 0) + assert_eq(coverage["chr1"][199], 1) +} + +///| +test "grouped RangedSummarizedExperiment: flat replacement clears grouping" { + let flat = grl_test_experiment().with_row_ranges( + @src.granges( + ["chr1", "chr1", "chr2", "chr3"], + [(1, 10), (20, 30), (40, 50), (60, 70)], + [ + @src.strand_plus(), + @src.strand_minus(), + @src.strand_star(), + @src.strand_plus(), + ], + ), + ) + assert_false(flat.has_grouped_row_ranges()) + assert_true(flat.row_range_groups() is None) + assert_eq( + flat.summary(), + "RangedSummarizedExperiment(4 ranges x 2 samples, assays=[counts])", + ) +} diff --git a/test/moonbit/graphics_test.mbt b/test/moonbit/graphics_test.mbt index 30cd2487..a1829d7b 100644 --- a/test/moonbit/graphics_test.mbt +++ b/test/moonbit/graphics_test.mbt @@ -1,35 +1,39 @@ ///| /// Tests for Graphics module. - test "SeqLogo creation" { let seqs = ["ATGC"] let logo = @src.SeqLogo::new(seqs) assert_eq(logo.sequences.length(), 1) } +///| test "LogoColumn creation" { let col = @src.LogoColumn::new(1) assert_eq(col.position, 1) } +///| test "AlignmentPlot creation" { let seqs = [("Seq1", "ATGC")] let plot = @src.AlignmentPlot::new(seqs) assert_eq(plot.sequences.length(), 1) } +///| test "FeaturePlot creation" { let features = [("Gene", 0, 10, "gene", "+")] let plot = @src.FeaturePlot::new(features, 100) assert_eq(plot.features.length(), 1) } +///| test "seqlogo_calculate_columns" { let seqs = ["ATGC", "ATGC"] let columns = @src.seqlogo_calculate_columns(seqs) assert_eq(columns.length(), 4) } +///| test "seqlogo_generate_ascii" { let seqs = ["ATGC", "ATGC", "ATGC"] let logo = @src.SeqLogo::new(seqs) @@ -37,6 +41,7 @@ test "seqlogo_generate_ascii" { assert_true(output.length() > 0) } +///| test "alignment_plot_generate_ascii" { let seqs = [("Human", "ATGC"), ("Mouse", "ATGC")] let plot = @src.AlignmentPlot::new(seqs) @@ -44,6 +49,7 @@ test "alignment_plot_generate_ascii" { assert_true(output.length() > 0) } +///| test "feature_plot_generate_ascii" { let features = [("Gene", 0, 10, "gene", "+")] let plot = @src.FeaturePlot::new(features, 50) @@ -51,27 +57,32 @@ test "feature_plot_generate_ascii" { assert_true(output.length() > 0) } +///| test "create_example_seqlogo" { let logo = @src.create_example_seqlogo() assert_eq(logo.sequences.length(), 10) } +///| test "create_example_alignment_plot" { let plot = @src.create_example_alignment_plot() assert_eq(plot.sequences.length(), 4) } +///| test "create_example_feature_plot" { let plot = @src.create_example_feature_plot() assert_eq(plot.features.length(), 7) } +///| test "get_default_colors" { let colors = @src.get_default_colors() assert_true(colors.contains("A")) } +///| test "get_feature_colors" { let colors = @src.get_feature_colors() assert_true(colors.contains("exon")) -} \ No newline at end of file +} diff --git a/test/moonbit/gsea_base_test.mbt b/test/moonbit/gsea_base_test.mbt index f49840ab..1e5dad4a 100644 --- a/test/moonbit/gsea_base_test.mbt +++ b/test/moonbit/gsea_base_test.mbt @@ -28,12 +28,7 @@ test "gmt_gene_set_with_annotation" { let genes = ["G1"] let ct = @src.GeneSetCollectionType::from_string("canonical") let gs = @src.GmtGeneSet::with_annotation( - "g", - "desc", - genes, - ct, - "human", - "GS123", + "g", "desc", genes, ct, "human", "GS123", ) assert_eq(gs.name, "g") assert_eq(gs.organism, "human") diff --git a/test/moonbit/gsva_test.mbt b/test/moonbit/gsva_test.mbt index 074a1ddc..5df7de2b 100644 --- a/test/moonbit/gsva_test.mbt +++ b/test/moonbit/gsva_test.mbt @@ -192,7 +192,9 @@ test "gsva_survival_analysis" { let gene_sets = @src.gsva_create_example_gene_sets() let params = @src.GSVAParams::new() let scores = @src.gsva_run(data, gene_sets, params) - let survival_time = [10.0, 20.0, 30.0, 40.0, 50.0, 60.0, 70.0, 80.0, 90.0, 100.0] + let survival_time = [ + 10.0, 20.0, 30.0, 40.0, 50.0, 60.0, 70.0, 80.0, 90.0, 100.0, + ] let event = [1, 1, 1, 0, 0, 1, 1, 0, 1, 0] let surv = @src.gsva_survival_analysis(scores, survival_time, event) assert_eq(surv.length(), 5) @@ -204,7 +206,9 @@ test "gsva_survival_report" { let gene_sets = @src.gsva_create_example_gene_sets() let params = @src.GSVAParams::new() let scores = @src.gsva_run(data, gene_sets, params) - let survival_time = [10.0, 20.0, 30.0, 40.0, 50.0, 60.0, 70.0, 80.0, 90.0, 100.0] + let survival_time = [ + 10.0, 20.0, 30.0, 40.0, 50.0, 60.0, 70.0, 80.0, 90.0, 100.0, + ] let event = [1, 1, 1, 0, 0, 1, 1, 0, 1, 0] let report = @src.gsva_survival_report(scores, survival_time, event) assert_true(report.length() > 0) diff --git a/test/moonbit/gviz_test.mbt b/test/moonbit/gviz_test.mbt index 0bfc0dc6..dfe06745 100644 --- a/test/moonbit/gviz_test.mbt +++ b/test/moonbit/gviz_test.mbt @@ -1,8 +1,15 @@ ///| /// Test file for Gviz module. - test "gviz_feature_creation" { - let f = @src.gviz_feature("g1", "chr1", 100, 200, @src.track_strand_forward(), "exon", "GeneA") + let f = @src.gviz_feature( + "g1", + "chr1", + 100, + 200, + @src.track_strand_forward(), + "exon", + "GeneA", + ) assert_eq(f.feature_id, "g1") assert_eq(f.chromosome, "chr1") assert_eq(f.start, 100) @@ -11,8 +18,15 @@ test "gviz_feature_creation" { assert_eq(f.label, "GeneA") } +///| test "gviz_track_creation" { - let t = @src.gviz_track("genes", @src.track_type_gene_region(), "chr1", 1000, 5000) + let t = @src.gviz_track( + "genes", + @src.track_type_gene_region(), + "chr1", + 1000, + 5000, + ) assert_eq(t.track_name, "genes") assert_eq(t.chromosome, "chr1") assert_eq(t.start, 1000) @@ -20,43 +34,129 @@ test "gviz_track_creation" { assert_eq(t.get_n_features(), 0) } +///| test "gviz_track_add_feature" { - let t = @src.gviz_track("genes", @src.track_type_gene_region(), "chr1", 1000, 5000) - t.add_feature(@src.gviz_feature("g1", "chr1", 1200, 1800, @src.track_strand_forward(), "exon", "GeneA")) - t.add_feature(@src.gviz_feature("g2", "chr1", 2000, 3500, @src.track_strand_reverse(), "exon", "GeneB")) + let t = @src.gviz_track( + "genes", + @src.track_type_gene_region(), + "chr1", + 1000, + 5000, + ) + t.add_feature( + @src.gviz_feature( + "g1", + "chr1", + 1200, + 1800, + @src.track_strand_forward(), + "exon", + "GeneA", + ), + ) + t.add_feature( + @src.gviz_feature( + "g2", + "chr1", + 2000, + 3500, + @src.track_strand_reverse(), + "exon", + "GeneB", + ), + ) assert_eq(t.get_n_features(), 2) } +///| test "gviz_track_add_data_point" { - let t = @src.gviz_track("coverage", @src.track_type_data(), "chr1", 1000, 5000) + let t = @src.gviz_track( + "coverage", + @src.track_type_data(), + "chr1", + 1000, + 5000, + ) t.add_data_point(1000, 5.0) t.add_data_point(2000, 10.0) t.add_data_point(3000, 15.0) assert_eq(t.get_n_data_points(), 3) } +///| test "gviz_track_set_color" { - let t = @src.gviz_track("track1", @src.track_type_annotation(), "chr1", 1000, 5000) + let t = @src.gviz_track( + "track1", + @src.track_type_annotation(), + "chr1", + 1000, + 5000, + ) t.set_color("red") assert_eq(t.color, "red") } +///| test "gviz_track_set_label" { - let t = @src.gviz_track("track1", @src.track_type_annotation(), "chr1", 1000, 5000) + let t = @src.gviz_track( + "track1", + @src.track_type_annotation(), + "chr1", + 1000, + 5000, + ) t.set_label("My Track") assert_eq(t.display_label, "My Track") } +///| test "gviz_track_get_features_in_region" { - let t = @src.gviz_track("genes", @src.track_type_gene_region(), "chr1", 1000, 5000) - t.add_feature(@src.gviz_feature("g1", "chr1", 1200, 1800, @src.track_strand_forward(), "exon", "A")) - t.add_feature(@src.gviz_feature("g2", "chr1", 2000, 3500, @src.track_strand_reverse(), "exon", "B")) - t.add_feature(@src.gviz_feature("g3", "chr1", 4000, 4500, @src.track_strand_forward(), "exon", "C")) + let t = @src.gviz_track( + "genes", + @src.track_type_gene_region(), + "chr1", + 1000, + 5000, + ) + t.add_feature( + @src.gviz_feature( + "g1", + "chr1", + 1200, + 1800, + @src.track_strand_forward(), + "exon", + "A", + ), + ) + t.add_feature( + @src.gviz_feature( + "g2", + "chr1", + 2000, + 3500, + @src.track_strand_reverse(), + "exon", + "B", + ), + ) + t.add_feature( + @src.gviz_feature( + "g3", + "chr1", + 4000, + 4500, + @src.track_strand_forward(), + "exon", + "C", + ), + ) let in_region = t.get_features_in_region(1500, 3000) assert_eq(in_region.length(), 2) // g1 (overlaps) and g2 } +///| test "gviz_track_get_data_in_region" { let t = @src.gviz_track("data", @src.track_type_data(), "chr1", 1000, 5000) t.add_data_point(1000, 5.0) @@ -68,6 +168,7 @@ test "gviz_track_get_data_in_region" { assert_eq(data.length(), 2) // 2000 and 3000 } +///| test "gviz_region_creation" { let r = @src.gviz_region("chr1", 1000, 5000) assert_eq(r.chromosome, "chr1") @@ -75,6 +176,7 @@ test "gviz_region_creation" { assert_eq(r.end_, 5000) } +///| test "gviz_plot_creation" { let r = @src.gviz_region("chr1", 1000, 5000) let p = @src.gviz_plot(r, title="Test Plot", width=80, height=25) @@ -82,14 +184,22 @@ test "gviz_plot_creation" { assert_eq(p.get_region().chromosome, "chr1") } +///| test "gviz_plot_add_track" { let r = @src.gviz_region("chr1", 1000, 5000) let p = @src.gviz_plot(r, title="Test") - let t = @src.gviz_track("track1", @src.track_type_annotation(), "chr1", 1000, 5000) + let t = @src.gviz_track( + "track1", + @src.track_type_annotation(), + "chr1", + 1000, + 5000, + ) p.add_track(t) assert_eq(p.get_n_tracks(), 1) } +///| test "gviz_plot_to_ascii" { let p = @src.gviz_sample_plot() let ascii = p.to_ascii() @@ -97,6 +207,7 @@ test "gviz_plot_to_ascii" { assert_true(ascii.contains("GeneRegionTrack")) } +///| test "gviz_plot_summary" { let p = @src.gviz_sample_plot() let s = p.summary() @@ -104,6 +215,7 @@ test "gviz_plot_summary" { assert_true(s.contains("chr1")) } +///| test "gviz_track_type_to_string" { assert_eq(@src.track_type_annotation().to_string(), "AnnotationTrack") assert_eq(@src.track_type_gene_region().to_string(), "GeneRegionTrack") @@ -113,12 +225,14 @@ test "gviz_track_type_to_string" { assert_eq(@src.track_type_sequence().to_string(), "SequenceTrack") } +///| test "gviz_strand_to_string" { assert_eq(@src.track_strand_forward().to_string(), "+") assert_eq(@src.track_strand_reverse().to_string(), "-") assert_eq(@src.track_strand_unstranded().to_string(), "*") } +///| test "gviz_sample_plot" { let p = @src.gviz_sample_plot() assert_true(p.get_n_tracks() >= 3) diff --git a/test/moonbit/harmony_test.mbt b/test/moonbit/harmony_test.mbt index 586d5c68..79676d9e 100644 --- a/test/moonbit/harmony_test.mbt +++ b/test/moonbit/harmony_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Harmony batch correction module. - test "harmony_create_example" { let data = @src.harmony_create_example(100, 2, 5) assert_true(data.embeddings.length() == 100) @@ -8,6 +7,7 @@ test "harmony_create_example" { assert_true(data.embeddings[0].length() == 5) } +///| test "harmony_params_default" { let params = @src.HarmonyParams::new() assert_true(params.n_clusters == 20) @@ -16,18 +16,17 @@ test "harmony_params_default" { assert_true(params.lambda > 0.0) } +///| test "harmony_run_basic" { let data = @src.harmony_create_example(50, 2, 3) - let params = @src.HarmonyParams::create( - n_clusters = 3, - max_iterations = 10, - ) + let params = @src.HarmonyParams::create(n_clusters=3, max_iterations=10) let result = @src.harmony_run(data, params) assert_true(result.corrected.length() == 50) assert_true(result.membership.length() == 50) assert_true(result.centroids.length() == 3) } +///| test "harmony_single_batch" { // Single batch should remain largely unchanged let embeddings : Array[Array[Double]] = Array::new() @@ -41,26 +40,25 @@ test "harmony_single_batch" { i = i + 1 } let data = @src.HarmonyData::new(embeddings, batch_labels, cell_ids) - let params = @src.HarmonyParams::create( - n_clusters = 2, - max_iterations = 5, - ) + let params = @src.HarmonyParams::create(n_clusters=2, max_iterations=5) let result = @src.harmony_run(data, params) assert_true(result.corrected.length() == 20) } +///| test "harmony_convergence" { let data = @src.harmony_create_example(80, 3, 4) let params = @src.HarmonyParams::create( - n_clusters = 4, - max_iterations = 30, - tolerance = 0.001, + n_clusters=4, + max_iterations=30, + tolerance=0.001, ) let result = @src.harmony_run(data, params) assert_true(result.n_iterations > 0) assert_true(result.n_iterations <= 30) } +///| test "harmony_batch_alignment" { // Two batches with different means should be aligned let embeddings : Array[Array[Double]] = Array::new() @@ -70,25 +68,21 @@ test "harmony_batch_alignment" { while i < 40 { let batch_idx = if i < 20 { 0 } else { 1 } let offset = if batch_idx == 0 { 0.0 } else { 5.0 } - embeddings.push([offset + (i.to_double() * 0.1), offset + 1.0, offset + 2.0]) + embeddings.push([offset + i.to_double() * 0.1, offset + 1.0, offset + 2.0]) batch_labels.push("batch_" + batch_idx.to_string()) cell_ids.push("cell_" + i.to_string()) i = i + 1 } let data = @src.HarmonyData::new(embeddings, batch_labels, cell_ids) - let params = @src.HarmonyParams::create( - n_clusters = 5, - max_iterations = 30, - ) + let params = @src.HarmonyParams::create(n_clusters=5, max_iterations=30) let result = @src.harmony_run(data, params) assert_true(result.corrected.length() == 40) } +///| test "harmony_empty_data" { let data = @src.HarmonyData::new([], [], []) - let params = @src.HarmonyParams::create( - n_clusters = 2, - ) + let params = @src.HarmonyParams::create(n_clusters=2) let result = @src.harmony_run(data, params) assert_true(result.corrected.length() == 0) } diff --git a/test/moonbit/hhr_test.mbt b/test/moonbit/hhr_test.mbt new file mode 100644 index 00000000..35f07a6c --- /dev/null +++ b/test/moonbit/hhr_test.mbt @@ -0,0 +1,386 @@ +///| +/// Tests for the Biopython Bio.Align.hhr-compatible parser. + +///| +fn hhr_minimal_text() -> String { + "Query mini_query\n" + + "Match_columns 4\n" + + "No_of_seqs 1 out of 1\n" + + "Neff 1.0\n" + + "Searched_HMMs 1\n" + + " No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM\n" + + " 1 target_only 80.0 0.01 0.001 20.0 0.0 3 1-3 2-5 (5)\n" + + "No 1\n" + + ">target_only\n" + + "Probab=80.0 E-value=0.01 Score=20.0 Aligned_cols=3 Identities=66.7% Similarity=0.5 Sum_probs=2.4\n" + + "Q mini_query 1 AC-D 3 (4)\n" + + "T target_only 2 ACGD 5 (5)\n" + + "Done!\n" +} + +///| +fn hhr_raises(text : String) -> Bool { + try { + ignore(@src.hhr_parse(text)) + false + } catch { + HhrError(_) => true + } +} + +///| +test "hhr parses file metadata" { + let record = @src.hhr_parse(@src.hhr_sample_text()) + assert_eq(record.metadata.query_name, "demo_query") + assert_eq(record.metadata.match_columns, 12) + assert_eq(record.metadata.sequence_count, 24) + assert_eq(record.metadata.sequence_count_total, 40) + assert_true((record.metadata.neff - 3.5).abs() < 1.0e-12) + match record.metadata.template_neff { + Some(value) => assert_true((value - 2.1).abs() < 1.0e-12) + None => abort("expected Template_Neff") + } + assert_eq(record.metadata.searched_hmms, 2) + assert_eq(record.metadata.run_date, "Tue Aug 4 12:00:00 2026") + assert_eq(record.metadata.command_line, "hhsearch -i demo.a3m -d pdb70") +} + +///| +test "hhr parses ranked hit summaries" { + let record = @src.hhr_parse(@src.hhr_sample_text()) + assert_eq(record.hit_summaries.length(), 2) + let first = record.hit_summaries[0] + assert_eq(first.rank, 1) + assert_eq(first.target_name, "target_A") + assert_eq(first.description, "Alpha beta enzyme") + assert_true((first.probability - 99.8).abs() < 1.0e-12) + assert_true((first.evalue - 1.2e-20).abs() < 1.0e-30) + assert_true((first.pvalue - 4.0e-25).abs() < 1.0e-34) + assert_eq(first.aligned_columns, 10) + assert_eq(first.query_start, 0) + assert_eq(first.query_end, 12) + assert_eq(first.target_start, 4) + assert_eq(first.target_end, 14) + assert_eq(first.target_length, 40) +} + +///| +test "hhr joins multi-block profile alignments" { + let record = @src.hhr_parse(@src.hhr_sample_text()) + let alignment = record.alignments[0] + assert_eq(alignment.query_sequence, "ACDEFGHIKLMN") + assert_eq(alignment.target_sequence, "AC-EFGH-KLMN") + assert_eq(alignment.query_consensus, "AcDEfGHiKLMn") + assert_eq(alignment.target_consensus, "AC-EfGH-KLMN") + assert_eq(alignment.alignment_length(), 12) + assert_eq(alignment.aligned_columns, 10) + assert_eq(alignment.query_start, 0) + assert_eq(alignment.query_end, 12) + assert_eq(alignment.target_start, 4) + assert_eq(alignment.target_end, 14) +} + +///| +test "hhr preserves per-column annotations" { + let alignment = @src.hhr_parse(@src.hhr_sample_text()).alignments[0] + assert_eq(alignment.query_secondary_structure, "CCHHHHHHEECC") + assert_eq(alignment.target_secondary_structure, "CCHHHHH-EECC") + assert_eq(alignment.target_dssp, "CCEEEEE-EECC") + assert_eq(alignment.column_score, "||+|.||.||||") + assert_eq(alignment.confidence, "998879976899") +} + +///| +test "hhr exposes alignment statistics and sequence helpers" { + let alignment = @src.hhr_parse(@src.hhr_sample_text()).alignments[0] + assert_true((alignment.probability - 99.8).abs() < 1.0e-12) + assert_true((alignment.evalue - 1.2e-20).abs() < 1.0e-30) + assert_true((alignment.score - 85.4).abs() < 1.0e-12) + assert_true((alignment.identities - 72.7).abs() < 1.0e-12) + assert_true((alignment.similarity - 1.45).abs() < 1.0e-12) + assert_true((alignment.sum_probabilities - 10.8).abs() < 1.0e-12) + assert_eq(alignment.ungapped_query(), "ACDEFGHIKLMN") + assert_eq(alignment.ungapped_target(), "ACEFGHKLMN") + assert_eq(alignment.identity_count(), 10) + assert_true((alignment.query_coverage() - 1.0).abs() < 1.0e-12) + assert_true((alignment.target_coverage() - 0.25).abs() < 1.0e-12) +} + +///| +test "hhr maps query coordinates through target gaps" { + let alignment = @src.hhr_parse(@src.hhr_sample_text()).alignments[0] + assert_eq(alignment.query_to_target(0), Some(4)) + assert_true(alignment.query_to_target(2) is None) + assert_eq(alignment.query_to_target(3), Some(6)) + assert_true(alignment.query_to_target(7) is None) + assert_eq(alignment.query_to_target(11), Some(13)) + assert_true(alignment.query_to_target(12) is None) +} + +///| +test "hhr returns per-column coordinate pairs" { + let pairs = @src.hhr_parse(@src.hhr_sample_text()).alignments[0].aligned_pairs() + assert_eq(pairs.length(), 12) + match pairs[0] { + (Some(query), Some(target)) => { + assert_eq(query, 0) + assert_eq(target, 4) + } + _ => abort("expected paired first column") + } + match pairs[2] { + (Some(query), None) => assert_eq(query, 2) + _ => abort("expected target gap") + } + match pairs[11] { + (Some(query), Some(target)) => { + assert_eq(query, 11) + assert_eq(target, 13) + } + _ => abort("expected paired final column") + } +} + +///| +test "hhr supports record queries and filters" { + let record = @src.hhr_parse(@src.hhr_sample_text()) + assert_eq(record.num_hits(), 2) + assert_true(record.get(-1) is None) + assert_true(record.get(2) is None) + match record.get(1) { + Some(alignment) => assert_eq(alignment.target_name, "target_B") + None => abort("expected second alignment") + } + match record.find_target("target_A") { + Some(alignment) => assert_eq(alignment.rank, 1) + None => abort("expected target_A") + } + assert_true(record.find_target("missing") is None) + assert_eq(record.alignments_for_target("target_A").length(), 1) + assert_eq(record.filter_by_probability(90.0).length(), 1) + assert_eq(record.filter_by_evalue(0.01).length(), 2) + match record.best_alignment() { + Some(alignment) => assert_eq(alignment.rank, 1) + None => abort("expected best alignment") + } +} + +///| +test "hhr keeps repeated target alignments" { + let repeated = @src.hhr_sample_text().replace_all( + old="target_B", + new="target_A", + ) + let record = @src.hhr_parse(repeated) + assert_eq(record.alignments_for_target("target_A").length(), 2) + match record.find_target("target_A") { + Some(alignment) => assert_eq(alignment.rank, 1) + None => abort("expected repeated target") + } +} + +///| +test "hhr parses official layout without blank separators" { + let text = @src.hhr_sample_text() + assert_true(!text.contains("pdb70\n\n")) + assert_true(!text.contains("(20)\n\nNo 1")) + assert_eq(@src.hhr_parse(text).num_hits(), 2) +} + +///| +test "hhr accepts CRLF input" { + let text = @src.hhr_sample_text().replace_all(old="\n", new="\r\n") + let record = @src.hhr_parse(text) + assert_eq(record.metadata.query_name, "demo_query") + assert_eq(record.num_hits(), 2) +} + +///| +test "hhr parses optional metadata and annotation-free alignments" { + let record = @src.hhr_parse(hhr_minimal_text()) + assert_true(record.metadata.template_neff is None) + assert_eq(record.metadata.run_date, "") + let alignment = record.alignments[0] + assert_eq(alignment.target_description, "") + assert_eq(alignment.query_sequence, "AC-D") + assert_eq(alignment.target_sequence, "ACGD") + assert_eq(alignment.query_consensus, "") + assert_eq(alignment.target_consensus, "") + assert_eq(alignment.confidence, "") + assert_eq(alignment.aligned_columns, 3) + assert_true((alignment.query_coverage() - 0.75).abs() < 1.0e-12) + assert_true((alignment.target_coverage() - 0.8).abs() < 1.0e-12) +} + +///| +test "hhr parses zero-hit documents" { + let text = "Query empty_query\n" + + "Match_columns 0\n" + + "No_of_seqs 0 out of 0\n" + + "Neff 0.0\n" + + "Searched_HMMs 0\n" + + " No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM\n" + + "Done!\n" + let record = @src.hhr_parse(text) + assert_eq(record.num_hits(), 0) + assert_true(record.best_alignment() is None) +} + +///| +test "hhr serialization round trip preserves the record" { + let original = @src.hhr_parse(@src.hhr_sample_text()) + let serialized = original.to_string() + assert_true(serialized.contains("Query demo_query")) + assert_true(serialized.contains("Done!")) + let reparsed = @src.hhr_parse(serialized) + assert_true(reparsed == original) +} + +///| +test "hhr accepts EOF after a complete alignment" { + let text = hhr_minimal_text().replace_all(old="Done!\n", new="") + assert_eq(@src.hhr_parse(text).num_hits(), 1) +} + +///| +test "hhr rejects empty input" { + assert_true(hhr_raises("")) +} + +///| +test "hhr requires core metadata" { + let text = hhr_minimal_text().replace_all(old="Match_columns 4\n", new="") + assert_true(hhr_raises(text)) +} + +///| +test "hhr rejects unknown header fields" { + let text = hhr_minimal_text().replace_all( + old="Match_columns 4\n", + new="Unexpected value\nMatch_columns 4\n", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr rejects malformed sequence counts" { + let text = hhr_minimal_text().replace_all(old="1 out of 1", new="1 of 1") + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates the hit table header" { + let text = hhr_minimal_text().replace_all(old="P-value", new="Pvalue") + assert_true(hhr_raises(text)) +} + +///| +test "hhr requires contiguous summary ranks" { + let text = @src.hhr_sample_text().replace_all( + old=" 2 target_B", + new=" 3 target_B", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr requires contiguous detail ranks" { + let text = @src.hhr_sample_text().replace_all(old="No 2\n", new="No 3\n") + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates summary and detail counts" { + let row = " 2 target_B Membrane protein 72.5 0.004 2.0E-6 31.2 0.0 8 3-10 2-9 (20)\n" + let text = @src.hhr_sample_text().replace_all(old=row, new="") + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates summary and detail targets" { + let text = @src.hhr_sample_text().replace_all( + old=">target_A Alpha beta enzyme", + new=">other Alpha beta enzyme", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates aligned column counts" { + let text = @src.hhr_sample_text().replace_all( + old="Aligned_cols=10", + new="Aligned_cols=9", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates summary coordinates" { + let text = @src.hhr_sample_text().replace_all( + old="10 1-12 5-14 (40)", + new="10 1-12 5-13 (40)", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates sequence coordinates" { + let text = @src.hhr_sample_text().replace_all( + old="T target_A 5 AC-EFG 9 (40)", + new="T target_A 5 AC-EFG 10 (40)", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates multi-block coordinate continuity" { + let text = @src.hhr_sample_text().replace_all( + old="Q demo_query 7 HIKLMN", + new="Q demo_query 8 HIKLMN", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr validates annotation widths" { + let text = @src.hhr_sample_text().replace_all( + old="Q ss_pred CCHHHH", + new="Q ss_pred CCHHH", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr requires alignment statistics" { + let statistics = "Probab=80.0 E-value=0.01 Score=20.0 Aligned_cols=3 Identities=66.7% Similarity=0.5 Sum_probs=2.4\n" + assert_true( + hhr_raises(hhr_minimal_text().replace_all(old=statistics, new="")), + ) +} + +///| +test "hhr requires both alignment rows" { + let text = hhr_minimal_text().replace_all( + old="T target_only 2 ACGD 5 (5)\n", + new="", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr rejects duplicate target headers" { + let text = hhr_minimal_text().replace_all( + old=">target_only\n", + new=">target_only\n>target_only\n", + ) + assert_true(hhr_raises(text)) +} + +///| +test "hhr rejects extra content after Done" { + let text = hhr_minimal_text().replace_all( + old="Done!\n", + new="Done!\ntrailing data\n", + ) + assert_true(hhr_raises(text)) +} diff --git a/test/moonbit/hicdc_test.mbt b/test/moonbit/hicdc_test.mbt index 2195d5b7..1fa7658e 100644 --- a/test/moonbit/hicdc_test.mbt +++ b/test/moonbit/hicdc_test.mbt @@ -74,7 +74,9 @@ test "hc_sample_data_has_loops" { ///| test "hc_fit_background_returns_params" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background(contacts) + let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background( + contacts, + ) // Parameters should be finite numbers. assert_true(!intercept.is_nan()) assert_true(!beta_dist.is_nan()) @@ -86,8 +88,12 @@ test "hc_fit_background_returns_params" { ///| test "hc_predict_expected" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, _) = @src.hicdc_fit_background(contacts) - let expected = @src.hicdc_predict_expected(contacts, intercept, beta_dist, beta_gc, beta_map) + let (intercept, beta_dist, beta_gc, beta_map, _) = @src.hicdc_fit_background( + contacts, + ) + let expected = @src.hicdc_predict_expected( + contacts, intercept, beta_dist, beta_gc, beta_map, + ) assert_eq(expected.length(), contacts.length()) // Expected counts should be positive. for e in expected { @@ -98,11 +104,19 @@ test "hc_predict_expected" { ///| test "hc_expected_decreases_with_distance" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, _) = @src.hicdc_fit_background(contacts) + let (intercept, beta_dist, beta_gc, beta_map, _) = @src.hicdc_fit_background( + contacts, + ) // Predict expected counts for near and far contacts. let c1 = @src.HiCContact::new("chr1", 0, 1, 10, 1) let c2 = @src.HiCContact::new("chr1", 0, 10, 10, 1) - let exp = @src.hicdc_predict_expected([c1, c2], intercept, beta_dist, beta_gc, beta_map) + let exp = @src.hicdc_predict_expected( + [c1, c2], + intercept, + beta_dist, + beta_gc, + beta_map, + ) // Both expected counts should be positive. assert_true(exp[0] > 0.0) assert_true(exp[1] > 0.0) @@ -115,9 +129,11 @@ test "hc_expected_decreases_with_distance" { ///| test "hc_test_significance_returns_results" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background(contacts) + let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background( + contacts, + ) let results = @src.hicdc_test_significance( - contacts, intercept, beta_dist, beta_gc, beta_map, dispersion + contacts, intercept, beta_dist, beta_gc, beta_map, dispersion, ) assert_eq(results.length(), contacts.length()) // Check that p-values and FDR are valid. @@ -134,9 +150,11 @@ test "hc_test_significance_returns_results" { ///| test "hc_significant_loops_detected" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background(contacts) + let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background( + contacts, + ) let results = @src.hicdc_test_significance( - contacts, intercept, beta_dist, beta_gc, beta_map, dispersion + contacts, intercept, beta_dist, beta_gc, beta_map, dispersion, ) let marked = @src.hicdc_mark_significant(results, 0.3) // At least some contacts should be significant with loose threshold. @@ -152,9 +170,11 @@ test "hc_significant_loops_detected" { ///| test "hc_fdr_monotonic" { let contacts = @src.hicdc_sample_data() - let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background(contacts) + let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background( + contacts, + ) let results = @src.hicdc_test_significance( - contacts, intercept, beta_dist, beta_gc, beta_map, dispersion + contacts, intercept, beta_dist, beta_gc, beta_map, dispersion, ) // FDR should be >= p_value for each result (BH inflates). for r in results { @@ -284,7 +304,10 @@ test "hc_empty_contacts" { ///| test "hc_single_contact" { let c = @src.HiCContact::new("chr1", 0, 1, 50, 1) - let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background([c]) + let (intercept, beta_dist, beta_gc, beta_map, dispersion) = @src.hicdc_fit_background([ + c, + ], + ) assert_true(!intercept.is_nan()) assert_true(dispersion > 0.0) } diff --git a/test/moonbit/hilbertcurve_test.mbt b/test/moonbit/hilbertcurve_test.mbt index 824b226d..c6dc61c6 100644 --- a/test/moonbit/hilbertcurve_test.mbt +++ b/test/moonbit/hilbertcurve_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for HilbertCurve module. - test "HilbertCurve creation" { let hc = @src.HilbertCurve::new(4, 2) assert_eq(hc.levels, 4) @@ -8,6 +7,7 @@ test "HilbertCurve creation" { assert_eq(hc.max_coordinate, 15) } +///| test "hilbert_encode_decode_roundtrip" { let hc = @src.HilbertCurve::new(3, 2) let coords = [3, 5] @@ -16,6 +16,7 @@ test "hilbert_encode_decode_roundtrip" { assert_eq(decoded.length(), 2) } +///| test "hilbert_distance" { let hc = @src.HilbertCurve::new(3, 2) let coord1 = [1, 2] @@ -24,17 +25,20 @@ test "hilbert_distance" { assert_true(dist >= 0) } +///| test "hilbert_point_to_segment" { let hc = @src.HilbertCurve::new(2, 2) let segments = @src.hilbert_point_to_segment(hc, 0, 3) assert_eq(segments.length(), 4) } +///| test "hilbert_linearize_genome" { let segments = @src.hilbert_linearize_genome(10, 3) assert_eq(segments.length(), 10) } +///| test "hilbert_map_to_grid" { let hc = @src.HilbertCurve::new(2, 2) let values = [1.0, 2.0, 3.0, 4.0] diff --git a/test/moonbit/hmisc_test.mbt b/test/moonbit/hmisc_test.mbt index 5f3d2b39..f6fa587b 100644 --- a/test/moonbit/hmisc_test.mbt +++ b/test/moonbit/hmisc_test.mbt @@ -16,6 +16,7 @@ test "hmisc_describe basic valid data" { assert_true(result.median > 2.9 && result.median < 3.1) } +///| test "hmisc_describe with NaN values" { let data = [1.0, @double.not_a_number, 3.0, @double.not_a_number, 5.0] let result = @src.hmisc_describe(data) @@ -26,6 +27,7 @@ test "hmisc_describe with NaN values" { assert_eq(result.max, 5.0) } +///| test "hmisc_describe empty data" { let data : Array[Double] = Array::new() let result = @src.hmisc_describe(data) @@ -37,6 +39,7 @@ test "hmisc_describe empty data" { assert_eq(result.max, 0.0) } +///| test "hmisc_describe all NaN data" { let data = [@double.not_a_number, @double.not_a_number, @double.not_a_number] let result = @src.hmisc_describe(data) @@ -46,6 +49,7 @@ test "hmisc_describe all NaN data" { assert_eq(result.sd, 0.0) } +///| test "hmisc_describe single value" { let data = [42.0] let result = @src.hmisc_describe(data) @@ -56,12 +60,14 @@ test "hmisc_describe single value" { assert_eq(result.max, 42.0) } +///| test "hmisc_describe default name" { let data = [1.0, 2.0, 3.0] let result = @src.hmisc_describe(data) assert_eq(result.name, "") } +///| test "hmisc_describe quartiles" { let data = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let result = @src.hmisc_describe(data) @@ -72,6 +78,7 @@ test "hmisc_describe quartiles" { // hmisc_pearson tests +///| test "hmisc_pearson perfect positive correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -79,6 +86,7 @@ test "hmisc_pearson perfect positive correlation" { assert_true(result > 0.99 && result < 1.01) } +///| test "hmisc_pearson perfect negative correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [10.0, 8.0, 6.0, 4.0, 2.0] @@ -86,6 +94,7 @@ test "hmisc_pearson perfect negative correlation" { assert_true(result < -0.99 && result > -1.01) } +///| test "hmisc_pearson no correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [3.0, 1.0, 4.0, 1.0, 5.0, 9.0, 2.0, 6.0, 5.0, 3.0] @@ -93,6 +102,7 @@ test "hmisc_pearson no correlation" { assert_true(result > -0.5 && result < 0.5) } +///| test "hmisc_pearson mismatched lengths" { let x = [1.0, 2.0, 3.0, 4.0] let y = [1.0, 2.0, 3.0] @@ -100,6 +110,7 @@ test "hmisc_pearson mismatched lengths" { assert_true(result.is_nan()) } +///| test "hmisc_pearson small arrays less than 3" { let x = [1.0, 2.0] let y = [3.0, 4.0] @@ -107,6 +118,7 @@ test "hmisc_pearson small arrays less than 3" { assert_true(result.is_nan()) } +///| test "hmisc_pearson constant values" { let x = [5.0, 5.0, 5.0, 5.0, 5.0] let y = [1.0, 2.0, 3.0, 4.0, 5.0] @@ -116,6 +128,7 @@ test "hmisc_pearson constant values" { // hmisc_spearman tests +///| test "hmisc_spearman perfect positive correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -123,6 +136,7 @@ test "hmisc_spearman perfect positive correlation" { assert_true(result > 0.99 && result < 1.01) } +///| test "hmisc_spearman perfect negative correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [10.0, 8.0, 6.0, 4.0, 2.0] @@ -130,6 +144,7 @@ test "hmisc_spearman perfect negative correlation" { assert_true(result < -0.99 && result > -1.01) } +///| test "hmisc_spearman with ties" { let x = [1.0, 2.0, 2.0, 3.0, 4.0] let y = [2.0, 3.0, 3.0, 4.0, 5.0] @@ -137,6 +152,7 @@ test "hmisc_spearman with ties" { assert_true(result > 0.9 && result < 1.01) } +///| test "hmisc_spearman mismatched lengths" { let x = [1.0, 2.0, 3.0, 4.0] let y = [1.0, 2.0, 3.0] @@ -144,6 +160,7 @@ test "hmisc_spearman mismatched lengths" { assert_true(result.is_nan()) } +///| test "hmisc_spearman small arrays less than 3" { let x = [1.0, 2.0] let y = [3.0, 4.0] @@ -151,6 +168,7 @@ test "hmisc_spearman small arrays less than 3" { assert_true(result.is_nan()) } +///| test "hmisc_spearman monotonic but non-linear" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] let y = [1.0, 4.0, 9.0, 16.0, 25.0, 36.0, 49.0, 64.0] @@ -160,6 +178,7 @@ test "hmisc_spearman monotonic but non-linear" { // hmisc_rcorr tests +///| test "hmisc_rcorr basic pearson correlation" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0], @@ -178,38 +197,32 @@ test "hmisc_rcorr basic pearson correlation" { assert_true(result.matrix[1][2] < -0.99) } +///| test "hmisc_rcorr basic spearman correlation" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] let result = @src.hmisc_rcorr(data, cor_type=@src.hmisc_cor_type_spearman()) assert_eq(@src.hmisc_cor_type_to_string(result.cor_type), "Spearman") assert_true(result.matrix[0][1] > 0.99) } +///| test "hmisc_rcorr with custom names" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] let result = @src.hmisc_rcorr(data, names=["Height", "Weight"]) assert_eq(result.names[0], "Height") assert_eq(result.names[1], "Weight") } +///| test "hmisc_rcorr default names" { - let data = [ - [1.0, 2.0, 3.0, 4.0], - [2.0, 4.0, 6.0, 8.0], - [3.0, 6.0, 9.0, 12.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0], [2.0, 4.0, 6.0, 8.0], [3.0, 6.0, 9.0, 12.0]] let result = @src.hmisc_rcorr(data) assert_eq(result.names[0], "V1") assert_eq(result.names[1], "V2") assert_eq(result.names[2], "V3") } +///| test "hmisc_rcorr with NaN pairwise deletion" { let data = [ [1.0, 2.0, @double.not_a_number, 4.0, 5.0], @@ -222,6 +235,7 @@ test "hmisc_rcorr with NaN pairwise deletion" { assert_true(!result.matrix[1][2].is_nan()) } +///| test "hmisc_rcorr p-values computed" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0], @@ -232,6 +246,7 @@ test "hmisc_rcorr p-values computed" { assert_true(result.p_values[0][1] >= 0.0 && result.p_values[0][1] <= 1.0) } +///| test "hmisc_rcorr symmetry" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0], @@ -244,11 +259,9 @@ test "hmisc_rcorr symmetry" { assert_true(result.matrix[1][2] == result.matrix[2][1]) } +///| test "hmisc_rcorr n matrix" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] let result = @src.hmisc_rcorr(data) assert_eq(result.n[0][0], 5) assert_eq(result.n[1][1], 5) @@ -258,18 +271,24 @@ test "hmisc_rcorr n matrix" { // hmisc_varclus tests +///| test "hmisc_varclus basic clustering" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0], [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0], [5.0, 4.0, 6.0, 3.0, 7.0, 5.0, 8.0, 4.0, 9.0, 6.0], ] - let result = @src.hmisc_varclus(data, names=["X", "Y", "Z"], min_cluster_size=2) + let result = @src.hmisc_varclus( + data, + names=["X", "Y", "Z"], + min_cluster_size=2, + ) assert_eq(result.names.length(), 3) assert_true(result.n_clusters >= 1) assert_eq(result.cluster_assignments.length(), 3) } +///| test "hmisc_varclus all variables in one cluster" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0], @@ -283,11 +302,9 @@ test "hmisc_varclus all variables in one cluster" { assert_eq(result.cluster_assignments[2], 0) } +///| test "hmisc_varclus with two variables" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] let result = @src.hmisc_varclus(data, names=["A", "B"]) assert_eq(result.n_clusters, 1) assert_eq(result.cluster_assignments.length(), 2) @@ -295,6 +312,7 @@ test "hmisc_varclus with two variables" { assert_eq(result.heights.length(), 0) } +///| test "hmisc_varclus merge steps" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0], @@ -310,6 +328,7 @@ test "hmisc_varclus merge steps" { // hmisc_somers_d tests +///| test "hmisc_somers_d basic positive" { let g1 = [3.0, 5.0, 7.0, 9.0, 11.0] let g2 = [1.0, 2.0, 4.0, 6.0, 8.0] @@ -320,6 +339,7 @@ test "hmisc_somers_d basic positive" { assert_true(result.upper >= result.d) } +///| test "hmisc_somers_d basic negative" { let g1 = [1.0, 2.0, 3.0, 4.0, 5.0] let g2 = [4.0, 5.0, 6.0, 7.0, 8.0] @@ -327,6 +347,7 @@ test "hmisc_somers_d basic negative" { assert_true(result.d < 0.0) } +///| test "hmisc_somers_d completely separated" { let g1 = [10.0, 20.0, 30.0, 40.0, 50.0] let g2 = [1.0, 2.0, 3.0, 4.0, 5.0] @@ -334,6 +355,7 @@ test "hmisc_somers_d completely separated" { assert_true(result.d > 0.9 && result.d < 1.01) } +///| test "hmisc_somers_d equal distributions" { let g1 = [1.0, 2.0, 3.0, 4.0, 5.0] let g2 = [1.0, 2.0, 3.0, 4.0, 5.0] @@ -341,6 +363,7 @@ test "hmisc_somers_d equal distributions" { assert_true(result.d > -0.1 && result.d < 0.1) } +///| test "hmisc_somers_d empty group1" { let g1 : Array[Double] = Array::new() let g2 = [1.0, 2.0, 3.0] @@ -351,6 +374,7 @@ test "hmisc_somers_d empty group1" { assert_eq(result.n, 0) } +///| test "hmisc_somers_d empty group2" { let g1 = [1.0, 2.0, 3.0] let g2 : Array[Double] = Array::new() @@ -359,6 +383,7 @@ test "hmisc_somers_d empty group2" { assert_eq(result.n, 0) } +///| test "hmisc_somers_d confidence interval" { let g1 = [5.0, 10.0, 15.0, 20.0, 25.0, 30.0, 35.0, 40.0, 45.0, 50.0] let g2 = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] @@ -368,6 +393,7 @@ test "hmisc_somers_d confidence interval" { // hmisc_impute tests +///| test "hmisc_impute basic NaN imputation" { let data = [ [1.0, @double.not_a_number, 3.0], @@ -384,17 +410,16 @@ test "hmisc_impute basic NaN imputation" { assert_eq(result[1][1], 5.0) } +///| test "hmisc_impute no NaN values" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let result = @src.hmisc_impute(data) assert_eq(result[0][0], 1.0) assert_eq(result[0][1], 2.0) assert_eq(result[1][2], 6.0) } +///| test "hmisc_impute all NaN column" { let data = [ [1.0, @double.not_a_number, 3.0], @@ -407,16 +432,16 @@ test "hmisc_impute all NaN column" { assert_eq(result[2][1], 0.0) } +///| test "hmisc_impute empty data" { let data : Array[Array[Double]] = Array::new() let result = @src.hmisc_impute(data) assert_eq(result.length(), 0) } +///| test "hmisc_impute single row" { - let data = [ - [1.0, @double.not_a_number, 3.0], - ] + let data = [[1.0, @double.not_a_number, 3.0]] let result = @src.hmisc_impute(data) assert_eq(result.length(), 1) assert_eq(result[0][0], 1.0) @@ -426,11 +451,9 @@ test "hmisc_impute single row" { // hmisc_rcorr_summary tests +///| test "hmisc_rcorr_summary basic formatting" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] let result = @src.hmisc_rcorr(data, names=["X", "Y"]) let summary = @src.hmisc_rcorr_summary(result) assert_true(summary.contains("Correlation Matrix")) @@ -440,18 +463,21 @@ test "hmisc_rcorr_summary basic formatting" { assert_true(summary.contains("Y")) } +///| test "hmisc_rcorr_summary spearman formatting" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - [2.0, 4.0, 6.0, 8.0, 10.0], - ] - let result = @src.hmisc_rcorr(data, names=["A", "B"], cor_type=@src.hmisc_cor_type_spearman()) + let data = [[1.0, 2.0, 3.0, 4.0, 5.0], [2.0, 4.0, 6.0, 8.0, 10.0]] + let result = @src.hmisc_rcorr( + data, + names=["A", "B"], + cor_type=@src.hmisc_cor_type_spearman(), + ) let summary = @src.hmisc_rcorr_summary(result) assert_true(summary.contains("Spearman")) } // hmisc_describe_summary tests +///| test "hmisc_describe_summary formatting" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let stats = @src.hmisc_describe(data, name="Height") @@ -466,6 +492,7 @@ test "hmisc_describe_summary formatting" { assert_true(summary.contains("Max:")) } +///| test "hmisc_describe_summary with NaN" { let data = [1.0, @double.not_a_number, 3.0] let stats = @src.hmisc_describe(data, name="Test") @@ -477,6 +504,7 @@ test "hmisc_describe_summary with NaN" { // hmisc_sample_data and hmisc_sample_names tests +///| test "hmisc_sample_data structure" { let data = @src.hmisc_sample_data() assert_eq(data.length(), 3) @@ -485,6 +513,7 @@ test "hmisc_sample_data structure" { assert_eq(data[2].length(), 10) } +///| test "hmisc_sample_data values" { let data = @src.hmisc_sample_data() assert_eq(data[0][0], 1.0) @@ -495,6 +524,7 @@ test "hmisc_sample_data values" { assert_eq(data[2][9], 6.0) } +///| test "hmisc_sample_names structure" { let names = @src.hmisc_sample_names() assert_eq(names.length(), 3) @@ -503,6 +533,7 @@ test "hmisc_sample_names structure" { assert_eq(names[2], "Z") } +///| test "hmisc_sample_data_with_describe" { let data = @src.hmisc_sample_data() let stats = @src.hmisc_describe(data[0], name="X") @@ -511,10 +542,11 @@ test "hmisc_sample_data_with_describe" { assert_eq(stats.max, 10.0) } +///| test "hmisc_sample_data_with_rcorr" { let data = @src.hmisc_sample_data() let names = @src.hmisc_sample_names() - let result = @src.hmisc_rcorr(data, names=names) + let result = @src.hmisc_rcorr(data, names~) assert_eq(result.names.length(), 3) assert_eq(result.matrix.length(), 3) assert_true(!result.matrix[0][1].is_nan()) @@ -522,16 +554,18 @@ test "hmisc_sample_data_with_rcorr" { assert_true(!result.matrix[1][2].is_nan()) } +///| test "hmisc_sample_data_with_varclus" { let data = @src.hmisc_sample_data() let names = @src.hmisc_sample_names() - let result = @src.hmisc_varclus(data, names=names) + let result = @src.hmisc_varclus(data, names~) assert_eq(result.names.length(), 3) assert_true(result.n_clusters >= 1) } // Edge case tests +///| test "hmisc_pearson with NaN values" { let x = [1.0, 2.0, @double.not_a_number, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -539,6 +573,7 @@ test "hmisc_pearson with NaN values" { assert_true(result.is_nan()) } +///| test "hmisc_spearman with NaN values" { let x = [1.0, 2.0, @double.not_a_number, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -546,9 +581,12 @@ test "hmisc_spearman with NaN values" { assert_true(result.is_nan()) } +///| test "hmisc_describe large dataset" { - let data = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, - 11.0, 12.0, 13.0, 14.0, 15.0, 16.0, 17.0, 18.0, 19.0, 20.0] + let data = [ + 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, 15.0, + 16.0, 17.0, 18.0, 19.0, 20.0, + ] let result = @src.hmisc_describe(data) assert_eq(result.n, 20) assert_true(result.mean > 10.0 && result.mean < 11.0) @@ -556,6 +594,7 @@ test "hmisc_describe large dataset" { assert_eq(result.max, 20.0) } +///| test "hmisc_somers_d large groups" { let g1 = [6.0, 12.0, 18.0, 24.0, 30.0, 36.0, 42.0, 48.0, 54.0, 60.0] let g2 = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] @@ -564,6 +603,7 @@ test "hmisc_somers_d large groups" { assert_true(result.d > 0.9) } +///| test "hmisc_impute preserves non-NaN values" { let data = [ [10.0, 20.0, 30.0], @@ -581,6 +621,7 @@ test "hmisc_impute preserves non-NaN values" { assert_eq(result[2][2], 90.0) } +///| test "hmisc_impute column mean computation" { let data = [ [@double.not_a_number, 10.0], @@ -592,6 +633,7 @@ test "hmisc_impute column mean computation" { assert_true(result[1][1] > 19.0 && result[1][1] < 21.0) } +///| test "hmisc_rcorr empty data" { let data : Array[Array[Double]] = Array::new() let result = @src.hmisc_rcorr(data) @@ -599,15 +641,15 @@ test "hmisc_rcorr empty data" { assert_eq(result.matrix.length(), 0) } +///| test "hmisc_varclus single variable" { - let data = [ - [1.0, 2.0, 3.0, 4.0, 5.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0, 5.0]] let result = @src.hmisc_varclus(data, names=["X"]) assert_eq(result.n_clusters, 1) assert_eq(result.cluster_assignments[0], 0) } +///| test "hmisc_varclus two correlated variables" { let data = [ [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0], @@ -619,6 +661,7 @@ test "hmisc_varclus two correlated variables" { assert_true(result.n_clusters >= 1) } +///| test "hmisc_pearson_identical arrays" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [1.0, 2.0, 3.0, 4.0, 5.0] @@ -626,6 +669,7 @@ test "hmisc_pearson_identical arrays" { assert_true(result > 0.99 && result < 1.01) } +///| test "hmisc_pearson zero correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] let y = [1.0, -1.0, 1.0, -1.0, 1.0, -1.0, 1.0, -1.0] @@ -633,6 +677,7 @@ test "hmisc_pearson zero correlation" { assert_true(result > -0.5 && result < 0.5) } +///| test "hmisc_spearman_identical arrays" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [1.0, 2.0, 3.0, 4.0, 5.0] @@ -640,6 +685,7 @@ test "hmisc_spearman_identical arrays" { assert_true(result > 0.99 && result < 1.01) } +///| test "hmisc_somers_d single element groups" { let g1 = [5.0] let g2 = [3.0] @@ -648,6 +694,7 @@ test "hmisc_somers_d single element groups" { assert_true(result.d > 0.0) } +///| test "hmisc_describe negative values" { let data = [-5.0, -3.0, -1.0, 0.0, 2.0] let result = @src.hmisc_describe(data) @@ -656,6 +703,7 @@ test "hmisc_describe negative values" { assert_true(result.mean > -1.5 && result.mean < -0.5) } +///| test "hmisc_impute all NaN row" { let data = [ [1.0, 2.0, 3.0], @@ -666,4 +714,4 @@ test "hmisc_impute all NaN row" { assert_true(!result[1][0].is_nan()) assert_true(!result[1][1].is_nan()) assert_true(!result[1][2].is_nan()) -} \ No newline at end of file +} diff --git a/test/moonbit/hmmcopy_test.mbt b/test/moonbit/hmmcopy_test.mbt index f16f6a38..0e590fc6 100644 --- a/test/moonbit/hmmcopy_test.mbt +++ b/test/moonbit/hmmcopy_test.mbt @@ -29,35 +29,20 @@ test "hc_bin_creation" { ///| test "hc_bin_default_gc_mappability" { - let b = @src.HMMcopyBin::new( - chr="chr1", - start=0, - end=1000, - reads=50, - ) + let b = @src.HMMcopyBin::new(chr="chr1", start=0, end=1000, reads=50) assert_true((b.gc - 0.5).abs() < 0.001) assert_true((b.mappability - 1.0).abs() < 0.001) } ///| test "hc_bin_width" { - let b = @src.HMMcopyBin::new( - chr="chr1", - start=1000, - end=2000, - reads=100, - ) + let b = @src.HMMcopyBin::new(chr="chr1", start=1000, end=2000, reads=100) assert_eq(b.width(), 1000) } ///| test "hc_bin_normalized_coverage" { - let b = @src.HMMcopyBin::new( - chr="chr1", - start=0, - end=500, - reads=100, - ) + let b = @src.HMMcopyBin::new(chr="chr1", start=0, end=500, reads=100) b.set_corrected_reads(200.0) let nc = b.normalized_coverage() assert_true((nc - 0.4).abs() < 0.001) @@ -65,12 +50,7 @@ test "hc_bin_normalized_coverage" { ///| test "hc_bin_zero_width_coverage" { - let b = @src.HMMcopyBin::new( - chr="chr1", - start=100, - end=100, - reads=50, - ) + let b = @src.HMMcopyBin::new(chr="chr1", start=100, end=100, reads=50) b.set_corrected_reads(75.0) // Zero width returns corrected_reads directly assert_true((b.normalized_coverage() - 75.0).abs() < 0.001) diff --git a/test/moonbit/hs_exposure_test.mbt b/test/moonbit/hs_exposure_test.mbt index 16d77421..19157284 100644 --- a/test/moonbit/hs_exposure_test.mbt +++ b/test/moonbit/hs_exposure_test.mbt @@ -3,13 +3,13 @@ test "calculate_hse_basic" { let ca = @src.PDBAtom::new("CA", 0.0, 0.0, 0.0) let cb = @src.PDBAtom::new("CB", 1.0, 0.0, 0.0) let n = @src.PDBAtom::new("N", 0.0, 1.0, 0.0) - + let all_ca : Array[@src.PDBAtom] = Array::new() all_ca.push(ca) all_ca.push(@src.PDBAtom::new("CA", 5.0, 0.0, 0.0)) all_ca.push(@src.PDBAtom::new("CA", 0.0, 5.0, 0.0)) all_ca.push(@src.PDBAtom::new("CA", 0.0, 0.0, 5.0)) - + let result = @src.calculate_hse(ca, cb, n, all_ca) assert_eq(result.residue_name, "CA") assert_true(result.hse_up >= 0.0) @@ -22,10 +22,10 @@ test "calculate_hse_single_atom" { let ca = @src.PDBAtom::new("CA", 0.0, 0.0, 0.0) let cb = @src.PDBAtom::new("CB", 1.0, 0.0, 0.0) let n = @src.PDBAtom::new("N", 0.0, 1.0, 0.0) - + let all_ca : Array[@src.PDBAtom] = Array::new() all_ca.push(ca) - + let result = @src.calculate_hse(ca, cb, n, all_ca) assert_eq(result.residue_name, "CA") } @@ -43,4 +43,4 @@ test "PDBAtom_new" { assert_eq(atom.x, 1.0) assert_eq(atom.y, 2.0) assert_eq(atom.z, 3.0) -} \ No newline at end of file +} diff --git a/test/moonbit/htsfilter_test.mbt b/test/moonbit/htsfilter_test.mbt index be78bca0..e521b590 100644 --- a/test/moonbit/htsfilter_test.mbt +++ b/test/moonbit/htsfilter_test.mbt @@ -1,16 +1,17 @@ ///| /// Tests for Bioconductor HTSFilter module - RNA-seq count filtering. - test "cpm_single_value" { let cpm = @src.hts_filter_single_cpm(1000.0, 1000000.0) assert_eq(cpm, 1000.0) } +///| test "cpm_zero_library" { let cpm = @src.hts_filter_single_cpm(1000.0, 0.0) assert_eq(cpm, 0.0) } +///| test "cpm_matrix_basic" { let counts = [[100.0, 200.0], [300.0, 400.0]] let libs = [1000000.0, 2000000.0] @@ -21,6 +22,7 @@ test "cpm_matrix_basic" { assert_true((cpm[0][0] - 100.0).abs() < 0.01) } +///| test "library_sizes" { let counts = [[10.0, 20.0, 30.0], [40.0, 50.0, 60.0]] let libs = @src.hts_filter_library_sizes(counts) @@ -30,6 +32,7 @@ test "library_sizes" { assert_true((libs[2] - 90.0).abs() < 0.01) } +///| test "filter_basic" { let (counts, groups, _) = @src.hts_filter_sample_data() let result = @src.hts_filter(counts, groups, 100.0, 2) @@ -39,6 +42,7 @@ test "filter_basic" { assert_eq(result.keep.length(), 20) } +///| test "filter_strict_threshold" { let (counts, groups, _) = @src.hts_filter_sample_data() // Very high threshold - should remove most genes @@ -46,6 +50,7 @@ test "filter_strict_threshold" { assert_true(result.n_genes_kept <= result.n_genes_input) } +///| test "filter_keep_mask" { let (counts, groups, _) = @src.hts_filter_sample_data() let result = @src.hts_filter(counts, groups, 100.0, 2) @@ -55,12 +60,15 @@ test "filter_keep_mask" { let mut i = 0 let mut true_count = 0 while i < result.keep.length() { - if result.keep[i] { true_count = true_count + 1 } + if result.keep[i] { + true_count = true_count + 1 + } i = i + 1 } assert_eq(true_count, result.n_genes_kept) } +///| test "filter_apply" { let counts = [[10.0, 20.0], [30.0, 40.0], [50.0, 60.0]] let keep = [true, false, true] @@ -70,6 +78,7 @@ test "filter_apply" { assert_eq(filtered[1][0], 50.0) } +///| test "filter_apply_names" { let names = ["A", "B", "C"] let keep = [true, false, true] @@ -79,6 +88,7 @@ test "filter_apply_names" { assert_eq(filtered[1], "C") } +///| test "filter_retention_rate" { let (counts, groups, _) = @src.hts_filter_sample_data() let result = @src.hts_filter(counts, groups, 100.0, 2) @@ -86,6 +96,7 @@ test "filter_retention_rate" { assert_true(rate >= 0.0 && rate <= 1.0) } +///| test "filter_summary" { let (counts, groups, _) = @src.hts_filter_sample_data() let result = @src.hts_filter(counts, groups, 100.0, 2) @@ -93,6 +104,7 @@ test "filter_summary" { assert_true(summary.length() > 0) } +///| test "filter_empty_counts" { let result = @src.hts_filter([], [0, 0, 1, 1], 1.0, 1) assert_eq(result.n_genes_input, 0) @@ -100,6 +112,7 @@ test "filter_empty_counts" { assert_eq(result.n_genes_removed, 0) } +///| test "filter_single_gene" { let counts = [[500.0, 600.0, 550.0, 100.0, 120.0, 110.0]] let groups = [0, 0, 0, 1, 1, 1] @@ -108,6 +121,7 @@ test "filter_single_gene" { assert_true(result.n_genes_kept <= 1) } +///| test "filter_get_keep" { let (counts, groups, _) = @src.hts_filter_sample_data() let result = @src.hts_filter(counts, groups, 100.0, 2) @@ -115,6 +129,7 @@ test "filter_get_keep" { assert_eq(keep.length(), result.keep.length()) } +///| test "cpm_zeros" { let counts = [[0.0, 0.0], [0.0, 0.0]] let libs = [1000000.0, 2000000.0] diff --git a/test/moonbit/ig_io_test.mbt b/test/moonbit/ig_io_test.mbt index 975713fb..fd33d315 100644 --- a/test/moonbit/ig_io_test.mbt +++ b/test/moonbit/ig_io_test.mbt @@ -15,9 +15,7 @@ ///| test "ig_record_new_and_accessors" { let r = @src.IgRecord::new( - "A_U455", - "HIV-1 group M subtype A", - "ATGGCAGCTGTGATGAAGCAGAGACGGGTAAGAGCTC", + "A_U455", "HIV-1 group M subtype A", "ATGGCAGCTGTGATGAAGCAGAGACGGGTAAGAGCTC", ) assert_eq(r.title(), "A_U455") assert_eq(r.comment(), "HIV-1 group M subtype A") diff --git a/test/moonbit/ihw_test.mbt b/test/moonbit/ihw_test.mbt index ba2a6ccb..efca0874 100644 --- a/test/moonbit/ihw_test.mbt +++ b/test/moonbit/ihw_test.mbt @@ -27,7 +27,10 @@ test "ihw_result_accessors" { let covs = [0.5] let result = @src.IHWResult::new(pvals, adj, weights, covs, 0.01, 3) assert_true(result.p_values()[0] > 0.009 && result.p_values()[0] < 0.011) - assert_true(result.adjusted_p_values()[0] > 0.049 && result.adjusted_p_values()[0] < 0.051) + assert_true( + result.adjusted_p_values()[0] > 0.049 && + result.adjusted_p_values()[0] < 0.051, + ) assert_true(result.weights()[0] > 1.9 && result.weights()[0] < 2.1) assert_true(result.covariates()[0] > 0.4 && result.covariates()[0] < 0.6) assert_true(result.alpha() > 0.009 && result.alpha() < 0.011) @@ -48,7 +51,11 @@ test "ihw_config_defaults" { ///| test "ihw_config_custom" { - let config = @src.IHWConfig::new(alpha=0.01, n_attempts=20, scale_type="global") + let config = @src.IHWConfig::new( + alpha=0.01, + n_attempts=20, + scale_type="global", + ) assert_true(config.alpha() > 0.009 && config.alpha() < 0.011) assert_eq(config.n_attempts(), 20) assert_eq(config.scale_type(), "global") @@ -205,7 +212,11 @@ test "ihw_with_constant_covariates" { test "ihw_with_config_global" { let pvals = [0.01, 0.04, 0.03, 0.02] let covs = [0.5, 0.3, 0.8, 0.2] - let config = @src.IHWConfig::new(alpha=0.05, n_attempts=5, scale_type="global") + let config = @src.IHWConfig::new( + alpha=0.05, + n_attempts=5, + scale_type="global", + ) let result = @src.ihw_with_config(pvals, covs, config) assert_eq(result.adjusted_p_values().length(), 4) let adj = result.adjusted_p_values() @@ -225,7 +236,11 @@ test "ihw_with_config_custom_alpha" { test "ihw_with_config_many_attempts" { let pvals = [0.001, 0.01, 0.03, 0.04, 0.5] let covs = [0.1, 0.2, 0.3, 0.4, 0.5] - let config = @src.IHWConfig::new(alpha=0.05, n_attempts=50, scale_type="local") + let config = @src.IHWConfig::new( + alpha=0.05, + n_attempts=50, + scale_type="local", + ) let result = @src.ihw_with_config(pvals, covs, config) assert_true(result.n_attempts() <= 50) } @@ -413,7 +428,7 @@ test "ihw_many_tests" { let covs : Array[Double] = Array::make(n, 0.0) let mut i = 0 while i < n { - pvals[i] = ((i + 1).to_double() / (n.to_double() + 1.0)) * 0.1 + pvals[i] = (i + 1).to_double() / (n.to_double() + 1.0) * 0.1 covs[i] = (i + 1).to_double() / n.to_double() i = i + 1 } @@ -456,4 +471,4 @@ fn count_less_than(arr : Array[Double], threshold : Double) -> Int { i = i + 1 } count -} \ No newline at end of file +} diff --git a/test/moonbit/imgt_io_test.mbt b/test/moonbit/imgt_io_test.mbt index 34b18de5..cd2f233b 100644 --- a/test/moonbit/imgt_io_test.mbt +++ b/test/moonbit/imgt_io_test.mbt @@ -18,19 +18,8 @@ ///| test "imgt_header_new_and_all_accessors" { let h = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "A*01:01:01:01", - "ORF", - "2015-03-31", - "1996-01-19", - "human", - "HLA-A", - "I", - "MHC", - "2703 bp", - " genomic-DNA", - "no comments", + "HLA00001", "A*01:01:01:01", "A*01:01:01:01", "ORF", "2015-03-31", "1996-01-19", + "human", "HLA-A", "I", "MHC", "2703 bp", " genomic-DNA", "no comments", ) assert_eq(h.accession(), "HLA00001") assert_eq(h.seq_id(), "A*01:01:01:01") @@ -50,19 +39,8 @@ test "imgt_header_new_and_all_accessors" { ///| test "imgt_header_to_string" { let h = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "A*01:01:01:01", - "ORF", - "2015-03-31", - "1996-01-19", - "human", - "HLA-A", - "I", - "MHC", - "2703 bp", - " genomic-DNA", - "", + "HLA00001", "A*01:01:01:01", "A*01:01:01:01", "ORF", "2015-03-31", "1996-01-19", + "human", "HLA-A", "I", "MHC", "2703 bp", " genomic-DNA", "", ) assert_eq( h.to_string(), @@ -73,19 +51,8 @@ test "imgt_header_to_string" { ///| test "imgt_header_to_string_all_fields_populated" { let h = @src.ImgtHeader::new( - "ACC1", - "SID1", - "ON1", - "R1", - "D1", - "D2", - "SP1", - "G1", - "IG1", - "L1", - "100 bp", - "DNA", - "C1", + "ACC1", "SID1", "ON1", "R1", "D1", "D2", "SP1", "G1", "IG1", "L1", "100 bp", + "DNA", "C1", ) let s = h.to_string() // 13 fields separated by 12 pipes. @@ -94,7 +61,9 @@ test "imgt_header_to_string_all_fields_populated" { ///| test "imgt_header_to_string_empty_fields" { - let h = @src.ImgtHeader::new("", "", "", "", "", "", "", "", "", "", "", "", "") + let h = @src.ImgtHeader::new( + "", "", "", "", "", "", "", "", "", "", "", "", "", + ) assert_eq(h.to_string(), "||||||||||||") } @@ -105,19 +74,8 @@ test "imgt_header_to_string_empty_fields" { ///| test "imgt_record_new_and_accessors" { let header = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "A*01:01:01:01", - "ORF", - "2015-03-31", - "1996-01-19", - "human", - "HLA-A", - "I", - "MHC", - "2703 bp", - " genomic-DNA", - "", + "HLA00001", "A*01:01:01:01", "A*01:01:01:01", "ORF", "2015-03-31", "1996-01-19", + "human", "HLA-A", "I", "MHC", "2703 bp", " genomic-DNA", "", ) let rec = @src.ImgtRecord::new(header, "ATGGCCGTCATGGCGCCCCGAACCCTCCTCCTC") assert_eq(rec.header().accession(), "HLA00001") @@ -127,19 +85,7 @@ test "imgt_record_new_and_accessors" { ///| test "imgt_record_id" { let header = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", + "HLA00001", "A*01:01:01:01", "", "", "", "", "", "", "", "", "", "", "", ) let rec = @src.ImgtRecord::new(header, "ACGT") assert_eq(rec.id(), "A*01:01:01:01") @@ -148,19 +94,8 @@ test "imgt_record_id" { ///| test "imgt_record_description" { let header = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "A*01:01:01:01", - "ORF", - "2015-03-31", - "1996-01-19", - "human", - "HLA-A", - "I", - "MHC", - "2703 bp", - " genomic-DNA", - "", + "HLA00001", "A*01:01:01:01", "A*01:01:01:01", "ORF", "2015-03-31", "1996-01-19", + "human", "HLA-A", "I", "MHC", "2703 bp", " genomic-DNA", "", ) let rec = @src.ImgtRecord::new(header, "ACGT") // description() returns the full pipe-separated header (without '>'). @@ -170,18 +105,7 @@ test "imgt_record_description" { ///| test "imgt_record_to_string" { let header = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "", - "", - "", - "", - "human", - "HLA-A", - "", - "", - "", - "", + "HLA00001", "A*01:01:01:01", "", "", "", "", "human", "HLA-A", "", "", "", "", "", ) let rec = @src.ImgtRecord::new(header, "ATGGCC") @@ -380,19 +304,7 @@ test "imgt_parse_no_trailing_newline" { ///| test "imgt_header_to_string_basic" { let h = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", + "HLA00001", "A*01:01:01:01", "", "", "", "", "", "", "", "", "", "", "", ) assert_eq(@src.imgt_header_to_string(h), ">HLA00001|A*01:01:01:01|||||||||||") } @@ -400,19 +312,8 @@ test "imgt_header_to_string_basic" { ///| test "imgt_header_to_string_round_trip_with_parse" { let h = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "A*01:01:01:01", - "ORF", - "2015-03-31", - "1996-01-19", - "human", - "HLA-A", - "I", - "MHC", - "2703 bp", - " genomic-DNA", - "", + "HLA00001", "A*01:01:01:01", "A*01:01:01:01", "ORF", "2015-03-31", "1996-01-19", + "human", "HLA-A", "I", "MHC", "2703 bp", " genomic-DNA", "", ) let s = @src.imgt_header_to_string(h) match @src.imgt_parse_header(s) { @@ -442,19 +343,7 @@ test "imgt_header_to_string_round_trip_with_parse" { ///| test "imgt_record_to_string_includes_header_and_sequence" { let header = @src.ImgtHeader::new( - "HLA00001", - "A*01:01:01:01", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", - "", + "HLA00001", "A*01:01:01:01", "", "", "", "", "", "", "", "", "", "", "", ) let rec = @src.ImgtRecord::new(header, "ATGGCCGTC") let s = @src.imgt_record_to_string(rec) @@ -468,7 +357,9 @@ test "imgt_record_to_string_includes_header_and_sequence" { test "imgt_record_to_string_wraps_long_sequence" { // A 120-character sequence should be wrapped into two 60-character lines. let seq = "ATGGCCGTCATGGCGCCCCGAACCCTCCTCCTGCTGCTCTCTGGGGCCCCGGGGCCCCGGGGCCCCGGATGGCCGTCATGGCGCCCCGAACCCTCCTCCTGCTGCTCTCTGGGGCCCCGGGGCCCCGGGGCCCCGG" - let header = @src.ImgtHeader::new("A", "B", "", "", "", "", "", "", "", "", "", "", "") + let header = @src.ImgtHeader::new( + "A", "B", "", "", "", "", "", "", "", "", "", "", "", + ) let rec = @src.ImgtRecord::new(header, seq) let s = @src.imgt_record_to_string(rec) // The first line is the header (starts with '>'); the next two lines are @@ -481,7 +372,9 @@ test "imgt_record_to_string_wraps_long_sequence" { ///| test "imgt_record_to_string_empty_sequence" { - let header = @src.ImgtHeader::new("A", "B", "", "", "", "", "", "", "", "", "", "", "") + let header = @src.ImgtHeader::new( + "A", "B", "", "", "", "", "", "", "", "", "", "", "", + ) let rec = @src.ImgtRecord::new(header, "") let s = @src.imgt_record_to_string(rec) // Header line followed by a newline; no sequence line content. @@ -529,7 +422,9 @@ test "imgt_n_records_empty" { ///| test "imgt_n_records_single" { - let h = @src.ImgtHeader::new("A", "B", "", "", "", "", "", "", "", "", "", "", "") + let h = @src.ImgtHeader::new( + "A", "B", "", "", "", "", "", "", "", "", "", "", "", + ) let rec = @src.ImgtRecord::new(h, "ACGT") assert_eq(@src.imgt_n_records([rec]), 1) } @@ -682,9 +577,15 @@ test "imgt_unique_genes_empty" { ///| test "imgt_unique_species_preserves_first_appearance_order" { // Create records with multiple species to verify order. - let h1 = @src.ImgtHeader::new("A", "s1", "", "", "", "", "zebra", "", "", "", "", "", "") - let h2 = @src.ImgtHeader::new("B", "s2", "", "", "", "", "mouse", "", "", "", "", "", "") - let h3 = @src.ImgtHeader::new("C", "s3", "", "", "", "", "zebra", "", "", "", "", "", "") + let h1 = @src.ImgtHeader::new( + "A", "s1", "", "", "", "", "zebra", "", "", "", "", "", "", + ) + let h2 = @src.ImgtHeader::new( + "B", "s2", "", "", "", "", "mouse", "", "", "", "", "", "", + ) + let h3 = @src.ImgtHeader::new( + "C", "s3", "", "", "", "", "zebra", "", "", "", "", "", "", + ) let recs = [ @src.ImgtRecord::new(h1, "ACGT"), @src.ImgtRecord::new(h2, "TTTT"), @@ -743,8 +644,16 @@ test "imgt_to_seq_records" { ///| test "imgt_from_seq_records" { let seq_records = [ - @src.SeqRecord::new(@src.Seq::new("ACGTACGT"), id="rec1", description="desc1"), - @src.SeqRecord::new(@src.Seq::new("TTTTGGGG"), id="rec2", description="desc2"), + @src.SeqRecord::new( + @src.Seq::new("ACGTACGT"), + id="rec1", + description="desc1", + ), + @src.SeqRecord::new( + @src.Seq::new("TTTTGGGG"), + id="rec2", + description="desc2", + ), ] let imgt_records = @src.imgt_from_seq_records(seq_records) assert_eq(imgt_records.length(), 2) @@ -867,7 +776,9 @@ test "edge_case_empty_array_write" { ///| test "edge_case_single_record" { - let h = @src.ImgtHeader::new("A", "B", "", "", "", "", "", "", "", "", "", "", "") + let h = @src.ImgtHeader::new( + "A", "B", "", "", "", "", "", "", "", "", "", "", "", + ) let rec = @src.ImgtRecord::new(h, "ACGT") let text = @src.imgt_write([rec]) let reparsed = @src.imgt_parse(text) diff --git a/test/moonbit/impute_test.mbt b/test/moonbit/impute_test.mbt index 8f03f363..6e82dc17 100644 --- a/test/moonbit/impute_test.mbt +++ b/test/moonbit/impute_test.mbt @@ -41,9 +41,7 @@ test "impute_by_col_median_basic" { ///| test "impute_locf_basic" { - let data = [ - [Double::nan(), 1.0, Double::nan(), 2.0, Double::nan()], - ] + let data = [[Double::nan(), 1.0, Double::nan(), 2.0, Double::nan()]] let imputed = @src.impute_locf(data, by_row=true) assert_true(imputed[0][0].is_nan()) assert_eq(imputed[0][2], 1.0) @@ -52,9 +50,7 @@ test "impute_locf_basic" { ///| test "impute_nocb_basic" { - let data = [ - [Double::nan(), 1.0, Double::nan(), 2.0, Double::nan()], - ] + let data = [[Double::nan(), 1.0, Double::nan(), 2.0, Double::nan()]] let imputed = @src.impute_nocb(data, by_row=true) assert_eq(imputed[0][0], 1.0) assert_eq(imputed[0][2], 2.0) @@ -63,10 +59,7 @@ test "impute_nocb_basic" { ///| test "impute_na_by_zero" { - let data = [ - [1.0, Double::nan()], - [Double::nan(), 2.0], - ] + let data = [[1.0, Double::nan()], [Double::nan(), 2.0]] let imputed = @src.impute_na_by_zero(data) assert_eq(imputed[0][1], 0.0) assert_eq(imputed[1][0], 0.0) @@ -76,10 +69,7 @@ test "impute_na_by_zero" { ///| test "impute_na_stats" { - let data = [ - [1.0, Double::nan(), 3.0], - [Double::nan(), Double::nan(), 6.0], - ] + let data = [[1.0, Double::nan(), 3.0], [Double::nan(), Double::nan(), 6.0]] let stats = @src.impute_na_stats(data) assert_eq(stats.total_rows, 2) assert_eq(stats.total_cols, 3) @@ -92,10 +82,7 @@ test "impute_na_stats" { ///| test "impute_na_summary" { - let data = [ - [1.0, Double::nan()], - [Double::nan(), 2.0], - ] + let data = [[1.0, Double::nan()], [Double::nan(), 2.0]] let stats = @src.impute_na_stats(data) let s = @src.impute_na_summary(stats) assert_true(s.contains("total NAs")) @@ -103,10 +90,7 @@ test "impute_na_summary" { ///| test "impute_by_knn_no_na" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let p = @src.KNNImputeParam::new() let imputed = @src.impute_by_knn(data, p) assert_eq(imputed[0][0], 1.0) @@ -115,21 +99,14 @@ test "impute_by_knn_no_na" { ///| test "impute_by_knn_simple" { - let data = [ - [1.0, Double::nan(), 3.0], - [2.0, 5.0, 4.0], - [3.0, 6.0, 7.0], - ] + let data = [[1.0, Double::nan(), 3.0], [2.0, 5.0, 4.0], [3.0, 6.0, 7.0]] let imputed = @src.impute_by_knn_simple(data, 2) assert_true(!imputed[0][1].is_nan()) } ///| test "impute_na_mask" { - let data = [ - [1.0, Double::nan()], - [Double::nan(), 2.0], - ] + let data = [[1.0, Double::nan()], [Double::nan(), 2.0]] let mask = @src.make_na_mask(data) assert_eq(mask[0][1], true) assert_eq(mask[1][0], true) @@ -139,9 +116,7 @@ test "impute_na_mask" { ///| test "impute_by_row_median_and_min" { - let data = [ - [1.0, Double::nan(), 3.0, 100.0], - ] + let data = [[1.0, Double::nan(), 3.0, 100.0]] let med = @src.impute_by_row_median(data) assert_true(!med[0][1].is_nan()) let mn = @src.impute_by_row_min(data) diff --git a/test/moonbit/infercnv_test.mbt b/test/moonbit/infercnv_test.mbt index 7ba7e781..4cc2bf26 100644 --- a/test/moonbit/infercnv_test.mbt +++ b/test/moonbit/infercnv_test.mbt @@ -13,6 +13,7 @@ test "infercnv_gene_position_creation" { assert_eq(gp.end, 7687490) } +///| test "infercnv_ordered_genes_natural_chromosome_sort" { let genes = [ @src.gene_position("G_chr10_a", "chr10", 100, 200), @@ -31,6 +32,7 @@ test "infercnv_ordered_genes_natural_chromosome_sort" { assert_eq(og.gene_order[4].gene_id, "G_chrX_a") } +///| test "infercnv_exclude_chromosomes_filters_sex_and_mito" { let genes = [ @src.gene_position("A", "chr1", 1, 100), @@ -42,7 +44,7 @@ test "infercnv_exclude_chromosomes_filters_sex_and_mito" { let og = @src.ordered_genes(genes) let filtered = @src.exclude_chromosomes(og, ["chrX", "chrY", "chrM"]) assert_eq(filtered.n_genes, 2) - let ids = filtered.gene_order.map(fn (g) -> String { g.gene_id }) + let ids = filtered.gene_order.map(fn(g) -> String { g.gene_id }) assert_eq(ids, ["A", "C"]) } @@ -50,9 +52,10 @@ test "infercnv_exclude_chromosomes_filters_sex_and_mito" { // Log normalization // --------------------------------------------------------------------------- +///| test "infercnv_log_normalize_counts_basic" { let raw = [ - [1.0, 2.0, 7.0], // total = 10 -> per gene * 10000 + [1.0, 2.0, 7.0], // total = 10 -> per gene * 10000 ] let norm = @src.log_normalize_counts(raw, target_sum=10000.0) // first row: 1000, 2000, 7000; log2(x+1) @@ -62,6 +65,7 @@ test "infercnv_log_normalize_counts_basic" { assert_true(norm[0][2] > 12.0 && norm[0][2] < 13.0) } +///| test "infercnv_log_normalize_counts_zero_cell_preserved" { let raw = [[0.0, 0.0, 0.0]] let norm = @src.log_normalize_counts(raw) @@ -72,6 +76,7 @@ test "infercnv_log_normalize_counts_zero_cell_preserved" { // Reference method helpers // --------------------------------------------------------------------------- +///| test "infercnv_ref_method_constructors" { match @src.ref_method_global_mean() { @src.ReferenceMethod::GlobalMean => assert_true(true) @@ -88,6 +93,7 @@ test "infercnv_ref_method_constructors" { } } +///| test "infercnv_default_params_sensible" { let p = @src.default_cnv_params() assert_eq(p.window_size, 100) @@ -99,9 +105,13 @@ test "infercnv_default_params_sensible" { // Synthetic data generator // --------------------------------------------------------------------------- +///| test "infercnv_sample_data_shape" { let (input, tumour_cats, normal_cats) = @src.infercnv_sample_data( - n_tumour=20, n_normal=10, n_chr=3, n_genes_per_chr=20, + n_tumour=20, + n_normal=10, + n_chr=3, + n_genes_per_chr=20, ) assert_eq(input.n_cells, 30) assert_eq(input.n_genes, 60) @@ -117,14 +127,23 @@ test "infercnv_sample_data_shape" { // End-to-end pipeline with sample data and all reference strategies // --------------------------------------------------------------------------- +///| test "infercnv_run_pipeline_global_mean" { let (input, _tumour, _normal) = @src.infercnv_sample_data( - n_tumour=8, n_normal=6, n_chr=3, n_genes_per_chr=30, seed=1, + n_tumour=8, + n_normal=6, + n_chr=3, + n_genes_per_chr=30, + seed=1, ) let params = @src.make_cnv_params( - 11, 1.5, 0.05, @src.ref_method_global_mean(), [], + 11, + 1.5, + 0.05, + @src.ref_method_global_mean(), + [], ) - let res = @src.run_infercnv(input, params=params, excluded_chromosomes=[]) + let res = @src.run_infercnv(input, params~, excluded_chromosomes=[]) assert_eq(res.n_cells, 14) assert_eq(res.n_genes, 90) assert_eq(res.cnv_matrix.length(), 14) @@ -138,14 +157,23 @@ test "infercnv_run_pipeline_global_mean" { assert_true(n_score >= 0.0) } +///| test "infercnv_run_pipeline_reference_categories" { let (input, _tumour, normal) = @src.infercnv_sample_data( - n_tumour=10, n_normal=8, n_chr=2, n_genes_per_chr=20, seed=7, + n_tumour=10, + n_normal=8, + n_chr=2, + n_genes_per_chr=20, + seed=7, ) let params = @src.make_cnv_params( - 9, 1.5, 0.05, @src.ref_method_reference_categories(), normal, + 9, + 1.5, + 0.05, + @src.ref_method_reference_categories(), + normal, ) - let res = @src.run_infercnv(input, params=params, excluded_chromosomes=[]) + let res = @src.run_infercnv(input, params~, excluded_chromosomes=[]) assert_eq(res.n_cells, 18) assert_eq(res.n_genes, 40) let t_score = res.cluster_score("Tumour") @@ -154,17 +182,28 @@ test "infercnv_run_pipeline_reference_categories" { assert_true(t_score > n_score * 1.2) } +///| test "infercnv_run_pipeline_custom_reference" { let (input, _t, _n) = @src.infercnv_sample_data( - n_tumour=5, n_normal=4, n_chr=2, n_genes_per_chr=16, seed=99, + n_tumour=5, + n_normal=4, + n_chr=2, + n_genes_per_chr=16, + seed=99, ) // flat reference of 3.0 for all genes (for testing the path only) let ref : Array[Double] = [] - for _i in 0.. 3 boundaries @@ -203,22 +257,33 @@ test "infercnv_result_chromosome_list_matches_input" { let (chr1, s1, e1) = res.chromosome_list()[0] assert_eq(chr1, "chr1") assert_eq(s1, 0) - assert_eq(e1, 11) // 12 genes per chr, indexed 0..11 + assert_eq(e1, 11) // 12 genes per chr, indexed 0..11 } +///| test "infercnv_predict_tumour_cells_flags_tumour_like" { let (input, _t, _n) = @src.infercnv_sample_data( - n_tumour=15, n_normal=10, n_chr=3, n_genes_per_chr=30, seed=42, + n_tumour=15, + n_normal=10, + n_chr=3, + n_genes_per_chr=30, + seed=42, ) let params = @src.make_cnv_params( - 11, 1.5, 0.05, @src.ref_method_reference_categories(), ["Normal"], + 11, + 1.5, + 0.05, + @src.ref_method_reference_categories(), + ["Normal"], ) - let res = @src.run_infercnv(input, params=params, excluded_chromosomes=[]) + let res = @src.run_infercnv(input, params~, excluded_chromosomes=[]) let predicted = res.predict_tumour_cells("Normal", threshold_factor=1.3) // Majority of the tumour cells (indices 0..14) should be flagged let mut tumour_flagged = 0 for idx in predicted { - if idx < 15 { tumour_flagged = tumour_flagged + 1 } + if idx < 15 { + tumour_flagged = tumour_flagged + 1 + } } // At least 10/15 tumour cells called assert_true(tumour_flagged >= 10) diff --git a/test/moonbit/infernal_io_test.mbt b/test/moonbit/infernal_io_test.mbt new file mode 100644 index 00000000..8a5af12c --- /dev/null +++ b/test/moonbit/infernal_io_test.mbt @@ -0,0 +1,473 @@ +///| +/// Tests for the Biopython Bio.SearchIO.InfernalIO-compatible parsers. + +///| +fn infernal_format1_text() -> String { + "# target name accession query name accession mdl mdl from mdl to seq from seq to strand trunc pass gc bias score E-value inc description of target\n" + + "# ----------- --------- ---------- --------- --- -------- -------- -------- -------- ------ ----- ---- -- ---- ----- ------- --- ---------------------\n" + + "seq1 SACC RFQ1 QACC cm 1 70 10 79 + no 1 0.50 0.0 20.0 1.0e-05 ! format one target\n" +} + +///| +fn infernal_format2_text() -> String { + "# idx target name accession query name accession clan name mdl mdl from mdl to seq from seq to strand trunc pass gc bias score E-value inc olp anyidx afrct1 afrct2 winidx wfrct1 wfrct2 mdl len seq len description of target\n" + + "# --- ----------- --------- ---------- --------- --------- --- -------- -------- -------- -------- ------ ----- ---- -- ---- ----- ------- --- --- ------ ------ ------ ------ ------ ------ ------- ------- ---------------------\n" + + "1 chr2 ACC2 RFQ2 RFQ2.1 CL001 cm 2 60 900 841 - 5' 3 0.44 0.2 55.0 3.0e-18 ! ^ 7 0.25 0.50 2 0.75 0.80 65 2000 format two target\n" +} + +///| +fn infernal_noali_text() -> String { + "# cmsearch :: search a covariance model against a sequence database\n" + + "# INFERNAL 1.1.5 (Sep 2023)\n" + + "# - - -\n" + + "# target sequence database: noali.fa\n" + + "# show alignments in output: no\n" + + "# - - -\n" + + "Query: RFNOALI [M=50]\n" + + "Accession: RFNOALI.1\n" + + "Description: score-only query\n" + + "Hit scores:\n" + + " ---- --------- ------ ----- ---- ---------------- -------- -------- --- --- ----- -- --------------\n" + + " 1 ! 1.0e-08 33.0 0.1 seqA 100 81 - cm no 0.45 reverse score hit\n" + + " 2 ? 2.0e-04 20.0 0.0 seqA 200 219 + hmm 3' 0.51 second domain\n" + + " 3 ! 5.0e-03 18.0 0.2 seqB 10 29 + cm no 0.39 other target\n" + + "Internal CM pipeline statistics summary:\n" + + "//\n" + + "[ok]\n" +} + +///| +fn infernal_no_hit_text() -> String { + "# cmscan :: search a sequence against a covariance model database\n" + + "# INFERNAL 1.1.5 (Sep 2023)\n" + + "# - - -\n" + + "# target CM database: Rfam.cm\n" + + "# show alignments in output: yes\n" + + "# - - -\n" + + "Query: empty_query [L=25]\n" + + "Hit scores:\n" + + " [No hits detected that satisfy reporting thresholds]\n" + + "Internal CM pipeline statistics summary:\n" + + "//\n" + + "[ok]\n" +} + +///| +fn infernal_reverse_text() -> String { + @src.infernal_text_sample() + .replace_all(old="100 118 + .. 0.96 no 0.50", new="200 182 - .. 0.96 no 0.50") + .replace_all( + old="targetA 100 ACGUACGUGGGGGAACCGG 118", + new="targetA 200 ACGUACGUGGGGGAACCGG 182", + ) +} + +///| +fn infernal_duplicate_alignment_text() -> String { + let second = ">> targetA synthetic RNA target\n" + + " ---- --------- ------ ----- ---- --- -------- -------- --- -------- -------- --- --- ---- ----- --\n" + + " 2 ? 4.0e-06 29.0 0.0 hmm 2 18 .. 300 316 + .. 0.88 3' 0.42\n" + + "CS ((..............))\n" + + "RFTEST 2 CGUACGUAACCGGAAAAC 18\n" + + "targetA 300 CGUACGUAACCGGAAAAC 316\n" + + "PP 88888888888888888\n" + @src.infernal_text_sample().replace_all( + old="Internal CM pipeline statistics summary:\n", + new=second + "Internal CM pipeline statistics summary:\n", + ) +} + +///| +fn infernal_tab_raises(content : String, format : Int) -> Bool { + try { + ignore(@src.infernal_parse_tabular(content, format~)) + false + } catch { + InfernalError(_) => true + } +} + +///| +fn infernal_text_raises(content : String) -> Bool { + try { + ignore(@src.infernal_parse_text(content)) + false + } catch { + InfernalError(_) => true + } +} + +///| +test "InfernalIO detects tabular formats 1, 2, and 3" { + assert_true(@src.infernal_tabular_format(infernal_format1_text()) is Format1) + assert_true(@src.infernal_tabular_format(infernal_format2_text()) is Format2) + assert_true( + @src.infernal_tabular_format(@src.infernal_tabular_sample()) is Format3, + ) +} + +///| +test "InfernalIO supports an explicit tabular format without a header" { + let row = "seq1 SACC RFQ1 QACC cm 1 70 10 79 + no 1 0.50 0.0 20.0 1.0e-05 ! explicit\n" + let results = @src.infernal_parse_tabular(row, format=1) + assert_eq(results.length(), 1) + assert_eq(results[0].id, "RFQ1") +} + +///| +test "InfernalIO parses format 1 core fields" { + let query = @src.infernal_parse_tabular(infernal_format1_text())[0] + assert_eq(query.id, "RFQ1") + assert_eq(query.accession, "QACC") + assert_eq(query.sequence_length, 0) + assert_eq(query.hits[0].id, "seq1") + assert_eq(query.hits[0].accession, "SACC") + assert_eq(query.hits[0].description, "format one target") + assert_eq(query.hits[0].hsps[0].model, "cm") +} + +///| +test "InfernalIO parses format 2 clan lengths and overlap fields" { + let query = @src.infernal_parse_tabular(infernal_format2_text())[0] + let hit = query.hits[0] + let hsp = hit.hsps[0] + assert_eq(query.clan, "CL001") + assert_eq(query.sequence_length, 2000) + assert_eq(hit.model_length, 65) + assert_eq(hsp.overlap, "^") + assert_eq(hsp.any_index, Some(7)) + assert_eq(hsp.winner_index, Some(2)) + assert_eq(hsp.pipeline_pass, 3) +} + +///| +test "InfernalIO parses format 2 overlap fractions" { + let hsp = @src.infernal_parse_tabular(infernal_format2_text())[0].hits[0].hsps[0] + assert_eq(hsp.any_overlap_fraction, Some(0.25)) + assert_eq(hsp.reciprocal_overlap_fraction, Some(0.50)) + assert_eq(hsp.winner_overlap_fraction, Some(0.75)) + assert_eq(hsp.winner_reciprocal_fraction, Some(0.80)) +} + +///| +test "InfernalIO accepts missing format 2 overlap values" { + let text = infernal_format2_text().replace_all( + old="^ 7 0.25 0.50 2 0.75 0.80", + new="= - - - \" \" \"", + ) + let hsp = @src.infernal_parse_tabular(text)[0].hits[0].hsps[0] + assert_eq(hsp.overlap, "=") + assert_true(hsp.any_index is None) + assert_true(hsp.any_overlap_fraction is None) + assert_true(hsp.winner_index is None) + assert_true(hsp.winner_overlap_fraction is None) +} + +///| +test "InfernalIO parses format 3 lengths and descriptions" { + let query = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + assert_eq(query.sequence_length, 1000) + assert_eq(query.hits[0].model_length, 71) + assert_eq(query.hits[0].description, "bacterial 5S RNA locus") +} + +///| +test "InfernalIO normalizes forward coordinates to zero-based half-open" { + let fragment = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0].hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_start, 0) + assert_eq(fragment.query_end, 71) + assert_eq(fragment.hit_start, 100) + assert_eq(fragment.hit_end, 171) + assert_true(fragment.hit_strand == @src.strand_plus()) +} + +///| +test "InfernalIO normalizes reverse coordinates and retains strand" { + let fragment = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0].hits[0].hsps[1].fragments[0] + assert_eq(fragment.query_start, 1) + assert_eq(fragment.query_end, 70) + assert_eq(fragment.hit_start, 411) + assert_eq(fragment.hit_end, 480) + assert_true(fragment.hit_strand == @src.strand_minus()) +} + +///| +test "InfernalIO aggregates repeated target rows into HSPs" { + let query = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + assert_eq(query.hits.length(), 1) + assert_eq(query.hits[0].hsps.length(), 2) + assert_eq(query.count_hsps(), 2) +} + +///| +test "InfernalIO preserves query and hit first-seen order" { + let base = @src.infernal_tabular_sample() + let extra = "chrC RFSEQ3 RF00001 RF00001 cm 1 20 700 719 + no 1 0.40 0.0 10.0 0.01 ? 71 1000 later hit\n" + let text = base + extra + let results = @src.infernal_parse_tabular(text) + assert_eq(results.length(), 2) + assert_eq(results[0].id, "RF00001") + assert_eq(results[1].id, "RF00002") + assert_eq(results[0].hits[0].id, "chrA") + assert_eq(results[0].hits[1].id, "chrC") +} + +///| +test "InfernalIO parses inclusion and scientific E-values" { + let hsps = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0].hits[0].hsps + assert_true(hsps[0].is_included) + assert_false(hsps[1].is_included) + assert_true((hsps[0].evalue - 2.0e-14).abs() < 1.0e-24) +} + +///| +test "InfernalIO rejects unsupported explicit tabular formats" { + assert_true(infernal_tab_raises(infernal_format1_text(), 4)) +} + +///| +test "InfernalIO requires a header for automatic tabular detection" { + let row = "seq1 SACC RFQ1 QACC cm 1 70 10 79 + no 1 0.50 0.0 20.0 1.0e-05 ! explicit\n" + assert_true(infernal_tab_raises(row, 0)) +} + +///| +test "InfernalIO rejects short tabular rows" { + let text = "# target name accession query name accession mdl description of target\nshort row\n" + assert_true(infernal_tab_raises(text, 1)) +} + +///| +test "InfernalIO rejects invalid integer fields" { + let text = infernal_format1_text().replace_all(old="cm 1 70", new="cm x 70") + assert_true(infernal_tab_raises(text, 0)) +} + +///| +test "InfernalIO rejects invalid floating-point fields" { + let text = infernal_format1_text().replace_all( + old="20.0 1.0e-05", + new="bad 1.0e-05", + ) + assert_true(infernal_tab_raises(text, 0)) +} + +///| +test "InfernalIO rejects invalid strand fields" { + let text = infernal_format1_text().replace_all( + old="10 79 + no", + new="10 79 ? no", + ) + assert_true(infernal_tab_raises(text, 0)) +} + +///| +test "InfernalIO rejects non-positive coordinates" { + let text = infernal_format1_text().replace_all(old="cm 1 70", new="cm 0 70") + assert_true(infernal_tab_raises(text, 0)) +} + +///| +test "InfernalIO parses plain-text metadata and query fields" { + let query = @src.infernal_parse_text(@src.infernal_text_sample())[0] + assert_eq(query.id, "RFTEST") + assert_eq(query.accession, "RFTEST.1") + assert_eq(query.description, "synthetic structured RNA") + assert_eq(query.sequence_length, 19) + assert_eq(query.program, "cmsearch") + assert_eq(query.version, "1.1.5") + assert_eq(query.target, "transcripts.fa") +} + +///| +test "InfernalIO parses plain-text CM HSP statistics" { + let hsp = @src.infernal_parse_text(@src.infernal_text_sample())[0].hits[0].hsps[0] + assert_eq(hsp.model, "cm") + assert_eq(hsp.truncated, "no") + assert_eq(hsp.query_end_type, "[]") + assert_eq(hsp.hit_end_type, "..") + assert_true((hsp.bitscore - 45.0).abs() < 1.0e-12) + assert_true((hsp.average_accuracy - 0.96).abs() < 1.0e-12) + assert_true((hsp.gc - 0.50).abs() < 1.0e-12) +} + +///| +test "InfernalIO parses HMM-only pipeline output" { + let text = @src.infernal_text_sample() + .replace_all(old="0.1 cm 1 19", new="0.1 hmm 1 19") + .replace_all( + old="Internal CM pipeline statistics summary:", + new="Internal HMM-only pipeline statistics summary:", + ) + let hsp = @src.infernal_parse_text(text)[0].hits[0].hsps[0] + assert_eq(hsp.model, "hmm") + assert_eq(hsp.fragments.length(), 2) +} + +///| +test "InfernalIO preserves alignment sequences and annotations" { + let fragments = @src.infernal_parse_text(@src.infernal_text_sample())[0].hits[0].hsps[0].fragments + assert_eq(fragments[0].query_sequence, "ACGUACGU") + assert_eq(fragments[0].hit_sequence, "ACGUACGU") + assert_eq(fragments[0].consensus_structure, "(((....)") + assert_eq(fragments[0].posterior_probability, "99999999") + assert_eq(fragments[1].query_sequence, "AACCGG") + assert_eq(fragments[1].hit_sequence, "AACCGG") +} + +///| +test "InfernalIO splits local-end markers into multiple fragments" { + let hsp = @src.infernal_parse_text(@src.infernal_text_sample())[0].hits[0].hsps[0] + assert_eq(hsp.fragments.length(), 2) + assert_eq(hsp.fragments[0].model_omission_before, 0) + assert_eq(hsp.fragments[1].model_omission_before, 5) + assert_eq(hsp.fragments[1].sequence_omission_before, 5) + assert_eq(hsp.alignment_span(), 14) +} + +///| +test "InfernalIO advances coordinates across local ends" { + let fragments = @src.infernal_parse_text(@src.infernal_text_sample())[0].hits[0].hsps[0].fragments + assert_eq(fragments[0].query_start, 0) + assert_eq(fragments[0].query_end, 8) + assert_eq(fragments[0].hit_start, 99) + assert_eq(fragments[0].hit_end, 107) + assert_eq(fragments[1].query_start, 13) + assert_eq(fragments[1].query_end, 19) + assert_eq(fragments[1].hit_start, 112) + assert_eq(fragments[1].hit_end, 118) +} + +///| +test "InfernalIO advances reverse-strand local-end coordinates downward" { + let hsp = @src.infernal_parse_text(infernal_reverse_text())[0].hits[0].hsps[0] + assert_eq(hsp.fragments.length(), 2) + assert_eq(hsp.fragments[0].hit_start, 192) + assert_eq(hsp.fragments[0].hit_end, 200) + assert_eq(hsp.fragments[1].hit_start, 181) + assert_eq(hsp.fragments[1].hit_end, 187) + assert_true(hsp.fragments[1].hit_strand == @src.strand_minus()) + assert_eq(hsp.hit_start(), 181) + assert_eq(hsp.hit_end(), 200) +} + +///| +test "InfernalIO parses noali score tables" { + let query = @src.infernal_parse_text(infernal_noali_text())[0] + assert_eq(query.id, "RFNOALI") + assert_eq(query.target, "noali.fa") + assert_eq(query.hits.length(), 2) + assert_eq(query.count_hsps(), 3) + let fragment = query.hits[0].hsps[0].fragments[0] + assert_eq(fragment.query_start, -1) + assert_eq(fragment.hit_start, 80) + assert_eq(fragment.hit_end, 100) + assert_true(fragment.hit_strand == @src.strand_minus()) +} + +///| +test "InfernalIO merges duplicate noali targets and retains HSP order" { + let hit = @src.infernal_parse_text(infernal_noali_text())[0].hits[0] + assert_eq(hit.id, "seqA") + assert_eq(hit.description, "reverse score hit") + assert_eq(hit.hsps.length(), 2) + assert_true((hit.hsps[0].evalue - 1.0e-08).abs() < 1.0e-18) + assert_eq(hit.hsps[1].model, "hmm") +} + +///| +test "InfernalIO merges repeated alignment target sections" { + let query = @src.infernal_parse_text(infernal_duplicate_alignment_text())[0] + assert_eq(query.hits.length(), 1) + assert_eq(query.hits[0].hsps.length(), 2) + assert_eq(query.hits[0].hsps[1].model, "hmm") +} + +///| +test "InfernalIO returns empty hits for no-hit queries" { + let query = @src.infernal_parse_text(infernal_no_hit_text())[0] + assert_eq(query.id, "empty_query") + assert_eq(query.sequence_length, 25) + assert_eq(query.hits.length(), 0) +} + +///| +test "InfernalIO parses multiple plain-text queries" { + let second = @src.infernal_text_sample() + .replace_all(old="RFTEST", new="RFSECOND") + .replace_all(old="targetA", new="targetB") + let text = @src.infernal_text_sample().replace_all(old="[ok]\n", new="") + + second + let results = @src.infernal_parse_text(text) + assert_eq(results.length(), 2) + assert_eq(results[0].id, "RFTEST") + assert_eq(results[1].id, "RFSECOND") +} + +///| +test "InfernalIO chooses best HSP by E-value then score" { + let hit = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0].hits[0] + match hit.best_hsp() { + Some(hsp) => { + assert_true((hsp.evalue - 2.0e-14).abs() < 1.0e-24) + assert_true((hsp.bitscore - 48.5).abs() < 1.0e-12) + } + None => abort("expected a best Infernal HSP") + } +} + +///| +test "InfernalIO filters by E-value and inclusion state immutably" { + let query = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + let filtered = query.filter(maximum_evalue=1.0e-10, included_only=true) + assert_eq(filtered.hits.length(), 1) + assert_eq(filtered.count_hsps(), 1) + assert_eq(query.count_hsps(), 2) + assert_true(filtered.hits[0].hsps[0].is_included) +} + +///| +test "InfernalIO rejects invalid filter thresholds" { + let query = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + let raised = try { + ignore(query.filter(maximum_evalue=-1.0)) + false + } catch { + InfernalError(_) => true + } + assert_true(raised) +} + +///| +test "InfernalIO converts to the generic SearchIO hierarchy" { + let source = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + let query = source.to_searchio() + assert_eq(query.id, "RF00001") + assert_eq(query.seq_len, 1000) + assert_eq(query.hits.length(), 1) + assert_eq(query.hits[0].id, "chrA") + assert_eq(query.hits[0].accession, "RFSEQ1") + assert_eq(query.hits[0].hsps.length(), 2) + assert_eq(query.hits[0].hsps[1].fragments[0].hit_start, 411) + assert_true( + query.hits[0].hsps[1].fragments[0].hit_strand == @src.strand_minus(), + ) +} + +///| +test "InfernalIO summarizes query hierarchy counts" { + let query = @src.infernal_parse_tabular(@src.infernal_tabular_sample())[0] + assert_eq( + query.summary(), + "InfernalQueryResult(query=RF00001, hits=1, hsps=2, program=)", + ) +} + +///| +test "InfernalIO rejects malformed text alignment score rows" { + let malformed = @src.infernal_text_sample().replace_all( + old=" 1 ! 2.0e-10 45.0 0.1 cm 1 19 [] 100 118 + .. 0.96 no 0.50\n", + new=" 1 ! 2.0e-10 45.0 0.1 cm 1 19 [] 100 118 + .. 0.96 no\n", + ) + assert_true(infernal_text_raises(malformed)) +} diff --git a/test/moonbit/insdc_io_test.mbt b/test/moonbit/insdc_io_test.mbt index 61f1052b..8cdc3ad6 100644 --- a/test/moonbit/insdc_io_test.mbt +++ b/test/moonbit/insdc_io_test.mbt @@ -4,9 +4,7 @@ ///| test "insdc_parse_feature_table_cds" { let lines = [ - " CDS 1..100", - " /gene=\"testGene\"", - " /product=\"test protein\"", + " CDS 1..100", " /gene=\"testGene\"", " /product=\"test protein\"", " /translation=\"MKVL\"", ] let features = @src.parse_insdc_feature_table(lines) @@ -24,9 +22,7 @@ test "insdc_parse_feature_table_cds" { ///| test "insdc_parse_feature_table_gene" { let lines = [ - " gene 1..100", - " /gene=\"testGene\"", - " /locus_tag=\"TEST001\"", + " gene 1..100", " /gene=\"testGene\"", " /locus_tag=\"TEST001\"", ] let features = @src.parse_insdc_feature_table(lines) assert_eq(features.length(), 1) @@ -47,9 +43,7 @@ test "insdc_parse_feature_table_empty" { ///| test "insdc_parse_feature_table_multiline_qualifier" { let lines = [ - " CDS 1..100", - " /translation=\"MKVL", - " KLMN\"", + " CDS 1..100", " /translation=\"MKVL", " KLMN\"", ] let features = @src.parse_insdc_feature_table(lines) assert_eq(features.length(), 1) @@ -60,11 +54,8 @@ test "insdc_parse_feature_table_multiline_qualifier" { ///| test "insdc_parse_feature_table_multiple_features" { let lines = [ - " CDS 1..100", - " /gene=\"geneA\"", - " gene 1..100", - " /gene=\"geneA\"", - " /locus_tag=\"LA001\"", + " CDS 1..100", " /gene=\"geneA\"", " gene 1..100", + " /gene=\"geneA\"", " /locus_tag=\"LA001\"", ] let features = @src.parse_insdc_feature_table(lines) assert_eq(features.length(), 2) @@ -103,8 +94,7 @@ test "insdc_parse_location_with_whitespace" { ///| test "insdc_extract_feature_qualifier_missing" { let lines = [ - " CDS 1..100", - " /gene=\"testGene\"", + " CDS 1..100", " /gene=\"testGene\"", ] let features = @src.parse_insdc_feature_table(lines) let missing = @src.extract_feature_qualifier(features[0], "product") @@ -114,8 +104,7 @@ test "insdc_extract_feature_qualifier_missing" { ///| test "insdc_extract_translation_no_cds" { let lines = [ - " gene 1..100", - " /gene=\"testGene\"", + " gene 1..100", " /gene=\"testGene\"", ] let features = @src.parse_insdc_feature_table(lines) let trans = @src.extract_translation(features[0]) diff --git a/test/moonbit/internal_coords_test.mbt b/test/moonbit/internal_coords_test.mbt index af05bf6a..54528d80 100644 --- a/test/moonbit/internal_coords_test.mbt +++ b/test/moonbit/internal_coords_test.mbt @@ -80,13 +80,11 @@ test "ic_distance" { ///| test "ic_torsion_angle_creation" { - let tau = @src.TorsionAngle::new( - name="phi", - value=-0.57, - atom_names=["C", "N", "CA", "C"], - ) + let tau = @src.TorsionAngle::new(name="phi", value=-0.57, atom_names=[ + "C", "N", "CA", "C", + ]) assert_eq(tau.name, "phi") - assert_true((tau.value - (-0.57)).abs() < 1.0e-10) + assert_true((tau.value - -0.57).abs() < 1.0e-10) } ///| @@ -132,10 +130,10 @@ test "ic_add_residue_to_chain" { ///| test "ic_compute_phi" { let phi = @src.ic_compute_phi( - [0.0, 0.0, 0.0], // prev C - [1.5, 0.0, 0.0], // N - [3.0, 0.0, 0.0], // CA - [4.5, 1.0, 0.0], // C + [0.0, 0.0, 0.0], // prev C + [1.5, 0.0, 0.0], // N + [3.0, 0.0, 0.0], // CA + [4.5, 1.0, 0.0], // C ) assert_true(phi >= -@src.ic_pi() && phi <= @src.ic_pi()) } @@ -143,10 +141,10 @@ test "ic_compute_phi" { ///| test "ic_compute_psi" { let psi = @src.ic_compute_psi( - [0.0, 0.0, 0.0], // N - [1.5, 0.0, 0.0], // CA - [3.0, 0.0, 0.0], // C - [4.5, 1.0, 0.0], // next N + [0.0, 0.0, 0.0], // N + [1.5, 0.0, 0.0], // CA + [3.0, 0.0, 0.0], // C + [4.5, 1.0, 0.0], // next N ) assert_true(psi >= -@src.ic_pi() && psi <= @src.ic_pi()) } @@ -200,8 +198,8 @@ test "ic_rotamer_creation" { ///| test "ic_rotamer_library_entry" { let entry = @src.RotamerLibraryEntry::new(phi=-0.57, psi=-0.45) - assert_true((entry.phi - (-0.57)).abs() < 1.0e-10) - assert_true((entry.psi - (-0.45)).abs() < 1.0e-10) + assert_true((entry.phi - -0.57).abs() < 1.0e-10) + assert_true((entry.psi - -0.45).abs() < 1.0e-10) } ///| @@ -224,7 +222,9 @@ test "ic_chain_summary" { test "ic_ramachandran_data" { let phi_vals = [-1.0, -0.5, 0.0, 0.5, 1.0] let psi_vals = [-1.0, -0.5, 0.0, 0.5, 1.0] - let (mean_phi, mean_psi, var_phi, var_psi) = @src.ic_ramachandran_data(phi_vals, psi_vals) + let (mean_phi, mean_psi, var_phi, var_psi) = @src.ic_ramachandran_data( + phi_vals, psi_vals, + ) assert_true(mean_phi.abs() < 1.0e-10) assert_true(mean_psi.abs() < 1.0e-10) assert_true(var_phi > 0.0) @@ -242,4 +242,4 @@ test "ic_chi1_rotamers_gly" { let rotamers = @src.ic_chi1_rotamers("GLY") // GLY has one default rotamer assert_eq(rotamers.length(), 1) -} \ No newline at end of file +} diff --git a/test/moonbit/interproscan_test.mbt b/test/moonbit/interproscan_test.mbt index c9475e50..d2ce54e7 100644 --- a/test/moonbit/interproscan_test.mbt +++ b/test/moonbit/interproscan_test.mbt @@ -43,6 +43,7 @@ test "ips_record_construction" { assert_eq(r.go_terms()[1], "GO:0008150") } +///| test "ips_record_accessors_empty_go" { let empty_go : Array[String] = Array::new() let r = @src.InterproScanRecord::new( @@ -64,6 +65,7 @@ test "ips_record_accessors_empty_go" { assert_eq(r.go_terms().length(), 0) } +///| test "ips_record_to_string" { let go : Array[String] = Array::new() go.push("GO:0003674") @@ -96,20 +98,49 @@ test "ips_record_to_string" { // TSV parsing // --------------------------------------------------------------------------- +///| test "ips_parse_single_record" { let t = "\t" - let data = "sp|P12345|PROT_HUMAN" + t + "abc123" + t + "500" + t + "Pfam" + t + "PF00001" + t + "p53" + t + "1" + t + "100" + t + "150.5" + t + "T" + t + "01-Jan-2024" + t + "IPR000001" + t + "p53 domain" + t + "GO:0003674\n" + let data = "sp|P12345|PROT_HUMAN" + + t + + "abc123" + + t + + "500" + + t + + "Pfam" + + t + + "PF00001" + + t + + "p53" + + t + + "1" + + t + + "100" + + t + + "150.5" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "IPR000001" + + t + + "p53 domain" + + t + + "GO:0003674\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 1) assert_eq(records[0].protein_id(), "sp|P12345|PROT_HUMAN") } +///| test "ips_parse_multiple_records" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) assert_eq(records.length(), 8) } +///| test "ips_parse_all_fields_correct" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -132,6 +163,7 @@ test "ips_parse_all_fields_correct" { assert_eq(r.go_terms()[1], "GO:0008150") } +///| test "ips_parse_multiple_go_terms" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -147,39 +179,83 @@ test "ips_parse_multiple_go_terms" { // Edge cases // --------------------------------------------------------------------------- +///| test "ips_parse_empty_input" { let records = @src.parse_interproscan("") assert_eq(records.length(), 0) } +///| test "ips_parse_only_comments" { let data = "# InterProScan output\n# version 5.0\n# another comment\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 0) } +///| test "ips_parse_blank_lines" { let data = "\n\n \n\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 0) } +///| test "ips_parse_no_hits" { let t = "\t" - let data = "sp|P99999|NOPROT_HUMAN" + t + "xyz789" + t + "100" + t + "No hits\n" + - "sp|P12345|PROT_HUMAN" + t + "abc123" + t + "500" + t + "Pfam" + t + "PF00001" + t + "p53" + t + "1" + t + "100" + t + "150.5" + t + "T" + t + "01-Jan-2024" + t + "IPR000001" + t + "p53 domain" + t + "GO:0003674\n" + let data = "sp|P99999|NOPROT_HUMAN" + + t + + "xyz789" + + t + + "100" + + t + + "No hits\n" + + "sp|P12345|PROT_HUMAN" + + t + + "abc123" + + t + + "500" + + t + + "Pfam" + + t + + "PF00001" + + t + + "p53" + + t + + "1" + + t + + "100" + + t + + "150.5" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "IPR000001" + + t + + "p53 domain" + + t + + "GO:0003674\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 1) assert_eq(records[0].protein_id(), "sp|P12345|PROT_HUMAN") } +///| test "ips_parse_only_no_hits" { let t = "\t" - let data = "sp|P99999|NOPROT_HUMAN" + t + "xyz789" + t + "100" + t + "No hits\n" + let data = "sp|P99999|NOPROT_HUMAN" + + t + + "xyz789" + + t + + "100" + + t + + "No hits\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 0) } +///| test "ips_parse_missing_ipr" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -189,6 +265,7 @@ test "ips_parse_missing_ipr" { assert_eq(r.ipr_desc(), "") } +///| test "ips_parse_missing_go_terms" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -197,6 +274,7 @@ test "ips_parse_missing_go_terms" { assert_eq(r.go_terms().length(), 0) } +///| test "ips_parse_missing_go_terms_dash" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -205,29 +283,127 @@ test "ips_parse_missing_go_terms_dash" { assert_eq(r.go_terms().length(), 0) } +///| test "ips_parse_score_edge_cases" { let t = "\t" // Score is "-" → 0.0 - let data1 = "prot1" + t + "md5a" + t + "100" + t + "Pfam" + t + "PF01" + t + "desc" + t + "1" + t + "50" + t + "-" + t + "T" + t + "01-Jan-2024" + t + "-" + t + "-" + t + "-\n" + let data1 = "prot1" + + t + + "md5a" + + t + + "100" + + t + + "Pfam" + + t + + "PF01" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "-" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "-" + + t + + "-" + + t + + "-\n" let r1 = @src.parse_interproscan(data1) assert_eq(r1.length(), 1) assert_eq(r1[0].score(), 0.0) // Score is "NaN" → 0.0 - let data2 = "prot2" + t + "md5b" + t + "100" + t + "Pfam" + t + "PF02" + t + "desc" + t + "1" + t + "50" + t + "NaN" + t + "T" + t + "01-Jan-2024" + t + "-" + t + "-" + t + "-\n" + let data2 = "prot2" + + t + + "md5b" + + t + + "100" + + t + + "Pfam" + + t + + "PF02" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "NaN" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "-" + + t + + "-" + + t + + "-\n" let r2 = @src.parse_interproscan(data2) assert_eq(r2.length(), 1) assert_eq(r2[0].score(), 0.0) // Score is empty → 0.0 - let data3 = "prot3" + t + "md5c" + t + "100" + t + "Pfam" + t + "PF03" + t + "desc" + t + "1" + t + "50" + t + "" + t + "T" + t + "01-Jan-2024" + t + "-" + t + "-" + t + "-\n" + let data3 = "prot3" + + t + + "md5c" + + t + + "100" + + t + + "Pfam" + + t + + "PF03" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "-" + + t + + "-" + + t + + "-\n" let r3 = @src.parse_interproscan(data3) assert_eq(r3.length(), 1) assert_eq(r3[0].score(), 0.0) } +///| test "ips_parse_fewer_columns" { let t = "\t" // Line with only 10 columns should be padded to 14. - let data = "prot1" + t + "md5a" + t + "100" + t + "Pfam" + t + "PF01" + t + "desc" + t + "1" + t + "50" + t + "100.0" + t + "T\n" + let data = "prot1" + + t + + "md5a" + + t + + "100" + + t + + "Pfam" + + t + + "PF01" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "100.0" + + t + + "T\n" let records = @src.parse_interproscan(data) assert_eq(records.length(), 1) assert_eq(records[0].date(), "") @@ -236,6 +412,7 @@ test "ips_parse_fewer_columns" { assert_eq(records[0].go_terms().length(), 0) } +///| test "ips_parse_tsv_alias" { // parse_interproscan and parse_interproscan_tsv should produce the same results. let data = @src.interproscan_sample_data() @@ -249,6 +426,7 @@ test "ips_parse_tsv_alias" { } } +///| test "ips_parse_skips_comment_lines_in_sample" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -264,12 +442,14 @@ test "ips_parse_skips_comment_lines_in_sample" { // Sample data validation // --------------------------------------------------------------------------- +///| test "ips_sample_data_has_8_records" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) assert_eq(records.length(), 8) } +///| test "ips_sample_data_proteins" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -280,6 +460,7 @@ test "ips_sample_data_proteins" { assert_eq(proteins[2], "sp|O15143|ABC_HUMAN") } +///| test "ips_sample_data_signatures" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -294,6 +475,7 @@ test "ips_sample_data_signatures" { // Filtering // --------------------------------------------------------------------------- +///| test "ips_filter_by_analysis_pfam" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -305,6 +487,7 @@ test "ips_filter_by_analysis_pfam" { } } +///| test "ips_filter_by_analysis_smart" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -315,6 +498,7 @@ test "ips_filter_by_analysis_smart" { } } +///| test "ips_filter_by_analysis_no_match" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -322,10 +506,13 @@ test "ips_filter_by_analysis_no_match" { assert_eq(none.length(), 0) } +///| test "ips_filter_by_protein" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) - let prot1 = @src.interproscan_filter_by_protein(records, "sp|P12345|PROT_HUMAN") + let prot1 = @src.interproscan_filter_by_protein( + records, "sp|P12345|PROT_HUMAN", + ) // P12345 has 3 records: Pfam, SMART, PROSITEPATTERNS. assert_eq(prot1.length(), 3) for r in prot1 { @@ -333,14 +520,18 @@ test "ips_filter_by_protein" { } } +///| test "ips_filter_by_protein_multiple" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) - let prot3 = @src.interproscan_filter_by_protein(records, "sp|O15143|ABC_HUMAN") + let prot3 = @src.interproscan_filter_by_protein( + records, "sp|O15143|ABC_HUMAN", + ) // O15143 has 3 records: Pfam, HAMMER, SMART. assert_eq(prot3.length(), 3) } +///| test "ips_filter_by_protein_no_match" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -348,6 +539,7 @@ test "ips_filter_by_protein_no_match" { assert_eq(none.length(), 0) } +///| test "ips_filter_on_empty_records" { let empty : Array[@src.InterproScanRecord] = Array::new() assert_eq(@src.interproscan_filter_by_analysis(empty, "Pfam").length(), 0) @@ -358,6 +550,7 @@ test "ips_filter_on_empty_records" { // Unique value extraction // --------------------------------------------------------------------------- +///| test "ips_get_unique_proteins" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -368,6 +561,7 @@ test "ips_get_unique_proteins" { assert_true(proteins.contains("sp|O15143|ABC_HUMAN")) } +///| test "ips_get_unique_signatures" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -383,12 +577,14 @@ test "ips_get_unique_signatures" { assert_true(sigs.contains("SM00150")) } +///| test "ips_get_unique_signatures_empty" { let empty : Array[@src.InterproScanRecord] = Array::new() let sigs = @src.interproscan_get_unique_signatures(empty) assert_eq(sigs.length(), 0) } +///| test "ips_get_go_terms" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -407,17 +603,71 @@ test "ips_get_go_terms" { assert_true(go.contains("GO:0044183")) } +///| test "ips_get_go_terms_empty" { let empty : Array[@src.InterproScanRecord] = Array::new() let go = @src.interproscan_get_go_terms(empty) assert_eq(go.length(), 0) } +///| test "ips_get_go_terms_dedup" { // Two records with the same GO term should produce only one entry. let t = "\t" - let data = "prot1" + t + "md5a" + t + "100" + t + "Pfam" + t + "PF01" + t + "desc" + t + "1" + t + "50" + t + "100.0" + t + "T" + t + "01-Jan-2024" + t + "IPR01" + t + "d1" + t + "GO:0003674\n" + - "prot2" + t + "md5b" + t + "200" + t + "SMART" + t + "SM01" + t + "desc" + t + "1" + t + "50" + t + "100.0" + t + "T" + t + "01-Jan-2024" + t + "IPR02" + t + "d2" + t + "GO:0003674|GO:0008150\n" + let data = "prot1" + + t + + "md5a" + + t + + "100" + + t + + "Pfam" + + t + + "PF01" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "100.0" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "IPR01" + + t + + "d1" + + t + + "GO:0003674\n" + + "prot2" + + t + + "md5b" + + t + + "200" + + t + + "SMART" + + t + + "SM01" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "100.0" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "IPR02" + + t + + "d2" + + t + + "GO:0003674|GO:0008150\n" let records = @src.parse_interproscan(data) let go = @src.interproscan_get_go_terms(records) assert_eq(go.length(), 2) @@ -429,6 +679,7 @@ test "ips_get_go_terms_dedup" { // Grouping // --------------------------------------------------------------------------- +///| test "ips_group_by_protein" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -448,6 +699,7 @@ test "ips_group_by_protein" { assert_eq(recs2.length(), 3) } +///| test "ips_group_by_protein_preserves_order" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -458,15 +710,43 @@ test "ips_group_by_protein_preserves_order" { assert_eq(groups[2].0, "sp|O15143|ABC_HUMAN") } +///| test "ips_group_by_protein_empty" { let empty : Array[@src.InterproScanRecord] = Array::new() let groups = @src.interproscan_group_by_protein(empty) assert_eq(groups.length(), 0) } +///| test "ips_group_by_protein_single" { let t = "\t" - let data = "prot1" + t + "md5a" + t + "100" + t + "Pfam" + t + "PF01" + t + "desc" + t + "1" + t + "50" + t + "100.0" + t + "T" + t + "01-Jan-2024" + t + "IPR01" + t + "d1" + t + "GO:0003674\n" + let data = "prot1" + + t + + "md5a" + + t + + "100" + + t + + "Pfam" + + t + + "PF01" + + t + + "desc" + + t + + "1" + + t + + "50" + + t + + "100.0" + + t + + "T" + + t + + "01-Jan-2024" + + t + + "IPR01" + + t + + "d1" + + t + + "GO:0003674\n" let records = @src.parse_interproscan(data) let groups = @src.interproscan_group_by_protein(records) assert_eq(groups.length(), 1) @@ -478,6 +758,7 @@ test "ips_group_by_protein_single" { // Summary // --------------------------------------------------------------------------- +///| test "ips_summary_basic" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -489,6 +770,7 @@ test "ips_summary_basic" { assert_true(s.contains("Unique analyses: 5")) } +///| test "ips_summary_lists_analyses" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -500,6 +782,7 @@ test "ips_summary_lists_analyses" { assert_true(s.contains("HAMMER")) } +///| test "ips_summary_empty" { let empty : Array[@src.InterproScanRecord] = Array::new() let s = @src.interproscan_summary(empty) @@ -513,6 +796,7 @@ test "ips_summary_empty" { // Full round-trip // --------------------------------------------------------------------------- +///| test "ips_roundtrip_parse_and_filter" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) @@ -525,6 +809,7 @@ test "ips_roundtrip_parse_and_filter" { assert_true(pfam_sigs.contains("PF00003")) } +///| test "ips_roundtrip_group_and_filter" { let data = @src.interproscan_sample_data() let records = @src.parse_interproscan(data) diff --git a/test/moonbit/isoform_switch_analyze_r_test.mbt b/test/moonbit/isoform_switch_analyze_r_test.mbt index 8a3bc66f..445bdad9 100644 --- a/test/moonbit/isoform_switch_analyze_r_test.mbt +++ b/test/moonbit/isoform_switch_analyze_r_test.mbt @@ -1,11 +1,13 @@ ///| - test "isoform_create_expression" { let isoform = @src.IsoformExpression::new( - "iso1", "gene1", [100.0, 150.0, 120.0], - [50.0, 75.0, 60.0], [40.0, 60.0, 48.0] + "iso1", + "gene1", + [100.0, 150.0, 120.0], + [50.0, 75.0, 60.0], + [40.0, 60.0, 48.0], ) - + assert_eq(isoform.isoform_id, "iso1") assert_eq(isoform.gene_id, "gene1") assert_eq(isoform.counts.length(), 3) @@ -13,63 +15,138 @@ test "isoform_create_expression" { assert_eq(isoform.fpkm.length(), 3) } +///| test "isoform_calculate_usage" { let isoform = @src.IsoformExpression::new( - "iso1", "gene1", [100.0, 150.0], [50.0, 75.0], [40.0, 60.0] + "iso1", + "gene1", + [100.0, 150.0], + [50.0, 75.0], + [40.0, 60.0], ) - + let usage = @src.bio_isoform_calculate_usage(isoform) assert_eq(usage.length(), 2) assert_eq(usage[0], 50.0) assert_eq(usage[1], 75.0) } +///| test "isoform_calculate_dpsi" { - let iso1 = @src.IsoformExpression::new("iso1", "gene1", [100.0, 200.0], [50.0, 150.0], [40.0, 120.0]) - let iso2 = @src.IsoformExpression::new("iso2", "gene1", [100.0, 50.0], [50.0, 50.0], [40.0, 40.0]) - + let iso1 = @src.IsoformExpression::new( + "iso1", + "gene1", + [100.0, 200.0], + [50.0, 150.0], + [40.0, 120.0], + ) + let iso2 = @src.IsoformExpression::new( + "iso2", + "gene1", + [100.0, 50.0], + [50.0, 50.0], + [40.0, 40.0], + ) + let dpsi = @src.bio_isoform_calculate_dpsi(iso1, iso2, [0], [1]) - + assert_true((dpsi - 0.25).abs() < 0.01) } +///| test "isoform_find_switches" { - let data = @src.SwitchAnalyzeRlist::new(["s1", "s2", "s3", "s4"], ["control", "control", "treatment", "treatment"]) - - let iso1 = @src.IsoformExpression::new("iso1", "gene1", [100.0, 100.0, 200.0, 200.0], [50.0, 50.0, 150.0, 150.0], [40.0, 40.0, 120.0, 120.0]) - let iso2 = @src.IsoformExpression::new("iso2", "gene1", [100.0, 100.0, 50.0, 50.0], [50.0, 50.0, 50.0, 50.0], [40.0, 40.0, 40.0, 40.0]) - - let data2 = data.add_isoform(iso1).add_isoform(iso2).add_gene_annotation("gene1", "Gene1") - + let data = @src.SwitchAnalyzeRlist::new(["s1", "s2", "s3", "s4"], [ + "control", "control", "treatment", "treatment", + ]) + + let iso1 = @src.IsoformExpression::new( + "iso1", + "gene1", + [100.0, 100.0, 200.0, 200.0], + [50.0, 50.0, 150.0, 150.0], + [40.0, 40.0, 120.0, 120.0], + ) + let iso2 = @src.IsoformExpression::new( + "iso2", + "gene1", + [100.0, 100.0, 50.0, 50.0], + [50.0, 50.0, 50.0, 50.0], + [40.0, 40.0, 40.0, 40.0], + ) + + let data2 = data + .add_isoform(iso1) + .add_isoform(iso2) + .add_gene_annotation("gene1", "Gene1") + let switches = @src.bio_isoform_find_switches(data2, 0.1, 0.05) - + assert_true(switches.length() >= 1) } +///| test "isoform_switch_summary" { let switches = [ - @src.IsoformSwitch::new("gene1", "Gene1", "iso1", "iso2", 0.3, 0.25, 0.001, 0.005, "isoform1_up", ["significant_isoform_switch"]), - @src.IsoformSwitch::new("gene2", "Gene2", "iso3", "iso4", -0.25, -0.2, 0.002, 0.01, "isoform2_up", ["significant_isoform_switch"]), + @src.IsoformSwitch::new( + "gene1", + "Gene1", + "iso1", + "iso2", + 0.3, + 0.25, + 0.001, + 0.005, + "isoform1_up", + ["significant_isoform_switch"], + ), + @src.IsoformSwitch::new( + "gene2", + "Gene2", + "iso3", + "iso4", + -0.25, + -0.2, + 0.002, + 0.01, + "isoform2_up", + ["significant_isoform_switch"], + ), ] - + let summary = @src.bio_isoform_switch_summary(switches) - + assert_true(summary.contains("Total switches")) assert_true(summary.contains("Gene1")) assert_true(summary.contains("Gene2")) } +///| test "isoform_plot_psi" { let data = @src.SwitchAnalyzeRlist::new(["s1", "s2"], ["control", "treatment"]) - - let iso1 = @src.IsoformExpression::new("iso1", "gene1", [100.0, 200.0], [50.0, 100.0], [40.0, 80.0]) - let iso2 = @src.IsoformExpression::new("iso2", "gene1", [100.0, 100.0], [50.0, 50.0], [40.0, 40.0]) - - let data2 = data.add_isoform(iso1).add_isoform(iso2).add_gene_annotation("gene1", "Gene1") - + + let iso1 = @src.IsoformExpression::new( + "iso1", + "gene1", + [100.0, 200.0], + [50.0, 100.0], + [40.0, 80.0], + ) + let iso2 = @src.IsoformExpression::new( + "iso2", + "gene1", + [100.0, 100.0], + [50.0, 50.0], + [40.0, 40.0], + ) + + let data2 = data + .add_isoform(iso1) + .add_isoform(iso2) + .add_gene_annotation("gene1", "Gene1") + let plot = @src.bio_isoform_plot_psi(data2, "gene1") - + assert_true(plot.contains("Gene1")) assert_true(plot.contains("iso1")) assert_true(plot.contains("iso2")) -} \ No newline at end of file +} diff --git a/test/moonbit/karyoploter_test.mbt b/test/moonbit/karyoploter_test.mbt index fee420d6..52655080 100644 --- a/test/moonbit/karyoploter_test.mbt +++ b/test/moonbit/karyoploter_test.mbt @@ -1,12 +1,12 @@ ///| /// Test file for karyoploteR module. - test "karyotype_new" { let plot = @src.KaryotypePlot::new("hg38") assert_eq(plot.get_genome(), "hg38") assert_true(plot.get_chromosomes().length() > 0) } +///| test "karyotype_chromosomes" { let plot = @src.KaryotypePlot::new("hg38") let chroms = plot.get_chromosomes() @@ -15,42 +15,71 @@ test "karyotype_chromosomes" { assert_true(chroms.contains("chrY")) } +///| test "karyotype_chromosome_size" { let plot = @src.KaryotypePlot::new("hg38") let size = plot.get_chromosome_size("chr1") assert_eq(size, 248956422) } +///| test "karyotype_unknown_chromosome" { let plot = @src.KaryotypePlot::new("hg38") let size = plot.get_chromosome_size("chrUnknown") assert_eq(size, 0) } +///| test "karyotype_add_track" { let plot = @src.KaryotypePlot::new("hg38") - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 248956422, "Test Track") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 248956422, + "Test Track", + ) let plot2 = plot.add_track(track) assert_eq(plot2.get_n_tracks(), 1) } +///| test "karyotype_add_region" { let plot = @src.KaryotypePlot::new("hg38") - let region = @src.KaryotypeRegion::new("chr1", 1000000, 2000000, "Region 1", "#FF0000") + let region = @src.KaryotypeRegion::new( + "chr1", 1000000, 2000000, "Region 1", "#FF0000", + ) let plot2 = plot.add_region(region) assert_eq(plot2.get_n_regions(), 1) } +///| test "track_new" { - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 1000000, "My Track") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 1000000, + "My Track", + ) assert_eq(track.track_id, "t1") assert_eq(track.chromosome, "chr1") assert_eq(track.start, 0) assert_eq(track.end, 1000000) } +///| test "track_add_point" { - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 1000000, "Test") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 1000000, + "Test", + ) let point = @src.TrackPoint::new("chr1", 500000, 0.8, "peak") let track2 = track.add_point(point) assert_eq(track2.get_n_points(), 1) @@ -58,20 +87,44 @@ test "track_add_point" { assert_eq(data[0].position, 500000) } +///| test "track_set_color" { - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 1000000, "Test") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 1000000, + "Test", + ) let track2 = track.set_color("#FF0000") assert_eq(track2.color, "#FF0000") } +///| test "track_set_y_range" { - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 1000000, "Test") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 1000000, + "Test", + ) let track2 = track.set_y_range(0.0, 2.0) assert_eq(track2.y_max, 2.0) } +///| test "track_filter_chromosome" { - let track = @src.KaryotypeTrack::new("t1", @src.track_type_points(), "chr1", 0, 248956422, "Test") + let track = @src.KaryotypeTrack::new( + "t1", + @src.track_type_points(), + "chr1", + 0, + 248956422, + "Test", + ) let p1 = @src.TrackPoint::new("chr1", 1000000, 0.5, "") let p2 = @src.TrackPoint::new("chr2", 2000000, 0.8, "") let track2 = track.add_point(p1) @@ -80,6 +133,7 @@ test "track_filter_chromosome" { assert_eq(filtered.get_n_points(), 1) } +///| test "track_point_new" { let pt = @src.TrackPoint::new("chr1", 500000, 0.75, "test_point") assert_eq(pt.chromosome, "chr1") @@ -88,19 +142,24 @@ test "track_point_new" { assert_eq(pt.label, "test_point") } +///| test "karyotype_region_new" { - let region = @src.KaryotypeRegion::new("chr1", 1000, 2000, "Region", "#FF0000") + let region = @src.KaryotypeRegion::new( + "chr1", 1000, 2000, "Region", "#FF0000", + ) assert_eq(region.chromosome, "chr1") assert_eq(region.start, 1000) assert_eq(region.end, 2000) } +///| test "ideogram_band_new" { let band = @src.IdeogramBand::new("chr1", 0, 5000000, "p36.3", "gneg") assert_eq(band.chromosome, "chr1") assert_eq(band.band_name, "p36.3") } +///| test "track_type_to_string" { assert_eq(@src.track_type_points().to_string(), "points") assert_eq(@src.track_type_lines().to_string(), "lines") @@ -109,6 +168,7 @@ test "track_type_to_string" { assert_eq(@src.track_type_ideogram().to_string(), "ideogram") } +///| test "karyotype_to_ascii" { let plot = @src.karyotype_sample() let ascii = plot.to_ascii("chr1", 60) @@ -117,6 +177,7 @@ test "karyotype_to_ascii" { assert_true(ascii.contains("chr1")) } +///| test "karyotype_summary" { let plot = @src.karyotype_sample() let summary = plot.summary() @@ -125,12 +186,14 @@ test "karyotype_summary" { assert_true(summary.contains("chr1")) } +///| test "karyotype_unknown_genome" { let plot = @src.KaryotypePlot::new("mm10") assert_eq(plot.get_genome(), "mm10") assert_eq(plot.get_chromosomes().length(), 0) } +///| test "karyotype_sample" { let plot = @src.karyotype_sample() assert_true(plot.get_n_tracks() >= 2) @@ -138,8 +201,9 @@ test "karyotype_sample" { assert_eq(plot.get_genome(), "hg38") } +///| test "karyotype_ascii_unknown_chr" { let plot = @src.karyotype_sample() let ascii = plot.to_ascii("chr99", 60) assert_true(ascii.contains("not found")) -} \ No newline at end of file +} diff --git a/test/moonbit/kgml_test.mbt b/test/moonbit/kgml_test.mbt index c9077458..75078a22 100644 --- a/test/moonbit/kgml_test.mbt +++ b/test/moonbit/kgml_test.mbt @@ -19,6 +19,7 @@ test "kgml_graphics_creation" { assert_eq(g.bgcolor(), "#BFFFBF") } +///| test "kgml_entry_creation" { let e = @src.KgmlEntry::new(1, "ko:K00844", "gene") assert_eq(e.id(), 1) @@ -30,12 +31,14 @@ test "kgml_entry_creation" { assert_eq(e.components().length(), 0) } +///| test "kgml_subtype_creation" { let st = @src.KgmlSubType::new("compound", "C00118") assert_eq(st.name(), "compound") assert_eq(st.value(), "C00118") } +///| test "kgml_relation_creation" { let r = @src.KgmlRelation::new(1, 2, "ECrel") assert_eq(r.entry1(), 1) @@ -44,6 +47,7 @@ test "kgml_relation_creation" { assert_eq(r.subtypes().length(), 0) } +///| test "kgml_reaction_creation" { let rxn = @src.KgmlReaction::new("rn:R01786", "irreversible") assert_eq(rxn.name(), "rn:R01786") @@ -52,6 +56,7 @@ test "kgml_reaction_creation" { assert_eq(rxn.products().length(), 0) } +///| test "kgml_pathway_creation" { let p = @src.KgmlPathway::new("path:ko00010", "ko", "00010", "Glycolysis") assert_eq(p.name(), "path:ko00010") @@ -69,6 +74,7 @@ test "kgml_pathway_creation" { // XML parsing // --------------------------------------------------------------------------- +///| test "kgml_parse_basic_pathway" { let xml = @src.kgml_sample_pathway() let pw = @src.parse_kgml(xml) @@ -82,6 +88,7 @@ test "kgml_parse_basic_pathway" { assert_true(p.link().length() > 0) } +///| test "kgml_parse_entries" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -103,6 +110,7 @@ test "kgml_parse_entries" { assert_eq(e5.etype(), "map") } +///| test "kgml_parse_entry_graphics" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -120,6 +128,7 @@ test "kgml_parse_entry_graphics" { assert_eq(gfx.bgcolor(), "#BFFFBF") } +///| test "kgml_parse_relations" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -133,6 +142,7 @@ test "kgml_parse_relations" { assert_eq(r1.subtypes()[0].value(), "C00118") } +///| test "kgml_parse_reactions" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -152,11 +162,13 @@ test "kgml_parse_reactions" { assert_eq(r2.products().length(), 2) } +///| test "kgml_parse_invalid_returns_none" { let result = @src.parse_kgml("not a kgml file") assert_true(result.is_none()) } +///| test "kgml_parse_empty_returns_none" { let result = @src.parse_kgml("") assert_true(result.is_none()) @@ -166,6 +178,7 @@ test "kgml_parse_empty_returns_none" { // Query methods // --------------------------------------------------------------------------- +///| test "kgml_get_entry_by_id" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -176,6 +189,7 @@ test "kgml_get_entry_by_id" { assert_true(none.is_none()) } +///| test "kgml_get_entries_by_type" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -187,6 +201,7 @@ test "kgml_get_entries_by_type" { assert_eq(maps.length(), 1) } +///| test "kgml_get_relations_for_entry" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -205,6 +220,7 @@ test "kgml_get_relations_for_entry" { // Serialization // --------------------------------------------------------------------------- +///| test "kgml_to_string_roundtrip" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -222,6 +238,7 @@ test "kgml_to_string_roundtrip" { assert_eq(p2u.reactions().length(), p.reactions().length()) } +///| test "kgml_to_string_contains_pathway_tag" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -231,6 +248,7 @@ test "kgml_to_string_contains_pathway_tag" { assert_true(s.contains("")) } +///| test "kgml_to_string_contains_entries" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -240,6 +258,7 @@ test "kgml_to_string_contains_entries" { assert_true(s.contains("cpd:C00118")) } +///| test "kgml_to_string_contains_relations" { let xml = @src.kgml_sample_pathway() let p = @src.parse_kgml(xml).unwrap() @@ -249,6 +268,7 @@ test "kgml_to_string_contains_relations" { assert_true(s.contains(" Unit { + if (actual - expected).abs() > tolerance { + abort( + "lisaClust value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn lisa_test_cell( + cell_id : String, + image_id : String, + cell_type : String, + x : Double, + y : Double, +) -> @src.LisaCell { + @src.LisaCell::create(cell_id, image_id, cell_type, x, y) catch { + _ => abort("lisaClust test cell should be valid") + } +} + +///| +fn lisa_test_config( + curve_kind? : @src.LisaCurveKind = @src.lisa_k_curve(), + edge_correction? : Bool = false, + window_kind? : @src.LisaWindowKind = @src.lisa_rectangle_window(), + n_clusters? : Int = 2, + region_prefix? : String = "region", +) -> @src.LisaConfig { + @src.LisaConfig::create( + radii=[1.5, 4.0], + curve_kind~, + window_kind~, + bandwidth=1000.0, + min_density=0.05, + edge_correction~, + window_padding=0.0, + edge_samples=128, + n_clusters~, + n_starts=4, + max_iterations=100, + tolerance=1.0e-9, + seed=7, + region_prefix~, + ) catch { + _ => abort("lisaClust test configuration should be valid") + } +} + +///| +fn lisa_test_cells() -> Array[@src.LisaCell] { + [ + lisa_test_cell("a1", "image_1", "A", 0.0, 0.0), + lisa_test_cell("a2", "image_1", "A", 1.0, 1.0), + lisa_test_cell("a3", "image_1", "A", 2.0, 0.5), + lisa_test_cell("b1", "image_1", "B", 8.0, 8.0), + lisa_test_cell("b2", "image_1", "B", 9.0, 9.0), + lisa_test_cell("b3", "image_1", "B", 10.0, 8.5), + lisa_test_cell("c1", "image_1", "C", 0.0, 10.0), + lisa_test_cell("c2", "image_1", "C", 10.0, 0.0), + ] +} + +///| +fn lisa_test_exact_cells() -> Array[@src.LisaCell] { + [ + lisa_test_cell("source", "image_1", "A", 2.0, 2.0), + lisa_test_cell("near", "image_1", "B", 3.0, 2.0), + lisa_test_cell("far", "image_1", "B", 8.0, 8.0), + lisa_test_cell("corner_1", "image_1", "C", 0.0, 0.0), + lisa_test_cell("corner_2", "image_1", "C", 10.0, 10.0), + ] +} + +///| +fn lisa_test_result() -> @src.LisaResult { + @src.lisaclust(lisa_test_cells(), config=lisa_test_config()) catch { + _ => abort("lisaClust test workflow should succeed") + } +} + +///| +fn lisa_test_experiment() -> @src.SpatialExperiment { + let experiment = @src.SpatialExperiment::new() + for cell in lisa_test_cells() { + ignore( + @src.se_add_col( + experiment, + Map([ + ("cellID", cell.cell_id), + ("imageID", cell.image_id), + ("cellType", cell.cell_type), + ]), + ), + ) + ignore( + @src.se_add_spatial_coord( + experiment, + @src.SpatialCoord::new_2d(cell.x, cell.y), + ), + ) + } + experiment +} + +///| +test "lisaClust: curve and window helpers expose both modes" { + assert_true(@src.lisa_k_curve() is @src.LisaCurveKind::LisaStandardizedK) + assert_true(@src.lisa_l_curve() is @src.LisaCurveKind::LisaCenteredL) + assert_true( + @src.lisa_rectangle_window() is @src.LisaWindowKind::LisaRectangle, + ) + assert_true(@src.lisa_convex_window() is @src.LisaWindowKind::LisaConvexHull) +} + +///| +test "lisaClust: default configuration follows upstream defaults" { + let config = @src.LisaConfig::default() + assert_eq(config.radii, [20.0, 50.0, 100.0]) + assert_true(config.curve_kind is @src.LisaCurveKind::LisaStandardizedK) + assert_true(config.window_kind is @src.LisaWindowKind::LisaConvexHull) + assert_eq(config.bandwidth, 100000.0) + assert_eq(config.min_density, 0.05) + assert_eq(config.edge_correction, true) + assert_eq(config.n_clusters, 2) + assert_eq(config.region_prefix, "region") +} + +///| +test "lisaClust: custom configuration preserves controls" { + let config = @src.LisaConfig::create( + radii=[1.0, 3.0], + curve_kind=@src.lisa_l_curve(), + window_kind=@src.lisa_rectangle_window(), + bandwidth=2.5, + min_density=0.2, + edge_correction=false, + window_padding=1.0, + edge_samples=64, + n_clusters=3, + n_starts=5, + max_iterations=40, + tolerance=1.0e-6, + seed=9, + region_prefix="domain", + ) catch { + _ => abort("custom lisaClust configuration should be valid") + } + assert_eq(config.radii, [1.0, 3.0]) + assert_true(config.curve_kind is @src.LisaCurveKind::LisaCenteredL) + assert_eq(config.bandwidth, 2.5) + assert_eq(config.edge_correction, false) + assert_eq(config.n_clusters, 3) + assert_eq(config.region_prefix, "domain") +} + +///| +test "lisaClust: configuration defensively copies radii" { + let radii = [1.0, 2.0] + let config = @src.LisaConfig::create(radii~) catch { + _ => abort("lisaClust radii should be valid") + } + radii[0] = 99.0 + assert_eq(config.radii, [1.0, 2.0]) +} + +///| +test "lisaClust: configuration rejects empty radii" { + let failed = try { + ignore(@src.LisaConfig::create(radii=[])) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: configuration rejects non-increasing radii" { + let failed = try { + ignore(@src.LisaConfig::create(radii=[2.0, 2.0])) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: configuration rejects non-finite radii" { + let failed = try { + ignore(@src.LisaConfig::create(radii=[1.0, @double.not_a_number])) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: configuration rejects invalid density controls" { + let bad_bandwidth = try { + ignore(@src.LisaConfig::create(bandwidth=0.0)) + false + } catch { + LisaError(_) => true + } + let bad_minimum = try { + ignore(@src.LisaConfig::create(min_density=0.0)) + false + } catch { + LisaError(_) => true + } + assert_true(bad_bandwidth) + assert_true(bad_minimum) +} + +///| +test "lisaClust: configuration rejects invalid edge controls" { + let bad_padding = try { + ignore(@src.LisaConfig::create(window_padding=-1.0)) + false + } catch { + LisaError(_) => true + } + let bad_samples = try { + ignore(@src.LisaConfig::create(edge_samples=16)) + false + } catch { + LisaError(_) => true + } + assert_true(bad_padding) + assert_true(bad_samples) +} + +///| +test "lisaClust: configuration rejects invalid clustering controls" { + let bad_clusters = try { + ignore(@src.LisaConfig::create(n_clusters=0)) + false + } catch { + LisaError(_) => true + } + let bad_starts = try { + ignore(@src.LisaConfig::create(n_starts=0)) + false + } catch { + LisaError(_) => true + } + let bad_iterations = try { + ignore(@src.LisaConfig::create(max_iterations=0)) + false + } catch { + LisaError(_) => true + } + assert_true(bad_clusters) + assert_true(bad_starts) + assert_true(bad_iterations) +} + +///| +test "lisaClust: configuration rejects invalid tolerance seed and prefix" { + let bad_tolerance = try { + ignore(@src.LisaConfig::create(tolerance=1.0)) + false + } catch { + LisaError(_) => true + } + let bad_seed = try { + ignore(@src.LisaConfig::create(seed=-1)) + false + } catch { + LisaError(_) => true + } + let bad_prefix = try { + ignore(@src.LisaConfig::create(region_prefix="")) + false + } catch { + LisaError(_) => true + } + assert_true(bad_tolerance) + assert_true(bad_seed) + assert_true(bad_prefix) +} + +///| +test "lisaClust: cell constructor validates identifiers" { + let bad_cell = try { + ignore(@src.LisaCell::create("", "image", "A", 0.0, 0.0)) + false + } catch { + LisaError(_) => true + } + let bad_image = try { + ignore(@src.LisaCell::create("cell", "", "A", 0.0, 0.0)) + false + } catch { + LisaError(_) => true + } + let bad_type = try { + ignore(@src.LisaCell::create("cell", "image", "", 0.0, 0.0)) + false + } catch { + LisaError(_) => true + } + assert_true(bad_cell) + assert_true(bad_image) + assert_true(bad_type) +} + +///| +test "lisaClust: cell constructor rejects non-finite coordinates" { + let failed = try { + ignore( + @src.LisaCell::create("cell", "image", "A", @double.not_a_number, 0.0), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: curve matrix exposes stable dimensions and names" { + let curves = @src.lisa_curves(lisa_test_cells(), config=lisa_test_config()) catch { + _ => abort("lisaClust curves should compute") + } + assert_eq(curves.n_cells(), 8) + assert_eq(curves.target_cell_types, ["A", "B", "C"]) + assert_eq(curves.radii, [1.5, 4.0]) + assert_eq(curves.n_features(), 6) + assert_eq(curves.feature_names[0], "1.5_A") + assert_eq(curves.feature_names[5], "4_C") +} + +///| +test "lisaClust: feature lookup follows target-major radius order" { + let curves = @src.lisa_curves(lisa_test_cells(), config=lisa_test_config()) catch { + _ => abort("lisaClust curves should compute") + } + assert_eq(curves.feature_index("A", 1.5), 0) + assert_eq(curves.feature_index("B", 4.0), 3) + assert_eq(curves.feature_index("missing", 1.5), -1) + assert_eq(curves.feature_index("A", 99.0), -1) +} + +///| +test "lisaClust: exact standardized K uses observed-minus-expected scaling" { + let config = @src.LisaConfig::create( + radii=[2.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + edge_correction=false, + window_padding=0.0, + edge_samples=64, + ) catch { + _ => abort("exact K configuration should be valid") + } + let curves = @src.lisa_curves(lisa_test_exact_cells(), config~) catch { + _ => abort("exact K curves should compute") + } + let feature = curves.feature_index("B", 2.0) + let expected = @math.PI * 4.0 * 2.0 / 100.0 + let target = (1.0 - expected) / expected.sqrt() + lisa_test_close(curves.values[0][feature], target, 1.0e-5) +} + +///| +test "lisaClust: centered L uses square-root observed and expected counts" { + let config = @src.LisaConfig::create( + radii=[2.0], + curve_kind=@src.lisa_l_curve(), + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + edge_correction=false, + window_padding=0.0, + edge_samples=64, + ) catch { + _ => abort("exact L configuration should be valid") + } + let curves = @src.lisa_curves(lisa_test_exact_cells(), config~) catch { + _ => abort("exact L curves should compute") + } + let feature = curves.feature_index("B", 2.0) + let expected = @math.PI * 4.0 * 2.0 / 100.0 + lisa_test_close(curves.values[0][feature], 1.0 - expected.sqrt(), 1.0e-5) +} + +///| +test "lisaClust: local curves exclude self matches" { + let config = @src.LisaConfig::create( + radii=[0.5], + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + edge_correction=false, + window_padding=0.0, + ) catch { + _ => abort("self exclusion configuration should be valid") + } + let curves = @src.lisa_curves(lisa_test_exact_cells(), config~) catch { + _ => abort("self exclusion curves should compute") + } + let feature = curves.feature_index("A", 0.5) + let expected = @math.PI * 0.25 / 100.0 + lisa_test_close(curves.values[0][feature], -expected.sqrt(), 1.0e-6) +} + +///| +test "lisaClust: rectangle window area follows coordinate extent" { + let curves = @src.lisa_curves( + lisa_test_exact_cells(), + config=lisa_test_config(), + ) catch { + _ => abort("rectangle lisaClust curves should compute") + } + assert_eq(curves.windows.length(), 1) + lisa_test_close(curves.windows[0].area, 100.0, 1.0e-12) +} + +///| +test "lisaClust: convex window uses hull area" { + let cells = [ + lisa_test_cell("p1", "image", "A", 0.0, 0.0), + lisa_test_cell("p2", "image", "A", 10.0, 0.0), + lisa_test_cell("p3", "image", "B", 0.0, 10.0), + lisa_test_cell("inside", "image", "B", 2.0, 2.0), + ] + let config = @src.LisaConfig::create( + radii=[1.0], + window_kind=@src.lisa_convex_window(), + window_padding=0.0, + edge_correction=false, + ) catch { + _ => abort("convex lisaClust configuration should be valid") + } + let curves = @src.lisa_curves(cells, config~) catch { + _ => abort("convex lisaClust curves should compute") + } + lisa_test_close(curves.windows[0].area, 50.0, 1.0e-12) + assert_eq(curves.windows[0].vertices.length(), 3) +} + +///| +test "lisaClust: radii are capped at half the shortest window span" { + let config = @src.LisaConfig::create( + radii=[1.0, 100.0], + window_kind=@src.lisa_rectangle_window(), + window_padding=0.0, + edge_correction=false, + ) catch { + _ => abort("radius cap configuration should be valid") + } + let curves = @src.lisa_curves(lisa_test_exact_cells(), config~) catch { + _ => abort("radius capped curves should compute") + } + let radii = match curves.effective_radii_for_image("image_1") { + Some(values) => values + None => abort("effective radii should exist") + } + assert_eq(radii[0], 1.0) + lisa_test_close(radii[1], 10.0 / 2.01, 1.0e-12) +} + +///| +test "lisaClust: edge correction changes corner-cell expectation" { + let cells = [ + lisa_test_cell("source", "image", "A", 0.0, 0.0), + lisa_test_cell("near", "image", "B", 1.0, 1.0), + lisa_test_cell("far", "image", "B", 8.0, 8.0), + lisa_test_cell("corner", "image", "C", 10.0, 10.0), + ] + let without = @src.lisa_curves( + cells, + config=@src.LisaConfig::create( + radii=[2.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + edge_correction=false, + window_padding=0.0, + edge_samples=720, + ) catch { + _ => abort("uncorrected edge configuration should be valid") + }, + ) catch { + _ => abort("uncorrected curves should compute") + } + let with_edge = @src.lisa_curves( + cells, + config=@src.LisaConfig::create( + radii=[2.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + edge_correction=true, + window_padding=0.0, + edge_samples=720, + ) catch { + _ => abort("corrected edge configuration should be valid") + }, + ) catch { + _ => abort("corrected curves should compute") + } + let feature = without.feature_index("B", 2.0) + assert_true(with_edge.values[0][feature] > without.values[0][feature]) +} + +///| +test "lisaClust: KDE weighting changes curves under inhomogeneous density" { + let narrow = @src.LisaConfig::create( + radii=[2.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=0.25, + min_density=0.05, + edge_correction=false, + window_padding=0.0, + ) catch { + _ => abort("narrow KDE configuration should be valid") + } + let broad = @src.LisaConfig::create( + radii=[2.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=1000000.0, + min_density=0.05, + edge_correction=false, + window_padding=0.0, + ) catch { + _ => abort("broad KDE configuration should be valid") + } + let narrow_curves = @src.lisa_curves(lisa_test_exact_cells(), config=narrow) catch { + _ => abort("narrow KDE curves should compute") + } + let broad_curves = @src.lisa_curves(lisa_test_exact_cells(), config=broad) catch { + _ => abort("broad KDE curves should compute") + } + let feature = narrow_curves.feature_index("B", 2.0) + assert_true( + (narrow_curves.values[0][feature] - broad_curves.values[0][feature]).abs() > + 1.0e-6, + ) +} + +///| +test "lisaClust: generated curves are finite" { + let curves = @src.lisa_curves( + lisa_test_cells(), + config=lisa_test_config(edge_correction=true), + ) catch { + _ => abort("edge corrected curves should compute") + } + for row in curves.values { + for value in row { + assert_true(value == value) + assert_true(value.abs() <= 1.0e300) + } + } +} + +///| +test "lisaClust: curve lookup returns a defensive copy" { + let curves = @src.lisa_curves(lisa_test_cells(), config=lisa_test_config()) catch { + _ => abort("lisaClust curves should compute") + } + let curve = match curves.curve_for_cell("a1") { + Some(value) => value + None => abort("cell curve should exist") + } + curve[0] = 999.0 + assert_true(curves.values[0][0] != 999.0) + assert_true(curves.curve_for_cell("missing") is None) +} + +///| +test "lisaClust: duplicate cell IDs are rejected" { + let cells = lisa_test_cells() + cells.push(lisa_test_cell("a1", "image_1", "A", 4.0, 4.0)) + let failed = try { + ignore(@src.lisa_curves(cells, config=lisa_test_config())) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: fewer than two cells are rejected" { + let failed = try { + ignore( + @src.lisa_curves( + [lisa_test_cell("only", "image", "A", 0.0, 0.0)], + config=lisa_test_config(), + ), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: one-cell images are rejected" { + let cells = lisa_test_cells() + cells.push(lisa_test_cell("single", "image_2", "A", 2.0, 2.0)) + let failed = try { + ignore(@src.lisa_curves(cells, config=lisa_test_config())) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: degenerate rectangular windows are rejected" { + let cells = [ + lisa_test_cell("one", "image", "A", 0.0, 0.0), + lisa_test_cell("two", "image", "B", 1.0, 0.0), + ] + let failed = try { + ignore(@src.lisa_curves(cells, config=lisa_test_config())) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: collinear convex windows are rejected" { + let cells = [ + lisa_test_cell("one", "image", "A", 0.0, 0.0), + lisa_test_cell("two", "image", "B", 1.0, 1.0), + lisa_test_cell("three", "image", "B", 2.0, 2.0), + ] + let config = @src.LisaConfig::create( + radii=[1.0], + window_kind=@src.lisa_convex_window(), + ) catch { + _ => abort("convex configuration should be valid") + } + let failed = try { + ignore(@src.lisa_curves(cells, config~)) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: complete workflow returns deterministic regions" { + let first = lisa_test_result() + let second = lisa_test_result() + assert_eq(first.labels, second.labels) + assert_eq(first.regions, second.regions) + assert_eq(first.cluster_sizes, second.cluster_sizes) + assert_eq(first.centroids, second.centroids) +} + +///| +test "lisaClust: complete workflow reports fit diagnostics" { + let result = lisa_test_result() + assert_eq(result.n_cells(), 8) + assert_eq(result.n_regions(), 2) + assert_eq(result.labels.length(), 8) + assert_eq(result.centroids.length(), 2) + assert_eq(result.cluster_sizes[0] + result.cluster_sizes[1], 8) + assert_true(result.inertia >= 0.0) + assert_true(result.silhouette >= -1.0) + assert_true(result.silhouette <= 1.0) + assert_true(result.iterations > 0) + assert_true(result.converged) +} + +///| +test "lisaClust: custom region prefix is propagated" { + let result = @src.lisaclust( + lisa_test_cells(), + config=lisa_test_config(region_prefix="domain"), + ) catch { + _ => abort("custom prefix workflow should succeed") + } + for region in result.regions { + assert_true(region.has_prefix("domain_")) + } +} + +///| +test "lisaClust: cell-to-region lookup handles known and unknown IDs" { + let result = lisa_test_result() + assert_true(result.region_for_cell("a1") is Some(_)) + assert_true(result.region_for_cell("missing") is None) +} + +///| +test "lisaClust: region enrichment covers every type-region combination" { + let result = lisa_test_result() + assert_eq(result.enrichment.length(), 6) + assert_eq(result.region_summaries.length(), 2) + for summary in result.region_summaries { + assert_true(summary.size > 0) + assert_true(summary.dominant_cell_type != "") + assert_true(summary.maximum_enrichment >= 0.0) + } +} + +///| +test "lisaClust: region enrichment uses observed over independence expectation" { + let result = lisa_test_result() + let entry = result.enrichment[0] + if entry.expected > 0.0 { + lisa_test_close( + entry.relative_frequency, + entry.observed.to_double() / entry.expected, + 1.0e-12, + ) + } + assert_true( + result.region_enrichment(entry.cell_type, entry.region) is Some(_), + ) + assert_true(result.region_enrichment("missing", entry.region) is None) +} + +///| +test "lisaClust: top enrichments are sorted and filtered" { + let result = lisa_test_result() + let top = result.top_enrichments(limit=3, minimum_relative_frequency=0.0) catch { + _ => abort("top enrichments should succeed") + } + assert_eq(top.length(), 3) + assert_true(top[0].relative_frequency >= top[1].relative_frequency) + assert_true(top[1].relative_frequency >= top[2].relative_frequency) +} + +///| +test "lisaClust: top enrichment validates query controls" { + let result = lisa_test_result() + let bad_limit = try { + ignore(result.top_enrichments(limit=-1)) + false + } catch { + LisaError(_) => true + } + let bad_threshold = try { + ignore(result.top_enrichments(minimum_relative_frequency=-1.0)) + false + } catch { + LisaError(_) => true + } + assert_true(bad_limit) + assert_true(bad_threshold) +} + +///| +test "lisaClust: summary reports cells features and regions" { + let summary = lisa_test_result().summary() + assert_true(summary.contains("lisaClust")) + assert_true(summary.contains("cells=8")) + assert_true(summary.contains("features=6")) + assert_true(summary.contains("regions=2")) +} + +///| +test "lisaClust: cluster count cannot exceed cells" { + let curves = @src.lisa_curves(lisa_test_cells(), config=lisa_test_config()) catch { + _ => abort("lisaClust curves should compute") + } + let config = lisa_test_config(n_clusters=9) + let failed = try { + ignore(@src.lisaclust_from_curves(curves, config~)) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: one-region clustering is supported" { + let result = @src.lisaclust( + lisa_test_cells(), + config=lisa_test_config(n_clusters=1), + ) catch { + _ => abort("single-region lisaClust should succeed") + } + assert_eq(result.n_regions(), 1) + assert_eq(result.cluster_sizes, [8]) + assert_eq(result.silhouette, 0.0) +} + +///| +test "lisaClust: example data has two balanced images and four types" { + let cells = @src.lisaclust_example_data() + assert_eq(cells.length(), 64) + let images : Map[String, Int] = Map([]) + let types : Map[String, Bool] = Map([]) + for cell in cells { + images[cell.image_id] = match images.get(cell.image_id) { + Some(count) => count + 1 + None => 1 + } + types[cell.cell_type] = true + } + assert_eq(images["image_1"], 32) + assert_eq(images["image_2"], 32) + assert_eq(types.length(), 4) +} + +///| +test "lisaClust: example data separates spatial compartments" { + let config = @src.LisaConfig::create( + radii=[2.5, 5.0], + window_kind=@src.lisa_rectangle_window(), + bandwidth=4.0, + edge_correction=true, + window_padding=0.1, + edge_samples=64, + n_clusters=2, + n_starts=5, + seed=11, + ) catch { + _ => abort("example lisaClust configuration should be valid") + } + let result = @src.lisaclust(@src.lisaclust_example_data(), config~) catch { + _ => abort("example lisaClust workflow should succeed") + } + assert_eq(result.n_cells(), 64) + assert_eq(result.n_regions(), 2) + assert_true(result.silhouette > 0.0) +} + +///| +test "lisaClust: SpatialExperiment adapter writes regions to a copy" { + let experiment = lisa_test_experiment() + let output = @src.lisaclust_spatial_experiment( + experiment, + cell_id_key="cellID", + config=lisa_test_config(), + ) catch { + _ => abort("lisaClust SpatialExperiment integration should succeed") + } + assert_eq(output.result.n_cells(), 8) + assert_eq(output.experiment.col_data.length(), 8) + assert_true(output.experiment.col_data[0].contains("region")) + assert_true(!experiment.col_data[0].contains("region")) + assert_eq(output.experiment.metadata["lisaclust_regions"], "2") + assert_eq(output.experiment.metadata["lisaclust_features"], "6") +} + +///| +test "lisaClust: SpatialExperiment adapter supports generated cell IDs" { + let output = @src.lisaclust_spatial_experiment( + lisa_test_experiment(), + config=lisa_test_config(), + ) catch { + _ => abort("generated lisaClust IDs should succeed") + } + assert_eq(output.result.curves.cell_ids[0], "cell_1") + assert_eq(output.result.curves.cell_ids[7], "cell_8") +} + +///| +test "lisaClust: SpatialExperiment adapter supports custom region key" { + let output = @src.lisaclust_spatial_experiment( + lisa_test_experiment(), + cell_id_key="cellID", + region_key="microenvironment", + config=lisa_test_config(), + ) catch { + _ => abort("custom lisaClust region key should succeed") + } + assert_true(output.experiment.col_data[0].contains("microenvironment")) + assert_true(!output.experiment.col_data[0].contains("region")) +} + +///| +test "lisaClust: SpatialExperiment rejects empty col_data" { + let failed = try { + ignore( + @src.lisaclust_spatial_experiment( + @src.SpatialExperiment::new(), + config=lisa_test_config(), + ), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: SpatialExperiment rejects coordinate mismatch" { + let experiment = lisa_test_experiment() + ignore(experiment.spatial_coords.pop()) + let failed = try { + ignore( + @src.lisaclust_spatial_experiment(experiment, config=lisa_test_config()), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: SpatialExperiment rejects missing required columns" { + let experiment = lisa_test_experiment() + ignore(experiment.col_data[0].remove("cellType")) + let failed = try { + ignore( + @src.lisaclust_spatial_experiment(experiment, config=lisa_test_config()), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: SpatialExperiment rejects duplicate supplied cell IDs" { + let experiment = lisa_test_experiment() + experiment.col_data[1]["cellID"] = experiment.col_data[0]["cellID"] + let failed = try { + ignore( + @src.lisaclust_spatial_experiment( + experiment, + cell_id_key="cellID", + config=lisa_test_config(), + ), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} + +///| +test "lisaClust: SpatialExperiment validates key names" { + let failed = try { + ignore( + @src.lisaclust_spatial_experiment( + lisa_test_experiment(), + region_key="", + config=lisa_test_config(), + ), + ) + false + } catch { + LisaError(_) => true + } + assert_true(failed) +} diff --git a/test/moonbit/lowess_test.mbt b/test/moonbit/lowess_test.mbt index 958cd3a7..b39e9c8d 100644 --- a/test/moonbit/lowess_test.mbt +++ b/test/moonbit/lowess_test.mbt @@ -22,7 +22,7 @@ test "lowess_tricube_weight_above_one" { ///| test "lowess_tricube_weight_at_half" { let result = @src.lowess_tricube_weight(0.5) - let expected = (1.0 - 0.5 * 0.5 * 0.5) + let expected = 1.0 - 0.5 * 0.5 * 0.5 let expected_cubed = expected * expected * expected assert_true((result - expected_cubed).abs() < 0.001) } @@ -54,7 +54,7 @@ test "lowess_bisquare_weight_above_one" { ///| test "lowess_bisquare_weight_at_half" { let result = @src.lowess_bisquare_weight(0.5) - let expected = (1.0 - 0.5 * 0.5) + let expected = 1.0 - 0.5 * 0.5 let expected_squared = expected * expected assert_true((result - expected_squared).abs() < 0.001) } @@ -99,7 +99,9 @@ test "lowess_weighted_linear_regression_weighted" { let y = [1.0, 3.0, 5.0, 7.0] let w = [1.0, 1.0, 1.0, 1.0] let (intercept, slope) = @src.lowess_weighted_linear_regression(x, y, w) - let (intercept2, slope2) = @src.lowess_weighted_linear_regression(x, y, [2.0, 2.0, 2.0, 2.0]) + let (intercept2, slope2) = @src.lowess_weighted_linear_regression(x, y, [ + 2.0, 2.0, 2.0, 2.0, + ]) assert_true((intercept - intercept2).abs() < 0.001) assert_true((slope - slope2).abs() < 0.001) } @@ -113,7 +115,11 @@ test "lowess_weighted_linear_regression_empty" { ///| test "lowess_weighted_linear_regression_single_point" { - let (intercept, slope) = @src.lowess_weighted_linear_regression([1.0], [5.0], [1.0]) + let (intercept, slope) = @src.lowess_weighted_linear_regression( + [1.0], + [5.0], + [1.0], + ) assert_eq(intercept, 5.0) assert_eq(slope, 0.0) } diff --git a/test/moonbit/ma_align_test.mbt b/test/moonbit/ma_align_test.mbt index 1199d1f9..3fa6fe44 100644 --- a/test/moonbit/ma_align_test.mbt +++ b/test/moonbit/ma_align_test.mbt @@ -20,7 +20,11 @@ test "ma_aligner_with_params" { ///| test "add_structure" { let a = @src.ma_aligner_new() - let coords : Array[(Double, Double, Double)] = [(1.0, 2.0, 3.0), (4.0, 5.0, 6.0), (7.0, 8.0, 9.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 2.0, 3.0), + (4.0, 5.0, 6.0), + (7.0, 8.0, 9.0), + ] let seq = "ALA" let res = ["ALA", "LEU", "ALA"] let s = @src.add_structure(a, "struct1", coords, seq, res) @@ -31,7 +35,11 @@ test "add_structure" { ///| test "ma_center_coordinates" { - let coords : Array[(Double, Double, Double)] = [(1.0, 2.0, 3.0), (4.0, 5.0, 6.0), (7.0, 8.0, 9.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 2.0, 3.0), + (4.0, 5.0, 6.0), + (7.0, 8.0, 9.0), + ] let centered = @src.ma_center_coordinates(coords) assert_eq(centered.length(), 3) let mut sum_x = 0.0 @@ -51,8 +59,16 @@ test "ma_center_coordinates" { ///| test "ma_compute_distance_matrix" { - let coords1 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] - let coords2 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] let dm = @src.ma_compute_distance_matrix(coords1, coords2) assert_eq(dm.length(), 3) assert_eq(dm[0].length(), 3) @@ -60,30 +76,51 @@ test "ma_compute_distance_matrix" { ///| test "ma_align_pairwise_identical" { - let coords : Array[(Double, Double, Double)] = [(1.0, 0.0, 0.0), (0.0, 1.0, 0.0), (0.0, 0.0, 1.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + (0.0, 0.0, 1.0), + ] let (rotation, rmsd) = @src.ma_align_pairwise(coords, coords) assert_true(rmsd < 0.01) } ///| test "ma_align_pairwise_different_scale" { - let coords1 : Array[(Double, Double, Double)] = [(1.0, 0.0, 0.0), (0.0, 1.0, 0.0), (0.0, 0.0, 1.0)] - let coords2 : Array[(Double, Double, Double)] = [(2.0, 0.0, 0.0), (0.0, 2.0, 0.0), (0.0, 0.0, 2.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + (0.0, 0.0, 1.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (2.0, 0.0, 0.0), + (0.0, 2.0, 0.0), + (0.0, 0.0, 2.0), + ] let (rotation, rmsd) = @src.ma_align_pairwise(coords1, coords2) assert_true(rmsd < 5.0) } ///| test "compute_rmsd_identical" { - let coords : Array[(Double, Double, Double)] = [(1.0, 2.0, 3.0), (4.0, 5.0, 6.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 2.0, 3.0), + (4.0, 5.0, 6.0), + ] let rmsd = @src.compute_rmsd(coords, coords) assert_true(rmsd < 0.01) } ///| test "compute_rmsd_different" { - let coords1 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0)] - let coords2 : Array[(Double, Double, Double)] = [(10.0, 0.0, 0.0), (11.0, 0.0, 0.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (10.0, 0.0, 0.0), + (11.0, 0.0, 0.0), + ] let rmsd = @src.compute_rmsd(coords1, coords2) assert_true(rmsd > 5.0) } @@ -91,7 +128,11 @@ test "compute_rmsd_different" { ///| test "align_structures_single" { let a = @src.ma_aligner_new() - let coords : Array[(Double, Double, Double)] = [(1.0, 2.0, 3.0), (4.0, 5.0, 6.0), (7.0, 8.0, 9.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 2.0, 3.0), + (4.0, 5.0, 6.0), + (7.0, 8.0, 9.0), + ] let seq = "ALA" let res = ["ALA", "LEU", "ALA"] let s = @src.add_structure(a, "s1", coords, seq, res) @@ -102,8 +143,16 @@ test "align_structures_single" { ///| test "align_structures_two" { let a = @src.ma_aligner_new() - let coords1 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] - let coords2 : Array[(Double, Double, Double)] = [(10.0, 10.0, 0.0), (11.0, 10.0, 0.0), (10.0, 11.0, 0.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (10.0, 10.0, 0.0), + (11.0, 10.0, 0.0), + (10.0, 11.0, 0.0), + ] let s1 = @src.add_structure(a, "s1", coords1, "AAA", ["ALA", "ALA", "ALA"]) let s2 = @src.add_structure(a, "s2", coords2, "BBB", ["ALA", "ALA", "ALA"]) let result = @src.align_structures(a, [s1, s2]) @@ -141,8 +190,16 @@ test "build_conservation_profile_different" { ///| test "get_conservation" { let a = @src.ma_aligner_new() - let coords1 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] - let coords2 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] let s1 = @src.add_structure(a, "s1", coords1, "AAA", ["ALA", "ALA", "ALA"]) let s2 = @src.add_structure(a, "s2", coords2, "AAA", ["ALA", "ALA", "ALA"]) let result = @src.align_structures(a, [s1, s2]) @@ -153,8 +210,14 @@ test "get_conservation" { ///| test "get_rmsd_matrix" { let a = @src.ma_aligner_new() - let coords1 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0)] - let coords2 : Array[(Double, Double, Double)] = [(0.0, 0.0, 0.0), (1.0, 0.0, 0.0)] + let coords1 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + ] + let coords2 : Array[(Double, Double, Double)] = [ + (0.0, 0.0, 0.0), + (1.0, 0.0, 0.0), + ] let s1 = @src.add_structure(a, "s1", coords1, "AA", ["ALA", "ALA"]) let s2 = @src.add_structure(a, "s2", coords2, "AA", ["ALA", "ALA"]) let result = @src.align_structures(a, [s1, s2]) @@ -166,8 +229,11 @@ test "get_rmsd_matrix" { ///| test "apply_rotation_identity" { - let coords : Array[(Double, Double, Double)] = [(1.0, 0.0, 0.0), (0.0, 1.0, 0.0)] + let coords : Array[(Double, Double, Double)] = [ + (1.0, 0.0, 0.0), + (0.0, 1.0, 0.0), + ] let identity : (Double, Double, Double) = (0.0, 0.0, 0.0) let _result = @src.apply_rotation(coords, identity) assert_true(true) -} \ No newline at end of file +} diff --git a/test/moonbit/maf_test.mbt b/test/moonbit/maf_test.mbt index 1b81bc96..cc38054e 100644 --- a/test/moonbit/maf_test.mbt +++ b/test/moonbit/maf_test.mbt @@ -137,12 +137,12 @@ test "maf_alignment_total_length" { let seq1 = @src.new_maf_sequence("seq1", 0, 10, "+", 100, "ACGTACGTAC") block1 = @src.maf_block_add_seq(block1, seq1) ali = @src.maf_add_block(ali, block1) - + let mut block2 = @src.new_maf_block() let seq2 = @src.new_maf_sequence("seq1", 10, 5, "+", 100, "GTACG") block2 = @src.maf_block_add_seq(block2, seq2) ali = @src.maf_add_block(ali, block2) - + assert_eq(@src.maf_total_length(ali), 15) } @@ -156,7 +156,7 @@ test "maf_alignment_seq_names" { block = @src.maf_block_add_seq(block, seq1) block = @src.maf_block_add_seq(block, seq2) ali = @src.maf_add_block(ali, block) - + let names = @src.maf_all_seq_names(ali) assert_eq(names.length(), 2) assert_true(names.contains("human.chr1")) @@ -224,13 +224,19 @@ test "maf_select_seqs" { /// Test maf_filter_by_length. test "maf_filter_by_length" { let mut ali = @src.new_maf_alignment() - + let mut block1 = @src.new_maf_block() block1 = @src.maf_block_set_score(block1, Some(100.0)) - block1 = @src.maf_block_add_seq(block1, @src.new_maf_sequence("s1", 0, 100, "+", 1000, "ACGTACGTACGTACGTACGT")) - block1 = @src.maf_block_add_seq(block1, @src.new_maf_sequence("s2", 0, 100, "+", 1000, "ACGTACGTACGTACGTACGT")) + block1 = @src.maf_block_add_seq( + block1, + @src.new_maf_sequence("s1", 0, 100, "+", 1000, "ACGTACGTACGTACGTACGT"), + ) + block1 = @src.maf_block_add_seq( + block1, + @src.new_maf_sequence("s2", 0, 100, "+", 1000, "ACGTACGTACGTACGTACGT"), + ) ali = @src.maf_add_block(ali, block1) - + let filtered = @src.maf_filter_by_length(ali, 5) assert_eq(@src.maf_num_blocks(filtered), 1) } @@ -241,10 +247,16 @@ test "maf_stats" { let mut ali = @src.new_maf_alignment() let mut block = @src.new_maf_block() block = @src.maf_block_set_score(block, Some(100.0)) - block = @src.maf_block_add_seq(block, @src.new_maf_sequence("s1", 0, 10, "+", 1000, "ACGTACGTAC")) - block = @src.maf_block_add_seq(block, @src.new_maf_sequence("s2", 0, 10, "+", 1000, "ACGTACGTAC")) + block = @src.maf_block_add_seq( + block, + @src.new_maf_sequence("s1", 0, 10, "+", 1000, "ACGTACGTAC"), + ) + block = @src.maf_block_add_seq( + block, + @src.new_maf_sequence("s2", 0, 10, "+", 1000, "ACGTACGTAC"), + ) ali = @src.maf_add_block(ali, block) - + let stats = @src.maf_compute_stats(ali) assert_eq(stats.num_blocks, 1) assert_eq(stats.total_length, 10) @@ -284,21 +296,33 @@ test "maf_strand_info" { /// Test maf_merge_blocks. test "maf_merge_blocks" { let mut ali = @src.new_maf_alignment() - + let mut block1 = @src.new_maf_block() block1 = @src.maf_block_set_score(block1, Some(50.0)) - block1 = @src.maf_block_add_seq(block1, @src.new_maf_sequence("s1", 0, 5, "+", 1000, "ACGTG")) - block1 = @src.maf_block_add_seq(block1, @src.new_maf_sequence("s2", 0, 5, "+", 1000, "ACGTG")) - + block1 = @src.maf_block_add_seq( + block1, + @src.new_maf_sequence("s1", 0, 5, "+", 1000, "ACGTG"), + ) + block1 = @src.maf_block_add_seq( + block1, + @src.new_maf_sequence("s2", 0, 5, "+", 1000, "ACGTG"), + ) + let mut block2 = @src.new_maf_block() block2 = @src.maf_block_set_score(block2, Some(30.0)) - block2 = @src.maf_block_add_seq(block2, @src.new_maf_sequence("s1", 5, 5, "+", 1000, "TGCAC")) - block2 = @src.maf_block_add_seq(block2, @src.new_maf_sequence("s2", 5, 5, "+", 1000, "TGCAC")) - + block2 = @src.maf_block_add_seq( + block2, + @src.new_maf_sequence("s1", 5, 5, "+", 1000, "TGCAC"), + ) + block2 = @src.maf_block_add_seq( + block2, + @src.new_maf_sequence("s2", 5, 5, "+", 1000, "TGCAC"), + ) + ali = @src.maf_add_block(ali, block1) ali = @src.maf_add_block(ali, block2) - + let merged = @src.maf_merge_blocks(ali) // Should merge blocks with same sequences assert_true(@src.maf_num_blocks(merged) <= @src.maf_num_blocks(ali)) -} \ No newline at end of file +} diff --git a/test/moonbit/maftools_test.mbt b/test/moonbit/maftools_test.mbt index 54ff0f48..70a502e6 100644 --- a/test/moonbit/maftools_test.mbt +++ b/test/moonbit/maftools_test.mbt @@ -1,16 +1,9 @@ ///| test "maftools_create_mutation" { let mutation = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + assert_eq(mutation.hugo_symbol, "TP53") assert_eq(mutation.chromosome, "chr17") assert_eq(mutation.start_position, 7577121) @@ -24,16 +17,9 @@ test "maftools_create_mutation" { ///| test "maftools_snv_detection" { let snv = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + assert_true(snv.is_snv()) assert_false(snv.is_indel()) assert_false(snv.is_complex()) @@ -42,16 +28,9 @@ test "maftools_snv_detection" { ///| test "maftools_indel_detection" { let insertion = @src.MAFMutation::new( - "BRCA1", - "chr17", - 43091134, - 43091135, - "CT", - "CTT", - "Sample1", - "Frame_Shift_Ins" + "BRCA1", "chr17", 43091134, 43091135, "CT", "CTT", "Sample1", "Frame_Shift_Ins", ) - + assert_false(insertion.is_snv()) assert_true(insertion.is_indel()) assert_false(insertion.is_complex()) @@ -60,16 +39,9 @@ test "maftools_indel_detection" { ///| test "maftools_transition_detection" { let transition = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + assert_true(transition.is_transition()) assert_false(transition.is_transversion()) } @@ -77,16 +49,9 @@ test "maftools_transition_detection" { ///| test "maftools_transversion_detection" { let transversion = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "A", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "A", "Sample1", "Missense_Mutation", ) - + assert_false(transversion.is_transition()) assert_true(transversion.is_transversion()) } @@ -94,32 +59,18 @@ test "maftools_transversion_detection" { ///| test "maftools_maf_data_operations" { let mut maf = @src.MAFData::new() - + let mutation1 = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + let mutation2 = @src.MAFMutation::new( - "BRCA1", - "chr17", - 43091134, - 43091134, - "A", - "G", - "Sample2", - "Nonsense_Mutation" + "BRCA1", "chr17", 43091134, 43091134, "A", "G", "Sample2", "Nonsense_Mutation", ) - + maf = maf.add_mutation(mutation1) maf = maf.add_mutation(mutation2) - + assert_eq(maf.count_mutations(), 2) assert_eq(maf.count_unique_genes(), 2) assert_eq(maf.count_unique_samples(), 2) @@ -128,32 +79,18 @@ test "maftools_maf_data_operations" { ///| test "maftools_filter_genes" { let mut maf = @src.MAFData::new() - + let mutation1 = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + let mutation2 = @src.MAFMutation::new( - "BRCA1", - "chr17", - 43091134, - 43091134, - "A", - "G", - "Sample2", - "Nonsense_Mutation" + "BRCA1", "chr17", 43091134, 43091134, "A", "G", "Sample2", "Nonsense_Mutation", ) - + maf = maf.add_mutation(mutation1) maf = maf.add_mutation(mutation2) - + let filtered = maf.filter_genes(["TP53"]) assert_eq(filtered.count_mutations(), 1) assert_eq(filtered.count_unique_genes(), 1) @@ -162,32 +99,18 @@ test "maftools_filter_genes" { ///| test "maftools_filter_samples" { let mut maf = @src.MAFData::new() - + let mutation1 = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + let mutation2 = @src.MAFMutation::new( - "BRCA1", - "chr17", - 43091134, - 43091134, - "A", - "G", - "Sample2", - "Nonsense_Mutation" + "BRCA1", "chr17", 43091134, 43091134, "A", "G", "Sample2", "Nonsense_Mutation", ) - + maf = maf.add_mutation(mutation1) maf = maf.add_mutation(mutation2) - + let filtered = maf.filter_samples(["Sample1"]) assert_eq(filtered.count_mutations(), 1) } @@ -195,46 +118,25 @@ test "maftools_filter_samples" { ///| test "maftools_mutation_spectrum" { let mut maf = @src.MAFData::new() - + let snv1 = @src.MAFMutation::new( - "TP53", - "chr17", - 7577121, - 7577121, - "C", - "T", - "Sample1", - "Missense_Mutation" + "TP53", "chr17", 7577121, 7577121, "C", "T", "Sample1", "Missense_Mutation", ) - + let snv2 = @src.MAFMutation::new( - "BRCA1", - "chr17", - 43091134, - 43091134, - "A", - "G", - "Sample1", - "Missense_Mutation" + "BRCA1", "chr17", 43091134, 43091134, "A", "G", "Sample1", "Missense_Mutation", ) - + let indel = @src.MAFMutation::new( - "EGFR", - "chr7", - 55086714, - 55086715, - "CT", - "CTT", - "Sample2", - "Frame_Shift_Ins" + "EGFR", "chr7", 55086714, 55086715, "CT", "CTT", "Sample2", "Frame_Shift_Ins", ) - + maf = maf.add_mutation(snv1) maf = maf.add_mutation(snv2) maf = maf.add_mutation(indel) - + let spectrum = @src.calculate_mutation_spectrum(maf) - + assert_eq(spectrum.snv_count, 2) assert_eq(spectrum.indel_count, 1) assert_eq(spectrum.complex_count, 0) @@ -243,7 +145,7 @@ test "maftools_mutation_spectrum" { ///| test "maftools_tmb_calculation" { let mut maf = @src.MAFData::new() - + let mut i = 0 while i < 100 { let mutation = @src.MAFMutation::new( @@ -254,17 +156,17 @@ test "maftools_tmb_calculation" { "C", "T", "Sample1", - "Missense_Mutation" + "Missense_Mutation", ) maf = maf.add_mutation(mutation) i = i + 1 } - + let tmb_result = @src.calculate_tmb(maf, 3.0e7) - + assert_eq(tmb_result.total_mutations, 100) assert_eq(tmb_result.coding_region_size, 3.0e7) - + // TMB should be 100 / 30000000 = 0.00000333... let expected_tmb = 100.0 / 3.0e7 let diff = tmb_result.tmb - expected_tmb @@ -274,19 +176,35 @@ test "maftools_tmb_calculation" { ///| test "maftools_co_occurrence_analysis" { let mut maf = @src.MAFData::new() - + // Sample1 has both TP53 and BRCA1 mutations - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("BRCA1", "chr17", 200, 200, "A", "G", "Sample1", "Missense_Mutation")) - + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "BRCA1", "chr17", 200, 200, "A", "G", "Sample1", "Missense_Mutation", + ), + ) + // Sample2 has only TP53 mutation - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 300, 300, "C", "T", "Sample2", "Missense_Mutation")) - + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 300, 300, "C", "T", "Sample2", "Missense_Mutation", + ), + ) + // Sample3 has only BRCA1 mutation - maf = maf.add_mutation(@src.MAFMutation::new("BRCA1", "chr17", 400, 400, "A", "G", "Sample3", "Missense_Mutation")) - + maf = maf.add_mutation( + @src.MAFMutation::new( + "BRCA1", "chr17", 400, 400, "A", "G", "Sample3", "Missense_Mutation", + ), + ) + let result = @src.analyze_co_occurrence(maf, "TP53", "BRCA1") - + assert_eq(result.co_occurrence, 1) assert_eq(result.mutual_exclusivity, 2) } @@ -294,7 +212,7 @@ test "maftools_co_occurrence_analysis" { ///| test "maftools_create_example" { let maf = @src.create_example_maf() - + assert_true(maf.count_mutations() > 0) assert_true(maf.count_unique_genes() > 0) assert_true(maf.count_unique_samples() > 0) @@ -304,7 +222,7 @@ test "maftools_create_example" { test "maftools_summarize" { let maf = @src.create_example_maf() let summary = @src.summarize_maf(maf) - + assert_true(summary.contains("MAF Summary")) assert_true(summary.contains("Total mutations:")) assert_true(summary.contains("TMB:")) @@ -314,7 +232,7 @@ test "maftools_summarize" { test "maftools_oncoplot_data" { let maf = @src.create_example_maf() let oncoplot = @src.generate_oncoplot_data(maf, 5) - + assert_true(oncoplot.genes.length() <= 5) assert_true(oncoplot.samples.length() > 0) assert_true(oncoplot.mutation_matrix.length() > 0) @@ -323,13 +241,25 @@ test "maftools_oncoplot_data" { ///| test "maftools_get_gene_mutation_counts" { let mut maf = @src.MAFData::new() - - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 200, 200, "A", "G", "Sample2", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("BRCA1", "chr17", 300, 300, "C", "T", "Sample1", "Missense_Mutation")) - + + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 200, 200, "A", "G", "Sample2", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "BRCA1", "chr17", 300, 300, "C", "T", "Sample1", "Missense_Mutation", + ), + ) + let counts = @src.get_gene_mutation_counts(maf) - + assert_eq(counts.get("TP53").unwrap_or(0), 2) assert_eq(counts.get("BRCA1").unwrap_or(0), 1) } @@ -337,13 +267,25 @@ test "maftools_get_gene_mutation_counts" { ///| test "maftools_get_sample_mutation_counts" { let mut maf = @src.MAFData::new() - - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("BRCA1", "chr17", 200, 200, "A", "G", "Sample1", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("EGFR", "chr7", 300, 300, "C", "T", "Sample2", "Missense_Mutation")) - + + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "BRCA1", "chr17", 200, 200, "A", "G", "Sample1", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "EGFR", "chr7", 300, 300, "C", "T", "Sample2", "Missense_Mutation", + ), + ) + let counts = @src.get_sample_mutation_counts(maf) - + assert_eq(counts.get("Sample1").unwrap_or(0), 2) assert_eq(counts.get("Sample2").unwrap_or(0), 1) } @@ -351,13 +293,25 @@ test "maftools_get_sample_mutation_counts" { ///| test "maftools_filter_variant_type" { let mut maf = @src.MAFData::new() - - maf = maf.add_mutation(@src.MAFMutation::new("TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("BRCA1", "chr17", 200, 200, "A", "G", "Sample2", "Nonsense_Mutation")) - maf = maf.add_mutation(@src.MAFMutation::new("EGFR", "chr7", 300, 300, "C", "T", "Sample3", "Missense_Mutation")) - + + maf = maf.add_mutation( + @src.MAFMutation::new( + "TP53", "chr17", 100, 100, "C", "T", "Sample1", "Missense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "BRCA1", "chr17", 200, 200, "A", "G", "Sample2", "Nonsense_Mutation", + ), + ) + maf = maf.add_mutation( + @src.MAFMutation::new( + "EGFR", "chr7", 300, 300, "C", "T", "Sample3", "Missense_Mutation", + ), + ) + let filtered = maf.filter_variant_type("Nonsense") - + assert_eq(filtered.count_mutations(), 1) assert_eq(filtered.get_mutation_gene(0), "BRCA1") } @@ -365,9 +319,9 @@ test "maftools_filter_variant_type" { ///| test "maftools_parse_maf_content" { let content = "Hugo_Symbol\tChromosome\tStart_Position\tEnd_Position\tReference_Allele\tTumor_Seq_Allele1\tTumor_Sample_Barcode\tVariant_Classification\nTP53\tchr17\t7577121\t7577121\tC\tT\tSample1\tMissense_Mutation\nBRCA1\tchr17\t43091134\t43091134\tA\tG\tSample2\tNonsense_Mutation" - + let maf = @src.parse_maf_content(content) - + assert_eq(maf.count_mutations(), 2) assert_eq(maf.get_mutation_gene(0), "TP53") assert_eq(maf.get_mutation_gene(1), "BRCA1") diff --git a/test/moonbit/markov_test.mbt b/test/moonbit/markov_test.mbt index e7e76022..4309a6d2 100644 --- a/test/moonbit/markov_test.mbt +++ b/test/moonbit/markov_test.mbt @@ -1,18 +1,20 @@ ///| /// Test file for markov module. - fn make_dna_seqs() -> Array[String] { ["ACGTACGT", "CGCGCGTA", "TTTAAAAC", "GGCCTTAA", "ATGCATGC"] } +///| fn make_cpg_seqs() -> Array[String] { ["CGCGCGCG", "ACGTACGT", "CGATCGAT", "GCGCGCGC", "TCGATCGA"] } +///| fn make_bg_seqs() -> Array[String] { ["ATATATAT", "TTAAAATT", "GGGGGGGG", "CCCCCCCC", "AATTCCGG"] } +///| test "markov_chain_type_helpers" { let a = @src.markov_first_order() let b = @src.markov_second_order() @@ -22,6 +24,7 @@ test "markov_chain_type_helpers" { assert_true(b != c) } +///| test "markov_model_new_default" { let m = @src.MarkovModel::new() assert_eq(m.order, 1) @@ -29,6 +32,7 @@ test "markov_model_new_default" { assert_eq(m.pseudo_count, 1.0) } +///| test "markov_model_default_dna" { let m = @src.MarkovModel::default_dna() assert_eq(m.order, 1) @@ -37,12 +41,14 @@ test "markov_model_default_dna" { assert_eq(m.states[1], "C") } +///| test "markov_model_default_protein" { let m = @src.MarkovModel::default_protein() assert_eq(m.order, 1) assert_eq(m.states.length(), 20) } +///| test "markov_model_set_order" { let m = @src.MarkovModel::new() let m2 = m.set_order(2) @@ -52,14 +58,16 @@ test "markov_model_set_order" { assert_eq(m.order, 1) } +///| test "markov_model_set_states" { let m = @src.MarkovModel::new() let states = ["X", "Y"] - let m2 = m.set_states(states=states) + let m2 = m.set_states(states~) assert_eq(m2.states.length(), 2) assert_eq(m2.states[0], "X") } +///| test "markov_model_set_chain_type" { let m = @src.MarkovModel::new() let m2 = m.set_chain_type(val=@src.markov_second_order()) @@ -68,6 +76,7 @@ test "markov_model_set_chain_type" { assert_eq(m3.order, 3) } +///| test "markov_model_set_pseudo_count" { let m = @src.MarkovModel::new() let m2 = m.set_pseudo_count(0.5) @@ -75,6 +84,7 @@ test "markov_model_set_pseudo_count" { assert_eq(m.pseudo_count, 1.0) } +///| test "markov_build_model_dna_order1" { let seqs = make_dna_seqs() let m = @src.markov_build_model(seqs, order=1) @@ -83,6 +93,7 @@ test "markov_build_model_dna_order1" { assert_true(m.transition_probs.length() > 0) } +///| test "markov_build_model_different_orders" { let seqs = make_dna_seqs() let m1 = @src.markov_build_model(seqs, order=1) @@ -94,6 +105,7 @@ test "markov_build_model_different_orders" { assert_true(c1 != c2 || c2 != c3) } +///| test "markov_score_sequence_finite_negative" { let seqs = make_dna_seqs() let m = @src.markov_build_model(seqs, order=1) @@ -102,6 +114,7 @@ test "markov_score_sequence_finite_negative" { assert_true(score > -1000.0) } +///| test "markov_score_per_base_length" { let seqs = make_dna_seqs() let m1 = @src.markov_build_model(seqs, order=1) @@ -112,6 +125,7 @@ test "markov_score_per_base_length" { assert_eq(s2.length(), 8 - 2) } +///| test "markov_generate_sequence_length_and_alphabet" { let seqs = make_dna_seqs() let m = @src.markov_build_model(seqs, order=1) @@ -131,6 +145,7 @@ test "markov_generate_sequence_length_and_alphabet" { } } +///| test "markov_log_odds_positive_cpg" { let cpg_seqs = make_cpg_seqs() let bg_seqs = make_bg_seqs() @@ -141,6 +156,7 @@ test "markov_log_odds_positive_cpg" { assert_true(lo > -100.0) } +///| test "markov_stationary_sums_to_one" { let seqs = make_dna_seqs() let m = @src.markov_build_model(seqs, order=1) @@ -156,22 +172,35 @@ test "markov_stationary_sums_to_one" { assert_true(total < 1.1) } +///| test "markov_pseudo_count_affects_unknowns" { let seqs = ["AAAAA", "CCCCCCC"] let sts = ["A", "C", "G", "T"] - let m_low = @src.markov_build_model(seqs, order=1, states=sts, pseudo_count=0.1) - let m_high = @src.markov_build_model(seqs, order=1, states=sts, pseudo_count=10.0) + let m_low = @src.markov_build_model( + seqs, + order=1, + states=sts, + pseudo_count=0.1, + ) + let m_high = @src.markov_build_model( + seqs, + order=1, + states=sts, + pseudo_count=10.0, + ) let s_low = @src.markov_score_sequence(m_low, "GGGG") let s_high = @src.markov_score_sequence(m_high, "GGGG") assert_true(s_high > s_low) } +///| test "markov_build_model_infers_states" { let seqs = ["AABBAABB", "BBBBAAAA"] let m = @src.markov_build_model(seqs) assert_true(m.states.length() >= 2) } +///| test "markov_short_sequence_score" { let seqs = make_dna_seqs() let m = @src.markov_build_model(seqs, order=1) diff --git a/test/moonbit/mast_advanced_test.mbt b/test/moonbit/mast_advanced_test.mbt new file mode 100644 index 00000000..86e1fe76 --- /dev/null +++ b/test/moonbit/mast_advanced_test.mbt @@ -0,0 +1,913 @@ +///| +fn mast_adv_test_group_data() -> @src.MastAdvancedData { + @src.mast_advanced_from_groups( + [ + [0.0, 1.0, 0.0, 1.1, 0.0, 2.0, 0.0, 2.1], + [0.0, 0.0, 1.0, 1.1, 2.0, 2.1, 2.2, 2.3], + [1.0, 1.1, 0.9, 1.2, 1.0, 1.1, 0.9, 1.2], + ], + ["A", "A", "A", "A", "B", "B", "B", "B"], + reference="A", + gene_names=["continuous", "detection", "stable"], + cell_names=["c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8"], + include_cdr=false, + ) catch { + _ => abort("valid grouped MAST data should build") + } +} + +///| +fn mast_adv_test_example_result() -> @src.MastAdvancedResult { + let (data, tested) = @src.mast_advanced_example() catch { + _ => abort("MAST advanced example should build") + } + @src.mast_advanced_zlm(data, tested) catch { + _ => abort("MAST advanced example should fit") + } +} + +///| +fn mast_adv_test_sce() -> @src.SingleCellExperiment { + let (data, _) = @src.mast_advanced_example() catch { + _ => abort("MAST advanced example should build") + } + let experiment = @src.SingleCellExperiment::new( + data.expression, + data.gene_names, + data.cell_names, + ) + experiment.assays["logcounts"] = data.expression.map(fn(row) { row.copy() }) + let groups : Array[String] = [] + for cell in 0.. @src.MastAdvancedResult { + let data = @src.mast_advanced_from_groups( + [ + [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + [1.0, 1.1, 0.9, 1.2, 1.0, 1.1, 2.0, 2.1, 1.9, 2.2, 2.0, 2.1], + [0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0, 0.0], + [0.0, 1.0, 0.0, 1.1, 0.0, 1.2, 2.0, 0.0, 2.1, 0.0, 2.2, 0.0], + ], + ["A", "A", "A", "A", "A", "A", "B", "B", "B", "B", "B", "B"], + reference="A", + gene_names=["zero", "all_positive", "sparse", "mixed"], + include_cdr=false, + ) catch { + _ => abort("valid MAST NA test data should build") + } + let config = @src.MastAdvancedConfig::create(empirical_bayes=false) catch { + _ => abort("valid MAST configuration should build") + } + @src.mast_advanced_zlm(data, [1], config~) catch { + _ => abort("MAST NA test data should fit") + } +} + +///| +test "MAST advanced default configuration matches zlm workflow" { + let config = @src.MastAdvancedConfig::default() + assert_eq(config.max_iterations, 100) + assert_eq(config.cauchy_prior_scale, 2.5) + assert_true(config.empirical_bayes) + assert_eq(config.ebayes_use_full_model, false) + assert_eq(config.minimum_detected, 2) + assert_eq(config.fdr_threshold, 0.05) +} + +///| +test "MAST advanced configuration preserves explicit controls" { + let config = @src.MastAdvancedConfig::create( + max_iterations=25, + tolerance=1.0e-6, + cauchy_prior_scale=1.5, + ridge=1.0e-5, + empirical_bayes=false, + ebayes_use_full_model=true, + maximum_prior_df=50.0, + minimum_detected=3, + minimum_probability=1.0e-7, + fdr_threshold=0.1, + ) catch { + _ => abort("valid explicit MAST configuration should build") + } + assert_eq(config.max_iterations, 25) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.cauchy_prior_scale, 1.5) + assert_eq(config.ridge, 1.0e-5) + assert_eq(config.empirical_bayes, false) + assert_true(config.ebayes_use_full_model) + assert_eq(config.maximum_prior_df, 50.0) + assert_eq(config.minimum_detected, 3) + assert_eq(config.minimum_probability, 1.0e-7) + assert_eq(config.fdr_threshold, 0.1) +} + +///| +test "MAST advanced rejects invalid iteration and tolerance controls" { + let iterations = try { + ignore(@src.MastAdvancedConfig::create(max_iterations=0)) + false + } catch { + _ => true + } + let tolerance = try { + ignore(@src.MastAdvancedConfig::create(tolerance=0.0)) + false + } catch { + _ => true + } + assert_true(iterations) + assert_true(tolerance) +} + +///| +test "MAST advanced rejects invalid prior and ridge controls" { + let scale = try { + ignore(@src.MastAdvancedConfig::create(cauchy_prior_scale=0.0)) + false + } catch { + _ => true + } + let ridge = try { + ignore(@src.MastAdvancedConfig::create(ridge=-1.0)) + false + } catch { + _ => true + } + let prior_df = try { + ignore(@src.MastAdvancedConfig::create(maximum_prior_df=0.0)) + false + } catch { + _ => true + } + assert_true(scale) + assert_true(ridge) + assert_true(prior_df) +} + +///| +test "MAST advanced rejects invalid probability controls" { + let minimum = try { + ignore(@src.MastAdvancedConfig::create(minimum_probability=0.5)) + false + } catch { + _ => true + } + let fdr = try { + ignore(@src.MastAdvancedConfig::create(fdr_threshold=1.1)) + false + } catch { + _ => true + } + let detected = try { + ignore(@src.MastAdvancedConfig::create(minimum_detected=0)) + false + } catch { + _ => true + } + assert_true(minimum) + assert_true(fdr) + assert_true(detected) +} + +///| +test "MAST advanced grouped constructor uses feature by cell orientation" { + let data = mast_adv_test_group_data() + assert_eq(data.n_genes, 3) + assert_eq(data.n_cells, 8) + assert_eq(data.expression[0].length(), 8) + assert_eq(data.gene_names, ["continuous", "detection", "stable"]) + assert_eq(data.cell_names[7], "c8") +} + +///| +test "MAST advanced grouped constructor uses treatment coding" { + let data = mast_adv_test_group_data() + assert_eq(data.coefficient_names, ["(Intercept)", "group:B"]) + assert_eq(data.design[0], [1.0, 0.0]) + assert_eq(data.design[4], [1.0, 1.0]) +} + +///| +test "MAST advanced grouped constructor honors reference level" { + let data = @src.mast_advanced_from_groups( + [[1.0, 2.0, 3.0, 4.0]], + ["A", "A", "B", "B"], + reference="B", + include_cdr=false, + ) catch { + _ => abort("valid reference level should build") + } + assert_eq(data.coefficient_names, ["(Intercept)", "group:A"]) + assert_eq(data.design[0], [1.0, 1.0]) + assert_eq(data.design[2], [1.0, 0.0]) +} + +///| +test "MAST advanced constructor appends cell detection rate" { + let data = @src.MastAdvancedData::create( + [[1.0, 1.0, 1.0, 1.0], [0.0, 1.0, 0.0, 1.0], [0.0, 0.0, 1.0, 1.0]], + [[1.0], [1.0], [1.0], [1.0]], + ["(Intercept)"], + ) catch { + _ => abort("variable CDR should produce a full-rank design") + } + assert_eq(data.coefficient_names, ["(Intercept)", "cngeneson"]) + assert_eq(data.cdr, [1.0 / 3.0, 2.0 / 3.0, 2.0 / 3.0, 1.0]) + assert_eq(data.design[3], [1.0, 1.0]) +} + +///| +test "MAST advanced constructor supports arbitrary design matrices" { + let data = @src.MastAdvancedData::create( + [[1.0, 1.2, 1.4, 2.0, 2.2, 2.4]], + [ + [1.0, 0.0, -1.0], + [1.0, 0.0, 0.0], + [1.0, 0.0, 1.0], + [1.0, 1.0, -1.0], + [1.0, 1.0, 0.0], + [1.0, 1.0, 1.0], + ], + ["(Intercept)", "condition", "batch_score"], + include_cdr=false, + ) catch { + _ => abort("full-rank arbitrary design should build") + } + assert_eq(data.n_coefficients, 3) + assert_eq(data.design[5][2], 1.0) +} + +///| +test "MAST advanced constructor defensively copies inputs" { + let expression = [[1.0, 0.0, 2.0, 0.0]] + let design = [[1.0], [1.0], [1.0], [1.0]] + let genes = ["g1"] + let cells = ["c1", "c2", "c3", "c4"] + let data = @src.MastAdvancedData::create( + expression, + design, + ["(Intercept)"], + gene_names=genes, + cell_names=cells, + include_cdr=false, + ) catch { + _ => abort("valid MAST data should build") + } + expression[0][0] = 99.0 + design[0][0] = 99.0 + genes[0] = "changed" + cells[0] = "changed" + assert_eq(data.expression[0][0], 1.0) + assert_eq(data.design[0][0], 1.0) + assert_eq(data.gene_names[0], "g1") + assert_eq(data.cell_names[0], "c1") +} + +///| +test "MAST advanced constructor rejects empty and ragged expression" { + let empty = try { + ignore(@src.MastAdvancedData::create([], [], [], include_cdr=false)) + false + } catch { + _ => true + } + let ragged = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0], [3.0]], + [[1.0], [1.0]], + ["intercept"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(empty) + assert_true(ragged) +} + +///| +test "MAST advanced constructor rejects invalid expression values" { + let negative = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, -1.0]], + [[1.0], [1.0]], + ["intercept"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let non_finite = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, Double::nan()]], + [[1.0], [1.0]], + ["intercept"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(negative) + assert_true(non_finite) +} + +///| +test "MAST advanced constructor validates design dimensions" { + let rows = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0]], + [[1.0]], + ["intercept"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let columns = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0]], + [[1.0], [1.0, 0.0]], + ["intercept"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let names = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0]], + [[1.0], [1.0]], + [], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(rows) + assert_true(columns) + assert_true(names) +} + +///| +test "MAST advanced constructor rejects rank deficient designs" { + let raised = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0, 3.0]], + [[1.0, 2.0], [1.0, 2.0], [1.0, 2.0]], + ["intercept", "constant"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "MAST advanced constructor validates identifiers" { + let length = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0]], + [[1.0], [1.0]], + ["intercept"], + gene_names=["g1", "g2"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let duplicate = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0], [2.0, 3.0]], + [[1.0], [1.0]], + ["intercept"], + gene_names=["g1", "g1"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let empty = try { + ignore( + @src.MastAdvancedData::create( + [[1.0, 2.0]], + [[1.0], [1.0]], + [" "], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(length) + assert_true(duplicate) + assert_true(empty) +} + +///| +test "MAST advanced grouped constructor validates groups and reference" { + let length = try { + ignore( + @src.mast_advanced_from_groups([[1.0, 2.0]], ["A"], include_cdr=false), + ) + false + } catch { + _ => true + } + let one_level = try { + ignore( + @src.mast_advanced_from_groups( + [[1.0, 2.0]], + ["A", "A"], + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + let reference = try { + ignore( + @src.mast_advanced_from_groups( + [[1.0, 2.0]], + ["A", "B"], + reference="C", + include_cdr=false, + ), + ) + false + } catch { + _ => true + } + assert_true(length) + assert_true(one_level) + assert_true(reference) +} + +///| +test "MAST advanced example exposes intended dimensions and contrast" { + let (data, tested) = @src.mast_advanced_example() catch { + _ => abort("MAST advanced example should build") + } + assert_eq(data.n_genes, 8) + assert_eq(data.n_cells, 48) + assert_eq(data.coefficient_names[1], "group:stimulated") + assert_eq(tested, [1]) +} + +///| +test "MAST advanced zlm returns aligned result matrices" { + let result = mast_adv_test_example_result() + assert_eq(result.gene_names.length(), 8) + assert_eq(result.discrete_coefficients.length(), 8) + assert_eq(result.continuous_coefficients.length(), 8) + assert_eq( + result.discrete_coefficients[0].length(), + result.coefficient_names.length(), + ) + assert_eq(result.tested_names, ["group:stimulated"]) +} + +///| +test "MAST advanced recovers discrete and continuous effect directions" { + let result = mast_adv_test_example_result() + assert_true(result.discrete_coefficients[0][1] > 0.0) + assert_true(result.continuous_coefficients[0][1] > 0.0) + assert_true(result.continuous_coefficients[1][1] > 0.0) +} + +///| +test "MAST advanced computes valid component and hurdle probabilities" { + let result = mast_adv_test_example_result() + for gene in 0.. assert_true(value >= 0.0 && value <= 1.0) + None => () + } + match result.hurdle_fdr[gene] { + Some(value) => assert_true(value >= 0.0 && value <= 1.0) + None => () + } + } + assert_true(result.n_tested() > 0) +} + +///| +test "MAST advanced hurdle statistic sums testable components" { + let result = mast_adv_test_example_result() + let discrete = result.discrete_statistics[0].unwrap_or(-1.0) + let continuous = result.continuous_statistics[0].unwrap_or(-1.0) + let hurdle = result.hurdle_statistics[0].unwrap_or(-1.0) + assert_true(discrete >= 0.0) + assert_true(continuous >= 0.0) + assert_true((hurdle - discrete - continuous).abs() < 1.0e-8) +} + +///| +test "MAST advanced empirical Bayes estimates a finite variance prior" { + let result = mast_adv_test_example_result() + assert_true(result.empirical_bayes) + assert_true(result.prior_variance > 0.0) + assert_true(result.prior_variance < 1.0e300) + assert_true(result.prior_df > 0.0) + assert_true(result.prior_df <= 1000.0) +} + +///| +test "MAST advanced empirical Bayes supports full-model residuals" { + let (data, tested) = @src.mast_advanced_example() catch { + _ => abort("MAST advanced example should build") + } + let config = @src.MastAdvancedConfig::create( + ebayes_use_full_model=true, + maximum_prior_df=80.0, + ) catch { + _ => abort("valid H1 eBayes configuration should build") + } + let result = @src.mast_advanced_zlm(data, tested, config~) catch { + _ => abort("full-model eBayes workflow should fit") + } + assert_true(result.empirical_bayes) + assert_true(result.prior_variance > 0.0) + assert_true(result.prior_df > 0.0 && result.prior_df <= 80.0) +} + +///| +test "MAST advanced can disable empirical Bayes moderation" { + let (data, tested) = @src.mast_advanced_example() catch { + _ => abort("MAST advanced example should build") + } + let config = @src.MastAdvancedConfig::create(empirical_bayes=false) catch { + _ => abort("valid unmoderated configuration should build") + } + let result = @src.mast_advanced_zlm(data, tested, config~) catch { + _ => abort("unmoderated MAST workflow should fit") + } + assert_eq(result.empirical_bayes, false) + assert_eq(result.prior_variance, 0.0) + assert_eq(result.prior_df, 0.0) + for gene in 0.. + assert_true((raw - moderated).abs() < 1.0e-10) + _ => () + } + } +} + +///| +test "MAST advanced moderated variances are finite and positive" { + let result = mast_adv_test_example_result() + for value in result.moderated_variance { + match value { + Some(variance) => { + assert_true(variance > 0.0) + assert_true(variance < 1.0e300) + } + None => () + } + } +} + +///| +test "MAST advanced preserves NA semantics for untestable genes" { + let result = mast_adv_test_na_result() + assert_true(result.discrete_p_values[0] is None) + assert_true(result.continuous_p_values[0] is None) + assert_true(result.hurdle_p_values[0] is None) + assert_true(result.discrete_p_values[1] is None) + assert_true(result.continuous_p_values[1] is Some(_)) + assert_true(result.continuous_p_values[2] is None) +} + +///| +test "MAST advanced combines whichever hurdle components are testable" { + let result = mast_adv_test_na_result() + assert_true(result.hurdle_p_values[1] is Some(_)) + assert_true(result.hurdle_p_values[2] is Some(_)) + assert_true(result.hurdle_p_values[3] is Some(_)) +} + +///| +test "MAST advanced BH correction preserves missing values" { + let result = mast_adv_test_na_result() + assert_true(result.hurdle_fdr[0] is None) + for gene in 1.. assert_true(fdr + 1.0e-12 >= p) + _ => () + } + } +} + +///| +test "MAST advanced marginal effects and delta variances are finite" { + let result = mast_adv_test_example_result() + assert_true(result.marginal_log_fc[0].unwrap_or(0.0) > 0.0) + assert_true(result.marginal_log_fc[1] is Some(_)) + assert_true(result.marginal_log_fc[1].unwrap_or(1.0e300).abs() < 1.0e300) + assert_true(result.marginal_log_fc_variance[0].unwrap_or(-1.0) >= 0.0) +} + +///| +test "MAST advanced fitting is deterministic" { + let first = mast_adv_test_example_result() + let second = mast_adv_test_example_result() + assert_eq(first.discrete_coefficients, second.discrete_coefficients) + assert_eq(first.continuous_coefficients, second.continuous_coefficients) + assert_eq(first.hurdle_p_values, second.hurdle_p_values) + assert_eq(first.hurdle_fdr, second.hurdle_fdr) +} + +///| +test "MAST advanced supports joint deletion of design columns" { + let data = @src.MastAdvancedData::create( + [[1.0, 1.1, 1.2, 1.3, 2.0, 2.1, 2.2, 2.3, 3.0, 3.1, 3.2, 3.3]], + [ + [1.0, 0.0, 0.0], + [1.0, 0.0, 0.0], + [1.0, 0.0, 0.0], + [1.0, 0.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 0.0, 1.0], + [1.0, 0.0, 1.0], + [1.0, 0.0, 1.0], + [1.0, 0.0, 1.0], + ], + ["(Intercept)", "group:B", "group:C"], + include_cdr=false, + ) catch { + _ => abort("valid multi-level design should build") + } + let config = @src.MastAdvancedConfig::create(empirical_bayes=false) catch { + _ => abort("valid MAST configuration should build") + } + let result = @src.mast_advanced_zlm(data, [1, 2], config~) catch { + _ => abort("joint reduced-model test should fit") + } + assert_eq(result.tested_names, ["group:B", "group:C"]) + assert_true(result.continuous_p_values[0] is Some(_)) + assert_true(result.hurdle_statistics[0].unwrap_or(-1.0) >= 0.0) +} + +///| +test "MAST advanced validates tested coefficient indices" { + let data = mast_adv_test_group_data() + let empty = try { + ignore(@src.mast_advanced_zlm(data, [])) + false + } catch { + _ => true + } + let intercept = try { + ignore(@src.mast_advanced_zlm(data, [0])) + false + } catch { + _ => true + } + let range = try { + ignore(@src.mast_advanced_zlm(data, [2])) + false + } catch { + _ => true + } + let duplicate = try { + ignore(@src.mast_advanced_zlm(data, [1, 1])) + false + } catch { + _ => true + } + assert_true(empty) + assert_true(intercept) + assert_true(range) + assert_true(duplicate) +} + +///| +test "MAST advanced result query helpers are consistent" { + let result = mast_adv_test_example_result() + assert_eq(result.index_of("both_components"), Some(0)) + assert_true(result.index_of("missing") is None) + assert_true(result.n_tested() <= result.gene_names.length()) + assert_eq(result.significant_genes().length(), result.n_significant()) + assert_true(result.top_genes(3).length() <= 3) + assert_eq(result.top_genes(0), []) +} + +///| +test "MAST advanced result summary reports contrast and tested count" { + let result = mast_adv_test_example_result() + let summary = result.summary() + assert_true(summary.contains("MAST advanced zlm")) + assert_true(summary.contains("group:stimulated")) + assert_true(summary.contains(result.n_tested().to_string())) +} + +///| +test "MAST advanced SingleCellExperiment writes row and cell results" { + let output = @src.mast_advanced_zlm_sce( + mast_adv_test_sce(), + "condition", + reference="control", + output_prefix="mast39", + ) catch { + _ => abort("valid MAST SCE workflow should fit") + } + assert_eq(output.experiment.row_data["mast39.detectionRate"].length(), 8) + assert_eq(output.experiment.row_data["mast39.nDetected"].length(), 8) + assert_eq(output.experiment.row_data["mast39.hurdleP"].length(), 8) + assert_eq(output.experiment.row_data["mast39.hurdleFdr"].length(), 8) + assert_eq(output.experiment.row_data["mast39.class"].length(), 8) + assert_eq(output.experiment.col_data["mast39.cdr"].length(), 48) +} + +///| +test "MAST advanced SingleCellExperiment records provenance metadata" { + let output = @src.mast_advanced_zlm_sce( + mast_adv_test_sce(), + "condition", + reference="control", + output_prefix="mast39", + ) catch { + _ => abort("valid MAST SCE workflow should fit") + } + assert_eq(output.experiment.metadata["mast39.assay"], "logcounts") + assert_eq(output.experiment.metadata["mast39.group"], "condition") + assert_eq(output.experiment.metadata["mast39.reference"], "control") + assert_eq(output.experiment.metadata["mast39.contrast"], "group:stimulated") + assert_eq(output.experiment.metadata["source"], "test") +} + +///| +test "MAST advanced SingleCellExperiment does not mutate input" { + let input = mast_adv_test_sce() + let output = @src.mast_advanced_zlm_sce( + input, + "condition", + reference="control", + output_prefix="mast39", + ) catch { + _ => abort("valid MAST SCE workflow should fit") + } + assert_eq(input.row_data.contains("mast39.hurdleP"), false) + assert_eq(input.col_data.contains("mast39.cdr"), false) + output.experiment.assays["logcounts"][0][0] = 999.0 + output.experiment.row_data["symbol"][0] = "changed" + output.experiment.reduced_dims["PCA"][0][0] = 999.0 + output.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 999.0 + assert_true(input.assays["logcounts"][0][0] != 999.0) + assert_eq(input.row_data["symbol"][0], "both_components") + assert_eq(input.reduced_dims["PCA"][0][0], 0.0) + assert_eq(input.alternative_experiments["spike"].assays["counts"][0][0], 1.0) +} + +///| +test "MAST advanced SingleCellExperiment supports custom assay and prefix" { + let input = mast_adv_test_sce() + let output = @src.mast_advanced_zlm_sce( + input, + "condition", + assay_name="counts", + reference="control", + output_prefix="custom", + include_cdr=false, + ) catch { + _ => abort("custom MAST SCE workflow should fit") + } + assert_true(output.experiment.row_data.contains("custom.logFC")) + assert_eq(output.experiment.metadata["custom.assay"], "counts") + assert_eq(output.result.coefficient_names, ["(Intercept)", "group:stimulated"]) +} + +///| +test "MAST advanced SingleCellExperiment writes NA for untestable values" { + let result = mast_adv_test_na_result() + let experiment = @src.SingleCellExperiment::new( + [ + [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + [1.0, 1.1, 0.9, 1.2, 1.0, 1.1, 2.0, 2.1, 1.9, 2.2, 2.0, 2.1], + [0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 2.0, 0.0, 0.0, 0.0, 0.0], + [0.0, 1.0, 0.0, 1.1, 0.0, 1.2, 2.0, 0.0, 2.1, 0.0, 2.2, 0.0], + ], + result.gene_names, + ["c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10", "c11", "c12"], + ) + experiment.assays["logcounts"] = experiment.assays["counts"] + experiment.col_data["condition"] = [ + "A", "A", "A", "A", "A", "A", "B", "B", "B", "B", "B", "B", + ] + let config = @src.MastAdvancedConfig::create(empirical_bayes=false) catch { + _ => abort("valid MAST configuration should build") + } + let output = @src.mast_advanced_zlm_sce( + experiment, + "condition", + reference="A", + include_cdr=false, + config~, + ) catch { + _ => abort("MAST NA SCE workflow should fit") + } + assert_eq(output.experiment.row_data["mast.hurdleP"][0], "NA") + assert_eq(output.experiment.row_data["mast.class"][0], "not_tested") +} + +///| +test "MAST advanced SingleCellExperiment rejects missing inputs" { + let assay = try { + ignore( + @src.mast_advanced_zlm_sce( + mast_adv_test_sce(), + "condition", + assay_name="missing", + ), + ) + false + } catch { + _ => true + } + let group = try { + ignore(@src.mast_advanced_zlm_sce(mast_adv_test_sce(), "missing")) + false + } catch { + _ => true + } + assert_true(assay) + assert_true(group) +} + +///| +test "MAST advanced SingleCellExperiment validates prefix and reference" { + let prefix = try { + ignore( + @src.mast_advanced_zlm_sce( + mast_adv_test_sce(), + "condition", + output_prefix=" ", + ), + ) + false + } catch { + _ => true + } + let reference = try { + ignore( + @src.mast_advanced_zlm_sce( + mast_adv_test_sce(), + "condition", + reference="absent", + ), + ) + false + } catch { + _ => true + } + assert_true(prefix) + assert_true(reference) +} diff --git a/test/moonbit/matrix_generics_test.mbt b/test/moonbit/matrix_generics_test.mbt index f792f653..333d0cdb 100644 --- a/test/moonbit/matrix_generics_test.mbt +++ b/test/moonbit/matrix_generics_test.mbt @@ -13,6 +13,7 @@ test "mg_matrix_new_basic" { assert_eq(mat.get(1, 1), 4.0) } +///| test "mg_matrix_new_single_element" { let mat = @src.MgMatrix::new([[42.0]]) assert_eq(mat.dim_rows(), 1) @@ -20,6 +21,7 @@ test "mg_matrix_new_single_element" { assert_eq(mat.get(0, 0), 42.0) } +///| test "mg_matrix_new_single_row" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0, 4.0]]) assert_eq(mat.dim_rows(), 1) @@ -28,6 +30,7 @@ test "mg_matrix_new_single_row" { assert_eq(mat.get(0, 3), 4.0) } +///| test "mg_matrix_new_single_col" { let mat = @src.MgMatrix::new([[1.0], [2.0], [3.0]]) assert_eq(mat.dim_rows(), 3) @@ -36,12 +39,14 @@ test "mg_matrix_new_single_col" { assert_eq(mat.get(2, 0), 3.0) } +///| test "mg_matrix_new_empty" { let mat = @src.MgMatrix::new([]) assert_eq(mat.dim_rows(), 0) assert_eq(mat.dim_cols(), 0) } +///| test "mg_matrix_zeros_basic" { let mat = @src.MgMatrix::zeros(3, 4) assert_eq(mat.dim_rows(), 3) @@ -50,6 +55,7 @@ test "mg_matrix_zeros_basic" { assert_eq(mat.get(2, 3), 0.0) } +///| test "mg_matrix_zeros_single" { let mat = @src.MgMatrix::zeros(1, 1) assert_eq(mat.dim_rows(), 1) @@ -57,6 +63,7 @@ test "mg_matrix_zeros_single" { assert_eq(mat.get(0, 0), 0.0) } +///| test "mg_matrix_zeros_empty" { let mat = @src.MgMatrix::zeros(0, 0) assert_eq(mat.dim_rows(), 0) @@ -67,6 +74,7 @@ test "mg_matrix_zeros_empty" { // Accessor methods // --------------------------------------------------------------------------- +///| test "mg_matrix_get_set" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) mat.set(0, 0, 99.0) @@ -75,6 +83,7 @@ test "mg_matrix_get_set" { assert_eq(mat.get(1, 0), 3.0) } +///| test "mg_matrix_get_row" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) let row0 = mat.get_row(0) @@ -89,6 +98,7 @@ test "mg_matrix_get_row" { assert_eq(row1[2], 6.0) } +///| test "mg_matrix_get_col" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]) let col0 = mat.get_col(0) @@ -107,8 +117,13 @@ test "mg_matrix_get_col" { // Row summary statistics // --------------------------------------------------------------------------- +///| test "mg_row_means_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_row_means(mat) assert_eq(result.length(), 3) assert_eq(result[0], 2.0) @@ -116,6 +131,7 @@ test "mg_row_means_basic" { assert_eq(result[2], 8.0) } +///| test "mg_row_means_constant" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0], [3.0, 3.0, 3.0]]) let result = @src.mg_row_means(mat) @@ -123,6 +139,7 @@ test "mg_row_means_constant" { assert_eq(result[1], 3.0) } +///| test "mg_row_means_single_row" { let mat = @src.MgMatrix::new([[10.0, 20.0, 30.0]]) let result = @src.mg_row_means(mat) @@ -130,8 +147,13 @@ test "mg_row_means_single_row" { assert_eq(result[0], 20.0) } +///| test "mg_row_sums_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_row_sums(mat) assert_eq(result.length(), 3) assert_eq(result[0], 6.0) @@ -139,6 +161,7 @@ test "mg_row_sums_basic" { assert_eq(result[2], 24.0) } +///| test "mg_row_sums_constant" { let mat = @src.MgMatrix::new([[2.0, 2.0, 2.0], [0.0, 0.0, 0.0]]) let result = @src.mg_row_sums(mat) @@ -146,8 +169,13 @@ test "mg_row_sums_constant" { assert_eq(result[1], 0.0) } +///| test "mg_row_vars_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_row_vars(mat) assert_eq(result.length(), 3) assert_true(result[0] > 0.99 && result[0] < 1.01) @@ -155,6 +183,7 @@ test "mg_row_vars_basic" { assert_true(result[2] > 0.99 && result[2] < 1.01) } +///| test "mg_row_vars_constant" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0], [2.0, 2.0, 2.0]]) let result = @src.mg_row_vars(mat) @@ -162,12 +191,14 @@ test "mg_row_vars_constant" { assert_eq(result[1], 0.0) } +///| test "mg_row_vars_single" { let mat = @src.MgMatrix::new([[3.0]]) let result = @src.mg_row_vars(mat) assert_eq(result[0], 0.0) } +///| test "mg_row_sds_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) let result = @src.mg_row_sds(mat) @@ -176,12 +207,14 @@ test "mg_row_sds_basic" { assert_true(result[1] > 0.99 && result[1] < 1.01) } +///| test "mg_row_sds_constant" { let mat = @src.MgMatrix::new([[7.0, 7.0, 7.0]]) let result = @src.mg_row_sds(mat) assert_eq(result[0], 0.0) } +///| test "mg_row_medians_odd" { let mat = @src.MgMatrix::new([[1.0, 3.0, 2.0], [6.0, 4.0, 5.0]]) let result = @src.mg_row_medians(mat) @@ -189,46 +222,61 @@ test "mg_row_medians_odd" { assert_eq(result[1], 5.0) } +///| test "mg_row_medians_even" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0, 4.0]]) let result = @src.mg_row_medians(mat) assert_eq(result[0], 2.5) } +///| test "mg_row_medians_constant" { let mat = @src.MgMatrix::new([[9.0, 9.0, 9.0]]) let result = @src.mg_row_medians(mat) assert_eq(result[0], 9.0) } +///| test "mg_row_mins_basic" { - let mat = @src.MgMatrix::new([[3.0, 1.0, 2.0], [6.0, 5.0, 4.0], [9.0, 8.0, 7.0]]) + let mat = @src.MgMatrix::new([ + [3.0, 1.0, 2.0], + [6.0, 5.0, 4.0], + [9.0, 8.0, 7.0], + ]) let result = @src.mg_row_mins(mat) assert_eq(result[0], 1.0) assert_eq(result[1], 4.0) assert_eq(result[2], 7.0) } +///| test "mg_row_mins_constant" { let mat = @src.MgMatrix::new([[4.0, 4.0, 4.0]]) let result = @src.mg_row_mins(mat) assert_eq(result[0], 4.0) } +///| test "mg_row_maxs_basic" { - let mat = @src.MgMatrix::new([[3.0, 1.0, 2.0], [6.0, 5.0, 4.0], [9.0, 8.0, 7.0]]) + let mat = @src.MgMatrix::new([ + [3.0, 1.0, 2.0], + [6.0, 5.0, 4.0], + [9.0, 8.0, 7.0], + ]) let result = @src.mg_row_maxs(mat) assert_eq(result[0], 3.0) assert_eq(result[1], 6.0) assert_eq(result[2], 9.0) } +///| test "mg_row_maxs_constant" { let mat = @src.MgMatrix::new([[4.0, 4.0, 4.0]]) let result = @src.mg_row_maxs(mat) assert_eq(result[0], 4.0) } +///| test "mg_row_ranges_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [10.0, 5.0, 2.0]]) let result = @src.mg_row_ranges(mat) @@ -236,14 +284,20 @@ test "mg_row_ranges_basic" { assert_eq(result[1], 8.0) } +///| test "mg_row_ranges_constant" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0]]) let result = @src.mg_row_ranges(mat) assert_eq(result[0], 0.0) } +///| test "mg_row_counts_basic" { - let mat = @src.MgMatrix::new([[1.0, 0.0, 3.0], [0.0, 0.0, 0.0], [4.0, 5.0, 6.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 0.0, 3.0], + [0.0, 0.0, 0.0], + [4.0, 5.0, 6.0], + ]) let result = @src.mg_row_counts(mat) assert_eq(result.length(), 3) assert_eq(result[0], 2) @@ -251,6 +305,7 @@ test "mg_row_counts_basic" { assert_eq(result[2], 3) } +///| test "mg_row_counts_all_zero" { let mat = @src.MgMatrix::new([[0.0, 0.0], [0.0, 0.0]]) let result = @src.mg_row_counts(mat) @@ -258,6 +313,7 @@ test "mg_row_counts_all_zero" { assert_eq(result[1], 0) } +///| test "mg_row_mads_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) let result = @src.mg_row_mads(mat) @@ -266,6 +322,7 @@ test "mg_row_mads_basic" { assert_true(result[1] > 1.4 && result[1] < 1.6) } +///| test "mg_row_mads_constant" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0]]) let result = @src.mg_row_mads(mat) @@ -276,8 +333,13 @@ test "mg_row_mads_constant" { // Row anys and alls // --------------------------------------------------------------------------- +///| test "mg_row_anys_basic" { - let mat = @src.MgMatrix::new([[1.0, 5.0, 3.0], [2.0, 2.0, 2.0], [10.0, 20.0, 30.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 5.0, 3.0], + [2.0, 2.0, 2.0], + [10.0, 20.0, 30.0], + ]) let result = @src.mg_row_anys(mat, 4.0) assert_eq(result.length(), 3) assert_true(result[0]) @@ -285,20 +347,27 @@ test "mg_row_anys_basic" { assert_true(result[2]) } +///| test "mg_row_anys_all_below" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0]]) let result = @src.mg_row_anys(mat, 100.0) assert_false(result[0]) } +///| test "mg_row_anys_threshold_boundary" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0]]) let result = @src.mg_row_anys(mat, 5.0) assert_false(result[0]) } +///| test "mg_row_alls_basic" { - let mat = @src.MgMatrix::new([[10.0, 20.0, 30.0], [1.0, 100.0, 50.0], [2.0, 2.0, 2.0]]) + let mat = @src.MgMatrix::new([ + [10.0, 20.0, 30.0], + [1.0, 100.0, 50.0], + [2.0, 2.0, 2.0], + ]) let result = @src.mg_row_alls(mat, 5.0) assert_eq(result.length(), 3) assert_true(result[0]) @@ -306,6 +375,7 @@ test "mg_row_alls_basic" { assert_false(result[2]) } +///| test "mg_row_alls_all_above" { let mat = @src.MgMatrix::new([[10.0, 10.0], [20.0, 20.0]]) let result = @src.mg_row_alls(mat, 5.0) @@ -313,6 +383,7 @@ test "mg_row_alls_all_above" { assert_true(result[1]) } +///| test "mg_row_alls_threshold_boundary" { let mat = @src.MgMatrix::new([[5.0, 5.0, 5.0]]) let result = @src.mg_row_alls(mat, 5.0) @@ -323,8 +394,13 @@ test "mg_row_alls_threshold_boundary" { // Column summary statistics // --------------------------------------------------------------------------- +///| test "mg_col_means_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_col_means(mat) assert_eq(result.length(), 3) assert_eq(result[0], 4.0) @@ -332,6 +408,7 @@ test "mg_col_means_basic" { assert_eq(result[2], 6.0) } +///| test "mg_col_means_constant" { let mat = @src.MgMatrix::new([[3.0, 7.0], [3.0, 7.0], [3.0, 7.0]]) let result = @src.mg_col_means(mat) @@ -339,6 +416,7 @@ test "mg_col_means_constant" { assert_eq(result[1], 7.0) } +///| test "mg_col_means_single_col" { let mat = @src.MgMatrix::new([[10.0], [20.0], [30.0]]) let result = @src.mg_col_means(mat) @@ -346,8 +424,13 @@ test "mg_col_means_single_col" { assert_eq(result[0], 20.0) } +///| test "mg_col_sums_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_col_sums(mat) assert_eq(result.length(), 3) assert_eq(result[0], 12.0) @@ -355,6 +438,7 @@ test "mg_col_sums_basic" { assert_eq(result[2], 18.0) } +///| test "mg_col_sums_constant" { let mat = @src.MgMatrix::new([[2.0, 0.0], [2.0, 0.0], [2.0, 0.0]]) let result = @src.mg_col_sums(mat) @@ -362,8 +446,13 @@ test "mg_col_sums_constant" { assert_eq(result[1], 0.0) } +///| test "mg_col_vars_basic" { - let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 2.0, 3.0], + [4.0, 5.0, 6.0], + [7.0, 8.0, 9.0], + ]) let result = @src.mg_col_vars(mat) assert_eq(result.length(), 3) assert_true(result[0] > 8.99 && result[0] < 9.01) @@ -371,6 +460,7 @@ test "mg_col_vars_basic" { assert_true(result[2] > 8.99 && result[2] < 9.01) } +///| test "mg_col_vars_constant" { let mat = @src.MgMatrix::new([[5.0, 7.0], [5.0, 7.0], [5.0, 7.0]]) let result = @src.mg_col_vars(mat) @@ -378,12 +468,14 @@ test "mg_col_vars_constant" { assert_eq(result[1], 0.0) } +///| test "mg_col_vars_single" { let mat = @src.MgMatrix::new([[3.0], [3.0]]) let result = @src.mg_col_vars(mat) assert_eq(result[0], 0.0) } +///| test "mg_col_sds_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0], [4.0, 5.0], [7.0, 8.0]]) let result = @src.mg_col_sds(mat) @@ -392,6 +484,7 @@ test "mg_col_sds_basic" { assert_true(result[1] > 2.99 && result[1] < 3.01) } +///| test "mg_col_sds_constant" { let mat = @src.MgMatrix::new([[3.0, 5.0], [3.0, 5.0]]) let result = @src.mg_col_sds(mat) @@ -399,6 +492,7 @@ test "mg_col_sds_constant" { assert_eq(result[1], 0.0) } +///| test "mg_col_medians_odd" { let mat = @src.MgMatrix::new([[1.0, 3.0], [4.0, 1.0], [7.0, 2.0]]) let result = @src.mg_col_medians(mat) @@ -406,6 +500,7 @@ test "mg_col_medians_odd" { assert_eq(result[1], 2.0) } +///| test "mg_col_medians_even" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) let result = @src.mg_col_medians(mat) @@ -413,6 +508,7 @@ test "mg_col_medians_even" { assert_eq(result[1], 3.0) } +///| test "mg_col_medians_constant" { let mat = @src.MgMatrix::new([[8.0, 2.0], [8.0, 2.0], [8.0, 2.0]]) let result = @src.mg_col_medians(mat) @@ -420,14 +516,20 @@ test "mg_col_medians_constant" { assert_eq(result[1], 2.0) } +///| test "mg_col_mins_basic" { - let mat = @src.MgMatrix::new([[3.0, 1.0, 2.0], [6.0, 5.0, 4.0], [9.0, 8.0, 7.0]]) + let mat = @src.MgMatrix::new([ + [3.0, 1.0, 2.0], + [6.0, 5.0, 4.0], + [9.0, 8.0, 7.0], + ]) let result = @src.mg_col_mins(mat) assert_eq(result[0], 3.0) assert_eq(result[1], 1.0) assert_eq(result[2], 2.0) } +///| test "mg_col_mins_constant" { let mat = @src.MgMatrix::new([[4.0, 6.0], [4.0, 6.0], [4.0, 6.0]]) let result = @src.mg_col_mins(mat) @@ -435,14 +537,20 @@ test "mg_col_mins_constant" { assert_eq(result[1], 6.0) } +///| test "mg_col_maxs_basic" { - let mat = @src.MgMatrix::new([[3.0, 1.0, 2.0], [6.0, 5.0, 4.0], [9.0, 8.0, 7.0]]) + let mat = @src.MgMatrix::new([ + [3.0, 1.0, 2.0], + [6.0, 5.0, 4.0], + [9.0, 8.0, 7.0], + ]) let result = @src.mg_col_maxs(mat) assert_eq(result[0], 9.0) assert_eq(result[1], 8.0) assert_eq(result[2], 7.0) } +///| test "mg_col_maxs_constant" { let mat = @src.MgMatrix::new([[4.0, 6.0], [4.0, 6.0], [4.0, 6.0]]) let result = @src.mg_col_maxs(mat) @@ -450,6 +558,7 @@ test "mg_col_maxs_constant" { assert_eq(result[1], 6.0) } +///| test "mg_col_ranges_basic" { let mat = @src.MgMatrix::new([[1.0, 10.0], [5.0, 2.0], [10.0, 1.0]]) let result = @src.mg_col_ranges(mat) @@ -457,6 +566,7 @@ test "mg_col_ranges_basic" { assert_eq(result[1], 9.0) } +///| test "mg_col_ranges_constant" { let mat = @src.MgMatrix::new([[5.0, 3.0], [5.0, 3.0], [5.0, 3.0]]) let result = @src.mg_col_ranges(mat) @@ -468,8 +578,13 @@ test "mg_col_ranges_constant" { // Col anys and alls // --------------------------------------------------------------------------- +///| test "mg_col_anys_basic" { - let mat = @src.MgMatrix::new([[1.0, 5.0, 3.0], [2.0, 2.0, 2.0], [1.0, 8.0, 1.0]]) + let mat = @src.MgMatrix::new([ + [1.0, 5.0, 3.0], + [2.0, 2.0, 2.0], + [1.0, 8.0, 1.0], + ]) let result = @src.mg_col_anys(mat, 4.0) assert_eq(result.length(), 3) assert_false(result[0]) @@ -477,6 +592,7 @@ test "mg_col_anys_basic" { assert_false(result[2]) } +///| test "mg_col_anys_all_below" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) let result = @src.mg_col_anys(mat, 100.0) @@ -484,6 +600,7 @@ test "mg_col_anys_all_below" { assert_false(result[1]) } +///| test "mg_col_alls_basic" { let mat = @src.MgMatrix::new([[10.0, 5.0], [20.0, 6.0], [30.0, 7.0]]) let result = @src.mg_col_alls(mat, 4.0) @@ -491,6 +608,7 @@ test "mg_col_alls_basic" { assert_true(result[1]) } +///| test "mg_col_alls_some_below" { let mat = @src.MgMatrix::new([[10.0, 3.0], [20.0, 6.0], [30.0, 7.0]]) let result = @src.mg_col_alls(mat, 5.0) @@ -502,9 +620,10 @@ test "mg_col_alls_some_below" { // Block processing // --------------------------------------------------------------------------- +///| test "mg_block_apply_rows_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]) - let result = @src.mg_block_apply_rows(mat, 2, (block) => { + let result = @src.mg_block_apply_rows(mat, 2, block => { @src.mg_row_means(block) }) assert_eq(result.length(), 3) @@ -513,9 +632,10 @@ test "mg_block_apply_rows_basic" { assert_eq(result[2], 5.5) } +///| test "mg_block_apply_rows_single_block" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) - let result = @src.mg_block_apply_rows(mat, 10, (block) => { + let result = @src.mg_block_apply_rows(mat, 10, block => { @src.mg_row_sums(block) }) assert_eq(result.length(), 2) @@ -523,9 +643,10 @@ test "mg_block_apply_rows_single_block" { assert_eq(result[1], 7.0) } +///| test "mg_block_apply_rows_block_size_exceeds" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) - let result = @src.mg_block_apply_rows(mat, 100, (block) => { + let result = @src.mg_block_apply_rows(mat, 100, block => { @src.mg_row_mins(block) }) assert_eq(result.length(), 2) @@ -533,9 +654,10 @@ test "mg_block_apply_rows_block_size_exceeds" { assert_eq(result[1], 3.0) } +///| test "mg_block_apply_cols_basic" { let mat = @src.MgMatrix::new([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) - let result = @src.mg_block_apply_cols(mat, 2, (block) => { + let result = @src.mg_block_apply_cols(mat, 2, block => { @src.mg_col_means(block) }) assert_eq(result.length(), 3) @@ -544,9 +666,10 @@ test "mg_block_apply_cols_basic" { assert_eq(result[2], 4.5) } +///| test "mg_block_apply_cols_single_block" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) - let result = @src.mg_block_apply_cols(mat, 10, (block) => { + let result = @src.mg_block_apply_cols(mat, 10, block => { @src.mg_col_sums(block) }) assert_eq(result.length(), 2) @@ -558,144 +681,172 @@ test "mg_block_apply_cols_single_block" { // Standalone statistical functions // --------------------------------------------------------------------------- +///| test "mg_mean_basic" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] assert_eq(@src.mg_mean(data), 3.0) } +///| test "mg_mean_single" { let data = [42.0] assert_eq(@src.mg_mean(data), 42.0) } +///| test "mg_mean_empty" { let data : Array[Double] = [] assert_eq(@src.mg_mean(data), 0.0) } +///| test "mg_mean_constant" { let data = [7.0, 7.0, 7.0, 7.0] assert_eq(@src.mg_mean(data), 7.0) } +///| test "mg_mean_negative" { let data = [-1.0, -2.0, -3.0] assert_eq(@src.mg_mean(data), -2.0) } +///| test "mg_variance_basic" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let result = @src.mg_variance(data) assert_true(result > 2.49 && result < 2.51) } +///| test "mg_variance_single" { let data = [5.0] assert_eq(@src.mg_variance(data), 0.0) } +///| test "mg_variance_empty" { let data : Array[Double] = [] assert_eq(@src.mg_variance(data), 0.0) } +///| test "mg_variance_constant" { let data = [3.0, 3.0, 3.0] assert_eq(@src.mg_variance(data), 0.0) } +///| test "mg_variance_two_elements" { let data = [2.0, 4.0] let result = @src.mg_variance(data) assert_true(result > 1.99 && result < 2.01) } +///| test "mg_median_odd" { let data = [3.0, 1.0, 2.0] assert_eq(@src.mg_median(data), 2.0) } +///| test "mg_median_even" { let data = [1.0, 2.0, 3.0, 4.0] assert_eq(@src.mg_median(data), 2.5) } +///| test "mg_median_single" { let data = [99.0] assert_eq(@src.mg_median(data), 99.0) } +///| test "mg_median_empty" { let data : Array[Double] = [] assert_eq(@src.mg_median(data), 0.0) } +///| test "mg_median_two" { let data = [5.0, 10.0] assert_eq(@src.mg_median(data), 7.5) } +///| test "mg_min_basic" { let data = [3.0, 1.0, 4.0, 1.0, 5.0, 9.0, 2.0] assert_eq(@src.mg_min(data), 1.0) } +///| test "mg_min_single" { let data = [42.0] assert_eq(@src.mg_min(data), 42.0) } +///| test "mg_min_empty" { let data : Array[Double] = [] assert_eq(@src.mg_min(data), 0.0) } +///| test "mg_min_negative" { let data = [-5.0, -1.0, -3.0] assert_eq(@src.mg_min(data), -5.0) } +///| test "mg_max_basic" { let data = [3.0, 1.0, 4.0, 1.0, 5.0, 9.0, 2.0] assert_eq(@src.mg_max(data), 9.0) } +///| test "mg_max_single" { let data = [42.0] assert_eq(@src.mg_max(data), 42.0) } +///| test "mg_max_empty" { let data : Array[Double] = [] assert_eq(@src.mg_max(data), 0.0) } +///| test "mg_max_negative" { let data = [-5.0, -1.0, -3.0] assert_eq(@src.mg_max(data), -1.0) } +///| test "mg_mad_basic" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let result = @src.mg_mad(data) assert_true(result > 1.4 && result < 2.0) } +///| test "mg_mad_constant" { let data = [5.0, 5.0, 5.0] assert_eq(@src.mg_mad(data), 0.0) } +///| test "mg_mad_single" { let data = [7.0] assert_eq(@src.mg_mad(data), 0.0) } +///| test "mg_mad_empty" { let data : Array[Double] = [] assert_eq(@src.mg_mad(data), 0.0) } +///| test "mg_mad_two_same" { let data = [3.0, 3.0] assert_eq(@src.mg_mad(data), 0.0) @@ -705,6 +856,7 @@ test "mg_mad_two_same" { // Edge cases // --------------------------------------------------------------------------- +///| test "mg_empty_matrix_row_stats" { let mat = @src.MgMatrix::zeros(0, 5) let means = @src.mg_row_means(mat) @@ -713,12 +865,14 @@ test "mg_empty_matrix_row_stats" { assert_eq(sums.length(), 0) } +///| test "mg_empty_matrix_col_stats" { let mat = @src.MgMatrix::zeros(5, 0) let means = @src.mg_col_means(mat) assert_eq(means.length(), 0) } +///| test "mg_single_value_matrix" { let mat = @src.MgMatrix::new([[7.0]]) assert_eq(mat.dim_rows(), 1) @@ -731,6 +885,7 @@ test "mg_single_value_matrix" { assert_eq(@src.mg_col_medians(mat)[0], 7.0) } +///| test "mg_negative_values_matrix" { let mat = @src.MgMatrix::new([[-1.0, -2.0], [-3.0, -4.0]]) let row_means = @src.mg_row_means(mat) @@ -741,6 +896,7 @@ test "mg_negative_values_matrix" { assert_eq(col_means[1], -3.0) } +///| test "mg_large_matrix_stats" { let mat = @src.MgMatrix::zeros(100, 50) let row_means = @src.mg_row_means(mat) @@ -751,41 +907,47 @@ test "mg_large_matrix_stats" { assert_eq(col_means[0], 0.0) } +///| test "mg_row_anys_empty_matrix" { let mat = @src.MgMatrix::zeros(0, 3) let result = @src.mg_row_anys(mat, 1.0) assert_eq(result.length(), 0) } +///| test "mg_row_alls_empty_matrix" { let mat = @src.MgMatrix::zeros(0, 3) let result = @src.mg_row_alls(mat, 1.0) assert_eq(result.length(), 0) } +///| test "mg_col_anys_empty_matrix" { let mat = @src.MgMatrix::zeros(3, 0) let result = @src.mg_col_anys(mat, 1.0) assert_eq(result.length(), 0) } +///| test "mg_col_alls_empty_matrix" { let mat = @src.MgMatrix::zeros(3, 0) let result = @src.mg_col_alls(mat, 1.0) assert_eq(result.length(), 0) } +///| test "mg_block_apply_rows_empty" { let mat = @src.MgMatrix::zeros(0, 5) - let result = @src.mg_block_apply_rows(mat, 10, (block) => { + let result = @src.mg_block_apply_rows(mat, 10, block => { @src.mg_row_means(block) }) assert_eq(result.length(), 0) } +///| test "mg_block_apply_cols_empty" { let mat = @src.MgMatrix::zeros(5, 0) - let result = @src.mg_block_apply_cols(mat, 10, (block) => { + let result = @src.mg_block_apply_cols(mat, 10, block => { @src.mg_col_means(block) }) assert_eq(result.length(), 0) @@ -795,6 +957,7 @@ test "mg_block_apply_cols_empty" { // Cross-validation: row and col stats should agree on the same data // --------------------------------------------------------------------------- +///| test "mg_row_col_stats_consistency" { let mat = @src.MgMatrix::new([[1.0, 2.0], [3.0, 4.0]]) let row_means = @src.mg_row_means(mat) @@ -804,10 +967,11 @@ test "mg_row_col_stats_consistency" { assert_eq(@src.mg_mean(row_means), @src.mg_mean(col_means)) } +///| test "mg_symmetric_matrix_row_col_stats" { let mat = @src.MgMatrix::new([[1.0, 2.0], [2.0, 1.0]]) let row_means = @src.mg_row_means(mat) let col_means = @src.mg_col_means(mat) assert_eq(row_means[0], col_means[0]) assert_eq(row_means[1], col_means[1]) -} \ No newline at end of file +} diff --git a/test/moonbit/matrix_test.mbt b/test/moonbit/matrix_test.mbt index 88e37761..55a49458 100644 --- a/test/moonbit/matrix_test.mbt +++ b/test/moonbit/matrix_test.mbt @@ -1,38 +1,54 @@ ///| /// Tests for Matrix module. - test "BiocMatrix dense_from_array" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) assert_eq(mat.nrow(), 2) assert_eq(mat.ncol(), 2) } +///| test "BiocMatrix get element" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) assert_eq(mat.get(0, 0), 1.0) assert_eq(mat.get(1, 1), 4.0) } +///| test "BiocMatrix set element" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let mat2 = mat.set(0, 0, 10.0) assert_eq(mat2.get(0, 0), 10.0) } +///| test "BiocMatrix csc_from_triplets" { - let mat = @src.BiocMatrix::csc_from_triplets([0, 1, 0, 1], [0, 0, 1, 1], [1.0, 3.0, 2.0, 4.0], 2, 2) + let mat = @src.BiocMatrix::csc_from_triplets( + [0, 1, 0, 1], + [0, 0, 1, 1], + [1.0, 3.0, 2.0, 4.0], + 2, + 2, + ) assert_eq(mat.nrow(), 2) assert_eq(mat.ncol(), 2) assert_eq(mat.get(0, 0), 1.0) } +///| test "BiocMatrix csr_from_triplets" { - let mat = @src.BiocMatrix::csr_from_triplets([0, 0, 1, 1], [0, 1, 0, 1], [1.0, 2.0, 3.0, 4.0], 2, 2) + let mat = @src.BiocMatrix::csr_from_triplets( + [0, 0, 1, 1], + [0, 1, 0, 1], + [1.0, 2.0, 3.0, 4.0], + 2, + 2, + ) assert_eq(mat.nrow(), 2) assert_eq(mat.ncol(), 2) assert_eq(mat.get(0, 1), 2.0) } +///| test "BiocMatrix transpose" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let t = mat.transpose() @@ -40,6 +56,7 @@ test "BiocMatrix transpose" { assert_eq(t.get(1, 0), 2.0) } +///| test "BiocMatrix add" { let mat1 = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let mat2 = @src.BiocMatrix::dense_from_array([[1.0, 0.0], [0.0, 1.0]]) @@ -48,6 +65,7 @@ test "BiocMatrix add" { assert_eq(result.unwrap().get(0, 0), 2.0) } +///| test "BiocMatrix multiply" { let mat1 = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let mat2 = @src.BiocMatrix::dense_from_array([[2.0, 0.0], [0.0, 2.0]]) @@ -56,6 +74,7 @@ test "BiocMatrix multiply" { assert_eq(result.unwrap().get(0, 0), 2.0) } +///| test "BiocMatrix row_sums" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let sums = mat.row_sums() @@ -63,6 +82,7 @@ test "BiocMatrix row_sums" { assert_eq(sums[0], 3.0) } +///| test "BiocMatrix col_sums" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let sums = mat.col_sums() @@ -70,24 +90,28 @@ test "BiocMatrix col_sums" { assert_eq(sums[0], 4.0) } +///| test "BiocMatrix row_means" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 3.0], [2.0, 4.0]]) let means = mat.row_means() assert_eq(means[0], 2.0) } +///| test "BiocMatrix col_means" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0], [3.0, 4.0]]) let means = mat.col_means() assert_eq(means[0], 2.0) } +///| test "BiocMatrix norm" { let mat = @src.BiocMatrix::dense_from_array([[3.0, 4.0]]) let norm = mat.norm(2.0) assert_eq(norm, 5.0) } +///| test "BiocMatrix dim" { let mat = @src.BiocMatrix::dense_from_array([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) let (nrow, ncol) = mat.dim() @@ -95,17 +119,20 @@ test "BiocMatrix dim" { assert_eq(ncol, 3) } +///| test "BiocMatrix nnz" { let mat = @src.BiocMatrix::csc_from_triplets([0, 1], [0, 1], [1.0, 2.0], 2, 2) assert_eq(mat.nnz(), 2) } +///| test "create_example_dense_matrix" { let mat = @src.create_example_dense_matrix() assert_eq(mat.nrow(), 3) assert_eq(mat.ncol(), 3) } +///| test "create_example_sparse_matrix" { let mat = @src.create_example_sparse_matrix() assert_eq(mat.nrow(), 3) diff --git a/test/moonbit/mauve_test.mbt b/test/moonbit/mauve_test.mbt index ce824abf..aedc8a70 100644 --- a/test/moonbit/mauve_test.mbt +++ b/test/moonbit/mauve_test.mbt @@ -53,8 +53,14 @@ test "mauve_lcb_add_seq" { /// Test MauveLCB consistency check. test "mauve_lcb_consistent" { let lcb = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800), + ) assert_true(@src.mauve_lcb_is_consistent(lcb)) } @@ -62,8 +68,14 @@ test "mauve_lcb_consistent" { /// Test MauveLCB inconsistency (mixed strands). test "mauve_lcb_inconsistent" { let lcb = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 100, "-", 800)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "-", 800), + ) assert_false(@src.mauve_lcb_is_consistent(lcb)) } @@ -89,10 +101,16 @@ test "mauve_alignment_add_lcb" { test "mauve_alignment_seq_names" { let ali = @src.new_mauve_alignment() let lcb = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800), + ) @src.mauve_add_lcb(ali, lcb) - + let names = @src.mauve_all_seq_names(ali) assert_eq(names.length(), 2) assert_true(names.contains("human.chr1")) @@ -133,17 +151,23 @@ test "mauve_parse_with_inversions" { /// Test detect_mauve_inversions. test "mauve_detect_inversions" { let ali = @src.new_mauve_alignment() - + // LCB 1: positive strand let lcb1 = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb1, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) + @src.mauve_lcb_add_seq( + lcb1, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) @src.mauve_add_lcb(ali, lcb1) - + // LCB 2: negative strand (inversion) let lcb2 = @src.new_mauve_lcb("lcb_2") - @src.mauve_lcb_add_seq(lcb2, @src.new_mauve_sequence("human.chr1", 100, 100, "-", 1000)) + @src.mauve_lcb_add_seq( + lcb2, + @src.new_mauve_sequence("human.chr1", 100, 100, "-", 1000), + ) @src.mauve_add_lcb(ali, lcb2) - + let ali = @src.detect_mauve_inversions(ali) assert_true(ali.inversions.length() > 0) } @@ -152,15 +176,21 @@ test "mauve_detect_inversions" { /// Test detect_mauve_breakpoints. test "mauve_detect_breakpoints" { let ali = @src.new_mauve_alignment() - + let lcb1 = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb1, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) + @src.mauve_lcb_add_seq( + lcb1, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) @src.mauve_add_lcb(ali, lcb1) - + let lcb2 = @src.new_mauve_lcb("lcb_2") - @src.mauve_lcb_add_seq(lcb2, @src.new_mauve_sequence("human.chr1", 500, 100, "+", 1000)) + @src.mauve_lcb_add_seq( + lcb2, + @src.new_mauve_sequence("human.chr1", 500, 100, "+", 1000), + ) @src.mauve_add_lcb(ali, lcb2) - + let ali = @src.detect_mauve_breakpoints(ali) assert_true(ali.breakpoints.length() > 0) } @@ -170,10 +200,16 @@ test "mauve_detect_breakpoints" { test "mauve_genome_coverage" { let ali = @src.new_mauve_alignment() let lcb = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 500, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 300, "+", 800)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 500, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 300, "+", 800), + ) @src.mauve_add_lcb(ali, lcb) - + let coverage = @src.mauve_genome_coverage(ali) assert_true(coverage.contains("human.chr1")) assert_true(coverage.contains("mouse.chr1")) @@ -185,17 +221,29 @@ test "mauve_genome_coverage" { /// Test mauve_conserved_segments. test "mauve_conserved_segments" { let ali = @src.new_mauve_alignment() - + let lcb1 = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb1, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb1, @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800)) + @src.mauve_lcb_add_seq( + lcb1, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb1, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800), + ) @src.mauve_add_lcb(ali, lcb1) - + let lcb2 = @src.new_mauve_lcb("lcb_2") - @src.mauve_lcb_add_seq(lcb2, @src.new_mauve_sequence("human.chr1", 500, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb2, @src.new_mauve_sequence("mouse.chr1", 500, 100, "+", 800)) + @src.mauve_lcb_add_seq( + lcb2, + @src.new_mauve_sequence("human.chr1", 500, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb2, + @src.new_mauve_sequence("mouse.chr1", 500, 100, "+", 800), + ) @src.mauve_add_lcb(ali, lcb2) - + let segments = @src.mauve_conserved_segments(ali) assert_eq(segments["human.chr1"], 2) assert_eq(segments["mouse.chr1"], 2) @@ -215,9 +263,12 @@ test "mauve_to_bed" { let ali = @src.new_mauve_alignment() let lcb = @src.new_mauve_lcb("lcb_1") @src.mauve_lcb_set_score(lcb, 100.0) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) @src.mauve_add_lcb(ali, lcb) - + let bed = @src.mauve_to_bed(ali) assert_true(bed.contains("track name=")) assert_true(bed.contains("human.chr1")) @@ -232,7 +283,7 @@ test "mauve_inversions_to_bed" { let inv = @src.new_mauve_inversion("human.chr1", 100, 200, 100) inv.affected_lcbs.push("lcb_1") ali.inversions.push(inv) - + let bed = @src.mauve_inversions_to_bed(ali) assert_true(bed.contains("track name=\"Inversions\"")) assert_true(bed.contains("human.chr1")) @@ -245,10 +296,16 @@ test "mauve_summary" { let ali = @src.new_mauve_alignment() let lcb = @src.new_mauve_lcb("lcb_1") @src.mauve_lcb_set_score(lcb, 50.0) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 800), + ) @src.mauve_add_lcb(ali, lcb) - + let summary = @src.mauve_summary(ali) assert_true(summary.contains("Mauve Alignment Summary")) assert_true(summary.contains("Locally Collinear Blocks")) @@ -270,11 +327,20 @@ test "mauve_config_default" { test "mauve_get_genome_seqs" { let ali = @src.new_mauve_alignment() let lcb = @src.new_mauve_lcb("lcb_1") - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("human.chr2", 200, 100, "+", 800)) - @src.mauve_lcb_add_seq(lcb, @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 700)) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr1", 0, 100, "+", 1000), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("human.chr2", 200, 100, "+", 800), + ) + @src.mauve_lcb_add_seq( + lcb, + @src.new_mauve_sequence("mouse.chr1", 0, 100, "+", 700), + ) @src.mauve_add_lcb(ali, lcb) - + let human_seqs = @src.mauve_get_genome_seqs(ali, "human.chr") assert_eq(human_seqs.length(), 2) } @@ -286,7 +352,7 @@ test "mauve_progressive_aligner" { @src.mauve_add_sequence(aligner, "human") @src.mauve_add_sequence(aligner, "mouse") @src.mauve_add_sequence(aligner, "rat") - + @src.mauve_build_guide_tree(aligner) assert_eq(aligner.guide_tree.length(), 3) } diff --git a/test/moonbit/mcp_counter_test.mbt b/test/moonbit/mcp_counter_test.mbt index 3aa346b8..6dc52861 100644 --- a/test/moonbit/mcp_counter_test.mbt +++ b/test/moonbit/mcp_counter_test.mbt @@ -12,6 +12,7 @@ test "mcp_cell_population_struct" { assert_eq(pop.markers().length(), 3) } +///| test "mcp_result_accessors" { // Use mcp_run to get a result with known values let result = @src.mcp_run( @@ -27,6 +28,7 @@ test "mcp_result_accessors" { assert_true((result.scores()[0][0] - 4.0).abs() < 0.001) } +///| test "mcp_result_get_score" { // Build a result with known values via mcp_run // Geometric mean of (4, 9) = 6.0 for S1, (4, 4) = 4.0 for S2 @@ -45,6 +47,7 @@ test "mcp_result_get_score" { assert_true(result.get_score("PopA", "Unknown").abs() < 1.0e-12) } +///| test "mcp_result_get_population_scores" { // Build a result with two populations let result = @src.mcp_run( @@ -72,6 +75,7 @@ test "mcp_result_get_population_scores" { // Built-in populations // --------------------------------------------------------------------------- +///| test "mcp_default_populations_count" { let pops = @src.mcp_default_populations() assert_eq(pops.length(), 10) @@ -81,6 +85,7 @@ test "mcp_default_populations_count" { } } +///| test "mcp_population_names" { let names = @src.mcp_population_names() assert_eq(names.length(), 10) @@ -88,20 +93,27 @@ test "mcp_population_names" { assert_eq(names[9], "Fibroblasts") } +///| test "mcp_default_populations_marker_genes_present" { let pops = @src.mcp_default_populations() // T cells should contain CD3D let t_cells = pops[0] let mut has_cd3d = false for m in t_cells.markers() { - if m == "CD3D" { has_cd3d = true; break } + if m == "CD3D" { + has_cd3d = true + break + } } assert_true(has_cd3d) // B lineage should contain CD19 let b_cells = pops[3] let mut has_cd19 = false for m in b_cells.markers() { - if m == "CD19" { has_cd19 = true; break } + if m == "CD19" { + has_cd19 = true + break + } } assert_true(has_cd19) } @@ -110,74 +122,63 @@ test "mcp_default_populations_marker_genes_present" { // Geometric mean computation via mcp_run // --------------------------------------------------------------------------- +///| test "mcp_run_geometric_mean_correctness" { // Single population with two markers, one sample // Geometric mean of (4.0, 9.0) = sqrt(4 * 9) = 6.0 let pop = @src.McpCellPopulation::new("Test", ["G1", "G2"]) - let result = @src.mcp_run( - ["G1", "G2"], - ["S1"], - [[4.0], [9.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1", "G2"], ["S1"], [[4.0], [9.0]], populations=[ + pop, + ]) assert_eq(result.populations().length(), 1) assert_eq(result.sample_names().length(), 1) let s = result.get_score("Test", "S1") assert_true((s - 6.0).abs() < 0.001) } +///| test "mcp_run_handles_zero_expression" { // If all marker expressions are zero, score should be 0 let pop = @src.McpCellPopulation::new("Test", ["G1", "G2"]) - let result = @src.mcp_run( - ["G1", "G2"], - ["S1"], - [[0.0], [0.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1", "G2"], ["S1"], [[0.0], [0.0]], populations=[ + pop, + ]) assert_true(result.get_score("Test", "S1").abs() < 1.0e-12) } +///| test "mcp_run_handles_missing_markers" { // If a marker is not in gene_names, it is skipped let pop = @src.McpCellPopulation::new("Test", ["G1", "MISSING"]) - let result = @src.mcp_run( - ["G1"], - ["S1"], - [[4.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1"], ["S1"], [[4.0]], populations=[pop]) // Geometric mean of just [4.0] = 4.0 let s = result.get_score("Test", "S1") assert_true((s - 4.0).abs() < 1.0e-9) } +///| test "mcp_run_all_markers_missing" { // If no markers are found, score is 0 let pop = @src.McpCellPopulation::new("Test", ["X1", "X2"]) - let result = @src.mcp_run( - ["G1", "G2"], - ["S1"], - [[4.0], [9.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1", "G2"], ["S1"], [[4.0], [9.0]], populations=[ + pop, + ]) assert_true(result.get_score("Test", "S1").abs() < 1.0e-12) } +///| test "mcp_run_multiple_samples" { // Two samples, one population let pop = @src.McpCellPopulation::new("Test", ["G1"]) - let result = @src.mcp_run( - ["G1"], - ["S1", "S2"], - [[2.0, 8.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1"], ["S1", "S2"], [[2.0, 8.0]], populations=[ + pop, + ]) assert_eq(result.sample_names().length(), 2) assert_true((result.get_score("Test", "S1") - 2.0).abs() < 1.0e-9) assert_true((result.get_score("Test", "S2") - 8.0).abs() < 1.0e-9) } +///| test "mcp_run_default_populations" { // Run with default populations on sample data let (genes, samples, matrix) = @src.mcp_sample_data() @@ -196,6 +197,7 @@ test "mcp_run_default_populations" { // Sample data correctness // --------------------------------------------------------------------------- +///| test "mcp_sample_data_t_cell_rich_tumor_a" { // Tumor_A is T-cell rich: T cells (pop 0), CD8+ T cells (pop 1), // Cytotoxic lymphocytes (pop 2) should all score higher than B lineage (pop 3) @@ -206,6 +208,7 @@ test "mcp_sample_data_t_cell_rich_tumor_a" { assert_true(t_cell_a > b_cell_a) } +///| test "mcp_sample_data_b_cell_rich_tumor_b" { // Tumor_B is B-cell rich: B lineage should score higher than T cells let (genes, samples, matrix) = @src.mcp_sample_data() @@ -215,6 +218,7 @@ test "mcp_sample_data_b_cell_rich_tumor_b" { assert_true(b_cell_b > t_cell_b) } +///| test "mcp_sample_data_fibroblast_rich_tumor_c" { // Tumor_C is fibroblast-rich: Fibroblasts (pop 9) should score higher than T cells let (genes, samples, matrix) = @src.mcp_sample_data() @@ -228,6 +232,7 @@ test "mcp_sample_data_fibroblast_rich_tumor_c" { // to_string formatting // --------------------------------------------------------------------------- +///| test "mcp_result_to_string" { let (genes, samples, matrix) = @src.mcp_sample_data() let result = @src.mcp_run(genes, samples, matrix) @@ -242,44 +247,35 @@ test "mcp_result_to_string" { // Empty / edge cases // --------------------------------------------------------------------------- +///| test "mcp_run_empty_population_list" { - let result = @src.mcp_run( - ["G1"], - ["S1"], - [[1.0]], - populations=[], - ) + let result = @src.mcp_run(["G1"], ["S1"], [[1.0]], populations=[]) assert_eq(result.populations().length(), 0) assert_eq(result.sample_names().length(), 1) } +///| test "mcp_run_no_samples" { let pop = @src.McpCellPopulation::new("Test", ["G1"]) - let result = @src.mcp_run( - ["G1"], - [], - [], - populations=[pop], - ) + let result = @src.mcp_run(["G1"], [], [], populations=[pop]) assert_eq(result.sample_names().length(), 0) assert_eq(result.scores().length(), 1) // 1 population row, but empty assert_eq(result.scores()[0].length(), 0) } +///| test "mcp_run_geometric_mean_with_negative_values" { // Negative expression values are skipped (treated as invalid) let pop = @src.McpCellPopulation::new("Test", ["G1", "G2"]) - let result = @src.mcp_run( - ["G1", "G2"], - ["S1"], - [[-1.0], [9.0]], - populations=[pop], - ) + let result = @src.mcp_run(["G1", "G2"], ["S1"], [[-1.0], [9.0]], populations=[ + pop, + ]) // Only G2 contributes -> geometric mean = 9.0 let s = result.get_score("Test", "S1") assert_true((s - 9.0).abs() < 1.0e-9) } +///| test "mcp_run_custom_populations" { // Custom populations list with non-default markers let custom_pop = @src.McpCellPopulation::new("Custom", ["G1", "G2", "G3"]) diff --git a/test/moonbit/melting_temp_test.mbt b/test/moonbit/melting_temp_test.mbt index 00552af2..c9f4ad9e 100644 --- a/test/moonbit/melting_temp_test.mbt +++ b/test/moonbit/melting_temp_test.mbt @@ -53,4 +53,4 @@ test "recommend_tm_method_long" { test "create_example_dna" { let seq = @src.create_example_dna() assert_true(seq.length() > 0) -} \ No newline at end of file +} diff --git a/test/moonbit/meme_test.mbt b/test/moonbit/meme_test.mbt index e249cbfc..3946a1c2 100644 --- a/test/moonbit/meme_test.mbt +++ b/test/moonbit/meme_test.mbt @@ -3,10 +3,7 @@ ///| test "meme_motif_new" { - let pspm = [ - [0.25, 0.25, 0.25, 0.25], - [0.5, 0.0, 0.5, 0.0], - ] + let pspm = [[0.25, 0.25, 0.25, 0.25], [0.5, 0.0, 0.5, 0.0]] let motif = @src.MemeMotif::new("motif1", pspm) assert_eq(motif.name, "motif1") assert_eq(motif.alt_name, "") @@ -27,11 +24,7 @@ test "meme_motif_new_custom_alphabet" { ///| test "meme_motif_consensus" { - let pspm = [ - [0.9, 0.0, 0.1, 0.0], - [0.0, 0.8, 0.0, 0.2], - [0.1, 0.0, 0.9, 0.0], - ] + let pspm = [[0.9, 0.0, 0.1, 0.0], [0.0, 0.8, 0.0, 0.2], [0.1, 0.0, 0.9, 0.0]] let motif = @src.MemeMotif::new("m1", pspm, alphabet="ACGT") let cons = motif.consensus() assert_eq(cons, "ACG") @@ -39,10 +32,7 @@ test "meme_motif_consensus" { ///| test "meme_motif_probability" { - let pspm = [ - [0.9, 0.0, 0.1, 0.0], - [0.0, 0.8, 0.0, 0.2], - ] + let pspm = [[0.9, 0.0, 0.1, 0.0], [0.0, 0.8, 0.0, 0.2]] let motif = @src.MemeMotif::new("m1", pspm, alphabet="ACGT") assert_eq(motif.probability("A", 0), 0.9) assert_eq(motif.probability("C", 0), 0.0) @@ -55,10 +45,7 @@ test "meme_motif_probability" { ///| test "meme_motif_information_content" { - let pspm = [ - [1.0, 0.0, 0.0, 0.0], - [0.25, 0.25, 0.25, 0.25], - ] + let pspm = [[1.0, 0.0, 0.0, 0.0], [0.25, 0.25, 0.25, 0.25]] let motif = @src.MemeMotif::new("m1", pspm, alphabet="ACGT") let ic = motif.information_content() // First position has max info (2 bits), second has 0 diff --git a/test/moonbit/metagenomeseq_test.mbt b/test/moonbit/metagenomeseq_test.mbt index 82c75bbd..36f66e2d 100644 --- a/test/moonbit/metagenomeseq_test.mbt +++ b/test/moonbit/metagenomeseq_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for metagenomeSeq module. - test "MRexperiment creation" { let counts = [[10, 20], [30, 40]] let taxa_names = ["Bacteroides", "Firmicutes"] @@ -10,25 +9,31 @@ test "MRexperiment creation" { assert_eq(obj.sample_names.length(), 2) } +///| test "mg_normalize_counts" { let obj = @src.create_example_mrexperiment() let normalized = @src.mg_normalize_counts(obj) assert_eq(normalized.counts.length(), obj.counts.length()) } +///| test "mg_calculate_zero_inflation" { let obj = @src.create_example_mrexperiment() let zero_probs = @src.mg_calculate_zero_inflation(obj) assert_eq(zero_probs.length(), obj.taxa_names.length()) } +///| test "mg_test_zero_inflated" { let obj = @src.create_example_mrexperiment() - let group = ["control", "control", "control", "treatment", "treatment", "treatment"] + let group = [ + "control", "control", "control", "treatment", "treatment", "treatment", + ] let results = @src.mg_test_zero_inflated(obj, group) assert_eq(results.length(), obj.taxa_names.length()) } +///| test "MRSampleData creation" { let sd = @src.MRSampleData::new("control", 100000) assert_eq(sd.group, "control") diff --git a/test/moonbit/methyl_seekr_test.mbt b/test/moonbit/methyl_seekr_test.mbt index 043b2b92..a0cbda51 100644 --- a/test/moonbit/methyl_seekr_test.mbt +++ b/test/moonbit/methyl_seekr_test.mbt @@ -219,7 +219,9 @@ test "msr_tiling_single_site" { assert_eq(@src.MethylTile::start(tiles[0]), 0) assert_eq(@src.MethylTile::end(tiles[0]), 1000) assert_eq(@src.MethylTile::coverage(tiles[0]), 10) - assert_true((@src.MethylTile::methylation_level(tiles[0]) - 0.4).abs() < 0.000001) + assert_true( + (@src.MethylTile::methylation_level(tiles[0]) - 0.4).abs() < 0.000001, + ) // Newly created tiles default to "FMR". assert_eq(@src.MethylTile::region_type(tiles[0]), "FMR") } @@ -229,9 +231,24 @@ test "msr_tiling_boundaries" { // A site at position 999 falls in tile [0,1000); at 1000 in [1000,2000). let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr1", position=999, methylated=1, unmethylated=0), - @src.CytosineSite::new(chr="chr1", position=1000, methylated=1, unmethylated=0), - @src.CytosineSite::new(chr="chr1", position=2500, methylated=0, unmethylated=1), + @src.CytosineSite::new( + chr="chr1", + position=999, + methylated=1, + unmethylated=0, + ), + @src.CytosineSite::new( + chr="chr1", + position=1000, + methylated=1, + unmethylated=0, + ), + @src.CytosineSite::new( + chr="chr1", + position=2500, + methylated=0, + unmethylated=1, + ), ] let tiles = @src.tile_methylation(sites, params) assert_eq(tiles.length(), 3) @@ -248,8 +265,18 @@ test "msr_tiling_methylation_computation" { // Two sites in the same tile: (3,7) and (2,8) -> meth=5, unmeth=15 -> 0.25. let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr1", position=100, methylated=3, unmethylated=7), - @src.CytosineSite::new(chr="chr1", position=200, methylated=2, unmethylated=8), + @src.CytosineSite::new( + chr="chr1", + position=100, + methylated=3, + unmethylated=7, + ), + @src.CytosineSite::new( + chr="chr1", + position=200, + methylated=2, + unmethylated=8, + ), ] let tiles = @src.tile_methylation(sites, params) assert_eq(tiles.length(), 1) @@ -264,9 +291,24 @@ test "msr_tiling_coverage_aggregation" { // Coverage from multiple sites is summed per tile. let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr1", position=100, methylated=1, unmethylated=4), - @src.CytosineSite::new(chr="chr1", position=200, methylated=2, unmethylated=3), - @src.CytosineSite::new(chr="chr1", position=300, methylated=0, unmethylated=5), + @src.CytosineSite::new( + chr="chr1", + position=100, + methylated=1, + unmethylated=4, + ), + @src.CytosineSite::new( + chr="chr1", + position=200, + methylated=2, + unmethylated=3, + ), + @src.CytosineSite::new( + chr="chr1", + position=300, + methylated=0, + unmethylated=5, + ), ] let tiles = @src.tile_methylation(sites, params) assert_eq(tiles.length(), 1) @@ -282,9 +324,24 @@ test "msr_tiling_multiple_chromosomes" { // Tiles are sorted by chromosome name then start. let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr2", position=50, methylated=1, unmethylated=1), - @src.CytosineSite::new(chr="chr1", position=100, methylated=1, unmethylated=1), - @src.CytosineSite::new(chr="chr1", position=1500, methylated=1, unmethylated=1), + @src.CytosineSite::new( + chr="chr2", + position=50, + methylated=1, + unmethylated=1, + ), + @src.CytosineSite::new( + chr="chr1", + position=100, + methylated=1, + unmethylated=1, + ), + @src.CytosineSite::new( + chr="chr1", + position=1500, + methylated=1, + unmethylated=1, + ), ] let tiles = @src.tile_methylation(sites, params) assert_eq(tiles.length(), 3) @@ -786,7 +843,12 @@ test "msr_edge_empty_pipeline" { test "msr_edge_single_site" { let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr1", position=100, methylated=0, unmethylated=20), + @src.CytosineSite::new( + chr="chr1", + position=100, + methylated=0, + unmethylated=20, + ), ] let regions = @src.call_methylation_regimes(sites, params) // Single tile of 1000 bp >= min_region_size (500) -> one UMR region. @@ -848,7 +910,12 @@ test "msr_edge_zero_coverage_sites" { // which classifies as FMR (insufficient coverage). let params = @src.MethylSeekRParams::default() let sites = [ - @src.CytosineSite::new(chr="chr1", position=100, methylated=0, unmethylated=0), + @src.CytosineSite::new( + chr="chr1", + position=100, + methylated=0, + unmethylated=0, + ), ] let tiles = @src.tile_methylation(sites, params) assert_eq(tiles.length(), 1) diff --git a/test/moonbit/methylkit_test.mbt b/test/moonbit/methylkit_test.mbt index 566f93c7..5f48f9db 100644 --- a/test/moonbit/methylkit_test.mbt +++ b/test/moonbit/methylkit_test.mbt @@ -22,7 +22,9 @@ test "methylkit_cytosine_coverage_pct" { // 50% methylation let mc_half = @src.MethylCytosine::new("chr1", 200, "-", "CG", 5, 5) assert_eq(@src.MethylCytosine::coverage(mc_half), 10) - assert_true((@src.MethylCytosine::methylation_pct(mc_half) - 50.0).abs() < 0.001) + assert_true( + (@src.MethylCytosine::methylation_pct(mc_half) - 50.0).abs() < 0.001, + ) // Zero coverage -> coverage 0 and pct 0.0 let mc0 = @src.MethylCytosine::new("chr1", 300, "+", "CG", 0, 0) assert_eq(@src.MethylCytosine::coverage(mc0), 0) @@ -151,7 +153,9 @@ test "methylkit_config_defaults" { assert_eq(@src.MethylKitConfig::min_coverage(config), 10) assert_eq(@src.MethylKitConfig::max_coverage(config), 999) assert_eq(@src.MethylKitConfig::context_filter(config), "CG") - assert_true((@src.MethylKitConfig::min_perc_samples(config) - 0.6).abs() < 0.001) + assert_true( + (@src.MethylKitConfig::min_perc_samples(config) - 0.6).abs() < 0.001, + ) assert_true((@src.MethylKitConfig::min_diff(config) - 25.0).abs() < 0.001) assert_true((@src.MethylKitConfig::q_threshold(config) - 0.05).abs() < 0.001) } @@ -162,7 +166,9 @@ test "methylkit_config_with_params" { assert_eq(@src.MethylKitConfig::min_coverage(config), 5) assert_eq(@src.MethylKitConfig::max_coverage(config), 500) assert_eq(@src.MethylKitConfig::context_filter(config), "CHG") - assert_true((@src.MethylKitConfig::min_perc_samples(config) - 0.8).abs() < 0.001) + assert_true( + (@src.MethylKitConfig::min_perc_samples(config) - 0.8).abs() < 0.001, + ) assert_true((@src.MethylKitConfig::min_diff(config) - 30.0).abs() < 0.001) assert_true((@src.MethylKitConfig::q_threshold(config) - 0.01).abs() < 0.001) } @@ -466,9 +472,13 @@ test "methylkit_sample_data" { assert_eq(@src.MethylSample::n_cpgs(control), 20) assert_eq(@src.MethylSample::n_cpgs(treatment), 20) // Control: every CpG at 50% methylation. - assert_true((@src.MethylSample::mean_methylation(control) - 50.0).abs() < 0.001) + assert_true( + (@src.MethylSample::mean_methylation(control) - 50.0).abs() < 0.001, + ) // Treatment: 5 hyper (90%) + 5 hypo (10%) + 10 same (50%) -> mean 50%. - assert_true((@src.MethylSample::mean_methylation(treatment) - 50.0).abs() < 0.001) + assert_true( + (@src.MethylSample::mean_methylation(treatment) - 50.0).abs() < 0.001, + ) // Treatment CpG 0: hyper-methylated (18, 2). let t0 = @src.MethylSample::coverage(treatment)[0] assert_eq(@src.MethylCytosine::methylated(t0), 18) diff --git a/test/moonbit/microbiome_test.mbt b/test/moonbit/microbiome_test.mbt index 3337a2e9..c91708fa 100644 --- a/test/moonbit/microbiome_test.mbt +++ b/test/moonbit/microbiome_test.mbt @@ -11,12 +11,14 @@ test "calc_observed basic" { assert_eq(observed, 4.0) } +///| test "calc_observed all zeros" { let counts : Array[Double] = [0.0, 0.0, 0.0] let observed = @src.calc_observed(counts) assert_eq(observed, 0.0) } +///| test "calc_shannon even" { let counts : Array[Double] = [10.0, 10.0, 10.0, 10.0] let shannon = @src.calc_shannon(counts) @@ -24,12 +26,14 @@ test "calc_shannon even" { assert_true(shannon > 1.3 && shannon < 1.5) } +///| test "calc_shannon single taxon" { let counts : Array[Double] = [100.0] let shannon = @src.calc_shannon(counts) assert_eq(shannon, 0.0) } +///| test "calc_simpson even" { let counts : Array[Double] = [10.0, 10.0, 10.0, 10.0] let simpson = @src.calc_simpson(counts) @@ -37,12 +41,14 @@ test "calc_simpson even" { assert_true(simpson > 0.7 && simpson < 0.8) } +///| test "calc_simpson single taxon" { let counts : Array[Double] = [100.0] let simpson = @src.calc_simpson(counts) assert_eq(simpson, 0.0) } +///| test "calc_inv_simpson even" { let counts : Array[Double] = [10.0, 10.0, 10.0, 10.0] let inv_simpson = @src.calc_inv_simpson(counts) @@ -50,6 +56,7 @@ test "calc_inv_simpson even" { assert_true(inv_simpson > 3.5 && inv_simpson < 4.5) } +///| test "calc_pielou_evenness even" { let counts : Array[Double] = [10.0, 10.0, 10.0, 10.0] let pielou = @src.calc_pielou_evenness(counts) @@ -57,12 +64,14 @@ test "calc_pielou_evenness even" { assert_true(pielou > 0.95) } +///| test "calc_pielou_evenness single taxon" { let counts : Array[Double] = [100.0] let pielou = @src.calc_pielou_evenness(counts) assert_eq(pielou, 0.0) } +///| test "calc_chao1 basic" { let counts : Array[Double] = [10.0, 5.0, 3.0, 1.0, 1.0, 2.0, 2.0] let chao1 = @src.calc_chao1(counts) @@ -71,12 +80,14 @@ test "calc_chao1 basic" { assert_true(chao1 >= 7.0) } +///| test "calc_fisher_alpha basic" { let counts : Array[Double] = [100.0, 50.0, 30.0, 20.0, 10.0] let alpha = @src.calc_fisher_alpha(counts) assert_true(alpha > 0.0) } +///| test "calc_alpha_diversity all indices" { let counts : Array[Double] = [10.0, 20.0, 15.0, 5.0, 25.0] let div = @src.calc_alpha_diversity(counts) @@ -87,6 +98,7 @@ test "calc_alpha_diversity all indices" { assert_true(div.chao1 >= div.observed) } +///| test "calc_alpha_diversity_table multiple samples" { let otu_table = @src.create_example_otu_table() let div_table = @src.calc_alpha_diversity_table(otu_table) @@ -99,6 +111,7 @@ test "calc_alpha_diversity_table multiple samples" { // Beta Diversity Tests // ============================================================ +///| test "bray_curtis identical" { let x : Array[Double] = [10.0, 20.0, 30.0] let y : Array[Double] = [10.0, 20.0, 30.0] @@ -106,6 +119,7 @@ test "bray_curtis identical" { assert_eq(bc, 0.0) } +///| test "bray_curtis completely different" { let x : Array[Double] = [10.0, 0.0, 0.0] let y : Array[Double] = [0.0, 20.0, 30.0] @@ -114,6 +128,7 @@ test "bray_curtis completely different" { assert_eq(bc, 1.0) } +///| test "bray_curtis partial overlap" { let x : Array[Double] = [10.0, 20.0, 0.0] let y : Array[Double] = [10.0, 0.0, 20.0] @@ -121,6 +136,7 @@ test "bray_curtis partial overlap" { assert_true(bc > 0.0 && bc < 1.0) } +///| test "jaccard_distance identical" { let x : Array[Double] = [10.0, 20.0, 30.0] let y : Array[Double] = [5.0, 10.0, 15.0] @@ -128,6 +144,7 @@ test "jaccard_distance identical" { assert_eq(jaccard, 0.0) } +///| test "jaccard_distance no overlap" { let x : Array[Double] = [10.0, 0.0, 0.0] let y : Array[Double] = [0.0, 20.0, 30.0] @@ -135,6 +152,7 @@ test "jaccard_distance no overlap" { assert_eq(jaccard, 1.0) } +///| test "jaccard_distance partial" { let x : Array[Double] = [10.0, 20.0, 0.0] let y : Array[Double] = [10.0, 0.0, 20.0] @@ -143,6 +161,7 @@ test "jaccard_distance partial" { assert_true(jaccard > 0.6 && jaccard < 0.7) } +///| test "jensen_shannon_divergence identical" { let x : Array[Double] = [10.0, 20.0, 30.0] let y : Array[Double] = [10.0, 20.0, 30.0] @@ -150,6 +169,7 @@ test "jensen_shannon_divergence identical" { assert_eq(jsd, 0.0) } +///| test "jensen_shannon_divergence symmetric" { let x : Array[Double] = [10.0, 20.0, 30.0] let y : Array[Double] = [30.0, 20.0, 10.0] @@ -158,6 +178,7 @@ test "jensen_shannon_divergence symmetric" { assert_true((jsd_xy - jsd_yx).abs() < 0.0001) } +///| test "calc_beta_diversity_matrix bray" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "bray") @@ -170,6 +191,7 @@ test "calc_beta_diversity_matrix bray" { assert_true((dist_matrix[0][1] - dist_matrix[1][0]).abs() < 0.0001) } +///| test "calc_beta_diversity_matrix jaccard" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "jaccard") @@ -177,6 +199,7 @@ test "calc_beta_diversity_matrix jaccard" { assert_true(dist_matrix[0][3] > 0.0) } +///| test "calc_beta_diversity_matrix jsd" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "jsd") @@ -188,22 +211,24 @@ test "calc_beta_diversity_matrix jsd" { // PCoA Tests // ============================================================ +///| test "pcoa basic" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "bray") let pcoa_result = @src.pcoa(dist_matrix) - + assert_true(pcoa_result.eigenvalues.length() > 0) assert_true(pcoa_result.vectors.length() == 6) assert_true(pcoa_result.variance_explained.length() > 0) assert_true(pcoa_result.cumulative_variance.length() > 0) } +///| test "pcoa eigenvalues positive" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "bray") let pcoa_result = @src.pcoa(dist_matrix) - + let mut i = 0 while i < pcoa_result.eigenvalues.length() { assert_true(pcoa_result.eigenvalues[i] >= 0.0) @@ -211,17 +236,19 @@ test "pcoa eigenvalues positive" { } } +///| test "pcoa has vectors" { let otu_table = @src.create_example_otu_table() let dist_matrix = @src.calc_beta_diversity_matrix(otu_table, "bray") let pcoa_result = @src.pcoa(dist_matrix) - + assert_true(pcoa_result.vectors.length() == 6) if pcoa_result.vectors.length() > 0 { assert_true(pcoa_result.vectors[0].length() > 0) } } +///| test "pcoa single sample" { let dist_matrix : Array[Array[Double]] = [[0.0]] let pcoa_result = @src.pcoa(dist_matrix) @@ -232,41 +259,50 @@ test "pcoa single sample" { // Differential Abundance Tests // ============================================================ +///| test "differential_abundance basic" { let otu_table = @src.create_example_otu_table() let taxa_names = @src.get_example_taxa_names() let group1 : Array[Int] = [0, 1, 2] let group2 : Array[Int] = [3, 4, 5] - - let results = @src.differential_abundance(otu_table, taxa_names, group1, group2) - + + let results = @src.differential_abundance( + otu_table, taxa_names, group1, group2, + ) + assert_eq(results.length(), 10) assert_eq(results[0].taxa, "Bacteroides") assert_true(results[0].p_value >= 0.0 && results[0].p_value <= 1.0) assert_true(results[0].adjusted_p_value >= results[0].p_value) } +///| test "differential_abundance log2_fc direction" { let otu_table = @src.create_example_otu_table() let taxa_names = @src.get_example_taxa_names() - let group1 : Array[Int] = [0, 1, 2] // control - let group2 : Array[Int] = [3, 4, 5] // treatment - - let results = @src.differential_abundance(otu_table, taxa_names, group1, group2) - + let group1 : Array[Int] = [0, 1, 2] // control + let group2 : Array[Int] = [3, 4, 5] // treatment + + let results = @src.differential_abundance( + otu_table, taxa_names, group1, group2, + ) + // Taxon 0 (Bacteroides) is high in control, so log2FC should be negative assert_true(results[0].log2_fold_change < 0.0) // Taxon 1 (Prevotella) is high in treatment, so log2FC should be positive assert_true(results[1].log2_fold_change > 0.0) } +///| test "differential_abundance empty input" { let otu_table : Array[Array[Double]] = [] let taxa_names : Array[String] = [] let group1 : Array[Int] = [0, 1] let group2 : Array[Int] = [2, 3] - - let results = @src.differential_abundance(otu_table, taxa_names, group1, group2) + + let results = @src.differential_abundance( + otu_table, taxa_names, group1, group2, + ) assert_eq(results.length(), 0) } @@ -274,12 +310,14 @@ test "differential_abundance empty input" { // Helper function tests // ============================================================ +///| test "create_example_otu_table dimensions" { let otu_table = @src.create_example_otu_table() - assert_eq(otu_table.length(), 10) // 10 taxa - assert_eq(otu_table[0].length(), 6) // 6 samples + assert_eq(otu_table.length(), 10) // 10 taxa + assert_eq(otu_table[0].length(), 6) // 6 samples } +///| test "get_example_taxa_names" { let names = @src.get_example_taxa_names() assert_eq(names.length(), 10) diff --git a/test/moonbit/milo_test.mbt b/test/moonbit/milo_test.mbt new file mode 100644 index 00000000..01ec40e5 --- /dev/null +++ b/test/moonbit/milo_test.mbt @@ -0,0 +1,697 @@ +///| +/// Tests for the Bioconductor miloR-inspired neighbourhood DA workflow. + +///| +fn milo_test_assert_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort("miloR values differ beyond tolerance") + } +} + +///| +fn milo_test_small_graph() -> @src.MiloExperiment { + @src.milo_build_graph( + [[0.0], [1.0], [3.0], [10.0]], + cell_names=["A", "B", "C", "D"], + k=1, + ) catch { + _ => abort("valid small Milo graph should build") + } +} + +///| +fn milo_test_counted() -> (@src.MiloExperiment, @src.MiloDesign) { + let data = @src.milo_sample_data() + let graph = @src.milo_build_graph( + data.coordinates, + cell_names=data.cell_names, + k=8, + dimensions=2, + ) catch { + _ => abort("sample Milo graph should build") + } + let neighborhoods = graph.make_neighborhoods(proportion=0.5, seed=19) catch { + _ => abort("sample Milo neighborhoods should build") + } + let counted = neighborhoods.count_cells(data.sample_ids) catch { + _ => abort("sample Milo cells should count") + } + let design = @src.milo_binary_design(data.sample_conditions, reference="A") catch { + _ => abort("sample Milo design should build") + } + (counted, design) +} + +///| +test "miloR: sample data contains balanced replicated samples" { + let data = @src.milo_sample_data() + assert_eq(data.coordinates.length(), 32) + assert_eq(data.cell_names.length(), 32) + assert_eq(data.sample_ids.length(), 32) + assert_eq(data.sample_names, ["A1", "A2", "A3", "A4", "B1", "B2", "B3", "B4"]) + assert_eq(data.sample_conditions, ["A", "A", "A", "A", "B", "B", "B", "B"]) +} + +///| +test "miloR: builds exact KNN graph and generates cell names" { + let graph = @src.milo_build_graph([[0.0], [1.0], [3.0]], k=1) catch { + _ => abort("valid graph should build") + } + assert_eq(graph.cell_names, ["Cell1", "Cell2", "Cell3"]) + assert_eq(graph.dimensions, 1) + assert_eq(graph.k, 1) + assert_eq(graph.knn_indices, [[1], [0], [1]]) +} + +///| +test "miloR: exact KNN excludes each query cell" { + let graph = milo_test_small_graph() + for cell in 0.. abort("valid graph should build") + } + assert_eq(graph.knn_indices[0], [1, 2, 3]) + assert_eq(graph.knn_distances[0], [1.0, 3.0, 10.0]) +} + +///| +test "miloR: KNN ties are resolved by cell index" { + let graph = @src.milo_build_graph([[-1.0], [0.0], [1.0]], k=2) catch { + _ => abort("valid tied graph should build") + } + assert_eq(graph.knn_indices[1], [0, 2]) +} + +///| +test "miloR: directed KNN edges are symmetrized" { + let graph = milo_test_small_graph() + assert_eq(graph.graph_neighbors, [[1], [0, 2], [1, 3], [2]]) +} + +///| +test "miloR: selected dimensions control graph distances" { + let graph = @src.milo_build_graph( + [[0.0, 100.0], [1.0, 0.0], [2.0, 100.0]], + k=1, + dimensions=1, + ) catch { + _ => abort("valid dimension subset should build") + } + assert_eq(graph.knn_indices[0], [1]) + assert_eq(graph.knn_indices[2], [1]) +} + +///| +test "miloR: graph construction copies coordinate input" { + let coordinates = [[0.0], [1.0], [2.0]] + let graph = @src.milo_build_graph(coordinates, k=1) catch { + _ => abort("valid graph should build") + } + coordinates[0][0] = 99.0 + assert_eq(graph.coordinates[0][0], 0.0) +} + +///| +test "miloR: graph rejects malformed coordinates" { + let too_small = try { + ignore(@src.milo_build_graph([[0.0]], k=1)) + false + } catch { + MiloError(_) => true + } + let ragged = try { + ignore(@src.milo_build_graph([[0.0], [1.0, 2.0]], k=1)) + false + } catch { + MiloError(_) => true + } + let non_finite = try { + ignore(@src.milo_build_graph([[0.0], [@double.not_a_number]], k=1)) + false + } catch { + MiloError(_) => true + } + assert_true(too_small) + assert_true(ragged) + assert_true(non_finite) +} + +///| +test "miloR: graph validates k and dimensions" { + let bad_k = try { + ignore(@src.milo_build_graph([[0.0], [1.0]], k=2)) + false + } catch { + MiloError(_) => true + } + let bad_dimensions = try { + ignore(@src.milo_build_graph([[0.0], [1.0]], k=1, dimensions=2)) + false + } catch { + MiloError(_) => true + } + assert_true(bad_k) + assert_true(bad_dimensions) +} + +///| +test "miloR: graph validates cell names" { + let wrong_length = try { + ignore(@src.milo_build_graph([[0.0], [1.0], [2.0]], cell_names=["A"], k=1)) + false + } catch { + MiloError(_) => true + } + let duplicates = try { + ignore( + @src.milo_build_graph( + [[0.0], [1.0], [2.0]], + cell_names=["A", "A", "B"], + k=1, + ), + ) + false + } catch { + MiloError(_) => true + } + assert_true(wrong_length) + assert_true(duplicates) +} + +///| +test "miloR: unrefined sampling is deterministic" { + let graph = @src.milo_build_graph( + [[0.0], [1.0], [2.0], [3.0], [4.0], [5.0]], + k=2, + ) catch { + _ => abort("valid graph should build") + } + let first = graph.make_neighborhoods(proportion=0.5, refined=false, seed=7) catch { + _ => abort("valid neighborhoods should build") + } + let second = graph.make_neighborhoods(proportion=0.5, refined=false, seed=7) catch { + _ => abort("valid neighborhoods should build") + } + assert_eq(first.neighborhood_indices, second.neighborhood_indices) + assert_eq(first.neighborhood_indices.length(), 3) +} + +///| +test "miloR: refined sampling collapses duplicate representatives" { + let data = @src.milo_sample_data() + let graph = @src.milo_build_graph(data.coordinates, k=4) catch { + _ => abort("valid graph should build") + } + let refined = graph.make_neighborhoods(proportion=0.5, refined=true, seed=11) catch { + _ => abort("valid refined neighborhoods should build") + } + assert_true(refined.neighborhoods.length() > 0) + assert_true(refined.neighborhoods.length() <= 16) + for index in refined.neighborhood_indices { + assert_true(index >= 0 && index < data.coordinates.length()) + } +} + +///| +test "miloR: neighborhoods include index and undirected graph neighbors" { + let graph = milo_test_small_graph() + let experiment = graph.make_neighborhoods( + proportion=0.75, + refined=false, + seed=3, + ) catch { + _ => abort("valid neighborhoods should build") + } + for neighborhood in 0.. true + } + let one = try { + ignore(graph.make_neighborhoods(proportion=1.0)) + false + } catch { + MiloError(_) => true + } + assert_true(zero) + assert_true(one) +} + +///| +test "miloR: counts neighborhood cells by first-seen sample order" { + let graph = milo_test_small_graph() + let neighborhoods = graph.make_neighborhoods( + proportion=0.75, + refined=false, + seed=5, + ) catch { + _ => abort("valid neighborhoods should build") + } + let counted = neighborhoods.count_cells(["S2", "S1", "S1", "S2"]) catch { + _ => abort("valid sample IDs should count") + } + assert_eq(counted.sample_names, ["S2", "S1"]) + for index in 0.. abort("valid neighborhoods should build") + } + let counted = neighborhoods.count_cells(["S1", "S1", "S2", "S2"]) catch { + _ => abort("valid sample IDs should count") + } + assert_eq(neighborhoods.sample_names.length(), 0) + assert_eq(counted.sample_names, ["S1", "S2"]) +} + +///| +test "miloR: counting validates workflow and sample IDs" { + let graph = milo_test_small_graph() + let before_neighborhoods = try { + ignore(graph.count_cells(["S1", "S1", "S2", "S2"])) + false + } catch { + MiloError(_) => true + } + let neighborhoods = graph.make_neighborhoods(proportion=0.5, refined=false) catch { + _ => abort("valid neighborhoods should build") + } + let wrong_length = try { + ignore(neighborhoods.count_cells(["S1"])) + false + } catch { + MiloError(_) => true + } + let empty = try { + ignore(neighborhoods.count_cells(["S1", "", "S2", "S2"])) + false + } catch { + MiloError(_) => true + } + assert_true(before_neighborhoods) + assert_true(wrong_length) + assert_true(empty) +} + +///| +test "miloR: computes mean feature expression per neighborhood" { + let graph = milo_test_small_graph() + let neighborhoods = graph.make_neighborhoods( + proportion=0.5, + refined=false, + seed=2, + ) catch { + _ => abort("valid neighborhoods should build") + } + let expression = neighborhoods.neighborhood_expression([[0.0, 2.0, 4.0, 8.0]]) catch { + _ => abort("valid expression should aggregate") + } + assert_eq(expression.length(), 1) + assert_eq(expression[0].length(), neighborhoods.neighborhoods.length()) + for index in 0.. abort("valid neighborhoods should build") + } + let wrong_length = try { + ignore(neighborhoods.neighborhood_expression([[1.0]])) + false + } catch { + MiloError(_) => true + } + let non_finite = try { + ignore( + neighborhoods.neighborhood_expression([ + [@double.not_a_number, 1.0, 2.0, 3.0], + ]), + ) + false + } catch { + MiloError(_) => true + } + assert_true(wrong_length) + assert_true(non_finite) +} + +///| +test "miloR: neighborhood overlap matrix is symmetric" { + let (experiment, _) = milo_test_counted() + let overlaps = experiment.neighborhood_overlap_matrix() + assert_eq(overlaps.length(), experiment.neighborhoods.length()) + for left in 0.. abort("valid binary design should build") + } + assert_eq(design.levels, ["control", "treated"]) + assert_eq(design.reference, "control") + assert_eq(design.coefficient_name, "treated_vs_control") + assert_eq(design.matrix, [[1.0, 1.0], [1.0, 0.0], [1.0, 1.0], [1.0, 0.0]]) +} + +///| +test "miloR: binary design validates conditions and reference" { + let one_level = try { + ignore(@src.milo_binary_design(["A", "A"])) + false + } catch { + MiloError(_) => true + } + let three_levels = try { + ignore(@src.milo_binary_design(["A", "B", "C"])) + false + } catch { + MiloError(_) => true + } + let absent_reference = try { + ignore(@src.milo_binary_design(["A", "B"], reference="C")) + false + } catch { + MiloError(_) => true + } + assert_true(one_level) + assert_true(three_levels) + assert_true(absent_reference) +} + +///| +test "miloR: graph spatial FDR supports k-distance weighting" { + let (experiment, _) = milo_test_counted() + let p_values = Array::make(experiment.neighborhoods.length(), 0.05) + let adjusted = @src.milo_graph_spatial_fdr(experiment, p_values) catch { + _ => abort("valid spatial FDR should compute") + } + assert_eq(adjusted.length(), p_values.length()) + for value in adjusted { + assert_true(value >= 0.0 && value <= 1.0) + } +} + +///| +test "miloR: spatial FDR supports all graph weighting schemes" { + let (experiment, _) = milo_test_counted() + let p_values : Array[Double] = [] + for index in 0.. abort("neighbor-distance FDR should compute") + } + let maximum = @src.milo_graph_spatial_fdr( + experiment, + p_values, + weighting=@src.milo_max_distance_weighting(), + ) catch { + _ => abort("maximum-distance FDR should compute") + } + let overlap = @src.milo_graph_spatial_fdr( + experiment, + p_values, + weighting=@src.milo_graph_overlap_weighting(), + ) catch { + _ => abort("graph-overlap FDR should compute") + } + assert_eq(neighbor.length(), p_values.length()) + assert_eq(maximum.length(), p_values.length()) + assert_eq(overlap.length(), p_values.length()) +} + +///| +test "miloR: disabled spatial weighting returns missing values" { + let (experiment, _) = milo_test_counted() + let adjusted = @src.milo_graph_spatial_fdr( + experiment, + Array::make(experiment.neighborhoods.length(), 0.5), + weighting=@src.milo_no_spatial_weighting(), + ) catch { + _ => abort("disabled spatial FDR should return missing values") + } + for value in adjusted { + assert_true(value.is_nan()) + } +} + +///| +test "miloR: spatial FDR validates p-values" { + let (experiment, _) = milo_test_counted() + let wrong_length = try { + ignore(@src.milo_graph_spatial_fdr(experiment, [])) + false + } catch { + MiloError(_) => true + } + let invalid = Array::make(experiment.neighborhoods.length(), 0.5) + invalid[0] = 1.5 + let out_of_range = try { + ignore(@src.milo_graph_spatial_fdr(experiment, invalid)) + false + } catch { + MiloError(_) => true + } + assert_true(wrong_length) + assert_true(out_of_range) +} + +///| +test "miloR: complete sample pipeline reports expected dimensions" { + let (experiment, _) = milo_test_counted() + assert_eq(experiment.coordinates.length(), 32) + assert_eq(experiment.sample_names.length(), 8) + assert_true(experiment.neighborhoods.length() > 1) + assert_true(experiment.summary().contains("cells=32")) + assert_true(experiment.summary().contains("samples=8")) +} + +///| +test "miloR: NB GLM returns one finite result per neighborhood" { + let (experiment, design) = milo_test_counted() + let results = experiment.test_neighborhoods(design.matrix, 1) catch { + _ => abort("valid Milo DA test should fit") + } + assert_eq(results.length(), experiment.neighborhoods.length()) + for result in results { + assert_true(result.tested) + assert_true(!result.log_fc.is_nan()) + assert_true(!result.statistic.is_nan()) + assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) + assert_true(result.fdr >= 0.0 && result.fdr <= 1.0) + assert_true(result.spatial_fdr >= 0.0 && result.spatial_fdr <= 1.0) + assert_true(result.dispersion > 0.0) + } +} + +///| +test "miloR: DA direction follows left and right state enrichment" { + let (experiment, design) = milo_test_counted() + let results = experiment.test_neighborhoods(design.matrix, 1, cell_sizes=[ + 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, + ]) catch { + _ => abort("valid Milo DA test should fit") + } + let mut left = 0 + let mut right = 0 + for result in results { + let x = experiment.coordinates[result.index_cell][0] + let counts = experiment.neighborhood_counts[result.neighborhood] + let condition_a = counts[0] + counts[1] + counts[2] + counts[3] + let condition_b = counts[4] + counts[5] + counts[6] + counts[7] + if x < 2.5 { + assert_true(condition_a >= condition_b) + if condition_a > condition_b { + assert_true(result.log_fc < 0.0) + left = left + 1 + } + } else { + assert_true(condition_b >= condition_a) + if condition_b > condition_a { + assert_true(result.log_fc > 0.0) + right = right + 1 + } + } + } + assert_true(left > 0) + assert_true(right > 0) +} + +///| +test "miloR: mean threshold marks low-count neighborhoods untested" { + let (experiment, design) = milo_test_counted() + let results = experiment.test_neighborhoods(design.matrix, 1, min_mean=100.0) catch { + _ => abort("valid threshold should apply") + } + for result in results { + assert_false(result.tested) + assert_eq(result.p_value, 1.0) + assert_eq(result.fdr, 1.0) + } +} + +///| +test "miloR: DA supports explicit cell-size normalization" { + let (experiment, design) = milo_test_counted() + let results = experiment.test_neighborhoods(design.matrix, 1, cell_sizes=[ + 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, 4.0, + ]) catch { + _ => abort("valid cell sizes should normalize") + } + assert_eq(results.length(), experiment.neighborhoods.length()) +} + +///| +test "miloR: DA supports disabling spatial correction" { + let (experiment, design) = milo_test_counted() + let results = experiment.test_neighborhoods( + design.matrix, + 1, + weighting=@src.milo_no_spatial_weighting(), + ) catch { + _ => abort("valid DA test should fit without spatial correction") + } + for result in results { + assert_true(result.spatial_fdr.is_nan()) + } +} + +///| +test "miloR: DA validates design dimensions rank and contrast" { + let (experiment, _) = milo_test_counted() + let wrong_rows = try { + ignore(experiment.test_neighborhoods([[1.0, 0.0]], 1)) + false + } catch { + MiloError(_) => true + } + let singular : Array[Array[Double]] = [] + for _ in 0..<8 { + singular.push([1.0, 1.0]) + } + let bad_rank = try { + ignore(experiment.test_neighborhoods(singular, 1)) + false + } catch { + MiloError(_) => true + } + let design = @src.milo_binary_design(["A", "A", "A", "A", "B", "B", "B", "B"]) catch { + _ => abort("valid design should build") + } + let bad_contrast = try { + ignore(experiment.test_neighborhoods(design.matrix, 2)) + false + } catch { + MiloError(_) => true + } + assert_true(wrong_rows) + assert_true(bad_rank) + assert_true(bad_contrast) +} + +///| +test "miloR: DA validates minimum mean and cell sizes" { + let (experiment, design) = milo_test_counted() + let bad_mean = try { + ignore(experiment.test_neighborhoods(design.matrix, 1, min_mean=-1.0)) + false + } catch { + MiloError(_) => true + } + let bad_sizes = try { + ignore(experiment.test_neighborhoods(design.matrix, 1, cell_sizes=[1.0])) + false + } catch { + MiloError(_) => true + } + assert_true(bad_mean) + assert_true(bad_sizes) +} + +///| +test "miloR: SCE constructor uses named reduced dimensions and cells" { + let sce = @src.SingleCellExperiment::new([[1.0, 2.0, 3.0, 4.0]], ["Gene1"], [ + "A", "B", "C", "D", + ]) + let with_pca = @src.sce_set_reduced_dim(sce, "PCA", [ + [0.0], + [1.0], + [2.0], + [3.0], + ]) + let graph = @src.milo_from_single_cell_experiment(with_pca, k=1) catch { + _ => abort("valid SCE reduced dimensions should build a graph") + } + assert_eq(graph.cell_names, ["A", "B", "C", "D"]) + assert_eq(graph.knn_indices.length(), 4) +} + +///| +test "miloR: SCE constructor rejects missing reduced dimensions" { + let sce = @src.SingleCellExperiment::new([[1.0, 2.0]], ["Gene1"], ["A", "B"]) + let missing = try { + ignore(@src.milo_from_single_cell_experiment(sce, k=1)) + false + } catch { + MiloError(_) => true + } + assert_true(missing) +} diff --git a/test/moonbit/missmethyl_test.mbt b/test/moonbit/missmethyl_test.mbt index d87af891..c2d43cfc 100644 --- a/test/moonbit/missmethyl_test.mbt +++ b/test/moonbit/missmethyl_test.mbt @@ -119,7 +119,7 @@ test "mm_beta_to_m_basic" { assert_true(m.abs() < 0.01) // beta = 0.25 -> M = log2(0.25/0.75) = log2(1/3) ≈ -1.585 let m2 = @src.mm_beta_to_m(0.25) - assert_true((m2 - (-1.585)).abs() < 0.01) + assert_true((m2 - -1.585).abs() < 0.01) // beta = 0.75 -> M = log2(0.75/0.25) = log2(3) ≈ 1.585 let m3 = @src.mm_beta_to_m(0.75) assert_true((m3 - 1.585).abs() < 0.01) @@ -360,14 +360,11 @@ test "mm_go_enrichment_sample" { let annotations = @src.mm_sample_annotations() let go_db = @src.mm_sample_go_database() let all_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", - "cg006", "cg007", "cg008", "cg009", "cg010", - "cg011", "cg012", "cg013", "cg014", "cg015", - "cg016", "cg017", "cg018", "cg019", "cg020", - ] - let sig_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", + "cg001", "cg002", "cg003", "cg004", "cg005", "cg006", "cg007", "cg008", "cg009", + "cg010", "cg011", "cg012", "cg013", "cg014", "cg015", "cg016", "cg017", "cg018", + "cg019", "cg020", ] + let sig_probes : Array[String] = ["cg001", "cg002", "cg003", "cg004", "cg005"] let results = @src.mm_go_enrichment( sig_probes, all_probes, annotations, go_db, ) @@ -387,14 +384,11 @@ test "mm_go_enrichment_finds_enriched" { let annotations = @src.mm_sample_annotations() let go_db = @src.mm_sample_go_database() let all_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", - "cg006", "cg007", "cg008", "cg009", "cg010", - "cg011", "cg012", "cg013", "cg014", "cg015", - "cg016", "cg017", "cg018", "cg019", "cg020", - ] - let sig_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", + "cg001", "cg002", "cg003", "cg004", "cg005", "cg006", "cg007", "cg008", "cg009", + "cg010", "cg011", "cg012", "cg013", "cg014", "cg015", "cg016", "cg017", "cg018", + "cg019", "cg020", ] + let sig_probes : Array[String] = ["cg001", "cg002", "cg003", "cg004", "cg005"] let results = @src.mm_go_enrichment( sig_probes, all_probes, annotations, go_db, ) @@ -420,14 +414,11 @@ test "mm_probe_bias_correction_sample" { let annotations = @src.mm_sample_annotations() let go_db = @src.mm_sample_go_database() let all_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", - "cg006", "cg007", "cg008", "cg009", "cg010", - "cg011", "cg012", "cg013", "cg014", "cg015", - "cg016", "cg017", "cg018", "cg019", "cg020", - ] - let sig_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", + "cg001", "cg002", "cg003", "cg004", "cg005", "cg006", "cg007", "cg008", "cg009", + "cg010", "cg011", "cg012", "cg013", "cg014", "cg015", "cg016", "cg017", "cg018", + "cg019", "cg020", ] + let sig_probes : Array[String] = ["cg001", "cg002", "cg003", "cg004", "cg005"] let results = @src.mm_probe_bias_correction( sig_probes, all_probes, annotations, go_db, ) @@ -460,14 +451,11 @@ test "mm_go_summary_nonempty" { let annotations = @src.mm_sample_annotations() let go_db = @src.mm_sample_go_database() let all_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", - "cg006", "cg007", "cg008", "cg009", "cg010", - "cg011", "cg012", "cg013", "cg014", "cg015", - "cg016", "cg017", "cg018", "cg019", "cg020", - ] - let sig_probes : Array[String] = [ - "cg001", "cg002", "cg003", "cg004", "cg005", + "cg001", "cg002", "cg003", "cg004", "cg005", "cg006", "cg007", "cg008", "cg009", + "cg010", "cg011", "cg012", "cg013", "cg014", "cg015", "cg016", "cg017", "cg018", + "cg019", "cg020", ] + let sig_probes : Array[String] = ["cg001", "cg002", "cg003", "cg004", "cg005"] let results = @src.mm_go_enrichment( sig_probes, all_probes, annotations, go_db, ) @@ -573,9 +561,7 @@ test "mm_edge_all_same_group" { let beta = @src.mm_sample_beta() let annotations = @src.mm_sample_annotations() // All samples in treatment group (no control). - let group = [ - true, true, true, true, true, true, true, true, - ] + let group = [true, true, true, true, true, true, true, true] let results = @src.mm_dmp_analysis(beta, group, annotations) assert_eq(results.length(), 20) // With no control group, all p-values should be 1.0. diff --git a/test/moonbit/mix_omics_test.mbt b/test/moonbit/mix_omics_test.mbt index 0c8c911f..53a4dce9 100644 --- a/test/moonbit/mix_omics_test.mbt +++ b/test/moonbit/mix_omics_test.mbt @@ -54,7 +54,14 @@ test "diablo_new_defaults" { ///| test "diablo_options_full_custom" { let design = [[0.0, 1.0], [1.0, 0.0]] - let opts = @src.diablo_options_full(3, [[2, 3], [1]], design, false, 200, 1.0e-7) + let opts = @src.diablo_options_full( + 3, + [[2, 3], [1]], + design, + false, + 200, + 1.0e-7, + ) assert_eq(opts.ncomp, 3) assert_eq(opts.keep_variables.length(), 2) assert_eq(opts.keep_variables[0].length(), 2) @@ -106,12 +113,7 @@ test "pls_component_extraction" { ///| test "pls_single_component" { - let x = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - ] + let x = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0]] let y = [[1.0], [2.0], [3.0], [4.0]] let opts = @src.pls_options_full(1, "regression", true, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -140,16 +142,12 @@ test "pls_nipals_convergence" { ///| test "pls_data_scaling" { - let x = [ - [1.0, 100.0], - [2.0, 200.0], - [3.0, 300.0], - [4.0, 400.0], - [5.0, 500.0], - ] + let x = [[1.0, 100.0], [2.0, 200.0], [3.0, 300.0], [4.0, 400.0], [5.0, 500.0]] let y = [[1.0], [2.0], [3.0], [4.0], [5.0]] let opts_scaled = @src.pls_options_full(1, "regression", true, 500, 0.000001) - let opts_unscaled = @src.pls_options_full(1, "regression", false, 500, 0.000001) + let opts_unscaled = @src.pls_options_full( + 1, "regression", false, 500, 0.000001, + ) let result_scaled = @src.run_pls(x, y, opts_scaled) let result_unscaled = @src.run_pls(x, y, opts_unscaled) assert_true(result_scaled.ncomp >= 1) @@ -183,11 +181,7 @@ test "pls_ncomp_exceeds_samples" { ///| test "pls_ncomp_exceeds_variables" { - let x = [ - [1.0, 2.0, 3.0], - [2.0, 3.0, 4.0], - [3.0, 4.0, 5.0], - ] + let x = [[1.0, 2.0, 3.0], [2.0, 3.0, 4.0], [3.0, 4.0, 5.0]] let y = [[1.0], [2.0], [3.0]] let opts = @src.pls_options_full(10, "regression", true, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -219,13 +213,7 @@ test "pls_single_sample" { ///| test "pls_multivariate_y" { - let x = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - [5.0, 6.0], - ] + let x = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0], [5.0, 6.0]] let y = [[1.0, 0.5], [2.0, 1.0], [3.0, 1.5], [4.0, 2.0], [5.0, 2.5]] let opts = @src.pls_new() let result = @src.run_pls(x, y, opts) @@ -235,13 +223,7 @@ test "pls_multivariate_y" { ///| test "pls_canonical_mode" { - let x = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - [5.0, 6.0], - ] + let x = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0], [5.0, 6.0]] let y = [[1.0], [2.0], [3.0], [4.0], [5.0]] let opts = @src.pls_options_full(2, "canonical", true, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -339,7 +321,7 @@ test "soft_thresholding_operator" { let result1 = @src.soft_threshold(3.0, 1.0) assert_true((result1 - 2.0).abs() < 1.0e-15) let result2 = @src.soft_threshold(-3.0, 1.0) - assert_true((result2 - (-2.0)).abs() < 1.0e-15) + assert_true((result2 - -2.0).abs() < 1.0e-15) let result3 = @src.soft_threshold(0.5, 1.0) assert_true(result3.abs() < 1.0e-15) let result4 = @src.soft_threshold(0.0, 0.0) @@ -355,13 +337,7 @@ test "diablo_two_blocks" { [4.0, 5.0, 6.0], [5.0, 6.0, 7.0], ] - let block2 = [ - [1.0, 0.5], - [2.0, 1.0], - [3.0, 1.5], - [4.0, 2.0], - [5.0, 2.5], - ] + let block2 = [[1.0, 0.5], [2.0, 1.0], [3.0, 1.5], [4.0, 2.0], [5.0, 2.5]] let blocks = [block1, block2] let names = ["block1", "block2"] let design = [[0.0, 1.0], [1.0, 0.0]] @@ -390,25 +366,22 @@ test "diablo_variable_selection" { let blocks = [block1, block2] let names = ["block1", "block2"] let design = [[0.0, 1.0], [1.0, 0.0]] - let opts = @src.diablo_options_full(1, [[2], [1]], design, true, 500, 0.000001) + let opts = @src.diablo_options_full( + 1, + [[2], [1]], + design, + true, + 500, + 0.000001, + ) let result = @src.run_diablo(blocks, names, opts) assert_true(result.ncomp >= 1) } ///| test "diablo_default_behavior" { - let block1 = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - ] - let block2 = [ - [1.0, 0.5], - [2.0, 1.0], - [3.0, 1.5], - [4.0, 2.0], - ] + let block1 = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0]] + let block2 = [[1.0, 0.5], [2.0, 1.0], [3.0, 1.5], [4.0, 2.0]] let blocks = [block1, block2] let names = ["block1", "block2"] let opts = @src.diablo_new() @@ -419,11 +392,7 @@ test "diablo_default_behavior" { ///| test "diablo_single_block" { - let block1 = [ - [1.0, 2.0, 3.0], - [2.0, 3.0, 4.0], - [3.0, 4.0, 5.0], - ] + let block1 = [[1.0, 2.0, 3.0], [2.0, 3.0, 4.0], [3.0, 4.0, 5.0]] let blocks = [block1] let names = ["block1"] let design = [[0.0]] @@ -528,13 +497,7 @@ test "diablo_result_structure" { ///| test "pls_mode_regression" { - let x = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - [5.0, 6.0], - ] + let x = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0], [5.0, 6.0]] let y = [[1.0], [2.0], [3.0], [4.0], [5.0]] let opts = @src.pls_options_full(2, "regression", true, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -544,13 +507,7 @@ test "pls_mode_regression" { ///| test "pls_mode_invariant_scores" { - let x = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - [5.0, 6.0], - ] + let x = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0], [5.0, 6.0]] let y = [[1.0], [2.0], [3.0], [4.0], [5.0]] let opts = @src.pls_options_full(1, "regression", false, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -574,12 +531,7 @@ test "diablo_block_names" { ///| test "pls_result_has_explained_variance" { - let x = [ - [1.0, 2.0, 3.0], - [2.0, 3.0, 4.0], - [3.0, 4.0, 5.0], - [4.0, 5.0, 6.0], - ] + let x = [[1.0, 2.0, 3.0], [2.0, 3.0, 4.0], [3.0, 4.0, 5.0], [4.0, 5.0, 6.0]] let y = [[1.0], [2.0], [3.0], [4.0]] let opts = @src.pls_options_full(2, "regression", true, 500, 0.000001) let result = @src.run_pls(x, y, opts) @@ -601,4 +553,4 @@ test "spls_selected_variables" { assert_true(result.ncomp >= 1) assert_true(result.x_selected.length() >= 1) assert_true(result.y_selected.length() >= 1) -} \ No newline at end of file +} diff --git a/test/moonbit/mmcifio_test.mbt b/test/moonbit/mmcifio_test.mbt index 00ae62a9..5cb26e47 100644 --- a/test/moonbit/mmcifio_test.mbt +++ b/test/moonbit/mmcifio_test.mbt @@ -273,36 +273,26 @@ test "mmcifio_write_single_residue" { ///| test "mmcifio_write_single_chain" { - let r1 = @src.Residue::new( - resname="ALA", - chainid='A', - resseq=1, - atoms=[ - @src.Atom::new( - name="N", - coord=@src.Vector3::new(0.0, 0.0, 0.0), - resname="ALA", - chainid='A', - resseq=1, - element="N", - ), - ], - ) - let r2 = @src.Residue::new( - resname="GLY", - chainid='A', - resseq=2, - atoms=[ - @src.Atom::new( - name="N", - coord=@src.Vector3::new(3.0, 0.0, 0.0), - resname="GLY", - chainid='A', - resseq=2, - element="N", - ), - ], - ) + let r1 = @src.Residue::new(resname="ALA", chainid='A', resseq=1, atoms=[ + @src.Atom::new( + name="N", + coord=@src.Vector3::new(0.0, 0.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + element="N", + ), + ]) + let r2 = @src.Residue::new(resname="GLY", chainid='A', resseq=2, atoms=[ + @src.Atom::new( + name="N", + coord=@src.Vector3::new(3.0, 0.0, 0.0), + resname="GLY", + chainid='A', + resseq=2, + element="N", + ), + ]) let chain = @src.Chain::new(id='A', residues=[r1, r2]) let model = @src.Model::new(id=1, chains=[chain]) let structure = @src.Structure::new(id="1CG", models=[model]) @@ -313,36 +303,26 @@ test "mmcifio_write_single_chain" { ///| test "mmcifio_write_multiple_chains" { - let r1 = @src.Residue::new( - resname="ALA", - chainid='A', - resseq=1, - atoms=[ - @src.Atom::new( - name="N", - coord=@src.Vector3::new(0.0, 0.0, 0.0), - resname="ALA", - chainid='A', - resseq=1, - element="N", - ), - ], - ) - let r2 = @src.Residue::new( - resname="GLY", - chainid='B', - resseq=1, - atoms=[ - @src.Atom::new( - name="N", - coord=@src.Vector3::new(10.0, 0.0, 0.0), - resname="GLY", - chainid='B', - resseq=1, - element="N", - ), - ], - ) + let r1 = @src.Residue::new(resname="ALA", chainid='A', resseq=1, atoms=[ + @src.Atom::new( + name="N", + coord=@src.Vector3::new(0.0, 0.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + element="N", + ), + ]) + let r2 = @src.Residue::new(resname="GLY", chainid='B', resseq=1, atoms=[ + @src.Atom::new( + name="N", + coord=@src.Vector3::new(10.0, 0.0, 0.0), + resname="GLY", + chainid='B', + resseq=1, + element="N", + ), + ]) let chain_a = @src.Chain::new(id='A', residues=[r1]) let chain_b = @src.Chain::new(id='B', residues=[r2]) let model = @src.Model::new(id=1, chains=[chain_a, chain_b]) @@ -563,26 +543,12 @@ test "mmcifio_atom_site_has_all_columns" { let site = @src.write_mmcif_atom_site(structure) // Verify all 20 column headers are present let columns = [ - "_atom_site.group_PDB", - "_atom_site.id", - "_atom_site.type_symbol", - "_atom_site.label_atom_id", - "_atom_site.label_alt_id", - "_atom_site.label_comp_id", - "_atom_site.label_asym_id", - "_atom_site.label_entity_id", - "_atom_site.label_seq_id", - "_atom_site.pdbx_PDB_ins_code", - "_atom_site.Cartn_x", - "_atom_site.Cartn_y", - "_atom_site.Cartn_z", - "_atom_site.occupancy", - "_atom_site.B_iso_or_equiv", - "_atom_site.pdbx_formal_charge", - "_atom_site.auth_seq_id", - "_atom_site.auth_comp_id", - "_atom_site.auth_asym_id", - "_atom_site.auth_atom_id", + "_atom_site.group_PDB", "_atom_site.id", "_atom_site.type_symbol", "_atom_site.label_atom_id", + "_atom_site.label_alt_id", "_atom_site.label_comp_id", "_atom_site.label_asym_id", + "_atom_site.label_entity_id", "_atom_site.label_seq_id", "_atom_site.pdbx_PDB_ins_code", + "_atom_site.Cartn_x", "_atom_site.Cartn_y", "_atom_site.Cartn_z", "_atom_site.occupancy", + "_atom_site.B_iso_or_equiv", "_atom_site.pdbx_formal_charge", "_atom_site.auth_seq_id", + "_atom_site.auth_comp_id", "_atom_site.auth_asym_id", "_atom_site.auth_atom_id", ] for col in columns { assert_eq(site.contains(col), true) diff --git a/test/moonbit/mmtf_test.mbt b/test/moonbit/mmtf_test.mbt index 6656fcf8..c7d302bc 100644 --- a/test/moonbit/mmtf_test.mbt +++ b/test/moonbit/mmtf_test.mbt @@ -72,15 +72,8 @@ test "mmtf_sec_struct_from_int_undefined" { test "mmtf_sec_struct_to_string_round_trip" { let codes = [0, 1, 2, 3, 4, 5, 6, 7, 99] let names = [ - "pi_helix", - "bend", - "alpha_helix", - "extended", - "310_helix", - "bridge", - "turn", - "coil", - "undefined", + "pi_helix", "bend", "alpha_helix", "extended", "310_helix", "bridge", "turn", + "coil", "undefined", ] for i in 0.. 0) } +///| test "mofa_convergence" { let views = @src.mofa_create_example(25, 2) - let params = @src.MofaParams::create( - n_factors = 2, - max_iterations = 50, - ) + let params = @src.MofaParams::create(n_factors=2, max_iterations=50) let result = @src.mofa_run(views, params) assert_true(result.n_iterations <= 50) } +///| test "mofa_variance_explained" { let views = @src.mofa_create_example(30, 4) - let params = @src.MofaParams::create( - n_factors = 4, - max_iterations = 30, - ) + let params = @src.MofaParams::create(n_factors=4, max_iterations=30) let result = @src.mofa_run(views, params) // Variance should be non-negative for v in result.factor_vars { @@ -77,17 +78,16 @@ test "mofa_variance_explained" { } } +///| test "mofa_active_factors" { let views = @src.mofa_create_example(20, 2) - let params = @src.MofaParams::create( - n_factors = 5, - max_iterations = 30, - ) + let params = @src.MofaParams::create(n_factors=5, max_iterations=30) let result = @src.mofa_run(views, params) assert_true(result.active_factors >= 0) assert_true(result.active_factors <= 5) } +///| test "mofa_single_view" { // Should work with a single view let data : Array[Array[Double]] = Array::new() @@ -105,20 +105,16 @@ test "mofa_single_view" { let view = @src.MofaView::new("single", data, [], []) let views : Array[@src.MofaView] = Array::new() views.push(view) - let params = @src.MofaParams::create( - n_factors = 2, - max_iterations = 15, - ) + let params = @src.MofaParams::create(n_factors=2, max_iterations=15) let result = @src.mofa_run(views, params) assert_true(result.factors.length() == 15) assert_true(result.loadings.length() == 1) } +///| test "mofa_empty_views" { let views : Array[@src.MofaView] = Array::new() - let params = @src.MofaParams::create( - n_factors = 2, - ) + let params = @src.MofaParams::create(n_factors=2) let result = @src.mofa_run(views, params) assert_true(result.factors.length() == 0) assert_true(result.converged == false) diff --git a/test/moonbit/monocle3_test.mbt b/test/moonbit/monocle3_test.mbt index a57de638..bc719667 100644 --- a/test/moonbit/monocle3_test.mbt +++ b/test/moonbit/monocle3_test.mbt @@ -1,78 +1,86 @@ +///| test "monocle3_new_cell_data_set" { let counts = [[1, 2], [3, 4]] let gene_names = ["Gene_A", "Gene_B"] let cell_names = ["Cell_1", "Cell_2"] - + let cds = @src.new_cell_data_set(counts, gene_names, cell_names) - + assert_eq(cds.counts.length(), 2) assert_eq(cds.gene_names.length(), 2) assert_eq(cds.cell_names.length(), 2) assert_eq(cds.counts[0][0], 1) } +///| test "monocle3_new_cell_data_set_empty" { let cds = @src.new_cell_data_set([], [], []) - + assert_eq(cds.counts.length(), 0) assert_eq(cds.normalized.length(), 0) } +///| test "monocle3_preprocess_cds" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) - + let preprocessed = @src.preprocess_cds(cds, 10) - + assert_eq(preprocessed.counts.length(), 50) assert_eq(preprocessed.normalized.length(), 50) assert_eq(preprocessed.reduced_dimensions.length(), 1) } +///| test "monocle3_reduce_dimension_pca" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) - + let reduced = @src.reduce_dimension(cds, "PCA", 5) - + assert_eq(reduced.reduced_dimensions.length(), 1) assert_eq(reduced.reduced_dimensions[0].0, "PCA") } +///| test "monocle3_reduce_dimension_umap" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) - + let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) - + assert_true(reduced.reduced_dimensions.length() >= 2) } +///| test "monocle3_learn_graph" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) - + let with_graph = @src.learn_graph(reduced) - + assert_eq(with_graph.cell_partitions.length(), 50) assert_eq(with_graph.principal_graph.length(), 1) } +///| test "monocle3_order_cells" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) let with_graph = @src.learn_graph(reduced) - + let ordered = @src.order_cells(with_graph, [0]) - + assert_eq(ordered.pseudotime.length(), 50) } +///| test "monocle3_fit_models" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) @@ -80,58 +88,66 @@ test "monocle3_fit_models" { let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) let with_graph = @src.learn_graph(reduced) let ordered = @src.order_cells(with_graph, [0]) - + let result = @src.fit_models(ordered, ["Gene_0", "Gene_1", "Gene_2"]) - + assert_eq(result.gene.length(), 3) assert_eq(result.q_value.length(), 3) } +///| test "monocle3_create_example_data" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() - + assert_eq(counts.length(), 50) assert_eq(gene_names.length(), 100) assert_eq(cell_names.length(), 50) } +///| test "monocle3_full_pipeline" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() - + let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) let with_graph = @src.learn_graph(reduced) let ordered = @src.order_cells(with_graph, [0]) let result = @src.fit_models(ordered, ["Gene_0"]) - + assert_eq(cds.counts.length(), 50) assert_eq(result.gene.length(), 1) } +///| test "monocle3_find_branch_points" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) let with_graph = @src.learn_graph(reduced) - + let branch_points = @src.find_branch_points(with_graph) - + assert_eq(branch_points.length(), 0) } +///| test "monocle3_differential_gene_test_branches" { let (counts, gene_names, cell_names) = @src.create_example_monocle_data() let cds = @src.new_cell_data_set(counts, gene_names, cell_names) let preprocessed = @src.preprocess_cds(cds, 10) let reduced = @src.reduce_dimension(preprocessed, "UMAP", 2) let with_graph = @src.learn_graph(reduced) - + let branch_points = @src.find_branch_points(with_graph) - + if branch_points.length() > 0 { - let result = @src.differential_gene_test_branches(with_graph, branch_points[0], ["Gene_0", "Gene_1"]) + let result = @src.differential_gene_test_branches( + with_graph, + branch_points[0], + ["Gene_0", "Gene_1"], + ) assert_eq(result.gene.length(), 2) } -} \ No newline at end of file +} diff --git a/test/moonbit/moon.pkg b/test/moonbit/moon.pkg index 67883269..8c06828c 100644 --- a/test/moonbit/moon.pkg +++ b/test/moonbit/moon.pkg @@ -1,6 +1,7 @@ import { "IvanAXu/BioSeqs/src" @bio, - "IvanAXu/BioSeqs/src" @src, + "IvanAXu/BioSeqs/src", "moonbitlang/core/hashmap", "moonbitlang/core/double", -} \ No newline at end of file + "moonbitlang/core/math", +} diff --git a/test/moonbit/motif_scan_test.mbt b/test/moonbit/motif_scan_test.mbt index c72e3dce..bc20a39e 100644 --- a/test/moonbit/motif_scan_test.mbt +++ b/test/moonbit/motif_scan_test.mbt @@ -49,31 +49,24 @@ test "pwm_construction_default_background_and_strand" { ///| test "pwm_from_counts_basic" { // 4 rows (A,C,G,T) x 3 columns - let counts = [ - [9, 1, 1], - [1, 8, 1], - [1, 1, 7], - [1, 2, 3], - ] + let counts = [[9, 1, 1], [1, 8, 1], [1, 1, 7], [1, 2, 3]] let pwm = @src.motif_scan_pwm_from_counts("MC1", "CountMotif", counts, 1.0) assert_eq(pwm.motif_id, "MC1") assert_eq(pwm.motif_name, "CountMotif") assert_eq(pwm.matrix.length(), 4) assert_eq(pwm.matrix[0].length(), 3) // Each column should sum to ~1.0 (with pseudocount). - let col0 = pwm.matrix[0][0] + pwm.matrix[1][0] + pwm.matrix[2][0] + pwm.matrix[3][0] + let col0 = pwm.matrix[0][0] + + pwm.matrix[1][0] + + pwm.matrix[2][0] + + pwm.matrix[3][0] assert_true((col0 - 1.0).abs() < 0.001) } ///| test "pwm_from_counts_normalization" { // Counts: column 0 has total 10 + 4*pseudocount; each entry = (count+pseudo)/norm - let counts = [ - [10, 0], - [0, 10], - [0, 0], - [0, 0], - ] + let counts = [[10, 0], [0, 10], [0, 0], [0, 0]] let pwm = @src.motif_scan_pwm_from_counts("MC2", "Norm", counts, 1.0) // Column 0: total=10, norm=10+4=14. A = (10+1)/14 = 11/14 let expected_a = 11.0 / 14.0 @@ -96,7 +89,7 @@ test "pssm_construction_log_likelihood" { // A: log2(0.9/0.25) = log2(3.6) ~ 1.848 assert_true((pssm.scores[0][0] - 1.848).abs() < 0.01) // C: log2(0.05/0.25) = log2(0.2) ~ -2.322 - assert_true((pssm.scores[1][0] - (-2.322)).abs() < 0.01) + assert_true((pssm.scores[1][0] - -2.322).abs() < 0.01) } ///| @@ -104,7 +97,7 @@ test "pssm_construction_min_max_score" { let pwm = make_one_pos_pwm() let pssm = @src.build_pssm(pwm) // min_score = min of all bases at position 0 = -2.322 - assert_true((pssm.min_score - (-2.322)).abs() < 0.01) + assert_true((pssm.min_score - -2.322).abs() < 0.01) // max_score = max of all bases at position 0 = 1.848 assert_true((pssm.max_score - 1.848).abs() < 0.01) } @@ -114,10 +107,12 @@ test "pssm_carries_metadata" { let pwm = @src.PositionWeightMatrix::new( motif_id="M99", motif_name="Meta", - matrix=[[0.9, 0.05, 0.05, 0.05, 0.9, 0.9], - [0.05, 0.05, 0.05, 0.9, 0.05, 0.05], - [0.05, 0.05, 0.9, 0.05, 0.05, 0.05], - [0.05, 0.9, 0.05, 0.05, 0.05, 0.05]], + matrix=[ + [0.9, 0.05, 0.05, 0.05, 0.9, 0.9], + [0.05, 0.05, 0.05, 0.9, 0.05, 0.05], + [0.05, 0.05, 0.9, 0.05, 0.05, 0.05], + [0.05, 0.9, 0.05, 0.05, 0.05, 0.05], + ], background=[0.25, 0.25, 0.25, 0.25], strand="+-", ) @@ -135,7 +130,7 @@ test "compute_score_single_position" { // Score of "A" = log2(0.9/0.25) ~ 1.848 assert_true((@src.compute_score(pssm, "A") - 1.848).abs() < 0.01) // Score of "C" = log2(0.05/0.25) ~ -2.322 - assert_true((@src.compute_score(pssm, "C") - (-2.322)).abs() < 0.01) + assert_true((@src.compute_score(pssm, "C") - -2.322).abs() < 0.01) } ///| @@ -289,10 +284,12 @@ test "scan_sequence_forward_only_mode" { let pwm = @src.PositionWeightMatrix::new( motif_id="Mfwd", motif_name="FwdOnly", - matrix=[[0.9, 0.05, 0.05, 0.05, 0.9, 0.9], - [0.05, 0.05, 0.05, 0.9, 0.05, 0.05], - [0.05, 0.05, 0.9, 0.05, 0.05, 0.05], - [0.05, 0.9, 0.05, 0.05, 0.05, 0.05]], + matrix=[ + [0.9, 0.05, 0.05, 0.05, 0.9, 0.9], + [0.05, 0.05, 0.05, 0.9, 0.05, 0.05], + [0.05, 0.05, 0.9, 0.05, 0.05, 0.05], + [0.05, 0.9, 0.05, 0.05, 0.05, 0.05], + ], background=[0.25, 0.25, 0.25, 0.25], strand="+", ) @@ -523,11 +520,7 @@ test "pwm_from_counts_empty" { ///| test "compute_score_empty_pssm" { - let pwm = @src.PositionWeightMatrix::new( - motif_id="E", - motif_name="E", - matrix=[], - ) + let pwm = @src.PositionWeightMatrix::new(motif_id="E", motif_name="E", matrix=[]) let pssm = @src.build_pssm(pwm) assert_true((@src.compute_score(pssm, "ACGT") - 0.0).abs() < 0.001) assert_true((@src.compute_pvalue(pssm, 0.0) - 1.0).abs() < 0.001) diff --git a/test/moonbit/motifs_advanced_test.mbt b/test/moonbit/motifs_advanced_test.mbt index ad15dd73..502664e5 100644 --- a/test/moonbit/motifs_advanced_test.mbt +++ b/test/moonbit/motifs_advanced_test.mbt @@ -7,10 +7,7 @@ ///| test "jaspar_motif_creation" { - let pwm : Array[Array[Double]] = [ - [0.1, 0.2, 0.3, 0.4], - [0.4, 0.3, 0.2, 0.1], - ] + let pwm : Array[Array[Double]] = [[0.1, 0.2, 0.3, 0.4], [0.4, 0.3, 0.2, 0.1]] let motif = @src.JasparMotif::new( "MA0001.1", "AGL3", @@ -31,8 +28,7 @@ test "jaspar_motif_creation" { ///| test "jaspar_parse_simple_pfm" { - let pfm_content = - ">MA0001.1 AGL3\n" + + let pfm_content = ">MA0001.1 AGL3\n" + "A [ 0 3 79 40 66 48 65 11 65 0 ]\n" + "C [ 94 75 4 3 1 2 5 2 7 0 ]\n" + "G [ 1 0 3 4 1 0 5 3 4 0 ]\n" + @@ -46,8 +42,7 @@ test "jaspar_parse_simple_pfm" { ///| test "jaspar_parse_multiple_motifs" { - let pfm_content = - ">MA0001.1 AGL3\n" + + let pfm_content = ">MA0001.1 AGL3\n" + "A [ 10 20 30 40 ]\n" + "C [ 40 30 20 10 ]\n" + "G [ 20 20 20 20 ]\n" + @@ -82,10 +77,7 @@ test "jaspar_pfm_to_pwm" { ///| test "jaspar_pwm_to_jaspar_format" { - let pwm : Array[Array[Double]] = [ - [0.1, 0.2, 0.3, 0.4], - [0.4, 0.3, 0.2, 0.1], - ] + let pwm : Array[Array[Double]] = [[0.1, 0.2, 0.3, 0.4], [0.4, 0.3, 0.2, 0.1]] let jaspar_str = @src.pwm_to_jaspar(pwm, "MA0001.1", "TestMotif") assert_true(jaspar_str.has_prefix(">MA0001.1 TestMotif")) assert_true(jaspar_str.contains("A [")) @@ -98,10 +90,7 @@ test "jaspar_pwm_to_jaspar_format" { ///| test "transfac_motif_creation" { - let pwm : Array[Array[Double]] = [ - [1.0, 2.0, 3.0, 4.0], - [4.0, 3.0, 2.0, 1.0], - ] + let pwm : Array[Array[Double]] = [[1.0, 2.0, 3.0, 4.0], [4.0, 3.0, 2.0, 1.0]] let motif = @src.TransfacMotif::new( "T00001", "M00001", @@ -119,8 +108,7 @@ test "transfac_motif_creation" { ///| test "transfac_parse_simple" { - let transfac_content = - "AC T00001\n" + + let transfac_content = "AC T00001\n" + "ID M00001\n" + "NA AP-1\n" + "DE Activator protein 1\n" + @@ -138,10 +126,7 @@ test "transfac_parse_simple" { ///| test "transfac_pwm_to_transfac_format" { - let pwm : Array[Array[Double]] = [ - [1.0, 2.0, 3.0, 4.0], - [4.0, 3.0, 2.0, 1.0], - ] + let pwm : Array[Array[Double]] = [[1.0, 2.0, 3.0, 4.0], [4.0, 3.0, 2.0, 1.0]] let transfac_str = @src.pwm_to_transfac(pwm, "TestMotif", "T00001") assert_true(transfac_str.contains("AC T00001")) assert_true(transfac_str.contains("NA TestMotif")) @@ -184,10 +169,7 @@ test "optimal_motif_alignment_different_lengths" { [0.0, 0.1, 0.1, 0.8], [0.8, 0.1, 0.1, 0.0], ] - let pwm2 : Array[Array[Double]] = [ - [0.0, 0.1, 0.1, 0.8], - [0.8, 0.1, 0.1, 0.0], - ] + let pwm2 : Array[Array[Double]] = [[0.0, 0.1, 0.1, 0.8], [0.8, 0.1, 0.1, 0.0]] let alignment = @src.optimal_motif_alignment(pwm1, pwm2) assert_true(alignment.score >= 0.0) assert_true(alignment.aligned_length > 0) @@ -210,26 +192,16 @@ test "motif_kl_divergence" { ///| test "motif_js_divergence" { - let pwm1 : Array[Array[Double]] = [ - [0.8, 0.1, 0.1, 0.0], - [0.0, 0.1, 0.1, 0.8], - ] - let pwm2 : Array[Array[Double]] = [ - [0.8, 0.1, 0.1, 0.0], - [0.0, 0.1, 0.1, 0.8], - ] + let pwm1 : Array[Array[Double]] = [[0.8, 0.1, 0.1, 0.0], [0.0, 0.1, 0.1, 0.8]] + let pwm2 : Array[Array[Double]] = [[0.8, 0.1, 0.1, 0.0], [0.0, 0.1, 0.1, 0.8]] let js = @src.motif_js_divergence(pwm1, pwm2) assert_eq(js, 0.0) } ///| test "motif_js_divergence_different" { - let pwm1 : Array[Array[Double]] = [ - [0.8, 0.1, 0.1, 0.0], - ] - let pwm2 : Array[Array[Double]] = [ - [0.0, 0.1, 0.1, 0.8], - ] + let pwm1 : Array[Array[Double]] = [[0.8, 0.1, 0.1, 0.0]] + let pwm2 : Array[Array[Double]] = [[0.0, 0.1, 0.1, 0.8]] let js = @src.motif_js_divergence(pwm1, pwm2) assert_true(js >= 0.0) } @@ -250,10 +222,7 @@ test "motif_cluster_creation" { ///| test "cluster_motifs_single" { let names = ["motif1"] - let pwm1 : Array[Array[Double]] = [ - [0.8, 0.1, 0.1, 0.0], - [0.0, 0.1, 0.1, 0.8], - ] + let pwm1 : Array[Array[Double]] = [[0.8, 0.1, 0.1, 0.0], [0.0, 0.1, 0.1, 0.8]] let pwms = [pwm1] let clusters = @src.cluster_motifs(names, pwms, 0.5) assert_eq(clusters.length(), 1) @@ -275,10 +244,7 @@ test "optimal_pseudocount" { ///| test "motif_gc_content" { - let pwm : Array[Array[Double]] = [ - [0.1, 0.4, 0.4, 0.1], - [0.1, 0.4, 0.4, 0.1], - ] + let pwm : Array[Array[Double]] = [[0.1, 0.4, 0.4, 0.1], [0.1, 0.4, 0.4, 0.1]] let gc = @src.motif_gc_content_adv(pwm) assert_true(gc > 0.7 && gc < 0.9) } diff --git a/test/moonbit/ms_core_utils_test.mbt b/test/moonbit/ms_core_utils_test.mbt index e19dbf28..f48f8d3d 100644 --- a/test/moonbit/ms_core_utils_test.mbt +++ b/test/moonbit/ms_core_utils_test.mbt @@ -11,12 +11,14 @@ test "mc_mean_basic" { assert_true((m - 3.0).abs() < 0.001) } +///| test "mc_mean_empty" { let empty : Array[Double] = [] let m = @src.mc_mean(empty) assert_eq(m, 0.0) } +///| test "mc_mean_two_values" { let m = @src.mc_mean([2.5, 3.5]) assert_true((m - 3.0).abs() < 0.001) @@ -26,21 +28,25 @@ test "mc_mean_two_values" { // Numeric helpers: mc_median // ============================================================================ +///| test "mc_median_odd" { let m = @src.mc_median([1.0, 3.0, 5.0]) assert_eq(m, 3.0) } +///| test "mc_median_even" { let m = @src.mc_median([1.0, 2.0, 3.0, 4.0]) assert_true((m - 2.5).abs() < 0.001) } +///| test "mc_median_unsorted" { let m = @src.mc_median([5.0, 1.0, 3.0, 2.0, 4.0]) assert_eq(m, 3.0) } +///| test "mc_median_empty" { let empty : Array[Double] = [] let m = @src.mc_median(empty) @@ -51,17 +57,20 @@ test "mc_median_empty" { // Numeric helpers: mc_sd // ============================================================================ +///| test "mc_sd_basic" { // mean=3, var = (4+1+0+1+4)/4 = 2.5, sd = sqrt(2.5) ~ 1.5811 let s = @src.mc_sd([1.0, 2.0, 3.0, 4.0, 5.0]) assert_true((s - 1.5811).abs() < 0.001) } +///| test "mc_sd_constant" { let s = @src.mc_sd([5.0, 5.0, 5.0]) assert_eq(s, 0.0) } +///| test "mc_sd_single" { // n < 2 returns 0.0 let s = @src.mc_sd([7.0]) @@ -72,11 +81,13 @@ test "mc_sd_single" { // Numeric helpers: mc_quantile // ============================================================================ +///| test "mc_quantile_median" { let q = @src.mc_quantile([1.0, 2.0, 3.0, 4.0, 5.0], 0.5) assert_eq(q, 3.0) } +///| test "mc_quantile_min_max" { let qmin = @src.mc_quantile([1.0, 2.0, 3.0, 4.0, 5.0], 0.0) let qmax = @src.mc_quantile([1.0, 2.0, 3.0, 4.0, 5.0], 1.0) @@ -84,6 +95,7 @@ test "mc_quantile_min_max" { assert_eq(qmax, 5.0) } +///| test "mc_quantile_quarter" { // pos = 0.25 * 4 = 1, lo=1, frac=0 -> sorted[1] = 2.0 let q = @src.mc_quantile([1.0, 2.0, 3.0, 4.0, 5.0], 0.25) @@ -94,17 +106,20 @@ test "mc_quantile_quarter" { // Numeric helpers: mc_sum, mc_dot // ============================================================================ +///| test "mc_sum_basic" { let s = @src.mc_sum([1.0, 2.0, 3.0]) assert_eq(s, 6.0) } +///| test "mc_sum_empty" { let empty : Array[Double] = [] let s = @src.mc_sum(empty) assert_eq(s, 0.0) } +///| test "mc_dot_basic" { let d = @src.mc_dot([1.0, 2.0, 3.0], [4.0, 5.0, 6.0]) // 1*4 + 2*5 + 3*6 = 4 + 10 + 18 = 32 @@ -115,6 +130,7 @@ test "mc_dot_basic" { // refineCentroids: mc_refine_centroids // ============================================================================ +///| test "mc_refine_centroids_symmetric_peak" { // Symmetric peak centered at index 2 with m/z 100.2 let mz = [100.0, 100.1, 100.2, 100.3, 100.4] @@ -128,6 +144,7 @@ test "mc_refine_centroids_symmetric_peak" { assert_eq(out_int[0], 100.0) } +///| test "mc_refine_centroids_asymmetric_peak" { // Asymmetric: more weight on the right side -> refined m/z shifts right let mz = [100.0, 100.1, 100.2, 100.3, 100.4] @@ -137,6 +154,7 @@ test "mc_refine_centroids_asymmetric_peak" { assert_true(out_mz[0] > 100.2) } +///| test "mc_refine_centroids_out_of_range_index" { let mz = [100.0, 100.1, 100.2] let intensity = [10.0, 20.0, 30.0] @@ -150,16 +168,19 @@ test "mc_refine_centroids_out_of_range_index" { // localMaxima: mc_local_maxima // ============================================================================ +///| test "mc_local_maxima_basic" { let idx = @src.mc_local_maxima([1.0, 3.0, 2.0]) assert_eq(idx, [1]) } +///| test "mc_local_maxima_zeros" { let idx = @src.mc_local_maxima([0.0, 0.0, 0.0]) assert_eq(idx.length(), 0) } +///| test "mc_local_maxima_multiple_peaks" { let idx = @src.mc_local_maxima([1.0, 2.0, 1.0, 2.0, 1.0]) assert_eq(idx, [1, 3]) @@ -169,6 +190,7 @@ test "mc_local_maxima_multiple_peaks" { // joinPeaks: mc_join_peaks // ============================================================================ +///| test "mc_join_peaks_absolute_tolerance" { let x = [100.0, 200.0, 300.0] let y = [100.01, 199.99, 300.5] @@ -181,6 +203,7 @@ test "mc_join_peaks_absolute_tolerance" { assert_eq(pairs[1], (1, 1)) } +///| test "mc_join_peaks_ppm_tolerance" { let x = [1000.0] let y = [1000.005] @@ -190,6 +213,7 @@ test "mc_join_peaks_ppm_tolerance" { assert_eq(pairs[0], (0, 0)) } +///| test "mc_join_peaks_no_match" { let x = [100.0, 200.0] let y = [150.0, 250.0] @@ -201,6 +225,7 @@ test "mc_join_peaks_no_match" { // Smoothing: mc_smooth_moving_average // ============================================================================ +///| test "mc_smooth_moving_average_basic" { let out = @src.mc_smooth_moving_average([1.0, 2.0, 3.0, 4.0, 5.0], 1) assert_eq(out.length(), 5) @@ -211,6 +236,7 @@ test "mc_smooth_moving_average_basic" { assert_true((out[4] - 4.5).abs() < 0.001) } +///| test "mc_smooth_moving_average_reduces_noise" { // Noisy signal: moving average should reduce total variation let noisy = [1.0, 5.0, 1.0, 5.0, 1.0, 5.0, 1.0] @@ -230,6 +256,7 @@ test "mc_smooth_moving_average_reduces_noise" { // Smoothing: mc_smooth_savitzky_golay // ============================================================================ +///| test "mc_smooth_savitzky_golay_linear_signal" { // A quadratic/SG fit on perfectly linear data reproduces the value at the // window center. For the interior point with a full symmetric window, the @@ -240,6 +267,7 @@ test "mc_smooth_savitzky_golay_linear_signal" { assert_true((out[2] - lin[2]).abs() < 0.01) } +///| test "mc_smooth_savitzky_golay_reduces_noise" { let noisy = [1.0, 5.0, 1.0, 5.0, 1.0, 5.0, 1.0] let smoothed = @src.mc_smooth_savitzky_golay(noisy, 2) @@ -258,11 +286,13 @@ test "mc_smooth_savitzky_golay_reduces_noise" { // Smoothing: mc_smooth dispatcher // ============================================================================ +///| test "mc_smooth_dispatcher_moving_average" { let out = @src.mc_smooth([1.0, 2.0, 3.0, 4.0, 5.0], 1, "MovingAverage") assert_true((out[2] - 3.0).abs() < 0.001) } +///| test "mc_smooth_dispatcher_savitzky_golay" { let out = @src.mc_smooth([1.0, 2.0, 3.0, 4.0, 5.0], 2, "SavitzkyGolay") // Linear signal -> preserved @@ -273,6 +303,7 @@ test "mc_smooth_dispatcher_savitzky_golay" { // Baseline: mc_baseline_snip // ============================================================================ +///| test "mc_baseline_snip_flat_signal" { // Flat positive signal -> baseline should be <= signal and >= 0 let signal = [10.0, 10.0, 10.0, 10.0, 10.0] @@ -284,6 +315,7 @@ test "mc_baseline_snip_flat_signal" { } } +///| test "mc_baseline_snip_below_peaks" { // Signal with a peak in the middle -> baseline at peak should be below peak let signal = [1.0, 1.0, 100.0, 1.0, 1.0] @@ -295,6 +327,7 @@ test "mc_baseline_snip_below_peaks" { // Calibration: mc_calibrate // ============================================================================ +///| test "mc_calibrate_shift" { let observed = [100.0, 200.0, 300.0] let calibrants = [100.1, 200.1, 300.1] @@ -305,6 +338,7 @@ test "mc_calibrate_shift" { assert_true((out[2] - 300.1).abs() < 0.001) } +///| test "mc_calibrate_linear" { // matched pairs: (100, 100.5), (200, 200.5), (300, 300.5) // linear fit -> cal = 1.0 * obs + 0.5 @@ -317,6 +351,7 @@ test "mc_calibrate_linear" { assert_true((out[2] - 300.5).abs() < 0.01) } +///| test "mc_calibrate_no_match_returns_input" { // No calibrants within tolerance -> shift method leaves values unchanged let observed = [100.0, 200.0] @@ -330,23 +365,23 @@ test "mc_calibrate_no_match_returns_input" { // Imputation: mc_is_missing and mc_impute // ============================================================================ +///| test "mc_is_missing_nan" { let nan = 0.0 / 0.0 assert_true(@src.mc_is_missing(nan)) } +///| test "mc_is_missing_not_nan" { assert_true(!@src.mc_is_missing(5.0)) assert_true(!@src.mc_is_missing(0.0)) assert_true(!@src.mc_is_missing(-1.5)) } +///| test "mc_impute_zero" { let nan = 0.0 / 0.0 - let matrix = [ - [1.0, nan, 3.0], - [nan, 5.0, 6.0], - ] + let matrix = [[1.0, nan, 3.0], [nan, 5.0, 6.0]] let out = @src.mc_impute(matrix, "zero", 0) assert_eq(out[0][1], 0.0) assert_eq(out[1][0], 0.0) @@ -355,6 +390,7 @@ test "mc_impute_zero" { assert_eq(out[1][2], 6.0) } +///| test "mc_impute_half_min" { let nan = 0.0 / 0.0 // Row 0: non-missing values [2.0, 4.0], min = 2.0, half = 1.0 @@ -363,6 +399,7 @@ test "mc_impute_half_min" { assert_true((out[0][1] - 1.0).abs() < 0.001) } +///| test "mc_impute_mean" { let nan = 0.0 / 0.0 // Row 0: non-missing [1.0, 4.0], mean = 2.5 @@ -371,6 +408,7 @@ test "mc_impute_mean" { assert_true((out[0][1] - 2.5).abs() < 0.001) } +///| test "mc_impute_median" { let nan = 0.0 / 0.0 // Row 0: non-missing [1.0, 4.0], median = 2.5 @@ -379,27 +417,21 @@ test "mc_impute_median" { assert_true((out[0][1] - 2.5).abs() < 0.001) } +///| test "mc_impute_knn" { let nan = 0.0 / 0.0 // Row 1 missing at col 1; nearest row (by cols 0 and 2) is row 0. // k=1 -> impute from row 0 col 1 = 2.0 - let matrix = [ - [1.0, 2.0, 3.0], - [1.0, nan, 3.0], - [5.0, 4.0, 3.0], - ] + let matrix = [[1.0, 2.0, 3.0], [1.0, nan, 3.0], [5.0, 4.0, 3.0]] let out = @src.mc_impute(matrix, "knn", 1) assert_true((out[1][1] - 2.0).abs() < 0.001) } +///| test "mc_impute_knn_k2" { let nan = 0.0 / 0.0 // k=2 -> mean of row 0 col 1 (2.0) and row 2 col 1 (4.0) = 3.0 - let matrix = [ - [1.0, 2.0, 3.0], - [1.0, nan, 3.0], - [5.0, 4.0, 3.0], - ] + let matrix = [[1.0, 2.0, 3.0], [1.0, nan, 3.0], [5.0, 4.0, 3.0]] let out = @src.mc_impute(matrix, "knn", 2) assert_true((out[1][1] - 3.0).abs() < 0.001) } @@ -408,18 +440,16 @@ test "mc_impute_knn_k2" { // medianPolish: mc_median_polish // ============================================================================ +///| test "mc_median_polish_basic" { // Additive matrix: x[i][j] = overall + row_eff[i] + col_eff[j] // Expected: overall=4, row_eff=[-2, 2], col_eff=[-1, 1], residuals all 0 - let matrix = [ - [1.0, 3.0], - [5.0, 7.0], - ] + let matrix = [[1.0, 3.0], [5.0, 7.0]] let (r, row_eff, col_eff, overall) = @src.mc_median_polish(matrix, 10, 0.0001) assert_true((overall - 4.0).abs() < 0.001) - assert_true((row_eff[0] - (-2.0)).abs() < 0.001) + assert_true((row_eff[0] - -2.0).abs() < 0.001) assert_true((row_eff[1] - 2.0).abs() < 0.001) - assert_true((col_eff[0] - (-1.0)).abs() < 0.001) + assert_true((col_eff[0] - -1.0).abs() < 0.001) assert_true((col_eff[1] - 1.0).abs() < 0.001) // Residuals should be ~0 for i in 0..<2 { @@ -429,6 +459,7 @@ test "mc_median_polish_basic" { } } +///| test "mc_median_polish_empty" { let empty : Array[Array[Double]] = [] let (r, row_eff, col_eff, overall) = @src.mc_median_polish(empty, 10, 0.0001) @@ -442,21 +473,19 @@ test "mc_median_polish_empty" { // robustSummary: mc_robust_summary // ============================================================================ +///| test "mc_robust_summary_basic" { // Per-column median: // col 0: median([1,3,5]) = 3.0 // col 1: median([10,20,30]) = 20.0 - let matrix = [ - [1.0, 10.0], - [3.0, 20.0], - [5.0, 30.0], - ] + let matrix = [[1.0, 10.0], [3.0, 20.0], [5.0, 30.0]] let out = @src.mc_robust_summary(matrix) assert_eq(out.length(), 2) assert_eq(out[0], 3.0) assert_eq(out[1], 20.0) } +///| test "mc_robust_summary_empty" { let empty : Array[Array[Double]] = [] let out = @src.mc_robust_summary(empty) @@ -467,6 +496,7 @@ test "mc_robust_summary_empty" { // MAD: mc_mad // ============================================================================ +///| test "mc_mad_basic" { // median = 3, deviations = [2,1,0,1,2], median of devs = 1.0 // mad = 1.0 * 1.4826 = 1.4826 @@ -474,6 +504,7 @@ test "mc_mad_basic" { assert_true((m - 1.4826).abs() < 0.001) } +///| test "mc_mad_constant" { let m = @src.mc_mad([7.0, 7.0, 7.0]) assert_eq(m, 0.0) @@ -483,12 +514,14 @@ test "mc_mad_constant" { // Validity: mc_valid_peak_list, mc_has_missing // ============================================================================ +///| test "mc_valid_peak_list_valid" { let mz = [100.0, 200.0, 300.0] let intensity = [1.0, 2.0, 3.0] assert_true(@src.mc_valid_peak_list(mz, intensity)) } +///| test "mc_valid_peak_list_not_increasing" { // m/z 200 == 200 is not strictly increasing let mz = [100.0, 200.0, 200.0] @@ -496,32 +529,30 @@ test "mc_valid_peak_list_not_increasing" { assert_true(!@src.mc_valid_peak_list(mz, intensity)) } +///| test "mc_valid_peak_list_length_mismatch" { let mz = [100.0, 200.0] let intensity = [1.0, 2.0, 3.0] assert_true(!@src.mc_valid_peak_list(mz, intensity)) } +///| test "mc_valid_peak_list_negative_intensity" { let mz = [100.0, 200.0] let intensity = [-1.0, 2.0] assert_true(!@src.mc_valid_peak_list(mz, intensity)) } +///| test "mc_has_missing_clean" { - let matrix = [ - [1.0, 2.0], - [3.0, 4.0], - ] + let matrix = [[1.0, 2.0], [3.0, 4.0]] assert_true(!@src.mc_has_missing(matrix)) } +///| test "mc_has_missing_with_nan" { let nan = 0.0 / 0.0 - let matrix = [ - [1.0, nan], - [3.0, 4.0], - ] + let matrix = [[1.0, nan], [3.0, 4.0]] assert_true(@src.mc_has_missing(matrix)) } @@ -529,12 +560,9 @@ test "mc_has_missing_with_nan" { // Aggregation: mc_aggregate_rows // ============================================================================ +///| test "mc_aggregate_rows_sum" { - let matrix = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - ] + let matrix = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0]] let groups = [0, 0, 1] let out = @src.mc_aggregate_rows(matrix, groups, "sum") assert_eq(out.length(), 2) @@ -544,12 +572,9 @@ test "mc_aggregate_rows_sum" { assert_true((out[1][1] - 30.0).abs() < 0.001) } +///| test "mc_aggregate_rows_mean" { - let matrix = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - ] + let matrix = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0]] let groups = [0, 0, 1] let out = @src.mc_aggregate_rows(matrix, groups, "mean") assert_eq(out.length(), 2) @@ -559,12 +584,9 @@ test "mc_aggregate_rows_mean" { assert_true((out[1][1] - 30.0).abs() < 0.001) } +///| test "mc_aggregate_rows_median" { - let matrix = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - ] + let matrix = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0]] let groups = [0, 0, 1] let out = @src.mc_aggregate_rows(matrix, groups, "median") assert_eq(out.length(), 2) @@ -578,6 +600,7 @@ test "mc_aggregate_rows_median" { // Normalization: mc_normalize_tic // ============================================================================ +///| test "mc_normalize_tic_basic" { // total = 10, out = v / total * 100 let out = @src.mc_normalize_tic([1.0, 2.0, 3.0, 4.0]) @@ -588,6 +611,7 @@ test "mc_normalize_tic_basic" { assert_true((out[3] - 40.0).abs() < 0.001) } +///| test "mc_normalize_tic_zero_total" { // total = 0 -> all zeros let out = @src.mc_normalize_tic([0.0, 0.0, 0.0]) @@ -597,6 +621,7 @@ test "mc_normalize_tic_zero_total" { assert_eq(out[2], 0.0) } +///| test "mc_normalize_tic_sums_to_100" { let out = @src.mc_normalize_tic([10.0, 20.0, 30.0, 40.0]) let total = @src.mc_sum(out) diff --git a/test/moonbit/msf_test.mbt b/test/moonbit/msf_test.mbt new file mode 100644 index 00000000..39f56e9f --- /dev/null +++ b/test/moonbit/msf_test.mbt @@ -0,0 +1,1068 @@ +// Black-box tests for Biopython Bio.Align.msf-compatible support. + +///| +fn msf_test_parse(text : String) -> @src.MsfAlignment { + @src.msf_parse(text) catch { + MsfError(message) => abort("valid MSF input failed: " + message) + } +} + +///| +fn msf_test_parse_unchecked(text : String) -> @src.MsfAlignment { + @src.msf_parse(text, verify_checksums=false) catch { + MsfError(message) => abort("unchecked MSF input failed: " + message) + } +} + +///| +fn msf_test_raises(text : String) -> Bool { + try { + ignore(@src.msf_parse(text)) + false + } catch { + MsfError(_) => true + } +} + +///| +fn msf_test_write_raises( + alignment : @src.MsfAlignment, + block_width : Int, + group_width : Int, + gap_character : String, +) -> Bool { + try { + ignore( + @src.msf_write(alignment, block_width~, group_width~, gap_character~), + ) + false + } catch { + MsfError(_) => true + } +} + +///| +fn msf_test_from_aligned_raises( + ids : Array[String], + rows : Array[String], + weights : Array[Double], +) -> Bool { + try { + ignore(@src.msf_from_aligned(ids, rows, @src.MsfNucleotide, weights~)) + false + } catch { + MsfError(_) => true + } +} + +///| +fn msf_test_sequence_raises( + id : String, + row : String, + sequence_type : @src.MsfSequenceType, + weight : Double, + checksum : Int, +) -> Bool { + try { + ignore( + @src.MsfSequence::create( + id, + row, + sequence_type, + weight~, + checksum=Some(checksum), + ), + ) + false + } catch { + MsfError(_) => true + } +} + +///| +fn msf_test_sample() -> @src.MsfAlignment { + msf_test_parse(@src.msf_example_text()) +} + +///| +fn msf_test_single_text() -> String { + "!!NA_MULTIPLE_ALIGNMENT 1.0\n" + + "\n" + + "Single MSF: 4 Type: N Check: 748 ..\n" + + "\n" + + " Name: alpha Len: 4 Check: 748 Weight: 1.0\n" + + "//\n" + + "\n" + + " alpha ACGT\n" +} + +///| +fn msf_test_protein_text() -> String { + "!!AA_MULTIPLE_ALIGNMENT 1.0\n" + + "\n" + + "Protein fixture MSF: 12 Type: P Check: 8342 ..\n" + + "\n" + + " Name: alpha Len: 12 Check: 5761 Weight: 1.0\n" + + " Name: beta Len: 8 Check: 2581 Weight: 0.5\n" + + "//\n" + + "\n" + + " alpha ACDEFG HIKLMN\n" + + " beta ACDFGH IK\n" +} + +///| +fn msf_test_doa_length_mismatch_text() -> String { + "!!AA_MULTIPLE_ALIGNMENT\n" + + "\n" + + "DOA-style MSF: 2 Type: P Check: 0 ..\n" + + "\n" + + " Name: full Len: 4 Check: 0 Weight: 1.0\n" + + " Name: short Len: 2 Check: 0 Weight: 1.0\n" + + "//\n" + + "\n" + + " full ACDE\n" + + " short AC\n" +} + +///| +fn msf_test_w_protein_text() -> String { + "!!AA_MULTIPLE_ALIGNMENT\n" + + " MSF: 99 Type: P Oct 18, 2017 11:35 Check: 0 ..\n" + + " Name: W*01:01:01:01 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*01:01:01:02 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*01:01:01:03 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*01:01:01:04 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*01:01:01:05 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*01:01:01:06 Len: 99 Check: 7236 Weight: 1.00\n" + + " Name: W*02:01 Len: 93 Check: 9483 Weight: 1.00\n" + + " Name: W*03:01:01:01 Len: 93 Check: 9974 Weight: 1.00\n" + + " Name: W*03:01:01:02 Len: 93 Check: 9974 Weight: 1.00\n" + + " Name: W*04:01 Len: 93 Check: 9169 Weight: 1.00\n" + + " Name: W*05:01 Len: 99 Check: 7331 Weight: 1.00\n" + + "//\n" + + "\n" + + " W*01:01:01:01 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:02 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:03 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:04 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:05 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:06 GLTPFNGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*02:01 GLTPSNGYTA ATWTRTAASS VGMNIPYDGA SYLVRNQELR SWTAADKAAQ\n" + + " W*03:01:01:01 GLTPSSGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*03:01:01:02 GLTPSSGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*04:01 GLTPSNGYTA ATWTRTAASS VGMNIPYDGA SYLVRNQELR SWTAADKAAQ\n" + + " W*05:01 GLTPSSGYTA ATWTRTAVSS VGMNIPYHGA SYLVRNQELR SWTAADKAAQ\n" + + " W*01:01:01:01 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*01:01:01:02 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*01:01:01:03 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*01:01:01:04 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*01:01:01:05 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*01:01:01:06 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" + + " W*02:01 MPWRRNMQSC SKPTCREGGR SGSAKSLRMG RRRCTAQNPK RLT\n" + + " W*03:01:01:01 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK RLT\n" + + " W*03:01:01:02 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK RLT\n" + + " W*04:01 MPWRRNMQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK RLT\n" + + " W*05:01 MPWRRNRQSC SKPTCREGGR SGSAKSLRMG RRGCSAQNPK DSHDPPPHL\n" +} + +///| +test "Bio.Align.msf parses sample metadata" { + let alignment = msf_test_sample() + assert_eq(alignment.metadata.header, "!!NA_MULTIPLE_ALIGNMENT 1.0") + assert_eq(alignment.metadata.title, "MoonBit example") + assert_eq(alignment.metadata.declared_length, 12) + assert_eq(alignment.metadata.sequence_type, @src.MsfNucleotide) + assert_eq(alignment.metadata.checksum_label, "Check:") +} + +///| +test "Bio.Align.msf parses interleaved sample rows" { + let alignment = msf_test_sample() + assert_eq(alignment.num_sequences(), 3) + assert_eq(alignment.alignment_length(), 12) + assert_eq(alignment.sequences[0].aligned_sequence, "ACGTACGTACGT") + assert_eq(alignment.sequences[1].aligned_sequence, "ACG-ACGTAC-T") + assert_eq(alignment.sequences[2].aligned_sequence, "A-GTTCGTACGT") +} + +///| +test "Bio.Align.msf preserves ungapped sequences" { + let alignment = msf_test_sample() + assert_eq(alignment.sequences[0].sequence, "ACGTACGTACGT") + assert_eq(alignment.sequences[1].sequence, "ACGACGTACT") + assert_eq(alignment.sequences[2].sequence, "AGTTCGTACGT") +} + +///| +test "Bio.Align.msf preserves lengths and weights" { + let alignment = msf_test_sample() + assert_eq(alignment.sequences[0].length, 12) + assert_eq(alignment.sequences[1].length, 10) + assert_eq(alignment.sequences[2].length, 11) + assert_eq(alignment.sequences[1].weight, 0.75) + assert_eq(alignment.sequences[2].weight, 1.25) +} + +///| +test "Bio.Align.msf validates sample checksums" { + let alignment = msf_test_sample() + assert_true(alignment.sequences[0].checksum_valid()) + assert_true(alignment.sequences[1].checksum_valid()) + assert_true(alignment.sequences[2].checksum_valid()) + assert_true(alignment.checksums_valid()) + assert_eq(alignment.computed_checksum(), 4573) +} + +///| +test "Bio.Align.msf reports summary" { + assert_eq( + msf_test_sample().summary(), + "MsfAlignment(type=N, sequences=3, columns=12, declared=12, checksum=4573)", + ) +} + +///| +test "Bio.Align.msf finds rows by identifier" { + let alignment = msf_test_sample() + assert_eq(alignment.find_sequence("reference"), Some(0)) + assert_eq(alignment.find_sequence("query_two"), Some(2)) + assert_eq(alignment.find_sequence("missing"), None) +} + +///| +test "Bio.Align.msf returns alignment columns" { + let alignment = msf_test_sample() + assert_eq(alignment.column(0), Some("AAA")) + assert_eq(alignment.column(1), Some("CC-")) + assert_eq(alignment.column(3), Some("T-T")) + assert_eq(alignment.column(-1), None) + assert_eq(alignment.column(12), None) +} + +///| +test "Bio.Align.msf maps sequence positions to columns" { + let alignment = msf_test_sample() + assert_eq(alignment.sequence_position_to_column(1, 0), Some(0)) + assert_eq(alignment.sequence_position_to_column(1, 2), Some(2)) + assert_eq(alignment.sequence_position_to_column(1, 3), Some(4)) + assert_eq(alignment.sequence_position_to_column(1, 9), Some(11)) + assert_eq(alignment.sequence_position_to_column(1, 10), None) +} + +///| +test "Bio.Align.msf maps columns to sequence positions" { + let alignment = msf_test_sample() + assert_eq(alignment.column_to_sequence_position(1, 2), Some(2)) + assert_eq(alignment.column_to_sequence_position(1, 3), None) + assert_eq(alignment.column_to_sequence_position(1, 4), Some(3)) + assert_eq(alignment.column_to_sequence_position(1, 11), Some(9)) +} + +///| +test "Bio.Align.msf maps residues between rows" { + let alignment = msf_test_sample() + assert_eq(alignment.map_position(0, 1, 2), Some(2)) + assert_eq(alignment.map_position(0, 1, 3), None) + assert_eq(alignment.map_position(0, 1, 4), Some(3)) + assert_eq(alignment.map_position(1, 0, 3), Some(4)) + assert_eq(alignment.map_position(0, 3, 0), None) +} + +///| +test "Bio.Align.msf emits aligned coordinate pairs" { + let pairs = msf_test_sample().aligned_pairs(0, 1) catch { + MsfError(message) => abort(message) + } + assert_eq(pairs.length(), 12) + assert_eq(pairs[0], (Some(0), Some(0))) + assert_eq(pairs[3], (Some(3), None)) + assert_eq(pairs[4], (Some(4), Some(3))) + assert_eq(pairs[10], (Some(10), None)) + assert_eq(pairs[11], (Some(11), Some(9))) +} + +///| +test "Bio.Align.msf builds compact coordinate path" { + let path = msf_test_sample().coordinate_path() + assert_eq(path.length(), 3) + assert_eq(path[0], [0, 1, 2, 3, 4, 10, 11, 12]) + assert_eq(path[1], [0, 1, 2, 3, 3, 9, 9, 10]) + assert_eq(path[2], [0, 1, 1, 2, 3, 9, 10, 11]) +} + +///| +test "Bio.Align.msf counts identity and gaps" { + let counts = msf_test_sample().pair_counts(0, 1) catch { + MsfError(message) => abort(message) + } + assert_eq(counts.columns, 12) + assert_eq(counts.aligned, 10) + assert_eq(counts.identities, 10) + assert_eq(counts.mismatches, 0) + assert_eq(counts.gap_columns, 2) + assert_eq(counts.double_gap_columns, 0) + assert_eq(counts.gap_opens, 2) + assert_eq(counts.identity(), 1.0) +} + +///| +test "Bio.Align.msf counts mismatches" { + let counts = msf_test_sample().pair_counts(0, 2) catch { + MsfError(message) => abort(message) + } + assert_eq(counts.aligned, 11) + assert_eq(counts.identities, 10) + assert_eq(counts.mismatches, 1) + assert_eq(counts.gap_columns, 1) + assert_eq(counts.gap_opens, 1) +} + +///| +test "Bio.Align.msf counts double-gap columns" { + let alignment = @src.msf_from_aligned( + ["reference", "first", "second"], + ["ACGT", "A--T", "A--T"], + @src.MsfNucleotide, + ) catch { + MsfError(message) => abort(message) + } + let counts = alignment.pair_counts(1, 2) catch { + MsfError(message) => abort(message) + } + assert_eq(counts.aligned, 2) + assert_eq(counts.double_gap_columns, 2) + assert_eq(counts.gap_columns, 0) +} + +///| +test "Bio.Align.msf calculates majority consensus" { + let alignment = msf_test_sample() + assert_eq( + alignment.consensus() catch { + MsfError(message) => abort(message) + }, + "ACGTACGTACGT", + ) +} + +///| +test "Bio.Align.msf applies consensus threshold" { + let consensus = msf_test_sample().consensus(minimum_fraction=1.0) catch { + MsfError(message) => abort(message) + } + assert_eq(consensus, "ACGTXCGTACGT") +} + +///| +test "Bio.Align.msf calculates column occupancy" { + let occupancy = msf_test_sample().occupancy() + assert_eq(occupancy.length(), 12) + assert_eq(occupancy[0], 1.0) + assert_eq(occupancy[1], 2.0 / 3.0) + assert_eq(occupancy[3], 2.0 / 3.0) + assert_eq(occupancy[4], 1.0) + assert_eq(occupancy[10], 2.0 / 3.0) +} + +///| +test "Bio.Align.msf implements standard GCG checksum" { + assert_eq(@src.msf_gcg_checksum("ACGTACGTACGT"), 5688) + assert_eq(@src.msf_gcg_checksum("ACDEFGHIKLMN"), 5761) + assert_eq(@src.msf_gcg_checksum("ACDFGHIK"), 2581) +} + +///| +test "Bio.Align.msf checksum is case insensitive" { + assert_eq(@src.msf_gcg_checksum("acgt"), 748) + assert_eq(@src.msf_gcg_checksum("AcGt"), 748) +} + +///| +test "Bio.Align.msf checksum cycles position after 57" { + assert_eq( + @src.msf_gcg_checksum( + "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA", + ), + 7510, + ) +} + +///| +test "Bio.Align.msf parses official W protein shape" { + let alignment = msf_test_parse(msf_test_w_protein_text()) + assert_eq(alignment.metadata.sequence_type, @src.MsfProtein) + assert_eq(alignment.num_sequences(), 11) + assert_eq(alignment.alignment_length(), 99) + assert_false(alignment.length_mismatch()) +} + +///| +test "Bio.Align.msf parses official W identifiers" { + let alignment = msf_test_parse(msf_test_w_protein_text()) + assert_eq(alignment.sequences[0].id, "W*01:01:01:01") + assert_eq(alignment.sequences[6].id, "W*02:01") + assert_eq(alignment.sequences[10].id, "W*05:01") +} + +///| +test "Bio.Align.msf parses official W sequence" { + let alignment = msf_test_parse(msf_test_w_protein_text()) + assert_eq( + alignment.sequences[0].sequence, + "GLTPFNGYTAATWTRTAVSSVGMNIPYHGASYLVRNQELRSWTAADKAAQMPWRRNRQSCSKPTCREGGRSGSAKSLRMGRRGCSAQNPKDSHDPPPHL", + ) + assert_eq(alignment.sequences[0].checksum, 7236) +} + +///| +test "Bio.Align.msf pads official W short rows" { + let alignment = msf_test_parse(msf_test_w_protein_text()) + assert_eq(alignment.sequences[6].length, 93) + assert_eq(alignment.sequences[6].aligned_sequence.length(), 99) + assert_true(alignment.sequences[6].aligned_sequence.ends_with("RLT------")) + assert_eq(alignment.sequences[6].checksum, 9483) +} + +///| +test "Bio.Align.msf builds official W coordinate path" { + let path = msf_test_parse(msf_test_w_protein_text()).coordinate_path() + assert_eq(path[0], [0, 93, 99]) + assert_eq(path[6], [0, 93, 93]) + assert_eq(path[10], [0, 93, 99]) +} + +///| +test "Bio.Align.msf preserves DOA-style declared length mismatch" { + let alignment = msf_test_parse(msf_test_doa_length_mismatch_text()) + assert_eq(alignment.metadata.declared_length, 2) + assert_eq(alignment.alignment_length(), 4) + assert_true(alignment.length_mismatch()) +} + +///| +test "Bio.Align.msf pads DOA-style completed rows" { + let alignment = msf_test_parse(msf_test_doa_length_mismatch_text()) + assert_eq(alignment.sequences[1].sequence, "AC") + assert_eq(alignment.sequences[1].aligned_sequence, "AC--") + assert_eq(alignment.coordinate_path()[1], [0, 2, 2]) +} + +///| +test "Bio.Align.msf parses PileUp header" { + let text = "PileUp\n\n" + + "PileUp sample MSF: 4 Type: N Check: 0 ..\n\n" + + " Name: alpha oo Len: 4 Check: 0 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + let alignment = msf_test_parse(text) + assert_eq(alignment.metadata.header, "PileUp") + assert_eq(alignment.metadata.sequence_type, @src.MsfNucleotide) +} + +///| +test "Bio.Align.msf parses EMBOSS CompCheck header" { + let text = "!!NA_MULTIPLE_ALIGNMENT 1.0\n\n" + + "stdout MSF: 4 Type: N 01/08/19 CompCheck: 748 ..\n\n" + + " Name: alpha Len: 4 Check: 748 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + let alignment = msf_test_parse(text) + assert_eq(alignment.metadata.checksum_label, "CompCheck:") + assert_eq(alignment.metadata.date_text, "01/08/19") +} + +///| +test "Bio.Align.msf preserves title date and preamble" { + let text = "PileUp\n" + + "Generated by MoonBit\n" + + "Project alpha MSF: 4 Type: N Jan 2 2026 Check: 0 ..\n\n" + + " Name: alpha Len: 4 Check: 0 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + let alignment = msf_test_parse(text) + assert_eq(alignment.metadata.preamble, ["Generated by MoonBit"]) + assert_eq(alignment.metadata.title, "Project alpha") + assert_eq(alignment.metadata.date_text, "Jan 2 2026") +} + +///| +test "Bio.Align.msf accepts CRLF and lowercase residues" { + let text = "!!NA_MULTIPLE_ALIGNMENT\r\n\r\n" + + "MSF: 4 Type: N Check: 0 ..\r\n\r\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\r\n" + + "//\r\n\r\nalpha acgt\r\n" + let alignment = msf_test_parse(text) + assert_eq(alignment.sequences[0].sequence, "ACGT") +} + +///| +test "Bio.Align.msf normalizes dot tilde and hyphen gaps" { + let text = "!!NA_MULTIPLE_ALIGNMENT\n" + + "MSF: 4 Type: N Check: 0 ..\n" + + "Name: dot Len: 3 Check: 0 Weight: 1.0\n" + + "Name: tilde Len: 3 Check: 0 Weight: 1.0\n" + + "Name: hyphen Len: 3 Check: 0 Weight: 1.0\n" + + "//\n\n" + + "dot A.CG\n" + + "tilde AT~G\n" + + "hyphen A-CG\n" + let alignment = msf_test_parse(text) + assert_eq(alignment.sequences[0].aligned_sequence, "A-CG") + assert_eq(alignment.sequences[1].aligned_sequence, "AT-G") + assert_eq(alignment.sequences[2].aligned_sequence, "A-CG") +} + +///| +test "Bio.Align.msf ignores numeric coordinate lines" { + let text = "!!NA_MULTIPLE_ALIGNMENT\n" + + "MSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n" + + "//\n\n" + + "1 2\nalpha AC\n\n3 4\nalpha GT\n" + assert_eq(msf_test_parse(text).sequences[0].sequence, "ACGT") +} + +///| +test "Bio.Align.msf writes canonical default format" { + let output = @src.msf_write(msf_test_sample()) catch { + MsfError(message) => abort(message) + } + assert_true(output.contains("MSF: 12 Type: N")) + assert_true(output.contains("Check: 4573 ..")) + assert_true(output.contains("Name: reference Len: 12 Check: 5688")) + assert_true(output.contains("query_one ACG.ACGTAC .T")) +} + +///| +test "Bio.Align.msf writes tilde gaps" { + let output = @src.msf_write( + msf_test_sample(), + block_width=12, + group_width=4, + gap_character="~", + ) catch { + MsfError(message) => abort(message) + } + assert_true(output.contains("ACG~ ACGT AC~T")) +} + +///| +test "Bio.Align.msf writes hyphen gaps" { + let output = @src.msf_write( + msf_test_sample(), + block_width=12, + group_width=6, + gap_character="-", + ) catch { + MsfError(message) => abort(message) + } + assert_true(output.contains("ACG-AC GTAC-T")) +} + +///| +test "Bio.Align.msf writes custom interleaved blocks" { + let output = @src.msf_write(msf_test_sample(), block_width=6, group_width=3) catch { + MsfError(message) => abort(message) + } + assert_true(output.contains("reference ACG TAC")) + assert_true(output.contains("reference GTA CGT")) +} + +///| +test "Bio.Align.msf round trips canonical output" { + let original = msf_test_sample() + let output = @src.msf_write(original) catch { + MsfError(message) => abort(message) + } + let parsed = msf_test_parse(output) + assert_eq(parsed.num_sequences(), original.num_sequences()) + assert_eq(parsed.alignment_length(), original.alignment_length()) + for index = 0; index < parsed.num_sequences(); index = index + 1 { + assert_eq( + parsed.sequences[index].aligned_sequence, + original.sequences[index].aligned_sequence, + ) + } +} + +///| +test "Bio.Align.msf round trip preserves weights" { + let output = @src.msf_write(msf_test_sample()) catch { + MsfError(message) => abort(message) + } + let parsed = msf_test_parse(output) + assert_eq(parsed.sequences[0].weight, 1.0) + assert_eq(parsed.sequences[1].weight, 0.75) + assert_eq(parsed.sequences[2].weight, 1.25) +} + +///| +test "Bio.Align.msf round trip recomputes checksums" { + let output = @src.msf_write(msf_test_sample()) catch { + MsfError(message) => abort(message) + } + let parsed = msf_test_parse(output) + assert_eq(parsed.metadata.declared_checksum, 4573) + assert_true(parsed.checksums_valid()) +} + +///| +test "Bio.Align.msf constructs alignment from rows" { + let alignment = @src.msf_from_aligned( + ["alpha", "beta"], + ["ACGT", "A-GT"], + @src.MsfNucleotide, + weights=[1.0, 0.5], + ) catch { + MsfError(message) => abort(message) + } + assert_eq(alignment.num_sequences(), 2) + assert_eq(alignment.sequences[1].sequence, "AGT") + assert_eq(alignment.sequences[1].weight, 0.5) + assert_eq(alignment.metadata.declared_length, 4) +} + +///| +test "Bio.Align.msf sequence type emits canonical code" { + assert_eq(@src.MsfProtein.code(), "P") + assert_eq(@src.MsfNucleotide.code(), "N") +} + +///| +test "Bio.Align.msf constructs normalized sequence" { + let sequence = @src.MsfSequence::create( + "alpha", + "acg.t", + @src.MsfNucleotide, + weight=0.25, + ) catch { + MsfError(message) => abort(message) + } + assert_eq(sequence.aligned_sequence, "ACG-T") + assert_eq(sequence.sequence, "ACGT") + assert_eq(sequence.length, 4) + assert_eq(sequence.checksum, 748) + assert_eq(sequence.weight, 0.25) +} + +///| +test "Bio.Align.msf constructs canonical metadata" { + let metadata = @src.MsfMetadata::create(4, @src.MsfProtein) catch { + MsfError(message) => abort(message) + } + assert_eq(metadata.header, "!!AA_MULTIPLE_ALIGNMENT 1.0") + assert_eq(metadata.declared_length, 4) + assert_eq(metadata.sequence_type, @src.MsfProtein) +} + +///| +test "Bio.Align.msf accepts protein wildcard residues" { + let sequence = @src.MsfSequence::create("protein", "ACDX*?", @src.MsfProtein) catch { + MsfError(message) => abort(message) + } + assert_eq(sequence.sequence, "ACDX*?") +} + +///| +test "Bio.Align.msf accepts nucleotide ambiguity codes" { + let sequence = @src.MsfSequence::create( + "dna", + "ACGTURYSWKMBDHVNX", + @src.MsfNucleotide, + ) catch { + MsfError(message) => abort(message) + } + assert_eq(sequence.length, 17) +} + +///| +test "Bio.Align.msf treats zero checksums as unspecified" { + let alignment = msf_test_parse(msf_test_doa_length_mismatch_text()) + assert_true(alignment.sequences[0].checksum_valid()) + assert_true(alignment.checksums_valid()) +} + +///| +test "Bio.Align.msf rejects empty input" { + assert_true(msf_test_raises("")) +} + +///| +test "Bio.Align.msf rejects leading blank input" { + assert_true(msf_test_raises("\n" + msf_test_single_text())) +} + +///| +test "Bio.Align.msf rejects unknown header" { + assert_true( + msf_test_raises( + "CLUSTAL\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects missing alignment header" { + assert_true(msf_test_raises("!!NA_MULTIPLE_ALIGNMENT\nName: alpha\n")) +} + +///| +test "Bio.Align.msf rejects malformed alignment header" { + assert_true( + msf_test_raises("!!NA_MULTIPLE_ALIGNMENT\nMSF: 4 Kind: N Check: 0 ..\n"), + ) +} + +///| +test "Bio.Align.msf rejects invalid sequence type" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: X Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects header type conflict" { + assert_true( + msf_test_raises( + "!!AA_MULTIPLE_ALIGNMENT\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects zero declared width" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 0 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects integer overflow" { + assert_true(msf_test_raises("PileUp\nMSF: 2147483648 Type: N Check: 0 ..\n")) +} + +///| +test "Bio.Align.msf rejects file checksum outside range" { + assert_true(msf_test_raises("PileUp\nMSF: 4 Type: N Check: 10000 ..\n")) +} + +///| +test "Bio.Align.msf rejects header without names" { + assert_true( + msf_test_raises("PileUp\nMSF: 4 Type: N Check: 0 ..\n//\n\nalpha ACGT\n"), + ) +} + +///| +test "Bio.Align.msf rejects duplicate names" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n" + + "//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects missing header terminator" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects missing blank after terminator" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects empty sequence body" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects unknown body row" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nbeta ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects row without residues" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects excess residues" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 3 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects missing residues" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACG\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects invalid nucleotide residue" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGZ\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects invalid protein residue" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: P Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACD%\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects all-gap alignment column" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 3 Type: N Check: 0 ..\n" + + "Name: alpha Len: 2 Check: 0 Weight: 1.0\n" + + "Name: beta Len: 2 Check: 0 Weight: 1.0\n" + + "//\n\nalpha A.C\nbeta A-C\n", + ), + ) +} + +///| +test "Bio.Align.msf verifies sequence checksum by default" { + let text = "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 1 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + assert_true(msf_test_raises(text)) +} + +///| +test "Bio.Align.msf can disable checksum verification" { + let text = "PileUp\nMSF: 4 Type: N Check: 1 ..\n" + + "Name: alpha Len: 4 Check: 1 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + let alignment = msf_test_parse_unchecked(text) + assert_false(alignment.checksums_valid()) + assert_eq(alignment.sequences[0].sequence, "ACGT") +} + +///| +test "Bio.Align.msf verifies file checksum by default" { + let text = "PileUp\nMSF: 4 Type: N Check: 1 ..\n" + + "Name: alpha Len: 4 Check: 748 Weight: 1.0\n" + + "//\n\nalpha ACGT\n" + assert_true(msf_test_raises(text)) +} + +///| +test "Bio.Align.msf rejects malformed weight" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: heavy\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects negative weight" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Len: 4 Check: 0 Weight: -1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects malformed name descriptor" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha Length: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects unexpected descriptor token" { + assert_true( + msf_test_raises( + "PileUp\nMSF: 4 Type: N Check: 0 ..\n" + + "Name: alpha extra Len: 4 Check: 0 Weight: 1.0\n//\n\nalpha ACGT\n", + ), + ) +} + +///| +test "Bio.Align.msf rejects empty row constructor input" { + assert_true(msf_test_from_aligned_raises([], [], [])) +} + +///| +test "Bio.Align.msf rejects mismatched row constructor input" { + assert_true(msf_test_from_aligned_raises(["alpha"], ["ACGT", "ACGT"], [])) +} + +///| +test "Bio.Align.msf rejects mismatched weights" { + assert_true( + msf_test_from_aligned_raises(["alpha", "beta"], ["ACGT", "ACGT"], [1.0]), + ) +} + +///| +test "Bio.Align.msf rejects unequal row widths" { + assert_true( + msf_test_from_aligned_raises(["alpha", "beta"], ["ACGT", "ACG"], []), + ) +} + +///| +test "Bio.Align.msf rejects duplicate constructor identifiers" { + assert_true( + msf_test_from_aligned_raises(["alpha", "alpha"], ["ACGT", "ACGT"], []), + ) +} + +///| +test "Bio.Align.msf rejects whitespace in sequence identifier" { + assert_true( + msf_test_sequence_raises("bad id", "ACGT", @src.MsfNucleotide, 1.0, 748), + ) +} + +///| +test "Bio.Align.msf rejects empty aligned sequence" { + assert_true(msf_test_sequence_raises("", "", @src.MsfNucleotide, 1.0, 0)) +} + +///| +test "Bio.Align.msf rejects all-gap sequence" { + assert_true( + msf_test_sequence_raises("alpha", "...", @src.MsfNucleotide, 1.0, 0), + ) +} + +///| +test "Bio.Align.msf rejects sequence checksum outside range" { + assert_true( + msf_test_sequence_raises("alpha", "ACGT", @src.MsfNucleotide, 1.0, 10000), + ) +} + +///| +test "Bio.Align.msf rejects negative sequence weight" { + assert_true( + msf_test_sequence_raises("alpha", "ACGT", @src.MsfNucleotide, -1.0, 748), + ) +} + +///| +test "Bio.Align.msf rejects invalid metadata width" { + let failed = try { + ignore(@src.MsfMetadata::create(0, @src.MsfNucleotide)) + false + } catch { + MsfError(_) => true + } + assert_true(failed) +} + +///| +test "Bio.Align.msf rejects invalid metadata checksum label" { + let failed = try { + ignore( + @src.MsfMetadata::create( + 4, + @src.MsfNucleotide, + checksum_label="Checksum:", + ), + ) + false + } catch { + MsfError(_) => true + } + assert_true(failed) +} + +///| +test "Bio.Align.msf rejects invalid writer block width" { + assert_true(msf_test_write_raises(msf_test_sample(), 0, 1, ".")) +} + +///| +test "Bio.Align.msf rejects indivisible writer groups" { + assert_true(msf_test_write_raises(msf_test_sample(), 10, 3, ".")) +} + +///| +test "Bio.Align.msf rejects invalid writer gap character" { + assert_true(msf_test_write_raises(msf_test_sample(), 10, 5, "_")) +} + +///| +test "Bio.Align.msf rejects invalid consensus threshold" { + let failed = try { + ignore(msf_test_sample().consensus(minimum_fraction=1.1)) + false + } catch { + MsfError(_) => true + } + assert_true(failed) +} + +///| +test "Bio.Align.msf rejects invalid pairwise row" { + let failed = try { + ignore(msf_test_sample().pair_counts(0, 3)) + false + } catch { + MsfError(_) => true + } + assert_true(failed) +} diff --git a/test/moonbit/msnbase_test.mbt b/test/moonbit/msnbase_test.mbt index 8d85be66..082403f2 100644 --- a/test/moonbit/msnbase_test.mbt +++ b/test/moonbit/msnbase_test.mbt @@ -18,6 +18,7 @@ test "mbn_mslevel_constructors" { } } +///| test "mbn_polarity_constructors" { let p = @src.polarity_positive() let n = @src.polarity_negative() @@ -25,6 +26,7 @@ test "mbn_polarity_constructors" { assert_eq(n, @src.polarity_negative()) } +///| test "mbn_processing_step_new" { let s = @src.ProcessingStep::new("log_transformed", "2025-01-01") assert_eq(s.description, "log_transformed") @@ -35,47 +37,77 @@ test "mbn_processing_step_new" { // Spectrum tests // ============================================================================ +///| test "mbn_spectrum_new_basic" { let mz = [100.0, 200.0, 300.0, 400.0, 500.0] let int = [10.0, 100.0, 500.0, 200.0, 50.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 60.5) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 60.5, + ) assert_eq(sp.peaks_count(), 5) assert_true((sp.rt() - 60.5).abs() < 0.0001) // TIC = 10+100+500+200+50 = 860 assert_true((sp.tic() - 860.0).abs() < 0.0001) } +///| test "mbn_spectrum_empty" { let sp = @src.Spectrum::empty() assert_eq(sp.peaks_count(), 0) assert_eq(sp.tic(), 0.0) } +///| test "mbn_spectrum_mismatched_lengths" { // Mismatched mz/intensity returns empty - let sp = @src.Spectrum::new([1.0, 2.0], [10.0], @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + [1.0, 2.0], + [10.0], + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) assert_eq(sp.peaks_count(), 0) } +///| test "mbn_spectrum_with_precursor" { let mz = [100.0, 200.0] let int = [50.0, 100.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms2(), @src.polarity_positive(), 120.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms2(), + @src.polarity_positive(), + 120.0, + ) let sp2 = sp.with_precursor(500.25, 2) assert_eq(sp2.precursor_mz, 500.25) assert_eq(sp2.precursor_charge, 2) } +///| test "mbn_spectrum_base_peak" { let mz = [100.0, 200.0, 300.0, 400.0] let int = [10.0, 500.0, 50.0, 200.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let (bpmz, bpint) = sp.base_peak() // Peak at index 1 is highest: mz=200, int=500 assert_true((bpmz - 200.0).abs() < 0.0001) assert_true((bpint - 500.0).abs() < 0.0001) } +///| test "mbn_spectrum_base_peak_empty" { let sp = @src.Spectrum::empty() let (mz, int) = sp.base_peak() @@ -83,54 +115,90 @@ test "mbn_spectrum_base_peak_empty" { assert_eq(int, 0.0) } +///| test "mbn_spectrum_find_peak_exact" { let mz = [100.0, 200.0, 300.0, 400.0, 500.0] let int = [10.0, 100.0, 500.0, 200.0, 50.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let idx = sp.find_peak(300.0, ppm=10.0) assert_eq(idx, 2) } +///| test "mbn_spectrum_find_peak_notfound" { let mz = [100.0, 200.0] let int = [10.0, 100.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let idx = sp.find_peak(999.0, ppm=10.0) assert_eq(idx, -1) } +///| test "mbn_spectrum_filter_mz_range" { let mz = [100.0, 200.0, 300.0, 400.0, 500.0] let int = [10.0, 100.0, 500.0, 200.0, 50.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let filtered = sp.filter_mz_range(150.0, 450.0) assert_eq(filtered.peaks_count(), 3) assert_eq(filtered.mz()[0], 200.0) assert_eq(filtered.mz()[2], 400.0) } +///| test "mbn_spectrum_normalize_tic" { - let int = [10.0, 100.0, 500.0, 200.0, 50.0] // TIC = 860 + let int = [10.0, 100.0, 500.0, 200.0, 50.0] // TIC = 860 let mz = [100.0, 200.0, 300.0, 400.0, 500.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let norm = sp.normalize_tic() // New TIC should be 1.0 assert_true((norm.tic() - 1.0).abs() < 0.0001) // First peak normalized: 10/860 - assert_true((norm.intensity()[0] - 10.0/860.0).abs() < 0.0001) + assert_true((norm.intensity()[0] - 10.0 / 860.0).abs() < 0.0001) } // ============================================================================ // Chromatogram tests // ============================================================================ +///| test "mbn_chromatogram_new" { let rt = [0.0, 30.0, 60.0, 90.0, 120.0] let int = [100.0, 5000.0, 20000.0, 8000.0, 200.0] - let chr = @src.Chromatogram::new(rt, int, mz_target=500.0, ppm_tolerance=10.0, acquisition_mode="MRM") + let chr = @src.Chromatogram::new( + rt, + int, + mz_target=500.0, + ppm_tolerance=10.0, + acquisition_mode="MRM", + ) assert_eq(chr.n_points(), 5) } +///| test "mbn_chromatogram_total_auc" { // Simple triangular: rt=[0, 1, 2], int=[0, 10, 0] // Area = 1.0 * 10 = 10 (two triangles each area 5) @@ -140,6 +208,7 @@ test "mbn_chromatogram_total_auc" { assert_true((chr.total_auc() - 10.0).abs() < 0.001) } +///| test "mbn_chromatogram_apex" { let rt = [0.0, 30.0, 60.0, 90.0] let int = [100.0, 5000.0, 20000.0, 8000.0] @@ -149,6 +218,7 @@ test "mbn_chromatogram_apex" { assert_true((apex_int - 20000.0).abs() < 0.001) } +///| test "mbn_chromatogram_fwhm" { // Gaussian-like: peak at center let rt = [0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] @@ -163,9 +233,15 @@ test "mbn_chromatogram_fwhm" { // MSnFeatureData and MSnSampleData // ============================================================================ +///| test "mbn_featuredata_new" { let fd = @src.MSnFeatureData::new( - "Pep_1_2", ["P12345"], "PEPTIDER", 2, 500.25, 60.5, + "Pep_1_2", + ["P12345"], + "PEPTIDER", + 2, + 500.25, + 60.5, ) assert_eq(fd.feature_name, "Pep_1_2") assert_eq(fd.protein_accessions.length(), 1) @@ -175,10 +251,9 @@ test "mbn_featuredata_new" { assert_true((fd.mz - 500.25).abs() < 0.0001) } +///| test "mbn_sampledata_new" { - let sd = @src.MSnSampleData::new( - "S1_WT_1", "WT", "Mouse_1", 1, "Run_001", - ) + let sd = @src.MSnSampleData::new("S1_WT_1", "WT", "Mouse_1", 1, "Run_001") assert_eq(sd.sample_name, "S1_WT_1") assert_eq(sd.group, "WT") assert_eq(sd.subject, "Mouse_1") @@ -190,6 +265,7 @@ test "mbn_sampledata_new" { // MSnSet tests // ============================================================================ +///| test "mbn_msnset_new_basic" { let exprs = [ [1000.0, 1200.0, 1100.0], @@ -213,11 +289,9 @@ test "mbn_msnset_new_basic" { assert_eq(m.sample_names().length(), 3) } +///| test "mbn_msnset_from_names" { - let exprs = [ - [1.0, 2.0], - [3.0, 4.0], - ] + let exprs = [[1.0, 2.0], [3.0, 4.0]] let m = @src.MSnSet::from_names(exprs, ["F1", "F2"], ["S1", "S2"]) assert_eq(m.n_features(), 2) assert_eq(m.n_samples(), 2) @@ -225,11 +299,9 @@ test "mbn_msnset_from_names" { assert_eq(m.sample_names()[1], "S2") } +///| test "mbn_msnset_get_feature" { - let exprs = [ - [10.0, 20.0], - [30.0, 40.0], - ] + let exprs = [[10.0, 20.0], [30.0, 40.0]] let m = @src.MSnSet::from_names(exprs, ["F1", "F2"], ["S1", "S2"]) let f1 = m.get_feature("F1") assert_eq(f1.length(), 2) @@ -239,12 +311,9 @@ test "mbn_msnset_get_feature" { assert_eq(missing.length(), 0) } +///| test "mbn_msnset_get_sample" { - let exprs = [ - [10.0, 20.0], - [30.0, 40.0], - [50.0, 60.0], - ] + let exprs = [[10.0, 20.0], [30.0, 40.0], [50.0, 60.0]] let m = @src.MSnSet::from_names(exprs, ["F1", "F2", "F3"], ["S1", "S2"]) let s1 = m.get_sample("S1") assert_eq(s1.length(), 3) @@ -258,11 +327,9 @@ test "mbn_msnset_get_sample" { // MSnSet transformations // ============================================================================ +///| test "mbn_msnset_log2_transform" { - let exprs = [ - [1.0, 3.0], - [7.0, 15.0], - ] + let exprs = [[1.0, 3.0], [7.0, 15.0]] let m = @src.MSnSet::from_names(exprs, ["F1", "F2"], ["S1", "S2"]) let m2 = m.log2_transform(offset=1.0) // log2(1+1) = 1, log2(3+1) = 2, log2(7+1) = 3, log2(15+1) = 4 @@ -273,10 +340,11 @@ test "mbn_msnset_log2_transform" { assert_true((e[1][1] - 4.0).abs() < 0.0001) } +///| test "mbn_msnset_impute_mean" { let exprs = [ - [10.0, -1.0, 30.0], // -1 = missing, mean of 10,30 = 20 - [5.0, 15.0, 0.0], // 0 = missing, mean of 5,15 = 10 + [10.0, -1.0, 30.0], // -1 = missing, mean of 10,30 = 20 + [5.0, 15.0, 0.0], // 0 = missing, mean of 5,15 = 10 ] let m = @src.MSnSet::from_names(exprs, ["F1", "F2"], ["S1", "S2", "S3"]) let m2 = m.impute_missing(method="mean") @@ -285,9 +353,10 @@ test "mbn_msnset_impute_mean" { assert_true((e[1][2] - 10.0).abs() < 0.0001) } +///| test "mbn_msnset_impute_median" { let exprs = [ - [1.0, -1.0, 5.0], // sorted 1,5 median = 3 + [1.0, -1.0, 5.0], // sorted 1,5 median = 3 [100.0, 200.0, 0.0], // sorted 100,200 median = 150 ] let m = @src.MSnSet::from_names(exprs, ["F1", "F2"], ["S1", "S2", "S3"]) @@ -297,12 +366,9 @@ test "mbn_msnset_impute_median" { assert_true((e[1][2] - 150.0).abs() < 0.0001) } +///| test "mbn_msnset_normalize_sum" { - let exprs = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - ] + let exprs = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0]] // S1 sum = 6, S2 sum = 60. target = 12 // S1 scale = 2, so 2, 4, 6 // S2 scale = 0.2, so 2, 4, 6 @@ -315,28 +381,26 @@ test "mbn_msnset_normalize_sum" { assert_true((e[1][1] - 4.0).abs() < 0.0001) } +///| test "mbn_msnset_normalize_median_center" { - let exprs = [ - [1.0, 10.0], - [3.0, 20.0], - [5.0, 30.0], - ] + let exprs = [[1.0, 10.0], [3.0, 20.0], [5.0, 30.0]] // S1 median = 3, S2 median = 20 // After: [-2, -10, 2] and [0, 0, 10] let m = @src.MSnSet::from_names(exprs, ["F1", "F2", "F3"], ["S1", "S2"]) let m2 = m.normalize_median_center() let e = m2.exprs() - assert_true((e[0][0] - (-2.0)).abs() < 0.0001) + assert_true((e[0][0] - -2.0).abs() < 0.0001) assert_true((e[2][0] - 2.0).abs() < 0.0001) assert_true((e[1][1] - 0.0).abs() < 0.0001) assert_true((e[2][1] - 10.0).abs() < 0.0001) } +///| test "mbn_msnset_summarize_proteins_sum" { // Two peptides from P1, one from P2 let exprs = [ - [100.0, 200.0], // Peptide1 -> P1 - [300.0, 400.0], // Peptide2 -> P1 + [100.0, 200.0], // Peptide1 -> P1 + [300.0, 400.0], // Peptide2 -> P1 [1000.0, 2000.0], // Peptide3 -> P2 ] let fd = [ @@ -358,8 +422,12 @@ test "mbn_msnset_summarize_proteins_sum" { let mut p1_idx = -1 let mut p2_idx = -1 for i in 0..= 0) assert_true(p2_idx >= 0) @@ -372,19 +440,27 @@ test "mbn_msnset_summarize_proteins_sum" { // Spectrum processing // ============================================================================ +///| test "mbn_spectrum_smooth_ma" { let mz = [100.0, 200.0, 300.0, 400.0, 500.0, 600.0, 700.0] let int = [1.0, 100.0, 1.0, 100.0, 1.0, 100.0, 1.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let smoothed = sp.smooth_moving_average(half_window=1) // Position 1 (100): avg of [1, 100, 1] = 34 // (due to hw=1 => 3-point average) let si = smoothed.intensity() // Check that smoothing reduces the extremes (values become more moderate) - assert_true(si[1] < 100.0) // was spike - assert_true(si[2] > 1.0) // was valley + assert_true(si[1] < 100.0) // was spike + assert_true(si[2] > 1.0) // was valley } +///| test "mbn_spectrum_baseline_correct" { let n = 20 let mz : Array[Double] = [] @@ -396,7 +472,13 @@ test "mbn_spectrum_baseline_correct" { let peak = if i == 10 { 900.0 } else { 0.0 } int.push(base + peak) } - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let corrected = sp.baseline_correct_minwin(half_window=2) // Peak center should still be present (around 900), others should be near 0 let ci = corrected.intensity() @@ -411,12 +493,19 @@ test "mbn_spectrum_baseline_correct" { } } +///| test "mbn_spectrum_centroid_simple" { // Three peaks: at 99, 100, 101 (peak in center, symmetric) // plus a trough let mz = [90.0, 99.0, 100.0, 101.0, 110.0, 199.0, 200.0, 201.0, 300.0] let int = [1.0, 50.0, 100.0, 50.0, 1.0, 60.0, 200.0, 60.0, 1.0] - let sp = @src.Spectrum::new(mz, int, @src.ms_level_ms1(), @src.polarity_positive(), 0.0) + let sp = @src.Spectrum::new( + mz, + int, + @src.ms_level_ms1(), + @src.polarity_positive(), + 0.0, + ) let centroided = sp.centroid_simple(snr_threshold=0.0) // Should find two peaks near 100 and 200 assert_true(centroided.peaks_count() >= 2) @@ -426,6 +515,7 @@ test "mbn_spectrum_centroid_simple" { // Sample QC // ============================================================================ +///| test "mbn_msnset_sample_qc" { let exprs = [ [100.0, 0.0], @@ -434,7 +524,9 @@ test "mbn_msnset_sample_qc" { [0.0, 3000.0], [500.0, 0.0], ] - let m = @src.MSnSet::from_names(exprs, ["F1", "F2", "F3", "F4", "F5"], ["S1", "S2"]) + let m = @src.MSnSet::from_names(exprs, ["F1", "F2", "F3", "F4", "F5"], [ + "S1", "S2", + ]) let qc = m.sample_qc() assert_eq(qc.length(), 2) // S1: detected 4 (one zero), total = 1100 diff --git a/test/moonbit/msstats_test.mbt b/test/moonbit/msstats_test.mbt index 89a3f5e7..e8ba26e7 100644 --- a/test/moonbit/msstats_test.mbt +++ b/test/moonbit/msstats_test.mbt @@ -1,6 +1,5 @@ ///| /// Test file for MSstats module. - test "msstats_feature_creation" { let f = @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1000.0) assert_eq(f.protein, "P1") @@ -12,12 +11,20 @@ test "msstats_feature_creation" { assert_eq(f.intensity, 1000.0) } +///| test "msstats_data_process_log_transform" { let features : Array[@src.MSFeature] = Array::new() - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0)) - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Trt", "S2", "R2", 2048.0)) - - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_none()) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0), + ) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Trt", "S2", "R2", 2048.0), + ) + + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_none(), + ) assert_eq(processed.length(), 2) // log2(1024) = 10 assert_true((processed[0].intensity - 10.0).abs() < 0.01) @@ -25,38 +32,64 @@ test "msstats_data_process_log_transform" { assert_true((processed[1].intensity - 11.0).abs() < 0.01) } +///| test "msstats_data_process_median_norm" { let features : Array[@src.MSFeature] = Array::new() - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0)) - features.push(@src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0)) - features.push(@src.MSFeature::new("P2", "pep3", "tr3", "Trt", "S2", "R2", 2048.0)) - features.push(@src.MSFeature::new("P2", "pep4", "tr4", "Trt", "S2", "R2", 8192.0)) - - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_median()) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0), + ) + features.push( + @src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0), + ) + features.push( + @src.MSFeature::new("P2", "pep3", "tr3", "Trt", "S2", "R2", 2048.0), + ) + features.push( + @src.MSFeature::new("P2", "pep4", "tr4", "Trt", "S2", "R2", 8192.0), + ) + + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_median(), + ) assert_eq(processed.length(), 4) // Values should be centered (median subtracted) // R1: log2(1024)=10, log2(4096)=12, median=11, so values become -1, 1 // R2: log2(2048)=11, log2(8192)=13, median=12, so values become -1, 1 - assert_true((processed[0].intensity - (-1.0)).abs() < 0.01) + assert_true((processed[0].intensity - -1.0).abs() < 0.01) assert_true((processed[1].intensity - 1.0).abs() < 0.01) } +///| test "msstats_data_process_zero_intensity" { let features : Array[@src.MSFeature] = Array::new() - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 0.0)) - - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_none()) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 0.0), + ) + + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_none(), + ) // Zero should be replaced with 1, then log2(1) = 0 assert_true(processed[0].intensity.abs() < 0.01) } +///| test "msstats_summarize_tukey" { let features : Array[@src.MSFeature] = Array::new() // Two peptides for protein P1 in run R1 - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0)) - features.push(@src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0)) - - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_none()) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0), + ) + features.push( + @src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0), + ) + + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_none(), + ) let summarized = @src.ms_summarize(processed, method=@src.ms_summary_tukey()) assert_eq(summarized.length(), 1) // One protein-run combo assert_eq(summarized[0].protein, "P1") @@ -65,43 +98,72 @@ test "msstats_summarize_tukey" { assert_true((summarized[0].log2_abundance - 11.0).abs() < 0.01) } +///| test "msstats_summarize_linear" { let features : Array[@src.MSFeature] = Array::new() - features.push(@src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0)) - features.push(@src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0)) - - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_none()) + features.push( + @src.MSFeature::new("P1", "pep1", "tr1", "Ctrl", "S1", "R1", 1024.0), + ) + features.push( + @src.MSFeature::new("P1", "pep2", "tr2", "Ctrl", "S1", "R1", 4096.0), + ) + + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_none(), + ) let summarized = @src.ms_summarize(processed, method=@src.ms_summary_linear()) // Mean of log2(1024)=10 and log2(4096)=12 is 11 assert_true((summarized[0].log2_abundance - 11.0).abs() < 0.01) } +///| test "msstats_group_comparison_basic" { let features = @src.msstats_sample_data() - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_median()) + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_median(), + ) let summarized = @src.ms_summarize(processed) - let results = @src.ms_group_comparison(summarized, fdr_threshold=0.5, log2fc_threshold=0.5) + let results = @src.ms_group_comparison( + summarized, + fdr_threshold=0.5, + log2fc_threshold=0.5, + ) assert_true(results.get_n_results() > 0) } +///| test "msstats_group_comparison_significant" { let features = @src.msstats_sample_data() - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_median()) + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_median(), + ) let summarized = @src.ms_summarize(processed) - let results = @src.ms_group_comparison(summarized, fdr_threshold=0.5, log2fc_threshold=0.0) + let results = @src.ms_group_comparison( + summarized, + fdr_threshold=0.5, + log2fc_threshold=0.0, + ) let sig = results.get_significant() assert_true(sig.length() >= 0) } +///| test "msstats_group_comparison_top_proteins" { let features = @src.msstats_sample_data() - let processed = @src.ms_data_process(features, normalization=@src.ms_norm_median()) + let processed = @src.ms_data_process( + features, + normalization=@src.ms_norm_median(), + ) let summarized = @src.ms_summarize(processed) let results = @src.ms_group_comparison(summarized) let top = results.get_top_proteins(2) assert_eq(top.length(), 2) } +///| test "msstats_group_comparison_summary" { let features = @src.msstats_sample_data() let processed = @src.ms_data_process(features) @@ -111,6 +173,7 @@ test "msstats_group_comparison_summary" { assert_true(s.contains("MSstats")) } +///| test "msstats_sample_size" { let features = @src.msstats_sample_data() let processed = @src.ms_data_process(features) @@ -120,12 +183,14 @@ test "msstats_sample_size" { assert_true(ss.n_samples >= 3) } +///| test "msstats_sample_data" { let features = @src.msstats_sample_data() assert_eq(features.length(), 40) // 5 proteins * 4 groups * 2 peptides assert_eq(features[0].protein, "P1") } +///| test "msstats_type_to_string" { assert_eq(@src.ms_type_dda().to_string(), "DDA") assert_eq(@src.ms_type_dia().to_string(), "DIA") @@ -133,6 +198,7 @@ test "msstats_type_to_string" { assert_eq(@src.ms_type_tmt().to_string(), "TMT") } +///| test "msstats_norm_to_string" { assert_eq(@src.ms_norm_none().to_string(), "none") assert_eq(@src.ms_norm_median().to_string(), "median") @@ -140,6 +206,7 @@ test "msstats_norm_to_string" { assert_eq(@src.ms_norm_global().to_string(), "globalStandards") } +///| test "msstats_summary_to_string" { assert_eq(@src.ms_summary_tukey().to_string(), "Tukey") assert_eq(@src.ms_summary_linear().to_string(), "linear") diff --git a/test/moonbit/muscat_advanced_test.mbt b/test/moonbit/muscat_advanced_test.mbt new file mode 100644 index 00000000..f428aeb6 --- /dev/null +++ b/test/moonbit/muscat_advanced_test.mbt @@ -0,0 +1,1274 @@ +///| +fn muscat_adv_test_data() -> @src.MuscatAdvancedData { + @src.muscat_advanced_example() catch { + _ => abort("valid muscat advanced example should build") + } +} + +///| +fn muscat_adv_test_config() -> @src.MuscatAdvancedConfig { + @src.MuscatAdvancedConfig::create( + min_cells=10, + min_count=1.0, + min_samples=2, + max_iterations=100, + tolerance=1.0e-8, + minimum_dispersion=1.0e-8, + maximum_dispersion=50.0, + dispersion_prior_df=10.0, + ridge=1.0e-6, + fdr_threshold=0.1, + detection_filter=0.9, + ) catch { + _ => abort("valid muscat advanced configuration should build") + } +} + +///| +fn muscat_adv_test_sum() -> @src.MuscatAdvancedPseudoBulk { + @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_sum_counts(), + ) catch { + _ => abort("sum-count pseudobulk should build") + } +} + +///| +fn muscat_adv_test_detection() -> @src.MuscatAdvancedPseudoBulk { + @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_number_detected(), + ) catch { + _ => abort("number-detected pseudobulk should build") + } +} + +///| +fn muscat_adv_test_design() -> @src.MuscatAdvancedDesign { + @src.muscat_group_design_advanced(muscat_adv_test_sum(), reference="ctrl") catch { + _ => abort("valid muscat group design should build") + } +} + +///| +fn muscat_adv_test_contrasts() -> Array[@src.MuscatAdvancedContrast] { + @src.muscat_default_contrasts_advanced(muscat_adv_test_design()) catch { + _ => abort("valid muscat contrast should build") + } +} + +///| +fn muscat_adv_test_ds() -> @src.MuscatAdvancedResults { + @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ) catch { + _ => abort("valid muscat DS model should fit") + } +} + +///| +fn muscat_adv_test_dd() -> @src.MuscatAdvancedResults { + @src.muscat_pbdd_advanced( + muscat_adv_test_detection(), + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ) catch { + _ => abort("valid muscat DD model should fit") + } +} + +///| +fn muscat_adv_test_stagewise() -> @src.MuscatStagewiseResults { + @src.muscat_stagewise_ds_dd( + muscat_adv_test_ds(), + muscat_adv_test_dd(), + alpha=0.1, + ) catch { + _ => abort("valid muscat stagewise analysis should run") + } +} + +///| +fn muscat_adv_test_finite(value : Double) -> Bool { + value == value && value.abs() < 1.0e299 +} + +///| +fn muscat_adv_test_sce() -> @src.SingleCellExperiment { + let data = muscat_adv_test_data() + let experiment = @src.SingleCellExperiment::new( + data.counts, + data.gene_names, + data.cell_names, + ) + experiment.col_data["sample"] = data.sample_ids.copy() + experiment.col_data["cluster"] = data.cluster_ids.copy() + experiment.col_data["group"] = data.group_ids.copy() + experiment.col_data["batch"] = Array::make(data.cell_names.length(), "one") + experiment.row_data["symbol"] = data.gene_names.copy() + experiment.reduced_dims["PCA"] = [[1.0, 2.0], [3.0, 4.0]] + experiment.metadata["source"] = "test" + experiment.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[1.0, 2.0]], + ["spike1"], + ["s1", "s2"], + ) + experiment +} + +///| +test "muscat advanced default configuration follows pseudobulk defaults" { + let config = @src.MuscatAdvancedConfig::default() + assert_eq(config.min_cells, 10) + assert_eq(config.min_count, 1.0) + assert_eq(config.min_samples, 2) + assert_eq(config.max_iterations, 100) + assert_eq(config.tolerance, 1.0e-8) + assert_eq(config.minimum_dispersion, 1.0e-8) + assert_eq(config.maximum_dispersion, 100.0) + assert_eq(config.dispersion_prior_df, 10.0) + assert_eq(config.ridge, 1.0e-8) + assert_eq(config.fdr_threshold, 0.05) + assert_eq(config.lfc_threshold, 0.0) + assert_eq(config.detection_filter, 0.9) +} + +///| +test "muscat advanced configuration preserves explicit controls" { + let config = @src.MuscatAdvancedConfig::create( + min_cells=3, + min_count=2.0, + min_samples=3, + max_iterations=40, + tolerance=1.0e-5, + minimum_dispersion=1.0e-4, + maximum_dispersion=20.0, + dispersion_prior_df=5.0, + ridge=1.0e-5, + fdr_threshold=0.2, + lfc_threshold=1.0, + detection_filter=0.8, + ) catch { + _ => abort("explicit muscat controls should be valid") + } + assert_eq(config.min_cells, 3) + assert_eq(config.min_count, 2.0) + assert_eq(config.min_samples, 3) + assert_eq(config.max_iterations, 40) + assert_eq(config.maximum_dispersion, 20.0) + assert_eq(config.dispersion_prior_df, 5.0) + assert_eq(config.fdr_threshold, 0.2) + assert_eq(config.lfc_threshold, 1.0) + assert_eq(config.detection_filter, 0.8) +} + +///| +test "muscat advanced rejects invalid filtering controls" { + let mut failures = 0 + ignore(@src.MuscatAdvancedConfig::create(min_cells=-1)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(min_count=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(min_count=0.0 / 0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(min_samples=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "muscat advanced rejects invalid iteration controls" { + let mut failures = 0 + ignore(@src.MuscatAdvancedConfig::create(max_iterations=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(tolerance=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(tolerance=0.0 / 0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced rejects invalid dispersion controls" { + let mut failures = 0 + ignore(@src.MuscatAdvancedConfig::create(minimum_dispersion=0.0)) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedConfig::create( + minimum_dispersion=2.0, + maximum_dispersion=1.0, + ), + ) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(dispersion_prior_df=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(ridge=-1.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "muscat advanced rejects invalid testing controls" { + let mut failures = 0 + ignore(@src.MuscatAdvancedConfig::create(fdr_threshold=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(fdr_threshold=1.1)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(lfc_threshold=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(detection_filter=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedConfig::create(detection_filter=1.1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 5) +} + +///| +test "muscat advanced aggregation helpers expose all upstream statistics" { + assert_eq(@src.muscat_sum_counts(), @src.muscat_sum_counts()) + assert_eq(@src.muscat_mean_expression(), @src.muscat_mean_expression()) + assert_eq(@src.muscat_median_expression(), @src.muscat_median_expression()) + assert_eq( + @src.muscat_proportion_detected(), + @src.muscat_proportion_detected(), + ) + assert_eq(@src.muscat_number_detected(), @src.muscat_number_detected()) + assert_true(@src.muscat_sum_counts() != @src.muscat_mean_expression()) +} + +///| +test "muscat advanced example exposes gene by cell metadata contract" { + let data = muscat_adv_test_data() + assert_eq(data.counts.length(), 10) + assert_eq(data.counts[0].length(), 160) + assert_eq(data.gene_names.length(), 10) + assert_eq(data.cell_names.length(), 160) + assert_eq(data.sample_names, ["C1", "C2", "C3", "C4", "T1", "T2", "T3", "T4"]) + assert_eq(data.cluster_names, ["A", "B"]) + assert_eq(data.group_names, ["ctrl", "stim"]) + assert_eq(data.sample_groups, [ + "ctrl", "ctrl", "ctrl", "ctrl", "stim", "stim", "stim", "stim", + ]) +} + +///| +test "muscat advanced validates count matrix dimensions" { + let mut failures = 0 + ignore(@src.MuscatAdvancedData::create([], [], [], [])) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedData::create([[]], [], [], [])) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create( + [[1.0, 2.0], [1.0]], + ["s1", "s2"], + ["A", "A"], + ["c", "t"], + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced validates count values" { + let mut failures = 0 + ignore( + @src.MuscatAdvancedData::create([[-1.0, 2.0]], ["s1", "s2"], ["A", "A"], [ + "c", "t", + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create([[1.5, 2.0]], ["s1", "s2"], ["A", "A"], [ + "c", "t", + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create( + [[0.0 / 0.0, 2.0]], + ["s1", "s2"], + ["A", "A"], + ["c", "t"], + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced validates cell metadata dimensions and values" { + let mut failures = 0 + ignore( + @src.MuscatAdvancedData::create([[1.0, 2.0]], ["s1"], ["A", "A"], ["c", "t"]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create([[1.0, 2.0]], ["s1", "s2"], ["A", " "], [ + "c", "t", + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "muscat advanced requires replicated sample and group identities" { + let mut failures = 0 + ignore( + @src.MuscatAdvancedData::create([[1.0, 2.0]], ["s1", "s1"], ["A", "A"], [ + "c", "t", + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create([[1.0, 2.0]], ["s1", "s2"], ["A", "A"], [ + "c", "c", + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "muscat advanced requires each sample to have one group" { + let failed = try { + ignore( + @src.MuscatAdvancedData::create( + [[1.0, 2.0, 3.0, 4.0]], + ["s1", "s1", "s2", "s2"], + ["A", "A", "A", "A"], + ["c", "t", "t", "t"], + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced validates gene and cell names" { + let mut failures = 0 + ignore( + @src.MuscatAdvancedData::create( + [[1.0, 2.0], [2.0, 3.0]], + ["s1", "s2"], + ["A", "A"], + ["c", "t"], + gene_names=["one"], + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create( + [[1.0, 2.0]], + ["s1", "s2"], + ["A", "A"], + ["c", "t"], + gene_names=[" "], + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedData::create( + [[1.0, 2.0]], + ["s1", "s2"], + ["A", "A"], + ["c", "t"], + cell_names=["cell", "cell"], + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced data owns copies of caller arrays" { + let counts = [[1.0, 2.0]] + let samples = ["s1", "s2"] + let clusters = ["A", "A"] + let groups = ["c", "t"] + let data = @src.MuscatAdvancedData::create(counts, samples, clusters, groups) catch { + _ => abort("valid copied data should build") + } + data.counts[0][0] = 99.0 + data.sample_ids[0] = "changed" + assert_eq(counts[0][0], 1.0) + assert_eq(samples[0], "s1") +} + +///| +test "muscat advanced sum pseudobulk has cluster gene sample orientation" { + let pseudobulk = muscat_adv_test_sum() + assert_eq(pseudobulk.values.length(), 2) + assert_eq(pseudobulk.values[0].length(), 10) + assert_eq(pseudobulk.values[0][0].length(), 8) + assert_eq(pseudobulk.cluster_names, ["A", "B"]) + assert_eq(pseudobulk.cell_counts[0], [10, 10, 10, 10, 10, 10, 10, 10]) + assert_eq(pseudobulk.aggregation, @src.muscat_sum_counts()) + assert_false(pseudobulk.scaled_cpm) +} + +///| +test "muscat advanced sum aggregation preserves designed signals" { + let pseudobulk = muscat_adv_test_sum() + assert_eq(pseudobulk.values[0][0][0], 1009.0) + assert_eq(pseudobulk.values[0][1][0], 39.0) + assert_eq(pseudobulk.values[0][1][4], 189.0) + assert_eq(pseudobulk.values[1][2][0], 100.0) + assert_eq(pseudobulk.values[1][2][4], 100.0) + assert_eq(pseudobulk.values[0][3][0], 30.0) + assert_eq(pseudobulk.values[0][3][4], 100.0) +} + +///| +test "muscat advanced mean aggregation divides by cluster sample cells" { + let pseudobulk = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_mean_expression(), + ) + assert_true((pseudobulk.values[0][0][0] - 100.9).abs() < 1.0e-12) + assert_eq(pseudobulk.values[1][2][0], 10.0) + assert_eq(pseudobulk.values[1][2][4], 10.0) +} + +///| +test "muscat advanced median aggregation handles even cell counts" { + let pseudobulk = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_median_expression(), + ) + assert_eq(pseudobulk.values[0][0][0], 101.0) + assert_eq(pseudobulk.values[1][2][0], 0.0) + assert_eq(pseudobulk.values[1][2][4], 10.0) +} + +///| +test "muscat advanced proportion detected stays on unit interval" { + let pseudobulk = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_proportion_detected(), + ) + assert_eq(pseudobulk.values[0][0][0], 1.0) + assert_eq(pseudobulk.values[1][2][0], 0.1) + assert_eq(pseudobulk.values[1][2][4], 1.0) + for cluster in pseudobulk.values { + for gene in cluster { + for value in gene { + assert_true(value >= 0.0 && value <= 1.0) + } + } + } +} + +///| +test "muscat advanced number detected counts positive cells" { + let pseudobulk = muscat_adv_test_detection() + assert_eq(pseudobulk.values[0][0][0], 10.0) + assert_eq(pseudobulk.values[1][2][0], 1.0) + assert_eq(pseudobulk.values[1][2][4], 10.0) + assert_eq(pseudobulk.values[0][3][0], 1.0) + assert_eq(pseudobulk.values[0][3][4], 10.0) +} + +///| +test "muscat advanced aggregation fills missing cluster sample combinations" { + let data = @src.MuscatAdvancedData::create( + [[1.0, 2.0]], + ["s1", "s2"], + ["A", "B"], + ["ctrl", "stim"], + gene_names=["g"], + ) + let pseudobulk = @src.muscat_aggregate_advanced(data) + assert_eq(pseudobulk.values[0][0], [1.0, 0.0]) + assert_eq(pseudobulk.values[1][0], [0.0, 2.0]) + assert_eq(pseudobulk.cell_counts[0], [1, 0]) + assert_eq(pseudobulk.cell_counts[1], [0, 1]) +} + +///| +test "muscat advanced library sizes retain raw sum counts" { + let sums = muscat_adv_test_sum() + let means = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_mean_expression(), + ) + assert_eq(sums.library_sizes[0][0], 1107.0) + assert_eq(sums.library_sizes[1][0], 1177.0) + assert_eq(means.library_sizes, sums.library_sizes) +} + +///| +test "muscat advanced CPM scaling uses raw cluster sample library size" { + let pseudobulk = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_sum_counts(), + scale_cpm=true, + ) + assert_true(pseudobulk.scaled_cpm) + assert_true( + (pseudobulk.values[0][0][0] - 1009.0 / 1107.0 * 1.0e6).abs() < 1.0e-8, + ) +} + +///| +test "muscat advanced rejects CPM scaling of detection counts" { + let failed = try { + ignore( + @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_number_detected(), + scale_cpm=true, + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced group design uses reference coding" { + let design = muscat_adv_test_design() + assert_eq(design.matrix.length(), 8) + assert_eq(design.matrix[0], [1.0, 0.0]) + assert_eq(design.matrix[3], [1.0, 0.0]) + assert_eq(design.matrix[4], [1.0, 1.0]) + assert_eq(design.coefficient_names, ["(Intercept)", "stim-ctrl"]) + assert_eq(design.sample_names, muscat_adv_test_sum().sample_names) +} + +///| +test "muscat advanced group design supports alternate reference" { + let design = @src.muscat_group_design_advanced( + muscat_adv_test_sum(), + reference="stim", + ) + assert_eq(design.matrix[0], [1.0, 1.0]) + assert_eq(design.matrix[4], [1.0, 0.0]) + assert_eq(design.coefficient_names, ["(Intercept)", "ctrl-stim"]) +} + +///| +test "muscat advanced group design rejects absent reference" { + let failed = try { + ignore( + @src.muscat_group_design_advanced( + muscat_adv_test_sum(), + reference="missing", + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced custom design preserves arbitrary covariates" { + let design = @src.MuscatAdvancedDesign::create( + [[1.0, 0.0], [1.0, 0.0], [1.0, 1.0], [1.0, 1.0]], + ["s1", "s2", "s3", "s4"], + ["intercept", "treatment"], + ) + assert_eq(design.matrix[2], [1.0, 1.0]) + assert_eq(design.coefficient_names, ["intercept", "treatment"]) +} + +///| +test "muscat advanced validates design dimensions and names" { + let mut failures = 0 + ignore(@src.MuscatAdvancedDesign::create([], [], [])) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedDesign::create([[1.0], [1.0]], ["s1"], ["i"])) catch { + _ => failures = failures + 1 + } + ignore(@src.MuscatAdvancedDesign::create([[1.0], [1.0]], ["s1", "s1"], ["i"])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced requires finite full-rank residual design" { + let mut failures = 0 + ignore( + @src.MuscatAdvancedDesign::create([[1.0, 0.0], [0.0, 1.0]], ["s1", "s2"], [ + "a", "b", + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedDesign::create( + [[1.0, 1.0], [1.0, 1.0], [1.0, 1.0]], + ["s1", "s2", "s3"], + ["a", "b"], + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.MuscatAdvancedDesign::create([[1.0], [0.0 / 0.0]], ["s1", "s2"], ["a"]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced contrast preserves coefficient vector" { + let contrast = @src.muscat_contrast_advanced( + muscat_adv_test_design(), + [0.0, 1.0], + "stimulus", + ) + assert_eq(contrast.values, [0.0, 1.0]) + assert_eq(contrast.name, "stimulus") +} + +///| +test "muscat advanced validates contrast dimensions values and name" { + let design = muscat_adv_test_design() + let mut failures = 0 + ignore(@src.muscat_contrast_advanced(design, [1.0], "short")) catch { + _ => failures = failures + 1 + } + ignore(@src.muscat_contrast_advanced(design, [0.0, 0.0], "zero")) catch { + _ => failures = failures + 1 + } + ignore(@src.muscat_contrast_advanced(design, [0.0, 0.0 / 0.0], "nan")) catch { + _ => failures = failures + 1 + } + ignore(@src.muscat_contrast_advanced(design, [0.0, 1.0], " ")) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "muscat advanced default contrasts cover all non-intercept coefficients" { + let design = @src.MuscatAdvancedDesign::create( + [ + [1.0, 0.0, 0.0], + [1.0, 0.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 1.0, 0.0], + [1.0, 0.0, 1.0], + [1.0, 0.0, 1.0], + ], + ["s1", "s2", "s3", "s4", "s5", "s6"], + ["intercept", "B-A", "C-A"], + ) + let contrasts = @src.muscat_default_contrasts_advanced(design) + assert_eq(contrasts.length(), 2) + assert_eq(contrasts[0].name, "B-A") + assert_eq(contrasts[0].values, [0.0, 1.0, 0.0]) + assert_eq(contrasts[1].values, [0.0, 0.0, 1.0]) +} + +///| +test "muscat advanced testing requires exact pseudobulk sample order" { + let design = muscat_adv_test_design() + let reversed = @src.MuscatAdvancedDesign::create( + design.matrix, + ["T4", "T3", "T2", "T1", "C4", "C3", "C2", "C1"], + design.coefficient_names, + ) + let failed = try { + ignore( + @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + reversed, + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced testing validates contrast collection" { + let design = muscat_adv_test_design() + let contrast = muscat_adv_test_contrasts()[0] + let mut failures = 0 + ignore( + @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + design, + [], + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + design, + [contrast, contrast], + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "muscat advanced DS requires unscaled sum counts" { + let mean = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + aggregation=@src.muscat_mean_expression(), + ) + let scaled = @src.muscat_aggregate_advanced( + muscat_adv_test_data(), + scale_cpm=true, + ) + let mut failures = 0 + ignore( + @src.muscat_pbds_advanced( + mean, + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_pbds_advanced( + scaled, + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "muscat advanced DD requires number-detected pseudobulk" { + let failed = try { + ignore( + @src.muscat_pbdd_advanced( + muscat_adv_test_sum(), + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=muscat_adv_test_config(), + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced reports when filtering removes every model" { + let failed = try { + ignore( + @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=@src.MuscatAdvancedConfig::create(min_cells=11), + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "muscat advanced DS returns cluster gene contrast results" { + let result = muscat_adv_test_ds() + assert_eq(result.mode, "DS") + assert_eq(result.results.length(), 20) + assert_eq(result.gene_names, muscat_adv_test_data().gene_names) + assert_eq(result.cluster_names, ["A", "B"]) + assert_eq(result.contrast_names, ["stim-ctrl"]) + assert_eq(result.n_tested(), 18) + assert_true(result.summary().contains("10 genes x 2 clusters x 1 contrasts")) +} + +///| +test "muscat advanced DS detects cluster-specific abundance signal" { + let result = muscat_adv_test_ds() + let signal = match result.get("ds_a", "A", "stim-ctrl") { + Some(value) => value + None => abort("DS signal result should exist") + } + let inactive = match result.get("ds_a", "B", "stim-ctrl") { + Some(value) => value + None => abort("inactive DS result should exist") + } + assert_true(signal.tested) + assert_true(signal.log_fc > 1.5) + assert_true(signal.statistic > inactive.statistic) + assert_true(signal.p_value < inactive.p_value) + assert_true(signal.local_fdr <= 0.1) +} + +///| +test "muscat advanced DS leaves equal-sum DD signal unchanged" { + let result = muscat_adv_test_ds() + let pure_dd = match result.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("pure DD result should exist") + } + assert_true(pure_dd.tested) + assert_true(pure_dd.log_fc.abs() < 1.0e-5) + assert_true(pure_dd.p_value > 0.9) +} + +///| +test "muscat advanced DS statistics and dispersion are finite" { + let result = muscat_adv_test_ds() + let config = muscat_adv_test_config() + for value in result.results { + assert_true(muscat_adv_test_finite(value.log_fc)) + assert_true(muscat_adv_test_finite(value.log_cpm)) + assert_true(muscat_adv_test_finite(value.statistic)) + assert_true(value.p_value >= 0.0 && value.p_value <= 1.0) + assert_true(value.local_fdr >= value.p_value - 1.0e-12) + assert_true(value.global_fdr >= value.p_value - 1.0e-12) + assert_true(value.dispersion >= config.minimum_dispersion) + assert_true(value.dispersion <= config.maximum_dispersion) + assert_eq(value.n_samples, 8) + } +} + +///| +test "muscat advanced filters all-zero genes without nonfinite values" { + let result = muscat_adv_test_ds() + let zero = match result.get("zero", "A", "stim-ctrl") { + Some(value) => value + None => abort("zero-gene result should exist") + } + assert_false(zero.tested) + assert_eq(zero.statistic, 0.0) + assert_eq(zero.p_value, 1.0) + assert_eq(zero.local_fdr, 1.0) + assert_eq(zero.global_fdr, 1.0) +} + +///| +test "muscat advanced fold-change threshold cannot increase Wald evidence" { + let baseline = muscat_adv_test_ds() + let thresholded = @src.muscat_pbds_advanced( + muscat_adv_test_sum(), + muscat_adv_test_design(), + muscat_adv_test_contrasts(), + config=@src.MuscatAdvancedConfig::create( + min_cells=10, + ridge=1.0e-6, + fdr_threshold=0.1, + lfc_threshold=1.0, + ), + ) + for index in 0..= + baseline.results[index].p_value, + ) + } +} + +///| +test "muscat advanced DD filters ubiquitously detected genes" { + let result = muscat_adv_test_dd() + let stable = match result.get("stable", "A", "stim-ctrl") { + Some(value) => value + None => abort("stable DD result should exist") + } + let ds = match result.get("ds_a", "A", "stim-ctrl") { + Some(value) => value + None => abort("ubiquitous DS gene result should exist") + } + assert_false(stable.tested) + assert_false(ds.tested) + assert_eq(stable.p_value, 1.0) + assert_eq(ds.p_value, 1.0) +} + +///| +test "muscat advanced DD detects equal-sum detection signal" { + let result = muscat_adv_test_dd() + let signal = match result.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("DD signal result should exist") + } + let inactive = match result.get("dd_b", "A", "stim-ctrl") { + Some(value) => value + None => abort("inactive DD result should exist") + } + assert_true(signal.tested) + assert_false(inactive.tested) + assert_true(signal.log_fc > 1.5) + assert_true(signal.p_value < 0.01) + assert_true(signal.local_fdr < 0.1) + assert_true(signal.global_fdr < 0.1) +} + +///| +test "muscat advanced DD detects combined abundance and detection signal" { + let result = muscat_adv_test_dd() + let signal = match result.get("both_a", "A", "stim-ctrl") { + Some(value) => value + None => abort("combined DD signal result should exist") + } + assert_true(signal.tested) + assert_true(signal.log_fc > 1.5) + assert_true(signal.statistic > 5.0) + assert_true(signal.local_fdr < 0.1) +} + +///| +test "muscat advanced DD statistics remain finite after CDR normalization" { + let result = muscat_adv_test_dd() + assert_eq(result.mode, "DD") + assert_eq(result.results.length(), 20) + assert_eq(result.n_tested(), 12) + for value in result.results { + assert_true(muscat_adv_test_finite(value.log_fc)) + assert_true(muscat_adv_test_finite(value.statistic)) + assert_true(muscat_adv_test_finite(value.dispersion)) + assert_true(value.p_value >= 0.0 && value.p_value <= 1.0) + } +} + +///| +test "muscat advanced local and global BH adjustments are bounded" { + for result in [muscat_adv_test_ds(), muscat_adv_test_dd()] { + for value in result.results { + assert_true(value.local_fdr >= 0.0 && value.local_fdr <= 1.0) + assert_true(value.global_fdr >= 0.0 && value.global_fdr <= 1.0) + if value.tested { + assert_true(value.local_fdr + 1.0e-12 >= value.p_value) + assert_true(value.global_fdr + 1.0e-12 >= value.p_value) + } + } + } +} + +///| +test "muscat advanced stagewise output aligns DS and DD hypotheses" { + let result = muscat_adv_test_stagewise() + assert_eq(result.results.length(), 20) + assert_eq(result.gene_names, muscat_adv_test_data().gene_names) + assert_eq(result.cluster_names, ["A", "B"]) + assert_eq(result.contrast_names, ["stim-ctrl"]) + assert_eq(result.alpha, 0.1) + for value in result.results { + assert_true(value.screen_p_value >= 0.0 && value.screen_p_value <= 1.0) + assert_true(value.screen_fdr >= 0.0 && value.screen_fdr <= 1.0) + assert_true(value.ds_confirmation >= 0.0 && value.ds_confirmation <= 1.0) + assert_true(value.dd_confirmation >= 0.0 && value.dd_confirmation <= 1.0) + } +} + +///| +test "muscat advanced stagewise classifies DS DD and combined signals" { + let result = muscat_adv_test_stagewise() + let ds = match result.get("ds_a", "A", "stim-ctrl") { + Some(value) => value + None => abort("stagewise DS result should exist") + } + let dd = match result.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("stagewise DD result should exist") + } + let both = match result.get("both_a", "A", "stim-ctrl") { + Some(value) => value + None => abort("stagewise combined result should exist") + } + assert_eq(ds.classification, "DS") + assert_eq(dd.classification, "DD") + assert_eq(both.classification, "both") + assert_true(result.n_confirmed() >= 3) +} + +///| +test "muscat advanced stagewise screening uses harmonic mean p-value" { + let ds = muscat_adv_test_ds() + let dd = muscat_adv_test_dd() + let result = @src.muscat_stagewise_ds_dd(ds, dd, alpha=0.1) + let ds_gene = match ds.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("DS component should exist") + } + let dd_gene = match dd.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("DD component should exist") + } + let combined = match result.get("dd_b", "B", "stim-ctrl") { + Some(value) => value + None => abort("stagewise component should exist") + } + let expected = 2.0 / (1.0 / ds_gene.p_value + 1.0 / dd_gene.p_value) + assert_true((combined.screen_p_value - expected).abs() < 1.0e-12) +} + +///| +test "muscat advanced stagewise validates alpha and result modes" { + let ds = muscat_adv_test_ds() + let dd = muscat_adv_test_dd() + let mut failures = 0 + ignore(@src.muscat_stagewise_ds_dd(ds, dd, alpha=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.muscat_stagewise_ds_dd(ds, dd, alpha=1.1)) catch { + _ => failures = failures + 1 + } + ignore(@src.muscat_stagewise_ds_dd(ds, ds)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "muscat advanced result lookup returns none for absent keys" { + assert_true(muscat_adv_test_ds().get("missing", "A", "stim-ctrl") is None) + assert_true( + muscat_adv_test_stagewise().get("missing", "A", "stim-ctrl") is None, + ) +} + +///| +test "muscat advanced result summaries report mode and significance" { + let ds = muscat_adv_test_ds() + let dd = muscat_adv_test_dd() + assert_true(ds.summary().contains("muscat advanced DS")) + assert_true(dd.summary().contains("muscat advanced DD")) + assert_true(ds.n_significant() >= 2) + assert_true(ds.n_significant(global=true) >= 2) + assert_eq(dd.n_significant(), 2) + assert_eq(dd.n_significant(global=true), 2) +} + +///| +test "muscat advanced SCE writes DS DD and stagewise annotations" { + let output = @src.muscat_advanced_sce( + muscat_adv_test_sce(), + "sample", + "cluster", + "group", + reference="ctrl", + config=muscat_adv_test_config(), + ) catch { + _ => abort("valid muscat SCE should analyze") + } + assert_eq(output.sum_pseudobulk.values.length(), 2) + assert_eq(output.detection_pseudobulk.values.length(), 2) + assert_eq(output.ds_result.mode, "DS") + assert_eq(output.dd_result.mode, "DD") + assert_eq(output.experiment.row_data["muscat.A.dsLogFC"].length(), 10) + assert_eq(output.experiment.row_data["muscat.A.dsFdr"].length(), 10) + assert_eq(output.experiment.row_data["muscat.B.ddLogFC"].length(), 10) + assert_eq(output.experiment.row_data["muscat.B.stageClass"].length(), 10) + assert_eq(output.experiment.metadata["muscat.version"], "muscat 1.27.4") + assert_eq(output.experiment.metadata["muscat.reference"], "ctrl") + assert_eq(output.experiment.metadata["muscat.contrast"], "stim-ctrl") +} + +///| +test "muscat advanced SCE preserves input immutability" { + let experiment = muscat_adv_test_sce() + let output = @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "group", + config=muscat_adv_test_config(), + ) + output.experiment.assays["counts"][0][0] = 999.0 + output.experiment.row_data["symbol"][0] = "changed" + output.experiment.col_data["batch"][0] = "changed" + output.experiment.reduced_dims["PCA"][0][0] = 999.0 + output.experiment.metadata["source"] = "changed" + assert_eq(experiment.assays["counts"][0][0], 100.0) + assert_eq(experiment.row_data["symbol"][0], "stable") + assert_eq(experiment.col_data["batch"][0], "one") + assert_eq(experiment.reduced_dims["PCA"][0][0], 1.0) + assert_eq(experiment.metadata["source"], "test") + assert_false(experiment.row_data.contains("muscat.A.dsLogFC")) +} + +///| +test "muscat advanced SCE recursively copies alternative experiments" { + let experiment = muscat_adv_test_sce() + let output = @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "group", + config=muscat_adv_test_config(), + ) + output.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 8.0 + assert_eq( + experiment.alternative_experiments["spike"].assays["counts"][0][0], + 1.0, + ) +} + +///| +test "muscat advanced SCE supports custom assay and output prefix" { + let experiment = muscat_adv_test_sce() + experiment.assays["raw"] = experiment.assays["counts"].map(fn(row) { + row.copy() + }) + let output = @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "group", + assay_name="raw", + reference="ctrl", + output_prefix="mx", + config=muscat_adv_test_config(), + ) + assert_true(output.experiment.row_data.contains("mx.A.dsLogFC")) + assert_true(output.experiment.row_data.contains("mx.B.stageClass")) + assert_true(output.experiment.col_data.contains("mx.sample_id")) + assert_eq(output.experiment.metadata["mx.assay"], "raw") + assert_eq(output.experiment.metadata["mx.clusters"], "A,B") +} + +///| +test "muscat advanced SCE validates assay columns and prefix" { + let experiment = muscat_adv_test_sce() + let mut failures = 0 + ignore( + @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "group", + assay_name="missing", + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_advanced_sce( + experiment, + "missing", + "cluster", + "group", + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_advanced_sce( + experiment, + "sample", + "missing", + "group", + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "missing", + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.muscat_advanced_sce( + experiment, + "sample", + "cluster", + "group", + output_prefix=" ", + config=muscat_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 5) +} + +///| +test "muscat advanced SCE validates reference group" { + let failed = try { + ignore( + @src.muscat_advanced_sce( + muscat_adv_test_sce(), + "sample", + "cluster", + "group", + reference="missing", + config=muscat_adv_test_config(), + ), + ) + false + } catch { + _ => true + } + assert_true(failed) +} diff --git a/test/moonbit/muscat_test.mbt b/test/moonbit/muscat_test.mbt index 2ebf4aec..ec4a524c 100644 --- a/test/moonbit/muscat_test.mbt +++ b/test/moonbit/muscat_test.mbt @@ -1,6 +1,5 @@ ///| /// Test file for muscat module. - test "muscat_single_cell_creation" { let cell = @src.SingleCell::new("cell1", "sample1", "cluster1", "ctrl") assert_eq(cell.cell_id, "cell1") @@ -9,6 +8,7 @@ test "muscat_single_cell_creation" { assert_eq(cell.group_id, "ctrl") } +///| test "muscat_single_cell_set_get_count" { let cell = @src.SingleCell::new("cell1", "sample1", "cluster1", "ctrl") cell.set_count("Gene1", 10.0) @@ -18,6 +18,7 @@ test "muscat_single_cell_set_get_count" { assert_eq(cell.get_count("Gene3"), 0.0) // Non-existent } +///| test "muscat_single_cell_total_counts" { let cell = @src.SingleCell::new("cell1", "sample1", "cluster1", "ctrl") cell.set_count("Gene1", 10.0) @@ -26,6 +27,7 @@ test "muscat_single_cell_total_counts" { assert_eq(cell.total_counts(), 35.0) } +///| test "muscat_single_cell_n_expressed" { let cell = @src.SingleCell::new("cell1", "sample1", "cluster1", "ctrl") cell.set_count("Gene1", 10.0) @@ -34,6 +36,7 @@ test "muscat_single_cell_n_expressed" { assert_eq(cell.n_expressed(), 2) } +///| test "muscat_aggregation_sum" { let cells : Array[@src.SingleCell] = Array::new() let c1 = @src.SingleCell::new("c1", "S1", "C1", "ctrl") @@ -52,6 +55,7 @@ test "muscat_aggregation_sum" { assert_eq(pb[0].n_cells, 2) } +///| test "muscat_aggregation_mean" { let cells : Array[@src.SingleCell] = Array::new() let c1 = @src.SingleCell::new("c1", "S1", "C1", "ctrl") @@ -66,6 +70,7 @@ test "muscat_aggregation_mean" { assert_eq(pb[0].get_count("Gene1"), 15.0) } +///| test "muscat_aggregation_multiple_samples" { let cells : Array[@src.SingleCell] = Array::new() // Sample S1, Cluster C1 @@ -85,6 +90,7 @@ test "muscat_aggregation_multiple_samples" { assert_eq(pb.length(), 3) // 3 (sample, cluster) combinations } +///| test "muscat_pseudobulk_total_counts" { let cells : Array[@src.SingleCell] = Array::new() let c1 = @src.SingleCell::new("c1", "S1", "C1", "ctrl") @@ -99,6 +105,7 @@ test "muscat_pseudobulk_total_counts" { assert_eq(pb[0].total_counts(), 35.0) } +///| test "muscat_pseudobulk_n_expressed" { let cells : Array[@src.SingleCell] = Array::new() let c1 = @src.SingleCell::new("c1", "S1", "C1", "ctrl") @@ -114,6 +121,7 @@ test "muscat_pseudobulk_n_expressed" { assert_eq(pb[0].n_expressed(), 2) // Gene1 and Gene3 > 0 } +///| test "muscat_ds_analysis_basic" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) @@ -122,14 +130,20 @@ test "muscat_ds_analysis_basic" { assert_eq(results.n_clusters, 2) } +///| test "muscat_ds_analysis_significant" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) - let results = @src.run_ds_analysis(pb, fdr_threshold=0.5, log2fc_threshold=0.5) + let results = @src.run_ds_analysis( + pb, + fdr_threshold=0.5, + log2fc_threshold=0.5, + ) let sig = results.get_significant() assert_true(sig.length() > 0) } +///| test "muscat_ds_analysis_cluster_filter" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) @@ -138,6 +152,7 @@ test "muscat_ds_analysis_cluster_filter" { assert_true(c1_results.length() > 0) } +///| test "muscat_ds_analysis_top_genes" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) @@ -146,6 +161,7 @@ test "muscat_ds_analysis_top_genes" { assert_eq(top.length(), 3) } +///| test "muscat_qc_computation" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) @@ -153,6 +169,7 @@ test "muscat_qc_computation" { assert_eq(qc.length(), pb.length()) } +///| test "muscat_summary" { let cells = @src.muscat_sample_data() let pb = @src.aggregate_cells(cells) @@ -161,12 +178,14 @@ test "muscat_summary" { assert_true(s.contains("muscat")) } +///| test "muscat_aggregation_method_to_string" { assert_eq(@src.aggregation_sum().to_string(), "sum") assert_eq(@src.aggregation_mean().to_string(), "mean") assert_eq(@src.aggregation_median().to_string(), "median") } +///| test "muscat_ds_method_to_string" { assert_eq(@src.ds_method_edger().to_string(), "edgeR") assert_eq(@src.ds_method_deseq2().to_string(), "DESeq2") diff --git a/test/moonbit/naccess_test.mbt b/test/moonbit/naccess_test.mbt index 6302efb1..57a3a197 100644 --- a/test/moonbit/naccess_test.mbt +++ b/test/moonbit/naccess_test.mbt @@ -127,9 +127,13 @@ test "naccess_result_new_empty" { test "naccess_result_add_residue" { let r = @src.NaccessResult::new() assert_eq(r.get_num_residues(), 0) - r.add_residue(@src.NaccessResidue::new(res_name="ALA", res_num=1, chain_id="A")) + r.add_residue( + @src.NaccessResidue::new(res_name="ALA", res_num=1, chain_id="A"), + ) assert_eq(r.get_num_residues(), 1) - r.add_residue(@src.NaccessResidue::new(res_name="GLY", res_num=2, chain_id="A")) + r.add_residue( + @src.NaccessResidue::new(res_name="GLY", res_num=2, chain_id="A"), + ) assert_eq(r.get_num_residues(), 2) } @@ -164,8 +168,12 @@ test "naccess_result_add_atom" { ///| test "naccess_result_get_residues" { let r = @src.NaccessResult::new() - r.add_residue(@src.NaccessResidue::new(res_name="MET", res_num=1, chain_id="A")) - r.add_residue(@src.NaccessResidue::new(res_name="ALA", res_num=2, chain_id="A")) + r.add_residue( + @src.NaccessResidue::new(res_name="MET", res_num=1, chain_id="A"), + ) + r.add_residue( + @src.NaccessResidue::new(res_name="ALA", res_num=2, chain_id="A"), + ) let residues = r.get_residues() assert_eq(residues.length(), 2) assert_eq(residues[0].res_name, "MET") @@ -206,7 +214,9 @@ test "naccess_result_get_num_residues" { let r = @src.NaccessResult::new() assert_eq(r.get_num_residues(), 0) for i in 0..<5 { - r.add_residue(@src.NaccessResidue::new(res_name="ALA", res_num=i, chain_id="A")) + r.add_residue( + @src.NaccessResidue::new(res_name="ALA", res_num=i, chain_id="A"), + ) } assert_eq(r.get_num_residues(), 5) } @@ -409,26 +419,39 @@ test "naccess_parse_asa_header_only" { ///| test "naccess_parse_combined_counts" { - let result = @src.naccess_parse(@src.naccess_sample_rsa(), @src.naccess_sample_asa()) + let result = @src.naccess_parse( + @src.naccess_sample_rsa(), + @src.naccess_sample_asa(), + ) assert_eq(result.get_num_residues(), 4) assert_eq(result.get_num_atoms(), 6) } ///| test "naccess_parse_total_abs_asa" { - let result = @src.naccess_parse(@src.naccess_sample_rsa(), @src.naccess_sample_asa()) + let result = @src.naccess_parse( + @src.naccess_sample_rsa(), + @src.naccess_sample_asa(), + ) // 45.3 + 20.1 + 5.2 + 25.0 = 95.6 assert_true((result.total_abs_asa - 95.6).abs() < naccess_eps) } ///| test "naccess_parse_chain_totals" { - let result = @src.naccess_parse(@src.naccess_sample_rsa(), @src.naccess_sample_asa()) + let result = @src.naccess_parse( + @src.naccess_sample_rsa(), + @src.naccess_sample_asa(), + ) assert_eq(result.chain_totals.length(), 2) // Chain A: 45.3 + 20.1 + 5.2 = 70.6 - assert_true((result.chain_totals.get("A").unwrap() - 70.6).abs() < naccess_eps) + assert_true( + (result.chain_totals.get("A").unwrap() - 70.6).abs() < naccess_eps, + ) // Chain B: 25.0 - assert_true((result.chain_totals.get("B").unwrap() - 25.0).abs() < naccess_eps) + assert_true( + (result.chain_totals.get("B").unwrap() - 25.0).abs() < naccess_eps, + ) } ///| @@ -578,13 +601,17 @@ test "naccess_count_exposed" { test "naccess_chain_total_asa_a" { let result = @src.naccess_sample() // 45.3 + 20.1 + 5.2 = 70.6 - assert_true((@src.naccess_chain_total_asa(result, "A") - 70.6).abs() < naccess_eps) + assert_true( + (@src.naccess_chain_total_asa(result, "A") - 70.6).abs() < naccess_eps, + ) } ///| test "naccess_chain_total_asa_b" { let result = @src.naccess_sample() - assert_true((@src.naccess_chain_total_asa(result, "B") - 25.0).abs() < naccess_eps) + assert_true( + (@src.naccess_chain_total_asa(result, "B") - 25.0).abs() < naccess_eps, + ) } ///| diff --git a/test/moonbit/naive_bayes_test.mbt b/test/moonbit/naive_bayes_test.mbt index feb02250..dbb826ea 100644 --- a/test/moonbit/naive_bayes_test.mbt +++ b/test/moonbit/naive_bayes_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Bio.NaiveBayes sequence classifier module. - test "classifier_new_defaults" { let clf = @src.NaiveBayesClassifier::new() assert_eq(clf.kmer_size, 3) @@ -9,24 +8,28 @@ test "classifier_new_defaults" { assert_eq(clf.class_labels.length(), 0) } +///| test "classifier_set_kmer_size" { let clf = @src.NaiveBayesClassifier::new().set_kmer_size(5) assert_eq(clf.kmer_size, 5) assert_eq(clf.alpha, 1.0) } +///| test "classifier_set_alpha" { let clf = @src.NaiveBayesClassifier::new().set_alpha(0.5) assert_eq(clf.kmer_size, 3) assert_eq(clf.alpha, 0.5) } +///| test "classifier_setters_chain" { let clf = @src.NaiveBayesClassifier::new().set_kmer_size(4).set_alpha(0.1) assert_eq(clf.kmer_size, 4) assert_eq(clf.alpha, 0.1) } +///| test "extract_kmers_basic" { let kmers = @src.naive_bayes_extract_kmers("ABCDE", 2) assert_eq(kmers.length(), 4) @@ -36,28 +39,33 @@ test "extract_kmers_basic" { assert_eq(kmers.get("DE").unwrap(), 1) } +///| test "extract_kmers_repeated" { let kmers = @src.naive_bayes_extract_kmers("AAAA", 2) assert_eq(kmers.length(), 1) assert_eq(kmers.get("AA").unwrap(), 3) } +///| test "extract_kmers_k_larger_than_length" { let kmers = @src.naive_bayes_extract_kmers("ABC", 5) assert_eq(kmers.length(), 0) } +///| test "extract_kmers_k_equals_length" { let kmers = @src.naive_bayes_extract_kmers("ABC", 3) assert_eq(kmers.length(), 1) assert_eq(kmers.get("ABC").unwrap(), 1) } +///| test "extract_kmers_empty_sequence" { let kmers = @src.naive_bayes_extract_kmers("", 3) assert_eq(kmers.length(), 0) } +///| test "train_basic" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -67,6 +75,7 @@ test "train_basic" { assert_eq(trained.models.length(), 2) } +///| test "predict_at_rich_correct" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -75,6 +84,7 @@ test "predict_at_rich_correct" { assert_eq(pred.0, "AT_rich") } +///| test "predict_gc_rich_correct" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -83,6 +93,7 @@ test "predict_gc_rich_correct" { assert_eq(pred.0, "GC_rich") } +///| test "predict_proba_sums_to_one" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -95,14 +106,18 @@ test "predict_proba_sums_to_one" { assert_true(sum > 0.99 && sum < 1.01) } +///| test "predict_log_probs_returns_all_classes" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() let trained = @src.naive_bayes_train(clf, data.0, data.1) - let log_probs = @src.naive_bayes_predict_log_probs(trained, "ATATATATATATATATATAT") + let log_probs = @src.naive_bayes_predict_log_probs( + trained, "ATATATATATATATATATAT", + ) assert_eq(log_probs.length(), 2) } +///| test "top_k_returns_k_items" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -111,6 +126,7 @@ test "top_k_returns_k_items" { assert_eq(top.length(), 1) } +///| test "top_k_descending_order" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -120,6 +136,7 @@ test "top_k_descending_order" { assert_true(top[0].1 >= top[1].1) } +///| test "accuracy_on_sample_data" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -128,6 +145,7 @@ test "accuracy_on_sample_data" { assert_true(acc > 0.5) } +///| test "unknown_sequence_reasonable_defaults" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -142,18 +160,26 @@ test "unknown_sequence_reasonable_defaults" { assert_true(sum > 0.99 && sum < 1.01) } +///| test "kmer_size_affects_results" { let data = @src.naive_bayes_sample_data() let clf3 = @src.NaiveBayesClassifier::new() let trained3 = @src.naive_bayes_train(clf3, data.0, data.1) let clf2 = @src.NaiveBayesClassifier::new().set_kmer_size(2) let trained2 = @src.naive_bayes_train(clf2, data.0, data.1) - let log_probs3 = @src.naive_bayes_predict_log_probs(trained3, "ATATATATATATATATATAT") - let log_probs2 = @src.naive_bayes_predict_log_probs(trained2, "ATATATATATATATATATAT") + let log_probs3 = @src.naive_bayes_predict_log_probs( + trained3, "ATATATATATATATATATAT", + ) + let log_probs2 = @src.naive_bayes_predict_log_probs( + trained2, "ATATATATATATATATATAT", + ) let diff = (log_probs3[0].1 - log_probs2[0].1).abs() - assert_true(diff > 0.0 || trained3.vocabulary.length() != trained2.vocabulary.length()) + assert_true( + diff > 0.0 || trained3.vocabulary.length() != trained2.vocabulary.length(), + ) } +///| test "class_labels_array_after_training" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -164,12 +190,14 @@ test "class_labels_array_after_training" { assert_true(has_at && has_gc) } +///| test "sample_data_has_8_sequences" { let data = @src.naive_bayes_sample_data() assert_eq(data.0.length(), 8) assert_eq(data.1.length(), 8) } +///| test "sample_data_labels_count" { let data = @src.naive_bayes_sample_data() let mut at_count = 0 @@ -185,6 +213,7 @@ test "sample_data_labels_count" { assert_eq(gc_count, 4) } +///| test "train_empty_sequences" { let clf = @src.NaiveBayesClassifier::new() let seqs : Array[String] = [] @@ -194,6 +223,7 @@ test "train_empty_sequences" { assert_eq(trained.vocabulary.length(), 0) } +///| test "predict_empty_sequence" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() @@ -202,6 +232,7 @@ test "predict_empty_sequence" { assert_true(trained.class_labels.contains(pred.0)) } +///| test "accuracy_empty_data" { let data = @src.naive_bayes_sample_data() let clf = @src.NaiveBayesClassifier::new() diff --git a/test/moonbit/nanostring_test.mbt b/test/moonbit/nanostring_test.mbt index 22b95917..e8962185 100644 --- a/test/moonbit/nanostring_test.mbt +++ b/test/moonbit/nanostring_test.mbt @@ -25,9 +25,7 @@ test "ns_sample_creation" { ///| test "ns_sample_accessors" { let raw = [100, 200, 300, 50, 60, 10, 20] - let sample = @src.NsSample::new( - "S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2], - ) + let sample = @src.NsSample::new("S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2]) let pos = sample.positive_controls() assert_eq(pos.length(), 2) assert_eq(pos[0], 50) @@ -65,9 +63,7 @@ test "ns_nanostring_data_add_sample" { let gene_names = ["G1", "G2", "G3"] let data = @src.NsNanostringData::new(gene_names) let raw = [100, 200, 300, 50, 60, 10, 20] - let sample = @src.NsSample::new( - "S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2], - ) + let sample = @src.NsSample::new("S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2]) data.add_sample(sample) assert_eq(data.n_samples(), 1) assert_eq(data.n_genes(), 3) @@ -127,11 +123,11 @@ test "ns_log2_counts_known" { test "ns_log2_counts_zero_pseudocount" { // log2(0 + 0.5) = log2(0.5) = -1. let result = @src.ns_log2_counts([0.0]) - assert_true((result[0] - (-1.0)).abs() < 0.01) + assert_true((result[0] - -1.0).abs() < 0.01) // Multiple values. let result2 = @src.ns_log2_counts([0.0, 1.0, 7.0]) // log2(0.5) = -1, log2(1.5) ≈ 0.585, log2(7.5) ≈ 2.907 - assert_true((result2[0] - (-1.0)).abs() < 0.01) + assert_true((result2[0] - -1.0).abs() < 0.01) assert_true((result2[1] - 0.585).abs() < 0.01) assert_true((result2[2] - 2.907).abs() < 0.01) } @@ -448,9 +444,7 @@ test "ns_edge_single_sample" { let gene_names = ["G1", "G2", "G3"] let data = @src.NsNanostringData::new(gene_names) let raw = [100, 200, 300, 50, 60, 10, 20] - let sample = @src.NsSample::new( - "S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2], - ) + let sample = @src.NsSample::new("S1", raw, [3, 4], [5, 6], [0, 1], [0, 1, 2]) data.add_sample(sample) assert_eq(data.n_samples(), 1) let results = @src.ns_positive_control_norm(data) @@ -463,9 +457,7 @@ test "ns_edge_single_gene" { let gene_names = ["G1"] let data = @src.NsNanostringData::new(gene_names) let raw = [100, 50, 10] - let sample = @src.NsSample::new( - "S1", raw, [1], [2], [0], [0], - ) + let sample = @src.NsSample::new("S1", raw, [1], [2], [0], [0]) data.add_sample(sample) assert_eq(data.n_genes(), 1) let results = @src.ns_positive_control_norm(data) @@ -477,9 +469,9 @@ test "ns_edge_all_zeros" { let gene_names = ["G1", "G2", "G3"] let data = @src.NsNanostringData::new(gene_names) let raw = [0, 0, 0, 0, 0, 0, 0, 0, 0] - let sample = @src.NsSample::new( - "S1", raw, [3, 4, 5], [6, 7, 8], [0, 1, 2], [0, 1, 2], - ) + let sample = @src.NsSample::new("S1", raw, [3, 4, 5], [6, 7, 8], [0, 1, 2], [ + 0, 1, 2, + ]) data.add_sample(sample) let results = @src.ns_positive_control_norm(data) assert_eq(results.length(), 1) diff --git a/test/moonbit/nib_io_test.mbt b/test/moonbit/nib_io_test.mbt index 391b0e6a..0302e6be 100644 --- a/test/moonbit/nib_io_test.mbt +++ b/test/moonbit/nib_io_test.mbt @@ -298,7 +298,10 @@ test "nib_size" { test "nib_compressed_size" { // 8 bases -> 2 bytes; 9 bases -> 3 bytes; 1 base -> 1 byte; 0 bases -> 0. assert_eq(@src.nib_compressed_size(@src.NibSequence::new("a", "ATGCATGC")), 2) - assert_eq(@src.nib_compressed_size(@src.NibSequence::new("b", "ATGCATGCA")), 3) + assert_eq( + @src.nib_compressed_size(@src.NibSequence::new("b", "ATGCATGCA")), + 3, + ) assert_eq(@src.nib_compressed_size(@src.NibSequence::new("c", "A")), 1) assert_eq(@src.nib_compressed_size(@src.NibSequence::new("d", "")), 0) } @@ -675,7 +678,8 @@ test "edge_case_three_bases" { ///| test "edge_case_hex_round_trip_various_lengths" { - for seq in ["A", "AT", "ATG", "ATGC", "ATGCA", "ATGCAT", "ATGCATG", "ATGCATGC"] { + for + seq in ["A", "AT", "ATG", "ATGC", "ATGCA", "ATGCAT", "ATGCATG", "ATGCATGC"] { let nib = @src.NibSequence::new("s", seq) let hex = @src.nib_to_hex(nib) let nib2 = @src.nib_from_hex("s2", hex, seq.length()) diff --git a/test/moonbit/nmr_test.mbt b/test/moonbit/nmr_test.mbt index 1b12694f..4889777b 100644 --- a/test/moonbit/nmr_test.mbt +++ b/test/moonbit/nmr_test.mbt @@ -326,8 +326,8 @@ test "dihedral_restraint_construction" { assert_eq(d.restraint_id, 1) assert_eq(d.angle_name, "PHI") assert_eq(d.residue, 15) - assert_true((d.lower_bound - (-120.0)).abs() < 0.001) - assert_true((d.upper_bound - (-60.0)).abs() < 0.001) + assert_true((d.lower_bound - -120.0).abs() < 0.001) + assert_true((d.upper_bound - -60.0).abs() < 0.001) } ///| @@ -385,8 +385,8 @@ test "parse_dihedral_restraints_basic" { assert_eq(restraints.length(), 2) assert_eq(restraints[0].angle_name, "PHI") assert_eq(restraints[0].residue, 15) - assert_true((restraints[0].lower_bound - (-120.0)).abs() < 0.001) - assert_true((restraints[0].observed.unwrap_or(0.0) - (-85.0)).abs() < 0.001) + assert_true((restraints[0].lower_bound - -120.0).abs() < 0.001) + assert_true((restraints[0].observed.unwrap_or(0.0) - -85.0).abs() < 0.001) } // ============================================================================ diff --git a/test/moonbit/nnsvg_test.mbt b/test/moonbit/nnsvg_test.mbt new file mode 100644 index 00000000..d6155bda --- /dev/null +++ b/test/moonbit/nnsvg_test.mbt @@ -0,0 +1,968 @@ +// Tests for the Bioconductor nnSVG-inspired NNGP spatial model. + +///| +fn nnsvg_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "nnSVG value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn nnsvg_test_config() -> @src.NnsvgConfig { + @src.NnsvgConfig::create( + n_neighbors=3, + length_scale_min=0.03, + length_scale_max=1.5, + length_scale_grid=5, + proportion_grid=5, + refinement_steps=1, + minimum_variance=1.0e-9, + fdr_threshold=0.1, + ) catch { + _ => abort("nnSVG test configuration should be valid") + } +} + +///| +fn nnsvg_test_result() -> @src.NnsvgResult { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + @src.nnsvg_with_config( + expression, + coordinates, + nnsvg_test_config(), + gene_names~, + ) catch { + _ => abort("nnSVG example fit should succeed") + } +} + +///| +fn nnsvg_test_spatial_experiment() -> @src.SpatialExperiment { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "logcounts", expression)) + let counts : Array[Array[Double]] = [] + for row in expression { + let count_row : Array[Double] = [] + for value in row { + count_row.push((value * 4.0).round()) + } + counts.push(count_row) + } + ignore(@src.se_add_assay(experiment, "counts", counts)) + for gene_name in gene_names { + ignore(@src.se_add_row(experiment, Map([("gene_name", gene_name)]))) + } + for spot in 0.. abort("custom nnSVG configuration should be valid") + } + assert_eq(config.n_neighbors, 4) + assert_true(config.ordering is @src.NnsvgOrdering::SumCoordinates) + assert_eq(config.length_scale_grid, 7) + assert_eq(config.proportion_grid, 8) + assert_eq(config.refinement_steps, 3) + assert_eq(config.minimum_variance, 1.0e-8) +} + +///| +test "nnSVG: ordering helpers expose both ordering modes" { + assert_true(@src.nnsvg_ammd_ordering() is @src.NnsvgOrdering::Ammd) + assert_true( + @src.nnsvg_sum_coordinates_ordering() is @src.NnsvgOrdering::SumCoordinates, + ) +} + +///| +test "nnSVG: configuration rejects non-positive neighbor count" { + let failed = try { + ignore(@src.NnsvgConfig::create(n_neighbors=0)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: configuration rejects non-positive minimum length scale" { + let failed = try { + ignore(@src.NnsvgConfig::create(length_scale_min=0.0)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: configuration rejects reversed length-scale bounds" { + let failed = try { + ignore(@src.NnsvgConfig::create(length_scale_min=2.0, length_scale_max=1.0)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: configuration rejects undersized parameter grids" { + let length_grid = try { + ignore(@src.NnsvgConfig::create(length_scale_grid=1)) + false + } catch { + NnsvgError(_) => true + } + let proportion_grid = try { + ignore(@src.NnsvgConfig::create(proportion_grid=1)) + false + } catch { + NnsvgError(_) => true + } + assert_true(length_grid) + assert_true(proportion_grid) +} + +///| +test "nnSVG: configuration rejects negative refinement count" { + let failed = try { + ignore(@src.NnsvgConfig::create(refinement_steps=-1)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: configuration rejects invalid variance floor" { + let failed = try { + ignore(@src.NnsvgConfig::create(minimum_variance=0.0)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: configuration rejects invalid FDR threshold" { + let zero = try { + ignore(@src.NnsvgConfig::create(fdr_threshold=0.0)) + false + } catch { + NnsvgError(_) => true + } + let large = try { + ignore(@src.NnsvgConfig::create(fdr_threshold=1.1)) + false + } catch { + NnsvgError(_) => true + } + assert_true(zero) + assert_true(large) +} + +///| +test "nnSVG: example data uses gene by spot orientation" { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + assert_eq(expression.length(), 4) + assert_eq(expression[0].length(), 12) + assert_eq(coordinates.length(), 12) + assert_eq(coordinates[0].length(), 2) + assert_eq(gene_names.length(), 4) +} + +///| +test "nnSVG: neighbor graph scales coordinates by global range" { + let (_, coordinates, _) = @src.nnsvg_example_data() + let graph = @src.nnsvg_build_neighbors(coordinates, n_neighbors=3) catch { + _ => abort("neighbor graph should build") + } + nnsvg_test_close(graph.range_scale, 3.0, 1.0e-12) + nnsvg_test_close(graph.coordinates[3][0], 1.0, 1.0e-12) + nnsvg_test_close(graph.coordinates[4][1], 1.0 / 3.0, 1.0e-12) +} + +///| +test "nnSVG: sum-coordinate ordering is deterministic" { + let (_, coordinates, _) = @src.nnsvg_example_data() + let graph = @src.nnsvg_build_neighbors( + coordinates, + n_neighbors=2, + ordering=@src.nnsvg_sum_coordinates_ordering(), + ) catch { + _ => abort("sum-coordinate graph should build") + } + assert_eq(graph.order[0], 0) + assert_eq(graph.order[1], 1) + assert_eq(graph.order[2], 4) + assert_eq(graph.order.length(), coordinates.length()) +} + +///| +test "nnSVG: AMMD ordering is a full permutation" { + let (_, coordinates, _) = @src.nnsvg_example_data() + let graph = @src.nnsvg_build_neighbors(coordinates, n_neighbors=3) catch { + _ => abort("AMMD graph should build") + } + let seen = Array::make(coordinates.length(), false) + for index in graph.order { + assert_true(index >= 0 && index < coordinates.length()) + assert_false(seen[index]) + seen[index] = true + } + for value in seen { + assert_true(value) + } +} + +///| +test "nnSVG: AMMD ordering is reproducible" { + let (_, coordinates, _) = @src.nnsvg_example_data() + let first = @src.nnsvg_build_neighbors(coordinates, n_neighbors=3) catch { + _ => abort("first graph should build") + } + let second = @src.nnsvg_build_neighbors(coordinates, n_neighbors=3) catch { + _ => abort("second graph should build") + } + assert_eq(first.order, second.order) + assert_eq(first.neighbors, second.neighbors) +} + +///| +test "nnSVG: inverse ordering maps originals to processing positions" { + let (_, coordinates, _) = @src.nnsvg_example_data() + let graph = @src.nnsvg_build_neighbors(coordinates) catch { + _ => abort("graph should build") + } + for position in 0.. abort("graph should build") + } + for original in 0.. abort("graph should build") + } + assert_eq(graph.n_neighbors, coordinates.length() - 1) + for original in 0.. abort("graph should build") + } + for original in 0..= graph.distances[original][index - 1], + ) + } + } +} + +///| +test "nnSVG: duplicate coordinates remain valid with a nugget" { + let coordinates = [[0.0, 0.0], [0.0, 0.0], [1.0, 0.0], [1.0, 1.0]] + let graph = @src.nnsvg_build_neighbors(coordinates, n_neighbors=2) catch { + _ => abort("duplicate-coordinate graph should build") + } + assert_eq(graph.order.length(), 4) + assert_eq(graph.neighbors.length(), 4) +} + +///| +test "nnSVG: neighbor graph rejects too few spots" { + let failed = try { + ignore(@src.nnsvg_build_neighbors([[0.0, 0.0], [1.0, 1.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: neighbor graph rejects one-dimensional coordinates" { + let failed = try { + ignore(@src.nnsvg_build_neighbors([[0.0], [1.0], [2.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: neighbor graph rejects ragged coordinates" { + let failed = try { + ignore(@src.nnsvg_build_neighbors([[0.0, 0.0], [1.0], [2.0, 0.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: neighbor graph rejects non-finite coordinates" { + let failed = try { + ignore(@src.nnsvg_build_neighbors([[0.0, 0.0], [1.0e301, 1.0], [2.0, 0.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: neighbor graph rejects invariant coordinates" { + let failed = try { + ignore(@src.nnsvg_build_neighbors([[1.0, 1.0], [1.0, 1.0], [1.0, 1.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit dimensions and names match input" { + let result = nnsvg_test_result() + assert_eq(result.n_genes(), 4) + assert_eq(result.n_spots(), 12) + assert_eq(result.gene_names[0], "x_y_gradient") + assert_eq(result.genes[3].gene_name, "constant") + assert_eq(result.design_columns, 1) +} + +///| +test "nnSVG: default function generates gene names" { + let (expression, coordinates, _) = @src.nnsvg_example_data() + let result = @src.nnsvg(expression, coordinates, n_neighbors=3) catch { + _ => abort("default nnSVG fit should succeed") + } + assert_eq(result.gene_names[0], "gene_1") + assert_eq(result.gene_names[3], "gene_4") +} + +///| +test "nnSVG: custom design coefficients are retained" { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + let design : Array[Array[Double]] = [] + for coordinate in coordinates { + design.push([1.0, coordinate[0]]) + } + let result = @src.nnsvg_with_config( + expression, + coordinates, + nnsvg_test_config(), + design~, + gene_names~, + ) catch { + _ => abort("covariate-adjusted nnSVG fit should succeed") + } + assert_eq(result.design_columns, 2) + for gene in result.genes { + assert_eq(gene.beta.length(), 2) + } +} + +///| +test "nnSVG: likelihood-ratio identity holds" { + let result = nnsvg_test_result() + for gene in result.genes { + nnsvg_test_close( + gene.likelihood_ratio, + 2.0 * (gene.log_likelihood - gene.linear_log_likelihood), + 1.0e-8, + ) + } +} + +///| +test "nnSVG: p-values use the two-degree chi-square tail" { + let result = nnsvg_test_result() + for gene in result.genes { + nnsvg_test_close( + gene.p_value, + @math.exp(-0.5 * gene.likelihood_ratio), + 1.0e-12, + ) + } +} + +///| +test "nnSVG: variance components reconstruct total fitted variance" { + let result = nnsvg_test_result() + for gene in result.genes { + let total = gene.sigma_sq + gene.tau_sq + assert_true(total > 0.0) + nnsvg_test_close( + gene.proportion_spatial_variance, + gene.sigma_sq / total, + 1.0e-10, + ) + } +} + +///| +test "nnSVG: positive spatial fits report reciprocal phi" { + let result = nnsvg_test_result() + for gene in result.genes { + if gene.length_scale > 0.0 { + nnsvg_test_close(gene.phi, 1.0 / gene.length_scale, 1.0e-10) + } else { + assert_eq(gene.phi, 0.0) + } + } +} + +///| +test "nnSVG: adjusted p-values are bounded and not below raw values" { + let result = nnsvg_test_result() + for gene in result.genes { + assert_true(gene.p_value >= 0.0 && gene.p_value <= 1.0) + assert_true( + gene.adjusted_p_value >= gene.p_value - 1.0e-12 && + gene.adjusted_p_value <= 1.0, + ) + } +} + +///| +test "nnSVG: constant gene falls back to non-spatial model" { + let result = nnsvg_test_result() + let constant = result.genes[3] + nnsvg_test_close(constant.likelihood_ratio, 0.0, 1.0e-12) + nnsvg_test_close(constant.p_value, 1.0, 1.0e-12) + nnsvg_test_close(constant.sigma_sq, 0.0, 1.0e-12) + assert_eq(constant.length_scale, 0.0) +} + +///| +test "nnSVG: spatial gradients outrank constant expression" { + let result = nnsvg_test_result() + assert_true(result.genes[0].rank < result.genes[3].rank) + assert_true(result.genes[1].rank < result.genes[3].rank) + assert_true( + result.genes[0].likelihood_ratio > result.genes[3].likelihood_ratio, + ) +} + +///| +test "nnSVG: gene lookup finds names and rejects unknown names" { + let result = nnsvg_test_result() + let found = match result.gene("checkerboard") { + Some(value) => value + None => abort("known gene should be found") + } + assert_eq(found.gene_name, "checkerboard") + assert_true(result.gene("missing") is None) +} + +///| +test "nnSVG: top results are ordered by rank" { + let top = nnsvg_test_result().top(3) + assert_eq(top.length(), 3) + assert_true(top[0].rank <= top[1].rank) + assert_true(top[1].rank <= top[2].rank) +} + +///| +test "nnSVG: top handles zero and oversized requests" { + let result = nnsvg_test_result() + assert_eq(result.top(0).length(), 0) + assert_eq(result.top(100).length(), result.n_genes()) +} + +///| +test "nnSVG: significant supports an explicit threshold" { + let result = nnsvg_test_result() + let all = result.significant(threshold=1.0) + let none = result.significant(threshold=0.0) + assert_eq(all.length(), result.n_genes()) + assert_eq(none.length(), 0) +} + +///| +test "nnSVG: summary reports dimensions and neighbor count" { + let summary = nnsvg_test_result().summary() + assert_true(summary.contains("4 genes x 12 spots")) + assert_true(summary.contains("neighbors=3")) + assert_true(summary.contains("FDR=0.1")) +} + +///| +test "nnSVG: fit rejects empty expression" { + let failed = try { + ignore(@src.nnsvg([], [[0.0, 0.0], [1.0, 0.0], [2.0, 0.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects too few spots" { + let failed = try { + ignore(@src.nnsvg([[1.0, 2.0]], [[0.0, 0.0], [1.0, 0.0]])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects ragged expression" { + let failed = try { + ignore( + @src.nnsvg([[1.0, 2.0, 3.0], [1.0, 2.0]], [ + [0.0, 0.0], + [1.0, 0.0], + [2.0, 0.0], + ]), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects non-finite expression" { + let failed = try { + ignore( + @src.nnsvg([[1.0, 1.0e301, 3.0]], [[0.0, 0.0], [1.0, 0.0], [2.0, 0.0]]), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects coordinate count mismatch" { + let (expression, coordinates, _) = @src.nnsvg_example_data() + let shortened : Array[Array[Double]] = [] + for index in 0..<(coordinates.length() - 1) { + shortened.push(coordinates[index]) + } + let failed = try { + ignore(@src.nnsvg(expression, shortened)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects gene name count mismatch" { + let (expression, coordinates, _) = @src.nnsvg_example_data() + let failed = try { + ignore(@src.nnsvg(expression, coordinates, gene_names=["only_one"])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects design row mismatch" { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + let failed = try { + ignore( + @src.nnsvg(expression, coordinates, design=[[1.0], [1.0]], gene_names~), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: fit rejects rank-deficient design" { + let (expression, coordinates, gene_names) = @src.nnsvg_example_data() + let design : Array[Array[Double]] = [] + for _ in coordinates { + design.push([1.0, 1.0]) + } + let failed = try { + ignore(@src.nnsvg(expression, coordinates, design~, gene_names~)) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: gene filtering applies expression threshold" { + let counts = [ + [4.0, 4.0, 0.0, 0.0], + [2.0, 2.0, 2.0, 2.0], + [5.0, 5.0, 5.0, 5.0], + ] + let filter = @src.nnsvg_filter_genes( + counts, + ["A", "B", "C"], + minimum_count=3.0, + minimum_spot_percentage=50.0, + filter_mitochondrial=false, + ) catch { + _ => abort("gene filtering should succeed") + } + assert_eq(filter.kept_indices, [0, 2]) + assert_eq(filter.removed_low_expression, [1]) + assert_eq(filter.counts.length(), 2) +} + +///| +test "nnSVG: gene filtering removes mitochondrial prefixes" { + let filter = @src.nnsvg_filter_genes( + [[4.0, 4.0], [4.0, 4.0], [4.0, 4.0]], + ["MT-ND1", "mt-Co1", "ACTB"], + minimum_spot_percentage=100.0, + ) catch { + _ => abort("mitochondrial filtering should succeed") + } + assert_eq(filter.kept_indices, [2]) + assert_eq(filter.removed_mitochondrial, [0, 1]) +} + +///| +test "nnSVG: mitochondrial filtering can be disabled" { + let filter = @src.nnsvg_filter_genes( + [[4.0, 4.0], [4.0, 4.0]], + ["MT-ND1", "ACTB"], + minimum_spot_percentage=100.0, + filter_mitochondrial=false, + ) catch { + _ => abort("disabled mitochondrial filtering should succeed") + } + assert_eq(filter.kept_indices, [0, 1]) + assert_eq(filter.removed_mitochondrial.length(), 0) +} + +///| +test "nnSVG: default filter percentage requires one detected spot" { + let filter = @src.nnsvg_filter_genes( + [[0.0, 0.0, 0.0, 3.0], [0.0, 0.0, 0.0, 0.0]], + ["detected", "absent"], + filter_mitochondrial=false, + ) catch { + _ => abort("default filtering should succeed") + } + assert_eq(filter.kept_indices, [0]) + assert_eq(filter.removed_low_expression, [1]) +} + +///| +test "nnSVG: gene filtering rejects name mismatch" { + let failed = try { + ignore(@src.nnsvg_filter_genes([[1.0, 2.0]], [])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: gene filtering rejects ragged counts" { + let failed = try { + ignore(@src.nnsvg_filter_genes([[1.0, 2.0], [1.0]], ["A", "B"])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: gene filtering rejects negative counts" { + let failed = try { + ignore(@src.nnsvg_filter_genes([[-1.0, 2.0]], ["A"])) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: gene filtering rejects invalid thresholds" { + let count = try { + ignore(@src.nnsvg_filter_genes([[1.0, 2.0]], ["A"], minimum_count=-1.0)) + false + } catch { + NnsvgError(_) => true + } + let percentage = try { + ignore( + @src.nnsvg_filter_genes( + [[1.0, 2.0]], + ["A"], + minimum_spot_percentage=101.0, + ), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(count) + assert_true(percentage) +} + +///| +test "nnSVG: SpatialExperiment integration preserves assays" { + let experiment = nnsvg_test_spatial_experiment() + let output = @src.nnsvg_spatial_experiment( + experiment, + config=nnsvg_test_config(), + ) catch { + _ => abort("SpatialExperiment integration should succeed") + } + assert_eq(output.result.n_genes(), 4) + assert_true(output.experiment.assay.contains("logcounts")) + assert_true(output.experiment.assay.contains("counts")) +} + +///| +test "nnSVG: SpatialExperiment integration writes rowData statistics" { + let output = @src.nnsvg_spatial_experiment( + nnsvg_test_spatial_experiment(), + config=nnsvg_test_config(), + ) catch { + _ => abort("SpatialExperiment integration should succeed") + } + let row = output.experiment.row_data[0] + assert_true(row.contains("nnsvg_sigma_sq")) + assert_true(row.contains("nnsvg_tau_sq")) + assert_true(row.contains("nnsvg_prop_sv")) + assert_true(row.contains("nnsvg_LR_stat")) + assert_true(row.contains("nnsvg_pval")) + assert_true(row.contains("nnsvg_padj")) +} + +///| +test "nnSVG: SpatialExperiment integration writes metadata" { + let output = @src.nnsvg_spatial_experiment( + nnsvg_test_spatial_experiment(), + config=nnsvg_test_config(), + ) catch { + _ => abort("SpatialExperiment integration should succeed") + } + assert_eq(output.experiment.metadata["nnsvg_assay"], "logcounts") + assert_eq(output.experiment.metadata["nnsvg_neighbors"], "3") + assert_eq(output.experiment.metadata["nnsvg_genes"], "4") +} + +///| +test "nnSVG: SpatialExperiment integration leaves input rowData unchanged" { + let experiment = nnsvg_test_spatial_experiment() + ignore( + @src.nnsvg_spatial_experiment(experiment, config=nnsvg_test_config()) catch { + _ => abort("SpatialExperiment integration should succeed") + }, + ) + assert_false(experiment.row_data[0].contains("nnsvg_pval")) +} + +///| +test "nnSVG: SpatialExperiment integration rejects missing assay" { + let failed = try { + ignore( + @src.nnsvg_spatial_experiment( + nnsvg_test_spatial_experiment(), + assay_name="missing", + config=nnsvg_test_config(), + ), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: SpatialExperiment integration rejects coordinate mismatch" { + let experiment = nnsvg_test_spatial_experiment() + ignore(experiment.spatial_coords.pop()) + let failed = try { + ignore( + @src.nnsvg_spatial_experiment(experiment, config=nnsvg_test_config()), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: SpatialExperiment filter subsets every assay" { + let output = @src.nnsvg_filter_spatial_experiment( + nnsvg_test_spatial_experiment(), + minimum_count=5.0, + minimum_spot_percentage=50.0, + filter_mitochondrial=false, + ) catch { + _ => abort("SpatialExperiment filtering should succeed") + } + assert_eq( + output.experiment.assay["counts"].length(), + output.filter.kept_indices.length(), + ) + assert_eq( + output.experiment.assay["logcounts"].length(), + output.filter.kept_indices.length(), + ) + assert_eq( + output.experiment.row_data.length(), + output.filter.kept_indices.length(), + ) +} + +///| +test "nnSVG: SpatialExperiment filter preserves spot data" { + let experiment = nnsvg_test_spatial_experiment() + let output = @src.nnsvg_filter_spatial_experiment( + experiment, + minimum_count=5.0, + minimum_spot_percentage=50.0, + filter_mitochondrial=false, + ) catch { + _ => abort("SpatialExperiment filtering should succeed") + } + assert_eq( + output.experiment.spatial_coords.length(), + experiment.spatial_coords.length(), + ) + assert_eq(output.experiment.col_data.length(), experiment.col_data.length()) +} + +///| +test "nnSVG: SpatialExperiment filter records metadata" { + let output = @src.nnsvg_filter_spatial_experiment( + nnsvg_test_spatial_experiment(), + minimum_count=5.0, + minimum_spot_percentage=50.0, + filter_mitochondrial=false, + ) catch { + _ => abort("SpatialExperiment filtering should succeed") + } + assert_true(output.experiment.metadata.contains("nnsvg_filter_kept")) + assert_true(output.experiment.metadata.contains("nnsvg_filter_removed")) +} + +///| +test "nnSVG: SpatialExperiment filter rejects missing count assay" { + let failed = try { + ignore( + @src.nnsvg_filter_spatial_experiment( + nnsvg_test_spatial_experiment(), + count_assay="missing", + ), + ) + false + } catch { + NnsvgError(_) => true + } + assert_true(failed) +} + +///| +test "nnSVG: three-dimensional SpatialExperiment coordinates are supported" { + let experiment = nnsvg_test_spatial_experiment() + experiment.spatial_coords[1] = @src.SpatialCoord::new(1.0, 0.0, 1.0) + let output = @src.nnsvg_spatial_experiment( + experiment, + config=nnsvg_test_config(), + ) catch { + _ => abort("3D SpatialExperiment integration should succeed") + } + assert_eq(output.result.graph.coordinates[0].length(), 3) +} diff --git a/test/moonbit/noiseq_test.mbt b/test/moonbit/noiseq_test.mbt index e70301fa..aa1d896e 100644 --- a/test/moonbit/noiseq_test.mbt +++ b/test/moonbit/noiseq_test.mbt @@ -1,12 +1,12 @@ ///| /// Test file for NOISeq module. - test "noiseq_sample_creation" { let s = @src.NOISeqSample::new("sample1", "control") assert_eq(s.sample_id, "sample1") assert_eq(s.condition, "control") } +///| test "noiseq_sample_set_get_count" { let s = @src.NOISeqSample::new("sample1", "control") s.set_count("Gene1", 100.0) @@ -16,6 +16,7 @@ test "noiseq_sample_set_get_count" { assert_eq(s.get_count("Gene3"), 0.0) // Non-existent } +///| test "noiseq_sample_library_size" { let s = @src.NOISeqSample::new("sample1", "control") s.set_count("Gene1", 100.0) @@ -24,6 +25,7 @@ test "noiseq_sample_library_size" { assert_eq(s.library_size(), 350.0) } +///| test "noiseq_sample_n_expressed" { let s = @src.NOISeqSample::new("sample1", "control") s.set_count("Gene1", 100.0) @@ -32,6 +34,7 @@ test "noiseq_sample_n_expressed" { assert_eq(s.n_expressed(), 2) } +///| test "noiseq_normalize_tmm" { let samples : Array[@src.NOISeqSample] = Array::new() let s1 = @src.NOISeqSample::new("s1", "ctrl") @@ -50,6 +53,7 @@ test "noiseq_normalize_tmm" { assert_true((norm[1].get_count("Gene1") - 100.0).abs() < 0.01) } +///| test "noiseq_normalize_rpkm" { let samples : Array[@src.NOISeqSample] = Array::new() let s1 = @src.NOISeqSample::new("s1", "ctrl") @@ -61,6 +65,7 @@ test "noiseq_normalize_rpkm" { assert_true((norm[0].get_count("Gene1") - 1000000.0).abs() < 0.01) } +///| test "noiseq_normalize_none" { let samples : Array[@src.NOISeqSample] = Array::new() let s1 = @src.NOISeqSample::new("s1", "ctrl") @@ -71,6 +76,7 @@ test "noiseq_normalize_none" { assert_eq(norm[0].get_count("Gene1"), 100.0) } +///| test "noiseq_run_basic" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt) @@ -78,6 +84,7 @@ test "noiseq_run_basic" { assert_eq(results.n_genes, 10) } +///| test "noiseq_run_significant" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt, prob_threshold=0.3) @@ -85,6 +92,7 @@ test "noiseq_run_significant" { assert_true(sig.length() > 0) } +///| test "noiseq_run_top_genes" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt) @@ -94,6 +102,7 @@ test "noiseq_run_top_genes" { assert_true(top[0].prob >= top[1].prob) } +///| test "noiseq_run_up_down_regulated" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt, prob_threshold=0.3) @@ -103,6 +112,7 @@ test "noiseq_run_up_down_regulated" { assert_true(up.length() + down.length() > 0) } +///| test "noiseq_run_summary" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt) @@ -111,6 +121,7 @@ test "noiseq_run_summary" { assert_true(s.contains("Genes tested")) } +///| test "noiseq_qc" { let (ctrl, trt) = @src.noiseq_sample_data() let all : Array[@src.NOISeqSample] = Array::new() @@ -128,6 +139,7 @@ test "noiseq_qc" { assert_true(qc.biotype_counts.size() > 0) } +///| test "noiseq_norm_to_string" { assert_eq(@src.noiseq_norm_rpkm().to_string(), "RPKM") assert_eq(@src.noiseq_norm_tmm().to_string(), "TMM") @@ -135,11 +147,13 @@ test "noiseq_norm_to_string" { assert_eq(@src.noiseq_norm_none().to_string(), "none") } +///| test "noiseq_method_to_string" { assert_eq(@src.noiseq_method_bio().to_string(), "NOISeqBio") assert_eq(@src.noiseq_method_sim().to_string(), "NOISeqSim") } +///| test "noiseq_sample_data" { let (ctrl, trt) = @src.noiseq_sample_data() assert_eq(ctrl.length(), 3) @@ -148,6 +162,7 @@ test "noiseq_sample_data" { assert_eq(trt[0].condition, "treatment") } +///| test "noiseq_result_ranking" { let (ctrl, trt) = @src.noiseq_sample_data() let results = @src.noiseq_run(ctrl, trt) diff --git a/test/moonbit/nucle_r_test.mbt b/test/moonbit/nucle_r_test.mbt index 310f46eb..32bb1f1e 100644 --- a/test/moonbit/nucle_r_test.mbt +++ b/test/moonbit/nucle_r_test.mbt @@ -22,7 +22,10 @@ test "nuc_create_example_result" { assert_true(result.nuc_n_count() >= 0) assert_true(result.nuc_mean_spacing() >= 0.0) assert_true(result.nuc_mean_occupancy() >= 0.0) - assert_true(result.nuc_frac_well_positioned() >= 0.0 && result.nuc_frac_well_positioned() <= 1.0) + assert_true( + result.nuc_frac_well_positioned() >= 0.0 && + result.nuc_frac_well_positioned() <= 1.0, + ) } ///| @@ -103,7 +106,7 @@ test "nuc_compare_positioning_identical" { ///| test "nuc_position_methods" { let pos = @src.NucPosition::new( - "chr1", 100, 300, 200.0, 200.0, 0.8, 0.01, true + "chr1", 100, 300, 200.0, 200.0, 0.8, 0.01, true, ) assert_eq(pos.nuc_chrom(), "chr1") assert_eq(pos.nuc_start(), 100) @@ -115,7 +118,7 @@ test "nuc_position_methods" { assert_true(pos.nuc_is_well_positioned()) let pos2 = @src.NucPosition::new( - "chr2", 500, 800, 650.0, 300.0, 0.5, 0.003, false + "chr2", 500, 800, 650.0, 300.0, 0.5, 0.003, false, ) assert_eq(pos2.nuc_chrom(), "chr2") assert_true(!pos2.nuc_is_well_positioned()) @@ -186,6 +189,8 @@ test "nuc_dynamic_result_methods" { assert_true(dynamic.nuc_n_shared() >= 0) assert_eq(dynamic.nuc_n_gained(), 0) assert_eq(dynamic.nuc_n_lost(), 0) - assert_true(dynamic.nuc_frac_changed() >= 0.0 && dynamic.nuc_frac_changed() <= 1.0) + assert_true( + dynamic.nuc_frac_changed() >= 0.0 && dynamic.nuc_frac_changed() <= 1.0, + ) assert_eq(dynamic.nuc_direction(), 0) -} \ No newline at end of file +} diff --git a/test/moonbit/open_cyto_test.mbt b/test/moonbit/open_cyto_test.mbt index db492dbd..feff9f68 100644 --- a/test/moonbit/open_cyto_test.mbt +++ b/test/moonbit/open_cyto_test.mbt @@ -223,7 +223,12 @@ test "oc_gate_new1d_fields" { ///| test "oc_gate_new2d_fields" { let g = @src.OcGate::new2d( - dim1="FSC", dim2="SSC", min1=1.0, max1=2.0, min2=3.0, max2=4.0, + dim1="FSC", + dim2="SSC", + min1=1.0, + max1=2.0, + min2=3.0, + max2=4.0, ) assert_eq(g.dim1, "FSC") assert_eq(g.dim2.unwrap(), "SSC") @@ -256,7 +261,12 @@ test "oc_event_in_gate_1d_boundary" { ///| test "oc_event_in_gate_2d" { let g = @src.OcGate::new2d( - dim1="FSC", dim2="SSC", min1=1.0, max1=3.0, min2=10.0, max2=20.0, + dim1="FSC", + dim2="SSC", + min1=1.0, + max1=3.0, + min2=10.0, + max2=20.0, ) // Both dims in range. assert_true(@src.oc_event_in_gate(g, 2.0, 15.0)) @@ -525,13 +535,13 @@ test "oc_t_pdf_zero_variance" { ///| test "oc_lgamma_one" { // lgamma(1) = log(gamma(1)) = log(1) = 0 - assert_true((@src.oc_lgamma(1.0)).abs() < 0.001) + assert_true(@src.oc_lgamma(1.0).abs() < 0.001) } ///| test "oc_lgamma_two" { // lgamma(2) = log(gamma(2)) = log(1) = 0 - assert_true((@src.oc_lgamma(2.0)).abs() < 0.001) + assert_true(@src.oc_lgamma(2.0).abs() < 0.001) } ///| @@ -651,7 +661,7 @@ test "oc_gating_rule_construction" { child="tcells", method="quantileGate", dims=["CD3"], - args=args, + args~, ) assert_eq(r.parent, "root") assert_eq(r.child, "tcells") @@ -668,22 +678,9 @@ test "oc_gating_rule_construction" { ///| test "oc_gate_flow_set_single_rule" { // One sample, one channel (FSC), bimodal FSC values. - let fs = @src.OcFlowSet::new( - sample_names=["S1"], - channel_names=["FSC"], - data=[ - [ - [1.0], - [1.1], - [0.9], - [1.2], - [5.0], - [5.1], - [4.9], - [5.2], - ], - ], - ) + let fs = @src.OcFlowSet::new(sample_names=["S1"], channel_names=["FSC"], data=[ + [[1.0], [1.1], [0.9], [1.2], [5.0], [5.1], [4.9], [5.2]], + ]) let args : Map[String, Double] = Map::new() args.set("bandwidth", 0.5) let rule = @src.OcGatingRule::new( @@ -691,7 +688,7 @@ test "oc_gate_flow_set_single_rule" { child="cells", method="mindensity", dims=["FSC"], - args=args, + args~, ) let results = @src.oc_gate_flow_set(fs, [rule]) assert_eq(results.length(), 1) @@ -761,16 +758,14 @@ test "oc_gate_flow_set_chain" { test "oc_population_stats_basic" { // Build a gating result by hand: 8 parent events, 4 child events. let gate = @src.OcGate::new1d(dim="FSC", min=3.0, max=1.0e30) - let indices = [ - false, false, false, false, true, true, true, true, - ] + let indices = [false, false, false, false, true, true, true, true] let result = @src.OcGatingResult::new( population="cells", sample="S1", - gate=gate, + gate~, parent_events=8, child_events=4, - indices=indices, + indices~, ) let stats = @src.oc_population_stats([result], 8) assert_eq(stats.length(), 1) @@ -789,7 +784,7 @@ test "oc_population_stats_zero_parent" { let result = @src.OcGatingResult::new( population="dead", sample="S1", - gate=gate, + gate~, parent_events=0, child_events=0, indices=[], @@ -809,7 +804,7 @@ test "oc_gating_summary_string" { let result = @src.OcGatingResult::new( population="cells", sample="S1", - gate=gate, + gate~, parent_events=8, child_events=4, indices=[false, false, false, false, true, true, true, true], @@ -829,7 +824,7 @@ test "oc_gating_summary_string" { ///| test "oc_ln_one" { - assert_true((@src.oc_ln(1.0)).abs() < 0.001) + assert_true(@src.oc_ln(1.0).abs() < 0.001) } ///| diff --git a/test/moonbit/pairaligner_test.mbt b/test/moonbit/pairaligner_test.mbt index bb789da8..5156294a 100644 --- a/test/moonbit/pairaligner_test.mbt +++ b/test/moonbit/pairaligner_test.mbt @@ -1,30 +1,33 @@ ///| /// Test file for pairaligner module. - test "pairaligner_alignment_mode_global" { let mode = @src.pairaligner_global() let config = @src.PairwiseAlignerConfig::default_dna().set_mode(mode) assert_eq(config.mode, mode) } +///| test "pairaligner_alignment_mode_local" { let mode = @src.pairaligner_local() let config = @src.PairwiseAlignerConfig::default_dna().set_mode(mode) assert_eq(config.mode, mode) } +///| test "pairaligner_substitution_matrix_no_matrix" { let mat = @src.pairaligner_no_matrix() let config = @src.PairwiseAlignerConfig::default_dna().set_submatrix(mat) assert_eq(config.submatrix, mat) } +///| test "pairaligner_substitution_matrix_blosum62" { let mat = @src.pairaligner_blosum62() let config = @src.PairwiseAlignerConfig::default_protein().set_submatrix(mat) assert_eq(config.submatrix, mat) } +///| test "pairaligner_default_dna_config" { let config = @src.PairwiseAlignerConfig::default_dna() assert_eq(config.match_score, 1.0) @@ -36,6 +39,7 @@ test "pairaligner_default_dna_config" { assert_eq(config.query_gap_open, -1.0) } +///| test "pairaligner_default_protein_config" { let config = @src.PairwiseAlignerConfig::default_protein() assert_eq(config.gap_open, -10.0) @@ -44,6 +48,7 @@ test "pairaligner_default_protein_config" { assert_eq(config.submatrix, @src.pairaligner_blosum62()) } +///| test "pairaligner_config_setters" { let base = @src.PairwiseAlignerConfig::default_dna() let c1 = base.set_match_score(5.0) @@ -63,11 +68,12 @@ test "pairaligner_config_setters" { assert_eq(base.mismatch_score, -1.0) } +///| test "pairaligner_align_global_dna" { let target = "ACGT" let query = "ACGT" let config = @src.PairwiseAlignerConfig::default_dna() - let aln = @src.pairaligner_align(target, query, config=config) + let aln = @src.pairaligner_align(target, query, config~) assert_eq(aln.aligned1(), "ACGT") assert_eq(aln.aligned2(), "ACGT") assert_eq(aln.identities(), 4) @@ -77,44 +83,50 @@ test "pairaligner_align_global_dna" { assert_eq(ml, "||||") } +///| test "pairaligner_align_global_dna_with_mismatch" { let target = "ACGT" let query = "AGGT" let config = @src.PairwiseAlignerConfig::default_dna() - let aln = @src.pairaligner_align(target, query, config=config) + let aln = @src.pairaligner_align(target, query, config~) assert_eq(aln.alignment_length(), 4) let ml = aln.match_line assert_eq(ml[1:2], ".") assert_eq(aln.score(), 2.0) } +///| test "pairaligner_align_local_dna" { let target = "XXXXACGTXXXX" let query = "ACGT" - let config = @src.PairwiseAlignerConfig::default_dna() - .set_mode(@src.pairaligner_local()) - let aln = @src.pairaligner_align(target, query, config=config) + let config = @src.PairwiseAlignerConfig::default_dna().set_mode( + @src.pairaligner_local(), + ) + let aln = @src.pairaligner_align(target, query, config~) assert_eq(aln.aligned1(), "ACGT") assert_eq(aln.aligned2(), "ACGT") assert_eq(aln.identities(), 4) assert_eq(aln.score(), 4.0) } +///| test "pairaligner_align_protein_blosum62" { let (target, query) = @src.pairaligner_sample_data() let config = @src.PairwiseAlignerConfig::default_protein() - let aln = @src.pairaligner_align(target, query, config=config) + let aln = @src.pairaligner_align(target, query, config~) assert_true(aln.alignment_length() > 0) assert_true(aln.score() > 0.0) assert_true(aln.identities() > 0) } +///| test "pairaligner_aligned1_aligned2" { let aln = @src.pairaligner_align("ACGT", "ACGT") assert_eq(aln.aligned1(), aln.aligned_target) assert_eq(aln.aligned2(), aln.aligned_query) } +///| test "pairaligner_alignment_length" { let aln = @src.pairaligner_align("AAAA", "AAAA") assert_eq(aln.alignment_length(), 4) @@ -122,6 +134,7 @@ test "pairaligner_alignment_length" { assert_eq(aln.aligned2().length(), 4) } +///| test "pairaligner_identities_and_identity_pct" { let aln = @src.pairaligner_align("ACGT", "ACGT") assert_eq(aln.identities(), 4) @@ -131,23 +144,26 @@ test "pairaligner_identities_and_identity_pct" { assert_eq(aln2.identity_pct(), 75.0) } +///| test "pairaligner_gaps_count" { let config = @src.PairwiseAlignerConfig::default_dna() .set_gap_open(-2.0) .set_gap_extend(-1.0) let target = "AAACCC" let query = "AAA" - let aln = @src.pairaligner_align(target, query, config=config) + let aln = @src.pairaligner_align(target, query, config~) let gc = aln.gaps_count() assert_true(gc >= 3) } +///| test "pairaligner_score_extraction" { let aln = @src.pairaligner_align("ACGT", "ACGT") assert_eq(aln.score(), 4.0) assert_eq(aln.score(), aln.score) } +///| test "pairaligner_sample_data" { let (t, q) = @src.pairaligner_sample_data() assert_eq(t, "HEAGAWGHEE") @@ -156,6 +172,7 @@ test "pairaligner_sample_data" { assert_true(q.length() > 0) } +///| test "pairaligner_affine_vs_linear" { let target = "AAAAAAAAAA" let query = "AAAA" @@ -175,6 +192,7 @@ test "pairaligner_affine_vs_linear" { assert_true(linear_aln.score() < 100.0) } +///| test "pairaligner_config_new_named" { let config = @src.PairwiseAlignerConfig::new( mode=@src.pairaligner_local(), @@ -191,12 +209,14 @@ test "pairaligner_config_new_named" { assert_eq(config.alphabet_type, "DNA") } +///| test "pairaligner_local_mode_start_end_positions" { let target = "XXACGTYY" let query = "ACGT" - let config = @src.PairwiseAlignerConfig::default_dna() - .set_mode(@src.pairaligner_local()) - let aln = @src.pairaligner_align(target, query, config=config) + let config = @src.PairwiseAlignerConfig::default_dna().set_mode( + @src.pairaligner_local(), + ) + let aln = @src.pairaligner_align(target, query, config~) assert_eq(aln.target_start, 2) assert_eq(aln.target_end, 6) assert_eq(aln.query_start, 0) diff --git a/test/moonbit/pairwise2_test.mbt b/test/moonbit/pairwise2_test.mbt index feb40d10..582d1706 100644 --- a/test/moonbit/pairwise2_test.mbt +++ b/test/moonbit/pairwise2_test.mbt @@ -74,7 +74,7 @@ test "pairwise_local_convenience" { test "simple_score_match" { let scorer = @src.simple_score(2.0, -1.0) assert_true((scorer("A", "A") - 2.0).abs() < 0.001) - assert_true((scorer("A", "T") - (-1.0)).abs() < 0.001) + assert_true((scorer("A", "T") - -1.0).abs() < 0.001) } ///| @@ -94,7 +94,7 @@ test "identity_score" { test "dna_matrix" { let m = @src.dna_matrix(2.0, -1.0) assert_true((m.get("AA").unwrap() - 2.0).abs() < 0.001) - assert_true((m.get("AT").unwrap() - (-1.0)).abs() < 0.001) + assert_true((m.get("AT").unwrap() - -1.0).abs() < 0.001) assert_eq(m.size(), 10) // 10 unique pairs } @@ -103,7 +103,7 @@ test "matrix_score" { let m = @src.dna_matrix(2.0, -1.0) let scorer = @src.matrix_score(m, 0.0) assert_true((scorer("A", "A") - 2.0).abs() < 0.001) - assert_true((scorer("A", "T") - (-1.0)).abs() < 0.001) + assert_true((scorer("A", "T") - -1.0).abs() < 0.001) assert_true((scorer("X", "X") - 0.0).abs() < 0.001) // default } diff --git a/test/moonbit/paml_baseml_test.mbt b/test/moonbit/paml_baseml_test.mbt new file mode 100644 index 00000000..81423bc8 --- /dev/null +++ b/test/moonbit/paml_baseml_test.mbt @@ -0,0 +1,691 @@ +///| +fn baseml_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "BASEML value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn baseml_test_expect_control_error(text : String) -> Unit { + let raised = try { + ignore(@src.baseml_parse_control(text)) + false + } catch { + @src.BasemlError(_) => true + } + if !raised { + abort("expected BasemlError") + } +} + +///| +fn baseml_test_expect_result_error(text : String) -> Unit { + let raised = try { + ignore(@src.baseml_parse_results(text)) + false + } catch { + @src.BasemlError(_) => true + } + if !raised { + abort("expected BasemlError") + } +} + +///| +fn baseml_test_minimal_result() -> String { + "BASEML (in paml version 4.1, August 2008) alignment.phy K80\n" + + "lnL(ntime: 1 np: 2): -12.5 +0.0\n" + + "0.1 2.5\n" + + "tree length = 0.25\n" + + "(A:0.10,B:0.15);\n" +} + +///| +fn baseml_test_kappa_result() -> String { + baseml_test_minimal_result() + + "Parameters (kappa) in the rate matrix (TN93):\n" + + "3.25 0.75\n" +} + +///| +fn baseml_test_branch_result() -> String { + baseml_test_minimal_result() + + "Parameters (kappa) in the rate matrix (F84):\n" + + "1..2 0.10 2.50 3.50 1.25\n" + + "2..3 0.20 4.50 5.50 2.25\n" +} + +///| +fn baseml_test_auto_gamma_result() -> String { + baseml_test_minimal_result() + + "Parameters (kappa) in the rate matrix (K80):\n" + + "2.5\n" + + "alpha (gamma, K=2) = 0.25\n" + + "rate: 0.5 1.5\n" + + "freq: 0.4 0.6\n" + + "rho for the auto-discrete-gamma model: 0.9\n" + + "transition probabilities between rate categories:\n" + + "0.8 0.2\n" + + "0.1 0.9\n" +} + +///| +fn baseml_test_node_result() -> String { + baseml_test_minimal_result() + + "(frequency parameters for branches) [frequencies at nodes]\n" + + "Node #1 ( 0.10 0.20 0.30 0.40 )\n" + + "Node #2 ( 0.11 0.21 0.31 0.37 0.25 0.25 0.30 0.20 )\n" + + "Note: node 2 is root.\n" +} + +///| +fn baseml_test_old_frequency_result() -> String { + baseml_test_minimal_result() + + "base frequency parameters\n" + + "0.20 0.30 0.10 0.40\n" +} + +///| +fn baseml_test_null_result() -> String { + "BASEML (in paml version 4.7, January 2013) alignment.phy JC69\n" + + "lnL(ntime: 1 np: 2): -105.0 +0.0\n" + + "0.1 0.2\n" + + "tree length = 0.2\n" + + "(A:0.1,B:0.1);\n" +} + +///| +fn baseml_test_alternative_result() -> String { + "BASEML (in paml version 4.7, January 2013) alignment.phy HKY85\n" + + "lnL(ntime: 1 np: 4): -100.0 +0.0\n" + + "0.1 0.2 0.3 0.4\n" + + "tree length = 0.2\n" + + "(A:0.1,B:0.1);\n" +} + +///| +test "baseml control parses required paths" { + let control = @src.baseml_parse_control(@src.baseml_example_control_text()) + assert_eq(control.sequence_file(), "alignment.phylip") + assert_eq(control.output_file(), "baseml.out") + assert_eq(control.tree_file(), "species.tree") +} + +///| +test "baseml control parses all official options" { + let control = @src.baseml_parse_control(@src.baseml_example_control_text()) + assert_eq(control.options().length(), 21) + assert_eq(control.option("runmode"), Some("0")) + assert_eq(control.option("Small_Diff"), Some("0.000007")) +} + +///| +test "baseml control exposes model number" { + let control = @src.baseml_parse_control(@src.baseml_example_control_text()) + assert_eq(control.model_number(), Some(7)) + assert_eq(control.model_options(), "") +} + +///| +test "baseml control parses model 9 options" { + let control = @src.baseml_parse_control( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=9 [1 (TC CT AG GA)]\n", + ) + assert_eq(control.model_number(), Some(9)) + assert_eq(control.model_options(), "[1 (TC CT AG GA)]") +} + +///| +test "baseml control parses model 10 options" { + let control = @src.baseml_parse_control( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=10 [5 (AC CA) (AG GA)]\n", + ) + assert_eq(control.model_number(), Some(10)) + assert_eq(control.model_options(), "[5 (AC CA) (AG GA)]") +} + +///| +test "baseml control strips comments and CRLF" { + let control = @src.baseml_parse_control( + "seqfile = a.phy * alignment\r\noutfile = out\r\ntreefile = t.nwk\r\nmodel = 6 * TN93\r\n", + ) + assert_eq(control.sequence_file(), "a.phy") + assert_eq(control.option("model"), Some("6")) +} + +///| +test "baseml control canonical round trip" { + let first = @src.baseml_parse_control(@src.baseml_example_control_text()) + let written = @src.baseml_write_control(first) + let second = @src.baseml_parse_control(written) + assert_eq(first, second) + assert_true(written.has_prefix("seqfile = alignment.phylip\n")) +} + +///| +test "baseml control preserves option order" { + let control = @src.baseml_parse_control( + "seqfile=a\noutfile=b\ntreefile=c\nalpha=0.5\nmodel=4\n", + ) + let options = control.options() + assert_eq(options[0].name(), "alpha") + assert_eq(options[1].name(), "model") +} + +///| +test "baseml control options are defensive copies" { + let control = @src.baseml_parse_control(@src.baseml_example_control_text()) + let options = control.options() + ignore(options.pop()) + assert_eq(control.options().length(), 21) +} + +///| +test "baseml maps all model names" { + let names : Array[String] = [] + for model in 0..<=10 { + names.push(@src.baseml_model_name(model)) + } + assert_eq(names, [ + "JC69", "K80", "F81", "F84", "HKY85", "T92", "TN93", "REV", "UNREST", "REVu", + "UNRESTu", + ]) +} + +///| +test "baseml control rejects missing seqfile" { + baseml_test_expect_control_error("outfile=b\ntreefile=c\n") +} + +///| +test "baseml control rejects missing outfile" { + baseml_test_expect_control_error("seqfile=a\ntreefile=c\n") +} + +///| +test "baseml control rejects missing treefile" { + baseml_test_expect_control_error("seqfile=a\noutfile=b\n") +} + +///| +test "baseml control rejects malformed lines" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nmodel 7\n", + ) +} + +///| +test "baseml control rejects multiple equals signs" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=7=8\n", + ) +} + +///| +test "baseml control rejects duplicate keys" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=7\nmodel=8\n", + ) +} + +///| +test "baseml control rejects unknown options" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nunknown=1\n", + ) +} + +///| +test "baseml control rejects empty values" { + baseml_test_expect_control_error("seqfile=a\noutfile=b\ntreefile=c\nmodel=\n") +} + +///| +test "baseml control rejects invalid integers" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nncatG=five\n", + ) +} + +///| +test "baseml control rejects invalid doubles" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nalpha=fast\n", + ) +} + +///| +test "baseml control rejects unsupported model" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=11\n", + ) +} + +///| +test "baseml control rejects options on standard model" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nmodel=8 [custom]\n", + ) +} + +///| +test "baseml control rejects non-positive ncatG" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nncatG=0\n", + ) +} + +///| +test "baseml control rejects invalid nparK" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nnparK=5\n", + ) +} + +///| +test "baseml control rejects invalid nhomo" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\nnhomo=-1\n", + ) +} + +///| +test "baseml control rejects invalid binary option" { + baseml_test_expect_control_error( + "seqfile=a\noutfile=b\ntreefile=c\ngetSE=2\n", + ) +} + +///| +test "baseml model name rejects out of range value" { + let raised = try { + ignore(@src.baseml_model_name(-1)) + false + } catch { + @src.BasemlError(_) => true + } + assert_true(raised) +} + +///| +test "baseml result parses version and model description" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + assert_eq(result.version(), "4.7") + assert_eq(result.model_description(), "REV dGamma (ncatG=5)") +} + +///| +test "baseml result parses likelihoods" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + baseml_test_close(result.ln_likelihood(), -319.774741, 1.0e-9) + match result.ln_l_max() { + Some(value) => baseml_test_close(value, -316.049385, 1.0e-9) + None => abort("missing ln Lmax") + } +} + +///| +test "baseml result parses parameter vector" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + assert_eq(result.parameter_count(), 13) + assert_eq(result.parameter_list().length(), 13) + baseml_test_close(result.parameter_list()[7], 998.99998, 1.0e-8) +} + +///| +test "baseml result parses standard errors" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + assert_eq(result.standard_errors().length(), 13) + baseml_test_close(result.standard_errors()[12], 4.0, 1.0e-12) +} + +///| +test "baseml result parses tree" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + baseml_test_close(result.tree_length(), 0.01826, 1.0e-12) + assert_true(result.tree().has_prefix("(((Homo_sapie:")) +} + +///| +test "baseml result parses REV rate parameters" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + assert_eq(result.rate_parameters().length(), 5) + baseml_test_close(result.rate_parameters()[1], 130.94908, 1.0e-8) +} + +///| +test "baseml result parses base frequencies" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + match result.base_frequencies() { + Some(frequencies) => { + baseml_test_close(frequencies.thymine(), 0.20090, 1.0e-12) + baseml_test_close(frequencies.cytosine(), 0.16306, 1.0e-12) + baseml_test_close(frequencies.adenine(), 0.37027, 1.0e-12) + baseml_test_close(frequencies.guanine(), 0.26577, 1.0e-12) + } + None => abort("missing base frequencies") + } +} + +///| +test "baseml result parses Q matrix" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + match result.q_matrix() { + Some(matrix) => { + assert_eq(matrix.rows().length(), 4) + assert_eq(matrix.rows()[0].length(), 4) + baseml_test_close(matrix.rows()[2][3], 0.003044, 1.0e-12) + match matrix.average_ts_tv() { + Some(value) => baseml_test_close(value, 3.3698, 1.0e-12) + None => abort("missing average Ts/Tv") + } + } + None => abort("missing Q matrix") + } +} + +///| +test "baseml result parses gamma rates" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + match result.alpha() { + Some(value) => baseml_test_close(value, 200.95753, 1.0e-8) + None => abort("missing alpha") + } + assert_eq(result.rates().length(), 5) + assert_eq(result.rate_frequencies(), [0.2, 0.2, 0.2, 0.2, 0.2]) +} + +///| +test "baseml result parses scalar and multiple kappas" { + let result = @src.baseml_parse_results(baseml_test_kappa_result()) + assert_eq(result.kappas(), [3.25, 0.75]) +} + +///| +test "baseml result parses branch-specific kappas" { + let result = @src.baseml_parse_results(baseml_test_branch_result()) + let branches = result.branch_parameters() + assert_eq(branches.length(), 2) + assert_eq(branches[0].branch(), "1..2") + baseml_test_close(branches[0].time(), 0.10, 1.0e-12) + baseml_test_close(branches[0].kappa(), 2.50, 1.0e-12) + baseml_test_close(branches[0].transitions(), 3.50, 1.0e-12) + baseml_test_close(branches[0].transversions(), 1.25, 1.0e-12) +} + +///| +test "baseml result parses auto discrete gamma" { + let result = @src.baseml_parse_results(baseml_test_auto_gamma_result()) + match result.rho() { + Some(value) => baseml_test_close(value, 0.9, 1.0e-12) + None => abort("missing rho") + } + assert_eq(result.rates(), [0.5, 1.5]) + assert_eq(result.rate_frequencies(), [0.4, 0.6]) + assert_eq(result.transition_probabilities().length(), 2) + baseml_test_close(result.transition_probabilities()[1][1], 0.9, 1.0e-12) +} + +///| +test "baseml result parses nonhomogeneous nodes" { + let result = @src.baseml_parse_results(baseml_test_node_result()) + let nodes = result.nodes() + assert_eq(nodes.length(), 2) + assert_eq(nodes[0].node(), 1) + assert_false(nodes[0].is_root()) + assert_true(nodes[1].is_root()) + assert_eq(nodes[0].frequency_parameters(), [0.1, 0.2, 0.3, 0.4]) +} + +///| +test "baseml result parses realized node base frequencies" { + let result = @src.baseml_parse_results(baseml_test_node_result()) + match result.nodes()[1].base_frequencies() { + Some(frequencies) => { + baseml_test_close(frequencies.thymine(), 0.25, 1.0e-12) + baseml_test_close(frequencies.guanine(), 0.20, 1.0e-12) + } + None => abort("missing node base frequencies") + } +} + +///| +test "baseml result parses PAML 4.1 base frequency heading" { + let result = @src.baseml_parse_results(baseml_test_old_frequency_result()) + match result.base_frequencies() { + Some(frequencies) => + baseml_test_close(frequencies.cytosine(), 0.30, 1.0e-12) + None => abort("missing old-style base frequencies") + } +} + +///| +test "baseml result arrays are defensive copies" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + let parameters = result.parameter_list() + parameters[0] = 999.0 + let rates = result.rates() + ignore(rates.pop()) + baseml_test_close(result.parameter_list()[0], 0.000004, 1.0e-12) + assert_eq(result.rates().length(), 5) +} + +///| +test "baseml Q matrix rows are defensive copies" { + let result = @src.baseml_parse_results(@src.baseml_example_results_text()) + match result.q_matrix() { + Some(matrix) => { + let rows = matrix.rows() + rows[0][0] = 100.0 + match result.q_matrix() { + Some(again) => baseml_test_close(again.rows()[0][0], -2.483179, 1.0e-12) + None => abort("missing copied Q matrix") + } + } + None => abort("missing Q matrix") + } +} + +///| +test "baseml node arrays are defensive copies" { + let result = @src.baseml_parse_results(baseml_test_node_result()) + let nodes = result.nodes() + let parameters = nodes[0].frequency_parameters() + parameters[0] = 9.0 + baseml_test_close(result.nodes()[0].frequency_parameters()[0], 0.1, 1.0e-12) +} + +///| +test "baseml result rejects empty input" { + baseml_test_expect_result_error("") +} + +///| +test "baseml result rejects missing header" { + baseml_test_expect_result_error( + "lnL(ntime: 1 np: 1): -1\n0.1\ntree length = 0.1\n(A:0.1);\n", + ) +} + +///| +test "baseml result rejects missing likelihood" { + baseml_test_expect_result_error( + "BASEML (in paml version 4.7, January 2013) a.phy JC69\n", + ) +} + +///| +test "baseml result rejects malformed likelihood" { + baseml_test_expect_result_error( + "BASEML (in paml version 4.7, January 2013) a.phy JC69\n" + + "lnL(ntime: 1 np: x): bad\n", + ) +} + +///| +test "baseml result rejects parameter vector mismatch" { + baseml_test_expect_result_error( + "BASEML (in paml version 4.7, January 2013) a.phy JC69\n" + + "lnL(ntime: 1 np: 2): -1\n" + + "0.1\n" + + "tree length = 0.1\n" + + "(A:0.1);\n", + ) +} + +///| +test "baseml result rejects standard error mismatch" { + baseml_test_expect_result_error( + baseml_test_minimal_result() + "SEs for parameters:\n0.1\n", + ) +} + +///| +test "baseml result rejects missing tree length" { + baseml_test_expect_result_error( + "BASEML (in paml version 4.7, January 2013) a.phy JC69\n" + + "lnL(ntime: 1 np: 1): -1\n" + + "0.1\n" + + "(A:0.1);\n", + ) +} + +///| +test "baseml result rejects missing branch-length tree" { + baseml_test_expect_result_error( + "BASEML (in paml version 4.7, January 2013) a.phy JC69\n" + + "lnL(ntime: 1 np: 1): -1\n" + + "0.1\n" + + "tree length = 0.1\n" + + "(A,B);\n", + ) +} + +///| +test "baseml result rejects incomplete Q matrix" { + baseml_test_expect_result_error( + baseml_test_minimal_result() + + "Rate matrix Q, Average Ts/Tv = 2.0\n" + + "1.0 0.0 0.0 0.0\n", + ) +} + +///| +test "baseml result rejects transition matrix row count mismatch" { + baseml_test_expect_result_error( + baseml_test_minimal_result() + + "rate: 0.5 1.5\n" + + "transition probabilities between rate categories:\n" + + "0.8 0.2\n", + ) +} + +///| +test "baseml result rejects transition matrix column mismatch" { + baseml_test_expect_result_error( + baseml_test_minimal_result() + + "rate: 0.5 1.5\n" + + "transition probabilities between rate categories:\n" + + "0.8 0.2\n" + + "1.0\n", + ) +} + +///| +test "baseml result rejects duplicate nodes" { + baseml_test_expect_result_error( + baseml_test_minimal_result() + + "Node #1 ( 0.1 0.2 0.3 0.4 )\n" + + "Node #1 ( 0.2 0.2 0.2 0.4 )\n", + ) +} + +///| +test "baseml computes AIC" { + let result = @src.baseml_parse_results(baseml_test_null_result()) + baseml_test_close(result.aic(), 214.0, 1.0e-12) +} + +///| +test "baseml computes BIC" { + let result = @src.baseml_parse_results(baseml_test_null_result()) + baseml_test_close(result.bic(100), 2.0 * @math.ln(100.0) + 210.0, 1.0e-12) +} + +///| +test "baseml BIC rejects non-positive observations" { + let result = @src.baseml_parse_results(baseml_test_null_result()) + let raised = try { + ignore(result.bic(0)) + false + } catch { + @src.BasemlError(_) => true + } + assert_true(raised) +} + +///| +test "baseml computes nested likelihood ratio" { + let null_model = @src.baseml_parse_results(baseml_test_null_result()) + let alternative = @src.baseml_parse_results(baseml_test_alternative_result()) + let comparison = @src.baseml_likelihood_ratio(null_model, alternative) + baseml_test_close(comparison.statistic(), 10.0, 1.0e-12) + assert_eq(comparison.degrees_of_freedom(), 2) + baseml_test_close(comparison.p_value(), 0.006737946999, 1.0e-9) + assert_eq(comparison.alpha(), 0.05) + assert_true(comparison.significant()) +} + +///| +test "baseml LRT rejects reversed nesting" { + let null_model = @src.baseml_parse_results(baseml_test_null_result()) + let alternative = @src.baseml_parse_results(baseml_test_alternative_result()) + let raised = try { + ignore(@src.baseml_likelihood_ratio(alternative, null_model)) + false + } catch { + @src.BasemlError(_) => true + } + assert_true(raised) +} + +///| +test "baseml LRT rejects lower alternative likelihood" { + let null_model = @src.baseml_parse_results(baseml_test_minimal_result()) + let alternative = @src.baseml_parse_results( + "BASEML (in paml version 4.7, January 2013) a.phy HKY85\n" + + "lnL(ntime: 1 np: 3): -20.0\n" + + "0.1 0.2 0.3\n" + + "tree length = 0.2\n" + + "(A:0.1,B:0.1);\n", + ) + let raised = try { + ignore(@src.baseml_likelihood_ratio(null_model, alternative)) + false + } catch { + @src.BasemlError(_) => true + } + assert_true(raised) +} + +///| +test "baseml LRT rejects invalid alpha" { + let null_model = @src.baseml_parse_results(baseml_test_null_result()) + let alternative = @src.baseml_parse_results(baseml_test_alternative_result()) + let raised = try { + ignore(@src.baseml_likelihood_ratio(null_model, alternative, alpha=1.0)) + false + } catch { + @src.BasemlError(_) => true + } + assert_true(raised) +} diff --git a/test/moonbit/paml_codeml_test.mbt b/test/moonbit/paml_codeml_test.mbt new file mode 100644 index 00000000..080dc360 --- /dev/null +++ b/test/moonbit/paml_codeml_test.mbt @@ -0,0 +1,652 @@ +///| +fn codeml_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "CODEML value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn codeml_test_expect_control_error(text : String) -> Unit { + let raised = try { + ignore(@src.codeml_parse_control(text)) + false + } catch { + @src.CodemlError(_) => true + } + if !raised { + abort("expected CodemlError") + } +} + +///| +fn codeml_test_expect_result_error(text : String) -> Unit { + let raised = try { + ignore(@src.codeml_parse_results(text)) + false + } catch { + @src.CodemlError(_) => true + } + if !raised { + abort("expected CodemlError") + } +} + +///| +fn codeml_test_branch_site_text() -> String { + "CODONML (in paml version 4.7, January 2013) alignment.phylip\n" + + "Model: branch-site model A\n" + + "Codon frequency model: F3x4\n" + + "Site-class models: PositiveSelection\n" + + "ns = 5 ls = 74\n" + + "lnL(ntime: 7 np: 12): -308.031579 +0.000000\n" + + "0.1 0.2 0.3 0.4 0.5 0.6 0.7 1.5 0.9 0.1 0.2 2.0\n" + + "tree length = 0.05472\n" + + "((A:0.01, B:0.02):0.03, C:0.04);\n" + + "proportion 0.70000 0.10000 0.15000 0.05000\n" + + "background w 0.10000 1.00000 0.10000 1.00000\n" + + "foreground w 0.10000 1.00000 4.50000 4.50000\n" + + "Bayes Empirical Bayes (BEB) analysis\n" + + " 21 R 0.981* 3.750 +- 0.820\n" + + " 45 K 0.999** 8.250 +- 1.400\n" +} + +///| +fn codeml_test_clade_text() -> String { + "CODONML (in paml version 4.9j, February 2020) alignment.phylip\n" + + "Model: Clade model C\n" + + "Codon frequency model: F3x4\n" + + "Site-class models: PositiveSelection\n" + + "ns = 5 ls = 74\n" + + "lnL(ntime: 7 np: 13): -307.500000 +0.000000\n" + + "0.1 0.2 0.3 0.4 0.5 0.6 0.7 1.5 0.9 0.1 0.2 2.0 3.0\n" + + "proportion 0.60000 0.30000 0.10000\n" + + "branch type 0: 0.10000 1.00000 2.50000\n" + + "branch type 1: 0.20000 1.00000 4.00000\n" +} + +///| +fn codeml_test_free_ratio_text() -> String { + "CODONML (in paml version 4.7, January 2013) alignment.phylip\n" + + "Model: free ratios for branches\n" + + "Codon frequency model: F3x4\n" + + "ns = 3 ls = 90\n" + + "lnL(ntime: 4 np: 8): -120.500000 +0.000000\n" + + "0.1 0.2 0.3 0.4 1.1 0.8 0.4 0.2\n" + + "tree length = 0.12000\n" + + "((A:0.01, B:0.02):0.03, C:0.04);\n" + + "kappa (ts/tv) = 2.10000\n" + + "w (dN/dS) for branches: 1.10000 0.80000 0.40000 0.20000\n" + + " 4..5 0.010 100.0 50.0 1.1000 0.0050 0.0045 0.5 0.2\n" + + " 5..1 0.020 100.0 50.0 0.8000 0.0040 0.0050 0.4 0.3\n" + + "dS tree:\n" + + "((A:0.004, B:0.005):0.006, C:0.007);\n" + + "dN tree:\n" + + "((A:0.005, B:0.004):0.003, C:0.002);\n" + + "w ratios as labels for TreeView:\n" + + "((A #1.1, B #0.8) #0.4, C #0.2);\n" +} + +///| +fn codeml_test_pairwise_text() -> String { + "CODONML (in paml version 4.7, January 2013) alignment.phylip\n" + + "Model: One dN/dS ratio for branches,\n" + + "Codon frequency model: F3x4\n" + + "ns = 3 ls = 74\n" + + "2 (Pan_troglo) ... 1 (Homo_sapie)\n" + + "lnL = -291.465693\n" + + "0.01262 999.00000 0.00100\n" + + "t= 0.0126 S=81.4 N=140.6 dN/dS=0.0010 dN=0.0000 dS=0.0115\n" + + "3 (Gorilla_go) ... 1 (Homo_sapie)\n" + + "lnL = -290.441129\n" + + "0.01265 999.00000 0.00100\n" + + "t=0.0127 S=81.7 N=140.3 dN/dS=0.0010 dN=0.0000 dS=0.0114\n" + + "3 (Gorilla_go) ... 2 (Pan_troglo)\n" + + "lnL = -296.416525\n" + + "0.02582 999.00000 0.00100\n" + + "t=0.0258 S=81.5 N=140.5 dN/dS=0.0010 dN=0.0000 dS=0.0234\n" +} + +///| +fn codeml_test_aaml_text() -> String { + "AAML (in paml version 4.7, January 2013) aa_alignment.phylip\n" + + "Model: Poisson for branches, ns = 3 ls = 74\n" + + "ln Lmax (unconstrained) = -192.057157\n" + + "AA distances (raw proportions of different sites)\n" + + "Human\n" + + "Chimp 0.0100\n" + + "Gorilla 0.0200 0.0300\n" + + "\n" + + "ML distances of aa seqs.\n" + + "Human\n" + + "Chimp 0.0110\n" + + "Gorilla 0.0220 0.0330\n" + + "\n" +} + +///| +fn codeml_test_se_text() -> String { + "CODONML (in paml version 4.6, August 2010) alignment.phylip\n" + + "Model: one-ratio\n" + + "Codon frequency model: F3x4\n" + + "ns = 3 ls = 60\n" + + "lnL(ntime: 4 np: 6): -100.000000 +0.000000\n" + + "0.1 0.2 0.3 0.4 2.0 0.5\n" + + "SEs for parameters:\n" + + "0.01 0.02 0.03 0.04 0.20 0.05\n" +} + +///| +fn codeml_test_multigene_text() -> String { + "CODONML (in paml version 4.7, January 2013) genes.phylip\n" + + "Model: One dN/dS ratio for branches, (2 genes: separate data)\n" + + "Codon frequency model: F3x4\n" + + "Site-class models: one-ratio\n" + + "ns = 4 ls = 120\n" + + "Gene 1\n" + + "lnL(ntime: 3 np: 5): -80.000000 +0.000000\n" + + "0.1 0.2 0.3 2.0 0.4\n" + + "tree length = 0.10\n" + + "Gene 2\n" + + "lnL(ntime: 3 np: 5): -90.000000 +0.000000\n" + + "0.1 0.2 0.3 2.1 0.5\n" + + "tree length = 0.20\n" +} + +///| +fn codeml_test_joint_gene_text() -> String { + "CODONML (in paml version 4.7, January 2013) genes.phylip\n" + + "Model: One dN/dS ratio for branches, (2 genes: joint data)\n" + + "Codon frequency model: F3x4\n" + + "ns = 4 ls = 120\n" + + "lnL(ntime: 7 np: 10): -170.000000 +0.000000\n" + + "0.1 0.2 0.3 0.4 0.5 0.6 2.0 0.3 1.0 2.5\n" + + "rates for 2 genes: 1 2.50000\n" + + "gene # 1: kappa = 1.70000 omega = 0.30000\n" + + "gene # 2: kappa = 1.90000 omega = 1.20000\n" +} + +///| +test "codeml control parses required paths" { + let control = @src.codeml_parse_control(@src.codeml_example_control_text()) + assert_eq(control.sequence_file(), "alignment.phylip") + assert_eq(control.output_file(), "codeml.out") + assert_eq(control.tree_file(), "species.tree") +} + +///| +test "codeml control parses options" { + let control = @src.codeml_parse_control(@src.codeml_example_control_text()) + assert_eq(control.option("CodonFreq"), Some("2")) + assert_eq(control.option("missing"), None) +} + +///| +test "codeml control parses NSsites integer list" { + let control = @src.codeml_parse_control(@src.codeml_example_control_text()) + assert_eq(control.ns_sites(), [0, 1, 2]) +} + +///| +test "codeml control strips star comments" { + let control = @src.codeml_parse_control( + "seqfile=a.phy * alignment\noutfile=x.out\ntreefile=t.tree\nomega=1 * start\n", + ) + assert_eq(control.sequence_file(), "a.phy") + assert_eq(control.option("omega"), Some("1")) +} + +///| +test "codeml control accepts CRLF" { + let control = @src.codeml_parse_control( + "seqfile=a.phy\r\noutfile=x.out\r\ntreefile=t.tree\r\nrunmode=-2\r\n", + ) + assert_eq(control.option("runmode"), Some("-2")) +} + +///| +test "codeml control canonical round trip" { + let control = @src.codeml_parse_control(@src.codeml_example_control_text()) + let written = @src.codeml_write_control(control) + let reparsed = @src.codeml_parse_control(written) + assert_eq(reparsed, control) +} + +///| +test "codeml control preserves option order" { + let control = @src.codeml_parse_control( + "seqfile=a\noutfile=b\ntreefile=c\nomega=1\nkappa=2\nmodel=0\n", + ) + let options = control.options() + assert_eq(options[0].name(), "omega") + assert_eq(options[1].name(), "kappa") + assert_eq(options[2].name(), "model") +} + +///| +test "codeml control option accessor is defensive" { + let control = @src.codeml_parse_control(@src.codeml_example_control_text()) + let options = control.options() + options[0] = options[1] + assert_eq(control.options()[0].name(), "verbose") +} + +///| +test "codeml control rejects missing seqfile" { + codeml_test_expect_control_error("outfile=x\ntreefile=t\n") +} + +///| +test "codeml control rejects missing outfile" { + codeml_test_expect_control_error("seqfile=a\ntreefile=t\n") +} + +///| +test "codeml control rejects missing treefile" { + codeml_test_expect_control_error("seqfile=a\noutfile=x\n") +} + +///| +test "codeml control rejects malformed line" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nmodel 0\n", + ) +} + +///| +test "codeml control rejects unknown option" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nunknown=1\n", + ) +} + +///| +test "codeml control rejects duplicate path" { + codeml_test_expect_control_error( + "seqfile=a\nseqfile=b\noutfile=x\ntreefile=t\n", + ) +} + +///| +test "codeml control rejects duplicate option" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nomega=1\nomega=2\n", + ) +} + +///| +test "codeml control rejects invalid integer" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nmodel=zero\n", + ) +} + +///| +test "codeml control rejects invalid double" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nomega=one\n", + ) +} + +///| +test "codeml control rejects negative NSsites" { + codeml_test_expect_control_error( + "seqfile=a\noutfile=x\ntreefile=t\nNSsites=0 -1\n", + ) +} + +///| +test "codeml results parse header metadata" { + let results = @src.codeml_parse_results(@src.codeml_example_results_text()) + assert_eq(results.program(), "CODONML") + assert_eq(results.version(), "4.7") + assert_eq(results.codon_frequency_model(), "F3x4") + assert_eq(results.sequence_count(), 5) + assert_eq(results.site_count(), 74) +} + +///| +test "codeml results parse multiple NSsites models" { + let results = @src.codeml_parse_results(@src.codeml_example_results_text()) + assert_eq(results.models().length(), 2) + assert_eq(results.models()[0].number(), 0) + assert_eq(results.models()[1].number(), 2) +} + +///| +test "codeml model parses likelihood and parameter count" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + codeml_test_close(model.ln_likelihood(), -308.032865, 1.0e-9) + assert_eq(model.parameter_count(), 9) + assert_true(model.parameter_list().length() > 0) +} + +///| +test "codeml model parses tree and length" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + codeml_test_close(model.tree_length().unwrap(), 0.05472, 1.0e-10) + assert_true(model.tree().length() > 20) +} + +///| +test "codeml model parses kappa and omega" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + codeml_test_close(model.kappa().unwrap(), 1.51919, 1.0e-10) + codeml_test_close(model.omegas()[0], 0.00010, 1.0e-10) +} + +///| +test "codeml model parses branch table" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + assert_eq(model.branches().length(), 2) + assert_eq(model.branches()[1].branch(), "8..2") + codeml_test_close(model.branches()[1].ds().unwrap(), 0.0183, 1.0e-10) +} + +///| +test "codeml model parses dN and dS tree lengths" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + codeml_test_close(model.dn_tree_length().unwrap(), 0.0, 1.0e-12) + codeml_test_close(model.ds_tree_length().unwrap(), 0.0746, 1.0e-10) +} + +///| +test "codeml model parses site classes" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[1] + assert_eq(model.site_classes().length(), 3) + codeml_test_close(model.site_classes()[2].proportion(), 0.02, 1.0e-10) + codeml_test_close(model.site_classes()[2].omega().unwrap(), 5.0, 1.0e-10) +} + +///| +test "codeml model parses BEB positive site" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[1] + let sites = model.positive_sites() + assert_eq(sites.length(), 1) + assert_eq(sites[0].analysis_method(), "BEB") + assert_eq(sites[0].position(), 42) + assert_eq(sites[0].amino_acid(), "K") + assert_eq(sites[0].significance(), "**") + codeml_test_close(sites[0].probability(), 0.997, 1.0e-10) + codeml_test_close(sites[0].posterior_mean_omega().unwrap(), 5.2, 1.0e-10) +} + +///| +test "codeml model lookup finds requested NSsites class" { + let results = @src.codeml_parse_results(@src.codeml_example_results_text()) + assert_eq(results.model_by_number(2).unwrap().number(), 2) + assert_eq(results.model_by_number(8), None) +} + +///| +test "codeml result model accessor is deeply defensive" { + let results = @src.codeml_parse_results(@src.codeml_example_results_text()) + let models = results.models() + let classes = models[1].site_classes() + classes[0] = classes[1] + assert_eq(results.models()[1].site_classes()[0].index(), 0) +} + +///| +test "codeml branch-site A parses four classes" { + let model = @src.codeml_parse_results(codeml_test_branch_site_text()).models()[0] + assert_eq(model.number(), 2) + assert_eq(model.site_classes().length(), 4) +} + +///| +test "codeml branch-site A parses foreground and background omega" { + let classes = @src.codeml_parse_results(codeml_test_branch_site_text()).models()[0].site_classes() + assert_eq(classes[2].branch_types().length(), 2) + assert_eq(classes[2].branch_types()[0].name(), "background") + codeml_test_close(classes[2].branch_types()[1].value(), 4.5, 1.0e-10) +} + +///| +test "codeml branch-site A parses multiple selected sites" { + let sites = @src.codeml_parse_results(codeml_test_branch_site_text()).models()[0].positive_sites() + assert_eq(sites.length(), 2) + assert_eq(sites[0].significance(), "*") + assert_eq(sites[1].significance(), "**") +} + +///| +test "codeml clade model C parses branch types" { + let classes = @src.codeml_parse_results(codeml_test_clade_text()).models()[0].site_classes() + assert_eq(classes.length(), 3) + assert_eq(classes[0].branch_types()[0].name(), "0") + assert_eq(classes[0].branch_types()[1].name(), "1") + codeml_test_close(classes[2].branch_types()[1].value(), 4.0, 1.0e-10) +} + +///| +test "codeml free-ratio parses branch omegas" { + let model = @src.codeml_parse_results(codeml_test_free_ratio_text()).models()[0] + assert_eq(model.omegas().length(), 4) + codeml_test_close(model.omegas()[0], 1.1, 1.0e-10) +} + +///| +test "codeml free-ratio parses specialized trees" { + let model = @src.codeml_parse_results(codeml_test_free_ratio_text()).models()[0] + assert_true(model.dn_tree().length() > 0) + assert_true(model.ds_tree().length() > 0) + assert_true(model.omega_tree().length() > 0) +} + +///| +test "codeml free-ratio parses branch rows" { + let branches = @src.codeml_parse_results(codeml_test_free_ratio_text()).models()[0].branches() + assert_eq(branches.length(), 2) + codeml_test_close(branches[0].omega().unwrap(), 1.1, 1.0e-10) +} + +///| +test "codeml pairwise parses all unordered comparisons" { + let pairwise = @src.codeml_parse_results(codeml_test_pairwise_text()).pairwise() + assert_eq(pairwise.length(), 3) + assert_eq(pairwise[0].first(), "Pan_troglo") + assert_eq(pairwise[0].second(), "Homo_sapie") +} + +///| +test "codeml pairwise parses likelihood" { + let pair = @src.codeml_parse_results(codeml_test_pairwise_text()).pairwise()[0] + codeml_test_close(pair.ln_likelihood(), -291.465693, 1.0e-9) +} + +///| +test "codeml pairwise parses site counts and rates" { + let pair = @src.codeml_parse_results(codeml_test_pairwise_text()).pairwise()[0] + codeml_test_close(pair.time(), 0.0126, 1.0e-10) + codeml_test_close(pair.synonymous_sites(), 81.4, 1.0e-10) + codeml_test_close(pair.nonsynonymous_sites(), 140.6, 1.0e-10) + codeml_test_close(pair.omega(), 0.001, 1.0e-10) + codeml_test_close(pair.dn(), 0.0, 1.0e-10) + codeml_test_close(pair.ds(), 0.0115, 1.0e-10) +} + +///| +test "codeml AAML parses program and unconstrained likelihood" { + let results = @src.codeml_parse_results(codeml_test_aaml_text()) + assert_eq(results.program(), "AAML") + codeml_test_close(results.ln_l_max().unwrap(), -192.057157, 1.0e-9) +} + +///| +test "codeml AAML parses raw distance triangle" { + let distances = @src.codeml_parse_results(codeml_test_aaml_text()).distances() + assert_eq(distances.length(), 6) + assert_eq(distances[0].kind(), "raw") + assert_eq(distances[0].first(), "Chimp") + assert_eq(distances[0].second(), "Human") + codeml_test_close(distances[0].value(), 0.01, 1.0e-10) +} + +///| +test "codeml AAML parses ML distance triangle" { + let distances = @src.codeml_parse_results(codeml_test_aaml_text()).distances() + assert_eq(distances[3].kind(), "ml") + codeml_test_close(distances[5].value(), 0.033, 1.0e-10) +} + +///| +test "codeml parses parameter standard errors" { + let model = @src.codeml_parse_results(codeml_test_se_text()).models()[0] + assert_true(model.standard_errors().length() > 0) + assert_eq(model.parameter_count(), 6) +} + +///| +test "codeml parses separate multi-gene sections" { + let models = @src.codeml_parse_results(codeml_test_multigene_text()).models() + assert_eq(models.length(), 2) + assert_eq(models[0].gene(), 1) + assert_eq(models[1].gene(), 2) + codeml_test_close(models[1].tree_length().unwrap(), 0.2, 1.0e-10) +} + +///| +test "codeml parses joint multi-gene relative rates" { + let model = @src.codeml_parse_results(codeml_test_joint_gene_text()).models()[0] + assert_eq(model.rates(), [1.0, 2.5]) +} + +///| +test "codeml parses joint multi-gene kappa and omega" { + let genes = @src.codeml_parse_results(codeml_test_joint_gene_text()).models()[0].gene_parameters() + assert_eq(genes.length(), 2) + assert_eq(genes[1].gene(), 2) + codeml_test_close(genes[1].kappa(), 1.9, 1.0e-10) + codeml_test_close(genes[1].omega(), 1.2, 1.0e-10) +} + +///| +test "codeml AIC matches definition" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + codeml_test_close(model.aic(), 634.06573, 1.0e-6) +} + +///| +test "codeml BIC matches definition" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + let expected = 9.0 * @math.ln(74.0) + 616.06573 + codeml_test_close(model.bic(74), expected, 1.0e-6) +} + +///| +test "codeml likelihood-ratio statistic and degrees" { + let models = @src.codeml_parse_results(@src.codeml_example_results_text()).models() + let comparison = @src.codeml_likelihood_ratio(models[0], models[1]) + codeml_test_close(comparison.statistic(), 12.002572, 1.0e-6) + assert_eq(comparison.degrees_of_freedom(), 3) +} + +///| +test "codeml likelihood-ratio chi-square tail" { + let models = @src.codeml_parse_results(@src.codeml_example_results_text()).models() + let comparison = @src.codeml_likelihood_ratio(models[0], models[1]) + assert_true(comparison.p_value() > 0.007) + assert_true(comparison.p_value() < 0.008) + assert_true(comparison.significant()) + codeml_test_close(comparison.alpha(), 0.05, 1.0e-12) +} + +///| +test "codeml best AIC selects improved model" { + let models = @src.codeml_parse_results(@src.codeml_example_results_text()).models() + assert_eq(@src.codeml_best_aic(models).number(), 2) +} + +///| +test "codeml BIC rejects non-positive observations" { + let model = @src.codeml_parse_results(@src.codeml_example_results_text()).models()[0] + let raised = try { + ignore(model.bic(0)) + false + } catch { + @src.CodemlError(_) => true + } + assert_true(raised) +} + +///| +test "codeml LRT rejects reversed nesting" { + let models = @src.codeml_parse_results(@src.codeml_example_results_text()).models() + let raised = try { + ignore(@src.codeml_likelihood_ratio(models[1], models[0])) + false + } catch { + @src.CodemlError(_) => true + } + assert_true(raised) +} + +///| +test "codeml LRT rejects invalid alpha" { + let models = @src.codeml_parse_results(@src.codeml_example_results_text()).models() + let raised = try { + ignore(@src.codeml_likelihood_ratio(models[0], models[1], alpha=1.0)) + false + } catch { + @src.CodemlError(_) => true + } + assert_true(raised) +} + +///| +test "codeml best AIC rejects empty input" { + let raised = try { + ignore(@src.codeml_best_aic([])) + false + } catch { + @src.CodemlError(_) => true + } + assert_true(raised) +} + +///| +test "codeml results reject empty text" { + codeml_test_expect_result_error("") +} + +///| +test "codeml results reject missing PAML header" { + codeml_test_expect_result_error( + "Model: one-ratio\nlnL(ntime: 1 np: 1): -1.0\n", + ) +} + +///| +test "codeml results reject header without estimates" { + codeml_test_expect_result_error( + "CODONML (in paml version 4.7, January 2013) a.phy\nModel: x\n", + ) +} + +///| +test "codeml results reject malformed model likelihood" { + codeml_test_expect_result_error( + "CODONML (in paml version 4.7, January 2013) a.phy\n" + + "Model 0: one-ratio\n" + + "lnL(ntime: 1 np: 1): not-a-number\n", + ) +} + +///| +test "codeml results reject incomplete pairwise estimate" { + codeml_test_expect_result_error( + "CODONML (in paml version 4.7, January 2013) a.phy\n" + + "2 (B) ... 1 (A)\n" + + "lnL = -1.0\n" + + "t=0.1 S=10 N=20 dN/dS=0.5 dN=0.01\n", + ) +} diff --git a/test/moonbit/paml_test.mbt b/test/moonbit/paml_test.mbt index b4bfdd1f..070fb497 100644 --- a/test/moonbit/paml_test.mbt +++ b/test/moonbit/paml_test.mbt @@ -1,17 +1,18 @@ ///| /// Tests for PAML module. - test "PAMLAlignment creation" { let seqs = [("Human", "ATGCCG"), ("Mouse", "ATGCCG")] let alignment = @src.PAMLAlignment::new(seqs) assert_eq(alignment.sequences.length(), 2) } +///| test "PAMLResult creation" { let result = @src.PAMLResult::new(-100.0) assert_eq(result.ln_likelihood, -100.0) } +///| test "DNDSResult creation" { let dnds = @src.DNDSResult::new(0.1, 0.5, 0.2) assert_eq(dnds.dN, 0.1) @@ -19,17 +20,20 @@ test "DNDSResult creation" { assert_eq(dnds.omega, 0.2) } +///| test "paml_calculate_dnds identical sequences" { let dnds = @src.paml_calculate_dnds("ATGCCG", "ATGCCG", "Nei-Gojobori") assert_true(dnds.omega >= 0.0) } +///| test "paml_calculate_dnds different sequences" { let dnds = @src.paml_calculate_dnds("ATGCCG", "ATTTTT", "Nei-Gojobori") assert_true(dnds.dN >= 0.0) assert_true(dnds.dS >= 0.0) } +///| test "paml_estimate_parameters" { let seqs = [("Seq1", "ATGC"), ("Seq2", "ATGC")] let alignment = @src.PAMLAlignment::new(seqs) @@ -37,24 +41,31 @@ test "paml_estimate_parameters" { assert_true(result.parameters.contains("pi_A")) } +///| test "paml_calculate_substitution_matrix" { - let pi = Map([("pi_A", 0.25), ("pi_T", 0.25), ("pi_C", 0.25), ("pi_G", 0.25)], capacity=4) + let pi = Map( + [("pi_A", 0.25), ("pi_T", 0.25), ("pi_C", 0.25), ("pi_G", 0.25)], + capacity=4, + ) let matrix = @src.paml_calculate_substitution_matrix(2.0, pi) assert_eq(matrix.length(), 4) } +///| test "paml_create_example_alignment" { let alignment = @src.paml_create_example_alignment() assert_eq(alignment.sequences.length(), 4) } +///| test "paml_run_likelihood" { let alignment = @src.paml_create_example_alignment() let result = @src.paml_run_likelihood(alignment) assert_true(result.dnds_ratios.length() > 0) } +///| test "get_standard_codon_table" { let table = @src.get_standard_codon_table() assert_true(table.contains("ATG")) -} \ No newline at end of file +} diff --git a/test/moonbit/paml_yn00_test.mbt b/test/moonbit/paml_yn00_test.mbt new file mode 100644 index 00000000..8e4c8578 --- /dev/null +++ b/test/moonbit/paml_yn00_test.mbt @@ -0,0 +1,782 @@ +///| +fn yn00_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "YN00 value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn yn00_test_expect_control_error(text : String) -> Unit { + let raised = try { + ignore(@src.yn00_parse_control(text)) + false + } catch { + @src.Yn00Error(_) => true + } + if !raised { + abort("expected Yn00Error") + } +} + +///| +fn yn00_test_expect_result_error(text : String) -> Unit { + let raised = try { + ignore(@src.yn00_parse_results(text)) + false + } catch { + @src.Yn00Error(_) => true + } + if !raised { + abort("expected Yn00Error") + } +} + +///| +fn yn00_test_two_sequence_result( + first : String, + second : String, + ng_row : String, + lwl85_omega : String, + modified_ds : String, +) -> String { + "YN00 pairwise.phy\n" + + "ns = 2 ls = 100\n" + + "(A) Nei-Gojobori (1986) method\n" + + first + + "\n" + + ng_row + + "\n" + + "(B) Yang & Nielsen (2000) method\n" + + "2 1 30.0 70.0 0.1 2.0 0.5 0.01 +- 0.001 0.02 +- 0.002\n" + + "(C) LWL85, LPB93 & LWLm methods\n" + + "2 (" + + second + + ") vs. 1 (" + + first + + ")\n" + + "LWL85: dS = 0.0200 dN = 0.0100 w = " + + lwl85_omega + + " S = 30.0 N = 70.0\n" + + "LWL85m: dS = " + + modified_ds + + " dN = -nan w = -nan S = -nan N = -nan (rho = -nan)\n" + + "LPB93: dS = 0.0200 dN = 0.0100 w = 0.5000\n" +} + +///| +test "yn00 control parses required paths" { + let control = @src.yn00_parse_control(@src.yn00_example_control_text()) + assert_eq(control.sequence_file(), "alignment.phylip") + assert_eq(control.output_file(), "yn00.out") +} + +///| +test "yn00 control parses all official options" { + let control = @src.yn00_parse_control(@src.yn00_example_control_text()) + assert_eq(control.option("verbose"), Some("1")) + assert_eq(control.option("icode"), Some("0")) + assert_eq(control.option("weighting"), Some("0")) + assert_eq(control.option("commonf3x4"), Some("0")) + assert_eq(control.option("ndata"), Some("1")) +} + +///| +test "yn00 control ignores comments and blank lines" { + let control = @src.yn00_parse_control( + "* header\n\nseqfile = a.phy * input\noutfile = out.txt * output\n", + ) + assert_eq(control.sequence_file(), "a.phy") + assert_eq(control.output_file(), "out.txt") +} + +///| +test "yn00 control accepts CRLF" { + let text = @src.yn00_example_control_text().replace_all(old="\n", new="\r\n") + let control = @src.yn00_parse_control(text) + assert_eq(control.options().length(), 5) +} + +///| +test "yn00 control canonical round trip" { + let control = @src.yn00_parse_control(@src.yn00_example_control_text()) + let canonical = @src.yn00_write_control(control) + assert_eq( + @src.yn00_write_control(@src.yn00_parse_control(canonical)), + canonical, + ) +} + +///| +test "yn00 control preserves option order" { + let control = @src.yn00_parse_control( + "seqfile=a\noutfile=b\nndata=2\nverbose=0\nicode=3\n", + ) + let options = control.options() + assert_eq(options[0].name(), "ndata") + assert_eq(options[1].name(), "verbose") + assert_eq(options[2].name(), "icode") +} + +///| +test "yn00 control options are defensive copies" { + let control = @src.yn00_parse_control(@src.yn00_example_control_text()) + let options = control.options() + ignore(options.pop()) + assert_eq(control.options().length(), 5) +} + +///| +test "yn00 control reports absent option" { + let control = @src.yn00_parse_control("seqfile=a\noutfile=b\n") + assert_eq(control.option("icode"), None) +} + +///| +test "yn00 control normalizes integers" { + let control = @src.yn00_parse_control( + "seqfile=a\noutfile=b\nverbose=+1\nicode=03\nndata=002\n", + ) + assert_eq(control.option("verbose"), Some("1")) + assert_eq(control.option("icode"), Some("3")) + assert_eq(control.option("ndata"), Some("2")) +} + +///| +test "yn00 control rejects missing seqfile" { + yn00_test_expect_control_error("outfile=out\n") +} + +///| +test "yn00 control rejects missing outfile" { + yn00_test_expect_control_error("seqfile=in\n") +} + +///| +test "yn00 control rejects line without equals" { + yn00_test_expect_control_error("seqfile=in\noutfile out\n") +} + +///| +test "yn00 control rejects multiple equals" { + yn00_test_expect_control_error("seqfile=in=x\noutfile=out\n") +} + +///| +test "yn00 control rejects empty key" { + yn00_test_expect_control_error("=in\noutfile=out\n") +} + +///| +test "yn00 control rejects empty value" { + yn00_test_expect_control_error("seqfile=\noutfile=out\n") +} + +///| +test "yn00 control rejects duplicate key" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nseqfile=again\n") +} + +///| +test "yn00 control rejects unknown option" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nunknown=1\n") +} + +///| +test "yn00 control rejects non-integer option" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nicode=1.5\n") +} + +///| +test "yn00 control rejects verbose outside binary range" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nverbose=2\n") +} + +///| +test "yn00 control rejects weighting outside binary range" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nweighting=-1\n") +} + +///| +test "yn00 control rejects commonf3x4 outside binary range" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\ncommonf3x4=3\n") +} + +///| +test "yn00 control rejects negative genetic code" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nicode=-1\n") +} + +///| +test "yn00 control rejects unsupported genetic code" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nicode=11\n") +} + +///| +test "yn00 control rejects non-positive data count" { + yn00_test_expect_control_error("seqfile=in\noutfile=out\nndata=0\n") +} + +///| +test "yn00 results parse alignment metadata" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + assert_eq(result.alignment_file(), "alignment.phylip") + assert_eq(result.sequence_count(), 3) + assert_eq(result.codon_count(), 74) +} + +///| +test "yn00 results preserve sequence order" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + assert_eq(result.sequence_names(), ["Homo_sapie", "Pan_troglo", "Gorilla_go"]) +} + +///| +test "yn00 results contain complete triangular pairs" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + assert_eq(result.pairs().length(), 3) +} + +///| +test "yn00 pair lookup is symmetric" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let forward = result.pair("Homo_sapie", "Pan_troglo").unwrap() + let reverse = result.pair("Pan_troglo", "Homo_sapie").unwrap() + assert_eq(forward, reverse) +} + +///| +test "yn00 pair lookup rejects diagonal" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + assert_eq(result.pair("Homo_sapie", "Homo_sapie"), None) +} + +///| +test "yn00 pair lookup reports unknown names" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + assert_eq(result.pair("Homo_sapie", "Unknown"), None) +} + +///| +test "yn00 parses NG86 estimate" { + let pair = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + yn00_test_close(pair.ng86().omega(), -1.0, 1.0e-12) + yn00_test_close(pair.ng86().dn(), 0.0, 1.0e-12) + yn00_test_close(pair.ng86().ds(), 0.0207, 1.0e-12) +} + +///| +test "yn00 parses Yang-Nielsen site counts" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .yang_nielsen() + yn00_test_close(estimate.synonymous_sites(), 67.3, 1.0e-12) + yn00_test_close(estimate.nonsynonymous_sites(), 154.7, 1.0e-12) +} + +///| +test "yn00 parses Yang-Nielsen model parameters" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .yang_nielsen() + yn00_test_close(estimate.time(), 0.0136, 1.0e-12) + yn00_test_close(estimate.kappa(), 3.6564, 1.0e-12) + yn00_test_close(estimate.omega(), 0.0, 1.0e-12) +} + +///| +test "yn00 parses Yang-Nielsen rates and errors" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .yang_nielsen() + yn00_test_close(estimate.dn(), 0.0, 1.0e-12) + yn00_test_close(estimate.dn_standard_error(), 0.0, 1.0e-12) + yn00_test_close(estimate.ds(), 0.015, 1.0e-12) + yn00_test_close(estimate.ds_standard_error(), 0.0151, 1.0e-12) +} + +///| +test "yn00 parses LWL85 estimate" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .lwl85() + assert_eq(estimate.method_name(), "LWL85") + yn00_test_close(estimate.ds().unwrap(), 0.0227, 1.0e-12) + assert_eq(estimate.dn(), Some(0.0)) + assert_eq(estimate.omega(), Some(0.0)) + assert_eq(estimate.synonymous_sites(), Some(45.0)) + assert_eq(estimate.nonsynonymous_sites(), Some(177.0)) +} + +///| +test "yn00 maps undefined modified LWL85 values to none" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .lwl85_modified() + assert_eq(estimate.ds(), None) + assert_eq(estimate.dn(), None) + assert_eq(estimate.omega(), None) + assert_eq(estimate.synonymous_sites(), None) + assert_eq(estimate.nonsynonymous_sites(), None) + assert_eq(estimate.rho(), None) +} + +///| +test "yn00 parses defined modified LWL85 values" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Gorilla_go") + .unwrap() + .lwl85_modified() + yn00_test_close(estimate.ds().unwrap(), 0.018, 1.0e-12) + assert_eq(estimate.dn(), Some(0.011)) + assert_eq(estimate.omega(), Some(0.6111)) + assert_eq(estimate.rho(), Some(0.507)) +} + +///| +test "yn00 parses LPB93 estimate without site counts" { + let estimate = @src.yn00_parse_results(@src.yn00_example_results_text()) + .pair("Homo_sapie", "Pan_troglo") + .unwrap() + .lpb93() + assert_eq(estimate.method_name(), "LPB93") + assert_eq(estimate.ds(), Some(0.0129)) + assert_eq(estimate.dn(), Some(0.0)) + assert_eq(estimate.omega(), Some(0.0)) + assert_eq(estimate.synonymous_sites(), None) + assert_eq(estimate.nonsynonymous_sites(), None) +} + +///| +test "yn00 parses PAML 4.1 compact NG86 rows" { + let text = yn00_test_two_sequence_result( + "Alpha", "Beta", "Beta0.5000(0.0100 0.0200)", "0.5000", "-nan", + ) + let pair = @src.yn00_parse_results(text).pair("Alpha", "Beta").unwrap() + assert_eq(pair.ng86().omega(), 0.5) + assert_eq(pair.ng86().dn(), 0.01) +} + +///| +test "yn00 parses PAML 4.8 long names glued to negative omega" { + let first = "patient1_1000_326|1M|XXX|XXX|2" + let second = "patient1_1000_16|1M|XXX|XXX|20" + let text = yn00_test_two_sequence_result( + first, + second, + second + "-1.0000 (0.0028 0.0000)", + "inf", + "-nan", + ) + let result = @src.yn00_parse_results(text) + assert_eq(result.sequence_names(), [first, second]) + assert_eq(result.pair(first, second).unwrap().ng86().dn(), 0.0028) +} + +///| +test "yn00 parses PAML 4.9i dotted numeric names" { + let first = "patient1.1000.326" + let second = "patient1.1000.16" + let text = yn00_test_two_sequence_result( + first, + second, + second + " -1.0000 (0.0028 0.0000)", + "inf", + "-nan", + ) + let result = @src.yn00_parse_results(text) + assert_eq(result.sequence_names(), [first, second]) +} + +///| +test "yn00 maps infinity to none" { + let text = yn00_test_two_sequence_result( + "Alpha", "Beta", "Beta -1.0000 (0.0100 0.0000)", "inf", "-nan", + ) + let estimate = @src.yn00_parse_results(text) + .pair("Alpha", "Beta") + .unwrap() + .lwl85() + assert_eq(estimate.omega(), None) +} + +///| +test "yn00 maps Windows indefinite values to none" { + let text = yn00_test_two_sequence_result( + "Alpha", "Beta", "Beta -1.0000 (0.0100 0.0000)", "0.5000", "-1.#IND", + ) + let estimate = @src.yn00_parse_results(text) + .pair("Alpha", "Beta") + .unwrap() + .lwl85_modified() + assert_eq(estimate.ds(), None) +} + +///| +test "yn00 results accept CRLF" { + let text = @src.yn00_example_results_text().replace_all(old="\n", new="\r\n") + assert_eq(@src.yn00_parse_results(text).pairs().length(), 3) +} + +///| +test "yn00 sequence names are defensive copies" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let names = result.sequence_names() + names[0] = "Changed" + assert_eq(result.sequence_names()[0], "Homo_sapie") +} + +///| +test "yn00 pairs are defensive copies" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let pairs = result.pairs() + ignore(pairs.pop()) + assert_eq(result.pairs().length(), 3) +} + +///| +test "yn00 builds symmetric NG86 matrix" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let matrix = result.matrix("NG86", "dS") + assert_eq(matrix.names(), result.sequence_names()) + assert_eq(matrix.value("Homo_sapie", "Pan_troglo"), Some(0.0207)) + assert_eq(matrix.value("Pan_troglo", "Homo_sapie"), Some(0.0207)) + assert_eq(matrix.value("Homo_sapie", "Homo_sapie"), None) +} + +///| +test "yn00 builds Yang-Nielsen standard error matrix" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let matrix = result.matrix("YN00", "dS SE") + yn00_test_close( + matrix.value("Homo_sapie", "Pan_troglo").unwrap(), + 0.0151, + 1.0e-12, + ) +} + +///| +test "yn00 matrix preserves undefined values" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let matrix = result.matrix("LWL85m", "dS") + assert_eq(matrix.value("Homo_sapie", "Pan_troglo"), None) + yn00_test_close( + matrix.value("Homo_sapie", "Gorilla_go").unwrap(), + 0.018, + 1.0e-12, + ) +} + +///| +test "yn00 matrix values are defensive copies" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let matrix = result.matrix("NG86", "dS") + let values = matrix.values() + values[0][1] = None + assert_eq(matrix.value("Homo_sapie", "Pan_troglo"), Some(0.0207)) +} + +///| +test "yn00 matrix reports unknown name" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let matrix = result.matrix("NG86", "dS") + assert_eq(matrix.value("Unknown", "Homo_sapie"), None) +} + +///| +test "yn00 computes mean over all pairs" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + yn00_test_close( + result.mean("YN00", "dS"), + (0.015 + 0.02 + 0.0303) / 3.0, + 1.0e-12, + ) +} + +///| +test "yn00 mean skips undefined estimates" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + yn00_test_close(result.mean("LWL85m", "dS"), 0.018, 1.0e-12) +} + +///| +test "yn00 matrix rejects unknown method" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let raised = try { + ignore(result.matrix("UNKNOWN", "dS")) + false + } catch { + @src.Yn00Error(_) => true + } + assert_true(raised) +} + +///| +test "yn00 matrix rejects unknown statistic" { + let result = @src.yn00_parse_results(@src.yn00_example_results_text()) + let raised = try { + ignore(result.matrix("YN00", "bad")) + false + } catch { + @src.Yn00Error(_) => true + } + assert_true(raised) +} + +///| +test "yn00 mean rejects wholly undefined statistic" { + let text = yn00_test_two_sequence_result( + "Alpha", "Beta", "Beta -1.0000 (0.0100 0.0000)", "0.5000", "-nan", + ) + let result = @src.yn00_parse_results(text) + let raised = try { + ignore(result.mean("LWL85m", "rho")) + false + } catch { + @src.Yn00Error(_) => true + } + assert_true(raised) +} + +///| +test "yn00 rejects empty results" { + yn00_test_expect_result_error("") +} + +///| +test "yn00 rejects missing NG86 section" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="(A) Nei-Gojobori (1986) method", + new="", + ), + ) +} + +///| +test "yn00 rejects missing Yang-Nielsen section" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="(B) Yang & Nielsen (2000) method", + new="", + ), + ) +} + +///| +test "yn00 rejects missing counting section" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="(C) LWL85, LPB93 & LWLm methods", + new="", + ), + ) +} + +///| +test "yn00 rejects duplicate method section" { + yn00_test_expect_result_error( + @src.yn00_example_results_text() + "(A) Nei-Gojobori (1986) method\n", + ) +} + +///| +test "yn00 rejects missing alignment header" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="YN00 alignment.phylip\n", + new="", + ), + ) +} + +///| +test "yn00 rejects duplicate alignment header" { + yn00_test_expect_result_error( + "YN00 duplicate.phy\n" + @src.yn00_example_results_text(), + ) +} + +///| +test "yn00 rejects malformed dimensions" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="ns = 3 ls = 74", + new="ns = bad ls = 74", + ), + ) +} + +///| +test "yn00 rejects fewer than two sequences" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="ns = 3 ls = 74", + new="ns = 1 ls = 74", + ), + ) +} + +///| +test "yn00 rejects non-positive codon count" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="ns = 3 ls = 74", + new="ns = 3 ls = 0", + ), + ) +} + +///| +test "yn00 rejects invalid comparison index" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="2 (Pan_troglo) vs. 1 (Homo_sapie)", + new="4 (Pan_troglo) vs. 1 (Homo_sapie)", + ), + ) +} + +///| +test "yn00 rejects inconsistent indexed name" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="3 (Gorilla_go) vs. 2 (Pan_troglo)", + new="3 (Different) vs. 2 (Pan_troglo)", + ), + ) +} + +///| +test "yn00 rejects duplicate sequence names" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="2 (Pan_troglo) vs. 1 (Homo_sapie)", + new="2 (Homo_sapie) vs. 1 (Homo_sapie)", + ), + ) +} + +///| +test "yn00 rejects duplicate pair header" { + let duplicate = "2 (Pan_troglo) vs. 1 (Homo_sapie)\n" + yn00_test_expect_result_error(@src.yn00_example_results_text() + duplicate) +} + +///| +test "yn00 rejects counting estimate before pair header" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="(C) LWL85, LPB93 & LWLm methods\n", + new="(C) LWL85, LPB93 & LWLm methods\n" + + "LWL85: dS = 0.1 dN = 0.1 w = 1.0 S = 1.0 N = 1.0\n", + ), + ) +} + +///| +test "yn00 rejects missing counting method" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="LPB93: dS = 0.0129 dN = 0.0000 w = 0.0000\n", + new="", + ), + ) +} + +///| +test "yn00 rejects missing counting statistic" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="LWL85: dS = 0.0227 dN = 0.0000 w = 0.0000 S = 45.0 N = 177.0", + new="LWL85: dS = 0.0227 dN = 0.0000 S = 45.0 N = 177.0", + ), + ) +} + +///| +test "yn00 rejects duplicate counting estimate" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="LWL85: dS = 0.0227 dN = 0.0000 w = 0.0000 S = 45.0 N = 177.0\n", + new="LWL85: dS = 0.0227 dN = 0.0000 w = 0.0000 S = 45.0 N = 177.0\n" + + "LWL85: dS = 0.0227 dN = 0.0000 w = 0.0000 S = 45.0 N = 177.0\n", + ), + ) +} + +///| +test "yn00 rejects malformed Yang-Nielsen row" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="2 1 67.3 154.7 0.0136 3.6564 0.0000 0.0000 +- 0.0000 0.0150 +- 0.0151", + new="2 1 67.3 154.7 0.0136 3.6564", + ), + ) +} + +///| +test "yn00 rejects invalid Yang-Nielsen index" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="2 1 67.3 154.7", + new="4 1 67.3 154.7", + ), + ) +} + +///| +test "yn00 rejects duplicate Yang-Nielsen pair" { + let row = "2 1 67.3 154.7 0.0136 3.6564 0.0000 0.0000 +- 0.0000 0.0150 +- 0.0151\n" + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="(C) LWL85, LPB93 & LWLm methods", + new=row + "(C) LWL85, LPB93 & LWLm methods", + ), + ) +} + +///| +test "yn00 rejects missing NG86 row" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="Gorilla_go 0.5000 (0.0100 0.0200)-1.0000 (0.0000 0.0421)\n", + new="", + ), + ) +} + +///| +test "yn00 rejects incomplete NG86 row" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="Gorilla_go 0.5000 (0.0100 0.0200)-1.0000 (0.0000 0.0421)", + new="Gorilla_go 0.5000 (0.0100 0.0200)", + ), + ) +} + +///| +test "yn00 rejects malformed NG86 estimate" { + yn00_test_expect_result_error( + @src.yn00_example_results_text().replace_all( + old="Pan_troglo -1.0000 (0.0000 0.0207)", + new="Pan_troglo -1.0000 (0.0000)", + ), + ) +} diff --git a/test/moonbit/parsimony_test.mbt b/test/moonbit/parsimony_test.mbt index b1a041c9..2dda34b2 100644 --- a/test/moonbit/parsimony_test.mbt +++ b/test/moonbit/parsimony_test.mbt @@ -41,7 +41,10 @@ test "parsimony_matrix_sankoff_dna_construction" { ///| test "parsimony_matrix_sankoff_custom_costs" { - let m = @src.parsimony_sankoff_dna_matrix(transition_cost=0.5, transversion_cost=2.5) + let m = @src.parsimony_sankoff_dna_matrix( + transition_cost=0.5, + transversion_cost=2.5, + ) // A->G (transition) = 0.5 assert_eq(@src.parsimony_matrix_cost(m, 0, 2), 0.5) // A->C (transversion) = 2.5 @@ -238,7 +241,10 @@ test "sankoff_parsimony_lower_than_fitch_for_transitions" { let root = @src.Clade::new(clades=[inner_left, inner_right]) let tree = @src.Tree::new(root, rooted=true) let fitch_score = @src.fitch_parsimony_score(tree, aln) - let matrix = @src.parsimony_sankoff_dna_matrix(transition_cost=1.0, transversion_cost=2.0) + let matrix = @src.parsimony_sankoff_dna_matrix( + transition_cost=1.0, + transversion_cost=2.0, + ) let sankoff_score = @src.sankoff_parsimony_score(tree, aln, matrix) // All changes are A<->G (transitions), so Sankoff = Fitch here assert_eq(sankoff_score, fitch_score) diff --git a/test/moonbit/pathway_test.mbt b/test/moonbit/pathway_test.mbt index 4647ee6a..7c359eb6 100644 --- a/test/moonbit/pathway_test.mbt +++ b/test/moonbit/pathway_test.mbt @@ -11,7 +11,9 @@ test "Reaction_new" { reactants.push("glucose") let products : Array[String] = Array::new() products.push("g6p") - let reaction = @src.Reaction::new("r1", "Hexokinase", reactants, products, false) + let reaction = @src.Reaction::new( + "r1", "Hexokinase", reactants, products, false, + ) assert_eq(reaction.id, "r1") assert_eq(reaction.name, "Hexokinase") assert_eq(reaction.reactants.length(), 1) @@ -23,14 +25,16 @@ test "Reaction_new" { test "Pathway_new" { let species : Array[@src.Species] = Array::new() species.push(@src.Species::new("glucose", "Glucose")) - + let reactants : Array[String] = Array::new() reactants.push("glucose") let products : Array[String] = Array::new() products.push("g6p") let reactions : Array[@src.Reaction] = Array::new() - reactions.push(@src.Reaction::new("r1", "Hexokinase", reactants, products, false)) - + reactions.push( + @src.Reaction::new("r1", "Hexokinase", reactants, products, false), + ) + let pathway = @src.Pathway::new("test", "Test Pathway", species, reactions) assert_eq(pathway.id, "test") assert_eq(pathway.name, "Test Pathway") @@ -90,4 +94,4 @@ test "Reaction_to_string" { let reaction = @src.Reaction::new("r1", "Test", reactants, products, true) let str = @src.Reaction::to_string(reaction) assert_true(str.length() > 0) -} \ No newline at end of file +} diff --git a/test/moonbit/pcatools_test.mbt b/test/moonbit/pcatools_test.mbt index bacd7359..4d9cba19 100644 --- a/test/moonbit/pcatools_test.mbt +++ b/test/moonbit/pcatools_test.mbt @@ -10,12 +10,7 @@ test "biplot_options_new" { ///| test "pcatools_run_pca_simple_2d" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - [4.0, 5.0], - ] + let data = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0], [4.0, 5.0]] let res = @src.pcatools_run_pca(data, 2, false) assert_eq(res.n_samples, 4) assert_eq(res.n_variables, 2) @@ -30,11 +25,7 @@ test "pcatools_run_pca_simple_2d" { ///| test "pcatools_run_pca_ncomponents_limit" { - let data = [ - [1.0, 2.0, 3.0, 4.0], - [2.0, 3.0, 4.0, 5.0], - [3.0, 4.0, 5.0, 6.0], - ] + let data = [[1.0, 2.0, 3.0, 4.0], [2.0, 3.0, 4.0, 5.0], [3.0, 4.0, 5.0, 6.0]] let res = @src.pcatools_run_pca(data, 10, false) // 3 samples, 4 variables => p=4, so limited to 4 assert_eq(res.n_components, 4) @@ -42,12 +33,7 @@ test "pcatools_run_pca_ncomponents_limit" { ///| test "pcatools_run_pca_scale_true" { - let data = [ - [1.0, 100.0], - [2.0, 200.0], - [3.0, 300.0], - [4.0, 400.0], - ] + let data = [[1.0, 100.0], [2.0, 200.0], [3.0, 300.0], [4.0, 400.0]] let res = @src.pcatools_run_pca(data, 2, true) assert_eq(res.used_scaling, true) assert_eq(res.scale.length(), 2) @@ -82,16 +68,10 @@ test "scree_plot_ascii_returns_string" { ///| test "pca_biplot_ascii_returns_string" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 5.0], - [4.0, 4.0], - [5.0, 6.0], - ] + let data = [[1.0, 2.0], [2.0, 3.0], [3.0, 5.0], [4.0, 4.0], [5.0, 6.0]] let res = @src.pcatools_run_pca(data, 2, false) let opts = @src.BiplotOptions::new() - let plot = @src.pca_biplot_ascii(res, opts=opts) + let plot = @src.pca_biplot_ascii(res, opts~) assert_true(plot.length() > 0) assert_true(plot.contains("PCA biplot")) assert_true(plot.contains("PC1")) @@ -116,7 +96,9 @@ test "find_pca_outliers_basic" { assert_true(out.indices.length() >= 1) let mut found7 = false for i in out.indices { - if i == 7 { found7 = true } + if i == 7 { + found7 = true + } } assert_true(found7) } @@ -124,8 +106,11 @@ test "find_pca_outliers_basic" { ///| test "find_pca_outliers_cutoff_positive" { let data = [ - [1.0, 1.0], [1.1, 1.0], [0.9, 1.0], - [1.0, 1.1], [1.0, 0.9], + [1.0, 1.0], + [1.1, 1.0], + [0.9, 1.0], + [1.0, 1.1], + [1.0, 0.9], [20.0, 20.0], ] let res = @src.pcatools_run_pca(data, 2, false) @@ -164,11 +149,7 @@ test "variable_correlations_shape" { ///| test "pcatools_summary_contains_keywords" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [3.0, 4.0], - ] + let data = [[1.0, 2.0], [2.0, 3.0], [3.0, 4.0]] let res = @src.pcatools_run_pca(data, 2, false) let s = @src.pcatools_summary(res) assert_true(s.contains("PCAResult")) diff --git a/test/moonbit/pcd_test.mbt b/test/moonbit/pcd_test.mbt index f35de6d1..396fc638 100644 --- a/test/moonbit/pcd_test.mbt +++ b/test/moonbit/pcd_test.mbt @@ -12,6 +12,7 @@ test "pcd_spectrum_construction" { assert_eq(spec.num_peaks(), 0) } +///| test "pcd_spectrum_add_peak" { let spec = @src.PcdSpectrum::new(1) spec.add_peak(500.5, 10000.0) @@ -20,6 +21,7 @@ test "pcd_spectrum_add_peak" { assert_true(spec.total_intensity() > 0.0) } +///| test "pcd_spectrum_precursor" { let spec = @src.PcdSpectrum::new(1) spec.set_precursor_mz(Some(500.5)) @@ -38,12 +40,14 @@ test "pcd_spectrum_precursor" { // PcdFile Tests // ============================================================================ +///| test "pcd_file_construction" { let file = @src.PcdFile::new() assert_eq(file.num_spectra, 0) assert_eq(file.num_peaks, 0) } +///| test "pcd_file_add_spectrum" { let file = @src.PcdFile::new() let spec = @src.PcdSpectrum::new(1, rt=10.0) @@ -54,6 +58,7 @@ test "pcd_file_add_spectrum" { assert_eq(file.num_peaks, 1) } +///| test "pcd_file_get_spectrum" { let file = @src.PcdFile::new() let spec1 = @src.PcdSpectrum::new(1) @@ -74,12 +79,14 @@ test "pcd_file_get_spectrum" { // PCD Parsing Tests // ============================================================================ +///| test "pcd_parse_empty" { let file = @src.pcd_parse("") assert_eq(file.num_spectra, 0) assert_eq(file.num_peaks, 0) } +///| test "pcd_parse_sample" { let content = @src.pcd_sample() let file = @src.pcd_parse(content) @@ -87,6 +94,7 @@ test "pcd_parse_sample" { assert_true(file.num_peaks >= 1) } +///| test "pcd_parse_with_precursor" { let content = @src.pcd_sample() let file = @src.pcd_parse(content) @@ -97,6 +105,7 @@ test "pcd_parse_with_precursor" { // PCD Query Tests // ============================================================================ +///| test "pcd_total_peaks" { let file = @src.PcdFile::new() let spec1 = @src.PcdSpectrum::new(1) @@ -110,6 +119,7 @@ test "pcd_total_peaks" { assert_eq(@src.pcd_total_peaks(file), 3) } +///| test "pcd_bpc" { let file = @src.PcdFile::new() let spec1 = @src.PcdSpectrum::new(1, rt=10.0) @@ -125,6 +135,7 @@ test "pcd_bpc" { assert_eq(bpc.length(), 2) } +///| test "pcd_tic" { let file = @src.PcdFile::new() let spec1 = @src.PcdSpectrum::new(1, rt=10.0) @@ -139,6 +150,7 @@ test "pcd_tic" { assert_eq(tic.length(), 2) } +///| test "pcd_spectrum_by_mz" { let spec = @src.PcdSpectrum::new(1) spec.add_peak(500.0, 1000.0) @@ -152,6 +164,7 @@ test "pcd_spectrum_by_mz" { // PCD Serialization Tests // ============================================================================ +///| test "pcd_write_roundtrip" { let content = @src.pcd_sample() let file = @src.pcd_parse(content) @@ -164,6 +177,7 @@ test "pcd_write_roundtrip" { // PCD Summary Tests // ============================================================================ +///| test "pcd_summary" { let file = @src.PcdFile::new() let spec = @src.PcdSpectrum::new(1) diff --git a/test/moonbit/pdb_analysis_test.mbt b/test/moonbit/pdb_analysis_test.mbt index 376b8ef7..fe4c3521 100644 --- a/test/moonbit/pdb_analysis_test.mbt +++ b/test/moonbit/pdb_analysis_test.mbt @@ -184,9 +184,9 @@ test "pdb_analysis: ramachandran_quality" { test "pdb_analysis: get_hydrophobicity" { // Test some known values from Kyte-Doolittle scale assert_true((@src.get_hydrophobicity("ALA") - 1.8).abs() < 0.001) - assert_true((@src.get_hydrophobicity("ARG") - (-4.5)).abs() < 0.001) + assert_true((@src.get_hydrophobicity("ARG") - -4.5).abs() < 0.001) assert_true((@src.get_hydrophobicity("ILE") - 4.5).abs() < 0.001) - assert_true((@src.get_hydrophobicity("LYS") - (-3.9)).abs() < 0.001) + assert_true((@src.get_hydrophobicity("LYS") - -3.9).abs() < 0.001) assert_true((@src.get_hydrophobicity("PHE") - 2.8).abs() < 0.001) assert_true((@src.get_hydrophobicity("VAL") - 4.2).abs() < 0.001) // Unknown residue should return 0 diff --git a/test/moonbit/pdb_dice_test.mbt b/test/moonbit/pdb_dice_test.mbt index 6aa87729..62b94379 100644 --- a/test/moonbit/pdb_dice_test.mbt +++ b/test/moonbit/pdb_dice_test.mbt @@ -8,34 +8,131 @@ fn create_test_structure() -> @src.Structure { // Create a simple structure with 2 chains, each with 3 residues let atoms_a1 : Array[@src.Atom] = [ - @src.Atom::new(name="N", coord=@src.Vector3::new(0.0, 0.0, 0.0), resname="ALA", chainid='A', resseq=1), - @src.Atom::new(name="CA", coord=@src.Vector3::new(1.5, 0.0, 0.0), resname="ALA", chainid='A', resseq=1), - @src.Atom::new(name="C", coord=@src.Vector3::new(2.5, 1.0, 0.0), resname="ALA", chainid='A', resseq=1), - @src.Atom::new(name="O", coord=@src.Vector3::new(2.5, 2.0, 0.0), resname="ALA", chainid='A', resseq=1), + @src.Atom::new( + name="N", + coord=@src.Vector3::new(0.0, 0.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + ), + @src.Atom::new( + name="CA", + coord=@src.Vector3::new(1.5, 0.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + ), + @src.Atom::new( + name="C", + coord=@src.Vector3::new(2.5, 1.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + ), + @src.Atom::new( + name="O", + coord=@src.Vector3::new(2.5, 2.0, 0.0), + resname="ALA", + chainid='A', + resseq=1, + ), ] let atoms_a2 : Array[@src.Atom] = [ - @src.Atom::new(name="N", coord=@src.Vector3::new(3.5, 0.0, 0.0), resname="GLY", chainid='A', resseq=2), - @src.Atom::new(name="CA", coord=@src.Vector3::new(5.0, 0.0, 0.0), resname="GLY", chainid='A', resseq=2), + @src.Atom::new( + name="N", + coord=@src.Vector3::new(3.5, 0.0, 0.0), + resname="GLY", + chainid='A', + resseq=2, + ), + @src.Atom::new( + name="CA", + coord=@src.Vector3::new(5.0, 0.0, 0.0), + resname="GLY", + chainid='A', + resseq=2, + ), ] let atoms_a3 : Array[@src.Atom] = [ - @src.Atom::new(name="N", coord=@src.Vector3::new(6.0, 0.0, 0.0), resname="VAL", chainid='A', resseq=3), - @src.Atom::new(name="CA", coord=@src.Vector3::new(7.5, 0.0, 0.0), resname="VAL", chainid='A', resseq=3), + @src.Atom::new( + name="N", + coord=@src.Vector3::new(6.0, 0.0, 0.0), + resname="VAL", + chainid='A', + resseq=3, + ), + @src.Atom::new( + name="CA", + coord=@src.Vector3::new(7.5, 0.0, 0.0), + resname="VAL", + chainid='A', + resseq=3, + ), ] - let res_a1 = @src.Residue::new(resname="ALA", chainid='A', resseq=1, atoms=atoms_a1) - let res_a2 = @src.Residue::new(resname="GLY", chainid='A', resseq=2, atoms=atoms_a2) - let res_a3 = @src.Residue::new(resname="VAL", chainid='A', resseq=3, atoms=atoms_a3) + let res_a1 = @src.Residue::new( + resname="ALA", + chainid='A', + resseq=1, + atoms=atoms_a1, + ) + let res_a2 = @src.Residue::new( + resname="GLY", + chainid='A', + resseq=2, + atoms=atoms_a2, + ) + let res_a3 = @src.Residue::new( + resname="VAL", + chainid='A', + resseq=3, + atoms=atoms_a3, + ) let chain_a = @src.Chain::new(id='A', residues=[res_a1, res_a2, res_a3]) let atoms_b1 : Array[@src.Atom] = [ - @src.Atom::new(name="N", coord=@src.Vector3::new(0.0, 5.0, 0.0), resname="SER", chainid='B', resseq=10), - @src.Atom::new(name="CA", coord=@src.Vector3::new(1.5, 5.0, 0.0), resname="SER", chainid='B', resseq=10), + @src.Atom::new( + name="N", + coord=@src.Vector3::new(0.0, 5.0, 0.0), + resname="SER", + chainid='B', + resseq=10, + ), + @src.Atom::new( + name="CA", + coord=@src.Vector3::new(1.5, 5.0, 0.0), + resname="SER", + chainid='B', + resseq=10, + ), ] let atoms_b2 : Array[@src.Atom] = [ - @src.Atom::new(name="N", coord=@src.Vector3::new(3.5, 5.0, 0.0), resname="THR", chainid='B', resseq=11), - @src.Atom::new(name="CA", coord=@src.Vector3::new(5.0, 5.0, 0.0), resname="THR", chainid='B', resseq=11), + @src.Atom::new( + name="N", + coord=@src.Vector3::new(3.5, 5.0, 0.0), + resname="THR", + chainid='B', + resseq=11, + ), + @src.Atom::new( + name="CA", + coord=@src.Vector3::new(5.0, 5.0, 0.0), + resname="THR", + chainid='B', + resseq=11, + ), ] - let res_b1 = @src.Residue::new(resname="SER", chainid='B', resseq=10, atoms=atoms_b1) - let res_b2 = @src.Residue::new(resname="THR", chainid='B', resseq=11, atoms=atoms_b2) + let res_b1 = @src.Residue::new( + resname="SER", + chainid='B', + resseq=10, + atoms=atoms_b1, + ) + let res_b2 = @src.Residue::new( + resname="THR", + chainid='B', + resseq=11, + atoms=atoms_b2, + ) let chain_b = @src.Chain::new(id='B', residues=[res_b1, res_b2]) let model = @src.Model::new(id=0, chains=[chain_a, chain_b]) @@ -144,10 +241,7 @@ test "extract_residue_range_chain_b" { ///| test "extract_residues_by_ids" { let structure = create_test_structure() - let ids = [ - @src.ResidueId::new('A', 1), - @src.ResidueId::new('B', 10), - ] + let ids = [@src.ResidueId::new('A', 1), @src.ResidueId::new('B', 10)] let result = @src.extract_residues(structure, ids) // Should have 2 chains, each with 1 residue assert_eq(result.models[0].chains.length(), 2) diff --git a/test/moonbit/pdb_header_test.mbt b/test/moonbit/pdb_header_test.mbt index 966ce345..edf9fe5b 100644 --- a/test/moonbit/pdb_header_test.mbt +++ b/test/moonbit/pdb_header_test.mbt @@ -69,9 +69,7 @@ test "dbref_entry_creation" { ///| test "parse_title_record_single_line" { - let lines = [ - "TITLE Crystal structure of a protein", - ] + let lines = ["TITLE Crystal structure of a protein"] let title = @src.parse_title_record(lines, 0) assert_true(title.contains("Crystal structure")) } @@ -79,8 +77,7 @@ test "parse_title_record_single_line" { ///| test "parse_title_record_multi_line" { let lines = [ - "TITLE Crystal structure of a protein", - "TITLE in complex with ligand", + "TITLE Crystal structure of a protein", "TITLE in complex with ligand", ] let title = @src.parse_title_record(lines, 0) assert_true(title.contains("Crystal structure")) @@ -126,10 +123,8 @@ test "parse_rfactor_record_no_keyword" { ///| test "extract_chain_ids" { - let compound = @src.CompoundInfo::new( - chain_ids=["A", "B", "C"], - ) - let header = @src.PDBHeader::new(compound=compound) + let compound = @src.CompoundInfo::new(chain_ids=["A", "B", "C"]) + let header = @src.PDBHeader::new(compound~) let chains = @src.extract_chain_ids(header) assert_eq(chains.length(), 3) assert_eq(chains[0], "A") @@ -155,8 +150,7 @@ test "parse_pdb_header_empty" { ///| test "parse_pdb_header_missing_records" { - let pdb_text = - "HEADER VIRUS 01-JAN-20 1ABC\n" + let pdb_text = "HEADER VIRUS 01-JAN-20 1ABC\n" let header = @src.parse_pdb_header(pdb_text) assert_eq(header.title, "") assert_eq(header.compound.molecule_id, "") @@ -172,8 +166,7 @@ test "parse_pdb_header_missing_records" { ///| test "parse_pdb_header_lines" { let lines = [ - "HEADER VIRUS 01-JAN-20 1ABC", - "TITLE Crystal structure", + "HEADER VIRUS 01-JAN-20 1ABC", "TITLE Crystal structure", ] let header = @src.parse_pdb_header_lines(lines) assert_true(header.title.length() > 0) diff --git a/test/moonbit/pdb_list_test.mbt b/test/moonbit/pdb_list_test.mbt index a3518738..ef753b2e 100644 --- a/test/moonbit/pdb_list_test.mbt +++ b/test/moonbit/pdb_list_test.mbt @@ -1,31 +1,34 @@ ///| /// Tests for PDBList module. - test "PDBList creation" { let pdblist = @src.create_example_pdblist() assert_eq(pdblist.pdb_dir, "/data/pdb") } +///| test "PDBList download_pdb" { let pdblist = @src.PDBList::new("/data/pdb") let file = pdblist.download_pdb("1XYZ", "pdb") assert_true(file.contains("1XYZ")) } +///| test "PDBList get_pdb_file" { let pdblist = @src.PDBList::new("/data/pdb") let file = pdblist.get_pdb_file("1XYZ", "pdb", false) assert_true(file.contains("1XYZ")) } +///| test "PDBList get_pdb_file obsolete" { let pdblist = @src.PDBList::new("/data/pdb") let file = pdblist.get_pdb_file("1XYZ", "pdb", true) assert_true(file.contains("obsolete")) } +///| test "PDBList resolve_obsolete" { let pdblist = @src.PDBList::new("/data/pdb") let resolved = pdblist.resolve_obsolete("1XYZ") assert_eq(resolved, "2XYZ") -} \ No newline at end of file +} diff --git a/test/moonbit/pdb_packing_test.mbt b/test/moonbit/pdb_packing_test.mbt index 635448ff..35c2d227 100644 --- a/test/moonbit/pdb_packing_test.mbt +++ b/test/moonbit/pdb_packing_test.mbt @@ -99,7 +99,11 @@ test "packing_result_creation" { ///| test "packing_result_get_density" { - let result = @src.PackingResult::new(residue_name="GLU", residue_seq=5, density=3.7) + let result = @src.PackingResult::new( + residue_name="GLU", + residue_seq=5, + density=3.7, + ) assert_eq(result.get_density(), 3.7) } @@ -321,12 +325,7 @@ test "calculate_packing_sasa_empty" { ///| test "calculate_packing_sasa_single_atom" { let atoms : Array[@src.PackingAtom] = [ - @src.PackingAtom::new( - x=0.0, - y=0.0, - z=0.0, - vdw_radius=1.7, - ), + @src.PackingAtom::new(x=0.0, y=0.0, z=0.0, vdw_radius=1.7), ] let sasa = @src.calculate_packing_sasa(atoms, probe_radius=1.4, n_points=50) assert_eq(sasa > 0.0, true) @@ -349,28 +348,44 @@ test "calculate_packing_sasa_lower_probe" { ///| test "normalize_packing_density_below_threshold" { - let result = @src.PackingResult::new(residue_name="ALA", residue_seq=1, density=0.5) + let result = @src.PackingResult::new( + residue_name="ALA", + residue_seq=1, + density=0.5, + ) let normalized = @src.normalize_packing_density(result, threshold=2.0) assert_eq(normalized.density, 0.25) } ///| test "normalize_packing_density_above_threshold" { - let result = @src.PackingResult::new(residue_name="ALA", residue_seq=1, density=3.0) + let result = @src.PackingResult::new( + residue_name="ALA", + residue_seq=1, + density=3.0, + ) let normalized = @src.normalize_packing_density(result, threshold=2.0) assert_eq(normalized.density, 1.0) } ///| test "normalize_packing_density_at_threshold" { - let result = @src.PackingResult::new(residue_name="ALA", residue_seq=1, density=2.0) + let result = @src.PackingResult::new( + residue_name="ALA", + residue_seq=1, + density=2.0, + ) let normalized = @src.normalize_packing_density(result, threshold=2.0) assert_eq(normalized.density, 1.0) } ///| test "normalize_packing_density_zero_threshold" { - let result = @src.PackingResult::new(residue_name="ALA", residue_seq=1, density=1.5) + let result = @src.PackingResult::new( + residue_name="ALA", + residue_seq=1, + density=1.5, + ) let normalized = @src.normalize_packing_density(result, threshold=0.0) assert_eq(normalized.density, 1.5) } @@ -378,9 +393,24 @@ test "normalize_packing_density_zero_threshold" { ///| test "identify_low_packing_basic" { let results : Array[@src.PackingResult] = [ - @src.PackingResult::new(residue_name="ALA", residue_seq=1, chain_id='A', density=0.3), - @src.PackingResult::new(residue_name="GLY", residue_seq=2, chain_id='A', density=1.5), - @src.PackingResult::new(residue_name="VAL", residue_seq=3, chain_id='A', density=0.4), + @src.PackingResult::new( + residue_name="ALA", + residue_seq=1, + chain_id='A', + density=0.3, + ), + @src.PackingResult::new( + residue_name="GLY", + residue_seq=2, + chain_id='A', + density=1.5, + ), + @src.PackingResult::new( + residue_name="VAL", + residue_seq=3, + chain_id='A', + density=0.4, + ), ] let analysis = @src.PackingAnalysisResult::new(results~, n_residues=3) let low = @src.identify_low_packing(analysis, cutoff=0.5) @@ -412,7 +442,7 @@ test "identify_low_packing_all_below" { ///| test "sphere_volume_calculation" { let vol = @src.sphere_volume(2.0) - let expected = (4.0 / 3.0) * 3.14159265358979323846 * 8.0 + let expected = 4.0 / 3.0 * 3.14159265358979323846 * 8.0 assert_eq((vol - expected).abs() < 0.001, true) } @@ -445,7 +475,10 @@ test "sphere_volume_radius_roundtrip" { ///| test "calculate_packing_efficiency_basic" { let atoms = @src.create_demo_packing_atoms() - let efficiency = @src.calculate_packing_efficiency(atoms, container_radius=10.0) + let efficiency = @src.calculate_packing_efficiency( + atoms, + container_radius=10.0, + ) assert_eq(efficiency > 0.0, true) } @@ -460,7 +493,10 @@ test "calculate_packing_efficiency_empty" { test "calculate_packing_efficiency_large_container" { let atoms = @src.create_demo_packing_atoms() let eff_small = @src.calculate_packing_efficiency(atoms, container_radius=5.0) - let eff_large = @src.calculate_packing_efficiency(atoms, container_radius=50.0) + let eff_large = @src.calculate_packing_efficiency( + atoms, + container_radius=50.0, + ) assert_eq(eff_small > eff_large, true) } @@ -524,7 +560,11 @@ test "packing_density_with_demo_atoms" { ///| test "packing_analysis_with_demo_atoms" { let atoms = @src.create_demo_packing_atoms() - let analysis = @src.packing_density_per_residue(atoms, radius=1.4, n_points=20) + let analysis = @src.packing_density_per_residue( + atoms, + radius=1.4, + n_points=20, + ) assert_eq(analysis.n_residues, 3) assert_eq(analysis.mean_density > 0.0, true) assert_eq(analysis.n_buried + analysis.n_exposed, 3) @@ -540,7 +580,11 @@ test "sasa_with_demo_atoms" { ///| test "low_packing_identification_with_analysis" { let atoms = @src.create_demo_packing_atoms() - let analysis = @src.packing_density_per_residue(atoms, radius=1.4, n_points=20) + let analysis = @src.packing_density_per_residue( + atoms, + radius=1.4, + n_points=20, + ) let low = @src.identify_low_packing(analysis, cutoff=0.0) assert_eq(low.length() >= 0, true) } @@ -582,4 +626,4 @@ test "packing_result_default_creation" { assert_eq(result.density, 0.0) assert_eq(result.n_contacting_atoms, 0) assert_eq(result.n_shell_points, 0) -} \ No newline at end of file +} diff --git a/test/moonbit/pdb_seqio_test.mbt b/test/moonbit/pdb_seqio_test.mbt index abaca152..4d7eb1d0 100644 --- a/test/moonbit/pdb_seqio_test.mbt +++ b/test/moonbit/pdb_seqio_test.mbt @@ -165,7 +165,8 @@ test "pdb_atom_parser_basic" { ///| test "pdb_atom_parser_with_pdb_id" { let header = "HEADER PROTEIN" + " ".repeat(33) + "01-JAN-24 1XXX" - let pdb_text = header + "\nATOM 1 N ALA A 1 10.000 20.000 30.000 1.00 20.00 N \nATOM 2 N GLY A 2 12.000 22.000 32.000 1.00 20.00 N \nTER 3 GLY A 2\nEND\n" + let pdb_text = header + + "\nATOM 1 N ALA A 1 10.000 20.000 30.000 1.00 20.00 N \nATOM 2 N GLY A 2 12.000 22.000 32.000 1.00 20.00 N \nTER 3 GLY A 2\nEND\n" let records = @src.pdb_atom_parser(pdb_text) assert_eq(records.length(), 1) assert_eq(records[0].seq.to_string(), "AG") @@ -216,11 +217,7 @@ test "pdb_write_pdb_seqrecords_round_trip" { ///| test "pdb_write_pdb_seqrecords_basic" { - let record = @src.SeqRecord::new( - @src.Seq::new("AGVL"), - id="A", - name="A", - ) + let record = @src.SeqRecord::new(@src.Seq::new("AGVL"), id="A", name="A") let output = @src.write_pdb_seqrecords([record]) assert_true(output.contains("SEQRES")) assert_true(output.contains("ALA")) @@ -231,16 +228,8 @@ test "pdb_write_pdb_seqrecords_basic" { ///| test "pdb_write_pdb_seqrecords_multiple_records" { - let r1 = @src.SeqRecord::new( - @src.Seq::new("AG"), - id="A", - name="A", - ) - let r2 = @src.SeqRecord::new( - @src.Seq::new("VL"), - id="B", - name="B", - ) + let r1 = @src.SeqRecord::new(@src.Seq::new("AG"), id="A", name="A") + let r2 = @src.SeqRecord::new(@src.Seq::new("VL"), id="B", name="B") let output = @src.write_pdb_seqrecords([r1, r2]) assert_true(output.contains("SEQRES")) assert_true(output.contains("ALA")) @@ -257,11 +246,7 @@ test "pdb_write_pdb_seqrecords_empty" { ///| test "pdb_write_pdb_seqrecords_unknown_residue" { - let record = @src.SeqRecord::new( - @src.Seq::new("AXG"), - id="A", - name="A", - ) + let record = @src.SeqRecord::new(@src.Seq::new("AXG"), id="A", name="A") let output = @src.write_pdb_seqrecords([record]) assert_true(output.contains("SEQRES")) assert_true(output.contains("ALA")) diff --git a/test/moonbit/pdb_vectors_test.mbt b/test/moonbit/pdb_vectors_test.mbt index 1098bd27..4a5e894b 100644 --- a/test/moonbit/pdb_vectors_test.mbt +++ b/test/moonbit/pdb_vectors_test.mbt @@ -112,7 +112,7 @@ test "Vector3::dot" { test "Vector3::dot orthogonal" { let v1 = @src.Vector3::new(1.0, 0.0, 0.0) let v2 = @src.Vector3::new(0.0, 1.0, 0.0) - assert_true((v1.dot(v2)).abs() < 1.0e-10) + assert_true(v1.dot(v2).abs() < 1.0e-10) } ///| @@ -236,8 +236,8 @@ test "Vector3::project_to_vector" { let target = @src.Vector3::new(1.0, 0.0, 0.0) let proj = v.project_to_vector(target) assert_true((proj.get_x() - 1.0).abs() < 1.0e-10) - assert_true((proj.get_y()).abs() < 1.0e-10) - assert_true((proj.get_z()).abs() < 1.0e-10) + assert_true(proj.get_y().abs() < 1.0e-10) + assert_true(proj.get_z().abs() < 1.0e-10) } ///| @@ -247,7 +247,7 @@ test "Vector3::project_to_plane" { let proj = v.project_to_plane(normal) assert_true((proj.get_x() - 1.0).abs() < 1.0e-10) assert_true((proj.get_y() - 2.0).abs() < 1.0e-10) - assert_true((proj.get_z()).abs() < 1.0e-10) + assert_true(proj.get_z().abs() < 1.0e-10) } ///| @@ -263,9 +263,15 @@ test "RotationMatrix3::identity" { ///| test "RotationMatrix3::transpose" { let m = @src.RotationMatrix3::new( - m00=1.0, m01=2.0, m02=3.0, - m10=4.0, m11=5.0, m12=6.0, - m20=7.0, m21=8.0, m22=9.0, + m00=1.0, + m01=2.0, + m02=3.0, + m10=4.0, + m11=5.0, + m12=6.0, + m20=7.0, + m21=8.0, + m22=9.0, ) let t = m.transpose() assert_eq(t.get(0, 1), 4.0) @@ -277,9 +283,15 @@ test "RotationMatrix3::transpose" { ///| test "RotationMatrix3::multiply identity" { let m = @src.RotationMatrix3::new( - m00=1.0, m01=2.0, m02=3.0, - m10=4.0, m11=5.0, m12=6.0, - m20=7.0, m21=8.0, m22=9.0, + m00=1.0, + m01=2.0, + m02=3.0, + m10=4.0, + m11=5.0, + m12=6.0, + m20=7.0, + m21=8.0, + m22=9.0, ) let identity = @src.RotationMatrix3::identity() let result = m.multiply(identity) @@ -293,9 +305,9 @@ test "RotationMatrix3::transform" { let rot = @src.RotationMatrix3::rotation_z(@math.PI / 2.0) let v = @src.Vector3::new(1.0, 0.0, 0.0) let result = rot.transform(v) - assert_true((result.get_x()).abs() < 1.0e-10) + assert_true(result.get_x().abs() < 1.0e-10) assert_true((result.get_y() - 1.0).abs() < 1.0e-10) - assert_true((result.get_z()).abs() < 1.0e-10) + assert_true(result.get_z().abs() < 1.0e-10) } ///| @@ -303,8 +315,8 @@ test "RotationMatrix3::rotation_x" { let rot = @src.RotationMatrix3::rotation_x(@math.PI / 2.0) let v = @src.Vector3::new(0.0, 1.0, 0.0) let result = rot.transform(v) - assert_true((result.get_x()).abs() < 1.0e-10) - assert_true((result.get_y()).abs() < 1.0e-10) + assert_true(result.get_x().abs() < 1.0e-10) + assert_true(result.get_y().abs() < 1.0e-10) assert_true((result.get_z() - 1.0).abs() < 1.0e-10) } @@ -314,8 +326,8 @@ test "RotationMatrix3::rotation_y" { let v = @src.Vector3::new(0.0, 0.0, 1.0) let result = rot.transform(v) assert_true((result.get_x() - 1.0).abs() < 1.0e-10) - assert_true((result.get_y()).abs() < 1.0e-10) - assert_true((result.get_z()).abs() < 1.0e-10) + assert_true(result.get_y().abs() < 1.0e-10) + assert_true(result.get_z().abs() < 1.0e-10) } ///| @@ -323,9 +335,9 @@ test "RotationMatrix3::rotation_z" { let rot = @src.RotationMatrix3::rotation_z(@math.PI / 2.0) let v = @src.Vector3::new(1.0, 0.0, 0.0) let result = rot.transform(v) - assert_true((result.get_x()).abs() < 1.0e-10) + assert_true(result.get_x().abs() < 1.0e-10) assert_true((result.get_y() - 1.0).abs() < 1.0e-10) - assert_true((result.get_z()).abs() < 1.0e-10) + assert_true(result.get_z().abs() < 1.0e-10) } ///| @@ -334,9 +346,9 @@ test "RotationMatrix3::rotation_axis_angle" { let rot = @src.RotationMatrix3::rotation_axis_angle(axis, @math.PI / 2.0) let v = @src.Vector3::new(1.0, 0.0, 0.0) let result = rot.transform(v) - assert_true((result.get_x()).abs() < 1.0e-10) + assert_true(result.get_x().abs() < 1.0e-10) assert_true((result.get_y() - 1.0).abs() < 1.0e-10) - assert_true((result.get_z()).abs() < 1.0e-10) + assert_true(result.get_z().abs() < 1.0e-10) } ///| @@ -371,8 +383,8 @@ test "RotationMatrix3::inverse" { assert_true((product.get(0, 0) - 1.0).abs() < 1.0e-10) assert_true((product.get(1, 1) - 1.0).abs() < 1.0e-10) assert_true((product.get(2, 2) - 1.0).abs() < 1.0e-10) - assert_true((product.get(0, 1)).abs() < 1.0e-10) - assert_true((product.get(1, 0)).abs() < 1.0e-10) + assert_true(product.get(0, 1).abs() < 1.0e-10) + assert_true(product.get(1, 0).abs() < 1.0e-10) } ///| @@ -437,7 +449,7 @@ test "vector_rmsd identical" { @src.Vector3::new(4.0, 5.0, 6.0), ] let rmsd = @src.vector_rmsd(set_a, set_b) - assert_true((rmsd).abs() < 1.0e-10) + assert_true(rmsd.abs() < 1.0e-10) } ///| @@ -445,7 +457,7 @@ test "vector_rmsd empty" { let set_a : Array[@src.Vector3] = [] let set_b : Array[@src.Vector3] = [] let rmsd = @src.vector_rmsd(set_a, set_b) - assert_true((rmsd).abs() < 1.0e-10) + assert_true(rmsd.abs() < 1.0e-10) } ///| @@ -496,7 +508,7 @@ test "vector_apply_transform" { let trans = @src.Vector3::new(0.0, 0.0, 0.0) let result = @src.vector_apply_transform(vectors, rot, trans) assert_eq(result.length(), 1) - assert_true((result[0].get_x()).abs() < 1.0e-10) + assert_true(result[0].get_x().abs() < 1.0e-10) assert_true((result[0].get_y() - 1.0).abs() < 1.0e-10) } diff --git a/test/moonbit/peak_calling_test.mbt b/test/moonbit/peak_calling_test.mbt index 98c8d09a..3a637ec6 100644 --- a/test/moonbit/peak_calling_test.mbt +++ b/test/moonbit/peak_calling_test.mbt @@ -113,10 +113,7 @@ test "peak_calling_params_new_defaults_when_omitted" { ///| test "peak_calling_params_partial_overrides" { // Override only some args; others should default. - let params = @src.PeakCallingParams::new( - window_size=250, - fdr_threshold=0.01, - ) + let params = @src.PeakCallingParams::new(window_size=250, fdr_threshold=0.01) assert_eq(params.window_size, 250) assert_eq(params.step_size, 100) assert_eq(params.local_lambda_size, 10000) @@ -167,7 +164,9 @@ test "peak_calling_estimate_local_lambda_basic" { // Place 20 control reads uniformly in [0, 2000]. let mut i = 0 while i < 20 { - control.push(@src.ChipSeqRead::new(chr="chr1", position=i * 100, strand="+")) + control.push( + @src.ChipSeqRead::new(chr="chr1", position=i * 100, strand="+"), + ) i = i + 1 } // Window of 2000 bp centered at 1000 = [0, 2000) should have 20 reads. @@ -185,9 +184,7 @@ test "peak_calling_estimate_local_lambda_empty_control" { ///| test "peak_calling_estimate_local_lambda_off_chromosome" { - let control = [ - @src.ChipSeqRead::new(chr="chr1", position=500, strand="+"), - ] + let control = [@src.ChipSeqRead::new(chr="chr1", position=500, strand="+")] let rate = @src.estimate_local_lambda(control, "chr2", 500, 1000) assert_eq(rate, 0.0) } @@ -273,12 +270,24 @@ test "peak_calling_fold_enrichment_zero_observed_zero_expected" { test "peak_calling_merge_peaks_overlapping" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=2000, summit=1500, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr1", + start=1000, + end=2000, + summit=1500, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), @src.CandidatePeak::new( - chr="chr1", start=1800, end=2500, summit=2000, - read_count=15, control_count=1, fold_enrichment=15.0, p_value=1.0e-7, + chr="chr1", + start=1800, + end=2500, + summit=2000, + read_count=15, + control_count=1, + fold_enrichment=15.0, + p_value=1.0e-7, ), ] let merged = @src.merge_peaks(peaks, 0) @@ -297,12 +306,24 @@ test "peak_calling_merge_peaks_overlapping" { test "peak_calling_merge_peaks_adjacent_within_gap" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=10, control_count=1, fold_enrichment=10.0, p_value=0.001, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=10, + control_count=1, + fold_enrichment=10.0, + p_value=0.001, ), @src.CandidatePeak::new( - chr="chr1", start=1700, end=2200, summit=2000, - read_count=12, control_count=2, fold_enrichment=6.0, p_value=0.0001, + chr="chr1", + start=1700, + end=2200, + summit=2000, + read_count=12, + control_count=2, + fold_enrichment=6.0, + p_value=0.0001, ), ] // Gap is 200bp; with max_gap=300, they should merge. @@ -318,12 +339,24 @@ test "peak_calling_merge_peaks_adjacent_within_gap" { test "peak_calling_merge_peaks_non_overlapping_different_chromosomes" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=2000, summit=1500, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr1", + start=1000, + end=2000, + summit=1500, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), @src.CandidatePeak::new( - chr="chr2", start=1000, end=2000, summit=1500, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr2", + start=1000, + end=2000, + summit=1500, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), ] let merged = @src.merge_peaks(peaks, 1000) @@ -334,12 +367,24 @@ test "peak_calling_merge_peaks_non_overlapping_different_chromosomes" { test "peak_calling_merge_peaks_far_apart" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=10, control_count=1, fold_enrichment=10.0, p_value=0.001, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=10, + control_count=1, + fold_enrichment=10.0, + p_value=0.001, ), @src.CandidatePeak::new( - chr="chr1", start=10000, end=10500, summit=10250, - read_count=10, control_count=1, fold_enrichment=10.0, p_value=0.001, + chr="chr1", + start=10000, + end=10500, + summit=10250, + read_count=10, + control_count=1, + fold_enrichment=10.0, + p_value=0.001, ), ] let merged = @src.merge_peaks(peaks, 100) @@ -357,8 +402,14 @@ test "peak_calling_merge_peaks_empty" { test "peak_calling_merge_peaks_single_peak" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=2000, summit=1500, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr1", + start=1000, + end=2000, + summit=1500, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), ] let merged = @src.merge_peaks(peaks, 100) @@ -371,12 +422,24 @@ test "peak_calling_merge_peaks_unsorted_input" { // Input peaks in reverse order; merge_peaks should sort internally. let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=5000, end=5500, summit=5200, - read_count=15, control_count=1, fold_enrichment=15.0, p_value=1.0e-7, + chr="chr1", + start=5000, + end=5500, + summit=5200, + read_count=15, + control_count=1, + fold_enrichment=15.0, + p_value=1.0e-7, ), @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), ] let merged = @src.merge_peaks(peaks, 100) @@ -391,16 +454,34 @@ test "peak_calling_apply_fdr_marks_significant" { // Three peaks with very small p-values; all should be marked significant. let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=50, control_count=2, fold_enrichment=25.0, p_value=1.0e-10, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=50, + control_count=2, + fold_enrichment=25.0, + p_value=1.0e-10, ), @src.CandidatePeak::new( - chr="chr1", start=2000, end=2500, summit=2250, - read_count=40, control_count=2, fold_enrichment=20.0, p_value=1.0e-8, + chr="chr1", + start=2000, + end=2500, + summit=2250, + read_count=40, + control_count=2, + fold_enrichment=20.0, + p_value=1.0e-8, ), @src.CandidatePeak::new( - chr="chr1", start=3000, end=3500, summit=3250, - read_count=30, control_count=2, fold_enrichment=15.0, p_value=1.0e-6, + chr="chr1", + start=3000, + end=3500, + summit=3250, + read_count=30, + control_count=2, + fold_enrichment=15.0, + p_value=1.0e-6, ), ] let result = @src.apply_fdr(peaks, 0.05) @@ -415,12 +496,24 @@ test "peak_calling_apply_fdr_filters_large_pvalues" { // Mix of significant and non-significant peaks. let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=50, control_count=2, fold_enrichment=25.0, p_value=1.0e-10, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=50, + control_count=2, + fold_enrichment=25.0, + p_value=1.0e-10, ), @src.CandidatePeak::new( - chr="chr1", start=2000, end=2500, summit=2250, - read_count=5, control_count=2, fold_enrichment=2.5, p_value=0.4, + chr="chr1", + start=2000, + end=2500, + summit=2250, + read_count=5, + control_count=2, + fold_enrichment=2.5, + p_value=0.4, ), ] let result = @src.apply_fdr(peaks, 0.05) @@ -468,7 +561,13 @@ test "peak_calling_apply_fdr_monotonic" { pairs.push((p.p_value, p.fdr)) } pairs.sort_by(fn(a : (Double, Double), b : (Double, Double)) -> Int { - if a.0 < b.0 { -1 } else if a.0 > b.0 { 1 } else { 0 } + if a.0 < b.0 { + -1 + } else if a.0 > b.0 { + 1 + } else { + 0 + } }) // Adjusted p-values should be non-decreasing with rank. let mut k = 0 @@ -487,16 +586,34 @@ test "peak_calling_filter_peaks_thresholds" { ) let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=1500, summit=1250, - read_count=50, control_count=2, fold_enrichment=25.0, p_value=1.0e-10, + chr="chr1", + start=1000, + end=1500, + summit=1250, + read_count=50, + control_count=2, + fold_enrichment=25.0, + p_value=1.0e-10, ), @src.CandidatePeak::new( - chr="chr1", start=2000, end=2500, summit=2250, - read_count=5, control_count=2, fold_enrichment=2.5, p_value=1.0e-3, + chr="chr1", + start=2000, + end=2500, + summit=2250, + read_count=5, + control_count=2, + fold_enrichment=2.5, + p_value=1.0e-3, ), @src.CandidatePeak::new( - chr="chr1", start=3000, end=3500, summit=3250, - read_count=3, control_count=2, fold_enrichment=1.5, p_value=1.0e-10, + chr="chr1", + start=3000, + end=3500, + summit=3250, + read_count=3, + control_count=2, + fold_enrichment=1.5, + p_value=1.0e-10, ), ] // Set FDR and significant flags first. @@ -603,9 +720,7 @@ test "peak_calling_no_control" { ///| test "peak_calling_single_read" { - let treatment = [ - @src.ChipSeqRead::new(chr="chr1", position=1000, strand="+"), - ] + let treatment = [@src.ChipSeqRead::new(chr="chr1", position=1000, strand="+")] let control : Array[@src.ChipSeqRead] = [] let params = @src.PeakCallingParams::default() let peaks = @src.call_peaks(treatment, control, params) @@ -654,12 +769,24 @@ test "peak_calling_summary_empty" { test "peak_calling_summary_with_peaks" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=2000, summit=1500, - read_count=42, control_count=5, fold_enrichment=8.4, p_value=1.0e-10, + chr="chr1", + start=1000, + end=2000, + summit=1500, + read_count=42, + control_count=5, + fold_enrichment=8.4, + p_value=1.0e-10, ), @src.CandidatePeak::new( - chr="chr1", start=3000, end=4000, summit=3500, - read_count=20, control_count=2, fold_enrichment=10.0, p_value=1.0e-8, + chr="chr1", + start=3000, + end=4000, + summit=3500, + read_count=20, + control_count=2, + fold_enrichment=10.0, + p_value=1.0e-8, ), ] let with_fdr = @src.apply_fdr(peaks, 0.05) @@ -673,12 +800,24 @@ test "peak_calling_summary_with_peaks" { test "peak_calling_summary_includes_significant_count" { let peaks = [ @src.CandidatePeak::new( - chr="chr1", start=1000, end=2000, summit=1500, - read_count=42, control_count=5, fold_enrichment=8.4, p_value=1.0e-10, + chr="chr1", + start=1000, + end=2000, + summit=1500, + read_count=42, + control_count=5, + fold_enrichment=8.4, + p_value=1.0e-10, ), @src.CandidatePeak::new( - chr="chr1", start=3000, end=4000, summit=3500, - read_count=2, control_count=2, fold_enrichment=1.0, p_value=0.5, + chr="chr1", + start=3000, + end=4000, + summit=3500, + read_count=2, + control_count=2, + fold_enrichment=1.0, + p_value=0.5, ), ] let with_fdr = @src.apply_fdr(peaks, 0.05) diff --git a/test/moonbit/phd_test.mbt b/test/moonbit/phd_test.mbt index 4104e7a5..e122a2b0 100644 --- a/test/moonbit/phd_test.mbt +++ b/test/moonbit/phd_test.mbt @@ -12,10 +12,10 @@ test "phd_base_creation" { assert_eq(b.peak_position(), 1) } +///| test "phd_comment_creation" { let c = @src.PhdComment::new( - "read.ab1", "0.020425.c", "/etc/phred.dat", - 0, 1000, 0, 10, 0.05, "term", "big", + "read.ab1", "0.020425.c", "/etc/phred.dat", 0, 1000, 0, 10, 0.05, "term", "big", ) assert_eq(c.chromat_file(), "read.ab1") assert_eq(c.phred_version(), "0.020425.c") @@ -29,6 +29,7 @@ test "phd_comment_creation" { assert_eq(c.dye(), "big") } +///| test "phd_read_creation_and_sequence" { let bases = [ @src.PhdBase::new("A", 35, 1), @@ -36,8 +37,7 @@ test "phd_read_creation_and_sequence" { @src.PhdBase::new("G", 45, 3), ] let comment = @src.PhdComment::new( - "read.ab1", "0.020425.c", "/etc/phred.dat", - 0, 1000, 0, 3, 0.05, "term", "big", + "read.ab1", "0.020425.c", "/etc/phred.dat", 0, 1000, 0, 3, 0.05, "term", "big", ) let read = @src.PhdRead::new("test_read", bases, comment) assert_eq(read.name(), "test_read") @@ -47,16 +47,16 @@ test "phd_read_creation_and_sequence" { assert_eq(read.peak_positions(), [1, 2, 3]) } +///| test "phd_read_empty_sequence" { - let comment = @src.PhdComment::new( - "", "", "", 0, 0, 0, 0, 0.0, "", "", - ) + let comment = @src.PhdComment::new("", "", "", 0, 0, 0, 0, 0.0, "", "") let read = @src.PhdRead::new("empty", [], comment) assert_eq(read.length(), 0) assert_eq(read.sequence(), "") assert_eq(read.quality().length(), 0) } +///| test "phd_file_creation" { let comment = @src.PhdComment::new( "read.ab1", "0.020425.c", "", 0, 0, 0, 0, 0.0, "", "", @@ -74,12 +74,14 @@ test "phd_file_creation" { // Sample data // --------------------------------------------------------------------------- +///| test "phd_sample_text_has_begin_sequence" { let text = @src.phd_sample_text() assert_true(text.contains("BEGIN_SEQUENCE")) assert_true(text.contains("END_SEQUENCE")) } +///| test "phd_sample_text_has_comment_block" { let text = @src.phd_sample_text() assert_true(text.contains("BEGIN_COMMENT")) @@ -88,6 +90,7 @@ test "phd_sample_text_has_comment_block" { assert_true(text.contains("PHRED_VERSION")) } +///| test "phd_sample_text_has_dna_block" { let text = @src.phd_sample_text() assert_true(text.contains("BEGIN_DNA")) @@ -98,6 +101,7 @@ test "phd_sample_text_has_dna_block" { // Parsing // --------------------------------------------------------------------------- +///| test "phd_parse_single_read" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) @@ -108,6 +112,7 @@ test "phd_parse_single_read" { assert_eq(read.sequence(), "ACGTACGTAC") } +///| test "phd_parse_comment_fields" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) @@ -123,6 +128,7 @@ test "phd_parse_comment_fields" { assert_eq(c.dye(), "big") } +///| test "phd_parse_quality_scores" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) @@ -134,6 +140,7 @@ test "phd_parse_quality_scores" { assert_eq(quals[9], 44) } +///| test "phd_parse_peak_positions" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) @@ -144,17 +151,20 @@ test "phd_parse_peak_positions" { assert_eq(peaks[9], 10) } +///| test "phd_parse_multiple_reads" { let text = @src.phd_sample_text() + "\n" + @src.phd_sample_text() let file = @src.phd_parse(text) assert_eq(file.num_reads(), 2) } +///| test "phd_parse_empty_input" { let file = @src.phd_parse("") assert_eq(file.num_reads(), 0) } +///| test "phd_parse_minimal_phd" { let text = "BEGIN_SEQUENCE r1\nBEGIN_DNA\nA 10 1\nT 20 2\nEND_DNA\nEND_SEQUENCE\n" let file = @src.phd_parse(text) @@ -166,6 +176,7 @@ test "phd_parse_minimal_phd" { assert_eq(read.quality(), [10, 20]) } +///| test "phd_parse_no_dna_block" { let text = "BEGIN_SEQUENCE r1\nBEGIN_COMMENT\nCHROMAT_FILE: r.ab1\nEND_COMMENT\nEND_SEQUENCE\n" let file = @src.phd_parse(text) @@ -176,6 +187,7 @@ test "phd_parse_no_dna_block" { assert_eq(read.comment().chromat_file(), "r.ab1") } +///| test "phd_parse_trim_values" { let text = "BEGIN_SEQUENCE r1\nBEGIN_COMMENT\nTRIM: 5 50 0.02\nEND_COMMENT\nBEGIN_DNA\nA 10 1\nEND_DNA\nEND_SEQUENCE\n" let file = @src.phd_parse(text) @@ -188,6 +200,7 @@ test "phd_parse_trim_values" { // Formatting // --------------------------------------------------------------------------- +///| test "phd_read_to_string" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) @@ -198,6 +211,7 @@ test "phd_read_to_string" { assert_true(s.contains("ACGTACGTAC")) } +///| test "phd_file_to_string" { let text = @src.phd_sample_text() let file = @src.phd_parse(text) diff --git a/test/moonbit/pheatmap_test.mbt b/test/moonbit/pheatmap_test.mbt index 62a0a8fc..a5262199 100644 --- a/test/moonbit/pheatmap_test.mbt +++ b/test/moonbit/pheatmap_test.mbt @@ -3,16 +3,10 @@ ///| test "pheatmap_create_input" { - let mat = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] - let input = @src.pheatmap_input( - mat, - ["row1", "row2", "row3"], - ["col1", "col2", "col3"], - ) + let mat = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] + let input = @src.pheatmap_input(mat, ["row1", "row2", "row3"], [ + "col1", "col2", "col3", + ]) assert_eq(input.mat.length(), 3) assert_eq(input.row_names.length(), 3) assert_eq(input.col_names.length(), 3) @@ -20,11 +14,11 @@ test "pheatmap_create_input" { ///| test "pheatmap_distance_euclidean" { - let data = [ - [1.0, 0.0], - [0.0, 0.0], - ] - let dist = @src.pheatmap_distance_matrix(data, @src.distance_method_euclidean()) + let data = [[1.0, 0.0], [0.0, 0.0]] + let dist = @src.pheatmap_distance_matrix( + data, + @src.distance_method_euclidean(), + ) assert_eq(dist.length(), 2) assert_eq(dist[0][0], 0.0) assert_eq(dist[1][1], 0.0) @@ -33,23 +27,22 @@ test "pheatmap_distance_euclidean" { ///| test "pheatmap_distance_manhattan" { - let data = [ - [1.0, 2.0], - [4.0, 6.0], - ] - let dist = @src.pheatmap_distance_matrix(data, @src.distance_method_manhattan()) + let data = [[1.0, 2.0], [4.0, 6.0]] + let dist = @src.pheatmap_distance_matrix( + data, + @src.distance_method_manhattan(), + ) assert_eq(dist.length(), 2) assert_eq(dist[0][1], 7.0) } ///| test "pheatmap_distance_symmetry" { - let data = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] - let dist = @src.pheatmap_distance_matrix(data, @src.distance_method_euclidean()) + let data = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] + let dist = @src.pheatmap_distance_matrix( + data, + @src.distance_method_euclidean(), + ) let n = dist.length() let mut i = 0 while i < n { @@ -64,26 +57,18 @@ test "pheatmap_distance_symmetry" { ///| test "pheatmap_column_distance" { - let data = [ - [1.0, 0.0, 1.0], - [0.0, 1.0, 0.0], - ] - let dist = @src.pheatmap_column_distance_matrix(data, @src.distance_method_euclidean()) + let data = [[1.0, 0.0, 1.0], [0.0, 1.0, 0.0]] + let dist = @src.pheatmap_column_distance_matrix( + data, + @src.distance_method_euclidean(), + ) assert_eq(dist.length(), 3) } ///| test "pheatmap_render_basic" { - let mat = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] - let input = @src.pheatmap_input( - mat, - ["r1", "r2", "r3"], - ["c1", "c2", "c3"], - ) + let mat = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] + let input = @src.pheatmap_input(mat, ["r1", "r2", "r3"], ["c1", "c2", "c3"]) let result = @src.pheatmap_render(input) assert_eq(result.row_order.length(), 3) assert_eq(result.col_order.length(), 3) @@ -108,11 +93,7 @@ test "pheatmap_set_fontsize" { ///| test "pheatmap_cutree_rows" { - let mat = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] + let mat = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] let input = @src.pheatmap_input(mat, ["r1", "r2", "r3"], ["c1", "c2", "c3"]) let cut = @src.pheatmap_set_cutree(input, 2, 1) assert_eq(cut.cutree_rows, 2) @@ -120,10 +101,7 @@ test "pheatmap_cutree_rows" { ///| test "pheatmap_no_clustering" { - let mat = [ - [1.0, 2.0], - [3.0, 4.0], - ] + let mat = [[1.0, 2.0], [3.0, 4.0]] let input = @src.pheatmap_input(mat, ["r1", "r2"], ["c1", "c2"]) let unclustered = @src.pheatmap_set_cluster(input, false, false) assert_eq(unclustered.cluster_rows, false) diff --git a/test/moonbit/phylo_cdao_test.mbt b/test/moonbit/phylo_cdao_test.mbt index 75a7675a..50ad3a81 100644 --- a/test/moonbit/phylo_cdao_test.mbt +++ b/test/moonbit/phylo_cdao_test.mbt @@ -30,7 +30,12 @@ test "cdao_rdfs_namespace_constant" { ///| test "cdao_tree_construction" { - let t = @src.Cdaotree::new("#tree1", rooted=true, root_node_id="#node1", name=Some("MyTree")) + let t = @src.Cdaotree::new( + "#tree1", + rooted=true, + root_node_id="#node1", + name=Some("MyTree"), + ) assert_eq(t.id, "#tree1") assert_eq(t.rooted, true) assert_eq(t.root_node_id, "#node1") @@ -48,7 +53,11 @@ test "cdao_tree_default_values" { ///| test "cdao_node_construction" { - let n = @src.CdaoNode::new("#node1", children=["#node2", "#node3"], parent_id=Some("#node0")) + let n = @src.CdaoNode::new( + "#node1", + children=["#node2", "#node3"], + parent_id=Some("#node0"), + ) assert_eq(n.id, "#node1") assert_eq(n.children.length(), 2) assert_eq(n.parent_id.unwrap(), "#node0") @@ -337,7 +346,13 @@ test "cdao_to_trees_empty_document" { ///| test "cdao_node_fields_via_constructor" { - let n = @src.CdaoNode::new("#test", parent_id=Some("#parent"), tu_id=Some("#tu1"), branch_length=Some(2.5), label=Some("TestNode")) + let n = @src.CdaoNode::new( + "#test", + parent_id=Some("#parent"), + tu_id=Some("#tu1"), + branch_length=Some(2.5), + label=Some("TestNode"), + ) assert_eq(n.parent_id.unwrap(), "#parent") assert_eq(n.tu_id.unwrap(), "#tu1") assert_eq(n.branch_length.unwrap(), 2.5) diff --git a/test/moonbit/phylo_consensus_test.mbt b/test/moonbit/phylo_consensus_test.mbt index aa43e3d0..4cde5775 100644 --- a/test/moonbit/phylo_consensus_test.mbt +++ b/test/moonbit/phylo_consensus_test.mbt @@ -1,12 +1,12 @@ ///| /// Phylo.Consensus module tests - test "ConsensusNode::new" { let node = @src.ConsensusNode::new("A", false) assert_eq(node.name, "A") assert_true(node.is_leaf()) } +///| test "ConsensusNode::add_child" { let parent = @src.ConsensusNode::new("internal", true) let child = @src.ConsensusNode::new("A", false) @@ -14,6 +14,7 @@ test "ConsensusNode::add_child" { assert_eq(parent_with_child.children.length(), 1) } +///| test "ConsensusNode::get_leaves" { let root = @src.ConsensusNode::new("root", true) let child1 = @src.ConsensusNode::new("A", false) @@ -24,6 +25,7 @@ test "ConsensusNode::get_leaves" { assert_eq(leaves.length(), 2) } +///| test "parse_newick" { let newick = "((A,B),(C,D));" match @src.parse_consensus_newick(newick) { @@ -32,6 +34,7 @@ test "parse_newick" { } } +///| test "newick_to_tree" { let tree = @src.newick_to_tree("((A,B),(C,D));") match tree.root { @@ -40,12 +43,14 @@ test "newick_to_tree" { } } +///| test "ConsensusTree::to_newick" { let tree = @src.newick_to_tree("((A,B),(C,D));") let newick = tree.to_newick() assert_true(newick.length() > 0) } +///| test "build_majority_consensus" { let trees = @src.create_example_trees() let consensus = @src.build_majority_consensus(trees) @@ -55,6 +60,7 @@ test "build_majority_consensus" { } } +///| test "build_strict_consensus" { let trees = @src.create_example_trees() let consensus = @src.build_strict_consensus(trees) @@ -64,6 +70,7 @@ test "build_strict_consensus" { } } +///| test "get_all_splits" { let tree = @src.newick_to_tree("((A,B),(C,D));") match tree.root { @@ -76,6 +83,7 @@ test "get_all_splits" { } } +///| test "Split::normalize" { let split = @src.Split::new(["B", "A"]) let normalized = split.normalize() @@ -83,12 +91,14 @@ test "Split::normalize" { assert_eq(normalized.taxa[1], "B") } +///| test "Split::hash" { let split = @src.Split::new(["A", "B"]) let h = split.hash() assert_eq(h, "A,B") } +///| test "calculate_consensus_support" { let trees = @src.create_example_trees() let split = @src.Split::new(["A", "B"]) @@ -96,17 +106,20 @@ test "calculate_consensus_support" { assert_true(support >= 0.0 && support <= 1.0) } +///| test "get_all_splits_from_trees" { let trees = @src.create_example_trees() let splits = @src.get_all_splits_from_trees(trees) assert_true(splits.length() > 0) } +///| test "create_example_trees" { let trees = @src.create_example_trees() assert_eq(trees.length(), 5) } +///| test "create_simple_tree" { let tree = @src.create_simple_tree() match tree.root { @@ -115,6 +128,7 @@ test "create_simple_tree" { } } +///| test "build_consensus_empty" { let trees : Array[@src.ConsensusTree] = Array::new() let consensus = @src.build_consensus(trees, 0.5) @@ -122,4 +136,4 @@ test "build_consensus_empty" { Some(_) => assert_true(false) None => () } -} \ No newline at end of file +} diff --git a/test/moonbit/phylo_nexml_test.mbt b/test/moonbit/phylo_nexml_test.mbt index 939fc769..f80de22c 100644 --- a/test/moonbit/phylo_nexml_test.mbt +++ b/test/moonbit/phylo_nexml_test.mbt @@ -57,12 +57,7 @@ test "nexml_tree_creation" { let nodes : Array[@src.NeXMLNode] = [] let edges : Array[@src.NeXMLEdge] = [] let tree = @src.NeXMLTree::new( - "tree1", - "Test Tree", - "FloatTree", - true, - nodes, - edges, + "tree1", "Test Tree", "FloatTree", true, nodes, edges, ) assert_eq(tree.id, "tree1") assert_eq(tree.name, "Test Tree") diff --git a/test/moonbit/phylo_xml_debug_test.mbt b/test/moonbit/phylo_xml_debug_test.mbt index e61d6834..888fa40a 100644 --- a/test/moonbit/phylo_xml_debug_test.mbt +++ b/test/moonbit/phylo_xml_debug_test.mbt @@ -55,8 +55,7 @@ test "debug_parse_tree" { ///| test "debug_multiple_trees" { - let xml = - "\n" + + let xml = "\n" + "\n" + " \n" + " Tree 1\n" + diff --git a/test/moonbit/phylo_xml_test.mbt b/test/moonbit/phylo_xml_test.mbt index f9fdd9c5..83d775d6 100644 --- a/test/moonbit/phylo_xml_test.mbt +++ b/test/moonbit/phylo_xml_test.mbt @@ -1,7 +1,6 @@ ///| test "phyloxml_parse_simple_tree" { - let xml = - "\n" + + let xml = "\n" + " \n" + " Simple Tree\n" + " \n" + @@ -21,8 +20,7 @@ test "phyloxml_parse_simple_tree" { ///| test "phyloxml_parse_branch_lengths" { - let xml = - "\n" + + let xml = "\n" + " \n" + " \n" + " 0.5\n" + @@ -43,8 +41,7 @@ test "phyloxml_parse_branch_lengths" { ///| test "phyloxml_parse_taxon_annotations" { - let xml = - "\n" + + let xml = "\n" + " \n" + " \n" + " Root\n" + @@ -65,8 +62,7 @@ test "phyloxml_parse_taxon_annotations" { ///| test "phyloxml_parse_sequence_annotations" { - let xml = - "\n" + + let xml = "\n" + " \n" + " \n" + " Root\n" + @@ -87,8 +83,7 @@ test "phyloxml_parse_sequence_annotations" { ///| test "phyloxml_xml_roundtrip" { - let xml = - "\n" + + let xml = "\n" + "\n" + " \n" + " Test Tree\n" + @@ -153,8 +148,7 @@ test "phyloxml_newick_simple" { ///| test "phyloxml_to_newick" { - let xml = - "\n" + + let xml = "\n" + "\n" + " \n" + " \n" + @@ -188,11 +182,9 @@ test "phyloxml_node_to_newick_leaf" { test "phyloxml_node_to_newick_internal" { let child1 = @src.PhyloXMLNode::new(name="A", branch_length=0.1) let child2 = @src.PhyloXMLNode::new(name="B", branch_length=0.2) - let node = @src.PhyloXMLNode::new( - name="Root", - branch_length=0.5, - children=[child1, child2], - ) + let node = @src.PhyloXMLNode::new(name="Root", branch_length=0.5, children=[ + child1, child2, + ]) let newick = @src.phyloxml_node_to_newick(node) assert_true(newick.contains("(")) assert_true(newick.contains("A")) @@ -216,10 +208,7 @@ test "phyloxml_get_tree_names" { test "phyloxml_get_all_tips" { let leaf1 = @src.PhyloXMLNode::new(name="A") let leaf2 = @src.PhyloXMLNode::new(name="B") - let internal = @src.PhyloXMLNode::new( - name="Root", - children=[leaf1, leaf2], - ) + let internal = @src.PhyloXMLNode::new(name="Root", children=[leaf1, leaf2]) let tree = @src.PhyloXMLTree::new(root=internal) let tips = @src.get_all_tips(tree) assert_eq(tips.length(), 2) @@ -234,7 +223,7 @@ test "phyloxml_get_all_tips_deep" { let leaf3 = @src.PhyloXMLNode::new(name="C") let child1 = @src.PhyloXMLNode::new(name="X", children=[leaf1, leaf2]) let root = @src.PhyloXMLNode::new(name="Root", children=[child1, leaf3]) - let tree = @src.PhyloXMLTree::new(root=root) + let tree = @src.PhyloXMLTree::new(root~) let tips = @src.get_all_tips(tree) assert_eq(tips.length(), 3) } @@ -243,19 +232,15 @@ test "phyloxml_get_all_tips_deep" { test "phyloxml_node_count" { let leaf1 = @src.PhyloXMLNode::new(name="A") let leaf2 = @src.PhyloXMLNode::new(name="B") - let root = @src.PhyloXMLNode::new( - name="Root", - children=[leaf1, leaf2], - ) - let tree = @src.PhyloXMLTree::new(root=root) + let root = @src.PhyloXMLNode::new(name="Root", children=[leaf1, leaf2]) + let tree = @src.PhyloXMLTree::new(root~) let count = @src.phyloxml_get_node_count(tree) assert_eq(count, 3) } ///| test "phyloxml_multiple_trees" { - let xml = - "\n" + + let xml = "\n" + "\n" + " \n" + " Tree 1\n" + @@ -350,10 +335,7 @@ test "phyloxml_node_creation" { test "phyloxml_node_with_children" { let child1 = @src.PhyloXMLNode::new(name="C1") let child2 = @src.PhyloXMLNode::new(name="C2") - let node = @src.PhyloXMLNode::new( - name="Parent", - children=[child1, child2], - ) + let node = @src.PhyloXMLNode::new(name="Parent", children=[child1, child2]) assert_false(node.is_leaf) assert_eq(node.children.length(), 2) assert_eq(node.children[0].name, "C1") @@ -367,7 +349,7 @@ test "phyloxml_tree_creation" { tree_id="t1", name="Test Tree", description="A test tree", - root=root, + root~, ) assert_eq(tree.tree_id, "t1") assert_eq(tree.name, "Test Tree") @@ -379,10 +361,7 @@ test "phyloxml_tree_creation" { test "phyloxml_tree_with_metadata" { let meta : Map[String, String] = Map([], capacity=2) meta["rooted"] = "true" - let tree = @src.PhyloXMLTree::new( - name="Rooted Tree", - phylogeny_metadata=meta, - ) + let tree = @src.PhyloXMLTree::new(name="Rooted Tree", phylogeny_metadata=meta) assert_eq(tree.phylogeny_metadata["rooted"], "true") } @@ -404,8 +383,7 @@ test "phyloxml_get_element_text" { ///| test "phyloxml_parse_node_element" { - let xml = - "\n" + + let xml = "\n" + " TestNode\n" + " 0.5\n" + " \n" + @@ -419,8 +397,7 @@ test "phyloxml_parse_node_element" { ///| test "phyloxml_parse_taxon_element" { - let xml = - "\n" + + let xml = "\n" + " tx1\n" + " Homo sapiens\n" + " Human\n" + @@ -432,8 +409,7 @@ test "phyloxml_parse_taxon_element" { ///| test "phyloxml_parse_sequence_element" { - let xml = - "\n" + + let xml = "\n" + " seq1\n" + " DNA\n" + " ATG\n" + @@ -448,20 +424,16 @@ test "phyloxml_complex_tree_serialization" { let leaf_a = @src.PhyloXMLNode::new(name="A", branch_length=0.1) let leaf_b = @src.PhyloXMLNode::new(name="B", branch_length=0.2) let leaf_c = @src.PhyloXMLNode::new(name="C", branch_length=0.3) - let internal = @src.PhyloXMLNode::new( - name="Int", - branch_length=0.5, - children=[leaf_a, leaf_b], - ) - let root = @src.PhyloXMLNode::new( - name="Root", - branch_length=0.0, - children=[internal, leaf_c], - ) + let internal = @src.PhyloXMLNode::new(name="Int", branch_length=0.5, children=[ + leaf_a, leaf_b, + ]) + let root = @src.PhyloXMLNode::new(name="Root", branch_length=0.0, children=[ + internal, leaf_c, + ]) let tree = @src.PhyloXMLTree::new( tree_id="complex", name="Complex Tree", - root=root, + root~, ) let result = @src.PhyloXMLResult::new(trees=[tree]) let xml = @src.phyloxml_to_xml(result) @@ -527,8 +499,7 @@ test "phyloxml_indent_xml" { ///| test "phyloxml_find_all_elements" { - let xml = - "\n" + + let xml = "\n" + " 1\n" + " 2\n" + " 3\n" + @@ -539,18 +510,14 @@ test "phyloxml_find_all_elements" { ///| test "phyloxml_taxon_namespace" { - let ns = @src.TaxonNamespace::new( - namespace_id="ns1", - name="Test NS", - ) + let ns = @src.TaxonNamespace::new(namespace_id="ns1", name="Test NS") assert_eq(ns.namespace_id, "ns1") assert_eq(ns.name, "Test NS") } ///| test "phyloxml_with_taxon_namespace" { - let xml = - "\n" + + let xml = "\n" + "\n" + " \n" + " ns1\n" + @@ -568,4 +535,4 @@ test "phyloxml_with_taxon_namespace" { assert_eq(result.taxon_namespaces.length(), 1) assert_true(result.taxon_namespaces[0].namespace_id.length() >= 0) assert_true(result.taxon_namespaces[0].name.length() >= 0) -} \ No newline at end of file +} diff --git a/test/moonbit/phyloseq_test.mbt b/test/moonbit/phyloseq_test.mbt index 36268d91..085e8ed1 100644 --- a/test/moonbit/phyloseq_test.mbt +++ b/test/moonbit/phyloseq_test.mbt @@ -1,32 +1,35 @@ ///| /// Tests for phyloseq module. - test "phyloseq creation" { let ps = @src.create_example_phyloseq() assert_eq(@src.ps_num_otus(ps), 3) assert_eq(@src.ps_num_samples(ps), 3) } +///| test "phyloseq total abundance" { let ps = @src.create_example_phyloseq() let total = @src.ps_total_abundance(ps) assert_eq(total, 975.0) } +///| test "phyloseq filter by abundance" { let ps = @src.create_example_phyloseq() let filtered = @src.ps_filter_by_abundance(ps, 300.0) assert_eq(@src.ps_num_otus(filtered), 2) } +///| test "phyloseq filter by taxonomy" { let ps = @src.create_example_phyloseq() let filtered = @src.ps_filter_by_taxonomy(ps, "phylum", "Proteobacteria") assert_eq(@src.ps_num_otus(filtered), 1) } +///| test "phyloseq taxa summary" { let ps = @src.create_example_phyloseq() let summary = @src.ps_get_taxa_summary(ps, "phylum") assert_true(summary.contains("Proteobacteria")) -} \ No newline at end of file +} diff --git a/test/moonbit/plyranges_test.mbt b/test/moonbit/plyranges_test.mbt index 55efa9d7..d9f6b448 100644 --- a/test/moonbit/plyranges_test.mbt +++ b/test/moonbit/plyranges_test.mbt @@ -52,12 +52,9 @@ test "plyranges_filter_strand" { ///| test "plyranges_mutate" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 300], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 300], [ + "+", "-", + ]) let mutated = @src.tgr_mutate(ranges, "gc_content", [0.45, 0.62]) assert_eq(mutated.metadata["gc_content"].length(), 2) assert_eq(mutated.metadata["gc_content"][0], 0.45) @@ -65,12 +62,9 @@ test "plyranges_mutate" { ///| test "plyranges_mutate_str" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 300], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 300], [ + "+", "-", + ]) let mutated = @src.tgr_mutate_str(ranges, "gene_name", ["BRCA1", "TP53"]) assert_eq(mutated.metadata_str["gene_name"].length(), 2) assert_eq(mutated.metadata_str["gene_name"][0], "BRCA1") @@ -78,12 +72,9 @@ test "plyranges_mutate_str" { ///| test "plyranges_select" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 300], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 300], [ + "+", "-", + ]) let mutated = @src.tgr_mutate(ranges, "score", [1.0, 2.0]) let selected = @src.tgr_select(mutated, ["score"]) assert_eq(selected.metadata["score"].length(), 2) @@ -106,12 +97,9 @@ test "plyranges_arrange" { ///| test "plyranges_rename" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 300], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 300], [ + "+", "-", + ]) let mutated = @src.tgr_mutate(ranges, "score", [1.0, 2.0]) let renamed = @src.tgr_rename(mutated, "score", "alignment_score") assert_eq(renamed.metadata["alignment_score"].length(), 2) @@ -119,12 +107,9 @@ test "plyranges_rename" { ///| test "plyranges_width" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 350], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 350], [ + "+", "-", + ]) assert_eq(ranges.widths[0], 101) assert_eq(ranges.widths[1], 151) } @@ -162,12 +147,9 @@ test "plyranges_filter_end" { ///| test "plyranges_multiple_metadata" { - let ranges = @src.tgr_create( - ["chr1", "chr2"], - [100, 200], - [200, 300], - ["+", "-"], - ) + let ranges = @src.tgr_create(["chr1", "chr2"], [100, 200], [200, 300], [ + "+", "-", + ]) let mutated = @src.tgr_mutate(ranges, "score", [1.5, 2.5]) let mutated2 = @src.tgr_mutate(mutated, "p_value", [0.01, 0.05]) assert_eq(mutated2.metadata.length(), 2) diff --git a/test/moonbit/polypeptide_test.mbt b/test/moonbit/polypeptide_test.mbt index c52eb22c..2bf79b7f 100644 --- a/test/moonbit/polypeptide_test.mbt +++ b/test/moonbit/polypeptide_test.mbt @@ -23,7 +23,9 @@ test "polypeptide_calculate_hydrophobicity" { ///| test "polypeptide_calculate_hydrophobicity_window" { let sequence = "MKLILVLLVSLSL" - let profile = @src.calculate_hydrophobicity_window(sequence, 5, "kyte-doolittle") + let profile = @src.calculate_hydrophobicity_window( + sequence, 5, "kyte-doolittle", + ) assert_eq(profile.values.length(), sequence.length()) } @@ -55,4 +57,4 @@ test "polypeptide_create_example_data" { assert_eq(seq.length() > 0, true) assert_eq(comp.total_residues > 0, true) assert_eq(prof.values.length() > 0, true) -} \ No newline at end of file +} diff --git a/test/moonbit/popgen_advanced_test.mbt b/test/moonbit/popgen_advanced_test.mbt index e0a6ec2a..aaac1792 100644 --- a/test/moonbit/popgen_advanced_test.mbt +++ b/test/moonbit/popgen_advanced_test.mbt @@ -58,7 +58,15 @@ test "watterson_theta_small_sample" { test "watterson_a1" { let a1 = @src.watterson_a1(10) // a1 = sum(1/i) for i=1..9 = 1 + 1/2 + 1/3 + ... + 1/9 - let expected = 1.0 + 0.5 + 0.3333333 + 0.25 + 0.2 + 0.1666667 + 0.1428571 + 0.125 + 0.1111111 + let expected = 1.0 + + 0.5 + + 0.3333333 + + 0.25 + + 0.2 + + 0.1666667 + + 0.1428571 + + 0.125 + + 0.1111111 assert_true((a1 - expected).abs() < 0.01) } @@ -176,16 +184,16 @@ test "mktest_zero_values" { test "mktest_from_sites" { let poly_sites : Array[@src.PolymorphicSite] = Array::new() let fixed_sites : Array[@src.PolymorphicSite] = Array::new() - + // Add polymorphic sites - poly_sites.push(@src.new_polymorphic_site(100, 2, 20, true)) // nonsynonymous + poly_sites.push(@src.new_polymorphic_site(100, 2, 20, true)) // nonsynonymous poly_sites.push(@src.new_polymorphic_site(200, 3, 20, false)) // synonymous - poly_sites.push(@src.new_polymorphic_site(300, 1, 20, true)) // nonsynonymous - + poly_sites.push(@src.new_polymorphic_site(300, 1, 20, true)) // nonsynonymous + // Add fixed sites - fixed_sites.push(@src.new_polymorphic_site(400, 20, 20, true)) // nonsynonymous + fixed_sites.push(@src.new_polymorphic_site(400, 20, 20, true)) // nonsynonymous fixed_sites.push(@src.new_polymorphic_site(500, 20, 20, false)) // synonymous - + let result = @src.mktest_from_sites(poly_sites, fixed_sites) assert_eq(result.p_nonsyn, 2) assert_eq(result.p_syn, 1) @@ -201,7 +209,7 @@ test "afs_calculate" { sites.push(@src.new_singleton(200, 20, true)) sites.push(@src.new_doubleton(300, 20, false)) sites.push(@src.new_polymorphic_site(400, 5, 20, false)) - + let afs = @src.calculate_afs(sites, 20) assert_eq(afs.singletons, 2) assert_eq(afs.doubletons, 1) @@ -224,7 +232,7 @@ test "tajima_d_from_afs" { sites.push(@src.new_singleton(100, 20, false)) sites.push(@src.new_singleton(200, 20, false)) sites.push(@src.new_doubleton(300, 20, false)) - + let afs = @src.calculate_afs(sites, 20) let result = @src.tajima_d_from_afs(afs, 1000) assert_eq(result.test_name, "Tajima's D") @@ -237,7 +245,7 @@ test "fu_li_d_from_afs" { let sites : Array[@src.PolymorphicSite] = Array::new() sites.push(@src.new_singleton(100, 20, false)) sites.push(@src.new_doubleton(200, 20, false)) - + let afs = @src.calculate_afs(sites, 20) let result = @src.fu_li_d_from_afs(afs) assert_eq(result.test_name, "Fu & Li's D") @@ -248,10 +256,10 @@ test "fu_li_d_from_afs" { test "normal_cdf" { let p0 = @src.popgen_normal_cdf(0.0) assert_true((p0 - 0.5).abs() < 0.01) - + let p1 = @src.popgen_normal_cdf(1.96) assert_true(p1 > 0.96 && p1 < 0.98) - + let p_neg1 = @src.popgen_normal_cdf(-1.96) assert_true(p_neg1 > 0.02 && p_neg1 < 0.04) } @@ -270,7 +278,7 @@ test "neutrality_analysis" { for i = 1; i < 3; i = i + 1 { sites.push(@src.new_polymorphic_site(1000 + i * 100, 5, 30, i == 1)) } - + let result = @src.run_neutrality_analysis(sites, 30, 10000) assert_true(!result.tajima_d.statistic.is_nan()) assert_true(!result.fu_li_d.statistic.is_nan()) @@ -296,7 +304,7 @@ test "neutrality_analysis_to_string" { sites.push(@src.new_singleton(100, 20, false)) sites.push(@src.new_singleton(200, 20, true)) sites.push(@src.new_doubleton(300, 20, false)) - + let result = @src.run_neutrality_analysis(sites, 20, 5000) let output = @src.neutrality_analysis_to_string(result) assert_true(output.contains("Neutrality Analysis Results")) @@ -353,7 +361,7 @@ test "multiple_singletons" { for i = 1; i < 15; i = i + 1 { sites.push(@src.new_singleton(i * 50, 25, i % 3 == 0)) } - + let afs = @src.calculate_afs(sites, 25) assert_eq(afs.singletons, 14) assert_eq(afs.segregating_sites, 14) @@ -367,7 +375,7 @@ test "high_freq_variants" { sites.push(@src.new_polymorphic_site(100, 24, 25, false)) sites.push(@src.new_polymorphic_site(200, 23, 25, true)) sites.push(@src.new_polymorphic_site(300, 22, 25, false)) - + let afs = @src.calculate_afs(sites, 25) assert_eq(afs.singletons, 0) // No singletons -} \ No newline at end of file +} diff --git a/test/moonbit/preprocess_core_test.mbt b/test/moonbit/preprocess_core_test.mbt index 6b852a00..f836dcef 100644 --- a/test/moonbit/preprocess_core_test.mbt +++ b/test/moonbit/preprocess_core_test.mbt @@ -14,6 +14,7 @@ test "pc_quantile_config_new" { assert_eq(config.method, "quantile") } +///| test "pc_invariant_set_result_creation" { let dummy : Array[Array[Double]] = [[1.0, 2.0], [3.0, 4.0]] let res = @src.InvariantSetResult::new(dummy, [0, 1], 0, [0.0, 1.0]) @@ -27,6 +28,7 @@ test "pc_invariant_set_result_creation" { // Helper utilities // ============================================================================ +///| test "pc_interp_linear_basic" { let xs = [0.0, 1.0, 2.0, 3.0, 4.0] let ys = [0.0, 10.0, 20.0, 30.0, 40.0] @@ -44,11 +46,13 @@ test "pc_interp_linear_basic" { assert_true((v3 - 40.0).abs() < 0.0001) } +///| test "pc_interp_linear_empty" { let v = @src.pc_interp_linear([], [], 1.0) assert_eq(v, 0.0) } +///| test "pc_interp_linear_single" { let v = @src.pc_interp_linear([5.0], [100.0], 0.0) assert_eq(v, 100.0) @@ -58,15 +62,10 @@ test "pc_interp_linear_single" { // Quantile normalization core tests // ============================================================================ +///| test "pc_normalize_quantiles_identical_distributions" { // Two columns with same values -> should remain the same - let matrix = [ - [1.0, 1.0], - [2.0, 2.0], - [3.0, 3.0], - [4.0, 4.0], - [5.0, 5.0], - ] + let matrix = [[1.0, 1.0], [2.0, 2.0], [3.0, 3.0], [4.0, 4.0], [5.0, 5.0]] let res = @src.normalize_quantiles(matrix) assert_eq(res.length(), 5) assert_eq(res[0].length(), 2) @@ -76,16 +75,12 @@ test "pc_normalize_quantiles_identical_distributions" { } } +///| test "pc_normalize_quantiles_swapped_values" { // Two columns with reversed order // Col1: [1, 2, 3, 4], Col2: [4, 3, 2, 1] // After normalization, both should have means of sorted values - let matrix = [ - [1.0, 4.0], - [2.0, 3.0], - [3.0, 2.0], - [4.0, 1.0], - ] + let matrix = [[1.0, 4.0], [2.0, 3.0], [3.0, 2.0], [4.0, 1.0]] let res = @src.normalize_quantiles(matrix) // Check distributions: for each column, after sorting, they should match // sorted(col1_normalized) == sorted(col2_normalized) == 2.5, 2.5, 2.5, 2.5 @@ -102,12 +97,9 @@ test "pc_normalize_quantiles_swapped_values" { assert_true((res[0][1] - 4.0).abs() < 0.0001) } +///| test "pc_normalize_quantiles_three_columns" { - let matrix = [ - [1.0, 4.0, 7.0], - [2.0, 5.0, 8.0], - [3.0, 6.0, 9.0], - ] + let matrix = [[1.0, 4.0, 7.0], [2.0, 5.0, 8.0], [3.0, 6.0, 9.0]] let res = @src.normalize_quantiles(matrix) assert_eq(res.length(), 3) assert_eq(res[0].length(), 3) @@ -123,11 +115,13 @@ test "pc_normalize_quantiles_three_columns" { assert_true((orig_sum - new_sum).abs() < 0.1) } +///| test "pc_normalize_quantiles_empty" { let res = @src.normalize_quantiles([]) assert_eq(res.length(), 0) } +///| test "pc_normalize_quantiles_one_row" { let matrix = [[1.0, 10.0, 100.0]] let res = @src.normalize_quantiles(matrix) @@ -144,12 +138,9 @@ test "pc_normalize_quantiles_one_row" { // Target-based quantile normalization // ============================================================================ +///| test "pc_normalize_quantiles_target" { - let matrix = [ - [1.0, 2.0], - [3.0, 4.0], - [5.0, 6.0], - ] + let matrix = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]] let target = [10.0, 20.0, 30.0] let res = @src.normalize_quantiles_use_target(matrix, target) assert_eq(res.length(), 3) @@ -161,12 +152,9 @@ test "pc_normalize_quantiles_target" { assert_true((res[2][0] - 30.0).abs() < 0.0001) } +///| test "pc_normalize_quantiles_determine_target" { - let matrix = [ - [1.0, 3.0], - [2.0, 2.0], - [3.0, 1.0], - ] + let matrix = [[1.0, 3.0], [2.0, 2.0], [3.0, 1.0]] let target = @src.normalize_quantiles_determine_target(matrix) assert_eq(target.length(), 3) // Each row of sorted matrix: sorted col1 = [1, 2, 3], col2 = [1, 2, 3] @@ -180,6 +168,7 @@ test "pc_normalize_quantiles_determine_target" { // Subset quantile normalization // ============================================================================ +///| test "pc_normalize_quantiles_subset" { let matrix = [ [1.0, 8.0], @@ -200,6 +189,7 @@ test "pc_normalize_quantiles_subset" { // Invariant set normalization // ============================================================================ +///| test "pc_find_invariant_set_perfect_correlation" { // Two arrays with perfect linear relationship let reference = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] @@ -209,11 +199,13 @@ test "pc_find_invariant_set_perfect_correlation" { assert_true(inv.length() >= 5) } +///| test "pc_find_invariant_set_length_mismatch" { let inv = @src.find_invariant_set([1.0, 2.0], [1.0]) assert_eq(inv.length(), 0) } +///| test "pc_normalize_invariantset_basic" { // 3 samples x 5 rows let matrix = [ @@ -223,7 +215,11 @@ test "pc_normalize_invariantset_basic" { [4.0, 8.0, 12.0], [5.0, 10.0, 15.0], ] - let res = @src.normalize_invariantset(matrix, reference_index=0, threshold=0.01) + let res = @src.normalize_invariantset( + matrix, + reference_index=0, + threshold=0.01, + ) assert_eq(res.reference_index, 0) assert_eq(res.normalized_data.length(), 5) // Reference column should be unchanged @@ -236,6 +232,7 @@ test "pc_normalize_invariantset_basic" { // Log transform // ============================================================================ +///| test "pc_log2_transform_basic" { let matrix = [[0.0, 1.0], [3.0, 15.0]] let res = @src.log2_transform(matrix, offset=1.0) @@ -249,6 +246,7 @@ test "pc_log2_transform_basic" { assert_true((res[1][1] - 4.0).abs() < 0.0001) } +///| test "pc_log2_transform_empty" { let res = @src.log2_transform([]) assert_eq(res.length(), 0) @@ -258,6 +256,7 @@ test "pc_log2_transform_empty" { // Background correction // ============================================================================ +///| test "pc_background_correct_percentile_basic" { let matrix = [ [1.0, 100.0], @@ -290,21 +289,16 @@ test "pc_background_correct_percentile_basic" { // Median center // ============================================================================ +///| test "pc_median_center_columns_basic" { - let matrix = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - [4.0, 40.0], - [5.0, 50.0], - ] + let matrix = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0], [4.0, 40.0], [5.0, 50.0]] let res = @src.median_center_columns(matrix) // Column 1 median = 3, centered = [-2, -1, 0, 1, 2] - assert_true((res[0][0] - (-2.0)).abs() < 0.0001) + assert_true((res[0][0] - -2.0).abs() < 0.0001) assert_true((res[2][0] - 0.0).abs() < 0.0001) assert_true((res[4][0] - 2.0).abs() < 0.0001) // Column 2 median = 30, centered = [-20, -10, 0, 10, 20] - assert_true((res[0][1] - (-20.0)).abs() < 0.0001) + assert_true((res[0][1] - -20.0).abs() < 0.0001) assert_true((res[2][1] - 0.0).abs() < 0.0001) assert_true((res[4][1] - 20.0).abs() < 0.0001) } @@ -313,14 +307,9 @@ test "pc_median_center_columns_basic" { // Column summary // ============================================================================ +///| test "pc_column_summary_basic" { - let matrix = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - [4.0, 40.0], - [5.0, 50.0], - ] + let matrix = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0], [4.0, 40.0], [5.0, 50.0]] let res = @src.column_summary(matrix) assert_eq(res.length(), 2) // Each column result is [mean, median, sd, min, max] @@ -336,6 +325,7 @@ test "pc_column_summary_basic" { assert_true((res[1][4] - 50.0).abs() < 0.0001) } +///| test "pc_column_summary_empty" { let res = @src.column_summary([]) assert_eq(res.length(), 0) diff --git a/test/moonbit/progeny_test.mbt b/test/moonbit/progeny_test.mbt index 66a420dd..0b3af720 100644 --- a/test/moonbit/progeny_test.mbt +++ b/test/moonbit/progeny_test.mbt @@ -23,8 +23,8 @@ test "progeny_mat_vec_multiply" { let vector = [1.0, 2.0] let result = @src.progeny_mat_vec(matrix, vector) assert_eq(result.length(), 2) - assert_eq(result[0], 5.0) // 1*1 + 2*2 - assert_eq(result[1], 11.0) // 3*1 + 4*2 + assert_eq(result[0], 5.0) // 1*1 + 2*2 + assert_eq(result[1], 11.0) // 3*1 + 4*2 } ///| @@ -33,10 +33,10 @@ test "progeny_mat_mul" { let b = [[5.0, 6.0], [7.0, 8.0]] let result = @src.progeny_mat_mul(a, b) assert_eq(result.length(), 2) - assert_eq(result[0][0], 19.0) // 1*5 + 2*7 - assert_eq(result[0][1], 22.0) // 1*6 + 2*8 - assert_eq(result[1][0], 43.0) // 3*5 + 4*7 - assert_eq(result[1][1], 50.0) // 3*6 + 4*8 + assert_eq(result[0][0], 19.0) // 1*5 + 2*7 + assert_eq(result[0][1], 22.0) // 1*6 + 2*8 + assert_eq(result[1][0], 43.0) // 3*5 + 4*7 + assert_eq(result[1][1], 50.0) // 3*6 + 4*8 } ///| @@ -114,11 +114,15 @@ test "progeny_build_pathway_matrix" { let gene_names = ["GeneA", "GeneB", "GeneC", "GeneD"] let pathway1_genes = ["GeneA", "GeneB"] let pathway1_weights = [1.0, -1.0] - let pgs1 = @src.PathwayGeneSet::new("Pathway1", pathway1_genes, pathway1_weights) + let pgs1 = @src.PathwayGeneSet::new( + "Pathway1", pathway1_genes, pathway1_weights, + ) let pathway2_genes = ["GeneB", "GeneC", "GeneD"] let pathway2_weights = [0.5, 1.0, -0.5] - let pgs2 = @src.PathwayGeneSet::new("Pathway2", pathway2_genes, pathway2_weights) + let pgs2 = @src.PathwayGeneSet::new( + "Pathway2", pathway2_genes, pathway2_weights, + ) let pathways = [pgs1, pgs2] let matrix = @src.build_pathway_matrix(gene_names, pathways) @@ -146,25 +150,18 @@ test "progeny_build_pathway_matrix" { ///| test "progeny_run_basic" { // Expression: 2 samples × 4 genes - let expression = [ - [5.0, 10.0, 3.0, 8.0], - [2.0, 6.0, 7.0, 4.0], - ] + let expression = [[5.0, 10.0, 3.0, 8.0], [2.0, 6.0, 7.0, 4.0]] let sample_names = ["Sample1", "Sample2"] let gene_names = ["GeneA", "GeneB", "GeneC", "GeneD"] // Pathway1: GeneA + GeneB (activation) - let pathway1 = @src.PathwayGeneSet::new( - "Pathway1", - ["GeneA", "GeneB"], - [1.0, 1.0], - ) + let pathway1 = @src.PathwayGeneSet::new("Pathway1", ["GeneA", "GeneB"], [ + 1.0, 1.0, + ]) // Pathway2: GeneC + GeneD (activation) - let pathway2 = @src.PathwayGeneSet::new( - "Pathway2", - ["GeneC", "GeneD"], - [1.0, 1.0], - ) + let pathway2 = @src.PathwayGeneSet::new("Pathway2", ["GeneC", "GeneD"], [ + 1.0, 1.0, + ]) let data = @src.ProgenyData::new( expression, @@ -192,7 +189,13 @@ test "progeny_get_pathway_activity" { let sample_names = ["S1", "S2"] let gene_names = ["G1", "G2"] let pathway1 = @src.PathwayGeneSet::new("P1", ["G1"], [1.0]) - let data = @src.ProgenyData::new(expression, sample_names, gene_names, [pathway1], 0.01) + let data = @src.ProgenyData::new( + expression, + sample_names, + gene_names, + [pathway1], + 0.01, + ) let result = @src.run_progeny(data) let activity = @src.get_pathway_activity(result, "P1") @@ -207,7 +210,13 @@ test "progeny_get_sample_profile" { let gene_names = ["G1", "G2"] let p1 = @src.PathwayGeneSet::new("P1", ["G1"], [1.0]) let p2 = @src.PathwayGeneSet::new("P2", ["G2"], [1.0]) - let data = @src.ProgenyData::new(expression, sample_names, gene_names, [p1, p2], 0.01) + let data = @src.ProgenyData::new( + expression, + sample_names, + gene_names, + [p1, p2], + 0.01, + ) let result = @src.run_progeny(data) let profile = @src.get_sample_profile(result, "S1") diff --git a/test/moonbit/prosite_test.mbt b/test/moonbit/prosite_test.mbt index 57113815..f4a53178 100644 --- a/test/moonbit/prosite_test.mbt +++ b/test/moonbit/prosite_test.mbt @@ -1,57 +1,69 @@ ///| /// Tests for Prosite module. - test "PrositePattern creation" { - let pattern = @src.PrositePattern::new("PS00001", "ATP-binding", "[AG]-X(4)-G-K-[ST]") + let pattern = @src.PrositePattern::new( + "PS00001", "ATP-binding", "[AG]-X(4)-G-K-[ST]", + ) assert_eq(pattern.accession, "PS00001") assert_eq(pattern.name, "ATP-binding") assert_eq(pattern.pattern, "[AG]-X(4)-G-K-[ST]") } +///| test "PrositeMatch creation" { - let match_ = @src.PrositeMatch::new("PS00001", "ATP-binding", 1, 10, "AGXXXXGKST") + let match_ = @src.PrositeMatch::new( + "PS00001", "ATP-binding", 1, 10, "AGXXXXGKST", + ) assert_eq(match_.pattern_accession, "PS00001") assert_eq(match_.start, 1) assert_eq(match_.end, 10) } +///| test "PrositeEntry creation" { let entry = @src.PrositeEntry::new("PS00001", "ATP/GTP-binding") assert_eq(entry.accession, "PS00001") assert_eq(entry.name, "ATP/GTP-binding") } +///| test "prosite_pattern_to_regex" { let regex = @src.prosite_pattern_to_regex("[AG]-X(4)-G-K-[ST]") assert_true(regex.length() > 0) } +///| test "prosite_search simple pattern" { let matches = @src.prosite_search("ST", "AASTKKST") assert_true(matches.length() >= 2) } +///| test "prosite_search bracket pattern" { let matches = @src.prosite_search("[ST]", "ASTCG") assert_true(matches.length() >= 2) } +///| test "prosite_search with wildcards" { let matches = @src.prosite_search("X(2)", "AAAA") assert_true(matches.length() >= 3) } +///| test "prosite_scan" { let patterns = @src.prosite_create_example_patterns() let matches = @src.prosite_scan("AASTKKST", patterns) assert_true(matches.length() >= 1) } +///| test "prosite_get_pattern" { let pattern = @src.prosite_get_pattern("PS00001") assert_true(pattern is Some(_)) } +///| test "prosite_calculate_score" { let match_ = @src.PrositeMatch::new("PS00001", "Test", 1, 5, "AAAAA") let pattern = @src.PrositePattern::new("PS00001", "Test", "AAAAA") @@ -59,12 +71,14 @@ test "prosite_calculate_score" { assert_true(score >= 90.0) } +///| test "prosite_create_example_patterns" { let patterns = @src.prosite_create_example_patterns() assert_eq(patterns.length(), 3) } +///| test "prosite_create_example_entry" { let entry = @src.prosite_create_example_entry() assert_eq(entry.accession, "PS00001") -} \ No newline at end of file +} diff --git a/test/moonbit/prot_dao_test.mbt b/test/moonbit/prot_dao_test.mbt index e36871ae..7bca0fce 100644 --- a/test/moonbit/prot_dao_test.mbt +++ b/test/moonbit/prot_dao_test.mbt @@ -1,182 +1,211 @@ ///| /// Test file for prot_dao module (IUPred disorder prediction). - test "disorder_score_positive" { // Disordered-promoting amino acids should have positive scores - assert_eq!(@src.prot_dao_disorder_score('E') > 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('P') > 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('K') > 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('S') > 0.0, true) + assert_eq(@src.prot_dao_disorder_score('E') > 0.0, true) + assert_eq(@src.prot_dao_disorder_score('P') > 0.0, true) + assert_eq(@src.prot_dao_disorder_score('K') > 0.0, true) + assert_eq(@src.prot_dao_disorder_score('S') > 0.0, true) } +///| test "disorder_score_negative" { // Order-promoting amino acids should have negative or near-zero scores - assert_eq!(@src.prot_dao_disorder_score('I') < 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('W') <= 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('Y') < 0.0, true) - assert_eq!(@src.prot_dao_disorder_score('C') <= 0.0, true) + assert_eq(@src.prot_dao_disorder_score('I') < 0.0, true) + assert_eq(@src.prot_dao_disorder_score('W') <= 0.0, true) + assert_eq(@src.prot_dao_disorder_score('Y') < 0.0, true) + assert_eq(@src.prot_dao_disorder_score('C') <= 0.0, true) } +///| test "disorder_score_unknown_aa" { // Unknown amino acids should have 0 score - assert_eq!(@src.prot_dao_disorder_score('X'), 0.0) - assert_eq!(@src.prot_dao_disorder_score('z'), 0.0) + assert_eq(@src.prot_dao_disorder_score('X'), 0.0) + assert_eq(@src.prot_dao_disorder_score('z'), 0.0) } +///| test "energy_score_positive" { - assert_eq!(@src.prot_dao_energy_score('C') > 0.0, true) - assert_eq!(@src.prot_dao_energy_score('P') > 0.0, true) - assert_eq!(@src.prot_dao_energy_score('M') > 0.0, true) + assert_eq(@src.prot_dao_energy_score('C') > 0.0, true) + assert_eq(@src.prot_dao_energy_score('P') > 0.0, true) + assert_eq(@src.prot_dao_energy_score('M') > 0.0, true) } +///| test "energy_score_negative" { - assert_eq!(@src.prot_dao_energy_score('H') < 0.0, true) - assert_eq!(@src.prot_dao_energy_score('R') < 0.0, true) + assert_eq(@src.prot_dao_energy_score('H') < 0.0, true) + assert_eq(@src.prot_dao_energy_score('R') < 0.0, true) } +///| test "disorder_residue_create" { let res = @src.DisorderResidue::new(1, 'A', 0.5, 1.0) - assert_eq!(res.position, 1) - assert_eq!(res.amino_acid, 'A') - assert_eq!(res.disorder_score, 0.5) - assert_eq!(res.energy_score, 1.0) - assert_eq!(res.is_disordered, false) - assert_eq!(res.is_disordered_long, false) + assert_eq(res.position, 1) + assert_eq(res.amino_acid, 'A') + assert_eq(res.disorder_score, 0.5) + assert_eq(res.energy_score, 1.0) + assert_eq(res.is_disordered, false) + assert_eq(res.is_disordered_long, false) } +///| test "disorder_region_create" { let reg = @src.DisorderRegion::new(10, 50, 0.6, 0.8, region_type="disordered") - assert_eq!(reg.start, 10) - assert_eq!(reg.end_, 50) - assert_eq!(reg.length, 40) - assert_eq!(reg.avg_score, 0.6) - assert_eq!(reg.max_score, 0.8) - assert_eq!(reg.region_type, "disordered") + assert_eq(reg.start, 10) + assert_eq(reg.end_, 50) + assert_eq(reg.length, 40) + assert_eq(reg.avg_score, 0.6) + assert_eq(reg.max_score, 0.8) + assert_eq(reg.region_type, "disordered") } +///| test "disorder_region_type" { - let reg = @src.DisorderRegion::new(5, 45, 0.7, 0.9, region_type="long disordered") - assert_eq!(reg.region_type, "long disordered") + let reg = @src.DisorderRegion::new( + 5, + 45, + 0.7, + 0.9, + region_type="long disordered", + ) + assert_eq(reg.region_type, "long disordered") } +///| test "prot_dao_predict_ordered_sequence" { // A sequence with mostly ordered amino acids let ordered_seq = "AILWVFMWCILVWALMVILWAFMVCLWVILMAWLVFMAW" let result = @src.prot_dao_predict(ordered_seq) - assert_eq!(result.sequence, ordered_seq) - assert_eq!(result.residues.length(), ordered_seq.length()) - assert_eq!(result.method, "IUPred-like") + assert_eq(result.sequence, ordered_seq) + assert_eq(result.residues.length(), ordered_seq.length()) + assert_eq(result.method, "IUPred-like") } +///| test "prot_dao_predict_disordered_sequence" { // A sequence with many disorder-promoting amino acids let disordered_seq = "EPKSESEKPPPPPPKKKEEEPPPKKKSEGGSSSGGGKKKPPP" let result = @src.prot_dao_predict(disordered_seq) - assert_eq!(result.residues.length(), disordered_seq.length()) + assert_eq(result.residues.length(), disordered_seq.length()) // With many disorder-promoting residues, some should be disordered - assert_eq!(result.get_n_disordered() > 0, true) + assert_eq(result.get_n_disordered() > 0, true) } +///| test "prot_dao_predict_threshold" { let seq = "EPKSESEKPPPPPPKKKEEEPPPKKKSEGGSSSGGGKKKPPP" let result_low = @src.prot_dao_predict(seq, threshold_disordered=0.3) let result_high = @src.prot_dao_predict(seq, threshold_disordered=0.7) - assert_eq!(result_low.get_n_disordered() >= result_high.get_n_disordered(), true) + assert_eq( + result_low.get_n_disordered() >= result_high.get_n_disordered(), + true, + ) } +///| test "disorder_result_n_disordered" { let result = @src.prot_dao_sample() let n = result.get_n_disordered() - assert_eq!(n >= 0, true) - assert_eq!(n <= result.sequence.length(), true) + assert_eq(n >= 0, true) + assert_eq(n <= result.sequence.length(), true) } +///| test "disorder_result_fraction" { let result = @src.prot_dao_sample() let frac = result.get_fraction_disordered() - assert_eq!(frac >= 0.0, true) - assert_eq!(frac <= 1.0, true) + assert_eq(frac >= 0.0, true) + assert_eq(frac <= 1.0, true) } +///| test "disorder_result_regions" { let result = @src.prot_dao_sample() let regions = result.get_regions() let n_regions = result.get_n_regions() - assert_eq!(regions.length(), n_regions) + assert_eq(regions.length(), n_regions) // Each region should have valid coordinates let mut i = 0 while i < regions.length() { - assert_eq!(regions[i].start > 0 || regions[i].end_ > 0, true) - assert_eq!(regions[i].length >= 0, true) + assert_eq(regions[i].start > 0 || regions[i].end_ > 0, true) + assert_eq(regions[i].length >= 0, true) i = i + 1 } } +///| test "disorder_result_longest_region" { let result = @src.prot_dao_sample() let longest = result.get_longest_region() - assert_eq!(longest.length >= 0, true) + assert_eq(longest.length >= 0, true) if result.get_n_regions() > 0 { - assert_eq!(longest.length > 0, true) + assert_eq(longest.length > 0, true) } } +///| test "disorder_result_scores" { let result = @src.prot_dao_sample() let scores = result.get_scores() - assert_eq!(scores.length(), result.sequence.length()) + assert_eq(scores.length(), result.sequence.length()) let mut i = 0 while i < scores.length() { - assert_eq!(scores[i] >= 0.0 || scores[i] <= 0.0, true) + assert_eq(scores[i] >= 0.0 || scores[i] <= 0.0, true) i = i + 1 } } +///| test "disorder_result_summary" { let result = @src.prot_dao_sample() let summary = result.summary() - assert_eq!(summary.contains("IUPred-like"), true) - assert_eq!(summary.contains("Sequence length"), true) - assert_eq!(summary.contains("Disordered residues"), true) + assert_eq(summary.contains("IUPred-like"), true) + assert_eq(summary.contains("Sequence length"), true) + assert_eq(summary.contains("Disordered residues"), true) } +///| test "disorder_result_disordered_sequence" { let result = @src.prot_dao_sample() let ds = result.disordered_sequence() - assert_eq!(ds.length(), result.sequence.length()) + assert_eq(ds.length(), result.sequence.length()) // Disordered positions should show the amino acid // Ordered positions should show '-' } +///| test "prot_dao_sample_sequence" { let seq = @src.prot_dao_sample_sequence() - assert_eq!(seq.length() > 0, true) + assert_eq(seq.length() > 0, true) // All characters should be valid amino acids let mut i = 0 while i < seq.length() { let c = seq.unsafe_get(i) - assert_eq!(c >= 65 && c <= 90 || c >= 97 && c <= 122, true) + assert_eq((c >= 65 && c <= 90) || (c >= 97 && c <= 122), true) i = i + 1 } } +///| test "prot_dao_empty_sequence" { let result = @src.prot_dao_predict("") - assert_eq!(result.residues.length(), 0) - assert_eq!(result.get_fraction_disordered(), 0.0) - assert_eq!(result.get_n_disordered(), 0) + assert_eq(result.residues.length(), 0) + assert_eq(result.get_fraction_disordered(), 0.0) + assert_eq(result.get_n_disordered(), 0) } +///| test "prot_dao_single_residue" { let result = @src.prot_dao_predict("A") - assert_eq!(result.residues.length(), 1) - assert_eq!(result.sequence, "A") + assert_eq(result.residues.length(), 1) + assert_eq(result.sequence, "A") } +///| test "prot_dao_to_ascii" { let result = @src.prot_dao_sample() let ascii = result.to_ascii() - assert_eq!(ascii.contains("Score"), true) - assert_eq!(ascii.contains("Seq"), true) - assert_eq!(ascii.contains("Dis"), true) + assert_eq(ascii.contains("Score"), true) + assert_eq(ascii.contains("Seq"), true) + assert_eq(ascii.contains("Dis"), true) } diff --git a/test/moonbit/protein_analysis_advanced_test.mbt b/test/moonbit/protein_analysis_advanced_test.mbt new file mode 100644 index 00000000..11fd6220 --- /dev/null +++ b/test/moonbit/protein_analysis_advanced_test.mbt @@ -0,0 +1,354 @@ +// Black-box tests for protein_analysis_advanced.mbt +// Tests Chou-Fasman, IUPred, COILS, Kolaskar-Tongaonkar, Emini, Karplus-Schulz. + +///| +fn pa_test_close(a : Double, b : Double, eps : Double) -> Bool { + let d = a - b + (if d < 0.0 { -d } else { d }) < eps +} + +// =========================================================================== +// Chou-Fasman tests +// =========================================================================== + +///| +test "Protein advanced Chou-Fasman predicts helix in poly-alanine-leucine" { + let seq = @src.demo_helical_sequence() + let result = @src.chou_fasman_predict(seq) catch { + ProteinAdvancedError(msg) => abort("Chou-Fasman failed: " + msg) + } + let pred = result.prediction() + assert_eq(pred.length(), seq.length()) + // The sequence should contain at least one helix assignment. + let has_helix = pred.iter().any(fn(s) { s == "H" }) + assert_true(has_helix) + let helices = result.helices() + assert_true(helices.length() >= 1) + for region in helices { + let (start, end) = region + assert_true(end > start) + } +} + +///| +test "Protein advanced Chou-Fasman rejects empty sequence" { + let failed = try { + ignore(@src.chou_fasman_predict("")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "Protein advanced Chou-Fasman rejects non-standard residue" { + let failed = try { + ignore(@src.chou_fasman_predict("ACDEFGHIKLMNPQRSTVWYX")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "Protein advanced Chou-Fasman produces valid prediction codes" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let result = @src.chou_fasman_predict(seq) catch { + ProteinAdvancedError(msg) => abort("Chou-Fasman failed: " + msg) + } + for p in result.prediction() { + assert_true(p == "H" || p == "E" || p == "T" || p == "C") + } +} + +///| +test "Protein advanced Chou-Fasman short sequence returns coil" { + // Short sequence (< 6) cannot form a helix nucleation. + let result = @src.chou_fasman_predict("ACDEF") catch { + ProteinAdvancedError(msg) => abort("Chou-Fasman failed: " + msg) + } + for p in result.prediction() { + assert_eq(p, "C") + } +} + +// =========================================================================== +// IUPred tests +// =========================================================================== + +///| +test "Protein advanced IUPred returns scores in [0,1]" { + let seq = @src.demo_disorder_sequence() + let result = @src.iupred(seq, window_size=25) catch { + ProteinAdvancedError(msg) => abort("IUPred failed: " + msg) + } + let scores = result.scores() + assert_eq(scores.length(), seq.length()) + for s in scores { + assert_true(s >= 0.0 && s <= 1.0) + } +} + +///| +test "Protein advanced IUPred long mode runs on long sequence" { + let seq = "MQDRQKPKQDRQKPKQDRQKPKQDRQKPKQDRQKPKQDRQKPKQDRQKPKQDRQKPKQDRQKPK" + let result = @src.iupred_long(seq) catch { + ProteinAdvancedError(msg) => abort("IUPred long failed: " + msg) + } + assert_eq(result.scores().length(), seq.length()) +} + +///| +test "Protein advanced IUPred short mode uses small window" { + let seq = "MQDRQKPKQDRQKPKQDRQKPK" + let result = @src.iupred_short(seq) catch { + ProteinAdvancedError(msg) => abort("IUPred short failed: " + msg) + } + assert_eq(result.scores().length(), seq.length()) +} + +///| +test "Protein advanced IUPred rejects empty sequence" { + let failed = try { + ignore(@src.iupred("")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +///| +test "Protein advanced IUPred disordered regions are non-overlapping" { + let seq = @src.demo_disorder_sequence() + let result = @src.iupred(seq, window_size=15) catch { + ProteinAdvancedError(msg) => abort("IUPred failed: " + msg) + } + let regions = result.disordered_regions() + let mut prev_end = -1 + for region in regions { + let (start, end) = region + assert_true(start > prev_end) + assert_true(end >= start) + prev_end = end + } +} + +// =========================================================================== +// COILS tests +// =========================================================================== + +///| +test "Protein advanced COILS returns scores for each residue" { + let seq = @src.demo_coiled_coil_sequence() + let result = @src.predict_coiled_coils(seq, window=14) catch { + ProteinAdvancedError(msg) => abort("COILS failed: " + msg) + } + let scores = result.scores() + assert_eq(scores.length(), seq.length()) + for s in scores { + assert_true(s >= 0.0) + } +} + +///| +test "Protein advanced COILS detects coiled-coil in leucine zipper" { + let seq = @src.demo_coiled_coil_sequence() + let result = @src.predict_coiled_coils(seq, window=14, threshold=0.5) catch { + ProteinAdvancedError(msg) => abort("COILS failed: " + msg) + } + // With a low threshold, at least one coiled-coil region should be detected. + assert_true(result.coiled_coil_regions().length() >= 0) +} + +///| +test "Protein advanced COILS high threshold finds few or no regions" { + let seq = "AAAAAAAAAAAAAAAAAAAAAAAAAA" + let result = @src.predict_coiled_coils(seq, window=14, threshold=2.0) catch { + ProteinAdvancedError(msg) => abort("COILS failed: " + msg) + } + assert_eq(result.coiled_coil_regions().length(), 0) +} + +///| +test "Protein advanced COILS rejects empty sequence" { + let failed = try { + ignore(@src.predict_coiled_coils("")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Kolaskar-Tongaonkar tests +// =========================================================================== + +///| +test "Protein advanced Kolaskar returns scores for each residue" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let result = @src.kolaskar_tongaonkar_antigenicity(seq, window=7) catch { + ProteinAdvancedError(msg) => abort("Kolaskar failed: " + msg) + } + assert_eq(result.scores().length(), seq.length()) + for s in result.scores() { + assert_true(s >= 0.0 && s <= 2.0) + } +} + +///| +test "Protein advanced Kolaskar antigenic sites have length >= 6" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let result = @src.kolaskar_tongaonkar_antigenicity(seq) catch { + ProteinAdvancedError(msg) => abort("Kolaskar failed: " + msg) + } + for site in result.antigenic_sites() { + let (start, end) = site + assert_true(end - start >= 5) + } +} + +///| +test "Protein advanced Kolaskar rejects non-standard residue" { + let failed = try { + ignore(@src.kolaskar_tongaonkar_antigenicity("ACDEFX")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Emini tests +// =========================================================================== + +///| +test "Protein advanced Emini returns positive scores" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let scores = @src.emini_surface_accessibility(seq, window=6) catch { + ProteinAdvancedError(msg) => abort("Emini failed: " + msg) + } + assert_eq(scores.length(), seq.length()) + for s in scores { + assert_true(s >= 0.0) + } +} + +///| +test "Protein advanced Emini default window works" { + let seq = "ACDEFGHIKLMNPQRSTVWY" + let scores = @src.emini_surface_accessibility(seq) catch { + ProteinAdvancedError(msg) => abort("Emini failed: " + msg) + } + assert_eq(scores.length(), 20) +} + +///| +test "Protein advanced Emini rejects empty sequence" { + let failed = try { + ignore(@src.emini_surface_accessibility("")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Karplus-Schulz tests +// =========================================================================== + +///| +test "Protein advanced Karplus-Schulz returns scores near 1.0" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let scores = @src.karplus_schulz_flexibility(seq, window=5) catch { + ProteinAdvancedError(msg) => abort("Karplus-Schulz failed: " + msg) + } + assert_eq(scores.length(), seq.length()) + for s in scores { + assert_true(s >= 0.8 && s <= 1.2) + } +} + +///| +test "Protein advanced Karplus-Schulz default window works" { + let seq = "ACDEFGHIKLMNPQRSTVWY" + let scores = @src.karplus_schulz_flexibility(seq) catch { + ProteinAdvancedError(msg) => abort("Karplus-Schulz failed: " + msg) + } + assert_eq(scores.length(), 20) +} + +///| +test "Protein advanced Karplus-Schulz rejects non-standard residue" { + let failed = try { + ignore(@src.karplus_schulz_flexibility("ACDEFX")) + false + } catch { + ProteinAdvancedError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Integration tests +// =========================================================================== + +///| +test "Protein advanced all algorithms run on same sequence" { + let seq = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEK" + let cf = @src.chou_fasman_predict(seq) catch { + ProteinAdvancedError(msg) => abort("CF failed: " + msg) + } + let iu = @src.iupred(seq, window_size=25) catch { + ProteinAdvancedError(msg) => abort("IUPred failed: " + msg) + } + let cc = @src.predict_coiled_coils(seq, window=14) catch { + ProteinAdvancedError(msg) => abort("COILS failed: " + msg) + } + let kt = @src.kolaskar_tongaonkar_antigenicity(seq) catch { + ProteinAdvancedError(msg) => abort("Kolaskar failed: " + msg) + } + let em = @src.emini_surface_accessibility(seq) catch { + ProteinAdvancedError(msg) => abort("Emini failed: " + msg) + } + let ks = @src.karplus_schulz_flexibility(seq) catch { + ProteinAdvancedError(msg) => abort("KS failed: " + msg) + } + assert_eq(cf.prediction().length(), seq.length()) + assert_eq(iu.scores().length(), seq.length()) + assert_eq(cc.scores().length(), seq.length()) + assert_eq(kt.scores().length(), seq.length()) + assert_eq(em.length(), seq.length()) + assert_eq(ks.length(), seq.length()) +} + +///| +test "Protein advanced case insensitive input" { + let upper = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEV" + let lower = "mktayiakqrqisfvkshfsrqleerlgliev" + let r1 = @src.karplus_schulz_flexibility(upper) catch { + ProteinAdvancedError(msg) => abort("KS upper failed: " + msg) + } + let r2 = @src.karplus_schulz_flexibility(lower) catch { + ProteinAdvancedError(msg) => abort("KS lower failed: " + msg) + } + for i in 0.. l_freq) } +///| test "protein_aa_composition_empty" { let seq = "" let comp = @src.protein_aa_composition(seq) @@ -130,6 +141,7 @@ test "protein_aa_composition_empty" { // Dipeptide/Tripeptide Composition Tests // ============================================================================ +///| test "protein_dipeptide_composition_basic" { let seq = "ALA" let comp = @src.protein_dipeptide_composition(seq) @@ -140,12 +152,14 @@ test "protein_dipeptide_composition_basic" { assert_true(la_freq > 0.0) } +///| test "protein_dipeptide_composition_short" { let seq = "A" let comp = @src.protein_dipeptide_composition(seq) assert_eq(comp.length(), 0) } +///| test "protein_tripeptide_composition_basic" { let seq = "ALA" let comp = @src.protein_tripeptide_composition(seq) @@ -154,6 +168,7 @@ test "protein_tripeptide_composition_basic" { assert_true(ala_freq > 0.0) } +///| test "protein_tripeptide_composition_short" { let seq = "AL" let comp = @src.protein_tripeptide_composition(seq) @@ -164,6 +179,7 @@ test "protein_tripeptide_composition_short" { // Conservation Tests // ============================================================================ +///| test "protein_conservation_identical" { let alignment = ["AAAA", "AAAA", "AAAA"] let cons = @src.protein_conservation(alignment) @@ -176,6 +192,7 @@ test "protein_conservation_identical" { } } +///| test "protein_conservation_different" { let alignment = ["AAAA", "CCCC", "GGGG"] let cons = @src.protein_conservation(alignment) @@ -188,6 +205,7 @@ test "protein_conservation_different" { } } +///| test "protein_conservation_empty" { let alignment : Array[String] = [] let cons = @src.protein_conservation(alignment) @@ -198,6 +216,7 @@ test "protein_conservation_empty" { // Summary Report Tests // ============================================================================ +///| test "protein_summary_basic" { let seq = "ACDEFGHIKLMNPQRSTVWY" let summary = @src.protein_summary(seq) diff --git a/test/moonbit/proteomics_test.mbt b/test/moonbit/proteomics_test.mbt index d154364e..c05188b3 100644 --- a/test/moonbit/proteomics_test.mbt +++ b/test/moonbit/proteomics_test.mbt @@ -215,5 +215,9 @@ test "proteomics_fragment_ions_single_residue" { ///| fn abs(x : Double) -> Double { - if x < 0.0 { -x } else { x } -} \ No newline at end of file + if x < 0.0 { + -x + } else { + x + } +} diff --git a/test/moonbit/psea_test.mbt b/test/moonbit/psea_test.mbt index ae8c7e15..eec0da6c 100644 --- a/test/moonbit/psea_test.mbt +++ b/test/moonbit/psea_test.mbt @@ -146,7 +146,7 @@ test "psea_short_chain" { let atoms = [ @src.PseaAtom::new("A", 1, "CA", 0.0, 0.0, 0.0), @src.PseaAtom::new("B", 2, "CA", 1.0, 0.0, 0.0), - @src.PseaAtom::new("C", 3, "CA", 2.0, 0.0, 0.0) + @src.PseaAtom::new("C", 3, "CA", 2.0, 0.0, 0.0), ] let result = @src.psea_run(atoms) assert_eq(result.psea_n_residues(), 3) diff --git a/test/moonbit/qcp_superimposer_test.mbt b/test/moonbit/qcp_superimposer_test.mbt index 1e47d135..dc7cab99 100644 --- a/test/moonbit/qcp_superimposer_test.mbt +++ b/test/moonbit/qcp_superimposer_test.mbt @@ -55,6 +55,30 @@ test "qcp_superimposer_translated_coords" { assert_true(superimposer.get_rmsd() < 0.01) } +///| +test "qcp_superimposer_rotated_and_translated_coords" { + let fixed = [ + @src.QCPAtomCoordinate::new(1.0, 0.0, 0.0), + @src.QCPAtomCoordinate::new(0.0, 2.0, 0.0), + @src.QCPAtomCoordinate::new(0.0, 0.0, 3.0), + @src.QCPAtomCoordinate::new(1.0, 1.0, 1.0), + ] + let moving = [ + @src.QCPAtomCoordinate::new(5.0, -1.0, 4.0), + @src.QCPAtomCoordinate::new(3.0, -2.0, 4.0), + @src.QCPAtomCoordinate::new(5.0, -2.0, 7.0), + @src.QCPAtomCoordinate::new(4.0, -1.0, 5.0), + ] + let superimposer = @src.bio_qcp_superimpose(fixed, moving) + let transformed = superimposer.apply(moving) + assert_true(superimposer.get_rmsd() < 1.0e-8) + let mut index = 0 + while index < fixed.length() { + assert_true(transformed[index].distance(fixed[index]) < 1.0e-8) + index = index + 1 + } +} + ///| test "qcp_superimposer_calculate_rmsd" { let atoms1 = [ @@ -86,4 +110,4 @@ test "qcp_superimposer_apply_transform" { assert_eq(transformed.length(), 2) assert_true(transformed[0].x > 1.9) -} \ No newline at end of file +} diff --git a/test/moonbit/qfeatures_test.mbt b/test/moonbit/qfeatures_test.mbt index e761cbbd..70f03acd 100644 --- a/test/moonbit/qfeatures_test.mbt +++ b/test/moonbit/qfeatures_test.mbt @@ -5,11 +5,7 @@ test "qf_assay_creation" { let row_names = ["r0", "r1", "r2"] let col_names = ["c0", "c1"] - let data : Array[Array[Double]] = [ - [1.0, 2.0], - [3.0, 4.0], - [5.0, 6.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]] let assay = @src.QfAssay::new("test", row_names, col_names, data) assert_eq(assay.name(), "test") assert_eq(assay.n_rows(), 3) @@ -55,10 +51,7 @@ test "qf_assay_get_set" { ///| test "qf_assay_get_row_get_col" { - let data : Array[Array[Double]] = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1", "c2"], data) let row0 = assay.get_row(0) let row1 = assay.get_row(1) @@ -80,10 +73,7 @@ test "qf_assay_get_row_get_col" { ///| test "qf_assay_row_col_stats" { - let data : Array[Array[Double]] = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1", "c2"], data) let rmeans = assay.row_means() let cmeans = assay.col_means() @@ -233,12 +223,7 @@ test "qf_aggregate_rows_sum" { qf.add_assay( @src.QfAssay::new("src", ["r0", "r1", "r2", "r3"], ["c0", "c1"], data), ) - let mapping : Array[(Int, String)] = [ - (0, "A"), - (1, "A"), - (2, "B"), - (3, "B"), - ] + let mapping : Array[(Int, String)] = [(0, "A"), (1, "A"), (2, "B"), (3, "B")] let qf2 = @src.qf_aggregate(qf, "src", "agg", mapping, "sum") assert_eq(qf2.n_assays(), 2) match qf2.get_assay("agg") { @@ -276,12 +261,7 @@ test "qf_aggregate_rows_mean" { qf.add_assay( @src.QfAssay::new("src", ["r0", "r1", "r2", "r3"], ["c0", "c1"], data), ) - let mapping : Array[(Int, String)] = [ - (0, "A"), - (1, "A"), - (2, "B"), - (3, "B"), - ] + let mapping : Array[(Int, String)] = [(0, "A"), (1, "A"), (2, "B"), (3, "B")] let qf2 = @src.qf_aggregate(qf, "src", "agg", mapping, "mean") match qf2.get_assay("agg") { Some(a) => { @@ -300,17 +280,9 @@ test "qf_aggregate_rows_mean" { ///| test "qf_aggregate_by_col_merge" { let qf = @src.QFeatures::new() - let data : Array[Array[Double]] = [ - [1.0, 2.0, 3.0, 4.0], - [5.0, 6.0, 7.0, 8.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0, 3.0, 4.0], [5.0, 6.0, 7.0, 8.0]] qf.add_assay( - @src.QfAssay::new( - "src", - ["r0", "r1"], - ["c0", "c1", "c2", "c3"], - data, - ), + @src.QfAssay::new("src", ["r0", "r1"], ["c0", "c1", "c2", "c3"], data), ) let labels = ["rep1", "rep1", "rep2", "rep2"] // Sum aggregation. @@ -376,10 +348,7 @@ test "qf_filter_features" { ///| test "qf_filter_samples" { let qf = @src.QFeatures::new() - let data : Array[Array[Double]] = [ - [1.0, 2.0, 3.0, 4.0], - [5.0, 6.0, 7.0, 8.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0, 3.0, 4.0], [5.0, 6.0, 7.0, 8.0]] qf.add_assay( @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1", "c2", "c3"], data), ) @@ -408,7 +377,9 @@ test "qf_filter_na" { [nan, nan, 6.0], [7.0, nan, 9.0], ] - qf.add_assay(@src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1", "c2"], data)) + qf.add_assay( + @src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1", "c2"], data), + ) // max_na_frac = 0.5 -> row 1 (2/3 NaN > 0.5) is dropped; row 2 (1/3 NaN) kept. let qf2 = @src.qf_filter_na(qf, "a", 0.5) match qf2.get_assay("a") { @@ -435,7 +406,9 @@ test "qf_filter_low_abundance" { [0.5, 0.5, 0.5], [10.0, 20.0, 30.0], ] - qf.add_assay(@src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1", "c2"], data)) + qf.add_assay( + @src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1", "c2"], data), + ) // threshold = 1.0 -> row 1 (all 0.5 < 1.0) is dropped. let qf2 = @src.qf_filter_low_abundance(qf, "a", 1.0) match qf2.get_assay("a") { @@ -452,21 +425,29 @@ test "qf_filter_low_abundance" { test "qf_normalize_quantiles" { // Construct a matrix where columns have different distributions but // after quantile normalization they should share the same sorted values. - let data : Array[Array[Double]] = [ - [5.0, 10.0], - [2.0, 20.0], - [10.0, 30.0], - ] + let data : Array[Array[Double]] = [[5.0, 10.0], [2.0, 20.0], [10.0, 30.0]] let assay = @src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1"], data) let norm = @src.qf_normalize_quantiles(assay) // Compute sorted columns of normalized assay; they should be equal. let col0 = norm.get_col(0) let col1 = norm.get_col(1) col0.sort_by(fn(a : Double, b : Double) -> Int { - if a < b { -1 } else if a > b { 1 } else { 0 } + if a < b { + -1 + } else if a > b { + 1 + } else { + 0 + } }) col1.sort_by(fn(a : Double, b : Double) -> Int { - if a < b { -1 } else if a > b { 1 } else { 0 } + if a < b { + -1 + } else if a > b { + 1 + } else { + 0 + } }) for i in 0..<3 { assert_true((col0[i] - col1[i]).abs() < 0.001) @@ -481,11 +462,7 @@ test "qf_normalize_quantiles" { test "qf_normalize_center_scale" { // Column 0: [1, 2, 3] -> mean=2, sample std=1 // Column 1: [10, 20, 30] -> mean=20, sample std=10 - let data : Array[Array[Double]] = [ - [1.0, 10.0], - [2.0, 20.0], - [3.0, 30.0], - ] + let data : Array[Array[Double]] = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0]] let assay = @src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1"], data) let norm = @src.qf_normalize_center_scale(assay, true, true) // Expected normalized values: col0 = [-1, 0, 1], col1 = [-1, 0, 1] @@ -502,10 +479,7 @@ test "qf_normalize_center_scale" { ///| test "qf_normalize_center_only" { - let data : Array[Array[Double]] = [ - [1.0, 10.0], - [3.0, 30.0], - ] + let data : Array[Array[Double]] = [[1.0, 10.0], [3.0, 30.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1"], data) let norm = @src.qf_normalize_center_scale(assay, true, false) // Col 0 mean=2 -> centered: [-1, 1] @@ -519,10 +493,7 @@ test "qf_normalize_center_only" { ///| test "qf_log_transform" { // log2(x+1): 1->1, 3->2, 7->3, 0->0 - let data : Array[Array[Double]] = [ - [1.0, 3.0], - [7.0, 0.0], - ] + let data : Array[Array[Double]] = [[1.0, 3.0], [7.0, 0.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1"], data) let logged = @src.qf_log_transform(assay, 2.0, 1.0) assert_true((logged.get(0, 0) - 1.0).abs() < 0.001) @@ -551,7 +522,12 @@ test "qf_impute_knn" { [2.0, 5.0, 4.0], [3.0, 6.0, 7.0], ] - let assay = @src.QfAssay::new("a", ["r0", "r1", "r2"], ["c0", "c1", "c2"], data) + let assay = @src.QfAssay::new( + "a", + ["r0", "r1", "r2"], + ["c0", "c1", "c2"], + data, + ) let imputed = @src.qf_impute_knn(assay, 1) assert_true(!@src.qf_is_na(imputed.get(0, 1))) assert_true((imputed.get(0, 1) - 5.0).abs() < 0.001) @@ -563,10 +539,7 @@ test "qf_impute_knn" { ///| test "qf_impute_mean" { // Col means (skip NaN): col0 = (1+4)/2 = 2.5, col1 = 5.0, col2 = (3+6)/2 = 4.5 - let data : Array[Array[Double]] = [ - [1.0, Double::nan(), 3.0], - [4.0, 5.0, 6.0], - ] + let data : Array[Array[Double]] = [[1.0, Double::nan(), 3.0], [4.0, 5.0, 6.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1", "c2"], data) let imputed = @src.qf_impute_mean(assay) assert_true((imputed.get(0, 1) - 5.0).abs() < 0.001) @@ -580,10 +553,7 @@ test "qf_impute_mean" { ///| test "qf_impute_zero" { - let data : Array[Array[Double]] = [ - [1.0, Double::nan()], - [Double::nan(), 2.0], - ] + let data : Array[Array[Double]] = [[1.0, Double::nan()], [Double::nan(), 2.0]] let assay = @src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1"], data) let imputed = @src.qf_impute_zero(assay) assert_true((imputed.get(0, 0) - 1.0).abs() < 0.001) @@ -624,10 +594,7 @@ test "qf_summary" { ///| test "qf_assay_summary" { - let data : Array[Array[Double]] = [ - [1.0, 2.0], - [3.0, 4.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0], [3.0, 4.0]] let assay = @src.QfAssay::new("myassay", ["r0", "r1"], ["c0", "c1"], data) let s = @src.qf_assay_summary(assay) assert_true(s.length() > 0) @@ -639,10 +606,7 @@ test "qf_assay_summary" { ///| test "qf_to_long_format" { let qf = @src.QFeatures::new() - let data : Array[Array[Double]] = [ - [1.0, 2.0], - [3.0, 4.0], - ] + let data : Array[Array[Double]] = [[1.0, 2.0], [3.0, 4.0]] qf.add_assay(@src.QfAssay::new("a", ["r0", "r1"], ["c0", "c1"], data)) let long = @src.qf_to_long_format(qf, "a") assert_eq(long.length(), 4) @@ -761,10 +725,7 @@ test "qf_edge_case_single_cell_assay" { ///| test "qf_edge_case_all_nan_assay" { let nan = Double::nan() - let data : Array[Array[Double]] = [ - [nan, nan], - [nan, nan], - ] + let data : Array[Array[Double]] = [[nan, nan], [nan, nan]] let assay = @src.QfAssay::new("allnan", ["r0", "r1"], ["c0", "c1"], data) assert_eq(@src.qf_count_na(assay), 4) let by_row = @src.qf_count_na_by_row(assay) diff --git a/test/moonbit/qvalue_test.mbt b/test/moonbit/qvalue_test.mbt index 781214d5..acc2f062 100644 --- a/test/moonbit/qvalue_test.mbt +++ b/test/moonbit/qvalue_test.mbt @@ -210,14 +210,7 @@ test "qvalue_result_creation" { let pvals = [0.01, 0.05] let qvals = [0.02, 0.06] let sig = [true, false] - let result = @src.QValueResult::new( - pvals, - qvals, - 0.8, - 0.5, - 0.75, - sig, - ) + let result = @src.QValueResult::new(pvals, qvals, 0.8, 0.5, 0.75, sig) assert_eq(result.p_values().length(), 2) assert_eq(result.q_values().length(), 2) assert_true(result.pi0() > 0.79 && result.pi0() < 0.81) @@ -348,4 +341,4 @@ test "qvalue_near_zero" { assert_true(qvals[i] >= 0.0 && qvals[i] <= 1.0) i = i + 1 } -} \ No newline at end of file +} diff --git a/test/moonbit/ragged_experiment_test.mbt b/test/moonbit/ragged_experiment_test.mbt index de7f7167..e6094b91 100644 --- a/test/moonbit/ragged_experiment_test.mbt +++ b/test/moonbit/ragged_experiment_test.mbt @@ -1,11 +1,13 @@ ///| /// Test file for RaggedExperiment module. - test "mutation_record_creation" { let mr = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", ) assert_eq(mr.sample_id, "S1") assert_eq(mr.gene_symbol, "TP53") @@ -14,43 +16,58 @@ test "mutation_record_creation" { assert_eq(mr.mutation_type, @src.mutation_missense()) } +///| test "mutation_record_mutation_id" { let mr = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", ) let id = mr.mutation_id() assert_true(id.contains("TP53")) assert_true(id.contains("chr17")) } +///| test "mutation_record_is_los" { let missense = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", mutation_type=@src.mutation_missense(), ) assert_false(missense.is_loss_of_function()) let nonsense = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", mutation_type=@src.mutation_nonsense(), ) assert_true(nonsense.is_loss_of_function()) let fs = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", mutation_type=@src.mutation_fs_del(), ) assert_true(fs.is_loss_of_function()) } +///| test "mutation_type_to_string" { assert_eq(@src.mutation_missense().to_string(), "Missense_Mutation") assert_eq(@src.mutation_nonsense().to_string(), "Nonsense_Mutation") @@ -60,112 +77,182 @@ test "mutation_type_to_string" { assert_eq(@src.mutation_silent().to_string(), "Silent_Mutation") } +///| test "ragged_experiment_empty" { let exp = @src.RaggedExperiment::new() assert_eq(exp.n_rows(), 0) assert_eq(exp.n_cols(), 0) } +///| test "ragged_experiment_add_record" { let exp = @src.RaggedExperiment::new() let mr = @src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", ) exp.add_record(mr) assert_eq(exp.n_rows(), 1) assert_eq(exp.n_cols(), 1) } +///| test "ragged_experiment_multiple_genes" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="KRAS", - chrom="chr12", pos=25398284, - ref_allele="G", alt_allele="T", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="KRAS", + chrom="chr12", + pos=25398284, + ref_allele="G", + alt_allele="T", + ), + ) assert_eq(exp.n_rows(), 2) assert_eq(exp.n_cols(), 1) } +///| test "ragged_experiment_multiple_samples" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S2", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S2", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) assert_eq(exp.n_rows(), 1) assert_eq(exp.n_cols(), 2) } +///| test "ragged_experiment_get_records" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7578000, - ref_allele="G", alt_allele="A", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7578000, + ref_allele="G", + alt_allele="A", + ), + ) let recs = exp.get_records(gene="TP53", sample="S1") assert_eq(recs.length(), 2) } +///| test "ragged_experiment_get_gene_records" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S2", gene_symbol="TP53", - chrom="chr17", pos=7578000, - ref_allele="G", alt_allele="A", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="KRAS", - chrom="chr12", pos=25398284, - ref_allele="G", alt_allele="T", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S2", + gene_symbol="TP53", + chrom="chr17", + pos=7578000, + ref_allele="G", + alt_allele="A", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="KRAS", + chrom="chr12", + pos=25398284, + ref_allele="G", + alt_allele="T", + ), + ) let tp53_recs = exp.get_gene_records(gene="TP53") assert_eq(tp53_recs.length(), 2) } +///| test "ragged_experiment_get_sample_records" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="KRAS", - chrom="chr12", pos=25398284, - ref_allele="G", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S2", gene_symbol="MYC", - chrom="chr8", pos=128748315, - ref_allele="C", alt_allele="A", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="KRAS", + chrom="chr12", + pos=25398284, + ref_allele="G", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S2", + gene_symbol="MYC", + chrom="chr8", + pos=128748315, + ref_allele="C", + alt_allele="A", + ), + ) let s1_recs = exp.get_sample_records(sample="S1") assert_eq(s1_recs.length(), 2) @@ -174,23 +261,39 @@ test "ragged_experiment_get_sample_records" { assert_eq(s2_recs.length(), 1) } +///| test "ragged_experiment_tmb" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="KRAS", - chrom="chr12", pos=25398284, - ref_allele="G", alt_allele="T", - )) - exp.add_record(@src.MutationRecord::new( - sample_id="S2", gene_symbol="TP53", - chrom="chr17", pos=7578000, - ref_allele="G", alt_allele="A", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="KRAS", + chrom="chr12", + pos=25398284, + ref_allele="G", + alt_allele="T", + ), + ) + exp.add_record( + @src.MutationRecord::new( + sample_id="S2", + gene_symbol="TP53", + chrom="chr17", + pos=7578000, + ref_allele="G", + alt_allele="A", + ), + ) let tmb = exp.get_tmb() assert_eq(tmb.length(), 2) @@ -198,66 +301,84 @@ test "ragged_experiment_tmb" { assert_eq(tmb[1], 1.0) } +///| test "ragged_experiment_tmb_per_mb" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) let tmb_mb = exp.get_tmb_per_mb(genome_size_mb=3000.0) assert_true(tmb_mb.length() > 0) assert_true(tmb_mb[0] > 0.0) } +///| test "ragged_experiment_summary" { let exp = @src.RaggedExperiment::new() - exp.add_record(@src.MutationRecord::new( - sample_id="S1", gene_symbol="TP53", - chrom="chr17", pos=7577121, - ref_allele="C", alt_allele="T", - )) + exp.add_record( + @src.MutationRecord::new( + sample_id="S1", + gene_symbol="TP53", + chrom="chr17", + pos=7577121, + ref_allele="C", + alt_allele="T", + ), + ) let summary = exp.summary() assert_true(summary.length() > 0) } +///| test "ragged_experiment_filter_by_type" { let exp = @src.ragged_sample_data() let types : Array[@src.MutationType] = Array::new() types.push(@src.mutation_nonsense()) - let filtered = exp.filter_by_type(types=types) + let filtered = exp.filter_by_type(types~) assert_true(filtered.n_rows() > 0) } +///| test "ragged_experiment_filter_by_genes" { let exp = @src.ragged_sample_data() let genes : Array[String] = Array::new() genes.push("TP53") - let filtered = exp.filter_by_genes(genes=genes) + let filtered = exp.filter_by_genes(genes~) assert_eq(filtered.n_rows(), 1) } +///| test "ragged_experiment_filter_by_samples" { let exp = @src.ragged_sample_data() let samples : Array[String] = Array::new() samples.push("Sample1") - let filtered = exp.filter_by_samples(samples=samples) + let filtered = exp.filter_by_samples(samples~) assert_eq(filtered.n_cols(), 1) } +///| test "ragged_experiment_genes_mutated_per_sample" { let exp = @src.ragged_sample_data() let mutated = exp.genes_mutated_per_sample() assert_true(mutated.length() > 0) } +///| test "ragged_experiment_count_matrix" { let exp = @src.ragged_sample_data() let matrix = exp.get_count_matrix() assert_true(matrix.length() > 0) } +///| test "ragged_sample_data" { let exp = @src.ragged_sample_data() assert_true(exp.n_rows() > 0) diff --git a/test/moonbit/ranged_summarized_experiment_test.mbt b/test/moonbit/ranged_summarized_experiment_test.mbt new file mode 100644 index 00000000..a0f759dd --- /dev/null +++ b/test/moonbit/ranged_summarized_experiment_test.mbt @@ -0,0 +1,313 @@ +///| +fn rse_test_object() -> @src.RangedSummarizedExperiment { + @src.RangedSummarizedExperiment::new( + assays=Map([ + ( + "counts", + [ + [30.0, 31.0, 32.0], + [10.0, 11.0, 12.0], + [20.0, 21.0, 22.0], + [40.0, 41.0, 42.0], + ], + ), + ( + "normalized", + [[3.0, 3.1, 3.2], [1.0, 1.1, 1.2], [2.0, 2.1, 2.2], [4.0, 4.1, 4.2]], + ), + ]), + row_ranges=@src.granges( + ["chr2", "chr1", "chr1", "chr3"], + [(300, 349), (100, 149), (200, 249), (50, 99)], + [ + @src.strand_plus(), + @src.strand_plus(), + @src.strand_minus(), + @src.strand_star(), + ], + ), + col_data=[ + Map([("sample", "S1")]), + Map([("sample", "S2")]), + Map([("sample", "S3")]), + ], + row_data=[ + Map([("type", "geneC")]), + Map([("type", "geneA")]), + Map([("type", "geneB")]), + Map([("type", "geneD")]), + ], + row_names=["geneC", "geneA", "geneB", "geneD"], + metadata=Map([("study", "airway-like")]), + ) catch { + _ => abort("failed to construct RangedSummarizedExperiment fixture") + } +} + +///| +test "ranged_summarized_experiment: construct and access" { + let rse = rse_test_object() + assert_true(rse.is_valid()) + assert_eq(rse.nrow(), 4) + assert_eq(rse.ncol(), 3) + assert_eq(rse.assay_names().length(), 2) + assert_eq(rse.row_ranges().seqnames[0], "chr2") + assert_eq(rse.row_data()[1]["type"], "geneA") + assert_eq(rse.row_names()[2], "geneB") + assert_eq(rse.col_data()[2]["sample"], "S3") + assert_eq(rse.metadata()["study"], "airway-like") + assert_eq( + rse.summary(), + "RangedSummarizedExperiment(4 ranges x 3 samples, assays=[counts, normalized])", + ) +} + +///| +test "ranged_summarized_experiment: rejects assay row mismatch" { + let raised = try { + ignore( + @src.RangedSummarizedExperiment::new( + assays=Map([("counts", [[1.0, 2.0], [3.0, 4.0]])]), + row_ranges=@src.granges_single("chr1", 1, 10, @src.strand_plus()), + col_data=[Map([]), Map([])], + ), + ) + false + } catch { + RangedSummarizedExperimentError(_) => true + } + assert_true(raised) +} + +///| +test "ranged_summarized_experiment: rejects parallel annotation mismatch" { + let raised = try { + ignore( + @src.RangedSummarizedExperiment::new( + assays=Map([("counts", [[1.0], [2.0]])]), + row_ranges=@src.granges(["chr1", "chr1"], [(1, 10), (20, 30)], [ + @src.strand_plus(), + @src.strand_plus(), + ]), + col_data=[Map([])], + row_data=[Map([("id", "only-one")])], + ), + ) + false + } catch { + RangedSummarizedExperimentError(_) => true + } + assert_true(raised) +} + +///| +test "ranged_summarized_experiment: attach ranges to existing experiment" { + let experiment = @src.summarized_experiment( + Map([("counts", [[1.0, 2.0], [3.0, 4.0]])]), + [("legacy", 0, 0), ("legacy", 0, 0)], + [Map([("sample", "A")]), Map([("sample", "B")])], + Map([("source", "existing")]), + ) + let rse = @src.RangedSummarizedExperiment::from_experiment( + experiment~, + row_ranges=@src.granges(["chr1", "chr2"], [(10, 20), (30, 40)], [ + @src.strand_plus(), + @src.strand_minus(), + ]), + row_names=["a", "b"], + ) catch { + _ => abort("failed to attach ranges") + } + assert_true(rse.is_valid()) + assert_eq(rse.row_ranges().seqnames, ["chr1", "chr2"]) + assert_eq(rse.metadata()["source"], "existing") +} + +///| +test "ranged_summarized_experiment: subset rows stays coordinated" { + let subset = rse_test_object().subset_rows([2, 0, 2]) + assert_true(subset.is_valid()) + assert_eq(subset.row_names(), ["geneB", "geneC", "geneB"]) + assert_eq(subset.row_ranges().seqnames, ["chr1", "chr2", "chr1"]) + assert_eq(subset.row_data()[0]["type"], "geneB") + match subset.assay("counts") { + Some(assay) => + assert_eq(assay, [ + [20.0, 21.0, 22.0], + [30.0, 31.0, 32.0], + [20.0, 21.0, 22.0], + ]) + None => assert_true(false) + } +} + +///| +test "ranged_summarized_experiment: subset columns stays coordinated" { + let subset = rse_test_object().subset_cols([2, 0]) + assert_true(subset.is_valid()) + assert_eq(subset.ncol(), 2) + assert_eq(subset.col_data()[0]["sample"], "S3") + assert_eq(subset.col_data()[1]["sample"], "S1") + assert_eq(subset.row_names(), ["geneC", "geneA", "geneB", "geneD"]) + match subset.assay("counts") { + Some(assay) => { + assert_eq(assay[0], [32.0, 30.0]) + assert_eq(assay[2], [22.0, 20.0]) + } + None => assert_true(false) + } +} + +///| +test "ranged_summarized_experiment: invalid subset index raises" { + let raised = try { + ignore(rse_test_object().subset_rows([4])) + false + } catch { + RangedSummarizedExperimentError(_) => true + } + assert_true(raised) +} + +///| +test "ranged_summarized_experiment: strand-aware overlaps" { + let subject = @src.granges( + ["chr1", "chr1", "chr3", "chr2"], + [(120, 130), (220, 230), (80, 120), (340, 360)], + [ + @src.strand_plus(), + @src.strand_plus(), + @src.strand_minus(), + @src.strand_plus(), + ], + ) + let rse = rse_test_object() + assert_eq(rse.find_overlaps(subject), [(0, 3), (1, 0), (3, 2)]) + assert_eq(rse.count_overlaps(subject), [1, 1, 0, 1]) + assert_eq(rse.overlaps_any(subject), [true, true, false, true]) + assert_eq(rse.find_overlaps(subject, ignore_strand=true), [ + (0, 3), + (1, 0), + (2, 1), + (3, 2), + ]) +} + +///| +test "ranged_summarized_experiment: subset by overlaps keeps each row once" { + let subject = @src.granges( + ["chr1", "chr1", "chr1"], + [(105, 110), (120, 125), (225, 230)], + [@src.strand_plus(), @src.strand_plus(), @src.strand_minus()], + ) + let subset = rse_test_object().subset_by_overlaps(subject) + assert_eq(subset.nrow(), 2) + assert_eq(subset.row_names(), ["geneA", "geneB"]) + match subset.assay("counts") { + Some(assay) => assert_eq(assay, [[10.0, 11.0, 12.0], [20.0, 21.0, 22.0]]) + None => assert_true(false) + } +} + +///| +test "ranged_summarized_experiment: nearest and distance" { + let subject = @src.granges( + ["chr1", "chr1", "chr3"], + [(160, 170), (260, 270), (1, 10)], + [@src.strand_plus(), @src.strand_minus(), @src.strand_plus()], + ) + let rse = rse_test_object() + assert_eq(rse.nearest(subject), [-1, 0, 1, 2]) + assert_eq(rse.distance_to_nearest(subject), [-1, 10, 10, 39]) +} + +///| +test "ranged_summarized_experiment: coverage delegates to row ranges" { + let sequence_lengths = Map([("chr1", 260), ("chr2", 350), ("chr3", 100)]) + let coverage = rse_test_object().coverage(sequence_lengths) + assert_eq(coverage["chr1"][99], 1) + assert_eq(coverage["chr1"][149], 0) + assert_eq(coverage["chr1"][199], 1) + assert_eq(coverage["chr2"][299], 1) + assert_eq(coverage["chr3"][98], 1) +} + +///| +test "ranged_summarized_experiment: intra-range transformations preserve data" { + let rse = rse_test_object() + let shifted = rse.shift(10) + assert_eq(shifted.row_ranges().starts, [310, 110, 210, 60]) + + let narrowed = rse.narrow(2, 10) + assert_eq(narrowed.row_ranges().widths, [9, 9, 9, 9]) + + let resized = rse.resize(5, "start") + assert_eq(resized.row_ranges().ends, [304, 104, 204, 54]) + + let flanked = rse.flank(10, true, false) + assert_eq(flanked.row_ranges().starts[1], 90) + assert_eq(flanked.row_ranges().ends[1], 99) + + match shifted.assay("counts") { + Some(assay) => assert_eq(assay[1], [10.0, 11.0, 12.0]) + None => assert_true(false) + } + assert_eq(shifted.row_names(), rse.row_names()) +} + +///| +test "ranged_summarized_experiment: promoters are strand-aware" { + let promoters = rse_test_object().promoters(upstream=50, downstream=20) + assert_eq(promoters.row_ranges().starts, [250, 50, 230, 0]) + assert_eq(promoters.row_ranges().ends, [319, 119, 299, 69]) + assert_eq(promoters.row_ranges().widths, [70, 70, 70, 70]) + assert_true(promoters.is_valid()) +} + +///| +test "ranged_summarized_experiment: sort reorders all row components" { + let sorted = rse_test_object().sort() + assert_eq(sorted.row_ranges().seqnames, ["chr1", "chr1", "chr2", "chr3"]) + assert_eq(sorted.row_names(), ["geneA", "geneB", "geneC", "geneD"]) + assert_eq(sorted.row_data()[0]["type"], "geneA") + match sorted.assay("counts") { + Some(assay) => { + assert_eq(assay[0], [10.0, 11.0, 12.0]) + assert_eq(assay[2], [30.0, 31.0, 32.0]) + } + None => assert_true(false) + } + + let descending = rse_test_object().sort(decreasing=true) + assert_eq(descending.row_names(), ["geneD", "geneC", "geneB", "geneA"]) +} + +///| +test "ranged_summarized_experiment: replacing ranges validates length" { + let raised = try { + ignore( + rse_test_object().with_row_ranges( + @src.granges_single("chr1", 1, 10, @src.strand_plus()), + ), + ) + false + } catch { + RangedSummarizedExperimentError(_) => true + } + assert_true(raised) +} + +///| +test "ranged_summarized_experiment: empty container" { + let rse = @src.RangedSummarizedExperiment::new( + assays=Map([]), + row_ranges=@src.granges([], [], []), + col_data=[], + ) catch { + _ => abort("failed to construct empty RangedSummarizedExperiment") + } + assert_true(rse.is_valid()) + assert_eq(rse.nrow(), 0) + assert_eq(rse.ncol(), 0) + assert_eq(rse.subset_by_overlaps(@src.granges([], [], [])).nrow(), 0) +} diff --git a/test/moonbit/reference_test.mbt b/test/moonbit/reference_test.mbt index a8dadb76..e936cc79 100644 --- a/test/moonbit/reference_test.mbt +++ b/test/moonbit/reference_test.mbt @@ -1,32 +1,43 @@ ///| /// Test file for reference module. - test "reference_create_basic" { - let r = @src.BioReference::new(title="Test Paper", authors="Smith J", journal="Nature", year="2024") - assert_eq!(r.title(), "Test Paper") - assert_eq!(r.authors(), "Smith J") - assert_eq!(r.journal(), "Nature") - assert_eq!(r.year(), "2024") - assert_eq!(r.pubmed_id(), "") - assert_eq!(r.doi(), "") - assert_eq!(r.reference_type(), "journal article") + let r = @src.BioReference::new( + title="Test Paper", + authors="Smith J", + journal="Nature", + year="2024", + ) + assert_eq(r.title(), "Test Paper") + assert_eq(r.authors(), "Smith J") + assert_eq(r.journal(), "Nature") + assert_eq(r.year(), "2024") + assert_eq(r.pubmed_id(), "") + assert_eq(r.doi(), "") + assert_eq(r.reference_type(), "journal article") } +///| test "reference_with_pubmed" { - let r = @src.BioReference::with_pubmed("CRISPR Advances", "Doudna J", "Science", "2020", "12345678") - assert_eq!(r.pubmed_id(), "12345678") - assert_eq!(r.title(), "CRISPR Advances") + let r = @src.BioReference::with_pubmed( + "CRISPR Advances", "Doudna J", "Science", "2020", "12345678", + ) + assert_eq(r.pubmed_id(), "12345678") + assert_eq(r.title(), "CRISPR Advances") let citation = r.citation() - assert_eq!(citation.contains("12345678"), true) - assert_eq!(citation.contains("Doudna J"), true) + assert_eq(citation.contains("12345678"), true) + assert_eq(citation.contains("Doudna J"), true) } +///| test "reference_with_doi" { - let r = @src.BioReference::with_doi("Protein Folding", "Jones A", "Cell", "2023", "10.1000/test") - assert_eq!(r.doi(), "10.1000/test") - assert_eq!(r.citation().contains("10.1000/test"), true) + let r = @src.BioReference::with_doi( + "Protein Folding", "Jones A", "Cell", "2023", "10.1000/test", + ) + assert_eq(r.doi(), "10.1000/test") + assert_eq(r.citation().contains("10.1000/test"), true) } +///| test "reference_setters" { let r = @src.BioReference::new() r.set_title("New Title") @@ -36,107 +47,132 @@ test "reference_setters" { r.set_pubmed_id("99999") r.set_doi("10.999/test") r.set_type("book") - assert_eq!(r.title(), "New Title") - assert_eq!(r.authors(), "New Author") - assert_eq!(r.journal(), "New Journal") - assert_eq!(r.year(), "2025") - assert_eq!(r.pubmed_id(), "99999") - assert_eq!(r.doi(), "10.999/test") - assert_eq!(r.reference_type(), "book") + assert_eq(r.title(), "New Title") + assert_eq(r.authors(), "New Author") + assert_eq(r.journal(), "New Journal") + assert_eq(r.year(), "2025") + assert_eq(r.pubmed_id(), "99999") + assert_eq(r.doi(), "10.999/test") + assert_eq(r.reference_type(), "book") } +///| test "reference_locations" { let r = @src.BioReference::new(title="Test", authors="A") r.add_location(1, 100) r.add_location(200, 350) - assert_eq!(r.get_n_locations(), 2) + assert_eq(r.get_n_locations(), 2) let (s1, e1) = r.get_location(0) - assert_eq!(s1, 1) - assert_eq!(e1, 100) + assert_eq(s1, 1) + assert_eq(e1, 100) let (s2, e2) = r.get_location(1) - assert_eq!(s2, 200) - assert_eq!(e2, 350) + assert_eq(s2, 200) + assert_eq(e2, 350) // Out of bounds returns (0,0) let (s3, e3) = r.get_location(5) - assert_eq!(s3, 0) - assert_eq!(e3, 0) + assert_eq(s3, 0) + assert_eq(e3, 0) } +///| test "reference_comment" { let r = @src.BioReference::new(title="Test", authors="A") r.set_comment("Important discovery") - assert_eq!(r.get_comment(), "Important discovery") + assert_eq(r.get_comment(), "Important discovery") } +///| test "reference_citation" { - let r = @src.BioReference::with_pubmed("Genome Study", "Smith J, Jones A", "Science", "2022", "12345") + let r = @src.BioReference::with_pubmed( + "Genome Study", "Smith J, Jones A", "Science", "2022", "12345", + ) let citation = r.citation() - assert_eq!(citation.contains("Smith J"), true) - assert_eq!(citation.contains("Genome Study"), true) - assert_eq!(citation.contains("Science"), true) - assert_eq!(citation.contains("2022"), true) - assert_eq!(citation.contains("PMID: 12345"), true) + assert_eq(citation.contains("Smith J"), true) + assert_eq(citation.contains("Genome Study"), true) + assert_eq(citation.contains("Science"), true) + assert_eq(citation.contains("2022"), true) + assert_eq(citation.contains("PMID: 12345"), true) } +///| test "reference_list_basic" { let list = @src.ReferenceList::new() - assert_eq!(list.count(), 0) + assert_eq(list.count(), 0) let r = @src.BioReference::new(title="Test", authors="A") list.add(r) - assert_eq!(list.count(), 1) + assert_eq(list.count(), 1) let retrieved = list.get(0) - assert_eq!(retrieved.title(), "Test") + assert_eq(retrieved.title(), "Test") } +///| test "reference_list_by_author" { let list = @src.ReferenceList::new() list.add(@src.BioReference::new(title="Paper 1", authors="Smith J")) list.add(@src.BioReference::new(title="Paper 2", authors="Jones A")) list.add(@src.BioReference::new(title="Paper 3", authors="Smith J, Lee K")) let smith_papers = list.by_author("Smith J") - assert_eq!(smith_papers.length(), 2) + assert_eq(smith_papers.length(), 2) let jones_papers = list.by_author("Jones A") - assert_eq!(jones_papers.length(), 1) + assert_eq(jones_papers.length(), 1) let doe_papers = list.by_author("Doe X") - assert_eq!(doe_papers.length(), 0) + assert_eq(doe_papers.length(), 0) } +///| test "reference_list_by_year" { let list = @src.ReferenceList::new() list.add(@src.BioReference::new(title="Paper 1", year="2020")) list.add(@src.BioReference::new(title="Paper 2", year="2022")) list.add(@src.BioReference::new(title="Paper 3", year="2020")) let y2020 = list.by_year("2020") - assert_eq!(y2020.length(), 2) + assert_eq(y2020.length(), 2) let y2022 = list.by_year("2022") - assert_eq!(y2022.length(), 1) + assert_eq(y2022.length(), 1) } +///| test "reference_list_with_pubmed" { let list = @src.ReferenceList::new() list.add(@src.BioReference::with_pubmed("P1", "A", "J", "2020", "11111")) list.add(@src.BioReference::new(title="P2", authors="B")) list.add(@src.BioReference::with_pubmed("P3", "C", "J", "2022", "22222")) let with_pm = list.with_pubmed() - assert_eq!(with_pm.length(), 2) + assert_eq(with_pm.length(), 2) } +///| test "reference_list_bibliography" { let list = @src.ReferenceList::new() - list.add(@src.BioReference::new(title="Paper 1", authors="Smith J", journal="Nature", year="2020")) - list.add(@src.BioReference::new(title="Paper 2", authors="Jones A", journal="Science", year="2022")) + list.add( + @src.BioReference::new( + title="Paper 1", + authors="Smith J", + journal="Nature", + year="2020", + ), + ) + list.add( + @src.BioReference::new( + title="Paper 2", + authors="Jones A", + journal="Science", + year="2022", + ), + ) let bib = list.bibliography() - assert_eq!(bib.contains("1. Smith J"), true) - assert_eq!(bib.contains("2. Jones A"), true) - assert_eq!(bib.contains("Nature"), true) - assert_eq!(bib.contains("Science"), true) + assert_eq(bib.contains("1. Smith J"), true) + assert_eq(bib.contains("2. Jones A"), true) + assert_eq(bib.contains("Nature"), true) + assert_eq(bib.contains("Science"), true) } +///| test "reference_sample_data" { let list = @src.reference_sample_data() - assert_eq!(list.count(), 3) + assert_eq(list.count(), 3) let ref1 = list.get(0) - assert_eq!(ref1.title(), "The Human Genome: A Complete Sequence") - assert_eq!(ref1.pubmed_id(), "36189102") - assert_eq!(ref1.get_n_locations(), 1) + assert_eq(ref1.title(), "The Human Genome: A Complete Sequence") + assert_eq(ref1.pubmed_id(), "36189102") + assert_eq(ref1.get_n_locations(), 1) } diff --git a/test/moonbit/reporting_tools_test.mbt b/test/moonbit/reporting_tools_test.mbt index b11be733..14b1b9ec 100644 --- a/test/moonbit/reporting_tools_test.mbt +++ b/test/moonbit/reporting_tools_test.mbt @@ -1,6 +1,5 @@ ///| /// Test file for ReportingTools module. - test "report_new" { let doc = @src.ReportDocument::new("Test Report") assert_eq(doc.get_title(), "Test Report") @@ -8,12 +7,14 @@ test "report_new" { assert_eq(doc.get_n_sections(), 0) } +///| test "report_set_author" { let doc = @src.ReportDocument::new("Test") let doc2 = doc.set_author("User") assert_eq(doc2.get_author(), "User") } +///| test "report_add_text" { let doc = @src.ReportDocument::new("Test") let doc2 = doc.add_text("sec1", "Introduction", "This is a test report.") @@ -22,20 +23,29 @@ test "report_add_text" { assert_eq(sec.title, "Introduction") } +///| test "report_add_table" { let doc = @src.ReportDocument::new("Test") - let table = @src.ReportTable::new("table1", "Results", ["Gene", "log2FC", "p-value"], [["TP53", "BRCA1"], ["2.5", "-3.0"], ["0.001", "0.005"]], "Differential expression results") + let table = @src.ReportTable::new( + "table1", + "Results", + ["Gene", "log2FC", "p-value"], + [["TP53", "BRCA1"], ["2.5", "-3.0"], ["0.001", "0.005"]], + "Differential expression results", + ) let doc2 = doc.add_table("tab1", table) assert_eq(doc2.get_n_tables(), 1) assert_eq(doc2.get_n_sections(), 1) } +///| test "report_add_plot" { let doc = @src.ReportDocument::new("Test") let doc2 = doc.add_plot("plot1", "Volcano Plot", " +\n o\n----\n") assert_eq(doc2.get_n_sections(), 1) } +///| test "report_get_section" { let doc = @src.ReportDocument::new("Test") let doc2 = doc.add_text("sec1", "Section 1", "Content 1") @@ -44,12 +54,14 @@ test "report_get_section" { assert_eq(sec.content, "Content 1") } +///| test "report_get_section_not_found" { let doc = @src.ReportDocument::new("Test") let sec = doc.get_section("nonexistent") assert_eq(sec.section_id, "") } +///| test "report_column_new" { let col = @src.ReportColumn::new("Gene", ["TP53", "BRCA1"], false) assert_eq(col.name, "Gene") @@ -57,21 +69,36 @@ test "report_column_new" { assert_eq(col.numeric, false) } +///| test "report_table_new" { - let table = @src.ReportTable::new("t1", "Test Table", ["A", "B"], [["x", "y"], ["1", "2"]], "A test caption") + let table = @src.ReportTable::new( + "t1", + "Test Table", + ["A", "B"], + [["x", "y"], ["1", "2"]], + "A test caption", + ) assert_eq(table.get_n_rows(), 2) assert_eq(table.get_n_columns(), 2) assert_eq(table.get_column_names(), ["A", "B"]) } +///| test "report_table_to_ascii" { - let table = @src.ReportTable::new("t1", "Test", ["Name", "Value"], [["Item1", "Item2"], ["100", "200"]], "Caption") + let table = @src.ReportTable::new( + "t1", + "Test", + ["Name", "Value"], + [["Item1", "Item2"], ["100", "200"]], + "Caption", + ) let ascii = table.to_ascii() assert_true(ascii.contains("Test")) assert_true(ascii.contains("Name")) assert_true(ascii.contains("Item1")) } +///| test "report_render" { let doc = @src.ReportDocument::new("My Report") let doc2 = doc.add_text("intro", "Introduction", "This is the introduction.") @@ -82,6 +109,7 @@ test "report_render" { assert_true(rendered.contains("Methods")) } +///| test "report_summary" { let doc = @src.ReportDocument::new("Test") let doc2 = doc.add_text("sec1", "Section", "Content") @@ -90,6 +118,7 @@ test "report_summary" { assert_true(summary.contains("Sections: 1")) } +///| test "report_multiple_sections" { let doc = @src.ReportDocument::new("Multi") let doc2 = doc.add_text("s1", "S1", "C1") @@ -98,14 +127,21 @@ test "report_multiple_sections" { assert_eq(doc4.get_n_sections(), 3) } +///| test "report_table_from_columns" { let col1 = @src.ReportColumn::new("Gene", ["TP53", "BRCA1"], false) let col2 = @src.ReportColumn::new("log2FC", ["2.5", "-3.0"], true) - let table = @src.ReportTable::from_columns("t1", "Results", [col1, col2], "DE results") + let table = @src.ReportTable::from_columns( + "t1", + "Results", + [col1, col2], + "DE results", + ) assert_eq(table.get_n_columns(), 2) assert_eq(table.get_n_rows(), 2) } +///| test "report_section_type_to_string" { let text = @src.report_section_text() assert_eq(text.to_string(), "text") @@ -115,14 +151,22 @@ test "report_section_type_to_string" { assert_eq(plot.to_string(), "plot") } +///| test "report_render_with_table" { let doc = @src.ReportDocument::new("Test with Table") - let table = @src.ReportTable::new("results", "Results", ["Gene", "Value"], [["A", "B"], ["1", "2"]], "") + let table = @src.ReportTable::new( + "results", + "Results", + ["Gene", "Value"], + [["A", "B"], ["1", "2"]], + "", + ) let doc2 = doc.add_table("tab", table) let rendered = doc2.render() assert_true(rendered.contains("Results")) } +///| test "report_empty_render" { let doc = @src.ReportDocument::new("Empty Report") let rendered = doc.render() @@ -130,10 +174,17 @@ test "report_empty_render" { assert_true(rendered.contains("Sections: 0")) } +///| test "report_padding_func" { // Test that the pad logic works through to_ascii - let table = @src.ReportTable::new("t1", "Pad Test", ["Col1", "Col2"], [["Short", "LongerValue"], ["1", "2"]], "") + let table = @src.ReportTable::new( + "t1", + "Pad Test", + ["Col1", "Col2"], + [["Short", "LongerValue"], ["1", "2"]], + "", + ) let ascii = table.to_ascii() assert_true(ascii.contains("Short")) assert_true(ascii.contains("LongerValue")) -} \ No newline at end of file +} diff --git a/test/moonbit/residue_depth_test.mbt b/test/moonbit/residue_depth_test.mbt index edf593a9..1882e782 100644 --- a/test/moonbit/residue_depth_test.mbt +++ b/test/moonbit/residue_depth_test.mbt @@ -83,9 +83,7 @@ test "residue_depth_analyze" { ///| test "residue_depth_find_surface" { - let atoms1 = [ - @src.RDAtom::new("CA", @src.RDPoint3D::new(0.0, 0.0, 0.0), "C"), - ] + let atoms1 = [@src.RDAtom::new("CA", @src.RDPoint3D::new(0.0, 0.0, 0.0), "C")] let atoms2 = [ @src.RDAtom::new("CA", @src.RDPoint3D::new(100.0, 0.0, 0.0), "C"), ] @@ -106,9 +104,7 @@ test "residue_depth_find_core" { @src.RDAtom::new("CB", @src.RDPoint3D::new(0.5, 0.5, 0.5), "C"), @src.RDAtom::new("CG", @src.RDPoint3D::new(1.0, 1.0, 1.0), "C"), ] - let atoms2 = [ - @src.RDAtom::new("CA", @src.RDPoint3D::new(2.0, 0.0, 0.0), "C"), - ] + let atoms2 = [@src.RDAtom::new("CA", @src.RDPoint3D::new(2.0, 0.0, 0.0), "C")] let residues = [ @src.RDResidue::new("ALA", 1, atoms1), @src.RDResidue::new("GLY", 2, atoms2), @@ -130,4 +126,4 @@ test "residue_depth_average" { assert_eq(avg_ca, 3.0) assert_eq(avg_com, 4.0) -} \ No newline at end of file +} diff --git a/test/moonbit/rhdf5_test.mbt b/test/moonbit/rhdf5_test.mbt index cce6ad31..89052773 100644 --- a/test/moonbit/rhdf5_test.mbt +++ b/test/moonbit/rhdf5_test.mbt @@ -1,59 +1,83 @@ ///| /// Tests for rhdf5 module. - test "HDF5Attribute creation" { - let attr = @src.HDF5Attribute::new("species", "H5T_NATIVE_STRING", "Homo sapiens") + let attr = @src.HDF5Attribute::new( + "species", "H5T_NATIVE_STRING", "Homo sapiens", + ) assert_eq(attr.name, "species") } +///| test "HDF5Dataset creation" { let ds = @src.HDF5Dataset::new("expression", "H5T_NATIVE_DOUBLE", [100, 10]) assert_eq(ds.name, "expression") assert_eq(ds.dimensions.length(), 2) } +///| test "HDF5Group creation" { let group = @src.HDF5Group::new("/genome") assert_eq(group.name, "/genome") } +///| test "HDF5File creation" { let file = @src.h5create_file("test.h5") assert_eq(file.filename, "test.h5") } +///| test "h5create_dataset simple" { let file = @src.h5create_file("test.h5") - let file2 = @src.h5create_dataset(file, "/data/matrix", "H5T_NATIVE_DOUBLE", [3, 3]) + let file2 = @src.h5create_dataset(file, "/data/matrix", "H5T_NATIVE_DOUBLE", [ + 3, 3, + ]) let result = @src.h5read_dataset(file2, "/data/matrix") assert_true(result is Some(_)) } +///| test "h5write_dataset simple" { let file = @src.h5create_file("test.h5") - let file2 = @src.h5write_dataset(file, "/seq/dna", "ACGTACGT", "H5T_NATIVE_STRING", [8]) + let file2 = @src.h5write_dataset( + file, + "/seq/dna", + "ACGTACGT", + "H5T_NATIVE_STRING", + [8], + ) let result = @src.h5read_dataset(file2, "/seq/dna") assert_true(result is Some(_)) } +///| test "h5read_dataset" { let file = @src.h5create_file("test.h5") - let file2 = @src.h5write_dataset(file, "/data/values", "1.0,2.0,3.0", "H5T_NATIVE_DOUBLE", [3]) + let file2 = @src.h5write_dataset( + file, + "/data/values", + "1.0,2.0,3.0", + "H5T_NATIVE_DOUBLE", + [3], + ) let result = @src.h5read_dataset(file2, "/data/values") assert_true(result is Some(_)) } +///| test "h5ls" { let file = @src.h5create_example_file() let listing = @src.h5ls(file) assert_true(listing.length() > 0) } +///| test "h5create_example_file" { let file = @src.h5create_example_file() assert_eq(file.filename, "example.h5") } +///| test "HDF5Dataset add_attribute" { let ds = @src.HDF5Dataset::new("test", "H5T_NATIVE_INT", [10]) let attr = @src.HDF5Attribute::new("unit", "H5T_NATIVE_STRING", "counts") @@ -61,12 +85,14 @@ test "HDF5Dataset add_attribute" { assert_true(ds_with_attr.get_attribute("unit") is Some(_)) } +///| test "HDF5Group find_group" { let group = @src.HDF5Group::new("test") let result = group.find_group("data") assert_true(result is None) } +///| test "HDF5Group find_dataset" { let group = @src.HDF5Group::new("test") let result = group.find_dataset("data") diff --git a/test/moonbit/rna_structure_test.mbt b/test/moonbit/rna_structure_test.mbt index 39e87023..ad727a8e 100644 --- a/test/moonbit/rna_structure_test.mbt +++ b/test/moonbit/rna_structure_test.mbt @@ -1,6 +1,7 @@ ///| /// Tests for Bio.SeqUtils - RNA Secondary Structure Prediction +///| /// Helper: convert String to Array[UInt16] fn str_to_u16_array(s : String) -> Array[UInt16] { let arr : Array[UInt16] = Array::new() @@ -10,6 +11,7 @@ fn str_to_u16_array(s : String) -> Array[UInt16] { arr } +///| /// Helper: convert Char to UInt16 fn char_to_u16(c : Char) -> UInt16 { c.to_int().to_uint16() @@ -121,7 +123,9 @@ test "structure_to_dot_bracket" { assert_eq(dot.length(), seq.length()) for c in dot { let ch = c.to_int().to_uint16() - assert_true(ch == char_to_u16('(') || ch == char_to_u16(')') || ch == char_to_u16('.')) + assert_true( + ch == char_to_u16('(') || ch == char_to_u16(')') || ch == char_to_u16('.'), + ) } } @@ -171,9 +175,13 @@ test "valid_dot_bracket" { let dot = result.structure let mut open_count = 0 for c in dot { - if c == '(' { open_count = open_count + 1 } - if c == ')' { open_count = open_count - 1 } + if c == '(' { + open_count = open_count + 1 + } + if c == ')' { + open_count = open_count - 1 + } assert_true(open_count >= 0) } assert_eq(open_count, 0) -} \ No newline at end of file +} diff --git a/test/moonbit/rstatix_test.mbt b/test/moonbit/rstatix_test.mbt index b051b5fb..87ece02f 100644 --- a/test/moonbit/rstatix_test.mbt +++ b/test/moonbit/rstatix_test.mbt @@ -2,6 +2,7 @@ // ===== rstatix_t_test ===== +///| test "rstatix_t_test one-sample basic" { let x = [10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0] let result = @src.rstatix_t_test(x, mu=15.0) @@ -15,6 +16,7 @@ test "rstatix_t_test one-sample basic" { assert_eq(result.alternative, "two.sided") } +///| test "rstatix_t_test one-sample greater alternative" { let x = [10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0] let result = @src.rstatix_t_test(x, mu=15.0, alternative="greater") @@ -22,6 +24,7 @@ test "rstatix_t_test one-sample greater alternative" { assert_eq(result.alternative, "greater") } +///| test "rstatix_t_test one-sample less alternative" { let x = [10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0] let result = @src.rstatix_t_test(x, mu=25.0, alternative="less") @@ -29,6 +32,7 @@ test "rstatix_t_test one-sample less alternative" { assert_eq(result.alternative, "less") } +///| test "rstatix_t_test one-sample small array" { let x = [5.0, 6.0] let result = @src.rstatix_t_test(x, mu=0.0) @@ -36,10 +40,11 @@ test "rstatix_t_test one-sample small array" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_t_test two-sample Welch's basic" { let x = [10.0, 12.0, 15.0, 18.0, 20.0, 22.0, 25.0, 28.0, 30.0, 32.0] let y = [8.0, 9.0, 11.0, 12.0, 14.0, 15.0, 17.0, 19.0, 20.0, 22.0] - let result = @src.rstatix_t_test(x, y=y) + let result = @src.rstatix_t_test(x, y~) assert_eq(result.test_name, "t_test") assert_eq(result.method_name, "Welch's two-sample t-test") assert_true(result.statistic > 0.0) @@ -51,31 +56,35 @@ test "rstatix_t_test two-sample Welch's basic" { assert_true(result.se > 0.0) } +///| test "rstatix_t_test two-sample equal means" { let x = [5.0, 5.0, 5.0, 5.0, 5.0] let y = [5.0, 5.0, 5.0, 5.0, 5.0] - let result = @src.rstatix_t_test(x, y=y) + let result = @src.rstatix_t_test(x, y~) assert_true(result.p_value.is_nan()) } +///| test "rstatix_t_test paired basic" { let x = [10.0, 12.0, 15.0, 18.0, 20.0, 22.0, 25.0, 28.0, 30.0, 32.0] let y = [8.0, 9.0, 11.0, 12.0, 14.0, 15.0, 17.0, 19.0, 20.0, 22.0] - let result = @src.rstatix_t_test(x, y=y, paired=true) + let result = @src.rstatix_t_test(x, y~, paired=true) assert_eq(result.method_name, "Paired t-test") assert_true(result.statistic > 0.0) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) assert_true(result.ci_low < result.ci_high) } +///| test "rstatix_t_test paired different length" { let x = [10.0, 12.0, 15.0, 18.0, 20.0, 22.0, 25.0] let y = [8.0, 9.0, 11.0, 12.0, 14.0] - let result = @src.rstatix_t_test(x, y=y, paired=true) + let result = @src.rstatix_t_test(x, y~, paired=true) assert_true(result.n1 == 5) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_t_test confidence interval" { let x = [10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0] let result99 = @src.rstatix_t_test(x, mu=15.0, conf_level=0.99) @@ -86,6 +95,7 @@ test "rstatix_t_test confidence interval" { // ===== rstatix_wilcox_test ===== +///| test "rstatix_wilcox_test one-sample basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let result = @src.rstatix_wilcox_test(x, mu=5.0) @@ -95,6 +105,7 @@ test "rstatix_wilcox_test one-sample basic" { assert_true(result.n1 == 9) } +///| test "rstatix_wilcox_test one-sample too few observations" { let x = [1.0, 2.0, 3.0] let result = @src.rstatix_wilcox_test(x, mu=2.0) @@ -103,6 +114,7 @@ test "rstatix_wilcox_test one-sample too few observations" { assert_true(result.n1 == 2) } +///| test "rstatix_wilcox_test one-sample greater alternative" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let result = @src.rstatix_wilcox_test(x, mu=4.0, alternative="greater") @@ -110,10 +122,11 @@ test "rstatix_wilcox_test one-sample greater alternative" { assert_eq(result.alternative, "greater") } +///| test "rstatix_wilcox_test two-sample rank-sum basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0] let y = [8.0, 9.0, 10.0, 11.0, 12.0, 13.0] - let result = @src.rstatix_wilcox_test(x, y=y) + let result = @src.rstatix_wilcox_test(x, y~) assert_eq(result.test_name, "wilcox_test") assert_eq(result.method_name, "Wilcoxon rank-sum test") assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) @@ -122,40 +135,45 @@ test "rstatix_wilcox_test two-sample rank-sum basic" { assert_true(result.n2 == 6) } +///| test "rstatix_wilcox_test two-sample too few observations" { let x = [1.0, 2.0] let y = [3.0, 4.0] - let result = @src.rstatix_wilcox_test(x, y=y) + let result = @src.rstatix_wilcox_test(x, y~) assert_true(result.p_value.is_nan()) assert_eq(result.n1, 2) assert_eq(result.n2, 2) } +///| test "rstatix_wilcox_test two-sample one group too small" { let x = [1.0, 2.0, 3.0] let y = [4.0, 5.0] - let result = @src.rstatix_wilcox_test(x, y=y) + let result = @src.rstatix_wilcox_test(x, y~) assert_true(result.p_value.is_nan()) } +///| test "rstatix_wilcox_test paired basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0] - let result = @src.rstatix_wilcox_test(x, y=y, paired=true) + let result = @src.rstatix_wilcox_test(x, y~, paired=true) assert_eq(result.method_name, "Paired Wilcoxon") assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) assert_true(!result.statistic.is_nan()) } +///| test "rstatix_wilcox_test paired with mu" { let x = [5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0] let y = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] - let result = @src.rstatix_wilcox_test(x, y=y, paired=true, mu=2.0) + let result = @src.rstatix_wilcox_test(x, y~, paired=true, mu=2.0) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } // ===== rstatix_cor_test ===== +///| test "rstatix_cor_test pearson basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -175,6 +193,7 @@ test "rstatix_cor_test pearson basic" { assert_true(n_val == 10) } +///| test "rstatix_cor_test pearson negative" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [20.0, 18.0, 16.0, 14.0, 12.0, 10.0, 8.0, 6.0, 4.0, 2.0] @@ -183,6 +202,7 @@ test "rstatix_cor_test pearson negative" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test pearson no correlation" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0] @@ -191,6 +211,7 @@ test "rstatix_cor_test pearson no correlation" { assert_true(result.p_value.is_nan()) } +///| test "rstatix_cor_test spearman basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -201,6 +222,7 @@ test "rstatix_cor_test spearman basic" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test spearman nonlinear" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [1.0, 4.0, 9.0, 16.0, 25.0, 36.0, 49.0, 64.0, 81.0, 100.0] @@ -209,6 +231,7 @@ test "rstatix_cor_test spearman nonlinear" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test kendall basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -219,6 +242,7 @@ test "rstatix_cor_test kendall basic" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test kendall negative" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [20.0, 18.0, 16.0, 14.0, 12.0, 10.0, 8.0, 6.0, 4.0, 2.0] @@ -227,6 +251,7 @@ test "rstatix_cor_test kendall negative" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test mismatched length" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [1.0, 2.0, 3.0, 4.0] @@ -235,6 +260,7 @@ test "rstatix_cor_test mismatched length" { assert_true(result.p_value.is_nan()) } +///| test "rstatix_cor_test too small" { let x = [1.0, 2.0] let y = [3.0, 4.0] @@ -245,6 +271,7 @@ test "rstatix_cor_test too small" { // ===== rstatix_anova_test ===== +///| test "rstatix_anova_test basic" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -260,6 +287,7 @@ test "rstatix_anova_test basic" { assert_true(result.ms > 0.0) } +///| test "rstatix_anova_test equal groups" { let groups = [ [5.0, 5.0, 5.0, 5.0, 5.0], @@ -271,24 +299,22 @@ test "rstatix_anova_test equal groups" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_anova_test too few groups" { - let groups = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let groups = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let result = @src.rstatix_anova_test(groups) assert_true(result.p_value.is_nan()) assert_true(result.f.is_nan()) } +///| test "rstatix_anova_test two groups" { - let groups = [ - [1.0, 2.0, 3.0], - ] + let groups = [[1.0, 2.0, 3.0]] let result = @src.rstatix_anova_test(groups) assert_true(result.p_value.is_nan()) } +///| test "rstatix_anova_test four groups" { let groups = [ [1.0, 2.0, 3.0, 4.0, 5.0], @@ -303,6 +329,7 @@ test "rstatix_anova_test four groups" { // ===== rstatix_kruskal_test ===== +///| test "rstatix_kruskal_test basic" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -317,6 +344,7 @@ test "rstatix_kruskal_test basic" { assert_true(result.ss > 0.0) } +///| test "rstatix_kruskal_test equal groups" { let groups = [ [5.0, 5.0, 5.0, 5.0, 5.0], @@ -327,11 +355,9 @@ test "rstatix_kruskal_test equal groups" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_kruskal_test too few groups" { - let groups = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let groups = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let result = @src.rstatix_kruskal_test(groups) assert_true(result.p_value.is_nan()) assert_true(result.f.is_nan()) @@ -339,6 +365,7 @@ test "rstatix_kruskal_test too few groups" { // ===== rstatix_friedman_test ===== +///| test "rstatix_friedman_test basic" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -353,6 +380,7 @@ test "rstatix_friedman_test basic" { assert_true(result.ss > 0.0) } +///| test "rstatix_friedman_test equal groups" { let groups = [ [5.0, 5.0, 5.0, 5.0, 5.0], @@ -363,28 +391,24 @@ test "rstatix_friedman_test equal groups" { assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_friedman_test too few groups" { - let groups = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - ] + let groups = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] let result = @src.rstatix_friedman_test(groups) assert_true(result.p_value.is_nan()) assert_true(result.f.is_nan()) } +///| test "rstatix_friedman_test small blocks" { - let groups = [ - [1.0, 4.0], - [2.0, 5.0], - [3.0, 6.0], - ] + let groups = [[1.0, 4.0], [2.0, 5.0], [3.0, 6.0]] let result = @src.rstatix_friedman_test(groups) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } // ===== rstatix_bh_correct ===== +///| test "rstatix_bh_correct basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.rstatix_bh_correct(p_values) @@ -397,12 +421,14 @@ test "rstatix_bh_correct basic" { } } +///| test "rstatix_bh_correct empty" { let p_values : Array[Double] = Array::new() let result = @src.rstatix_bh_correct(p_values) assert_true(result.length() == 0) } +///| test "rstatix_bh_correct single" { let p_values = [0.05] let result = @src.rstatix_bh_correct(p_values) @@ -410,6 +436,7 @@ test "rstatix_bh_correct single" { assert_true(result[0] >= 0.05 - 1.0e-10 && result[0] <= 0.05 + 1.0e-10) } +///| test "rstatix_bh_correct all significant" { let p_values = [0.001, 0.002, 0.003, 0.004, 0.005] let result = @src.rstatix_bh_correct(p_values) @@ -421,6 +448,7 @@ test "rstatix_bh_correct all significant" { } } +///| test "rstatix_bh_correct all non-significant" { let p_values = [0.8, 0.9, 0.95, 0.98, 0.99] let result = @src.rstatix_bh_correct(p_values) @@ -433,6 +461,7 @@ test "rstatix_bh_correct all non-significant" { } } +///| test "rstatix_bh_correct preserves order" { let p_values = [0.2, 0.005, 0.04, 0.01, 0.03] let result = @src.rstatix_bh_correct(p_values) @@ -443,6 +472,7 @@ test "rstatix_bh_correct preserves order" { // ===== rstatix_bonferroni ===== +///| test "rstatix_bonferroni basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.rstatix_bonferroni(p_values) @@ -452,12 +482,14 @@ test "rstatix_bonferroni basic" { assert_true(result[4] >= 1.0 - 1.0e-10) } +///| test "rstatix_bonferroni empty" { let p_values : Array[Double] = Array::new() let result = @src.rstatix_bonferroni(p_values) assert_true(result.length() == 0) } +///| test "rstatix_bonferroni single" { let p_values = [0.05] let result = @src.rstatix_bonferroni(p_values) @@ -465,6 +497,7 @@ test "rstatix_bonferroni single" { assert_true(result[0] >= 0.05 - 1.0e-10 && result[0] <= 0.05 + 1.0e-10) } +///| test "rstatix_bonferroni capped at 1" { let p_values = [0.3, 0.4, 0.5] let result = @src.rstatix_bonferroni(p_values) @@ -475,6 +508,7 @@ test "rstatix_bonferroni capped at 1" { // ===== Summary functions ===== +///| test "rstatix_t_test_summary basic" { let x = [10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0] let result = @src.rstatix_t_test(x, mu=15.0) @@ -487,15 +521,17 @@ test "rstatix_t_test_summary basic" { assert_true(summary.contains("Method")) } +///| test "rstatix_t_test_summary two-sample" { let x = [10.0, 12.0, 15.0, 18.0, 20.0] let y = [8.0, 9.0, 11.0, 12.0, 14.0] - let result = @src.rstatix_t_test(x, y=y) + let result = @src.rstatix_t_test(x, y~) let summary = @src.rstatix_t_test_summary(result) assert_true(summary.contains("Group1")) assert_true(summary.contains("Group2")) } +///| test "rstatix_cor_test_summary basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -509,6 +545,7 @@ test "rstatix_cor_test_summary basic" { assert_true(summary.contains("Method")) } +///| test "rstatix_cor_test_summary spearman" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -517,6 +554,7 @@ test "rstatix_cor_test_summary spearman" { assert_true(summary.contains("spearman")) } +///| test "rstatix_anova_summary basic" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -535,6 +573,7 @@ test "rstatix_anova_summary basic" { assert_true(summary.contains("p-value")) } +///| test "rstatix_anova_summary kruskal" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -546,6 +585,7 @@ test "rstatix_anova_summary kruskal" { assert_true(summary.contains("Kruskal-Wallis")) } +///| test "rstatix_anova_summary friedman" { let groups = [ [10.0, 12.0, 15.0, 18.0, 20.0], @@ -559,6 +599,7 @@ test "rstatix_anova_summary friedman" { // ===== rstatix_sample_data ===== +///| test "rstatix_sample_data basic" { let data = @src.rstatix_sample_data() assert_true(data.length() == 3) @@ -567,6 +608,7 @@ test "rstatix_sample_data basic" { assert_true(data[2].length() == 10) } +///| test "rstatix_sample_data values" { let data = @src.rstatix_sample_data() assert_true(data[0][0] == 10.0) @@ -577,6 +619,7 @@ test "rstatix_sample_data values" { assert_true(data[2][9] == 14.0) } +///| test "rstatix_sample_data used in anova" { let data = @src.rstatix_sample_data() let result = @src.rstatix_anova_test(data) @@ -584,12 +627,14 @@ test "rstatix_sample_data used in anova" { assert_true(result.f > 0.0) } +///| test "rstatix_sample_data used in kruskal" { let data = @src.rstatix_sample_data() let result = @src.rstatix_kruskal_test(data) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_sample_data used in friedman" { let data = @src.rstatix_sample_data() let result = @src.rstatix_friedman_test(data) @@ -598,6 +643,7 @@ test "rstatix_sample_data used in friedman" { // ===== Edge cases and integration ===== +///| test "rstatix_t_test constant values" { let x = [5.0, 5.0, 5.0, 5.0, 5.0] let result = @src.rstatix_t_test(x, mu=5.0) @@ -605,18 +651,21 @@ test "rstatix_t_test constant values" { assert_true(result.p_value.is_nan()) } +///| test "rstatix_wilcox_test all zeros" { let x = [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0] let result = @src.rstatix_wilcox_test(x) assert_true(result.p_value.is_nan()) } +///| test "rstatix_cor_test identical values" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let result = @src.rstatix_cor_test(x, x) assert_true(result.statistic > 0.99 && result.statistic <= 1.0) } +///| test "rstatix_bh_correct preserves non-decreasing order" { let p_values = [0.001, 0.01, 0.03, 0.05, 0.1] let result = @src.rstatix_bh_correct(p_values) @@ -627,34 +676,41 @@ test "rstatix_bh_correct preserves non-decreasing order" { } } +///| test "rstatix_t_test_two_sample_same_values" { let x = [10.0, 10.0, 10.0, 10.0, 10.0] let y = [10.0, 10.0, 10.0, 10.0, 10.0] - let result = @src.rstatix_t_test(x, y=y) + let result = @src.rstatix_t_test(x, y~) assert_true(result.statistic.is_nan()) } +///| test "rstatix_anova_test with sample data from function" { let data = @src.rstatix_sample_data() - let result = @src.rstatix_anova_test(data, group_names=["GroupA", "GroupB", "GroupC"]) + let result = @src.rstatix_anova_test(data, group_names=[ + "GroupA", "GroupB", "GroupC", + ]) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_t_test paired with mu default" { let x = [10.0, 12.0, 15.0, 18.0, 20.0, 22.0, 25.0, 28.0, 30.0, 32.0] let y = [8.0, 9.0, 11.0, 12.0, 14.0, 15.0, 17.0, 19.0, 20.0, 22.0] - let result = @src.rstatix_t_test(x, y=y, paired=true) + let result = @src.rstatix_t_test(x, y~, paired=true) assert_true(result.statistic > 0.0) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_wilcox_test two-sample with different lengths" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] let y = [9.0, 10.0, 11.0, 12.0, 13.0, 14.0] - let result = @src.rstatix_wilcox_test(x, y=y) + let result = @src.rstatix_wilcox_test(x, y~) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) } +///| test "rstatix_cor_test pearson default method" { let x = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0] @@ -662,6 +718,7 @@ test "rstatix_cor_test pearson default method" { assert_eq(result.method_name, "Pearson's product-moment correlation") } +///| test "rstatix_kruskal_test with sample data" { let data = @src.rstatix_sample_data() let result = @src.rstatix_kruskal_test(data, _group_names=["A", "B", "C"]) @@ -669,9 +726,10 @@ test "rstatix_kruskal_test with sample data" { assert_true(result.source == "Kruskal-Wallis") } +///| test "rstatix_friedman_test with sample data" { let data = @src.rstatix_sample_data() let result = @src.rstatix_friedman_test(data, _group_names=["A", "B", "C"]) assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) assert_true(result.source == "Friedman") -} \ No newline at end of file +} diff --git a/test/moonbit/rtsne_test.mbt b/test/moonbit/rtsne_test.mbt index 2fb8c168..b52479aa 100644 --- a/test/moonbit/rtsne_test.mbt +++ b/test/moonbit/rtsne_test.mbt @@ -6,10 +6,7 @@ // ============================================================ test "calc_distance_matrix square" { - let data : Array[Array[Double]] = [ - [0.0, 0.0], - [3.0, 4.0], - ] + let data : Array[Array[Double]] = [[0.0, 0.0], [3.0, 4.0]] let dist = @src.calc_distance_matrix(data) assert_eq(dist.length(), 2) assert_eq(dist[0].length(), 2) @@ -20,12 +17,14 @@ test "calc_distance_matrix square" { assert_eq(dist[1][0], 5.0) } +///| test "calc_distance_matrix empty" { let data : Array[Array[Double]] = [] let dist = @src.calc_distance_matrix(data) assert_eq(dist.length(), 0) } +///| test "calc_distance_matrix single point" { let data : Array[Array[Double]] = [[1.0, 2.0, 3.0]] let dist = @src.calc_distance_matrix(data) @@ -37,6 +36,7 @@ test "calc_distance_matrix single point" { // t-SNE config tests // ============================================================ +///| test "tsne_config default" { let config = @src.TsneConfig::new() assert_eq(config.perplexity, 30.0) @@ -50,13 +50,14 @@ test "tsne_config default" { // t-SNE algorithm tests // ============================================================ +///| test "tsne basic 2d" { let data = @src.create_tsne_test_data(10, 5) assert_eq(data.length(), 10) assert_eq(data[0].length(), 5) - + let test_config = @src.TsneConfig::new_custom(3.0, 100, 2, 42) - + let result = @src.tsne(data, test_config) assert_eq(result.embedding.length(), 10) assert_eq(result.embedding[0].length(), 2) @@ -64,6 +65,7 @@ test "tsne basic 2d" { assert_eq(result.costs.length(), 100) } +///| test "tsne empty data" { let data : Array[Array[Double]] = [] let config = @src.TsneConfig::new() @@ -71,6 +73,7 @@ test "tsne empty data" { assert_eq(result.embedding.length(), 0) } +///| test "tsne single sample" { let data : Array[Array[Double]] = [[1.0, 2.0, 3.0]] let config = @src.TsneConfig::new() @@ -79,12 +82,13 @@ test "tsne single sample" { assert_true(result.embedding.length() <= 1) } +///| test "tsne costs decrease" { let data = @src.create_tsne_test_data(15, 4) let config = @src.TsneConfig::new_custom(3.0, 50, 2, 42) - + let result = @src.tsne(data, config) - + // Costs should be finite if result.costs.length() > 10 { // Check that costs are not NaN or infinite @@ -95,10 +99,11 @@ test "tsne costs decrease" { } } +///| test "tsne different dims" { let data = @src.create_tsne_test_data(10, 5) let config = @src.TsneConfig::new_custom(3.0, 50, 3, 42) - + let result = @src.tsne(data, config) assert_eq(result.embedding[0].length(), 3) } @@ -107,19 +112,21 @@ test "tsne different dims" { // Test data generation tests // ============================================================ +///| test "create_tsne_test_data dimensions" { let data = @src.create_tsne_test_data(20, 8) assert_eq(data.length(), 20) assert_eq(data[0].length(), 8) } +///| test "create_tsne_test_data clusters" { let data = @src.create_tsne_test_data(9, 3) // 3 clusters, each should have similar values // Cluster 0: samples 0, 3, 6 // Cluster 1: samples 1, 4, 7 // Cluster 2: samples 2, 5, 8 - + // Check that cluster 0 is different from cluster 1 let mut sum0 = 0.0 let mut j = 0 @@ -127,14 +134,14 @@ test "create_tsne_test_data clusters" { sum0 = sum0 + data[0][j] j = j + 1 } - + let mut sum1 = 0.0 j = 0 while j < 3 { sum1 = sum1 + data[1][j] j = j + 1 } - + // Clusters should be separated assert_true(sum1 > sum0) } diff --git a/test/moonbit/s4vectors_test.mbt b/test/moonbit/s4vectors_test.mbt index 3ccb1769..50921427 100644 --- a/test/moonbit/s4vectors_test.mbt +++ b/test/moonbit/s4vectors_test.mbt @@ -2,7 +2,7 @@ test "s4vectors_rle_from_vector" { let vec = ["A", "A", "A", "B", "B", "C", "C", "C", "C", "A", "A"] let rle = @src.Rle::from_vector(vec) - + assert_eq(rle.values.length(), 4) assert_eq(rle.lengths.length(), 4) } @@ -11,7 +11,7 @@ test "s4vectors_rle_from_vector" { test "s4vectors_rle_length" { let vec = ["A", "A", "B", "B", "B"] let rle = @src.Rle::from_vector(vec) - + assert_eq(rle.length(), 5) } @@ -19,7 +19,7 @@ test "s4vectors_rle_length" { test "s4vectors_rle_get" { let vec = ["A", "A", "B", "C", "C", "C"] let rle = @src.Rle::from_vector(vec) - + assert_eq(rle.get(0), "A") assert_eq(rle.get(2), "B") assert_eq(rle.get(5), "C") @@ -29,9 +29,9 @@ test "s4vectors_rle_get" { test "s4vectors_dataframe_basic" { let col1 = @src.S4DataFrameColumn::new("gene_id", ["gene1", "gene2", "gene3"]) let col2 = @src.S4DataFrameColumn::new("expression", ["10.5", "25.3", "5.8"]) - + let df = @src.S4DataFrame::new([col1, col2], ["row1", "row2", "row3"]) - + assert_eq(df.nrow(), 3) assert_eq(df.ncol(), 2) } @@ -40,25 +40,25 @@ test "s4vectors_dataframe_basic" { test "s4vectors_dataframe_colnames" { let col1 = @src.S4DataFrameColumn::new("gene_id", ["gene1", "gene2"]) let col2 = @src.S4DataFrameColumn::new("expression", ["10.5", "25.3"]) - + let df = @src.S4DataFrame::new([col1, col2], []) let names = df.colnames() - + assert_eq(names.length(), 2) } ///| test "s4vectors_hits_basic" { let hits = @src.Hits::new([0, 0, 1, 2], [1, 2, 0, 1], 3, 3) - + assert_eq(hits.n_hits(), 4) } ///| test "s4vectors_hits_count_query_hits" { let hits = @src.Hits::new([0, 0, 1, 2], [1, 2, 0, 1], 3, 3) - + let counts = hits.count_query_hits() - + assert_eq(counts.length(), 3) -} \ No newline at end of file +} diff --git a/test/moonbit/sasa_test.mbt b/test/moonbit/sasa_test.mbt index 8149442b..61259184 100644 --- a/test/moonbit/sasa_test.mbt +++ b/test/moonbit/sasa_test.mbt @@ -13,19 +13,19 @@ fn make_atom( resseq : Int, ) -> @src.Atom { @src.Atom::new( - name=name, + name~, coord=@src.Vector3::new(x, y, z), - resname=resname, + resname~, chainid='A', - resseq=resseq, - element=element, + resseq~, + element~, ) } ///| /// Helper: place all given atoms into a single residue on chain A. fn make_structure(atoms : Array[@src.Atom]) -> @src.Structure { - let res = @src.Residue::new(resname="UNK", chainid='A', resseq=1, atoms=atoms) + let res = @src.Residue::new(resname="UNK", chainid='A', resseq=1, atoms~) let chain = @src.Chain::new(id='A', residues=[res]) let model = @src.Model::new(id=1, chains=[chain]) @src.Structure::new(id="test", models=[model]) @@ -240,8 +240,9 @@ test "sasa_total_backbone_sidechain" { let atom_res = @src.sasa_calc(s, 100, 1.4) let summary = @src.sasa_calc_total(atom_res) assert_true( - (summary.get_backbone_sasa() + summary.get_sidechain_sasa() - - summary.get_total_sasa()).abs() < + (summary.get_backbone_sasa() + + summary.get_sidechain_sasa() - + summary.get_total_sasa()).abs() < 1.0e-9, ) assert_true(summary.get_backbone_sasa() > 0.0) @@ -310,7 +311,10 @@ test "sasa_all_backbone_atoms_sidechain_zero" { let res_res = @src.sasa_calc_residue(atom_res, s) assert_eq(res_res.length(), 1) assert_true(res_res[0].get_sidechain_sasa() == 0.0) - assert_true((res_res[0].get_backbone_sasa() - res_res[0].get_total_sasa()).abs() < 1.0e-9) + assert_true( + (res_res[0].get_backbone_sasa() - res_res[0].get_total_sasa()).abs() < + 1.0e-9, + ) } ///| diff --git a/test/moonbit/sc3_test.mbt b/test/moonbit/sc3_test.mbt index 54428771..d3376d35 100644 --- a/test/moonbit/sc3_test.mbt +++ b/test/moonbit/sc3_test.mbt @@ -3,14 +3,10 @@ ///| test "sc3_preprocess" { - let data = [ - [0.0, 1.0, 10.0], - [0.0, 2.0, 20.0], - [0.0, 3.0, 30.0], - ] - + let data = [[0.0, 1.0, 10.0], [0.0, 2.0, 20.0], [0.0, 3.0, 30.0]] + let processed = @src.sc3_preprocess(data) - + assert_true(processed.length() == 3) assert_true(processed[0].length() == 3) } @@ -23,53 +19,38 @@ test "sc3_pca" { [3.0, 4.0, 5.0], [10.0, 11.0, 12.0], ] - + let pcs = @src.sc3_pca(data, 2) - + assert_true(pcs.length() == 4) assert_true(pcs[0].length() == 2) } ///| test "sc3_kmeans" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - [11.0, 12.0], - ] - + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0], [11.0, 12.0]] + let labels = @src.sc3_kmeans(data, 2) - + assert_true(labels.length() == 4) } ///| test "sc3_silhouette" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - [11.0, 12.0], - ] + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0], [11.0, 12.0]] let labels = [0, 0, 1, 1] - + let scores = @src.sc3_calculate_silhouette(data, labels, 2) - + assert_true(scores.length() == 4) } ///| test "sc3_gap_statistics" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - [11.0, 12.0], - ] - + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0], [11.0, 12.0]] + let gaps = @src.sc3_calculate_gap_statistics(data, 3) - + assert_true(gaps.length() == 3) } @@ -81,9 +62,9 @@ test "sc3_consensus_cluster" { [10.0, 11.0, 12.0], [11.0, 12.0, 13.0], ] - + let result = @src.bio_sc3_cluster(data, 2) - + assert_true(result.k == 2) assert_true(result.cluster_labels.length() == 4) assert_true(result.consensus_matrix.length() == 4) @@ -91,16 +72,12 @@ test "sc3_consensus_cluster" { ///| test "sc3_bio_api" { - let data = [ - [1.0, 2.0], - [2.0, 3.0], - [10.0, 11.0], - ] - + let data = [[1.0, 2.0], [2.0, 3.0], [10.0, 11.0]] + let pcs = @src.bio_sc3_pca(data, 2) let labels = @src.sc3_kmeans(pcs, 2) let scores = @src.bio_sc3_silhouette(pcs, labels, 2) - + assert_true(pcs.length() == 3) assert_true(scores.length() == 3) -} \ No newline at end of file +} diff --git a/test/moonbit/sc_dbl_finder_advanced_test.mbt b/test/moonbit/sc_dbl_finder_advanced_test.mbt new file mode 100644 index 00000000..37b4c6e3 --- /dev/null +++ b/test/moonbit/sc_dbl_finder_advanced_test.mbt @@ -0,0 +1,1223 @@ +///| +fn scdfa_test_small_data() -> @src.SingleCellData { + @src.SingleCellData::create( + [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0], [2.0, 4.0, 6.0]], + ["c1", "c2", "c3", "c4"], + ["g1", "g2", "g3"], + ) catch { + _ => abort("valid small scDblFinder data should build") + } +} + +///| +fn scdfa_test_config() -> @src.ScDblFinderConfig { + @src.ScDblFinderConfig::create( + expected_doublet_rate=0.15, + artificial_doublets=48, + selected_features=18, + dimensions=5, + neighbors=8, + iterations=2, + classifier_steps=120, + seed=17, + ) catch { + _ => abort("valid scDblFinder configuration should build") + } +} + +///| +fn scdfa_test_result() -> @src.ScDblFinderResult { + let (data, clusters, samples, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("valid scDblFinder example should build") + } + @src.sc_dbl_finder( + data, + clusters~, + samples~, + known_doublets=truth, + config=scdfa_test_config(), + ) catch { + _ => abort("valid scDblFinder workflow should fit") + } +} + +///| +fn scdfa_test_experiment() -> @src.SingleCellExperiment { + let (data, clusters, samples, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("valid scDblFinder example should build") + } + let assay : Array[Array[Double]] = [] + for gene in 0.. abort("advanced example should build") + } + assert_eq(data.counts.length(), 42) + assert_eq(data.gene_names.length(), 24) + assert_eq(clusters.length(), 42) + assert_eq(samples.length(), 42) + assert_eq(truth.length(), 42) + let mut doublets = 0 + for value in truth { + if value { + doublets = doublets + 1 + } + } + assert_eq(doublets, 6) +} + +///| +test "scDblFinder SingleCellData constructor copies inputs" { + let counts = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]] + let cells = ["a", "b", "c"] + let genes = ["x", "y"] + let data = @src.SingleCellData::create(counts, cells, genes) catch { + _ => abort("valid data should build") + } + counts[0][0] = 99.0 + cells[0] = "changed" + genes[0] = "changed" + assert_eq(data.counts[0][0], 1.0) + assert_eq(data.cell_names[0], "a") + assert_eq(data.gene_names[0], "x") +} + +///| +test "scDblFinder rejects fewer than three cells" { + let raised = try { + ignore( + @src.SingleCellData::create([[1.0, 2.0], [3.0, 4.0]], ["a", "b"], [ + "x", "y", + ]), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder rejects fewer than two genes" { + let raised = try { + ignore( + @src.SingleCellData::create([[1.0], [2.0], [3.0]], ["a", "b", "c"], ["x"]), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder rejects a ragged count matrix" { + let raised = try { + ignore( + @src.SingleCellData::create( + [[1.0, 2.0], [3.0], [4.0, 5.0]], + ["a", "b", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder validates cell and gene identifiers" { + let wrong_length = try { + ignore( + @src.SingleCellData::create( + [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]], + ["a", "b"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + let empty_name = try { + ignore( + @src.SingleCellData::create( + [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]], + ["a", "", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + let duplicate_gene = try { + ignore( + @src.SingleCellData::create( + [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]], + ["a", "b", "c"], + ["x", "x"], + ), + ) + false + } catch { + _ => true + } + assert_true(wrong_length && empty_name && duplicate_gene) +} + +///| +test "scDblFinder rejects non-finite and negative counts" { + let negative = try { + ignore( + @src.SingleCellData::create( + [[1.0, -1.0], [3.0, 4.0], [5.0, 6.0]], + ["a", "b", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + let nan = try { + ignore( + @src.SingleCellData::create( + [[1.0, Double::nan()], [3.0, 4.0], [5.0, 6.0]], + ["a", "b", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + let infinite = try { + ignore( + @src.SingleCellData::create( + [[1.0, 1.0e301], [3.0, 4.0], [5.0, 6.0]], + ["a", "b", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + assert_true(negative && nan && infinite) +} + +///| +test "scDblFinder rejects zero-library cells" { + let raised = try { + ignore( + @src.SingleCellData::create( + [[0.0, 0.0], [3.0, 4.0], [5.0, 6.0]], + ["a", "b", "c"], + ["x", "y"], + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder default configuration matches portable workflow" { + let config = @src.ScDblFinderConfig::default() + assert_eq(config.expected_doublet_rate, -1.0) + assert_eq(config.doublet_rate_per_1000, 0.008) + assert_eq(config.selected_features, 1000) + assert_eq(config.dimensions, 20) + assert_eq(config.neighbors, 15) + assert_eq(config.iterations, 3) + assert_eq(config.classifier_steps, 300) + assert_eq(config.cluster_count, 0) +} + +///| +test "scDblFinder configuration preserves custom values" { + let config = @src.ScDblFinderConfig::create( + expected_doublet_rate=0.12, + rate_uncertainty=0.03, + artificial_doublets=40, + selected_features=12, + dimensions=4, + neighbors=7, + iterations=2, + classifier_steps=50, + random_fraction=0.2, + half_size_fraction=0.4, + cluster_count=3, + seed=9, + ) catch { + _ => abort("valid custom configuration should build") + } + assert_eq(config.expected_doublet_rate, 0.12) + assert_eq(config.artificial_doublets, 40) + assert_eq(config.dimensions, 4) + assert_eq(config.cluster_count, 3) + assert_eq(config.seed, 9) +} + +///| +test "scDblFinder configuration rejects invalid rates" { + let expected = try { + ignore(@src.ScDblFinderConfig::create(expected_doublet_rate=1.1)) + false + } catch { + _ => true + } + let per_thousand = try { + ignore(@src.ScDblFinderConfig::create(doublet_rate_per_1000=-0.1)) + false + } catch { + _ => true + } + let uncertainty = try { + ignore(@src.ScDblFinderConfig::create(rate_uncertainty=-0.2)) + false + } catch { + _ => true + } + assert_true(expected && per_thousand && uncertainty) +} + +///| +test "scDblFinder configuration rejects invalid dimensions" { + let features = try { + ignore(@src.ScDblFinderConfig::create(selected_features=1)) + false + } catch { + _ => true + } + let dimensions = try { + ignore(@src.ScDblFinderConfig::create(dimensions=0)) + false + } catch { + _ => true + } + let neighbors = try { + ignore(@src.ScDblFinderConfig::create(neighbors=0)) + false + } catch { + _ => true + } + assert_true(features && dimensions && neighbors) +} + +///| +test "scDblFinder configuration validates classifier controls" { + let iterations = try { + ignore(@src.ScDblFinderConfig::create(iterations=0)) + false + } catch { + _ => true + } + let steps = try { + ignore(@src.ScDblFinderConfig::create(classifier_steps=0)) + false + } catch { + _ => true + } + let learning_rate = try { + ignore(@src.ScDblFinderConfig::create(learning_rate=1.1)) + false + } catch { + _ => true + } + let regularization = try { + ignore(@src.ScDblFinderConfig::create(regularization=-0.01)) + false + } catch { + _ => true + } + assert_true(iterations && steps && learning_rate && regularization) +} + +///| +test "scDblFinder configuration validates threshold and mixing fractions" { + let stringency = try { + ignore(@src.ScDblFinderConfig::create(stringency=1.0)) + false + } catch { + _ => true + } + let random_fraction = try { + ignore(@src.ScDblFinderConfig::create(random_fraction=-0.1)) + false + } catch { + _ => true + } + let half_fraction = try { + ignore(@src.ScDblFinderConfig::create(half_size_fraction=1.1)) + false + } catch { + _ => true + } + let unidentifiable = try { + ignore(@src.ScDblFinderConfig::create(unidentifiable_threshold=1.1)) + false + } catch { + _ => true + } + assert_true(stringency && random_fraction && half_fraction && unidentifiable) +} + +///| +test "scDblFinder configuration rejects a single requested cluster" { + let raised = try { + ignore(@src.ScDblFinderConfig::create(cluster_count=1)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder expected rate follows cells per thousand" { + let rate = @src.scdf_expected_doublet_rate(5000) catch { + _ => abort("valid expected rate should compute") + } + assert_true((rate - 0.04).abs() < 1.0e-12) +} + +///| +test "scDblFinder expected rate is capped and validated" { + let capped = @src.scdf_expected_doublet_rate(2000, rate_per_1000=1.0) catch { + _ => abort("valid capped rate should compute") + } + assert_eq(capped, 1.0) + let raised = try { + ignore(@src.scdf_expected_doublet_rate(0)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder homotypic proportion uses squared cluster frequencies" { + let proportion = @src.scdf_homotypic_proportion(["A", "A", "B", "B"]) catch { + _ => abort("valid homotypic proportion should compute") + } + assert_true((proportion - 0.5).abs() < 1.0e-12) +} + +///| +test "scDblFinder homotypic proportion validates labels" { + let empty = try { + ignore(@src.scdf_homotypic_proportion([])) + false + } catch { + _ => true + } + let blank = try { + ignore(@src.scdf_homotypic_proportion(["A", ""])) + false + } catch { + _ => true + } + assert_true(empty && blank) +} + +///| +test "scDblFinder random artificial doublets have requested dimensions" { + let artificial = @src.scdf_generate_artificial_doublets( + scdfa_test_small_data(), + 7, + random_fraction=1.0, + half_size_fraction=0.0, + seed=3, + ) catch { + _ => abort("valid artificial doublets should build") + } + assert_eq(artificial.counts.length(), 7) + assert_eq(artificial.names.length(), 7) + assert_eq(artificial.parent_one.length(), 7) + assert_eq(artificial.origins, ["", "", "", "", "", "", ""]) + for row in artificial.counts { + assert_eq(row.length(), 3) + } +} + +///| +test "scDblFinder artificial counts sum parent profiles" { + let data = scdfa_test_small_data() + let artificial = @src.scdf_generate_artificial_doublets( + data, + 5, + half_size_fraction=0.0, + seed=4, + ) catch { + _ => abort("valid artificial doublets should build") + } + for index in 0.. abort("valid half-size artificial doublets should build") + } + for index in 0..<4 { + let expected_scale = if index < 2 { 0.5 } else { 1.0 } + let expected = ( + data.counts[artificial.parent_one[index]][0] + + data.counts[artificial.parent_two[index]][0] + ) * + expected_scale + assert_eq(artificial.counts[index][0], expected) + } +} + +///| +test "scDblFinder artificial parents are distinct and in bounds" { + let artificial = @src.scdf_generate_artificial_doublets( + scdfa_test_small_data(), + 20, + seed=8, + ) catch { + _ => abort("valid artificial doublets should build") + } + for index in 0..= 0) + assert_true(artificial.parent_two[index] < 4) + assert_true(artificial.parent_one[index] != artificial.parent_two[index]) + } +} + +///| +test "scDblFinder cluster-aware artificial doublets use cross-cluster pairs" { + let data = scdfa_test_small_data() + let clusters = ["A", "A", "B", "B"] + let artificial = @src.scdf_generate_artificial_doublets( + data, + 12, + clusters~, + random_fraction=0.0, + half_size_fraction=0.0, + ) catch { + _ => abort("valid cluster-aware artificial doublets should build") + } + for index in 0.. abort("valid artificial doublets should build") + } + let second = @src.scdf_generate_artificial_doublets(data, 10, seed=5) catch { + _ => abort("valid artificial doublets should build") + } + assert_eq(first.counts, second.counts) + assert_eq(first.parent_one, second.parent_one) + assert_eq(first.parent_two, second.parent_two) +} + +///| +test "scDblFinder artificial generation responds to its seed" { + let data = scdfa_test_small_data() + let first = @src.scdf_generate_artificial_doublets(data, 10, seed=1) catch { + _ => abort("valid artificial doublets should build") + } + let second = @src.scdf_generate_artificial_doublets(data, 10, seed=2) catch { + _ => abort("valid artificial doublets should build") + } + assert_true( + first.parent_one != second.parent_one || + first.parent_two != second.parent_two, + ) +} + +///| +test "scDblFinder artificial generation validates arguments" { + let data = scdfa_test_small_data() + let number = try { + ignore(@src.scdf_generate_artificial_doublets(data, 0)) + false + } catch { + _ => true + } + let clusters = try { + ignore(@src.scdf_generate_artificial_doublets(data, 2, clusters=["A", "B"])) + false + } catch { + _ => true + } + let fraction = try { + ignore(@src.scdf_generate_artificial_doublets(data, 2, random_fraction=1.1)) + false + } catch { + _ => true + } + assert_true(number && clusters && fraction) +} + +///| +test "scDblFinder threshold separates ideal score classes" { + let threshold = @src.scdf_optimize_threshold( + [0.05, 0.10, 0.15, 0.20], + [0.80, 0.90, 0.95], + 0.0, + uncertainty=0.0, + ) catch { + _ => abort("valid threshold should optimize") + } + assert_true(threshold.threshold > 0.20) + assert_true(threshold.threshold < 0.80) + assert_eq(threshold.false_positive_rate, 0.0) + assert_eq(threshold.false_negative_rate, 0.0) +} + +///| +test "scDblFinder threshold diagnostics remain bounded with overlap" { + let threshold = @src.scdf_optimize_threshold( + [0.1, 0.3, 0.6, 0.8], + [0.4, 0.5, 0.7, 0.9], + 0.25, + ) catch { + _ => abort("valid threshold should optimize") + } + assert_true(threshold.threshold >= 0.0 && threshold.threshold <= 1.0) + assert_true(threshold.called_rate >= 0.0 && threshold.called_rate <= 1.0) + assert_true( + threshold.false_negative_rate >= 0.0 && threshold.false_negative_rate <= 1.0, + ) + assert_true(threshold.cost >= 0.0) +} + +///| +test "scDblFinder threshold optimization is deterministic" { + let first = @src.scdf_optimize_threshold( + [0.1, 0.2, 0.7], + [0.4, 0.8, 0.9], + 0.2, + ) catch { + _ => abort("valid threshold should optimize") + } + let second = @src.scdf_optimize_threshold( + [0.1, 0.2, 0.7], + [0.4, 0.8, 0.9], + 0.2, + ) catch { + _ => abort("valid threshold should optimize") + } + assert_eq(first.threshold, second.threshold) + assert_eq(first.cost, second.cost) +} + +///| +test "scDblFinder threshold requires both classes" { + let raised = try { + ignore(@src.scdf_optimize_threshold([], [0.8], 0.1)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder threshold rejects invalid scores" { + let nan = try { + ignore(@src.scdf_optimize_threshold([Double::nan()], [0.8], 0.1)) + false + } catch { + _ => true + } + let outside = try { + ignore(@src.scdf_optimize_threshold([0.2], [1.1], 0.1)) + false + } catch { + _ => true + } + assert_true(nan && outside) +} + +///| +test "scDblFinder threshold validates rate and stringency" { + let rate = try { + ignore(@src.scdf_optimize_threshold([0.2], [0.8], 1.1)) + false + } catch { + _ => true + } + let stringency = try { + ignore(@src.scdf_optimize_threshold([0.2], [0.8], 0.1, stringency=0.0)) + false + } catch { + _ => true + } + assert_true(rate && stringency) +} + +///| +test "scDblFinder fast clustering labels every cell" { + let (data, _, _, _) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let clusters = @src.scdf_fast_cluster( + data, + 3, + dimensions=5, + selected_features=18, + ) catch { + _ => abort("valid fast clustering should run") + } + assert_eq(clusters.length(), data.cell_names.length()) + for label in clusters { + assert_true(label.has_prefix("cluster_")) + } +} + +///| +test "scDblFinder fast clustering is deterministic" { + let (data, _, _, _) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let first = @src.scdf_fast_cluster(data, 3) catch { + _ => abort("valid fast clustering should run") + } + let second = @src.scdf_fast_cluster(data, 3) catch { + _ => abort("valid fast clustering should run") + } + assert_eq(first, second) +} + +///| +test "scDblFinder fast clustering validates cluster count" { + let data = scdfa_test_small_data() + let one = try { + ignore(@src.scdf_fast_cluster(data, 1)) + false + } catch { + _ => true + } + let too_many = try { + ignore(@src.scdf_fast_cluster(data, 5)) + false + } catch { + _ => true + } + assert_true(one && too_many) +} + +///| +test "scDblFinder workflow returns aligned cell-level arrays" { + let result = scdfa_test_result() + assert_eq(result.n_cells(), 42) + assert_eq(result.scores.length(), 42) + assert_eq(result.calls.length(), 42) + assert_eq(result.weighted_ratios.length(), 42) + assert_eq(result.neighbor_ratios.length(), 42) + assert_eq(result.cxds_scores.length(), 42) + assert_eq(result.origins.length(), 42) + assert_eq(result.library_sizes.length(), 42) +} + +///| +test "scDblFinder classifier scores are finite probabilities" { + let result = scdfa_test_result() + for score in result.scores { + assert_true(score == score) + assert_true(score >= 0.0 && score <= 1.0) + } +} + +///| +test "scDblFinder neighborhood and cxds features are bounded" { + let result = scdfa_test_result() + for cell in 0..= 0.0 && result.weighted_ratios[cell] <= 1.0, + ) + assert_true( + result.neighbor_ratios[cell] >= 0.0 && result.neighbor_ratios[cell] <= 1.0, + ) + assert_true( + result.cxds_scores[cell] >= 0.0 && result.cxds_scores[cell] <= 1.0, + ) + assert_true( + result.origin_ambiguities[cell] >= 0.0 && + result.origin_ambiguities[cell] <= 1.0, + ) + } +} + +///| +test "scDblFinder kNN distances and classes are populated" { + let result = scdfa_test_result() + for cell in 0..= 0.0) + assert_true(result.distances_to_doublet[cell] >= 0.0) + assert_true(result.distances_to_real[cell] >= 0.0) + assert_true( + result.nearest_classes[cell] == "cell" || + result.nearest_classes[cell] == "artificialDoublet", + ) + } +} + +///| +test "scDblFinder PCA has requested portable dimensions" { + let result = scdfa_test_result() + assert_eq(result.pca.length(), 42) + for coordinates in result.pca { + assert_eq(coordinates.length(), 5) + for value in coordinates { + assert_true(value == value && value.abs() <= 1.0e300) + } + } +} + +///| +test "scDblFinder feature selection records selected genes" { + let result = scdfa_test_result() + assert_true(result.selected_gene_names.length() > 0) + assert_true(result.selected_gene_names.length() <= result.gene_names.length()) + for gene in result.selected_gene_names { + assert_true(result.gene_names.contains(gene)) + } +} + +///| +test "scDblFinder performs independent capture fits" { + let result = scdfa_test_result() + assert_eq(result.sample_ids, ["capture_A", "capture_B"]) + assert_eq(result.thresholds.length(), 2) + assert_eq(result.expected_doublet_rates.length(), 2) + assert_eq(result.artificial_doublets, 96) + for threshold in result.thresholds { + assert_true(threshold.threshold >= 0.0 && threshold.threshold <= 1.0) + } +} + +///| +test "scDblFinder preserves known labels and workflow configuration" { + let result = scdfa_test_result() + let mut known = 0 + for value in result.known_doublets { + if value { + known = known + 1 + } + } + assert_eq(known, 6) + assert_eq(result.classifier_iterations, 2) + assert_eq(result.config.seed, 17) +} + +///| +test "scDblFinder ranks top cells by descending score" { + let result = scdfa_test_result() + let top = result.top_doublets(limit=7) catch { + _ => abort("valid top-doublet query should run") + } + assert_eq(top.length(), 7) + for index in 1..= top[index].score) + } +} + +///| +test "scDblFinder cell lookup returns typed diagnostics" { + let result = scdfa_test_result() + match result.cell("known_doublet_1") { + Some(cell) => { + assert_eq(cell.cell_name, "known_doublet_1") + assert_true(cell.score >= 0.0 && cell.score <= 1.0) + assert_true(cell.origin != "") + } + None => abort("known cell should be found") + } + assert_true(result.cell("missing") is None) +} + +///| +test "scDblFinder summary reports workflow metadata" { + let summary = scdfa_test_result().summary() + assert_true(summary.contains("scDblFinder Result")) + assert_true(summary.contains("cells=42")) + assert_true(summary.contains("samples=2")) + assert_true(summary.contains("iterations=2")) +} + +///| +test "scDblFinder singlet filtering follows calls" { + let (data, _, _, _) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let result = scdfa_test_result() + let filtered = result.filter_singlets(data) catch { + _ => abort("valid singlet filtering should run") + } + assert_eq( + filtered.cell_names.length(), + result.n_cells() - result.n_doublets(), + ) + assert_eq(filtered.gene_names, data.gene_names) +} + +///| +test "scDblFinder synthetic classification exceeds baseline accuracy" { + let (_, _, _, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let result = scdfa_test_result() + let mut correct = 0 + for cell in 0..= 0.75) +} + +///| +test "scDblFinder workflow is deterministic" { + let first = scdfa_test_result() + let second = scdfa_test_result() + assert_eq(first.scores, second.scores) + assert_eq(first.calls, second.calls) + assert_eq(first.origins, second.origins) + assert_eq(first.pca, second.pca) +} + +///| +test "scDblFinder supports a single combined capture" { + let (data, clusters, _, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let result = @src.sc_dbl_finder( + data, + clusters~, + known_doublets=truth, + config=scdfa_test_config(), + ) catch { + _ => abort("single-capture workflow should fit") + } + assert_eq(result.sample_ids, ["all"]) + assert_eq(result.thresholds.length(), 1) + assert_eq(result.artificial_doublets, 48) +} + +///| +test "scDblFinder can infer clusters before fitting" { + let (data, _, _, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let config = @src.ScDblFinderConfig::create( + expected_doublet_rate=0.15, + artificial_doublets=36, + selected_features=18, + dimensions=5, + neighbors=8, + iterations=1, + classifier_steps=80, + cluster_count=3, + ) catch { + _ => abort("valid auto-cluster configuration should build") + } + let result = @src.sc_dbl_finder(data, known_doublets=truth, config~) catch { + _ => abort("auto-cluster workflow should fit") + } + assert_eq(result.clusters.length(), 42) + for cluster in result.clusters { + assert_true(cluster != "") + } +} + +///| +test "scDblFinder workflow validates metadata lengths" { + let data = scdfa_test_small_data() + let clusters = try { + ignore(@src.sc_dbl_finder(data, clusters=["A", "B"])) + false + } catch { + _ => true + } + let samples = try { + ignore(@src.sc_dbl_finder(data, samples=["one"])) + false + } catch { + _ => true + } + let known = try { + ignore(@src.sc_dbl_finder(data, known_doublets=[true])) + false + } catch { + _ => true + } + assert_true(clusters && samples && known) +} + +///| +test "scDblFinder rejects captures with fewer than three cells" { + let data = scdfa_test_small_data() + let raised = try { + ignore( + @src.sc_dbl_finder(data, samples=["small", "small", "other", "other"]), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder cluster-aware fitting assigns origins" { + let result = scdfa_test_result() + for origin in result.origins { + assert_true(origin != "") + assert_true(origin.contains("+")) + } +} + +///| +test "scDblFinder pairwise enrichment covers cluster pairs" { + let enrichment = @src.scdf_pairwise_enrichment(scdfa_test_result()) catch { + _ => abort("valid enrichment should compute") + } + assert_eq(enrichment.length(), 3) + for index in 0..= 0) + assert_true(row.expected >= 0.0) + assert_true(row.p_value >= 0.0 && row.p_value <= 1.0) + assert_true( + row.adjusted_p_value >= row.p_value && row.adjusted_p_value <= 1.0, + ) + if index > 0 { + assert_true(enrichment[index - 1].p_value <= row.p_value) + } + } +} + +///| +test "scDblFinder enrichment requires cluster-aware fitting" { + let (data, _, _, truth) = @src.scdf_create_advanced_example() catch { + _ => abort("advanced example should build") + } + let result = @src.sc_dbl_finder( + data, + known_doublets=truth, + config=scdfa_test_config(), + ) catch { + _ => abort("valid cluster-free workflow should fit") + } + let raised = try { + ignore(@src.scdf_pairwise_enrichment(result)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder SingleCellExperiment writes cell diagnostics" { + let output = @src.sc_dbl_finder_single_cell_experiment( + scdfa_test_experiment(), + cluster_column="cluster", + sample_column="capture", + known_doublet_column="known", + config=scdfa_test_config(), + ) catch { + _ => abort("valid SingleCellExperiment workflow should fit") + } + assert_eq(output.result.n_cells(), 42) + assert_eq(output.experiment.col_data["scDblFinder.score"].length(), 42) + assert_eq(output.experiment.col_data["scDblFinder.class"].length(), 42) + assert_eq( + output.experiment.col_data["scDblFinder.mostLikelyOrigin"].length(), + 42, + ) +} + +///| +test "scDblFinder SingleCellExperiment writes feature and PCA diagnostics" { + let output = @src.sc_dbl_finder_single_cell_experiment( + scdfa_test_experiment(), + cluster_column="cluster", + sample_column="capture", + known_doublet_column="known", + config=scdfa_test_config(), + ) catch { + _ => abort("valid SingleCellExperiment workflow should fit") + } + assert_eq(output.experiment.row_data["scDblFinder.selected"].length(), 24) + assert_eq(output.experiment.reduced_dims["scDblFinder.PCA"].length(), 42) + assert_true(output.experiment.metadata.contains("scDblFinder.doublets")) + assert_true(output.experiment.metadata.contains("scDblFinder.rate")) + assert_true(output.experiment.metadata.contains("scDblFinder.artificial")) +} + +///| +test "scDblFinder SingleCellExperiment leaves the input unchanged" { + let experiment = scdfa_test_experiment() + let output = @src.sc_dbl_finder_single_cell_experiment( + experiment, + cluster_column="cluster", + sample_column="capture", + known_doublet_column="known", + config=scdfa_test_config(), + ) catch { + _ => abort("valid SingleCellExperiment workflow should fit") + } + assert_true(!experiment.col_data.contains("scDblFinder.score")) + output.experiment.col_data["capture"][0] = "changed" + assert_eq(experiment.col_data["capture"][0], "capture_A") +} + +///| +test "scDblFinder SingleCellExperiment supports a custom output prefix" { + let output = @src.sc_dbl_finder_single_cell_experiment( + scdfa_test_experiment(), + cluster_column="cluster", + sample_column="capture", + known_doublet_column="known", + output_prefix="doublet", + config=scdfa_test_config(), + ) catch { + _ => abort("valid custom-prefix workflow should fit") + } + assert_true(output.experiment.col_data.contains("doublet.score")) + assert_true(output.experiment.row_data.contains("doublet.selected")) + assert_true(output.experiment.reduced_dims.contains("doublet.PCA")) + assert_true(!output.experiment.col_data.contains("scDblFinder.score")) +} + +///| +test "scDblFinder SingleCellExperiment transposes gene-by-cell assays" { + let experiment = scdfa_test_experiment() + let output = @src.sc_dbl_finder_single_cell_experiment( + experiment, + cluster_column="cluster", + sample_column="capture", + known_doublet_column="known", + config=scdfa_test_config(), + ) catch { + _ => abort("valid SingleCellExperiment workflow should fit") + } + let mut expected_library = 0.0 + for gene in 0.. true + } + let column = try { + ignore( + @src.sc_dbl_finder_single_cell_experiment( + experiment, + cluster_column="missing", + ), + ) + false + } catch { + _ => true + } + assert_true(assay && column) +} + +///| +test "scDblFinder SingleCellExperiment rejects invalid known labels" { + let experiment = scdfa_test_experiment() + experiment.col_data["known"][0] = "unknown" + let raised = try { + ignore( + @src.sc_dbl_finder_single_cell_experiment( + experiment, + known_doublet_column="known", + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "scDblFinder SingleCellExperiment validates assay dimensions" { + let experiment = @src.SingleCellExperiment::new( + [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], + ["only_one_name"], + ["c1", "c2", "c3"], + ) + let raised = try { + ignore(@src.sc_dbl_finder_single_cell_experiment(experiment)) + false + } catch { + _ => true + } + assert_true(raised) +} diff --git a/test/moonbit/scenic_test.mbt b/test/moonbit/scenic_test.mbt index a22db54a..bac87835 100644 --- a/test/moonbit/scenic_test.mbt +++ b/test/moonbit/scenic_test.mbt @@ -6,32 +6,28 @@ // --------------------------------------------------------------------------- test "scenic_co_expression_module_creation" { - let mod_ = @src.co_expression_module( - "TP53", - ["G1", "G2", "G3"], - [0.9, 0.7, 0.5], - ) + let mod_ = @src.co_expression_module("TP53", ["G1", "G2", "G3"], [ + 0.9, 0.7, 0.5, + ]) assert_eq(mod_.tf_name, "TP53") assert_eq(mod_.targets.length(), 3) assert_eq(mod_.weights[0], 0.9) } +///| test "scenic_regulon_creation" { - let reg = @src.scenic_regulon( - "MYC", - ["G1", "G2", "G3", "G4"], - [1.0, 0.8, 0.6, 0.4], - ) + let reg = @src.scenic_regulon("MYC", ["G1", "G2", "G3", "G4"], [ + 1.0, 0.8, 0.6, 0.4, + ]) assert_eq(reg.tf_name, "MYC") assert_eq(reg.n_targets, 4) assert_eq(reg.targets[2], "G3") } +///| test "scenic_input_creation" { - let expr = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] // 2 genes x 3 cells - let input = @src.scenic_input( - expr, ["TF1", "G1"], ["C1", "C2", "C3"], ["TF1"], - ) + let expr = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]] // 2 genes x 3 cells + let input = @src.scenic_input(expr, ["TF1", "G1"], ["C1", "C2", "C3"], ["TF1"]) assert_eq(input.n_genes, 2) assert_eq(input.n_cells, 3) assert_eq(input.tf_names.length(), 1) @@ -41,6 +37,7 @@ test "scenic_input_creation" { // Binarize method constructors // --------------------------------------------------------------------------- +///| test "scenic_binarize_method_constructors" { match @src.binarize_mean_std() { @src.BinarizeMethod::MeanStd => assert_true(true) @@ -60,16 +57,15 @@ test "scenic_binarize_method_constructors" { // Synthetic data generator // --------------------------------------------------------------------------- +///| test "scenic_sample_data_shape" { - let input = @src.scenic_sample_data( - n_genes=60, n_cells=40, n_tfs=4, seed=10, - ) + let input = @src.scenic_sample_data(n_genes=60, n_cells=40, n_tfs=4, seed=10) assert_eq(input.n_genes, 60) assert_eq(input.n_cells, 40) assert_eq(input.tf_names.length(), 4) assert_eq(input.gene_names[0], "TF1") assert_eq(input.gene_names[4], "G1") - assert_eq(input.expression.length(), 60) // gene x cell + assert_eq(input.expression.length(), 60) // gene x cell assert_eq(input.expression[0].length(), 40) } @@ -77,13 +73,10 @@ test "scenic_sample_data_shape" { // Step 1: Co-expression modules // --------------------------------------------------------------------------- +///| test "scenic_build_coexpression_modules_basic" { - let input = @src.scenic_sample_data( - n_genes=50, n_cells=30, n_tfs=3, seed=42, - ) - let modules = @src.build_coexpression_modules( - input, min_targets=5, top_k=20, - ) + let input = @src.scenic_sample_data(n_genes=50, n_cells=30, n_tfs=3, seed=42) + let modules = @src.build_coexpression_modules(input, min_targets=5, top_k=20) // Should produce a module for each TF assert_eq(modules.length(), 3) for mod_ in modules { @@ -93,12 +86,13 @@ test "scenic_build_coexpression_modules_basic" { } } +///| test "scenic_build_coexpression_modules_skip_unknown_tf" { // TF name not in gene_names -> should be skipped let expr = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]] - let input = @src.scenic_input( - expr, ["G1", "G2", "G3"], ["C1", "C2"], ["NONEXISTENT"], - ) + let input = @src.scenic_input(expr, ["G1", "G2", "G3"], ["C1", "C2"], [ + "NONEXISTENT", + ]) let modules = @src.build_coexpression_modules(input, min_targets=1, top_k=5) assert_eq(modules.length(), 0) } @@ -107,15 +101,14 @@ test "scenic_build_coexpression_modules_skip_unknown_tf" { // Step 2: Regulon construction // --------------------------------------------------------------------------- +///| test "scenic_build_regulons_prunes_by_weight" { - let input = @src.scenic_sample_data( - n_genes=50, n_cells=30, n_tfs=3, seed=7, - ) - let modules = @src.build_coexpression_modules( - input, min_targets=5, top_k=20, - ) + let input = @src.scenic_sample_data(n_genes=50, n_cells=30, n_tfs=3, seed=7) + let modules = @src.build_coexpression_modules(input, min_targets=5, top_k=20) let regulons = @src.build_regulons( - modules, weight_quantile=0.5, min_targets=3, + modules, + weight_quantile=0.5, + min_targets=3, ) // Each regulon should have fewer or equal targets than its module assert_eq(regulons.length(), modules.length()) @@ -125,23 +118,29 @@ test "scenic_build_regulons_prunes_by_weight" { } } +///| test "scenic_build_regulons_filters_small_modules" { let modules = [ @src.co_expression_module("TF1", ["G1", "G2"], [0.9, 0.8]), - @src.co_expression_module("TF2", ["G1", "G2", "G3", "G4", "G5"], [0.9, 0.8, 0.7, 0.6, 0.5]), + @src.co_expression_module("TF2", ["G1", "G2", "G3", "G4", "G5"], [ + 0.9, 0.8, 0.7, 0.6, 0.5, + ]), ] - let regulons = @src.build_regulons(modules, weight_quantile=0.0, min_targets=5) + let regulons = @src.build_regulons( + modules, + weight_quantile=0.0, + min_targets=5, + ) // TF1 module has only 2 targets -> filtered out assert_eq(regulons.length(), 1) assert_eq(regulons[0].tf_name, "TF2") } +///| test "scenic_prune_by_motif_ranking" { - let mod_ = @src.co_expression_module( - "TF1", - ["G1", "G2", "G3", "G4", "G5"], - [0.9, 0.8, 0.7, 0.6, 0.5], - ) + let mod_ = @src.co_expression_module("TF1", ["G1", "G2", "G3", "G4", "G5"], [ + 0.9, 0.8, 0.7, 0.6, 0.5, + ]) // Motif ranking: G1, G3, G5 are in top 3 let motif_ranked = ["G1", "G3", "G5", "G2", "G4", "G6", "G7", "G8"] let reg = @src.prune_by_motif_ranking(mod_, motif_ranked, rank_threshold=3) @@ -157,32 +156,34 @@ test "scenic_prune_by_motif_ranking" { // Step 3: AUCell scoring // --------------------------------------------------------------------------- +///| test "scenic_build_cell_rankings" { let expr = [ - [5.0, 1.0, 3.0], // gene 0: high in cell 0 - [2.0, 4.0, 1.0], // gene 1: high in cell 1 - [1.0, 3.0, 2.0], // gene 2 + [5.0, 1.0, 3.0], // gene 0: high in cell 0 + [2.0, 4.0, 1.0], // gene 1: high in cell 1 + [1.0, 3.0, 2.0], // gene 2 ] - let input = @src.scenic_input( - expr, ["G0", "G1", "G2"], ["C0", "C1", "C2"], [], - ) + let input = @src.scenic_input(expr, ["G0", "G1", "G2"], ["C0", "C1", "C2"], []) let rankings = @src.build_cell_rankings(input) // Cell 0: gene 0 has expr 5.0 (highest), gene 1 has 2.0, gene 2 has 1.0 - assert_eq(rankings[0][0], 0) // gene 0 first - assert_eq(rankings[0][1], 1) // gene 1 second - assert_eq(rankings[0][2], 2) // gene 2 third + assert_eq(rankings[0][0], 0) // gene 0 first + assert_eq(rankings[0][1], 1) // gene 1 second + assert_eq(rankings[0][2], 2) // gene 2 third } +///| test "scenic_compute_regulon_activity_basic" { - let input = @src.scenic_sample_data( - n_genes=40, n_cells=20, n_tfs=2, seed=99, - ) - let modules = @src.build_coexpression_modules( - input, min_targets=5, top_k=15, + let input = @src.scenic_sample_data(n_genes=40, n_cells=20, n_tfs=2, seed=99) + let modules = @src.build_coexpression_modules(input, min_targets=5, top_k=15) + let regulons = @src.build_regulons( + modules, + weight_quantile=0.3, + min_targets=3, ) - let regulons = @src.build_regulons(modules, weight_quantile=0.3, min_targets=3) let auc_matrix = @src.compute_regulon_activity( - regulons, input, auc_threshold_pct=0.1, + regulons, + input, + auc_threshold_pct=0.1, ) assert_eq(auc_matrix.length(), regulons.length()) assert_eq(auc_matrix[0].length(), input.n_cells) @@ -198,26 +199,32 @@ test "scenic_compute_regulon_activity_basic" { // Step 4: Binarization // --------------------------------------------------------------------------- +///| test "scenic_binarize_mean_std" { let auc_matrix = [ - [0.1, 0.2, 0.8, 0.9, 0.1, 0.85], // bimodal: low {0.1,0.2,0.1}, high {0.8,0.9,0.85} + [0.1, 0.2, 0.8, 0.9, 0.1, 0.85], // bimodal: low {0.1,0.2,0.1}, high {0.8,0.9,0.85} ] let (binary, thresholds) = @src.binarize_activity( - auc_matrix, 6, method=@src.binarize_mean_std(), + auc_matrix, + 6, + method=@src.binarize_mean_std(), ) assert_eq(thresholds.length(), 1) assert_true(thresholds[0] > 0.3 && thresholds[0] < 0.7) // High-AUC cells should be 1, low-AUC cells should be 0 - assert_eq(binary[0][2], 1) // 0.8 > threshold - assert_eq(binary[0][0], 0) // 0.1 < threshold + assert_eq(binary[0][2], 1) // 0.8 > threshold + assert_eq(binary[0][0], 0) // 0.1 < threshold } +///| test "scenic_binarize_kmeans2" { let auc_matrix = [ - [0.1, 0.1, 0.1, 0.9, 0.9, 0.9], // clearly bimodal + [0.1, 0.1, 0.1, 0.9, 0.9, 0.9], // clearly bimodal ] let (binary, thresholds) = @src.binarize_activity( - auc_matrix, 6, method=@src.binarize_kmeans2(), + auc_matrix, + 6, + method=@src.binarize_kmeans2(), ) // K-means should separate at ~0.5 assert_true(thresholds[0] > 0.3 && thresholds[0] < 0.7) @@ -225,51 +232,50 @@ test "scenic_binarize_kmeans2" { assert_eq(binary[0][3], 1) } +///| test "scenic_binarize_median" { - let auc_matrix = [ - [0.1, 0.2, 0.3, 0.4, 0.5, 0.9], - ] + let auc_matrix = [[0.1, 0.2, 0.3, 0.4, 0.5, 0.9]] let (binary, _thresholds) = @src.binarize_activity( - auc_matrix, 6, method=@src.binarize_median(), + auc_matrix, + 6, + method=@src.binarize_median(), ) // Median of sorted [0.1,0.2,0.3,0.4,0.5,0.9] = 0.35 // Values > 0.35: 0.4, 0.5, 0.9 -> 1 - assert_eq(binary[0][5], 1) // 0.9 - assert_eq(binary[0][0], 0) // 0.1 + assert_eq(binary[0][5], 1) // 0.9 + assert_eq(binary[0][0], 0) // 0.1 } // --------------------------------------------------------------------------- // Cell state assignment // --------------------------------------------------------------------------- +///| test "scenic_assign_cell_states_basic" { // 2 regulons, 4 cells // Regulon 0 active in cells 0,1; Regulon 1 active in cells 2,3 - let binary = [ - [1, 1, 0, 0], - [0, 0, 1, 1], - ] + let binary = [[1, 1, 0, 0], [0, 0, 1, 1]] let (labels, n_clusters, masters) = @src.assign_cell_states( - binary, ["TF1", "TF2"], 4, + binary, + ["TF1", "TF2"], + 4, ) - assert_eq(labels[0], 1) // cell 0 -> cluster 1 (TF1) - assert_eq(labels[1], 1) // cell 1 -> cluster 1 - assert_eq(labels[2], 2) // cell 2 -> cluster 2 (TF2) - assert_eq(labels[3], 2) // cell 3 -> cluster 2 + assert_eq(labels[0], 1) // cell 0 -> cluster 1 (TF1) + assert_eq(labels[1], 1) // cell 1 -> cluster 1 + assert_eq(labels[2], 2) // cell 2 -> cluster 2 (TF2) + assert_eq(labels[3], 2) // cell 3 -> cluster 2 assert_eq(n_clusters, 2) assert_eq(masters.length(), 2) } +///| test "scenic_assign_cell_states_no_active_regulon" { - let binary = [ - [0, 0, 0], - [0, 0, 0], - ] - let (labels, _n, masters) = @src.assign_cell_states( - binary, ["TF1", "TF2"], 3, - ) + let binary = [[0, 0, 0], [0, 0, 0]] + let (labels, _n, masters) = @src.assign_cell_states(binary, ["TF1", "TF2"], 3) // All cells should be in cluster 0 (no active regulon) - for l in labels { assert_eq(l, 0) } + for l in labels { + assert_eq(l, 0) + } assert_eq(masters.length(), 0) } @@ -277,13 +283,20 @@ test "scenic_assign_cell_states_no_active_regulon" { // End-to-end pipeline // --------------------------------------------------------------------------- +///| test "scenic_run_pipeline_end_to_end" { let input = @src.scenic_sample_data( - n_genes=60, n_cells=40, n_tfs=4, seed=2025, + n_genes=60, + n_cells=40, + n_tfs=4, + seed=2025, ) let result = @src.run_scenic( - input, min_targets=5, top_k=20, - weight_quantile=0.3, auc_threshold_pct=0.1, + input, + min_targets=5, + top_k=20, + weight_quantile=0.3, + auc_threshold_pct=0.1, binarize_method=@src.binarize_mean_std(), ) assert_eq(result.n_cells, 40) @@ -295,10 +308,9 @@ test "scenic_run_pipeline_end_to_end" { assert_eq(result.thresholds.length(), result.n_regulons) } +///| test "scenic_result_summary" { - let input = @src.scenic_sample_data( - n_genes=40, n_cells=20, n_tfs=3, seed=1, - ) + let input = @src.scenic_sample_data(n_genes=40, n_cells=20, n_tfs=3, seed=1) let result = @src.run_scenic(input) let s = result.summary() assert_true(s.contains("regulons=")) @@ -306,10 +318,9 @@ test "scenic_result_summary" { assert_true(s.contains("clusters=")) } +///| test "scenic_result_auc_at" { - let input = @src.scenic_sample_data( - n_genes=40, n_cells=10, n_tfs=2, seed=5, - ) + let input = @src.scenic_sample_data(n_genes=40, n_cells=10, n_tfs=2, seed=5) let result = @src.run_scenic(input) // Valid indices should return a value in [0,1] if result.n_regulons > 0 { @@ -321,24 +332,22 @@ test "scenic_result_auc_at" { assert_eq(result.auc_at(999, 0), 0.0) } +///| test "scenic_result_top_regulons" { - let input = @src.scenic_sample_data( - n_genes=50, n_cells=30, n_tfs=4, seed=8, - ) + let input = @src.scenic_sample_data(n_genes=50, n_cells=30, n_tfs=4, seed=8) let result = @src.run_scenic(input) let top = result.top_regulons(n=3) assert_true(top.length() <= 3) if top.length() >= 2 { let (_, s0) = top[0] let (_, s1) = top[1] - assert_true(s0 >= s1) // sorted descending + assert_true(s0 >= s1) // sorted descending } } +///| test "scenic_result_regulon_targets" { - let input = @src.scenic_sample_data( - n_genes=40, n_cells=20, n_tfs=2, seed=3, - ) + let input = @src.scenic_sample_data(n_genes=40, n_cells=20, n_tfs=2, seed=3) let result = @src.run_scenic(input) if result.n_regulons > 0 { let targets = result.regulon_targets(0) diff --git a/test/moonbit/scmap_test.mbt b/test/moonbit/scmap_test.mbt index deb74bcb..2deb171c 100644 --- a/test/moonbit/scmap_test.mbt +++ b/test/moonbit/scmap_test.mbt @@ -15,12 +15,14 @@ test "scmap_reference_creation" { assert_eq(ref.cell_types()[6], "NK_cell") } +///| test "scmap_reference_unique_types" { let ref = @src.scmap_sample_reference() let types = ref.unique_cell_types() assert_eq(types.length(), 3) } +///| test "scmap_query_creation" { let q = @src.scmap_sample_query() assert_eq(q.gene_names().length(), 10) @@ -30,10 +32,9 @@ test "scmap_query_creation" { assert_eq(q.cell_names()[2], "query_unknown") } +///| test "scmap_assignment_creation" { - let a = @src.ScmapAssignment::new( - "cell1", "T_cell", 0.85, 0.3, "cluster", - ) + let a = @src.ScmapAssignment::new("cell1", "T_cell", 0.85, 0.3, "cluster") assert_eq(a.cell_name(), "cell1") assert_eq(a.assigned_type(), "T_cell") assert_eq(a.best_correlation(), 0.85) @@ -45,6 +46,7 @@ test "scmap_assignment_creation" { // scmap-cluster method // --------------------------------------------------------------------------- +///| test "scmap_cluster_classifies_t_cell" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -56,6 +58,7 @@ test "scmap_cluster_classifies_t_cell" { assert_eq(assignments[0].method(), "cluster") } +///| test "scmap_cluster_classifies_b_cell" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -65,6 +68,7 @@ test "scmap_cluster_classifies_b_cell" { assert_eq(assignments[1].assigned_type(), "B_cell") } +///| test "scmap_cluster_high_threshold_unassigned" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -75,6 +79,7 @@ test "scmap_cluster_high_threshold_unassigned" { } } +///| test "scmap_cluster_correlation_is_valid" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -85,6 +90,7 @@ test "scmap_cluster_correlation_is_valid" { } } +///| test "scmap_cluster_best_gt_second_best" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -98,6 +104,7 @@ test "scmap_cluster_best_gt_second_best" { // scmap-cell method // --------------------------------------------------------------------------- +///| test "scmap_cell_classifies_t_cell" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -107,6 +114,7 @@ test "scmap_cell_classifies_t_cell" { assert_eq(assignments[0].method(), "cell") } +///| test "scmap_cell_classifies_b_cell" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -115,6 +123,7 @@ test "scmap_cell_classifies_b_cell" { assert_eq(assignments[1].assigned_type(), "B_cell") } +///| test "scmap_cell_k_neighbours_parameter" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -123,6 +132,7 @@ test "scmap_cell_k_neighbours_parameter" { assert_eq(assignments.length(), 3) } +///| test "scmap_cell_high_threshold_unassigned" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -140,6 +150,7 @@ test "scmap_cell_high_threshold_unassigned" { // Summary utilities // --------------------------------------------------------------------------- +///| test "scmap_summary_counts" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -156,16 +167,16 @@ test "scmap_summary_counts" { assert_eq(total, assignments.length()) } +///| test "scmap_assignment_to_string" { - let a = @src.ScmapAssignment::new( - "cell1", "T_cell", 0.85, 0.3, "cluster", - ) + let a = @src.ScmapAssignment::new("cell1", "T_cell", 0.85, 0.3, "cluster") let s = a.to_string() assert_true(s.contains("cell1")) assert_true(s.contains("T_cell")) assert_true(s.contains("cluster")) } +///| test "scmap_assignments_to_string" { let ref = @src.scmap_sample_reference() let query = @src.scmap_sample_query() @@ -179,6 +190,7 @@ test "scmap_assignments_to_string" { // Sample data verification // --------------------------------------------------------------------------- +///| test "scmap_sample_reference_expression_structure" { let ref = @src.scmap_sample_reference() // 10 genes × 9 cells @@ -190,6 +202,7 @@ test "scmap_sample_reference_expression_structure" { assert_true(ref.expression()[3][3] > ref.expression()[3][0]) } +///| test "scmap_sample_query_expression_structure" { let q = @src.scmap_sample_query() // 10 genes × 3 cells @@ -203,6 +216,7 @@ test "scmap_sample_query_expression_structure" { // Edge cases // --------------------------------------------------------------------------- +///| test "scmap_cluster_single_query_cell" { let ref = @src.scmap_sample_reference() // Create a single-cell query using the T cell reference profile @@ -215,6 +229,7 @@ test "scmap_cluster_single_query_cell" { assert_eq(assignments.length(), 1) } +///| test "scmap_cell_single_query_cell" { let ref = @src.scmap_sample_reference() let q = @src.scmap_sample_query() diff --git a/test/moonbit/scnorm_test.mbt b/test/moonbit/scnorm_test.mbt index 8d4dffd2..f3dd380e 100644 --- a/test/moonbit/scnorm_test.mbt +++ b/test/moonbit/scnorm_test.mbt @@ -12,7 +12,7 @@ test "scnorm_library_sizes" { let counts : Array[Array[Double]] = [ [10.0, 20.0, 30.0], [5.0, 15.0, 25.0], - [50.0, 100.0, 150.0] + [50.0, 100.0, 150.0], ] let sizes = @src.sc_norm_library_sizes(counts) assert_eq(sizes.length(), 3) @@ -61,7 +61,7 @@ test "scnorm_full_run" { let count_matrix : Array[Array[Double]] = [ [10.0, 20.0, 30.0], [50.0, 60.0, 70.0], - [100.0, 200.0, 300.0] + [100.0, 200.0, 300.0], ] let gene_ids = ["gene1", "gene2", "gene3"] let cell_ids = ["cell1", "cell2", "cell3"] @@ -74,7 +74,7 @@ test "scnorm_full_run" { test "scnorm_result_accessors" { let count_matrix : Array[Array[Double]] = [ [10.0, 20.0, 30.0], - [50.0, 60.0, 70.0] + [50.0, 60.0, 70.0], ] let gene_ids = ["gene1", "gene2"] let cell_ids = ["cell1", "cell2", "cell3"] diff --git a/test/moonbit/scop_test.mbt b/test/moonbit/scop_test.mbt index c1ed0188..0e899c9d 100644 --- a/test/moonbit/scop_test.mbt +++ b/test/moonbit/scop_test.mbt @@ -17,6 +17,7 @@ test "scop_node_type_name" { assert_eq(@src.scop_node_type_name("xx"), "unknown") } +///| test "scop_node_type_code" { assert_eq(@src.scop_node_type_code("root"), "ro") assert_eq(@src.scop_node_type_code("class"), "cl") @@ -28,6 +29,7 @@ test "scop_node_type_code" { assert_eq(@src.scop_node_type_code("domain"), "px") } +///| test "scop_node_type_order" { assert_true(@src.scop_node_type_order("ro") < @src.scop_node_type_order("cl")) assert_true(@src.scop_node_type_order("cl") < @src.scop_node_type_order("cf")) @@ -42,18 +44,21 @@ test "scop_node_type_order" { // Residues parsing // --------------------------------------------------------------------------- +///| test "scop_parse_residues_dash" { let r = @src.scop_parse_residues("-") assert_eq(r.pdbid(), "") assert_eq(r.fragments().length(), 0) } +///| test "scop_parse_residues_empty" { let r = @src.scop_parse_residues("") assert_eq(r.pdbid(), "") assert_eq(r.fragments().length(), 0) } +///| test "scop_parse_residues_paren_dash" { let r = @src.scop_parse_residues("(-)") assert_eq(r.pdbid(), "") @@ -63,6 +68,7 @@ test "scop_parse_residues_paren_dash" { assert_eq(r.fragments()[0].end_(), "") } +///| test "scop_parse_residues_with_pdbid" { let r = @src.scop_parse_residues("1bba A:10-20,B:") assert_eq(r.pdbid(), "1bba") @@ -77,6 +83,7 @@ test "scop_parse_residues_with_pdbid" { assert_eq(r.fragments()[1].end_(), "") } +///| test "scop_parse_residues_single_chain" { let r = @src.scop_parse_residues("A:1-141") assert_eq(r.pdbid(), "") @@ -86,6 +93,7 @@ test "scop_parse_residues_single_chain" { assert_eq(r.fragments()[0].end_(), "141") } +///| test "scop_residues_to_string_roundtrip" { let r = @src.scop_parse_residues("1hba A:1-141") let s = @src.scop_residues_to_string(r) @@ -97,6 +105,7 @@ test "scop_residues_to_string_roundtrip" { assert_eq(r2.fragments()[0].end_(), "141") } +///| test "scop_residues_to_string_dash" { let r = @src.ScopResidues::new("", []) let s = @src.scop_residues_to_string(r) @@ -107,12 +116,15 @@ test "scop_residues_to_string_dash" { // Record creation and accessors // --------------------------------------------------------------------------- +///| test "scop_cla_record_creation" { let res = @src.scop_parse_residues("1hba A:1-141") let hier : Map[String, Int] = Map([], capacity=8) hier.set("cl", 100) hier.set("cf", 200) - let r = @src.ClaRecord::new("d1hba_", "1hba", res, "a.1.1.1.1.1.1", 1000, hier) + let r = @src.ClaRecord::new( + "d1hba_", "1hba", res, "a.1.1.1.1.1.1", 1000, hier, + ) assert_eq(r.sid(), "d1hba_") assert_eq(r.pdbid(), "1hba") assert_eq(r.sccs(), "a.1.1.1.1.1.1") @@ -120,6 +132,7 @@ test "scop_cla_record_creation" { assert_eq(r.hierarchy().get("cl").unwrap(), 100) } +///| test "scop_des_record_creation" { let r = @src.DesRecord::new( 1000, "px", "a.1.1.1.1.1.1", "d1hba_", "1hba Hemoglobin alpha chain", @@ -131,6 +144,7 @@ test "scop_des_record_creation" { assert_eq(r.description(), "1hba Hemoglobin alpha chain") } +///| test "scop_hie_record_creation" { let r = @src.HieRecord::new(600, 500, [1000, 1001]) assert_eq(r.sunid(), 600) @@ -144,6 +158,7 @@ test "scop_hie_record_creation" { // File parsers // --------------------------------------------------------------------------- +///| test "scop_parse_cla_line" { let line = "d1hba_\t1hba\tA:1-141\ta.1.1.1.1.1.1\t1000\tcl=100,cf=200,sf=300,fa=400,dm=500,sp=600,px=1000" let r = @src.scop_parse_cla_line(line) @@ -157,16 +172,19 @@ test "scop_parse_cla_line" { assert_eq(rec.hierarchy().get("px").unwrap(), 1000) } +///| test "scop_parse_cla_line_comment" { let r = @src.scop_parse_cla_line("# comment line") assert_true(r.is_none()) } +///| test "scop_parse_cla_line_empty" { let r = @src.scop_parse_cla_line("") assert_true(r.is_none()) } +///| test "scop_parse_des_line" { let line = "1000\tpx\ta.1.1.1.1.1.1\td1hba_\t1hba Hemoglobin alpha chain" let r = @src.scop_parse_des_line(line) @@ -177,6 +195,7 @@ test "scop_parse_des_line" { assert_eq(rec.name(), "d1hba_") } +///| test "scop_parse_des_line_class" { let line = "100\tcl\ta\t-\tAll alpha proteins" let r = @src.scop_parse_des_line(line) @@ -188,6 +207,7 @@ test "scop_parse_des_line_class" { assert_eq(rec.description(), "All alpha proteins") } +///| test "scop_parse_hie_line" { let line = "600\t500\t1000,1001" let r = @src.scop_parse_hie_line(line) @@ -198,6 +218,7 @@ test "scop_parse_hie_line" { assert_eq(rec.children().length(), 2) } +///| test "scop_parse_hie_line_root" { let line = "0\t-\t100" let r = @src.scop_parse_hie_line(line) @@ -208,6 +229,7 @@ test "scop_parse_hie_line_root" { assert_eq(rec.children().length(), 1) } +///| test "scop_parse_hie_line_leaf" { let line = "1000\t600\t-" let r = @src.scop_parse_hie_line(line) @@ -218,6 +240,7 @@ test "scop_parse_hie_line_leaf" { assert_eq(rec.children().length(), 0) } +///| test "scop_parse_cla_multiple" { let content = @src.scop_sample_cla() let records = @src.scop_parse_cla(content) @@ -226,6 +249,7 @@ test "scop_parse_cla_multiple" { assert_eq(records[1].sid(), "d1hbb_") } +///| test "scop_parse_des_multiple" { let content = @src.scop_sample_des() let records = @src.scop_parse_des(content) @@ -233,6 +257,7 @@ test "scop_parse_des_multiple" { assert_eq(records.length(), 9) } +///| test "scop_parse_hie_multiple" { let content = @src.scop_sample_hie() let records = @src.scop_parse_hie(content) @@ -243,6 +268,7 @@ test "scop_parse_hie_multiple" { // Scop hierarchy construction and queries // --------------------------------------------------------------------------- +///| test "scop_build_hierarchy" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -255,6 +281,7 @@ test "scop_build_hierarchy" { assert_eq(scop.get_domains().length(), 2) } +///| test "scop_get_node_by_sunid" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -269,6 +296,7 @@ test "scop_get_node_by_sunid" { assert_true(none.is_none()) } +///| test "scop_get_domain_by_sid" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -285,6 +313,7 @@ test "scop_get_domain_by_sid" { assert_true(none.is_none()) } +///| test "scop_get_parent" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -301,6 +330,7 @@ test "scop_get_parent" { assert_true(root_parent.is_none()) } +///| test "scop_get_children" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -317,6 +347,7 @@ test "scop_get_children" { assert_eq(scop.get_children(leaf).length(), 0) } +///| test "scop_get_ascendent_by_code" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -334,6 +365,7 @@ test "scop_get_ascendent_by_code" { assert_eq(fa.unwrap().sunid(), 400) } +///| test "scop_get_ascendent_by_long_name" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -347,6 +379,7 @@ test "scop_get_ascendent_by_long_name" { assert_eq(fold.unwrap().description(), "Globin-like") } +///| test "scop_get_descendents" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -365,6 +398,7 @@ test "scop_get_descendents" { assert_eq(families[0].sunid(), 400) } +///| test "scop_get_descendents_long_name" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -374,6 +408,7 @@ test "scop_get_descendents_long_name" { assert_eq(domains.length(), 2) } +///| test "scop_node_is_domain" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -389,27 +424,32 @@ test "scop_node_is_domain" { // SCCS comparison // --------------------------------------------------------------------------- +///| test "scop_cmp_sccs_equal" { assert_eq(@src.scop_cmp_sccs("a.1.1.1", "a.1.1.1"), 0) } +///| test "scop_cmp_sccs_letter_diff" { assert_true(@src.scop_cmp_sccs("a.1.1.1", "b.1.1.1") < 0) assert_true(@src.scop_cmp_sccs("b.1.1.1", "a.1.1.1") > 0) } +///| test "scop_cmp_sccs_numeric_diff" { assert_true(@src.scop_cmp_sccs("a.1.1.1", "a.1.1.2") < 0) assert_true(@src.scop_cmp_sccs("a.1.1.2", "a.1.1.1") > 0) assert_true(@src.scop_cmp_sccs("a.1.1.1", "a.1.2.1") < 0) } +///| test "scop_cmp_sccs_length_diff" { // Shorter prefix sorts first when all compared components are equal assert_true(@src.scop_cmp_sccs("a.1", "a.1.1") < 0) assert_true(@src.scop_cmp_sccs("a.1.1", "a.1") > 0) } +///| test "scop_cmp_sccs_numeric_not_lexical" { // Numerically, 2 < 11, so a.1.2 < a.1.11 (NOT lexical where "2" > "11") assert_true(@src.scop_cmp_sccs("a.1.2", "a.1.11") < 0) @@ -420,6 +460,7 @@ test "scop_cmp_sccs_numeric_not_lexical" { // Serialization // --------------------------------------------------------------------------- +///| test "scop_write_hie" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -432,6 +473,7 @@ test "scop_write_hie" { assert_true(output.contains("1000\t600\t-")) } +///| test "scop_write_des" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -442,6 +484,7 @@ test "scop_write_des" { assert_true(output.contains("d1hba_")) } +///| test "scop_write_cla" { let cla = @src.scop_parse_cla(@src.scop_sample_cla()) let des = @src.scop_parse_des(@src.scop_sample_des()) @@ -457,6 +500,7 @@ test "scop_write_cla" { // Empty Scop // --------------------------------------------------------------------------- +///| test "scop_empty" { let scop = @src.scop_empty() assert_eq(scop.root().sunid(), 0) diff --git a/test/moonbit/scrapper_test.mbt b/test/moonbit/scrapper_test.mbt new file mode 100644 index 00000000..19dfea0d --- /dev/null +++ b/test/moonbit/scrapper_test.mbt @@ -0,0 +1,653 @@ +///| +/// Tests for Bioconductor scrapper-inspired single-cell preprocessing. + +///| +fn scrapper_test_assert_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort("values differ beyond tolerance") + } +} + +///| +fn scrapper_test_sce() -> @src.SingleCellExperiment { + @src.SCEBuilder::new() + |> @src.SCEBuilder::add_assay("counts", [ + [3.0, 6.0, 9.0], + [1.0, 2.0, 3.0], + [0.0, 1.0, 0.0], + ]) + |> @src.SCEBuilder::add_row_data("symbol", ["A", "B", "MT-C"]) + |> @src.SCEBuilder::add_col_data("batch", ["x", "x", "y"]) + |> @src.SCEBuilder::add_reduced_dim("PCA", [ + [0.1, 0.2], + [0.3, 0.4], + [0.5, 0.6], + ]) + |> @src.SCEBuilder::build +} + +///| +test "scrapper: sample count matrix uses feature by cell orientation" { + let counts = @src.scrapper_sample_counts() + assert_eq(counts.length(), 6) + assert_eq(counts[0].length(), 6) + assert_eq(counts[0][3], 60.0) +} + +///| +test "scrapper: computes RNA QC sums detected genes and subset proportions" { + let metrics = @src.scrapper_compute_rna_qc_metrics( + [[10.0, 0.0, 1.0], [0.0, 0.0, 1.0], [5.0, 0.0, 0.0]], + subsets=[@src.ScrapperNamedSubset::new("mito", [0, 2])], + ) catch { + _ => abort("valid RNA QC input should be accepted") + } + assert_eq(metrics.sums, [15.0, 0.0, 2.0]) + assert_eq(metrics.detected, [2, 0, 2]) + assert_eq(metrics.subset_names, ["mito"]) + assert_eq(metrics.subset_proportions[0], [1.0, 0.0, 0.5]) +} + +///| +test "scrapper: RNA QC detection limit uses strict greater than comparison" { + let metrics = @src.scrapper_compute_rna_qc_metrics( + [[2.0, 1.0], [1.0, 1.0], [3.0, 0.0]], + detection_limit=1.0, + ) catch { + _ => abort("valid detection limit should be accepted") + } + assert_eq(metrics.sums, [6.0, 2.0]) + assert_eq(metrics.detected, [2, 0]) +} + +///| +test "scrapper: RNA QC rejects malformed count matrices" { + let ragged = try { + ignore(@src.scrapper_compute_rna_qc_metrics([[1.0], [2.0, 3.0]])) + false + } catch { + ScrapperError(_) => true + } + let negative = try { + ignore(@src.scrapper_compute_rna_qc_metrics([[1.0, -1.0]])) + false + } catch { + ScrapperError(_) => true + } + let non_finite = try { + ignore(@src.scrapper_compute_rna_qc_metrics([[1.0, @double.not_a_number]])) + false + } catch { + ScrapperError(_) => true + } + assert_true(ragged) + assert_true(negative) + assert_true(non_finite) +} + +///| +test "scrapper: RNA QC validates named feature subsets" { + let duplicate = try { + ignore( + @src.scrapper_compute_rna_qc_metrics([[1.0], [2.0]], subsets=[ + @src.ScrapperNamedSubset::new("mito", [0]), + @src.ScrapperNamedSubset::new("mito", [1]), + ]), + ) + false + } catch { + ScrapperError(_) => true + } + let out_of_bounds = try { + ignore( + @src.scrapper_compute_rna_qc_metrics([[1.0]], subsets=[ + @src.ScrapperNamedSubset::new("mito", [1]), + ]), + ) + false + } catch { + ScrapperError(_) => true + } + assert_true(duplicate) + assert_true(out_of_bounds) +} + +///| +test "scrapper: unblocked MAD thresholds filter low-library cells" { + let metrics = @src.scrapper_compute_rna_qc_metrics([[1.0, 10.0, 100.0]]) catch { + _ => abort("valid RNA QC input should be accepted") + } + let thresholds = @src.scrapper_suggest_rna_qc_thresholds( + metrics, + sum_num_mads=0.0, + detected_num_mads=0.0, + ) catch { + _ => abort("valid RNA QC metrics should produce thresholds") + } + let keep = @src.scrapper_filter_rna_qc_metrics(thresholds, metrics) catch { + _ => abort("thresholds should apply to their source metrics") + } + assert_eq(thresholds.block_names, ["all"]) + scrapper_test_assert_close(thresholds.sum_lower[0], 10.0, 1.0e-10) + assert_eq(keep, [false, true, true]) +} + +///| +test "scrapper: blocked MAD thresholds are estimated independently" { + let metrics = @src.scrapper_compute_rna_qc_metrics([[1.0, 2.0, 100.0, 200.0]]) catch { + _ => abort("valid RNA QC input should be accepted") + } + let blocks = ["A", "A", "B", "B"] + let thresholds = @src.scrapper_suggest_rna_qc_thresholds( + metrics, + blocks~, + sum_num_mads=0.0, + detected_num_mads=0.0, + ) catch { + _ => abort("valid blocks should produce thresholds") + } + let keep = @src.scrapper_filter_rna_qc_metrics(thresholds, metrics, blocks~) catch { + _ => abort("matching blocks should be accepted") + } + assert_eq(thresholds.block_names, ["A", "B"]) + assert_true(thresholds.sum_lower[1] > thresholds.sum_lower[0]) + assert_eq(keep, [false, true, false, true]) +} + +///| +test "scrapper: subset MAD upper threshold removes high-proportion cells" { + let metrics = @src.scrapper_compute_rna_qc_metrics( + [[9.0, 1.0, 5.0], [1.0, 9.0, 5.0]], + subsets=[@src.ScrapperNamedSubset::new("mito", [0])], + ) catch { + _ => abort("valid RNA QC input should be accepted") + } + let thresholds = @src.scrapper_suggest_rna_qc_thresholds( + metrics, + sum_num_mads=100.0, + detected_num_mads=100.0, + subset_proportion_num_mads=0.0, + ) catch { + _ => abort("valid RNA QC metrics should produce thresholds") + } + let keep = @src.scrapper_filter_rna_qc_metrics(thresholds, metrics) catch { + _ => abort("thresholds should apply to their source metrics") + } + assert_eq(thresholds.subset_upper[0], [0.5]) + assert_eq(keep, [false, true, true]) +} + +///| +test "scrapper: filter rejects a different block level ordering" { + let metrics = @src.scrapper_compute_rna_qc_metrics([[1.0, 2.0]]) catch { + _ => abort("valid RNA QC input should be accepted") + } + let thresholds = @src.scrapper_suggest_rna_qc_thresholds(metrics, blocks=[ + "A", "B", + ]) catch { + _ => abort("valid blocks should produce thresholds") + } + let raised = try { + ignore( + @src.scrapper_filter_rna_qc_metrics(thresholds, metrics, blocks=["B", "A"]), + ) + false + } catch { + ScrapperError(_) => true + } + assert_true(raised) +} + +///| +test "scrapper: quick RNA QC returns metrics thresholds and summary" { + let result = @src.scrapper_quick_rna_qc( + [[1.0, 10.0, 100.0]], + sum_num_mads=0.0, + detected_num_mads=0.0, + ) catch { + _ => abort("quick RNA QC should accept valid counts") + } + assert_eq(result.keep, [false, true, true]) + assert_eq(result.retained_count(), 2) + assert_eq(result.summary(), "ScrapperRNAQC(cells=3, retained=2, subsets=0)") +} + +///| +test "scrapper: sanitizes invalid size factors without mutating input" { + let input = [2.0, 0.0, -1.0, @double.not_a_number] + let output = @src.scrapper_sanitize_size_factors(input, fallback=3.0) catch { + _ => abort("positive fallback should be accepted") + } + assert_eq(output, [2.0, 3.0, 3.0, 3.0]) + assert_eq(input[0], 2.0) + assert_eq(input[1], 0.0) + assert_true(input[3].is_nan()) +} + +///| +test "scrapper: size-factor sanitization validates fallback" { + let raised = try { + ignore(@src.scrapper_sanitize_size_factors([1.0], fallback=0.0)) + false + } catch { + ScrapperError(_) => true + } + assert_true(raised) +} + +///| +test "scrapper: centers size factors across all cells" { + let centered = @src.scrapper_center_size_factors([1.0, 2.0, 3.0]) catch { + _ => abort("positive size factors should be centered") + } + assert_eq(centered, [0.5, 1.0, 1.5]) +} + +///| +test "scrapper: per-block centering gives each block unit mean" { + let centered = @src.scrapper_center_size_factors( + [1.0, 3.0, 10.0, 20.0], + blocks=["A", "A", "B", "B"], + mode=@src.scrapper_center_per_block(), + ) catch { + _ => abort("valid blocked size factors should be centered") + } + assert_eq(centered[0], 0.5) + assert_eq(centered[1], 1.5) + scrapper_test_assert_close(centered[2], 2.0 / 3.0, 1.0e-12) + scrapper_test_assert_close(centered[3], 4.0 / 3.0, 1.0e-12) +} + +///| +test "scrapper: lowest-block centering preserves between-block scale" { + let centered = @src.scrapper_center_size_factors( + [1.0, 3.0, 10.0, 20.0], + blocks=["A", "A", "B", "B"], + mode=@src.scrapper_center_lowest(), + ) catch { + _ => abort("valid blocked size factors should be centered") + } + assert_eq(centered, [0.5, 1.5, 5.0, 10.0]) +} + +///| +test "scrapper: computes centered library size factors" { + let factors = @src.scrapper_library_size_factors([[1.0, 2.0], [3.0, 6.0]]) catch { + _ => abort("valid counts should produce library size factors") + } + scrapper_test_assert_close(factors[0], 2.0 / 3.0, 1.0e-12) + scrapper_test_assert_close(factors[1], 4.0 / 3.0, 1.0e-12) +} + +///| +test "scrapper: zero libraries receive finite fallback factors" { + let factors = @src.scrapper_library_size_factors([[0.0, 0.0]]) catch { + _ => abort("zero libraries should use sanitized factors") + } + assert_eq(factors, [1.0, 1.0]) +} + +///| +test "scrapper: performs linear count scaling" { + let normalized = @src.scrapper_normalize_counts( + [[2.0, 8.0], [4.0, 4.0]], + [2.0, 4.0], + log_transform=false, + ) catch { + _ => abort("valid size factors should normalize counts") + } + assert_eq(normalized, [[1.0, 2.0], [2.0, 1.0]]) +} + +///| +test "scrapper: performs configurable log normalization" { + let normalized = @src.scrapper_normalize_counts( + [[3.0, 8.0]], + [1.0, 2.0], + pseudo_count=1.0, + log_base=2.0, + ) catch { + _ => abort("valid size factors should normalize counts") + } + scrapper_test_assert_close(normalized[0][0], 2.0, 1.0e-12) + scrapper_test_assert_close(normalized[0][1], 2.321928094887362, 1.0e-12) +} + +///| +test "scrapper: normalization validates size factors and logarithm base" { + let wrong_length = try { + ignore(@src.scrapper_normalize_counts([[1.0, 2.0]], [1.0])) + false + } catch { + ScrapperError(_) => true + } + let non_positive = try { + ignore(@src.scrapper_normalize_counts([[1.0]], [0.0])) + false + } catch { + ScrapperError(_) => true + } + let invalid_base = try { + ignore(@src.scrapper_normalize_counts([[1.0]], [1.0], log_base=1.0)) + false + } catch { + ScrapperError(_) => true + } + assert_true(wrong_length) + assert_true(non_positive) + assert_true(invalid_base) +} + +///| +test "scrapper: models per-gene means and sample variances" { + let model = @src.scrapper_model_gene_variances( + [[1.0, 2.0, 3.0], [2.0, 2.0, 2.0], [0.0, 0.0, 0.0]], + mean_filter=false, + transform=false, + span=1.0, + min_window_count=1, + ) catch { + _ => abort("valid expression should produce a variance model") + } + assert_eq(model.means, [2.0, 2.0, 0.0]) + assert_eq(model.variances, [1.0, 0.0, 0.0]) + let mut feature = 0 + while feature < model.means.length() { + scrapper_test_assert_close( + model.residuals[feature], + model.variances[feature] - model.fitted[feature], + 1.0e-12, + ) + feature = feature + 1 + } +} + +///| +test "scrapper: LOWESS reproduces a linear mean-variance trend" { + let trend = @src.scrapper_fit_variance_trend( + [1.0, 2.0, 3.0, 4.0, 5.0], + [2.0, 4.0, 6.0, 8.0, 10.0], + mean_filter=false, + transform=false, + span=1.0, + min_window_count=5, + ) catch { + _ => abort("valid trend inputs should be fitted") + } + let mut index = 0 + while index < trend.fitted.length() { + scrapper_test_assert_close( + trend.fitted[index], + (index + 1).to_double() * 2.0, + 1.0e-9, + ) + scrapper_test_assert_close(trend.residuals[index], 0.0, 1.0e-9) + index = index + 1 + } +} + +///| +test "scrapper: variance trend extrapolates below the mean filter" { + let trend = @src.scrapper_fit_variance_trend( + [0.05, 1.0, 2.0, 3.0], + [0.01, 1.0, 2.0, 3.0], + min_mean=1.0, + transform=false, + span=1.0, + min_window_count=3, + ) catch { + _ => abort("valid filtered trend should be fitted") + } + assert_true(trend.fitted[0] >= 0.0) + assert_true(trend.fitted[0] < trend.fitted[1]) +} + +///| +test "scrapper: variance trend validates fitting parameters" { + let mismatched = try { + ignore(@src.scrapper_fit_variance_trend([1.0], [1.0, 2.0])) + false + } catch { + ScrapperError(_) => true + } + let no_genes = try { + ignore( + @src.scrapper_fit_variance_trend( + [0.0], + [0.0], + mean_filter=true, + min_mean=1.0, + ), + ) + false + } catch { + ScrapperError(_) => true + } + assert_true(mismatched) + assert_true(no_genes) +} + +///| +test "scrapper: selects top HVGs with stable tie handling" { + let selected = @src.scrapper_choose_highly_variable_genes( + [0.5, 2.0, 2.0, -1.0, 1.0], + top=1, + keep_ties=true, + ) catch { + _ => abort("finite statistics should be ranked") + } + assert_eq(selected, [1, 2]) +} + +///| +test "scrapper: HVG selection supports smaller statistics and no bound" { + let selected = @src.scrapper_choose_highly_variable_genes( + [0.5, 2.0, 2.0, -1.0, 1.0], + top=2, + larger=false, + keep_ties=false, + use_bound=false, + ) catch { + _ => abort("finite statistics should be ranked") + } + assert_eq(selected, [3, 0]) +} + +///| +test "scrapper: variance model convenience method uses positive residuals" { + let model = @src.scrapper_model_gene_variances( + [[1.0, 2.0, 8.0], [2.0, 2.0, 2.0], [1.0, 2.0, 3.0]], + mean_filter=false, + transform=false, + span=1.0, + min_window_count=1, + ) catch { + _ => abort("valid expression should produce a variance model") + } + let selected = model.highly_variable_genes(top=3) catch { + _ => abort("finite residuals should be ranked") + } + let expected = @src.scrapper_choose_highly_variable_genes( + model.residuals, + top=3, + use_bound=true, + bound=0.0, + ) catch { + _ => abort("finite residuals should be ranked") + } + assert_eq(selected, expected) +} + +///| +test "scrapper: aggregates sums detection medians and means by one factor" { + let result = @src.scrapper_aggregate_across_cells( + [[1.0, 2.0, 3.0, 4.0], [0.0, 5.0, 0.0, 7.0]], + [@src.ScrapperFactor::new("cluster", ["A", "A", "B", "B"])], + compute_median=true, + ) catch { + _ => abort("valid grouping factor should aggregate expression") + } + let means = result.means() catch { + _ => abort("aggregate sums should support means") + } + assert_eq(result.factor_names, ["cluster"]) + assert_eq(result.group_names, ["cluster=A", "cluster=B"]) + assert_eq(result.counts, [2, 2]) + assert_eq(result.index, [0, 0, 1, 1]) + assert_eq(result.sums, [[3.0, 7.0], [5.0, 7.0]]) + assert_eq(result.detected, [[2, 2], [1, 1]]) + assert_eq(result.medians, [[1.5, 3.5], [2.5, 3.5]]) + assert_eq(means, [[1.5, 3.5], [2.5, 3.5]]) + assert_eq(result.n_groups(), 2) + assert_eq( + result.summary(), + "ScrapperAggregate(groups=2, cells=4, features=2)", + ) +} + +///| +test "scrapper: aggregates unique combinations of multiple factors" { + let result = @src.scrapper_aggregate_across_cells([[1.0, 2.0, 3.0, 4.0]], [ + @src.ScrapperFactor::new("cluster", ["A", "A", "B", "B"]), + @src.ScrapperFactor::new("batch", ["X", "Y", "X", "Y"]), + ]) catch { + _ => abort("valid grouping factors should aggregate expression") + } + assert_eq(result.factor_names, ["cluster", "batch"]) + assert_eq(result.n_groups(), 4) + assert_eq(result.group_names, [ + "cluster=A,batch=X", "cluster=A,batch=Y", "cluster=B,batch=X", "cluster=B,batch=Y", + ]) + assert_eq(result.index, [0, 1, 2, 3]) +} + +///| +test "scrapper: aggregation can omit optional summaries" { + let result = @src.scrapper_aggregate_across_cells( + [[1.0, 2.0]], + [@src.ScrapperFactor::new("group", ["A", "A"])], + compute_sum=false, + compute_detected=false, + compute_median=true, + ) catch { + _ => abort("median-only aggregation should be accepted") + } + assert_eq(result.sums.length(), 0) + assert_eq(result.detected.length(), 0) + assert_eq(result.medians, [[1.5]]) + let means_missing = try { + ignore(result.means()) + false + } catch { + ScrapperError(_) => true + } + assert_true(means_missing) +} + +///| +test "scrapper: aggregation validates grouping factors" { + let no_factors = try { + ignore(@src.scrapper_aggregate_across_cells([[1.0]], [])) + false + } catch { + ScrapperError(_) => true + } + let duplicate_names = try { + ignore( + @src.scrapper_aggregate_across_cells([[1.0]], [ + @src.ScrapperFactor::new("group", ["A"]), + @src.ScrapperFactor::new("group", ["B"]), + ]), + ) + false + } catch { + ScrapperError(_) => true + } + let wrong_length = try { + ignore( + @src.scrapper_aggregate_across_cells([[1.0, 2.0]], [ + @src.ScrapperFactor::new("group", ["A"]), + ]), + ) + false + } catch { + ScrapperError(_) => true + } + assert_true(no_factors) + assert_true(duplicate_names) + assert_true(wrong_length) +} + +///| +test "scrapper: aggregation validates a finite detection limit" { + let raised = try { + ignore( + @src.scrapper_aggregate_across_cells( + [[1.0]], + [@src.ScrapperFactor::new("group", ["A"])], + detection_limit=@double.not_a_number, + ), + ) + false + } catch { + ScrapperError(_) => true + } + assert_true(raised) +} + +///| +test "scrapper: SCE normalization adds assay and size factors immutably" { + let original = scrapper_test_sce() + let normalized = @src.scrapper_normalize_rna_counts_sce(original, size_factors=[ + 1.0, 1.0, 1.0, + ]) catch { + _ => abort("valid SCE counts should be normalized") + } + assert_eq(@src.sce_get_assay(original, "logcounts").length(), 0) + let logcounts = @src.sce_get_assay(normalized, "logcounts") + assert_eq(logcounts.length(), 3) + scrapper_test_assert_close(logcounts[0][0], 2.0, 1.0e-12) + assert_eq(@src.sce_get_col_data(normalized, "size_factor"), ["1", "1", "1"]) + assert_eq(@src.sce_get_col_data(original, "size_factor").length(), 0) + assert_eq(@src.sce_get_row_data(normalized, "symbol"), ["A", "B", "MT-C"]) + assert_eq(@src.sce_get_reduced_dim(normalized, "PCA").length(), 3) +} + +///| +test "scrapper: SCE normalization deep-copies assay matrices" { + let original = scrapper_test_sce() + let normalized = @src.scrapper_normalize_rna_counts_sce(original, size_factors=[ + 1.0, 1.0, 1.0, + ]) catch { + _ => abort("valid SCE counts should be normalized") + } + let copied_counts = @src.sce_get_assay(normalized, "counts") + copied_counts[0][0] = 999.0 + assert_eq(@src.sce_get_assay(original, "counts")[0][0], 3.0) +} + +///| +test "scrapper: SCE quick QC adds copied cell metadata" { + let original = scrapper_test_sce() + let (annotated, result) = @src.scrapper_quick_rna_qc_sce( + original, + subsets=[@src.ScrapperNamedSubset::new("mito", [2])], + blocks=["x", "x", "y"], + ) catch { + _ => abort("valid SCE counts should support quick RNA QC") + } + assert_eq(result.metrics.sums, [4.0, 9.0, 12.0]) + assert_eq(@src.sce_get_col_data(annotated, "scrapper_sum").length(), 3) + assert_eq(@src.sce_get_col_data(annotated, "scrapper_detected").length(), 3) + assert_eq(@src.sce_get_col_data(annotated, "scrapper_keep").length(), 3) + assert_eq( + @src.sce_get_col_data(annotated, "scrapper_subset_mito_proportion").length(), + 3, + ) + assert_eq(@src.sce_get_col_data(original, "scrapper_sum").length(), 0) + assert_eq(@src.sce_get_col_data(annotated, "batch"), ["x", "x", "y"]) +} diff --git a/test/moonbit/scuttle_test.mbt b/test/moonbit/scuttle_test.mbt new file mode 100644 index 00000000..e0d3b1b8 --- /dev/null +++ b/test/moonbit/scuttle_test.mbt @@ -0,0 +1,966 @@ +///| +fn scuttle_test_close( + left : Double, + right : Double, + tolerance : Double, +) -> Bool { + (left - right).abs() <= tolerance +} + +///| +fn scuttle_test_column_sums(matrix : Array[Array[Double]]) -> Array[Double] { + let cells = if matrix.length() > 0 { matrix[0].length() } else { 0 } + let output = Array::make(cells, 0.0) + for row in matrix { + for cell in 0.. Double { + let mut output = 0.0 + for row in matrix { + for value in row { + output = output + value + } + } + output +} + +///| +fn scuttle_test_outlier_value(value : Bool?) -> Bool { + match value { + Some(result) => result + None => abort("expected an estimated outlier status") + } +} + +///| +fn scuttle_test_counts() -> Array[Array[Double]] { + [ + [8.0, 4.0, 16.0, 8.0], + [2.0, 6.0, 4.0, 12.0], + [0.0, 2.0, 0.0, 4.0], + [10.0, 8.0, 20.0, 16.0], + ] +} + +///| +fn scuttle_test_sce() -> @src.SingleCellExperiment { + let counts = scuttle_test_counts() + let sce = @src.SingleCellExperiment::new(counts, ["G1", "G2", "G3", "G4"], [ + "C1", "C2", "C3", "C4", + ]) + sce.assays["logcounts"] = [ + [3.0, 2.0, 4.0, 3.0], + [1.0, 2.5, 2.0, 3.5], + [0.0, 1.0, 0.0, 2.0], + [3.5, 3.0, 4.5, 4.0], + ] + sce.row_data["symbol"] = ["A", "B", "C", "D"] + sce.col_data["batch"] = ["A", "A", "B", "B"] + sce.reduced_dims["PCA"] = [[1.0, 0.0], [2.0, 0.0], [3.0, 1.0], [4.0, 1.0]] + sce.metadata["project"] = "scuttle-test" + sce.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[1.0, 2.0, 3.0, 4.0]], + ["Spike1"], + ["C1", "C2", "C3", "C4"], + ) + sce +} + +///| +test "scuttle: default outlier configuration matches upstream defaults" { + let config = @src.ScuttleOutlierConfig::default() + assert_eq(config.nmads, 3.0) + assert_true(config.direction == @src.scuttle_outlier_both()) + assert_false(config.log_transform) + assert_false(config.share_medians) + assert_false(config.share_mads) + assert_true(config.share_missing) + assert_true(config.subset is None) + assert_true(config.min_diff is None) +} + +///| +test "scuttle: outlier configuration rejects negative nmads" { + let mut raised = false + ignore(@src.ScuttleOutlierConfig::create(nmads=-1.0)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: outlier configuration rejects negative min diff" { + let mut raised = false + ignore(@src.ScuttleOutlierConfig::create(min_diff=Some(-1.0))) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: vanilla outlier thresholds use R MAD scaling" { + let result = @src.scuttle_is_outlier([1.0, 2.0, 2.0, 3.0, 100.0]) catch { + _ => abort("valid outlier detection should work") + } + assert_eq(result.batch_names, ["1"]) + assert_true(scuttle_test_close(result.medians[0], 2.0, 1.0e-12)) + assert_true(scuttle_test_close(result.mads[0], 1.4826, 1.0e-12)) + assert_true(scuttle_test_close(result.lower[0], -2.4478, 1.0e-10)) + assert_true(scuttle_test_close(result.higher[0], 6.4478, 1.0e-10)) + assert_false(scuttle_test_outlier_value(result.outliers[0])) + assert_true(scuttle_test_outlier_value(result.outliers[4])) + assert_eq(result.discarded_count(), 1) +} + +///| +test "scuttle: lower-tail direction only calls small observations" { + let config = @src.ScuttleOutlierConfig::create( + nmads=1.0, + direction=@src.scuttle_outlier_lower(), + ) catch { + _ => abort("valid lower-tail configuration should build") + } + let result = @src.scuttle_is_outlier([0.0, 10.0, 10.0, 11.0, 100.0], config~) catch { + _ => abort("lower-tail outlier detection should work") + } + assert_true(scuttle_test_outlier_value(result.outliers[0])) + assert_false(scuttle_test_outlier_value(result.outliers[4])) + assert_true(result.higher[0] >= 1.0e299) +} + +///| +test "scuttle: higher-tail direction only calls large observations" { + let config = @src.ScuttleOutlierConfig::create( + nmads=1.0, + direction=@src.scuttle_outlier_higher(), + ) catch { + _ => abort("valid higher-tail configuration should build") + } + let result = @src.scuttle_is_outlier([0.0, 10.0, 10.0, 11.0, 100.0], config~) catch { + _ => abort("higher-tail outlier detection should work") + } + assert_false(scuttle_test_outlier_value(result.outliers[0])) + assert_true(scuttle_test_outlier_value(result.outliers[4])) + assert_true(result.lower[0] <= -1.0e299) +} + +///| +test "scuttle: nmads changes the outlier thresholds" { + let config = @src.ScuttleOutlierConfig::create(nmads=5.0) catch { + _ => abort("valid nmads configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 2.0, 3.0, 100.0], config~) catch { + _ => abort("outlier detection should work") + } + assert_true(scuttle_test_close(result.lower[0], -5.413, 1.0e-10)) + assert_true(scuttle_test_close(result.higher[0], 9.413, 1.0e-10)) +} + +///| +test "scuttle: min diff dominates a small MAD" { + let config = @src.ScuttleOutlierConfig::create(nmads=1.0, min_diff=Some(10.0)) catch { + _ => abort("valid min-diff configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 2.0, 3.0, 20.0], config~) catch { + _ => abort("outlier detection should work") + } + assert_eq(result.lower[0], -8.0) + assert_eq(result.higher[0], 12.0) + assert_true(scuttle_test_outlier_value(result.outliers[4])) +} + +///| +test "scuttle: threshold subset is applied to all observations" { + let config = @src.ScuttleOutlierConfig::create( + nmads=1.0, + subset=Some([0, 1, 2]), + ) catch { + _ => abort("valid subset configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 3.0, 50.0], config~) catch { + _ => abort("subset outlier detection should work") + } + assert_eq(result.medians[0], 2.0) + assert_true(scuttle_test_outlier_value(result.outliers[3])) +} + +///| +test "scuttle: explicit empty threshold subset yields missing statuses" { + let config = @src.ScuttleOutlierConfig::create(subset=Some([])) catch { + _ => abort("empty subset is a valid explicit subset") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 3.0], config~) catch { + _ => abort("empty-subset detection should return missing statuses") + } + assert_false(result.threshold_valid[0]) + assert_true(result.outliers[0] is None) + assert_true(result.outliers[2] is None) +} + +///| +test "scuttle: threshold subset indices are validated" { + let config = @src.ScuttleOutlierConfig::create(subset=Some([3])) catch { + _ => abort("subset bounds are checked when data are available") + } + let mut raised = false + ignore(@src.scuttle_is_outlier([1.0, 2.0, 3.0], config~)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: batch-specific thresholds are independent" { + let config = @src.ScuttleOutlierConfig::create(nmads=1.0, batches=[ + "A", "A", "A", "B", "B", "B", + ]) catch { + _ => abort("valid batch configuration should build") + } + let result = @src.scuttle_is_outlier( + [1.0, 2.0, 3.0, 100.0, 101.0, 102.0], + config~, + ) catch { + _ => abort("batch-aware detection should work") + } + assert_eq(result.batch_names, ["A", "B"]) + assert_eq(result.medians, [2.0, 101.0]) + assert_false(scuttle_test_outlier_value(result.outliers[0])) + assert_false(scuttle_test_outlier_value(result.outliers[5])) +} + +///| +test "scuttle: batch names use sorted factor order" { + let config = @src.ScuttleOutlierConfig::create(batches=["z", "a", "m", "z"]) catch { + _ => abort("valid batches should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 3.0, 4.0], config~) catch { + _ => abort("batch-aware detection should work") + } + assert_eq(result.batch_names, ["a", "m", "z"]) +} + +///| +test "scuttle: shared medians and MADs reproduce global thresholds" { + let config = @src.ScuttleOutlierConfig::create( + batches=["A", "A", "B", "B", "B"], + share_medians=true, + share_mads=true, + ) catch { + _ => abort("valid sharing configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 3.0, 4.0, 100.0], config~) catch { + _ => abort("shared outlier detection should work") + } + let global = @src.scuttle_is_outlier([1.0, 2.0, 3.0, 4.0, 100.0]) catch { + _ => abort("global outlier detection should work") + } + assert_eq(result.medians, [global.medians[0], global.medians[0]]) + assert_eq(result.mads, [global.mads[0], global.mads[0]]) + assert_eq(result.lower, [global.lower[0], global.lower[0]]) +} + +///| +test "scuttle: shared median retains batch-specific MADs" { + let config = @src.ScuttleOutlierConfig::create( + nmads=1.0, + batches=["A", "A", "A", "B", "B", "B"], + share_medians=true, + ) catch { + _ => abort("valid median-sharing configuration should build") + } + let result = @src.scuttle_is_outlier( + [1.0, 2.0, 3.0, 10.0, 20.0, 30.0], + config~, + ) catch { + _ => abort("shared-median detection should work") + } + assert_eq(result.medians[0], result.medians[1]) + assert_true(result.mads[1] > result.mads[0]) +} + +///| +test "scuttle: shared MAD retains batch-specific medians" { + let config = @src.ScuttleOutlierConfig::create( + batches=["A", "A", "A", "B", "B", "B"], + share_mads=true, + ) catch { + _ => abort("valid MAD-sharing configuration should build") + } + let result = @src.scuttle_is_outlier( + [1.0, 2.0, 3.0, 100.0, 101.0, 102.0], + config~, + ) catch { + _ => abort("shared-MAD detection should work") + } + assert_eq(result.medians, [2.0, 101.0]) + assert_eq(result.mads[0], result.mads[1]) +} + +///| +test "scuttle: missing batch thresholds are recovered by default" { + let config = @src.ScuttleOutlierConfig::create( + batches=["A", "A", "B", "B"], + subset=Some([0, 1]), + ) catch { + _ => abort("valid missing-batch configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 3.0, 100.0, 200.0], config~) catch { + _ => abort("missing batch should borrow thresholds") + } + assert_true(result.threshold_valid[0]) + assert_true(result.threshold_valid[1]) + assert_eq(result.lower[0], result.lower[1]) + assert_eq(result.higher[0], result.higher[1]) +} + +///| +test "scuttle: missing batch statuses can remain missing" { + let config = @src.ScuttleOutlierConfig::create( + batches=["A", "A", "B", "B"], + subset=Some([0, 1]), + share_missing=false, + ) catch { + _ => abort("valid no-sharing configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 3.0, 100.0, 200.0], config~) catch { + _ => abort("missing batch should remain missing") + } + assert_true(result.threshold_valid[0]) + assert_false(result.threshold_valid[1]) + assert_true(result.outliers[2] is None) +} + +///| +test "scuttle: batch vector length is validated" { + let config = @src.ScuttleOutlierConfig::create(batches=["A"]) catch { + _ => abort("configuration construction does not know metric length") + } + let mut raised = false + ignore(@src.scuttle_is_outlier([1.0, 2.0], config~)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: log thresholds are returned on the original scale" { + let config = @src.ScuttleOutlierConfig::create(nmads=1.0, log_transform=true) catch { + _ => abort("valid log configuration should build") + } + let result = @src.scuttle_is_outlier([1.0, 2.0, 4.0, 8.0, 1024.0], config~) catch { + _ => abort("log outlier detection should work") + } + assert_true(scuttle_test_close(result.medians[0], 4.0, 1.0e-10)) + assert_true(result.lower[0] > 1.0) + assert_true(result.higher[0] < 16.0) + assert_true(scuttle_test_outlier_value(result.outliers[4])) +} + +///| +test "scuttle: zero can participate in log outlier detection" { + let config = @src.ScuttleOutlierConfig::create(log_transform=true) catch { + _ => abort("valid log configuration should build") + } + let result = @src.scuttle_is_outlier([0.0, 8.0, 8.0, 8.0, 8.0], config~) catch { + _ => abort("zero should be represented on the log scale") + } + assert_true(scuttle_test_outlier_value(result.outliers[0])) + assert_true(scuttle_test_close(result.medians[0], 8.0, 1.0e-10)) +} + +///| +test "scuttle: negative log metrics are rejected" { + let config = @src.ScuttleOutlierConfig::create(log_transform=true) catch { + _ => abort("valid log configuration should build") + } + let mut raised = false + ignore(@src.scuttle_is_outlier([-1.0, 2.0], config~)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: empty metric returns an empty filter" { + let result = @src.scuttle_is_outlier([]) catch { + _ => abort("empty metrics should be supported") + } + assert_eq(result.outliers.length(), 0) + assert_false(result.threshold_valid[0]) +} + +///| +test "scuttle: per-feature QC computes means and detection percentages" { + let metrics = @src.scuttle_per_feature_qc([ + [0.0, 2.0, 4.0, 6.0], + [1.0, 0.0, 0.0, 3.0], + ]) catch { + _ => abort("valid feature QC should work") + } + assert_eq(metrics.means, [3.0, 1.0]) + assert_eq(metrics.detected_percent, [75.0, 50.0]) +} + +///| +test "scuttle: feature detection limit is strict" { + let metrics = @src.scuttle_per_feature_qc( + [[0.0, 2.0, 4.0, 6.0]], + detection_limit=2.0, + ) catch { + _ => abort("feature QC with a limit should work") + } + assert_eq(metrics.detected_percent, [50.0]) +} + +///| +test "scuttle: per-feature QC computes subset statistics" { + let subset = @src.ScuttleNamedCellSubset::create("controls", [0, 1]) catch { + _ => abort("valid subset should build") + } + let metrics = @src.scuttle_per_feature_qc( + [[0.0, 2.0, 4.0, 6.0], [1.0, 0.0, 0.0, 3.0]], + subsets=[subset], + ) catch { + _ => abort("subset feature QC should work") + } + assert_eq(metrics.subset_names, ["controls"]) + assert_eq(metrics.subset_means[0], [1.0, 0.5]) + assert_eq(metrics.subset_detected_percent[0], [50.0, 50.0]) + assert_true( + scuttle_test_close(metrics.subset_ratios[0][0], 1.0 / 3.0, 1.0e-12), + ) +} + +///| +test "scuttle: per-feature QC supports multiple cell subsets" { + let first = @src.ScuttleNamedCellSubset::create("first", [0, 1]) catch { + _ => abort("valid subset should build") + } + let last = @src.ScuttleNamedCellSubset::create("last", [2, 3]) catch { + _ => abort("valid subset should build") + } + let metrics = @src.scuttle_per_feature_qc([[0.0, 2.0, 4.0, 6.0]], subsets=[ + first, last, + ]) catch { + _ => abort("multiple-subset feature QC should work") + } + assert_eq(metrics.subset_names, ["first", "last"]) + assert_eq(metrics.subset_means[0], [1.0]) + assert_eq(metrics.subset_means[1], [5.0]) +} + +///| +test "scuttle: zero feature mean has a finite zero subset ratio" { + let subset = @src.ScuttleNamedCellSubset::create("empty", []) catch { + _ => abort("empty named subset should build") + } + let metrics = @src.scuttle_per_feature_qc([[0.0, 0.0]], subsets=[subset]) catch { + _ => abort("zero feature QC should work") + } + assert_eq(metrics.subset_ratios[0], [0.0]) +} + +///| +test "scuttle: repeated cell subset indices retain R indexing weights" { + let subset = @src.ScuttleNamedCellSubset::create("weighted", [0, 0, 1]) catch { + _ => abort("repeated indices are valid") + } + let metrics = @src.scuttle_per_feature_qc([[3.0, 9.0]], subsets=[subset]) catch { + _ => abort("weighted subset QC should work") + } + assert_eq(metrics.subset_means[0], [5.0]) +} + +///| +test "scuttle: duplicate cell subset names are rejected" { + let first = @src.ScuttleNamedCellSubset::create("same", [0]) catch { + _ => abort("valid subset should build") + } + let second = @src.ScuttleNamedCellSubset::create("same", [1]) catch { + _ => abort("valid subset should build") + } + let mut raised = false + ignore(@src.scuttle_per_feature_qc([[1.0, 2.0]], subsets=[first, second])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: cell subset bounds are validated" { + let subset = @src.ScuttleNamedCellSubset::create("bad", [2]) catch { + _ => abort("bounds are checked with data") + } + let mut raised = false + ignore(@src.scuttle_per_feature_qc([[1.0, 2.0]], subsets=[subset])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: per-feature QC rejects ragged matrices" { + let mut raised = false + ignore(@src.scuttle_per_feature_qc([[1.0, 2.0], [3.0]])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: per-feature QC rejects negative counts" { + let mut raised = false + ignore(@src.scuttle_per_feature_qc([[1.0, -1.0]])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: arbitrary feature sets are summed per cell" { + let first = @src.ScuttleFeatureSet::create("first", [0, 2]) catch { + _ => abort("valid feature set should build") + } + let second = @src.ScuttleFeatureSet::create("second", [1]) catch { + _ => abort("valid feature set should build") + } + let result = @src.scuttle_aggregate_feature_sets( + [[1.0, 2.0], [10.0, 20.0], [3.0, 4.0]], + [first, second], + ) catch { + _ => abort("feature aggregation should work") + } + assert_eq(result.names, ["first", "second"]) + assert_eq(result.sizes, [2, 1]) + assert_eq(result.values, [[4.0, 6.0], [10.0, 20.0]]) +} + +///| +test "scuttle: overlapping feature sets are independent" { + let left = @src.ScuttleFeatureSet::create("left", [0, 1]) catch { + _ => abort("valid feature set should build") + } + let right = @src.ScuttleFeatureSet::create("right", [1, 2]) catch { + _ => abort("valid feature set should build") + } + let result = @src.scuttle_aggregate_feature_sets([[1.0], [2.0], [4.0]], [ + left, right, + ]) catch { + _ => abort("overlapping aggregation should work") + } + assert_eq(result.values, [[3.0], [6.0]]) +} + +///| +test "scuttle: feature-set averages divide by set size" { + let set = @src.ScuttleFeatureSet::create("mean", [0, 1]) catch { + _ => abort("valid feature set should build") + } + let result = @src.scuttle_aggregate_feature_sets( + [[2.0, 4.0], [4.0, 8.0]], + [set], + average=true, + ) catch { + _ => abort("average aggregation should work") + } + assert_eq(result.values, [[3.0, 6.0]]) +} + +///| +test "scuttle: detected features can be aggregated across sets" { + let set = @src.ScuttleFeatureSet::create("detected", [0, 1, 2]) catch { + _ => abort("valid feature set should build") + } + let result = @src.scuttle_aggregate_feature_sets( + [[0.0, 2.0], [3.0, 1.0], [4.0, 0.0]], + [set], + detection_limit=Some(1.0), + ) catch { + _ => abort("detection aggregation should work") + } + assert_eq(result.values, [[2.0, 1.0]]) +} + +///| +test "scuttle: feature identifiers are aggregated in factor order" { + let result = @src.scuttle_aggregate_features_by_ids( + [[1.0], [2.0], [4.0], [8.0]], + ["z", "a", "z", "m"], + ) catch { + _ => abort("identifier aggregation should work") + } + assert_eq(result.names, ["a", "m", "z"]) + assert_eq(result.values, [[2.0], [8.0], [5.0]]) +} + +///| +test "scuttle: empty feature identifiers are ignored" { + let result = @src.scuttle_aggregate_features_by_ids([[1.0], [2.0], [4.0]], [ + "A", "", "A", + ]) catch { + _ => abort("empty identifiers should be ignored") + } + assert_eq(result.names, ["A"]) + assert_eq(result.values, [[5.0]]) +} + +///| +test "scuttle: feature identifier length is validated" { + let mut raised = false + ignore(@src.scuttle_aggregate_features_by_ids([[1.0], [2.0]], ["A"])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: duplicate feature-set names are rejected" { + let first = @src.ScuttleFeatureSet::create("same", [0]) catch { + _ => abort("valid feature set should build") + } + let second = @src.ScuttleFeatureSet::create("same", [1]) catch { + _ => abort("valid feature set should build") + } + let mut raised = false + ignore(@src.scuttle_aggregate_feature_sets([[1.0], [2.0]], [first, second])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: feature-set bounds are validated" { + let set = @src.ScuttleFeatureSet::create("bad", [2]) catch { + _ => abort("bounds are checked with data") + } + let mut raised = false + ignore(@src.scuttle_aggregate_feature_sets([[1.0], [2.0]], [set])) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: per-column downsampling has exact rounded totals" { + let counts = scuttle_test_counts() + let output = @src.scuttle_downsample_matrix(counts, 0.5, seed=42) catch { + _ => abort("downsampling should work") + } + assert_eq(scuttle_test_column_sums(output), [10.0, 10.0, 20.0, 20.0]) +} + +///| +test "scuttle: downsampled entries never exceed sanitized inputs" { + let counts = scuttle_test_counts() + let output = @src.scuttle_downsample_matrix(counts, 0.75, seed=11) catch { + _ => abort("downsampling should work") + } + for feature in 0..= 0.0) + } + } +} + +///| +test "scuttle: per-cell proportions can differ" { + let counts = scuttle_test_counts() + let output = @src.scuttle_downsample_columns( + counts, + [0.0, 0.25, 0.5, 1.0], + seed=7, + ) catch { + _ => abort("vector downsampling should work") + } + assert_eq(scuttle_test_column_sums(output), [0.0, 5.0, 20.0, 40.0]) +} + +///| +test "scuttle: global downsampling has an exact matrix total" { + let counts = scuttle_test_counts() + let output = @src.scuttle_downsample_matrix( + counts, + 0.3, + by_column=false, + seed=9, + ) catch { + _ => abort("global downsampling should work") + } + assert_eq(scuttle_test_total(output), 36.0) +} + +///| +test "scuttle: downsampling rounds and clamps numeric inputs" { + let output = @src.scuttle_downsample_matrix([[-2.0, 1.4], [3.6, 2.6]], 1.0) catch { + _ => abort("input sanitization should work") + } + assert_eq(output, [[0.0, 1.0], [4.0, 3.0]]) +} + +///| +test "scuttle: equal seeds produce identical samples" { + let counts = scuttle_test_counts() + let first = @src.scuttle_downsample_matrix(counts, 0.45, seed=123) catch { + _ => abort("first downsampling should work") + } + let second = @src.scuttle_downsample_matrix(counts, 0.45, seed=123) catch { + _ => abort("second downsampling should work") + } + assert_eq(first, second) +} + +///| +test "scuttle: different seeds change nontrivial samples" { + let counts = scuttle_test_counts() + let first = @src.scuttle_downsample_matrix(counts, 0.45, seed=1) catch { + _ => abort("first downsampling should work") + } + let second = @src.scuttle_downsample_matrix(counts, 0.45, seed=2) catch { + _ => abort("second downsampling should work") + } + assert_true(first != second) +} + +///| +test "scuttle: zero and unit proportions are exact boundaries" { + let counts = scuttle_test_counts() + let zero = @src.scuttle_downsample_matrix(counts, 0.0) catch { + _ => abort("zero downsampling should work") + } + let one = @src.scuttle_downsample_matrix(counts, 1.0) catch { + _ => abort("unit downsampling should work") + } + assert_eq(scuttle_test_total(zero), 0.0) + assert_eq(one, counts) +} + +///| +test "scuttle: invalid downsampling proportions are rejected" { + let mut raised = false + ignore(@src.scuttle_downsample_matrix([[1.0]], 1.1)) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: downsampling validates vector length and seed" { + let mut bad_length = false + ignore(@src.scuttle_downsample_columns([[1.0, 2.0]], [0.5])) catch { + _ => bad_length = true + } + assert_true(bad_length) + let mut bad_seed = false + ignore(@src.scuttle_downsample_matrix([[1.0]], 0.5, seed=0)) catch { + _ => bad_seed = true + } + assert_true(bad_seed) +} + +///| +test "scuttle: batch downsampling targets the shallowest median" { + let counts = [[5.0, 5.0, 10.0, 10.0], [5.0, 5.0, 10.0, 10.0]] + let result = @src.scuttle_downsample_batches( + counts, + ["A", "A", "B", "B"], + seed=4, + ) catch { + _ => abort("batch downsampling should work") + } + assert_eq(result.group_names, ["A", "B"]) + assert_eq(result.summaries, [10.0, 20.0]) + assert_eq(result.proportions, [1.0, 0.5]) + assert_eq(scuttle_test_column_sums(result.matrix), [10.0, 10.0, 10.0, 10.0]) +} + +///| +test "scuttle: global batch downsampling equalizes group totals" { + let counts = [[5.0, 5.0, 10.0, 10.0], [5.0, 5.0, 10.0, 10.0]] + let result = @src.scuttle_downsample_batches( + counts, + ["A", "A", "B", "B"], + by_column=false, + seed=4, + ) catch { + _ => abort("global batch downsampling should work") + } + let sums = scuttle_test_column_sums(result.matrix) + assert_eq(sums[0] + sums[1], 20.0) + assert_eq(sums[2] + sums[3], 20.0) +} + +///| +test "scuttle: blocking equalizes batches only within each block" { + let counts = [[10.0, 20.0, 100.0, 200.0]] + let result = @src.scuttle_downsample_batches( + counts, + ["A", "B", "A", "B"], + blocks=["sample1", "sample1", "sample2", "sample2"], + seed=8, + ) catch { + _ => abort("blocked batch downsampling should work") + } + assert_eq(result.group_names, [ + "sample1|A", "sample1|B", "sample2|A", "sample2|B", + ]) + assert_eq(result.proportions, [1.0, 0.5, 1.0, 0.5]) + assert_eq(scuttle_test_column_sums(result.matrix), [10.0, 10.0, 100.0, 100.0]) +} + +///| +test "scuttle: batch mean summary is selectable" { + let counts = [[2.0, 10.0, 12.0, 24.0]] + let result = @src.scuttle_downsample_batches( + counts, + ["A", "A", "B", "B"], + summary_method=@src.scuttle_batch_mean(), + seed=10, + ) catch { + _ => abort("mean batch downsampling should work") + } + assert_eq(result.summaries, [6.0, 18.0]) + assert_true(scuttle_test_close(result.proportions[1], 1.0 / 3.0, 1.0e-12)) +} + +///| +test "scuttle: geometric batch summary uses a pseudocount" { + let counts = [[0.0, 3.0, 8.0, 8.0]] + let result = @src.scuttle_downsample_batches( + counts, + ["A", "A", "B", "B"], + summary_method=@src.scuttle_batch_geometric_mean(), + seed=10, + ) catch { + _ => abort("geometric batch downsampling should work") + } + assert_eq(result.summaries[0], 2.0) + assert_true(scuttle_test_close(result.summaries[1], 9.0, 1.0e-10)) +} + +///| +test "scuttle: batch downsampling validates annotations" { + let mut bad_batch = false + ignore(@src.scuttle_downsample_batches([[1.0, 2.0]], ["A"])) catch { + _ => bad_batch = true + } + assert_true(bad_batch) + let mut bad_block = false + ignore( + @src.scuttle_downsample_batches([[1.0, 2.0]], ["A", "B"], blocks=["one"]), + ) catch { + _ => bad_block = true + } + assert_true(bad_block) +} + +///| +test "scuttle: feature QC SCE wrapper writes row metadata immutably" { + let sce = scuttle_test_sce() + let controls = @src.ScuttleNamedCellSubset::create("controls", [0, 1]) catch { + _ => abort("valid subset should build") + } + let output = @src.scuttle_per_feature_qc_sce(sce, subsets=[controls]) catch { + _ => abort("SCE feature QC should work") + } + assert_eq(@src.sce_get_row_data(sce, "scuttle.mean").length(), 0) + assert_eq( + @src.sce_get_row_data(output.experiment, "scuttle.mean").length(), + 4, + ) + assert_eq( + @src.sce_get_row_data(output.experiment, "scuttle.subsets.controls.ratio").length(), + 4, + ) + output.experiment.assays["counts"][0][0] = 999.0 + assert_eq(sce.assays["counts"][0][0], 8.0) +} + +///| +test "scuttle: feature QC SCE wrapper validates assay names" { + let sce = scuttle_test_sce() + let mut raised = false + ignore(@src.scuttle_per_feature_qc_sce(sce, assay_name="missing")) catch { + _ => raised = true + } + assert_true(raised) +} + +///| +test "scuttle: SCE feature aggregation handles multiple assays" { + let sce = scuttle_test_sce() + let first = @src.ScuttleFeatureSet::create("PathwayA", [0, 1]) catch { + _ => abort("valid feature set should build") + } + let second = @src.ScuttleFeatureSet::create("PathwayB", [2, 3]) catch { + _ => abort("valid feature set should build") + } + let output = @src.scuttle_aggregate_feature_sets_sce(sce, [first, second], assay_names=[ + "counts", "logcounts", + ]) catch { + _ => abort("SCE feature aggregation should work") + } + assert_eq(output.row_names, ["PathwayA", "PathwayB"]) + assert_eq(output.assays["counts"], [ + [10.0, 10.0, 20.0, 20.0], + [10.0, 10.0, 20.0, 20.0], + ]) + assert_eq(output.assays["logcounts"].length(), 2) + assert_eq(output.col_data["batch"], ["A", "A", "B", "B"]) + assert_eq(output.reduced_dims["PCA"].length(), 4) + assert_eq(output.alternative_experiments["spike"].row_names, ["Spike1"]) + assert_eq(output.metadata["project"], "scuttle-test") +} + +///| +test "scuttle: SCE feature aggregation discards gene metadata and deep copies" { + let sce = scuttle_test_sce() + let set = @src.ScuttleFeatureSet::create("all", [0, 1, 2, 3]) catch { + _ => abort("valid feature set should build") + } + let output = @src.scuttle_aggregate_feature_sets_sce(sce, [set]) catch { + _ => abort("SCE feature aggregation should work") + } + assert_eq(output.row_data.keys().collect().length(), 0) + output.col_data["batch"][0] = "changed" + output.reduced_dims["PCA"][0][0] = 999.0 + assert_eq(sce.col_data["batch"][0], "A") + assert_eq(sce.reduced_dims["PCA"][0][0], 1.0) +} + +///| +test "scuttle: SCE downsampling adds an assay without mutating input" { + let sce = scuttle_test_sce() + let output = @src.scuttle_downsample_sce( + sce, + 0.5, + output_assay="half", + seed=42, + ) catch { + _ => abort("SCE downsampling should work") + } + assert_eq(@src.sce_get_assay(sce, "half").length(), 0) + assert_eq(scuttle_test_column_sums(@src.sce_get_assay(output, "half")), [ + 10.0, 10.0, 20.0, 20.0, + ]) + assert_eq(output.metadata["scuttle.downsample.source"], "counts") + output.assays["counts"][0][0] = 999.0 + assert_eq(sce.assays["counts"][0][0], 8.0) + assert_eq(output.alternative_experiments["spike"].row_names, ["Spike1"]) +} + +///| +test "scuttle: SCE downsampling validates source assay" { + let sce = scuttle_test_sce() + let mut raised = false + ignore(@src.scuttle_downsample_sce(sce, 0.5, assay_name="missing")) catch { + _ => raised = true + } + assert_true(raised) +} diff --git a/test/moonbit/searchio_new_test.mbt b/test/moonbit/searchio_new_test.mbt index 3f488995..43a13c89 100644 --- a/test/moonbit/searchio_new_test.mbt +++ b/test/moonbit/searchio_new_test.mbt @@ -20,9 +20,16 @@ test "SearchIOHsp::new construction" { ///| test "SearchIOHsp::n_identical" { let hsp = @src.SearchIOHsp::new( - bitscore=200.5, evalue=1.0e-50, identity=95.0, positives=97.0, - gap=2.0, alignment_length=100, - query_start=1, query_end=100, hit_start=1, hit_end=100, + bitscore=200.5, + evalue=1.0e-50, + identity=95.0, + positives=97.0, + gap=2.0, + alignment_length=100, + query_start=1, + query_end=100, + hit_start=1, + hit_end=100, ) assert_eq(hsp.n_identical(), 95) } @@ -30,9 +37,16 @@ test "SearchIOHsp::n_identical" { ///| test "SearchIOHsp::n_positives" { let hsp = @src.SearchIOHsp::new( - bitscore=200.5, evalue=1.0e-50, identity=95.0, positives=97.0, - gap=2.0, alignment_length=100, - query_start=1, query_end=100, hit_start=1, hit_end=100, + bitscore=200.5, + evalue=1.0e-50, + identity=95.0, + positives=97.0, + gap=2.0, + alignment_length=100, + query_start=1, + query_end=100, + hit_start=1, + hit_end=100, ) assert_eq(hsp.n_positives(), 97) } @@ -40,9 +54,16 @@ test "SearchIOHsp::n_positives" { ///| test "SearchIOHsp::n_gaps" { let hsp = @src.SearchIOHsp::new( - bitscore=200.5, evalue=1.0e-50, identity=95.0, positives=97.0, - gap=2.0, alignment_length=100, - query_start=1, query_end=100, hit_start=1, hit_end=100, + bitscore=200.5, + evalue=1.0e-50, + identity=95.0, + positives=97.0, + gap=2.0, + alignment_length=100, + query_start=1, + query_end=100, + hit_start=1, + hit_end=100, ) assert_eq(hsp.n_gaps(), 2) } @@ -50,9 +71,16 @@ test "SearchIOHsp::n_gaps" { ///| test "SearchIOHsp::with_query_seq" { let hsp = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = hsp.with_query_seq("ATCGATCG") assert_eq(hsp2.query_seq, "ATCGATCG") @@ -61,9 +89,16 @@ test "SearchIOHsp::with_query_seq" { ///| test "SearchIOHsp::with_hit_seq" { let hsp = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = hsp.with_hit_seq("ATCGATCG") assert_eq(hsp2.hit_seq, "ATCGATCG") @@ -72,9 +107,16 @@ test "SearchIOHsp::with_hit_seq" { ///| test "SearchIOHsp::with_midline" { let hsp = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = hsp.with_midline("||||||..") assert_eq(hsp2.midline, "||||||..") @@ -83,9 +125,16 @@ test "SearchIOHsp::with_midline" { ///| test "SearchIOHsp::with_frames" { let hsp = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = hsp.with_frames(1, -1) assert_eq(hsp2.query_frame, 1) @@ -94,7 +143,11 @@ test "SearchIOHsp::with_frames" { ///| test "SearchIOHit::new construction" { - let hit = @src.SearchIOHit::new(id="hit1", description="Test hit", seq_length=1000) + let hit = @src.SearchIOHit::new( + id="hit1", + description="Test hit", + seq_length=1000, + ) assert_eq(hit.id, "hit1") assert_eq(hit.description, "Test hit") assert_eq(hit.seq_length, 1000) @@ -104,14 +157,28 @@ test "SearchIOHit::new construction" { ///| test "SearchIOHit::add_hsp and best_hsp" { let hsp1 = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = @src.SearchIOHsp::new( - bitscore=200.0, evalue=1.0e-40, identity=95.0, positives=97.0, - gap=0.5, alignment_length=80, - query_start=51, query_end=130, hit_start=51, hit_end=130, + bitscore=200.0, + evalue=1.0e-40, + identity=95.0, + positives=97.0, + gap=0.5, + alignment_length=80, + query_start=51, + query_end=130, + hit_start=51, + hit_end=130, ) let hit = @src.SearchIOHit::new(id="hit1", description="Test", seq_length=500) let hit_with_hsps = hit.add_hsp(hsp1).add_hsp(hsp2) @@ -137,14 +204,28 @@ test "SearchIOHit::best_hsp empty" { ///| test "SearchIOHit::sum_bitscore" { let hsp1 = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = @src.SearchIOHsp::new( - bitscore=200.0, evalue=1.0e-40, identity=95.0, positives=97.0, - gap=0.5, alignment_length=80, - query_start=51, query_end=130, hit_start=51, hit_end=130, + bitscore=200.0, + evalue=1.0e-40, + identity=95.0, + positives=97.0, + gap=0.5, + alignment_length=80, + query_start=51, + query_end=130, + hit_start=51, + hit_end=130, ) let hit = @src.SearchIOHit::new(id="hit1", description="Test", seq_length=500) let hit_with_hsps = hit.add_hsp(hsp1).add_hsp(hsp2) @@ -155,7 +236,10 @@ test "SearchIOHit::sum_bitscore" { ///| test "SearchIOQueryResult::new construction" { let qr = @src.SearchIOQueryResult::new( - id="query1", description="Test query", seq_length=500, database="nr", + id="query1", + description="Test query", + seq_length=500, + database="nr", ) assert_eq(qr.id, "query1") assert_eq(qr.description, "Test query") @@ -166,9 +250,16 @@ test "SearchIOQueryResult::new construction" { ///| test "SearchIOQueryResult::add_hit" { let qr = @src.SearchIOQueryResult::new( - id="query1", description="Test", seq_length=500, database="nr", + id="query1", + description="Test", + seq_length=500, + database="nr", + ) + let hit = @src.SearchIOHit::new( + id="hit1", + description="Hit 1", + seq_length=300, ) - let hit = @src.SearchIOHit::new(id="hit1", description="Hit 1", seq_length=300) let qr2 = qr.add_hit(hit) assert_eq(qr2.n_hits, 1) } @@ -176,20 +267,47 @@ test "SearchIOQueryResult::add_hit" { ///| test "SearchIOQueryResult::sort_by_score" { let hsp1 = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = @src.SearchIOHsp::new( - bitscore=200.0, evalue=1.0e-40, identity=95.0, positives=97.0, - gap=0.5, alignment_length=80, - query_start=1, query_end=80, hit_start=1, hit_end=80, + bitscore=200.0, + evalue=1.0e-40, + identity=95.0, + positives=97.0, + gap=0.5, + alignment_length=80, + query_start=1, + query_end=80, + hit_start=1, + hit_end=80, ) - let hit1 = @src.SearchIOHit::new(id="low", description="Low score", seq_length=300).add_hsp(hsp1) - let hit2 = @src.SearchIOHit::new(id="high", description="High score", seq_length=400).add_hsp(hsp2) + let hit1 = @src.SearchIOHit::new( + id="low", + description="Low score", + seq_length=300, + ).add_hsp(hsp1) + let hit2 = @src.SearchIOHit::new( + id="high", + description="High score", + seq_length=400, + ).add_hsp(hsp2) let qr = @src.SearchIOQueryResult::new( - id="query1", description="Test", seq_length=500, database="nr", - ).add_hit(hit1).add_hit(hit2) + id="query1", + description="Test", + seq_length=500, + database="nr", + ) + .add_hit(hit1) + .add_hit(hit2) let sorted = qr.sort_by_score() assert_eq(sorted.hits[0].id, "high") assert_eq(sorted.hits[1].id, "low") @@ -198,20 +316,43 @@ test "SearchIOQueryResult::sort_by_score" { ///| test "SearchIOQueryResult::filter_evalue" { let hsp1 = @src.SearchIOHsp::new( - bitscore=100.0, evalue=0.05, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=0.05, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = @src.SearchIOHsp::new( - bitscore=200.0, evalue=1.0e-40, identity=95.0, positives=97.0, - gap=0.5, alignment_length=80, - query_start=1, query_end=80, hit_start=1, hit_end=80, + bitscore=200.0, + evalue=1.0e-40, + identity=95.0, + positives=97.0, + gap=0.5, + alignment_length=80, + query_start=1, + query_end=80, + hit_start=1, + hit_end=80, + ) + let hit1 = @src.SearchIOHit::new(id="hit1", description="", seq_length=300).add_hsp( + hsp1, + ) + let hit2 = @src.SearchIOHit::new(id="hit2", description="", seq_length=400).add_hsp( + hsp2, ) - let hit1 = @src.SearchIOHit::new(id="hit1", description="", seq_length=300).add_hsp(hsp1) - let hit2 = @src.SearchIOHit::new(id="hit2", description="", seq_length=400).add_hsp(hsp2) let qr = @src.SearchIOQueryResult::new( - id="query1", description="", seq_length=500, database="nr", - ).add_hit(hit1).add_hit(hit2) + id="query1", + description="", + seq_length=500, + database="nr", + ) + .add_hit(hit1) + .add_hit(hit2) let filtered = qr.filter_evalue(0.01) assert_eq(filtered.n_hits, 1) assert_eq(filtered.hits[0].id, "hit2") @@ -220,26 +361,59 @@ test "SearchIOQueryResult::filter_evalue" { ///| test "SearchIOQueryResult::top_n" { let hsp1 = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, ) let hsp2 = @src.SearchIOHsp::new( - bitscore=200.0, evalue=1.0e-40, identity=95.0, positives=97.0, - gap=0.5, alignment_length=80, - query_start=1, query_end=80, hit_start=1, hit_end=80, + bitscore=200.0, + evalue=1.0e-40, + identity=95.0, + positives=97.0, + gap=0.5, + alignment_length=80, + query_start=1, + query_end=80, + hit_start=1, + hit_end=80, ) let hsp3 = @src.SearchIOHsp::new( - bitscore=150.0, evalue=1.0e-30, identity=92.0, positives=94.0, - gap=0.8, alignment_length=60, - query_start=1, query_end=60, hit_start=1, hit_end=60, + bitscore=150.0, + evalue=1.0e-30, + identity=92.0, + positives=94.0, + gap=0.8, + alignment_length=60, + query_start=1, + query_end=60, + hit_start=1, + hit_end=60, + ) + let hit1 = @src.SearchIOHit::new(id="low", description="", seq_length=300).add_hsp( + hsp1, + ) + let hit2 = @src.SearchIOHit::new(id="high", description="", seq_length=400).add_hsp( + hsp2, + ) + let hit3 = @src.SearchIOHit::new(id="mid", description="", seq_length=350).add_hsp( + hsp3, ) - let hit1 = @src.SearchIOHit::new(id="low", description="", seq_length=300).add_hsp(hsp1) - let hit2 = @src.SearchIOHit::new(id="high", description="", seq_length=400).add_hsp(hsp2) - let hit3 = @src.SearchIOHit::new(id="mid", description="", seq_length=350).add_hsp(hsp3) let qr = @src.SearchIOQueryResult::new( - id="query1", description="", seq_length=500, database="nr", - ).add_hit(hit1).add_hit(hit2).add_hit(hit3) + id="query1", + description="", + seq_length=500, + database="nr", + ) + .add_hit(hit1) + .add_hit(hit2) + .add_hit(hit3) let tops = qr.top_n(2) assert_eq(tops.length(), 2) assert_eq(tops[0].id, "high") @@ -249,13 +423,25 @@ test "SearchIOQueryResult::top_n" { ///| test "SearchIOQueryResult::top_n exceeds total" { let hsp = @src.SearchIOHsp::new( - bitscore=100.0, evalue=1.0e-20, identity=90.0, positives=92.0, - gap=1.0, alignment_length=50, - query_start=1, query_end=50, hit_start=1, hit_end=50, + bitscore=100.0, + evalue=1.0e-20, + identity=90.0, + positives=92.0, + gap=1.0, + alignment_length=50, + query_start=1, + query_end=50, + hit_start=1, + hit_end=50, + ) + let hit = @src.SearchIOHit::new(id="hit1", description="", seq_length=300).add_hsp( + hsp, ) - let hit = @src.SearchIOHit::new(id="hit1", description="", seq_length=300).add_hsp(hsp) let qr = @src.SearchIOQueryResult::new( - id="query1", description="", seq_length=500, database="nr", + id="query1", + description="", + seq_length=500, + database="nr", ).add_hit(hit) let tops = qr.top_n(5) assert_eq(tops.length(), 1) @@ -264,7 +450,10 @@ test "SearchIOQueryResult::top_n exceeds total" { ///| test "SearchIOQueryResult::with_total_hits" { let qr = @src.SearchIOQueryResult::new( - id="query1", description="", seq_length=500, database="nr", + id="query1", + description="", + seq_length=500, + database="nr", ) let qr2 = qr.with_total_hits(100) assert_eq(qr2.total_hits, 100) @@ -298,7 +487,10 @@ test "search_io_mock_blast_tabular" { ///| test "SearchIOIterator::parse_blast_tabular basic" { let content = @src.search_io_mock_blast_tabular() - let iter = @src.SearchIOIterator::new(content, @src.search_io_format_from_str("blast-tab")) + let iter = @src.SearchIOIterator::new( + content, + @src.search_io_format_from_str("blast-tab"), + ) let results = iter.parse_blast_tabular() assert_true(results.length() >= 0) } @@ -306,7 +498,10 @@ test "SearchIOIterator::parse_blast_tabular basic" { ///| test "SearchIOIterator::parse_blast_tabular hit count" { let content = @src.search_io_mock_blast_tabular() - let iter = @src.SearchIOIterator::new(content, @src.search_io_format_from_str("blast-tab")) + let iter = @src.SearchIOIterator::new( + content, + @src.search_io_format_from_str("blast-tab"), + ) let results = iter.parse_blast_tabular() assert_true(results.length() >= 0) } @@ -314,21 +509,30 @@ test "SearchIOIterator::parse_blast_tabular hit count" { ///| test "SearchIOIterator::parse_blast_tabular scores" { let content = @src.search_io_mock_blast_tabular() - let iter = @src.SearchIOIterator::new(content, @src.search_io_format_from_str("blast-tab")) + let iter = @src.SearchIOIterator::new( + content, + @src.search_io_format_from_str("blast-tab"), + ) let results = iter.parse_blast_tabular() assert_true(results.length() >= 0) } ///| test "SearchIOIterator::parse empty content" { - let iter = @src.SearchIOIterator::new("", @src.search_io_format_from_str("blast-tab")) + let iter = @src.SearchIOIterator::new( + "", + @src.search_io_format_from_str("blast-tab"), + ) let results = iter.parse_blast_tabular() assert_eq(results.length(), 0) } ///| test "SearchIOIterator::parse empty content text" { - let iter = @src.SearchIOIterator::new("", @src.search_io_format_from_str("blast-text")) + let iter = @src.SearchIOIterator::new( + "", + @src.search_io_format_from_str("blast-text"), + ) let results = iter.parse_blast_text() assert_eq(results.length(), 0) } @@ -340,7 +544,10 @@ test "SearchIOIterator::parse blast text basic" { ">hit1 description\n" + "Score = 200.5 (520 bits), Expect = 1.0e-50\n" + "Identities = 95/100 (95%)\n" - let iter = @src.SearchIOIterator::new(text, @src.search_io_format_from_str("blast-text")) + let iter = @src.SearchIOIterator::new( + text, + @src.search_io_format_from_str("blast-text"), + ) let results = iter.parse_blast_text() assert_true(results.length() >= 0) } @@ -348,7 +555,10 @@ test "SearchIOIterator::parse blast text basic" { ///| test "SearchIOIterator::parse dispatch" { let content = @src.search_io_mock_blast_tabular() - let iter = @src.SearchIOIterator::new(content, @src.search_io_format_from_str("blast-tab")) + let iter = @src.SearchIOIterator::new( + content, + @src.search_io_format_from_str("blast-tab"), + ) let results = iter.parse() assert_true(results.length() >= 0) } @@ -356,7 +566,10 @@ test "SearchIOIterator::parse dispatch" { ///| test "SearchIOIterator::parse unknown format returns empty" { let content = "some content" - let iter = @src.SearchIOIterator::new(content, @src.search_io_format_from_str("unknown")) + let iter = @src.SearchIOIterator::new( + content, + @src.search_io_format_from_str("unknown"), + ) let results = iter.parse() assert_eq(results.length(), 0) } diff --git a/test/moonbit/seq_complexity_test.mbt b/test/moonbit/seq_complexity_test.mbt index ff10eaa2..3a3125bc 100644 --- a/test/moonbit/seq_complexity_test.mbt +++ b/test/moonbit/seq_complexity_test.mbt @@ -134,6 +134,7 @@ test "seq_complexity: sequence_similarity_empty" { // LCC (Wooton-Federhen) tests +///| test "seq_complexity: lcc_basic" { let scores = @src.lcc("ATCGATCGATCGATCG", window=12, k=3) assert_true(scores.length() > 0) @@ -144,6 +145,7 @@ test "seq_complexity: lcc_basic" { } } +///| test "seq_complexity: lcc_low_complexity" { let scores = @src.lcc("AAAAAAAAAAAAAAAA", window=12, k=3) assert_true(scores.length() > 0) @@ -153,20 +155,28 @@ test "seq_complexity: lcc_low_complexity" { } } +///| test "seq_complexity: lcc_empty" { let scores = @src.lcc("", window=12, k=3) assert_eq(scores.length(), 0) } +///| test "seq_complexity: lcc_short" { let scores = @src.lcc("AC", window=12, k=3) assert_eq(scores.length(), 0) } +///| test "seq_complexity: lcc_low_complexity_regions" { // Mix of low complexity and high complexity let seq = "AAAAAAAAAAAAAACGATCGATCG" - let regions = @src.lcc_low_complexity_regions(seq, window=12, k=3, threshold=0.15) + let regions = @src.lcc_low_complexity_regions( + seq, + window=12, + k=3, + threshold=0.15, + ) // Should find low complexity region at the beginning assert_true(regions.length() > 0) } diff --git a/test/moonbit/seq_location_test.mbt b/test/moonbit/seq_location_test.mbt index 6326c089..59b6fe7d 100644 --- a/test/moonbit/seq_location_test.mbt +++ b/test/moonbit/seq_location_test.mbt @@ -1,161 +1,227 @@ ///| /// Test file for seq_location module. - test "exact_position" { let pos = @src.exact_position(42) - assert_eq!(pos.get(), 42) - assert_eq!(pos.to_string(), "42") - assert_eq!(pos.is_exact(), true) + assert_eq(pos.get(), 42) + assert_eq(pos.to_string(), "42") + assert_eq(pos.is_exact(), true) } +///| test "before_position" { let pos = @src.before_position(50) - assert_eq!(pos.get(), 50) - assert_eq!(pos.to_string(), "<50") - assert_eq!(pos.is_exact(), false) + assert_eq(pos.get(), 50) + assert_eq(pos.to_string(), "<50") + assert_eq(pos.is_exact(), false) } +///| test "after_position" { let pos = @src.after_position(100) - assert_eq!(pos.get(), 100) - assert_eq!(pos.to_string(), ">100") - assert_eq!(pos.is_exact(), false) + assert_eq(pos.get(), 100) + assert_eq(pos.to_string(), ">100") + assert_eq(pos.is_exact(), false) } +///| test "one_of_position" { let opts : Array[@src.Pos] = Array::new() opts.push(@src.exact_position(5)) opts.push(@src.exact_position(7)) opts.push(@src.exact_position(9)) let pos = @src.one_of_position(7, opts) - assert_eq!(pos.get(), 7) - assert_eq!(pos.is_exact(), false) + assert_eq(pos.get(), 7) + assert_eq(pos.is_exact(), false) let s = pos.to_string() - assert_eq!(s.starts_with("{"), true) - assert_eq!(s.contains("5"), true) - assert_eq!(s.contains("7"), true) - assert_eq!(s.contains("9"), true) + assert_eq(s.starts_with("{"), true) + assert_eq(s.contains("5"), true) + assert_eq(s.contains("7"), true) + assert_eq(s.contains("9"), true) } +///| test "within_position" { - let pos = @src.within_position(50, @src.exact_position(40), @src.exact_position(60)) - assert_eq!(pos.get(), 50) - assert_eq!(pos.is_exact(), false) + let pos = @src.within_position( + 50, + @src.exact_position(40), + @src.exact_position(60), + ) + assert_eq(pos.get(), 50) + assert_eq(pos.is_exact(), false) let s = pos.to_string() - assert_eq!(s.starts_with("("), true) - assert_eq!(s.contains("40"), true) - assert_eq!(s.contains("60"), true) + assert_eq(s.starts_with("("), true) + assert_eq(s.contains("40"), true) + assert_eq(s.contains("60"), true) } +///| test "simple_location_basic" { - let loc = @src.simple_location(@src.exact_position(10), @src.exact_position(50), strand="+") - assert_eq!(loc.start(), 10) - assert_eq!(loc.end(), 50) - assert_eq!(loc.strand(), "+") - assert_eq!(loc.len(), 40) - assert_eq!(loc.is_compound(), false) - assert_eq!(loc.to_string(), "10..50") + let loc = @src.simple_location( + @src.exact_position(10), + @src.exact_position(50), + strand="+", + ) + assert_eq(loc.start(), 10) + assert_eq(loc.end(), 50) + assert_eq(loc.strand(), "+") + assert_eq(loc.len(), 40) + assert_eq(loc.is_compound(), false) + assert_eq(loc.to_string(), "10..50") } +///| test "simple_location_reverse_strand" { - let loc = @src.simple_location(@src.exact_position(100), @src.exact_position(200), strand="-") - assert_eq!(loc.strand(), "-") - assert_eq!(loc.len(), 100) + let loc = @src.simple_location( + @src.exact_position(100), + @src.exact_position(200), + strand="-", + ) + assert_eq(loc.strand(), "-") + assert_eq(loc.len(), 100) } +///| test "simple_location_overlaps" { - let loc1 = @src.simple_location(@src.exact_position(10), @src.exact_position(50)) - let loc2 = @src.simple_location(@src.exact_position(30), @src.exact_position(70)) - let loc3 = @src.simple_location(@src.exact_position(100), @src.exact_position(200)) - assert_eq!(loc1.overlaps(loc2), true) - assert_eq!(loc1.overlaps(loc3), false) + let loc1 = @src.simple_location( + @src.exact_position(10), + @src.exact_position(50), + ) + let loc2 = @src.simple_location( + @src.exact_position(30), + @src.exact_position(70), + ) + let loc3 = @src.simple_location( + @src.exact_position(100), + @src.exact_position(200), + ) + assert_eq(loc1.overlaps(loc2), true) + assert_eq(loc1.overlaps(loc3), false) } +///| test "simple_location_contains" { - let loc = @src.simple_location(@src.exact_position(10), @src.exact_position(50)) - assert_eq!(loc.contains(10), true) - assert_eq!(loc.contains(25), true) - assert_eq!(loc.contains(49), true) - assert_eq!(loc.contains(5), false) - assert_eq!(loc.contains(50), false) - assert_eq!(loc.contains(100), false) + let loc = @src.simple_location( + @src.exact_position(10), + @src.exact_position(50), + ) + assert_eq(loc.contains(10), true) + assert_eq(loc.contains(25), true) + assert_eq(loc.contains(49), true) + assert_eq(loc.contains(5), false) + assert_eq(loc.contains(50), false) + assert_eq(loc.contains(100), false) } +///| test "compound_location_basic" { - let exon1 = @src.simple_location(@src.exact_position(0), @src.exact_position(100)) - let exon2 = @src.simple_location(@src.exact_position(200), @src.exact_position(350)) + let exon1 = @src.simple_location( + @src.exact_position(0), + @src.exact_position(100), + ) + let exon2 = @src.simple_location( + @src.exact_position(200), + @src.exact_position(350), + ) let locs : Array[@src.Loc] = Array::new() locs.push(exon1) locs.push(exon2) let cl = @src.compound_location(locs, strand="+") - assert_eq!(cl.is_compound(), true) - assert_eq!(cl.parts().length(), 2) - assert_eq!(cl.len(), 250) + assert_eq(cl.is_compound(), true) + assert_eq(cl.parts().length(), 2) + assert_eq(cl.len(), 250) } +///| test "compound_from_simple" { let starts : Array[Int] = [0, 200, 500] let ends : Array[Int] = [100, 350, 600] let cl = @src.compound_from_simple(starts, ends, strand="+") - assert_eq!(cl.is_compound(), true) - assert_eq!(cl.parts().length(), 3) - assert_eq!(cl.len(), 350) // 100 + 150 + 100 + assert_eq(cl.is_compound(), true) + assert_eq(cl.parts().length(), 3) + assert_eq(cl.len(), 350) // 100 + 150 + 100 } +///| test "compound_location_overlaps" { - let exon1 = @src.simple_location(@src.exact_position(0), @src.exact_position(100)) - let exon2 = @src.simple_location(@src.exact_position(200), @src.exact_position(350)) + let exon1 = @src.simple_location( + @src.exact_position(0), + @src.exact_position(100), + ) + let exon2 = @src.simple_location( + @src.exact_position(200), + @src.exact_position(350), + ) let locs : Array[@src.Loc] = Array::new() locs.push(exon1) locs.push(exon2) let cl = @src.compound_location(locs) - let overlapping = @src.simple_location(@src.exact_position(50), @src.exact_position(150)) - let not_overlapping = @src.simple_location(@src.exact_position(400), @src.exact_position(500)) - assert_eq!(cl.overlaps(overlapping), true) - assert_eq!(cl.overlaps(not_overlapping), false) + let overlapping = @src.simple_location( + @src.exact_position(50), + @src.exact_position(150), + ) + let not_overlapping = @src.simple_location( + @src.exact_position(400), + @src.exact_position(500), + ) + assert_eq(cl.overlaps(overlapping), true) + assert_eq(cl.overlaps(not_overlapping), false) } +///| test "compound_location_contains" { - let exon1 = @src.simple_location(@src.exact_position(0), @src.exact_position(100)) - let exon2 = @src.simple_location(@src.exact_position(200), @src.exact_position(350)) + let exon1 = @src.simple_location( + @src.exact_position(0), + @src.exact_position(100), + ) + let exon2 = @src.simple_location( + @src.exact_position(200), + @src.exact_position(350), + ) let locs : Array[@src.Loc] = Array::new() locs.push(exon1) locs.push(exon2) let cl = @src.compound_location(locs) - assert_eq!(cl.contains(50), true) - assert_eq!(cl.contains(250), true) - assert_eq!(cl.contains(150), false) - assert_eq!(cl.contains(400), false) + assert_eq(cl.contains(50), true) + assert_eq(cl.contains(250), true) + assert_eq(cl.contains(150), false) + assert_eq(cl.contains(400), false) } +///| test "location_genbank_format" { - let loc = @src.simple_location(@src.exact_position(9), @src.exact_position(50)) + let loc = @src.simple_location( + @src.exact_position(9), + @src.exact_position(50), + ) let gb = @src.location_to_genbank(loc) - assert_eq!(gb.contains("10"), true) - assert_eq!(gb.contains("50"), true) + assert_eq(gb.contains("10"), true) + assert_eq(gb.contains("50"), true) } +///| test "from_genbank_coords" { let loc = @src.from_genbank_coords(10, 50, strand="+") - assert_eq!(loc.start(), 9) - assert_eq!(loc.end(), 50) - assert_eq!(loc.strand(), "+") - assert_eq!(loc.len(), 41) + assert_eq(loc.start(), 9) + assert_eq(loc.end(), 50) + assert_eq(loc.strand(), "+") + assert_eq(loc.len(), 41) } +///| test "parse_genbank_location_simple" { let loc = @src.parse_genbank_location("10..50") - assert_eq!(loc.start(), 9) // 0-based - assert_eq!(loc.end(), 50) - assert_eq!(loc.len(), 41) + assert_eq(loc.start(), 9) // 0-based + assert_eq(loc.end(), 50) + assert_eq(loc.len(), 41) } +///| test "seq_location_sample" { let samples = @src.seq_location_sample() - assert_eq!(samples.length(), 5) + assert_eq(samples.length(), 5) // First is simple forward - assert_eq!(samples[0].is_compound(), false) + assert_eq(samples[0].is_compound(), false) // Third is compound - assert_eq!(samples[2].is_compound(), true) + assert_eq(samples[2].is_compound(), true) } diff --git a/test/moonbit/seq_quality_trim_test.mbt b/test/moonbit/seq_quality_trim_test.mbt index a27c52c1..c61d1a39 100644 --- a/test/moonbit/seq_quality_trim_test.mbt +++ b/test/moonbit/seq_quality_trim_test.mbt @@ -113,7 +113,10 @@ test "sqt_trim_adapter_perfect_match" { let qual = "IIIIIIIIIIIIIIII" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=0, + seq, + qual, + adapter, + allowed_mismatches=0, ) assert_eq(t_seq, "ACGTACGT") assert_eq(t_seq.length(), 8) @@ -125,7 +128,10 @@ test "sqt_trim_adapter_with_mismatches" { let qual = "IIIIIIIIIIIIIIII" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=2, + seq, + qual, + adapter, + allowed_mismatches=2, ) assert_eq(t_seq, "ACGTACGT") } @@ -136,7 +142,10 @@ test "sqt_trim_adapter_no_match" { let qual = "IIIIIIIIIIIIIIII" let adapter = "GGGGGGGG" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=0, + seq, + qual, + adapter, + allowed_mismatches=0, ) assert_eq(t_seq, "ACGTACGTACGTACGT") assert_eq(t_seq.length(), 16) @@ -146,9 +155,7 @@ test "sqt_trim_adapter_no_match" { test "sqt_trim_adapter_empty_adapter" { let seq = "ACGTACGT" let qual = "IIIIIIII" - let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, "", allowed_mismatches=0, - ) + let (t_seq, _) = @src.sqt_trim_adapter(seq, qual, "", allowed_mismatches=0) assert_eq(t_seq, "ACGTACGT") } @@ -158,7 +165,10 @@ test "sqt_trim_adapter_short_sequence" { let qual = "II" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=0, + seq, + qual, + adapter, + allowed_mismatches=0, ) assert_eq(t_seq, "AC") } @@ -169,7 +179,10 @@ test "sqt_trim_adapter_at_beginning" { let qual = "IIIIIIIIIIIIIIII" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=0, + seq, + qual, + adapter, + allowed_mismatches=0, ) assert_eq(t_seq.length(), 0) } @@ -180,7 +193,10 @@ test "sqt_trim_adapter_partial_overlap" { let qual = "IIIIIIIIIIIIIIII" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=1, + seq, + qual, + adapter, + allowed_mismatches=1, ) assert_eq(t_seq, "ACGTACGT") } @@ -191,7 +207,10 @@ test "sqt_trim_adapter_at_3_prime_end" { let qual = "IIIIIIIIIIIIIIIIIIIII" let adapter = "AGATCGGAA" let (t_seq, _) = @src.sqt_trim_adapter( - seq, qual, adapter, allowed_mismatches=0, + seq, + qual, + adapter, + allowed_mismatches=0, ) assert_eq(t_seq, "ACGTACGTACGT") } @@ -344,9 +363,7 @@ test "sqt_trim_reads_batch_with_adapter_config" { ] let qual_config = @src.QualityConfig::new(20, 4, 4, 20) let adapter_config = @src.AdapterConfig::new(["AGATCGGAA"], 4, 0) - let results = @src.sqt_trim_reads( - reads, qual_config, Some(adapter_config), - ) + let results = @src.sqt_trim_reads(reads, qual_config, Some(adapter_config)) assert_eq(results.length(), 1) assert_eq(results[0].trimmed_seq, "ACGTACGT") assert_true(results[0].trim_type == "adapter:AGATCGGAA") @@ -354,9 +371,7 @@ test "sqt_trim_reads_batch_with_adapter_config" { ///| test "sqt_trim_reads_quality_trim_only" { - let reads = [ - @src.FastqRead::new("r1", "ACGTACGT", "IIII####"), - ] + let reads = [@src.FastqRead::new("r1", "ACGTACGT", "IIII####")] let qual_config = @src.QualityConfig::new(20, 4, 4, 20) let results = @src.sqt_trim_reads(reads, qual_config, None) assert_eq(results.length(), 1) @@ -365,9 +380,7 @@ test "sqt_trim_reads_quality_trim_only" { ///| test "sqt_trim_reads_poly_a_trim" { - let reads = [ - @src.FastqRead::new("r1", "ACGTACGGGAAAA", "IIIIIIIIIIIIII"), - ] + let reads = [@src.FastqRead::new("r1", "ACGTACGGGAAAA", "IIIIIIIIIIIIII")] let qual_config = @src.QualityConfig::new(20, 4, 4, 20) let results = @src.sqt_trim_reads(reads, qual_config, None) assert_eq(results.length(), 1) @@ -381,9 +394,7 @@ test "sqt_trim_reads_adapter_at_end" { ] let qual_config = @src.QualityConfig::new(20, 4, 4, 20) let adapter_config = @src.AdapterConfig::new(["AGATCGGAA"], 4, 0) - let results = @src.sqt_trim_reads( - reads, qual_config, Some(adapter_config), - ) + let results = @src.sqt_trim_reads(reads, qual_config, Some(adapter_config)) assert_eq(results.length(), 1) assert_eq(results[0].trimmed_seq, "ACGTACGTACGT") } @@ -398,9 +409,7 @@ test "sqt_trim_reads_empty_batch" { ///| test "sqt_trim_reads_preserves_original_data" { - let reads = [ - @src.FastqRead::new("read_orig", "ACGTACGT", "IIIIIIII"), - ] + let reads = [@src.FastqRead::new("read_orig", "ACGTACGT", "IIIIIIII")] let qual_config = @src.QualityConfig::new(20, 4, 4, 20) let results = @src.sqt_trim_reads(reads, qual_config, None) assert_eq(results.length(), 1) @@ -416,7 +425,9 @@ test "sqt_compute_stats_basic" { @src.FastqRead::new("r2", "TTTTTTTT", "IIIIIIII"), ] let results = [ - @src.TrimResult::new("r1", "ACGTACGT", "ACGTACGT", "IIIIIIII", "IIIIIIII", "none"), + @src.TrimResult::new( + "r1", "ACGTACGT", "ACGTACGT", "IIIIIIII", "IIIIIIII", "none", + ), @src.TrimResult::new("r2", "TTTTTTTT", "", "IIIIIIII", "", "poly_A"), ] let stats = @src.sqt_compute_stats(results, reads) @@ -434,8 +445,12 @@ test "sqt_compute_stats_gc_content" { @src.FastqRead::new("at_high", "AAAATTTT", "IIIIIIII"), ] let results = [ - @src.TrimResult::new("gc_high", "GGGGCCCC", "GGGGCCCC", "IIIIIIII", "IIIIIIII", "none"), - @src.TrimResult::new("at_high", "AAAATTTT", "AAAATTTT", "IIIIIIII", "IIIIIIII", "none"), + @src.TrimResult::new( + "gc_high", "GGGGCCCC", "GGGGCCCC", "IIIIIIII", "IIIIIIII", "none", + ), + @src.TrimResult::new( + "at_high", "AAAATTTT", "AAAATTTT", "IIIIIIII", "IIIIIIII", "none", + ), ] let stats = @src.sqt_compute_stats(results, reads) assert_true(stats.gc_content_before > 0.4) @@ -455,12 +470,8 @@ test "sqt_compute_stats_empty" { ///| test "sqt_compute_stats_all_discarded" { - let reads = [ - @src.FastqRead::new("r1", "ACGT", "IIII"), - ] - let results = [ - @src.TrimResult::new("r1", "ACGT", "", "IIII", "", "quality"), - ] + let reads = [@src.FastqRead::new("r1", "ACGT", "IIII")] + let results = [@src.TrimResult::new("r1", "ACGT", "", "IIII", "", "quality")] let stats = @src.sqt_compute_stats(results, reads) assert_eq(stats.total_reads, 1) assert_eq(stats.kept_reads, 0) @@ -532,9 +543,7 @@ test "sqt_fastq_parse_multiple_with_variable_length" { ///| test "sqt_fastq_serialize_single" { - let reads = [ - @src.FastqRead::new("test1", "ACGT", "IIII"), - ] + let reads = [@src.FastqRead::new("test1", "ACGT", "IIII")] let output = @src.sqt_fastq_serialize(reads) assert_eq(output, "@test1\nACGT\n+\nIIII\n") } @@ -561,4 +570,4 @@ test "sqt_fastq_roundtrip" { assert_eq(reparsed[0].sequence, "ACGTACGT") assert_eq(reparsed[1].id, "sample2") assert_eq(reparsed[1].sequence, "TGCA") -} \ No newline at end of file +} diff --git a/test/moonbit/seqfeature_advanced_test.mbt b/test/moonbit/seqfeature_advanced_test.mbt index 63ac3e2b..625fbfda 100644 --- a/test/moonbit/seqfeature_advanced_test.mbt +++ b/test/moonbit/seqfeature_advanced_test.mbt @@ -580,4 +580,4 @@ test "seq_feature_extended_qualifiers_independent" { let feat2 = feat1.add_qualifier("gene", "BRCA1") assert_eq(feat1.qualifiers().length(), 0) assert_eq(feat2.qualifiers().length(), 1) -} \ No newline at end of file +} diff --git a/test/moonbit/seqio_advanced_test.mbt b/test/moonbit/seqio_advanced_test.mbt index d249c719..ffd0df6f 100644 --- a/test/moonbit/seqio_advanced_test.mbt +++ b/test/moonbit/seqio_advanced_test.mbt @@ -10,8 +10,7 @@ ///| test "parse_embl_basic_single" { - let embl_text = - "ID HSBGLOD; SV 1; linear; genomic DNA; STD; HUM; 500 BP.\n" + + let embl_text = "ID HSBGLOD; SV 1; linear; genomic DNA; STD; HUM; 500 BP.\n" + "XX\n" + "AC M12345;\n" + "XX\n" + @@ -26,13 +25,15 @@ test "parse_embl_basic_single" { assert_eq(records[0].id, "M12345") assert_eq(records[0].name, "HSBGLOD") assert_eq(records[0].description, "Human beta-globin gene region") - assert_eq(records[0].seq.to_string(), "ATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGC") + assert_eq( + records[0].seq.to_string(), + "ATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGCATGC", + ) } ///| test "parse_embl_multiple_records" { - let embl_text = - "ID SEQ1; SV 1; linear; DNA; STD; 10 BP.\n" + + let embl_text = "ID SEQ1; SV 1; linear; DNA; STD; 10 BP.\n" + "XX\n" + "AC ACC001;\n" + "XX\n" + @@ -63,8 +64,7 @@ test "parse_embl_multiple_records" { ///| test "parse_embl_no_ac_uses_id" { - let embl_text = - "ID MYSEQ; SV 1; linear; DNA; STD; 4 BP.\n" + + let embl_text = "ID MYSEQ; SV 1; linear; DNA; STD; 4 BP.\n" + "XX\n" + "DE Sequence without accession.\n" + "XX\n" + @@ -82,8 +82,7 @@ test "parse_embl_no_ac_uses_id" { ///| test "parse_pir_basic" { - let pir_text = - ">P1;ALB_HUMAN\n" + + let pir_text = ">P1;ALB_HUMAN\n" + "Serum albumin precursor (Human).\n" + "MKWVTFISLLFLFSSAYSRGVFRRDTHKSEIAHRFKDLGEEHFKGLVLIAFSQYLQQCPFDEHVKLVNELTEFAK*" let records = @src.parse_pir(pir_text) @@ -97,8 +96,7 @@ test "parse_pir_basic" { ///| test "parse_pir_multiple" { - let pir_text = - ">P1;PROT1\n" + + let pir_text = ">P1;PROT1\n" + "Protein one.\n" + "MVLSQDEVVCF* \n" + ">P1;PROT2\n" + @@ -114,10 +112,7 @@ test "parse_pir_multiple" { ///| test "parse_pir_sequence_spaces_ignored" { - let pir_text = - ">DL;DNA1\n" + - "Linear DNA fragment.\n" + - "atgc atgc atgc *" + let pir_text = ">DL;DNA1\n" + "Linear DNA fragment.\n" + "atgc atgc atgc *" let records = @src.parse_pir(pir_text) assert_eq(records.length(), 1) assert_eq(records[0].id, "DNA1") @@ -227,8 +222,18 @@ test "write_genbank_sequence_format" { ///| test "write_genbank_multiple_records" { - let r1 = @src.SeqRecord::new(@src.Seq::new("ACGT"), id="A1", name="L1", description="First") - let r2 = @src.SeqRecord::new(@src.Seq::new("TGCA"), id="B2", name="L2", description="Second") + let r1 = @src.SeqRecord::new( + @src.Seq::new("ACGT"), + id="A1", + name="L1", + description="First", + ) + let r2 = @src.SeqRecord::new( + @src.Seq::new("TGCA"), + id="B2", + name="L2", + description="Second", + ) let out = @src.write_genbank([r1, r2]) // Should have two LOCUS and two terminators assert_eq(out.split("LOCUS").length(), 3) // includes 1 before first match @@ -239,8 +244,7 @@ test "write_genbank_multiple_records" { ///| test "seqio_parse_embl_through_unified" { - let embl = - "ID S1; SV 1; linear; DNA; STD; 4 BP.\n" + + let embl = "ID S1; SV 1; linear; DNA; STD; 4 BP.\n" + "XX\n" + "AC A001;\n" + "XX\n" + @@ -287,9 +291,7 @@ test "seqio_parse_tab_through_unified" { ///| test "seqio_write_genbank_tab_through_unified" { - let records = [ - @src.SeqRecord::new(@src.Seq::new("ACGT"), id="r1"), - ] + let records = [@src.SeqRecord::new(@src.Seq::new("ACGT"), id="r1")] try { let gb_text = @src.seqio_write(records, "genbank") assert_true(gb_text.contains("LOCUS")) diff --git a/test/moonbit/seqlogo_test.mbt b/test/moonbit/seqlogo_test.mbt index 6e6b7019..1f958e2e 100644 --- a/test/moonbit/seqlogo_test.mbt +++ b/test/moonbit/seqlogo_test.mbt @@ -9,12 +9,7 @@ ///| test "seqlogo_pwm_new_basic" { - let matrix = [ - [0.25, 0.25], - [0.25, 0.25], - [0.25, 0.25], - [0.25, 0.25], - ] + let matrix = [[0.25, 0.25], [0.25, 0.25], [0.25, 0.25], [0.25, 0.25]] let pwm = @src.SeqLogoPwm::new(matrix) assert_eq(pwm.width(), 2) assert_eq(pwm.alphabet_size(), 4) @@ -22,12 +17,7 @@ test "seqlogo_pwm_new_basic" { ///| test "seqlogo_pwm_new_get_values" { - let matrix = [ - [0.9, 0.1], - [0.03, 0.1], - [0.04, 0.7], - [0.03, 0.1], - ] + let matrix = [[0.9, 0.1], [0.03, 0.1], [0.04, 0.7], [0.03, 0.1]] let pwm = @src.SeqLogoPwm::new(matrix) assert_true((pwm.get(0, 0) - 0.9).abs() < 0.001) assert_true((pwm.get(0, 1) - 0.1).abs() < 0.001) @@ -39,10 +29,10 @@ test "seqlogo_pwm_new_get_values" { test "seqlogo_pwm_new_get_out_of_bounds" { let matrix = [[0.25], [0.25], [0.25], [0.25]] let pwm = @src.SeqLogoPwm::new(matrix) - assert_true((pwm.get(-1, 0)).abs() < 0.001) - assert_true((pwm.get(4, 0)).abs() < 0.001) - assert_true((pwm.get(0, -1)).abs() < 0.001) - assert_true((pwm.get(0, 100)).abs() < 0.001) + assert_true(pwm.get(-1, 0).abs() < 0.001) + assert_true(pwm.get(4, 0).abs() < 0.001) + assert_true(pwm.get(0, -1).abs() < 0.001) + assert_true(pwm.get(0, 100).abs() < 0.001) } ///| @@ -89,8 +79,8 @@ test "seqlogo_from_sequences_basic" { // Position 3: all T -> freq 1.0 assert_true((pwm.get(3, 3) - 1.0).abs() < 0.001) // Off-diagonal frequencies are 0 - assert_true((pwm.get(0, 1)).abs() < 0.001) - assert_true((pwm.get(1, 0)).abs() < 0.001) + assert_true(pwm.get(0, 1).abs() < 0.001) + assert_true(pwm.get(1, 0).abs() < 0.001) } ///| @@ -134,7 +124,7 @@ test "seqlogo_letter_new_and_accessors" { test "seqlogo_letter_zero_height" { let letter = @src.SeqLogoLetter::new('T', 0.0, "#CC0000") assert_eq(letter.letter(), 'T') - assert_true((letter.height()).abs() < 0.001) + assert_true(letter.height().abs() < 0.001) assert_eq(letter.color(), "#CC0000") } @@ -168,28 +158,18 @@ test "seqlogo_column_add_letter" { ///| test "seqlogo_ic_uniform_pwm_zero" { // All positions have uniform 0.25 frequency -> IC = 0 - let matrix = [ - [0.25, 0.25], - [0.25, 0.25], - [0.25, 0.25], - [0.25, 0.25], - ] + let matrix = [[0.25, 0.25], [0.25, 0.25], [0.25, 0.25], [0.25, 0.25]] let pwm = @src.SeqLogoPwm::new(matrix) let ic = @src.seqlogo_information_content(pwm) assert_eq(ic.length(), 2) - assert_true((ic[0]).abs() < 0.001) - assert_true((ic[1]).abs() < 0.001) + assert_true(ic[0].abs() < 0.001) + assert_true(ic[1].abs() < 0.001) } ///| test "seqlogo_ic_conserved_position" { // Fully conserved A -> IC = log2(4) = 2.0 - let matrix = [ - [1.0], - [0.0], - [0.0], - [0.0], - ] + let matrix = [[1.0], [0.0], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let ic = @src.seqlogo_information_content(pwm) assert_eq(ic.length(), 1) @@ -201,12 +181,7 @@ test "seqlogo_ic_mixed_position" { // A=0.5, C=0.25, G=0.125, T=0.125 // IC = 0.5*log2(2) + 0.25*log2(1) + 0.125*log2(0.5) + 0.125*log2(0.5) // = 0.5 - 0.125 - 0.125 = 0.25 - let matrix = [ - [0.5], - [0.25], - [0.125], - [0.125], - ] + let matrix = [[0.5], [0.25], [0.125], [0.125]] let pwm = @src.SeqLogoPwm::new(matrix) let ic = @src.seqlogo_information_content(pwm) assert_true((ic[0] - 0.25).abs() < 0.001) @@ -223,12 +198,7 @@ test "seqlogo_ic_bg_non_uniform" { // IC = 0.5*log2(0.5/0.1) + 0.5*log2(0.5/0.4) // = 0.5*log2(5) + 0.5*log2(1.25) // = 0.5*2.321928 + 0.5*0.321928 = 1.321928 - let matrix = [ - [0.5], - [0.5], - [0.0], - [0.0], - ] + let matrix = [[0.5], [0.5], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let ic = @src.seqlogo_information_content_bg(pwm, [0.1, 0.4, 0.4, 0.1]) assert_true((ic[0] - 1.321928).abs() < 0.001) @@ -236,17 +206,10 @@ test "seqlogo_ic_bg_non_uniform" { ///| test "seqlogo_ic_bg_uniform_equals_default" { - let matrix = [ - [0.9], - [0.03], - [0.04], - [0.03], - ] + let matrix = [[0.9], [0.03], [0.04], [0.03]] let pwm = @src.SeqLogoPwm::new(matrix) let ic_default = @src.seqlogo_information_content(pwm) - let ic_bg = @src.seqlogo_information_content_bg(pwm, [ - 0.25, 0.25, 0.25, 0.25, - ]) + let ic_bg = @src.seqlogo_information_content_bg(pwm, [0.25, 0.25, 0.25, 0.25]) assert_eq(ic_default.length(), ic_bg.length()) assert_true((ic_default[0] - ic_bg[0]).abs() < 0.001) } @@ -258,12 +221,7 @@ test "seqlogo_ic_bg_uniform_equals_default" { ///| test "seqlogo_total_ic_conserved" { // Two fully conserved positions -> total IC = 4.0 - let matrix = [ - [1.0, 1.0], - [0.0, 0.0], - [0.0, 0.0], - [0.0, 0.0], - ] + let matrix = [[1.0, 1.0], [0.0, 0.0], [0.0, 0.0], [0.0, 0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let total = @src.seqlogo_total_information_content(pwm) assert_true((total - 4.0).abs() < 0.001) @@ -303,8 +261,8 @@ test "seqlogo_max_ic_protein" { ///| test "seqlogo_max_ic_edge_cases" { // alphabet_size <= 1 -> 0.0 - assert_true((@src.seqlogo_max_information_content(0)).abs() < 0.001) - assert_true((@src.seqlogo_max_information_content(1)).abs() < 0.001) + assert_true(@src.seqlogo_max_information_content(0).abs() < 0.001) + assert_true(@src.seqlogo_max_information_content(1).abs() < 0.001) // Binary alphabet -> log2(2) = 1.0 assert_true((@src.seqlogo_max_information_content(2) - 1.0).abs() < 0.001) } @@ -315,12 +273,7 @@ test "seqlogo_max_ic_edge_cases" { ///| test "seqlogo_compute_logo_columns_count" { - let matrix = [ - [0.25, 1.0], - [0.25, 0.0], - [0.25, 0.0], - [0.25, 0.0], - ] + let matrix = [[0.25, 1.0], [0.25, 0.0], [0.25, 0.0], [0.25, 0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) assert_eq(logo.width(), 2) @@ -330,12 +283,7 @@ test "seqlogo_compute_logo_columns_count" { ///| test "seqlogo_compute_logo_conserved_heights" { // Fully conserved A: IC=2.0, A height=2.0, others=0.0 - let matrix = [ - [1.0], - [0.0], - [0.0], - [0.0], - ] + let matrix = [[1.0], [0.0], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) let col = logo.columns()[0] @@ -345,37 +293,27 @@ test "seqlogo_compute_logo_conserved_heights" { assert_eq(col.letters()[0].letter(), 'A') assert_true((col.letters()[0].height() - 2.0).abs() < 0.001) // Other letters have height 0 - assert_true((col.letters()[1].height()).abs() < 0.001) - assert_true((col.letters()[2].height()).abs() < 0.001) - assert_true((col.letters()[3].height()).abs() < 0.001) + assert_true(col.letters()[1].height().abs() < 0.001) + assert_true(col.letters()[2].height().abs() < 0.001) + assert_true(col.letters()[3].height().abs() < 0.001) } ///| test "seqlogo_compute_logo_uniform_zero_ic" { - let matrix = [ - [0.25], - [0.25], - [0.25], - [0.25], - ] + let matrix = [[0.25], [0.25], [0.25], [0.25]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) let col = logo.columns()[0] - assert_true((col.info_content()).abs() < 0.001) + assert_true(col.info_content().abs() < 0.001) // All letter heights are 0 since IC=0 for letter in col.letters() { - assert_true((letter.height()).abs() < 0.001) + assert_true(letter.height().abs() < 0.001) } } ///| test "seqlogo_compute_logo_total_ic" { - let matrix = [ - [1.0, 0.25], - [0.0, 0.25], - [0.0, 0.25], - [0.0, 0.25], - ] + let matrix = [[1.0, 0.25], [0.0, 0.25], [0.0, 0.25], [0.0, 0.25]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) // Position 0: IC=2.0, Position 1: IC=0.0 -> total = 2.0 @@ -384,12 +322,7 @@ test "seqlogo_compute_logo_total_ic" { ///| test "seqlogo_compute_logo_letter_colors" { - let matrix = [ - [1.0], - [0.0], - [0.0], - [0.0], - ] + let matrix = [[1.0], [0.0], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) let col = logo.columns()[0] @@ -451,12 +384,12 @@ test "seqlogo_to_ascii_contains_letters" { ///| test "seqlogo_to_ascii_zero_height" { - let ascii = @src.seqlogo_to_ascii(@src.seqlogo_compute_logo(@src.SeqLogoPwm::new([ - [0.25], - [0.25], - [0.25], - [0.25], - ])), 5) + let ascii = @src.seqlogo_to_ascii( + @src.seqlogo_compute_logo( + @src.SeqLogoPwm::new([[0.25], [0.25], [0.25], [0.25]]), + ), + 5, + ) // IC=0 so no letters rendered, only spaces and newlines assert_true(ascii.length() > 0) assert_false(ascii.contains("A")) @@ -480,12 +413,7 @@ test "seqlogo_to_text_table_has_header" { ///| test "seqlogo_to_text_table_tab_separated" { - let matrix = [ - [1.0], - [0.0], - [0.0], - [0.0], - ] + let matrix = [[1.0], [0.0], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) let table = @src.seqlogo_to_text_table(logo) @@ -528,12 +456,7 @@ test "seqlogo_consensus_from_sequences" { ///| test "seqlogo_consensus_tie_first_base" { // When frequencies are tied, the lowest-index base wins - let matrix = [ - [0.25], - [0.25], - [0.25], - [0.25], - ] + let matrix = [[0.25], [0.25], [0.25], [0.25]] let pwm = @src.SeqLogoPwm::new(matrix) let consensus = @src.seqlogo_consensus_sequence(pwm) assert_eq(consensus, "A") @@ -556,12 +479,7 @@ test "seqlogo_logo_summary_non_empty" { ///| test "seqlogo_logo_summary_width" { - let matrix = [ - [1.0, 0.25], - [0.0, 0.25], - [0.0, 0.25], - [0.0, 0.25], - ] + let matrix = [[1.0, 0.25], [0.0, 0.25], [0.0, 0.25], [0.0, 0.25]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo(pwm) let summary = @src.seqlogo_logo_summary(logo) @@ -610,12 +528,7 @@ test "seqlogo_sample_pwm_from_sample_sequences" { ///| test "seqlogo_edge_single_position" { - let matrix = [ - [1.0], - [0.0], - [0.0], - [0.0], - ] + let matrix = [[1.0], [0.0], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) assert_eq(pwm.width(), 1) let ic = @src.seqlogo_information_content(pwm) @@ -654,7 +567,7 @@ test "seqlogo_edge_all_equal_frequencies" { assert_true(v.abs() < 0.001) } let logo = @src.seqlogo_compute_logo(pwm) - assert_true((logo.total_ic()).abs() < 0.001) + assert_true(logo.total_ic().abs() < 0.001) } ///| @@ -695,12 +608,7 @@ test "seqlogo_edge_to_ascii_zero_max_height" { ///| test "seqlogo_compute_logo_bg_non_uniform" { - let matrix = [ - [0.5], - [0.5], - [0.0], - [0.0], - ] + let matrix = [[0.5], [0.5], [0.0], [0.0]] let pwm = @src.SeqLogoPwm::new(matrix) let logo = @src.seqlogo_compute_logo_bg(pwm, [0.1, 0.4, 0.4, 0.1]) let col = logo.columns()[0] diff --git a/test/moonbit/seqxml_io_test.mbt b/test/moonbit/seqxml_io_test.mbt index ecb563e4..58ae344e 100644 --- a/test/moonbit/seqxml_io_test.mbt +++ b/test/moonbit/seqxml_io_test.mbt @@ -10,14 +10,8 @@ ///| test "seqxml_type_from_tag" { // Known tags map to their corresponding types. - assert_eq( - @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("dna")), - "dna", - ) - assert_eq( - @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("rna")), - "rna", - ) + assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("dna")), "dna") + assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("rna")), "rna") assert_eq( @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("protein")), "protein", @@ -27,24 +21,33 @@ test "seqxml_type_from_tag" { @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("xyz")), "unknown", ) - assert_eq( - @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("")), - "unknown", - ) + assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("")), "unknown") } ///| test "seqxml_type_to_tag" { assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("dna")), "dna") assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("rna")), "rna") - assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("protein")), "protein") - assert_eq(@src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("unknown")), "unknown") + assert_eq( + @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("protein")), + "protein", + ) + assert_eq( + @src.SeqXmlType::to_tag(@src.SeqXmlType::from_tag("unknown")), + "unknown", + ) } ///| test "seqxml_type_description" { - assert_eq(@src.SeqXmlType::description(@src.SeqXmlType::from_tag("dna")), "DNA") - assert_eq(@src.SeqXmlType::description(@src.SeqXmlType::from_tag("rna")), "RNA") + assert_eq( + @src.SeqXmlType::description(@src.SeqXmlType::from_tag("dna")), + "DNA", + ) + assert_eq( + @src.SeqXmlType::description(@src.SeqXmlType::from_tag("rna")), + "RNA", + ) assert_eq( @src.SeqXmlType::description(@src.SeqXmlType::from_tag("protein")), "Protein", @@ -234,8 +237,7 @@ test "seqxml_document_full_metadata" { ///| test "seqxml_parse_dna_entry" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " ATGCATGC\n" + @@ -253,8 +255,7 @@ test "seqxml_parse_dna_entry" { ///| test "seqxml_parse_rna_entry" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " AUGCAUGC\n" + @@ -271,8 +272,7 @@ test "seqxml_parse_rna_entry" { ///| test "seqxml_parse_protein_entry" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " MKLVGV\n" + @@ -289,8 +289,7 @@ test "seqxml_parse_protein_entry" { ///| test "seqxml_parse_with_species_sourcedb" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " \n" + @@ -311,8 +310,7 @@ test "seqxml_parse_with_species_sourcedb" { ///| test "seqxml_parse_with_properties" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " \n" + @@ -336,8 +334,7 @@ test "seqxml_parse_with_properties" { ///| test "seqxml_parse_multiple_entries" { - let content = - "\n" + + let content = "\n" + "\n" + " ACGT\n" + " AUGU\n" + @@ -450,7 +447,9 @@ test "seqxml_document_to_xml" { assert_true(xml.contains("")) assert_true(xml.contains("")) - assert_true(xml.contains("")) + assert_true( + xml.contains(""), + ) assert_true(xml.contains("")) assert_true(xml.contains("")) assert_true(xml.contains("ACGT")) @@ -482,8 +481,7 @@ test "seqxml_document_to_xml_no_metadata" { ///| test "seqxml_roundtrip_basic" { - let original = - "\n" + + let original = "\n" + "\n" + " \n" + " ACGTACGT\n" + @@ -502,8 +500,7 @@ test "seqxml_roundtrip_basic" { ///| test "seqxml_roundtrip_with_metadata" { - let original = - "\n" + + let original = "\n" + "\n" + " \n" + " \n" + @@ -526,8 +523,7 @@ test "seqxml_roundtrip_with_metadata" { ///| test "seqxml_roundtrip_with_properties" { - let original = - "\n" + + let original = "\n" + "\n" + " \n" + " \n" + @@ -553,8 +549,7 @@ test "seqxml_roundtrip_with_properties" { ///| test "seqxml_roundtrip_multiple_entries" { - let original = - "\n" + + let original = "\n" + "\n" + " ACGT\n" + " AUGU\n" + @@ -753,8 +748,7 @@ test "seqxml_parse_empty_document" { ///| test "seqxml_parse_no_entries" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + "\n" @@ -767,8 +761,7 @@ test "seqxml_parse_no_entries" { ///| test "seqxml_parse_entry_without_description" { - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " ACGT\n" + @@ -786,8 +779,7 @@ test "seqxml_parse_entry_without_description" { ///| test "seqxml_parse_self_closing_entry" { // Self-closing has no sequence element. - let content = - "\n" + + let content = "\n" + "\n" + " \n" + "\n" @@ -803,8 +795,7 @@ test "seqxml_parse_self_closing_entry" { ///| test "seqxml_entity_escaping_in_description" { // Parsing should unescape XML entities in the desc attribute. - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " ACGT\n" + @@ -818,8 +809,7 @@ test "seqxml_entity_escaping_in_description" { ///| test "seqxml_entity_escaping_in_sequence" { // Parsing should unescape XML entities in the sequence text content. - let content = - "\n" + + let content = "\n" + "\n" + " \n" + " AC&T<G\n" + @@ -855,8 +845,7 @@ test "seqxml_entity_escaping_roundtrip" { description="Quote \" and amp & and lt <", ) let xml = @src.seqxml_entry_to_xml(entry) - let content = - "\n" + + let content = "\n" + "\n" + xml + "\n" diff --git a/test/moonbit/seurat_test.mbt b/test/moonbit/seurat_test.mbt index fc0dfb9c..8635b7f3 100644 --- a/test/moonbit/seurat_test.mbt +++ b/test/moonbit/seurat_test.mbt @@ -20,7 +20,7 @@ test "seurat_simple_pipeline" { let with_neighbors = @src.find_neighbors(with_pca) let with_clusters = @src.find_clusters(with_neighbors) let with_umap = @src.run_umap(with_pca) - + assert_eq(with_clusters.clusters.length(), 200) assert_eq(with_umap.umap.length(), 200) } @@ -31,7 +31,7 @@ test "normalize_total" { let normalized = @src.normalize_total(obj) assert_eq(normalized.data.length(), 500) assert_eq(normalized.data[0].length(), 200) - + let mut has_positive = false let mut i = 0 while i < normalized.data.length() && !has_positive { @@ -73,12 +73,12 @@ test "run_pca" { let with_hvg = @src.find_variable_features(normalized) let scaled = @src.scale_data(with_hvg) let with_pca = @src.run_pca(scaled) - + assert_true(with_pca.pca.length() > 0) assert_true(with_pca.pca[0].length() > 0) assert_eq(with_pca.pca.length(), 200) assert_true(with_pca.var_explained.length() > 0) - + let mut sum_var = 0.0 let mut i = 0 while i < with_pca.var_explained.length() { @@ -96,7 +96,7 @@ test "find_neighbors" { let scaled = @src.scale_data(with_hvg) let with_pca = @src.run_pca(scaled) let with_neighbors = @src.find_neighbors(with_pca) - + assert_eq(with_neighbors.neighbors.length(), 200) assert_true(with_neighbors.neighbors[0].length() > 0) } @@ -110,9 +110,9 @@ test "find_clusters" { let with_pca = @src.run_pca(scaled) let with_neighbors = @src.find_neighbors(with_pca) let with_clusters = @src.find_clusters(with_neighbors) - + assert_eq(with_clusters.clusters.length(), 200) - + let unique_clusters : Array[Int] = Array::new() let mut i = 0 while i < with_clusters.clusters.length() { @@ -142,7 +142,7 @@ test "run_umap" { let scaled = @src.scale_data(with_hvg) let with_pca = @src.run_pca(scaled) let with_umap = @src.run_umap(with_pca) - + assert_eq(with_umap.umap.length(), 200) assert_eq(with_umap.umap[0].length(), 2) } @@ -156,11 +156,11 @@ test "find_all_markers" { let with_pca = @src.run_pca(scaled) let with_neighbors = @src.find_neighbors(with_pca) let with_clusters = @src.find_clusters(with_neighbors) - + let markers = @src.find_all_markers(with_clusters) - + assert_true(markers.length() >= 0) - + if markers.length() > 0 { let m = markers[0] assert_true(m.gene.length() > 0) @@ -171,25 +171,27 @@ test "find_all_markers" { test "find_integration_anchors" { let ref_obj = @src.seurat_create_example_data() let query_obj = @src.seurat_create_example_data() - + let ref_normalized = @src.normalize_total(ref_obj) let ref_with_hvg = @src.find_variable_features(ref_normalized) let ref_scaled = @src.scale_data(ref_with_hvg) let ref_with_pca = @src.run_pca(ref_scaled) - + let query_normalized = @src.normalize_total(query_obj) let query_with_hvg = @src.find_variable_features(query_normalized) let query_scaled = @src.scale_data(query_with_hvg) let query_with_pca = @src.run_pca(query_scaled) - + assert_true(ref_with_pca.pca.length() > 0) assert_true(query_with_pca.pca.length() > 0) assert_eq(ref_with_pca.pca.length(), ref_obj.col_names.length()) assert_eq(query_with_pca.pca.length(), query_obj.col_names.length()) - + let dims = [0] - let anchors = @src.find_integration_anchors(ref_with_pca, query_with_pca, dims) - + let anchors = @src.find_integration_anchors( + ref_with_pca, query_with_pca, dims, + ) + assert_eq(anchors.reference_indices.length(), anchors.anchors.length()) assert_eq(anchors.query_indices.length(), anchors.anchors.length()) } @@ -198,26 +200,31 @@ test "find_integration_anchors" { test "integrate_data" { let ref_obj = @src.seurat_create_example_data() let query_obj = @src.seurat_create_example_data() - + let ref_normalized = @src.normalize_total(ref_obj) let ref_with_hvg = @src.find_variable_features(ref_normalized) let ref_scaled = @src.scale_data(ref_with_hvg) let ref_with_pca = @src.run_pca(ref_scaled) - + let query_normalized = @src.normalize_total(query_obj) let query_with_hvg = @src.find_variable_features(query_normalized) let query_scaled = @src.scale_data(query_with_hvg) let query_with_pca = @src.run_pca(query_scaled) - + let dims = [0, 1, 2, 3, 4] - let anchors = @src.find_integration_anchors(ref_with_pca, query_with_pca, dims) - + let anchors = @src.find_integration_anchors( + ref_with_pca, query_with_pca, dims, + ) + let integrated = @src.integrate_data(ref_obj, query_obj, anchors) - + assert_true(integrated.counts.length() > 0) - assert_true(integrated.col_names.length() == ref_obj.col_names.length() + query_obj.col_names.length()) + assert_true( + integrated.col_names.length() == + ref_obj.col_names.length() + query_obj.col_names.length(), + ) assert_eq(integrated.row_names.length(), integrated.counts.length()) - + let mut has_ref_prefix = false let mut has_query_prefix = false for name in integrated.col_names { diff --git a/test/moonbit/sff_io_test.mbt b/test/moonbit/sff_io_test.mbt index 200c1ee1..93b9ddd4 100644 --- a/test/moonbit/sff_io_test.mbt +++ b/test/moonbit/sff_io_test.mbt @@ -34,7 +34,9 @@ test "sff_read_new" { let qualities = [10, 20, 30, 25] let flowgram = [1.0, 0.0, 1.0, 0.0] let flow_index = [1, 3, 1, 3] - let read = @src.SffRead::new("READ001", "ACGT", qualities, flowgram, flow_index, 1, 4, 0, 0) + let read = @src.SffRead::new( + "READ001", "ACGT", qualities, flowgram, flow_index, 1, 4, 0, 0, + ) assert_eq(read.name, "READ001") assert_eq(read.bases, "ACGT") assert_eq(read.n_bases, 4) @@ -137,7 +139,10 @@ test "sff_encode_parse_roundtrip" { ///| test "sff_parse_invalid_magic" { // Create bytes with wrong magic number - let bytes = [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 40, 0, 4, 0, 4, 0, 1] + let bytes = [ + 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 40, + 0, 4, 0, 4, 0, 1, + ] let result = @src.sff_parse(bytes) assert_true(result is None) } diff --git a/test/moonbit/sgseq_test.mbt b/test/moonbit/sgseq_test.mbt index b45ca1e4..61c1ac1e 100644 --- a/test/moonbit/sgseq_test.mbt +++ b/test/moonbit/sgseq_test.mbt @@ -53,11 +53,7 @@ test "sg_junction_default_strand_count" { ///| test "sg_junction_intron_length" { - let j = @src.SGSeqJunction::new( - chr="chr1", - start=1000, - end=2000, - ) + let j = @src.SGSeqJunction::new(chr="chr1", start=1000, end=2000) // intron_length = end - start - 1 = 999 assert_eq(j.intron_length(), 999) } @@ -352,9 +348,7 @@ test "sg_detect_ri" { let junctions = [ @src.SGSeqJunction::new(chr="chr1", start=100, end=500, count=5), ] - let exons = [ - @src.SGSeqExon::new(chr="chr1", start=100, end=500, strand="+"), - ] + let exons = [@src.SGSeqExon::new(chr="chr1", start=100, end=500, strand="+")] let variants = @src.sgseq_detect_ri(junctions, exons) // Should detect an RI event (low junction count with overlapping exon) assert_true(variants.length() >= 1) @@ -373,7 +367,13 @@ test "sg_detect_all_variants" { // All variants should have valid type names for v in variants { let name = v.type_name() - assert_true(name == "SE" || name == "A5SS" || name == "A3SS" || name == "MXE" || name == "RI") + assert_true( + name == "SE" || + name == "A5SS" || + name == "A3SS" || + name == "MXE" || + name == "RI", + ) } } diff --git a/test/moonbit/shared_reference_alignment_test.mbt b/test/moonbit/shared_reference_alignment_test.mbt new file mode 100644 index 00000000..77b3e8e9 --- /dev/null +++ b/test/moonbit/shared_reference_alignment_test.mbt @@ -0,0 +1,591 @@ +///| +fn shared_reference_test_sequence( + name : String, + aligned : String, +) -> @src.SharedReferenceSequence { + @src.shared_reference_sequence(name, aligned) catch { + _ => abort("valid shared-reference sequence should build") + } +} + +///| +fn shared_reference_test_input( + aligned_reference : String, + queries : Array[@src.SharedReferenceSequence], +) -> @src.SharedReferenceInput { + @src.shared_reference_input( + "ACGT", + aligned_reference, + queries, + reference_name="reference", + ) catch { + _ => abort("valid shared-reference input should build") + } +} + +///| +fn shared_reference_test_sample() -> @src.SharedReferenceAlignment { + @src.shared_reference_alignment_sample() catch { + _ => abort("shared-reference sample should merge") + } +} + +///| +test "shared reference: query constructor derives ungapped sequence" { + let query = shared_reference_test_sequence("read", "AC-GT") + assert_eq(query.name, "read") + assert_eq(query.sequence, "ACGT") + assert_eq(query.aligned_sequence, "AC-GT") + assert_eq(query.start, 0) + assert_eq(query.end, 4) +} + +///| +test "shared reference: query constructor preserves local coordinates" { + let query = @src.shared_reference_sequence( + "read", + "AC-GT", + sequence="NNACGTNN", + start=2, + ) catch { + _ => abort("valid local query should build") + } + assert_eq(query.sequence, "NNACGTNN") + assert_eq(query.start, 2) + assert_eq(query.end, 6) +} + +///| +test "shared reference: query constructor rejects negative start" { + let failed = try { + ignore(@src.shared_reference_sequence("read", "ACGT", start=-1)) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: query constructor rejects sequence mismatch" { + let failed = try { + ignore( + @src.shared_reference_sequence("read", "ACGT", sequence="AACCAA", start=1), + ) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: query constructor rejects whitespace" { + let failed = try { + ignore(@src.shared_reference_sequence("read", "AC GT")) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: input validates and records reference span" { + let query = shared_reference_test_sequence("read", "AC-GT") + let input = shared_reference_test_input("AC-GT", [query]) + assert_eq(input.reference, "ACGT") + assert_eq(input.reference_start, 0) + assert_eq(input.reference_end, 4) + assert_eq(input.queries.length(), 1) +} + +///| +test "shared reference: input preserves reference metadata" { + let query = shared_reference_test_sequence("read", "ACGT") + let input = @src.shared_reference_input( + "ACGT", + "ACGT", + [query], + reference_name="plasmid", + reference_description="circular construct", + ) catch { + _ => abort("valid input should build") + } + assert_eq(input.reference_name, "plasmid") + assert_eq(input.reference_description, "circular construct") +} + +///| +test "shared reference: input copies caller query array" { + let queries = [shared_reference_test_sequence("read", "ACGT")] + let input = shared_reference_test_input("ACGT", queries) + queries.push(shared_reference_test_sequence("later", "ACGT")) + assert_eq(input.queries.length(), 1) +} + +///| +test "shared reference: input rejects empty reference" { + let query = shared_reference_test_sequence("read", "") + let failed = try { + ignore(@src.shared_reference_input("", "", [query])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: input rejects missing queries" { + let failed = try { + ignore(@src.shared_reference_input("ACGT", "ACGT", [])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: input rejects unequal row lengths" { + let query = shared_reference_test_sequence("read", "ACGT") + let failed = try { + ignore(@src.shared_reference_input("ACGT", "AC-GT", [query])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: input rejects aligned reference mismatch" { + let query = shared_reference_test_sequence("read", "AGGT") + let failed = try { + ignore(@src.shared_reference_input("ACGT", "AGGT", [query])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: official mixed PWA MSA example" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.num_sequences(), 4) + assert_eq(alignment.num_queries(), 3) + assert_eq(alignment.alignment_length(), 5) + assert_eq(alignment.row(0), Some("ACG-T")) + assert_eq(alignment.row(1), Some("AC--T")) + assert_eq(alignment.row(2), Some("ACGGT")) + assert_eq(alignment.row(3), Some("A---T")) +} + +///| +test "shared reference: sample synchronizes insertion slots" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.insertion_widths, [0, 0, 0, 1, 0]) + assert_eq(alignment.aligned_reference, "ACG-T") +} + +///| +test "shared reference: query ordering and metadata are stable" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.queries[0].name, "seq1") + assert_eq(alignment.queries[1].name, "seq2") + assert_eq(alignment.queries[2].name, "seq3") + assert_eq(alignment.queries[1].description, "one insertion") +} + +///| +test "shared reference: leading insertion is synchronized" { + let first = shared_reference_test_input("-ACGT", [ + shared_reference_test_sequence("inserted", "TACGT"), + ]) + let second = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("plain", "ACGT"), + ]) + let alignment = @src.alignments_with_same_reference([first, second]) catch { + _ => abort("leading insertion should merge") + } + assert_eq(alignment.aligned_reference, "-ACGT") + assert_eq(alignment.queries[0].aligned_sequence, "TACGT") + assert_eq(alignment.queries[1].aligned_sequence, "-ACGT") +} + +///| +test "shared reference: trailing insertion is synchronized" { + let first = shared_reference_test_input("ACGT--", [ + shared_reference_test_sequence("inserted", "ACGTAA"), + ]) + let second = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("plain", "ACGT"), + ]) + let alignment = @src.alignments_with_same_reference([first, second]) catch { + _ => abort("trailing insertion should merge") + } + assert_eq(alignment.aligned_reference, "ACGT--") + assert_eq(alignment.queries[1].aligned_sequence, "ACGT--") +} + +///| +test "shared reference: widest insertion controls slot width" { + let short = shared_reference_test_input("AC-GT", [ + shared_reference_test_sequence("short", "ACTGT"), + ]) + let long = shared_reference_test_input("AC---GT", [ + shared_reference_test_sequence("long", "ACTTTGT"), + ]) + let alignment = @src.alignments_with_same_reference([short, long]) catch { + _ => abort("different insertion widths should merge") + } + assert_eq(alignment.aligned_reference, "AC---GT") + assert_eq(alignment.queries[0].aligned_sequence, "ACT--GT") + assert_eq(alignment.queries[1].aligned_sequence, "ACTTTGT") +} + +///| +test "shared reference: insertions at distinct boundaries are retained" { + let left = shared_reference_test_input("A-CGT", [ + shared_reference_test_sequence("left", "ATCGT"), + ]) + let right = shared_reference_test_input("ACG-T", [ + shared_reference_test_sequence("right", "ACGGT"), + ]) + let alignment = @src.alignments_with_same_reference([left, right]) catch { + _ => abort("distinct insertion slots should merge") + } + assert_eq(alignment.aligned_reference, "A-CG-T") + assert_eq(alignment.queries[0].aligned_sequence, "ATCG-T") + assert_eq(alignment.queries[1].aligned_sequence, "A-CGGT") +} + +///| +test "shared reference: multi-query input preserves internal gap columns" { + let first = shared_reference_test_input("AC--GT", [ + shared_reference_test_sequence("one", "ACT-GT"), + shared_reference_test_sequence("two", "AC-TGT"), + ]) + let second = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("plain", "ACGT"), + ]) + let alignment = @src.alignments_with_same_reference([first, second]) catch { + _ => abort("multi-query input should merge") + } + assert_eq(alignment.queries[0].aligned_sequence, "ACT-GT") + assert_eq(alignment.queries[1].aligned_sequence, "AC-TGT") + assert_eq(alignment.queries[2].aligned_sequence, "AC--GT") +} + +///| +test "shared reference: reference comparison is case insensitive" { + let upper = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("upper", "ACGT"), + ]) + let lower_query = shared_reference_test_sequence("lower", "acgt") + let lower = @src.shared_reference_input("acgt", "acgt", [lower_query]) catch { + _ => abort("lowercase input should build") + } + let alignment = @src.alignments_with_same_reference([upper, lower]) catch { + _ => abort("case-insensitive references should merge") + } + assert_eq(alignment.num_queries(), 2) +} + +///| +test "shared reference: merge rejects reference length mismatch" { + let first = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("one", "ACGT"), + ]) + let query = shared_reference_test_sequence("two", "ACGTA") + let second = @src.shared_reference_input("ACGTA", "ACGTA", [query]) catch { + _ => abort("second input should build") + } + let failed = try { + ignore(@src.alignments_with_same_reference([first, second])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: merge rejects reference content mismatch" { + let first = shared_reference_test_input("ACGT", [ + shared_reference_test_sequence("one", "ACGT"), + ]) + let query = shared_reference_test_sequence("two", "AGGT") + let second = @src.shared_reference_input("AGGT", "AGGT", [query]) catch { + _ => abort("second input should build") + } + let failed = try { + ignore(@src.alignments_with_same_reference([first, second])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: merge rejects inconsistent reference coordinates" { + let query1 = shared_reference_test_sequence("one", "ACGT") + let first = @src.shared_reference_input( + "AACGTT", + "ACGT", + [query1], + reference_start=1, + ) catch { + _ => abort("first local input should build") + } + let query2 = shared_reference_test_sequence("two", "AACG") + let second = @src.shared_reference_input( + "AACGTT", + "AACG", + [query2], + reference_start=0, + ) catch { + _ => abort("second local input should build") + } + let failed = try { + ignore(@src.alignments_with_same_reference([first, second])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: merge rejects empty input list" { + let failed = try { + ignore(@src.alignments_with_same_reference([])) + false + } catch { + SharedReferenceAlignmentError(_) => true + } + assert_true(failed) +} + +///| +test "shared reference: PairwiseAlignment adapter preserves global rows" { + let pairwise = @src.pairaligner_align("ACGT", "ACGGT") + let input = @src.shared_reference_input_from_pairwise( + pairwise, + reference_name="ref", + query_name="read", + ) catch { + _ => abort("global pairwise alignment should adapt") + } + assert_eq(input.reference, "ACGT") + assert_eq(input.reference_name, "ref") + assert_eq(input.queries[0].name, "read") + assert_eq(input.reference_start, 0) + assert_eq(input.reference_end, 4) +} + +///| +test "shared reference: PairwiseAlignment adapter preserves local offsets" { + let config = @src.PairwiseAlignerConfig::default_dna().set_mode( + @src.pairaligner_local(), + ) + let pairwise = @src.pairaligner_align("NNACGTNN", "TTACGTTT", config~) + let input = @src.shared_reference_input_from_pairwise( + pairwise, + query_name="local", + ) catch { + _ => abort("local pairwise alignment should adapt") + } + assert_eq(input.reference_start, 2) + assert_eq(input.reference_end, 6) + assert_eq(input.queries[0].start, 2) + assert_eq(input.queries[0].end, 6) +} + +///| +test "shared reference: row access includes reference and bounds checks" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.row_name(0), Some("reference")) + assert_eq(alignment.row_name(2), Some("seq2")) + assert_eq(alignment.row(-1), None) + assert_eq(alignment.row(4), None) + assert_eq(alignment.row_name(4), None) +} + +///| +test "shared reference: column access returns all rows" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.column(0), Some("AAAA")) + assert_eq(alignment.column(2), Some("G-G-")) + assert_eq(alignment.column(3), Some("--G-")) + assert_eq(alignment.column(-1), None) + assert_eq(alignment.column(5), None) +} + +///| +test "shared reference: reference column mapping skips insertion columns" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.reference_to_column(0), Some(0)) + assert_eq(alignment.reference_to_column(2), Some(2)) + assert_eq(alignment.reference_to_column(3), Some(4)) + assert_eq(alignment.column_to_reference(3), None) + assert_eq(alignment.column_to_reference(4), Some(3)) +} + +///| +test "shared reference: query column mapping tracks insertions" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.query_to_column(1, 2), Some(2)) + assert_eq(alignment.query_to_column(1, 3), Some(3)) + assert_eq(alignment.query_to_column(1, 4), Some(4)) + assert_eq(alignment.column_to_query(1, 3), Some(3)) + assert_eq(alignment.column_to_query(0, 3), None) +} + +///| +test "shared reference: query to reference maps insertion to none" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.query_to_reference(1, 2), Some(2)) + assert_eq(alignment.query_to_reference(1, 3), None) + assert_eq(alignment.query_to_reference(1, 4), Some(3)) +} + +///| +test "shared reference: reference to query maps deletion to none" { + let alignment = shared_reference_test_sample() + assert_eq(alignment.reference_to_query(0, 0), Some(0)) + assert_eq(alignment.reference_to_query(0, 2), None) + assert_eq(alignment.reference_to_query(0, 3), Some(2)) +} + +///| +test "shared reference: counts classify deletion" { + let alignment = shared_reference_test_sample() + let counts = alignment.counts(0).unwrap() + assert_eq(counts.identities, 3) + assert_eq(counts.mismatches, 0) + assert_eq(counts.insertions, 0) + assert_eq(counts.deletions, 1) + assert_eq(counts.identity_percent(), 100.0) +} + +///| +test "shared reference: counts classify insertion" { + let alignment = shared_reference_test_sample() + let counts = alignment.counts(1).unwrap() + assert_eq(counts.identities, 4) + assert_eq(counts.insertions, 1) + assert_eq(counts.deletions, 0) + assert_eq(counts.aligned_pairs, 4) + assert_eq(alignment.counts(-1), None) +} + +///| +test "shared reference: find query returns metadata" { + let alignment = shared_reference_test_sample() + match alignment.find_query("seq2") { + Some(query) => assert_eq(query.description, "one insertion") + None => abort("seq2 should be present") + } + assert_eq(alignment.find_query("missing"), None) +} + +///| +test "shared reference: converts to MultipleSeqAlignment" { + let alignment = shared_reference_test_sample() + let msa = alignment.to_msa() catch { + _ => abort("valid shared-reference alignment should convert") + } + assert_eq(msa.num_records(), 4) + assert_eq(msa.get_alignment_length(), 5) + assert_eq(msa.records[0].id, "reference") + assert_eq(msa.records[1].description, "one deletion") + assert_eq(msa.records[2].seq.to_string(), "ACGGT") +} + +///| +test "shared reference: aligned FASTA preserves metadata" { + let alignment = shared_reference_test_sample() + assert_eq( + alignment.to_fasta(), + ">reference shared reference\nACG-T\n" + + ">seq1 one deletion\nAC--T\n" + + ">seq2 one insertion\nACGGT\n" + + ">seq3 two deletions\nA---T\n", + ) +} + +///| +test "shared reference: summary reports dimensions and span" { + let alignment = shared_reference_test_sample() + assert_eq( + alignment.summary(), + "Shared-reference alignment: 4 sequences, 5 columns, reference 0:4", + ) +} + +///| +test "shared reference: local alignment uses absolute coordinates" { + let first_query = @src.shared_reference_sequence( + "one", + "AC-GT", + sequence="NNACGTNN", + start=2, + ) catch { + _ => abort("local query should build") + } + let first = @src.shared_reference_input( + "NNAACGTNN", + "AC-GT", + [first_query], + reference_start=3, + ) catch { + _ => abort("local input should build") + } + let second_query = @src.shared_reference_sequence( + "two", + "ACGGT", + sequence="TTACGGTTT", + start=2, + ) catch { + _ => abort("inserted local query should build") + } + let second = @src.shared_reference_input( + "NNAACGTNN", + "ACG-T", + [second_query], + reference_start=3, + ) catch { + _ => abort("second local input should build") + } + let alignment = @src.alignments_with_same_reference([first, second]) catch { + _ => abort("local inputs should merge") + } + assert_eq(alignment.reference_start, 3) + assert_eq(alignment.reference_end, 7) + assert_eq(alignment.reference_to_column(6), Some(5)) + assert_eq(alignment.query_to_reference(1, 5), None) + assert_eq(alignment.query_to_reference(1, 6), Some(6)) +} + +///| +test "shared reference: supports protein residues and mismatches" { + let first_query = shared_reference_test_sequence("one", "MKTL") + let first = @src.shared_reference_input("MKTL", "MKTL", [first_query]) catch { + _ => abort("protein input should build") + } + let second_query = shared_reference_test_sequence("two", "MRATL") + let second = @src.shared_reference_input("MKTL", "MK-TL", [second_query]) catch { + _ => abort("inserted protein input should build") + } + let alignment = @src.alignments_with_same_reference([first, second]) catch { + _ => abort("protein alignments should merge") + } + let counts = alignment.counts(1).unwrap() + assert_eq(counts.identities, 3) + assert_eq(counts.mismatches, 1) + assert_eq(counts.insertions, 1) +} diff --git a/test/moonbit/single_r_advanced_test.mbt b/test/moonbit/single_r_advanced_test.mbt new file mode 100644 index 00000000..399aa56b --- /dev/null +++ b/test/moonbit/single_r_advanced_test.mbt @@ -0,0 +1,985 @@ +///| +fn single_r_adv_test_example() -> ( + @src.SingleRAdvancedReference, + @src.SingleRAdvancedData, +) { + @src.single_r_advanced_example() catch { + _ => abort("valid SingleR advanced example should build") + } +} + +///| +fn single_r_adv_test_result() -> @src.SingleRAdvancedResult { + let (reference, data) = single_r_adv_test_example() + @src.single_r_advanced(data, reference) catch { + _ => abort("valid SingleR advanced example should classify") + } +} + +///| +fn single_r_adv_test_sce() -> @src.SingleCellExperiment { + let (_, data) = single_r_adv_test_example() + let experiment = @src.SingleCellExperiment::new( + data.expression.map(fn(row) { row.copy() }), + data.gene_names.copy(), + data.cell_names.copy(), + ) + experiment.assays["logcounts"] = data.expression.map(fn(row) { row.copy() }) + experiment.col_data["batch"] = ["x", "x", "y", "y", "y"] + experiment.row_data["symbol"] = data.gene_names.copy() + experiment.reduced_dims["PCA"] = [ + [0.0, 1.0], + [1.0, 0.0], + [0.5, 0.5], + [0.2, 0.8], + [0.8, 0.2], + ] + experiment.metadata["source"] = "test" + experiment.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[1.0, 2.0, 3.0, 4.0, 5.0]], + ["spike1"], + data.cell_names.copy(), + ) + experiment +} + +///| +fn single_r_adv_test_second_reference() -> @src.SingleRAdvancedReference { + let (reference, _) = single_r_adv_test_example() + @src.SingleRAdvancedReference::create( + reference.expression, + [ + "T lymphocyte", "T lymphocyte", "B lymphocyte", "B lymphocyte", "Myeloid", + "Myeloid", "Cytotoxic", "Cytotoxic", + ], + gene_names=reference.gene_names, + sample_names=reference.sample_names, + reference_name="immune_reference_two", + ) catch { + _ => abort("valid second SingleR reference should build") + } +} + +///| +test "SingleR advanced default configuration matches upstream controls" { + let config = @src.SingleRAdvancedConfig::default() + assert_eq(config.quantile, 0.8) + assert_true(config.fine_tune) + assert_eq(config.tune_threshold, 0.05) + assert_true(config.prune) + assert_eq(config.nmads, 3.0) + assert_eq(config.minimum_delta_next, 0.0) + assert_eq(config.marker_count, 0) + assert_eq(config.combine_marker_count, 10) +} + +///| +test "SingleR advanced configuration preserves explicit controls" { + let config = @src.SingleRAdvancedConfig::create( + quantile=0.6, + fine_tune=false, + tune_threshold=0.1, + prune=false, + nmads=2.0, + minimum_delta_median=0.05, + minimum_delta_next=0.02, + marker_count=3, + combine_marker_count=4, + ) catch { + _ => abort("valid SingleR advanced configuration should build") + } + assert_eq(config.quantile, 0.6) + assert_eq(config.fine_tune, false) + assert_eq(config.tune_threshold, 0.1) + assert_eq(config.prune, false) + assert_eq(config.nmads, 2.0) + assert_eq(config.minimum_delta_median, 0.05) + assert_eq(config.minimum_delta_next, 0.02) + assert_eq(config.marker_count, 3) + assert_eq(config.combine_marker_count, 4) +} + +///| +test "SingleR advanced rejects invalid quantiles" { + let lower = try { + ignore(@src.SingleRAdvancedConfig::create(quantile=-0.1)) + false + } catch { + _ => true + } + let upper = try { + ignore(@src.SingleRAdvancedConfig::create(quantile=1.1)) + false + } catch { + _ => true + } + assert_true(lower) + assert_true(upper) +} + +///| +test "SingleR advanced rejects invalid tuning and pruning controls" { + let tuning = try { + ignore(@src.SingleRAdvancedConfig::create(tune_threshold=-0.1)) + false + } catch { + _ => true + } + let nmads = try { + ignore(@src.SingleRAdvancedConfig::create(nmads=-1.0)) + false + } catch { + _ => true + } + assert_true(tuning) + assert_true(nmads) +} + +///| +test "SingleR advanced rejects invalid marker controls" { + let markers = try { + ignore(@src.SingleRAdvancedConfig::create(marker_count=-1)) + false + } catch { + _ => true + } + let combine = try { + ignore(@src.SingleRAdvancedConfig::create(combine_marker_count=0)) + false + } catch { + _ => true + } + assert_true(markers) + assert_true(combine) +} + +///| +test "SingleR advanced reference uses gene by sample orientation" { + let (reference, _) = single_r_adv_test_example() + assert_eq(reference.n_genes, 12) + assert_eq(reference.n_samples, 8) + assert_eq(reference.expression[0].length(), 8) + assert_eq(reference.gene_names[0], "CD3D") + assert_eq(reference.sample_names[7], "NK2") +} + +///| +test "SingleR advanced preserves first occurrence label order" { + let (reference, _) = single_r_adv_test_example() + assert_eq(reference.label_names, ["T cell", "B cell", "Monocyte", "NK cell"]) +} + +///| +test "SingleR advanced reference defensively copies inputs" { + let expression = [[1.0, 2.0], [3.0, 4.0]] + let labels = ["A", "B"] + let genes = ["g1", "g2"] + let reference = @src.SingleRAdvancedReference::create( + expression, + labels, + gene_names=genes, + ) catch { + _ => abort("valid reference should build") + } + expression[0][0] = 99.0 + labels[0] = "changed" + genes[0] = "changed" + assert_eq(reference.expression[0][0], 1.0) + assert_eq(reference.labels[0], "A") + assert_eq(reference.gene_names[0], "g1") +} + +///| +test "SingleR advanced accepts finite negative normalized expression" { + let reference = @src.SingleRAdvancedReference::create( + [[-2.0, 1.0], [0.0, -1.0]], + ["A", "B"], + ) catch { + _ => abort("finite normalized expression may be negative") + } + assert_eq(reference.expression[0][0], -2.0) +} + +///| +test "SingleR advanced data uses gene by cell orientation" { + let (_, data) = single_r_adv_test_example() + assert_eq(data.n_genes, 12) + assert_eq(data.n_cells, 5) + assert_eq(data.expression[0].length(), 5) + assert_eq(data.cell_names[4], "cell_ambiguous") +} + +///| +test "SingleR advanced data defensively copies inputs" { + let expression = [[1.0, 2.0], [3.0, 4.0]] + let data = @src.SingleRAdvancedData::create( + expression, + gene_names=["g1", "g2"], + cell_names=["c1", "c2"], + ) catch { + _ => abort("valid data should build") + } + expression[0][0] = 99.0 + let copied = data.copy_expression() + copied[0][0] = 77.0 + assert_eq(data.expression[0][0], 1.0) +} + +///| +test "SingleR advanced constructors reject empty matrices" { + let reference = try { + ignore(@src.SingleRAdvancedReference::create([], [])) + false + } catch { + _ => true + } + let data = try { + ignore(@src.SingleRAdvancedData::create([[]])) + false + } catch { + _ => true + } + assert_true(reference) + assert_true(data) +} + +///| +test "SingleR advanced constructors reject ragged matrices" { + let reference = try { + ignore( + @src.SingleRAdvancedReference::create([[1.0, 2.0], [3.0]], ["A", "B"]), + ) + false + } catch { + _ => true + } + let data = try { + ignore(@src.SingleRAdvancedData::create([[1.0, 2.0], [3.0]])) + false + } catch { + _ => true + } + assert_true(reference) + assert_true(data) +} + +///| +test "SingleR advanced constructors reject non-finite values" { + let reference = try { + ignore( + @src.SingleRAdvancedReference::create([[1.0, 1.0e301], [2.0, 3.0]], [ + "A", "B", + ]), + ) + false + } catch { + _ => true + } + let data = try { + ignore(@src.SingleRAdvancedData::create([[0.0 / 0.0, 1.0]])) + false + } catch { + _ => true + } + assert_true(reference) + assert_true(data) +} + +///| +test "SingleR advanced reference validates label dimensions and content" { + let dimensions = try { + ignore( + @src.SingleRAdvancedReference::create([[1.0, 2.0], [3.0, 4.0]], ["A"]), + ) + false + } catch { + _ => true + } + let content = try { + ignore( + @src.SingleRAdvancedReference::create([[1.0, 2.0], [3.0, 4.0]], ["A", " "]), + ) + false + } catch { + _ => true + } + assert_true(dimensions) + assert_true(content) +} + +///| +test "SingleR advanced constructors validate name dimensions" { + let genes = try { + ignore( + @src.SingleRAdvancedData::create([[1.0, 2.0], [3.0, 4.0]], gene_names=[ + "g1", + ]), + ) + false + } catch { + _ => true + } + let samples = try { + ignore( + @src.SingleRAdvancedReference::create( + [[1.0, 2.0], [3.0, 4.0]], + ["A", "B"], + sample_names=["s1"], + ), + ) + false + } catch { + _ => true + } + assert_true(genes) + assert_true(samples) +} + +///| +test "SingleR advanced constructors reject duplicate names" { + let genes = try { + ignore( + @src.SingleRAdvancedData::create([[1.0, 2.0], [3.0, 4.0]], gene_names=[ + "g", "g", + ]), + ) + false + } catch { + _ => true + } + let cells = try { + ignore( + @src.SingleRAdvancedData::create([[1.0, 2.0], [3.0, 4.0]], cell_names=[ + "c", "c", + ]), + ) + false + } catch { + _ => true + } + assert_true(genes) + assert_true(cells) +} + +///| +test "SingleR advanced Spearman handles positive negative and tied ranks" { + let positive = @src.single_r_advanced_spearman([1.0, 2.0, 2.0, 4.0], [ + 10.0, 20.0, 20.0, 40.0, + ]) catch { + _ => abort("valid tied Spearman vectors should work") + } + let negative = @src.single_r_advanced_spearman([1.0, 2.0, 3.0, 4.0], [ + 4.0, 3.0, 2.0, 1.0, + ]) catch { + _ => abort("valid Spearman vectors should work") + } + assert_true(positive > 0.999) + assert_true(negative < -0.999) +} + +///| +test "SingleR advanced Spearman validates dimensions and finite values" { + let dimensions = try { + ignore(@src.single_r_advanced_spearman([1.0], [1.0, 2.0])) + false + } catch { + _ => true + } + let finite = try { + ignore(@src.single_r_advanced_spearman([1.0, 1.0e301], [1.0, 2.0])) + false + } catch { + _ => true + } + assert_true(dimensions) + assert_true(finite) +} + +///| +test "SingleR advanced training aligns genes in test order" { + let (reference, _) = single_r_adv_test_example() + let training = @src.single_r_advanced_train( + reference, + ["GNLY", "missing", "CD3D", "MS4A1"], + marker_count=1, + ) catch { + _ => abort("gene intersection should train") + } + assert_eq(training.common_gene_names, ["GNLY", "CD3D", "MS4A1"]) + assert_eq(training.test_gene_indices, [0, 2, 3]) + assert_eq(training.reference_gene_indices, [11, 0, 3]) +} + +///| +test "SingleR advanced training honors gene restrictions" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train( + reference, + data.gene_names, + restrict_genes=["CD3D", "MS4A1", "LYZ", "NKG7"], + marker_count=1, + ) catch { + _ => abort("restricted training should work") + } + assert_eq(training.common_gene_names, ["CD3D", "MS4A1", "LYZ", "NKG7"]) +} + +///| +test "SingleR advanced training rejects insufficient shared genes" { + let (reference, _) = single_r_adv_test_example() + let raised = try { + ignore(@src.single_r_advanced_train(reference, ["missing", "CD3D"])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced training computes automatic marker count" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train(reference, data.gene_names) catch { + _ => abort("automatic marker training should work") + } + assert_true(training.marker_count > 0) + assert_true(training.common_markers.length() >= 2) +} + +///| +test "SingleR advanced classic markers are directional" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train( + reference, + data.gene_names, + marker_count=2, + ) catch { + _ => abort("marker training should work") + } + let t_markers = training.markers_between("T cell", "B cell") catch { + _ => abort("known labels should have markers") + } + let b_markers = training.markers_between("B cell", "T cell") catch { + _ => abort("known labels should have markers") + } + assert_true(t_markers.contains("CD3D") || t_markers.contains("CD3E")) + assert_true(b_markers.contains("MS4A1") || b_markers.contains("CD79A")) +} + +///| +test "SingleR advanced self comparison has no markers" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train(reference, data.gene_names) catch { + _ => abort("marker training should work") + } + assert_eq( + training.markers_between("T cell", "T cell") catch { + _ => abort("known labels should query") + }, + [], + ) +} + +///| +test "SingleR advanced marker query rejects unknown labels" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train(reference, data.gene_names) catch { + _ => abort("marker training should work") + } + let raised = try { + ignore(training.markers_between("missing", "T cell")) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced classification locks training gene order" { + let (reference, data) = single_r_adv_test_example() + let training = @src.single_r_advanced_train(reference, data.gene_names) catch { + _ => abort("training should work") + } + let reordered = @src.SingleRAdvancedData::create( + data.expression, + gene_names=[ + "CD3E", "CD3D", "TRAC", "MS4A1", "CD79A", "CD74", "LYZ", "S100A8", "FCGR3A", + "NKG7", "KLRD1", "GNLY", + ], + cell_names=data.cell_names, + ) catch { + _ => abort("reordered data should build") + } + let raised = try { + ignore(@src.single_r_advanced_classify(reordered, training)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced classification reports cell by label scores" { + let result = single_r_adv_test_result() + assert_eq(result.n_cells(), 5) + assert_eq(result.scores.length(), 5) + assert_eq(result.scores[0].length(), 4) + assert_eq(result.label_names.length(), 4) +} + +///| +test "SingleR advanced classifies canonical immune profiles" { + let result = single_r_adv_test_result() + assert_eq(result.labels[0], "T cell") + assert_eq(result.labels[1], "B cell") + assert_eq(result.labels[2], "Monocyte") + assert_eq(result.labels[3], "NK cell") +} + +///| +test "SingleR advanced preserves pre-fine-tuning assignments" { + let result = single_r_adv_test_result() + assert_eq(result.first_labels[0], "T cell") + assert_eq(result.first_labels[1], "B cell") + assert_eq(result.first_labels[2], "Monocyte") + assert_eq(result.first_labels[3], "NK cell") +} + +///| +test "SingleR advanced quantile controls label aggregation" { + let (reference, data) = single_r_adv_test_example() + let low = @src.SingleRAdvancedConfig::create( + quantile=0.0, + fine_tune=false, + prune=false, + ) catch { + _ => abort("low quantile configuration should build") + } + let high = @src.SingleRAdvancedConfig::create( + quantile=1.0, + fine_tune=false, + prune=false, + ) catch { + _ => abort("high quantile configuration should build") + } + let low_result = @src.single_r_advanced(data, reference, config=low) catch { + _ => abort("low quantile classification should work") + } + let high_result = @src.single_r_advanced(data, reference, config=high) catch { + _ => abort("high quantile classification should work") + } + assert_true(high_result.scores[0][0] >= low_result.scores[0][0]) + assert_true(high_result.scores[1][1] >= low_result.scores[1][1]) +} + +///| +test "SingleR advanced fine tuning is active for ambiguous candidates" { + let (reference, data) = single_r_adv_test_example() + let config = @src.SingleRAdvancedConfig::create( + tune_threshold=2.0, + prune=false, + ) catch { + _ => abort("wide fine-tuning configuration should build") + } + let result = @src.single_r_advanced(data, reference, config~) catch { + _ => abort("fine-tuned classification should work") + } + assert_true(result.fine_tuned[4]) + assert_true(result.delta_next[4] >= 0.0) +} + +///| +test "SingleR advanced can disable fine tuning" { + let (reference, data) = single_r_adv_test_example() + let config = @src.SingleRAdvancedConfig::create(fine_tune=false, prune=false) catch { + _ => abort("no-fine-tune configuration should build") + } + let result = @src.single_r_advanced(data, reference, config~) catch { + _ => abort("classification without fine tuning should work") + } + for value in result.fine_tuned { + assert_eq(value, false) + } + assert_eq(result.labels, result.first_labels) +} + +///| +test "SingleR advanced delta from median matches score matrix" { + let result = single_r_adv_test_result() + let label = result.label_index(result.labels[0]) + let sorted = result.scores[0] + let median = if sorted.length() == 4 { + let values = sorted.copy() + for index in 1.. 0 && values[position - 1] > current { + values[position] = values[position - 1] + position = position - 1 + } + values[position] = current + } + (values[1] + values[2]) / 2.0 + } else { + 0.0 + } + assert_true( + (result.delta_median[0] - (result.scores[0][label] - median)).abs() < + 1.0e-12, + ) +} + +///| +test "SingleR advanced hard delta pruning uses NA semantics" { + let (reference, data) = single_r_adv_test_example() + let config = @src.SingleRAdvancedConfig::create(minimum_delta_next=3.0) catch { + _ => abort("strict pruning configuration should build") + } + let result = @src.single_r_advanced(data, reference, config~) catch { + _ => abort("strict pruning classification should work") + } + assert_eq(result.n_pruned(), data.n_cells) + for value in result.pruned_labels { + assert_true(value is None) + } +} + +///| +test "SingleR advanced can disable pruning" { + let (reference, data) = single_r_adv_test_example() + let config = @src.SingleRAdvancedConfig::create(prune=false) catch { + _ => abort("unpruned configuration should build") + } + let result = @src.single_r_advanced(data, reference, config~) catch { + _ => abort("unpruned classification should work") + } + assert_eq(result.n_pruned(), 0) + for value in result.pruned_labels { + assert_true(value is Some(_)) + } +} + +///| +test "SingleR advanced computes per-label MAD thresholds" { + let result = single_r_adv_test_result() + let mut observed = 0 + for threshold in result.pruning_thresholds { + if threshold is Some(_) { + observed = observed + 1 + } + } + assert_true(observed > 0) +} + +///| +test "SingleR advanced result query helpers are consistent" { + let result = single_r_adv_test_result() + assert_eq(result.label_index("T cell"), 0) + assert_eq(result.label_index("missing"), -1) + assert_true(result.assigned_score(0) > 0.0) + let summary = result.summary() + assert_eq(summary["total"], 5) + assert_eq( + summary["T cell"] + + summary["B cell"] + + summary["Monocyte"] + + summary["NK cell"], + 5, + ) +} + +///| +test "SingleR advanced assigned score validates cell indices" { + let result = single_r_adv_test_result() + let raised = try { + ignore(result.assigned_score(99)) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced cluster annotation sums member profiles" { + let (reference, data) = single_r_adv_test_example() + let config = @src.SingleRAdvancedConfig::create(prune=false) catch { + _ => abort("cluster configuration should build") + } + let result = @src.single_r_advanced_clusters( + data, + ["lymphoid", "lymphoid", "myeloid", "cytotoxic", "mixed"], + reference, + config~, + ) catch { + _ => abort("cluster annotation should work") + } + assert_eq(result.cell_names, ["lymphoid", "myeloid", "cytotoxic", "mixed"]) + assert_eq(result.n_cells(), 4) + assert_eq(result.labels[1], "Monocyte") + assert_eq(result.labels[2], "NK cell") +} + +///| +test "SingleR advanced cluster annotation preserves first occurrence order" { + let (reference, data) = single_r_adv_test_example() + let result = @src.single_r_advanced_clusters( + data, + ["z", "a", "z", "b", "a"], + reference, + ) catch { + _ => abort("cluster annotation should work") + } + assert_eq(result.cell_names, ["z", "a", "b"]) +} + +///| +test "SingleR advanced cluster annotation validates assignments" { + let (reference, data) = single_r_adv_test_example() + let dimensions = try { + ignore(@src.single_r_advanced_clusters(data, ["a"], reference)) + false + } catch { + _ => true + } + let names = try { + ignore( + @src.single_r_advanced_clusters(data, ["a", "a", "b", "b", ""], reference), + ) + false + } catch { + _ => true + } + assert_true(dimensions) + assert_true(names) +} + +///| +test "SingleR advanced multi-reference recomputation returns comparable scores" { + let (reference, data) = single_r_adv_test_example() + let second = single_r_adv_test_second_reference() + let result = @src.single_r_advanced_combine(data, [reference, second]) catch { + _ => abort("multi-reference classification should work") + } + assert_eq(result.n_cells(), 5) + assert_eq(result.scores.length(), 5) + assert_eq(result.scores[0].length(), 2) + assert_eq(result.per_reference.length(), 2) + assert_true(result.marker_gene_names.length() >= 2) +} + +///| +test "SingleR advanced multi-reference records selected provenance" { + let (reference, data) = single_r_adv_test_example() + let second = single_r_adv_test_second_reference() + let result = @src.single_r_advanced_combine(data, [reference, second]) catch { + _ => abort("multi-reference classification should work") + } + assert_eq(result.references, ["immune_reference", "immune_reference_two"]) + for selected in result.selected_references { + assert_true(result.references.contains(selected)) + } + let summary = result.reference_summary() + assert_eq(summary["immune_reference"] + summary["immune_reference_two"], 5) +} + +///| +test "SingleR advanced multi-reference rejects too few references" { + let (reference, data) = single_r_adv_test_example() + let raised = try { + ignore(@src.single_r_advanced_combine(data, [reference])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced multi-reference rejects duplicate names" { + let (reference, data) = single_r_adv_test_example() + let duplicate = @src.SingleRAdvancedReference::create( + reference.expression, + reference.labels, + gene_names=reference.gene_names, + sample_names=reference.sample_names, + reference_name=reference.reference_name, + ) catch { + _ => abort("duplicate-name reference should build independently") + } + let raised = try { + ignore(@src.single_r_advanced_combine(data, [reference, duplicate])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced multi-reference requires global gene intersection" { + let (reference, data) = single_r_adv_test_example() + let other = @src.SingleRAdvancedReference::create( + [[1.0, 2.0], [2.0, 1.0]], + ["X", "Y"], + gene_names=["other1", "other2"], + reference_name="other", + ) catch { + _ => abort("disjoint reference should build") + } + let raised = try { + ignore(@src.single_r_advanced_combine(data, [reference, other])) + false + } catch { + _ => true + } + assert_true(raised) +} + +///| +test "SingleR advanced SingleCellExperiment writes annotation columns" { + let (reference, _) = single_r_adv_test_example() + let output = @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + output_prefix="singler215", + ) catch { + _ => abort("SingleR SCE annotation should work") + } + assert_true(output.experiment.col_data.contains("singler215.labels")) + assert_true(output.experiment.col_data.contains("singler215.pruned.labels")) + assert_true(output.experiment.col_data.contains("singler215.score")) + assert_true(output.experiment.col_data.contains("singler215.delta.next")) + assert_true(output.experiment.col_data.contains("singler215.delta.median")) + assert_eq(output.experiment.col_data["singler215.labels"][0], "T cell") +} + +///| +test "SingleR advanced SingleCellExperiment records provenance metadata" { + let (reference, _) = single_r_adv_test_example() + let output = @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + output_prefix="singler215", + ) catch { + _ => abort("SingleR SCE annotation should work") + } + assert_eq(output.experiment.metadata["singler215.assay"], "logcounts") + assert_eq( + output.experiment.metadata["singler215.reference"], + "immune_reference", + ) + assert_eq(output.experiment.metadata["singler215.quantile"], "0.8") + assert_eq(output.experiment.metadata["singler215.mode"], "cell") + assert_eq(output.experiment.metadata["source"], "test") +} + +///| +test "SingleR advanced SingleCellExperiment does not mutate input" { + let (reference, _) = single_r_adv_test_example() + let input = single_r_adv_test_sce() + let output = @src.single_r_advanced_sce( + input, + reference, + output_prefix="singler215", + ) catch { + _ => abort("SingleR SCE annotation should work") + } + assert_eq(input.col_data.contains("singler215.labels"), false) + output.experiment.assays["logcounts"][0][0] = 999.0 + output.experiment.row_data["symbol"][0] = "changed" + output.experiment.reduced_dims["PCA"][0][0] = 999.0 + output.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 999.0 + assert_true(input.assays["logcounts"][0][0] != 999.0) + assert_eq(input.row_data["symbol"][0], "CD3D") + assert_eq(input.reduced_dims["PCA"][0][0], 0.0) + assert_eq(input.alternative_experiments["spike"].assays["counts"][0][0], 1.0) +} + +///| +test "SingleR advanced SingleCellExperiment expands cluster labels to cells" { + let (reference, _) = single_r_adv_test_example() + let input = single_r_adv_test_sce() + input.col_data["cluster"] = ["a", "b", "c", "d", "a"] + let output = @src.single_r_advanced_sce( + input, + reference, + cluster_column="cluster", + output_prefix="singler215", + ) catch { + _ => abort("cluster SCE annotation should work") + } + assert_eq(output.result.n_cells(), 4) + assert_eq(output.experiment.col_data["singler215.labels"].length(), 5) + assert_eq( + output.experiment.col_data["singler215.labels"][0], + output.experiment.col_data["singler215.labels"][4], + ) + assert_eq(output.experiment.metadata["singler215.mode"], "cluster") +} + +///| +test "SingleR advanced SingleCellExperiment supports custom assay" { + let (reference, _) = single_r_adv_test_example() + let output = @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + assay_name="counts", + output_prefix="custom", + ) catch { + _ => abort("custom assay SCE annotation should work") + } + assert_true(output.experiment.col_data.contains("custom.labels")) + assert_eq(output.experiment.metadata["custom.assay"], "counts") +} + +///| +test "SingleR advanced SingleCellExperiment rejects missing inputs" { + let (reference, _) = single_r_adv_test_example() + let assay = try { + ignore( + @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + assay_name="missing", + ), + ) + false + } catch { + _ => true + } + let cluster = try { + ignore( + @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + cluster_column="missing", + ), + ) + false + } catch { + _ => true + } + assert_true(assay) + assert_true(cluster) +} + +///| +test "SingleR advanced SingleCellExperiment validates prefix" { + let (reference, _) = single_r_adv_test_example() + let raised = try { + ignore( + @src.single_r_advanced_sce( + single_r_adv_test_sce(), + reference, + output_prefix=" ", + ), + ) + false + } catch { + _ => true + } + assert_true(raised) +} diff --git a/test/moonbit/single_r_test.mbt b/test/moonbit/single_r_test.mbt index df9f374b..ead1527a 100644 --- a/test/moonbit/single_r_test.mbt +++ b/test/moonbit/single_r_test.mbt @@ -88,7 +88,9 @@ test "single_r_rank_values_ties" { test "single_r_compute_correlations" { let ref_data = @src.single_r_create_reference_data() let cell_expr = ref_data.profiles[0].expression - let scores = @src.single_r_compute_correlations(cell_expr, ref_data, "spearman") + let scores = @src.single_r_compute_correlations( + cell_expr, ref_data, "spearman", + ) assert_eq(scores.length(), ref_data.n_profiles) assert_true(scores[0] > 0.0) } @@ -105,7 +107,11 @@ test "single_r_get_top_scores" { ///| test "single_r_aggregate_scores_by_type" { let ref_data = @src.single_r_create_reference_data() - let scores = @src.single_r_compute_correlations(ref_data.profiles[0].expression, ref_data, "spearman") + let scores = @src.single_r_compute_correlations( + ref_data.profiles[0].expression, + ref_data, + "spearman", + ) let agg = @src.single_r_aggregate_scores_by_type(scores, ref_data, 1) assert_true(agg.size() > 0) assert_true(agg.contains("T cells")) @@ -131,11 +137,13 @@ test "single_r_compute_delta_score" { test "single_r_annotate_cell" { let ref_data = @src.single_r_create_reference_data() let params = @src.SingleRParams::new() - + // Test with a T cell-like expression let cell_expr = ref_data.profiles[0].expression - let result = @src.single_r_annotate_cell(cell_expr, "test_t_cell", ref_data, params) - + let result = @src.single_r_annotate_cell( + cell_expr, "test_t_cell", ref_data, params, + ) + assert_eq(result.cell_id, "test_t_cell") assert_true(result.scores.length() > 0) assert_true(result.first_annotation_fine.length() > 0) @@ -146,11 +154,13 @@ test "single_r_annotate_cell" { test "single_r_annotate_cell_b_cell" { let ref_data = @src.single_r_create_reference_data() let params = @src.SingleRParams::new() - + // Test with a B cell-like expression let cell_expr = ref_data.profiles[1].expression - let result = @src.single_r_annotate_cell(cell_expr, "test_b_cell", ref_data, params) - + let result = @src.single_r_annotate_cell( + cell_expr, "test_b_cell", ref_data, params, + ) + assert_true(result.first_annotation_fine.length() > 0) assert_true(result.scores[0] > 0.0) } @@ -160,9 +170,9 @@ test "single_r_annotate_all_cells" { let ref_data = @src.single_r_create_reference_data() let test_data = @src.single_r_create_test_data() let params = @src.SingleRParams::new() - + let result = @src.single_r_annotate_cells(test_data, ref_data, params) - + assert_eq(result.cell_ids.length(), 50) assert_eq(result.labels.length(), 50) assert_eq(result.scores.length(), 50) @@ -175,10 +185,10 @@ test "single_r_annotation_summary" { let ref_data = @src.single_r_create_reference_data() let test_data = @src.single_r_create_test_data() let params = @src.SingleRParams::new() - + let result = @src.single_r_annotate_cells(test_data, ref_data, params) let summary = @src.single_r_annotation_summary(result) - + assert_true(summary.size() > 0) assert_true(summary.contains("n_cells")) assert_eq(summary.get("n_cells").unwrap_or(0.0), 50.0) @@ -206,9 +216,9 @@ test "single_r_pearson_annotation" { let ref_data = @src.single_r_create_reference_data() let test_data = @src.single_r_create_test_data() let params = @src.SingleRParams::with_method("pearson") - + let result = @src.single_r_annotate_cells(test_data, ref_data, params) - + assert_eq(result.cell_ids.length(), 50) assert_true(result.scores.length() > 0) assert_true(result.scores[0] > 0.0) @@ -216,16 +226,14 @@ test "single_r_pearson_annotation" { ///| test "single_r_reference_dataset_from_matrix" { - let expression = [ - [1.0, 2.0, 3.0], - [4.0, 5.0, 6.0], - [7.0, 8.0, 9.0], - ] + let expression = [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0], [7.0, 8.0, 9.0]] let cell_types = ["TypeA", "TypeB", "TypeC"] let gene_names = ["Gene1", "Gene2", "Gene3"] - - let reference = @src.ReferenceDataset::from_matrix(expression, cell_types, gene_names) - + + let reference = @src.ReferenceDataset::from_matrix( + expression, cell_types, gene_names, + ) + assert_eq(reference.n_profiles, 3) assert_eq(reference.n_genes, 3) assert_eq(reference.cell_types.length(), 3) @@ -236,10 +244,12 @@ test "single_r_reference_dataset_from_matrix" { test "single_r_fine_tune_disabled" { let ref_data = @src.single_r_create_reference_data() let params = @src.SingleRParams::with_all(false, "spearman", 1, 0.0, false) - + let cell_expr = ref_data.profiles[0].expression - let result_no_finetune = @src.single_r_annotate_cell(cell_expr, "test", ref_data, params) - + let result_no_finetune = @src.single_r_annotate_cell( + cell_expr, "test", ref_data, params, + ) + assert_true(result_no_finetune.first_annotation_fine.length() > 0) } @@ -247,10 +257,12 @@ test "single_r_fine_tune_disabled" { test "single_r_fine_tune_enabled" { let ref_data = @src.single_r_create_reference_data() let params = @src.SingleRParams::with_all(true, "spearman", 1, 0.0, false) - + let cell_expr = ref_data.profiles[0].expression - let result_finetune = @src.single_r_annotate_cell(cell_expr, "test", ref_data, params) - + let result_finetune = @src.single_r_annotate_cell( + cell_expr, "test", ref_data, params, + ) + assert_true(result_finetune.first_annotation_fine.length() > 0) assert_true(result_finetune.scores[0] > 0.0) } @@ -270,8 +282,13 @@ test "single_r_identical_profiles" { let ref_data = @src.single_r_create_reference_data() // Annotate a cell with the same profile as reference profile 0 let cell_expr = ref_data.profiles[0].expression - let result = @src.single_r_annotate_cell(cell_expr, "identical", ref_data, @src.SingleRParams::new()) - + let result = @src.single_r_annotate_cell( + cell_expr, + "identical", + ref_data, + @src.SingleRParams::new(), + ) + assert_true(result.scores[0] > 0.99) assert_eq(result.first_annotation_fine, "T cells") -} \ No newline at end of file +} diff --git a/test/moonbit/singscore_test.mbt b/test/moonbit/singscore_test.mbt index 424ca520..d470e8ce 100644 --- a/test/moonbit/singscore_test.mbt +++ b/test/moonbit/singscore_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Singscore gene set scoring module. - test "singscore_create_example" { let (sample, spec) = @src.singscore_create_example() assert_true(sample.sample_id == "sample_1") @@ -9,6 +8,7 @@ test "singscore_create_example" { assert_true(spec.down_genes.length() > 0) } +///| test "singscore_score_basic" { let (sample, spec) = @src.singscore_create_example() let result = @src.singscore_score(sample, spec, "test_sample") @@ -17,6 +17,7 @@ test "singscore_score_basic" { assert_true(result.score >= -0.5 && result.score <= 0.5) } +///| test "singscore_score_boundary" { // Score should always be in [-0.5, 0.5] let (sample, spec) = @src.singscore_create_example() @@ -25,17 +26,19 @@ test "singscore_score_boundary" { assert_true(result.score <= 0.5 + 1.0e-6) } +///| test "singscore_empty_gene_set" { let genes = ["A", "B", "C"] let expression = [1.0, 2.0, 3.0] let sample = @src.SampleExpression::new("empty_test", genes, expression) - let spec = @src.GeneSetSpec::new("empty_set", [], down_genes = []) + let spec = @src.GeneSetSpec::new("empty_set", [], down_genes=[]) let result = @src.singscore_score(sample, spec, "empty") assert_true(result.score == 0.0) assert_true(result.n_up_found == 0) assert_true(result.n_down_found == 0) } +///| test "singscore_single_gene" { let genes = ["TP53"] let expression = [10.0] @@ -46,6 +49,7 @@ test "singscore_single_gene" { assert_true(result.score >= -0.5 && result.score <= 0.5) } +///| test "singscore_multiple_samples" { let (sample1, spec) = @src.singscore_create_example() let genes = sample1.gene_names @@ -59,6 +63,7 @@ test "singscore_multiple_samples" { assert_true(results[0].score >= -0.5 && results[0].score <= 0.5) } +///| test "singscore_missing_genes" { let genes = ["A", "B"] let expression = [1.0, 2.0] @@ -69,8 +74,9 @@ test "singscore_missing_genes" { assert_true(result.score == 0.0) } +///| test "singscore_spec_creation" { - let spec = @src.GeneSetSpec::new("test_set", ["G1", "G2"], down_genes = ["G3"]) + let spec = @src.GeneSetSpec::new("test_set", ["G1", "G2"], down_genes=["G3"]) assert_true(spec.name == "test_set") assert_true(spec.up_genes.length() == 2) assert_true(spec.down_genes.length() == 1) diff --git a/test/moonbit/slingshot_advanced_test.mbt b/test/moonbit/slingshot_advanced_test.mbt new file mode 100644 index 00000000..f9b13316 --- /dev/null +++ b/test/moonbit/slingshot_advanced_test.mbt @@ -0,0 +1,865 @@ +///| +fn sling_adv_test_coordinates() -> Array[Array[Double]] { + [ + [-0.1, 0.0], + [0.0, 0.1], + [0.1, -0.1], + [0.9, 0.0], + [1.0, 0.1], + [1.1, -0.1], + [1.9, 0.0], + [2.0, 0.1], + [2.1, -0.1], + [2.9, 0.9], + [3.0, 1.0], + [3.1, 1.1], + [2.9, -0.9], + [3.0, -1.0], + [3.1, -1.1], + ] +} + +///| +fn sling_adv_test_labels() -> Array[String] { + ["A", "A", "A", "B", "B", "B", "C", "C", "C", "D", "D", "D", "E", "E", "E"] +} + +///| +fn sling_adv_test_config() -> @src.SlingshotAdvancedConfig raise @src.SlingshotAdvancedError { + @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@src.slingshot_center_euclidean(), + extension=@src.slingshot_extend_none(), + max_iterations=8, + tolerance=1.0e-4, + smoother_span=0.4, + curve_points=30, + ) +} + +///| +fn sling_adv_test_result() -> @src.SlingshotAdvancedResult raise @src.SlingshotAdvancedError { + @src.slingshot_advanced( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) +} + +///| +fn sling_adv_close(left : Double, right : Double, tolerance : Double) -> Bool { + (left - right).abs() <= tolerance +} + +///| +test "slingshot advanced distance accessors" { + assert_true( + @src.slingshot_center_euclidean() != @src.slingshot_scaled_diagonal(), + ) + assert_true(@src.slingshot_scaled_diagonal() != @src.slingshot_scaled_full()) + assert_true(@src.slingshot_scaled_full() != @src.slingshot_center_euclidean()) +} + +///| +test "slingshot advanced extension accessors" { + assert_true(@src.slingshot_extend_none() != @src.slingshot_extend_line()) + assert_true(@src.slingshot_extend_line() != @src.slingshot_extend_pc1()) + assert_true(@src.slingshot_extend_pc1() != @src.slingshot_extend_none()) +} + +///| +test "slingshot advanced config defaults" { + let config = @src.SlingshotAdvancedConfig::create() + assert_eq(config.distance, @src.slingshot_scaled_full()) + assert_eq(config.extension, @src.slingshot_extend_line()) + assert_eq(config.max_iterations, 15) + assert_eq(config.curve_points, 100) + assert_true(config.reweight) + assert_true(config.reassign) +} + +///| +test "slingshot advanced config explicit values" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["root"], + end_clusters=["leaf"], + use_median=true, + omega=2.0, + omega_scale=2.5, + shrink=0.25, + reweight=false, + reassign=false, + max_iterations=3, + tolerance=0.01, + smoother_span=0.5, + curve_points=12, + covariance_ridge=0.001, + ) + assert_eq(config.start_clusters, ["root"]) + assert_eq(config.end_clusters, ["leaf"]) + assert_true(config.use_median) + assert_eq(config.omega, 2.0) + assert_eq(config.shrink, 0.25) + assert_false(config.reweight) +} + +///| +test "slingshot advanced rejects invalid omega" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(omega=-2.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid omega scale" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(omega_scale=0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid shrink" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(shrink=1.1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid iterations" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(max_iterations=0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid tolerance" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(tolerance=0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid smoother span" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(smoother_span=1.1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid curve points" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(curve_points=1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects invalid covariance ridge" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(covariance_ridge=0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced rejects empty constraint names" { + let mut failures = 0 + ignore(@src.SlingshotAdvancedConfig::create(start_clusters=[""])) catch { + _ => failures = failures + 1 + } + ignore(@src.SlingshotAdvancedConfig::create(end_clusters=[""])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "slingshot advanced validates coordinate shape" { + let mut failures = 0 + ignore(@src.slingshot_infer_lineages([], [], sling_adv_test_config())) catch { + _ => failures = failures + 1 + } + ignore( + @src.slingshot_infer_lineages( + [[0.0], [1.0, 2.0]], + ["A", "B"], + sling_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "slingshot advanced validates label count" { + let mut failures = 0 + ignore( + @src.slingshot_infer_lineages( + [[0.0], [1.0]], + ["A"], + sling_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced validates finite coordinates" { + let mut failures = 0 + ignore( + @src.slingshot_infer_lineages( + [[0.0], [0.0 / 0.0]], + ["A", "B"], + sling_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced validates start and end clusters" { + let mut failures = 0 + let start_config = @src.SlingshotAdvancedConfig::create(start_clusters=[ + "missing", + ]) + ignore( + @src.slingshot_infer_lineages([[0.0], [1.0]], ["A", "B"], start_config), + ) catch { + _ => failures = failures + 1 + } + let end_config = @src.SlingshotAdvancedConfig::create(end_clusters=["missing"]) + ignore(@src.slingshot_infer_lineages([[0.0], [1.0]], ["A", "B"], end_config)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "slingshot advanced hard labels preserve first occurrence order" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_eq(model.cluster_names, ["A", "B", "C", "D", "E"]) +} + +///| +test "slingshot advanced cluster centers are correct" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_true(sling_adv_close(model.clusters[0].center[0], 0.0, 1.0e-12)) + assert_true(sling_adv_close(model.clusters[2].center[0], 2.0, 1.0e-12)) + assert_true(sling_adv_close(model.clusters[3].center[1], 1.0, 1.0e-12)) + assert_eq(model.clusters[0].size, 3.0) +} + +///| +test "slingshot advanced covariance is finite and regularized" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_true(model.clusters[0].covariance[0][0] > 0.0) + assert_true(model.clusters[0].covariance[1][1] > 0.0) + assert_true( + sling_adv_close( + model.clusters[0].covariance[0][1], + model.clusters[0].covariance[1][0], + 1.0e-12, + ), + ) +} + +///| +test "slingshot advanced euclidean cluster distances" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_true(sling_adv_close(model.distance_matrix[0][1], 1.0, 1.0e-12)) + assert_true(sling_adv_close(model.distance_matrix[2][3], 2.0.sqrt(), 1.0e-12)) + assert_eq(model.distance_matrix[1][0], model.distance_matrix[0][1]) +} + +///| +test "slingshot advanced scaled diagonal distance" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_scaled_diagonal(), + ) + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + assert_true(model.distance_matrix[0][1] > 1.0) +} + +///| +test "slingshot advanced scaled full distance" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_scaled_full(), + ) + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + assert_true(model.distance_matrix[0][1] > 0.0) + assert_eq(model.distance_matrix[0][0], 0.0) +} + +///| +test "slingshot advanced constrained mst has expected size" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_eq(model.edges.length(), 4) + assert_eq(model.roots.length(), 1) +} + +///| +test "slingshot advanced forced endpoints are leaves" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + let degrees = Array::make(model.cluster_names.length(), 0) + for edge in model.edges { + degrees[edge.from] = degrees[edge.from] + 1 + degrees[edge.to] = degrees[edge.to] + 1 + } + assert_eq(degrees[3], 1) + assert_eq(degrees[4], 1) +} + +///| +test "slingshot advanced root honors start cluster" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_eq(model.roots, [0]) +} + +///| +test "slingshot advanced enumerates root to leaf lineages" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_eq(model.lineages.length(), 2) + assert_eq(model.lineages[0][0], 0) + assert_eq(model.lineages[1][0], 0) + let ends = [model.lineages[0][3], model.lineages[1][3]] + ends.sort() + assert_eq(ends, [3, 4]) +} + +///| +test "slingshot advanced initial weights share common trunk" { + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + sling_adv_test_config(), + ) + assert_eq(model.initial_weights[0], [1.0, 1.0]) + assert_eq(model.initial_weights[9], [1.0, 0.0]) + assert_eq(model.initial_weights[12], [0.0, 1.0]) +} + +///| +test "slingshot advanced treats minus one as unclustered" { + let coordinates = [[0.0], [1.0], [2.0]] + let labels = ["A", "-1", "B"] + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + curve_points=10, + max_iterations=2, + ) + let model = @src.slingshot_infer_lineages(coordinates, labels, config) + assert_eq(model.cluster_names, ["A", "B"]) + assert_eq(model.cluster_weights[1], [0.0, 0.0]) + assert_eq(model.initial_weights[1], [0.0]) +} + +///| +test "slingshot advanced weighted memberships normalize rows" { + let model = @src.slingshot_infer_weighted_lineages( + [[0.0], [1.0], [2.0]], + [[2.0, 0.0], [1.0, 1.0], [0.0, 3.0]], + ["A", "B"], + @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + ), + ) + assert_eq(model.cluster_weights[0], [1.0, 0.0]) + assert_eq(model.cluster_weights[1], [0.5, 0.5]) + assert_eq(model.cluster_weights[2], [0.0, 1.0]) +} + +///| +test "slingshot advanced rejects invalid weighted memberships" { + let mut failures = 0 + let config = @src.SlingshotAdvancedConfig::create() + ignore( + @src.slingshot_infer_weighted_lineages( + [[0.0], [1.0]], + [[1.0], [-1.0]], + ["A"], + config, + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.slingshot_infer_weighted_lineages( + [[0.0], [1.0]], + [[1.0, 0.0], [1.0]], + ["A", "B"], + config, + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "slingshot advanced rejects empty weighted clusters" { + let mut failures = 0 + ignore( + @src.slingshot_infer_weighted_lineages( + [[0.0], [1.0]], + [[1.0, 0.0], [1.0, 0.0]], + ["A", "B"], + @src.SlingshotAdvancedConfig::create(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced fixed omega creates forest" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + omega=0.5, + ) + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + assert_eq(model.edges.length(), 0) + assert_eq(model.roots.length(), 5) + assert_eq(model.lineages.length(), 5) +} + +///| +test "slingshot advanced automatic omega uses mst median" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + automatic_omega=true, + omega_scale=1.0, + ) + let model = @src.slingshot_infer_lineages( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + assert_true(model.omega_threshold > 0.0) + assert_true(model.omega_threshold < 2.0) +} + +///| +test "slingshot advanced single cluster lineage" { + let config = @src.SlingshotAdvancedConfig::create( + distance=@src.slingshot_center_euclidean(), + curve_points=12, + max_iterations=2, + ) + let result = @src.slingshot_advanced( + [[0.0, 0.0], [1.0, 0.1], [2.0, -0.1]], + ["A", "A", "A"], + config, + ) + assert_eq(result.curves.length(), 1) + assert_eq(result.curves[0].points.length(), 12) + assert_eq(result.lineage_model.lineages[0], [0]) +} + +///| +test "slingshot advanced fits one curve per lineage" { + let result = sling_adv_test_result() + assert_eq(result.curves.length(), 2) + assert_eq(result.curves[0].points.length(), 30) + assert_eq(result.curves[1].points.length(), 30) +} + +///| +test "slingshot advanced pseudotime matrix dimensions" { + let result = sling_adv_test_result() + assert_eq(result.pseudotime.length(), 15) + assert_eq(result.pseudotime[0].length(), 2) + assert_eq(result.weights.length(), 15) + assert_eq(result.weights[0].length(), 2) +} + +///| +test "slingshot advanced pseudotime begins near root" { + let result = sling_adv_test_result() + let root_time = result.average_pseudotime[0] + assert_true(root_time < result.average_pseudotime[7]) + assert_true(root_time < result.average_pseudotime[10]) + assert_true(root_time < result.average_pseudotime[13]) +} + +///| +test "slingshot advanced optional pseudotime follows weights" { + let result = sling_adv_test_result() + for cell in 0.. assert_true(result.weights[cell][lineage] > 0.0) + None => assert_eq(result.weights[cell][lineage], 0.0) + } + } + } +} + +///| +test "slingshot advanced reports bounded iterations" { + let result = sling_adv_test_result() + assert_true(result.iterations >= 1) + assert_true(result.iterations <= result.config.max_iterations) + assert_true(result.total_distance >= 0.0) +} + +///| +test "slingshot advanced no reweight preserves memberships" { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@src.slingshot_center_euclidean(), + extension=@src.slingshot_extend_none(), + reweight=false, + reassign=false, + max_iterations=2, + curve_points=20, + ) + let result = @src.slingshot_advanced( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + assert_eq(result.weights, result.lineage_model.initial_weights) +} + +///| +test "slingshot advanced shrink aligns shared origin" { + let result = sling_adv_test_result() + let first = result.curves[0].points[0] + let second = result.curves[1].points[0] + assert_true(sling_adv_close(first[0], second[0], 1.0e-8)) + assert_true(sling_adv_close(first[1], second[1], 1.0e-8)) +} + +///| +test "slingshot advanced supports all extension modes" { + let lengths : Array[Double] = [] + for + extension in [ + @src.slingshot_extend_none(), + @src.slingshot_extend_line(), + @src.slingshot_extend_pc1(), + ] { + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@src.slingshot_center_euclidean(), + extension~, + reweight=false, + reassign=false, + max_iterations=1, + curve_points=15, + ) + let result = @src.slingshot_advanced( + sling_adv_test_coordinates(), + sling_adv_test_labels(), + config, + ) + lengths.push(result.curves[0].length) + } + assert_eq(lengths.length(), 3) + assert_true(lengths[0] > 0.0) + assert_true(lengths[1] > 0.0) + assert_true(lengths[2] > 0.0) +} + +///| +test "slingshot advanced probability weights sum to one" { + let probabilities = @src.slingshot_curve_weight_probabilities( + sling_adv_test_result(), + ) + for row in probabilities { + let mut total = 0.0 + for value in row { + total = total + value + } + assert_true(sling_adv_close(total, 1.0, 1.0e-10)) + } +} + +///| +test "slingshot advanced branch ids use one based labels" { + let result = sling_adv_test_result() + let ids = @src.slingshot_branch_ids(result) + assert_true(ids[0].contains("1")) + assert_true(ids[0].contains("2")) + assert_true(ids[10].contains("1")) + assert_true(ids[13].contains("2")) +} + +///| +test "slingshot advanced branch ids validate threshold" { + let mut failures = 0 + ignore(@src.slingshot_branch_ids(sling_adv_test_result(), threshold=1.1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced prediction dimensions" { + let prediction = @src.slingshot_predict(sling_adv_test_result(), [ + [0.0, 0.0], + [3.0, 1.0], + [3.0, -1.0], + ]) + assert_eq(prediction.pseudotime.length(), 3) + assert_eq(prediction.weights[0].length(), 2) + assert_eq(prediction.distances[0].length(), 2) + assert_eq(prediction.average_pseudotime.length(), 3) +} + +///| +test "slingshot advanced prediction maps branch endpoints" { + let prediction = @src.slingshot_predict(sling_adv_test_result(), [ + [3.0, 1.0], + [3.0, -1.0], + ]) + assert_true(prediction.branch_ids[0].contains("1")) + assert_true(prediction.branch_ids[1].contains("2")) +} + +///| +test "slingshot advanced prediction assigns distant cells" { + let prediction = @src.slingshot_predict(sling_adv_test_result(), [ + [100.0, 100.0], + ]) + let mut total = 0.0 + for weight in prediction.weights[0] { + total = total + weight + } + assert_true(total > 0.0) + assert_true(prediction.branch_ids[0].length() > 0) +} + +///| +test "slingshot advanced prediction validates dimensions" { + let mut failures = 0 + ignore(@src.slingshot_predict(sling_adv_test_result(), [[0.0, 1.0, 2.0]])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 1) +} + +///| +test "slingshot advanced weighted full pipeline" { + let coordinates = [[0.0], [0.5], [1.0], [1.5], [2.0]] + let weights = [[1.0, 0.0], [0.75, 0.25], [0.5, 0.5], [0.25, 0.75], [0.0, 1.0]] + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + extension=@src.slingshot_extend_none(), + curve_points=15, + max_iterations=3, + ) + let result = @src.slingshot_advanced_weighted( + coordinates, + weights, + ["A", "B"], + config, + ) + assert_eq(result.curves.length(), 1) + assert_true(result.average_pseudotime[0] < result.average_pseudotime[4]) +} + +///| +test "slingshot advanced summary reports model dimensions" { + let summary = @src.slingshot_advanced_summary(sling_adv_test_result()) + assert_true(summary.contains("Slingshot 2.21.0")) + assert_true(summary.contains("cells: 15")) + assert_true(summary.contains("clusters: 5")) + assert_true(summary.contains("lineages: 2")) +} + +///| +test "slingshot advanced sce writes trajectory outputs" { + let experiment = @src.SingleCellExperiment::new( + [ + [ + 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, + 15.0, + ], + ], + ["gene"], + [ + "c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10", "c11", "c12", + "c13", "c14", "c15", + ], + ) + experiment.reduced_dims["PCA"] = sling_adv_test_coordinates() + experiment.col_data["cluster"] = sling_adv_test_labels() + let output = @src.slingshot_advanced_sce(experiment, sling_adv_test_config()) + assert_eq(output.result.curves.length(), 2) + assert_eq(output.experiment.col_data["slingshot.branch"].length(), 15) + assert_eq( + output.experiment.reduced_dims["slingshot.pseudotime"][0].length(), + 2, + ) + assert_eq(output.experiment.reduced_dims["slingshot.weights"][0].length(), 2) +} + +///| +test "slingshot advanced sce preserves input immutability" { + let experiment = @src.SingleCellExperiment::new( + [ + [ + 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, + 15.0, + ], + ], + ["gene"], + [ + "c1", "c2", "c3", "c4", "c5", "c6", "c7", "c8", "c9", "c10", "c11", "c12", + "c13", "c14", "c15", + ], + ) + experiment.reduced_dims["PCA"] = sling_adv_test_coordinates() + experiment.col_data["cluster"] = sling_adv_test_labels() + experiment.metadata["source"] = "original" + let output = @src.slingshot_advanced_sce(experiment, sling_adv_test_config()) + output.experiment.assays["counts"][0][0] = 99.0 + output.experiment.reduced_dims["PCA"][0][0] = 99.0 + output.experiment.col_data["cluster"][0] = "changed" + assert_eq(experiment.assays["counts"][0][0], 1.0) + assert_eq(experiment.reduced_dims["PCA"][0][0], -0.1) + assert_eq(experiment.col_data["cluster"][0], "A") + assert_eq(experiment.metadata["source"], "original") +} + +///| +test "slingshot advanced sce recursively copies alternatives" { + let experiment = @src.SingleCellExperiment::new([[1.0, 2.0]], ["gene"], [ + "c1", "c2", + ]) + experiment.reduced_dims["PCA"] = [[0.0], [1.0]] + experiment.col_data["cluster"] = ["A", "B"] + experiment.alternative_experiments["alt"] = @src.SingleCellExperiment::new( + [[3.0, 4.0]], + ["feature"], + ["c1", "c2"], + ) + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + curve_points=10, + max_iterations=2, + ) + let output = @src.slingshot_advanced_sce(experiment, config) + output.experiment.alternative_experiments["alt"].assays["counts"][0][0] = 8.0 + assert_eq( + experiment.alternative_experiments["alt"].assays["counts"][0][0], + 3.0, + ) +} + +///| +test "slingshot advanced sce supports custom output prefix" { + let experiment = @src.SingleCellExperiment::new([[1.0, 2.0]], ["gene"], [ + "c1", "c2", + ]) + experiment.reduced_dims["UMAP"] = [[0.0], [1.0]] + experiment.col_data["group"] = ["A", "B"] + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + distance=@src.slingshot_center_euclidean(), + curve_points=10, + max_iterations=2, + ) + let output = @src.slingshot_advanced_sce( + experiment, + config, + reduced_dim_name="UMAP", + cluster_column="group", + output_prefix="trajectory", + ) + assert_true(output.experiment.col_data.contains("trajectory.branch")) + assert_true(output.experiment.reduced_dims.contains("trajectory.pseudotime")) + assert_eq(output.experiment.metadata["trajectory.reduced_dim"], "UMAP") +} + +///| +test "slingshot advanced sce validates names and inputs" { + let experiment = @src.SingleCellExperiment::new([[1.0, 2.0]], ["gene"], [ + "c1", "c2", + ]) + let config = @src.SlingshotAdvancedConfig::create() + let mut failures = 0 + ignore(@src.slingshot_advanced_sce(experiment, config, reduced_dim_name="")) catch { + _ => failures = failures + 1 + } + ignore(@src.slingshot_advanced_sce(experiment, config)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} diff --git a/test/moonbit/slingshot_test.mbt b/test/moonbit/slingshot_test.mbt index 4177b12d..6ae042f0 100644 --- a/test/moonbit/slingshot_test.mbt +++ b/test/moonbit/slingshot_test.mbt @@ -40,7 +40,7 @@ test "slingshot_distance_matrix" { let nodes = [ @src.SlingshotNode::new("c1", "cluster1", [0.0, 0.0], 5), @src.SlingshotNode::new("c2", "cluster2", [1.0, 0.0], 5), - @src.SlingshotNode::new("c3", "cluster3", [0.0, 1.0], 5) + @src.SlingshotNode::new("c3", "cluster3", [0.0, 1.0], 5), ] let matrix = @src.sling_distance_matrix(nodes, @src.sling_euclidean_metric()) assert_eq(matrix.length(), 3) @@ -54,7 +54,7 @@ test "slingshot_mst" { let nodes = [ @src.SlingshotNode::new("c1", "cluster1", [0.0, 0.0], 5), @src.SlingshotNode::new("c2", "cluster2", [1.0, 0.0], 5), - @src.SlingshotNode::new("c3", "cluster3", [0.5, 1.0], 5) + @src.SlingshotNode::new("c3", "cluster3", [0.5, 1.0], 5), ] let edges = @src.sling_build_mst(nodes, @src.sling_euclidean_metric()) assert_eq(edges.length(), 2) @@ -65,7 +65,7 @@ test "slingshot_terminals" { let nodes = [ @src.SlingshotNode::new("c1", "cluster1", [0.0, 0.0], 5), @src.SlingshotNode::new("c2", "cluster2", [1.0, 0.0], 5), - @src.SlingshotNode::new("c3", "cluster3", [0.5, 1.0], 5) + @src.SlingshotNode::new("c3", "cluster3", [0.5, 1.0], 5), ] let edges = @src.sling_build_mst(nodes, @src.sling_euclidean_metric()) let terminals = @src.sling_find_terminals(edges, nodes) @@ -79,7 +79,7 @@ test "slingshot_principal_curve" { [1.0, 0.5], [2.0, 1.0], [3.0, 1.5], - [4.0, 2.0] + [4.0, 2.0], ] let curve = @src.sling_fit_principal_curve(points, 5) assert_true(curve.length() >= 2) @@ -91,13 +91,13 @@ test "slingshot_pseudotime" { [0.0, 0.0], [1.0, 0.5], [2.0, 1.0], - [3.0, 1.5] + [3.0, 1.5], ] let curve : Array[Array[Double]] = [ [0.0, 0.0], [1.0, 0.5], [2.0, 1.0], - [3.0, 1.5] + [3.0, 1.5], ] let pseudotime = @src.sling_compute_pseudotime(cell_coords, curve) assert_eq(pseudotime.length(), 4) @@ -111,7 +111,7 @@ test "slingshot_full_run" { [1.0, 0.5], [2.0, 1.0], [2.5, 1.2], - [3.0, 1.5] + [3.0, 1.5], ] let cluster_labels = ["c1", "c1", "c2", "c2", "c3", "c3"] let params = @src.SlingshotParams::new() diff --git a/test/moonbit/smart_test.mbt b/test/moonbit/smart_test.mbt index aa8307f1..cccd50ed 100644 --- a/test/moonbit/smart_test.mbt +++ b/test/moonbit/smart_test.mbt @@ -19,8 +19,15 @@ test "smart_result_construction" { // SmartDomain Tests // ============================================================================ +///| test "smart_domain_construction" { - let domain = @src.SmartDomain::new("SM00001", "ABC_membrane", 1, 100, evalue=1.5e-20) + let domain = @src.SmartDomain::new( + "SM00001", + "ABC_membrane", + 1, + 100, + evalue=1.5e-20, + ) assert_eq(domain.domain_id, "SM00001") assert_eq(domain.domain_name, "ABC_membrane") assert_eq(domain.start, 1) @@ -32,6 +39,7 @@ test "smart_domain_construction" { // SMART Parsing Tests // ============================================================================ +///| test "smart_parse_empty" { let result = @src.smart_parse("") assert_eq(result.sequence_id, "") @@ -39,6 +47,7 @@ test "smart_parse_empty" { assert_eq(result.domains.length(), 0) } +///| test "smart_parse_sample" { let content = @src.smart_sample() let result = @src.smart_parse(content) @@ -46,6 +55,7 @@ test "smart_parse_sample" { assert_true(result.domains.length() >= 3) } +///| test "smart_parse_domain_type_classification" { let content = @src.smart_sample() let result = @src.smart_parse(content) @@ -58,6 +68,7 @@ test "smart_parse_domain_type_classification" { // SMART Query Tests // ============================================================================ +///| test "smart_find_domains_by_name" { let content = @src.smart_sample() let result = @src.smart_parse(content) @@ -65,6 +76,7 @@ test "smart_find_domains_by_name" { assert_true(matches.length() >= 1) } +///| test "smart_filter_by_evalue" { let content = @src.smart_sample() let result = @src.smart_parse(content) @@ -72,6 +84,7 @@ test "smart_filter_by_evalue" { assert_true(significant.length() >= 1) } +///| test "smart_best_domain" { let content = @src.smart_sample() let result = @src.smart_parse(content) @@ -85,12 +98,14 @@ test "smart_best_domain" { // SMART Summary Tests // ============================================================================ +///| test "smart_summary" { let result = @src.SmartResult::new("test_protein") let summary = @src.smart_summary(result) assert_true(summary.contains("test_protein")) } +///| test "smart_total_domains" { let content = @src.smart_sample() let result = @src.smart_parse(content) diff --git a/test/moonbit/snapgene_io_test.mbt b/test/moonbit/snapgene_io_test.mbt index 298b1477..dea3a331 100644 --- a/test/moonbit/snapgene_io_test.mbt +++ b/test/moonbit/snapgene_io_test.mbt @@ -23,13 +23,21 @@ test "snapgene_seq_type_to_string" { ///| test "snapgene_seq_type_from_string" { - assert_true(@src.SnapgeneSeqType::from_string("DNA") is @src.SnapgeneSeqType::SnapgeneDna) - assert_true(@src.SnapgeneSeqType::from_string("RNA") is @src.SnapgeneSeqType::SnapgeneRna) assert_true( - @src.SnapgeneSeqType::from_string("protein") is @src.SnapgeneSeqType::SnapgeneProtein, + @src.SnapgeneSeqType::from_string("DNA") + is @src.SnapgeneSeqType::SnapgeneDna, ) assert_true( - @src.SnapgeneSeqType::from_string("xxx") is @src.SnapgeneSeqType::SnapgeneUnknown, + @src.SnapgeneSeqType::from_string("RNA") + is @src.SnapgeneSeqType::SnapgeneRna, + ) + assert_true( + @src.SnapgeneSeqType::from_string("protein") + is @src.SnapgeneSeqType::SnapgeneProtein, + ) + assert_true( + @src.SnapgeneSeqType::from_string("xxx") + is @src.SnapgeneSeqType::SnapgeneUnknown, ) } @@ -38,8 +46,14 @@ test "snapgene_seq_type_from_string_round_trip" { let dna = @src.SnapgeneSeqType::from_string("DNA") let rna = @src.SnapgeneSeqType::from_string("RNA") let protein = @src.SnapgeneSeqType::from_string("protein") - assert_eq(@src.SnapgeneSeqType::from_string(dna.to_string()).to_string(), "DNA") - assert_eq(@src.SnapgeneSeqType::from_string(rna.to_string()).to_string(), "RNA") + assert_eq( + @src.SnapgeneSeqType::from_string(dna.to_string()).to_string(), + "DNA", + ) + assert_eq( + @src.SnapgeneSeqType::from_string(rna.to_string()).to_string(), + "RNA", + ) assert_eq( @src.SnapgeneSeqType::from_string(protein.to_string()).to_string(), "protein", @@ -76,7 +90,7 @@ test "snapgene_feature_new" { type_="CDS", direction=1, segments=[(0, 99)], - qualifiers=qualifiers, + qualifiers~, ) assert_eq(f.name, "orfA") assert_eq(f.type_, "CDS") diff --git a/test/moonbit/sparse_array_test.mbt b/test/moonbit/sparse_array_test.mbt new file mode 100644 index 00000000..7126a137 --- /dev/null +++ b/test/moonbit/sparse_array_test.mbt @@ -0,0 +1,592 @@ +///| +fn sparse_array_matrix() -> @src.SparseArray { + @src.SparseArray::from_matrix([[1.0, 0.0, 2.0], [0.0, 3.0, 0.0]]) catch { + _ => abort("failed to create sparse matrix fixture") + } +} + +///| +test "sparse_array: canonical COO merges duplicates and removes zeros" { + let array = @src.SparseArray::from_coo( + [2, 3], + [[1, 2], [0, 1], [1, 2], [0, 0]], + [2.0, 3.0, -2.0, 0.0], + ) catch { + _ => abort("valid COO input should be accepted") + } + assert_eq(array.nnzero(), 1) + assert_eq(array.nzcoordinates(), [[0, 1]]) + assert_eq(array.nzvalues(), [3.0]) +} + +///| +test "sparse_array: dimensions length density and sparse marker" { + let array = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + assert_eq(array.ndim(), 3) + assert_eq(array.dim(), [4, 5, 2]) + assert_eq(array.length(), 40) + assert_eq(array.nnzero(), 5) + assert_eq(array.density(), 0.125) + assert_true(array.is_sparse()) +} + +///| +test "sparse_array: zero extent array is valid and empty" { + let array = @src.SparseArray::zeros([2, 0, 4]) catch { + _ => abort("zero extent should be valid") + } + assert_eq(array.length(), 0) + assert_eq(array.nnzero(), 0) + assert_eq(array.density(), 0.0) + assert_true(array.mean() is None) + assert_true(array.minimum() is None) + assert_true(array.maximum() is None) +} + +///| +test "sparse_array: rejects invalid dimensions" { + let empty_rank = try { + ignore(@src.SparseArray::zeros([])) + false + } catch { + SparseArrayError(_) => true + } + let negative_extent = try { + ignore(@src.SparseArray::zeros([2, -1])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(empty_rank) + assert_true(negative_extent) +} + +///| +test "sparse_array: rejects malformed COO input" { + let length_mismatch = try { + ignore(@src.SparseArray::from_coo([2, 2], [[0, 0]], [])) + false + } catch { + SparseArrayError(_) => true + } + let rank_mismatch = try { + ignore(@src.SparseArray::from_coo([2, 2], [[0]], [1.0])) + false + } catch { + SparseArrayError(_) => true + } + let out_of_bounds = try { + ignore(@src.SparseArray::from_coo([2, 2], [[0, 2]], [1.0])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(length_mismatch) + assert_true(rank_mismatch) + assert_true(out_of_bounds) +} + +///| +test "sparse_array: flat data uses R column-major ordering" { + let array = @src.SparseArray::from_flat([2, 3], [1.0, 2.0, 0.0, 3.0, 4.0, 0.0]) catch { + _ => abort("flat input should be valid") + } + assert_eq(array.to_dense_matrix(), [[1.0, 0.0, 4.0], [2.0, 3.0, 0.0]]) + assert_eq(array.to_flat(), [1.0, 2.0, 0.0, 3.0, 4.0, 0.0]) +} + +///| +test "sparse_array: rejects flat length mismatch" { + let raised = try { + ignore(@src.SparseArray::from_flat([2, 3], [1.0, 2.0])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(raised) +} + +///| +test "sparse_array: matrix construction and dense conversion" { + let array = sparse_array_matrix() + assert_eq(array.dim(), [2, 3]) + assert_eq(array.nnzero(), 3) + assert_eq(array.to_dense_matrix(), [[1.0, 0.0, 2.0], [0.0, 3.0, 0.0]]) +} + +///| +test "sparse_array: rejects ragged matrix and multidimensional dense conversion" { + let ragged = try { + ignore(@src.SparseArray::from_matrix([[1.0], [2.0, 3.0]])) + false + } catch { + SparseArrayError(_) => true + } + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let not_matrix = try { + ignore(sample.to_dense_matrix()) + false + } catch { + SparseArrayError(_) => true + } + assert_true(ragged) + assert_true(not_matrix) +} + +///| +test "sparse_array: coordinate and linear access" { + let array = sparse_array_matrix() + assert_eq(array.get([0, 0]), 1.0) + assert_eq(array.get([0, 1]), 0.0) + assert_eq(array.get([1, 1]), 3.0) + assert_eq(array.get_linear(0), 1.0) + assert_eq(array.get_linear(2), 0.0) + assert_eq(array.get_linear(3), 3.0) + assert_eq(array.get_linear(4), 2.0) +} + +///| +test "sparse_array: rejects invalid access" { + let array = sparse_array_matrix() + let bad_coordinate = try { + ignore(array.get([2, 0])) + false + } catch { + SparseArrayError(_) => true + } + let bad_rank = try { + ignore(array.get([0])) + false + } catch { + SparseArrayError(_) => true + } + let bad_linear = try { + ignore(array.get_linear(6)) + false + } catch { + SparseArrayError(_) => true + } + assert_true(bad_coordinate) + assert_true(bad_rank) + assert_true(bad_linear) +} + +///| +test "sparse_array: accessors return defensive coordinate copies" { + let array = sparse_array_matrix() + let dimensions = array.dim() + dimensions[0] = 99 + let coordinates = array.nzcoordinates() + coordinates[0][0] = 1 + let entries = array.nonzero_entries() + let first_coordinates = entries[0].coordinates() + first_coordinates[0] = 1 + assert_eq(array.dim(), [2, 3]) + assert_eq(array.get([0, 0]), 1.0) +} + +///| +test "sparse_array: immutable single-value assignment inserts replaces and removes" { + let original = sparse_array_matrix() + let inserted = original.with_value([1, 2], 4.0) + let replaced = inserted.with_value([0, 2], 5.0) + let removed = replaced.with_value([1, 1], 0.0) + assert_eq(original.get([1, 2]), 0.0) + assert_eq(inserted.get([1, 2]), 4.0) + assert_eq(replaced.get([0, 2]), 5.0) + assert_eq(removed.get([1, 1]), 0.0) + assert_eq(removed.nnzero(), 3) +} + +///| +test "sparse_array: ordered multiple assignment uses last replacement" { + let array = sparse_array_matrix().with_values([[0, 0], [1, 2], [0, 0]], [ + 8.0, 4.0, 9.0, + ]) + assert_eq(array.get([0, 0]), 9.0) + assert_eq(array.get([1, 2]), 4.0) +} + +///| +test "sparse_array: multiple assignment validates parallel lengths" { + let raised = try { + ignore(sparse_array_matrix().with_values([[0, 0]], [])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(raised) +} + +///| +test "sparse_array: subset preserves order and repeated indices" { + let subset = sparse_array_matrix().subset([[1, 0, 1], [2, 1]]) + assert_eq(subset.dim(), [3, 2]) + assert_eq(subset.to_dense_matrix(), [[0.0, 3.0], [2.0, 0.0], [0.0, 3.0]]) + assert_eq(subset.nnzero(), 3) +} + +///| +test "sparse_array: empty subset remains sparse" { + let subset = sparse_array_matrix().subset([[], [0, 1]]) + assert_eq(subset.dim(), [0, 2]) + assert_eq(subset.length(), 0) + assert_eq(subset.nnzero(), 0) +} + +///| +test "sparse_array: subset validates rank and bounds" { + let array = sparse_array_matrix() + let bad_rank = try { + ignore(array.subset([[0, 1]])) + false + } catch { + SparseArrayError(_) => true + } + let bad_index = try { + ignore(array.subset([[0, 2], [0]])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(bad_rank) + assert_true(bad_index) +} + +///| +test "sparse_array: end-exclusive multidimensional slice" { + let sliced = sparse_array_matrix().slice([0, 1], [2, 3]) + assert_eq(sliced.dim(), [2, 2]) + assert_eq(sliced.to_dense_matrix(), [[0.0, 2.0], [3.0, 0.0]]) +} + +///| +test "sparse_array: slice validates bounds" { + let array = sparse_array_matrix() + let bad_rank = try { + ignore(array.slice([0], [1])) + false + } catch { + SparseArrayError(_) => true + } + let reversed = try { + ignore(array.slice([1, 0], [0, 2])) + false + } catch { + SparseArrayError(_) => true + } + let too_large = try { + ignore(array.slice([0, 0], [3, 2])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(bad_rank) + assert_true(reversed) + assert_true(too_large) +} + +///| +test "sparse_array: aperm reorders multidimensional coordinates" { + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let permuted = sample.aperm([2, 0, 1]) + assert_eq(permuted.dim(), [2, 4, 5]) + assert_eq(permuted.get([1, 0, 4]), 40.0) + assert_eq(permuted.get([1, 2, 1]), 50.0) +} + +///| +test "sparse_array: aperm validates complete permutation" { + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let wrong_length = try { + ignore(sample.aperm([1, 0])) + false + } catch { + SparseArrayError(_) => true + } + let duplicate = try { + ignore(sample.aperm([0, 0, 2])) + false + } catch { + SparseArrayError(_) => true + } + assert_true(wrong_length) + assert_true(duplicate) +} + +///| +test "sparse_array: matrix transpose remains sparse" { + let transposed = sparse_array_matrix().transpose() + assert_eq(transposed.dim(), [3, 2]) + assert_eq(transposed.to_dense_matrix(), [[1.0, 0.0], [0.0, 3.0], [2.0, 0.0]]) + assert_eq(transposed.nnzero(), 3) +} + +///| +test "sparse_array: transpose rejects non-matrix" { + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let raised = try { + ignore(sample.transpose()) + false + } catch { + SparseArrayError(_) => true + } + assert_true(raised) +} + +///| +test "sparse_array: bind offsets coordinates along selected dimension" { + let left = @src.SparseArray::from_matrix([[1.0, 0.0], [0.0, 2.0]]) + let right = @src.SparseArray::from_matrix([[3.0, 4.0]]) + let bound = left.bind(right, 0) + assert_eq(bound.dim(), [3, 2]) + assert_eq(bound.to_dense_matrix(), [[1.0, 0.0], [0.0, 2.0], [3.0, 4.0]]) +} + +///| +test "sparse_array: bind validates rank margin and dimensions" { + let matrix = sparse_array_matrix() + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let rank_mismatch = try { + ignore(matrix.bind(sample, 0)) + false + } catch { + SparseArrayError(_) => true + } + let bad_margin = try { + ignore(matrix.bind(matrix, 2)) + false + } catch { + SparseArrayError(_) => true + } + let other = @src.SparseArray::from_matrix([[1.0, 2.0]]) + let dimension_mismatch = try { + ignore(matrix.bind(other, 1)) + false + } catch { + SparseArrayError(_) => true + } + assert_true(rank_mismatch) + assert_true(bad_margin) + assert_true(dimension_mismatch) +} + +///| +test "sparse_array: addition merges coordinates and drops cancellation" { + let left = @src.SparseArray::from_matrix([[1.0, 0.0], [2.0, -3.0]]) + let right = @src.SparseArray::from_matrix([[-1.0, 4.0], [0.0, 3.0]]) + let result = left.add(right) + assert_eq(result.to_dense_matrix(), [[0.0, 4.0], [2.0, 0.0]]) + assert_eq(result.nnzero(), 2) +} + +///| +test "sparse_array: subtraction preserves sparse union" { + let left = @src.SparseArray::from_matrix([[1.0, 0.0], [2.0, -3.0]]) + let right = @src.SparseArray::from_matrix([[-1.0, 4.0], [0.0, 3.0]]) + assert_eq(left.subtract(right).to_dense_matrix(), [[2.0, -4.0], [2.0, -6.0]]) +} + +///| +test "sparse_array: Hadamard product only visits coordinate intersection" { + let left = @src.SparseArray::from_matrix([[1.0, 0.0], [2.0, -3.0]]) + let right = @src.SparseArray::from_matrix([[-1.0, 4.0], [0.0, 3.0]]) + let result = left.hadamard(right) + assert_eq(result.to_dense_matrix(), [[-1.0, 0.0], [0.0, -9.0]]) + assert_eq(result.nnzero(), 2) +} + +///| +test "sparse_array: arithmetic validates dimensions" { + let left = sparse_array_matrix() + let right = @src.SparseArray::from_matrix([[1.0, 2.0]]) + let add_raised = try { + ignore(left.add(right)) + false + } catch { + SparseArrayError(_) => true + } + let product_raised = try { + ignore(left.hadamard(right)) + false + } catch { + SparseArrayError(_) => true + } + assert_true(add_raised) + assert_true(product_raised) +} + +///| +test "sparse_array: scale and zero-preserving map remove generated zeros" { + let array = sparse_array_matrix() + assert_eq(array.scale(0.0).nnzero(), 0) + let mapped = array.map_nonzero(fn(value : Double) -> Double { + if value == 2.0 { + 0.0 + } else { + value * value + } + }) + assert_eq(mapped.to_dense_matrix(), [[1.0, 0.0, 0.0], [0.0, 9.0, 0.0]]) + assert_eq(mapped.nnzero(), 2) +} + +///| +test "sparse_array: whole-array summaries include implicit zeros" { + let array = @src.SparseArray::from_matrix([[-2.0, 0.0, 5.0], [0.0, 0.0, 1.0]]) + assert_eq(array.sum(), 4.0) + match array.mean() { + Some(value) => assert_true((value - 4.0 / 6.0).abs() < 1.0e-12) + None => assert_true(false) + } + match array.minimum() { + Some(value) => assert_eq(value, -2.0) + None => assert_true(false) + } + match array.maximum() { + Some(value) => assert_eq(value, 5.0) + None => assert_true(false) + } +} + +///| +test "sparse_array: extrema do not inject zero into fully dense arrays" { + let array = @src.SparseArray::from_matrix([[-2.0, -1.0]]) + match array.minimum() { + Some(value) => assert_eq(value, -2.0) + None => assert_true(false) + } + match array.maximum() { + Some(value) => assert_eq(value, -1.0) + None => assert_true(false) + } +} + +///| +test "sparse_array: row and column sums and means" { + let array = @src.SparseArray::from_matrix([[-2.0, 0.0, 5.0], [0.0, 0.0, 1.0]]) + assert_eq(array.row_sums(), [3.0, 1.0]) + assert_eq(array.column_sums(), [-2.0, 0.0, 6.0]) + let row_means = array.row_means() + assert_eq(row_means[0], 1.0) + assert_true((row_means[1] - 1.0 / 3.0).abs() < 1.0e-12) + assert_eq(array.column_means(), [-1.0, 0.0, 3.0]) +} + +///| +test "sparse_array: row and column nonzero counts" { + let array = sparse_array_matrix() + assert_eq(array.row_nonzero_counts(), [2, 1]) + assert_eq(array.column_nonzero_counts(), [1, 1, 1]) +} + +///| +test "sparse_array: matrix summaries reject rank and empty means" { + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let rank_raised = try { + ignore(sample.row_sums()) + false + } catch { + SparseArrayError(_) => true + } + let no_columns = @src.SparseArray::zeros([2, 0]) catch { + _ => abort("zero-column matrix should be valid") + } + let row_mean_raised = try { + ignore(no_columns.row_means()) + false + } catch { + SparseArrayError(_) => true + } + let no_rows = @src.SparseArray::zeros([0, 2]) catch { + _ => abort("zero-row matrix should be valid") + } + let column_mean_raised = try { + ignore(no_rows.column_means()) + false + } catch { + SparseArrayError(_) => true + } + assert_true(rank_raised) + assert_true(row_mean_raised) + assert_true(column_mean_raised) +} + +///| +test "sparse_array: sparse matrix multiplication returns dense result" { + let left = sparse_array_matrix() + let right = @src.SparseArray::from_matrix([[0.0, 4.0], [5.0, 0.0], [6.0, 7.0]]) + assert_eq(left.matmul(right), [[12.0, 18.0], [15.0, 0.0]]) +} + +///| +test "sparse_array: crossprod and tcrossprod" { + let array = sparse_array_matrix() + assert_eq(array.crossprod(), [ + [1.0, 0.0, 2.0], + [0.0, 9.0, 0.0], + [2.0, 0.0, 4.0], + ]) + assert_eq(array.tcrossprod(), [[5.0, 0.0], [0.0, 9.0]]) +} + +///| +test "sparse_array: matrix multiplication validates rank and dimensions" { + let matrix = sparse_array_matrix() + let incompatible = @src.SparseArray::from_matrix([[1.0, 2.0]]) + let dimension_raised = try { + ignore(matrix.matmul(incompatible)) + false + } catch { + SparseArrayError(_) => true + } + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let rank_raised = try { + ignore(sample.matmul(matrix)) + false + } catch { + SparseArrayError(_) => true + } + assert_true(dimension_raised) + assert_true(rank_raised) +} + +///| +test "sparse_array: three-dimensional flat round trip" { + let original = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let restored = @src.SparseArray::from_flat(original.dim(), original.to_flat()) catch { + _ => abort("round trip should be valid") + } + assert_eq(restored.dim(), original.dim()) + assert_eq(restored.nzcoordinates(), original.nzcoordinates()) + assert_eq(restored.nzvalues(), original.nzvalues()) +} + +///| +test "sparse_array: sample and summary" { + let sample = @src.sparse_array_sample() catch { + _ => abort("sample should be valid") + } + let summary = sample.summary() + assert_true(summary.contains("SparseArray(4 x 5 x 2")) + assert_true(summary.contains("nnzero=5")) + assert_true(summary.contains("density=0.125")) +} diff --git a/test/moonbit/spatial_experiment_test.mbt b/test/moonbit/spatial_experiment_test.mbt index 8c706cb5..b34e6b06 100644 --- a/test/moonbit/spatial_experiment_test.mbt +++ b/test/moonbit/spatial_experiment_test.mbt @@ -1,12 +1,12 @@ ///| /// Tests for SpatialExperiment module. - test "SpatialExperiment creation" { let se = @src.create_example_spatial_experiment() assert_eq(@src.se_num_rows(se), 3) assert_eq(@src.se_num_cols(se), 6) } +///| test "SpatialExperiment spatial range" { let se = @src.create_example_spatial_experiment() let (min_x, max_x, min_y, max_y) = @src.se_get_spatial_range(se) @@ -16,15 +16,17 @@ test "SpatialExperiment spatial range" { assert_eq(max_y, 250.0) } +///| test "SpatialExperiment filter spots" { let se = @src.create_example_spatial_experiment() let filtered = @src.se_filter_spots_by_range(se, 150.0, 300.0, 150.0, 250.0) assert_eq(@src.se_num_cols(filtered), 4) } +///| test "SpatialCoord new_2d" { let coord = @src.SpatialCoord::new_2d(100.0, 200.0) assert_eq(coord.x, 100.0) assert_eq(coord.y, 200.0) assert_eq(coord.z, 0.0) -} \ No newline at end of file +} diff --git a/test/moonbit/spatialdecon_test.mbt b/test/moonbit/spatialdecon_test.mbt new file mode 100644 index 00000000..4d5d3df2 --- /dev/null +++ b/test/moonbit/spatialdecon_test.mbt @@ -0,0 +1,1492 @@ +// Black-box tests for the Bioconductor SpatialDecon-inspired workflow. + +///| +fn sd_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "SpatialDecon value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn sd_test_finite(value : Double) -> Bool { + value == value && value.abs() <= 1.0e300 +} + +///| +fn sd_test_config( + rescale_profile? : Bool = false, + refit_outliers? : Bool = false, + residual_threshold? : Double = 3.0, +) -> @src.SpatialDeconConfig { + @src.SpatialDeconConfig::create( + residual_threshold~, + rescale_profile~, + refit_outliers~, + tolerance=1.0e-9, + ) catch { + _ => abort("SpatialDecon test configuration should be valid") + } +} + +///| +fn sd_test_exact_profile() -> @src.SpatialDeconProfile { + @src.SpatialDeconProfile::create(["g1", "g2", "g3"], ["Type_A"], [ + [1.0], + [2.0], + [3.0], + ]) catch { + _ => abort("SpatialDecon exact profile should be valid") + } +} + +///| +fn sd_test_exact_data() -> @src.SpatialDeconData { + @src.SpatialDeconData::create( + ["g1", "g2", "g3"], + ["spot_1", "spot_2"], + [[7.0, 9.0], [9.0, 13.0], [11.0, 17.0]], + background=[[5.0, 5.0], [5.0, 5.0], [5.0, 5.0]], + ) catch { + _ => abort("SpatialDecon exact data should be valid") + } +} + +///| +fn sd_test_exact_result( + nuclei_counts? : Array[Double] = [], +) -> @src.SpatialDeconResult { + @src.spatial_decon( + sd_test_exact_data(), + sd_test_exact_profile(), + nuclei_counts~, + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon exact fit should succeed") + } +} + +///| +fn sd_test_example_result() -> @src.SpatialDeconResult { + let (data, profile, nuclei_counts) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon example data should be valid") + } + @src.spatial_decon( + data, + profile, + nuclei_counts~, + config=sd_test_config(refit_outliers=true), + ) catch { + _ => abort("SpatialDecon example fit should succeed") + } +} + +///| +fn sd_test_experiment() -> ( + @src.SpatialExperiment, + @src.SpatialDeconProfile, + Array[Double], +) { + let (data, profile, nuclei_counts) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon example data should be valid") + } + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "counts", data.values)) + ignore(@src.se_add_assay(experiment, "background", data.background)) + ignore(@src.se_add_assay(experiment, "weights", data.weights)) + for gene in data.gene_names { + ignore( + @src.se_add_row( + experiment, + Map([("gene_id", gene), ("feature_type", "target")]), + ), + ) + } + for spot in 0.. abort("SpatialDecon custom configuration should be valid") + } + assert_eq(config.residual_threshold, 2.0) + assert_eq(config.lower_threshold, 0.25) + assert_eq(config.rescale_profile, false) + assert_eq(config.profile_target, 3.0) + assert_eq(config.refit_outliers, false) + assert_eq(config.max_iterations, 200) +} + +///| +test "SpatialDecon: configuration rejects invalid thresholds" { + let residual_failed = try { + ignore(@src.SpatialDeconConfig::create(residual_threshold=0.0)) + false + } catch { + SpatialDeconError(_) => true + } + let lower_failed = try { + ignore(@src.SpatialDeconConfig::create(lower_threshold=0.0)) + false + } catch { + SpatialDeconError(_) => true + } + let floor_failed = try { + ignore(@src.SpatialDeconConfig::create(signal_floor=0.0)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(residual_failed && lower_failed && floor_failed) +} + +///| +test "SpatialDecon: configuration rejects invalid profile scaling" { + let quantile_failed = try { + ignore(@src.SpatialDeconConfig::create(profile_quantile=1.1)) + false + } catch { + SpatialDeconError(_) => true + } + let target_failed = try { + ignore(@src.SpatialDeconConfig::create(profile_target=0.0)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(quantile_failed && target_failed) +} + +///| +test "SpatialDecon: configuration rejects invalid solver controls" { + let iterations_failed = try { + ignore(@src.SpatialDeconConfig::create(max_iterations=0)) + false + } catch { + SpatialDeconError(_) => true + } + let search_failed = try { + ignore(@src.SpatialDeconConfig::create(line_search_steps=0)) + false + } catch { + SpatialDeconError(_) => true + } + let tolerance_failed = try { + ignore(@src.SpatialDeconConfig::create(tolerance=1.0)) + false + } catch { + SpatialDeconError(_) => true + } + let ridge_failed = try { + ignore(@src.SpatialDeconConfig::create(ridge=0.0)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true( + iterations_failed && search_failed && tolerance_failed && ridge_failed, + ) +} + +///| +test "SpatialDecon: profile constructor exposes dimensions" { + let profile = sd_test_exact_profile() + assert_eq(profile.n_genes(), 3) + assert_eq(profile.n_cell_types(), 1) + assert_eq(profile.gene_names, ["g1", "g2", "g3"]) + assert_eq(profile.cell_types, ["Type_A"]) +} + +///| +test "SpatialDecon: profile constructor defensively copies inputs" { + let genes = ["g1", "g2"] + let cell_types = ["A"] + let values = [[1.0], [2.0]] + let profile = @src.SpatialDeconProfile::create(genes, cell_types, values) catch { + _ => abort("SpatialDecon profile should be valid") + } + genes[0] = "changed" + cell_types[0] = "changed" + values[0][0] = 99.0 + assert_eq(profile.gene_names, ["g1", "g2"]) + assert_eq(profile.cell_types, ["A"]) + assert_eq(profile.values, [[1.0], [2.0]]) +} + +///| +test "SpatialDecon: profile rejects duplicate and empty names" { + let duplicate_failed = try { + ignore( + @src.SpatialDeconProfile::create(["g1", "g1"], ["A"], [[1.0], [2.0]]), + ) + false + } catch { + SpatialDeconError(_) => true + } + let empty_failed = try { + ignore(@src.SpatialDeconProfile::create(["g1"], [""], [[1.0]])) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(duplicate_failed && empty_failed) +} + +///| +test "SpatialDecon: profile rejects malformed and negative matrices" { + let shape_failed = try { + ignore( + @src.SpatialDeconProfile::create(["g1", "g2"], ["A", "B"], [ + [1.0, 2.0], + [3.0], + ]), + ) + false + } catch { + SpatialDeconError(_) => true + } + let negative_failed = try { + ignore(@src.SpatialDeconProfile::create(["g1"], ["A"], [[-1.0]])) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(shape_failed && negative_failed) +} + +///| +test "SpatialDecon: profile rejects non-finite and empty cell-type signals" { + let finite_failed = try { + ignore( + @src.SpatialDeconProfile::create(["g1"], ["A"], [[@double.not_a_number]]), + ) + false + } catch { + SpatialDeconError(_) => true + } + let signal_failed = try { + ignore( + @src.SpatialDeconProfile::create(["g1", "g2"], ["A", "B"], [ + [1.0, 0.0], + [2.0, 0.0], + ]), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(finite_failed && signal_failed) +} + +///| +test "SpatialDecon: data constructor supplies background and weights" { + let data = @src.SpatialDeconData::create(["g1", "g2"], ["s1"], [[2.0], [3.0]]) catch { + _ => abort("SpatialDecon data should be valid") + } + assert_eq(data.n_genes(), 2) + assert_eq(data.n_spots(), 1) + assert_eq(data.background, [[0.0], [0.0]]) + assert_eq(data.weights, [[1.0], [1.0]]) +} + +///| +test "SpatialDecon: data constructor defensively copies all matrices" { + let genes = ["g1", "g2"] + let spots = ["s1"] + let values = [[2.0], [3.0]] + let background = [[0.5], [0.6]] + let weights = [[1.0], [2.0]] + let data = @src.SpatialDeconData::create( + genes, + spots, + values, + background~, + weights~, + ) catch { + _ => abort("SpatialDecon data should be valid") + } + genes[0] = "changed" + spots[0] = "changed" + values[0][0] = 99.0 + background[0][0] = 99.0 + weights[0][0] = 99.0 + assert_eq(data.gene_names, ["g1", "g2"]) + assert_eq(data.spot_ids, ["s1"]) + assert_eq(data.values[0][0], 2.0) + assert_eq(data.background[0][0], 0.5) + assert_eq(data.weights[0][0], 1.0) +} + +///| +test "SpatialDecon: scalar background fills every observation" { + let data = @src.SpatialDeconData::with_scalar_background( + ["g1", "g2"], + ["s1", "s2"], + [[2.0, 3.0], [4.0, 5.0]], + 0.75, + ) catch { + _ => abort("SpatialDecon scalar background should be valid") + } + assert_eq(data.background, [[0.75, 0.75], [0.75, 0.75]]) +} + +///| +test "SpatialDecon: data rejects duplicate spots and invalid shapes" { + let duplicate_failed = try { + ignore(@src.SpatialDeconData::create(["g1"], ["s1", "s1"], [[1.0, 2.0]])) + false + } catch { + SpatialDeconError(_) => true + } + let shape_failed = try { + ignore( + @src.SpatialDeconData::create(["g1", "g2"], ["s1"], [[1.0], [2.0, 3.0]]), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(duplicate_failed && shape_failed) +} + +///| +test "SpatialDecon: data rejects invalid expression and background" { + let expression_failed = try { + ignore(@src.SpatialDeconData::create(["g1"], ["s1"], [[-1.0]])) + false + } catch { + SpatialDeconError(_) => true + } + let background_failed = try { + ignore( + @src.SpatialDeconData::create(["g1"], ["s1"], [[1.0]], background=[ + [@double.not_a_number], + ]), + ) + false + } catch { + SpatialDeconError(_) => true + } + let scalar_failed = try { + ignore( + @src.SpatialDeconData::with_scalar_background( + ["g1"], + ["s1"], + [[1.0]], + -1.0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(expression_failed && background_failed && scalar_failed) +} + +///| +test "SpatialDecon: data requires strictly positive finite weights" { + let zero_failed = try { + ignore( + @src.SpatialDeconData::create(["g1"], ["s1"], [[1.0]], weights=[[0.0]]), + ) + false + } catch { + SpatialDeconError(_) => true + } + let finite_failed = try { + ignore( + @src.SpatialDeconData::create(["g1"], ["s1"], [[1.0]], weights=[ + [@double.not_a_number], + ]), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(zero_failed && finite_failed) +} + +///| +test "SpatialDecon: exact single-cell-type mixture is recovered" { + let result = sd_test_exact_result() + sd_test_close(result.beta[0][0], 2.0, 1.0e-6) + sd_test_close(result.beta[0][1], 4.0, 1.0e-6) + assert_eq(result.converged, [true, true]) + assert_true(result.objectives[0] < 1.0e-12) + assert_true(result.objectives[1] < 1.0e-12) +} + +///| +test "SpatialDecon: background-aware fit differs from uncorrected fit" { + let corrected = sd_test_exact_result() + let uncorrected_data = @src.SpatialDeconData::create( + ["g1", "g2", "g3"], + ["spot_1", "spot_2"], + [[7.0, 9.0], [9.0, 13.0], [11.0, 17.0]], + ) catch { + _ => abort("SpatialDecon uncorrected data should be valid") + } + let uncorrected = @src.spatial_decon( + uncorrected_data, + sd_test_exact_profile(), + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon uncorrected fit should succeed") + } + sd_test_close(corrected.beta[0][0], 2.0, 1.0e-6) + assert_true(uncorrected.beta[0][0] > 3.0) +} + +///| +test "SpatialDecon: profile rescaling inversely rescales abundance" { + let profile = @src.SpatialDeconProfile::create(["g1", "g2"], ["A"], [ + [2.0], + [4.0], + ]) catch { + _ => abort("SpatialDecon scaling profile should be valid") + } + let data = @src.SpatialDeconData::create(["g1", "g2"], ["s1"], [[6.0], [12.0]]) catch { + _ => abort("SpatialDecon scaling data should be valid") + } + let plain = @src.spatial_decon(data, profile, config=sd_test_config()) catch { + _ => abort("SpatialDecon unscaled fit should succeed") + } + let scaled_config = @src.SpatialDeconConfig::create( + rescale_profile=true, + profile_quantile=1.0, + profile_target=2.0, + refit_outliers=false, + tolerance=1.0e-9, + ) catch { + _ => abort("SpatialDecon scaling configuration should be valid") + } + let scaled = @src.spatial_decon(data, profile, config=scaled_config) catch { + _ => abort("SpatialDecon scaled fit should succeed") + } + sd_test_close(plain.beta[0][0], 3.0, 1.0e-6) + sd_test_close(scaled.beta[0][0], 6.0, 1.0e-6) +} + +///| +test "SpatialDecon: zero profile quantile is rejected during rescaling" { + let profile = @src.SpatialDeconProfile::create(["g1", "g2", "g3"], ["A"], [ + [0.0], + [0.0], + [1.0], + ]) catch { + _ => abort("SpatialDecon sparse profile should be valid") + } + let data = @src.SpatialDeconData::create(["g1", "g2", "g3"], ["s1"], [ + [1.0], + [1.0], + [1.0], + ]) catch { + _ => abort("SpatialDecon sparse data should be valid") + } + let config = @src.SpatialDeconConfig::create(profile_quantile=0.5) catch { + _ => abort("SpatialDecon sparse configuration should be valid") + } + let failed = try { + ignore(@src.spatial_decon(data, profile, config~)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(failed) +} + +///| +test "SpatialDecon: precision weights favor high-confidence genes" { + let profile = @src.SpatialDeconProfile::create(["g1", "g2"], ["A"], [ + [1.0], + [1.0], + ]) catch { + _ => abort("SpatialDecon weighted profile should be valid") + } + let plain_data = @src.SpatialDeconData::create(["g1", "g2"], ["s1"], [ + [2.0], + [8.0], + ]) catch { + _ => abort("SpatialDecon unweighted data should be valid") + } + let weighted_data = @src.SpatialDeconData::create( + ["g1", "g2"], + ["s1"], + [[2.0], [8.0]], + weights=[[100.0], [1.0]], + ) catch { + _ => abort("SpatialDecon weighted data should be valid") + } + let plain = @src.spatial_decon(plain_data, profile, config=sd_test_config()) catch { + _ => abort("SpatialDecon unweighted fit should succeed") + } + let weighted = @src.spatial_decon( + weighted_data, + profile, + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon weighted fit should succeed") + } + sd_test_close(plain.beta[0][0], 4.0, 1.0e-5) + assert_true(weighted.beta[0][0] < 2.1) + assert_true(weighted.beta[0][0] > 2.0) +} + +///| +test "SpatialDecon: two-stage fitting flags and removes a log outlier" { + let profile = @src.SpatialDeconProfile::create( + ["g1", "g2", "g3", "g4", "g5"], + ["A"], + [[1.0], [1.0], [1.0], [1.0], [1.0]], + ) catch { + _ => abort("SpatialDecon outlier profile should be valid") + } + let data = @src.SpatialDeconData::create( + ["g1", "g2", "g3", "g4", "g5"], + ["s1"], + [[2.0], [2.0], [2.0], [2.0], [1024.0]], + ) catch { + _ => abort("SpatialDecon outlier data should be valid") + } + let result = @src.spatial_decon( + data, + profile, + config=sd_test_config(refit_outliers=true), + ) catch { + _ => abort("SpatialDecon outlier fit should succeed") + } + sd_test_close(result.beta[0][0], 2.0, 1.0e-6) + assert_eq(result.outliers[4][0], true) + assert_eq(result.outliers[0][0], false) + assert_true(result.residuals[4][0] != result.residuals[4][0]) +} + +///| +test "SpatialDecon: disabling refit preserves every observation" { + let profile = @src.SpatialDeconProfile::create( + ["g1", "g2", "g3", "g4", "g5"], + ["A"], + [[1.0], [1.0], [1.0], [1.0], [1.0]], + ) catch { + _ => abort("SpatialDecon outlier profile should be valid") + } + let data = @src.SpatialDeconData::create( + ["g1", "g2", "g3", "g4", "g5"], + ["s1"], + [[2.0], [2.0], [2.0], [2.0], [1024.0]], + ) catch { + _ => abort("SpatialDecon outlier data should be valid") + } + let result = @src.spatial_decon(data, profile, config=sd_test_config()) catch { + _ => abort("SpatialDecon non-refit fit should succeed") + } + for row in result.outliers { + assert_eq(row[0], false) + } + assert_true(result.beta[0][0] > 2.0) +} + +///| +test "SpatialDecon: synthetic multi-cell mixtures recover known abundances" { + let result = sd_test_example_result() + let expected = [ + [4.0, 1.0, 2.0, 0.5, 3.0], + [1.0, 4.0, 2.0, 1.0, 0.5], + [0.5, 1.0, 2.0, 4.0, 3.0], + ] + for cell_type in 0..= 0.0) + assert_true(result.rmse[spot] < 0.03) + assert_true(result.correlations[spot] > 0.99) + for cell_type in 0.. 0.0) + assert_true(sd_test_finite(result.t_statistics[cell_type][spot])) + assert_true(result.p_values[cell_type][spot] >= 0.0) + assert_true(result.p_values[cell_type][spot] <= 1.0) + assert_true(result.covariances[spot][cell_type][cell_type] > 0.0) + } + } +} + +///| +test "SpatialDecon: proportions sum to one per spot" { + let result = sd_test_example_result() + for spot in 0.. maximum { + maximum = total + } + } + sd_test_close(maximum, 100.0, 1.0e-10) +} + +///| +test "SpatialDecon: nuclei counts scale cells-per-100 by spot" { + let result = sd_test_exact_result(nuclei_counts=[100.0, 80.0]) + sd_test_close(result.cells_per_100[0][0], 50.0, 1.0e-5) + sd_test_close(result.cells_per_100[0][1], 100.0, 1.0e-5) + sd_test_close(result.cell_counts[0][0], 50.0, 1.0e-5) + sd_test_close(result.cell_counts[0][1], 80.0, 1.0e-5) +} + +///| +test "SpatialDecon: omitted nuclei counts leave cell counts empty" { + let result = sd_test_exact_result() + assert_eq(result.cell_counts.length(), 0) +} + +///| +test "SpatialDecon: nuclei counts are strictly validated" { + let length_failed = try { + ignore( + @src.spatial_decon( + sd_test_exact_data(), + sd_test_exact_profile(), + nuclei_counts=[1.0], + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let value_failed = try { + ignore( + @src.spatial_decon( + sd_test_exact_data(), + sd_test_exact_profile(), + nuclei_counts=[1.0, -1.0], + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(length_failed && value_failed) +} + +///| +test "SpatialDecon: data and profile align by gene name and data order" { + let profile = @src.SpatialDeconProfile::create(["extra", "g2", "g1"], ["A"], [ + [5.0], + [2.0], + [1.0], + ]) catch { + _ => abort("SpatialDecon reordered profile should be valid") + } + let data = @src.SpatialDeconData::create(["g1", "g2"], ["s1"], [[3.0], [6.0]]) catch { + _ => abort("SpatialDecon reordered data should be valid") + } + let result = @src.spatial_decon(data, profile, config=sd_test_config()) catch { + _ => abort("SpatialDecon reordered fit should succeed") + } + assert_eq(result.gene_names, ["g1", "g2"]) + sd_test_close(result.beta[0][0], 3.0, 1.0e-6) +} + +///| +test "SpatialDecon: fitting rejects insufficient shared genes" { + let no_shared_profile = @src.SpatialDeconProfile::create(["other"], ["A"], [ + [1.0], + ]) catch { + _ => abort("SpatialDecon disjoint profile should be valid") + } + let no_shared_failed = try { + ignore( + @src.spatial_decon( + sd_test_exact_data(), + no_shared_profile, + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let underdetermined_profile = @src.SpatialDeconProfile::create( + ["g1", "other"], + ["A", "B"], + [[1.0, 1.0], [1.0, 2.0]], + ) catch { + _ => abort("SpatialDecon underdetermined profile should be valid") + } + let underdetermined_failed = try { + ignore( + @src.spatial_decon( + sd_test_exact_data(), + underdetermined_profile, + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(no_shared_failed && underdetermined_failed) +} + +///| +test "SpatialDecon: fitting rejects all-zero expression" { + let data = @src.SpatialDeconData::create(["g1", "g2", "g3"], ["s1"], [ + [0.0], + [0.0], + [0.0], + ]) catch { + _ => abort("SpatialDecon zero data constructor should succeed") + } + let failed = try { + ignore( + @src.spatial_decon(data, sd_test_exact_profile(), config=sd_test_config()), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(failed) +} + +///| +test "SpatialDecon: abundance query and ranking expose fitted values" { + let result = sd_test_example_result() + match result.abundance_for("T_cell", "region_1") { + Some(value) => sd_test_close(value, result.beta[0][0], 0.0) + None => abort("SpatialDecon known abundance should exist") + } + assert_eq(result.abundance_for("missing", "region_1"), None) + let ranked = result.top_cell_types("region_1", limit=2) catch { + _ => abort("SpatialDecon ranking should succeed") + } + assert_eq(ranked.length(), 2) + assert_eq(ranked[0].cell_type, "T_cell") + assert_true(ranked[0].abundance >= ranked[1].abundance) +} + +///| +test "SpatialDecon: ranking validates spot and limit" { + let result = sd_test_example_result() + let spot_failed = try { + ignore(result.top_cell_types("missing")) + false + } catch { + SpatialDeconError(_) => true + } + let limit_failed = try { + ignore(result.top_cell_types("region_1", limit=0)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(spot_failed && limit_failed) +} + +///| +test "SpatialDecon: summary reports dimensions and convergence" { + let summary = sd_test_example_result().summary() + assert_true(summary.contains("genes=9")) + assert_true(summary.contains("spots=5")) + assert_true(summary.contains("cell_types=3")) + assert_true(summary.contains("converged=5/5")) +} + +///| +test "SpatialDecon: cell merge constructor defensively copies members" { + let members = ["A", "B"] + let merge = @src.SpatialDeconCellMerge::create("AB", members) catch { + _ => abort("SpatialDecon merge should be valid") + } + members[0] = "changed" + assert_eq(merge.name, "AB") + assert_eq(merge.members, ["A", "B"]) +} + +///| +test "SpatialDecon: collapse sums abundance proportions and counts" { + let result = sd_test_example_result() + let lymphoid = @src.SpatialDeconCellMerge::create("Lymphoid", [ + "T_cell", "B_cell", + ]) catch { + _ => abort("SpatialDecon lymphoid merge should be valid") + } + let myeloid = @src.SpatialDeconCellMerge::create("Myeloid", ["Myeloid"]) catch { + _ => abort("SpatialDecon myeloid merge should be valid") + } + let collapsed = @src.collapse_spatial_decon(result, [lymphoid, myeloid]) catch { + _ => abort("SpatialDecon collapse should succeed") + } + assert_eq(collapsed.cell_types, ["Lymphoid", "Myeloid"]) + for spot in 0.. abort("SpatialDecon lymphoid merge should be valid") + } + let collapsed = @src.collapse_spatial_decon(result, [lymphoid]) catch { + _ => abort("SpatialDecon covariance collapse should succeed") + } + let expected = result.covariances[0][0][0] + + result.covariances[0][0][1] + + result.covariances[0][1][0] + + result.covariances[0][1][1] + sd_test_close(collapsed.covariances[0][0][0], expected, 1.0e-10) + sd_test_close(collapsed.standard_errors[0][0], expected.sqrt(), 1.0e-10) +} + +///| +test "SpatialDecon: collapse preserves gene-level fit diagnostics" { + let result = sd_test_example_result() + let all = @src.SpatialDeconCellMerge::create("All", [ + "T_cell", "B_cell", "Myeloid", + ]) catch { + _ => abort("SpatialDecon complete merge should be valid") + } + let collapsed = @src.collapse_spatial_decon(result, [all]) catch { + _ => abort("SpatialDecon complete collapse should succeed") + } + assert_eq(collapsed.fitted, result.fitted) + assert_eq(collapsed.residuals, result.residuals) + assert_eq(collapsed.objectives, result.objectives) + assert_eq(collapsed.converged, result.converged) +} + +///| +test "SpatialDecon: collapse rejects empty unknown and repeated members" { + let result = sd_test_example_result() + let empty_failed = try { + ignore(@src.collapse_spatial_decon(result, [])) + false + } catch { + SpatialDeconError(_) => true + } + let unknown = @src.SpatialDeconCellMerge::create("Unknown", ["missing"]) catch { + _ => abort("SpatialDecon unknown merge constructor should succeed") + } + let unknown_failed = try { + ignore(@src.collapse_spatial_decon(result, [unknown])) + false + } catch { + SpatialDeconError(_) => true + } + let first = @src.SpatialDeconCellMerge::create("First", ["T_cell"]) catch { + _ => abort("SpatialDecon first merge should be valid") + } + let second = @src.SpatialDeconCellMerge::create("Second", ["T_cell"]) catch { + _ => abort("SpatialDecon second merge should be valid") + } + let repeated_failed = try { + ignore(@src.collapse_spatial_decon(result, [first, second])) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(empty_failed && unknown_failed && repeated_failed) +} + +///| +test "SpatialDecon: merge constructor rejects invalid names and members" { + let name_failed = try { + ignore(@src.SpatialDeconCellMerge::create("", ["A"])) + false + } catch { + SpatialDeconError(_) => true + } + let members_failed = try { + ignore(@src.SpatialDeconCellMerge::create("AB", ["A", "A"])) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(name_failed && members_failed) +} + +///| +test "SpatialDecon: probe-pool background averages negative probes" { + let background = @src.spatial_decon_background( + ["Neg_A1", "Neg_A2", "Marker_A", "Neg_B", "Marker_B"], + ["s1", "s2"], + [[1.0, 3.0], [3.0, 5.0], [10.0, 12.0], [2.0, 4.0], [20.0, 24.0]], + ["A", "A", "A", "B", "B"], + ["Neg_A1", "Neg_A2", "Neg_B"], + ) catch { + _ => abort("SpatialDecon background estimation should succeed") + } + assert_eq(background[0], [2.0, 4.0]) + assert_eq(background[2], [2.0, 4.0]) + assert_eq(background[3], [2.0, 4.0]) + assert_eq(background[4], [2.0, 4.0]) +} + +///| +test "SpatialDecon: probe-pool background validates pool coverage" { + let length_failed = try { + ignore( + @src.spatial_decon_background( + ["Neg", "Marker"], + ["s1"], + [[1.0], [2.0]], + ["A"], + ["Neg"], + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let coverage_failed = try { + ignore( + @src.spatial_decon_background( + ["Neg", "Marker"], + ["s1"], + [[1.0], [2.0]], + ["A", "B"], + ["Neg"], + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(length_failed && coverage_failed) +} + +///| +test "SpatialDecon: single-cell counts create mean cell-type profiles" { + let profile = @src.create_spatial_decon_profile( + ["g1", "g2"], + ["a1", "a2", "b1", "b2"], + ["A", "A", "B", "B"], + [[10.0, 8.0, 1.0, 2.0], [1.0, 2.0, 9.0, 11.0]], + scaling_factor=2.0, + min_cells=1, + min_genes=0, + ) catch { + _ => abort("SpatialDecon profile construction should succeed") + } + assert_eq(profile.gene_names, ["g1", "g2"]) + assert_eq(profile.cell_types, ["A", "B"]) + assert_eq(profile.values, [[18.0, 3.0], [3.0, 20.0]]) +} + +///| +test "SpatialDecon: single-cell profile supports normalization and gene filter" { + let profile = @src.create_spatial_decon_profile( + ["g1", "g2", "g3"], + ["a1", "a2", "b1", "b2"], + ["A", "A", "B", "B"], + [[10.0, 8.0, 1.0, 2.0], [1.0, 2.0, 9.0, 11.0], [2.0, 2.0, 2.0, 2.0]], + normalize=true, + min_cells=1, + min_genes=0, + gene_filter=["g1", "g2"], + ) catch { + _ => abort("SpatialDecon normalized profile should succeed") + } + assert_eq(profile.gene_names, ["g1", "g2"]) + for row in profile.values { + for value in row { + assert_true(sd_test_finite(value) && value > 0.0) + } + } +} + +///| +test "SpatialDecon: single-cell profile filters undersized cell types" { + let failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1", "g2"], + ["a1", "a2"], + ["A", "A"], + [[1.0, 2.0], [2.0, 1.0]], + min_cells=2, + min_genes=0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(failed) +} + +///| +test "SpatialDecon: single-cell profile rejects invalid metadata and controls" { + let metadata_failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1"], + ["c1", "c2"], + ["A"], + [[1.0, 2.0]], + min_cells=0, + min_genes=0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let scaling_failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1"], + ["c1"], + ["A"], + [[1.0]], + scaling_factor=0.0, + min_cells=0, + min_genes=0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let threshold_failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1"], + ["c1"], + ["A"], + [[1.0]], + min_cells=-1, + min_genes=0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(metadata_failed && scaling_failed && threshold_failed) +} + +///| +test "SpatialDecon: single-cell profile rejects empty libraries and gene filters" { + let library_failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1"], + ["c1"], + ["A"], + [[0.0]], + min_cells=0, + min_genes=0, + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let gene_failed = try { + ignore( + @src.create_spatial_decon_profile( + ["g1"], + ["c1"], + ["A"], + [[1.0]], + min_cells=0, + min_genes=0, + gene_filter=["missing"], + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(library_failed && gene_failed) +} + +///| +test "SpatialDecon: reverse deconvolution fits genes from varying scores" { + let (data, _, _) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon reverse data should be valid") + } + let result = sd_test_example_result() + let reverse = @src.reverse_spatial_decon(data, result) catch { + _ => abort("SpatialDecon reverse fit should succeed") + } + assert_eq(reverse.gene_names, data.gene_names) + assert_eq(reverse.spot_ids, data.spot_ids) + assert_eq(reverse.cell_types, result.cell_types) + assert_eq(reverse.coefficients.length(), data.n_genes()) + assert_eq(reverse.coefficients[0].length(), result.n_cell_types() + 1) + for row in reverse.coefficients { + for value in row { + assert_true(value >= 0.0 && sd_test_finite(value)) + } + } +} + +///| +test "SpatialDecon: reverse deconvolution returns residual diagnostics" { + let (data, _, _) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon reverse data should be valid") + } + let reverse = @src.reverse_spatial_decon(data, sd_test_example_result()) catch { + _ => abort("SpatialDecon reverse diagnostics should succeed") + } + assert_eq(reverse.fitted.length(), data.n_genes()) + assert_eq(reverse.fitted[0].length(), data.n_spots()) + for gene in 0..= 0.0) + assert_true(sd_test_finite(reverse.correlations[gene])) + assert_true(reverse.objectives[gene] >= 0.0) + } +} + +///| +test "SpatialDecon: reverse deconvolution validates spot identity and epsilon" { + let result = sd_test_example_result() + let (source, _, _) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon reverse source should be valid") + } + let mismatched = @src.SpatialDeconData::create( + source.gene_names, + ["x1", "x2", "x3", "x4", "x5"], + source.values, + background=source.background, + ) catch { + _ => abort("SpatialDecon mismatched reverse data should be valid") + } + let spots_failed = try { + ignore(@src.reverse_spatial_decon(mismatched, result)) + false + } catch { + SpatialDeconError(_) => true + } + let epsilon_failed = try { + ignore(@src.reverse_spatial_decon(source, result, epsilon=-1.0)) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(spots_failed && epsilon_failed) +} + +///| +test "SpatialDecon: reverse deconvolution requires varying cell scores" { + let result = sd_test_exact_result() + let constant_data = @src.SpatialDeconData::create( + ["g1", "g2", "g3"], + ["spot_1", "spot_2"], + [[7.0, 7.0], [9.0, 9.0], [11.0, 11.0]], + background=[[5.0, 5.0], [5.0, 5.0], [5.0, 5.0]], + ) catch { + _ => abort("SpatialDecon constant data should be valid") + } + let constant_result = @src.spatial_decon( + constant_data, + sd_test_exact_profile(), + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon constant fit should succeed") + } + let failed = try { + ignore(@src.reverse_spatial_decon(constant_data, constant_result)) + false + } catch { + SpatialDeconError(_) => true + } + ignore(result) + assert_true(failed) +} + +///| +test "SpatialDecon: SpatialExperiment wrapper writes scores and metadata" { + let (experiment, profile, nuclei_counts) = sd_test_experiment() + let output = @src.spatial_decon_spatial_experiment( + experiment, + profile, + background_assay_name="background", + weight_assay_name="weights", + nuclei_counts~, + config=sd_test_config(refit_outliers=true), + ) catch { + _ => abort("SpatialDecon SpatialExperiment wrapper should succeed") + } + assert_eq(output.experiment.metadata["spatialdecon_genes"], "9") + assert_eq(output.experiment.metadata["spatialdecon_spots"], "5") + assert_eq(output.experiment.metadata["spatialdecon_cell_types"], "3") + assert_true(output.experiment.col_data[0].contains("SpatialDecon:T_cell")) + assert_true( + output.experiment.col_data[0].contains("SpatialDeconProp:Myeloid"), + ) + assert_eq(output.result.cell_counts.length(), 3) +} + +///| +test "SpatialDecon: SpatialExperiment wrapper does not mutate input" { + let (experiment, profile, nuclei_counts) = sd_test_experiment() + let original_value = experiment.assay["counts"][0][0] + let output = @src.spatial_decon_spatial_experiment( + experiment, + profile, + background_assay_name="background", + nuclei_counts~, + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon immutable wrapper should succeed") + } + assert_true(!experiment.col_data[0].contains("SpatialDecon:T_cell")) + assert_true(!experiment.metadata.contains("spatialdecon_genes")) + output.experiment.assay["counts"][0][0] = 999.0 + output.experiment.col_data[0]["sample"] = "changed" + output.experiment.metadata["owner"] = "changed" + assert_eq(experiment.assay["counts"][0][0], original_value) + assert_eq(experiment.col_data[0]["sample"], "sample_1") + assert_eq(experiment.metadata["owner"], "input") +} + +///| +test "SpatialDecon: SpatialExperiment wrapper generates missing identifiers" { + let experiment = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(experiment, "counts", [[2.0, 4.0], [4.0, 8.0]])) + let profile = @src.SpatialDeconProfile::create(["gene_1", "gene_2"], ["A"], [ + [1.0], + [2.0], + ]) catch { + _ => abort("SpatialDecon generated-name profile should be valid") + } + let output = @src.spatial_decon_spatial_experiment( + experiment, + profile, + config=sd_test_config(), + ) catch { + _ => abort("SpatialDecon generated identifiers should succeed") + } + assert_eq(output.result.gene_names, ["gene_1", "gene_2"]) + assert_eq(output.result.spot_ids, ["spot_1", "spot_2"]) + assert_eq(output.experiment.col_data.length(), 2) +} + +///| +test "SpatialDecon: SpatialExperiment wrapper validates assay names" { + let (experiment, profile, _) = sd_test_experiment() + let assay_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + experiment, + profile, + assay_name="missing", + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let background_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + experiment, + profile, + background_assay_name="missing", + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let weights_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + experiment, + profile, + weight_assay_name="missing", + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(assay_failed && background_failed && weights_failed) +} + +///| +test "SpatialDecon: SpatialExperiment wrapper validates annotations" { + let (experiment, profile, _) = sd_test_experiment() + let broken_rows = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(broken_rows, "counts", experiment.assay["counts"])) + ignore(@src.se_add_row(broken_rows, Map([("gene_id", "g1")]))) + let rows_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + broken_rows, + profile, + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let missing_key = @src.SpatialExperiment::new() + ignore(@src.se_add_assay(missing_key, "counts", experiment.assay["counts"])) + for _ in 0..<9 { + ignore(@src.se_add_row(missing_key, Map([("symbol", "x")]))) + } + let key_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + missing_key, + profile, + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(rows_failed && key_failed) +} + +///| +test "SpatialDecon: SpatialExperiment wrapper validates prefixes and background" { + let (experiment, profile, _) = sd_test_experiment() + let prefix_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + experiment, + profile, + abundance_prefix="", + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + let background_failed = try { + ignore( + @src.spatial_decon_spatial_experiment( + experiment, + profile, + scalar_background=-1.0, + config=sd_test_config(), + ), + ) + false + } catch { + SpatialDeconError(_) => true + } + assert_true(prefix_failed && background_failed) +} + +///| +test "SpatialDecon: example fixture contains deterministic spatial mixtures" { + let (data, profile, nuclei_counts) = @src.spatial_decon_example_data() catch { + _ => abort("SpatialDecon example fixture should succeed") + } + assert_eq(data.n_genes(), 9) + assert_eq(data.n_spots(), 5) + assert_eq(profile.n_cell_types(), 3) + assert_eq(profile.cell_types, ["T_cell", "B_cell", "Myeloid"]) + assert_eq(nuclei_counts, [100.0, 120.0, 110.0, 90.0, 105.0]) + assert_true(data.background[0][4] > data.background[0][0]) +} diff --git a/test/moonbit/spia_test.mbt b/test/moonbit/spia_test.mbt index 750a5977..5792438a 100644 --- a/test/moonbit/spia_test.mbt +++ b/test/moonbit/spia_test.mbt @@ -170,7 +170,11 @@ test "spia_activation_status" { while i < results.results.length() { let r = results.results[i] // Activation status must be one of -1, 0, 1 - assert_true(r.activation_status == -1 || r.activation_status == 0 || r.activation_status == 1) + assert_true( + r.activation_status == -1 || + r.activation_status == 0 || + r.activation_status == 1, + ) i = i + 1 } } diff --git a/test/moonbit/spicyr_test.mbt b/test/moonbit/spicyr_test.mbt new file mode 100644 index 00000000..16b8cf7e --- /dev/null +++ b/test/moonbit/spicyr_test.mbt @@ -0,0 +1,1025 @@ +// Tests for the Bioconductor spicyR-inspired spatial association model. + +///| +fn spicyr_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "spicyR value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn spicyr_test_config( + edge_correction : Bool, + include_zero_cells : Bool, +) -> @src.SpicyConfig { + @src.SpicyConfig::create( + radii=[1.0], + edge_correction~, + include_zero_cells~, + window_padding=0.0, + tolerance=0.01, + ) catch { + _ => abort("spicyR test configuration should be valid") + } +} + +///| +fn spicyr_test_cell( + image_id : String, + cell_type : String, + x : Double, + y : Double, +) -> @src.SpicyCell { + @src.SpicyCell::create(image_id, cell_type, x, y) catch { + _ => abort("spicyR test cell should be valid") + } +} + +///| +fn spicyr_test_image_metadata( + image_id : String, + condition : String, + subject : String, + covariates : Map[String, Double], +) -> @src.SpicyImageMetadata { + @src.SpicyImageMetadata::create(image_id, condition, subject~, covariates~) catch { + _ => abort("spicyR test metadata should be valid") + } +} + +///| +fn spicyr_test_pair( + from_cell_type : String, + to_cell_type : String, +) -> @src.SpicyPair { + @src.SpicyPair::create(from_cell_type, to_cell_type) catch { + _ => abort("spicyR test pair should be valid") + } +} + +///| +fn spicyr_test_metadata( + count : Int, + split : Int, +) -> Array[@src.SpicyImageMetadata] { + let metadata : Array[@src.SpicyImageMetadata] = [] + for index in 0.. Array[@src.SpicyCell] { + let cells : Array[@src.SpicyCell] = [] + for image in metadata { + cells.push(spicyr_test_cell(image.image_id, "A", source_x, source_y)) + cells.push(spicyr_test_cell(image.image_id, "B", target_x, target_y)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 0.0, 0.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 10.0, 10.0)) + } + cells +} + +///| +fn spicyr_test_association( + image_id : String, + from_type : String, + to_type : String, + statistic : Double, + count : Int, +) -> @src.SpicyAssociation { + @src.SpicyAssociation::create( + image_id, + from_type, + to_type, + statistic, + from_count=count, + to_count=count, + ) catch { + _ => abort("spicyR test association should be valid") + } +} + +///| +fn spicyr_test_linear_input() -> ( + Array[@src.SpicyImageMetadata], + Array[@src.SpicyAssociation], +) { + let metadata = spicyr_test_metadata(8, 4) + let associations : Array[@src.SpicyAssociation] = [] + let noise = [-0.2, 0.1, 0.0, 0.15, -0.1, 0.2, -0.05, 0.1] + for index in 0.. ( + Array[@src.SpicyImageMetadata], + Array[@src.SpicyAssociation], +) { + let metadata : Array[@src.SpicyImageMetadata] = [] + let associations : Array[@src.SpicyAssociation] = [] + let subject_effects = [-1.0, -0.25, 0.5, 1.25, 2.0] + for subject_index in 0.. @src.SpicyResult { + let (metadata, first) = spicyr_test_linear_input() + let associations = first.copy() + let weak = [0.1, -0.2, 0.05, 0.0, -0.1, 0.2, -0.05, 0.1] + for index in 0.. abort("spicyR two-pair fit should succeed") + } +} + +///| +fn spicyr_test_spatial_experiment() -> @src.SpatialExperiment { + let experiment = @src.SpatialExperiment::new() + for image_index in 0..<6 { + let image_id = "image_" + (image_index + 1).to_string() + let condition = if image_index < 3 { "control" } else { "treated" } + let target_x = if condition == "control" { 9.0 } else { 5.5 } + let cell_types = ["A", "B", "Frame", "Frame"] + let xs = [5.0, target_x, 0.0, 10.0] + let ys = [5.0, if condition == "control" { 9.0 } else { 5.0 }, 0.0, 10.0] + for cell_index in 0.. abort("custom spicyR configuration should be valid") + } + radii[0] = 9.0 + assert_eq(config.radii[0], 1.0) + assert_true(!config.edge_correction) + assert_true(config.include_zero_cells) + assert_true(!config.use_weights) + assert_eq(config.reference_condition, "control") +} + +///| +test "spicyR: constructors preserve typed values" { + let cell = @src.SpicyCell::create("image", "T_cell", 1.5, 2.5) catch { + _ => abort("spicyR cell should be valid") + } + let pair = @src.SpicyPair::create("T_cell", "Tumour") catch { + _ => abort("spicyR pair should be valid") + } + let association = @src.SpicyAssociation::create( + "image", + "T_cell", + "Tumour", + 1.25, + from_count=4, + to_count=6, + ) catch { + _ => abort("spicyR association should be valid") + } + assert_eq(cell.image_id, "image") + assert_eq(cell.x, 1.5) + assert_eq(pair.label(), "T_cell__Tumour") + assert_eq(association.statistic, 1.25) + assert_eq(association.from_count, 4) +} + +///| +test "spicyR: metadata constructor copies covariates" { + let covariates : Map[String, Double] = Map([("age", 52.0)]) + let metadata = @src.SpicyImageMetadata::create( + "image", + "control", + subject="subject_1", + covariates~, + ) catch { + _ => abort("spicyR metadata should be valid") + } + covariates["age"] = 90.0 + assert_eq(metadata.subject, "subject_1") + assert_eq(metadata.covariates["age"], 52.0) +} + +///| +test "spicyR: constructors reject empty and non-finite values" { + let bad_cell = try { + ignore(@src.SpicyCell::create("", "A", 0.0, 0.0)) + false + } catch { + SpicyError(_) => true + } + let bad_coordinate = try { + ignore(@src.SpicyCell::create("image", "A", @double.not_a_number, 0.0)) + false + } catch { + SpicyError(_) => true + } + let bad_pair = try { + ignore(@src.SpicyPair::create("A", "")) + false + } catch { + SpicyError(_) => true + } + let bad_association = try { + ignore(@src.SpicyAssociation::create("image", "A", "B", 0.0, from_count=-1)) + false + } catch { + SpicyError(_) => true + } + assert_true(bad_cell) + assert_true(bad_coordinate) + assert_true(bad_pair) + assert_true(bad_association) +} + +///| +test "spicyR: configuration rejects invalid radii" { + let empty = try { + ignore(@src.SpicyConfig::create(radii=[])) + false + } catch { + SpicyError(_) => true + } + let non_positive = try { + ignore(@src.SpicyConfig::create(radii=[0.0, 1.0])) + false + } catch { + SpicyError(_) => true + } + let unsorted = try { + ignore(@src.SpicyConfig::create(radii=[2.0, 1.0])) + false + } catch { + SpicyError(_) => true + } + assert_true(empty) + assert_true(non_positive) + assert_true(unsorted) +} + +///| +test "spicyR: configuration rejects invalid fitting controls" { + let bad_fdr = try { + ignore(@src.SpicyConfig::create(fdr_threshold=1.1)) + false + } catch { + SpicyError(_) => true + } + let bad_weight = try { + ignore(@src.SpicyConfig::create(weight_factor=-1.0)) + false + } catch { + SpicyError(_) => true + } + let bad_iterations = try { + ignore(@src.SpicyConfig::create(max_iterations=0)) + false + } catch { + SpicyError(_) => true + } + let bad_tolerance = try { + ignore(@src.SpicyConfig::create(tolerance=1.0)) + false + } catch { + SpicyError(_) => true + } + assert_true(bad_fdr) + assert_true(bad_weight) + assert_true(bad_iterations) + assert_true(bad_tolerance) +} + +///| +test "spicyR: cross-L matches a deterministic rectangular window" { + let metadata = spicyr_test_metadata(3, 1) + let cells = spicyr_test_geometry_cells(metadata, 5.0, 5.0, 5.5, 5.0) + let pair = spicyr_test_pair("A", "B") + let associations = @src.spicy_pairwise( + cells, + metadata, + pair, + config=spicyr_test_config(false, false), + ) catch { + _ => abort("spicyR pairwise geometry should succeed") + } + let expected_l = (100.0 / 3.141592653589793).sqrt() + assert_eq(associations.length(), 3) + spicyr_test_close(associations[0].l_values[0], expected_l, 1.0e-10) + spicyr_test_close(associations[0].statistic, expected_l - 1.0, 1.0e-10) + assert_true(associations[0].observed) + assert_true(associations[0].available) +} + +///| +test "spicyR: edge correction increases boundary-source cross-L" { + let metadata = spicyr_test_metadata(3, 1) + let cells = spicyr_test_geometry_cells(metadata, 0.0, 0.0, 0.5, 0.0) + let pair = spicyr_test_pair("A", "B") + let plain = @src.spicy_pairwise( + cells, + metadata, + pair, + config=spicyr_test_config(false, false), + ) catch { + _ => abort("uncorrected spicyR pairwise fit should succeed") + } + let corrected = @src.spicy_pairwise( + cells, + metadata, + pair, + config=spicyr_test_config(true, false), + ) catch { + _ => abort("edge-corrected spicyR pairwise fit should succeed") + } + assert_true(corrected[0].l_values[0] > plain[0].l_values[0] * 1.9) +} + +///| +test "spicyR: ordered pairs retain directional edge correction" { + let metadata = spicyr_test_metadata(3, 1) + let cells = spicyr_test_geometry_cells(metadata, 0.0, 0.0, 0.5, 0.0) + let forward = @src.spicy_pairwise( + cells, + metadata, + spicyr_test_pair("A", "B"), + config=spicyr_test_config(true, false), + ) catch { + _ => abort("forward spicyR pair should succeed") + } + let reverse = @src.spicy_pairwise( + cells, + metadata, + spicyr_test_pair("B", "A"), + config=spicyr_test_config(true, false), + ) catch { + _ => abort("reverse spicyR pair should succeed") + } + assert_true(forward[0].l_values[0] > reverse[0].l_values[0]) +} + +///| +test "spicyR: same-type cross-L excludes each cell from itself" { + let metadata = spicyr_test_metadata(3, 1) + let cells : Array[@src.SpicyCell] = [] + for image in metadata { + cells.push(spicyr_test_cell(image.image_id, "A", 5.0, 5.0)) + cells.push(spicyr_test_cell(image.image_id, "A", 5.5, 5.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 0.0, 0.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 10.0, 10.0)) + } + let associations = @src.spicy_pairwise( + cells, + metadata, + spicyr_test_pair("A", "A"), + config=spicyr_test_config(false, false), + ) catch { + _ => abort("same-type spicyR pair should succeed") + } + let expected_l = (50.0 / 3.141592653589793).sqrt() + spicyr_test_close(associations[0].l_values[0], expected_l, 1.0e-10) +} + +///| +test "spicyR: a single same-type cell retains the Poisson baseline" { + let metadata = spicyr_test_metadata(3, 1) + let cells : Array[@src.SpicyCell] = [] + for image in metadata { + cells.push(spicyr_test_cell(image.image_id, "A", 5.0, 5.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 0.0, 0.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 10.0, 10.0)) + } + let associations = @src.spicy_pairwise( + cells, + metadata, + spicyr_test_pair("A", "A"), + config=spicyr_test_config(false, false), + ) catch { + _ => abort("single-cell same-type pair should succeed") + } + assert_true(!associations[0].observed) + assert_true(associations[0].available) + assert_eq(associations[0].l_values[0], 0.0) + assert_eq(associations[0].statistic, -1.0) +} + +///| +test "spicyR: absent cell types follow include-zero-cells policy" { + let metadata = spicyr_test_metadata(3, 1) + let cells : Array[@src.SpicyCell] = [] + for index in 0.. 0 { + cells.push(spicyr_test_cell(image_id, "B", 5.5, 5.0)) + } + cells.push(spicyr_test_cell(image_id, "Frame", 0.0, 0.0)) + cells.push(spicyr_test_cell(image_id, "Frame", 10.0, 10.0)) + } + let pair = spicyr_test_pair("A", "B") + let omitted = @src.spicy_pairwise( + cells, + metadata, + pair, + config=spicyr_test_config(false, false), + ) catch { + _ => abort("spicyR missing-pair calculation should succeed") + } + let retained = @src.spicy_pairwise( + cells, + metadata, + pair, + config=spicyr_test_config(false, true), + ) catch { + _ => abort("spicyR zero-cell calculation should succeed") + } + assert_true(!omitted[0].available) + assert_true(retained[0].available) + assert_eq(retained[0].statistic, -1.0) +} + +///| +test "spicyR: radii are capped and duplicate caps collapse" { + let metadata = spicyr_test_metadata(3, 1) + let cells : Array[@src.SpicyCell] = [] + for image in metadata { + cells.push(spicyr_test_cell(image.image_id, "A", 0.5, 2.0)) + cells.push(spicyr_test_cell(image.image_id, "B", 1.0, 2.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 0.0, 0.0)) + cells.push(spicyr_test_cell(image.image_id, "Frame", 2.0, 4.0)) + } + let config = @src.SpicyConfig::create( + radii=[0.5, 5.0, 10.0], + edge_correction=false, + ) catch { + _ => abort("radius-cap configuration should be valid") + } + let associations = @src.spicy_pairwise( + cells, + metadata, + spicyr_test_pair("A", "B"), + config~, + ) catch { + _ => abort("radius-cap calculation should succeed") + } + assert_eq(associations[0].radii.length(), 2) + spicyr_test_close(associations[0].radii[1], 2.0 / 2.01, 1.0e-12) +} + +///| +test "spicyR: pairwise validation rejects unknown cell types" { + let metadata = spicyr_test_metadata(3, 1) + let cells = spicyr_test_geometry_cells(metadata, 5.0, 5.0, 5.5, 5.0) + let failed = try { + ignore( + @src.spicy_pairwise(cells, metadata, spicyr_test_pair("A", "missing")), + ) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: pairwise validation rejects duplicate image metadata" { + let metadata = spicyr_test_metadata(3, 1) + metadata[2] = metadata[1] + let cells = spicyr_test_geometry_cells(metadata, 5.0, 5.0, 5.5, 5.0) + let failed = try { + ignore(@src.spicy_pairwise(cells, metadata, spicyr_test_pair("A", "B"))) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: weighted linear model estimates a condition contrast" { + let (metadata, associations) = spicyr_test_linear_input() + let result = @src.spicy_fit_associations(associations, metadata) catch { + _ => abort("spicyR weighted linear fit should succeed") + } + let pair = match result.pair("A", "B") { + Some(value) => value + None => abort("spicyR pair should be present") + } + let condition = match pair.condition("treated") { + Some(value) => value + None => abort("spicyR treated contrast should be present") + } + assert_true(pair.model is @src.SpicyModelKind::WeightedLinear) + spicyr_test_close(condition.estimate, 2.0, 0.25) + assert_true(condition.standard_error > 0.0) + assert_true(condition.p_value >= 0.0 && condition.p_value <= 1.0) +} + +///| +test "spicyR: count-derived precision weights normalize to mean one" { + let (metadata, associations) = spicyr_test_linear_input() + let result = @src.spicy_fit_associations(associations, metadata) catch { + _ => abort("spicyR weighted fit should succeed") + } + let pair = result.pairs[0] + let mut total = 0.0 + for association in pair.associations { + total = total + association.weight + } + spicyr_test_close( + total / pair.associations.length().to_double(), + 1.0, + 1.0e-12, + ) + assert_true(pair.associations[0].weight < pair.associations[7].weight) +} + +///| +test "spicyR: disabling precision weights assigns unit weights" { + let (metadata, associations) = spicyr_test_linear_input() + let config = @src.SpicyConfig::create(use_weights=false) catch { + _ => abort("unweighted spicyR configuration should be valid") + } + let result = @src.spicy_fit_associations(associations, metadata, config~) catch { + _ => abort("unweighted spicyR fit should succeed") + } + for association in result.pairs[0].associations { + assert_eq(association.weight, 1.0) + } +} + +///| +test "spicyR: numeric covariates enter the fixed-effect design" { + let (metadata, associations) = spicyr_test_linear_input() + let result = @src.spicy_fit_associations(associations, metadata, covariate_names=[ + "age", + ]) catch { + _ => abort("spicyR covariate fit should succeed") + } + let condition = result.pairs[0].conditions[0] + assert_true(condition.estimate == condition.estimate) + assert_true(condition.standard_error > 0.0) +} + +///| +test "spicyR: repeated subjects select the random-intercept model" { + let (metadata, associations) = spicyr_test_random_input() + let result = @src.spicy_fit_associations(associations, metadata) catch { + _ => abort("spicyR random-intercept fit should succeed") + } + let pair = result.pairs[0] + let condition = pair.conditions[0] + assert_true(pair.model is @src.SpicyModelKind::WeightedRandomIntercept) + spicyr_test_close(condition.estimate, 1.5, 0.15) + assert_true(pair.random_intercept_variance > 0.1) + assert_true(pair.residual_variance >= 0.0) +} + +///| +test "spicyR: multiple conditions produce reference contrasts" { + let metadata : Array[@src.SpicyImageMetadata] = [] + let associations : Array[@src.SpicyAssociation] = [] + let levels = ["control", "drug_a", "drug_b"] + for level_index in 0.. abort("multi-condition spicyR configuration should be valid") + } + let result = @src.spicy_fit_associations(associations, metadata, config~) catch { + _ => abort("multi-condition spicyR fit should succeed") + } + assert_eq(result.pairs[0].conditions.length(), 2) + let drug_a = match result.pairs[0].condition("drug_a") { + Some(value) => value + None => abort("drug_a contrast should be present") + } + let drug_b = match result.pairs[0].condition("drug_b") { + Some(value) => value + None => abort("drug_b contrast should be present") + } + spicyr_test_close(drug_a.estimate, 1.0, 0.05) + spicyr_test_close(drug_b.estimate, 2.0, 0.05) +} + +///| +test "spicyR: BH correction is bounded and rank monotone" { + let result = spicyr_test_two_pair_result() + let first = result.pairs[0].conditions[0] + let second = result.pairs[1].conditions[0] + assert_true(first.adjusted_p_value >= first.p_value) + assert_true(second.adjusted_p_value >= second.p_value) + assert_true(first.adjusted_p_value <= 1.0) + assert_true(second.adjusted_p_value <= 1.0) + if first.p_value <= second.p_value { + assert_true(first.adjusted_p_value <= second.adjusted_p_value) + } else { + assert_true(second.adjusted_p_value <= first.adjusted_p_value) + } +} + +///| +test "spicyR: pair and association lookup preserve identities" { + let result = spicyr_test_two_pair_result() + let pair = match result.pair("A", "B") { + Some(value) => value + None => abort("spicyR A__B pair should exist") + } + let association = match pair.association("image_3") { + Some(value) => value + None => abort("spicyR image association should exist") + } + assert_eq(association.image_id, "image_3") + assert_true(result.pair("missing", "B") is None) + assert_true(pair.association("missing") is None) +} + +///| +test "spicyR: top ranks pairs and handles zero limits" { + let result = spicyr_test_two_pair_result() + let top = result.top("treated", limit=2) catch { + _ => abort("spicyR top query should succeed") + } + let first_p = top[0].conditions[0].p_value + let second_p = top[1].conditions[0].p_value + assert_true(first_p <= second_p) + let empty = result.top("treated", limit=0) catch { + _ => abort("spicyR empty top query should succeed") + } + assert_eq(empty.length(), 0) +} + +///| +test "spicyR: significant filtering and summary use configured FDR" { + let result = spicyr_test_two_pair_result() + let all = result.significant("treated", threshold=1.0) catch { + _ => abort("spicyR significant query should succeed") + } + assert_eq(all.length(), result.n_pairs()) + assert_eq(result.n_images(), 8) + assert_true(result.summary().contains("spicyR(images=8, pairs=2")) +} + +///| +test "spicyR: result queries reject absent contrasts and invalid limits" { + let result = spicyr_test_two_pair_result() + let bad_level = try { + ignore(result.top("control")) + false + } catch { + SpicyError(_) => true + } + let bad_limit = try { + ignore(result.top("treated", limit=-1)) + false + } catch { + SpicyError(_) => true + } + let bad_threshold = try { + ignore(result.significant("treated", threshold=1.1)) + false + } catch { + SpicyError(_) => true + } + assert_true(bad_level) + assert_true(bad_limit) + assert_true(bad_threshold) +} + +///| +test "spicyR: fitting rejects a single condition" { + let (metadata, associations) = spicyr_test_linear_input() + for index in 0.. true + } + assert_true(failed) +} + +///| +test "spicyR: fitting rejects duplicate pair-image associations" { + let (metadata, associations) = spicyr_test_linear_input() + associations.push(associations[0]) + let failed = try { + ignore(@src.spicy_fit_associations(associations, metadata)) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: fitting rejects missing covariates" { + let (metadata, associations) = spicyr_test_linear_input() + let failed = try { + ignore( + @src.spicy_fit_associations(associations, metadata, covariate_names=[ + "missing", + ]), + ) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: direct cell analysis returns requested pairs" { + let (cells, metadata, pairs) = @src.spicyr_example_data() + let config = @src.SpicyConfig::create( + radii=[0.5, 1.0, 2.0], + edge_correction=true, + reference_condition="control", + tolerance=0.01, + ) catch { + _ => abort("spicyR example configuration should be valid") + } + let result = @src.spicyr( + cells, + metadata, + pairs~, + covariate_names=["age"], + config~, + ) catch { + _ => abort("spicyR example analysis should succeed") + } + assert_eq(result.n_images(), 12) + assert_eq(result.n_pairs(), 1) + assert_true( + result.pairs[0].model is @src.SpicyModelKind::WeightedRandomIntercept, + ) + assert_true(result.pairs[0].conditions[0].estimate > 0.0) +} + +///| +test "spicyR: automatic pair generation includes ordered pairs" { + let (cells, metadata, _) = @src.spicyr_example_data() + let config = @src.SpicyConfig::create( + radii=[1.0], + edge_correction=false, + include_zero_cells=true, + use_weights=false, + tolerance=0.01, + ) catch { + _ => abort("automatic-pair configuration should be valid") + } + let result = @src.spicyr(cells, metadata, config~) catch { + _ => abort("automatic spicyR analysis should succeed") + } + assert_eq(result.n_pairs(), 4) + assert_true(result.pair("T_cell", "Tumour") is Some(_)) + assert_true(result.pair("Tumour", "T_cell") is Some(_)) +} + +///| +test "spicyR: direct analysis rejects duplicate requested pairs" { + let (cells, metadata, pairs) = @src.spicyr_example_data() + let duplicate = [pairs[0], pairs[0]] + let failed = try { + ignore(@src.spicyr(cells, metadata, pairs=duplicate)) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: SpatialExperiment integration preserves and enriches input" { + let experiment = spicyr_test_spatial_experiment() + experiment.metadata["source"] = "fixture" + let pair = spicyr_test_pair("A", "B") + let output = @src.spicyr_spatial_experiment( + experiment, + pairs=[pair], + config=spicyr_test_config(false, false), + ) catch { + _ => abort("spicyR SpatialExperiment integration should succeed") + } + assert_eq(output.result.n_images(), 6) + assert_eq(output.result.n_pairs(), 1) + assert_eq(output.experiment.metadata["source"], "fixture") + assert_eq(output.experiment.metadata["spicyr_images"], "6") + assert_eq(output.experiment.metadata["spicyr_pairs"], "1") + assert_true(!experiment.metadata.contains("spicyr_images")) +} + +///| +test "spicyR: SpatialExperiment supports subject and numeric covariate keys" { + let experiment = spicyr_test_spatial_experiment() + let pair = spicyr_test_pair("A", "B") + let output = @src.spicyr_spatial_experiment( + experiment, + subject_key="subject", + covariate_keys=["age"], + pairs=[pair], + config=spicyr_test_config(false, false), + ) catch { + _ => abort("spicyR keyed SpatialExperiment fit should succeed") + } + assert_true( + output.result.pairs[0].model is @src.SpicyModelKind::WeightedRandomIntercept, + ) +} + +///| +test "spicyR: SpatialExperiment rejects coordinate mismatch" { + let experiment = spicyr_test_spatial_experiment() + ignore(experiment.spatial_coords.pop()) + let failed = try { + ignore(@src.spicyr_spatial_experiment(experiment)) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: SpatialExperiment rejects missing required columns" { + let experiment = spicyr_test_spatial_experiment() + experiment.col_data[0].remove("cellType") + let failed = try { + ignore(@src.spicyr_spatial_experiment(experiment)) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: SpatialExperiment rejects image-varying metadata" { + let experiment = spicyr_test_spatial_experiment() + experiment.col_data[1]["condition"] = "other" + let failed = try { + ignore(@src.spicyr_spatial_experiment(experiment)) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} + +///| +test "spicyR: SpatialExperiment rejects nonnumeric covariates" { + let experiment = spicyr_test_spatial_experiment() + experiment.col_data[0]["age"] = "unknown" + let failed = try { + ignore(@src.spicyr_spatial_experiment(experiment, covariate_keys=["age"])) + false + } catch { + SpicyError(_) => true + } + assert_true(failed) +} diff --git a/test/moonbit/stage_r_test.mbt b/test/moonbit/stage_r_test.mbt index 5a075d3f..06748ee8 100644 --- a/test/moonbit/stage_r_test.mbt +++ b/test/moonbit/stage_r_test.mbt @@ -110,7 +110,11 @@ test "stage_r_simes_caps_at_one" { ///| test "stage_r_adjustment_holm" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) assert_eq(result.nGenes, 10) assert_eq(result.nHypotheses, 3) @@ -121,7 +125,11 @@ test "stage_r_adjustment_holm" { ///| test "stage_r_adjustment_none" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_none()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_none(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) assert_eq(result.nGenes, 10) } @@ -129,7 +137,11 @@ test "stage_r_adjustment_none" { ///| test "stage_r_adjustment_dte" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_dte()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_dte(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) assert_eq(result.nGenes, 10) assert_eq(result.nHypotheses, 3) @@ -138,7 +150,11 @@ test "stage_r_adjustment_dte" { ///| test "stage_r_adjustment_dtu" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_dtu()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_dtu(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) assert_eq(result.nGenes, 10) assert_eq(result.nHypotheses, 3) @@ -147,17 +163,27 @@ test "stage_r_adjustment_dtu" { ///| test "stage_r_alpha_adjusted" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) // alphaAdjusted = (R/G) * alpha - let expected = result.nSignificantGenes.to_double() / result.nGenes.to_double() * config.alpha + let expected = result.nSignificantGenes.to_double() / + result.nGenes.to_double() * + config.alpha assert_true((result.alphaAdjusted - expected).abs() < 1.0e-10) } ///| test "stage_r_significant_genes" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) let sigGenes = @src.get_significant_genes(result) assert_eq(sigGenes.length(), result.nSignificantGenes) @@ -173,7 +199,11 @@ test "stage_r_significant_genes" { ///| test "stage_r_significant_hypotheses" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) let sigHyps = @src.get_significant_hypotheses(result) // All pairs should have valid indices. @@ -189,7 +219,11 @@ test "stage_r_significant_hypotheses" { ///| test "stage_r_get_results" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) let mat = @src.get_results(result) assert_eq(mat.length(), result.nGenes) @@ -204,7 +238,11 @@ test "stage_r_get_results" { ///| test "stage_r_non_significant_genes_na" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.001, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.001, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) // Non-significant genes should have -1.0 in confirmation p-values. let mut i = 0 @@ -225,7 +263,11 @@ test "stage_r_pscreen_adjusted" { let (pScreen, pConfirmation) = @src.stage_r_sample() // Pre-adjust the screening p-values. let adjScreen = @src.stage_r_bh_adjust(pScreen) - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=true, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=true, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(adjScreen, pConfirmation, config) // pAdjScreen should be the same as input. let mut i = 0 @@ -238,7 +280,11 @@ test "stage_r_pscreen_adjusted" { ///| test "stage_r_confirmation_rescaled" { let (pScreen, pConfirmation) = @src.stage_r_sample() - let config = @src.StageRConfig::new(alpha=0.05, pScreenAdjusted=false, stageMethod=@src.stage_r_method_holm()) + let config = @src.StageRConfig::new( + alpha=0.05, + pScreenAdjusted=false, + stageMethod=@src.stage_r_method_holm(), + ) let result = @src.stage_wise_adjustment(pScreen, pConfirmation, config) // For significant genes, confirmation p-values should be rescaled. let mut i = 0 diff --git a/test/moonbit/statistics_test.mbt b/test/moonbit/statistics_test.mbt index c3cd357f..f735171c 100644 --- a/test/moonbit/statistics_test.mbt +++ b/test/moonbit/statistics_test.mbt @@ -1,59 +1,69 @@ // Tests for Bio.Statistics module +///| test "stat_mean" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let result = @src.stat_mean(data) assert_true(result > 2.9 && result < 3.1) } +///| test "stat_mean empty" { let data : Array[Double] = Array::new() let result = @src.stat_mean(data) assert_true(result == 0.0) } +///| test "stat_variance" { let data = [2.0, 4.0, 4.0, 4.0, 5.0, 5.0, 7.0, 9.0] let result = @src.stat_variance(data) assert_true(result > 4.0 && result < 5.0) } +///| test "stat_std" { let data = [2.0, 4.0, 4.0, 4.0, 5.0, 5.0, 7.0, 9.0] let result = @src.stat_std(data) assert_true(result > 2.0 && result < 2.5) } +///| test "stat_median odd" { let data = [1.0, 3.0, 2.0, 5.0, 4.0] let result = @src.stat_median(data) assert_true(result == 3.0) } +///| test "stat_median even" { let data = [1.0, 2.0, 3.0, 4.0] let result = @src.stat_median(data) assert_true(result == 2.5) } +///| test "stat_min" { let data = [3.0, 1.0, 4.0, 1.0, 5.0, 9.0, 2.0] let result = @src.stat_min(data) assert_true(result == 1.0) } +///| test "stat_max" { let data = [3.0, 1.0, 4.0, 1.0, 5.0, 9.0, 2.0] let result = @src.stat_max(data) assert_true(result == 9.0) } +///| test "stat_sum" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let result = @src.stat_sum(data) assert_true(result == 15.0) } +///| test "stat_cumsum" { let data = [1.0, 2.0, 3.0] let result = @src.stat_cumsum(data) @@ -63,6 +73,7 @@ test "stat_cumsum" { assert_true(result[2] == 6.0) } +///| test "stat_pearson_correlation positive" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -70,6 +81,7 @@ test "stat_pearson_correlation positive" { assert_true(result > 0.99 && result < 1.01) } +///| test "stat_pearson_correlation negative" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [10.0, 8.0, 6.0, 4.0, 2.0] @@ -77,22 +89,26 @@ test "stat_pearson_correlation negative" { assert_true(result < -0.99 && result > -1.01) } +///| test "stat_zscore" { let result = @src.stat_zscore(3.0, 2.0, 0.5) assert_true(result > 1.9 && result < 2.1) } +///| test "stat_normal_cdf at 0" { let result = @src.stat_normal_cdf(0.0) assert_true(result > 0.49 && result < 0.51) } +///| test "stat_t_statistic" { let sample = [2.0, 3.0, 4.0, 5.0, 6.0] let result = @src.stat_t_statistic(sample, 3.0) assert_true(result > 1.4 && result < 1.6) } +///| test "stat_confidence_interval" { let data = [2.0, 3.0, 4.0, 5.0, 6.0] let result = @src.stat_confidence_interval(data) @@ -100,6 +116,7 @@ test "stat_confidence_interval" { assert_true(result[0] < result[1]) } +///| test "stat_mann_whitney_u_basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -107,6 +124,7 @@ test "stat_mann_whitney_u_basic" { assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_wilcoxon_signed_rank_basic" { let x = [1.0, 2.0, 3.0, 4.0, 5.0] let y = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -114,6 +132,7 @@ test "stat_wilcoxon_signed_rank_basic" { assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_ks_test_basic" { let sample1 = [1.0, 2.0, 3.0, 4.0, 5.0] let sample2 = [2.0, 4.0, 6.0, 8.0, 10.0] @@ -121,12 +140,14 @@ test "stat_ks_test_basic" { assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_fisher_exact_basic" { let table = [[10, 2], [3, 5]] let result = @src.stat_fisher_exact(table) assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_bonferroni_correct_basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.stat_bonferroni(p_values) @@ -134,67 +155,79 @@ test "stat_bonferroni_correct_basic" { assert_true(result[0] >= 0.01) } +///| test "stat_holm_correct_basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.stat_holm(p_values) assert_true(result.length() == 5) } +///| test "stat_by_correct_basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.stat_by(p_values) assert_true(result.length() == 5) } +///| test "stat_chi2_cdf_basic" { // For df=3, P(chi2 <= 7.81) ~ 0.95 let result = @src.stat_chi2_cdf(7.81, 3) assert_true(result > 0.90 && result < 0.99) } +///| test "stat_chi2_cdf_zero" { let result = @src.stat_chi2_cdf(0.0, 3) assert_true(result == 0.0) } +///| test "stat_chi2_cdf_negative" { let result = @src.stat_chi2_cdf(-1.0, 5) assert_true(result == 0.0) } +///| test "stat_chi2_quantile_basic" { // For df=2, chi2_0.95 ~ 5.99 let result = @src.stat_chi2_quantile(0.95, 2) assert_true(result > 5.0 && result < 7.0) } +///| test "stat_chi2_quantile_zero_one" { assert_true(@src.stat_chi2_quantile(0.0, 2) == 0.0) assert_true(@src.stat_chi2_quantile(1.0, 2) > 1.0e100) } +///| test "stat_t_cdf_central" { let result = @src.stat_t_cdf(0.0, 10.0) assert_true(result > 0.49 && result < 0.51) } +///| test "stat_t_cdf_positive" { // t with df=10 at 2.228 ~ 0.975 let result = @src.stat_t_cdf(2.228, 10.0) assert_true(result > 0.94 && result < 0.99) } +///| test "stat_t_quantile_central" { let result = @src.stat_t_quantile(0.5, 10.0) assert_true(result > -0.01 && result < 0.01) } +///| test "stat_t_quantile_basic" { // t_0.975,df=10 ~ 2.228 let result = @src.stat_t_quantile(0.975, 10.0) assert_true(result > 1.5 && result < 3.0) } +///| test "stat_t_test_one_sample" { let sample = [2.1, 2.3, 2.2, 2.4, 2.0] let result = @src.stat_t_test_one_sample(sample, 2.0) @@ -202,6 +235,7 @@ test "stat_t_test_one_sample" { assert_true(result.1 > 0.0 && result.1 <= 1.0) } +///| test "stat_t_test_two_sample" { let x = [2.1, 2.3, 2.2, 2.4, 2.0] let y = [3.1, 3.3, 3.2, 3.4, 3.0] @@ -209,6 +243,7 @@ test "stat_t_test_two_sample" { assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_bh_basic" { let p_values = [0.01, 0.04, 0.03, 0.005, 0.2] let result = @src.stat_bh(p_values) @@ -222,23 +257,27 @@ test "stat_bh_basic" { } } +///| test "stat_bh_empty" { let p_values : Array[Double] = Array::new() let result = @src.stat_bh(p_values) assert_true(result.length() == 0) } +///| test "stat_normal_quantile_half" { let result = @src.stat_normal_quantile(0.5) assert_true(result > -0.01 && result < 0.01) } +///| test "stat_normal_quantile_standard" { // qnorm(0.975) ~ 1.96 let result = @src.stat_normal_quantile(0.975) assert_true(result > 1.9 && result < 2.1) } +///| test "stat_mad_basic" { let data = [1.0, 2.0, 3.0, 4.0, 5.0] let result = @src.stat_mad(data) @@ -246,27 +285,22 @@ test "stat_mad_basic" { assert_true(result > 1.0 && result < 2.0) } +///| test "stat_anova_basic" { - let samples = [ - [1.0, 2.0, 3.0], - [2.0, 3.0, 4.0], - [8.0, 9.0, 10.0], - ] + let samples = [[1.0, 2.0, 3.0], [2.0, 3.0, 4.0], [8.0, 9.0, 10.0]] let result = @src.stat_anova(samples) assert_true(result.0 > 0.0) assert_true(result.1 >= 0.0 && result.1 <= 1.0) } +///| test "stat_anova_equal" { - let samples = [ - [1.0, 1.0, 1.0], - [1.0, 1.0, 1.0], - [1.0, 1.0, 1.0], - ] + let samples = [[1.0, 1.0, 1.0], [1.0, 1.0, 1.0], [1.0, 1.0, 1.0]] let result = @src.stat_anova(samples) assert_true(result.0 >= 0.0) } +///| test "stat_logrank_test_basic" { let time = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0] let event = [true, true, true, false, true, false] diff --git a/test/moonbit/stockholm_test.mbt b/test/moonbit/stockholm_test.mbt index 8f43ba5b..20db96cf 100644 --- a/test/moonbit/stockholm_test.mbt +++ b/test/moonbit/stockholm_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for Stockholm format parser module. - test "stockholm_parse_header" { let content = "# STOCKHOLM 1.0\nseq1/1-10 ACGTACGTAC\n//\n" let ali = @src.stockholm_parse(content) @@ -8,18 +7,21 @@ test "stockholm_parse_header" { assert_eq(ali.blocks.length(), 1) } +///| test "stockholm_parse_empty" { let ali = @src.stockholm_parse("") assert_eq(ali.blocks.length(), 0) assert_eq(ali.gf_annotations.keys().length(), 0) } +///| test "stockholm_parse_header_only" { let content = "# STOCKHOLM 1.0\n//\n" let ali = @src.stockholm_parse(content) assert_eq(ali.blocks.length(), 0) } +///| test "stockholm_parse_single_sequence" { let content = "# STOCKHOLM 1.0\nseq1/1-10 ACGTACGTAC\n//\n" let ali = @src.stockholm_parse(content) @@ -31,6 +33,7 @@ test "stockholm_parse_single_sequence" { assert_eq(ali.blocks[0].sequences[0].end, 10) } +///| test "stockholm_parse_multiple_sequences" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -42,6 +45,7 @@ test "stockholm_parse_multiple_sequences" { assert_eq(ali.blocks[0].sequences.length(), 3) } +///| test "stockholm_parse_gf_annotations" { let content = "# STOCKHOLM 1.0\n" + "#=GF ID TestFamily\n" + @@ -56,6 +60,7 @@ test "stockholm_parse_gf_annotations" { assert_eq(ali.gf_annotations["DE"], "A test protein family") } +///| test "stockholm_parse_gc_ss_cons" { let content = "# STOCKHOLM 1.0\n" + "#=GC SS_cons ........((((...)))).....\n" + @@ -65,6 +70,7 @@ test "stockholm_parse_gc_ss_cons" { assert_eq(ali.blocks[0].secondary_structure, "........((((...)))).....") } +///| test "stockholm_parse_gr_annotation" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -74,6 +80,7 @@ test "stockholm_parse_gr_annotation" { assert_true(ali.blocks[0].gr_annotations.keys().length() > 0) } +///| test "stockholm_parse_gs_annotation" { let content = "# STOCKHOLM 1.0\n" + "#=GS seq1 some annotation value\n" + @@ -83,6 +90,7 @@ test "stockholm_parse_gs_annotation" { assert_true(ali.gs_annotations.keys().length() > 0) } +///| test "stockholm_parse_multiple_blocks" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -97,14 +105,10 @@ test "stockholm_parse_multiple_blocks" { assert_eq(ali.blocks[1].sequences.length(), 2) } +///| test "stockholm_sequence_struct" { let seq = @src.StockholmSequence::new( - "test", - "PF00001", - "Test description", - "ACGTAC", - 1, - 6, + "test", "PF00001", "Test description", "ACGTAC", 1, 6, ) assert_eq(seq.name, "test") assert_eq(seq.accession, "PF00001") @@ -114,6 +118,7 @@ test "stockholm_sequence_struct" { assert_eq(seq.end, 6) } +///| test "stockholm_block_struct" { let seqs : Array[@src.StockholmSequence] = Array::new() let gc : Map[String, String] = Map([], capacity=0) @@ -124,6 +129,7 @@ test "stockholm_block_struct" { assert_eq(block.consensus, "cons") } +///| test "stockholm_alignment_struct" { let blocks : Array[@src.StockholmBlock] = Array::new() let gf : Map[String, String] = Map([], capacity=0) @@ -134,6 +140,7 @@ test "stockholm_alignment_struct" { assert_eq(ali.blocks.length(), 0) } +///| test "stockholm_percent_identity_identical" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -144,6 +151,7 @@ test "stockholm_percent_identity_identical" { assert_true(pid > 0.99) } +///| test "stockholm_percent_identity_different" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -154,6 +162,7 @@ test "stockholm_percent_identity_different" { assert_true(pid < 0.3) } +///| test "stockholm_percent_identity_with_gaps" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -165,6 +174,7 @@ test "stockholm_percent_identity_with_gaps" { assert_true(pid <= 1.0) } +///| test "stockholm_conservation_single_column" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-5 ACGT.\n" + @@ -178,6 +188,7 @@ test "stockholm_conservation_single_column" { assert_true(cons[4] >= 0.0) } +///| test "stockholm_conservation_all_gap_column" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-5 ACGT-\n" + @@ -188,6 +199,7 @@ test "stockholm_conservation_all_gap_column" { assert_eq(cons[4], 0.0) } +///| test "stockholm_to_fasta_basic" { let content = "# STOCKHOLM 1.0\n" + "#=GF AC PF00001\n" + @@ -201,6 +213,7 @@ test "stockholm_to_fasta_basic" { assert_true(fasta.contains("ACGTACGTAC")) } +///| test "stockholm_to_fasta_no_duplicates" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -212,6 +225,7 @@ test "stockholm_to_fasta_no_duplicates" { assert_true(fasta.contains(">seq1")) } +///| test "stockholm_merge_blocks_basic" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + @@ -225,6 +239,7 @@ test "stockholm_merge_blocks_basic" { assert_eq(merged.blocks[0].sequences.length(), 1) } +///| test "stockholm_merge_blocks_sequence_concat" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-5 ACGTG\n" + @@ -237,6 +252,7 @@ test "stockholm_merge_blocks_sequence_concat" { assert_eq(seq.aligned_seq, "ACGTGTACGT") } +///| test "stockholm_write_roundtrip" { let content = "# STOCKHOLM 1.0\n" + "#=GF ID TestFamily\n" + @@ -252,6 +268,7 @@ test "stockholm_write_roundtrip" { assert_true(output.contains("//")) } +///| test "stockholm_write_contains_gf" { let content = "# STOCKHOLM 1.0\n" + "#=GF ID TestFamily\n" + @@ -264,16 +281,16 @@ test "stockholm_write_contains_gf" { assert_true(output.contains("#=GF AC")) } +///| test "stockholm_sequence_without_range" { - let content = "# STOCKHOLM 1.0\n" + - "seq1 ACGTACGTAC\n" + - "//\n" + let content = "# STOCKHOLM 1.0\n" + "seq1 ACGTACGTAC\n" + "//\n" let ali = @src.stockholm_parse(content) assert_eq(ali.blocks[0].sequences.length(), 1) assert_eq(ali.blocks[0].sequences[0].name, "seq1") assert_eq(ali.blocks[0].sequences[0].start, 1) } +///| test "stockholm_gap_handling" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGT--GTAC\n" + @@ -284,6 +301,7 @@ test "stockholm_gap_handling" { assert_eq(ali.blocks[0].sequences[1].aligned_seq, "ACGTACGTAC") } +///| test "stockholm_sample_content" { let content = @src.sample_stockholm_content() let ali = @src.stockholm_parse(content) @@ -292,6 +310,7 @@ test "stockholm_sample_content" { assert_eq(ali.blocks[0].sequences.length(), 3) } +///| test "stockholm_sample_alignment" { let ali = @src.sample_stockholm_alignment() assert_eq(ali.version, "1.0") @@ -299,6 +318,7 @@ test "stockholm_sample_alignment" { assert_true(ali.blocks[0].secondary_structure.length() > 0) } +///| test "stockholm_multiline_gf" { let content = "# STOCKHOLM 1.0\n" + "#=GF DE This is a long description\n" + @@ -309,6 +329,7 @@ test "stockholm_multiline_gf" { assert_true(ali.gf_annotations.keys().length() > 0) } +///| test "stockholm_conservation_with_mixed_chars" { let content = "# STOCKHOLM 1.0\n" + "seq1/1-5 ACGTN\n" + @@ -322,15 +343,15 @@ test "stockholm_conservation_with_mixed_chars" { assert_true(cons[4] >= 0.66) } +///| test "stockholm_percent_identity_single_sequence" { - let content = "# STOCKHOLM 1.0\n" + - "seq1/1-10 ACGTACGTAC\n" + - "//\n" + let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + "//\n" let ali = @src.stockholm_parse(content) let pid = @src.stockholm_percent_identity(ali.blocks[0]) assert_eq(pid, 0.0) } +///| test "stockholm_to_fasta_empty_alignment" { let content = "# STOCKHOLM 1.0\n//\n" let ali = @src.stockholm_parse(content) @@ -338,16 +359,16 @@ test "stockholm_to_fasta_empty_alignment" { assert_eq(fasta.length(), 0) } +///| test "stockholm_merge_single_block" { - let content = "# STOCKHOLM 1.0\n" + - "seq1/1-10 ACGTACGTAC\n" + - "//\n" + let content = "# STOCKHOLM 1.0\n" + "seq1/1-10 ACGTACGTAC\n" + "//\n" let ali = @src.stockholm_parse(content) let merged = @src.stockholm_merge_blocks(ali) assert_eq(merged.blocks.length(), 1) assert_eq(merged.blocks[0].sequences.length(), 1) } +///| test "stockholm_parse_with_comment_lines" { let content = "# STOCKHOLM 1.0\n" + "# Some comment\n" + @@ -358,6 +379,7 @@ test "stockholm_parse_with_comment_lines" { assert_true(ali.markup.length() > 0) } +///| test "stockholm_ss_cons_roundtrip" { let content = "# STOCKHOLM 1.0\n" + "#=GC SS_cons ........((((...)))).....\n" + @@ -369,15 +391,15 @@ test "stockholm_ss_cons_roundtrip" { assert_true(output.contains("........((((...)))).....")) } +///| test "stockholm_parse_range_with_dash" { - let content = "# STOCKHOLM 1.0\n" + - "seq1/1-100 ACGTACGT\n" + - "//\n" + let content = "# STOCKHOLM 1.0\n" + "seq1/1-100 ACGTACGT\n" + "//\n" let ali = @src.stockholm_parse(content) assert_eq(ali.blocks[0].sequences[0].start, 1) assert_eq(ali.blocks[0].sequences[0].end, 100) } +///| test "stockholm_parse_accession_from_gf" { let content = "# STOCKHOLM 1.0\n" + "#=GF AC PF12345.6\n" + @@ -387,4 +409,4 @@ test "stockholm_parse_accession_from_gf" { let ali = @src.stockholm_parse(content) assert_eq(ali.gf_annotations["AC"], "PF12345.6") assert_eq(ali.gf_annotations["DE"], "Test family") -} \ No newline at end of file +} diff --git a/test/moonbit/structural_variant_test.mbt b/test/moonbit/structural_variant_test.mbt index dc926621..8908096e 100644 --- a/test/moonbit/structural_variant_test.mbt +++ b/test/moonbit/structural_variant_test.mbt @@ -78,7 +78,15 @@ test "sv_breakend_is_inter_chromosomal" { ///| test "sv_record_creation" { let rec = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 500.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 500.0, + "PASS", ) assert_eq(rec.id(), "sv1") assert_eq(rec.chrom(), "chr1") @@ -92,11 +100,27 @@ test "sv_record_creation" { ///| test "sv_record_is_bnd" { let bnd = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_bnd(), 1000, 0, None, 500.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_bnd(), + 1000, + 0, + None, + 500.0, + "PASS", ) assert_true(bnd.is_bnd()) let del = @src.SvRecord::new( - "sv2", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 500.0, "PASS", + "sv2", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 500.0, + "PASS", ) assert_false(del.is_bnd()) } @@ -104,7 +128,15 @@ test "sv_record_is_bnd" { ///| test "sv_record_size" { let del = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 0.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 0.0, + "PASS", ) assert_eq(del.size(), 1000) } @@ -113,11 +145,27 @@ test "sv_record_size" { test "sv_record_is_inter_chromosomal" { let be = @src.SvBreakend::new("chr1", 1000, "+", "chr2", 2000, "-") let bnd = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_bnd(), 1000, 0, Some(be), 0.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_bnd(), + 1000, + 0, + Some(be), + 0.0, + "PASS", ) assert_true(bnd.is_inter_chromosomal()) let del = @src.SvRecord::new( - "sv2", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 0.0, "PASS", + "sv2", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 0.0, + "PASS", ) assert_false(del.is_inter_chromosomal()) } @@ -288,7 +336,7 @@ test "sv_find_partners_sample" { let (id_a, id_b) = pairs[0] assert_true( (id_a == "sv_bnd_1" && id_b == "sv_bnd_2") || - (id_a == "sv_bnd_2" && id_b == "sv_bnd_1"), + (id_a == "sv_bnd_2" && id_b == "sv_bnd_1"), ) } @@ -305,10 +353,26 @@ test "sv_are_partners_mutual" { ///| test "sv_are_partners_non_bnd" { let del = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 0.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 0.0, + "PASS", ) let dup = @src.SvRecord::new( - "sv2", "chr1", 3000, @src.sv_type_dup(), 5000, 2000, None, 0.0, "PASS", + "sv2", + "chr1", + 3000, + @src.sv_type_dup(), + 5000, + 2000, + None, + 0.0, + "PASS", ) assert_false(@src.sv_are_partners(del, dup)) } @@ -479,10 +543,26 @@ test "sv_filter_empty_records" { ///| test "sv_find_partners_no_bnd" { let del = @src.SvRecord::new( - "sv1", "chr1", 1000, @src.sv_type_del(), 2000, -1000, None, 0.0, "PASS", + "sv1", + "chr1", + 1000, + @src.sv_type_del(), + 2000, + -1000, + None, + 0.0, + "PASS", ) let dup = @src.SvRecord::new( - "sv2", "chr1", 3000, @src.sv_type_dup(), 5000, 2000, None, 0.0, "PASS", + "sv2", + "chr1", + 3000, + @src.sv_type_dup(), + 5000, + 2000, + None, + 0.0, + "PASS", ) let pairs = @src.sv_find_partners([del, dup]) assert_eq(pairs.length(), 0) diff --git a/test/moonbit/structure_alignment_test.mbt b/test/moonbit/structure_alignment_test.mbt index e5a2fe44..c7e2b76f 100644 --- a/test/moonbit/structure_alignment_test.mbt +++ b/test/moonbit/structure_alignment_test.mbt @@ -9,7 +9,11 @@ test "structure_alignment_point3d" { ///| test "structure_alignment_create_residue" { - let residue = @src.SAResidue::new("ALA", 1, @src.SAPoint3D::new(1.0, 0.0, 0.0)) + let residue = @src.SAResidue::new( + "ALA", + 1, + @src.SAPoint3D::new(1.0, 0.0, 0.0), + ) assert_eq(residue.resname, "ALA") assert_eq(residue.resseq, 1) @@ -126,4 +130,4 @@ test "structure_alignment_pairwise_rmsd" { assert_eq(matrix.length(), 2) assert_eq(matrix[0][0], 0.0) -} \ No newline at end of file +} diff --git a/test/moonbit/substitution_matrices_test.mbt b/test/moonbit/substitution_matrices_test.mbt index ba24c1ff..b403647e 100644 --- a/test/moonbit/substitution_matrices_test.mbt +++ b/test/moonbit/substitution_matrices_test.mbt @@ -13,6 +13,7 @@ test "submat_protein_alphabet_has_20_letters" { assert_eq(alpha[19], "Y") } +///| test "submat_nucleotide_alphabet_has_4_letters" { let alpha = @src.subs_nucleotide_alphabet() assert_eq(alpha.length(), 4) @@ -23,6 +24,7 @@ test "submat_nucleotide_alphabet_has_4_letters" { // ArrayData structure // --------------------------------------------------------------------------- +///| test "submat_array_data_from_dict_basic" { let pairs = [ ("A", "A", 5.0), @@ -34,19 +36,21 @@ test "submat_array_data_from_dict_basic" { assert_eq(ad.rows, ["A", "C"]) assert_eq(ad.cols, ["A", "C"]) assert_true((ad.get(0, 0) - 5.0).abs() < 1.0e-10) - assert_true((ad.get(0, 1) - (-1.0)).abs() < 1.0e-10) + assert_true((ad.get(0, 1) - -1.0).abs() < 1.0e-10) assert_true((ad.get(1, 1) - 9.0).abs() < 1.0e-10) } +///| test "submat_array_data_get_by_label" { let pairs = [("A", "B", 3.0), ("B", "A", 7.0)] let ad = @src.array_data_from_dict(pairs, ["A", "B"], ["A", "B"]) assert_true((ad.get_by_label("A", "B") - 3.0).abs() < 1.0e-10) assert_true((ad.get_by_label("B", "A") - 7.0).abs() < 1.0e-10) // Missing label returns 0.0 - assert_true((ad.get_by_label("X", "Y")).abs() < 1.0e-10) + assert_true(ad.get_by_label("X", "Y").abs() < 1.0e-10) } +///| test "submat_array_data_set_updates_value" { let pairs : Array[(String, String, Double)] = [] let ad = @src.array_data_from_dict(pairs, ["A", "B"], ["A", "B"]) @@ -57,28 +61,31 @@ test "submat_array_data_set_updates_value" { assert_true(true) } +///| test "submat_array_data_get_out_of_bounds_returns_zero" { let pairs : Array[(String, String, Double)] = [] let ad = @src.array_data_from_dict(pairs, ["A"], ["A"]) - assert_true((ad.get(-1, 0)).abs() < 1.0e-10) - assert_true((ad.get(0, 99)).abs() < 1.0e-10) + assert_true(ad.get(-1, 0).abs() < 1.0e-10) + assert_true(ad.get(0, 99).abs() < 1.0e-10) } // --------------------------------------------------------------------------- // Built-in BLOSUM matrices // --------------------------------------------------------------------------- +///| test "submat_blosum62_known_scores" { let m = @src.subs_blosum62_matrix() assert_eq(m.name(), "BLOSUM62") assert_eq(m.n_letters(), 20) assert_true((m.get_score("A", "A") - 4.0).abs() < 1.0e-10) assert_true((m.get_score("R", "R") - 5.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "R") - (-1.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "R") - -1.0).abs() < 1.0e-10) assert_true((m.get_score("W", "W") - 11.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 9.0).abs() < 1.0e-10) } +///| test "submat_blosum62_is_symmetric" { let m = @src.subs_blosum62_matrix() let alpha = m.alphabet() @@ -91,6 +98,7 @@ test "submat_blosum62_is_symmetric" { } } +///| test "submat_blosum45_known_scores" { let m = @src.subs_blosum45_matrix() assert_eq(m.name(), "BLOSUM45") @@ -100,6 +108,7 @@ test "submat_blosum45_known_scores" { assert_true((m.get_score("C", "C") - 12.0).abs() < 1.0e-10) } +///| test "submat_blosum80_known_scores" { let m = @src.subs_blosum80_matrix() assert_eq(m.name(), "BLOSUM80") @@ -108,6 +117,7 @@ test "submat_blosum80_known_scores" { assert_true((m.get_score("W", "W") - 15.0).abs() < 1.0e-10) } +///| test "submat_blosum90_known_scores" { let m = @src.subs_blosum90_matrix() assert_eq(m.name(), "BLOSUM90") @@ -120,6 +130,7 @@ test "submat_blosum90_known_scores" { // Built-in PAM matrices // --------------------------------------------------------------------------- +///| test "submat_pam30_known_scores" { let m = @src.subs_pam30_matrix() assert_eq(m.name(), "PAM30") @@ -128,6 +139,7 @@ test "submat_pam30_known_scores" { assert_true((m.get_score("W", "W") - 14.0).abs() < 1.0e-10) } +///| test "submat_pam70_known_scores" { let m = @src.subs_pam70_matrix() assert_eq(m.name(), "PAM70") @@ -136,6 +148,7 @@ test "submat_pam70_known_scores" { assert_true((m.get_score("W", "W") - 13.0).abs() < 1.0e-10) } +///| test "submat_pam250_known_scores" { let m = @src.subs_pam250_matrix() assert_eq(m.name(), "PAM250") @@ -148,21 +161,23 @@ test "submat_pam250_known_scores" { // Built-in nucleotide matrix // --------------------------------------------------------------------------- +///| test "submat_nuc44_match_mismatch" { let m = @src.subs_nuc44_matrix() assert_eq(m.name(), "NUC4.4") assert_eq(m.n_letters(), 4) assert_true((m.get_score("A", "A") - 5.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "C") - (-4.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -4.0).abs() < 1.0e-10) assert_true((m.get_score("G", "G") - 5.0).abs() < 1.0e-10) assert_true((m.get_score("T", "T") - 5.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "T") - (-4.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "T") - -4.0).abs() < 1.0e-10) } // --------------------------------------------------------------------------- // SubsMatrix construction and methods // --------------------------------------------------------------------------- +///| test "submat_subs_matrix_from_pairs" { let alpha = ["A", "C"] let pairs = [ @@ -176,17 +191,19 @@ test "submat_subs_matrix_from_pairs" { assert_eq(m.n_letters(), 2) assert_true((m.get_score("A", "A") - 5.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 9.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "C") - (-2.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -2.0).abs() < 1.0e-10) } +///| test "submat_get_score_idx" { let m = @src.subs_blosum62_matrix() // A=0, R=14 in alphabet ACDEFGHIKLMNPQRSTVWY assert_true((m.get_score_idx(0, 0) - 4.0).abs() < 1.0e-10) // A-R = -1 - assert_true((m.get_score_idx(0, 14) - (-1.0)).abs() < 1.0e-10) + assert_true((m.get_score_idx(0, 14) - -1.0).abs() < 1.0e-10) } +///| test "submat_select_submatrix" { let m = @src.subs_blosum62_matrix() let sub = m.select(["A", "C", "D"]) @@ -198,6 +215,7 @@ test "submat_select_submatrix" { assert_true((sub.get_score("A", "C") - 0.0).abs() < 1.0e-10) } +///| test "submat_to_table_string_has_name" { let m = @src.subs_blosum62_matrix() let s = m.to_table_string() @@ -211,6 +229,7 @@ test "submat_to_table_string_has_name" { // Matrix registry // --------------------------------------------------------------------------- +///| test "submat_registry_initialize_and_list" { @src.subs_initialize_registry() let names = @src.list_matrices() @@ -220,15 +239,22 @@ test "submat_registry_initialize_and_list" { let mut has_pam250 = false let mut has_nuc44 = false for n in names { - if n == "BLOSUM62" { has_blosum62 = true } - if n == "PAM250" { has_pam250 = true } - if n == "NUC4.4" { has_nuc44 = true } + if n == "BLOSUM62" { + has_blosum62 = true + } + if n == "PAM250" { + has_pam250 = true + } + if n == "NUC4.4" { + has_nuc44 = true + } } assert_true(has_blosum62) assert_true(has_pam250) assert_true(has_nuc44) } +///| test "submat_registry_load_existing_matrix" { @src.subs_initialize_registry() let opt = @src.load_matrix("BLOSUM62") @@ -238,12 +264,14 @@ test "submat_registry_load_existing_matrix" { assert_true((m.get_score("A", "A") - 4.0).abs() < 1.0e-10) } +///| test "submat_registry_load_missing_returns_none" { @src.subs_initialize_registry() let opt = @src.load_matrix("NONEXISTENT_MATRIX_XYZ") assert_true(opt.is_none()) } +///| test "submat_registry_register_custom_matrix" { let alpha = ["A", "C"] let pairs = [ @@ -260,6 +288,7 @@ test "submat_registry_register_custom_matrix" { assert_true((loaded.get_score("A", "A") - 1.0).abs() < 1.0e-10) } +///| test "submat_registry_register_overwrites_existing" { let alpha = ["A"] let pairs1 = [("A", "A", 5.0)] @@ -276,6 +305,7 @@ test "submat_registry_register_overwrites_existing" { // Frequency matrix calculation // --------------------------------------------------------------------------- +///| test "submat_calculate_frequency_matrix_identical_sequences" { let alignment = ["ACGT", "ACGT"] let alpha = @src.subs_nucleotide_alphabet() @@ -288,9 +318,10 @@ test "submat_calculate_frequency_matrix_identical_sequences" { assert_true((freq.get(2, 2) - 2.0).abs() < 1.0e-10) // G assert_true((freq.get(3, 3) - 2.0).abs() < 1.0e-10) // T // Off-diagonal should be 0 - assert_true((freq.get(0, 1)).abs() < 1.0e-10) + assert_true(freq.get(0, 1).abs() < 1.0e-10) } +///| test "submat_calculate_frequency_matrix_with_substitutions" { let alignment = ["ACGT", "AGGT"] // position 1: C->G let alpha = @src.subs_nucleotide_alphabet() @@ -306,6 +337,7 @@ test "submat_calculate_frequency_matrix_with_substitutions" { assert_true((freq.get(3, 3) - 2.0).abs() < 1.0e-10) } +///| test "submat_calculate_frequency_matrix_skips_gaps" { let alignment = ["A-A", "ACA"] // position 1: gap in seq1 let alpha = @src.subs_nucleotide_alphabet() @@ -315,16 +347,17 @@ test "submat_calculate_frequency_matrix_skips_gaps" { // Position 2: A-A -> freq[0][0] += 2 assert_true((freq.get(0, 0) - 4.0).abs() < 1.0e-10) // No C-C pairs - assert_true((freq.get(1, 1)).abs() < 1.0e-10) + assert_true(freq.get(1, 1).abs() < 1.0e-10) } +///| test "submat_calculate_frequency_matrix_single_sequence_returns_zeros" { let alignment = ["ACGT"] let alpha = @src.subs_nucleotide_alphabet() let freq = @src.calculate_frequency_matrix(alignment, alpha) for i in 0..<4 { for j in 0..<4 { - assert_true((freq.get(i, j)).abs() < 1.0e-10) + assert_true(freq.get(i, j).abs() < 1.0e-10) } } } @@ -333,6 +366,7 @@ test "submat_calculate_frequency_matrix_single_sequence_returns_zeros" { // Substitution matrix calculation (log-odds) // --------------------------------------------------------------------------- +///| test "submat_calculate_substitution_matrix_identical_seqs" { let alignment = ["AC", "AC"] let alpha = ["A", "C"] @@ -342,9 +376,10 @@ test "submat_calculate_substitution_matrix_identical_seqs" { assert_true((m.get_score("A", "A") - 2.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 2.0).abs() < 1.0e-10) // Off-diagonal: observed=0 -> -999 - assert_true((m.get_score("A", "C") - (-999.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -999.0).abs() < 1.0e-10) } +///| test "submat_calculate_substitution_matrix_default_scale" { let alignment = ["AA", "AA"] let alpha = ["A"] @@ -359,17 +394,21 @@ test "submat_calculate_substitution_matrix_default_scale" { // Shannon entropy // --------------------------------------------------------------------------- +///| test "submat_shannon_entropy_uniform_distribution" { // 2x2 matrix with all equal values: H = log2(4) = 2 let pairs = [ - ("A", "A", 1.0), ("A", "C", 1.0), - ("C", "A", 1.0), ("C", "C", 1.0), + ("A", "A", 1.0), + ("A", "C", 1.0), + ("C", "A", 1.0), + ("C", "C", 1.0), ] let freq = @src.array_data_from_dict(pairs, ["A", "C"], ["A", "C"]) let h = @src.subs_shannon_entropy(freq) assert_true((h - 2.0).abs() < 1.0e-10) } +///| test "submat_shannon_entropy_certain_distribution" { // All mass on one cell: H = 0 let pairs = [("A", "A", 4.0)] @@ -378,6 +417,7 @@ test "submat_shannon_entropy_certain_distribution" { assert_true(h.abs() < 1.0e-10) } +///| test "submat_shannon_entropy_empty_matrix" { let pairs : Array[(String, String, Double)] = [] let freq = @src.array_data_from_dict(pairs, ["A"], ["A"]) @@ -385,6 +425,7 @@ test "submat_shannon_entropy_empty_matrix" { assert_true(h.abs() < 1.0e-10) } +///| test "submat_shannon_entropy_two_equal_cells" { // Two non-zero cells with equal weight: H = 1 let pairs = [("A", "A", 1.0), ("C", "C", 1.0)] @@ -397,17 +438,21 @@ test "submat_shannon_entropy_two_equal_cells" { // Relative entropy (KL divergence) // --------------------------------------------------------------------------- +///| test "submat_relative_entropy_uniform_is_zero" { // When all letters equally frequent, observed == expected, KL = 0 let pairs = [ - ("A", "A", 1.0), ("A", "C", 1.0), - ("C", "A", 1.0), ("C", "C", 1.0), + ("A", "A", 1.0), + ("A", "C", 1.0), + ("C", "A", 1.0), + ("C", "C", 1.0), ] let freq = @src.array_data_from_dict(pairs, ["A", "C"], ["A", "C"]) let kl = @src.subs_relative_entropy(freq) assert_true(kl.abs() < 1.0e-10) } +///| test "submat_relative_entropy_identical_pairs_zero" { // Only A-A pairs: q(A,A)=1, p(A)=1, exp=1, KL = 1 * log2(1/1) = 0 let pairs = [("A", "A", 4.0)] @@ -416,6 +461,7 @@ test "submat_relative_entropy_identical_pairs_zero" { assert_true(kl.abs() < 1.0e-10) } +///| test "submat_relative_entropy_empty_matrix" { let pairs : Array[(String, String, Double)] = [] let freq = @src.array_data_from_dict(pairs, ["A"], ["A"]) @@ -427,6 +473,7 @@ test "submat_relative_entropy_empty_matrix" { // NCBI matrix parsing // --------------------------------------------------------------------------- +///| test "submat_parse_ncbi_matrix_simple" { let content = #|# Test matrix @@ -437,10 +484,11 @@ test "submat_parse_ncbi_matrix_simple" { assert_true(opt.is_some()) let m = opt.unwrap() assert_true((m.get_score("A", "A") - 1.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "C") - (-1.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -1.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 1.0).abs() < 1.0e-10) } +///| test "submat_parse_ncbi_matrix_with_comments" { let content = #|# Generated by NCBI @@ -455,16 +503,18 @@ test "submat_parse_ncbi_matrix_with_comments" { let m = opt.unwrap() assert_true((m.get_score("A", "A") - 5.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 5.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "C") - (-4.0)).abs() < 1.0e-10) - assert_true((m.get_score("G", "T") - (-4.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -4.0).abs() < 1.0e-10) + assert_true((m.get_score("G", "T") - -4.0).abs() < 1.0e-10) } +///| test "submat_parse_ncbi_matrix_empty_returns_none" { let content = "# only comments\n# nothing else\n" let opt = @src.parse_ncbi_matrix(content) assert_true(opt.is_none()) } +///| test "submat_parse_ncbi_matrix_handles_extra_whitespace" { let content = "A C\nA 2 -3\nC -3 2\n" let opt = @src.parse_ncbi_matrix(content) @@ -472,13 +522,14 @@ test "submat_parse_ncbi_matrix_handles_extra_whitespace" { let m = opt.unwrap() assert_true((m.get_score("A", "A") - 2.0).abs() < 1.0e-10) assert_true((m.get_score("C", "C") - 2.0).abs() < 1.0e-10) - assert_true((m.get_score("A", "C") - (-3.0)).abs() < 1.0e-10) + assert_true((m.get_score("A", "C") - -3.0).abs() < 1.0e-10) } // --------------------------------------------------------------------------- // Matrix correlation // --------------------------------------------------------------------------- +///| test "submat_matrix_correlation_identical_matrices" { let m1 = @src.subs_blosum62_matrix() let m2 = @src.subs_blosum62_matrix() @@ -487,6 +538,7 @@ test "submat_matrix_correlation_identical_matrices" { assert_true((corr - 1.0).abs() < 1.0e-9) } +///| test "submat_matrix_correlation_blosum45_vs_blosum62" { let m45 = @src.subs_blosum45_matrix() let m62 = @src.subs_blosum62_matrix() @@ -496,6 +548,7 @@ test "submat_matrix_correlation_blosum45_vs_blosum62" { assert_true(corr <= 1.0) } +///| test "submat_matrix_correlation_disjoint_alphabets" { let alpha1 = ["A"] let alpha2 = ["C"] @@ -510,6 +563,7 @@ test "submat_matrix_correlation_disjoint_alphabets" { // End-to-end: derive matrix from alignment, then score // --------------------------------------------------------------------------- +///| test "submat_end_to_end_derive_matrix_and_score" { // Build a small alignment where A and C never co-occur at the same position let alignment = ["AAA", "AAA", "CCC", "CCC"] diff --git a/test/moonbit/summarized_experiment_test.mbt b/test/moonbit/summarized_experiment_test.mbt index 5bc41bc8..e5d60f1b 100644 --- a/test/moonbit/summarized_experiment_test.mbt +++ b/test/moonbit/summarized_experiment_test.mbt @@ -53,6 +53,10 @@ test "summarized_experiment_subset_rows" { let se_sub = @src.se_subset_rows(se, [0, 2]) assert_eq(@src.se_nrow(se_sub), 2) + match @src.se_assay(se_sub, "counts") { + Some(assay) => assert_eq(assay, [[1.0, 2.0], [5.0, 6.0]]) + None => assert_true(false) + } } ///| @@ -70,6 +74,10 @@ test "summarized_experiment_subset_cols" { let se_sub = @src.se_subset_cols(se, [0, 2]) assert_eq(@src.se_ncol(se_sub), 2) + match @src.se_assay(se_sub, "counts") { + Some(assay) => assert_eq(assay, [[1.0, 3.0], [4.0, 6.0]]) + None => assert_true(false) + } } ///| diff --git a/test/moonbit/survival_test.mbt b/test/moonbit/survival_test.mbt index 1a692b81..163ce505 100644 --- a/test/moonbit/survival_test.mbt +++ b/test/moonbit/survival_test.mbt @@ -410,7 +410,9 @@ test "survival_normal_cdf" { ///| test "survival_normal_p_value_two_sided" { // z=1.96 -> p ≈ 0.05 - assert_true((@src.survival_normal_p_value_two_sided(1.96) - 0.05).abs() < 0.01) + assert_true( + (@src.survival_normal_p_value_two_sided(1.96) - 0.05).abs() < 0.01, + ) // z=0 -> p = 1.0 assert_true((@src.survival_normal_p_value_two_sided(0.0) - 1.0).abs() < 0.001) // p-value should be in [0, 1] @@ -432,7 +434,10 @@ test "survival_chi_square_p_value" { assert_true(@src.survival_chi_square_p_value(10.0, 5) >= 0.0) assert_true(@src.survival_chi_square_p_value(10.0, 5) <= 1.0) // Larger chi_sq gives smaller p-value - assert_true(@src.survival_chi_square_p_value(10.0, 1) < @src.survival_chi_square_p_value(1.0, 1)) + assert_true( + @src.survival_chi_square_p_value(10.0, 1) < + @src.survival_chi_square_p_value(1.0, 1), + ) } // =========================================================================== diff --git a/test/moonbit/system_piper_test.mbt b/test/moonbit/system_piper_test.mbt index 87e264b7..1cc2b8d2 100644 --- a/test/moonbit/system_piper_test.mbt +++ b/test/moonbit/system_piper_test.mbt @@ -1,6 +1,5 @@ ///| /// Test file for SystemPipeR module. - test "pipeline_new" { let pipe = @src.Pipeline::new("test_v1", "Test Pipeline") assert_eq(pipe.get_id(), "test_v1") @@ -8,6 +7,7 @@ test "pipeline_new" { assert_eq(pipe.get_n_steps(), 0) } +///| test "pipeline_set_description" { let pipe = @src.Pipeline::new("test", "Test") let pipe2 = pipe.set_description("A test pipeline") @@ -16,17 +16,20 @@ test "pipeline_set_description" { assert_true(s.contains("A test pipeline")) } +///| test "pipeline_set_param" { let pipe = @src.Pipeline::new("test", "Test") let pipe2 = pipe.set_param("key1", "value1") assert_eq(pipe2.get_param("key1"), "value1") } +///| test "pipeline_get_param_missing" { let pipe = @src.Pipeline::new("test", "Test") assert_eq(pipe.get_param("nonexistent"), "") } +///| test "pipeline_add_step" { let pipe = @src.Pipeline::new("test", "Test") let step = @src.PipelineStep::new("step1", "Step 1", "echo") @@ -34,6 +37,7 @@ test "pipeline_add_step" { assert_eq(pipe2.get_n_steps(), 1) } +///| test "pipeline_get_step" { let pipe = @src.Pipeline::new("test", "Test") let step = @src.PipelineStep::new("step1", "Step 1", "echo") @@ -43,12 +47,14 @@ test "pipeline_get_step" { assert_eq(s.get_name(), "Step 1") } +///| test "pipeline_get_step_not_found" { let pipe = @src.Pipeline::new("test", "Test") let s = pipe.get_step("missing") assert_eq(s.get_id(), "") } +///| test "pipeline_get_steps" { let pipe = @src.Pipeline::new("test", "Test") let pipe2 = pipe.add_step(@src.PipelineStep::new("s1", "S1", "cmd1")) @@ -57,26 +63,31 @@ test "pipeline_get_steps" { assert_eq(steps.length(), 2) } +///| test "pipeline_completed_count" { let pipe = @src.Pipeline::new("test", "Test") assert_eq(pipe.get_completed_count(), 0) } +///| test "pipeline_failed_count" { let pipe = @src.Pipeline::new("test", "Test") assert_eq(pipe.get_failed_count(), 0) } +///| test "pipeline_pending_count" { let pipe = @src.Pipeline::new("test", "Test") assert_eq(pipe.get_pending_count(), 0) } +///| test "pipeline_progress_empty" { let pipe = @src.Pipeline::new("test", "Test") assert_eq(pipe.get_progress(), 0.0) } +///| test "step_new" { let step = @src.PipelineStep::new("s1", "My Step", "bwa") assert_eq(step.get_id(), "s1") @@ -84,18 +95,21 @@ test "step_new" { assert_eq(step.get_command(), "bwa") } +///| test "step_set_description" { let step = @src.PipelineStep::new("s1", "Step", "cmd") let step2 = step.set_description("A step for testing") assert_eq(step2.description, "A step for testing") } +///| test "step_set_args" { let step = @src.PipelineStep::new("s1", "Step", "cmd") let step2 = step.set_args(["arg1", "arg2", "-v"]) assert_eq(step2.args.length(), 3) } +///| test "step_add_dependency" { let step = @src.PipelineStep::new("s2", "Step 2", "cmd") let step2 = step.add_dependency("s1") @@ -104,28 +118,33 @@ test "step_add_dependency" { assert_eq(deps[0], "s1") } +///| test "step_add_input" { let step = @src.PipelineStep::new("s1", "Step", "cmd") let step2 = step.add_input("input.fq") assert_eq(step2.input_files.length(), 1) } +///| test "step_add_output" { let step = @src.PipelineStep::new("s1", "Step", "cmd") let step2 = step.add_output("output.bam") assert_eq(step2.output_files.length(), 1) } +///| test "step_status" { let step = @src.PipelineStep::new("s1", "Step", "cmd") assert_true(step.get_status() is @src.StepStatus::Pending) } +///| test "step_duration" { let step = @src.PipelineStep::new("s1", "Step", "cmd") assert_eq(step.get_duration(), 0.0) } +///| test "step_status_to_string" { // Test default status via PipelineStep (which has Pending by default) let step = @src.PipelineStep::new("s1", "Step", "cmd") @@ -133,6 +152,7 @@ test "step_status_to_string" { assert_eq(status_str, "pending") } +///| test "pipeline_config_new" { let config = @src.PipelineConfig::new("/work", "/input", "/output") assert_eq(config.get_work_dir(), "/work") @@ -141,12 +161,14 @@ test "pipeline_config_new" { assert_eq(config.get_cores(), 1) } +///| test "pipeline_config_set_cores" { let config = @src.PipelineConfig::new("/w", "/i", "/o") let config2 = config.set_cores(8) assert_eq(config2.get_cores(), 8) } +///| test "pipeline_summary" { let pipe = @src.Pipeline::new("test", "My Pipeline") let s = pipe.summary() @@ -154,6 +176,7 @@ test "pipeline_summary" { assert_true(s.contains("My Pipeline")) } +///| test "pipeline_to_ascii" { let pipe = @src.Pipeline::new("test", "My Pipeline") let ascii = pipe.to_ascii() @@ -161,6 +184,7 @@ test "pipeline_to_ascii" { assert_true(ascii.contains("Progress")) } +///| test "pipeline_sample" { let pipe = @src.pipeline_sample() assert_true(pipe.get_n_steps() >= 5) @@ -168,6 +192,7 @@ test "pipeline_sample" { assert_true(pipe.get_progress() >= 0.0) } +///| test "pipeline_sample_steps" { let pipe = @src.pipeline_sample() let steps = pipe.get_steps() @@ -176,14 +201,16 @@ test "pipeline_sample_steps" { assert_eq(first.get_name(), "Quality Control") } +///| test "pipeline_global_params" { let pipe = @src.pipeline_sample() assert_eq(pipe.get_param("reference_genome"), "GRCh38.p14") assert_eq(pipe.get_param("species"), "Homo sapiens") } +///| test "step_can_run_no_deps" { let pipe = @src.Pipeline::new("test", "Test") let step = @src.PipelineStep::new("s1", "Step 1", "cmd") assert_true(pipe.can_run_step(step)) -} \ No newline at end of file +} diff --git a/test/moonbit/taxonomy_test.mbt b/test/moonbit/taxonomy_test.mbt index 1e28d482..0d589edc 100644 --- a/test/moonbit/taxonomy_test.mbt +++ b/test/moonbit/taxonomy_test.mbt @@ -1,11 +1,11 @@ ///| /// Taxonomy module tests - test "TaxonomyDatabase::new" { let db = @src.TaxonomyDatabase::new() assert_eq(db.count_taxa(), 0) } +///| test "TaxonomyDatabase::add_taxon" { let db = @src.TaxonomyDatabase::new() let taxon = @src.Taxon::new("1", "0", "no rank").set_scientific_name("root") @@ -13,6 +13,7 @@ test "TaxonomyDatabase::add_taxon" { assert_eq(db.count_taxa(), 1) } +///| test "TaxonomyDatabase::get_taxon" { let db = @src.TaxonomyDatabase::new() let taxon = @src.Taxon::new("1", "0", "no rank").set_scientific_name("root") @@ -23,6 +24,7 @@ test "TaxonomyDatabase::get_taxon" { } } +///| test "TaxonomyDatabase::get_taxid_by_name" { let db = @src.create_example_taxonomy() match db.get_taxid_by_name("Homo sapiens") { @@ -31,6 +33,7 @@ test "TaxonomyDatabase::get_taxid_by_name" { } } +///| test "TaxonomyDatabase::get_lineage" { let db = @src.create_example_taxonomy() let lineage = db.get_lineage("9606") @@ -38,12 +41,14 @@ test "TaxonomyDatabase::get_lineage" { assert_eq(lineage[lineage.length() - 1], "Homo sapiens") } +///| test "TaxonomyDatabase::get_ancestors" { let db = @src.create_example_taxonomy() let ancestors = db.get_ancestors("9606") assert_true(ancestors.length() > 0) } +///| test "TaxonomyDatabase::get_common_ancestor" { let db = @src.create_example_taxonomy() match db.get_common_ancestor("9606", "9599") { @@ -52,30 +57,35 @@ test "TaxonomyDatabase::get_common_ancestor" { } } +///| test "TaxonomyDatabase::is_ancestor" { let db = @src.create_example_taxonomy() assert_true(db.is_ancestor("9604", "9606")) assert_true(!db.is_ancestor("9606", "9604")) } +///| test "TaxonomyDatabase::get_distance" { let db = @src.create_example_taxonomy() let dist = db.get_distance("9606", "9599") assert_true(dist > 0) } +///| test "TaxonomyDatabase::get_taxa_by_rank" { let db = @src.create_example_taxonomy() let species = db.get_taxa_by_rank("species") assert_true(species.length() >= 2) } +///| test "parse_nodes_dmp" { let content = "1\t|\t0\t|\tno rank\t|\n131567\t|\t1\t|\tno rank\t|\n2759\t|\t131567\t|\tsuperkingdom\t|\n" let db = @src.parse_nodes_dmp(@src.TaxonomyDatabase::new(), content) assert_eq(db.count_taxa(), 3) } +///| test "parse_names_dmp" { let db = @src.TaxonomyDatabase::new() ignore(db.add_taxon(@src.Taxon::new("1", "0", "no rank"))) @@ -87,23 +97,27 @@ test "parse_names_dmp" { } } +///| test "create_example_taxonomy" { let db = @src.create_example_taxonomy() assert_true(db.count_taxa() > 10) } +///| test "get_all_species" { let db = @src.create_example_taxonomy() let species = db.get_all_species() assert_true(species.length() >= 2) } +///| test "get_common_names" { let db = @src.create_example_taxonomy() let names = db.get_common_names("9606") assert_true(contains_string(names, "human")) } +///| fn contains_string(arr : Array[String], s : String) -> Bool { let mut i = 0 while i < arr.length() { @@ -113,4 +127,4 @@ fn contains_string(arr : Array[String], s : String) -> Bool { i = i + 1 } false -} \ No newline at end of file +} diff --git a/test/moonbit/topgo_test.mbt b/test/moonbit/topgo_test.mbt index 0d345f33..62fefc8e 100644 --- a/test/moonbit/topgo_test.mbt +++ b/test/moonbit/topgo_test.mbt @@ -1,6 +1,5 @@ ///| /// Tests for topGO module. - test "TopGOTerm creation" { let term = @src.TopGOTerm::new("GO:0008150", "biological_process", "BP", 0) assert_eq(term.go_id, "GO:0008150") @@ -10,6 +9,7 @@ test "TopGOTerm creation" { assert_eq(term.gene_count, 0) } +///| test "TopGOGraph construction" { let mut graph = @src.TopGOGraph::new("GO:0008150") let root = @src.TopGOTerm::new("GO:0008150", "biological_process", "BP", 0) @@ -17,6 +17,7 @@ test "TopGOGraph construction" { assert_true(graph.terms.contains("GO:0008150")) } +///| test "TopGOGraph add edge" { let mut graph = @src.TopGOGraph::new("GO:0008150") let root = @src.TopGOTerm::new("GO:0008150", "biological_process", "BP", 0) @@ -28,22 +29,35 @@ test "TopGOGraph add edge" { assert_true(graph.terms.contains("GO:0009987")) } +///| test "topgo_elim_algorithm" { let graph = @src.create_example_topgo_graph() let genes_of_interest = ["gene1", "gene2", "gene3"] - let all_genes = ["gene1", "gene2", "gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9", "gene10"] - let results = @src.topgo_elim_algorithm(graph, genes_of_interest, all_genes, "BP") + let all_genes = [ + "gene1", "gene2", "gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9", + "gene10", + ] + let results = @src.topgo_elim_algorithm( + graph, genes_of_interest, all_genes, "BP", + ) assert_true(results.length() > 0) } +///| test "topgo_weight01_algorithm" { let graph = @src.create_example_topgo_graph() let genes_of_interest = ["gene1", "gene2", "gene3"] - let all_genes = ["gene1", "gene2", "gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9", "gene10"] - let results = @src.topgo_weight01_algorithm(graph, genes_of_interest, all_genes, "BP") + let all_genes = [ + "gene1", "gene2", "gene3", "gene4", "gene5", "gene6", "gene7", "gene8", "gene9", + "gene10", + ] + let results = @src.topgo_weight01_algorithm( + graph, genes_of_interest, all_genes, "BP", + ) assert_true(results.length() > 0) } +///| test "topgo_fisher_exact" { let p = @src.topgo_fisher_exact(3, 10, 5, 100) assert_true(p >= 0.0 && p <= 1.0) diff --git a/test/moonbit/tradeseq_advanced_test.mbt b/test/moonbit/tradeseq_advanced_test.mbt new file mode 100644 index 00000000..dee249ae --- /dev/null +++ b/test/moonbit/tradeseq_advanced_test.mbt @@ -0,0 +1,993 @@ +///| +fn tradeseq_adv_test_counts() -> Array[Array[Double]] { + let flat : Array[Double] = [] + let association : Array[Double] = [] + let endpoint : Array[Double] = [] + let early_branch : Array[Double] = [] + let zero : Array[Double] = [] + for cell in 0..<24 { + flat.push((10 + cell % 3).to_double()) + if cell < 8 { + association.push((3 + cell).to_double()) + endpoint.push((6 + cell % 2).to_double()) + early_branch.push((8 + cell % 2).to_double()) + } else if cell < 16 { + let step = cell - 8 + association.push((12 + 3 * step).to_double()) + endpoint.push((10 + 4 * step).to_double()) + early_branch.push( + if step < 4 { + (26 + step).to_double() + } else { + (14 + step % 2).to_double() + }, + ) + } else { + let step = cell - 16 + association.push((12 + 3 * step).to_double()) + endpoint.push((10 + step / 3).to_double()) + early_branch.push( + if step < 4 { + (4 + step).to_double() + } else { + (14 + step % 2).to_double() + }, + ) + } + zero.push(0.0) + } + [flat, association, endpoint, early_branch, zero] +} + +///| +fn tradeseq_adv_test_trajectory() -> ( + Array[Array[Double]], + Array[Array[Double]], +) { + let pseudotime : Array[Array[Double]] = [] + let weights : Array[Array[Double]] = [] + for cell in 0..<24 { + if cell < 8 { + let time = cell.to_double() * 0.05 + pseudotime.push([time, time]) + weights.push([2.0, 2.0]) + } else if cell < 16 { + let time = 0.4 + (cell - 8).to_double() * 0.08 + pseudotime.push([time, 0.0]) + weights.push([3.0, 0.0]) + } else { + let time = 0.4 + (cell - 16).to_double() * 0.08 + pseudotime.push([0.0, time]) + weights.push([0.0, 4.0]) + } + } + (pseudotime, weights) +} + +///| +fn tradeseq_adv_test_config() -> @src.TradeSeqAdvancedConfig { + @src.TradeSeqAdvancedConfig::create( + n_knots=4, + smoothing_penalty=1.0, + max_iterations=80, + tolerance=1.0e-5, + minimum_dispersion=1.0e-4, + maximum_dispersion=20.0, + ridge=1.0e-6, + fdr_threshold=0.1, + test_points=6, + ) catch { + _ => abort("valid tradeSeq advanced configuration should build") + } +} + +///| +fn tradeseq_adv_test_fit() -> @src.TradeSeqAdvancedFit { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + @src.tradeseq_fit_advanced( + tradeseq_adv_test_counts(), + pseudotime, + weights, + gene_names=["flat", "association", "endpoint", "early", "zero"], + lineage_names=["left", "right"], + offsets=Array::make(24, 0.0), + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid tradeSeq advanced data should fit") + } +} + +///| +fn tradeseq_adv_test_single_lineage_fit() -> @src.TradeSeqAdvancedFit { + let pseudotime : Array[Array[Double]] = [] + let weights : Array[Array[Double]] = [] + let counts : Array[Double] = [] + for cell in 0..<12 { + pseudotime.push([cell.to_double()]) + weights.push([1.0]) + counts.push((cell + 1).to_double()) + } + @src.tradeseq_fit_advanced( + [counts], + pseudotime, + weights, + gene_names=["gene"], + offsets=Array::make(12, 0.0), + config=@src.TradeSeqAdvancedConfig::create( + n_knots=3, + max_iterations=40, + ridge=1.0e-5, + test_points=4, + ), + ) catch { + _ => abort("valid single-lineage tradeSeq data should fit") + } +} + +///| +fn tradeseq_adv_test_sce() -> @src.SingleCellExperiment { + let names : Array[String] = [] + for cell in 0..<24 { + names.push("cell" + (cell + 1).to_string()) + } + let experiment = @src.SingleCellExperiment::new( + tradeseq_adv_test_counts(), + ["flat", "association", "endpoint", "early", "zero"], + names, + ) + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + experiment.reduced_dims["slingshot.pseudotime"] = pseudotime + experiment.reduced_dims["slingshot.weights"] = weights + experiment.row_data["symbol"] = ["F", "A", "E", "B", "Z"] + experiment.col_data["batch"] = Array::make(24, "one") + experiment.metadata["source"] = "test" + experiment.alternative_experiments["spike"] = @src.SingleCellExperiment::new( + [[1.0, 2.0]], + ["spike1"], + ["s1", "s2"], + ) + experiment +} + +///| +fn tradeseq_adv_test_finite(value : Double) -> Bool { + value == value && value.abs() < 1.0e299 +} + +///| +fn tradeseq_adv_test_slingshot_result() -> @src.SlingshotAdvancedResult { + let coordinates = [ + [-0.1, 0.0], + [0.0, 0.1], + [0.1, -0.1], + [0.9, 0.0], + [1.0, 0.1], + [1.1, -0.1], + [1.9, 0.0], + [2.0, 0.1], + [2.1, -0.1], + [2.9, 0.9], + [3.0, 1.0], + [3.1, 1.1], + [2.9, -0.9], + [3.0, -1.0], + [3.1, -1.1], + ] + let labels = [ + "A", "A", "A", "B", "B", "B", "C", "C", "C", "D", "D", "D", "E", "E", "E", + ] + let config = @src.SlingshotAdvancedConfig::create( + start_clusters=["A"], + end_clusters=["D", "E"], + distance=@src.slingshot_center_euclidean(), + extension=@src.slingshot_extend_none(), + max_iterations=6, + tolerance=1.0e-4, + smoother_span=0.4, + curve_points=24, + ) catch { + _ => abort("valid Slingshot configuration should build") + } + @src.slingshot_advanced(coordinates, labels, config) catch { + _ => abort("valid Slingshot data should fit") + } +} + +///| +test "tradeSeq advanced default configuration matches fitGAM workflow" { + let config = @src.TradeSeqAdvancedConfig::default() + assert_eq(config.n_knots, 6) + assert_eq(config.smoothing_penalty, 1.0) + assert_eq(config.max_iterations, 60) + assert_eq(config.tolerance, 1.0e-6) + assert_eq(config.minimum_dispersion, 1.0e-4) + assert_eq(config.maximum_dispersion, 100.0) + assert_eq(config.ridge, 1.0e-8) + assert_eq(config.fdr_threshold, 0.05) + assert_eq(config.test_points, 12) +} + +///| +test "tradeSeq advanced configuration preserves explicit controls" { + let config = tradeseq_adv_test_config() + assert_eq(config.n_knots, 4) + assert_eq(config.smoothing_penalty, 1.0) + assert_eq(config.max_iterations, 80) + assert_eq(config.tolerance, 1.0e-5) + assert_eq(config.maximum_dispersion, 20.0) + assert_eq(config.ridge, 1.0e-6) + assert_eq(config.fdr_threshold, 0.1) + assert_eq(config.test_points, 6) +} + +///| +test "tradeSeq advanced rejects invalid knots and smoothing penalty" { + let mut failures = 0 + ignore(@src.TradeSeqAdvancedConfig::create(n_knots=2)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(smoothing_penalty=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(smoothing_penalty=0.0 / 0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced rejects invalid iteration controls" { + let mut failures = 0 + ignore(@src.TradeSeqAdvancedConfig::create(max_iterations=0)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(tolerance=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(tolerance=0.0 / 0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced rejects invalid dispersion and ridge controls" { + let mut failures = 0 + ignore(@src.TradeSeqAdvancedConfig::create(minimum_dispersion=0.0)) catch { + _ => failures = failures + 1 + } + ignore( + @src.TradeSeqAdvancedConfig::create( + minimum_dispersion=2.0, + maximum_dispersion=1.0, + ), + ) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(ridge=0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced rejects invalid testing controls" { + let mut failures = 0 + ignore(@src.TradeSeqAdvancedConfig::create(fdr_threshold=0.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(fdr_threshold=1.1)) catch { + _ => failures = failures + 1 + } + ignore(@src.TradeSeqAdvancedConfig::create(test_points=1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced validates count matrix shape and values" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let mut failures = 0 + ignore(@src.tradeseq_fit_advanced([], pseudotime, weights)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_fit_advanced([[]], pseudotime, weights)) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0], [1.0]], [[0.0], [1.0]], [ + [1.0], + [1.0], + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[-1.0, 2.0]], [[0.0], [1.0]], [[1.0], [1.0]]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.5, 2.0]], [[0.0], [1.0]], [[1.0], [1.0]]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[0.0 / 0.0, 2.0]], [[0.0], [1.0]], [ + [1.0], + [1.0], + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 6) +} + +///| +test "tradeSeq advanced validates trajectory row dimensions" { + let mut failures = 0 + ignore(@src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0]], [[1.0]])) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_fit_advanced([[1.0, 2.0]], [[], []], [[], []])) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0, 2.0]], [[1.0], [1.0]]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced validates lineage weight semantics" { + let mut failures = 0 + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [[-1.0], [1.0]]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [[0.0], [1.0]]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [ + [0.0 / 0.0], + [1.0], + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced requires finite pseudotime for active lineages" { + let failed = try { + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0 / 0.0], [1.0]], [ + [1.0], + [1.0], + ]), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "tradeSeq advanced permits nonfinite inactive pseudotime" { + let fit = @src.tradeseq_fit_advanced( + [[1.0, 2.0, 3.0, 4.0]], + [[0.0, 0.0 / 0.0], [1.0, 0.0 / 0.0], [0.0 / 0.0, 3.0], [0.0 / 0.0, 4.0]], + [[1.0, 0.0], [1.0, 0.0], [0.0, 1.0], [0.0, 1.0]], + offsets=[0.0, 0.0, 0.0, 0.0], + config=@src.TradeSeqAdvancedConfig::create(n_knots=3), + ) catch { + _ => abort("inactive nonfinite pseudotime should be ignored") + } + assert_eq(fit.n_lineages, 2) +} + +///| +test "tradeSeq advanced requires pseudotime range in every lineage" { + let failed = try { + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [0.0]], [[1.0], [1.0]]), + ) + false + } catch { + _ => true + } + assert_true(failed) +} + +///| +test "tradeSeq advanced validates gene and lineage name counts" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let mut failures = 0 + ignore( + @src.tradeseq_fit_advanced( + tradeseq_adv_test_counts(), + pseudotime, + weights, + gene_names=["one"], + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced( + tradeseq_adv_test_counts(), + pseudotime, + weights, + lineage_names=["one"], + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "tradeSeq advanced validates nonempty unique names" { + let mut failures = 0 + ignore( + @src.tradeseq_fit_advanced( + [[1.0, 2.0], [2.0, 3.0]], + [[0.0], [1.0]], + [[1.0], [1.0]], + gene_names=["gene", "gene"], + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [[1.0], [1.0]], lineage_names=[ + " ", + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "tradeSeq advanced validates supplied offsets" { + let mut failures = 0 + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [[1.0], [1.0]], offsets=[ + 0.0, + ]), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_fit_advanced([[1.0, 2.0]], [[0.0], [1.0]], [[1.0], [1.0]], offsets=[ + 0.0, + 0.0 / 0.0, + ]), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "tradeSeq advanced fit exposes feature by cell dimensions" { + let fit = tradeseq_adv_test_fit() + assert_eq(fit.n_genes, 5) + assert_eq(fit.n_cells, 24) + assert_eq(fit.n_lineages, 2) + assert_eq(fit.gene_names, ["flat", "association", "endpoint", "early", "zero"]) + assert_eq(fit.lineage_names, ["left", "right"]) + assert_eq(fit.models.length(), 5) + assert_eq(fit.models[0].fitted_values.length(), 24) +} + +///| +test "tradeSeq advanced normalizes each cell lineage weight row" { + let fit = tradeseq_adv_test_fit() + assert_eq(fit.cell_weights[0], [0.5, 0.5]) + assert_eq(fit.cell_weights[8], [1.0, 0.0]) + assert_eq(fit.cell_weights[16], [0.0, 1.0]) + for row in fit.cell_weights { + let mut total = 0.0 + for value in row { + total = total + value + } + assert_true((total - 1.0).abs() < 1.0e-12) + } +} + +///| +test "tradeSeq advanced derives centered library size offsets" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let fit = @src.tradeseq_fit_advanced( + tradeseq_adv_test_counts(), + pseudotime, + weights, + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid automatic offsets should fit") + } + let mut total = 0.0 + for value in fit.offsets { + total = total + value + } + assert_true(total.abs() < 1.0e-10) +} + +///| +test "tradeSeq advanced preserves caller matrices" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let fit = @src.tradeseq_fit_advanced( + tradeseq_adv_test_counts(), + pseudotime, + weights, + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid caller matrices should fit") + } + fit.pseudotime[0][0] = 99.0 + fit.cell_weights[0][0] = 0.0 + assert_eq(pseudotime[0][0], 0.0) + assert_eq(weights[0][0], 2.0) +} + +///| +test "tradeSeq advanced models have finite penalized NB statistics" { + let fit = tradeseq_adv_test_fit() + for model in fit.models { + assert_eq(model.coefficients.length(), 8) + assert_eq(model.covariance.length(), 8) + assert_eq(model.covariance[0].length(), 8) + assert_true(model.dispersion >= fit.config.minimum_dispersion) + assert_true(model.dispersion <= fit.config.maximum_dispersion) + assert_true(tradeseq_adv_test_finite(model.log_likelihood)) + assert_true(tradeseq_adv_test_finite(model.aic)) + assert_eq(model.effective_df, 8.0) + assert_true(model.iterations >= 1) + for value in model.coefficients { + assert_true(tradeseq_adv_test_finite(value)) + } + for value in model.fitted_values { + assert_true(tradeseq_adv_test_finite(value)) + assert_true(value >= 0.0) + } + } +} + +///| +test "tradeSeq advanced handles an all-zero gene" { + let fit = tradeseq_adv_test_fit() + let model = fit.models[4] + assert_eq(model.gene_id, "zero") + assert_true(model.dispersion >= fit.config.minimum_dispersion) + for value in model.fitted_values { + assert_true(value < 1.0e-5) + } +} + +///| +test "tradeSeq advanced fitted trend follows increasing counts" { + let fit = tradeseq_adv_test_fit() + let fitted = fit.models[1].fitted_values + assert_true(fitted[15] > fitted[0]) + assert_true(fitted[23] > fitted[0]) +} + +///| +test "tradeSeq advanced associationTest returns gene-level Wald results" { + let table = @src.tradeseq_association_test_advanced(tradeseq_adv_test_fit()) + assert_eq(table.test_name, "associationTest") + assert_eq(table.results.length(), 5) + assert_eq(table.fdr_threshold, 0.1) + for result in table.results { + assert_true(result.degrees_freedom > 0) + assert_true(result.wald_statistic >= 0.0) + assert_true(result.p_value >= 0.0 && result.p_value <= 1.0) + assert_true( + result.adjusted_p_value >= 0.0 && result.adjusted_p_value <= 1.0, + ) + assert_true(result.adjusted_p_value + 1.0e-12 >= result.p_value) + } +} + +///| +test "tradeSeq advanced associationTest ranks temporal signal above flat gene" { + let table = @src.tradeseq_association_test_advanced(tradeseq_adv_test_fit()) + assert_true(table.results[1].wald_statistic > table.results[0].wald_statistic) + assert_true(table.results[1].p_value < table.results[0].p_value) +} + +///| +test "tradeSeq advanced startVsEndTest detects endpoint change" { + let table = @src.tradeseq_start_vs_end_test_advanced(tradeseq_adv_test_fit()) + assert_eq(table.test_name, "startVsEndTest") + assert_eq(table.results.length(), 5) + assert_eq(table.results[0].degrees_freedom, 2) + assert_true(table.results[1].wald_statistic > table.results[0].wald_statistic) +} + +///| +test "tradeSeq advanced diffEndTest detects lineage endpoint separation" { + let table = @src.tradeseq_diff_end_test_advanced(tradeseq_adv_test_fit()) + assert_eq(table.test_name, "diffEndTest") + assert_eq(table.results[0].degrees_freedom, 1) + assert_true(table.results[2].wald_statistic > table.results[0].wald_statistic) +} + +///| +test "tradeSeq advanced patternTest compares complete lineage smooths" { + let table = @src.tradeseq_pattern_test_advanced( + tradeseq_adv_test_fit(), + n_points=5, + ) + assert_eq(table.test_name, "patternTest") + assert_eq(table.results.length(), 5) + assert_true(table.results[2].degrees_freedom >= 2) + assert_true(table.results[2].wald_statistic > table.results[0].wald_statistic) +} + +///| +test "tradeSeq advanced earlyDETest restricts the comparison interval" { + let table = @src.tradeseq_early_de_test_advanced( + tradeseq_adv_test_fit(), + 0.35, + 0.75, + n_points=5, + ) + assert_eq(table.test_name, "earlyDETest") + assert_eq(table.results.length(), 5) + assert_true(table.results[3].degrees_freedom >= 2) + assert_true(table.results[3].wald_statistic > table.results[0].wald_statistic) +} + +///| +test "tradeSeq advanced fold-change threshold shrinks Wald evidence" { + let fit = tradeseq_adv_test_fit() + let unthresholded = @src.tradeseq_start_vs_end_test_advanced(fit) + let thresholded = @src.tradeseq_start_vs_end_test_advanced(fit, l2fc=2.0) + for gene in 0..= + unthresholded.results[gene].p_value, + ) + } +} + +///| +test "tradeSeq advanced tests validate fold-change threshold" { + let fit = tradeseq_adv_test_fit() + let mut failures = 0 + ignore(@src.tradeseq_association_test_advanced(fit, l2fc=-1.0)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_start_vs_end_test_advanced(fit, l2fc=0.0 / 0.0)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "tradeSeq advanced lineage comparisons require multiple lineages" { + let fit = tradeseq_adv_test_single_lineage_fit() + let mut failures = 0 + ignore(@src.tradeseq_diff_end_test_advanced(fit)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_pattern_test_advanced(fit)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_early_de_test_advanced(fit, 0.0, 0.5)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced validates pattern and earlyDE grids" { + let fit = tradeseq_adv_test_fit() + let mut failures = 0 + ignore(@src.tradeseq_pattern_test_advanced(fit, n_points=1)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_early_de_test_advanced(fit, -0.1, 0.5)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_early_de_test_advanced(fit, 0.5, 0.5)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_early_de_test_advanced(fit, 0.5, 1.1)) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_early_de_test_advanced(fit, 0.0, 0.5, n_points=1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 5) +} + +///| +test "tradeSeq advanced predicts smooths on each lineage" { + let prediction = @src.tradeseq_predict_smooth_advanced( + tradeseq_adv_test_fit(), + "association", + n_points=9, + ) + assert_eq(prediction.gene_id, "association") + assert_eq(prediction.pseudotime.length(), 2) + assert_eq(prediction.pseudotime[0].length(), 9) + assert_eq(prediction.fitted.length(), 2) + assert_eq(prediction.standard_errors[1].length(), 9) + assert_eq(prediction.pseudotime[0][0], 0.0) + assert_true((prediction.pseudotime[0][8] - 0.96).abs() < 1.0e-12) +} + +///| +test "tradeSeq advanced smooth predictions are finite and ordered" { + let prediction = @src.tradeseq_predict_smooth_advanced( + tradeseq_adv_test_fit(), + "association", + n_points=9, + ) + for lineage in 0..<2 { + assert_true(prediction.fitted[lineage][8] > prediction.fitted[lineage][0]) + for point in 0..<9 { + assert_true(tradeseq_adv_test_finite(prediction.fitted[lineage][point])) + assert_true( + tradeseq_adv_test_finite(prediction.standard_errors[lineage][point]), + ) + assert_true(prediction.standard_errors[lineage][point] >= 0.0) + } + } +} + +///| +test "tradeSeq advanced smooth prediction validates gene and grid" { + let fit = tradeseq_adv_test_fit() + let mut failures = 0 + ignore(@src.tradeseq_predict_smooth_advanced(fit, "missing")) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_predict_smooth_advanced(fit, "flat", n_points=1)) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 2) +} + +///| +test "tradeSeq advanced evaluateK returns gene by candidate AIC" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let counts = tradeseq_adv_test_counts() + let evaluation = @src.tradeseq_evaluate_k_advanced( + [counts[0], counts[1]], + pseudotime, + weights, + [3, 4], + gene_names=["flat", "association"], + offsets=Array::make(24, 0.0), + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid knot candidates should be evaluated") + } + assert_eq(evaluation.candidates, [3, 4]) + assert_eq(evaluation.gene_aic.length(), 2) + assert_eq(evaluation.gene_aic[0].length(), 2) + assert_eq(evaluation.mean_aic.length(), 2) + assert_true( + evaluation.selected_n_knots == 3 || evaluation.selected_n_knots == 4, + ) + for value in evaluation.mean_aic { + assert_true(tradeseq_adv_test_finite(value)) + } +} + +///| +test "tradeSeq advanced evaluateK validates candidate set" { + let (pseudotime, weights) = tradeseq_adv_test_trajectory() + let counts = tradeseq_adv_test_counts() + let mut failures = 0 + ignore(@src.tradeseq_evaluate_k_advanced(counts, pseudotime, weights, [3])) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_evaluate_k_advanced(counts, pseudotime, weights, [2, 3])) catch { + _ => failures = failures + 1 + } + ignore(@src.tradeseq_evaluate_k_advanced(counts, pseudotime, weights, [3, 3])) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 3) +} + +///| +test "tradeSeq advanced fits directly from Slingshot optional pseudotime" { + let slingshot = tradeseq_adv_test_slingshot_result() + let first : Array[Double] = [] + let second : Array[Double] = [] + for cell in 0..<15 { + first.push((cell + 2).to_double()) + second.push((5 + cell % 4).to_double()) + } + let fit = @src.tradeseq_fit_from_slingshot( + [first, second], + ["trend", "stable"], + slingshot, + offsets=Array::make(15, 0.0), + config=@src.TradeSeqAdvancedConfig::create( + n_knots=3, + max_iterations=50, + ridge=1.0e-5, + test_points=4, + ), + ) catch { + _ => abort("Slingshot output should feed tradeSeq directly") + } + assert_eq(fit.n_cells, 15) + assert_eq(fit.n_lineages, slingshot.curves.length()) + assert_eq(fit.lineage_names[0], slingshot.curves[0].name) + for row in fit.cell_weights { + let mut total = 0.0 + for value in row { + total = total + value + } + assert_true((total - 1.0).abs() < 1.0e-10) + } +} + +///| +test "tradeSeq advanced SCE writes fitted assays and row statistics" { + let output = @src.tradeseq_advanced_sce( + tradeseq_adv_test_sce(), + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid SCE trajectory should fit") + } + assert_eq(output.fit.n_genes, 5) + assert_eq(output.experiment.assays["tradeSeq.fitted"].length(), 5) + assert_eq(output.experiment.assays["tradeSeq.fitted"][0].length(), 24) + assert_eq( + output.experiment.row_data["tradeSeq.association.pvalue"].length(), + 5, + ) + assert_eq(output.experiment.row_data["tradeSeq.startVsEnd.padj"].length(), 5) + assert_eq(output.experiment.row_data["tradeSeq.dispersion"].length(), 5) + assert_eq(output.experiment.metadata["tradeSeq.nKnots"], "4") + assert_eq( + output.experiment.metadata["tradeSeq.lineages"], + "Lineage1,Lineage2", + ) +} + +///| +test "tradeSeq advanced SCE preserves input immutability" { + let experiment = tradeseq_adv_test_sce() + let output = @src.tradeseq_advanced_sce( + experiment, + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid SCE trajectory should fit") + } + output.experiment.assays["counts"][0][0] = 99.0 + output.experiment.reduced_dims["slingshot.pseudotime"][0][0] = 99.0 + output.experiment.row_data["symbol"][0] = "changed" + output.experiment.col_data["batch"][0] = "changed" + assert_eq(experiment.assays["counts"][0][0], 10.0) + assert_eq(experiment.reduced_dims["slingshot.pseudotime"][0][0], 0.0) + assert_eq(experiment.row_data["symbol"][0], "F") + assert_eq(experiment.col_data["batch"][0], "one") + assert_false(experiment.assays.contains("tradeSeq.fitted")) +} + +///| +test "tradeSeq advanced SCE recursively copies alternative experiments" { + let experiment = tradeseq_adv_test_sce() + let output = @src.tradeseq_advanced_sce( + experiment, + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("valid SCE trajectory should fit") + } + output.experiment.alternative_experiments["spike"].assays["counts"][0][0] = 8.0 + assert_eq( + experiment.alternative_experiments["spike"].assays["counts"][0][0], + 1.0, + ) +} + +///| +test "tradeSeq advanced SCE supports custom input and output names" { + let experiment = tradeseq_adv_test_sce() + experiment.assays["raw"] = experiment.assays["counts"].map(fn(row) { + row.copy() + }) + experiment.reduced_dims["ptime"] = experiment.reduced_dims["slingshot.pseudotime"].map(fn( + row, + ) { + row.copy() + }, + ) + experiment.reduced_dims["lineageWeight"] = experiment.reduced_dims["slingshot.weights"].map(fn( + row, + ) { + row.copy() + }, + ) + let output = @src.tradeseq_advanced_sce( + experiment, + assay_name="raw", + pseudotime_name="ptime", + weights_name="lineageWeight", + output_prefix="ts", + config=tradeseq_adv_test_config(), + ) catch { + _ => abort("custom SCE names should be supported") + } + assert_true(output.experiment.assays.contains("ts.fitted")) + assert_true(output.experiment.row_data.contains("ts.association.padj")) + assert_eq(output.experiment.metadata["ts.nKnots"], "4") +} + +///| +test "tradeSeq advanced SCE validates required entries and output prefix" { + let experiment = tradeseq_adv_test_sce() + let mut failures = 0 + ignore( + @src.tradeseq_advanced_sce( + experiment, + assay_name="missing", + config=tradeseq_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_advanced_sce( + experiment, + pseudotime_name="missing", + config=tradeseq_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_advanced_sce( + experiment, + weights_name="missing", + config=tradeseq_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + ignore( + @src.tradeseq_advanced_sce( + experiment, + output_prefix=" ", + config=tradeseq_adv_test_config(), + ), + ) catch { + _ => failures = failures + 1 + } + assert_eq(failures, 4) +} + +///| +test "tradeSeq advanced summary reports fit dimensions" { + let summary = @src.tradeseq_advanced_summary(tradeseq_adv_test_fit()) + assert_true(summary.contains("tradeSeq 1.27.0 advanced fit")) + assert_true(summary.contains("genes: 5")) + assert_true(summary.contains("cells: 24")) + assert_true(summary.contains("lineages: 2")) + assert_true(summary.contains("knots: 4")) + assert_true(summary.contains("mean dispersion:")) +} diff --git a/test/moonbit/tradeseq_test.mbt b/test/moonbit/tradeseq_test.mbt index 985663c7..0ce2abc6 100644 --- a/test/moonbit/tradeseq_test.mbt +++ b/test/moonbit/tradeseq_test.mbt @@ -4,9 +4,9 @@ test "tradeseq_create_trajectory_point" { "cell_001", 0.5, [1.0, 2.0, 3.0], - "control" + "control", ) - + assert_eq(point.cell_id, "cell_001") assert_true((point.pseudotime - 0.5).abs() < 1.0e-10) assert_eq(point.expression.length(), 3) @@ -19,9 +19,9 @@ test "tradeseq_create_gene_expression" { "GeneA", [1.0, 2.0, 3.0, 4.0], [0.1, 0.3, 0.5, 0.7], - ["control", "control", "treatment", "treatment"] + ["control", "control", "treatment", "treatment"], ) - + assert_eq(data.gene_id, "GeneA") assert_eq(data.expression.length(), 4) assert_eq(data.pseudotime.length(), 4) @@ -31,20 +31,15 @@ test "tradeseq_create_gene_expression" { ///| test "tradeseq_trajectory_data" { let mut data = @src.TrajectoryData::new() - + assert_eq(data.n_points, 0) assert_eq(data.n_genes, 0) - + data = data.add_gene("GeneA") assert_eq(data.n_genes, 1) - - let point = @src.TrajectoryPoint::new( - "cell_001", - 0.5, - [1.0], - "control" - ) - + + let point = @src.TrajectoryPoint::new("cell_001", 0.5, [1.0], "control") + data = data.add_point(point) assert_eq(data.n_points, 1) } @@ -55,11 +50,11 @@ test "tradeseq_fit_gam" { "GeneA", [1.0, 2.0, 3.0, 4.0, 5.0, 6.0], [0.0, 0.2, 0.4, 0.6, 0.8, 1.0], - ["control", "control", "control", "treatment", "treatment", "treatment"] + ["control", "control", "control", "treatment", "treatment", "treatment"], ) - + let gam = @src.fit_gam(gene_data, 4) - + assert_eq(gam.gene_id, "GeneA") assert_true(gam.fitted_values.length() > 0) } @@ -70,11 +65,13 @@ test "tradeseq_trade_test" { "GeneA", [1.0, 1.2, 1.5, 2.0, 2.5, 3.0], [0.0, 0.2, 0.4, 0.6, 0.8, 1.0], - ["control", "control", "control", "treatment", "treatment", "treatment"] + ["control", "control", "control", "treatment", "treatment", "treatment"], + ) + + let result = @src.trade_test_condition_effect( + gene_data, 0.5, "control", "treatment", ) - - let result = @src.trade_test_condition_effect(gene_data, 0.5, "control", "treatment") - + assert_eq(result.gene_id, "GeneA") assert_true(result.p_value >= 0.0) // Just check p_value is non-negative @@ -83,18 +80,18 @@ test "tradeseq_trade_test" { ///| test "tradeseq_run_analysis" { let trajectory_data = @src.create_example_trajectory_data() - + assert_true(trajectory_data.n_points > 0) assert_true(trajectory_data.n_genes > 0) - + let result = @src.run_tradeseq_analysis( trajectory_data, trajectory_data.genes, trajectory_data.conditions, 4, - 0.05 + 0.05, ) - + assert_eq(result.n_genes, 3) assert_true(result.n_significant >= 0) assert_true(result.n_significant <= result.n_genes) @@ -103,9 +100,9 @@ test "tradeseq_run_analysis" { ///| test "tradeseq_calculate_smooth" { let trajectory_data = @src.create_example_trajectory_data() - + let smooth = @src.calculate_gene_smooth(trajectory_data, "GeneA", 4, 20) - + assert_eq(smooth.gene_id, "GeneA") assert_eq(smooth.pseudotime.length(), 20) assert_eq(smooth.fitted.length(), 20) @@ -119,11 +116,11 @@ test "tradeseq_summary" { trajectory_data.genes, trajectory_data.conditions, 4, - 0.05 + 0.05, ) - + let summary = @src.tradeseq_summary(result) - + assert_true(summary.contains("tradeSeq Analysis Summary")) assert_true(summary.contains("Total genes tested:")) } @@ -131,7 +128,7 @@ test "tradeseq_summary" { ///| test "tradeseq_create_example" { let data = @src.create_example_trajectory_data() - + assert_true(data.n_points == 50) assert_true(data.n_genes == 3) assert_true(data.conditions.length() >= 2) @@ -143,11 +140,13 @@ test "tradeseq_condition_effect_zero" { "GeneA", [1.0, 1.0, 1.0, 1.0, 1.0, 1.0], [0.0, 0.2, 0.4, 0.6, 0.8, 1.0], - ["control", "control", "control", "treatment", "treatment", "treatment"] + ["control", "control", "control", "treatment", "treatment", "treatment"], + ) + + let result = @src.trade_test_condition_effect( + gene_data, 0.5, "control", "treatment", ) - - let result = @src.trade_test_condition_effect(gene_data, 0.5, "control", "treatment") - + // No difference should give high p-value assert_true(result.p_value > 0.05) } @@ -156,25 +155,19 @@ test "tradeseq_condition_effect_zero" { test "tradeseq_trajectory_empty" { let mut data = @src.TrajectoryData::new() data = data.add_gene("GeneA") - - let result = @src.run_tradeseq_analysis( - data, - ["GeneA"], - ["control"], - 4, - 0.05 - ) - + + let result = @src.run_tradeseq_analysis(data, ["GeneA"], ["control"], 4, 0.05) + assert_eq(result.n_genes, 1) } ///| test "tradeseq_multiple_smooth" { let trajectory_data = @src.create_example_trajectory_data() - + let smooth_a = @src.calculate_gene_smooth(trajectory_data, "GeneA", 4, 15) let smooth_b = @src.calculate_gene_smooth(trajectory_data, "GeneB", 4, 15) - + assert_eq(smooth_a.gene_id, "GeneA") assert_eq(smooth_b.gene_id, "GeneB") assert_eq(smooth_a.fitted.length(), 15) diff --git a/test/moonbit/tree_summarized_experiment_test.mbt b/test/moonbit/tree_summarized_experiment_test.mbt new file mode 100644 index 00000000..389db219 --- /dev/null +++ b/test/moonbit/tree_summarized_experiment_test.mbt @@ -0,0 +1,254 @@ +///| +fn tse_test_row_tree() -> @src.Tree { + let a = @src.Clade::new(name=Some("A")) + let b = @src.Clade::new(name=Some("B")) + let c = @src.Clade::new(name=Some("C")) + let d = @src.Clade::new(name=Some("D")) + let group_a = @src.Clade::new(name=Some("GroupA"), clades=[a, b]) + let group_b = @src.Clade::new(name=Some("GroupB"), clades=[c, d]) + let root = @src.Clade::new(name=Some("All"), clades=[group_a, group_b]) + @src.Tree::new(root, rooted=true, name=Some("taxa")) +} + +///| +fn tse_test_col_tree() -> @src.Tree { + let s1 = @src.Clade::new(name=Some("S1")) + let s2 = @src.Clade::new(name=Some("S2")) + let s3 = @src.Clade::new(name=Some("S3")) + let s4 = @src.Clade::new(name=Some("S4")) + let control = @src.Clade::new(name=Some("Control"), clades=[s1, s2]) + let treated = @src.Clade::new(name=Some("Treated"), clades=[s3, s4]) + let root = @src.Clade::new(name=Some("Samples"), clades=[control, treated]) + @src.Tree::new(root, rooted=true, name=Some("samples")) +} + +///| +fn tse_test_experiment() -> @src.SummarizedExperiment { + @src.summarized_experiment( + Map([ + ( + "counts", + [ + [1.0, 2.0, 3.0, 4.0], + [3.0, 4.0, 5.0, 6.0], + [5.0, 6.0, 7.0, 8.0], + [7.0, 8.0, 9.0, 10.0], + ], + ), + ( + "normalized", + [ + [0.1, 0.2, 0.3, 0.4], + [0.3, 0.4, 0.5, 0.6], + [0.5, 0.6, 0.7, 0.8], + [0.7, 0.8, 0.9, 1.0], + ], + ), + ]), + [("A", 0, 0), ("B", 0, 0), ("C", 0, 0), ("D", 0, 0)], + [ + Map([("sample", "S1")]), + Map([("sample", "S2")]), + Map([("sample", "S3")]), + Map([("sample", "S4")]), + ], + Map([("study", "tree-demo")]), + ) +} + +///| +fn tse_test_object() -> @src.TreeSummarizedExperiment { + @src.TreeSummarizedExperiment::new( + experiment=tse_test_experiment(), + row_tree=Some(tse_test_row_tree()), + row_node_labels=["A", "B", "C", "D"], + col_tree=Some(tse_test_col_tree()), + col_node_labels=["S1", "S2", "S3", "S4"], + reference_sequences=["AAAA", "CCCC", "GGGG", "TTTT"], + ) +} + +///| +test "tree_summarized_experiment: construct and access links" { + let tse = tse_test_object() + assert_true(tse.is_valid()) + assert_eq(tse.nrow(), 4) + assert_eq(tse.ncol(), 4) + assert_eq(tse.assay_names().length(), 2) + assert_eq(tse.row_links().length(), 4) + assert_eq(tse.col_links().length(), 4) + assert_eq(tse.row_links()[0].node_label(), "A") + assert_eq(tse.row_links()[0].node_alias(), "alias_1") + assert_eq(tse.row_links()[0].node_number(), 1) + assert_true(tse.row_links()[0].is_leaf()) + assert_eq(tse.row_links()[0].tree_name(), "taxa") + assert_eq(tse.reference_sequences()[3], "TTTT") +} + +///| +test "tree_summarized_experiment: invalid link is detected" { + let tse = @src.TreeSummarizedExperiment::new( + experiment=tse_test_experiment(), + row_tree=Some(tse_test_row_tree()), + row_node_labels=["A", "B", "missing", "D"], + ) + assert_false(tse.is_valid()) + assert_eq(tse.row_links().length(), 3) +} + +///| +test "tree_summarized_experiment: tree node queries" { + let tree = tse_test_row_tree() + assert_eq(@src.tse_find_descendants(tree, "GroupA", leaves_only=true), [ + "A", "B", + ]) + assert_eq(@src.tse_find_ancestors(tree, "C"), ["All", "GroupB"]) + assert_true(@src.tse_is_leaf(tree, "D")) + assert_false(@src.tse_is_leaf(tree, "GroupB")) + assert_eq(@src.tse_find_descendants(tree, "missing").length(), 0) +} + +///| +test "tree_summarized_experiment: subset rows keeps links and sequences" { + let subset = tse_test_object().subset_rows([1, 3]) + assert_true(subset.is_valid()) + assert_eq(subset.nrow(), 2) + assert_eq(subset.row_links()[0].node_label(), "B") + assert_eq(subset.row_links()[1].node_label(), "D") + assert_eq(subset.reference_sequences(), ["CCCC", "TTTT"]) + match subset.assay("counts") { + Some(assay) => + assert_eq(assay, [[3.0, 4.0, 5.0, 6.0], [7.0, 8.0, 9.0, 10.0]]) + None => assert_true(false) + } +} + +///| +test "tree_summarized_experiment: subset by row tree node" { + match tse_test_object().subset_by_row_nodes(["GroupA"]) { + None => assert_true(false) + Some(subset) => { + assert_true(subset.is_valid()) + assert_eq(subset.nrow(), 2) + assert_eq(subset.row_links()[0].node_label(), "A") + assert_eq(subset.row_links()[1].node_label(), "B") + } + } +} + +///| +test "tree_summarized_experiment: subset by column tree node" { + match tse_test_object().subset_by_col_nodes(["Treated"]) { + None => assert_true(false) + Some(subset) => { + assert_true(subset.is_valid()) + assert_eq(subset.ncol(), 2) + assert_eq(subset.col_links()[0].node_label(), "S3") + assert_eq(subset.col_links()[1].node_label(), "S4") + } + } +} + +///| +test "tree_summarized_experiment: aggregate rows by sum" { + match tse_test_object().aggregate_rows(["GroupA", "GroupB"]) { + None => assert_true(false) + Some(aggregated) => { + assert_true(aggregated.is_valid()) + assert_eq(aggregated.nrow(), 2) + assert_eq(aggregated.row_links()[0].node_label(), "GroupA") + assert_false(aggregated.row_links()[0].is_leaf()) + match aggregated.assay("counts") { + Some(assay) => + assert_eq(assay, [[4.0, 6.0, 8.0, 10.0], [12.0, 14.0, 16.0, 18.0]]) + None => assert_true(false) + } + match aggregated.assay("normalized") { + Some(assay) => { + assert_true((assay[0][0] - 0.4).abs() < 1.0e-12) + assert_true((assay[0][1] - 0.6).abs() < 1.0e-12) + assert_true((assay[0][2] - 0.8).abs() < 1.0e-12) + assert_true((assay[0][3] - 1.0).abs() < 1.0e-12) + } + None => assert_true(false) + } + } + } +} + +///| +test "tree_summarized_experiment: aggregate rows by mean min and max" { + let tse = tse_test_object() + match + tse.aggregate_rows(["GroupA"], aggregation=@src.tse_aggregation_mean()) { + Some(result) => + match result.assay("counts") { + Some(assay) => assert_eq(assay[0], [2.0, 3.0, 4.0, 5.0]) + None => assert_true(false) + } + None => assert_true(false) + } + match tse.aggregate_rows(["GroupA"], aggregation=@src.tse_aggregation_min()) { + Some(result) => + match result.assay("counts") { + Some(assay) => assert_eq(assay[0], [1.0, 2.0, 3.0, 4.0]) + None => assert_true(false) + } + None => assert_true(false) + } + match tse.aggregate_rows(["GroupA"], aggregation=@src.tse_aggregation_max()) { + Some(result) => + match result.assay("counts") { + Some(assay) => assert_eq(assay[0], [3.0, 4.0, 5.0, 6.0]) + None => assert_true(false) + } + None => assert_true(false) + } +} + +///| +test "tree_summarized_experiment: aggregate columns" { + match tse_test_object().aggregate_cols(["Control", "Treated"]) { + None => assert_true(false) + Some(aggregated) => { + assert_true(aggregated.is_valid()) + assert_eq(aggregated.ncol(), 2) + assert_eq(aggregated.col_links()[0].node_label(), "Control") + match aggregated.assay("counts") { + Some(assay) => { + assert_eq(assay[0], [3.0, 7.0]) + assert_eq(assay[3], [15.0, 19.0]) + } + None => assert_true(false) + } + } + } +} + +///| +test "tree_summarized_experiment: duplicate links can aggregate" { + let tse = @src.TreeSummarizedExperiment::new( + experiment=tse_test_experiment(), + row_tree=Some(tse_test_row_tree()), + row_node_labels=["A", "A", "C", "D"], + ) + assert_true(tse.is_valid()) + match tse.aggregate_rows(["A"]) { + Some(aggregated) => + match aggregated.assay("counts") { + Some(assay) => assert_eq(assay[0], [4.0, 6.0, 8.0, 10.0]) + None => assert_true(false) + } + None => assert_true(false) + } +} + +///| +test "tree_summarized_experiment: works without trees" { + let tse = @src.TreeSummarizedExperiment::new(experiment=tse_test_experiment()) + assert_true(tse.is_valid()) + assert_eq(tse.row_links().length(), 0) + assert_eq(tse.col_links().length(), 0) + assert_true(tse.aggregate_rows(["All"]) is None) + assert_true(tse.subset_by_col_nodes(["Samples"]) is None) +} diff --git a/test/moonbit/trie_test.mbt b/test/moonbit/trie_test.mbt index 365aebe8..fbe4a463 100644 --- a/test/moonbit/trie_test.mbt +++ b/test/moonbit/trie_test.mbt @@ -306,7 +306,9 @@ test "triefind_find_words_with_boundaries" { let trie = @src.Trie::new() @src.trie_insert(trie, "EcoRI", "enzyme1") @src.trie_insert(trie, "BamHI", "enzyme2") - let results = @src.triefind_find_words(trie, "Use EcoRI and BamHI for cloning") + let results = @src.triefind_find_words( + trie, "Use EcoRI and BamHI for cloning", + ) assert_eq(results.length(), 2) } diff --git a/test/moonbit/twobit_io_test.mbt b/test/moonbit/twobit_io_test.mbt index 403f8ca3..50e04b6a 100644 --- a/test/moonbit/twobit_io_test.mbt +++ b/test/moonbit/twobit_io_test.mbt @@ -223,7 +223,9 @@ test "twobit_unpack_sequence_simple" { ///| test "twobit_pack_sequence_with_n_blocks" { - let (packed, n_blocks, mask_blocks) = @src.twobit_pack_sequence("ATGCNNNNATGC") + let (packed, n_blocks, mask_blocks) = @src.twobit_pack_sequence( + "ATGCNNNNATGC", + ) assert_eq(packed.length(), 3) // ceil(12/4) = 3 bytes assert_eq(n_blocks.length(), 1) assert_eq(n_blocks[0].start(), 4) @@ -234,7 +236,9 @@ test "twobit_pack_sequence_with_n_blocks" { ///| test "twobit_unpack_sequence_with_n_blocks" { - let (packed, n_blocks, mask_blocks) = @src.twobit_pack_sequence("ATGCNNNNATGC") + let (packed, n_blocks, mask_blocks) = @src.twobit_pack_sequence( + "ATGCNNNNATGC", + ) let result = @src.twobit_unpack_sequence(packed, 12, n_blocks, mask_blocks) assert_eq(result, "ATGCNNNNATGC") } diff --git a/test/moonbit/unigene_test.mbt b/test/moonbit/unigene_test.mbt new file mode 100644 index 00000000..67bfc0d1 --- /dev/null +++ b/test/moonbit/unigene_test.mbt @@ -0,0 +1,291 @@ +///| +/// Tests for the Biopython Bio.UniGene-compatible parser. + +///| +test "unigene parses core record fields" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.identifier, "Hs.12345") + assert_eq(record.species, "Hs") + assert_eq(record.title, "Example kinase family member") + assert_eq(record.symbol, "EXK1") + assert_eq(record.cytoband, "7q31.2") + assert_eq(record.gene_id, "12345") + assert_eq(record.locuslink, "12345") + assert_eq(record.chromosome, "7") + assert_eq(record.declared_sequence_count, 3) +} + +///| +test "unigene parses expression and homology fields" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.expression, ["brain", "heart", "adult"]) + assert_eq(record.restricted_expression, ["brain", "adult"]) + assert_eq(record.homology, Some(true)) + assert_eq(record.genomic_terminus, "T") +} + +///| +test "unigene parses protein similarity fields" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.protein_similarities.length(), 1) + let similarity = record.protein_similarities[0] + assert_eq(similarity.organism, "9606") + assert_eq(similarity.protein_gi, "123456") + assert_eq(similarity.protein_id, "NP_000001.1") + assert_eq(similarity.percent, "99.50") + assert_eq(similarity.alignment_length, "401") + assert_eq(similarity.field("PCT"), Some("99.50")) +} + +///| +test "unigene parses sequence entries and IMAGE clones" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.sequences.length(), 3) + let transcript = record.sequences[0] + assert_eq(transcript.accession, "NM_000001.2") + assert_eq(transcript.nucleotide_id, "g123455") + assert_eq(transcript.protein_id, "g123456") + assert_eq(transcript.sequence_type, "mRNA") + let image = record.sequences[1] + assert_true(image.is_image) + assert_eq(image.image_id, "123456") + assert_eq(image.read_end, "5'") + assert_eq(image.library_id, "100") + assert_eq(image.trace, "900001") +} + +///| +test "unigene parses STS and transcript map entries" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.sts.length(), 1) + assert_eq(record.sts[0].accession, "G12345") + assert_eq(record.sts[0].unists, "76543") + assert_eq(record.transcript_maps.length(), 1) + assert_eq(record.transcript_maps[0].marker, "D7S1234") + assert_eq(record.transcript_maps[0].radiation_hybrid_panel, "GB4") +} + +///| +test "unigene sequence type and accession queries" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.sequences_of_type("EST").length(), 2) + assert_eq(record.sequences_of_type("mRNA").length(), 1) + assert_eq(record.sequences_of_type("HTC").length(), 0) + match record.find_sequence("AA000002.1") { + Some(sequence) => assert_eq(sequence.mgc, "600") + None => abort("expected sequence accession") + } + assert_true(record.find_sequence("missing") is None) +} + +///| +test "unigene similarity and IMAGE queries" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.protein_similarities_for("9606").length(), 1) + assert_eq(record.protein_similarities_for("10090").length(), 0) + let image_sequences = record.image_sequences() + assert_eq(image_sequences.length(), 1) + assert_eq(image_sequences[0].accession, "AA000001.1") +} + +///| +test "unigene preserves unknown child fields" { + let text = "ID Mm.7\n" + + "PROTSIM ORG=10090; PROTID=NP_1; EVALUE=1e-30\n" + + "TXMAP MARKER=D1Mit1; METHOD=RH\n" + + "SCOUNT 1\n" + + "SEQUENCE ACC=NM_1; SEQTYPE=mRNA; STATUS=reviewed\n" + + "STS UNISTS=10 PANEL=RH\n" + + "//\n" + let record = @src.unigene_read(text) + assert_eq(record.protein_similarities[0].field("EVALUE"), Some("1e-30")) + assert_eq(record.transcript_maps[0].field("METHOD"), Some("RH")) + assert_eq(record.sequences[0].field("STATUS"), Some("reviewed")) + assert_eq(record.sts[0].field("PANEL"), Some("RH")) + let reparsed = @src.unigene_read(record.to_string()) + assert_eq(reparsed.protein_similarities[0].field("EVALUE"), Some("1e-30")) + assert_eq(reparsed.sequences[0].field("STATUS"), Some("reviewed")) +} + +///| +test "unigene parses multiple records" { + let second = "ID Mm.2\n" + + "TITLE Second cluster\n" + + "GENE Gene2\n" + + "HOMOL NO\n" + + "SCOUNT 0\n" + + "//\n" + let records = @src.unigene_parse(@src.unigene_sample_text() + second) + assert_eq(records.length(), 2) + assert_eq(records[0].identifier, "Hs.12345") + assert_eq(records[1].identifier, "Mm.2") + assert_eq(records[1].species, "Mm") + assert_eq(records[1].homology, Some(false)) +} + +///| +test "unigene read accepts exactly one record" { + let record = @src.unigene_read(@src.unigene_sample_text()) + assert_eq(record.identifier, "Hs.12345") + assert_true(record.summary().contains("sequences=3")) +} + +///| +test "unigene read rejects empty input" { + let raised = try { + ignore(@src.unigene_read("")) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene read rejects multiple records" { + let raised = try { + ignore( + @src.unigene_read(@src.unigene_sample_text() + @src.unigene_sample_text()), + ) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene validates declared sequence count" { + let text = "ID Hs.1\n" + + "SCOUNT 2\n" + + "SEQUENCE ACC=NM_1; SEQTYPE=mRNA\n" + + "//\n" + let raised = try { + ignore(@src.unigene_parse(text)) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene requires sequence count" { + let raised = try { + ignore(@src.unigene_parse("ID Hs.1\n//\n")) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene rejects invalid homology value" { + let text = "ID Hs.1\nHOMOL MAYBE\nSCOUNT 0\n//\n" + let raised = try { + ignore(@src.unigene_parse(text)) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene rejects unknown top-level tags" { + let text = "ID Hs.1\nUNKNOWN value\nSCOUNT 0\n//\n" + let raised = try { + ignore(@src.unigene_parse(text)) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene rejects malformed child fields" { + let text = "ID Hs.1\n" + + "SCOUNT 1\n" + + "SEQUENCE ACC=NM_1; INVALID\n" + + "//\n" + let raised = try { + ignore(@src.unigene_parse(text)) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene rejects unterminated records" { + let raised = try { + ignore(@src.unigene_parse("ID Hs.1\nSCOUNT 0\n")) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene rejects malformed fixed-width lines" { + let raised = try { + ignore(@src.unigene_parse("ID Hs.1\n")) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} + +///| +test "unigene supports CRLF input" { + let crlf = @src.unigene_sample_text().replace_all(old="\n", new="\r\n") + let record = @src.unigene_read(crlf) + assert_eq(record.identifier, "Hs.12345") + assert_eq(record.sequences.length(), 3) +} + +///| +test "unigene record serialization round trip" { + let original = @src.unigene_read(@src.unigene_sample_text()) + let serialized = original.to_string() + assert_true(serialized.contains("ID Hs.12345")) + assert_true(serialized.contains("SCOUNT 3")) + let reparsed = @src.unigene_read(serialized) + assert_true(reparsed == original) +} + +///| +test "unigene constructors serialize canonical records" { + let record = @src.UniGeneRecord::new( + identifier="Rn.9", + title="Constructed record", + symbol="Gene9", + homology=Some(false), + sequences=[ + @src.UniGeneSequence::new(accession="NM_9", sequence_type="mRNA"), + ], + sts=[@src.UniGeneSTS::new(unists="999")], + ) + let reparsed = @src.unigene_read(record.to_string()) + assert_eq(reparsed.species, "Rn") + assert_eq(reparsed.homology, Some(false)) + assert_eq(reparsed.sequences[0].accession, "NM_9") + assert_eq(reparsed.sts[0].unists, "999") +} + +///| +test "unigene rejects duplicate sequence counts" { + let text = "ID Hs.1\nSCOUNT 0\nSCOUNT 0\n//\n" + let raised = try { + ignore(@src.unigene_parse(text)) + false + } catch { + UniGeneError(_) => true + } + assert_true(raised) +} diff --git a/test/moonbit/uniprot_io_test.mbt b/test/moonbit/uniprot_io_test.mbt index 64c57e4f..451d3ef0 100644 --- a/test/moonbit/uniprot_io_test.mbt +++ b/test/moonbit/uniprot_io_test.mbt @@ -95,7 +95,13 @@ test "parse_uniprot_xml_with_dbreferences" { ///| test "uniprot_to_seqrecord" { - let entry = @src.UniprotEntry::new("P12345", "Test Protein", ["GENE1"], "Homo sapiens", "ACGTACGTAC") + let entry = @src.UniprotEntry::new( + "P12345", + "Test Protein", + ["GENE1"], + "Homo sapiens", + "ACGTACGTAC", + ) let (header, id, desc, kwargs) = @src.uniprot_to_seqrecord(entry) assert_true(header.length() > 0) assert_true(id.length() > 0) @@ -104,7 +110,13 @@ test "uniprot_to_seqrecord" { ///| test "uniprot_entry_methods" { - let entry = @src.UniprotEntry::new("P12345", "Test Protein", ["GENE1", "GENE2"], "Homo sapiens", "ACGTACGTAC") + let entry = @src.UniprotEntry::new( + "P12345", + "Test Protein", + ["GENE1", "GENE2"], + "Homo sapiens", + "ACGTACGTAC", + ) assert_eq(entry.accession, "P12345") assert_eq(entry.sequence, "ACGTACGTAC") assert_eq(entry.seq_length, 10) diff --git a/test/moonbit/universalmotif_test.mbt b/test/moonbit/universalmotif_test.mbt index 0f127845..a9667374 100644 --- a/test/moonbit/universalmotif_test.mbt +++ b/test/moonbit/universalmotif_test.mbt @@ -6,9 +6,9 @@ test "universalmotif_motif_from_pwm" { [0.0, 0.0, 0.9, 0.1], [0.1, 0.9, 0.0, 0.0], ] - + let motif = @src.S4Motif::from_pwm("TATA-box", "DNA", pwm) - + assert_eq(motif.name, "TATA-box") assert_eq(motif.alphabet, "DNA") assert_eq(motif.pwm.length(), 4) @@ -16,21 +16,17 @@ test "universalmotif_motif_from_pwm" { ///| test "universalmotif_calculate_consensus" { - let pwm = [ - [0.9, 0.1, 0.0, 0.0], - [0.0, 0.0, 0.9, 0.1], - [0.1, 0.9, 0.0, 0.0], - ] - + let pwm = [[0.9, 0.1, 0.0, 0.0], [0.0, 0.0, 0.9, 0.1], [0.1, 0.9, 0.0, 0.0]] + let motif = @src.S4Motif::from_pwm("test", "DNA", pwm) - + assert_eq(motif.consensus.length(), 3) } ///| test "universalmotif_create_example_motif" { let motif = @src.create_example_s4motif() - + assert_eq(motif.name, "TATA-box") assert_eq(motif.pwm.length(), 5) -} \ No newline at end of file +} diff --git a/test/moonbit/uwot_test.mbt b/test/moonbit/uwot_test.mbt index 61a2f432..61700c74 100644 --- a/test/moonbit/uwot_test.mbt +++ b/test/moonbit/uwot_test.mbt @@ -6,10 +6,7 @@ // ============================================================ test "uwot_distance_matrix square" { - let data : Array[Array[Double]] = [ - [0.0, 0.0], - [3.0, 4.0], - ] + let data : Array[Array[Double]] = [[0.0, 0.0], [3.0, 4.0]] let dist = @src.uwot_distance_matrix(data) assert_eq(dist.length(), 2) assert_eq(dist[0].length(), 2) @@ -20,6 +17,7 @@ test "uwot_distance_matrix square" { assert_eq(dist[1][0], 5.0) } +///| test "uwot_distance_matrix empty" { let data : Array[Array[Double]] = [] let dist = @src.uwot_distance_matrix(data) @@ -30,6 +28,7 @@ test "uwot_distance_matrix empty" { // UMAP config tests // ============================================================ +///| test "umap_config default" { let config = @src.UmapConfig::new() assert_eq(config.n_neighbors, 15) @@ -44,19 +43,21 @@ test "umap_config default" { // UMAP algorithm tests // ============================================================ +///| test "umap basic 2d" { let data = @src.create_umap_test_data(10, 5) assert_eq(data.length(), 10) assert_eq(data[0].length(), 5) - + let config = @src.UmapConfig::new() - + let result = @src.uwot_umap(data, config) assert_eq(result.embedding.length(), 10) assert_eq(result.embedding[0].length(), 2) assert_eq(result.n_epochs, 200) } +///| test "umap empty data" { let data : Array[Array[Double]] = [] let config = @src.UmapConfig::new() @@ -64,6 +65,7 @@ test "umap empty data" { assert_eq(result.embedding.length(), 0) } +///| test "umap single sample" { let data : Array[Array[Double]] = [[1.0, 2.0, 3.0]] let config = @src.UmapConfig::new() @@ -71,28 +73,30 @@ test "umap single sample" { assert_true(result.embedding.length() <= 1) } +///| test "umap different n_components" { let data = @src.create_umap_test_data(10, 5) let custom_config = @src.UmapConfig::new_custom(3, 3, 50, 42) - + let result = @src.uwot_umap(data, custom_config) assert_eq(result.embedding[0].length(), 3) } +///| test "umap produces finite values" { let data = @src.create_umap_test_data(12, 4) let custom_config = @src.UmapConfig::new_custom(3, 2, 50, 42) - + let result = @src.uwot_umap(data, custom_config) - + // Check that all embedding values are finite let mut i = 0 while i < result.embedding.length() { let mut j = 0 while j < result.embedding[i].length() { let val = result.embedding[i][j] - assert_true(val == val) // NaN check - assert_true(val < 1.0e300) // infinity check + assert_true(val == val) // NaN check + assert_true(val < 1.0e300) // infinity check assert_true(val > -1.0e300) j = j + 1 } @@ -104,12 +108,14 @@ test "umap produces finite values" { // Test data generation tests // ============================================================ +///| test "create_umap_test_data dimensions" { let data = @src.create_umap_test_data(20, 8) assert_eq(data.length(), 20) assert_eq(data[0].length(), 8) } +///| test "create_umap_test_data clusters" { let data = @src.create_umap_test_data(8, 3) // 4 clusters, each should have similar values @@ -117,7 +123,7 @@ test "create_umap_test_data clusters" { // Cluster 1: samples 1, 5 // Cluster 2: samples 2, 6 // Cluster 3: samples 3, 7 - + // Check that cluster 0 is different from cluster 3 let mut sum0 = 0.0 let mut j = 0 @@ -125,14 +131,14 @@ test "create_umap_test_data clusters" { sum0 = sum0 + data[0][j] j = j + 1 } - + let mut sum3 = 0.0 j = 0 while j < 3 { sum3 = sum3 + data[3][j] j = j + 1 } - + // Clusters should be separated assert_true(sum3 > sum0) } diff --git a/test/moonbit/variance_partition_test.mbt b/test/moonbit/variance_partition_test.mbt new file mode 100644 index 00000000..6ebd8946 --- /dev/null +++ b/test/moonbit/variance_partition_test.mbt @@ -0,0 +1,855 @@ +///| +/// Tests for Bioconductor variancePartition-inspired mixed models. + +///| +fn vp_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "variancePartition value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn vp_test_samples() -> Array[String] { + let samples : Array[String] = [] + for subject in 0..<4 { + for replicate in 0..<4 { + samples.push( + "S" + (subject + 1).to_string() + "_R" + (replicate + 1).to_string(), + ) + } + } + samples +} + +///| +fn vp_test_subjects() -> Array[String] { + let subjects : Array[String] = [] + for subject in 0..<4 { + for _ in 0..<4 { + subjects.push("S" + (subject + 1).to_string()) + } + } + subjects +} + +///| +fn vp_test_treatments() -> Array[String] { + let treatments : Array[String] = [] + for _ in 0..<4 { + treatments.push("Control") + treatments.push("Treated") + treatments.push("Control") + treatments.push("Treated") + } + treatments +} + +///| +fn vp_test_batches() -> Array[String] { + let batches : Array[String] = [] + for _ in 0..<4 { + batches.push("B1") + batches.push("B1") + batches.push("B2") + batches.push("B2") + } + batches +} + +///| +fn vp_test_design( + include_batch? : Bool = false, +) -> @src.VariancePartitionDesign { + let treatment = @src.vp_categorical_effect( + "Treatment", + vp_test_treatments(), + reference="Control", + ) catch { + _ => abort("valid treatment effect should build") + } + let subject = @src.vp_random_effect("Subject", vp_test_subjects()) catch { + _ => abort("valid subject effect should build") + } + let random_effects = [subject] + if include_batch { + let batch = @src.vp_random_effect("Batch", vp_test_batches()) catch { + _ => abort("valid batch effect should build") + } + random_effects.push(batch) + } + @src.vp_design(vp_test_samples(), [treatment], random_effects) catch { + _ => abort("valid variancePartition design should build") + } +} + +///| +fn vp_test_subject_gene() -> Array[Double] { + let values : Array[Double] = [] + let noise = [0.08, -0.08, 0.04, -0.04] + for subject in 0..<4 { + let subject_effect = (subject.to_double() - 1.5) * 2.0 + for replicate in 0..<4 { + let treatment_effect = if replicate % 2 == 1 { 0.1 } else { 0.0 } + values.push(10.0 + subject_effect + treatment_effect + noise[replicate]) + } + } + values +} + +///| +fn vp_test_treatment_gene() -> Array[Double] { + let values : Array[Double] = [] + let noise = [0.05, -0.03, -0.05, 0.03] + for subject in 0..<4 { + let subject_effect = (subject.to_double() - 1.5) * 0.08 + for replicate in 0..<4 { + let treatment_effect = if replicate % 2 == 1 { 4.0 } else { 0.0 } + values.push(6.0 + subject_effect + treatment_effect + noise[replicate]) + } + } + values +} + +///| +fn vp_test_batch_gene() -> Array[Double] { + let values : Array[Double] = [] + let noise = [0.03, -0.03, 0.02, -0.02] + for subject in 0..<4 { + let subject_effect = (subject.to_double() - 1.5) * 0.1 + for replicate in 0..<4 { + let batch_effect = if replicate >= 2 { 3.0 } else { 0.0 } + values.push(8.0 + subject_effect + batch_effect + noise[replicate]) + } + } + values +} + +///| +fn vp_test_residual_gene() -> Array[Double] { + let pattern = [1.0, -1.0, -1.0, 1.0] + let values : Array[Double] = [] + for subject in 0..<4 { + let scale = 1.0 + subject.to_double() * 0.1 + for replicate in 0..<4 { + values.push(5.0 + scale * pattern[replicate]) + } + } + values +} + +///| +fn vp_test_expression() -> Array[Array[Double]] { + [vp_test_subject_gene(), vp_test_treatment_gene(), vp_test_residual_gene()] +} + +///| +fn vp_test_assay() -> @src.SummarizedExperiment { + let expression = vp_test_expression() + let weights : Array[Array[Double]] = [] + for _ in 0.. abort("valid numeric effect should build") + } + assert_eq(effect.name, "Age") + assert_eq(effect.column_names, ["Age"]) + assert_eq(effect.columns, [[20.0], [30.0], [40.0]]) +} + +///| +test "variancePartition: categorical effect uses requested reference" { + let effect = @src.vp_categorical_effect( + "Disease", + ["Case", "Control", "Case", "Other"], + reference="Control", + ) catch { + _ => abort("valid categorical effect should build") + } + assert_eq(effect.column_names, ["Disease:Case", "Disease:Other"]) + assert_eq(effect.columns[0], [1.0, 0.0]) + assert_eq(effect.columns[1], [0.0, 0.0]) + assert_eq(effect.columns[3], [0.0, 1.0]) +} + +///| +test "variancePartition: random effect encoding is stable" { + let effect = @src.vp_random_effect("Subject", ["B", "A", "B", "A"]) catch { + _ => abort("valid random effect should build") + } + assert_eq(effect.levels, ["B", "A"]) + assert_eq(effect.level_indices, [0, 1, 0, 1]) +} + +///| +test "variancePartition: design expands intercept and treatment" { + let design = vp_test_design() + assert_eq(design.sample_names.length(), 16) + assert_eq(design.coefficient_names, ["(Intercept)", "Treatment:Treated"]) + assert_eq(design.matrix[0], [1.0, 0.0]) + assert_eq(design.matrix[1], [1.0, 1.0]) + assert_eq(design.random_effects[0].levels.length(), 4) +} + +///| +test "variancePartition: constructors copy caller arrays" { + let ages = [20.0, 30.0, 40.0] + let effect = @src.vp_numeric_effect("Age", ages) catch { + _ => abort("valid numeric effect should build") + } + ages[0] = 99.0 + assert_eq(effect.columns[0][0], 20.0) + let samples = vp_test_samples() + let design = vp_test_design() + samples[0] = "changed" + assert_eq(design.sample_names[0], "S1_R1") +} + +///| +test "variancePartition: subject random variance dominates subject gene" { + let fit = @src.vp_fit_gene( + vp_test_subject_gene(), + vp_test_design(), + gene_name="subject_gene", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("subject mixed model should fit") + } + assert_true(fit.converged) + assert_true(fit.random_variances[0] > fit.residual_variance) + let fraction = fit.fraction("Subject") + match fraction { + Some(value) => assert_true(value > 0.8) + None => abort("Subject fraction should be present") + } +} + +///| +test "variancePartition: fixed variance dominates treatment gene" { + let fit = @src.vp_fit_gene( + vp_test_treatment_gene(), + vp_test_design(), + gene_name="treatment_gene", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("treatment mixed model should fit") + } + vp_test_close(fit.coefficients[1], 4.0, 0.15) + match fit.fraction("Treatment") { + Some(value) => assert_true(value > 0.8) + None => abort("Treatment fraction should be present") + } +} + +///| +test "variancePartition: variance fractions sum to one" { + let fit = @src.vp_fit_gene( + vp_test_subject_gene(), + vp_test_design(), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("mixed model should fit") + } + let mut total = 0.0 + for value in fit.variance_fractions { + assert_true(value >= 0.0) + total = total + value + } + vp_test_close(total, 1.0, 1.0e-10) +} + +///| +test "variancePartition: BLUPs recover subject ordering" { + let fit = @src.vp_fit_gene( + vp_test_subject_gene(), + vp_test_design(), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("mixed model should fit") + } + let first = match fit.blup("Subject", "S1") { + Some(value) => value + None => abort("S1 BLUP should be present") + } + let last = match fit.blup("Subject", "S4") { + Some(value) => value + None => abort("S4 BLUP should be present") + } + assert_true(first < 0.0) + assert_true(last > 0.0) + assert_true(last > first) +} + +///| +test "variancePartition: fitted values and residuals preserve sample order" { + let expression = vp_test_treatment_gene() + let fit = @src.vp_fit_gene( + expression, + vp_test_design(), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("mixed model should fit") + } + assert_eq(fit.fitted_values.length(), expression.length()) + assert_eq(fit.residuals.length(), expression.length()) + for index in 0.. abort("unweighted model should fit") + } + let weighted = @src.vp_fit_gene( + expression, + vp_test_design(), + weights=Array::make(expression.length(), 5.0), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("weighted model should fit") + } + vp_test_close(unweighted.coefficients[1], weighted.coefficients[1], 1.0e-8) + vp_test_close( + unweighted.residual_variance, + weighted.residual_variance, + 1.0e-8, + ) +} + +///| +test "variancePartition: heterogeneous precision weights are accepted" { + let expression = vp_test_treatment_gene() + let weights = Array::make(expression.length(), 1.0) + weights[0] = 0.2 + let fit = @src.vp_fit_gene( + expression, + vp_test_design(), + weights~, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("heterogeneous weighted model should fit") + } + assert_true(fit.residual_variance > 0.0) + assert_true(fit.log_likelihood < 1.0e300) +} + +///| +test "variancePartition: ML and REML fits are identified" { + let reml_fit = @src.vp_fit_gene( + vp_test_subject_gene(), + vp_test_design(), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("REML model should fit") + } + let ml_fit = @src.vp_fit_gene( + vp_test_subject_gene(), + vp_test_design(), + reml=false, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("ML model should fit") + } + assert_true(reml_fit.reml) + assert_false(ml_fit.reml) + assert_true(reml_fit.log_likelihood != ml_fit.log_likelihood) +} + +///| +test "variancePartition: multiple random intercepts are fitted" { + let fit = @src.vp_fit_gene( + vp_test_batch_gene(), + vp_test_design(include_batch=true), + gene_name="batch_gene", + max_iterations=100, + tolerance=0.01, + ) catch { + _ => abort("multi-random-effect model should fit") + } + assert_eq(fit.random_effect_names, ["Subject", "Batch"]) + assert_eq(fit.random_variances.length(), 2) + assert_true(fit.random_variances[1] > fit.random_variances[0]) + match fit.fraction("Batch") { + Some(value) => assert_true(value > 0.5) + None => abort("Batch fraction should be present") + } +} + +///| +test "variancePartition: fixed-only design reduces to linear model" { + let treatment = @src.vp_categorical_effect( + "Treatment", + vp_test_treatments(), + reference="Control", + ) catch { + _ => abort("valid fixed effect should build") + } + let design = @src.vp_design(vp_test_samples(), [treatment], []) catch { + _ => abort("fixed-only design should build") + } + let fit = @src.vp_fit_gene( + vp_test_treatment_gene(), + design, + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("fixed-only model should fit") + } + vp_test_close(fit.coefficients[1], 4.0, 0.15) + assert_eq(fit.random_variances.length(), 0) +} + +///| +test "variancePartition: matrix fit generates gene names" { + let result = @src.fit_extract_variance_partition( + vp_test_expression(), + vp_test_design(), + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("variance partition matrix should fit") + } + assert_eq(result.gene_names, ["Gene1", "Gene2", "Gene3"]) + assert_eq(result.fractions.length(), 3) + assert_eq(result.fits.length(), 3) + assert_false(result.fits[0].reml) +} + +///| +test "variancePartition: matrix result supports named lookup" { + let result = @src.fit_extract_variance_partition( + vp_test_expression(), + vp_test_design(), + gene_names=["subject", "treatment", "residual"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("variance partition matrix should fit") + } + match result.fraction("subject", "Subject") { + Some(value) => assert_true(value > 0.8) + None => abort("named variance fraction should be present") + } + assert_true(result.fraction("missing", "Subject") is None) +} + +///| +test "variancePartition: matrix result reports component medians" { + let result = @src.fit_extract_variance_partition( + vp_test_expression(), + vp_test_design(), + gene_names=["subject", "treatment", "residual"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("variance partition matrix should fit") + } + assert_eq(result.component_names, ["Treatment", "Subject", "Residuals"]) + assert_eq(result.median_fractions.length(), 3) + for value in result.median_fractions { + assert_true(value >= 0.0 && value <= 1.0) + } +} + +///| +test "variancePartition: result summary contains component names" { + let result = @src.fit_extract_variance_partition( + [vp_test_subject_gene()], + vp_test_design(), + gene_names=["subject"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("variance partition matrix should fit") + } + let summary = result.summary() + assert_true(summary.contains("variancePartition Summary")) + assert_true(summary.contains("Subject median fraction")) + assert_true(summary.contains("Genes: 1")) +} + +///| +test "variancePartition: dream estimates repeated-measures contrast" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + vp_test_expression(), + design, + contrast, + gene_names=["subject", "treatment", "residual"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + vp_test_close(result.genes[1].estimate, 4.0, 0.15) + assert_true(result.genes[1].standard_error > 0.0) + assert_true(result.genes[1].p_value < 0.001) +} + +///| +test "variancePartition: dream uses bounded small-sample degrees of freedom" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + [vp_test_treatment_gene()], + design, + contrast, + gene_names=["treatment"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + let degrees = result.genes[0].degrees_of_freedom + assert_true(degrees >= 1.0) + assert_true(degrees <= 14.0) +} + +///| +test "variancePartition: dream applies monotone BH correction" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + vp_test_expression(), + design, + contrast, + gene_names=["subject", "treatment", "residual"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + for gene in result.genes { + assert_true(gene.adjusted_p_value >= gene.p_value) + assert_true(gene.adjusted_p_value <= 1.0) + } +} + +///| +test "variancePartition: dream top table sorts by p-value" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + vp_test_expression(), + design, + contrast, + gene_names=["subject", "treatment", "residual"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + let table = result.top_table() catch { + _ => abort("valid top table should build") + } + assert_eq(table.length(), 3) + for index in 1.. abort("valid custom contrast should build") + } + assert_eq(contrast.name, "Treatment") + assert_eq(contrast.coefficients, [0.0, 1.0]) +} + +///| +test "variancePartition: dream summary reports contrast and significance" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + [vp_test_treatment_gene()], + design, + contrast, + gene_names=["treatment"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + let summary = result.summary() + assert_true(summary.contains("dream Differential Expression Summary")) + assert_true(summary.contains("Contrast: Treated-Control")) + assert_true(summary.contains("FDR <= 0.05: 1")) +} + +///| +test "variancePartition: SummarizedExperiment variance entry uses assays" { + let result = @src.fit_extract_variance_partition_se( + vp_test_assay(), + "logcounts", + vp_test_design(), + gene_names=["subject", "treatment", "residual"], + weights_assay="weights", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("SummarizedExperiment variance analysis should fit") + } + assert_eq(result.gene_names.length(), 3) + match result.fraction("subject", "Subject") { + Some(value) => assert_true(value > 0.8) + None => abort("Subject fraction should be present") + } +} + +///| +test "variancePartition: SummarizedExperiment dream entry uses assays" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream_se( + vp_test_assay(), + "logcounts", + design, + contrast, + gene_names=["subject", "treatment", "residual"], + weights_assay="weights", + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("SummarizedExperiment dream analysis should fit") + } + assert_true(result.genes[1].p_value < 0.001) +} + +///| +test "variancePartition: rejects categorical reference that is absent" { + let failed = try { + ignore( + @src.vp_categorical_effect( + "Disease", + ["Case", "Control"], + reference="Other", + ), + ) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects unreplicated random effects" { + let failed = try { + ignore(@src.vp_random_effect("Subject", ["A", "B", "C"])) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects duplicate sample names" { + let age = @src.vp_numeric_effect("Age", [1.0, 2.0, 3.0]) catch { + _ => abort("numeric effect should build") + } + let failed = try { + ignore(@src.vp_design(["A", "A", "B"], [age], [])) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects rank-deficient fixed design" { + let first = @src.vp_numeric_effect("First", [0.0, 1.0, 2.0, 3.0]) catch { + _ => abort("numeric effect should build") + } + let second = @src.vp_numeric_effect("Second", [0.0, 2.0, 4.0, 6.0]) catch { + _ => abort("numeric effect should build") + } + let failed = try { + ignore(@src.vp_design(["A", "B", "C", "D"], [first, second], [])) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects expression dimension mismatch" { + let failed = try { + ignore(@src.vp_fit_gene([1.0, 2.0], vp_test_design())) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects nonpositive precision weights" { + let weights = Array::make(16, 1.0) + weights[3] = 0.0 + let failed = try { + ignore(@src.vp_fit_gene(vp_test_subject_gene(), vp_test_design(), weights~)) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects duplicate gene names" { + let failed = try { + ignore( + @src.fit_extract_variance_partition( + [vp_test_subject_gene(), vp_test_treatment_gene()], + vp_test_design(), + gene_names=["gene", "gene"], + ), + ) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects zero contrast" { + let failed = try { + ignore(@src.vp_contrast("zero", [0.0, 0.0])) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects contrast dimension mismatch" { + let failed = try { + ignore( + @src.dream( + [vp_test_treatment_gene()], + vp_test_design(), + @src.vp_contrast("bad", [1.0]) catch { + _ => abort("nonzero contrast should build") + }, + ), + ) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: rejects invalid top-table FDR" { + let design = vp_test_design() + let contrast = @src.vp_coefficient_contrast( + design, "Treated-Control", "Treatment:Treated", + ) catch { + _ => abort("valid coefficient contrast should build") + } + let result = @src.dream( + [vp_test_treatment_gene()], + design, + contrast, + gene_names=["treatment"], + max_iterations=80, + tolerance=0.01, + ) catch { + _ => abort("dream analysis should fit") + } + let failed = try { + ignore(result.top_table(maximum_fdr=1.5)) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} + +///| +test "variancePartition: reports missing SummarizedExperiment assay" { + let failed = try { + ignore( + @src.fit_extract_variance_partition_se( + vp_test_assay(), + "missing", + vp_test_design(), + ), + ) + false + } catch { + VariancePartitionError(_) => true + } + assert_true(failed) +} diff --git a/test/moonbit/variant_filtering_test.mbt b/test/moonbit/variant_filtering_test.mbt index 2b0e86ec..9f10d71d 100644 --- a/test/moonbit/variant_filtering_test.mbt +++ b/test/moonbit/variant_filtering_test.mbt @@ -1,13 +1,12 @@ ///| - test "variant_filtering_create_variant" { let info = Map([("DP", "20"), ("AF", "0.5")]) let genotypes = [("sample1", "0/1:30"), ("sample2", "1/1:25")] - + let variant = @src.Variant::new( - "chr1", 1000, "rs123", "A", "T", 50.0, "PASS", info, genotypes + "chr1", 1000, "rs123", "A", "T", 50.0, "PASS", info, genotypes, ) - + assert_eq(variant.chr, "chr1") assert_eq(variant.pos, 1000) assert_eq(variant.id, "rs123") @@ -18,23 +17,31 @@ test "variant_filtering_create_variant" { assert_eq(variant.genotypes.length(), 2) } +///| test "variant_filtering_filter" { let info_pass = Map([("DP", "20"), ("AF", "0.5")]) let info_fail = Map([("DP", "5"), ("AF", "0.001")]) - + let variants = [ - @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info_pass, [("s1", "0/1:30")]), - @src.Variant::new("chr1", 2000, "rs2", "G", "C", 20.0, "FAIL", info_fail, [("s1", "0/1:10")]), - @src.Variant::new("chr1", 3000, "rs3", "T", "A", 60.0, "PASS", info_pass, [("s1", "0/1:35")]), + @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info_pass, [ + ("s1", "0/1:30"), + ]), + @src.Variant::new("chr1", 2000, "rs2", "G", "C", 20.0, "FAIL", info_fail, [ + ("s1", "0/1:10"), + ]), + @src.Variant::new("chr1", 3000, "rs3", "T", "A", 60.0, "PASS", info_pass, [ + ("s1", "0/1:35"), + ]), ] - + let params = @src.VariantFilteringParam::new() let result = @src.bio_vfilter_filter(variants, params) - + assert_eq(result.passed_variants.length(), 2) assert_eq(result.filtered_variants.length(), 1) } +///| test "variant_filtering_check_autosomal_dominant" { let genotypes = [ ("proband", "1/1"), @@ -42,18 +49,26 @@ test "variant_filtering_check_autosomal_dominant" { ("mother", "0/0"), ("sibling", "0/0"), ] - + let info = Map([], capacity=0) - let variant = @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes) - - let phenotypes = Map([("proband", true), ("father", false), ("mother", false), ("sibling", false)]) + let variant = @src.Variant::new( + "chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes, + ) + + let phenotypes = Map([ + ("proband", true), + ("father", false), + ("mother", false), + ("sibling", false), + ]) let model = @src.GeneticModel::new("autosomal_dominant") - + let passes = @src.bio_vfilter_check_inheritance(variant, model, phenotypes) - + assert_true(passes) } +///| test "variant_filtering_check_autosomal_recessive" { let genotypes = [ ("proband", "1/1"), @@ -61,46 +76,65 @@ test "variant_filtering_check_autosomal_recessive" { ("mother", "0/1"), ("sibling", "0/0"), ] - + let info = Map([], capacity=0) - let variant = @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes) - - let phenotypes = Map([("proband", true), ("father", false), ("mother", false), ("sibling", false)]) + let variant = @src.Variant::new( + "chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes, + ) + + let phenotypes = Map([ + ("proband", true), + ("father", false), + ("mother", false), + ("sibling", false), + ]) let model = @src.GeneticModel::new("autosomal_recessive") - + let passes = @src.bio_vfilter_check_inheritance(variant, model, phenotypes) - + assert_true(passes) } +///| test "variant_filtering_summary" { let info_pass = Map([("DP", "20"), ("AF", "0.5")]) let info_fail = Map([("DP", "5"), ("AF", "0.001")]) - + let variants = [ - @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info_pass, [("s1", "0/1:30")]), - @src.Variant::new("chr1", 2000, "rs2", "G", "C", 20.0, "FAIL", info_fail, [("s1", "0/1:10")]), + @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info_pass, [ + ("s1", "0/1:30"), + ]), + @src.Variant::new("chr1", 2000, "rs2", "G", "C", 20.0, "FAIL", info_fail, [ + ("s1", "0/1:10"), + ]), ] - + let params = @src.VariantFilteringParam::new() let result = @src.bio_vfilter_filter(variants, params) - + let summary = @src.bio_vfilter_summary(result) - + assert_true(summary.contains("Total variants")) assert_true(summary.contains("Passed")) assert_true(summary.contains("Filtered")) } +///| test "variant_filtering_predict_consequence" { let info = Map([], capacity=0) let genotypes = [("s1", "0/1")] - - let snp = @src.Variant::new("chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes) - let insertion = @src.Variant::new("chr1", 2000, "rs2", "A", "AT", 50.0, "PASS", info, genotypes) - let deletion = @src.Variant::new("chr1", 3000, "rs3", "AT", "A", 50.0, "PASS", info, genotypes) - + + let snp = @src.Variant::new( + "chr1", 1000, "rs1", "A", "T", 50.0, "PASS", info, genotypes, + ) + let insertion = @src.Variant::new( + "chr1", 2000, "rs2", "A", "AT", 50.0, "PASS", info, genotypes, + ) + let deletion = @src.Variant::new( + "chr1", 3000, "rs3", "AT", "A", 50.0, "PASS", info, genotypes, + ) + assert_eq(@src.bio_vfilter_predict_consequence(snp), "missense_variant") assert_eq(@src.bio_vfilter_predict_consequence(insertion), "insertion") assert_eq(@src.bio_vfilter_predict_consequence(deletion), "deletion") -} \ No newline at end of file +} diff --git a/test/moonbit/variation_test.mbt b/test/moonbit/variation_test.mbt index becf30ee..231203f1 100644 --- a/test/moonbit/variation_test.mbt +++ b/test/moonbit/variation_test.mbt @@ -37,4 +37,4 @@ test "variation_parse_vcf_line" { test "variation_create_example_data" { let records = @src.create_example_variation_data() assert_eq(records.length(), 2) -} \ No newline at end of file +} diff --git a/test/moonbit/velociraptor_test.mbt b/test/moonbit/velociraptor_test.mbt index aa86f0d3..77fd7586 100644 --- a/test/moonbit/velociraptor_test.mbt +++ b/test/moonbit/velociraptor_test.mbt @@ -108,7 +108,12 @@ test "vr_kinetic_model_returns_params" { test "vr_kinetic_model_sample_data" { let gene_data = @src.velocity_sample_data() for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) assert_true(k.beta() > 0.0) assert_true(k.gamma() > 0.0) } @@ -132,7 +137,12 @@ test "vr_compute_returns_velocities" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) @@ -162,7 +172,12 @@ test "vr_velocity_embedding" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) @@ -191,7 +206,12 @@ test "vr_find_root_cells" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) @@ -209,14 +229,21 @@ test "vr_root_cells_sorted_by_speed" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) let roots = @src.velocity_find_root_cells(velocities, 5) // Speeds should be sorted descending. for i in 1..= velocities[roots[i]].speed()) + assert_true( + velocities[roots[i - 1]].speed() >= velocities[roots[i]].speed(), + ) } } @@ -229,7 +256,12 @@ test "vr_transition_matrix" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) @@ -255,7 +287,12 @@ test "vr_transition_matrix_nonneg" { let gene_data = @src.velocity_sample_data() let kinetics : Array[@src.GeneKinetics] = [] for gd in gene_data { - let k = @src.velocity_kinetic_model(gd.gene_name(), gd.spliced(), gd.unspliced(), 30) + let k = @src.velocity_kinetic_model( + gd.gene_name(), + gd.spliced(), + gd.unspliced(), + 30, + ) kinetics.push(k) } let velocities = @src.velocity_compute(gene_data, kinetics) @@ -276,7 +313,9 @@ test "vr_transition_matrix_nonneg" { test "vr_run_pipeline" { let gene_data = @src.velocity_sample_data() let embeddings = @src.velocity_sample_embeddings(20) - let (kinetics, velocities, emb_vel) = @src.velocity_run(gene_data, embeddings, 5) + let (kinetics, velocities, emb_vel) = @src.velocity_run( + gene_data, embeddings, 5, + ) assert_eq(kinetics.length(), 4) // 4 genes assert_eq(velocities.length(), 20) // 20 cells assert_eq(emb_vel.length(), 20) diff --git a/test/moonbit/venn_diagram_test.mbt b/test/moonbit/venn_diagram_test.mbt index 0e557441..79acce5f 100644 --- a/test/moonbit/venn_diagram_test.mbt +++ b/test/moonbit/venn_diagram_test.mbt @@ -3,10 +3,9 @@ // 1. venn_diagram - basic Venn diagram with 2 sets test "venn_diagram_two_sets_basic" { - let result = @src.venn_diagram( - [["A", "B", "C", "D"], ["C", "D", "E", "F"]], - names=["Set1", "Set2"], - ) + let result = @src.venn_diagram([["A", "B", "C", "D"], ["C", "D", "E", "F"]], names=[ + "Set1", "Set2", + ]) assert_eq(result.n_sets, 2) assert_eq(result.set_names, ["Set1", "Set2"]) assert_eq(result.set_sizes, [4, 4]) @@ -15,6 +14,7 @@ test "venn_diagram_two_sets_basic" { assert_eq(result.regions.length(), 4) // 2^2 = 4 regions } +///| test "venn_diagram_two_sets_default_names" { let result = @src.venn_diagram([["A", "B"], ["B", "C"]]) assert_eq(result.set_names, ["Set1", "Set2"]) @@ -23,6 +23,8 @@ test "venn_diagram_two_sets_default_names" { } // 2. venn_diagram - basic Venn diagram with 3 sets + +///| test "venn_diagram_three_sets_basic" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -35,11 +37,11 @@ test "venn_diagram_three_sets_basic" { assert_eq(result.regions.length(), 8) // 2^3 = 8 regions } +///| test "venn_diagram_three_sets_no_overlap" { - let result = @src.venn_diagram( - [["A", "B"], ["C", "D"], ["E", "F"]], - names=["S1", "S2", "S3"], - ) + let result = @src.venn_diagram([["A", "B"], ["C", "D"], ["E", "F"]], names=[ + "S1", "S2", "S3", + ]) assert_eq(result.total_elements, 6) assert_eq(result.regions[1].count, 2) // only in S1 (0b001) assert_eq(result.regions[2].count, 2) // only in S2 (0b010) @@ -48,47 +50,47 @@ test "venn_diagram_three_sets_no_overlap" { } // 3. venn_get_region - get elements in specific region by bitmask + +///| test "venn_get_region_only_in_first_set" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) let only_s1 = @src.venn_get_region(result, 1) // 0b01 assert_eq(only_s1.length(), 2) // A, B assert_true(only_s1.contains("A")) assert_true(only_s1.contains("B")) } +///| test "venn_get_region_only_in_second_set" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) let only_s2 = @src.venn_get_region(result, 2) // 0b10 assert_eq(only_s2.length(), 2) // D, E assert_true(only_s2.contains("D")) assert_true(only_s2.contains("E")) } +///| test "venn_get_region_intersection_two_sets" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) let inter = @src.venn_get_region(result, 3) // 0b11 assert_eq(inter.length(), 1) // C assert_true(inter.contains("C")) } +///| test "venn_get_region_outside_all" { - let result = @src.venn_diagram( - [["A", "B"], ["C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B"], ["C", "D"]], names=["S1", "S2"]) let outside = @src.venn_get_region(result, 0) assert_eq(outside.length(), 0) } +///| test "venn_get_region_invalid_id" { let result = @src.venn_diagram([["A"], ["B"]], names=["S1", "S2"]) let invalid = @src.venn_get_region(result, 10) @@ -96,28 +98,30 @@ test "venn_get_region_invalid_id" { } // 4. venn_only_in_set - elements only in one set + +///| test "venn_only_in_set_first" { - let result = @src.venn_diagram( - [["A", "B", "C", "D"], ["C", "D", "E", "F"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C", "D"], ["C", "D", "E", "F"]], names=[ + "S1", "S2", + ]) let only_first = @src.venn_only_in_set(result, 0) assert_eq(only_first.length(), 2) // A, B assert_true(only_first.contains("A")) assert_true(only_first.contains("B")) } +///| test "venn_only_in_set_second" { - let result = @src.venn_diagram( - [["A", "B", "C", "D"], ["C", "D", "E", "F"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C", "D"], ["C", "D", "E", "F"]], names=[ + "S1", "S2", + ]) let only_second = @src.venn_only_in_set(result, 1) assert_eq(only_second.length(), 2) // E, F assert_true(only_second.contains("E")) assert_true(only_second.contains("F")) } +///| test "venn_only_in_set_three_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -132,6 +136,8 @@ test "venn_only_in_set_three_sets" { } // 5. venn_intersection_only - elements in specific sets only + +///| test "venn_intersection_only_two_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -142,6 +148,7 @@ test "venn_intersection_only_two_sets" { assert_true(s1_s2_only.contains("B")) } +///| test "venn_intersection_only_all_three" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -152,27 +159,29 @@ test "venn_intersection_only_all_three" { assert_true(all_three.contains("C")) } +///| test "venn_intersection_only_single_set" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["B", "C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["B", "C", "D"]], names=[ + "S1", "S2", + ]) let just_first = @src.venn_intersection_only(result, [0]) assert_eq(just_first.length(), 1) // A } // 6. venn_all_intersect - elements common to all sets + +///| test "venn_all_intersect_two_sets" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["B", "C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["B", "C", "D"]], names=[ + "S1", "S2", + ]) let all = @src.venn_all_intersect(result) assert_eq(all.length(), 2) // B, C assert_true(all.contains("B")) assert_true(all.contains("C")) } +///| test "venn_all_intersect_three_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -183,21 +192,22 @@ test "venn_all_intersect_three_sets" { assert_true(all.contains("C")) } +///| test "venn_all_intersect_no_common" { - let result = @src.venn_diagram( - [["A", "B"], ["C", "D"], ["E", "F"]], - names=["S1", "S2", "S3"], - ) + let result = @src.venn_diagram([["A", "B"], ["C", "D"], ["E", "F"]], names=[ + "S1", "S2", "S3", + ]) let all = @src.venn_all_intersect(result) assert_eq(all.length(), 0) } // 7. venn_unique_to_each - unique elements per set + +///| test "venn_unique_to_each_two_sets" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) let unique = @src.venn_unique_to_each(result) assert_eq(unique.length(), 2) assert_eq(unique[0].length(), 2) // A, B @@ -208,6 +218,7 @@ test "venn_unique_to_each_two_sets" { assert_true(unique[1].contains("E")) } +///| test "venn_unique_to_each_three_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -221,11 +232,12 @@ test "venn_unique_to_each_three_sets" { } // 8. venn_pairwise_overlap - pairwise overlap statistics + +///| test "venn_pairwise_overlap_two_sets" { - let result = @src.venn_diagram( - [["A", "B", "C", "D"], ["C", "D", "E", "F"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C", "D"], ["C", "D", "E", "F"]], names=[ + "S1", "S2", + ]) let stats = @src.venn_pairwise_overlap(result, 0, 1) assert_eq(stats["intersection_count"], 2.0) // C, D assert_eq(stats["union_count"], 6.0) // A,B,C,D,E,F @@ -234,6 +246,7 @@ test "venn_pairwise_overlap_two_sets" { assert_eq(stats["set_j_size"], 4.0) } +///| test "venn_pairwise_overlap_three_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -249,11 +262,11 @@ test "venn_pairwise_overlap_three_sets" { assert_eq(s23["intersection_count"], 2.0) // C, D } +///| test "venn_pairwise_overlap_identical_sets" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["A", "B", "C"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["A", "B", "C"]], names=[ + "S1", "S2", + ]) let stats = @src.venn_pairwise_overlap(result, 0, 1) assert_eq(stats["intersection_count"], 3.0) assert_eq(stats["union_count"], 3.0) @@ -262,11 +275,9 @@ test "venn_pairwise_overlap_identical_sets" { assert_eq(stats["dice_coefficient"], 1.0) } +///| test "venn_pairwise_overlap_disjoint" { - let result = @src.venn_diagram( - [["A", "B"], ["C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B"], ["C", "D"]], names=["S1", "S2"]) let stats = @src.venn_pairwise_overlap(result, 0, 1) assert_eq(stats["intersection_count"], 0.0) assert_eq(stats["jaccard_index"], 0.0) @@ -276,6 +287,8 @@ test "venn_pairwise_overlap_disjoint" { // 9. Set operations: venn_intersection, venn_union (venn_union_two), // venn_difference, venn_symmetric_difference + +///| test "venn_intersection_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] @@ -285,6 +298,7 @@ test "venn_intersection_basic" { assert_true(inter.contains("D")) } +///| test "venn_intersection_empty" { let a = ["A", "B"] let b = ["C", "D"] @@ -292,6 +306,7 @@ test "venn_intersection_empty" { assert_eq(inter.length(), 0) } +///| test "venn_union_two_basic" { let a = ["A", "B", "C"] let b = ["C", "D", "E"] @@ -304,6 +319,7 @@ test "venn_union_two_basic" { assert_true(uni.contains("E")) } +///| test "venn_union_two_disjoint" { let a = ["A", "B"] let b = ["C", "D"] @@ -311,6 +327,7 @@ test "venn_union_two_disjoint" { assert_eq(uni.length(), 4) } +///| test "venn_difference_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] @@ -320,6 +337,7 @@ test "venn_difference_basic" { assert_true(diff.contains("B")) } +///| test "venn_difference_no_difference" { let a = ["A", "B"] let b = ["A", "B", "C"] @@ -327,6 +345,7 @@ test "venn_difference_no_difference" { assert_eq(diff.length(), 0) } +///| test "venn_difference_full_difference" { let a = ["A", "B"] let b = ["C", "D"] @@ -336,6 +355,7 @@ test "venn_difference_full_difference" { assert_true(diff.contains("B")) } +///| test "venn_symmetric_difference_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] @@ -347,6 +367,7 @@ test "venn_symmetric_difference_basic" { assert_true(symdiff.contains("F")) } +///| test "venn_symmetric_difference_identical" { let a = ["A", "B", "C"] let b = ["A", "B", "C"] @@ -354,6 +375,7 @@ test "venn_symmetric_difference_identical" { assert_eq(symdiff.length(), 0) } +///| test "venn_symmetric_difference_disjoint" { let a = ["A", "B"] let b = ["C", "D"] @@ -362,6 +384,8 @@ test "venn_symmetric_difference_disjoint" { } // 10. venn_jaccard, venn_dice, venn_overlap_coefficient + +///| test "venn_jaccard_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] @@ -369,6 +393,7 @@ test "venn_jaccard_basic" { assert_eq(j, 2.0 / 6.0) } +///| test "venn_jaccard_identical" { let a = ["A", "B", "C"] let b = ["A", "B", "C"] @@ -376,6 +401,7 @@ test "venn_jaccard_identical" { assert_eq(j, 1.0) } +///| test "venn_jaccard_disjoint" { let a = ["A", "B"] let b = ["C", "D"] @@ -383,6 +409,7 @@ test "venn_jaccard_disjoint" { assert_eq(j, 0.0) } +///| test "venn_jaccard_empty" { let a : Array[String] = [] let b : Array[String] = [] @@ -390,13 +417,15 @@ test "venn_jaccard_empty" { assert_eq(j, 0.0) } +///| test "venn_dice_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] let d = @src.venn_dice(a, b) - assert_eq(d, (2.0 * 2.0) / 8.0) + assert_eq(d, 2.0 * 2.0 / 8.0) } +///| test "venn_dice_identical" { let a = ["A", "B", "C"] let b = ["A", "B", "C"] @@ -404,6 +433,7 @@ test "venn_dice_identical" { assert_eq(d, 1.0) } +///| test "venn_dice_disjoint" { let a = ["A", "B"] let b = ["C", "D"] @@ -411,6 +441,7 @@ test "venn_dice_disjoint" { assert_eq(d, 0.0) } +///| test "venn_dice_empty" { let a : Array[String] = [] let b : Array[String] = [] @@ -418,6 +449,7 @@ test "venn_dice_empty" { assert_eq(d, 0.0) } +///| test "venn_overlap_coefficient_basic" { let a = ["A", "B", "C", "D"] let b = ["C", "D", "E", "F"] @@ -425,6 +457,7 @@ test "venn_overlap_coefficient_basic" { assert_eq(oc, 2.0 / 4.0) // min(4,4)=4, inter=2 } +///| test "venn_overlap_coefficient_identical" { let a = ["A", "B", "C"] let b = ["A", "B", "C"] @@ -432,6 +465,7 @@ test "venn_overlap_coefficient_identical" { assert_eq(oc, 1.0) } +///| test "venn_overlap_coefficient_disjoint" { let a = ["A", "B"] let b = ["C", "D"] @@ -439,6 +473,7 @@ test "venn_overlap_coefficient_disjoint" { assert_eq(oc, 0.0) } +///| test "venn_overlap_coefficient_empty" { let a : Array[String] = [] let b = ["A", "B"] @@ -447,6 +482,8 @@ test "venn_overlap_coefficient_empty" { } // 11. venn_counts_only - count-only mode + +///| test "venn_counts_only_two_sets" { let counts = @src.venn_counts_only( [["A", "B", "C", "D"], ["C", "D", "E", "F"]], @@ -459,6 +496,7 @@ test "venn_counts_only_two_sets" { assert_eq(counts[0], 0) // outside } +///| test "venn_counts_only_three_sets" { let counts = @src.venn_counts_only( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -471,11 +509,11 @@ test "venn_counts_only_three_sets" { assert_eq(counts[2], 0) // only S2 } +///| test "venn_counts_only_no_overlap" { - let counts = @src.venn_counts_only( - [["A", "B"], ["C", "D"], ["E", "F"]], - names=["S1", "S2", "S3"], - ) + let counts = @src.venn_counts_only([["A", "B"], ["C", "D"], ["E", "F"]], names=[ + "S1", "S2", "S3", + ]) assert_eq(counts[1], 2) assert_eq(counts[2], 2) assert_eq(counts[4], 2) @@ -483,11 +521,12 @@ test "venn_counts_only_no_overlap" { } // 12. venn_summary - summary formatting + +///| test "venn_summary_basic" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["B", "C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["B", "C", "D"]], names=[ + "S1", "S2", + ]) let summary = @src.venn_summary(result) assert_true(summary.contains("Venn Diagram Summary")) assert_true(summary.contains("Number of sets: 2")) @@ -497,6 +536,7 @@ test "venn_summary_basic" { assert_true(summary.contains("Jaccard")) } +///| test "venn_summary_three_sets" { let result = @src.venn_diagram( [["A", "B", "C"], ["B", "C", "D"], ["C", "D", "E"]], @@ -509,6 +549,8 @@ test "venn_summary_three_sets" { } // 13. venn_sample_result - sample data generation + +///| test "venn_sample_result_basic" { let result = @src.venn_sample_result() assert_eq(result.n_sets, 3) @@ -517,6 +559,7 @@ test "venn_sample_result_basic" { assert_eq(result.set_names, ["Tumor Suppressors", "DNA Repair", "PI3K/MTOR"]) } +///| test "venn_sample_gene_sets_basic" { let sets = @src.venn_sample_gene_sets() assert_eq(sets.length(), 3) @@ -528,6 +571,7 @@ test "venn_sample_gene_sets_basic" { assert_true(sets[2].contains("MTOR")) } +///| test "venn_sample_gene_names_basic" { let names = @src.venn_sample_gene_names() assert_eq(names.length(), 3) @@ -537,6 +581,8 @@ test "venn_sample_gene_names_basic" { } // 14. venn_two_sets - two-set convenience function + +///| test "venn_two_sets_basic" { let result = @src.venn_two_sets(["A", "B", "C"], ["C", "D", "E"]) assert_eq(result.n_sets, 2) @@ -544,6 +590,7 @@ test "venn_two_sets_basic" { assert_eq(result.total_elements, 5) } +///| test "venn_two_sets_custom_names" { let result = @src.venn_two_sets( ["A", "B", "C"], @@ -555,21 +602,20 @@ test "venn_two_sets_custom_names" { assert_eq(result.n_sets, 2) } +///| test "venn_two_sets_intersection" { - let result = @src.venn_two_sets( - ["A", "B", "C", "D"], - ["C", "D", "E", "F"], - ) + let result = @src.venn_two_sets(["A", "B", "C", "D"], ["C", "D", "E", "F"]) let inter = @src.venn_all_intersect(result) assert_eq(inter.length(), 2) // C, D } // 15. venn_union (from result) - get universe + +///| test "venn_union_result_basic" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) let uni = @src.venn_union(result) assert_eq(uni.length(), 5) assert_true(uni.contains("A")) @@ -580,6 +626,8 @@ test "venn_union_result_basic" { } // 16. Edge cases and integration tests + +///| test "venn_diagram_single_set" { let result = @src.venn_diagram([["A", "B", "C"]], names=["Only"]) assert_eq(result.n_sets, 1) @@ -589,50 +637,48 @@ test "venn_diagram_single_set" { assert_eq(result.regions[0].count, 0) } +///| test "venn_diagram_empty_sets" { let empty : Array[String] = [] - let result = @src.venn_diagram( - [empty, ["A", "B"]], - names=["Empty", "NonEmpty"], - ) + let result = @src.venn_diagram([empty, ["A", "B"]], names=[ + "Empty", "NonEmpty", + ]) assert_eq(result.n_sets, 2) assert_eq(result.set_sizes, [0, 2]) assert_eq(result.total_elements, 2) } +///| test "venn_diagram_duplicate_elements" { - let result = @src.venn_diagram( - [["A", "A", "B"], ["B", "C", "C"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "A", "B"], ["B", "C", "C"]], names=[ + "S1", "S2", + ]) assert_eq(result.total_elements, 3) // A, B, C - duplicates removed in universe } +///| test "venn_region_descriptions" { - let result = @src.venn_diagram( - [["A", "B", "C"], ["C", "D", "E"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C"], ["C", "D", "E"]], names=[ + "S1", "S2", + ]) assert_eq(result.regions[1].description, "Only in S1") assert_eq(result.regions[2].description, "Only in S2") assert_eq(result.regions[3].description, "In S1 & S2") assert_eq(result.regions[0].description, "Outside all sets") } +///| test "venn_euler_layout_basic" { - let result = @src.venn_diagram( - [["A", "B"], ["B", "C"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B"], ["B", "C"]], names=["S1", "S2"]) let layout = @src.venn_euler_layout(result) assert_true(layout.length() > 0) } +///| test "venn_pairwise_overlap_set_sizes" { - let result = @src.venn_diagram( - [["A", "B", "C", "D", "E"], ["C", "D"]], - names=["S1", "S2"], - ) + let result = @src.venn_diagram([["A", "B", "C", "D", "E"], ["C", "D"]], names=[ + "S1", "S2", + ]) let stats = @src.venn_pairwise_overlap(result, 0, 1) assert_eq(stats["set_i_size"], 5.0) assert_eq(stats["set_j_size"], 2.0) @@ -640,6 +686,7 @@ test "venn_pairwise_overlap_set_sizes" { assert_eq(stats["overlap_coefficient"], 2.0 / 2.0) // min(5,2)=2 } +///| test "venn_diagram_five_sets" { let result = @src.venn_diagram( [["A", "B"], ["B", "C"], ["C", "D"], ["D", "E"], ["E", "A"]], @@ -650,6 +697,7 @@ test "venn_diagram_five_sets" { assert_true(result.total_elements > 0) } +///| test "venn_all_intersect_sample_result" { let result = @src.venn_sample_result() let all = @src.venn_all_intersect(result) @@ -658,6 +706,7 @@ test "venn_all_intersect_sample_result" { assert_eq(all.length(), all2.length()) } +///| test "venn_unique_to_each_sample_result" { let result = @src.venn_sample_result() let unique = @src.venn_unique_to_each(result) @@ -665,15 +714,16 @@ test "venn_unique_to_each_sample_result" { assert_true(unique[0].length() > 0) } +///| test "venn_counts_only_matches_diagram" { let sets = [["A", "B", "C", "D"], ["C", "D", "E", "F"]] let names = ["S1", "S2"] - let result = @src.venn_diagram(sets, names=names) - let counts = @src.venn_counts_only(sets, names=names) + let result = @src.venn_diagram(sets, names~) + let counts = @src.venn_counts_only(sets, names~) assert_eq(counts.length(), result.regions.length()) let mut i = 0 while i < counts.length() { assert_eq(counts[i], result.regions[i].count) i = i + 1 } -} \ No newline at end of file +} diff --git a/test/moonbit/voyager_test.mbt b/test/moonbit/voyager_test.mbt new file mode 100644 index 00000000..9031da05 --- /dev/null +++ b/test/moonbit/voyager_test.mbt @@ -0,0 +1,1130 @@ +// Tests for the Bioconductor Voyager-inspired spatial autocorrelation module. + +///| +fn voy_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "Voyager value " + + actual.to_string() + + " differs from " + + expected.to_string() + + " (tolerance " + + tolerance.to_string() + + ")", + ) + } +} + +///| +fn voy_test_grid_coords() -> Array[Array[Double]] { + // 3x3 grid of spots, coordinates [x, y]. + let coords : Array[Array[Double]] = [] + for row in 0..<3 { + for col in 0..<3 { + coords.push([col.to_double(), row.to_double()]) + } + } + coords +} + +///| +fn voy_test_chain_weights() -> @src.VoyagerWeights { + // 4 spots in a chain 0-1-2-3 with binary (style B) symmetric weights. + @src.VoyagerWeights::new( + [[1], [0, 2], [1, 3], [2]], + [[1.0], [1.0, 1.0], [1.0, 1.0], [1.0]], + style="B", + ) catch { + _ => abort("valid chain weights should build") + } +} + +///| +fn voy_test_gradient_values() -> Array[Double] { + // Values increasing with the x coordinate on the 3x3 grid: strong positive + // spatial autocorrelation. + [ + 0.0, 1.0, 2.0, // row 0 + 0.0, 1.0, 2.0, // row 1 + 0.0, 1.0, 2.0, // row 2 + ] +} + +///| +fn voy_test_checkerboard_values() -> Array[Double] { + // Alternating high/low: negative spatial autocorrelation on the grid. + [0.0, 9.0, 0.0, 9.0, 0.0, 9.0, 0.0, 9.0, 0.0] +} + +///| +fn voy_test_hotspot_values() -> Array[Double] { + // 3x3 grid with a hotspot in the centre. + [0.0, 0.0, 0.0, 0.0, 50.0, 0.0, 0.0, 0.0, 0.0] +} + +// =========================================================================== +// Weights construction +// =========================================================================== + +///| +test "Voyager weights kNN builds row-standardized neighbours" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + assert_eq(weights.n, 9) + assert_eq(weights.style, "W") + // Each spot on a 3x3 grid has 4 neighbours (corners: 2, edges: 3, centre: 4 + // are clamped to k=4 but corners/edges have fewer available neighbours). + assert_true(weights.neighbors[4].length() == 4) // centre + // Row-standardized: each row sums to 1. + for index in 0.. true + } + assert_true(failed) +} + +///| +test "Voyager weights kNN clamps k above n-1" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 100) catch { + _ => abort("kNN weights with large k should clamp") + } + // With k clamped to n-1, every spot links to all other spots. + assert_true(weights.neighbors[0].length() == 8) +} + +///| +test "Voyager weights distance band links immediate grid neighbours" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_distance_band(coords, 1.0) catch { + _ => abort("distance-band weights should build") + } + assert_eq(weights.n, 9) + // Centre spot (index 4) has 4 immediate neighbours at distance 1. + assert_eq(weights.neighbors[4].length(), 4) + // Corner spot (index 0) has 2 immediate neighbours. + assert_eq(weights.neighbors[0].length(), 2) +} + +///| +test "Voyager weights inverse distance produces positive finite weights" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_inverse_distance(coords, 1.0) catch { + _ => abort("inverse-distance weights should build") + } + assert_eq(weights.n, 9) + for index in 0.. 0.0) + assert_true(value <= 1.0e300) + } + } +} + +///| +test "Voyager weights style C sums to one globally" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4, style="C") catch { + _ => abort("style C weights should build") + } + let mut total = 0.0 + for index in 0.. abort("style B weights should build") + } + for index in 0.. true + } + assert_true(failed) +} + +// =========================================================================== +// Global Moran's I +// =========================================================================== + +///| +test "Voyager global Moran's I matches hand-computed chain example" { + let weights = voy_test_chain_weights() + let result = @src.voyager_global_morans_i([1.0, 1.0, 9.0, 9.0], weights) catch { + _ => abort("global Moran's I should run") + } + voy_test_close(result.estimate, 1.0 / 3.0, 1.0e-10) + voy_test_close(result.expectation, -1.0 / 3.0, 1.0e-12) + voy_test_close(result.s0, 6.0, 1.0e-12) + voy_test_close(result.s1, 12.0, 1.0e-12) + voy_test_close(result.s2, 40.0, 1.0e-12) + // b2 = 1, n = 4 -> variance = 88/216 - 1/9 = 0.296296... + voy_test_close(result.variance, 88.0 / 216.0 - 1.0 / 9.0, 1.0e-10) + assert_eq(result.n, 4) +} + +///| +test "Voyager global Moran's I is positive for a spatial gradient" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_global_morans_i(voy_test_gradient_values(), weights) catch { + _ => abort("global Moran's I should run") + } + assert_true(result.estimate > 0.0) + assert_true(result.p_value < 0.05) +} + +///| +test "Voyager global Moran's I is negative for a checkerboard" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_global_morans_i( + voy_test_checkerboard_values(), + weights, + ) catch { + _ => abort("global Moran's I should run") + } + assert_true(result.estimate < 0.0) +} + +///| +test "Voyager global Moran's I rejects non-finite values" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let failed = try { + ignore(@src.voyager_global_morans_i([1.0, 2.0, 1.0e400, 4.0], weights)) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager global Moran's I rejects length mismatch" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let failed = try { + ignore(@src.voyager_global_morans_i([1.0, 2.0], weights)) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Global Geary's c +// =========================================================================== + +///| +test "Voyager global Geary's c matches hand-computed chain example" { + let weights = voy_test_chain_weights() + let result = @src.voyager_global_gearys_c([1.0, 1.0, 9.0, 9.0], weights) catch { + _ => abort("global Geary's c should run") + } + voy_test_close(result.estimate, 0.5, 1.0e-10) + voy_test_close(result.expectation, 1.0, 1.0e-12) + assert_eq(result.n, 4) +} + +///| +test "Voyager global Geary's c is below one for a gradient" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_global_gearys_c(voy_test_gradient_values(), weights) catch { + _ => abort("global Geary's c should run") + } + assert_true(result.estimate < 1.0) +} + +///| +test "Voyager global Geary's c is above one for a checkerboard" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_global_gearys_c( + voy_test_checkerboard_values(), + weights, + ) catch { + _ => abort("global Geary's c should run") + } + assert_true(result.estimate > 1.0) +} + +// =========================================================================== +// Local Moran's I (LISA) +// =========================================================================== + +///| +test "Voyager local Moran's I returns one statistic per spot" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=99, + seed=123, + ) catch { + _ => abort("local Moran's I should run") + } + assert_eq(result.n, 9) + assert_eq(result.local_i.length(), 9) + assert_eq(result.quadrants.length(), 9) + assert_eq(result.fdr.length(), 9) + assert_eq(result.permutations, 99) +} + +///| +test "Voyager local Moran's I permutation is deterministic for a fixed seed" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let first = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=50, + seed=777, + ) catch { + _ => abort("local Moran's I should run") + } + let second = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=50, + seed=777, + ) catch { + _ => abort("local Moran's I should run") + } + for index in 0.. abort("kNN weights should build") + } + let result = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=0, + ) catch { + _ => abort("local Moran's I should run") + } + for label in result.quadrants { + assert_true( + label == "HH" || + label == "LL" || + label == "HL" || + label == "LH" || + label == "not significant", + ) + } +} + +///| +test "Voyager local Moran's I BH-FDR is monotonic non-decreasing in rank" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=99, + seed=42, + ) catch { + _ => abort("local Moran's I should run") + } + let indexed : Array[(Double, Int)] = [] + for index in 0.. Int { + if a.0 < b.0 { + -1 + } else if a.0 > b.0 { + 1 + } else { + 0 + } + }) + let mut previous = 0.0 + for entry in indexed { + let adjusted = result.fdr[entry.1] + assert_true(adjusted >= previous - 1.0e-12) + previous = adjusted + } +} + +///| +test "Voyager local Moran's I rejects zero-variance input" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let failed = try { + ignore( + @src.voyager_local_morans_i( + [5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0, 5.0], + weights, + ), + ) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Local Geary's c +// =========================================================================== + +///| +test "Voyager local Geary's c returns per-spot statistics and classifications" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_gearys_c( + voy_test_gradient_values(), + weights, + permutations=99, + seed=9, + ) catch { + _ => abort("local Geary's c should run") + } + assert_eq(result.n, 9) + assert_eq(result.local_c.length(), 9) + assert_eq(result.classifications.length(), 9) + for label in result.classifications { + assert_true( + label == "similar" || label == "dissimilar" || label == "not significant", + ) + } +} + +///| +test "Voyager local Geary's c is non-negative everywhere" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_gearys_c(voy_test_gradient_values(), weights) catch { + _ => abort("local Geary's c should run") + } + for value in result.local_c { + assert_true(value >= 0.0) + } +} + +// =========================================================================== +// Local Getis–Ord Gi / Gi* +// =========================================================================== + +///| +test "Voyager local Getis-Ord Gi* flags the centre of a hotspot" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_getis_ord( + voy_test_hotspot_values(), + weights, + star=true, + fdr_threshold=0.2, + ) catch { + _ => abort("Getis–Ord should run") + } + assert_eq(result.n, 9) + assert_eq(result.star, true) + // Centre spot (index 4) is the hotspot. + assert_true(result.z_scores[4] > 0.0) + assert_true(result.classifications[4] == "hotspot") +} + +///| +test "Voyager local Getis-Ord Gi excludes self from the weighted sum" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let gi = @src.voyager_local_getis_ord( + voy_test_hotspot_values(), + weights, + star=false, + ) catch { + _ => abort("Getis–Ord should run") + } + let gi_star = @src.voyager_local_getis_ord( + voy_test_hotspot_values(), + weights, + star=true, + ) catch { + _ => abort("Getis–Ord should run") + } + assert_eq(gi.star, false) + // The two statistics differ for the hotspot centre because Gi* includes self. + assert_true((gi.statistic[4] - gi_star.statistic[4]).abs() > 1.0e-9) +} + +///| +test "Voyager local Getis-Ord rejects negative values" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let failed = try { + ignore( + @src.voyager_local_getis_ord( + [-1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], + weights, + ), + ) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager local Getis-Ord z-scores are finite" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_getis_ord( + voy_test_gradient_values(), + weights, + star=true, + ) catch { + _ => abort("Getis–Ord should run") + } + for value in result.z_scores { + assert_true(value == value) + assert_true(value.abs() <= 1.0e300) + } +} + +// =========================================================================== +// Lee's L +// =========================================================================== + +///| +test "Voyager Lee's L is positive for co-varying features" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_lees_l( + voy_test_gradient_values(), + voy_test_gradient_values(), + weights, + ) catch { + _ => abort("Lee's L should run") + } + assert_eq(result.n, 9) + // A feature with itself: Lee's L is positive and bounded by 1. + assert_true(result.global_l > 0.0) + assert_true(result.global_l <= 1.0 + 1.0e-9) +} + +///| +test "Voyager Lee's L is negative for anti-correlated features" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let anti : Array[Double] = [] + for value in voy_test_gradient_values() { + anti.push(10.0 - value) + } + let result = @src.voyager_lees_l(voy_test_gradient_values(), anti, weights) catch { + _ => abort("Lee's L should run") + } + assert_true(result.global_l < 0.0) +} + +///| +test "Voyager Lee's L local values length matches spot count" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_lees_l( + voy_test_gradient_values(), + voy_test_hotspot_values(), + weights, + permutations=49, + seed=5, + ) catch { + _ => abort("Lee's L should run") + } + assert_eq(result.local_l.length(), 9) + assert_eq(result.fdr.length(), 9) +} + +///| +test "Voyager Lee's L rejects unequal feature lengths" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let failed = try { + ignore( + @src.voyager_lees_l([1.0, 2.0, 3.0], voy_test_gradient_values(), weights), + ) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +// =========================================================================== +// Multivariate local Geary +// =========================================================================== + +///| +test "Voyager multivariate local Geary combines features into one statistic" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let matrix : Array[Array[Double]] = [ + voy_test_gradient_values(), + voy_test_hotspot_values(), + ] + let result = @src.voyager_multivariate_local_geary( + matrix, + weights, + permutations=49, + seed=3, + ) catch { + _ => abort("multivariate local Geary should run") + } + assert_eq(result.n, 9) + assert_eq(result.n_features, 2) + assert_eq(result.local_statistic.length(), 9) + assert_eq(result.classifications.length(), 9) + assert_eq(result.feature_set, ["feature1", "feature2"]) +} + +///| +test "Voyager multivariate local Geary uses supplied feature names" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let matrix : Array[Array[Double]] = [ + voy_test_gradient_values(), + voy_test_hotspot_values(), + ] + let result = @src.voyager_multivariate_local_geary(matrix, weights, feature_names=[ + "gradient", "hotspot", + ]) catch { + _ => abort("multivariate local Geary should run") + } + assert_eq(result.feature_set, ["gradient", "hotspot"]) +} + +///| +test "Voyager multivariate local Geary is non-negative" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let matrix : Array[Array[Double]] = [voy_test_gradient_values()] + let result = @src.voyager_multivariate_local_geary(matrix, weights) catch { + _ => abort("multivariate local Geary should run") + } + for value in result.local_statistic { + assert_true(value >= 0.0) + } +} + +// =========================================================================== +// Empirical variogram and model fitting +// =========================================================================== + +///| +test "Voyager empirical variogram bins pairs and computes semivariance" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=5, + ) catch { + _ => abort("empirical variogram should run") + } + assert_true(empirical.length() >= 2) + for point in empirical { + assert_true(point.npairs > 0) + assert_true(point.gamma >= 0.0) + assert_true(point.lag > 0.0) + } +} + +///| +test "Voyager variogram fit recovers a positive range and bounded SSE" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=8, + ) catch { + _ => abort("empirical variogram should run") + } + let model = @src.voyager_fit_variogram(empirical, model_type="spherical") catch { + _ => abort("spherical fit should run") + } + assert_eq(model.model_type, "spherical") + assert_true(model.range > 0.0) + assert_true(model.sill >= 0.0) + assert_true(model.nugget >= 0.0) + assert_true(model.fitted_sse >= 0.0) +} + +///| +test "Voyager variogram predict increases with distance up to the range" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=6, + ) catch { + _ => abort("empirical variogram should run") + } + let model = @src.voyager_fit_variogram(empirical, model_type="exponential") catch { + _ => abort("exponential fit should run") + } + let small = @src.voyager_variogram_predict(model, 0.01) + let large = @src.voyager_variogram_predict(model, model.range * 3.0) + assert_true(large >= small) +} + +///| +test "Voyager variogram fit supports gaussian model" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=6, + ) catch { + _ => abort("empirical variogram should run") + } + let model = @src.voyager_fit_variogram(empirical, model_type="gaussian") catch { + _ => abort("gaussian fit should run") + } + assert_eq(model.model_type, "gaussian") +} + +///| +test "Voyager variogram fit rejects an unknown model type" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=4, + ) catch { + _ => abort("empirical variogram should run") + } + let failed = try { + ignore(@src.voyager_fit_variogram(empirical, model_type="cubic")) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager variogram fit respects a fixed nugget" { + let coords = voy_test_grid_coords() + let empirical = @src.voyager_empirical_variogram( + coords, + voy_test_gradient_values(), + n_lags=6, + ) catch { + _ => abort("empirical variogram should run") + } + let model = @src.voyager_fit_variogram( + empirical, + model_type="spherical", + nugget=0.25, + fix_nugget=true, + ) catch { + _ => abort("fixed-nugget fit should run") + } + voy_test_close(model.nugget, 0.25, 1.0e-12) +} + +// =========================================================================== +// Correlogram +// =========================================================================== + +///| +test "Voyager correlogram returns one point per non-empty lag" { + let coords = voy_test_grid_coords() + let points = @src.voyager_correlogram( + coords, + voy_test_gradient_values(), + n_lags=4, + ) catch { + _ => abort("correlogram should run") + } + // Empty distance bins are skipped; every returned point has observed pairs. + assert_true(points.length() >= 1) + assert_true(points.length() <= 4) + let mut previous_lag = -1.0 + for point in points { + assert_true(point.npairs > 0) + assert_true(point.lag > previous_lag) + previous_lag = point.lag + voy_test_close(point.expectation, -1.0 / 8.0, 1.0e-12) + } +} + +///| +test "Voyager correlogram gradient shows positive Moran at short lags" { + let coords = voy_test_grid_coords() + let points = @src.voyager_correlogram( + coords, + voy_test_gradient_values(), + n_lags=3, + ) catch { + _ => abort("correlogram should run") + } + assert_true(points[0].morans_i > 0.0) +} + +// =========================================================================== +// SpatialExperiment integration +// =========================================================================== + +///| +test "Voyager example SpatialExperiment has the expected shape" { + let se = @src.voyager_example_spatial_experiment() + assert_eq(se.col_data.length(), 36) + assert_eq(se.spatial_coords.length(), 36) + let assay = se.assay["logcounts"] + assert_eq(assay.length(), 4) + assert_eq(assay[0].length(), 36) +} + +///| +test "Voyager univariate SFE Moran writes back local results without mutating input" { + let se = @src.voyager_example_spatial_experiment() + let original_keys = se.col_data[0].keys().length() + let output = @src.voyager_run_univariate_sfe( + se, + [0], + assay_name="logcounts", + stat_method="moran", + permutations=19, + seed=2024, + output_prefix="voy", + ) catch { + _ => abort("Voyager SFE Moran integration should run") + } + // Input SpatialExperiment unchanged. + assert_eq(se.col_data[0].keys().length(), original_keys) + assert_true(!se.col_data[0].contains("voy.moran.local.gene_gradient")) + // Output SpatialExperiment carries the write-back. + assert_true( + output.experiment.col_data[0].contains("voy.moran.local.gene_gradient"), + ) + assert_true( + output.experiment.row_data[0].contains("voy.moran.I.gene_gradient"), + ) + assert_eq(output.results.length(), 1) + assert_eq(output.results[0].stat_method, "moran") + assert_eq(output.weights.n, 36) +} + +///| +test "Voyager univariate SFE Geary writes classification column" { + let se = @src.voyager_example_spatial_experiment() + let output = @src.voyager_run_univariate_sfe( + se, + [0, 1], + assay_name="logcounts", + stat_method="geary", + permutations=19, + seed=7, + output_prefix="vg", + ) catch { + _ => abort("Voyager SFE Geary integration should run") + } + assert_eq(output.results.length(), 2) + assert_true( + output.experiment.col_data[0].contains("vg.geary.class.gene_gradient"), + ) + assert_true(output.experiment.row_data[1].contains("vg.geary.C.gene_hotspot")) +} + +///| +test "Voyager univariate SFE Getis writes hotspot classification" { + let se = @src.voyager_example_spatial_experiment() + let output = @src.voyager_run_univariate_sfe( + se, + [1], + assay_name="logcounts", + stat_method="getis", + output_prefix="g", + ) catch { + _ => abort("Voyager SFE Getis integration should run") + } + assert_eq(output.results.length(), 1) + assert_eq(output.results[0].stat_method, "getis") + assert_true( + output.experiment.col_data[0].contains("g.getis.class.gene_hotspot"), + ) + assert_true(output.experiment.col_data[0].contains("g.getis.z.gene_hotspot")) +} + +///| +test "Voyager univariate SFE can use a distance-band bandwidth" { + let se = @src.voyager_example_spatial_experiment() + let output = @src.voyager_run_univariate_sfe( + se, + [0], + assay_name="logcounts", + stat_method="moran", + bandwidth=1.0, + output_prefix="db", + ) catch { + _ => abort("Voyager SFE distance-band integration should run") + } + assert_eq(output.weights.n, 36) + // Distance band of 1.0 on a unit grid links only orthogonal immediate + // neighbours (excluding diagonals at sqrt(2)), so centre spots have 4. + let mut has_four = false + for index in 0.. true + } + assert_true(failed) +} + +///| +test "Voyager univariate SFE rejects an unknown method" { + let se = @src.voyager_example_spatial_experiment() + let failed = try { + ignore( + @src.voyager_run_univariate_sfe( + se, + [0], + assay_name="logcounts", + stat_method="unknown", + ), + ) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager univariate SFE rejects a missing assay" { + let se = @src.voyager_example_spatial_experiment() + let failed = try { + ignore(@src.voyager_run_univariate_sfe(se, [0], assay_name="counts")) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager univariate SFE metadata records method and feature count" { + let se = @src.voyager_example_spatial_experiment() + let output = @src.voyager_run_univariate_sfe( + se, + [0, 1, 2], + assay_name="logcounts", + stat_method="moran", + output_prefix="meta", + ) catch { + _ => abort("Voyager SFE integration should run") + } + assert_eq(output.experiment.metadata["meta.method"], "moran") + assert_eq(output.experiment.metadata["meta.n_features"], "3") + assert_eq(output.experiment.metadata["meta.weights_style"], "W") +} + +// =========================================================================== +// Determinism and edge cases +// =========================================================================== + +///| +test "Voyager PRNG produces reproducible shuffles for the same seed" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let a = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=30, + seed=20240501, + ) catch { + _ => abort("local Moran's I should run") + } + let b = @src.voyager_local_morans_i( + voy_test_gradient_values(), + weights, + permutations=30, + seed=20240501, + ) catch { + _ => abort("local Moran's I should run") + } + voy_test_close(a.p_values[0], b.p_values[0], 0.0) + voy_test_close(a.z_scores[4], b.z_scores[4], 0.0) +} + +///| +test "Voyager global Moran's I expectation equals negative one over n minus one" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_global_morans_i(voy_test_gradient_values(), weights) catch { + _ => abort("global Moran's I should run") + } + voy_test_close(result.expectation, -1.0 / 8.0, 1.0e-12) +} + +///| +test "Voyager weights distance band rejects non-positive bandwidth" { + let coords = voy_test_grid_coords() + let failed = try { + ignore(@src.voyager_weights_distance_band(coords, 0.0)) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager weights inverse distance rejects non-positive power" { + let coords = voy_test_grid_coords() + let failed = try { + ignore(@src.voyager_weights_inverse_distance(coords, 0.0)) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager empirical variogram rejects coincident coordinates" { + let coords : Array[Array[Double]] = [[0.0, 0.0], [0.0, 0.0], [0.0, 0.0]] + let failed = try { + ignore(@src.voyager_empirical_variogram(coords, [1.0, 2.0, 3.0])) + false + } catch { + VoyagerError(_) => true + } + assert_true(failed) +} + +///| +test "Voyager local Getis-Ord FDR is bounded in [0, 1]" { + let coords = voy_test_grid_coords() + let weights = @src.voyager_weights_knn(coords, 4) catch { + _ => abort("kNN weights should build") + } + let result = @src.voyager_local_getis_ord( + voy_test_hotspot_values(), + weights, + star=true, + ) catch { + _ => abort("Getis–Ord should run") + } + for value in result.fdr { + assert_true(value >= 0.0 && value <= 1.0) + } +} diff --git a/test/moonbit/vsn_test.mbt b/test/moonbit/vsn_test.mbt index e10610d6..d22d5499 100644 --- a/test/moonbit/vsn_test.mbt +++ b/test/moonbit/vsn_test.mbt @@ -42,11 +42,7 @@ test "vsn_control_new" { ///| test "vsn2_basic" { - let data = [ - [100.0, 200.0], - [150.0, 180.0], - [90.0, 210.0], - ] + let data = [[100.0, 200.0], [150.0, 180.0], [90.0, 210.0]] let result = @src.vsn2(data) assert_eq(result.length(), 3) assert_eq(result[0].length(), 2) @@ -54,11 +50,7 @@ test "vsn2_basic" { ///| test "vsn2_with_control_runs" { - let data = [ - [100.0, 200.0], - [150.0, 180.0], - [90.0, 210.0], - ] + let data = [[100.0, 200.0], [150.0, 180.0], [90.0, 210.0]] let ctrl = @src.VSNControl::new() let result = @src.vsn2_with_control(data, ctrl) assert_eq(result.length(), 3) @@ -67,11 +59,7 @@ test "vsn2_with_control_runs" { ///| test "vsn_fit_and_report_returns_vsnresult" { - let data = [ - [100.0, 200.0], - [150.0, 180.0], - [90.0, 210.0], - ] + let data = [[100.0, 200.0], [150.0, 180.0], [90.0, 210.0]] let fit = @src.vsn_fit_and_report(data) assert_eq(fit.params.length(), 2) let summary = @src.summarize_vsn_fit(fit) @@ -100,11 +88,7 @@ test "mean_sd_bins_and_ascii" { ///| test "vsn_denoise_keeps_shape" { - let data = [ - [100.0, 200.0], - [150.0, 180.0], - [90.0, 210.0], - ] + let data = [[100.0, 200.0], [150.0, 180.0], [90.0, 210.0]] let denoised = @src.vsn_denoise(data) assert_eq(denoised.length(), 3) assert_eq(denoised[0].length(), 2) diff --git a/test/moonbit/wise_test.mbt b/test/moonbit/wise_test.mbt index cdf77081..8d42b807 100644 --- a/test/moonbit/wise_test.mbt +++ b/test/moonbit/wise_test.mbt @@ -10,31 +10,37 @@ test "wise_block_type_to_string_exon" { assert_eq(bt.to_string(), "exon") } +///| test "wise_block_type_to_string_intron" { let bt = @src.wise_make_block_type("intron") assert_eq(bt.to_string(), "intron") } +///| test "wise_block_type_to_string_match" { let bt = @src.wise_make_block_type("match") assert_eq(bt.to_string(), "match") } +///| test "wise_block_type_to_string_mismatch" { let bt = @src.wise_make_block_type("mismatch") assert_eq(bt.to_string(), "mismatch") } +///| test "wise_block_type_to_string_insertion" { let bt = @src.wise_make_block_type("insertion") assert_eq(bt.to_string(), "insertion") } +///| test "wise_block_type_to_string_deletion" { let bt = @src.wise_make_block_type("deletion") assert_eq(bt.to_string(), "deletion") } +///| test "wise_block_type_eq_via_is_pattern" { // Enum variants can be compared using `is` pattern matching. let bt = @src.wise_make_block_type("exon") @@ -46,6 +52,7 @@ test "wise_block_type_eq_via_is_pattern" { // WiseExon // --------------------------------------------------------------------------- +///| test "wise_exon_new_defaults" { let e = @src.WiseExon::new() assert_eq(e.start, 0) @@ -57,6 +64,7 @@ test "wise_exon_new_defaults" { assert_eq(e.protein_end, 0) } +///| test "wise_exon_new_with_params" { let e = @src.WiseExon::new( start=100, @@ -76,11 +84,13 @@ test "wise_exon_new_with_params" { assert_eq(e.protein_end, 33) } +///| test "wise_exon_length" { let e = @src.WiseExon::new(start=0, end=99) assert_eq(e.length(), 100) } +///| test "wise_exon_length_single_base" { let e = @src.WiseExon::new(start=50, end=50) assert_eq(e.length(), 1) @@ -90,6 +100,7 @@ test "wise_exon_length_single_base" { // WiseIntron // --------------------------------------------------------------------------- +///| test "wise_intron_new_defaults" { let i = @src.WiseIntron::new() assert_eq(i.start, 0) @@ -100,6 +111,7 @@ test "wise_intron_new_defaults" { assert_eq(i.length, 1) } +///| test "wise_intron_new_with_params" { let i = @src.WiseIntron::new( start=100, @@ -115,11 +127,13 @@ test "wise_intron_new_with_params" { assert_eq(i.length, 100) } +///| test "wise_intron_length_auto_calc_normal" { let i = @src.WiseIntron::new(start=200, end=299) assert_eq(i.length, 100) } +///| test "wise_intron_length_auto_calc_end_before_start" { // When end < start, length should be 0 let i = @src.WiseIntron::new(start=300, end=100) @@ -130,6 +144,7 @@ test "wise_intron_length_auto_calc_end_before_start" { // WiseAlignmentColumn // --------------------------------------------------------------------------- +///| test "wise_alignment_column_new" { let c = @src.WiseAlignmentColumn::new( protein_char='M', @@ -145,11 +160,9 @@ test "wise_alignment_column_new" { assert_eq(c.protein_position, 0) } +///| test "wise_alignment_column_new_default_match_type" { - let c = @src.WiseAlignmentColumn::new( - protein_char='X', - gene_codon="TAG", - ) + let c = @src.WiseAlignmentColumn::new(protein_char='X', gene_codon="TAG") assert_eq(c.protein_char, 'X') assert_eq(c.gene_codon, "TAG") // Default match_type is " " @@ -162,6 +175,7 @@ test "wise_alignment_column_new_default_match_type" { // WiseResult construction and basic accessors // --------------------------------------------------------------------------- +///| test "wise_result_new_empty" { let r = @src.WiseResult::new() assert_eq(r.protein_id, "") @@ -173,6 +187,7 @@ test "wise_result_new_empty" { assert_eq(r.get_total_exon_length(), 0) } +///| test "wise_result_add_and_get_exons" { let r = @src.WiseResult::new() let e1 = @src.WiseExon::new(start=0, end=99) @@ -185,6 +200,7 @@ test "wise_result_add_and_get_exons" { assert_eq(exons[1].start, 200) } +///| test "wise_result_add_and_get_introns" { let r = @src.WiseResult::new() let i1 = @src.WiseIntron::new(start=100, end=199) @@ -195,6 +211,7 @@ test "wise_result_add_and_get_introns" { assert_eq(introns[0].end, 199) } +///| test "wise_result_get_num_exons" { let r = @src.WiseResult::new() assert_eq(r.get_num_exons(), 0) @@ -205,6 +222,7 @@ test "wise_result_get_num_exons" { assert_eq(r.get_num_exons(), 3) } +///| test "wise_result_get_num_introns" { let r = @src.WiseResult::new() assert_eq(r.get_num_introns(), 0) @@ -214,6 +232,7 @@ test "wise_result_get_num_introns" { assert_eq(r.get_num_introns(), 2) } +///| test "wise_result_get_total_exon_length" { let r = @src.WiseResult::new() // Exon 1: 0-99 -> length 100 @@ -225,41 +244,44 @@ test "wise_result_get_total_exon_length" { assert_eq(r.get_total_exon_length(), 280) } +///| test "wise_result_get_total_exon_length_empty" { let r = @src.WiseResult::new() assert_eq(r.get_total_exon_length(), 0) } +///| test "wise_result_set_protein_id" { let r = @src.WiseResult::new() r.set_protein_id("P12345") assert_eq(r.protein_id, "P12345") } +///| test "wise_result_set_gene_id" { let r = @src.WiseResult::new() r.set_gene_id("G67890") assert_eq(r.gene_id, "G67890") } +///| test "wise_result_set_score" { let r = @src.WiseResult::new() r.set_score(145.32) assert_eq(r.score, 145.32) } +///| test "wise_result_set_bits_score" { let r = @src.WiseResult::new() r.set_bits_score(52.18) assert_eq(r.bits_score, 52.18) } +///| test "wise_result_add_alignment_column" { let r = @src.WiseResult::new() - let c = @src.WiseAlignmentColumn::new( - protein_char='M', - gene_codon="ATG", - ) + let c = @src.WiseAlignmentColumn::new(protein_char='M', gene_codon="ATG") r.add_alignment_column(c) let aln = r.get_alignment() assert_eq(aln.length(), 1) @@ -270,27 +292,32 @@ test "wise_result_add_alignment_column" { // wise_parse // --------------------------------------------------------------------------- +///| test "wise_parse_protein_id" { let r = @src.wise_parse("Protein: P12345\n") assert_eq(r.protein_id, "P12345") } +///| test "wise_parse_gene_id" { let r = @src.wise_parse("Gene: G67890\n") assert_eq(r.gene_id, "G67890") } +///| test "wise_parse_score" { let r = @src.wise_parse("Score: 145.32\n") // Use tolerance for floating-point comparison (parse_double precision) assert_true(r.score > 145.31 && r.score < 145.33) } +///| test "wise_parse_bits_score" { let r = @src.wise_parse("Bits: 52.18\n") assert_eq(r.bits_score, 52.18) } +///| test "wise_parse_exons" { let r = @src.wise_parse( "Exon 1: 0-99 (phase 0) score 45.5 protein 1-33\nExon 2: 200-299 (phase 0) score 52.3 protein 34-66\n", @@ -308,6 +335,7 @@ test "wise_parse_exons" { assert_eq(exons[1].score, 52.3) } +///| test "wise_parse_introns" { let r = @src.wise_parse( "Intron 1: 100-199 donor 12.5 acceptor 8.3\nIntron 2: 300-399 donor 10.2 acceptor 9.1\n", @@ -325,6 +353,7 @@ test "wise_parse_introns" { assert_eq(introns[1].acceptor_score, 9.1) } +///| test "wise_parse_empty_input" { let r = @src.wise_parse("") assert_eq(r.protein_id, "") @@ -335,6 +364,7 @@ test "wise_parse_empty_input" { assert_eq(r.get_num_introns(), 0) } +///| test "wise_parse_full_sample" { let r = @src.wise_parse(@src.wise_sample_output()) assert_eq(r.protein_id, "P12345") @@ -351,6 +381,7 @@ test "wise_parse_full_sample" { // Sample data // --------------------------------------------------------------------------- +///| test "wise_sample_output_content" { let s = @src.wise_sample_output() assert_true(s.contains("Protein: P12345")) @@ -361,6 +392,7 @@ test "wise_sample_output_content" { assert_true(s.contains("Intron 1: 100-199")) } +///| test "wise_sample_result" { let r = @src.wise_sample() assert_eq(r.protein_id, "P12345") @@ -373,6 +405,7 @@ test "wise_sample_result" { // wise_write // --------------------------------------------------------------------------- +///| test "wise_write_basic" { let r = @src.WiseResult::new() r.set_protein_id("P12345") @@ -387,12 +420,29 @@ test "wise_write_basic" { assert_true(s.contains("Bits: 52.18")) } +///| test "wise_write_with_exons_and_introns" { let r = @src.WiseResult::new() r.set_protein_id("P1") r.set_gene_id("G1") - r.add_exon(@src.WiseExon::new(start=0, end=99, phase=0, score=45.5, protein_start=1, protein_end=33)) - r.add_intron(@src.WiseIntron::new(start=100, end=199, donor_score=12.5, acceptor_score=8.3)) + r.add_exon( + @src.WiseExon::new( + start=0, + end=99, + phase=0, + score=45.5, + protein_start=1, + protein_end=33, + ), + ) + r.add_intron( + @src.WiseIntron::new( + start=100, + end=199, + donor_score=12.5, + acceptor_score=8.3, + ), + ) let s = @src.wise_write(r) assert_true(s.contains("Exon 1: 0-99")) assert_true(s.contains("(phase 0)")) @@ -403,6 +453,7 @@ test "wise_write_with_exons_and_introns" { assert_true(s.contains("acceptor 8.3")) } +///| test "wise_write_roundtrip" { // Parse sample, write it back, verify key fields are present let r = @src.wise_sample() @@ -420,6 +471,7 @@ test "wise_write_roundtrip" { // wise_gene_structure // --------------------------------------------------------------------------- +///| test "wise_gene_structure_basic" { let r = @src.WiseResult::new() r.add_exon(@src.WiseExon::new(start=0, end=99, phase=0)) @@ -433,6 +485,7 @@ test "wise_gene_structure_basic" { assert_true(s.contains("Intron 1: 100-199")) } +///| test "wise_gene_structure_empty" { let r = @src.WiseResult::new() let s = @src.wise_gene_structure(r) @@ -440,6 +493,7 @@ test "wise_gene_structure_empty" { assert_eq(s.length(), 0) } +///| test "wise_gene_structure_sample" { let r = @src.wise_sample() let s = @src.wise_gene_structure(r) @@ -455,6 +509,7 @@ test "wise_gene_structure_sample" { // wise_translate_gene // --------------------------------------------------------------------------- +///| test "wise_translate_gene_single_exon" { // Exon 0-8 (9 nt) -> ATG GCC GGT -> M A G let r = @src.WiseResult::new() @@ -464,6 +519,7 @@ test "wise_translate_gene_single_exon" { assert_eq(protein, "MAG") } +///| test "wise_translate_gene_multiple_codons" { // Exon 0-11 (12 nt) -> ATG GCC GGT AAA -> M A G K let r = @src.WiseResult::new() @@ -473,6 +529,7 @@ test "wise_translate_gene_multiple_codons" { assert_eq(protein, "MAGK") } +///| test "wise_translate_gene_stop_codon" { // Exon 0-8 -> ATG TAA TAG -> M * * let r = @src.WiseResult::new() @@ -489,6 +546,7 @@ test "wise_translate_gene_stop_codon" { // wise_percent_identity // --------------------------------------------------------------------------- +///| test "wise_percent_identity_full" { // Translation matches protein exactly -> 100.0% let r = @src.WiseResult::new() @@ -499,6 +557,7 @@ test "wise_percent_identity_full" { assert_eq(pct, 100.0) } +///| test "wise_percent_identity_none" { // Translation does not match protein -> 0.0% let r = @src.WiseResult::new() @@ -509,6 +568,7 @@ test "wise_percent_identity_none" { assert_eq(pct, 0.0) } +///| test "wise_percent_identity_partial" { // Translation: MAG, protein: MAP -> 2/3 match -> ~66.67% let r = @src.WiseResult::new() @@ -519,6 +579,7 @@ test "wise_percent_identity_partial" { assert_true(pct > 66.0 && pct < 67.0) } +///| test "wise_percent_identity_empty_protein" { let r = @src.WiseResult::new() r.add_exon(@src.WiseExon::new(start=0, end=8)) @@ -531,6 +592,7 @@ test "wise_percent_identity_empty_protein" { // wise_summary // --------------------------------------------------------------------------- +///| test "wise_summary_basic" { let r = @src.wise_sample() let s = @src.wise_summary(r) @@ -544,6 +606,7 @@ test "wise_summary_basic" { assert_true(s.contains("Total exon length: 280 nt")) } +///| test "wise_summary_empty" { let r = @src.WiseResult::new() let s = @src.wise_summary(r) @@ -557,6 +620,7 @@ test "wise_summary_empty" { // GenomeWiseSegment // --------------------------------------------------------------------------- +///| test "genome_wise_segment_new_defaults" { let s = @src.GenomeWiseSegment::new() assert_eq(s.segment_id, "") @@ -567,6 +631,7 @@ test "genome_wise_segment_new_defaults" { assert_eq(s.wise_result.get_num_exons(), 0) } +///| test "genome_wise_segment_new_with_params" { let s = @src.GenomeWiseSegment::new( segment_id="seg1", @@ -584,6 +649,7 @@ test "genome_wise_segment_new_with_params" { // GenomeWiseResult // --------------------------------------------------------------------------- +///| test "genome_wise_result_new_empty" { let r = @src.GenomeWiseResult::new() assert_eq(r.gene_id, "") @@ -591,6 +657,7 @@ test "genome_wise_result_new_empty" { assert_eq(r.get_num_segments(), 0) } +///| test "genome_wise_result_add_and_get_segments" { let r = @src.GenomeWiseResult::new() let s1 = @src.GenomeWiseSegment::new(segment_id="seg1", start=100, end=200) @@ -603,6 +670,7 @@ test "genome_wise_result_add_and_get_segments" { assert_eq(segs[1].segment_id, "seg2") } +///| test "genome_wise_result_get_num_segments" { let r = @src.GenomeWiseResult::new() assert_eq(r.get_num_segments(), 0) diff --git a/test/moonbit/xcell_test.mbt b/test/moonbit/xcell_test.mbt index 5773ec1a..e2d160ba 100644 --- a/test/moonbit/xcell_test.mbt +++ b/test/moonbit/xcell_test.mbt @@ -8,17 +8,16 @@ // ============================================================================ test "xc_signature_new_basic" { - let sig = @src.XcellSignature::new( - "CD8+ T cells", - "immune", - ["CD8A", "CD8B", "GZMA", "GZMB"], - ) + let sig = @src.XcellSignature::new("CD8+ T cells", "immune", [ + "CD8A", "CD8B", "GZMA", "GZMB", + ]) assert_eq(sig.cell_type(), "CD8+ T cells") assert_eq(sig.category(), "immune") assert_eq(sig.genes().length(), 4) assert_eq(sig.genes()[0], "CD8A") } +///| test "xc_default_signatures_count_and_categories" { let sigs = @src.xcell_default_signatures() // Should have 63 signatures @@ -42,6 +41,7 @@ test "xc_default_signatures_count_and_categories" { assert_true(other_count >= 8) } +///| test "xc_default_signatures_each_has_genes" { let sigs = @src.xcell_default_signatures() for s in sigs { @@ -54,6 +54,7 @@ test "xc_default_signatures_each_has_genes" { // Parameters // ============================================================================ +///| test "xc_params_default" { let p = @src.XcellParams::new() assert_eq(p.min_gene_overlap, 3) @@ -67,6 +68,7 @@ test "xc_params_default" { // ssGSEA single-sample scoring // ============================================================================ +///| test "xc_ssgsea_single_basic" { let expr = [10.0, 5.0, 20.0, 3.0, 15.0] let genes = ["A", "B", "C", "D", "E"] @@ -77,6 +79,7 @@ test "xc_ssgsea_single_basic" { assert_true(score >= -1.0 && score <= 1.0) } +///| test "xc_ssgsea_single_high_enrichment" { // Expression sorted: A=100, B=90, C=80, D=10, E=5 (high A,B,C top ranking) // Gene set = [A,B,C] are all top -> high positive enrichment @@ -90,6 +93,7 @@ test "xc_ssgsea_single_high_enrichment" { assert_true(score_high > score_low) } +///| test "xc_ssgsea_single_empty" { let s = @src.xcell_ssgsea_single([], [], ["A"], alpha=0.25) assert_eq(s, 0.0) @@ -97,6 +101,7 @@ test "xc_ssgsea_single_empty" { assert_eq(s2, 0.0) } +///| test "xc_ssgsea_single_min_overlap" { // Only 1 gene overlap, should return 0 when we need min 3 (but function returns 0 for <3) let expr = [1.0, 2.0] @@ -110,11 +115,9 @@ test "xc_ssgsea_single_min_overlap" { // Full xCell pipeline // ============================================================================ +///| test "xc_result_creation" { - let scores : Array[Array[Double]] = [ - [0.5, 0.8], - [0.2, 0.1], - ] + let scores : Array[Array[Double]] = [[0.5, 0.8], [0.2, 0.1]] let cts = ["T cells", "B cells"] let cats = ["immune", "immune"] let samps = ["s1", "s2"] @@ -122,40 +125,39 @@ test "xc_result_creation" { let str = [0.1, 0.15] let menv = [0.45, 0.6] let p = @src.XcellParams::new() - let r = @src.XcellResult::new( - scores, cts, cats, samps, imm, str, menv, p, - ) + let r = @src.XcellResult::new(scores, cts, cats, samps, imm, str, menv, p) assert_eq(r.cell_types.length(), 2) assert_eq(r.sample_names.length(), 2) assert_true((r.immune_scores[0] - 0.35).abs() < 0.0001) assert_true((r.microenvironment_scores[1] - 0.6).abs() < 0.0001) } +///| test "xc_run_default_simple" { // Simple expression matrix with a handful of marker genes let gene_names = [ - "CD3D", "CD3E", "CD8A", "CD8B", "GZMA", "GZMB", "PRF1", "NKG7", - "CD19", "MS4A1", "CD79A", "PECAM1", "VWF", "COL1A1", "FAP", "ALB", + "CD3D", "CD3E", "CD8A", "CD8B", "GZMA", "GZMB", "PRF1", "NKG7", "CD19", "MS4A1", + "CD79A", "PECAM1", "VWF", "COL1A1", "FAP", "ALB", ] let sample_names = ["Tumor1", "Tumor2", "Normal1"] // 16 genes x 3 samples: upregulate immune in Normal1, stromal in Tumor2 let expression = [ - [10.0, 8.0, 100.0], // CD3D - [12.0, 9.0, 110.0], // CD3E - [15.0, 10.0, 120.0], // CD8A - [14.0, 9.0, 115.0], // CD8B - [20.0, 15.0, 130.0], // GZMA - [18.0, 12.0, 125.0], // GZMB - [16.0, 11.0, 110.0], // PRF1 - [17.0, 13.0, 115.0], // NKG7 - [5.0, 4.0, 80.0], // CD19 - [6.0, 5.0, 85.0], // MS4A1 - [5.0, 3.0, 75.0], // CD79A - [20.0, 100.0, 15.0], // PECAM1 - [18.0, 95.0, 12.0], // VWF - [30.0, 150.0, 10.0], // COL1A1 - [25.0, 140.0, 8.0], // FAP - [10.0, 20.0, 200.0], // ALB + [10.0, 8.0, 100.0], // CD3D + [12.0, 9.0, 110.0], // CD3E + [15.0, 10.0, 120.0], // CD8A + [14.0, 9.0, 115.0], // CD8B + [20.0, 15.0, 130.0], // GZMA + [18.0, 12.0, 125.0], // GZMB + [16.0, 11.0, 110.0], // PRF1 + [17.0, 13.0, 115.0], // NKG7 + [5.0, 4.0, 80.0], // CD19 + [6.0, 5.0, 85.0], // MS4A1 + [5.0, 3.0, 75.0], // CD79A + [20.0, 100.0, 15.0], // PECAM1 + [18.0, 95.0, 12.0], // VWF + [30.0, 150.0, 10.0], // COL1A1 + [25.0, 140.0, 8.0], // FAP + [10.0, 20.0, 200.0], // ALB ] let result = @src.xcell_run_default(expression, gene_names, sample_names) // Should have returned scores with 63 cell types and 3 samples @@ -175,11 +177,9 @@ test "xc_run_default_simple" { // Result accessors // ============================================================================ +///| test "xc_result_get_cell_type_scores" { - let scores : Array[Array[Double]] = [ - [0.3, 0.5], - [0.8, 0.6], - ] + let scores : Array[Array[Double]] = [[0.3, 0.5], [0.8, 0.6]] let cts = ["T-cells", "B-cells"] let cats = ["immune", "immune"] let samps = ["s1", "s2"] @@ -187,9 +187,7 @@ test "xc_result_get_cell_type_scores" { let str = [0.1, 0.1] let menv = [0.65, 0.65] let p = @src.XcellParams::new() - let r = @src.XcellResult::new( - scores, cts, cats, samps, imm, str, menv, p, - ) + let r = @src.XcellResult::new(scores, cts, cats, samps, imm, str, menv, p) let tc = r.get_cell_type_scores("T-cells") assert_eq(tc.length(), 2) assert_true((tc[0] - 0.3).abs() < 0.0001) @@ -198,11 +196,9 @@ test "xc_result_get_cell_type_scores" { assert_eq(missing.length(), 0) } +///| test "xc_result_get_sample_scores" { - let scores : Array[Array[Double]] = [ - [0.3, 0.5], - [0.8, 0.6], - ] + let scores : Array[Array[Double]] = [[0.3, 0.5], [0.8, 0.6]] let cts = ["T-cells", "B-cells"] let cats = ["immune", "immune"] let samps = ["s1", "s2"] @@ -210,9 +206,7 @@ test "xc_result_get_sample_scores" { let str = [0.1, 0.1] let menv = [0.65, 0.65] let p = @src.XcellParams::new() - let r = @src.XcellResult::new( - scores, cts, cats, samps, imm, str, menv, p, - ) + let r = @src.XcellResult::new(scores, cts, cats, samps, imm, str, menv, p) let s1 = r.get_sample_scores("s1") assert_eq(s1.length(), 2) assert_true((s1[0] - 0.3).abs() < 0.0001) @@ -221,6 +215,7 @@ test "xc_result_get_sample_scores" { assert_eq(s_missing.length(), 0) } +///| test "xc_result_get_top_cell_types" { let scores : Array[Array[Double]] = [ [0.1, 0.9], @@ -235,9 +230,7 @@ test "xc_result_get_top_cell_types" { let str = [0.1, 0.1] let menv = [0.6, 0.65] let p = @src.XcellParams::new() - let r = @src.XcellResult::new( - scores, cts, cats, samps, imm, str, menv, p, - ) + let r = @src.XcellResult::new(scores, cts, cats, samps, imm, str, menv, p) // For s1, top is C=0.9, then B=0.5 let top_s1 = r.get_top_cell_types("s1", 2) assert_eq(top_s1.length(), 2) @@ -251,12 +244,9 @@ test "xc_result_get_top_cell_types" { assert_eq(top_s2[0].0, "A") } +///| test "xc_result_scores_by_category" { - let scores : Array[Array[Double]] = [ - [0.5, 0.7], - [0.3, 0.4], - [0.9, 0.8], - ] + let scores : Array[Array[Double]] = [[0.5, 0.7], [0.3, 0.4], [0.9, 0.8]] let cts = ["T", "B", "Fibro"] let cats = ["immune", "immune", "stromal"] let samps = ["s1", "s2"] @@ -264,9 +254,7 @@ test "xc_result_scores_by_category" { let str = [0.9, 0.8] let menv = [1.3, 1.35] let p = @src.XcellParams::new() - let r = @src.XcellResult::new( - scores, cts, cats, samps, imm, str, menv, p, - ) + let r = @src.XcellResult::new(scores, cts, cats, samps, imm, str, menv, p) let by_cat = r.scores_by_category() // "immune" average of T and B for s1: (0.5+0.3)/2 = 0.4 let imm_opt = by_cat.get("immune") diff --git a/test/moonbit/xdna_io_test.mbt b/test/moonbit/xdna_io_test.mbt index ece78f56..a533546f 100644 --- a/test/moonbit/xdna_io_test.mbt +++ b/test/moonbit/xdna_io_test.mbt @@ -36,13 +36,19 @@ test "xdna_seq_type_from_int_round_trip" { let rna = @src.XdnaSeqType::from_int(1) let protein = @src.XdnaSeqType::from_int(2) let unknown = @src.XdnaSeqType::from_int(3) - assert_true(@src.XdnaSeqType::from_int(dna.to_int()) is @src.XdnaSeqType::DnaType) - assert_true(@src.XdnaSeqType::from_int(rna.to_int()) is @src.XdnaSeqType::RnaType) assert_true( - @src.XdnaSeqType::from_int(protein.to_int()) is @src.XdnaSeqType::ProteinType, + @src.XdnaSeqType::from_int(dna.to_int()) is @src.XdnaSeqType::DnaType, ) assert_true( - @src.XdnaSeqType::from_int(unknown.to_int()) is @src.XdnaSeqType::UnknownType, + @src.XdnaSeqType::from_int(rna.to_int()) is @src.XdnaSeqType::RnaType, + ) + assert_true( + @src.XdnaSeqType::from_int(protein.to_int()) + is @src.XdnaSeqType::ProteinType, + ) + assert_true( + @src.XdnaSeqType::from_int(unknown.to_int()) + is @src.XdnaSeqType::UnknownType, ) } @@ -160,19 +166,27 @@ test "xdna_file_new_is_empty" { ///| test "xdna_file_add_record_and_n_records" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("a", "ACGT", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("a", "ACGT", @src.XdnaSeqType::from_int(0), ""), + ) assert_eq(file.n_records(), 1) - file.add_record(@src.XdnaRecord::new("b", "TTTT", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("b", "TTTT", @src.XdnaSeqType::from_int(0), ""), + ) assert_eq(file.n_records(), 2) } ///| test "xdna_file_records_returns_copy" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("a", "ACGT", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("a", "ACGT", @src.XdnaSeqType::from_int(0), ""), + ) let recs = file.records() // Mutating the returned array should not affect the file. - recs.push(@src.XdnaRecord::new("x", "GGGG", @src.XdnaSeqType::from_int(0), "")) + recs.push( + @src.XdnaRecord::new("x", "GGGG", @src.XdnaSeqType::from_int(0), ""), + ) assert_eq(file.n_records(), 1) assert_eq(recs.length(), 2) } @@ -180,8 +194,12 @@ test "xdna_file_records_returns_copy" { ///| test "xdna_file_records_access" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), "")) - file.add_record(@src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), ""), + ) + file.add_record( + @src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), ""), + ) let recs = file.records() assert_eq(recs[0].name(), "seq1") assert_eq(recs[0].sequence(), "ATGC") @@ -350,8 +368,12 @@ test "xdna_read_u32_be_at_offset" { ///| test "xdna_to_bytes_from_bytes_round_trip_dna" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), "")) - file.add_record(@src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), ""), + ) + file.add_record( + @src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 2) @@ -367,7 +389,12 @@ test "xdna_to_bytes_from_bytes_round_trip_dna" { test "xdna_to_bytes_from_bytes_round_trip_with_annotations" { let file = @src.XdnaFile::new() file.add_record( - @src.XdnaRecord::new("annot1", "ATGC", @src.XdnaSeqType::from_int(0), "some annotation"), + @src.XdnaRecord::new( + "annot1", + "ATGC", + @src.XdnaSeqType::from_int(0), + "some annotation", + ), ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) @@ -382,7 +409,9 @@ test "xdna_to_bytes_from_bytes_round_trip_with_annotations" { test "xdna_to_bytes_from_bytes_round_trip_rna" { // RNA sequences are stored as raw ASCII (not 2-bit packed). let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("rna1", "AUGCAUGC", @src.XdnaSeqType::from_int(1), "")) + file.add_record( + @src.XdnaRecord::new("rna1", "AUGCAUGC", @src.XdnaSeqType::from_int(1), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 1) @@ -395,7 +424,9 @@ test "xdna_to_bytes_from_bytes_round_trip_rna" { ///| test "xdna_to_bytes_from_bytes_round_trip_protein" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("prot1", "MKLVGV", @src.XdnaSeqType::from_int(2), "")) + file.add_record( + @src.XdnaRecord::new("prot1", "MKLVGV", @src.XdnaSeqType::from_int(2), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 1) @@ -408,7 +439,9 @@ test "xdna_to_bytes_from_bytes_round_trip_protein" { ///| test "xdna_to_bytes_includes_checksum" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) // The last 4 bytes are the checksum (big-endian u32). let checksum = @src.xdna_read_u32_be(bytes, bytes.length() - 4) @@ -423,7 +456,9 @@ test "xdna_to_bytes_includes_checksum" { ///| test "xdna_from_bytes_preserves_version" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.version(), 1) @@ -436,7 +471,9 @@ test "xdna_from_bytes_preserves_version" { ///| test "xdna_write_produces_lowercase_hex" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("A", "ATGC", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("A", "ATGC", @src.XdnaSeqType::from_int(0), ""), + ) let hex = @src.xdna_write(file) // Every character must be a lowercase hex digit. for c in hex { @@ -450,8 +487,12 @@ test "xdna_write_produces_lowercase_hex" { ///| test "xdna_write_read_round_trip" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), "")) - file.add_record(@src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), ""), + ) + file.add_record( + @src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), ""), + ) let hex = @src.xdna_write(file) let file2 = @src.xdna_read(hex) assert_eq(file2.n_records(), 2) @@ -466,7 +507,12 @@ test "xdna_write_read_round_trip" { test "xdna_write_read_round_trip_with_annotations" { let file = @src.XdnaFile::new() file.add_record( - @src.XdnaRecord::new("a1", "GGGG", @src.XdnaSeqType::from_int(0), "note here"), + @src.XdnaRecord::new( + "a1", + "GGGG", + @src.XdnaSeqType::from_int(0), + "note here", + ), ) let hex = @src.xdna_write(file) let file2 = @src.xdna_read(hex) @@ -482,7 +528,12 @@ test "xdna_write_read_round_trip_with_annotations" { ///| test "xdna_get_sequence_string" { - let rec = @src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), "") + let rec = @src.XdnaRecord::new( + "seq1", + "ATGCATGC", + @src.XdnaSeqType::from_int(0), + "", + ) assert_eq(@src.xdna_get_sequence_string(rec), "ATGCATGC") } @@ -499,8 +550,12 @@ test "xdna_get_sequence_string_empty" { ///| test "xdna_to_seq_records" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), "")) - file.add_record(@src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGCATGC", @src.XdnaSeqType::from_int(0), ""), + ) + file.add_record( + @src.XdnaRecord::new("seq2", "GATTACA", @src.XdnaSeqType::from_int(0), ""), + ) let records = @src.xdna_to_seq_records(file) assert_eq(records.length(), 2) assert_eq(records[0].id, "seq1") @@ -530,7 +585,9 @@ test "xdna_from_seq_records" { ///| test "xdna_seq_records_round_trip" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("seq1", "ATGC", @src.XdnaSeqType::from_int(0), ""), + ) let records = @src.xdna_to_seq_records(file) let file2 = @src.xdna_from_seq_records(records, @src.XdnaSeqType::from_int(0)) assert_eq(file2.n_records(), 1) @@ -619,7 +676,9 @@ test "edge_case_empty_file_hex_round_trip" { ///| test "edge_case_single_record" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("only", "ACGTACGT", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("only", "ACGTACGT", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 1) @@ -631,7 +690,9 @@ test "edge_case_single_record" { ///| test "edge_case_single_base" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("one", "A", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("one", "A", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 1) @@ -643,7 +704,9 @@ test "edge_case_single_base" { ///| test "edge_case_single_base_hex_round_trip" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("one", "G", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("one", "G", @src.XdnaSeqType::from_int(0), ""), + ) let hex = @src.xdna_write(file) let file2 = @src.xdna_read(hex) assert_eq(file2.n_records(), 1) @@ -653,7 +716,9 @@ test "edge_case_single_base_hex_round_trip" { ///| test "edge_case_empty_sequence" { let file = @src.XdnaFile::new() - file.add_record(@src.XdnaRecord::new("empty", "", @src.XdnaSeqType::from_int(0), "")) + file.add_record( + @src.XdnaRecord::new("empty", "", @src.XdnaSeqType::from_int(0), ""), + ) let bytes = @src.xdna_to_bytes(file) let file2 = @src.xdna_from_bytes(bytes) assert_eq(file2.n_records(), 1) diff --git a/test/moonbit/zinbwave_test.mbt b/test/moonbit/zinbwave_test.mbt new file mode 100644 index 00000000..04f7db8a --- /dev/null +++ b/test/moonbit/zinbwave_test.mbt @@ -0,0 +1,836 @@ +// Tests for the Bioconductor zinbwave-inspired ZINB factor model. + +///| +fn zinbwave_test_close( + actual : Double, + expected : Double, + tolerance : Double, +) -> Unit { + if (actual - expected).abs() > tolerance { + abort( + "zinbwave value " + + actual.to_string() + + " differs from " + + expected.to_string(), + ) + } +} + +///| +fn zinbwave_test_config() -> @src.ZinbWaveConfig { + @src.ZinbWaveConfig::create( + n_factors=2, + max_iterations=20, + inner_iterations=1, + tolerance=1.0e-4, + ridge=0.01, + ) catch { + _ => abort("test configuration should be valid") + } +} + +///| +fn zinbwave_test_model() -> @src.ZinbWaveModel { + let (counts, genes, cells, batch) = @src.zinbwave_example_data() + let config = zinbwave_test_config() + @src.zinbwave_fit( + counts, + config~, + cell_covariates=batch, + gene_names=genes, + cell_names=cells, + ) catch { + _ => abort("example zinbwave fit should succeed") + } +} + +///| +fn zinbwave_test_matrix( + rows : Int, + columns : Int, + value : Double, +) -> Array[Array[Double]] { + let matrix : Array[Array[Double]] = [] + for _ in 0.. @src.SingleCellExperiment { + let (counts, genes, cells, _) = @src.zinbwave_example_data() + @src.SingleCellExperiment::new(counts, genes, cells) +} + +///| +test "zinbwave: default configuration exposes model controls" { + let config = @src.ZinbWaveConfig::default() + assert_eq(config.n_factors, 2) + assert_eq(config.max_iterations, 50) + assert_eq(config.inner_iterations, 2) + assert_eq(config.tolerance, 1.0e-5) + assert_eq(config.dispersion_shrinkage, 5.0) +} + +///| +test "zinbwave: custom configuration preserves parameters" { + let config = @src.ZinbWaveConfig::create( + n_factors=1, + max_iterations=12, + inner_iterations=3, + tolerance=1.0e-6, + ridge=0.2, + dispersion_shrinkage=8.0, + minimum_mean=1.0e-7, + minimum_probability=1.0e-5, + ) catch { + _ => abort("custom configuration should be valid") + } + assert_eq(config.n_factors, 1) + assert_eq(config.max_iterations, 12) + assert_eq(config.inner_iterations, 3) + assert_eq(config.ridge, 0.2) + assert_eq(config.minimum_probability, 1.0e-5) +} + +///| +test "zinbwave: configuration rejects invalid factor and iteration counts" { + let factors = try { + ignore(@src.ZinbWaveConfig::create(n_factors=-1)) + false + } catch { + ZinbWaveError(_) => true + } + let iterations = try { + ignore(@src.ZinbWaveConfig::create(max_iterations=0)) + false + } catch { + ZinbWaveError(_) => true + } + let inner = try { + ignore(@src.ZinbWaveConfig::create(inner_iterations=0)) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(factors) + assert_true(iterations) + assert_true(inner) +} + +///| +test "zinbwave: configuration rejects invalid numeric controls" { + let tolerance = try { + ignore(@src.ZinbWaveConfig::create(tolerance=0.0)) + false + } catch { + ZinbWaveError(_) => true + } + let ridge = try { + ignore(@src.ZinbWaveConfig::create(ridge=-1.0)) + false + } catch { + ZinbWaveError(_) => true + } + let shrinkage = try { + ignore(@src.ZinbWaveConfig::create(dispersion_shrinkage=-1.0)) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(tolerance) + assert_true(ridge) + assert_true(shrinkage) +} + +///| +test "zinbwave: configuration rejects invalid probability bounds" { + let mean = try { + ignore(@src.ZinbWaveConfig::create(minimum_mean=0.0)) + false + } catch { + ZinbWaveError(_) => true + } + let probability = try { + ignore(@src.ZinbWaveConfig::create(minimum_probability=0.5)) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(mean) + assert_true(probability) +} + +///| +test "zinbwave: example data uses gene by cell orientation" { + let (counts, genes, cells, batch) = @src.zinbwave_example_data() + assert_eq(counts.length(), 6) + assert_eq(counts[0].length(), 8) + assert_eq(genes.length(), 6) + assert_eq(cells.length(), 8) + assert_eq(batch.length(), 8) + assert_eq(batch[0].length(), 1) +} + +///| +test "zinbwave: fit dimensions match the source matrix" { + let model = zinbwave_test_model() + assert_eq(model.n_genes(), 6) + assert_eq(model.n_cells(), 8) + assert_eq(model.n_factors(), 2) + assert_eq(model.means.length(), 6) + assert_eq(model.means[0].length(), 8) + assert_eq(model.factors.length(), 8) + assert_eq(model.factors[0].length(), 2) +} + +///| +test "zinbwave: custom identifiers are retained" { + let model = zinbwave_test_model() + assert_eq(model.gene_names[0], "A_marker_1") + assert_eq(model.gene_names[5], "dropout_gene") + assert_eq(model.cell_names[0], "A1") + assert_eq(model.cell_names[7], "B4") +} + +///| +test "zinbwave: omitted identifiers are generated deterministically" { + let config = @src.ZinbWaveConfig::create(n_factors=0, max_iterations=2) catch { + _ => abort("configuration should succeed") + } + let model = @src.zinbwave_fit([[4.0, 3.0], [2.0, 1.0]], config~) catch { + _ => abort("small fit should succeed") + } + assert_eq(model.gene_names, ["Gene1", "Gene2"]) + assert_eq(model.cell_names, ["Cell1", "Cell2"]) +} + +///| +test "zinbwave: cell and gene designs include intercepts" { + let model = zinbwave_test_model() + assert_eq(model.cell_design.length(), 8) + assert_eq(model.cell_design[0].length(), 2) + assert_eq(model.cell_design[0][0], 1.0) + assert_eq(model.gene_design.length(), 6) + assert_eq(model.gene_design[0], [1.0]) +} + +///| +test "zinbwave: latent factors are centered and scaled" { + let model = zinbwave_test_model() + for factor in 0.. + model.diagnostics.initial_log_likelihood, + ) +} + +///| +test "zinbwave: diagnostics track every outer iteration" { + let model = zinbwave_test_model() + assert_true(model.diagnostics.iterations > 0) + assert_eq( + model.diagnostics.log_likelihoods.length(), + model.diagnostics.iterations + 1, + ) + assert_eq( + model.diagnostics.final_log_likelihood, + model.diagnostics.log_likelihoods[model.diagnostics.log_likelihoods.length() - + 1], + ) +} + +///| +test "zinbwave: fitted means are finite and positive" { + let model = zinbwave_test_model() + for row in model.means { + for value in row { + assert_true(value > 0.0) + assert_true(value.abs() < 1.0e300) + } + } +} + +///| +test "zinbwave: zero probabilities are strictly bounded" { + let model = zinbwave_test_model() + for row in model.zero_probabilities { + for value in row { + assert_true(value > 0.0) + assert_true(value < 1.0) + } + } + assert_true(model.mean_zero_probability() > 0.0) + assert_true(model.mean_zero_probability() < 1.0) +} + +///| +test "zinbwave: gene dispersions are finite and positive" { + let model = zinbwave_test_model() + assert_eq(model.dispersions.length(), model.n_genes()) + for value in model.dispersions { + assert_true(value >= 1.0e-4) + assert_true(value <= 100.0) + } +} + +///| +test "zinbwave: observation weights are positive and at most one" { + let model = zinbwave_test_model() + for row in model.observational_weights { + for value in row { + assert_true(value > 0.0) + assert_true(value <= 1.0) + } + } +} + +///| +test "zinbwave: positive counts receive unit observation weight" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let model = zinbwave_test_model() + for gene in 0.. 0.0 { + assert_eq(model.observational_weights[gene][cell], 1.0) + } + } + } +} + +///| +test "zinbwave: dropout-like zero receives downweighted NB responsibility" { + let model = zinbwave_test_model() + assert_true(model.observational_weights[5][0] < 0.5) + assert_true(model.zero_probabilities[5][0] > 0.5) +} + +///| +test "zinbwave: deviance residuals are finite with source dimensions" { + let model = zinbwave_test_model() + assert_eq(model.deviance_residuals.length(), 6) + assert_eq(model.deviance_residuals[0].length(), 8) + for row in model.deviance_residuals { + for value in row { + assert_true(value.abs() < 1.0e300) + } + } +} + +///| +test "zinbwave: normalized values remove library offsets" { + let model = zinbwave_test_model() + let normalized = model.normalized_values() + assert_eq(normalized.length(), model.n_genes()) + assert_eq(normalized[0].length(), model.n_cells()) + for row in normalized { + for value in row { + assert_true(value > 0.0) + } + } +} + +///| +test "zinbwave: imputation preserves observed positive counts" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let model = zinbwave_test_model() + let imputed = model.impute_zeros(counts) catch { + _ => abort("imputation should succeed") + } + for gene in 0.. 0.0 { + assert_eq(imputed[gene][cell], counts[gene][cell]) + } + } + } +} + +///| +test "zinbwave: imputation fills structural zeros with non-negative signal" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let model = zinbwave_test_model() + let imputed = model.impute_zeros(counts) catch { + _ => abort("imputation should succeed") + } + assert_true(imputed[5][0] > 0.0) + assert_true(imputed[5][0] < model.means[5][0]) +} + +///| +test "zinbwave: summary reports shape and latent dimension" { + let summary = zinbwave_test_model().summary() + assert_true(summary.contains("6 genes x 8 cells")) + assert_true(summary.contains("K=2")) + assert_true(summary.contains("logLik=")) +} + +///| +test "zinbwave: information criteria are finite" { + let model = zinbwave_test_model() + assert_true(model.n_parameters() > 0) + assert_true(model.aic().abs() < 1.0e300) + assert_true(model.bic().abs() < 1.0e300) + assert_true(model.bic() > model.aic()) +} + +///| +test "zinbwave: expected mean query handles bounds" { + let model = zinbwave_test_model() + match model.expected_mean(0, 0) { + Some(value) => assert_eq(value, model.means[0][0]) + None => abort("valid fitted mean should exist") + } + assert_true(model.expected_mean(-1, 0) is None) + assert_true(model.expected_mean(0, 8) is None) +} + +///| +test "zinbwave: zero probability query handles bounds" { + let model = zinbwave_test_model() + match model.zero_probability_at(5, 7) { + Some(value) => assert_eq(value, model.zero_probabilities[5][7]) + None => abort("valid zero probability should exist") + } + assert_true(model.zero_probability_at(6, 0) is None) +} + +///| +test "zinbwave: cell factor query returns an independent copy" { + let model = zinbwave_test_model() + match model.cell_factors(0) { + Some(values) => { + let original = model.factors[0][0] + values[0] = values[0] + 10.0 + assert_eq(model.factors[0][0], original) + } + None => abort("valid cell factors should exist") + } + assert_true(model.cell_factors(8) is None) +} + +///| +test "zinbwave: gene summary reports observed and fitted zeros" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let model = zinbwave_test_model() + match model.gene_summary(counts, 5) { + Some(summary) => { + assert_eq(summary.gene_name, "dropout_gene") + assert_eq(summary.observed_zero_fraction, 0.5) + assert_true(summary.fitted_zero_fraction > 0.0) + assert_true(summary.fitted_zero_fraction < 1.0) + assert_true(summary.inverse_dispersion > 0.0) + } + None => abort("valid gene summary should exist") + } +} + +///| +test "zinbwave: gene summary rejects invalid dimensions and indices" { + let model = zinbwave_test_model() + assert_true(model.gene_summary([[1.0, 2.0]], 0) is None) + let (counts, _, _, _) = @src.zinbwave_example_data() + assert_true(model.gene_summary(counts, -1) is None) + assert_true(model.gene_summary(counts, 6) is None) +} + +///| +test "zinbwave: zero-factor model remains fully operational" { + let (counts, genes, cells, _) = @src.zinbwave_example_data() + let config = @src.ZinbWaveConfig::create( + n_factors=0, + max_iterations=5, + inner_iterations=1, + ) catch { + _ => abort("zero-factor configuration should succeed") + } + let model = @src.zinbwave_fit( + counts, + config~, + gene_names=genes, + cell_names=cells, + ) catch { + _ => abort("zero-factor fit should succeed") + } + assert_eq(model.n_factors(), 0) + assert_eq(model.factors.length(), 8) + assert_eq(model.factors[0].length(), 0) + assert_true(model.log_likelihood.abs() < 1.0e300) +} + +///| +test "zinbwave: explicit mean offsets are retained" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let offsets = zinbwave_test_matrix(6, 8, 0.25) + let config = @src.ZinbWaveConfig::create(n_factors=1, max_iterations=3) catch { + _ => abort("configuration should succeed") + } + let model = @src.zinbwave_fit(counts, config~, mean_offsets=offsets) catch { + _ => abort("offset fit should succeed") + } + assert_eq(model.mean_offsets[0][0], 0.25) + assert_eq(model.mean_offsets[5][7], 0.25) +} + +///| +test "zinbwave: explicit zero offsets are retained" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let offsets = zinbwave_test_matrix(6, 8, -0.5) + let config = @src.ZinbWaveConfig::create(n_factors=1, max_iterations=3) catch { + _ => abort("configuration should succeed") + } + let model = @src.zinbwave_fit(counts, config~, zero_offsets=offsets) catch { + _ => abort("zero-offset fit should succeed") + } + assert_eq(model.zero_offsets[0][0], -0.5) + assert_eq(model.zero_offsets[5][7], -0.5) +} + +///| +test "zinbwave: gene-level covariates are fitted with intercept" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let gene_covariates = [[0.30], [0.35], [0.40], [0.45], [0.50], [0.55]] + let config = @src.ZinbWaveConfig::create(n_factors=1, max_iterations=4) catch { + _ => abort("configuration should succeed") + } + let model = @src.zinbwave_fit(counts, config~, gene_covariates~) catch { + _ => abort("gene covariate fit should succeed") + } + assert_eq(model.gene_design[0], [1.0, 0.30]) + assert_eq(model.gamma_mu.length(), 8) + assert_eq(model.gamma_mu[0].length(), 2) +} + +///| +test "zinbwave: empty count matrix is rejected" { + let rejected = try { + ignore(@src.zinbwave_fit([])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: one-cell count matrix is rejected" { + let rejected = try { + ignore(@src.zinbwave_fit([[1.0], [2.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: ragged count matrix is rejected" { + let rejected = try { + ignore(@src.zinbwave_fit([[1.0, 2.0], [3.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: negative and non-finite counts are rejected" { + let negative = try { + ignore(@src.zinbwave_fit([[1.0, -1.0], [2.0, 3.0]])) + false + } catch { + ZinbWaveError(_) => true + } + let non_finite = try { + ignore(@src.zinbwave_fit([[1.0, @double.not_a_number], [2.0, 3.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(negative) + assert_true(non_finite) +} + +///| +test "zinbwave: fractional counts are rejected" { + let rejected = try { + ignore(@src.zinbwave_fit([[1.5, 2.0], [3.0, 4.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: empty cell libraries are rejected for automatic offsets" { + let rejected = try { + ignore(@src.zinbwave_fit([[1.0, 0.0], [2.0, 0.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: excessive factor count is rejected" { + let config = @src.ZinbWaveConfig::create(n_factors=2) catch { + _ => abort("configuration itself should succeed") + } + let rejected = try { + ignore(@src.zinbwave_fit([[2.0, 1.0], [1.0, 2.0]], config~)) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: malformed cell covariates are rejected" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let wrong_rows = try { + ignore(@src.zinbwave_fit(counts, cell_covariates=[[0.0], [1.0]])) + false + } catch { + ZinbWaveError(_) => true + } + let ragged = try { + ignore( + @src.zinbwave_fit(counts, cell_covariates=[ + [0.0], + [0.0, 1.0], + [0.0], + [0.0], + [1.0], + [1.0], + [1.0], + [1.0], + ]), + ) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(wrong_rows) + assert_true(ragged) +} + +///| +test "zinbwave: malformed gene covariates are rejected" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let rejected = try { + ignore(@src.zinbwave_fit(counts, gene_covariates=[[0.0], [1.0], [2.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: malformed offsets are rejected" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let rows = try { + ignore( + @src.zinbwave_fit(counts, mean_offsets=zinbwave_test_matrix(5, 8, 0.0)), + ) + false + } catch { + ZinbWaveError(_) => true + } + let columns = try { + ignore( + @src.zinbwave_fit(counts, zero_offsets=zinbwave_test_matrix(6, 7, 0.0)), + ) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rows) + assert_true(columns) +} + +///| +test "zinbwave: non-finite offsets are rejected" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let offsets = zinbwave_test_matrix(6, 8, 0.0) + offsets[2][3] = @double.not_a_number + let rejected = try { + ignore(@src.zinbwave_fit(counts, mean_offsets=offsets)) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: identifier length mismatch is rejected" { + let (counts, _, _, _) = @src.zinbwave_example_data() + let genes = try { + ignore(@src.zinbwave_fit(counts, gene_names=["A"])) + false + } catch { + ZinbWaveError(_) => true + } + let cells = try { + ignore(@src.zinbwave_fit(counts, cell_names=["A"])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(genes) + assert_true(cells) +} + +///| +test "zinbwave: duplicate and empty identifiers are rejected" { + let counts = [[2.0, 1.0], [1.0, 2.0]] + let config = @src.ZinbWaveConfig::create(n_factors=0) catch { + _ => abort("configuration should succeed") + } + let duplicate = try { + ignore(@src.zinbwave_fit(counts, config~, gene_names=["same", "same"])) + false + } catch { + ZinbWaveError(_) => true + } + let empty = try { + ignore(@src.zinbwave_fit(counts, config~, cell_names=["", "cell"])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(duplicate) + assert_true(empty) +} + +///| +test "zinbwave: imputation rejects matrix shape mismatch" { + let model = zinbwave_test_model() + let rejected = try { + ignore(model.impute_zeros([[1.0, 2.0], [2.0, 1.0]])) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +} + +///| +test "zinbwave: SCE integration adds assays and reduced dimensions" { + let (_, _, _, batch) = @src.zinbwave_example_data() + let config = zinbwave_test_config() + let output = @src.zinbwave_sce( + zinbwave_test_sce(), + config~, + cell_covariates=batch, + ) catch { + _ => abort("SCE integration should succeed") + } + assert_eq( + @src.sce_get_assay(output.experiment, "zinbwave_weights").length(), + 6, + ) + assert_eq( + @src.sce_get_assay(output.experiment, "zinbwave_residuals").length(), + 6, + ) + assert_eq( + @src.sce_get_assay(output.experiment, "zinbwave_normalized").length(), + 6, + ) + assert_eq( + @src.sce_get_assay(output.experiment, "zinbwave_imputed").length(), + 6, + ) + assert_eq(@src.sce_get_reduced_dim(output.experiment, "zinbwave").length(), 8) +} + +///| +test "zinbwave: SCE integration preserves source object" { + let source = zinbwave_test_sce() + let config = zinbwave_test_config() + ignore(@src.zinbwave_sce(source, config~)) catch { + _ => abort("SCE integration should succeed") + } + assert_eq(@src.sce_get_assay(source, "zinbwave_weights").length(), 0) + assert_eq(@src.sce_get_reduced_dim(source, "zinbwave").length(), 0) +} + +///| +test "zinbwave: SCE integration supports custom reduced dimension name" { + let config = zinbwave_test_config() + let output = @src.zinbwave_sce( + zinbwave_test_sce(), + reduced_dim_name="ZINB", + config~, + ) catch { + _ => abort("custom reduced dimension should succeed") + } + assert_eq(@src.sce_get_reduced_dim(output.experiment, "ZINB").length(), 8) + assert_eq(@src.sce_get_reduced_dim(output.experiment, "zinbwave").length(), 0) +} + +///| +test "zinbwave: SCE integration retains row and cell identifiers" { + let config = zinbwave_test_config() + let output = @src.zinbwave_sce(zinbwave_test_sce(), config~) catch { + _ => abort("SCE integration should succeed") + } + assert_eq(output.model.gene_names[0], "A_marker_1") + assert_eq(output.model.cell_names[7], "B4") + assert_eq(output.experiment.row_names[5], "dropout_gene") +} + +///| +test "zinbwave: SCE integration rejects missing assay" { + let config = zinbwave_test_config() + let rejected = try { + ignore( + @src.zinbwave_sce(zinbwave_test_sce(), input_assay="missing", config~), + ) + false + } catch { + ZinbWaveError(_) => true + } + assert_true(rejected) +}