Part of the arXiv preprint is based upon a bulk RNA-seq experiment newly performed for the preprint. The data (raw or processed) are not available; only a small subset of summary statistics (the DE test) is disclosed. However, the interpretation of the DE test depends on the methodology, which is in turn described in ways that are occasionally opaque.
If I understand correctly, only one RNA-seq experiment was performed, with two conditions and multiple replicates.
- Size is unclear.
- Libraries ... were sequenced on an Illumina NextSeq 2000 with 75 base-pair paired-end reads... (4.4.5)
- Raw paired-end RNA-Seq reads (2 × 150 bp) (4.5.3), which means 150 base-pair paired-end reads.
- The number of replicates is unclear.
- twelve samples (including Y27632-treated, wild-type, and bead-treated conditions) (4.5.3), which reads as though "bead-treated" is a distinct condition and there are three conditions total. 4.4.3 suggests all are bead-treated.
- yielding a matrix of 6 samples across two conditions (4.5.5) literally means that there were a total of six samples, which is not twelve. For an example of the common usage: if the census has data on three hundred million people across fifty states, it does not mean there are three hundred million people per state.
- The preprocessing analysis section is unclear.
- 4.5.2 provides no method information, except alluding to demultiplexing (which is not explored elsewhere).
- 4.5.3 is malformed:
-- (intended) vs – (in manuscript, possibly a LaTeX artifact). I cannot otherwise speak to the HISAT2 procedure, except to say that dta ("to retain splice junction information for downstream transcriptome assembly") seems excessive for a procedure not actually intended to produce a de novo transcriptome assembly. The nf-core/rnaseq workflow abstains from quantifying the results of HISAT2 outputs altogether "due to the lack of an appropriate option to calculate accurate expression estimates from HISAT2 derived genomic alignments".
- 4.5.4 states that "[t]his workflow provides a reproducible framework for high-throughput RNA-Seq data processing from raw reads through gene-level quantification". It is hard to see what is meant by this, for the following reasons:
- The example of the workflow is not high-throughput.
- Code is not available.
- The description is not sufficient to reconstruct the procedure. The entire HISAT2 workflow only mentions one material variable. Version numbers are not disclosed.
- Various other metrics related to the experiment are not disclosed.
Part of the arXiv preprint is based upon a bulk RNA-seq experiment newly performed for the preprint. The data (raw or processed) are not available; only a small subset of summary statistics (the DE test) is disclosed. However, the interpretation of the DE test depends on the methodology, which is in turn described in ways that are occasionally opaque.
If I understand correctly, only one RNA-seq experiment was performed, with two conditions and multiple replicates.
--(intended) vs – (in manuscript, possibly a LaTeX artifact). I cannot otherwise speak to the HISAT2 procedure, except to say thatdta("to retain splice junction information for downstream transcriptome assembly") seems excessive for a procedure not actually intended to produce a de novo transcriptome assembly. Thenf-core/rnaseqworkflow abstains from quantifying the results of HISAT2 outputs altogether "due to the lack of an appropriate option to calculate accurate expression estimates from HISAT2 derived genomic alignments".