From 912c2190b5313942e46bf982e03fad9d8bc1baea Mon Sep 17 00:00:00 2001 From: ClareRobin Date: Fri, 21 Aug 2026 16:53:13 +0800 Subject: [PATCH 1/2] readthedocs updates (content updates) --- docs/source/cmd.rst | 8 +- docs/source/index.rst | 2 +- docs/source/preparation.rst | 40 -------- docs/source/quickstart.rst | 161 +++++++++++++++++++----------- docs/source/running_demo_data.rst | 107 ++++++++++++++++++++ 5 files changed, 212 insertions(+), 106 deletions(-) delete mode 100644 docs/source/preparation.rst create mode 100644 docs/source/running_demo_data.rst diff --git a/docs/source/cmd.rst b/docs/source/cmd.rst index 0c8e37a..4518301 100644 --- a/docs/source/cmd.rst +++ b/docs/source/cmd.rst @@ -10,16 +10,16 @@ We provide 2 main scripts to run the analysis of differential RNA modifications * Input -Output files from ``nanopolish eventalgin``. Please refer to :ref:`Data preparation ` for the full Nanopolish command. +Output files from ``nanopolish eventalign`` or ``f5c eventalign``. Please refer to :ref:`Quickstart ` for the full commands. ================================= ========== =================== ============================================================================================================ Argument name Required Default value Description ================================= ========== =================== ============================================================================================================ ---eventalign=FILE Yes NA Eventalign filepath, the output from nanopolish. +--eventalign=FILE Yes NA Eventalign filepath, the output from nanopolish or f5c eventalign. --out_dir=DIR Yes NA Output directory. --gtf_or_gff=FILE No NA GTF or GFF file path used for mapping transcriptomic to genomic coordinates. --transcript_fasta=FILE No NA Transcript FASTA path used for mapping transcriptomic to genomic coordinates. ---skip_eventalign_indexing No False To skip indexing the eventalign nanopolish output. +--skip_eventalign_indexing No False To skip indexing the eventalign output. --genome No False To run on Genomic coordinates. Without this argument, the program will run on transcriptomic coordinates. --kmer_source=STR No reference_kmer Which kmer column to use from the eventalign file: ``reference_kmer`` (default, for transcriptome alignments) or ``model_kmer`` (for genome alignments, which contain reverse-strand reads). --n_processes=NUM No 1 Number of processes to run. @@ -33,7 +33,7 @@ Argument name Required Default value Descriptio ====================== ============== =============================================================================================================================================================== File name File type Description ====================== ============== =============================================================================================================================================================== -eventalign.index csv File index indicating the position in the `eventalign.txt` file (the output of nanopolish eventalign) where the segmentation information of each read index is stored, allowing a random access. +eventalign.index csv File index indicating the position in the `eventalign.txt` file (the output of nanopolish or f5c eventalign) where the segmentation information of each read index is stored, allowing a random access. data.json json Intensity level mean for each position. data.index csv File index indicating the position in the `data.json` file where the intensity level means across positions of each gene is stored, allowing a random access. data.log txt Gene ids being processed. diff --git a/docs/source/index.rst b/docs/source/index.rst index 6ed5fa7..28cfc21 100644 --- a/docs/source/index.rst +++ b/docs/source/index.rst @@ -26,9 +26,9 @@ Contents installation quickstart + running_demo_data outputtable configuration - preparation data cmd citing diff --git a/docs/source/preparation.rst b/docs/source/preparation.rst deleted file mode 100644 index ab4456c..0000000 --- a/docs/source/preparation.rst +++ /dev/null @@ -1,40 +0,0 @@ -.. _preparation: - -Data preparation from raw reads -=================================== - -1. After obtaining the raw signal files (POD5, or FAST5 for older runs), the first step is to basecall them. Below is an example using `Dorado `_, Oxford Nanopore's current basecaller (it replaces the older Guppy/Albacore basecallers and supports both RNA002 and RNA004 chemistries — select the model that matches your chemistry). You can find more detail about basecalling at `Oxford Nanopore Technologies `_:: - - dorado basecaller --emit-fastq > - - For RNA004 data use an ``rna004`` model (e.g. ``rna004_130bps_sup@v5.1.0``); for RNA002 data use an ``rna002`` model. See the `Dorado documentation `_ for the available models and options. - -2. Align the basecalled reads with `minimap2 `_. xPore supports both **transcriptome** and **genome** alignments — choose one depending on which coordinate system you want in the output. - - **Transcriptome alignment** (align to a transcriptome reference). xPore reports transcriptomic coordinates, or genomic coordinates if you also pass ``--genome`` together with ``--gtf_or_gff`` and ``--transcript_fasta`` to ``xpore dataprep``:: - - minimap2 -ax map-ont -uf -t 3 --secondary=no > 2>> - samtools view -Sb | samtools sort -o - &>> - samtools index &>> - - **Genome alignment** (align directly to a genome reference; use spliced alignment so that reads spanning introns map correctly). Genome alignments contain reverse-strand reads, so run ``xpore dataprep`` with ``--kmer_source model_kmer`` for these (see :ref:`Command line arguments `):: - - minimap2 -ax splice -uf -k14 -t 3 --secondary=no > 2>> - samtools view -Sb | samtools sort -o - &>> - samtools index &>> - -3. Resquiggle (align the raw signal to the reference) to produce the eventalign file. We recommend `f5c `_, an optimised, CPU/GPU-accelerated re-implementation of ``nanopolish eventalign`` that produces equivalent output much faster on large datasets:: - - # index the raw signal: use -d for FAST5, or --slow5 for SLOW5/BLOW5 - f5c index -d - f5c eventalign --reads \ - --bam \ - --genome \ - --rna \ - --signal-index \ - --scale-events \ - --threads 32 > - - For **RNA004** data, add ``--kmer-model ``: recent versions of f5c auto-select the 9-mer model for RNA004, which xPore cannot use — xPore requires the 5-mer model (see `xPore issue #215 `_). - - ``nanopolish eventalign`` can be used instead with the same arguments. Note that the ``--genome`` argument here refers to the **alignment reference** (the transcriptome or genome FASTA used in step 2), not xPore's ``--genome`` flag. diff --git a/docs/source/quickstart.rst b/docs/source/quickstart.rst index 4d5ca44..039aecf 100644 --- a/docs/source/quickstart.rst +++ b/docs/source/quickstart.rst @@ -4,9 +4,11 @@ Quickstart - Detection of differential RNA modifications ========================================================= .. note:: - **Updates in xPore v2.2:** xPore is now compatible with genome alignments and RNA004 data — see the table below and the :ref:`Data preparation from raw reads ` section for more information. + **Updates in xPore v2.2:** xPore is now compatible with genome alignments and RNA004 data — see the table below and the steps for more information. -xPore is now compatible with genome alignment. See below for the minimal commands to run xPore on transcriptome- or genome-aligned data: +This page walks through the full pipeline, from raw nanopore signal (POD5, or FAST5 for older runs) through basecalling, alignment, resquiggling, ``xpore dataprep`` and ``xpore diffmod``, to a ranked table of differentially modified sites. If you would rather try xPore on a ready-made dataset first, see :ref:`Running demo data `. + +xPore is compatible with both transcriptome- and genome-aligned data. See below for the minimal commands to run xPore on each: .. list-table:: :header-rows: 1 @@ -38,89 +40,125 @@ xPore is now compatible with genome alignment. See below for the minimal command - | ``xpore diffmod`` | ``--config `` -Running Example Demo Data -------------------------- +1. Basecalling +-------------- + +Typically MinKNOW runs `Dorado `_ and basecalls live during acquisition. + +If you need to (re)basecall offline — e.g. a different model, a newer Dorado version, or basecalling wasn't enabled during the run — run Dorado directly on the raw signal files (POD5, or FAST5 for older runs). Dorado replaces the older Guppy/Albacore basecallers and supports both RNA002 and RNA004 chemistries — select the model that matches your chemistry. You can find more detail about basecalling at `Oxford Nanopore Technologies `_:: + + dorado basecaller --emit-fastq > + +For RNA004 data use an ``rna004`` model (e.g. ``rna004_130bps_sup@v5.1.0``); for RNA002 data use an ``rna002`` model. See the `Dorado documentation `_ for the available models and options. + +2. Alignment +------------ + +Align the basecalled reads with `minimap2 `_. xPore supports both **transcriptome** and **genome** alignments — choose one depending on which coordinate system you want in the output. -Download and extract the demo dataset from our `zenodo `_:: +**Transcriptome alignment** (align to a transcriptome reference). xPore reports transcriptomic coordinates, or genomic coordinates if you also pass ``--genome`` together with ``--gtf_or_gff`` and ``--transcript_fasta`` to ``xpore dataprep``:: - wget https://zenodo.org/record/5162402/files/demo.tar.gz - tar -xvf demo.tar.gz + minimap2 -ax map-ont -uf -t 3 --secondary=no > 2>> + samtools view -Sb | samtools sort -o - &>> + samtools index &>> -After extraction, you will find:: - - demo - |-- Hek293T_config.yml # configuration file - |-- data - |-- HEK293T-METTL3-KO-rep1 # dataset dir - |-- HEK293T-WT-rep1 # dataset dir - |-- demo.gtf # GTF (general transfer format) file for transcript-to-gene mapping - |-- demo.fa # transcriptome reference FASTA file for transcript-to-gene mapping +**Genome alignment** (align directly to a genome reference; use spliced alignment so that reads spanning introns map correctly). Genome alignments contain reverse-strand reads, so run ``xpore dataprep`` with ``--kmer_source model_kmer`` for these (see :ref:`Command line arguments `):: -Each dataset under the ``data`` directory contains the following directories: + minimap2 -ax splice -uf -k14 -t 3 --secondary=no > 2>> + samtools view -Sb | samtools sort -o - &>> + samtools index &>> -* ``fast5`` : Raw signal, FAST5 files -* ``fastq`` : Basecalled reads, FASTQ file -* ``bamtx`` : Transcriptome-aligned sequence, BAM file -* ``nanopolish``: Eventalign files obtained from `nanopolish eventalign `_ +3. Resquiggling +---------------- -Note that the FAST5, FASTQ and BAM files are required to obtain the eventalign file with Nanopolish, xPore only requires the eventalign file. See our :ref:`Data preparation page ` for details to obtain the eventalign file from raw reads. +Resquiggle (align the raw signal to the reference) to produce the eventalign file. We recommend `f5c `_, an optimised, CPU/GPU-accelerated re-implementation of ``nanopolish eventalign`` that produces equivalent output much faster on large datasets:: -1. Preprocess the data for each data set using ``xpore dataprep``. Note that the ``--gtf_or_gff`` and ``--transcript_fasta`` arguments are required to map transcriptomic to genomic coordinates when the ``--genome`` option is chosen, so that xPore can run based on genome coordinates. However, GTF is the recommended option. If GFF is the only file available, please note that the GFF file works with GENCODE or ENSEMBL FASTA files, but not UCSC FASTA files. We plan to remove the requirement of FASTA files in a future release.(This step will take approximately 5h for 1 million reads):: + # index the raw signal: use -d for FAST5, or --slow5 for SLOW5/BLOW5 + f5c index -d + f5c eventalign --reads \ + --bam \ + --genome \ + --rna \ + --signal-index \ + --scale-events \ + --threads 32 > - # Within each dataset directory i.e. demo/data/HEK293T-METTL3-KO-rep1 and demo/data/HEK293T-WT-rep1, run - xpore dataprep \ - --eventalign nanopolish/eventalign.txt \ - --gtf_or_gff ../../demo.gtf \ - --transcript_fasta ../../demo.fa \ - --out_dir dataprep \ - --genome +For **RNA004** data, add ``--kmer-model ``: recent versions of f5c auto-select the 9-mer model for RNA004, which xPore cannot use — xPore requires the 5-mer model (see `xPore issue #215 `_). -The output files are stored under ``dataprep`` in each dataset directory: +``nanopolish eventalign`` can be used instead with the same arguments. Note that the ``--genome`` argument here refers to the **alignment reference** (the transcriptome or genome FASTA used in step 2). -* ``eventalign.index`` : Index file to access ``eventalign.txt``, the output from nanopolish eventalign -* ``data.json`` : Preprocessed data for ``xpore-diffmod`` +4. Preprocess with ``xpore dataprep`` +-------------------------------------- + +Preprocess the eventalign file with ``xpore dataprep``, using the command below that matches your alignment mode from step 2. This step will take approximately 5h for 1 million reads. + +.. list-table:: + :header-rows: 1 + :widths: 33 34 33 + + * - **Transcriptome alignment (output transcriptome coordinates)** + - **Transcriptome alignment (output genome coordinates)** + - **Genome alignment (output genome coordinates)** + * - | ``xpore dataprep`` + | ``--eventalign `` + | ``--out_dir `` + - | ``xpore dataprep`` + | ``--eventalign `` + | ``--out_dir `` + | ``--genome`` + | ``--transcript_fasta `` + | ``--gtf_or_gff `` + - | ``xpore dataprep`` + | ``--eventalign `` + | ``--out_dir `` + | ``--kmer_source model_kmer`` + +The ``--gtf_or_gff`` and ``--transcript_fasta`` arguments map transcriptomic to genomic coordinates; GTF is the recommended option — if GFF is the only file available, note that it works with GENCODE or ENSEMBL FASTA files, but not UCSC FASTA files. We plan to remove the requirement of FASTA files in a future release. + +The output files are stored under ````: + +* ``eventalign.index`` : Index file to access ``eventalign.txt``, the output from f5c/nanopolish eventalign +* ``data.json`` : Preprocessed data for ``xpore diffmod`` * ``data.index`` : File index of ``data.json`` for random access per gene * ``data.readcount`` : Summary of readcounts per gene * ``data.log`` : Log file -Run ``xpore dataprep -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. +Run this for each of your samples (e.g. each condition/replicate). Run ``xpore dataprep -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. -2. Prepare a ``.yml`` configuration file. With this YAML file, you can specify the information of your design experiment, the data directories, the output directory, and the method options. -In the demo directory, there is an example configuration file ``Hek293T_config.yaml`` available that you can use as a starting template. -Below is how it looks like:: +5. Configuration file +----------------------- - notes: Pairwise comparison without replicates with default parameter setting. +Prepare a ``.yml`` configuration file. With this YAML file, you can specify the information of your design experiment, the data directories, the output directory, and the method options:: data: - KO: - rep1: ./data/HEK293T-METTL3-KO-rep1/dataprep - WT: - rep1: ./data/HEK293T-WT-rep1/dataprep + : + : + : + : - out: ./out # output dir + out: # output dir - # The demo data is RNA002. Since v2.2 the default prior is the RNA004 model, - # so point xpore-diffmod at the bundled RNA002 model to reproduce the demo: - prior: /path/to/xpore/diffmod/RNA002_5mer_model.csv + # Since xPore v2.2 the default unmodified-signal prior is the RNA004 model. + # For RNA002 data, set prior to the bundled RNA002 model instead: + # prior: /path/to/xpore/diffmod/RNA002_5mer_model.csv +See the :ref:`Configuration file page ` for more details. -See the :ref:`Configuration file page ` for more details. Note that since xPore v2.2 the default unmodified-signal prior is the RNA004 model; for RNA002 data (like this demo) set ``prior:`` to the bundled ``RNA002_5mer_model.csv`` as shown above. +6. Run ``xpore diffmod`` +-------------------------- -3. Now that we have the data and the configuration file ready for modelling differential modifications using ``xpore-diffmod``. +Now that we have the data and the configuration file ready, model differential modifications using ``xpore diffmod``:: -:: + xpore diffmod --config - # At the demo directory where the configuration file is, run. - xpore diffmod --config Hek293T_config.yml - -The output files are generated within the ``out`` directory: +The output files are generated within the output directory specified in the config: * ``diffmod.table`` : Result table of differential RNA modification across all tested positions * ``diffmod.log`` : Log file Run ``xpore diffmod -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. -We can rank the significantly differentially modified sites based on ``pval_HEK293T-KO_vs_HEK293T-WT``. The results are shown below.:: +We can rank the significantly differentially modified sites based on the p-value column, e.g. ``pval__vs_``. The results look like:: id position kmer diff_mod_rate_KO_vs_WT pval_KO_vs_WT z_score_KO_vs_WT ... sigma2_unmod sigma2_mod conf_mu_unmod conf_mu_mod mod_assignment t-test ENSG00000114125 141745412 GGACT -0.823318 4.241373e-115 -22.803411 ... 5.925238 18.048687 0.968689 0.195429 lower 1.768910e-19 @@ -129,15 +167,16 @@ We can rank the significantly differentially modified sites based on ``pval_HEK2 ENSG00000159111 47824137 GGACA -0.604056 7.614675e-24 -10.068479 ... 7.164075 4.257725 0.553929 0.353160 lower 2.294337e-10 ENSG00000114125 141745249 GGACT -0.514980 2.779122e-19 -8.977134 ... 5.215243 20.598471 0.954968 0.347174 lower 1.304111e-06 -4. (Optional) We can consider only one modification type per k-mer by finding the majority ``mod_assignment`` of each k-mer. -For example, the majority of the modification means of ``GGACT`` (``mu_mod``) is lower than the non-modification counterpart (``mu_unmod``). -We can filter out those positions whose ``mod_assigment`` values are not in line with those of the majority in order to restrict ourselves with one modification type per kmer in the analysis. -This can be done by running ``xpore postprocessing``. +7. (Optional) Postprocessing +------------------------------ -:: +We can consider only one modification type per k-mer by finding the majority ``mod_assignment`` of each k-mer. +For example, the majority of the modification means of ``GGACT`` (``mu_mod``) is lower than the non-modification counterpart (``mu_unmod``). +We can filter out those positions whose ``mod_assigment`` values are not in line with those of the majority in order to restrict ourselves with one modification type per kmer in the analysis. +This can be done by running ``xpore postprocessing``:: - xpore postprocessing --diffmod_dir out + xpore postprocessing --diffmod_dir -With this command, we will get the final file in which only kmers with their ``mod_assignment`` different from the majority assigment of the corresponding kmer are removed. The output file ``majority_direction_kmer_diffmod.table`` is generated in the ``out`` directtory. You can find more details in our paper. +With this command, we will get the final file in which only kmers with their ``mod_assignment`` different from the majority assigment of the corresponding kmer are removed. The output file ``majority_direction_kmer_diffmod.table`` is generated in the output directory. You can find more details in our paper. Run ``xpore postprocessing -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. diff --git a/docs/source/running_demo_data.rst b/docs/source/running_demo_data.rst new file mode 100644 index 0000000..8d77cd6 --- /dev/null +++ b/docs/source/running_demo_data.rst @@ -0,0 +1,107 @@ +.. _demo: + +Running demo data +=================== + +.. note:: + This demo dataset is **RNA002**, **transcriptome-aligned** data from the original xPore paper. It predates genome alignment support — for the full raw-data-to-diffmod pipeline, including genome alignment, see :ref:`Quickstart `. + +Download and extract the demo dataset from our `zenodo `_:: + + wget https://zenodo.org/record/5162402/files/demo.tar.gz + tar -xvf demo.tar.gz + +After extraction, you will find:: + + demo + |-- Hek293T_config.yml # configuration file + |-- data + |-- HEK293T-METTL3-KO-rep1 # dataset dir + |-- HEK293T-WT-rep1 # dataset dir + |-- demo.gtf # GTF (general transfer format) file for transcript-to-gene mapping + |-- demo.fa # transcriptome reference FASTA file for transcript-to-gene mapping + +Each dataset under the ``data`` directory contains the following directories: + +* ``fast5`` : Raw signal, FAST5 files +* ``fastq`` : Basecalled reads, FASTQ file +* ``bamtx`` : Transcriptome-aligned sequence, BAM file +* ``nanopolish``: Eventalign files obtained from `nanopolish eventalign `_ + +Note that the FAST5, FASTQ and BAM files are required to obtain the eventalign file with Nanopolish, xPore only requires the eventalign file. This demo already provides the eventalign file, so it skips the basecalling, alignment and resquiggling steps — see the :ref:`Quickstart ` page for how to produce an eventalign file from raw reads. + +1. Preprocess the data for each data set using ``xpore dataprep``. Note that the ``--gtf_or_gff`` and ``--transcript_fasta`` arguments are required to map transcriptomic to genomic coordinates when the ``--genome`` option is chosen, so that xPore can run based on genome coordinates. However, GTF is the recommended option. If GFF is the only file available, please note that the GFF file works with GENCODE or ENSEMBL FASTA files, but not UCSC FASTA files. We plan to remove the requirement of FASTA files in a future release. (This step will take approximately 5h for 1 million reads):: + + # Within each dataset directory i.e. demo/data/HEK293T-METTL3-KO-rep1 and demo/data/HEK293T-WT-rep1, run + xpore dataprep \ + --eventalign nanopolish/eventalign.txt \ + --gtf_or_gff ../../demo.gtf \ + --transcript_fasta ../../demo.fa \ + --out_dir dataprep \ + --genome + +The output files are stored under ``dataprep`` in each dataset directory: + +* ``eventalign.index`` : Index file to access ``eventalign.txt``, the output from nanopolish eventalign +* ``data.json`` : Preprocessed data for ``xpore diffmod`` +* ``data.index`` : File index of ``data.json`` for random access per gene +* ``data.readcount`` : Summary of readcounts per gene +* ``data.log`` : Log file + +Run ``xpore dataprep -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. + +2. Prepare a ``.yml`` configuration file. With this YAML file, you can specify the information of your design experiment, the data directories, the output directory, and the method options. +In the demo directory, there is an example configuration file ``Hek293T_config.yaml`` available that you can use as a starting template. +Below is how it looks like:: + + notes: Pairwise comparison without replicates with default parameter setting. + + data: + KO: + rep1: ./data/HEK293T-METTL3-KO-rep1/dataprep + WT: + rep1: ./data/HEK293T-WT-rep1/dataprep + + out: ./out # output dir + + # The demo data is RNA002. Since v2.2 the default prior is the RNA004 model, + # so point xpore-diffmod at the bundled RNA002 model to reproduce the demo: + prior: /path/to/xpore/diffmod/RNA002_5mer_model.csv + +See the :ref:`Configuration file page ` for more details. Note that since xPore v2.2 the default unmodified-signal prior is the RNA004 model; for RNA002 data (like this demo) set ``prior:`` to the bundled ``RNA002_5mer_model.csv`` as shown above. + +3. Now that we have the data and the configuration file ready for modelling differential modifications using ``xpore diffmod``. + +:: + + # At the demo directory where the configuration file is, run. + xpore diffmod --config Hek293T_config.yml + +The output files are generated within the ``out`` directory: + +* ``diffmod.table`` : Result table of differential RNA modification across all tested positions +* ``diffmod.log`` : Log file + +Run ``xpore diffmod -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. + +We can rank the significantly differentially modified sites based on ``pval_HEK293T-KO_vs_HEK293T-WT``. The results are shown below.:: + + id position kmer diff_mod_rate_KO_vs_WT pval_KO_vs_WT z_score_KO_vs_WT ... sigma2_unmod sigma2_mod conf_mu_unmod conf_mu_mod mod_assignment t-test + ENSG00000114125 141745412 GGACT -0.823318 4.241373e-115 -22.803411 ... 5.925238 18.048687 0.968689 0.195429 lower 1.768910e-19 + ENSG00000159111 47824212 GGACT -0.828023 1.103790e-88 -19.965293 ... 2.686549 13.820089 0.644436 0.464059 lower 5.803242e-18 + ENSG00000159111 47824138 GGGAC -0.757891 1.898161e-73 -18.128515 ... 3.965195 9.877299 0.861480 0.359984 lower 9.708552e-08 + ENSG00000159111 47824137 GGACA -0.604056 7.614675e-24 -10.068479 ... 7.164075 4.257725 0.553929 0.353160 lower 2.294337e-10 + ENSG00000114125 141745249 GGACT -0.514980 2.779122e-19 -8.977134 ... 5.215243 20.598471 0.954968 0.347174 lower 1.304111e-06 + +4. (Optional) We can consider only one modification type per k-mer by finding the majority ``mod_assignment`` of each k-mer. +For example, the majority of the modification means of ``GGACT`` (``mu_mod``) is lower than the non-modification counterpart (``mu_unmod``). +We can filter out those positions whose ``mod_assigment`` values are not in line with those of the majority in order to restrict ourselves with one modification type per kmer in the analysis. +This can be done by running ``xpore postprocessing``. + +:: + + xpore postprocessing --diffmod_dir out + +With this command, we will get the final file in which only kmers with their ``mod_assignment`` different from the majority assigment of the corresponding kmer are removed. The output file ``majority_direction_kmer_diffmod.table`` is generated in the ``out`` directtory. You can find more details in our paper. + +Run ``xpore postprocessing -h`` or visit our :ref:`Command line arguments ` to explore the full usage description. From 456b1a6be8b9d540bccee83dc06cd78afaad9d78 Mon Sep 17 00:00:00 2001 From: ClareRobin Date: Mon, 24 Aug 2026 09:32:11 +0800 Subject: [PATCH 2/2] docs update: added f5c release tag (version compatible with genome alignment) --- README.md | 2 +- docs/source/quickstart.rst | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index 7e4cde8..2e0a697 100644 --- a/README.md +++ b/README.md @@ -33,7 +33,7 @@ xPore is described in details in [Pratanwanich et al. *Nat Biotechnol* (2021)](h ### Release History -The current release is xPore v2.2, which adds support for genome-aligned eventalign output (via the new `--kmer_source model_kmer` flag) and RNA004 data (now the default `xpore-diffmod` prior). +The current release is xPore v2.2, which adds support for genome-aligned eventalign output (via the new `--kmer_source model_kmer` flag) and RNA004 data (now the default `xpore-diffmod` prior). For genome-aligned data, use [f5c v1.7](https://github.com/hasindu2008/f5c/releases/tag/v1.7) or later to generate the eventalign file — it is the version tested for enabling genome-aligned RNA with xPore v2.2. Please refer to the github release history for previous releases: https://github.com/GoekeLab/xpore/releases diff --git a/docs/source/quickstart.rst b/docs/source/quickstart.rst index 039aecf..68fb1db 100644 --- a/docs/source/quickstart.rst +++ b/docs/source/quickstart.rst @@ -85,6 +85,8 @@ Resquiggle (align the raw signal to the reference) to produce the eventalign fil For **RNA004** data, add ``--kmer-model ``: recent versions of f5c auto-select the 9-mer model for RNA004, which xPore cannot use — xPore requires the 5-mer model (see `xPore issue #215 `_). +For **genome-aligned** data, use `f5c v1.7 `_ or later — it is the version tested for enabling genome-aligned RNA with xPore v2.2. + ``nanopolish eventalign`` can be used instead with the same arguments. Note that the ``--genome`` argument here refers to the **alignment reference** (the transcriptome or genome FASTA used in step 2). 4. Preprocess with ``xpore dataprep``