Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Data Processing Pipeline for Small RNA Analysis in HtRNAPan

Introduction

HtRNAPan (Herbal tRNA Panorama), an integrated platform dedicated to the panoramic analysis and functional annotation of tRNAs in herbal medicines, by integrating multi-omics data with AI-powered prediction models, along with reported tRNA sequences, modifications, and modification enzyme information. This database supports the large-scale, systematic annotation of tRNA sequences, modification sites and probabilities, tsRNAs and their targets, and modification enzymes in herbal medicines. HtRNAPan is designed to provide researchers in horticulture, plant science, and pharmaceutical development with an efficient and comprehensive analytical platform, facilitating innovative exploration of tRNAs and their derivatives in both fundamental and applied research.

cite from: HtRNAPan: A comprehensive platform for tRNA landscape analysis and functional annotation in herbal medicines

Pipeline Workflow

Version: 1.0

pipline

1. Data Download and Cleaning

Retrieve datasets from the following sources:

  • NCBI

  • Modomics

    For Modomics data, we downloaded the original HTML files and performed data scraping using the Python script in the 01.data_scrape directory as follows.

    #protein 
    python 01.data_scrape/01.modifications/run.py -i 01.data_scrape/01.modifications/test/1.html -o ./1.json
    
    #modification 
    python 01.data_scrape/02.proteins/run.py -i 01.data_scrape/02.proteins/test/1.html -o ./1.json

2. tRNA Prediction

Use tRNAscan-SE to predict tRNA from Genomes.

tRNAscan-SE  -E -o ${genome}.tRNA -f ${genome}.structure --thread 16 ${genome} 

3. tRNA 2D and 3D Structure Prediction

Predict 2D and 3D Structures Using the VfoldPipeline_alone.

4. tRNA Modification Prediction

After downloading modified and unmodified tRNA sequences from Modomics, first use a script to extract the positions and types of modifications, then classify them according to amino acid types. Next, perform data alignment between the sequence to be predicted and sequences of the same amino acid type, and extract the modification ratios.

# total pipline
bash 03.tRNA-mod/pipeline.sh

# Only One amino acid
bash 03.tRNA-mod/pipeline.sh --aa Gly

5. Analysis of Modification-Related Isozymes

Predict modified-related isozymes using miniprot.

miniprot --gff -Iut50 ${genome} ${pep} > miniprot.gff
grep -v '^#' miniprot.gff|gffread -  -g ${genome} -y ${species}.proteins.fa

6. tsRNA Mining and Target Gene Prediction

Use blastn to align the miRNA-Seq data to the tRNA sequences, and then use a script to organize the tsRNA data.

makeblastdb -in tRNA.fa -dbtype nucl -out tRNA_db
blastn -query miRNA_R1.fasta -db tRNA_db -out R1_vs_tRNA.blastout -outfmt 6  -a 32  -evalue 1e-5
blastn -query miRNA_R2.fasta -db tRNA_db -out R1_vs_tRNA.blastout -outfmt 6  -a 32  -evalue 1e-5 
perl 04.tsRNA/tsRNA-pipeline.pl miRNA tRNA.fa

RNAhybrid -c -p 0.05 -s 3utr_human -t 3utr_sequences.fa -f 2,7 -e -30 -b 1  -q tsRNA.fa >out

About

herbal tRNA database

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages