-
Notifications
You must be signed in to change notification settings - Fork 10
Ingest #6
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
Ingest #6
Changes from all commits
Commits
Show all changes
12 commits
Select commit
Hold shift + click to select a range
55e13da
copy zika ingest
0233699
adapt to ebola
5600aa7
update ebola build to pull new ingest data
477979b
refactor: flag intermediate files as temp
e80815f
refactor: serotypes
a17b2f0
docs: ingest directory
fad2af9
edit: build accepts and decompresses zst inputs
5c060a0
docs: included a data provisioning section
a1f36e2
refactor: used explicit paths instead of references to the rules vari…
a157df7
edit: Use a permalink for each script
5a2c5c4
fixup: change dengue to ebola
6ad1636
Pick curl or wget based on availability
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,2 @@ | ||
| strain_id_field: "accession" | ||
| display_strain_field: "strain_original" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,31 @@ | ||
| strain accession date region country division city authors url title | ||
| G5016.1 KR105293 2014-08-16 subsaharan_africa sierra_leone kenema kenema Goba et al https://www.ncbi.nlm.nih.gov/nuccore/KR105293 "Ebola virus epidemiology, transmission, and viral evolution from four months of sequencing in Sierra Leone" | ||
| EM_080193 KR817239 2014-06-29 subsaharan_africa liberia liberia liberia Carroll et al https://www.ncbi.nlm.nih.gov/nuccore/KR817239 Temporal and spatial analysis of the 2014-2015 Ebola virus outbreak in West Africa | ||
| EM_074351 KR817110 2014-07-22 subsaharan_africa liberia lofa lofa Carroll et al https://www.ncbi.nlm.nih.gov/nuccore/KR817110 Temporal and spatial analysis of the 2014-2015 Ebola virus outbreak in West Africa | ||
| LIBR10253 KT725282 2014-09-16 subsaharan_africa liberia margibi margibi Ladner et al https://www.ncbi.nlm.nih.gov/nuccore/KT725282 "Evolution and spread of Ebola virus in Liberia, 2014-2015" | ||
| G3848 KM233110 2014-06-18 subsaharan_africa sierra_leone kailahun kailahun Gire et al https://www.ncbi.nlm.nih.gov/nuccore/KM233110 Genomic surveillance elucidates Ebola virus origin and transmission during the 2014 outbreak | ||
| EM_000027 KR817068 2014-09-01 subsaharan_africa guinea guinea guinea Carroll et al https://www.ncbi.nlm.nih.gov/nuccore/KR817068 Temporal and spatial analysis of the 2014-2015 Ebola virus outbreak in West Africa | ||
| LIBR10047 KT725345 2014-10-01 subsaharan_africa liberia margibi margibi Ladner et al https://www.ncbi.nlm.nih.gov/nuccore/KT725345 "Evolution and spread of Ebola virus in Liberia, 2014-2015" | ||
| G4350.1 KR105242 2014-07-14 subsaharan_africa sierra_leone kenema kenema Goba et al https://www.ncbi.nlm.nih.gov/nuccore/KR105242 "Ebola virus epidemiology, transmission, and viral evolution from four months of sequencing in Sierra Leone" | ||
| J0165 KP759703 2014-11-08 subsaharan_africa sierra_leone western_urban western_urban Tong et al https://www.ncbi.nlm.nih.gov/nuccore/KP759703 Genetic diversity and evolutionary dynamics of Ebola virus in Sierra Leone | ||
| G3926.2 KR105208 2014-06-22 subsaharan_africa sierra_leone kailahun kailahun Goba et al https://www.ncbi.nlm.nih.gov/nuccore/KR105208 "Ebola virus epidemiology, transmission, and viral evolution from four months of sequencing in Sierra Leone" | ||
| EM_080223 KR817241 2014-07-04 subsaharan_africa liberia liberia liberia Carroll et al https://www.ncbi.nlm.nih.gov/nuccore/KR817241 Temporal and spatial analysis of the 2014-2015 Ebola virus outbreak in West Africa | ||
| G4345.1 KR105239 2014-07-18 subsaharan_africa sierra_leone kenema kenema Goba et al https://www.ncbi.nlm.nih.gov/nuccore/KR105239 "Ebola virus epidemiology, transmission, and viral evolution from four months of sequencing in Sierra Leone" | ||
| LIBR10218 KT725284 2014-09-29 subsaharan_africa liberia liberia liberia Ladner et al https://www.ncbi.nlm.nih.gov/nuccore/KT725284 "Evolution and spread of Ebola virus in Liberia, 2014-2015" | ||
| G3807 KM233084 2014-06-15 subsaharan_africa sierra_leone kailahun kailahun Gire et al https://www.ncbi.nlm.nih.gov/nuccore/KM233084 Genomic surveillance elucidates Ebola virus origin and transmission during the 2014 outbreak | ||
| LIBR10224 KT725302 2014-09-26 subsaharan_africa liberia montserrado montserrado Ladner et al https://www.ncbi.nlm.nih.gov/nuccore/KT725302 "Evolution and spread of Ebola virus in Liberia, 2014-2015" | ||
| 13031_EMLH KU296665 2015-05-18 subsaharan_africa sierra_leone western_urban western_urban Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296665 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| G5996.1 KR105337 2014-09-25 subsaharan_africa sierra_leone sierra_leone sierra_leone Goba et al https://www.ncbi.nlm.nih.gov/nuccore/KR105337 "Ebola virus epidemiology, transmission, and viral evolution from four months of sequencing in Sierra Leone" | ||
| DML24708 KT357845 2015-01-28 subsaharan_africa sierra_leone kono kono Smits et al https://www.ncbi.nlm.nih.gov/nuccore/KT357845 "Hypermutated Ebola virus circulating in Magazine Wharf area, Freetown, Sierra Leone" | ||
| 15686_EMLK KU296555 2015-06-30 subsaharan_africa sierra_leone kambia kambia Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296555 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| EM_COY_2015_016617 EM_COY_2015_016617 2015-05-16 subsaharan_africa guinea dubreka dubreka Quick et al https://github.com/nickloman/ebov/ ? | ||
| Makona-UK3 KR025228 2015-03-12 subsaharan_africa sierra_leone western_area western_area Lewandowski et al https://www.ncbi.nlm.nih.gov/nuccore/KR025228 Direct Submission | ||
| EM_000968 KR817085 2014-10-01 subsaharan_africa guinea macenta macenta Carroll et al https://www.ncbi.nlm.nih.gov/nuccore/KR817085 Temporal and spatial analysis of the 2014-2015 Ebola virus outbreak in West Africa | ||
| 490_EMLH KU296746 2015-01-30 subsaharan_africa sierra_leone sierra_leone sierra_leone Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296746 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| PL4696 KU296513 2015-03-13 subsaharan_africa sierra_leone port_loko port_loko Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296513 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| MK2788 KU296573 2015-03-07 subsaharan_africa sierra_leone bombali bombali Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296573 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| LIBR10110 KT725387 2014-08-20 subsaharan_africa liberia liberia liberia Ladner et al https://www.ncbi.nlm.nih.gov/nuccore/KT725387 "Evolution and spread of Ebola virus in Liberia, 2014-2015" | ||
| J0102 KP759657 2014-10-27 subsaharan_africa sierra_leone port_loko port_loko Tong et al https://www.ncbi.nlm.nih.gov/nuccore/KP759657 Genetic diversity and evolutionary dynamics of Ebola virus in Sierra Leone | ||
| EM_COY_2015_017574 EM_COY_2015_017574 2015-06-10 subsaharan_africa guinea boke boke Quick et al https://github.com/nickloman/ebov/ ? | ||
| 18636_EMLH KU296460 2015-07-08 subsaharan_africa sierra_leone sierra_leone sierra_leone Arias et al https://www.ncbi.nlm.nih.gov/nuccore/KU296460 Rapid outbreak sequencing of Ebola virus in Sierra Leone identifies transmission chains linked to sporadic cases | ||
| V768 KR534588 2014-08-27 subsaharan_africa guinea conakry conakry Simon-Loriere et al https://www.ncbi.nlm.nih.gov/nuccore/KR534588 Distinct lineages of Ebola virus in Guinea during the 2014 West African epidemic |
Large diffs are not rendered by default.
Oops, something went wrong.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,89 @@ | ||
| # nextstrain.org/ebola/ingest | ||
|
|
||
| This is the ingest pipeline for Ebola virus sequences. | ||
|
|
||
| ## Usage | ||
|
|
||
| > NOTE: All command examples assume you are within the `ingest` directory. | ||
| > If running commands from the outer `ebola` directory, please replace the `.` with `ingest` | ||
|
|
||
| Fetch sequences with | ||
|
|
||
| ```sh | ||
| nextstrain build --cpus 1 . data/sequences.ndjson | ||
| ``` | ||
|
|
||
| Run the complete ingest pipeline with | ||
|
|
||
| ```sh | ||
| nextstrain build --cpus 1 . | ||
| ``` | ||
|
|
||
| This will produce two files (within the `ingest` directory): | ||
|
|
||
| - data/metadata.tsv | ||
| - data/sequences.fasta | ||
|
|
||
| Run the complete ingest pipeline and upload results to AWS S3 with | ||
|
|
||
| ```sh | ||
| nextstrain build . --configfiles config/config.yaml config/optional.yaml | ||
| ``` | ||
|
|
||
| ### Adding new sequences not from GenBank | ||
|
|
||
| #### Static Files | ||
|
|
||
| Do the following to include sequences from static FASTA files. | ||
|
|
||
| 1. Convert the FASTA files to NDJSON files with: | ||
|
|
||
| ```sh | ||
| wget https://raw.githubusercontent.com/nextstrain/monkeypox/master/ingest/bin/fasta-to-ndjson -O bin/fasta-to-ndjson | ||
| chmod 755 bin/fasta-to-ndjson | ||
| ./bin/fasta-to-ndjson \ | ||
| --fasta {path-to-fasta-file} \ | ||
| --fields {fasta-header-field-names} \ | ||
| --separator {field-separator-in-header} \ | ||
| --exclude {fields-to-exclude-in-output} \ | ||
| > ingest/data/{file-name}.ndjson | ||
| ``` | ||
|
|
||
| 2. Add the following to the `.gitignore` to allow the file to be included in the repo: | ||
|
|
||
| ```gitignore | ||
| !ingest/data/{file-name}.ndjson | ||
| ``` | ||
|
|
||
| 3. Add the `file-name` (without the `.ndjson` extension) as a source to `ingest/config/config.yaml`. This will tell the ingest pipeline to concatenate the records to the GenBank sequences and run them through the same transform pipeline. | ||
|
|
||
| ## Configuration | ||
|
|
||
| Configuration takes place in `config/config.yaml` by default. | ||
| Optional configs for uploading files and Slack notifications are in `config/optional.yaml`. | ||
|
|
||
| ### Environment Variables | ||
|
|
||
| The complete ingest pipeline with AWS S3 uploads and Slack notifications uses the following environment variables: | ||
|
|
||
| #### Required | ||
|
|
||
| - `AWS_DEFAULT_REGION` | ||
| - `AWS_ACCESS_KEY_ID` | ||
| - `AWS_SECRET_ACCESS_KEY` | ||
| - `SLACK_TOKEN` | ||
| - `SLACK_CHANNELS` | ||
|
|
||
| #### Optional | ||
|
|
||
| These are optional environment variables used in our automated pipeline for providing detailed Slack notifications. | ||
|
|
||
| - `GITHUB_RUN_ID` - provided via [`github.run_id` in a GitHub Action workflow](https://docs.github.com/en/actions/learn-github-actions/contexts#github-context) | ||
| - `AWS_BATCH_JOB_ID` - provided via [AWS Batch Job environment variables](https://docs.aws.amazon.com/batch/latest/userguide/job_env_vars.html) | ||
|
|
||
| ## Input data | ||
|
|
||
| ### GenBank data | ||
|
|
||
| GenBank sequences and metadata are fetched via NCBI Virus. | ||
| The exact URL used to fetch data is constructed in `bin/genbank-url`. | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I tested this approach using sequences we have in fauna.
These steps all ran as expected. When I tried to dry-run the ingest after modifying the config file to include this new source, I got the following error:
I tried renaming the NDJSON file I created above to
data/local_ebola_all.ndjsonand ran the workflow again. It ran as expected. In cases where the strain names were different between fauna and GenBank for the same accession (e.g.,Ebolavirus/H_sapiens_wt/SLE/2014/Makona_J0005andJ0005for accessionKP759628), both records appeared in the output metadata and sequences. While we can disambiguate the metadata records by strain name, the sequences use only the accession, so we end up with duplicate records in sequences.These duplicate sequences cause augur filter to exit with an error (with a list of duplicate strains) in the top-level workflow. This outcome is probably close to what we would expect, since we want users to know duplicates exist and handle these accordingly. It might be nicer for users for the ingest to warn them about duplicates, but I don't think we have any tools to check for dups currently, do we?
Regarding the issue with the NDJSON filename mismatch above, we could probably address this problem by removing the serotype wildcard from the sources, since we plan to collect all Ebola serotypes into a single pair of metadata and sequence files and split by serotype in the Nextstrain workflow anyway.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Done with commit: fc12467
Would be wonderful if we did...I tend to use smof uniq , uniq_merge.py, or write up a one-off perl scripts to remove duplicate sequences. There are many possible ways to remove sequence duplicates. Was there a de-dups rule you had in mind? Maybe from a nextstrain repo?
I've seen some gnarly strain names in my time (and don't get me started on resequenced Isolates, or sequenced segments from the same strain)... so highly recommend keeping with the Monkeypox method of IDs by Accession. I see you're checking old fauna datasets.... My method of checking fauna was writing a quick script to replace
with its associated:
I can spec out a quick perl/python script to do the swap where it (1) reads in a metadata file to generate strain_name-to-accession_number hash and (2) read in sequence file to replace
strain_namewithAccession_number.Ah this is exactly where I use uniq_merge.py to flag accessions with multiple strain names and create a fix-stain-name.pl script. I'll think about it more. Thank you for the feedback.