Skip to content

Repository files navigation

Genotyping

DOI

This workflow creates a VCF file based on the SNP microarray data in a set of .CEL files. This VCF file can be used for demultiplexing 10X scRNA-seq data where multiple samples have been "pooled" and captured together.

Authors

This workflow was originally developed by @nbartonicek as a collection of bash and SGE scripts. These scripts were then refactored as part of the Swarbrick Lab souporcell workflow (VPN required) by @dlroden, as an extension of the original souporcell workflow by @wheaton5 et al. Parts of the Swarbrick Lab souporcell workflow were then transferred to the Swarbrick Lab demuxafy workflow by @dlroden and @BeataKiedik, as an extension of the original demuxafy workflow by @drneavin. Finally, the genotyping steps in the Swarbrick Lab demuxafy workflow were extracted and refactored as a standalone workflow (this repo) by @johnyaku.

Overview

SNP microarray data is stored in .CEL files, with one .CEL file per donor sample, as defined in a sample sheet (donors.csv). This workflow converts the genotyping information in these .CEL files into a VCF file containing high-confidence variant calls for all donor samples. Actually, two versions of this VCF file are produced: one for the hg19 assembly (build 37) and another for the hg38 assembly (build 38).

The workflow also creates a .psam file, ready for use with plink. The donor IDs in this .psam file are reformatted (if necessary) by replacing hyphens with underscores, so id_map.tsv is provided for switching back and forth between the original and reformatted donor IDs.

This workflow uses the Analysis Power Tools (APT) by ThermoFisher Scientific. Annotation of the VCF file is based on the NetAffx annotation for the Axiom UK Biobank array (also by ThermoFisher Scientific), which in turn is based on the hg19 assembly of the human genome. This worklow formats the VCF file by adding header contigs for hg19 and reformatting the chomosome names from 1, 2, 3, ... to chr1, chr2, chr3, ... Finally, the formatted hg19 VCF file is lifted over to the hg38 assembly.

The workflow, defined by the Snakefile runs as shown by the following rule graph:

Rulegraph

See the individual rule definitions to understand the function of each rule.

Further reading: Axiom Genotyping Solution Data Analysis User Guide

Installation and use

Third-party software

The genotyping rules call the Analysis Power Tools (APT) and SNPolisher, which are proprietary software distributed by Life Technologies Corporation (ThermoFisher Scientific) under an End User License Agreement. The EULA grants a non-transferable licence, with no right to sublicense, to use the software on computers you own or control, for research use only, and prohibits distributing or electronically transmitting it -- alone or combined with other products. APT is not open source: the 2.x downloads ship no source code, and the GPL terms that search results often surface apply to the retired 1.x series.

APT is therefore not covered by the licence for this repository, and cannot be redistributed with it. To run the apt, ps_metrics, ps_classification, otv_caller and make_vcf rules you must obtain APT from ThermoFisher yourself, under your own acceptance of their terms, and build a container image using the Dockerfile in containers/apt/. Point the workflow at the result with containers.apt in your config file, as either a reference to a registry you control or the path to a local Singularity image:

containers:
  apt: "docker://your-private-registry/apt:2.12.0"

The prep.sh script does this for you — it builds the image, converts it for Singularity if needed, puts it where containers.apt points, and checks that APT runs inside it:

./modules/genotyping/prep.sh --configfile config/genotyping/config.yaml

Because APT cannot be redistributed, any registry you use must be one you control (a private or institutional registry) -- do not push the image to a public registry such as Docker Hub, GHCR, or Quay. See preparing the APT container for the options, including clusters such as NCI Gadi that provide Singularity but not Docker, and containers/apt/README.md for the licensing background.

The array library and annotation files referenced under refs.apt in the config file (Axiom_UKB_WCSG.*) are vendor-supplied and likewise cannot be redistributed here. ThermoFisher reserves all rights in them, and they are not covered by the APT EULA, which applies only to the software. prep.sh downloads them for you and checks them against recorded checksums:

./modules/genotyping/prep.sh --configfile config/genotyping/config.yaml --what resources

Note that the config file names three of these files, but APT reads ten: the arg file is an XML document that references seven further library files by bare filename, which APT resolves against the directory the arg file lives in. They therefore all have to be downloaded into that one directory. See fetching the Axiom array files.

The remaining rules use openly licensed containers: bcftools (MIT/Expat) and samtools (MIT/Expat) via BioContainers, and the bcftools +liftover plugin (MIT).

Tools, references and citations

This section records the provenance a citing manuscript needs. Full citations, with DOIs/PMIDs, and how to cite tools that have no dedicated paper (APT, SNPolisher) are in docs/prior-art.md.

Platform

Item Value Source
Array Applied Biosystems™ UK Biobank Axiom™ Array (Axiom_UKB_WCSG) CEL headers (affymetrix-array-type)
Array design citation Bycroft C, et al. Nature 2018;562:203–209. doi:10.1038/s41586-018-0579-z
Library package Axiom_UKB_WCSG r5 refs.apt filenames
Annotation build na35 (GRCh37/hg19) Axiom_UKB_WCSG.na35.annot.db
Genotyping facility Ramaciotti Centre for Genomics, UNSW Sydney service records

Tools

Tool Version Container Role
Analysis Power Tools (APT) apt-genotype-axiom 2.10.0 (calling step; the manuscript version) built from containers/apt/ genotype calling, metrics, classification, OTV, VCF export
SNPolisher bundled in the APT image as above probeset classification
bcftools 1.21 quay.io/biocontainers/bcftools:1.21--h8b25389_0 VCF formatting
bcftools +liftover bcftools 1.18 + liftover plugin yangyxt/bcftools_liftover:1.18 hg19 → hg38 liftover
samtools 1.21 quay.io/biocontainers/samtools:1.21--h50ea8bc_0 FASTA indexing
Snakemake 7.32.4 conda (env/snakemake_7.32.4.yaml) workflow engine

See containers/apt/README.md for how to build the container, including pinning the manuscript's APT 2.10.0.

Reference data

Reference Build How it is obtained
hg19 FASTA UCSC hg19 dvc import-url (UCSC goldenPath)
hg38 FASTA GRCh38, Ensembl release-98 (main chromosomes) 25 dvc import-url stages + the prepare_hg38 rule
liftover chain UCSC hg19ToHg38 dvc import-url (UCSC gbdb)
Axiom library Axiom_UKB_WCSG r5 Thermo Fisher (prep.sh --what resources)
Axiom annotation na35 Thermo Fisher

Citing this workflow

A CITATION.cff at the repository root provides citation metadata, which GitHub surfaces via Cite this repository.

Known caveats of the workflow's output are documented in docs/limitations.md.

Test data

The test dataset is four human induced pluripotent stem cell lines from GEO series GSE224950, run on the same Axiom UK Biobank array as our own data and deposited by the Powell lab at the Garvan Institute for Neavin et al. 2023. Two are female and two male, so the sex calls in the .psam are exercised.

These .CEL files are not redistributed here. They are fetched from GEO and checked against recorded checksums by prep.sh:

./prep.sh --what testdata

See fetching the test data.

Earlier releases tested against SNP microarray measurements from our own donors, which are potentially identifiable and not publicly redistributable. Those files are no longer referenced here; researchers wanting them should refer to the data availability statement of the associated publication.

The bundled test dataset (config/test.yaml) is reproducible outside the lab: prep.sh --what all --configfile config/test.yaml fetches the APT container, the public Axiom array files, and the public test CEL files (GEO GSE224950), and the reference genomes are public DVC import-url stages pulled with dvc update. Every input is publicly obtainable -- the repository has no private data imports.

The workflow itself does not depend on any private data -- it can be run on any set of Axiom .CEL files by pointing the config file at them.

Releases

Projects using this workflow pin a specific revision of it as a git submodule, so that published results can always be traced back to the exact code that produced them. Revisions used for published analyses are tagged with a paper/ prefix -- see the tags.

Standalone releases are tagged with a semantic version (e.g. v1.0.0) and archived on Zenodo with a citable DOI -- see Citing this workflow.

License

This repository is released under the MIT License, with the exception of the third-party software noted above.

About

CEL to VCF

Topics

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages