YTAB = Yeast Transposon Analysis Browser
A reproducible platform for exploring yeast transposon-seq data, including:
- library and sample QC
- insertion mapping summaries
- gene-level abundance and fitness summaries
- essentiality and screen analysis
- reproducible workflow outputs for downstream visualization
app/— browser / UIpipeline/— Nextflow workflowsrc/— core analysis coderesources/— reference files and annotationsconfigs/— parameter filescontainers/— Dockerfilesdocs/— architecture and migration notesexamples/— small demo inputstests/— unit and integration tests
cd "<repository-root>"
./scripts/local/ytab_setup_env.sh
mamba activate ytab-local
./scripts/local/ytab_launch_app.sh \
--project-config output/projects/H2O2_screen_v1/config/project.yamlytab_setup_env.sh creates or updates the ytab-local environment from
environment.local.yml. It uses mamba when available, installs mamba into
an existing conda base environment when needed, or falls back to a repo-local
micromamba bootstrap if neither conda nor mamba is installed. Use
./scripts/local/ytab_setup_env.sh --launch-app to set up the environment and
start the app in one command.
Before starting a development session, collaborators should sync from the current shared branch with:
./scripts/local/ytab_sync_start.shThis script refuses to run if local staged, unstaged, or untracked files exist,
then fetches and pulls with rebase from origin on the current branch. It does
not stage, commit, push, clean, reset, or delete files.
The app binds to 127.0.0.1:3838 by default. Another host or port must be requested explicitly.
Use python scripts/local/ytab_project_status.py --project-config <project.yaml> --show-next
to inspect resumable stage state without running analysis.
YTAB has two separate scientific branches: parent-only MidLC normalization feeds the essentiality classifier, while treated-versus-parent fitness uses raw per-sample SummaryTable output and performs CPM normalization inside R. The orchestrator preserves this separation and reuses stage caches.
The Fitness Screen exposes generated comparison designs, temporary interactive subsets, explicit Preview and Run modes, and design-aware cache reuse. Classifier labels are optional and do not change fitness calculations. Matching, historical, and legacy results remain distinct; existing calls can be searched, filtered, and viewed as a call distribution or effect-versus-z plot. No P values or volcano plot are introduced.
YTAB now opens on a local landing page. For low-memory computers or large FASTQ inputs, preprocess first with two threads and then open the app:
mamba activate ytab-local
python scripts/local/ytab_run_pipeline.py \
--project-config output/projects/<PROJECT_ID>/config/project.yaml \
--profile core --threads 2 --keep-going
./scripts/local/ytab_launch_app.sh \
--project-config output/projects/<PROJECT_ID>/config/project.yamlAlternatively, launch ./scripts/local/ytab_launch_app.sh, choose Create new project, select a local FASTQ directory, initialize it, and continue preprocessing. FASTQs stay in place and are not uploaded. Mapping runs one sample at a time, can take substantial time, and can be resumed later; completed cached stages are skipped. Raw SummaryTable completion makes a project analysis-ready. MidLC belongs only to the essentiality-classifier branch, while treated-versus-parent fitness uses raw SummaryTables and performs CPM normalization inside its R analysis.
The workspace keeps an active-job banner visible across tabs. Persistent progress files record the current sample, confirmed completions, elapsed time, recent state, and an approximate ETA after the first non-cached item finishes. Estimates vary with storage and system load. Cancelling preserves completed outputs so later cache-aware runs can resume without repeating successful samples.
Quality Control presents compact Mapping QC and raw Summary QC tables; secondary summary metrics and project-relative source paths remain available under Detailed metrics. Library Diagnostics has an independent searchable sample selection with all-eligible, parent, treated, and custom presets. Preview command validates inputs without scientific execution; Run diagnostics performs or reuses an exact sample-set-aware cached run, while Force rerun only bypasses a matching cache. Concise results remain separate from collapsed technical logs and paths.