diff --git a/.prettierignore b/.prettierignore
index f0a6037..c908443 100644
--- a/.prettierignore
+++ b/.prettierignore
@@ -6,3 +6,4 @@ results/
.DS_Store
*.code-workspace
assets/*.html
+docs/*.md
diff --git a/CHANGELOG.md b/CHANGELOG.md
index afc95d5..401b6f4 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,5 +1,6 @@
## ASPEN development version
+- Document successful `dryrun` and `run` console output in `docs/deployment.md`, including example screenshots, an explanation of `STEP`/`OK`/`INFO`/`NEXT`, and explicit links between wrapper messages and `WORKDIR` status files. (#147, @kopardev)
- Add attempt-aware resource scaling for the main Snakemake rules, document how baseline `cluster.json` resources interact with retry-based `resources.gres` scaling, and fix dry-run failure handling so failed dry-runs now exit non-zero instead of reporting success. (#152, @kopardev)
- Detect SLURM-killed Snakemake child jobs in classic cluster mode by adding a `--cluster-status` hook and `scancel` integration, so timed out, OOM-killed, cancelled, or otherwise failed child jobs now retry/fail cleanly instead of leaving the master workflow hanging indefinitely. (#148, @kopardev)
- Clarify motif enrichment outputs and interpretation in the documentation: fix the reported location of HOMER/AME result folders, explain how ASPEN reuses HOMER-generated `target.fa`/`background.fa` in AME, document the AME parallelization strategy, and add simpler guidance for interpreting `knownResults.txt` and `ame_results.txt`. (#146, @kopardev)
diff --git a/docs/assets/images/aspen_console_output_examples.html b/docs/assets/images/aspen_console_output_examples.html
new file mode 100644
index 0000000..0966854
--- /dev/null
+++ b/docs/assets/images/aspen_console_output_examples.html
@@ -0,0 +1,210 @@
+
+
+
+
+
+ ASPEN Console Output Examples
+
+
+
+
+
+
ASPEN Docs Assets
+
ASPEN console output examples
+
Terminal-style screenshots for the deployment docs, showing what successful dryrun and run wrapper output look like in practice.
+
+
+
+
+
Successful dry-run summary
+
Use in deployment docs near the aspen -m=dryrun example.
+
+
+
+
+
bash • ASPEN dryrun example
+
+
$ aspen --workdir=<WORKDIR> --runmode=dryrun
+STEP [dryrun] Running preflight checks
+INFO Validating config, samples, and runtime dependencies
+INFO Rendering the planned Snakemake DAG without execution
+OK Dry-run completed.
+INFO Full dry-run output captured in <WORKDIR>/dryrun.log
+NEXT Submit run with: aspen --workdir=<WORKDIR> --runmode=run
+
+
+
+
+
+
Successful run submission summary
+
Use in deployment docs near the aspen -m=run and monitoring guidance.
+
+
+
+
+
bash • ASPEN run example
+
+
$ aspen --workdir=<WORKDIR> --runmode=run
+OK Dry-run was successful.
+STEP [run] Preparing workdir for new execution
+OK Created <WORKDIR>/submit_script.sbatch
+INFO State marker updated: pipeline.running (reason=submission_started, slurm_job_id=NA)
+INFO State marker updated: pipeline.running (reason=sbatch_submitted, slurm_job_id=<jobid>)
+---------------------------------------------------------------------
+OK Job submitted successfully (SLURM job ID: <jobid>)
+NEXT Monitor: squeue -u $USER
+NEXT Progress: tail -f <WORKDIR>/snakemake.log
+NEXT Status: ls -1 <WORKDIR>/pipeline.*
+NEXT Sidecar: <WORKDIR>/pipeline.status.json
+
+
+
+
+
diff --git a/docs/assets/images/aspen_dryrun_console.jpeg b/docs/assets/images/aspen_dryrun_console.jpeg
new file mode 100644
index 0000000..b06792f
Binary files /dev/null and b/docs/assets/images/aspen_dryrun_console.jpeg differ
diff --git a/docs/assets/images/aspen_dryrun_console.svg b/docs/assets/images/aspen_dryrun_console.svg
new file mode 100644
index 0000000..cb3542d
--- /dev/null
+++ b/docs/assets/images/aspen_dryrun_console.svg
@@ -0,0 +1,55 @@
+
diff --git a/docs/assets/images/aspen_run_console.jpeg b/docs/assets/images/aspen_run_console.jpeg
new file mode 100644
index 0000000..d33242a
Binary files /dev/null and b/docs/assets/images/aspen_run_console.jpeg differ
diff --git a/docs/assets/images/aspen_run_console.svg b/docs/assets/images/aspen_run_console.svg
new file mode 100644
index 0000000..f91a717
--- /dev/null
+++ b/docs/assets/images/aspen_run_console.svg
@@ -0,0 +1,62 @@
+
diff --git a/docs/deployment.md b/docs/deployment.md
index 3ed5f6a..6756db5 100644
--- a/docs/deployment.md
+++ b/docs/deployment.md
@@ -31,20 +31,21 @@ ASPEN requires a sample manifest file (`samples.tsv`) to identify and organize y
- `path_to_R2_fastq`: Absolute path to the Read 2 FASTQ file (required for paired-end data).
!!! note
-Symlinks for R1 and R2 files will be created in the results directory, named as `.R1.fastq.gz` and `.R2.fastq.gz`, respectively. Therefore, original filenames do not need to be altered.
+ Symlinks for R1 and R2 files will be created in the results directory, named as `.R1.fastq.gz` and `.R2.fastq.gz`, respectively. Therefore, original filenames do not need to be altered.
!!! note
-The `replicateName` is used as a prefix for individual peak calls, while the `sampleName` serves as a prefix for consensus peak calls.
+ The `replicateName` is used as a prefix for individual peak calls, while the `sampleName` serves as a prefix for consensus peak calls.
!!! warning "Biological vs. technical replicates"
-ASPEN expects **one row per biological replicate**. If you sequenced the same sample across multiple lanes or sequencing runs (technical replicates), you must **concatenate those FASTQ files into a single file** before creating your manifest — ASPEN does not merge lanes internally.
+ ASPEN expects **one row per biological replicate**. If you sequenced the same sample across multiple lanes or sequencing runs (technical replicates), you must **concatenate those FASTQ files into a single file** before creating your manifest — ASPEN does not merge lanes internally.
- | Replicate type | Definition | What to do |
- |---|---|---|
- | **Biological** | Independent biological samples (separate cultures, animals, patients, etc.) | One row per sample in `samples.tsv` |
- | **Technical** | Same sample re-sequenced across multiple lanes or runs | `cat` the FASTQs together first, then one row |
+ Biological replicates are independent samples:
+
+ - **Biological**: independent biological samples (separate cultures, animals, patients, etc.). Use one row per sample in `samples.tsv`.
+ - **Technical**: the same sample re-sequenced across multiple lanes or runs. `cat` the FASTQs together first, then use one row.
Example of concatenating technical replicates before running ASPEN:
+
```bash
cat sample1_L001_R1.fastq.gz sample1_L002_R1.fastq.gz > sample1_R1.fastq.gz
cat sample1_L001_R2.fastq.gz sample1_L002_R2.fastq.gz > sample1_R2.fastq.gz
@@ -53,7 +54,7 @@ ASPEN expects **one row per biological replicate**. If you sequenced the same sa
DESeq2 (used in `diffatac`) requires **at least 2 biological replicates per group**. Technical replicates do not count as biological replicates and will not satisfy this requirement.
!!! note
-For differential ATAC analysis, create a `contrasts.tsv` file with two columns (Group1 and Group2 ... aka Sample1 and Sample2, without headers) and place it in the output directory after initialization. Ensure each group/sample in the contrast has at least two biological replicates, as DESeq2 requires this for accurate contrast calculations.
+ For differential ATAC analysis, create a `contrasts.tsv` file with two columns (Group1 and Group2 ... aka Sample1 and Sample2, without headers) and place it in the output directory after initialization. Ensure each group/sample in the contrast has at least two biological replicates, as DESeq2 requires this for accurate contrast calculations.
## 🏃 Running the ASPEN Pipeline
@@ -70,11 +71,12 @@ aspen -m=init -w=
This command generates a config.yaml and a placeholder `samples.tsv` in the specified directory. Edit these files to reflect your experimental setup, replacing the placeholder `samples.tsv` with your prepared manifest. If performing differential analysis, include the `contrasts.tsv` file at this stage.
!!! note
-To explore all possible options of the `aspen` command you can either run it without any arguments or run `aspen --help`
+ To explore all possible options of the `aspen` command you can either run it without any arguments or run `aspen --help`
Here is what help looks like:
-> **Note**: This is illustrative example output captured at doc-writing time — exact values (e.g. `pipeline_home`, `git commit/tag`, `aspen_version`) will differ depending on which ASPEN version/branch is installed at your site. Run `aspen --help` yourself to see the current values for your installation.
+!!! note
+ This is illustrative example output captured at doc-writing time — exact values (e.g. `pipeline_home`, `git commit/tag`, `aspen_version`) will differ depending on which ASPEN version/branch is installed at your site. Run `aspen --help` yourself to see the current values for your installation.
```bash
@@ -250,6 +252,21 @@ Before executing the full analysis and after editing the `config.yaml` as needed
aspen -m=dryrun -w=
```
+#### What successful dry-run output looks like
+
+
+
+The dry-run summary confirms that ASPEN finished the preflight checks successfully, points you to `dryrun.log` for the full transcript, and ends with the exact `run` command to use next.
+
+In practice, the wrapper prefixes mean:
+
+- `STEP` starts a major wrapper phase.
+- `OK` confirms that phase completed successfully.
+- `INFO` reports status details such as where output was written.
+- `NEXT` gives the follow-up command or monitoring action ASPEN expects you to take.
+
+For `dryrun`, the short terminal summary is only the high-level result. The full dry-run transcript is written to `dryrun.log`; see the [outputs page](outputs.md) for the complete workdir artifact reference.
+
This step outlines the sequence of tasks (Directed Acyclic Graph - DAG) without actual execution, allowing you to verify the planned operations.
### 🚀 Execute the Pipeline
@@ -260,13 +277,28 @@ If the dry run output is satisfactory, proceed to execute the pipeline:
aspen -m=run -w=
```
+#### What successful run submission output looks like
+
+
+
+The run summary shows that ASPEN created the submission script, wrote the `pipeline.running` marker updates during submission, and printed the `NEXT` monitoring hints for `squeue`, `snakemake.log`, and `pipeline.status.json`.
+
+Those lines map directly to files in `WORKDIR`:
+
+- `submit_script.sbatch` is created before the master Slurm job is submitted.
+- `pipeline.running` is the human-readable status marker updated during execution.
+- `pipeline.status.json` is the machine-readable sidecar for scripting and automation.
+- `snakemake.log` is the detailed workflow execution log once the run starts.
+
+Treat the `NEXT` lines as the wrapper's built-in "what should I do now?" guidance. They point you to the same files and commands documented below and in the [outputs page](outputs.md), without requiring you to remember the monitoring commands yourself.
+
This command submits a master job to the Slurm workload manager, which orchestrates the entire analysis workflow, managing job submissions and monitoring progress.
ASPEN submits the generated Snakemake job with rule-level resources exposed to Slurm. The heavy rules that may need extra scratch space now define an attempt-aware `resources.gres` value, so the first submission uses the baseline `cluster.json` request and later retries can scale the requested `lscratch` allocation automatically if a rule is retried by Snakemake. This keeps the default configuration in `cluster.json` while still allowing individual rules to become more conservative on repeated failures.
- 🛠️ **Optional Argument**:
-`--singcache` or `-c`: Specify a Singularity cache directory. The default is `/data/${USER}/.singularity` if available; otherwise, it defaults to `${WORKDIR}/snakemake/.singularity`.
+`--singcache` or `-c`: Override the Singularity cache directory. On Biowulf, when you load the `ccbrpipeliner` module, ASPEN already uses `SIFCACHE=/data/CCBR_Pipeliner/SIFS` for you, so you usually do not need `-c`. Use `-c` only on another HPC system if you want to point ASPEN at your own cache directory or pull containers yourself.
**💡 Example Command**:
@@ -274,13 +306,31 @@ ASPEN submits the generated Snakemake job with rule-level resources exposed to S
aspen -m=run -w= -c /data/${USER}/.singularity
```
-This command runs the pipeline with the specified working directory and Singularity cache directory.
-
-> **Note**: If deploying on Biowulf, try setting the `--singcache` to `/data/CCBR_Pipeliner/SIFS` to reuse the pre-pulled containers and save time.
+This example is for a non-Biowulf HPC system where you want to manage your own Singularity cache location.
## 📊 Monitor ASPEN Runs
-To monitor the status of your ASPEN pipeline and its associated jobs on a Slurm-managed system, you can utilize the squeue and scontrol commands. The squeue command provides information about jobs in the scheduling queue, while scontrol offers detailed insights into specific jobs.
+For day-to-day status checks, use the sidecar and `pipeline.*` files below first. They are the fastest and most reliable way to see whether ASPEN is running, completed, or failed. Reach for `squeue` and `scontrol` only when you want an advanced scheduler-level view or need to inspect an individual SLURM job.
+
+If you do need to inspect the cluster directly, `squeue` shows the queue state and `scontrol` exposes detailed job metadata.
+
+### 📝 Pipeline State Markers: the primary status check
+
+ASPEN writes a set of state-tracking files directly into `WORKDIR` while a `run` is executing, so you can check status from the sidecar and `pipeline.*` files first, even without Slurm access (for example, from a laptop over `ssh`):
+
+- `pipeline.running`, `pipeline.completed`, `pipeline.failed`, `pipeline.canceled` — exactly one of these marker files exists at a time, reflecting the current state. While the pipeline is running, `pipeline.running` is periodically refreshed by a background progress monitor with a human-readable summary, including the percentage of Snakemake steps completed so far:
+
+ ```bash
+ cat /pipeline.running
+ ```
+
+- `pipeline.status.json` — a machine-readable sidecar with the same information (`state`, `reason`, `slurm_job_id`, start/end timestamps, `duration_seconds`, `tasks_done`/`tasks_total`, `exit_code`), useful for scripting/automation:
+
+ ```bash
+ cat /pipeline.status.json
+ ```
+
+- `snakemake.log.jobby` / `snakemake.log.jobby.short` — a `jobby` TSV summary of per-rule/job resource usage, generated as a best-effort step after the run finishes (even if the Slurm submission itself failed before Snakemake started).
To view all your active and pending jobs, execute:
@@ -303,21 +353,3 @@ To quickly gauge the process of the entire pipeline run:
```bash
grep "done$" /snakemake.log
```
-
-### 📝 Lightweight Status Checks via Pipeline State Markers
-
-In addition to `squeue`/`scontrol`, ASPEN writes a set of state-tracking files directly into `WORKDIR` while a `run` is executing, so you don't need Slurm access (e.g. from a laptop over `ssh`) to check on a run:
-
-- `pipeline.running`, `pipeline.completed`, `pipeline.failed`, `pipeline.canceled` — exactly one of these marker files exists at a time, reflecting the current state. While the pipeline is running, `pipeline.running` is periodically refreshed by a background progress monitor with a human-readable summary, including the percentage of Snakemake steps completed so far:
-
- ```bash
- cat /pipeline.running
- ```
-
-- `pipeline.status.json` — a machine-readable sidecar with the same information (`state`, `reason`, `slurm_job_id`, start/end timestamps, `duration_seconds`, `tasks_done`/`tasks_total`, `exit_code`), useful for scripting/automation:
-
- ```bash
- cat /pipeline.status.json
- ```
-
-- `snakemake.log.jobby` / `snakemake.log.jobby.short` — a `jobby` TSV summary of per-rule/job resource usage, generated as a best-effort step after the run finishes (even if the Slurm submission itself failed before Snakemake started).
diff --git a/docs/index.md b/docs/index.md
index 67714c2..978d5e0 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -1,8 +1,8 @@
## Background
!!! tip "New here?"
-If you just want to **run the pipeline**, skip ahead to
-[Running ASPEN](deployment.md). Come back here later for the
-scientific background on ATAC-seq and why ASPEN is built the way it is.
+ If you just want to **run the pipeline**, skip ahead to
+ [Running ASPEN](deployment.md). Come back here later for the
+ scientific background on ATAC-seq and why ASPEN is built the way it is.
The Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) has revolutionized genomics by providing a rapid and sensitive method to assess chromatin accessibility across the genome. This technique offers profound insights into gene regulation, epigenetic modifications, and the dynamic landscape of the chromatin environment. To facilitate the comprehensive analysis of ATAC-seq data, the Center for Cancer Research (CCR) Collaborative Bioinformatics Resource (CCBR) has developed ASPEN (**A**tac **S**eq **P**ip**E**li**N**e), an automated, robust, reproducible pipeline designed using the Snakemake pipelining framework to streamline the complexities inherent in ATAC-seq data processing.
diff --git a/docs/limitations.md b/docs/limitations.md
index 7eed620..59c86d3 100644
--- a/docs/limitations.md
+++ b/docs/limitations.md
@@ -18,32 +18,32 @@
| hs1 | Human | _Homo sapiens_ (T2T-CHM13) |
| hs1_chrR | Human | _Homo sapiens_ (T2T-CHM13 + chrR rDNA unit; see note below) |
- !!! tip "Use `hs1_chrR` for ribosomal DNA chromatin accessibility studies"
- Standard genome assemblies mask or collapse ribosomal DNA (rDNA) repeat
- loci, making it impossible to call ATAC-seq peaks in these regions.
- `hs1_chrR` is a customized version of the T2T-CHM13 (`hs1`) assembly
- where endogenous rDNA-like sequences are masked throughout the canonical
- chromosomes, and a single consensus rDNA unit (NCBI KY962518.1, modified)
- is inserted as an additional chromosome **chrR**. This design ensures that
- rDNA-mapping reads are captured unambiguously on chrR rather than being
- lost to multi-mapping or suppressed by masking.
-
- **When to choose `hs1_chrR` over `hs1`:**
-
- - Your experiment involves a perturbation that may affect nucleolar
- chromatin or ribosomal gene accessibility (e.g., RNA Pol I inhibition,
- nucleolar stress, epigenetic reprogramming).
- - You want to quantify ATAC-seq peaks at rDNA promoters, the transcribed
- region, or the intergenic spacer (IGS) of the ribosomal repeat unit.
- - You need FRiP scores, TSS enrichment, or differential accessibility
- analysis specifically for rDNA loci alongside the rest of the genome.
-
- **Reference:**
- George SS, Pimkin M, Paralkar VR.
- *"Construction and validation of customized genomes for human and mouse ribosomal DNA mapping."*
- J Biol Chem (2023).
-
- Genome files:
+!!! tip "Use `hs1_chrR` for ribosomal DNA chromatin accessibility studies"
+ Standard genome assemblies mask or collapse ribosomal DNA (rDNA) repeat
+ loci, making it impossible to call ATAC-seq peaks in these regions.
+ `hs1_chrR` is a customized version of the T2T-CHM13 (`hs1`) assembly
+ where endogenous rDNA-like sequences are masked throughout the canonical
+ chromosomes, and a single consensus rDNA unit (NCBI KY962518.1, modified)
+ is inserted as an additional chromosome **chrR**. This design ensures that
+ rDNA-mapping reads are captured unambiguously on chrR rather than being
+ lost to multi-mapping or suppressed by masking.
+
+ **When to choose `hs1_chrR` over `hs1`:**
+
+ - Your experiment involves a perturbation that may affect nucleolar
+ chromatin or ribosomal gene accessibility (e.g., RNA Pol I inhibition,
+ nucleolar stress, epigenetic reprogramming).
+ - You want to quantify ATAC-seq peaks at rDNA promoters, the transcribed
+ region, or the intergenic spacer (IGS) of the ribosomal repeat unit.
+ - You need FRiP scores, TSS enrichment, or differential accessibility
+ analysis specifically for rDNA loci alongside the rest of the genome.
+
+ **Reference:**
+ George SS, Pimkin M, Paralkar VR.
+ *"Construction and validation of customized genomes for human and mouse ribosomal DNA mapping."*
+ J Biol Chem (2023).
+
+ Genome files:
- **Spike-in genomes supported**: Spike-in genomes supported is limited to:
diff --git a/docs/outputs.md b/docs/outputs.md
index 579b884..82af38c 100644
--- a/docs/outputs.md
+++ b/docs/outputs.md
@@ -144,13 +144,13 @@ Content details:
| tmp | various | - Can be deleted. - Blacklist index. - Intermediate FASTQs. - Genrich output reads. |
!!! note
-BAM files from `dedupBam` can be used for downstream footprinting analysis using [CCBR_TOBIAS](https://github.com/CCBR/CCBR_Tobias) pipeline
+ BAM files from `dedupBam` can be used for downstream footprinting analysis using [CCBR_TOBIAS](https://github.com/CCBR/CCBR_Tobias) pipeline
!!! note
-[bamCompare](https://deeptools.readthedocs.io/en/develop/content/tools/bamCompare.html) from deeptools can be run to compare BAMs from `dedupBam` for comprehensive BAM comparisons.
+ [bamCompare](https://deeptools.readthedocs.io/en/develop/content/tools/bamCompare.html) from deeptools can be run to compare BAMs from `dedupBam` for comprehensive BAM comparisons.
!!! note
-BAM files from `dedupBam` can also be converted to BED format and processed with [chromVAR](https://github.com/GreenleafLab/chromVAR) to identify variability in motif accessibility across samples and assess differentially active transcription factors from the JASPAR database.
+ BAM files from `dedupBam` can also be converted to BED format and processed with [chromVAR](https://github.com/GreenleafLab/chromVAR) to identify variability in motif accessibility across samples and assess differentially active transcription factors from the JASPAR database.
#### How consensus peaks are generated
@@ -212,13 +212,13 @@ sample3.consensus.bed ─► fixed-width peaks ─┘
!!! tip "Config knobs that control consensus"
-| Parameter | Default | Round | Effect |
-| -------------------------- | ------- | ----- | ---------------------------------------------------------------------------------------- |
-| `consensus_min_replicates` | `2` | 1 | Min. replicates a peak must appear in to be retained in per-sample consensus |
-| `consensus_min_spm` | `5` | 1 | Min. signal-per-million reads threshold for a peak to be included |
-| `roi_min_replicates` | `1` | 2 | Min. samples/replicates a fixed-width peak must appear in to be kept in the ROI set |
-| `roi_min_spm` | `2` | 2 | Min. signal-per-million reads threshold for a fixed-width peak to be kept in the ROI set |
-| `fixed_width` | `500` | 1 | Width (bp) of fixed-width peaks used to build the ROI set |
+ | Parameter | Default | Round | Effect |
+ | -------------------------- | ------- | ----- | ---------------------------------------------------------------------------------------- |
+ | `consensus_min_replicates` | `2` | 1 | Min. replicates a peak must appear in to be retained in per-sample consensus |
+ | `consensus_min_spm` | `5` | 1 | Min. signal-per-million reads threshold for a peak to be included |
+ | `roi_min_replicates` | `1` | 2 | Min. samples/replicates a fixed-width peak must appear in to be kept in the ROI set |
+ | `roi_min_spm` | `2` | 2 | Min. signal-per-million reads threshold for a fixed-width peak to be kept in the ROI set |
+ | `fixed_width` | `500` | 1 | Width (bp) of fixed-width peaks used to build the ROI set |
#### Counts matrices: reads vs Tn5 nicking sites, and `dedup` vs `nondedup`
@@ -448,10 +448,10 @@ while sample-level `*.consensus.bed` inputs use all consensus peaks. If your
replicate and consensus motif results differ, this is one reason why.
!!! tip
-If you need to confirm the exact HOMER settings used in a finished run,
-start with `motifFindingParameters.txt`. If you want to reproduce the AME
-input precisely, reuse the `target.fa` and `background.fa` files in the
-same output folder.
+ If you need to confirm the exact HOMER settings used in a finished run,
+ start with `motifFindingParameters.txt`. If you want to reproduce the AME
+ input precisely, reuse the `target.fa` and `background.fa` files in the
+ same output folder.
#### Interpreting motif enrichment results
diff --git a/docs/overview.md b/docs/overview.md
index 3678a5a..0a4267a 100644
--- a/docs/overview.md
+++ b/docs/overview.md
@@ -59,23 +59,23 @@ Genrich is integrated into the pipeline to complement MACS2. This tool is partic
By combining the strengths of MACS2 and Genrich, ASPEN delivers a comprehensive and reliable peak detection framework, facilitating downstream analyses and enabling researchers to uncover critical insights into chromatin accessibility and gene regulation.
!!! tip "Which peak caller should I use?"
-Both MACS2 and Genrich run automatically for every sample — you
-don't have to choose upfront. **Genrich is generally
-recommended for ATAC-seq** because it was purpose-built for
-assays like ATAC-seq/DNase-seq: it models Tn5 transposase cut
-sites directly (`-j` ATAC-seq mode) rather than adapting a
-ChIP-seq fragment-shift model, and it combines biological
-replicates natively via Fisher's method instead of requiring a
-separate consensus step. **MACS2** was originally designed for
-ChIP-seq and requires ATAC-specific parameter workarounds to
-approximate cut-site signal, but remains the field standard —
-it's included because many reviewers and downstream tools
-expect to see MACS2 peaks, and comparing both gives an extra
-sanity check. CCBR's internal benchmarking on ASPEN's ATAC-seq
-data has generally found Genrich peaks to be higher quality. If
-your MACS2 and Genrich DiffATAC results disagree substantially
-for a given region, treat that region as lower-confidence
-rather than assuming one caller is unconditionally "right".
+ Both MACS2 and Genrich run automatically for every sample — you
+ don't have to choose upfront. **Genrich is generally
+ recommended for ATAC-seq** because it was purpose-built for
+ assays like ATAC-seq/DNase-seq: it models Tn5 transposase cut
+ sites directly (`-j` ATAC-seq mode) rather than adapting a
+ ChIP-seq fragment-shift model, and it combines biological
+ replicates natively via Fisher's method instead of requiring a
+ separate consensus step. **MACS2** was originally designed for
+ ChIP-seq and requires ATAC-specific parameter workarounds to
+ approximate cut-site signal, but remains the field standard —
+ it's included because many reviewers and downstream tools
+ expect to see MACS2 peaks, and comparing both gives an extra
+ sanity check. CCBR's internal benchmarking on ASPEN's ATAC-seq
+ data has generally found Genrich peaks to be higher quality. If
+ your MACS2 and Genrich DiffATAC results disagree substantially
+ for a given region, treat that region as lower-confidence
+ rather than assuming one caller is unconditionally "right".
### 🤝 **Consensus Peaks**
@@ -116,21 +116,21 @@ ASPEN employs custom scripts to analyze the distribution of fragment lengths wit
To evaluate the sufficiency of sequencing depth and detect potential biases introduced during PCR amplification, ASPEN utilizes Preseq to estimate library complexity, reporting the Non-Redundant Fraction (`NRF`) and PCR Bottlenecking Coefficients (`PBC1`, `PBC2`). This metric helps determine whether the sequencing effort is adequate to capture the diversity of the library, ensuring that the data is representative of the underlying chromatin landscape. By identifying potential saturation or over-representation of certain fragments, researchers can assess the reliability of their sequencing results.
!!! tip "Rule of thumb"
-Per [ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards), the preferred values are **`NRF` > 0.9, `PBC1` > 0.9, and `PBC2` > 3**. Lower values indicate a less complex library (e.g. over-amplified by PCR), which can inflate apparent signal at a subset of loci rather than reflecting true biological accessibility.
+ Per [ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards), the preferred values are **`NRF` > 0.9, `PBC1` > 0.9, and `PBC2` > 3**. Lower values indicate a less complex library (e.g. over-amplified by PCR), which can inflate apparent signal at a subset of loci rather than reflecting true biological accessibility.
### 🧬 **Transcription Start Site (TSS) Enrichment**
ASPEN calculates TSS enrichment scores, a widely recognized quality metric for ATAC-seq data. These scores measure the accumulation of sequencing reads around transcription start sites (TSS), which are hallmark regions of open chromatin. High TSS enrichment scores indicate well-prepared libraries with minimal technical artifacts, as they reflect the accessibility of promoter regions and the integrity of the chromatin preparation process.
!!! tip "Rule of thumb"
-[ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards) define annotation-dependent TSS enrichment cutoffs — for example, using a GRCh38 RefSeq TSS annotation: **< 5 is concerning, 5-7 is acceptable, and > 7 is ideal**. **Caveat:** ASPEN builds its TSS bins from GENCODE (not RefSeq) gene annotations (see `resources/tssBed/`), so ENCODE's exact per-annotation cutoffs may not transfer precisely to ASPEN's TSS enrichment values — treat these numbers as directional guidance (aim for high single digits or higher) rather than an exact pass/fail threshold.
+ [ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards) define annotation-dependent TSS enrichment cutoffs — for example, using a GRCh38 RefSeq TSS annotation: **< 5 is concerning, 5-7 is acceptable, and > 7 is ideal**. **Caveat:** ASPEN builds its TSS bins from GENCODE (not RefSeq) gene annotations (see `resources/tssBed/`), so ENCODE's exact per-annotation cutoffs may not transfer precisely to ASPEN's TSS enrichment values — treat these numbers as directional guidance (aim for high single digits or higher) rather than an exact pass/fail threshold.
### 📊 **Fraction of Reads in Peaks (FRiP)**
The Fraction of Reads in Peaks (FRiP) score quantifies the proportion of sequencing reads that fall within identified peaks, serving as a measure of the signal-to-noise ratio in the dataset. Higher FRiP scores indicate datasets with strong, biologically meaningful signals and minimal background noise. Additionally, ASPEN computes the fraction of reads localized to specific genomic features, such as promoters, enhancers, and DNase hypersensitive sites (DHS). These feature-specific FRiP scores provide further insights into the quality and biological relevance of the data.
!!! tip "Rule of thumb"
-Per [ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards), a FRiP score **> 0.3** indicates high-quality data, though values **> 0.2** may still be acceptable. Consistently lower scores suggest poor signal-to-noise and warrant a closer look at library prep or peak-calling parameters.
+ Per [ENCODE's ATAC-seq data standards](https://www.encodeproject.org/atac-seq/#standards), a FRiP score **> 0.3** indicates high-quality data, though values **> 0.2** may still be acceptable. Consistently lower scores suggest poor signal-to-noise and warrant a closer look at library prep or peak-calling parameters.
---
@@ -151,11 +151,11 @@ If a project later needs de novo motif discovery, that can be added as a
separate workflow enhancement rather than mixed into the default run.
!!! tip "Rule of thumb"
-Start by looking for motif families that are strong in **both** HOMER and
-AME. In HOMER, focus on motifs with very small `p-value`/`q-value` values
-and a clear increase in `% of Target Sequences with Motif` relative to
-background. In AME, focus on low `adj_p-value`/`E-value` hits where `%TP`
-is clearly higher than `%FP`.
+ Start by looking for motif families that are strong in **both** HOMER and
+ AME. In HOMER, focus on motifs with very small `p-value`/`q-value` values
+ and a clear increase in `% of Target Sequences with Motif` relative to
+ background. In AME, focus on low `adj_p-value`/`E-value` hits where `%TP`
+ is clearly higher than `%FP`.
Detailed file locations and interpretation notes for `knownResults.txt`,
`ame_results.txt`, `target.fa`, and `background.fa` are documented in
@@ -212,18 +212,19 @@ In ASPEN, if spike-in data is present:
This spike-in-derived scaling factor allows the comparison of chromatin accessibility across conditions even when global chromatin accessibility levels differ (e.g., treatment-induced repression or global decondensation).
!!! tip "Should I turn on spike-in normalization?"
-Ask yourself: do I expect a **global, genome-wide shift** in chromatin accessibility between my conditions — rather than just **localized** changes at a handful of specific regulatory elements?
+ Ask yourself: do I expect a **global, genome-wide shift** in chromatin accessibility between my conditions — rather than just **localized** changes at a handful of specific regulatory elements?
- - **Yes** (e.g. a chromatin remodeler inhibitor, a broad transcription factor knockdown/knockout, drug-induced chromatin modulation) → turn on spike-in normalization. Standard depth-based normalization (DESeq2 size factors) assumes *most* regions are unchanged between conditions — that assumption breaks down under a genome-wide shift, and spike-in gives you an external, biology-independent scale instead.
- - **No** (you expect differences to be confined to specific loci/pathways, with most of the genome unchanged) → spike-in is probably unnecessary. It adds experimental complexity (extra reagents, a second alignment step, and a "0 spike-in reads" failure mode to manage — see the warning below) for a scenario DESeq2's built-in normalization already handles well.
+ **Yes** (for example, a chromatin remodeler inhibitor, a broad transcription factor knockdown/knockout, or drug-induced chromatin modulation) → turn on spike-in normalization. Standard depth-based normalization (DESeq2 size factors) assumes *most* regions are unchanged between conditions, and that assumption breaks down under a genome-wide shift. Spike-in gives you an external, biology-independent scale instead.
+
+ **No** (you expect differences to be confined to specific loci/pathways, with most of the genome unchanged) → spike-in is probably unnecessary. It adds experimental complexity (extra reagents, a second alignment step, and a "0 spike-in reads" failure mode to manage — see the warning below) for a scenario DESeq2's built-in normalization already handles well.
See [Enabling Spike-In Normalization](deployment.md#enabling-spike-in-normalization-optional) for the config steps once you've decided to use it.
ASPEN performs spike-in-aware normalization transparently, and reports both raw and normalized counts in the final output matrix for differential analysis. This ensures flexibility in downstream interpretation while preserving the ability to adjust for systemic experimental artifacts.
!!! warning "What if a replicate has 0 spike-in reads?"
-If `spikein: true` is set but a replicate has **zero reads** aligned to the spike-in genome (e.g. the spike-in material wasn't actually added to that library, the spike-in genome/index is misconfigured, or a genuinely contamination-free host-only library), ASPEN cannot compute a scaling factor for it and the run will **fail with a clear error message** naming the affected replicate(s) rather than silently producing `Inf`/`NaN` normalized counts. To resolve this: verify the spike-in genome/index path in `config.yaml`, confirm spike-in material was actually included during library prep for that replicate, or set `spikein: false` if none of your samples have spike-in material.
+ If `spikein: true` is set but a replicate has **zero reads** aligned to the spike-in genome (e.g. the spike-in material wasn't actually added to that library, the spike-in genome/index is misconfigured, or a genuinely contamination-free host-only library), ASPEN cannot compute a scaling factor for it and the run will **fail with a clear error message** naming the affected replicate(s) rather than silently producing `Inf`/`NaN` normalized counts. To resolve this: verify the spike-in genome/index path in `config.yaml`, confirm spike-in material was actually included during library prep for that replicate, or set `spikein: false` if none of your samples have spike-in material.
### 📊 **Reporting**
-ASPEN enhances differential chromatin accessibility analysis by providing an interactive HTML report generated from DESeq2 results, featuring various visualizations. Additionally, it offers a TSV (tabl-delimited) file with integrated gene annotations, compatible with Microsoft Excel, enabling efficient data manipulation and facilitating the identification of genes near regions with altered accessibility. The Excel file is aggregated across all different contrasts queried in the project.
+ASPEN enhances differential chromatin accessibility analysis by providing an interactive HTML report generated from DESeq2 results, featuring various visualizations. Additionally, it offers a TSV (tab-delimited) file with integrated gene annotations, compatible with Microsoft Excel, enabling efficient data manipulation and facilitating the identification of genes near regions with altered accessibility. The Excel file is aggregated across all different contrasts queried in the project.