Skip to content
jsonmadPublic

About

Python package for converting multiplex immunofluorescence single-cell exports from QuPath into AnnData objects for spatial analysis.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

125 Commits

Folders and files

Repository files navigation

QXYCell

QXYCell badge

QXYCell acts as a Python bridge to convert cell measurements and spatial data from multiplex immunofluorescence images processed in QuPath (QuPath; GPLv3) into an AnnData object for cell typing, visualization, and spatial analysis. The resulting .h5ad object can be used with downstream spatial analysis tools.

Workflow

QXYCell workflow from multiplex tissue imaging through QuPath, QXYCell, AnnData, spatial plots, and downstream analysis

Installation

QXYCell requires Python 3.10 or newer.

Install in new environment

git clone https://github.com/jsonmad/QXYCell.git
cd QXYCell
conda env create -f environment.yml
conda activate qxycell

Verify the installation

python -c "import qxycell; print('QXYCell import OK')"
qxycell --help

Update if required (Navigate to the cloned repository before running)

conda activate qxycell
git pull
conda env update -f environment.yml --prune

Prepare data in QuPath

Before running QXYCell, follow the QuPath preparation guide. to create these four inputs:

  1. Sample, tissue-feature, and imaging-artifact annotations exported as GeoJSON
    • QuPath > File > Export objects as GeoJSON
  2. Cell segmentation and cell measurements exported as .tsv or .csv
    • QuPath > Measure > Export measurements
  3. Single-object classifiers for channel/marker thresholds (.json)
    • QuPath > Classify > Object classification > Create single measurement classifier
  4. Cell-boundary geometry exported as GeoJSON
    • QuPath > Objects > Select > Select detections > Select cells
    • QuPath > File > Object data.. > Export as GeoJSON

Quick start

  • Run this quickstart in a Marimo notebook.
  • Each data-processing stage checkpoints (saves) the current adata object to the active output folder.
  • Run each stage below in order, using a separate Marimo cell.
  • Pause where noted to review thresholds and cell-type logic YAML.
  • To begin, define the path to the QuPath project directory and an output directory for the QXYCell results.
  • The output directory can be anywhere but should not sit inside of the QuPath project directory.

Activate the qxycell environment and start Marimo.

conda activate qxycell
marimo edit

Marimo opens in a browser. Save the new notebook as a Python file, for example qxycell_workflow.py, then run the Python code below one stage at a time in separate cells. On the OVD, start Marimo after changing to the folder that contains the notebook and project files so relative paths resolve there.

import qxycell as qxy

project_dir = r"\path\to\qupath_project"
output_dir = r"\path\to\outputs\run_1"
# Stage 1: import cell measurements and create the AnnData checkpoint.

adata = qxy.import_cells(project_dir, output_dir=output_dir)
# Stage 2: add or refresh annotations and optional cell polygons.
# Annotation names containing "sample" automatically define adata.obs["Sample"].

qxy.add_annotations(adata, pixel_size_um=0.28)

# Optional Stage 2b: choose the identifier string used in annotation names to remove cells in regions with imaging or tissue artifacts.

qxy.remove_cells(adata, remove_cells="ignore")
qxy.remove_cells(adata, remove_cells="folded_tissue")
# Stage 3: use classifier JSON values directly, or use a reviewed table as the
# final source. Classifier thresholding saves applied values to
# thresholds/classifier_thresholds.tsv; copy and edit that file to refine them.

# Stage 3A: apply thresholds from QuPath object-classifier JSON files.

qxy.threshold_from_classifiers(adata)

# Stage 3B: generate or manually refine a table, then apply it as the final source.

threshold_table = qxy.generate_threshold_table(
    project_dir,
    output_dir=output_dir,
    )

# Pause here to review and fill every per-image threshold at <output_dir>/thresholds/thresholds_YYMMDD-HHMM.tsv.

qxy.threshold_from_table(adata, threshold_table)
# Stage 4: generate the prompt used to draft celltype_logic.yaml.

qxy.celltype_prompt(
    adata,
    context="Describe the tissue and expected populations",
    )

# Pause for biology domain expert review, save the reviewed cell type YAML, then continue.
# Stage 5: assign cell types using the reviewed cell-type logic.

celltype_summary = qxy.celltype(
    adata,
    "/path/to/celltype_logic.yaml",
)

# If no path is specified, QXYCell defaults to the newest .yaml or .yml file in the active output folder’s celltype/ directory.
# Stage 6: Visaully sanity check the assigned cell types spatially.

qxy.plot_spatial(adata, category_col="celltype", show=True)
# Stage 7: plot marker positivity and intensity by assigned cell type.

qxy.plot_marker_positivity_heatmap(
    adata,
    category_col="celltype",
    show=True,
    )

qxy.plot_marker_intensity_heatmap(
    adata,
    category_col="celltype",
    show=True,
    )
# Stage 8: Use sample annotations instead of whole images.

qxy.plot_marker_positivity_heatmap(
    adata,
    category_col="Sample",
    show=True,
    )

qxy.plot_marker_intensity_heatmap(
    adata,
    category_col="Sample",
    show=True,
    )
  • Each successful stage updates the active .h5ad and refreshes tables/cells_obs.csv and tables/markers_var.csv.
  • If annotations are updated after cells have been removed, rerun adata = qxy.import_cells(project_dir, output_dir=output_dir) before refreshing annotations and removing cells again.
  • Classifier thresholding qxy.threshold_from_classifiers(adata) saves the applied values to <output_dir>/thresholds/classifier_thresholds.tsv.
  • Table thresholding qxy.threshold_from_table(adata, "<output_dir>/thresholds/thresholds_YYMMDD-HHMM.tsv") uses only the named reviewed table. To refine classifier values, copy or rename classifier_thresholds.tsv, update the copied table, then apply that copy with table thresholding.
  • You can exit exit() after any stage finishes successfully. To restart, activate the same environment, start a new interactive session, recreate the path variables, and load the .h5ad from the output_dir or the exact .h5ad path.
# restarting a session
import qxycell as qxy

project_dir = "/path/to/qupath_project"
output_dir = "/path/to/outputs/run_1"
adata = qxy.load(output_dir)

# or exact path to .h5ad
adata = qxy.load("/path/to/outputs/run_1/h5ad/qxycell.h5ad")

Documentation

Guide Use it for
QuPath quick start Streamlined exports, threshold choice, and cell typing
QuPath preparation Preparing images, segmenting cells, measuring features, and exporting QuPath assets
Running the staged workflow Checkpoints, rerun rules, output folders, validation, and the optional single-call workflow
Sample metadata Matching experimental, clinical, and batch metadata to images, samples, or TMA cores
Cell typing Prompt generation, reviewed YAML rules, assignment diagnostics, validation, and reruns
Plotting Spatial figures, cell boundaries, annotation polygons, bars, heatmaps, formats, and palettes
Cellular neighbourhoods Local composition profiles, clustering, naming, parameter review, and neighbourhood plots
AnnData structure and outputs Stored fields, dataset summaries, provenance, output files, and save/load behavior

Additional reference material:

Support and license

Report reproducible bugs and feature requests through GitHub Issues. Report suspected vulnerabilities privately as described in SECURITY.md. QXYCell is released under the MIT License.

About

Python package for converting multiplex immunofluorescence single-cell exports from QuPath into AnnData objects for spatial analysis.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages