Spatial Proteomics: In Situ PLA data
The input to the notebooks is the output from the pipex_bigfish pipeline (a fork of pipex). This pipeline transforms CODEX-derived images and isPLA data into AnnData files (.h5ad). Each .h5ad file contains, for each cell, the mean marker intensity as well as the number of isPLA dots.
Here is a summary of the data processing steps:
Input:
- qptiff files (8-bit compressed)
PIPEX steps:
- Cell segmentation using Stardist (default settings: nuclei diameter = 20, cytoplasm expansion = 20)
- Computation of mean intensities per cell and per marker
BigFish dot detection:
- isPLA image is denoised and spots are enhanced using a Laplacian of Gaussian (LoG) filter
- Peaks are detected in the filtered image with a local maximum detection algorithm
- An intensity threshold is applied to discriminate actual spots from noisy background
- Dense regions decomposition: detects dense and bright regions with potential clustered spots, then uses gaussian simulations to correct misdetection in these regions
- Cluster detection (cluster spots in point cloud and detect relevant aggregated structures) — not used in this study
- Stores number of isPLA spots and clusters per cell
Output:
- AnnData objects with cell information (marker intensities, isPLA spots)
H&E manual alignment:
- For each tissue sample, a consecutive slice with H&E was manually aligned to the DAPI image to assist with manual annotations
Manual annotations:
- Regions with incorrect isPLA signal (edge effects, tears, over-saturation, etc.) were manually annotated and all cells within these regions are removed from the AnnData objects for the rest of the study
Tumor regions:
- Tumor and non-tumor regions were manually annotated. All cells within tumor regions are labeled as “in_tumor” for downstream analysis.
The first notebook, notebooks/1_analyse_h5ad.ipynb, processes spatial proteomics data stored in AnnData (.h5ad) files. It performs the following steps:
- Loads processed data for each sample.
- Applies geometric transformations for spatial alignment.
- Removes edge effects and adds tumor region annotations if available.
- Computes positive cells for each marker using Otsu thresholding.
- (Optional) Classifies immune cells based on marker expression.
- Saves the processed data for downstream analysis.
The third notebook, notebooks/2_generate_tmap.ipynb, prepares interactive TissUUmaps projects for spatial exploration of the processed data. It performs the following steps:
- Loads processed
.h5adfiles and associated image data for each sample. - Converts and rescales TIFF images for visualization, including biomarker and H&E images.
- Updates TissUUmaps project state with available image layers and metadata.
- Adjusts visualization settings (e.g., layer visibility, opacity, marker display).
- Writes updated
.h5adand.tmapproject files for each sample. - Optionally extracts individual biomarker images from multiplexed TIFF files.
- Copies generated TissUUmaps project files to a separate directory for sharing or deployment.
The fourth notebook, notebooks/3_concatenate_anndata.ipynb, combines processed AnnData (.h5ad) files from all samples into a single dataset for downstream analysis. It performs the following steps:
- Searches for all processed
.h5adfiles in the output directories. - Loads each file and adds a sample identifier to the metadata.
- Concatenates all AnnData objects into one unified AnnData object.
- Ensures unique observation names across samples.
- Saves the combined dataset as
all_adata.h5adfor further analysis.
The fifth notebook, notebooks/4_figures.ipynb and notebooks/5_figures_supp.ipynb, create the figures used in the main paper and supplementary materials, respectively.