Product Documentation

Overview

In the CM4AI project, 200 genes/proteins are the subject of coordinated experiments in three data modalities. The gene/proteins comprise 100 chromatin modifiers and 100 metabolic enzymes. Data has been collected on the chromatin modifiers in Year 1 of the project. We selected coherent groups of genes that are known to interact in processes involved in cancer, neuropsychiatric, and cardiac disorders. Data is being acquired in the triple-negative breast cancer cell line MDA-MB-468, including upon treatment with paclitaxel or vorinostat, and in two iPSC lines in the undifferentiated state as well as in differentiated neurons and cardiomyocytes using complementary mapping approaches. We are using proteomic mass spectrometry to map protein-protein interactions, cellular imaging to map the spatial subcellular organization of proteins, and genetic perturbation screens via CRISPR/Cas9 to assess the transcriptome-wide impact of gene perturbations.

Affinity Purification Mass spectrometry (AP-MS) Data

Affinity Purification-Mass Spectrometry (AP-MS) is a powerful technique used to uncover the intricate networks of protein interactions within cells. Different methods can be used but the most popular approach is to attach a purifiable tag to a protein of interest, the “bait”. In the context of CM4AI the tag is directly attached to the endogenous protein of interest using gene editing to maintain physiological level of expression. The tagged bait is then purified from cell extract with its interacting partners. Once separated, the captured proteins are identified using a mass spectrometer allowing to infer protein-protein interactions. Understanding these interactions allows can unravel complex cellular processes, discover new roles for proteins, and gain insights into diseases caused by disrupted protein networks.

AP-MS Data acquisition and processing

  • The raw data are acquired on an Orbitrap Fusion™ Lumos™ Tribrid™ Mass Spectrometer (Thermo Scientific)
  • The raw data are processed using Maxquant v1.6.12.0 using default parameter with LFQ quantification and the “202209_uniprot_human_reviewed.fasta” file as reference.
  • The evidence file from maxquant is processed using ArtMS in R to generate SAINT input files using the following command lines. 
  • SAINTexpress was run locally using version SAINTexpress_v3.6.3 

SAINT is designed to assess the reliability and significance of protein-protein interactions identified through mass spectrometry experiments. It takes into account both the presence and abundance of proteins in different experimental conditions or replicates. The method uses a probabilistic model to estimate the probability that a given interaction is a true positive rather than a random or non-specific interaction. This approach helps to distinguish true interactions from noise in large-scale protein interaction datasets.

https://pubmed.ncbi.nlm.nih.gov/21131968

https://saint-apms.sourceforge.net/Main.html

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4102138

Immunofluorescence Imaging Data

The cell imaging takes place in the Emma Lundberg lab (https://ell-core.stanford.edu/). We will provide insight into the expression and spatiotemporal distribution of up to 700 proteins, stained in the breast cancer cell line MDA-MB-468 or the iPSC line KOLF2.1J, either under untreated conditions or treated with the chemotherapeutic drug Paclitaxel or Vorinostat (SAHA). For each protein, the subcellular distribution of the protein is investigated by immunofluorescence staining and confocal microscopy. After image acquisition, we classify the subcellular localization of the protein in one or more of 35 different organelles and fine subcellular structures. The generated data set will allow us to see how treatment with Paclitaxel or Vorinostat influences the subcellular distribution of proteins. Finally, this data set will be contributed to build MuSIC maps, as demonstrated in this publication (https://www.nature.com/articles/s41586-021-04115-9).

The immunofluorescence method in detail

In the immunofluorescence staining procedure (ICC-IF), each protein of interest is visualized by using an antibody that detects the protein. The antibodies are derived from The Human Protein Atlas (https://proteinatlas.org) project. The antibody-based staining of the protein is shown in green in the generated images. During the staining procedure, the sample is also co-stained for additional structural markers of the cell: nuclei (stained with DAPI, shown in blue in the generated images), Microtubuli (MT) (stained with a Tubulin antibody, shown in red in the generated images), and endoplasmic reticulum (ER) (stained with a Calreticulin antibody, shown in yellow in the generated images).

Current stage (Version 0.1alpha)

We have mapped the protein localization for the first 95 genes in the cell line MDA-MB-468, either under untreated conditions or treated with Paclitaxel or treated with Vorinostat. We noticed that the drug treatment largely influenced the cellular morphology (see example images below). After paclitaxel treatment, nuclear morphology drastically changed. The Paclitaxel and the Vorinostat treatment each reduced the number of cells present at the end of the experiment. Thus, we have seeded more cells in those wells that were treated with Paclitaxel or Vorinostat compared to the wells that were left untreated. Nonetheless, fewer cells than in the untreated condition were observed at the end of the experiment.

For some proteins we observed similar subcellular localization in all conditions:

Other proteins appeared with heterogeneous subcellular protein localization patterns and showed a shift in the percentage of cells showing a certain type of subcellular protein distribution:

In the next steps, we will analyze what the specific differences are in more detail and apply this workflow for staining more proteins to get an even more comprehensive picture of the subcellular protein landscape in breast cancer cells treated with Paclitaxel or Vorinostat, compared to untreated cells.

 

CRISPR Screening Data

  • Single-cell capture with Chromium Next GEM Chip M Single Cell Kit (10x Genomics)
  • FASTQs sequencing data was acquired on a NovaSeq S4 (Illumina)
  • Alignment of the FASTQ files on Human Genome GRCh38 and the CRISPR library was done with Cell Ranger 7.1.0 Software (10x Genomics)

 

Gene NCBI Gene ID Molecular function Associated Disease
PARP1 142 ADPRibosylation Intellectual disability / cancer
BAP1 8314 Ubiquitin regulation Cancer
USP7 7874 Ubiquitin regulation Intellectual disability / cancer
NEDD4 4734 Ubiquitin regulation Brugada syndrome
BRCA1 672 Ubiquitin regulation Intellectual disability / Heart disease / Cancer
BARD1 580 Ubiquitin regulation Intellectual disability / Heart disease / Cancer
BRE1A 56254 Ubiquitin regulation Autism / heart disease
CBL 867 Ubiquitin regulation Cancer
MSL1 339287 Ubiquitin regulation Intellectual disability / Heart disease / Cancer
PHF6 84295 Ubiquitin regulation Intellectual disability
TAF1 6872 Ubiquitin regulation Congenital heart disease / Mental retardation
UHRF1 29128 Ubiquitin regulation Autism
UHRF2 115426 Ubiquitin regulation Congenital heart disease
KAT2A 2648 Histone Acetyl Transferase Autism
KAT3A 1387 Histone Acetyl Transferase Intellectual disability / Heart disease / Cancer
KAT3B 2033 Histone Acetyl Transferase Intellectual disability / Heart disease / Cancer
KAT6A 7994 Histone Acetyl Transferase Intellectual disability / Heart disease / Cancer
KAT6B 23522 Histone Acetyl Transferase Intellectual disability / Heart disease / Cancer
HDAC1 3065 Deacetylase Cancer / Cardiac defect
HDAC2 3066 Deacetylase Cancer / Cardiac defect
HDAC3 8841 Deacetylase Cancer
HDAC4 9759 Deacetylase Cancer
HDAC5 10014 Deacetylase Cancer / Cardiac defect
HDAC6 10013 Deacetylase Cancer
HDAC7 51564 Deacetylase Cancer / Vascular integrity
HDAC8 55869 Deacetylase Cancer
HDAC9 9734 Deacetylase Cancer / Cardiac defect
HDAC10 83933 Deacetylase Cancer
HDAC11 79885 Deacetylase cancer
SIRT2 22933 Deacetylase Cancer / Acute myocardial infarction 
KDM3B 51780 Lysine Demetylase Schizophrenia
KDM4C 23081 Lysine Demetylase Upper aerodigestive tract cancer, association with
KDM5A 5927 Lysine Demetylase Intellectual disability / Heart disease / breast cancer
KDM5B 10765 Lysine Demetylase Intellectual disability / Heart disease
KDM5C 8242 Lysine Demetylase Intellectual disability
KDM6A 7403 Lysine Demetylase Kabuki syndrome / Autism / Cancer
KDM6B 23135 Lysine Demetylase Intellectual disability
TET1 80312 Lysine Demetylase Autism?
TET2 54790 Lysine Demetylase Myeloproliferative neoplasms / prostate cancer
TET3 200424 Lysine Demetylase Intellectual disability
ATM 472 Kinase Ataxia telandiectasia / Cancer
ATR 545 Kinase Intellectual disability / Cancer
AURKB 9212 Kinase Cancer
AURKA 6790 Kinase Cancer
CSNK2A1 1457 Kinase neurological diseases / inflamation / Cancer
HASPIN 83903 Kinase Cancer
JAK2 3717 Kinase Myeloproliferative neoplasms / Erythrocytosis
MSK1 855651 Kinase Cancer
MSK2 8986 Kinase Cancer
5292 5292 Kinase Autism
PRKCA 5578 Kinase Cardiomyopathy, Autism / Cancer
PRKCB 5579 Kinase Autism
PRKCD 5580 Kinase Immunodeficiency
BAZ1B 9031 Kinase Autism / heart disease
PRKAA1 5562 AMPK Cancer / Cardio vascular syndrome
PRKAA2 5563 AMPK Cancer
PRKAB1 5564 AMPK Cancer / Neurological disease
PRKAB2 5565 AMPK Cancer
PRKAG1 5571 AMPK Cancer / Cardio vascular syndrome
PRKAG2 51422 AMPK Cancer / Cardio vascular syndrome
PRKAG3 53632 AMPK Cancer / Cardio vascular syndrome
DNMT1 1786 Lysine Methyltransferase Neuropathy / Cerebellar ataxia / cancer
DNMT3A 1788 Lysine Methyltransferase Intellectual disability / cancer
DNMT3B 1789 Lysine Methyltransferase Kabuki syndrome / Cancer
KMT1D 79813 Lysine Methyltransferase Intellectual disability / Heart disease / cancer
KMT1F 83852 Lysine Methyltransferase Autism / Cancer
KMT2A 4297 Lysine Methyltransferase Intellectual disability / Heart disease / cancer
KMT2C 58508 Lysine Methyltransferase Intellectual disability / Heart disease / cancer
KMT2D 8085 Lysine Methyltransferase Kabuki syndrome 
KMT2E 55904 Lysine Methyltransferase Intellectual disability / Heart disease
KMT2H 55870 Lysine Methyltransferase Intellectual disability / Heart disease
KMT3B 64324 Lysine Methyltransferase Intellectual disability / Heart disease / cancer
PRMT1 3276 Lysine Methyltransferase Cancer
EYA1 2138 Phosphatase Cancer / BOR syndrome
EYA3 2140 Phosphatase Cancer
PPP1CA 5499 Phosphatase Encephalitis / Cancer
PPP1CB 5500 Phosphatase Cancer
PPP1CC 5501 Phosphatase Cancer
PPP4C 5531 Phosphatase Autism / Cancer
BPTF 2186 Reader Neurodevelopmental disorder / Cancer
BRD4 23476 Reader Heart disease / neuro-dev delay / cancer
BRD7 29117 Reader Cancer
BRPF1 7862 Reader Intellectual disability / Cardiac disorder / Cancer
CHD4 1108 Reader Intellectual disability
ING1 3621 Reader Cancer
TAF1 6872 Reader Intellectual development disorder / Cancer
TRIM24 8805 Reader Cancer
YWHAZ 7534 Reader Intellectual disability / Cardiac disorder / Cancer
YWHAE 7531 Reader Cancer
YWHAQ 10971 Reader Creutzfeldt-Jakob Disease / Cancer
YWHAB 7529 Reader Cancer
YWHAH 7533 Reader Schizophrenia
YWHAG 7532 Reader Intellectual disability
BRG1 6597 Reader Nervous an cardiovascular diseases / Cancer
CBX2 84733 Reader Cancer
CBX3 11335 Reader Cancer
MRG15 10933 Reader Cancer
P/CAF 8850 Reader Cancer
53BP1 7158 Reader Microcephaly / Cancer
BRD3 8019 Reader Intellectual disability / Cancer
CHD1 1105 Reader Neurodevelopmental disorder / Cancer
MBTD1 54799 Reader Cancer
PDP1 54704 Reader Cancer
RAG2 5897 Reader Cancer / Immunodeficiency
Gene NCBI Gene ID Molecular function
FOLR1 2348 Folate/Methionine Cycle
MTHFR 4524 Folate/Methionine Cycle
MTR 4548 Folate/Methionine Cycle
MTHFD1 4522 Folate/Methionine Cycle
MAT1A 4143 Folate/Methionine Cycle
MAT2A 4144 Folate/Methionine Cycle
MAT2B 27430 Folate/Methionine Cycle
AHCY 191 Folate/Methionine Cycle
SHMT1 6470 Folate/Methionine Cycle
RFK 55312 FAD
FLAD1 80308 FAD
TKFC 26007 FAD
ACAD9 28976 Acethyl-CoA
ACADM 34 Acethyl-CoA
ACADS 35 Acethyl-CoA
ACADSB 36 Acethyl-CoA
ACADVL 37 Acethyl-CoA
ACAT1 38 Acethyl-CoA
ACOT7 11332 Acethyl-CoA
ACOX3 8310 Acethyl-CoA
ACSS2 55902 Acethyl-CoA
ACSS3 79611 Acethyl-CoA
GMPS 8833 Purine metabolism
LDHA 3939 Lactate metabolism
LDHB 3945 Lactate metabolism
LDHD 197257 Lactate metabolism
ACLY 47 TCA cycle
ALDH5A1 7915 TCA cycle
DLAT 1737 TCA cycle
DLD 1738 TCA cycle
GLUD1 2746 TCA cycle
GLUD2 2747 TCA cycle
IDH1 3417 TCA cycle
IDH3A 3419 TCA cycle
IDH3B 3420 TCA cycle
IDH3G 3421 TCA cycle
MDH2 4191 TCA cycle
OGDH 4967 TCA cycle
PDHA1 5160 TCA cycle
PDHA2 5161 TCA cycle
PDHB 5162 TCA cycle
SDHA 6389 TCA cycle
G6PD 2539 NADH
GAPDH 2597 NADH
NADK 65220 NADH
NDUFS4 4724 NADH
NDUFAB1 4706 NADH
NMNAT2 23057 NADH
NNT 23530 NADH
PGD 5226 NADH
PGM1 5236 NADH
PHGDH 26227 NADH
PYGL 5836 NADH
SRC 6714 Kinases
KIT 3815 Kinases
ITPKA 3706 Kinases
AKT1 207 Kinases
BRAF 673 Kinases
CDK4 1019 Kinases
DAPK3 1613 Kinases
FYN 2534 Kinases
PIK3CA 5290 Kinases
PIK3C2A 5286 Kinases
NUAK1 9891 Kinases
STK4 6789 Kinases
MAPK1 5594 Kinases
MTOR 2475 Kinases
PIP5K1C 23396 Kinases
RIPK1 8737 Kinases
RPS6KA3 6197 Kinases
KRAS 3845 Kinases
TGFBR2 7048 Kinases
WEE1 7465 Kinases
PPA2 27068 other
AHCYL2 23382 other
CYP27A1 1593 other
DIO3 1735 other
GSTM1 2944 other
HEXA 3073 other
NAGS 162417 other
NCOA1 8648 other
PANK2 80025 other
PARP4 143 other
PARP6 56965 other
PLCD4 84812 other
PLCG1 5335 other
SLC2A1 6513 other
UGDH 7358 other
UGT8 7368 other
UBE3C 9690 other
CYP24A1 1591 other
GCLM 2730 other
GYG2 8908 other
PDE4A 5141 other
PLPP7 84814 other
AMDHD2 51005 other
ABHD6 57406 other

Each dataset is stored as a RO-Crate package. Datasets contain the data for a gene set measured in a specific cell line under a specific treatment. (Including “untreated”)

The RO-Crate format

Research Object Crate (RO-Crate) is a lightweight standard for packaging research data with their metadata. It contains schema.org and other standard structured vocabulary annotations in JSON-LD, of all included or referenced data objects, and aims to make best-practice in formal metadata description accessible and practical for use in a wide variety of situations. Each RO-Crate is a folder that contains, along with one or more data files, a ro-crate-metadata.json file, containing structured information about the data (metadata) in JSON-LD format. It is good practice to also include a readme.md text file describing the contents.

Image Datasets

Each RO-Crate contains:

  • A cell line – antibody – treatment table describing the experiment
  • Four channel data folders: red, green, blue yellow
    • Each folder contains images in .jpeg format.
  • ro-crate-metadata.json
    • Structured metadata for the dataset
    • README

Example README file:

This data set has been generated in the Emma Lundberg lab (https://ell-core.stanford.edu/) in July-August 2023 as part of the CM4AI Research Program (https://digitaltumors.org/).

This data set displays the spatial localization of 95 proteins in the breast cancer cell line MDA-MB-468. MDA-MB-468 cells were seeded into 96-well glass bottom plates and treated either with the chemotherapeutic drug Paclitaxel or with the chemotherapeutic drug Vorinostat (SAHA) or not treated. For drug treatment, higher cell numbers were seeded than for the cells that were not treated to reach a sufficient amount of remaining cells at the end of the experiment (treatment with Paclitaxel or Vorinostat disturbs cell growth). After fixation of the cells with PFA, for each protein, the subcellular distribution of the protein was investigated by immunofluorescence-based staining (ICC-IF) and confocal microscopy. Each well was stained with DAPI (to label nuclei, “blue” channel), a Calreticulin antibody (to label the endoplasmic reticulum, “yellow” channel), a Tubulin antibody (to label Microtubuli, “red” channel), and an antibody against a protein of interest (a different antibody / protein for each well, “green” channel). In the green channel, for each well, one of 95 proteins was stained with an antibody from the proteinatlas.org project. 

The .tsv file describes each image in the data set. Each row represents one image. The columns describe the staining from which the image was taken:

“Antibody ID” describes the antibody ID for the antibody applied to stain the protein visible in the “green” channel. The antibody ID can be looked up at proteinatlas.org to find out more information about the antibody.

“ENSEMBL ID” indicates the ENSEMBL ID(s) of the gene(s) of the proteins visualized in the “green” channel.

Treatment refers to how the cells that are depicted in the image were treated (with Paclitaxel, Vorinostat, or untreated)

“Well” refers to the well coordinate on the 96-well plate

“Region” is a unique identifier for the position in the well, where the cells were acquired.

 

AP-MS Datasets

Each RO-Crate contains three files:

  • readme.md
    • Documentation for the dataset
  • apms.tsv
    • Processed AP-MS data as a table of bait-prey interactions
  • ro-crate-metadata.json
    • Structured metadata for the dataset

Columns of the apms.tsv interaction data

  • Bait: Name of the pull downed protein
  • Prey: Uniprot ID number of identified proteins by MS in pull down (putative bait interactor).
  • PreyGene.x: Uniprot protein name of identified protein by MS in pull down (putative bait interactor).
  • Spec: Number of spectral count in each test biological replicates (separated by | ).
  • SpecSum: Sum of Spectral counts in test samples.
  • AvgSpec: Average Spectral counts across replicates in test samples.
  • NumReplicates.x: Number of replicates in test samples.
  • ctrlCounts: Number of spectral count in each control replicates (separated by | ).
  • AvgP.x: Average probability that an interaction is true, measure of the likelihood that a given interaction is a true positive rather than a random or non-specific interaction. A lower AvgP indicates higher confidence in the interaction being genuine.
  • MaxP.x: maximum probability associated with a protein interaction in the context of its prey-bait pair. Similar to AvgP, a lower MaxP suggests a higher likelihood of the interaction being true.
  • TopoAvgP.x: extension of the AvgP score that also takes into consideration the topology of the interaction network. It incorporates information about the hierarchical structure of the interaction data to provide a refined assessment of the interactions.
  • TopoMaxP.x: topology-aware score that considers the maximum probability of an interaction in the context of the interaction network’s topology.
  • SaintScore.x: composite score that integrates multiple aspects of the interaction data, including spectral counts and probability estimates. It’s designed to prioritize interactions based on their strength and reliability. Higher SaintScores indicate interactions that are more likely to be true.
  • logOddsScore: Logarithm of the odds ratio between test and control conditions for each prey as a measure of interaction significance. The LogOddsScore is a statistical score that represents the logarithm of the odds ratio for a protein-protein interaction. It’s used to quantify the strength and significance of the association between two proteins in an interaction network. The odds ratio compares the likelihood of the interaction occurring to the likelihood of it not occurring. Taking the logarithm of the odds ratio often helps to transform the score into a more symmetric and interpretable form, making it easier to compare and analyze the interactions. Higher LogOddsScores typically indicate stronger evidence for the interaction.
  • FoldChange.x: represents the ratio of the abundance of a protein or interaction in one experimental condition (Test) compared to another (control). It helps assess whether the abundance of a protein changes significantly between different conditions.
  • BFDR.x: Bayesian False Discovery Rate

 

CRISPR perturbation scRNA-Seq Datasets

Each RO-Crate contains:

h5ad file
———————————————————–
This file is a .h5ad file. Details about the data structure of this format
can be found here: https://anndata.readthedocs.io/en/latest/. This data is
a cell x gene matrix, with all pre-procesing steps done in accordance to
single cell best practices guide
https://www.sc-best-practices.org/conditions/perturbation_modeling.html#analysing-single-pooled-crispr-screens)
up to section 19.4.5. The .X layer is set to the ‘X_pert’ output of the
mixscape pipeline.

Cell map

The cell map generated by the cell maps toolkit is a hierarchy of nodes and edges. Nodes represent cellular subsystems that contain proteins; edges represent containment of lower system (bottom) by upper (top). 

Columns of the hierarchy nodes attributes
  • Name: Name of the subsystem
  • CD_MemberList: List of gene names in the subsystem
  • CD_MemberList_Size: Number of genes in the subsystem
  • CD_MemberList_LogSize: Log of the number of genes in the subsystem
  • CD_Labeled: Boolean denoting if the CD_CommunityName is set to a value 
  • CD_CommunityName: Name of community set by invocation of Run Functional Enrichment in Community Detection APplication and Service (CDAPS)
  • CD_AnnotatedMembers: String of space delimited node names used to set value in CD_CommunityName
  • CD_AnnotatedMembers_Overlap: CD_AnnotatedMembers_Size divided by CD_MemberList_Size
  • CD_AnnotatedMembers_Pvalue: Pvalue obtained from term mapping algorithm invoked by Run Functional Enrichment in CDAPS
  • HiDeF_persistence: community persistence value for pan-resolution community detection

Cell map RO-Crate content

Each RO-Crate contains three files:

  • readme.md
    • Documentation for the cell map.
  • A .cx file representing the cell map as a network 
  • ro-crate-metadata.json
    • Structured metadata for the cell map.

Cell maps on NDEx

CM4AI cell maps are also stored on NDEx, the Network Data Exchange, in network format. On NDEx, cell maps can be visualized and inspected. Nodes representing systems within the map can link to the corresponding interaction network linking the proteins in the system. The cell maps can also be opened in the Cytoscape desktop application and access programmatically from Python and R.

The goal of the CM4AI Standards group is to enable packaging and presentation of all original and derived datasets from CM4AI, up to and including the final Cell Map results, in FAIR (Findable-Accessible-Interoperable-reusable) format, with complete provenance graphs showing their derivation and providing a foundation for pre-model AI Explainability (XAI) as well as subsequent post-model XAI analytics. The architectural requirements for FAIR data are described in (Wilkinson MD, et al. 2016; https://doi.org/10.1038/sdata.2016.18). 
All datasets are packaged in standard RO-Crate format data+metadata packages (https://www.researchobject.org/ro-crate/ ), including: the software used to compute them if they are results data; a complete provenance graph of the dataset derivations specified in the Evidence Graph Ontology (EVI – https://w3id.org/EVI), an expansion of the W3C PROV ontology (http://www.w3.org/ns/prov-o) for biomedical research; and dataset- or software release-level metadata serialized in JSON-LD using the schema.org and EVI vocabularies. The metadata for each dataset includes a URI linked to one or more datasets or software packages which the metadata describes.
These packages are created using the FAIRSCAPE-CLI client-side toolkit. FAIRSCAPE-CLI is called in the Tools pipeline at each significant data derivation or computation, starting from the loading of data from each of the Data Acquisition modules and ending with the generation of the Cell Maps. 
FAIRSCAPE-CLI-generated RO-Crate packages may be loaded into the FAIRSCAPE digital commons environment at the University of Virginia, and/or repackaged in Bagit (IETF RFC 8493) wrappers for upload to instances of Dataverse, where Bagit package upload is configured for the installation, and/or made available directly. Packages loaded to FAIRSCAPE digital commons will have fully instantiated ark ids.

The documentation for the cellmaps pipeline can be found at https://cellmaps-pipeline.readthedocs.io/en/latest/. The pipeline invokes six tools in the cell maps toolkit that each create an output directory with results and the RO-Crates.  The toolkit steps are currently 1) image downloader to download the image data, 2) AP-MS downloader to download the image data, 3) generate image embedding tool to create image embeddings using a densenet model 4) generate AP-MS embedding tool to create AP-MS embeddings using node2vec, 5) co-embedding tool to create an integrated embedding from the image and AP-MS embeddings and 6) generate hierarchy tool to create a cell map from the integrated embeddings. 

Each step in the cell maps toolkit including the pipeline are distributed as a PyPI package. The source code is available on GitHub and documentation is available with usage examples and steps for installation and tool development.