Overview
In the CM4AI project, 200 genes/proteins are the subject of coordinated experiments in three data modalities. The gene/proteins comprise 100 chromatin modifiers and 100 metabolic enzymes. Data has been collected on the chromatin modifiers in Year 1 of the project. We selected coherent groups of genes that are known to interact in processes involved in cancer, neuropsychiatric, and cardiac disorders. Data is being acquired in the triple-negative breast cancer cell line MDA-MB-468, including upon treatment with paclitaxel or vorinostat, and in two iPSC lines in the undifferentiated state as well as in differentiated neurons and cardiomyocytes using complementary mapping approaches. We are using proteomic mass spectrometry to map protein-protein interactions, cellular imaging to map the spatial subcellular organization of proteins, and genetic perturbation screens via CRISPR/Cas9 to assess the transcriptome-wide impact of gene perturbations.
Affinity Purification Mass spectrometry (AP-MS) Data
Affinity Purification-Mass Spectrometry (AP-MS) is a powerful technique used to uncover the intricate networks of protein interactions within cells. Different methods can be used but the most popular approach is to attach a purifiable tag to a protein of interest, the “bait”. In the context of CM4AI the tag is directly attached to the endogenous protein of interest using gene editing to maintain physiological level of expression. The tagged bait is then purified from cell extract with its interacting partners. Once separated, the captured proteins are identified using a mass spectrometer allowing to infer protein-protein interactions. Understanding these interactions allows can unravel complex cellular processes, discover new roles for proteins, and gain insights into diseases caused by disrupted protein networks.
AP-MS Data acquisition and processing
- The raw data are acquired on an Orbitrap Fusion™ Lumos™ Tribrid™ Mass Spectrometer (Thermo Scientific)
- The raw data are processed using Maxquant v1.6.12.0 using default parameter with LFQ quantification and the “202209_uniprot_human_reviewed.fasta” file as reference.
- The evidence file from maxquant is processed using ArtMS in R to generate SAINT input files using the following command lines.
- SAINTexpress was run locally using version SAINTexpress_v3.6.3
SAINT is designed to assess the reliability and significance of protein-protein interactions identified through mass spectrometry experiments. It takes into account both the presence and abundance of proteins in different experimental conditions or replicates. The method uses a probabilistic model to estimate the probability that a given interaction is a true positive rather than a random or non-specific interaction. This approach helps to distinguish true interactions from noise in large-scale protein interaction datasets.
https://pubmed.ncbi.nlm.nih.gov/21131968
https://saint-apms.sourceforge.net/Main.html
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4102138
Immunofluorescence Imaging Data
The cell imaging takes place in the Emma Lundberg lab (https://ell-core.stanford.edu/). We will provide insight into the expression and spatiotemporal distribution of up to 700 proteins, stained in the breast cancer cell line MDA-MB-468 or the iPSC line KOLF2.1J, either under untreated conditions or treated with the chemotherapeutic drug Paclitaxel or Vorinostat (SAHA). For each protein, the subcellular distribution of the protein is investigated by immunofluorescence staining and confocal microscopy. After image acquisition, we classify the subcellular localization of the protein in one or more of 35 different organelles and fine subcellular structures. The generated data set will allow us to see how treatment with Paclitaxel or Vorinostat influences the subcellular distribution of proteins. Finally, this data set will be contributed to build MuSIC maps, as demonstrated in this publication (https://www.nature.com/articles/s41586-021-04115-9).
The immunofluorescence method in detail
In the immunofluorescence staining procedure (ICC-IF), each protein of interest is visualized by using an antibody that detects the protein. The antibodies are derived from The Human Protein Atlas (https://proteinatlas.org) project. The antibody-based staining of the protein is shown in green in the generated images. During the staining procedure, the sample is also co-stained for additional structural markers of the cell: nuclei (stained with DAPI, shown in blue in the generated images), Microtubuli (MT) (stained with a Tubulin antibody, shown in red in the generated images), and endoplasmic reticulum (ER) (stained with a Calreticulin antibody, shown in yellow in the generated images).
Current stage (Version 0.1alpha)
We have mapped the protein localization for the first 95 genes in the cell line MDA-MB-468, either under untreated conditions or treated with Paclitaxel or treated with Vorinostat. We noticed that the drug treatment largely influenced the cellular morphology (see example images below). After paclitaxel treatment, nuclear morphology drastically changed. The Paclitaxel and the Vorinostat treatment each reduced the number of cells present at the end of the experiment. Thus, we have seeded more cells in those wells that were treated with Paclitaxel or Vorinostat compared to the wells that were left untreated. Nonetheless, fewer cells than in the untreated condition were observed at the end of the experiment.

For some proteins we observed similar subcellular localization in all conditions:

Other proteins appeared with heterogeneous subcellular protein localization patterns and showed a shift in the percentage of cells showing a certain type of subcellular protein distribution:

In the next steps, we will analyze what the specific differences are in more detail and apply this workflow for staining more proteins to get an even more comprehensive picture of the subcellular protein landscape in breast cancer cells treated with Paclitaxel or Vorinostat, compared to untreated cells.
CRISPR Screening Data

- Single-cell capture with Chromium Next GEM Chip M Single Cell Kit (10x Genomics)
- FASTQs sequencing data was acquired on a NovaSeq S4 (Illumina)
- Alignment of the FASTQ files on Human Genome GRCh38 and the CRISPR library was done with Cell Ranger 7.1.0 Software (10x Genomics)
| Gene | NCBI Gene ID | Molecular function | Associated Disease |
| PARP1 | 142 | ADPRibosylation | Intellectual disability / cancer |
| BAP1 | 8314 | Ubiquitin regulation | Cancer |
| USP7 | 7874 | Ubiquitin regulation | Intellectual disability / cancer |
| NEDD4 | 4734 | Ubiquitin regulation | Brugada syndrome |
| BRCA1 | 672 | Ubiquitin regulation | Intellectual disability / Heart disease / Cancer |
| BARD1 | 580 | Ubiquitin regulation | Intellectual disability / Heart disease / Cancer |
| BRE1A | 56254 | Ubiquitin regulation | Autism / heart disease |
| CBL | 867 | Ubiquitin regulation | Cancer |
| MSL1 | 339287 | Ubiquitin regulation | Intellectual disability / Heart disease / Cancer |
| PHF6 | 84295 | Ubiquitin regulation | Intellectual disability |
| TAF1 | 6872 | Ubiquitin regulation | Congenital heart disease / Mental retardation |
| UHRF1 | 29128 | Ubiquitin regulation | Autism |
| UHRF2 | 115426 | Ubiquitin regulation | Congenital heart disease |
| KAT2A | 2648 | Histone Acetyl Transferase | Autism |
| KAT3A | 1387 | Histone Acetyl Transferase | Intellectual disability / Heart disease / Cancer |
| KAT3B | 2033 | Histone Acetyl Transferase | Intellectual disability / Heart disease / Cancer |
| KAT6A | 7994 | Histone Acetyl Transferase | Intellectual disability / Heart disease / Cancer |
| KAT6B | 23522 | Histone Acetyl Transferase | Intellectual disability / Heart disease / Cancer |
| HDAC1 | 3065 | Deacetylase | Cancer / Cardiac defect |
| HDAC2 | 3066 | Deacetylase | Cancer / Cardiac defect |
| HDAC3 | 8841 | Deacetylase | Cancer |
| HDAC4 | 9759 | Deacetylase | Cancer |
| HDAC5 | 10014 | Deacetylase | Cancer / Cardiac defect |
| HDAC6 | 10013 | Deacetylase | Cancer |
| HDAC7 | 51564 | Deacetylase | Cancer / Vascular integrity |
| HDAC8 | 55869 | Deacetylase | Cancer |
| HDAC9 | 9734 | Deacetylase | Cancer / Cardiac defect |
| HDAC10 | 83933 | Deacetylase | Cancer |
| HDAC11 | 79885 | Deacetylase | cancer |
| SIRT2 | 22933 | Deacetylase | Cancer / Acute myocardial infarction |
| KDM3B | 51780 | Lysine Demetylase | Schizophrenia |
| KDM4C | 23081 | Lysine Demetylase | Upper aerodigestive tract cancer, association with |
| KDM5A | 5927 | Lysine Demetylase | Intellectual disability / Heart disease / breast cancer |
| KDM5B | 10765 | Lysine Demetylase | Intellectual disability / Heart disease |
| KDM5C | 8242 | Lysine Demetylase | Intellectual disability |
| KDM6A | 7403 | Lysine Demetylase | Kabuki syndrome / Autism / Cancer |
| KDM6B | 23135 | Lysine Demetylase | Intellectual disability |
| TET1 | 80312 | Lysine Demetylase | Autism? |
| TET2 | 54790 | Lysine Demetylase | Myeloproliferative neoplasms / prostate cancer |
| TET3 | 200424 | Lysine Demetylase | Intellectual disability |
| ATM | 472 | Kinase | Ataxia telandiectasia / Cancer |
| ATR | 545 | Kinase | Intellectual disability / Cancer |
| AURKB | 9212 | Kinase | Cancer |
| AURKA | 6790 | Kinase | Cancer |
| CSNK2A1 | 1457 | Kinase | neurological diseases / inflamation / Cancer |
| HASPIN | 83903 | Kinase | Cancer |
| JAK2 | 3717 | Kinase | Myeloproliferative neoplasms / Erythrocytosis |
| MSK1 | 855651 | Kinase | Cancer |
| MSK2 | 8986 | Kinase | Cancer |
| 5292 | 5292 | Kinase | Autism |
| PRKCA | 5578 | Kinase | Cardiomyopathy, Autism / Cancer |
| PRKCB | 5579 | Kinase | Autism |
| PRKCD | 5580 | Kinase | Immunodeficiency |
| BAZ1B | 9031 | Kinase | Autism / heart disease |
| PRKAA1 | 5562 | AMPK | Cancer / Cardio vascular syndrome |
| PRKAA2 | 5563 | AMPK | Cancer |
| PRKAB1 | 5564 | AMPK | Cancer / Neurological disease |
| PRKAB2 | 5565 | AMPK | Cancer |
| PRKAG1 | 5571 | AMPK | Cancer / Cardio vascular syndrome |
| PRKAG2 | 51422 | AMPK | Cancer / Cardio vascular syndrome |
| PRKAG3 | 53632 | AMPK | Cancer / Cardio vascular syndrome |
| DNMT1 | 1786 | Lysine Methyltransferase | Neuropathy / Cerebellar ataxia / cancer |
| DNMT3A | 1788 | Lysine Methyltransferase | Intellectual disability / cancer |
| DNMT3B | 1789 | Lysine Methyltransferase | Kabuki syndrome / Cancer |
| KMT1D | 79813 | Lysine Methyltransferase | Intellectual disability / Heart disease / cancer |
| KMT1F | 83852 | Lysine Methyltransferase | Autism / Cancer |
| KMT2A | 4297 | Lysine Methyltransferase | Intellectual disability / Heart disease / cancer |
| KMT2C | 58508 | Lysine Methyltransferase | Intellectual disability / Heart disease / cancer |
| KMT2D | 8085 | Lysine Methyltransferase | Kabuki syndrome |
| KMT2E | 55904 | Lysine Methyltransferase | Intellectual disability / Heart disease |
| KMT2H | 55870 | Lysine Methyltransferase | Intellectual disability / Heart disease |
| KMT3B | 64324 | Lysine Methyltransferase | Intellectual disability / Heart disease / cancer |
| PRMT1 | 3276 | Lysine Methyltransferase | Cancer |
| EYA1 | 2138 | Phosphatase | Cancer / BOR syndrome |
| EYA3 | 2140 | Phosphatase | Cancer |
| PPP1CA | 5499 | Phosphatase | Encephalitis / Cancer |
| PPP1CB | 5500 | Phosphatase | Cancer |
| PPP1CC | 5501 | Phosphatase | Cancer |
| PPP4C | 5531 | Phosphatase | Autism / Cancer |
| BPTF | 2186 | Reader | Neurodevelopmental disorder / Cancer |
| BRD4 | 23476 | Reader | Heart disease / neuro-dev delay / cancer |
| BRD7 | 29117 | Reader | Cancer |
| BRPF1 | 7862 | Reader | Intellectual disability / Cardiac disorder / Cancer |
| CHD4 | 1108 | Reader | Intellectual disability |
| ING1 | 3621 | Reader | Cancer |
| TAF1 | 6872 | Reader | Intellectual development disorder / Cancer |
| TRIM24 | 8805 | Reader | Cancer |
| YWHAZ | 7534 | Reader | Intellectual disability / Cardiac disorder / Cancer |
| YWHAE | 7531 | Reader | Cancer |
| YWHAQ | 10971 | Reader | Creutzfeldt-Jakob Disease / Cancer |
| YWHAB | 7529 | Reader | Cancer |
| YWHAH | 7533 | Reader | Schizophrenia |
| YWHAG | 7532 | Reader | Intellectual disability |
| BRG1 | 6597 | Reader | Nervous an cardiovascular diseases / Cancer |
| CBX2 | 84733 | Reader | Cancer |
| CBX3 | 11335 | Reader | Cancer |
| MRG15 | 10933 | Reader | Cancer |
| P/CAF | 8850 | Reader | Cancer |
| 53BP1 | 7158 | Reader | Microcephaly / Cancer |
| BRD3 | 8019 | Reader | Intellectual disability / Cancer |
| CHD1 | 1105 | Reader | Neurodevelopmental disorder / Cancer |
| MBTD1 | 54799 | Reader | Cancer |
| PDP1 | 54704 | Reader | Cancer |
| RAG2 | 5897 | Reader | Cancer / Immunodeficiency |
| Gene | NCBI Gene ID | Molecular function |
| FOLR1 | 2348 | Folate/Methionine Cycle |
| MTHFR | 4524 | Folate/Methionine Cycle |
| MTR | 4548 | Folate/Methionine Cycle |
| MTHFD1 | 4522 | Folate/Methionine Cycle |
| MAT1A | 4143 | Folate/Methionine Cycle |
| MAT2A | 4144 | Folate/Methionine Cycle |
| MAT2B | 27430 | Folate/Methionine Cycle |
| AHCY | 191 | Folate/Methionine Cycle |
| SHMT1 | 6470 | Folate/Methionine Cycle |
| RFK | 55312 | FAD |
| FLAD1 | 80308 | FAD |
| TKFC | 26007 | FAD |
| ACAD9 | 28976 | Acethyl-CoA |
| ACADM | 34 | Acethyl-CoA |
| ACADS | 35 | Acethyl-CoA |
| ACADSB | 36 | Acethyl-CoA |
| ACADVL | 37 | Acethyl-CoA |
| ACAT1 | 38 | Acethyl-CoA |
| ACOT7 | 11332 | Acethyl-CoA |
| ACOX3 | 8310 | Acethyl-CoA |
| ACSS2 | 55902 | Acethyl-CoA |
| ACSS3 | 79611 | Acethyl-CoA |
| GMPS | 8833 | Purine metabolism |
| LDHA | 3939 | Lactate metabolism |
| LDHB | 3945 | Lactate metabolism |
| LDHD | 197257 | Lactate metabolism |
| ACLY | 47 | TCA cycle |
| ALDH5A1 | 7915 | TCA cycle |
| DLAT | 1737 | TCA cycle |
| DLD | 1738 | TCA cycle |
| GLUD1 | 2746 | TCA cycle |
| GLUD2 | 2747 | TCA cycle |
| IDH1 | 3417 | TCA cycle |
| IDH3A | 3419 | TCA cycle |
| IDH3B | 3420 | TCA cycle |
| IDH3G | 3421 | TCA cycle |
| MDH2 | 4191 | TCA cycle |
| OGDH | 4967 | TCA cycle |
| PDHA1 | 5160 | TCA cycle |
| PDHA2 | 5161 | TCA cycle |
| PDHB | 5162 | TCA cycle |
| SDHA | 6389 | TCA cycle |
| G6PD | 2539 | NADH |
| GAPDH | 2597 | NADH |
| NADK | 65220 | NADH |
| NDUFS4 | 4724 | NADH |
| NDUFAB1 | 4706 | NADH |
| NMNAT2 | 23057 | NADH |
| NNT | 23530 | NADH |
| PGD | 5226 | NADH |
| PGM1 | 5236 | NADH |
| PHGDH | 26227 | NADH |
| PYGL | 5836 | NADH |
| SRC | 6714 | Kinases |
| KIT | 3815 | Kinases |
| ITPKA | 3706 | Kinases |
| AKT1 | 207 | Kinases |
| BRAF | 673 | Kinases |
| CDK4 | 1019 | Kinases |
| DAPK3 | 1613 | Kinases |
| FYN | 2534 | Kinases |
| PIK3CA | 5290 | Kinases |
| PIK3C2A | 5286 | Kinases |
| NUAK1 | 9891 | Kinases |
| STK4 | 6789 | Kinases |
| MAPK1 | 5594 | Kinases |
| MTOR | 2475 | Kinases |
| PIP5K1C | 23396 | Kinases |
| RIPK1 | 8737 | Kinases |
| RPS6KA3 | 6197 | Kinases |
| KRAS | 3845 | Kinases |
| TGFBR2 | 7048 | Kinases |
| WEE1 | 7465 | Kinases |
| PPA2 | 27068 | other |
| AHCYL2 | 23382 | other |
| CYP27A1 | 1593 | other |
| DIO3 | 1735 | other |
| GSTM1 | 2944 | other |
| HEXA | 3073 | other |
| NAGS | 162417 | other |
| NCOA1 | 8648 | other |
| PANK2 | 80025 | other |
| PARP4 | 143 | other |
| PARP6 | 56965 | other |
| PLCD4 | 84812 | other |
| PLCG1 | 5335 | other |
| SLC2A1 | 6513 | other |
| UGDH | 7358 | other |
| UGT8 | 7368 | other |
| UBE3C | 9690 | other |
| CYP24A1 | 1591 | other |
| GCLM | 2730 | other |
| GYG2 | 8908 | other |
| PDE4A | 5141 | other |
| PLPP7 | 84814 | other |
| AMDHD2 | 51005 | other |
| ABHD6 | 57406 | other |
Each dataset is stored as a RO-Crate package. Datasets contain the data for a gene set measured in a specific cell line under a specific treatment. (Including “untreated”)
The RO-Crate format
Research Object Crate (RO-Crate) is a lightweight standard for packaging research data with their metadata. It contains schema.org and other standard structured vocabulary annotations in JSON-LD, of all included or referenced data objects, and aims to make best-practice in formal metadata description accessible and practical for use in a wide variety of situations. Each RO-Crate is a folder that contains, along with one or more data files, a ro-crate-metadata.json file, containing structured information about the data (metadata) in JSON-LD format. It is good practice to also include a readme.md text file describing the contents.
Image Datasets
Each RO-Crate contains:
- A cell line – antibody – treatment table describing the experiment
- Four channel data folders: red, green, blue yellow
- Each folder contains images in .jpeg format.
- ro-crate-metadata.json
- Structured metadata for the dataset
- README
Example README file:
This data set has been generated in the Emma Lundberg lab (https://ell-core.stanford.edu/) in July-August 2023 as part of the CM4AI Research Program (https://digitaltumors.org/).
This data set displays the spatial localization of 95 proteins in the breast cancer cell line MDA-MB-468. MDA-MB-468 cells were seeded into 96-well glass bottom plates and treated either with the chemotherapeutic drug Paclitaxel or with the chemotherapeutic drug Vorinostat (SAHA) or not treated. For drug treatment, higher cell numbers were seeded than for the cells that were not treated to reach a sufficient amount of remaining cells at the end of the experiment (treatment with Paclitaxel or Vorinostat disturbs cell growth). After fixation of the cells with PFA, for each protein, the subcellular distribution of the protein was investigated by immunofluorescence-based staining (ICC-IF) and confocal microscopy. Each well was stained with DAPI (to label nuclei, “blue” channel), a Calreticulin antibody (to label the endoplasmic reticulum, “yellow” channel), a Tubulin antibody (to label Microtubuli, “red” channel), and an antibody against a protein of interest (a different antibody / protein for each well, “green” channel). In the green channel, for each well, one of 95 proteins was stained with an antibody from the proteinatlas.org project.
The .tsv file describes each image in the data set. Each row represents one image. The columns describe the staining from which the image was taken:
“Antibody ID” describes the antibody ID for the antibody applied to stain the protein visible in the “green” channel. The antibody ID can be looked up at proteinatlas.org to find out more information about the antibody.
“ENSEMBL ID” indicates the ENSEMBL ID(s) of the gene(s) of the proteins visualized in the “green” channel.
Treatment refers to how the cells that are depicted in the image were treated (with Paclitaxel, Vorinostat, or untreated)
“Well” refers to the well coordinate on the 96-well plate
“Region” is a unique identifier for the position in the well, where the cells were acquired.
AP-MS Datasets
Each RO-Crate contains three files:
- readme.md
- Documentation for the dataset
- apms.tsv
- Processed AP-MS data as a table of bait-prey interactions
- ro-crate-metadata.json
- Structured metadata for the dataset
Columns of the apms.tsv interaction data
- Bait: Name of the pull downed protein
- Prey: Uniprot ID number of identified proteins by MS in pull down (putative bait interactor).
- PreyGene.x: Uniprot protein name of identified protein by MS in pull down (putative bait interactor).
- Spec: Number of spectral count in each test biological replicates (separated by | ).
- SpecSum: Sum of Spectral counts in test samples.
- AvgSpec: Average Spectral counts across replicates in test samples.
- NumReplicates.x: Number of replicates in test samples.
- ctrlCounts: Number of spectral count in each control replicates (separated by | ).
- AvgP.x: Average probability that an interaction is true, measure of the likelihood that a given interaction is a true positive rather than a random or non-specific interaction. A lower AvgP indicates higher confidence in the interaction being genuine.
- MaxP.x: maximum probability associated with a protein interaction in the context of its prey-bait pair. Similar to AvgP, a lower MaxP suggests a higher likelihood of the interaction being true.
- TopoAvgP.x: extension of the AvgP score that also takes into consideration the topology of the interaction network. It incorporates information about the hierarchical structure of the interaction data to provide a refined assessment of the interactions.
- TopoMaxP.x: topology-aware score that considers the maximum probability of an interaction in the context of the interaction network’s topology.
- SaintScore.x: composite score that integrates multiple aspects of the interaction data, including spectral counts and probability estimates. It’s designed to prioritize interactions based on their strength and reliability. Higher SaintScores indicate interactions that are more likely to be true.
- logOddsScore: Logarithm of the odds ratio between test and control conditions for each prey as a measure of interaction significance. The LogOddsScore is a statistical score that represents the logarithm of the odds ratio for a protein-protein interaction. It’s used to quantify the strength and significance of the association between two proteins in an interaction network. The odds ratio compares the likelihood of the interaction occurring to the likelihood of it not occurring. Taking the logarithm of the odds ratio often helps to transform the score into a more symmetric and interpretable form, making it easier to compare and analyze the interactions. Higher LogOddsScores typically indicate stronger evidence for the interaction.
- FoldChange.x: represents the ratio of the abundance of a protein or interaction in one experimental condition (Test) compared to another (control). It helps assess whether the abundance of a protein changes significantly between different conditions.
- BFDR.x: Bayesian False Discovery Rate
CRISPR perturbation scRNA-Seq Datasets
Each RO-Crate contains:
| h5ad file ———————————————————– This file is a .h5ad file. Details about the data structure of this format can be found here: https://anndata.readthedocs.io/en/latest/. This data is a cell x gene matrix, with all pre-procesing steps done in accordance to single cell best practices guide https://www.sc-best-practices.org/conditions/perturbation_modeling.html#analysing-single-pooled-crispr-screens) up to section 19.4.5. The .X layer is set to the ‘X_pert’ output of the mixscape pipeline. |
Cell map
The cell map generated by the cell maps toolkit is a hierarchy of nodes and edges. Nodes represent cellular subsystems that contain proteins; edges represent containment of lower system (bottom) by upper (top).
Columns of the hierarchy nodes attributes
- Name: Name of the subsystem
- CD_MemberList: List of gene names in the subsystem
- CD_MemberList_Size: Number of genes in the subsystem
- CD_MemberList_LogSize: Log of the number of genes in the subsystem
- CD_Labeled: Boolean denoting if the CD_CommunityName is set to a value
- CD_CommunityName: Name of community set by invocation of Run Functional Enrichment in Community Detection APplication and Service (CDAPS)
- CD_AnnotatedMembers: String of space delimited node names used to set value in CD_CommunityName
- CD_AnnotatedMembers_Overlap: CD_AnnotatedMembers_Size divided by CD_MemberList_Size
- CD_AnnotatedMembers_Pvalue: Pvalue obtained from term mapping algorithm invoked by Run Functional Enrichment in CDAPS
- HiDeF_persistence: community persistence value for pan-resolution community detection
Cell map RO-Crate content
Each RO-Crate contains three files:
- readme.md
- Documentation for the cell map.
- A .cx file representing the cell map as a network
- ro-crate-metadata.json
- Structured metadata for the cell map.
Cell maps on NDEx

CM4AI cell maps are also stored on NDEx, the Network Data Exchange, in network format. On NDEx, cell maps can be visualized and inspected. Nodes representing systems within the map can link to the corresponding interaction network linking the proteins in the system. The cell maps can also be opened in the Cytoscape desktop application and access programmatically from Python and R.
The goal of the CM4AI Standards group is to enable packaging and presentation of all original and derived datasets from CM4AI, up to and including the final Cell Map results, in FAIR (Findable-Accessible-Interoperable-reusable) format, with complete provenance graphs showing their derivation and providing a foundation for pre-model AI Explainability (XAI) as well as subsequent post-model XAI analytics. The architectural requirements for FAIR data are described in (Wilkinson MD, et al. 2016; https://doi.org/10.1038/sdata.2016.18).
All datasets are packaged in standard RO-Crate format data+metadata packages (https://www.researchobject.org/ro-crate/ ), including: the software used to compute them if they are results data; a complete provenance graph of the dataset derivations specified in the Evidence Graph Ontology (EVI – https://w3id.org/EVI), an expansion of the W3C PROV ontology (http://www.w3.org/ns/prov-o) for biomedical research; and dataset- or software release-level metadata serialized in JSON-LD using the schema.org and EVI vocabularies. The metadata for each dataset includes a URI linked to one or more datasets or software packages which the metadata describes.
These packages are created using the FAIRSCAPE-CLI client-side toolkit. FAIRSCAPE-CLI is called in the Tools pipeline at each significant data derivation or computation, starting from the loading of data from each of the Data Acquisition modules and ending with the generation of the Cell Maps.
FAIRSCAPE-CLI-generated RO-Crate packages may be loaded into the FAIRSCAPE digital commons environment at the University of Virginia, and/or repackaged in Bagit (IETF RFC 8493) wrappers for upload to instances of Dataverse, where Bagit package upload is configured for the installation, and/or made available directly. Packages loaded to FAIRSCAPE digital commons will have fully instantiated ark ids.
The documentation for the cellmaps pipeline can be found at https://cellmaps-pipeline.readthedocs.io/en/latest/. The pipeline invokes six tools in the cell maps toolkit that each create an output directory with results and the RO-Crates. The toolkit steps are currently 1) image downloader to download the image data, 2) AP-MS downloader to download the image data, 3) generate image embedding tool to create image embeddings using a densenet model 4) generate AP-MS embedding tool to create AP-MS embeddings using node2vec, 5) co-embedding tool to create an integrated embedding from the image and AP-MS embeddings and 6) generate hierarchy tool to create a cell map from the integrated embeddings.
Each step in the cell maps toolkit including the pipeline are distributed as a PyPI package. The source code is available on GitHub and documentation is available with usage examples and steps for installation and tool development.

