# Overview

Partek software enables researchers to easily perform genomic data analysis without ever needing to write a single line of code or sacrificing statistical power or advanced functionality. From alignment to pathway analysis, Partek provides a seamless, integrated analysis solution on a single platform that provides the power of a cloud or cluster when needed, and the convenience of desktop software for less compute intensive tasks.

Here you will find documentation on how to use and administer our products.

|                                                                                                                                                   |                                                                                                                                                                       |                                                                                                                                                         |
| :-----------------------------------------------------------------------------------------------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------: |
| <p><img src="/files/EmR0sY9xJAeUBZmIw6Jf" alt="Partek Flow icon" data-size="line"><br><a href="/pages/KC3PNjdUnua3LvQF3vjt">Partek™ Flow™</a></p> | <p><img src="/files/LPf4fjq4TKMP3Xy12KgW" alt="Partek Genomics Suite icon" data-size="line"><br><a href="/pages/8uLbaChJKoKITfIu2Fa0">Partek™ Genomics Suite™</a></p> | <p><img src="/files/NXhkGipiy8HciAMpkijZ" alt="Partek Pathway icon" data-size="line"><br><a href="/pages/IY30IVA6PtL1W0If05uQ">Partek™ Pathway™</a></p> |


# Partek Flow

|                                                                                                                                                                          |                                                                                                                                                                                               |                                                                                                                                                       |
| :----------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------: |
|     <p><img src="/files/Z3zZfC9Vg0sxodNGyAl1" alt="Installation guide icon" data-size="original"><br><a href="/pages/Qb2SgqUzCOlEg79AjgKQ">Installation Guide</a></p>    |                  <p><img src="/files/4ReIwAmfdc66EUAZXLmr" alt="Getting started icon" data-size="original"><br><a href="/pages/Iv3tNmz68smCdx85jmKQ">Getting Started</a></p>                  |    <p><img src="/files/AnVftJyiN6lRLTlNvByo" alt="Task menu icon" data-size="original"><br><a href="/pages/MZlRFANs4pwtQlzEDmUQ">Task Menu</a></p>    |
| <p><img src="/files/ADcOcq4Q0RXcqducyro1" alt="Tutorial RNA-Seq data icon" data-size="original"><br><a href="/pages/G1JxRUlx1e871tByIkrK">Tutorial: RNA-Seq Data</a></p> | <p><img src="/files/8aT3Hj9pG7y2mxSSxYm0" alt="Creating and analyzing a project icon" data-size="original"><br><a href="/pages/jOZlZgOjejwD3lOWSsF7">Creating and Analyzing a Project</a></p> |     <p><img src="/files/g7hxZEfeIphJWZUGtkdj" alt="Webinars icon" data-size="original"><br><a href="/pages/tfFv8sbS4eah5xLQGapa">Webinars</a></p>     |
|          <p><img src="/files/xMOeGYXsHvoIctEbzvsw" alt="Release notes icon" data-size="original"><br><a href="/pages/RsNN6nKglqAggcPVsjD3">Release Notes</a></p>         |          <p><img src="/files/S8wpqyf2SNKYY2FzTh8r" alt="Library file management icon" data-size="original"><br><a href="/pages/x0c3jebudasvcga55yzr">Library File Management</a></p>          | <p><img src="/files/0lMwnC3W1p3OAfJP8Nty" alt="White papers icon" data-size="original"><br><a href="/pages/dcdfB1bvW2Ht0ZYYDHQd">White Papers</a></p> |


# Frequently Asked Questions

## General

### How to create a project?

To create a project, you first need to [transfer files to the Partek Flow server](/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server), and then import the files into your project using the import data wizard, here is the [video ](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial)and [more information](/partek-flow/user-manual/importing-data).

### Can I change my user avatar?

Yes, navigate to [My profile](/partek-flow/user-manual/settings/personal/my-profile) and click the "Change image" button. Do this by clicking your avatar at the top right corner of the interface, select [Settings](/partek-flow/user-manual/settings), then choose [Profile](/partek-flow/user-manual/settings/personal/my-profile).

### How do I add and use my own lists?

Click your avatar in the top right corner of the Partek Flow interface, choose [Settings ](/partek-flow/user-manual/settings)in the menu, and select [Lists ](/partek-flow/user-manual/settings/components/lists)from the left panel of the [Components](/partek-flow/user-manual/settings/components) section. Lists can also be generated from result tables using the "Save as managed list" button. For more information please click [here](/partek-flow/user-manual/settings/components/lists).

### Can I repeat a task and everything downstream of it, while changing only one/a few parameters?

Yes, click on the rectangular task that you want to change the parameters. On the context-specific menu on the right, under Task actions, select ‘Rerun with downstream tasks’, this will bring you to the task set up page where you can edit the parameters for the task, then click Finish to run the task with the new parameters. The tasks downstream of it will be initiated automatically.

### What can I use to identify cells that are actively expressing genes within a gene list?

Use [AUCell ](/partek-flow/user-manual/task-menu/exploratory-analysis/aucell)to identify cells with active gene sets; this task calculates a value for each cell by ranking all genes by their expression level in the cell and identifying what proportion of the genes from the gene list fall within the top 5% (default cutoff) of genes. An alternative option is to use the Gene score for a feature list to select and filter populations based on the distribution; [click here for more information](/partek-flow/user-manual/settings/components/lists#use-a-feature-list-to-select-and-filter-populations).

### Can I build and use pipelines for my analysis?

Yes, click on [Import a pipeline](/partek-flow/user-manual/pipelines/importing-a-pipeline) on the bottom of the [Analyses ](/partek-flow/tutorials/creating-and-analyzing-a-project/the-analyses-tab)tab dashboard. This will help you import either our hosted pipelines or your own saved pipeline which can be found under [Settings ](/partek-flow/user-manual/settings)-> [Components ](/partek-flow/user-manual/settings/components)-> [Pipelines](/partek-flow/user-manual/settings/components/pipeline-management). Click [here ](/partek-flow/tutorials/bulk-rna-seq/saving-and-running-a-pipeline)for steps to save and run a pipeline. For more information related to navigating pipelines click [here](/partek-flow/user-manual/pipelines).

### How do I classify cells?

Classification in Partek Flow can be performed manually or with automatic cell classification which is explained in more detail [here](/partek-flow/user-manual/task-menu/classification). Users often want to classify cells by gene expression threshold(s), for details on classification by marker expression click [here](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/perform-exploratory-analysis#classify-cells-based-on-a-marker-for-expression). Automatic classification needs to be performed on a non-normalized single cell data node; once complete, [publish cell attributes to project](/partek-flow/user-manual/task-menu/annotation-metadata/publish-cell-attributes-to-project) then use this classification in visualizations and tasks. You may choose to perform [Graph-based clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/graph-based-clustering) and [K-means clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/k-means-clustering) to help identify biomarkers that can then be used to identify the clusters and we also provide hosted lists for different cell types.

### My server is full, how do I make more space?

We recommend cleaning up projects as well as removing library files that you do not need, then removing the orphaned files. You can also export analyzed projects and save them on an external machine, then when you need them again you can import them to the server. Please see this information for more details related to: [Project management](/partek-flow/tutorials/creating-and-analyzing-a-project/project-management), [Removing library files](/partek-flow/user-manual/settings/components/library-file-management/removing-library-files), and [Orphaned files](/partek-flow/user-manual/settings/access/orphaned-files). Right click on the data node to delete files from projects that are not needed (e.g. fastqs from project pipelines that are analyzed); you will not be able to perform tasks from this node once the files are deleted.

### How do I add library files if I am not studying human or mouse?

To add a new assembly, click on [Settings ](/partek-flow/user-manual/settings)-> [Library files](/partek-flow/user-manual/settings/components/library-file-management). From the Assembly drop-down list, select Add assembly and specify the species. If the species name is not in the list, choose Other and type in the name with the assembly version (multiple assembly versions can exist for one species, e.g. hg19 and hg38 for Homo Sapiens). You need to add the reference file which is a .fasta file containing sequence information. Once the reference file is added, you can build any aligner index to perform the alignment task.

The Annotation model is a file containing feature location. This file can be used to quantify to annotation model in RNA-Seq analysis, or annotate variant or peaks in a DNA-Seq or ATAC-Seq/ChIP-Seq data analysis pipeline. The file format should be .gtf/.gff/.bed.

We recommend looking for the species files on the [Ensembl ](https://useast.ensembl.org/index.html)website. There is no need to unzip or save these files to your local machine, instead right click and copy the link address of the specific file (not a link to a folder). For more details, here is the documentation chapter: [Library File Management](/partek-flow/user-manual/settings/components/library-file-management).

### Are Genome coordinates 1-based or 0-based?

Genome coordinates for annotation models stored in Partek Flow are 1-based, start-inclusive, and stop-exclusive. This means that the first base position starts from one, the start coordinate for a feature is included in the feature and the stop/end coordinate is not included in the feature. These are the genome coordinates that are printed in various task reports and output files when an annotation model is involved in the task. When custom annotation files are added to Partek Flow, the genome coordinates are converted into this format. The coordinates are converted back if necessary for a specific task. shows how the genome coordinates vary between different annotation formats.

![](/files/ZhXWWRvglqMMGvZWyJtN)

### Can I add transgenes to my reference files?

Yes, to add transgenes (including gfp or related) to the references files, first choose an assembly, create the transgene reference, and merge the references together (e.g. combine mm10 with dttomato). This is the same process for the annotation file.

### How do I export data from the result nodes?

Left click to select the data node you want to export. In the bottom of the task menu there will be an option to [Download data](/partek-flow/quick-start-guide#downloading-your-data).

### Why can’t I find the RPKM method on the normalization menu (I see FPKM)? <a href="#why-cant-i-find-the-rpkm-method-on-the-normalization-menu-i-see-fpkm" id="why-cant-i-find-the-rpkm-method-on-the-normalization-menu-i-see-fpkm"></a>

When working with paired data it should be the case that FPKM is available, and when working with single end data RPKM should be available. These metrics are essentially analogous, but based on the underlying method used for calculation (accounting for two reads mapping to 1 fragment and not counting twice for paired end data). Here is a simple description of the differences in calculation between RPKM and FPKM: <http://www.rna-seqblog.com/rpkm-fpkm-and-tpm-clearly-explained>.

### Why don't you have RPM in your normalization? <a href="#why-dont-you-have-rpm-in-your-normalization" id="why-dont-you-have-rpm-in-your-normalization"></a>

RPM (reads per million) is the same as Total Count. Please use Total Count.

### What is a canonical transcript? <a href="#what-is-a-canonical-transcript" id="what-is-a-canonical-transcript"></a>

For genes with multiple transcripts, one of the transcripts is picked as the canonical transcript. Based on the UCSC definition from the table browser,

> **knownCanonical** - identifies the canonical isoform of each cluster ID, or gene. Generally, this is the longest isoform.

we define the canonical transcript as either the longest CDS (coding DNA sequence) if the gene has translated transcripts, or the longest cDNA.

### Why are there decimal values in the Partek E/M quantification output? <a href="#why-are-there-decimal-values-in-the-partek-e-m-quantification-output" id="why-are-there-decimal-values-in-the-partek-e-m-quantification-output"></a>

The Partek E/M quantification algorithm can give decimal values because of multi-mapping reads (the same read potentially aligning to multiple locations) and overlapping transcripts/genes (a read that maps to a location with multiple transcripts or genes at that location). In these scenarios, the read count will be split.

For example, if a read maps to two potential locations, then that read contributes 0.5 counts to the first location and 0.5 counts to the second location. Similarly, if a read maps to one location with two overlapping genes, then that read contributes 0.5 counts to the first gene and 0.5 counts to the second gene.

If you need to remove the decimal points for downstream analysis outside of Partek Flow, you can round the values to the nearest integer.

### Why is the number of variants listed in the variant reports and summarize cohort mutations report different? <a href="#why-is-the-number-of-variants-listed-in-the-variant-reports-and-summarize-cohort-mutations-report-di" id="why-is-the-number-of-variants-listed-in-the-variant-reports-and-summarize-cohort-mutations-report-di"></a>

For variants with multiple alternative alleles, the variant has one row for all alternative alleles, while the summarize cohort mutations report lists each alternative allele on a separate rows. The number of variants listed at the top of the each report is calculated from the number of rows in the report.

### How long is data retained for expired Partek Flow subscriptions?

Partek Flow data is retained for one year after the subscription expires. We recommend exporting your projects and data to your local machine prior to letting the subscription expire.

## Visualization

### How do I order my heatmap by the cell types?

If you would like specific groups (e.g. cell types) in a certain order, do not perform [Hierarchical clustering](/partek-flow/tutorials/bulk-rna-seq/generating-a-hierarchical-clustering-heatmap) on these cells and instead choose to assign order, then use click and drag to reorder the groups. If you want to remove a group, you can choose to exclude this group in the filtering section. You can still perform [Hierarchical clustering](/partek-flow/tutorials/bulk-rna-seq/generating-a-hierarchical-clustering-heatmap) on the features if you would like to. Hierarchical clustering will force the heatmap to cluster and you would need to click the dendrogram nodes to switch the order. Click [here ](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering)for more information.

### How do I display UMAP for each sample in the Data Viewer?

For a multi-sample project, all of the downstream tasks will be run separately if 'Split by sample' was checked when performing the PCA task. Visualization of different samples can be displayed by 'Sample' using the 'Misc' section in the Axes card. To show different samples side by side, one can click 'Duplicate plot' first, then use the 'Sample' option to switch the samples.

### Can I visualize fold change values on a heatmap without using a z-score?

Yes, the default settings can be modified by clicking "Configure" in the Advanced settings during task set-up, then change the "feature scaling" option to "none" to plot the values without scaling. For more information related to to the heatmap click [here](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering).

### Why don't I see Flip mode on the heatmap? Why can't I download all of the data after zooming?

The Flip mode and download all data options are disabled if there are more than 2.5 million values (rows x columns) in the heatmap.

### How to label gene names on volcano plot?

By default, genes are selected if the p-value is <=0.05 and |fold change| >=2 and when the number of selected genes is less than 2000 genes, they will be labeled. You can click on Style button in Configure section, choose a gene annotation field from the Label by drop-down list to change the label. If you number of selected genes is select less than or equal to 100, Partek Flow will try to spread out labels as much as possible to clearly display the labels. If number of selected genes is more than 100, labels will be next to the selected genes, there will be overlaps where genes are close together. If there are more than 2000 genes selected, no label will be displayed.

If you click any blank space, you can turn off select and use different selection mode button on the vertical bar on the upper-right corner of the plot to manually select dots on the plot.

## Statistics

### Why do I get "?" for FDR p-values in my Deseq2 result?

When a feature (gene) has low expression, it will be filtered by automatic independent filtering. To avoid this, you can either [filter features](/partek-flow/tutorials/bulk-rna-seq/filtering-features) to exclude low expression features before Deseq2, or in the Deseq2 advanced options, choose apply [independent filtering](/partek-flow/user-manual/task-menu/differential-analysis/deseq2-r-vs-deseq2#deseq2-r-vsdeseq2-troubleshooting) setting. Details about independent filtering can be found at the [Deseq2 documentation](https://bioconductor.org/packages/release/bioc/vignettes/DESeq2/inst/doc/DESeq2.html#indfilt).

[Click here for troubleshooting other differential analysis models and "?" results](/partek-flow/user-manual/task-menu/differential-analysis/troubleshooting)

### What is fold change?

Fold change indicates the extent of increase or decrease in feature expression in a comparison. In Partek Flow, fold change is in linear scale (even if the input data is in log scale). It is converted from ratio, which is the LSmean of group one divided by LSmean of group two in your comparison. When the ratio is greater than 1, fold change is identical to ratio; when the ratio is less than 1, fold change is -1/ratio. There is no fold change value between -1 to 1. When ratio/fold change is 1, that means there is no change between the two groups.

Log ratio option in Partek Flow is converted from ratio, this is a value comparable to log fold change in some other tools.

### Can I label a Volcano plot with gene names?

Yes, go to Style in the [Data Viewer](/partek-flow/user-manual/data-viewer) and make sure Gene name is selected under "Labeling". Next, go to the in plot selection tools (right side of the graphic) and use any of the selection tools to select the cells that you would like to label. You can use ctrl or shift to select multiple populations at once. For more information on the Volcano plot click [here](/partek-flow/user-manual/visualizations/volcano-plot).

### In Volcano plot, what is inconclusive group mean?

By default, Flow is using the p value <= 0.05 and |fold change|>=2 as the significance cutoff. If genes meet both p value and fold change cutoff, they are significantly up or down regulated genes. If they only meet one criteria, they are called inconclusive. If genes won't pass either criteria, they are not significant. Click on the Statistics button in the Configure section in the left control panel, you can change the cutoff. Click on the Style button to change the color of significance categories.

### What is the difference between FDR and FDR step up?

FDR is the expected proportion of false discoveries among all discoveries. FDR Step-up is a particular method to keep FDR under a given level, alpha, that was proposed in this [paper](https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/j.2517-6161.1995.tb02031.x). In Partek Flow, if one calls all of the features with p-values 0.02 or less, the FDR is less or equal to 0.41.

### How to perform a paired t-Test in Flow

You should have at least the following two attributes in the Metadata, treatment (including two subgroups) and subject ID (to pair the two samples). When performing differential analysis, choose ANOVA and include both attributes into the ANOVA model, the two-way ANOVA is mathematically equivalent to paired t-Test.

### Can I compare one attribute at a time versus all of the others combined?

Yes, you can use the [Compute biomarkers](/partek-flow/user-manual/task-menu/differential-analysis/compute-biomarkers) task to compare one subgroup at a time to all of the others combined. An alternative option is to set up the differential analysis model in this way; for more information please see the information here for each model.

### I downloaded gene counts from the output data node generated by the Quantify to annotation model task, why can't I find my genes of interest?

In the [Quantifying to an annotation model](/partek-flow/tutorials/bulk-rna-seq/quantifying-to-an-annotation-model) dialog, by default, Partek Flow filters features based on the total count across all of the samples and features with a total count greater than 10 will be reported. If you want to report all of the genes in the annotation file, change the Filter features value to **0**.

## Biological Interpretation

### What is the difference between GSEA and Gene Set Enrichment?

In Partek Flow, GSEA should be performed on a sample/cell and feature matrix data node (e.g. normalization count data). GSEA is used to detect a gene set/a pathway which is significantly different between two groups. Gene set enrichment should be performed on a filtered gene list; it is used to identify overrepresented gene set/pathway based the filtered gene list using Fisher's exact test. The input data is a filtered list using gene names.

### What is the enrichment score shown in the Gene Set Enrichment report?

The enrichment score shown in the enrichment report is the negative natural log of the enrichment p-value derived from Fisher Exact test. The higher the enrichment score, the more overrepresented our list of genes in the gene set of a GO/pathway category.

### In KEGG pathway, genes can be colored by Fold change and p-value etc, how are the gene statistics calculated?

For Gene set enrichment analysis, only genes from the input data node (filtered gene list) will be colored in the KEGG pathway gene network, using the statistics in the data node.

During GSEA (or Gene set ANOVA) computation, we also perform ANOVA on each gene based on the attributed selected independent from GESA computation (at gene set level). The results of ANOVA is only used to color the genes in the KEGG gene network. If GSEA is computed using another other database, e.g. GO, we don't compute ANOVA on each gene since GO databased doesn't have gene network information.

### When should I use GSEA or Gene set ANOVA?

Both methods should be performed on a normalized matrix data node, and requires gene symbol in feature annotation. Both methods are detecting a differentially expressed Gene set (pathway) instead of each individual gene. The algorithms are different. GSEA is a popular method from the [Broad institute](https://www.gsea-msigdb.org/gsea/index.jsp). Gene Set ANOVA is based on generalized linear model, [here ](/partek-flow/white-papers/gene-set-anova)are the details.


# General

## How to create a project? <a href="#how-to-create-a-project" id="how-to-create-a-project"></a>

To create a project, you first need to [transfer files to the Partek Flow server](/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server), and then import the files into your project using the import data wizard, here is the [video](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial) and [more information](/partek-flow/user-manual/importing-data).

## Can I change my user avatar? <a href="#can-i-change-my-user-avatar" id="can-i-change-my-user-avatar"></a>

Yes, navigate to [My profile](/partek-flow/user-manual/settings/personal/my-profile) and click the "Change image" button. Do this by clicking your avatar at the top right corner of the interface, select [Settings](/partek-flow/user-manual/settings), then choose [Profile](/partek-flow/user-manual/settings/personal/my-profile).

## How do I add and use my own lists? <a href="#how-do-i-add-and-use-my-own-lists" id="how-do-i-add-and-use-my-own-lists"></a>

Click your avatar in the top right corner of the Partek Flow interface, choose [Settings](/partek-flow/user-manual/settings) in the menu, and select [Lists](/partek-flow/user-manual/settings/components/lists) from the left panel of the [Components](/partek-flow/user-manual/settings/components) section. Lists can also be generated from result tables using the "Save as managed list" button. For more information please click [here](/partek-flow/user-manual/settings/components/lists).

## Can I repeat a task and everything downstream of it, while changing only one/a few parameters? <a href="#can-i-repeat-a-task-and-everything-downstream-of-it-while-changing-only-one-a-few-parameters" id="can-i-repeat-a-task-and-everything-downstream-of-it-while-changing-only-one-a-few-parameters"></a>

Yes, click on the rectangular task that you want to change the parameters. On the context-specific menu on the right, under Task actions, select ‘Rerun with downstream tasks’, this will bring you to the task set up page where you can edit the parameters for the task, then click Finish to run the task with the new parameters. The tasks downstream of it will be initiated automatically.

## What can I use to identify cells that are actively expressing genes within a gene list? <a href="#what-can-i-use-to-identify-cells-that-are-actively-expressing-genes-within-a-gene-list" id="what-can-i-use-to-identify-cells-that-are-actively-expressing-genes-within-a-gene-list"></a>

Use [AUCell](/partek-flow/user-manual/task-menu/exploratory-analysis/aucell) to identify cells with active gene sets; this task calculates a value for each cell by ranking all genes by their expression level in the cell and identifying what proportion of the genes from the gene list fall within the top 5% (default cutoff) of genes. An alternative option is to use the Gene score for a feature list to select and filter populations based on the distribution; [click here for more information](/partek-flow/user-manual/settings/components/lists).

## Can I build and use pipelines for my analysis? <a href="#can-i-build-and-use-pipelines-for-my-analysis" id="can-i-build-and-use-pipelines-for-my-analysis"></a>

Yes, click on [Import a pipeline](/partek-flow/user-manual/pipelines/importing-a-pipeline) on the bottom of the [Analyses](/partek-flow/tutorials/creating-and-analyzing-a-project/the-analyses-tab) tab dashboard. This will help you import either our hosted pipelines or your own saved pipeline which can be found under [Settings](/partek-flow/user-manual/settings) -> [Components](/partek-flow/user-manual/settings/components) -> [Pipelines](/partek-flow/user-manual/settings/components/pipeline-management). Click [here](/partek-flow/tutorials/bulk-rna-seq/saving-and-running-a-pipeline) for steps to save and run a pipeline. For more information related to navigating pipelines click [here](/partek-flow/user-manual/pipelines).

## How do I classify cells? <a href="#how-do-i-classify-cells" id="how-do-i-classify-cells"></a>

Classification in Partek Flow can be performed manually or with automatic cell classification which is explained in more detail [here](/partek-flow/user-manual/task-menu/classification). Users often want to classify cells by gene expression threshold(s), for details on classification by marker expression click [here](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/perform-exploratory-analysis). Automatic classification needs to be performed on a non-normalized single cell data node; once complete, [publish cell attributes to project](/partek-flow/user-manual/task-menu/annotation-metadata/publish-cell-attributes-to-project) then use this classification in visualizations and tasks. You may choose to perform [Graph-based clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/graph-based-clustering) and [K-means clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/k-means-clustering) to help identify biomarkers that can then be used to identify the clusters and we also provide hosted lists for different cell types.

## My server is full, how do I make more space? <a href="#my-server-is-full-how-do-i-make-more-space" id="my-server-is-full-how-do-i-make-more-space"></a>

We recommend cleaning up projects as well as removing library files that you do not need, then removing the orphaned files. You can also export analyzed projects and save them on an external machine, then when you need them again you can import them to the server. Please see this information for more details related to: [Project management](/partek-flow/tutorials/creating-and-analyzing-a-project/project-management), [Removing library files](/partek-flow/user-manual/settings/components/library-file-management/removing-library-files), and [Orphaned files](/partek-flow/user-manual/settings/access/orphaned-files). Right click on the data node to delete files from projects that are not needed (e.g. fastqs from project pipelines that are analyzed); you will not be able to perform tasks from this node once the files are deleted.

## How do I add library files if I am not studying human or mouse? <a href="#how-do-i-add-library-files-if-i-am-not-studying-human-or-mouse" id="how-do-i-add-library-files-if-i-am-not-studying-human-or-mouse"></a>

To add a new assembly, click on [Settings](/partek-flow/user-manual/settings) -> [Library files](/partek-flow/user-manual/settings/components/library-file-management). From the Assembly drop-down list, select Add assembly and specify the species. If the species name is not in the list, choose Other and type in the name with the assembly version (multiple assembly versions can exist for one species, e.g. hg19 and hg38 for Homo Sapiens). You need to add the reference file which is a .fasta file containing sequence information. Once the reference file is added, you can build any aligner index to perform the alignment task.

The Annotation model is a file containing feature location. This file can be used to quantify to annotation model in RNA-Seq analysis, or annotate variant or peaks in a DNA-Seq or ATAC-Seq/ChIP-Seq data analysis pipeline. The file format should be .gtf/.gff/.bed.

We recommend looking for the species files on the [Ensembl](https://useast.ensembl.org/index.html) website. There is no need to unzip or save these files to your local machine, instead right click and copy the link address of the specific file (not a link to a folder). For more details, here is the documentation chapter: [Library File Management - Partek® Documentation](/partek-flow/user-manual/settings/components/library-file-management).

## Are Genome coordinates 1-based or 0-based? <a href="#are-genome-coordinates-1-based-or-0-based" id="are-genome-coordinates-1-based-or-0-based"></a>

Genome coordinates for annotation models stored in Partek Flow are 1-based, start-inclusive, and stop-exclusive. This means that the first base position starts from one, the start coordinate for a feature is included in the feature and the stop/end coordinate is not included in the feature. These are the genome coordinates that are printed in various task reports and output files when an annotation model is involved in the task. When custom annotation files are added to Partek Flow, the genome coordinates are converted into this format. The coordinates are converted back if necessary for a specific task. shows how the genome coordinates vary between different annotation formats.

![](https://help.partek.illumina.com/~gitbook/image?url=https%3A%2F%2F1384254481-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FJVEESmJAPppJ3ijFq5aR%252Fuploads%252Fgit-blob-62ec1d304bd259c3e91b7d46041b3ed3ed9abbda%252Fare-genome-coordinates-1-based-or-0-based.png%3Falt%3Dmedia\&width=768\&dpr=4\&quality=100\&sign=feaa56b3\&sv=1)

## Can I add transgenes to my reference files? <a href="#can-i-add-transgenes-to-my-reference-files" id="can-i-add-transgenes-to-my-reference-files"></a>

Yes, to add transgenes (including gfp or related) to the references files, first choose an assembly, create the transgene reference, and merge the references together (e.g. combine mm10 with dttomato). This is the same process for the annotation file.

## How do I export data from the result nodes? <a href="#how-do-i-export-data-from-the-result-nodes" id="how-do-i-export-data-from-the-result-nodes"></a>

Left click to select the data node you want to export. In the bottom of the task menu there will be an option to [Download data](/partek-flow/quick-start-guide#downloading-your-data).

## Why can’t I find the RPKM method on the normalization menu (I see FPKM)? <a href="#why-cant-i-find-the-rpkm-method-on-the-normalization-menu-i-see-fpkm" id="why-cant-i-find-the-rpkm-method-on-the-normalization-menu-i-see-fpkm"></a>

When working with paired data it should be the case that FPKM is available, and when working with single end data RPKM should be available. These metrics are essentially analogous, but based on the underlying method used for calculation (accounting for two reads mapping to 1 fragment and not counting twice for paired end data). Here is a simple description of the differences in calculation between RPKM and FPKM: <http://www.rna-seqblog.com/rpkm-fpkm-and-tpm-clearly-explained>.

## Why don't you have RPM in your normalization? <a href="#why-dont-you-have-rpm-in-your-normalization" id="why-dont-you-have-rpm-in-your-normalization"></a>

RPM (reads per million) is the same as Total Count. Please use Total Count.

## What is a canonical transcript? <a href="#what-is-a-canonical-transcript" id="what-is-a-canonical-transcript"></a>

For genes with multiple transcripts, one of the transcripts is picked as the canonical transcript. Based on the UCSC definition from the table browser,

> **knownCanonical** - identifies the canonical isoform of each cluster ID, or gene. Generally, this is the longest isoform.

we define the canonical transcript as either the longest CDS (coding DNA sequence) if the gene has translated transcripts, or the longest cDNA.

## Why are there decimal values in the Partek E/M quantification output? <a href="#why-are-there-decimal-values-in-the-partek-e-m-quantification-output" id="why-are-there-decimal-values-in-the-partek-e-m-quantification-output"></a>

The Partek E/M quantification algorithm can give decimal values because of multi-mapping reads (the same read potentially aligning to multiple locations) and overlapping transcripts/genes (a read that maps to a location with multiple transcripts or genes at that location). In these scenarios, the read count will be split.

For example, if a read maps to two potential locations, then that read contributes 0.5 counts to the first location and 0.5 counts to the second location. Similarly, if a read maps to one location with two overlapping genes, then that read contributes 0.5 counts to the first gene and 0.5 counts to the second gene.

If you need to remove the decimal points for downstream analysis outside of Partek Flow, you can round the values to the nearest integer.

## Why is the number of variants listed in the variant reports and summarize cohort mutations report different? <a href="#why-is-the-number-of-variants-listed-in-the-variant-reports-and-summarize-cohort-mutations-report-di" id="why-is-the-number-of-variants-listed-in-the-variant-reports-and-summarize-cohort-mutations-report-di"></a>

For variants with multiple alternative alleles, the variant has one row for all alternative alleles, while the summarize cohort mutations report lists each alternative allele on a separate rows. The number of variants listed at the top of the each report is calculated from the number of rows in the report.

## How long is data retained for expired Partek Flow subscriptions?

Partek Flow data is retained for one year after the subscription expires. We recommend exporting your projects and data to your local machine prior to letting the subscription expire.


# Visualization

## How do I order my heatmap by the cell types? <a href="#how-do-i-order-my-heatmap-by-the-cell-types" id="how-do-i-order-my-heatmap-by-the-cell-types"></a>

If you would like specific groups (e.g. cell types) in a certain order, do not perform [Hierarchical clustering](/partek-flow/tutorials/bulk-rna-seq/generating-a-hierarchical-clustering-heatmap) on these cells and instead choose to assign order, then use click and drag to reorder the groups. If you want to remove a group, you can choose to exclude this group in the filtering section. You can still perform [Hierarchical clustering](/partek-flow/tutorials/bulk-rna-seq/generating-a-hierarchical-clustering-heatmap) on the features if you would like to. Hierarchical clustering will force the heatmap to cluster and you would need to click the dendrogram nodes to switch the order. Click [here](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering) for more information.

## How do I display UMAP for each sample in the Data Viewer? <a href="#how-do-i-display-umap-for-each-sample-in-the-data-viewer" id="how-do-i-display-umap-for-each-sample-in-the-data-viewer"></a>

For a multi-sample project, all of the downstream tasks will be run separately if 'Split by sample' was checked when performing the PCA task. Visualization of different samples can be displayed by 'Sample' using the 'Misc' section in the [Axes](/partek-flow/user-manual/data-viewer) card. To show different samples side by side, one can click 'Duplicate plot' first, then use the 'Sample' option to switch the samples.

## Can I visualize fold change values on a heatmap without using a z-score? <a href="#can-i-visualize-fold-change-values-on-a-heatmap-without-using-a-z-score" id="can-i-visualize-fold-change-values-on-a-heatmap-without-using-a-z-score"></a>

Yes, the default settings can be modified by clicking "Configure" in the Advanced settings during task set-up, then change the "feature scaling" option to "none" to plot the values without scaling. For more information related to the heatmap click [here](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering).

## Why don't I see Flip mode on the heatmap? Why can't I download all of the data after zooming? <a href="#why-dont-i-see-flip-mode-on-the-heatmap-why-cant-i-download-all-of-the-data-after-zooming" id="why-dont-i-see-flip-mode-on-the-heatmap-why-cant-i-download-all-of-the-data-after-zooming"></a>

The Flip mode and download all data options are disabled if there are more than 2.5 million values (rows x columns) in the heatmap.

## How to label gene names on volcano plot? <a href="#how-to-label-gene-names-on-volcano-plot" id="how-to-label-gene-names-on-volcano-plot"></a>

By default, genes are selected if the p-value is <=0.05 and |fold change| >=2 and when the number of selected genes is less than 2000 genes, they will be labeled. You can click on Style button in Configure section, choose a gene annotation field from the Label by drop-down list to change the label. If you number of selected genes is select less than or equal to 100, Partek Flow will try to spread out labels as much as possible to clearly display the labels. If number of selected genes is more than 100, labels will be next to the selected genes, there will be overlaps where genes are close together. If there are more than 2000 genes selected, no label will be displayed.

If you click any blank space, you can turn off select and use different selection mode button on the vertical bar on the upper-right corner of the plot to manually select dots on the plot.


# Statistics

## Why do I get "?" for FDR p-values in my Deseq2 result? <a href="#why-do-i-get-for-fdr-p-values-in-my-deseq2-result" id="why-do-i-get-for-fdr-p-values-in-my-deseq2-result"></a>

When a feature (gene) has low expression, it will be filtered by automatic independent filtering. To avoid this, you can either [filter features](/partek-flow/tutorials/bulk-rna-seq/filtering-features) to exclude low expression features before Deseq2, or in the Deseq2 advanced options, choose [apply independent filtering](/partek-flow/user-manual/task-menu/differential-analysis/deseq2-r-vs-deseq2#deseq2-r-vsdeseq2-troubleshooting) setting. Details about independent filtering can be found at the [Deseq2 documentation](https://bioconductor.org/packages/release/bioc/vignettes/DESeq2/inst/doc/DESeq2.html#indfilt).

[Click here for troubleshooting other differential analysis models and "?" results](/partek-flow/user-manual/task-menu/differential-analysis/troubleshooting)

## What is fold change? <a href="#what-is-fold-change" id="what-is-fold-change"></a>

Fold change indicates the extent of increase or decrease in feature expression in a comparison. In Partek Flow, fold change is in linear scale (even if the input data is in log scale). It is converted from ratio, which is the LSmean of group one divided by LSmean of group two in your comparison. When the ratio is greater than 1, fold change is identical to ratio; when the ratio is less than 1, fold change is -1/ratio. There is no fold change value between -1 to 1. When ratio/fold change is 1, that means there is no change between the two groups.

Log ratio option in Partek Flow is converted from ratio, this is a value comparable to log fold change in some other tools.

## Can I label a Volcano plot with gene names? <a href="#can-i-label-a-volcano-plot-with-gene-names" id="can-i-label-a-volcano-plot-with-gene-names"></a>

Yes, go to Style in the [Data Viewer](/partek-flow/user-manual/data-viewer) and make sure Gene name is selected under "Labeling". Next, go to the in plot selection tools (right side of the graphic) and use any of the selection tools to select the cells that you would like to label. You can use ctrl or shift to select multiple populations at once. For more information on the Volcano plot click [here](/partek-flow/user-manual/visualizations/volcano-plot).

## In Volcano plot, what is inconclusive group mean? <a href="#in-volcano-plot-what-is-inconclusive-group-mean" id="in-volcano-plot-what-is-inconclusive-group-mean"></a>

By default, Flow is using the p value <= 0.05 and |fold change|>=2 as the significance cutoff. If genes meet both p value and fold change cutoff, they are significantly up or down regulated genes. If they only meet one criteria, they are called inconclusive. If genes won't pass either criteria, they are not significant. Click on the Statistics button in the Configure section in the left control panel, you can change the cutoff. Click on the Style button to change the color of significance categories.

## What is the difference between FDR and FDR step up? <a href="#what-is-the-difference-between-fdr-and-fdr-step-up" id="what-is-the-difference-between-fdr-and-fdr-step-up"></a>

FDR is the expected proportion of false discoveries among all discoveries. FDR Step-up is a particular method to keep FDR under a given level, alpha, that was proposed in this [paper](https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/j.2517-6161.1995.tb02031.x). In Partek Flow, if one calls all of the features with p-values 0.02 or less, the FDR is less or equal to 0.41.

## How to perform a paired t-Test in Flow <a href="#how-to-perform-a-paired-t-test-in-flow" id="how-to-perform-a-paired-t-test-in-flow"></a>

You should have at least the following two attributes in the Metadata, treatment (including two subgroups) and subject ID (to pair the two samples). When performing differential analysis, choose ANOVA and include both attributes into the ANOVA model, the two-way ANOVA is mathematically equivalent to paired t-Test.

## Can I compare one attribute at a time versus all of the others combined? <a href="#can-i-compare-one-attribute-at-a-time-versus-all-of-the-others-combined" id="can-i-compare-one-attribute-at-a-time-versus-all-of-the-others-combined"></a>

Yes, you can use the [Compute biomarkers](/partek-flow/user-manual/task-menu/differential-analysis/compute-biomarkers) task to compare one subgroup at a time to all of the others combined. An alternative option is to set up the differential analysis model in this way; for more information please see the information [here](/partek-flow/user-manual/task-menu/differential-analysis) for each model.

## I downloaded gene counts from the output data node generated by the Quantify to annotation model task, why can't I find my genes of interest? <a href="#i-downloaded-gene-counts-from-the-output-data-node-generated-by-the-quantify-to-annotation-model-tas" id="i-downloaded-gene-counts-from-the-output-data-node-generated-by-the-quantify-to-annotation-model-tas"></a>

In the[ Quantifying to an annotation model](/partek-flow/tutorials/bulk-rna-seq/quantifying-to-an-annotation-model) dialog, by default, Partek Flow filters features based on the total count across all of the samples and features with a total count greater than 10 will be reported. If you want to report all of the genes in the annotation file, change the Filter features value to 0.

### &#x20;<a href="#biological-interpretation" id="biological-interpretation"></a>


# Biological Interpretation

## What is the difference between GSEA and Gene Set Enrichment? <a href="#what-is-the-difference-between-gsea-and-gene-set-enrichment" id="what-is-the-difference-between-gsea-and-gene-set-enrichment"></a>

In Partek Flow, GSEA should be performed on a sample/cell and feature matrix data node (e.g. normalization count data). GSEA is used to detect a gene set/a pathway which is significantly different between two groups. Gene set enrichment should be performed on a filtered gene list; it is used to identify overrepresented gene set/pathway based the filtered gene list using Fisher's exact test. The input data is a filtered list using gene names.

## What is the enrichment score shown in the Gene Set Enrichment report? <a href="#what-is-the-enrichment-score-shown-in-the-gene-set-enrichment-report" id="what-is-the-enrichment-score-shown-in-the-gene-set-enrichment-report"></a>

The enrichment score shown in the enrichment report is the negative natural log of the enrichment p-value derived from Fisher Exact test. The higher the enrichment score, the more overrepresented our list of genes in the gene set of a GO/pathway category.

## In KEGG pathway, genes can be colored by Fold change and p-value etc, how are the gene statistics calculated? <a href="#in-kegg-pathway-genes-can-be-colored-by-fold-change-and-p-value-etc-how-are-the-gene-statistics-calc" id="in-kegg-pathway-genes-can-be-colored-by-fold-change-and-p-value-etc-how-are-the-gene-statistics-calc"></a>

For Gene set enrichment analysis, only genes from the input data node (filtered gene list) will be colored in the KEGG pathway gene network, using the statistics in the data node.

During GSEA (or Gene set ANOVA) computation, we also perform ANOVA on each gene based on the attributed selected independent from GESA computation (at gene set level). The results of ANOVA is only used to color the genes in the KEGG gene network. If GSEA is computed using another other database, e.g. GO, we don't compute ANOVA on each gene since GO databased doesn't have gene network information.

## When should I use GSEA or Gene set ANOVA? <a href="#when-should-i-use-gsea-or-gene-set-anova" id="when-should-i-use-gsea-or-gene-set-anova"></a>

Both methods should be performed on a normalized matrix data node, and requires gene symbol in feature annotation. Both methods are detecting a differentially expressed Gene set (pathway) instead of each individual gene. The algorithms are different. GSEA is a popular method from the [Broad institute](https://www.gsea-msigdb.org/gsea/index.jsp). Gene Set ANOVA is based on generalized linear model, [here](/partek-flow/white-papers/gene-set-anova) are the details.


# How to cite Partek software

**In-text:**

Proper in-text citation for Partek software must include the trademarked name of the Partek software product used and the version number at the time of data analysis.

**Examples:**

Gene specific analysis was performed using *Partek™ Flow™* software, v11.0.

Batch removal was performed using *Partek™ Genomics Suite™* software, v7.0.

**Reference or bibliography list:**

Citing Partek software in a bibliography or reference list should include the company name, copyright date, trademarked name of the Partek software product, version number, format, and link to product website.

**Example:**

Partek, an Illumina company (2024). Partek™ Flow™ (Version 11.0) \[Computer software]. <https://www.illumina.com/products/by-type/informatics-products/partek-flow.html>


# Quick Start Guide

## Overview

This guide outlines the basics of Partek Flow usage. Partek Flow can be installed in either a server, computer cluster or on the cloud. Regardless of where it's installed, it can be viewed using any web browser. We recommend using Google Chrome.

This guide covers:

* [Overview](#overview)
* [Starting a new project](#starting-a-new-project)
* [Basic Partek Flow layout](#basic-partek-flow-layout)
* [Saving visualizations](#saving-visualizations)
* [Downloading your data](#downloading-your-data)
* [Viewing task details](#viewing-task-details)
* [Additional navigation](#additional-navigation)
* [Access tutorial data from the interface](#access-tutorial-data-from-the-interface)
* [Partek Flow in action](#partek-flow-in-action)

Logging in to your Partek Flow account will bring up the Home page. This page shows recent projects you've worked on and pertinent details about each project. Click the project name to open the project and continue the analysis.

<figure><img src="/files/OmAdECTPsfUocYEJlyCg" alt=""><figcaption><p><em>Access projects from the homepage by clicking on the project name</em></p></figcaption></figure>

## Starting a new project

From the Home Page, click the **Add project** button.

<div align="center"><figure><img src="/files/ecmrbgw0FgS1kNzvKYJd" alt=""><figcaption><p><em>Click the Add project button to start a new projec</em>t</p></figcaption></figure></div>

Assign a name to the project and click the **Create project** button.

<figure><img src="/files/lzHdSPB4S6Jl9nRgtHxW" alt="" width="228"><figcaption><p><em>Name the project and click Create project</em></p></figcaption></figure>

### Uploading your dataset

Upon creation of a new project, the Analyses tab will appear, prompting you to add samples to your project.

Click the blue Add data button <img src="/files/kaJfyXF5Oqr1BoaSF10X" alt="Add data button" data-size="original">

Select the type of data (**Single cell, Bulk, Other**), choose the **assay** type, and select the data **format** for the files you would like to create samples from. Partek Flow accepts various data types. Use the **Next** button to proceed with import.

There are three ways you can upload the data:

1. From your Partek Flow server ([click here for more information](https://help.partek.illumina.com/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server))
2. From a URL
3. From a GEO / ENA Bioproject ([click here for more information](https://help.partek.illumina.com/partek-flow/user-manual/importing-data/import-a-geo-ena-project))

Because genomics datasets are generally large, it is ideal to have the data copied in a folder directly accessible to the Partek Flow server. Make sure that the directory has the appropriate permissions for Partek Flow to read and write files in that folder. You may wish to seek assistance from your system administrator in uploading your data directly.

## Basic Partek Flow layout

### The Metadata Tab

Once the samples have been created, assign the corresponding sample attributes for each sample using the **Metadata** tab.

<figure><img src="/files/6j5Uv3yAnwkmP1Yc3CD8" alt=""><figcaption><p><em>Sample and cell attributes are managed in the Metadata tab</em></p></figcaption></figure>

The most efficient way to assign sample attributes is by clicking **Assign sample attributes from a file** and uploading a tab delimited text file. The file should contain a table with the following:

* The first row lists the attribute names (e.g. Treatment, Exposure) and
* The first column of the table lists the sample names (the sample names in the file must be identical to the ones listed in the Sample name column in the Metadata tab)
* List the corresponding attributes for each sample in the succeeding columns

If you already have attribute information about your cells from a single cell experiment, you can use the **Annotate cells** task in the **Analyses tab** of Partek Flow to apply this information to the data. [Please click here for more information on adding a priori cell annotation. ](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/annotation-metadata/annotate-cells)Once applied, these can be used like any other attribute in Partek Flow, and thus can be used for cell selection, classification and differential analysis. Data nodes in the Analyses pipeline that contain this cell level information (e.g. Graph-based clusters data node) can be published to the project and will then be available to **Manage** under Cell attributes. [Please click here for more information on publishing cell attributes to the project.](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/annotation-metadata/publish-cell-attributes-to-project)

### The Analyses Tab

After samples have been added and associated with valid data files, a data node will appear in the Analyses tab. The Analyses tab is where you can invoke tasks, using the toolbox on the right, and view the results of your analysis.

To add more data, use the context sensitive menu on the right (toolbox) and choose **Add data** under *Import.* There is also an option in the **Metadata** tab to **Add data**.

Data can no longer be added to the project once the analyses begins and tasks are performed from the initial data node.

<figure><img src="/files/TDFbgWJClpcDMNahHDkl" alt=""><figcaption><p><em>The Analyses tab showing a data node of unaligned reads. Use the Context-sensitive menu on the right to run tasks.</em></p></figcaption></figure>

### Data and task nodes

The Analyses tab contains two elements: data nodes (circles) and task nodes (rectangles) connected by lines and arrows . Collectively, they represent a data analysis pipeline.

<figure><img src="/files/QbU2upszzK9JQGk1iKaJ" alt=""><figcaption><p><em>Example of a data analysis pipeline</em></p></figcaption></figure>

### Performing tasks

Clicking a data node brings up a context sensitive menu on the right. This menu changes depending on the type of data node. It will only present tasks which can be performed on that specific data type. Hover over the task to obtain additional information regarding each option.

<figure><img src="/files/KAQQb7iZoFcH8vsNXQR4" alt=""><figcaption><p><em>Use the context sensitive menu to run a task on the data node</em></p></figcaption></figure>

Select the task you wish to perform from the menu. When configuring task options, additional information regarding each option is available. Click **Finish** to perform the task.

<figure><img src="/files/3jwfKqKyeH4YWuJwIaot" alt=""><figcaption><p><em>Configure the task options to meet your needs and click Finish</em></p></figcaption></figure>

When available, hover over <img src="/files/NVjRqSk652F2pR4YQ4xY" alt="Tooltip" data-size="line"> Tooltips or click the <img src="/files/j4U3Qyz71lz2KsNXlv0B" alt="Video icon" data-size="line"> video help for decision making.

Depending on the task, a new data node may automatically be created and connected to the original data node. This contains the data resulting from the task. Tasks that do not produce new data types, such as Pre-alignment QA/QC, will not produce an additional data node.

To view the results of a task, click the data node and choose the **Task report** option on the menu.

## Saving visualizations

Some visualizations are available in the task report and others are sent to the Data viewer for modification and export.

### Save visualizations from a task report

Click Save <img src="/files/aeNqRq6OWkPDl4G0ozzW" alt="Save button" data-size="line"> on any visualization to export a publication-quality image.

<figure><img src="/files/dSX2NPkzAfx5gO6DghjC" alt=""><figcaption><p><em>Save visualizations using the save icon</em></p></figcaption></figure>

### Save visualizations from the Data viewer

To save an individual image within the Data viewer to your machine, click <img src="/files/aeNqRq6OWkPDl4G0ozzW" alt="Save button" data-size="line"> **Export image** in the top right corner of the image and select the format, size, and resolution then click **Save**.

<figure><img src="/files/c93jk4D8EKapSygnPWRt" alt=""><figcaption><p><em>Click export image to Save the image in different formats, sizes and resolutions</em></p></figcaption></figure>

For more information on using the Data viewer [please click here](https://help.partek.illumina.com/partek-flow/user-manual/data-viewer).

## Downloading your data

Data associated with any data node can be downloaded by clicking the node and choosing **Download data** at the bottom of the task menu. Compressed files will be downloaded to the local computer where the user is accessing the Partek Flow server. Note that bigger files (such as unaligned reads) would take longer to download. For guidance, a file size estimate is provided for each data node.

<figure><img src="/files/68TpuINS3AC0HvDmE8av" alt=""><figcaption><p><em>Select the data node and navigate to the bottom of the task menu then click Download data</em></p></figcaption></figure>

## Viewing task details

There are three ways to view details about the tasks performed in Partek Flow.

* Click the individual task node (rectangle) and select **Task details** from the task menu
* Click the data node (circle) and select **Data summary report** from the task menu
* Click the **Log** tab

### View individual task details

This is used to view the task details from an individual task, including the commands. Navigate to the the task node (rectangle) and select **Task details** from the task menu.

<figure><img src="/files/9wKWzF9aS7YpC8e0dZMp" alt=""><figcaption><p><em>Select the task node (rectangle) and click Task details to view information about the task including the commands</em></p></figcaption></figure>

### View task details for all performed tasks upstream of the data node

This is used to access the task details for all of the steps performed upstream from the data node that has been selected. Click the data node (circle) and select **Data summary report** from the task menu.

<figure><img src="/files/48ZG3K8No6Csu2MXhhwQ" alt=""><figcaption><p><em>Select the data node (circle) and click Data summary report to show information for all upstream tasks</em></p></figcaption></figure>

[Please click here for more information on the Data summary report](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/data-summary-report)

### View all tasks performed on the project

The **Log** tab is used to view all tasks performed within the project. It contains a table of the tasks that are running, scheduled, or those that have been completed within the Partek Flow. It provides an overview of the task progress, enables task management, and links to detailed reports for each task.

<figure><img src="/files/p2Eb44HB4ugTZnu9JCUs" alt=""><figcaption><p><em>The Log tab is used to view all tasks within the project and enables task management including task progress and links to detailed reports for each task</em></p></figcaption></figure>

[Please click here for information about the Log tab.](https://help.partek.illumina.com/partek-flow/tutorials/creating-and-analyzing-a-project/the-log-tab)

## Additional navigation

### Project settings tab

To modify the name of the project and other details like adding collaborators, navigate to the **Project settings** tab.

<figure><img src="/files/3YqGJCwcZQuqD5MCZNMH" alt=""><figcaption><p><em>Navigate to the Project settings tab to modify the name of the project and other details like adding collaborators</em></p></figcaption></figure>

[Please click here for information about the Project settings tab.](https://help.partek.illumina.com/partek-flow/tutorials/creating-and-analyzing-a-project/the-project-settings-tab)

### Task node actions

Task node (rectangle) actions on the task menu (toolbox) allow you to rerun tasks, rerun with downstream tasks, edit the description, and change the color.

<figure><img src="/files/3NSewLqLE8rd5L7bkKDH" alt=""><figcaption><p><em>Click the task node (rectangle) to perform task actions in the toolbox (task menu)</em></p></figcaption></figure>

[Please click here for more information on Task actions.](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/task-actions)

### Pipelines

Pipelines can be created for future use. Select **Create new pipeline** near the bottom left-hand side of the browser window. [Please click here for a tutorial which covers saving and running a pipeline.](https://help.partek.illumina.com/partek-flow/tutorials/bulk-rna-seq/saving-and-running-a-pipeline)

<figure><img src="/files/jcezn9VLBvFZ8fFVFRlt" alt=""><figcaption><p><em>Pipelines can be created for future use by clicking Create new pipeline</em></p></figcaption></figure>

[Please click here for more information about Pipeline management.](https://help.partek.illumina.com/partek-flow/user-manual/settings/components/pipeline-management)

## Access tutorial data from the interface

Partek Flow also provides demo data within the interface so you can automatically begin a project with this data type for bulk RNA-Seq or single cell RNA-Seq and follow along with the tutorials listed below. Navigate to your avatar in the top right corner and click **Settings** then click the data type of interest.

* [Bulk RNA Seq](https://help.partek.illumina.com/partek-flow/tutorials/bulk-rna-seq) (RNA-Seq 5-AZA)
* [Single cell RNA-Seq ](https://help.partek.illumina.com/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples)(Single cell glioma (multi-sample))

<figure><img src="/files/f3rSQmRCu2Xf1kcJ2BID" alt=""><figcaption><p><em>Navigate to the tutorial data by clicking your Avatar, Settings, then the data of interest</em></p></figcaption></figure>

## Partek Flow in action

Watch [live training event recordings for different types of data.](https://help.partek.illumina.com/partek-flow/live-training-event-recordings)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Getting Started with Your Partek Flow Hosted Trial

Ready to start work on your Partek Flow hosted trial? This page has some helpful videos to get you started!

* [Uploading Your Data to a Hosted Instance of Partek Flow](#uploading-your-data-to-a-hosted-instance-of-partek-flow)
* [How to Get Started on your First Project](#how-to-get-started-on-your-first-project)

Don't have a trial server yet? Request one on [our website](http://www.partek.com/free-trial).

## Uploading Your Data to a Hosted Instance of Partek Flow

This short video shows you how to import your data into a hosted instance of Partek Flow. Adjust your device's volume for optimal sound.

**Note: When upload large size of data, it might take a while, please turn off the computer sleep mode settings!**

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2F8bKXMyRQNpyWjfY4mkJi%2FTransferFiles.mp4?alt=media&token=21c8a85c-07b8-4c1d-9545-5bc568e1dc4b>" %}

## How to Get Started on your First Project

In this short video, we'll give you an overview of the interface and how to get started with your analysis.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2FQrpbCoOQqr0ID1rhWRvg%2FGetting_started_2.mp4?alt=media&token=a752139a-da52-4879-b5b8-c2311fec518d>" %}

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Installation Guide

Partek Flow is a web-based application for genomic data analysis and visualization. It can be installed on a desktop computer, computer cluster or cloud. Users can then access Partek Flow from any browser-enabled device, such as a personal computer, tablet or smartphone.

Read on to learn about the following installation topics:

* [Minimum System Requirements](/partek-flow/installation-guide/minimum-system-requirements)
* [Single Cell Toolkit System Requirements](/partek-flow/installation-guide/single-cell-toolkit-system-requirements)
* [Single Node Installation](/partek-flow/installation-guide/single-node-installation)
* [Single Node Amazon Web Services‎ Deployment](/partek-flow/installation-guide/single-node-amazon-web-services-deployment)
* [Multi-Node Cluster Installation](/partek-flow/installation-guide/multi-node-cluster-installation)
* [Creating Restricted User Folders within the Partek Flow Server](/partek-flow/installation-guide/creating-restricted-user-folders-within-the-partek-flow-server)
* [Updating Partek Flow](/partek-flow/installation-guide/updating-partek-flow)
* [Uninstalling Partek Flow](/partek-flow/installation-guide/uninstalling-partek-flow)
* [Dependencies](/partek-flow/installation-guide/dependencies)
* [Docker and Docker-compose](/partek-flow/installation-guide/docker-and-docker-compose)
* [Java KeyStore and Certificates](/partek-flow/installation-guide/java-keystore-and-certificates)
* [Kubernetes](/partek-flow/installation-guide/kubernetes)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Minimum System Requirements

## Web Browser Requirements

Regardless of whether Partek Flow is installed on a server or the cloud, users will be interacting with the software using a web browser. We support the latest Google Chrome, Mozilla Firefox, Microsoft Edge and Apple Safari browsers. While we make an effort to ensure that Partek Flow is robust, note that some browser plugins may affect the way the software is viewed on your browser.

* [Web Browser Requirements](#web-browser-requirements)
* [Hardware Requirements (Single-node Linux)](#hardware-requirements-single-node-linux)
* [Hardware Requirements (Cluster or Cloud)](#hardware-requirements-cluster-or-cloud)
* [Storage Recommendations](#storage-recommendations)

## Hardware Requirements (Single-node Linux)

If you are installing Partek Flow on your own single-node server, we require the following for successful installation:

* **Linux:** Ubuntu® 18.04, Redhat® 8, CentOS® 8 or later versions of these distributions
* 64-bit 2GHz quad-core processor1
* 48GB of RAM2
* \> 2TB of storage available for data
* \> 100GB on the root partition
* A broadband internet connection

We support Docker-based installations. Please contact <support@partek.com> for more information.

1Note that some analyses have higher system requirements for example to run the STAR aligner on a reference genome of size \~3 GB (such as human, mouse or rat), 16 cores are required.

2Input sample file size can also impact memory usage, which is particularly the case for TopHat alignments.

Increasing hardware resources (cores, RAM, disk space, and speed) will allow for faster processing of more samples.

If you are licensed for the Single Cell Toolkit, please see Single Cell Toolkit System Requirements for amended hardware requirements.

## Hardware Requirements (Cluster or Cloud)

Please contact [Partek Technical Support](http://www.partek.com/support) if you would like to install Partek Flow on your own HPC or cloud account. We will assist in assessing your hardware needs and can make recommendations regarding provisioning sufficient resources to run the software.

## Storage Recommendations

Proper storage planning is necessary to avoid future frustration and costly maintenance. Here are several DO's and DO NOT's:

DO:

* Plan for at least 3 to 5 times more storage than you think is necessary. Investing extra funds in storage and storage performance is always worth it.
* Keep all Flow data on a single partition that is expandable, such as RAID or LVM.
* Back up your data, especially the Partek Flow database.

DO NOT:

* Store data on 'removable' USB drives. Partek Flow will not be able to see these drives.
* Store data across multiple partitions or folder locations. This will increase the maintenance burden substantially.
* Use non-Linux file systems like NTFS.

## Additional Assistance

If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.


# Single Cell Toolkit System Requirements

* [Up to 100,000 cells per analysis](#up-to-100000-cells-per-analysis)
* [More than 100,000 cells per analysis](#more-than-100000-cells-per-analysis)

Because of the large size of single cell RNA-Seq data sets and the computationally-intensive tools used in single cell analysis, we have amended our system requirements and recommendations for installations of Partek Flow with the Single Cell toolkit.

## Up to 100,000 cells per analysis

#### Required

* Linux: Ubuntu® 18.04, Redhat® 8, CentOS® 8, or newer
* CPU: 64-bit 2 GHz quad-core processor
* Memory: 64 GB of RAM
* Local scratch space\*: 1 TB with cached or native speeds of 2GB/s or higher
* Storage: > 2 TB available for data and > 100 GB on the root partition

#### Recommended

* Linux: Ubuntu® 18.04, Redhat® 8, CentOS® 8, or newer
* CPU: 64-bit 2 GHz quad-core processor
* Memory: 128 GB of RAM
* Local scratch space1: 2 TB with cached or native speeds of 2GB/s or higher
* Storage: > 2 TB available for data and > 100 GB on the root partition

## More than 100,000 cells per analysis

#### Required

* Linux: Ubuntu® 18.04, Redhat® 8, CentOS® 8, or newer
* CPU: 64-bit 2 GHz quad-core processor
* Memory: 256 GB of RAM
* Local scratch space1: 2 TB with cached or native speeds of 2GB/s or higher
* Storage: > 4 TB available for data

#### Recommended

* Linux: Ubuntu® 18.04, Redhat® 8, CentOS® 8, or newer
* CPU: 64-bit 2 GHz quad-core processor
* Memory: 512 GB of RAM
* Local scratch space1: 10 TB with cached or native speeds of 2GB/s or higher
* Storage: 10 TB available for data

For fastest performance:

* Newer generation CPU cores with avx2 or avx-512 are recommended.
* Performance scales proportionality to the number of CPU cores available.
* Hyper thread cores (threads) scales performance for most operations other than principal component analysis.

*\*Contact Partek support for recommended setup of local scratch storage*

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Single Node Installation

This document describes how to set up and configure a **single-node** Partek Flow license.

* [Docker Compose Deployment](/partek-flow/installation-guide/docker-and-docker-compose)

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Installing on Linux

This guide is specific to the **single-node** Partek Flow installation and uses the **Linux** package managers and repositories.

If you had installed older versions of Partek Flow using a zip file (run the flowstatus.sh to verify), follow the steps to switch to the package manager before proceeding.

This section covers the following topics:

* [Default Installation Directory](#default-installation-directory)
* [Checking your Linux Distribution](#checking-your-linux-distribution)
* [Installation on Debian/Ubuntu Distributions](#installation-on-debianubuntu-distributions)
* [Installation on RedHat/CentOS Distributions](#installation-on-redhatcentos-distributions)
* [End User Tools](#end-user-tools)
* [Access Partek Flow on a Web Browser](#access-partek-flow-on-a-web-browser)
* [Enter the License Key](#enter-the-license-key)
* [Create an Administrator Account](#create-an-administrator-account)
* [Select Library File Directory](#select-library-file-directory)

## Default Installation Directory

By default, Partek Flow is installed under /opt/partek\_flow and temporary files are housed in /opt/partek\_flow/temp.

## Checking your Linux Distribution

The installation procedure varies per Linux distribution. To check your distribution, open a terminal and run:

```
$ cat /etc/issue
```

## Installation on Debian/Ubuntu Distributions

1. Add the public key for the Partek package repository\*

```
$ sudo apt-key adv --keyserver keyserver.ubuntu.com --recv-keys C82B61BF
```

2. Add the Partek package list to your repository\*

```
$ sudo wget -P /etc/apt/sources.list.d/ http://packages.partek.com/debian/partek-flow.list
```

\*Steps 1 and 2 only need to be performed once prior to the first installation. Re-installation and updates do not require this step.

3. Update the list of available packages

```
$ sudo apt-get update
```

4. Install Partek Flow

```
$ sudo apt-get install partekflow
```

When asked to continue, type the letter **Y** and press **Enter**.

During the installation, you will be prompted for the Flow server port (Figure 1). Unless necessary, accept the default HTTP port: 8080 by pressing **Enter**.

<figure><img src="/files/eCgx1WDknPRMKjnB3GgJ" alt=""><figcaption><p><em>Figure 1. Configuring HTTP port for Partek Flow during installation</em></p></figcaption></figure>

5. If additional configuration is needed, use the **reconfigure** command below. This can be run any time after Partek Flow is installed. For details regarding each setting, contact the Partek Licensing Department. In most cases, this step can be skipped.

```
$ sudo dpkg-reconfigure partekflow
```

6. (Optional) To manually configure Partek Flow or to set additional advanced options or environment variables, edit the following configuration file:

```
/etc/partekflow.conf
```

Note that all changes made by the reconfigure command in step 5 are also stored in this configuration file.

7. Start the Partek Flow server

```
$ sudo service partekflowd start
```

A message should indicate that Partek Flow is now running: Starting Partek Flow server: OK

Step 7 needs to be performed only once after installation. Partek Flow will start automatically whenever the server restarts.

## Installation on RedHat/CentOS Distributions

1. Retrieve the Partek yum repo configuration

```
$ sudo wget -P /etc/yum.repos.d http://packages.partek.com/redhat/stable/partekflow.repo
```

Step 1 needs to be performed only once prior to the first installation. Re-installation and updates do not require this step.

2. Install Partek Flow

```
$ sudo yum install partekflow
```

3. When asked to continue, type the letter **Y** and press **Enter**
4. (Optional) To manually configure Partek Flow or to set additional advanced options or environment variables, edit the file located at:

```
/etc/partekflow.conf
```

5. Start the Partek Flow server

```
$ sudo service partekflowd restart
```

The following message indicates that Partek Flow is now running:Starting Partek Flow server: OK\
Step 5 needs to be performed only once after installation. Partek Flow will start automatically whenever a server restarts.

## End User Tools

A user can access Partek Flow using a web browser on any browser-enabled device, such as a personal computer, tablet, smartphone etc. We recommend using Google Chrome™ or, alternatively, Mozilla Firefox™. The screen resolution should be set to 1024 × 768 pixels or higher. This is particularly important for the use of visualization tools such as Chromosome Viewer.

### Access Partek Flow on a Web Browser

Once Partek Flow server has been started, access the interface using a web browser.

* If you are on the computer running Partek Flow, go to <http://localhost:8080/flow>\
  Note that if the Flow server port was assigned a different number during installation, replace 8080 with the correct port.
* If you are on a computer other than the Partek Flow server computer, localhost should be replaced with the IP address of the Partek Flow server computer

## Enter the License Key

When Partek Flow is launched for the first time, the user is prompted to provide a license key (Figure 2). **Copy** and **paste** the license key received from Partek Licensing Support in the License key box and select **Next**\*.

\*You may need to refresh your page if you are not automatically directed to the "Create an administrator account" after selecting "Next".

If you have not received the license key, contact your account representative or [request a trial](http://www.partek.com/free-trial/).

<figure><img src="/files/ezX6t8f17vEWGpDvMFfk" alt=""><figcaption><p><em>Figure 2. Setting up the Partek Flow license during installation</em></p></figcaption></figure>

## Create an Administrator Account

Partek Flow supports multiple users, each of which can either be classified as administrator or regular user, based on access privileges. The first account to be created is always an administrator account. Additional users may be added after installation. To set up the first (administrator) account, specify a username, password, and email address (Figure 3), then click **Next**.

<figure><img src="/files/IllwJZ3cfbjPwIWSbnVB" alt=""><figcaption><p><em>Figure 3. Setting up the Partek Flow 'admin' account during installation</em></p></figcaption></figure>

## Select Library File Directory

Select a directory folder to store the library files that will be downloaded or generated by Partek Flow (Figure 4). All Partek Flow users share library files and the size of the library folder can grow significantly. We recommend at least 100GB of free space should be allocated for library files. The free space in the selected library file directory is shown. Click **Next** to proceed. You can change this directory after installation by changing system preferences. For more information, see Library file management.

<figure><img src="/files/tjexuiP6erxyTux16uI8" alt=""><figcaption><p><em>Figure 4. Selecting the library file directory</em></p></figcaption></figure>

After selecting the library file directory, the installation is done. Click the **Finish** button (Figure 5).

<figure><img src="/files/5GNI0wt89FjWPdGVC4e6" alt=""><figcaption><p><em>Figure 5. Completing Partek Flow installation</em></p></figcaption></figure>

The browser will display the Partek Flow Homepage (Figure 6). Since there are no projects available, links to the download tutorial data as well as the link to the documentation pages are displayed.

<figure><img src="/files/tmqLEcGZNkOnrARXMhHV" alt=""><figcaption><p><em>Figure 6. The Partek Flow homepage</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.


# Single Node Amazon Web Services Deployment

* [Creating a New Elastic Compute Cloud Instance for Partek Flow Software](#creating-a-new-elastic-compute-cloud-instance-for-partek-flow-software)
* [Enabling External Access to the Partek Flow Elastic Compute Cloud Instance)](#enabling-external-access-to-the-partek-flow-elastic-compute-cloud-instance)
* [Attaching the Amazon Elastic Block Store Volume for Partek Flow Data Storage)](#attaching-the-amazon-elastic-block-store-volume-for-partek-flow-data-storage)
* [Installing Partek Flow on a New Elastic Compute Cloud Instance](#installing-partek-flow-on-a-new-elastic-compute-cloud-instance)
* [Partek Amazon Web Services Support](#partek-amazon-web-services-support)
* [General Recommendations](#general-recommendations)
* [Amazon Web Services Instance Type Resources and Costs](#amazon-web-services-instance-type-resources-and-costs)
* [Elastic Block Store Volumes](#elastic-block-store-volumes)

## Creating a New Elastic Compute Cloud Instance for Partek Flow Software

Note: This guide assumes all items necessary for the Amazon elastic Comput Clout (EC2) instance does not exist, such as Amazon Virtual Private Cloud (VPC), subnets, and security groups, thus their creation is covered as well.

Log in to the Amazon Web Services (AWS) management console at <https://console.aws.amazon.com>

Click on **EC2**

Switch to the region intended to deploy Partek Flow software. This tutorial uses US East (N. Virginia) as an example.

On the left menu, click on **Instances**, then click the **Launch Instance** button. The **Choose an Amazon Machine Image (AMI)** page will appear.

Click the Select button next to **Ubuntu Server 16.04 LTS (HVM), SSD Volume Type - ami-f4cc1de2**. *NOTE:* Please use the latest Ubuntu AMI. It is likely that the AMI listed here will be out of date.

**Choose an Instance Type**, the selection depends on your budget and the size of the Partek Flow deployment. We recommend **m4.large** for testing or cluster front-end operation, **m4.xlarge** for standard deployments, and **m4.2xlarge** for alignment-heavy workloads with a large user-base. See the section AWS instance type resources and costs for assistance with choosing the right instance. In most cases, the instance type and associated resources can be changed after deployment, so one is not locked into the choices made for this step.

*NOTE:* New instance types will become available. Please use the latest mX instance type provided as it will likely perform better and be more cost effective than older instance types.

On the **Configure Instance Details** page, make the following selections:

* Set the **number of instances** to 1. An **autoscaling group** is not necessary for single-node deployments
* **Purchasing Option:** Leave **Request Spot Instances** unchecked. This is relevant for cost-minimization of Partek Flow cluster deployments.
* **Network:** If you do not have a virtual private cloud (VPC) already created for Partek Flow, click **Create New VPC**. This will open a new browser tab for VPC management.
  * Use the following settings for the VPC:
    * **Name Tag:** *Flow-VPC*
    * **IPv4 CIDR block:** *10.0.0.0/16*
    * Select **No IPv6 CIDR Block**
    * **Tenancy:** *Default*
  * Click **Yes, Create**. You may be asked to select a **DHCP Option set**. If so, then make sure the dynamic host configuration protocol (DHCP) option set has the following properties:
    * **Options:** *domain-name = ec2.internal;domain-name-servers = AmazonProvidedDNS;*
    * **DNS Resolution:** leave the defaults set to *yes*
    * **DNS Hostname:** change this to yes as internal DNS resolution may be necessary depending on the Partek Flow deployment
  * Once created, the new Flow-VPC will appear in the list of available VPCs. The VPC needs additional configuration for external access. To continue, right click on Flow-VPC and select **Edit DNS Resolution**, select Yes, and then Save. Next, right click the Flow-VPC and select **Edit DNS Hostnames**, select Yes, then Save.
  * Make sure the DHCP option set is set to the one created above. If it is not, right-click on the row containing Flow-VPC and select **Edit DHCP Option Sets**.
  * Close the **VPC Management** tab and go back to the **EC2 Management Console**.
* Click the refresh arrow next to **Create New VPC** and select *Flow-VPC*.
* Click **Create New Subnet** and a new browser tab will open with a list of existing subnets. Click **Create Subnet** and set the following options:
  * **Name Tag:** *Flow-Subnet*
  * **VPC:** *Flow-VPC*
  * **VPC CIDRs:** This should be automatically populated with the information from Flow-VPC
  * **Availability Zone:** It is OK to let Amazon choose for you if you do not have a preference
  * **IPv4 CIDR block:** *10.0.1.0/24*
* Stay on the **VPC Dashboard Tab** and on the left navigation menu, click **Internet Gateways**, then click **Create Internet Gateway** and use the following options:
  * **Name Tag:** *Flow-IGW*
  * Click **Yes, Create**
* The new gateway will be displayed as **Detached**. Right click on the Flow-IGW gateway and select **Attach to VPC**, then select Flow-VPC and click **Yes, Attach**.
* Click on **Route Tables** on the left navigation menu.
* If it exists, select the route table already associated with *Flow-VPC*. If not, make a new route table and associate it with *Flow-VPC*. Click on the new route table, then click the **Routes** tab toward the bottom of the page. The route *Destination = 10.0.0.0/16 Target = local* should already be present. Click **Edit**, then **Click Add another route** and set the following parameters:
  * **Destination:** *0.0.0.0/0*
  * **Target** set to *Flow-IGW* (the internet gateway that was just created)
* Click **Save**
* Close the **VPC Dashboard** browser tab and go back to the **EC2 Management Console** tab. Note that you should still be on **Step 3: Configure Instance Details**.

Click the refresh arrow next to Create New Subnet and select Flow-Subnet.

**Auto-assign Public IP:** Use subnet setting (Disable)

**Placement Group:** No placement group

**IAM role:** None.

*Note: For multi-node Partek Flow deployments or instances where you would like Partek to manage AWS resources on*\
\&#xNAN;*your behalf, please see Partek AWS support and set up an IAM role for your Partek Flow EC2 instance. In most cases*\
\&#xNAN;*a specialized IAM role is unnecessary and we only need instance ssh keys.*

**Shutdown Behaviour:** *Stop*

**Enable Termination Protection:** select *Protect against accidental termination*

**Monitoring:** leave **Enable CloudWatch Detailed Monitoring** disabled

**EBS-optimized Instance:** Make sure **Launch as EBS-optimized Instance** is enabled. Given the recommended choice of an m4 instance type, EBS optimization should be enabled at no extra cost.

**Tenancy:** *Shared - Run a shared hardware instance*

**Network Interfaces:** leave as-is

**Advanced Details:** leave as-is

Click **Next: Add Storage**. You should be on **Step 4: Add Storage**

For the existing root volume, set the following options:

* **Size:** *8 GB*
* **Volume Type:** *Magnetic*
* Select **Delete on Termination**
  * Note: All Partek Flow data is stored on a non-root EBS volume. Since only the OS is on the root volume and not frequently re-booted, a fast root volume is probably not necessary or worth the cost. For more information about EBS volumes and their performance, see the section EBS volumes.

Click Add New Volume and set the following options:

* **Volume Type:** *EBS*
* **Device:** */dev/sdb* (take the default)
* Do not define a snapshot
* **Size (GiB):** 500
  * Note: This is the minimum for ST1 volumes, see: EBS volumes
* **Volume Type:** *Throughput optimized HDD (ST1)*
* Do not delete on terminate or encrypt

Click **Next: Add Tags**

* You do not need to define any tags for this new EC2 instance, but you can if you would like.

Click **Next: Configure Security Group**

* For **Assign a Security Group** select **Create a New Security Group**
* **Security Group Name:** *Flow-SG*
* **Description:** *Security group for Partek Flow server*
* Add the following rules:
  * SSH set **Source** to *My IP* (or the address range of your company or institution)
  * Click **Add Rule:**
  * Set **Type** to *Custom TCP Rule*
  * Set **Port Range** to *8080*
  * Set **Source** to anywhere *(0.0.0.0/0, ::/0)*
    * Note: It is recommended to restrict **Source** to just those that need access to Partek Flow.

Click **Review and Launch**

* The AWS console will suggest this server not be booted from a magnetic volume. Since there is not a lot of IO on the root partition and reboots are will be rare, choosing **Continue with Magnetic** will reduce costs. Choosing an SSD volume will not provide substantial benefit but it OK if one wishes to use an SSD volume. See the EBS Volumes section for more information.

Click **Launch**

Create a new keypair:

* Name the keypair Flow-Key
* Download this keypair, the run **chmod 600 Flow-Key**.pem (the downloaded key) so it can be used.
* Backup this key as one may lose access to the Partek Flow instance without it.

The new instance will now boot. Use the left navigation bar and click on **Instances**. Click the pencil icon and assign the instance the name Partek Flow Server

## Enabling External Access to the Partek Flow Elastic Compute Cloud Instance

The server should be assigned a fixed IP address. To do this, click on **Elastic IPs** on the left navigation menu from the **EC2 Management Console**.

* Click **Allocate New Address**
* Assign **Scope** to *VPC*
* Click **Allocate**

On the table containing the newly allocated elastic IP, right click and select Associate Address

* For **Instance**, select the instance name *Flow Test Server*
* For **Private IP**, select the one private IP available for the Partek Flow EC2 instance, then click **Associate**

Note: For the remaining steps, we refer to the elastic ip as elastic.ip

SSH to the new Flow-Server instance:

```
$ chmod 600 Flow-Key.pem
```

```
$ ssh -i Flow-Testing.pem ubuntu@elastic.ip
```

## Attaching the Amazon Elastic Block Store Volume for Partek Flow Data Storage

Attach, format, and move the ubuntu home directory onto the large ST1 elastic block store (EBS) volume. All Partek Flow data will live in this volume. Consult the AWS EC2 documentation for further information about attaching EBS volumes to your instance.

```
$ sudo su
```

```
$ mkfs -t ext4 /dev/xvdb
```

Note: Under Volumes in the EC2 management console, inspect Attachment Information. It will likely list the large ST1 EBS volume as attached to /dev/sdb. Replace "s" with "xv" to find the device name to use for this mkfs command.

Make a note of the newly created UUID for this volume

Copy the *ubuntu* home directory onto the EBS volume using a temporary mount point:

```
$ mount -t ext4 /dev/xvdb /mnt/
```

```
$ rsync -avr /home/ /mnt/
```

```
$ umount /mnt/
```

Make the EBS volume mount at system boot:

Add the following to /etc/fstab: *UUID=the-UUID-from-the-mkfs-command-above /home ext4 defaults,nofail 0 2*

```
$ mount -a
```

Disconnect the ssh session, then log in again to make sure all is well

## Installing Partek Flow on a New Elastic Compute Cloud Instance

Note: For additional information about Partek Flow installations, see our generic Installation Guide

Before beginning, send the media access control (MAC) address of the EC2 instance to MAC address of the EC2 instance to <licensing@partek.com>. The output of ifconfig will suffice. Given this information, Partek employees will create a license for your AWS server. MAC addresses will remain the same after stopping and starting the Partek Flow EC2 instance. If the MAC address does change, let our licensing department know and we can add your license to our floating license server or suggest other workarounds.

Install required packages for Partek Flow:

```
$ sudo apt-get update
```

```
$ sudo apt-get install software-properties-common
```

```
$ sudo add-apt-repository -y ppa:openjdk-r/ppa
```

```
$ sudo apt-get install openjdk-8-jdk python python-pip python-dev zlib1g-dev python-matplotlib r-base python-htseq libxml2-dev perl make gcc g++ zlib1g libbz2-1.0 libstdc++6 libgcc1 libncurses5 libsqlite3-0 libfreetype6 libpng12-0 zip unzip libgomp1 libxrender1 libxtst6 libxi6 debconf 
```

```
$ sudo pip install --upgrade pip && pip install --upgrade --upgrade-strategy eager --force-reinstall virtualenv numpy pysam cnvkit
```

Install Partek Flow:

Note: Make sure you are running as the ubuntu user.

```
$ cd (we will install Partek Flow to ubuntu's home directory)
```

```
$ wget --content-disposition packages.partek.com/linux/flow-release
```

```
$ unzip PartekFlow*.zip
```

```
$ ./partek_flow/start_flow.sh
```

Partek Flow has finished loading when you see *INFO: Server startup in xxxxxxx ms in the partek\_flow/logs/catalina.out* log file. This takes \~30 seconds.

Alternative: Install Flow with Docker. Our base packages are located here: [https://hub.docker.com/r/partekinc/flow/tags](https://hub.docker.com/r/partekinc/flow/tags/)

Open Partek Flow with a web browser: [http://elastic.ip:8080/](http://awstest.partek.com:8080/)

Enter license key

Set up the Partek Flow admin account

Leave the library file directory at its default location and check that the free space listed for this directory is consistent with what was allocated for the ST1 EBS volume.

Done! Partek Flow is ready to use.

## Partek Amazon Web Services Support

After the EC2 instance is provisioned, we are happy to assist with setting up Partek Flow or address other issues you encounter with the usage of Partek Flow. The quickest way to receive help is to allow us remote access to your server by sending us Flow-Key.pem and amending the SSH rule for Flow-SG to include access from IP 97.84.41.194 (Partek HQ). We recommend sending us the Flow-Key.pem via secure means. The easiest way to do this is with the following command:

```
$ curl -F "file=@FlowKey.pem" https://installfeedback.partek.com/fupload
```

We also provide live assistance via GoTo meeting or TeamViewer if you are uncomfortable with us accessing your EC2 instance directly. Before contacting us, please run $ ./partek\_flow/flowstatus.sh to send us logs and other information that will assist with your support request.

## General Recommendations

With newer EC2 instance types, it is possible to change the instance type of an already deployed Partek Flow EC2 server. We recommend doing several rounds of benchmarks with production-sized workloads and evaluate if the resources allocated to your Partek Flow server are sufficient. You may find that reducing resources allocated to the Partek Flow server may come with significant cost savings, but can cause UI responsiveness and job run-times to reach unacceptable levels. Once you have found an instance type that works, you may wish to use reserved instance pricing which is significantly cheaper than on-demand instance pricing. Reserved instances come with 1 or 3-year usage terms. Please see the [EC2 Reserved Instance Marketplace](https://aws.amazon.com/ec2/purchasing-options/reserved-instances/marketplace/) to sell or purchase existing reserved instances at reduced rates.

The network performance of the EC2 instance type becomes an important factor if the primary usage of Partek Flow is for alignment. For this use case, one will have to move copious amounts of data back (input fastq files) and forth (output bam files) between the Partek Flow server and the end users, thus it is important to have as what AWS refers to as high network performance which for most cases is around 1 Gb/s. If the focus is primarily on downstream analysis and visualization (e.g. the primary input files are ADAT) then network performance is less of a concern.

We recommend HVM virtualization as we have not seen any performance impact from using them and non-HVM instance types can come with significant deployment barriers.

Make sure your instance is EBS optimized by default and you are not charged a surcharge for EBS optimization.

T-class servers, although cheap, may slow responsiveness for the Partek Flow server and generally do not provide sufficient resources.

We do not recommend placing any data on instance store volumes since all data is lost on those volumes after an instance stops. This is too risky as there are cases where user tasks can take up unexpected amounts of memory forcing a server stop/reboot.

## Amazon Web Services Instance Type Resources and Costs

The values below were updated April 2017. The latest pricing and EC2 resource offerings can be found at [http://www.ec2instances.info](http://www.ec2instances.info/)

| Instance Type | Memory   | Cores   | EBS throughput      | Network Performance   | Monthly cost |
| ------------- | -------- | ------- | ------------------- | --------------------- | ------------ |
| m4.large      | 8.0 GB   | 2 vCPUs | 56.25 MB/s M        | Medium                | $78.840      |
| r4.large      | 15.25 GB | 2 vCPUs | 50 MB/s H(10G int)  | High (+10G interface) | $97.09       |
| m4.xlarge     | 16.0 GB  | 4 vCPUs | 93.75 MB/s H        | High                  | $156.950     |
| r4.xlarge     | 30.5 GB  | 4 vCPUs | 100 MB/s H          | High                  | $194.180     |
| m4.2xlarge    | 32.0 GB  | 8 vCPUs | 125 MB/s H          | High                  | $314.630     |
| r4.2xlarge    | 61.0 GB  | 8 vCPUs | 200 MB/s H(10G int) | High (+10G interface) | $388.360     |

Single server recommendation: **m4.xlarge** or **m4.2xlarge**

[Network performance values](http://epamcloud.blogspot.com.br/2013/03/testing-amazon-ec2-network-speed.html) for US-EAST-1 correspond to: Low \~ 50Mb/s, Medium \~ 300Mb/s, High \~ 1Gb/s.

## Elastic Block Store Volumes

Choice of a volume type and size:

This is dependent on the type of workload. For must users, the Partek Flow server tasks are alignment-heavy so we recommend a throughput optimized HDD (ST1) EBS volume since most aligner operations are sequential in nature. For workloads that focus primarily on downstream analysis, a general purpose SSD volume will suffice but the costs are greater. For those who focus on alignment or host several users, the storage requirements can be high. ST1 EBS volumes have the following characteristics:

Max throughput 500 MiB/s

$0.045 per GB-month of provisioned storage ($22.5 per month for a 500 GB of storage).

Note that EBS volumes can be grown or performance characteristics changed. To minimize costs, start with a smaller EBS volume allocation of 0.5 - 2 TB as most mature Partek Flow installations generate roughly this amount of data. When necessary, the EBS volume and the [underlying file system can be grown on-line](http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-expand-volume.html#recognize-expanded-volume-linux) (making ext4 a good choice). Shrinking is also possible but may require the Partek Flow server to be offline.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Multi-Node Cluster Installation

Partek Flow is a genomics data analysis and visualization software product designed to run on compute clusters. The following instructions assume the most basic setup of Partek Flow and must only be attempted by system administrators who are familiar with Linux-based commands. These instructions are not intended to be comprehensive. Cluster environments are largely variable, thus there are no 'one size fits all' instructions. The installation procedure on a computer cluster is highly dependent on the type of computer cluster and the environment it is located. We can to support a large array of Linux distributions and configurations. In all cases, Partek Technical Support will be available to assist with cluster installation and maintenance to ensure compatibility with any cluster environment. Please consult with Partek Licensing Support (<licensing@partek.com>) for additional information.

Prior to installation, make sure you have the license key related to the host-ID of the compute cluster the software will be installed in. Contact <licensing@partek.com> for key generation.

* [Installation on a Computer Cluster](#installation-on-a-computer-cluster)
* [Integration with your queueing system](#integration-with-your-queueing-system)
* [Bringing up workers](#bringing-up-workers)
* [Shutting down workers](#shutting-down-workers)
* [Updating Partek Flow](#updating-partek-flow)

## Installation on a Computer Cluster

Make a standard linux user account that will run the Partek Flow server and all associated processes. It is assumed this account is synced between the cluster head node and all compute nodes. For this guide, we name the account *flow*

1. Log into the flow account and proceed to the cd to the flow home directory

```
cd home/flow
```

2. Download Partek Flow and the remote worker package

```
wget --content-disposition http://packages.partek.com/linux/flow
```

```
wget --content-disposition http://packages.partek.com/linux/flow-worker
```

3. Unzip these files into the flow home directory */home/flow*. This yields two directories: *partek\_flow* and P\_artekFlowRemoteWorker\_
4. Partek Flow can generate large amounts of data, so it needs to be configured to the bulk of this data in the largest shared data store available. For this guide we assume that the directory is located at */shared*. Adjust this path accordingly.
5. It is required that the Partek Flow server (which is running on the head node) and remote workers (which is running on the compute nodes) see identical file system paths for any directory Partek Flow has read or write access to. Thus /shared and /home/flow must be mounted on the Flow server and all compute nodes. Create the directory */shared/FlowData* and allow the *flow* linux account write access to it
6. It is assumed the head node is attached to at least two separate networks: (1) a public network that allows users to log in to the head node and (2) a private backend network that is used for communication between compute nodes and the head node. Clients connect to the Flow web server on port *8080* so adjust the firewall to allow inbound connections to *8080* over the public network of the head node. Partek Flow will connect to remote workers over your private network on port *2552* and *8443*, so make sure those ports are open to the private network on the flow server and workers.
7. Partek Flow needs to be informed of what private network to use for communication between the server and workers. It is possible that there are several private networks available (gigabit, infiniband, etc.) so select one to use. We recommend using the fastest network available. For this guide, let's assume that private network is *10.1.0.0/16*. Locate the headnode hostname that resolves to an address on the *10.1.0.0/16* network. This must resolve to the same address on all compute nodes.
8. For example:

host ***head-node.local*** yields ***10.1.1.200***

Open /home/flow/.bashrc and add this as the last line:

```
export CATALINA_OPTS="$CATALINA_OPTS -Djava.awt.headless=true
-DflowDispatcher.flow.command.hostname=head-node.local
-DflowDispatcher.akka.remote.netty.tcp.hostname=head-node.local"
```

Source .bashrc so the environment variable ***CATALINA\_OPTS*** is accessible.

NOTE: If workers are unable to connect (below), then replace all hostnames with their respective IPs.

9. Start Partek Flow

```
~/partek_flow/start_flow.sh
```

10. You can monitor progress by tailing the log file partek\_flow/logs/catalina.out. After a few minutes, the server should be up.
11. Make sure the correct ports are bound

```
netstat -tulpn
```

12. You should see *10.1.1.200:2552* and *:::8080* as *LISTENing*. Inspect *catalina.out* for additional error messages.
13. Open a browser and go to [http://localhost:8080](http://localhost:8080/) on the head node to configure the Partek Flow server.
14. Enter the license key provided (Figure 1)

![Setting up the Partek Flow license during installation](/files/5jFUBVRILuQCgxFWn7dQ)\
\&#xNAN;*Figure 1. Setting up the Partek Flow license during installation*

15. If there appears to be an issue with the license or there is a message about 'no workers attached', then restart Partek Flow. It may take 30 sec for the process to shut down. Make sure the process is terminated before starting the server back up:

```
~/partek_flow/stop_flow.sh
```

Then run:

```
~/partek_flow/start_flow.sh
```

16. You will now be prompted to setup the Partek Flow admin user (Figure 2). Specify the username (admin), password and email address for the administrator account and click **Next**

![Setting up the Partek Flow admin account during installation](/files/IZCFPvjcFr7QvkDpmKu8)\
\&#xNAN;*Figure 2. Setting up the Partek Flow 'admin' account during installation*

17. Select a directory folder to store the library files that will be downloaded or generated by Partek Flow (Figure 3). All Partek Flow users share library files and the size of the library folder can grow significantly. We recommend at least 100GB of free space should be allocated for library files. The free space in the selected library file directory is shown. Click **Next** to proceed. You can change this directory after installation by changing system preferences. For more information, see Library file management.

![Selecting the library file directory](/files/CQkLl3stUlBPJyiO2eH9)\
\&#xNAN;*Figure 3. Selecting the library file directory*

18. To set up the Partek Flow data paths, click on **Settings** located on the top-right of the Flow server webpage. On the left, click on **Directory permissions** then **Permit access to a new directory**. Add */shared/PartekFlow* and allow all users access.
19. Next click on **System preferences** on the left menu and change **data download directory** and **default project output directory** to */shared/PartekFlow/downloads* and */shared/PartekFlow/project\_output* respectively

Note: If you do not see the ***/shared**folder* listed, click on the **Refresh folder list** link that is toward the bottom of the download directory dialog

20. Since you do not want to run any work on the head node, go to **Settings>System preferences>Task queue and job processing** and uncheck **Start internal worker at Partek Flow server startup**.
21. Restart the Flow server:

```
~/partek_flow/stop_flow.sh
```

After 30 seconds, run:

```
~/partek_flow/start_flow.sh
```

This is needed to disable the internal worker.

22. Test that remote workers can connect to the Flow server
23. Log in as the flow user to one of your compute nodes. Assume the hostname is compute-0. Since your home directory is exported to all compute nodes, you should be able to go to /home/flow/PartekFlowRemoteWorker/
24. To start the remote worker:

```
./partekFlowRemoteWorker.sh head-node.local compute-0
```

25. These two addresses should both be in the *10.1.0.0/16* address space. The remote worker will output to stdout when you run it. Scan for any errors. You should see the message woot! I'm online.
26. A successfully connected worker will show up on the *Resource management* page on the Partek Flow server. This can be reached from the main homepage or by clicking **Resource management** from the *Settings* page. Once you have confirmed the worker can connect, kill the remote worker (CTRL-C) from the terminal in which you started it.
27. Once everything is working, return to library file management and add the genomes/indices required by your research team. If Partek hosts these genomes/indices, these will automatically be downloaded by Partek Flow

## Integration with your queueing system

1. In effect, all you are doing is submitting the following command as a batch job to bring up remote workers:

```
/home/flow/PartekFlowRemoteWorker/partekFlowRemoteWorker.sh head-node.local compute-0
```

2. The second parameter for this script can be obtained automatically via:

```
$(hostname -s)
```

## Bringing up workers

Bring up workers by running the command below. You only need to run one worker per node:

```
/home/flow/PartekFlowRemoteWorker/partekFlowRemoteWorker.sh head-node.local compute-0
```

## Shutting down workers

Go to the *Resource management page* and click on the **Stop button (red square)** next to the worker you wish to shut down. The worker will shut down gracefully, as in it will wait for currently running work on that node to finish, then it will shut down.

## Updating Partek Flow

For the cluster update, you will get a link of .zip file for Partek Flow and remote Flow worker respectively from Partek support, all of the following actions should be performed as the Linux user that runs Flow. Do NOT run Flow as root.

1. Go to the Flow installation directory. This is usually the home directory of the Linux user that runs Flow and it should contain a directory named "**partek\_flow**". The location of the Flow install can also be obtained by running ***ps aux | grep flow*** and examining the path of the running Flow executable.
2. Shut down Flow:

```
./partek_flow/stop_flow.sh
```

3. Download the new version of Flow and the Flow worker:

```
wget --content-disposition http://packages.partek.com/linux/flow-release
```

```
wget --content-disposition http://packages.partek.com/linux/flow-worker-release
```

4. Make sure Flow has exited:

```
ps aux | grep flow
```

The flow process should no longer be listed.

5. Unpack the new version of Flow install and backup the old install:

```
mv partek_flow partek_flow_prev
```

```
mv PartekFlowRemoteWorker PartekFlowRemoteWorker_prev
```

6. Backup the Flow database folder. This should be located in the home directory of the user that runs Flow.

```
tar -czvf partek-db-bkp-date.tgz ~/.partekflow
```

7. Start the updated version of Flow:

```
./partek_flow/start_flow.sh
```

```
tail -f partek_flow/logs/catalina.out
```

(make sure there is nothing of concern in this file when starting up Flow. You can stop the file tailing by typing: CTRL-C)

You may also want to examine the the main Flow log for errors:

```
~/.partekflow/logs/flow.log
```

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Creating Restricted User Folders within the Partek Flow server

Partek Flow provides the infrastructure to isolate data from different users within the same server. This guide will provide general instructions on how to create this environment within Partek Flow. This can be modified to accommodate existing file systems already accessible to the server.

Go to **Settings > Directory** permissions and restrict parent folder access (typically */home/flow*) to *Administrator* accounts only

<figure><img src="/files/FzKRuA1tbGimCoZys29D" alt=""><figcaption><p><em>Figure 1. Setting directory permission for administrators</em></p></figcaption></figure>

Click the Permit access to a new directory button and navigate to the folder with your library files (typically /home/flow/FlowData/library\_files). Select the All users (automatically updated) checkbox to permit all users (including those that will be added in the future) to see the library files associated with the Partek Flow server

<figure><img src="/files/CJaLVlxnp6zJpkbmaSiG" alt=""><figcaption><p><em>Figure 2. Allow all users permission to see the library files</em></p></figcaption></figure>

Then go to System preferences > Filesystem and storage and set the Default project output directory to "Sample file directory"

<figure><img src="/files/pvqYl33IuhXT7YVpdLRY" alt=""><figcaption><p><em>Figure 3. Set default project output directory</em></p></figcaption></figure>

Create your first user and select the Private directory checkbox. Specify where the private directory for that user is located

<figure><img src="/files/vNIHGUIF9BcHf7frv7Xr" alt=""><figcaption><p><em>Figure 4. Adding a user with a private directory</em></p></figcaption></figure>

If needed, you can create a user directory by clicking Browse > Create new folder

<figure><img src="/files/5O36ZW1qDQQqbE0QZ4kF" alt=""><figcaption><p><em>Figure 5. Create a new private user folder</em></p></figcaption></figure>

This automatically sets browsing permissions for that private directory to that user

<figure><img src="/files/yPqnCB2JuuCDlIexRycG" alt=""><figcaption><p><em>Figure 6. Private directories automatically get restricted permissions</em></p></figcaption></figure>

When a user creates a project. The default project output directory is now within their own restricted folder

<figure><img src="/files/ST94k7ugD5XcxLdPZC1e" alt=""><figcaption><p><em>Figure 7. Project output directory will now be within private directory</em></p></figcaption></figure>

More importantly, other users cannot see them

<figure><img src="/files/5O36ZW1qDQQqbE0QZ4kF" alt=""><figcaption><p><em>Figure 8. Other users' directories are not visible</em></p></figcaption></figure>

Add additional users as needed

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Updating Partek Flow

Before performing updates, we recommend Backing Up the Database.

For tomcat build update, download the latest version from below:

```
wget --content-disposition packages.partek.com/linux/flow
```

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Uninstalling Partek Flow

## Linux

Open a terminal window and enter the following command.

Debian/Ubuntu:

```
$ sudo apt-get remove partekflow
```

RedHat/Fedora/CentOS:

```
$ sudo yum remove partekflow
```

The uninstall removes binaries only (*/opt/partek\_flow*). The logs, database (*partek\_db*) and files in the *home/flow/.partekflow* folder will remain unaffected.

## MacOS

1. Stop and quit Partek Flow using the Partek Flow app in the menu.
2. Using Finder, delete Flow application from the Applications menu.

**Missing image**\
\&#xNAN;*Figure 1. Control of Partek Flow through the menu bar*

This process does not delete data or the library files. Users who wish to delete those can delete them using Finder or terminal. The default location of project output files and library files is the /FlowData directory under the user's home folder. However, the actual location may vary depending on your System or Project specific settings.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Dependencies

* [Select Library File Directory](#select-library-file-directory)
* [CNVkit](#cnvkit)
* [DESeq2](#deseq2)
* [HTSeq](#htseq)
* [MACS3](#macS3)
* [Python](#python)
* [R](#r)
* [Variant Effect Predictor](#variant-effect-predictor)
* [DECoN](#decon)

Flow ships with tasks that do not have all of their dependencies included. On startup Flow will attempt to install the dependencies, but not every system is equipped to install them.

In the case of any difficulties, it is highly recommended to instead use a docker deployment (cluster installations may require singularity instead, which is somewhat still a work-in-progress)Z

## CNVkit

Requires Python 2.7 or later.

On startup Flow will attempt to install additional python packages using the command

```
pip install --user cnvkit==0.9.5
```

\
\
Requires R 3.2.3 or later.

On startup Flow will attempt to install additional R packages.

There are cascading dependencies, but you can view the core libraries in partek\_flow/bin/cnvkit-0.8.5/install.R

If these packages can't be built locally, it may be possible for the user to download them from us (see below).

## DESeq2

Requires R 3.0 or later.

On startup Flow will attempt to install additional R packages.

There are cascading dependencies, but you can view the core libraries in partek\_flow/bin/deseq\_two-3.5/install.R

If these packages can't be built locally, it may be possible for the user to download them from us (see below).

RcppArmadillo may also have dependencies on multi-threading shared objects that may not be on the LD\_LIBRARY\_PATH

The recommendation is to copy those .so files to a folder and make sure it is available from the LD\_LIBRARY\_PATH when the server/worker starts.

Additional dynamic libraries (such as libxml2.so) may be missing and we can provide a copy appropriate for the target OS.

## HTSeq

Requires Python 2.7 or 3.4 or above

On startup Flow attempts to install using pip

## MACS3

Requires python 3.0 or above

```
pip install --user numpy==1.19.5 Cython==0.29.30 cykhash==2.0.0 macs3==3.0.0a7
```

## Python

If there are any conflicts with preinstalled python packages, Flow should be configured to run with its own virtual environment:

```
pip install virtualenv
```

```
virtualenv ~/.partekflow/.local
```

```
source ~/.partekflow/.local/bin/activate
```

```
pip install HTSeq==0.11.0
```

```
pip install cnvkit==0.9.5
```

or

```
wget customer.partek.com/python-dependencies.zip
```

```
unizp -d ~/.partekflow/ python-dependencies.zip
```

## R

R can usually be installed from the package manager. If the user installs Flow via apt or yum it should already be installed.

For older operating systems R is not available and will need to be installed from [source](http://rweb.crmda.ku.edu/cran/src/base/R-3/)

Currently, we offer a set of R packages compatible with some versions of R

* [3.2.3](http://customer.partek.com/R_libs.3.2.3.zip)
* [3.4.0](http://customer.partek.com/R_libs-3.4.0.zip)
* [3.4.3](http://customer.partek.com/R_libs-3.4.3.zip)

Extract this file in the home directory. (Make .R a symlink if the home directory doesn't have enough free space)

These packages include the dependencies for both CNVkit and DESeq2

When running R diagnostic commands outside flow, it can simplify things if the environment includes a reference to the \~/.R folder:

```
export R_LIBS_USER=$HOME/.R
```

or load

```
.libPaths("~/.R")
```

in \~/.Rprofile

list loaded packages:

```
(.packages())
```

get the version:

```
packageVersion("packageName")
```

```
R_HOME=/path/to/R
```

## Variant Effect Predictor

This is a compiled Perl script (so it has no direct dependency on Perl itself) we have had one report (istem.fr) of it failing to run.

## DECoN

DECoN comes pre-installed in the flow\_dna container [registry.partek.com/flow\_dna](http://registry.partek.com/flow_dna)

Documentation on installing DECoN is available here: <https://github.com/RahmanTeam/DECoN/blob/master/DECoN-v1.0.2.pdf>

DECoN requires R version 3.1.2

It must be installed under /opt/R-3.1.2 or set the DECON\_R environment variable to its folder

```
wget http://cran.wustl.edu/src/base/R-3/R-3.1.2.tar.gz
```

```
tar xfz R-3.1.2.tar.gz  
```

```
cd R-3.1.2
```

```
./configure --with-x=no && make
```

Download DECoN

<https://github.com/RahmanTeam/DECoN/archive/refs/tags/v1.0.2.zip>

and install it under /opt/DECoN or set the DECON\_PATH environment variable to its folder

You may need to add

```
symlink.system.packages: TRUE
```

to Linux/packrat/packrat.opts

See also: [Minimum System Requirements](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/installation-guide/installation-guide/minimum-system-requirements.md)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Docker and Docker-compose

* [Docker](#docker)
* [Useful commands](#useful-commands)
* [Docker-compose](#docker-compose)

## Docker

Container-based deployments using Docker are recommended for single-node deployments. The container image for Partek Flow comes with all dependencies and third party utilities needed for all software features. This results in trivial initial deployments, upgrades, and maintenance.

Before proceeding, follow the [Docker documentation](https://docs.docker.com/engine/install/) to install Docker.

## Useful commands

```
docker ps
```

This command will output the details of currently running containers including port forwarding, container name/id, and uptime.

```
docker exec -it containername bash
```

This command will execute a shell in the container’s environment. This is useful for troubleshooting application or deployment issues.

## Docker-compose

“Compose is a tool for defining and running multi-container Docker applications. With Compose, you use a YAML file to configure your application’s services. Then, with a single command, you create and start all the services from your configuration.“

[Documentation](https://docs.docker.com/compose/reference/overview/)

Illumina support can assist customers with the creation of a docker-compose.yml file with all customer environment-specific configuration necessary to run Partek Flow on any server that meets our [Minimum system requirements](https://help.partek.illumina.com/partek-flow/installation-guide/minimum-system-requirements).

Below is an example docker-compose.yml file which after modification can be used to deploy Partek Flow. Such modifications are detailed in the description of important values following the example docker-compose.yml file.

```yaml
services:
  flowheadnode:
    restart: unless-stopped
    image: public.ecr.aws/partek-flow/rtw:$tag
    hostname: flowheadnode
    environment:
      # Set license location. Can be a server ex. @lic-test.example.com or /home/flow/.partekflow/license/Partek.lic
      - PARTEKLM_LICENSE_FILE=/home/flow/.partekflow/license/Partek.lic
    ports:
      - "8080:8080"
    # The MAC must match what is in the license file or defined on the license server. Please change this.
    networks:
      flexlm:
        mac_address: aa:bb:cc:dd:ee:ff
    # The internal path must be /home/flow
    volumes:
      # This uses an external path for the Flow data. The 'flow' user inside the container must be able to read and write to this directory.
      - /home/flow:/home/flow
networks:
  flexlm:
```

Some important values in docker-compose.yml:

* **image:** The container image tag. The value for $tag corresponds to the Partek Flow version. Please visit the [Partek Flow Release Notes](/partek-flow/release-notes) to find the latest version number and replace $tag with this value.
* **environment:** Environment variables that configure additional Partek Flow features.
* **port:** The exposed port used to access Partek Flow via a web browser. In this example <http://localhost:8080/>
* **mac\_address:** The mac\_address needs to match the value for HOSTID in the provided Partek Flow license file. Note that the ":" character is required here, but may not be present in the HOSTID value.
* **volumes:** This specifies the filesystem path on the server to be shared with the Partek Flow. All user data will be read and written to this path. Additional paths can be added. Where write access is required, e.g. the /home/flow mount, the path should be set to uid:gid 1000:1000.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Java KeyStore and Certificates

* [Java Keystore](#java-keystore)
* [Adding a certificate to the KeyStore](#adding-a-certificate-to-the-keystore)

## Java Keystore

JKS or Java KeyStore is used in Flow for some very specific scenarios where encryption is involved and there is a need for asymmetric encryption.

Partek Flow is shipped with a Java Keystore on its own, the file is found at .../partek\_flow/distrib/flowkeystore where you may want to add your public and private certificates.

## Adding a certificate to the KeyStore

If you already have a certificate please skip to the next step.

### Create a certificate

Please place the key in a secure folder. (it is advisable to place in Flow's home directory. eg. /home/flow/keys

```
[~] openssl genrsa -out flow.key 2048
```

```
[~] openssl ecparam -genkey -name secp384r1 -out flow.key
```

```
[~] openssl req -new -x509 -sha256 -key flow.key -out flow.crt -days 3650
```

These commands above are meant to be used in a terminal. There are other ways to help you make a certificate but they will not going to be mentioned here.

If you wish to understand the flags used above please refer to the OpenSSL documentation.

### Import a certificate into flowkeystore

For this step you will have to find where the cacerts file is located, it is under the Java installation, if you do not know how to do it contact us and we can help.

In the example the cacerts file is located at /usr/lib/jvm/java-11-openjdk-amd64/lib/security/cacerts

```
[~] keytool -import -file /home/flow/.partekflow/keys/flow.key -alias someName -keystore /usr/lib/jvm/java-11-openjdk-amd64/lib/security/cacerts -storepass changeit -noprompt
```

### Tell the JVM where to find the key

We need to tell Partek Flow where the key is located, to do this we will edit a file which contains some of the Flow settings.

The file is usually located at /etc/partekflow\.conf if you do not have this file we would advise to use the bashrc file from the system user that runs Partek Flow.

At the end of that file please add:

```
export CATALINA_OPTS="$CATALINA_OPTS -Djavax.net.ssl.trustStore=${HOME}/keys"
```

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Kubernetes

Below are the yaml documents which describe the bare minimum infrastructure needed for a functional Flow server. It is best to start with a single-node proof of concept deployment. Once that works, the deployment can be extended to multi-node with elastic worker allocation. Each section is explained below.

## The Flow headnode pod

```
apiVersion: v1
kind: Pod
metadata:
  name: flowheadnode
  namespace: partek-flow
  labels:
    app.kubernetes.io/name: flowheadnode
    deployment: dev
spec:
  securityContext:
    fsGroup: 1000
  containers:
    - name: flowheadnode
      image: xxxxxxxxxxxx.dkr.ecr.us-west-2.amazonaws.com/partek-flow:current-23.0809.22
      resources:
        requests:
          memory: "16Gi"
          cpu: 8
      env:
        - name: PARTEKLM_LICENSE_FILE
          value: "@flexlmserver"
        - name: PARTEK_COMMON_NO_TOTAL_LIMITS
          value: "1"
        - name: CATALINA_OPTS
          value: "-DFLOW_WORKER_MEMORY_MB=1024 -DFLOW_WORKER_CORES=2 -Djavax.net.ssl.trustStore=/etc/flowconfig/cacerts -Xmx14g"
      volumeMounts:
        - name: home-flow
          mountPath: /home/flow
        - name: flowconfig
          readOnly: true
          mountPath: "/etc/flowconfig"
  volumes:
    - name: home-flow
      persistentVolumeClaim:
        claimName: partek-flow-pvc
    - name: flowconfig
      secret:
        secretName: flowconfig
```

### Pod metadata

On a kubernetes cluster, all Flow deployments are placed in their own namespace, for example namespace: partek-flow. The label app.kubernetes.io/name: flowheadnode allows binding of a service or used to target other kubernetes infrastructure to this headnode pod. The label deployment: dev allows running multiple Flow instances in this namespace (dev, tst, uat, prd, etc) if needed and allows workers to connect to the correct headnode. For stronger isolation, running each Flow instance in its own namespace is optimal.

### Data storage

The Flow docker image requires 1) a writable volume mounted to /home/flow 2) This volume needs to be readable and writable by UID:GID 1000:1000 3) For a multi-node setup, this volume needs to be cross mounted to all worker pods. In this case, the persistent volume would be backed by some network storage device such as EFS, NFS, or a mounted FileGateway.

This section achieves goal 2)

```
spec:
  securityContext:
    fsGroup: 1000
```

The flowconfig volume is used to override behavior for custom Flow builds and custom integrations. It is generally not needed for vanilla deployments.

### The Flow docker image

Partek Flow is shipped as a single docker image containing all necessary dependencies. The same image is used for worker nodes. Most deployment-related configuration is set as environment variables. Auxiliary images are available for additional supporting infrastructure, such as flexlm and worker allocator images.

Official Partek Flow images can be found on our release notes page: Release Notes The image tags assume the format: registry.partek.com/rtw:YY.MMMM.build New rtw images are generally released several times a month. The image in the example above references a private ECR. It is highly recommended that the target image from registry.partek.com be loaded into your ECR. Image pulls will be much faster from AWS - this reduces the time to dynamically allocate workers. It also removes a single point of failure - if registry.partek.com were down it would impact your ability to launch new workers on demand.

### Flow headnode resource request

Partek Flow uses the head node to handle all interactive data visualization. Additional CPU resources are needed for this, the more the better and 8 is a good place to start. As for memory, we recommend 8 to 16 GiB. Resource limits are not included here, but are set to large values globally:

```
# This allows us to create pods with only a request set, but not a limit set. Further tuning is recommended. 
apiVersion: v1
kind: LimitRange
metadata:
  name: partek-flow-limit-range
spec:
  limits:
    - max:
        memory: 512Gi
        cpu: 64
      default:
        memory: 512Gi
        cpu: 64
      defaultRequest:
        memory: 4Gi
        cpu: 2
      type: Container
```

### Relevant Flow headnode environment variables

#### PARTEKLM\_LICENSE\_FILE

Partek Flow uses FlexLM for licensing. Currently we do not offer or have implemented any alternative. Values for this environment variable can be:

#### @flexlmserveraddress

An external flexlm server. We provide a Partek specific container image and detail a kubernetes deployment for this below. This license server can also live outside the kubernetes cluster - the only requirement being that it is network accessible. /home/flow/.partekflow/license/Partek.lic - Use this path exactly. This path is internal to the headnode container and is persisted on a mounted PVC.

Unfortunately, FlexLM is MAC address based and does not quite fit in with modern containerized deployments. There is no straightforward or native way for kubernetes to set the MAC address upon pod/container creation, so using a license file on the flowheadnode pod (/home/flow/.partekflow/license/Partek.lic ) could be problematic (but not impossible). In further examples below, we provide a custom FlexLM container that can be instantiated as a pod/service. This works by creating a new network interface with the requested MAC address inside the FlexLM pod.

#### PARTEK\_COMMON\_NO\_TOTAL\_LIMITS

Please leave this set at "1". Partek Flow need not enforce any limits as that is the responsibility of kubernetes. Setting this to anything else may result in Partek executables hanging.

#### CATALINA\_OPTS

This is a hodgepodge of Java/Tomcat options. Parts of interest:

```
-DFLOW_WORKER_MEMORY_MB=1024 -DFLOW_WORKER_CORES=2
```

It is possible for the Flow headnode to execute jobs locally in addition to dispatching them to remote workers. These two options set resource limits on the Flow internal worker to prevent resource contention with the Flow server. If remote workers are not used and this remains a single-node deployment, meaning ALL jobs will execute on the internal worker, then it is best to remove the CPU limit (-DFLOW\_WORKER\_CORES) and only set -DFLOW\_WORKER\_MEMORY\_MB equal to the kubernetes memory resource request.

```
-Djavax.net.ssl.trustStore=/etc/flowconfig/cacerts
```

If Flow connects to a corporate LDAP server for authentication, it will need to trust the LDAP certificates.

```
-Xmx14g
```

JVM heap size. If the internal worker is not used, set this to be a little less than the kubernetes memory resource request. If the internal worker is an use, and the intent is to stay with a single-node deployment, then set this to be \~ 25% of the kubernetes memory resource request, but no less than \~ 4 GiB.

## The Flow headnode service definition

```
apiVersion: v1
kind: Service
metadata:
  name: flowheadnode
spec:
  type: ClusterIP
  ports:
    - port: 80
      targetPort: 8080
      protocol: TCP
      name: http
    - port: 2552
      targetPort: 2552
      protocol: TCP
      name: akka
    - port: 8443
      targetPort: 8443
      protocol: TCP
      name: licensing
  selector:
    app.kubernetes.io/name: flowheadnode
```

The flowheadnode service is needed 1) so that workers have a DNS name (flowheadnode) to connect to when they start and 2) so that we can attach an ingress route to make the Flow web interface accessible to end users. The app.kubernetes.io/name: flowheadnode selector is what binds this to the flowheadnode pod.

* 80:8080 - Users interact with Flow entirely over a web browser
* 2552:2552 - Workers communicate with the Flow server over port 2552
* 8443:8443 - Partek executed binaries connect back to the Flow server over port 8443 to do license checks

## Ingress to flowheadnode

```
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: flowheadnode
  annotations:
    kubernetes.io/ingress.class: "nginx"
    nginx.ingress.kubernetes.io/force-ssl-redirect: "true"
spec:
  rules:
    - host: flow.dev-devsvc.domain.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: flowheadnode
                port:
                  number: 80
```

This provides external users HTTPS access to Flow at host: flow\.dev-devsvc.domain.com Your details will vary. This is where we bind to the flowheadnode service.

## The flexlm service pod

```
# On a NEW deployment, you need to exec into this pod and add the license file
# to /usr/local/flexlm/licenses
# After a license file is present, the flexlm daemon will start automatically

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: flexlmserver-pvc
spec:
  accessModes:
    - ReadWriteOnce
  resources:
    requests:
      storage: 10Gi     # flex.log is the only thing that slowly grows here
  storageClassName: gp2-ebs-sc
  volumeMode: Filesystem
---
apiVersion: v1
kind: Service
metadata:
  name: flexlmserver
spec:
  type: ClusterIP
  ports:
    - port: 27000
      targetPort: 27000
      protocol: TCP
      name: flexmain
    - port: 27001
      targetPort: 27001
      protocol: TCP
      name: flexvendor
  selector:
    app.kubernetes.io/name: flexlmserver
---
apiVersion: v1
kind: Pod
metadata:
  name: flexlmserver
  namespace: partek-flow
  labels:
    app.kubernetes.io/name: flexlmserver
spec:
  containers:
    - name: flexlmserver
      image: public.ecr.aws/partek-flow/kube-flexlm-server
      ports:
        - containerPort: 27000
        - containerPort: 27001
      resources:
        limits:
          memory: "256Mi"
          cpu: 1
      securityContext:
        capabilities:
          add: ["NET_ADMIN"]
      volumeMounts:
        - name: flexlmserver-pvc
          mountPath: /usr/local/flexlm/licenses
  volumes:
    - name: flexlmserver-pvc
      persistentVolumeClaim:
        claimName: flexlmserver-pvc
```

The yaml documents above will bring up a complete Partek-specific license server.

Note that the service name is flexlmserver. The flowheadnode pod connects to this license server via the PARTEKLM\_LICENSE\_FILE="@flexlmserver" environment variable.

You should deploy this flexlmserver first, since the flowheadnode will need it available in order to start in a licensed state.

Partek will send a Partek.lic file licensed to some random MAC address. When this license is (manually) written to /usr/local/flexlm/licenses, the pod will continue execution by creating a new network interface using the MAC address in Partek.lic, then it will start the licensing service. This is why the NET\_ADMIN capability is added to this pod.

The license from Partek must contain VENDOR parteklm PORT=27001 so the vendor port remains at 27001 in order to match the service definition above. Without this, this port is randomly set by FlexLM.

This image is currently available from public.ecr.aws/partek-flow/kube-flexlm-server but this may change in the future.


# Live Training Event Recordings

Here you will find videos of past live training webinars.

* [Bulk RNA-Seq Analysis Training](/partek-flow/live-training-event-recordings/bulk-rna-seq-analysis-training)
* [Basic scRNA-Seq Analysis & Visualization Training](/partek-flow/live-training-event-recordings/basic-scrna-seq-analysis-and-visualization-training)
* [Advanced scRNA-Seq Data Analysis Training](/partek-flow/live-training-event-recordings/advanced-scrna-seq-data-analysis-training)
* [Bulk RNA-Seq and ATAC-Seq Integration Training](/partek-flow/live-training-event-recordings/bulk-rna-seq-and-atac-seq-integration-training)
* [Spatial Transcriptomics Data Analysis Training](/partek-flow/live-training-event-recordings/spatial-transcriptomics-data-analysis-training)
* [scRNA and scATAC Data Integration Training](/partek-flow/live-training-event-recordings/scrna-and-scatac-data-integration-training)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Bulk RNA-Seq Analysis Training

{% embed url="<https://youtu.be/8Usk5XkVvXg>" %}

**Agenda with start times:**

* Transfer file to the server (00:02:28)
* Import fastq files (00:03:55)
* Add sample attributes (00:06:00)
* Pre-alignment QA/QC (00:10:45)
* Alignment QC (00:15:30)
* Quantification (00:21:45)
* Filter and normalization (00:32:13)
* Differential analysis and visualization (00:38:50)
* Filter genes to generate a heatmap (00:53:50)
* Biological interpretation (01:08:00)


# Basic scRNA-Seq Analysis & Visualization Training

{% embed url="<https://youtu.be/4Ca724klx3w>" %}

**Agenda with start times:**

* Data import (00:01:07)
* QA/QC (00:10:52)
* Normalization (00:33:28)
* Generating a UMAP (00:35:22)
* Graph-based clustering (00:38:24)
* Cell type classification (00:42:31)
* Differential analysis (01:12:00)


# Advanced scRNA-Seq Data Analysis Training

{% embed url="<https://youtu.be/902jCe3ATQE>" %}

**Agenda with start times:**

* Data introduction (00:01:45)
* Batch removal (00:03:40)
* Cell type classification (00:10:57)
* Differential gene expression detection (00:26:30)
* Cell type abundance analysis (00:32:00)
* Heatmap and bubble map (00:41:25)
* Automated cell type classification (00:53:20)
* Pseudo bulk analysis (01:01:40)
* CITE-Seq data analysis (01:09:20)
* Spatial transcriptome data analysis (01:15:30)


# Bulk RNA-Seq and ATAC-Seq Integration Training

{% embed url="<https://youtu.be/-hPq-av57Sk>" %}

**Agenda with start times:**

* Data set overview (00:01:14)
* Bulk RNA-Seq analysis pipeline (00:02:40)
* ATAC-Seq alignment (00:39:47)
* Filter alignment (00:42:30)
* Peak detection and chromosome view (00:46:10)
* Filter peaks (00:57:00)
* Peak annotation (01:00:20)
* Compare peak regions across samples (01:07:43)
* Compare peak regions across different samples (01:12:24)
* Motif detection (01:19:05)
* RNA-Seq and ATAC-Seq integration (01:22:25)


# Spatial Transcriptomics Data Analysis Training

{% embed url="<https://youtu.be/_WCBjbipLr0>" %}

**Agenda with start times:**

* What is Partek Flow (00:1:39)
* Importing spatial transcriptome data (00:5:20)
* QA/QC, filter, and normalize data (00:13:25)
* Exploratory analysis (00:20:05)
* Batch effect removal and graph-based clustering (00:27.53)
* Integrate spatial and gene expression data (00:35.50)
* Detect differentially expressed genes (00:41:15)
* Pathway analysis and biological interpretation (00:45:15)
* Spatial transcriptome pipeline (00:50:13)


# scRNA and scATAC Data Integration Training

{% embed url="<https://youtu.be/761H_FYvy1M>" %}

**Agenda with start times:**

* Data file format (00:00:43)
* Data import (00:10:43)
* QC and filter (00:13:26)
* scRNA analysis (00:20:10)
* scATAC peak annotation (00:20:03)
* scATAC promoter analysis (00:26:22)
* scATAC peak analysis (00:29:18)
* scRNA and scATAC integration (00:33:00)
* scATAC motif detection (00:57:09)


# Tutorials

Partek Flow tutorials provide step-by-step instructions using a supplied data set to teach you how to use the software tools. Upon completion of each tutorial, you will be able to apply your knowledge in your own studies.

* [Creating and Analyzing a Project](/partek-flow/tutorials/creating-and-analyzing-a-project)
* [Bulk RNA-Seq](/partek-flow/tutorials/bulk-rna-seq)
* [Analyzing Single Cell RNA-Seq Data](/partek-flow/tutorials/analyzing-single-cell-rna-seq-data)
* [Analyzing CITE-Seq Data](/partek-flow/tutorials/analyzing-cite-seq-data)
* [10x Genomics Visium Spatial Data Analysis](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis)
* [10x Genomics Xenium Data Analysis](/partek-flow/tutorials/10x-genomics-xenium-data-analysis)
* [Single Cell RNA-Seq Analysis (Multiple Samples)](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples)
* [Analyzing Single Cell ATAC-Seq data](/partek-flow/tutorials/analyzing-single-cell-atac-seq-data)
* [Analyzing Illumina Infinium Methylation array data](https://help.partek.illumina.com/partek-flow/tutorials/analyzing-illumina-infinium-methylation-array-data)
* [NanoString CosMx Tutorial](/partek-flow/tutorials/nanostring-cosmx-tutorial)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Creating and Analyzing a Project

Partek Flow software manages separate experiments as projects. A complete project consists of input data, tasks used to analyze the data, the resulting output files, and a list of users involved in the analysis.

![Creating and analyzing a project analyses tab](/files/otMwtxLY9Haz1rSfEOav)

This chapter provides instructions in creating and analyzing a project and covers:

* [Creating a New Project](/partek-flow/tutorials/creating-and-analyzing-a-project/creating-a-new-project)
* [The Metadata Tab](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab)
* [The Analyses Tab](/partek-flow/tutorials/creating-and-analyzing-a-project/the-analyses-tab)
* [The Log Tab](/partek-flow/tutorials/creating-and-analyzing-a-project/the-log-tab)
* [The Project Settings Tab](/partek-flow/tutorials/creating-and-analyzing-a-project/the-project-settings-tab)
* [The Attachments Tab](/partek-flow/tutorials/creating-and-analyzing-a-project/the-attachments-tab)
* [Project Management](/partek-flow/tutorials/creating-and-analyzing-a-project/project-management)
* [Importing a GEO / ENA project](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/creating-and-analyzing-a-project/importing-a-geo-ena-project/README.md)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Creating a New Project

Using a web browser, log in to Partek Flow. From the Home page click the **New Project** button; enter a project name (Figure 1) and then click **Create project**.

<figure><img src="/files/rjFfO1rRqJ6XiT2AY9kP" alt=""><figcaption><p><em>Figure 1. Partek Flow Home page and the dialog box for naming a project (inset)</em></p></figcaption></figure>

The Project name is the basis of the default name of the output directory for this project. Project names are unique, thus a new project cannot have the same name as an existing project within the same Partek Flow server.

Once a new project has been created, the user is automatically directed to the Analysis tab of the Project View (Figure 2).

<figure><img src="/files/9kvCaQIc6CIx438PufA8" alt=""><figcaption><p><em>Figure 2. The Analyses tab is used to Add data after a project has been created</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# The Metadata Tab

The Partek Flow *Metadata* Tab has an option to import data, and is where sample/cell attributes are managed. This is also where users can modify the location of the project output folder.

* [Import data](#import-data)
  * [Automatically create samples from files](#automatically-create-samples-from-files)
  * [Create a new blank sample](#create-a-new-blank-sample)
  * [Importing count matrix data](#importing-count-matrix-data)
* [Project output directory](#project-output-directory)
* [Sample Annotation](#sample-annotation)
  * [Sample attributes](#sample-attributes)
  * [Adding a categorical attribute](#adding-a-categorical-attribute)
  * [Adding a numeric attribute](#adding-a-numeric-attribute)
  * [Adding a system-wide attribute](#adding-a-system-wide-attribute)
  * [Assigning categories or values to attributes](#assigning-categories-or-values-to-attributes)
  * [Assigning attributes using a Sample Annotation Text File](#assigning-attributes-using-a-sample-annotation-text-file)
  * [Guidelines for preparing the sample annotation text file](#guidelines-for-preparing-the-sample-annotation-text-file)
  * [Use of attributes as Optional columns in task report tables](#use-of-attributes-as-optional-columns-in-task-report-tables)
* [Deleting or Renaming samples within a Project](#deleting-or-renaming-samples-within-a-project)

## Import data

The *Metadata* tab can be used to import data. To add samples to the project, click **Add data** under Import, different import options are displayed using the cascading menu (Figure 1).

<figure><img src="/files/hDVsrOOhgB45CxinKqO0" alt=""><figcaption><p><em>Figure 1. The Partek Flow Metadata tab and selecting options for adding samples</em></p></figcaption></figure>

### Automatically create samples from files

This method adds samples by creating them simultaneously as the data gets imported into a project. The sample names are assigned automatically based on filenames.

*Before proceeding, it is ideal that you have already transfer the data you wish to analyze in a folder (with appropriate permissions) within the Partek Flow server. Please seek assistance from your system administrator in uploading your data directly.*

Select the **Automatically create samples from files** button. The next screen will feature a file browser that will show any folders you have access to in the Partek Flow server (Figure 2). Select a folder by clicking the folder name. Files in the selected folder that have file formats that can be imported by Partek Flow will be displayed and tick-marked on the right panel. You can exclude some files from the folder by unselecting the check mark on the left side of the filename. When you have made your selections, click the **Create sample** button.

<figure><img src="/files/Ow7S5UqVIQAodAFz61p0" alt=""><figcaption><p><em>Figure 2. Selecting files in the Partek Flow server to be imported in a project</em></p></figcaption></figure>

Alternatively, files can also be uploaded and imported into the project from the user's local computer -only use this option if your file size is less than 500MB. Select the **My computer** radio button (Figure 3) and the options of selecting the local file and the upload (destination) directory will appear. Only one file at a time can be imported to a project using this method.

<figure><img src="/files/mDT5dg34G4VbNgRQePrZ" alt=""><figcaption><p><em>Figure 3. Selecting files from the user's local computer for upload and import</em></p></figcaption></figure>

Multiple data files can be compressed a single .zip file before uploading. Partek Flow will automatically unzip the files and put them in the upload directory.

Please be aware that the use of the method illustrated in Figure 3 highly depends on the speed and latency of the Internet connection between the user's computer and the server where Partek Flow is installed. Given the large size of most genomics data sets, is not recommended in most cases.

After successful creation of samples from files, the Data tab now contains a Sample management table (Figure 4). The Sample name column in the table is automatically generated based on the filenames and the table is sorted in alphabetical order.

Clicking the on the\*\* Show data files\*\* link on the lower right side of the sample management table will expand the table and reveal the filenames of the files associated with each sample. Conversely, clicking on **Hide data files** will hide the file information.

The columns in the expanded view show the files associated with each sample. Files are organized by file type. Any filename extensions that indicate compression (such as .gz) are not shown.

<figure><img src="/files/5yBOYGdaPAe3tx3ZqOZ8" alt=""><figcaption><p><em>Figure 4. The sample management table with data files shown</em></p></figcaption></figure>

Once a sample is created in a project, the files associated with it can be modified. In the expanded view, mouse over the +/- column of a sample. The highlighted icons will correspond to the options for the sample on that row.

Click the **green icon** ( ![Green icon](/files/Co7v2sWm0np2kaKy6JxC) ) to associate additional files or the **red icon** ( ![Red icon](/files/d2cX8eXhMu7p22BKCMBr) ) to dissociate a file from a sample. You can manually associate multiple files with one sample. Dissociating a file from a sample does not delete the file from the Partek Flow server.

### Create a new blank sample

Samples can be added one at a time by selecting the **Create a new blank sample** option (Figure 5). In the following dialog box, type a sample name and click **Create**. This process creates a sample entry in the sample management table but there is no associated file with it, hence it is a "blank sample."

Expanding the Sample management table by clicking **Show data files** on the lower left corner of the table will reveal the option to associate files to the blank sample.

Mouse over the +/- column and click the **green icon** ( ![Green icon](/files/Co7v2sWm0np2kaKy6JxC) ) to associate a file(s) to the sample. Perform the process for every sample in your project.

<figure><img src="/files/K0UYqqOdqx2gk8E75Xhd" alt=""><figcaption><p><em>Figure 5. Adding a blank sample</em></p></figcaption></figure>

### Importing count matrix data

Alternatively, if you have a matrix of data, such as raw read count data in text format, select Import count matrix. The requirements of this text file are listed below:

* The file contains numeric values in a tab-delimited format, samples can be on rows while features (e.g.gene names) are in columns, or vice versa
* The file contains unique sample IDs and feature IDs
* If the data contains sample attribute information, all these attributes have to be ether
  * The leftmost columns when samples are on rows (Figure 6)
  * The first few rows when samples are on columns (Figure 7)

<figure><img src="/files/wAcLKF0sAPCifuvcEuB6" alt=""><figcaption><p><em>Figure 6. Example of sample on row, the first column is sample ID, the 2nd and 3rd columns contain sample attribute information, feature count starts from column 4</em></p></figcaption></figure>

<figure><img src="/files/ULPEXfaeq44P0ApbJ5IZ" alt=""><figcaption><p><em>Figure 7. Example of sample on column, the first row is sample ID, the 2nd and 3rd rows contain sample attribute information, feature count starts from column 4</em></p></figcaption></figure>

Like all other input files, you can upload the file from the Partek Flow server, My Computer or via a URL. Uploading the file brings up a file preview window (Figure 8). The preview of the first few rows and columns of the text file should help you determine on which rows/columns the relevant counts are located (the preview will display up to 100 rows and 100 columns). Inspect the text preview and indicate the orientation of the text file under **File format>Input format**.

If the read counts are based on a compatible annotation file in Partek Flow, you can specify that annotation file under **Gene/feature annotation**. Select the appropriate genome build and annotation model for your count data. Select the **Contain sample attributes** checkbox if your data includes additional sample information.

<figure><img src="/files/vxGOnpfxbChRx62lyQIs" alt=""><figcaption><p><em>Figure 8. Importing count matrix data preview. This example is showing a count matrix text file has samples on rows with sample attributes colummns</em></p></figcaption></figure>

The example above is showing an example text file with samples listed on rows. The gene ID is compatible with the *hg19 RefSeq Transcripts - 2016-08-01* annotation model. Under the **Column information** and **Row information** sections, indicate the location of the Sample ID, which in this case is on *Column 1*. Indicate the sample attribute location by marking where it starts, which in the example is at *Column 2*. Mark the Feature ID, which in this case are gene IDs and starts at *Column 4*.

If the data has been log transformed, specify the base under **Counts format**.

## Project output directory

The project output directory is the folder within the Partek Flow server where all output files produced during analysis will be stored.

*The default directory is configured by the Partek Flow Administrator under the Settings menu (under **System Preferences > Default project output directory**).*

If the user does not override the default, the task output will go to a subdirectory with the name of the Project.

The user has the option of specifying an existing folder or creating a new one as the project output directory. To do so, click the ![Pencil icon](/files/0FsC0ACMhZsMGmRrOvoi) icon next to the directory and specify or create a new folder in the dialog box.

## Sample Annotation

After samples have been added in the project, additional information about the samples can be added. Information such as disease type, age, treatment, or sex can be annotated to the data by assigning the Attributes for each sample.

Certain tasks in Partek Flow, such as Gene-Specific Analysis, require that samples be assigned attributes in order to do statistical comparisons between groups or samples. As attributes are added to the project, additional columns in the sample management table will be created.

### Sample attributes

Attributes can be managed or created within a project. Under the Data tab, click the button to open the *Manage attributes* page (Figure 9).

<figure><img src="/files/dEREYYRiEFvppEyrlOq4" alt=""><figcaption><p><em>Figure 9. Managing attributes</em></p></figcaption></figure>

To prepare for later data analysis using statistical tools, attributes can either be categorical or numeric (i.e., continuous).

### Adding a categorical attribute

For categorical attributes, there are two levels of visibility. *Project-specific* categorical attributes are visible only within the current project. *System-wide categorical attributes* are visible across all the projects within the Partek Flow server, and are useful for maintaining uniformity of terms. Importing samples in a new project will retain the system-wide attributes, but not the project-specific attributes.

A feature of Partek Flow is the use of controlled vocabulary for *categorical attributes*, allowing samples to be assigned only within pre-defined categories. It was designed to effectively manage content and data and allow teams to share common definitions. The use of standard terms minimizes confusion.

To add a categorical attribute in the *Manage attributes* page, click the **Add new attribute** (Figure 10). In the dialog box, type a **Name** for the attribute, select the **Categorical** radio button next to **Attribute type**, select the visibility of the attribute and then click the **Add** button.

<figure><img src="/files/HogRqZNwdsOZ6SwWToJO" alt=""><figcaption><p><em>Figure 10. Adding a categorical attribute and defining the categories</em></p></figcaption></figure>

Individual categories for the attribute must then be entered. Enter a name of the **New category** in the New category text box and click **Add** (Figure 11). The Name of the new category will show up in the table. The category can also be edited by clicking ![Pencil icon](/files/0FsC0ACMhZsMGmRrOvoi) or deleted by clicking ![Red x icon](/files/AwpurpGwRw4U9gEMUdna) (visible on mouse-over). Repeat to add additional categories within the attribute.

Repeat the process for additional attributes of the samples in your study. When done, click **Back to sample management table**. Categorical attributes will default to Project-specific visibility.

Click an attribute name to drag and drop can change the order of the attributes displayed on the data tab. Click on group name to drag and drop vertically can change the order of the group name, which can be reflected on visualization.

<figure><img src="/files/IthmgwFgpbeAdIFZGWoA" alt=""><figcaption><p><em>Figure 11. Adding categories within an attribute</em></p></figcaption></figure>

### Adding a numeric attribute

To add a numeric attribute in the Manage attributes page, click the **Add new attribute**. In the dialog box (Figure 13), type a **Name** for the attribute, select the **Numeric** radio button next to **Attribute type**, and then click the **Add** button. Some optional parameters for numeric attributes include the Minimum value, Maximum value, and Units. When done, click **Add** to return to the Manage attributes page. Repeat the process add more numeric attributes. When done, click **Back to sample management table**.

<figure><img src="/files/Sqe4MeaPzd7pukmcohCr" alt=""><figcaption><p><em>Figure 12. Adding a numeric attribute and specifying the units</em></p></figcaption></figure>

### Adding a system-wide attribute

Since system-wide attributes do not have to be created by the current user, they only need to be added to the sample management table in a project.

In the *Data* tab, click **Add a system-wide attribute** button. In the dialog box that follows (Figure 14), a drop down menu is located next to **Add attribute** where you can select the *System-wide attribute* you would like to add to the project. Once selected, it will be recognized automatically as either *Categorical, system-wide* or the *Numeric attribute*.

For an *System-wide categorical attribute*, the different categories are listed and you have the option of pre-filling the columns with N/A (or any other category within the attribute). Click **Add column** and you will return to the *Data* Tab.

<figure><img src="/files/zU5QlHOP5yaeACc3Vnfn" alt=""><figcaption><p><em>Figure 13. Adding a system-wide categorical attribute column</em></p></figcaption></figure>

### Assigning categories or values to attributes

After adding all the desired attributes to a project, the sample management table will show a new column for each attribute (Figure 15). The columns will initially as "N/A", as the samples have not yet been categorized or assigned a value. To edit the table, click **Edit attributes**. Assign the sample attributes by using a drop down for categorical attributes (controlled vocabulary) or typing with a keyboard for numeric attributes.

<figure><img src="/files/sRbOB0KiyeNYWpDVb1ty" alt=""><figcaption><p><em>Figure 14. A sample management table prior to assigning attribute columns for a new attribute</em></p></figcaption></figure>

When all the attributes have been entered, click **Apply changes** and the sample management table will be updated. After editing the sample table, make sure there are no fields with blank or N/A values before proceeding. To rename or delete attributes, click **Manage attributes** from the *Data* tab to access the *Manage attributes* page.

### Assigning attributes using a Sample Annotation Text File

Another way to assign attributes to samples in the Data tab is to use a text file that contains the table of attributes and categories/values. This table is prepared outside of Partek Flow using any text editing software capable of saving tab-separated text files.

Using a text editor, prepare a table containing the attributes. An example is shown in Figure 16. There should only be one tab between columns with no extra tabs after the last column. In this particular example, the first column contains the filename and the text file is saved as *Sampleinfo.txt*.

<figure><img src="/files/zQbJlad2CKtyyO35U86I" alt=""><figcaption><p><em>Figure 15. A sample annotation text file. This view shows tab stops</em></p></figcaption></figure>

The first row of the table in the text file contains the attributes (as headers). The first *column* of the table in the text file, regardless of the header of the first column, should contain either the sample names or the file names of the samples already added in Partek Flow. The first column is the unique identifier that will match the samples to the correct values or categories.

To upload sample attributes, click **Assign sample attributes from a file** in the *Data* tab. Then indicate where the attribute text file is stored and navigate to it. Partek Flow will parse the text file and present attributes that will be available for import (Figure 17).

Select the attributes you want to import by clicking the Import check box. Imported attributes that do not currently exist in the project will create new *project-specific attributes*.

<figure><img src="/files/DNHp9b2exSWdhqhjQXUB" alt=""><figcaption><p><em>Figure 16. Assigning attributes of samples using a sample annotation text file</em></p></figcaption></figure>

You can change the name of a specific attribute by editing the *Attribute* name text box. Columns containing letter characters are automatically selected as *categorical attributes*. Columns containing numbers are suggested to be *numeric attributes* and can be changed to categorical using the drop down menu under *Attribute* type.

### Guidelines for preparing the sample annotation text file

* The first column is always the unique identifier and can refer only to *File names* or *Sample names*.
* If using *Sample names* in the first column, they must match the entries of the Sample name column in the Sample management table.
* If using *File names* in the first column, use the filenames shown in the *fastq* column of the expanded sample management table (see Figure 4) then add the extension .gz. All filenames must include the complete file extension (e.g., *Samplename.fastq.gz*).
* The header name of the first column of the table (top left cell of our text table) is irrelevant but should not be left blank. Whether the first column contains *File names* or *Sample names* will be chosen during the process.
* The last column cannot have empty values
* Missing data (blank cells) can only be handled if the attribute is numeric. If it is categorical, please put a character in it.

It is advisable to use Sample name as the first column identifier when:

* Samples are associated with more than one file (for instance, paired-end reads and/or technical replicates).
* The files were imported in the SRA format (from the Sequence Read Archive database). In Partek Flow, they are automatically converted to the FASTQ format. Consequently, their filenames would change once they are imported. The new file names can be seen by expanding the sample management table, the new extension would be *.fastq.gz*.

If attributes are assigned from two different text files, the following will happen:

* If the previous attributes have the same header and type (both are either categorical or numeric), the values are overwritten.
* If there are different/additional headers on the "second round" of assignment, these new attributes will be appended to the table.
* For *numeric attributes*, a "blank" value will not override a previous value.

### Use of attributes as Optional columns in task report tables

The attributes assigned to the samples within the Data Tab will be associated with the samples throughout the project. During the course of analysis, Partek Flow tasks generate various tables and any attributes associated with a sample can be included in the table as optional columns. An example is shown in Figure 18 for a pre-alignment QA/QC report where the Optional columns link on the top left of the table reveal the different sample attributes.

<figure><img src="/files/NKm78lSfsAtFC6MN41fu" alt=""><figcaption><p><em>Figure 17. Optional columns include sample attributes</em></p></figcaption></figure>

## Deleting or Renaming samples within a Project

In the *Data* tab, each sample can be renamed or deleted from the project by clicking the gear icon next to the sample name. The gear icon is readily visible upon mouse over (Figure 19). Sample can only be deleted if no analysis has been performed on the data yet. If any analysis has been performed on the data node, then delete sample operation is invisible. You can perform filter samples in downstream analysis if you want to exclude certain samples in further analysis. Deleting a sample from a project does not delete the associated files, which will remain on the disk.

<figure><img src="/files/gQwQwF8Q4ZuUYCVd0ckq" alt=""><figcaption><p><em>Figure 18. Renaming or deleting a sample</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# The Analyses Tab

After samples have been added and associated with valid data files, in a Partek Flow project, a data node will appear in the Analyses tab (Figure 1). The Analyses tab is where different analysis tools and the corresponding reports are accessed.

* [Data and Task Nodes](#data-and-task-nodes)
* [The Context Sensitive Menu](#the-context-sensitive-menu)
* [Running a Task](#running-a-task)
* [Cancelling and Deleting Tasks](#cancelling-and-deleting-tasks)
* [Task Results and Task Actions](#task-results-and-task-actions)
* [Layers](#layers)
* [Collapsing Tasks](#collapsing-tasks)
* [Downloading Data](#downloading-data)

<figure><img src="/files/Ey0Z2axBB2IUqN2inqS2" alt=""><figcaption><p><em>Figure 1. The Analyses tab showing a data node of unaligned reads</em></p></figcaption></figure>

## Data and Task Nodes

The Analyses tab contains two elements: data nodes (circles) and task nodes (rounded rectangles) connected by lines and arrows. Collectively, they represent a data analysis pipeline.

Data nodes (Figure 2) may represent a file imported into the project, or a file generated by Partek Flow as an output of a task (e.g., alignment of FASTQ files generates BAM files).

Missing image Figure 2. The Analyses tab showing a data node of unaligned reads

Task nodes (Figure 3) represent the analysis steps performed on the data within a project. For details on the tasks available in Partek Flow, see the specific chapters of this user manual dedicated to the different tasks.

<figure><img src="/files/9lQUxsclIBaAncUaFokO" alt=""><figcaption><p><em>Figure 3. Examples of different types of data nodes</em></p></figcaption></figure>

## The Context Sensitive Menu

Clicking on a node reveals the context sensitive menu, on the right side of the screen.

<figure><img src="/files/OvFaAGNty4dfUEJwshNm" alt=""><figcaption><p><em>Figure 4. The context sensitive menu is revealed when a node is selected</em></p></figcaption></figure>

Only the tasks that are available for the selected data node will be listed in the menu. For data nodes, actions that can be performed on that specific data type will appear.

In Figure 4, a node that contains Unaligned reads is selected (bold black line). The tasks listed are the ones that can be performed on unaligned data (QA/QC, Pre-analysis tools, and Aligners).

To hide the context sensitive menu, simply click the ![Grey x icon](/files/GWxbMMiYjXIgRwgkH0bo) symbol on the upper left corner of the context sensitive menu. Clicking the triangles will collapse ( ![Collapse icon](/files/umBCTBMFaN50QKZc1aHh) ) or expand ( ![Expand icon](/files/Hgd3GKMuEEYReATv08os) ) the different categories of tasks that are shown.

After a task is performed on a data node, a new task node is created and connected to the original data node. Depending on the task, a new data node may automatically be generated that contains the resulting data. For details of individual tasks, see Task Menu.

In Figure 5, alignment was performed on the unaligned reads. Two additional nodes were added: a task node for Align reads and an output data node containing the Aligned reads.

<figure><img src="/files/5Z2tLh3MjS8jTNpEPeaz" alt=""><figcaption><p>Figure 5. Certain tasks performed on a data node generate additional data nodes. The example shows the Aligned reads node, which was generated upon alignment of the Unaligned reads node</p></figcaption></figure>

## Running a Task

To run a task, select a data node and then locate the task you wish to perform from the context sensitive menu. Mouse over to see a description of the action to be performed. Click the **specific task**, set the additional parameters (Figure 6), and click **Finish**. The task will be scheduled, the display will refresh, and the screen will return to the project's Analyses tab.

In Figure 6, the STAR aligner was selected and the choices for the aligner index and additional alignment options appeared.

<figure><img src="/files/yJ9Zi0GTv6vpJHimi5Zm" alt=""><figcaption><p><em>Figure 6. Running a STAR alignment task in Partek Flow. Dialog boxes to set the parameters appear</em></p></figcaption></figure>

Tasks that are currently running (or scheduled in the queue) appear as translucent nodes. The progress of the task is indicated by the progress bar within the task node. Hovering the mouse pointer over the node will highlight the related nodes (with a thin black outline) and display the status of the task (Figure 7).

If a task is expected to generate data nodes, expected nodes appear even before the task is completed. They will have a lighter shade of color to indicate that they have not yet been generated as the task is still being performed. Once all tasks are done, all nodes would appear in the same shade.

<figure><img src="/files/JUyjiwWWEGNhNXrBWL3F" alt=""><figcaption><p><em>Figure 7. A running task showing the progress indicator. The output data node, also in the lighter shade of color, appears even before the task completes. This enables the user to launch additional tasks while an upstream task is still in progress</em></p></figcaption></figure>

## Cancelling and Deleting Tasks

Tasks can only be cancelled by the user that started the task or by the owner of the project. Running or pending tasks can be canceled by clicking the right mouse button on the task node and then selecting **Cancel** (Figure 8). Alternatively, the task node may be selected and the **Cancel task** selected from the context sensitive menu.

<figure><img src="/files/xRVBoyQUZUbw1Nr2Veyh" alt=""><figcaption><p><em>Figure 8. Canceling a task may be done by right clicking on the running task or by selecting Cancel task in the task panel</em></p></figcaption></figure>

A verification dialog will appear (Figure 9) asking to confirm the task cancellation, the cancelled tasks will remain in the Analyses tab but will be flagged by gray x circles on the nodes (Figure 10).

<figure><img src="/files/UdFpETOUQGtwwmosyT57" alt=""><figcaption><p><em>Figure 9. Verification of Task Cancellation</em></p></figcaption></figure>

Data nodes connected to incomplete tasks are also incomplete as no output can be generated (Figure 10).

<figure><img src="/files/2NnxUwNtZk9QCr88WHYI" alt=""><figcaption><p>Figure 10. Warnings indicate that the task failed (or was cancelled) and the data node is empty</p></figcaption></figure>

To delete tasks from the project click the right mouse button on the task node and then select **Delete** (Figure 11). Alternatively, click the task node and select **Delete task** from the context sensitive menu.

<figure><img src="/files/xRVBoyQUZUbw1Nr2Veyh" alt=""><figcaption><p><em>Figure 11. A task can be deleted by right clicking on the task and selecting Delete or selecting Delete task in the context sensitive menu</em></p></figcaption></figure>

A verification dialog will appear (Figure 12). A yellow ![Caution icon](/files/iATHmX2chdikiYZ4RtLX) warning sign will show up if there some downstream tasks performed by collaborators will be affected. Deleting the tasks output files optional. If this is not selected, the task nodes will disappear from the Analyses tab but the output files will remain in the project output directory.

<figure><img src="/files/M1iA6M2cDvl2ZLC5QYGY" alt=""><figcaption><p><em>Figure 12. Verification of Task Deletion</em></p></figcaption></figure>

## Task Results and Task Actions

Selecting a task node will reveal a menu pane with two sections: Task results and Task actions (Figure 13).

<figure><img src="/files/tsGiRWtVM9kgHhWeHOih" alt=""><figcaption><p><em>Figure 13. Context sensitive menu after selecting a task node</em></p></figcaption></figure>

Items from the Task results section inform on the action performed in that node. Certain tasks generate a Task report (Figure 14), which include any tables or charts that the task may have produced.

<figure><img src="/files/JsRpVcS6wZncjvuboZ5o" alt=""><figcaption><p><em>Figure 14. An example of a Task report for the Trim bases task</em></p></figcaption></figure>

The Task details shows detailed diagnostic information about the task (Figure 15). It includes the duration and parameters of the task, lists of input and output files, and the actual commands (in the command line) that were run.

<figure><img src="/files/zJhNLY0nCfwxIE0ePn5p" alt=""><figcaption><p><em>Figure 15. An example of a Task details page for a Pre-alignment QA/QC task</em></p></figcaption></figure>

Additionally, the Task details page would contain the error logs of unsuccessful runs. The user can download the logs or send them directly to Partek. This page plays an important role in diagnosing and troubleshooting issues related to task.

Double clicking on a task node will show the Task report page. However, if no report was generated, the user will be directed to the Task details page.

In the Task actions sections, the selected task can be **Re-run w/new parameters**, and in case it is part of a pipeline that includes additional tasks after it, running the Downstream tasks is an option. Re-running tasks will result in a new layer being made in the Analyses tab.

Another action available for a task node is **Add task description** (Figure 16), which is a way to add notes to the project. The user can enter a description, which will be displayed when the mouse pointer is hovered over the task node.

<figure><img src="/files/zVNX6JIFEE4DQ7PsIYr5" alt=""><figcaption><p><em>Figure 16. Adding a task description</em></p></figcaption></figure>

## Layers

It is common for next-generation sequencing data analysis to examine different task parameters for optimization. Users may want to modify an upstream step (e.g. alignment parameters) and inspect its effect on downstream results (e.g. percent aligned reads).

The implementation of Layers makes optimizations easy and organized. Instead of creating separate nodes in a pipeline, another set of nodes with a different color is stacked on top of previous analyses (Figure 17). To see the parameters that were changed between runs, hover the mouse icon over the set of stacked task nodes and a pop-up balloon will display them. The text color signifies the layer corresponding to a specific parameter.

<figure><img src="/files/BDMFa6rK4RqhtEUYAxB0" alt=""><figcaption><p><em>Figure 17. Layers and balloon text correspond to different parameters</em></p></figcaption></figure>

Layers are formed when the same task is performed on the same data node more than once. They are also formed when a task node is selected and the **Re-run it w/new parameter**s is selected in the context sensitive menu. This will allow the users to change the options only for the selected task. The user may choose to re-run the task to which the changes have been made, as well as all the downstream tasks until the analysis is completed. To do so, select **Re-run w/new parameters, downstream tasks** from the context sensitive menu.

To select a different layer, use the left mouse button to click on any node of the desired layer. All the nodes associated with the selected layer have the same color and when clicked will be displayed on the top of the stack.

## Collapsing Tasks

Addition of task and resulting data nodes to project may lead to creation of long pipelines, extending well beyond the border of the canvas (Figure 18).

<figure><img src="/files/YhZGyW3x0qLkk9P3yz1y" alt=""><figcaption><p><em>Figure 18. Addition of tasks may extend the pipeline beyond the borders of the canvas</em></p></figcaption></figure>

In that case, the pipeline can be collapsed, to hide the steps that are (no longer) relevant. For example, once the single-cell RNA-seq data has been quantified, Single cell counts data node will be a new analysis start point, as the subsequent analyses will not focus on alignment, UMI deduplication etc. To start, **right-click** on the task node which should become a boundary of the collapsed portion of your pipeline and select **Collapse tasks** (Figure 19).

<figure><img src="/files/TZ8bwlj3gEnk1ErxyzEn" alt=""><figcaption><p><em>Figure 19. To collapse a part of the pipeline, right-click on a task node</em></p></figcaption></figure>

All the tasks on that layer will turn purple. Then **left-click** the task which should be the other boundary of the collapsed portion. All the tasks that will be collapsed will turn green and a dialog will appear (Figure 19). In the example shown in Figure 19, the tasks between Trim tags and Quantify barcodes will be collapsed. Give the collapsed section a name (up to 30 characters) and select **Save** (Figure 20).

<figure><img src="/files/NkfddWTxxPSZlKeDXhQh" alt=""><figcaption><p><em>Figure 20. Tasks within the pipeline that will be collapsed are highlighted in green. Instead of them, a single new task node will appear, with the custom label</em></p></figcaption></figure>

The collapsed portion of the pipeline is replaced by single task node, with a custom label ("Single cell preprocessing"; Figure 21).

<figure><img src="/files/vThJqdvkDy9JSA7n5LtF" alt=""><figcaption><p><em>Figure 21. "Single cell preprocessing" task node represents five collapsed tasks. Pipeline can be expanded by double clicking on the collapsed node</em></p></figcaption></figure>

To re-expand the pipeline **double click** on the task node representing the collapsed portion of the pipeline. Alternatively, **single click** on the node and select **Expand**... on the context-sensitive menu. Within the same menu, you can also preview the contents of the collapsed task by selecting **View\...** (Figure 22).

<figure><img src="/files/F8wzcR9iWamrsvp1ToUM" alt=""><figcaption><p><em>Figure 22. Options for a collapsed task (in this example: "Single cell preprocessing"): View..., Expand..., Change color</em></p></figcaption></figure>

## Downloading data

Data associated with any data node can be downloaded using the ![Download icon](/files/rl7iLD85BbMXL5pr3DYs) **Download data** link in the context sensitive menu (Figure 23). Compressed files will be downloaded to the local computer where the user is accessing the Partek Flow server. Note that bigger files (such as unaligned reads) would take longer to download. For guidance, a file size estimate is provided for each data node. These zipped files can easily be imported by the Partek® Genomics Suite® software.

<figure><img src="/files/p212kJUP9zQe03GEoGoR" alt=""><figcaption><p><em>Figure 23. Downloading the data from a data node using the task pane (an example is shown)</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# The Log Tab

The Log tab contains a table of the tasks that are running, scheduled, or those that have been completed within the Partek Flow (Figure 1). It provides an overview of the task progress, enables task management, and links to detailed reports for each task.

<figure><img src="/files/CDg54KB4vIdDgRDVeuMX" alt=""><figcaption><p><em>Figure 1. The Queue tab showing running, waiting, and done tasks</em></p></figcaption></figure>

Each row of the table corresponds to a task node in the Analyses tab. The list can be sorted according to a specific column using the sort icon ![Sort icon](/files/04C2lk7xYptkgPUqOGE0).

The Task column lists the name of the tasks. On the left of the task name is a colored circle indicating the layer of this task. The column is searchable by task name. Clicking the task name will open the Task report page. If the task did not generate a report, the link will go to the Task details page.

The User column identifies the task owner. Aside from the user who created the project, only collaborators and users with admin privileges can start tasks in a project. Clicking a name in the User column will display the corresponding User profile.

The End column shows when the task was completed. It will show the actual time for completed tasks, and the estimated time for running tasks. These estimates improve in accuracy as more tasks are completed in the current Partek Flow instance.

The Status column displays the current status of the task, such as Waiting, Running, Done, Canceled. If the task is currently running, a status progress bar will appear in the column. Once completed, the status of a task will be Done and the End column will be updated with the completion time.

A waiting task may be waiting for upstream tasks to complete ( ![Waiting T icon](/files/k6ZZZcJgmeMqAD2oYOTF) ) or waiting for more computing resources to be available ( ![Waiting R icon](/files/I7LGpxUdRgOu3ddbQ8rr) ).

The Action column contains the cancel button ( ![Cancel red x icon](/files/v517zoM0Auyo5CeSYtQz) ) while a task is queued or running. Clicking this button will cancel the task. A trash icon ( ![Trash icon](/files/6AJiF3QG5ma3UKPAIr3v) ) will appear in the Action column for completed, canceled or failed tasks, and will allow the task to be deleted from the project. Deleting a task in the Queue tab will remove the corresponding nodes in the Analyses tab. Unless the user has admin privileges, a user may only cancel and delete a task that he/she started. The User, End, and Status columns may be used to filter ( ![Filter icon](/files/a5nfgK3X8m8KwP4PQ53b) ) the table.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# The Project Settings Tab

* [Project details](#project-details)
* [Members](#members)

## Project details

The Project details section shows the Name of the project as well as (optional) a textual project Description and a Thumbnail (picture) (Figure 1).

<figure><img src="/files/MYQAFZk6XHAW2QjKKqJV" alt=""><figcaption><p><em>Figure 1. Project settings tab contains optional details on the project and controls the user accounts that are permitted to work in the current project (all the details are usage examples)</em></p></figcaption></figure>

The owner and collaborators (if any) can customize the Description and Thumbnail entries by pushing the orange **Edit project details** button (Figure 2). The fields can now be edited to:

* Rename the project (names are limited to 30 characters). The original project Name is the one selected when creating the project
* Add or change a project description (up to 2000 characters)
* Add or change a thumbnail of the project (supported formats are .jpg, .bmp, .gif, .png; the maximum size of the image file is 10 MB)

The Description accepts hyperlinks starting with "http\://" or "https\://" and if selected, will open a new tab browser to navigate to the website. This description will be also displayed to collaborators and administrators on the Partek® Flow® Home Page. **Choose File** button launches a file browser showing directory structure of the local computer from which the thumbnail image file will be uploaded. Alternatively, **Clear thumbnail** button removes the current thumbnail.

Once all the edits have been made, push **Save** to accept (or **Cancel** to reject).

<figure><img src="/files/lPATXlY4xclYJpwb4eep" alt=""><figcaption><p><em>Figure 2. Project details can be customized to rename the project, add a description of the project, or add a thumbnail (project name is an example)</em></p></figcaption></figure>

If a thumbnail has been added, it will appear on the Project details tab (Figure 3) and on the home page of Partek Flow, on the Details tab of the project.

<figure><img src="/files/eNTRSlQvPecwxBhJ3vR8" alt=""><figcaption><p><em>Figure 3. Project settings tab, example with thumbnail (image of the boškarin buffalo taken from Wikipedia)</em></p></figcaption></figure>

## Members

The Members section provides an overview of users associated with a particular project and enables project creators (owners) and administrators to add collaborators (Figure 1). A user (without administrator status) has to be specified as a collaborator in a project to be able to access the project in his/her home folder and to perform tasks.

To add a collaborator, use the **Add member** drop-down menu. The drop-down menu will list users you are collaborating with on any project on the current instance of Partek Flow. Click a user name in the drop-down list and then click the ![Green plus icon](/files/fa7wq1RxQ6oV0mx3wf1W) button to add them as a collaborator. To add a user you have not collaborated with before, type their exact username (e.g., jsmith) and click the ![Green plus icon](/files/fa7wq1RxQ6oV0mx3wf1W) button to add them. Depending on the collaborator's preference settings, an email notification may be sent to the email address associated with their user account. To delete a collaborator, select the ![Red x icon](/files/AwpurpGwRw4U9gEMUdna) next to their username (you will be asked for confirmation).

Pushing the **pencil icon** (Pencil icon]\(../../../.gitbook/assets/pencil-icon.png)) by a project member can result in two dialogs, depending on the status of the member. For a collaborator or a viewer, the pencil icon changes the member's role (e.g. from a Viewer to a Collaborator) (Figure 4).

<figure><img src="/files/gjlZ4CJh90qEupTbkU7l" alt=""><figcaption><p><em>Figure 4. Changing the role of a project member</em></p></figcaption></figure>

Moreover, the project owner can transfer the ownership to another user account (one of the accounts already available at the current instance of Partek Flow) using the New owner dropdown list (Figure 5). The previous (old) owner can remain as a project collaborator, with the help of the matching option.

<figure><img src="/files/3CFIPnbRTsNl59at87SD" alt=""><figcaption><p><em>Figure 5. Appointing a new project owner</em></p></figcaption></figure>

If email notifications are turned on for project ownership transfer, an email dialog box appears. This can be used to add additional text to the notification email body (Figure 6).

<figure><img src="/files/GvHhN1B5jBDwdNeZmahQ" alt=""><figcaption><p><em>Figure 6. Configure notification email upon transferring project ownership</em></p></figcaption></figure>


# The Attachments Tab

The Attachments tab allows the project owner to add external (i.e. non-Partek Flow) files to a project (for instance, spreadsheets, word documents, manuscripts). To attach a file, go to the **Attachments** tab (Figure 1). **Choose File** button invokes the file browser showing the directory structure of the local computer. Select the file that you want to attach and then click on the **Upload attachment** button. For security reasons, Partek Flow will not allow you to add an executable file.

<figure><img src="/files/dz6JawRcrShktu5e9hN7" alt=""><figcaption><p><em>Figure 1. Adding an attachment to a project</em></p></figcaption></figure>

All added files will be listed in the table under the Attachments tab (Figure 2). The tab will also display file sizes, the user name of the person who uploaded the file and the time it was uploaded. Note that uploaded files will count towards the total size of the project, and thus, if available, to the disk quota of the project owner.

To remove a file, click the ![Red x icon](/files/AwpurpGwRw4U9gEMUdna) icon. To download the attachment, click the ![Red download icon](/files/oC3DxVY0zdZWQ8knInKs) icon.

<figure><img src="/files/5DlC6NiMzibV8XQmTazk" alt=""><figcaption><p><em>Figure 2. Attached files listed under the Attachments button (the image is just an illustration). To remove a file, use the red cross button</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Project Management

* [Project Deletion](#project-deletion)
* [Selecting Files for Deletion](#selecting-files-for-deletion)
* [Project Import and Export](#project-import-and-export)

## Project Deletion

A project may be deleted from the Partek Flow server using the ![Cog icon](/files/WPqMDrW1LkdCKABwMDUa) button on the upper right side of the Project View page (Figure 1).

<figure><img src="/files/5hPryvNGljhidS6ROrRl" alt=""><figcaption><p><em>Figure 1. Archiving or deleting a project</em></p></figcaption></figure>

Alternatively, you can also delete your projects directly from the Homepage by clicking ![Trash icon](/files/6AJiF3QG5ma3UKPAIr3v) **Delete project** under the Actions column (Figure 2).

<figure><img src="/files/oBf9QbMMGMkidNvxP23C" alt=""><figcaption><p><em>Figure 2. Deleting project from the Home page</em></p></figcaption></figure>

## Selecting Files for Deletion

After clicking the ![Red trash icon](/files/qvcSFrVUy4o9jglwquQm) **Delete project**, a page displaying all the files associated with the project appears. Clicking the triangle ![Triangle expand icon](/files/UcRjJKZDGJvmbl3SLHeg) will expand the list. Select the files to be deleted from the server by clicking the corresponding checkboxes next to each file (Figure 3). By default, all output files generated by the project will be deleted.

<figure><img src="/files/vCotApT2xrqp2nS8EbgS" alt=""><figcaption><p><em>Figure 3. Select files to delete</em></p></figcaption></figure>

If you wish to delete the input files associated with the project, you can do that as well by clicking the **Input files** checkbox. Note that a warning icon ![Caution icon](/files/iATHmX2chdikiYZ4RtLX) appears next to input files that are used in other projects (Figure 4). These cannot be deleted until all projects associated with them are deleted.

<figure><img src="/files/pZQtRE0aZjbsWd82triI" alt=""><figcaption><p><em>Figure 4. Input files used in different projects cannot be deleted</em></p></figcaption></figure>

## Project Import and Export

Every project can be exported before it is removed from the server. By exporting old projects, you can free up some storage on your server. You can import the exported project back into Partek Flow later if needed.

### Exporting a Project

Open a project and on the analysis page, click on the gear ![Cog icon](/files/WPqMDrW1LkdCKABwMDUa) button and choose ![Purple export icon](/files/QY0DjhjebcMORJv3uwDJ) **Export project** (Figure 1). You can also export the project directly from Partek Flow home page by clicking the ![Purple export icon](/files/QY0DjhjebcMORJv3uwDJ) icon under the Action column (Figure 2).

When you export a project, you will be asked whether to include library files to export or not. If you choose **Yes**, the current version of library files used in the project will be archived and you can reproduce the result when you later import the project and re-do the analysis. However, it will make the archive size bigger. If you choose **No**, the library files will not be exported. Note that when you import the project later, you can only use the available version of needed library files to re-do the same analysis, and the results might not be the same.

### Importing a Project

The **Import project** option is under Projects drop-down menu on the top of the Partek Flow page (Figure 5). This can be accessed on any Partek Flow page.

<figure><img src="/files/iHDMQv4OLAVs52AQTYNA" alt=""><figcaption><p><em>Figure 5. Import project menu</em></p></figcaption></figure>

The input of this option is the zipped file of the exported project. Browse to the file location which can either be the Partek Flow server, a local machine, or a URL. The zip file first needs to be uploaded to the Partek Flow server (if it is not on the server already), and then Partek Flow will unpack the zip file into a project. The project name will be the same as the exported project name. If the project with the same name already exists, the imported project will have a number appended to it (e.g., ProjectName\_1).

The owner of the imported project will be the user that imported it.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing a GEO / ENA project

* [How to import a study from GEO / ENA](#how-to-import-a-study-from-geo--ena)
* [Common Issues](#common-issues)
  * [Error Message - The project did not yield any data. Double-check the project ID, or try importing the data manually](#error-message---the-project-did-not-yield-any-data-double-check-the-project-id-or-try-importing-the-data-manually)
  * [The project was imported, but the Analyses tab is empty and there are no FASTQ files](#the-project-was-imported-but-the-analyses-tab-is-empty-and-there-are-no-fastq-files)
  * [Something is missing or the import failed](#something-is-missing-or-the-import-failed)
* [FAQ](#faq)
  * [What are GEO and ENA?](#what-are-geo-and-ena)
  * [How do I know if a GEO project is also in ENA?](#how-do-i-know-if-a-geo-project-is-also-in-ena)

## How to import a study from GEO / ENA

If a project is publicly available in the Gene Expression Omnibus (GEO) and European Nucleotide Archive (ENA) databases, you can import associated FASTQ files, sample attributes, and project details automatically into Partek Flow.

* Click **Projects** at the top of the page
* Click **Import project**

<figure><img src="/files/7bCuRUFjEn1mcdPjqD6y" alt=""><figcaption><p><em>Figure 1. Importing project invoke</em></p></figcaption></figure>

* Choose **GEO / ENA project** for Select files from
* Type the BioProject ID or the GEO Accession number

<figure><img src="/files/8S3gSbLWXeolFKKWe5H0" alt=""><figcaption><p><em>Figure 2. Enter the Bioproject ID in the Import project dialog</em></p></figcaption></figure>

The format of a BioProject ID is PRJNA followed by one to six numbers (e.g., PRJNA291540). The format of a GEO Accession number is GSE followed by one to five numbers (e.g., GSE71578).

* Click **Import project** at the bottom

The **Analyses tab** will include an Unaligned reads data node once the data download has started (Figure 3). It may take a while for the download to complete depending on the size of the data. FASTQ files are downloaded from the ENA BioProject page.

<figure><img src="/files/mSt2bRyph9HhDKA3ld4Y" alt=""><figcaption><p><em>Figure 3. FASTQ files will be added as an Unaligned reads data node in the Analyses tab</em></p></figcaption></figure>

## Common Issues

### Error Message - The project did not yield any data. Double-check the project ID, or try importing the data manually

If the study is not publicly available in both GEO and ENA, project import will not succeed.

### The project was imported, but the Analyses tab is empty and there are no FASTQ files

If there is an ENA project, but the FASTQ files are not available through ENA, the project will be created, but data will not be imported.

### Something is missing or the import failed

A variety of other issues and irregularities can cause imports to not succeed or partially succeed, including, but not limited to, a BioProject having multiple associated GSE IDs, incomplete information on the GEO or ENA page, and either the GEO or ENA project not being publicly available.

## FAQ

### What are GEO and ENA?

The Gene Expression Omnibus (GEO) and the European Nucleotide Archive (ENA) are web-accessible public repositories for genomic data and experiments. Access and learn more about their resources at their respective websites:

* GEO - <https://www.ncbi.nlm.nih.gov/geo/>
* ENA - <https://www.ebi.ac.uk/ena>

### How do I know if a GEO project is also in ENA?

You can search ENA using the GEO ID (e.g., GSE71578) to check if there is a matching ENA project (Figure 6).

<figure><img src="/files/l7cr8pwUhmMCwT3tPK41" alt=""><figcaption><p><em>Figure 4. Searching ENA using a GEO ID</em></p></figcaption></figure>

<figure><img src="/files/CItdPwMXuiydGatsDi0v" alt=""><figcaption><p><em>Figure 5. ENA Study page</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Bulk RNA-Seq

This tutorial gives an overview of RNA-Seq analysis with Partek Flow. It will guide you through creating an RNA-Seq analysis pipeline. The goals of the analysis are to create a list of differentially expressed genes, visualize these gene expression signatures by hierarchical clustering, and interpret the gene lists using gene ontology (GO) enrichment.

This tutorial will illustrate:

* [Importing the tutorial data set](/partek-flow/tutorials/bulk-rna-seq/importing-the-tutorial-data-set)
* [Adding sample attributes](/partek-flow/tutorials/bulk-rna-seq/adding-sample-attributes)
* [Running pre-alignment QA/QC](/partek-flow/tutorials/bulk-rna-seq/running-pre-alignment-qa-qc)
* [Trimming bases and filtering reads](/partek-flow/tutorials/bulk-rna-seq/trimming-bases-and-filtering-reads)
* [Aligning to a reference genome](/partek-flow/tutorials/bulk-rna-seq/aligning-to-a-reference-genome)
* [Running post-alignment QA/QC](/partek-flow/tutorials/bulk-rna-seq/running-post-alignment-qa-qc)
* [Quantifying to an annotation model](/partek-flow/tutorials/bulk-rna-seq/quantifying-to-an-annotation-model)
* [Filtering features](/partek-flow/tutorials/bulk-rna-seq/filtering-features)
* [Normalizing counts](/partek-flow/tutorials/bulk-rna-seq/normalizing-counts)
* [Exploring the data set with PCA](/partek-flow/tutorials/bulk-rna-seq/exploring-the-data-set-with-pca)
* [Performing differential expression analysis with DESeq2](/partek-flow/tutorials/bulk-rna-seq/performing-differential-expression-analysis-with-deseq2)
* [Viewing DESeq2 results and creating a gene list](/partek-flow/tutorials/bulk-rna-seq/viewing-deseq2-results-and-creating-a-gene-list)
* [Viewing a dot plot for a gene](/partek-flow/tutorials/bulk-rna-seq/viewing-a-dot-plot-for-a-gene)
* [Visualizing gene expression in Chromosome view](/partek-flow/tutorials/bulk-rna-seq/visualizing-gene-expression-in-chromosome-view)
* [Generating a hierarchical clustering heatmap](/partek-flow/tutorials/bulk-rna-seq/generating-a-hierarchical-clustering-heatmap)
* [Performing biological interpretation](/partek-flow/tutorials/bulk-rna-seq/performing-biological-interpretation)
* [Saving and running a pipeline](/partek-flow/tutorials/bulk-rna-seq/saving-and-running-a-pipeline)

## Description of the Data Set

This tutorial uses a subset of the data set published in Xu et al. 2013 (PMID: 23902433). In the experiment, mRNA was isolated from HT29 colon cancer cells treated with the drug 5-aza-deoxy-cytidine (5-aza) at three different doses: 0μM (control), 5μM, or 10μM. The mRNA was sequenced using Illumina HiSeq (paired end reads). The goal of the experiment was to identify differentially expressed genes between the different treatment groups.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing the tutorial data set

The tutorial data set includes 9 samples equally divided into 3 treatment groups. Sequencing was performed by an Illumina HiSeq (paired-end reads), but the workflow can be easily adapted for data generated by other sequencers. Each sample has 2 fastq files for a total of 18 fastq files.

You can obtain the tutorial data set through Partek Flow.

* Click your **avatar**
* Click **Settings** in the drop-down menu (Figure 1)

<figure><img src="/files/CNmxpC61lAh0Fx58mPBT" alt=""><figcaption><p><em>Figure 1. Location of the Settings link on the main page of Partek Flow</em></p></figcaption></figure>

At the top of the System information page, there is a section labeled Tutorial data (Figure 2).

<figure><img src="/files/jdvEGPfVyrmrkgzJUg96" alt=""><figcaption><p><em>Figure 2. Click RNA-Seq 5-AZA to download the sample data set</em></p></figcaption></figure>

* Click **RNA-Seq 5-AZA** to download the bulk RNA-Seq tutorial data set

A new project will be created and you will be directed to the Analyses Tab. The data will be downloaded automatically (Figure 3) and imported into your project. Because this is a tutorial project, there is no need to click on Add data as it will be done automatically.

<figure><img src="/files/LDjTxU1vb6ySGKyMMrNM" alt=""><figcaption><p><em>Figure 3. A new project will be automatically created and populated with the tutorial data set</em></p></figcaption></figure>

At first the project is empty, but the file download will start automatically in the background. You can wait a few minutes then refresh your browser or you can monitor the download progress using the Queue.

* Click **Queue**
* Click **View Queued Tasks** in the drop-down menu

The Queued tasks page will open (Figure 4).

<figure><img src="/files/pTm68YBVjOJDtlswa6St" alt=""><figcaption><p><em>Figure 4. Viewing download progress in the Queued tasks page</em></p></figcaption></figure>

* Click **Projects**
* Click **RNA-Seq 5-AZA** in the drop-down menu

The Analyses tab will open (Figure 5). If the download has completed, you will see a blue circle titled mRNA.

The project name can be changed at any time by clicking the *Project settings* tab.

<figure><img src="/files/TBg1P50avzZWvxvECuBV" alt=""><figcaption><p><em>Figure 5. Analyses tab showing a mRNA data node</em></p></figcaption></figure>

Once the download completes, the sample table will appear in the Metadata tab.

* Click the **Metadata** tab. The Metadata tab includes the sample table with the names of each imported sample (Figure 6).

<figure><img src="/files/iVl3BNbUtNNGKxkBpC4X" alt=""><figcaption><p><em>Figure 6. Sample table in the Data tab after downloading the tutorial data set</em></p></figcaption></figure>

In the next section of the tutorial, we will add a sample attribute that indicates the treatment group of each sample.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Adding sample attributes

Attributes describe samples. Examples of sample attributes include treatment group, age, sex, and time point. Attributes can be added individually in the Metadata tab or in bulk [using a text file](/partek-flow/quick-start-guide#the-metadata-tab). In this tutorial, we will add one attribute, 5-AZA Dose, manually.

* Click the **Metadata** tab
* Click **Manage** under Sample attributes (Figure 1)

<figure><img src="/files/bZQKt4ET73dAAdHRZHHN" alt=""><figcaption><p><em>Figure 1. Adding sample attributes manually</em></p></figcaption></figure>

* Click **Add new attribute** (Figure 2)

<figure><img src="/files/CW5KP8xICNa3EmvIBQ19" alt=""><figcaption><p><em>Figure 2. Adding a new attribute</em></p></figcaption></figure>

* To configure a new attribute, at Name, type in **5-AZA Dose** as the name of the attribute
* Click **Add** to add 5-AZA Dose as a categorical, project-specific attribute (Figure 3)

<figure><img src="/files/ELt404Ty4RJOFbtXTd4E" alt=""><figcaption><p><em>Figure 3. Configuring a new attribute</em></p></figcaption></figure>

* Name the first New category **0uM**
* Click the plus icon to add category (Figure 4)
* Repeat for two additional categories, **5uM** and **10uM** (Figure 5)

<figure><img src="/files/00QAvYSQlmKaABMyo7mm" alt=""><figcaption><p><em>Figure 4. Creating attribute category</em></p></figcaption></figure>

<figure><img src="/files/nFbo5525f60ev82sC2a1" alt=""><figcaption><p><em>Figure 5. All category attributes for 5-AZA Dose</em></p></figcaption></figure>

* Click **Back to metadata tab**

The data table now includes an Attribute column for 5-AZA Dose (Figure 6). Next, we need to assign samples attribute categories for 5-AZA Dose.

<figure><img src="/files/YEmeAvu4PqsbOUDDEhNb" alt=""><figcaption><p><em>Figure 6. Data table updated with column for 5-AZA Dose</em></p></figcaption></figure>

* Select Assign values

The option to edit the 5-AZA Dose field for each sample will appear as a drop-down menu (Figure 7).

<figure><img src="/files/PdgTZNDdHhq8DXFnQ49y" alt=""><figcaption><p><em>Figure 7. Dropdown menu to select treatment for each sample</em></p></figcaption></figure>

* Select the 5-AZA Dose text box for a sample to bring up a drop-down menu with the 5-AZA Dose attribute categories (0uM, 5uM, 10uM)
* Use the drop-down menus to add a treatment group for each sample

The first three samples (SRR592573-5) should be **0uM**, the next three samples (SRR592576-8) should be **5uM**, and the final three samples (SRR592579-81) should be 10uM (Figure 8).

* Click **Apply changes**

<figure><img src="/files/ApxMD9DNRkoptvjIT8JO" alt=""><figcaption><p><em>Figure 8. Add 5-AZA dose to each sample as shown</em></p></figcaption></figure>

The data table will now show each a 5-AZA Dose attribute for each sample.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Running pre-alignment QA/QC

With attributes added, we can begin building our pipeline.

* Click the Analyses tab

In the Analysis tab, data are represented as circles, termed data nodes. One data node, mRNA, should be visible in the Analysis tab (Figure 1).

<figure><img src="/files/p0RKqqvpCv9qLpQdyLzt" alt=""><figcaption><p><em>Figure 1. Data are represented as circles</em></p></figcaption></figure>

* Click the mRNA node

Clicking a data node brings up the context-sensitive task menu on the right with tasks that can be performed on the data node (Figure 2).

<figure><img src="/files/NaPWSM4691gTc0nYILhd" alt=""><figcaption><p><em>Figure 2. Select a data node to open the context-sensitive task menu</em></p></figcaption></figure>

Pre-alignment QA/QC assesses the quality of the unaligned reads and will help us determine whether trimming or filtering is necessary.

* Click **Pre-alignment QA/QC** in the QA/QC section of the task menu
* Click **Finish** to run the task with default settings

Running a task creates a task node, e.g. the blue rectangle labeled Pre-alignment QA/QC (Figure 3), which contains details on the task and a report. While tasks have been queued or are in progress they have a lighter color. Any output nodes that the task will generate are also displayed in a lighter color until the task completes. Once the task begins running, a progress bar is displayed on the task node.

<figure><img src="/files/oPRsa4WkK6Vm8mYJZutK" alt=""><figcaption><p><em>Figure 3. Tasks are represented as rectangles</em></p></figcaption></figure>

* Click the **Pre-alignment QA/QC** node

The context-sensitive task menu (Figure 4) shows the option to view the Task report and the Task details. You can also access a task report by double-clicking on a task node.

<figure><img src="/files/SdxbRLQnXREgs5vi4gx7" alt=""><figcaption><p><em>Figure 4. Navigating to the task report using the context-sensitive menu</em></p></figcaption></figure>

* Click **Task report**

Pre-aligment QA/QC provides information about the sequencing quality of unaligned reads (Figure 5). Both project level summaries and sample-level summaries are provided.

<figure><img src="/files/yP9u8z7an9HspxZ1Xyav" alt=""><figcaption><p><em>Figure 5. Viewing the pre-alignment QA/QC report</em></p></figcaption></figure>

* Click sample **SSR592573** in the data table of the report to open its sample-level report

The Average base quality score per position graph (Figure 6) gives the average Phred score for each position in the reads.

<figure><img src="/files/djcjitcSLF2mJyCsXcRo" alt=""><figcaption><p><em>Figure 6. Average base quality score per position for sample SRR592573</em></p></figcaption></figure>

A Phred score is a measure of base call accuracy with a higher score indicating greater accuracy.

| Phred Quality Score | Probability of incorrect base call | Base call accuracy |
| :-----------------: | :--------------------------------: | :----------------: |
|          20         |              1 in 100              |         99%        |
|          30         |              1 in 1000             |        99.9%       |
|          40         |             1 in 10,000            |       99.99%       |
|          50         |            1 in 100,000            |       99.999%      |

By convention, a score above 20 is considered adequate. As you can see, the standard error bars in the graph show that some reads have quality scores below 20 for some of their base pair calls near the 3' end.

Based on the results of Pre-alignment QA/QC, while most of the reads are high quality, we will need to perform read trimming and filtering. For more information about the information included in the task report, please see the [Pre-alignment QA/QC user guide](/partek-flow/user-manual/task-menu/qa-qc/pre-alignment-qa-qc).

* Click **RNA-Seq 5-AZA** (the project name) to return to the Analyses tab

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Trimming bases and filtering reads

Based on pre-alignment QA/QC, we need to trim low quality bases from the 3' end of reads.

* Click the **Unaligned reads** data node
* Click **Pre-alignment tools** in the task menu
* Click **Trim bases** (Figure 1)

<figure><img src="/files/dONExtdjS4I3u5URwuvx" alt=""><figcaption><p><em>Figure 1. Invoking the Trim bases task</em></p></figcaption></figure>

By default, Trim bases removes bases starting at the 3' end and continuing until it finds a base pair call with a Phred score of equal to or greater than 35 (Figure 2).

* Click **Finish** to run Trim bases with default settings

<figure><img src="/files/yBhA9li63vOYySWEV1Hw" alt=""><figcaption><p><em>Figure 2. Configure the Trim bases task</em></p></figcaption></figure>

The Trim bases task will generate a new data node, Trimmed reads (Figure 3). We can view the task report for Trim bases by double-clicking either the Trim bases task node or the Trimmed reads data node or choosing Task report from the task menu.

<figure><img src="/files/whQWos2eOlKEnY3ynmJT" alt=""><figcaption><p><em>Figure 3. A task and a data node are created from the Trim bases task</em></p></figcaption></figure>

* Double-click the **Trimmed reads** data node to open the task report

The report shows the percentage of trimmed reads and reads removed in a table and two graphs (Figure 4).

<figure><img src="/files/CrDsNc4lBf8c4GLE40v0" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/AKxCgBBX0qAywIGAveGY" alt=""><figcaption><p><em>Figure 4. Results of the Trim bases task</em></p></figcaption></figure>

The results are fairly consistent across samples with \~2% of reads untrimmed, \~86% trimmed, and \~12% removed for each. The average quality score for each sample is increased with higher average quality scores at the 3' ends.

* Click **RNA-Seq 5-AZA** to return to the Analyses tab

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Aligning to a reference genome

With our reads trimmed, we now have high-quality reads for each sample. The next step is to align the reads to a reference genome. Alignment matches each of the short sequencing reads to a location in the reference genome.

* Click **Trimmed reads**
* Click **Aligners** in the task menu to display available aligners (Figure 1)

<figure><img src="/files/QRfQCtogSiQOgqwjQXoE" alt=""><figcaption><p><em>Figure 1. A variety of different aligners are available in Partek Flow. In this tutorial we will use STAR aligner, a popular choice for aligning RNA-Seq data.</em></p></figcaption></figure>

Partek Flow offers a variety of different aligners. Mouse over any option for a short description. For this tutorial, we will use STAR, a fast and accurate aligner commonly used for RNA-Seq data. For more information about STAR and the other aligners, please consult the [Aligners ](/partek-flow/user-manual/task-menu/aligners)user guide.

* Click **STAR**

The STAR aligner options allow us to select the genome build (assembly) and index. For this tutorial, our data set contains only reads that map to chromosome 22 to minimize the time required for resource-intensive tasks, such as alignment.

* Click **Finish** to run with **hg38** selected for Assembly and **Whole genome** for the Aligner index (Figure 2)

<figure><img src="/files/l9M2hYCL5K5IP97qjCH5" alt=""><figcaption><p><em>Figure 2. Configuring the STAR aligner</em></p></figcaption></figure>

Alignment is a resource-intensive task and may take around 20 minutes to complete, even when mapping only reads from a single chromosome. Task and data nodes that have been queued, but not completed, are shown in a lighter color than completed tasks (Figure 3).

<figure><img src="/files/bgB707zTjTiMXJKh6rx4" alt=""><figcaption><p><em>Figure 3. Queued tasks and their output data nodes are shown in a lighter color</em></p></figcaption></figure>

The Align reads task generates an Aligned reads data node once complete. You can wait for the alignment task to finish or you can continue building the pipeline while the results of alignment are pending; additional tasks can be added to the pipeline and queued before the current task has completed.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Running post-alignment QA/QC

After alignment has completed, we can view the quality of alignment by performing post-alignment QA/QC.

* Click the **Aligned reads** data node
* Click **QA/QC** inthe task menu
* Click **Post-alignment QA/QC** from the QA/QC section of the task menu (Figure 1)

<figure><img src="/files/fP2N5sPqthsWC7duQ7Th" alt=""><figcaption><p><em>Figure 1. Invoking Post-alignment QA/QC</em></p></figcaption></figure>

A Post-alignment QA/QC task node will be generated (Figure 2).

<figure><img src="/files/gEbp34Qn7UltVo4GzA1V" alt=""><figcaption><p><em>Figure 2. Post-alignment QA/QC task node</em></p></figcaption></figure>

* Double-click the **Post-alignment QA/QC** task node to view the task report

Similar to the Pre-alignment QA/QC task report, general quality information about the whole data set is displayed and sample-level reports can be opened by clicking a sample name in the table.

The first two graphs in the report (Figure 3) show the alignment breakdown and total reads in each sample.

<figure><img src="/files/4DNxweMXhlP14dpNPlw2" alt=""><figcaption><p><em>Figure 3. Alignment statistics for the data set</em></p></figcaption></figure>

From these graphs, we can see that more than 95% of reads were aligned, but the total number of reads for each sample varies. Normalizing for the variability in total read counts will be addressed in a later section of the tutorial.

For more information about the graphs and information presented in the Post-alignment QA/QC task report, see the [Post-alignment QA/QC](/partek-flow/user-manual/task-menu/qa-qc/post-alignment-qa-qc) user guide.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Quantifying to an annotation model

RNA-Seq uses the number of sequencing reads per gene or transcript to quantify gene expression. Once reads are aligned to a reference genome, we need to assign each read to a known transcript or gene to give a read-count per transcript or gene.

* Click the **Aligned reads** data node
* Click **Quantification** in the task menu

We will use Partek E/M to quantify reads to an annotation model in this tutorial. For more information about the other quantification options, please see the [Quantification](/partek-flow/user-manual/task-menu/quantification) user guide.

* Click **Quantify to an annotation model (Partek E/M)** (Figure 1)

<figure><img src="/files/TNiCweaDiCEDBY13BvX6" alt=""><figcaption><p><em>Figure 1. Invoking Quantify to an annotation model (Partek E/M)</em></p></figcaption></figure>

We will use the default options for quantification. To learn more about the different options, please see the [Quantify to annotation model (Partek E/M)](/partek-flow/user-manual/task-menu/quantification/quantify-to-annotation-model-partek-em) user guide.

* Choose Ensembl Transcripts release 105 annotation from the Annotation model drop-down menu (you may need to download it first, via [Library File Management](/partek-flow/user-manual/settings/components/library-file-management/library-file-management-page) or click *Add annotation model* from the bottom of the drop-down)
* Click **Finish** (Figure 2)

<figure><img src="/files/M2QFMd7CQ3voRbJZDUIL" alt=""><figcaption><p><em>Figure 2. Configuring Quantify to annotation model (Partek E/M)</em></p></figcaption></figure>

The Quantify to annotation model task node outputs two data nodes, Gene counts and Transcript counts (Figure 3).

<figure><img src="/files/0Dgl65ATOXPtNoyESDbS" alt=""><figcaption><p><em>Figure 3. The Quantify to annotation model task produces annotation model dependent outputs (e.g. Gene counts and Transcript counts node)</em></p></figcaption></figure>

To view the results of quantification, double-click an output data node to view the task report.

The task report details the number of reads within exons, introns, and intergenic regions. For detailed information about the quantification results, see the [Quantify to annotation model (Partek E/M)](/partek-flow/tutorials/bulk-rna-seq/quantifying-to-an-annotation-model) user guide.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Filtering features

Low expression genes may be indistinguishable from noise and will decrease the sensitivity of differential expression analysis.

* Click the **Gene counts** node
* Click **Filtering** in the task menu
* Click **Filter features** (Figure 1)

<figure><img src="/files/Pvu8j2GIKqoBamAUnax8" alt=""><figcaption><p><em>Figure 1. Selecting Filter features</em></p></figcaption></figure>

* Click **Noise reduction filter**
* Set the filter to **maximum <= 1**
* Click **Finish** (Figure 2)

<figure><img src="/files/CBFAvyoGs9N66fEHnry3" alt=""><figcaption><p><em>Figure 2. Filtering low expressed genes</em></p></figcaption></figure>

A new Filtered counts data node will be created to que additional tasks (Figure 3) like Normalization.

<figure><img src="/files/xnGdAYvRaSl3GJTjxs92" alt=""><figcaption><p><em>Figure 3. Filtered counts node</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Normalizing counts

Because different samples have different total numbers of reads, it would be misleading to calculate differential expression by comparing read count numbers for genes across samples without normalizing for the total number of reads.

* Click the **Filtered counts** data node
* Click **Normalization and scaling** in the task menu
* Click **Normalization** (Figure 1)

<figure><img src="/files/adq1jJGCf8PY0NH4PPy8" alt=""><figcaption><p><em>Figure 1. Invoking Normalization task</em></p></figcaption></figure>

The Normalization task set up page will open (Figure 2).

<figure><img src="/files/8YURULRyBWWsCEi7stR2" alt=""><figcaption><p><em>Figure 2. Normalization task set up page</em></p></figcaption></figure>

Normalization can be performed by samples or by features. By samples is selected by default; this is appropriate for the tutorial data set.

Available normalization methods are listed in the left-hand panel. For more information about these options, please see the [Normalize counts](/partek-flow/tutorials/bulk-rna-seq/normalizing-counts) user guide.

For this tutorial, we will use the recommended default normalization settings.

* Select ![Recommended button](/files/549LQxqP0BrU6VF4X30t)

This adds the Median ratio normalization method, which is suitable for performing differential expression analysis using DESeq2 (Figure 3).

<figure><img src="/files/IMP0HWyI7IxAtSpnHkg8" alt=""><figcaption><p><em>Figure 3. Select recommended normalization method</em></p></figcaption></figure>

* Click **Finish** to perform normalization

A Normalize counts task node and a Normalized counts data node are added to the pipeline (Figure 4)

<figure><img src="/files/9jCzyjBIMQiww8thP399" alt=""><figcaption><p><em>Figure 4. Normalize counts task node and Normalized counts data node</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Exploring the data set with PCA

The principal components analysis (PCA) scatter plot allows us to visualize similarities and differences between the samples in a data set.

* Click the **Normalized counts** data node
* Click **Exploratory analysis** in the task menu
* Click **PCA**
* Click **Finish** to run PCA with the default options

A PCA task node and a PCA data node will be added to the pipeline (Figure 1)

<figure><img src="/files/YbzOFzWjGn629lUSldrP" alt=""><figcaption><p><em>Figure 1. PCA task node and data node are added to the pipeline</em></p></figcaption></figure>

* Double click the **PCA** data node to open the PCA scatter plot and corresponding information in the Data Viewer (Figure 2)

<figure><img src="/files/IEyhw2BpBGi4z0X3a1N6" alt=""><figcaption><p><em>Figure 2. Viewing the PCA plot in Data Viewer</em></p></figcaption></figure>

In the Data Viewer, select the 3D PCA scatter plot then click **Style** under Configure and set the Color by drop-down to **5-AZA Dose**. The scatter plot shows each sample as a sphere, colored by treatment group, in a three dimensional plot. The x, y, and z axes are the first three principal components. The percentage of total variance explained by each is listed next to the axis label. The size of each axis is determined by the variance along that axis. The plot is fully interactive; it can be rotated and points selected.

Here, we can see that samples separate based on treatment, but there is noticeable separation within treatment groups, particularly the 0μM and 10μM treatment groups.

For more detailed information about the PCA scatter plot, please see the [PCA](/partek-flow/user-manual/task-menu/exploratory-analysis/pca) user guide.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Performing differential expression analysis with DESeq2

After normalizing the data, we can perform differential analysis to identify genes that are differentially expressed based on treatment.

* Click the **Normalized counts** node
* Click **Statistics** in the task menu
* Click **Differential analysis** in the task menu (Figure 1)

<figure><img src="/files/a6QDmtot2SrHf88S6ohx" alt=""><figcaption><p><em>Figure 1. Navigating to the differential analysis options</em></p></figcaption></figure>

Select the appropriate differential analysis method (Figure 2). In this tutorial we are going to use DESeq2, but Partek Flow offers a number of alternatives. Hover the mouse over the <img src="/files/mUWsVjOIC84ncxeITnVa" alt="Info tip" data-size="line"> symbol for more information on each differential analysis method, or see our [Differential Analysis](/partek-flow/user-manual/task-menu/differential-analysis) user guide for a more in-depth look.

<figure><img src="/files/lXdJgyhCPkYlXnNf2C2E" alt=""><figcaption><p><em>Figure 2. Select the method for differential analysis from the options provided.</em></p></figcaption></figure>

* Check **5-AZA Dose** and click **Add factors** to add the attribute to the statistical model.

<figure><img src="/files/Jv3utAAqA3UGtXHw4vCE" alt=""><figcaption><p><em>Figure 3. Selecting attributes to include in analysis</em></p></figcaption></figure>

* Select **Next** to continue with 5-AZA Dose as the selected attribute

The Comparisons page will open (Figure 4).

<figure><img src="/files/T2UlT4tI7FMvROL52NHv" alt=""><figcaption><p><em>Figure 4. The Comparison selector allows multiple comparisons to be designed and added</em></p></figcaption></figure>

It is easiest to think about comparisons as the questions we are asking. In this case, we want to know what are the differentially expressed genes between untreated and treated cells. We can ask this for each dose individually and for both collectively.

The upper box will be the numerator and the lower box will be the denominator in the comparison calculation so we will select the 0μM control in the lower box.

* Drag 5μM to the upper box
* Drag 0μM to the lower box
* Click **Add comparison** to add **5μM vs. 0μM** to the comparison table (Figure 5)

<figure><img src="/files/9ZywusfMWapomlb2TmE1" alt=""><figcaption><p><em>Figure 5. Designing a comparison to add</em></p></figcaption></figure>

* Repeat to create comparisons for **10μM vs. 0μM** and **10μM,5μM vs. 0μM** (Figure 6)

<figure><img src="/files/xX80oZrBrR197JaUvE8p" alt=""><figcaption><p><em>Figure 6. Comparisons for 5uM vs. 0uM, 10uM vs. 0uM, and 5uM:10uM vs. 0uM have been added</em></p></figcaption></figure>

* Click **Finish** to perform DESeq2 as configured

A DESeq2 task node and a DESeq2 data node will be added to the pipeline (Figure 7).

<figure><img src="/files/PD56GEEoFcQutgO1DH0F" alt=""><figcaption><p><em>Figure 7. DESeq2 task node and data node</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Viewing DESeq2 results and creating a gene list

Once we have performed DESeq2 to identify differentially expressed genes, we can create a list of significantly differentially expressed genes using cutoff thresholds.

* Double click the **DESeq2** data node to open the task report

The task report shows genes on rows and the results of the DESeq2 on columns (Figure 1).

<figure><img src="/files/wrO3SjPU1r8nBgUw3yFK" alt=""><figcaption><p><em>Figure 1. Viewing the DESeq2 task report</em></p></figcaption></figure>

To get a sense of what filtering thresholds to set, we can view a volcano plot for a comparison.

* Click ![Volcano icon](/files/keaDTokBfjJKEUxgY2qO) next to the 5uM vs. 0uM comparison

A volcano plot will open showing p-value on the y-axis and fold-change on the x-axis (Figure 2). If the gene labels are on (not shown), click on the plot to turn them off.

<figure><img src="/files/KkRAm698EO8JTK7PVOq1" alt=""><figcaption><p><em>Figure 2. Viewing DESeq2 results with a volcano plot</em></p></figcaption></figure>

Thresholds for the cutoff lines are set using the Statistics card (Left panel > Configure > Statistics). The default thresholds are |2| for the X axis and 0.05 for the Y axis.

* Switch to the browser tab showing the DESeq2 report
* Click **FDR step up**
* Click the triangle next to FDR step up to open the FDR step up options
* Leave **All contrasts** selected
* Set the cutoff value to **0.05**. Hit **Enter**.

This will include genes that have a FDR step up value of less than or equal to 0.05 for all three contrasts, 5μM vs. 0μM, 10μM vs. 0μM and 5μM:10μM vs. 0μM. FDR step up is the false discovery rate adjusted p-value used by convention in microarray and next generation sequencing data sets in place of unadjusted p-value.

* Click **Fold-change**
* Click the triangle next to Fold-change to open the Fold-change options
* Leave **All contrasts** selected
* Set to From **-2** to **2** with **Exclude range** selected. Hit **Enter**.

Note that the number of genes that pass the filter is listed at the top of the filter menu next to Results: and will update to reflect any changes to the filter. Here, 28 genes pass the filter (Figure 3). Depending on your settings, the number may be slightly different.

<figure><img src="/files/7taEfBgtZcRDxlwZvQ0X" alt=""><figcaption><p><em>Figure 3. Applying filters to the DESeq2 results</em></p></figcaption></figure>

* Click <img src="/files/evhbH8KfAIikNXm9qrLV" alt="Generate filtered node button" data-size="line"> to create a data node with only the genes that pass the filter

This creates a Filter list task node and a Filtered feature list data node (Figure 4).

<figure><img src="/files/PMoBYpVZ2obcRqefUnAI" alt=""><figcaption><p><em>Figure 4. Filter list and a new Feature list node are added to the pipeline</em></p></figcaption></figure>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Viewing a dot plot for a gene

In addition to the volcano plot showing all genes, we can view expression levels of each gene on a dot plot.

* Double-click the **Filtered feature list** data node to open the task report
* Click the **FDR step up** header in the 5uM vs. 0uM section to sort by ascending FDR step up

In the task report table, there is a column labeled View with three icons ![Chromosome icon](/files/ASUOKINSvxmV7WaE9E3F) ![Scatter plot icon](/files/lVLROyObYQRWa25mr62I) ![Notebook icon](/files/KOPT8zhCgePaYKmxR9Ph) in each row.

* Select ![Scatter plot icon](/files/lVLROyObYQRWa25mr62I) to open a dot plot for the gene KLHDC7B

The dot plot for KLHDC7B (Figure 1) shows each sample as a point with normalized reads count on the y-axis. Samples are separated and colored by treatment group.

<figure><img src="/files/2NpdQhxlGSm1iZxm9YZR" alt=""><figcaption><p><em>Figure 1. Dot plot for KLHDC7B</em></p></figcaption></figure>

For more information about the dot plot, please see the [Dot Plot](/partek-flow/user-manual/visualizations/dot-plot) user guide. To return to the DESeq2 report, switch to the browser table with filtered feature list.

In the DESeq2 report, you can select ![Notebook icon](/files/KOPT8zhCgePaYKmxR9Ph) to view additional information about the statistical results for a gene or select ![Chromosome icon](/files/ASUOKINSvxmV7WaE9E3F) to view the region in Chromosome View. Chromosome View is discussed in the next section of the tutorial.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Visualizing gene expression in Chromosome view

From the DESeq2 task report, we can browse to any gene in the Chromosome view.

* Click ![Chromosome icon](/files/ASUOKINSvxmV7WaE9E3F) in the KLHDC7B row to open Chromosome view (Figure 1)

<figure><img src="/files/7Qv18AMLQ091QZQ2GdVo" alt=""><figcaption><p><em>Figure 1. Browsing to a location in Chromosome View</em></p></figcaption></figure>

A new tab will open showing KLHDC7B in the Chromosome view (Figure 2).

<figure><img src="/files/qsOoScNozcCfbCKPTcrd" alt=""><figcaption><p><em>Figure 2. Viewing in Chromosome view</em></p></figcaption></figure>

Chromosome View shows reference genome, annotation, and data set information together aligned at genomic coordinates.

Each track has Configure track ![Configure track icon](/files/c4TFNDha8KPNb0GJYgQ5) and Move track ![Move track icon](/files/Uv2otjdqHt0TKnlS0E3d) buttons that can be used to modify each track.

The top track shows average number of total count normalized reads for each of the three treatment groups in a stacked histogram. The second track shows the Ensembl annotation.

We can add tracks from any data node using Select Tracks.

* Click **Select tracks**

A pop-up dialog showing the pipeline allows us to choose which data to display as tracks in Chromosome view (Figure 3).

<figure><img src="/files/SWKa4eNG2YHImeG1ZTKX" alt=""><figcaption><p><em>Figure 3. Choosing tracks to display in Chromosome view</em></p></figcaption></figure>

* Click **Reads pileup** under Aligned reads on the left-hand side of the dialog
* Click **Display selection** to make the change

The reads pileup track is now included (Figure 4).

<figure><img src="/files/UtQS7fzN3fPZnIawOGbu" alt=""><figcaption><p><em>Figure 4. The reads pileup track shows every read in its aligned position on the reference genome</em></p></figcaption></figure>

To learn more about Chromosome view, please consult the [Chromosome View](/partek-flow/user-manual/visualizations/chromosome-view) user guide.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Generating a hierarchical clustering heatmap

To check how well our list of differentially expressed genes distinguishes one treatment group from another, we can perform hierarchical clustering based on the gene list. Clustering can also be used to discover novel groups within your data set, identify gene expression signatures distinguishing groups of samples, and to identify genes with similar patterns of gene expression across samples.

* Click the **Filtered feature list** data node
* Click **Exploratory analysis in** the task menu
* Click **Hierarchical clustering** (Figure 1)

![Figure 1. Invoking Hierarchical clustering task](/files/0BqWbJefwXgsKe4q9a4m)

The *Hierarchical clustering* menu will open (Figure 2). Hierarchical clustering can be performed with a heatmap or bubble map plot. **Cluster** must be selected under *Ordering* for both *Feature order* and *Sample order* if both the features (columns) and samples (rows) are to be clustered.

![Figure 2. Configuring Hierarchical clustering](/files/ogC9P4mdwWejtXoXPgfb)

* Click **Finish** to run with default settings

A *Hierarchical clustering* task node will be added to the pipeline (Figure 3).

![Figure 3. Hierarchical clustering task node](/files/D6ZLDHX8Ly2Vw1RD1qZH)

* Double-click the **Hierarchical clustering / heatmap** task node to view the heatmap

The *Dendrogram view* will open showing a heatmap with the hierarchical clustering results (Figure 4).

![Figure 4. Viewing the hierarchical clustering heatmap](/files/NJTCLsFt2NBBVBvWFQDP)

Samples are shown on rows and genes on columns. Clustering for samples and genes is shown through the dendrogram trees. More similar samples/genes are separated by fewer branch points of the dendrogram tree.

The heatmap displays standardized expression values with a mean of zero and standard deviation of one.

The heatmap can be customized to improve data visualization using the menu on the *Configuration* panel on the left.

* Click *Annotations* under *Configure* section in the left panel
* Select **5-AZA Dose** from *Row annotation* drop-down
* Click *Axes* under *Configure*
* Change *Column labels* **Feature** to gene\_name using the drop-down

Samples are now labeled with their *5-AZA Dose* group and column labels are Gene names (Figure 5).

<figure><img src="/files/fx0P9lGHNVJ8pTzdrhpW" alt=""><figcaption><p><em>Figure 5. Samples labeled with their 5-AZA Dose group</em></p></figcaption></figure>

Samples from the 5μM and 10μM groups are more similar to each other than to the 0μM group.

We can save the heatmap as a publication-quality image.

* Click the **Export image** icon in the top right corner of the plot <img src="/files/VmRrm77FmWqg8MDhQnMO" alt="" data-size="original">
* Choose format, size and resolution in the *Export image* dialog (Figure 6)

![Figure 6. Choosing format, size and resolution for export](/files/gkz4Rkygw0xNJNNo0tdp)

* Click **Save**

The heatmap will be saved as a .png file and downloaded in your web browser.

For more information about hierarchical clustering and the *Dendrogram view*, please see the [Hierarchical Clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering) user guide.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Performing biological interpretation

* [Enrichment analysis](#enrichment-analysis)
* [KEGG enrichment analysis](#kegg-enrichment-analysis)

To learn more about the biology underlying gene expression changes, we can use gene ontology (GO) or pathway enrichment analysis. Enrichment analysis identifies over-represented GO terms or pathways in a filtered list of genes.

## Enrichment analysis

* Click the **Filtered feature list** data node
* Click **Biological interpretation** in the task menu
* Click **Gene set enrichment** then select **Gene set database** to perform GO enrichment analysis (Figure 1)

![Figure 1. Invoking enrichment analysis](/files/zQPD7FkZpwD0zIfZIAAe)

* Select the latest gene set from geneontology.org from the *Gene set database* drop-down menu
* Click **Finish**

A gene set enrichment task node and data node will be added to the pipeline (Figure 2).

![Figure 2. Gene set enrichment task node and data node](/files/0pqLi4xKtqjb0Lio4imm)

* Double-click the **Gene set enrichment** task node to open the task report (Figure 3)

![Figure 3. Viewing the GO enrichment task report](/files/T4ZPC1GiWQmx3z4CnOpt)

The *gene set enrichment* task report lists GO terms by ascending p-value with the most significant GO term at the top of the list. Also included are the enrichment score, the number of genes from that GO term in the list, and the number of genes from that GO term that are not in the list.

To view the genes associated with each GO term, select ![](/files/KOPT8zhCgePaYKmxR9Ph) to open the extra details page. To view additional information about a GO term, click the blue gene set ID in the first column to open the linked geneontology.org entry in a new browser tab.

For more information about GO enrichment analysis, please see the [Gene Set Enrichment](/partek-flow/user-manual/task-menu/biological-interpretation/gene-set-enrichment) user guide.

## KEGG enrichment analysis

KEGG enrichment analysis identifies pathways that are over-represented in a gene list data node.

* Click the **Filtered feature list** data node
* Click **Biological interpretation** in the task menu
* Click **Gene set enrichment** then select **KEGG database**
* Click **Finish** in the configuration dialog to run KEGG analysis with the *Homo sapiens* KEGG database

A *Pathway* enrichment task node and data node will be added to the pipeline (Figure 4).

![Figure 4. Pathway enrichment task node and data node](/files/hD3ZH8FnDwQ8FI8kzPOu)

* Double-click the **Pathway enrichment** task node to open the task report

The *Pathway enrichment* task report is similar to the *gene set* e*nrichment analysis* task report (Figure 5).

![Figure 5. Pathway enrichment task report](/files/FzXKphMnZhAPiYrvrXKm)

To view an interactive KEGG pathway map, click the pathway ID (*Gene set* column).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Saving and running a pipeline

By following the steps in this tutorial, you have built a pipeline. You can save this pipeline for future use.

* Select **Create new pipeline** near the bottom left-hand side of the *Analyses* tab
* Select the **Pre-alignment QA/QC,** **Trim bases**, **STAR**, **Post-alignment QA/QC,** **Quantify to annotation model, Filter features, Normalize counts**, **PCA**, and **DESeq2** task nodes to include them in the pipeline. Note that skipping tasks within a pipeline is not allowed if the task uses the upstream output
* Name the pipeline as **RNA-Seq basic analysis**
* Give a description for the pipeline; we have noted **trimming, alignment, filter genes, normalization, DESeq2**
* Click **Create pipeline** (Figure 1)

![Figure 1. Select tasks to include then provide a name and description for the new pipeline](/files/tNRF6Rrxvgfau5HCU3Wv)

To access this pipeline in the future, select an unaligned reads data node and choose **Pipelines** from the task menu. Available saved pipelines will be available to choose from the *Pipelines* section of the task menu (Figure 2).

![Figure 2. Created pipelines will appear in the Pipelines section of the task menu](/files/BgZpRkRK7B1ZK6OUQmoo)

After selecting the pipeline, you will be prompted to choose the reference genome for alignment, the annotation for quantification, and the comparisons for DESeq2. After selecting these options, the pipeline will automatically run. See [Pipelines](/partek-flow/user-manual/pipelines) for more information.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Analyzing Single Cell RNA-Seq Data

* [Filtering cells](#filtering-cells)
* [Filter features](#filter-features)
* [Normalization](#normalization)
* [PCA](#pca)
* [Graph-based clustering](#graph-based-clustering)
* [t-SNE](#t-sne)
* [Coloring the t-SNE scatter plot](#coloring-the-t-sne-scatter-plot)
* [Selecting cells on the t-SNE scatter plot](#selecting-cells-on-the-t-sne-scatter-plot)
* [Filtering cells on the t-SNE scatter plot](#filtering-cells-on-the-t-sne-scatter-plot)
* [Classifying cells](#classifying-cells)
* [Comparing gene expression between cell types](#comparing-gene-expression-between-cell-types)
* [Generating a heatmap](#generating-a-heatmap)
* [Performing enrichment analysis](#performing-enrichment-analysis)
* [Pipeline](#pipeline)

This tutorial presents an outline of the basic series of steps for analyzing a single cell RNA-Seq experiment in Partek Flow starting with the count matrix file.

This tutorial includes only one sample, but the same steps will be followed when analyzing multiple samples. For notes on a few aspects specific to a multi-sample analysis, please see our [Single Cell RNA-Seq Analysis (Multiple Samples)](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples) tutorial.

If you are new to Partek Flow, please see [Getting Started with Your Partek Flow Hosted Trial](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial) for information about data transfer and import and [Creating and Analyzing a Project](/partek-flow/tutorials/creating-and-analyzing-a-project) for information about the Partek Flow user interface.

## Filtering cells

An important step in analyzing single cell RNA-Seq data is to filter out low quality cells. A few examples of low-quality cells are doublets, cells damaged during cell isolation, or cells with too few reads to be analyzed. You can do this in Partek Flow using the Single cell QA/QC task.

* Click on the **Single cell data** node
* Click on the **QA/QC** section of the task menu
* Click on **Single cell QA/QC**

A task node, *Single cell QA/QC*, is produced. Initially, the node will be semi-transparent to indicate that it has been queued, but not completed. A progress bar will appear on the *Single cell QA/QC* task node to indicate that the task is running (Figure 1).

![Figure 1. Analyses tab](/files/tZ6rToJuKZqoMzlW12xp)

* Click the **Single cell QA/QC** node once it finishes running
* Double-click the **Task report** in the task menu

The *Single cell QA/QC* report includes interactive violin plots showing the value of every cell in the project on several quality measures (Figure 2).

![](/files/eP2rA0G8De28W2ZiEryb)

![Figure 2. Single cell QA/QC plot](/files/AbCqZQUihH0sUK9PCbmV)

There can be four plots: number of read counts per cell, number of detected genes per cell, the percentage of mitochondrial reads per cell, and the percentage of ribosomal counts.

Each point on the plots is a cell and the violins illustrate the distribution of values for the y-axis metric. Cells can be filtered either with the plot controls by using the selection tools on the right of the plot (**rectangle mode** ![](/files/fWhZ77MUDSExC9nDK7pO), ellipse mode ![](/files/vwCI60rLBgURWD7IJFPj), or **lasso mode** ![](/files/5OX0B6UoQW7mTYErtHSi) ) and selecting a region on one of the plots or by setting thresholds using the **Select & Filter** ![](/files/oxZ6qU9XrxtZqCzl2njJ) tool. Here, we will apply a filter for the number of read counts.

The plots will be shaded to reflect the selection. Cells that are excluded will be shown as dim dots on all plots.

The read counts per cell and number of detected genes per cell are typically used to filter out potential doublets - if a cell as an unusually high number of total counts or detected genes, it may be a doublet. The mitochondrial reads percentage can be used to identify cells damaged during cell isolation - if a cell has a high percentage of mitochondrial counts, it is likely damaged or dying and may need to be excluded.

* Open the ![](/files/oxZ6qU9XrxtZqCzl2njJ) **Select & Filter** icon in the left panel. The histograms can be pinned while fine tuning the selections (Figure 2). Set the filters to represent the majority of the population (violin width)
* Click the filter icon ![](/files/EsUWu45aqb78DXdxTYik) and **Apply observation filter** then select the *Single cell counts* data node to run the Filter cells task on the first Single cell counts data node, it generates a *Filtered counts* task node that generates a *Filtered cells* results node
* Use ![](/files/88oChDQjutw7qb8FRzyo) **Save as** to give this Data Viewer session a new name (e.g. QA/QC filter) so you can return to this filter at any time and see the exact criteria that has been selected and filtered.

## Filter features

A common task in bulk and single-cell RNA-Seq analysis is to filter the data to include only informative genes (features). Because there is no gold standard for what makes a gene informative or not and ideal gene filtering criteria depends on your experimental design and research question, Partek Flow has a wide variety of flexible filtering options. The Filter features step can also be performed before normalization or after normalization.

* Click the data node containing count matrix
* Click **Filtering** in the task menu
* Click **Filter features**

There are four categories of filter available - noise reduction, statistics based, feature metadata, and feature list.

The noise reduction filter allows you to exclude genes considered background noise based on a variety of criteria. The statistics based filter is useful for focusing on a certain number or percentile of genes based on a variety of metrics, such as variance. The feature list filter allows you to filter your data set to include or exclude particular genes.

For example, you can use a noise reduction filter to exclude genes that are not expressed by any cell in the data set, but were included in the matrix file.

* Click the **Noise reduction filter** check box
* Set the *Noise reduction filter to Exclude features where* **value <= 0 in at least 99.9% of cells** using the drop-down menus and text boxes
* Click **Finish** to apply the filter (Figure 3)

![Figure 3. Filter features](/files/NTwru0iE4rHDcvOijcQ5)

This results node, *Filtered counts*, will be the starting point for the next stage of analysis.

## Normalization

Because different cells will have a different number of total counts, it is important to normalize the data prior to downstream analysis. For droplet-based single cell isolation and library preparation methods that use a 3' counting strategy, where only the 3' end of each transcript is captured and sequenced, we recommend the following normalization - 1. CPM (counts per million), 2. Add 1, 3. Log2. This accounts for differences in total UMI counts per cell and log transforms the data, which makes the data easier to visualize.

* Click the *Filtered cells* results node produced by the *Filtered counts* task
* Click **Normalization and scaling** in the context-sensitive task menu on the right
* Click **Normalization**
* Click ![](/files/SJ5Moeqnb2LdCG5wtMyo)to add the recommended normalization scheme

This adds *CPM (counts per million), Add 1,* and *Log2* to the *Normalization order* panel. Normalization steps are performed in descending order.

* Click **Finish** to apply the normalization (Figure 4 )

![Figure 4. Normalization settings](/files/K7THzs0Jk5KUlJTEUMaM)

A new *Normalized counts* data node will be produced. You can choose to change the color of this node by right-clicking on the task node then clicking **Change color** and/or rename the result node by right-clicking and selecting **Rename data node**.

In the example below, I have changed the color to dark blue and renamed the results node based on the scheme.

![](/files/Rarc06RWwrQqFaQYxnU9)

For more information on normalizing data in Partek Flow, please see the Normalize Counts section of the user manual.

## PCA

Principal components (PC) analysis (PCA) is an exploratory technique that is used to describe the structure of high dimensional data by reducing its dimensionality. Because PCA is used to reduce the dimensionality of the data prior to clustering as part of a standard single cell analysis workflow, it is useful to examine the results of PCA for your data set prior to clustering.

* Click the *Filtered counts* node
* Click **Exploratory analysis** in the task menu
* Click **PCA** from the drop-down list

You can choose *Features contribute* **equally** to standardize the genes prior to PCA or allow more variable genes to have a larger effect on the PCA by choosing **by variance**. By default, we take variance into account and focus on the most variable genes.

If you have multiple samples, you can choose to run PCA for each sample individually or for all samples together by selecting or not selecting the *Split by sample* option (Figure 5).

![Figure 5. Configuring PCA](/files/5RYY5S7tqbuJAdXNoD0l)

* Click **Finish** to run

A new *PCA* task node will be produced.

* Double-click the **PCA** task node to open the 3D PCA scatter plot in data viewer (Figure 6)

![Figure 6. PCA scatter plot, each dot is a cell](/files/haR2QfzcDSv1yxuOzU2Z)

Beside PCA coordinates of the cells, PCA task report also includes, the Scree plot, the component loadings table, and the PC projections table.

The Scree plot lists PCs on the x-axis and the amount of variance explained by each PC on the y-axis, measured in Eigenvalue. The higher the Eigenvalue, the more variance is explained by the PC. Typically, after an initial set of highly informative PCs, the amount of variance explained by analyzing additional PCs is minimal. By identifying the point where the Scree plot levels off, you can choose an optimal number of PCs to use in downstream analysis steps like graph-based clustering, UMAP and t-SNE.

* To draw a Scree plot, in Data viewer, choose **Scree plot** icon available in **New plot** under *Setup* on the left panel ![](/files/cVkmea0x0mKV2tAPVU5Z), choose the PCA data node (Figure 7)

Note that Partek Flow suggests appropriate data for each plot type that is chosen so only PCA results will be available to select from for the Scree plot.

![Figure 7. PCA Scree plot](/files/9zfX5wnncc6gB3kvTJbO)

* Mouse over the Scree plot to identify the point where additional PCs offer little additional information

In this data set, a reasonable cut-off could be set anywhere between 7 and 20 PCs.

Viewing the genes correlated with each PC can be useful when choosing how many PCs to include.

* Click the **Table** ![](/files/oSSeuD6idY1SaO9gx9cd) option in the **New plot** icon under *Setup* and select the PCA data node to open the Component loadings table (Figure 8)

![Figure 8. Component loadings](/files/ygcRJB7M9sBv5ys3Naxv)

This table lists genes on rows and PCs on columns, the value in this table is correlation coefficient r. The table can be downloaded as a text file by clicking on the Export table data icon ![](/files/GnqlgIP5pO0RMMO6czpx) on the upper-right corner of the plot.

To display PCA projects table, click on the Table drop-down list in the **Content** icon under *Configure* and choose **PCA projections** (Figure 9)

![Figure 9. PC projection configuration dialog](/files/7l7QRNBn5EbbULrIfUCa)

PCA projections table contains each row as an observation (a cell in this case), each column represents one principal component (Figure 10). This table can be downloaded as text file, the same way as the component loading table.

![Figure 10. PCA project table](/files/NthGlvvpOor9U74V1g4d)

## Graph-based clustering

Graph-based clustering identifies groups of similar cells using PC values as the input. By including only the most informative PCs, noise in the data set is excluded, improving the results of clustering.

* Click the *PCA* data node
* Click **Exploratory analysis** in the task menu
* Click **Graph-based clustering**

Clustering can be performed on each sample individually or on all samples together. Here, we are working with a single sample.

* Check *Compute biomarkers* to compute features that are highly expressed when comparing each cluster (Figure 11)
* Click **Configure** to access the *Advanced options* and change the *Number of nearest neighbors* to **50** and *Nearest Neighbor Type* to **K-NN** for this example tutorial.

![Figure 11. Configure Graph-based clustering](/files/TA7DrM1UwDdxPAZfwf0t)

The *Number of principal components* should be set based on the your examination of the Scree plot and component loadings table. The default value of 100 is likely exhaustive for most data sets, but may introduce noise that reduces the number of clusters that can be distinguished.

* Click **Finish** to run the task

A new *Graph-based clusters* data and *Biomarkers* data node will be generated along with the task nodes.

* Double-click the **Graph-based clusters** node to see the cluster results and statistics (left screenshot on Figure 12)
* Double-click the **Biomarkers** node to see the computed biomarkers if you have selected this option (right screenshot on Figure 12)

The *Graph-based clustering result* lists the *Total number of clusters* and what proportion of cells fall into each cluster as well as *Maximum modularity* which is a measurement of the quality of the clustering result where optimal modularity is 1. The *Biomarkers* node includes the top features for each graph-based cluster. It displays the top-10 genes that distinguish each cluster from the others. **Download** at the bottom right of the table can be used to view and save more features. These are calculated using an ANOVA test comparing the cells in each group to all the other cells, filtering to genes that are 1.5 fold upregulated, and sorting by ascending p-value. This ensures that the top-10 genes of each cluster are highly and disproportionately expressed in that cluster.

![](/files/mTXq8PRSaNE9h6fXxLkr) ![Figure 12. Graph-based clustering results](/files/2Mj7xGowtNXi966deoSp)

We will use t-SNE to visualize the results of Graph-based clustering.

## t-SNE

t-Distributed Stochastic Neighbor Embedding (t-SNE) is a dimensional reduction technique that prioritizes local relationships to build a low-dimensional representation of the high-dimensional data that places objects that are similar in high-dimensional space close together in the low-dimensional representation. This makes t-SNE well suited for analyzing high-dimensional data when the goal is to identify groups of similar objects, such as cell types in single cell RNA-Seq data.

* Click the **Graph-based clusters** node
* Click **Exploratory analysis** in the task menu
* Click **t-SNE**

If you have multiple samples, you can choose to run t-SNE for each sample individually or for all samples together using the *Split cells by sample* option. Please note that this option will not be present if you are running t-SNE on a clustering result. For clarity, clustering results run with all samples together must be viewed together and clustering results run by sample must be viewed by sample.

Like Graph-based clustering, t-SNE takes PC values as its input and further reduces the data down to two or three dimensions. For consistency, you should use the same number of PCs as the input for t-SNE that you used for Graph-based clustering.

* Click **Apply**
* Click **Finish** to run (Figure 13)

![Figure 13. t-SNE configuration](/files/UWUbQw47v80C3989bnli)

A new *t-SNE* task node will be produced.

* Double-click the **t-SNE** node to open the t-SNE task report (Figure 14). Use the panel on the left to modify the plot or add more plots to this Data viewer session.

![Figure 14. t-SNE plot](/files/ZGgjzD5jt56K2vuM9v0T)

The t-SNE scatter plot is interactive and can be viewed for 2D or 3D. The t-SNE plot is 3D by default. You can rotate the 3D plot by left-clicking and dragging your mouse or using **Control** under *Configure*. You can zoom in and out using your mouse wheel. You can pan by right-clicking and dragging your mouse. You can use **Style** to modify *color*, *shape*, *size*, and *labeling* (e.g. add a fog effect to improve depth perception on the plot). Add a 2D plot clicking **New plot,** selecting **2D Scatter plot** and selecting t-SNE as the source of the data.

## Coloring the t-SNE scatter plot

Click on the plot to ensure that the plot window is selected. Click **Style** under *Configure* to color the t-SNE.

* *Color by* the options in the drop-down menu under *Color.* You should be on the normalized counts node which can be seen by hovering over or clicking the circle (node) to the right of the drop-down.
* Click the text field in the drop-down and start typing CD79A then select the gene by clicking on it (Figure 15)

![Figure 15. Coloring by a gene](/files/9oZ572NoSh1WtnoVf2mf)

The cells on the plot will be colored based on their expression level of CD79A (Figure 16). In the example in Figure 16, the **Style** icon has been dragged to a different location on the screen and the legend has also been resized and moved. Resizing the legend can either be done on the legend itself or using the **Description** icon under *Configure*.

![Figure 16. Coloring by CD79A expression](/files/7CrpQ2jLcJUKunuyiehv)

Coloring by one gene uses the two-color numeric palette, which can be customized by clicking ![](/files/B5InuzEotSNBiZ9AWxXH). To color by more than one gene use the **Numeric triad** option in the drop-down. If you color by more than one gene, the color palette switches to a Green-Red-Blue color scheme with the balance between the three color channels determined by the values of the three genes. For example, a cell that expresses all three genes would be white, a cell that expresses the first two genes would be yellow, and a cell that expresses none of the genes would be black (Figure 17).

![Figure 17. Coloring by three genes](/files/CudsG5RQmngJww7ka368)

Clicking a cell on the plot shows the expression values of the cell in the legend. Hovering over a cell on the plot also shows this information and related details (Figure 18).

![Figure 18. Viewing expression values of a cell](/files/Tn0t6uf9AIN7epDA2qWz)

If you want to color by more than three genes at time, such as by a list of genes that distinguish a particular cell type, you can use the *color by* **Feature list** option.

* Select **Feature** **List** from the *Color by* drop-down
* Choose **Cytotoxic cells** from the *List* drop-down (use **List management** in **Settings** to add lists to Partek Flow which will automatically make them available here)
* Choose **PCA** from the *Metric* drop-down

Coloring by a list, in this way, calculates the first three principal components for the gene list and colors the cells on the plot by their values along those three PCs with green for PC1, red for PC2, and blue for PC3 (Figure 19).

![Figure 19. Coloring by a list](/files/bHASS2b9tOFawFMJSX8S)

Typically, the expression of a set of marker genes will be highly correlated, allowing the first PC to account for a large percentage of the variance between cells for that gene list. As a result, the group of cells characterized by their expression of the genes on the list will separate from the rest of the cells along PC1 and will be colored green (Figure 16). If the gene list is more complex, for example, including marker genes for multiple cell types, there may be several sets of correlated genes accounting for significant amounts of variance, leading to groups of cells being distinguishable along PC2 and PC3 as well. In that case, there may be green, blue, and red groups of cells on the plot. If the gene list does not distinguish any group of cells, all cells will have similar PC values, leading to similarly colored cells on the plot.

In addition to coloring by gene expression and by gene lists, the points can be colored by any cell or sample attribute. Available attributes are listed as options in the *Color by* drop-down menu. Note that any available options are dependent upon the selected data node. In the following section we will use the attribute **Graph-based** to color our cells by the clusters identified in the Graph-based clustering task (Figure 20).

![Figure 20. Coloring by Graph-based attribute](/files/FzvxmWQpNKeFXKq2X8Q5)

## Selecting cells on the t-SNE scatter plot

The most basic way to select a point on the scatter plot is to click it with the mouse while in pointer mode. To select multiple cells, you can hold Ctrl on your keyboard and click the cells. To select larger groups of cells, you can switch to Lasso mode by clicking ![](/files/5OX0B6UoQW7mTYErtHSi) in the plot controls on the right hand side. The lasso lets you freely draw a shape to select a cluster of cells.

* Click ![](/files/5OX0B6UoQW7mTYErtHSi) to activate Lasso mode
* Left-click and hold to draw a lasso around a cluster of cells
* Release and click the starting circle to close the lasso and select the enclosed cells (Figure 21)

You can also create a lasso with straight lines using Lasso mode by clicking, releasing, and clicking again to draw a shape.

![Figure 21. Lassoing cells](/files/ZnadZm9eRCzBxDinXFw6)

By default, selected cells are shown in bold while unselected cells are dimmed (Figure 22). This can be changed to gray selected cells using the *Select & Filter* tool in the left panel as shown in Figure 22.

* Double-click any blank section of the scatter plot to clear the selection

![Figure 22. Selected cells](/files/PBWBCnrSNCyggNQwbFmt)

Alternatively, you can select cells using any criteria available for the data node that is selected in the *Select & Filter* tool. To change the data selection click the circle (node) and select the data.

* Choose **Graph-based** from the **Criteria** drop-down menu in the **Select & Filter** tool after ensuring you on are on the Graph-based cluster node by hovering on the circle (Figure 23). If you are not on the correct node, you need to click the circle and select the data.

![Figure 23. Picking an attribute](/files/x9wS43Ok7Juo7l0KmPWS)

This adds check boxes for each level of the attribute (i.e., clusters). Click a check box to select the cells with that attribute level.

* Click only **2** and **3**

This selects cells from Graph-based clusters 2 and 3 (Figure 24). The number of selected cells is listed in the Legend on the plot.

![Figure 24. Selecting by attribute](/files/zjYtQX56hffeQ0b7YycR)

Cells can also be selected based on their gene expression values in the **Select & Filter** section.

* Click the circle and select the *Normalized counts* node which has gene expression data
* Type **cd3d** in the text field of the drop-down
* Click on **CD3D** to add it as criteria to select from and use the slider or text field to adjust the selected values. Pin the histogram to visualize the distribution during selection.

Very specific selections can be configured by adding criteria in this way. In the example below, Clusters 2 and 3 and high CD3D expression is selected (Figure 25).

![Figure 25. Selecting by gene expression level](/files/LZ1iJgYn6CZRJCQ8rsbv)

## Filtering cells on the t-SNE scatter plot

Once a cell has been selected on the plot, it can be filtered. The filter controls can exclude or include (only) any selected cell. Filtering can be particularly useful when you want to use a gene expression threshold to classify a group of cells, but the gene in question is not exclusively expressed by your cell type of interest.

In this example we can filter to include just cells from the selection we have already made.

* Click ![](/files/a5nfgK3X8m8KwP4PQ53b) (filter include) to filter to just the selected cells (Figure 26).

The plot will update to show only the included cells as seen in Figure 26.

Cells that are not shown on the plot cannot be selected, allowing you to focus on the visible cells. The number of cells shown on the plot out of the total number of original cells is shown in the Legend. You can adjust the view to focus on only the included cells.

* Click ![](/files/I3tCJnMy2GlM462x7TFx) on the plot controls or toggle on **Fit visible** in the **Axes** configuration to rescale the axes to the filtered points

To revert to the original scaling, click the ![](/files/P1BEJGgKGh0pHr8NpNsY) button again or turn off Fit visible with the toggle.

![Figure 26. Activating the filter](/files/VVU2AOCe6QTrEIX3LlCI)

* Alternatively, to exclude selected cells, click ![](/files/PMIRI9cDHX7OlasuG1p4) (filter exclude) (Figure 27)

Additional inclusion or exclusion filters can be added to focus on a smaller subset of cells.

![Figure 27. Filtered t-SNE scatter plot](/files/vCyViOU8tDJ99UNkdQU1)

* Click **Clear filters** to remove applied filters

The plot will update to show all cells and return to the original scaling.

## Classifying cells

Classifying cells allows to you assign cells to groups that can be used in downstream analysis and visualizations. Commonly, this is used to describe cell types, such as B cells and T cells, but can be used to describe any group of cells that you want to consider together in your analysis, such as cycling cells or CD14 high expressing cells. Each cell can only belong to one class at a time so you cannot create overlapping classes.

To classify a cell, just select it then click **Classify selection** in the **Classify** tool.

For example, we can classify a cluster of cells expressing high levels of CD79A as B cells.

* Set *Color by* in the **Style** configuration to the normalized counts node
* Type **CD79A** in the search box and select it. Rotate the 3D plot if you need to see this cluster more clearly.
* Click ![](/files/5OX0B6UoQW7mTYErtHSi) to activate Lasso mode
* Draw a lasso around the cluster of CD79A-expressing cells (Figure 28)

![Figure 28. Selecting a cluster of CD79A-expressing cells](/files/3oaD70AQMQ4sg3kVJ6Ah)

Because most of these cells express CD79A, a B cell marker, and because they cluster together on the t-SNE, suggesting they have similar overall gene expression, we believe that all these cells are B cells.

* Click **Classify** under *Tools* in the left panel
* Type **B cells** for the Name
* Click **Save** (Figure 29)

![Figure 29. Classifying cells](/files/mww8VONo5W63iU3lKpq6)

You can edit the name of a classification or delete it. In this project we use the hosted feature lists for "NK cells", "T cells" and "Monocytes" to classify these cell types by coloring the cells in the t-SNE plot and selecting the cells expressing those genes as shown above. See the [list management](/partek-flow/user-manual/settings/components/lists) documentation for more information on how to add these lists. The classifications you have made are saved as a working draft so if you close the plot and return to it, the classifications will still be there and can be visualized on the plot as **"New classification"**. However, classifications are not available for downstream tasks until you apply them. Continue classifying the clusters and save the Data viewer session until you are ready to apply the classification to the data project.

* *Color by* **New classifications** under **Style** (Figure 30) while you are still working on the classifications

![Figure 30. Color by New classification](/files/PA6uwkhfrter3bD2zivk)

To use the classifications in downstream tasks and visualizations, you must first apply them.

* Click **Apply classifications**
* **Name** the classification (e.g. Classified Cell Types)
* Click **Run** to confirm

Once you have added a classification to the project, you can color the t-SNE plot by the Classification.

Here, I classified a few additional cell types using a combination of known marker genes and the clustering results then applied the classification (Figure 31).

![Figure 31. Color by Applied classification](/files/ZZsdfTiUJ6A0lPqXwfBE)

Summarize *Classifications* with the number and percentage of cells from each sample that belong to each classification using an **Attribute table** under **New plot**. This is particularly useful when you are classifying cells from multiple samples.

* Click **New plot**
* Select **Attribute table** and the source of data (Figure 32) which in this case is called *Classify result*

![Figure 32. Attribute table](/files/H04g1g37DitXj0mgcgYI)

* Click on the **Normalized counts"** node
* Navigate to the **Compute biomarkers** task under **Statistics** in the task menu
* Follow the task dialogue and click **Finish** (Figure 33)
* Double click the **Biomarkers** node to view the Biomarkers results

![Figure 33. Compute biomarkers](/files/igai28OVnz1Yrwhe9b0r)

## Comparing gene expression between cell types

A common goal in single cell analysis is to identify genes that distinguish a cell type. To do this, you can use the differential analysis tools in Partek Flow. I will show how to use the ANOVA test in Partek Flow, a statistical test shown to be highly effective for differential analysis of single cell RNA-Seq data.

* Click the **Normalized counts** results node
* Click **Statistics** in the toolbox
* Click **Differential Analysis**
* Select **ANOVA** as the M\_ethod to use for differential analysis\_

The first page of the configuration dialog asks what attributes you want to include in the statistical test. Here, we only want to consider the Classifications, but in a more complex experiment, you could also include experimental conditions or other sample attributes.

* Click **Classified Cell Types**
* Click **Next** (Figure 34)

![Figure 34. Choosing attributes to include in the statistical test](/files/om8rxhymXRA5PuHSzVgv)

We will make a comparison between NK cells and all the other cell types to identify genes that distinguish NK cells. You can also use this tool to identify genes that differ between two cell types or genes that differ in the same cell type between experimental conditions.

* Drag **NK cells** to the top panel

The top panel is the numerator for fold-change calculations so the experimental or test groups should be selected in the top panel.

* Click all the other classifications in the bottom panel

The bottom panel is the denominator for fold-change calculations so the control group should be selected in the bottom panel.

* Click **Add comparison**

This adds the comparison to the statistical test.

* Click **Finish** to run the ANOVA task (Figure 35)

![Figure 35. Configuring comparisons in the GSA task](/files/ALlggSIvAICIWBuWBTFu)

* Double-click the newly generated data node to open the ANOVA task report

The ANOVA task report lists genes on rows and the results of the statistical test (p-value, fold change, etc.) on columns (Figure 36). For more information, please see our documentation page on the [ANOVA task report](/partek-flow/user-manual/task-menu/differential-analysis/anova-limma-trend-limma-voom).

![Figure 36. Viewing GSA results](/files/FkHPbvioCwJMCkluJZjU)

Genes are listed in ascending order by the p-value of the first comparison so the most significant gene is listed first. To view a volcano plot for any comparison, click ![](/files/keaDTokBfjJKEUxgY2qO). To view a violin plot for a gene, click ![](/files/BZ0OjuJ8mO2fyo9vWKCW) next to the Gene ID.

* Click ![](/files/BZ0OjuJ8mO2fyo9vWKCW) for CCL4

The Feature plot viewer will open showing a dot plot for CCL4 which can be modified to summarize the data in different ways (Figure 37). In the image below, the red boxes highlight the changes that were made to configure the plot. This includes overlaying the violins (density plots with the width corresponding to frequency) on the dot plot represented by the Classified Cell Types.

![Figure 37. Gene expression plot](/files/u6HDRcKJUpLYwUMHiZIa)

You can switch the grouping of cells. To do this, show the X axis labels then click and drag the labels to reposition the cell types on the plot.

* Click **ANOVA report** to return to the table

The table lists all of genes in the data set; using the filter control panel on the left, we can filter to just the genes that are significantly different for the comparison.

* Click **FDR step up** and click the arrow next to it
* Set to **1e-8**

Here, we are using a very stringent cutoff to focus only on genes that are specific to NK cells, but other applications may require a less stringent cutoff.

* Click **Fold change** and click the arrow next to it
* Set to **-2** to **2**

The number of genes at the top of the filter control panel updates to indicate how many genes are left after the filters are applied.

* Click ![](/files/KTIuGBSW2eVUhh3ZWtI8) to generate a filtered version of the table for downstream analysis

The ANOVA report will close and a new task, the Differential analysis filter, will run and generate a filtered Feature list data node.

For more information about the ANOVA task, please see the [Differential Gene Expression - ANOVA](/partek-flow/user-manual/task-menu/differential-analysis/anova-limma-trend-limma-voom) section of our user manual.

## Generating a heatmap

Once we have filtered to a list of significantly different genes, we can visualize these genes by generating a heatmap.

* Click the **Filtered feature list** data node produced by the Differential analysis filter
* Click **Exploratory analysis** in the toolbox
* Click **Hierarchical clustering** **/ heatmap**

The hierarchical clustering task will generate the heatmap; choose **Heatmap** as the plot type. You can choose to **Cluster** features (genes) and cells (samples) under *Feature order* and *Cell order* in the *Ordering* section. You will almost always want to cluster features as this generates the clear blocks of color that make heatmaps comprehensible. For single cell data sets, you may choose to forgo clustering the cells in favor of ordering them by the attribute of interest. Here, we will not filter the cells, but instead order them by their classification.

* Click **Assign order** under *Cell order*

You can filter samples using the Filtering section of the configuration dialog. Here, we will not filter out any samples or cells.

* Choose **Classification** from the Ordering drop-down menu
* Drag **NK cells** to the top of the Sample order
* Click **Finish** to run (Figure 38)

![Figure 38. Configuring hierarchical clustering](/files/CJqMR678hl8OfJ70YoqO)

* Double-click the **Hierarchical cluster** task node to open the task report

It may initially be hard to distinguish striking differences in the heatmap. This is common in single cell RNA-Seq data because outlier cells will skew the high and low ends. We can adjust the minimum and maximum of the color scheme to improve the appearance of the heatmap.

* Click **Heatmap**
* Toggle on the *Range* **Min** and set to -2
* Toggle on the *Range* **Max** and set to 2

Distinct blocks of red and blue are now more pronounced on the plot. Cells are on rows and genes are on columns. Because of the limited number of pixels on the screen, genes are grouped. You can zoom in using the zoom controls or your mouse wheel if you want to view individual gene rows. We can annotate the plot with cell attributes.

* Choose **Classified Cell Types** from the **Annotations** drop-down menu
* Change the **Annotation font** size under *Style* in the **Annotations** section

The plot now includes blocks of color along the left edge indicating the classification of the cells. We can transpose the plot to give the cell labels a bit more space.

* Click **Transposed** under **Axes** or use the transpose button ![](/files/WHp5lR20imwiCRXzjQcC) on the plot to flip the axes
* Toggle off the **Row labels** under **Axes** to remove the sample labels

![Figure 39. Configurable heat map](/files/EtcI1s9G4p8FfmuqcH8D)

As with any visualization in Partek Flow, the image can be saved as a publication-quality image to your local machine by clicking ![](/files/oDrw6FbEdm5yjzDcPEbe) or sent to a page in the project notebook by clicking ![](/files/lijXR1qRKTWQI8S6j8rC). For more information about Hierarchical clustering, please see the [Hierarchical Clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering) section of the user manual.

## Performing enrichment analysis

While a long list of significantly different genes is important information about a cell type, it can be difficult to identify what the biological consequences of these changes might be just by looking at the genes one at a time. Using enrichment analysis, you can identify gene sets and pathways that are over-represented in a list of significant genes, providing clues to the biological meaning of your results.

* Click the **Feature list** data node produced by the Differential analysis filter
* Click **Biological interpretation**
* Click **Gene set enrichment**

We distribute the gene sets from the Gene Ontology Consortium, but Gene set enrichment can work with any custom or public gene set database.

* Choose the latest assembly available from the Gene set drop-down
* Click **Finish**
* Double-click the **Gene set enrichment** task node to open the task report

The Gene set enrichment task report lists gene sets on rows with an enrichment score and p-value for each. It also lists how many genes in the gene set were in the input gene list and how many were not (Figure 40). Clicking the Gene set ID links to the geneontology.org page for the gene set.

![Figure 40. Gene set enrichment report](/files/QvueRGlF4v9PiKZIeA2G)

In Partek Flow, you can also check for enrichment of KEGG pathways using the Pathway enrichment task. The task is quite similar to the Gene set enrichment task, but uses KEGG pathways as the gene sets.

The task report is similar to the Gene set enrichment task report with enrichment scores, p-values, and the number of genes in and not in the list (Figure 41).

![Figure 41. Pathway enrichment report](/files/u4MMh9pcWME6WBIjqcPc)

Clicking the KEGG pathway ID in the Pathway enrichment task report opens a KEGG pathway map (Figure 42). The KEGG pathway maps have fold-change and p-value information from the input gene list overlaid on the map, adding a layer of additional information about whether the pathway was upregulated or downregulated in the comparison.

![Figure 42. KEGG pathway map](/files/5fuLHWW1aAs6l7rZMBvD)

Color are customizable using the control panel on the left and the plot is interactive. Mousing over gene boxes gives the genes accounted for by the box, with genes present in the input list shown in bold, and the coloring gene shown in red (Figure 43).

![Figure 43. Viewing pathway map details](/files/pKZYjpmHmoOm2UPgFWoJ)

Clicking a pathway box opens the map of that pathway, providing an easy way to explore related gene networks.

## Pipeline

![Figure 44. Described pipeline shown in the Analyses tab](/files/gY7Rs8CkFWf2F3o87dh0)

For information about automating steps in this analysis workflow, please see our documentation page on [Making a Pipeline](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/user-manual/pieplines/making-a-pipeline.md).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Analyzing CITE-Seq Data

* [Importing Feature Barcoding Data](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/analyzing-cite-seq-data/importing-barcode-data.md)
* [Data Processing](/partek-flow/tutorials/analyzing-cite-seq-data/data-processing)
* [Dimensionality Reduction and Clustering](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/analyzing-cite-seq-data/dimensionality-reduction-and-clustering/README.md)
* [Classifying Cells](/partek-flow/tutorials/analyzing-cite-seq-data/classifying-cells)
* [Differentially Expressed Proteins and Genes](/partek-flow/tutorials/analyzing-cite-seq-data/differentially-expressed-proteins-and-genes)

This tutorial presents an outline of the basic series of steps for analyzing a 10x Genomics Gene Expression with Feature Barcoding (antibody) data set in Partek Flow starting with the output of Cell Ranger.

If you have Cell Hashing data, please see our documentation on [Hashtag demultiplexing](/partek-flow/user-manual/task-menu/pre-analysis-tools/hashtag-demultiplexing).

This tutorial includes only one sample, but the same steps will be followed when analyzing multiple samples. For notes on a few aspects specific to a multi-sample analysis, please see our [Single Cell RNA-Seq Analysis (Multiple Samples)](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples) tutorial.

If you are new to Partek Flow, please see [Getting Started with Your Partek Flow Hosted Trial](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial) for information about data transfer and import and [Creating and Analyzing a Project](/partek-flow/tutorials/creating-and-analyzing-a-project) for information about the Partek Flow user interface.

## Data set

The data set for this tutorial is a demonstration data set from 10x Genomics. The sample includes cells from a dissociated Extranodal Marginal Zone B-Cell Tumor (MALT: Mucosa-Associated Lymphoid Tissue) stained with BioLegend TotalSeq-B antibodies. We are starting with the [Feature / cell matrix HDF5 (filtered)](https://cf.10xgenomics.com/samples/cell-exp/3.0.0/malt_10k_protein_v3/malt_10k_protein_v3_filtered_feature_bc_matrix.h5) produced by Cell Ranger. Prior to beginning, transfer this file to your Partek Flow using the **Transfer files** button on the homepage.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing Feature Barcoding Data

* [Create a new Project](#create-a-new-project)
* [Import data](#import-data)

## Create a new Project

Let's start by creating a new project.

* On the *Home page*, click **New project** (Figure 1)
* Give the project a name
* Click **Create project**

![Figure 1. Create a new project and give it a meaningful name (e.g. CITE-Seq tutorial)](/files/LlBkIK4xkvctbPAP6mmj)

## Import data

* In the Analyses tab, click **Add data**
* Click **10x Genomics Cell Ranger counts h5** (Figure 2)
* Choose the filtered HDF5 file for the MALT sample produced by Cell Ranger

![Figure 2. Import options for CITE-Seq tutorial data](/files/ymaAqBF3g9p98vVTrV3G)

Move the .h5 file to where Partek Flow is installed using ![](/files/vxKKv51qniTOjIlzV2Z8), then browse to its location.

![Figure 3. Import options for CITE-Seq tutorial data](/files/GwHtm1xVshsRys0JKtAW)

Note that Partek Flow also supports the feature-barcode matrix output (barcodes.tsv, features.tsv, matrix.mtx) from Cell Ranger. The import steps for a feature-barcode matrix are identical to this tutorial.

* Click **Next**
* Name the sample **MALT** (the default is the file name)
* Specify the annotation used for the gene expression data (here, we choose **Homo sapiens (human) - hg38** and **Ensembl Transcripts release 109**). If Ensembl 109 is not available from the drop-down list, choose **Add annotation** and download it.
* Check **Features with non-zero values across all samples** in the *Report* section
* Click **Finish** (Figure 3)

![Figure 4. File format options for MALT data set](/files/Vt2rcbTXfu19SY7xh8EY)

A *Single cell counts* data node will be created under the *Analyses* tab after the file has been imported. We can move on to processing the data.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Data Processing

* [Split matrix](#split-matrix)
* [Filter low-quality cells](#filter-low-quality-cells)
* [Normalization](#normalization)
* [Merge Protein and mRNA data](#merge-protein-and-mrna-data)
* [Collapsing tasks to simplify the pipeline](#collapsing-tasks-to-simplify-the-pipeline)
* [References](#references)

## Split matrix

The *Single cell counts* data node contains two different types of data, mRNA expression and protein expression. So that we can process these two different types of data separately, we will split the data by data type.

* Click the **Single cell counts** data node
* Click **Pre-analysis tools** in the toolbox
* Click **Split by feature type**

A rectangular task node will be created along with two circular data nodes, one for each data type (Figure 1). The labels for these data types are determined by features.csv file used when processing the data with Cell Ranger. Here, our data is labeled *Gene Expression*, for the mRNA data, and *Antibody Capture*, for the protein data.

![Figure 1. Split by feature type produces two data nodes, one for each data type](/files/yP0Bmt7Ywn7U55TO67dO)

## Filter low-quality cells

An important step in analyzing single cell RNA-Seq data is to filter out low-quality cells. A few examples of low-quality cells are doublets, cells damaged during cell isolation, or cells with too few counts to be analyzed. In a CITE-Seq experiment, protein aggregation in the antibody staining reagents can cause a cell to have a very high number of counts. These are low-quality cells that can be excluded. Additionally, if all cells in a data set are expected to show a baseline level of expression for one of the antibodies used, it may be appropriate to filter out cells with very low counts or a low number of detected features. You can do this in Partek Flow using the *Single cell QA/QC* task.

We will start with the protein data.

* Click the **Antibody Capture** data node
* Click **QA/QC** in the toolbox
* Click **Single Cell QA/QC**

This produces a *Single-cell QA/QC* task node (Figure 2).

![Figure 2. Single cell QA/QC produces a task node](/files/jz3x2sQ7zuOlzxvgnTqa)

* Double-click the **Single cell QA/QC** task node to open the task report

The *Single cell QA/QC* report opens in a new data viewer session. There are interactive violin plots showing the most commonly used quality metrics for each cell: the total count per cell and the number of detected features per cell (Figure 3). Each point on the plots is a cell and the violins illustrate the distribution of values for the y-axis metric.

![Figure 3. Each cell is shown as a point on the plot.](/files/rOP3cyQKWD4Ew5LMJHP6)

For this analysis, we will set a maximum counts threshold to exclude potential protein aggregates and, because we expect every cell to be bound by several antibodies, we will also set a minimum counts threshold.

* Select one of the plots on the canvas
* In the **Select & Filter** icon on the left under *Tools*, set the *Counts* threshold to keep cells between **500** and **20000** (Figure 4)

![Figure 4. Filtering low quality cells based on protein expression data](/files/SNHz50Os2WylQiRhtOgn)

* Click ![](/files/zdAVcV8BSPFAx45UsZIH) under *Filter* on the right
* Click **Apply observation filter...**
* Select the **Antibody Capture** data node as input in the pipeline preview (Figure 5)
* Click **Select**

![Figure 5. After the Apply filter button is selected, you will be presented with a preview of your pipeline. You need to select the appropriate data node to apply the filtering to. In this case, the Antibody capture node](/files/J8VCrieFO64foxfRKWmR)

You will see a message telling you a new task has been enqueued.

* Click **OK** to dismiss the message
* Click the **project name** at the top to go back to the *Analyses* tab
* Your browser may warn you that any unsaved changes to the data viewer session will be lost. Ignore this message and proceed to the *Analyses* tab

A new task, *Filter counts*, is added to the *Analyses* tab. This task produces a new *Filter counts* data node.

Next, we can repeat this process for the *Gene Expression* data node.

* Click the **Gene Expression** data node
* Click the **QA/QC** section in the toolbox
* Click **Single Cell QA/QC**

This produces a *Single-cell QA/QC* task node

* Double-click the **Single cell QA/QC** task node to open the task report

The task report lists the number of counts per cell, the number of detected features per cell, the percentage of mitochondrial reads per cell, and the percentage of ribosomal counts per cell in four violin plots (Figure 6). For this analysis, we will set maximum and minimum thresholds for total counts and detected genes to exclude potential doublets and a maximum mitochondrial reads percentage filter to exclude potential dead or dying cells. There is no need to apply a filter based on the percentage of ribosomal counts in this tutorial.

* In the *Selection* card on the right, set the *Counts* threshold to keep cells between **1500** and **15000**
* Set the *Detected features* to keep cells between **400** and **4000**
* Set the % *Mitochondrial counts* to keep cells between **0%** and **20%** (Figure 6)

![Figure 6. Filtering low quality cells based on gene expression data](/files/S6s1IPltcxa65AnpROe3)

* Click ![](/files/a5nfgK3X8m8KwP4PQ53b) under *Filter* on the right
* Click **Apply observations filter**
* Select the **Gene Expression** data node as input in the pipeline preview
* Click **Select**
* Click **OK** to dismiss the message about the task being enqueued
* Click the **project name** at the top to go back to the *Analyses* tab
* Your browser may warn you that any unsaved changes to the data viewer session will be lost. Ignore this message and proceed to the *Analyses* tab

A new task, *Filter counts*, is added to the *Analyses* tab. This task produces a new *Filter counts* data node (Figure 7)

![Figure 7. Antibody Capture and Gene Expression data have been filtered to remove low quality cells](/files/E2p60SLaCcMq9FEwrFke)

## Normalization

After excluding low-quality cells, we can normalize the data.

We will start with the protein data.

* Click the **Filtered counts** data node produced by filtering the *Antibody Capture* data node
* Click **Normalization and scaling** in the toolbox
* Click **Normalization**
* Click the  button
* Click **Finish** to run (Figure 8)

![Figure 8. Recommended normalization for protein count data](/files/P6Q2nT2rfo9MKdMOEfu3)

The recommended normalization for protein data includes the following steps: Add 1, Divide by Geometric mean, Add 1, Log base 2. This is a variant of Centered log-ratio (CLR), which was used to normalize antibody capture protein counts data in the paper that introduced CITE-Seq \[1] and in subsequent publications on similar assays \[2. 3]. CLR normalization includes the following steps: Add 1, Divide by Geometric mean, Add 1, log base e. Normalizing the protein data to base 2 instead of e allows for better integration with gene expression data further downstream. If you would prefer to use CLR, click and drag CLR from the panel on the left to the right. If you do choose to use CLR, we recommend making sure the gene expression data is normalized to the base e, to allow for smoother integration further downstream.

Normalization produces a *Normalized counts* data node on the *Antibody Capture* branch of the pipeline.

Next, we can normalize the mRNA data. We will use the recommended normalization method in Partek Flow, which accounts for differences in library size, or the total number of UMI counts, per cell, adds 1 and log2 transforms the data.

* Click the **Filtered counts** data node produced by filtering the *Gene Expression* data node
* Click the **Normalization and scaling** section in the toolbox
* Click **Normalization**
* Click the ![](/files/SJ5Moeqnb2LdCG5wtMyo) button
* Click **Finish** to run (Figure 9)

![Figure 9. Recommended normalization for single cell gene expression data](/files/gCCyEcaizWaIFO67j61l)

Normalization produces a *Normalized counts* data node on the *Gene Expression* branch of the pipeline (Figure 10).

![Figure 10. The two normalization tasks produce Normalized counts data nodes](/files/kXS4yNyDiHFskUpPcyVY)

## Merge Protein and mRNA data

For quality filtering and normalization, we needed to have the two data types separate as the processing steps were distinct. For downstream analysis, we want to be able to analyze protein and mRNA data together. To bring the two data types back together, we will merge the two normalized counts data nodes.

* Click the **Normalized counts** data node on the *Antibody Capture* branch of the pipeline
* Click **Pre-analysis tools** in the toolbox
* Click **Merge matrices**
* Click **Select data node** to launch the data node selector

Data nodes that can be merged with the *Antibody Capture* branch *Normalized counts* data node are shown in color (Figure 11).

![Figure 11. Select the normalizated gene expression counts to merge the protein counts with](/files/OGCEhqeVgkKY3rqpPOE1)

* Click the **Normalized counts** data node on the *Gene Expression* branch of the pipeline (Figure 11)
* Click **Select**
* Click **Finish** to run the task

The output is a *Merged counts* data node (Figure 12). This data node will include the normalized counts of our protein and mRNA data. The intersection of cells from the two input data nodes is retained so only cells that passed the quality filter for both protein and mRNA data will be included in the *Merged counts* data node.

![Figure 12. Merged counts output](/files/UKZ2QBeSLLDoH3ndNmbf)

## Collapsing tasks to simplify the pipeline

To simplify the appearance of the pipeline, we can group task nodes into a single collapsed task. Here, we will collapse the filtering and normalization steps.

* Right-click the **Split by feature type** task node
* Choose **Collapse tasks** from the pop-up dialog (Figure 13)

![Figure 13. Choosing the first task node to generate a collapsed task](/files/vQqpArAOOXuEN0ao2MOS)

Tasks that can be selected for the beginning and end of the collapsed section of the pipeline are highlighted in purple (Figure 14). We have chosen the *Split matrix* task as the start and we can choose *Merge matrices* as the end of the collapsed section.

![Figure 14. Tasks that can be the start or end of a collapsed task are shown in purple](/files/XVcFT3UxovSnVuG0fGfA)

* Click the **Merge matrices task** to choose it as the end of the collapsed section
* Name the Collapsed task **Data processing**
* Click **Save** (Figure 15)

![Figure 15. Naming the collapsed task](/files/OmYmJRYPkOoUphrM5Ctq)

The new collapsed task, *Data processing*, appears as a single rectangular task node (Figure 16).

![Figure 16. Collapsed tasks are represented by a single task node](/files/I5h0KVUusNCbppApNk1W)

To view the tasks in *Data processing,* we can expand the collapsed task.

* Double-click **Data processing** to expand it or right-click and choose **Expand collapsed task**

When expanded, the collapsed task is shown as a shaded section of the pipeline with a title bar (Figure 17).

![Figure 17. Expanding a collapsed task to show its components](/files/E2byU7KcpiPyhjP5nlD6)

To re-collapse the task, you can double click the title bar or click the ![](/files/3y6GRZyg4eK5ajrp1hCn) icon in the title bar. To remove the collapsed task, you can click the ![](/files/ipOd3jnCN34HHiiCK98K). Please note that this will not remove tasks, just the grouping.

* Double-click the *Data processing* title bar to re-collapse

## References

\[1] Stoeckius, M., Hafemeister, C., Stephenson, W., Houck-Loomis, B., Chattopadhyay, P. K., Swerdlow, H., ... & Smibert, P. (2017). Simultaneous epitope and transcriptome measurement in single cells. Nature methods, 14(9), 865.

\[2] Stoeckius, M., Zheng, S., Houck-Loomis, B., Hao, S., Yeung, B. Z., Mauck, W. M., ... & Satija, R. (2018). Cell hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics. Genome biology, 19(1), 224.

\[3] Mimitou, E., Cheng, A., Montalbano, A., Hao, S., Stoeckius, M., Legut, M., ... & Satija, R. (2018). Expanding the CITE-seq tool-kit: Detection of proteins, transcriptomes, clonotypes and CRISPR perturbations with multiplexing, in a single assay. bioRxiv, 466466.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Dimensionality Reduction and Clustering

* [PCA](#pca)
* [Graph-based clustering](#graph-based-clustering)
* [UMAP](#umap)
* [Notes on Performing Exploratory Analysis with Protein or Gene Expression Data Only](#notes-on-performing-exploratory-analysis-with-protein-or-gene-expression-data-only)

## PCA

Next, we will perform some exploratory analysis on the merged mRNA and protein expression data and visualize the data in preparation to identify cell populations. Because the merged count matrix has thousands of features, it is a good idea to reduce the dimensionality of the data for more efficient downstream processing.

* Click the **Merged counts** data node
* Click **Exploratory analysis** in the toolbox
* Click **PCA**
* Click **Finish** to run the PCA with default settings (Figure 1)

![Figure 1. Run PCA with default settings](/files/1MLdIGZcW38oUsXyCYW4)

A PCA task node will be added to the pipeline under the *Analyses* tab and a circular PCA output data node will be produced (Figure 2).

![Figure 2. PCA task run on the merged counts data node](/files/D6KE96LrEHnfKdwcIKku)

Once the task completes, we will inspect the results to decide the optimal number of principal components (PCs) to use in downstream analyses. To do this, we will use a Scree plot.

* Double click the **PCA** data node to open the task report

The PCA plot will open in a new data viewer session. A 3D scatterplot will be displayed on the canvas (Figure 3).

![Figure 3. Each dot is a different cell. Cells are clustered based on how similar their expression profile is across the combined mRNA and protein data](/files/aVsOQGi0YOAqdJJEbTlG)

* Click and drag the **Scree plot** from **New plot** under *Setup* on the left onto the canvas
* Drop it over the **Replace** option (Figure 4)

![Figure 4. Click and drag the Scree plot to replace the PCA plot on the canvas](/files/i6HJBr4WOBgwRJB2iEQj)

* Select **PCA** as data for the new Scree plot (Figure 5)

![Figure 5. The PCA data node contains the data to draw the Scree plot](/files/OgR9vkNXHK32JUQke9Yu)

The Scree plot (Figure 6) shows the eigenvalues on the y-axis for each of the 100 PCs on the x-axis. The higher the eigenvalue, the more variance explained by each PC. Typically, after an initial set of highly informative PCs, the amount of variance explained by analyzing additional components is minimal. By identifying the point where the Scree plot levels off, you can choose an optimal number of PCs to use in downstream analysis steps like graph-based clustering and UMAP.

![Figure 6. Scree plot shows the amount of variation explained by each principal component](/files/apcbZZhUGes9wgfdBkLC)

* Click and drag over the first set of PCs to zoom in (Figure 7)

![Figure 7. Click and drag on the Scree plot to zoom in and see the first set of principal components](/files/FG8GtG6NvzVORcRqoJlW)

* Mouse over the Scree plot to identify the point where additional PCs offer little additional information (Figure 8)

In this data set, a reasonable cut-off could be set anywhere between around 10 and 30 PCs. We will use 15 in downstream steps.

![Figure 8. Identifying the optimal number of PCs](/files/XqLD9mN8DquulT1LofCw)

## Graph-based clustering

We can use Graph-based clustering to group similar cells together in an unsupervised manner.

* Click the **project name** near the top to go back to the *Analyses* tab
* Click the circular **PCA** data node
* Click **Exploratory analysis** in the toolbox
* Click **Graph-based clustering**
* Click to **Compute biomarkers**
* Set the number of principal components to **15** (Figure 9)
* Click **Configure** under *Advanced options* and change the *Resolution* to **1.0**
* Click **Finish** to run the task

![Figure 9. Graph-based clustering task set up. Reduce the number of PCs to 15](/files/SAzWYu3gAiMt9PjRoUhv)

A *Graph-based clustering* task node will be added to the pipeline under the *Analyses* tab and a circular *Graph-based clusters* output data node will be produced (Figure 10)

![Figure 10. Graph-based clustering task and output data nodes](/files/DJYfBFRrQQXdubxOFPud)

## UMAP

Once the graph-based clustering task has completed, we can visualize the results with a UMAP plot. You could use the same steps here to generate a t-SNE plot. For this tutorial, we will use UMAP, as it is faster on several thousand cells.

* Click the circular **PCA** data node
* Click **Exploratory analysis** in the toolbox
* Click **UMAP**
* Set the number of principal components to **15** (Figure 11)
* Click **Finish** to run the task

![Figure 11. UMAP task set up. Reduce the number of PCs to 15.](/files/YEhiCO4fLhpOGElWkT23)

A *UMAP* task node will be added to the pipeline under the *Analyses* tab and a circular *UMAP* output data node will be produced (Figure 12)

![Figure 12. UMAP task and output data node](/files/de5E5QYfoUXYy5Z5irfh)

## Notes on Performing Exploratory Analysis with Protein or Gene Expression Data Only

In this tutorial, we have performed exploratory analysis on merged protein and gene expression data, and we will perform classification on the merged data in the next step.

It can be interesting to perform exploratory analysis on the two feature types separately. For example, you might be interested to see how the clustering of the same cells differs between protein expression profiles vs. gene expression profiles.

To perform exploratory analysis on the two feature types separately, select the *Merged counts* data node, click *Pre-analysis tools*, followed by *Split by feature type* from the toolbox. A new task, *Split by feature type,* will be added to the pipeline resulting in two output data nodes: *Antibody capture* (protein data) and *Gene expression* (mRNA data). Both contain the same high-quality cells.

Performing exploratory analysis with gene expression data is the same as for the merged counts. Because there are a large number of genes, you will need to reduce the dimensionality with PCA, choose an optimal number of PCs and perform downstream clustering and visualization (e.g. graph-based clustering and UMAP/t-SNE). Performing exploratory analysis with protein data is different. There is no need to reduce the dimensionality as there are only a handful of features (17 proteins in this case), so you can proceed straight to downstream clustering and visualization. Figure 13 shows an example of how the pipeline might look if the data is split and analyzed separately.

![Figure 13. Example of how the pipeline might look if you split the merged counts and perform exploratory analysis for protein and gene expression data separately](/files/vUiOsPPnyBf4wNB4a8tr)

You can then use the *Data viewer* to bring together multiple plots for comparison (Figure 14).

![Figure 14. Comparison of 2D UMAP plots for the same cells clustered on protein, mRNA and merged data. All cells are coloured based on their expression of the CD3D gene (in blue). Note, the plots in this figure may differ from the default UMAP plots because these are 2D plots. Default UMAP plots re in 3D.](/files/vIjvkynuprSExYl5HwGm)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Classifying Cells

* [Exploratory Analysis Results](#exploratory-analysis-results)
* [T cells](#t-cells)
* [B cells](#b-cells)

We will now examine the results of our exploratory analysis and use a combination of techniques to classify different subsets of T and B cells in the MALT sample.

## Exploratory Analysis Results

* Double click the merged **UMAP** data node
* Under *Configure* on the left, click **Style,** select the **Graph-based cluster** node, and color by the **Graph-based** attribute (Figure 1)

![Figure 1. Color the cells in the UMAP plot by their graph-based cluster assignment](/files/mgLFJQJBvj0Jc0EdtXPY)

The 3D UMAP plot opens in a new data viewer session (Figure 2). Each point is a different cell and they are clustered based on how similar their expression profiles are across proteins and genes. Because a graph-based clustering task was performed upstream, a biomarker table is also displayed under the plot. This table lists the proteins and genes that are most highly expressed in each graph-based cluster. The graph-based clustering found 11 clusters, so there are 11 columns in the biomarker table.

* Click and drag the **2D scatter plot** icon from **New plot** onto the canvas (Figure 2)
* Drop the 2D scatter plot to the **right** of the UMAP plot

![Figure 2. Add a 2D scatter plot and place it to the right of the UMAP plot](/files/z1ZndDxdZSyTGgalK8qC)

* Click **Merged counts** to use as data for the 2D scatter plot (Figure 3)

![Figure 3. Choose Merged counts data to draw the 2D scatter plot](/files/c60vYONwKs3lXHiqUrmF)

A 2D scatter plot has been added to the right of the UMAP plot. The points in the 2D scatter plot are the same cells as in the UMAP, but they are positioned along the x- and y-axes according to their expression level for two protein markers: CD3\_TotalSeqB and CD4\_TotalSeqB, respectively (Figure 4).

![Figure 4. The canvas now has a 2D scatter plot next to the UMAP](/files/zMcb1UGMl2qk4Jq2CnVW)

* In **Select & Filter**, click **Criteria** to change the selection mode
* Click the **blue circle** next to the *Add rule* drop-down menu (Figure 5)

![Figure 5. Click the blue circle to change the data source for the rule selector](/files/ikRViBQWqg2D46DJbcEr)

* Click **Merged counts** to change the data source
* Choose **CD3\_TotalSeqB** from the drop-down list (Figure 6)

![Figure 6. Choose the CD3\_TotalSeqB protein marker as a selection rule](/files/pg4wvZEkkPXFkl5g0SBF)

* Click and drag the **slider** on the CD3D\_TotalSeqB selection rule to include the CD3 positive cells (Figure 7)

![Figure 7. Use the slider to select cells with positive expression for the CD3 protein marker](/files/L3mXHxVgpsoEzutBBHhJ)

As you move the slider up and down, the corresponding points on both plots will dynamically update. The cells with a high expression for the CD3 protein marker (a marker for T cells) are highlighted and the deselected points are dimmed (Figure 8).

![Figure 8. CD3+ cells are selected on both plots](/files/gorAdBSJGvS6WHBnQ9Lj)

* Click **Merged counts** in **Get data** on the left under *Setup*
* Click and drag **CD8a\_TotalSeqB** onto the 2D scatter plot (Figure 9)
* Drop CD8\_TotalSeqB onto the **x-axis** configuration option

![Figure 9. Change the feature plotted on the x-axis to CD8\_TotalSeqB](/files/Dxm7CRJcV3ppxt2EyXvh)

The CD3 positive cells are still selected, but now you can see how they separate into CD4 and CD8 positive populations (Figure 10).

![Figure 10. 2D scatter plot with CD4\_TotalSeqB and CD8\_TotalSeqB features on the axes](/files/LUdSPX2h3tA56pGykUFO)

The simplest way to classifying cell types is to look for the expression of key marker genes or proteins. This approach is more effective with CITE-Seq data than with gene expression data alone as the protein expression data has a better dynamic range and is less sparse. Additionally, many cell types have expected cell surface marker profiles established using other technologies such as flow cytometry or CyTOF. Let's compare the resolution power of the CD4 and CD8A gene expression markers compared to their protein counterparts.

* Click the **duplicate plot** icon above the 2D scatter plot (Figure 11)

![Figure 11. Click the duplicate plot icon to make a copy of the 2D scatter plot](/files/U7ugT2fqeOZfYrq8xz2L)

* Click **Merged counts** in the **Get Data** icon under *Setup*
* Search for the **CD4** gene
* Click and drag **CD4** onto the duplicated 2D scatter plot
* Drop the CD4 gene onto the **y-axis** option
* Search for the **CD8A** gene
* Click and drag **CD8A** onto the duplicated 2D scatter plot
* Drop the CD8A gene onto the **x-axis** option

The second 2D scatter plot has the CD8A and CD4 mRNA markers on the x- and y-axis, respectively (Figure 12). The protein expression data has a better dynamic range than the gene expression data, making it easier to identify sub-populations.

![Figure 12. The second 2D scatter plot (bottom) has the CD8 and CD4 genes plotted against each other](/files/qIph511tvWW1HY5E581h)

* On the first 2D scatter plot (with protein markers), click ![](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner
* Manually select the cells with high expression of the CD4\_TotalSeqB protein marker (Figure 13)

More than 2000 cells show positive expression for the CD4 cell surface protein.

![Figure 13. Draw a lasso to manually select CD4+ cells, based on protein expression](/files/zQBwp8L8gPGO2rKPjG1b)

Let's perform the same test on the gene expression data.

* Click ![](/files/sSBYPOuizlIJf7OXvD1R) in the top right of the plot to switch back to pointer mode
* Click on a blank spot on the plot to clear the selection
* On the second 2D scatter plot (with mRNA markers), click ![](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner
* Manually select the cells with high expression of the CD4 gene marker (Figure 14)

![Figure 14. Draw a lasso to manually select CD4+ (mRNA) cells](/files/qBlD5G0kL1HiwhguoJEi)

This time, only 500 cells show positive expression for the CD4 marker gene. This means that the protein data is less sparse (i.e. there fewer zero counts), which further helps to reliably detect sub-populations.

## T cells

Based on the exploratory analysis above, most of the CD3 positive cells are in the group of cells in the right side of the UMAP plot. This is likely to be a group of T cells. We will now examine this group in more detail to identify T cell sub-populations.

* Click ![](/files/ipOd3jnCN34HHiiCK98K) in the top right corner of both 2D scatter plots, to remove them from the canvas
* Click ![](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner of the 3D UMAP plot
* Draw a lasso around the group of putative T cells (Figure 15)

![Figure 15. Select the group of putative T cells](/files/C09ZhbKyCqbFLEd0PIxT)

* Click ![](/files/zdAVcV8BSPFAx45UsZIH) in the **Select & Filter** tool to include the selected points
* Click ![](/files/sSBYPOuizlIJf7OXvD1R) in the top right of the plot to switch back to pointer mode
* Click and drag the plot to rotate it around

![Figure 16. Group of putative T-cells](/files/uRmvB2Y842GVO7O2jWXU)

This group of putative T cells predominantly consists of cells assigned to graph-based clusters 3, 4, and 6, indicated by the colors. Examining the biomarker table for these clusters can help us infer different types of T cell.

* Add the Biomarkers table using the **Table** option in the **New plot** menu, you can drag and reposition the table using the button in the top left corner of the plot ![](/files/y5O9JHnCiFWKaMemy47J).
* Click and drag the bar between the UMAP plot and the biomarker table to resize the biomarker table to see more of it (Figure 17)

If you need to create more space on the canvas, hide the panel words on the left using the arrow ![](/files/30wFobreUYLPYGs4O4QX).

![Figure 17. Resize plots to see more of the biomarker table](/files/V2MlxGxnrl5j8zzTt0xY)

Cluster 6 has several interesting biomarkers. The top biomarker is CXCL13, a gene expressed by follicular B helper T cells (Tfh cells). Another biomarker is the PD-1 protein, which is expressed in Tfh cells. This protein promotes self-tolerance and is a target for immunotherapy drugs. The TIGIT protein is also expressed in cluster 6 and is another immunotherapy drug target that promotes self-tolerance.

Cluster 4 expresses several marker genes associated with cytotoxicity (e.g. NKG7 and GZMA) and both CD3 and CD8 proteins. Thus, these are likely to be cytotoxic cells.

We can visually confirm these expression patterns and assess the specificity of these markers by coloring the cells on the UMAP plot based on their expression of these markers.

* Click the **duplicate plot** icon above the UMAP plot

We will color the cells on the duplicate by their expression of marker genes, while keeping the original plot colored by graph-based cluster assignment.

* Click and drag the **CXCL13** gene from the biomarker table onto the duplicate UMAP plot
* Drop the CXCL13 gene onto the **Green (feature)** option (Figure 18)

![Figure 18. Click and drag the gene from the biomarker table onto the plot](/files/sBtJhNSoblNcdBCyaCxl)

* Click and drag the **NKG7** gene from the biomarker table onto the duplicate UMAP plot
* Drop the NKG7 gene onto the **Red (feature)** option

The cells with higher CXCL13 and NKG7 expression are now colored green and red, respectively. By looking at the two UMAP plots side by side, you can see these two marker genes are localized in graph-based clusters 6 and 4, respectively (Figure 19).

![Figure 19. The cells in the UMAP plot on the right are colored by their expression of CXCL13 (green) and NKG7 (red) marker genes. These cells belong to graph-based clusters 6 and 4, respectively, shown in the plot on the left](/files/h4iOEAGqTdKG6e63QJIi)

* In **Select & Filter**, click ![](/files/3ri6MoSm3tqb4QpoRgf7) to remove the CD3\_TotalSeqB filtering rule
* Click the **blue circle** next to the *Add criteria* drop-down list
* Search for **Graph** to search for a data source
* Select **Graph-based clustering** (derived from the Merged counts > PCA data nodes)
* Click the **Add criteria** drop-down list and choose **Graph-based** to add a selection rule (Figure 20)

![Figure 20. Change the data source to Graph-based clustering and choose Graph-based from the drop-down list](/files/NcfQOIkfLF0kvKNEvP4Q)

* In the *Graph-based* filtering rule, click **All** to deselect all cells
* Click cluster **6** to select all cells in cluster 6
* Using the **Classify** tool, click **Classify selection**
* Label the cells as **Tfh** **cells** (Figure 21)
* Click **Save**

![Figure 21. Select all cluster 6 cells and classify them as Tfh cells](/files/keMgfrRKgDcbtO1GQH7C)

* Click ![](/files/PMIRI9cDHX7OlasuG1p4) in **Select & Filter** to exclude the cluster 6/Tfh cells
* Click cluster **4** to select all cells in cluster 4
* In the **Classify** icon, click **Classify selection**
* Label the cells as **Cytotoxic cells**
* Click **Save**
* Click ![](/files/PMIRI9cDHX7OlasuG1p4) in **Select & Filter** to exclude the cluster 4/Cytotoxic cells

We can classify the remaining cells as helper T cells, as they predominantly express the CD4 protein marker.

* Click on the **invert selection** icon in either of the UMAP plots (Figure 22)

![Figure 22. Invert the selection to select all remaining cells](/files/QxcskMpAQc3exUJZ0uRL)

* In **Classify**, click **Classify selection**
* Label the cells as **Helper T cells**
* Click **Save**

Let's look at our progress so far, before we classify subsets of B-cells.

* Click the **Clear filters** link in **Select & Filter**
* Select the duplicate UMAP plot (with the cell colored by marker genes)
* Under *Configure* on the left, open **Style** and color the cells by **New classifications** (Figure 23)

![Figure 23. Color by New classifications (T cell subsets)](/files/tYr4m5vf6tbdSFZ7UqRC)

## B cells

In addition to T-cells, we would expect to see B lymphocytes, at least some of which are malignant, in a MALT tumor sample. We can color the plot by expression of a B cell marker to locate these cells on the UMAP plot.

* In the **Get data** icon on the left, click **Merged counts**
* Scroll down or use the search bar to find the **CD19\_TotalSeqB** protein marker
* Click and drag the **CD19\_TotalSeqB** marker over to the UMAP plot on the right
* Drop the CD19\_TotalSeqB marker over the **Color** configuration option on the plot

The cells in the UMAP plot are now colored from grey to blue according to their expression level for the CD19 protein marker (Figure 24). The CD19 positive cells correspond to several graph-based clusters. We can filter to these cells to examine them more closely,

![Figure 24. Cells in UMAP plot colored by their expression of CD19 protein](/files/J05ck7H4ySXrfaRmilxb)

* Click ![](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner of the UMAP plot
* Lasso around the CD19 positive cells (Figure 25)
* Click ![](/files/zdAVcV8BSPFAx45UsZIH) in **Select & Filter** to include the selected points

![Figure 25. Lasso around CD19 positive cells](/files/b8am3iLNIG9zmGeFSwoV)

The plots will rescale to include the selected points. The CD19 positive cells include cells from graph-based clusters 1, 2 and 7 (Figure 26).

![Figure 26. Filtered CD19 positive cells](/files/Q3fD2JNmAdjeHv5NnuIW)

* Find the **CD3\_TotalSeqB** protein marker in the biomarker table
* Click and drag the **CD3\_TotalSeqB** onto the UMAP plot on the right
* Drop the CD3\_TotalSeqB protein marker onto the **Color** configuration option on the plot (Figure 27)

While these cells express T cell markers, they also group closely with other putative B cells and express B cell markers (CD19). Therefore, these cells are likely to be doublets.

![Figure 27. Some cells within the CD19 positive clusters show signs of expressing T-cells markers](/files/kwGTIXQSg7sk6ftQpFiD)

* Select either of the UMAP plots
* Click on the **Select & Filter**
* Find the **CD3\_TotalSeqB** protein marker in the biomarker table
* Click and drag **CD3\_TotalSeqB** onto the **Add criteria** drop-down list in **Select & Filter** (Figure 28)
* Set the minimum threshold to **3** in the CD3\_TotalSeqB selection (Figure 29)
* Click the **Classify** icon then click **Classify selection**
* Label the cells as **Doublets**
* Click **Save**
* Click ![](/files/PMIRI9cDHX7OlasuG1p4) in **Select & Filter** to exclude the selected points

![Figure 28. Click and drag the CD3 protein marker directly onto the Add criteria drop-down list to create a selection criteria](/files/30fF0KVya6iZOwgyB78P)

![Figure 29. Select the remaining CD3 positive doublet cells](/files/IHxQSqVBBY73BRZ7g79i)

The biomarkers for clusters 1 and 2 also show an interesting pattern. Cluster 1 lists IGHD as its top biomarker, while cluster 2 lists IGHA1 as the fourth most significant. Both IGHD (Immunoglobulin Heavy Constant Delta) and IGHA1 (Immunoglobulin Heavy Constant Alpha 1) encode classes of the immunoglobulin heavy chain constant region. IGHD is part of IgD, which is expressed by mature B cells, and IGHA1 is part of IgA1, which is expressed by activated B cells. We can color the plot by both of these genes to visualize their expression.

* Click, drag and drop **IGHD** from the biomarker table onto the **Green (feature)** configuration option on the UMAP plot on the right
* Click, drag and drop **IGHA1** from the biomarker table onto the **Red (feature)** configuration option on the UMAP plot on the right (Figure 30)

![Figure 30. The B cells colored by IGHD (green) and IGHA1 (red) gene expression](/files/obrmyQN2DLpxtlp3snaA)

We can use the lasso tool to select and classify these populations.

* Click ![](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner of the UMAP plot
* Lasso around the IGHD positive cells (Figure 31)
* In the **Classify** icon on the left, click **Classify selection**
* Label the cells as **Mature B cells**
* Click **Save**

![Figure 31. Lasso around the IGHD positive cells](/files/HwZKkICsCwgBjpdAdvv0)

* Lasso around the IGHA1 positive cells (Figure 32)
* In the **Classify** icon on the left, click **Classify selection**
* Label the cells as **Activated B cells**
* Click **Save**

![Figure 32. Select IGHA1 positive cells](/files/ZqOhQOqnGpJ5lMRHtX36)

We can now visualize our classifications.

* Click the **Clear filters** link in the **Select & Filter** icon on the left
* Select the duplicate UMAP plot (with the cell colored by marker genes)
* Under *Configure* on the left, click the **Style** icon and color the cells by **New classifications** (Figure 33)

![Figure 33. UMAP with cells colored by cell types](/files/YuUdUnJB5fo5DV42xp9R)

* Click **Apply classifications** in the **Classify** icon
* Name the attribute **Cell type**
* Click **Run**
* Click **OK** to close the message about a classification task being enqueued

Optionally, you may wish to save this data viewer session if you need to go back and reclassify cells later. To save the session, click the ![](/files/ZsuYfFNhtj9gY9W3dPui) icon on the left and name the session.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Differentially Expressed Proteins and Genes

* [Filter Groups](#filter-groups)
* [Re-split the Matrix](#re-split-the-matrix)
* [Differential Analysis and Visualization - Protein Data](#differential-analysis-and-visualization---protein-data)
* [Differential Analysis, Visualization, and Pathway analysis - Gene Expression Data](#differential-analysis-visualization-and-pathway-analysis---gene-expression-data)

Next, we will filter out certain cells and re-split the data. Re-splitting the data can be useful if you want to perform differential analysis and downstream analysis separately for proteins and genes. For your own analyses, re-splitting the data is optional. You could just as well continue with differential analysis with the merged data if you prefer.

## Filter Groups

Because we have classified our cells, we can now filter based on those classifications. This can be used to focus on a single cell type for re-clustering and sub-classification or to exclude cells that are not of interest for downstream analysis.

* Click the **Merged counts** data node
* Click **Filtering**
* Click **Filter cells**
* Set to **exclude** **Cell type is Doublets** using the drop-down menus
* Click **OR**
* Set the second filter to **exclude Cell type is N/A** using the drop-down menus
* Click **Finish** to apply the filter (Figure 1)

![Figure 1. Set up the Filter groups task to exlcude Doublets and cells that are not classified](/files/zSjHXMzdZLYEkBb0xQ3w)

This produces a *Filtered counts* data node (Figure 2).

![Figure 2. Filter groups output](/files/Mc3KodxqqVMpTKBVNc1S)

## Re-split the Matrix

* Click the **Filtered counts** data node
* Click **Pre-analysis tools**
* Click **Split by feature type**

This will produce two data nodes, one for each data type (Figure 3). The split data nodes will both retain cell classification information.

![Figure 3. It is possible to re-split the merged matrix once again](/files/dQ4Gpo9wHoLnXqJXdes1)

## Differential Analysis and Visualization - Protein Data

Once we have classified our cells, we can use this information to perform comparisons between cell types or between experimental groups for a cell type. In this project, we only have a single sample, so we will compare cell types.

* Click the **Antibody Capture** data node
* Click **Statistics**
* Click **Differential analysis**
* Click **ANOVA** then click **Next**

The first step is to choose which attributes we want to consider in the statistical test.

* Click **Cell type**
* Click **Add factor**
* Click **Next**

Next, we will set up the comparison we want to make. Here, we will compare the Activated and Mature B cells.

* Drag **Activated B cells** in the top panel
* Drag **Mature B cells** in the bottom panel
* Click **Add comparison**

The comparison should appear in the table as *Activated B cells vs. Mature B cells*.

* Click **Finish** to run the statistical test (Figure 4)

![Figure 4. Setting up a comparison for differentially expressed proteins](/files/cI6A1twrrrC6XQuyFdP7)

The *ANOVA* task produces an *ANOVA* data node.

* Double-click the **ANOVA** data node to open the task report

The report lists each feature tested, giving p-value, false discovery rate adjusted p-value (FDR step up), and fold change values for each comparison (Figure 5).

![Figure 5. GSA report for protein expression data](/files/a2Adzx8SQfVZdFZhx3XW)

In addition to the listed information, we can access dot and violin plots for each gene or protein from this table.

* Click ![image2019-5-24 14\_50\_50](/files/lVLROyObYQRWa25mr62I) in the *CD45RA\_TotalSeqB* row

This opens a dot plot in a new data viewer session, showing CD45A expression for cells in each of the classifications (Figure 6). First, we exclude *Doublets* and *N/A* cells from the plot:

* Open **Select and filter**, select **Criteria**
* Drag "Cell type" from the legend title to the **Add criteria** box
* Uncheck **Doublets** and **N/A**
* Click to include selected points

![Figure 6. CD45RA dot plot for all cells](/files/o0hbE6Ike6NZyivjxncR)

We can use the *Configuration* panel on the left to edit this plot.

* Open the **Style** icon
* Switch on **Violins** under *Summary*
* Switch on **Overlay** under *Summary*
* Switch on **Colored** under *Summary*
* Select the *Graph-based clustering* node in the **Color by** section
* **Color by** Graph-based clusters under **Color** and use the slider to decrease the **Opacity**
* Open the **Axes** icon
* Select the *Graph-based clustering* node in the **X axis** section
* Change the *X axis data* to Graph-based clusters
* Use the slider to increase the **Jitter** on the *X axis* (Figure 7)

![Figure 7. Configure the dot plot using the tools on the left](/files/jyQemo2ti0beYJfSuiz9)

* Click the **project name** to return to the *Analyses* tab

To visualize all of the proteins at the same time, we can make a hierarchical clustering heat map.

* Click the **ANOVA** data node
* Click **Exploratory analysis** in the toolbox
* Click **Hierarchical clustering/heatmap**
* In the *Cell order* section, choose **Graph-based clusters** from the *Assign order* drop-down list
* Click **Finish** to run with the other default settings
* Double-click the **Hierarchical clustering** task node to open the heatmap

The heatmap can easily be customized using the tools on the left.

* Open the **Axes** icon
* Switch off *Show* **Row labels**
* Increase the **Font** to 16 (Figure 8)

![Figure 8. Heatmap showing altered Axes labels](/files/QxJtvzHEtCKmzEIdIbRn)

* Activate the **Transpose** switch which will switch the Row and Column labels, so now the Row labels will be shown (Figure 9)

![Figure 9. Transpose the Heatmap to switch the columns and rows](/files/9DDdSYZef68yjCcLPWA2)

* Open the **Dendrograms** icon
* Choose *Row color* **By cluster** and change *Row clusters* to **4**
* Change *Row dendrogram size* to **80** (Figure 10)

![Figure 10. Configure the Dendrograms settings](/files/yFFqftdoywL9OqjFzuSu)

* In the **Heatmap** icon
* Navigate to *Range* under *Color*
* Set the Min and Max to **-1.2** and **1.2**, respectively
* Change the *Shape* to **Circle** (Figure 11)

![Figure 11. Configure the Heatmap icon](/files/JLRAJZ87V4XSTwpXv2PK)

* Switch the *Shape* back to **Rectangle**
* Change the *Color Palette* by clicking on the color squares and selecting colors from the rainbow. Click outside of the selection box to exit this selection. The color options can be dragged alone the Palette to highlight value differences (Figure 12).

![Figure 12. Heatmap showing expression of protein markers after changing the Heatmap settings further](/files/p1pbjy074efuue3a0D2N)

Feel free to explore the other tool options on the left to customize the plot further.

## Differential Analysis, Visualization, and Pathway analysis - Gene Expression Data

We can use a similar approach to analyze the gene expression data.

* Click the **project name** to return to the *Analyses* tab
* Click the **Gene Expression** data node
* Click the **Antibody Capture** data node
* Click **Statistics**
* Click **Differential analysis**
* Click **ANOVA** then click **Next**
* Click **Cell type**
* Click **Add factor**
* Click **Next**
* Drag **Activated B cells** in the top panel
* Drag **Mature B cells** in the bottom panel
* Click **Add comparison**

The comparison should appear in the table as *Activated B cells vs. Mature B cells*.

* Click **Finish** to run the statistical test

As before, this will generate an *ANOVA* task node and n *ANOVA* data node.

* Double-click the **ANOVA** task node to open the task report (Figure 13)

![Figure 13. GSA report for the gene expression data](/files/FTzXC9veRfLp6KNRiGCp)

Because more than 20,000 genes have been analyzed, it is useful to use a volcano plot to get an idea about the overall changes.

* Click ![image2019-5-24 15\_5\_13](/files/keaDTokBfjJKEUxgY2qO) in the top right corner of the table to open a volcano plot

The Volcano plot opens in a new data viewer session, in a new tab in the web browser. It shows each gene as a point with cutoff lines set for P-value (y-axis) and fold-change (x-axis). By default, the P-value cutoff is set to 0.05 and the fold-change cutoff is set at |2| (Figure 14).

The plot can be configured using various tools on the left. For example, the **Style** icon can be used to change the appearance of the points. The X and Y-axes can be changed in the **Axes** icon. The **Statistics** icon can be used to set different Fold-change and P-value thresholds for coloring up/down-regulated genes. The in plot controls can be used to transpose ![image2022-8-30\_9-58-46](/files/WHp5lR20imwiCRXzjQcC) the volcano plot (Figure 14).

![Figure 14. The volcano plot can be Configured using the icons on the left and in plot controls](/files/CT1bPnj0LSzM4T1sKrCn)

* Click the **ANOVA report** tab in your web browser to return to the full report

We can filter the full set of genes to include only the significantly different genes using the filter panel on the left.

* Click **FDR step up**
* Type **0.05** for the cutoff and press **Enter** on your keyboard
* Click **Fold change**
* Set to From **-2** to **2** and press **Enter** on your keyboard

The number at the top of the filter will update to show the number of included genes (Figure 15).

![Figure 15. Use the panel on the left to filter the list for significant genes](/files/MX64g8mBGYaFK6wkz1xY)

* Click ![Screenshot 2023-09-25 at 10 01 54](/files/KTIuGBSW2eVUhh3ZWtI8) to create a new data node including only these significantly different genes

A task, *Differential analysis filter*, will run and generate a new *Filtered* *Feature list* data node. We can get a better idea about the biology underlying these gene expression changes using gene set or pathway enrichment. Note, you need to have the Pathway toolkit enabled to perform the next steps.

* Click the **Filtered feature list** data node
* Click **Biological interpretation** in the toolbox
* Click **Pathway enrichment**
* Make sure that **Homo sapiens** is selected in the *Species* drop-down menu
* Click **Finish** to run
* Double-click the **Pathway enrichment** task node to open the task report

The pathway enrichment results list KEGG pathways, giving an enrichment score and p-value for each (Figure 16).

![Figure 16. Results of pathway enrichment test](/files/LyrzG7GFEr3xl0of28SD)

To get a better idea about the changes in each enriched pathway, we can view an interactive KEGG pathway map.

* Click **path:hsa05202** in the Transcriptional misregulation in cancer row

The KEGG pathway map shows up-regulated genes from the input list in red and down-regulated genes from the input list in green (Figure 17).

![Figure 17. Transcriptional misregulation in cancer pathway with significant genes highlighted in green and red](/files/uhqFypD1QpYcr9P32HPc)

![Figure 18. Final CITE-Seq pipeline](/files/14oXdAeymzWLQ71TnI6u)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# 10x Genomics Visium Spatial Data Analysis

In this tutorial, we demonstrate how to:

* [Start with pre-processed Space Ranger output files](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/start-with-pre-processed-space-ranger-output-files)
* [Start with 10x Genomics Visium fastq files](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/start-with-10x-genomics-visium-fastq-files/README.md)
* [Spatial data analysis steps](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/spatial-data-analysis-steps/README.md)
* [View tissue images](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/view-tissue-images/README.md)

## Tutorial Data Set

The tutorial data is based on [10x Genomics Datasets](https://www.10xgenomics.com/resources/datasets?query=\&page=1\&configure%5BhitsPerPage%5D=50\&configure%5BmaxValuesPerFacet%5D=1000).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Start with pre-processed Space Ranger output files

Space Ranger output files are pre-processed 10x Genomics Visium data. The steps covered here will show you how to import and continue analyses with this pre-processed data from the Space Ranger pipeline. Partek Flow refers to this high cellular resolution data as *Single cell counts*; each point (spot) can be 1-10 cell resolution depending on the tissue type\*.

## Add files to the project

The project includes [Human Colon Cancer (Replicate 1)](https://www.10xgenomics.com/resources/datasets/visium-cytassist-gene-expression-libraries-of-post-xenium-human-colon-cancer-ffpe-using-the-human-whole-transcriptome-probe-set-2-standard) and [Human Colon Cancer (Replicate 2)](https://www.10xgenomics.com/resources/datasets/visium-cytassist-gene-expression-libraries-of-post-xenium-human-colon-cancer-ffpe-using-the-human-whole-transcriptome-probe-set-2-standard) output files in one project.

* Obtain the filtered **Count matrix files** **(h5 or HDF5) files** and **Spatial outputs** for each sample

![](/files/kxY8uO3He8y85EWMq27r)

The spatial imaging outputs should be in compressed format.

* Navigate the options to select **10x Genomics Visium Space Ranger output** as the file format for input

![](/files/bqifu6SvPnMwCiz0Npms)

* Click [**Transfer files**](/partek-flow/user-manual/importing-data) on the homepage, under settings, or during import

Proceed to transfer files as shown below using the **10x Genomics Visium Space Ranger outputs** importer.

![](/files/u63nlWDLStlok43esBHI)

* Navigate to the appropriate files for each sample.

***Please note that the 10x Genomics Space Ranger output can be count matrix data as 1 filtered .h5 file per sample or sparse matrix files for each sample as 3 files (two .csv with one .mtx or two .tsv with one .mtx for each sample).***

***The spatial output files should be compressed in one .gz or zip file when uploaded to Partek Flow server. This should include these related image files: tissue\_hires\_image.png, tissue\_lowres\_image.png, aligned\_fiducials.jpg, detected\_tissue\_image.jpg, tissue\_positions\_list.csv, scalefactors\_json.json.***

***The high resolution image can be uploaded and is optional.***

![](/files/6wp1Ua9YrjhGanc7UF8t)

Count matrix files and spatial outputs should be included for each sample. Once added, the *Cells* and *Features* values will update.

Once the download completes, the sample table will appear in the *Metadata* tab, with one row per sample.

![](/files/njVDPurjjuRptRpDf3jU)

The sample table is pre-populated with sample attributes, *# Cells*. Sample attributes can be added and edited manually by clicking *Manage* in the *Sample attributes* menu on the left. If a new attribute is added, click *Assign values* to assign samples to different groups. Alternatively, you can use the *Assign values from a file* option to assign sample attributes using a tab-delimited text file. For more information about sample attributes, see [here](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab).

For this tutorial, we do not need to edit or change any sample attributes.

## Resources

* [More about 10x Genomics Visium Spatial Gene Expression](https://www.10xgenomics.com/products/spatial-gene-expression?utm_medium=search\&utm_medium=search\&utm_medium=search\&utm_source=google\&utm_source=google\&utm_source=google\&utm_campaign=sem-goog-2022-website-page-ra_g-p_visium-brand-core\&utm_campaign=sem-goog-2022-website-page-ra_g-p_visium-brand-core\&utm_campaign=sem-goog-2022-website-page-ra_g-p_visium-brand-core\&useroffertype=website-page\&useroffertype=website-page\&useroffertype=website-page\&userresearcharea=ra_g\&userresearcharea=ra_g\&userresearcharea=ra_g\&userregion=multi\&userregion=multi\&userregion=multi\&userrecipient=customer\&userrecipient=customer\&userrecipient=customer\&usercampaignid=7011P0000013tOiQAI\&usercampaignid=7011P0000013tOiQAI\&usercampaignid=7011P0000013tOiQAI\&gad_source=1\&gclid=CjwKCAiApuCrBhAuEiwA8VJ6JvRsyo49FQ_fBqAiVGxwcPDR2KqyU2tHZpuKYqwQPbMP9Rf8PnswPxoCJaYQAvD_BwE)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Start with 10x Genomics Visium fastq files

* [Add files to the project](#add-files-to-the-project)
* [Pre-processing the unaligned fastq files with Space Ranger](#pre-processing-the-unaligned-fastq-files-with-space-ranger)
* [Annotate Visium image](#annotate-visium-image)

The fastq files are not pre-processed. The steps covered here will show you how to import and pre-process of the Visium Spatial Gene Expression data with brightfield and fluorescence microscope images.

## Add files to the project

The sample used for this tutorial can be found in the [10x Genomics Datasets](https://www.10xgenomics.com/resources/datasets?query=\&page=1\&configure%5Bfacets%5D%5B0%5D=chemistryVersionAndThroughput\&configure%5Bfacets%5D%5B1%5D=pipeline.version\&configure%5BhitsPerPage%5D=50\&configure%5BmaxValuesPerFacet%5D=1000\&menu%5Bproducts.name%5D=Single%20Cell%20Immune%20Profiling\&refinementList%5Bproduct.name%5D=). We will use the [Control, replicate 1 mouse brain sample](https://www.10xgenomics.com/resources/datasets/visium-cytassist-gene-expression-libraries-of-post-xenium-mouse-brain-ff-using-the-mouse-whole-transcriptome-probe-set-2-standard).

* Choose the **10x Genomics Visium fastq** import format
* Click **Next**

![](/files/qWEsxUNyLfx2uqqNAZYU)

If you have not [transferred files to the server](/partek-flow/user-manual/importing-data) already, [click here for more details](/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server) and choose to **Transfer files to the server**.

* Select the fastq files in the upload folder used for file transfer (select all sample files at one time; including R1 and R2 for each sample)
* Click **Finish**

The prefix used for R1 and R2 fastq files should match; one sample is shown in this example.

![](/files/BKSWTMQlwACNvDfaMZ9V)

The fastq files will be imported into the project as an *Unaligned reads* node.

![](/files/ObyF9873RYiS9d0LnLmj)

## Pre-processing the unaligned fastq files with Space Ranger

The unaligned reads must be preprocessed before proceeding with the analysis steps covered here: [Spatial data analysis](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/spatial-data-analysis-steps).

* From the unaligned reads node, select **Space Ranger** from the *10x Genomics* drop-down in the toolbox.

![](/files/RgH3GscNQjyKuNcm6Wr9)

[For more information about Space Ranger click here.](/partek-flow/user-manual/task-menu/10x-genomics/space-ranger)

* Specify the type of 10x Visium assay; this tutorial uses the *Visium CytAssist gene expression* library as the assay type
* If you have not done so already, a Cell Ranger reference should be created
* Specify the Reference assembly
* Select the Image and Probe set files that have already been transferred to the server for all samples
* Choose *visium-2-large* as the *Slide parameter* because this Visium CystAssist sample used a 11 x 11 slide capture area
* Click Finish

![](/files/rwSnPp7oCm6vttTok08J)

The *Space Ranger* task output results in a *Single cell counts* node.

![](/files/DhaGcsdd4JMA1YcfQz0M)

## Annotate Visium image

The tissue image must be annotated to associate the microscopy image with the expression data.

* Click the newly created *Single cell counts* data node
* Click the **Annotation/Metadata** section in the toolbox
* Click **Annotate Visium image**
* Click on the **Browse** button to open the file browser and point to the file **\_spatial.zip**, created by the *Space Ranger* task
* Click **Finish**

![](/files/iYlitSs4g0Fk6hgLvDUt)

Select the zipped image folder for each sample. The image zip file should contain 6 files including image files and tissue position text file with a scale factor json file. The setup page shows the sample table (one sample per row).

You can find the location of the *\_spatial.zip* file using the following steps. Select the **Space Ranger** task node (i.e. the rectangle) and then click on the **Task Details** (toolbox). Click on the **Output files** link to open the page with the list of files created by the *Space Ranger* task. **Mouse over** any of the files to see the directory in which the file is located. The figure below shows the path to the .zip file which is required for *Annotate Visium image*.

![](/files/FfYTK4o4vmfs1WQcnwqF)

Mousing over a file on the Output files page shows a balloon with the file location.

A new data node, *Annotated counts*, will be generated.

![](/files/61gxZV1xYMRR57W64SIq)

The *Annotated counts* node is **Split by sample**. This means that any tasks performed from this node will also be split by sample. Invoke tasks from the Single cell counts node to combine samples for analyses.

*Annotate Visium image* task creates a new node, *Annotated counts*. Double click on the **Annotated counts** node to invoke the *Data Viewer* showing data points overlaid on top of the microscopy image.

![](/files/4QuD3mMofMtNGYJo2nVH)

Data Viewer session as a result of opening an Annotated counts data node. Each data point is a tissue spot.

Proceed with analysis from the *Single cell counts* node. [Click here to learn about viewing the multiple tissue images in the Data Viewer.](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/view-tssue-images.md)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Spatial data analysis steps

* [Visium data analysis pipeline](#visium-data-analysis-pipeline)
* [Performing tasks in the Analyses tab](#performing-tasks-in-the-analyses-tab)
* [Filter Features](#filter-features)
* [Normalization](#normalization)
* [Exploratory analysis](#exploratory-analysis)
* [Automatic classification](#automatic-classification)
* [Publish cell attributes to project](#publish-cell-attributes-to-project)
* [Modify cell attribute](#modify-cell-attribute)

Here we are [starting with Spacer Ranger outputs as the Single cell counts node](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/start-with-pre-processed-space-ranger-output-files).

## Visium data analysis pipeline

A basic example of a spatial data analysis, starting from the *Single cell counts* node, is shown below and is similar to [a Single cell RNA-Seq analysis pipeline](/partek-flow/tutorials/analyzing-single-cell-rna-seq-data) with the addition of the *Spatial report task* (shown) or *Annotate Visium image* task (not shown).

![](/files/ZEUCqOrUyoUpAaPnbbhZ)

Note that *QA/QC* has not been performed in this example, to visualize all spots (points) on the tissue image. *Single cell QA/QC* can be performed from the *Single cell counts* node with the filtered cells applied to the *Single cell counts* before the Filter features task. [Click here for more information on Single cell QA/QC (see the pipeline in Figure 11)](/partek-flow/user-manual/task-menu/qa-qc/single-cell-qa-qc).

## Performing tasks in the Analyses tab

A context-sensitive menu will appear on the right side of the pipeline. Use the drop-downs in the toolbox to open available tasks for the selected data node.

![](/files/ScTJ1unny0bqZgMPwBOK)

Low-quality cells can be filtered out during the spatial data analysis using *QA/QC* and will not be viewed on the tissue image. [Click here for more information on Single cell QA/QC](/partek-flow/user-manual/task-menu/qa-qc/single-cell-qa-qc). We will not perform *Single cell QA/QC* in this tutorial; this task would be invoked from the *Single cell counts* node and the *Filter features* task discussed below would be invoked from this output node (*Filtered counts*).

## Filter Features

Remove gene expression counts that are not relevant to the analysis.

* Click the **Filtering** drop-down in the toolbox
* Click the **Filter Features** task
* Choose **Noise reduction**
* Exclude features where *value* <= *0.0* in at least *99.0%* of the cells
* Click **Finish**

![](/files/oteJb0FMS22PcQEHsMDt)

Remove gene expression values that are zero in the majority of the cells.

A task node, *Filtered counts*, is produced. Initially, the node will be semi-transparent to indicate that it has been queued, but not completed. A progress bar will appear on the *Filter features* task node to indicate that the task is running.

![](/files/qGb5lVkkmsuDcDhnbVfa)

## Normalization

Normalize (transform) the cells to account for variability between cells.

* Select the **Filtered Counts** result node
* Choose the **Normalization** task from the toolbox

![](/files/a4ouJtjiJJkHuAsKrgO0)

* Click **Use recommended**
* Click **Finish**

![](/files/7LGsl5lT36wEisvX1iMF)

## Exploratory analysis

Explore the data by dimension reduction and clustering methods.

* Click the *Normalized counts* result node
* Select the **PCA** task under *Exploratory analysis* in the toolbox
* Unselect **Split by Sample**
* Click **Finish**

![](/files/tVAMqhAX3GOpkKyq2GTI)

The *PCA* result node generated by the *PCA* task can be visualized by double-clicking the circular node.

* Single click the *PCA* result node
* Select the **Graph-based clustering** task from the toolbox
* Click **Finish**

![](/files/SQDEC5mPE2RW97kewnX0)

The results of graph-based clustering can be viewed by PCA, UMAP, or t-SNE. Follow the steps outlined below to generate a UMAP.

* Select the *Graph-based clustering* result node by single click
* Select the **UMAP** task from the toolbox
* Click **Finish**

![](/files/XBLVRqxCpHZObfX0Y9h5)

* Double-click the *UMAP* result node

![](/files/BF0TgB0YieNKkTEegZpd)

The UMAP is automatically colored by the graph-based clustering result in the previous node. To change the color, click Style.

## Automatic classification

Classify the cells using [Garnett automatic classification](/partek-flow/user-manual/task-menu/classification) to determine cell types.

* Click the *Filtered counts* node
* From the *Classification* drop-down in the toolbox, select **Classify cell type**
* Using the *Managed classifiers*, select the *human Intestine* Garnett classifier
* Click **Finish**

The output of this task produces the *Classify result* node.

Double-click the *Classify result* node to view the cell count for each cell type and the top marker features for each cell type.

![](/files/I7WcBBnYSuf4C2l7UOnm)

## Publish cell attributes to project

Publish cell attributes to the project to make this attribute accessible for downstream applications.

* Click the *Classify result* node
* Select **Publish cell attributes** **to project** under *Annotation/Metadata*
* Select cell\_type from the drop-down and click the green ![image2023-12-11\_21-50-5](/files/fa7wq1RxQ6oV0mx3wf1W) **Add** button
* Name the cell attribute
* Click Finish

![](/files/qw2FlWHPqzCkHglMAzoh)

Publish cell attributes can be applied to result nodes with cell annotation (e.g. click the *graph-based clustering* result node and follow the same steps).

An example of this completed task is shown below.

![](/files/XOdcjkh9N0XWg0RWcszO)

Since this attribute has been published, we can choose to right-click the *Publish cell attributes to project* node and remove this from the pipeline. This attribute will be managed in the Metadata tab (discussed below).

## Modify cell attribute

The name of the Cell attribute can be changed in the *Metadata* tab (right of the *Analyses* tab).

![](/files/tk4lw8JaDNeR2LlOkOHh)

* Click **Manage**
* Click the *Action dots*
* Choose **Modify attribute**
* Rename the attribute *Cell Type*
* Click **Save**
* Click **Back to metadata tab**

![](/files/Rb5yOeDDselcX63Y7aDz)

Drag and drop the categories to rearrange the order of these categories, The order here will determine the plotted order and legend in visualizations.

We can use these Cell attributes in analyses tasks such as Statistics (e.g. differential analysis comparisons) as well as to Style the visualizations in the Data Viewer.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# View tissue images

* [Visualize the annotated image from the automatically generated Spatial report task](#visualize-the-annotated-image-from-the-automatically-generated-spatial-report-task)
* [Visualize the annotated image from the Annotate Visium image task](#visualize-the-annotated-image-from-the-annotate-visium-image-task)
* [Modify Style](#modify-style)
* [Modify Axes](#modify-axes)
* [Color by gene expression and attributes](#color-by-gene-expression-and-attributes)

## Visualize the annotated image from the automatically generated Spatial report task

[With the pre-processed samples imported, we can begin analysis.](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/start-with-pre-processed-space-ranger-output-files)

* Click **Analyses** to switch to the *Analyses* tab

For now, the *Analyses* tab has a starting node, a circular node called *Single cell* counts and also a rectangular task node called *Spatial report* which was automatically generated for this type of data. As you [perform analyses, additional nodes representing tasks and new data will be created](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/spatial-data-analysis-steps), forming a visual representation of your analysis pipeline.

* Click the **Spatial report** node
* Click **Task report** on the task menu

![](/files/wLXMMzAT7MedhvZmSt13)

The spatial report will display the first sample (Replicate 1). We want to visualize all of the samples using the steps below.

* Duplicate the plot by clicking the **Duplicate plot** button in the upper right controls (arrow 1)
* Open the **Axes** configuration option (arrow 2)
* Change the Sample on the duplicated image under *Misc* (arrow 3)

![](/files/Tz7gfm020wsjbblNxy42)

Each data point is a tissue spot. Duplicate and change the sample to view multiple samples.

## Visualize the annotated image from the Annotate Visium image task

If starting with unprocessed [fastq files](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/start-with-10x-genomics-visium-fastq-files/README.md), the *Annotate Visium image* task will create a new result node, *Annotated counts*.

* Double click on the **Annotated counts** node to invoke the *Data Viewer* showing data points overlaid on top of the microscopy image
* Follow the steps outlined above by duplicating the image to visualize the multiple samples

## Modify Style

To modify the points on the image to show more of the background image use the Style configuration option.

* Press and hold Ctrl or Shift to select both plots
* Click **Style** in the left panel
* Move the *Opacity* slider to the left
* Change the *Point size* to 3

![](/files/xG233t7qoOB5aQedkg2Y)

* Click **Save** in the left panel and give the session an appropriate name

![](/files/2SOLWnaKsVabhA7v6KSq)

## Modify Axes

Modify the axes to remove the X and Y coordinates from the tissue image.

* Press and hold Ctrl or Shift to select both plots
* Click **Axes** in the left panel
* Toggle off *Show lines* for both the X & Y axis
* Toggle off *Show title* and *Show axis* for both the X & Y axis

![](/files/JXxRtnm1ZmR7GLyE5b8v)

## Color by gene expression and attributes

Style the image and color by normalized gene expression using three genes of interest.

* Press and hold Ctrl or Shift to select both plots
* Click **Style**
* Click the blue circle node ![image2023-12-12\_13-21-34](/files/2DpHsNhOrREG01GakzfU) to the right of the *Color by* drop-down
* Select the *Normalized counts* node as the source
* Choose to **Color by** *Numeric triad*
* Use the Green drop-down to select IL32, Red drop-down to select DES, and Blue drop-down to select PTGDS genes (type in name of gene in drop-down)
* Increase the *Point size* to 11

![](/files/wBRbbTvvFniasFd5wSNr)

To color by the [Cell attribute "Cell Type"](/partek-flow/tutorials/10x-genomics-visium-spatial-data-analysis/spatial-data-analysis-steps#publish-cell-attributes-to-project) which we previously determined in this tutorial, use the **Color by** drop-down and select Cell Type. Cell type is a blue categorical attribute while green attributes are numerical.

![](/files/kuaLgswvMH9eYAze43Yj)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# 10x Genomics Xenium Data Analysis

In this tutorial, we demonstrate how to:

* [Import 10x Genomics Xenium Analyzer output](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/import-10x-genomics-xenium-analyzer-output)
* [Process Xenium data](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/process-xenium-data)
* [Perform Exploratory analysis](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/perform-exploratory-analysis)
* [Make comparisons using Compute biomarkers and Biological interpretation](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/tutorials/10x-genomics-xenium-data-analysis/make-comparisons-using-compute-biomarkers-and-biological-interpretation/README.md)

## Tutorial Data Set

The tutorial data is based on [10x Genomics Datasets](https://www.10xgenomics.com/resources/datasets?query=\&page=1\&configure%5BhitsPerPage%5D=50\&configure%5BmaxValuesPerFacet%5D=1000).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Import 10x Genomics Xenium Analyzer output

* [Obtain and add files to the project](#obtain-and-add-files-to-the-project)

## Obtain and add files to the project

The project includes [Human Breast Cancer (In Situ Replicate 1)](https://www.10xgenomics.com/products/xenium-in-situ/preview-dataset-human-breast) and [Human Breast Cancer (In Situ Replicate 2)](https://www.10xgenomics.com/products/xenium-in-situ/preview-dataset-human-breast) files in one project.

* Obtain the **Xenium Output Bundles** (Figure 1) for each sample.

![Figure 1. Obtain the Xenium Output Bundle on your machine](/files/Uvxqb1OHskrMwjXGPdZ8)

* Navigate the options to select **10x Genomics Xenium Output Bundle** as the file format for input. Choose to import **10x Genomics Xenium** for your project (Figure 2).

![Figure 2. Transfer files using the 10x Genomics Xenium importer](/files/3ASbdio8F4cIqm61pgFT)

* Click [**Transfer files**](/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server) on the homepage, under settings, or during import.
* Click the blue **+ Add sample button** ![](/files/m4rooUgafEIGH8R98KlB)then use the green **Add sample** ![image2023-8-21\_12-7-17](/files/LfxK7ctIWZqXlmfixfLD) button to add each sample's Xenium output bundle folder. If you have not already transferred the folder to the server, this can be done using **Transfer files to the server**![image2023-8-21\_12-18-22](/files/XzofuvXdtMtN4n4Dk88U) (Figure 3).

<figure><img src="/files/Vra90OGwAL294b81sNuE" alt=""><figcaption><p>Figure 3. Transfer files using the 10x Genomics Xenium importer</p></figcaption></figure>

* You will need to **decompress** the Xenium Output Bundle zip file before they are uploaded to the server. After decompression, you can **drag and drop** the entire folder into the Transfer files dialog, all individual files in the folder will be listed in the Transfer files dialog after drag & drop, with no folder structure (Figure 4). The folder structure will be restored after upload is completed.

![Figure 4. Drag & drop unzipped Xenium Output Bundle folder into Transfer files dialog](/files/mCkEyNNZ3PfAmoo1gyvq)

* Once uploaded the folder to the server, navigate to the appropriate folder for each sample using **Add sample** ![image2023-8-21\_12-8-24](/files/LfxK7ctIWZqXlmfixfLD) (Figure 5).

The Xenium output bundle should be included for each sample (Figure 5). Each sample requires the whole sample folder or a folder containing these 6 files: cell\_feature\_matrix.h5, cells.csv.gz, cell\_boundaries.csv.gz, nucleus\_boundaries.csv.gz, transcripts.csv.gz, morphology\_focus.ome.tif. Once added, the *Cells* and *Features* values will update. You can choose an annotation file during import that matches what was used to generate the feature count.

Do not limit cells with a total read count since Xenium data is targeted to less features.

![Figure 5. Add Xenium output bundle](/files/R06UNTODRGT2rdZ00GTC)

Once the download completes, the sample table will appear in the *Metadata* tab, with one row per sample (Figure 6).

![Figure 6. Each sample should be present in the Metadata tab](/files/Iv7l27od2ETN20CfI2fd)

The sample table is pre-populated with one sample attributes: # Cells. Sample attributes can be added and edited manually by clicking *Manage* in the *Sample attributes* menu on the left. If a new attribute is added, click *Assign values* to assign samples to different groups. Alternatively, you can use the *Assign values from a file* option to assign sample attributes using a tab-delimited text file. For more information about sample attributes, see [here](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab#sample-annotation). Cell attributes are found under Sample attributes and can be added by [publishing cell attributes to a project](/partek-flow/user-manual/task-menu/annotation-metadata/publish-cell-attributes-to-project).

For this tutorial, we do not need to edit or change sample attributes.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Process Xenium data

* [Obtain and add files to the project](#obtain-and-add-files-to-the-project)
* [Perform quality analysis / quality control and filtering](#perform-quality-analysis--quality-control-and-filtering)
* [Normalize the data](#normalize-the-data)

## Obtain and add files to the project

Follow along to add files to your Xenium project: [Add files to the project](/partek-flow/tutorials/10x-genomics-xenium-data-analysis/import-10x-genomics-xenium-analyzer-output).

## Perform quality analysis / quality control and filtering

* Filter the data including control probes using the **Filter features** task
* Choose **Feature metadata filter**
* Include the **Gene Expression** features
* Click **Finish**

![](/files/GcMKgw6zoF5uoWxGy1ma)

* This results in a *Filter features* task (rectangle) and node (circle) results.
* Right click the circle to **Rename data node** to "*Filtered to only gene expression*"

![](/files/S5Ro4Eaw333uSYdcnHBz)

* Click the circular "*Filtered to only gene expression*" results node and select the **Single cell QA/Q**C task from the context sensitive menu on the right

![](/files/XFEUcDjHuFxalOnglTD9)

* When the task completes it will be opaque and no longer transparent with a progress bar

![](/files/Ay2SMmlh0uQIf8fnl6JH)

Double click the opaque rectangle task to open and [filter cells as described here.](/partek-flow/user-manual/task-menu/qa-qc/single-cell-qa-qc) Apply the observation filter to the "*Filtered to only gene expression*" results node. This results in a "*Filtered cells node*".

## Normalize the data

* Select the "*Filtered cells node*" and choose the **Normalization** task from the *Normalization and scaling* drop-down in the task menu

![](/files/Y6Q1qxHqSI40TvECbS3f)

* Click the **Use recommended** button to proceed with these settings
* Click **Finish**

![](/files/Czp3atA7exvSX7haL92g)

This results in a Normalized counts node as shown below in the pipeline.

![](/files/M6itGeuWXcb3dIW4NJxA)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Perform Exploratory analysis

* [Use Principle Components Analysis (PCA) to reduce dimensions](#use-principle-components-analysis-pca-to-reduce-dimensions)
* [Classify cells based on a marker for expression](#classify-cells-based-on-a-marker-for-expression)

## Use Principle Components Analysis (PCA) to reduce dimensions

* Click the **Normalized counts** data node
* Expand the **Exploratory analysis** section of the task menu
* Click **PCA**

![](/files/IjqeJck3eXJWDUwSGK06)

In this tutorial we will modify the PCA task parameters, to not split by sample, to keep the cells from both samples on the PCA output.

* Uncheck (de-select) the **Split by sample** checkbox under *Grouping*
* Click **Finish**

![](/files/8oDgsyebQ8i7eWdHGSAV)

* Double-click the circular **PCA** node to view the results

![](/files/eU3gio4noRZDKfoAQR0W)

From this PCA node, further exploratory tasks can be performed (e.g. t-SNE, UMAP, and Graph-based clustering).

## Classify cells based on a marker for expression

* Choose **Style** under *Configure*
* **Color by** and search for *fasn* by typing the name
* Select *FASN* from the drop-down

![](/files/kibtFfKnZlW8mAz9Cpuw)

The colors can be customized by selecting the **color palette** then using the color drop-downs as shown below.

![](/files/LknfQr0ull03aOFBIzxf)

Ensure the colors are distinguishable such as in the image above using a blue and green scale for **Maximum** and **Minimum**, respectively.

* Click **FASN** in the legend to make it draggable (pale green background) and continue to drag and drop FASN to **Add criteria** within the **Select & Filter** *Tool*
* Hover over the slider to see the distribution of FASN expression

![](/files/SercMqtv5dYjKeCerE0O)

Multiple gene thresholds can be used in this type of classification by performing this step with multiple markers.

* Drag the slider to select the population of cells expressing high FASN (the cutoff here is 10 or the middle of the distribution).

![](/files/bsy10gWZrqlI7aS5E0we)

* Click **Classify** under *Tools*
* Click **Classify selection**

![](/files/ndqmnPia9e7kibSWXVxa)

* Give the classification a name "FASN high"

![](/files/DEMEdYdJ77Ld71Touwcg)

* Under the Select & Filter tool, choose **Filter** to **exclude** the selected cells

![](/files/SAU4w9ZBeN1tQzB3A2WG)

Exit all Tools and Configure options

* Click the "X" in the right corner
* Use the **rectangle** selection mode on the PCA to select all of the points on the image

![](/files/AwgRldYcytyQzF0m35bP)

This results in 147538 cells selected.

![](/files/fDyk5Xd6SZYl6LMM2d9X)

* Open **Classify**
* Click **Classify selection** and name this population of cells "FASN low"
* Click **Apply classifications** and give the classification a name "FASN expression"

![](/files/Gv4qjSI0erlpXx8vzw44)

Now we will be able to use this classification in downstream applications (e.g. differential analysis).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Make comparisons using Compute biomarkers and Biological interpretation

* [Compute biomarkers](#compute-biomarkers)
* [Create a list](#create-a-list)
* [Biological Interpretation](#biological-interpretation)

We will compare the classification (FASN expression) we previously made based on expression levels of the FASN gene. Here, we will compare FASN high and FASN low cells to identify genes and pathways.

## Compute biomarkers

* Select the Normalized counts node and choose **Compute biomarkers** from the *Statistics* drop-down

![](/files/K3CPD45mQ440fkDDTRio)

* Choose the "FASN expression" attribute
* Do not select **Split by sample**
* Click Finish

![](/files/wTVbgMNTle0QwCpBiGQ8)

This results in a Biomarkers report.

* Double-click the **Biomarkers** results node to open the report

![](/files/IcnNg4scaR9SAEQYpEaa)

The top features are reported for the comparison.

* **Download** this table with more than 10 features using the Download option ![image2023-7-26\_9-38-10](/files/0chncmNC8AvpPIiCtIFG)

![](/files/TT5Ys5PBKA6ziZleqzyh)

[Please click here for more information on differential analysis methods.](/partek-flow/user-manual/task-menu/differential-analysis)

## Create a list

We will create a list using [List management](/partek-flow/user-manual/settings/components/lists) with these 10 genes, so that we can use this list in the Gene set enrichment task.

* Click your **username** in the top right corner
* Select **Settings** from the drop-down
* Choose **Lists** from the **Components** drop-down in the menu on the left

![](/files/AInNa0utgaTs9gCCBjYC)

* Use the **+ New list** button to add these 10 genes

![](/files/KqwS8Egs7xFrPllm3ypr)

* Choose **Text** as the list option
* Give the list a **Name** and **Description**
* Enter the 10 genes in column format as shown below
* Click **Add list**

![](/files/FlUlCXpNT78KWmIr30RK)

The list has been added and can now be used for further analysis. The **Actions** button can be used to modify this list if necessary, as shown below.

![](/files/bKa120pEND67WYQgiGir)

## Biological Interpretation

Here we are going to perform [Gene Set Enrichment](/partek-flow/user-manual/task-menu/biological-interpretation/gene-set-enrichment) on our top 10 features for the FASN high group that we have added as a list called "Top 10 FASN high Features".

* Go to the **Analyses tab**
* Select the Normalized counts node
* Choose **Gene set enrichment** from the *Biological interpretation* drop-down in the task menu

![](/files/HvxFYuygzqC3lCAOeMXJ)

* Use the KEGG database for pathway enrichment
* Check **Specify background gene list**
* Select "Top 10 FASN high Features" as the **Background gene list**
* Click **Finish**

![](/files/B9nUgpNy8dHiNGG3qfbJ)

This results in a Pathway enrichment report, as shown below.

* Double-click the report to view the pathways involved in this list of genes

![](/files/dE92e8hVXRaF2xFK14cq)

Please click [here](/partek-flow/user-manual/task-menu/biological-interpretation) for more information on Biological interpretation.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Single Cell RNA-Seq Analysis (Multiple Samples)

In this tutorial, we demonstrate how to:

* [Getting started with the tutorial data set](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples/getting-started-with-the-tutorial-data-set)
* [Classify cells from multiple samples using t-SNE](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples/classify-cells-from-multiple-samples-using-t-sne)
* [Compare expression between cell types with multiple samples](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples/compare-expression-between-cell-types-with-multiple-samples)

## Tutorial Data Set

The tutorial is based on the work published by [Venteicher and co-workers](http://science.sciencemag.org/content/355/6332/eaai8478), on isocitrate dehydrogenase-mutant gliomas. Single cells from tumor biopsies were processed by flow cytometry and the libraries were prepared by [Smart-seq2](https://www.nature.com/articles/nmeth.2639) protocol. The tutorial data set consists of eight expression matrix files, one per patient sample. The tumors were categorized as either astrocytoma or oligodendroglioma glioma subtype by histology. The matrix files contain gene expression values normalized by the following transformation log2\[(TPM/10)+1].

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Getting started with the tutorial data set

* [Creating a new project and importing the tutorial data set](#creating-a-new-project-and-importing-the-tutorial-data-set)
* [Filtering cells in single cell RNA-Seq data](#filtering-cells-in-single-cell-rna-seq-data)
* [Filtering genes in single cell RNA-Seq data](#filtering-genes-in-single-cell-rna-seq-data)
* [Normalizing single cell RNA-Seq data](#normalizing-single-cell-rna-seq-data)

## Creating a new project and importing the tutorial data set

The tutorial data set is available through Partek Flow.

* Click your **avatar** (Figure 1)

![Figure 1. Location of the Settings link on the main page of Partek Flow](/files/eo8y02NDeCpKhUNNXg0e)

* Click **Settings**

On the *System information* page, the *Download tutorial data* section includes pre-loaded data sets used by Partek Flow tutorials (Figure 2).

![Figure 2. Tutorial data sets available through Partek Flow](/files/fWhS1DcKijBe7CWmWJaI)

* Click **Single cell glioma (multi-sample)**

The tutorial data set will be downloaded onto your Partek Flow server and a new project, *Glioma (multi-sample),* will be created. You will be directed to the *Data* tab of the new project. Because this is a tutorial project, there is no need to click on *Import data*, as the import is handled automatically (Figure 3).

![Figure 3. The data tab during tutorial data import](/files/T2m838ks1hAnVYrFsUC7)

You can wait a few minutes for the download to complete, or check the download progress by selecting **Queue** then **View queued tasks...** to view the *Queue* (Figure 4).

![Figure 4. Viewing the queue](/files/YtWOUKGkQ1ElVo1HJRtq)

Once the download completes, the sample table will appear in the *Data* tab, with one row per sample (Figure 5).

![Figure 5. Sample data table listing the name and the number of cells for each sample](/files/ZiZyRvUmH4nhr6cWboWx)

The sample table is pre-populated with two sample attributes: # Cells and Subtype. Sample attributes can be added and edited manually by clicking *Manage* in the *Sample attributes* menu on the left. If a new attribute is added, click *Assign values* to assign samples to different groups. Alternatively, you can use the *Assign values from a file* option to assign sample attributes using a tab-delimited text file. For more information about sample attributes, see [here](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab#sample-annotation).

For this tutorial, we do not need to edit or change any sample attributes.

## Filtering cells in single cell RNA-Seq data

With samples imported and annotated, we can begin analysis.

* Click **Analyses** to switch to the *Analyses* tab

For now, the *Analyses* tab has only a single node, *Single cell counts.* As you perform the analysis, additional nodes representing tasks and new data will be created, forming a visual representation of your analysis pipeline.

* Click on the **Single cell counts** node

A context-sensitive menu will appear on the right-hand side of the pipeline (Figure 9). This menu includes tasks that can be performed on the selected counts data node.

An important step in analyzing single cell RNA-Seq data is to filter out low-quality cells. A few examples of low-quality cells are doublets, cells damaged during cell isolation, or cells with too few counts to be analyzed.

* Expand the **QA/QC** section of the task menu
* Click on **Single cell QA/QC** (Figure 6)

![Figure 6. Selecting the Single cell QA/QC task from the task menu](/files/k8uAiwMqKBq4irULXqy0)

A task node, *Single cell QA/QC*, is produced. Initially, the node will be semi-transparent to indicate that it has been queued, but not completed. A progress bar will appear on the *Single cell QA/QC* task node to indicate that the task is running.

* Click the **Single cell QA/QC** node once it finishes running
* Click **Task report** on the task menu (Figure 7)

![Figure 7. Selecting the task report for any task node opens a report with any tables or charts the task produced](/files/PbbCU89dqEB3MCQpgP57)

The *Single cell QA/QC* report opens in a new data viewer session. There are interactive violin plots showing the most commonly used quality metrics for each cell from all samples combined (Figure 8). For this data set, there are two relevant plots: the total count per cell and the number of detected genes per cell. Each point on the plots is a cell and the violins illustrate the distribution of values for the y-axis metric. Typically, there is a third plot showing the percentage of mitochondrial counts per cell, but mitochondrial transcripts were not included in the data set by the study authors, so this plot is not informative for this data set.

* Remove the % mitochondrial counts and the extra text box in the bottom right by clicking **Remove plot** in the top right corner of each plot (Figure 8).

![Figure 8. Each cell is shown as a point on the plot. Remove the % mitochondrial counts and empty text box using the X icons](/files/MbLGGkTjxbATDbfiR6Ub)

The plots are highly customizable and can be used to explore the quality of cells in different samples.

* Click on **Single cell counts** in the **Get Data** icon on the left (Figure 9)
* Click and drag the **Sample name** attribute onto the *Counts plot* and drop it onto the *X-axis*
* Repeat this for the *Detected genes* plot

![Figure 9. Click and drag the Sample name attribute onto the X-axis for each plot](/files/F22wO8hUDP1fVaoEnmlE)

The cells are now separated into different samples along the x-axis (Figure 10)

* Hold Control and left-click to select both plots
* Open the **Style** icon on the left under *Configure*
* Under *Color*, use the slider to reduce the **Opacity**
* Open the **Axis** icon on the left
* Adjust the **X-rotation** on the plots to **90**

![Figure 10. Counts and detected genes plots can be customized to compare cells from different samples](/files/KzxAZUyk5P64yeq22i4S)

Note how both plots were modified at the same time.

Cells can be selected by setting thresholds using the **Select & Filter** tool. Here, we will select cells based on the total count

* Open **Select & Filter** under *Tools* on the left
* Under *Criteria*, Click **Pin histogram** to see the distribution of counts
* Set the *Counts* thresholds to **8000 and 20500**

Selected cells will be in blue and deselected cells will be dimmed (Figure 11).

![Figure 11. Previewing a filter using the Single cell QA/QC violin plots](/files/4mW0tJRhykCFdnY02Kez)

Because this data set was already filtered by the study authors to include only high-quality cells, this count filter is sufficient.

* Click ![Filter\_include\_icon](/files/zdAVcV8BSPFAx45UsZIH) under *Filter* to include the selected cells
* Click **Apply observation filter**
* Click the **Single cell counts** data node in the pipeline preview (Figure 12)
* Click **Select**

![Figure 12. After the Apply filter button is selected, you will be presented with a preview of your pipeline. You need to select the appropriate data node to apply the filtering to.](/files/Tt1Fwbsvl7PGnEJ7Trl2)

A new task, *Filter counts*, is added to the *Analyses* tab. This task produces a new *Filter counts* data node (Figure 13).

* Click on the **Glioma (multi-sample)** project name at the top to go back to the *Analyses* tab
* Your browser may warn you that any unsaved changes to the data viewer session will be lost. Ignore this message and proceed to the *Analyses* tab

![Figure 13. Applying a cell quality filter](/files/2hlZL6VWBFQjvx0uLboy)

Most tasks can be queued up on data nodes that have not yet been generated, so you can wait for filtering step to complete, or proceed to the next section.

## Filtering genes in single cell RNA-Seq data

A common task in bulk and single-cell RNA-Seq analysis is to filter the data to include only informative genes. Because there is no gold standard for what makes a gene informative or not, ideal gene filtering criteria depends on your experimental design and research question. Thus, Partek Flow has a wide variety of flexible filtering options.

* Click the **Filter counts** node produced by the *Filter counts* task
* Click **Filtering** in the task menu
* Click **Filter features** (Figure 14)

![Figure 14. Invoking Filter features](/files/UVHbmYCmTnDWuv7Tx01h)

There are four categories of filter available - noise reduction, statistics based, feature metadata, and feature list (Figure 15).

![Figure 15. Viewing the filtering options](/files/8PUO980EKt6Uv6DsXfgA)

The noise reduction filter allows you to exclude genes considered background noise based on a variety of criteria. The statistics based filter is useful for focusing on a certain number or percentile of genes based on a variety of metrics, such as variance. The feature list filter allows you to filter your data set to include or exclude particular genes.

We will use a noise reduction filter to exclude genes that are not expressed by any cell in the data set but were included in the matrix file.

* Click the **Noise reduction filter** checkbox
* Set the *Noise reduction filter to* **Exclude features where value <= 0 in 99% of cells** using the drop-down menus and text boxes (Figure 16)
* Click **Finish** to apply the filter

![Figure 16. Configuring a noise reduction filter to exclude genes not expressed in the data set](/files/cXu51q8YVajljdQzC05d)

This produces a *Filtered counts* data node. This will be the starting point for the next stage of analysis - identifying cell types in the data using the interactive t-SNE plot.

## Normalizing single cell RNA-Seq data

We are omitting normalization in this tutorial because the data has already been normalized.

The tutorial data set is taken from a published study and has already been normalized using TPM (Transcripts per million), which normalizes for the length of feature and total reads, and transformed as log2(TPM/10+1). This normalization and transformation scheme can be performed in Partek Flow, along with other commonly used RNA-Seq data normalization methods.

For more information on normalizing data in Partek Flow, please see the [Normalization](/partek-flow/user-manual/task-menu/normalization-and-scaling/normalization) section of the user manual.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Classify cells from multiple samples using t-SNE

* [Multiple single-sample t-SNE plots](#multiple-single-sample-t-sne-plots)
* [One multi-sample t-SNE plot](#one-multi-sample-t-sne-plot)

t-SNE (t-distributed stochastic neighbor embedding) is a visualization method commonly used to analyze single-cell RNA-Seq data. Each cell is shown as a point on the plot and each cell is positioned so that it is close to cells with similar overall gene expression. When working with multiple samples, a t-SNE plot can be drawn for each sample or all samples can be combined into a single plot. Viewing samples individually is the default in Partek Flow because sample to sample variation and outlier samples can obscure cell type differences if all samples are plotted together. However, as you will see in this tutorial, in some data sets, cell type differences can be visualized even when samples are combined.

Using the t-SNE plot, cells can be classified based on clustering results and differences in expression of key marker genes.

## Multiple single-sample t-SNE plots

Prior to performing t-SNE, it is a good idea to reduce the dimensionality of the data using principal components analysis (PCA).

* Click the **Filtered counts** data node after the *Filter features* task
* Select **PCA** from the *Exploratory analysis* section of the task menu (Figure 1)

![Figure 1. Select the PCA task from the Exploratory analysis menu](/files/yRK1jfvST3utgBE0v9qQ)

* Click **Finish** to run PCA with default settings (Figure 2)

Note, the default settings include the *Split by sample* checkbox being selected. This means that the dimensionality reduction will be performed on each sample separately.

![Figure 2. PCA task set up page with default settings](/files/ohoTpOR7HVdCYMyTUkaO)

PCA task and data nodes will be generated.

* Click the **PCA** data node
* Select **t-SNE** from the *Exploratory analysis* section of the task menu (Figure 3)

![Figure 3. Invoking t-SNE from the task menu](/files/b7r1yqTlrM6YOEOjnp2a)

* Click **Finish** from the *t-SNE* dialog to run t-SNE with the default settings (Figure 4)

![Figure 4. t-SNE task set up with default settings](/files/jtQPu0Rza2Wl318gFKTc)

Because the upstream PCA task was performed separately for each sample, the t-SNE task will also be performed separately for each sample. t-SNE task and data nodes will be generated (Figure 5).

![Figure 5. t-SNE task node](/files/WDGZi7yqrcBl0cvQibvs)

Once the t-SNE task has completed, we can view the t-SNE plots

* Click the **t-SNE** node
* Click **Task report** from the task menu or double click the *t-SNE* node

The t-SNE will open in a new data viewer session. The t-SNE plot for the first sample in the data set, MGH36 (Figure 6), will open on the canvas. Please note that the appearance of the t-SNE plot may differ each time it is drawn so your t-SNE plots may look different than those shown in this tutorial. However, the cell-to-cell relationships indicated will be the same.

![Figure 6. Viewing t-SNE plot of a single sample](/files/r1QLHU1WzLLgxc1DmjPp)

The t-SNE plot is in 3D by default. To change the default, click your avatar in the top right > *Settings* > *My Preferences* and edit your graphics preferences and change the default scatter plot format from 3D to 2D.

You can rotate the 3D plot by left-clicking and dragging your mouse. You can zoom in and out using your mouse wheel. The 2D t-SNE is also calculated and you can switch between the 2D and 3D plots on the canvas. We will do this later on in the tutorial.

Each sample has its own plot. We can switch between samples.

* Open the **Axes** icon on the left under *Configure* (Figure 7)
* Navigate to *Misc*
* Select the ![arrow\_right\_icon\_expand\_triangle\_gray](/files/zcT1L2ZANNl5q2BM8wE5) icon below the *Sample* name to go to the next sample

The t-SNE plot has switched to show the next sample, MGH42 (Figure 7).

![Figure 7. Viewing t-SNE plot of MGH42](/files/lRwDYRYYmFPLxSxBybqD)

The goal of this analysis is to compare malignant cells from two different glioma subtypes, astrocytoma and oligodendroglioma. To do this, we need to identify the malignant cells we want to include and which cells are the normal cells we want to exclude.

The t-SNE plot in Partek Flow offers several options for identifying, selecting, and classifying cells. In this tutorial, we will use the expression of known marker genes to identify cell types.

To visualize the expression of a marker gene, we can color cells on the t-SNE plot by their expression level.

* Select any of the count data nodes from **Get data** on the left (Single cell counts, or any of the Filtered counts, Figure 8)
* Search for the **BCAN** gene
* Click and drag the BCAN gene onto the plot and drop it over the **Green (feature)** option

![Figure 8. Coloring cells by BCAN expression](/files/S9bZNTeB8cKdYYNSPVYH)

The cells will be colored from black to green based on their expression level of BCAN, with cells expressing higher levels more green (Figure 9). BCAN is highly expressed in glioma cells.

![Figure 9. Cells colored by BCAN expression](/files/tZ1QAD5FriIxSmj2AmKO)

In Partek Flow, we can color cells by more than one gene. We will now add a second glioma marker gene, GPM6A.

* Select any of the count data nodes from the *Data* card on the left (Single cell counts, or any of the Filtered counts)
* Search for the **GPM6A** gene
* Click and drag the GPM6A gene onto the plot and drop it over the **Red (feature)** option

Cells expressing GPM6A are now colored red and cells expressing BCAN are colored green. Cells expressing both genes are colored yellow, while cells expressing neither are colored black (Figure 10).

![Figure 10. Coloring cells by BCAN and GPM6A](/files/sQ9IZj1DLqUs6yVYA39V)

Numerical expression levels for each gene can be viewed for individual cells.

* Switch to pointer mode by clicking ![image2017-12-28 13\_12\_15](/files/sSBYPOuizlIJf7OXvD1R) in the top right corner of the plot
* Select a cell by pointing and clicking

The expression level for that cell is displayed on the legend for each gene. Expression values can also be viewed by mousing over a cell (Figure 11).

* Deselect the cell by clicking on any blank space on the plot

![Figure 11. Viewing expression levels for an individual cell. The dots on the legend indicate the expression level of the selected cell. The expression levels also appear in the label when you mouse over a cell](/files/fivKpHSwHWNz388u2zrc)

Now that cells are colored by the expression of two glioma cell markers, we can classify any cell that expresses these genes as glioma cells. Because t-SNE groups cells that are similar across the high-dimensional gene expression data, we will consider cells that form a group where the majority of cells express BCAN and/or GPM6A as the same cell type, even if they do not express either marker gene.

* Switch to lasso mode by clicking ![image2017-12-28 13\_3\_50](/files/5OX0B6UoQW7mTYErtHSi) in the top right of the plot
* Draw the lasso around the cluster of green, red, and yellow cells and click the circle to close the lasso (Figure 12)

![Figure 12. Selecting a group of cells using the 3D lasso tool](/files/ZzOeOxbD4GSLN0kb5Xzg)

Selected cells are shown in bold and unselected cells are dimmed. The number of selected cells is indicated in the figure legend. The cells are plotted on the color scale depending on their relative expression levels of the two marker genes (Figure 13)

![Figure 13. Viewing expression levels for a group of cells](/files/qZVjqTavliyxLDDocKUH)

* Click **Classify selection** in the **Classify** icon under *Tools*

A dialog to give the classification a name will appear.

* Name the classification **Glioma**
* Click **Save** (Figure 14)

![Figure 14. Classifying selection](/files/uUTYBW7XoeZ8fQgZaRP9)

Once cells have been classified, the classification is added to **Classify**. The number of cells belonging to the classification is listed. In MGH42, there are 460 glioma cells (Figure 15).

![Figure 15. The number of cells in each classification is displayed](/files/r0bvXANp5Nid0LS847AG)

Classifications made on the t-SNE plot are retained as a draft as part of the data viewer session. In this tutorial, we will classify malignant cells for each sample before we save and apply the classifications, but if necessary, you can save the data viewer session by clicking the ![image2022-8-30\_11-32-53](/files/A6CRd9rlqbWWWNef2jng) **Save** icon on the left to retain all of the formatting and draft classifications. The data viewer session will be stored under the *Data viewer* tab and can be re-opened to continue making classifications at a later time.

* Switch to pointer mode by clicking ![image2017-12-28 13\_12\_15](/files/sSBYPOuizlIJf7OXvD1R) in the top right corner of the plot
* Deselect the cells by clicking on any blank space on the plot
* Open **Axes** and navigate to *Sample* under *Misc*
* Select the ![arrow\_right\_icon\_expand\_triangle\_gray](/files/zcT1L2ZANNl5q2BM8wE5) icon below the sample name to go to the next sample, MGH45
* Rotate the 3D t-SNE plot to get a better view of cells from the green, red, and yellow cluster
* Switch to lasso mode by selecting ![image2017-12-28 13\_3\_50](/files/5OX0B6UoQW7mTYErtHSi) in the top right corner of the plot
* Draw the lasso around the cluster of colored cells and click the circle to close the lasso (Figure 16)

![Figure 16. Classifying malignant cells in sample MGH45](/files/I3Ao114TAXMyWNpPKe7H)

* Select **Classify selection** in the **Classify** icon
* Type **Glioma** or select **Glioma** from the drop-down list (Figure 17)
* Click **Save**

![Figure 17. Adding cells in a second sample to an existing classification](/files/mVq9qpPJGotX8aEmMpCj)

* Repeat these steps for each of the 6 remaining samples. Remember to go back to the first sample (MGH36) to classify the glioma cells in that samples too.

There should be 5,322 glioma cells in total across all 8 samples.

* The classification name can be edited or deleted (Figure 18).

![Figure 18. Edit or delete the classification](/files/6YNn3fBRKnXZFhy7Ea1R)

With the malignant cells in every sample classified, it is time to save the classifications.

* Click **Apply classifications** in the **Classify** icon
* Name the classification attribute **Cell type (sample level)**
* Click **Run** **(Figure 19)**

![Figure 19. Name the cell-level attribute](/files/lziu3pnuErShMjuES9S5)

The new attribute is stored in the *Data* tab and is available to any node in the project.

* Click on the **Glioma (multi-sample)** project name at the top to go back to the *Analyses* tab
* Your browser may warn you that any unsaved changes to the data viewer session will be lost. Ignore this message and proceed to the *Analyses* tab

## One multi-sample t-SNE plot

For some data sets, cell types can be distinguished when all samples can be visualized together on one t-SNE plot. We will use a t-SNE plot of all samples to classify glioma, microglia, and oligodendrocyte cell types.

* Click on the **Glioma (multi-sample)** project name at the top to go back to the *Analyses* tab
* Click the **Filtered counts** data node after the *Filter features* task
* Click **PCA** in the *Exploratory analysis* section of the task menu
* Uncheck the **Split by sample** checkbox (Figure 22)
* Click **Finish**

![Figure 20. Combine all cells into one plot by unchecking the Split by sample box](/files/q7qnHhIcpcNbW7Wv0RS1)

The PCA task will run as a new green layer.

* Click the new **PCA** data node
* Select **t-SNE** from the *Exploratory analysis* section of the task menu
* Click **Finish** to run the t-SNE task with default settings

The t-SNE task will be added to the green layer (Figure 23). Layers are created in Partek Flow when the same task is run on the same data node.

![Figure 21. Multi-sample PCA and t-SNE tasks are added as a new layer](/files/S3oI5KwyON7NodA7zV5G)

Once the task has completed, we can view the plot.

* Double-click the **green** **t-SNE** data node to open the t-SNE scatter plot
* Click and drag the **2D scatter plot** icon onto the canvas and **replace** the 3D scatter plot (Figure 24)

![Figure 22. Viewing the multi-sample t-SNE plot](/files/PZ13CEYUbeMSNcxN1npP)

* Search for and select **green t-SNE** data node (Figure 25)

![Figure 23. Select the green multi-sample t-SNE data node to draw the 2D t-SNE plot](/files/IEvdjl47hXJ6Z9RJKXc8)

* In the **Style** icon, choose **Sample** **name** from the *Color by* drop-down list under *Color*

Viewing the 2D t-SNE plot, while most cells cluster by sample, there are a few clusters with cells from multiple samples (Figure 26).

![Figure 24. Viewing the multi-sample t-SNE plot in 2D](/files/NlrpaVlBdP7PYFexdqli)

Using marker genes, BCAN (glioma), CD14 (microglia), and MAG (oligodendrocytes), we can assess whether these multi-sample clusters belong to our known cell types.

* Select any of the count data nodes from the *Data* card on the left (Single cell counts, or any of the Filtered counts)
* Search for the **BCAN** gene
* Click and drag the BCAN gene onto the plot and drop it over the **Green (feature)** option
* Search for the **CD14** gene
* Click and drag the **CD14** gene onto the plot and drop it over the **Red (feature)** option
* Search for the **MAG** gene
* Click and drag the **MAG** gene onto the plot and drop it over the **Blue (feature)** option

After coloring by these marker genes, three cell populations are clearly visible (Figure 27).

![Figure 25. Overlaying marker gene expression on the multi-sample t-SNE plot](/files/IxuwBRMGygHyBoyNF41F)

The red cells are CD14 positive, indicating that they are the microglia from every sample.

* Switch to lasso mode by clicking the ![image2017-12-28 13\_3\_50](/files/5OX0B6UoQW7mTYErtHSi) icon in the top right of the plot
* Draw the lasso around the cluster of red cells and click the circle to close the lasso (Figure 28)
* Open the **Classify** tool and click **Classify selection**
* Name the classification **Microglia**
* Click **Save**

![Figure 26. Classifying microglia (red)](/files/0AGdruAAbh111WfAhgYL)

The blue cells are MAG positive, indicating that they are the oligodendrocytes from every sample.

* Switch to pointer mode by clicking ![image2017-12-28 13\_12\_15](/files/sSBYPOuizlIJf7OXvD1R) in the top right corner of the plot
* Deselect the cells by clicking on any blank space on the plot
* Switch to lasso mode again by clicking the ![image2017-12-28 13\_3\_50](/files/5OX0B6UoQW7mTYErtHSi) icon in the top right of the plot
* Draw the lasso around the cluster of blue cells and click the circle to close the lasso
* Open the **Classify** tool and click **Classify selection**
* Name the classification **Oligodendrocytes**
* Click **Save**

Finally, we will classify the BCAN expressing cells on the plot as glioma cells from every sample.

* Switch to pointer mode by clicking ![image2017-12-28 13\_12\_15](/files/sSBYPOuizlIJf7OXvD1R) in the top right corner of the plot
* Deselect the cells by clicking on any blank space on the plot
* Switch to lasso mode again by clicking the ![image2017-12-28 13\_3\_50](/files/5OX0B6UoQW7mTYErtHSi) icon in the top right of the plot
* Draw the lasso around the cluster of green cells and click the circle to close the lasso
* Open the **Classify** tool and click **Classify selection**
* Name the classification **Glioma**
* Click **Save**
* Switch to pointer mode by clicking ![image2017-12-28 13\_12\_15](/files/sSBYPOuizlIJf7OXvD1R) in the top right corner of the plot
* Deselect the cells by clicking on any blank space on the plot

The number of cells classified as microglia, oligodendrocytes, and glioma are shown in **Classify** (Figure 29)

![Figure 27. The number of cells for each cell type](/files/sCmLMYMEJgQaP1gY8dqm)

* Click **Apply classifications** in the **Classify** icon (Figure 30)

![Figure 28. Apply classifications to the data project](/files/lNd17lXOcE12mrRB0cIs)

* Name the classification attribute **Cell type (multi-sample)** (Figure 31)
* Click **Run**

![Figure 29. Name the cell-level attribute](/files/UG0txvkGNiCI9kXcA75F)

The new attribute is now available for downstream analysis.

* Click on the **Glioma (multi-sample)** project name at the top to go back to the *Analyses* tab
* Your browser may warn you that any unsaved changes to the data viewer session will be lost. Ignore this message and proceed to the *Analyses* tab

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Compare expression between cell types with multiple samples

* [Filter cells](#filter-cells)
* [Identify differentially expressed genes](#identify-differentially-expressed-genes)
* [Exploring differentially expressed genes](#exploring-differentially-expressed-genes)

Differential expression analysis can be used to compare cell types. Here, we will compare glioma and oligodendrocyte cells to identify genes differentially regulated in glioma cells from the oligodendroglioma subtype. Glioma cells in oligodendroglioma are thought to originate from oligodendrocytes, thus directly comparing the two cell types will identify genes that distinguish them.

## Filter cells

To analyze only the oligodendroglioma subtype, we can filter the samples.

* Click the **Filtered counts** data node
* Expand **Filtering** in the task menu
* Click **Filter cells** (Figure 1)

![Figure 1. Invoking the sample filter](/files/KjWe1rZZ7eWxamXuqH4P)

The filter lets us include or exclude samples based on sample ID and attribute.

* Set the filter to **Include** samples where **Subtype is Oligodendroglioma**
* Click **AND**
* Set the second filter to **exclude Cell type (multi-sample) is Microglia**
* Click **Finish** to apply the filter (Figure 2)

![Figure 2. Configuring the group filter](/files/DtRQze2YhIwvc2bEJIaE)

A *Filtered counts* data node will be created with only cells that are from oligodendroglioma samples (Figure 3).

![Figure 3. Filtering groups generates a Filtered counts data node](/files/9ZP4bBURZRxGPv8BRU5M)

## Identify differentially expressed genes

* Click the new **Filtered counts** data node
* Click **Statistics** > **Differential analysis** in the task menu
* Click **GSA**

The configuration options (Figure 4) includes sample and cell-level attributes. Here, we want to compare different cell types so we will include *Cell type (multi-sample)*.

* Click **Cell type (multi-sample)**
* Click **Next**

![Figure 4. Choosing attributes to include in the statistical test](/files/oTNYQu9vX1zKSi6wkoC4)

Next, we will set up a comparison between glioma and oligodendrocyte cells.

* Click **Glioma** in the top panel
* Click **Oligodendrocytes** in the bottom panel
* Click **Add comparison** (Figure 5)

This will set up fold calculations with glioma as the numerator and oligodendrocytes as the denominator.

![Figure 5. Defining the comparison between Glioma and Oligodendrocytes](/files/SlMslqP86bJD2Uv4z5xy)

* Click **Finish** to run the GSA

A green *GSA* data node will be generated containing the results of the GSA.

* Double-click the **green** **GSA** data node to open the GSA report

Because of the large number of cells and large differences between cell types, the p-values and FDR step up values are very low for highly significant genes. We can use the volcano plot to preview the effect of applying different significance thresholds.

* Click ![image2018-2-16 13\_8\_45](/files/OyPT7C228Af1rWjNBfbv) to view the **Volcano plot**
* Open the **Style** icon on the left, change *Size* **point size** **to 6**
* Open the **Axes** icon on the left and change the Y-axis to **FDR step up (Glioma vs Oligodendrocytes)**
* Open the **Statistics** icon and change the *Significance* of **X threshold to -10 and 10** and the **Y threshold to 0.001**
* Open the **Select & Filter** icon, set the **Fold change thresholds to -10 and 10**
* In **Select & Filter**, click ![Remove\_icon](/files/ipOd3jnCN34HHiiCK98K) to remove the **P-value (Glioma vs Oligodendrocytes)** selection rule. From the drop-down list, add **FDR step up (Glioma vs Oligodendrocytes)** as a selection rule and set the maximum to 0.001

Note these changes in the icon settings and volcano plot below (Figure 6).

![Figure 6. Previewing a filter by adjusting the size of the points, changing the Y-axis, adjusting the X & Y significance thresholds and changing the selection criteria](/files/yfJmScJ6ZPU8WxMKka9Q)

We can now recreate these conditions in the GSA report filter.

* Click **GSA report** tab in your web browser to return to the GSA report
* Click **FDR step up**
* Set the **FDR step up** filter to **Less than or equal to** **0.001**
* Press **Enter**
* Click **Fold change**
* Set the **Fold change** filter to **From -10 to 10**
* Press **Enter**

The filter should include 291 genes.

* Click ![image2018-2-16 13\_19\_34](/files/O4wVIxP9siGBxz4iW0GN) to apply the filter and generate a *Filtered Feature list* node

## Exploring differentially expressed genes

To visualize the results, we can generate a hierarchical clustering heatmap.

* Click the **Filtered feature list** produced by the *Differential analysis filter* task
* Click **Exploratory analysis** in the task menu
* Click **Hierarchical clustering/heatmap**

Using the hierarchical clustering options we can choose to include only cells from certain samples. We can also choose the order of cells on the heatmap instead of clustering. Here, we will include only glioma cells and order the samples by sample name (Figure 7).

* Make sure **Cluster** is unchecked for *Cell order*
* Click **Filter cells** under *Filtering* and set the filter to **include Cell type (multi-sample) is Glioma**
* Choose **Sample name** from the *Cell order* drop-down menu in the *Assign order* section
* Click **Finish**

![Figure 7. Configuring hierarchical clustering](/files/UVai81WfMfmG0cA3oSzc)

* Double click the green **Hierarchical clustering** node to open the heatmap

The heatmap differences may be hard to distinguish at first; the range from red to blue with a white midpoint is set very wide because of a few outlier cells. We can adjust the range to make more subtle differences visible. We can also adjust the color.

* Set the **Range** toggle **Min** to **-1.5**
* Set the **Range** toggle **Max** to **1.5**

The heatmap now shows clear patterns of red and blue.

* Click **Axis titles** and deselect the **Row labels** and **Column labels** of the panel to hide sample and feature names, respectively.
* Select **Sample name** from the *Annotations* drop-down menu

Cells are now labeled with their sample name. Interestingly, samples show characteristic patterns of expression for these genes (Figure 8).

![Figure 8. Hierarchical clustering heatmap with cells on rows (ordered by sample name) and genes on columns (clustered)](/files/HJ9pP3GdruC49MyyAtpB)

* Click **Glioma (multi-sample)** to return to the *Analyses* tab.

We can use gene set enrichment to further characterize the differences between glioma and oligodendrocyte cells.

* Click the **Filtered feature list** node
* Click **Biological interpretation** in the task menu
* Click **Gene set enrichment**
* Change *Database* to **Gene set database** and click Finish to continue with the most recent gene set (Figure 9)

![Figure 9. Gene set enrichment dialogue](/files/VCkojobNLYOImX0qyEmJ)

A *Gene set* *enrichment* node will be added to the pipeline .

* Double-click the **Gene set enrichment** task node to open the task report

Top GO terms in the enrichment report include "ensheathment of neurons" and "axon ensheathment" (Figure 10), which corresponds well with the role of oligodendrocytes in creating the myelin sheath that supports and protect axons in the central nervous system.

![Figure 10. GO enrichment task report](/files/seQD9MWi2rODcmzvWRIN)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Analyzing Single Cell ATAC-Seq data

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2FebDxwbqfRGJbMOZdoLBS%2FscATACSeq%20demo.mp4?alt=media&token=06ab49cc-a517-4230-972b-0753da51adfc>" %}

For your convenience, here is a video showing the below steps.

* [Transfer files and create a new project](#transfer-files-and-create-a-new-project)
* [Import the FASTQ files](#import-the-fastq-files)
* [Convert FASTQ to count](#convert-fastq-to-count)
* [QA/QC](#qa-qc)
* [Filter cells](#filter-cells)
* [Filter features](#filter-features)
* [Annotate regions](#annotate-regions)
* [TF-IDF (frequency-inverse document frequency) normalization](#tf-idf-frequency-inverse-document-frequency-normalization)
* [SVD (singular value decomposition)](#svd-singular-value-decomposition)
* [Graph-based clustering](#graph-based-clustering)
* [UMAP](#umap)
* [Promoter sum matrix](#promoter-sum-matrix)
* [Classifying cells](#classifying-cells)
* [Differential analysis](#differential-analysis)
* [Pipeline](#pipeline)
* [References](#references)

This guide illustrates how to process FASTQ files produced using the 10x Genomics Chromium Single Cell ATAC assay to obtain a Single cell counts data node, which is the starting point for analysis of single-cell ATAC experiments.

If you are new to Partek Flow, please see [Getting Started with Your Partek Flow Hosted Trial](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial) for information about data transfer and import and [Creating and Analyzing a Project](/partek-flow/tutorials/creating-and-analyzing-a-project) for information about the Partek Flow user interface.

This tutorial uses a [10X 5k PBMC dataset](https://www.10xgenomics.com/resources/datasets/5-k-peripheral-blood-mononuclear-cells-pbm-cs-from-a-healthy-donor-next-gem-v-1-1-1-1-standard-2-0-0) if you would like to follow along exactly.

## Transfer files and create a new project

We recommend uploading your FASTQ files (fastq.gz) to a folder on your Partek Flow server before importing them into a project. Data files can be transferred into Flow from the *Home* page by clicking the **Transfer file** button (Figure 1). Following the instruction In Figure 1 to complete the data transfer. Users have the option to change the **Upload directory** by clicking the **Browse** button and either select another existing directory or create a new directory.

To create a new project, from the *Home* page click the **New Projec**t button; enter a project name and then click **Create project**. Once a new project has been created, click the **Add data** button in the *Analyses* tab.

![Figure 1. Transfer file in Partek Flow.](/files/q663UqLiBJ32G2LenMj7)

## Import the FASTQ files

To proceed, click the **Add data** button in the *Analyses* tab. In the *Single cell > scATAC-Seq* section select **fastq** and click **Next**. The file browser interface will open (Figure 3). Select the FASTQ files using the file browser interface and push the **Finish** button to complete the task. Paired end reads will be automatically detected and multiple lanes for the same sample will be automatically combined into a single sample. We encourage users to include all the FASTQ files including the index files although they are optional.

When the FASTQ files have finished importing, the *Unaligned reads* data node will appear in the *Analyses* tab.

![Figure 2. Data tab in Partek Flow.](/files/tpVScgz0JawD2VQ71Is5)

![Figure 3. Input FASTQ files for scATAC-Seq data in Flow.](/files/ld2bUDnt2YujMM6jLC9K)

## Convert FASTQ to count

To deal with the single cell ATAC-seq FASTQ data, Partek Flow has wrapped the 'cellranger-atac count' pipeline from Cell Ranger ATAC v2.0\[1]. It takes FASTQ files and performs multiple analysis simultaneously including reads filtering and alignment, barcode counting, identification of transposase cut sites, peak and cell calling, and generates the count matrix.

To run Cell Ranger - ATAC task:

* Click the *Unaligned reads* data node
* Select **Cell Ranger - ATAC** in the **10x Genomics** section in the task menu on the right
* Select **Single cell ATAC** in *Assay type* for ATAC-Seq data only
* Choose the proper *Reference assembly* for the data (you may have to create the reference)
* Press the **Finish** button to run the task with default settings (Figure 4)

![Figure 4. Convert FASTQ by Cell Ranger - ATAC task in Flow.](/files/vCxelBaFejk8Jipo2M7B)

To learn more about how to run [Cell Ranger - ATAC](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/user-manual/task-menu/10x-genomics/cell-ranger---atac.md) task in Flow, please refer to our online [documentation](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/user-manual/task-menu/10x-genomics/cell-ranger---atac.md).

The output of the count matrix then becomes the starting point for downstream analysis for scATAC-seq data in Flow (Figure 5).

![Figure 5. Single cell QA/QC task for scATAC-Seq data in Flow.](/files/lHP1n0RpsVJllCfpbzyO)

## QA/QC

An important step in analyzing single cell ATAC data is to filter out low quality cells. A few examples of low-quality cells are doublets, cells with a low TSS enrichment score, cells with a high proportion of reads mapping to the genomic blacklist regions, or cells with too few reads to be analyzed. Users are able to do this in Partek Flow using the Single cell QA/QC task.

* Click on the *Single cell counts* node
* Click on the **QA/QC** section in the task menu
* Click on **Single cell QA/QC**

A task node, *Single cell QA/QC*, is produced. Initially, the node will be semi-transparent to indicate that it has been queued, but not completed. A progress bar will appear on the *Single cell QA/QC* task node to indicate that the task is running (Figure 5).

* Click the *Single cell QA/QC* node once it finishes running
* Click **Task report** in the task menu

![Figure 6. QA/QC task report for scATAC - Seq data in Flow.](/files/XYQh4qJgomYhDFCWACBQ)

The *Single cell QA/QC* report includes interactive violin plots showing the value of every cell in the project on several quality measures (Figure 6).

There are five plots: Nucleosome signal, TSS enrichment, % reads in peaks, Blacklist ratio, and Peak region fragments. Each point on the plots is a cell and the violins illustrate the distribution of values for the y-axis metric. Cells can be filtered either by clicking and dragging to select a region on one of the plots or by setting thresholds using the filters below the plots. Here, we will apply a filter for the number of read counts. The plot will be shaded to reflect the filter. Cells that are excluded will be shown as black dots on both plots.

Descriptions of QC metrics:

**Nucleosome signal**: calculated per single cell, which quantifies the approximate ratio of mononucleosomal to nucleosome-free fragments. The histogram of DNA fragment sizes (determined from the paired-end sequencing reads) should exhibit a strong nucleosome banding pattern which corresponds to the length of DNA wrapped around a single nucleosome.

**TSS enrichment**: Transcriptional start site (TSS) enrichment score. The ENCODE project has defined an ATAC-seq targeting score based on the ratio of fragments centered at the TSS to fragments in TSS-flanking regions (see <https://www.encodeproject.org/data-standards/terms/>). Poor ATAC-seq experiments typically will have a low TSS enrichment score.

**Peak region fragments**: total number of fragments in peaks which is a measure of cellular sequencing depth/complexity. Cells with very few reads may need to be excluded due to low sequencing depth. Cells with extremely high levels may represent doublets, nuclei clumps, or other artifacts.

**% reads in peaks**: Represents the fraction of all fragments that fall within ATAC-seq peaks. Cells with low values (i.e. <15-20%) often represent low-quality cells or technical artifacts that should be removed. Note that this value can be sensitive to the set of peaks used.

**Blacklist ratio**: The ENCODE project has provided a list of blacklist regions, representing reads which are often associated with artifactual signals. Cells with a high proportion of reads mapping to these areas (compared to reads mapping to peaks) often represent technical artifacts and should be removed.

## Filter cells

To filter out low quality cells (Figure 7),

* Open the **Select & Filter** menu
* Set the filters on nucleosome signal **< 4;** Peak region fragment **500-30000**; leave the rest as they are
* Click the filter icon ![Screenshot 2023-09-29 at 11 11 11](/files/4Iv8eQwsJU8RzJ7XHDZD) and **Apply observation filter** to run the Filter cells task on the first *Single cell ATAC counts* data node, it generates a *Filtered cells* node

![Figure 7. Filter low quality cells in Partek Flow.](/files/agYEi1pibnArqfVqJMb4)

## Filter features

Another common task is to filter the data to include only informative features. Partek Flow has a wide variety of flexible filtering options.

Filter features task can be invoked from any counts or single cell data node. Noise Reduction and Statistics Based filters take each feature and perform the specified calculation across all the cells. The filter is applied to the values in the selected data node and the output is a filtered version of the input data node.

In the task dialog, click the check box to activate one or more of the filter types, configure the filter(s), and click **Finish** to run (Figure **8**).

![Figure 8. Filter features in Partek Flow.](/files/CCY0HAPk1woGeTlg3gTJ)

## Annotate regions

To understand the importance of enriched regions in regulating gene expression, Flow uses **Annotate regions** task to add information about overlapping or nearby genomic features. That gives regulatory context for enriched regions.

The input for *Annotate peaks* is a Peaks type data node.

* Click the **Filtered features** data node
* Click the **Peak analysis** section in the toolbox
* Click **Annotate regions**
* Set the *Genomic overlaps* parameter

The *Genomics overlaps* parameter lets you choose one of two options (Figure 9).

* *Report one gene region per peak (precedence applies)* chooses one gene section for each peak using the precedence order to settle cases where more than one gene section overlaps a peak. The order of precedence is TSS, TTS, CDS Exon, 5' UTR Exon, 3' UTR Exon, Intron, Intergenic.
* *Report all gene regions per peak* creates a row for each gene section that overlaps a peak in the task report.

![Figure 9. Annotate regions in Partek Flow.](/files/FTKqbrhFMgUVMSlulFc2)

Users are able to define the transcription start site (TSS) and transcription termination site (TTS) limit in the unit of bp.

* Choose a gene/feature annotation from the drop-down menu
* Click **Finish** to run

## TF-IDF (frequency-inverse document frequency) normalization

Latent semantic indexing (LSI) was first introduced for the analysis of scATAC-seq data by [Cusanovich *et al.* 2018](https://www.nature.com/articles/nature25981)\[2]. LSI combines steps of frequency-inverse document frequency (TF-IDF) normalization followed by singular value decomposition (SVD). Partek Flow wrapped Signac's TF-IDF normalization for single cell ATAC-seq dataset. It is a two-step normalization procedure that both normalizes across cells to correct for differences in cellular sequencing depth, and across peaks to give higher values to more rare peaks\[3].

**TF-IDF normalization** in Flow can be invoked in *Normalization and scaling* section by clicking any *single cell counts* data node (Figure 10).

![Figure 10. TF-IDF normalization for scATAC-Seq in Flow.](/files/6hXWbwZL4cIi1f21dO6g)

To run **TF-IDF normalization**,

* Click a S**ingle cell counts** data node, in this case the **Annotated regions** node
* Click the **Normalization and scaling** section in the toolbox
* Click **TF-IDF normalization**

The output of **TF-IDF normalization** is a new data node that has been normalized by log(*TF x IDF*)*.*

## SVD (singular value decomposition)

Singular value decomposition (SVD) will be applied to *TF-IDF* output in scATAC-Seq data. It returns a reduced dimension representation of a matrix. Although SVD and Principal components analysis (PCA) are two different techniques, the SVD has a close connection to PCA. Because PCA is simply an application of the SVD. For users who are more familiar with scRNA-Seq, you can think of SVD as analogous to the output of PCA. And similarly, the statistical interpretation of singular values is in the form of variance in the data explained by the various components.

To run **SVD** task,

* Click a **Normalized counts** data node
* Click the **Exploratory analysis** section in the toolbox
* Click **SVD**

The GUI is simple and easy to understand. The **SVD** dialog is only asking to select the number of singular values to compute (Figure 11). By default 100 singular values will be computed if users don't want to compute all of them. However, the number could be adjusted manually or typed in directly. Simply click the **Finish** button if you want to run the task as default.

The task report for **SVD** is similar to PCA\_.\_ Its output will be used for downstream analysis and visualization, including Harmony and WNN.

![Figure 11. SVD task configuration dialog in Partek Flow.](/files/Zwm8mVaULLiaDbnfp8wy)

## Graph-based clustering

Graph-based clustering (Figure 12) identifies groups of similar cells using SVD values as the input. By including the informative SVDs, noise in the data set is excluded, improving the results of clustering.

* Click the **SVD** output
* Click **Exploratory analysis** in the task menu
* Click **Graph-based clustering**
* Check **Compute biomarkers**
* Click **Finish** to run as default

![Figure 12. Configure Graph-based clustering in Flow.](/files/pt57zpitP08gDEZpBmep)

A new *Graph-based clusters* data and a *Biomarkers* data node will be generated.

* Double-click the **Graph-based clusters** node to see the cluster results and statistics (Figure 13)
* Double-click the **Biomarkers** node to see the computed biomarkers if you have selected this option (Figure 14)

The *Graph-based clustering result* (Figure 13) lists the *Total number of clusters* and what proportion of cells fall into each cluster as well as *Maximum modularity* which is a measurement of the quality of the clustering result where optimal modularity is 1. The *Biomarkers* report (Figure 14) includes the top features for each graph-based cluster. It displays the top-10 genes that distinguish each cluster from the others. **Download** at the bottom right of the table can be used to view and save more features. These are calculated using an ANOVA test comparing the cells in each group to all the other cells, filtering to genes that are 1.5 fold upregulated, and sorting by ascending *p-value*. This ensures that the top-10 genes of each cluster are highly and disproportionately expressed in that cluster.

![Figure 13. Graph-based clustering results in Flow.](/files/T4xHav8tyTJEG4lhh6hI)

![Figure 14. Computer biomarkers results in Flow.](/files/DgQJrtUYVtBAMtmfHD6z)

## UMAP

Similar to t-SNE, Uniform Manifold Approximation and Projection (UMAP) is a dimensional reduction technique. UMAP aims to preserve the essential high-dimensional structure and present it in a low-dimensional representation. UMAP is particularly useful for visually identifying groups of similar samples or cells in large high-dimensional data sets.

To run UMAP (Figure **15**):

* Click the **SVD** data node
* Click the **Exploratory analysis** section of the toolbox
* Click **UMAP**
* Click **Finish** to run with default settings

UMAP produces a *UMAP* task node. Opening the task report launches a scatter plot showing the *UMAP* results. Each point on the plot is a cell for single cell data. The plot will open in 2D or 3D depending on the user preference.

![Figure 15. UMAP configuration in Partek Flow.](/files/LhPitpkEAcXDigqCpijC)

## Promoter sum matrix

The Annotate regions task in Flow labels individual peaks as promoters for a particular gene if the peak falls 1000 bases upstream from a gene's transcription start site, or 100 bases downstream from a gene's transcription start site by default (Figure 9). A promoter sum for a given gene is the number of cut sites per cell that fall within all the peaks labeled as promoters (-1000bp \~ 100bp by default or user defined through Annotate regions) for that gene. Higher promoter sum values indicate higher chromatin accessibility in the promoter region \[4].

Flow task **Promoter sum matrix** summarizes each promoter sum and outputs a cell x gene matrix. In the matrix, only genes that have peaks within its promoter region have been included. In Flow **Promoter sum matrix** can be invoked in the Peak analysis section by clicking the Annotated regions data node (Figure 16).

To run **Promoter sum matrix** in Flow,

* Click the **Annotated regions** data node
* Click the **Peak analysis** section in the toolbox
* Click **Promoter sum matrix**

Once the task has been finished, a new data node will be produced where the promoter sum value for each feature can be used to color UMAP/t-SNE and to determine cell type with raw data. We recommend users normalize its output prior to color the UMAP just like the scRNA-seq data.

![Figure 16. Promoter sum matrix in Flow.](/files/ZEBAwkLJdb7gqPxLc75f)

## Classifying cells

Double-clicking the *UMAP* task node will open the task report in the Data Viewer.

To classify a cell, just select it then click **Classify selection** in the **Classify** tool.

For example, we can classify a cluster of cells expressing high levels of ***MS4A1*** as B cells.

* Make sure the right data source has been selected. For scATAC-seq data, it shall be the normalized counts of promoter sum values in most cases (Figure 17)
* Set *Color by* in the **Style** configuration to the normalized counts node
* Type ***MS4A1*** in the search box and select it. Rotate the 3D plot if you need to see this cluster more clearly.
* Click ![Screenshot 2023-09-29 at 14 48 59](/files/5OX0B6UoQW7mTYErtHSi) to activate Lasso mode
* Draw a lasso around the cluster of *MS4A1*-expressing cells
* Click **Classify selection** under *Tools* in the left panel
* Type **B cells** for the Name
* Click **Save** (Figure 18)

Repeat the above steps to finish the other cell type classifications. To be able to use the classifications in downstream tasks and visualizations, you must first apply them.

* Click **Apply classifications**
* **Name** the classification (e.g. Cell type)
* Check the **Compute biomarkers** if needed
* Click **Run** to complete the task

Once the classifications have been added to the project, one can color the UMAP/t-SNE plot by the Classification or compare the differentially expressed genes between different cell types.

![Figure 17. Select the data source in Data Viewer.](/files/JAy8sr2qlHdf2ZZkpGzm)

![Figure 18. Color cells in UMAP by MS4A1 in Flow.](/files/N0OtHQxdESdLtZBpWsKq)

## Differential analysis

To identify genes that distinguish a cell type, one can use the differential analysis tools in Partek Flow.

* Click the **TF-IDF normalized counts** data node
* Click the **Differential analysis** section in the toolbox
* Click **Hurdle model**
* Select the factors and interactions to include in the statistical test (Figure 19). Cell type has been selected here as an example.

![Figure 19. Hurdle model for differential analysis in Flow.](/files/3hojzRm0cTsZC7LJNVZP)

* Click **Next**
* Define comparisons between factor or interaction levels (Figure 20)
* Click **Add comparison** to add the comparison to the *Comparisons* table.
* Click **Finish** to run the statistical test as default

![Figure 20. Define comparisons in Hurdle model.](/files/EIHp5eRxkORNhMPsDx47)

*Hurdle model* produces a *Feature list* task node. The results table and options are the same as the [GSA](/partek-flow/user-manual/task-menu/differential-analysis/gsa) task report except the last two columns. The percentage of cells where the feature is detected (value is above the background threshold) in different groups (Pct(group1), Pct(group2)) are calculated and included in the Hurdle model report.

A filtered Feature list data node can be produced by running the Differential analysis filter in the Hurdle model task report (Figure 21) .

![Figure 21. Generate filtered node for differential analysis results in Flow.](/files/ZhYCt5ku1cdDSE8xVI4d)

Once we have filtered a list of differentially expressed genes, we can visualize these genes by generating a [heatmap](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering), or perform the Gene set enrichment analysis and [motif detection](/partek-flow/user-manual/task-menu/motif-detection).

## Pipeline

![Figure 22. Described pipeline shown in the Analyses tab](/files/1U2ZhbzwrQ8nVJq6dfSW)

For information about automating steps in this analysis workflow, please see our documentation page on [Making a Pipeline](/partek-flow/user-manual/pipelines/making-a-pipeline).

## References

1. <https://support.10xgenomics.com/single-cell-atac/software/pipelines/latest/what-is-cell-ranger-atac>
2. Cusanovich, D., Reddington, J., Garfield, D. *et al.* The *cis*-regulatory dynamics of embryonic development at single-cell resolution. *Nature* **555,** 538–542 (2018). <https://doi.org/10.1038/nature25981>
3. <https://satijalab.org/signac/index.html>
4. <https://support.10xgenomics.com/single-cell-atac/software/visualization/latest/tutorial-celltypes>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Analyzing Illumina Infinium Methylation array data

This guide provides instructions for analyzing Illumina Infinium Methylation array data.

The tutorial uses this [Infinium Methylation Screening Array Demo Data Set](https://support.illumina.com/downloads/infinium-methylation-demo-data-set.html) if following along exactly with the analyses pipeline.

If you are new to Partek Flow, please see the [Quick start guide](https://help.partek.illumina.com/partek-flow/quick-start-guide) for information about the Partek Flow user interface.

## Import Illumina methylation array data

We recommend uploading the microarray data to a folder on your Partek Flow server before importing into a project. Data files can be transferred to your server from the *Home* *page* by clicking the **Transfer file** button. Users have the option to change the *Upload directory* by clicking the **Browse** button and either select another existing directory or create a new directory. [Please click here for more information on transferring files to the server](https://help.partek.illumina.com/partek-flow/user-manual/importing-data#navigating-the-file-browser-to-transfer-files-to-the-server).

To create a new project on the *Home page* click the +**Add data** button, enter a project name, and click **Create project**.

<figure><img src="/files/FEcEvZsW0umBEk12l2lP" alt="" width="226"><figcaption><p><em>Give the project a name then click Create project</em></p></figcaption></figure>

Upon project creation you will land in the Analyses tab, prompting the addition of sample data to the project. Click the blue **Add data** button <img src="/files/kaJfyXF5Oqr1BoaSF10X" alt="Add data button" data-size="line">

When available, hover over <img src="/files/NVjRqSk652F2pR4YQ4xY" alt="Tooltip" data-size="line"> Tooltips or click the <img src="/files/j4U3Qyz71lz2KsNXlv0B" alt="Video icon" data-size="line"> video help for decision making.

<figure><img src="/files/ZzykwA8FfsN5eUoL0U2n" alt=""><figcaption><p><em>Add data to the project by clicking Add data (blue circle)</em></p></figcaption></figure>

Select **Microarray**, **Methylation** and **Illumina methylation idat** as the file format for import then click **Next**.

<figure><img src="/files/qkeTBenjxwBx3LFHZICf" alt=""><figcaption><p><em>Choose Microarray, Methylation, and Illumina methylation idat then click Next</em></p></figcaption></figure>

Navigate to the idat files that have been uploaded to the server. For this tutorial, there are two paired idat files per sample.

If you have not already transferred the files to the server you can choose to do this within the import task by clicking the **Transfer files to the server** button.

<figure><img src="/files/kuI7TT6mnOsFJX5BkMd7" alt=""><figcaption><p>T<em>ransfer idat files to the server by clicking Transfer files to the server</em></p></figcaption></figure>

This will bring you to the Transfer files page. Click the **Transfer files** button, add the files for transfer then click **Upload**. Do not terminate the browser or let your computer go to sleep during transfer. A time estimate for upload is provided but may change.

<figure><img src="/files/hv33Fgukio8dc6o77wgd" alt=""><figcaption><p><em>Transfer the files to the server where you can find them</em></p></figcaption></figure>

When the transfer completes the upload window will close. In the *Transfer files* page, the transferred files *Status* is *Complete*.

<figure><img src="/files/GUFgmpAvGVWoY6NWV0HP" alt=""><figcaption><p>Th<em>e Transfer files page shows the idat file transfer is complete</em></p></figcaption></figure>

Now, the selected files are on the server in the folder specified during transfer.

In the project import task, navigate to the files on your server. In the example below, only these files were saved to this folder so I will check the top box to select all then click **Finish**.

<figure><img src="/files/RN866SbLrI8DwHI0VyLC" alt=""><figcaption><p><em>Select the paired idat sample files for import into the project</em></p></figcaption></figure>

This starts the *Importing* of selected data to the project. The transparent task bar will complete as import progresses.

<figure><img src="/files/nvjYTR6X9iDPvt9EQ6IE" alt=""><figcaption><p><em>Importing the data to the project</em></p></figcaption></figure>

When the import completes, the *Microarray methylation* data node appears in the *Analyses* tab. Hovering of this node, we see 60 samples (572.28 MB data) are contained in this data node.

<figure><img src="/files/8ZaqOFze7d2HpnXHFIWD" alt=""><figcaption><p><em>The Microarray methylation data node contains the imported data, hover over the node to see details</em></p></figcaption></figure>

## Assign sample metadata

Add sample metadata to the project by navigating to the *Metadata* tab. Select the **Assign values from file** button as an efficient way to assign sample attributes using a tab delimited text file.

<figure><img src="/files/uZ175jgkAG5yh3zfZYVi" alt=""><figcaption><p><em>Add sample metadata in the Metadata tab</em></p></figcaption></figure>

If the samples metadata file is not already on the server, click **Transfer files to the server**.

<figure><img src="/files/WBFSE103zTx1ZcG6FrRf" alt=""><figcaption><p>Transfer the file to the server</p></figcaption></figure>

Select the file then click **Next**.

The tab delimited file should contain a table with the following:

* The first column of the table lists the sample names (the sample names in the file must be identical to the ones listed in the project *Sample name* column in the *Metadata* tab)
* The first row lists the attribute names (e.g. Treatment, Exposure)
* List any corresponding attributes for each sample in succeeding columns

Make any wanted modifications and click **Import**.

<figure><img src="/files/NRAdi7Ru7DlM9SiPuMVt" alt=""><figcaption><p><em>Add sample attributes from a file</em></p></figcaption></figure>

This adds the defined attribute information to the *Metadata* tab. **Manage** and **Assign values** buttons can be used to further modify sample attributes.

<figure><img src="/files/6Z0vIYoGqoB1tEINkZxe" alt=""><figcaption><p><em>Manage and Assign values in the Metadata tab to further modify sample attributes</em></p></figcaption></figure>

Click the left **Analyses** tab to navigate back to the analyses pipeline.

Single-click the **Microarray methylation** data node to run the first task using the context sensitive menu on the right. No tasks have been performed on this data so there is still an option to *Add data* to the project; once analysis tasks are performed from this data node, data can no longer be added to the project.

<figure><img src="/files/aGo7mMgOoGDaVQXdInF6" alt=""><figcaption><p><em>Run the first task on the Microarray methylation data node using the context sensitive menu on the right</em></p></figcaption></figure>

## Generate beta value

Using the task menu on the right under *Methylation analysis* select the **Generate beta value** task.

Choose the Chip name as *Infinium Methylation Screening Array*, If the Chip name is not listed use the dropdown to select **New Chip** then add the Illumina manifest file.

Keep the default settings and click **Finish**.

<figure><img src="/files/rfneT5y9HB7w82RwtS1g" alt=""><figcaption><p><em>Choose the appropriate manifest file (Chip name), keep the default settings then click Finish</em></p></figcaption></figure>

This task output is the *Methylation beta* data node. Methylation Beta-values are continuous variables between 0 and 1.

Single-click the *Methylation beta* data node node and choose the next **PCA** task under *Exploratory analysis* from the task menu.

<figure><img src="/files/0JesHHbPqD1nRTWj4d5t" alt=""><figcaption><p><em>Perform PCA task</em></p></figcaption></figure>

## Generate PCA

In the *Analyses* tab, select the *Methylation beta* data node then use the task menu *Exploratory analyses* options to run the **PCA** task *.* Keep the default settings the same and click **Finish**.

<figure><img src="/files/2SP7Lyip5S26RBVpSFSR" alt=""><figcaption><p><em>Que the PCA task with default settings and click Finish</em></p></figcaption></figure>

In the *Analyses* tab, double click the *PCA* data node (circle) or single click the data node and select **Task report** under *Task results* in the task menu to view the PCA results in the Data Viewer.

<figure><img src="/files/egWHJesWlZGmUyCmzbuq" alt=""><figcaption><p><em>PCA task results in the Data Viewer</em></p></figcaption></figure>

## Detect differential methylation

Single-click the *Methylation beta* data node node and perform the **Detect differential methylation** task under *Methylation analysis* in the task menu.

<figure><img src="/files/kV9mppVRFqBICZRE159r" alt=""><figcaption><p><em>Select the Methylation beta data node then que the Detect differential methylation task</em></p></figcaption></figure>

This task converts the Beta-values to M-values and uses these to perform ANOVA differential expression analysis. [Please click here for more information on the ANOVA model.](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/differential-analysis/anova-limma-trend-limma-voom)

Follow along with the task to make one-way or two-way ANOVA comparisons. The configured ANOVA model is performed on both Beta-value and M-value matrices. Click **Finish**.

This outputs the *Detect differential methylation* task report list. Open the **Task report** from the task menu or double-click the data node.

<figure><img src="/files/KbcErAxfj2Jq0rUfUS2S" alt=""><figcaption><p><em>Open the Detect differential methylation report</em></p></figcaption></figure>

The outputs of this task include significance as *P-value* and *FDR step up* which is from the M-values. The *LSMeans* of the groups and the *Difference* are of the Beta-values.

Click the Optional columns button to add more column data including annotation from the Illumina manifest file. For more information on these optional columns [please see the "Infinium Methylation Screening Array Manifest Column Headings pdf here.](https://support.illumina.com/downloads/infinium-methylation-screening-manifest-files.html)

Use the left filter panel to filter the results then click Generate filtered node.

<figure><img src="/files/WRjFH2zIGALXRTqPYiSN" alt=""><figcaption><p><em>Use the filter panel to filter the results then click Generate filtered node</em></p></figcaption></figure>

The filtered node is now available in the *Analyses* pipeline.

## Visualize filtered results with Hierarchical clustering / heatmap

Select the generated *Filtered feature list* data node and use the *Exploratory analysis* task menu dropdown to perform the **Hierarchical clustering / heatmap** task.

<figure><img src="/files/tlKoIsuhteJY2X011GOw" alt=""><figcaption></figcaption></figure>

In the *Hierarchical clustering / heatmap* task settings, change *Sample order* to **Assign order**, select the attribute, then click **Finish**. This will order the heatmap rows based on the attribute order assigned.

<figure><img src="/files/wvhgC7wpiIeNZDYl4kAN" alt=""><figcaption><p><em>Modify the heatmap task settings and click Finish</em></p></figcaption></figure>

[Click here for more information on Hierarchical clustering task settings to visualize a Heatmap or Bubble map](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering#invoking-hierarchical-clustering)

Completion of this task will output the *Hierarchical clustering / heatmap* task results. Double-click this node or select **Task report** under *Task results* from the task menu to open the results in the Data viewer.

<figure><img src="/files/k3CqUdrny4HU7rRIjk5F" alt=""><figcaption><p><em>Open the Hierarchical clustering / heatmap task results</em></p></figcaption></figure>

The heatmap visualization can be altered within the Data viewer using both the left menu and the in-plot controls.

<figure><img src="/files/mh1NS1IyCwoaING31sh0" alt=""><figcaption><p><em>Modify the task results using the Data viewer controls</em></p></figcaption></figure>

[Please click here for more information on using the Data viewer to modify visualizations.](https://help.partek.illumina.com/partek-flow/user-manual/data-viewer)

Click **Save as** in the left menu to save the data viewer session to return and make changes.

To save the full individual image within the Data viewer to your machine, click the in-plot <img src="/files/aeNqRq6OWkPDl4G0ozzW" alt="Save button" data-size="line"> **Export image** icon in the top right corner of the image, choose **All data** then select the format, size, and resolution and click **Save**.

<figure><img src="/files/w6qc9oOGiWZmIEbBwruw" alt=""><figcaption><p><em>Save the heatmap to your machine by selecting All data</em></p></figcaption></figure>

### Perform biological interpretation

Select the *Filtered feature list* data node within the Analyses tab and choose the **Gene set enrichment** task under *Biological interpretation* in the task menu.

<figure><img src="/files/wFZpiE6vkBFlxHR7fMGq" alt=""><figcaption><p><em>Run the Gene set enrichment task on the filtered feature list data node</em></p></figcaption></figure>

Use the *Gene set enrichment* task settings to change default parameters.

Choose the KEGG database. Click the dropdown to change the selection; select **New library to** add the most recent library available.

Check **Select feature identifier** as *Gene Symbol* or *Gene IDs* to ensure genes are used instead of probe IDs for this task.

After optimizing the task settings, click Finish.

<figure><img src="/files/xhwi5uIcW1C891rMn2nd" alt=""><figcaption><p><em>Optimize the settings for the Gene set enrichment task</em></p></figcaption></figure>

The output of the *Gene set enrichment* task node is the *Pathway enrichment* data node.

### Filter the Gene set enrichment result

Filter the KEGG pathway gene sets by selecting the Pathway enrichment data node then clicking the **Filter gene sets** task from the task menu.

<figure><img src="/files/GDn4w8OhIIbEJVBWbFgI" alt=""><figcaption><p><em>Select the Pathway enrichment and click the Filter gene sets task</em></p></figcaption></figure>

Modify the Filter by parameters to include P-value < 0.050 then click Finish.

<figure><img src="/files/oXBxV4NaBZ0Vw9mfV5kf" alt=""><figcaption><p><em>Modify the Filter by parameters to include P-value &#x3C; 0.050 and click Finish</em></p></figcaption></figure>

Completion of this task will output a *Filter gene sets* task node and a *Filtered list* data node in the Analyses tab. Open the filtered gene sets by selecting the *Filtered list* data node and clicking the task menu **Task report** from the right task menu.

<figure><img src="/files/RNRDAzOXs5e9oOErIYW4" alt=""><figcaption><p><em>Open the filtered gene sets by double clicking on the Filtered list data node</em></p></figcaption></figure>

Because we filtered the gene sets to fewer than 100 rows, click the button to **View plots in the Data Viewer**. The filtering step can also be performed within the Pathway enrichment report.

<figure><img src="/files/qi4q9uZbVQPUM6yh4dEF" alt=""><figcaption><p><em>Filtering to fewer than 100 rows allows the View plots in the Data Viewer button</em></p></figcaption></figure>

This opens the filtered list report in the Data viewer for further modification.

<figure><img src="/files/Yhl0AV2sqS0L9FmJCVQK" alt=""><figcaption><p><em>Filtered list of KEGG pathway results in the Data viewer</em></p></figcaption></figure>

[Please click here for more information on using the Data viewer to modify visualizations.](https://help.partek.illumina.com/partek-flow/user-manual/data-viewer)

Save this data viewer session, with a meaningful name to revisit for future analysis, using the **Save as** button in the left menu.

This saved session is accessible by selecting the project **Data viewer** tab or by clicking *Data Viewer* within the breadcrumb (shown above the Data viewer canvas).

[Please click here for more information on the Gene set enrichment](https://help.partek.illumina.com/partek-flow/user-manual/task-menu/biological-interpretation/gene-set-enrichment) task like the interactive KEGG pathway maps.

## Analyses pipeline for Infinium methylation array data analysis

This completes the example analyses pipeline.

<figure><img src="/files/W4DQI8XRU797rCEgYAp0" alt=""><figcaption><p><em>Analyses pipeline for Infinium methylation array</em></p></figcaption></figure>


# NanoString CosMx Tutorial

This tutorial, will demonstrate:

* [Importing CosMx data](broken://spaces/ZG5p2nnl0YGy3GEZrqCO/pages/xMEg47lav98jYQ7xaN2t)
* [QA/QC, data processing, and dimension reduction](broken://spaces/ZG5p2nnl0YGy3GEZrqCO/pages/Wz2aBJveCUrZbxtIagjG)
* [Cell Typing](broken://spaces/ZG5p2nnl0YGy3GEZrqCO/pages/jAlV1IPdfKYSaGtZzz5m)
* [Classify subpopulations & differential expression analysis](broken://spaces/ZG5p2nnl0YGy3GEZrqCO/pages/RfobB7Lu2WFKMNFqGXI1)

### Tutorial Data Set <a href="#nanostringcosmxtutorial-tutorialdataset" id="nanostringcosmxtutorial-tutorialdataset"></a>

The tutorial data is based on the [CosMx Human Frontal Cortex FFPE Dataset](https://nanostring.com/products/cosmx-spatial-molecular-imager/ffpe-dataset/human-frontal-cortex-ffpe-dataset/)

\\


# Importing CosMx data

### Obtain and add files to the project <a href="#importingcosmxdata-obtainandaddfilestotheproject" id="importingcosmxdata-obtainandaddfilestotheproject"></a>

There are 5 data files needed for CosMx data import: polygons.csv, exprMat\_file.csv, fov\_positions\_file.csv, metadata\_file.csv, tx\_file.csv. The last file (tx\_file.csv) contains the transcript information and is not required when importing protein data. Additionally the user will need an image folder, called either CellComposite or CellOverlay (**please ensure the images in this folder are not blank before uploading them to Flow**). This image folder contains one image per FOV and will be used for the visualisation of the spatial data.

Create a folder per sample, the folder will contain the 5 files and the image folder.

Transfer the data to your Flow server before moving onto the import step.

### Data import <a href="#importingcosmxdata-dataimport" id="importingcosmxdata-dataimport"></a>

* Create a new project and click the **'Add data'** button in the *Analyses* tab

<figure><img src="/files/jS0JbM1szdp8e2XXSGc5" alt=""><figcaption></figcaption></figure>

* Select *Single cell > Spatial > NanoString CosMx* and click **Next**

<figure><img src="/files/p8bd6RCRsh4jGI6VbdLj" alt=""><figcaption></figcaption></figure>

* Click **Add sample**, name it and select the sample folder
* The importer will automatically select the image folder. If you have uploaded more than one image folder, select the one you want to use for the project.
* Select the appropriate annotation files (in this case we'll use hg38 - Ensembl Transcripts release 109)
* Leave the rest of the settings as default
* Click **Finish**

<figure><img src="/files/CaVpkxzjNzEg7z0XmInL" alt=""><figcaption></figcaption></figure>

### Additional Assistance <a href="#importingcosmxdata-additionalassistance" id="importingcosmxdata-additionalassistance"></a>

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.

![](https://documentation.partek.com/download/resources/com.adaptavist.confluence.rate:rate/resources/themes/v2/gfx/loading_mini.gif)


# QA/QC, data processing, and dimension reduction

### QA/QC & Data Processing <a href="#qa-qc-dataprocessing-anddimensionreduction-qa-qc-and-dataprocessing" id="qa-qc-dataprocessing-anddimensionreduction-qa-qc-and-dataprocessing"></a>

<figure><img src="/files/mJGNiOTRkQYdrnoCxDjE" alt=""><figcaption></figcaption></figure>

Once the data has been imported in the project we can start pre-processing the data:

We will first remove all non-expression features in the data (e.g. NegProbes).

* Click on *Filtering > Filter features* from the menu on the right
* Select *Metadata* and set the task settings as follows
* Click **Finish**

<figure><img src="/files/sMMZ4CrkXe0LpdKn4oXq" alt=""><figcaption></figcaption></figure>

In the Analyses tab

* Click on the resulting filtered counts node
* Select *QA/QC > Single cell* QA/QC from toolbox, once the task has completed we can open the report by double-clicking the node:

<figure><img src="/files/bY02LQezMW8rn1VlBeGb" alt=""><figcaption></figcaption></figure>

We will remove the cells with low counts and number of detected features.

* Click on *Select & Filter* and set lower threshold to 50 for both (remember that this is data-dependent and will change based on your dataset)
* Click Filter![](https://documentation.partek.com/download/thumbnails/98206053/Screenshot%202024-07-04%20at%2017.18.31.png?version=1\&modificationDate=1720109917705\&api=v2) include
* Click *Apply observation filter* to the filtered counts node:

<figure><img src="/files/aefwZw0iGDRYUS6mPUWH" alt=""><figcaption></figcaption></figure>

Click on the node generated by the filtering task in the Analyses tab.

* Click *Filtering > Filter features.* Apply a noise reduction filter:

<figure><img src="/files/ey8682ULGfx5VNG0bN1P" alt=""><figcaption></figcaption></figure>

We can now normalize our filtered data.

* Click *Normalization and scaling > Normalization.* Use the recommended settings by clicking ![](https://documentation.partek.com/download/thumbnails/98206053/Screenshot%202024-07-04%20at%2017.29.01.png?version=1\&modificationDate=1720110545187\&api=v2):

<figure><img src="/files/UKFzkHGg7FU8m3csst4o" alt=""><figcaption></figcaption></figure>

### Data Exploration <a href="#qa-qc-dataprocessing-anddimensionreduction-dataexploration" id="qa-qc-dataprocessing-anddimensionreduction-dataexploration"></a>

Now that we have filtered low quality cells and normalized our data, we can start clustering to identify cell populations.

* Click on the normalized data node
* From the menu on the right select *Exploratory analysis > PCA.* We are going to use the top 2000 features by variance and calculate the first 50 principal components (PCs):

<figure><img src="/files/Z6z3dW1oXPlPlaDnxY1p" alt=""><figcaption></figcaption></figure>

Once the PCA has run, click on the PCA result node in the Analyses tab.

* Select *Exploratory analysis > UMAP* from the toolbox. Set the UMAP parameters as follows:
  * Top **20** PCs
  * Local neighborhood size **60**
  * Minimal distance **0.20**

<figure><img src="/files/EYRNEz0c0UJA4mKvReO2" alt=""><figcaption></figcaption></figure>

While the UMAP is running we can also queue a clustering task. Click on the PCA result node in the Analyses tab, select *Exploratory analysis > Graph-based clustering.*

* We are going to use the Leiden algorithm to cluster our data (make sure to select the radio button for it)
* Set the number of PCs to **10**
* In the advanced settings, set the resolution parameter to **8e-5** and click **Apply**:

<figure><img src="/files/2IJSZSJHz3acbZNqKsys" alt=""><figcaption></figcaption></figure>

\\

### Additional Assistance <a href="#qa-qc-dataprocessing-anddimensionreduction-additionalassistance" id="qa-qc-dataprocessing-anddimensionreduction-additionalassistance"></a>

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.

![](https://documentation.partek.com/download/resources/com.adaptavist.confluence.rate:rate/resources/themes/v2/gfx/loading_mini.gif)


# Cell typing

Now that we have clustered our data we can start identifying cell populations present in our tissue. Using the results of our dimension reduction analysis, we will look at the clustering results on the UMAP and the spatial map:

* Double click the Spatial report node. This will automatically open a new Data viewer session and plot the spatial map with the high resolution image in the background.
* Click on *New plot > 3D Scatter plot* and select the UMAP node.

<figure><img src="/files/YOS179bsXELaRvrFxmB8" alt=""><figcaption></figcaption></figure>

Add the biomarkers table from the graph-based clustering task.

* Click on *New plot > Table* and select the *Biomarkers* table generated after the *Graph-based clustering* task

Change the color style to the graph-based clusters.

* Click anywhere on the Spatial plot, then click *Style,* click on the node selector next to the *Color by* drop-down, select the *Graph-based clustering* node
* Select 'Graph based' from the drop-down options

<figure><img src="/files/gBbHPyeiyHjxF1WjVzu8" alt=""><figcaption></figcaption></figure>

Copy these settings on the UMAP.

* Click on the legend title, drag and drop it on the color option on the UMAP plot:

<figure><img src="/files/PewG76ZakBt3U3SoFGfg" alt=""><figcaption></figcaption></figure>

Now that we have plotted the clustering results we can start exploring them. We have identified markers for 13 clusters, with apparent spatial separation. The sample is human front cortex, thus we can expect our clusters to represent some of the major cell types found in the tissue (e.g. astrocytes, microglia, oligodendrocytes etc). Looking at the markers for Cluster 4, we can see there are a few known astrocyte markers: AQP4, AGT, FGFR3.

* Drag and drop each one of the 3 genes from the biomarkers table onto of the 'Green, Red, Blue' features of the spatial plot to visualize the in-situ expression:

<figure><img src="/files/QQLPMbdvIVWp99pRMagK" alt=""><figcaption></figcaption></figure>

Zoom in to better observe the cell-level expression patterns.

* Use the mouse wheel control to zoom in and out on the image

<figure><img src="/files/2C1qCYdDMsExJbbLR0Jf" alt=""><figcaption></figcaption></figure>

Let's classify the cells from cluster 4 as astrocytes:

* Click *Select & Filter > Criteria,* drag the Graph-based attribute on the *Add criteria* box and select only Cluster 4.
* Click *Classify > Classify selection.* Type 'Astrocytes' and **Save.**
* Click **Apply classifications**, type 'Cell type' in the box and click **Run.**

<figure><img src="/files/0Rf1gZAuAtB9L6MsMqds" alt=""><figcaption></figcaption></figure>

### Additional Assistance <a href="#celltyping-additionalassistance" id="celltyping-additionalassistance"></a>

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.

![](https://documentation.partek.com/download/resources/com.adaptavist.confluence.rate:rate/resources/themes/v2/gfx/loading_mini.gif)


# Classify subpopulations & differential expression analysis

### Identification of A1/A2 Astrocytes <a href="#classifysubpopulations-and-differentialexpressionanalysis-identificationofa1-a2astrocytes" id="classifysubpopulations-and-differentialexpressionanalysis-identificationofa1-a2astrocytes"></a>

Astrocytes populations are often found in two activation states: the neurotoxic or pro-inflammatory phenotype (A1) and the neuroprotective or anti-inflammatory phenotype (A2). Now that we have identified our astrocyte population we can move onto the sub-classification of A1/A2 astrocytes. For this we are going to use the GFAP marker, a commonly used marker for A1 astrocytes.

* Double click the Spatial report node in the *Analyses* tab of your project to open a new data viewer session.
* Click *Select & Filter ,* select *Criteria,* and add the 'Cell type' attribute in the criteria box.
* Select only the *Astrocytes* and click ![](https://documentation.partek.com/download/thumbnails/98206089/Screenshot%202024-07-19%20at%2010.47.50.png?version=1\&modificationDate=1721382473109\&api=v2) to include only the selected points
* Now add *GFAP* to the criteria box and toggle the *Pin histogram* option
* The histogram shows the presence of two distinct populations that segregate based on *GFAP* expression
* Set the upper threshold as 5

<figure><img src="/files/wISuWD1YuwrRbvRJnoon" alt=""><figcaption></figcaption></figure>

Classify the population

* *Classify > Classify selection,* name the cells 'A2' and **Save**
* Now slide the upper threshold to the max and set the lower threshold to 5, then *Classify > Classify selection,* name the cells 'A1' and **Save**
* We are now ready to **Apply classifications,** name the attribute 'Astrocyte sub-population' and **Run**

<figure><img src="/files/5sE8XJfjGDR4NFILlZkk" alt=""><figcaption></figcaption></figure>

Having classified our sub-populations we can now use that information to identify genes and biologically processes differentially activated between the two.

* Click on the *Normalized counts* node, select *Statistics > Differential analysis > Hurdle model,* click **Next**
* Select the *Astrocyte sub-population* > **Add factors,** then click **Next**
* Drag *A1* to the Numerator and *A2* to the *Denominator* box, **Add comparison,** then click **Finish**

<figure><img src="/files/CnC4SsK8FBM8rnABsPKQ" alt=""><figcaption></figcaption></figure>

Once the differential expression analysis task has completed, you can explore the report and the subsequent analysis steps following this tutorial: [Compare expression between cell types with multiple samples](/partek-flow/tutorials/single-cell-rna-seq-analysis-multiple-samples/compare-expression-between-cell-types-with-multiple-samples)

\\

\\

### Additional Assistance <a href="#classifysubpopulations-and-differentialexpressionanalysis-additionalassistance" id="classifysubpopulations-and-differentialexpressionanalysis-additionalassistance"></a>

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# User Manual

Partek Flow software is designed specifically for the analysis needs of large genomic data. It has a simple-to-use graphical interface, context sensitive menus, powerful statistics, and interactive visualizations to easily find biological meaning in genomic data.

The information found in this manual will help you to get the most of your Partek Flow software license.

* [Interface](/partek-flow/user-manual/interface)
* [Importing Data](/partek-flow/user-manual/importing-data)
* [Task Menu](/partek-flow/user-manual/task-menu)
* [Data Viewer](/partek-flow/user-manual/data-viewer)
* [Visualizations](/partek-flow/user-manual/visualizations)
* [Pipelines](/partek-flow/user-manual/pipelines)
* [Large File Viewer](/partek-flow/user-manual/large-file-viewer)
* [Settings](/partek-flow/user-manual/settings)
* [Server Management](/partek-flow/user-manual/server-management)
* [Enterprise Features and Toolkits](/partek-flow/user-manual/enterprise-features-and-toolkits)
* [Microarray Toolkit](/partek-flow/user-manual/microarray-toolkit)
* [Glossary](/partek-flow/user-manual/glossary)
* [Partek Flow FAQs](/partek-flow/user-manual/partek-flow-faqs)
* [Getting Help](/getting-help)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Interface

* [Home Page](#home-page)
* [Home Page Icons](#home-page-icons)
* [System Options](#system-options)
* [Progress Indicator and Queue](#progress-indicator-and-queue)
* [Searching the Home Page and Project Details](#searching-the-home-page-and-project-details)

## Home Page

The Partek Flow Home page is the first page displayed upon login. It provides a quick overview of recent activities and provides access to several system options (Figure 1).

To access a project, click the blue project name. The projects can be sorted by the column title and searched by criteria using the search options on the right and the search bar above the project table. The Repository summary shows the total number of projects, samples, and cells in your Partek Flow server (Figure 1).

![Partek Flow Home page](/files/Kd7iY2iCKolE2QBrnp7R)\
\&#xNAN;*Figure 1. Partek Flow Home page*

Shown at the top of the Home page, the **New project** button ![New project button](/files/SteFvNj1VZ1xbrAorJaW) provides a quick link to create a new project in your Partek Flow server.

The **Transfer files** button ![New project button](/files/SteFvNj1VZ1xbrAorJaW) is used to transfer data to the server.

The **Optional columns** ![Optional columns button](/files/MKmOtaV5He5ghRBOgkPN) button can be used to add additional column information to the project table.

The **Search button** ![Search icon](/files/sGNjbY8T4Ig4QgIJjsLM) will search for project names and descriptions that have been typed into the search bar.

Additional project details can also be opened and closed individually using the **arrow** ![Collapse icon](/files/umBCTBMFaN50QKZc1aHh) to the left of the project name. To open additional project details for all projects at once, use the arrow ![Triangle xpand icon](/files/UcRjJKZDGJvmbl3SLHeg) to the left of the Project name column header.

The table listing all the projects can be sorted by clicking the **sort icon** ![Sort icon 2](/files/22XeiLMb8QwdhalMu8K7) to the right of the table headers. By default, the table is sorted by the date when the project was last modified.

Under the Actions column, **three vertical dots** ![Three vertical dots](/files/jrOo1iz9SbP4MaCLSY9A) will open the project actions options: **Open in new tab** ![Open in new tab icon](/files/I6O49EUyhBVvnwqeGs1K), **Export project** , and **Delete project** ![Red trash icon](/files/qvcSFrVUy4o9jglwquQm) (Figure 2).

![Main icons of the Home page](/files/DIqZfvjxhw8TRapAqhkv)\
\&#xNAN;*Figure 2. Main icons of the Home page*

## System Options

The drop-­down menu in the upper right corner of the Home page (Figure 3) displays options that are not related to a project or a task, but to the Partek Flow application as a whole. This links to the **Settings** and **Profile** and gives you the option to log out of your server.

![Accessing the system options](/files/kgnzdpNzFUjvfaahHjBo)\
\&#xNAN;*Figure 3. Accessing the system options*

## Progress Indicator and Queue

The left­-most icon ![Home page button](/files/T7YkalzFjJWzZ6CvvrIs) will bring you back to the Home page with one click.

The next icon is the progress indicator, summarizing the current status of the Partek Flow server. If no tasks are being processed, the icon is grey and static and the idle message is shown upon mouse over. Clicking on this icon will direct you to the **System resources** under **Settings** (Figure 4).

![Progress indicator showing no tasks in progress](/files/j7yhgXlbZWGZBvBTITjl)\
\&#xNAN;*Figure 4. Progress indicator showing no tasks in progress*

If the server is running, the progress indicator will depict green bars animated on the server icon (Figure 5).

![Progress indicator showing Partek Flow server is active](/files/KxfIYqh4KWkFOw4qLdwn)\
\&#xNAN;*Figure 5. Progress indicator showing Partek Flow server is active*

Selecting the **Queue** drop­down will list the number of running tasks launched by the user. Additional information about the queue including the estimated completion time as well the total number of queued tasks launched can be obtained by selecting **View queued tasks** (Figure 6).

![Viewing queued tasks](/files/Zmg9aak5VjOxnIrZ0zBD)\
\&#xNAN;*Figure 6. Viewing queued tasks*

To view all previous tasks, select the **View recent activity** link. Clicking this link loads the Activity log page (Figure 7). It displays all the tasks within the projects accessible to the user (either as the owner or as a collaborator), including the tasks launched by other users of the Partek Flow instance.

![Activity log pages shows all the tasks within the projects accessible to the user](/files/EWbYL2VKwDjqK7WTfvMc)\
\&#xNAN;*Figure 7. Activity log pages shows all the tasks within the projects accessible to the user*

The Display radio buttons enable filtering of the log. **All activity** will show all the tasks (irrespective of the task owner, i.e. the user starting the task), while **My activity** lists only the tasks started by the current user. In contrast to the latter, **Collaborator’s activity** displays the tasks that are not owned by the current user (but to which the user has access as a collaborator).

The Activity log page also contains a search function that can help find a particular task (Figure 8). Search can be performed through the entire log (**All columns**), or narrowed down to one of the columns (using the **drop­-down** list).

![A search term entered in the search box](/files/L9R2JeeH78BTgeB8RYeX)\
\&#xNAN;*Figure 8. A search term entered in the search box*

## Searching the Home Page and Project Details

The Home page lists the most recent projects that have been performed on the server. By default, the table contains all the projects owned by the current user or ones where the user is a collaborator. The list entries are links that automatically load the selected project.

The **Search** box can be used to find specific projects based on project titles and descriptions. You can also **Search by** individual or multiple criteria based on projects **Owners, Members, Cell count, Sample count, Cell type, Organ,** and **Organism** (Figure 9). Cell type, Organ, Organism, cell count, and sample count are sourced from the metadata tab. Owner and members are sourced from the project settings tab.

![Filtering the projects](/files/WkKt7nUjjdcU1L5bCVZa)\
\&#xNAN;*Figure 9. Filtering the projects*

The criteria used for the search are listed above the table along with the number of projects containing the criteria.

The project table displays optional columns as the project name, owner, your role, the members that have access to the project, date last modified, size, number of samples, and number of cells. The drop down contains a thumbnail, short description, samples contained within the project and queue status (Figure 10).

The project settings tab within each project can be used to modify details such as the project name, thumbnail, description, species, members, and owner.

![The project table details](/files/Qt6jZTM3LBJnI90RZGzb)\
\&#xNAN;*Figure 10. The project table details*

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing Data

Partek Flow can import a wide variety of data types including raw counts, matrices, microarray, variant call files as well as unaligned and aligned NGS data.

* [Navigating the file browser to transfer files to the server](#navigating-the-file-browser-to-transfer-files-to-the-server)
* [Associate fastq files for multi-omic data](#associate-fastq-files-for-multi-omic-data)
* [SFTP File Transfer Instructions](/partek-flow/user-manual/importing-data/sftp-file-transfer-instructions)
* [Sample Table from a Text File](https://help.partek.illumina.com/partek-flow/quick-start-guide#the-metadata-tab)
* [Import single cell data](/partek-flow/user-manual/importing-data/import-single-cell-data)
* [Importing 10x Genomics Matrix Files](/partek-flow/user-manual/importing-data/importing-10x-genomics-matrix-files)
* [Importing and Demultiplexing Illumina BCL Files](/partek-flow/user-manual/importing-data/importing-and-demultiplexing-illumina-bcl-files)
* [Partek Flow Uploader for Ion Torrent](/partek-flow/user-manual/importing-data/partek-flow-uploader-for-ion-torrent)
* [Importing 10x Genomics .bcl Files](/partek-flow/user-manual/importing-data/importing-10x-genomics-bcl-files)
* [Import a GEO / ENA project](https://github.com/illumina-swi/partek-docs/blob/main/docs/partek-flow/user-manual/importing-data/import-a-geo--ena-project.md)

The following file types are valid and will be recognized by the Partek Flow file browser.

* bam
* bcf
* bcl
* bgx
* bpm
* cbcl
* CEL
* csv
* fa
* fasta
* fastq
* fcs
* fna
* fq
* gz
* h5ad
* h5 matrix
* idat
* loom
* mtx
* probe\_tab
* qual
* raw
* rds
* sam
* sff
* sra
* tar
* tsv
* txt
* vcf
* zip

In cases where paired end fastq data is present, files will also be automatically recognized and their paired relationship will be maintained throughout the analysis.

Matching on paired end files is based on file names: every character in both file names must match, except for the section that determines whether a file is the first or the second file. For instance, if the first file contains "\_R1", "\_1", "\_F3", "\_F5" in the file name, the second file must contain something in the lines with the following: "\_R2", "\_2", "\_F5", "\_F5-P2", "\_F5-BC", "\_R3", "\_R5" etc. The identifying section must be separated from the rest of the filename with underscores or dots. If two conflicting identifiers are present then the file is treated as single end. For example, s\_1\_1 matches s\_1\_2, as described above. However, s\_2\_1 does not mate with s\_1\_2 and the files will be treated as two single-end files.

Apart from paired-end data, files with conventional filename suffixes that indicate that they belong to the same sample are consolidated. These suffixes include:

* Adapter sequences
  * "*bbbbbb" followed by "*" or at the end of the file name, where each "b" is "A", "C", "G", or "T"
* Lane numbers
  * "*L###" followed by "*" or at the end of the file name, where each "#" is a digit 0 to 9
* Dates
  * in the form "####-##-##" preceded or followed by a period or underscore
* Set number
  * of the form "\_###" from the end

## Navigating the file browser to transfer files to the server

The file browser is used to transfer files to the server so that these files can be added to a project for analysis. If you are importing a Bioproject from GEO/ENA or using URLs for data import, there is no need to transfer the files to the server.

To access the file browser and upload data to the server, use any of these options:

* access **Transfer files** ![Transfer files no fill](/files/8HWgFupSnHzbuZhwzBLC) on the Partek® Flow® homepage
* within a project, after selecting the file type to transfer, using the **transfer files link** ![Transfer files to the server](/files/XzofuvXdtMtN4n4Dk88U) available within all file import options
* from the settings, go to **Access management > Transfer files**

Using the file browser to transfer files to the server:

* Click **Transfer files** ![Transfer files with fill](/files/t6Is2KK601hLtI8yJxQP) to access the file browser
* Drag and drop or click **My Device** ![My device icon](/files/pnMFpvivknhRKeNdiNTC) to add files from your machine
* Click **Browse** ![Search icon](/files/sGNjbY8T4Ig4QgIJjsLM) to modify the Upload directory or create a new folder. The **Upload directory** ![Upload directory field](/files/l8Mnc8aNUVWyjohaCb8k) should be specified, known, and distinguishable for project file management. You will return to this directory and access the files to import them into a project
* To continue to add more files use **+ Add more** ![Add more button](/files/e3zf7JxqJVHvs3lQEJKv) in the top right corner. To cancel the process select **Cancel** in the top left corner
* Click **Upload** ![Upload 1 file button](/files/19Ang9DXgSjgVxSBl4TP) to complete the file upload
* Do not exit the browser tab or let the computer go to sleep or shut down until the transfer has completed

File size displayed in the table is binary format, not decimal format (e.g. GB displayed in the table is gigibyte not gigabyte. 1 gigibyte is 1,073,741,824 bytes. 1 gigabyte is 1,000,000,000 bytes. 1 gigibyte is 1.074 gigabytes).

## Associate fastq files for multi-omic data

There are projects with more than one file type, such as single cell multi-omic assays that generate protein and RNA together. In these cases, files need to be associated with each other if starting in fastq format. If we start with processed data, there is no need to associate these files.

* Define the type of data the file represents when importing files into the project.

![Select data type](/files/NCtlRVz8ZaT4TKAPLlCy)

If this step is skipped, the data type can be changed after import by right clicking the data node.

![Change data type](/files/ArGoi3UsnHaDC7y9hhEh)

* After importing both types of data, associate fastqs with the already imported data. (e.g. associate RNA fastqs with ATAC fastqs already imported into Partek Flow.

![Associate files with this sample](/files/uLGpRwzMmtmfB512e3cK)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# SFTP File Transfer Instructions

* [Introduction](#introduction)
* [SFTP with WinSCP](#sftp-with-winscp)
* [SFTP with FileZilla](#sftp-with-filezilla)
* [SFTP command line usage](#sftp-command-line-usage)
* [Points of Caution](#points-of-caution)

Only use this method if encounter issues using [Transfer file](/partek-flow/quick-start-guide/getting-started-with-your-partek-flow-hosted-trial) function on Flow homepage, contact <support@partek.com> to obtain private key.

## Introduction

The following instructions detail the use of SFTP (Secured File Transfer Protocol) to transfer data to and from your Partek Flow instance. SFTP offers significant performance and security enhancements over FTP for file transfers. It also enables the use of robust file syncing utilities, e.g. RSYNC, and is compatible with common file transfer programs such as FileZilla and WinSCP.

To transfer files with SFTP, you will need to have your Partek Flow:

* Server Name. Example: myname.partek.com
* Username. Example: flowloginname
* Private authentication key

This information should have been e-mailed to you from the Partek licensing team. If you lose this information, contact Partek support and we will resend your authentication key to you.

## SFTP with WinSCP

WinSCP is an open source, free SFTP client for Windows. Its main function is file transfer between a local and a remote computer.

### Downloading WinSCP

To download WinSCP, visit WinSCP's official site: <https://winscp.net/eng/download.php> On the WinSCP page you may need to scroll a bit down, to reach the green button Download WinSCP.

![Download button on WinSCP's official web site](/files/zHOI6wvobODg5lORXFl8)\
\&#xNAN;*Figure 1. Download button on WinSCP's official web site. Note: the version may change from the time of writing of this document*

### Connecting to your Partek server with WinSCP

Download and install WinSCP on your local computer and then launch the program.

* On the Login page click on the **New Site** icon.

![Adding NewSite on WinSCP's Login page](/files/7a1nM8stzOpiFXAKFcOg)\
\&#xNAN;*Figure 2. Adding NewSite on WinSCP's Login page*

* Type in the Host Name, which is the same as the web address that you use to access your instance of Partek Flow

The web address for your instance of Partek Flow has been sent to you by Partek's Licensing team. In this example, the web address is ilukic5i.partek.com.

* Type in the User name, that has also been sent to you (and is the same user name that you use to log on to Partek Flow). In this example, the web address is lukici.

![Adding Host name and User name information. Use the host and user name that has been sent to you by Partek's licensing team](/files/bZOmYbKZJcCtHOfOjlb9)\
\&#xNAN;*Figure 3. Adding Host name and User name information. Use the host and user name that has been sent to you by Partek's licensing team*

* To proceed click on the **Advanced...** button
* Then select the **Authentication** in the **SSH** section of the Advanced Site Settings dialog

![Adding the id\_rsa file to WinSCP. Use the Advanced Site Settings tab and select Authentication](/files/UeczEMOjcO4jsDSm2uXs)\
\&#xNAN;*Figure 4. Adding the id\_rsa file to WinSCP. Use the Advanced Site Settings tab and select Authentication*

Select the ... button (under Private key file) to browse for the id\_rsa file.

* The file has been sent to you by Partek's licensing team attached to the same email that gave you your URL and username.

If you do not see it in the Select private key file browser, switch to **All Files** (*.*)

![Showing all files in the Select private key file dialog](/files/ktKIRJfaLViy1dTYJtDv)\
\&#xNAN;*Figure 5. Showing all files in the Select private key file dialog*

* Click **Open**

WinSCP will ask you to confirm file format conversion

* Click **OK**.

![Converting key file format](/files/BW2zBDtHw6m94TqHx3bn)\
\&#xNAN;*Figure 6. Converting key file format*

WinSCP will create a file in .ppk format.

* Click **Save** to save the converted key file, id\_rsa.ppk, to a secure location on your local computer.
* Click **OK** again to confirm the change.

Your private key has been saved in .ppk format and added to WinSCP

Click **OK** to proceed

![Private key in .ppk format added to WinSCP](/files/8IxO5m5WH7IZmbtB2Ojy)\
\&#xNAN;*Figure 7. Private key in .ppk format added to WinSCP*

* Click **Save** to save the new WinSCP settings.

This will open the Save session as site dialog. You can accept the default name (in this example <lukici@ilukic5i.partek.com>) or add a custom name. The name that you specify here will appear in the left panel of the Login dialog.

* Once you have made your edits, click **OK**.

![Customising the name for the new site on the Login dialog. In this example, the name is lukici@ilukic5i.partek.com](/files/Rcn17wxVmkEXKbrN7Rh3)\
\&#xNAN;*Figure 8. Customising the name for the new site on the Login dialog. In this example, the name is <lukici@ilukic5i.partek.com>*

* On the Login page, select your newly created site (in this example: <lukici@ilukic5i.partek.com>) and click the **Login** button.

The first time you connect, a warning message will appear, asking you whether you want to connect to an unknown server.

* Click **Yes** to proceed.

![The first time you connect to your Partek server, WinSCP will present a warning message. Click Yes to connect to the server](/files/9vlpPL6r2lR3gb79ZA0E)\
\&#xNAN;*Figure 9. The first time you connect to your Partek server, WinSCP will present a warning message. Click Yes to connect to the server*

The progress towards establishing a connection will be displayed in a dialog. This process is automatic and you do not need to do anything.

![Progress of the connection to your Partek server will be displayed on the screen](/files/rcUBiiCmBPNZeGB7k8pE)\
\&#xNAN;*Figure 10. Progress of the connection to your Partek server will be displayed on the screen*

The WinSCP interface includes is split into two panels. The panel on the left shows the directory structure of your local computer and the panel on the right shows the directory structure of your Partek Flow file server.

![WinSCP screen after connection divides into two panels: the files one the left are on the local computer, while the files on the right are on Partek server](/files/zx5oXKGeavasIfKt3ofY)\
\&#xNAN;*Figure 11. WinSCP screen after connection divides into two panels: the files one the left are on the local computer, while the files on the right are on Partek server*

To transfer a file, just **drag and drop** the file from one panel to the other. The progress of your transfer will be shown on the screen.

![Progress of file transfer is shown on screen](/files/M3wH5bcpY8XmENKM6DoS)\
\&#xNAN;*Figure 12. Progress of file transfer is shown on screen*

## SFTP with FileZilla

FileZilla is a graphical file transfer tool that runs on Windows, OSX, and Linux. It is great when needing to do bulk transfers as all transfers are added to a queue and processed in the background. It is possible to browse your files on the Partek Flow server while transfers are active. This is also the best solution when you are not on a computer with command line access or you are uncomfortable with command line operations.

### Downloading FileZilla

We recommend downloading the FileZilla install packages from us. They are also available from download aggregator sites (e.g. CNET, download.com, sourceforge) but these sites have been known to bundle adware and other unwanted software products into the downloads they provide, so avoid them.

* Mac OSX: <http://packages.partek.com/bin/filezilla/fz-osx.app.tar.bz2>
* Windows 32-bit: <http://packages.partek.com/bin/filezilla/fz-win32.exe>
* Windows 64-bit: <http://packages.partek.com/bin/filezilla/fz-win64.exe>
* Linux (Please use your distribution's package manager to install Filezilla):
* Ubuntu:

```
$ sudo apt-get update
```

```
$ sudo apt-get install filezilla
```

* RedHat, see the following guide: <http://juventusitprofessional.blogspot.com/2013/09/linux-install-filezilla-on-centos-or.html>
* OpenSuse, see: <https://software.opensuse.org/package/filezilla>

### Connecting to your Partek server with FileZilla

After starting FileZilla, click on the Site Manager icon located at the top left corner of the FileZilla window.

![Figure 13](/files/muV9kUXuoX1KEYI27xct)\
\&#xNAN;*Figure 13*

Click on the New Site button on the left of the popup dialog.

![Figure 14](/files/H5FoyCxaOltnb77LOhS0)\
\&#xNAN;*Figure 14*

Type in a name for the connection. Example: “Partek SFTP”.

![Figure 15](/files/KaBuRNsRls8IAIDZJUZR)\
\&#xNAN;*Figure 15*

The connection details to the right need to be changed to reflect the information you received via email. The default settings will NOT work.

![Figure 16](/files/reYK2VW9n3U5EpuPrMrR)\
\&#xNAN;*Figure 16*

* Set Host: to your partek server name
* Leave Port: blank
* Change Protocol: to SFTP - SSH File Transfer Protocol
* Change User: to your Partek Flow login name
* Change Logon Type: to Key File and select the key file received via email.

![Figure 17](/files/56L2mxtcc4JGVCx9sS3P)\
\&#xNAN;*Figure 17*

When selecting your key file, change the file selection from its default of PPK files to All files. Otherwise you key file will not be visible in the file browser.

After selecting your key file, click the Connect button.

![Figure 18](/files/oIZKcsFhd8xxWpdrlB7D)\
\&#xNAN;*Figure 18*

Click the checkbox to always trust this host and click OK. Once connected, you can begin to browse and transfer files. The files and folders to the left are on your computer, the ones on the right are on the Flow server.

![Figure 19](/files/HbiEqPrh5n8ASCi8otzI)\
\&#xNAN;*Figure 19*

## SFTP command line usage

### Importing your private authentication key

You will receive a file called id\_rsa via email. Download this file, note where you downloaded it to, then use ssh-add to import the key. If you logout or reboot your computer, you will need to re-run the commands below. After key import, you will not be asked a password when transferring files to your Partek Flow server.

```
$ cd directory/with/key
```

```
$ chmod 600 id_rsa
```

```
$ eval $(ssh-agent)
```

```
$ ssh-add id_rsa
```

### Copying files and folders between your Partek Flow server and local computer

#### RSYNC usage

RSYNC is useful when resuming a failed transfer. Instead of re-uploading or downloading what has already been transferred, RSYNC will copy only what it needs.

The command below will sync the folder "local\_folder" with the "remote\_folder" on Partek's servers. To transfer in the other direction, reverse the last two parameters.

```
$ rsync -avr --progress ./local_folder/ flowloginname@myname.partek.com:~/remote_folder/
```

With rsync, don't forget the trailing '/' on directory names.

Before moving the files, we strongly advise you to use FileZilla to explore the directory structure of the Partek server and then create a new directory to transfer the files to.

## Points of Caution

* When you delete files from the Partek Flow server they are gone and can not be recovered.
* Please use Partek Flow to delete projects and results. Manually removing data using SFTP could break your server.
* Wait until ALL input data for a particular project has been transferred to the Partek Flow server before importing data via Partek Flow. If you try to import samples while the upload is occurring the import job will crash.
* When upload raw data to Partek hosted Flow server, we recommend to a create subfolder for each experiment at the same level of "FlowData" folder or inside "FlowData" folder.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Import single cell data

* [Import single cell data for different assay types and formats](#import-single-cell-data-for-different-assay-types-and-formats)
* [Import single cell data from count matrix text file(s) using Full count matrix as the data format](#import-single-cell-data-from-count-matrix-text-file-s-using-full-count-matrix-as-the-data-format)

## Import single cell data for different assay types and formats

Select **Single cell**, choose the **assay** type (scRNA-Seq, Spatial transcriptomics, scATAC-Seq, V(D)J, Flow/Mass cytometry), and select the data **format** (Figure 1). Use the **Next** button to proceed with import.

![Figure 1. Choose import single cell data option](/files/MkeQhwPGwLan7gXx11k9)

## Import single cell data from count matrix text file(s) using Full count matrix as the data format

Partek Flow supports single cell data analysis in count matrix text format using the Full count matrix data format (Figure 2). Each matrix text file is assumed to represent on sample, each value in the matrix represents expression value of a feature (e.g. a gene, or a transcript) in a cell. The expression value can be raw count, or normalized count. The requirement of the format of each text file should be the same as [count matrix data](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab#importing-count-matrix-data).

Specify text file location, only one text file (in other words one sample) can be imported at once, preview of the file will be displayed, configuration of the file format is the same as [Import count matrix data](/partek-flow/tutorials/creating-and-analyzing-a-project/the-metadata-tab#importing-count-matrix-data). In addition, you need to specify the details about this file.

![Figure 2. Choose Full count matrix as the data format](/files/4VcqIBWl4rS2cO5t4gw0)

Click **Finish**, the sample will be imported, on the data tab, number of cells in the sample will be displayed.

To import multiple samples, repeat the above steps by clicking **Import data** on the *Metadata* tab or within the task menu (toolbox) on the *Analyses* tab. Make the same previous selections using the cascading menu.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing 10x Genomics Matrix Files

## Importing 10x Genomics Matrix Files

* [Importing single cell data](#importing-single-cell-data)
  * [Importing matrices into Partek Flow (this Market Exchange Format is popular for public repositories)](#importing-matrices-into-partek-flow-this-market-exchange-format-is-popular-for-public-repositories)
  * [Importing matrices in h5 format (this Hierarchical Data Format is recommended for multiple samples)](#importing-matrices-in-h5-format-this-hierarchical-data-format-is-recommended-for-multiple-samples)
* [Importing spatial data](#importing-spatial-data)
  * [Importing Xenium Output Bundle](#importing-xenium-output-bundle)

## Importing single cell data

Partek Flow supports the import of [filtered gene-barcode matrices](https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/output/matrices) generated by 10x Genomics' [Cell Ranger pipeline](https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/what-is-cell-ranger).

Below is a video summarizing the import of these files:

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2FDlVdJ0PVbUsCsosCvuFH%2F10X_matrix_import.mp4?alt=media&token=b9f81a28-ce69-4af9-8d86-208439a72134>" %}

### Importing matrices into Partek Flow (this Market Exchange Format is popular for public repositories)

To import the matrices into Partek Flow, create a new project and click **Add data** then select **Import scRNA count feature-barcode-mtx** under **Single cell > scRNA-Seq**.

![Figure 1. Importing single cell data](/files/4ieGKBzRJm478DZ4vD2I)

Samples can be added using the **Add sample** button. Each sample should be given a name and three files should be uploaded per sample using the **Browse** button.

![Figure 2. Import three feature-barcode matrix files for each sample](/files/0kZOIU2Ocb95QbXMddn6)

If you have not already, transfer the files to the server to be accessed when you click **Browse**. Follow the directions [here](/partek-flow/user-manual/importing-data) to add files to the server. Make sure the files are **decompressed** before they are uploaded to the server.

By default, the Cell Ranger pipeline output will have a folder called filtered\_gene\_bc\_matrices (Figure 3). It is helpful to rename and organize the files prior to transfer using the File browser.

There are folders nested within the matrix folder, typically representing the reference genome it was aligned to. Navigate to the lowest subfolder, this should contain three files:

* barcodes.tsv
* genes.tsv
* matrix.mtx

**Select all 3 files** for import into Partek Flow

![](/files/21VTe9MTQivtYvaCAJp1)

![Figure 3. Filtered matrix folder from Cell Ranger pipeline](/files/RyUPnbwQ9q168y7xfL4z)

Specify the annotation file used when running the pipeline for additional information such as mitochondrial counts (Figure 4). Other information can also be specified, such as the count value format. All features can be reported or features with non-zero values across all samples can be reported and the read count threshold can be modified to make the import more efficient.

![Figure 4. Configuring matrix metadata](/files/qW1mScBvAV9D7pKniWon)

Click **Finish** when you have completed configuration. This will queue the import task.

### Importing matrices in h5 format (this Hierarchical Data Format is recommended for multiple samples)

The Cell Ranger pipeline can also generate the same filtered gene barcode matrix in h5 format. This gives you the ability to select just one file per matrix and select multiple matrices to import in batch. To import an h5 matrix, select the **Import scRNA full count matrix or h5** option (Figure 1). **Browse** for the files and modify any configuration options. Remember the files need to be transferred to the server.

![Figure 5. Importing matrix in h5 format](/files/Hmz6IArQSWAkPGjgp5oR)

This feature is also useful for importing multiple samples in batch. Simply put all h5 files from your experiment on a single folder, navigate to the folder and select all the matrices you would like to import.

Configure all the relevant sample metadata, including sample name and the annotation that was used to generate the matrices, and click **Finish** when completed. Note that all matrices *must have been generated using the same reference genome and annotation* to be imported into the same project.

## Importing spatial data

### Importing Xenium Output Bundle

Raw output data generated by the 10x Genomics' Xenium Onboard Analysis pipeline consists of decoded transcript counts and morphology images. The raw output and other standard output files derived from them are compiled into a zipped file called Xenium Output Bundle.

To import the Xenium Output Bundle into Partek Flow, create a new project and click **Add data**, then select **Import 10x Genomics Xenium** under **Single cell > Spatial**, click **Next.**

![Figure 6. Importing Xenium spatial data](/files/FdJFXjK4X5xvVRZP9Wjw)

Samples can be added using the **Add sample** button. Each sample should be given a name and a folder containing the required 6 files: cell\_feature\_matrix.h5, cells.csv.gz, cell\_boundaries.csv.gz, nucleus\_boundaries.csv.gz, transcripts.csv.gz, morphology\_focus.ome.tif should be uploaded per sample using the **Browse** button. The required 6 files should be all included in the Xenium Output Bundle folder.

![Figure 7. Import Xenium Output Bundle folder for each sample](/files/LNix4sYlO7i1vnghLDFR)

If you have not already, transfer the files to the server to be accessed when you click **Browse**. Follow the directions [here](/partek-flow/user-manual/importing-data#nagivating-the-file-browser-to-transfer-files-to-the-server) to add files to the server. You will need to **decompress** the Xenium Output Bundle zip file before they are uploaded to the server. After decompression, you can **drag and drop** the entire folder into the Transfer files dialog, all individual files in the folder will be listed in the Transfer files dialog after drag & drop, with no folder structure. The folder structure will be restored after upload is completed.

![Figure 8. Drag & drop unzipped Xenium Output Bundle folder into Transfer files dialog](/files/KwyrUQRco5e7lUNhivmb)

Once you have uploaded the folder into the server, you can continue to select the folder for each sample from **Browse**. Once the folder is selected, the *Cells* and *Features* values will auto-populate. You can choose an annotation file that matches what was used to generate the feature count. Then, click **Finish** to start importing the data into your project.

![Figure 9. Add Xenium Output Bundle and select annotation](/files/ipC4qxGq1U7lPbnWZNBh)

### Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing and Demultiplexing Illumina BCL Files

Primary sequencing output of an Illumina sequencer are per-cycle base call (bcl) files, which first need to be converted to fastq format, so that the data can be pushed to downstream applications. Partek Flow software comes with a conversion tool that can be used to import data in the bcl file format . In addition to the file conversion, this tool also demultiplexes the bcl files in the same step and outputs demultiplexed fastq files as the result.

We recommend you start by transferring the entire Illumina run folder to the Partek Flow server. To start a new project with bcl files, first select **bcl** under the **Other** import tab (Figure 1)

![Figure 1. Bcl file import setup dialog. Required input includes: RunInfo.xml file, SampleSheet.csv file, and a directory hosting .bcl files](/files/d1ZI6qdiCdjw0X7OxCbt)

The resulting window shows the configuration dialog (Figure 2).

![Figure 2. Bcl file import setup dialog. Required input includes: RunInfo.xml file, SampleSheet.csv file, and a directory hosting .bcl files](/files/13KGLErBP4vteXKyaCjz)

The bcl files hold the base calls and are in the *Data directory* within the whole Illumina run folder. Note that the *Data directory* file path needs to point to the directory, not to an individual bcl file.

The RunInfo.xml file is generated by the primary analysis software and contains information on the run, flow cell, instrument, time stamp, and the read structure (number of reads, number of cycles per read, whether a read is an index read). This file is typically stored at the top level in the Illumina run folder.

The SampleSheet.csv file provides the information on the relationship between the samples and indices specified during library creation. Although it has four sections, two sections (Settings and Data) are important for the data import and conversion. For more information on the files, consult Illumina documentation.

Selecting the **Configure** option under the *Advanced options* section enables a granular control of the import (Figure 3).

![Figure 3. Advanced options of bcl importer](/files/mTAIVfGC9AAkQFOe4Q4a)

The *Select tiles* option (--tiles) enables the user to process only a subset of tiles available in the flow cell. The input for this option is a comma-separated list of regular expressions.

*Min trimmed read length* (--minimum-trimmed-read-length) specifies the minimum read length after adapter removal.

*Mask short adapter reads* (--mask-short-adapter-reads) applies when a read is trimmed below the length specified by *Min trimmed read length*. If the number of bases after adapter removal is less than *Min trimmed read length*, it forces the read length to be equal to *Min trimmed read length* by replacing the adapter bases that fall below the specified length by Ns. If the number of remaining bases falls below *Mask short adapter sequences*, then it replaces all the bases in a read with Ns.

*Adapter stringency* (--adapter-stringency) specifies the minimum match rate that triggers the masking or trimming of adapters. The rate is calculated as MatchCount / (MatchCount + MismatchCount). Only the reads exceeding the specified rate of sequence identity with adapters are trimmed.

*Barcode mismatches* (--barcode-mismatches) controls the number of allowed mismatches per index sequence.

*Use bases mask* (--use-bases-mask) defines a custom read structure that may be different to the structure specified in the RunInfo.xml file. The input for this option is a comma-separated list where Y and I are followed by a number indicating how many sequencing cycles to include in the fastq file. For example, if the option is set to Y26,I8,Y98, 26 cycles (26bp) will be used to generate the R1 sequence, 8 cycles (8bp) will be used for the sample index, and 98 cycles (98bp) will be used to generate the R2 sequence.

*Do not split files by lane* (--no-lane-splitting) prevents splitting of fastq files by lane, i.e. the converter will merge multiple lanes and generate one fastq file per sample.

*Create fastq for index reads* (--create-fastq-for-index-reads) creates an extra fastq file for each sample containing the sample index sequence for each read. This will be imported as an extra sample into the project.

*Ignore missing bcls* (--ignore-missing-bcl) will interpret missing base call files as N.

*Ignore missing filter* (--ignore-missing-filter) will ignore missing filter files and assume all clusters pass the filter.

*Ignore missing positions* (--ignore-missing-positions) will write new, unique coordinates into the header line if the cluster location files are missing.

*Ignore missing controls* (--ignore-missing-control) will interpret missing control files as missing not-set control bits.

*Save undetermined* *fastq* will take the reads that could not be assigned to a sample index and collect them into an Undetermined\_S0.fastq file, which will be imported as a new sample.

The result of the import is an *Unaligned reads* data node, containing demultiplexed fastq files.

For more information about the BCL to FASTQ conversion tool, including information on the proper folder structure and instructions for formatting the SampleSheet.csv file, please consult the [bcl2fastq2 Conversion Software Guide](https://emea.support.illumina.com/content/dam/illumina-support/documents/documentation/software_documentation/bcl2fastq/bcl2fastq2-v2-20-software-guide-15051736-03.pdf).

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Partek Flow Uploader for Ion Torrent

The Partek Flow Uploader is a Torrent Browser plugin that lets users upload run results to Partek Flow for further analysis.

* [Quick Video Tutorial on running the Plugin](#PartekFlowUploaderforIonTorrent-QuickVideoTutorialonrunningthePlugin)
* [Downloading the Plugin](#PartekFlowUploaderforIonTorrent-DownloadingthePlugin)
* [Adding the Plugin to your Run Plan](#PartekFlowUploaderforIonTorrent-AddingthePlugintoyourRunPlan)
* [Running the Plugin from a Report](#PartekFlowUploaderforIonTorrent-RunningthePluginfromaReport)
* [Sample table created by the Plugin](#PartekFlowUploaderforIonTorrent-SampletablecreatedbythePlugin)
* [Conversion of UBAM to FASTQ files](#PartekFlowUploaderforIonTorrent-ConversionofUBAMtoFASTQfiles)

## Quick Video Tutorial on running the Plugin

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2FbdEPeXiW4z9dg7ZeuF2D%2FTSPluginv102.mp4?alt=media&token=bea0ff88-982e-492b-a4a7-8399b3fdfbe4>" %}

The clip above (video only, no audio) shows the Partek Flow Uploader plugin in action.

## Downloading the Plugin

Download the Partek® Flow® Uploader from the links below:

<https://customer.partek.com/plugins/PFU/PartekFlowUploader-1.03.zip>

This is a compressed zipped file. Do not unzip.

Installation of the Plugin

Installation only needs to be performed once per Torrent Browser. All users of the same instance of Torrent Browser will be able to use the plugin. For future versions of the plugin, the steps below can also be used for updating.

To install the plugin, first log into Torrent Browser (Figure 1).

![Figure 1. Torrent Browser login page](/files/2PX3AGFHseVunadF339B)

Navigate to **Plugins** under dropdown menu in upper-right corner under the gear icon (Figure 2).

![Figure 2. Accessing the Plugins](/files/UgPTqdK4iIhtvWvUrpl8)

Click the **Install or Upgrade Plugin** button (Figure 3).

![Figure 3. Installing a new plugin in the Torrent Browser](/files/tPVGZDb4RHCzsvjTAAcp)

Click **Select File** (Figure 4) and use the file browser to select the zip file you downloaded from the [download link](#downloading-the-plugin). Click **Upload and Install**.

![Figure 4. Uploading and installing the zip file of the plugin](/files/pWH7IDOhiVO1peQAm4cl)

Verify that the Partek Flow Uploader is listed and that the **Enabled** checkbox is selected (Figure 5).

![Figure 5. Table showing the Partek Flow Uploader successfully installed](/files/hADUMC1u9oSnNinOnA2U)

From the plugins table, click the (Manage) gear icon for the Partek Flow Uploader and select **Configure** in the drop-down menu (Figure 6).

![Figure 6. Accessing the Global configuration of the Partek Flow Uploader](/files/xs8GJVZbe7Tdbj1z0SkY)

Global Partek Flow configuration settings can be entered into the plugin. When set, it will serve as the default for all users of the Torrent Browser. *If multiple Partek Flow users are expected to run the plugin, it is recommended to leave the username and password fields blank so that individual users can enter them as needed.*

In the configuration dialog (Figure 7), enter the Partek Flow URL, your username and your password. Clicking on **Check configuration** would verify your credentials and indicate if a valid username and password has been entered. Click **Save** when done.

![Figure 7. Global Partek Flow Uploader configuration settings](/files/WFsY32DTHs389O3HwCm1)

Click the **Rescan Plugins for Changes** button. Rescanning the plugins will finish the installation and save the configuration.

## Adding the Plugin to your Run Plan

In the Torrent Browser, you can configure a Run Plan to include the Partek Flow Uploader. You can create a new Run Plan (from a Sample or a Template) or edit an existing Run Plan. In the example in Figure 8, the Partek Flow Uploader will be included in an existing Run Plan. From the Planned Runs page, click the gear icon in the last column, and choose **Edit.**

![Figure 8. Selecting an existing Run Plan to Edit](/files/ZFT5LB14BwmzGDCPsIcH)

In the Edit Plan page, go to the **Plugins** tab (Figure 9) and select the checkbox next to the **PartekFlowUploader**.

![Figure 9. Editing the Plugins section of the Run Plan](/files/NOZqTvcP4J82typfdaUt)

Click the **Configure** hyperlink next to the **PartekFlowUploader** (Figure 10). If necessary, enter the Partek Flow URL, your username and your password. These are the same credentials you use to access Partek Flow directly on a web browser. Note that some fields may already be pre-populated depending on the global plugin configuration, you can edit the entries as needed. *All fields are required to successfully run the plugin.*

The **Project Name** field will be used in Partek Flow to create a new project where the run results will be exported. However, if a project with that name already exists, the samples will be added to that existing project. This enables you to combine multiple runs into one project. Project Names are limited to 30 characters. If not specified, the plugin will use the Run Name as the Project Name.\
Click the **Check configuration** button to see if you typed a valid username and password. When ready, click **Save Changes** to proceed.

![Figure 10. Configuring the Partek Flow Uploader as part of a Run Plan](/files/x4cg9gCj0qqkySeY88dj)

Proceed with your Run Plan. The plugin will wait for the base calling to be finished before exporting the data to Partek Flow.

Once the Run Plan is executed, data will be automatically exported to the Partek Flow Server. In the Run Report, go to the **Plugin Summary** tab and the plugin status will be displayed. An example of a successful Plugin upload is shown in Figure 11.

![Figure 11. Partek Flow Uploader showing successful transfer](/files/8LOClaFzJXmbM75IHrzB)

To access the project, click on the **Partek Flow** hyperlink in the plugin results (Figure 11). You can also go directly to Partek Flow in a new browser window and access your account. In your Partek Flow homepage (Figure 12), you will now see the project created by the Partek Flow Uploader.

![Figure 12. Partek Flow Homepage with the new Project created by Partek Flow Uploader](/files/VfMauW7YnpWvBeNke0CK)

## Running the Plugin from a Report

You can manually invoke the plugin from a completed run report. This allows you to export the data from the Torrent Server if you did not include the plugin in the original run plan. This also gives you the flexibility to export the same run results onto different project(s). Open the run report and scroll down to the bottom of the page (Figure 13). In the **Plugin Summary** tab, click the **Select Plugins to Run** button.

![Figure 13. Running the Plugin from a completed run](/files/GJ5ZUQ4A3CwqSLXHv1Di)

From the plugin list (Figure 14), select the **PartekFlowUploader** plugin.

![Figure 14. Selecting the Partek Flow Uploader Plugin](/files/gco80Isqu5NA8L1B5mqO)

Configure the Partek Flow Uploader. Enter the Partek Flow URL, your username and your password (Figure 15). These are the same credentials you use to access Partek Flow directly on a web browser. Although some fields may already be pre-populated depending on the global plugin configuration, you can edit the entries as needed. *All fields are required to successfully run the plugin.*

![Figure 15. Configuring the Partek Flow Uploader from a Report](/files/OumH5JwcsDrak6brHS7p)

The **Project Name** field will be used in Partek Flow to create a new project where the run results will be exported. However, if a project with that name already exists, the samples will be added to that existing project. This enables you to combine multiple runs into one project. Project Names are limited to 30 characters. The default project name is the Run Name.

When ready, click **Export to Partek Flow** to proceed. If you wish to cancel, click on the X on the lower right of the dialog box.

Note that configuring the Plugin from a report (Figure 15) is very similar to configuring it as part of a Run Plan (Figure 10) with two notable differences:

1. The **Check configuration** has been replaced by **Export to Partek Flow** button, which when clicked, immediately proceeds to the export.
2. The **Save changes** button has been removed so any change in the configuration cannot be saved (compared to editing a run plan where plugin settings are saved)

Once the plugin starts running, it will indicate that it is *Queued* on the upper right corner of the Plugins Summary (Figure 16). There will also be a blue **Stop** button to cancel the operation.

![Figure 16. The plugin is Queued as indicated and a blue Stop button is available](/files/GSBfcQUVRnUByjYACilE)

Click the **Refresh plugin status** button to update. The plugin status will show *Completed* once the export is done and the data is available in Partek Flow (Figure 11).

## Sample table created by the Plugin

The Partek Flow Uploader plugin sends the unaligned bam files to the Partek Flow server. For each file, a Sample of the same name will be created in the Data tab (Figure 17).

![Figure 17. Expanded Sample Table showing Data files](/files/bFtXwE1YAk5p1p2uxxkM)

Reads that had no detectable barcodes have been combined in a sample with a prefix: *nomatch\_rawlib.basecaller.bam* (Row 17 in Figure 17). You can removed this sample from your analysis by clicking the gear icon ![gear\_icon\_settings\_gray](/files/eaP9vVsveLwkC2gmEW3n) next to the sample name and choosing **Delete sample**.

The data transferred by the Partek Flow Uploader is stored in a directory created for the Project within the user's default project output directory. For example, in Figure 17, the data for this project is stored in: **/home/flow/FlowData/jsmith/Project\_CHP Hotspot Panel.**

## Conversion of UBAM to FASTQ files

The plugin transfers the Unaligned BAM data from the Torrent Browser. The UBAM file format retains all the information of the Ion Torrent Sequencer. In the Partek Flow Project, the *Analyses* tab would show a circular data node named *Unaligned bam*. Click on the data node and the context-sensitive task menu will appear on the right (Figure 18).

Unaligned BAM files are only compatible with the TMAP aligner, which can be selected in the *Aligners* section of the Task Menu. If you wish to use other aligners, you can convert the unaligned BAM files to FASTQ using the **Convert to FASTQ** task under Pre-alignment tools. Some information specific to Ion Torrent Data (such as Flow Order) are not retained in the FASTQ format. However, those are only relevant to Ion Torrent developed tools (such as the Torrent Variant Caller) and are not relevant to any other analysis tools.

![Figure 18. Analyses tab of Partek Flow showing the task menu available for the selected Unaligned BAM file data node](/files/OpIOUzS8Eq8MmGGr1ysC)

Once converted, the reads can then be aligned using a variety of aligners compatible with FASTQ input (Figure 19). You can also perform other tasks such as Pre-alignment QAQC or run an existing pipeline. Another option is to include the *Convert to FASTQ* task in your pipeline and you can invoke the pipeline directly from an *Unaligned bam* data node.

![Figure 19. Unaligned reads in the FASTQ format are compatible with more tasks](/files/TjhJ8mPnX72wsjzgaW2G)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Importing 10x Genomics .bcl Files

Partek Flow supports .bcl files based on 10x Genomics library preparation. The following document will guide you through the steps.

To start the import, create a new project and then select **Import Data > Import bcl files**. The *Import bcl* dialog will come up (Figure 1).

![Figure 1. Import bcl dialog](/files/16vmrpLBcMVhV8wfpKHI)

Use the *Data directory* option to point to the location of the directory holding the data. It is located at the top level of the run directory and is typically labeled *Data*. Please see the tool tip for more info.

Use the *Run info file* option to point to the *RunInfo.xml* file. It is located at the top level of the run directory.

Use the *Sample sheet file* to point to the sample sheet file, which is usually a .csv file. Partek Flow can accept 10X Genomics' ["simple" and Illumina Experiment Manager (IEM) sample sheet format](https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/using/mkfastq#example_data), which utilize 10X Genomics' sample index set codes. Each index set code corresponds to a mixture of four sample index sequences per sample. Alternatively, Partek Flow will also accept a sample sheet file that has been correctly formatted using the [sample sheet generator](https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/using/bcl2fastq-direct) provided by 10X Genomics.

The click on the **Configure** link and make the following changes (Figure 2).

* *Min trimmed read length*: **8**
* *Mask short adapter reads*: **8**
* *Use bases mask*: see below
* *Create fastq for index reads*: **OFF**
* *Ignore missing bcls*: **ON**
* *Ignore missing filter*: **ON**
* *Ignore missing positions*: **ON**
* *Ignore missing controls*: **ON**

For the *Use bases mask* option, the read structure for Chromium Single cell 3' v2 prep kit is typically **Y26,I8,Y98.** The settings for Chromium Single cell 3' v3/v3.1 is typically **Y28,I8,Y91**. Please check the read structure detailed in the *RunInfo.xml* file and adjust the values to match your data.

![Figure 2. Setting the advanced options to import bcl files. The Use bases mask settings shown here are for Chromium v2 chemistry](/files/4QYnxfucqeukekmZuhyn)

Click **Apply** to accept and then **Finish** to import your files.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Import a GEO / ENA project

* [How to import a study from GEO / ENA](#how-to-import-a-study-from-geo--ena)
* [Common Issues](#common-issues)
  * [Error Message - The project did not yield any data. Double-check the project ID, or try importing the data manually](#error-message---the-project-did-not-yield-any-data-double-check-the-project-id-or-try-importing-the-data-manually)
  * [The project was imported, but the Analyses tab is empty and there are no FASTQ files](#the-project-was-imported-but-the-analyses-tab-is-empty-and-there-are-no-fastq-files)
  * [Something is missing or the import failed](#something-is-missing-or-the-import-failed)
* [FAQ](#faq)
  * [What are GEO and ENA?](#what-are-geo-and-ena)
  * [How do I know if a GEO project is also in ENA?](#how-do-i-know-if-a-geo-project-is-also-in-ena)

## How to import a study from GEO / ENA

If a project is publicly available in the Gene Expression Omnibus (GEO) and European Nucleotide Archive (ENA) databases, you can import associated FASTQ files and sample attributes automatically into Partek Flow.

* On the Homepage click **New Project** to create a project and give the project a name

![](/files/vTsW5AQA5WAOtInbsgx2)

* Click **Add data**

![](/files/iBnD3c7ZcRT3PRZNVNYF)

* Select **fastq** as the file type after choosing between **Single cell** or **Bulk** as the assay types

![](/files/DYLLMUMbNMWc6YFRgKDD)

* Click **Next**
* Choose GEO / ENA
* Enter the BioProject ID of the data set you would like to download. The format of a BioProject ID is PRJNA followed by one to six numbers (e.g. PRJNA381606)

![](/files/4lfdlI4jeS9NclKG1f3a)

A GEO ID can also be used in the format GSE followed by one to five numbers (e.g. GSE71578).

* Click **Finish**

It may take a while for the download to complete depending on the size of the data. FASTQ files are downloaded from the ENA BioProject page.

* FASTQ files will be added as an Unaligned reads data node in the Analyses tab

![](/files/VfVSimfZC9N2MxNAXwEZ)

## Common Issues

### Error Message - The project did not yield any data. Double-check the project ID, or try importing the data manually

If the study is not publicly available in both GEO and ENA, project import will not succeed.

### The project was imported, but the Analyses tab is empty and there are no FASTQ files

If there is an ENA project, but the FASTQ files are not available through ENA, the project will be created, but data will not be imported.

### Something is missing or the import failed

A variety of other issues and irregularities can cause imports to not succeed or partially succeed, including, but not limited to, a BioProject having multiple associated GSE IDs, incomplete information on the GEO or ENA page, and either the GEO or ENA project not being publicly available.

## FAQ

### What are GEO and ENA?

The Gene Expression Omnibus (GEO) and the European Nucleotide Archive (ENA) are web-accessible public repositories for genomic data and experiments. Access and learn more about their resources at their respective websites:

GEO - <https://www.ncbi.nlm.nih.gov/geo/>

ENA - <https://www.ebi.ac.uk/ena>

### How do I know if a GEO project is also in ENA?

* You can search ENA using the GEO ID (e.g., GSE71578) to check if there is a matching ENA project.

![](/files/5R1YuMTYfcARiXHX52yp)

* Open the Study result to view the BioProject ID (e.g., PRJNA381606) and a table with information about the samples and files included in the project

![](/files/7GVxVj6w3EN9WdptEnu9)

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Task Menu

The Task Menu lists all the tasks that can be performed on a specific node. It can be invoked from either a **Data** or **Task node** and appears on the right hand side of the *Analyses* tab. It is *context-sensitive*, meaning that it will only present tasks that the user can perform on the selected node. For example, selecting an *Aligned reads* data node will not present aligners as options.

Clicking a **Data node** presents a variety of tasks:

* [Data summary report](/partek-flow/user-manual/task-menu/data-summary-report)
* [QA/QC](/partek-flow/user-manual/task-menu/qa-qc)
  * [Pre-alignment QA/QC](/partek-flow/user-manual/task-menu/qa-qc/pre-alignment-qa-qc)
  * [ERCC Assessment](/partek-flow/user-manual/task-menu/qa-qc/ercc-assessment)
  * [Post-alignment QA/QC](/partek-flow/user-manual/task-menu/qa-qc/post-alignment-qa-qc)
  * [Coverage Report](/partek-flow/user-manual/task-menu/qa-qc/coverage-report)
  * [Validate Variants](/partek-flow/user-manual/task-menu/qa-qc/validate-variants)
  * [Feature distribution](/partek-flow/user-manual/task-menu/qa-qc/feature-distribution)
  * [Single-cell QA/QC](/partek-flow/user-manual/task-menu/qa-qc/single-cell-qa-qc)
  * [Cell barcode QA/QC](/partek-flow/user-manual/task-menu/qa-qc/cell-barcode-qa-qc)
* [Pre-alignment tools](/partek-flow/user-manual/task-menu/pre-alignment-tools)
  * [Trim bases](/partek-flow/user-manual/task-menu/pre-alignment-tools/trim-bases)
  * [Trim adapters](/partek-flow/user-manual/task-menu/pre-alignment-tools/trim-adapters)
  * [Filter reads](/partek-flow/user-manual/task-menu/pre-alignment-tools/filter-reads)
  * [Trim tags](/partek-flow/user-manual/task-menu/pre-alignment-tools/trim-tags)
* [Post-alignment tools](/partek-flow/user-manual/task-menu/post-alignment-tools)
  * [Filter alignments](/partek-flow/user-manual/task-menu/post-alignment-tools/filter-alignments)
  * [Convert alignments to unaligned reads](/partek-flow/user-manual/task-menu/post-alignment-tools/convert-alignments-to-unaligned-reads)
  * [Combine alignments](/partek-flow/user-manual/task-menu/post-alignment-tools/combine-alignments)
  * [Deduplicate UMIs](/partek-flow/user-manual/task-menu/post-alignment-tools/deduplicate-umis)
  * [Downscale alignments](/partek-flow/user-manual/task-menu/post-alignment-tools/downscale-alignments)
* [Annotation/Metadata](/partek-flow/user-manual/task-menu/annotation-metadata)
  * [Annotate cells](/partek-flow/user-manual/task-menu/annotation-metadata/annotate-cells)
  * [Annotation report](/partek-flow/user-manual/task-menu/annotation-metadata/annotation-report)
  * [Publish cell attributes to project](/partek-flow/user-manual/task-menu/annotation-metadata/publish-cell-attributes-to-project)
  * [Attribute report](/partek-flow/user-manual/task-menu/annotation-metadata/attribute-report)
  * [Annotate Visium image](/partek-flow/user-manual/task-menu/annotation-metadata/annotate-visium-image)
* [Pre-analysis tools](/partek-flow/user-manual/task-menu/pre-analysis-tools)
  * [Generate group cell counts](/partek-flow/user-manual/task-menu/pre-analysis-tools/generate-group-cell-counts)
  * [Pool cells](/partek-flow/user-manual/task-menu/pre-analysis-tools/pool-cells)
  * [Split matrix](/partek-flow/user-manual/task-menu/pre-analysis-tools/split-matrix)
  * [Hashtag demultiplexing](/partek-flow/user-manual/task-menu/pre-analysis-tools/hashtag-demultiplexing)
  * [Merge matrices](/partek-flow/user-manual/task-menu/pre-analysis-tools/merge-matrices)
  * [Descriptive statistics](/partek-flow/user-manual/task-menu/pre-analysis-tools/descriptive-statistics)
  * [Spot clean](/partek-flow/user-manual/task-menu/pre-analysis-tools/spot-clean)
* [Aligners](/partek-flow/user-manual/task-menu/aligners)
* [Quantification](/partek-flow/user-manual/task-menu/quantification)
  * [Quantify to annotation model (Partek E/M)](/partek-flow/user-manual/task-menu/quantification/quantify-to-annotation-model-partek-em)
  * [Quantify to transcriptome (Cufflinks)](/partek-flow/user-manual/task-menu/quantification/quantify-to-transcriptome-cufflinks)
  * [Quantify to reference (Partek E/M)](/partek-flow/user-manual/task-menu/quantification/quantify-to-reference-partek-em)
  * [Quantify regions](/partek-flow/user-manual/task-menu/quantification/quantify-regions)
  * [HTSeq](/partek-flow/user-manual/task-menu/quantification/htseq)
  * [Count feature barcodes](/partek-flow/user-manual/task-menu/quantification/count-feature-barcodes)
  * [Salmon](/partek-flow/user-manual/task-menu/quantification/salmon)
* [Filtering](/partek-flow/user-manual/task-menu/filtering)
  * [Filter features](/partek-flow/user-manual/task-menu/filtering/filter-features)
  * [Filter groups (samples or cells)](/partek-flow/user-manual/task-menu/filtering/filter-groups-samples-or-cells)
  * [Filter barcodes](/partek-flow/user-manual/task-menu/filtering/filter-barcodes)
  * [Split by attribute](/partek-flow/user-manual/task-menu/filtering/split-by-attribute)
  * [Downsample Cells](/partek-flow/user-manual/task-menu/filtering/downsample-cells)
* [Normalization and scaling](/partek-flow/user-manual/task-menu/normalization-and-scaling)
  * [Impute low expression](/partek-flow/user-manual/task-menu/normalization-and-scaling/impute-low-expression)
  * [Impute missing values](/partek-flow/user-manual/task-menu/normalization-and-scaling/impute-missing-values)
  * [Normalization](/partek-flow/user-manual/task-menu/normalization-and-scaling/normalization)
  * [Normalize to baseline](/partek-flow/user-manual/task-menu/normalization-and-scaling/normalize-to-baseline)
  * [Normalize to housekeeping genes](/partek-flow/user-manual/task-menu/normalization-and-scaling/normalize-to-housekeeping-genes)
  * [Scran deconvolution](/partek-flow/user-manual/task-menu/normalization-and-scaling/scran-deconvolution)
  * [SCTransform](/partek-flow/user-manual/task-menu/normalization-and-scaling/sctransform)
  * [TF-IDF normalization](/partek-flow/user-manual/task-menu/normalization-and-scaling/tf-idf-normalization)
* [Batch removal](/partek-flow/user-manual/task-menu/batch-removal)
  * [General linear model](/partek-flow/user-manual/task-menu/batch-removal/general-linear-model)
  * [Harmony](/partek-flow/user-manual/task-menu/batch-removal/harmony)
  * [Seurat3 integration](/partek-flow/user-manual/task-menu/batch-removal/seurat3-integration)
* [Differential Analysis](/partek-flow/user-manual/task-menu/differential-analysis)
  * [GSA](/partek-flow/user-manual/task-menu/differential-analysis/gsa)
  * [ANOVA/LIMMA-trend/LIMMA-voom](/partek-flow/user-manual/task-menu/differential-analysis/anova-limma-trend-limma-voom)
  * [Kruskal-Wallis](/partek-flow/user-manual/task-menu/differential-analysis/kruskal-wallis)
  * [Detect alt-splicing (ANOVA)](/partek-flow/user-manual/task-menu/differential-analysis/detect-alt-splicing-anova)
  * [DESeq2(R) vs DESeq2](/partek-flow/user-manual/task-menu/differential-analysis/deseq2-r-vs-deseq2)
  * [Hurdle model](/partek-flow/user-manual/task-menu/differential-analysis/hurdle-model)
  * [Compute biomarkers](/partek-flow/user-manual/task-menu/differential-analysis/compute-biomarkers)
  * [Transcript Expression Analysis - Cuffdiff](/partek-flow/user-manual/task-menu/differential-analysis/transcript-expression-analysis-cuffdiff)
  * [Troubleshooting](/partek-flow/user-manual/task-menu/differential-analysis/troubleshooting)
* [Survival Analysis with Cox regression and Kaplan-Meier analysis - Partek Flow](/partek-flow/user-manual/task-menu/survival-analysis-with-cox-regression-and-kaplan-meier-analysis-partek-flow)
* [Exploratory Analysis](/partek-flow/user-manual/task-menu/exploratory-analysis)
  * [Graph-based Clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/graph-based-clustering)
  * [K-means Clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/k-means-clustering)
  * [Compare Clusters](/partek-flow/user-manual/task-menu/exploratory-analysis/compare-clusters)
  * [PCA](/partek-flow/user-manual/task-menu/exploratory-analysis/pca)
  * [t-SNE](/partek-flow/user-manual/task-menu/exploratory-analysis/t-sne)
  * [UMAP](/partek-flow/user-manual/task-menu/exploratory-analysis/umap)
  * [Hierarchical Clustering](/partek-flow/user-manual/task-menu/exploratory-analysis/hierarchical-clustering)
  * [AUCell](/partek-flow/user-manual/task-menu/exploratory-analysis/aucell)
  * [Find multimodal neighbors](/partek-flow/user-manual/task-menu/exploratory-analysis/find-multimodal-neighbors)
  * [SVD](/partek-flow/user-manual/task-menu/exploratory-analysis/svd)
  * [CellPhoneDB](/partek-flow/user-manual/task-menu/exploratory-analysis/cellphonedb)
* [Trajectory Analysis](/partek-flow/user-manual/task-menu/trajectory-analysis)
  * [Trajectory Analysis (Monocle 2)](/partek-flow/user-manual/task-menu/trajectory-analysis/trajectory-analysis-monocle-2)
  * [Trajectory Analysis (Monocle 3)](/partek-flow/user-manual/task-menu/trajectory-analysis/trajectory-analysis-monocle-3)
* [Variant Callers](/partek-flow/user-manual/task-menu/variant-callers)
  * [SAMtools](/partek-flow/user-manual/task-menu/variant-callers/samtools)
  * [FreeBayes](/partek-flow/user-manual/task-menu/variant-callers/freebayes)
  * [LoFreq](/partek-flow/user-manual/task-menu/variant-callers/lofreq)
* [Variant Analysis](/partek-flow/user-manual/task-menu/variant-analysis)
  * [Fusion Gene Detection](/partek-flow/user-manual/task-menu/variant-analysis/fusion-gene-detection)
  * [Annotate Variants](/partek-flow/user-manual/task-menu/variant-analysis/annotate-variants)
  * [Annotate Variants (SnpEff)](/partek-flow/user-manual/task-menu/variant-analysis/annotate-variants-snpeff)
  * [Annotate Variants (VEP)](/partek-flow/user-manual/task-menu/variant-analysis/annotate-variants-vep)
  * [Filter Variants](/partek-flow/user-manual/task-menu/variant-analysis/filter-variants)
  * [Summarize Cohort Mutations](/partek-flow/user-manual/task-menu/variant-analysis/summarize-cohort-mutations)
  * [Combine Variants](/partek-flow/user-manual/task-menu/variant-analysis/combine-variants)
* [Copy Number Analysis (CNVkit)](/partek-flow/user-manual/task-menu/copy-number-analysis-cnvkit)
* [Peak Callers (MACS2)](/partek-flow/user-manual/task-menu/peak-callers-macs2)
* [Peak analysis](/partek-flow/user-manual/task-menu/peak-analysis)
  * [Annotate Peaks](/partek-flow/user-manual/task-menu/peak-analysis/annotate-peaks)
  * [Promoter sum matrix](/partek-flow/user-manual/task-menu/peak-analysis/promoter-sum-matrix)
* [Motif Detection](/partek-flow/user-manual/task-menu/motif-detection)
* [Metagenomics](/partek-flow/user-manual/task-menu/metagenomics)
  * [Kraken](/partek-flow/user-manual/task-menu/metagenomics/kraken)
  * [Alpha & beta diversity](/partek-flow/user-manual/task-menu/metagenomics/alpha-and-beta-diversity)
  * [Choose taxonomic level](/partek-flow/user-manual/task-menu/metagenomics/choose-taxonomic-level)
* [10x Genomics](/partek-flow/user-manual/task-menu/10x-genomics)
  * [Cell Ranger - Gene Expression](/partek-flow/user-manual/task-menu/10x-genomics/cell-ranger-gene-expression)
  * [Cell Ranger - ATAC](/partek-flow/user-manual/task-menu/10x-genomics/cell-ranger-atac)
  * [Space Ranger](/partek-flow/user-manual/task-menu/10x-genomics/space-ranger)
  * [STARsolo](/partek-flow/user-manual/task-menu/10x-genomics/starsolo)
* [V(D)J Analysis](/partek-flow/user-manual/task-menu/vdj-analysis)
* [Biological Interpretation](/partek-flow/user-manual/task-menu/biological-interpretation)
  * [Gene Set Enrichment](/partek-flow/user-manual/task-menu/biological-interpretation/gene-set-enrichment)
  * [GSEA](/partek-flow/user-manual/task-menu/biological-interpretation/gsea)
* [Correlation](/partek-flow/user-manual/task-menu/correlation)
  * [Correlation analysis](/partek-flow/user-manual/task-menu/correlation/correlation-analysis)
  * [Sample Correlation](/partek-flow/user-manual/task-menu/correlation/sample-correlation)
  * [Similarity matrix](/partek-flow/user-manual/task-menu/correlation/similarity-matrix)
* [Export](/partek-flow/user-manual/task-menu/export)
* [Classification](/partek-flow/user-manual/task-menu/classification)
* [Task actions](/partek-flow/user-manual/task-menu/task-actions)
* [Feature linkage analysis](/partek-flow/user-manual/task-menu/feature-linkage-analysis)

Clicking a **Task node** gives you the option to view the *Task results* or perform *Task actions* such as rerunning the task (Figure 1).

<div align="left"><figure><img src="/files/MXbaVjxYde4zTY1qLS27" alt=""><figcaption><p>Figure 1. Task menu invoked from a Task node</p></figcaption></figure></div>

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Task actions

Left single clicking on any task (the rectangles) in the analysis pipeline will cause a Task Actions section to appear in the pop-up menu. This allows users to:

* Rerun tasks: rerun the selected task, the task dialog will pop-up and users can change parameters of the task. Previous downstream analysis of the selected task will not be rerun.
* Rerun with downstream tasks: rerun the selected task, the task dialog will pop-up, users can change the parameters of the current task and the downstream analysis will be rerun with the same configuration as the previous one.
* Edit description: the description of the task can be replaced by manually typing in string.
* Change color: choose a color to apply only on the selected task by clicking on **Apply.** Click **Apply to downstream** to change the selected task and the downstream pipeline color to the newly selected color.
* Delete task: this option is only available if the user is the owner of the project or the owner of the task. When a task is deleted, all downstream tasks, including tasks from other users, will be deleted. Users may check the box to choose to delete the task's output files. If delete output files is not checked, the task will be removed from the pipeline, but the output files of the task will remain on the disk.
* Restart task: this option is only available on failed tasks and requires an admin role to perform, but does not require that you have a user account. Since you are logged in as an admin, restarting a task will not take up a concurrent seat and the disk space consumed by the output files will count towards the original owner of the task's storage space.

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.


# Data summary report

The *Data summary report* in Partek Flow provides an overview of all tasks performed as part of a pipeline. This is particularly useful for report writing, record keeping and revisiting projects after a long period of time.

This user guide will cover the following topics:

* [Viewing the Data Summary Report](#viewing-the-data-summary-report)
* [Saving the Data Summary Report](#saving-the-data-summary-report)
* [Quick Video Demo of the Data Summary Report](#quick-video-demo-of-the-data-summary-report)

## Viewing the Data Summary Report

Click on an output data node under the *Analyses* tab of a project and choose **Data summary report** from the context sensitive menu on the right (Figure 1). The report will include details from all of the tasks upstream of the selected the node. If tasks have been performed downstream of the selected data node, they will not be included in the report.

![Figure 1. Accessing the data summary report. In this example, the report will include details on Sample data and the Trim bases, Align reads, Quantify to transcriptome and Gene analysis tasks](/files/IKCW7Kmgh2ryG8ZFDM1y)

Each task will appear as a separate section on the *Data summary report* (Figure 2). The first section of the report (*Sample data*) will summarize the input samples information. Click the **grey arrows** ( ![arrow\_down\_icon\_collapse\_triangle\_gray](/files/w2LVyET8X38H8sCR9J9g) / ![arrow\_right\_icon\_expand\_triangle\_gray](/files/zcT1L2ZANNl5q2BM8wE5) ) to expand and collapse each section. When expanded, the task name, user that performed the task, start date and time, duration and the output file size are displayed (Figure 2). To view or hide a table of task settings, click **Show/hide details** (Figure 3).

![Figure 2. The Data summary report](/files/SuLqgolCfwK9ELTnwEZe)

![Figure 3. Click the Show/hide details link to reveal the task settings. Note how non-default settings are highlighted in red](/files/gfyA0Jri5fdhKFg1yoFn)

## Saving the Data Summary Report

The *Data summary report* can be saved in different formats via the web browser. The instructions below are for Google Chrome. If you are using a different browser, consult your browser's help for equivalent instructions.

### Save as a PDF

On the *Data summary report*, expand all sections and show all task details. Right-click anywhere on the page and choose **Print...** from the menu (Figure 4) or use **Ctrl+P** (**Command+P** on Mac). In the print dialog, click **Change…** (Figure 5) and set the destination to **Save as PDF**. Select the **Background graphics** checkbox (optional), click the blue **Save** button (Figure 5) and choose a file location on your local machine.

The PDF can be attached to an email and/or opened in a PDF viewer of your choice.

![Figure 4. To save the Data summary report as a PDF in Google Chrome, choose Print.](/files/DYHtwGzJB8hOQrMFTfW2)

![Figure 5. Print dialog in Google Chrome](/files/2hwipPBknbaIursH44hO)

### Save as HTML

On the *Data summary report*, right-click anywhere on the page and choose **Save as…** from the menu (Figure 6) or use **Ctrl+S** (**Command+S** on Mac). Choose a file location on your local machine and set the file type to **Web Page, Complete**.

The HTML file can be opened in a browser of your choice.

<figure><img src="/files/PfpuMzUfWpfjI5XrUN8B" alt=""><figcaption><p>Figure 6. To save the Data summary report as HTML in Google Chrome, choose Save as...</p></figcaption></figure>

## Quick Video Demo of the Data Summary Report

The short video clip below (with audio) shows a tutorial of looking at the Data Summary Report

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZG5p2nnl0YGy3GEZrqCO%2Fuploads%2FeskKNhvRarbkfTV67hPI%2FAudit_trail.mp4?alt=media&token=e0aebd3f-bf9f-4ee7-9583-2f8186ab5d74>" %}

## Additional Assistance

If you need additional assistance, please visit [our support page](http://www.partek.com/support) to submit a help ticket or find phone numbers for regional support.




---

[Next Page](/llms-full.txt/1)

