All pages
Powered by GitBook
Couldn't generate the PDF for 130 pages, generation stopped at 100.
Extend with 50 more pages.
1 of 100

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Loading...

Task actions

Left single clicking on any task (the rectangles) in the analysis pipeline will cause a Task Actions section to appear in the pop-up menu. This allows users to:

  • Rerun tasks: rerun the selected task, the task dialog will pop-up and users can change parameters of the task. Previous downstream analysis of the selected task will not be rerun.

  • Rerun with downstream tasks: rerun the selected task, the task dialog will pop-up, users can change the parameters of the current task and the downstream analysis will be rerun with the same configuration as the previous one.

  • Edit description: the description of the task can be replaced by manually typing in string.

  • Change color: choose a color to apply only on the selected task by clicking on Apply. Click Apply to downstream to change the selected task and the downstream pipeline color to the newly selected color.

  • Delete task: this option is only available if the user is the owner of the project or the owner of the task. When a task is deleted, all downstream tasks, including tasks from other users, will be deleted. Users may check the box to choose to delete the task's output files. If delete output files is not checked, the task will be removed from the pipeline, but the output files of the task will remain on the disk.

  • Restart task: this option is only available on failed tasks and requires an admin role to perform, but does not require that you have a user account. Since you are logged in as an admin, restarting a task will not take up a concurrent seat and the disk space consumed by the output files will count towards the original owner of the task's storage space.

If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

Data summary report

The Data summary report in Partek Flow provides an overview of all tasks performed as part of a pipeline. This is particularly useful for report writing, record keeping and revisiting projects after a long period of time.

This user guide will cover the following topics:

QA/QC

Partek Flow contains a number of quality control tools and reports that can be used to evaluate the current status of your analysis and decide downstream steps. Quality control tools are organized under the Quality Assurance / Quality Control (QA/QC) section of the context-sensitive menu and are available for unaligned and aligned reads data nodes.

This section will illustrate:

Filter reads

Next generation sequencing (NGS) data is notably huge in file size. Dealing with NGS data is not only time consuming but also puts constraints on hard disk space. This is especially true if analysis parameters need to be optimized. The Filter reads task is a very useful tool to get a subset of the raw data upon which optimization can be performed. The optimized parameters can then be saved and applied to the whole dataset.

Filter reads is only available for unaligned reads of FASTQ format. Select the Unaligned Reads data node then select Filter reads from the Pre-alignment tools section on the menu.

There are two options to filter reads: Subsample reads and Filter by read length.

To Subsample reads, specify how many reads you want to keep for every nth reads. For example: if the user specifies to "Keep one read for every 10 reads" (Figure 1), this means that for every 10 reads, the program will keep only 1 read. This is equivalent to keeping 10% of the data.

To Filter by read length, set the read length limits by choosing the minimum and maximum read length(s) to keep.

Post-alignment tools

Post-alignment tools involve tasks that can be performed on aligned data. These are typically used in preparing aligned data for other downstream analyses, such as DNA-Seq or RNA-Seq analysis.

To invoke Post-alignment tools, click on an Aligned reads data node (Figure 1). There are three functions available in Post-alignment tools:

Combine alignments

The Combine alignments task can potentially maximize the number of reads that map to a region. When unaligned reads resulting from one aligner are then aligned using a different one, the two resulting alignments can combined together for downstream analysis within Partek Flow. Note that this can only be performed if they were aligned to the same reference genome.

To invoke this task, click an Aligned reads data node and select Combine alignments (Figure 1).

A list of compatible alignments will appear. The color of the text signifies the layer the alignment corresponds to (Figure 2). Select the alignment you would like to combine and click Finish.

The resulting Aligned reads data node can now be used for downstream analysis (Figure 3).

Note that this task combines the files in the data node within Partek Flow but does not merge the BAM files. Downloading an aligned reads data node from a Combine alignments task will result to multiple BAM files per sample.

Downscale alignments

The Downscale alignments task can be invoked on data node containing bam/sam files, e.g. Aligned reads data node. The only parameter to specify is the percentage of the alignments to keep in the results, the range of the parameter should be between 0 to 100 (Figure 1). All the samples in the input data node will use the same parameter. The output data node contains bam files with a subset of the alignments.

If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

Annotation/Metadata

This section has tools that are useful in managing and understanding single cell data, especially for downstream analysis. To invoke Annotation/Metadata tools, click on any Single cell counts data node. These include the following tasks:

Annotation report

The Annotation report provides a table summarizing the cell-level attributes of a single cell counts data node.

To run Annotation report:

  • Click a Single cell counts data node

  • Click the Annotation/Metadata section of the toolbox

  • Click Annotation report

Pre-analysis tools

This section constitutes tools that are useful in preparing single cell data for downstream analysis, such as multi-sample comparison or multi-omics analysis. To invoke Pre-analysis tools, click on any Single cell counts data node. These include the following tasks:

Generate group cell counts

If a single cell data node contains cell attribute information, e.g., clustering results, classifications, or imported attributes, a counts-type data node containing the number of cells from each attribute group for each sample can be generated and used for downstream analysis.

To invoke Generate group cell counts:

  • Click a single cell count data node with cell-level attribute information

  • Click Pre-analysis tools in the toolbox

Quantification

In RNA-seq data analysis, after alignment, the most common step is to estimate gene or/and transcript expression abundance, the expression level is represented by read counts. There are three options in this step:

Quantify to transcriptome (Cufflinks)

Cufflinks assembles transcripts and estimates transcript abundances on aligned reads. Implementation details are explained in Trapnell et al. [1]

The Cufflinks task has three options that can be configured (Figure 1):

  • Novel transcript: this option does not require any annotation reference, it will do de novo assembly to reconstruct transcripts and estimate their abundance

  • Annotation transcript: this option requires an annotation model to quantify the aligned reads to known transcripts based on the annotation file.

Impute low expression

Single cell RNA-seq gene expression counts are zero inflated due to inefficient mRNA capture. This normalization task is based on MAGIC[1]–MArkov Affinity-based Graph Imputation of Cells), to recover gene expression lost due to drop-out. The limitation on using this method is up to 50K cells in the input data node.

To invoke this task, click on a normalized data node which has less than 50K cells, it will first compute PCA to use the number of PCs specified to impute (Figure 1).

Click Finish to run the task, it will output low expression imputed matrix in the output report node.

References

  1. Dijk D et al. MAGIC: A diffusion-based imputation method reveals gene-gene interactions in single-cell RNA-sequencing data

\

Impute missing values

This task is to replace missing data in the data with estimated values based on selected method.

First select the computation is based on samples/cells or features, and click Finish to replace missing values. Some functions will generate the same results no matter which transform option is selected, e.g. constant value. Others will generate different results:

  • Constant values: specify a value to replace the missing data

  • Maximum: use maximum value of samples/cells or features to replace missing data depends transform option

Additional Assistance

our support page

Quantify regions

  • HTSeq

  • Count feature barcodes

  • Salmon

  • Quantify to annotation model (Partek E/M)
    Quantify to transcriptome (Cufflinks)
    Quantify to reference (Partek E/M)

    Additional Assistance

    our support page
    Figure 1. Default filter alignments settings
    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Additional Assistance

    Figure 1. Select number of PCs to use for imputation

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. Combining alignments
    Figure 2. Select the alignment you would like to combine with
    Figure 3. Combined alignments will appear as a new Aligned reads data node

    Additional Assistance

    Attribute report

  • Annotate Visium image

  • If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Annotate cells
    Annotation report
    Publish cell attributes to project

    Additional Assistance

    Hashtag demultiplexing

  • Merge matrices

  • Descriptive statistics

  • Spot clean

  • If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Generate group cell counts
    Pool cells
    Split matrix

    Additional Assistance

    Novel transcript with annotation as guide: this option requires an annotation file to quantify the aligned reads to known transcripts as well as assemble aligned reads to novel transcripts. The results include all transcripts in the annotation file plus any novel transcripts that are assembled.

    When the Use bias correction check box is selected, it will use the genome sequence information to look for overrepresented sequences and improve the accuracy of transcript abundance estimates.

    1. Trapnell C, Williams B, Pertea G, et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nature Biotech. 2010; 28:511-515.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. Cufflinks configuration dialog

    References

    Additional Assistance

    Mean: use mean value of samples/cells or features to replace missing data depends transform option

  • Median: use median value of samples/cells or features to replace missing data depends transform option

  • Minimum: use minimum value of samples/cells or features to replace missing data depends transform option

  • K-nearest neighbor (mean): specify number of neighbors (N), Euclidean metric is used to compute neighbors, use mean of (N) neighbors to replace missing data

  • K-nearest neighbor (median): specify number of neighbors (N), Euclidean metric is used to compute neighbors, use median of (N) neighbors to replace missing data

  • If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. Select methods to replace missing values in the data

    Additional Assistance

  • Coverage Report

  • Validate Variants

  • Feature distribution

  • Single-cell QA/QC

  • Cell barcode QA/QC

  • In addition to the tools listed above, many other functionalities can also be interpreted in sense of quality control. For instance, principal components analysis, hierarchical clustering (on sample level), variant detection report, and quantification report.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Pre-alignment QA/QC
    ERCC Assessment

    Additional Assistance

    Post-alignment QA/QC

    Click on an output data node under the Analyses tab of a project and choose Data summary report from the context sensitive menu on the right (Figure 1). The report will include details from all of the tasks upstream of the selected the node. If tasks have been performed downstream of the selected data node, they will not be included in the report.

    Figure 1. Accessing the data summary report. In this example, the report will include details on Sample data and the Trim bases, Align reads, Quantify to transcriptome and Gene analysis tasks

    Each task will appear as a separate section on the Data summary report (Figure 2). The first section of the report (Sample data) will summarize the input samples information. Click the grey arrows ( / ) to expand and collapse each section. When expanded, the task name, user that performed the task, start date and time, duration and the output file size are displayed (Figure 2). To view or hide a table of task settings, click Show/hide details (Figure 3).

    Figure 2. The Data summary report
    Figure 3. Click the Show/hide details link to reveal the task settings. Note how non-default settings are highlighted in red

    The Data summary report can be saved in different formats via the web browser. The instructions below are for Google Chrome. If you are using a different browser, consult your browser's help for equivalent instructions.

    On the Data summary report, expand all sections and show all task details. Right-click anywhere on the page and choose Print... from the menu (Figure 4) or use Ctrl+P (Command+P on Mac). In the print dialog, click Change… (Figure 5) and set the destination to Save as PDF. Select the Background graphics checkbox (optional), click the blue Save button (Figure 5) and choose a file location on your local machine.

    The PDF can be attached to an email and/or opened in a PDF viewer of your choice.

    Figure 4. To save the Data summary report as a PDF in Google Chrome, choose Print.
    Figure 5. Print dialog in Google Chrome

    On the Data summary report, right-click anywhere on the page and choose Save as… from the menu (Figure 6) or use Ctrl+S (Command+S on Mac). Choose a file location on your local machine and set the file type to Web Page, Complete.

    The HTML file can be opened in a browser of your choice.

    Figure 6. To save the Data summary report as HTML in Google Chrome, choose Save as...

    The short video clip below (with audio) shows a tutorial of looking at the Data Summary Report

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Viewing the Data Summary Report
    Saving the Data Summary Report

    Viewing the Data Summary Report

    Saving the Data Summary Report

    Save as a PDF

    Save as HTML

    Quick Video Demo of the Data Summary Report

    Additional Assistance

    Quick Video Demo of the Data Summary Report
    Figure 2. Filter by read length(s) by setting the parameters for minimum and maximum read length

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. By default, subsample raw data by keeping one read for every 10 reads

    Additional Assistance

    Combine alignments
  • Deduplicate UMIs

  • Downscale alignments

  • Figure 1. Invoking Post-alignment tools

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Filter alignments
    Convert alignments to unaligned reads

    Additional Assistance

    The task will run and generate a task report.

    The Annotation report includes two tables (Figure 1) - the top table summarizes the categorical attributes, giving the number of levels of each attribute, and the bottom table summarizes the numeric attributes, providing some basic summary statistics about the distribution of each attribute.

    Figure 1. Annotation report tables for categorical and numeric attributes

    To download a text-file version of one of the tables, click Download in lower right-hand corner of the table.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Additional Assistance

    Click Generate group cell counts
  • Select the attribute to group the cells from the Group by drop-down menu (Figure 1)

  • Click Finish

  • Figure 1. Select an attribute to group cells

    A group cell counts node will be generated. The data node contains a matrix of cell counts in each sample for each group. You can view the counts results in the Group cell counts report (Figure 2).

    Figure 2. Group cell counts for each sample are listed in the task report

    The Cell counts data node is a counts type data node and downstream analysis tasks, such as normalization, PCA, and ANOVA, can be used to analyze the group cell counts data.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Additional Assistance

    Pre-alignment tools

    Partek Flow provides Pre-alignment tools that allow the user to process next-generation sequencing data before proceeding to alignment. These tools are not only useful for controlling the quality of data, but can also be used for subsampling prior to analyzing the full dataset. There are three functions available in Pre-alignment tools:

    • Trim bases

    • Trim adapters

    • Filter reads

    User is expected to have preliminary understanding of:

    • File formats for next generation sequencing data

    • Phred-quality score

    In order to show the Pre-alignment tools, select an Unaligned reads or Trimmed reads data node. They will appear on the context-sensitive menu on the right of the screen (Figure 1).

    Different Pre-alignment tools are available for different formats of unaligned reads. For example: if the reads are in FASTQ format, then all four tools are available. On the other hand, if the unaligned reads are in FASTA or SFF format, then the Filter reads option is not available.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Trim adapters

    The existence of adapter sequences at the 5'-end or 3'-end of the reads has shown to be one of the major problems during alignment, causing the reads to be unaligned. Thus, removing adapter sequence is of utmost importance if the sequenced read length is longer than the molecule of interest, such as microRNA. The fact that mature microRNAs are short in length makes it almost certain that the adapter sequence will be sequenced at the 3'-end of the miRNA.

    In order to know whether the data has been adapter-trimmed for microRNA data, we can look at the pre-alignment QA/QC of the raw data, specifically the read length distribution. If the read length distribution peaks at approximately 22-23 bases, this usually means the data has been adapter-trimmed. However, if you have a fixed length distribution, then very likely the data is not adapter-trimmed and you will need to get the adapter sequence from your vendor or service provider and use the Trim adapter function to trim away the adapter sequence.

    Partek Flow software wraps Cutadapt [1], a widely used tool for adapter trimming. It can be used to trim adapter sequences in nucleotide-space data as well as color-space data.

    In order to use Trim adapters function, you will need to know the adapter sequences. To trigger the Trim adapters function, please select Unaligned Reads node and then select Trim adapters from the Pre-alignment tools section of the task pane. In the Trim adapters page (Figure 1), paste the adapter sequences into the textbox and select the button.

    There are three options when it comes to trimming the adapter sequence:

    • Trimming for adapter ligated to 3'-end: the adapter sequence and anything that follows it will be trimmed away from the 3'-end.

    • Trimming for adapter ligated to 5'-end or 3'-end: the adapter sequence is identified within the read or overlapping the 3'-end, then the adapter sequence and anything that follows it will be trimmed away. However, if the adapter sequence partially overlaps the 5'-end of the read, the initial portion of the read matching the adapter sequence is trimmed and anything that follows it is kept.

    • Trimming for adapter ligated to 5'-end: if the adapter sequence appears partially at the 5'-end or within the read, the preceding sequence including the adapter sequence is trimmed. User has the option to use a special character '^' at the beginning of the adapter sequence, meaning the adapter is 'anchored'. An anchored adapter must appear in its entirety at the 5'-end of the read (i.e. it is a prefix of the read).

    For Trim adapters, more than one adapter sequences can be specified at once. When multiple adapters are provided, all adapters are evaluated based on how many bases it overlaps the read as well as the error rate. Adapters which have a lower number of overlapped nucleotides or high error rates are removed from consideration.

    After that, the best adapter will be chosen based on the number of matching bases to the read. If there is a tie, adapters of the same type will be chosen in the order they are provided and adapters of different types will be chosen by type in the following order: first 3', then 5' or 3', and lastly 5' adapters.

    There are cases when the Trim adapters function does not work properly, for example: the existence of N's base in the read, etc. Therefore, there are advanced options which allows user to configure how the matching is done to trim adapter sequence. The advanced options dialog box is shown in Figure 2.

    The first section of advanced options is the Adapter options. This is used to configure how the matching between the adapter sequence and the read will be performed. This includes the maximum error rate allowed, the number of matched times, minimum length of overlapped bases, allowing Ns (ambiguous base) in adapter and whether N will be treated as wildcards. User can roll-over mouse cursor to the info button to get more information of each parameter.

    The second section of advanced options is the Filtering options. This is used to filter adapter-trimmed reads which are shorter than the minimum read length. This is to avoid having reads too short because short reads gives non-unique alignment and we would like to avoid that.

    The third section of advanced options is the Additional modification to reads. The quality cutoff is used to trim bad quality bases from the reads before trimming adapter. Quality encoding tells the quality score encoding for the raw data. The Reads names prefix and suffix is used to add prefix and suffix to the read ID. Lastly, the Negative quality zero if checked will convert all negative quality score base to zero.

    1. Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 2011; 17: 10-12.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Filter alignments

    • Introduction

    • Removing duplicates

    • Remove alignments with mismatches

    Introduction

    The Filter alignments task can be used to filter aligned reads data using specified parameters. To invoke the task, click on an Aligned reads data node and select Filter alignments. By default, this task removes low-quality reads, singletons and unaligned read information stored within the BAM/SAM file (Figure 1).

    Users also have the option to remove duplicate reads in aligned data. For DNA-Seq analysis, this is typically performed to minimize redundant variant calling information. To remove duplicates, click on the Remove duplicates checkbox (Figure 2).

    Select the number of reads you want to keep. Then specify when alignments are treated as duplicates. This can either be reads that map to the same start position or, additionally, have the same sequence. You can also select whether to keep the read with the highest mapping score or a randomly-selected duplicate.

    To remove alignments with mismatches, select the Remove alignments with mismatches check box. Using the selector, specify the number the number of mismatched bases that need to be exceeded for the alignment to be excluded (Figure 3). Note that mismatches also include insertions and deletions.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Convert alignments to unaligned reads

    Aligned reads can be converted to unaligned reads in Partek Flow. The task is available under Post-alignment tools in the task menu when any Aligned reads data node is selected, which can be a result of an aligner in Partek Flow or data already aligned before import.

    Generating unaligned reads from aligned data gives you the flexibility to remap the reads using either a different aligner, a different set of alignment parameters, or a different genome reference. This is particularly useful in analyzing sequences from xenograft models where the same set of reads can be aligned two different species. It may also be useful if the original unaligned FASTQ files are not as easily accessible to the user as the aligned BAM files.

    To perform the task, select an Aligned reads data node and click Convert alignments to unaligned reads task in the task menu (Figure 1).

    Figure 1. Convert alignments to unaligned reads
    Figure 1. Convert alignments to unaligned reads

    During the conversion, the BAM files are converted to FASTQ files and a new Unaligned reads data node will be generated (Figure 2) .

    Figure 2. Unaligned reads generated from aligned reads

    The filenames of the FASTQ files will be based on the sample names in the Data tab. The files generated are compressed with the extension *.fq.gz. For samples containing BAM files with paired end reads, two FASTQ files will be generated for each, and the files names will be appended with _1 and _2. An example in Figure 3 shows 18 .fq.gz output files that came from 9 BAM files.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Annotate cells

    If you have attribute information about your cells, you can use the Annotate cells task in Partek Flow to apply this information to the data. Once applied, these can be used like any other attribute in Partek Flow, and thus can be used for cell selection, classification and differential analysis.

    To run Annotate cells:

    • Click a Single cell counts data node

    • Click the Annotation/Metadata section in the toolbox

    • Click Annotate cells

    You will be prompted to specify annotation input options:

    • Single file (all samples): requires one .txt file for all cells in all samples. Each row in the file represents a barcode and at least one barcode column which will match the barcodes in your data. It also requires a column containing Sample ID which must match the Sample name in the Metadata tab of your project.

    • File per sample: requires the format of all of the annotation files to be the same. Each file has barcodes on rows, it requires one barcode column that will match the barcodes in your data in that sample. All files should have the same set of columns, column headers are case sensitive.

    Browse to each sample file on the Partek Flow server to specify annotation files for all of the samples in the dialog (Figure 1).

    To view a preview of the file, click Show Preview (Figure 2).

    If you would like to annotate your matrix features with a gene annotation file, you can choose an annotation file at the bottom on the dialog. You can choose any gene/feature annotation available on the Partek Flow server. If a feature annotation is selected, the percentage of mitochondrial reads will be calculated using the selected annotation file.

    • Click Next to continue

    The next dialog page previews the attributes found in the annotations text file (Figure 3).

    You can choose which attributes to import using the check-boxes, change the names of attributes using the text fields, and indicate whether an attribute with numbers is categorical or numeric.

    • Click Finish to import the attributes

    A new data node, Annotated single cell counts, will be generated (Figure 4).

    You annotations will be available in downstream analysis tasks.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Spot clean

    The terminology spot swapping describes the artifact for spatial data that mRNA bleed from nearby spots causes substantial contamination of UMI counts[1]. Spot clean in Flow is a task that aims to improve estimates of expression by correcting for spot swapping.

    The task can only be invoked from the Space Ranger task output data node since it takes the raw count matrix as input. To run the Spot clean task in Flow:

    • Click the Single cell counts outputted from Space Ranger (Figure 1)

    • Click Pre-analysis tools in the toolbox

    • Click Spot clean

    • Click Finish to run the task with default settings

    Another single cell counts node will be generated. The data node contains a matrix of cell counts with the decontaminated gene expressions (Figure 2). Downstream analysis tasks, such as normalization, PCA, and ANOVA, can be performed on the new single cell counts node.

    Parameters in this task that you can adjust include:

    Gene cutoff: Filter out genes with average expressions among tissue spots below or equal to this cutoff. Default: 0.1.

    Max iteration: Maximum iteration for EM parameter updates. Default: 10. Set a smaller number to save computation time.

    1. Ni, Z., Prasad, A., Chen, S. et al. SpotClean adjusts for spot swapping in spatial transcriptomics data. Nat Commun 13, 2971 (2022).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Quantify to reference (Partek E/M)

    This task does not need an annotation model file, since the annotation is retrieved from the BAM file itself. The sequence names in the BAM files constitute the features with which the reads are quantified against.

    This task is generally performed on reads aligned to a transcriptome, e.g when a species does not have a genome reference, and the bam files contain transcriptome information. In this case, the features for this quantification task are the reference sequence names in the input bam files.

    There are two parameters in Quantify to reference (Figure 1):

    Figure 1. Quantify to reference dialog
    • Min coverage: will filter out any features (sequence names) that have fewer reads across all samples than the value specified

    • Strict paired-end compatibility: this only affects paired end data. When it is checked, only reads that have two ends aligned to the same feature will be counted. Otherwise, reads will still be counted as exonic compatible reads even if the mate is not compatible with the feature

    During quantification:

    1. We scan through each of the BAM files and find all the transcripts that meet the minimum coverage threshold.

    2. With those transcripts, we "create" an annotation file that has the transcript name as the sequence name and the Gene ID and the Transcript ID have the same transcript name. The start position is 1 and the end position is the length of the transcript.

    3. Effectively, what the annotation file does is filter out the low coverage transcripts.

    4. Since we don't know where the transcripts are in the genome, chromosome view will display only one transcript at a time (i.e., the transcript names are treated like "chromosomes").

    The output data node will display a similar Task report as the task.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Salmon

    • Running Salmon

    • References

    Salmon1 is a method for quantifying transcript abundance from raw read sequence fastq files, it will generate transcript level count matrix as output.

    https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5600148/

    Running Salmon

    Salmon will run on any FASTQ file input.

    • Click the data node containing your FASTQ files

    • Click the Quantification section in the toolbox

    • Click Salmon

    • Choose Assembly and Aligner index on gene annotation model (Figure 1)

    Salmon index can't be built on reference assembly, you need to provide transcript annotation file beforehand, the index will be built on transcript annotation model. For more information about adding aligner index based on annotation model, see this in library file management.

    The task generates two data nodes: Transcript counts node contains NumReads which is the estimate of number of reads mapping to each transcript; the best estimate is often not integer. Gene counts node contains the sum of all the reads from the corresponding transcripts from each gene.

    Note: If you want to perform differential analysis, you need to add offset to deal with 0 values.

    1. Rob Patro, Geet Duggal, Michael Love, Rafeal A Irizarry, Carl Kingsford. Salmon: fast and bias-aware quantification of transcript expression using dual-phase inference. Published online 2017 Mar 6. doi:

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Validate Variants

    Validate variants is available for data nodes containing variants (Variants, Filtered variants, or Annotated Variants). The purpose of this task is to understand the performance of the variant calling pipeline by comparing variant calls from a sample within the project to known “gold standard” variant data that already exist for that sample. This “gold standard” data can encompass variants identified with high confidence by other experimental or computational approaches.

    Setting up the task (Figure 1) involves identifying the Genome Build used for variant detection and the Sample to validate within the project. Target specific regions allow for specification of the Target regions for this study, relating to the regions sequenced for all samples in the project. Benchmark target regions represent the regions that have been previously interrogated to identify “gold standard” variant calls in the sample of interest. These parameters are important to ensure that only overlapping regions are compared, avoiding the identification of false positives or false negative variants in regions covered by only the project sample or the “gold standard” sample. Both sections utilize a Gene/feature annotation file, which can be previously associated with Partek Flow via or added on the fly. The Validated variants file is a single sample vcf file containing the “gold standard” variant calls for the sample of interest and can be previously associated with Partek Flow as a Variant Annotation Database via Library File Management or added on the fly.

    Deduplicate UMIs

    The Deduplicate UMIs task identifies and removes reads mapped to the same chromosomal location with duplicate unique molecular identifiers (UMIs). The details of our UMI deduplication methods are outlined in the white paper.

    To invoke Deduplicate UMIs:

    Hashtag demultiplexing

    Cell Hashing enables sample multiplexing and super-loading in single cell RNA-Seq isolation by labeling each sample with a sample-specific oligo-tagged antibody against a ubiquitously expressed cell surface protein.

    The Hashtag demultiplexing

    Quantify regions

    In ChIP-seq or ATAC-seq analysis, a major challenge after detecting enriched regions or peaks is to compare samples and identify differentially enriched regions. In order to compare samples, a common set of regions must be identified and the number of reads mapping to each region quantified. The Quantify regions task addresses this challenge by generating a union set of unique regions and reporting the number of reads from each sample mapping to each region.

    To run Quantify regions:

    • Click a Peaks data node

    • Click the Quantification section in the toolbox

    Filter barcodes

    In droplet-based single cell isolation and library prep methods, each droplet is labeled by a unique nucleotide barcode. Because not all droplets will contain cells, it is important to filter out nucleotide barcodes that correspond to empty droplets prior to downstream analysis.

    In Partek Flow, you can filter barcodes interactively using a knee plot after UMI deduplication in the or after quantification in the . Alternatively, you can filter using preset options in the Filter barcodes task.

    To invoke Filter barcodes:

    Batch removal

    When a project contains multiple libraries, the data might contains variabilities due to technical differences (e.g. sequencing machine, library prep kit etc.) in addition to biological differences (like treatment, genotype etc.). Batch removal is essential to remove the noise and discover biological variations.

    Filtering

    Partek Flow has the flexibility to subsample your data for further downstream analyses. Filter data by:

    Differential Analysis

    Powerful Partek Flow statistical analysis tools help identify differential expression patterns in the dataset. These can take into account a wide variety of data types and experimental designs.

    Normalization and scaling

    To ensure that different data sets are comparable, several normalization and scaling options are available in Partek Flow. These include newly-developed algorithms specifically tailored for genomic analysis.

    Split matrix

    Split matrix can be invoked on any counts data node with more than one feature type. For example, a CITE-Seq experiment would have Gene Expression counts and Antibody Capture counts in the single cell counts data node. Datasets generated by experiments also utilize this task to split different feature measurements for downstream analysis.

    There are no parameters to configure, to run:

    • Click the counts data node you want to split

    • Click the Pre-analysis tools section of the toolbox

    Exploratory Analysis

    Partek Flow offers a wide variety of tools to help you explore your data. Which tools are available depends on the type of data node selected.

    Attribute report

    This task is only available on single cell matrix data node. It will summarize the cell level attributes of a data node and the result displayed in two tables with one table containing the categorical attributes while the other table contains the numerical attribute.

    To download a text-file version of one of the tables, click Download in lower right-hand corner of the table.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Downsample Cells

    The Downsample Cells task is used to randomly downsample the number of cells in a single cell data set. This task can be used to reduce the size of large single cell datasets to small and manageable sizes for quick analysis. Another use case for this task is for a project with multi samples with each sample having different number of cells. Downsample Cells can be used to randomly select an equal number of cells for all the samples in the project. For the default setting, the sample with the minimum number of cells is used with the number of cells in that sample set as the number of cells to be selected in the other samples. However, this default setting can be changed to a preferred number by the user. If the number selected by the user is greater than the number of cells in one or more samples, those samples will not be downsampled and all the cells in those samples will be returned. If the number selected by the user is greater than the number of cells in all the samples, then none of the samples will be downsampled.

    To run a downsample task first click on a single cell count data node. Go to the Filtering section and select Downsample cells task (Figure 1).

    Clicking on the Downsample cells task will lead to a dialogue menu with the number of cells to be downsampled set to the minimum number of cells in the project. In the figure below, the minimum number of cells in any of the samples was 2658 and this is used in the default settings. Click Finish to run the task (Figure 2).

    Split by attribute

    The Split by Attribute task is used to split a data node into different nodes based on the groups in a categorical attribute, each data node only includes the samples/cells from one group. It is a more efficient way to filter your data if you plan to perform downstream analysis on each and every group separately in an attribute.

    Click on the data node and select split by attribute from the Filtering section in task menu (Figure 1).

    Select the attribute to split the data on. In this case, data will be split according to the Age attribute (Figure 2).

    Result of the split by attribute task will be two separate data nodes, each contains samples from one age group (Figure 3).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    General linear model
    Harmony
    Seurat3 integration
    Split by attribute
  • Downsample Cells

  • Filter features
    Filter groups (samples or cells)
    Filter barcodes

    Detect alt-splicing (ANOVA)

  • DESeq2(R) vs DESeq2

  • Hurdle model

  • Compute biomarkers

  • Transcript Expression Analysis - Cuffdiff

  • Troubleshooting

  • GSA
    ANOVA/LIMMA-trend/LIMMA-voom
    Kruskal-Wallis

    Normalize to baseline

  • Normalize to housekeeping genes

  • Scran deconvolution

  • SCTransform

  • TF-IDF normalization

  • Impute low expression
    Impute missing values
    Normalization

    PCA

  • t-SNE

  • UMAP

  • Hierarchical Clustering

  • AUCell

  • Find multimodal neighbors

  • SVD

  • CellPhoneDB

  • Graph-based Clustering
    K-means Clustering
    Compare Clusters

    Showing Pre-alignment tools

    Additional Assistance

    Trim tags
    our support page
    Figure 1. Showing Pre-alignment Tools from an unaligned reads node

    Removing duplicates

    Remove alignments with mismatches

    Additional Assistance

    our support page
    Figure 1. Default filter alignments settings
    Figure 2. Removing duplicate reads
    Figure 3. Removing alignments with mismatches

    Additional Assistance

    our support page
    Figure 1. Selecting a cell annotation file
    Figure 2. Annotation files should have cells on rows and cell attributes on columns
    Figure 3. Previewing attributes to import
    Figure 4. Output of annotate cells

    References

    Additional Assistance

    https://doi.org/10.1038/s41467-022-30587-y
    our support page
    Figure 1. Invoke Spot clean in Partek Flow.
    Figure 2. Spot clean task output in Flow.

    References

    Additional Assistance

    chapter
    10.1038/nmeth.4197
    our support page
    Figure 1. Configure Salmon task--select index file

    Attribute report

    Running attribute report

    Click on any single cell count data node and select Attribute report from the Annotation/Metadata task menu

    Double click on the result node to view the Attribute report table

    Result of the Attribute report task showing categorical and numerical attributes

    Additional Assistance

    our support page
    Figure 1. Click on data node (in red circle) and select attribute report (red rectangle)
    Figure 2. Double click on the Attribute task report (red oval)
    Figure 3. Attribute report tables for categorical and numeric attributes
    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Additional Assistance

    Figure 1. Single Click on Data node to be downsampled (red) and select the Downsample cells task (red rectangle) from the filtering menu
    Figure 2. Downsample cell number to 2658 for each sample in project

    Additional Assistance

    our support page
    Figure 1. Single Click on Data node to be split (red) and select the split by attribute task (red rectangle) from the filtering menu
    Figure 2. Splitting data into >110 and <110 groups of the Age category. Click Finish
    Figure 3. Count data node split into two :>110 and <110 data nodes (red rectangle)

    Additional Assistance

    our support page
    Figure 2. Unaligned reads generated from aligned reads
    Figure 3. Task details of unaligned reads generated from aligned paired end samples
    Figure 3. Task details of unaligned reads generated from aligned paired end samples

    Additional Assistance

    Quantify to annotation model
    our support page
    Figure 1. Task dialog for Validate variants

    The Validate variants results page contains statistics related to the comparison of variants in the project sample compared to the validated variant calls for the sample (Figure 2). The results are split into two sections, one based on metrics calculated from the comparison of SNVs and the other from the comparison of INDELs.

    Figure 2. Example of the Variant validation report, with analysis at the level of both SNVs and INDELs. Note that the table is truncated due to the number of columns.

    The following SNP-level metrics are contained within the report, comparing the sample in the project to the validated variant data:

    • No genotypes: the number of missing genotypes from the sample in the Flow project

    • Same as reference: the number of homozygous reference genotypes from the sample in the Flow project

    • True positives: the number of variant genotypes from the sample in the Flow project that match the validated variants file

    • False positives: the number of variant genotypes from the sample in the Flow project that are not found in the validated variants file

    • True negatives: the number of loci that do not have variant genotypes in the sample in the Flow project and the validated variants file

    • False negatives: the number of genotypes that do not have variant genotypes in the sample in the Flow project but do have variant genotypes in the validated variants file

    • Sensitivity: the proportion of variant genotypes in the validated variants file that are correctly identified in the sample in the Flow project (true positive rate)

    • Specificity: the proportion of non-variant loci in the validated variants file that are non-variant in the sample in the Flow project (true negative rate)

    • Precision: the number of true positive calls divided by the number of all variant genotypes called in the sample in the Flow project (positive predictive value),

    • F-measure: a measure of the accuracy of the calling in the Flow pipeline relative to the validated variants. It considers both the precision and the recall of the test to compute the score. The best value at 1 (perfect precision and recall) and worst at 0.

    • Matthews correlation: a measure of the quality of classification, taking into account true and false positives and negatives. The Matthews correlation is a correlation coefficient between the observed and predicted classifications, ranging from −1 and +1. A coefficient of +1 represents a perfect prediction, 0 no better than random prediction and −1 indicates completely wrong prediction.

    • Transitions: variant allele interchanges of purines or pyrimidines in the sample in the Flow project relative to the reference

    • Transversions: variant allele interchanges of purines to/from pyrimidines in the sample in the Flow project relative to the reference

    • Ti/Tv ratio: ratio of transition to transversions in the sample in the Flow project

    • Heterozygous/Homozygous ratio: the ratio of heterzygous and homozygous genotypes in the sample in the Flow project

    • Percentage of sites with depth < 5: the percentage of variant genotypes in the sample in the Flow project that have fewer than 5 supporting reads

    • Depth, 5th percentile: 5% of sequencing depth found across all variant genotypes in the sample in the Flow project

    • Depth, 50th percentile: 50% of sequencing depth found across all variant genotypes in the sample in the Flow project

    • Depth, 95th percentile: 95% of sequencing depth found across all variant genotypes in the sample in the Flow project

    The INDEL-level metrics columns contained within the report are identical, with the exception of a lack of information with regards to transitions and transversion.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Library File Management

    Additional Assistance

    Click an Aligned reads data node
  • Click Post-alignment tools in the toolbox

  • Click Deduplicate UMIs

  • The task configuration dialog content depends on whether you imported FASTQ files or BAM files into Partek Flow.

    UMIs and barcodes are detected and recorded by the Trim tags task in Partek Flow. You can choose whether to retain only one alignment per UMI or not (Figure 1). The default will depend on which prep kit was used in the Trim tags task.

    Figure 1. UMI deduplication dialog if UMIs and barcodes were processed by Trim tags

    If you select Retain only one alignment per UMI, you will be asked to choose an assembly and gene/feature annotation file. The annotation file is used to check whether a read overlaps an exonic region. Only reads that have 50% overlap with an exon will be retained.

    If you do not select Retain only one alignment per UMI, UMI deduplication will proceed without filtering to exonic reads. Other differences between the two options are outlined in the UMI Deduplication in Partek Flow white paper.

    Imported BAMs generated by other tools can be imported into Partek Flow and deduplicated by the software. Additional options are available in the task configuration dialog to allow you to specify the location of the UMI and cell barcode information typically stored in the BAM header. Specify the BAM header tags in the text fields. For example, when processing a BAM file produced by CellRanger 3.0.1, the BAM identifier tag for the UMI sequence is UR and the BAM identifier for the barcode sequence is CR (Figure 2).

    Figure 2. Specify the location of the BAM UMI and barcode tags

    The option to Retain only one alignment per UMI is also available when starting from a BAM file.

    The Deduplicate UMIs task report includes a knee plot showing the number of deduplicated reads per barcode. This plot is used to filter the barcodes to include only barcodes corresponding to cells. For more information about using the knee plot to filter barcodes, please see the Cell Barcode QA/QC page. One difference between the Deduplication report and the Cell Barcode QA/QC report is that the Deduplication report gives the number of initial alignments and the number of deduplicated alignments for each sample (Figure 3). This indicates how many of your aligned reads were PCR duplicates and how many were unique molecules.

    The initial number of cells is set by our automatic filter. You can set the filter manually by clicking on the plot or by typing a cutoff number in the Cells or Reads in cells text boxes. If there are multiple samples, each sample receives a plot and filters are set per sample.

    Figure 3. The deduplication report shows the number of UMIs per barcode

    The number of cells, reads in cells, median reads per cell, number of initial alignments, and number of deduplicated alignments are listed for each sample in the summary table (Figure 4).

    Figure 4. Deduplication report summary table

    Clicking Apply filter at either the knee plot or the summary table will run the Filter barcodes task and generate a Filtered reads data node.

    To return to the knee plot, click Back to filter.

    To reset the filters for all sample to the automatic cutoff, click Reset all filters.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Configuring Deduplicate UMIs
    Deduplicate UMIs task report
    UMI Deduplication in Partek Flow

    Configuring Deduplicate UMIs

    Imported FASTQ

    Imported BAM

    Deduplicate UMIs task report

    Additional Assistance

    task is an implementation of the algorithm used in Stoeckius et al. 20181 for multiplexing cell hashing data. The task adds cell-level attributes
    Sample of origin
    and
    Cells in droplet
    .

    To run Hashtag demultiplexing, your data must meet the following criteria:

    • Data node contains number of features less than number of observations

    • Data node must be output from normalization task (recommended normalization method for hashtag is CLR)

    If you are processing your FASTQ files in Partek Flow, be sure to specify a different Data type for your Cell Hashing FASTQ files on import than the FASTQ files for your gene expression and any other antibody data.

    If you are processing your FASTQ files using Cell Ranger, be sure to specify a different feature_type for your Cell Hashing antibodies than any other antibodies in the Feature Reference CSV File.

    If you want to specify sample IDs instead of using hashtag feature IDs as the sample IDs, you will need to prepare a tab-delimited text file (.txt) with hashtag feature IDs in the first column and the corresponding sample IDs in the second column (Figure 1). A header row is required.

    Figure 1. Sample ID .txt file example. The first column is the hashtag feature ID and the second column is the corresponding sample ID
    • Click the Normalized counts data node for your cell hashing data

    • Click Hashtag demultiplexing in the Pre-analysis tools section of the toolbox

    • Click Browse to select your Sample ID file (Optional)

    • Click Finish to run

    The output is a Demultiplexed counts data node (Figure 2).

    Figure 2. Hashtag demultiplexing can be run on a Normalized counts data node and outputs a Demultiplexed counts data node

    Two cell-level attributes, Cells in droplet and Sample of origin, are added by this task and are available for use in downstream tasks. You can download the attribute values for each cell by clicking the Demultiplexed counts data node, clicking Download, and choosing to download Attributes only.

    We recommend using Annotate cells to transfer the new attributes to other sections of your project after downloading the attributes text file.

    It is also possible to use the Merge matrices task to combine your data types and attributes.

    1. Stoeckius, M., Zheng, S., Houck-Loomis, B., Hao, S., Yeung, B.Z., Mauck, W.M., Smibert, P. and Satija, R., 2018. Cell Hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics. Genome biology, 19(1), p.224.

    \

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Prerequisites for running Hashtag demultiplexing
    Running Hashtag demultiplexing
    References

    Prerequisites for running Hashtag demultiplexing

    Running Hashtag demultiplexing

    References

    Additional Assistance

    Click Quantify regions

    The Quantify regions task takes MACS2 output, a Peaks or Annotated Peaks data node, as its input. In a typical ATAC-Seq or ChIP-Seq analysis, MACS2 is configured to output a set of enriched regions or peaks for each experimental sample or group individually. Quantify regions takes these sets of regions and merges them into a union set of unique regions that it saves as a .bed file. To combine the region sets, overlapping regions between samples/groups are merged. Where overlap ends, a break point is created and a new region defined. All non-overlapping or unique regions from each sample/group are also included.

    For example, consider an experiment where MACS2 detected enriched regions for two samples, Sample A and Sample B. In Sample A, a region is detected on chromosome 1 from 100bp to 300bp, chr1:100-300. In Sample B, a region is detected at chr1:160-360. The Quantify regions task will give the following union set of unique regions for these partially overlapping regions:

    chr1:100-160 (region detected in Sample A only)

    chr1:160-300 (region detected in both Sample A and Sample B)

    chr1:300-360 (region detected in Sample B only)

    After generating a .bed file with the union set of unique regions, Quantify regions performs quantification using the same algorithm as Quantify to annotation model (Partek E/M) with the .bed file as the annotation model.

    The Quantify regions dialog includes configuration options for generating the union set of unique regions and quantifying reads to the regions (Figure 1).

    When regions from multiple samples are combined, a small offset in position between enriched regions in different samples can result in many very short unique regions in the union set. The Minimum region size option lets you filter out these very short regions. If a region is smaller than the specified cutoff, the region is excluded. By default, this is set to 50bp, but may need to be adjusted depending on the size of regions you expect to see in your assay.

    Quantification options are the same as in the Quantify to annotation model (Partek E/M) dialog. The Percent of read length is set to 50% by default to account for small offsets in position between enriched regions in different samples.

    Figure 1. Quantify regions dialog

    Quantify regions generates a counts data node with the number of counts in each region for each sample. This data node can be annotated with gene information using the Annotate regions task and analyzed using tasks that take counts data as input, such as normalization, PCA, and ANOVA. For ChIP-Seq experiments with input control samples, the Normalize to baseline task can be used prior to downstream analysis.

    Similar to the Quantify to annotation model (Partek E/M) task report, the Quantify regions task report includes feature distribution information including a descriptive stats table, a distribution bar chart, a sample box plot, and sample histogram (Figure 2).

    Figure 2. Quantify regions task report

    To download the .bed file with the union set of unique regions, click the Quantify regions task node, click Task details, click the regions.bed file in the Output files section, and click Download.

    1. Xing Y, Yu T, Wu YN, Roy M, Kim J, Lee C. An expectation-maximization algorithm for probabilistic reconstructions of full-length isoforms from splice graphs. Nucleic Acids Res. 2006; 34(10):3150-60.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Quantify regions method

    Configuring Quantify regions

    Quantify regions output

    References

    Additional Assistance

    Click a Deduplicated reads data node

  • Click the Filtering section of the toolbox

  • Click Filter barcodes

  • You can choose to filter barcodes using three or four options, depending on whether you have already run a Filter barcodes task for your samples in the project (Figure 1).

    Figure 1. Configuring Filter barcodes

    The automatic filter threshold is set for each sample individually. It picks the cutoff between cells and empty droplets by identifying where the UMI content per barcode drops precipitously when moving in descending order from the barcode with the highest number of UMIs.

    Set the number of cells per sample to include. This is set for all samples; if set to 100, the top 100 barcodes by total UMI count for each sample will be retained.

    Set the percent of reads in cells per sample to include. The number of barcodes included will be set to match the specified percent of reads in cells for each sample. Barcodes are included starting with the barcode with the highest number of total UMIs and proceeding in descending order of total UMIs per barcode until the specified percent of reads has been met or exceeded.

    If you have already run a Filter barcodes task for your samples in the project, the Previous filter option will be available. This option lets you filter to the same cell barcodes that were included by the previous filter. This option is particularly useful for CITE-Seq data, where antibody barcodes and gene expression data must be processed separately, but you will want to analyze the same cell barcodes in downstream steps.

    Selecting Previous filter opens a table with information about the previous barcode filters in your project (Figure 2).

    Figure 2. Previous filter options

    To help you identify which previous filter you want to apply, the color of the task node on the Analyses tab, the number of cell barcodes retained (summed for all samples), and the time/date the previous filter task was submitted are included in the table. To view the number of cells and percentage of reads in cells for each sample in a previous filter, mouse over the button (Figure 3).

    Figure 3. Checking details of a previous filter

    Use the radio buttons in the first column to pick which filter you want to use.

    After configuring the task, click Finish to run.

    Filter barcodes produces a Filtered reads data node (Figure 4). Filter barcodes does not have a task report.

    Figure 4. Filter barcodes produces a Filtered reads data node

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Configuring Filter barcodes
    Output of Filter barcodes
    UMI deduplication task report
    Cell barcode QA/QC task report

    Configuring Filter barcodes

    Automatic

    Number of cells

    Percent of reads in cell

    Previous filter

    Output of Filter barcodes

    Additional Assistance

    Click Split matrix

    The Split matrix task will run and generate output data nodes for each of the feature types. For example, if there are Antibody Capture and Gene Expression feature types in the input, Split matrix will generate two data nodes (Figure 1). Every sample is included in both matrices.

    Figure 1. Split matrix generates separate data nodes for each feature type in the input matrix

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    10X Genomics' Feature Barcoding

    Additional Assistance

    Advanced Options for Trim Adapters

    References

    Additional Assistance

    our support page
    Figure 1. Trim adapters setup page
    Figure 2. Advanced options dialog box for Trim adapters function

    ERCC Assessment

    The ERCC (External RNA Control Consortium) developed a set of RNA standards for quality control in microarray, qPCR, and sequencing applications. These RNA standards are spiked-in RNA with known concentrations and composition (i.e. sequence length and GC content). They can be used to evaluate sensitivity and accuracy of RNA-seq data.

    The ERCC analysis is performed on unaligned data, if the ERCC RNA standards have been added to the samples. There are 92 ERCC spiked-in sequences with different concentrations and different compositions. The idea is that the raw data will be aligned (with Bowtie) to the known ERCC-RNA sequences to get the count of each ERCC sequence. This information is available within Partek Flow and will be used to plot the correlation between the observed counts and the expected concentration. If there is a high correlation between the observed counts versus the expected concentration, you can be confident that the quantified RNA-seq data are reliable. Partek Flow supports Mix1 and Mix 2 ERCC formulations. Both formulations use the same ERCC sequences, but each sequence is present at different expected concentrations. If both Mix 1 and Mix 2 formulations have been used, ExFold comparison can be performed to compare the observed and expected Mix1:Mix2 ratio for each spike-in.

    To start ERCC assessment, select an unaligned reads node and choose ERCC in the context sensitive menu. If all samples in the project have used the Mix 1 or Mix 2 formulation, choose the appropriate radio button at the top (Figure 1).

    Figure 1. Setting up advanced options for the alignment of ERCC controls using Bowtie

    If some samples have been treated with the Mix 1 formulation and others have been treated with the Mix 2 formulation, choose the ExFold comparison radio button (Figure 2). Set up the pairwise comparisons by choosing the Mix 1 and Mix 2 samples that you wish to compare from the drop-down lists, followed by the green plus ( ) icon. The selected pair of samples will be added to the table below.

    You can change the Bowtie parameters by clicking Configure before the alignment (Figure 1), although the default parameters work fine for most data. Once the task has been set up correctly, select Finish.

    ERCC task report starts with a table (Figure 3), which summarizes the result on the project level. The table shows which samples use the Mix 1 or Mix 2 formulation. The total number of alignments to the ERCC controls are also shown, which is further divided into the total number of alignments to the forward strand and the reverse strand. The summary table also gives the percentage of ERCC controls that contain alignment counts (i.e. are present). Generally, the fraction of present controls should be as high as possible, however, there are certain ERCC controls that may not contain alignment counts due to their low concentration; that information is useful for evaluation of the sensitivity of the RNA-seq experiment. The coefficient of determination (R squared) of the present ERCC controls is listed in the next column. As a rule of a thumb, you should expect a good correlation between the observed alignment counts and the actual concentration, or else the RNA-seq quantification results may not be accurate. Finally, the last two columns give estimates of bias with regards to sequence length and GC content, by giving the correlation of the alignment counts with the sequence length and the GC content, respectively.

    If ExFold comparison was enabled, an extra table will be produced in the ERCC task report (Figure 4). Each row in the table is a pairwise comparison. This table lists the percentage of ERCC controls present in the Mix 1 and Mix 2 samples and the R squared for the observed vs expected Mix1:Mix2 ratios.

    The ERCC spike-ins plot (Figure 5) shows the regression lines between the actual spike-in concentration (x-axis, given in log2 space) and the observed alignment counts (y-axis, given in log2 space), for all the samples in the project. The samples are depicted as lines, and the probes with the highest and lowest concentration are highlighted as dots. The regression line for a particular sample can be turned off by simply clicking on the sample name in the legend beneath the plot.

    Optionally, you can invoke a principal components analysis plot (View PCA), which is based on RPKM-normalised counts, using the ERCC sequences as the annotation file (not shown).

    For more details, go to the sample-level report (Figure 6) by selecting a sample name on the summary table. First, you will get a comprehensive scatter plot of observed alignment counts (y-axis, in log2 space) vs. the actual spike-in concentration (x-axis, in log2 space). Each dot on the plot represents an ERCC sequence, coloured based on GC content and sized by sequence length (plot controls are on the right).

    The table (Figure 7) lists individual controls, with their actual concentration, alignment counts, sequence length, and % GC content. The table can be downloaded to the local computer by selecting the Download link.

    For more details on ExFold comparisons, select a comparison name in the ExFold summary table (Figure 8). First, you will get a comprehensive scatter plot of observed Mix1:Mix2 ratios (y-axis, in log2 space) vs. the expected Mix1:Mix2 ratio (x-axis, in log2 space). Each dot on the plot represents an ERCC sequence, coloured based on GC content and sized by sequence length (plot controls are on the right).

    The table (Figure 9) lists individual controls, with each samples' alignment counts, together with the observed and expected Mix1:Mix2 ratios. The table can be downloaded to the local computer by selecting the Download link.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Post-alignment QA/QC

    Post-alignment QA/QC is available for data nodes containing aligned reads (Aligned reads) and has no special control dialog. Similar to the pre-alignment QA/QC report, the post-alignment contains two tiers, i.e. project-level report and sample-level report.

    The project-level report starts with a summary table (Figure 1). Unlike pre-alignment QA/QC report, each row now corresponds to a sample (sample names are hyperlinks to sample-level report). Table allows for a quick comparison across all the samples within the project. Any outlying sample can, therefore, easily be spotted.

    Note that the summary table reflects the underlying chemistry. While Figure 1 shows a summary table for single-end sequencing, an example table for paired-end sequencing is given in Figure 2. Common features are discussed first.

    The first two columns contain total number of reads (Total reads) and total number of alignments (Total alignments). Theoretically, for single-end chemistry, total number of reads equals total number of alignments. For double-end reads, theoretical result is to have twice as many alignments as reads (the term “read” refers to the fragment being sequenced, and since each fragment is sequenced from two directions, one can expect to get two alignments per fragment). When counting the actual number of alignments (Total alignments), however, reads that align more than once (multimappers) are also taken into account. Next, the Aligned column contains the fraction of all the reads that were aligned to the reference assembly.

    The Coverage column shows the fraction (%) of the reference assembly that was sequenced and the average sequencing coverage (×) of the covered regions is in the Avg. coverage depth column. The Avg. quality is mapping quality, as reported by the aligner (not all aligners support this metric). Avg. length is the average read length and average read quality is given in Avg. quality column. Finally, %GC is the fraction of G or C calls.

    In addition, the Post-alignment QA/QC report for single-end reads (Figure 1) contains the Unique column. This refers to the fraction of uniquely aligned reads.

    On the other hand, the Post-alignment QA/QC report for paired-end reads (Figure 2) contains these columns:

    • Unique singleton

      • fraction of alignments corresponding to the reads where only one of the paired reads can be uniquely aligned

    • Unique paired

      • fraction of alignments corresponding to the reads where both of the paired reads can be uniquely aligned

    Note: for paired-end reads, if one end is aligned, the mate is not aligned, the alignment rate calculating will include the read as the numerator, also since the mate is not aligned, we will also include this read in the unaligned data node (if the generate unaligned reads data node option is selected) for 2nd stage alignment, this will generate discrepancy between total reads and "unaligned reads + total reads * alignment rate", because reads with only one mate aligned are counted twice.

    In addition to the summary table, several graphs are plotted to give a comparison across multiple samples in the project. Those graphs are Alignment breakdown, Coverage, Genomic Coverage, Average base quality per position, Average base quality score per read, and Average alignments per read. Two of those (Average base quality plots) have already been described.

    The alignment breakdown chart (Figure 3) presents each sample as a column, and has two vertical axes (i.e. Alignment percent and Total reads). The percentage of reads with different alignment outcomes (Unique paired, Unique singleton, Non-unique, Unaligned) is represented by the left-side y-axis and visualized by stacked columns. The total number of reads in each sample is given using the black line and shown on the right-side y-axis.

    The Coverage plot (Figure 4) shows the Average read depth (in covered regions) for each sample using columns and can be red off the left-hand y-axis. Similarly, the Genomic coverage plot shows genome coverage in each sample, expressed as a fraction of the genome.

    The last graph is Average alignments per read (Figure 5) and shows the average number of alignments for each read, with samples as columns. For single-end data, the expected average alignments per read is one, while for paired-end data, the expected average alignments per read is two.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Coverage Report

    Coverage report is also available for data nodes containing aligned reads (Aligned reads, Trimmed reads, or Filtered reads). The purpose of the report is to understand how well the genomic regions of interest are covered by sequencing reads for a particular analysis.

    When setting up the task (Figure 1), you first need to specify the Genome build and then a Gene/feature annotation file, which defines the genomic regions you are interested in (e.g. exome or genes within a panel). The Gene/feature annotation can be previously associated with Partek® Flow® via Library File Management or added on the fly.

    Complete coverage report will contain percentage of bases within the specified genes / features with coverage greater than or equal to the coverage levels defined under Add minimum coverage levels. To add a level, click on the green plus . Alternatively, to remove it, click on the red cross icon.

    As for the Advanced options, if Strand-specificity is turned on, only reads which match the strand of a given region will be considered for that region’s coverage statistics.

    Generate target enrichment graphs will generate a graphical overview of coverage across each feature.

    When Use multithreading is checked, the computation will utilize multiple CPUs. However, if the input or output data is on some file systems like GPFS file system, which doesn't support well on multi-thread tasks, unchecking this option will prevent task failures.

    Coverage report result page contains project-level overview and starts with a summary table, with one sample per row (Figure 2) The first few columns show the percentage of bases in the genomic features which are covered at the specified level (or higher) (default: 1×, 20×, 100×). Average coverage is defined as the sum of base calls of each base in the genomic features divided by the length of the genomic features. Similarly, Average quality is defined as the sum of average quality of those bases that cover the genomic features, divided by the length of covered genomic features. The last two columns show the number of On-tarted reads (overlapping the genomic features) and Off-target reads (not overlapping the features). The Optional columns link enables import of any meta-data present in the data table ().

    Quantification of on- and off-target reads is also displayed in the column chart below the table (Figure 3), showing each sample as a separate column and fraction of on-/off-target reads on the y-axis.

    Region coverage summary hyperlink opens a new page, with a table showing average coverage for each region (rows), across the samples (columns) (Figure 4).

    The browser icon in the right-most column () of the Region average coverage summary table opens the Coverage graph for the respective region (Figure 5). The horizontal axis is the normalized position within the genomic feature, represented as 1st to 100th percentile of the length of the feature. The vertical axis is coverage. Each line on the plot is a single sample, and the samples are listed below the plot.

    The Coverage summary (Figure 6) plot is an overview of coverage across of the targeted genomic features for all the samples in the project. Each line within the plot is a single sample, the horizontal axis is the normalized position within the genomic feature, represented as 1st to 100th percentile of the length of the feature, while the vertical axis show the average coverage (across all the features for a given sample).

    If you need more details about a sample, click on the sample name in the Coverage report table (Figure 7). The columns are as follows:

    • Region name: the genomic feature identifier (as specified in the annotation file)

    • Chromosome: the chromosome of the genomic feature (or region)

    • Start: the start position of the genomic feature (1-based)

    • Stop: the stop position of the genomic feature (2-based, which means the stop position is exclusive)

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Feature distribution

    The Feature distribution plot visualizes the distribution of features in a counts data node.

    Running Feature distribution

    To run Feature distribution:

    • Click a counts data node

    • Click the QA/QC section of the toolbox

    • Click Feature distribution

    A new task node is generated with the Feature distribution report.

    The Feature distribution task report plots the distribution of all features (genes or transcripts) in the input data node with one feature per row (Figure 1). Features are ordered by average value in descending order.

    The plot can be configured using the panel of the left-hand side of the page.

    Using the filter, you can choose which features are shown in the task report.

    The Manual filter lets you type a feature ID (such as a gene symbol) and filter to matching features by clicking + . You can add multiple feature IDs to filter to multiple features (Figure 2).

    The List filter lets you filter to the features included in a feature list. To learn more about feature lists in Partek Flow, please see .

    Distributions can be plotted as histograms, with the x-axis being the expression value and the y-axis the frequency, or as a strip plot, where the x-axis is the expression value and the position of each cell/sample is shown as a thin vertical line, or strip, on the plot (Figure 3).

    To switch between plot types, use the Plot type radio buttons.

    Mousing over a dot in the histogram plot gives the range of feature values that are being binned to generate the dot and the number of cells/samples for that bin in a pop-up (Figure 4).

    Mousing over a strip shows the sample ID and feature value in a pop-up. If there are multiple cells/samples with the same value, only one strip will be visible for those cells/samples and the mouse-over will indicate how many cells/samples are represented by that one strip (Figure 5).

    Clicking a strip will highlight that cell/sample in all of the plots on the page (Figure 6).

    The grey dot in each strip plot shows the median value for that feature. To view the median value, mouse over the dot (Figure 7).

    To navigate between pages, use the Previous and Next buttons or type the page number in the text field and click Enter on your keyboard.

    The number of features that appear in the plot on each page is set by the Items per page drop-down menu (Figure 8). You can choose to show 10, 25, or 50 features per page.

    When Plot type is set to Histogram, you can choose to configure the Y-axis scale using the Scale Y-axis radio buttons. Feature max sets each feature plot y-axis individually. Page max sets the same y-axis range for every feature plot on the page, with the range determined by the feature with the highest frequency value.

    You can add attribute information to the plots using the Color by drop-down menu.

    For histogram plots, the histograms will be split and colored by the levels of the selected attribute (Figure 9). You can choose any categorical attribute.

    For strip plots, the sample/cell strips will be colored by the levels or values of the selected attribute (Figure 10). You can choose any categorical or numeric attribute.

    Trim bases

    The Trim bases task is used to trim bases from the 5'-end or 3'-end of the reads. The most obvious reason for Trim bases is to trim away poor quality bases from the read prior to alignment because these can potentially affect alignment rate.

    The task allows user to trim reads in different ways (Figure 1), including:

    • Trim bases based on quality score

    • Trim bases from 3'-end

    • Trim bases from 5'-end

    • Trim bases from both ends

    Trim bases from 5'-or 3'-end (Figures 2-3) allows a fixed number of bases to be trimmed away from the 5'- or 3'-end of the reads. These two functions are useful for when your read length is constant. This is not recommended if the read length is not constant, since good quality bases from shorter reads are likely trimmed away by these functions.

    Trim bases from both ends (Figure 4) allows user to keep only bases from a fixed start and end position of the reads. This is particularly useful if poor quality bases are observed on both ends of the read. So instead of performing trim bases successively from the 5'- and 3'-end, the trim bases will only be performed once by trimming from both ends.

    Trim bases based on quality score (Figure 5) is probably the most useful function to trim poor quality bases from the 5'- or 3'-ends of reads. This function allows dynamic trimming of bases depending on quality score. The trimming can be done from either 5'-end, 3'-end or both ends of the reads. The function evaluates each base from the end of the read and trims it away until the last base has a quality score greater than the specified threshold. For an extensive evaluation of read trimming effects on Illumina NGS data analysis, see Del Fabbro et. al. [1].

    In some cases, the reads that result from base trimming can have very short read lengths and thus are not recommended for alignment. Thus, Partek® Flow® Flow provides the option to set a Min read length after base trimming. This discards reads that are shorter than the set length.

    Also, reads could have a high percentage of N's or ambiguous bases. Thus, the Max N setting is available to discard reads with %Ns higher than the set threshold

    The Quality encoding option refers to the Phred quality score encoded within the FASTQ input file. The list of available options are: Phred+33, Phred+64, Solexa+64 and Integers. Selecting Auto-detect will determine whether the quality encoding is Phred+33 or Phred+64. For Solexa data, you will need to select Solexa+64. For most of datasets, auto-detect option works very well with a few exception cases where the base quality score falls into the grey zone (ambiguous zone) of Phred+33 and Phred+64 score. However, if the quality-encoding scheme is known, we recommend to selecting the encoding format directly from the quality encoding list.

    Figure 6 shows the options available for all the different selection of Trim bases function. Note the default Min read length is 25bp. For micro RNA sequencing data, this default Min read length needs to be set to a smaller value (we recommend 15) to account for mature microRNAs.

    The Task Details page for Trim bases can be accessed by selecting the task node Trim bases, and subsequently selecting Task Details from the Task results section. In the Task details page, several sections are available:

    • General task information: contains information such as the task name, owner, status, submitted time, start, end and duration of the task

    • Output Files: contains the description of each output file. If you roll-over your mouse cursor to the file name, you will get the exact location of the file on the server. If you click on the file name, you will have the option to view up to 999 lines of the raw data. You can also download the file from the server.

    • Input Files: contains the information of input files. This section lists down all the input files used in the Trim bases task.

    The Trim bases Task Report page can be accessed by selecting either the Trim bases task node or Trimmed reads data node and then selecting the Task Report from the Task results section of the context sensitive menu. There is a link at the bottom of the page to directly go to the Task Details page. The page displays the following components:

    • Summary table: gives the total number of reads in each sample, the total number of reads trimmed (i.e. with at least one base trimmed from the read), total number of reads removed (due to Min read length and Max N parameters), the average number of bases trimmed per read, the average read quality before trim bases and finally the average read quality after trim bases.

    • Stacked bar-chart: shows percentage of untrimmed reads, trimmed reads and removed reads are shown in a stacked bar-chart to compare all the samples.

    • Average base quality score per position of trimmed reads: shows the average base quality score at each position of the trimmed reads for all samples in the project.

    The Trim bases function produces trimmed unaligned reads which is named as Trimmed Reads data node. The Trimmed Reads node will have the "trimmed" word appended to the filename. The Trimmed Reads data can be downloaded by selecting the Trimmed Reads node and then select Download data from the context sensitive menu. However, if you have access to the Partek Flow server, you can go to the Task Details page and identify the location of the output files from the Output Files section as described on the Trim Bases Task Details section above. The Trimmed Reads data node will have the same format as the raw data.

    1. Del Fabbro C, Scalabrin S, Moragante M, Giorgi FM. An extensive evaluation of read trimming effects on Illumina NGS data analysis. PLoS ONE. 2013; 8(12): e85024.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.none;">43rates

    Annotate Visium image

    • Space Ranger output

    • Annotate Visium image

    • References

    Space Ranger output

    The Visium spatial gene expression solution from 10X Genomics allows us to spatially resolve RNA-Seq expression data in individual tissue sections. For the analysis of Visium spatial gene expression data in Partek Flow, you will need the following output files from the Space Ranger _outs_1 subdirectory:

    • The filtered count matrix -- either the .h5 file (one file) or feature-barcode matrix (three files).

    • spatial – Outputs of spatial pipeline (Figure 1) (Spatial imaging data)

    We recommend import count matrix file using .h5 file format, which allow you to import multiple samples at the same time. You need to rename the file names and put all the files in the sample folder to import in one go.

    The spatial subdirectory contains image related files, they are:

    tissue_hires_image.png

    tissue_lowres_image.png

    aligned_fiducials.jpg

    detected_tissue_image.jpg

    tissue_positions_list.csv

    scalefactors_json.json

    The folder should be compressed in one .gz or zip file when you upload to Flow server. You can pick the file for each sample from the Partek Flow server, your local computer, or a URL using the file browser .

    To run Annotate cells:

    • Click a Single cell counts data node

    • Click the Annotation/Metadata section in the toolbox

    • Click Annotate Visium image

    You will be prompted to pick a Spatial image file for each sample you want to annotate (Figure 2).

    • Click Finish

    A new data node, Annotated counts, will be generated (Figure 3).

    When the task report of the annotated counts node is opened (or double click on the Annotated counts node), the images will be displayed in data viewer (Figure 4)

    It is a 2D plot and XY axes are the tissue spot coordinates. Tissue spots are on top of the slide image. The opacity of the tissue spots can be changed using the slider to show more of the image under Configure>Style>Color.

    To view different samples in the Data viewer, navigate to Axes under Configure Click on the button under Content in the left panel (Figure 5).

    From the Configure>Background>Image drop-down list in the Data viewer, different formats of the image can be selected (Figure 6).

    Click on Show image to turn on or off the background image.

    Note that the "Annotate Visium image" task splits by sample so all of the downstream tasks will also do this (e.g. if the pipeline is built from this node all of the downstream tasks will also be split and viewed per sample). To generate plots with multiple samples on one plot, build the pipeline from the "Single cell counts" node.

    [1]

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    HTSeq

    • Configurable options

    • Output

    • References

    HTSeq is set of tools for processing high-throughput sequencing data1. In Partek Flow, we have implemented the htseq-count script from HTSeq for quantifying aligned reads to an annotation model.

    The input for HTSeq is an Aligned reads data node and a Gene/Feature annotation file.

    To run HTSeq:

    • Click an Aligned reads data node

    • Click the Quantification section of the toolbox

    • Click HTSeq

    Please note that HTSeq has not been optimized for performance and can take a very long time to run compared with on the same data.

    HTSeq includes basic options (Figure 1) and advanced options accessible by clicking Configure (Figure 2).

    The annotation file contains the features the aligned reads will be quantified to. For more information about adding an annotation model, please see .

    Depending on the library preparation method, information about the strand of the original transcript may be faithfully preserved or lost. This setting controls whether HTSeq considers strand during quantification. Consult your library preparation method user manual if you are unsure about if and how the method preserved strand information.

    If set to no, a read is considered to be overlapping a feature regardless of whether it maps to the same or opposite strand as the feature.

    If set to yes, the read has to be matched to the same strand as the feature if single-end reads and matched to the same strand for the first read and the opposite strand for the second read if paired-end reads.

    If set to reverse, the read has to be matched to the opposite strand as the feature if single-end reads and matched to the opposite strand for the first read and the same strand for the second read if paired-end reads.

    By default, only features (e.g., genes) with one or more aligned read will be included in the output. If this option is selected, all features from the annotation model, including those without any matching aligned reads, will be included.

    HTSeq will skip reads with an alignment quality lower than the value specified here. The default is 10.

    This option determines how HTSeq handles reads that partially overlap features. The default is union.

    If set to union, any read that is partially overlapped by a feature will be assigned to that feature. Assignment is non-exclusive if multiple feature overlap.

    If set to intersection-strict, only reads that are fully overlapped by a feature will be assigned to that feature. Assignment is non-exclusive if multiple feature fully overlap.

    If set to intersection-nonempty, reads are assigned to the feature that has the greatest overlap. Assignment is non-exclusive if multiple feature overlap the same amount.

    This option determines how HTSeq counts reads that are assigned to more than one feature. The default is none.

    If set to none, reads that are assigned to more than one feature are not counted for any feature.

    If set to all, reads are counted for all features they are assigned to.

    HTSeq outputs a Gene counts data node (Figure 3). There is no task report.

    The ribosomal reads % column, present when the data is downloaded, is identified by searching ribosomal genes by their gene symbol against a list of 89 L & S ribosomal genes taken from . -; it calculates 'the sum of counts across these genes / the total count' * 100 to give the % of ribosomal counts.

    1. Simon Anders, Paul Theodor Pyl, Wolfgang Huber. HTSeq — A Python framework to work with high-throughput sequencing data. Bioinformatics (2014), in print, online at doi:10.1093/bioinformatics/btu638

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Filter features

    • Noise reduction filter

    • Statistics based filter

    • Feature metadata filter

    • Feature list filter

    A common task in bulk and single-cell RNA-Seq analysis is to filter the data to include only informative genes. Because there is no gold standard for what makes a gene informative or not, and ideal gene filtering criteria depend on your experimental design and research question, Partek Flow has a wide variety of flexible filtering options.

    Filter features task can be invoked from any counts or single cell data node. Noise Reduction and Statistics Based filters take each feature and perform the specified calculation across all of the cells. The filter is applied to the values in the selected data node and the output is a filtered version of the input data node.

    In the task dialog, select the filter option to activate the filter type and configure the filter, then click Finish to run.

    The Noise reduction filter lets you exclude features that meet basic criteria (Figure 1).

    Descriptive statistics you can choose are:

    • Coefficient of variation: std. dev divided by mean of the feature

    • Geometric mean: nth root of the product of the n numbers, n is the number of features

    • Maximum: the highest value of a feature

    • Mean: the average value of a feature

    For each of these you can choose to exclude features that are:

    • <: less than

    • <=: less than or equal to

    • == equal to

    • >: greater than

    The threshold is set using the text box. The input must be a number; it can be an integer or decimal, positive or negative.

    If you select value, you can also choose a percentage of samples or cells that must meet the criteria for the feature to be excluded (Figure 2).

    The Statistics based filter lets you include a number or percentile of genes based on descriptive statistics (Figure 3).

    Select Counts to specify a number of top features to include or select Percentiles to specify the top percentile of features to include.

    Descriptive statistics you can choose are:

    • Coefficient of variance

    • Geometric mean

    • Maximum

    • Mean

    If the data linked to feature (gene) annotation, different fields in the annotation can be used to filter, e.g. genomic location information, gene biotype information etc. (Figure 4)

    You can specify logical operation using different annotation field information.

    You can filter features based on a feature lists (Figure 5).

    If you have added feature lists in Partek Flow using the feature, the filter using Saved list option will be available. Otherwise, you can specify a Manual list by typing in the Filter criteria box.

    If you choose Saved list, the drop-down list will display all the feature lists added in ; If you choose Manual list, you can manually type in the feature IDs/names in the box, one feature per row.

    You can choose to include or exclude features in any list that you have added.

    Use the Feature identifier option to choose which identifier from your annotation matches the values in the feature list.

    The filter features task report lists the filter criteria, reports distribution statistics for the remaining features, and indicates the number and percentage of features that passed the filter (Figure 6).

    If the input was a count matrix data node, sample box plot and sample histograms are provided to show the distribution of features after filtering. These plots are not available if the input was a single cell counts data node.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    General linear model

    • Advanced options

    This method is based on general linear model, much like ANOVA in reverse, it calculates the variation attributed to the factor(s) being removed then adjusting the original values to remove the variation.

    By including batch in the differential analysis model, the variability due to the batch effect is accounted for when calculating p-values. In this sense, batch effects are best handled as part of the differential analysis model. However, clustering data or visualizing biological effects can be very difficult if batch effects are present in the original data. We transform the original values to remove the batch effect using this tool.

    We recommend normalizing your data prior to removing batch effects, but the task will run on any counts data node.

    • Click the counts data node

    • Click the Batch removal section of the toolbox

    • Click General linear model

    The batch effect removal dialog is similar to the dialog for ANOVA. To set up the model, you need to choose which attributes should be considered. Here, you should include the batch attribute, any attributes that interact with batch, and the interactions between these attributes.

    For example, in the case where you have different cell types and batch may have a different effect on different cell types, you would need to include both batch, cell type, and the interaction between batch and cell type. Here, batch is the attribute Version and cell type is the attribute Classification.

    • Click Version and Classification

    • Click Add factors

    • Click Version and Classification

    • Click Add interaction (Figure 1)

    To remove the batch effect and its interaction with cell type, we can click the Remove checkbox for both Version and Version*Classification.

    • Click the Remove checkbox for Version and Version*Classification

    • Click Finish to run (Figure 2)

    The output of is a new data node, Batch effect adjusted counts. This data node contains the batch effect corrected values can be used as the input for downstream tasks such as clustering and UMAP (Figure 3).

    The advanced options for Remove batch effect are shared by .

    Normalize to housekeeping genes

    This normalization is performed on observations (samples) using internal control features (genes). The internal control features, usually housekeeping genes, should not vary among samples[1]. The implementation details is as follows:

    1. Compute geometric mean of all the control genes (features) e.g. (g1 to gm) in each sample S (f means feature, S means sample, 1-m are control features), represented by GS1 to GSn (n number of samples).

    2. Compute geometric mean of across all samples (GS1 to GSn), represented by GS

    3. Compute the scaling factor for each sample, S1=GS1/GS, S2=GS2/GS ... Sn=GSn/GS

    4. Normalize all the gene expression by divided by its sample scaling factor

    Note: The input data node must contain all positive values to compute geometric mean.

    Select Normalize to housekeeping genes task in Normalization and scaling section in the pop-up menu when you select a count matrix data node, the dialog will list all the features included in the data node on the left panel

    Figure 1. When a data node containing a count matrix is selected, Normalize to baseline is available in the toolbox

    Select control genes on the left panel and move them to the right panel. You can also use search box to find the feature and click the plus button to add it to the right panel.

    Click Finish

    1. Frank Speleman. Accurate normalization of real-time quantitative RT_PCR data by geometric averaging of multiple internal control genes. Genome Biology. 2002.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Compare Clusters

    What is Compare clusters?

    Compare clusters is a tool to identify the optimal number of clusters for K-means Clustering using the Davies-Bouldin index. The Davies-Bouldin index is a measure of cluster quality where a lower value indicates better clustering, i.e., the separation between points within the clusters is low (tight clusters) and separation between clusters is high (distinct clusters).

    Running Compare clusters

    We recommend normalizing your data prior to running Compare clusters, but the task will run on any counts data node.

    • Click the counts data node

    • Click the Exploratory analysis section of the toolbox

    • Click Compare clusters

    • Configure the parameters

    • Click Finish to run (Figure 1)

    The parameters for Compare clusters are the same as for K-means clustering.

    The Compare clusters task report is an interactive line chart with the number of clusters on the x-axis and the Davies-Bouldin index on the y-axis (Figure 2).

    The Compare clusters task report can be used to run K-means clustering.

    • Click a point on the plot to select it or type the number of clusters in the text box Partition data into clusters

    Selecting a point sets it as the number of clusters to partition the data into. The number of clusters with the lowest Davies-Bouldin index value is chosen by default.

    • Click Generate clusters to run K-means clustering with the selected number of clusters

    A K-means clustering task node and a Clustering result data node are produced. Please see our documentation on K-means Clustering for more details.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Compute biomarkers

    This task can be invoked from count matrix data node or clustering task report (Statistics > Compute biomarkers). It performs Student's t-tests on the selected attribute, comparing one subgroup at a time vs all the others combined. By default, the up-regulated genes are reported as biomarkers.

    • Compute biomarker dialog

    Compute biomarker dialog

    In the set-up dialog, select the attribute from the drop down list. The available attributes are categorical attributes which can be seen on the Data tab (i.e. project-level attributes) as well as and data node-specific annotation, e.g. graph-based clustering result (Figure 1). If the task is run on graph-based clustering output data node, the calculation is using upstream data node which contains feature counts – typically the input data node of PCA.

    Figure 1. Compute biomarker dialog: selecting attribute

    By default, the result outputs the features that are up-regulated by at least 1.5 fold change (in linear scale) for each subgroup comparing to the others. The result is displayed in a table with each column is a subgroup name, each row is a feature. Features are ranked by the ascending p-values within each sub-category. An example is shown in Figure 2. If a subgroup has fewer biomarkers than the others, the "extra" fields for that subgroup will be left blank.

    Figure 3. Biomarkers table (example). Top 10 biomarkers for each cluster are shown. Download link provides the full results table

    Furthermore, the Download link (upper-left corner of the table report) downloads a .txt file to the local computer (default file name: Biomarkers.txt), which contains the full report: all the genes with fold change > 1.5, with corresponding fold change and p-values.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Merge matrices

    In complex projects, different data matrices (e.g. observations on rows and features on columns) need to be merged in order to achieve the analysis goals. For example, two cell populations were identified on separate branches of the analysis pipeline and to combine them before any joint downstream steps, the expression matrices have to be combined. Alternatively, two assays (gene expression and protein expression) were performed on the same cells so the expression matrices have to be merged for joint analysis.

    Merge matrices task is located in the Pre-analysis tools section of the toolbox and it can handle two scenarios: Merge cells/samples and Merge features (Figure 1). To start, select the first data node on the pipeline (e.g. single cell counts) and then select the Merge matrices task.

    Figure 1. Setup page for Merge matrices task

    Merge Cells/Samples

    To use the Merge cells option, the data matrices (one or more) that are to be merged with the currently selected one should have the same features (e.g. genes), but distinct cells. Push the Select data nodes button and Partek Flow will display a preview of the pipeline; the data nodes that can be merged are shown in color of the branch, other data nodes are disabled (greyed out). Left click on the data node that you want to merge with the current one and push the Select button, you can select multiple data nodes to merge. The selected node(s) will be shown under the Select data nodes button (Figure 2). If you made a mistake, use the Clear selection icon. Push Finish to proceed.

    To use the Merge features option, the data matrices (one or more) that are to be merged with the currently selected one should have the same cells (or samples), but distinct features (e.g. gene and protein expression). Push the Select data nodes button and Partek Flow will display you a preview of the pipeline; the data nodes that can be merged are shown in color of the branch, others are disabled (greyed out). Left lick on the data node that you want to merge with the current one and push the Select button. The selected node will be shown under the Select data nodes button. Repeat the procedure if you would like to merge additional nodes. If you made a mistake, use the Clear selection icon. Push Finish to proceed.

    The output of the Merge matrices task is a Merged counts data node (Figure 3).

    For a practical example using Merge matrices, please see our tutorial on .

    Depending on your goal, you may want to consider a different approach. For example, data matrices based on two different assays (e.g. gene and protein expression) can be combined using . Instead of merging two (or more) cell populations by using Merge cells, you may want to use filtering (Filtering > Filter groups) to filter out the populations that you do not consider relevant / filter in the populations of your interest.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Scran deconvolution

    Library size normalization is the simplest strategy for performing scaling normalization. But composition biases will be present when any unbalanced differential expression exists between samples. The removal of composition biases is a well-studied problem for bulk RNA sequencing data analysis. However, single-cell data can be problematic for these bulk normalization methods due to the dominance of low and zero counts[1]. To overcome this, Partek Flow wrapped the calculateSumFactors() function from R package scran. It pools counts from many cells to increase the size of the counts for accurate size factor estimation. Pool-based size factors are then “deconvolved” into cell-based factors for normalization of each cell’s expression profile[1].

    Scran deconvolution in Flow can be invoked in Normalization and scaling section by clicking any single cell counts data node (Figure 1).

    Figure 1. Scran deconvolution task in Normalization and scaling section in Flow.

    To run Scran deconvolution,

    • Click a single cell counts data node

    • Click the Normalization and scaling section in the toolbox

    • Click Scran deconvolution

    The GUI is simple and easy to understand. The first Scran deconvolution dialog is asking to select the cluster name from a drop-down list that includes all the attributes for this dataset. The selected cluster is an optional factor specifying which cells belong to which cluster, for deconvolution within clusters (Figure 2). Simply click the Finish button if you want to run the task as default.

    The output of Scran deconvolution is a new data node that has been normalized by the pool-based size factors of each cell and log2 transformed_._ We can then use this new normalized matrix for downstream analysis and visualization (Figure 3).

    Other parameters in this task that you can adjust include:

    Pool size: A numeric vector of pool sizes, i.e., number of cells per pool.

    Max cluster size: An integer scalar specifying the maximum number of cells in each cluster.

    Enforce positive estimates: A logical scalar indicating whether linear inverse models should be used to enforce positive estimates.

    Scaling factor: A numeric scalar containing scaling factors to adjust the counts prior to computing size factors.

    1. Lun, A. T., K. Bach, and J. C. Marioni. Pooling across cells to normalize single-cell RNA sequencing data with many zero counts. Genome Biol. 2016.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    SCTransform

    SC transform task performs the variance stabilizing normalization proposed in [1]. The task's interface follows that of SCTransform() function in R [2]. SCTransform v2 [3] provides the ability to perform downstream differential expression analyses besides the improvements on running speed and memory consumption. v2 is the default method in Flow.

    We recommend performing the normalization on a single cell raw count data node. Select SCTransform task in Normalization and scaling section on the pop-up menu to invoke the dialog (Figure 1).

    Figure 1. When a data node containing a single cell count matrix is selected, SCTransform is available in the toolbox

    By default, it will generate report on all the input features. Unchecking the Report all features, user can limit the results to a certain number of features with highest variance.

    In Advanced options, users can the click Configure to change the default settings (Figure 2).

    Figure 2. Advanced configure options

    Scale results: Whether to scale residuals to have unit variance; default is FALSE

    Center results_:_ When set to Yes, center all the transformed features to have zero mean expression. Default is TRUE.

    VST v2: Default is TRUE. When set to 'v2', it sets method = glmGamPoi_offset, n_cells=2000, and exclude_poisson = TRUE which causes the model to learn theta and intercept only besides excluding poisson genes from learning and regularization; If default is unchecked, it uses the original sctransform model (v1), it will only generate SC scaled data node.

    There are two data nodes generated from this task (if VST v2 option is checked as default):

    SC scaled data: it is a matrix of normalized values (residuals) that by default has the same size as the input data set. This data node is used to perform downstream exploratory analysis e.g. PCA, Seurat3 integration etc (Figure 3), this data node is not recommend to use for differential analysis.

    SC corrected data: is equivalent to the ‘corrected counts’ in data slot generated after PrepSCTFindMarkers task in the SCT assay in Seurat object. It is used for downstream differential expression(DE) analyses (Figure 3).

    Note: When perform DE analysis with Hurdle, the 'shrinkage of error term variance' option might need to turn off depending on the dataset. Similarly, the 'Lognormal with shrinkage/voom' option needs to turn off when run DE with GSA.

    References

    1. Christoph Hafemeister, Rahul Satija. Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression.

    2. SCTransform() documentation

    3. Saket Choudhary, Rahul Satija. Comparison and evaluation of statistical error models for scRNA-seq.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Publish cell attributes to project

    This task is only available on single cell matrix data node. It will publish one or more cell level attributes to the project level, so the attribute can be edited and seen from all single cell count data nodes within a project. This function can be used on annotate cell task output data node, graph based cluster data node etc.

    Running publish cell attributes to project

    • Click on a single cell counts data node

    • Choose Publish cell attributes to project in the Annotation/Metadata section of the toolbox

    This invokes the dialog as (Figure 1)

    Figure 1. Select attributes to publish

    From the drop-down list to select one or more attributes to publish. Only numeric attributes and categorical attributes with less than 1000 levels will be available in the list.

    After selection, click on the green plus button () to add, change the attribute by typing in the New name box (Figure 2).

    Click Finish at the bottom of the page, all of the attributes will be available to edit on the Data tab > Cell attributes Manage. All data nodes in the project will be able to use those attributes.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    PCA

    Principal components (PC) analysis (PCA) is an exploratory technique that is used to describe the structure of high dimensional data by reducing its dimensionality. It is a linear transformation that converts n original variables (typically: genes or transcripts) into n new variables, which are called PCs, they have three important properties:

    • PCs are ordered by the amount of variance explained

    • PCs are uncorrelated

    • PCs explain all variation in the data

    PCA is a principal axis rotation of the original variables that preserves the variation in the data. Therefore, the total variance of the original variables is equal to the total variance of the PCs.

    If read quantification (i.e. mapping to a transcript model) was performed by Partek E/M algorithm, PCA can be invoked on a quantification output data node (Gene counts or Transcript counts) or, after normalization, on a Normalized counts data node. Select a node on the canvas and then PCA in the Exploratory analysis section of the context sensitive menu.

    There are two options for features contribute (Figure 1):

    equally: all the features are standardized to mean of 0 and standard deviation of 1 . This option will give all the features equal weight in the analysis, this is the default option for e.g bulk RNA-seq data.

    by variance: the analysis will give more emphasis to the features with higher variances. This is the default option for e.g. single cell RNA-seq data

    If the input data node is in linear scale, you can perform log transformation on PCA calculation.

    The PCA task creates a new task node, and to open it and see the result, do one of the following: select the PCA task node, proceed to the context sensitive menu and go to the Task result; or double-click on the PCA task node. The report containing eigenvalues, PC projections, component loadings, and mapping error information for the first three PCs.

    When the PCA node is opened in Data viewer, by default, it is the 3D scatterplot, Scree plot with Eigenvalues, and Component loadings table (Figure 2). Each dot on the 3D scatter plot represents an observation, while the first three PCs are shown on the X-, Y-, and Z-axis respectively, with the information content of an individual PC is in the parenthesis.

    As an exploratory tool, the the PCA scatterplot is applied to view any groupings in the data set and generate hypotheses based on the outcome, or to spot possible outliers.

    To rotate the 3D scatter plot left click & drag. To zoom in or out, use the mouse wheel. Click and drag the legend can move the legend to different location on the viewer.

    Detailed configuration on PCA plot can be found by clicking Help>How-to videos>Data viewer section.

    In the Data viewer, when a PCA data node is selected from Get Data under Setup (left panel), the node can be dragged and dropped to the screen (Figure 3), then you will have the option to select a scree plot and tables.

    When choose Scree plot icon, it will plot a 2D viewer, X-axis represents PCs, Y-axis represents eigenvalues (Figure 4)

    When mouse over on a point on the line, it will display detailed information of the PC. The scree plot shows how much variation each PC represents, so it is often used to determine the number of principal components to keep for downstream analysis (e.g. tSNE, UMAP, graph-base clustering). The "elbow" point of the graph where the eigenvalues seem to level off should be considered as a cutoff point for downstream analysis.

    PCA data node can also be draw as tables, when choose Table icon , it will display the component loadings matrix in the viewer (Figure 5). The Content can be modified using the Content configuration option; the table can be paged through here or from the lower right corner.

    In the table, each row is a feature, the column represent PCs, the value is the correlation coefficient. Under Content, there is a PCA projections option, change to this option to display the projection table (Figure 6). In this table, each row is an observation, each column is a PC, the values are the PC scores.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Kruskal-Wallis

    • Running the task

    • Report

    The Kruskal-Wallis and Dunn's tests (Non-parametric ANOVA) task is used to identify deferentially expressed genes among two or more groups. Note that such rank-based tests are generally advised for use with larger sample sizes.

    Running the task

    To invoke the Kruskal-Wallis test, select any count-based data nodes, these include:

    • Gene counts

    • Transcript counts

    • Normalized counts

    Select Statistics > Differential analysis in the context-sensitive menu, then select Kruskal-Wallis (Figure 1).

    Select a specific factor for analysis and click the Next button (Figure 2). Note that this task can only take into account one factor at a time.

    For more complicated experimental designs, go back to the original count data that will be used as input and perform Rank normalization at the Features level (Figure 3). The resulting Normalized counts data node can then be analyzed using the Detect differential expression (ANOVA) task, which can take into account multiple factors as well as interactions.

    Define the desired comparisons between groups and click the Finish button (Figure 4). Note that comparisons can only be added between single group (i.e. one group per box).

    The results of the analysis will appear similar to other differential expression analysis results. However, the column to indicate mean expression levels for each group will display the median instead (Figure 5).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    TF-IDF normalization

    Latent semantic indexing (LSI) was first introduced for the analysis of scATAC-seq data by Cusanovich et al. 2018[1]. LSI combines steps of frequency-inverse document frequency (TF-IDF) normalization followed by singular value decomposition (SVD). Partek Flow wrapped Signac's TF-IDF normalization for single cell ATAC-seq dataset. It is a two-step normalization procedure that both normalizes across cells to correct for differences in cellular sequencing depth, and across peaks to give higher values to more rare peaks[2].

    TF-IDF normalization in Flow can be invoked in Normalization and scaling section by clicking any single cell counts data node (Figure 1).

    Figure 1. TF-IDF normalization task in Normalization and scaling section in Flow.

    To run TF-IDF normalization,

    • Click a single cell counts data node

    • Click the Normalization and scaling section in the toolbox

    • Click TF-IDF normalization

    The output of TF-IDF normalization is a new data node that has been normalized by log(TF x IDF). We can then use this new normalized matrix for downstream analysis and visualization (Figure 2).

    1. Cusanovich, D., Reddington, J., Garfield, D. et al. The cis-regulatory dynamics of embryonic development at single-cell resolution. Nature 555, 538–542 (2018).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Normalize to baseline

    If your experimental design includes a sample or a group of samples serving as a baseline control, you can normalize the experimental samples by subtracting or dividing by the baseline sample(s) using the Normalize to baseline task in Partek Flow. For example, in PCR experiments, the delta Ct values of control samples are subtracted from the delta Ct values of experimental samples to obtain delta-delta Ct values for the experimental samples.

    The Normalize to baseline option is available in the Normalization and Scaling section of the context-sensitive menu (Figure 1) upon selection of any count matrix data node.

    Figure 1. When a data node containing a count matrix is selected, Normalize to baseline is available in the toolbox

    There are three options to choose the baseline samples:

    • use all samples

    • use a group

    • use matched pairs

    To normalize data to all the samples, choose to calculate the baseline using the mean or median of all samples for each feature, and choose to subtract baseline or ratio to baseline for the normalization method (Figure 2), and click Finish.

    Use a group to create baseline

    When there is a subset of samples that serve as the baseline in the experiment, select use group for Choose baseline samples. The specific group should be specified using sample attributes (Figure 3).

    Choose use group, select the attribute containing the baseline group information, e.g. Treatment in this example, with the samples with the group Control for the Treatment attribute used as the baseline. The control samples can be filtered out after normalization by selecting the Remove baseline samples after normalization check box.

    When using matched pairs, one sample from each pair serves as the control. An attribute specifying the pairs must be selected in addition to an attribute designating which sample in each pair is the baseline sample (Figure 4).

    After normalization, all values for the control sample will be either 0 or 1 depending on the normalization method chosen, so we recommend removing baseline samples when using matched pairs.

    The output of Normalize to baseline is a Normalized counts data node.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    SVD

    To analyze scATAC-seq data, Partek Flow introduced a new technique - LSI (latent semantic indexing )[1]. LSI combines steps of frequency-inverse document frequency (TF-IDF) normalization followed by singular value decomposition (SVD). This returns a reduced dimension representation of a matrix. Although SVD and Principal components analysis (PCA) are two different techniques, the SVD has a close connection to PCA. Because PCA is simply an application of the SVD. For users who are more familiar with scRNA-seq, you can think of SVD as analogous to the output of PCA. And similarly, the statistical interpretation of singular values is in the form of variance in the data explained by the various components. The singular values produced by the SVD are in order from largest to smallest and when squared are proportional the amount of variance explained by a given singular vector.

    SVD task in Flow can be invoked in Exploratory analysis section by clicking any single cell counts data node (Figure 1). We recommend running SVD on the normalized data, particularly the TF-IDF normalized counts for scATAC-seq analysis.

    Figure 1. SVD task in Flow

    To run SVD task,

    • Click a single cell counts data node

    • Click the Exploratory analysis section in the toolbox

    • Click SVD

    The GUI is simple and easy to understand. The SVD dialog is only asking to select the number of singular values to compute (Figure 2). By default 100 singular values will be computed if users don't want to compute all of them. However, the number could be adjusted manually or typed in directly. Simply click the Finish button if you want to run the task as default.

    The task report for SVD is similar to PCA. Its output will be used for downstream analysis and visualization, including Harmony (Figure 3).

    1. Cusanovich, D., Reddington, J., Garfield, D. et al. The cis-regulatory dynamics of embryonic development at single-cell resolution. Nature 555, 538–542 (2018).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Annotate Variants

    The Annotate variants task in Partek Flow provides a means to add information with regards to genomic features, such as transcript models, and existing variant databases to the variants contained in the projects. This information can be useful for filtering, interpreting, and prioritizing variants for downstream investigation. The Annotate variants task can be invoked from any Variants or Annotated variants data node, and the task will be added to and supplement any existing annotation in the underlying vcf files. Annotation information will also be visible in the downstream View variants Variant report.

    Annotate variants dialog

    The task dialog for Annotate variants contains three sections: Assembly, Annotate with genomic features, and Annotate with known variants (Figure 1). If variant detection was performed in Partek Flow, the Assembly will be displayed as text in the section, and you do not have the option to change the reference. In the event that variant detection was performed outside of Partek Flow, you will need to select the appropriate Assembly utilized for variant detection in the drop-down list. Assemblies previously added to library files (see Library File Management) will be available for selection or New assembly… can be utilized to import the reference sequence from within the task.

    Figure 1. Compoents of the annotate variants dialog

    Selecting Annotate with genomic feature provides the means to add gene/feature information to the variants (Figure 2). This typically takes the form of overlaying a transcript model (such as Ensembl). Annotation models previously added to library files (see ) will be available for selection or Add annotation model in the drop-down list can be utilized to import an annotation model to library files within the task. Promoter upstream limit and Promoter downstream limit provides a means to set the number of bases flanking the transcription start site, and this region will considered the promoter of a feature.

    Selecting Annotate with known variants will provide the ability to specify a Variant annotation database (Figure 2). Known variant databases in vcf format, such as dbSNP1 and 1000 Genomes2 for human variants, can be used in the task. Additional databases not provided for automated download in Partek® Flow®, such as the Catalogue of Somatic Mutations in Cancer (COSMIC)3, can be obtained and employed by the user. Variant databases previously added to library files (see ) will be available for selection or Add variant database in the menu can be utilized to import the variant database to library files from within the task.

    1. Sherry ST. dbSNP: the NCBI database of genetic variation. Nucleic Acids Research. 2001;29(1):308-311. doi:10.1093/nar/29.1.308

    2. Auton A, Abecasis GR, Altshuler DM, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68-74. doi:10.1038/nature15393.

    3. Forbes SA, Bhamra G, Bamford S, et al. The Catalogue of Somatic Mutations in Cancer (COSMIC). In: Haines JL, Korf BR, Morton CC, Seidman CE, Seidman JG, Smith DR, eds. Current Protocols in Human Genetics. Hoboken, NJ, USA: John Wiley & Sons, Inc.; 2008. http://doi.wiley.com/10.1002/0471142905.hg1011s57.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    FreeBayes

    A very popular variant detection approach that performs well in many situations, FreeBayes (version 1.0.1) employs a Bayesian statistical framework to determine the most likely combination of genotypes in the sample(s) at each position in a reference genome for any number of individuals from a population. It is haplotype-based, calling variants based on the literal sequences of reads aligned to a particular target and not their precise alignment. This method can identify both single nucleotide variants and insertions/deletion events. Information on the model underlying the variant detection are detailed by Garrison et al.1

    FreeBayes dialog

    Selecting FreeBayes from the context sensitive menu will bring up the Freebayes task dialog (Figure 1), which contains two sections: Select Reference sequence and Advanced options.

    Figure 1. Components of the FreeBayes task dialog

    Select Reference sequence will specify the reference assembly to utilize for variant detection. If the alignment was generated in Partek Flow, the Assembly will be displayed as text in the section, and you do not have the option to change the reference. In the event that alignment was performed outside of Partek Flow, you will need to select the appropriate Assembly utilized for alignment in the drop-down list. Assemblies previously added to library files (see Library File Management) will be available for selection or New assembly… can be utilized to import the reference sequence to library files from within the task.

    Advanced options provides a means to tune parameters in the variant detection for optimal performance. Upon invoking the task dialog, Option set is set to Default, and these parameters are provided by the FreeBayes developers. Clicking Configure button will open a window to tune advanced options. Freebayes has advanced options for Population model, Allele scope, Indel realignment, Input filters, Mappability priors, Genotype likelihoods, Algorithmic features, and Report options. Moving the mouse cursor over the info button will provide details for each parameter. Please refer to the for further information on tuning these parameters.

    1. Garrison E, Marth G. Haplotype-based variant detection from short-read sequencing. July 2012. .

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Variant Analysis

    Variations in nucleotide sequence, in the form of single nucleotide variants (SNVs) and insertion and deletion events (INDELs), can either be neutral in nature or can have functional effects. Partek Flow provides all the tools necessary to interrogate and prioritize variants for further analysis. Variants stored in Variant Call Format (vcf) files can be analyzed to filter, annotate, summarize, visualize, and validate your panel of identified variants. Multiple vcf processing tools are available under the Variant analysis section of the context sensitive menu

    • Fusion Gene Detection

    • Annotate Variants

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    LoFreq

    LoFreq (version 2.1.3a) is a very sensitive and fast variant caller that can be employed to robustly call low-frequency variants. Uilizing sources of sequencing error in the detection model, LoFreq can identify variants below the sequencing error rate. The significance of each variant is calculated to allow for control of false positives. This method can identify both single nucleotide variants and insertions/deletion events, although the current implementation does not produce discrete genotype calls. Information on the model underlying the variant detection is detailed by Wilm et al.1

    LoFreq dialog

    Selecting LoFreq from the context sensitive menu will bring up the LoFreq task dialog (Figure 1), which contains two sections: Select Reference sequence and Advanced options.

    Figure 1. Components of the LoFreq task dialog

    Select Reference sequence will specify the reference assembly to utilize for variant detection. If the alignment was generated in Partek Flow, the Assembly will be displayed as text in the section, and you do not have the option to change the reference. In the event that alignment was performed outside of Partek Flow, you will need to select the appropriate Assembly utilized for alignment in the drop-down list. Assemblies previously added to library files (see Library File Management) will be available for selection or New assembly… can be utilized to import the reference sequence to library files from within the task.

    Advanced options provides a means to tune parameters in the variant detection for optimal performance. Upon invoking the task dialog, Option set is set to Default, and these parameters are provided by the LoFreq developers. Clicking Configure will open a window to tune advanced options. LoFreq has advanced options for Region control, Base-call quality, Base-alignment (BAQ) and indel-alignment (IDAQ) qualities, Mapping quality, Indels, Source quality, P-values, and Other. Moving the mouse cursor over the info button will provide details for each parameter. Please refer to the for further suggestions on tuning these parameters.

    1. Wilm A, Aw PPK, Bertrand D, et al. LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets. Nucleic Acids Res. 2012;40(22):11189-11201.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Pool cells

    • Running Pool cells

    • Output of Pool cells

    Pool cells combines RNA-Seq data from all cells of a particular cell type classification for each sample. In essence, Pool cells creates virtual bulk RNA-Seq data from single cell RNA-Seq data. Because it is virtual bulk RNA-Seq data, all the same tasks that can be performed on bulk RNA-Seq gene counts data in Partek Flow can be performed on the output of Pool cells.

    Pool cells makes it easy to compare gene expression for a cell type of interest between experimental groups.

    Running Pool cells

    Before running Pool cells, you must classify the cells. To run Pool cells, select the data node with your classified cells and select Pool cells from the QA/QC section of the task menu (Figure 1).

    Options for Pool cells are Sum, Maximum, Mean, and Median. Expression values for cells from the same sample with the same cell type classification will be merged using the chosen pooling method (Figure 2). Sum is selected by default. After choosing a pooling method, select Finish to run the Pool cells task.

    Pool cells generates a counts data node for each classified cell type in the data set (Figure 3).

    Each counts data node is equivalent to simulated bulk RNA-Seq counts data for a cell type. The same tasks that can be performed on bulk RNA-Seq counts data can be performed on Pool cells output data nodes, including normalization, filtering, PCA, and differential expression analysis.

    The counts data of a cell type for each sample can be downloaded by clicking the counts data node and selecting Download data from the task menu. The counts data text file lists each sample and its pooled counts values (sum, maximum, mean, or median) for each feature (gene/transcript) in alphabetical order (Figure 4).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Troubleshooting

    If a Partek Flow task fails (no report is produced), please follow the directions in Reporting a problem.

    If the task report is produced, but the results are missing for some features represented by "?" (Figure 1), it may be because something went wrong with the estimation procedure. To better understand this, use the information available in the Extra details report (Figure 1). This type of information is present for many tasks, including Differential Analysis and Survival Analysis.

    Figure 1. Use “View extra details report” button to see why the results are missing.

    Click the Extra details report for the feature of interest (Figure 1). This will display the Extra details report (Figure 2). When the estimation procedure fails, a red triangle will be present next to the information criteria value. Hover over the triangle to see a detailed error message.

    Figure 2. . The error message shows up if one hovers over the colored triangle

    In many cases, estimation failure is due to low expression, filter out low expression features or choose a reasonable normalization method will resolve this issue.

    Sometimes the estimation results are not missing but the reported values look inadequate. If this is the case, the Extra details report may show that the estimation procedure generated a warning, and the triangle is yellow. To remove suspicious results in the report, set Use only reliable estimation results to Yes in the Advanced Options (Figure 3). The warnings will then be treated the same way as estimation failures.

    To see the results for as many features as possible, regardless of how reliable they are, set Use only reliable estimation results to No and the result will be reported unless there is an estimation failure. For example, DESeq2 uses Cook’s distances to flag features with outlying expression values and if “Use reliable results” is set to Yes (Figure 3) the p-values for such features are not reported which may lead to some missing values in the report (set Use only reliable estimation results to No to avoid this).

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Variant Callers

    Variations in nucleotide sequence, in the form of single nucleotide variants (SNVs) and insertion and deletion events (INDELs), can exist within the germline or can be acquired by somatic alterations. Partek Flow provides pipeline creation tools to identify both SNVs and INDELs using aligned reads generated from targeted, whole exome, or whole genome DNA-Seq (or RNA-Seq) data. Detection of these variants can be performed by comparison against either the reference sequence utilized for alignment or among paired samples in a project. Tools for variant detection are performed on either Aligned reads or Filtered reads data nodes (Figure 1), and the Detect variants task node will produce a Variants data node. The Variants data node will contain Variant Call Format (vcf) files for each sample in the project. Three detection tools, each employing unique algorithms to identify variants in aligned sequence data, are available under the Variant callers section of the context sensitive menu:

    • SAMtools

    Figure 1. Showing Variant callers from an aligned reads node

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    Trajectory Analysis

    Cells undergo changes to transition from one state to another as part of development, disease, and throughout life. Because these changes can be gradual, trajectory analysis attempts to describe progress through a biological process as a position along a path. Because biological processes are often complex, trajectory analysis builds branching trajectories where different paths can be chosen at different points along the trajectory. The progress of a cell along a trajectory from the starting point or root, can be quantified as a numeric value, pseudotime.

    Partek Flow offers Monocle 2 and Monocle 3 methods.

    • Trajectory Analysis (Monocle 2)

    • Trajectory Analysis (Monocle 3)

    Difference Between Monocle 3 and Monocle 2

    Major updates in Monocle 3 (compared to Monocle 2) include:

    • Monocle 3 learns the principal trajectory graph in the UMAP space;

    • the principal graph is smoothened and small branches are excluded;

    • support for principal graphs with loops and convergence points;

    • support for multiple root nodes.

    If you need additional assistance, please visit to submit a help ticket or find phone numbers for regional support.

    arrow_down_icon_collapse_triangle_gray
    arrow_right_icon_expand_triangle_gray

    Task Menu

    The Task Menu lists all the tasks that can be performed on a specific node. It can be invoked from either a Data or Task node and appears on the right hand side of the Analyses tab. It is context-sensitive, meaning that it will only present tasks that the user can perform on the selected node. For example, selecting an Aligned reads data node will not present aligners as options.

    Clicking a Data node presents a variety of tasks:

    Pre-alignment QA/QC

    Selecting a node with unaligned reads (either Unaligned reads or Trimmed reads) shows the QA/QC section in the context sensitive menu, with two options (Figure 1). To assess the quality of your raw reads, use Pre-alignment QA/QC.

    Pre-alignment QA/QC setup dialog is given in Figure 2. Examine reads allows you to control the number of reads processed by the tool; All reads, or a subset (One of every n reads). The latter option is obviously not as thorough, but is much faster than All reads.

    If selected, K-mer length creates a per-sample report with the position of the most frequent k-mers (i.e. sequences of k nucleotides) of the length specified in the dialog. The range of input values is from one to 10.

    The last control refers to .fastq files. Partek® Flow® can automatically detect the quality encoding scheme (Auto detect) or you can use one of the options available in the drop-down list. However, the auto-detection is only applicable for Phred+33 and Phred+64 type of quality encoding score. For early version of Solexa quality encoding score, select Solexa+64 from the Quality encoding drop down list. For a paired-end data, the pre-alignment QA/QC will be done on each read in pair separately and the results will be shown separately as well.

    Single-cell QA/QC

    The Single-cell QA/QC task in Partek Flow enables you to visualize several useful metrics that will help you include only high-quality cells. To invoke the Single-cell QA/QC task:

    • Click a Single cell counts data node

    • Click the QA/QC section of the task menu

    • Click Single cell QA/QC

    Cell barcode QA/QC

    The Cell barcode QA/QC task lets you determine whether a given cell barcode is associated with a cell. This is an important QC step in all droplet-based single cell RNA-seq experiments, such as Drop-seq, where all barcodes are sequenced.

    To invoke Cell barcode QA/QC:

    Harmony

    It is challenging to analyze scRNA-seq data, particularly when they are assayed with different technologies. Because biological and technical differences are interspersed. Harmony[1] is an algorithm that projects cells into a shared embedding where cells group by cell type rather than dataset-specific conditions. Harmony is able to simultaneously account for multiple experimental and biological factors while integrating different datasets.

    Harmony in Flow can be invoked in Batch removal section only if

    1. the data has some categorical attributes (only categorical attributes can be included in the model)

    2. PCA data node is selected (Figure 1).

    To run Harmony,

    t-SNE

    t-Distributed Stochastic Neighbor Embedding (t-SNE) is a dimensional reduction technique [1]. t-SNE aims to preserve the essential high-dimensional structure and present it in a low-dimensional representation. t-SNE is particularly useful for visually identifying groups of similar samples or cells in large high-dimensional data sets such as single cell RNA-Seq.

    We recommend normalizing your data prior to running t-SNE, but the task will run on any counts data node.

    • Click the counts data node

    • Click the Exploratory analysis section of the toolbox

    K-means Clustering

    K-means clustering is a method for identifying groups of similar observations, i.e. cells or samples. K-means clustering aims to group observations into a pre-determined number of clusters (k) so that each observation belongs to the cluster with the nearest mean. An important aspect of K-means clustering is that it expects clusters to be of similar size (equal variance) and shape (distribution of variance is spherical). The Compare Clusters task can also be used to help determine the optimal number of K-means clusters.

    We recommend normalizing your data prior to running K-means clustering, but the task will run on any counts data node.

    • Click the counts data node

    • Click the Exploratory analysis section of the toolbox

    Seurat3 integration

    Seurat v3[1] introduced new methods for the integration of multiple single-cell datasets, no matter whether they were collected from different individuals, experimental conditions, technologies, etc. Seurat 3 integration method aims to use a subset of the data as reference for the integrate analysis. The method integrates all other data with the reference subset. The subset can be one sample or a subgroup of samples defined by the factor attribute.

    Seurat3 integration in Flow can be invoked in Batch removal section if a Normalized counts data node is selected (Figure 1).

    To run Seurat3 integration,

    • Click a Normalized counts data node

    Transcript Expression Analysis - Cuffdiff

    This option is only available when Cufflinks quantification node is selected. Detailed implementation information can be found in the Cuffdiff manual [5].

    When the task is selected, the dialog will display all the categorical attributes more than one subgroups (Figure 1).

    When an attribute is selected, pairwise comparisons of all the levels will be performed independently.

    Click on Configure button in the Advanced options to configure normalization method and library types (Figure 2).

    There are three library normalization methods:

    • Class-fpkm: library size factor is set to 1, no scaling applied to FPKM values

    AUCell

    SAMtools

    SAMtools1 (version 1.2) utilizes the mpileup command to look at observed bases in the reads covering every genomic position represented in the aligned sequence data and calculate the likelihood of every possible genotype at a locus. Subsequently, bcftools applies the prior probability and uses Bayesian inference to call actual genotypes, outputting variant information in Variant Call Format (vcf). This method can identify both single nucleotide variants and insertions/deletion events. General information about the underlying algorithm utilized by SAMtools is detailed by Li. 2,3

    Selecting SAMtools from the context sensitive menu will bring up the SAMtools task dialog, which contains three default sections: Variant detection method, Select Reference sequence, and Advanced options.

    In the Variant detection method drop-down list, Against reference will compare base composition for each sample against the reference sequence assembly, independently (Figure 1).

    In the event paired samples exist within the project, detection Paired samples can be utilized to identify loci with differing genotypes between the pair once each sample has been compared to the reference sequence assembly. In instances where there is limited information to accurately determine genotypes in one or both of the samples, the same genotype may be called for case and control if it differs from the reference. The Filter variants task can be used to exclude these spurious loci. To perform this analysis, sample attributes must be added in the Data tab of the project (Figure 2). Specifically, an attribute must be added for sample ID (shared between the paired samples) and an attribute must also be added for sample type that differentiates the paired samples.

    Find multimodal neighbors

    Multi-omics single cell analysis is based on simultaneous detection of different types of biological molecules on the same cells. Common multi-omics techniques include feature barcoding or CITE-seq (cellular indexing of transcriptomes and epitopes by sequencing) technologies, which enable parallel assessment of gene and protein expression. Specific bioinformatics tools have been developed to enable scientists to integrate results of multiple assays and learn relative importance of each type (or each biological molecule) in identification of cell types. Partek Flow supports weighted nearest neighbor (WNN) analysis (1), which can help combine output of two molecular assays.

    This task can only be performed on data nodes containing PCA scores – which are PCA output and graph based clustering output nodes generated from PCA nodes. To start, select a PCA data node of one of the assays (e.g. gene expression) and go to

    Annotate Variants (SnpEff)

    An important aspect of variant analysis is the ability to prioritize specific variants for further investigation. As variant detection can often identify a large number of variants, it may be difficult to determine which variants may impact phenotypes. SnpEff (version 4.1k) provides a means to annotate and predict the effects of variants on genes, allowing for prioritization of variants within the project. In addition, the SnpEff databases utilized for prediction support a large number of genome assemblies. Information regarding the implementation of the predictions is detailed by Cingolani et al.1 The predicted effect of the variant is categorized by impact:

    • HIGH - frame shifts, addition/deletion of stop codons, etc;

    • MODERATE – codon change/deletion/insertion, etc;

    Annotate Variants (VEP)

    An important aspect of variant analysis is the ability to prioritize variants for downstream analysis. As variant detection can often identify a large number of variants, it may be difficult to determine which variants may impact phenotypes. As implemented in Partek Flow, the Ensembl Variant Effect Predictor (VEP, version 84)1 provides a means to add detailed annotation to variants in the project such as discrete aspects of transcript models and variant databases not available in the task. For variants identified in human data, information from popular tools that predict the impact of variants that cause amino acid changes, SIFT2–4 and PROVEAN5 (available for the hg19 genome assembly), will be included. VEP databases can be obtained for multiple species, and content will be dependent on available transcript and variant information for that organism. The Annotate variants (VEP) task can be invoked from any Variants or Annotated variants data node, and the task will supplement any existing annotation in the vcf files. Annotation information will also be visible in the downstream View variants Variant report .

    The task dialog for Annotate variants (VEP) contains two sections: Select Variant Effect Predictor database and Advanced options (Figure 1). Select Variant Effect Predictor database will specify the reference assembly to utilize for variant detection. If the variant detection was performed in Partek Flow, the Assembly will be displayed as text in the section. Upon initial task usage, click the Create variant effect predictor database button to import a database. The VEP database for hg19 is available for automated download in Partek Flow, and information regarding obtaining additional databases for other species and genome assemblies can be found in the

    Additional Assistance

    Annotate Variants (SnpEff)
    Annotate Variants (VEP)
    Filter Variants
    Summarize Cohort Mutations
    Combine Variants
    our support page

    Additional Assistance

    our support page

    Feature distribution plot configuration

    Filter

    Plot type

    Page

    Scale Y-axis

    Color by

    List management
    Figure 1. The feature distribution plot shows each feature as a row. Features are ordered by average value in descending order.
    Figure 2. Filtering features on the Feature distribution plot
    Figure 3. Histogram (top) and Strip (bottom) plots
    Figure 4. Histogram tooltip
    Figure 5. Strip tooltip
    Figure 6. Selecting a on one strip highlights it for all features.
    Figure 7. Viewing the median for a feature
    Figure 8. Setting items per page
    Figure 9. Splitting histograms by an attribute
    Figure 10. Coloring strips by an attribute. Categorical (top) or Numeric (bottom) attributes can be chosen.

    Input Parameters: contains the parameters used for running Trim bases function. This section tells what option has been selected for the Trim bases task. It includes all the parameters used for the task, such as minimum read length, maximum percentage of N's base, quality encoding, quality score threshold (if applicable) and how trimming is performed.

  • Command Lines: shows the commands used for running Trim bases function by the software Partek Flow

  • Advanced options

    Trim Bases Task Details Page

    Trim Bases Task Report Page

    Trim Bases Output Files

    References

    Additional Assistance

    our support page
    Figure 1. Select a Trim mode to trim poor quality bases from reads
    Figure 2. Trim bases from 5'-end
    Figure 3. Trim bases from 3'-end
    Figure 4. Trim bases from both ends
    Figure 5. Trim bases based on quality score
    Figure 6. Trim bases options. A) Trim from 3'-end; B) Trim from 5'-end; C) Trim from both ends; D) Trim based on base quality score

    Configurable options

    Basic Options

    Annotation file

    Strand specificity

    Include features with no counts

    Advanced Options

    Min qual

    Overlap mode

    Nonunique mode

    Output

    References

    Additional Assistance

    Quantify to annotation model (Partek E/M)
    Adding an Annotation Model
    HGNC
    https://help.partek.illumina.com/partek-flow/user-manual/task-menu/qa-qc/single-cell-qa-qc
    our support page
    Figure 1. Basic HTSeq options
    Figure 2. Advanced options for HTSeq
    Figure 3. HTSeq output

    Median: value of the mid point of a feature

  • Minimum: lowest value of a feature

  • Range: the difference between the highest and lowest values of a feature

  • Std dev.: the square root of the variance

  • Sum: total value of the feature

  • Variance: the average of the squared differences from the mean

  • Dispersion: variance divided by mean of the feature

  • >=: greater than or equal to
    Median
  • Minimum

  • Range

  • Standard deviation (std dev)

  • Sum

  • Variance

  • Dispersion

  • Noise reduction filter

    Statistics based filter

    Feature metadata filter

    Feature list filter

    Filter features task report

    Additional Assistance

    Filter features task report
    List management
    List management
    our support page
    Figure 1. Noise reduction filter
    Figure 2. Selecting value exposes additional options
    Figure 3. Statistics based filter
    Figure 4. Filter features based on feature annotation fields
    Figure 5. Feature list filter
    Figure 6. Feature filter report table
    Figure 7. Filter features report figures

    Advanced options

    ANOVA/LIMMA-trend/LIMMA-voom
    Figure 1. Adding factors and interactions to the model
    Figure 2. Configuring the batch effect removal tool
    Figure 3. Before (left) and after (right) batch effect removal of Version and Version*Classification. Cells are sized by Version and colored by Classification.

    Compare clusters task report

    Additional Assistance

    Figure 1. Compare clusters configuration dialog
    Figure 2. The Compare clusters task report shows the Davies-Bouldin index for each number of clusters.

    Report

    Additional Assistance

    our support page
    Figure 1. Select any count node to invoke the Non-parametric ANOVA task
    Figure 2. Select one factor for analysis
    Figure 3. Normalize your count data by rank to do non-parametric testing on more complicated experimental designs
    Figure 4. Set-up desired comparisons
    Figure 5. The task's ANOVA report will display the median instead of the LSmean

    Output of Pool cells

    Additional Assistance

    our support page
    Figure 1. Invoking Pool cells
    Figure 2. Choosing the pooling method
    Figure 3. Output of Pool cells. This data set had 3 cell types, Malignant, Microglia/macrophages, and Oligodendrocytes. Pool cells has creates a counts data node for each cell type.
    Figure 4. Pool cells counts data can be downloaded as a tab-delimited text file. The downloaded text file can be opened in any spreadsheet viewing program or text editor.

    Additional Assistance

    our support page
    Figure 1. PCA setup dialog
    Figure 2. Double-click on the PCA result node to open the Principal components analysis result
    Figure 3. Drag PCA data node to plot
    Figure 4. Scree plot
    Figure 5. Component loadings are the correlation coefficients between the features and PCs.
    Figure 6. PCA project table

    Additional Assistance

    our support page
    Figure 2. ExFold comparison can be performed between specified Mix 1 and Mix 2 pairs of samples
    Figure 3. Summary of ERCC assessment. Each row is a sample (an example is shown)
    Figure 4. Summary of ExFold comparison. Each row is a different pairwise comparison
    Figure 5. ERCC spike-ins plot. Lines (one per sample) correspond to regression lines between actual spike-in concentrations and observed number of alignments. Dots represent present ERCC sequences with the lowest and the highest concentration
    Figure 6. Scatter plot of actual observed alignment counts vs. probe concentration for each ERCC control within a sample. Each dot is an ERCC control, coloured by GC content and sized by concentration
    Figure 7. Table report for ERCC controls within a sample. The default sort order is by column Control; the example in the figure is sorted by Actual (Concentration) to highlight the relationship between the control concentration and number of alignments
    Figure 8. Scatter plot of actual observed Mix1:Mix2 ratios vs. expected Mix1: Mix2 ratios for each ERCC control within a sample. Each dot is an ERCC control, coloured by GC content and sized by concentration
    Figure 9. Table report for ERCC ExFold comparison within a sample
    add_plus_icon_green

    References

    Additional Assistance

    our support page

    Additional Assistance

    our support page
    Figure 3. Biomarkers table (example). Top 10 biomarkers for each cluster are shown. Download link provides the full results table

    Merge Features

    Task Output

    Alternative paths

    Additional Assistance

    Analyzing CITE-Seq Data
    Find multimodal neighbors
    our support page
    Figure 2. Using Merge cells option. The selected data node is shown under the Select data nodes button (in this example: Normalized counts). More than one node can be selected for merging
    Figure 3. Example of a Merged counts data node. The node was created by using the Merge features option of the Merge matrices task. In this example, results of two assays (gene and protein expression) were merged

    References

    Additional Assistance

    https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0947-7
    our support page
    Figure 2. Interface of Scran deconvolution task in Partek Flow. Example attributes are indicated in the drop-down list if Cluster name is checked.
    Figure 3. Example workflows to demonstrate downstream analysis and visualization of Scran deconvolution output.

    Additional Assistance

    https://doi.org/10.1101/576827
    https://www.rdocumentation.org/packages/Seurat/versions/3.1.4/topics/SCTransform
    https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02584-9
    our support page
    Figure 3. Downstream analysis after SCTransform v2 task in Flow.

    Additional Assistance

    our support page
    Figure 2. Select one or more attributes to add
    image2021-2-2_17-20-14

    References

    Additional Assistance

    https://doi.org/10.1038/nature25981
    https://satijalab.org/signac/index.html
    our support page
    Figure 2. Example workflows to demonstrate downstream analysis and visualization of TF-IDF normalization output.

    Use all samples to create baseline

    Use matched pairs

    Additional Assistance

    our support page
    Figure 2. Use the mean or median of all samples as the baseline to normalize the data
    Figure 3. Use a subgroup of samples to create baseline to normalize the data
    Figure 4. Designated pairs and the baseline sample in each pair to normalize by matched pairs

    References

    Additional Assistance

    https://doi.org/10.1038/nature25981
    our support page
    Figure 2 Interface of SVD task in Partek Flow
    Figure 3. Example workflows to demonstrate downstream analysis and visualization of SVD output for scATAC-seq data.

    References

    Additional Assistance

    Library File Management
    Library File Management
    our support page
    Figure 2. Configuration of Annotate with genomic features

    References

    Additional Assistance

    FreeBayes documentation
    https://arxiv.org/abs/1207.3907
    our support page

    References

    Additional Assistance

    LoFreq documentation
    our support page

    Additional Assistance

    our support page
    Figure 3. If only reliable results are used, features with warnings are not reported.
  • Non-unique singleton

    • fraction of singletons that align to multiple locations

  • Non-unique paired

    • fraction of paired reads that align to multiple locations

  • Additional Assistance

    our support page
    Figure 1. Post-alignment QA/QC output table for single-end sequencing
    Figure 2. Post-alignment QA/QC output table for paired-end sequencing
    Figure 3. Alignment breakdown plot. Each sample is a column. Fraction of reads with respect to their alignment status is given on the left-side y-axis and colour coded. Total number of reads is depicted by the black line and can be red on the right-side y-axis
    Figure 4. Coverage plot. Average read depth (times, in covered regions) is shown by columns and quantified on the left y-axis. Genome coverage (%) is shown by the black line and quantified on the right y-axis
    Figure 5. Average alignments per read. Each sample is a column, average number of alignments per read is on the y-axis

    Strand: the strand of the genomic feature

  • Total exon length: the length of the genomic feature

  • Reads: the total number of reads aligning to the genomic feature

  • % GC: the percentage of GC contents of those reads aligning to the genomic feature

  • % N: the percentage of ambiguous bases (N) of those reads aligning to the genomic feature

  • (n)x: the proportion of the genomic feature which is covered by at least n number of alignments. [Note: n is the coverage level that you specified when submitting Coverage report task, defaults are 1×, 20×, 100×]

  • Average coverage: the average sequencing depth across all bases in the genomic feature

  • Average quality: the average quality score across covered bases in the genomic feature

  • : the invokes the Coverage graph across the genomic feature, showing the current sample only (Figure 23) (or mouse over to get a preview of the plot)

  • : the invokes the Chromosome view and browses to the genomic location

  • Additional Assistance

    Data tab
    our support page
    Figure 1. Setting up Coverage report. The example on the figure shows the use of a custom annotation file, CRCTargets, which defines target regions for a targeted resequencing panel
    Figure 2. Coverage report table showing one sample per row and key coverage metrics on columns
    Figure 3. On-target vs. off-target chart. Each column is a sample. Fraction of reads on- or off-target is on the y-axis
    Figure 4. Region average coverage summary (truncated). Regions are on rows, samples on columns. The icon in the right-most column opens a coverage graph for that region
    Figure 5. Coverage report of a genomic region (in this example NM_000038). The horizontal axis is the normalized position within the genomic feature, represented as 1st to 100th percentile of the length of the feature. The vertical axis is coverage. Each line on the plot is a single sample, and the samples are listed below the plot
    Figure 6. Coverage summary plot with sample shown as lines, normalised genomic position on the horisontal axis and average coverage on the vertical axis
    Figure 7. Coverage report table, sample-level (truncated). Each genomic feature is a row, coverage metrics are on columns (an example is shown)
    bar_chart_icon_gray
    add_plus_icon_green
    delete_icon_x_cross_red

    Pre-alignment QA/QC

  • ERCC Assessment

  • Post-alignment QA/QC

  • Coverage Report

  • Validate Variants

  • Feature distribution

  • Single-cell QA/QC

  • Cell barcode QA/QC

  • Pre-alignment tools

    • Trim bases

    • Trim adapters

    • Filter reads

  • Post-alignment tools

    • Filter alignments

    • Convert alignments to unaligned reads

    • Combine alignments

  • Annotation/Metadata

    • Annotate cells

    • Annotation report

    • Publish cell attributes to project

  • Pre-analysis tools

    • Generate group cell counts

    • Pool cells

    • Split matrix

  • Aligners

  • Quantification

    • Quantify to annotation model (Partek E/M)

    • Quantify to transcriptome (Cufflinks)

    • Quantify to reference (Partek E/M)

  • Filtering

    • Filter features

    • Filter groups (samples or cells)

    • Filter barcodes

  • Normalization and scaling

    • Impute low expression

    • Impute missing values

    • Normalization

  • Batch removal

    • General linear model

    • Harmony

    • Seurat3 integration

  • Differential Analysis

    • GSA

    • ANOVA/LIMMA-trend/LIMMA-voom

    • Kruskal-Wallis

  • Survival Analysis with Cox regression and Kaplan-Meier analysis - Partek Flow

  • Exploratory Analysis

    • Graph-based Clustering

    • K-means Clustering

    • Compare Clusters

  • Trajectory Analysis

    • Trajectory Analysis (Monocle 2)

    • Trajectory Analysis (Monocle 3)

  • Variant Callers

    • SAMtools

    • FreeBayes

    • LoFreq

  • Variant Analysis

    • Fusion Gene Detection

    • Annotate Variants

    • Annotate Variants (SnpEff)

  • Copy Number Analysis (CNVkit)

  • Peak Callers (MACS2)

  • Peak analysis

    • Annotate Peaks

    • Promoter sum matrix

  • Motif Detection

  • Metagenomics

    • Kraken

    • Alpha & beta diversity

    • Choose taxonomic level

  • 10x Genomics

    • Cell Ranger - Gene Expression

    • Cell Ranger - ATAC

    • Space Ranger

  • V(D)J Analysis

  • Biological Interpretation

    • Gene Set Enrichment

    • GSEA

  • Correlation

    • Correlation analysis

    • Sample Correlation

    • Similarity matrix

  • Export

  • Classification

  • Task actions

  • Feature linkage analysis

  • Clicking a Task node gives you the option to view the Task results or perform Task actions such as rerunning the task (Figure 1).

    Figure 1. Task menu invoked from a Task node

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Data summary report
    QA/QC

    Additional Assistance

    Figure 2. Pre-alignment QA/QC setup dialog (defaults)

    Most sequencing applications now use the phred quality score. This score indicates the probability that the base was accurately identified. The table below shows the corresponding base call accuracies for each score:

    Phred Quality Score
    Base Call Accuracy

    10

    90%

    20

    99%

    30

    The task report is organised in two tiers. The initial view shows project-level report with all the samples. An overview table is at the top, while matching plots are below.

    The Pre-alignment QA/QC output table contains one input file per row, with typical metrics on columns (%GC: fraction of GC content; %N: fraction of no-calls) (Figure 3). The file names are hyperlinks, leading to the sample-level reports. To save the table as a txt file to a local computer, push the Download link. Table columns can be sorted using double arrows icon ( ).

    Figure 3. Pre-alignment QA/QC output table (project-level). Each row is an input file. %N: proportion of no-calls, %GC: GC content

    Two project-level plots are Average base quality per position and Average base quality score per read (Figure 4). The latter plot presents the proportion of reads (y-axis) with certain average quality score (meaning all the base qualities within a read are averaged; x-axis). Mouse over a data point to get the matching readouts. The Save icon saves the plot in a .svg format to the local machine. Each line on the plot represents a data file and you can select the sample names from the legend to hide/un-hide individual lines.

    Figure 4. Pre-alignment QA/QC project-level plots (each line is a file)

    A sample-level report begins with a header, which is a collection of typical quality metrics (Figure 5).

    Figure 5. Header of a sample-level pre-alignment QA/QC report

    Below the header you will find four plots: Base composition, Average base quality score per position (same as above, but on the sample level), Distribution of base quality scores (the same as Average base quality score per read, but on the sample level), and Distribution of read lengths.

    Base composition plot specifies relative abundance of each base per position (Figure 6), with N standing for no-calls. By selecting individual bases on the legend, you can remove them from the plot / bring them back on. To zoom in, left-click & drag over a region of interest. To zoom out, use the Reset button ( ) to recreate the original view, or the magnifier glass ( ) to zoom out one level.

    Figure 6. Base composition plot: fraction of each base at a given position within a read. N: no call

    Distribution of read lengths shows a single column for fixed length data (e.g. Illumina sequencing). However, for quality-trimmed data or non-fixed length data (like Ion Torrent sequencing), expect to see a read’s length distribution (Figure 7).

    Figure 7. Distribution of read lengths (an example processed by Ion Torrent sequencer is shown)

    If K-mer length option was turned on when setting up the task, an additional plot will be added to the sample-level report, i.e. K-mer Content (Figure 8). For each position, K-mer composition is given, but only the top six most frequent K-mers are reported; high frequency of a K-mer at a given site (enrichment) indicates a possible presence of sequencing adapters in the short reads.

    Figure 8. K-mer Content plot. Position of a K-mer is on the horizontal axis, while the K-mer frequency is on the y-axis. Top six K-mers are listed below the plot. In this example, the value of K was set to eight. The most frequent reported K-mer (CTGTCTCT) is a reverse complement of a commonly used adapter (AGAGACAG)

    The pre-alignment QA/QC report as described above is generally available for the NGS data of fastq format. For other types of data, the report may differ depending on the availability of information. For example, for fasta format, there is no base quality score information and therefore all the figures or graphs related to base or read quality score will be unavailable.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. QA/QC options on unaligned reads

    Additional Assistance

    By default, all samples are used to perform QA/QC. You can choose to split the sample and perform QA/QC separately for each sample.

    If your Single cell counts data node has been annotated with a gene/transcript annotation, the task will run without a task configuration dialog. However, if you imported a single cell counts matrix without specifying a gene/transcript annotation file, you will be prompted to choose the genome assembly and annotation file by the Single cell QA/QC configuration dialog (Figure 1). Note, it is still possible to run the task without specifying an annotation file. If you choose not to specify an annotation file, the detection of mitochondrial counts will not be possible.

    Figure 1. If an annotation was not specified on import, you can choose the assembly and annotation after invoking Single-cell QA/QC

    The Single cell QA/QC task report opens in a new data viewer session. Four dot and violin plots showing the value of every cell in the project are displayed on the canvas: counts per cell, detected features per cell, the percentage of mitochondrial counts per cell, and the percentage of ribosomal counts per cell (Figure 2).

    Figure 2. Dot & violin plots showing values of each cell for several quality measures

    If your cells do not express any mitochondrial genes or an appropriate annotation file was not specified, the plot for the percentage of mitochondrial counts per cell will be non-informative (Figure 3).

    Figure 3. A non-informative percentage of mitochondrial counts plot

    Mitochondrial genes are defined as genes located on a mitochondrial chromosome in the gene annotation file. The mitochondrial chromosome is identified in the gene annotation file by having "M" or "MT" in its chromosome name. If the gene annotation file does not follow this naming convention for the mitochondrial chromosome, Partek Flow will not be able to identify any mitochondrial genes. If your single cell RNA-Seq data was processed in another program and the count matrix was imported into Partek Flow, be sure that the annotation field that matches your feature IDs was chosen during import; Partek Flow will be unable to identify any mitochondrial genes if the gene symbols in the imported single cell data and the chosen gene/feature annotation do not match.

    Ribosomal genes are defined as genes that code for proteins in the large and small ribosomal subunits. Ribosomal genes are identified by searching their gene symbol against a list of 89 L & S ribosomal genes taken from HGNC. The search is case-insensitive and includes all known gene name aliases from HGNC. Identifying ribosomal genes is performed independent of the gene annotation file specified.

    Total counts are calculated as the sum of the counts for all features in each cell from the input data node. The number of detected features is calculated as the number of features in each cell with greater than zero counts. The percentage of mitochondrial counts is calculated as the sum of counts for known mitochondrial genes divided by the sum of counts for all features and multiplied by 100. The percentage of ribosomal counts are calculated as the sum of counts for known ribosomal genes divided by the sum of counts for all features and multiplied by 100.

    Each point on the plots is a cell. All cells from all samples are shown on the plots. The overlaid violins illustrate the distribution of cell values for the y-axis metric.

    The appearance of a plot can be configured by selecting a plot and adjusting the Configure settings in the panel on the left (Figure 4). Here are some suggestions, but feel free to explore the other options available:

    • Open Axes and change the Y-axis scale to Logarithmic. This can be helpful to view the range of values better, although it is usually better to keep the Ribosomal counts plot in linear scale.

    • Open Style and reduce the Color Opacity using the slider. For data sets with very many cells, it may be helpful to decrease the dot opacity to better visualize the plot density.

    • Within Style switch on Summary Box & Whiskers. Inspecting the median, Q1, Q3, upper 90%, and lower 10% quantiles of the distributions can be helpful in deciding appropriate thresholds.

    Figure 4. Use the Configuration options on the left customize the appearance of plots

    High-quality cells can be selected using Select & Filter, which is pre-loaded with the selection criteria, one for each quality metric (Figure 5).

    Figure 5. Select & Filter criteria

    Hovering the mouse over one of the selection criteria reveals a histogram showing you the frequency distribution of the respective quality metric. The minimum and maximum thresholds can be adjusted by clicking and dragging the sliders or by typing directly into the text boxes for each selection criteria (Figure 6).

    Figure 6. Use the sliders and histograms to adjust the selection criteria

    Alternatively, Pin histogram to view all of the distributions at one time to determine thresholds with ease (Figure 7).

    Figure 7. Pin the histogram to view all of the distributions at one time

    Adjusting the selection criteria will select and deselect cells in all three plots simultaneously. Depending on your settings, the deselected points will either be dimmed or gray. The filters are additive. Combining multiple filters will include the intersection of the three filters. The number of cells selected is shown in the figure legend of each plot (Figure 8).

    Figure 8. The filters are additive and deselected cells are dimmed

    To filter the high-quality cells, click the include selected cells icon in Filter in the top right of Select & Filter, and click Apply observation filter... (Figure 9).

    Figure 9. Include the selected points and apply the filter

    Select the input data node for the filtering task and click Select (Figure 10).

    Figure 10. After the Apply filter button is selected, you will be presented with a preview of your pipeline. You need to select the appropriate data node to apply the filtering to

    A new data node, Filtered counts, will be generated under the Analyses tab (Figure 11).

    Figure 11. Filter cells task runs from the Single cell QA/QC report

    Double click the Filtered counts data node to view the task report. The report includes a summary of the count distribution across all features for each sample; a detailed breakdown of the number of cells included in the filter for each sample; and the minimum and maximum values for each quality metric (expressed genes, total counts, etc) across the included cells for each sample (Figure 12).

    Figure 12. Filtered counts task report

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Additional Assistance

  • Click a Single cell counts data node

  • Click the QA/QC section of the task menu

  • Click Cell barcode QA/QC

  • The task can be performed with or without the EmptyDrops method enabled.

    To perform the task without the EmptyDrops method enabled, leave the checkbox unchecked and click Finish (Figure 1).

    Figure 1. Leave the checkbox unchecked to run the task without the EmptyDrops method

    The Cell barcode QA/QC task report is a plot (Figure 2). Barcodes are ordered on the X-axis by the number reads such that the barcode closest to the Y-axis has the most reads and the barcode furthest from the Y-axis has the fewest reads. The Y-axis value is the number of mapped reads corresponding to each barcode. This type of plot is often referred to as a knee plot.

    Figure 2. Cell barcode QA/QC task report is used to filter barcodes

    The knee plot is used to choose a cutoff point between barcodes that correspond to cells and barcodes that do not. Partek Flow automatically calculates an inflection point, shown by the vertical line on the graph. Barcodes designated as cells are shown in blue while barcodes designated as without cells (background) are shown in grey.

    The cutoff can be adjusted by dragging the vertical line across the graph or by using the text fields in the Filter panel on the left-hand side of the plot. Using the Filter panel, you can specify the number of cells or the percentage of reads in cells and the cutoff point will be adjusted to match your criteria. The number of cells and the percentage of counts in cells is adjusted as the cutoff point is changed. To return to the automatically calculated cutoff, click Reset sample filter.

    The percentage of counts in cells and median counts per cell are useful technical quality metrics that can be consulted when optimizing sample handling, cell isolation techniques, and library preparation.

    One knee plot is generated for each sample. In projects with multiple samples, Next and Back buttons will appear at the top to enable navigation between sample knee plots. Manual filters must be set separately for each sample. This is typically used when the user expects a certain number of cells to be processed, like in experiments where droplets were loaded with a predefined number of cells.

    To view a summary of the currently selected filter settings for all samples, click Summary table. This opens a table showing key metrics for each sample in the project (Figure 3).

    Figure 3. Barcode QA/QC summary table lists filtering information for each sample

    To return to the knee plot view, click Back to filter. To apply the filter and run the Filter barcodes task, click Apply filter. A Filtered counts data node will be generated.

    The EmptyDrops method (1) uses a statistical test to identify which barcodes correspond to real cells and empty droplets. An ambient RNA expression profile is estimated from barcodes below a specified total UMI count threshold, using the Good-Turing algorithm. The expression profile of each barcode above the low-count threshold is then tested for deviations from the ambient profile. Real cells are expected to have a low p-value, indicating a significant deviation from the expected background noise level. False discovery rate (FDR) correction is applied to all the p-values and those falling equal to or below the specified FDR level are detected as real cells. This can allow for the detection of additional cells that would otherwise be discarded due to a low total UMI count.

    This method requires empty barcodes to be present in the single cell count matrix, in order to estimate the ambient RNA profile. If your data has already been filtered to remove barcodes with low total counts, this method will not be suitable. For example, if you are working with 10X Genomics data, the EmptyDrops method can only be run on the raw counts, not the filtered counts.

    In addition, a knee point threshold will be calculated to identify cells with a very high total UMI count. It's possible that some barcodes with a high total UMI count will not pass the EmptyDrops significance test. This could be due to biases in the ambient RNA profile, leading to a non-significant difference between a barcode's expression profile vs the ambient profile. To protect against this issue, it is advisable to use the EmptyDrops results in conjunction with the knee point filter, on the assumption that barcodes with a very high total UMI count will always correspond to real cells. Note, the knee point will be more conservative than the inflection point calculated by Partek Flow when the EmptyDrops method is not enabled.

    To perform the task with the EmptyDrops method, check the checkbox, configure the additional options, and click Finish (Figure 4)

    Figure 4. Check the box to run the task with the EmptyDrops method and configure the other settings

    Ambient count threshold

    Barcodes with a total UMI count equal to or below this threshold will be used to create the ambient RNA expression profile to estimate background noise. The default is set to 100, which is reasonable for most data.

    FDR threshold

    Barcodes equal to or below this FDR threshold show a significant deviation from the ambient profile and can therefore be considered real cells. Increasing this value will result in more cells, but will also increase the number of potential false positives.

    Random generator seed

    This is used for performing Monte Carlo simulations to determine p-values. To reproduce results, use the same random seed for all runs.

    The task report will appear similar to Figure 2, with additional metrics on the left (Figure 5).

    Figure 5. Cell barcode QA/QC task report with EmptyDrops enabled

    The number of actual cells detected by the EmptyDrops test and the knee point filter are shown above the Venn diagram on the left. In Figure 5, 3,189 barcodes are above the knee point filter (represented by the vertical blue line on the plot) and 2,657 barcodes passed the significance test in EmptyDrops. The overlap between these sets of barcodes is represented by the Venn diagram. In Figure 5, 1,583 barcodes pass the significance test in EmptyDrops and have a high total UMI count above the knee point filter; 1,606 barcodes have a very high total UMI count with no significant difference from the ambient profile in EmptyDrops; 1,074 barcodes fall below the knee point but are still significantly different from the ambient profile.

    The number of cells included by the knee point filter can be adjusted either by click on the plot to change the position of the vertical blue line or by typing a different number of cells into the text box on the left.

    The total number of cells is shown in the text box on the left. By default, this will be all of the cells detected by the knee point filter plus the extra cells detected by EmptyDrops. In Figure 5, this means the 3,189 cells with a high total UMI count plus the additional 1,074 cells from EmptyDrops (total = 4,263).

    Different sections of the Venn diagram can be selected/deselected to include/exclude barcodes. For example, in Figure 5, clicking the '1,606' section of the Venn diagram will deselect those barcodes. Now, the only cells that will pass the filter will be the significant ones from EmptyDrops (Figure 6).

    Figure 6. Click a section of the Venn diagram to deselect it. In this case, only the 2,657 cells that pass the EmptyDrops test will be included
    1. Lun, A., Riesenfeld, S., Andrews, T. et al. EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data. Genome Biol. 2019; 20: 63.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Cell Barcode QA/QC without EmptyDrops
    Cell Barcode QA/QC with EmptyDrops
    References

    Cell Barcode QA/QC without EmptyDrops

    Cell Barcode QA/QC with EmptyDrops

    References

    Additional Assistance

  • Click a PCA data node

  • Click the Batch removal section in the toolbox

  • Click Harmony

  • You will be prompted to pick some attribute(s) for analysis. The Harmony dialog is similar to the General linear model batch removal. To set up the model, you need to choose which attributes should be considered. For example, in the case of one dataset that has different cell types from multiple batches, the batch may have divergent impacts on different cell types. Here, batch is the attribute Sample name and cell type is the attribute Cell type (Figure 2).

    To remove batch effects with default settings,

    • Click Sample name

    • Click Add factors

    • Click Finish

    Figure 2. Select factors to remove.

    The output of Harmony is a new data node. This data node contains the Harmony corrected values and can be used as the input for downstream tasks such as Graph-based clustering, UMAP and T-SNE (Figure 3).

    Figure 3. UMAP displays the cells colored by Cell type before (left) and after (right) Harmony integration.

    Users can click Configure to change the default settings In Advanced options (Figure 4).

    Figure 4. Advanced configure options for Harmony in Flow.

    Diversity clustering penalty (theta): Default theta=2. Higher value of penalty will have stronger correction, which results in better mixing . Zero penalty means no correction. The range of this value is from 0 to positive infinity.

    Number of clusters (nclust): Number of clusters in model. Set this to the distinct count of cell types. nclust=1 equivalent to simple linear regression. Use 0 to enable Seurat’s RunHarmony() default setting.

    Width of soft kmeans clusters (sigma): The range of this value is from 0 to positive infinity. When set it to 0, an observation will be assigned to 1 cluster (hard clustering). When the value is greater than 0, the observation will be potentially belong to multiple clusters (soft clustering, or fuzzy clustering). Default sigma=0.1. Sigma scales the distance from a cell to cluster centroids. Larger values of sigma result in observations assigned to more clusters. Smaller values of sigma make soft kmeans cluster approach hard clustering.

    Ridge regression penalty (lambda): Default lambda=1. Lambda must be strictly positive. Smaller values result in more aggressive correction.

    Random seed: Use the same random seed to reproduce the results.

    1. Korsunsky I, Millard N, Fan J, Slowikowski K, Zhang F, Wei K, Baglaenko Y, Brenner M, Loh P-r, Raychaudhuri S. Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods; 2019. https://doi.org/10.1038/s41592-019-0619-0.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. Harmony task in Batch removal section in Flow.

    References

    Additional Assistance

    Click t-SNE
  • Click Finish to run

  • t-SNE produces a t-SNE task node. Opening the task report launches a scatter plot showing the t-SNE results. Each point on the plot is a cell for single cell data or a sample for bulk data. The plot will open in 2D or 3D depending on the user preference.

    Figure 1. t-SNE task report with interactive scatter plot

    Chose whether to run t-SNE on all samples together or on each sample individually.

    Checking the box will run t-SNE on each sample individually.

    This option appears when there are multiple feature types in the input data node (e.g., CITE-Seq data).

    Select Any to run on all features or pick a feature type.

    t-SNE preserves the local structure of the data by focusing on the distances between each point and its nearest neighbors. Perplexity can be thought of as the number of nearest neighbors being considered. The optimal perplexity depends on the size and density of the data. Generally, a larger and/or more dense data set will benefit from a higher perplexity (Figure 2). Default is 30. The range of possible values is 3 to 100.

    Figure 2. Setting t-SNE perplexity to 5 (left), 30 (middle), 100 (right)

    t-SNE uses an iterative algorithm to optimize the low-dimensional representation. More iterations will result in a more accurate embedding to an extent, but will take longer to run. Default is 1000.

    Several parts of t-SNE utilize a random number generator to provide an initial value. Default is 1. To reproduce the results, use the same random seed at all runs.

    If selected, t-SNE initializes from random initial positions for each point. If disabled, the initial values for each point are assigned using the largest principal components extracted from the raw data. Default is enabled.

    The metric to use when computing distances in high-dimensional space. Options are Euclidean, Manhattan, Chebyshev, Canberra, Bray Curtis, and Cosine. Default is Euclidean.

    If checked, mapping error information will be available in the task report. Default is disabled.

    Output a t-SNE table data node that can be downloaded. The 2D t-SNE coordinates are labeled Feature 1 and Feature 2; the 3D t-SNE coordinates are labeled Feature 3, 4, and 5. Default is disabled.

    t-SNE uses principal components as its input. The number of principal components to use is set here.

    We recommend using the PCA task to determine the optimal number of principal components for your data. Default is 50.

    Options are equally or by variance. Feature values can be standardized prior to PCA so that the contribution of each feature does not depend on its variance. To standardize, choose equally. To take variance into account and focus on the most variable features, choose by variance. Default is by variance.

    You can choose to log transform the data prior to running PCA as part of t-SNE. Default is disabled.

    If you are normalizing the data, choose a log base. Default is 2 when Log transform data is enabled.

    If you are normalizing the data, choose an offset. Default is 1 when Log transform data is enabled.

    [1] L.J.P. van der Maaten and G.E. Hinton. Visualizing High-Dimensional Data Using t-SNE. Journal of Machine Learning Research 9(Nov):2579-2605, 2008.

    What is t-SNE?

    Running t-SNE

    Basic t-SNE parameters

    Split cells by sample

    Include features where "Feature type" is

    Advanced t-SNE parameters

    Perplexity

    Number of iterations

    Random generator seed

    Initialize output values at random

    Distance metric

    Generate mapping error statistics

    Generate t-SNE table

    PCA: Number of principal components

    PCA: Features contribute

    Normalization: Log transform data

    Normalization: Log base

    Normalization: Log offset

    References

    Click K-means clustering

  • Configure the parameters

  • Click Finish to run (Figure 1)

  • K-means clustering produces a K-means Clusters result data node; double-click to open the task report which lists the cluster statistics (Figure 2). If Compute biomarkers was enabled, top markers will be available by double-clicking the Biomarkers result data node. If clustering was run with Split by sample enabled on a single cell counts data node, the cluster results table displays the number of clusters found for each sample and clicking the sample name opens the sample-level report.

    Figure 2. K-means clustering task report

    The total number of clusters is listed along with the number and percentage of cells in each cluster.

    The K-means Clustering result data node includes the input values and adds cluster assignment as a new attribute, K-means, for each observation.

    Choose which distance metric to use for cluster distance calculations. Options include Euclidean, Absolute Value, Euclidean Squared, Kendall Correlation, Max Value, Min Value, Pearson Correlation, Rank Correlation, Average Euclidean, Shape, Cosine, Canberra, Bray Curtis, Tanimoto, Pearson Correlation Absolute, Rank Correlation Absolute, and Kendall Correlation Absolute. The default is Euclidean.

    Choose between specifying a set number of clusters or a range to test for the best fit number of clusters. The best fit is determined by the number of clusters with the lowest Davies–Bouldin index. The default is set to 10 for a fixed number of clusters. The initial values for the range option are 3 to 20 clusters.

    Choose whether to run the ANOVA test comparing each cluster to all other observations to identify features that have higher values in that cluster. Default is Enabled.

    This option is present in single cell data. If enabled, K-means clustering will be run separately for each sample. If disabled, K-means clustering will be run on all cells from the input data. Default is set by the Split single cell by sample option in the user preference page.

    If enabled, the initial cluster centroids will be selected randomly from among the data points. If disabled, the initial cluster centroids will be selected to optimize distance between clusters. Default is Disabled.

    This sets the random seed used if Random cluster initialization is enabled. Use the same random seed to reproduce results.

    If enabled, all cluster centroids will be recomputed at the end of each iteration. If disabled, each cluster centroid will be recomputed as the members of the cluster change. Default is Enabled.

    The maximum number of iterations to perform before setting on a set of clusters. Default is 1000.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    What is K-means clustering?

    Running K-means clustering

    <figure><img src="../../../../.gitbook/assets/image (92).png" alt=""><figcaption><p>Figure 1. K-means clustering configuration dialog</p></figcaption></figure>

    Cluster statistics

    Basic K-means clustering parameters

    Distance metric

    Number of clusters

    Compute biomarkers

    Split cells by sample

    Advanced K-means clustering parameters

    Random cluster initialization

    Random seed

    Batch centroid computations

    Max iterations

    Additional Assistance

    Click the Batch removal section in the toolbox
  • Click Seurat3 Integration

  • You will be promoted to pick some attribute(s) for analysis. The first Seurat3 integration dialog is a drop-down list that includes the factors for data integration. To set up the model, you need to choose which attribute should be considered. For example, in the case of one dataset that has different cell types from multiple technologies(Tech), different technology may have divergent impacts on different cell types. Hence, the attribute Tech should be considered to be the batch factor. The attribute celltype represents different cell types in this dataset (Figure 2).

    To integrate data with default settings,

    • Select Tech from the dropdown list

    • Click Finish

    Figure 2. Interface of Seurat3 integration in Flow. Example attributes are indicated in the drop-down list.

    The output of Seurat3 integration is a new data node - Integrated counts (Figure 1). We can then use this new integrated matrix for downstream analysis and visualization (Figure 3).

    Figure 3. UMAP displays the cells colored by Tech attribute before(left) and after(right) Seurat3 integration. Both images in the figure share the same legend.

    Users can click Configure to change the default settings In Advanced options (Figure 4).

    Figure 4. Advanced configure options for Seurat3 integration in Flow.

    Use reference to find anchors: when this box is checked, the first group of the selected attribute is used as reference to find anchors. To use a different group as reference, change the order of subgroups of the attribute in the attribute management page on Data tab. When the box is unchecked, anchors will be identified by comparing all pairs of subgroups, this option is very computationally intensive.

    Perform L2 normalization: Perform L2 normalization on the CCA cell embeddings after dimensional reduction.

    Pick anchors: How many neighbors (k) to use when picking anchors.

    Filter anchors: How many neighbors (k) to use when filtering anchors.

    Score anchors: How many neighbors (k) to use when scoring anchors.

    Nearest neighbor finding methods: Method for nearest neighbor finding. Options include: rann, annoy.

    \

    1. Stuart T, Butler A, Hoffman P, et al. Comprehensive integration of single-cell data. Cell, 2019. DOI:https://doi.org/10.1016/j.cell.2019.05.031 \

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    \

    \

    Figure 1. Seurat3 integration task in Batch removal section in Flow.

    References

    Additional Assistance

  • Geometric: FPKM are scaled via the median of the geometric means of the fragment counts across all libraries [6]. This is the default option (and is identical to the one used by DESeq)

  • Quartile: FPKMs are scaled via the ratio of the 75 quartile fragment counts to the average 75 quartile value across all libraries

  • The library types have three options:

    • Fr-unstranded: reads from the left-most end of the fragment in transcript coordinates map to the transcript strand, and the right-most end maps to the opposite strand. E.g. standard Illlumina

    • Fr-firststrand: reads from the left-most end of the fragment in transcript coordinates map to the transcript strand, and the right-most end maps to the opposite strand. The right-most end of the fragment is the first sequenced or only sequenced for single-end reads. It is assumed that only the strand generated during first strand synthesis is sequenced. E.g. dUPT, NSR, NNSR

    • Fr-secondstrand: reads from the left-most end of the fragment in transcript coordinates map to the transcript strand, and the right-most end maps to the opposite strand. The left-most end of the fragment is the first sequenced or only sequenced for single-end reads. It is assumed that only the strand generated during second strand synthesis is sequenced. E.g. Directional Illumina, standard SOLiD.

    The report of the cuffdiff task is a table of a feature list p-values, q-value and log2 fold-change information for all the comparisons (Figure 3).

    Figure 3. Figure 20: Cuffdiff task report. Each row is a feature, p-value, q-value and log2 fold change columns are display for each comparison

    In the p-value column, besides an actual p-value, which means the test was performed successfully, there is also the following flags which indicate the test was not successful:

    • NOTEST: not enough alignments for testing

    • LOWDATA: too complex or shallowly sequences

    • HIGHDATA: too many fragments in locus

    • FAIL: when an ill-conditioned covariance matrix or other numerical exception prevents testing

    The table can be downloaded as a text file when clicking the Download button on the lower-right corner of the table.

    1. Benjamini, Y., Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing, JRSS, B, 57, 289-300.

    2. Storey JD. (2003) The positive false discovery rate: A Bayesian interpretation and the q-value. Annals of Statistics, 31: 2013-2035.

    3. Auer, 2011, A two-stage Poisson model for testing RNA-Seq

    4. Burnham, Anderson, 2010, Model selection and multimodel inference

    5. Law C, Voom: precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biology, 2014 15:R29.

    6. Anders S, Huber W: Differential expression analysis for sequence count data. Genome Biology, 2010

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Figure 1. Cuffdiff setup dialog. “Select attributes(s) to groups samples” lists the categorical attributes which have at least two levels (e.g. “Cell type” and “Time”)
    Figure 2. Advanced option of cuffdiff

    References

    Additional Assistance

    AUCell is a tool to identify cells that are actively expressing genes within a gene list [1]. For each input gene list, AUCell calculates a value for each cell by ranking all genes by their expression level in the cell and identifying what proportion of the genes from the gene list fall within the top 5% (default cutoff) of genes. This method allows the AUCell value to represent the proportion of genes from the gene list that are expressed in the cell and their relative expression compared to other genes within the cell. Because this is a rank-based method and is calculated for each cell individually, AUCell can be run on raw or normalized data. As an AUCell value is the proportion of genes from the list that are within the top percentile of expressed genes, AUCell values can range from 0 to 1, but may have a more restricted range.

    AUCell values can be used directly as input for downstream analysis, such as clustering. Another common use is to set an AUCell value cutoff for expressing vs. not and used this to classify cells. AUCell values will separate cells most effectively when the genes in the list are highly and specifically expressed in a population of cells. If the genes are specifically expressed, but not highly expressed, the AUCell value will not be as useful.

    AUCell can be run on any single cell counts data node.

    • Click the single cell counts data node

    • Click the Exploratory analysis section in the toolbox

    • Click AUCell

    • Choose gene lists by clicking and dragging them to the panel on the right or clicking the green plus that appears after mousing over a gene list (Figure 1)\

    • Click Finish to run

    AUCell produces an AUCell result data node. The AUCell result data node includes the input counts data and adds the AUCell scores to the original data as a new data type, AUCell Values. AUCell values for each input feature list are included as features in the AUCell result data node. These features created by AUCell are named after the feature list (e.g., B cells, Cytotoxic cells).

    Because the AUCell values are added as features, they can be used as input for clustering, differential analysis, and visualization tasks.

    To produce a data node containing only the AUCell values, use Split matrix to split the AUCell result data node into separate data nodes for each of its data types. This can be helpful if you intend on performing downstream analysis on the AUCell values. To perform differential analysis, it is advisable to normalize the values by adding a small offset (e.g. 1E-9) and Logit transformation to the base Log2 using the Normalization task. This will make the values continuous and suitable for differential analysis with methods such as ANOVA/LIMMA-trend/LIMMA-voom, Non-parametric ANOVA or Welch's ANOVA. For differential analysis, please check the Low-value filter is set to None and the values are correctly recognized as Log2 transformed in the Advanced settings.

    If an AUCell result data node or other downstream data node containing AUCell Values is used as the input for AUCell, the additional AUCell values will be added as additional features of the AUCell values data type in the new AUCell result data node.

    For each gene set, AUCell computes the intersection between the gene list and the input data set. If the intersection size is below the specified threshold, the gene set is ignored and no AUCell score is calculated for it. Default is 5.

    To calculate the AUCell value, genes are ranked and the fraction of genes from the gene list that are above the percentile cutoff is the AUCell value. This parameter sets the percentile cutoff. Default is 5.

    [1] Aibar, S., González-Blas, C. B., Moerman, T., Imrichova, H., Hulselmans, G., Rambow, F., ... & Atak, Z. K. (2017). SCENIC: single-cell regulatory network inference and clustering. Nature methods, 14(11), 1083.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    What is AUCell?

    What is AUCell?
    Running AUCell
    Advanced AUCell parameters
    References

    Running AUCell

    Advanced AUCell parameters

    Minimum gene set size

    Top features in percentiles

    References

    Additional Assistance

    Figure 2. Example of attributes required for paired variant analysis

    Examples of the latter can include case and control or tumor and normal. If these attributes are present, a section for Analysis options will be displayed below the Variant detection method (Figure 3). To utilize this feature, select Paired analysis. Match ID must then be specified and should correspond to the attribute that references the sample ID shared between the pair. Selecting Case/control will allow for discriminating genotypes between paired samples in downstream tasks. Attribute should correspond to the attribute that defines type within sample pairs, and Control can be specified for whatever category relates to the reference sample.

    Figure 3. Specifying options for paired variant detection

    Select Reference sequence will specify the reference assembly to utilize for variant detection. If the alignment was generated in Partek® Flow®, the Assembly will be displayed as text in the section, and you do not have the option to change the reference. In the event that alignment was performed outside of Partek Flow, you will need to select the appropriate Assembly utilized for alignment in the drop-down list. Assemblies previously added to library files (see Library File Management) will be available for selection or New assembly… can be utilized to import the reference sequence to library files from within the task.

    Advanced options provides a means to tune parameters in the variant detection for optimal performance. Upon invoking the task dialog, Option set is set to Default, and these parameters are provided by the SAMtools developers. Clicking Configure will open a window to tune advanced options Moving the mouse cursor over the info button will provide details for each parameter. Please refer to the SAMtools documentation for further details on any of these parameters.

    1. Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinforma Oxf Engl. 2009;25(16):2078-2079.

    2. Li H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics. 2011;27(21):2987-2993.

    3. Li H. Improving SNP discovery by base alignment quality. Bioinformatics. 2011;27(8):1157-1158.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    SAMtools dialog

    Figure 1. Selecting a variant detection method in the Samtools dialog

    References

    Additional Assistance

    Exploratory analysis
    >
    Find multimodal neighbors
    in the toolbox. On the task setup page, use the
    Select data node
    button to point to the PCA data node of the other assay (e.g. protein expression), by default, there is a node selected (Figure 1).
    Figure 1. Setting up the Find multimodal neighbors task. Use the Select data node button to point to the second (of two) assay

    When you click the Select data node button, Partek Flow will open another dialog, showing your current pipeline (Figure 2). Data nodes that can be used for WNN are in color of the branch, other nodes are disabled (greyed out). To pick a node, left-click on it and then push the Select button.

    Figure 2. Select data node dialog. A node that can be used for WNN is highlighted in blue (in this example, the PCA node on the right)

    The selected data node is shown under the Select data node button. If you made a mistake, use the Clear selection link (Figure 1).

    If there are graph-based clustering task performed on PCA data node, the output of graph-based clustering node also has PCA score from the input data, so the output graph-based clustering data nodes also can be candidate of WNN task.

    To customize the Advanced options, select the Configure link (Figure 1). At present you can only change the number of nearest neighbors for each modality (-k.nn option of the Seurat package); the default value is 20 (Figure 3). An illustration on how to use that option to assess the robustness of WNN analysis can be found in Hao et al. (1). The nearest neighbor search method is K-NN and distance metric is Euclidean.

    Figure 3. Advanced options of the Find multimodal neighbors

    To launch the Find multimodal neighbors task, click the Finish button on the task setup page (Figure 1). For each cell, the WNN algorithm calculates its closest neighbors based on a weighted combination of RNA and protein similarities. The output of the Find multimodal neighbors task is a WNN data node.

    For downstream analysis, you can launch a UMAP or graph-based clustering tasks on a WNN node. For example, Figure 4 shows a snippet of analysis of a feature barcoding data set; gene expression and protein expression data were processed separately, and then Find multimodal neighbors was invoked on two respective PCA data nodes. UMAP and graph-based clustering tasks were performed on WNN node.

    Figure 4. Example of a pipeline with a Find multimodal neighbors task and the resulting WNN data node

    For an excellent illustration on advantages of WNN algorithm for identification of cell types, please see this blog post.

    1. Hao Y, Hao S, Andersen-Nissen E, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184(13):3573-3587.e29. doi:10.1016/j.cell.2021.04.048

    2. https://satijalab.org/seurat/articles/weighted_nearest_neighbor_analysis.html

    Invoking Find Multimodal Neighbours

    Invoking Find Multimodal Neighbours
    References

    References

    LOW – synonymous changes, etc;
  • MODIFIER – changes outside coding regions, etc.

  • Further details about output metrics can be found in the SnpEff documentation. The Annotate variants (SnpEff) task can be invoked from any Variants or Annotated variants data node, and the task will supplement any existing annotation in the vcf files. Annotation information will also be visible in the View variants Variant report and the Summarize cohort mutations Cohort mutation summary report

    The task dialog for Annotate variants (SnpEff) contains two sections: Select SnpEff database and Advanced options (Figure 1). Select SnpEff database will specify the reference assembly to utilize for variant detection. If the variant detection was performed in Partek Flow, the Assembly will be displayed as text in the section, and you do not have the option to change the reference. In the event that variant detection was performed outside of Partek Flow, you will need to select the appropriate Assembly utilized for variant detection in the drop-down list. Assemblies previously added to library files (see Library File Management) will be available for selection or New assembly… can be utilized to import the reference sequence to library files from within the task. Select SnpEff database will allow selection of databases utilized for prediction, and Partek Flow provides automated download of a limited number of these databases. Databases previously added to library files (see Library File Management) will be available for selection or Add SnpEff variant database in the menu can be utilized to import the reference sequence to library files from within the task. Additional information of SnpEff databases can be found in the SnpEff documentation.

    Figure 1. Components of the SnpEff dialog

    Advanced options provides a means to tune parameters for annotation generated from the SnpEff database. Upon invoking the task dialog, Option set is set to Default, and these parameters are prescribed by the developers of SnpEff. Clicking Configure will open a window to tune advanced options (Figure 2). SnpEff has Advanced options for Results filter options, Annotation options, and Database options. Moving the mouse cursor over the info button will provide details for each parameter.

    Figure 2. Configuration of SnpEff advanced options
    1. Cingolani P, Platts A, Wang LL, et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin). 2012;6(2):80-92.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Annotate variants (SnpEff) dialog

    References

    Additional Assistance

    documentation.
    Figure 1. Components of the VEP dialog

    Advanced options provides a means to specify aspects of the annotation generated from the VEP annotation task. Upon invoking the task dialog, Option set is set to Default. Clicking Configure will open a window to specify additional components of annotation (Figure 2). VEP has Advanced options for Identifiers, Output options, and Co-located variants. Moving the mouse cursor over the info button will provide details for each parameter.

    Figure 2. Configuration of VEP advanced options

    In the report, there variant impact information, it is a subjective classification of the severity of the variant consequence:

    Low: a variant that is assumed to be mostly harmless or unlikely to change protein behavior

    Moderate: a non-disruptive variant that might change protein effectiveness

    Modifier: usually non-coding variants or variants affecting non-coding genes, where predictions are difficult or there is no evidence of impact

    High: a variant is assumed to have high disruptive impact in the protein, probably causing protein truncation, loss of function or triggering nonsense mediated decay.

    If you need additional assistance, please visit our support page to submit a help ticket or find phone numbers for regional support.

    Annotate variants (VEP) dialog

    Annotate Variants

    Additional Assistance

    VEP

    Annotate Visium image

    References

    Additional Assistance

    https://support.10xgenomics.com/spatial-gene-expression/software/pipelines/latest/output/overview
    our support page
    Figure 1. The content in outputs directory from Space Ranger. The folder and file highlighted in red are the required files