How To Run

Requirements Currently supported Nextflow versions: 23.04.2

Below is a summary of how to run the pipeline. See here for full instructions.

Pipelines should be run WITH A SINGLE SAMPLE AT TIME. Otherwise resource allocation and Nextflow errors could cause the pipeline to fail.

  1. The recommended way of running the pipeline is to download and unpack a release with submodules.

  2. The source code should never be modified when running our pipelines

  3. Create a config file for input, output, and parameters. An example for a config file can be found here. See Inputs for the detailed description of each variable in the config file. The config file can be generated using a python script (see below).

  4. Do not directly modify the source template.config, but rather you should copy it from the pipeline release folder to your project-specific folder and modify it there

  5. Create the input csv using the template. The example csv is a single-lane sample, however this pipeline can take multi-lane sample as well, with each record in the csv file representing a lane (a pair of fastq). All records must have the same value in the sample column. See Inputs for detailed description of each column. All columns must exist in order to run the pipeline successfully.

  6. Again, do not directly modify the source template csv file. Instead, copy it from the pipeline release folder to your project-specific folder and modify it there.

  7. Inputs can also be provided in YAML format using the YAML template and passed to the pipeline using -params-file.

  8. The pipeline can be executed locally using the command below:

nextflow run path/to/main.nf -config path/to/sample-specific.config [-params-file path/to/input.yaml] [-profile <selected_profile>]
  • For example, path/to/main.nf could be: /hot/software/pipeline/pipeline-align-DNA/Nextflow/release/8.0.0/pipeline/align-DNA.nf
  • path/to/sample-specific.config is the path to where you saved your project-specific copy of template.config

BWA-MEM2 Genome Index The reference genome index must be generated by BWA-MEM2 with the correct version. Genome index generated by old BWA-MEM2 versions or the original BWA is not accepted. The reference genome index can be generated using the generate-genome-index.nf nextflow pipeline. To run this pipeline, you need to create a config file using this template to specify the path of reference_fasta and the temp_dir. The temp_dir is used to store intermediate files of Nextflow. The genome index files are saved to the same directory of the input reference FASTA by the pipeline. Use the command below to run this generate genome index pipeline:

nextflow run path/to/generate-genome-index.nf -config path/to/genome-specific.config

BWA-MEM2 expects the reference genome index to be at the same directory as the reference genome FASTA, so it's important to keep them together.

minibwa Genome Index

The minibwa reference must be indexed with the same minibwa version used for alignment. Generate the index with minibwa index reference.fa. With the pipeline's simple reference contract, do not supply a separate index prefix: the command must produce reference.fa.l2b and reference.fa.mbw next to the FASTA, and reference_fasta_minibwa must point to reference.fa. BWA-MEM2 indexes cannot be reused by minibwa.

minibwa currently does not properly support alternate contigs. Use a reference without alternate contigs.

HISAT2 Genome Index The reference genome index must be generated from HISAT2 using hisat2-build. When passing the hisat2 index to the config, only the path up to the prefix(basename) must be specified:

The basename is the name of any of the index files up to but not including the final .1.ht2 / etc. hisat2 looks for the specified index first in the current directory, then in the directory specified in the HISAT2_INDEXES environment variable.