The properly rendered version of this document can be found at Read The Docs.
If you are reading this on github, you should instead click here.
This pipeline tests a set of reads for contamination. It takes as input:
- a set of ReadGroupSets to test
- statistics on reference allele frequencies for SNPs with a single alternative from a set of VariantSets
and combines these to produce an estimate of the amount of contamination.
Uses the sequence data alone approach described in:
The pipeline is implemented on Google Cloud Dataflow.
Most users launch Dataflow jobs from their local machine. This is unrelated to where the job itself actually runs (which is controlled by the
--runner parameter). Either way, Java 8 is needed to run the Jar that kicks off the job.
If you do not have Java on your local machine, the following setup instructions will allow you to launch Dataflow jobs using the Google Cloud Shell:
- If you have not already done so, follow the Genomics Quickstart.
- If you have not already done so, follow the Dataflow Quickstart.
- Use the Cloud Console to activate the Google Cloud Shell.
- Run the following commands in the Cloud Shell to install Java 8.
sudo apt-get update sudo apt-get install --assume-yes openjdk-8-jdk maven sudo update-alternatives --config java sudo update-alternatives --config javac
Depending on the pipeline, Cloud Shell may not not have sufficient memory to run the pipeline locally (e.g., without the
--runner command line flag). If you get error
java.lang.OutOfMemoryError: Java heap space, follow the instructions to run the pipeline using Compute Engine Dataflow workers instead of locally (e.g. use
If you want to run a small pipeline on your machine before running it in parallel on Compute Engine, you will need ALPN since many of these pipelines require it. When running locally, this must be provided on the boot classpath but when running on Compute Engine Dataflow workers this is already configured for you. You can download it from here. For example:
wget -O alpn-boot.jar \ http://central.maven.org/maven2/org/mortbay/jetty/alpn/alpn-boot/8.1.8.v20160420/alpn-boot-8.1.8.v20160420.jar
Download the latest GoogleGenomics dataflow runnable jar from the Maven Central Repository. For example:
wget -O google-genomics-dataflow-runnable.jar \ https://search.maven.org/remotecontent?filepath=com/google/cloud/genomics/google-genomics-dataflow/v1-0.1/google-genomics-dataflow-v1-0.1-runnable.jar
The following command will calculate the contamination estimate for a given ReadGroupSet and specific region in the 1,000 Genomes dataset. It also uses the VariantSet within 1,000 Genomes for retrieving the allele frequencies.
java -Xbootclasspath/p:alpn-boot.jar \ -cp google-genomics-dataflow-runnable.jar \ com.google.cloud.genomics.dataflow.pipelines.VerifyBamId \ --references=17:41196311:41277499 \ --readGroupSetIds=CMvnhpKTFhDq9e2Yy9G-Bg \ --variantSetId=10473108253681171589 \ --output=gs://YOUR-BUCKET/dataflow-output/verifyBamId-platinumGenomes-BRCA1-readGroupSet-CMvnhpKTFhCAv6TKo6Dglgg.txt
The above command line runs the pipeline locally over a small portion of the genome, only taking a few minutes. If modified to run over a larger portion of the genome or the entire genome, it may take a few hours depending upon how many virtual machines are configured to run concurrently via
--numWorkers. Add the following additional command line parameters to run the pipeline on Google Cloud instead of locally:
--runner=DataflowPipelineRunner \ --project=YOUR-GOOGLE-CLOUD-PLATFORM-PROJECT-ID \ --stagingLocation=gs://YOUR-BUCKET/dataflow-staging \ --numWorkers=#
To run this pipeline over the entire genome, use
--allReferences instead of
To run the pipeline on a different group of read group sets:
* Change the
--readGroupSetIds or the
* Update the
--references as appropriate (e.g., add/remove the ‘chr’ prefix on reference names).
To configure the pipeline more to fit your needs in terms of the minimum allele frequency to use or the fraction of positions to check, change the
If the Application Default Credentials are not sufficient, use
--client-secrets PATH/TO/YOUR/client_secrets.json. If you do not already have this file, see the authentication instructions to obtain it.
--help to get more information about the command line options. Change
the pipeline class name below to match the one you would like to run.
java -cp google-genomics-dataflow*runnable.jar \ com.google.cloud.genomics.dataflow.pipelines.VariantSimilarity --help
See the source code for implementation details: https://github.com/googlegenomics/dataflow-java