Skip to main content

tools to support genome and metagenome analysis

Project description

genome-grist - map Illumina metagenomes to GenBank genomes

PyPI License: 3-Clause BSD

In brief

genome-grist is a toolkit to do the following:

  1. download a metagenome
  2. process it into trimmed reads, and make a sourmash signature
  3. search the sourmash signature with 'gather' against sourmash databases, e.g. all of genbank
  4. download the matching genomes from genbank
  5. map all metagenome reads to genomes using minimap
  6. extract matching reads iteratively based on gather, successively eliminating reads that matched to previous gather matches
  7. run mapping on “leftover” reads to genomes
  8. summarize all mapping results

Installation

The command:

python -m pip install genome-grist

will install the latest version. Plase use python3.7 or later. We suggest using an isolated conda environment; the following commands should work for conda:

conda create -n grist python=3.7 pip
conda activate grist
python -m pip install genome-grist

Quick start:

Run the following three commands.

First, download SRA sample HSMA33MX, trim reads, and build a sourmash signature:

genome-grist process HSMA33MX smash_reads

Next, run sourmash signature against genbank:

genome-grist process HSMA33MX gather_genbank

(NOTE, this depends on the latest genbank genomes and won't work for most people just yet; for now, use cached results from the repo:

cp tests/test-data/HSMA33MX.x.genbank.gather.csv outputs/genbank/
touch outputs/genbank/HSMA33MX.x.genbank.gather.out

)

Finally, download the reference genomes, map reads and produce a summary report:

genome-grist process HSMA33MX summarize -j 8

(You can run all of the above with make test in the repo.)

The summary report will be in outputs/reports/report-HSMA33MX.html.

You can see some example reports for this and other data sets online:

Compute requirements

You'll need enough disk space to store about 5 copies of your raw metagenome.

The peak memory requirement is in the k-mer trimming and sourmash gather steps. You'll probably want between 30 and 60 GB of RAM for those, although for smaller or less diverse metagenomes, you will use a lot less.

Full set of top-level process targets

  • download_reads
  • trim_reads
  • smash_reads
  • gather_genbank
  • download_matching_genomes
  • map_reads
  • summarize

Support

genome-grist is alpha-level software. Please be patient and kind :).

Please ask questions and add comments by filing github issues.

Why the name grist?

'grist' is in the sourmash family of names (sourmash, wort, distillerycats, etc.) See Grist.

(It is not the computing grist!)


CTB Nov 8, 2020

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

genome-grist-0.3.2.tar.gz (4.4 MB view details)

Uploaded Source

File details

Details for the file genome-grist-0.3.2.tar.gz.

File metadata

  • Download URL: genome-grist-0.3.2.tar.gz
  • Upload date:
  • Size: 4.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.0 requests/2.24.0 setuptools/50.3.2 requests-toolbelt/0.9.1 tqdm/4.50.2 CPython/3.7.6

File hashes

Hashes for genome-grist-0.3.2.tar.gz
Algorithm Hash digest
SHA256 667818ec6c805fd5ac233bc13e35e1faac8eed82e1d0b9117021a3e26b9b2de1
MD5 61b3664925e3adc8261e0e64bbbf2bbb
BLAKE2b-256 260ae6d6dcef8aa3b02b6ba8531a635044470e2091064afc74e18ffb10dd1900

See more details on using hashes here.

Provenance

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page