Python for Big Genomics Datasets

In this chapter, we will cover the following recipes:

  • Using high-performance data formats – HDF5
  • Doing parallel computing with Dask
  • Using high-performance data formats – Parquet
  • Computing sequencing statistics using Spark
  • Optimizing code with Cython and Numba

Get Bioinformatics with Python Cookbook - Second Edition now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.