Showing posts with label synbio. Show all posts
Showing posts with label synbio. Show all posts

Saturday, April 22, 2017

Sunday Morning Insight: "No Need for the Map of a Cat, Mr Feynman" or The Long Game in Nanopore Sequencing.



About 5 weeks ago, we wondered how we could tell if the world was changing right before our eyes ?  well, this is happening, instance #2 just got more real:


Nanopore sequencing is a promising technique for genome sequencing due to its portability, ability to sequence long reads from single molecules, and to simultaneously assay DNA methylation. However until recently nanopore sequencing has been mainly applied to small genomes, due to the limited output attainable. We present nanopore sequencing and assembly of the GM12878 Utah/Ceph human reference genome generated using the Oxford Nanopore MinION and R9.4 version chemistry. We generated 91.2 Gb of sequence data (~30x theoretical coverage) from 39 flowcells. De novo assembly yielded a highly complete and contiguous assembly (NG50 ~3Mb). We observed considerable variability in homopolymeric tract resolution between different basecallers. The data permitted sensitive detection of both large structural variants and epigenetic modifications. Further we developed a new approach exploiting the long-read capability of this system and found that adding an additional 5x-coverage of "ultra-long" reads (read N50 of 99.7kb) more than doubled the assembly contiguity. Modelling the repeat structure of the human genome predicts extraordinarily contiguous assemblies may be possible using nanopore reads alone. Portable de novo sequencing of human genomes may be important for rapid point-of-care diagnosis of rare genetic diseases and cancer, and monitoring of cancer progression. The complete dataset including raw signal is available as an Amazon Web Services Open Dataset at: https://github.com/nanopore-wgs-consortium/NA12878.
Here is some context:

And previously on Nuit Blanche:
 
Credit: NASA, JPL




Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Sunday, December 18, 2016

Sunday Morning Insight: And You, What Are You Waiting For ?



For different reasons, the winter break/solstice is a good time for getting stuff done and focusing on small or large projects. Some of them lead to discoveries and/or momentous firsts. This year is no exception with a few days before the solstice Clive Brown, the CTO of Oxford Nanopore decided to build his own genome through self sequencing. From the read me on his experiment on Github.
So far as I am aware this is the first full coverage Human Genome sequenced by the individual who provided the input sample (ONT-HG1). This may prove significant in future.
The last sentence is obviously a rather typical self-effacing affirmation also known as "British understatement". 

 And you, what are you waiting for ?
 

    Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Wednesday, October 05, 2016

    Universal microbial diagnostics using random DNA probes

    Using compressive sensing for pathogen detection.



    Universal microbial diagnostics using random DNA probes by Amirali Aghazadeh, Adam Y. Lin, Mona A. Sheikh, Allen L. Chen, Lisa M. Atkins, Coreen L. Johnson, Joseph F. Petrosino, Rebekah A. Drezek and  Richard G. Baraniuk

    Early identification of pathogens is essential for limiting development of therapy-resistant pathogens and mitigating infectious disease outbreaks. Most bacterial detection schemes use target-specific probes to differentiate pathogen species, creating time and cost inefficiencies in identifying newly discovered organisms. We present a novel universal microbial diagnostics (UMD) platform to screen for microbial organisms in an infectious sample, using a small number of random DNA probes that are agnostic to the target DNA sequences. Our platform leverages the theory of sparse signal recovery (compressive sensing) to identify the composition of a microbial sample that potentially contains novel or mutant species. We validated the UMD platform in vitro using five random probes to recover 11 pathogenic bacteria. We further demonstrated in silico that UMD can be generalized to screen for common human pathogens in different taxonomy levels. UMD’s unorthodox sensing approach opens the door to more efficient and universal molecular diagnostics.




     
    Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Wednesday, July 27, 2016

    Streaming algorithms for identification of pathogens and antibiotic resistance potential from real-time MinION (TM) sequencing

    About four years ago, I tried to predict the future for August 25, 2030. In order to do this, I first mentioned The Steamrollers i.e. technologies that were exponential in nature. Nanopore sequencing was one of them. In the second installment, I mentioned different algorithms that could help in making sense of the data generated by these steamrollers (Predicting the Future: Randomness and Parsimony). Streaming was one of them. It is really no surprise, if like in hyperspectral imaging or nanopore sequencing you are producing a lot of data, your interest switch from the modeling aspect of things to how can it be helpful now and how fast.  How does it change science ? well you just need to read the following article:

    The main contribution of this article is to demonstrate that despite the higher error rate, it is possible to return clinical actionable information, including species and strain identification from as few as 500 reads. We achieved this by developing novel approaches that are less sensitive to base-calling errors and which use whatever subset of genome-wide information is observed up to a point in time, rather than a panel of pre-defined markers or genes. For example, the strain typing presence/absence approach relies only on being able to identify homology to genes and also allows for a level of incorrect gene annotation.



    Streaming algorithms for identification of pathogens and antibiotic resistance potential from real-time MinIONTMsequencing by Minh Duc Cao, Devika Ganesamoorthy, Alysha G. Elliott, Huihui Zhang, Matthew A. Cooper and Lachlan J.M. Coin
    The recently introduced Oxford Nanopore MinION platform generates DNA sequence data in real-time. This has great potential to shorten the sample-to-results time and is likely to have benefits such as rapid diagnosis of bacterial infection and identification of drug resistance. However, there are few tools available for streaming analysis of real-time sequencing data. Here, we present a framework for streaming analysis of MinION real-time sequence data, together with probabilistic streaming algorithms for species typing, strain typing and antibiotic resistance profile identification. Using four culture isolate samples, as well as a mixed-species sample, we demonstrate that bacterial species and strain information can be obtained within 30 min of sequencing and using about 500 reads, initial drug-resistance profiles within two hours, and complete resistance profiles within 10 h. While strain identification with multi-locus sequence typing required more than 15x coverage to generate confident assignments, our novel gene-presence typing could detect the presence of a known strain with 0.5x coverage. We also show that our pipeline can process over 100 times more data than the current throughput of the MinION on a desktop computer.






    Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Saturday, June 25, 2016

    Saturday Morning Video: Machine Learning in Computational Biology Workshop, @NIPS2015

     
     
     
    The NIPS2015 Workshops videos are out. In particular, we have that of the Machine Learning in Computational Biology workshop today (I am trying to organize a workshop for this coming NIPS and if it accepted I'll make sure that all the videos are taken). Enjoy !

    Credit photo: Date: 24 June 2016, Satellite: Rosetta, Depicts: Comet 67P/Churyumov-Gerasimenko
    Copyright: ESA/Rosetta/NAVCAM, CC BY-SA IGO 3.0

    Rosetta navigation camera (NavCam) image taken on 17 June 2016 at 30.8 km from the centre of comet 67P/Churyumov-Gerasimenko. The image measures 2.7 km across and has a scale of about 2.6 m/pixel.
    The image has been cleaned to remove the more obvious bad pixels and cosmic ray artefacts, and intensities have been scaled.
    Another version of this image, which has been contrast enhanced, is available here.
    More images of comet 67P/Churyumov-Gerasimenko can be found in the '67P - by Rosetta' collection.

    This work is licensed under a Creative Commons Attribution-ShareAlike 3.0 IGO License.
     
    Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Monday, May 02, 2016

    Low Dimensionality in Gene Expression Data Enables the Accurate Extraction of Transcriptional Programs from Shallow Sequencing - implementation -

    Rajat just sent me the following:

      Dear Igor,

    Thank you for your excellent blog Nuit Blanche!  I have learned a lot from your postings over the years.

    Would you be interested in passing along to your readers my recent paper with Graham Heimberg, Hana El-Samad, and Matt Thomson?  It is on one of the few topics you have covered in your blog outside of compressive sensing, namely the growing availability and decreasing cost of RNA sequencing data.  

    In our paper, we show that the high read depths conventionally used in RNA sequencing are not needed in cases where the primary results rely on clustering or classification (or other low-dimensional representations) of RNA-seq data.  This is of particular importance for single-cell RNA-seq, where read depth is inherently low due to fundamental limits in the chemistry of capturing RNA.

    Our title and abstract is included below.  If this is too far outside the scope of your blog, we completely understand.

    Best,
    Rajat
    I don't think talking about how biology is low dimensional and therefore certain bounds apply for sampling is outside the scope of Nuit Blanche :-) Thanks Rajat ! Here is how the paper starts:

    The modern engineering discipline of signal processing has demonstrated that structural properties of natural signals can often be exploited to enable new classes of low cost measurements. The central insight is that many natural signals are effectively ‘‘low dimensional.’’ Geometrically, this means that these signals lie on a noisy, low-dimensional manifold embedded in the observed, high-dimensional measurement space. Equivalently, this property indicates that there is a basis representation in which these signals can be accurately captured by a small number of basis vectors relative to the original measurement dimension (Donoho, 2006; Candès et al., 2006; Hinton and Salakhutdinov, 2006). Modern algorithms exploit the fact that the number of measurements required to reconstruct a low-dimensional signal can be far fewer than the apparent number of degrees of freedom. For example, in images of natural scenes, correlations between neighboring pixels induce an effective low dimensionality that allows high-accuracy image reconstruction even in the presence of considerable measurement noise such as point defects in many camera pixels (Duarte et al., 2008). Like natural images, it has long been appreciated that biological systems contain structural features that can lead to an effective low dimensionality in data. Most notably, genes are commonly co-regulated within transcriptional modules; this produces covariation in the expression of many genes (Eisen et al., 1998; Segal et al., 2003; Bergmann et al., 2003). The widespread presence of such modules indicates that the natural dimensionality of gene expression is determined not by the number of genesin the genome but by the number of regulatory modules

    Strangely enough, this figure looks like a sharp phase transition of sorts:


     
    Here is the (open) paper: Low Dimensionality in Gene Expression Data Enables the Accurate Extraction of Transcriptional Programs from Shallow Sequencing

    Summary: A tradeoff between precision and throughput constrains all biological measurements, including sequencing-based technologies. Here, we develop a mathematical framework that defines this tradeoff between mRNA-sequencing depth and error in the extraction of biological information. We find that transcriptional programs can be reproducibly identified at 1% of conventional read depths. We demonstrate that this resilience to noise of “shallow” sequencing derives from a natural property, low dimensionality, which is a fundamental feature of gene expression data. Accordingly, our conclusions hold for ∼350 single-cell and bulk gene expression datasets across yeast, mouse, and human. In total, our approach provides quantitative guidelines for the choice of sequencing depth necessary to achieve a desired level of analytical resolution. We codify these guidelines in an open-source read depth calculator. This work demonstrates that the structure inherent in biological networks can be productively exploited to increase measurement throughput, an idea that is now common in many branches of science, such as image processing.

     The read depth calculator is here: https://thomsonlab.github.io/html/formula.html


    Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Friday, June 12, 2015

    Sparse Proteomics Analysis - A compressed sensing-based approach for feature selection and classification of high-dimensional proteomics mass spectrometry data - implementation -

    The knowledge generated by the genome only makes sense if one can do a good job at figuring out what sort of proteins are producing by different elements of the DNA. Hence, while GWAS studies are great and the alignement issues in sequencing is becoming easier to handle, the big unknown is now to connect information with the zoo of proteins produced by the body. That effort includes sensing the right thing in a very large dimensional space, the subject of today's paper. Let us hope it guides us into producing better sensors . And yes as the article says

    "a [Machine Learning] classification problem is equivalent to 1-bit CS " 

    Without further ado Sparse Proteomics Analysis - A compressed sensing-based approach for feature selection and classification of high-dimensional proteomics mass spectrometry data by Tim Conrad, Martin Genzel, Nada Cvetkovic, Niklas Wulkow, Alexander Leichtle, Jan Vybiral, Gitta Kutyniok, Christof Schütte

    Motivation: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested how MS spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients having a particular disease. Machine learning algorithms are needed to (a) identify these discriminating features and (b) classify unknown spectra based on this feature set. Since the acquired data is usually noisy, the algorithms should be robust to noise and outliers, and the identified feature set should be as small as possible.
    Results: We present a new algorithm, Sparse Proteomics Analysis (SPA), based on the theory of Compressed Sensing that allows to identify a minimal discriminating set of features from mass spectrometry data-sets. We show how our method performs on artificial and real-world data-sets.
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Tuesday, June 09, 2015

    Compressive phase transition, a million genomes and 3rd generation sequencing


    Steve Hsu recently made a video presentation in Genetic architecture and predictive modeling of quantitative traits that explains why reaching a million genomes is becoming important in light of what we know of compressive sensing phase transition [1]. Go watch it. In the meantime, today I'll be attending a workshop on Bioinformatics for Third Generation Sequencing in Lille. It willfeature presentations from Clive Brown from Oxford Nanopore Technologies) on the latest methods and devices for nanopore sensing and Nick Loman on Bioinformatics approaches for real-time nanopore sequencing. Nanopore sequencing has been mentioned here before as it is the technology that makes sequencing not-NP-Hard anymore [2]



    N00241314.jpg was taken on May 25, 2015 and received on Earth May 26, 2015. The camera was pointing toward SATURN, and the image was taken using the CL1 and MT3 filters.
    Image Credit: NASA/JPL-Caltech/Space Science Institute 





     
     
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Friday, May 15, 2015

    Hamming's time: The Important Things after Commodity Sequencing




    If you recall Friday afternoon Hamming's time takes its roots in one of the comment Dick Hamming made in his famous "You and Your Research". There, he mentioned that while at Bell Labs, he kept some time in the week for thinking of big problems so that he could try to make a go at them. We are also indulging in this similar exercise today, with a few thoughts triggered by the following: A vision for ubiquitous sequencing by Yaniv Erlich here is the abstract:
    Genomics has recently celebrated reaching the \$1000 genome milestone, making affordable DNA sequencing a reality. This goal of the sequencing revolution has been successfully completed. Looking forward, the next goal of the revolution can be ushered in by the advent of sequencing sensors - miniaturized sequencing devices that are manufactured for real time applications and deployed in large quantities at low costs. The first part of this manuscript envisions applications that will benefit from moving the sequencers to the samples in a range of domains. In the second part, the manuscript outlines the critical barriers that need to be addressed in order to reach the goal of ubiquitous sequencing sensors.
    we've mentioned Yaniv's earlier work connecting population genomics and compressive sensing related techniques before here.

    Also right now in London there is a conference by users of the Oxford nanopore technology, here are a few presentations:
    and  
    and Sequencing ultra-long DNA molecules with the Oxford Nanopore MinION by John M Urban, Jacob Bliss, Charles E Lawrence, Susan A Gerbi 
    Oxford Nanopore Technologies’ nanopore sequencing device, the MinION, holds the promise of sequencing ultra-long DNA fragments superior to 100kb. An obstacle to realizing this promise is delivering ultra-long DNA molecules to the nanopores. We present our progress in developing cost-effective ways to overcome this obstacle and our resulting MinION data, including multiple reads superior to 100kb. 
    Since sequencing is not NP-hard anymore,  then following Dick's thought process:

    Along those lines at some urging from John Tukey and others, I finally adopted what I called ``Great Thoughts Time.'' When I went to lunch Friday noon, I would only discuss great thoughts after that. By great thoughts I mean ones like: ``What will be the role of computers in all of AT&T?'', ``How will computers change science?'' For example, I came up with the observation at that time that nine out of ten experiments were done in the lab and one in ten on the computer. I made a remark to the vice presidents one time, that it would be reversed, i.e. nine out of ten experiments would be done on the computer and one in ten in the lab. They knew I was a crazy mathematician and had no sense of reality. I knew they were wrong and they've been proved wrong while I have been proved right. They built laboratories when they didn't need them. I saw that computers were transforming science because I spent a lot of time asking ``What will be the impact of computers on science and how can I change it?'' I asked myself, ``How is it going to change Bell Labs?'' I remarked one time, in the same address, that more than one-half of the people at Bell Labs will be interacting closely with computing machines before I leave. Well, you all have terminals now. I thought hard about where was my field going, where were the opportunities, and what were the important things to do. Let me go there so there is a chance I can do important things.
    and much like Dick and Yaniv, we should ask ourselves:

    What will be the impact of commodity sequencing on science and how can we change it? Let us go there so there is a chance we can do important things.

    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Tuesday, February 24, 2015

    Crossing the P-river (follow-up)

    A follow up to yesterday's blog entry. How much time did it take to gather all the information from the entire bacterial genome using a USB connected instrument that looks like a large memory stick and featured in the paper in Crossing the P-river: A complete bacterial genome assembled de novo using only nanopore sequencing data ?

     


    Wow ! You're seeing history in the making.
     
    Credit photo: this forum thread
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Monday, February 23, 2015

    Crossing the P-river: A complete bacterial genome assembled de novo using only nanopore sequencing data

     
     
    We interrupt our regularly scheduled program, you probably remember this entry from fifteen days ago:
     Here is a new piece of that puzzle, with the right sensor (the Oxford Nanopre MinION instrument a long read sequencer), things that used to be difficult are enabling the flooding of a new tsunami: A complete bacterial genome assembled de novo using only nanopore sequencing data by Nicholas James Loman , Joshua Quick , Jared T Simpson

    A method for de novo assembly of data from the Oxford Nanopore MinION instrument is presented which is able to reconstruct the sequence of an entire bacterial chromosome in a single contig. Initially, overlaps between nanopore reads are detected. Reads are then subjected to one or more rounds of error correction by a multiple alignment process employing partial order graphs. After correction, reads are assembled using the Celera assembler. We show that this method is able to assemble nanopore reads from Escherichia coli K-12 MG1655 into a single contig of length 4.6Mb permitting a full reconstruction of gene order. The resulting assembly has 98.4% nucleotide identity compared to the finished reference genome.
    The software pipeline used to generate these assemblies is freely available online at https://github.com/jts/nanocorrect.
     
     
     
    Related blog entries:

     

     

     
     
     
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Tuesday, January 06, 2015

    REMBO : Bayesian Optimization in a Billion Dimensions via Random Embeddings - implementation -

    I am a little late on this one, but here is a Random Features approach used to enable large scale bayesian computations:




    Bayesian optimization techniques have been successfully applied to robotics, planning, sensor placement, recommendation, advertising, intelligent user interfaces and automatic algorithm configuration. Despite these successes, the approach is restricted to problems of moderate dimension, and several workshops on Bayesian optimization have identified its scaling to high-dimensions as one of the holy grails of the field. In this paper, we introduce a novel random embedding idea to attack this problem. The resulting Random EMbedding Bayesian Optimization (REMBO) algorithm is very simple, has important invariance properties, and applies to domains with both categorical and continuous variables. We present a thorough theoretical analysis of REMBO, including regret bounds that only depend on the problem's intrinsic dimensionality. Empirical results confirm that REMBO can effectively solve problems with billions of dimensions, provided the intrinsic dimensionality is low. They also show that REMBO achieves state-of-the-art performance in optimizing the 47 discrete parameters of a popular mixed integer linear programming solver.
    Bayesian Optimization in High Dimensions via Random Embeddings by Ziyu Wang, Masrour Zoghi, Frank Hutter, David Matheson, Nando de Freitas
    Bayesian optimization techniques have been successfully applied to robotics, planning, sensor placement, recommendation, advertising, intelligent user interfaces and automatic algorithm configuration. Despite these successes, the approach is restricted to problems of moderate dimension, and several workshops on Bayesian optimization have identified its scaling to high dimensions as one of the holy grails of the field. In this paper, we introduce a novel random embedding idea to attack this problem. The resulting Random EMbedding Bayesian Optimization (REMBO) algorithm is very simple and applies to domains with both categorical and continuous variables. The experiments demonstrate that REMBO can effectively solve high-dimensional problems, including automatic parameter configuration of a popular mixedinteger linear programming solver.

    An implementation of REMBO can be found at: https://github.com/ziyuw/rembo

    Let us note the potential use of this technique for synthetic gene design

     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Sunday, December 14, 2014

    Sunday Morning Insight: The Stuff of Discovery



    In machine learning, they call the connundrum "exploitation versus exploration", in other circles we talk about "improving stuff versus discovery". There are many ways discovery can be defined. For instance in Crossing into P territory, we noted that a new kind of sensor (a genome sequencer) could enable experimentations in polynomial time thereby clearly expanding the possibilities to do real discovery. While both the PacBio and Oxford Nanopore technologies have just been made available before the summer, they are already changing the nature of discovery in that field [1]. Before, people would wonder how one could put pieces of DNA together, now that this complexity is mostly gone the new norm is now becoming: since genomes can be assembled easily, what sort of discovery can be done with a collection of genomes.
    As I've said before, we have had this exact same explosion unraveling in compressive sensing ten years ago. What happened since ? Many polynomial time algorithms were developed with the emphasis of being faster than the previous ones and soon enough more complex data structures began to be exploited by the algorithms. There is really no reason to believe why this should not happen in genome sequencing: We are going to have many algorithms that do alignement using either PacBio or Oxford Nanopore technologies.

    But in compressive sensing, something else happened: We began to discover a few things. This past Friday ( Hamming's time: Scientific Discovery Enabled by Compressive Sensing and related fields ) I provided two candidates. One candidate used the structure of the problem and its limitation to make predictions, while the other used the new paradigm to show how to nix a theory. Here are two other: An inference based on sparsity priors [2] (as featured in Catching an aha moment with compressive sensing ) and another one [3] about finding a needle in a haystack in an exponential families of solutions featured in ( Cluster expansion made easy with Bayesian compressive sensing ).

    In actuality, those four examples fall into two categories: One category is where one uses the new prior as a way to find that needle in an exponential haystack while the other category uses empirical complexity bounds to reduce the phase space of what is feasible.

    Either exploit the newfound capability or reduce the exploration horizon: different sides of the same discovery coin. Both are important.


    References:
    [1] Genomic sequencing: Recent tweets, papers, blog posts and attendant comment on that blog post:
    Widespread polycistronic transcripts in mushroom-forming fungi revealed by single-molecule long-read mRNA sequencing by Sean Gordon, Elizabeth Tseng, Asaf Salamov, Jiwei Zhang, Xiandong Meng, Zhiying Zhao, Dongwan Don Kang, Jason Underwood, Igor V Grigoriev, Melania Figueroa, Jonathan S Schilling, Feng Chen, Zhong Wang


    MinION nanopore sequencing identifies the position and structure of a bacterial antibiotic resistance island by Philip M Ashton, Satheesh Nair, Tim Dallman, Salvatore Rubino, Wolfgang Rabsch, Solomon Mwaigwisya, John Wain & Justin O'Grady
    And a comment following this article: USB-sized DNA sequencer is error-prone, but still useful

    Hi! Thanks for the write up! Getting spoken about on Ars Technica is definitely crossed something off my bucket list (I'm first author on the paper discussed).

    I would just like to say a few things about the MinION/Oxford Nanopore:

    1) While the error rate we observed is high compared to e.g. Illumina, it is comparable to PacBio (the main high throughput, long read tech).

    2) You say 'the great promise of nanopore sequencing has been very difficult to match in practice'. However, I don't really think that is true. What ONT have done is amazing!

    2a) First of all, the form factor is revolutionary. I'm not sure what your definition of a USB product is, but mine would be 'something where the only connection is a USB connection'. The MinION meets this.

    2b) Rather than limiting the MinION device to a small number of elite institutes, they sent it to hundreds of 'normal' people. This is a brave move that speaks to the confidence they have in their technology. We had a positive experience with it, some people probably less so, others more so. This approach to letting everyone have a crack is surely one to applaud?

    2c) The technology is just fantastic - single molecule sequencing using a biological pore! Think about how hard that must be to engineer! In a way that can be shipped to and used by hundreds of non-specialists! I should say that Illumina and PacBio also have awesome devices/technologies, but this one is newer ;-)

    3) A slight technical issue, but the short reads weren't used to correct the long reads. The long reads were used to join contigs made using the short reads.

    4) Another slight technical issue, the Illumina technology with bias is specifically the Nextera protocol. This has been known since this technology was developed by Jay Shendure's lab.

    Thanks again for writing us up! 

     
    [2] Direct inference of protein–DNA interactions using compressed sensing methods by Mohammed AlQuraishi, and Harley H. McAdams (featured in Catching an aha moment with compressive sensing )

    [3] Lance J. Nelson*, Vidvuds Ozolins, C. Shane Reese, Fei Zhou, Gus L. W. Hart, "Cluster expansion made easy with Bayesian compressive sensing," Phys. Rev. B 88, 155105 (Oct. 2013). [pdf] and Lance J. Nelson*, Gus L. W. Hart, Fei Zhou, and Vidvuds Ozolins, "Compressive sensing as a paradigm for building physics models," Phys. Rev. B 87 035125 (2013). [pdf] featured in ( Cluster expansion made easy with Bayesian compressive sensing )
     
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Sunday, October 19, 2014

    Sunday Morning Insight: Crossing into P territory

     



    To recap, in compressive sensing, it's been known for a while that some solutions can be found thanks to l_1 (P or Polynomial time) relaxation of combinatorial problems (NP). In fact, the whole field of compressive sensing took off when people realized one could be on the P side most of the time.

    In genome sequencing the latest long read technology have enabled the whole field to transport itself  from an NP territory into one where polynomial-time algorithms (P) will do OK. The threshold to cross is about 2K. Here is what we can read from the PacBio technology



    When you go in P territory, many things change, here is one:

    and here is what people say about the Oxford Nanopore technology.
     
     
     
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Thursday, October 09, 2014

    MONGOOSE: An exact arithmetic toolbox for a consistent and reproducible structural analysis of metabolic network models

    While genomics gets all the publicity, the next most important step is figuring out how proteins are produced and used into what is called metabolic networks. It is actually a key toward making personalized medicine a reality. A few readers might remember these entry entitled: Instances of Null Spaces: Can Compressive Sensing Help Study Non Steady State Metabolic Networks ?. followed by And so it begins ... Compressive Genomics, Well today's paper that has a co-author from the latter paper look at the first problem by figuring out specific metabolic networks using optimization tools.





    Constraint-based models are currently the only methodology that allows the study of metabolism at the whole-genome scale. Flux balance analysis is commonly used to analyse constraint-based models. Curiously, the results of this analysis vary with the software being run, a situation that we show can be remedied by using exact rather than floating-point arithmetic. Here we introduce MONGOOSE, a toolbox for analysing the structure of constraint-based metabolic models in exact arithmetic. We apply MONGOOSE to the analysis of 98 existing metabolic network models and find that the biomass reaction is surprisingly blocked (unable to sustain non-zero flux) in nearly half of them. We propose a principled approach for unblocking these reactions and extend it to the problems of identifying essential and synthetic lethal reactions and minimal media. Our structural insights enable a systematic study of constraint-based metabolic models, yielding a deeper understanding of their possibilities and limitations.
     
     
    An implementation of MONGOOSE is available from here.
     
     
    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Saturday, August 30, 2014

    Saturday Morning Video: Life at the Speed of Light - Craig Venter ( and some remarks )

    We mentioned Craig Venter before ( How Can Compressive Sensing and Advanced Matrix Factorizations enable Synthetic Biology ? ), he is a very inspiring speaker. This presentation by him is no exception. In fact, I believe there are a lot many issues that need to be addressed in this exploration and I have decided to create the Paris BioSciences Meetup group as a result. We'll about the scientific and technical aspect of some of the issues mentioned in this video and more, come join us if you are in the Paris area. Without further ado:



    A fascinating aspect of their work is how Craig's team devise a way to figure out which gene is good for life and which one it is not. Mostly by knocking them one by one and finding out which ones allow the cell to survive. This looks like a combinatorial approach. Maybe a group testing approach ( connected to comprressive sensing) might help. At about 38 minutes, he also makes the case that much of the genome is oversampled so that a specific capacity can be replaced by a different set of genes. 

    I note that his team uses a mix of previous and the new PacBio long read technology. 

    It was in the video we mentioned back in 2012 that sometimes a 1 base pair means the difference between life and death of the cell. There may be redundancy at the gene level but a one base pair difference between life and death point to other issues. 

    From an engineering standpoint, I note that the optics of the Illumina had to be reframed after the road trip (at 58 minutes).


    Relevant blog entries:
    1. Videos and Slides: Next-Generation Sequencing Technologies - Elaine Mardis (2014)
    2. Improving Pacific Biosciences' Single Molecule Real Time Sequencing Technology through Advanced Matrix Factorization ?
    3. DNA Sequencing, Information Theory, Advanced Matrix Factorization and all that...
    4. It's quite simply, the stuff of Life...


    Join the CompressiveSensing subreddit or the Google+ Community and post there !
    Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

    Printfriendly