Showing posts with label GreatThoughtsFriday. Show all posts
Showing posts with label GreatThoughtsFriday. Show all posts

Friday, March 02, 2012

Catching an aha moment with compressive sensing

Here is another example of Great Thoughts Friday a la Hamming: I mentioned it here and here, it is one of my favorite paper of 2011 ( it is behind a paywall  will shortly be is on Mohammed AlQuraishi's site at http://www.stanford.edu/~quraishi/protein-dna.html - thanks Mohammed !)    Vijay Pande  makes a compelling case as to why this approach is new in (Compressed) sensing and sensibility from which I'll quote the end:

.... AlQuraishi and  McAdams  (3) demonstrate an impressive advance in predictive capability: when applied to the prediction of the binding specificity of proteins to DNA, they found approximately 90% accuracy, compared with approximately 60% for the best-performing alternative computational methods. It will be exciting to see future applications of this method to other areas. With all its strengths, it is important to stress that this model does not completely solve the problem of transferability. Although these models are sufficiently regularized to avoid overfitting, they are still limited by the fundamental nature of the data used as input. This differs from a physics-based approach that, in principle, does not suffer from this issue of transferability, if all the relevant degrees of freedom are included in the model (9).
This suggests that a fusion of both approaches is particularly appealing, whereby physical properties are used as prior knowledge (i.e., as a prior in a Bayesian formulation) but then the model is derived from existing data. The ability to fuse data-driven and physics based approaches could push these types of models even further.
Another article relates the aha moment that led to this beautiful paper (if you have been doing computational chemistry you know what I am talking about)

AlQuraishi’s “aha” moment came when he realized that determining atomic-level energy potentials in protein–DNA complexes could be treated mathematically as a signal acquisition problem. By using the crystal structures of the complexes as the “camera” or sensors, and using the experimentally determined binding affinities of the complexes as the compressive measurements, he could determine the “signal”—the attraction between specific pairs of atoms in the protein and DNA. The crystal structures are typically obtained by X-ray crystallography or nuclear magnetic resonance (NMR) spectroscopy and are archived in the Protein Data Bank, a community resource for biological research.
“Measurements that we take for granted as being one kind of measurement, a structural measurement, can actually be seen as a different kind of measurement,” AlQuraishi explains. “In a sense it’s two different mathematical formulations, one that’s of use for the structural stuff, and one that’s used for the statistical compressed sensing stuff. On face value they look different. But I had been staring at each individually for a long time, and it occurred to me that with a simple transformation, you could get one to look like the other.”
AlQuraishi’s advisor and collaborator at Stanford, Harley McAdams, a research professor of developmental biology, explains why this transformation works: “You have to ask the question, Why does a protein bind to DNA? It’s because there’s a set of atoms in the protein and a set of atoms in the DNA that have an attraction to each other—that’s what we call a potential. The sum of these individual atomic attractions gives you the net binding energy of the protein to the DNA. If you were able to figure out the set of all atomic relationships or proximities within this structure—which you can from the crystal structures, and that’s where a lot of computation comes in—and you did that for a lot of different cases, then you could statistically determine which of these interactions are important.”
That’s exactly what the compressed sensing computation does, says McAdams: “Given what we know about the set of atomic proximities between the protein and the DNA from these different cases, and what we know about their binding energies, we can take that information and infer what are the atom-to-atom potentials.”
AlQuraishi sums it up by saying, “The idea is that these protein–DNA complexes serve as natural experimental apparatus, because we know where the two pieces, the protein and DNA, are; and we also know the energy of that interaction. So each complex is effectively a probe into the underlying biophysics of protein–DNA interactions at the atomic level.
“And it’s a direct measurement of the atomic-level interactions,” McAdams adds. “All other previous methods have not been that direct—they’ve been just purely statistical correlations, or based upon some hypothesized physical theory that would be applicable. But this method doesn’t use any hypothetical theory, it says OK, let’s just go in and measure it. That’s its power.”

The protein-DNA element is the (compressive) measurement. wow, just wow, a beautiful aha moment. The other thing I like about this paper is that they deifne this new De Novo potentials while substantially removing themselves from the traditional Van der Walls deterministic approaches. My bet is that only computational chemistry outsiders would take this risk. 



Compressed sensing has revolutionized signal acquisition, by enabling complex signals to be measured with remarkable fidelity using a small number of so-called incoherent sensors. We show that molecular interactions, e.g., protein–DNA interactions, can be analyzed in a directly analogous manner and with similarly remarkable results. Specifically, mesoscopic molecular interactions act as incoherent sensors that measure the energies of microscopic interactions between atoms. We combine concepts from compressed sensing and statistical mechanics to determine the interatomic interaction energies of a molecular system exclusively from experimental measurements, resulting in a “de novo” energy potential. In contrast, conventional methods for estimating energy potentials are based on theoretical models premised on a priori assumptions and extensive domain knowledge. We determine the de novo energy potential for pairwise interactions between protein and DNA atoms from (i) experimental measurements of the binding affinity of protein–DNA complexes and (ii) crystal structures of the complexes. We show that the de novo energy potential can be used to predict the binding specificity of proteins to DNA with approximately 90% accuracy, compared to approximately 60% for the best performing alternative computational methods applied to this fundamental problem. This de novo potential method is directly extendable to other biomolecule interaction domains (enzymes and signaling molecule interactions) and to other classes of molecular interactions.

Supporting information is here. The supplemental information (no paywall for that) starts with:

SI Methods
SI Methods describe the general methodology for de novo potential determination using compressed sensing, the determination of protein–DNA potentials, and the prediction of protein–DNAbinding sites and testing methodology. There are three sections. Section I contains the derivation of the general methodology for de novo potential determination using compressed sensing. The formulation in Section I is not specific to a particular application, but is applicable to a range of biological and chemical systems. The mathematical notation and terminology used throughout is in Section I.
This work brings together disparate concepts from information theory, statistical mechanics, and structural biology. The following background references may be useful to the reader:
• Linear and logistic regression (1).
• Compressed sensing (2, 3).
• Statistical mechanical ensembles (4, 5).
• Structural basis of protein–DNA interactions (6, 7).
Section II describes the application of the general methodology from Section I to the determination of a de novo protein–DNA potential. We reformulate the abstract constructs of the general methodology to the specifics of protein–DNA interactions and introduce several modifications that exploit the unique properties of protein–DNA interactions. Our choices for metaparameters and implementation details are also described in
Section II.
Section III contains a description of the use of the protein–DNA potentials described in Section II to predict protein–DNA-binding sites. We detail our structure-based approach toprotein–DNA-binding site prediction, the dataset used for training and testing, and the quantitative metrics used to compare results between our de novo potential method and previously published methods..



Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, June 03, 2011

Why Data Mining the Cell Phone Records In Germany is Important in Identifying the HUS Outbreak Source ?

This is a follow-up to Using the Cell Phone Data of Foreigners/Visitors to Find the German E.Coli Source

You will find below two documents from the European Center for Disease Prevention and Control and the Robert Koch Institute (RKI) that point to their level of understanding back on May 26/27th. The Health Ministry of Germany still seems to say that they don't know the source of the outbreak eight days later. The other news are hat the specialists are still asking people what they have eaten when in fact, the two reports below show you that you just need to be in physical contact (hand to hand) with people that have been exposed to it to get sick. In short, getting the cell phone records of sick people might be a much better way to evaluating second hand exposure than asking people if they have eaten raw tomatoes. If the exposure to a large part of the population is due to uncleaned surfaces touched by first exposed people then asking whether they have eaten tomatoes won't help. We now have 12 countries with citizens who have been exposed to this outbreak and so we are beginning to have enough statistics to find out the culprit. 


From a May 27th European Center for Disease Prevention and Control risk assessment:
"...The update provided by Germany on 27 May reports 276 cases of HUS since 25 April. While HUS cases are usually observed in children under 5 years of age, in this outbreak 87% are adults, with a clear predominance of women (68%). Cases in children of school age are also reported. Two people affected by HUS have died. The onset of disease relating to the latest reported case was 25 May. New cases are still being reported.
Laboratory results from samples taken from patients have identified STEC strain of serotype O104:H4 (Stx2-positive, eae-negative). A German study has shown that eae-negative STEC strains generally affect adults more than children. Two strains isolated from patients from Hesse and Bremerhaven were shown to be highly resistant against third  generation cephalosporins (ESBL) and resistant to trimethoprim/sulfonamid and tetracyclines. Most cases are from, or have a history of travel to, northern Germany (mainly Hamburg, Northern Lower Saxony, Schleswig-Holstein). Clusters of cases were reported from Hesse and linked to a catering company that supplies cafeterias. These most likely constitute a satellite outbreak.
The source of the outbreak has not yet been confirmed and intensive investigations are ongoing. German health authorities suspect that contaminated food is the vehicle of the outbreak, based on the epidemiological description (e.g. age and geographical distribution) of the cases. Current investigations are focused on raw vegetables. Preliminary results of a case-control study (with 25 cases and 96 controls) conducted by the Robert Koch Institute (RKI) and the health authorities in Hamburg demonstrate a significant association between disease and the consumption of raw tomatoes, fresh cucumbers and lettuce. Considering that the ongoing outbreak included many cases with a severe course of disease, the RKI and the Federal Institute for Risk Assessment (BfR) recommend people in Germany to abstain from consuming raw tomatoes, fresh cucumbers and leafy salads, especially in the northern part of the country, until further notice. Regular food hygiene rules remain in effect. .."



"...
Preliminary results of the STEC/HUS Case Control StudyPreliminary results of the epidemiological case-control study conducted jointly by the Robert Koch Institute and Hamburg Health Authorities show that the patients affected by the current EHEC outbreak consumed raw tomatoes, cucumbers and lettuce significantly more often compared to the healthy controls. However, whether only one or more of these three vegetables are associated with the outbreak remains unclear.Although the consumption of the described food items could explain the majority of the HUS-cases, other food items cannot be definitely excluded as the source of infection. The study was carried out in Hamburg only; therefore results cannot be generalized to other affected areas in Germany.In total 25 cases with HUS and 96 controls from Hamburg were included in the study since Friday, 20th of May 2011. They were matched according to gender, age group and area of residence. Detailed information concerning consumed food, eating habits and other possible sources of infection was compared between the hospitalized patients and healthy individuals (controls).The outbreak has affected the Northern part of Germany most severely, suggesting that the contaminated food items were mainly distributed there. Nevertheless, as HUS-cases were reported from other parts of Germany as well, contaminated food items could be present in other regions as well.As the outbreak is still ongoing and the public health impact is serious, the RKI and the BfR recommend as a precaution until further notice - in addition to the usual hygienic measures concerning handling of fruits and vegetables – not to eat raw tomatoes, cucumbers and lettuce, especially in the Northern Germany.It remains of vital importance for all persons with diarrhea to follow strict hand hygiene, especially if in contact with small children and immunocompromised individuals. Recommendations for good kitchen hygiene practice, as described in the BfR information sheet (www.bfr.bund.de), remain valid.Date: 26.05.2011.."



Great Thoughts Friday: Using the Cell Phone Data of Foreigners/Visitors to Find the German E.Coli Source

Dirk sent as a comment to the suggestion to use cell phone data to uncover the source (or sources)

Well that approach seems interesting. However, a first issue is that Germany and especially the Germans are peculiar with data privacy and probably there will be an outcry in the media about such an approach. A second issue: How should one interpret the outcome of such an approach. Isn't is probable that the outcome will be something like: "These people visited public restroom frequently" or "have been at Pharmacies or Doctors". In short: How to distinguish cause and effect?

On the first issue: In a democracy, the use of data that has already been collected can always be used for other purposes as long as you have the right filtering tool in effect. I am sure Germany can find ways to set up a privacy commission that oversees the use of this data in times of crisis. Five hundred sick people and  some twenty dead  count as a crisis in my book. 

On the second issue, I recall how John Snow discovered the source of a Cholera source in London: By mapping the location of the sick/dead people by their habitation.



Even he had to do additional thinking after having collected the data (use a diffusion model) to find the epicenter of the disease ( a water handle on Broad street). The way I see I see it, the authorities have no clue. We still don't know if it's in the water system, transportation system, airborne, foodstuff, everything is on the table. Throwing this new layer of information might put additional constraint on the infinite number of solutions they already have (an underdetermined system of some sort).


Of specific interest are the outliers: the foreigners and the Germans not living in Hamburg. Some of them were there for a short time and so are likely the ones whose locations are most interesting. They are probably the ones putting the most constraints on this underdetermined system.


Saturday, May 28, 2011

Great Thoughts Friday: Kickstarter as a way to fund some of these crazy projects

I was thinking about this yesterday, but did not get around to writing this entry before today. So it really is a Great Thoughts Friday entry

Some of you may not know about Kickstarter. it is a site where one features project that requests funding from people on the interwebs. Examples for the technology or photography section include:











Following some of the themes mentioned here, what could be some of the crazy projects thast could be funded through Kickstarter ?

  • an iPad/iPod app that downloads automatically all the pdfs featured on nuit blanche for offline viewing
  • fly a high altitude balloon to get a hemispherical view of the horizon
  • development of a platform to be flown on a high altitude balloon to allow for an on board camera to be directed from the ground
  • support travel for travel and cost of living if needed for Sparseman. if he sends something to the Google Exacycle program
  • build a multispectral imager out of an iPhone4.
  • ?????


Friday, May 20, 2011

Great Thoughts Friday: Unsupervised Feature Learning of Exotic Camera Systems

The issue of sparse coding is not new to our readers since it falls in the dictionary learning issue of interest in the reconstruction stage in compressive sensing. Thanks to Bob, I was reminded of the following video of Andrew Ng about unsupervised learning (one way of performing dictionary learning). As Andrew shows some of these techniques are getting better and better.


By necessity, people are using the same databases. But if we take the argument further, some of these databases are somehow dependent on the very camera parameters that took these shots and one could make the argument that the calibration issue is hidden behind some well thought out benchmark databases. While one of the need for decomposing images is important, eventually, we are often not knowledgeable about the features that were learned (an issue Andrew points to at the very end of his video/slides.) Given all this, I wonder if exotic camera systems such as multiple Spherical Catadioptric Cameras ( as featured in Axial-Cones: Modeling Spherical Catadioptric Cameras for Wide-Angle Light Field Rendering by Yuichi Taguchi, Amit Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam and Ramesh Raskar) or the random lens imager that produce highly redundant information of the scene of interest, can provide better classification rate than the current techniques ?


If we take this argument further, what would then be a good way of comparing these exotic camera systems for classification purposes ?...

Note:
For those of you interested: At the end of the Andrew Ng's presentation, there is a link to this course and handouts of the corresponding course at Stanford. It features two videos.

Random Earthquake Thoughts

The following are random thoughts that may be good for someone who is on a Great Thoughts Friday mode.  Large Earthquakes are few and could really be considered as sparse objects but since we are talking about time series, shouldn't some of the tools used in detecting those in high dimensional time series be using rank minimization techniques mentioned this week ? By high dimensional time series, I don't just mean acceleration traces from different geographical locations but rather more complex data from other types of sensor networks.

By reading a little more about the potential connection between earthquakes and ionospheric measurements in Ionospheric Precursors of Earthquakes; Recent Advances in Theory and Practical Applications by Sergey Pulinets, I decided to watch for a day the ionospheric TEC measurements from JPL updated every five minutes




The Ionospheric TEC map are constructed from tomographic measurements performed between GPS satellites and some ground stations. For a day, I noticed something obvious: direct sunlight on a region increases the electron content in the ionosphere over that region albeit not uniformly. Where is the non-uniformity coming from ? I am not sure but if one follows the theme of the paper above, these non-uniformities may be linked to solar and/or geomagnetic activity and gases releases in the cases of earthquakes and therefore, they might be co-located to earth faults.




Both the TEC map and the night map seem to correlate. Now I wonder how these maxima average over time and correlate with the known faults ? or


Earthquake Density Maps for the World from the USGS)


one wonders if any of these maxima can also be correlated with geostationary observation of earth or the activity on the Sun. Over the past day, I didn't see an obvious connection with the last two. 

Let us note that the general public cannot seem to have access to previous ionospheric TEC maps

In a different direction, some folks are also looking at earthquake prediction based on a network analysis of past events. The paper is Earthquake networks based on similar activity patterns by Joel N. Tenenbaum, Shlomo Havlin, H. Eugene Stanley. The abstract reads:
Earthquakes are a complex spatiotemporal phenomenon the underlying mechanism for which is still not fully understood despite decades of research and analysis. We propose and develop a network approach to earthquake events. In this network, a node represents a spatial location while a link between two nodes represents similar activity patterns in the two different locations. The strength of a link is proportional to the strength of the cross-correlation. We apply our network approach to a Japanese earthquake catalog spanning the 14-year period 1985-1998. We find strong links representing large correlations between patterns in locations separated by more than 1000 km. We also find significant similarities in the network structure upon comparing different time periods.

The results are not overly convincing but I wonder if this type of analysis would not be amenable to a diffusion wavelet approach.

Liked this entry ? subscribe to the Nuit Blanche feed, there's more where that came from

Friday, May 06, 2011

It's Friday, it's ... Great Thoughts Time

Richard Hamming said it better in his speech entitled You and Your Research:

Along those lines at some urging from John Tukey and others, I finally adopted what I called ``Great Thoughts Time.'' When I went to lunch Friday noon, I would only discuss great thoughts after that. By great thoughts I mean ones like: ``What will be the role of computers in all of AT&T?'', ``How will computers change science?'' For example, I came up with the observation at that time that nine out of ten experiments were done in the lab and one in ten on the computer. I made a remark to the vice presidents one time, that it would be reversed, i.e. nine out of ten experiments would be done on the computer and one in ten in the lab. They knew I was a crazy mathematician and had no sense of reality. I knew they were wrong and they've been proved wrong while I have been proved right. They built laboratories when they didn't need them. I saw that computers were transforming science because I spent a lot of time asking ``What will be the impact of computers on science and how can I change it?'' I asked myself, ``How is it going to change Bell Labs?'' I remarked one time, in the same address, that more than one-half of the people at Bell Labs will be interacting closely with computing machines before I leave. Well, you all have terminals now.

Ok, I don't know if those a great thoughts, but here are some of my questions:


  • Can we have dumb AI ? (especially since memories and CPUs are cheap).
  • If I were training a brain inspired computational visual system on photos and text and I were to get that system to read webpages: How much time would it take to get a similar heat map as that of a human ?
  • What does it mean to be past Peak Oil ? (figure from here) while at the same time potentially facing global warming ?



With one billion internet enabled cameras:

  • Can we do a better job at diagnosing at 6 months old, conditions like Autistic Spectrum Disorder ?
  • Can we use star tracker algorithms for finding stars to perform regular check for skin cancer detection?
  • How long before people share their stools for diagnosing purposes ?
  • How will people share private data for these new diagnosing tools to emerge ?

And you, what are your thoughts ? they don't have to be questions!

Printfriendly