Showing posts with label Genetics. Show all posts
Showing posts with label Genetics. Show all posts

Thursday, November 16, 2017

Ginkgo; Good or Bad

I have a dozen or more Ginkgo trees, all from seed from the New York Botanical Garden trees, which were allegedly grown from seeds from the surviving Hiroshima ginkgo. The seed have just dropped and they do smell a bit rancid but after a month or two they disappear almost over night, the squirrels. Now allegedly the ginkgo increases brain capacity. So am I allowing for the evolution of brilliant squirrels?

In a recent paper by Oliveira et al they note:

Despite the ancient use in Chinese popular medicine and, more recently, in western modern medicine in many European countries, the biological effects of extracts of G. biloba (GBE) are still not clearly known. In modern medicine GBE has been used for tinnitus, to reverse memory loss, for dementia, and Alzheimer’s and Parkinson’s diseases in elderly people. Besides reports on improvement of blood circulation in the brain, there are a number of studies pointing to complex cellular effects, involving signal transduction pathways and epigenetic modifications. Evidence are presented from recent reports concerning genotoxic and antigenotoxic properties and the corresponding mechanisms underlying such activities, mostly regarding the prooxidant and antioxidant activities of the extract. However, several examples of direct interaction of the extract and its components with specific proteins are provided, especially for DNA damage repair, contributing for antigenotoxicity. Evidence of epigenetic effects of GBE are also presented from approaches involving transcriptomics, detection of activity of histone deacetylases, and screening of plant extracts with cell-based systems for detection of posttranslational modifications. The modulation of chromatin-remodeling enzymes by GBE and their interaction with proteins involved in DNA damage repair, apoptosis, and signal transduction are discussed in the context of neurodegeneration.

It is not clear what the implications are. But I have noticed my squirrels borrowing some books from my library. The works all related to editing human DNA. Should we be concerned?

Wednesday, May 31, 2017

What is a Species?

That question has been around for a while. Mayr had proposed a definition that essentially said a species was a collection of animals or plants which had the capacity of breeding among each other. As Kevin de Queiroz has stated:

Ernst Mayr played a central role in the establishment of the general concept of species as metapopulation lineages, and he is the author of one of the most popular of the numerous alternative definitions of the species category. Reconciliation of incompatible species definitions and the development of a unified species concept require rejecting the interpretation of various contingent properties of metapopulation lineages, including intrinsic reproductive isolation in Mayr's definition, as necessary properties of species. On the other hand, the general concept of species as metapopulation lineages advocated by Mayr forms the foundation of this reconciliation, which follows from a corollary of that concept also advocated by Mayr: the proposition that the species is a fundamental category of biological organization. Although the general metapopulation lineage species concept and Mayr's popular species definition are commonly confused under the name “the biological species concept,” they are more or less clearly distinguished in Mayr's early writings on the subject. Virtually all modern concepts and definitions of the species category, not only those that require intrinsic reproductive isolation, are to be considered biological according to the criterion proposed by Mayr.

In the current Nature there is an ongoing debate as to the many conflicting definitions. The authors state:

The assumption that species are fixed entities underpins every international agreement on biodiversity conservation, all national environmental legislation and the efforts of many individuals and organizations to safeguard plants and animals. Yet for a discipline aiming to impose order on the natural world, taxonomy (the classification of complex organisms) is remarkably anarchic. There is reasonable agreement among taxonomists that a species should represent a distinct evolutionary lineage. But there is none about how a lineage should be defined. 'Species' are often created or dismissed arbitrarily, according to the individual taxonomist's adherence to one of at least 30 definitions. Crucially, there is no global oversight of taxonomic decisions — researchers can 'split or lump' species with no consideration of the consequences.

 The problem is that especially in plants we have a significant amount of inbreeding potential between what we call species. The plants have the same number of chromosomes of the same length with genes in the same positions, almost. Then why call them different species? All too frequently it is in the eye of the beholder.

Friday, July 22, 2016

Assigning Credit

Back in the early 50s with Watson and Crick, one could argue over perhaps a dozen people at best who were in the fray. The paper had two authors.

Today we have papers with in excess of a thousand authors. The recent CRISPR debate provides focus on this issue.

In the recent Nature article there is an excellent discussion of how best to attribute what to whom.

They note:

The history of CRISPR–Cas9 gene editing has become a subject of fierce debate and a bitter, high-stakes patent battle. Researchers and institutes have been jostling aggressively to make sure that they are credited for their share of the work in everything from academic papers to news stories.

 They continue:

In January, Eric Lander, president of the Broad Institute of MIT and Harvard in Cambridge, Massachusetts, tossed into this minefield a historical portrait called 'The Heroes of CRISPR'2. It was instantly controversial. Some said that it marginalized the contributions of certain researchers, and they questioned the decision to publish the article without a conflict-of-interest statement noting that the Broad Institute is embroiled in a patent dispute that hinges on determining who invented CRISPR–Cas9 gene editing.

Lander may very well have stepped into a hornets nest. One suspects he was just trying to lay out the best understanding of what occurred. That is always useful. But in today's world we have a proliferation if not explosion of post docs and junior faculty. How does one best account for all the conversations, insights, bench work, and the like.

It continues:

Outside that community, however, the accolades continue to be heaped on senior investigators. “We need to invent ways to expand the medals podium,” says Lander. “The idea that scientific discovery involves just one, two or three people is so nineteenth-century.”

Science is no the Oscars. There is no Best Director or Best Film. It is an incremental process of incremental contributions until that one moment that some it comes together. The biological sciences especially is a team effort, and for better or worse the team may be what gets recognition. The recent Silicon Valley attempt to make this Hollywood just intensifies the star issue, a politically correct stardom at that.

Sunday, June 26, 2016

The Ignorance of Some Academics

In a recent review of the book Life’s Greatest Secret: The Race to Crack the Genetic Code by Matthew Cobb, a book I admit I have neither read nor do I have any intent to waste my time thereupon, there is however a letter to the editor of the New York Review of Books, NYRB, a generally leftist oriented opinion publication, which stated some rather revealing opinions.

In the above mentioned letter to the NYRB the author states regarding the original review[1]:

He notes that developments in mathematical information theory and cybernetics soon after World War II had a strong influence on the way biologists began to talk about life and heredity —specifically the idea that DNA contains a “genetic code” replete with “information.” Much the same point can be made about psychology and neuroscience: they too were influenced by these developments in the mathematical theory of information, introducing this notion into the heart of theories of the mind and the brain. But Orr goes on to remark:

"In the end, the information sciences provided biologists with loose but useful metaphors and analogies, a language that allowed scientists to think and speak in new ways. But the high-powered mathematics of these fields proved mostly impotent in biology [I would add in psychology and neuroscience too]. No one, for instance, used Shannon’s equations to say anything especially interesting about organisms…"
.
This raises a troubling question: If these recent ways of talking in biology, psychology, and neuroscience are really just loose but useful metaphors, now deeply ingrained in these sciences, what is the literally true way of speaking for which they substitute? How can we reformulate these sciences in such a way that the information metaphors are replaced by sober statement of fact? And do scientists now agree that the borrowed way of talking really is just loose metaphor, or have they come to take it for literal truth? This question seems to me not sufficiently addressed, though very important.
Now the original Reviewer notes[2]:

A second theme concerns the respective roles of theory (of any sort) versus experiment in biology. In the early 1960s, mathematicians confidently declared that “it will be interesting to see how much of the final solution [to the coding problem] will be proposed by mathematicians before the experimentalists find it.” As Cobb concludes, the “answer. . was simple: not one single part of it.” The interesting question is why theory failed here. Part of the answer, as Cobb emphasizes, is related to Crick’s idea of the frozen accident. The genetic code seems at least partly arbitrary. It represents a half-decent arrangement arrived at by the imperfect, tinkering process of evolution by natural selection and, once settled on, it couldn’t be “improved,” or made somehow more systematic. In such a situation theory is likely useless.

Let me examine these statement a bit in detail.

First:

Let me reiterate Wiener and Cybernetics. He made basically the following observations:

1. The world is filled with uncertainty. Things are random, and we have to acknowledge and accept that.

2. Many organic entities are systems. Namely they have actuators and effectors. They have cause and effect. They are in effect a system which means we can model cause and effect, albeit under condition 1 above it may very well be random.

3. Systems have feedback elements, namely inputs yield outputs which in turn can effect inputs. That means we have systems whose dynamics are uncertain systems with dynamic effects.

In simple terms the Cybernetic world is a stochastic dynamic system. Now what about cells, DNA and their functions? They are stochastic dynamic systems. Ligands attached to receptors which activate pathways which start DNA reading via a promoter and conversion which produces RNA and then produce proteins which then become ligands. Some proteins actually modulate the pathways and receptors. Thus the dynamics of cells is a stochastic dynamic system. Almost all studies in cancer pathways revolve about that fact. Cybernetics from a Wienerian mindset is fully accepted in systems biology. It is the very heart of systems biology.

The letter writer is thus in error. The Cybernetic model is hardly a loose model. It is at the very heart of understanding cancer dynamics. It is necessary. The Reviewer, the Author and the Letter Writer are in gross error in their understanding and articulation of cybernetics.

Furthermore, as we have shown in repeated malignancies, one can view cancer as a separate dynamic stochastic organism, growing apart from the human host. Yet if one accepts such a model, then it is possible using systems approaches to use systems identification theory to identify the system and systems control theory to mitigate the threat from this organism. One misses the point in examining single cells, one must view the amalgam, albeit a heterogeneous organism genetically.

Second:

Now for Shannon. Frankly his approach is quite limited. I started teaching Information Theory at MIT more than fifty years ago, and even earlier with a Wienerian view. Shannon and Information work on communications. Frankly the very term information is a misnomer, mainly since we do not know what we mean by information from an epistemological basis so by applying this terms to data bits was cute and catchy but does it a disservice because people who fail to have any understanding of it will err in its application. Shannon was interested in signals sent over a noisy channel where there were limits in transmission capacity, say bandwidth.

He has two main theorems. The Channel Coding Theorem which says how fast or at what rate you cane transmits a signal over a channel at a rate where the channel has a capacity in say bandwidth and an interference in terms of noise. There frankly is no real "information" here. Secondly is the Source Coding Theorem which says how much you can compress some signal with excess stuff, called information. For example, we can compress voice to a few hundreds of bits, 0 and 1, per second or we can likewise do the same with video at so many thousand bits per second. The Source and Channel Coding Theorems of Shannon describe dealing with redundancy and dealing with limited capacity and noise respectively. That's it folks! It does not tell you about "information", whatever that means to someone. To Shannon it was changing a sine wave to a bunch of on and off signals. That's all folks.

So when one sees articles, letters, books like this one shudders and understands why we have so many poorly educated students. It’s the teachers stupid!

Now; do biologists use any of these approaches? Think Eric Lander and the human genome. How do you think he got to match all the broken pieces of DNA into a genome; mathematics? After all he was a trained engineer in coding, that Shannon stuff. So he may not have used the two Theorems but he did employ the ideas resulting therefrom.

Pity we have people who make these statements. But alas this is all too common.

Thursday, May 5, 2016

Lysenko and The Soviet Academy



Graham has written a wonderful book on Lysenko and the Russian School of Genetics during the Stalin era. Lysenko viewed inheritance in the sense that certain characteristics could be handed down in generations based upon environmental factors experienced by parents. That is the change in a genetic makeup was not solely due to genetic changes per se. He could turn summer wheat to winter wheat by getting it used to a change in weather. Thus he did not need a genetic alteration but an environmental alteration was sufficient. In a sense the concept did play into the hands of the Marxist reasoning.

Graham blends the understanding of epigenetic changes that are currently being understood with the ideas of Lysenko and asks if this new understand then justifies Lysenko's ideas. On the other hand, Graham details Lysenko's way of dealing with his academic adversaries often resulting in their imprisonment and demise. The current understanding of gene expression and thus phenotype is that genes can be turned on and off by such epigenetic factors as methylation. Methyl groups bind to the nucleotides and also suppress expression directly by blocking the gene or indirectly by blocking transcription factors.

This is somatic epigenetics. Germ line epigenetics, parent to child has also been observed. Namely effects on the parent causing epigenetic changes can be handed down to the child, where it was assumed that the methylation of certain bases was eliminate but somehow they can be preserved. Thus, in a simplistic sense, an environmental change imprinting the parent can imprint the offspring. This may or may not be consistent in a broad sense with Lysenko but the author discusses it in some detail. Graham's discussion is limited as one would expect in a short book of this type but he does explain some of the issues well including the event of the "Dutch Winter", an epigenetic benchmark.

Graham has a wonderful discussion of his opportunistic meeting with Lysenko at a lunch table in the Russian Academy, and the brief attempt to elicit some explanation from Lysenko. Lysenko was as one would expect defensive since this occurred after he was taken down from his perch yet retained his academic credentials. This discussion is quintessential east meets west based upon my personal experiences in Russia when first meeting some notable. It was clear from Graham's description that Lysenko was still wary especially since Graham had been critical of him in Graham's prior writings.

Graham also presents a clear and coherent discussion of the players in this tragedy, the geneticists following the true path and how Lysenko and his actions resulted in their fall.

The only point that would have been useful to explore would be the need by the Marxist theorists to have a Lysenko position versus a Darwinian one. I had seen this battle with the probabilists. Marxist theory is deterministic and probability is its enemy. Yet many probabilists managed to work and prosper. Individuals like Gnedenko, Kolmogorov, Stratonovich, Markov and others developed the basis for stochastic processes that we see used in fields as broad as finance with the Black-Scholes theorem in options trading, a thought anathema to the Marxists. Graham does provide some insight but it would be worthwhile to have a more in depth discussion of this potential conflict.

Overall the book is an excellent addition to understanding both the Russian Academy and its functioning, the Stalinist management of the overall society, and a petri dish model of Academic infighting. It is very worthwhile for those seeking to understand both Russia as well as the politics of Science, albeit in a different vein.

Saturday, October 31, 2015

Genetic Network Modelling

The bench biologist is in a continual search mode for some new gene interaction. What new gene can be a, not the, cause of say prostate cancer. At the other extreme is the systems biologists who use their mathematical models to propose reactions. The intersection of these two is not that fruitful as of yet.

In Science the authors conclude:

Models are simplified (but not simplistic) representations of real systems, and this is precisely the property that makes them attractive to explore the consequences of our assumptions, and to identify where we lack understanding of the principles governing a biological system. Models are tools to uncover mechanisms that cannot be directly observed, akin to microscopes or nuclear magnetic resonance machines. Used and interpreted appropriately, with due attention paid to inherent uncertainties, the mathematical and computational modeling of biological systems allows the exploration of hypotheses. But the relevance of these models depends on the ability to assess, communicate, and, ultimately, understand their uncertainties. 

 The process is iterative. Models are built, tested, found lacking, and then reiterated. The challenge is that these are quite complex and of massive dimensions. Perhaps methods akin to 19th Century thermodynamics may play a role, gross constructs like enthalpy and Gibbs free energy, but perhaps not.

It will take time to get these models to work properly, but they are needed as a cornerstone of true science.

Tuesday, November 4, 2014

A Fantastic Book

There is a draft of a book, Cell Biology by the Numbers, also see dropbox, by Milo and Phillips, that is just amazing! It is a set of questions with answers briefly given, such as how fast is translation and transcription. Also how big is a cell. It is really a fantastic amalgam of questions, answers, and technique. It is not yet published but I would argue that it should be on the desk of anyone trying to do real work in the life sciences. The questions and answers alone are a gold mine but the way they arrive at them and explain them are sine qua non.

Keep a watch out for this one, it is a true gem!

Monday, May 19, 2014

Alan Turing: Calico Cats, Zebras, and Daylilies


Turing in 1954 wrote his paper on Tessellation, the patterns we often see on animals and plants. The question is; is Turing’s approach the right way to understand these patterns? Namely Turing proposed a hypothetical method whereby cells communicate with one another and that this communications is akin to flows of some, as of 1954, yet to be defined chemical substance or substances. Depending on the concentrations of these substances the cells then turned, for example, black or white, as in a zebra, and that this flow being coordinated in some manner yielded a pattern, not just a mass of black and white hairs.

Turing hypothesized, for example, that there may be two controlling molecules in varying densities and if one molecule was denser than the other it would turn on white and otherwise it would turn on black. But Turing said more, that the flow of the molecular density was not just random but that cells somehow participated in a distributed manner so that the densities flowed as waves, with peaks and valleys. Thus the Zebra stripes were a reflection of this flow. When the black molecule was at a high the hairs were black and when the white ones were high the hairs were white. The net result looked like waves of white and black.

Now a second model that has become of recent popularity is the explanation for the Calico Cat.


The explanation for this is dramatically different. Here Calico Cats are all female. The way it works is the epigenetic silencing of one of the X chromosomes. This allegedly is totally done at random. As seen above this of course is hardly the case. If every hair cell were random then we would expect to have a blend of two or more colors and not the patterns that we have above. This means that if this is epigenetic that spatially there is some mechanism that is not totally random. There is some form of cell to cell memory and cell to cell thresholding. Namely the hair stays black, brown or white for some spatial period and then switches to the other state. That is the Turing Tessellation effect. What then does that?

Now consider a third example; the daylily. We show a typical example below. Here we have an eyezone, the dark red around the milled and we even have red on the edges. This is again a tessellation type as described by Turing. There are areas where there is dark red and areas where there is light red or pink. Is this epigenetic, genetic, or a Turing tessellation.


In fact daylilies can show dramatic patterning as they get more sophisticated. The above is a simple form of patterning, a wave of dark and light red.

Thus what enables these patterns? In epigenetics of cats, it is the turning on and off of one X chromosome, but not really randomly, so there must be some mechanism that selects which one gets wrapped in an lncRNA and which does not. What is that activator? Not yet known.

In daylilies we know that color is driven by anthocyanin production. More of one and we get one color and more of the other another color. We also know that certain proteins, gene products, act as catalysts facilitating one anthocyanin path or another. Thus if one gene is producing then it may drive up pone anthocyanin or another. What turns these genes on and off? Perhaps epigenetic methylation as we see in many other examples. But that begs the question of what causes the methylation? It appears to be a pattern like that of the Calico Cat.

Thus the Turing Tessellation is a process that explains these patterns but the facilitator of that process, the extracellular or intercellular molecule is not known. Thus we have an interesting area of exploration. In addition we see similar effects in the field of metastatic cancers, where we get clusters of metastatic cells, and not just random aberrant one.

Saturday, March 29, 2014

Genome Size

There is a piece in Nature which has some interest. It is the determination of the genome size of the pine tree, Loblolly. This is a pine which may make it as far north as where I am in New Jersey. It is a bit strange in that it grow branches only on the side where there is strong sunlight.

Nature states:

A species of pine tree native to the southeastern United States has a genome with 23 billion base pairs, more than 7 times the length of the human genome....Another team, made up of many of the same researchers and led by Jill Wegrzyn at the University of Connecticut in Storrs, characterized around 50,000 of the genes and estimated that 82% of the loblolly genome is made from repetitive elements. This work, the first pine genome assembled so far, provides a foundation to study the biology of conifers, the authors say.

As a question to pose: How would a plot of gene length versus lifetime of species appear? Namely would say a Ginkgo have lots or excess base pairs or what? Just a thought.

Thursday, March 13, 2014

Too Much Information

In a recent JAMA paper on whole genome sequencing the authors examine 12 patients in detail and the results were mixed. One patient had BRCA mutation which was beneficial. The others were a mixed bag.

As the authors conclude:  

In this exploratory study of 12 volunteer adults, the use of WGS was associated with incomplete coverage of inherited disease genes, low reproducibility of detection of genetic variation with the highest potential clinical effects, and uncertainty about clinically reportable findings. In certain cases, WGS will identify clinically actionable genetic variants warranting early medical intervention. These issues should be considered when determining the role of WGS in clinical medicine.

The problem is several fold. First there are many know genes with uncertain effects. Second there are many unknown genes with totally uncertain effects. Yes the genome has been mapped for over a decade but the unknow genes are "known" but their effects are uncertain. Third there are many epigenetic effects which are uncertain. Fourth many cancers are the result of subtle in lesion changes not reflective of a large scale sample.

The authors continue:  

As technical barriers to human DNA sequencing decrease and the cost of whole-genome sequencing (WGS) approaches $1000, WGS and protein-coding genome sequencing (whole-exome sequencing [WES]) are increasingly used in clinical medicine. Both WGS/WES can successfully aid clinical diagnosis, reveal the genetic basis of rare familial diseases, and explicate novel disease biology. Regardless of context, even in apparently healthy individuals, WGS/WES are expected to uncover genetic findings of potential clinical importance. However, comprehensive clinical interpretation and reporting of clinically significant findings are seldom performed. As WGS/WES are applied more broadly, questions have been raised about the duty for discovery, interpretation, and reporting of clinical findings. Recently published recommendations define genetic variant types in a minimum list of inherited disease genes that are suggested to be subject to discovery, reporting, and clinical follow-up regardless of the primary indication for sequencing, patient preference, or patient age. Despite this, the technical sensitivity and reproducibility of clinical genetic findings using WGS and the clinical opportunities and costs associated with discovery and reporting of these and other clinical findings in WGS data remain undefined.

 The problems that can be seen is that many patients may demand the tests or physicians can see a way to "sell" the tests and the result will be an added load on the already burdened Health Care system to deal with what is at best conjecture. One wonders how this fits into the ACA?

Saturday, February 22, 2014

Neanderthals

The Neanderthal Man, by Svante, is a compelling recount by a principal in the discovery of genes of the Neanderthals. It starts with the interest in recovering DNA from old sources, and in this case some liver bought at the local market and then desiccated in an oven at 50C. The tale spans over some twenty years, with diversions typical of science, and ultimately ends with the publishing of some of the most interesting results in understanding man and his evolution.

Svante is an exceptionally good writer and the tale flows quite smoothly. If one understands the science, then one can fill in the gaps and the tales is well presented. If one does not understand the science then one can still appreciate what is happening by taking the results presented at face value.

The tale works back and forth from the fundamental science to the interrelationships between various players in the overall search. Svante shows how he managed to deal with the anthropologists and others to get samples of Neanderthals from as far away as Siberia. It also demonstrates some of the more cooperative nature of science as new techniques is shared and how Svante is assisted by many others who are but in related fields.

The efforts span from California to Eastern Russia and it shows that in today's environment the ability to communicate changed what would have been multi-lifetime efforts into a fast paced move to provide the final answers.

This book is a stark contrast to Watson's Double Helix. The Helix is a strong interplay of personalities; it portrays competitiveness and at times pettiness that is common in certain scientific endeavors. Helix was a true race, a sprint to get DNA right, and a succinct set of observations which became the underpinnings of Svante's efforts. Svante is the opposite of Watson. The ego is missing; the collegiality if present, yet one still sense the pace. Yet it is not a pace with an edge, it is a steady pace to get it right.

This is definitely a great book for those seeking to understand the Neanderthal developments as well as understanding perhaps how the research community has matured as it has expanded.

Also, upon some reflection, I recall when I first read Watson's Double Helix just after it was published I could recognize the highly competitive world of research since I was still at MIT. In contrast Svante portrays a totally different world, one more of communications and cooperation. The worlds of Watson and Svante are separated by some half century, and the difference is startling, one is near ruthless and the other collegial, with a sense of cooperation moving forward. Great job!

Monday, December 2, 2013

Science: Principles or Cook Book?

There is a recent paper by Prof. Dougherty from Texas A&M which bemoans the state of some parts of science in the current environment[1]. As Dougherty so clearly states: 

…science concerns relations between measurable variables and it is these relations that constitute the subject matter of science, scientific knowledge ipso facto is mathematically constituted… 

 Let me give a couple of examples of how this applies.

First, let us look at the world of genomics which I have been discussing herein for a while. The introduction of the microarray has allowed an explosion of data that has then allowed scientists to putatively argue some relationship between genes and cancers. Namely they go about examining say 9,000 prostate cancer patients and using microarrays primed for say 500 genes they conclude that say some 50 of these gene are seen in prostate cancer. They then allege that there is some actionable clinical relationship between the presence of the gene and the cancer. There is no underlying system model identifying this, just a microarray demonstrating that “oftentimes” these genes are under or over expressed.

Second, let us look at the BRAF V600 melanoma cases. Here unlike the above we have a case where one knows the RAF pathway and that loss of control of certain elements of that pathway lead to gene instabilities and thus a malignant expression. Therefore one targets the mutated RAF gene, the BRAF V600, and it results in a suppression of the malignancy, for a while. Then we had squamous cell carcinomas, but since the full pathway was known, go down one step and there was MEK and controlling it controlled the sequella. In this case there was a model, a system, and by logically following the system one found what the next step should be.

The above are two examples of how “science” is being done today in the area of gene related results. The second example is a Dougherty like science, namely it connects data to an underlying model which is predictable, and by using that the cancer is controllable, at least until another instability results. The first model, data collecting, is not really science as we accept it today. It is more akin to 19th century Botany, at best, where one goes out and collects specimens of plants and then tries to sew together a quilt of understanding to explain nature.

What Dougherty is focusing on is the Why question. When I recall Medical School, one is taught What and How. What disease is it and How do I treat it. In contrast Engineering is first Why and then How. There is a strong dissonance when an Engineer is studying Medicine. At least forty years ago. An Engineer all too often keeps asking Why, what is the underlying set of basic scientific principles that explain the phenomenon and how can I express them in a manner in which they can be used on a predictable basis. Why would drive many a Medical Professor to apoplexy. Medicine was for a long while the transfer down of “facts” and not validatable principles. The old adage at graduation that fifty percent of what one had just learned in Medical School was now invalid was a bit of a joke but sadly it was also true.

But as we move to Genomics we sadly see this trait arise again. There is a tension between those who want to have basic repeatable principles to build upon and those who believe that collecting data is the sine qua non. Let me give an example of a recent experience. Prof Lander at MIT is teaching an EdX course on Biology. Now Lander is brilliant and his style of teaching is in many ways classic MIT. Namely he highlights the basic principles, and then the student works through the Problem Sets developing the details for themselves. So far so good. His first two three fourths of the course was fantastic. Then I noticed a subtle change, a change that, unless you were prepared to recognize would have slipped through the cracks. He slowly started giving a mixture or core predictable principles and cook book recipes. For example, we know that we can denature DNA because the base pair bonds are Hydrogen bonds, relatively weak, and the backbone Phosphate bonds are strong because they are ionic. Thus by heating the molecule we break the Hydrogen bonds first and then before we break the ionic bonds we can do our complementary additions, thus PCR works well.

On the other hand as he progressed to a discussion of Knock Out genes there were a collection of “tricks” or cook book recipes that were used. Why, for example did one get the modified DNA into the denatured gene the way he said? Well it just happens. Well nothing just happens. Fortunately bench Biologists have developed many “tricks”, like alchemists, and as a result they have become a bit too comfortable with this unexplained bevy of tools, albeit indispensable, but in the long run self-defeating.

As Dougherty states when he examines data mining as an example of the Biologist’s flair for data at all costs: 

Data mining and Copernicus share a lack of experimental design; however, in contradistinction to data mining, Copernicus thought about unplanned data and changed the world, the key word being ‘thought.’ Copernicus was not an algorithm numerically crunching data until some stopping point, very often with no adequate theory of convergence or accuracy. Copernicus had a mind and ideas. William Barrett writes, ‘The absence of an intelligent idea in the grasp of a problem cannot be redeemed by the elaborateness of the machinery one subsequently employs’. Or as M. L. Bittner and I have asked, ‘Does anyone really believe that data mining could produce the general theory of relativity’? Data mining represents a regression from the achievements of three and a half centuries of epistemological progress to a radical empiricism, in regard to which Reichenbach writes, ‘A mere report of relations observed in the past cannot be called knowledge. If knowledge is to reveal objective relations of physical objects, it must include reliable predictions. A radical empiricism, therefore, denies the possibility of knowledge’. A collection of measurements together with statements about the measurements is not scientific knowledge, unless those statements are tied to verifiable predictions concerning the phenomena to which the measurements pertain. 

What is Dougherty getting at? Simply, to reiterate the first quote: Science demands a marriage between data and models, to be true science it must be predictable and predictable based upon an embodiment in an abstraction.

Let me now apply this to genomics. Consider prostate cancer. The question is complex but can be asked; what is the first set of steps that lead to prostate cancer? Let us examine what we know:

First, we know many of the pathways. We know that the AKT pathway is critical, we know that c-MYC is a critical control element, we know that PTEN is often mutated, and we know that AR (Androgen Receptors) ultimately get mutated and we have metastatic growth. We pathways, we have relationships; we can demonstrate causality and results. Thus a modicum of a basis in reality exists. If one would use this pathway model and then search using microarrays matched against the model one arguable could iterate to improved models and improved predictability. The data without the model is useless and the model without the data is unverifiable.

Second, we can ask what sets the process off. Are all the changes due to mutations or more likely due to epigenetic insults? Thus when we look at MDS for example, we are looking at a hypermethylated set of blood stem cells. Something hypermethylated them and we know that since they are hypermethylated that the gene expression is repressed and thus cell proliferation of immature cells is a result. In prostate cancer, is the control mechanism lost because of a mutation, methylation, both, and in what order? Having a model allows one to validate and then iterate along a consistent trajectory of reality.

What does Dougherty have to say here? 

While ignorance of basic scientific method is a serious problem, it is necessary to probe further than simply methodological ignorance to get at the full depth of the educational problem. Science does not stand alone, disjoint from the rest of culture. Science takes place within the general human intellectual condition. Biology cannot be divorced from physics, nor can either be divorced from mathematics and philosophy. One’s total intellectual repertoire affects the direction of inquiry: the richer one’s knowledge, the more questions that can be asked. Schrodinger comments, ‘A selection has been made on which the present structure of science is built. That selection must have been influenced by circumstances that are other than purely scientific’  

The point I believe he is making is that in the new world of Genomics, it is necessary to have a foundation that exceeds just the Laboratory and its tricks. One must understand that no matter what we think that every time we look at a cell, at an organism, we are looking at a system, at some stochastic dynamical process wherein things move forward, albeit randomly, but in a way controlled by principles. We must look at the world wherein data is used not as an end in itself but as an iterative process with our mathematical world view. Thus the tools needed to view this world are extensive yet available. Engineers are trained to use them daily. Perhaps Genomics will grow to appreciate their essential import.

Saturday, November 30, 2013

Great Genomics Book

The book by Dale, Von Schantz and Platt, From Genes to Genomes, is almost perfect. It is a 350 or so page exceptionally well written book describing all the introductory materials one would need to become current with genomes and genomics efforts. As with many of the other books I had around I first looked at this and at a glance set it aside. Then came the moment when I wanted to re-understand something and I opened this book up and I was hooked. It in a clean and clear manner takes the reader from basic DNA principles and through all of the key techniques used in genomic studies today. It avoids getting to complex into any one area and it reads in a straightforward and consistent manner. It is a superb asset for “catching up” and I suspect for first learning the materials.

Chapter 1 is the basic introduction to genes and genomes. It is DNA 101 but it contains little tidbits of essential materials that are all well integrated. One thus starts with a clear understanding where the authors are taking the reader.

Chapter 2 is the material on basic gene cloning. It uses the plasmid approach with bacteriophage and does so without burdening the reader with too much overhead and history. This Chapter discusses technique and technology and the reader is given a logical approach to the basics of cloning. Restriction enzymes are introduced and the material is adequate to have enough depth to see how they can be applied. There are, of course, a lot of implementation questions that are left hanging but that is typical of this study. There is a section on ligation and I would like to have seen this carried over a bit when discussing gene knock-outs. We can understand how the genes ligate but the question is how well does this carry-over the later processes.

Chapter 3 discusses DNA libraries. This is a wonderful summary of the concept. The graphics supplement the text without over powering it. One example of what I call the cook book facts is demonstrated in p 95 when discussing hybridization. Here is the curve showing how as temperature increases the DNA starts to break apart. This is the denaturing of DNA, a concept again used with the PCR analysis. This is less a theoretical or structure issue but one of those cook-book facts that have been added to the tool chest of the Genome builder.

Chapter 4 is the PCR process. Simply it is the separating of DNA, then tagging one end and the other end and going through a temperature sensitive denaturing and rebuilding until what is left is millions of copies of a desired DNA segment. My only complaint here is that the graphic, good, albeit it could be made a bit better with color.

Chapter 5 discusses sequencing and it gives a superb discussion of the Sanger approach. Namely ddNTPs are used with segments and then measured in a gel electrophoresis. I assume that the reader may have had some understanding of the physical details but overall it is clear and exceptionally useful.

The text continues developing other elements used in current day genomics. Chapter 9 is an attempt to provide an overview of microarrays, SNP, and even GWAS and phylogenetics. My problem here is that they are trying to stretch it a bit too far. These are reasonable summaries but to do microarrays justice it may take a bit more detail, and yes color, and the phylogenetics is much too much just a high level summary.

Finally the Glossary is fantastic and worth every page.

The strengths and weakness of the book are simple. On the strength side it covers all the key issues superbly. On the negative side, and this may be perhaps me, I find that almost like Organic Chemistry, in Gene manipulation there still are many cookbook rules that are scattered between the facts and logical constructs. If somehow there could be a clarification of the cook book rule and the well understood logical steps that would be a help.

Overall I would highly recommend this book for almost anyone, from beginner to professional. My focus is clinical and theoretical modelling and analysis, and I have avoided bench work as much as possible. But by reading this book I can see again how much work has been done over the past few decades.

Monday, November 25, 2013

Personal Genome

There has been an explosive growth in people wanting to get their genome analyzed. The FDA last Friday issued what in my opinion appears to be almost a cease and desist order, or as they phrase it:

must immediately discontinue marketing the PGS until such time as it receives FDA marketing authorization for the device.

 Specifically the FDA claims:

This product is a device within the meaning of section 201(h) of the FD&C Act, 21 U.S.C. 321(h), because it is intended for use in the diagnosis of disease or other conditions or in the cure, mitigation, treatment, or prevention of disease, or is intended to affect the structure or function of the body. For example, your company’s website at ... markets the PGS for providing “health reports on 254 diseases and conditions,” including categories such as “carrier status,” “health risks,” and “drug response,” and specifically as a “first step in prevention” that enables users to “take steps toward mitigating serious diseases” such as diabetes, coronary heart disease, and breast cancer. Most of the intended uses for PGS listed on your website, a list that has grown over time, are medical device uses under section 201(h) of the FD&C Act. Most of these uses have not been classified and thus require premarket approval or de novo classification, as FDA has explained to you on numerous occasions.

 There always has been a set of concerns here. Most people have no clue what some of these readings mean. Also the reading may be at best reflective a a still questionable result. One should ask even if most physicians have the ability to ascertain the results.

There is also the question of what rights have you lost by sending your DNA to some third party.

This has always been an issue and the FDA may very well have taken a proper and necessary step. It will be interesting to follow this.

Thursday, October 3, 2013

Designer Children

Apparently the PTO has considered issuing a Patent to a US company for the ability to select male and female gametes to "design" a child with certain features. From a Nature Genetics in Medicine  article the authors state:

...the US Patent and Trademark Office in June 2013, and it will issue as US Patent No. 8543339 on
24 September 2013. It contains claims to a computer system and to a computer program, but our focus here is on the patent’s claims to a method for gamete donor selection: 


“A method for gamete donor selection,1 comprising (i) receiving a specification including a phenotype of interest that can  b be present in a hypothetical offspring; (ii) receiving a genotype of a recipient and a plurality of genotypes of a respective plurality of donors; (iii) using one or more computer processors coupled to one or more memories configured to provide one or more computer processors with instructions to determine statistical information including probabilities of observing the phenotype of interest resulting from different combinations of the genotype of the recipient and genotypes of the plurality of donors; and (iv) identifying a preferred donor among the plurality of donors, based at least in part on the statistical information determined, including comparing the probabilities of observing the phenotype of interest resulting from different combinations of the genotype of the recipient and the genotypes of the plurality of donors to identify the preferred donor.”

It appears as if the company believes that by selecting male and female gametes based upon a full genetic determination of the parentage that one can determine the characteristics of the offspring. Surprise. Welcome to epigenetics. Things do not always work out that way.

Now perhaps the authors of the piece doth protest a bit too much. In fact one of them apparently got into a comment fight on a PLOS Blog of a highly respected Genetics Blogger. Thus this topic seems to have raised the ire of many folks.

The problem is that the epigenetic elements associated with combining male and female gametes is quite complex. Thus just choosing them assuming the genes are fine is at best an interesting first start but no guarantee.

This will be an interesting battle to watch.