Showing posts with label science-book-reviews. Show all posts
Showing posts with label science-book-reviews. Show all posts

May 10, 2017

Embryology Special Issue

Me and my colleagues are pleased to announce an upcoming special issue of the journal Biology (Basel). The topic is "Computational, Theoretical, and Experimental Approaches to Embryogenesis" (see announcement). Our view of what constitutes embryogenesis research is rather broad, spanning experimental studies, cellular reprogramming, bioinformatics, and artficial life. Therefore, we seek submissions from a wide variety of researchers and article types.


As the lead editor, I will take any questions you might have about interesting ideas, types of articles, or if you are interested in peer-review. As noted on the poster, the deadline for submissions is August 31, 2017. Looking forward to an excellent issue.

UPDATED (5/17):
With the initial dealine fast approaching, we have decided to extend the submission deadline to December 31. 

June 8, 2016

200K, 1K (or less) At A Time

While I have not been keeping up with my blogging habit over the last 18 months or so, Synthetic Daisies is still reaching milestones. As of today, we have reached 200,000 reads! While this took 7 years and 6 months (as of June 15), it is quite a milestone. A few historical points.

Frequency of posts over time

The bulk of posts (particularly the longer posts) were written during a period from late 2011 to early 2015. Some of them have had longer lifetimes than others, as you will see below.

Logos over time
2008-2012. Classic version

2012-2015. New design, pretentious slogan.

2016-present. Cleaned up new design.


Here is a reading list with some of the lesser-read but perhaps most interesting posts in the blog's history.

Network Science:
Six Degrees of the Alpha Male: breeding networks to understand population structure. August 22, 2014.

Fireside Science: Inspired by a visit to the Network's Frontier.... December 16, 2013.


Meta-science:
Scientific Paradigm Network. February 8, 2015.

Academic Connectivity and the Future of Scientific Ideas. September 9, 2011.

Fireside Science: The Representation of Representations. June 21, 2014.


Book Reviews:
Review of "Arrival of the Fittest". March 9, 2015.

Metabiology and the Evolutionary Proof. January 11, 2013.

Review of "Intelligent Movement Machine". April 19, 2009.


Cognition, Biology, Technology, and Innovation:
Merging electronics and biology: the future of touch. November 1, 2012.

The "nature" of materials: evolution and biomimetics. December 26, 2011.

I, Automaton. September 16, 2013.


Evolution, Alife, and Complexity:
Artificial Life meets Geodynamics (EvoGeo). November 21, 2012.

Reflections on Chaos in Biological Evolution. May 25, 2013.

The Neuromechanics and Evolution of Very Slow Movements. April 18, 2012.


Systems Biology:
Modeling Processes with No Beginning, an Adaptive Middle, and No End. October 27, 2013.

Facilitated Variation (FV): a random (walk) tour. October 29, 2011.

"Reining" in Diabetes. January 10, 2011.


Game Theory and Complexity:
Games, Noise, and Science-related Obscure References. April 8, 2013.

Makin' Pha-ses. March 11, 2013.


Although the blog got off to a slow start, I learned a lot about "how to blog" (use interactive media to greater effect) over the course of time. Nevertheless, hooray for 200K!

March 9, 2015

Review of "Arrival of the Fittest"

"Arrival of the Fittest" is the latest book by Andreas Wagner, a professor at the University of Zurich in Switzerland. The book [1] tackles a subtly complicated topic: the evolution and evolvability of innovations using a biochemical and computational perspective. For the most part, Wagner succeeds at presenting an elegant case for how the ability to naturally evolve innovations lies at the heart of the evolutionary process. To get there, however, Wagner must introduce us to a number of semi-obscure concepts (at least to the layman). The book can be summarized in four parts: survival of the most novel (I), the concept of innovability (II), the safe and the risky (III), and multiple origins, multiple solutions (IV). I will give a technical review of the book by highlighting these four interrelated themes.



I. Survival of the most novel.
To understand the propagation of innovative solutions throughout the tree of life, it is important to understand the difference between the force of evolution versus the role of selection. Wagner proposes that, rather than creating innovations, the role of natural selection is to preserve them. Innovations themselves are the result mutation, recombination and historical contingencies. The first two mechanisms are capable of randomly generating either very simple innovations or the components of more complex innovations. But to achieve the "tinkering" that seems to be prevalent in complex genetic pathways and phenotypes, we need to have standard meta-components that can lock in previous changes. This allows for the limited exploration of a fitness space without incurring the cost of losing previous advances.


Wagner teases historical contingencies this into two classes: building blocks and standards. In the case of building blocks, the working components that result from previous innovation are modularized into larger units. This allows for increasing complexity to be had from a stochastic process (evolution) as well as accelerating the process of finding and retaining novelties (innovation). The coordination of building blocks often leads to standards, which are much more flexible than hyper-specialized systems but are much more able to deal with hyper-complexity.

Building blocks assembled into a complex structure.

Wagner illustrates this by comparing metabolic engines (a high-tolerance biological engine) with the internal combustion engine (a high-precision mechanical innovation). While metabolic engines can use an interchangeable set of fuels and reactants, internal combustion engines require specific specifications to operate. While one might argue that this implies mechanical engines will be "optimized" and metabolic engines will be made "good enough", this is indeed the point. Rather than survival of the fittest, we observe a survival of the fit enough, with selection acting strongest on functional novelties that augment fitness.

A high-tolerance but highly-specific engine (internal combustion) type.

II. Concept of innovability.
In the case of complex metabolic networks, the basic function has not changed throughout evolutionary history. What has changed is the number of reactions, which scales with evolutionary complexity. This scaling requires very general standards. But to make these standards interoperable across divergent evolution, a diversity of building blocks regulatory mechanisms are also required. This leads us to a set of principles which can explain general trends in innovation rather than on a case-by-case basis. Another such principle involves the number of parts and identity of a specific innovation. The number of interacting parts and their configuration provides a means to compare specific innovations in a hypothetical manner. Wagner proposes that this be done using a high-dimensional structure such as a hypercube [2] or a neutral network [3].

In a hypercube representation, each node represents a specific genotype, while the edges represent pathways (or the accumulation of mutations) between genotypes. In short, the shortest paths are the most probable. In a highly-connected space, a long distance can be traveled across the space in a short number of mutations (or edges).

III. The safe and the risky.
How can the structure of a biological system make things safe for innovation? Certainly, blindly changing key components of a metabolic network or developmental scaffolding without regard for essential function can result in lethality. The presence of the building blocks and standards principles provides a failsafe means to tinker or change without lethal disruption. But there are two other principles at work: redundancy and connectivity. These principles are arguably more important in acting as gatekeepers for the fitness benefits reaped by innovations.

A machine finding its way through a maze. One way to search for novel solutions in a de novo fashion.

Redundancy involves the presence of multiple parts that play a similar or interchangeable role in the functioning of a system. For example, both genetic and metabolic networks can be robust to the removal of single components. When components get removed without having a consequence on function [4], we can say that such a component is redundant. Gene duplications can play a similar role: single copies can be knocked out without a large detriment to fitness. Related to redundant function is, of course, robustness. Robustness can be thought of as how much redundancy exists in a specific system. Wagner gives the example of phenocopying as a way in which developmental robustness can lead to redundant phenotypic configurations.

Connectivity is a less appreciated aspect of evolution. Yet in terms of functional and conceptual unity, connectivity is an essential component of evolving systems that produce innovation. Three aspects of connectivity are most important here: interaction type, density, and connection order. There are a number of interaction types in biological networks that can give rise to innovation. For example, we can look to genetic networks (interaction between genes), breeding networks (interactions between conspecifics), or even the aforementioned neutral networks (interactions between evolutionary configurations) for ways in which connectivity can yield pathways that lead to innovations of high fitness.

An example of a genetic (gene interaction) network. Sometimes this approach is called "hairball science". COURTESY: Figure 1 in [5].

Network density (or the number of interconnections between nodes) allows for a greater number of potential innovative solutions to be explored in a shorter number of steps. This can increase the number of potential solutions, which can be reached in a shorter number of evolutionary steps and/or time (depending on your perspective). In the neutral network, a number of "safe" pathways of equivalent fitness are created for the innovating pathway or organism to explore. While the network density is important in facilitating innovation, the connection order of the network is also important. In networks with a short connection order (e.g. small-world networks), even random paths through the network can yield large-scale, non-deleterious change.

An example of constituent genotypes in a neutral network with reference to robustness. COURTESY: Ricardo Azevedo, Wikipedia.

IV. Multiple origins, multiple solutions.
The final point of Wagner's book is to remind us that innovations can be achieved through multiple unique solutions, and that existing innovations may have had multiple origin points. While evolution also has no unique solution, innovation is particularly subject to parallel evolution. Since this book does not shy away from computational representations, Wagner applies the idea of evolution strategies [6] to illustrate how innovations can be modeled as a five-part process. Specifically, innovation involves trial-and-error, population-based exploration, solutions with multiple origins, a combinatorial structure, and a stochastic process with a mutation-selection structure. While these factors also define the evolutionary process, we might say that innovation is inseparable from evolution by natural selection. Thinking more broadly and considering social systems, cultural innovation might also be best characterized as an evolutionary process. Wagner's approach to innovation as evolvability is a useful and accessible introduction to the topic, and places this new set of concepts and terminology directly into the context of evolution by natural selection.

NOTES:
[1] For other, briefer reviews, please see: Hoppe, R.B.   Andreas Wagner: Arrival of the Fittest: Solving Evolution’s Greatest Puzzle. Panda's Thumb blog, November 4 (2014) AND Pagel, M.   The Neighborly Nature of Evolution. Nature, 514, 34 (2014).

[2] For an example, please see: Gavrilets, S. and Gravner, J.   Percolation on the Fitness Hypercube and the Evolution of Reproductive Isolation. Journal of Theoretical Biology, 184(1), 51–64 (1997).

[3] For an example, please see: Wagner, A.   Robustness and Evolvability in Living Systems. Princeton University Press (2005).

[4] Li, J., Yuan, Z., and Zhang, Z.   The Cellular Robustness by Genetic Redundancy in Budding Yeast. PLoS Genetics, 6(11), e1001187 (2010).

[5] Magtanong, L. et.al   Dosage suppression genetic interaction networks enhance functional wiring diagrams of the cell. Nature Biotechnology, 29, 505-511 (2011).

[6] Evolution strategies assumes that evolution is an optimizing process that results from the iterative application of a mutation-selection model. For more, please see: Schwefel, H-P.   Numerical optimazaion of computer models. Wiley Press, Chichester (1981) AND Beyer, H-G. and Schwefel, H-P.   Evolution Strategies: A Comprehensive Introduction. Journal of Natural Computing, 1(1), 3–52 (2002).

August 1, 2013

Universal Patterns and Origins of Innovation

Here are two recent posts on innovation originally featured on my micro-blog, Tumbld Thoughts. Each post reviews a comtemporary book on the patterns inherent in the innovation process. The first (I) features several different archetypes, while the second (II) features the kinds of environments that are key for maximizing innovation.

I. Universal Patterns of Innovation


Interesting book I ran across recently on the universal "patterns" that seem to underlie innovation and invention [1]. While the book is about much more than this, one core theme is the practice of discovery and what we might learn from looking at the practices of different inventors. 


One way to take advantage of these patterns is to learn the rules of innovation. These rules are defined as the underlying talent, knowledge, and allocation of resources neccessary for innovation. Another lesson learned is to recognize there are at least three principles (better understood as personal styles) that define great inventions. These are:

1. Serendipity, or being able to exploit chance discoveries. William Shockley's work with semiconductors (leading to the transistor) best exemplifies this principle.

2. Proof-of-principle, or the 99% perspiration, 1% inspiration approach. Thomas Edison's work on the incandescent lightbulb best exemplifies this principle.

3. Inspired Exertion, or the greater than 1% inspiration approach. Jeff Hawkins' work in developing the Palm mobile computer best exemplifies this principle.

The third lesson that leads us to innovation is to study the designs of great innovators. See these Synthetic Daisies post from 2009 and 2011 on my own (evolving) thoughts on this topic.


II. The Origins of Innovation


Here is a link to a sped-up whiteboard animation video featuring content from Steven Johnson's book "Where Good Ideas Come From: the natural history of innovation" [3]. His main thesis is that innovation tends to occur in connected spaces such as cities, reefs, and webs [4]. As mentioned at the end of the video: "chance favors the connected mind".



Innovation also occurs as a process. One of these processes is called the slow hunch. The example Johnson gives for this is Tim Berners Lee and invention of the internet. At first, the proto-internet was conceived as a way to organize personal data. The next stage involved extending the connectivity aspect to interpersonal tangles. Finally, a version of the internet we all recognize came to fruition as a dynamic set of interconnected documents and links. This process of successive iteration took years to achieve.


NOTES:

[1] Alesso, H.P., Smith, C., and Burke, J.   Connections: Patterns of Discovery. Wiley/IEEE Press (2008).

[2]  Alicea, B.   Innovation Class/Book. Synthetic Daisies blog, June 30 (2009) AND Alicea, B.   In praise of repetition? Synthetic Daisies blog, April 11 (2011).

[3] Also see his TED talk on the book. For more sped-up whiteboard animations on innovation and the process of invention, please see: Alicea, B.   New Directions in Making Innovation Pay. Synthetic Daisies blog, June 1 (2013).

[4] The city is a literal city (particularly the mixing that occurs on city streets), the reef is a space that metaphorically resembles a coral reef (diverse individuals visit to feed and mingle), and the web is a network (made explicit in the internet).

July 13, 2013

Maps, Models, and Concepts, July edition

Here are some maps, models, and concepts reposted from my micro-blog, Tumbld Thoughts. This is a mash-up of recent books and articles in my reading queue, plus recent features from around the web. The order is: I (Nearly-decomposable Systems), II (Rube Goldberg Mechanisms), III (a profile of Elon Musk's Hyperloop), and IV (Maps + Metadata).


I. Nearly-decomposable Systems

This concept was originally proposed in Chapter 4 of "Sciences of the Artificial" by Herbert A. Simon (Chapter 4 is entitled "Architecture of Complexity"). The focus is on a concept called "nearly-decomposable systems".

In the modern practice of algorithmic representation, decomposability enables one to represent a complex system as a strict hierarchy. By contrast, nearly-decomposable systems can be found in systems where short-run behavior is statistically independent but long-run behavior is dependent in an aggregate fashion. While discrete states appear to exist at static intervals, examining the dynamics reveal interactions (or overlap) between these states. 

One example of this can be seen in the above image, which is adapted from Figure 7. A system is partitioned into 8 spatially non-overlapping components (A1, A2, A3, B2, B2, C1, C2, C3). A sparse matrix (top left) can then be constructed to model the selective functional interactivity between these components (discrete states). In this case, short-run behavior is restricted to interactions within each state, while long-run behavior characterizes the interactions between states. Overall, intra-component linkages (e.g. interactions within A1) are greater than inter-component linkages (e.g. interactions between A1 and B2). 

In Chapter 4, Simon applies this concept to the behavior of diffusing particles in physiochemical systems. However, in systems with autonomous intelligence (e.g. social systems), agents (the equivalent of diffusing particles) can influence and communicate with each other. This concept can also be applied to hierarchical systems (e.g. social and biological complexity). In such cases, the distinction between "broad" vs. "narrow" hierarchies (e.g. hierarchical span) becomes important. 

For a related concept as applied to systems biology, please see: 

Alicea, B.   The Curse of Orthogonality. Synthetic Daisies blog, October 3 (2011).

Further Reading:

Agre, P.E.   Hierarchy and History in Simon's "Architecture of Complexity". Journal of the Learning Sciences, 12(3) 2003.

Bentley, J.L. and Saxe, J.B.   Decomposable searching problems I. Static-to-dynamic transformation. Journal of Algorithms, 1(4), 301-358 (1980).

Feigenbaum, E.A.   Retrospective: Herbert A. Simon, 1916-2001. Science, 291(5511), 2107 (2001).

Simon, H.A.   Near-decomposability and the speed of evolution. Industrial and Corporate Change, 11(3), 587-599.

Simon, H.A.   Sciences of the Artificial. MIT Press, Cambridge, MA (1969).


II. Rube Goldberg (e.g. convoluted, non-optimal) Mechanisms

Happy (posthumous) Birthday (July 4th) to Rube Goldberg, the father of the Rube Goldberg machine [1]. Rube Goldberg machines provide a convoluted way to accomplish something that is otherwise simple. For example, to get sand out of a pair of shoes, one could take their shoes off, turn them upside down, and tap them.

In the Rube Goldberg universe, however, you would have to build a complex machine with many degrees of freedom to accomplish the same feat. His creations were massively inefficient, and that's the whole point -- if your worldview is one of parsimony, you will find his comics humorously absurd.

Given that his birthday falls on July 4 (US Independence Day), the associated Google Doodle (from 2010) features a seven-step machine that lights a firecracker. But can Rube Goldberg machines be useful? For more on this, check out a few Synthetic Daisies blog posts [2] on the application of Rube Goldberg-like machines to systems biology and evolution.

In this case, we are evaluating viable function in the context of maximal convolution -- in other case, the more steps to accomplish a task, the better. Perhaps this [3] is a biologically-plausible alternative to the view of evolution as parsimony.


III. Conceptual Porn for Technology Visionaries

Tech visionaries unite! I have run across a critical mass of articles on Elon Musk's proposal to build a high-speed transportation system called the "Hyperloop" [4]. Musk has described it as "a cross between the Concorde, a railgun, and an air hockey table". Supposedly better than high-speed rail, airplanes, or electric vehicles.

The hyperloop concept seems to build off of a number of existing technologies [5], some more developed for commercial use than others. It is not a vacuum tube nor a conventional rail system, although it incorporates both of these design elements. A version of what Musk is envisioning is similar to the Evacuated Tube Transport system, patented by Daryl Oster (founder of ET3 technologies) [6].

How does it work and is it simply hype? See these popular news features from Gizmag.com, Autoblog
Green, and The Atlantic Wire for more information. Is this fringe science? Read the book "Physics on the Fringe" [7] and decide for yourself.


IV. Maps + Metadata Tell Interesting Tales

The first set of layered maps are from the NASA/NOAA Green project. The goal of this joint project is to make a remotely-sensed map of the earth's vegetation (land mass vegetation only) using data from the Suomi NPP satellite.

In this YouTube video, it is explained that land cover and vegetation changes can have an effect on weather variability. Viewing these processes over the course of a year (animated in the video) can help us understand interactions between the biosphere and atmosphere. 

Examples of this are shown above. The pictures (from top) represent: the red deserts of Australia (shown as white expanses), snowcover in North America, the grasslands in the Florida everglades, and deforestation in East Africa.


The second set of layered maps are brought to us in the form of a really cool video animation from Cube Cities and Google Earth. Specifically, this is a time lapse of Chicago's skyline and its growth from 1865 (perspective view) to 2014 (birds-eye view). To illustrate the differences, I have screen-captured selected years and placed them in series. The buildings are 3-D models superimposed on a 2-D map of the city. Quite impressive.

NOTES: 

[1] Leibach, J.   Rube Goldberg Mashup. Science Friday blog, July 4 (2013).

[2] Alicea, B.   Machinery of Biocomplexity, new arXiv paper. Synthetic Daisies blog, April 19 (2011) AND Non-razors, unite! Synthetic Daisies blog, January 30 (2009)

[3] Bottom figure (minimal biological model) is from: Alicea, B.   The "Machinery" of Biocomplexity: understanding non-optimal architectures in biological systems. arXiv: 1104.3559 [q-bio.QM] (2011).

[4] Basulto, D.   Is the Hyperloop the Future of Transportation? BigLoop blog, June 12 (2013).

[5] For more information on these technologies, please see the following list:




d) VHST technology. For more information, please see: Salter, R.M. The Very High-speed Transit (VHST) System. Rand Corporation (1972).

See also the following patent document: Oster, D.   Evacuated tube transport. US5950543. USPTO (1999).

January 11, 2013

Metabiology and the Evolutionary Proof

I have recently read the book "Proving Darwin" by Gregory Chaitin. Greg is a mathematician who is best known for his work in computational complexity [1]. "Proving Darwin" is not only a Mathematician's take on evolution by natural selection, but also a rebuttal of intelligent design. Despite its slimness [2], the book is quite interesting. Chaitin is particularly interested in the creative and algorithmic complexity aspects of natural selection. While this would seem a bit obscure to a biologist, this perspective provides crucial evidence for the viability of natural selection, and even parallels some empirical results from artificial life and evolutionary biology. Aside from providing a review of this book, I will also provide a conceptual evaluation of of the core thesis by using simulated data.


To orient the biological audience, a Mathematical proof is a formal logical construct one uses to demonstrate the plausibility of a given solution. In the case of an unsolved mathematical problem (e.g. Riemann, P vs. NP), the proof is what is required to claim that a problem is solved [3]. In this case, Chaitin does not lay out his proof directly, but does draw his conclusions from it. This is an approach he calls "Metabiology" [4]. The metabiological approach characterizes living systems as programs. Therefore, the primary difference between living and non-living entities in metabiology is the ability to transmit information. This is not a new observation, but is key to his characterization of evolutionary systems. A flame is used to highlight this difference: while a flame can have metabolism and can self-reproduce, it cannot evolve like a living entity because it cannot transmit information [5]. While this may not be a completely reasonable analogy, it does allow for us to further understand the nature of creativity in living systems.

The algorithm of evolution proposed in the book is a k-bit program which can be modified via mutation [6]. This is similar to in silico approaches to evolutionary biology such as Tierra and Avida. The relative information content (or program size complexity) of a mutation in the k-bit program is the information content of a non-mutated program (e.g. genotype) given the information content of a mutant program [7]. The k-bit program (Figure 1), as a highly stylized and abstract population of genomes, of results in a partially-stochastic dynamical system, the relevance of which will soon become clear.

Figure 1. Schematic of a k-bit program and how it evolves

In the context of evolution, creativity is the ability to create new combinations -- and ultimately forms. And according to Metabiology, the most creative lineages and forms that contain the most information-rich mutations. But what does it mean to possess "informative" mutations? To better understand this, Chaitin provides us with three evolutionary regimes to test his mathematical theory. These are: 1) intelligent design, 2) brainless exhaustive search, and 3) cumulative random evolution. Intelligent design is provided as a representative of evolution being a deterministic process that is limited to fixed number of outcomes. By contrast, brainless exhaustive search assumes that evolution is entirely random, with no constraints. Furthermore, this model has no memory, and does not allow for evolutionary conservation or modularity. Cumulative random evolution provides a middle path between these two scenarios: random mutation given a
limited fitness landscape, resulting in evolutionary trajectories that are highly path-dependent [8].

By solely considering the Big O metric of program complexity [9], it is shown that the intelligent design scenario achieves a maximum fitness value much faster than either brainless exhaustive search or cumulative random evolution (Figure 2). Why is this and how does such a result prove natural selection? Evolutionary algorithms tend to be much slower than comparable techniques in contexts such as multivariate optimization [10], even for a convex (e.g. where the goal is clearly defined) search space. However, the speed of attaining a maximum fitness value does not translate into a robust result. This is because in Metabiology, fitness [11] is defined as the rate of biological creativity rather than as being related to differential reproduction. The speed (or tempo) of evolution is measured by how fast fitness grows [12]. While this does not provide a specific test of selective forces, it does move us away from the naive adage of "survival of the fittest". Instead, we find lineages with varying amounts of creativity, and a situation akin to "survival of the most creative".



Figure 2. How quickly maximum fitness is achieved for three evolutionary scenarios. Inset: the earliest portion of evolution (indicated by bracket). Hypothetical scenario modeled using simulated (pseudo) data.

The maximum fitness for the intelligent design scenario (as shown in Figure 3) occurs very early in the evolutionary trajectory. Essentially, it shows a big initial gain for intelligent design. Presumably, this strategy finds the known strategies first. The other two strategies reach their maximal fitness later on. However, it is the secondary properties, not simply speed of solution, which actually make for an adaptive outcome. Using these secondary properties as a guide, we can assess exactly why cumulative random evolution is more powerful than the alternatives. As characterized in the book, there are four additional ways to evaluate these models: creativity, time, robustness, and coherence of function. Figure 4 shows continua for each of these.

Figure 3. Table that shows the degree of creativity, time taken, robustness exhibited, and systemic coherence exhibited by each evolutionary scenario. Hypothetical scenario modeled using simulated (pseudo) data.

Finally, let us consider how these results square away with findings from the artificial life literature. When I first started to read about Metabiology, I was struck by how parallel this approach is to the Avida platform. Both approaches use programs (in the form of instruction sets) as a stand-in for the genome. In particular, the tendency for fitness (and complexity) to grow over evolutionary time. This either suggests that there is a fundamental insight to be had here, or that systems like these are very good at maximizing yield.


Figure 4. Table that shows the degree of creativity, time taken, robustness exhibited, and systemic coherence exhibited by each evolutionary scenario. Values along continuum are arbitrary. 1 = Intelligent Design, 2 = Brainless Exhaustive Search, 3 = Cumulative Random Search.

Instruction-based evolutionary models are derived from Turing machines [12], for which the machine's state is indeterminate. This is based on the halting mechanism, which does not allow for guided behavior in any way. Yet evolution must be guided in some fashion, and intelligent design provides brittle results at best. This is where Chaitin introduces readers the concept of an Oracle [13], which can be thought of as a top-down guidance mechanism for the halting (or mutational) behavior of a random system. Think of this as a natural version of site-directed mutagenesis, so that lethal combinations would be much less likely than otherwise. This natural (albeit hypothetical) oracle could also resolve some of the evolutionary paradoxes we often witness in natural populations. For example, an oracle might prevent populations from getting stuck in low-fitness regimes, or make large portions of a rugged fitness landscape more accessible [14]. In metabiological explorations, throwing out lethal mutants maintains high levels of fitness and allows for evolutionary dynamics such as competition and sustained arms races.

A naturalistic mechanism for the proposed oracle is still very much a mystery. However, the architecture of execution routines and the structure of genetic regulatory networks (GRNs) is very similar. From a programming perspective, a hierarchical structure could serve as a heuristic stand-in for this oracle. In particular, an approach called subrecursion (e.g the use of subrecursive hierarchies -- [15]) might provide a mechanism for the supervised maximization of fitness. This is the direction the end of the book should have taken -- however, this might also be work in progress. A more general mechanism may lie in the mysteries of mathematical incompleteness and randomness itself. Mathematical incompleteness (discussed at length in this book in the form of Chaitin's Omega number) suggests that all consistent statements must include undecidable propositions which neither be proven nor disproven. This might help us understand why although the forces of evolution are well characterized, the explanatory power of evolutionary laws is limited [16].

So what are the theorems of evolution? Chaitin does not make them explicit in this book, which is a letdown. And a causal reader gets the sense that Metabiology is simply a reformulation of results already confirmed in the areas of population genetics and digital biology. However, it is interesting that Chaitin has come at these results using a parallel methodology. Perhaps this convergence is the best evidence for evolution by natural selection.

NOTES:

[1] The field of Algorithmic Information Theory, to be exact.

[2] Contrast this with the thickness of "The Structure of Evolutionary Theory" by Steven J. Gould.

[3] As opposed to a test of the null hypothesis (NHST) or transitional fossil.

[4] This is based on a course at UFRJ and a lecture at the Santa Fe Institute (which I cannot locate on YouTube). For more information, please view this blog post at Logical Atomist.

[5] Another difference between a flame and a living population is the scale of its replicators. A flame, while hard to characterize as a population, behaves more like a slime mold than an intermittently-coordinated population (e.g. flock of birds or herd of sheep). This is an interesting point which is beyond the scope of the book.

[6] The probability of mutation (or probability of M) scales asymptotically to program length 2-k.

[7] In other words, the lineage resembles a probabilistic graphical model (PGM - e.g. Bayesian network). The more complex lineages are simply those with more informative mutations.

For a tutorial on PGMs, please see the following reference: Airoldi, E.M.  Getting Started in Probabilistic Graphical Models. PLoS Computational Biology, 3(12), e252 (2007).

[8] For simplicity's sake we can assume that all programs maximize their fitness on a fitness landscape. The actual topology is not important in this example.

[9] A Beginner's Guide to Big O Notation. Rob Bell blog, June 2009. See also Big O Notation, Wikipedia page.

[10] For the speed of evolutionary algorithms in the context of multiobjective optimization, please see: Coello-Coello, C.A., Lamont, G.B., and Van Veldhuizen, D.A.  Evolutionary Algorithms for Solving Multi-objective Problems. Springer, Berlin (2007).

[11] Fitness is characterized using the Busy Beaver (BB) function. In metabiology, BB(N) is the maximum fitness function attainable given the template program, suite of mutations, and strategy of evolution pursued.

For comparative purposes, another person doing work on the role of creativity in evolution (in silico) is Joel Lehman at UT-Austin.

[12] For more on the speed (or rate) of evolution through the accumulation of mutations in E. coli (e.g. wet-lab experimentation), please see: Kryazhimskiy, S., Tkacik, G., and Plotkin, J.B.  The dynamics of adaptation on correlated fitness landscapes. PNAS, 106(44), 18638-18643 (2009) AND Desai, M.M., Fisher, D.S., and Murray, A.W.   The Speed of Evolution and Maintenance of Variation in Asexual Populations. Current Biology, 17(5), 385-394 (2007).

For an Avidian study on the rate of evolution in the context of fitness landscape ruggedness (e.g. in silico experimentation), please see: Clune, J., Misevic, D., Ofria, C., Lenski, R.E., Elena, S.F., Sanjuan, R.  Natural Selection Fails to Optimize Mutation Rates for Long-Term Adaptation on Rugged Fitness Landscapes. PLoS Computational Biology, 4(9), e100087.

[13] In thermodynamics, such an oracle is referred to as Maxwell's Demon. While this demon is rendered hypothetically impossible, a computational oracle is less bound by the laws of physics. For a primer on Maxwell's Demon, please see this Java app (shown below).


[14] Franke, J., Klozer, A., de Visser, J.A.G.M., and Krug, J.  Evolutionary Accessibility of Mutational Pathways. PLoS Computational Biology, 7(8), e1002134 (2011).

[15] For more information, please see:  Calude, C.  Theories of Computational Complexity. Elsevier, Amsterdam (1988) AND Calude, C.  Incompleteness, Complexity, Randomness, and Beyond. Minds and Machines, 12(4), 503-517 (2002).

[16] For those who are unfamiliar, the best-known evolutionary law is allometry, or the scaling of size and growth across phylogeny.

December 5, 2012

Triangulating Scientific “Truths”: an ignorant perspective

I have recently read the book "Ignorance: how it drives science" by Stuart Firestein, a Neuroscientist at NYU. The title "Ignorance" refers to the ubiquity of how what we don't know influences the scientific facts we teach, reference, and hold up as popular examples. In fact, he teaches a course at New York University (NYU) on how "Ignorance" can guide research, with guest lecturers from a variety of fields [1]. This is of particular interest to me, as I hosted a workshop last summer (at Artificial Life 13) called "Hard-to-Define Events" (HTDE). From what I have seen in the past year or so [2], people seem to be converging on this idea.


In my opinion, two trends are converging that seem to be generating interest in this topic. One is the rise of big data and the internet, which make the communication of results and rendering of the "research landscape" easier. Literature mining tools [3] are enabling discovery in and of itself, but also revealing the shortcomings of previously-published studies. There has also been a good deal of controversy raised over the last 10 years in terms of a replication crisis [4] coupled with the realization that the scientific method is not as rigorous [5] as previously thought.

The work of Jeff Hawkins, head of Numenta, Inc. [6], addresses many of these issues from the perspective of formulating a theoretical synthesis. For a number of years, he has been interested in a unified theory of the human brain. While there are challenges in terms of both testing such a theory AND getting the field to fit it into their conceptual schema, Jeff has nevertheless found success building technological artifacts based on these ideas.


Jeff's work illustrates the balance between integrating what we do know about a scientific field with what we don't know. This involves using novel neural network models to generate "intelligent" and predictive behavior. Computational abstraction is a useful tool in this regard, but in the case of empirical science the challenge is to include what we do know and exclude what we don't from our models.

According to this viewpoint, the success of scientific prediction (e.g. the extent to which a theory is useful) is dependent upon whether findings and deductions are convergent or divergent. By convergent and divergent, I mean how independent findings can be used to triangulate a single set of principles or predict similar outcomes. Examples of convergent findings include Darwin’s finch beak diversity to understand natural selection [7] and the use of behavioral and neuroimaging assays [8] to understand attention. 


There are two ways in which the author proposes that ignorance operates in the domain of science. For one, ignorance defines our knowledge. The more we discover, the more we discover we don't know. This is often the case with problems and in fields for which little a priori knowledge exists. The field of neuroscience has certainly has its share of landmark discoveries that ultimately raise more questions than provide answers in terms of function or mechanism. Many sciences often go through a "stamp-collecting" or "natural history" phase, during which characterization is the primary goal [9]. Only later does hypothesis-driven, predictive science even seem appropriate.

The second role of ignorance is a caveat, based on the first part of the word: "to ignore". In this sense, scientific models can be viewed as tools of conformity. There is a tendency to ignore what does not fit the model, treating these data as outliers or noise. You can think about this as a challenge to the traditional use of curve-fitting and normalization models [10], both of which are biased towards treating normalcy as a statistical signal signature.

If we think about this algorithmically [11], it requires a constantly growing problem space, but in a manner typically associated with reflexivity. What would an algorithm defining "what we don't know" and “reflexive” science look like? Perhaps this can be better understood using a metaphor of the Starship Enterprise embedded in a spacetime topology. Sometimes, the Enterprise must venture into uncharted regions of space (but one that still corresponds to spacetime). While the newly-discovered features are embedded in the existing metric, these features are unknown a priori [12]. Now consider features that exist beyond the spacetime framework (beyond the edge of the known universe) [13]. How does a faux spacetime get extrapolated to features found here? The word extrapolation is key, since the features will not necessarily be classified in a fundamentally new way (e.g. prior experience will dictate what the extended space will look like).



With this in mind, there are several points that occurred to me as I was reading “Ignorance” that might serve as heuristics for doing exploratory and early-stage science:

1) Instead of focusing on convexity (optimal points), examine the trajectory:

* problem spaces which are less well-known have a higher degree of nonconvexity, and have a moving global optimum.

* this allows us to derive trends in problem space instead of merely finding isolated solutions, especially for an ill-defined problem. It also prevents an answer marooned in solution space.

2) Instead of getting the correct answer, focus on defining the correct questions:

* according to Stuart Firestein, David Hilbert's (Mathematician) approach was to predict best questions rather than best answers (e.g. what a futurist would do).

* a new book by Michael Brooks [14] focuses on outstanding mysteries across a number of scientific fields, from dark matter to abiogenesis.

3) People tend to ask question where we have the most complete information (e.g. look where the light shines brightest, not where the answer actually is):

* this leads us to make the comparison between prediction (function) and phenomenology (structure). Which is better? Are the relative benefits for each mode of investigation problem-dependent?


Stepping back from the solution space-as-Starfleet mission metaphor, we tend to study what we can measure well, and in turn what we can characterize well. But what is the relationship between characterization, measurement, and a solution (or ground-breaking scientific finding)? There are two components in the progression from characterization to solution. The first is to characterize enough of the problem space so as to create a measure, and then use that measure for further characterization, ultimately arriving at a solution. The second is to characterize enough of the problem space so that coherent questions can be asked, which allows a solution to be derived. When combined, this may provide the best tradeoff between precision and profundity. 

Yet which should come first, the measurement or the question? This largely depends on the nature of the measurement. In some cases, measures are more universal than single questions or solutions (e.g. information entropy, fMRI, optical spectroscopy). Metric spaces are also subject to this universality. If a measure can lead to any possible solution, then it is much more independent of question. In “Ignorance”, von Neumann's universal constructor, as applied to NKS theory by Wolfram [15] are discussed as a potentially universal measurement scheme. 


There are two additional points I found intriguing. The first is a fundamental difference between scientific fields where there is a high degree of "ignorance" (e.g. neuroscience) versus those where there is a relatively low degree (e.g. particle physics). This is not a new observation, but has implications for applied science. For example, the interferometer is a tool used in the physical sciences to build inferences between and find information among signals in a system. Would it be possible to build an interferometer based on neural data? Yes and no. While there is an emerging technology called brain-machine interfaces (BMI), these interfaces are limited to well-characterized signals and favorable electrophysiological conditions [16]. Indeed, as we uncover increasingly more about brain function, perhaps brain-machine interface technology will become closer to being like an interferometer. Or perhaps not, which would reveal a lot about how intractable ignorance (e.g. abundance of unknowable features) might be in this field. 

The second point involves the nature of innovation, or rather, innovations which lead to useful inventions. It is generally thought that engaging in applied science is the shortest route to success in this area. After all, pure research (e.g. asking questions about the truly unknown) invoves blind trial-and-error and ad-hoc experiments which lead to hard-to-interpret results. Yet "Ignorance" author Firestein argues that pure research might be more useful in terms of generating future innovation than we might recognize. This is becuase while there are many blind alleyways in the land of pure research, there are also many opportunities for serendipity (e.g. luck). It is the experiments that benefit from luck which potentially drive innovation along the furthest.

NOTES:

[1] One example is Brian Greene, a popularizer of string theory and theoretical physics and faculty member at NYU.

[2] via researching the literature and internal conversations among colleagues.

[3] for an example, please see: Frijters, R., van Vugt, M., Smeets, R., van Schaik, R., de Vlieg, J., and Alkema, W. (2010). Literature Mining for the Discovery of Hidden Connections between Drugs, Genes and Diseases. PLoS Computational Biology, 6(9), e1000943.

[4] Yong, E. (2012). Bad Copy. Nature, 485, 298.

[5] Ioannidis, J.P. (2005). Why Most Published Research Findings Are False. PLoS Medicine, 2(8), e124 (2005).

[6] Also the author of the following book: Hawkins, J. and Blakeslee, S. (2004). On Intelligence. Basic Books.

[7] here is a list of examples for adaptation and natural selection (COURTESY: PBS).

[8] for an example of how this has affected the field of Psychology, please see: Sternberg, R.J. and Grigorenko, E.L. (2001). Unified Psychology. American Psychologist, 56(12), 1069-1079.

[9] this was true of biology in the 19th century, and neuroimaging in the late 20th century. There are likely other examples I have not included here.

[10] here are more information on curve fitting (demo) and normalization (Wiki).

[11] this was discussed in the HTDE Workshop (2012). What is the computational complexity of a scientific problem? Can this be solved via parallel computing, high-throughput simulation, or other strategies?

Here are some additional insights from the philosophy of science (A) and the emerging literature on solving well-defined problems through optimizing experimental design (B):

A] Casti, J.L. (1989). Paradigms Lost: images of man in the mirror of science. William Morrow.

B] Feala, J.D., Cortes, J., Duxbury, P.M., Piermarocchi, C., McCulloch, A.D., and Paternostro, G. (2010). Systems approaches and algorithms for discovery of combinatorial therapies. Wiley Interdisciplinary Reviews: Systems Biology and Medicine, 2, 127.  AND  Lee, C.J. and Harper, M. (2012). Basic Experiment Planning via Information Metrics: the RoboMendel Problem. arXiv, 1210.4808 [cs.IT].

[12] our spacetime topology corresponds to a metric space, a common context, or conceptual framework. In an operational sense, this could be the dynamic range of a measurement device or the logical structure of a theory.

[13] I have no idea how this would square away with current theory. For now, let’s just consider the possibility.....

[14] the citation is: Brooks, M. (2009). 13 Things That Don't Make Sense. Doubleday.

[15] NKS (New Kind of Science) demos are featured at Wolfram's NKS website.

[16] this is a theme throughout the BMI and BCI literature, and includes variables such as type of signal used, patient population, and what is being controlled.

Printfriendly