Showing posts with label parallelism. Show all posts
Showing posts with label parallelism. Show all posts

September 28, 2013

Fireside Science: Bits of Blue-sky Scientific Computing

This is the first post in a series being cross-posted to Fireside Science (the new group blog sponsored by SciFund).


For my first post to Fireside Science, I would like to discuss some advances in scientific computing from a "blue sky" perspective. I will approach this exploration by talking about how computing can improve both the the modeling of our world and the analysis of data.


Better Science Through Computation

The traditional model of science has been (ideally) an interplay between theory and data. With the rise of high-fidelity and high-performance computing, however, simulation and data analysis has become a critical component in this dialogue. 

Simulation: controllable worlds, better science?

The need to deliver high-quality simulations of real-world scenarios in a controlled manner has many scientific benefits. Uses for these virtual environments include simulating hard-to-observe events (Supernovae or other events in stellar evolution) or provide highly-controlled environments for cognitive neuroscience experimentation (simulations relevant to human behavior).

A CAVE environment, being used for data visualization.

Virtual environments that achieve high levels of realism and customizability are rapidly becoming an integral asset to experimental science. Not only can stimuli be presented in a controlled manner, but all aspects of the environment (and even human interactions with the environment) can be quantified and tracked. This allows for three main improvements on the practice of science (discussed in greater detail in [1]):

1) Better ecological validity. In psychology and other experimental sciences, high ecological validity allows for the results of a given experiment to be generalized across contexts. High ecological validity results from environments which do not differ greatly from conditions found in the real-world.

Modern virtual settings allow for high degrees of environmental complexity to be replicated in a way that does not impede normal patterns of interaction. Modern virtual worlds allows for interaction using gaze, touch, and other means often used in the real-world. Contrast this with a 1980s era video game: we have come a long way since crude interactions with 8-bit characters using a joystick. And it will only get better in the future.


Virtual environments have made the cover of major scientific journals, and have great potential in scientific discovery as well [1].

2) The customization of environmental variables. While behavioral and biological scientists often talk about the effects of environment, these effects must often remain qualitative (or at best crudely quantitative). With virtual environments, environmental variables be added, subtracted, and manipulated in a controlled fashion.

Not only can the presence/absence and intensities of these variables be directly measured, but the interactions between virtual environment objects and an individual (e.g. human or animal subject) can be directly detected and quantified as well.

3) Greater compatibility with big data and computational dynamics: The continuous tracking of all environmental and interaction information results in the immediate conversion of this information to computable form [2]. This allows us to build more complete models of the complex processes underlying behavior or discover subtle patterns in the data.

Big Data Models

Once you have data, what do you do with it? That's a question that many social scientists and biologists have traditionally taken for granted. With the concurrent rise of high-throughput data collection (e.g. next-gen sequencing) and high-performance computing (HPC), however, this is becoming an important issue for reconsideration. Here I will briefly highlight some recent developments in big data-related computing.

Big data can come from many sources. High-throughput experiments in biology (e.g. next-generation sequencing) is one such example. The internet and sensor networks also provide a source of large datasets. Big datasets and difficult problems [3] require computing resources that are many times more powerful than what is currently available to the casual computer user. Enter petabyte (or petascale) computing.

National Petascale Computing Facility (Blue Waters, UIUC). COURTESY: Wikipedia.

Most new laptop computers (circa 2013) are examples of gigabyte computing. These computers utilize 2 to 4 processors (often using only one at a time). Supercomputers such as the Blue Waters computer at UIUC have many more processors, and operate at the petabyte scale [4]. Supercomputers such as IBM's Roadrunner, had well over 10,000 processors. Some of the most powerful computers even run at the exascale (e.g. 1000x faster than petascale). The point of all this computing power is to perform many calculations quickly, as the complexity of a very large dataset can make its analysis impractical using small-scale devices.

Even using petascale machines, difficult problems (such as drug discovery or very-large phylogenetic analyses) can take an unreasonable amount of time when run serially. So increasingly, scientists are also using parallel computing as a strategy for analyzing and processing big data. Parallel computing involves dividing up the task of computation amongst multiple processors so as to reduce the overall amount of compute time. This requires specialized hardware and advances in software, as the algorithms and tools designed for small-scale computing (e.g. analyses done on a laptop) are often inadequate to take full advantage of the parallel processing that supercomputers enable.

Physical size of the Cray Jaguar supercomputer. Petascale computing courtesy of the Oak Ridge National Lab.

Media-based Computation and Natural Systems Lab

This is an idea I presented to a Social Simulation conference (hosted in Second Life) back in 2007. The idea involves building a virtual world that would be accessible to people from around the world. Experiments could then be conducted through the use of virtual models, avatars, secondary data, and data capture interfaces (e.g. motion sensors, physiological state sensors).

The Media-based Computation and Natural Systems (CNS) Lab, in its original Second Life location, circa 2007.

The CNS Lab (as proposed) features two components related to experiments not easily done in the real-world [5]. This is an extension of virtual environments to a domain that is relatively unexplored using virtual environments: the interface between the biological world and the virtual world. With increasingly sophisticated I/O devices and increases in computational power, we might be able to simulate and replicate the black box of physiological processes and the hard-to-observe process of long-term phenotypic adaptation.

Component #1: A real-time experiment demonstrating the effect of extreme environments on the human body. 

This would be a simulation to demonstrate and understand the limits of human physiological capacity usually observed in limited contexts [6]. In the virtual world, an avatar would enter a long tube or tank, the depth of which would serve as a environmental gradient. As the avatar moves deeper into the length of the tube, several parameters representing variables such as atmospheric pressure, temperature, and medium would increase or decrease accordingly.

There should also be ways to map individual-level variation to the avatar in order to provide some connection between the participant and the simulation of human physiology. Because this experience is distributed on the internet (originally proposed as a Second Life application) a variety of individuals could experience and participate in an experiment once limited to a physiology laboratory.

Examples of deep-sea fishes (from top): Barreleye (Macropinna microstoma), Fangtooth (Anoplogaster cornuta), Frilled Shark (Chlamydoselachus anguineus)COURTESY: National Geographic and Monterey Bay Aquarium.

Component #2: An exploration of deep sea fish anatomy and physiology. 

Deep sea fishes are used as an example of organisms that adapted to deep sea environments that may have evolved from ancestral forms originating in shallow, coastal environments [7]. The object of this simulation is to observe a “population” change over from ancestral pelagic fishes to derived deep sea fishes as environmental parameters within the tank change. The participant will be able to watch evolution “in progress” through a time-elapsed overview of fish phylogeny.

This would be an opportunity to observe adaptation as it happens, in a way not necessarily possible in real-world experimentation. The key components of the simulation would be: 1) time-elapsed morphological change and 2) the ability to examine a virtual model of the morphology before and after adaptation. While these capabilities would be largely (and in some cases wholly) inferential, it would provide an interactive means to better appreciate the effects of macroevolution.

A highly stylized (e.g. scala naturae) view of improving techniques in human discovery, culminating in computing.

A tongue-in-cheek cartoon showing the evolution of computer storage (as opposed to processing power). Nevertheless, this is pretty rapid evolution.

NOTES:

[1] These journal covers are in reference to the following articles: Science cover, Bainbridge, W.S.   The Scientific Research Potential of Virtual Worlds. Science, 317, 412 (2007). Nature Reviews Neuroscience cover, Bohil, C., Alicea, B., and Biocca, F. Virtual Reality in Neuroscience Research and Therapy. Nature Reviews Neuroscience, 12, 752-762 (2011).

[2] Raw numeric data, measurement indices, and, ultimately, zeros and ones.

[3] Garcia-Risueno, P. and Ibanez, P.E.   A review of High Performance Computing foundations for scientists. arXiv, 1205.5177 (2012).

For a very basic introduction to big data, please see: Mayer-Schonberger, V. and Cukier, K.   Big Data: a revolution that will transform how we live, work, and think. Eamon Dolan (2013).

[4] Hemsoth, N.   Inside the National Petascale Computing Facility. HPCWire blog, May 12 (2011).

[5] Alicea, B.   Reverse Distributed Computing: doing science experiments in Second Life. European Social Simulation Association/Artificial Life Group (2007).

[6] Downey, G.   Human (amphibious model): living in and on the water. Neuroanthropology blog, February 3 (2011).

For an example of how human adaptability in extreme environments has traditionally been quantified, please see: LeScanff, C., Larue, J., and Rosnet, E.   How to measure human adaptation in extreme environments: the case of Antarctic wintering-over. Aviation, Space, and Environmental Medicine, 68(12), 1144-1149 (1997).

[7] For more information on deep sea fishes, please see: Romero, A.   The Biology of Hypogean Fishes. Developments in Environmental Biology of Fishes, Vol. 21. Springer (2002).

January 9, 2013

Playing for Science


This is being re-posted from my microblog, Tumbld Thoughts.


Here is an interesting article from "The Scientist" called "Games for Science". It features a number of games designed for massively parallel data analysis [1]. The idea is to distribute data to personal computers, perform operations there, and them re-assemble the analyzed data at its source. Two projects are EteRNA (RNA structural design) and Phylo (multiple sequence alignments).

[1] one example is BOINC (admined at UC-Berkeley). An example of distributed computing.

August 27, 2012

Degeneracy: a central mechanism in evolution

In this post, I will provide an overview of a concept in evolutionary biology called degeneracy. Degeneracy has been defined by [1, 2] as structurally different entities that perform identical functions or yields identical outputs. From a complex systems perspective, degeneracy is closely related to redundancy and robustness [1]. Yet not as often cited as the other two terms [3], degeneracy is still an important feature of evolved biological systems (see Figure 1). For example, degeneracy may explain the absence of key proteins in up to 30% of healthy patients, or the absence of growth defects in yeast with deleted genes of known functional consequence [1]. It is a prominant property of genetic regulatory networks, as most examples characterized in the literature are tied to gene regulatory function in one way or another. For a wider range of examples, Edelman and Gally [1] provide a list of 22 examples from a range of biological systems (see Figure 2 for how degeneracy fits into the broader context of biological complexity).

Figure 1. A comparison of the number of times each term ("degeneracy", "redundancy", and "robustness") appears in the scientific literature.

Evidence for degeneracy can be found in the existence of multiple routes to a specific physiological function. According to [2], members of the adhesins gene family in yeast can interchangably perform functional roles when expression is elevated. This is in contrast to lock-and-key systems (e.g. receptor-ligand binding), which interact with a high degree of specificity [4]. Another common signature of degeneracy is cross-talk, which is common in gene regulatory and neural pathways. The presence of degenerate pathways (or at least the name for them) suggests a "degeneration" from some previous state. This is at least partially correct: it appears that degenerate relationships involve a generalization of function over evolutionary time. However, this need not result from a mechanism that was originally functionally specialized.

To understand this in context, let us return to the "lock-and-key" model. Lock-and-key systems are highly specialized with regard to function. In fact, if one were to consider only the end product, we might be tempted to conclude a purposeful design. However, if we consider that both the lock AND key have shared evolutionary histories, it becomes more possible that this arrangement is not only the product of mutation-selection dynamics but historical contingency as well [5]. Lock-and-key phenomena result from selection for extreme specialization [1]. While this might have a fitness advantage in some contexts, in highly-veriable environments it is not particularly advantageous. Thus, the historical lock-in [6] that inadvertently results from selection can results in evolutionary dead-ends. What degeneracy provides, then, is a means to either rescue a phenotype from or circumvent entirely such instances.

There are also potential evolutionary tradeoffs between network functionality and function of the individual components of this network. One example of this is the role of an individual gene in a genetic regulatory network. Whether degeneracy result from a true evolutionary tradeoff or as a signature of a complex system's emergent properties [8] is not clear. However, Whitacre and Bender [9] modeled biological networks as a complex adaptive system (CAS). Using this approach, robustness was found to result from both diversity and degeneracy. In this case, invariance to perturbation (e.g. robustness) results from many possible ways to achieve a common function. Diversity allows for the biological system to move away from the extreme specificity required of the "lock-and-key" model [10], while a degenerate architecture incorporates this diversity into a functional mechanism.


Figure 2. A schematic showing the relationship between degeneracy and the related concepts of complexity, robustness, and evolvability. Adapted from Figure 1 in [7].


Tononi, Sporns, and Edelman [11] have proposed ways to quantify degeneracy in biological networks. The quantification is based on the notion that redundancy and degeneracy stand in contrast to all outputs of a system (e.g. gene network) being statistically independent of one another. The concept of mutual information [12] is used to quantify the degree of a shared functional role between output. If two or more output share information, the related functions are said to exhibit degeneracy. Another quantitative approach is to think of degenerate biological systems as degenerate sets of overlapping functions [1, 13]. Given its mathematical similarity to phylogenetic theory, such an approach might reveal new insights into convergent evolution.

What can degeneracy teach us about complex biological systems? One lesson is that in some cases there may be a fitness benefit for maintaining a parallel architecture. While not clearly beneficial in every context, parallelism can perform critical functions such as buffer against mutations or act as a noise filter [14]. A second lesson is that degeneracy is not always degenerate: far from being a failure of optimization, degeneracy provides a means to incorporate the stuff of evolutionary time (mutation) into a system that does not become reliant on any single pathway (specificity). In this way, degenerate biological systems are often the most adaptable, which means that some outcomes of the evolutionary process can truly be described as "survival of the most degenerate".

NOTES:

[1] Edelman, G.M. and Gally, J.A.   Degeneracy and complexity in biological systems. PNAS, 98(24) 13763–13768 (2001).

[2] Whitacre, J. and Bender, A.   Degeneracy: A design principle for achieving robustness and evolvability. Journal of Theoretical Biology 263 (2010) 143–153.

[3] Data for graph courtesy of PubMed, search data August 27, 2012.

[4] Adami, C.   Reducible Complexity. Science, 312(5770), 61-63 (2006)  AND  Brouat, C., Garcia, N., Andary, C., and McKey, D.   Plant lock and ant key: pairwise coevolution of an exclusion filter in an ant-plant mutualism. Proceedings of the Royal Society of London B, 268, 2131-2141 (2001).

[5] For concept of historical contingency, please see: Swartz, B.A.   On the “Duel” Nature of History: Revisiting Contingency versus Determinism. PLoS Biology, 7(12), e1000259 (2009). AND Fontana, W. and Schuster, P.   Continuity in Evolution: On the Nature of Transitions. Science, 280(5368), 1451-1455 (1998).

[6] References to historical lock-in can be found in: Nelson, R. and Winter, S.   An evolutionary theory of economic change, Harvard University Press, Cambridge, MA (1982). This term is often used in the economics and business literature, but also has relevance to the biological world.

[7] Whitacre, J.M.   Degeneracy: a  link  between  evolvability, robustness and complexity in biological systems. Theoretical Biology and Medical  Modeling, 7(6), 6 (2010).

[8] A system with emergent properties produces an output which is greater than the sum of its parts. Weak emergence is a case where the collective effects can be reduced to its individual components, while strong emergence results in collective effects that are irreducible. In many cases, biological evolution can be thought of as resulting from strong emergence.

For information specific to evolution, please see: Blitz, D.   Emergent Evolution: Qualitative Novelty and the Levels of Reality. Kluwer Academic, Dordrecht (1992) AND Bedau, M.A.   Downward causation and autonomy in weak emergence. Principia, 6, 5-50 (2003).

[9] Whitacre, J.M. and Bender, A.   Networked buffering: a basic mechanism for distributed robustness in complex adaptive systems. Theoretical Biology and Medical Modeling, 7, 20 (2010).

* the complex adaptive systems (CAS) approach is a way to model systems of high complexity using a series of interacting agents that originated out of the Santa Fe Institute.

For a general overview, please see: Holland, J.H.   Studying Complex Adaptive Systems. Journal of Systems Science and Complexity, 19, 1–8 (2006) AND Miller, J.H. and Page, S.E.   Complex Adaptive Systems: An Introduction to Computational Models of Social Life. Princeton University Press, Princeton, NJ.

[10] Gomez-Gardenes, J., Moreno, Y., and Floria, L.M.   On the robustness of complex heterogeneous gene expression networks. Biophysical Chemistry, 115, 225-228 (2005).

[11] Tononi, G., Sporns, O., and Edelman, G.   Measures of degeneracy and redundancy in biological networks. PNAS, 96, 3257-3262 (1999). This work was developed for studying brain networks, but theoretically can be applied to a wide range of biological systems, including genetic regulatory networks.

[12] For a general introduction, please see: Latham, P.E. and Roudi, Y.   Mutual information. Scholarpedia, 4(1), 1658 (2009). Link.

[13] I could find no references to this outside of the original paper (Edelman and Gally). I imagine it combines the mathematical concept of degeneracy (when objects change their set membership over time) and conventional set theory.

[14] For mutational and phenotypic buffering, please see: Braendle, C. and Felix, M.A.   Plasticity and Errors of a Robust Developmental System in Different Environments. Developmental Cell, 15(5), 714-724 (2008) AND Rutherford, S.L. and Lindquist, S.   Hsp90 as a capacitor for morphological evolution. Nature, 336-342 (1998).

For noise filtering in a genetic regulatory network, please see: Orrell, D. and Bolouri, H.   Control of internal and external noise in genetic regulatory networks. Journal of Theoretical Biology, 230(3), 301-312 (2004) AND Lestas, I., Paulsson, J., Ross, N.E., and Vinnicombe, G.   Noise in gene regulatory networks. IEEE Transactions in Automation and Control, 53, 189-200 (2008). .

** For the latest information on degeneracy research, please see the work of James Whitacre, admin of the Degeneracy and Selection online community. Also, the Wikipedia page for Degeneracy (biology) features a comprehensive bibliography featuring papers from many areas of biology.

March 30, 2012

A graphical, parallel biological world.....

This Wednesday (April 4) at 2pm Pacific time, I will be presenting a lecture entitled "Scenes from a grahical, parallel biological world" to the Embryo Physics groups in Second Life. The lecture (see slides here) will be an excursion into the world of GPU (graphical processing unit) computing. I have been interested in GPU computing for a few years now, and interested in computer graphics for significantly longer.

 Scene from the Embryo Physics course (between classes).


This is a different approach to using CUDA architectures, which for most scientists is all about solving their mathematically-intensive problems faster. My interest is a bit more philosophical. I wish to understand the connections between parallelism and high-powered graphics processing (e.g. image rendering) in the service of designing novel data structures for analyzing scientific data. In this case, I am interested in what we can learn from a host of biological systems (ranging from population biology to cell biology).

Perhaps this approach can be leveraged to better understand the self-organization and underlying context that characterizes many biological systems. This talk represents a first step in this synthesis, but should be interesting in any case.

UPDATE (4/4): A blogscript of the talk is available on this page.

August 21, 2011

Algorithms for Parallel Universes (potentially)

This last week, I attended a seminar called "Proven Algorithmic Techniques for Manycore Processors", presented by the Virtual School of Computational Science and Engineering (VSCSE), which is sponsored by the NCSA at the University of Illinois Urbana-Champaign (the same people who head up the Blue Waters supercomputer initiative). The topic was developing algorithms suitable for the architecture of many-core processors. This seminar opened by eyes to the challenges imposed by large, non-uniform datasets, not only in the context of parallel computing, but also as a broader research topic for statisticians and applied mathematics.


The focus of this seminar was on applications for the Graphical Processing Unit (GPU) technology developed by nVIDIA. The level and extent of instruction, especially from Dr. Hwu, was excellent. This is part of the NCSA summer school initiative, which is very useful for covering niche topics. I attended the "Big Data for Science" seminar last year, which was also very good). Here are my written notes from the session (I always find that written notes helps along the learning process).

February 3, 2011

Paper(s) of the Week (Colloids and Parallelism)


This week, I found two papers in Science that are worth taking a look at:

One is a materials science paper called "Supracollidal Reaction Kinetics of Janus Spheres". Science, 331, 199 (2011). link

In the first paper, particles called "Janus spheres" are used to build clusters and other structures. The term "Janus sphere" refers to their two-faced nature -- hydrophobic (water-repelling) on one hemisphere, and hydrophilic (water-loving) on the other. In this sense, they behave like dipoles. The authors did molecular dynamics simulations to show that these particles can self-assemble into higher-order structures such as fibrillar triple helices. A good article for anyone interested in molecular self-assembly.

The other is a biomimetics paper called "A Biological Solution to a Fundamentally Distributed Computing Problem". Science, 331, 183 (2011). link

The second paper is an attempt to design an algorithm that mimics cell autonomy in insect development for the purpose of designing "smart" autonomous distributed systems (such as networks or robot teams). The model system is molecular signaling during the development of sensory organ precursors in insects. Sensory organ precursor (SOP) cells are selected in development from many other like cells in a single proneural cluster to become bristles. This occurs through an "election" process whereby every cell in the cluster is connected to an SOP with no two SOPs being adjacent. This is similar to the maximal independent set problem in computing, where a network of processors are sorted into a topology similar to that of insect development using an election criterion. The authors come up with a 15-step algorithm which may be useful for students of optimization and many other topical areas.

Printfriendly