Showing posts with label fireside-science. Show all posts
Showing posts with label fireside-science. Show all posts

August 4, 2014

The Ukraine is Strong (for Synthetic Daisies)!

This post was written using Ubuntu 13.10 (Saucy Salamander), GIMP 2.6 and Blogilo. No Bitcoins (or their open-source alternatives) were transacted in its creation.


According to the game of Risk, the Ukraine is weak. But for Synthetic Daisies blog and as an example of viral content, the Ukraine is strong! In the past few days, a Synthetic Daisies (and Fireside Science) blog post called "Bitcoin Angst with an Annotated Blogroll" has gone viral in the Ukraine.
 


The associated pictures demonstrate how one can simply but effectively triangulate viral content from basic analytic data. This is also confirmed by the number of Pageviews made by users with alternative browsers and operating systems, which is either a Ukraine thing, a Bitcoin community thing, or both. In any case, keep up the diffusion!





June 21, 2014

Fireside Science: The Representation of Representations

This content is being cross-posted to Fireside Science, and is the third in a three-part series on the "science of science".


This is the final in a series of posts on the science of science and analysis. In past posts, we have covered theory and analysis. However, there is a third component of scientific inquiry: representation. So this post is about the representation of representations, and how representations shape science in more ways than the casual observer might believe.

The three-pronged model of science (theory, experiment, simulation). Image is adapted from Fermi Lab Today Newsletter, April 27 (2012).

For the uninitiated, science is mostly analysis and data collection with theory being a supplement at best and necessary evil at worst. Ideally, modern science rests on three pillars: experiment, theory, and simulation. For these same uninitiated, the representation of scientific problems is a mystery. But in fact, it has been the most important motivation for much of the scientific results we celebrate today. Interestingly, the field of computer science relies heavily on representation, but this concern generally does not carry over into the empirical sciences.

Ideagram (e.g. representation) of complex problem solving. Embedded are a series of Hypotheses and the processes that link them together. COURTESY: Diagram from [1].

Problem Representation
So exactly what is scientific problem representation? In short, it is the basis for designing experiments and conceiving of models. It is the sieve through which scientific inquiry flows, restricting the typical "question to be asked" to the most plausible or fruitful avenues. It is often the basis of consensus and assumptions. On the other hand, representation is quite a bit more subjective than people typically would like their scientific inquiry to be. Yet this subjectivity need not lead to an endless debate about the validity of one point of view versus another. There are heuristics one can use to ensure that problems are represented in a consistent and non-leading way.

3-D Chess: a high-dimensional representation of warfare and strategy.

Models that Converge
Convergent models speaks to something I alluded to in "Structure and Theory of Theories" when I discussed the theoretical landscape of different academic fields. The first way is whether or not allied sciences or models point in the same direction. To do this, I will use a semi-hypothetical example. The hypothetical case is to consider three models (A, B, and C) of the same phenomenon. Each of these models make different assumptions and includes different factors, but should at least be consistent with each other. One real-world example of this is the use of gene trees (phylogenies) and species trees (phylogenies) to understand evolution in a lineage [2]. In this case, each model uses the same taxa (evolutionary scenario), but includes incongruent data. While there are a host of empirical reasons why these two models can exhibit incongruence [3], models that are as representationally complete as possible might resolve these issues.

Orientation of Causality
The second way is to ensure that the one's representation gets the source of causality right. For problems that are not well-posed or poorly characterized, this can be an issue. Let's take Type III errors [4] as an example of this. In hypothesis testing, type III errors involve using the wrong explanation for a significant result. In layman's terms, this is getting the right answer for the wrong reasons. Even more than in the  case of type I and II errors, focusing on the correct problem representation plays a critical role in resolving potential type III errors.

Yet problem representation does not always help resolve these types of errors. Skeptical interpretation of the data can also be useful [5]. To demonstrate this, let us turn to the over-hyped area of epigenetics and its larger place in evolutionary theory. Clearly, epigenetics plays some role in the evolution of life, but is not deeply established in terms of models and theory. Because of this representational ambiguity, some interpretations play a trick. In a conceptual representation that embodies this trick, scarcely-understood high-level phenomena such as epigenetics will usurp the role of related phenomena such as genetic diversity and population processes. When the thing in your representation is not well-defined or quite popular (e.g. epigenetics), it can take on a causal life of its own. Posing the problem in this way allows us to obscure known dependencies between genes, genetic regulation, and the environment without proving exceptions to these established relationships.

Popularity is Not Sufficiency
The third way is to understand that popular conceptions do not translate into representational sufficiency. In logical deduction, it is often pointed out that necessity does not equal sufficiency. But as with the epigenetics example, it also holds that popularity cannot make something sufficient in and of itself. In my opinion, this is one of the problems with using narrative structures in the communication of science: sometimes an appealing narrative does more to obscure scientific findings than it does in making things accessible to lay people.

Fortunately, this can be shown by looking at media coverage of any big news story. The CNN plane coverage [6] shows this quite clearly: coverage of rampant speculation and conspiracy theory was a way to emphasize an increasingly popular story. In such cases, speculation is the order of the day, while thoughtful analysis gets pushed aside. But is this simply a sin of the uninitiated, or can we see parallels of this in science? Most certainly, there is a problem with recognizing the difference between "popular" science and worthwhile science [7]. There is also precedence from the way in which certain studies or areas of study are hyped. Some in the scientific community [8] have argued that Nature's hype of the ENCODE project [9] results fell into this category.

One example of a mesofact: ratings for the TV show The Simpsons over the course of several hundred episodes. COURTESY: Statistical analysis in [10].

Mesofacts
Related to these points is the explicit relationship between data and problem representation. In some ways, this brings us back to a computational view of science, where data do not make sense unless it is viewed in the context of a data structure. But sometimes the factual aspect of data varies over time in a way that obscures our mental models, and in turn obscures problem representation.

To make this explicit, Sam Arbesman has coined the term "mesofact" [11]. A mesofact is knowledge that changes slowly over time given new data. Populations of specific places (e.g. Minneapolis, Bolivia, Africa) has changed in both absolute and relative terms over the past 50 years. But when problems and experimental designs are formulated assuming that facts related to these data (e.g. rank of cities by population) do not change over time, we can get the analysis fundamentally wrong.

This may seem like a trivial example. However, mesofacts have relevance to a host of problems in science, from experimental replication to inferring the proper order of causation. The problem comes down to an interaction between data's natural variance (variables) and the constructs used to represent our variables (facts). When the data exhibit variance against an unchanging mean, it is much easier to use this variable as a stand-in for facts. But when this is not true, scientifically-rigorous facts are much harder to come by. Instead of getting into an endless discussion about the nature of facts, we can instead look to how facts and problem representation might help us tease out the more metaphysical aspects of experimentation.

Applying Problem Representation to Experimental Manipulation
When we do experiments, how do we know what our experimental manipulations really mean? The question itself seems self-evident, but perhaps it is worth exploring. Suppose that you wanted to explore the causes of mental illness, but did not have the benefits of modern brain science as a guide. In defining mental illness itself, you might work from a behavioral diagnosis. But the mechanisms would still be a mystery. Is it a supernatural mechanism (e.g. demons) [12], an ultimate form of causation (reductionism), or a global but hard-to-see mechanism (e.g. quantum something) [13]? An experiment done the same way but assuming three different architectures could conceivably yield statistical significance for all of them.

In this case, a critical assessment of problem representation might be able to resolve this ambiguity. This is something that as modelers and approximators, computational scientists deal with all of the time. Yet it is also an implicit (and perhaps even more fundamental) component of experimental science. For most of the scientific method's history, we have gotten around this fundamental concern by relying on reductionism. But in doing so, this restricts us to doing highly-focused science without appealing to the big picture. In a sense, we are blinded by science by doing science.

Focusing on problem representation allows us a way out of this. Not only does it allow us to break free from the straightjacket of reductionism, but also allows us to address the problem of experimental replication more directly. As has been discussed in many other venues [14], the lack of an ability to replicate experiments has plagued both Psychological and Medical research. But it is in these areas which representation is most important, primarily because it is hard to get right. Even in cases where the causal mechanism is known, the underlying components and the amount of variance they explain can vary substantially from experiment to experiment.

Theoretical Shorthand as Representation
Problem representation also allows us to make theoretical statements using mathematical shorthand. In this case, we face the same problem as the empiricist: are we focusing on the right variables? More to the point, are these variables fundamental or superficial? To flesh this out, I will discuss two examples of theoretical shorthand, and whether or not they might be concentrating on the deepest (and most generalizable) constructs possible.

The first example comes from Hamilton's rule, derived by the behavioral ecologist W.D. Hamilton [15]. Hamilton's rule describes altruistic behavior in terms of kin selection. The rule is a simple linear equation that assumes adaptive outcomes will be optimal ones. In terms of a representation, these properties provide a sort of elegance that makes it very popular.


In this short representation, an individual's relatedness to a conspecific contributes more to their behavioral motivation to help that individual than a typical trade-off between costs and benefits. Thus, a closely-related conspecific (e.g. a brother) will invest more into a social relationship with their kin than with non-kin. In general, they will take more personal risks in doing so. While more math is used to support the logic of this statement [15], this inequality is often treated as a widely applicable theoretical statement. However, some observers [16] have found the parsimony of this representation to be both too incomplete and intellectually unsatisfying. And indeed, sometimes an over-simplistic model does not deal with exceptions well.

The second example comes from Thomas Piketty's work. Piketty, economist and author of "Capital in the 21rst Century" [17], has proposed something he calls the "First Law" which explains how income inequality relates to economic growth. The formulation, also a simple inequality, characterizes the relationship between economic growth, inherited wealth, and income inequality within a society.


In this equally short representation, inequality is driven by the relative dominance of two factors: inherited wealth and economic growth. When growth is very low, and inherited wealth exists at a nominal level, inequality persists and dampens economic mobility. In Piketty's book, other equations and a good amount of empirical investigation is used to support this statement. Yet, despite its simplicity, it has held up (so far) to the scrutiny of peer review [18]. In this case, representation through variables that generalize greatly but do not handle exceptional behavior well produce a highly-predictive model. On the other hand, this form of representation also makes it hard to distinguish between a highly unequal post-industrial society and a feudal, agrarian one.

Final Thoughts
I hope to have shown you that representation is an underappreciated component of doing and understanding science. While the scientific method is our best strategy for discovering new knowledge about the natural world, it is not without its burden of conceptual complexity. In the theory of theories, we learned that formal theories are based on both deep reasoning and are (by necessity) often incomplete. In the analysis of analyses, we learned that the data are not absolute. Much reflection and analytical detail must be taken to ensure that an analysis represents meaningful facets of reality. And in this post, these loose ends were tied together in the form of problem representation. While an underappreciated aspect of practicing science, representing problems in the right way is essential for separating out science from pseudoscience, reality from myth, and proper inference from hopeful inference.

NOTES:
[1] Eldrett, G.   The art of complex problem-solving. MediaExplored blog, July 10 (2010).

[2] Nichols, R.   Gene trees and species trees are not the same. Trends in Ecology and Evolution, 16(7), 358-364 (2001).

[3] Gene trees and species trees can be incongruent for many reasons. Nature Knowledge Project (2012).

[4] Schwartz, S. and Carpenter, K.M.   The right answer for the wrong question: consequences of type III error for public health research. American Journal of Public Health, 89(8), 1175–1180 (1999).

[5] It is important here to distinguish between careful skepticism and contrarian skepticism. In addition, skeptical analysis is not always compatible with the scientific method.

For more, please see: Myers, P.Z.   The difference between skeptical thinking and scientific thinking. Pharyngula blog, June 18 (2014) AND Hugin   The difference between "skepticism" and "critical thinking"? RationalSkepticism.org, May 19 (2010).

[6] Abbruzzese, J.   Why CNN is obsessed with Flight 370: "The Audience has Spoken". Mashable, May 9 (2014).

[7] Biba, E.   Why the government should fund unpopular science. Popular Science, October 4 (2013).

[8] Here are just a few examples of the pushback against the ENCODE hype:


a) Mount, S.   ENCODE: Data, Junk and Hype. On Genetics blog, September 8 (2012).

b) Boyle, R.   The Drama Over Project Encode, And Why Big Science And Small Science Are Different. Popular Science, February 25 (2013).

c) Moran, L.A.   How does Nature deal with the ENCODE publicity hype that it created? Sandwalk blog, May 9 (2014).

[9] For an example of the nature of this hype, please see: The Story of You: ENCODE and the human genome. Nature Video, YouTube, September 10 (2012).

[10] Fernihough, A.   Kalkalash! Pinpointing the Moments “The Simpsons” became less Cromulent. DiffusePrior blog, April 30 (2013).

[11] Arbesman, S.   Warning: your reality is out of date. Boston Globe, February 28 (2010). Also see the following website: http://www.mesofacts.org/

[12] Surprisingly, this is a contemporary phenomenon: Irmak, M.K.   Schizophrenia or Possession? Journal of Religion and Health, 53, 773-777 (2014). For a thorough critique, please see: Coyne, J.   Academic journal suggests that schizophrenia may be caused by demons. Why Evolution is True blog, June 10 (2014).

[13] This is an approach favored by Deepak Chopra. He borrows the rather obscure idea of "nonlocality" (yes, basically a wormhole in spacetime) to explain higher levels of conscious awareness with states of brain activity.

[14] Three (divergent) takes on this:

a) Unreliable Research: trouble at the lab. Economist, October 19 (2013).

b) Ioannidis, J.P.A.   Why Most Published Research Findings Are False. PLoS Med 2(8): e124 (2005).

c) Alicea, B.   The Inefficiency (and Information Content) of Scientific Discovery. Synthetic Daisies blog, November 19 (2013).

[15] Hamilton, W. D.   The Genetical Evolution of Social Behavior. Journal of Theoretical Biology, 7(1), 1–16 (1964). See also: Brembs, B.   Hamilton's Theory. Encyclopedia of Genetics.

[16] Goodnight, C.   Why I Don’t like Kin Selection. Evolution in Structured Populations blog, April 23 (2014).

[17] Piketty, T.   Capital in the 21st Century. Belknap Press (2014). See also: Galbraith, J.K.   Unpacking the First Fundamental Law. Economist's View blog, May 25 (2014).

[18] DeLong, B.   Trying, yet again, to communicate the arithmetic scaffolding of Piketty's "capital in the Twenty-First Century". Washington Center for Equitable Growth blog, June 5 (2014).

May 12, 2014

Fireside Science: The Analysis of Analyses

This material is cross-posted to Fireside Science. This is part of a continuing series on the science of science (or meta-science, if you prefer). The last post was about the structure and theory of theories.


In this post, I will discuss the role of data analysis and interpretation. Why do we need data, as opposed to simply observing the world or making up stories? The simple answer: it gives us a systematic accounting of the world in general and experimental manipulations in particular. As opposed to the apparition on a piece of toast, it provides a systematic accounting of the natural world independent of our sensory and conceptual biases. But as we saw in the theory of theories post, and as we will see in this post, it takes a lot of hard work and thoughtfulness. What we end up with is an analysis of analyses.

Data take many forms, so approach analysis with caution. COURTESY: [1].

Introduction
What exactly is data, anyways? We hear a lot about it, but rarely stop to consider why it is so potentially powerful. Data are both an abstraction of and incomplete sampling (approximation) of the real world. While the data are not absolute (e.g. you can always have more data or more completely sample the world), the data provide a means of generalization that is partially free from stereotyping. And as we can see in the cartoon above, not all data that influence our hypothesis can even be measured. Some of it is beyond the scope of our current focus and technology (e.g hidden variables), while some of it consists of interactions between variables.

In the context of the theory of theories, data has the same advantage over anecdote that deep, informed theories have over naive theories. In the context of the analysis of analyses, data does not speak for itself. To conduct a successful analysis of analysis, it is important to be both interpretive and objective. Finding the optimal balance between each of these gives us an opportunity to reason more clearly and completely. If this causes some people to lose their view of data as infallible, then so be it. Sometimes the data fails us, and other times we fail ourselves.

When it comes to interpreting data, the social psychologist Jon Haidt suggests that "we think we are scientists, but we are actually laywers" [2]. But I would argue this is where the difference between the untrained eyes sharing Infographics and the truly informed acts of analysis and data interpretation becomes important. The latter is an example of a meta-meta-analysis, or a true analysis of analyses.

The implications of Infographics are clear (or are they?) COURTESY: Heatmap, xkcd.

NHST: the incomplete analysis?
I will begin our discussion with a current hot topic in the field of analysis. It involves interpreting statistical "significance" using an approach called Null Hypothesis Statistical Testing (or NHST). If you have even done a t-test or ANOVA, you have used this approach. The current discussion about the scientific replication crisis is tied to the use (and perhaps overuse) of these types of tests. The basic criticism involves the inability of NHST statistics to conduct multiple tests properly and properly deal with experimental replication.

Example of the NHST and its implications. COURTESY: UC Davis StatWiki.

This has even led scientists such as John Ioannidis to demonstrate why "most significant results are wrong". But perhaps this is just to make a rhetorical point. The truth is, our data are inherently noisy. Too many assumptions/biases go into collecting most datasets, all for data which has too little known structure. Not only are our data noisy, but in some cases may also possess hidden structure which violates the core assumptions of many statistical tests [3]. Some people have rashly (and boldly) proposed that this points to flaws in the entire scientific enterprise. But, like most things, this does not take into account the nature of the empirical enterprise and reification of the word significance.

A bimodal (e.g. non-normal) distribution, being admonished by its unimodal brethren. Just one case in which the NHST might fail us.

The main problem with the NHST is that it relies upon distinguishing signal from noise [4], but not always in the broader context of effects size or statistical power. In a Nature News correspondence [5], Regina Nuzzo discusses the shortcomings of the NHST approach and tests of statistical significance (e.g. p-values). Historical context of the so-called frequentist approach [6] is provided, and its connection to assessing the validity of experimental replications are discussed. One possible solution is the use of Bayesian techniques [7] to assess something called statistical power. The Bayesian approach allows one to use a prior distribution (or historical conditioning) to better assess the meaningfulness of one's statistically significant result. But the construction of priors relies on the existence of reliable data. If these data do not exist for some reason, we are back to square one.

Big Data and its Discontents
Another challenge to conventional analysis involves the rise of so-called big data. Big data is the collection and analysis of very large datasets, which come from sources such as high-throughput biology experiments, computational social science, open-data repositories, and sensor networks. Considering their size, big data analyses should allow for good power and ability to distinguish signal from noise. Yet due to their structure, we are often required to rely upon correlative analyses. While correlation is equated with relational information, it (as it always has) does not equate to causation [8]. Innovations in machine learning and other data modeling techniques can sometimes overcome this limitation, but correlative analyses are still the easiest way to deal with these data.

IBM's Watson: powered by large databases and correlative inference. Sometimes this cognitive heuristic works well, sometimes not so much.

Given a large enough collection of variables with a large number of observations, correlations can lead to accurate generalizations about the world [9]. The large number of variables are needed to extract relationships, while the large number of observations are needed to understand the true variance. This can be a problem where subtle, higher-order relationships (e.g. feedbacks, time-dependent saturations) exist or when the variance is not uniform with respect to the mean (e.g. bimodal distributions).

Complex Analyses
Sometimes large datasets require more complicated methods to find relevant and interesting features. These features can be thought of as solutions. How do we use complex analysis to find these features? In the world of analysis of analyses, large datasets can be mapped to solution spaces with a defined shape. This strategy uses convergence/triangulation as a guiding principle, but does so through the rules of metric geometry and computational complexity. A related and emerging approach called topological data analysis [10] can be used to conduct rigorous relational analyses. Topological data analysis takes datasets and maps them to a geometric shape (e.g. topology) such as a tree or in this case a surface.


A portrait of convexity (quadratic function). A gently sloping dataset, a gently sloping hypothesis space. And nothing could be further from the truth......

In topological data analyses, the solution space encloses all possible answers on a surface, while the surface itself has a shape that represents how easy it is to move from one portion of the solution space to another.
One common assumption is that this solution space is known and finite, while the shape is convex (e.g. a gentle curve). If that were always true, then analysis would be easy: we could use a moderate large-sized dataset to get the gist of patterns in the data. any additional scientific inquiry would constitute filling in the gaps. And indeed sometimes it works out this way.

One example of a topological data analysis of most likely Basketball positions (includes both existing and possible positions). COURTESY: Ayasdi Analytics and [10].

The Big Data Backlash.....Enter Meta-Analysis
Despite its successes, there is nevertheless a big data backlash. Ernest Davis and Gary Marcus [11] present us with nine reasons why big data are problematic. Some of these have been covered in the last section, while others suggest that there can be too much data. This is an interesting position, since it is common wisdom that more data always give you more resolution and insight. Insight and information can be obscured by noisy or irrelevant data. But even the most informative of datasets can yield misinformed analyses if the analyst is not thoughtful.

Of course, ever-bigger datasets by themselves do not give us the insights necessary to determine whether or not a generalized relationship is significant. The ultimate goal of data analysis should be to gain deep insights into whatever the data represent. While this does involve a degree of interpretive subjectivity, it also requires an intimate dialogue between analysis, theory, and simulation. Perhaps the latter is much more important, particularly in cases where the data are politically or socially sensitive. These considerations are missing from much contemporary big data analysis [12]. This vision goes beyond the conventional "statistical test on a single experiment" kind of experimental investigation, and leads us to meta-analysis.

The basic premise of a meta-analysis is to use a strategy of convergence/triangulation to converge upon results using a series of studies. The logic here involves using the power of consensus and statistical power to arrive at a solution. The problem is represented as a series of experiments with an effect size for each. For example, if I believe that eating oranges causes cancer, how should I arrive at a sound conclusion? One study with a very large effect size, or many studies with various effect sizes and experimental contexts. According to the meta-analysis view, the latter should be most informative. In the case of potential factors in myocardial infarction [13], significant results that all point in the same direction (with minimum effect size variability) lend the strongest support to a given hypothesis.

Example of a meta-analysis. COURTESY: [13].

The Problem with Deep Analysis
We can go even further down the rabbit hole of analysis, for better or for worse. However, this often leads to problems of interpretation, as deep analyses are essentially layered abstractions. In other words, they are higher-level abstractions dependent upon lower-level abstractions. This leads us to a representation of representations, which will be covered in an upcoming post. Here, I will propose and briefly explore two phenomena: significant pattern extraction and significant reconstructive mimesis.

One form of deep analysis involves significant pattern extraction. While the academic field of pattern recognition has made great strides [14], sometimes the collection of data (which involve pre-processing and personal bias) is flawed. Other times, it is the subjective interpretation of these data which are flawed. In either case, this results in the extraction patterns that make no sense that are then assigned significance. Worse yet, some of these patterns are also thought to be of great symbolic significance [15]. The Bible Code is one example of such pseudo-analysis. Patterns (in this case secret codes) are extracted from a database (a book), and then these data are probed for novel but coincidental pattern formation (codes formed by the first letter of every line of text). As this is usually interpreted as decryption (or deconvolution) of an intentionally placed message, significant pattern extraction is related to the deep, naive theories discussed in "Structure and Theory of Theories".

Congratulations! Your pattern recognition algorithm came up with a match. Although if it were a computer instead of a mind, it might do a more systematic job of rejecting it as a false positive. LESSON: the confirmatory criteria for a significant result needs to be rigorous.

But suppose that our conclusions are not guided by unconscious personal biases or ignorance. We might intentionally leverage biases in the service of parsimony (or making things simpler). Sometimes, the shortcuts we take in representing natural processes present difficulties in understanding what is really going on. This is a problem of significant reconstructive mimesis. In the case of molecular animations, this has been pointed out by Carl Zimmer [16] and PZ Myers [17] for molecular animations. In most molecular animations, processes occur smoothly (without error) and within full view of the human observer. Contrast this with the inherent noisiness and spatially-crowded environment of the cell, which is highly realistic but not very understandable. In such cases, we construct a model which consists of data, but that model is selective and the data is deliberately sparse (in this case smoothed). This is an example of a representation (the model) that informs an additional representation (the data). For purposes of simplicity, the model and data are somehow compressed to preserve signal and remove noise. And in the case of a digital image file (e.g. .jpg, .gif) such schemes work pretty well. But in other cases, the data are not well-known, and significant distortions are actually intentional. This is where big challenges arise in getting things right.

An multi-layered abstraction from a highly-complex multivariate dataset? Perhaps. COURTESY: Salvador Dali, Three Sphinxes of Bikini.

Conclusions
Data analysis is hard. But in the world of everyday science, we often forget how complex and difficult this endeavor is. Modern software packages have made the basic and well-established analysis techniques deceptively simple to employ. In moving to big data and multivariate datasets, however, we begin to face head-on the challenges of analysis. In some cases, highly effective techniques have simply not been developed yet. This will require creativity and empirical investigation, things we do not often associate with statistical analysis. It will also require a role for theory, and perhaps even the theory of theories.

As we can see from our last few examples, advanced data analysis can require conceptual modeling (or representations). And sometimes, we need to map between domains (from models to other, higher-order models) to make sense of a dataset. This, the most complex of analyses, can be considered representations of representations. Whether a particular representation of a representation is useful or not depends upon how much noiseless information can be extracted from the available data. Particularly robust high-level models can take very little data and provide us with a very reliable result. But this is an ideal situation, and often even the best models presented with large amounts of data can fail to given a reasonable answer. Representations of a representations also provide us with the opportunity to imbue an analysis with deep meaning. In a subsequent post, I will this out in more detail. For now, I leave you with this quote:
“An unsophisticated forecaster uses statistics as a drunken man uses lampposts — for support rather than for illumination.” Andrew Lang.

NOTES:
[1] Learn Statistics with Comic Books. CTRL Lab Notebook, April 14 (2011).

[2] Mooney, C.   The Science of Why We Don't Believe Science. Mother Jones, May/June (2011).

[3] Kosko, B.   Statistical Independence: What Scientific Idea Is Ready For Retirement. Edge Annual Question (2014).

[4] In order to separate signal from noise, we must first define noise. Noise is consistent with processes that occur at random, such as the null hypothesis or a coin flip. Using this framework, a significant result (or signal) is a result that deviates from random chance to some degree. For example, a p-value of 0.05 represents a 95% chance that the replicates observed could not have occurred due to chance. This is, of course, an incomplete account of the relationship between signal and noise. Models such as Signal Detection Theory (SDT) or data smoothing techniques can also be used to improve the signal-to-noise ratio.

[5] Nuzzo, R.   Scientific Method: Statistical Errors. Nature News and Comment, February 12 (2014).

[6] Fox, J.   Frequentist vs. Bayesian Statistics: resources to help you choose. Oikos blog, October 11 (2011).

[7] Gelman, A.   So-called Bayesian hypothesis testing is just as bad as regular hypothesis testing. Statistical Modeling, Causal Inference, and Social Science blog, April 2 (2011).

[8] For some concrete (and satirical) examples of how correlation does not equal causation, please see Tyler Vigen's Spurious Correlations blog.

[9] Voytek, B.   Big Data: what's it good for? Oscillatory Thoughts blog, January 30 (2014).

[10] Beckham, J.   Analytics Reveal 13 New Basketball Positions. Wired, April 30 (2012).

[11] Davis, E. and Marcus, G.   Eight (No, Nine!) Problems with Big Data. NYTimes Opinion, April 6 (2014).

[12] Leek, J.   Why big data is in trouble - they forgot applied statistics. Simply Statistics blog, May 7 (2014).

[13] Egger, M.   Bias in meta-analysis detected by a simple, graphical test. BMJ, 315 (1997).

[14] Jain, A.K., Duin, R.P.W., and Mao, J.   Statistical Pattern Recognition: a review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(1), 4-36 (2000).

It is interesting to note that the practice of statistical pattern recognition (training a statistical model with data to evaluate additional instances of data) has developed techniques and theories related to rigorously rejecting false positives and other spurious results.

[15] McCardle, G.     Pareidolia, or Why is Jesus on my Toast? Skeptoid blog, June 6 (2011).

[16] Zimmer, C.     Watch Proteins Do the Jitterbug. NYTimes, April 10 (2014).

[17] Myers, P.Z.   Molecular Machines! Pharyngula blog, September 3 (2006).

March 2, 2014

Fireside Science: Logical Fallacy vs. Logical Fallacy

This content is cross-posted to Fireside Science. To get the most out of this post, please review the following materials:

Alicea, B.   Informed Intuition > Pure Logic, Reason + No Information = Fallacy? Synthetic Daisies blog, January 4 (2014).

The peer-review committee for pure rationality. For more, please see [1].


Awhile back, I posted some critiques of and modifications to the conventional approach to logical fallacies [1] here on Synthetic Daisies. It seems as though every debate of the issues on the internet involves an accusation that one side is engaging in some sort of "fallacy". This is especially true of topics of broader societal relevance, where the notion of logical fallacies has become entangled with denialism [2] and epistemic closure [3].

Social Media argumentation, one person's take.

To recap (full version of the post here), I proposed that we replace six fallacies on the chart above and replace them with seven fallacies that are more inclusive of moral (e.g. emotional) and cultural biases. To me, the "Skeptic's Guide to the Universe" model feels like a 12-step program of rationality. It may help you think in a desirable way (e.g. pure rationality). However, pure rationality does not provide you with a means to place conditions on an objective argument. The triumph of logical rigor ultimately becomes a straight-jacket of the mind, reducing one's ability to think situationally.

Are the arbiters of deduction wrong on six counts?

Now it appears that I'm not alone in my concerns. Big Think now has a theme "The Fallacy Fallacy" on the fallacies of logical fallacies [4], with contributions from Alex Berezow, Julia Galef, Daniel Honan, and James Lawrence Powell


In this collection of essays and interviews, the overuse of logical fallacies itself is cited as a fallacy of composition, and provides better ways to construct arguments. These include several general observations related to the validity of reason itself. These transcend the popular "identify the fallacy" model.

One theme involves making the case for consensus through joint argumentation. Correct answers are not to be found via the most rigorous argument, but by exploring many complementary arguments, each with their own flaws.  

Another theme involves being mindful of cognitive biases such as confirmation bias or subconscious cultural preferences. Even when an argument is highly rigorous by the standards of logical consistency, they may still suffer from a lack of perspective. 

The third major theme involves the recognition that ignorance is a valid starting point [5] for many arguments. It is impossible to know everything about a topic, so any principled argument is bound to be incomplete. And the traditional fallacy model [6] is likely to make things worse.



NOTES:
[1] This is a list of 24 common logical fallacies, courtesy of Yourlogicalfallacyis.com (Jesse Richardson, Andy Smith, and Som Meadon). Also, most of these are individually found on Wikipedia with a more detailed explanation.

[2] Reinert, C.   Denialism vs. Skepticism. Institute for Ethics and Emerging Technologies blog, February 23 (2014).

[3] Cohen, P.   "Epistemic Closure"? Those are fighting words. NY Times Books, April 27 (2010).

[4] This is not a tautology! But it's not the same thing as the formal version of the fallacy fallacy (a.k.a. argumentum ad logicam).

[5] Contrast with: Argument from Ignorance. RationalWiki.

[6] A nice resource for better understanding all possible logical fallacies: The Fallacy-a-Day-Podcast. A fallacy a day, in readable and podcast form.

December 16, 2013

Fireside Science: Inspired by a visit to the Network's Frontier....

This post has been cross-posted to Fireside Science.


Recently, I attended the Network Frontiers Workshop at Northwestern University in Evanston, IL. This was a three-day session in which researchers engaged in network science from around the world gathered to present their work. They also came from many home disciplines, including computational biology, applied math and physics, economics and finance, neuroscience, and more.


The schedule (all researcher names and talk titles) can be found here. I was among one of the first presenters on the first day, presenting “From Switches to Convolution to Tangled Webs” [1], which involves network science from a evolutionary systems biology perspective.


One Field, Many Antecedents
For many people who have a passing familiarity with network science, it may not be clear as to how people from so many disciplines can come together around a single theme. Unlike more conventional (e.g. causal) approaches to science, network (or hairball) science is all about finding the interactions between the objects of analysis. Network science is the large-scale application of graph theory to complex systems and ever-bigger datasets. These data can come from social media platforms, high-throughput biological experiments, and observations of statistical mechanics. 

The visual definition of a scientific "hairball". This is not causal at all.....

25,000 foot View of Network Science
But what does a network science analysis look like? To illustrate, I will use an example familiar to many internet users. Think of a social network with many contacts. The network consists of nodes (e.g. friends) and edges (e.g. connections) [2]. Although there may be causal phenomena in the network (e.g. influence, transmission), the structure of the network is determined by correlative factors. If two individuals interact in some way, this increases the correlation between the nodes they represent. This gives us a web of connections in which the connectivity can range from random to highly-ordered, and the structure can range from homogeneous to heterogeneous.

Friend data from my Facebook account, represented as a sizable (N=64) heterogeneous network. COURTESY: Wolfram|Alpha Facebook app.

Continuing with the social network example, you may be familiar with the notion of “six degrees of separation” [3].  This describes one aspect (e.g. something that enables nth-order connectivity) of the structure inherent in complex networks. Again consider the social network: if there are preferences for who contacts whom, a randomly-connected network results. The path between any two individuals in such a network is generally high, as there are no reliable short-cuts. This path across the network is also known as the network diameter, and is an important feature of a network's topology.

Example of a social network. This example is homogeneous, but with highly-regular structure (e.g. non-random). 

Let us further assume that in the same network, there happen to be strong preferences for inter-node communication, which leads to changes in connectivity. In such cases, we get connectivity patterns that range from scale-free [4] to small-world [5]. In social networks, small-world networks have been implicated in the “six degrees” phenomenon, as the path between any two individuals is much shorter than in the random case. Scale-free and especially small-world networks have a heterogeneous structure, which can include local subnetworks (e.g. modules or communities) and small subpopulations of nodes with many more connections than other nodes (e.g. network hubs). Statistically, heterogeneity can be determined using a number of measures, including betweenness centrality and network diameter.

Example of a small-world network, in the scheme of things. 

Emerging Themes
While this example was made using a social network, the basic methodological and statistical approach can be applied to any system of strongly-interacting agents that can provide a correlation structure [6]. For example, high-throughput measurements of gene expression can be used to form a gene-gene interaction network. Genes that correlate with each other (above a pre-determined threshold) are consider connected in a first-order manner. The connections, while indirectly observed, can be statistically robust and validated via experimentation. And since all assayed genes (or the order of 103 genes) are likewise connected, second and third-order connections are also possible. The topology of a given gene-gene interaction network may be informative about the general effects of knockout experiments, environmental perturbations, and more [7].

This combination of exploratory and predictive power is just one reason why the network approach has been applied to many disciplines, and has even formed a discipline in and of itself [8]. At the Network Frontiers Workshop, the talks tended to coalesce around several themes that define potential future directions for this new field. These include:

A) general mechanisms: there are a number of mechanisms that allow for the network to adaptively change, stay the same in the face of pressure to change, or function in some way. These mechanisms include robustness, the identification of switches and oscillators, and the emergence of self-organized criticality among the interacting nodes. Papers representing this theme may be found in [9].

The anatomy of a forest fire's spread, from a network perspective.

B) nestedness, community detection, and clustering: Along with the concept of core-periphery organization, these properties may or may not exist in a heterogeneous network. But such techniques allow us to partition a network into subnetworks (modules) that may operate with a certain degree of independence. Papers representing this theme may be found in [10].

C) multilevel networks: even in the case of social networks, each "node" can represent a number of parallel processes. For example, while a single organism possesses both a genotype and a phenotype, the correlational structure for genotypic and phenotypic interactions may not always be identical. To solve this problem, a bipartite (two independent) graph structure may be used to represent different properties of the population of interest. While this is just a simple example, multilevel networks have been used creatively to attack a number of problems [11].

D) cascades, contagions: the diffusion of information in a network can be described in a number of ways. While the common metaphor of "spreading" may be sufficient in homogeneous networks, it may be insufficient to describe more complex processes. Cascades occur when transmission is sustained beyond first-order interactions. In a social network, messages that gets passed to a friend of a friend of a friend (e.g. third-order interactions) illustrate the potential of the network topology to enable cascade. Papers representing this theme may be found in [12].

E) hybrid models: as my talk demonstrates, the power and potential of complex networks can be extended to other models. For example, the theoretical "nodes" in a complex network can be represented as dynamic entities. Aside from real-world data, this can be achieved using point processes, genetic algorithms, or cellular automata. One theme I detected in some of the talks was the potential for a game-theoretic approach, while others involved using Google searches and social media activity to predict markets and disease outbreaks [13].

Here is a map of connectivity across three social media platforms: Facebook, Twitter, and Mashable. COURTESY: Figure 13 in [14].

NOTES:
[1] Here is the abstract and presentation. The talk centered around a convolution architecture, my term for a small-scale physical flow diagram that can be evolved to yield not-so-efficient (e.g. sub-optimal) biological processes. These architectures can be embedded into large, more complex networks as subnetworks (in a manner analogous to functional modules in gene-gene interaction or gene regulatory networks).

One person at the conference noted that this had strong parallels with the book “Plausibility of Life” (excerpts here) by Marc Kirschner and John Gerhart. Indeed, this book served as inspiration for the original paper and current talk.

[2] In practice, "nodes" can represent anything discrete, from people to cities to genes and proteins. For an example from brain science, please see: Stanley, M.L., Moussa, M.N., Paolini, B.M., Lyday, R.G., Burdette, J.H. and Laurienti, P.J.   Defining nodes in complex brain networks. Frontiers in Computational Neuroscience, doi:10.3389/fncom.2013.00169 (2013).

[3] the "six degrees" idea is based on an experiment conducted by Stanley Milgram, in which he sent out and tracked the progression of a series of chain letters through the US Mail system (a social network). 

The potential power of this phenomenon (the opportunity to identify and exploit weak ties in a network) was advanced by the sociologist Mark Granovetter: Granovetter, M.   The Strength of Weak Ties: A Network Theory Revisited. Sociological Theory, 1, 201–233 (1983).

The small-world network topology (the Watts-Strogatz model), which embodies the "six degrees" principle, was proposed in the following paper: Watts, D. J. and Strogatz, S. H.   Collective dynamics of 'small-world' networks. Nature, 393(6684), 440–442 (1998).

[4] Scale-free networks can be defined as a network with no characteristic number of connections across all nodes. Connectivity tends to scale with growth in the number of nodes and/or edges. Whereas connectivity in a random network can be characterized using a Gaussian (e.g. normal) distribution, connectivity in a scale-free network can be characterized using a Power Law (e.g. exponential) distribution.

[5] Small-world networks are defined by their hierarchical (e.g. strongly heterogeneous) structure and a short path length across the network. This is a special case of the more general scale-free pattern, and can be characterized with a strong power law (e.g. the distribution has a thicker tail). Because any one node can reach any other node in a relatively small number of steps, there are a number of organizational consequences to this type of configuration.

[6] Here are two foundational papers on network science [a, b] enlightening primers on complexity and network science [c, d]:
[a] Albert, R. and Barabasi, A-L.   Statistical mechanics of complex networks. Reviews in Modern Physics, 74, 47–97 (2002).

[b] Newman, M.E.J.   The structure and function of complex networks. SIAM Review, 45, 167–256 (2003).

[c] Shalizi, C.   Community Discovery Methods for Complex Networks. Cosma Shalizi's Notebooks - Center for the Study of Complex Systems, July 12 (2013).

[d] Voytek, B.   Non-linear Systems. Oscillatory Thoughts blog, June 28 (2013).

[7] For an example, please see: Cornelius, S.P., Kath, W.L., and Motter, A.E.   Controlling complex networks with compensatory perturbations. arXiv:1105.3726 (2011).

[8] Guimera, R., Uzzi, B., Spiro, J., and Amaral, L.A.N   Team Assembly Mechanisms Determine Collaboration Network Structure and Team Performance. Science, 308, 697 (2005).

[9] References for general mechanisms (e.g. switches and oscillators):
[a] Taylor, D., Fertig, E.J., and Restrepo, J.G.   Dynamics in hybrid complex systems of switches and oscillators. Chaos, 23, 033142 (2013).

[b] Malamud, B.D., Morein, G., and Turcotte, D.L.   Forest Fires: an example of self-organized critical behavior. Science, 281, 1840-1842 (1998).

[c] Ellens, W. and Kooij, R.E.   Graph measures and network robustness. arXiv: 1311.5064 (2013).

[d] Francis, M.R. and Fertig, E.J.   Quantifying the dynamics of coupled networks of switches and oscillators. PLoS One, 7(1), e29497 (2012).

[10] References for clustering [a], community detection [b-e], core-periphery structure detection [f], and nestedness [g]:
[a] Malik, N. and Mucha, P.J.   Role of social environment and social clustering in spread of opinions in co-evolving networks. Chaos, 23, 043123 (2013).

[b] Rosvall, M. and Bergstrom, C.T.   Maps of random walks on complex networks reveal community structure. PNAS, 105(4), 1118-1123 (2008).


* the image above was taken from Figure 3 of [a]. In [a], an information-theoretic approach to discovering network communities (or subgroups) is introduced.

[c] Colizza, V., Pastor-Satorras, R. and Vespignani, A.   Reaction–diffusion processes and metapopulation models in heterogeneous networks. Nature Physics, 3, 276-282 (2007).

[d] Bassett, D.S., Porter, M.A., Wymbs, N.F., Grafton, S.T., Carlson, J.M., and Mucha, P.J.   Robust detection of dynamic community structure in networks. Chaos, 23, 013142 (2013).

* the authors characterize the dynamic properties of temporal networks using methods such as optimization variance and randomization variance.

[e] Nishikawa, T. and Motter, A.E.   Discovering network structure beyond communities, Scientific Reports, 1, 151 (2011).

[f] Bassett, D.S., Wymbs, N.F., Rombach, M.P., Porter, M.A., Mucha, P.J., and Grafton,
S.T.   Task-Based Core-Periphery Organization of Human Brain Dynamics. PLoS Computational Biology, 9(9), e1003171 (2013).

* a good exampkle of how core-periphery structure is extracted from brain networks constructed from fMRI data.

[g] Staniczenko, P.P.A., Kopp, J.C., and Allesina, S.   The ghost of nestedness on ecological networks. Nature Communications, doi:10.1038/ncomms2422 (2012).

[11] References for multilevel networks:
[a] Szell, M., Lambiotte, R., Thurner, S.   Multirelational organization of large-scale social networks in an online world. PNAS, doi/10.1073/pnas.1004008107 (2010).

[b] Ahn, Y-Y., Bagrow, J.P., and Lehmann, S.   Link communities reveal multiscale complexity in networks. Nature, 466, 761-764 (2010).

[12] References for cascades and contagions:
[a] Centola, D.   The Spread of Behavior in an Online Social Network Experiment. Science, 329, 1194-1197 (2010).

[b] Brummitt, C.D., D’Souza, R.M., and Leicht, E.A.   Suppressing cascades of load in interdependent networks. PNAS, doi:10.1073/pnas.1110586109 (2011).

[c] Brockmann, D. and Helbing, D.   The Hidden Geometry of Complex, Network-Driven Contagion Phenomena. Science, 342(6164), 1337-1342 (2013).

[d] Glasserman, P. and Young, H.P.   How Likely is Contagion in Financial Networks? Oxford University Department of Economics Discussion Papers, #642 (2013).

[13] Reference for hybrid networks and other themes, including network evolution [a,b] and the use of big data in network analysis [c,d]:
[a] Pang, T.Y. and Maslov, S.   Universal distribution of component frequencies in biological and technological systems. PNAS, doi:10.1073/pnas.1217795110 (2012).

[b] Bassett, D.S., Wymbs, N.F., Porter, M.A., Mucha, P.J., and Grafton, S.T.   Cross-Linked Structure of Network Evolution. arXiv: 1306.5479 (2013).

[c] Ginsberg, J., Mohebbi, M.H., Patel, R.S., Brammer, L., Smolinski, M.S., and Brilliant, L.   Detecting influenza epidemics using search engine query data. Nature, 457, 1012–1014 (2008).

[d] Michel, J-B., Shen, Y.K., Aiden, A.P., Veres, A., Gray, M.K., Google Books Team, Pickett, J.P., Hoiberg, D., Clancy, D., Norvig, P., Orwant, J., Pinker, S., Nowak, M.A., Aiden, E.L.   Quantitative Analysis of Culture Using Millions of Digitized Books. Science, 331(6014), 176-182 (2011).

[14] Ferrara, E.   A large-scale community structure analysis in Facebook. EPJ Data Science, 1:9 (2012).

Printfriendly