Showing posts with label experimental-replication. Show all posts
Showing posts with label experimental-replication. Show all posts

September 14, 2015

Hodgepodge Blogpost, September 2015

Welcome to the blogging hodgepodge for this month. I wanted to clear up by reading queue, and present some of these ideas and articles in an entertaining way. The topics include: modeling, significant results, and hidden variables (but perhaps not discussed in a conventional manner). As a bonus, we get career advice for scientific researchers and relevant discussion.

Mutant phenptypes from the Fukushima area of Japan. COURTESY: National Geographic.


Flawed Models Cannot Be Made Idealistic

"Essentially, all models are wrong, but some are useful" -- George Box. What makes for a bad model? Poor assumptions, oversimplication/vagueness, or underfitting with respect to available data? These articles address some of these issues, with particular relevance to societal consequences.

Kirchner, L.   When Big Data Becomes Bad. ProPublica, September 2 (2015).

O'Neil, C.   Big Data, Disparate Impact, and the Neoliberal Mindset. Mathbabe blog, September 7 (2015).

Schuster, P.   Models: From Exploration to Prediction -- Bad Reputation of Modeling in Some Disciplines Results from Nebulous Goals. Complexity, doi:10.1002/cplx.21729 (2015).

Rickert, J.   How do you know if your model is going to work? Part 2: Intraining set measures. R-bloggers, September 8 (2015).


Once upon a time, this was a viable model of how nature worked. COURTESY: Geocentric Model, Redorbit.


The Real World is Complex, Idealized Methods Notwithstanding

The debate over replicability in Psychology (and by extension sciences that are not particle physics) rages on. This month, a shot was fired from the "Psychology is not very replicable" camp. The Open Science Collaboration published a paper in Science showing that many replications of experiments fail to reproduce the same levels of statistical significance and power as the original studies.

Critics have blamed this lack of replicability on a number of culprits, including shortcomings of the NHST approach itself. Two potential culprits I have pointed to previously include complexity and cultural context, the latter which we will return to in a bit.



What explains these replicated results? . COURTESY: Figure 1, Science, 349, doi:10.1126/science.aac4716 (2015) AND Loria, TechInsider.

Open Science Collaboration.   Estimating the reproducibility of psychological science. Science, doi:10.1126/science.aac4716 (2015).

Loria, K.   Everything that's wrong with psychology studies in 2 simple charts. TechInsider, August 28 (2015).

Barrett, L.F.   Psychology is not in Crisis. NYTimes Opinion, September 1 (2015).


The Unreasonable Effectiveness of Cultural Context*

* a play on: Wigner, E.   The Unreasonable Effectiveness of Mathematics in the Natural Sciences.


Vanderbilt, T.   Why Futurism Has a Cultural Blindspot. We predicted cell phones, but not women in the workplace. Nautil.us blog, September 10 (2015).

* the latest critique of futurism, this time from a sociological perspective.



Yau, N.   Bourdieu’s Food Space chart, from fast food to French Laundry. Flowing Data blog, June
21 (2012).

* our contemporary Economic World, according to Pierre Bourdieu (as told by Leigh Wells).


Career Advice (Not Avarice):

Hossenfelder, S.   How to publish your first scientific paper. Backreaction blog, September 11 (2015).

* this blog post not only provides advice on how to get started as a published researcher, but also gives advice on how to formulate research ideas and structure manuscripts that will garner the interest of editors and reviewers.

Curry, S.   Peer review, preprints and the speed of science. Guardian, September 7 (2015).

* yet another article in favor of the open-science movement, in this case advocating for mechanisms (e.g. preprint servers, open peer review) that have the potential to speed up and otherwise improve the research enterprise.

McDonnell, J.J.   Creating a Research Brand. Science, 349, 758 (2015).

This author uses a marketing metaphor to help imporive the efficiency of a researcher's efforts. The advice bolis down to the following:

* promote results, publications, and lectures all around a central theme.

* find the right breadth of research. This should be greater than a hyper-specialized topic, but narrow enough to constitute a unique niche.

September 17, 2014

Heuristic Haystacks and the Messy Lesson

As an admitted and self-styled parsimony skeptic, I was interested to see a discussion in the blogosphere on the seductive allure of simple explanations [1]. This was in the context of economic policy and decision-making, with Paul Krugman even offering an H.L. Mencken quote: "For every complex problem there is an answer that is clear, simple, and wrong" [2]. Yet while parsimony was never brought up, I suspect that hypotheses and arguments related to the efficient markets hypothesis were always somewhat in mind.


There are, of course, broader parallels between seductive simplicity and parsimony. As I have pointed out before, I find parsimony to be a overly-seductive null model [3]. The simplest explanation often leads us not to the truth, but to what is most conceptually consistent. In some cases (where theory is well-established) this works out well. Intuition in support of serendipity and serendipity in support of discovery is an unassuming (and often underplayed) pillar of science [4]. However, in cases where our intuitions get in the way of objective analysis, this becomes problematic. And this seeming exception is actually quite common. In a related manner, this brings up an interesting problem of the relationship between parsimony as a decision-making criterion and the epistomology of a scientific phenomenon.

An appalling lack of faith in both Occam's and Einstein's worldviews. More horrifying details in my Ignite! talk on the topic.

This relationship, or more accurately inconsistency, is due to argumentatively-influenced judgments on a naturalistic search space. Even in children, it is observed that argumentation is rife with confirmation bias and logically arguing to absurd positions [5]. While argumentation allows us to build hypotheses, it also gets us stuck in a conceptual minimum (my own ad-hoc phrase). In a previous post, I pointed to recent work on how belief systems and associated systems of argumentation can shape our perception of reality. But, of course, this cannot will the natural world to our liking. In fact, it often serves to muddy the conceptual and theoretical waters [6]. Therefore, you often have a conceptual gap unrelated to problem incompleteness which we will flesh out in the rest of this post.

The first point to be made here is that such an inconsistency introduces two biases that shape how we think about the simplest explanation, and more generally about what is optimal. First of all, can we even find the true simplest explanation? Perhaps the simplest possible statement that can be constructed cannot capture the true complexity of a given situation. This is particularly true when there are competing dimensions (or layers or levels) of complexity. Secondly, and particularly in the face of complexity, simplicity can often be a foil to deep understanding. Unfortunately, this is often conceptualized of and practiced upon in a destructive way, favoring simple and homogeneous mental models over more subtle ones.

How to dream of complex sheep....

In the parlance of decision-making theory, parsimony is consistent with the notion of good-enough heuristics. In the work of Gigerenzer [7], such heuristics are claimed to be nearly optimal when compared to formal analysis of a problem. This can also be seen with statistical prediction rules that outperform human judgments in a number of everyday contexts [8]. But is this a statement of problem "wickedness", or a statement of superiority with respect to human cognition? When compared to problems that require needle in a haystack criteria, fast and frugal heuristics (and hence parsimony) is severely lacking.

So complexity introduces a secondary bias at best and serves as a severe limitation to achieving parsimony at worst. One might expect that experimentally verifying a prediction made in conjunction with Occam's Razor requires finding an exact analytical solution. Finding this proverbial "needle in a haystack" requires both a multi-criterion, algorithmically-friendly heuristic solution in addition to a formal strategy that often defies intuition. Seemingly, the simple solution cannot keep up.

I found it! It was quick, but I was also quite lucky,

NOTES:
[1] The Simplicity Paradox. Stumbling and Mumbling blog, September 9 (2014) AND Krugman, P.   Simply Unacceptable. The Conscience of a Liberal blog, September 5 (2014).

[2] This is not to equate parsimony with methodological snake oil -- in fact, I am arguing quite the opposite. But I am merely pointing out that parsimony is an incomplete hypothesis for acquiring knowledge.

[3] For more, please see this Synthetic Daisies post: Alicea, B.   Argument from Non-Optimality: what does it mean to be optimal? Synthetic Daisies blog, July 28 (2013).

[4] Kantorovich, A.   Scientific Discovery: Logic and Tinkering. SUNY Press, Albany (1993).

[5] I say "even" in children even though the latter (logically arguing to absurd conclusions) is often expected from children. But we see these things in adults as well, and such is the point of argumentation theory. For more, please see: Mercier, H.   Reasoning Serves Argumentation in Children. Cognitive Development, 26(3), 177–191 (2011).

[6] Wolchover, N.   Is Nature Unnatural? Quanta Magazine, May 24 (2013).

[7] While there are likely other (and perhaps better) examples, I am using a reference cited in [1]: Gigerenzer, G.   Bounded and Rational. In "Contemporary Debates in Cognitive Science", R.J. Stainton eds. Blackwell, Oxford, UK (2006).

[8] lukeprog   Statistical Prediction Rules Out-Perform Expert Human Judgments. LessWrong blog, January 18 (2011).

June 21, 2014

Fireside Science: The Representation of Representations

This content is being cross-posted to Fireside Science, and is the third in a three-part series on the "science of science".


This is the final in a series of posts on the science of science and analysis. In past posts, we have covered theory and analysis. However, there is a third component of scientific inquiry: representation. So this post is about the representation of representations, and how representations shape science in more ways than the casual observer might believe.

The three-pronged model of science (theory, experiment, simulation). Image is adapted from Fermi Lab Today Newsletter, April 27 (2012).

For the uninitiated, science is mostly analysis and data collection with theory being a supplement at best and necessary evil at worst. Ideally, modern science rests on three pillars: experiment, theory, and simulation. For these same uninitiated, the representation of scientific problems is a mystery. But in fact, it has been the most important motivation for much of the scientific results we celebrate today. Interestingly, the field of computer science relies heavily on representation, but this concern generally does not carry over into the empirical sciences.

Ideagram (e.g. representation) of complex problem solving. Embedded are a series of Hypotheses and the processes that link them together. COURTESY: Diagram from [1].

Problem Representation
So exactly what is scientific problem representation? In short, it is the basis for designing experiments and conceiving of models. It is the sieve through which scientific inquiry flows, restricting the typical "question to be asked" to the most plausible or fruitful avenues. It is often the basis of consensus and assumptions. On the other hand, representation is quite a bit more subjective than people typically would like their scientific inquiry to be. Yet this subjectivity need not lead to an endless debate about the validity of one point of view versus another. There are heuristics one can use to ensure that problems are represented in a consistent and non-leading way.

3-D Chess: a high-dimensional representation of warfare and strategy.

Models that Converge
Convergent models speaks to something I alluded to in "Structure and Theory of Theories" when I discussed the theoretical landscape of different academic fields. The first way is whether or not allied sciences or models point in the same direction. To do this, I will use a semi-hypothetical example. The hypothetical case is to consider three models (A, B, and C) of the same phenomenon. Each of these models make different assumptions and includes different factors, but should at least be consistent with each other. One real-world example of this is the use of gene trees (phylogenies) and species trees (phylogenies) to understand evolution in a lineage [2]. In this case, each model uses the same taxa (evolutionary scenario), but includes incongruent data. While there are a host of empirical reasons why these two models can exhibit incongruence [3], models that are as representationally complete as possible might resolve these issues.

Orientation of Causality
The second way is to ensure that the one's representation gets the source of causality right. For problems that are not well-posed or poorly characterized, this can be an issue. Let's take Type III errors [4] as an example of this. In hypothesis testing, type III errors involve using the wrong explanation for a significant result. In layman's terms, this is getting the right answer for the wrong reasons. Even more than in the  case of type I and II errors, focusing on the correct problem representation plays a critical role in resolving potential type III errors.

Yet problem representation does not always help resolve these types of errors. Skeptical interpretation of the data can also be useful [5]. To demonstrate this, let us turn to the over-hyped area of epigenetics and its larger place in evolutionary theory. Clearly, epigenetics plays some role in the evolution of life, but is not deeply established in terms of models and theory. Because of this representational ambiguity, some interpretations play a trick. In a conceptual representation that embodies this trick, scarcely-understood high-level phenomena such as epigenetics will usurp the role of related phenomena such as genetic diversity and population processes. When the thing in your representation is not well-defined or quite popular (e.g. epigenetics), it can take on a causal life of its own. Posing the problem in this way allows us to obscure known dependencies between genes, genetic regulation, and the environment without proving exceptions to these established relationships.

Popularity is Not Sufficiency
The third way is to understand that popular conceptions do not translate into representational sufficiency. In logical deduction, it is often pointed out that necessity does not equal sufficiency. But as with the epigenetics example, it also holds that popularity cannot make something sufficient in and of itself. In my opinion, this is one of the problems with using narrative structures in the communication of science: sometimes an appealing narrative does more to obscure scientific findings than it does in making things accessible to lay people.

Fortunately, this can be shown by looking at media coverage of any big news story. The CNN plane coverage [6] shows this quite clearly: coverage of rampant speculation and conspiracy theory was a way to emphasize an increasingly popular story. In such cases, speculation is the order of the day, while thoughtful analysis gets pushed aside. But is this simply a sin of the uninitiated, or can we see parallels of this in science? Most certainly, there is a problem with recognizing the difference between "popular" science and worthwhile science [7]. There is also precedence from the way in which certain studies or areas of study are hyped. Some in the scientific community [8] have argued that Nature's hype of the ENCODE project [9] results fell into this category.

One example of a mesofact: ratings for the TV show The Simpsons over the course of several hundred episodes. COURTESY: Statistical analysis in [10].

Mesofacts
Related to these points is the explicit relationship between data and problem representation. In some ways, this brings us back to a computational view of science, where data do not make sense unless it is viewed in the context of a data structure. But sometimes the factual aspect of data varies over time in a way that obscures our mental models, and in turn obscures problem representation.

To make this explicit, Sam Arbesman has coined the term "mesofact" [11]. A mesofact is knowledge that changes slowly over time given new data. Populations of specific places (e.g. Minneapolis, Bolivia, Africa) has changed in both absolute and relative terms over the past 50 years. But when problems and experimental designs are formulated assuming that facts related to these data (e.g. rank of cities by population) do not change over time, we can get the analysis fundamentally wrong.

This may seem like a trivial example. However, mesofacts have relevance to a host of problems in science, from experimental replication to inferring the proper order of causation. The problem comes down to an interaction between data's natural variance (variables) and the constructs used to represent our variables (facts). When the data exhibit variance against an unchanging mean, it is much easier to use this variable as a stand-in for facts. But when this is not true, scientifically-rigorous facts are much harder to come by. Instead of getting into an endless discussion about the nature of facts, we can instead look to how facts and problem representation might help us tease out the more metaphysical aspects of experimentation.

Applying Problem Representation to Experimental Manipulation
When we do experiments, how do we know what our experimental manipulations really mean? The question itself seems self-evident, but perhaps it is worth exploring. Suppose that you wanted to explore the causes of mental illness, but did not have the benefits of modern brain science as a guide. In defining mental illness itself, you might work from a behavioral diagnosis. But the mechanisms would still be a mystery. Is it a supernatural mechanism (e.g. demons) [12], an ultimate form of causation (reductionism), or a global but hard-to-see mechanism (e.g. quantum something) [13]? An experiment done the same way but assuming three different architectures could conceivably yield statistical significance for all of them.

In this case, a critical assessment of problem representation might be able to resolve this ambiguity. This is something that as modelers and approximators, computational scientists deal with all of the time. Yet it is also an implicit (and perhaps even more fundamental) component of experimental science. For most of the scientific method's history, we have gotten around this fundamental concern by relying on reductionism. But in doing so, this restricts us to doing highly-focused science without appealing to the big picture. In a sense, we are blinded by science by doing science.

Focusing on problem representation allows us a way out of this. Not only does it allow us to break free from the straightjacket of reductionism, but also allows us to address the problem of experimental replication more directly. As has been discussed in many other venues [14], the lack of an ability to replicate experiments has plagued both Psychological and Medical research. But it is in these areas which representation is most important, primarily because it is hard to get right. Even in cases where the causal mechanism is known, the underlying components and the amount of variance they explain can vary substantially from experiment to experiment.

Theoretical Shorthand as Representation
Problem representation also allows us to make theoretical statements using mathematical shorthand. In this case, we face the same problem as the empiricist: are we focusing on the right variables? More to the point, are these variables fundamental or superficial? To flesh this out, I will discuss two examples of theoretical shorthand, and whether or not they might be concentrating on the deepest (and most generalizable) constructs possible.

The first example comes from Hamilton's rule, derived by the behavioral ecologist W.D. Hamilton [15]. Hamilton's rule describes altruistic behavior in terms of kin selection. The rule is a simple linear equation that assumes adaptive outcomes will be optimal ones. In terms of a representation, these properties provide a sort of elegance that makes it very popular.


In this short representation, an individual's relatedness to a conspecific contributes more to their behavioral motivation to help that individual than a typical trade-off between costs and benefits. Thus, a closely-related conspecific (e.g. a brother) will invest more into a social relationship with their kin than with non-kin. In general, they will take more personal risks in doing so. While more math is used to support the logic of this statement [15], this inequality is often treated as a widely applicable theoretical statement. However, some observers [16] have found the parsimony of this representation to be both too incomplete and intellectually unsatisfying. And indeed, sometimes an over-simplistic model does not deal with exceptions well.

The second example comes from Thomas Piketty's work. Piketty, economist and author of "Capital in the 21rst Century" [17], has proposed something he calls the "First Law" which explains how income inequality relates to economic growth. The formulation, also a simple inequality, characterizes the relationship between economic growth, inherited wealth, and income inequality within a society.


In this equally short representation, inequality is driven by the relative dominance of two factors: inherited wealth and economic growth. When growth is very low, and inherited wealth exists at a nominal level, inequality persists and dampens economic mobility. In Piketty's book, other equations and a good amount of empirical investigation is used to support this statement. Yet, despite its simplicity, it has held up (so far) to the scrutiny of peer review [18]. In this case, representation through variables that generalize greatly but do not handle exceptional behavior well produce a highly-predictive model. On the other hand, this form of representation also makes it hard to distinguish between a highly unequal post-industrial society and a feudal, agrarian one.

Final Thoughts
I hope to have shown you that representation is an underappreciated component of doing and understanding science. While the scientific method is our best strategy for discovering new knowledge about the natural world, it is not without its burden of conceptual complexity. In the theory of theories, we learned that formal theories are based on both deep reasoning and are (by necessity) often incomplete. In the analysis of analyses, we learned that the data are not absolute. Much reflection and analytical detail must be taken to ensure that an analysis represents meaningful facets of reality. And in this post, these loose ends were tied together in the form of problem representation. While an underappreciated aspect of practicing science, representing problems in the right way is essential for separating out science from pseudoscience, reality from myth, and proper inference from hopeful inference.

NOTES:
[1] Eldrett, G.   The art of complex problem-solving. MediaExplored blog, July 10 (2010).

[2] Nichols, R.   Gene trees and species trees are not the same. Trends in Ecology and Evolution, 16(7), 358-364 (2001).

[3] Gene trees and species trees can be incongruent for many reasons. Nature Knowledge Project (2012).

[4] Schwartz, S. and Carpenter, K.M.   The right answer for the wrong question: consequences of type III error for public health research. American Journal of Public Health, 89(8), 1175–1180 (1999).

[5] It is important here to distinguish between careful skepticism and contrarian skepticism. In addition, skeptical analysis is not always compatible with the scientific method.

For more, please see: Myers, P.Z.   The difference between skeptical thinking and scientific thinking. Pharyngula blog, June 18 (2014) AND Hugin   The difference between "skepticism" and "critical thinking"? RationalSkepticism.org, May 19 (2010).

[6] Abbruzzese, J.   Why CNN is obsessed with Flight 370: "The Audience has Spoken". Mashable, May 9 (2014).

[7] Biba, E.   Why the government should fund unpopular science. Popular Science, October 4 (2013).

[8] Here are just a few examples of the pushback against the ENCODE hype:


a) Mount, S.   ENCODE: Data, Junk and Hype. On Genetics blog, September 8 (2012).

b) Boyle, R.   The Drama Over Project Encode, And Why Big Science And Small Science Are Different. Popular Science, February 25 (2013).

c) Moran, L.A.   How does Nature deal with the ENCODE publicity hype that it created? Sandwalk blog, May 9 (2014).

[9] For an example of the nature of this hype, please see: The Story of You: ENCODE and the human genome. Nature Video, YouTube, September 10 (2012).

[10] Fernihough, A.   Kalkalash! Pinpointing the Moments “The Simpsons” became less Cromulent. DiffusePrior blog, April 30 (2013).

[11] Arbesman, S.   Warning: your reality is out of date. Boston Globe, February 28 (2010). Also see the following website: http://www.mesofacts.org/

[12] Surprisingly, this is a contemporary phenomenon: Irmak, M.K.   Schizophrenia or Possession? Journal of Religion and Health, 53, 773-777 (2014). For a thorough critique, please see: Coyne, J.   Academic journal suggests that schizophrenia may be caused by demons. Why Evolution is True blog, June 10 (2014).

[13] This is an approach favored by Deepak Chopra. He borrows the rather obscure idea of "nonlocality" (yes, basically a wormhole in spacetime) to explain higher levels of conscious awareness with states of brain activity.

[14] Three (divergent) takes on this:

a) Unreliable Research: trouble at the lab. Economist, October 19 (2013).

b) Ioannidis, J.P.A.   Why Most Published Research Findings Are False. PLoS Med 2(8): e124 (2005).

c) Alicea, B.   The Inefficiency (and Information Content) of Scientific Discovery. Synthetic Daisies blog, November 19 (2013).

[15] Hamilton, W. D.   The Genetical Evolution of Social Behavior. Journal of Theoretical Biology, 7(1), 1–16 (1964). See also: Brembs, B.   Hamilton's Theory. Encyclopedia of Genetics.

[16] Goodnight, C.   Why I Don’t like Kin Selection. Evolution in Structured Populations blog, April 23 (2014).

[17] Piketty, T.   Capital in the 21st Century. Belknap Press (2014). See also: Galbraith, J.K.   Unpacking the First Fundamental Law. Economist's View blog, May 25 (2014).

[18] DeLong, B.   Trying, yet again, to communicate the arithmetic scaffolding of Piketty's "capital in the Twenty-First Century". Washington Center for Equitable Growth blog, June 5 (2014).

May 12, 2014

Fireside Science: The Analysis of Analyses

This material is cross-posted to Fireside Science. This is part of a continuing series on the science of science (or meta-science, if you prefer). The last post was about the structure and theory of theories.


In this post, I will discuss the role of data analysis and interpretation. Why do we need data, as opposed to simply observing the world or making up stories? The simple answer: it gives us a systematic accounting of the world in general and experimental manipulations in particular. As opposed to the apparition on a piece of toast, it provides a systematic accounting of the natural world independent of our sensory and conceptual biases. But as we saw in the theory of theories post, and as we will see in this post, it takes a lot of hard work and thoughtfulness. What we end up with is an analysis of analyses.

Data take many forms, so approach analysis with caution. COURTESY: [1].

Introduction
What exactly is data, anyways? We hear a lot about it, but rarely stop to consider why it is so potentially powerful. Data are both an abstraction of and incomplete sampling (approximation) of the real world. While the data are not absolute (e.g. you can always have more data or more completely sample the world), the data provide a means of generalization that is partially free from stereotyping. And as we can see in the cartoon above, not all data that influence our hypothesis can even be measured. Some of it is beyond the scope of our current focus and technology (e.g hidden variables), while some of it consists of interactions between variables.

In the context of the theory of theories, data has the same advantage over anecdote that deep, informed theories have over naive theories. In the context of the analysis of analyses, data does not speak for itself. To conduct a successful analysis of analysis, it is important to be both interpretive and objective. Finding the optimal balance between each of these gives us an opportunity to reason more clearly and completely. If this causes some people to lose their view of data as infallible, then so be it. Sometimes the data fails us, and other times we fail ourselves.

When it comes to interpreting data, the social psychologist Jon Haidt suggests that "we think we are scientists, but we are actually laywers" [2]. But I would argue this is where the difference between the untrained eyes sharing Infographics and the truly informed acts of analysis and data interpretation becomes important. The latter is an example of a meta-meta-analysis, or a true analysis of analyses.

The implications of Infographics are clear (or are they?) COURTESY: Heatmap, xkcd.

NHST: the incomplete analysis?
I will begin our discussion with a current hot topic in the field of analysis. It involves interpreting statistical "significance" using an approach called Null Hypothesis Statistical Testing (or NHST). If you have even done a t-test or ANOVA, you have used this approach. The current discussion about the scientific replication crisis is tied to the use (and perhaps overuse) of these types of tests. The basic criticism involves the inability of NHST statistics to conduct multiple tests properly and properly deal with experimental replication.

Example of the NHST and its implications. COURTESY: UC Davis StatWiki.

This has even led scientists such as John Ioannidis to demonstrate why "most significant results are wrong". But perhaps this is just to make a rhetorical point. The truth is, our data are inherently noisy. Too many assumptions/biases go into collecting most datasets, all for data which has too little known structure. Not only are our data noisy, but in some cases may also possess hidden structure which violates the core assumptions of many statistical tests [3]. Some people have rashly (and boldly) proposed that this points to flaws in the entire scientific enterprise. But, like most things, this does not take into account the nature of the empirical enterprise and reification of the word significance.

A bimodal (e.g. non-normal) distribution, being admonished by its unimodal brethren. Just one case in which the NHST might fail us.

The main problem with the NHST is that it relies upon distinguishing signal from noise [4], but not always in the broader context of effects size or statistical power. In a Nature News correspondence [5], Regina Nuzzo discusses the shortcomings of the NHST approach and tests of statistical significance (e.g. p-values). Historical context of the so-called frequentist approach [6] is provided, and its connection to assessing the validity of experimental replications are discussed. One possible solution is the use of Bayesian techniques [7] to assess something called statistical power. The Bayesian approach allows one to use a prior distribution (or historical conditioning) to better assess the meaningfulness of one's statistically significant result. But the construction of priors relies on the existence of reliable data. If these data do not exist for some reason, we are back to square one.

Big Data and its Discontents
Another challenge to conventional analysis involves the rise of so-called big data. Big data is the collection and analysis of very large datasets, which come from sources such as high-throughput biology experiments, computational social science, open-data repositories, and sensor networks. Considering their size, big data analyses should allow for good power and ability to distinguish signal from noise. Yet due to their structure, we are often required to rely upon correlative analyses. While correlation is equated with relational information, it (as it always has) does not equate to causation [8]. Innovations in machine learning and other data modeling techniques can sometimes overcome this limitation, but correlative analyses are still the easiest way to deal with these data.

IBM's Watson: powered by large databases and correlative inference. Sometimes this cognitive heuristic works well, sometimes not so much.

Given a large enough collection of variables with a large number of observations, correlations can lead to accurate generalizations about the world [9]. The large number of variables are needed to extract relationships, while the large number of observations are needed to understand the true variance. This can be a problem where subtle, higher-order relationships (e.g. feedbacks, time-dependent saturations) exist or when the variance is not uniform with respect to the mean (e.g. bimodal distributions).

Complex Analyses
Sometimes large datasets require more complicated methods to find relevant and interesting features. These features can be thought of as solutions. How do we use complex analysis to find these features? In the world of analysis of analyses, large datasets can be mapped to solution spaces with a defined shape. This strategy uses convergence/triangulation as a guiding principle, but does so through the rules of metric geometry and computational complexity. A related and emerging approach called topological data analysis [10] can be used to conduct rigorous relational analyses. Topological data analysis takes datasets and maps them to a geometric shape (e.g. topology) such as a tree or in this case a surface.


A portrait of convexity (quadratic function). A gently sloping dataset, a gently sloping hypothesis space. And nothing could be further from the truth......

In topological data analyses, the solution space encloses all possible answers on a surface, while the surface itself has a shape that represents how easy it is to move from one portion of the solution space to another.
One common assumption is that this solution space is known and finite, while the shape is convex (e.g. a gentle curve). If that were always true, then analysis would be easy: we could use a moderate large-sized dataset to get the gist of patterns in the data. any additional scientific inquiry would constitute filling in the gaps. And indeed sometimes it works out this way.

One example of a topological data analysis of most likely Basketball positions (includes both existing and possible positions). COURTESY: Ayasdi Analytics and [10].

The Big Data Backlash.....Enter Meta-Analysis
Despite its successes, there is nevertheless a big data backlash. Ernest Davis and Gary Marcus [11] present us with nine reasons why big data are problematic. Some of these have been covered in the last section, while others suggest that there can be too much data. This is an interesting position, since it is common wisdom that more data always give you more resolution and insight. Insight and information can be obscured by noisy or irrelevant data. But even the most informative of datasets can yield misinformed analyses if the analyst is not thoughtful.

Of course, ever-bigger datasets by themselves do not give us the insights necessary to determine whether or not a generalized relationship is significant. The ultimate goal of data analysis should be to gain deep insights into whatever the data represent. While this does involve a degree of interpretive subjectivity, it also requires an intimate dialogue between analysis, theory, and simulation. Perhaps the latter is much more important, particularly in cases where the data are politically or socially sensitive. These considerations are missing from much contemporary big data analysis [12]. This vision goes beyond the conventional "statistical test on a single experiment" kind of experimental investigation, and leads us to meta-analysis.

The basic premise of a meta-analysis is to use a strategy of convergence/triangulation to converge upon results using a series of studies. The logic here involves using the power of consensus and statistical power to arrive at a solution. The problem is represented as a series of experiments with an effect size for each. For example, if I believe that eating oranges causes cancer, how should I arrive at a sound conclusion? One study with a very large effect size, or many studies with various effect sizes and experimental contexts. According to the meta-analysis view, the latter should be most informative. In the case of potential factors in myocardial infarction [13], significant results that all point in the same direction (with minimum effect size variability) lend the strongest support to a given hypothesis.

Example of a meta-analysis. COURTESY: [13].

The Problem with Deep Analysis
We can go even further down the rabbit hole of analysis, for better or for worse. However, this often leads to problems of interpretation, as deep analyses are essentially layered abstractions. In other words, they are higher-level abstractions dependent upon lower-level abstractions. This leads us to a representation of representations, which will be covered in an upcoming post. Here, I will propose and briefly explore two phenomena: significant pattern extraction and significant reconstructive mimesis.

One form of deep analysis involves significant pattern extraction. While the academic field of pattern recognition has made great strides [14], sometimes the collection of data (which involve pre-processing and personal bias) is flawed. Other times, it is the subjective interpretation of these data which are flawed. In either case, this results in the extraction patterns that make no sense that are then assigned significance. Worse yet, some of these patterns are also thought to be of great symbolic significance [15]. The Bible Code is one example of such pseudo-analysis. Patterns (in this case secret codes) are extracted from a database (a book), and then these data are probed for novel but coincidental pattern formation (codes formed by the first letter of every line of text). As this is usually interpreted as decryption (or deconvolution) of an intentionally placed message, significant pattern extraction is related to the deep, naive theories discussed in "Structure and Theory of Theories".

Congratulations! Your pattern recognition algorithm came up with a match. Although if it were a computer instead of a mind, it might do a more systematic job of rejecting it as a false positive. LESSON: the confirmatory criteria for a significant result needs to be rigorous.

But suppose that our conclusions are not guided by unconscious personal biases or ignorance. We might intentionally leverage biases in the service of parsimony (or making things simpler). Sometimes, the shortcuts we take in representing natural processes present difficulties in understanding what is really going on. This is a problem of significant reconstructive mimesis. In the case of molecular animations, this has been pointed out by Carl Zimmer [16] and PZ Myers [17] for molecular animations. In most molecular animations, processes occur smoothly (without error) and within full view of the human observer. Contrast this with the inherent noisiness and spatially-crowded environment of the cell, which is highly realistic but not very understandable. In such cases, we construct a model which consists of data, but that model is selective and the data is deliberately sparse (in this case smoothed). This is an example of a representation (the model) that informs an additional representation (the data). For purposes of simplicity, the model and data are somehow compressed to preserve signal and remove noise. And in the case of a digital image file (e.g. .jpg, .gif) such schemes work pretty well. But in other cases, the data are not well-known, and significant distortions are actually intentional. This is where big challenges arise in getting things right.

An multi-layered abstraction from a highly-complex multivariate dataset? Perhaps. COURTESY: Salvador Dali, Three Sphinxes of Bikini.

Conclusions
Data analysis is hard. But in the world of everyday science, we often forget how complex and difficult this endeavor is. Modern software packages have made the basic and well-established analysis techniques deceptively simple to employ. In moving to big data and multivariate datasets, however, we begin to face head-on the challenges of analysis. In some cases, highly effective techniques have simply not been developed yet. This will require creativity and empirical investigation, things we do not often associate with statistical analysis. It will also require a role for theory, and perhaps even the theory of theories.

As we can see from our last few examples, advanced data analysis can require conceptual modeling (or representations). And sometimes, we need to map between domains (from models to other, higher-order models) to make sense of a dataset. This, the most complex of analyses, can be considered representations of representations. Whether a particular representation of a representation is useful or not depends upon how much noiseless information can be extracted from the available data. Particularly robust high-level models can take very little data and provide us with a very reliable result. But this is an ideal situation, and often even the best models presented with large amounts of data can fail to given a reasonable answer. Representations of a representations also provide us with the opportunity to imbue an analysis with deep meaning. In a subsequent post, I will this out in more detail. For now, I leave you with this quote:
“An unsophisticated forecaster uses statistics as a drunken man uses lampposts — for support rather than for illumination.” Andrew Lang.

NOTES:
[1] Learn Statistics with Comic Books. CTRL Lab Notebook, April 14 (2011).

[2] Mooney, C.   The Science of Why We Don't Believe Science. Mother Jones, May/June (2011).

[3] Kosko, B.   Statistical Independence: What Scientific Idea Is Ready For Retirement. Edge Annual Question (2014).

[4] In order to separate signal from noise, we must first define noise. Noise is consistent with processes that occur at random, such as the null hypothesis or a coin flip. Using this framework, a significant result (or signal) is a result that deviates from random chance to some degree. For example, a p-value of 0.05 represents a 95% chance that the replicates observed could not have occurred due to chance. This is, of course, an incomplete account of the relationship between signal and noise. Models such as Signal Detection Theory (SDT) or data smoothing techniques can also be used to improve the signal-to-noise ratio.

[5] Nuzzo, R.   Scientific Method: Statistical Errors. Nature News and Comment, February 12 (2014).

[6] Fox, J.   Frequentist vs. Bayesian Statistics: resources to help you choose. Oikos blog, October 11 (2011).

[7] Gelman, A.   So-called Bayesian hypothesis testing is just as bad as regular hypothesis testing. Statistical Modeling, Causal Inference, and Social Science blog, April 2 (2011).

[8] For some concrete (and satirical) examples of how correlation does not equal causation, please see Tyler Vigen's Spurious Correlations blog.

[9] Voytek, B.   Big Data: what's it good for? Oscillatory Thoughts blog, January 30 (2014).

[10] Beckham, J.   Analytics Reveal 13 New Basketball Positions. Wired, April 30 (2012).

[11] Davis, E. and Marcus, G.   Eight (No, Nine!) Problems with Big Data. NYTimes Opinion, April 6 (2014).

[12] Leek, J.   Why big data is in trouble - they forgot applied statistics. Simply Statistics blog, May 7 (2014).

[13] Egger, M.   Bias in meta-analysis detected by a simple, graphical test. BMJ, 315 (1997).

[14] Jain, A.K., Duin, R.P.W., and Mao, J.   Statistical Pattern Recognition: a review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(1), 4-36 (2000).

It is interesting to note that the practice of statistical pattern recognition (training a statistical model with data to evaluate additional instances of data) has developed techniques and theories related to rigorously rejecting false positives and other spurious results.

[15] McCardle, G.     Pareidolia, or Why is Jesus on my Toast? Skeptoid blog, June 6 (2011).

[16] Zimmer, C.     Watch Proteins Do the Jitterbug. NYTimes, April 10 (2014).

[17] Myers, P.Z.   Molecular Machines! Pharyngula blog, September 3 (2006).

July 23, 2013

WARNING: Data and Narratives may lead to Bias

This post contains three features cross-posted to my micro-blog, Tumbld Thoughts. I am carving something at its joints here (see note #5 for reference), but am not sure exactly what. I guess it is the role of belief and human nature in things we usually (from a common-sense perspective, at least) consider to be objective and/or conceptually attractive (e.g. data analysis, scientific theory, and intuitions about complexity).

Featured are three loosely-related topics: the uncritical interpretation of big data (I), shortcomings of the narrative explanation (II), and subtle but important biases in scientific thinking (III). Nothing is sacred (or at least reverent) here, but then again it shouldn't be.

I. Uncritical Interpretations of "Big" Data


This is what a religion based on big data would look like. Or rather, this is what the uncritical interpretation of big data [1] currently looks like. A few readings on the topic:

1) a news article [2] on how the mis-interpretation of big data (and over-reliance on its correlative relationships) threaten to undermine effective decision-making. Of greatest importance is distinguishing between correlation and causation, which is a data analysis (and logical reasoning) issue that predates big data.


2) an op-ed [3] featuring the work of Seth Stephens-Davidowitz, an economist from Google who is debunking the idea that child abuse and neglect decreased in the aftermath of the 2008 economic crisis. He is accomplishing this using a novel methodology, an example of how using different methods to address the same problem yields different results. In this case, aggregate Google searches (based on unobserved, online activity) were used. An example of how we might better extract causal relationships from "big" datasets.



II. Going beyond "the narrative" explanation

Is "the narrative" the best way to convey complex ideas, especially when it comes to social or scientific explanation? Barry Ritholtz and Cullen Roche [4] remind us that in contemporary economics, the prevailing narratives often fail to capture the complexity or even the outcomes of real-world situations. 

This failure is one of explanations that have no capacity to take into account conflicting evidence. This ultimately results in cognitive dissonance, which becomes more pronounced as the narrative explanations continue to fail.

That being said, using a narrative structure to describe the function and complexity of scientific and social concepts is not always to be avoided. Here is my list of reasons why narratives "work" in this context:

1) narratives contain common structures that are shared across cultures.

2) narrative structures are consistent with naive models [5], which represent an intuitive view of the world [6].

3) narratives are compact ways to encode information, such as oral traditions or feature films. Compare the amount of potential information contained in these with a Wiki or a blog post.

Why narratives do not work in this context:

1) narratives can perpetuate naive models of the world in the face of contradictory evidence. Examples of this are given in [4].

2) naive models (and thus narratives) are often conservative, and do not encode nonlinear effects or the parallel progression of events. Narrative thinking favors simple cause-and-effect mechanisms over mechanisms that favor multiple causes or long-term, delayed outcomes [7].

3) like most specialized information, they often require intersubjectivity. Unlike most specialized information, they require a moral logic.This may or may not obfuscate the interpretation of events.

4) narratives often utilize allegorical arguments, in which a single string of text can be interpreted in many ways. For example, if a message is passed around a circle, if often changes due to intrasubjectivity (e.g. individual interpretations). While this is good for cultural diversity, it is not so good for scientific replication.

Whereas the narrative is valuable to a communicator, it may be less valuable from a technical standpoint. Overall, using "plain language" and "narrative structure" can actually undercut the scientific content. 


III. Bias in Scientific Thinking


Is it bias, or is it good science? A series of recent papers/talks may give us some insight [8]. The first is a Skepticon IV talk [9] and Measure of Doubt blog post by Juila Galef on the "Straw Vulcan" phenomenon. A Straw Vulcan is shorthand for the popular misconceptions surrounding logical decisionmaking and how what may seem logical may be co-opted by emotionally-driven biases.


Interesting enough, but what does this have to do with science? Well, recently published papers in PLoS Biology [10] suggest that bias is a natural feature of scientific thinking, which results from a tension between the need to shift paradigms (Kuhnian science) and the need to falsify hypotheses (Popperian science)


NOTES:

[1] Whitehorn, M.   The Parable of Beer and Diapers. The Register, August 15 (2006).

[2] Asay, M.   Big Data's Dehumanizing Impact on Public Policy. ReadWrite content aggregator, July 12 (2013).

[3] Stephens-Davidowitz, S.   How Googling Unmasks Child Abuse. NYT Opinion, July 13 (2013).

[4] Examples from so-called "common sense knowledge" in economic issues includes:

a) Ritholtz, B.   The Narrative Fails. The Big Picture blog, July 19 (2013).

b) Roche, C.   The Fear Trade has been Demolished. Pragmatic Capitalism blog, July 19 (2013).

[5] For a structure learning perspective, please see: Gershman, S.J. and Niv, Y.   Learning latent structure: carving nature at its joints. Current Opinion in Neurobiology, 20, 251-256 (2010).

[6] for more information in the role of narratives vs. data in the wealth inequality debates, see the following:

a) Norton, M.I. and Ariely, D.   Building a better America -- one wealth quintile at a time. Perspectives on Psychological Science, 6(1), 9-12 (2011). Bottom image is taken from Figure 2.

b) Noah, T.   Theoretical Egalitarians. Slate.com, September 27 (2010).

[7] Wexler, M.   Invisible Hands: intelligent design and free markets. Journal of Ideology, 33 (2011).


[9] here is video of Julia Galef's lecture at Skepticon IV, and here is her post on the topic at Measure of Doubt blog.

[10] relevant papers from this issue include the following:

a) Chase, J.M.   The Shadow of Bias. PLoS Biology, 11(7), e1001608 (2013).

b) Tsilidis, K.K., Panagiotou, O.A., Sena, E.S., Aretouli, E., Evangelou, E., Howells, D.W., Al-Shahi Salman, R., Macleod, M.R., and Ioannidis, J.P.A.   Evaluation of Excess Significance Bias in Animal Studies of Neurological Diseases. PLoS Biology, 11(7), e1001609 (2013).

Printfriendly