Showing posts with label arXiv. Show all posts
Showing posts with label arXiv. Show all posts

August 9, 2022

New Paper on Developmental Braitenberg Vehicles now live!

 

The special issue of Artificial Life on Embodied Intelligence is now live! Inside you will find our paper "Braitenberg Vehicles as Developmental Neurosimulation", which has lived on the arXiv since 2020. This paper lays out an approach to Developmental Neurosimulation, involving three adversarial approaches to the agent-based development of embodied brains and embodied cognition. Here is the abstract:

Connecting brain and behavior is a longstanding issue in the areas of behavioral science, artificial intelligence, and neurobiology. As is standard among models of artificial and biological neural networks, an analogue of the fully mature brain is presented as a blank slate. However, this does not consider the realities of biological development and developmental learning. Our purpose is to model the development of an artificial organism that exhibits complex behaviors. We introduce three alternate approaches to demonstrate how developmental embodied agents can be implemented. The resulting developmental Braitenberg vehicles (dBVs) will generate behaviors ranging from stimulus responses to group behavior that resembles collective motion. We will situate this work in the domain of artificial brain networks along with broader themes such as embodied cognition, feedback, and emergence. Our perspective is exemplified by three software instantiations that demonstrate how a BV-genetic algorithm hybrid model, a multisensory Hebbian learning model, and multi-agent approaches can be used to approach BV development. We introduce use cases such as optimized spatial cognition (vehicle-genetic algorithm hybrid model), hinges connecting behavioral and neural models (multisensory Hebbian learning model), and cumulative classification (multi-agent approaches). In conclusion, we consider future applications of the developmental neurosimulation approach.

There are many themes to follow up on in this paper. Just of few examples include:

* brain/body scaling in an embodied agent.

* the role of multisensory integration in the development of cognition.

* ways to classify shapes and motifs in the emergence of multi-agent collectives. 

* spatial cognition and transfer learning in developmental embodied systems.

Congratulations to Stefan Dvoretskii, Ziyi Gong, Ankit Gupta, Jesse Parent, and Bradly Alicea for their hard work.

October 24, 2016

Open Access Week: How Am I Doing, Altmetrics?

This is one of two posts in celebration of Open Access week (on Twitter: #oaweek, #open access, #OpenScience). To kick things off, we will go through an informal evaluation of Altmetrics and other indicators of research paper usership.

In this post, I will discuss some quick investigations I did using the Altmetric metric system (known visually as the number within the multicolored donut). Altmetrics go beyond academic metrics based solely on academic journal prestige or number of formal citations in academic papers (e.g. h-index). In this post, I will discuss how these metrics might be used to help better understand the full impact on one's work.


The Altmetric donut and its diversity of input sources. The Altmetric score is based on how many interactions your content received from each source medium.

The first exercise I did was to acquire Altmetric donuts for journal articles and preprints for which I did not have such data. This includes venues such as arXiv, Stem Cells and Development, and Principles of Cloning II, which do not feature Altmetric donuts on their pages. Interestingly, the bioRxiv preprint server does, in addition to tracking .pdf download and abstract view counts.



Example of an Altmetric donut in context (top) and readership stats (bottom) from a recent Biology paper for which I am an author. 

Retrieving a donut and data summary from the Altmetric database is easy. You embed a few line of code (see inset below) into an HTML document, and the donut and score appear where desired. While the donut is most useful for augmenting a publication list, in this case I simply created a test document for collating data from across many papers.

// Formal journal article citation
Alicea, B., Murthy, S., Keaton, S.A., Cobbett, P., Cibelli, J.B., and Suhr, S.T.  Defining phenotypic respecification diversity using multiple cell lines and reprogramming regimens. Stem Cells and Development, 22(19), 2641-2654 (2013).
// Code for donut and database call; possible data subclasses include:
// data-arxiv-id
// data-handle
// data-doi

In context, the donut can provide useful information about how a given paper is diffusing through the academic internet. In the case of the Stem Cells and Development paper (see code), the paper has an Altmetric score of 9. While the Journal website does not have Altmetric or download data, it does provide a doi identifier and select forward citations.


Examples of the Altmetric database entry (top) and the Journal website (bottom) for the Stem Cells and Development paper.

Similar data exist for a follow-up paper to the Stem Cells and Development paper -- in this case, a preprint involving a specialized quantitative analysis (based on Signal Detection Theory) of the same data. For this paper, we have an arXiv identifier, which provides us with a donut and statistics on the relative popularity of the paper based on age and other similar documents in the Altmetric database.

A typical arXiv article page, in this case for an arXiv preprint related to the Stem Cells and Development paper.

This arXiv preprint comes with code for the analysis, which is posted to Github.

For this particular paper, there is an associated Github repository. Even for preprint repositories with Altmetric and readership data (such as bioRxiv), the integration of Github materials is rather poor, particularly in generating an Altmetric. Alternately, there is an opportunity for Github to This is an area for which user statistics linked back to the original paper would be appreciated. 

Altmetrics for the same arXiv preprint. We can access data on the sources of the Altmetric score, as well as the attention score in the context of all other tracked documents in the Altmetrics database.

We can also integrate readership data across sources to come up with a picture of how our academic work is being shared, consumed, and diffused. In this example, I will show how data from a blog analytics engine and Altmetric data can be combined. Research blogs are an up-and-coming area of research in Altmetric statistics capture. I have taken two blogrolls (Carnival of Evolution #46 and Carnival of Evolution #70), for which citable versions were posted to Figshare immediately after going live. My blogging platform (Blogger) has readership stats but no Altmetrics, while Figshare has Altmetrics and readership stats for the Figshare version only.



Altmetric data for two blogrolls cross-posted to Figshare, which provides both a doi identifier and an Altmetric donut. There is also view and download information for the Figshare version, which may or may not be inclusive of people viewing such content on the blog site.

Let's look at the Figshare data first. Carnival #46 has an Altmetrics score of 10 with 188 views and 58 downloads. By contrast, Carnival #70 has an Altmetrics score of 6 with 331 views and 82 downloads. Clearly, there is some variation in direct engagement between the two datasets that is proportional to the score.


Readership statistics for Carnival of Evolution #46 (top) and Carnival of Evolution 
#70 (bottom). Blog analytics only provides the number of "reads" on the home site since publication.

There is also little relationship between the number of Blogger reads and the Altmetric score (as the Altmetric score does not directly capture this number). Carnival #46 has 7928 reads over roughly 4 years and 7 months. Carnival #70 has 1602 reads over roughly 2 years and 7 months. 

Even in cases where no Altmetric donut can be generated (such as for book chapters), there are still ways to evaluate an article's reach. In the case of Academia.edu, a new feature has been added that allows people to leave a comment when they interact with a document. This is a more qualitative assessment of engagement, but also provides authors an idea of whether or not "reads" or "views" translate into more than just a passing glance.

Two consumers of a book chapter took time to express their gratitude. Other reasons can be quite interesting as well, particularly when they have to do with educational purposes.

Hope you have enjoyed this exercise. It is not meant to be an exhaustive discussion of the Altmetric evaluation system, nor is it the limit of what can be done with Altmetrics and other tools for tracking you work. While there is clearly more technical work to be done on this front, tools such as Altmetric APIs are available. The biggest challenge is to building a social economy based on a variety of research outputs. The field is moving quite rapidly, so what I have shown here is likely to be just the beginning. 

September 29, 2015

Reconsidering the Model as a Unit of Regulation: cybernetics and the adaptive outcome

Here is a preview of an essay Robert Stone and I have been working on as part of the Orthogonal Research initiative during the course of the last year. The formal title is: "The Foundations of Control and Cognition: The Every Good Regulator Theorem". This essay takes a classical tool from the cybernetics literature and applies it to game theoretic and other problems of our interest.

Robert Stone, cybernetics enthusiast

Robert Stone and myself, bringing cybernetics back to the "soft" but immensely-complex (social, brain, and biological) sciences. The full version (with notes, definitions, and additional references) can be found here.

A seemingly simple discrete system with feedback (which makes it not so simple during future iterations). COURTESY: intgr, Wikimedia commons.

I. Introduction
            In the history of scientific discovery, there have been examples of certain persons or facets of their work being considered ‘out of step’ with the dominant scientific or philosophical trends of the time. As such, they risk falling down a deep well in our cultural landscape, with their work’s efficacy lost to subsequent generations. If their work has merit, it may be considered ahead of its’ time by future generations. The timing of a given theory or great idea is largely determined by cultural and cognitive biases that favor the dominant paradigm [1]. In other cases, ideas at the paradigmatic vanguard end up resurrected in a more pragmatic way. The acceptance of such ideas occurs either gradually or in one fell swoop at a later point in time. Let us keep this in mind as we discuss Ronald C. Conant and W. Ross Ashby’s seminal work “Every Good Regulator Theorem” [2] (EGRT):

“[The EGRT is]….a theorem is presented which shows, under very broad conditions, that any regulator that is maximally both successful and simple must be isomorphic with the system being regulated…….Making a model is thus necessary.” [2]

The EGRT characterizes regulation with respect to cybernated control systems. In the case of Ashby and Conant [3], the EGRT developed within the context of several intersecting traditional fields. These include algorithmics, information theory, systems theory, and behavioral science. In such a context, models are exceedingly important. Given the reliance of the EGRT concept on inference and propositional thinking, there is an essential reliance on models. In fact, the EGRT exists at such a high level of abstraction that even with a high degree of specification may not be directly applicable in the real world [4]. However, there are certain advantages of cybernetic modeling that make their cross-contextual application useful.

Ashby's graphical formulation of the EGRT Theorem with original notation. COURTESY: [2].


II. Background
Let us return to the notion of modeling as phenomenology. Systems engage in modeling not simply to purposely regulate their environments, but rather to reactively respond to input stimuli in a way that maintains higher-level states [5]. This ability to model becomes part of their structure at the most basic of levels, though it would be fair to say most modeling (in the way we will use the word) is the result of cognitive processes. The constructivist might argue that such metacognitive dynamics [6] would influence one’s proposed scientific model. Like Shakespeare’s Hamlet, however, the question of whether or not to model (or be) is one of survival, whether that survival be genetic or memetic. Rather than reviewing the proof step-by-step, let’s discuss its potential significance in a variety of use-cases. In the process, we will be transcending the traditional boundaries of autonomic, ‘choice’, or even cognitive.

          Simply put, the Every Good Regulator Theorem says that regulators operate on approximations (e.g. models) of the thing they are regulating. This requires a mapping of the natural world to the model. While one might consider the activities of encoding and translation to be inherently cognitive, genomic systems also perform biological control functions in the absence of cognition [7, 8]. In the biological control example, what matters is not intent, but accuracy. Rather than an actively goal-oriented criterion, what we observe here is passively goal-oriented system output. Accuracy of the approximated model influences the quality of regulation. Thus, there need not be agency on the part of any single system component. Indeed, to survive as a unit in an interrelated system, a regulating machine must construct an interactive model that includes inputs, outputs, and feedback.

Let us consider a couple cases of regulatory dynamics, which may be valuable in understanding the importance of this theorem. We can then move on to what could this mean for both further theoretical development and practical application. A good place to begin in cognitive science is game theory [9]. One of the most simple, effective, and most explanatory strategies in the Prisoners’ Dilemma game is the tit-for-tat strategy [10]. In this 2-player, 2x2 game, the tit-for-tat strategy is simple: ‘Do unto others as they have done unto you’ after an initial good faith move of cooperation. The strategy is simply to copy your opponent's behavior. If the opposing agent cooperates, so does the tit-for-tat strategizing agent; if they defect, the tit-for-tat strategist follows suit. The intended outcome of the strategy is to move the exchange towards an equilibrium (though this is not the only possible outcome, nor is the strategy perfect).

Of specific interest here is that the mechanics of the strategy requires a model to be held in memory by the agent employing tit-for tat (a 1-bit cooperate/defect model), regardless of the strategy employed by the other agent (whether that be a more sophisticated maximizing strategy, or random selections). While an economist might view this as free-riding behavior by one of the two agents, the selection of tit-for-tat by both players can produce a cooperative equilibrium, such as in the evolution of reciprocal altruism in biological systems. The EGRT suggests that the greater the memory for an agent, and the longer it has the opportunity to observe and integrate the moves of its opponent, the greater its’ potential for effective regulation.

Over time, this can lead to greater accuracy for the agent’s cognitive model and a more stable equilibrium game outcome. Further, this equilibrium state can be long-lasting, given extended memory capacity for more detailed models, and may evolve towards ‘a conspiracy of doves’, within a game of homo lupus homini. An agent with a greater memory capacity can also employ more elaborate (or deeper) strategies over time. This development of deeper strategies may also feedback into modifying its model of the external world [11]. Overall, the capability to regulate behavior of other players depends on the inferential and predictive capacities of each player’s model: in a highly complex competitive game environment, a good regulator has a superior model, or it will find itself regulated by a competing agent in the game, especially as the behaviors get more complex.

“The theorem has the interesting corollary that the living brain, so far as it is to be successful and efficient as a regulator for survival, must proceed, in learning, by the formation of a model (or models) of its environment.” [2]

An example of a basic 2x2 payoff matrix characterizing the Prisoner's Dilemma. COURTESY: "Extortion in Prisoner's Dilemma", Blank on the Map blog, September 19 (2012).

III. Further Considerations
Let us now consider a more complicated scenario where we might be able to uncover the universal components of the EGRT phenomenology. The context will be two people on a blind date (this can actually be a complicated scenario). If one has been in one of these (terrifying) contexts, then one can already see where we are going. The cognitive agents are continually competing to increase the efficacy of their models of the other agent, while also attempting to constrain the modeling of the other agent towards a compact image they prefer. Although rarely implemented successfully, winning strategies include accurately modeling the other actor and influencing the state of their mental model. This can include both elaborate, multi-step strategies, and simpler strategies, the complexity of which is does not indicate their effectiveness. If the goal is a continuation of relations, the acquisition and intentional obfuscation of information occurs at appropriate times and in appropriate ways. Furthermore, this information has contextual value. As in most scenarios involving imperfect or asymmetrical information [12], your model must be superior to become the leader of the interaction [13], and thus control of regulation.

Does regulation even require what we would call cognition? This of course depends on our definition of cognition and regulation. However, let us consider that a bacterium does not have a “cognitive” or mental model of its environment, yet appears to have little trouble getting around and controlling some aspects of its landscape. The similarities between chemotactic sensation and mental models built upon multisensory stimuli serve as evidence for the universal character of the EGRT. In fact, Heylighen [14] has proposed that cybernetic regulation is a highly-generalized form of cognition. Yet do thermostats or other mechanical systems possess anything approaching what we consider cognition? While none of these has the cognitive capacity of a brain, they do have information processing capabilities from their physical or electronic structure, memory states, and crude models of how things ‘should’ be, towards which they regulate conditions. Non-cognitive systems possessing these characteristics are obviously still capable of rudimentary communication, control, decision making, and regulation, at least abstractly. We should also expect some degree of continuity that crosses the boundary of the cognitive and non-cognitive, since cognitive systems evolved from less intentional ones with more rudimentary forms of behavioral control.

“...success in regulation implies that a sufficiently similar model must have been built, whether it was done explicitly, or simply developed as the regulator was improved.” [2]


IV. Conclusion
Earlier, we had touched upon the history of scientific discovery, and contextual model building. A scientific theory is simply a model, and its value lies in its efficacy and repeatability (thus its’ trustworthiness and ability to aid in regulation). Theoretical models have tended, historically, to shift from informal, conceptual models towards formal mathematical ones (consider Comte’s Philosophy of Science). As a given model acquires more data, and as those data create ever-more accurate model revisions with higher fidelity. The overall capacity to aid regulation increases via feedback. Thus, the model’s value to humans increases. However, as noted by the example of ahead of their time thinking, scientific thought does not exist in a vacuum, and the landscape conditions need to be aligned so that the model can prove fruitful. Consider how we are witnessing an explosion in robust formal mathematical and/or computer models either aiding or besting human cognitive efforts [15, 16]. Informational revisions of the model often occur faster than the landscape conditions change, so adaptive cross-contextual models may prove more successful in dynamic situations, such as ones which are developed by human thought and human cultural systems.

            This ability to cross the boundary between cognitive and non-cognitive with models may challenge either our informal, colloquial conception of cognition or the universality criterion of the formal EGRT. As both features of cognition and more universal mechanisms, information processing, memory, communication, and selection can occur without any kind of cognitive superstructure. Perhaps the context of what we call “cognition” is too limiting. What about human cognition then is truly universal, and what is unique to a certain set mechanisms and representational models? For example, are models of so-called cellular decision-making [17] an unduly anthropomorphic representation of cellular differentiation and metabolism, or is it drawing upon a common set of universal properties that can only be abstracted from the system by an appropriate model?

Rather than trying to solve this philosophical puzzle now, let us take leave to consider that a deep truth like the one perhaps contained within the formalism of the EGRT should make us question scientific knowledge in a manner akin to reconsidering our firmly-held beliefs. It should make us reconsider how well we understand the relationship between nature and our own conceptual models. In that, it kindles the same spark from which all great scientific theories alight: It leads us to more questions, new ways of thinking about things, and guides us towards more accurate, repeatable, and otherwise ‘good’ models.

“Now that we know that any regulator (if it conforms to the qualifications given) must model what it regulates, we can proceed to measure how efficiently the brain carries out this process. There can no longer be question about whether the brain models its environment: it must.” [2]

References:
[1] Kuhn, T.   Structure of Scientific Revolutions. University of Chicago Press (1962). 

[2] Conant, R.C. and Ashby, W.R.   Every good regulator of a system must be a model of that system. International Journal of Systems Science, 1(2), 89–97 (1970).

[3] Ashby, W.R.   Introduction to Cybernetics. Chapman and Hall (1962).

[4] Fishwick, P.   The Role of Process Abstraction in Simulation. IEEE Transactions on Systems, Man, and Cybernetics, 18(1), 18-39 (1988).

[5] Brooks, R.   Intelligence Without Representation. Artificial Intelligence, 47, 139-159 (1991).

[6] Kornell, N. Metacognition in Humans and Animals. Current Directions in Psychological Science, 18(1), 11-15 (2009).

[7] Ertel, A. and Tozeren, A.   Human and mouse switch-like genes share common transcriptional regulatory mechanisms for bimodality. BMC Genomics, 23(9), 628 (2008).

[8] Gormley, M. and Tozeren, A.   Expression profiles of switch-like genes accurately classify tissue and infectious disease phenotypes in model-based classification. BMC Bioinformatics, 9, 486 (2008).

[9] Gintis, H.   Game Theory Evolving. Princeton University Press (2000).

[10] Imhof, L.A., Fudenberg, D., and Nowak, M.A.   Tit-for-tat or Win-stay, Lose-shift? Journal of Theoretical Biology, 247(3), 574–580 (2007).

[11] Liberatore, P. and Schaerf, M.   Belief Revision and Update: Complexity of Model Checking. Journal of Computer and System Sciences, 62(1), 43–72 (2001).

[12] Rasmussen, E.   Games and Information: an introduction to game theory. Blackwell Publishing (2006).

[13] Simaan, M. and Cruz, J.B.   On the Stackleberg Strategy in Nonzero-Sum Games. Journal of Optimization Theory and Applications, 11(5), 533-555 (1973).

[14] Heylighen, F.   Principles of Systems and Cybernetics: an evolutionary perspective. CiteSeerX, doi:10.1.1.32.7220 http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.32.7220 (1992).

[15] LeCun, Y., Bengio, Y., and Hinton, G. Deep Learning. Nature, 521, 436-444 (2015).

[16] Ferrucci, D., Brown, E., Chu-Carroll, J., Fan, J., Gondek, D., Kalyanpur, A.A., Lally, A., Murdock, J.W., Nyberg, E., Prager, J., Schlaefer, N., and Welty, C.   Building Watson: an overview of the DeepQA project. AI Magazine, Fall (2010).

[17] Kobayashi, T.J., Kamimura, A.   Theoretical aspects of cellular decision-making and information-processing. Advances in Experimental Medicine and Biology, 736, 275-291 (2012).


UPDATE (9/30): During the editorial process, Rob and I had a discussion about using the word "alight" (in the final paragraph). I was not sure about the correct word usage, but Rob assured me that it was being used correctly in this context. But to back this up even further (and to gratuitously insert an informatics Easter Egg), here is the Google Ngram history of "alight" usage since 1800. 




March 30, 2015

Causality, part II (was it caused by Part I?)

This post serves as a follow-up to a Synthetic Daisies post written in 2012 on new methods to detect causality in data.

Here are a few interesting readings at the intersection of data analysis and the philosophy of science. The first [1] is a new arXiv paper [2] that evaluates two approaches to evaluating causality using two machine learning techniques. A plethora of discriminative machine learning techniques have emerged in recent years to address relatively simple relationships. In terms of cause and effect itself, the distinguishing signal is often subtle and unclear even for seemingly obvious sets of relationships. In [2], techniques called Additive Noise Methods [3] and Information Geometric Causal Influence [4]. A dataset called CauseEffectPairs [5] was used to benchmark each method, and show that causal relationships can be uncovered from a wide variety of data.


The second paper (or rather series of papers) is on the topic of strong inference [6]. Strong inference is an alternative to hyper-reductionism and the use of over-simplified models. Strong inference involves the use of a conditional inductive tree to examine the possible causes for a given phenomenon [7]. Potential causes (or hypotheses) represent nodes of the tree, and these hypotheses are falsified as one moves through the tree using either inductive or empirical criteria. Unlike the machine learning models we discussed, the goal is to lead a researcher to key experiments that help to uncover the sources of variation. In general, this process of elimination lead us to the best answers, Yet according to Platt in [2], this approach can ultimarely provide us with axiomatic statements.

Conceptual steps involved in strong inference. COURTESY: Figure 1 in [8].

While this seems to be a fruitful methodology, it has turned out to be more inspirational than as a source of analytical rigor [9]. Strong inference hs inflenced a variety of scientific fields concentrated in the biological and social sciences. Platt predicted [2] that sciences that concurred with strong inference would be fields that experienced a greater number of breakthrough advances. However, in testing Platt's predictions regarding the efficacy of Strong Inference, is have been found that advances are not directly related to the adoption of the method [10]. This could be due to our incomplete understanding of the factors that drive scientific discovery and the rate of advancement. 


[2] Mooij, J.M., Peters, J., Janzing, D., Zscheischler, J., and Scholkopf, B.   Distinguishing cause from effect using observational data: methods and benchmarks. arXiv, 1412.3773 (2014).

[3] Hoyer, P.O., Janzing, D., Mooij, J.M., Peters, J., and Scholkopf, B.   Nonlinear causal discovery with additive noise models. In Advances in Neural Information Processing Systems (NIPS), 21, 689-696 (2009).

[4] Daniusis, P., Janzing, D., Mooij, J.M., Zscheischler, J., Steudel, B., Zhang, K., and Scholkopf, B. Inferring deterministic causal relations. In Proceedings of the 26th Annual Conference on Uncertainty in Artificial Intelligence (UAI), 143-150 (2010).

[5] This work was part of the CauseEffect Pairs Challenge and was presented at NIPS 2013.

[6] Platt, J.R.   Strong Inference: certain systematic methods of scientific thinking may produce much more rapid progress than others. Science, 146(3642), 347-352 (1964).

[7] Neuroskeptic   Is Science Broken? Let's Ask Carl Popper. Neuroskeptic blog, March 15 (2015).

[8] Fudge, D.S.   Fifty years of J.R. Platt's Strong Inference. Journal of Experimental Biology, 217, 1202-1204 (2014).

[9] Davis, R.H.   Strong Inference: rationale or inspiration? Perspectives in Biology and Medicine, 49(2), 238-250 (2006).

[10] O'Donohue, W. and Buchanan, J.A.   The Weaknesses of Strong Inference. Behavior and Philosophy, 29, 1-20 (2001).

March 14, 2015

A Modest Framework for Scientific Transparency

Here are six points for the integration of open-access science publishing and open data. This was developed from personal practice and research in addition to interactions with the Research Data Service (University of Illinois) and the SciFund challenge. This pipeline begins at the write-up stage, but some points rely on practice prior to analysis and write-up.


A)   Preprint (e.g. kernel of hypothesis- or question-driven results).

A number of options exist for this, including arXiv, bioRxiv, PLoS One, or another permanent location that provides a formal archival address or digital object identifier (doi). The core paper should be brief (6-12 pgs) and formal.


B)   Advanced methods/theory.

These can be submitted as supplemental materials, either in the same repository as the preprint itself or on another permanent server. As opposed to simple auxillary files, this should be set up more along the lines of an iPython notebook.


C)   Advanced Analysis.

This can be treated in the same manner as the advanced methods/theory. This will include transformational datasets (e.g. time-frequency decompositions, log transforms, combinations of data from multiple sources in a common framework) and the associated data tables and figures/graphs.


D)   Datasets.

1)   Raw Data: images, unprocessed vectorial or matricial output.

These will be stored as formatted image files, ASCII files, or tabular files.

2)   Processed Data: numeric variables, simple annotation.

These will be appended to the raw data either in the file or as linked files in the same directory.

3)   Higher-level Data: correlational, data fusion, decompositional.

These will include the transformational datasets mentioned in the section on Advanced Analysis. These datasets are to be linked to the raw and processed data directory. Simple annotation methods will confirm the identity.

4) Higher-level Representation: RDF/XML descriptive models, algorithmic (e.g. data landscapes, possibility spaces).

These types of representations can help us go beyond the typical reliance on “statistical significance” and “future directions” to provide a rigorous approach to guide future investigations. An example of this is parameterization models from existing data.


E)   Blogging Publicity.

All materials should be promoted through a blog post. This can be in the form of a feature article, or as a series of annotated links. This can be followed up with reposting key features of the initial post to a social blog like Tumblr or sharing a link via Twitter.


F)   Peer Commentary.

While this is typically kept confidential, there are so-called post-peer-review venues that provide a means to review work (e.g. PeerJ, F1000). This includes both formal (actionable) statements and informal statements in the form of critiques. 


This outline represents the entirely of a scientific reporting pipeline (from formal write-up to published items), although I am no doubt missing something. I will be fleshing each of these points out in future posts with real data and examples from Orthogonal Research and my work at the University of Illinois.

January 5, 2015

The Flow of Time, Science, and Archives

Here are a few milestones and interesting items to report for the New Year:

2014 was a good year for both space science and science in general. It's safe to say that the biggest story in science for 2014 was the successful landing of a probe (the Philae lander) on the surface of a comet (67p) by the European Space Agency. But the year was also not without dissapointments. Overall, many breakthrough findings and excellent papers occured in a number of fields. In the years ahead, it will be interesting to see what kinds of advances are made in 2015 from emerging work done during 2014.

If you are tired of celebrating another New Year according to the Gregorian calendar, here is an article from Futurity to make you consider an alternative. While the focus in this article is on the Hanke-Henry Permenant Calendar, there indeed is more than one way of dividing up the time it takes our planet to make a complete revolution around the Sun. This may or may not include calendars from other cultures, of course.

First milestone: The 24-year-old preprint server [1] arXiv published its one millionth (10^6th) paper on December 29, just in time for the new year. Despite being around for a quarter century, arXiv has become the template for an open access publishing revolution. Originally founded by Paul Ginsparg in 1991 [2], the bulk of the million paper total reflects impressive growth in the past several years.

The lifespan of the arXiv in terms of growth over time and abundance of articles by field. COURTESY: arXiv and [2].

Second milestone: By the end of Monday, January 5th, Synthetic Daisies will have reached 120,000 readers. Much like the arXiv, the bulk of this growth has occured in the last few years. The blog was started in December, 2008, so I also wish the blog a Happy 6th Birthday [3].


NOTES:
[1] Tomaiuolo, N.G. and Packer, J.G.   Pushing the Envelope of Electronic Scholarly Publishing. Searcher, 8(9), October (2000).

[2] Ginsparg, P.   arXiv at 20. Nature, 476, 145-147 (2011).

[3] Is this actually possible, or is it more like worshipping a fetish? I guess for purposes of good form, I should create an avatar that represents "the blog".



July 21, 2014

Four Readings and an Open Science Argument

Here are some papers from my reading queue. Four readings on human culture, behavior, and evolution, and one feature (set of readings) on Open Science. 

Four Readings......

Here are four readings from the reading queue on human culture, behavior, and evolution. The picture below (only tangentially related to the first paper) is from [1].


The first paper [2] is on the genetic architecture of economic and political preferences. Using a SNP analysis, the authors demonstrate that such traits have a polygenicarchitecture (e.g. many genes, small effect size for each). Studies that are underpowered (and no one really knows what the appropriate sample sizes should be) can potentially generate many false positive associations between genes and behavior. Nevertheless, understanding the presence of key variants for social preferences might help us understand why some people seem to be inherently "liberal" or "conservative".

The second paper [3] presents us with a premise that equates (or perhaps confounds) the psychophysiology of political ideologies with the roots of more general ideological bias. Are we really looking at "natural" differences between liberals and conservatives? Or does this simply demonstrate that high-profile social issues with already polar liberal and conservative positions [4] are undergirded by strong emotional responses? The standard evolutionary psychology explanation is a bit contrived as well. But it goes well with the previous article.


Crossmodal and cross-cultural comparisons, unite! In this study [5], people from several different cultures were asked to make both "congruent" and "incongruent" associations between smells and colors. The authors come to the conclusion that cultural context through experience has both statistical (covariance) and semantic (linguistic) components.

The fourth article [6] is a gateway article to several recent studies in the area of neuroplasticity. The gateway leads to the work being done in the laboratory of Michael Stryker [7]. Learn about the "neural volume control knob" and much, much more.


.....And An Open Science Argument


Here are some additional readings on networking and open science from my reading queue. The first is a paper on the life-cycle of a preprint on the arXiv [8] The top image is Figure 2 in the paper. The other two readings advocate for the use of open access protocols and social media to disseminate research [9] and counter cultural biases towards keeping research behind laboratory doors [10].


NOTES:

[2] Benjamin, D.J. et.al    The genetic architecture of economic and political preferences. PNAS, 10:1073/ pnas.1120666109 (2014).


[4] Related to this is the concept of the news filter bubble. One recent paper on this phenomenon: Koutra, D., Bennett, P., and Horvitz, E.   Events and Controversies: influences of a shocking news event on information seeking. arXiv, 1405.1486 (2014).

[5] Ren et.al   Cross-Cultural Color-Odor Associations. PLoS One, 9(7), e101651 (2014).

[6] Stix, G.   Neuroplasticity: new clues to just how much the adult brain can change. Scientific American blog, July 14 (2014).

[7] Two notable publications:
a) Fu, Y., Tucciarone, J.M., Espinosa, S., Sheng, N., Darcy, D.P., Nicoll, R.A., Huang, J., and Stryker, M.P.   A Cortical Circuit for Gain Control by Behavioral State. Cell, 156, 1139–1152 (2014).

b) Niell, C.M. and Stryker, M.P.   Modulation of Visual Responses by Behavioral State in Mouse Visual Cortex. Neuron, 65, 472-479 (2010).

[8] Shuai, X., Pepe, A., and Bollen, J.   How the Scientific Community Reacts to Newly Submitted Preprints: Article Downloads, Twitter Mentions, and Citations. PLoS One, 7(11), e47523 (2012).

[9] Allen , E.   “All research should be OA”. We agree! ScienceOpen blog, July 14 (2014).

[10] Konkiel, S.   How to become an academic networking pro on LinkedIn. ImpactStory blog, April 24 (2014).

July 8, 2014

Contributions to the bioRxiv, Summer 2014

I have been busy finishing up some work done in the Cellular Reprogramming Lab between 2010 and 2013. These two papers were submitted to and rejected as a single paper from PLoS One (after two rounds of revision). They were subsequently split them into their wet-lab molecular biology (written as an extended protocol) and computational components (written as a more conventional manuscript) for publication on bioRxiv.



The first paper (wet-lab molecular biology) is called "Using Polysome Isolation with Mechanism Alteration to Uncover Transcriptional and Translational Dynamics in Key Genes", in which we explore the world of mRNA regulation during adaptive cellular processes. The first part of the title (polysome isolation) involves harvesting mRNA from the polysome (translation-related mRNA). Harvesting this in tandem with mRNA associated with transcription provide us with a direct comparison between the transcriptome (TST) and translatome (TLT).


The second part of the title involves administering drug treatments to fibroblast populations which have systematic effects on transcription and protein production. These treatments are called "mechanism alteration" because they mimic changes that occur in a dying or transforming cell.


The third part of the title involves looking at transcriptional and translational dynamics for key genes. One criticism of the combined paper involved the use of candidate genes instead of high-throughput data. High-throughput data is great if one can afford it. On the other hand, large datasets can leave you with more questions than answers, which might be particularly true of this work. 


These figures demonstrate the analysis of experiments which validate the polysome recovery technique and the effects of drug treatments (mechanism disruption) on both transcriptome- and translatome- related mRNA (TST and TLT, respectively).

The second paper (computational) is called "Modeling Cellular Information Processing Using a Dynamical Approximation of Cellular mRNA". There is also a Github repository that contains associated Matlab code and simulations. Work from the first and second papers have been presented previous to their bioRxiv release, notably at the Stem Cell and Regenerative Medicine Conference held on the Oakland University (MI) campus in 2012.


We will go through this title in backwards order this time. The last part (cellular mRNA) refers to its connection to the first paper. Gene expression measured at both the transcriptome (TST) and translatome (TLT) will be used to model the cell's general response to mechanism alteration. In this case, an assumption is made: fluctuations of both mRNA fractions and at multiple points in time represents a regulatory process. 

Thus, a first-order feedback model can be constructed, with a simplified set of feedforward, feedback, and decay components. While there are a multitude of mRNA decay pathways and processing functions, this model focuses on a much simpler abstraction: the path from DNA to protein with single sources of decay and feedback. Each fraction of mRNA can be represented as a point process controlled by inputs and outputs.


The middle part of the title (dynamical approximation) then refers to the simulation of mRNA dynamics using the model, its components, and biological data. The idea is to approximate meaningful trends at certain points in a biological process, which is expected to differ by gene and by model component. This is where the proof-of-concept nature of this paper is most evident. 

While a somewhat contrived means to approximate a complex biological process is used in both papers, the original application was to be for understanding the early stages of cellular reprogramming. However, it proved to be exceedingly difficult to go from iPS cultures to meaningful computational inference.

An example of the first-order feedback model. A: a graphical example of the model components. B: an example of activity among the components over time.

Finally, the first part of the title (modeling cellular information processing) is based on the interpretation of the model output. the notion of cellular information processing treats the regulation of mRNA as an information processing problem. That is, you have an input, a process, and an output. systematic noise can also be added to the model, depending on the application. The process itself (mRNA processing from DNA transcription to RNA translation) is a transformation of information. 

When a cell is challenged by an environmental stimulus or the need to change phenotype, information provided by mRNA can be operated upon in a number of ways (linear responses, accumulation, delayed responses). Three information processing principles are used to interpret phenomena such as linear decay, the sequestration of mRNA at either the TST or TLT, and the differential response among individual genes.


Finally, I must point out how cool the altimetric support is on the bioRxiv. Here is a screen shot showing the number of tweets, abstract views, and .pdf downloads for the wet-lab paper:


Almost as functional as the analytics data on Blogger (which is not saying much, but for a formal publication venue, it's pretty impressive). The people at Cold Spring Harbor Lab have done a good job on this, and are ahead of the people at arXiv on this. As it turns out, however, the arXiv has a principled policy on this. Whether viewership stats are a sign of vanity and worthy of scare quotes is another matter.


Printfriendly