I spent some hours last days browsing through Edward Tufte's nice book Visual Explanations. Although sometimes it gets on graphical issues way more complex than I normally need, it is a great material both on learning how to present results and on using data to support analytical thinking. So that I try to keep some of the lessons I just learned, I decided to keep a few notes here:
- On presenting data visually, ensure there is a scale and a reference for the reader
- It is often useful to look at data on scales one order of magnitude larger and smaller than the actual quantities
- Place data in an appropriate context for assessing cause and effect. This includes reasoning about reasonable explanatory variables and expected effects.
- Make quantitative comparisons. "The deep question of statistical analysis is compared to what?"
- Consider alternative explanations and contrary cases.
- Assess possible errors in the number reported in the graphics.
- In particular, aggregations on time and space, although sometimes necessary, can mask or distort data.
- Make all visual distinctions as subtle as possible, but still clear and effective. Think of elements in your displays as obeying degrees of contrast: if all (bg, axis, data, ...) have the same contrast, they'll all get the same attention. In particular, applying this to background elements clarifies data.
- Keep criticizing and learning from visual displays you find useful or not.
Paul Ehrlich has an exciting story on the SEED magazine (which I find worthwhile to follow) detailing past and recent progresses on the field of cultural evolutionism.
It is enlightening to follow Paul's journey through the parallels which can be made on the way our culture and our genes evolve as a response to our environment. Even more interesting, however, is his discussion on the differences between these phenomenons and his experiments investigating these differences using canoe-building techniques as a case study. His final remarks do a good job in summing up it all:
We directly tested a theory of cultural evolution. Our work has helped to uncover a piece of the larger, more complex process of culture change and has shown that it is reasonable to think of that change as evolution. Natural selection can operate in cultural evolution as well as in genetic evolution. Though canoe features may not be related to the genetic attributes of people who construct and use them, nor is natural selection likely the central force in cultural evolution, a comprehensive view of cultural evolution does now seem possible. And despite the daunting complexity, I believe we will one day understand how cultures evolve, and that it will help us all to survive.
We move reasonably differently from albatrosses and monkeys. An impressive work on Nature this month uses a trace of a very large number (6 million) of cell phone users to model patterns in the mobility of human beings. The following is an excerpt from the abstract of it("Understanding individual human mobility patterns", by Golzález, Hidalgo and Barabási) and to the side is a link to one of their graphs just because it looks cool:
We find that, in contrast with the random trajectories predicted by the prevailing Lévy flight and random walk models7, human trajectories show a high degree of temporal and spatial regularity, each individual being characterized by a time-independent characteristic travel distance and a significant probability to return to a few highly frequented locations. After correcting for differences in travel distances and the inherent anisotropy of each trajectory, the individual travel patterns collapse into a single spatial probability distribution, indicating that, despite the diversity of their travel history, humans follow simple reproducible patterns. This inherent similarity in travel patterns could impact all phenomena driven by human mobility, from epidemic prevention to emergency response, urban planning and agent-based modelling.
Most interesting, Nature published on the same issue an editorial which, although praises this paper, discusses an interesting aspect of the modelling approach it takes:
To some extent this 'physicalization' of the social sciences is healthy for the field; it has already brought in many new ideas and perspectives. But it also needs to be regarded with some caution.
As many social scientists have pointed out, the goal of their discipline is not simply to understand how people behave in large groups, but to understand what motivates individuals to behave the way they do. The field cannot lose focus on that — even as it moves to exploit the power of these new technological tools, and the mathematical regularities they reveal. Comprehending capricious and uncertain human events at every level remains one of the most challenging questions in science.
A recently launched effort called scientists without borders is trying to ease networking between scientists in poor countries and their fellows worldwide.
This certainly has the interesting potential of peering scientists in developing countries with people in the centers of excellence in their fields around the world. Nevertheless, I think there are exciting possibilities also in easing developing-world researchers to discover each other. At least in computer science, it happens often that most research we have access to is that which is legitimated in US conferences, where US-made research prevails.
I believe increasing acknowledgment of research being done in countries which have similar issues can allow new initiatives and approaches which have a genuinely developing world perspective.
I've been recently investing time in using activity similarity graphs as tools to understand the structure of sharing in BitTorrent and tagging communities in two works in collaboration with Elizeu, Matei and Adriana (all three of them have done interesting work on characterizing system usage patterns using similarity graphs in the past).
I'm still toddling on this, but some of the graphs we're looking at are pretty big, what calls for graph visualization tools which are both versatile and efficient. I've been playing with two which are quite interesting:
This was made with GUESS, which is quite easy to use and has the great functionality of understanding python scripts to interact with the graph: 
The problem I ran into was that for very large graphs, the fact that GUESS does its rendering in a background thread renders it very uninteractive (you don't really know whether what you asked the GUI to do is going to take long, and you don't have a button to cancel it).
I then found Cytoscape, which deals very nicely with very large graphs and even has some nice plugins for doing topology analysis. I only managed to plot this in Cytoscape: 