Showing posts with label Earth Science. Show all posts
Showing posts with label Earth Science. Show all posts

Monday, February 6, 2012

The Data Provenance Project

It is a scientist's job to ask a lot of questions and to search for answers. Often this means collecting extensive data and studying it to generate meaning. Along the way, scientists may see an unusual output, such as described in the hydrology research project in my last post. To follow the trail leading eventually to our friend the moose, questions had to be asked: which sensor produced the anomalous data, when did this happen, what day, what time, how long did the change last, was there a similar rise in water level at nearby sensors in other streams? These answers come in the form of data: Provenance Data. Provenance comes from the French verb "provenir", meaning "to come from, to come forth".

But Provenance Data is not just for getting to the bottom of mysteries. Provenance Data is a key contributor to proving and justifying scientific conclusions.  This is where the challenging concept of "raw data" comes into play. What data exactly are we talking about when we ask for the "raw data"? How do we present that data to others such that it has meaning, given that a bunch of numbers without any interpretation is often meaningless? But once we interpret (manipulate) it, is it still "raw"? It is easy to get trapped in a circular conundrum.

An Example, returning to the hydrology project: is raw data the data about stream outflow at a given location? This outflow information has meaning, but was generated by a synthesis and filtering of other data. So, is raw data the average water weight generated every 15 minutes at various onshore loggers? Maybe. We can get yet more specific: is raw data the individual underwater sensor readings taken every few seconds? Maybe...but at this point would those readings make any sense to anyone other than a few highly trained specialists and engineers?

Probably not.  So how helpful would it actually be in proving and justifying claims of stream outflow volume to the concerned external evaluator or critic? We didn't even discuss the fact that there are enormous technological hurdles to maintaining every single sensor reading for any length of time. Not to mention that if you ask an ecologist they would probably present even more alternatives for the title of "raw data".

More than ever, in this day and age of constant challenging and questioning of scientific claims, something is needed to assist with obtaining a full picture of where results come from and what they mean.

As explained to me by Barbara Lerner, computer science faculty at Mount Holyoke College, Provenance Data is useful for answering many questions related to understanding, validation and accountability: to provide tracking of data, to enable a study of interacting actions inherent to any complex process, to facilitate investigation of deeper and broader questions generated by data inherent to complex processes.

Barbara is part of the multi-institutional Data Provenance Project which is developing a process system to aid scientists in collecting, storing and analyzing Provenance Data. She works with faculty at the University of Massachusetts at Amherst (Lee Osterweil) and at Harvard Forest (Emery Boose). The tool they are creating will provide a disciplined method to track how and when data was collected, and how it has been manipulated, all the way through to the development of descriptive models.  There are applications in diverse domains; her focus is the Harvard Forest ecology project measuring stream volume outflow we have been discussing. When the project is complete, the ecologists her team works with will be able to extensively query and manipulate their data - without having to learn a query language such as SQL. The current prototype is already able to produce Data Derivation Graphs (DDG) for the scientists.

Here is a very simple example of a  DDG describing the process for obtaining one stream discharge value, using a specialized processing language called Little-JIL:

Detailed explanations can be found in the team's published papers.*

There are challenges on many levels to building a Data Provenance tool. One of the biggest concerns is with balancing technical flexibility with ease of use for the non computer scientist. For this reason the computer scientists work closely with the ecologists, who think this project is "cool" and are happy to provide ongoing feedback. There are other challenges: those inherent to graph problems in general; all sorts of challenges to developing process systems that will be functional across disciplines. Other areas of interest range from processes tied to climate modeling, emergency room care, chemotherapy delivery and labor negotiations. Clearly, the long term benefits extend far beyond the ecology project. Theoretically, any science process, research or otherwise, will be able to use this system once it is fully developed.

As Barbara Lerner says, it is extremely rewarding to do outward looking things and obtain concrete results.  It is inspiring to work with other scientists who think this work is exciting. The field of computer science benefits, the overall cause of science benefits, and society benefits. Hard to argue with any of that.



*Barbara Lerner, Emery Boose, Leon Osterweil, Aaron Ellison and Lori Clarke, "Provenance and Quality Control in Sensor Networks", Environmental Information Managemet 2011 Conference, Santa Barbara, California, September 2011.

Wednesday, February 1, 2012

Preparing for the Unexpected Moose in Your Hydrology Research Study

Let's say you are an ecologist studying a watershed high in the mountains. You care about the water flow through several streams that feed a pristine lake. Ultimately you want to understand the stream discharge process in this region - what is the volume? You need to collect a lot of complex data in order to build a realistic model of what happens day in and day out.

Water can enter the environment several ways, including rain and another body of water; water can exit the environment in several ways including evaporation, entering another body of water, or seeping underground. So you create small dams and place sensors in the water at well chosen locations. Each sensor measures the weight of the water (among other things) and feeds that data every few seconds to a data logger on the nearby shore. The data logger computes an average every 15 minutes and saves those values for you. Every so often you trek up the trail to your sensors, Palm Pilot in hand, download the data, take it back to the lab.

(Compressing the description of the scientific process for purposes of brevity) Run statistical analyses on the data, generate defensible behavioral models, write up the results and publish them.

Until the day that you notice a very strange reading. The water level is suddenly unusually high. Why might this be...

By running a few standard checks and conducting a little investigation you discover that a moose stepped in the water. If you are like me, when you first heard this all too real scenario, you almost fell off the chair laughing at the thought of a moose blithely wandering into the middle of a serious research project.

One unexpected and undetected moose could really mess up your data driven model of stream flow. Fortunately, the moose is reasonably easy to figure out. But other scenarios are a lot harder to get to the bottom of when you are dealing with complex natural phenomena and processes. What if you are measuring and modeling atmospheric carbon flow and sequestration in trees over that same expanse of forest? What if you include variables related to climate change, which is sure to bring in-depth scrutiny from peers and critics? You absolutely need to be able to explain and justify your conclusions to science and perhaps even to the wider public.

What you need is Provenance Data: the data about the data; the meta-data, whatever you want to call it. Provenance Data is the data that describes how those stream values were obtained, when they were obtained, what was done to that data. The contextual information surrounding the so-called Raw Data.

Computer Scientists are involved in a series of research projects to enable the gathering and clear presentation of Provenance Data. Next post, I will explain what they are doing, as well as why I said "so-called" Raw Data.


Monday, November 21, 2011

The Computer Sand Society

This week begins the holiday season for many people and often with it, unfortunately, way too much stress. Maybe that is why this week (Thanksgiving week in the US) brings out some strangeness. There was the blog post forwarded by a friend about stuffing a turkey with twinkies. It is hard not to feel fondness for something that springs back into shape when you step on it, but ... make a meat glaze out of the so-called "creme filling"? Yuck.

I also received an advertisement in the mail addressed to me at "The Computer Sand Society". Walking back from the mailbox, I thought: this could be a new interdisciplinary application! In response to this inspiration, a friend sent a lovely video link about sand animation; another friend suggested that, nice as it sounds, poolside would be much better than beachside because sand is hard on laptops. Could it be worse than three bouts of sick video cards?  Perhaps, yes it could. Although I once told a cell carrier that my phone had mysteriously died when it had in fact fallen in the toilet (would you want to explain that one?), I'm not sure the computer manufacturer would buy into the notion that a gritty substance floating around the motherboard stemmed from disintegrating integrated circuits.

However, as I do believe in the power of creative thinking to spur innovation, I suspect there is opportunity for The Computer Sand Society to come into its own.

Modeling and simulation of sand castles. Has anyone developed a system, similar to those used by architectural design firms, to analyze the possibilities for ever more complex creations, factoring in the properties of sand - fineness of particles, distribution of various well crumbled crustacean shells, positioning relative to the high tide mark, mineral components?

Sand castle building is serious business for some people. The U.S. Open Sandcastle Competition bit the dust this year and its demise has many people very upset. I wonder if profits from my envisioned application might have helped hold it together? We could have perhaps drawn on the nearby expertise of the famous Scripps Institution of Oceanography, NOAA (the National Oceanographic and Atmospheric Administration) and the local surfers who are out every morning, afternoon and evening rain or shine, 365 days a year. Who understands the interactions between sand and surf better than these subject matter experts?

I have always wanted to combine my love of the outdoors with the potential of computing. A typo by some overworked marketing employee has given me the inspiration for a new hi-tech startup. All that is needed now is a really dedicated team to get it off its feet - and an angel investor.

Any takers?

Happy Thanksgiving - have some fun and forget the stressful stuff for a few days.


Thursday, March 17, 2011

Computing has an Important Role to Play in Earthquake Preparedness and Response

When devastation as large as that currently happening in Japan occurs, it can be hard to know what to say or do. If you are like me, you have been reading the news daily (or more often), caught up in a mix of complicated reactions. This morning for example I watched computer generated weather simulations of possible flow patterns of radioactive contamination (via the BBC).  As the simulation looped over and over I couldn't help but be transfixed by one large multicolored plume as it slid like a mutant amoeba over Southern California. Right here in other words. The colors registered different levels of radiation. Computers generated those simulations and unsettling as they were, I'm glad to be able to see them. It is better to have knowledge from a reliable source than no knowledge, even when that knowledge is based on probabilities and a great deal of the unknown.

I was very grateful for computer science when the recent earthquake struck New Zealand. A friend lives in Christchurch and it was only a matter of a few nerve wracking days before a brief post appeared on Facebook telling all of us that she was ok - no doubt considerably freaked out, but ok. Thank you to the computer scientist creators of social networking.

The situation was very different in 2004 when the tsunami hit Sri Lanka and someone I know was very near the coast. It was over a week before we learned that he and his family were alive. There was no email access, no smart phones, no Facebook page, nothing but waiting and telephone calls to the US State Department (who were terrific by the way). The computing communication infrastructure either was not there to start with or had been completely disabled by the dual natural disasters of earthquake and tsunami.

Watching the developing situation in Japan, the triple disasters unfolding as nuclear contamination possibilities are added to realities of earthquake and tsumani, watching the weather models, I was reminded of the researchers around the world who work full time developing models of earthquake simulation and  who perform seismic hazard analysis. They work on these models so that we can know as much as possible about what can happen, how it can happen, how we can best prepare, where to erect buildings and other structures and how to protect them as best we can.

Developing 3-D and 4-D maps and models are classic computing problems of large scale data analysis: selecting and applying the "best" constraints, knowing that the model you develop will depend upon choices about possible epicenter (location of the earth directly above the underground origin of the earthquake), focal depth (how deep the origin is), magnitude (amount of energy released) and possible paths the seismic waves may follow. There are innumerable factors to include or leave out of this type of model such as local and regional variations, ground type, land masses, rock type...just for starters. It is all about improving probabilities and predictions.

The paths of seismic waves are not always what one might expect. For example, one reason Los Angeles gets hit so hard by some earthquakes on the famous San Andreas fault is because there is a natural "funnel" that directs ground motion directly into the city from a section of the fault well east of the city. Complex modeling and a solid knowledge of the land revealed this important information. 

You have to know your hardware, firmware and software; you have to know how to work with the latest and most sophisticated networks of high performance computing. Operating system, algorithm and programming language optimization. Databases to hold all those data and sophisticated networks to link the distributed grids of computers.

If you have an interest in earth science and scientific computing (you don't need expertise in both - this is where collaboration between fields comes in) then here is an area where you can work to make a difference in people's lives.