Showing posts with label ecology. Show all posts
Showing posts with label ecology. Show all posts

Monday, November 26, 2012

What is Your Opinion About Sustainability in the Computing Curriculum?

I'm making final edits to my blurb on Sustainability for a panel presentation description about social and professional issues in the computer science curriculum*. More precisely, I'm thinking about conflict. There isn't much time when you are on a panel. What to focus on, what to focus on...so many choices and I'm conflicted. And oh boy, so will be some of my audience. Conflicted. Perhaps many of them, if past performance is any predictor of future performance. Which market analysts remind us all the time is not the case.

Yet, enlightened educators and psychologists tell us about the beneficial opportunities for managed conflict. Not the kind where you duke it out and slug your neighbor, but the kind where something provocative lands in your lap and you wrestle with it in a civilized manner as a group.

Sustainability in the computing curriculum is my little piece of the panel presentation**. When I wrote my original blurb I ended it with "Lisa will discuss the sometimes controversial sustainability knowledge unit in the social and professional practice knowledge area".

One of the anonymous reviewers asked: what is controversial? I wasn't sure if s/he was positively inclined and surprised by the statement or didn't know much about the issue and was just curious. I am 99.99% sure the reviewer was not negatively inclined towards the idea of sustainability in the computing curriculum. Anyone who gets all p.o.'d about the idea knows they are in conflict with a growing movement.

Another reviewer suggested that I bring up to speed members of the audience who are not familiar with the fundamental issues. In light of the first reviewer, this makes good sense. Ok, will do - if you wrote that and are reading this, then yes, I will make sure when I speak to cover the fundamentals for those who are not already deeply embroiled in everything.

And embroiled many people are. I was momentarily surprised to read the question asking what is controversial about infusing sustainability into the undergraduate computing curriculum. Perhaps because I routinely encounter professional colleagues who have strong opinions on the matter. In prior outreach on this issue I have encountered everything from:

"Thank goodness AT LAST this issue is being taken seriously!" 
to 
"Oh no, not ANOTHER mandate being shoved down my throat!"

(Mandate? Mandate? They are called "recommendations" for a reason).

There are also people in the professional community of computing education who are curious, curious, to hear about what the controversy is all about. Not ready to bite my head off nor to shower me with roses. Just curious.

I am reminded by this reminder that rather than presuming either roses or rotten tomatoes when I speak next March, I can view this as a micro-classroom opportunity. Perhaps challenge the crowd with comments such as these:

Sustainability is part and parcel of computer science and you ignore it at your peril

Sustainability is more than reducing your electricity load - which we suck at by the way

The "solution" isn't sending your old electronics off to a developing country for recycling  and patting yourself on the back

I believe all of the these, and I could continue with some evidence, but the point isn't (and won't be) for people to sit and take solemn notes about the pros and cons and the logic of it all. The whole point here will not be for me to talk talk talk but to get people off their comfy little conference hall chairs and engaging with the challenge of sustainability in their classrooms. What a panel can provide is an opportunity for constructively dealing with a difficult, challenging, conflicted topic with one's peers. In person. Where it is a lot harder to flame someone.

If you are a computing educator and think sustainability in the classroom is great stuff but haven't overcome the challenges of curricular rubber hitting the road let's all wrestle with your excitement and questions.

If you are a computing educator and think infusing sustainability in the classroom is silly or impossible let's all wrestle with your skepticism.

If you are a computing educator and not sure what you think - even better. I would like to put you right smack in between your opinionated peers and let's all talk about it.

*The Panel will be presented at the SIGCSE 2013 Symposium in Denver, Colorado and is called: "Computer Science Curriculum 2013: Social and Professional Recommendations from the ACM/IEEE-CS Task Force".

**My fellow panelists and wonderful colleagues are: Beth Hawthorne - bravely moderating this adventurous panel, along with Flo Appel and Carol Spradling, both of whom are battle seasoned veterans of the social and professional issues world of computing.







Monday, February 6, 2012

The Data Provenance Project

It is a scientist's job to ask a lot of questions and to search for answers. Often this means collecting extensive data and studying it to generate meaning. Along the way, scientists may see an unusual output, such as described in the hydrology research project in my last post. To follow the trail leading eventually to our friend the moose, questions had to be asked: which sensor produced the anomalous data, when did this happen, what day, what time, how long did the change last, was there a similar rise in water level at nearby sensors in other streams? These answers come in the form of data: Provenance Data. Provenance comes from the French verb "provenir", meaning "to come from, to come forth".

But Provenance Data is not just for getting to the bottom of mysteries. Provenance Data is a key contributor to proving and justifying scientific conclusions.  This is where the challenging concept of "raw data" comes into play. What data exactly are we talking about when we ask for the "raw data"? How do we present that data to others such that it has meaning, given that a bunch of numbers without any interpretation is often meaningless? But once we interpret (manipulate) it, is it still "raw"? It is easy to get trapped in a circular conundrum.

An Example, returning to the hydrology project: is raw data the data about stream outflow at a given location? This outflow information has meaning, but was generated by a synthesis and filtering of other data. So, is raw data the average water weight generated every 15 minutes at various onshore loggers? Maybe. We can get yet more specific: is raw data the individual underwater sensor readings taken every few seconds? Maybe...but at this point would those readings make any sense to anyone other than a few highly trained specialists and engineers?

Probably not.  So how helpful would it actually be in proving and justifying claims of stream outflow volume to the concerned external evaluator or critic? We didn't even discuss the fact that there are enormous technological hurdles to maintaining every single sensor reading for any length of time. Not to mention that if you ask an ecologist they would probably present even more alternatives for the title of "raw data".

More than ever, in this day and age of constant challenging and questioning of scientific claims, something is needed to assist with obtaining a full picture of where results come from and what they mean.

As explained to me by Barbara Lerner, computer science faculty at Mount Holyoke College, Provenance Data is useful for answering many questions related to understanding, validation and accountability: to provide tracking of data, to enable a study of interacting actions inherent to any complex process, to facilitate investigation of deeper and broader questions generated by data inherent to complex processes.

Barbara is part of the multi-institutional Data Provenance Project which is developing a process system to aid scientists in collecting, storing and analyzing Provenance Data. She works with faculty at the University of Massachusetts at Amherst (Lee Osterweil) and at Harvard Forest (Emery Boose). The tool they are creating will provide a disciplined method to track how and when data was collected, and how it has been manipulated, all the way through to the development of descriptive models.  There are applications in diverse domains; her focus is the Harvard Forest ecology project measuring stream volume outflow we have been discussing. When the project is complete, the ecologists her team works with will be able to extensively query and manipulate their data - without having to learn a query language such as SQL. The current prototype is already able to produce Data Derivation Graphs (DDG) for the scientists.

Here is a very simple example of a  DDG describing the process for obtaining one stream discharge value, using a specialized processing language called Little-JIL:

Detailed explanations can be found in the team's published papers.*

There are challenges on many levels to building a Data Provenance tool. One of the biggest concerns is with balancing technical flexibility with ease of use for the non computer scientist. For this reason the computer scientists work closely with the ecologists, who think this project is "cool" and are happy to provide ongoing feedback. There are other challenges: those inherent to graph problems in general; all sorts of challenges to developing process systems that will be functional across disciplines. Other areas of interest range from processes tied to climate modeling, emergency room care, chemotherapy delivery and labor negotiations. Clearly, the long term benefits extend far beyond the ecology project. Theoretically, any science process, research or otherwise, will be able to use this system once it is fully developed.

As Barbara Lerner says, it is extremely rewarding to do outward looking things and obtain concrete results.  It is inspiring to work with other scientists who think this work is exciting. The field of computer science benefits, the overall cause of science benefits, and society benefits. Hard to argue with any of that.



*Barbara Lerner, Emery Boose, Leon Osterweil, Aaron Ellison and Lori Clarke, "Provenance and Quality Control in Sensor Networks", Environmental Information Managemet 2011 Conference, Santa Barbara, California, September 2011.

Wednesday, February 1, 2012

Preparing for the Unexpected Moose in Your Hydrology Research Study

Let's say you are an ecologist studying a watershed high in the mountains. You care about the water flow through several streams that feed a pristine lake. Ultimately you want to understand the stream discharge process in this region - what is the volume? You need to collect a lot of complex data in order to build a realistic model of what happens day in and day out.

Water can enter the environment several ways, including rain and another body of water; water can exit the environment in several ways including evaporation, entering another body of water, or seeping underground. So you create small dams and place sensors in the water at well chosen locations. Each sensor measures the weight of the water (among other things) and feeds that data every few seconds to a data logger on the nearby shore. The data logger computes an average every 15 minutes and saves those values for you. Every so often you trek up the trail to your sensors, Palm Pilot in hand, download the data, take it back to the lab.

(Compressing the description of the scientific process for purposes of brevity) Run statistical analyses on the data, generate defensible behavioral models, write up the results and publish them.

Until the day that you notice a very strange reading. The water level is suddenly unusually high. Why might this be...

By running a few standard checks and conducting a little investigation you discover that a moose stepped in the water. If you are like me, when you first heard this all too real scenario, you almost fell off the chair laughing at the thought of a moose blithely wandering into the middle of a serious research project.

One unexpected and undetected moose could really mess up your data driven model of stream flow. Fortunately, the moose is reasonably easy to figure out. But other scenarios are a lot harder to get to the bottom of when you are dealing with complex natural phenomena and processes. What if you are measuring and modeling atmospheric carbon flow and sequestration in trees over that same expanse of forest? What if you include variables related to climate change, which is sure to bring in-depth scrutiny from peers and critics? You absolutely need to be able to explain and justify your conclusions to science and perhaps even to the wider public.

What you need is Provenance Data: the data about the data; the meta-data, whatever you want to call it. Provenance Data is the data that describes how those stream values were obtained, when they were obtained, what was done to that data. The contextual information surrounding the so-called Raw Data.

Computer Scientists are involved in a series of research projects to enable the gathering and clear presentation of Provenance Data. Next post, I will explain what they are doing, as well as why I said "so-called" Raw Data.