Saturday, November 13, 2010

Risk as a population concept

An interesting theme that has emerged in the RURS project is the distinction between risk as a population level concept - the probability of an event occurring - and uncertainty as an individual level concept - does the event happen to me, personally. This first arose in the Psychology session, and also in the Political Science session. Andrew Gelman (a statistician) made a similar observation while commenting on an article about medical trials and ethics:
As a doctor, Elliott focuses on individual patients, whereas, as a statistician, I've been trained to focus on the goal of accurately estimate treatment effects.
It would seem that this idea has some legs, if not much in the way of actual work devoted to it.

Foodweb theory solves the Afghanistan Problem

Ecology has a long history of borrowing nifty ideas from other disciplines and making theoretical hay out of them - just think game theory, optimal foraging etc. So it's pretty neat to see the arrow of theory pointing the other way: here's a short video on how food web theory can help US strategy in Afghanistan.

Thursday, November 11, 2010

What is an experiment?

The other day the meaning of the term "experiment" was called into question. It matters, because a student and I have a paper in press in which we use the emphasis placed on experimentation to distinguish between two schools of thought in Adaptive Management. I don't want to pre-empt Jamie's paper here, but I did want to talk about what I think an experiment is.
Well, according to Wikipedia, an experiment is "...the step in the scientific method that arbitrates between competing models or hypotheses." (Aside: it is interesting that there is a distinction made between model and hypothesis - for another time perhaps.) OK, I can't see anything wrong with that, but we need to dig a little bit deeper. There are two additional attributes that are important for distinguishing between methods of arbitrating among hypotheses: the number of simultaneous experimental treatments and the amount of replication within treatments.

The number of simultaneous experimental treatments is fairly obvious - how many different manipulations of the system under study are in use? This could range from one (an observational study of existing conditions) to many (a laboratory study with positive and negative control treatments and a dose response). Is the term "experiment" appropriate across this entire range?

The second attribute is the amount of replication within a treatment - in how many different places and times was the effect of the treatment observed? This too can range from one to many.

I believe that the term experiment is appropriate when the number of simultaneous experimental treatments is greater than one, regardless of how much replication is present. Replication does matter, but it doesn't affect the ability of the experimenter to determine causation. Rather it affects the scope of the causation - with only one replicate per treatment it is not possible to generalize beyond the set of objects studied. Within that set it is still possible to determine if a hypothesis is consistent with the data, and attribute the differences between treatment responses to the manipulation.

This attribution of causation is the reason for the treatments to be "simultaneous", because this reduces the extent of unmeasured differences between the observational units. Simultaneous has the usual temporal meaning, but also carries a certain spatial component. Clearly, two patches of grassland on different continents are unlikely to serve as reasonable replicates of each other - there are simply too many things changing. However, two grassland patches in the same ecoregion may well work for comparing different burning practices.

In AM, the idea of an experiment is to use the management action itself to create the experimental treatments, and in that case the desire to determine causation beyond the current set of objects (e.g. management sites) is less important than figuring out which management actions work the best. An experiment will figure that distinction out quicker than applying treatments sequentially to a single object, because the simultaneity of the treatments helps to reduce the number of alternative explanations.

While it is true that society often conducts large scale manipulations of ecosystems without simultaneous alternative treatments, I do not believe it is helpful to describe these manipulations as experiments. If we do, then everything is an experiment, and the word ceases to have any value, much like the word sustainability or indeed, adaptive management. An experiment with simultaneous treatments is not the only way to distinguish between competing hypotheses, but it is a very good way when it is possible.

Speaking Romulan

A big part of the RURS exercise is about building interdisciplinary understanding of common concepts - breaking down jargon. Here is a neat story about why that matters from Christopher Reddy, a chemist that worked on the Deep Horizon oil spill.

Thursday, November 4, 2010

RURS - Psychology Edition

The RURS team hit the psychology department today, and a stimulating discussion was generated by the 18 participants, a mix of students and faculty. One interesting feature of this group was the presence of participants with professional responsibility – clinical psychologists – and this group had some very interesting things to say about risk.

Participants, both clinicians and others, inevitably discussed risk in the context making a decision such as whether to admit a client to hospital or send them home. A second kind of paradigmatic decision involved visual detection tasks, such as examining a radar screen or x-ray machine looking for bombs in luggage. In all cases they clearly identified two kinds of errors, false positives and false negatives, and associated different possible negative outcomes with each error. Participants also discussed critical thresholds in this decision making context: when an indicator of a risky behavior moves past a critical threshold, then a clinical action will be taken.

Exactly how such critical thresholds arise was a topic of some discussion. Data are often continuous, but for convenience are broken into categories for display and analysis. In some cases these arbitrary breaks become decision thresholds by default.

There was much discussion of “implicit cognition inaccessible to verbal processing”; decision making where people have trouble articulating how their decisions are reached. The opposite type of decision making involves consciously analyzing a set of steps to reach a conclusion. This distinction was reflected in discussion about the utility of information from the population scale versus the single individual scale.

Participants distinguished between population scale trends in the likelihood of an adverse outcome, estimated from actuarial data on large populations, versus the clinical setting where a single individual is being treated. Regardless of population level risk factors that are present in that patient, there is uncertainty about the particular outcome for that patient. For example, the best population level predictor of immediate suicide risk is a previous suicide attempt. However, that particular information does not rule out either outcome for the patient right now. Thus clinicians rely more on qualitative heuristics, i.e. implicit decision making, including the situational context for a patient. The context of the patient is important source of uncertainty, because many aspects of that context are unknown. This notion that the exact future outcome is not known was the dominant definition of uncertainty for this group.

This contrast between the population level and the individual level was expanded on in a discussion of how risk factors are used in diagnosis – “Not all risk factors are created equal.” For example, diagnosis of Attention Deficit Hyperactivity Disorder (ADHD) is done by examining a list of “neurological soft signs” that vary in their predictive ability at the population level. Simply adding up the number of these signs that are present and using that as the heuristic guide leads to over-prediction of ADHD, whereas only using a single strong predictor would lead to under-prediction of ADHD.

Participants described risk as two dimensional – the likelihood of an event and the magnitude of the adverse outcome. They also mentioned that adverse outcomes are difficult to define quantitatively, which creates uncertainty. A participant offered an anecdote illustrating these two dimensions – the tornado and the trick-or-treaters. A trained storm spotter was dispatched to examine a cloud on Halloween – a time of year at which tornadoes are rare. On arrival, the spotter noted “… a rotating wall cloud …”, an indicator that a tornado was possible, although it appeared weak. However, there was a nearby town with many trick-or-treaters out on the streets, so even though the likelihood of an event was low, the potential adverse outcome for even a small storm was great. The decision was made to activate the tornado alarms in the town.

Another repeated point was that the consequences of errors are a shift in critical thresholds, affecting the sensitivity and specificity of decision makers. For example, viewing a radar or x-ray screen for a long time reduces sensitivity, increasing false negative decisions. This can be mitigated by changing personnel regularly. A second example involved what happens during the training of security screeners. After a screener misses a simulated bomb, i.e. makes a false negative decision, their rate of false positive decisions increases – they overcompensate. This occurs even though the actual adverse consequence is very small. A poignant additional example with a larger adverse outcome was offered by a clinician – “You never forget your first suicide.”

Participants also made a distinction between risk to an individual patient, to the clinician, and to third parties. Suzie Q may be suicidal, which creates a risk to her, but if she is also homicidal then this creates a risk to third parties. This third party risk creates additional uncertainty, because who the third parties are is unknown to the clinician.

Risk to the clinician arises primarily from accountability – if the client injures themselves or someone else, is the clinician legally responsible? The concrete example offered involved a clinician treating a couple, and during the treatment it becomes clear to the clinician that domestic violence is an issue. The clinician is not legally obligated to report the domestic violence, and so is not accountable. However, if there is a child in the home, then there is a legal obligation to report the possibility of child abuse, creating a risk to the clinician if the potential is not reported.

Participants identified an additional trade-off between resource need and availability in the face of uncertainty – for example there are not enough hospital beds for everyone who meets a given level of homicidal tendencies. This was one area where participants agreed that population level data had a role to play, in figuring out whether resources allocated to particular needs were sufficient.

Some participants had studied the Anterior cingulate cortex, and found that pretty important and fascinating – although they didn’t expand on it. (That was one of those times when one is abruptly reminded that interdisciplinary work is hard and takes time!) Participants also raised the observation that the ability to perceive and act on risk is something that develops over time – teenagers are particularly bad at it – and in addition that studies show this ability to be variable among people, and genetically heritable.

Participants identified some additional sources of uncertainty arising from data. Measurement uncertainty arises because psychological instruments don’t measure underlying constructs exactly. Alternatively, relevant information on a risk factor or of a client’s context may be missing, creating additional uncertainty. In a slightly different context, there is a desire to be able to eliminate human judgment from risky decisions, for example by using functional brain imaging to detect if someone is being deceptive. This could create a false sense of security, which would be unjustified because of the inability of the instrumentation to attribute a given response to a particular cause in the subject – it is difficult to operationalize the assessment of risk.

Tuesday, November 2, 2010

RURS - Sociology Edition

As I am still recovering from my recent transcritical bifurcation (well, more of a multiple perforation, but its done now), Sarah Michaels contributed the following guest post on yesterdays visit to the Sociology department by the RURS team:

Sociology was the first social science stop in the Risk and Uncertainty Road Show. What came through in the discussion was the concern sociologists had for the implications of what they were doing. There was the major push in the discipline to be responsible to the population being researched. The sociologists were acutely aware of the potential for harm from how they dealt with risk and uncertainty as academics researching at risk populations. The engagement of academic sociologists in “real world” research sets up tradeoffs and conflicts. One of the first tradeoffs is the pursuit of “Truth” can be in tension with the concern for what is of value to the population being investigated. For example, while academics may be interested in causality, the community may be more interested in solutions or effects of the problems being investigated.

The sociologists identified the risk of getting the story wrong, of not fully understanding the social dynamics they were trying to uncover. At the same time, they were concerned with the risk of their findings being intentionally or unintentionally misinterpreted and that such a misinterpretation could be used to harm a community. They were also mindful that incorrect and/or inappropriate conclusions or implications would be drawn from their research.

Sociologists were concerned about the risk of asking the wrong questions. This could take the form of asking questions that the community being investigated didn’t think were important or that the questions asked might be heavily biased or potentially harmful to respondents.

Sociologists may face risk in conducting their research. This may involve risk of physical danger from working with a violent population, risk of going to jail for not sharing data generated and risk of harm to one’s career from studying controversial topics, marginalized populations or generating controversial conclusions.

Sociologists noted that uncertainty arose from bias. They were particularly mindful of the bias students perceived in course instruction and observed that it was easier to sort out bias in research than in teaching. The sociologists recognized bias in the questions they pose and the inherent uncertainty in what they do. In response to the latter they are careful not to speak beyond their data.

Sociologists confront uncertainty that stems from the gap between “Truth” and what they observe. As such, sociologists face uncertainty of measurement. For example, since they are well aware of how problematic self reporting by respondents is they use multiple measurements.

Two critical thresholds were highlighted. The first was ideally to do no harm in conducting research and at a minimum to try and make sure benefits outweigh harm. The second was the subjective boundary between ethical and unethical research behavior.

Saturday, October 16, 2010

Critical Slow Downs

Ugh - Saturday morning and I'm feeling critically slowed down. Hopefully, its not an early warning of an approaching transcritical bifurcation! John Drake and Blaine Griffen have a really nice paper in the Sept 10 Issue of Nature on Early warning signals of extinction in deteriorating environments. This work follows up on the ideas of Stephen Carpenter, Will Brock and Martin Scheffer from the Resilience Alliance of trying to detect when a system is approaching a "critical threshold" before the threshold is reached. This would be useful, to say the least.
Drake and Griffen used an experimental system of Daphnia magna in microcosms to examine what would happen when populations experienced deteriorating environmental quality - in their case represented by a decrease in the amount of freeze dried algae supplied. In an environment with a constant food supply, this system behaves as a logistic population with dn/dt = rn(1-n/k), where both r and k change as the amount of food changes. The bifurcation happens as r -> 0 from either direction. The notion of "Critical Slowing Down" is a dynamical phenomenon that occurs in the vicinity of the threshold - increases in temporal autocorrelation, spatial correlation, and coefficient of variation. They found that a variety of single and composite indicators showed responses several generations before the "tipping point" was reached. Wow!
And I do mean, WOW, really. This is cool stuff. What I'm not sure is whether it will help in the real world. The difficulty arises because of the accuracy of the measurements in the experiment. They know exactly how many individuals in each population. In most cases, we will have only statistical estimates of that number, and usually, only an index to abundance. In addition, we will rarely have the luxury of data that precede the deterioration of the environment, and usually only one population instead of 30 replicates. Still, its worth thinking about more, for sure.