Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Thursday, November 1, 2012

More Survey Analysis

I'm repeating my previous post, since it seems relevant to review the data again before the election. I accessed all these sites at about 9:00 PM CST. I

Real Clear Politics
Obama 201
Romney 191
Toss Ups 146
Obama by +1 ECV

Intrade Presidential Election 2012
Obama ~67% win
Romney~33% win

FiveThirthyEight 
Obama 303.2 +/-56 ECV
Romney 234.8 -/+56 ECV
~81% Obama wins

HuffPost Pollster

Obama  277 ECV
Romney  206 ECV
Toss Ups 55 ECV
Obama has enough certain electoral votes to win.

Obama 285 ECV
Romney 191 ECV
Toss Ups 62 ECV (but 44 of those lean strongly towards Romney)

Election Analytics 
Obama 296.7
Romney 241.3
~99.4% Obama wins

As before, I have arranged these in roughly increasing order of favor for Obama winning the 2012 election. The last three sites (HuffPost, TPM, EA) are making strong claims that Obama has the electoral votes to win already. 538 is not far behind that claim.
RCP seems to be sitting on the fence, not making strong claim about the toss-up States, and there is nothing wrong with that.

Other notes:
HuffPost Pollster has been added to the results. (but you saw that already.)

FiveThirtyEight: Last time I made the claim that Nate Silver's analysis is as close to neutral as can be found. Today though, I saw this: Nate Silver bets Joe Scarborough $1000 that Obama wins. It is not clear if this was intended as a partisan statement, or simply a good bet. It's not wrong to claim Obama is a good bet.

Should I update again tomorrow?

Tuesday, October 30, 2012

Survey Says ...

After discussing politics with a friend this morning, I was inspired to make the rounds of sites that accumulate political polling results and see what they all are saying. All results accessed on the web late afternoon on November 6th.

Real Clear Politics
Obama 201 +/-61
Romney 191 -/+61
Toss Ups 146
Obama by +10 ECV
(edit: those error bound probably belong to the 538 analysis, not RCP).

Intrade Presidential Election 2012
Obama 64.6% win
Romney 35.4% win

FiveThirthyEight 
Obama 294.6 ECV
Romney 243.3 ECV
~72.9% Obama wins

Obama 265
Romney 206
(Comment: therefore 67 are not certain)

Election Analytics 
Obama 291.4
Romney 246.6
~95% Obama wins

I have arranged these in increasing order of favor for Obama winning the 2012 election. Real Clear Politics and Talking Points Memo are sites I do not regularly follow, so I have no clear opinion of their methods. I suspect these sites have a conservative and liberal biases (respectively) but I have not direct evidence for that. Certainly if they are both presenting statistical results in an unbiased manner, they ought to have about the same conclusion. All I can really say here is that RCP doesn't offer much of a direct election forecast.

Edit: Intrade is another site I don't follow, but it has been getting a lot of media attention too.

I personally favor the methods Nate Silver uses on FiveThirtyEight, as I consider it to be the most technically sophisticated analysis, and it takes economic data into account as well. Silver offers a lot of day-to-day commentary on daily polling and his predictive model. Since the Democrats have lead the elections since the blog started in 2008, many of those prediction have been that Democrats are going to win. Some people interpret this an a liberal bias, but Silver's Senate predictions have been very good, and it not a bias if he is correct.  I have never seen Silver make anything close to an endorsement, and so on that basis I think FiveThirtyEight gives the most unbiased political analysis available.

Election Analytics is different from the others. It's a small academic page rather than a news site. Their statistical methods are sound, but they make some strong assumptions, which seems to be why they are able to make such a strong prediction for Obama winning. I haven't looked into just what these assumptions are, so I can't say if they are justified or not.

All this might not make my friend happy, but all the data is saying pretty much the same thing, with varying degrees on certainty.

Update: Added Intrade prediction at the suggestion of +Kevin Clift. Accessed the morning of 10/31/12.
Edit: next time I'll include this site too: http://electoral-vote.com/

Wednesday, February 8, 2012

Suddenly, Statistics in the Recall is in the News

I don't know that I can take credit for this, but this morning Phil Scarr at Blogging Blue points me to an article at the Milwaukee Journal Sentinel:
Analysis: Invalid signatures likely not enough to halt Walker recall

This seems to be just what I suggested in my previous post, two days ago. Whether or not I had the idea first, I applaud the effort.

[Edit: Fixed a link. I originally linked to a different but related post by Phil.]

Sunday, February 5, 2012

The Statistics of Verifying Recall Signatures

There is a huge political battle raging to Wisconsin, but you probably know this already. The drive to recall Governor Scott Walker has gotten plenty of media attention. Some 540,000 signatures are needed, and the Democrats turned in well over one million signatures to be verified. Walker supporters are hard at work trying to identify false signatures, to get as many petition thrown out as they can. That means about half of all signatures could be bad and the Democrats would still have enough to force a recall election. Given this rather daunting situation, how hard should Walker's supporters try? Do they really need to check 500,000 signature, a million signatures? How many is enough to be confident of a reliable result?

Regardless of your opinion about Scott Walker and the recall, some simple statistical sampling can help answer the question, and it requires checking far fewer than 500,000 signatures. I'm going to assume that each petition form contains 20 signatures, and 50,000 petition forms to total one million signatures. Essential to this process is a simple random sample**, where we can select a sample of petitions so that each form has an equal chance of being selected. There are fancier schemes, but this is the easiest way get an unbiased sample, and for me to explain.

Starting with a generous assumption that half of all signatures are fake, and that this number varies with a standard deviation of 2.6 bad signatures per petition, meaning most petitions have between 5 and 15 bad signatures (also generous). For a sample of size n petitions that gives a total count of x bad signatures the formula for the percentage bad is x/(20 * n) [x divided by (20 times n)] with a standard error of 2.6/sqrt(n)  [that's 2.6 divided by the square-root of n]. Based on the mathematical law of statistical averages from a random sample, we can say that the actual number of bad signatures is close to x/(20 * n), where close means it is within about 2 standard errors in 95% of all such samples performed in this way. For a sample of size n = 1000 petitions, our assumptions and statistics say we should observe 50% plus or minus about 3.2%, or between 46.8% and 53.2% bad signatures - IF the assumptions are correct. We can say there is a confidence level of about 95% (19 times out of 20) that the interval generated this way will capture the true rate of false signatures.


Now the good news: If the actual percentage is more or less than 50%, the standard error should be a little smaller either way, meaning the estimate will be a little more accurate. More good news: this setting is what statisticians call a "finite population sample", which means this estimate will be a little more accurate yet, because the recount is sampling a significant fraction of the total population.


Long story short, if you want to verify recall petitions, take a sample of about 1000 petitions, check them carefully, and calculate the percentage of bad signatures. If that percentage is less than 45% or so, then it is time to stop counting and start campaigning.


** In practice random samples are not always "simple", but this is what statisticians call it.

Sunday, December 18, 2011

Statistics for Badgers

I just discovered BadgerStatsan organization that presents data driven commentary on Wisconsin's economy, education, business climate, and other topics.

http://badgerstat.org/2011/jobs/


Our work is motivated by our belief that:
  • Many Wisconsinites want crediblenonpartisan information about their state, including insights into what’s working and what’s not, and about our state’s challenges and opportunities.
  • Meeting Wisconsin’s challenges will require government that is more efficient andeffective, producing better results for citizens and better value for taxpayers. 
  • Wisconsin government, at the state and local levels, would benefit from a more performance-oriented culture that focuses on results and uses performance data to manage.
  • Every citizen deserves to know how their government is doing in key policy areas. Toward that end, every level of government (and every agency) should provide, online for citizens, a set of clear, timely, and accurate performance measures and goals. 
  • Wisconsin’s future depends on an informed citizenry, since meeting our state’s challenges — and seizing our opportunities — will require people of all political stripes to come together in informed public dialogue to help chart our future.


Speaking as a data-guy, I appreciate and encourage this sort of information oriented reporting. This could become Wisconsin's own version of 538.com.
http://badgerstat.org/

Wednesday, August 24, 2011

Science Blogs Survey

I just participated in the Science Bloggers Census. There was some question in my mind if this blog really qualifies, but Edward was kind enough to link me to an article answering that question of "Just what is a Science Blog?".

Looking back at my recent posts, I haven't been posting very much science related content recently, but I'm also counting some of my mathematical posts at my other blog. That blog gets more traffic, and in a way it's even more scientific than what I write here, but most people don't consider games as science.

The survey results will appear in September at http://labs.fieldofscience.com/ .

Saturday, January 8, 2011

How to tell the difference between a theoretical statistician and an applied statistician

How to tell the difference between a theoretical statistician and an applied statistician?
A theoretical statistician knows all about measure theory but has never seen a measurement whereas the actual use of measure theory by the applied statistician is a set of measure zero.
--- Stephen Senn
Found on Statistical Modeling, Causal Inference, and Social Science
Dread Tomato Addiction blog signature

Tuesday, November 16, 2010

Statistics Fail


image found here. source?
History suggests that if you rescale, shift, and truncate two time series, it's usually quite easy to make them look very similar. This does not mean anything at all, it's simply fishing***. The graphic is suggesting that the US is following Japan into economic deflation, which would be bad. However, this is showing a subset of inflation data (food and energy*) compared on an annual percent change scale* which is shifted by 12.5 years** and truncated before 1989/2001**, and suggesting they are somehow the same. Yeah, right.

* possibly arbitrary.
** almost certainly arbitrary.
*** Statistical jargon alert: "Fishing" is a word for the practice of looking at your data in so many different ways that you eventually find an association simply by chance, and then reporting that association as if it was what you were looking for in the first place. See also: bullshit, cheating, lying, multiple comparisons.

Found on NYT, but the source of the graphic is not clear. I sure as hell wouldn't put my name on it.
[Hat-Tip]
Dread Tomato Addiction blog signature

Wednesday, October 13, 2010

Data Science, Science Data

Image Nature News
The latest Nature News has an article titled Computational Science: ... Error, describing the increasing difficulty scientists face as the computer programming required for research becomes more and more complex. I face a similar prob in my work. Much of the work I do requires a fair amount of database management as a precursor to the analysis. Most of this is basic SAS programming, with some tricky bits here and there, but it is all within reach of my programming skills. Most researchers don't have my programming skills (and with a BS in CoSci, many statisticians don't have my programming skills), which is one of the reasons they might come in for statistical help in the first place. Database management is an important skill for an applied statistician, but it is not my primary skill ("Dammit Jim, I've a statistician, not a bricklayer").

The point of this is not to blow my own horn, but that I have a set of skills for managing databases that is nearly independent of my statistical knowledge. The Nature News article points out the problems with programming skills, but the same problem exist with database skills: Some researchers don't understand the basics of recording data in an organized manner, and disorganized data can lead to as many problems as disorganized programming.

It is not too unusual for researchers to bring me data (typically in a spreadsheet), and sometime I spot specific problems that could be error in how they collected and recorded the data. This is fairly important, because if the data is wrong then my analysis will be too. Sometimes I can fix these errors for them, other times I have to have the researcher fix the problems, because it requires medical knowledge and familiarity with (or access to) the original data source to make the correction. Once these bugs have been ironed out, all it well and I do my statistical thing.

There is another sort of error though, and it is much more subtle. These are the errors in the data that don't really look like errors. When someone brings me their data and there is nothing obviously wrong, I probably don't question it, and proceed with the analysis. There are some common ways this might happen: cut & paste errors, "bad" sorts that scramble the data, inconsistency in data entry, all simple mistakes. Sometimes evidence of these errors shows up during my database management prep work or during the analysis itself. Obviously if a mistake is found, it gets fixed. However, if my experience with finding errors in the late stages of analysis is any indicator, then if seem likely that some of these errors are never found. The "garbage-in, garbage-out" principle applies, and some of the analyses I've produced were likely garbage, because the data was garbage.

The good news is this sort of error is unlikely to contribute much to the larger of body of scientific knowledge. By the nature of statistics (and with an assumption of some randomness) these subtle errors are unlikely to produce significant results, less likely to be in agreement with other published studies, and certainly unlikely to be verified by follow-up studies. The bad news is that some simple, perhaps even careless mistakes can ruin months or even years of research effort, which is a waste of effort.

Finally, this brings me that other set of skills: teaching. Whenever I have the opportunity to work with people who are starting off on new research projects, I try to teach the basic data-skills, the do's and don'ts, to help them get good data and do good research. Not everyone is interested in spreadsheets and databases, but it is not too hard to convince researchers that a little extra effort up front to get good data will pay dividends down the road when it comes to publications. It certainly pays me dividends when it comes to actually doing the statistical analysis - my primary skill - rather than spending hours (or days, or weeks) trying to track down what went wrong with the data, or unknowingly analyzing junk data.
Dread Tomato Addiction blog signature

Friday, August 27, 2010

Biostatistics vs. Lab Research

I had this same conversation just this morning!



Somehow I doubt the full humor of the situation will be apparent to most people, but this conversation occurs at my job on an occasional basis. A sample size of N=3 is a barest minimum for even a t-test, and on that basis alone probably isn't enough, but I'm willing to set that aside for the bigger issue, because it depends on the question.

It is a matter of the question asked. Not all experiments are the same, nor are all samples the same. In my conversation this morning, there was a basic misunderstanding of the sample unit. There was a sample size of N=3 in one group (treatment), and another group of 3 serving as a control. The problem (well one of the problems) was that the controls were not used as an independent group, but rather as a way to normalize each of the first 3. Instead of having two groups of 3 each, that we really had was a single group of 3 pairs of subjects (matched pairs). This lead to a few hours of trying to untangle What had been done versus what needed to be done. Frustrating, but then education is an important part of my job too.
Dread Tomato Addiction blog signature

Tuesday, May 11, 2010

Second Verse: The Census Will Be Wrong

This is a follow-on to my post of three days ago: The Census Will Be Wrong. We Could Fix it. My friend Matt sort of set me off with his comments - I think he likes to do that :-)  - and my "longer response" has turned into a post of it's own, and inspired another yet to come.

Matt writes: I think the author underplays the risk of political/agenda hijinks... The first thing that came to my mind when I read this was that craptastic Lancet paper about civilian deaths in Iraq. Perception is reality, and statistics have gotten a bad rap from a few bad actors. Them's the breaks.

First of all, this is by no means a flame aimed at my friend, and anyone who says otherwise is itching for a fight. This is intended to be my professional opinion with some lightly researched examples to illustrate the problem.
  1. The current census "head count" is known to be flawed (1), and both Democrats and Republicans attempt to take advantage of this (1). 
  2. The mathematics of statistical resampling are apolitical. It's simple a better way to do it, and less subject to error and bias.
  3. Statistical resampling is a simple concept wrapped in a lot of boring math. The basic idea is to go back re-check some of the original counts, and fix them.
  4. The Lancet surveys of war casualties in Iraq is arguably flawed, but a flawed paper in no way invalidates a field of mathematics, or for that matter even the methods of that paper. By way of equally flawed logic, we should abandon all automobiles because Chrysler made the "K" car (That last bit makes more sense than I thought it would.).
  5. The Lancet paper is, if anything, an example that of a study that would benefit from resampling. At a glance - and that's literally all I've given it - the war causalities estimates may suffer more from a lack of precision than a lack of accuracy.

There seems to be a common thread here that many people just don't get what statistics can really tell us. Part of that problem is the growing pains of a fairly recent area of mathematics working it's way into a culture already stressed with information overload. Another part of the problem is that statistics have been poorly taught, frightening students with the math and failing to convey the meaning behind it.

That last sentence expresses one of my original motivation for this blog: to hell with the math, I want people to understand the meaning. Keeping with this theme, my next post will be about the statistical meaning of Accuracy, Precision, and Bias.

Here are some odds and ends I dugs up while researching this post:
-Article about 1999 Supreme Court decision on statistical resampling.
-The 1999 Supreme Court decision on statistical resampling.
-There are some additional comments on the blog of Jordan S. Ellenberg, author of the Washington Post Op-Ed.
-Unrelated Census Hijinx
Dread Tomato Addiction blog signature

Saturday, May 8, 2010

The census will be wrong. We could fix it.

This is sort of a professional pet peeve among statisticians, and the issue comes up with every census;

Jordan Ellenberg writes: Opponents say that statistical adjustment would violate the constitutional requirement of an "actual enumeration" of the population. Justice Antonin Scalia wrote in 1998 that the Constitution's language was "arguably incompatible . . . with gross statistical estimates." The sampling adjustment is indeed an estimate of the population -- but so is the unadjusted number, which estimates that the number of Americans missed is zero! To choose the raw count is to be wrong on purpose in order to avoid being wrong by accident.

Emphasis added. There are demonstrably better statistical methods to perform census estimates, and they should be used.


[Hat Tip 2 Terence Tao]
Dread Tomato Addiction blog signature

Saturday, January 9, 2010

Hula-Hoops, Confidence Intervals, and Invisible Men

In addition to doing statistical analysis, a big part of my job is explaining the meaning of the results. I am privileged to work with a lot of really smart people on all manner of biomedical and translational research. However, "really smart" does not always mean that everyone is highly numerate; some people just are not good at expressing themselves that way. That's OK too, not everyone can be expected to be good at everything, but it does present a special challenge sometimes when I need to explain something statistical to someone who has difficulty understanding. Somewhat surprisingly, the toughest questions are not about the most complex mathematics. Instead, the hardest questions to answer are often on the most basic concepts.

Case in point, a few days ago I was asked to define a confidence interval for a client who needed to be able it explain the concept to yet another colleague who was asking her. Here I have someone not-too-numerate needing to explain it to someone else likely not-too-numerate, and it's important, so I needed to give a clear and simple definition for her to understand and pass on.

A bit of background before I give my definition; this was relating to an observational study on clinical data, and we have a large number of means, standard errors, and confidence intervals to report. All intervals presented at at a 95% confidence level.

A confidence interval (CI) is a range of values that is likely to contain the actual value we want to know. For the purposes of this paper we have made a lot of single value estimates, or point estimates, of the means and slopes in which we are interested. Although this is not a random sample, we hope that this patient sample is an unbiased representation of a larger group of similar patients. If we were to repeat the study with another group, we would hope to get similar results, as opposed to very different findings. Confidence intervals represent a reasonable range of values that might be the true value that represents the entire population, and not just the single sample we happen to have. The confidence level is the probability, here 95%, that the true value we are trying to determine is actually within the interval.

To make sure I got the full meaning across, I added a bit about how to interpret confidence intervals.

Confidence intervals are very useful for interpreting the clinical importance of a finding. If the low and low ends of the interval are both clinically meaningful and not too far apart, then we can be fairly certain of a sound result even though we have some uncertainty due to sampling error. If the CI is “wide” there might be a lot of room for interpretation of what the result really means. The width of a confidence interval is directly related to the statistical significance. There is also a matter of “clinical significance” – not everything that is statistically significant is clinically meaningful, and vice-versa.


I wrote that up and fired it off in an email, but I felt like I still hadn't gotten it simple enough. I needed an easy non-technical example, and a moment of inspiration hit me; the Hula-Hoop as an analogy to a confidence interval, and an invisible man as the population the interval is trying to capture:



Invisible Man Photo

A simpler/sillier definition: You are throwing a hula-hoop at an invisible man. You can be pretty sure that you have actually caught the invisible man inside the hoop (95%), but you can’t be certain. If it is a big hula-hoop, you still don’t know exactly where he is. If the hula-hoop is small enough then you know almost exactly where he is, or close enough that you don’t care.


Now think about throwing “statistical hula-hoops” at the values you want to know about the general population to report in your study. You have captured most of them inside the hoops (95%), but you still don’t know precisely where they are.



Perhaps not the finest hour of statistical science, but if it gets the point across I'll still be happy with it.

[image - 100 hula hoopsDread Tomato Addiction blog signature

Saturday, December 12, 2009

Dipping my toes into the turbid waters of AGW


This is an appeal to my readers and fellow bloggers for some advice. I've already pitched this to two prominent bloggers I occasionally correspond with, but I don't have direct contact with everyone I'd like to poll via email or Facebook, so this is my open call for responses. I would like your opinions, in a Science blogger/Dear Abby sort of way, and anyone else that is likely to read this is welcome to chip in too.


A friend has asked me to participate in a blog/project to conduct an open source attempt to replicate some climate modeling results. This is likely to be an amateur effort at best, but the stated intention is to educate about what really goes into climate modeling. Now I believe my friend to be a reasonable sort of skeptic, but it turns out he has some connections with people like Steve McIntyre and Eric Raymond. This gives me some concern, and I am leery of getting involved in anything that even gives the appearance of supporting the AGW deniers.

Oh yeah, AGW = Anthropogenic Global Warming, if you didn't know already.


I would appreciate your opinions on whether I should become involved, or stay the hell away from it.


Some other information relevant to my participation:

  1.  I have a good mathematics and statistics background, and did some relevant modeling of physical processes (hydrology) in grad school, but have no background in climate science.
  2. I really haven't been following the AGW debate past some casual reading, but not being caught up in the old arguments might be a good thing.
  3. I believe the scientific reports and news that AGW is real. I do have a some of skeptical thoughts about the extent of climate change the models predict, but even a small change is a valid cause for serious concern. I am concerned.
  4. I think this might actually be a valid and useful educational effort, and I'm pretty sure I could make some good contributions.
  5. I need another blog to write for like I need another hole in my head. ;-)

One friend already commented to me ...

"Not to discourage you from having fun, but there are a plethora of people stepping into the debate without sufficient preparation."

Another, who is self-described as very conservative, encourages me to go for it.

[Images Wikipedia, downloaded 12/12/2009] 

Dread Tomato Addiction blog signature

Dread Tomato Addiction blog signature

Wednesday, December 2, 2009

Things that make you go "Hmmm": drug deaths

[From Information is Beautiful]

Visualising the Guardian Datablog:

I’m doing a regular weekly visualisation for the excellent Guardian Datablog, the front-end for an amazing library of statistics and data, lovingly hand-gathered by The Guardian.

My first post is about Deadly Drugs.

[...]
 IiB presents this chart:



Check out the article on The Guardian blog for detail and data. You want both right?

I'll second that.

Sunday, November 29, 2009

193%

From Probably Bad News. Bad math from Fox News is not news either.
Also at FlowingData, with flame war discussion to boot!

source

Breaking News: 193% Is The New 100%:

[post repaired November 2010]

Friday, November 27, 2009

Many Exciting Tables

Found in an online SPSS guide (very near the bottom):

SPSS gives you many exciting tables for repeated measures ANOVA, most of which you can ignore in whole or at least in part. [emphasis added]

Wow! Not only is it an exciting statistical result, but you can probably just ignore it??? WHAT!?!

Likely anyone who isn't a statistician won't see the humor in that -  But to see these described on one hand as "exciting tables" (somehow unlikely), and "can just ignore" on the other hand is unexpected, to the the least. Trust me; the statisticians are ROTFLAO. Dread Tomato Addiction blog signature

Wednesday, November 11, 2009

A Periodic Table of Visualization Methods




Each element represents a different type of data visualization, and hovering the mouse-pointer over any of them will pop-up an example. "Hi" stands for Histogram as demonstrated in the captured pop-up below.



Very nicely done! See for yourself.

This happy accident occurred while I was searching for "element-pun" material in response to comments on a recent post at The Endeavor. Thanks John!
Dread Tomato Addiction blog signature

Friday, October 30, 2009

Diesel Exhaust is a Weighty Matter

Scientific American has this online article -



- the point of which is to point out the enormous amounts of pollutants produced by idling truck engines, and that New York City has a law regarding this that really ought to be enforced. I have little doubt that this really is a significant source of pollutions, but this statement made me raise my eyebrows:

"Idling buses, cars and trucks may not seem like a big deal, but in New York City they spew out as much pollution as nine million diesel trucks driving from the Bronx to Staten Island, according to the Environmental Defense Fund. That’s roughly 130,000 tons of carbon dioxide, 940 tons of nitrogen oxide, 24 tons of soot particles, and 6,400 tons of carbon monoxide each year"

Can that be right? That's a lot of trucks making the trip. I'm a statistician, and I wonder about such things like the accuracy of statistics like this. Watching TV on a Friday night, I started doing so some Googling and back-of-the-envelope calculations during the commercials.

The claim of the article: 130,000 + 940 + 24 + 6,400 = 137364 of pollutants released each year, that's 274,728,000 pounds. The distance between the Bronx and Staten Island is 33.9 miles (so says Google Maps), and in 9 million trips that works out to 305,100,000 miles. 274,728,000 pounds divided by 305,100,000 miles is 0.9 pounds of pollutants per mile.

The Economy of Diesel trucks: The average big diesel truck pulling a load gets about 5.5 miles per gallon of fuel (so says Wiki Answers), or 0.182 gallons of diesel per mile. Diesel weighs about 7.2 pounds per gallon (in cold weather yet, so says faqs.org), so 0.182 gallons/mile times 7.2 lbs/gallon is about 1.31 pounds of diesel per mile driven.

Now 0.9 lbs/mile of pollutants divided by 1.31 ponds per mile of diesel work out to 0.6873, or about 70% of the total mass of the diesel fuel converted to pollution. At this point, I ran into difficulty finding a source for the actual composition of diesel exhaust. 70% might be reasonable; after all, the mass of the fuel has to go somewhere.

But wait, I've made a mistake - most of the 130,000 tons of carbon dioxide is oxygen, without looking up the molecular weights of carbon and oxygen, AT LEAST two-thirds of that mass is coming from the atmosphere and not the diesel fuel. 70% now seems more reasonable.

There is another possible error, which I don't know how to resolve. In my reading I discovered (lost the link, sorry) that idling diesel engines are relatively clean, but produce heavy pollution when under a load (you see this on the road all the time). Therefore the number of idling vehicles must be really huge to make of this difference. This is also quite possible; NYC is a BIG place, but the article does not offer any information about how many vehicles this might be.

I would assume that if NYC has a law that trucks should not idle for more than a minute, there must be some good evidence somewhere to back this up. The original claim seems to come from the Environmental Defense Fund, but I'm too tired for more digging tonight. This has been an interesting bit of fact checking, except that I ran into a wall with limited knowledge of chemistry. Maybe I should ask a Chemist?

[UPDATE, 11/01/2009]
I received this response to my question at About.Chemistry, Sean writes:
the average chemical formula for diesel is C12H23. with that said, the mass of a carbon atom is 12.01 g, and hydrogen is 1.008. so, mathematically, diesel has the molecular weight, on average of 167.304 grams per mole of fuel.
the weight of the oxygen atom is 15.99g (mostly rounded to 16g), so carbon dioxide is 44.01 grams per mole.
in general, this relates to something of the sort:
2 moles of C12H23 +  O2 gas in excess -makes-> 12moles CO2+ 12moles CO+ 23moles H2O. as the formula for the burning of the diesel, if it was a very complete and clean reaction, however, we all know that's never the case. ;.;
as for the 130,000 tons of CO2, that comes to
 117 934 016 200g of CO2

and of that gram mass, 27% is carbon, while 73% is oxygen. that's what...
31842184374 grams of carbon and 86091831826 grams of oxygen.
however, diesel is more of a blend of things and not just the carbon and hydrogen, which pretty much takes all that I have written and makes it almost useless. In my research, though, I've seen more about the fact that sulfur is present in the fuel than the carbon emissions, as that will inhibit the use of catalytic filters to scrub the exhaust clean, as in most vehicles.
also, in response to your lost link, I dug this up
http://busbuilding.com/bus-conversion/diesel-engine-idling-from-an-authority-detroit-diesel/
which says that idling is bad for diesel because it produces more exhaust via incomplete combustion.
anywho, I'm not too advanced in chemistry, so forgive me if I supplied you with random nonsense, I was trying to think of a way to equate the mass of fuel to pollutants produced, but as I can't find an exact formula for what the reactions are this is the best I think I can do. I'm hoping someone else can chime in from here and make more sense of things, and of anything, I wish you luck with your search.

Thanks Sean!

After some more digging I turned up an abstract and a paper on wind tunnel experiments (Part 1, Part 2) that are related, but none seem to be the origin of the original statistics. I might have more luck searching from work tomorrow.
-- OR --
I could ask the author of the SciAm article, Mr Brett Israel. Would that make it too easy?

[Further update 12/09/2009]
I never heard back from the author and I've pretty much given up. Sean's comment about catalytic converters is something else I had not thought of that might affect the results. So many windmills to tilt at, and so little time.

Dread Tomato Addiction blog signature

Sunday, October 4, 2009

Hammy the Hamster Goes Organic

[Reposting with updated links]

Is Organic food really better? Ask any hamster ...



... and be sure to visit The Cooks Den for the out-takes video and comments too. Me though, I want the data!


Here is the response I posted over at The Cooks Den.

First, I applaud this experiment. Flaws and all, this is great fun, and any discussion taking it seriously or impprove suggesting how to improve the experiment serves to illustrate what is involved in the scientific method. Bravo!

Second, if we must take it seriously, then we ought to do it right: I strongly disagree with the statistical comparisons presented so far. A t-test or ANOVA applied here is wrong. This is categorical data and there is no a priori reason to disregard the indifferent results. In a Chi-square test of homogeneity with a null hypothesis of equal proportion of responses (25 each; conventional, organic, indifferent) the exact p-value is 0.1890. Even if we exclude the indifferent responses, the p-value is still 0.0854. Neither of these meet the typical standard for statistical significance (p less than 0.05). This is not conclusive evidence that Hammy’s choices are anything other than random.

Crap. I hate when I don't spot the typo until after I have submitted the post.
[Update: Updated links for The Cooks Den. Originally poster March 2009]

Dread Tomato Addiction blog signature