Monday, May 24, 2010

Effective Sample Size

It is Monday morning and you are traveling from home to work, a 20 miles journey, when you hit a traffic jam. Things are slow, traffic is doing about 20 mph; you start getting anxious about that meeting they scheduled for 8. After 10 miles of slow traffic the road suddenly clears up and you decide to travel the remaining 10 miles at a speed of 80 mph. You are late to the meeting and tell your boss what happened, she listen quietly and asks, what was your average speed?

Some people, when they hear the word average, automatically think of adding the two numbers, 20+80, and dividing the result by 2. The arithmetic mea of these two speeds is 50 mph. However, miles per hour (mph) is a ratio of two quantities for which the arithmetic average is not appropriate. The first 10 miles, traveling @20 mph, took you 30 minutes, while the last 10 miles, traveling @80 mph, took you 7.5 minutes. Since you traveled 20 miles in 37.5 minutes your average speed is 20*(60/37.5) = 32 mph.

The "average" of these two speeds is given not by the arithmetic mean but by the harmonic mean: the product of the two numbers divided by the arithmetic average of the two. The "average" speed is then [20x80] / [(20+80)/2] = 32. In general, for a set of n positive numbers, the harmonic mean H resembles the arithmetic average but in an inverted sort of way,



The harmonic mean is useful for determining the effective sample size when comparing two populations means. This is because the precision of the average, X̄, is given by the standard deviation, which is inversely related to the sample size. The effective sample size for comparing two means with sample sizes n1 and n2 is given by


Let's say you are conducting a study two compare two products, A and B, using a two-sample t-test, and you take a random sample of 4 product A, and a random sample of 12 product B units (since you want to be "sure" about product B you take "extra" samples). The effective sample size for comparing the average of the 4 product A samples with the average of the 12 product B samples is then [4x12]/[(4+12)/2] = 6. What this means is that your study with 16 samples is equivalent to a study with 6 samples each from populations A and B, for a total of 12 samples. In other words, you are using 4 extra samples.

As you can see, having a balanced number of samples (n1=n2) per population is not just a statistical nicety, but can save you materials, time, and money.


Monday, May 10, 2010

Statistics Is the New Grammar

That is how Clive Thompson ends his article Do You Speak Statistics? in the May issue of Wired magazine. I like this sentence, it makes you think of statistics in a different way. Grammar has to do with language, with syntax, or "the principles or rules of an art, science, or technique" (Merriam-Webster). Statistics being the "New Grammar" implies a language that we may not totally understand, or even be aware of: a new way of looking at, and interpreting the world around us. Through several examples Thompson makes the point that statistics is crucial for public life.

…our inability to grasp statistics — and the mother of it all, probability — makes us believe stupid things. (Clive Thompson)

Back in the late 80s Prof. John A. Paulos wrote a book Innumeracy: Mathematical Illiteracy and Its Consequences, using a term, innumeracy, coined by Prof. Douglas R Hofstadter to denote a "person's inability to make sense of numbers" (Thompson quotes Prof. Allen but it does not mention innumeracy). In my post Lack of Statistical Reasoning I wrote about Prof. Pinker's observation that lack of statistical reasoning is the most important scientific concept that lay people fail to understand. When someone asks me what I do for a living I tell them that I help people "make sense of data", and when I collaborate with, and teach, engineers and scientists I help them realize what may seem obvious:

variation exists in everything we do;
understanding and reducing variation is key for success.

Thompson argues that "thinking statistically is tricky". Perhaps, but Statistical Thinking starts with the realization that, as Six-Sigma practitioners know full well, variation is everywhere.

A lot of controversy has been generated after the hacking of emails of the Climatic Research Unit (CRU) at the University of East Anglia. The university set up a Scientific Appraisal Panel to "assess the integrity of the research published by the Climatic Research Unit in the light of various external assertions". The first conclusion of the report is "no evidence of any deliberate scientific malpractice in any of the work of the Climatic Research Unit". From the statistics point of view is very interesting to read the second conclusion:

We cannot help remarking that it is very surprising that research in an area that depends so heavily on statistical methods has not been carried out in close collaboration with professional statisticians.

The panel's opinion is that the work of the scientists doing climate research "is fundamentally statistical", echoing Thompson's argument.

A few years ago Gregory F. Treverton, Director, RAND Center for Global Risk and Security, wrote an interesting article, Risks and Riddles, where he made a wonderful distinction between puzzles and mysteries: "Puzzles can be solved; they have answers", "A mystery cannot be answered; it can only be framed". But the connection he made between puzzles and mysteries to information is compelling,

Puzzle-solving is frustrated by a lack of information.
By contrast, mysteries often grow out of too much information. (Gregory F. Treverton)

There is so much information these days that another Wired magazine writer, Chris Anderson, calls it the Petabyte Age. A petabyte is a lot of data, a petabyte is a quadrillion bytes (1015), or the equivalent to about 13 years of HD-TV video. Google handles so much information that in 2008 was processing over 20 petabytes of data per day. In this data deluge, how do we know what to keep (signal) and what to throw away (noise)? That is why statistics is the new grammar.




Monday, May 3, 2010

What If Einstein Had JMP?

Last week I had the opportunity to speak at the third Mid-Atlantic JMP Users Group (MAJUG) conference, hosted by the University of Delaware in Newark, Delaware (UDaily). The opening remarks were given by the Dean of the College of Engineering, Michael Chajes, who used the theme of the event, "JMP as Catalyst for Engineering and Science", to emphasize that the government and the general public need to realize that science alone is not enough for the technological advances that are required in the future, we also need engineering. It was a nice lead to my opening slide, the quote by Theodore Von Karman:

Scientists discover the world that exists;
engineers create the world that never was.

I talked about some of the things that made Einstein famous (traveling on a beam of light, turning off gravity, mass being a form of energy), as well as some things that may not be well known. Do you know that he used statistics in his first published paper, Conclusions Drawn from the Phenomena of Capillarity (Annalen der Physik 4 (1901), 513-523)?

Einstein "started from the simple idea of attractive forces amongst the molecules", which allowed him to write the relative potential of two molecules, assuming they are the same, as a sum of all pair interactions P = P - ½ c2 ∑∑φ(r). In this equation "c is a characteristic constant of the molecule, φ(r), however, is a function of the distance of the molecules, which is independent of the nature of the molecule". "In analogy with gravitational forces", he postulated that the constant c is an additive function of the number of atoms in the molecule; i.e., c =∑cα, giving a further simplification for the potential energy per unit volume as P = P - K(∑cα)22. In order to study the validity of this relationship Einstein "took all the data from W. Ostwald's book on general chemistry", and determined the values of c for 17 different carbon (C), hydrogen (H), and oxygen (O) molecules.

It is interesting that not a single graph was provided in the paper. Had he had JMP he could have used my favorite tool, the graphics canvas of the graph builder, to display the relationship between the constant c and the number of atoms in each of the elements in the molecule.


We can clearly see that c increases linearly with the number of atoms in C and H (the last point, corresponding to the molecule with 16 hydrogen atoms, seems to fall off the line), but not so much with the number of oxygen atoms. We can also see that for the most part, the 17 molecules chosen have more hydrogen atoms than carbon atoms, and carbon atoms than oxygen atoms. This "look test", as my friend Jim Ford calls it, is the catalyst that can spark new discoveries, or the generation of new hypotheses about our data.

In order to find the fitting constant in the linear equation c =∑cα, he "used for the calculation of cα for C, H, and O [by] the least squares method". Einstein was clearly familiar with the method and, as Iglewicz (2007) puts it, this was "an early use of statistical arguments in support of a scientific proposition". These days it is very easy for us to fit a linear model using least squares. We just "add" the data to JMP and with a few clicks voilà, we get the "fitting constants", or parameter estimates as we statisticians call them, of our linear equation.


For a given molecule we can, using the parameter estimates from the above table, write the constant c = 48.05×#Carbon Atoms + 3.63×#Hydrogen Atoms + 45.56×#Oxygen Atoms.

For young Einstein it was another story. Calculations had to be done by hand, and round off and arithmetic errors produced some mistakes. Iglewicz notes "simple arithmetic errors and clumsy data recordings, including use of the floor function rather than rounding". He also didn't have available the wealth of regression diagnostics that have been developed to assess the goodness of the least squares fit. JMP provides many of them within the "Save Columns" menu.


One the first plots we should look at is a plot of the studentized residuals vs. predicted values. Points that scatter around the zero line like "white noise" are an indication that the model has done a good job at extracting the signal from the data leaving behind just noise. For Einstein's data the studentized residuals plot clearly shows a no so "white noise" pattern, putting into question the adequacy of the model. Points beyond ±3 in the y axis, are a possible indication of response "outliers". Note that two points are more than 2 units away from 0 center line. One of them corresponds to the high hydrogen point that we saw in the graph builder plot.


In addition to the studentized residuals and predicted values there are two diagnostics, the Hats (leverage) and the Cook's D Influence, that are particularly useful for determining if a particular observation is influential on the fitted model. A high Hat value indicates a point that is distant from the center of the independent variables (X outliers), while a large Cook's D indicates an influential point, in the sense that excluding from the fit may cause changes in the parameter estimates.

A great way to combine the studentized residuals, the Hats, and the Cook's D is by means of a bubble plot. We plot the studentized residuals on the Y-axis, the Hats on the X-axis, and use the Cook's D to define the size of the bubble. We overlay a line at 0 in the Y-axis, to help the eye assess randomness and gauge distances from 0, and a line at 2×3/17 ≃ 0.35 in the X-axis, to assess how large the Hats are. Here 3 is the number of parameters or "fitting constants" in the model and 17 corresponds to the number of observations.


Right away we can see that molecule #1, Limonene, the one with the largest number of hydrogen atoms (16), is a possible outlier in both the Y (studentized residual=-3.02) and X (Hat=0.47), and also and influential observation (Cook's D=1.99). Molecule #16, Valeraldehyde, could also be an influential observation. What to do? One can fit the model without Limonene and Valeraldehyde to study their effect on the parameter estimates, or perhaps consider a different model.

In the final paragraph of his paper Einstein notes that "Finally, it is noteworthy that the constants cα in general increase with increasing atomic weight, but not always proportionally", and goes on to explain some of the questions that need further investigation, pointing out that ‘We also assume that the potential of the molecular forces is the same as if the matter were to be evenly distributed in space. This, however, is an assumption which we can expect to be only approximately true." Murrell and Grobert (2002) indicate that this is a very poor approximation because it "greatly overestimate the attractive forces" between the molecules.

For more on Einstein first published paper you can look at the papers by Murrell and Grobert (2002) and Iglewicz (2007), and Chapter 7 of our book Analyzing and Interpreting Continuos Data Using JMP where we analyze Einstein's data in more detail.





Monday, April 12, 2010

Sample Size Statements

One of our readers, Gary Kelly, who is carefully reading our book Analyzing and Interpreting Continuous Data Using JMP: A Step-by-Step Guide, had, what he called, a curious question:

I am wondering how you typically report out the sample size and power in a statement, given that it is a function of both alpha and beta.

He pointed out that in our book we don't explicitly state it, which is true. He wanted to know

Using your example in chapter 5, starting on page 244, how would you write a statement about the sample size output?

The example in Section 5.3.5, pages 243-248, describes the sample size calculations for comparing two sample means. Figure 5.12, page 247, shows the JMP output from the Sample Size and Power calculator

In my previous post, How Many Do I Need?, I went over the required four "ingredients" for the calculations. We use JMP's default value of Alpha= 0.05 (5%), as the significance level for the test, we state that the noise in the system Std Dev=1 unit, or 1 sigma, that we want to detect a difference of at least 1.5 standard deviations (Difference to detect), and that we want the test to have Power of 0.9 (90%). The calculator indicates that we need a total of 21 samples, or 21/2 = 11 (rounding up) per group, for our study.

So how do we frame our statement about this sample size calculation? We can say that

A total sample size of 22 experimental units, 11 per group, provides a 90% chance of detecting a difference ≥ 1.5 standard deviations between the two populations means with 95% confidence.

Thanks Gary for bringing this up.


Monday, April 5, 2010

How Many Do I Need?

This seems to be one of the most popular questions faced by statisticians, and one that, although may seem simple, always requires additional information. Let's say we are designing a study to investigate if there is a statistically significant difference between the average performance of two populations, like the average mpg of two types of cars, or the average DC resistance of two cable designs. In this two-sample test of significance scenario the sample size calculation depends on four "ingredients":

1. The smallest difference between the two averages that we want to be able to detect
2. The estimate of the standard deviation of the two populations
3. The probability of declaring that there is a difference when there is none
4. The probability of detecting a difference when the difference exists

The third ingredient is known as the significance level of the test, and is the probability of making a Type I error; i.e., declaring that there is a difference between the populations when there is none. The value of the significance level (Alpha) is usually taken as 5%. It was Sir Ronald Fisher, one of the founding fathers of modern statistics, who suggested the value of 0.05 (1 in 20)

as a limit in judging whether a deviation ought to be considered significant or not.

However, I do not believe Fisher intended 5% to become the default value in tests of significance. Notice what he wrote in his 1956 book Statistical Methods and Scientific Inference.

No scientific worker has a fixed level of significance from year to year, and in all circumstances, he rejects hypothesis; he rather gives his mind to each particular case in the light of his evidence of ideas. (3rd edition, Page 41)

The last ingredient reflects the ability of the test to detect a difference when it exists; i.e. its power. We want our studies to have good power, a suggested value is 80%. But be careful, the more power you want the more samples you are going to need. In a future post I will show you how a Sample Size vs. Power plot is a great tool to evaluate how many samples are needed to achieve certain power.

The research hypothesis under consideration should drive the sample size calculations. Let's say that we want to see if there is a difference in the the average DC resistance performance of two cable designs. Given that the standard deviation for these cable designs is about 0.5 Ohm, the question of "how many samples do we need?" now becomes:

how many samples do we need to be able to detect a difference of at least 1 Ohm between the two cable designs, with a significance level of 5% and a power of 80%.

In JMP it is very easy to calculate sample sizes as a function of the four ingredients described above. From the DOE menu select Sample Size and Power, and then the type of significance test to be performed. The figure below shows the Sample Size and Power dialog window for the DC resistance two-sample test of significance. Note that by default JMP populates Alpha, the significance level of the test, with 0.05. The highlighted values are the required inputs and the "Sample Size" the total required sample size for the study.























The results indicate that we need about 11 samples, or 6 per cable design, to be able to detect a difference of at least 1 Ohm between the two cable designs.

Next time you ask yourself, or your local statistician, how many samples are needed, remember that additional information is required, and that the calculations only tell you how many samples you need but not how and where to take the samples, the sampling scheme (more about sampling schemes on a future post).

Monday, February 22, 2010

Old Dog, New Tricks

The statistics profession has gotten some good hype over the past year. In the summer the New York Times published "For Todays Graduate, Just One Word: Statistics". In this article, they discuss "the new breed of statisticians. . ." which ". . . use powerful computers and sophisticated mathematical models to hunt for meaningful patterns and insights in vast troves of data." Some of these statisticians can earn a whopping 6 figure salary in their first year after graduating and they get to analyze data from areas which include ". . . sensor signals, surveillance tapes, social network chatter, public records and more." And, I have to agree with the chief economist at Google, Hal Varian, the job does sound kind of "sexy".

As a 20 year veteran supporting engineers and scientists in the Industrial sector, I feel a little left behind when I read such articles or peruse the job openings section of journals and see the type of statisticians being recruited, month after month, and year after year. It is interesting to see how statistics, and statisticians, have adapted to the changing world around us. During the technology boom in the 1980's, the industrial statistician was king (or queen). I consider myself privileged to have actually worked in the Semiconductor industry during its boom, where the need to look for patterns and signals in vast amounts of data were, and still are, common place. If you were pursuing a statistics degree in the past decade, you would be foolish not to consider the specialization of Biostatistics. With the explosion of direct-to-consumer marketing of drugs, drug companies needed these types of statisticians to design and analyze clinical trials to determine the efficacy and safety of the drugs. As we look to the past 5 years or so, we see a new hybrid of statistician, one that combines statistics, mathematics, and computer science to better deal with digital data in a variety of areas, such as finance, web traffic, and marketing.

As I already mentioned, it is hard not to want to "jump ship" to be part of the latest exciting surge of statistics to come along. That is, until I get a dose of reality which brings me back to center. I guess troves of data also require droves of statisticians and data analysts. If you go on to read the New York Times article mentioned above, you will see that these new super statisticians may work in a group with 250 other data analyst, all hoping for that big break through mathematical/statistical algorithm that will better predict consumer behavior or web traffic patterns.

I should consider myself lucky that I actually get to interact with the engineers and scientist that run the experiments that I have designed and take action on the outcomes of the analysis that I presented. Unfortunately, the days where the industrial statistician reined supreme are long past. But luckily, there is still enough manufacturing in the United States to keep the few of us who remain busy and, even though I'm an old dog, I can still learn some new tricks!


Tuesday, February 2, 2010

What Kind of Trouble Are You In?

Well, I guess that depends on what you have done and more importantly if you have gotten caught! Many of you are probably wondering what this has to do with statistics. In his book, The Six Sigma Practitioner's Guide to Data Analysis (2005), Wheeler aptly describes the nature of "trouble" as it relates to the stability (predictability) and capability (product conformance) of manufacturing processes. Using these two dimensions, a process can be in one of four states:

1. No Trouble: Conforming product (capable) & predictable process (stable)
2. Process Trouble: Conforming product (capable) & unpredictable process (unstable)
3. Product Trouble: Nonconforming product (incapable) & predictable process (stable)
4. Double Trouble: Nonconforming product (incapable) & unpredictable process (unstable).

Hopefully, most of you have experienced a process that is in the 'No Trouble' zone, which is also referred to as the 'ideal state'. The focus of these processes should be on maintaining and sustaining a state of stability and capability. A process which is in 'Process Trouble' is unstable, but producing product which is within specification limits most of the time. That is, measurements may be out of the control limits but within specification limits. Unless this type of process requires heroic feats, by operators and engineers, to make conforming product, its instability most likely goes undetected because we have not yet gotten caught by a yield bust. Moving on, a process that is in 'Product Trouble' can be thought of as a process that is predictably "bad". In other words, because it is stable the process is doing the best that it can in in currents state, but its best performance results in a consistent amount of nonconforming product. While nonconforming product is undesirable, if the losses are consistent, the job of planning and logistics will be much easier. Finally, the 'Double Trouble' process is both unstable and incapable, which can result in a big headache for the business and for those supporting it!

In order to determine the state of your process, you will first need to determine the key output attributes and measurements and then assess their process stability and capability. Recall from my last post, "Is a Control Chart Enough to Evaluate Process Stability", the stability of a process can be determined using a control chart and looking for nonrandom patterns or trends in the data and unusual points that plot outside of the control limits. In addition, the SR ratio can be added to provide a more objective assessment of the process stability, with a stable process producing an SR ratio close to 1 and an unstable process resulting in an SR ratio > 1. The control chart below is the Tensile strength data presented in my post, "Why Are My Control Limits So Narrow?", with the control limits adjusted for the large batch-to-batch variation. There are no points out of control for this process and the SR ratio = 2.272 / 2.2452 = 1.02, indicating a stable process parameter.


How do we assess the capability of a process? This can be done by evaluating the process capability index, Cpk, and determining if it meets our stated goal of "at least 1". For the Tensile data, what if our specification limits for any individual measurement are LSL =45 and USL = 61 and note the target value = 53. From the output below, we see that Cpk = 0.832 with a 95% confidence interval of (0.713, 0.951). Since the upper bound of our confidence interval is less than 1, we have not shown that we can meet our goal of "at least 1". Therefore, by this definition, we would assess this process parameter as incapable.



Based upon results from the process stability assessment (stable) and process capability assessment (incapable) this would put this process parameter in the "Product Trouble" zone. In other words, our process is predictably "bad" and makes out-of-spec product on a regular basis. If we recenter the average Tensile Strength closer to the target value of 53, then we can achieve a Cpk = 1.137, as is shown by Cp in the output above and possible achieve the "ideal state" or "no trouble" zone.

Periodically conducting these assessments to understand the state of your processes is advisable. There is no reason to wait until you are in "double trouble" to pay attention to the health of your processes, because in this context, one of two scenarios has probably occurred. Either your customer was the recipient of bad product and informed you of the problem or, you discovered a rash of bad product through an unexpected yield bust at final inspection. Yes, you got caught! In either event, working through these types of process upsets is draining to the business and potentially dissatisfying to the customer. Remember, ignoring an unstable and incapable process, will eventually catch up with you.