I figured some people here would be interested to read this interview with Clarissa Smith and Feona Attwood, conductors of a study on pornography use (5,490 respondents, 2/3 male and 1/3 female). They are still in the process of sifting through data, but there are some interesting tentative results contained in the interview and here at the official site (under "results").
I think these researchers need to include a statistician in their research team. Because they are using a non-randomly chosen sample and generalizing from that on the whole population. Which is a big no-no in terms of basic statistics. This kind of generalization doesn't work at all in terms of logic. They have a lot fewer older people responding to their survey than younger people. And they say that this implies younger people watch more porn than older people. But another possibility they can't rule out is that younger people are more willing to talk about sex and porn than older people are. It's willingness of people to talk about their porn, rather their use of porn that they doing statistics on.
^ On the website there is wordage that seems to limit the purview to those that utilize online porn, resolving the problem you raised. Also, you can't get a reputable PhD without taking a lot of statistics; they probably don't need a real stat to confirm their research objectives unless its something quite ambitious or there is something wrong with the data that they discovered while the were collecting it. Failing that, the techniques to bring data to a hypothesis are the first things taught in a doctoral program, and everyone should know them back to front and front to back.
^ Thanks. Also, related to this, Susanna Paasonen (University of Turku) has been working on a really fascinating and exciting project in which they gather data on memories of first exposure/experience with pornography. I had the opportunity to chat with her about it, and the responses sounded so rich and diverse. I am eagerly anticipating the finished project. You can read more about it here: Porn Memories.
^^ & ^^^ I didn't see a hint of statistics in their analysis. In describing their research methodology, they explicitly said they were using "snowball" sampling, which in and of itself precludes doing any sort of statistical analysis beyond simply describing the population that responded to their survey. But let's see if we can reject the null hypothesis that this sample IS representative of the US population. (I suspect we can reject the null hypothesis, because the typical American won't fill out an online poll about porn.) Here's a study on sexual orientation in the US. Because it's hard to know detailed sexual information about the population, I will reframe my hypothesis to say that the two surveyed populations are significantly different. If you look on page six, and assume that half the population is male and half is female, you can infer that about 1.8% of the population self-identifies as "Bi" about 2.2% of men self-identify as "Gay", about 1.1% of women identify as "Lesbian", and about 0.3% self-identifies as "Trans". On page 1 of the porn study, they described their respondents. "Of the 5,490 responses, the great majority named themselves as heterosexual, as these figures show: Heterosexual = 3842 (70.1%); Gay = 186 (3.4%); Lesbian = 56 (1.0%); Bisexual = 905 (6.5%); Queer = 303 (5.5%); and Unsure = 189 (3.4%)" The two studies deal with their "gay" and "lesbian" categories differently. The larger statistical study on "identity" treat "Lesbians" as a fraction of the female population, while the porn study treats them as a fraction of the human population, with similar differences for "Gays". I will deal with this by assuming that half of the human population is female and half is male, even though this is clearly not true, given– among other factors– that 0.3% of the human population is transgender. On its face, the Porn study seems to include an excess of Gays, Lesbians and Bis. Moreover, there are large number of "Queers" and "Unsures." I will count these people as "Heterosexual" because doing so would tend to validate the null hypothesis, and I am trying to invalidate it. So while this is most likely a false assumption, it is the more cautious assumption from a statistical point of view. If I can show that the null hypothesis is false, even with this highly conservative assumption, then that would only tend to strengthen my argument. (The "Trans" group in the statistical study will also be counted as heterosexual for consistency.) If I can not invalidate the null hypothesis making this assumption, then so-be-it; I'm not trying to get tenure, so who cares? So given a population of 5490, we would expect: Hetero: 5210 Gay: 60.4 Lesbian: 30.2 Bi: 98.8 So I fed all of this into Excel: So we have a chi-squared of 7004 and (4-1=) 3 degrees of freedom. I plugged these into a Chi Squared Calculator, and got p< 0.00001. So there's less than a 1 in 100,000 chance that the Porn Survey is representative of the population at large. (Basically the chi-squared calculator broke: A chi-squared of 12 would be highly significant. A chi-squared of 7000 is literally off the charts.) So, I can say with an extremely high level of statistical confidence that the porn study is NOT representative of the general population. Of course, they could have hired a statistician to tell them this, but they weren't doing a statistical study, and they understood that fact.
The way it works is that first you need to choose a defined population to study, such as online porn users, for example. And then you need to choose a random sample from this population, so that it's representative of your defined population. The sample needs to be random in order to be representative. And only when the sample is representative of the population, then you can use it to draw statistical conclusions about your defined population. There is always a need for a random sample, when it's a sample you are studying and not the whole population, regardless of what this population is. And if you read the information from the website describing the above study, then nowhere there you will find that they did something to make sure their sample was random. Which is a mistake in basic statistics and logic. And the thing about trusting researchers, just because they have a PhD. It's a bad idea. Because there is plenty of misconduct going on among research scientists in various fields. You really need to look at their research methods and techniques and not just trust them that they did the right thing. For example, a recent study claims that 1 in 4 cancer research papers contains faked data. Here are some other articles about science fraud: http://arstechnica.com/science/2012/10/research-fraud-exploded-over-the-last-decade/ http://www.wired.com/2015/06/follow-friday-science-fraud-watchdogs/
Do they claim that it is representative though? And even studies that follow the rules of quantitative research are flawed for similar reasons. This reminds me of a book I read called Watching Sex which had a similar data gathering technique but even less broad and more reliant on self-reporting via email (and a few in-person, if I recall correctly). The sample size is relatively small. The author addresses this in the introduction, stating that his findings should not be taken as representative of the population etc. and analyzes how the reports may be affected by various factors. The reports are interesting in their own right. I don't think they should be dismissed simply because they do not reveal some kind of across the board "truth" which would be near impossible to discover anyway. As long as the methodology is thoroughly explained and analyzed in itself, I think the research and discussion is worthy.
No, they didn't claim to be representative. Sorry, I was trying to engage in a bit of satire, and it didn't work. Asking whether a sample is statistically valid when statistical research isn't being done is like asking whether the battle of Gettysburg would have reached a different conclusion if Robert E. Lee had a time-travel machine: It offers no insight into history, and asking whether a qualitative study used a statistically–valid sample offers no insight into the research. The poorly-done satire used statistics to show that it wasn't a statistically valid sample.
Now that's what I call a Gross Display of Power :) Nice work CL. Similarly, I can't believe that chi-squared was not literally the first thing that they did, and why they don't assert that their sample is representative of the population.
Dan, that is pretty much the opposite of the way that it "works". What you've described is the ideal. No one risks their dissertation on a research objective that they don't already know that the data, which they already have or get it from someone who already told them about it, kinda-sorta supports right off the bat. That's what doctoral advisors are for: to steer you away from the sexy objective you're really interested in and towards the boring one you can actually get done on time. I used to think that advisors were just jerks wanting you to fill out their research for them; now I understand that they steer you so that they can help point you at easy data that they already have or are in a better position to get. Your position is pretty cynical. Not necessarily false, but you seem to imply that there's another demographic out there making assertions more deserving of our trust than those that follow the scientific method which of course includes peer review?