>> found that 62% of respondents reported no or limited statistical knowledge
The other 38% didn't understand the question...(my extrapolation)
I did a couple semesters of statistics at uni. And I can confidently say that the number of people who can answer 3 simple questions on statistics (like say mean versus medians, confidence levels or margins of error) is, well, a rounding error from 0.
Indeed, statistically, no-one has a clue how statistics work.
I did however learn enough to know that statistics can tell you absolutely anything you want them to say. Assuming you don't just make them up, they're trivial to manipulate to generate the headline you want.
When used to evaluate risk, the comprehension goes down further (a fact willfully exploited by any decent marketing.)
One of my favourite jokes: 93 % of statistics are made up on the spot, and 61 % of people believe them.
Bonus points for changing the figures every time you retell the joke.
I’d recommend reading “How to make the world count” instead of a work of a morally bankrupt guy who used the very same tricks he criticized to discredit the cancer studies of tobacco usage.
Yes, but less pithy. Placing the "average person" next to the "average number of legs" makes it clear that the idiomatic and intuitive sense of "average" that most people have is wrong. The more precise "typical person" doesn't make that same point as well, though an intelligent reader would certainly be expected to be able to read the same from it.
You _have_ to change the numbers each time, because you're constantly measuring and these things have basic variation. It would be unrealistic if the fake statistics were the same every time!
> the quote popularized by Mark Twain "Lies, damned lies, and statistics"
I’m increasingly convinced this is a thought-terminating cliche. Understanding the difference between a median, mean and mode is fundamentally empowering. Enough people, however, will stop themselves from trying to understand that by quoting such a joke.
Its more a foreshadowing of Simpson's paradox and a cluster of similarly counterintuitive effects. Statistical evidence that is technically of high quality can cause decision makers to believe the exact opposite of what is true.
> Statistical evidence that is technically of high quality can cause decision makers to believe the exact opposite of what is true
High-quality statistics allow themselves to be inspected. When a statistic is being passed around unadorned, particularly with a call to action, that’s no different from an uncited fact being accepted at face value. (Most of the complexity of statistics is in sampling. Most of the errors in statistical judgment tend to come down to misunderstanding how a sample’s statistics translate to the population you actually care about.)
> mean versus medians, confidence levels or margins of error
It's basic when you are attending an undergraduate course but most people can understand mean (as a dictionary might generically define it) and have a general feeling for margin of error (again, not in the mathematical way.)
Statistics has been the most difficult course in my CS course. For some reason when I start counting events to get a probability I find several perfectly plausible ways to count them, get five different probabilities and none of them is the correct answer.
And about being "trivial to manipulate [numbers] to generate the headline you want" a politician once told me that you can show the same numbers in any way you want, as in to demonstrate a thesis or its opposite.
With large sets or sequences of numbers approaching almost infinity, counting as a method of determining how large things are ... doesn't work well as you found out. To determine if ginormous set A is larger than ginormous set B, you have to do things like matching-pairing. For every element in set A, pair it with an element of set B; if you exhaust all of the elements in B but still have elements in A... then set A is larger than set B.
To formulate the probability of B happening, for example, you have to make an educated guess. Create a functional equation that map A (all possibilities) to B (desired possibility). Then you make a ratio-fraction with the equation in the numerator and set A in the denominator. Using Algebra, try to eliminate references to A in both the numerator and denominator to give you an educated guess of what the probability is.
It seems like you are mostly talking about cardinality, or proofs of problems pertaining to them, and the person you are responding to is saying they struggle with combinatorics
In college, studying mathematics, probability was the course in did best in. The funny part is that I have a form of ilnumeria. Effectively I can’t count. As a result I have no desire to try to actually count things, and the easiest way for me to cope is by understanding the theory behind counting.
I wish people had an intuitive understanding of probabilities. It seems the average person can only think in terms of "basically never happens", "fifty fifty" and "sure thing".
They do have an intuitive understanding! You just outlined it, and it works well enough for most things. Statistics is generally very useful, but for individuals it's not really important or useful to understand, because the data is rarely ever clear or trustworthy enough to use for making decisions about specific situations.
Probabilities might be more intuitive, statistics are intuitive only for the most straight forward topics. That's why it's so easy to present entirely correct but very misleading statistics.
And how people don’t get how probabilities work when run over repeated instances. Like it is only 0,1% risk that X happens everytime Y is done which sounds low but then it turns out Y is done Z number of times per day…
The word “most” is a mess. Strictly speaking the person with “the most” has the plurality, not even half. “Most people” can mean more than half, but is often used to mean or imply anything from a supermajority to “almost everyone”.
Maybe people would have a better understanding of percentages if 20-sided dice were used more. "This has a 95% chance to succeed" -> roll a d20, anything but 1 is a success
Meanwhile 6-sided dice don't map to nice numbers, and having two dice makes the maths even harder to intuit
But (assuming two dice rolled at the same time) 2/36 is something like "roll a five and a six" or "roll either a pair of sixes or a pair of fives". Neither of which is quite as intuitive or commonly used in game design as "using one d20, roll a natural 20"
And yet they struggle with the abstract concept. They'll wonder why the candidate with a projected 20% chance of winning won. They won't know what to do with the information of a 70% chance of rain tomorrow. "Either it will rain or it won't, which is it!" Perhaps it's the frequentist approach that comes more natural than the Bayesian approach, who knows.
One of the first managers I had who studied management theory commented once that if estimates were accurate, we’d get them done earlier than expected as often as we got them done later than expected, which is not true and thus we are doing something wrong.
A probability that isn't computed (or reasonably computable) is just a personal feeling. Your 40% or 70% aren't probabilities if you didn't compute the probability of at least the key events involved to arrive at the number. People say a number as their confidence to meet the deadline on a scale from 0-100.
I'm not saying they're useless, they're an educated guess, a valuable means of communicating an expert opinion. It can be replaced with "I have high confidence of success".
The assessment is subjective, includes personal beliefs, assumptions, etc. not just measurable objective points. The measurement uncertainty at every step and the confidence interval are disappointing despite people using percentages, suggesting confidence to the percent. A scale from 1 to 5 is most of the times just as accurate.
2 people looking at the same situation could have very different degrees of belief, which is why 2 engineers could give very different estimates. But if I say I believe the likelihood of something is 73.8% I might be more credible to the listener. Looks like I calculated, not like the guy who said "3 out of 5 we get it done".
I'm not very familiar with how they work. Don't prediction markets give the individual gambler a boolean choice and this results in a picture of the split over the population? There's no granular option for the individual to bet 40% vs. 60% for something happening, the vote with 100% confidence that Spain wins the World Cup, they don't give 85% chance of winning.
The way to express a 85% confidence is to size your investment accordingly. For example, if the market is at 75%, then you buy yes-shares until the price is at 85%. If the market is at 90%, then buy no-shares instead until the price is down to 85%.
Statistics might be the most eye opening course I took in university. Like, I don't think I understood the concept of distribution before it and regarded an average of measurements as one measurement. I had a very strange concept of averages.
I started thinking about this percentage on decision makers around the world and got a chuckle.
I've been trying to talk about the median vs. mean with a bunch of politicians on themes around the zillionaires, wealth or consumption and for most parts they're clueless. Or the ones with degrees still go with the normal (mean that skews the normal) as they're afraid.
I just happened to get a copy of naked statistics yesterday, as I feel the human traits of poor comprehension of probabilities can be enhanced to at least some extent. And my degree from the social side didn't include statistics.
If there are better entry-level books on the matter I'm happy to take some recommendations.
Related, this result on statistics comprehension is another stepping stone on my path away from believing in technocracy - a journey I’ve been on for the past 10 years and have felt more passionately about over the past 2 years.
I was recently convinced by a rather compelling argument that technocracy (rule by the technically capable) often falls to a local minima where the people in power find the first scrap of evidence that supports their preconceived agenda. Wildly racist? Natural selection “proves” that all of your decisions are “technically” justified.
But more directly to your point about being effective communicators of statistics, I’ve found that it’s difficult to help someone form their own conclusion from statistical results. Most people want the conclusion fed to them, with the statistics supplied as evidence. It’s up to us (proper statisticians) to act ethically: provide analyses in good faith and challenge the poorly conceived analyses of amateurs and statisticians acting in bad faith.
> And my degree from the social side didn't include statistics.
In the university I went to, Department of Statistics used to be part of the Faculty of Social Sciences. It was impossible to graduate from social sciences without taking a few classes of statistics. Then in some reorganization, Department of Statistics merged with Department of Mathematics. Except that old-school statisticians didn't want to move to the Faculty of Science, and they somehow managed to keep their offices in the Faculty of Social Sciences.
And I took some math in university before dropping out of CSC decades ago, so it isn't too hard, but my hunch is that with the BSs students the statistics in decent level would've been too much anyway.
Rhetoric is persuasion under the presupposition that what you are trying to persuade someone of is true. Sophistry is indifferent to the truth and is merely concerned with results. Rhetoric respects the humanity of the interlocutor. Sophistry seeks to exploit him.
Exactly. 38% is way too high. Are we sure about that? Statistics is not some required learning in school. Even those who learned (me for example) cannot confidently claim what it actually means. I would say most of the people don't have a clue what statistics mean
Here in Norway, statistics has been introduced in 9th and 10th grade (15 and 16 year olds). Here[1] is what a 9th grader is supposed to know at the end of the school year. Some key points translated to English:
- Interpret and critically evaluate statistical representations from the media and the local community.
- Calculate measures of central tendency and measures of dispersion in custom and real datasets, and use the results to describe the data.
- Calculate and evaluate probability in statistics and games.
This came after my time, and our PIRSA score aren't the best so perhaps practice is lacking.
We are just as bad or worse with charts. I don’t know whether finding Tufte when I was still relatively young saved or harmed my sanity. Maybe both. Everyone’s charts are awful, including at least half of mine, and I was trying to be objective. Many people are just trying to prove a theory they had before the chart was made.
When I have seen instances of this, it's usually because there is another variable. Example from wikipedia:
>A common example of Simpson's paradox involves the batting averages of players in professional baseball. It is possible for one player to have a higher batting average than another player each year for a number of years, but to have a lower batting average across all of those years. This phenomenon can occur when there are large differences in the number of at bats between the years.
The per-year values aren't weighted in the combined total average.
Congrats for having a shot, sadly that doesn't answer the questions posed in the slightest.
Again, why does the GP assume that a full half are distinctly lower than "average understanding" ? Why the assumption that multiple numbers cannot share an average understanding?
I’m interested, in what cases would it not be true if we assume “average” here to be “median”? I can only think of one: if we consider “understanding” to be discrete rather than continuous.
the title, in respect to the content, is super lame "more than half"... dude what is "more" and how much more, and are we talking the upper bound or lower bound of the "more". perhaps we can conclude, the author demonstrates perfectly the principle that he tries to convene... few people are ready to work with statistics.
A side note - I blame that on our education. In my anecdotal n=1 case, the statistics were never taught in school (and I was in math class all the way and the school was higher tier locally, a "lyceum"). And in the local Polytechnic the only course which taught the subject was a Probability Theory and Math Statistics. It was one of those "intimidating" courses on my faculty, the ones which older students scare freshmen with. And it was indeed as crazy as they said, I remember nothing from it outside of the sheer horror of rote remembering hundreds of pages of, well, something. I scrapped by with a lowest passing mark and promptly cleared my brain cache of that.
Nowadays I often stumble upon this or that statistical topic, watch an educational video, read a wiki explanation and it makes sense. Plus I've picked up unstructured and chaotically a lot of terms just by reading IT articles and forums.
tl;dr - my point is, our education is severely lacking a simple, short and concise statistics course on ELI5 level, for middle schoolers. And it is a huge gap in skills people actually do need in common life, outside of STEM. And the skills I'm talking about aren't even hard, a core set of basic concepts, without any math, can likely fit in a tiny brochure written in simple literary English with a few illustrations. Or a set of YT videos along the same lines.
It would be interesting to probe the political orientation of the people who understand and those who don’t and see if there is some correlation there.
My thesis is that no-one understands. So there's no political orientation.
Your suggestion suggests that a little bit of technical education would change political allegiance. Alas that is incorrect. Politics is about worldview (me and mine versus you and yours) and a world view cannot be changed by something as mundane as education.
Each person has both a self-centered side, and a community-centered side. For some it's mostly self centered. For others mostly community centered. Politics is about finding out which side the population has swung to.
A bit of a tangent, but I don't think it's really fair to say it's about self vs community. I think the main reason that people start to lean in a different direction as they age is because you start to see cracks in the ideals you held when younger. The internet is a good example. Many of us lived through the birth of the mainstreaming of the internet. And our perspectives of what the future held at that moment was generally just exceptionally naive. For those who are a bit younger, a similar phenomena played out with the rise of social media as well.
And as you see how reality and society tend to 'really' work, wisdom in other words, it tends to result in a shifting perspective on the ideal way forward for society. I'm certainly much more socially minded than when I was younger now. And I also think that's fairly typical - like most young people I was completely self absorbed which makes 'real' social mindedness quite uncommon. Yet my political leaning is gradually shifting ever more towards the 'self centered' orientation, in your terminology.
The problem is that convincing yourself that you are smarter than your younger self with opposite views, means that you are right about the politics now, is just a fallacy.
The change doesn't come from thinking one is smarter or dumber, but you simply having much more information to work with. We form extensive opinions and views about the world relatively young, but it's largely based on things we have 0 personal knowledge or experience of.
Again I think the internet/social media are perfect examples. Predicting what would happen isn't about being smart or dumb. It's a matter of just not having enough real world experience to even begin to formulate meaningful views. And so the views and expectations we did form were quite out of touch with what really happened and indeed what we probably 'should' have expected to happen all along.
I'd separate worldview from ideology. I think worldviews, for people who have all seen a comparable share of the world over a comparable time era, are going to fall quite close. But ideologies are something distinct from that. Worldviews should influence ideology, but in practice it's the other way around for many people - even as they age. In any case, without that world experience - the values we formulate when younger are going to be based on assumptions that will ultimately prove to be false.
But definitions are very much important when discussing statistics. Being able to communicate using correct terminology is important. Part of knowledge is being able to ingest statistics produce by someone else.
i mean, the headline statistic can be misleading, but you gotta dig in and wrestle with the details. just like anything, we cannot boil down complex things to single numbers and expect any sort of meaningful signal. we gotta roll up our sleeves, look at definitions, think about what our actual questions are, how we might answer those questions through measurements and observations, and what the confounders are. i think a common issue folks have with stats is that they expect a tidy answer, and it just doesn't do that: it's more of a way to prove the world...the results still need some interpretation.
I was having a conversation with my dad about how if a virus had a 10% death rate it would be terrifying.
His reply was that 10% was nothing, that means you've got 90% chance to survive. Then said you've probably got a 10% chance of being hit by a bus every time you cross the road. Just absolutely no idea what probabilities are at all or ability to extrapolate probabilities from real world lived experience.
I'd say in reality it's way more than that. The first statistics course in uni was a very humbling experience. I realised that while I thought I understood a lot (and I was coming from a CS heavy background, olympiads and such) real statistics is way harder and a lot more counterintuitive than I thought. Granted, this talks about "basic" statistical understanding, but even that is way more complicated than most people assume.
Yeah. Like most people, my only experience of statistics is studying as part of maths in secondary school. I was really strong in maths generally and most of it came relatively easily to me, but I struggled with statistics. Examples like the Monty Hall problem show just how unintuitive the basic concepts are.
Given how effective stories and anecdotes are in convincing people, it really seems like the human brain is not wired to grasp these concepts easily.
yeah, i've studied a lot of math and a lot of cs, and stats is tough. part of the problem is the terminology, and just giving a ton of complex machinery without telling you what it's actually doing. i've never learned that way, and it is very easy to feel like you're doing some sort of dark magic.
also, probability theory vs statistics is an important distinction: prob theory is a nice clean mathematical subject, while statistics is almost the philosophy of applying probability theory to the world.
Statistical reasoning isn't really motivated when it's taught, at least in the US. My schooling (I did jump around a bunch) assumed the student to have picked it up through vague balls-and-bins style problems taught in various units in various grade levels.
By the time I took my statistics class in undergrad math, coming from a similar background to you, they just sort of assumed you had a head for combinatorics and used that to develop everything else. I was a really good student in undergrad and statistics was my hardest class, I spent like 2x time on that class than any other class.
In grad school I took a class on complex system failure analysis and was quite apprehensive. My hope was that I could team up with a classmate to help with the math while I could work on the systems levels analysis. Turns out that because I understood systems really well, system failure offered me the intuition I needed to really understand statistics. I aced the class, published a moderately popular paper in distributed systems using what I learned, then went and took our graduate level statistics class widely known to be very difficult and aced it.
I think tacking statistical thinking on as an afterthought in curriculum is a huge mistake in the school system, especially so in the age of machine learning. I think for the average student statistical thinking is even more important than a lot of trigonometry.
Ugh no. You and the sibling commenter pushed me to do some searching but it looks like that class isn't offered anymore. It was a special class anyway (which is kinda common in grad schools.) I'll see if I can dig up some lecture notes from a long time ago.
This is also a political problem. Most of the debates about racism for instance are between people who are illiterate in term of basic statistics, on both sides of the debate. They take extreme tails of distributions as benchmarks, think in term of stereotypes rather than distributions, infer rules from a handful of samples, etc. It makes the whole exercise sterile.
It is not just journalists and politicians that are guilty of that. I see that from many academics who should know better, particularly in humanities departments.
Are you talking about actual academic work or the pop- version of it? Because most of the damage I see is in pop-psychology, pop-sociology. In my limited experience looking into the topic humanities researchers make specific balanced claims backed with some numbers and reasoning, and all the nuance gets lost when translated to news articles and pop-sci books or videos
I don't think it's a question of design. Some ideas in stats are just fundamentally unintuitive. Pedagogy can help mitigate this, but it would take a revolution to fit it into a neat and intuitive framework, like calculus.
What does it really mean to properly understand p-value? What I remember is that if the p-value is less than 0.05, the research result is considered statistically significant. That's about as far as my memory goes. I know that's actually a misunderstanding, but that's how most people understand it. I'm not sure how much I need to know to say I truly understand it.
P-values are almost always taught poorly, but it's not actually that difficult of a concept. I took statistics in high school, again in undergrad, and it wasn't until the third time in grad school that it actually made intuitive sense (thank you Julia Yang!). When you're testing a hypothesis in statistics, it's easier to formulate a "null hypothesis" which is the opposite of what you're testing, and then try to disprove that null hypothesis.
A p-value is the probability, assuming the null hypothesis is true, of obtaining a result at least as extreme as the one actually observed.
Put differently: if the null hypothesis were true, then for p=0.05 you'd see <things at least as far from the test statistic as what you just observed> at most 5% of the time.
Put differently again: If the null hypothesis you are testing is true, then for p=0.05 random sampling would not return an observation as far from the test statistic as you just observed, 95% of the time.
Agree. I think it's more important to understand statistical fallacies (selection bias, regression to the mean, survivorship bias, etc). Those are extremely common trip wires but to recognize them you don't need to memorize formal definitions.
It is mumbo jumbo (also people hearing “significant” treat this as “effect/change is large and important” which is in no relation to the actual amount of change).
Low p-value basically means how surprising your data would be if there were actually no effect (ie less than 5% of the time you’ll get this due to randomness if there is no change — which is rather impossible)
Sample size matters heavily. With more observations, estimates become more precise, so increasingly small differences can become statistically significant. With a large sample, you can therefore get a tiny, practically meaningless effect with a very small p-value.
Eg effect of $1 can be statistically significant (not random) which does not matter in practical terms if average is like $10000.
So the key point here is not only to look at the p-value but also at an actual change. If a drug gives you only 0.01% more hair, it doesn’t matter to you that it is guaranteed.
One of the really awkward points of stats is that not only does sample size matter, but also model specification. Very small misspecifications can easily lead to infinitesimal P values over large sample sizes.
Similar things are true of Bayesian stats, leading to things like predictively oriented posteriors being studied nowadays.
Say you go out and you see a sports bar and the FIFA world cup is on. France vs Sweden is playing on the TVs in the bar and the place is packed with fans of both countries. You look at a few people (wearing the national football colours) and you think to yourself "the French fans seem shorter than the Swedish fans". You decide to run an experiment to test this hypothesis. Lucky for you at half-time exactly n fans of each side agree to let you measure their height. You do this and of your sample of n fans, the French are 2cms shorter on average than the Swedish fans.
Now: the p-value is the probability of you getting a result at least this extreme assuming the null hypothesis (that there is no difference in average height between the general population of French and Swedish people).
Fair play to you for answering, but that’s not what a P value is! A p value is the chance of getting a result greater than the result you actually got, _assuming you are testing a hypothesis_.
I’m also pretty sure I would fall in the camp of saying “nope don’t understand P values” as I can’t remember anything else about them.
Same. I think it involves doing some other stuff very much properly... Like formulating whole thing... And even then you might luck out if you try enough of things...
So I admit I only know that P values are somewhat useful some of the time.
There is the old joke that 87 percent of all statistics is made up on the spot. I’ve told it many times but a fair amount of people seemed to believe it hook, line, and sinker.
Even leaving aside the quantitative stuff (p-values, medians, whatever) and can't crunch the math, IMHO at least you should have seen how statistics can lead you to the exact opposite conclusion from reality, so that you at least know whether to think twice about a conclusion drawn from statistics thrown at you. Yet I was recently quite surprised to learn that even many folks in tech had never encountered Simpson's paradox before. All it takes to start explaining that is a scatterplot and a few lines, and yet it doesn't seem to be taught widely. It's rather terrifying, given that most people (myself often included) will be happy to believe "obvious" conclusions drawn from seeing one percentage greatly exceed another.
I would say the overwhelming majority of even university faculty members lack a basic understanding of statistics.
In their heads, they’re doing something closer to informal Bayesian reasoning: they have prior beliefs based past observation, previous research, experience, etc., wich is then all baked into every choice and decision
But when it comes time to write the paper, they translate that into frequentist statistics: p-values. As if all should be compared to random chance: significance thresholds, confidence intervals.
They do this to give their conclusions formal credibility: oh look how compelling my result is, I found all this new surprising stuff.
Close to 100% of life sciences research works that way.
As a grad student at a public university in the USA, I worked as a teaching assistant for the sociology class "Introduction to Statistics." The course used SPSS software to enable the students to perform statistical analysis on data sets. Although we offered an excellent textbook (listed below), most of the undergraduate students struggled with the material, so my colleagues and I added a supplemental reading list to the syllabus for those students interested in studying the main concepts. The following books are the ones we recommended to the students. I have read all these texts, and they are pretty good at explaining statistics to anyone interested in learning the material at a conceptual level; anyone can read and understand these books, regardless of their educational background.
Ian Ayres (2007) Super Crunchers: Why Thinking by Numbers is the New Way to Be Smart
Joseph Healey (2020) Statistics: A Tool for Social Research
[This is the textbook used in the class, which I highly recommend]
Catherine Marsh and Jane Elliott (2008) Exploring Data: An Introduction to Data Analysis for Social Scientists, second edition
John Allen Paulos (1988) Innumeracy: Mathematical Illiteracy and Its Consequences
Nate Silver (2012) The Signal and the Noise: Why So Many Predictions Fail, but Some Don't
Nassim Nicholas Taleb (2005) Fooled by Randomness, second edition;
_______ (2010) The Black Swan: The Impact of the Highly Improbable, second edition
One of the simplest ways to measure risk in stocks, bonds, crypto, mfs etc for ordinary people is Maximum Drawdown. Once you understand how it works, it takes the stress out of making and managing your own portfolio P — espexially your non retirement account. The question you have to ask your self is: "Given some portfolio with returns of X%, am I ok with this asset being down by Z% over T years — i.e the _computed/inferred_ drawdown?" If the answer on one end is no I cannot afford P to have any drawdown at all, then just put your money in a cash/MMF and call it a day. Generally people are ok with some risk on some percent of P and stash the rest in cash, and you _risk adjust_ for Z and T. The portfolio choices are surprisingly simple.
No other question matters. There are some risks but the key element is the understanding of the statistical variance because that allows me to say that "ok 50% of P can afford to be down for 3 years and I won't be homeless".
Most people don't know this but most money managers(managing money for ordinary Americans) that you hire compute this number _once_ and make tiny adjustments to your portfolio(I am talking once a year maybe) and take 1-2%.
I follow this guy called Dave Stein who has a B2C product called money for the rest of us that taught me this. I am not affiliated with them in any way.
I think the posted article is also the description of America's labor market. Many people don't use it on a day today basis and most people cannot engage their system 2 to assess risks over a few weeks out, let alone 5-10 years.
But I found This is a real practical everyday example of why statistical uncertainty is important to know about. It teaches you how to compare — risk adjust — two vastly different investments e.g. BTC vs S&P or indexing vs value strategies. Then you sleep at night.
It is also not intuitive and many people are very anxious and FOMO driven when it comes to money. So you need to internalize the idea for it to sink.
The drawdowns are a statistical measure and you could be unlucky to catch a big depression style thing once in 30 years. And T is typically longer than 5 years, typically 10 or more.
This was abundantly clear when people, even on HN, were upset about the bureau of labor statistics revising their numbers tendentially downwards, probably confusing the notion of statistical bias for that of political bias.
If the figures are biased, just estimate the bias and correct for that, what is the big deal they said, as if the bias variance tradeoff was not a thing.
I would be more likely to believe the results if they tested these adults and not asked them.
There's a big ego hit in admitting you don't know something. And many people are brought up thinking that it's a shame not to know something and that someone is better for knowing something. Like, a better person, not just better in some field.
And then the second question was:
“How often would you base decisions on reported statistics if you understood it better?”
Talk about a leading question. I get that they were tacking on their “study” to a larger question pool, but they could have put more effort in to the questions or done some actual testing.
I recommend every youth take a week to a month to try to read and internalize the book "Once upon a number." It's full of those odd statistical paradoxes that seem impossible but once you really wrap your head around them you realize how oversimplified your prior understanding of probability was.
And even the headline gets it wrong. They asked 1,000 people, over half of whom say they lacked basic statistical understanding.
The first question was:
“How much do you understand about statistics and p-values?”
not probability, p-values.
I had to look up p-values to check it meant what I thought it did to avoid making a fool of myself in this comment. I knew I was probably right, but possibly wrong.
The same goes for my assumption that this survey was engineered to deliver a bad result.
Why? because it makes it easier to get funding for a the author's solution to the problem. Again, probably right, possibly wrong.
Statistics, like math, is a language, first learn the vocabulary, then you can judge the meaning
Everyone is hating on statistics, yet HN posts studies all the time.
Yes, stats can be used deceptively or incompetently used, but they are generally useful. You just have to look at the full study to see what the starting data was and all the actions taken on it. Then you have to make a judgements call on if/how you want to apply that to your own choices, political policies, etc.
A great example of a stats being used by two sides of an argument and political positions would be gun control. See Everytown Research and Crime Prevention Reseach Center. You can see what gets measured, included, excluded, etc by each side and how those results get fitted for support of political policy.
A coarse understanding like "the smaller p-value is the more likely a headline is true" is worse than no understanding at all. I bet most people who believed they understood what p-value is are like that though.
Makes me wonder what the numbers are like in other countries. There's a lot of talk here about issues with the American education system and what not, but I suspect the numbers would be about the same in many other places too.
Like, if I asked the average person here in the UK what a p-value was, I suspect the majority either wouldn't know or would have barely heard of the idea. I suspect those numbers wouldn't change much in other countries, whether in Europe, Asia, Africa or South America.
> How much do you understand about statistics and p-values?
I suspect that it’s far far less than 38% of people who actually understand statistics to this level. I suspect if someone on HN went around and asked their co workers to explain what a P value is in 2 sentences, it would be less than 10% of a (presumably) highly educated workforce.
I suspect about 40% of adults are unable to tell the difference between mean/median/ mode, or could answer the Monty hall problem, or even “if I flip a coin 3 times are the chances I get heads 3 times in a row”
10% may be a magnitude too high at least if we expect the explanation to be correct. For example in a sample of statisticians/epidemiologists only 12.5% got two basic features of p-values correct [1]. And there are a lot of studies showing similar levels of misunderstanding among people using p-values professionally.
I use p-values daily. I can and do compute them using various methods, including by hand. But I'd still not be confident in my two sentence explanation. P-values are very unintuitive and very easy to get subtly wrong.
In my experience, a not insignificant chunk of students actively in a stats class think that a p value is the probability of the null hypothesis. My money is on more like 3% of educated professionals.
If you asked me for the definition of p value I could give it to you but if you asked me what it meant before this thread I would have said the same thing as you’ve said here.
I mean, the Monty Hall problem is a literal gotcha that trips up literal professors, that's a horrendous example to use as the baseline for "basic understanding of statistics".
Other than that choice of example, I do agree in that I doubt anywhere near 40% of adults have basic statistical literacy. I've played in card game tournaments semi-professionally and just gambler's fallacy + results-oriented thinking alone make it so easy to take other people's money, and if you can't figure out such basic concepts as "getting tails once doesn't mean I'm due for a heads next flip" even when you're literally losing money, what are the chances of anyone else caring about understanding it when they're not even being given the hands-on reward-based reinforcement learning opportunity?
Yeah I maybe should have ignored the Monty hall problem - although I’d guess if you’ve studied enough stats to know what a P value is, and how to measure it, you’ve come up against the Monty hall problem!
Most HNers don't even understand what percentages are and throw out dumb statements like "mega corp can treat 1% of users like crap, it's a small number!"
I can beat a gorilla in a fist fight because it's gonna rip my body in half, but lose by disqualification due to rules violations. A posthumous win is still a win.
The fact that 62% of Americans have little or no understanding of statistics may be related to those 40% of Americans that believe in Creationism, i.e. human (and fossils) were created by God a few thousands of years ago, and not of randomness and evolution over millions of years. I think no other industrial country has that disbelief in science.
https://news.gallup.com/poll/261680/americans-believe-creati...
This problem with science is apparent in another survey: in 2009, a Pew Center publication showed that 33% of scientists in the USA believed in God (and 18% in a transient power), which is much lower than the 80% belief of the general American population at the time. Of course, this is not a proof of causality in either direction, but scientific knowledge is seemingly inversely correlated to religiosity. And the USA are still more religious than any other industrial more-or-less-democratic country.
Your link goes beyond your 40% figure as well if you add in the 33% who don't believe in strict Creationism but believe that Christian God guided the process of evolution.
Basic statistics and basic economics/finance are two subjects that are completely underdeveloped within the general population. Coincidentally those two effectively rule our entire lives. As an aside, STAT 200 looks like a solid stats course.
What's going on in Pennsylvania? Is it so highly educated a place that they don't understand that 50% of U.S. adults can't read at a 6th grade level, let alone understand "confidence intervals"?
I would estimate that 90% of software developers I have spoken to lack basic statistical understanding. Basic concepts like kurtosis, skewness, moments, the core Greeks all seem to be alien concepts to most.
You can probably lead a decent life without understanding kurtosis and skewness. You probably can’t if you don’t understand when a normal distribution applies.
> The survey showed that 62% of U.S. adults self-report having little to no idea what statistics or statistical concepts like p-values are, but 90% of them would base decisions on reported statistics at least sometimes if they understood them better.
You're kidding, right? 38% self-report more than that? If their self-report were accurate it would imply an education system that has truly excelled.
Yeah I figure I'm "well educated" and I never took a real statistics class and so never learned the definition of a p-value. I also don't know about the Wars of the Roses or Millard Fillmore. I did learn about Hildegard von Bingen and Laguerre polynomials.
I actually think this is very good. People thinking they understand statistics is far more dangerous and causes far more issues that people honestly saying they don’t…
In common English, I would say that "average" as used in the quote is understood to be something closer to mode, hence why the joke works, because it trades on the ambiguity of "average" and the difference between mode and mean.
these are the self-aware ones. I worked with and trained hundreds of scientists; I would say 95% of them had at best a shaky understanding of basic probability and statistical concepts!!
I think the average person is definitely intelligent. It's just that we are very prone to mixing intelligence with knowledge in addition to also forgetting how long it took us to learn things that are so automatic for us we see them as trivial.
Given that most measures of intelligence follow nearly normal distributions (as long as you stay away from the tails), I think it doesn’t matter that much.
Stats is one of those subjects which displays the dunning kruger effect really well.
Sure on the surface stuffile averages, medians, standard deviations etc are quite easy but the moment you start going deeper you realize how difficult and un-intuitive things start to get.
I'm still in the valley of despair, and frankly it's probably the best place to be in for the average person. Understand the basics and know just enough to realize how easy it is to mislead people with it.
Statistical literacy, or any literacy for that matter, requires hard work that is seen as a normal way of life. When machines do that hard work for us, it is no longer seen as a normal thing, but some unnecessary hard work (when we have cars, why walk 20 miles? And we have machines that do statistics).
Just like how those college kids saw my work as horrifically weird hard work, while I saw it as a normal thing.
My interpretation is how the risks (statistics) of those job functions are more known, people choose to avoid them. But the people that are-the-statistic don't necessarily agree with or understand those statistics.
For example, if OP is a coal miner that hasn't had health issues yet, they may choose to discount statistics that declare x% of coal miners have negative health outcomes.
Nope. It's a true story. And how it relates to thread is given below in another of my comments. "Things change with time" should have given you some relation. Statistics is seen by the current generation as hard work, rightly so, due to availability of easier ways of dealing with it.
Oh we used to dream of only walking 20 miles a day.
We used to get up at 3am half an hour before we went to bed, eat a lump of cold poison then walk FOURTY miles uphill to school then when we got home our Dad would slice us in two with a bread knife
The other 38% didn't understand the question...(my extrapolation)
I did a couple semesters of statistics at uni. And I can confidently say that the number of people who can answer 3 simple questions on statistics (like say mean versus medians, confidence levels or margins of error) is, well, a rounding error from 0.
Indeed, statistically, no-one has a clue how statistics work.
I did however learn enough to know that statistics can tell you absolutely anything you want them to say. Assuming you don't just make them up, they're trivial to manipulate to generate the headline you want.
When used to evaluate risk, the comprehension goes down further (a fact willfully exploited by any decent marketing.)
Statistically, most statistics are meaningless.
reply