Whenever this comes up, I think about the conjunction fallacy https://en.m.wikipedia.org/wiki/Conjunction_fallacy. The observation that human subjects seem to assign higher probability to joint events than a single event. Which is weird because the probability of two events at the same time (conjunction) is always less than or equal to the probability of a single event on its own.
How does the Bayesian brain hypothesis deal with this fallacy? It seems to me that nothing based on classical probability can explain this fallacy. So either the observation that humans can assign higher probability to joint events is wrong or human decision making isn't exactly probabilistic (in the classical sense, can't rule out exotic probabilistic approaches).
EDIT: As several folks have commented that the conjunction fallacy can be explained away by different arguments based on interpretation and semantic issues. Indeed, the original Linda problem was susceptible to these issues. However, since then several researchers have tried to study this effect more carefully and it seems to still persist. An example that I'm aware of is the following https://link.springer.com/article/10.3758/BF03195280 where the authors used unambiguous language and a betting paradigm, but still found the effect. Again, this is most likely not fool proof. Regardless, I do not think the fallacy can be trivially explained away as an effect of ambiguous language.
No, that's not what's going on. The reason for the fallacy is that we tend to find more detailed stories more convincing than less detailed stories.
However, not everybody, you can be trained against that. I've heard that police interrogators are less prone to this fallacy because they know that liars who had time to prepare often add dozens of details to their story that no ordinary person would remember.
> 'fallacy is that we tend to find more detailed stories more convincing'
Anecdotic: Somebody asked: "Why do I feel often that angry, when I get a 'Typ5-Answer'?" (With the background an TYp5-Answer is located in the field of the manipulation of (someones) reality.
HINT: Typ1-Answer: labeling of 'Objects' / Type2-Answer: naming of coherencies / Ty ....(-;
Good point. But... this particular objection has a flaw (which may be copied by the BBH).
The Bayesian machinery tells you how to update your beliefs given evidence. It doesn't tell you what shape those beliefs should be in the first place.
My theory is that we carry around a deck of personas or sterotypes. We hear of a new person, and some behaviour of theirs. We then predict which persona was likely to generate that behaviour. Conditional on the distribution over personas we answer the questions about that person's predicted behaviour.
In the `Linda` example the theory above suggests we take the background information about all their political commitments at college, which from the description would seem to give strong evidence about their persona.
The common and wrong answer to the question stems from predicting, based on the persona, that the person would continue with their political commitments.
The 'shape being wrong' issue here is that maybe personas are not the right way to structure the problem. But Bayes's theorem doesn't tell you that. That's a whole load of extra machinery that people have additionally developed and should deploy when using Bayes's theorem.
Back to the word problem. An issue is that the options:
- A
- A&B
Both include A. In a world where A always happens the only non-trivial way to read the question is that the first option must implicitly mean (A&!B).
If A is assumed to be true, the question then becomes is B more likely to be true or not true. The background information given is a reasonable explanation for the common answer of people selecting (A&B).
In more strictly mathematical terms, having a "normative" update rule (Bayes' rule) doesn't tell you what topology of latent variables the generative model "ought" to have, only how to link new information into a preexisting generative model.
Using the KL divergence of the posterior predictive distribution as a target to optimize does a bit better, but still isn't a "solution".
>Seriously, what doeos "topology of latent variables" even mean?
The simple answer is: the graph topology of the resulting program traces, equivalent to the topology of a graphical model sampled from a distribution over graphical models. The complicated answer is: the Scott topology of the program-trace space.
This is a fallacy of interpretation (translating verbal representations of events into mental representations of events), not of probabilistic reasoning.
The crux I think is that when
A. bank teller
is juxtaposed with
B. bank teller and feminist,
we vaguely and falsely interpret A as "bank teller and not feminist", while the correct explicit interpretation is "bank teller and possibly feminist, but possibly not".
Just to add something that I found interesting when my logic professor said it at the time: while that “bank teller and not feminist” interpretation is strictly incorrect in a theoretical world, IRL it’s useful for humans to just assume pieces of information. Most people are not particularly feminist, so it’s probably safe to assume someone - especially a bank teller, maybe - isn’t feminist if it’s not mentioned.
Rather than a fallacy of some kind, people may actually be so good at correct inferences that they are bad at leaving that intuition out of their reasoning process when thinking about a weird, outlying theoretical world.
It doesn't seem surprising at all that we fail to understand abstract questions like this. The experts only learned to answer them correctly after millennia worth of mathematical development!
To study how accurately Bayesian some animals are, I think you need to find ways to pose them questions which are relevant to them, and read off their inferences from their behaviour. This is obviously harder to do, and you have to worry a great deal about the animal optimising for something that differs from your first guess (e.g. it doesn't want maximum food on a good day, it wants not to starve on a bad day). I don't have references to hand but I think that when we can do this, the results are quite good.
That's still a fallacy, even if the mechanism behind is successful on average. Assuming a bank teller is non-feminist is rational only if the feminism matters, which it does not. Even if 0 bank tellers are feminist, it's still unhelpful (and harmful is even 1 bank teller is feminist) to assume that an unknown teller is feminist for the purpose of the question. Prejudices are often statistically more correct than incorrect (though socially problematic), but in some cases are still flat-out inorrect, as they are in this example. Too to much of a "good" thing is toxic.
The Wikipedia article on the effect mentions that the experiment has been done to correct that by explicitly wording it better and it reduces the magnitude of the effect but it still exists.
These are structured as: Do you choose (A) alone, or do you choose (A) and (B)? (A) is often benign and (B) is often something the subject thinks is a good fit. So, in some ways these seem like "trick questions" - I wonder what happens if both options are benign. As the Wikipedia article mentions, in the typical case the subject may end up choosing to use an easy heuristic rather than thinking hard about the actual definition of probability and such.
There are other ways the brain could be Bayesian though. For example, at the lowest level our neurons could use Bayesian inference. How this manifests itself at much higher levels such as with language and planning might be difficult to predict.
I would describe that heuristic as Bayesian, since the subject evaluates the prior evidence for each option, and then chooses the one with the most supporting information, instead of considering the options' logical or mathematical properties.
Most of the decision making while shooting for a hoop in basketball (the example from the article) is subconscious and reflexive, while decisions made related to the conjunction fallacy are conscious. So it's quite possible different mechanisms are used in subconscious motor -control reflexes and conscious 'logical' decision making.
I realise both are implemented using neural networks, but as an analogy digital computers can implement bayesian algorithms, boolean algorithms, predicate logic, etc. It's quite possible our neural systems have optimised to different behaviours to solve different problems, and in _some_ cases these implement or approximate to Bayesian methods.
I don't think that's right, or at least not so binary. The average person (and even highly rational people, when they are being informal) uses "logical" reasononing that is intuitive, not calculated, not so different fro "muscle memory" of shooting hoops. Even professional mathematicians, (in published papers!) sometimes make statements that "feel" correct even though they are disproven under scrutiny (and not due to small (but important) mistakes like getting a sign wrong in a calculation)
> How does the Bayesian brain hypothesis deal with this fallacy?
Maybe brain operates here on a syntactical level: which sentence is more likely to be present in a narrative about a subject and system 2 ([0]) fails to preempt the answer.
I just went through the preprint and I do not understand your comment. What specifically ticked you off? The preprint is well written, arguments are clear and there's enough background for an expert to work things out.
As Atiyah says in the preprint. The magic is the Todd function and the Mathematical framework that comes with it. It seems Atiyah has developed a new framework (which he calls Arithmetic Physics) and a side product of the framework you get a simple proof of RH. I don't know if the proof is correct. But I don't see any signs of crackpottery in the preprint.
Finally, this is in the style of Atiyah. He is known to be a "theory builder" rather than a "problem solver". True to that, he's claiming a whole new way of looking at number theory. So even if the proof turns out to be false. Mathematicians still get some new ideas.
No, it is not "well written". I'm no expert in analytic number theory, but here are some sanity checks:
His definition of the critical strip (2.4) is wrong.
He works with some family of polynomial functions who agree on the sets K[a] that have open interior (2.1). Of course, two polynomials that agree on infinitely many points are identical. So there really is not much to his "Todd-function". It is just a polynomial.
From his claims 2.3 and 2.4 then follows T(n)=n, for all natural n and hence T(s)=s, as T is a polynomial.
What does "T is compatible with any analytic formula" in (2.4) even mean? Does it mean "for f(X) a everywhere converging power series, then T(f(s))=f(T(s)), for s in C"? This can only hold for T(s)=s, again. So maybe it means something else? He applies it to f(X)=Im(X-1/2), which is not a power series, so what does he mean?
The Hirzebruch reference is a 250pp book. The paragraph on Todd-Polynomials (which are a family of multivariate polynomials, btw. There is no "Todd-polynomial" T in Hirzebruch!) does not contain a formula as claimed in (2.6).
Considering the last two breakthrough claims, that Atiyah made (no complex S^6 sphere and a new proof of Feit-Thompson) vanished in thin air, I remain more than sceptical that this "preprint" can be salvaged.
> He works with some family of polynomial functions who agree on the sets K[a] that have open interior (2.1). Of course, two polynomials that agree on infinitely many points are identical.
Consider f(x, y) := xy and g(x, y) := xy^2
Fixing x=0 note that f and g agree along {(x, y) | x = 0, y in R}. But f is not identical to g. There is no open subset of R^2 such that f and g agree throughout the subset.
Would rewording "two polynomials that agree on infinitely many points are identical" as "two polynomials that agree on any open set" fix this? Or restrict the statement to polynomials in one variable only?
You're right, I was thinking about univariate polynomials. The "preprint" is only concerned with functions of one complex argument, so this should suffice.
You'll have to go to the previous paper[1], where this Todd function is supposedly defined and...it's not defined there either. I mean, the whole thing is really sad and not something that I feel like talking much. Just give things a read.
Did you do an undergraduate math degree? There are some very real issues with the Todd function, such as no polynomial behaving the way he suggests for pretty elementery reasons. This doesn't even come close to meeting the standards of what a paper should look like, especially one that is claiming such a big result.
They're probably the most fundamental kind of reinforcement learning algorithms. Understanding bandit algorithms is crucial to developing a good understanding of RL.
>7. I saw that the only group of people able to preserve a minimum of humanity in conditions of starvation and abuse were the religious believers, the sectarians (almost all of them), and most priests.
This, to me is the point of religion. We need religion when things are hard and unpredictable. In a world where most things are certain and predictable, religion has no value.
Hard to tell. Some of the people who went through GULAG system were convinced atheists. Sergey Korolev, the father of Soviet space programme. Or Dmitry Likhachev. Lev Landau who was wrecked physically yet not in spirit. And certainly thousands of less notable people we never heard about.
It's not about being an atheist per se, but about believing. The others I'm sure strongly believed in something, which kept them going. Landau for example had physics to occupy him and keep him sane.
Here are some basic things, the predictability of which most people in the developed world enjoy.
1. Clean water
2. Roof over the head and the sense of safety that comes with it.
3. Basic food.
A sense of community and friendship is probably the only basic human necessity that is not certain in the developed world.
Now imagine a world where none of this is can be taken for granted. You live under a constant threat from various sources (disease, wild animals, other humans out there to steal, loot from you etc) , there's not enough food, water and on top of that you're lonely. Religion is one driving force that helps people through these.
And then - BAM! You've (or someone from your relatives) got cancer. Or something.
- True, man is mortal, but that is itself only half the evil. The trouble is that man is sometimes suddenly mortal, that's the tricky part! Basically, he can never say what will happen to him this evening.
'What an idiotic way of putting it...” thought Berlioz, and objected:
- Certainly, that is an exaggeration. I know more or less exactly what will happen this evening. Of course, if a brick falls on my head on Bronnaya...
- Bricks are out of the question, - the stranger broke him off sharply, - not a single brick will ever fall on anybody's head. Under no circumstances, I assure you, does this constitute a threat. You will die a different death.
- And perhaps you know just which? - inquired Berlioz with the most natural irony, he had clearly been drawn into some kind of absurd conversation, - and can tell me?
- Certainly, - responded the stranger. He measured Berlioz with his gaze, as if he were sewing him a suit, and mumbled through his teeth, something like: 'One, two... Mercury in the second house... the moon is down... six - misfortune... evening - seven...” - then he loudly and delightedly proclaimed: - You'll have your head cut off!”
If a rich person's got cancer - he'll be able to pay off his dying and live in acceptable conditions before he passes.
While someone poor would probably end up being unable to pay for his painkillers and dying while praying to his relatives so that someone killed him not to endure his death.
So no, rich people do have a better ability to hedge their risks.
This has nothing to do with the initial statement though.
Not to mention that the result is the same anyway.
>live in acceptable conditions before he passes.
And no, this is not as simple as you are trying to picture. Sometimes - yes, you can deal with it this way, in many other scenarios painkillers and other shit is not enough and the only way to stop the suffering is either forced coma or suicide.
> * The more white and privileged you are - the more predictable your life is.*
That's a quantifiable claim, especially in response to an article about a gulag system that killed and tortured many white people, many of whom you may have called privileged before they were imprisoned.
I would love to see your justification. Seeing a scatter plots with whiteness and privilege on the horizontal axis and predictability on the vertical would seem persuasive, though I'm not sure how you'd quantify the variables.
Gulag system tortured political dissidents, who could or could not have been rich or privileged.
Those who were privileged in the Soviet Union have always had a better hand. Don't have a job? Here's a job for you. Your brother's got cancer? We'll find the best doctors for him. Can't find a TV set for your nice flat? Here's a guy you could call. Your idiot child has been speeding and killed someone walking down the street? Well, we may speak to those policemen so that his life would not be destroyed by such a small mistake.
I understand why you picked on the phrase "white and privileged", but it has nothing to do with the substance of my statements.
I am incredibly certain that a privileged person can afford to have a better treatment than a poor person. Which is literally what "hedging your risks" means.
Look at someone speaking about cancer over there: the guy's clearly missing the point that if you're rich and you die from cancer - then you're probably still going to be better off than if you're poor and die from cancer.
Very little. He did some work on what would now be called mutational "recovery" of frame shift insertion/deletion variants in phages, for about a year on sabbatical. Which at the time would have been enough for him to have predicted several aspects about the genetic code, such as DNA triplets forming codons. Which was funnily enough anticipated by George Gamow around the same time looking at the number of naturally occurring amino acids and the number of nucleobases.
The numbers are so tiny that comparisons are pretty much irrelevant. Since 1996 there have been just 8 mass k-12 shootings (incidents involving 4 or more school deaths, excluding the assailant), in a nation of 330 million people.
That's so far into small numbers territory that any comparisons are guaranteed to be overwhelmed by noise.
When making a decision as to whether something is small, large, or tiny one needs a scale. What's the scale here? For example, if the rest of the world combined has had 12 mass k-12 shootings, then American accounts for 40% of the total shootings which is a huge fraction. So in my original question I was trying to find the right scale.
That number seems suspicious; you'd get a very different answer if your criterion was "any gunshot wound at a school", for example. But that definition might be a better match for what people think of as a school shooting.
(e.g. Northern Ireland had a very large number of terrorist attacks where a warning was given allowing evacuation - would they not count as terrorist incidents even if nobody was killed? I suspect not)
Any gunshot would at a school probably includes a very different pattern of behavior. If someone wants to murder (or maim) a specific person, or small group of people, they may carry that out with a gun, because it's expedient, but might switch to a knife if guns aren't available. If someone wants to murder (or maim) a large number of people, guns or explosives are really the only practical means; mass knifings have happened, but are even more rare than mass shootings.
How does the Bayesian brain hypothesis deal with this fallacy? It seems to me that nothing based on classical probability can explain this fallacy. So either the observation that humans can assign higher probability to joint events is wrong or human decision making isn't exactly probabilistic (in the classical sense, can't rule out exotic probabilistic approaches).
EDIT: As several folks have commented that the conjunction fallacy can be explained away by different arguments based on interpretation and semantic issues. Indeed, the original Linda problem was susceptible to these issues. However, since then several researchers have tried to study this effect more carefully and it seems to still persist. An example that I'm aware of is the following https://link.springer.com/article/10.3758/BF03195280 where the authors used unambiguous language and a betting paradigm, but still found the effect. Again, this is most likely not fool proof. Regardless, I do not think the fallacy can be trivially explained away as an effect of ambiguous language.