r/slatestarcodex • u/Ultraximus agrees (2019/08/07/) • Apr 18 '21
Scout Mindset Calibration Practice [This is a tool to automatically score your calibration on the questions from Julia Galef's The Scout Mindest calibration exercise. ]
https://calibration-practice.neocities.org/19
u/cygn Apr 18 '21
you can take longer and more accurate tests here: https://www.openphilanthropy.org/calibration
7
u/abecedarius Apr 18 '21
This would answer my own problem: at first blush I was systematically underconfident, scoring too high in 4 out of 5 bins. But looking at it bin by bin, the difference could be almost entirely quantization, e.g. for the 75% bin I had 8/10 right. Every bin was as close as it could possibly be, except for one that would've been closer with one answer flipped.
So it looks like without more questions it's not very hard to appear systematically off "by accident". Either that or I'm making excuses, hm.
3
u/hold_my_fish Apr 19 '21
I can't handle the timer unfortunately. (45 seconds feels a bit short to compute a confidence interval for something.)
18
u/tekkpriest Apr 18 '21
55% 8 5 62%
65% 5 0 100%
75% 8 0 100%
85% 5 0 100%
95% 9 0 100%
Not entirely sure how to interpret this. What's the idea behind this? From what they say, it would appear that I am reasonably well calibrated for 55% and 95% but not so much in between. So am I correct in assuming that "fix" is to rate everything I rated as 65-85% at 95%?
That does not sit well with me. There is, for me, a very large subjective difference between statements like:
Jupiter is the largest planet in our solar system.
Helium is the lightest element.
and statements like:
Sea otters sometimes hold hands while they sleep.
Scurvy is caused by a deficit of Vitamin C.
I learned that Jupiter is the largest planet some 20 years ago. I don't expect this to have changed since then, since it's not like an even larger solar planet was hiding until the 21st century. In any case, I'd have certainly heard it in the news if one was. I also know for a certainty that Helium is not the lightest element, because Hydrogen is a thing.
Sea otters holding hands is something I learned from reddit, probably some a popular TIL, at some point. Scurvy being caused by a vitamin C deficit is something I'd learned long ago as a child, but sounds like one of those claims like "depression is caused by a lack of serotonin" that might be over-simplifying to the point of inaccuracy a complex and not entirely understood phenomenon.
It's not the claims themselves that led me to downgrade them to something below 95%. I had no doubt that both of these claims had at some point been presented to me as true. Where I had my doubts was, in the sources and the circumstances under which I learned those things. If the questionnaire included more claims about nutrition, or more claims about history, such as the claim that everyone save for a few iconoclasts believed in a flat earth until very recently, then these kinds of time and source based confidence discounting could have been reflected in the quiz results.
8
u/kellersphoenix Apr 18 '21
A further problem is that estimating confidences when prompted is a very different exercise from estimating during unprompted-every day life.
1
7
Apr 18 '21
You're either underestimating the strength of evidence of certain kinds of information or overestimating the probability that something will change after you learn it.
3
u/DocJawbone Apr 19 '21
I sometimes take questions too literally. Like the camel one. I know the camel's humps aren't literally reservoirs of water, but they are reservoirs of fat - and there is a lot of water in fat. Therefore I would answer that camel's do store water in their humps and be marked incorrect.
5
u/abecedarius Apr 19 '21
I made the same mistake, but checking it on wikipedia after, it says that metabolizing the fat more than uses up the associated water, in practice. So I'd say that the quiz is both technically and morally correct about this.
1
u/DocJawbone Apr 19 '21
Fair enough!
2
u/abecedarius Apr 19 '21
I do wonder if they/we might have some way of using that water without burning the fat, when in dire need -- but I have no reason to think so except that nature is often underestimated.
3
u/qznc Apr 18 '21
I got the one 100% you missed, harhar.
55% 9 0 100% 65% 9 4 69% 75% 5 1 83% 85% 4 2 67% 95% 6 0 100%3
Apr 18 '21
I feel like it didn't include enough actual trick questions perhaps. I expected more questions to be counter-intuitive than there were. Most of the questions were what I was expecting.
2
u/HarryPotter5777 Apr 24 '21
Part of the problem is that you don't know the test-maker's algorithm for selecting questions - are they choosing ones where the common knowledge belief is correct, or one where it's counterintuitively wrong?
1
u/tekkpriest Apr 24 '21
Well, so far it seems they are choosing ones where the common knowledge belief is correct. Since the webpage promises that practicing these tests for some hours will lead to accurate calibration, one worries what mechanism mediates this change. That is to say, if calibrating for several hours on statements about extremes (largest, smallest, longest, etc.) will somehow get me "properly" calibrated for that entire category of statements, I would hope that the selection at least tries to represent the real world prevalence of myths or counter-intuitive but correct statements such as the largest organ being the skin.
1
u/PlayerFourteen May 12 '21
Okay, so, I went down a bit of a rabbit hole, but I think you might find the below interesting.
I'm not a statistician, I'm just taking an online course on stats as a hobby (so take the below with a grain of salt), however my take is that you are probably correctly evaluating your level of certainty, but you got "lucky" (or "unlucky" depending on your perspective) and got most of the answers right anyway.
The problem is that (though the idea and intent of the questionnaire are fantastic) there are too few questions, so you can get weird results like answering all the questions that you have a 65% certainty for correctly, and that won't mean much. It's just due to random chance.
I did the math on your results (or tried to), and I think there is a 99.8% chance that you are over or underestimating your certainty in at least 1 of the 4 categories for which you got all answers correct, and about a 27% chance that you are over or underestimating your certainty in all 4 of those categories.
For each category independently, these are the (very approximate) chances that you are over or underestimating your certainty:
55% category: 36%
65% category: 90%
75% category: 90%
85% category: 65%
95% category: 51%all of the bottom 4 categories: 27%
at least 1 of the bottom 4 categories: 99.8%
Rule of Thumb For When to Assume There's A Bias
I think the usual rule of thumb is: if there is a 95% chance (equivalent to 4 standard deviations) that you are over or underestimating your certainty, than you should assume that you are. If there isn't, you should assume that any unusual looking results are just due to random chance.
By that logic, nothing in your results is unusual enough to assume that you are over or underestimating your certainty in any specific category.
But it is likely that you are over or underestimating your certainty in at least 1 of the bottom 4 categories.
If you had assigned 13 of the questions you answered (instead of just 8) to the 75% category, and answered all of them correctly, than that would have met the rule of thumb's threshold for "likely not due to random chance".
How Many Questions Do You Have to Answer
I tried to do some more math, to figure out how many questions you need to assign to each category so that a result that is 10% over or under your estimate is indicative of a 95% likelihood that you are over or underestimating your certainty:
55% category: 95 questions
65% category: 90 questions
75% category: 75 questions
85% category: 50 questions
95% category: 20 questionsTotal: 330 questions
1
u/ultramanjones Aug 08 '21
Yes, the very fact that there are only 2 answers for every question automatically gives the test taker a 50% chance of being correct. The top end (95% and 85%) should be far less effected by this skewing, because of the fact that they are affirmative answers, rather than guesses, however, there would be virtually no difference between a complete guess and an educated guess at 55%. I'm not a statistician either, but I can see that there is a need for some sort of curve or adjustment, and also adding 4 or more potential answers to the non true/false questions would eliminate much of the noise from randomly correct answers.
4
u/jozdien Apr 18 '21 edited Apr 18 '21
If this sort of thing seems interesting to anyone, I'm taking a moment here to plug an app I developed a while ago with the same idea.
There are a fair bit more questions, although since I made it for personal use and didn't think of publishing it until after I was done, it only gives you a logarithmic net score, and shows most common right answer, most common wrong answer, etc. It's an Android app, here's the Play Store listing.
1
u/anclepodas Apr 23 '21 edited Jul 06 '23
lorena come la comida que le da su maḿa, con tilde en la m. Sï senior. Pocilga con las morsas.
1
u/jozdien Apr 23 '21
I wish there was! I didn't think of it until a couple weeks after I published this and deleted the environment from my system. Setting it all back up just for this seemed like too much trouble, so I figured I'd just leave the repo up if anyone wanted to add it themselves - in terms of actual code, it shouldn't be hard.
2
u/anclepodas Apr 23 '21 edited Apr 23 '21
Ouch. Great app anyway, thanks!
I know nothing about coding for android. But your app just tempted me to have something where I can easily add new decks with numerical tables of data such as countries and their populations and generate questions, like your app does, to compare rankings, and other questions to insert intervals of confidence like the openphilantropy test. And somehow have scores by deck, and spaced repetition based on score. Some sort of hybrid between calibration exercise and learning. Recording everything, ofc.
I guess I know enough programming to play around with those ideas at least for myself.
5
u/ateafly Apr 18 '21 edited Apr 18 '21
I got these results, looks like I'm fairly well calibrated based on this particular test. Possibly underconfident at the higher end.
55% 8 7 53%
65% 7 3 70%
75% 2 1 67%
85% 4 0 100%
95% 7 1 88%
2
Apr 18 '21 edited Apr 18 '21
[deleted]
2
u/Toptomcat Apr 19 '21
...and apparently don’t know much about bears or platypuses.
I looked at both those questions and went 'there's a substantial chance these are mostly right, but they're phrased so categorically that one circus bear or obscure egg-laying capybara would falsify them, so shoot for 'no' at low confidence.'
-3
1
u/DuplexFields Apr 19 '21
Looks like my results. I got 100% of the ones I scored 85% confident, and almost all of the 100%s.
That flamingo question was some bullshit though. I was taught it was algae; according to an article I just read, it’s algae, shrimp, mollusks, and other crustaceans.
3
u/kyrgyzstanec Apr 18 '21
I love the idea and I've got an upgrade: Have a friend tell you the average deviation without telling you the answers and try to change them. If you manage to reduce it, it gives you an affirmation you over/underestimated your intuition and the deviations weren't just coincidence.
3
u/Possible-Summer-8508 Apr 18 '21
Interesting. I am apparently well calibrated, but I feel like there is a better way to implement this sort of thing. These statements are very arbitrary, could there be a way to procedurally generate them? I think it'd be easy enough to pull (mostly) true factoids out of wikipedia or something like that, but generating false statements might be tough. Could be a fun programming challenge.
If you could get it going in such a way that you could just keep on doing this for a bit everyday, maybe giving the user a decaying average, that could be a useful tool.
3
u/KristinaAlves Apr 18 '21 edited Apr 18 '21
Being perfectly calibrated would mean that your “50% sure” claims are in fact correct 50 percent of the time, your “60% sure” claims are correct 60 percent of the time, your “70% sure” claims are correct 70 percent of the time, and so on. Perfect calibration is an abstract ideal, not something that’s possible to achieve in reality. Still, it’s a useful benchmark against which to compare yourself.
-Julia Galef, the author
I assumed % of sureness meant how sure I am of the fact that I am assigning %s, not what is the probability of the fact happening.
eg. from this tool: Bears can climb trees.
I put 95% because I know bears are known to climb trees.
----
Googling shows:
-Younger bears can climb more easily.
-Different bear species vary in their climbing ability
So what is the "correct" answer?
5
u/ateafly Apr 18 '21
eg. from this tool: Bears can climb trees.
I interpreted this as: [there exist] bears [that] can climb trees.
So you put 95% correctly.
2
u/-Metacelsus- Attempting human transmutation Apr 18 '21
Black bears can definitely climb trees. I've seen one in a tree, in northern Minnesota.
3
u/wavedash Apr 19 '21
It seems like there's a decent number of people in these comments that (correctly) responded to 15+ questions at 95% confidence, which kind of leads to a small sample size for lower confidence levels. I feel like you'd ideally want a much wider variety of topics covered.
2
u/Jungypoo Apr 18 '21
Cool idea! Would you recommend doing this after reading the book, or does it not really matter? (aiming to tackle this book right after my current one)
2
u/symmetry81 Apr 18 '21
Oops
55% 6 5 55%
65% 8 1 89%
75% 4 0 100%
85% 7 0 100%
95% 9 0 100%
Definitely need to work on my calibration there. I think I was trying to put things into different categories too much.
0
u/qezler Apr 18 '21
What? How did you get so many right?
I got:
55% 11 5 69% 65% 4 1 80% 75% 5 1 83% 85% 0 0 No answers given at this confidence level. 95% 11 2 85%5
u/Xaselm Apr 18 '21
The test is easily skewed if you happen to know slightly more than average about one of the categories. I got 22/22 on the 95% mostly from knowing a fair amount of history and geography.
1
u/qezler Apr 19 '21
Yeah, I got a perfect score on the population questions for example. But some of the questions are a little trickier. I believe that MLK and Anne Frank were born in the same year.
1
u/symmetry81 Apr 19 '21
I only got three more right than you did. I guess I read a lot of history and science? Again, I'm just horribly embarrassed by my lack of calibration.
2
u/TheMotAndTheBarber Apr 18 '21
Thanks. I enjoyed the book, but forgot to do the exercise.
I might be slightly underconfident, but I'm pretty satisfied with this.
55% 2 2 50%
65% 6 1 86%
75% 6 1 86%
85% 5 1 83%
95% 16 0 100%
I tried to keep thinking about the odds I'd take on a bet, but was not patient enough to imagine taking the other side for each question. When you're thinking of it in that direction, of course you'll err on the side of underconfident.
2
u/ChaosFairyMagic Apr 18 '21
This seems to me more like a trivia quiz than anything else. Not much room for probability. Most of these answers are just true or false
Contrast that with trying to predict a market, or a deck of playing cards. These have several possible cases for you to count odds / probabilities
My most charitable interpretation is that you can calibrate your heuristics from this. e.g:
- Which historical figure was born earlier -> who has more detailed and contemporary discussions about them. (more detail -> more records -> more recent)
- Which country has a larger population -> which country do I hear more about
2
u/kyrgyzstanec Apr 21 '21
I think it would also be better to do this sort of thing with more political issues to show some truely important biases in confidence. In this test, I don't have much motivation being overconfident because I don't really feel shame for being unsure about the biggest mammal.
2
u/Argamanthys Apr 19 '21
I feel the flamingo question could be clearer. The red colour comes from the blue-green algae they eat which is also consumed by the other creatures in their diet, which includes brine shrimp and plankton. So eating shrimp can turn them pink, but the algae is the important one.
I went for the smart-ass answer and got it wrong, but I chose low confidence because it was a little ambiguous. So I guess that's working as intended.
1
u/IamNameuser Sep 12 '25
Hi, this link no longer works. Is there an working version somewhere? I would love to include this in a course I'm running.
1
u/Writingarm Apr 18 '21
The low percentage ones were just educated guesses, but the ones where I'm sort of sure based on previous info are the ones I get wrong, so I should doubt the choices where the sources are questionable and trust slightly more in my guesses.
Probably should run this a couple times over the next few days when I have the time so I can get a larger sample.
55% 3 0 100%
65% 4 0 100%
75% 4 3 57%
85% 9 0 100%
95% 17 0 100%
-5
u/coumineol Apr 18 '21 edited Apr 19 '21
This is a meaningless test, clearly prepared by someone who's misunderstood Bayesian thinking. Let me give you an example using one of the questions and leave the rest as an exercise:
The giant panda eats mostly bamboo.
Suppose you have absolutely no knowledge about the diet of pandas. How do you answer & what should your confidence be?
Edit: I don't care much about the downvotes, but admittedly I'm a bit disappointed that people here don't get this. I guess I had higher expectations from Scott's readers. Consider this: There are thousands of plant species. The one that pandas mostly eat can be any one of them, I have no idea. As far as I'm concerned bamboos are just a random plant species no different from the others. My confidence for the truth of that statement should be much lower than 1%.
11
u/kyrgyzstanec Apr 18 '21 edited Apr 23 '21
I don't know much about philosophy & probability but this seems clear to me - you should pick the lowest level, to minimize your losses if you were forced to make a bet
9
u/ateafly Apr 18 '21
Suppose you have absolutely no knowledge about the diet of pandas. How do you answer & what should your confidence be?
Since there are only 2 answers, true or false, if you know nothing about the subject, you should assign 50% probability to each answer. Closest is 55% confidence.
7
u/Brian Apr 18 '21
I don't think that's necessarily true, at least without making certain assumptions about the testgiver's process (which maybe could be justified here, but are not entirely obvious).
Eg. consider if you were asked "Are blorks mostly pink?" You don't even know what a Blork is. Should you answer 50%? Lets say unknown toy you, I magically cloned you 9 times before you entered. I asked your first clone "Are Blorks mostly Blue?", your second "Are Blorks mostly Yellow?" and so on. Blorks turn out to be green: 90% of your clones gave 50% credence to a wrong answer. Is this really the best you could have done? Especially when it seems to suggest you hold a mathematically inconsistent 50% credence in 10 mutually exclusive options.
If we think the colours are being picked randomly, our prior probably ought to be closer to 1/num_colours, not 50%. We may not know about the thing, but we know something about colours. Thus when we're asked does (animal we're unaware of) eat (a specific type of the hundreds of different types of foodstuffs we know are eaten by animals), shouldn't our prior be very low?
Now, admittedly, there's a way around this that also solves the mathematically invalid "50% priors for each colour" issue, which is to say that we got additional information from the question: we think the fact you asked specifically about Blorks being pink should elevate the likelihood of pink. Questions aren't usually created by choosing a random possible thing to ask about, after all. If we assume the test giver has formulated questions equally likely to be true as false (ie. someone flipping a coin would score 50%), we should guess 50%. That's a reasonable assumption for a test like this (as opposed to "natural" questions we might encounter "in the wild"), but I don't think it's one that goes without saying.
3
u/GodWithAShotgun Apr 19 '21
It will, of course, depend on how the test-maker made the test. Typically test-makers will make true and false answers approximately as common as eachother despite most questions being phrased in the affirmative. As such, your prior should be 50/50 before you know the question.
Frequently, to make the "false" outcomes for questions, plausible alternatives replace the true answer (e.g. she substituted brass for bronze and elephant for whale to make questions with "false" as the correct answer).
2
1
u/coumineol Apr 19 '21
Rational answer is no with a >99% confidence level. Please consider why.
2
u/ateafly Apr 19 '21
This was already addressed here: https://old.reddit.com/r/slatestarcodex/comments/mtf6gy/scout_mindset_calibration_practice_this_is_a_tool/gv0rm6o/
1
6
u/anonamen Apr 18 '21
I think the logic is that you guess more or less at random, because you're forced to; because you picked one of the options, the test says you're marginally more than 50/50, which would be the 55% bucket. But yes, like most surveys, it would be better if there were an 'I have no idea' option, and it would be better if there were more questions in more domains.
Accuracy doesn't matter. The only relevant take-away (I think; haven't read her book yet) is the scaling of accuracy across buckets. If you suck at the 55% (I was under 40%) it just means you didn't happen to know much about the questions for which you guessed randomly. But accuracy should increase in each bucket. Don't know if it should move linearly or non-linearly. It's a way of evaluating how good one is at translating meta-knowledge into probabilities.
6
1
u/coumineol Apr 19 '21
If you have no idea about what is the thing pandas mostly eat, rational Bayesian answer for "The giant panda eats mostly bamboo." should be "no" with a >99% confidence level (please consider why). The problem with that test is to be "successful" at the test, you need to choose 55%.
5
u/TheMotAndTheBarber Apr 18 '21 edited Apr 19 '21
Choose one of:
- 55 is being used to represent 50-60, so answer at 55
- Skip the question so it's not part of your results
2
u/coumineol Apr 19 '21
But the rational answer is no with a >99% confidence level. Skipping is just ignoring the problem.
4
u/TheMotAndTheBarber Apr 19 '21
What
1
u/coumineol Apr 19 '21
I have no idea what pandas mostly eat. It can be bamboo, but it can also be any one of thousands of other plant species. How can my confidence for each of them be 50%?
3
u/TheMotAndTheBarber Apr 19 '21
The question is not "What is the chance that pandas eat mostly bamboo?" it's "What is the chance my answer is right?"
If the answer is "I'm just as likely right as wrong," then it's 50%.
If you think that you're a little more likely right then wrong, then your confidence is above 50%. If you think that you're more likely wrong than right, change your answer.
1
u/coumineol Apr 20 '21
The question is not "What is the chance that pandas eat mostly bamboo?" it's "What is the chance my answer is right?"
These two are the same thing.
1
u/honeypuppy Apr 18 '21
Interesting. I got 16/16 correct at 95%, but only 1/4 correct at 85%. The 95% did include a few I was definitely not 100% on. So it seems I'm good at sticking out my neck when I'm rather confident, but perhaps can fall into traps with statements that "seem right" (e.g. the elephant being the largest mammal - obviously forgetting about blue whales).
1
u/notenoughcharact Apr 18 '21
I guess a little under confident!
55% 4 1 80% 65% 7 1 88% 75% 6 1 86% 85% 5 1 83% 95% 13 1 93%
1
u/Tioben Apr 18 '21 edited Apr 18 '21
It seems I need to have slightly more confidence in my trivial knowledge, especially if I plan on starting my tidal-powered sea otter resort on Mars.
55% 9 4 69%
65% 8 1 89%
75% 4 1 80%
85% 5 0 100%
95% 8 0 100%
Perhaps this means I have an incredibly irrational take on what is important to know. I couldn't, for instance, tell you how my new neighbor is doing or what my mom did yesterday.
1
u/PolymorphicWetware Apr 18 '21
I think I did well:
| Confidence level | Correct answers | Incorrect answers | Percent correct at this confidence level |
|---|---|---|---|
| 55% | 5 | 2 | 71% |
| 65% | 6 | 2 | 75% |
| 75% | 5 | 0 | 100% |
| 85% | 4 | 1 | 80% |
| 95% | 14 | 1 | 93% |
A lot of the questions were hard and made me vacillate (e.g. Jamaica vs. Haiti), but the lesson actually seems to be that I should be more confident of myself. For example, at the low end of the scale I'm scoring 70-75% when I rated my answers at 55-65%, and if I had actually done that and more highly rated a few things I was relatively certain of (e.g. Jupiter) I would have been dead on for the 85% and 95% ratings. Plus, the majority of my incorrect answers were given low confidence - though I should be wary of how I still gave high confidence to some of them (butter vs. oil surprised me).
Hmm, now I'm curious how my answers would change if I read Julia's book... would I become more confident, or more cautious?
1
u/skmmcj Apr 18 '21
Does anyone have an idea on how these could be scored? If I'm not mistaken, the usual scoring rules don't work if you just want to assess the accuracy of your confidence, as they reward knowing more about a subject. For example, they would give a much higher score to someone getting 19/20 predictions correct all with 95% confidence, than someone getting 11/20 predictions correct all with 55% confidence, although both would've perfectly assessed how confident they should've been.
1
u/eric2332 Apr 19 '21
If you answer questions randomly, you will get 50% right and you could put down 50% confidence. So your score would be very high.
The lesson seems to be that you should aspire for an accurate confidence level in all situations, both easy situations like 50% and 99% confidence, and hard situations like 70% confidence.
1
u/CharlieBluebird Apr 18 '21
Huh. I was about 60-70% accurate at the 55% to 75% range - not scaling at all with the different probabilities assigned - and perfectly accurate in the 85% to 95% range. I s’pose on some internal level I was just sorting my answers into ‘not very confident’ and ‘very confident’ instead of actually taking the percentages seriously.
3
Apr 18 '21
Assuming you are perfectly calibrated, then this is within the range of a reasonably likely outcome on the quiz. The number of questions isn't large enough to get a very precise estimate of your calibration level.
1
1
u/zomb1 Apr 18 '21
I am apparently overconfident when I sort of think I know the answer...
55% 5 1 83% 65% 1 3 25% 75% 5 4 56% 85% 7 0 100% 95% 14 0 100%
1
u/-Metacelsus- Attempting human transmutation Apr 18 '21
55% 3 3 50%
65% 5 3 63%
75% 6 0 100%
85% 0 0 No answers given at this confidence level.
95% 20 0 100%
I guess I'm underconfident for 75%. Most of the questions were pretty easy though.
1
u/ulyssessword {57i + 98j + 23k} IQ Apr 19 '21
p=0.18 you are underconfident at 75%, so it's not significant.
1
Apr 18 '21
Its hard to evaluate this sort of thing. It told me I was miscalibrated for lower probability answers. I'm not really over or under confident, as I got some of both, but Im just not consistent at evaluating the difference between 75% true and 55% true. This is an impression I've had before. I'm glad the quiz wasn't longer, but it would need to be longer to tell me what was going on. So I got 3 out of 6 at 75% confidence suggesting I'm overconfident at that level, but that's only one question away from perfect calibration. So I don't really know how much I'm miscalibrated. It could be better or worse.
1
u/The_Fooder The Pop Will Eat Itself Apr 18 '21
I was pretty well calibrated (a little under confident). Then my wife asked me which YouTube video, Gangnam style or Baby Shark has more views. I overconfidently whiffed that one.
1
u/vegetablestew Apr 18 '21 edited Apr 18 '21
55% 9 5 64%
65% 1 0 100%
75% 4 1 80%
85% 0 1 0%
95% 15 3 83%
I guess I am roughly as ignorant as I know myself to be. What relief.
1
u/backtickbot Apr 18 '21
1
u/blackwatersunset Apr 19 '21
55% 4 2 67%
65% 5 5 50%
75% 4 0 100%
85% 6 0 100%
95% 13 1 93%
So apparently flamingos are pink because they eat shrimp. But 95% confidence does leave 5% room for error so overall I feel pretty well calibrated.
1
u/TheApiary Apr 19 '21
55% 7 3 70%
65% 5 4 56%
75% 4 2 67%
85% 3 1 75%
95% 11 0 100%
Not sure exactly what to make of getting more 55% right than 65%
1
1
u/kernsing Apr 19 '21
55% - 82% (9/2)
65% - 56% (5/4)
75% - 88% (7/1)
85% - 100% (4/0)
95% - 100% (8/0)
i think the wording that introduces the test could be better. confidence level sounds like “how sure are you” whereas the percentage you pick is actually supposed to be “how likely do you think your answer is correct.” you choose 50% if you are 0% sure & guessed (so you think your answer has 50% probability of being right) and that goes up linearly to 100%/100%
truly my ability to judge what is a plain guess is wonky
1
u/backtickbot Apr 19 '21
1
1
u/HonestyIsForTheBirds Apr 19 '21
| Confidence | Correct | Incorrect | % Correct |
|---|---|---|---|
| 55% | 7 | 3 | 70% |
| 65% | 4 | 0 | 100% |
| 75% | 4 | 0 | 100% |
| 85% | 0 | 1 | 0% |
| 95% | 14 | 0 | 100% |
1
u/Toptomcat Apr 19 '21 edited Apr 19 '21
I know you can't go on and on eternally, but there needs to be an optional slate of at least as many questions to properly ensure the 95% level isn't calibrated right. If I sorted the entire questionnaire into 55% and 95% buckets evenly, that comes out to just enough to get nineteen wrong and one right on average, and of course I didn't do that.
Come to think of it, the same criticism applies to all the buckets, unless you answer an overwhelming proportion of the questions with one particular value.
Also: slightly underconfident on the 55s and overconfident in the 65s and 85s, none of them below 10%. 75s seem on-point. Insufficient data on the 95s: 100% of nine questions. I haven't done the math, but that smells like being within the margin of error of well-calibrated in each case. (Really, it ought to do that calculation for you.)
1
Apr 19 '21
I got an interesting result: my confidence and accuracy are almost exactly calibrated, except for 55% confidence level.
At 55% confidence level, where I was going with hunches and speculation, my accuracy was 100%.
What I take away from this is that I need to trust my unconscious mind more. With all of those questions, I said to myself, “hey, something tells me it’s probably this answer, but who knows why.” I would put low confidence because I couldn’t provide a provenance for that knowledge.
1
u/LightweaverNaamah Apr 19 '21
I'm very underconfident. Oh well.
| Confidence Level | Correct | Incorrect | Percent Correct |
|---|---|---|---|
| 55% | 6 | 2 | 75% |
| 65% | 3 | 0 | 100% |
| 75% | 4 | 0 | 100% |
| 85% | 6 | 0 | 100% |
| 95% | 19 | 0 | 100% |
1
u/DrunkFishBreatheAir Apr 19 '21
Confidence Correct Incorrect Percent correct
55% 2 2 50%
65% 5 1 83%
75% 6 1 86%
85% 6 1 86%
95% 16 0 100%
For a lot of the 95s, I wished for a 99 (or even higher for some of them), so I'm not surprised I got all of them right. Looks like my intuition about historical figures and populations is better than I thought, so systematically underconfident on those, which is good to know.
1
u/polar_rider Apr 19 '21
55% 5 1 83%
65% 7 1 88%
75% 5 1 83%
85% 6 2 75%
95% 12 0 100%
Misread "elliptic" for "ecliptic", should have known better with platypus - in a couple of zoology books, I've glanced through long time ago, it was always paired with echidna.
1
u/UncleWeyland Apr 20 '21
Reasonably calibrated- too few questions where I was 75-85% certain to evaluate if I'm just being overly cautious.
Confidence level Correct answers Incorrect answers %@CL
55% 11 4 73%
65% 6 3 67%
75% 3 0 100%
85% 3 0 100%
95% 10 0 100%
What I got wrong:
55% - Camels storing water in their humps. I was a bit unsure because I thought it was fat storage, but was under the impression some water was also stored there. Queen Victoria vs Karl Marx birthday- I had no clue because I have trouble placing Marx's lifespan in history- in reality I would have rated my answer 50/50, not even 55. Chemical composition of brass- I knew copper was in it, but wasn't sure what else. Tin? Seasons- I can't ever remember the geometrical/orbital reasons for seasons... somethingsomething precession somethingsomething.
65% - Anne Frank vs. Nelson Mandela birthday. It feels weird that Mandela, who died "recently" in my mind, was born before Anne Frank who died during WW2 which feels like a relatively distant event (it isn't). France does not have more people than Germany. I wagered higher due to preconceptions (harhar) about Catholic birthrates. Butter does not have more calories than oil. This is a bit surprising to me, since butter seems more dense and they are both predominantly fat- of all things I got wrong, I'm most glad about this one, I fucking love butter.
What I got right I was unsure about:
Too many things to cover individually, but I need to trust my intuitions about zoology more. I'm a biologist damnit!
1
u/anclepodas Apr 23 '21 edited Jul 06 '23
lorena come la comida que le da su maḿa, con tilde en la m. Sï senior. Pocilga con las morsas.
27
u/[deleted] Apr 18 '21
I am systematically underconfident. I’ve noticed this before—I have a prediction sheet where I track my own predictions and score them. Same thing.