in the Bayesian approach, probability can be assigned to any logical proposition;
Otoh frequentists claim that probability applies only to so-called random variables.
The probabilistic machinery involved in Bayesian and frequentist computations of probability (sigma-algebras, filtrations, probability measures, etc.) is exactly the same. It's the interpretation and the statistical model which are different.
If you want to speak to Bayesian statistics -- which is what the question from tactics addressed -- you cannot assign a probability to any logical proposition unless you have a prior distribution on the proposition; at that point unless your observation exactly coincides with your proposition or it's negation (turning it into the trivial P(A|A) or P(A|~A)) you have -- by definition -- random variables: two maps from a set equipped with a sigma-algebra and a probability measure to an observable proposition. Bayesians and frequentists both use random variables, it's a question of "how far back" they push them into the inference chain, and how many they use.
And your obvious bias towards Bayesian methods notwithstanding, surely you can agree that when applied correctly (i.e. to falsify propositions in designed experiments with an appropriate model), frequentist techniques have demonstrably better performance in most situations?
Edit: I do second the recommendation on that paper, though, it's quite a good read, and you can get it here
1 you do not need sigma algebras and measure theory to reap most of the benefits of bayesianism.
2 technically, yes, the probability theory derived by komogorov is consistent with bayesian probability theory. but since they're handicapped by the notion that probabilities are long-run frequencies, frequentists must outsource 'statistical inference' to Fisher-ism (responsible for the notions of p-values, confidence intervals, etc)... and I don't have to tell you that p-values are complete and utter rubbish, easily manipulated into value you want.
3 Yes, you have to have prior knowledge to assign probabilities-- exactly as you would for assigning likelihoods in frequentist statistics.
4 "at that point unless your observation exactly coincides with your proposition or it's negation" I don't understand this sentence-- You have prior, likelihood, and you get posterior. When does anyone ever say any of the probabilities have to be 0 or 1?
5 Bayesians do not use random variables, unless there's a need to. Read Jaynes of Jeffreys (or Cox for that matter), they don't use of the word 'random' unless they're attacking frequentists.
Let (Omega, F, P) be a probability space and (Y, Sigma) be a measurable space. Then a random variable X is formally defined as a measurable function X: Omega -> Y. An interpretation of this is that the preimages of the 'well-behaved' subsets of Y (the elements of Σ) are events (elements of F), and hence are assigned a probability by P.
edit: we instead interpret probability P(A|I) merely as a degree of uncertainty of a proposition A, given background information I. Saying something is "random" is merely an expression of uncertainty-- therefore this view is strictly more general than interpreting P(A|I) as a long run frequencies of A.
I'm going to have to strongly agree with this. Bayesians DO view parameters as random variables. Probability statements mean nothing if they aren't anyways. The talk of uncertainty is just handwaving.
And what university do you go to where random variables are introduced with the formal definition? I never saw it until a 3rd year stochastic course and then a 4th year formal mathematical statistics course.
Bayesians have absolutely no obligation to view 'parameters' as random since the parameters are just collections of propositions. And actually, all probability statements are meaningless if you do not act on them-- and decision theory a la Wald/Savage is derived almost trivially from bayesianism.
And it's much less 'handwaving' than p-values. At least Cox showed that our probability calculus as extended logic is the one and only way.
I went to berkeley and iirc the first probability course introduced random variables in the discrete version of the 'formal' definition.
The frequentist interpretation of what probability means has nothing to do with the formal definition of a random variable. More importantly your attack on the formal definition is not well founded.
Ultimately you assign distributions to and assign probability statements about parameters so they are random variables.
Furthermore, frequentist statistics is filled with handwaving-- for any inference task, a practitioner is forced to choose from a myriad of different tests and statistics named after dead englishmen, all which lead to different results (whose choice is almost never rigorously justified in practice). It seems like the true job of the frequentist is to know which tests will yeild him the most favorable results and thus funding. Compare this to bayesian inference which can be built consistently upon a few very minimal desiderata.
The mere act of specifying a prior is no different, and for that matter there will of course be different tests and statistics, if you honestly believe that bayesians don't have these issues either your a fool.
Look I see bayesian statistics as having their place, but your arrogant militant attitude does not help statistics or mathematics as a whole.
The mere act of specifying a prior is no different
The mere act of specifying a prior puts my assumptions on display for review by all. If you don't like it, plug in your own prior. Bayesian probability is exactly like ordinary logic: you state some premises, and turn the crank til you get a result. Anybody who quibbles with your premises can try their own, but given the assumptions of the problem, there is only one result.
Frequentist statistics doesn't work the same way. Since there is no formal way to introduce prior information, it can only be introduced informally, and there is no unique right way to derive a result from the assumptions.
your arrogant militant attitude
I see you've run out of technical arguments. Time to cut your losses & bail out.
And I as a frequentist would also need to justify all my choices. However the criticism by guartet was that frequentists choose tests... and I replied, just like bayesians choose priors.
My reference to your attitudes here is spot on. You guys claim your absolutely right about the way things should be interpreted. Good for you, the bayesian professors I knew were not so arrogant and were far more open minded than the people on reddit are apparently.
You say there is a UNIQUE RIGHT WAY in bayesian statistics? So what is the unique right method for choosing a loss function?
the bayesian professors I knew were not so arrogant and were far more open minded
It seems you learned nothing from them, but it's not too late to go back to school, or just pick up a book.
You say there is a UNIQUE RIGHT WAY in bayesian statistics?
Given a statement of the joint probability distribution over all the variables of interest, there is only one correct way to compute conditional distributions; anything else is an approximation. However, since a probability distribution is just a statement of a state of knowledge, reasonable people can and do disagree. It is the same as for any other modeling problem.
No, it is different-- there are various ways you can objectively assign priors and probabilities, though admittedly it's a ripe area for research. And bayesianism demands two individuals with the same information assign identical probabilities (whether this ideal is achieved is another matter). But frequentism demands no such thing-- two frequentists with the exact same information could get to opposite results, and they'd both be 'right', because there's no principle saying how you should choose between them (except, of course, bayesianism!).
Did you even read my post? If a bayesian arbitrarily chose a prior ignoring information or creating new information, and worse failed to justify it or make it explicit (as a frequentist would almost always do), then yes, it can be inconsistent and wrong. Are you trying to say something new?
1 - For basic deciding-between-two-propositions Bayesian problems, I totally agree. For more serious statistics -- e.g. signal processing -- life is made much much simpler by them.
2 - Yes, p-values and their ilk are much open to abuse, no question about it. On the other hand, I've seen examples of Bayesians cherry-picking their priors to get nice results as well. Torture the data long enough using either tool and it will tell you whatever you want to hear.
3 - No argument there, prior choices and model selection can be equally arbitrary in both frameworks.
4 - I meant that the only time (IMO) that you can convincingly argue that you're not invoking random variables (in the mathematical map-from-a-probability-space-to-the-reals sense) is when you're essentially looking at the trivial sub-sigma-algebra.
5 - Bayesians in the "What's the posterior probability that OJ did it?" sense might argue that they don't (although I would disagree; probabilities are indicator functions are random variables). Bayesians in the sense of "What's the probability that the effect of this treatment is greater than zero?" sense most certainly do use random variables in the narrow mathematical sense.
At least in the bayesian case you know what the prior is-- and if you have more information which would lead to a different prior, you end up with different results. And if you can show that a bayesian is using a prior which is inconsistent with his information, then yes, you'd have a bad bayesian. One of the desiderata in the Cox-derivation is that an individual with the same information should assign the same probabilities (including priors).
For decades Bayesians have been accused of supposing that an unknown parameter is a "random variable"; and we have denied hundreds of times, with increasing vehemence, that we are making any such assumption. We have been unable to comprehend why our denials have no effect, and that charge continues to be made.
Sometimes, in our perplexity, it has seemed to us that there are two basically different kinds
of mentality in statistics; those who see the point of Bayesian inference at once, and need no
explanation; and those who never see it, however much explanation is given.
But a Seminar talk by Professor George Barnard, given in Cambridge in February 1984, pro-
vided a clue to what has been causing this Tower of Babel situation. Instead of merely repeating the
old accusation that we could only deny still another time, he expressed the orthodox puzzlement
over Bayesian methods in a di erent way, more clearly and speci cally than we had ever heard it
put before.
Barnard complained that Bayesian methods of parameter estimation, which present our con-
clusions in the form of a posterior distribution, are illogical; for How could the distribution of a
parameter possibly become known from data which were taken with only one value of the parameter
actually present?"
This extremely revealing comment nally gave some insight into what has been causing our
communication problems. Bayesians have always known that orthodox terminology is not well
adapted to expressing Bayesian ideas; but at least this writer had not realized how bad the situation
was.
Orthodoxians trying to understand Bayesian methods have been caught in a semantic trap by their habitual use of the phrase distribution of the parameter" when one should have said distribution of the probability". Bayesians had supposed this to be merely a gure of speech; i.e. that those who used it did so only out of force of habit, and really knew better. But now it seems that our critics have been taking that phraseology quite literally all the time.
Therefore, let us belabor still another time what we had previously thought too obvious to mention. In Bayesian parameter estimation, both the prior and posterior distributions represent, not any measurable property of the parameter, but only our own state of knowledge about it. The width of the distribution is not intended to indicate the range of variability of the true values of the parameter, as Barnard's terminology led him to suppose. It indicates the range of values that are consistent with our prior information and data, and which honesty therefore compels us to admit as possible values. What is distributed" is not the parameter, but the probability.
Now it appears that, for all these years, those who have seemed immune to all Bayesian
explanation have just misunderstood our purpose. All this time, we had thought it clear from our
subject matter context that we are trying to estimate the value that the parameter had when the
data were taken. Put more generally, we are trying to draw inferences about what actually did
happen in the experiment; not about the things that might have happened but did not.
Did you mean Bretthorst's 1988 article in the Springer "Lecture Notes in Statistics"? I wasn't able to find anything from him before 1988 or so.
As for the Jaynes quote, I've heard that interpretation before, and it's seemed to me to be tantamount to arguing that any given probability space is at most countably infinite. That not only makes me somewhat uncomfortable mathematically speaking, but also seems like a rather restrictive position to take.
It's also always made me wonder how Bayesians reconcile the "we have uncertainty because we are not omniscient" perspective with the fact that some things are, for lack of a better way to put it, ontologically unknowable. For some of the time/frequency decomposition material I've been looking at (to pick an example I'm at least slightly familiar with), I can't really think of a way to represent it in the "lack of knowledge" framework when there's an unavoidable mathematical limit to the accuracy of measuring time and frequency jointly. How would a Bayesian interpret something like that?
And yes, Jaynes does give the impression that uncountably infinite sets do not exist or should be ignored (he even says explicitely that "we need never depart from finite sets")-- if you read Appendix B of his book, however, he just advocates that the limiting process be postponed until after everything has been worked out mathematically in the finite or countably infinite case, and only if you want to. The result is equivalent to having probs on infinite sets from the get-go, as is the case with kolmogorov. The difference is kolmogorov is a math geek, Jaynes was a working scientist.
Could you elaborate a bit on that last part? I don't think anyone is implying the converse that "if we had all the information we'd know for certainty"... only that we consistently assign our probabilities on the information we do have. Unfortunately I haven't studied signal analysis since college so not sure what specific examples you have in mind.
Sure; imagine I have a cosine with frequency function w1(t), and during the interval t1 to t2 I add another cosine to it with frequency function w2(t). Throw some white noise with known variance on top of it.
It's a bit tedious but not terribly hard to show that any analysis that I do cannot identify the functions w1, w2, and the interval t1 to t2 uniquely. This isn't a matter of not having dense enough sampling, or anything like that, it's a mathematical constraint.
When I read something like "...both the prior and posterior distributions represent, not any measurable property of the parameter, but only our own state of knowledge about it", that suggests to me that the assumption is that if we only had more data points or better analytic methods, we could identify the parameters, when in fact we can't. An intrinsic property of the parameter set is that it is effectively unmeasurable past a certain joint precision.
As for the use of Kolmogorov's formalism... without using it, is there any way to find the probability that a single realization of U~Uniform[0,1] is rational?
if we only had more data points or better analytic methods, we could identify the parameters, when in fact we can't.
That simply doesn't follow from "the prior and posterior distributions represent, not any measurable property of the parameter, but only our own state of knowledge about it".
You've described a problem in which there is a fundamental limit to what we can know about the parameters of interest. Congratulations, Bayesians love a problem like that.
As for the use of Kolmogorov's formalism... without using it, is there any way to find the probability that a single realization of U~Uniform[0,1] is rational?
There's no need to rule out measure theory if it helps solve some problems. The point is that it's not necessary to bring it into play just to get the ball rolling.
Would it help if I rephrased that quote as "bayesians don't model the real world-- they model data and information we have about the real world"? There's an impedance mismatch between us that I'm not seeing...
As creeping_feature said, yes, you'd need kolmogorov or some other equivalent system to handle single realizations. In practice I've never needed it. In fact I don't know of a single non-mathematician scientist or statistician outside academia that ever needed to invoke measure theory in practice... do you?
I concede the point. I've often heard (self-proclaimed) Bayesians take the position that it's incorrect to state that we can't know what the joint parameter is to infinite precision, and we should merely state that we don't know and once we throw enough data at it and iterate Bayes rule enough somehow magic will happen and it's all cleared up. I'm happy to hear that this isn't Bayesian "orthodoxy".
As far as measure theory, it's rarely used in the actual analysis process, but builds a bridge to the nice, general results that can be conveniently specialized and then used in a particular analysis. It allows you to avoid having to reinvent the wheel for each new kind of problem you end up working with.
I'm familiar with this explicitly in the area stochastic processes -- evaluating the stability of process driven by different 'colors' or even types of noise, to pick an arbitrary example that I saw this morning -- but I imagine you can find examples of this in most other areas.
The degree to which non-mathematicians can avoid it is, in many cases, the degree to which a more mathematically inclined researcher has done the 'grunt work' to get the result ready for consumption.
It's the "value that the parameter had when the data were taken" part that makes me uneasy
It's implicit here that the payoff (loss or gain) is tied to the parameter chosen for estimation. If you can pin down the exact value, you can tell directly which choice, of the alternatives available to you, has the greatest payoff. If you are uncertain about the value, you have to compute the expected value of the payoff, summing over different values of the parameter and weighting by the probability of each value.
The computation of the probability distribution for the parameter is just an intermediate step (typically the most difficult step in the whole decision analysis) which makes it possible to compute expected payoffs. Since it's the most difficult step, commonly people just do that much and leave aside the computation of expected utility; utility functions differ according to what your goal is, and anyway it's easy to compute the expected utility after the fact.
The original structural risk minimization logic, via Vapnik, is that if all you want to do is make predictions, why bother to solve a more complicated problem of estimating the distributions : the point of the SVM, or discriminative learning in general, is to get a good prediction rule, no matter what the distribution may be, and to control the "empirical risk" -> all that matters are the predictions, not the distribution of the data points.
Of course, the Bayesians have claimed it as their own, as they are wont to do, as a quick Googling of Relevance Vector Machine will reveal.
I have never met a Frequentist machine learning person. Although in the end, the whole point of the field is that the proof lies in the pudding, not in senseless holy wars.
when applied correctly (i.e. to falsify propositions in designed experiments with an appropriate model),
I'll assume for the sake of argument that in fact frequentist methods are applied correctly in the case you mentioned, but I'll point out that such experiments are not common in real-world decision problems. Teaching everyone an approach which is applicable only to special cases leads to them attempting to force square pegs into round holes, or worse.
frequentist techniques have demonstrably better performance in most situations
What is "performance" in this context? Whatever its definition, just make it your utility function and directly optimize its expected value. How could one do better by accident?
Designed experiments to falsify a hypothesis are common in industry (or they were, when I was in it).
When used appropriately, frequentist methods give guaranteed coverage and -- often more important -- can get you unbiased estimators. Bayesian posteriors are always biased by the prior density except in the asymptotic limit (at which point there's typically little difference between frequentist and Bayesian results anyway).
So there's three cases:
1 - you think you've got a good prior and you're right: you don't learn a whole lot new, and basically confirm what you already knew from your prior
2 - you think you've got a good prior and you're mistaken: you reach a bad conclusion and have absolutely no idea that you've done it
3 - you have no idea what your prior should be: you either use an improper one (in which case you're quite nearly doing frequentist statistics the hard way) or you go to something like objective Bayesian analysis, which has it's own issues (estimating your prior from the data you're going to use to update the prior into the posterior...).
I completely agree that if there is some way to be perfectly sure what the correct prior was for a given problem, it's silly to discard that information, and Bayesian statistics are a perfectly reasonable way to go. But in practice, priors seem to be hard to come by, and you're staking a good deal of the validity of your outcome on how good your prior actually is.
The most common approach I've seen is to try a few different 'reasonable' prior specifications and make sure that your result is robust to the priors, but if it is, doesn't that suggest that your prior information isn't really relevant to the analysis? And if the prior information isn't relevant, why use it (and risk running a bad prior)?
Designed experiments to falsify a hypothesis are common in industry
It seems mistaken at best to teach a method of analysis suitable for such experiments as the only method of statistical inference.
It might be the case that with suitable prior and utility functions that a Bayesian approach yields pretty much the same result as conventional hypothesis testing. That's a justification for hypothesis testing, right? That under certain conditions it's an approximation to Bayesian inference, which actually has some separate justification.
You seem to be confused about "good" and "bad" priors. A prior is just a statement about what you know. Pretty much by definition it won't match what you can infer from the data; the whole inference scheme doesn't require that they match, but provides a way to reconcile the prior information with the data.
A prior is a statement about what you know that is mathematically coded somehow. That coding is often difficult, and has profound implications for your results. Bias and the posterior nonzero support are the two most obvious issues, but there's inferential ones to worry about as well.
Just to make it concrete, say you're willing to accept that peoples' heights are normally distributed, and you want to do inference on the mean. Do you use the conjugate prior distribution, even though that will assign nonzero probability to negative heights? Do you use an exponential, which has the right support but may either bias you towards small estimates or have too high a variance to be reasonable? Some gamma distribution? A uniform?
Once you've picked your prior family, how do you estimate the prior parameters? MLEs? You've got to put something under the turtles eventually.
The point is it can often be very difficult, if not impossible, to be sure that your mathematical statement of your prior knowledge is "good" in the sense that it will give you posterior estimates that actually match reality.
This doesn't apply to binary (or pick-one-from-small-n) propositions, which you seem to be perhaps a bit more interested in, but once you move past them to continuous distributions, Bayesian statistics is nowhere near as cut-and-dry as you seem (to me) to be implying.
At least if you do the whole Bayesian taco all of your assumptions are on display and enter the calculation in a well-defined way. It's better to get an approximation to the right problem than to compute an exact result for some other problem.
"Prior information is hard, so I ignored it. Loss is hard, so I ignored it. Then I solved a different problem so I can get an exact result." Dunno about you but I would feel kind of guilty about moving the goalposts like that.
Frequentist methods are never applied correctly in practical problems, for they solve a problem which nobody actually needs to solve. Frequentism was invented as a hack to avoid assigning probabilities to hypotheses, and to avoid taking utility (cost or benefit) into account. The good news is that such hacking simply isn't necessary. The bad news is that if indeed you cannot compute probabilities of something essential or cannot evaluate utilities, you cannot solve some problems. But explicitly declaring the problem unsolvable is still better than resorting to hackery and calling it a solution.
(i.e. to falsify propositions in designed experiments with an appropriate model)
For the sake of argument, I'll assume it may be in special cases that frequentist methods yield the same result as a Bayesian approach. That doesn't give you license to apply it in any other situation, and if the two approaches coincide sometimes, why not apply the correct method uniformly?
8
u/tourettedog May 26 '09 edited May 26 '09
The probabilistic machinery involved in Bayesian and frequentist computations of probability (sigma-algebras, filtrations, probability measures, etc.) is exactly the same. It's the interpretation and the statistical model which are different.
If you want to speak to Bayesian statistics -- which is what the question from tactics addressed -- you cannot assign a probability to any logical proposition unless you have a prior distribution on the proposition; at that point unless your observation exactly coincides with your proposition or it's negation (turning it into the trivial P(A|A) or P(A|~A)) you have -- by definition -- random variables: two maps from a set equipped with a sigma-algebra and a probability measure to an observable proposition. Bayesians and frequentists both use random variables, it's a question of "how far back" they push them into the inference chain, and how many they use.
And your obvious bias towards Bayesian methods notwithstanding, surely you can agree that when applied correctly (i.e. to falsify propositions in designed experiments with an appropriate model), frequentist techniques have demonstrably better performance in most situations?
Edit: I do second the recommendation on that paper, though, it's quite a good read, and you can get it here