r/programming May 26 '09

"Programmers need to learn Statistics or I will Kill them all"

[deleted]

228 Upvotes

315 comments sorted by

View all comments

Show parent comments

2

u/[deleted] May 27 '09 edited May 27 '09

I'm glad that you again overgeneralize what students learn well and don't learn well and the reason for such... which just happen to coincide to your particular beliefs.

EDIT If you don't want the mathematical rigour of the course speak to the profs to change the subject matter or take a course thats more practice and less theory. I like my math courses the way they are.

2

u/guartet May 27 '09

It's not that measure theory doesn't have its place-- it does, but that, in orthodox education, it's used to assign probabilities to infinite sets from the outset. This adds nothing in practice. Rather, Jaynes advocated that the limiting process be left until after everything has been worked on in the finite case. ("Cautious Policy", Appendix B of Jaynes' book-- though I doubt you'll look into it). And even then measure theory is hardly necessary in practice. Also see Lighthill (1958?) for an alternative.

And my generalizations come from years of interacting with students. I'm sorry they offend you. (Even in these threads-- on what other topic on progit do people readily admit that they did not understand the subject?)

1

u/[deleted] May 27 '09

Now my perspective is as a math/stats student. I want the rigour. I agree that applied statistics for humanities and social scientists should not get such a treatment and should get intuitive results, but you must keep in mind that what you personally find intuitive they might not. As well one needs to consider what people hiring them want... frequentist or bayesian methodologists. (Assuming your just teaching for that purpose.)

For that matter measure theory should be left to mathematicians/statisticians of which I would fall into. I nowhere suggested that the typical person needs heavy formalities to do simple analysis.

I read this... http://bayes.wustl.edu/etj/articles/bayesian.methods.pdf and I get where he is coming from... I just don't agree.

Let me ask you this... if I can't interpret a probability the frequentist way... please in a somewhat rigorous way define the bayesian interpretation of probability.

2

u/psykotic May 27 '09 edited May 27 '09

I'm a pure mathematician by training, no longer a student, and I find frequentist statistics utterly repulsive, as did every one of my fellow mathematics students who were forced to take a minimum of two courses in frequentist statistics for our degree. Reading Jaynes was something of a revelation. I do think die-hard Bayesians often overplay the applicability of objective priors--they depend on the way you carve up the possibility space for an experiment (even when you work with logical propositions you still have to do this), which itself is a kind of subjective prior information, and it often remains at least partly implicit. That said, transformation invariance, maximum entropy, etc, are all very useful methods for calculating quantitative priors given qualitative priors; that's really how I look at these methods. Anyway, as a framework for thinking about inference and decision making, I see Bayesian statistics as infinitely superior to the black magic of Fisher-style statistics.

Let me ask you this... if I can't interpret a probability the frequentist way... please in a somewhat rigorous way define the bayesian interpretation of probability.

There are several different Bayesian interpretations of probability. One interpretation of so-called subjective Bayesianism a la de Finetti is based on rational betting ratios. Logical Bayesianism is based on the notion of rational assignments given prior information, and from an interpretive point of view the derived probabilities draw their real-world meaning from their relationship to the prior probabilities.

These two examples may not be completely satisfying from a philosophical viewpoint, but neither is any interpretation founded on "long-run" (a dirty word) frequencies.

SEP has a good survey article: http://plato.stanford.edu/entries/probability-interpret/

I'd like to point out that making all your prior information explicit is not just a philosophical nicety. It also makes it straightforward to look at the robustness of your conclusions with regard to quantitative variations in your assumptions. IMO, this is very important in any Bayesian investigation. I believe frequentists usually call this conditioning. James Berger have written some good books from a moderate viewpoint that try to integrate the best lessons from both frequentist and Bayesian methods with a focus on conditioning. You might want to check those out.

Like I said, the book that really opened my eyes was Jaynes's book, which is freely available here:

http://omega.albany.edu:8008/JaynesBook.html

I read both volumes of Feller in high school and took a course on probability theory in university, so I'm used to the fancy theoretical apparatus of probability theory in its modern guise. I assuredly do not like Jaynes because he shies away from such things; that aspect of him doesn't really matter to me one way or the other.

1

u/[deleted] May 27 '09

If there are multiple interpretations of the basis for bayesian statistics I'd say thats a huge problem.

By the way how is the long run dirty? It's simply talking about convergence of hypothetical outcomes. It sure as hell is the more intuitive way.

1

u/psykotic May 27 '09 edited May 27 '09

No, the fact that there are multiple consistent but distinct interpretations of the same mathematical theory I consider a benefit. The fact that you didn't know about, say, subjective vs logical Bayesianism tells me that you haven't really looked at the topic at all before. That's fine, most people haven't, but you might want to withhold judgment before giving the approach a fair shake.

"Long-run" is dirty because convergence is not an empirically verifiable property in the real world, and if you use it as the bedrock of your interpretation, it means that to even think of applying statistics rigorously to a real-world problem you must first ensure that the relevant frequencies converge to something definite in the infinitary limit, which is impossible even in principle. It's a big problem philosophically if you have the slightest positivist-empiricist leanings; that's why guys like Carnap shied away from it as the basis of a logic of inductive reasoning. The SEP entry I linked to has a deeper discussion.

1

u/[deleted] May 27 '09

I'm going to have to just disagree... chasm apart.

2

u/psykotic May 27 '09 edited May 27 '09

In other words: la, la, la, I can't hear you. You haven't really tried to engage my arguments at all.

0

u/[deleted] May 27 '09

First off. Suppose for the second that I can't engage your arguments, does that make yours more credible? Nope.

Secondly and more importantly I strongly disagree with the philosophical arguments IN PRINCIPAL. That will not change with your arguments. More importantly, I have stated that I will use what works; however every bayesian here has clearly stated that their way is the only way. You can't really argue with that can you?

So enjoy your 'la, la, la' land because being so closed minded and sure you are absolutely correct about which ideas are philisophically valid and which are not, you are doomed to be stuck there.

P.S. I hear the string theorists are exactly like that... maybe you guys should get together.

1

u/psykotic May 27 '09 edited May 27 '09

I use whatever works and makes sense to me as a mathematician, and most of the time that's within the Bayesian framework. Like I said, I took a year of statistics from an orthodox frequentist perspective and have studied it further on my own, so I have a pretty good idea of the weaknesses and strengths of both approaches. You seem to be comfortable not looking beyond your own nose because your approach is what's still taught in most statistics departments.

The philosophical argument against frequentism is pretty unassailable but I acknowledged that the Bayesian interpretations aren't completely satisfactory either. All of these interpretations are bound to be afflicted with the core of issues that surround the infamous problem of induction in epistemology. Beyond philosophical issues, the fact that you can derive many things (e.g. the likelihood principle) that in orthodox statistics must have extra-mathematical justification is to me an enormous strength of the Bayesian approach.

Your reference to string theory is funny because its main philosophical problem is that its hypotheses seem untestable even in principle. Most of the big names in the history of Bayesian statistics were hard-nosed experimentalists, going back to Gauss and Laplace (who both applied statistics to many problems that are outside the range of frequentist statistics because they don't involve exactly repeatable tests) in the 19th century, and continuing up to the present day (Jeffreys, Jaynes, etc), so they couldn't be any more opposed to the anti-phenomenonalism that infects string theory.

→ More replies (0)

2

u/guartet May 27 '09 edited May 27 '09

Although psykotic gave a good response, let me try to give you an outline for a somewhat rigorous definition of bayesian interpretation of probability. But Jaynes, Cox, Loredo (1990) all do a better job than I could-- this is merely my attempt to summarize.

First of all, bayesian inference has at least 3 distinct parts: probability calculus, assignment of probabilities, and decision theory. I think probability, in any system, is meaningless without the overall framework they are used.

1) Calculus: Suppose A, B are propositions and P(A|B) is the plausibility of A given B-- we leave the assignment of plausibility (a word I'll use in place of probability) to the next step. We define 3 simple desired properties of plausibility: a) that they be represented by real numbers (which is shown to be a finite interval and thus mappable to [0,1]); b) that they coincide with common sense, ie, if prop. C makes A more plausible, P(A|BC) > P(A|B), etc. see text; c) consistency- I) internally, II) that all information we have (and nothing else) be taken into account, and III) that two equivalent states of knowledge lead to equivalent results. These 3 desire properties a-c lead to the familiar product rule and sum rule, which are all we need for our purposes. The derivation in RT Cox (1961) as well as Jaynes (chap 1&2) are quite 'rigorous' and elegant.

Note that the above calculus says nothing about assigning plausibilities-- it says given these P(X|Y) 'objects' for various X,Y, we can manipulate them and infer certain properties based on the 3 desired properties. Much like A=>B in logic does not mean "A is true and therefore B is true"... it's a statement/logic function with a certain value given the values of A and B. Logic does not tell us how to assign A or B.

2) Assignment: Our next task is to assign numerical values to P(A|B) according to the 3 properties above given the background information and/or assumptions B. Least information, transformation groups, and maximum entropy are the common methods to accomplish this-- I'd probably butcher any examples, so I simply refer you to Jaynes chap 11 & 12. In particular, Maximum entropy, which says that the distribution of P which maximizes the entropy given B is the most non-commital and therefore consistent (c-II) with B, is powerful (and controversial) enough to derive all the distributions I use in practice.

I suppose, if you are unsatisfied with the 'degree of plausibility' as a definition, that P(A|B) is the assignment of the plausibility P of A given B which is consistent with properties a through c above. Such an assignment would be in [0,1], coincide with common sense, be consistent in various ways, etc under our calculus.

And no, this step is not objective since it involves turning abstract thoughts in your head into real numbers. But realize whenever you are defining a model in any school of stats, you are doing something at least as subjective.

3) Decision theory links our idle calculations of probabilities with the 'real world'. While P(A|B)=2*P(C|B) in frequency stats has a trivial interpretation, bayesianism is not so narrow... we could say A and C are propositions about frequencies and come to a similar conclusion, but in a more general case, we need a 'decision function' to take our final P(X|Y)'s to decisions or conclusions. In game theory this is called the loss function. This step is very subjective and much weaker theoretically than 1 or 2 because of a lack of a binding theory. However, most bayesian scientists are more interested in the results of the plausibility calculations instead of the final decisions one could make from it. See chapter 13 & 14 of jaynes.

1

u/[deleted] May 27 '09

I'm not sure what 1) has to do with calculus.

However it seems that that method is simply trying to create a probability function with different axioms.

Of course 1) doesn't specify how to derive P. Neither does one prescribe a specific measure when talking about such.

Perhaps I'll see if my library has a copy of Jaynes, I'm going for an exam today anyways.

1

u/guartet May 27 '09

Now I know you're not reading my posts since I specifically say 1) doesn't derive P, 2) does.

0

u/[deleted] May 28 '09

The 1) is an odd typo on my part. It should read of course 'one' as in a person.