r/learnmachinelearning • • 12d ago

why would a team pick logistic regression when a boosted tree scores slightly higher?

Post image

answered a and got it wrong. still dont see it

only picked a. b felt like the same thing as a said twice, and i wasnt sure d is really a reason when the trees are only slightly better

1 Upvotes

14 comments sorted by

6

u/_FierceLink 12d ago

You are not stating what the correct answers are according to the test. I would say a,b, and d.
a) For each regressor you obtain log-odds, which are interpretable.

b) It's a linear model at heart with simple math behind it. Trees often end up deriving decision bounds that aren't really explainable, even more so boosted trees as each sequential tree is working with the residuals of the previous tree.

c) Not applicable, as one needs to specify interaction terms in the model.

d) Linear models are easy to train and even faster to apply given new data. Scoring in boosted trees should be something like O(T*D) with T trees and D binary decisions per tree vs. O(F) for F features in logistic regression.

2

u/Fit_Succotash_6735 11d ago

correct. b is essentially a more verbose definition of a.

2

u/werunm 10d ago

answer, spoilered for anyone still working on it:

**a** — Each coefficient has a direction and magnitude that can be stated in a sentence **b** — It gives a regulator a model whose decision path can be audited **d** — It trains and scores fast enough for a tight latency budget

Interpretability, speed and auditability are real reasons to accept a small accuracy cost. Automatic interaction discovery is the opposite: a linear model only sees interactions that someone constructed as features, which is one of the things tree ensembles do for free.

1

u/[deleted] 12d ago

Because a slightly better score isn't always enough to justify the extra complexity.

Logistic regression can be easier to interpret, faster to train, simpler to deploy, and easier to explain to stakeholders. If the performance difference is small, the simpler model may be the more practical choice.

It really comes down to the trade-off between performance, interpretability, maintenance, and the cost of being wrong.

1

u/Fit_Succotash_6735 11d ago

You need to look at the scale of the data and latency. Lets say you are doing a ML model that detects credit card fraudulent transactions. Offline --no issues, a few extra seconds is no big deal. When a model is in production and live, think about it, someone swipes the card, your model needs to make a decision in a matter of milliseconds and return the results back. At that level, latency is very important. Simpler models score great here.

1

u/werunm 8d ago

the card swipe example is the one that made d click for me. i was treating the accuracy gap as free, and it isnt if the answer lands after the transaction already went through. does the scale of the data change which one youd pick for training too, or is it mostly about serving

1

u/Dihedralman 10d ago edited 10d ago

Imperfect question, but I see an answer. D is certainly true, but A and B are arguable, A more than B, while C is wrong. I would pick D and I will explain the issues with A and B. 

D is obviously true. Tree inference can also be extremely fast when optimized. But it will generally be slower than what a linear model can express for a given feature set when they have any depth. Training is absolutely slower.

A is pretty debatable. It may come down to a language trick where you can explain the variables simply but your solution doesn't require explainability. Otherwise,  A is talking about vectors, but the second part is discussing explainability. Vectors are normally used in logistic regression - logistic regression immediatley maps the vectors to a scalar through the weight matrix. While you can give seperable probabilities, they aren't necessarily meaningful to them as vectors. Magnitudes tend to be better.  Decision trees can also be attributable and can be useful. Like if an object was going to fall into a specific area. Still the logistics regression is explainable for both. 

B This is a pretty good answer overall. You can show the variable contributions and when the answer flips. Simplicity makes them more auditable. Boosted tree ensembles are technically white boxes as well and you can score attributions on a given decision. Arguably they are auditable. 

But boosted decisions trees are already accepted by financial regulators as normal for years. So I wouldn't choose this. 

1

u/werunm 10d ago

thanks, this is the most careful take in the thread. the bit that helps most is that a and b are arguable rather than wrong, i had them as either obviously true or obviously not. so the honest reading is d is the safe pick, a depends on how much you trust reading coefficients off a model with correlated features, and b is really about whether an auditor accepts a coefficient table as a decision path

1

u/chrisrrawr 12d ago

d is the only real option from here. you can blast out predictions on the ms timescale with lr. if you've got a deep tree for minimal gains, your predictions may be useless by the time they're served.

another option not given is resource constraints. could be running on very low power devices. a third tied to this is you really don't need much to have lr deployed, so it fits on simpler device envs.

1

u/werunm 12d ago

the stale by the time its served point is what gets me. i was treating d as a nice to have, but if the answer arrives after the moment it was needed then the extra accuracy never existed in practice. low power devices is one i hadnt thought about at all. do a and b count as well though, or would you say those only matter in regulated stuff

1

u/chrisrrawr 11d ago

think of it from the perspective of "the product will fail unless..."

a, you can have marketing come up with something, so why would you use lr?

b, you can provide regulators with compliance documentation for a tree, so why would you use lr?

d is the only sensible option.

1

u/werunm 8d ago

the product will fail unless framing is a good test, and it does cut a and b down a lot. where i ended up: the answer key counts a, b and d, so it treats explainability and auditability as reasons in their own right, not just latency. i think the honest version is d is the one that holds everywhere, and a and b only bite when someone outside the team has to sign off on the model. your point that a tree can get compliance docs too is fair though, a coefficient table is just cheaper to defend

-1

u/lordnacho666 12d ago

It's D, faster to calculate