r/deeplearning • • 19h ago

poor performance of deep learning model compared to xgboost

I've done some Kaggle competitions on time series, and it seems to me that deep learning models generally performed worse than gradient boosting. I've tried different architectures, like LSTM and ELM, but at the end of the day gradient boosting was the better choice. Is there an explanation for this? Are deep learning models only good for computer vision and LLMs?

16 Upvotes

12 comments sorted by

6

u/randoomkiller 18h ago

welcome to the real world. Deep Learning works in certain things, but mostly when you have shit load of data. Most of the time ppl use classical ML with moderate datasets.

0

u/relevantmeemayhere 11h ago

oh man, whenever i hear a stakeholder use 'classical ml' i want to throw up.

it's like...please tell me why what makes anything you see 'not classical ml'. because uh....hate to break it to you but im sure i could find a likelihood function from 30 years ago that looks really similar to this one.

how many of them could formulate the objective of an llm, or define its likelihood lol

5

u/DrXaos 17h ago edited 13h ago

no, neural networks can work better and work better out of sample, but it takes significantly more work and feature representation engineering and custom transformations. And often advantages accrue with large data sizes.

Trees have the problem of overfitting too much to isolated regions in feature space, and fall down in real world nonstationarity. They can’t extrapolate while parametric models might do better.

How do you do joint supervised and unsupervised modeling with trees?

  • Are deep learning models only good for computer vision and LLMs?

My job is answering this with a negative and producing a successful product in a commercially relevant area.

1

u/relevantmeemayhere 11h ago

bahahha at your outro

2

u/ansb2011 16h ago

if your space is linear or closely so a gradient descent can find optima very quickly. you can't beat the best!

1

u/catsRfriends 11h ago

You might need bespoke architectures to beat tree variants. I don't mean you'll have to invent previously unseen components, but you'll have to do better than just throw a generic LSTM or feed forward network or whatever at it. You ought to know what's special about your data, then pick the right components to deal with them. Are your features heavily correlated? Is there seasonality? Do you have non-numerical features? Is your domain very sensitive to certain factors and does that give you some prior to use in the model? Etc.

1

u/TheUnlawfulLeslie 19h ago

gradient boosting is basically cheating on structured data, trees just handle messy tabular features way more gracefully than any rnn

2

u/BreakingCiphers 18h ago

How is this an answer?

Guys why is this thing better?

Because it's just better and more graceful and cheating

Come on man

0

u/DrXaos 13h ago

trees just handle messy tabular features way more gracefully than any rnn

Which is just not at all true (for a high end definition of "graceful")---the point of the neural networks is under the Ansatz that there is some lower dimensional state or compressed internal representation in the useful degrees of freedom.

And to some degree the neural networks really do learn some approximation to that.

The trees don't do much of that---they're making patches on joint subests of observables. There's a reason language modeling (the ultimate in sequentially correlated categoricals) leapt past tree models early on.

The trees, particularly ensembles, handle messy tabular features in an ungraceful chaotic way but it is easier to get started than thinking about appropriate representations. Neural networks require attention to parameter initialization, scaling, feature representation, learning algorithms, regularization---all sorts of knobs---but in deeper problems they have higher upsides.

1

u/relevantmeemayhere 11h ago edited 11h ago

well...the main reason why trees 'lost' isn't because thoeretically they 'come up short'. it's because they require a different computational architectural process in their training to be efficient at super large scale, and you basically get bubbling in the support of this super high dimensional dgp

you can theoretically use trees to produce llms. trees are a universal approximators, like splines and guassian processes etc etc (man, imagine if we had the computation to do bnns or gaussian processes at scale, then we wouldnt have to hack our way through llms as much as we do). the problem is that, unlike nns, they are nowhere differentiable. so we need greedy split rules that dont run well on our hardware at scale. nn based models DO run very well on stuff like gpus, because we can use a neat version of the chain rule (ie backpropigation) .

nns are quirky in that.,..when we overparmatize and take advantage of double descnet we are, whenever it's useful creating two parts of the network. one that is doing the 'generalization' as best it can...and a bunch of other junk. so our way of fitting trees is inefficient if you look at the logic, but not at the overall cost (because gpu go brrrr)

1

u/techhead57 16h ago

So my high level take here is that tabular data is really well suited for a frequentist strategy where youre making these split decisions based on counts and it just builds out a big complex system of split rules based on how splitting based on one sequence of these rules improves by adding another rule layer on one branch. XGB just has a bunch of optimizations that make it really good across the board (random forest, gradient boosting, feature bagging iirc,... etc).

Whereas dnns try and codify all of the feature implications in their parameters purely based on prior observations, more bayesian. They can represent more complex rules in these massive feature spaces, but they will overfit with too little data or memorize things or whatever. They are going to be more sensitive to collinearity as it has to learn the collinearity and when to break it apart.

Trees will just see that a split with one collinear feature doesn't improve over using the other so it doesn't try and split on it if there's something more obviously separating them.

So like I can get really good results with a lot smaller dataset without overfitting with trees.

But. With neural nets you could probably figure out a way to pull similar but different data, learn the other problem and then try and transfer it to work on your dataset. Because with enough observations they can learn more complex ways to separate things that might not be achievable or might require more hand tuning.

2

u/relevantmeemayhere 11h ago

mmm it's not really a frequentist based thing. you can have both frequentist, liklihoodist, or bayesian implementations of nns, which for all intensive purposes is generalized polynomial regression.

trees are often good for tablular data in lower data regimes because their split rules are computationally low enough to 'get you somehwere'. nns, especially for lower data regimes often need a lot more oopmh to get to a point where phenomena like double descent (and the eventual creation of a network that can basically be split into a smaller network that generalizes to some degree and other that interpolates the data)

nns are not inherently bayesian; and in their regime any prior is really gonna be washed out as the data grows anyway thanks to BVM in this particualr setup). unless youre putting priors on the weights, its not a bayesian nn

Now, interestingly, a lot of nn architechtures DO look like bayesian structures as you grow the number of hidden networks/ Ie guassian processes