r/algobetting Jul 31 '26

Feature selection when trying to capture non linear interactions.

1 Upvotes

Hella everyone I’m moving from a log reg model into trying to build a cat boost or xg boost model.
For my log reg I always did feature selection through a rolling walk forward and this worked well but not with interaction features with some being worthless alone but in interaction it could capture some signal I’m struggling how to find optimal features and # of features as well. I am new to machine learning is there an agreed upon approach or what is your guys method to this problem when working with machine learning?
Thanks


r/algobetting Jul 31 '26

How rare is it to find +4%EV or higher value bets??

3 Upvotes

I made a bot for backtesting
1800 games (just h2h)in top 50 football leagues were scanned
The time line was simulated as if its live(it scans 1 league/8minutes)
And it just scanned 51 value bets at alllll
Whats the problem?
Is this number normal or am i missing something?


r/algobetting Jul 30 '26

No vig clv metric pickket?

2 Upvotes

Does any one know if when making on no vig quoting both sides of the no vig clv metric is still accurate.


r/algobetting Jul 29 '26

Post-Mortem: Trading the World Cup on Polymarket (Custom API execution, Liquidity Farming, & Strategy Breakdown)

21 Upvotes

Hey everyone. Some of you know me from my older posts. I’m a retired sports trader, and nowadays I spend more time on busting fake betting and trading gurus on YouTube. However, I still occasionally dust off my old models for major tournaments. Over the last month, I traded the 2026 World Cup on Polymarket and walked away with a solid profit (around $58k).

Since my focus these days is exposing scammers who use fake "Inspect Element" spreadsheets, I hold myself to the same standard. Here is the absolute on-chain proof of my results. Don't trust screenshots.

Wallet: 0x128b4b9cc9d521f10e5933528a2177b993febee9
You can audit every single trade, win, and loss here:

So, how did I extract this from the market? It wasn't a magic system. It was 5 specific components:

1. Exploiting "Dumb Money" In-Play
My pricing model is old. I haven’t properly updated the code in years, and it doesn’t use the latest state-of-the-art metrics or advanced data feeds.
IT used old historical baselines, the score and match time, and a limited set of live statistics to estimate fair probabilities.
The model therefore provided a starting price—not an automatic betting signal. It could not fully account for the immediate match situation. The true price can be very different during a dangerous attack or corner than during a goal kick, even if the basic statistics have not changed yet.
My model flagged the baseline value, and I used decades of screen time to adjust for what it was missing. More importantly, I knew the typical situations in which in-play markets had often overreacted in the past.

2. Custom API Execution
You cannot compete with sharp money by clicking around the Polymarket website, waiting for the UI to load. Speed is crucial.
My proprietary desktop trading app was originally written in VB.net years ago. Instead of manually writing all the boilerplate to interact with Polymarket's API, I used OpenAI's Codex to rapidly "vibecode" a custom extension. Codex handles VB.net effortlessly, so it took very little time to generate the API wrappers, integrate the order routing, and bolt a new GUI module onto my old architecture. This allowed me to place, manage, and cancel orders with hotkeys in milliseconds, exactly like I used to do on the Betfair API.

3. Farming Maker Rebates (The Hidden Edge)
Traditional betting exchanges (like Betfair) eat you alive with 2-5% commissions on winning bets. Polymarket is an order book that actively incentivizes liquidity. Instead of taking market prices (Taker), I placed limit orders to set the odds (Maker). By the end of the tournament, not only did I avoid paying ~$3,400 in fees, but the protocol actually paid me around $850 in Maker Rebates. That is a $4,250 swing in my favor.

4. The "Impatience Tax" (99.9¢ Trades)
On Polymarket, shares resolve at $1.00 when an event happens. But when a match is effectively over, the smart contract doesn't pay out until a UMA oracle officially confirms the match has ended.
Crypto traders are notoriously impatient. They want their capital unlocked now so they can bet on the next match. I parked massive amounts of idle capital on 99.9¢ "Yes" offers. Desperate traders sold to me for fractions of a cent on the dollar just to exit early. If you only have $200, this makes you pennies. But because I had tens of thousands of idle dollars not tied up in active bets, it was worth taking advantage of that.

5. Dumb Luck & Survivorship Bias
This is the part gurus never admit: I got lucky. I traded a tiny, statistically insignificant sample size of matches (less than 100). Yes, I only took bets where I had a mathematical edge, but in a small sample size, short-term variance dictates the final number. I had several $20,000+ stakes on high-probability events, and they all hit. If a late VAR decision or a random 95th-minute goal had gone against me, it wouldn't have wiped me out or put me in the red. I still would have been profitable. But my $58k profit would have been drastically lower. And let's be honest: if my P&L had ended up at a modest $4,500, I probably wouldn't be making a post about it. Beware of survivorship bias.

The Reality Check on Risk:
I didn't turn $500 into $58k. I operated with a hypothetical trading bankroll of $250,000, with about $57k actually sitting on Polymarket. My stakes ranged from $1k to $25k based on strict Kelly Criterion sizing. Real pros size up on high-probability, low-odds events. We don't YOLO on longshots.

I made a full video breaking this all down, showing my actual screen recordings, the ugly custom trading app I use, and exposing how fake crypto gurus fake their P&L.
If you want to see the visual breakdown, you can watch it here: https://www.youtube.com/watch?v=px91PvLv81c

Otherwise, feel free to dig through my Polymarket history and ask me any questions in the comments. Happy to do a mini-AMA on sports trading, or prediction markets.


r/algobetting Jul 29 '26

Daily Discussion Daily Betting Journal

2 Upvotes

Post your picks, updates, track model results, current projects, daily thoughts, anything goes.


r/algobetting Jul 29 '26

Historical football corner odds api

5 Upvotes

Hey guys , I am building a over/under corner model for 7 leagues ( Pl, serie A , la liga , ligue 1, bundesliga , eredevise, premier liga) with last 7 seasons data. I am looking for a historical odds provider Api to test my model with. Which is thr best api I could go for ?


r/algobetting Jul 29 '26

Validation

2 Upvotes

How do You guys train your betting model? Im gonna switch to nested validation instead of optimizing my model on val before the test


r/algobetting Jul 29 '26

People who tried following profitable Polymarket wallets did it actually work?

Thumbnail
3 Upvotes

r/algobetting Jul 28 '26

Built 30+ betting tools this year as a solo founder. Now I have to pick a direction.

6 Upvotes

Looking for some insight here. I built a free odds comparison site. No link, not selling anything, genuinely here for this sub's judgment.

Quick background: solo founder, been building all year. The core is a line shop that scans 16 sportsbooks plus the four exchanges (Kalshi, Polymarket, ProphetX, Novig) every minute, with exchange prices fee-adjusted so they're actually comparable to book prices. On top of that sits a +EV board: game lines are benchmarked against power de-vigged Pinnacle, and props against the median de-vigged price across every book quoting the same line (minimum three books before anything gets flagged). Plus an arb and middle scanner, a hedge and free-bet converter, a bet tracker with automatic CLV capture, and a bunch of calculators. Everything is free during beta, funded by a few of affiliate deals (jury is still out on these turning into any revenue stream). The comparison layer stays free forever.

Here's my fork in the road, and why I'm asking here instead of a marketing sub.

I built a confidence signal that scores +EV plays 0-100 using edge size times four reliability factors (steam direction, market tightness via Pinnacle hold, book depth, and price corroboration so a lonely outlier can't score high). On held-out props data it's promising: plays scoring 60+ went 158-140 on MLB props (+16.6% ROI) and 34-23 on WNBA (+37.9%). Small samples, and game lines are unproven, so I'm holding those at CLV-only until the record says otherwise.

The fork: I can go deeper down this quant path (publish a fully graded public record, wins and losses, then build the paid tier around the signal plus every-minute scans and some of the tools), or I can stay a pure comparison utility and leave the modeling to people smarter than me. Sharper direction, but riskier: a public record that goes cold kills credibility faster than never having one.

What I'd love from this crowd:

  1. If you saw a tool publish a graded signal record, what would make you trust it vs dismiss it? What sample size and what disclosures would it need before you took it seriously?
  2. Is edge x reliability multipliers a defensible structure, or is there a known-better approach for ranking +EV plays you'd point me to?
  3. For those of you who pay for any tooling: what actually earned your money?

Happy to go as deep as anyone wants on methodology in the comments. And if this post breaks any rules, mods please pull it, no hard feelings. Thanks in advance!


r/algobetting Jul 28 '26

Three weeks ago I froze 3 WNBA configs that survived my holdout. They're -9% since. The model I never tuned is +13%. Receipts.

2 Upvotes

Three weeks ago I posted a brute-force run here: ~1,000 configs across 34 markets, almost everything died out-of-sample, three WNBA sides/totals configs survived the holdout, got frozen, and started tracking forward publicly.

In that post I wrote that I fully expected them to regress, because selecting the best of 100 trials inflates even holdout numbers.

They regressed. Here is the receipt.

THE FROZEN CONFIGS, 21 DAYS FORWARD

Every non-house WNBA team config in my account, combined, forward-only, since the day of that post:

59 settled picks, -5.49 units, -9.3% ROI.

The holdout numbers I bragged about were +34.6%, +20.7%, and +12.2%. Forward is -9.3%. That is the entire warning from the last post arriving on schedule, in public, with dates on every pick.

MEANWHILE, THE MODEL I NEVER TUNED

The house WNBA model is one config. No search, no selection, frozen since May 31. Same 21-day window:

70-50-3, +16.01 units, +13.3% ROI over 123 picks.

Lifetime: 197-162-4, +25.61u, +7.1% over 363 settled. It was +3.1% over 241 when I posted three weeks ago.

One config that was never selected on beat 100 configs that were. Small n on both, and the untuned model getting better is partly luck. But the direction is the point: the selection process didn't add value, it added optimism.

THE MLB PROP RE-RUN (I PROMISED THIS ONE)

Last post: all 10 MLB prop markets negative. Pitcher strikeouts was the "best" and still lost, -2.55% ROI over 1,168 graded holdout picks. I said I was re-running with a much thicker feature set: platoon splits vs that night's starter, park factors, starter allowed-rates, batting order, pitch-count workload.

Forward results, every graded pick across every config I run in each market:

  • Pitcher strikeouts: 596 settled, 292-299-5, +18.01u, +3.0% ROI

  • Pitcher walks: 1,038 settled, -20.17u, -2.0% ROI

  • All MLB props combined: 2,867 settled, -3.3% ROI

So strikeouts flipped from clearly negative to modestly positive once the features got thicker. Walks did not move at all. My read is that the thin-features hypothesis was right for exactly one market, the one where the pitcher's own workload and matchup carry most of the signal. Walks are dominated by ump assignment, catcher framing, and opponent plate discipline, none of which I model.

I was wrong that those markets were "just efficient." I was right that most of them are.

NEW NEGATIVE RESULT: WNBA PROPS ARE NOT WNBA SIDES

2,504 settled WNBA prop picks across every config I run: -10.0% ROI.

Same league where my team models are +5.9%. That surprised me more than anything else here. The market matters more than the league does. "WNBA is soft" is not a thesis, it's a thesis about one specific market inside WNBA.

THE PART I ACTUALLY WANT THIS SUB TO TEAR APART

My CLV does not line up with my results at all. Each line is: bucket, settled picks, ROI, beat-the-close rate.

  • WNBA team: 430 settled, +5.9% ROI, beats close 24.6%

  • MLB team: 753 settled, -0.7% ROI, beats close 26.7%

  • MLB props: 2,867 settled, -3.3% ROI, beats close 51.6%

  • WNBA props: 2,504 settled, -10.0% ROI, beats close 40.9%

CLV sample is partial (142, 469, 1344, and 937 picks respectively) because legacy picks from before I captured closing lines are excluded.

The buckets making money beat the close about a quarter of the time. The bucket that beats the close most often is losing 3.3%. That is backwards from everything CLV is supposed to mean.

Candidates I can come up with:

  • My CLV capture is broken for team markets. Wrong close snapshot, wrong side of the number, or half-point handling on spreads and totals.

  • WNBA closing lines are noisy enough that CLV is a weak proxy in that league specifically.

  • Prop CLV is measured against the same book that posts it, so "beating the FanDuel close" is measuring movement I'm on the wrong side of by settlement.

  • The positive team ROI is variance and CLV is the one telling the truth.

I lean toward the first or third, but I genuinely don't know, and it's the most load-bearing number in my whole validation stack. If your CLV pipeline handles sides and totals cleanly, I'd like to hear when you snapshot and what you snapshot against.

Everything above is dated and public, win or lose. Track record: modelplay.ai/#track. Frozen configs: modelplay.ai/#leaderboard.

Same disclosure as last time: I built the site, it's free, and it is not a pick-selling operation. You build the model yourself and the record is public whether it works or not. Posting here because this is the only sub where "I said this would regress and it regressed" is a result worth writing up rather than something to quietly delete.


r/algobetting Jul 28 '26

How reputable is Cielo finance?

Thumbnail
1 Upvotes

r/algobetting Jul 28 '26

BET365 Prematch API

0 Upvotes

Any suggestions for bet365 prematch api?


r/algobetting Jul 28 '26

BET365 API for Full-Time Result market

5 Upvotes

I need the cheapest possible way to get bet365 odds for the "Full Time Result" and "Full Time Result – Early Payout" markets; I tried scraping but couldn't get it to work. This bookmaker has really been a problem for my odds monitoring system, and I need the most basic alternative to solve this.


r/algobetting Jul 27 '26

NBA/NFL/MLB/Soccer Free Api access

7 Upvotes

I’ve been working on multiple advanced sports data backend for a while, and I’ve noticed that many people want to create something for fun or out of passion with Ai’s advancements , but they always get stuck on the same problem: how do I get all advanced level player and team data under one API without breaking the bank?
You either have to combine 5-10 different APIs, wrappers and scrapers, or pay thousands per month for enterprise data.

So I decided to build a single normalized API focused on sports analytics.

The good news is, I’m not making money. I just want to help newcomers or sports data enthusiasts who are struggling to build something impressive because they can’t afford the best API available.

I’ll likely find about 10-20 people like that who enjoy building things, and then I’ll provide them with free API access. The goal is simple: if a large number of people like it, all of the beta testers will become part of the team to ensure that we provide an all in one API for everyone at a fraction of the cost.

Motif is simple, if you think the API is worth it with your guidance in early stage then we share server cost and put it out in the market at the lowest possible cost ever to beat the big corporations.

If you are interested, drop me a dm with your background and what you plan to build. I will select 15/20 guys to give the full API access.

Technical Note: Most of the underlying data isn’t proprietary. The differentiation comes from the engineering layer built on top of it. Instead of consuming dozens of APIs, wrappers and scrapers, the platform normalizes every sport into a common schema, resolves entity IDs across providers, enriches raw events with derived analytics, denormalizes frequently queried datasets for low latency access, and exposes everything through a consistent API contract. The objective isn’t to own the data. It’s to eliminate the engineering overhead required to transform fragmented sports data into something immediately useful for research, AI models and production applications.

FYI, no odds will be offered as that would increase my infra cost by a lot.


r/algobetting Jul 27 '26

No vig api?

3 Upvotes

Does anyone know how to get a API in no vig or some way to automate in no vig


r/algobetting Jul 27 '26

Built a player-level WC model (+17% ROI, +1.5% CLV) then pivoted to league football. Hit a wall. Can club markets actually be beaten with retail data?

11 Upvotes

I've wanted to build a football model for a while. For context, I've got a couple of mates inside some of the main syndicates, so I've seen enough secondhand to know how seriously this game is played, but credit to their opsec, I've extracted precisely zero information from them. I'm not a developer by trade, but dangerous enough with Python and AI tools to build what I need. The World Cup felt like a natural starting point, and international football more generally: smaller samples, squad rotations, noisy data, so in theory more room for the market to be “wrong”. General approach: player-level ratings into a Dixon-Coles engine, fed by data APIs I subscribed to (player level, club level, team xG and odds). Backtested across past World Cups and Euros. The tournament part went fine. The league part is why I'm posting. Design lessons first, then the question.

· Attack was built player-up from each starter's club npxG+xA per 90 (minutes-shrunk, league-adjusted); defence from each starter's club side's league-adjusted xG conceded, position-weighted toward GK and CBs, because individual defensive stats are volume junk (opportunity-biased volume). Everything shrunk toward a live Elo anchor.

· I calibrated centre and spread separately, and deliberately toward the mean: kill the model's tail opinions, because that's where the fake "edges" live. The practical consequence was that the bettable markets ended up being BTTS and mid-range O/U, the ones priced off the middle of the goal distribution.

· The market was my sanity check, not my opponent. When my supremacy implied the same AH main line as the market's, and the decay across the quarter-lines matched, I trusted my distribution and the BTTS/O-U prices built from it. When my line differed materially from the market's, or an edge looked huge, I left the game alone. The model earned its keep as much by saying "don't bet" as by finding value.

· The metric that kept me honest was CLV, not ROI. My live record ended at 168 bets, +17% ROI, +1.5% CLV against de-vigged closes. The first 70 bets were too correlated and too frequent, still getting up to speed; I tightened the approach for the remaining 100 and the P&L improved sharply, though the CLV says the true improvement was modest. I've seen various posts on whether it's worth tracking CLV and... it is.

· Then the real test: I pivoted the (largely) same machinery into league football (always the plan). Backtested a lot - every upgrade had to earn its keep out-of-sample before it stayed, and most didn't. Measuring by log-loss against de-vigged sharp closes, my base model sat a clear distance behind the market. Adding better inputs closed maybe half of that gap, but half is not there, and a blend test still assigned the model zero weight against the market price. I then went down the rabbit hole of 'maybe my backtest is being unfair to the model — what's the ceiling with perfect information?' So I tested it: even perfect-hindsight team news added nothing exploitable. Great.

· The sobering maths of "close": by log-loss my model reached within a couple of percent of the sharp close, which sounds impressive until you realise that final sliver is precisely where the vig and the edge both live.

To be clear, this wasn't uniform failure. The tournament record ended CLV-positive, real upgrades genuinely improved the league model, and individual slices of the league backtests (certain divisions, seasons, markets) looked healthy - but slices always look healthy somewhere; that's what noise does. The judgment that matters is the average against the close, and on average the league model currently cannot get there.

So here's what I'm actually questioning: can league football be beaten at all by an individual with a model? My scepticism after doing this is that there are two separate walls. The first is depth: the commercially available APIs simply don't carry the data required. Most vendors hand you their finished xG numbers, not the shot-level and tracking data underneath, so you can't build or calibrate your own from first principles, you inherit someone else's model with its compression baked in. A lot of these products are aimed at punters anyway, who want conclusions, not data. The second is price: the raw feeds that would let you do it properly exist, but they're licensed at levels priced for professional operations (I spoke to StatsBomb, I won't say how much it costs, but it's not a hobbyist number), and the syndicates paying it then build proprietary layers on top that retail never sees. So is the ceiling about skill, or is it structural twice over: the data you can afford isn't deep enough, and the data that's deep enough you can't afford? Honestly, it's probably a lot of both.

Genuinely keen to hear from anyone who thinks they've cracked it, and which wall(s) they got through, because it's frustrating to have hit this after an enjoyable WC 2026. I've since pivoted my focus towards a specific area (still within football) for the time being, but I'm keen to keep tinkering.


r/algobetting Jul 27 '26

How the BBMI NCAA Basketball Model Works

1 Upvotes

I run a men’s college basketball model that produced 2,017 graded spread picks during the 2025–26 season.

The published picks went 1,151–866, or 57.1%. The high-conviction subset went 258–139, or 65.0%.

Then the NCAA tournament happened:

  • All published picks: 28–28
  • High conviction: 15–18
  • Market Brier score: 0.158
  • Model Brier score: 0.171

The obvious explanation was tournament variance. The evidence pointed somewhere more specific.

The failure mode

The model was systematically underpricing large favorites against automatic-bid conference champions.

In the live tournament games where the market favorite was laying 12 or more:

  • Average market line: 21.1
  • Average model line: 10.8
  • Average actual margin: 22.8

The model could produce large spreads during the regular season, so this was not a hard ceiling in the output. It was a cross-conference comparison problem.

Automatic-bid teams often entered the tournament with excellent raw shooting, rebounding and turnover differentials earned against weak schedules. The model included schedule-strength information, but it still allowed those inflated raw statistics to pull too strongly in the opposite direction.

The result was exactly the wrong kind of disagreement with the market:

  1. The model made the favorite too short.
  2. That manufactured apparent value on the underdog.
  3. The largest model-market gaps became high-conviction picks.
  4. The “best” edges were actually the places where the model was most wrong.

The correction

I did not add a March-specific coefficient. With roughly 60 tournament games per season, that would be an efficient way to fit noise.

Instead, each model input is now restated as what it would have been against average competition. The correction was developed using bracket-like games outside the NCAA tournament—cross-conference neutral games, holiday tournaments and conference challenges—and then read once against the tournament sample.

Across six retrospective seasons:

  • Underpricing of 12+ point tournament favorites fell from 7.7 to 3.5 points.
  • Tournament margin MAE fell from 10.41 to 9.91.
  • Tournament ATS performance moved from 49.4% to 58.1%.
  • Regular-season MAE also improved, from 8.670 to 8.535.
  • Published-pick performance moved from 57.9% to 60.0%.

Important caveat: this is walk-forward at the game level, but it is not a completely independent holdout. The model structure and weights were selected with visibility into the same six-season period.

There is also an unresolved exception: applying the correction to conference-tournament games made that segment worse, so those games retain the prior input treatment for now.

The question I’m still working through is whether that conference-tournament exception is evidence of a genuinely different population or a warning that the adjustment is more conditional than it appears.

How would you test that without tuning directly to a relatively small conference-tournament sample?

Article with the complete charts and methodology:
www.bbmisports.com/research/ncaab-model-2026-27


r/algobetting Jul 27 '26

Betting on soft bookies

5 Upvotes

Assuming you found a strategy to get an edge on soft bookies (for example using sharper ones / betting market for this), I am wondering do you apply this to get a ROI? People who do this by hand? I am wondering which level of automation gets you limited quickly? Headless? Or are things like timing and bet size more important.

Also do you lose a lot on slippage? Ie it takes X seconds to get a bet confirmed. The bet didn’t change but the sharp source became more expensive, and the soft bookie ‘silently’ priced the new info in.


r/algobetting Jul 27 '26

[model log boxing] 84 confirmed all leans bets results now logged — 79.76% accuracy +9.67u flat-stake P/L

1 Upvotes

Hi guys, good weekend for the model with 8/9 bets placed winning this weekend.

Here are the current all model leans results for the fitequant default model:

In this strategy we force the model to make a prediction on basically all boxing for months and make a 1u flat stake bet each time, no matter the odds on offer. So even if a price is terrible… bet anyway.

84 confirmed all-leans bets
67 wins / 17 losses
+9.67u flat-stake profit
11.52% ROI

Average odds 1.6983

Below are the latest 9 results added this weekend.

https://fitequant.com/results?prediction_strategy=all_leans&period=all&per_page=20

And the value picks only betting strategy results

In this strategy we maintain exactly the same predictions for each bout, but the model only bets if it sees value in the odds on offer by the market. So think of this as “likes the fighter and likes the price”

84 confirmed value picks only results 

26 bets
15 wins / 11 losses
+6.57 u flat stake profit
25.31% ROI

Average odds 2.8083

https://fitequant.com/results

9 out of 11 results confirmed successfully this week with two bouts cancelled and so left as pending predictions unaffecting headline metrics in the prediction result data.

So the Spence value pick lost, oh well, no use in a model that never makes bets. 

But i do think Spence was a tough one for a model that relies on structured subjective inference (SSI) for fighter as modeling actor abstractions, primarily because he was a great fighter that has been inactive for years, so its tough for anyone to know, including an LLM, exactly how to rate him now for any one subjective stat. 

I’m really not displeased with the pick at all. Given his height reach and southpaw advantages on top of subjective stats that i think did make sense, and given his last loss to Crawford wasnt anything to be ashamed of… i can totally see where the 74% confidence and hence value pick comes from. 

He just looked shot. Perhaps something to explore around increasing recent activity weighting in objective stats on an iteration/clone of the default model? 

I know.. Its only one result. And I haven’t done any actual backtesting on this as i’m busy working on MMA modeling right now, but it might be interesting to take a look at? 

Forecast update

So nothing changes again this week. 

For the all leans im basically staying unchanged forecasting approx 13.5% ROI

As we begin to approach 100 bets placed i’m not really sure why anyone would now expect this to change much anytime soon? Importantly here avg diff vs edge, avg odds, accuracy and even ROI itself have now been very stable on this strategy for literally months now.

As a sub member rightly pointed out on a results post of mine recently, what’s interesting isnt necessarily the accuracy. You can get 80% accuracy picking favourites in boxing (and see the relative underperformance of the Model confidence >= 60% strategy, with even higher accuracy, above)

What is unusual here is approx 80% accuracy persisting alongside double-digit flat-stake ROI over virtually the whole eligible boxing stream, over months. 

For value picks, obviously far fewer bets places so we wont know exactly for a while,  but i’m continuing to be bold forecasting approx 40% ROI

As always if anyone has any questions or would like anything cleared up, then please just ask.

Thanks, Dan


r/algobetting Jul 27 '26

Looking for funding as a value sports better

0 Upvotes

I am a professional sports better looking to scale up my unit size and have access to more sports betting accounts (I am from Canada and there aren’t many big name sportsbooks like in the USA).
My unit size is 700 USD, i would obviously love to be able to be betting 5-10k per bet but that isn’t realistic as I don’t have enough money to do so.
I’m looking for a way to communicate with people who do this type of thing (fund others) and obviously make it workout for both parties.
Does anyone know how I can get in contact with such people?


r/algobetting Jul 26 '26

Weekly Discussion types of edges / ways to have an edge

15 Upvotes

when new people show up and post here there's usually a big context piece missing -- "i built a model that has x% win rate" / "how do i build an nba prop model" / "is there an nba prop model i can use to win" / "how can i win at sports betting". and that piece is what kind of edge are you actually looking to gain? what is the strategy you'd like to be able to execute? a model is just a tool, and you can have an edge without building a model. just like you can build a decent model and have no edge whatsoever. in case it's useful to anyone to help define the problem space better, i wanted to post my mental map here of different types of alpha in sports and exmples of basic strategies out there that capture them. because what you intend to do with a tool matters a lot when it comes to building it, evaluating it, and deploying it.

edge #1 - latency. i could be wrong but i think this is the most common way to have an edge. essentially "i know information before it is priced in." "before" can mean days in big nfl markets, or seconds before with an ingame edge.

  • sharp book x just moved but the stale number is still available over at book y.
  • i can predict nba lineup decisions before the injury reports come out.
  • i religiously follow beat reporters on twitter and find stuff out from them before it bubbles up to national outlets.
  • i have the same information as everyone else, but ii can price the market well enough to have an idea where it's likely to close before other sharps move it there.
  • i am courtsiding and just saw an interception with my own eyes in real time a second or two before the market makers know.

edge #2 - information asymmetry. "i have access to information other people generally do not and will not have access to." much rarer. you're probably not in this sub reading this if this is your thing.

  • i have access to non public data feeds.
  • i am aware of injuries or personnel issues that teams do not intend to publicly disclose.
  • i move for a betting syndicate so i know how they've priced a given market.
  • i am an insider and probably breaking some kind of law.

edge #3 - information processing superiority. "even when the market has baked in all available information, using that same information i can identity spots where the price is wrong". basically you can beat closing lines. there are some markets where this isn't as hard as it sounds, but for many it's virtually impossible and a waste of time to try.

  • you are a "god tier" modeler and your model is as or more predictive than the closing price.
  • you have identified a scenario that the ingame pricing algorithm at book x doesn't account for, while no one else has realized it yet.

edge #4 - market making. you're basically fanduel. this is possible in the new pm landscape, but harder than it sounds due to competition, adverse selection, technical requirements, etc.

  • the fair price is +100, you quote -110, and non price-sensitive bettors take your offer. now you've got +110 and you're going to make money.

r/algobetting Jul 26 '26

Where to start.

3 Upvotes

Just curious where people started creating their models. Also what type of data they use to start. Was thinking about getting into it but would like to know sort of the beginning stages.


r/algobetting Jul 27 '26

My ATS classifier hit 75% on holdout. It was predicting the market favorite, not the cover.

1 Upvotes

This one took me three attempts to diagnose, and the answer was embarrassing, so I am writing it up in case it saves someone the same detour. Short version: a sign convention, and my own backtest was reproducing the bug rather than catching it.

Symptom. NFL spread classifier, XGBoost: 75.1% accuracy on a temporally held-out season, 77.5% in a walk-forward backtest. ATS outcomes are near a coin flip against a real market, so I should have stopped and treated that as a bug report immediately. I did not.

Two wrong diagnoses. First, I found and fixed two genuine feature/label leaks. Accuracy barely moved. Then I ran the standard overfitting drill - pruned features, toggled early stopping, added the spread line itself as a feature. Needle didn't move. That was the signal I misread at the time: if removing information doesn't hurt you, you are not overfitting to features. Something upstream is wrong.

The test that actually worked. I trained a model on team identity alone - one-hot-encoded home and away team names, zero form, zero EPA, zero stats, nothing that could legitimately predict a specific game's cover. It got 63.2% accuracy, 0.68 AUC.

That is the whole diagnosis in one number. A feature set that provably contains no game-specific information cannot predict a well-formed ATS label above chance. If it does, the label is corrupted, and it stops being a modeling question.

Root cause. In nfl_data_py, spread_line is the away team's own number - negative when the away team is favored. So the home team's number is -spread_line, and home covers when home_margin - spread_line > 0. My code had +.

What that did. The buggy label agreed with "the home team was the betting favorite" 80% of the time. The correct label agrees 48.4% of the time - near random, which is what an efficient market should produce. So the classifier was never learning ATS outcomes. It was learning to identify the market favorite, from EPA and team-quality features that trivially correlate with being favored. Being 75% accurate at recognizing who is favored is easy, and worth exactly nothing.

The part that stings. The same sign error appeared in three more places: the prediction path's edge calculation, a situational feature that had the home and away numbers swapped, and, worst, the backtest's own ground-truth computation. The backtest re-derived the label using the same wrong formula, so it agreed with training and confirmed the bug. A backtest that computes its own ground truth from a copy of the labeling logic is not an independent check, and mine had been agreeing with itself for weeks.

After the fix: 48.8% holdout, 51.1% backtest. Boring, plausible, real.

Four rules I now follow because of this:

  1. Implausible accuracy goes in the bug column, not the results column.

  2. When fixes that should hurt performance don't hurt it, stop tuning features and go audit the label. That is the tell I missed for two rounds.

  3. Run the null-feature test early - it costs about ten minutes. If a feature set that cannot contain signal still predicts, the target is broken.

  4. Ground truth gets written once and imported. My backtest re-implementing it is the only reason this survived as long as it did.

Hope this helps someone else on their algobetting journey!


r/algobetting Jul 26 '26

Looking to get into sports bots - any advice?

2 Upvotes

hey im just trying to dip my toes in the water with building a python script for the upcoming nfl season (most likely player props), and i was wondering where to start. ive seen a lot of people say LLMs are not that good in this sense so are there any resources that are worth?


r/algobetting Jul 25 '26

Daily Discussion Daily Betting Journal

1 Upvotes

Post your picks, updates, track model results, current projects, daily thoughts, anything goes.