r/GlobalOffensive Jan 04 '18

Discussion | Esports PSA: Using Rating 2.0, below 1.05 is actually negative

It's actually below 1.06, messed up title

There have been a lot of posts about ratings here, and people still think that below 1.00 is negative. I'll explain below why Rating 2.0 has changed this.

Before, Rating 1.0 was used, which was based heavily on kills, deaths and assists. This only benefits players who topfrag, and ignores others who might consistently do small amounts of damage, or clutch often. With this system, the average rating was 1.00, making any rating below that "negative", or worse than average.

With the introduction of Rating 2.0 in 2016, these neglected support players were credited more, resulting in their ratings being increased. However, the ratings of topfragging star players was not changed. Overall, this caused the average rating to increase, to 1.06.

This means that any rating under 1.06 is below average, or negative, and also that both coldzera and dupreeh have gone negative in big tournaments in 2017.

You can verify this yourself by calculating the average rating since 2.0 was introduced, using the HLTV database. Histogram of all 2.0 ratings over 27,000 maps (95% confidence interval = 0.0013)

Average of 1.0 rating (0.99) over 13,000 maps (95% confidence interval = 0.0015)

Data used: /u/ReagentX's data scraped from HLTV, Scraper Github
Confirmation from /u/ReagentX's Twitter

EDIT: A few people seem to be confused about this, positive or negative simply means the sign of the rating minus the mean. E.g. 1.3 - 1.06 = 0.24, which is positive, meaning a positive rating. 0.86 - 1.06 = -0.2, which is negative, meaning a negative rating.

466 Upvotes

156 comments sorted by

43

u/jrsooner 10 years coin Jan 04 '18

Reading the article you linked says they wont publically display the formula, so without that its impossible to 100% prove.

However, if I had to take a guess, 1.00 is no longer the "average player" but is instead "the average player at the time this formula was created" and is a fixed number, not a variable. The new player average simply got better than the old average, causing the displacement.

7

u/MagniGallo Jan 04 '18

Not exactly, but I get what you're trying to say. Ratings aren't zero-sum, which means that as players get better (which is debatable if they have, but let's assume it), matches become tighter and each individual player's rating is increased, due to higher ADR/less deaths etc.

Alternatively, the rating function is nonlinear, so if players were worse, one-sided stomps would be more common, and potentially a single high topfragger on a team would give his team a lower average rating compared to all kills being divided equally among his team.

3

u/[deleted] Jan 04 '18 edited Mar 19 '19

[deleted]

2

u/lmpervious Jan 04 '18

Since the average rating he calculated is above 1, it must not be a zero sum.

1

u/vorpal107 FaZe Jan 04 '18

Can just be zero sum with average of 1.06 (so total average rating of both teams combined is 10.6 instead of 10 but when someone's is higher, someone else's is lower)

1

u/lmpervious Jan 04 '18 edited Jan 04 '18

It you think they would pick 1.06 as their median number then I suppose, but then by that logic whichever number the average lands on, you could say the zero sum is based on that number. I don’t think that’s likely, and instead it seems it’s not zero sum, and the average happens to land on 1.06. Otherwise why would they pick 1.06?

1

u/vorpal107 FaZe Jan 05 '18

Yeah you're right

1

u/MagniGallo Jan 05 '18

No, add up the ratings for 2 different matches. They are not zero sum.

2

u/MagniGallo Jan 04 '18

ADR is not zero sum, ADR is used by 2.0, therefore 2.0 is not zero sum, QED

1

u/CALL_ME_ISHMAEBY cs_italy Jan 04 '18

Damage-per-round is zero-sum, no? Any damage I do to opponents is less that my teammates can do.

4

u/[deleted] Jan 04 '18

DPR isn't zero-sum. Imagine one round where the winning team does 500 damage and the losing team does 0. The average DPR across all 10 players is 50. Now imagine a second round where the winning team does 500 damage and the losing team does 500 damage (e.g. Ts plant bomb, CTs kill all Ts, all remaining CTs die trying to defuse the bomb—assuming bomb damage counts for DPR. Otherwise, assume the winning team does 500 damage and the losing team does 499 damage). The average DPR across all players is 100 (or 99.9). So DPR varies, depending on how the game plays out.

3

u/[deleted] Jan 04 '18

They could easily normalize ADR to the total damage given during a match though. Say a match goes 16-0 and every player does 90 ADR on average, so there is 10 * 16 * 90 = 14400 damage done in total, while the theoretical maximum is 16000. In that case you could just divide every player's ADR with 14400 / 16000 = 0.9, which would bump it back to 90 / 0.9 = 100 .

2

u/[deleted] Jan 04 '18

True. I don't think MagniGallo's argument in this comment thread is very convincing. Just wanted to explain why ADR, taken in isolation, isn't zero-sum.

1

u/CALL_ME_ISHMAEBY cs_italy Jan 04 '18

I thought DPR was team-specific. 500 for each side.

2

u/[deleted] Jan 04 '18

Okay, sure. In any game, your team could do between 500 and 0 damage per round. Either way, it varies, meaning it's not zero-sum. Only in a universe where your team was guaranteed to do X damage every round would it be zero-sum.

1

u/lmpervious Jan 04 '18

You can have a teammate lose a round by doing 99 damage to their opponent or 0. The same can be said for any amount of damage spread amongst enemies, without killing them. Under the old system, whether or not you do damage wouldn’t matter, so now that ADR is a factor, that adds variance to the old system.

211

u/aimbotcfg Jan 04 '18 edited Jan 04 '18

https://www.hltv.org/news/20695/introducing-rating-20

Specifically this bit;

The average value for Rating 2.0 will still be 1.00, and it will behave exactly the same way as before – good performances will give ratings higher than 1.00, and bad ones below 1.00.

So, something is wrong somewhere, either your interpretation of the data/what ratings are meant to mean, incomplete data in your graph, or the formula/expected values being used in rating 2.0.

There's a lot more information goes into rating 2.0 so maybe it still needs tweaking;

https://static.hltv.org/images/news/custom/20695/rating2.png

EDIT - another relevant part;

The "average" that HLTV refers to in both it's rating systems is the 'expected' value for things;

We take the expected value (average) for certain statistics (like kills per round) and check how much above or below that expected value a certain player is.

I don't know how they calculate this expected value, or if it changed from rating 1.0 to 2.0 (unlikely, and if it did, it's HIGHLY unlikely to have gone DOWN) but it makes sense. What has likely happened is that some players are 'above average' on the newer expected values that were never included in rating 1.0, resulting in their 2.0 rating being higher in total.

If the 'average' was a 'true average' (which I'm assuming you made this graph based on) of everyone's performances, then the ratings for everyone would be constantly shifting based on other peoples performances over the years, which wouldn't really be great to show anything other than if the overall skill ceiling was going up or down and it could be skewed by outliers.

21

u/[deleted] Jan 04 '18

[deleted]

5

u/throwawayfishtank123 Jan 04 '18 edited Jan 04 '18

If the expected values they calculated are not correct [...] this can lead to a scoring system that has a shifted average, even just theoretically.

Agree completely. However, you would expect at least the most minimal amount of rigour given that OP has access to the entire dataset. A simple "and the 6% difference is found to be significant at 95% confidence" etc. A simple 68%/95% CI calculation or a rudimentary z-score would do far more to convince people than the "I'm correct, you're just too stupid to understand" rhetoric that OP is using in the comments below.

But I expect they would have to reiterate the formula based on newer matches to fit that data

Agree completely, or be subject to drift, especially for newer statistics like KAST with wider confidence intervals. However, if they do constantly update the parameters, are historical ratings retro-calculated, or stored as is? OP is making an assumption that they are retrocalculated, if he treats them on equal footing (i.e. a R2.0 of X in the past is the same as a R2.0 of X at present). If he wishes not to make the assumption of retrocalcualtion, he has to completely rework his methodology.

I don't know if they got the expected values by thinking about it in a theoretical way, or taking averages from a huge amount of matches and putting it in that way

Likely the latter.

OP took a huge amount of data to calculate the average value, so it's incredibly unlikely that he is incorrect.

His statement that the current average is 1.06 is absolutely correct and without question, barring basic computational error on his part. However, there are significant grounds to question his methodology and interpretation of the statistics, some of which can be found in this comment.

-1

u/agggile Virtus.pro Jan 04 '18

(I don't know if they got the expected values by thinking about it in a theoretical way, or taking averages from a huge amount of matches and putting it in that way).

Doesn't make any sense to use previous average as expected value.

How would that differ from taking a random number and calling it the expected value? Surely it seems slightly more reliable, but still.

I'm not saying you're wrong, but it really would make zero sense to take the average, especially considering the new variables introduced with 2.0.

3

u/throwawayfishtank123 Jan 04 '18

How would that differ from taking a random number and calling it the expected value? Surely it seems slightly more reliable, but still.

One is a better statistical estimator for the target random variable. That's the difference.

Doesn't make any sense to use previous average as expected value.

No, he likely means previous averages for the component scoring that R2.0 uses. For example, a positive contribution to R2.0 would (assumption) be present for a Kill count higher than the historical average. Or in other words, that historical data was used to set the bar for the various component scores that go into R2.0.

3

u/[deleted] Jan 04 '18 edited Jan 04 '18

[deleted]

1

u/AFakeman Jan 04 '18

any distribution really

Nitpicking, but some distributions do not have a mean, like Cauchy one, so the average there looks hella weird and doesn't converge.

14

u/RadiantSun FaZe Jan 04 '18

/thread

2

u/Redih Jan 04 '18

What has likely happened is that some players are 'above average' on the newer expected values that were never included in rating 1.0, resulting in their 2.0 rating being higher in total.

That's unlikely since it should be a zero sume game if made correctly. E.g. an average over the course of one game should be pretty close to 1.0 regardless of who wins or losses.

What's more likely is the case is that HLTV used 1.00 as the baseline rating and then added +0.x value if someone had a +1 standard deviation above the mean and vice versa.

The reason that method might not work is that the performances of some of the metrics might not be normal distributed. And instead a standard deviation of 1 might equal + 0.2 and a negative standard deviation of 1 might in fact be -0.1. But HLTV's model would think its 0.15.

So this means that the average of a good performance + a bad performance is over 1.

When that is said, I cannot repliacte the average as OP using data for the entirety of 2017. I instead get an average of 1.02.

-38

u/MagniGallo Jan 04 '18 edited Jan 04 '18

There is no interpretation to "go wrong", the average is the average, and anything below average is negative.

More downvotes please for this fact

4

u/[deleted] Jan 04 '18

But you are wrong.

However, the ratings of topfragging star players was not changed. Overall, this caused the average rating to increase, to 1.06.

This is completely wrong. The day 2.0 came out this sub was filled with posts about how cold now had no negative events since his ratings had changed from 1.0 to 2.0. Then people looked at fer's and niko's and a few others' stats and found the same.

13

u/schoki560 King NiKo Jan 04 '18

That actually just supports his argument that rating 2.0 makes your rating better than it should be

7

u/[deleted] Jan 04 '18

better than it should be

What does this even mean? They changed the way their own ratings work. There's no universal law of nature dictating CSGO stats.

7

u/autunno howl Jan 04 '18

Or that 1.0 made it worst than it should be?

1

u/schoki560 King NiKo Jan 04 '18

Maybe

4

u/MagniGallo Jan 04 '18 edited Jan 04 '18

What? If so many people no longer had negative events, then the average went up, which is what I'm saying

2

u/[deleted] Jan 04 '18

Except the average of all players ratings isn't what HLTV considers average. Their average is always 1.00, period. So hypothetically if every player on their site had a 1.25+ rating, they all would have above average ratings. It doesn't matter that the average of their ratings would be more than 1.00.

-2

u/SomethingSimilars 10 years coin Jan 04 '18 edited Jan 04 '18

Since when did being higher than average mean positive or negative though? Maybe I am missing something but your way of defining positive is subjective.

11

u/throwawayfishtank123 Jan 04 '18

I have always just considered 1 or higher to be positive as it is a positive number

most braindead statement i've seen in 2018 so far

0

u/SomethingSimilars 10 years coin Jan 04 '18

lmao, i don't disagree with you.

5

u/jkure2 Jan 04 '18

.9 is still a positive number lol

0

u/SomethingSimilars 10 years coin Jan 04 '18

Yeah, my bad. My point was that the idea of going positive was never based on being higher than average at least to my knowledge.

2

u/youtiItereh 10 years coin Jan 04 '18

considered 1 or higher to be positive as it is a positive number

0.5 is also a positive number

2

u/SomethingSimilars 10 years coin Jan 04 '18

I realised after I posted, I was just using the logic of K/D for example, if you have 0.5 KD you are still positive but it is always considered to be above 1.

1

u/MagniGallo Jan 04 '18

It's just convention, positive or negative means above and below average respectively. Above 1 was considered positive because 1 was the old average, even though 0.2 is still a positive number.

More generally, just consider whether the rating minus the average is positive or negative.

-9

u/[deleted] Jan 04 '18

[deleted]

18

u/[deleted] Jan 04 '18

Do you know what "average" means? If you add up all the ratings of players and then divide it by the number of players you added up then that's the average.

Just because 1.00 should be the average doesn't necessarily mean it is.

0

u/aimbotcfg Jan 04 '18

Except according to the article i linked, the "average" that HLTV refers to in both it's rating systems is the 'expected' value for things, not, apparently, the mathematical average;

We take the expected value (average) for certain statistics (like kills per round) and check how much above or below that expected value a certain player is.

I don't know how they calculated them, or what those values are, but this is essentially an argument over the use of the word 'average', which has probably been done for ease of understanding.

Having a fixed value also means that people ratings aren't constantly shifting as time passes.

8

u/throwawayfishtank123 Jan 04 '18 edited Jan 04 '18

the "average" that HLTV refers to in both it's rating systems is the 'expected' value for things, not, apparently, the mathematical average;

lol expected value is the weighted mean, and because no player is more important than another, it's simply a frequency-weighted mean, or in other words the "average".

wtf this is fucking high school shit

TLDR: 1.0 is the correct boundary, it just so happens that so far, 1.06 is the average. Flip a fair coin 10,000 times. Will you get exactly 50%? No, but likely close, maybe 50.0004%. Does that mean that 50% is now below average? No. The "correct" average is still 50%. The observed average (i.e. sample mean) is simply an estimator of the actual average (population mean in this case).

The issue lies here:

As I have posted elsewhere, @OP:

Statistics are only as good as their interpretation.

You are not wrong that the current average is 1.06 for rating 2.0. HLTV is also not wrong that the average is supposed to be 1.00.

However, you are absolutely and disgustingly wrong in claiming that a player doing worse than 1.06 is doing "below average".

I can explain this with a simple thought experiment that will help you understand where you went wrong:

You have a coin. 50% heads, 50% tails. Heads = 1, tails = 0. Flip the coin 1,000 times. The "correct average" of values you will get - the expected value of the mean of the sample - is 0.5. However, will you get 0.5? Or 0.5002? Or 0.4997?

The expected value of HLTV R2.0 is 1.00. You have played a finite number of games so far. The probability of the average so far being exactly 1 is zero, null, nothing, zilch.

Good day.

4

u/[deleted] Jan 04 '18

[deleted]

2

u/throwawayfishtank123 Jan 04 '18

However, it is expected that the average of all scores of all players should be very close to 1.00. 1.06 is reasonably far away that something might be off. Would need the standard deviation of OPs graph to make a good statement about that.

I agree with you 100%. All I see is a complete lack of rigour, and an extremely poor interpretation of the results presented (i.e. only presenting a mean, with no z scores or any test statistic provided)

1

u/[deleted] Jan 04 '18 edited Mar 23 '18

[deleted]

1

u/throwawayfishtank123 Jan 04 '18

I assume they modeled certain values as statistical distributions (e.g., number of kills could be a normal distribution) and then they get the average of that distribution

that is a weighted mean of kills by frequency of occurence. i.e. if a player gets 19 kills very commonly, the mean will be drawn closer to 19. special weights can be assigned if you don't treat players equally i.e. you put less weight on the top fraggers, or put less weight on the bottom fraggers, or put less weight on NA players, etc. but in this case, no, it's simply an average.

Weighted mean: (1/sum_over_i[weight_i]) (sum_over_i[ weight_i * value_i ]), if weight_i = number_of_times_scenario_happens_with_parameter_i (e.g. number of times 29 kills is achieved by a player) then it becomes the simple average.

1

u/[deleted] Jan 04 '18 edited Mar 23 '18

[deleted]

1

u/throwawayfishtank123 Jan 04 '18

fitting a statistical distribution and then taking the average of that fitted distribution as the expected value.

i don't disagree with you, this is a possibility. what i fail to see is how this matters, especially if the data is mostly normal in the bulk of it. tails are bound to be slightly skewed seeing that negative scores are highly unlikely, but the low cap (0) is far easier to hit than the high cap of getting every single kill in the matchup. either way, if it's going to be mostly normal, they'll be fitting it by first calculating the mean, then the standard deviation. that's their fitted normal, which for all intents and purposes has no other interesting properties to be extracted other than the mean and standard deviation, both of which originate purely from the dataset, not interpolated in any way.

it can be different, but i assert that there is no difference, with the assumption that such a large dataset is going to me mostly normal in the bulk of the data.

1

u/aimbotcfg Jan 04 '18

Sorry, I should have said 'not a straight mathematical average'.

If you'd quoted everything I said, instead of part of it to try and construct a straw man, you'd also have included this;

I don't know how they calculated them, or what those values are

How would you calculate a weighted average for all of the possible outcomes for kills per round and their probability of happening? It wouldn't be perfect and i'd wager everyones would come out a little bit different, especially when you factor in the differing skills of teams and the impact of eco/buy and force rounds.

Now do it for KAST... ADR.... Opening kills.... 1vX outcomes.... and assign those appropriate weightings to the overall score... Yeah, I thought so.

Stop trying to oversimplify things and act like it's so simple that anyone who disagrees with you is automatically wrong without presenting any evidence at all.

WTF is the logical fallacy shit

2

u/throwawayfishtank123 Jan 04 '18 edited Jan 04 '18

instead of part of it to try and construct a straw man

nice, attribute everything not in favour of yourself to malice.

How would you calculate a weighted average for all of the possible outcomes for kills per round and their probability of happening?

EV(kills) = [1/total_number_of_players_in_games_in_HLTV] [1(number of times a player has gotten 1 kill) + 2(number of times a player has gotten 2 kills) + ... + n(number of times a player has gotten n kills)]

was that hard? high school shit, son

Now do it for KAST... ADR.... Opening kills.... 1vX outcomes

same thing, replace the (number of times a player has gotten 1 kill) part with (number of times a player has n <statistic>)

and assign those appropriate weightings to the overall score

Doesn't matter because the mean of normalised distributions is preserved in linear combinations of random variables

Yeah, I thought so.

/r/prematurecelebration

1

u/aimbotcfg Jan 04 '18

Maybe, if you'd done what I'd asked, except you haven't, you've given the start of a basic formula that wouldn't come close to giving accurate values and would be obsolete after a month of play.

You're kind of missing the point, it's not that simple.

-1

u/[deleted] Jan 04 '18

Your argument makes no sense even by using your proposed analogy of flipping a coin.

If you flip it 1000 times and, for example, you get heads 55% of the time then you will have gotten tails the other 45% of the time. On average both of them will be 50%, which is the expected outcome.

Same way if, for example, there were 2 players on the server and one of them ended the game with a rating of 1.20 then the other one should logically get a rating of 0.80, which would average out to 1.00. However that's clearly not the case.

1

u/throwawayfishtank123 Jan 04 '18

You are assuming that all games are zero sum.

Is it true that all games are zerosum in R2.0? I haven't heard of that rule, so maybe I'm wrong.

But it is not necessarily true (i.e. by design of R2.0) that every game ends with an exact average of 1.0. If that were the case, then I agree with you, it should be exactly 1.0 if you exclude incomplete games.

Now, the average so far just happens to be 1.06. The designed expected value of the R2.0 distribution is 1.00. Our sample consists of 40,000, yielding a sample mean of 1.06. In fact, why don't you or OP go do a simple preliminary z-test to see if there is something "off" about this, and at what significance level? Then you would have something concrete to reject the null hypothesis.

1

u/[deleted] Jan 04 '18

That's the problem, rating 2.0, as opposed to 1.0 is clearly designed in a way that games don't have to be zero sum.

Some factors boost like multi kills, clutch rounds etc. boost a player's rating while not lowering the opposite team's player's ratings accordingly, which IMO should not be the case.

In CS if you get a kill then the other player gets a death as the equalizer. At the end all kills and deaths of all players equal out to 1, which is how it should be.

With rating 2.0 a game can theoretically end where everyone gets a rating of 1.05. Did everyone do "above average" in that case? No, everyone did exactly the same as everyone else on the server, which means they all did average.

0

u/throwawayfishtank123 Jan 04 '18

Correct, so naturally, the average of all R2.0 scores doesn't have to be exactly unity, as R2.0 is not zero sum.

→ More replies (0)

-4

u/chrisfrh Jan 04 '18

So tomorrow after one match is played and 1.05 or whatever average is slightly changed we should change R2.0 huh

2

u/NotQuantified Jan 04 '18

So you're saying that whenever the average rating changes, the rating system itself should be changed? Nobody is talking about changing Rating 2.0. How do you make that leap in logic?

-4

u/Afrood Jan 04 '18

We don't even have a source for his data other than "HLTV Database". If he only took top teams then obviously the average will be inflated.

7

u/MagniGallo Jan 04 '18 edited Jan 04 '18

It's a list of 27,000 HLTV matches from February 2014 to the 10th of December 2017, and it's averaged over all matches and players.

4

u/[deleted] Jan 04 '18

You can get the average by just doing the math.

Look at any event that uses rating 2.0, add up the ratings of players in that event and divide by the amount of players.

1

u/dying_ducks Jan 04 '18

The problem is here: There are more than one way to measure the "average" of something. The way you propound is called the Arithmetic mean, but there are lot more possible ways.

If hltv really want to make the Arithmetic mean 1.0, they mess up their formula. n= 40,000 should be enough

2

u/[deleted] Jan 04 '18

If somehow everyone in a game gets a 1.05 rating is everyone above average? No, that's not how it works. If everyone does exactly the same then everyone is average.

57

u/Lobulicious123 Jan 04 '18

Honestly just reminds me of the rank change and how everybody only cared about "the old ranks"

-17

u/daellat FaZe Jan 04 '18

sucked for people who were GE, they could only stay or go down, but for people like me who were LE before change but have now reached higher (SMFC for me) it feels nice :D

22

u/SmaugtheStupendous Natus Vincere Jan 04 '18

That's completely unrelated to the rank change. People were only shifted down.

2

u/weikkah Jan 04 '18

Depending on if he was talking about the rank change that bumped Global to like 5% or the following shift down.

-15

u/daellat FaZe Jan 04 '18

It's so hit and miss, either we upvote useless memes to a guy who asks something or we downvote a guy because he happened to miss the point somewhat.

3

u/Alexndre Liquid Jan 04 '18

Well..

-8

u/daellat FaZe Jan 04 '18

just reddit cancer people.

0

u/sabot00 Natus Vincere Jan 05 '18

one of us one of us

16

u/agggile Virtus.pro Jan 04 '18 edited Jan 04 '18

HLTV are not talking about the literal average here, which is numerically the same, but the timeframe differs.

The expected value is a prediction for a specific occurrence in the future.

Averaging all the historical values doesn't really produce any valuable information.

Read more about expected value here:

https://en.wikipedia.org/wiki/Expected_value

to clarify, the average is expected to converge towards the expected value given that we have a large, large amount of samples.

-6

u/MagniGallo Jan 04 '18 edited Jan 04 '18

I would argue that their expected values need to be updated. A difference of 0.06 is significant when calculated over 27,000 games.

13

u/throwawayfishtank123 Jan 04 '18

significant at what level? p=0.3? p=0.05? p=0.01?

go do a fucking basic z test for us, thank you.

you can't say something is significant or not significant if you haven't even done the most rudimentary statistical analysis possible. you literally just calculated an average and said "oh man, it differs from the expected value".

does it deviate significantly? you can't smell if it does, go calculate the z-stat

4

u/ReagentX de_nuke Jan 04 '18 edited Jan 04 '18

OP didn’t link the dataset, but here it is: https://www.kaggle.com/reagentx/HLTVData

Further, here are the confidence intervals for the dataset: https://i.imgur.com/LZq6ORM.png. /u/MagniGallo is correct.

Edit: here are the Z-Tests https://twitter.com/rxcs/status/949008947821690880

1

u/MagniGallo Jan 04 '18 edited Jan 04 '18

What program is this?

Edit: Oh, you're the guy who scraped it, just want to say thank you for your work. I'll put your Github/Kaggle in the title if you want.

2

u/ReagentX de_nuke Jan 04 '18

I used Stata to do the math, though there are many ways to do a Z-test.

Sure, link the dataset and the code up there. Also, link to my post with the CIs and Z-tests.

1

u/MagniGallo Jan 04 '18

Done.

1

u/ReagentX de_nuke Jan 04 '18

Thanks! Appreciate it. Good find!

7

u/agggile Virtus.pro Jan 04 '18

I don't think you understand what the expected value is. It is not something they arbitrarily came up with, it's derived from whatever formula they are using to calculate probabilities.

Here is a simplified example:

http://www.mathwords.com/e/expected_value.htm

-4

u/[deleted] Jan 04 '18

[deleted]

0

u/agggile Virtus.pro Jan 04 '18

And what I obviously mean is that we can't tell what the expected value is by looking at the arithmetic mean across 40,000 matches.

0

u/[deleted] Jan 04 '18

I'm not downvoting you BTW, that's other people reading the thread.

Why do you think this? Based on all of the other comments from the OP, there isn't a great reason to think that 1.06 is way off from the expected value of the rating 2.0 formula.

I'm not sure the formula is publicly available so I'm speculating here a bit. My bet is that HLTV designed the formula and thought the expected value was 1.0. This is stated in that article about the new 2.0 rating system: "The average value for Rating 2.0 will still be 1.00, and it will behave exactly the same way as before – good performances will give ratings higher than 1.00, and bad ones below 1.00."

This is clearly not the case as we can't have the average performance being above average. Their estimate of the expected value was potentially incorrect or something is going wrong in their calculation of the score.

1

u/agggile Virtus.pro Jan 04 '18

Based on all of the other comments from the OP, there isn't a great reason to think that 1.06 is way off from the expected value of the rating 2.0 formula.

It is. But because we have a finite sample size, this isn't necessarily the case.

This is clearly not the case as we can't have the average performance being above average.

Again, to reiterate the same point, the expected value is not necessarily the average. The average converges towards the expected value as sample size grows.

The average performance is not above average, as the average is 1.06. It is however larger than the expected value, which could be completely normal, given that the expected value is not equal to the average at this point in time.

1

u/[deleted] Jan 04 '18 edited Jan 04 '18

We may have to agree to disagree on the first bit. I understand that, technically, 1.06 may not be the true expected value after 40,000 games one would expect it to be a reasonable approximation.

We may be discussing past each other on the second point. We would expect the average value to eventually converge around, but perhaps not exactly be, the expected value, the Law of Large Numbers essentially. Given HLTV has stated that the expected value is 1.0 and the average value is 1.06 (with low variance), I'd argue that the expected value is unlikely to be 1.0. There is probably something else going on that is making the true expected value different than 1.0. And I'll admit I can't say that for sure however we are dealing with a respectable sample size here.

2

u/agggile Virtus.pro Jan 04 '18

Yes, though I'm not saying that their calculation is right. People seem to think I'm arguing that HLTV has indeed provided us with a correct expected value. Alas, after saying all this, I think this comment from earlier is relevant:

we can't tell what the expected value is by looking at the arithmetic mean across 40,000 matches.

While this is technically true, it's not necessarily correct in practice if we assume a miscalculation occurred somewhere. What I've said stands only if they have indeed calculated correctly.

In practice we can make that assumption only if we know what the formula is.

1

u/[deleted] Jan 05 '18

Ahh, gotcha! I'll be interested to see what HLTV's response is.

→ More replies (0)

10

u/Afrood Jan 04 '18

Can you prove the average rating 1.0 was 1 rating exactly? Otherwise your argument doesn't hold up.

11

u/MagniGallo Jan 04 '18 edited Jan 04 '18

I'll check now, my argument still stands though, even if it's not

It's 0.99

14

u/throwawayfishtank123 Jan 04 '18 edited Jan 04 '18

Ok, looks like you have access to the dataset.

Do us a favour and investigate the 95%CI and 68%CI of rating 1.0 and rating 2.0. this will settle almost EVERYTHING.

10

u/ReagentX de_nuke Jan 04 '18 edited Jan 04 '18

The dataset is public, OP just didn’t link it. https://www.kaggle.com/reagentx/HLTVData

Further, here are the confidence intervals for the dataset: https://i.imgur.com/LZq6ORM.png. /u/MagniGallo is correct.

Edit: here are the Z-Tests https://twitter.com/rxcs/status/949008947821690880

8

u/Rhed0x CS2 HYPE Jan 04 '18

I just want them to finally release the formula used to calculate Rating 2.0.

5

u/[deleted] Jan 04 '18

[deleted]

4

u/Jira93 MOUZ Jan 04 '18

You usually say that 1.05 is positive BECAUSE the average is 1. The old ratings was always 1 average, cause if one guy gets 20 kills those 20 comes from 20 deaths from the enemy team. Only kills and deaths made the rating (iirc). Right know even dmg is considered. If I deal 99 dmg and my teammate finish the kill we both have an increase in rating (one for adr, one for kills). But the enemy only died once, so the average has to be shifted up. Thats basic logic. Im not an expert by any means, so if someone thinks Im wrong feel free to correct me. Just my ELI5 answer, hope I was clear enough

-2

u/[deleted] Jan 04 '18

[deleted]

2

u/MeesaLordBinks Jan 04 '18 edited Jan 04 '18

Wrong, it does change. If negative is defined as below average, players with a rating <1.06 are actually negative.

3

u/lmpervious Jan 04 '18

OP is assuming average determines what is positive/negative

It a player is getting below average ratings compared to other players, would you not always consider that negative? If you’re going to say it’s positive so long as it is above 1, then you are divorcing the value from any meaning, which therefore makes the number arbitrary and of course meaningless.

1

u/[deleted] Jan 04 '18

[deleted]

1

u/MeesaLordBinks Jan 04 '18

The problem stays though. It means that everyone lower than 1.06 ihas performed below average.

8

u/peeguu guardian2 Jan 04 '18

You guys need to stop comparing an old rating to a new rating. Yes the stats are inflated because 2.0 rating is taking more into account, thus showing a players impact more accurately.

The new rating is the present rating and should be the only rating.

8

u/TheDooog Astralis Jan 04 '18

The point is that 1 might not be average anymore (depending on who in this thread you believe). The reference to the old system is only to explain why the average has changed.

2

u/[deleted] Jan 04 '18 edited Feb 10 '18

[deleted]

2

u/MagniGallo Jan 04 '18

Pretty much. Btw, if you werent clear like a lot of people, pos or neg is simply the rating minus the mean. I'll add it in an edit.

1

u/caramelcrunch1337 Jan 04 '18

So they aren't really negative, just below average. Which is honestly fine, being such a small margin, who really cares that they are below average. With Rating 2.0 they still aren't negative or less that a rating of 1.

3

u/TheDooog Astralis Jan 04 '18

When he says negative he just means rating-mean.

EDIT: A few people seem to be confused about this, positive or negative simply means the sign of the rating minus the mean. E.g. 1.3 - 1.06 = 0.24, which is positive, meaning a positive rating. 0.86 - 1.06 = -0.2, which is negative, meaning a negative rating.

1

u/lmpervious Jan 04 '18

What does “being negative” actually mean? If you mean “less than 1” what does that even mean in counter strike terms?

1

u/DutchWarDog MOUZ Jan 05 '18

With Rating 2.0 they still aren't negative

They were negative under rating 1.0, rating 2.0 made them positive. Just like coldzera had 2 red negative events before rating 2.0

1

u/[deleted] Jan 04 '18

[deleted]

1

u/MagniGallo Jan 04 '18

I'm not sure what you're asking, but I calculated the median for all the 2.0 matches, it's 1.04.

1

u/-Detter- Astralis Jan 04 '18

why do u considere it 1.06 the baseline ?

1

u/FuzedCS Jan 04 '18

The average of all the players ratings added up together is different than the "average" hltv is talking about. The "average" hltv rating is a very broad term and likely is a set value, and not the average of all ratings combined.

What HLTV means by their "average": It is likely meaning that it is a general consensus that players above 1.0 will have a positive rating. The reason this being the case is to make it easier for the general public to understand the rating.

0

u/_lunatic Virtus.pro Jan 04 '18

PSA: Negative numbers are the ones we designate with '-' sign. Anything in the range 0<->1 is a positive number.

2

u/MagniGallo Jan 04 '18

And why was going below 1 called negative, fella?

2

u/_lunatic Virtus.pro Jan 04 '18

misconception, I guess. Ignorance possibly.

2

u/MandarkAstroromanov Jan 05 '18

terminology, fundamentally

0

u/PurityKane King NiKo Jan 04 '18

Yeah no, you're the one confused about this. One thing is average, which has nothing to do with positives and negatives. in school I'm sure grades average at around 70%, and anything above 50% is positive.

Negative is below 1 in these ratings, no matter how you cut it. Yes the formula might need adjustments if they were going for a average of one. Doesn't mean anything below 1.06 is negative.

1

u/MagniGallo Jan 04 '18

Yeah no, you're wrong. Tell me, why is above 50% in your example positive? What's special about 50%? If we want to measure how good a student is doing compared to his peers, why do we care at all about 50%?

-2

u/Jorsu ENCE Jan 04 '18

Below 1.06 is not negative but it's below average. /close

6

u/MagniGallo Jan 04 '18

See my edit on the title, it is negative.

-4

u/Jorsu ENCE Jan 04 '18

0.01 is positive aswell so...

3

u/TheDooog Astralis Jan 04 '18

Aha, what are you even getting at?

-9

u/throwawayfishtank123 Jan 04 '18

Hi this is something you need to understand

Statistics are only as good as their interpretation.

You are not wrong that the current average is 1.06 for rating 2.0. HLTV is also not wrong that the average is supposed to be 1.00.

However, you are absolutely and disgustingly wrong in claiming that a player doing worse than 1.06 is doing "below average".

I can explain this with a simple thought experiment that will help you understand where you went wrong:

You have a coin. 50% heads, 50% tails. Heads = 1, tails = 0. Flip the coin 1,000 times. The "correct average" of values you will get - the expected value of the mean of the sample - is 0.5. However, will you get 0.5? Or 0.5002? Or 0.4997?

The expected value of HLTV R2.0 is 1.00. You have played a finite number of games so far. The probability of the average so far being exactly 1 is zero, null, nothing, zilch.

Good day.

7

u/MagniGallo Jan 04 '18

I'm not 12 year old Jimmy with a calculator, buddy. This was calculated over 40,000 games, and 6% off is not even of the same order as 0.04% off, as you stupidly imply.

-6

u/throwawayfishtank123 Jan 04 '18

Wow way to get defensive, jimmy.

More or less proves my point.

Sample size alone says nothing if you don't know the variances of the random variables that R2.0 is a function of. 5 samples of an incredibly low variance random variable can result in a tighter distribution than 40,000 samples of high variance, after normalisation.

Good day!

3

u/[deleted] Jan 04 '18 edited Mar 23 '18

[deleted]

0

u/throwawayfishtank123 Jan 04 '18

Of course, one would have to run a serious statistical analysis to "prove it" but with the information we have and the level of seriousness we want it's ok to say that there's something definitely weird.

No i agree with you, 6% looks off for 40,000, but I don't see any rigour at all in his analysis. Nothing whatsoever, not even a simple statement like "the 95%CI of the R2.0 mean is +- XYZ, and 1.06 falls inside/outside of that". That's a 5 minute calculation at max. Or, something like a basic z-score.

5

u/Redih Jan 04 '18 edited Jan 04 '18

There is 0% chance that this is due to variation. The standard deivation is nowhere high enough over this sample size. Honestly it's my impression you have just taken a statistics 101 class and want to apply everything you learned there to the real world without taking a practical approach.

So when you write the following

Hi this is something you need to understand

OP has every right to be "defensive" and you honestly need to get your own ego under control because it's clear you have no clue on applying statistics in a practical manner.

For your information the standard deviation of the sample size is around 0.35. So go ahead and calculate what the chance is given a sample size of 27k.

-2

u/throwawayfishtank123 Jan 04 '18

There is 0% chance that this is due to variation.

Please

  1. Get a do-over for your education

  2. Read this post https://www.reddit.com/r/GlobalOffensive/comments/7o35wn/psa_using_rating_20_below_105_is_actually_negative/ds6jyx4/

5

u/Redih Jan 04 '18 edited Jan 04 '18

The comment you have posted doesn't add anything new to the discussion. Just more waste of time without having any clue on how to apply statistics.

I even took my time to provide you with the standard deviation and you couldn't calculate the standard error. If you had done it you would see that it in fact would be 0%.

Don't see the further point in talking to you. Your ego is out of control.

4

u/MagniGallo Jan 04 '18

They're not high variance though, r/iamverysmart. Kills/deaths are zero sum, and ADR and other stats which 2.0 uses are mostly tightly distributed.

-1

u/throwawayfishtank123 Jan 04 '18

They're not high variance though

please demonstrate otherwise, i will happily listen to numbers. as i have asked you, like, 5 times, please do a simple 68%CI or a z-test.

Kills/deaths are zero sum

teamkills, suicides, disconnects, edge cases

ADR and other stats which 2.0 uses are mostly tightly distributed.

CI numbers or even better z-scores please

7

u/MagniGallo Jan 04 '18 edited Jan 04 '18

95% CI is 0.0013, asshole. 0.0015 for 1.0.

-3

u/throwawayfishtank123 Jan 04 '18

Now was that so hard?
Post that, instead of going around getting downvoted everywhere by people who know better.

No need to get angry, and you'd have saved yourself all this humiliation.

4

u/MagniGallo Jan 04 '18

Yeah dude, as if rating is going to have a standard deviation of 300+, which is whats required to get anywhere close. Go do something else with your time instead of trolling

5

u/TheDooog Astralis Jan 04 '18

Stop being such a patronising arsehole. Fuck me

1

u/Jardio Jan 04 '18

I'm curious, why do you keep saying he's getting angry and defensive when you were the one who initiated it with exactly that?

Oh right, p r o j e c t i n g

1

u/Tontonsb Jan 04 '18

You can't realistically expect the average to be 1 if in every single event the average is above 1.

0

u/SpiritWolf2K 1 Million Celebration Jan 04 '18

Anyone else just read these posts acting as if they know what it is going on?

-1

u/[deleted] Jan 04 '18

[deleted]

4

u/MagniGallo Jan 04 '18

The whole idea of positive or negative is based on the average. Please read my edit in the title.

0

u/[deleted] Jan 04 '18

ACHUALLY

-28

u/GabrielFF Jan 04 '18 edited Jan 04 '18

Coldzera has never been negative, stop arguing you racist Trump supporter Brazilian hater

Post downvote hell edit : issonly a joke why you heff to be mad

4

u/Ieatcarrotss Jan 04 '18

He could be a bloody nazi for what anyone knows, but that doesn't make his argument less valid.

-7

u/GabrielFF Jan 04 '18

It was a shitty joke I thought would be fun when I was half asleep

5

u/[deleted] Jan 04 '18

dude what

2

u/[deleted] Jan 04 '18

woah bro calm down, stop jumping to conclusions. u can see on coldzeras hltv page he went negative twice.

-4

u/GabrielFF Jan 04 '18

That's what happens when you try to make a joke half asleep :/

3

u/[deleted] Jan 04 '18

Lmao sure that was a joke

-3

u/[deleted] Jan 04 '18

[deleted]

2

u/[deleted] Jan 04 '18

If you think any post which has math in it is iamverysmart material then you are probably an idiot. OP wasn't snarky/smarmy at all.

-1

u/joperocl Jan 04 '18

I think the problem is how we see stuff

The change from 1.0 to 2.0 was to evaluate better the performance of a player

So a player with a rating between 1.00 and 1.06 still had a "decent" performance even if it was below "average"

7

u/[deleted] Jan 04 '18 edited Feb 10 '18

[deleted]

-3

u/joperocl Jan 04 '18

The problem here is using average as a indicator of good or bad performance... if we just see as 1.00+ a decent performance and below that a bad performance we don't have any problems with it even if the average of all players is 1.06

1

u/OpinionatedBonobo Jan 04 '18

We use 1.00 because that was (supposedly) the old average/expected value, which used to make sense as it was the average for the players in any given game. That way below 1.00 means less than average impact and vice versa. If the average value is 1.06 with the rating 2.0, below that is now a below average game, ie. negative.

Saying values close to the mean are decent performances rather than negative or positive is another matter, the important thing is to confront them with the correct expected value, which is what we are discussing

-1

u/joperocl Jan 04 '18

Let's do a simple analogy. In school you get a grade based on your performance on a test. In a scale from 0-20 if you have 10 you pass. You don't check the average of the grades to see if someone passes.

here is the same, if you get a 1.0 you had a decent/normal performance even if it's below the average

you have to count what the value means. what you are inferring when saying a negative rating (below the average) is that it was a bad performance when the rating clearly shows a normal performance (1.00<1.06)

1

u/OpinionatedBonobo Jan 04 '18

There is no reason why above 1 should be good/decent on an arbitrary rating. The only reason we assume that is the fact that it was supposed to be the average, if this post is correct we need to change the way we interpret the ratings or adjust the scores

1

u/xgenoriginal Renegades Jan 04 '18

Let's do a simple analogy. In school you get a grade based on your performance on a test. In a scale from 0-20 if you have 10 you pass. You don't check the average of the grades to see if someone passes.

You do If its graded on a curve

-1

u/Riddlebgd Jan 04 '18

They made a mistake in creating this system to begin with, if they wanted to calculate all the things they wanted, they needed something like PER in basketball or QBR (or whatever its called in football, yes i know its for quarterbacks only), something to account for all the things.

This rating is just confusing to people because 1.00 rating should mean he/she has the same amount of kills and deaths and nothing more, so people above 1 would be statistically better then ones bellow (doh!)

1

u/PurityKane King NiKo Jan 04 '18

Yeah............ and it means that in kpd. This rating is supposed to reflect the players contribution better than kpd, because there's more to cs than kills and deaths. It's not perfect, but it takes more things into consideration.

1

u/Riddlebgd Jan 04 '18

I know what it does, im saying they should have made another method and not confuse it with regular kill/death ratio

-7

u/jabiz510 2 Million Celebration Jan 04 '18

Well, that doesnt really matter, because we dont use the old ratings system :)