Showing posts with label math. Show all posts
Showing posts with label math. Show all posts

Tuesday, April 22, 2014

On Regression to the Mean

I didn't play in the Barbu game at Lounge Day, but when I came back from eating at Mel's after my Titan game I was greeted by Andrew yelling at me about how regression to the mean doesn't exist because Kevin won the Barbu game and Andrew thinks Kevin was the worst player at the table and therefore if regression to the mean was a real thing Kevin should have lost the game.

I struggle to handle this. Is Andrew being intentionally obtuse? (Something he was rather vocal about disliking earlier in the weekend.) Maybe Andrew has absolutely no understanding about probability and statistics? (It may explain that history degree...) Maybe there's some other explanation for why he's being so vocal about ignoring what I see as basic facts about randomness, but I'm not seeing it. I feel like he's probably misusing this term because 'stats geeks' who follow professional hockey used it when his preferred NHL team, The Toronto Maple Leafs, got off to a good start to this latest season. They all predicted gloom and doom for the Leafs, and said gloom and doom came true with the Leafs missing the playoffs.

In the hopes that maybe he's just not understanding what is going on I'm going to explain what the term actually means. At its core regression to the mean is a concept which basically says when you have two independent events it doesn't matter what the outcome was for the first one; the second one expects to be 'average'. It's used when dealing with a large sample of outcomes, and is something experimenters need to keep in mind in order to deal with the innate randomness that may be going on under the surface.

With the Maple Leafs the 'stats geeks' were looking at the way the Leafs were actually playing at the start of the season and not on the actual outcome of the games they played. They look at things like how Phil Kessel shot 18% in the first 14 games of the season but shot more like 11% over his career. You can look at his hot start (9 goals in 14 games) and extrapolate that something happened to Phil Kessel this offseason that made him awesome. You could then assume he's going to put up 54 goals on the season and the Leafs were going to keep winning games and win the Stanley Cup and PLAN THE PARADE. The 'stats geeks' on the other hand looked at it, saw that he was probably just getting lucky bounces since his shooting percentage was significantly higher than expected, and that he was apt to cool off a little. I believe the same was true of their goaltending as well, who were playing better than they had historically. As far as puck possession, which the 'stats geeks' have found to be a better predictor of future success than goals or wins, Toronto was something like 29th out of 30 teams this season. The underlying stats showed they were really bad, and the most likely reason for their early success was simple dumb luck.

Now, it's important to point out that regression to the mean doesn't say we expect Kessel to miss a ton of shots so he ends up back at 11% shooting percentage again. He's not 'due' a cold streak to go with his hot streak. That's the gambler's fallacy. No, all regression to the mean is stating is that for the last 83% of the season we expect them to play around average. It just happens that average for the Leafs is worse than everyone but the Sabres and it wasn't likely that their hot start was going to salvage their season. All the hot start meant was that they were likely going to finish above their expected finishing spot. (Which should be a scary thought for Leafs fans, since they have the #8 pick in the draft with a team that's probably 'good enough' to have earned the #2 pick.)

So how should regression to the mean be applied to a money Barbu game in the Lounge? Realistically, not at all. There's a lot of randomness in any given hand of cards and you only play 32 deals total in a game of Barbu. Regression to the mean tells us nothing at all about how a given hand will play out except that we can expect it to be 'average'. But average in a game where one lucky card fall or opponent misplay takes you from scoring -252 to scoring +72? That's really not telling you a lot. What it does tell us is that should a 'bad player' get lucky in the first hard they do get to keep all those 'unjustified' points. Regression to the mean means they're apt to be average for the remaining 31 hands but they get to keep the good result from the first hand so overall they rate to finish higher than expected.

Nevermind that I'm not convinced that Kevin is actually the worst player at the table, let alone the worst player by such a large skill margin that he rates to be super negative. As far as I am aware the people at the table have played Barbu something like once every year or two; I'm sure everyone was making some silly mistakes which just serves to amplify the randomness. A bad player is less likely to get punished when the other players are screwing up. I don't think I've played a game of Barbu since Byung's bachelor party back in 2008 and while my brain is still telling me I'm awesome at the game I'm sure the reality is that I'm rusty and will make some silly mistakes rounding back into form.

We could apply regression to the mean if we set up a large enough sample size, though. Like, say, if we were to build out a money Barbu circuit. Tag me in with the 4 people who played the game on the weekend and let's play 50 games total with each player sitting out 10 games. Now we're talking each person playing 1280 hands. Now we're talking enough current experience to shake off a lot of the rust for most of the games. Someone getting lucky in any given game would impact their overall final result, but we expect the other 39 games to fall in line with the expected result and someone who is actually the worst player at the table really expects to finish negative and near the bottom of the pack.

I think if we actually did such a circuit I would finish at the top. I think Sky would also be positive. I think the other three would all be negative. I'm actually pretty sure that's how it would shake out. Because regression to the mean should kick in over that large a span of games and I am the best.

Tuesday, July 10, 2012

TrueSkill vs TrueSkill Mean

I finally got suckered into reading the Yucata.de forums about the recent change to the metagame system. Or rather, I read a couple of the numerous threads with countless posts on the topic. The Yucata forum is more civilized than most of the gaming forums I've read over the years but it doesn't make my head hurt any less. For the most part it's people making up numbers they think furthers their point of view and talking past each other in an effort to make their point known. I got dragged into it partially because I can't stand that someone might be wrong on the internet and partially because I truly believe the system is flawed and want somehow to convince the people in charge so they change it. So I parsed some data from my recent games and went flying in headfirst to provide my own made up numbers. *sigh*

At any rate, on the bus ride home I actually had an epiphany on my biggest issue. I was trying to talk myself out of starting a new account when I figured out just what it is that bothers me so much about the new system. I even talked about it a little bit in my post last week when I mentioned that no one is ever overvalued so you never get to take advantage of someone else's TrueSkill being inflated. This is the big problem! New players start off super-deflated and everyone slowly works their way towards where they should be. Two people who are where they should be can play just fine and have positive metagame EV. The problem comes because you sometimes face someone deflated and never face anyone inflated so if you keep playing random people you get hurt. The core issue is that the new ranking system compares TrueSkill values instead of TrueSkill means. I'm going to go into some details about the two and why I strongly feel the wrong one is being used.

I intend to post this on their forums tomorrow, assuming no one finds any flaws in my reasoning between now and then...


The basic idea behind TrueSkill is some engineers at MicroSoft wanted to find a way to boil down how good someone was at Halo to a single number. The reason they wanted to do this was threefold: better matchmaking in online games, bragging rights among the players, and so they could award invites/byes to big events. It's a great idea, and not really anything new since chess and Magic had a similar rating scheme going (Elo rating). The problem with using Elo was it was designed as a two-player system and Halo can have huge free-for-all games. So the MicroSoft guys crunched their numbers and came up with a new system that would work for multiplayer games and which actually gave each person two relevant numbers, not one. Instead of the system giving someone a single definite rating it gives each person a range of ratings where their actual rating is likely to reside. This range is a Gaussian distribution (bell curve) and is defined by an estimated average rating (mean) and how certain the system is that your actual rating is near the mean (sigma or standard deviation).

If you play a game for the first time the system knows nothing at all about you. As such it assigns your TrueSkill mean to be the average of the entire population (an arbitrary number defined by the system, on Yucata that number is 1200) and gives you a huge sigma since it has no clue how good you might actually be. This value is again arbitrarily defined by the system and determines how big the range of ratings will actually be. On Yucata sigma starts at 400. Because the range is a bell curve we know that there's a 99.7% chance that your actual rating will lie within 3 sigmas of your mean. On Yucata this means that 99.7% of people will have an actual rating somewhere between 0 and 2400.

Now, numbers tend to confuse people. I'd expect people who play board games online are less likely to be confused than your average person but they're still confusing. As such, leaderboards don't want to show you two numbers and expect you to understand Gaussian distributions. People want a single number to look at and brag about. People like to see numbers that go up! Having half your population go down from the starting value (which is expected; half the people are below average) isn't cool. So while the TrueSkill system behind the scenes knows about the mean and the sigma you still want a single number that people can look at and see get bigger. As such, for leaderboards, the system is designed to spit out the very lowest value of that 99.7% range. In formula form, TrueSkill = mean-3*sigma.

Note on Yucata this value for a new player to a game will be 0. And will tend to get bigger, even if someone is bad at a game, because the system will become more and more sure of their actual rating as they play games. Someone with a real rating of 600 could expect to see a mean of 700 with a sigma of 50, for example. This still gives a TrueSkill number of 550 (we're 99.85% certain they're at least 550) but they're really subpar at the game.

This is fantastic for getting people to keep playing the game. They're bad, but they won't get discouraged by having their shown TrueSkill value keep sinking. Instead it will actually creep upwards as the system hones in on just where they should be. It's just fine for leaderboards as well since in order to get a really high value you need to play a lot of games (to shrink your sigma) and win a large percentage of them (to grow your mean). I really like it for both of those reasons.

It's really important to note that the formulas used by the TrueSkill system behind the scenes don't use this displayed TrueSkill number for any reason. After a game ends it uses the old means and sigmas of the players to compute the new means and sigmas. Those are the values that actually mean something as far as how good someone is. Knowing what their extreme lower bound is doesn't actually help a whole lot. Comparing two extreme lower bounds is really of questionable use.

It's this piece of information that I think is key to my problem with the new metagame ranking system. In order to determine how many metagame points are earned after a game the system compares the leaderboard TrueSkill number instead of the TrueSkill mean. This sticks three times each person's sigma into the mix and a new player's sigma is so massive it dominates the entire formula. My sigma in Roll Through The Ages, for example, is 18. A new player's is 400. That's a difference of 1146! Let's look at the difference of the two different comparisons:

TrueSkill mean:

u1-u2 = 1394-1200 = 194

TrueSkill leaderboard value:
(u1-3*s1)-(u2-3*s2)=(1394-3*18)-(1200-3*400)=194+1146=1340

Do note that in the first formula we're not very confident in the 1200 and that's not being taken into account at all. We're assuming the new guy is average when it comes to how he'll impact the metagame ranking. In the second formula we're actually assuming he's one of the worst players on the planet. 99.85% of people rate to be better than this guy!

If the first formula was used sometimes we'd punish the new guy. Sometimes we'd give him a boost. And his opponent in that game would sometimes get a boost and would sometimes get punished. In fact, since we're starting with an assumption he's exactly average, these two will cancel out in the long run. Short term there will be some fluctuations but those sorts of things are expected in this sort of system.

If the second formula was used we'd almost always boost the new guy. Only one person in 667 is actually worse than we're assuming here. The other 666 people are getting a big boost from this formula. And by extension, the people they play against are getting a big penalty.

I believe the new metagame system is quite reasonable when two people with established rating play against each other. It isn't a coincidence that the two formula above actually approach each other in that situation. bk375 (the #4 guy on the RTTA leaderboard and someone who has started a new account) only has a sigma of 15. Compare me to him using the two formula again:

TrueSkill mean:

u1-u2 = 1394-1418 = -24

TrueSkill leaderboard value:

(u1-3*s1)-(u2-3*s2)=(1394-3*18)-(1418-3*15)=-24-9=-33

Comparing me to the new guy resulted in 194 vs 1340 or a 590% increase. Comparing me to bk375 resulted in -24 vs -33 or a 38% increase. That's a fantastically huge disparity!


My question is, why does the system use TrueSkill leaderboard value instead of TrueSkill mean value? Is there something I'm missing that makes one want to benefit people with a high sigma? (And, by extension, punish those with low sigma who play against them.) I was trying to figure out why starting a new account felt so appealing and it's this one issue that's why.

Wednesday, December 07, 2011

Help With Correlations

I've been playing a lot of Roll Through The Ages on Yucata recently. It's a short dice game with a few reasonable paths to victory. I've cycled through some different plans as I've played games with reasonable success but I don't know what the optimal plan is. When I get beat I try to figure out what I did wrong and adapt my play in the future but I'm not sure what really works.

One of the neat things with Yucata is you can replay any game you've played in the past. I was thinking it would be useful to build a spreadsheet or database with the key parts of every game I've played and then crunch some numbers to figure out which choices correlate with winning.

One of the things you do in the game is spend turns building infrastructure for better future turns, but the game is really short. When is it right to shift from building more cities into scoring points? At 5 cities? 6? 7? Does it depend on what your opponent is doing?

If you had the choice should you choose to go first or second? There are definite advantages to both. Personally I like going second but I can't justify that stance. (Some aspects of the game are first come, first served so going first is good. But the end of the game is variable and the person who goes second can often end the game if they'll win or extend it one more turn if they won't.)

How about techs? Quarrying into engineering seems good. Quarrying into empire seems good. Agriculture and masonry seem good. But which wins more often? Does it matter what the opponent is doing?

If your opener is 4 goods and 3 food should you buy a 10 cost tech or save up? What if your opener is 2 goods, 3 food, and a coin?


I know from my sports nerdery that this sort of analysis is possible. (Did you know in the NFL that defensive penalties have no correlation with winning percentage? Offensive penalties on the other hand are negatively correlated, especially false start penalties.) But I don't know how to do it. I'm sure I could have taken some third or fourth year stats course that would have taught this sort of thing but I didn't take enough stats courses it seems...

So, my question to anyone reading this... Anyone know of any books I could buy that would help out?

Tuesday, January 11, 2011

Testing Profession Skill-Ups

When I was preparing to work on the Realm First: Illustrious Cooking achievement I found a formula to work out the odds of getting a skill-up depending on when a recipe turned yellow, when it turned grey, and current skill level. I didn't find any proof of this formula anywhere or anyone really giving any justification for it. I found it in a few spots and ran with it, but I needed far less than I thought I would to max cooking. Sky needed far fewer metagems than he thought he would to max jewelcrafting, too. Did we just get a little lucky or is the formula wrong?

I want to find out by testing the formula to see if it stands up for at least one recipe. The issue then becomes that I need a lot of repetition at the same skill level to have any confidence in estimating the true odds of a skill-up at that skill level. My first thought was to just unlearn cooking over and over since it has a recipe at skill 1 that uses only vendor materials. Unfortunately it turns out you can't unlearn secondary professions, so that won't work. I could delete the character every time but you need to be level 5 to learn cooking so I'd need to level the new character every time in addition to having to transfer funds to buy the cheap vendor mats. Alternatively I could run tailoring off linen cloth or engineering off rough stone if I could gather up enough stuff. Tailoring has the advantage that I could make actual greens to disenchant to mitigate some of the cost.

How much stuff is enough stuff? The answer to that question depends on how confident I want to be. If I was testing a single event (like a die or a coin) I'd probably want to be sure within half a percent 99% of the time. Now, n = Z^2/(4E^2) where Z is pulled from a table of values for normal distributions and E is the maximum error I'd accept so I'd need to flip that coin n = (2.5759)^2/(4*(.005)^2) = 66353. That's a very big number. There's an added wrinkle for me as well in that I'm not testing a single event, I'm testing like 60 of them. There's a different probability at every skill level, so I'm estimating the odds of 60 numbers, not 1 number. To maintain that level of confidence would require 133k linen cloth per skill level. Even worse, the first few yellow points will give me a skill-up right away so getting 66k trials on them requires powering the first 24 orange ticks almost that many times as well.

What I can do, though, is relax my confidence requirements. Bolt of Linen Cloth has 25 skills between yellow and grey, so I'm expecting a 4% drop between each skill level. So if I set my maximum error to 2% I should still see the gradation between different skill levels. With 25 numbers here I should see if the formula is way wrong or not as long as most of my estimates are good, so I don't need the 99% confidence either. 95.45% (two standard deviations) seems like it should be more than good enough. With those new numbers n = 2500. Ugh. Nevermind the fact I cant force the same number of trials for each skill level, plowing 5k linen into each of the 25 skill points is absurd. We're talking a third of the linen needed to open the gates of Ahn'Qiraj just to run this test. That's not going to happen.

What if I assume I want the last skill point to have the 2500 iterations? Then the first one only gets 100 iterations. That really hurts my confidence in the first few goes but does bring my needed linen down to 70k. Is this something I can do or should I just give up? I do also get to make something from the bolts, so I get to run half again as many trials on Brown Linen Pants. 70k seems semi-reasonable but not something I'm going to farm myself. I think I'll just start buying cheap linen on the AH for a while and see how many I end up with.

I can't start now anyway, as I need a way to collect the data. The easiest way to do this has to be writing a LUA mod to output information each time something is crafted. Does the language have access to the right information? Only one way to find out, and that's to actually learn to write a WoW mod.

Sunday, December 05, 2010

Profession Skill Up Chance

I've acquired 50+ of each of my chosen meats for the upcoming push for realm first: illustrious cooking. Do I need more? Do I want more? How safe am I, really?

Those are the questions I was asking today and turned to google to see if anyone actually knew the answer. I couldn't find any links to data from any actual studies but I did find in a few spots the same formula which makes some intuitive sense. If the formula is right then it also explains the weirdness in terms of when the different cooking recipes turned green according to tier.

Orange is always 100% to skill-up (other than skinning which Blizzard intentionally changed a long time ago). Each recipe now has two relevant numbers. The last skill level which is 100% to skill-up and the first skill level with is 0% to skill-up. Every level between those two has a chance of providing a skill-up which is linearly related to those two numbers. The formula is (grey skill - current skill) / (grey skill - yellow skill + 1). The green level is artificial and represents a less than 50% chance to skill-up, and hence lies halfway between the grey skill and the yellow skill. (This is why I can get 10 yellow ticks out of a tier 1 food and only 5 from a tier 3 food. They have a wider range before they go grey.)

Anyone who bothered to read my two part series on Markov chains should see where I'm headed here. Skilling up is strictly increasing and I now have quantified probabilities for advancing from state to state. How much food do I need? I can build a matrix and find out! (Well, I can build 3 matrices and find out since they all have different ranges between yellow and grey.) Each of the three tiers has 15 points in orange, so for each tier I need 15 meat plus 10 skill-ups.

Building up my matrix and running it on the online tool I found shows that for tier 1 I expect to need 12.941 meat to get those final 10 points, for a total of 28 meat. The 64 snake eyes I have should be more than sufficient. In fact, it's a practical guarantee. The odds of not getting 25 skill points with 64 meat is 0. Experimenting a bit, the 95% threshold is 32 pieces of meat.

For tier 2 we're looking at needing 16.5583 extra pieces of meat. The 95% threshold (including the orange 15) is 38. I have 57 blood shrimp which is 99.9988% likely to get me to 500.

For tier 3 we're looking at needing 32.2187 extra pieces of meat. The 95% threshold (again including the orange 15) is 71. I have 59 crocolisk tails which is 85.55% likely to get me to 525.

I'm happy with my practical guarantee to get through tier 1. Missing 1 of every 100000 tries at passing tier 2 is acceptable. But failing 1 in 7 at getting through tier 3 is not something I'm happy with. Fortunately for me I have an easy out. Vek has had enough double token days that I can afford to buy a 4th recipe. I was hoping I wouldn't need to but if I do then I have 114 combined tier 3 meats which would give me a 99.92% to make it. I was thinking I should save the tokens in case I needed to buy a tier 2 recipe but this shows me that's wrong. Even with 2 recipes I'm still more likely to fail at tier 3 than at tier 2. I should be good and I will absolutely make sure to have that bonus skill from a daily sitting around in case disaster strikes. (And beg for more food!)

On the plus side, I doubt other people are thinking they need to stock up hundreds of the tier 3 meat so if I need to make that run other people might have to as well. Better make sure to log out in frost spec for the extra mounted run speed!

Saturday, October 30, 2010

San Juan: Gold Mine Math

When I was writing up the post on Prospector related cards I wanted to quantify just how good the Gold Mine was but I didn't feel like I had time to really get the 'right' answer so I floated out a good enough answer by just ignoring the fact that cards were in play and that once a card was flipped up it couldn't be flipped up again. (A different copy of the same card, sure, but not the same card itself.) So I just worked out the odds of all different from a 110 card deck with no replacement and got:

  • 24% chance of getting any card
  • 15% chance of getting any 5
  • 11% chance of getting any 6
  • 20% chance of getting a 5 or a 6
  • 85% of the time when you get anything, you get a 5 or a 6

The way I obtained these numbers was to build a spreadsheet with every possible combination of the numbers 1 through 6. This gave me 1296 rows. I also counted how many of each cost exist in a fresh deck and used a vlookup command to bring those values over. Then it was a simple matter of multiplying the odds for each cost on each row to determine how many of those specific combinations existed. Sum all those numbers up to get 146,410,000 which is 110^4. Then it was a simple matter to write a couple formulae checking to see if that row matched any of my above criteria and summing the rows which did and dividing by the total number.

Now, part of my knows there's some way to do this with nPrs and nCrs and factorials. Knowing that such things exist is what made me think it would be hard/fun to work out preciser values taking account of the fact we don't have replacement and that order doesn't matter. However I had an idea while I was trying to go to sleep last night which I tried this morning and it seemed to have worked, giving the following numbers:

  • 25% chance of getting any card
  • 16% chance of getting any 5
  • 11% chance of getting any 6
  • 21% chance of getting a 5 or a 6
  • 86% of the time when you get anything, you get a 5 or a 6
The first thing I did was subtract 3 from my number of 1s in the deck. Each player starts with an Indigo and you've already built a Gold Mine. I'm ignoring any other possible buildings though if you cared about the odds from a given game state you could take that into account. For the general case this should work. I then slightly modified my vlookup formulae. Without loss of generality I assumed the card in the first column was pulled first. The second column then checked to see if it was pulling the same value as the first or not. If it was then I subtracted 1 from the count of such cards in the deck. The third column checked the first 2 for doubles. The fourth column checked all others for doubles.

The effect of doing this on the numbers is it didn't reduce the quantity of any of the good rows (they already had no duplicates so they could never have pulled the same card twice with replacement) but did decrease the chance of a given 'bad' row from occurring. (For example, the old table said pulling 4 6s at once happened 4096 times while the new one says 1680.) Doing this decreases our overall sum of possible outcomes down to 123,854,640. (107*106*105*104) Then it was a simple matter of redividing by the new denominator to get the new odds. I made the two changes, of course, but if I'd just made the replacement change then the odds of everything good happening would have gone up by 110*110*110/(109*108*107) or up by about 5.7%. 

What does this mean for how good Gold Mine actually is? Well, it'll hit a quarter of the time, so if you build it early you'll probably get 4 cards out of it. Of those 4 cards you can expect 3 of them to be very good (but possibly duplicating previous good pulls). You have no way to influence if you'll get those cards on the first pulls, or on the last ones, or every time, or never. But the cost is actually pretty small to build it and you have actually a pretty reasonable chance of winning the game from it as a result. Low risk, high reward? I think you have to go for that, sadly. Is it worth delaying a Library to build it? If you think your opponent is better than you then I think you have to. If you don't then maybe not.

Sunday, October 24, 2010

Markov Chains 2

Yesterday I investigated using Markov Chains to work out how long the average person would take to get the A Mask for All Occasions achievement in World of Warcraft. However, in order to get the problem into a simple Markov Chain format I had to assume the player was only getting masks by Trick or Treating which is not the only way of getting them. On top of the 16% chance you have of getting a mask by trick or treating you also have a 14% chance of getting one by defending a town from the Headless Horseman and (as of Wednesday) what seems to be a 100% chance of getting one from the Headless Horseman daily dungeon.

To save the casual reader from Wall of Math, details are after the jump.

Saturday, October 23, 2010

Markov Chains

Hallowe'en is approaching which means the Hallow's End event is up in World of Warcraft. One of the achievements for this event is to collect 20 different masks (one for each race/gender combination in the game) and the odds of actually getting it done are pretty absurd as Sthenno posted about here. I have no reason to believe his numbers are anything but correct but I've heard people talk about Markov chains in the past, think they could be used to generate the odds, and want to learn about them so I'm going to crunch the numbers myself for fun.

Unfortunately I don't have any mathematics software so I'm stuck with Excel, my brain, and brute force for now. Maybe I'll try to find a trial version of Maple or Matlab or something when I get home. So I'm going to have to start small, which is good, since it should hopefully make things easier to understand. This is going to be a pretty big wall of crunching the numbers text, so I've added in a jump break here. Jump on if you don't value your sanity!

Monday, June 28, 2010

That's Combinatorics



I had some people over after the Serenity movie on Saturday to play board games. I learned two new games and played a third I’d only played once before (Louis XIV, Endeavor, Ra:Dice Game). All three were games I enjoyed and would like to play again. Louis XIV had an interesting ending that sparked a bit of debate due to a hidden scoring mechanic in the game that I thought was a little interesting. I ended up writing a script to simulate the ending which I ran about half a million times to crunch some numbers about how the game ‘should’ end from the game state we got to.

Without going into the specifics of the game itself, know that there’s only three way to score points in the game. The first is accomplishing a task and is worth 5 points. The second is to collect tokens each worth 1 point. (These tokens have pictures on them but you get them blindly so you have no control over the pictures.) The third is to have the most of a given picture of token. Each picture that you have the most of is worth 1 bonus point. The first two sources are easy to calculate, the question becomes how many of the bonus points do you expect to get?

Now, the token pile is split into 6 pictures, and there are 10 of each picture. We played a three player game and ended up with a 9-11-13 split. The debate broke down into how reasonable is it for a given player in that position to score up enough bonus points to win. I was the 9 and had completed an extra task over the course of the game, so I had an extra 5 points. Pounder had the 11 and Duncan had the 13. Since each token is worth 1 and I had 5 extra points I was in the lead going into the bonus phase. I was up by 1 point on Duncan and 3 on Pounder. The bonus points ended up being split 2-2-2 (ties go to no one) so I won, but ‘should’ I have won? How reasonable would it have been for Pounder to come back from down 3 to pass me? (This came about from a debate about who to ‘screw’ if you have the choice.) Assuming no tiebreakers, the approximate odds from my simulation for each result are:

Nick Wins 24.06%
Pounder Wins 0.02%
Duncan Wins 39.50%
Nick-Duncan Tie 35.88%
Nick-Pounder Tie 0.03%
Pounder-Duncan Tie 0.02%
Nick-Pounder-Duncan Tie 0.50%

Turns out being down 3 with the 11 of a 9-11-13 split is a very bad spot to be. Pounder has practically no chance of outright winning and not much better of pulling off a tie. Even if he wins all tiebreakers (in the specifics of the game he won tiebreakers, then I beat Duncan) he doesn’t even win 1 game in 100. I outright win 24% of the time which is pretty reasonable, so despite it being a little favourable to have gotten a 2-2-2 split it’s not like I stole the win. When you consider I had the tiebreaker on Duncan my win chance shoots up to almost 60% which is pretty reasonable. And indicates a very close game took place. If Pounder had some way to give Duncan or I a point during the game it would have had a pretty big impact on the outcome of the game. (All those ties turn into straight wins for Duncan and some number of my wins turn into ties.)

A question then is how many points is a chit worth? It’s worth 1 plus the amount of bonus you earn divided by the number of them it took to get there. In the specific game my chits were worth more because it only took 9 of them to get the same number of bonus points Duncan got with 13. But how much are they worth in general? This is a question I should be able to answer with combinatorics but it’s been a good 10 years since I took an enumeration course and I just can’t work it out right now. (Aidan said he was going to try!) I should probably get a book and refresh myself on it. (And on stats, so I could work out confidence intervals for my simulations instead of just assuming they’re good enough…) But for the specific breakdown of 9-11-13 I can say that the 9 expects to get .86 bonus points, the 11 expects to get 1.4 bonus points and the 13 expects to get 2.1 bonus points. (With 1.6 wasted to ties.) On a per tile basis, my chits were worth 1.09, Pounder’s were worth 1.13, and Duncan’s were worth 1.16. The more you get the more they’re worth, but they still aren’t worth very much extra.

The final question then becomes: “Is this a good mechanic”? (The initial debate on Saturday started with this very question.) My viewpoint hasn’t changed given the numbers but I’m curious if anyone else has an opinion on the matter before I go into details.