Showing posts with label analytics. Show all posts
Showing posts with label analytics. Show all posts

Monday, April 14, 2014

Rough Translations: Is the traditional OBP formula the most suitable?

Thanks to a recent Sports Illustrated article and Baseball Prospectus interview, I stumbled across the website for the Grupo Independiente para la Investigacion del Beisbol (GIIB), a group interested in applying sabermetric principles to Cuban baseball. Their website is very interesting, but it's in Spanish. In order to spread their very useful approach, I'm putting my loose translation here, in the hopes that (even if it's not 100% accurate) it will at least be an improvement over what you can get for free through something like Google Translate.

DISCLAIMER

  • All content is property of the GIIB and is not my own. I claim zero rights to it. If they get mad about this translation, they just have to contact me and I will absolutely take it down.
  • I also claim zero responsibility for the accuracy of these translations. I am not a native Spanish speaker, but I did take Spanish in high school and am currently a level-17 Duolingo user, for whatever that's worth. Any missing Spanish knowledge (of which there is a lot) will be supplied by Google Translate.
  • Because my understanding of the original is limited, the translations will probably not be exact. I hope to at least capture the spirit of the article, so others who don't know Spanish can still read the GIIB's research. Sentences or phrases I can't get a good handle on will be denoted by italics. You're welcome to leave corrections or other constructive feedback in the comments.

***

¿Es la fórmula tradicional del OBP la más idónea? (Is the traditional OBP formula the most suitable?)

Is the OBP formula the most correct way to measure the probability that a batter reaches base? Is not the sacrifice bunt an opportunity to get on base? Why does the OBP formula ignore when a batter reaches base on an error?

OBP = (H+BB+HBP)/(AB+BB+HBP+SF)

H: Hits
BB: Base on Balls
HBP: Hit by Pitch
AB: At Bats
SF: Sacrifice Flies

The formula for OBP is too focused on the analysis of the individual hitter and not what he really contributes to the team. One piece of evidence for this last statement is the fact that the OBP formula excludes the sacrifice bunt from the denominator. Suppose a batter comes up with runners on first and second with no outs, and grounds out to second, allowing the runners to advance. What is the difference between this at bat and a sacrifice bunt? Are the two actions not worth the same to their team? Then why does OBP make a distinction between them? Defenders of OBP argue that, because the sacrifice bunt is ordered by the bench, it should not be seen as an opprotunity to get on base and therefore should be excluded from the denominator. But is this true? Are all sacrifice bunts ordered? Should OBP distinguish between sacrifice bunts put on by the manager and other sacrifices? Of course note; and even if we consider all sacrifices as ordered from the bench, the sacrifice isn't an opportunity to reach base? Really?

Let's return to the previous situation: runners on first and second, no outs. Suppose the batter bunts the ball and reaches first safely, getting credit for a hit; is this bunt not an opportunity to get on base? In other words, if the bunt goes for a hit, it is a positive action, but if it advances the runners (which could also be accomplished by other means as shown above) and the batter is thrown out at first, it doesn't count as an opportunity to get on base? This is incongruous.

Now we analyze another important aspect of OBP: the exclusion of times the batter reached on an error.

Suppose a batter reaches on an error. The batter accomplished one of his goals (to get on base), and has become a runner and an opportunity for his team to score. Any runner represents a great opportunity for a team to manufacture runs, regardless of whether he has reached by error, hit, or walk. But the basic OBP formula doesn't see it this way. According to the basic OBP formula, reaching on an error is a negative action. When a batter reaches first by an error, according to this formula he is not credited with reaching first safely but is credited with an at bat. In other words, the times when a batter reaches on an error, it is counted as a failed opportunity to reach first. This is difficult to understand.

Let's analyze this from a different perspective. For some, the error has nothing to do with the offensive player; it is true that the error is a bad defensive acction, but does this mean the batter has no influence? What is the difference between a hit and an error? Subjective concepts; first, the positioning of the defensive player, and then if the scorer considered there to be an opportunity for an out. Is a Texas leaguer to center field worth more than to connect on a shot to third that the fielder can't handle? In general, baseball rewards speed and placement and not force, but is this really important? Imagine your favorite team, losing by one in the ninth, with two outs and the tying run on third. The batter reaches on an error and the game is tied. Does it really matter how the game was tied? Would you rather lose?

It is true that errors are sometimes just bad defensive plays, where there is not a hard-hit ball or fast players to rush the defense and force them to make risky throws. But is the HBP not an error by the pitcher? The HBP is in most cases a mistake by the pitcher, a pitch that gets away from him; but it is nevertheless counted in the traditional OBP formula as a positive offensive action. It takes a lot for us to understand how a HBP is worth more for the batter than to reach on an error.

We now return to the initial question: is the traditional OBP formula the best way to measure the probability of a batter reaching base? Briefly, OBP does not count when a batter reaches on an error and doesn't consider a sacrifice bunt as an opportunity to get on base. We therefore calculate OBP in an alternative way and call it gOBP.

gOBP = (H+BB+HBP+ROE)/(AB+BB+HBP+SF+SAC)

ROE: Reached on error [Trans. note: abbreviated EE in original]
SAC: Sacrifice hit

To capture the reality of the concepts we must move away from the moralistic way of thinking that still exists in baseball analysis. If we want to know the real probability that a batter reaches base, then we have to count all opportunities and all successful actions. Each turn at bat is an opportunity to reach base; and of course, every time the batter reaches first safely, it is a positive result for him. We exclude only interference or obstruction from this analysis, because the probabilities of these actions occuring in a baseball game is around 0.0027 [~0.27 percent]. That is to say, it is such a small sample that it is negligible.

For now, enough philosophical discussions about baseball. We will concentrate on mathematical tools that allow us to demonstrate which of these two statistics is more useful when analyzing the offensive prowess of a team.

For this we compiled statistics from the last 15 Cuban National Series. We calculated both variants of OBP and found the linear correlation of both with runs scored per game.


This is the scatterplot of the traditional OBP vs. runs per game. These statistics show a linear correlation of 0.93. But we are not questioning the proven utility of the classic formula, but we are analyzing which formula is the most suitable.

We next show the scatterplot of gOBP with respect to runs per game.


These statistics show a linear correlation of 0.95.

Although the difference is not very high, it is not possible to deny that gOBP is slightly closer to reality than OBP, at least numerically. We remember that correlation does not explain causation between variables. But we consider gOBP the optimal indicator to measure the probability that a batter will reach base. This is why STRIKE, as well as other projects that derive from it including StatsPlay, includes gOBP in its reports and recommends it over OBP to measure this concept.

Regards, friends.

Rough Translations: Analyzing Industriales

Thanks to a recent Sports Illustrated article and Baseball Prospectus interview, I stumbled across the website for the Grupo Independiente para la Investigacion del Beisbol (GIIB), a group interested in applying sabermetric principles to Cuban baseball. Their website is very interesting, but it's in Spanish. In order to spread their very useful approach, I'm putting my loose translation here, in the hopes that (even if it's not 100% accurate) it will at least be an improvement over what you can get for free through something like Google Translate.

DISCLAIMER

  • All content is property of the GIIB and is not my own. I claim zero rights to it. If they get mad about this translation, they just have to contact me and I will absolutely take it down.
  • I also claim zero responsibility for the accuracy of these translations. I am not a native Spanish speaker, but I did take Spanish in high school and am currently a level-17 Duolingo user, for whatever that's worth. Any missing Spanish knowledge (of which there is a lot) will be supplied by Google Translate.
  • Because my understanding of the original is limited, the translations will probably not be exact. I hope to at least capture the spirit of the article, so others who don't know Spanish can still read the GIIB's research. Sentences or phrases I can't get a good handle on will be denoted by italics. You're welcome to leave corrections or other constructive feedback in the comments.

***

Analyzing Industriales (Analizando a Industriales)

Via the attachment that this article contains and you can download [Trans. note: available through thegiib.com.], the first preseason analysis that our group performed for Industriales is made public (dated July 2013). In this analysis, very controversial issues surrounding the blue team are touched on. Below is a summary of what you will find; we hope you enjoy it:

- Regular players?
- Short field?
- Catchers?
- Yulieski or Rudy at 3B?
- How to replace Odrisamer?
- Predictions?
- Some players and their characteristics
- Advice
- Graph of the principal Industriales players' wOBA over time.

Rough Translations: Announcement

Thanks to a recent Sports Illustrated article and Baseball Prospectus interview, I stumbled across the website for the Grupo Independiente para la Investigacion del Beisbol (GIIB), a group interested in applying sabermetric principles to Cuban baseball. Their website is very interesting, but it's in Spanish. In order to spread their very useful approach, I'm putting my loose translation here, in the hopes that (even if it's not 100% accurate) it will at least be an improvement over what you can get for free through something like Google Translate.

DISCLAIMER

  • All content is property of the GIIB and is not my own. I claim zero rights to it. If they get mad about this translation, they just have to contact me and I will absolutely take it down.
  • I also claim zero responsibility for the accuracy of these translations. I am not a native Spanish speaker, but I did take Spanish in high school and am currently a level-17 Duolingo user, for whatever that's worth. Any missing Spanish knowledge (of which there is a lot) will be supplied by Google Translate.
  • Because my understanding of the original is limited, the translations will probably not be exact. I hope to at least capture the spirit of the article, so others who don't know Spanish can still read the GIIB's research. Sentences or phrases I can't get a good handle on will be denoted by italics. You're welcome to leave corrections or other constructive feedback in the comments.

***

Convocatoria (Announcement)

This website is the official way the StatsPlay application is distributed, along with the necessary databases and accessories. We believe that this is the best way to make the project available to everyone, without borders, and with the most independence possible. We regret that many users in our country can't download all that they need for the project to work for them. We know the limited connectivity options and Internet access that exist. Nevertheless, many people do have access to a variety of sites on the national network, such as through the intranets of their workplaces.

For that reason:

I call

on all sites hosted on servers on the national territory under the .cu domain, especially Cubadebate, Granma, Juventud Rebelde, Infomed, and BeisbolCubano, to contribute to the dissemination of this project. All national sites have total freedom to host and allow the downloads of everything related to the StatsPlay applications and databases from their servers. We hope that this way, this project will have a very broad scope, and that a greater number of people can enjoy it and use it.

-Camilo Quintas, leader of the GIIB.

Rough Translations: Introduction/The GIIB Web Site

Thanks to a recent Sports Illustrated article and Baseball Prospectus interview, I stumbled across the website for the Grupo Independiente para la Investigacion del Beisbol (GIIB), a group interested in applying sabermetric principles to Cuban baseball. Their website is very interesting, but it's in Spanish. In order to spread their very useful approach, I'm putting my loose translation here, in the hopes that (even if it's not 100% accurate) it will at least be an improvement over what you can get for free through something like Google Translate.

DISCLAIMER

  • All content is property of the GIIB and is not my own. I claim zero rights to it. If they get mad about this translation, they just have to contact me and I will absolutely take it down.
  • I also claim zero responsibility for the accuracy of these translations. I am not a native Spanish speaker, but I did take Spanish in high school and am currently a level-17 Duolingo user, for whatever that's worth. Any missing Spanish knowledge (of which there is a lot) will be supplied by Google Translate.
  • Because my understanding of the original is limited, the translations will probably not be exact. I hope to at least capture the spirit of the article, so others who don't know Spanish can still read the GIIB's research. Sentences or phrases I can't get a good handle on will be denoted by italics. You're welcome to leave corrections or other constructive feedback in the comments.

***

Sitio Web del GIIB (The GIIB Web Site)

With great joy we launch our group's official website today. We want to thank everyone who has kindly offered us their help. Although it is not necessary, it is recommended that all users register to be able to consume all of our site's services without difficulty. We believe that this is the best way to make our first proposal public and official: the launch of the StatsPlay project. In this site you will find everything you need to inform yourself about StatsPlay. We hope that with this project, all those linked to baseball can supplement the existing information needed in our country.

Welcome to all. GIIB.

Monday, December 30, 2013

Site News: Movin' on up in 2014!

Last year, I set a New Year's resolution to do more analytics work, including 60 blog posts. I got through 15.

But it's not all bad! Two big pieces of news for 2014:
  • I will be attending the Sloan Sports Analytics Conference again this year.  I submitted an abstract to the research paper competition, which was accepted.  Unfortunately, the results of my research disproved my hypothesis, and the whole thing's come crashing down.  Ordinarily, that means you would see it repurposed here as a blog post but...
  • I've been hired as a contributing writer to Beyond the Box Score, "a saber-slanted baseball community", where I will be writing articles on a regular basis.
I'll keep this blog open for non-baseball stuff, but most of my writing will appear over there.

Best wishes to all my reader(s) for a happy and healthy 2014!

Friday, December 27, 2013

Sorting Through a Million Bags of M&Ms

As a kid, I used to sort bags of M&Ms by color. This was my first sort of data science project, and my parents' first clue that this one was a little off. Every once in awhile, I'll revert to that habit (especially with those fun-size bags you get around Halloween), which led to a long, drawn-out discussion with a friend about the probability of getting a fun-size bag of Skittles with no purples.


So when a recent trip to the vending machine produced a free bag of M&Ms, I found myself asking a number of questions about the distribution of the different colors in a bag of M&Ms. A quick Google search produced no official statement from the company, except this one from 2008:


Our color blends were selected by conducting consumer preference tests, which indicate the assortment of colors that pleased the greatest number of people and created the most attractive overall effect.

On average, our mix of colors for M&M'S MILK CHOCOLATE CANDIES is 24% cyan blue, 20% orange, 16% green, 14% bright yellow, 13% red, 13% brown.

Each large production batch is blended to those ratios and mixed thoroughly. However, since the individual packages are filled by weight on high-speed equipment, and not by count, it is possible to have an unusual color distribution.

Well, we have two bags of M&Ms here. We can check to see whether these proportions are still accurate using a chi-square goodness of fit test. Since there are six colors, we will be looking at a distribution with five degrees of freedom.

These tables show the result of the chi-square calculation for each bag.

BAG 1 Red Orange Yellow Green Blue Brown Total
Observed 8 12 4 10 12 8 54
Expected 7.02 10.8 7.56 8.64 12.96 7.02 54
(O-E)^2/E 0.137 0.133 1.676 0.214 0.071 0.137 2.369

BAG 2 Red Orange Yellow Green Blue Brown Total
Observed 8 7 2 14 12 11 54
Expected 7.02 10.8 7.56 8.64 12.96 7.02 54
(O-E)^2/E 0.137 1.337 4.089 3.325 0.071 2.256 11.22

For a 95% confidence interval (alpha = 0.05), x^2 = 11.0705 for a distribution with 5 degrees of freedom. This suggests that we have one normal bag and one outlier. This is not especially conclusive evidence for or against the 2008 distribution, but the significant lack of yellows makes me suspect the distribution has changed.

We can also ask questions about how unusual is a bag with only two M&Ms of a given color. Unfortunately, this is non-trivial to solve theoretically. But we can estimate these probabilities by simulating a large number of bags of M&Ms. I built a MATLAB script (available on request) to simulate an arbitrary number of 1.69-oz M&M bags. For convenience, I assumed each bag had a consistent number of M&Ms (54). I then drew 54 random numbers uniformly distributed on the interval [0,1], splitting up the number line to match the 2008 proprtions (i.e., a random number less than 0.13 meant a red M&M, a number between 0.13 and 0.33 meant orange, and so on). I repeated this process one million times, because "a million bags of M&Ms" sounded cool.


"A million bags of M&Ms isn't cool. You know what's cool? A billion bags of M&Ms."

Shut up, Justin.

Anyway, you end up with this graph showing the distribution for each color. It's discretized, of course, because 0.4 of an M&M is nonsensical. But, thanks to the central limit theorem, all of the distributions are normal. Note that the red and brown curves are essentially right on top of each other*.

* - Apologies to those of you, like my advisor, who are red-green colorblind, and thought M&Ms came in blue, yellow, orange, and a couple shades of brown.


Now that we have this data set, we can answer a whole bunch of questions.

What are the odds that I get a bag with less than N yellows?
For this, we can make cumulative distribution functions based on that figure. So, if we assume the 2008 distribution is accurate, the probability of getting four or fewer yellows in a bag is approximately 11%. And if each bag represents an independent sample (which might not be true, depending on the manufacturing process), the probability of getting two consecutive bags with four or fewer yellows is 1.2%.

What are the odds that I get a bag with less than N of one color?
Here we have to use a different curve. For instance, a bag with only four yellows seems rare from the previous graph, but remember: that just deals with the probability that you have four or fewer yellows, or four or fewer reds. This question deals with the probability you have four or fewer of one color, regardless of which color it is. And now we see that the probability you have a bag with no more than four of one color is about 48%. For two consecutive bags (assuming independence), the probability is a still-reasonable 23%. So, while getting two bags with a small number of yellows is unusual, getting two bags with a small number of any color is pretty common.


What are the odds that I get a bag missing a color?
I'm sure the process of setting those percentages involves minimizing this possibility: if you were six, and your favorite color was red, you might get upset if you went through a whole bag of M&Ms with no reds. As a result, this is a pretty uncommon occurrence: in my data set, the odds were approximately 1-in-690.

What are the odds that I get a bag that's entirely one color?
This never happened in the million trials I ran. In fact, the greatest number of any single color in one bag was 30 blues (out of 54).

What are the odds that I get a bag with equal numbers of all colors?
You would think this wouldn't be too crazy, but in fact it's very rare. I estimate that the odds are about 1-in-42,000.

What are the odds I get more blues than any other color?
Before I present these results, it's important to note that MATLAB's min and max functions don't deal with ties very well. Ideally, you'd had a function such that a two-way tie would count as 0.5 for each color, and a three-way tie as 1/3, but what actually happens is the left-most column gets 1, and everyone else gets zero. This means that red, orange, and yellow will be skewed a little high, and green, blue, and brown will be skewed a little low. But this is already a 1,000-word entry on candy-coated chocolate, so the min/max functions are good enough for me.

Thursday, August 22, 2013

How Consistent is Fantasy Football Consistency?

The end of summer means the imminent start of the NFL season. And while players prepare with grueling workouts in 100-degree heat, fans are preparing by spending hours staring at fantasy football preview magazines and webpages and cheat sheets.

The problem facing the fantasy football player is one of prediction: which statistics from the previous season best predict value in the upcoming season. One such measure of performance is a player's consistency -- the variation in the number of points he scores in a given week. Pro Football Reference has previously shown that good teams should prefer more consistent lineups, while weaker teams should prefer less consistent lineups on the theory that their best chance of winning involves a few "lightning in a bottle" weeks.

But how do you determine which players are consistent?

Tuesday, January 29, 2013

Super Bowl Hype Drive: New Orleans' Bad Mojo

This year marks the Super Bowl's return to New Orleans for the first time since 2002, when the Patriots upset the Rams on a last-second Adam Vinatieri field goal. But the Superdome has a reputation for hosting blowouts, even given the fact that many Super Bowls are one-sided affairs. Is this reputation deserved?

Wednesday, January 16, 2013

How Much Is a Win Worth to an NBA Team?

Last month, I used J.C. Bradbury's free agent valuation method to determine how many wins the Red Sox expected Mike Napoli and Shane Victorino to contribute to the team in 2013. That worked fine, but suppose we want to build a similar model for the NBA. Again, we'll use the basic system Bradbury outlines in "The Baseball Economist" (ch. 13). Here, Bradbury found a relationship between revenue, wins, and the size of the city a franchise plays in.

All three of those variables are readily available. For city size, we'll use the population of the metropolitan statistical area (MSA) each team plays its home games in, as reported in the 2010 U.S. Census*. Revenue is available through Forbes' Business of Basketball listings. This data is almost exactly one year old -- suggesting that it covers the 2010-2011 season, and not the recent lockout-shortened 2011-2012 season. This is better for our purposes; I don't want the compressed schedules and reduced number of games to interfere with my results.

* - And the Canadian equivalent for Toronto, with the hope that the two have very similar methodologies.

Monday, December 10, 2012

Evaluating MLB Signings, Part 2: Madness

Last time out, we asked how good the Napoli and Victorino signings were for the Boston Red Sox. Using J.C. Bradbury's method, we established that we need to do the following:
1. Figure out how much a win is worth,
2. Figure out how much an individual player contributed to his team's wins, and
3. Convert that number of wins into a dollar value.

Thursday, December 6, 2012

Evaluating MLB Signings, Part 1: Methods

After their 2012 season went down in flames, the Boston Red Sox were active during the recent winter meetings, signing 31-year-old first baseman/catcher Mike Napoli to a 3-year, $39 million contract, and 32-year-old outfielder Shane Victorino to a 3-year, $37.5 million contract. The moves were modest when compared to past offseasons, but the question remains: will the Sox get value from their new acquisitions?

Monday, December 3, 2012

A Richly-Deserved Beating: CFB ATS Update

At the beginning of the college football season, I asked whether you could use a team's recent against-the-spread (ATS) history to predict how they would do against the spread this season. I came to the conclusion that
[T]he perpetually underperforming teams (like Tulane) and perpetually overperforming teams (like Boise State) are a function of luck* rather than some underlying market inefficiency.

Well, the season's over (except for the bowl games). Was I right?

Monday, September 17, 2012

Thinking Out Loud: Replacement Referees and Home-Field Advantage

Home-field advantage, like the Cubs' curse and the substandard jumping abilities of Caucasians, is one of those sports truisms that has been accepted for decades as a given. The 2011 book Scorecasting investigated this phenomenon and ventured to explain why home-field advantage still existed in the era of free agency, chartered jets, and five-star hotels. Consider this slideshow, taken from the presentation given by the authors at the 2011 Sloan Sports Analytics Conference:

Friday, September 14, 2012

Bad Beats: The Predictive Power of Past Years' ATS Record

Last week, I used this space to complain about losing money to my friend Dave by betting against Tulane football.

Faaascinating, I know. But it gives me a chance to make an Important Point about the predictive value of statistics.

My confidence in my bet was based on the fact that, from 2003 to 2011, Tulane covered just under 40% of their games against the spread. Winning 60 percent of your bets would make the average professional bettor salivate, so I was happy to bet based on this big trend.

There were two things I ignored: first, that one game is the smallest of sample sizes, and second, that past results are no guarantee of future performance. The second point is the interesting one, so let's focus on that: if a team has done better/worse than average against the spread in the past, does that tell us anything about its performance against the spread in the future?

Friday, September 7, 2012

Bad Beats: Rutgers at Tulane, Sep. 1, 2012

I owe Tulane football an apology. It seems I underestimated this year's team.

When the opening line for the season opener against Rutgers was listed at 17, I immediately jumped on Twitter.
Even when confronted with the one other Tulane football fan on Twitter, I refused to back down.
And it seemed Vegas agreed with me, kind of: by the week of the game, the line had moved to 20, though the over/under still suggested Tulane's score would be a natural number.

Saturday, July 7, 2012

WAR Stars

A long time ago, in a galaxy far, far away, the All-Star Game meant something. You hear this all the time from Werther's-loving, fedora-wearing, chair-rocking, old-timey sportswriters, so it must be true. It's not fair to call this narrative a "dead-horse": this is like some proto-horse ancestor, found primarily in cave paintings, that died out in the last Ice Age.

Pictured: Mariano Rivera's rookie season


These days, the players treat the game as a glorified exhibition game -- which it is -- and have some fun with it -- which they should -- while trying to protect themselves from injury -- which is just smart*, considering the size of their contracts, but for some reason is treated as sacrilege.

* - Ask Ted Williams, who shattered his elbow crashing into a wall in 1950, or Pedro Martinez, who pitched his arm off in 1999.

But I don't want to talk about that.

Wednesday, June 20, 2012

Euro 2012 Tournament Odds

The European Championship pits the best 16* international soccer teams on the continent against each other, producing better matchups than the World Cup and fun Eurozone debt crisis proxy fights. The 16 teams are divided into four groups of four. Each group plays a round robin, with the top two in each group advancing to a single-elimination tournament.

* - Okay, 14 best teams and 2 host teams. But even the host nations are pretty good: There's no North Korea losing games 7-0.

The group stage just finished, and the tournament is about to get underway. Of the eight teams remaining, who has the best chance of winning it all?

For baseball and basketball games, you can calculate the expected win probability for a single game from the two teams' Pythagorean records using the log5 method. For international soccer, you can use the Elo ratings to determine the expected win probabilities. The formula looks like this:
We = 1 / (10(-dr/400) + 1),
where dr is the difference in Elo rankings between the two teams.

Once you know the probability that a team wins its first game, you can compute the probability that team wins its second game using conditional probabilities. Like this:
P(wins first game)*(P(beats potential opponent A)*P(potential opponent A wins first game)+P(beats potential opponent B)*P(potential opponent B wins first game))

Keep doing that and you eventually get win probabilities for the entire tournament:
Elo Semis Finals Champs Odds (x-to-1)
Spain 2110 83% 67% 44% 2.3
Germany 2063 85% 59% 32% 3.1
England 1950 61% 25% 10% 10.1
Portugal 1883 65% 18% 6.1% 16.4
Italy 1871 39% 12% 3.7% 27
France 1839 17% 9% 2.5% 40
Czech Rep. 1779 35% 6% 1.5% 68.3
Greece 1763 15% 4% 0.9% 112

And you're never going to believe this, but the teams with the best ratings are the ones with the highest chances to win. You can see that the probabilities depend a bit on matchups: France has only a 17% chance of knocking off Spain in the first round, but a 2.5% chance of winning the championship, whereas the Czech Republic has a 35% chance of beating Portugal but a 1.5% chance of making the finals, since they'd have a lower win probability against the remaining teams if they did advance.

Comparing the odds to the betting markets, it looks like Spain might actually be a good wager: if you believe this method, their odds to win are 2.3-to-1, but the books have them listed at 2.6-to-1. Practically everyone else is overvalued, but laughably so for Portugal (listed at 7.5-to-1), Italy (9-to-1), and France (12).

UPDATE #1: June 21, 4:45 p.m.
Here's what the odds look like after Portugal's win over the Czech Republic:
Elo Semis Finals Champs Odds (x-to-1)
Spain 2110 83% 64% 41% 2.4
Germany 2063 85% 59% 32% 3.1
Portugal 1901 100% 29% 11% 9.3
England 1950 61% 25% 10% 10.3
Italy 1871 39% 12% 3.6% 27.3
France 1839 17% 7% 2.1% 47.6
Greece 1763 15% 4% 0.9% 117
Note that everyone else's odds change, despite not playing, because of the improvement in Portugal's ranking. Portugal now has an 11% chance of winning the tournament, leapfrogging England, who obviously still has to survive the semifinal matchup. France, a bad bet to begin with, becomes a slightly worse bet now that they have to beat both Iberian Peninsula teams to get to the finals.

UPDATE #2: June 22, 4:45 p.m.
Here's what the odds look like after Germany's demolishing of the Greeks:
Elo Semis Finals Champs Odds (x-to-1)
Germany 2074 100% 70% 39% 2.6
Spain 2110 83% 64% 39% 2.6
Portugal 1901 100% 29% 9.6% 10.5
England 1956 62% 21% 8.4% 11.9
Italy 1871 38% 9% 2.7% 36.9
France 1839 17% 7% 1.8% 54.8
Germany are* technically the favorites by half a percentage point over Spain. But, again, Spain still has a first round match to play; with their odds somewhere in the 11-to-4 range (2.75-to-1), they still look like the best bet.

* - This always looks wrong to me, but everyone else does it. My least favorite thing about soccer.

UPDATE #3: June 25, 9:15 a.m.
Here's what the odds look like after this weekend's games, including the first upset of the knockout stages as the favored England side screwed up some penalty kicks to lose to Italy in a shootout:
Elo Semis Finals Champs Odds (x-to-1)
Spain 2123 100% 78% 49% 2.0
Germany 2074 100% 73% 36% 2.8
Italy 1902 100% 27% 7.6% 13.2
Portugal 1901 100% 22% 7.2% 13.8

This is the last update here, but I'll still post updates on Twitter. Italy's odds are slightly better than Portugal's, just because Germany has a lower rating than Spain's. I do like the fact that every German game from here out (vs. Italy, then vs. Spain/Portugal winner) is another Bailout Bowl, though some of that has to do with the woeful state of the Eurozone.

Thursday, May 31, 2012

Tangled in the Rigging: Defending the NBA Draft Lottery

The NBA conference finals brings with it one of the best sideshows in sports: the NBA draft lottery, in which 14 grown men stand around awkwardly for half an hour to figure out how a bunch of ping pong balls bounced. We*, the viewing audience, are treated to a half-hour special containing some 15 minutes of talking heads speculating wildly, 2 minutes of commisioner David Stern reading franchise names, and 5 minutes of awkward interviews with team representatives. Fascinating.

* - Maybe "we" is the wrong pronoun; I mean, I didn't watch it.

But while the presentation of the lottery may not be especially compelling, the lottery itself sure is. The lottery teams (i.e., those that miss the playoffs) are ranked in inverse order of record, with the worst teams receiving the best chances of a high pick. So the team with the worst record has a 25% chance of winning the lottery, the second-worst team has a 19.9% chance of winning the lottery, and so on down to the 14th-worst team (the last team out of the playoffs) who has a 0.5% chance of winning the lottery. The whole list of probabilities for this year's draft is available here.

Some have argued (with varying degrees of seriousness) that the lottery system is rigged*, and point to the fact that the worst team in the league hasn't won a lottery since the Orlando Magic won and picked Dwight Howard in 2004. But I want to stress this again, because it's important: the team with the highest probability will still lose the lottery (i.e., not get the first overall pick) 75% of the time.

Monday, May 7, 2012

...I mean, uh, ALBERT PUJOLS IS AWESOME AT BASEBALL

"That one's for you, Bryan." (via)

I typed up my most recent article, investigating the causes behind Albert Pujols' home run drought, with a growing sense of dread. Visions of Pujols homering in his first at bat Thursday night hung over my frantic analysis like some baseball bat of Damocles. I watched Thursday and Friday's games with bated breath.

Considering the power Pujols has displayed throughout his career, the fact that he made it until Sunday afternoon was as good as I could have hoped for. Plus, now I get to write an easy follow-up post whining about my miserable luck*.

After the home run, Keith Law tweeted that he couldn't wait for the whiplash-incuding speed with which the narrative on Pujols would change. As a free service to those busy producers of the talking head panel shows on ESPN, we here at The Feats of Strength wish to provide a general recipe for a ready-made Albert Pujols discussion.

Thursday, May 3, 2012

We're So Sorry, Uncle Albert

Albert Pujols would be very happy if we all just forgot about the last month of baseball.

Pujols signed a 10-year, $240M contract over the winter with the Los Angeles Angels of Anaheim*. The statistically-minded balked at the deal, which would pay him $30m in 2021, when Pujols will be 41 years old. Still, it wasn't hard to imagine that the deal could still be worthwhile, so long as Pujols crushed in the first half of the deal.

...Whoops. Pujols, projected by FanGraphs to hit 28 home runs this season, is still waiting for his first homer some 107 plate appearances into the season. That's the longest drought of the season, and naturally, everyone's asking, "What's wrong with The Machine?"