The Signal and the Noise | Nate Silver | Talks at Google
By Talks at Google
Full Transcript
Okay. Well, I guess we should get started. My task is to introduce our speaker. Now, I suppose it's theoretically possible that everybody just wandered into this room by chance, expecting that something interesting might be happening. I don't know. What do you think, Nate? Is that plausible? You could have herding. Right. A few people actually know what they're doing and everyone else kind of follows along? It could be, but I'm guessing that event has a relatively low probability. So I'm guessing that most of the people know who our speaker is and you know his story. So I'm not going to introduce him. I'm just going to turn you over to Nate. Well, thank you. I'll spare you the kind of wedding toast speech, and we'll get to the questions pretty soon. But it's a real honor to be here, to be here at Google, in my view, probably the smartest company in the world right now, about working with data and using it to actually transform people's lives. And I hope that some of what we were able to accomplish. And I tend to use we when. I mean, I sometimes. Right. But I do see it as kind of a collective project that 538 is a part of to make journalism smarter, to make political coverage better. And it just had people be a little bit more data literate and data savvy. And when that kind of hits into the punditry, then interesting things, I guess, can happen sometimes. But we have a lot of great questions, it looks like, and I think we'll leave time. There'll be time for questions at the end from people in the room as well. So let's get started. Absolutely. So we produced all these questions using Dory. So these are questions submitted by members of the audience. The techniques in 538's methodology section were invented ages ago. Someone could have done what you were doing, but at least the 1990s or earlier. Why did it take the world so long to produce and appreciate Nate Silver? I was busy playing poker. Yeah. Sign your mom. My mom. Right. No. If the Congress hadn't basically banned Internet poker, then maybe I'd still be doing that. Right? No. So, in all seriousness, though, it's only been a few election cycles, maybe the last four or five, where you had as many state polls as we do right now, and it's only been more recently than that. We can really collect them pretty easily on the Internet, where you do a Google News search or you steal from Rickler Politics or whatever else. Right. But collecting all that data and have it be essentially free isn't something that's been able to be true for all that long, really. But also, I think people are taking more of a do it yourself mentality toward aggregation and toward producing content. When a news organization, and the Times is no exception, produces a poll, they want you to think it's the only poll in the world. Right. They don't provide the context of, well, this poll says that Romney's ahead, But the other 12 polls say he's not. Right. Which is what a reporter probably should do. So you have, I think, some perverse incentives in news organizations who produce their own polls, which are, which are expensive, and I'm a fan of good, expensive telephone polling to pretend that no one else is doing the same thing. Right. And so I think that's part of it. It kind of cuts against the newsroom's first instinct. And part of it is that this is actually a relatively new thing where you have so much data available for free. So right before the election, the Romney campaign claimed to be confident that they would win. What was wrong with their internal polling? Why were their predictions so far off or were they just keeping a straight face when they knew that defeat was coming? Yeah, I think maybe more. Well, there are different layers of the Romney organization, Right. And one question is, what did Mitt Romney himself believe? And, you know, I'm not sure about that. There's a question of what did Mitt Romney's pollsters really believe? And, you know, it's probably not what they told the press. Exactly. Right. So I think there's some culpability here where the press is fed spin and bullshit, basically. Right. And then they act surprised later on when the bullshit, in fact turns out to be bullshit. Right. And they blame the messenger when they should be blaming themselves. But, you know, look, when campaign pollsters talk to the press, they're not trying to produce accurate information. Right. They're trying to put things in the most favorable light they can. You know, I had very few conversations with the campaigns. I try and avoid that, almost go out of my way to avoid that. But when I talk to the Romney folks, they tended to have a. A pretty realistic view of what the landscape looked like. Maybe it's because when they're talking to me, then I'm not going to go actually report that. It's just kind of a friendly on background conversation. So maybe you get less spin and a little bit more actual insight. But I think it's not hard to tell. It's hard for a campaign to operate under the assumption that it's probably going to lose, then it seems like it's wasting itself. And so there are all kinds of different biases that can creep in from confirmation bias in the way you kind of construct your samples to what news actually gets reported to the candidate. I think any organization, whether it's a political campaign or anything else, if you're an organization that can do a reasonably good or reasonably honest self assessment, that's a big edge. That's why, by the way, consultants make millions of dollars per year. Right. McKinsey or whatever else, they're not really adding a lot of value necessarily. I'm sure they're very smart people. Right. But with the actual work products so much as they provide mediation services to resolve internal conflicts in companies. Right. Where you can do politically incorrect things if you hire McKinsey to go in and say something obvious, basically. Right. So it's kind of the same agency problem you have, I think, in a political campaign where. And by the way, one thing the Obama campaign did in a way I'm not sure about this year, is they had four or five different pollsters that worked independently from one another, at least to some degree. So they would have maybe the same two different pollsters in House survey Ohio. And then if you have a difference of opinion in one candidate, one pollster might show a tie In Ohio, one might have Obama up 5, then you can start to iron that out. And those discussions are often productive. But within a campaign, it can become something quite insular. I think in general, people overrate the value of, oh, the campaign's internal polls, even if you actually saw the real internal polls, which you don't even then. I think even though some of the pollsters are very smart, they're operating under worse incentives than you have for a public pollster. And it's hard to add value relative to simple methods with the extra things you might do. And so if you have bad incentives, then often that that leads to worse overall results, even if you're a more capable person in some abstract sense. Your remark reminds me of the definition of a consultant. Somebody who borrows your watch to tell you what time it is. Exactly. So probability is hard to grasp for the layperson. So what do you think the best way is to visualize the prediction data? You know, I think if you have something which is like geographical data. So the National Weather Service actually has worked on this problem a lot with its hurricane forecast. If you can actually show people the cone of uncertainty, as they call it, then they gradually come to understand what that means. They've also Kind of worked on problems like how much uncertainty do you want to show in that cone? Right. If you show the 90% confidence interval, in essence, it would still be very, very wide. Right. If you show the 50% confidence interval, well, it's kind of useful, but you're going to be outside that range half the time. They kind of settled by trial and error on showing two thirds the possible outcomes. But probability, it is something where you can say all the time that, well, this kind of has an 80% chance of winning. And that's a straightforward kind of objective statement. But people don't really grasp what that means. Exactly. That's why I say, having played poker, it comes more naturally to me because you know that if your opponent has to catch a flush with one card to come, he'll make his hand 20% of the time, you'll win 80% of the time. But believe me, you play plenty of hands and you know, you take those bad beats 20% of the time. Right. You know what it feels like you busted out of tournaments or lost a lot of money on those hands. So you're very aware of what kind of 80, 20 really kind of feels like. I think that's harder for people when they're not encountering probability on a day to day basis. I think in general, the news media has a tough time and maybe people in general have a tough time with everything is either 50, 50 or 100 0, right. It's either, oh, too close to call. Right. Or it's like, we're sure what's going to happen. It'd be a shock if anything else transpired. Right. When usually we're operating on the margins here. Right. And if you guys are engineers, I know some of you probably are, you know, when you're building a program, you're making little incremental improvements. Most of the time, more often than you have this aha moment, you're debugging things and you're making these marginal gains. Right. And you know, that kind of process doesn't necessarily make for great stories. Necessarily. And so that also clashes with the way that campaigns tend to be covered. I just read a review of a new textbook called Introduction to Probability and Statistics using Texas. Yeah, yeah, yeah. True story, true story. So maybe literacy will increase if we get the right material. Well, I know some investment banks and hedge funds are now training their people by having them play Texas hold'. Em. Right. That somehow that attitude. Right. And the instinct to take new information that you perceive and have a better first judgment. I'm a guy who says, oh, don't do the Malcolm Gladwell blink thing. Right. You want to think through problems, think a little bit more slowly. Right. At the same time to have a better first cut when you encounter a new piece of data, a better first instinct for whether that's meaningful or not, whether it should change your summation of a problem or not. That's helpful as well. And poker is a very good game for developing that skill, I think. So as you and your model gain renown and credibility, how will your forecast affect the very events they aim to predict? And how can you adjust for this in your model? Well, I was talking with some people over lunch and this can be a problem that you guys might encounter as well because Google search itself is so influential on what people identify and maybe what trends go viral, then it becomes a very challenging problem potentially. I would hope that we didn't influence people all that much. It worries me maybe a little bit more in something like, in something like a primary campaign where voters are being more tactical if you have to pick between five or six options. There's some evidence, by the way, has worked toward Rick Santorum's benefit in the Iowa caucus where you basically had three interchangeable candidates, more or less Bachmann, Perry and Santorum who were all polling at about 10%. Right. That's kind of inefficient. An inefficient outcome where people who would want Bachman to be win the Iowa caucus for some reason would probably be happy with Santorum as well. So once one candidate had any edge, you saw a lot of tactical shifting. And in this case there was some CNN poll that was published, they had Santorum gaining. It could have been real, it could have been a statistical fluke. Right. But that gave him momentum and then people joined the bandwagon and the surge became a self fulfilling prophecy after a while. In general elections where you have just two parties, you have less of that tactical behavior. So I'm not as worried about it, I don't think. You know, there can be things though, like when you have when we had US House forecast in 2010, we'll do it again in 2014, there are scarce resources. So if you give a candidate only a 5% chance of winning, then maybe her congressional committee won't fund her as much as another candidate. Right. And so that can be complicated. Or when we do things like say, well, here's the optimal allocation of resources between different states, we had our return on investment index. Right. Implicitly that kind of assumes that the campaigns are behaving the way they have in the past. Right. It's kind of like the, actually the kind of sort of like the Lucas critique almost in economics. Right. Where you're modeling things based on past data but you're implicitly accounting for decision making of past policymakers. Right. So the extent to which certain states might move relative to other states implicitly includes the assumptions that campaigns were making before. And if they change the way they behave, maybe target particular states differently than they did in the past, then yeah, you have some structural uncertainty in the problem. So these things are really complicated. They kind of keep you up late at night. I think they're one reason in general to when you are making a choice as far as specifying a confidence interval to err on the side of including more uncertainty rather than less. Right. Although you can have the reverse case where if you have a self fulfilling prophecy, then you can be accurate just because that's what you predict is going to happen. And lo and behold, it does. When I was working for did some work for a movie studio a couple of years ago and they had a belief in their culture that some particular random date, say the first weekend in October, right. Was a very good time to release a movie. I think they had some movie that just had a good script one year. Right. Was a well directed movie that did well in that slot. So every year thereafter they put some breakout film in that October 7th slot, let's call it. And because they believed it was going to do well, it would pump a lot of marketing budget behind it and lo and behold, it did do well. Right. So it became more and more entrenched and they seemed very, very smart because they predicted that every year this film would be successful. But of course you pick one of your better films and put lots and lots of marketing muscle behind it, then it'll work. But it has nothing to do with the date on which it's released. Yes, we see this at Google. I talked to an advertiser once who said I know my advertising works because I increase my expenditure every December and what do you know, my sales go up. So what do you think is the best way to evaluate the performance of forecasters who are computing probability estimates for discrete events and in particular what of these events are rare? If they're rare, it's tough. Right. If they're common, then you know, calibration in my view is the best test. Meaning of the things you say will happen 80% of the time, do they happen 80% of the time. And it's also, it's Problematic if they happen half the time. Right. But also they happen 100% of the time. Right. People don't realize that, like if you make these 80, 20 calls. Right. You're supposed to get that wrong 20% of the time or you're doing something wrong. Right. You're being under confident, I guess. And sometimes we've been accused of being too conservative in kind of the confidence intervals that we specify. But there are also cases where the Senate candidate, North Dakota Heidi Heitkamp, we had her with only an 8% chance of winning and she did. Right. I think we had Rick Santorum with a 3% chance of winning the Iowa caucus at one point before his surge, or a 2% chance. Right. And. And he did. So we look at that over the long run, but over the short run, then you have to be. I don't know. Right. I mean, there's no way to. If you don't have a large enough sample size to test things truly out of sample, then yeah, there aren't any great substitutes for it. Right. And I tended to thought more to I guess I call them Bayesian priors. But to think that. Does the structure of this model make a lot of sense? You can make inferences. One thing I found helpful in writing my book is that you talk about forecasting practices in different fields. And so you can learn implicitly from that. Right. Well, here are some good principles that generally produce more reliable forecasts. So am I following those guidelines? And if you are then having one additional case might not tell you all that much necessarily. If you have 100 additional cases or 1,000, then sure, it tells you an awful lot. And one thing that Google gets to do is because you guys can kind of collect data on demand, basically from billions of users. Right. Then it's a very. Here you can solve a lot of things through trial and error. And if you have that much data, then that is probably a better way to do it. Right. More kind of a brute force approach. But I think people don't necessarily aren't sensitive enough to how much your whole attitude and approach should change when you're in an environment that isn't data rich. You almost have to do some of the opposite things really, I think, and not be overly fixated on how any one result goes. But it's tricky for people. So tell us about your tech setup and workflow. So what programming and stats languages do you use? And one question I know this audience wants to know is do you work in a cubicle with three other people? Do you get Free massages. I do not get free massages. The food at the New York Times cafeteria is not as good as the food at the Google cafeteria. So I don't really have as much of a routine as a lot of people do. Right. In part because I go through different phases where I'm traveling a lot or promoting or writing the book or promoting the book or doing media or not doing media or working on a model or going to the office. Right. So there's not really a routine all that much. I do have a desk at the Times, which is shared with Micah Cohen, also the538 blog. But I go in probably two or three times a week if I have something else to do in Midtown. Otherwise not. I find it hard actually to do writing when you're in a newsroom. To write, I need total quiet or music playing. Right. But I think writing is the thing that requires more concentration than anything else I do. Although programming also, I'm not a very good programmer. So when we actually have to program the static code, that requires a lot of effort. And for that stuff, I think you need some silence and some quiet. Where offices are great places to go to socialize. Right. But I think for some types of tasks, you need to think on your own as well. But my setup's not very advanced. I have a Asus laptop, which I need to replace because it's basically falling apart right now. I use Stata for most stuff and Microsoft Excel. I used to do everything, try to do everything in Excel, and that kind of was a bad idea. I mean, Excel's not a bad program. It's a pretty good program. Right. You can kind of think of it as like a visual programming language. Right. But. But in 2008, it took like two hours to actually run the model on Excel because it was so slow. Now you can do it in about five minutes. And I'm sure you could get down to a minute and a half if I were a more competent Stata programmer. But to have some kind of code is helpful. What's the next thing you want to learn in this area? In this area? The tools. The tools area. I don't know. I mean, maybe R, but it's hard. That's a good answer. Yeah. The switching cost is pretty high when you're used to having the 15 static commands that I know. I know how to make them do anything in ways that are probably really inefficient in some sense. But it's hard to switch. But it might be good to go to R or something. I think that would be the one that would make the most sense. I would think. So. What's the most statistically unsound tactic in professional sports? Well, I think in terms of being something which is routinely done badly and which cost teams quite a bit of win expectation is the way the NFL teams operate on fourth down, right? They punt way too much and to a lesser extent go for field goals instead of going for it more often than they should. I've always found this ironic, by the way, because the good, like, masculine thing to do would be to like, keep your offense on the field. Right. And go for it more. Right. You think the bias would almost work the other way, but it doesn't. Teams punt way too often. They're much too risk averse. The announcers tend to be really punitive to a guy when he decides to go for it on basically 4th and 2 or 4th and 1, you should be going for it almost all the time, unless you're deep in your own territory and you have a big lead already, but your incentives again become. Become distorted, where if you go for it and the quarterback gets sacked or something, you'll come in for a lot of heat, where the normal thing to do is to punt. And so you don't risk as much there. But there are some teams, the Patriots are very good about how they manage that stuff. And those decisions are worth about half a win per year, which doesn't sound like much. Right. But over a 16 game schedule, that's worth quite a bit. It's worth basically millions of millions of dollars if you look for the price an NFL team pays for a marginal win. And just based on reading his literature and keeping your offense on the field more, it should be fun. But teams have screwed that up for years now. This is the Paul Romer file. The Paul Romer paper. Yeah, yeah, yeah. And he was the one who found that it's actually half a win a year just based on this one type of decision, which will come up three or four times in a game. But it's odd that so few teams have done that reading. I suppose. They don't read the journal Political Economy for some reason. No. Yeah. So in the North Dakota, Montana Senate races, which your algorithm got wrong, you excluded polls that turned out to be accurate. So what went wrong? And do you plan to make any changes in the future? So I guess in the North Dakota case, the Heidi Heitkamp campaign, she was a Democrat in that race, right? Released polls that had her ahead, and usually internal polls exaggerate a little Bit, although the rule is kind of like if you're actually winning, you don't need to lie in the internal poll. If you're losing though, you do. So you kind of have these asymmetries. Right. So people said the fact that she was releasing polls and her opponent wasn't, but in general, we don't use polls released by the campaigns we think in the long run. For example, in the Wisconsin recall election earlier this year, you had a number of campaigns, polls released by Democratic groups that had the race tied, whereas the public polls had Scott Walker up by six and Scott Walker won by six pretty much. Right. And so that's the case more often than not. But North Dakota was a case where the model, if it doesn't have polls, it defaults to what we call state fundamentals, which means that, hey, it's North Dakota, the Republican will probably win. Right. And so there wasn't very much polling data at all. So it gave the fundamentals a fairly heavy weight. Right. And I don't know, I think it was. I think it's a well designed, sensible model, basically. Although that was a case where people who were using the high camp internal polls often got the right call and we did not. But, but I'm willing to live with that. One other thing too is that we model that we calibrate it based on how well polls have done since 1990. Looking back on every Senate race since 1990, and it's not like that the fundamentals calculations are perfect, but the polls are also often pretty bad in Senate races. So it helps to hedge your bets a little bit more. If it's true that the polls are on a secular trend toward becoming more accurate, then you could trust the polls even more. I'm still a little bit skeptical that polls are really on a long term upward trend. Right. When you've had, when it's becoming harder and harder to reach people on the phone. Internet polls had, including Google's, had a pretty good year this year. But. So I don't know, I think we're going to have, look, we're going to have some year where the polls do bad probably and someone will look really smart for saying trust the polls. I mean, frankly, the polls in the GP primaries this year were a pretty wild ride. But look, if the polls are right, then just trust the polls and you'll do pretty well. And it's hard to add necessarily a ton of value on top of that. So 538 is exceedingly transparent about its processes and inputs, but it's not 100% open. Is there any path to being able to open up the source, allow others to run simulations as well, or is that not in the cards? So there are kind of three different levels here, right? One is the level of transparency that we have in practice, right. Where things are described very fully. And I think there's room. So sometimes you'll have, say the academic standard is maybe you release your source code and all your data and there's a lot to be said for that. Right. I think it's also important though to explain not just what you're doing, but why you're doing it. Right. You might have a fully specified model that's publicly available, but that's also a terribly designed model. Right. An overfit regression model or something that just doesn't make any sense or makes very questionable assumptions. So the fact that we're kind of walking and talking through the model and explaining why we made the choices that we did, I think is helpful. With that said, I would ideally like to have fuller, not complete, but more transparency and more kind of formal language describing what we're doing. Part of what happens is that writing these articles that are methodologically based take a lot of time. Right. And so I'll have an aspiration to have, oh, we'll have this basically 10,000 word article that describes the 538 model more fully. And then there's some breaking news story, right. Some Sanders makes some joke about rape or something. Right. And. And you wind up covering that and not catching up to the methodological stuff. But I would like to move toward a place where we have more disclosure and transparency and it's fairly explicit and yes, people can reverse engineer it. I don't think we're literally going to release source code. It is a proprietary product. And one thing I have is because I license it to the New York Times, if you release that source code, then I'm paying out of pocket, in essence, in terms of when I next have the deal come up. But the idea is not really to hide anything. And I've always answered all questions about the model. I just don't want to quite make it literally so someone can run the source code and permute it themselves. You also have this perverse thing where maybe you guys have experienced this, different teams here, where sometimes when you disclose more, it won't shut up the trolls basically. Right, it won't. They'll just find other things, other little details to nitpick. Right. So I think it just means what's your own kind of ethical Practice as far as, just so you're clear about if you disclose things more fully, you're like, well, you guys know how this is working. You guys know that I'm not manipulating these numbers, right? Like, some people thought that we would actually go in and manually give a weight to every poll, right? Where I go, here's the We Ask America poll in Nebraska. It's 1.23653, right? We're just making up these like long numbers, right, Arbitrarily. Whereas now it's all done algorithmically. But I think things could have been a little bit better explained. I would strive to do more of that going forward, but you do have to kind of carve out time, I think, for that. And one thing also, sorry, I'm kind of giving long winded answers here, but we're also hoping that by the next election that I'll have a bigger team around 5:38. I mean, it really is. We have great interactive graphics people at the Times and great editors and so forth. But in terms of content producers, it really is just me and Micah, right? There was not a big 538 team. And the Times understands that the 538 team needs to grow or if I go somewhere else, then probably the same thing. So that'll be different, I think, next year. But I get, you know, a lot of it just is like, well, I ran out of time to do things I would otherwise want to do. So on this methodological point, what do you think about market approaches? Political stock markets like the Iowa Electronic Market or in trade, or these other cases where you're looking basically at crowdsourcing the forecasts. So, you know, in trade, look, in general, I think that markets are the least worst solution a lot of time. Right. It doesn't mean they're perfect, though. I think there is a question of how informed is the average trader at Intrade, for example. And by the way, the government decided, the CFTC decided to sue in trade today. I'm not sure if that's really the best use of taxpayer dollars. So that could be an interesting development. But look, I think that sophisticated political observers could probably beat in trade. I'm not sure they could though, if in trade, if they were more real money on the line, I guess I'd say. I don't mean that to sound totally pejorative or anything, but look, if you really had a good play, if you really thought you had a good model of the election, then you wouldn't be betting for a few thousand bucks on Intrade and You'd be instead rather working at a hedge fund developing a portfolio where it's pretty clear that, for example, coal stocks were correlated with Romney's chance of winning at Intrade, or green energy stocks were correlated conversely with Obama's chance of winning or different types of health care stocks. There are enough companies where they have predictable enough economic effects in the election, where even if you have some fuzziness in building your portfolio, you can make a lot more that way potentially. But the other question is you can always have hurting in markets where you have the blind leading the blind potentially. I think there were times in the end of the election where there was some odd activity on Intrade where Romney's price would shift a lot with no real reason. In the fundamentals. There were some theories that you had people, wealthy Republicans or whatever, conservatives kind of buying these shares to influence the way the race was covered by the media. And you had divergences where at in trade, Romney might be priced with a 35% chance of winning, but at other markets he'd have a 15% chance or 20% chance. Really big spreads that are not consistent with highly rational behavior. So you have this kind of abstract question of what if you really had a highly liquid stock market or prediction market, how would that do? Then you have the question of, well, how well does n trade do in practice? Giving the constraints it has on people depositing money and the lack of. And I think the answer is it's useful. It's a lot better than the pundits, right? For example, But I think once people start to invest it with godlike power, like, oh, it's the market, it must be. Right? That's exactly when markets fail. It's when you think markets are infallible that markets become much less useful tools. So can you put historical polls into your model and simulate the predictions from 2000, 2004, et cetera? And it would be nice to know if the elections are ever as close as the pundits say. Well, I think both at the end of the campaign, both the 2000 and the 2004 elections would have been closer, right? But remember in 2008, on the weekend before the election, the McLaughlin Group, they asked their panel, who's going to win the election next Tuesday, right? And three of the four panelists said it was too close to call. And in 2008, this wasn't even debatable, right? It was. Obama was up like seven points in every single poll of every single swing state, Right? I mean, just no excuse for that at all. If you're not just constantly trafficking in bullshit, basically. Right. So the fact that in 2012 you met some resistance, I guess, shouldn't have been surprising. Exactly. I'm not sure. In 1984 there'd probably be some columnists saying, well, Walter Mondale. Don't count Walter Mondale out. People haven't voted yet. Anything could happen. We're hearing reports of more yard signs in Roanok, but there's always someone to make the contrarian argument. Right. And it's hard for people to kind of accept, I think, reality in some sense for the end of a campaign. So let's switch to a sports question. Coaches and managers are generally judged by how many games are won and whether or not they win championships. Do you believe there will be a reliable way to measure coaching performance from a game to game basis? It's tricky. I mean, in the NFL maybe you can do a little bit more with those play calling things I talked about. But, you know, but I do think that one bias that stat heads sometimes have is to assume that if you can't measure something easily, that it isn't important. Right. Hypothetically, if you had a baseball manager who got his players to perform 5% better than they otherwise would. And how do you define 5%? I'm not sure. Right. Otherwise, because they're better rested or you have a better clubhouse environment or anything else, then that would be extremely valuable and that would outweigh the cost of laying down too many sacrifice bunts or whatever else. Now, can you measure that? It would probably be pretty hard. Right. You have too many degrees of freedom in the problem to really pin anything down in a more reliable way. But, but both baseball teams and basketball teams, I think have learned, for example, that as it's become easier to measure defensive contributions from individual players, that they were probably underrating defense, if anything, and that maybe stat heads in particular. Billy Beane's earliest kind of mid-90s A's teams had terrible defensive outfields because you couldn't measure it. Right. And so it was kind of assumed that, oh, it doesn't matter, but, but when you had Matt Stairs playing center field or something, that created big problems. And so they kind of learned the hard way that you had to try and measure this stuff and gain some advantage from that. But I think there could be coaches that make more of a difference, it seems, certainly, if you look at, for example, college basketball programs, that a coach can be hugely valuable. That gets tied into recruiting and everything else. So it's hard to say. But, but I don't know. So I think are we ever going to be able to have a coach metric? Like probably not, unfortunately. Right. It's just a hard problem to address. But that doesn't mean that their values should be dismissed out of hand, I think. So. Which past elections in the modern era would you most want to apply your methods to and why range over time and space here? I think it would be fun to apply your method to an election that had three viable candidates. Right. The 1992 election, although that close to the end would have been a lot of fun potentially. Although also the variance is much higher. In a three candidate race, a lot more crazy things can happen. So you could wind up with more egg on your face. But yeah, that would be the fantasy, I think scenario is to deal with an election where you had three viable candidates. Right. And you had states that could shift between blue and red and that would be fun. Yeah. Have you thought about looking at parliamentary elections? Well, I tried this for the UK and our model kind of sucked actually where Clegg mania wore off. Right. And the model was really bullish on Nick Clegg. Right. But that's also a case where in the UK you don't have a lot of district by district polls. Right. So you basically have national poll. Imagine if we only had national polls and then you had to infer from that what the electoral college would look like. Right. Any assumption you make is going to be, is going to be potentially crude or it's going to be crude and potentially flawed. That's kind of what you had to do in the uk. So that plus the fact that you did have three parties and you did have strategic behavior at the end where the Liberal Democrat voters were basically more of them were labor voters than Conservative voters. Right. And so when, when that bubble bursts, then labor over performed our predictions there. But it is interesting. Look, anything can be modeled as long as you can accurately predict the confidence intervals basically to say, well, in the primaries, for example, we would often say, well, we have no idea what's going to happen, but that Rick Santorum can get anywhere from 10% of the vote to 40% of the vote. Right. Because that's what a realistic confidence interval looks like based on how accurate the polls have been in the past. Although I would like to do better. I think our model for working on the primaries this year, I was kind of in a minimalist phase. Right. So I'm like, well, I know the polls stink, but let's just take the polls and see how well they can do and have Them be well calibrated. Right. And I don't know, I think I'd like to add a few more elements into the primary. So for next cycle, we might look at things like, for example, endorsements are a really good predictor in primaries, not because they influence voters, but because they show who the party supports and the party can tilt the playing field in all sorts of ways. Right. Where basically there was a week at which Newt Gingrich looked like he was actually a threat to win the Republican nomination. So he basically became a pinata. Right. Where anyone who would ever vote a Republican in their life and went on television, it was their job that week to. To bash Newt Gingrich. Right. And so that was how a party operates strategically to make sure it gets the nominee that it wants. And believe me, you know, Romney might not have been in the best candidate, but, you know, Obama versus Gingrich would not have been a better result for Republicans. I don't think so. Looking at those other factors in primaries might be more interesting. And we'll have, I guess, two opportunities next cycle. It's selfishly, journalists should be waiting for an Obama win because that way you get two open primaries next year. Right. Instead of just one. Those years are a lot more interesting to cover. I think with the success that we saw from Internet polling in the last election, polling is going to be a lot cheaper. England, for example, has very high Internet penetration and you could get some much better data in the future than you have in the past. Well, yeah, and it was good that the Google poll and YouGov and a couple other of the Internet polls did well this year because for better or for worse, this is going to be the future. I think inevitably people don't use their cell phones or, excuse me, their phones, certainly their landlines in the same way that they used to. I have a landline because it came free with my cable package. Right. And I use it if I have to do a radio interview or during hurricanes and stuff, basically. Right. But you're not going to reach me. The Quinnipiac poll is not going to reach me on my landline. Right. Where people screen their calls, even with cell phones. Now, increasingly, a lot of younger people are using mostly text messages to communicate, and you kind of use the phone as a backup almost. Right. So that's becoming a challenge. And it's kind of more natural now to encounter people as they're kind of browsing online. Now the Google Consumer Surveys project almost approaches being kind of random Internet poke, where it's like that's what you do ideally, right. It's just someone's every X number of seconds you have a random chance of a pop up window saying OK, well if you're going to vote for then we'll let you keep searching. That's the equivalent of the random digit dial method in essence. And Google's not doing that. Exactly. But the fact that Internet polling did well is encouraging. And the fact that you have companies that know how to use data doing this. A lot of pollsters are fairly primitive, right. They're using likely voter models that were designed 50 years ago. Right. And they're very defensive and stubborn when something goes wrong. I mean, for example, the Gallup organization, right. I mean their polls have been screwy for years, right. People focus on the end result which also haven't been that good for Gallup. But the fact that they'll have these wild gyrations from day to day, right. I mean that's usually a sign of a badly designed. When you have wild fluctuations that aren't dictated by the fundamentals, you're probably doing something wrong with the way they're taking their information. And by the way, because you are only reaching a fraction of voters now that all this really is, it is a modeling question. Right. It's not a pure polling question anymore where even Pew, even the best pollsters only get about 10% of people on the phone. That data has to be massaged. Right. And that means that you need statisticians to do it in ways that are smart. Right. And accurate. And if you're doing the same things you were 50 years ago, then you probably won't be doing very well. Okay, I'm going to ask the last question on my list here and then we'll open it up to the floor. Both campaigns use tech for their get out the vote efforts, but the Democrats seem to have been substantially more effective than the Republicans in this department mentioned. So that wouldn't show up in the polls. What would be a good way to measure the effect of those efforts? Well, I mean, you know, over performing the polls could be one potential indication of that. But you know, you could, I think if you looked at local voting patterns. Right. Did cities, for example, in Ohio where Obama had a field office, did he get more of the vote there than comparable cities where. Where he didn't. And those can be hard to design because maybe there are reasons why Obama picked certain cities. And so you have selection bias problems. But I think you want to look at more macro level data. When campaigns are testing the effectiveness of advertising that's what they do. You run markets in Raleigh, but not Durham or whatever, North Carolina, although you couldn't probably do that in practice. But you can set up natural experiments and see how well ads work. But it's hard to say. I mean people assume this Obama ground game made a lot of difference and I'm sure it made some difference. But also if you look at, go back now and look at states like California, Well, Obama beat his polls in California by several points on election day. Right. So it wasn't purely, he beat his polls in like Connecticut, states like that. So it wasn't purely just in the swing states where he over performed his polls. It was, it was kind of everywhere. Right. And so the question is to what extent is the ground game edge priced into the polling already? In theory, if you have voters who are more enthusiastic or more engaged by the campaigns and they're going to meet a likely voter screen to begin with and so they'll actually be in the poll already. So I don't know the difficult problems to figure out. I do think that the fact that, you know, look, Bush was very data, Bush and Roe were very data literate in 2000, way ahead of Gore that year. And, and they had a swing state advantage then where they lost popular vote, won the electoral College. Right. Why Republicans didn't keep that up? McCain really abandoned the ground game in 2008 in part because of a lack of funding. Romney had plenty of money, but I think he was just maybe a little bit stubborn. Right. I think maybe there's a belief that, well, our message is so powerful, probably some startup companies that think our product is so good that we don't have to market ourselves kind of thing. Right. That's usually a sign of trouble. Right. And maybe the Romney campaign thought that, you know, once people see how terrible the economy is, well then, well then they'll vote for us no matter what. Right. When you really, you have to work on the marginal voter pretty hard to get them to turn out or to change their vote. Sounds like you're saying that Romney should have hired a consultant. Well, so the other question is, the other question is or two, does the Obama campaign have a systematic advantage? Because people who are data literate tend to be liberal or Democratic leaning. Right. And you know, I mean, how many people vote for Romney in the Bay Area? Right. Like maybe 20% or something. Right. So you potentially have some of these issues as well where if you have a party that has a reputation for being anti science or anti empiricism. Right. Then it's maybe not going to necessarily get the best and brightest. And look, there are some bright people working for Romney. But I think it actually is like a little bit of a challenge potentially. One of my colleagues at Berkeley once said, I've heard there's such a thing as a Republican, but I've never actually met one. So let's open it up. Now. The acoustics are bad and we have to repeat your question. So the floor back there loudly, please. So the question was, what's wrong with the polling companies? Are they not effective? Are they politically biased? What do you see the problems there? I mean, I tend to think for the most part polling firms aren't deliberately weighing the data to match their political bias. There are incentives to be accurate. I think you might have some exceptions where, you know, for example or not, for example. So like Rasmussen Reports, for instance. Right. Their business model seems to be to provide good news for Republicans. Basically, it's within a certain distance of being realistic. Right. And so. Well, I'm just saying, like that seems to be almost a conscious decision on their behalf. Right. Because they had polls and every year their polls have been had a Republican lean right. Relative to consensus. So when Republicans do well, then their polls are flying. But when they have a bad year, they're really far off. Right. But that would be something that would, I think a company who was prizing accuracy would go in and look at and say, well, why are we always Republican leaning and often wrong relative to everyone else? And you would fix it. Right. But they seem to have a different idea about the way that they can make the most income and get their name in the news the most. But I think, you know, I don't think Gallup has a Republican partisan bias or anything. I think they just, they're just a stubborn old company, Right. That, you know, wasn't willing to open up the hood enough and wasn't willing to admit that they had a problem. Right. And so we kind of get back to this whole question of honest self assessment. But you know, people assume when you had the unskewed polls guy, he was convinced that the Fox News poll was part of the conspiracy to rig the election for Obama. Right. And believe me, when you had the New York Times, I think had some poll where last year Obama had like a 39% approval rating, just an outlier, really. That kind of happens sometimes. People are like, well, you know, the New York Times is, you know, they've become a conservative paper now. Right. This proves it. I mean, the polling wing of a newspaper is Usually the least biased part. Right. Of a news organization, including often the, you know, reporters themselves, where there are some great reporters, but there's a lot of subjectivity in how the news is reported, even on the front page. Right. And polling, at least you have a relatively. If you do the right things right, you kind of follow the recommended practices, then you come closer to objectivity, at least potentially. Okay, so the question was, you've shown how statistics can help dramatically in sports and in politics. So what's the next area? If I knew, I would be taking meetings with VCs right now. I don't know. I think there are a lot of areas where education is an area that fascinates me because you're seeing more data used. And I don't know a lot about how it's being used, but I bet it's not being used all that well, necessarily. There are a lot of things, I think everything from. From kind of. I did a project for New York Magazine where we looked at different neighborhoods in New York and tried to rate them. We found, by the way, actually, that New York rental apartment prices were very efficiently calibrated on a neighborhood to neighborhood basis where basically, if you look at distance to midtown Manhattan plus Parkland plus schools, then you can explain. You get an r squared of 93% and explaining prices per square foot. Right, right. But things like that might be fun. But look, I mean, for every problem where statistics have made a lot of progress possible, you also have areas where they've been probably misapplied in different ways. Where in the book you talk about how earthquake prediction's not gotten a lot better, for example, I think in economics, although some of the stuff that Hal is doing. Right. I mean, how often do you. I mean, do you talk with, like, the White House and stuff, or do you talk with. Well, I can't. Yeah. I don't mean the White House particularly, but I mean, do you talk with like, the Bureau of Labor Statistics? Because they should be. Yeah, yeah, yeah, yeah. Like, I mean, this is a rare space where you actually have new economic data being produced that might be underutilized. Right. But in general, most economic people have been trying to, you know, do economic forecasts for a very long time and not really gotten anywhere. So it might be a question of saying, okay, well, we kind of. We're crying uncle, we're giving up, but we're going to specify more accurately how much uncertainty there really is in the forecast. Right. So there are kind of two ends that you can burn. We can say, well, if we want to bring our predictions more in line with reality. Right. Sometimes it's a matter of applying better techniques and sometimes it's a matter of admitting what we don't know and problems that are intractable, at least as far as the near term goes. Yeah, I think there's a whole area, economists sometimes called nowcasting, where the objective isn't so much to forecast, but to try to get an accurate measure of what's going on now rather wait for the government statistics to come out six weeks or a quarter later. And with all the real time data that's available in the private sector, I think we do much, much better job of nowcasting the economy, tracking what's going on on a moment by moment basis. But the one thing I did want to suggest as a very attractive area, at least for us here at Google, is marketing. Because marketing is a subject that's full of mythology and shamans and all this stuff. But if you do it right, as we try to do here at Google, I think you really make some progress. Yeah, and probably industries like TV and film are cases where the studios are very smart and very scientific about some things. But in terms of how do you build it, what's the right portfolio of different film genres to build or questions like that. Right. Or to, you know, if you have a particular product, to whom should you show that? Advertising. That's where Google probably does, working on analogous types of problems. But I think there's probably a lot of low hanging fruit in those sorts of areas. The question is, when you put economic variables in your model, for example, you just put them in with equal weight and the explanation was you wanted to avoid overfitting, but how do you know you didn't distort the model in some other way? Why was that an attractive choice in that circumstance? There's some literature kind of in the forecasting literature that says that as a default, if you don't have enough empirical data, if the data is not rich enough to develop weights from your sample, then just tending to weight factors equally is not a bad habit. And at least that way you avoid convoluted schemes where you're weighting something at 80%, something else at.02%. Look, if you have we use seven economic variables and we're trying to fit them based on 10 past elections. Right. So there's nowhere near enough and they're highly correlated with one another. Right. So you know, we're in the ballpark of having enough data to develop reliable weights based on the in sample result. But really philosophy too is just that we're not trying to. We didn't want to develop an economic or an election specific measure of the economy. We just wanted an overall economic index. Right. Ideally you would have some economic index that someone else developed. I mean, some banks will have these. For example, like Goldman Sachs, their forecasting team has a version of GDP that's like better than gdp, right, because it includes more data series and reduces some of the noise you get in the GDP estimate. Right. So if we had something like that, then that would be ideal. But there's not a public product that quite fit the bill. So we just want an index that basically captured kind of what's going on in the economy in one number the extent you can, which of course is inherently a little bit foolish. But it's probably better to have something that accounts for different major aspects. So jobs, income, industrial production, inflation, different major categories. And we didn't think you had to get a lot fancier than that in terms of weighting them because it's not clear what you're weighting to. Right. We wanted the weight to correspond to how important is something to the economy, basically. And then say, then you got to answer the question of, well, how does the economy predict elections? Instead of kind of saying, well, this one obscure variable that we found on, on Fred perfectly correlates with elections from. I mean, those models have not done all that well. People get too cutesy. But the principle that you want to aggregate information and have more than one indicator is pretty helpful. And some models, like our model, actually would have had just the economic component, actually would have had Obama winning by two and a half points. Right. Which is pretty close to the actual outcome. And other models people use that took indices or that aggregated different economic data points together also did pretty well. Whereas some of the ones that just fixated on one day at a series were pretty far off the mark. One had Romney winning by six points or something, which was not a very good forecast this year. So again, I want to repeat the question was what's the craziest thing or funniest thing a fan has said to you? Pretty tough question here. I'm not sure. I was on the plane the other day, I was flying coach in a middle seat out in San Francisco. Right. And some woman in first class said, I guess they don't have any poles and coach or something weird like that. Right. But there are all these weird benefits. So I found now that this will wear off, I'm sure. Right. But if I'm going out to dinner with my friends. Right. I'm booking the reservation in my name. So there's like a 30% chance that you get like free round of champagne or something. At some point I should reveal that secret. Right. Yeah. I think there's going to be a lot of people trying that. Yeah. Trying that plan in San Francisco. So it may way off. No, I went out to some restaurant which it was called Local's Corner, which was very good, by the way, in the Mission District. Very good. But they said they had Googled me to make sure that I was. Or Twitter followed me. Right. To see that I was actually in the Bay Area. Right. Because they didn't want to have. Because they've had this scam before where someone's like, I'm Peter Fonda. Right. And maybe they weren't quite sure what Peter Fonda looked like. And so they try and play that game to get a better seat. Right. So they actually followed the paper trail and did their research. What kind of seat do you suppose Mitt Romney would get in San Francisco? Well, I was thinking what would happen if. Sorry, we're full. Exactly. Because look, in 2016, in 2016, if the Republican's likely to win, then hopefully we'll have the Republican winning in our forecasting model. Right. So that's like I'd have to go to like, I don't know, like Dallas or something. Right. Give you a free plate of barbecue ribs for correcting. I like ribs. I like Texas food, but yeah. All right, last question. How can we improve the teaching of statistics? That's a big question. Right? An important question, I think. Important. And I'm not sure why statistics education and math education in general is so abstract in this country. I think you just kind of give kids data and kind of let them play around with it at a certain level. And I don't know exactly what type of data you give to a high school class. Exactly. But to have kids conduct a survey. Right. Or to have kids just start counting things and come to some conclusions about them. I mean, the fact that people are phobic of math is odd. You'll get the comments sometimes from readers on my blog, like, oh, I thought I hated math, or I'm terrible at math. But I like your blog. You'll hear that. Right. And probably that person is not terrible at math. Probably they had some teacher who made them dislike math for some reason along the way because it wasn't taught in a way that had practical, real world applications. And I hope that will change. I'm not sure I don't know enough about what different curriculum teachers are using. I think probably the fact that you have more teaching to the test now is harmful, of course, to any type of innovation in math and science education. But that could be a subject that I want to write a book on or something down the line is how math is is taught and how that can be improved. Okay. Nate, thanks for coming. Thank you, guys. Great.
