Geoff Huston 0:00 A lot of routing protocols at the time that didn't use TCP had this short minded or amnesiac counter every 90 seconds, every period, it said you could be losing your marbles. Let me refresh your memory. Here's everything I know, and those broadcast storms of just refreshing stuff you already were told was a great fail safe in an unreliable transport, but ultimately was incredibly inefficient making this TCP say, I've told you, I'm never going to tell you again, unless it changes. That was good. That cut down the amount of routing traffic like crazy. The next thing is looping. BGP had a very clever way of detecting looping, because if A talks to B, talks to C, talks to D, talks to A, how do you know it's a loop? It's kind of hard sometimes, if you think about it, but what BGP did is it attached, almost like a snail trail, a line of thread through the maze. Every time a network, its prefix went through a network, it added the number of that network to its trail. So let's say it's network A, the path is A. When it gets sent to B, it goes well, the path is BA. When that set sent to C, the path is CBA. In other words, the networks that have seen this advertisement sign it and attach it, that signature, that their number, their identifier, to that update. So if you have loop it, you're seeing yourself. And thought, seen this. This is not real. I've already seen it. I'm going to reject up front path vector was the big difference between BGP and rip the Routing Information Protocol, which did not have that and consequently, the only way it figured out a loop was to actually count to infinity. George Michaelson 1:57 You're listening to ping. A podcast by APNIC discussing all things related to measuring the Internet. I'm your host, George Michaelson, this time, I'm talking to Geoff Huston from APNIC labs again in his regular monthly spot on ping. for over two decades, Geoff has been publishing the CIDR report, a daily web update on the state of BGP worldwide with information on routing table size and where it comes from, along with information about anomalous announcements started by Tony Bates at Cisco in the mid 90s, then carried on by Phil Smith, who was also at Cisco at the time, and then Geoff. It's a remarkable continuous measurement of the state of routing. The CIDR report grew out of concerns with the consequences of worldwide growth in the size of the routing table, and especially a trend towards de-aggregation, when as holders deliberately announce more specific prefixes from Their blocks to optimize routing behaviors in a process called traffic engineering. But what is CIDR or classless Internet domain routing? What problem was it designed to solve and in a modern Internet, Is this the problem we think it is? Has the CIDR report been overtaken by other events in routing? Geoff, welcome back to ping. What shall we talk about today? Geoff Huston 3:23 Hi George. Look, I think it's time we talked about routing. It's been a continuous Internet subject for the last eon or two. George Michaelson 3:31 Oh, it's been a while since we did that, at least two episodes. Geoff Huston 3:35 Oh, this time I actually want to kick into a report that I've been producing for at least 20 odd years, and has been produced for more than 30, I think about a sort of a summary of the state of the routing system. It's called the CIDR report, George Michaelson 3:50 CIDR Geoff Huston 3:51 But what I'd like to do first is to Yes, now it's not about something you drink, truly. I'd like to sort of talk about the routing problem, as it was, right from the word go, and how the CIDR report sort of snuck into all this George Michaelson 4:07 Okey dokey, roll the clock back 20 years. 30 years? Is that enough? Geoff Huston 4:12 Oh, keep going. We're talking 1980s and I suppose one of the things that was totally different about the desire of computer networks, as distinct from, say, telephone networks and similar ones of their day, is that computer networks were meant to be self learning, that if you had one device, it wasn't a network. If you had two, they were meant to be able to discover each other, [George: right] And if you had a collection of these things, some of them connected over local area networks, big thick Ethernet wires, and some connected over thinner wires. Y'all meant to discover each other. George Michaelson 4:54 That was very much driven in the experience of pre existing structural networks that were built with much more like a god designer who said, Thou shalt connect this to that, and thou shalt fan it out in these ways, to these remote entry vehicles and these concentrators and will gather data from these locations and send it along these links. It's like somebody laid out a map of a network designed around devices and a computer. And the 80s people were saying, we don't want to have to do that. Geoff Huston 5:23 Well, that's right. And I think in some ways, the 80s saw the beginning of Ethernet as a local area network. And the thing about Ethernet, which was, I make it sounds like nothing today, but at the time, you just plugged your computer into that common wire, that bus cable, originally with a thick yellow cable you got at your drill and your very expensive transceiver, drilled into the coaxial cable on the little black mark, tapped it, connected your computer, and you're able to talk to everyone else, other than the need for a drill and an expensive transceiver. There wasn't a controller of these Ethernet there was no one in charge. You just connected your network, and while we sort of tried to push them further out into the wider area, we wanted to take that plug and play convenience with us. We wanted to use networks that did not have a controller, that were actually an amalgam of almost self organized networks that became off a self organizing beast. And that was the desire of the routing protocols at the time [George: right] to exchange the role of the network controller, the Big Fat Controller, if you're a fan of Thomas the Tank Engine, to exchange that role into one that was realistically one which was brokered by the protocol itself, [George: right] So that was the magic desire. And there are a number of attempts at this by various vendors in various ways. The one that a lot of us had experience with was actually a vendor network called DECnet. George Michaelson 6:59 Yeah, used much as you did Geoff. It was a regular part of my life. The thing Geoff Huston 7:04 about DECnet was It was organized into two sort of tiers. There was an area where machines inside an area had a detailed conversation about their connections. You're on the same Ethernet as me. I'm connected by a dedicated wire to you, etc. And so inside an area, all the computers needed to be interconnected one way or another. Didn't need to be a mesh, but any computer in a network in an area needed to get to any other computer in the same area without leaving that area right, just sort of like a local city or something. And then DECnet had a second upper level protocol, almost like an inter area protocol, where there were area borders, and the area borders spoke via dedicated connections to other area borders. And they said, Hi, I can reach all of area one. How about you? I can reach all of area two. And so to send a packet between a computer in area one to one in area two, you directed your packet to your local area border router, who passed it over to the right area, who then sent it on to the right computer in that area. George Michaelson 8:20 So Geoff the way you're describing this, it's like there's a bit from Box A and a bit from Box B, because if you're inside the area, you're making it sound as if it used the technique to discover all the things within that area, the way the protocol worked, you'd eventually discover every printer, every other host, every other computer. But if you want to go to another area, there was a higher structural layer that was having to do some management and passing right, Geoff Huston 8:45 right. Well, things almost as a an atlas, a mapping system. If you're in a city, it's nice to know every street, but if you're trying to get a packet to a different city, you really don't care. You want to get it to that city and let the destination city sort itself out. You didn't want to start the network and the protocol interactions with all the detail of remote location, so you literally did divide and conquer. Divide into areas, and at the inter area level, you only looked after the connections between these area border routers, and within an area, you looked after the detail, George Michaelson 9:21 you've already introduced the ideas of something local and a boundary with a border and a border router, and the concept of an area. And to me, these are four terms that are very interesting to establish inside this model this early on in people's thinking. Geoff Huston 9:37 I think DECnet was a model that actually we could have borrowed and pushed into the Internet, almost subconsciously, I think so many of us had cut our teeth from that particular vendor network, and this concept of areas Radia Perlman had a lot to do with it when she was working in digital was actually really seductively, right? It sort of pushed detail into areas where you need a detail. But when you're simply trying to move around a country or around a region, you didn't need to do that. You just simply talked about areas. And there were some pretty massive DECnet networks, the NASA Science Internet, the high energy physics network. They really did span most of the Northern Henry one way or another. They weren't fantastically big networks, you know, kilobits per second, but they did have a lot, a lot of members, even digital's own corporate network, George Michaelson 10:29 even this semi automatic mechanism of things being discovered and things forming areas and communicating with each other, worked at the scale of the whole of North America, the whole of Britain, places of that kind of scale could quite comfortably build one of these things. Geoff Huston 10:44 So, yeah, digital built a network with 100,000 computers in it as their corporate network. It was possible to scale it up. And so when we started to build this Internet beyond the limited Research Project Agency, the ARPAnet exercise, into the bigger field, that routing concept we carried with us, and the Internet itself, I think it really took off when the National Science Foundation transformed the Internet by supporting a thing called the NSF net across the United States to connect their super computers. And, you know, all of a sudden, instead of a few 100 computers, which was the ARPANET, the NSF net was talking in terms of 5000 10,000 20,000 network, there 20,000 computers inside a few months. Not only that, but they wanted to interconnect to the NASA network, to the high energy physics network. They wanted to interconnect to the sort of quickly dying ARPANET, the old defense network, old defense sponsored network. So the problem was complex. It wasn't simple, and they needed to use much the same trick. But what they did, I think, was quite insightful. It was that area and time when divide and conquer made sense, that there was no single corporate and what they very quickly realized was that the networking, routing protocol you used inside a network an area didn't need to be the same one as you used outside that area. George Michaelson 12:14 That is quite a radical idea when you compare it to other behaviors extant in the Internet in these days, the idea that things didn't have to be the same, in some ways, ran counter to a culture that said, we're going to standardize the format of email addresses across the entire surface of this network. We're going to make telnet and ftp pretty much use the same kind of logical structure. And here you are saying divide and conquer allowed people to stand up and say, let's actually think about doing it differently inside an area and differently between areas. That's interesting. An interesting difference Geoff Huston 12:50 actually, like a difference between a commercial service offering and the so called, you know, the research people at the time, the commercial thought, I solve all your problems and so DECnet really was, I think, two different routing protocols smashed into one because, you know, DECnet needed to do everything, whereas in the research world, you know, the academic research world, it really was. Look, we're not going to tell you what to do. Here are some things, but if you don't like them, go run your own. And there were, in the 80s, a lot of university campuses, I know I lived in one where driving your own and building your own local area network protocol was a sport, a competitive hobby. I worked at the ANU, home of ANU net at the University of Monash in Victoria, Monnet and on and on and on. George Michaelson 13:38 And UQ net, which is where I was and csironet, which is where I was before that. And you could build a surprisingly large amount of interconnected state using incredibly simple protocols, like RIP, the protocol we use to construct the network at UQ, right, Geoff Huston 13:53 right. So there was this sort of thing of, roll your own local network, roll your own network within sort of your network, and then a case of, how do you hook them together? And that's kind of well, in some ways, it's quite easy, if you think that a routing protocol is simply to tell you where addresses are, address prefixes. So each network has a bundle of IP addresses, in this case, and hopefully they're contiguous. Well, at the time they were and they came in three sacks, the Class C sack, and that had 256 addresses, the Class B sack, which had 65,000 and for the privileged few, there was the class A sack with 17 million addresses. George Michaelson 14:39 Yeah, but this kind of anti Goldilocks moment, because the class C sack that anyone could get easily was really too small, and if you were big right on the Class B sack was running a bit, and the class A sack, well, you couldn't get it. Geoff Huston 14:52 We will get there. But I want to talk about this for a second, because what the routing protocol needed to do was in the exterior sense, to say to all the other networks, or at least to the ones that were immediately adjacent, as networks, not as computers, as networks. Hey, I can reach these following Class C addresses bundles and a Class B bundle. Here have the bundles that I can reach. You've got a packet address. Any of those bundles, send them to me. I'm your network. You tell me what you can get to George Michaelson 15:28 I'm your network, which is a magic moment in information sharing, isn't it, because it takes a thing that could be an enumerated list. Here's 255 things, and it says, Nah I'm just going to refer to it as a common bit and how many there are, and that's all I have to tell you. Geoff Huston 15:30 And so this idea of bundling up and talking about reachability by network sacks or prefixes appeared, and the original implementations of this was the EGP, interestingly, the exterior gateway protocol, and there were a few sort of things around that, but it was all relatively simple. But oddly enough, the problem they were solving was not simple, because already we had nine regional networks, a National Science backbone network, a NASA network, a high energy physics network, a network in the United Kingdom, a network, I think in Germany, very quickly, there are a bunch of these coming up and getting from one place to another was a non trivial problem, and you couldn't just send the details of every connection to every part of the net. There was no bandwidth, no capability. It would work. So we needed a better form of exterior network routing. And you know, the answer was pretty simple, as it turned out, it came out in 1989 when IBM's Jakov Rechter, IBM was the contractor to NSF, and Yaakov was the principal scientist over there at IBM, and Cisco's early employee, Kurt Lockheed, came up with this protocol called the Border Gateway routing protocol. Now it was simple because it relied on something that I think was originally in a paper in 1956 the Bellman Ford, distance vector protocol. I'll tell you everything I know when you hear what I say, add one to every metric about what I know. I can see this address prefix in one hop so you go, very good, Geoff, I see it. Therefore in two hops. I see this one in five, because people have told me and told me, and you say, very good. I'll make it six. George Michaelson 17:29 So the essential quality here is that we've stayed in the space of we don't want to have to have a big boss that tells us what to say. We all shout out into the world the stuff about us. But instead of shouting the entire list of what we got, we now come to consider, well, I consider myself an area. I shout out efficiently the prefixes the block of things I have. And you've just said, we added a twist to that, that when you do it, you use this technique from the 50s called Bellman Ford, shouting it out is cost of one, and the guy who hears it shouts it out again, but adds one to make it a cost of two. And if you were the third in the chain and you heard it, you'd shout it out, and you'd add a cost to make it three, right? Geoff Huston 18:09 right And if you had two connections, people shouting at you, and one of them says, I see this address prefix with a cost of one, and the other one says, I too see it, but my cost is eight. Obviously you'd pick the cost of one, you know, so you naturally prefer the shorter cost. George Michaelson 18:26 So we have one unit, one unit of cost and its distance. And if everyone obeys the rules, you can rationally determine the lowest, the best cost. Geoff Huston 18:36 Right? Simple distance. You and I could connect over anything we wanted, everything was cost one if it moved from one network to another. Wasn't very sophisticated, but I want my big fat wire to be more, more preferable than your tiny little wire, and it's kind of No, no. BGP doesn't do that. The distance between a network is one stop trying to be clever. George Michaelson 18:59 There's no dimension for that. Geoff Huston 19:01 However, there were a few things that were clever. The first thing they did, actually, I think it was, was one of the big fundamental things, is that the transport that BGP used was actually TCP, which means, when I tell you, and you send me an acknowledgement, you know it, I know you know it because you acknowledge the packet. I never need to tell you again, only if it changes. George Michaelson 19:27 So BGP didn't have to construct its own way of saying. Did you hear that? Are you sure about that? Can I believe that because TCP told you they got it. Geoff Huston 19:37 A lot of routing protocols at the time that didn't use TCP had this short minded or amnesiac counter every 90 seconds, every period it said you could be losing your marbles. Let me refresh your memory. Here's everything I know and those broadcast storms of just refreshing stuff you already were told was a great fail safe in an unreliable transfer. Bought, but ultimately was incredibly inefficient making this TCP said, I've told you, I'm never going to tell you again, unless it changes. That was good. That cut down the amount of routing traffic like crazy. The next thing is looping. BGP had a very clever way of detecting looping, because if A talks to B, talks to C, talks to D, talks to a how do you know it's a loop. It's kind of hard sometimes, if you think about it, but what BGP did is it attached, almost like a snail trail, a line of thread through the maze. Every time a network, its prefix went through a network, it added the number of that work to its trail. So let's say it's network A, the path is A. When it gets sent to B, it goes well, the path is BA. When that set sent to C, the path is CBA. In other words, the networks that have seen this advertisement sign it and attach it. That signature that their number, their identifier, to that update. So if you ever loop it, you're seeing yourself and thought, seen this. This is not real. I've already seen it. I'm going to reject up front path. It was the big difference between BGP and rip the Routing Information Protocol, which did not have that. And consequently, the only way it figured out a loop was to actually count to infinity. George Michaelson 21:24 That takes a long time. Geoff Huston 21:26 Well, they made it 16, but it's still not very good. So we've talked about the separation between inside your network and between networks. We've talked about TCP, we've talked about path, vector and transport. But the other thing I think we've just alluded to, but it is very, very important. Is no one's in control. There are no permissions. It is a network of peers. George Michaelson 21:47 Well, we could say there's a kind of moment here that we're not going to call control of routing. But if you said, put a number, B puts its number, those numbers do have to be kept distinct. And so there's a function outside of routing, handing out the numbers to make sure you only hand them out once, and you don't make two people think they both claim to be the same unique identity. But that aside, nothing in routing determines certain parts are built. Geoff Huston 22:15 We had a registry run by Stanford research that handed out address prefixes, those bags of addresses, and it handed out what we called autonomous system numbers, these identifiers for networks. And there are very few qualifications to those. You sent a fax because it was the time of faxes over to these folk, and they fax back saying your autonomous system number is this number, and if you ask for address prefix, here's the address prefix, Class A, B or C, and off we all went. And this kind of worked through relatively well. The early days of the expansion of the Internet, when the NSF started work, there were a few 100 such networks. And over the ensuing years, you know, 90 91 92 then NSF heyday, it really did swing up to 15,000 component networks, which is amazingly fast. George Michaelson 23:11 And this mechanism of simple BGP, a path mechanism to detect loops, assignment of unique numbers, this was essentially working, and meant that prior state you've described, the several networks in America, the discrete networks in Britain and Germany, the emerging network in Australia, they were able to come into a union of exchanging information with each other. No central controller. It was just fundamentally working. Geoff Huston 23:37 You just needed to connect to someone who was connected. It was this was that you just added yourself to the edge, and that edge based system was phenomenally successful, brilliantly successful. By 1994 we were faced, and had been facing, for some years, some predictions, which were truly awesome, because all of a sudden this was the consumption problem. We were going to consume all the available addresses. And let me qualify that we weren't really there are 4 billion addresses in IPV4, but there are only only 16 bits of the class B bags, only 65,000 and by the mid 90s, in fact, early 90s, we were consuming them at a great rate. By 1994 we had about two years to go when we'd run out no more middle size networks. We had lots of C's, lots of little networks, but, you know, that was its own problem. We had an astonishing number of those, but the routers weren't big enough to route them all, and that was just going to drown us in noise. We had a small number of the big networks, but there were only 25 .. 127 of them. And so at some point we're going to run out if we ever started using them heavily. So how do we stop this? Well, the answer was, dispense with the pre packaged bag sizes, get rid of the Goldilocks problem. And. Say, look, here's a bunch of addresses that all share a common prefix, and here's the size of that prefix. So you could have networks with 256 hosts, or you could move it one bit to the left and have 512, hosts, or 1024, hosts. You could customize the size of your network. That's a great idea. It scales brilliantly. But this routing protocol, BGP only thought about class A, class B, class C, George Michaelson 25:30 so that magic of how the Goldilocks bags was designed meant you only had to look at top bits, a small number of bits at the front of the address to work out the size of the bag. And you've said, if we're running out of bags, we're going to have to find a way to split things into different sizes, but you're no longer going to be able to use just the front bits to work out how big is the bag, right? Geoff Huston 25:56 right. Every sort of routing object is now a common prefix value and the length of that prefix value. So a Class C network in the old speak was actually the first 24 bits of an address plus the size 24 [George: right]. But that also allowed me to talk about, say, a bunch of four Class Cs as a bunch of 22 bits of common prefix followed by the number 22 which is the length of that prefix. George Michaelson 26:30 Yeah, it's a kind of slightly odd moment for people, because what you're measuring is the length of the thing that is set. And the natural way a lot of people think about this is it couldn't be a count of things that are set but a count of things that are free. But the fact is, we made a decision. It's the count of how long is set. Geoff Huston 26:49 right. So you're not out of the woods yet. You defined a new address plan, which will get us out of the problem of running out of class B addresses really neatly. We're not going to run out anymore, because we haven't got that slop of trying to fit everything into a class B, we can customize. But if the routing protocol doesn't recognize that notation, that way of describing addresses, you're no better off, literally no better off. And so what we had to do was to change the routing protocols for RIP, we had the inventive name rip v2 which included George Michaelson 27:24 good name, Don't knock it, Geoff, that's a very good name. Geoff Huston 27:28 It included prefix sizes. And for BGP, we went from BGP three to the imaginatively named BGP four, George Michaelson 27:36 good name, Geoff, Don't knock it. Geoff Huston 27:39 It included classless inter domain routing, and this is where the acronym CIDR George Michaelson 27:45 comes from, C, I, D, R, CIDR. Geoff Huston 27:48 CIDR. Now there was an IETF meeting, an Internet engineering task force meeting, in March 1994 and the whole thing about let's all run BGP four was given a thorough airing, and at the time, Cisco, who had been instrumental in pushing this out, and was one of the major router vendors of the day, deployed it and pushed it out with their customers. And effects were dramatic and immediate. In early March 94 we had just kicked 20,000 routes in the routing table, class, A's, B's and C's a magic number, because one of the networks, the US military network, Milnet, could only take 20,000 routes, and it just couldn't handle any more connectivity. Oops. In the next sort of eight weeks after that meeting, that number of routing entries dropped from 2000 down to 17,500 because we're able to summarize more effectively, like a whole bunch of little C's George Michaelson 28:48 It had an immediate benefit that it was able to allow you to put forward more efficiently statements analogous to the old bag sizes, but now representing bigger chunks. Geoff Huston 29:00 Er Yes and by doing more accurate and single chunks for larger and larger addresses, what we were having because everyone was making up for the shortfall in Bs with lots and lots of Cs, we were having exponential growth in the routing table. And it wasn't that there was an exponential number of new networks as such. There really was a growth in the number of C networks, because the medium sized networks were expressed as eight CS and 10 CS and so on. And we were trying to relate that by making that a single route object again. And interestingly, in the ensuing years, you know, 94 95 96 the exponential growth in the routing table disappeared and became largely linear again. And we all thought, well, that's pretty cool, but we kind of didn't quite think it through, because George Michaelson 29:54 simplify, yes, simplifying assumptions, where you don't think about all the data. Dimensions of problem and risk. And after you've been running it for a while, probably you now depend on it. You learn that it carries the seeds of its own destruction. Geoff Huston 30:10 You know, when I said the distance, the metric between two adjacent networks is always one, and the issue is I could connect to you with a big, multi mega bit connection, and that had a distance of one I could connect to you with a dial up modem with the speed of, you know, I don't know, a few bits per fortnight, and that would too, have a metric of one. How do I say to the world, please use the link over there. Please don't use this tiny link. It's only there for emergencies. How do I perform what we currently call traffic engineering and load up a network whose links aren't all the same, they're not uniform quality and size? Somehow I want to express that, and BGP says, Well, don't look at me, dude. And the reason why BGP can't do it is that, if I'm trying to express a metric that's common to 20,000 independent different networks, what's the metric? What's the unit? What does it mean? George Michaelson 31:15 Yeah, there's no controller. There's no one in charge. There's no one to say from Monday, if you write this magic value in this is what it means Geoff Huston 31:24 that link is of cost five. This link is of cost 10. No one knows that, and we kind of gave up on that, but we had another trick. And this is again, clever in a, I suppose, useful way, but almost perverse at the same time. So I have a bunch of Class Cs. Let's say I have 16 of them, 16, that's a prefix of length, 20 -16 Class Cs so I have a bunch of networks, but I can also advertise as well as the 20 a couple of slash 24s drawn from inside that first network. What do you mean? As it says, I can take more specific prefixes and also advertise them. You know? Why would you want to do that? Geoff, ah, because of a quirk in the way BGP actually works. Because not only does BGP try to select the route with the minimum metric, the minimum number of networks to traverse it also, and absolutely prefers the most specific. George Michaelson 32:32 So if you talk about something two ways, and one of them is, oh, just casually in passing, there's a million things I know about hosts in this block, but you also specifically say, here's some magic info about a small section of it. You're meant to prefer the magic info if you're walking along and that's the one you're heading to, whatever was said about that smaller chunk that's more important to you. Geoff Huston 32:57 That's more important. So if I want to bias traffic being into me. So if you need to buy us the information about how to reach me, because I've got multiple paths, because, you know, networking is getting serious, I can do just selectively advertising my more specific prefixes that I think are going to attract traffic down links which are either cheaper or better capacity, I can engineer my traffic to match my policies as to how I connect to the rest of the Internet. George Michaelson 33:28 But Geoff, I'm starting to think you said earlier that the network had been growing as an exponential curve and increase in the amount of things being announced, and when this CIDR trick was announced, it magically became both smaller and a bit more linear, and you've now said but we found a way to do traffic engineering if you announce lots and lots and lots and lots more than you were announcing. Geoff Huston 33:53 Yeah. Shame about the temper growth of BGP, because we got into this wholesale of announcing more specifics, and yet again, the race was on that we're actually finding that the routing table wasn't just a bunch of addresses that you can reach, but a large amount of that table size and a large amount of the routing protocol was actually processing refinements to that more specifics George Michaelson 34:22 specific things that are specific to you because you are attempting to manage engineering aspects of how you need to be reached. So the whole of the world has to see this so that you can manage something that's probably quite close to you, but that's the only mechanism we have. Geoff Huston 34:39 That's the only mechanism we have. So it didn't take long or half of the routing table, [George half] be populated by more specifics aggregates in the other half of the routing table, half George Michaelson 34:53 of all the information we shared was this attempt to engineer management of link utilization. Wow. Geoff Huston 34:59 Because it was easier to make, if you will, lots of inadequate links than it was to do a small number of monster links, because they cost a fortune. And so it was quite common practice to have richer connectivity, more connections, but then try and balance your traffic across them using BGP in creative ways. So what happens if everyone does it? Well, we all have to carry routers with larger and larger routing tables. That's interesting. How long have I got to look up an address inside a routing table? Well, you've got the time it takes for one packet because you need to get ready for the next packet to make a look up decision. So as our links increase in speed, we've got less time to do a lookup, and as the number of entries in that table increases, we've got more things we need to look through to find a match. And we're kind of struggling at the point because routing is getting more expensive per unit George Michaelson 36:02 right? So money has entered the room, money expressed as the capitalized cost of making a machine that you can buy now and foreseeably cope with this increase in the rate of speed and the increase in the amount of information. This is not a trivial decision. Geoff, this is quite a consequential one. Geoff Huston 36:24 Well, it's it's a classic tragedy of the commons. I can optimize and engineer my position in the inter domain routing space by slicing and dicing my advertised networks into a fine grained storm of prefixes and selectively advertise it. And indeed, some folk made a business of optimize advertisements to actually engineer across multiple links, constantly changing those advertise more specifics to balance traffic. So I'm optimizing myself, but the cost is everybody. And I mean, everybody else has to carry those additional routes, so no one else benefits just me, and that's the tragedy of the commons in a sort of a simple restatement, collectively, income is bad, individually happy, but only at the cost of everyone else. George Michaelson 37:21 So this becomes a problem that's quite important for people to understand the shape of the problem. How big is this problem? How big is this problem growing? How does it impact me individually, making my capital acquisitions? How do I plan this? That's a classic situation where measurement comes to the fore, isn't it? Geoff Huston 37:41 Well, it is, but I'm going to phrase it somewhat differently. How can you get folk to stop it? I don't mention no one's in charge. So it's kind of well, you know, hit me. There's no forcing functions. Nobody's in control. So trying to put a lid on that behavior, try to say, look, that is really anti social. You know, I can appreciate you've got a problem to solve, but can you appreciate the rest of us have to carry your load? Try and exercise some constraint. George Michaelson 38:13 So it's measurement. It's measurement to a different outcome. It's kind of like the shame list of people who've left empty bottles in the fridge at uni. It's name and shame, isn't it? Geoff Huston 38:24 You've got it name and shame. Try and find the worst abusers of this practice and say, just please how much you're doing to the routing system, how bad your individual contribution is. And look, I hope you, or maybe your people around you chastise you and moderate what you're doing, because everyone is suffering from you. And oddly enough, that is the CIDR report. George Michaelson 38:48 Blimey, Geoff Huston 38:49 Naming and Shaming, George Michaelson 38:50 It's social engineering. Geoff Huston 38:52 Oh yes, very much. So, very much. So originally started by Tony Bates when he was in Cisco, I believe it was after he'd gone through InternetMCI, and taken over by Philip Smith, also at Cisco. And then I ended up picking it up when I was, I was in Telstra, I think, at the time, and carrying it through. And it was, it was a report that simply said, This is how big the routing table is, and this is the daily size for the last seven days. Ooh, look it rowing. And by the way, if we got rid of all those more specifics, if we aggregated like crazy, here's the size it could have been. It's got the same information content, the same reachability content, but here's how much less of a load it could be. Please appreciate this. But the report actually went one further. Not only did it list those networks which were really, really abusing this, it also listed what each of them could do in terms of withdrawing more specifics and even adding one or two aggregates to help that which would have the. Same traffic engineering outcomes would preserve all of their properties of incoming traffic, but drastically reduce their footprint in the routing system. George Michaelson 40:09 So it's both a name and shame, and here's how to get yourself out of the mess. It's a classic the way out of the hole is to stop digging, but adding with it, here's a ladder to lift you out of the hole. Geoff Huston 40:22 Yes, here's a way to actually, if you invest some effort, you can achieve the same outcomes. And this is how. So, you know, please consider this is a way through. So how good was it? Broadly enough, there are actually a few folk who studied the CIDR report as an exercise either in chastisement, Han Nusbacher and Barry green reported to NANOG, because they manually followed up each week. CIDR report got in touch with the big 10 and said, Hey, what are you doing? Can you do better? Here's what the CIDR report says about you, and they reported some limited levels of success, [George: right] But there was actually a student at MIT in 2011 Stephen Woodrow, who got his PhD on the CIDR report, yay. Stephen, who actually did analyze in some depth how effective was naming and shaming. George Michaelson 41:14 Geoff, you're talking about the CIDR report in the past tense, but you haven't actually stopped publishing. Have you? Geoff Huston 41:20 still coming out every day, still coming out. But this report was in 2011 there's some time ago. And Stephen's report was actually interesting. He said, Look, in the first few years it was effective, it really did have some traction with the routing community. And folk did exercise constraint after they had been their attention to be drawn to it, and, you know, lifted their game. But even in 2011 he said, Look, that's declined a lot, and by 2011 it really didn't make much difference. And if you look at the stats, George Michaelson 41:54 but still we go on Geoff Huston 41:56 you look at the stats that more specific good news is it hasn't got worse. The bad news is it hasn't gotten any better. And as we get to 1 million and just exceeded routing table entries in v4 half a million do not add reachability, do not add information, they simply refine the traffic. George Michaelson 42:16 And they could be done more effectively and be a bit more efficient, Geoff Huston 42:21 you'd have smaller, cheaper routers. And what about V6? Well, the answer was, V6 did start with the best of intentions, and in 2004 2005 when, admittedly, the routing table was tiny, only about 20% of the routes were more specific, and there was this feeling even by 2011 at the time of Steven's report that we could do better, and we're doing better, because in 2011 around 20% of the V6 routes were more specific. It was pretty stable, but George Michaelson 42:53 good place to be. Geoff Huston 42:54 When was V6? day, June the sixth 2012 or something? Lets be serious about V6? Well, what does getting serious really mean? It means using it for real traffic. It means using it for traffic engineering. It means Yes, you guessed it, advertising more specifics to actually bias the flow of traffic down your wires in V6. And not only did we achieve George Michaelson 43:18 But surely people don't disaggregate to the same extent, please. Geoff, tell me they're slightly more effective in six. Geoff Huston 43:25 Oh God, no, it's up to 60% worse. Now I don't know if these six were more serious about V6. I don't know what the story is, but we actually were doing worse. That was a peak. By the way. We have got better. We're now down at 58% and that's pretty stable over the last few years. It's really no it's worse. It really isn't much better. George Michaelson 43:46 This is quite an interesting situation, because we've done a historical walk through, getting to the point where we understand why people do what they do, and an emerging tragedy of the commons. And we have three people, Tony Hain Phil Smith and you who carry forward this report the CIDR report that someone in 2011 has said, Yeah, it did work, but it's not so effective now in terms of its impact on things. But you haven't stopped publishing the report, and we still know the shape of the hole and the depth of the hole that we're digging ourselves into that interest me. Geoff, does this still matter? Geoff Huston 44:25 Well, the beauty of the report, of course, is just a program. It doesn't need human intervention. So publishing it is without cost, drawing people's attention to it is much harder, and getting folk to act on it is much harder. But I suspect there's something else at play as well. You see, at the time, 90s, early 2000s we had a network which was used to connect people to service, people to content, whatever was there was over there, and you were here, and the network was the way you got to there. And. What it assumed, interestingly, was that storage, computation, processing, everything that makes a service is expensive and difficult, and comms is cheap, because if I can get your packets over there, you can access the service. We're all happy. You know, Microsoft were busy doing updates of their windows operating system from a massive barn of servers located in Seattle, in Washington and in North West America. George Michaelson 45:29 So routing was absolutely vital to deliver service models that at that time, you did it by using a routing fabric to get there. Geoff Huston 45:38 Routing really mattered. But as it turned out, the underlying assumption, computation is expensive, storage is expensive, you know, mounting services is expensive, communications is cheap. Is the exact opposite of what we see today. Fiber optics has, you know, completely changed things around to some extent, but what's really changed, really changed is Moore's law. Reputation is dirt cheap. Storage is dirt cheap. Why didn't I just replicate this service 4000 times around the entire planet and position my content beside every single major eyeball network says Netflix, says YouTube, says Akamai. And so now we've actually found it easier not to route. And this is this whole argument. George Michaelson 46:27 Oh, that is a strong statement. Geoff, no routing necessary to get the things you want to see and do. Geoff Huston 46:35 Well, if I can get my content really close to you all I've got to go through is the last mile transit, the last sorry, the last mile access. There's no routing. It's just switching. Because cost equals distance squared. If I can deliver that service very, very close to you, it's really cheap because I don't have to go through long, skinny pipes around the world. It's fast, so I can not only take on the old broadcast television networks, but I can up the ante by a factor of 10, faster, cheaper, better by simply locating that everywhere. And so now about 90% of most eyeball networks have their traffic being passed from a data center down the road to an eyeball the other side of the road, and the total traffic, the total distance phone, is tiny. There's no transit. If there's no transit, there's no routing. And so in some ways, the Internet's routing system is a curious historical artifact that a few poor people have to rely on, but everyone else is over it. The vast pool of money has moved onward into data centers and localized delivery, and the CIDR report has accurately documented that in ways that I think we didn't envisage at the time. So we've just moved on from routing, and that's why the CIDR report is no longer fundamentally an important report. It's just a historical artifact, George Michaelson 48:01 interesting, fascinating, but not actually driving the engine room of what makes the Internet special. Geoff Huston 48:08 We thought routing was everything. We were wrong. Interesting. George Michaelson 48:11 So is there somewhere people can go to read the CIDR report on the web? Geoff Huston 48:17 www.cidr-report.org, and there's even a V6 version as well on that page. Yes, it's all there. Thanks, Geoff, thank you, and thank you listeners. George Michaelson 48:29 If you've got a story or research to share here on ping, why not get in contact by email to ping@apnic.net or via the APNIC social media channels, also remember the measurement@apnic.net mailing list on orbit is there to discuss and share relevant collaborative opportunities, grants and funding opportunities, jobs and graduate placings, or to seek feedback from the community on your own measurement projects. Be sure to check out the APNIC website for all your resource and community needs until next time you.