Ilyas Rahimi 0:00 As mentioned, we have a baseline. So at the moment, we do have a baseline of how much bytes per day do we have from the each route. Actually, that that's an established at the moment. And then we would compare that to how many updates each of those bind unbound and not resolver will take each day. So we counted the total traffic would be created by a local route serving locally compared to the hierarchical main way, which actually there is. So it's just kind of looking at the small query-based queries driven daily compared to a local route. How much traffic it costs. George Michaelson 0:47 You're listening to Ping, a podcast by APNIC discussing all things related to measuring the Internet. I'm your host George Michaelson. This time, I'm talking to Ilyas Rahimi, who has just graduated from the University of Amsterdam and Willem Toorop from NLNet Labs. Ilyas did a master's thesis project on the impact of Local Root zone serving. This was supervised by Willem and Marios Avgeris, who is an assistant professor at UvA, where Ilyas studied. Local Root is an operational practice of downloading the state of the root zone from a publication point or in-band inside the DNS protocol. Once the zone has been downloaded and validated, the resolver can know authoritatively from locally held data some queries simply can't resolve because their final top-level domain element, the rightmost one, like .net in the fully qualified domain name www.apnic.net, they can know this isn't in the root zone. It could be a typo, or a misconfiguration, or a deliberate attempt to cause a denial of service attack with reflection traffic. Aside from minimizing the time and the traffic to refuse to serve these unresolvable domain names, local route also permits query minimization. Leakage of what name is being looked up can be kept to a minimum. There is one step in the chain of query resolution that doesn't have to be performed outside the local resolver for a commonly used top-level domain like .com. This won't be much different to having the cached copy locally, but for less frequently used top-level domains, and there are now 1000s of them, it's a small additional benefit. Ilyas's work looked at the impact of different ways of fetching the top-level root zone and the impact that this had on traffic from the resolver under different models of how to fetch, how often to fetch, and the cost benefit of fetching the entire root zone compared to the atomic query through the cache as normal. Willem. Ilyas, welcome to Ping. Willem Toorop 3:02 Thank you, George. Nice to be here again. Ilyas Rahimi 3:05 Yes, indeed, it's an honor. George Michaelson 3:06 People know Willem because he's presented on Ping a couple of times before. But Ilyas, could you introduce yourself to the audience? Ilyas Rahimi 3:14 I'm Ilyas Rahimi. I recently graduated from OS3 of UvA University of Amsterdam, and Willem used to present DNS in the class, and I really loved that. And at some point, I just went to him and I wanted to do a research based on DNS to him, and he offered me this project, which I'm actually really really happy about. And I learned a lot during this project, so that's the kind of the start of this project. George Michaelson 3:40 Willem, your office is on the campus of Willem Toorop 3:43 yes, that's right. George Michaelson 3:44 The university, so although NLNet isn't formally in the teaching framework, you have quite a high degree of engagement with masters and PhD students. Willem Toorop 3:52 Yes, absolutely. The security and network engineering master of the University of Amsterdam is fantastic. They do two short research projects a year, or they used to do that. I'm not sure if they do that anymore. And they are ideal for if I want to investigate something to involve a student, because then there's the obligation that it has to happen, right? This research. George Michaelson 4:19 Yeah, it's part of their marks!. Willem Toorop 4:21 If I need to do it myself, then there's always something else with priority. George Michaelson 4:25 Yeah, Willem Toorop 4:25 but I do do the guest lecture into DNS and DNSSEC. Yeah, on the master. George Michaelson 4:32 So Ilyas, we're getting a context here that this is a master's thesis where you were looking at aspects of the DNS service. Yes. Ilyas Rahimi 4:40 Yeah. George Michaelson 4:41 You gave a presentation on this to the University of Amsterdam as like a thesis defense, and it's on this change of behavior in the DNS, where instead of all queries going to the root, the 13 canonical root servers, people are encouraged to download or access a copy of the root zone and have it locally in the local cache state of the resolver, and you were looking at behaviors of this system. Ilyas Rahimi 5:07 True. Yeah, as you mentioned, the local root serving by default is actually a very hot topic, and we try to quantify the traffic trade off between the queries and the zone distribution. George Michaelson 5:17 How exactly did you go about doing this? Ilyas Rahimi 5:20 Actually, two things to do that we took the RRSIG data, where we have all the data of all the root, the identical IP addresses, and the total amount of the data served. So you can actually calculate the amount of bytes, say per root per day, by taking the identical IPv4 addresses and aggregation of IPv6, combining them with the total bytes which you have, and dividing them to each other, then you will get per resolver a baseline of how much the data is per resolver per day. Kind of George Michaelson 5:50 It's kind of a way of taking the aggregate traffic that's visible at the root and say, if I have this many discretely unique addresses coming to me with queries, I can make a model based on the distribution of those queries across those addresses. Ilyas Rahimi 6:06 Yeah, you have to search in the data. You have the UDP traffic, TCP traffic, and the data coming, data going, and you have to select their the their data which you are actually looking for. And based on that, and we actually calculated the total amount of MBs per day per route, as mentioned at the beginning, it was around two MB per resolver per day. George Michaelson 6:27 Two megabytes per resolver per day. Ilyas Rahimi 6:31 Yeah, George Michaelson 6:31 two men as a volume or as a query count. That's a megabytes measure, not a count of query. Willem Toorop 6:37 Yeah, the amount of traffic. So this is the baseline measurement, right? [George: Yeah]. Question that LEAs asked himself is: How would traffic on the Internet or per resolver be impacted if it would follow local root, so serving the root zone locally? So instead of sending queries to the route, it would do a transfer of the root zone, and so how does that balance out? How many traffic is there currently per letter per resolver to the root compared to what the amount of traffic would then be? George Michaelson 7:14 So, in in this model, Ilyas, how often would a resolver be expected to fetch and update its belief in what the root zone Ilyas Rahimi 7:22 is. Actually, two times a day, but this is the hierarchical way where you will just do the query and the root will answer that. So at this level, updating doesn't matter that much. George Michaelson 7:34 Two fetches a day would represent the equivalent of doing a complete fetch of the state of the root zone at that time, there are what 4000 labels in the root at present. Plus, you have to fetch all the associated signatures, all of the NS's, all of the check sums, all of the key information, all the associated records. It's not nothing, is it? I mean, doing this fetch is of itself quite a large blob of data. Ilyas Rahimi 8:01 That's true. That's the second baseline where we actually compare the data which we actually gathered. I think the count will actually be there on the VMs which we actually run to gather the data of the local root. While the query driven is a small amount of the query which you will actually easily get from the root. George Michaelson 8:17 Did you, in effect, do a replay against the synthetic construct, so you have the real root, an instance, and you are aware of all the queries made to it and all of the addresses of things querying it. Did you create like a test bed and rerun the equivalent of that traffic from a set of addresses to represent that query load under local root. How did you actually perform the measurement of change between the two states? Ilyas Rahimi 8:45 Yeah, for the local root serving is actually we took the four VMs. What they actually do is they are having a copy of of the root instead of querying the root each time. They will have their own local copy of the root. But what is the difference in this setup is that they don't query each time to the root, but they will do often update their local copy. So the difference is actually how often do they copy their local copy, so to say. And by copying that, they had different approaches. We had four PMs, two doing the HTTPS based transfer, which was a full transfer without looking at the SOA record. While the other two, they did an incremental update each time when the route actually changed and the SOA changed. They did an update. George Michaelson 9:31 You were testing real systems, real systems code like the bind code or the NLNet product or knot. This wasn' a synthetic DNS resolver. You were using production code. Ilyas Rahimi 9:31 There's research going on where they actually show you how do even just do the whole configuration, and we followed that lead to configure the VMs, so we don't made any mistakes from there. Willem Toorop 9:56 Yeah, there's an Internet draft currently by amongst others Warren Kumari that proposes to make local root a best current practice. Earlier version had in an appendix different configuration how to configure different software to do local boot, and it has a configuration for bind, for unbound and for knot resolver. And Ilyas installed bind and unbound and knot resolver. The software that could also fetch the root over HTTPS, which is unbound and knot, had two versions: one doing DNS and the other one during HTTPS. I have to correct myself. knot -knot resolver does not support incremental transfer of the root, so it can do only HTTPS. George Michaelson 10:44 This experiment is kind of dividing into two parts in my ideation, my thinking. There is a part which is about a measurement, and I believe from some stuff you've told me, Willem, that there's actually Root Server Security and Stability Committee and the Route Operations Group (RRSAC) have an idea of how they believe measurements should be done for traffic at the route, and so a component of this ILAS, I'm assuming you were looking at providing data that would align with that methodology for counting. Yes, Ilyas Rahimi 11:18 as mentioned, we have a baseline. So at the moment, we do have a baseline of how much bytes per day do we have from the each root. Actually, that that's an established at the moment, and then we would compare that to how many updates each of those bind unbound and knot resolver will take each day. So we counted the total traffic would be created by a local root serving locally compared to the hierarchical main way, which actually there is. So it's just kind of looking at the small query-based queries driven daily compared to a local root. How much traffic it costs? George Michaelson 11:18 This is using real-world traffic, using real-world server code, using a substantive volume of queries that's representative of the volume across the day, and two mechanisms for fetching local root: one in band using DNS and one out of band using web. Willem Toorop 11:30 Indeed, so the so it's comparing. Indeed, assuming that currently no resolvers do local root, which is obviously not true because there's a small number that would do it already. But the root servers provide metrics about their own letters. This is in a report. The metrics they provide it's in report RSSAC 02, and one of the things they provide is yeah, not not so much the traffic volume of each letter per day, but it's the number of queries of certain sizes. So from that you can deduce the number of traffic volume per day, and they also provide the unique IP addresses they see at that letter per day. So if you divide the total number of volume traffic by the number of unique IP addresses seen, then you have an estimate of the number of of the traffic volume per resolver per day for that specific letter. George Michaelson 13:17 It's likely that that would actually be something like a Gaussian distribution, and there would be some number of resolvers that would do phenomenally few queries, and some number of resolvers that would saturate. But in strict sense, Ilyas Rahimi 13:30 yeah, George Michaelson 13:30 that's probably equivalent to an average across all of them. Willem Toorop 13:33 On the other side, the real world measurement that we do is so the four different configurations with actual resolvers configured to do local routes, George Michaelson 13:44 and that was the other half of my bifurcating model. Because one part is strictly measurement, but the other is like operational experience and proof of functionality, isn't it? [Willem: Yes]. You were actually testing these instructions and these systems did what they say on the label, Willem Toorop 14:01 yes, and we then assume or can also observe that no queries will be sent to the root anymore because they can all be deduced from local data. George Michaelson 14:13 Well, this actually is an interesting question because, as a theory, I think that's wonderful. In principle, no queries, but I'm starting to wonder: Is there no leakage? Is there no assertion of queries that bypass the cache state in the resolver? Did it actually zero traffic? Willem Toorop 14:32 There is for some resolvers like bind and unbound. There will be no traffic other than to look up the actual root servers and occasional SOA query to see if there's a fresh version of the root zone, but knot knot resolver has a different approach because they have a different approach which is also currently proposed as an Internet draft, by Paul Hoffman to fetch the root zone over HTTPS and then fill the resolves with cache, right? So if certain names would be evicted from the cache because there is no place anymore, or then it's not used as often, such so, then those names will be queried to the root. George Michaelson 15:20 Yeah, Willem Toorop 15:20 this this is actually not what Ilyas was looking into. George Michaelson 15:24 Ilyas, you have a test bed, you have these resolver instances, you establish a baseline based on traffic, and you therefore have a model with rough edges around some number of local roots in the existing state and the nature of a distribution of traffic load, and you ran your experiment. Can you tell us a little bit about the kind of data you saw? Ilyas Rahimi 15:47 Yeah. So actually, we measure the total amount on occurring on DNS, and what you actually see is quite accurate. Let's say the bind results, the update per day was each time accurate about one point 45 megabytes, and depending on the update, differed from day to day. For example, on June 20 we had two updates, while on 21 and 22 and 23 we had three and four updates. So actually have three updates depending on the time that you start, and that counts actually. George Michaelson 16:17 So it's not a fixed clock cycle zone update. It's a variant depending on need. Willem Toorop 16:23 Well, what we could see is that the root is actually always updated, refined, so to say, twice a day. That's just a plain refined. And if there are updates to the root, like some delegation has new name servers or has new DS record from TLD. [George: Yeah]. Then beside beside those two clients that happens. George Michaelson 16:45 So they don't stage all queued updates until a checkpoint. They do on demand re-sign if there's a need. Willem Toorop 16:53 Yes. So therefore, there will always be two re-signs a day, and if there's a change that day, then there will be a third one, or if there are accidently two changes, what happened during the measurement period that Ilyas did. Ilyas Rahimi 17:08 Yeah, and it did actually differs if you are looking at DNS based update or you are looking at HTTPS based update. On DNS based update, it purely looked at the SOA changes. On the SOA changes, it did an update, while on HTTPS based it doesn't do that. That was also by unbound, but that is whether it was fixed already, right? Willem Toorop 17:28 Indeed. So maybe you should explain first what's wrong with inbound or how naively inbound was implemented, so to say. Ilyas Rahimi 17:37 Yeah, the polling full zone was actually based on 30 minutes standard on HTTPS, and it didn't look if there was a change. So what actually happened at the background each time with each update of two point 19 megabytes, it updated after each 30 minutes. So at the end of the day, you would have 48 updates each day on Unbound HTTPS base. George Michaelson 17:58 Right, the HTTP protocol has a if changed functional behavior that you can in effect get the equivalent of a time serial to know whether or not the thing you last fetched has undergone change. And it sounds like you're saying it just didn't bother looking at that and it just did a prescriptive fetch. Willem Toorop 18:17 That's right. And the root zone has a SOA records telling resolvers or telling secondaries, secondary authoritative servers to refresh or to have the zone every 30 minutes. And for secondary, that means do a query for the SOA record. Has the serial number changed? No. Then you can leave it at that. But if it has changed, then you have to fetch the whole zone. But there is no way to do a SOA query over HTTPS. George Michaelson 18:49 Yeah. Willem Toorop 18:49 Right. If the well, yeah, DNS over HTTPS, but that's not how the root zone is served. I'd say. George Michaelson 18:57 say. Ilyas, you ran this experiment. Is there broadly comparable behavior from all of the resolvers in this model of testing? Noting that unbound had this difference, is there in general broadly the same behavior? Ilyas Rahimi 19:12 Depends again. What we actually, as generally, can establish is that the bytes per day per update that is fixed for each resolver. So that depends on the amount of updates, how much each day happens. That will make the difference in the MBs, so to say. George Michaelson 19:27 I am a true believer in local root. At this point, as a naive party, my assumption is you have been able to demonstrate that there is a significant saving in traffic by deploying local root. Ilyas Rahimi 19:40 No, that's actually the the other way around. George Michaelson 19:44 Seriously, Ilyas Rahimi 19:45 yeah. While local root actually provides some security in dependencies, like if there's a small outage, local root is really good at that. So hiding that if there's a small outage between the root, local root would take care of that. Or if there's a DDoS attack lets say. George Michaelson 20:00 Yeah. So, although this is something that people want to do and that has qualities of preserving privacy, visibility of queries to the root, and potentially has an impact in terms of root server load serving queries then are removed from the system, you appear to be saying that on the client side, people have to do more work. Ilyas Rahimi 20:21 Yeah. Yeah. It is almost two times the amount. But if we put it in perspective, it's much more than that. If we just assume that a 10% of the total resolvers will use local root, at the best part, which is not actually, it would add a 3.8 terabytes per day. In worst case, just adopting 10% of local root by unbound, while the bug is solved, which we previously discussed, it would be around 27 terabytes a day additionally on the local root side. George Michaelson 20:56 Wow! Ilyas Rahimi 20:56 So the amount which it will be adding is really, really huge. George Michaelson 21:01 But within the capacity of the root service system as a whole to serve. I mean, this isn't actually a barrier that means cannot proceed. It's just you can't proceed on a basis it reduces traffic quite the way you might think. Willem Toorop 21:15 Yeah, that's right. This was only looking at the amount, the the traffic increase, so to say, or the change in traffic, and whether it would increase or decrease, as well, with how the root is currently updated, there will be a substantial traffic increase to the root, which they probably can handle fine. But yeah, George Michaelson 21:37 but this actually is a shift in logic because there was a component of belief, certainly on my part, that this would reduce traffic and therefore have beneficial impact in terms of future growth, and that could still potentially be true, but it's not immediately apparent. There has to be a second reason that would justify making a deployment like this. So, what do you think is the outcome that could make it rational to continue with doing a local root activity. Is it in some other aspect of behavior of the system as a whole? Ilyas Rahimi 22:09 I think there comes the magic of Willem. He was looking at signing the zone incrementally, so then instead of one huge update each time, you will have an incremental part of that, which would reduce significantly. So that would be the magic part of it. Maybe Willem can expand on that. Willem Toorop 22:27 That is true. But besides, even if there would be substantial traffic increase, there are still many advantages of having the route local to the resolver. As you mentioned, there's a privacy or increased privacy that the root does not see as many queries as it does see without local root. [George: Yeah] there's reduced reliance on the root. George Michaelson 22:53 So there's a resiliency story in this. There's a potential that this could reduce weak points in the structural behaviors of the system as a whole. Willem Toorop 23:01 Yeah, and it's also if the resolver is verifying the zone MD and validating the signature of zone MD, which is a hash of the whole root zone. There is a security benefit as well because currently referrals in the root zone are not signed. Only the delegation signer resource record is signed, making the DNSSEC chain of trust. But glue and name server data is not fine. So George Michaelson 23:34 right, and Glue is not independently checked of necessity. It's used because it has to be when it's an in Bailewick reference to the subsidiary record being referred to, but it doesn't invoke DNS checks. Where having the zone fetched and the zone MD fetched and all of that computation done means you can have inherently much higher trust in that Willem Toorop 23:55 data. Absolutely, George Michaelson 23:57 there is also a quality that if you use the out of band fetch mechanism. You do still have to fetch data, but you take that data out of the DNS UDP plane and you move it into the HTTP TLS plane. It could come from completely unassociated sources. Willem Toorop 24:16 Yes, I think that's indeed also an idea that I've heard about to let the content distribution network deal with the problem of distributing the root zone instead of the the root serve operators. George Michaelson 24:31 So, Ilyas, did you see further work that could be carried out in this field? Do you think there are other projects that master students could look at in this? Ilyas Rahimi 24:39 Yeah, my research is actually very time bound. I only had one month to carry this whole experiment. George Michaelson 24:45 Yeah, Ilyas Rahimi 24:45 and as Willem said, the research are being changed at our university. They will be getting three months, so that will be much more time. And then you can actually look on the wire. We actually what we did, we looked at the whole data on the wire. We didn't look in depth how much is the pure DNS cost actually, [George: right] Excluding the handshake, excluding all the overhead, maybe that will give us a whole other insight, which I'm sure it would. Removing all the handshakes and all the over wires, it might give us a whole other aspect of what we are seeing at the moment. George Michaelson 25:19 Yeah, I think it's great when you do work like this that has the potential not only to answer immediate questions but also points at further work to be done. I think that's really fantastic. Willem, you're going to follow up with this? You see potential for more? Willem Toorop 25:32 Yes, absolutely. Well, I actually have done some follow-up work myself in what Ilyas already mentioned in incrementally signing the root, yeah, looking at how that would influence the amount of traffic. George Michaelson 25:47 Perhaps that is something we could discuss another time on Ping, because I think that would make quite a nice little story in and of itself. Willem Toorop 25:55 Absolutely, yes, definitely, we can do that. George Michaelson 25:58 So, Willem, Ilyas, thank you very much for coming on Ping and telling us about this measurement activity. That's fantastic. Is there somewhere people can read about this? Ilyas Rahimi 26:07 I think Willem already updated on his page, and he exchanged it actually also in the Google Doc. George Michaelson 26:14 Oh, I'll make sure that that is available on the blog that we post alongside this recording. Willem Toorop 26:19 Yes, Ilyas's of his work is also available on the NLNetLab's website in the research and then student projects. He is right on top of that list there, so you can download the paper there. George Michaelson 26:34 Great, thank you very much. Willem Toorop 26:35 You're welcome. Ilyas Rahimi 26:36 Thank you for having us here. George Michaelson 26:41 If you've got a story or research to share here on Ping, why not get in contact by email to ping@apnic.net or via the APNIC social media channels? Also, remember the measurement@apnic.net mailing list on orbit is there to discuss and share relevant collaborative opportunities, grants and funding opportunities, jobs and graduate placings, or to seek feedback from the community on your own measurement projects. Be sure to check out the APNIC website for all your resource and community needs. Until next time.