Geoff Huston 0:00 So how do I know if you won't or can't or are having trouble getting it back to me? What happens if the answer gets lost? Because the IP is not a reliable protocol. We say to ourselves and remember it: every packet is an adventure. And if you're out there on radio and your mobile and so on, it's an even bigger adventure. Packets do get lost for all kinds of reasons. So if you're in UDP, how do you know that you're not going to get an answer? It's not going to tell you. There's no trailing data. You know, you're just sitting there going, "Hurumph! I'm getting pretty bored now. What do you do? Send it again. George Michaelson 0:48 You're listening to Ping, a podcast by APNIC discussing all things related to measuring the Internet. I'm your host George Michaelson. This time, I'm talking again to Geoff Huston from APNIC Labs in his regular monthly spot on ping, and yes, this time Geoff and I are talking about the DNS again, continuing the exploration of the excess queries Geoff has been seeing in labs, and two questions about how the DNS operates. Firstly, what is the impact of caching and negative caching in the context of DNS queries to unique labels. Is there any way to reliably inform a resolver that the label it's asking for simply doesn't exist, and that related neighbor labels cannot exist either? Would this help reduce the query burden? Secondly, Geoff has been exploring the idea of moving the DNS transport off UDP. What impact might this have on the amount of DNS query and related workload for both Resolver and the authoritative DNS server? Geoff, welcome back to Ping. What should we talk about this time? Geoff Huston 1:59 Well, George, I can't resist. I just can't. I've been poking around in the DNS, and [George: again], again, each time I poke, there's something else down there. This is a remarkable protocol and a remarkable environment. It's kind of like, as a colleague suggested to me, it's kind of like chess. Every single piece on the board has a very small number of really simple moves, you know. [George: Yeah], pawns can do this, etc. etc. But you put it together, and all of a sudden, it's blindingly complex. And almost the same thing exists in the DNS. It's not a complex protocol. You have questions, I have answers. Ask me a question, and I'll give you an answer. You go well. What could possibly go wrong? George Michaelson 2:43 Oh, quite a lot. Quite a lot. Geoff Huston 2:46 Well, the answer is indeed it is very bizarre, but things do go wrong a lot. This particular story actually starts back in 2019. George Michaelson 2:56 Right. Geoff Huston 2:57 What we were looking at at the time was one of the more insidious ways to attack in the DNS was to ask for names that didn't exist. George Michaelson 3:08 So why was that a particularly strong method of attack? Geoff Huston 3:12 Well, they don't get cached, and even if they do, if I ask you for fruzzle bot and then fruzzle bit and then fruzzle bat, you will cache the fact that the first, you know, random name didn't exist. But when I ask for a second, slightly different name, your cache won't work. And so every single time I generate a new random name, DNS cache is not there because it's a random name. So the queries go all the way back to the authoritative server. George Michaelson 3:41 There's two dimensions. If you ask for a name you're confident hasn't been seen before, it has to go to server, and there is a negative cacheing parameter. But you seem to be saying the effect of the negative cache in parameter is either it's not respected and it's thrown away, or the cache lifetime is so short that it doesn't matter, and for a third thing, you're never coming back to that label. So the cache is useless. Geoff Huston 4:05 Well, you're never coming back. Now you kind of go, but that's okay. That's all in the protocol. But you're forgetting something. Cashing is what makes the DNS work. George Michaelson 4:15 Yes, we covered that in one of our last pieces, didn't we? Geoff Huston 4:18 If you don't cache, you're dead. [George: Yeah] Because quite frankly, those points of authority, those points where the answers are, if they are exposed to the vagaries of 4 billion users all the time, things would melt really quickly. Caching is good, but we were exploring this whole issue of the random name attack defeats caching, and we were looking at a particular technology that might help, and at the time we were looking at what they call aggressive NSEC cacheing, [George: right] Where there was this curious feature in the DNS that if you use cryptography and sign your name so that you can test the authenticity, one of the responses was what we call a span. So in your zone, if you have A, B, and C, then if you ask for AA, it doesn't say AA doesn't exist. No, no, no, no, no. It says all of the names between A and B don't exist. George Michaelson 5:16 All of the possible strings that exist starting with A and starting with B in that region, all of them do not exist. Geoff Huston 5:25 So, if I had a zone that only had A, B, and C, then in three quick queries, I actually know that not only do the names A, B, and C exist, but these negative answers tell me that every other possible name you could possibly dream of does not exist, and I don't need to ask the authoritative name server. It's brilliant. George Michaelson 5:48 It's a good technique, but you said this depends on DNSSEC, so this depends on being given authenticated signed statements and saying when you ask questions, tell me the authenticated sign. Say you've got to be in the game. Geoff Huston 6:05 You're right. I will go a little bit further here and go exploring that part of this particular game. of chess is really a topic for another another podcast here. What I wanted to explore was when we actually started to test this and did it live? That we used a mechanism in online ads to seed these kinds of queries against our authoritative name servers. It's almost as if we were DOS-sing ourselves. [George: Yeah], we made 60 million of these queried names, all unique, and that's a big number from all over the Internet, George Michaelson 6:41 you're normally handling somewhere around 20 to 30 million advert experiments a day, aren't you? So making 60 million uniques on top, you're doubling your load. Geoff Huston 6:51 Well, at the time, it was a good experiment, though, and we kind of thought we will see back because we were only just asking for a single type of record. We'll see back about 60 million queries. [George: Yeah] great. We thought we saw 142 million queries. George Michaelson 7:06 Right. So we're now coming close to this behavior. We talked about a couple of EPs back. [Geoff: Yeah] You're doing a thing. You think you'll only see one. You're seeing a lot more. You're seeing two for every one thing you expect to see. Geoff Huston 7:20 To be precise, and you know, numbers matter. 2.37 on average. George Michaelson 7:25 Wow. Geoff Huston 7:26 We thought, okay, that's interesting, and kind of put that away and went on with what we were doing. But we had reason a few weeks ago to come back to this experiment, and over seven days at the start of August in 2026, we did it again. Now this time we ran 115 point 8 million tests because we can George Michaelson 7:46 Because more is better. I mean, come on, what could possibly go wrong? Geoff Huston 7:50 I I love computers. That's right, 115 point 8 million tests, and we found 512 point 8 million queries. George Michaelson 8:00 That's not 2.37. Geoff Huston 8:02 No, it's 4.43. And you sit there and go, has it taken seven years for the DNS to sort of go rotten almost to twice the number of of sort of noise points? George Michaelson 8:14 Yeah, Geoff Huston 8:14 that's significant because this is really something. If something is messy, then it's very messy. So there are some things we need to talk about, and one of the things we need to talk about is this dual-stack network that we're in. George Michaelson 8:30 Yeah, Geoff Huston 8:31 and you know, again, it's a topic of some other conversation. But despite the fervent hopes of some and the resigned, the resigned expectations of others, we're stack in dual-stack forever. George Michaelson 8:43 Yeah, even a network as large as Reliance Jio, which at scale is like 95% pure V6 you see dual-stack events happening, don't you? Geoff Huston 8:53 Well, the poor old clients in Jio still go to services that are singly stacked in v4, so no matter what they would like, they still need to supply some mechanism to get to v4. So there's a dual-stack world. Secondly, to make it more interesting, we're trying to reduce the time from click to result in a browser. That sort of first time response going. I clicked on a name. Why is it taking forever? And that's a decent kind of question. [George: Yeah]. And so what we actually did was use a rather weird record called a SVC record. For those who are familiar with the DNS, it's the HTTPS record, which was the variant of SRVC-B. And the issue is, it doesn't just give you an address, an IP address. It gives you well, what level of protocol should I use? Should I use QUIC? Should I use TLS? Should I use HTTP two? Should I use you know, and so on? Let me tell you about the service in a DNS response, so that you don't waste time. George Michaelson 10:01 Yeah, this is because the higher level protocols incur cost establishment costs, and probing to find out which one can be used means you might actually be chaining up multiple connects to arrive at the one you're going to use. And if it has bootstrapping priming data, you've got additional packet flows before you even start communicating, if you can use the DNS to say hi, I'm here. Talk to me on TLS two. Here's my bootstrapping info. Go for your life. You have short circuited a huge amount of future traffic at a cost of more DNS. Geoff Huston 10:36 Well, yeah. Do you have an IPV4 address record? Do you have an IPV6 address record? Do you have a SVC record? Do you have? Do you have? Now, the astute amongst you of the listeners would say, "Well, come on. Why don't you just use a compound DNS query? Haven't we got any? Why don't you just say, 'Tell me everything you know about this name. Tell me the lot. George Michaelson 10:57 You know, I don't often get a chance to get on my hobby horse and ride it around the table. Geoff Huston 11:03 Mount away, George, clipity clop. Let's go. George Michaelson 11:06 The thing with the DNS is that DNS geeks love to make decisions that are subsequently going to have implications. So any, what a wonderful question type. Any is, but is any the same as all. Well, no. Geoff Huston 11:23 It used to be in a number of implementations that any was the same as all. Tell me everything you know. But pretty soon, folk realized that that's a really good attack. George Michaelson 11:34 We started on attack profiles in the DNS, and here we are back at attack profiles. Geoff Huston 11:40 If you look at some of these records, and I remember BBC BBC.co.uk, if you asked it for a text record TXT, you see a whole suite of authentication answers. It's kind of crypto, but there are a lot of them. So one single tiny query generates a huge number, an amount of answer. [George: Yeah] any is an attack. George Michaelson 12:04 It's a multiplier effect that boosts. If you can make someone think the question has come from over there and send this tiny query over there gets flooded with a massive answer, Geoff Huston 12:18 and therein, dear listener, lies the recipe of a very effective number of DDoS denial of service attacks that have been used for decades, and because it's UDP, it's actually really hard to stop them. But I wanted to get back to this issue that the way we counted this is: you sort of go, "Can I do compounds? And the answer is no, you can't. You can't say, "Give me the A, give me the V6 the quad A. Oh, and give me the SVC record, the HTTPS. George Michaelson 12:44 Yeah, Geoff Huston 12:45 the DNS says not RFC9619 1 per query. Slow down. George Michaelson 12:52 No, this is what we do. We do one. Not here's a list of six things to tell me. Ask one thing. Geoff Huston 12:57 Now, if you were really silly and you were living in an age where you know continental drift was considered fast. You would ask for the A, wait for the answer, and then ask for the next one, the quad A, the V6 and wait for an answer, and then ask for the SVC record. That's kind of crazy. George Michaelson 13:13 Australia has moved 10 centimeters to the left. Geoff Huston 13:17 In the meantime, you know people don't live that long, so what we actually do is gang them up. So you'll see in today's world we launch a number of queries with the same query name. Bang bang bang. Give me the IPV4 Give me the IPV6 Give me the service record. Really fast. And now let's go back to our problem because we've upped the number of queries with the same query name by almost double 512 million over you know 115 million tests or whatever it was right. George Michaelson 13:48 The effect of this is a magnifier and it has potential to represent a road to a DDoS, but compared to a response that is a ratio of maybe 50 to one or 100 to one, this is smaller. So it's not great, Geoff. But as DDoS attack vectors go, this one is not actually that good. Geoff Huston 14:08 Well, what I'm saying is, let's first understand what's changed in seven years to see why [George: yeah] that number has gone up. This repetition of queries with the same query name V6 has almost doubled Q to V6 advocates to cheer madly from the sidelines. Yay! It's gone up from 24% of users in October 2019 to 43% of users yesterday. So you know, cause for celebration in some way. There's more dual-stack What does that mean? George Michaelson 14:36 More DNS. Geoff Huston 14:37 There's more quad A queries as well as A, so we've almost doubled the amount of people who have V6 who then want to ask for the V6 address, but you don't know if the thing you're going to has it or not. So to be safe, to be sure, you ask for both the A record and the quad A record v4 and V6 at the same time. You couldn't use the same query RFC 961, you know. So you've got to use two queries back to back. So that's one cause. George Michaelson 15:07 And then this additional thing that you said was a variant of the SVC record. Now that's not about DNS. That's about what web protocol I should use, right? Geoff Huston 15:18 Well, it's predominantly used for that, but it can also be used for aliases prioritization. I'd prefer you to go to my primary server, but over here for you, go to my other server. Oh, actually, don't use this name; use that name. So there's a whole bunch of things you can do with this SVC record, and they're very good. But it didn't exist in 2019. Not exist. There were no queries for HTTPS other than maybe some dedicated enthusiasts on their local machine. But today, yesterday, 32% of users query for this HTTPS record, George Michaelson 15:55 and this is a query that's functionally embedded in the browser. Yes, so this is the browser technology releasing this capability. Geoff Huston 16:04 Oh, things get complicated. There's a library on most machines, a DNS stub resolver library, we call it, and the application says to that library, "I want to know a V4 address. I want to know this. I want to know that. I want to know an HTTPS record. That gets passed to the local library, which then passes it across via the network into the DNS system, and the query takes on. George Michaelson 16:29 So back then, the browser didn't ask the local library, "Can I see an HTTPS record? But now, Geoff Huston 16:36 wasn't there? Yeah, George Michaelson 16:37 30% of them. Geoff Huston 16:38 32% of users asked for that. George Michaelson 16:41 Wow! Geoff Huston 16:41 So with our 115 million tests, if I take those proportions of V6 users and HTTPS aware, etc. and I do a very naive sort of add them up, I get to 217 million queries straight away. So if there was no repetition whatsoever, you ask for a V4 and you ask again. If there was none of that, we'd still have 217 million queries. George Michaelson 17:07 But that's a gap from the amount you saw. Geoff Huston 17:09 You're missing 300 million. George Michaelson 17:12 What's 300 million between friends Geoff Huston 17:14 between friends? Yes, it's just the DNS. It can cope. 57% of the queries are repetitions, and so now we get into the transport issue. Why do you repeat a query? Now there's a certain amount of what would I call it theology. Some people worked it out about 30 years ago, and we trust what they did. So now we just repeat it to ourselves without going back through their working. We just accept it. It's almost an article of faith, and one of those articles is that UDP is lightweight, efficient, and really, really good for short transactions. And the DNS is a short transaction application. George Michaelson 17:54 Now you have been, if I may use the metaphor, singing in that choir quite substantially, Geoff. You've often said it's the protocol of choice when you want to get something done fast, and you don't need to establish and maintain a long-lived connection. You've been there. Geoff Huston 18:12 Absolutely. Think about this for a second. TCP to protect both ends uses what we call a three-way handshake. The client says, "Hi, I want to speak to you. The server says, "I acknowledge you want to speak to me. Here is a token. If you get back this token, give it back to me, and we will start exchanging data. So the exchange works in protocol shorthand: SYN, SYN ACK, ACK, three-way handshake, and that takes one round trip time. Now, if you're across the room one millisecond, it takes one millisecond. But if you're on the other side of the world, it takes some time, right? So that overhead doesn't exist in UDP because UDP is just I send you a packet, you send me an answer. Well, did I send you that packet? Well, I'm going to send an answer to whoever's in the source address of that IP thing I receive. I don't know if it's you or not. We haven't shaken hands, but I don't really care. There's no state. You send me something, I send it back to you. That's it. [George: Yeah]. So how do I know if you won't or can't or are having trouble getting it back to me? What happens if the answer gets lost? Because the IP is not a reliable protocol. We say to ourselves and remember it: every packet is an adventure. And if you're out there on radio and your mobile and so on, it's an even bigger adventure. Packets do get lost for all kinds of reasons. So if you're in UDP, how do you know that you're not going to get an answer? It's not going to tell you. There's no trailing data. You know, you're just sitting there going, "Har, I'm getting pretty bored now. What do you do? Send it again. [George: You gotta]. You have a clock. The issue is, though, I haven't done an initial handshake. I have no idea. Dear, how far away you are! You could be God on Mars. Wow, that's a round trip time of some some hours, if not days. Or you could be across the roof. I don't know. You're just a IP address. So not only do I need to have a timer, but the timer needs to be guess. [George: Yeah] because you don't know what it should be, right? Now, what if I'm a conservative person and I I feel that every packet is worth handling carefully, nurtured, and duplication is evil? Then what I will do is use a pretty long timer to give you leeway in answering me. Now I look at something like a geostationary satellite, where the round trip time up there and back, up there again and back, is two thirds of a second. And I go, well, you know, if you don't get back to me within a second, let's call it dead. George Michaelson 20:50 Wow, a second in the context of roll the clock 15 minutes when we talked about reducing the delay between clicking off the name, getting the answer. Wow, you've just blown that out of the water. Geoff Huston 21:05 A second doesn't work this this time, does it? It's a continental rift timescale. George Michaelson 21:09 I think you need to treat your packets a bit more. Try again. This no repeats don't work for me. I'm repeating. How about Geoff Huston 21:16 one millisecond? George Michaelson 21:18 It's very quick. Geoff Huston 21:19 I'll wait for a millisecond, that's kind of you're on some weird drugs because almost nothing other than if you're in the dave same data center is just a millisecond away. George Michaelson 21:28 Yeah, Geoff Huston 21:29 even to head out into the big band Internet and cross a city, you're going to spend at least 10 milliseconds getting there and back. George Michaelson 21:36 I could imagine because I'm good at imagining. I could imagine protocols putting their hands up and saying, "I need to have some hysteresis of the context you typically operate in. If I know that you're speaking the Mars, I'll whack the timer up. But given I know you're sitting in your home office in Brisbane and you're speaking to the local data center, I'm going to pick a number reflective of prior history. But at this point, we've now changed from every packet is an adventure to your past history has to influence making you the right shape to fit. And the easier path is just pick a number, dude. Geoff Huston 22:13 I'm sorry, I'm writing some code for a recursive resolver. Now that keeps on firing UDP packets, and all of the authoritative servers out there. Ultimately, you know, you don't know where the next packet's going. So there's no repeated factor. You're not asking the same authoritative name server again and again and again. There is no history. Now I'm writing code. I'm not in a context. I don't know if it's going to be used in a mobile device, in a laptop, in a server, I have no idea. I'm writing code for anything, everything. George Michaelson 22:49 You got to pick a number. Geoff Huston 22:50 I reach for my dartboard and my dart. I close my eyes and I throw it. George Michaelson 22:56 Where did it land? Where did it? Geoff Huston 22:58 Where did that land? Let's use that number. And it's kind of there is a sort of if it's too short, then you're going to fire off a query before the first answer had a hope of getting to, even if the Internet was perfect. If you pick too big a number, the user is going to be sitting there going, "God, this is so slow today. I think I need to change my recursive resolver because this one is crap. You need to somehow get a number that's sort of like Goldilocks- not too small, not too big, but what's just right, nobody knows. So you need to figure out something. What do we err on? Well, this is interesting. If I had to pay per packet, I would be very, very conservative because the poor old user is going to get landed with a huge bill if I'm very aggressive in unnecessarily repeating packets. If George Michaelson 23:52 If there is a constraint, it's a forcing function in economics to say avoid payment that pushes your number one direction. But Geoff, as we have said before, nobody's paying up front here, right? Packets are free. Geoff Huston 24:06 Packets are free. So, of those 500 odd million packets that we observe, one particular query over 24 hours was repeated 38,679 times, and it's kind of yeah, I guess it was free, but you know, really, dude, really, are you serious? George Michaelson 24:25 Are you sure? Geoff Huston 24:25 Thankfully, George Michaelson 24:27 Are you sure? Are you sure? Geoff Huston 24:29 Every two seconds, [George: Oh go on,] incredible. Computers can be like that. Yes, what did you say? What did you say? What did you say? We actually started looking at this, and we actually found that it's not as bad as you think. It's not everyone is crowding out there in the stupid case. 60% of these repeated queries, if I see it twice, and 60% of the time I only see it twice. 20% occur three times. 10% slightly less. 8% four times etc. So there's a pretty quick drop off, and most of the time repetition is actually mildly well behaved. And so the real question is, well, what are those timers? Well, it depends on what code you're running, both in the recursive resolver and in the stub. George Michaelson 25:15 So if I imagine what this might look like as a measurement guy looking at a chart, it's going to be a classic stacked bar, and you've got time on the x-axis, and there's this huge column somewhere on the left, which is the aggregate most people are in this space. And there might be a couple of extra bumps further out that are outliers that, for whatever reason, just seem to be out there doing weird times, but it should be strong signal once, couple of blips, and then long tail. That's what I'd expect. Geoff Huston 25:48 Every time we see a new query name coming at us, we start a clock on that name. George Michaelson 25:55 Yeah, Geoff Huston 25:55 and if we see another query for the same name, we say "dunk, George Michaelson 26:00 stop the clock. Geoff Huston 26:01 And no, don't stop the clock. Keep it running, but mark that interval, yeah, and record a plus one of that interval. And now draw a graph of all these times that we got hit after the initial query. So yes, what we do is then we have these sets of how long did it take for a repeat query to arrive, and we keep on listening for up to 24 hours each day. Now, oddly enough, George, if you listen to the DNS for long enough, you will find repeats of queries that are months or even years old. But again, that's a different topic. We call them zombies, and they honestly have taken life of their own. But let's look at the, if you will, more more prevalent repeats, and there are a number of timers that people have just said, "Oh, it's a good time. I don't want to think anymore. And the most obvious one is one second. One second is conservative, as we've said. People don't wait for a second, so it's sort of safe, but not very fast. George Michaelson 26:57 So a second is also more than the path, the RTT, the round trip between me and a geosynchronous, and it's probably more than the end-to-end delay for me to send a packet pretty well anywhere on the planet. Geoff Huston 27:11 You can go around the world on fiber in a second, no problem. In fact, I think you can probably do it twice. You know, if that's what you wanted to do, unlike an airplane where it's going to take you 36 hours inside a second, zooming, zoom, zoom. 300,000 kilometers per second, I think, is the speed of light in a vacuum. So you know you move. So what times do these implementations use? So this is a little bit interesting in terms of folks sitting there in their darkened room with their screens, dreaming up a time, and that each implementation has dreamt up something different. The most aggressive one is 270 milliseconds, just over a quarter of a second. And you kind of go, well, what's that? Well, that's sort of halfway around the world on fiber. It's sort of Europe to Asia and back again, or Australia to America and back again, and a bit more. So 270 is the first of these kind of. It's a time. It's it's longer than you'd see in most cases. It's not bad. Then there's 370 milliseconds, just over a third of a second. Similar, one that's prevalent, oddly enough, is 800 milliseconds, which I actually think is long. George Michaelson 28:20 Oh, that's very long. Geoff Huston 28:21 It is long in this kind of day and age. And the next one is one second, which used to be a favorite of Microsoft software. Kind of, who cares about time? You're not impatient. I'm not impatient. Let's wait for a second. And bizarrely, and I'm not sure where this has come from, but there was a slight peak at 1.2 seconds. So, what we have is resolvers in UDP repeat the query, and the ones at 270 milliseconds and at 370 milliseconds have a higher probability of jumping in before the answers come from the first one, they might be a bit aggressive, and in some ways it doesn't matter because if the first one was going to come anyway, you'll take the answer you get and keep on going. And if it's lost, you haven't lost too much time. You've only lost approximately a third of a second. George Michaelson 29:16 Now you're kind of in an interesting place there because we started this conversation, saying how do you pick a number, and you can't use it based on past transactional knowledge. The protocol has to be written for every case. You're picking it based on beliefs about the right choice, and you said the one that you're seeing predominantly very slightly precedes you completing getting the answer back to them, so the round trip time-that's a millisecond measure in a certain scale-work time to actually find an answer and commit it and send it incurs more delay. Geoff Huston 29:52 Yes, George Michaelson 29:53 and that delay-it appears the retransmit timer is kind of edging close to a race and has got in front. You'd started to reply. You'd said, "Here's an answer, and they're already, at some extent, saying, "I'm impatient". Geoff Huston 30:09 Let's look at this for a second and kind of uncover some more detail here. I send a query and wait for a quarter of a second, and you're under a lot of load. You're a busy lad. It might take a third of a second to generate an answer. So you discover the answer and send it on its way back to me. Now it might well take more than a third of a second. Why? Because you might not have the answer in your cache. So you're going to have to do the whole. You know, what do I know? Which is the right server to ask. Hang on a second, I'm working here. Hang on with it. And I quarter a second later go. I'm getting impatient. I'm getting impatient. Have another query. Now n the best of all possible worlds, the answer is already on its way back to me by the time I send the second query. It doesn't matter. I'll take the first answer and go. Thank you very much. Even if the answer isn't ready yet, the second query is a repeat of the query the resolver is working on. So what the resolver says is, "Dude, hang on. I'm working on this. I am not going to answer you until the first task is finished, and I will answer your second query with my cache, and I will basically send two answers back to back. George Michaelson 31:27 So, if I had to do work to go ask someone else, the second query should not make me go ask someone else. I localize all that cost in me and go answer answer. I use cache. Geoff Huston 31:40 Yes, exactly. So it's it's kind of a victimless crime. You have a bit more traffic in the DNS, but it's not harmful. George Michaelson 31:50 Theory Geoff Huston 31:51 the theory, and we'll return to that. So this is why folk tend to go short, not aggressively short, because that would cause meltdown. But a quarter of a second, a third of a second-it's not bad. It's not bad. So what we see are peaks where the interval between receiving the first query and receiving the second tends to be at this area of fractions of a second, you know, third of a second, etc. But that's not all we see. Oh no, no, no. What we also see are two queries back to back within 10 milliseconds, and it's kind of well, well, hang on. Are you saying you're not even waiting at all? You're just banging, but but but that's not. And we see a lot of that of the repeated queries. Around about 60% or so occur within the first 10 milliseconds, and we talked about that last time. This query duplication, bang bang. So UDP is part of the issue, but the other side is a more insidious issue, and I think it's about scaling. I think it's actually about DNS farms. I think it's the fact that when you've got a front end and a bunch of worker bees scheduling which worker bee takes on which query and how you handle the answers is actually really quite a difficult problem. There are no standards, and when there are no standards, engineers get creative, and that's the last thing you want engineers to do on you because they're really bad at that. George Michaelson 33:21 This is what the people who own these assets would call their air quotes secret sauce. How they shard, how they scale, how they coordinate cohesion, so the single site acquires a consistent cache. In their heads, they'd be thinking, "Oh, that trick we do where we go bang bang. That's really good for us. That's our air quotes secret sauce. So why would they standardize it? They think it's probably a differentiator. Geoff Huston 33:50 Let's talk about this again. Put on the brakes and go back to what we were looking at. We were looking at an area where the answer was that name does not exist. Annex to main. You've asked for something. I don't care if it's a v4 record or a V6 record or whatever. George Michaelson 34:06 Doesn't exist. Geoff Huston 34:06 That name does not exist. So if you've got other questions about that name, forget it. George Michaelson 34:12 You're wasting everybody's time, including your own. Geoff Huston 34:15 Right. Stuff the fact that name doesn't exist in your cache. Turn off all the other pending queries were overdue, but the DNS is weird. You can say no a whole bunch of ways. There is the weird no. Why would you do this? Because in DNSSEC land, no is big. I need to give you signatures of the start and the end of the span and the signature of the authority record. That's three signatures in one answer. And if I'm using, well, let's go overboard and go RSA 4096 or some other enormous enormous key. The answer is I've just blown my UDP size budget. You know, I can't do that. George Michaelson 34:59 Yeah. Geoff Huston 34:59 So. NX domain can be a big answer. George Michaelson 35:02 Signed NX domain incurs cost on everyone. You, the sender, you've got to compute SIGs over things, send them. Me, the receiver, I'm meant to check SIGs, and I've then got to do compute work on that to do the crypto. And I had to receive a big data stream. So I might say my secret sauce is not doing this trick, Geoff Huston 35:24 and and there's more because in your mind and mine is this model that when I'm generating an answer, I have access to the entire zone. In our trick case, it was I know that the zone contains A, B, and C. But what if I didn't? What if I'm a front end and you've asked me about BB and I know it doesn't exist, but I don't know what does exist? What I want to do is kind of go look. I'm going to punt on whether BB doesn't exist or not. I'm not going to tell you that, but you've asked for an IPv4 record, an A. I will tell you there is no A record. What about the others? Oh, ask. You know, not going to tell you anything about the others. It's called no error, no data. So it's not that there's no such domain. I'm not going to tell you that. [George: Yeah] I'm just going to tell you that what you've asked for isn't exist. The query type. George Michaelson 36:15 It's like that very old game that was the desktop game of Mastermind. [Geoff: Oh yes], real Mastermind is 20 open-ended questions. But the desktop game, you are hunting a specific answer. Cow with a brown side on the left, visible from a train on a Tuesday at six a.m. No, sir, not a car on the right side. On the left side, and you're in that space. You're not saying there couldn't be a SRVC-B record for this thing. You're just saying there's no four here. Geoff Huston 36:47 Have you any sixes? Well, I'm not going to tell you whether they have any cards or not, but I haven't got a six. [George: Yeah]. Have you any sevens? You know, and that's the kind of game that the DNS is playing with no error, no data. It's very specific. What you asked for is not here. [George: Yeah], I'm not going to tell you about anything else. That's gratuitous answer. I'm just going to tell you what you asked for isn't there. No error, no data. There's two other. In fact, there are three other ways of answering you. The next one is I don't feel very well. I'm going to not answer you because I'm failed. SERVFAIL. Come back later. George Michaelson 37:19 Yeah, the Zen answer. The Zen answer of I do not answer. Geoff Huston 37:23 Well, this is weird. This serve file says I'm not feeling very good for whatever reason. But what it does say is I'm not going to tell you for how long. So you might want to ask me again, and I might feel better. There's another error code which I love is REFUSED. I'm feeling fine, but I don't like you. It's you, not me. George Michaelson 37:41 I'm doing some DNS measurement work for you, Geoff, and one of the things I've been fighting is the level of REFUSED because I'm issuing 567, 800 queries a second to do DNS measurement, and the back end service are going, mate, no, Geoff Huston 37:57 mate, REFUSED It's not me, it's you, and it really is you, and I'm shut down for you. George Michaelson 38:04 In this space is the interesting behavior of EDNS0 extended error messages because in Hypothesis you could do a REFUSED and give me an EDE code saying rate exceeded or an EDE code saying outside policy or a number of things, you could refine your refusal. Is anyone looking for those extended codes? I don't think so. Geoff Huston 38:30 Will it change your query behavior if I come back with an error code going? It's a Tuesday. George Michaelson 38:35 No, it's only for diagnostics for geeks. It's not like coders are going to use it. Geoff Huston 38:40 Otherwise, called a waste of time. So no, I'm not going to give you gratuitous information. A no is a no, George Michaelson 38:46 right? Geoff Huston 38:46 And REFUSED is one of those no's, a bit like in next domain, but it's you. George Michaelson 38:51 Yeah, Geoff Huston 38:51 I might be answering other people. I don't care. I'm not answering you. And the last of the no's is actually again a classic UDP. I heard your query, but I'm going to tell you that. I'm not going to tell you anything. Silence. [George: Yeah] which is weird. So again, we rammed about 120 million queries against one of our servers for each of those cases. So for NX domain, the DNS is remarkably good, and if you think about it, the name in a query can be interpreted as micro code to instruct the server how to behave. So, if you use a certain name with one of our servers, it will always respond with NX domain. The same server will respond with no data if you give it a subtly different query name. [George: Yeah] you're changing the implicit instruction code. George Michaelson 39:42 This was used for part of the verification of DNSSEC, wasn't it? They created namespaces that responded based on the nature of the query. Back in the day that I was involved in writing code that had to process Visa cards, Visa and Mastercard published card numbers, and sa. If you use these cards doing merchant tests, we will accept it from this one. Nobody does money and refuse it from that one. Nobody does money. You can test your software with one each of a good card and a bad card, and that's kind of where you are. Geoff Huston 40:15 Right. That was where we are in programming these responses. So we didn't need six different servers. We use one off we go approximately 120 million in each case, but the number of queries we got, wow! You know, we thought NX domain was kind of bad, 4.43 queries against the non-existent name. But if you analyze the number of repeats, just under 300 million, that's a ratio of 2.5 repeats for each query. Okay, what if there's a real answer? Yes. Interestingly, it's 2.7, slightly higher. Why? Because in the NX domain case, if you're going to ask three queries and you get back the answer to the first before you ask the third, you're not going to ask the third one. You know the name doesn't exist. You're going to short circuit the process. So I'm not going to ask for a quote a HTTPS. If I get an answer from the a already that says forget it, dude, no name, I'll chunk the rest. So it's slightly lower. It's not that much different. How about no data? It's there, but it's not you. It's 3.2 repeats on average per query, slightly higher. So. Are you sure? I'm going to try a different recursive resolver because I don't trust that answer. So it's slightly noisier. How about refuse? George Michaelson 41:34 Interesting that it's less than four because my personal belief would have been if you only have one resolver, you'll try at least twice, and if you have two resolvers, which I have been told is on average the number of resolvers people configure, you'll try each of them twice. I would have expected you to see an average somewhere in the fours, and you're saying the average was somewhere in the threes. Interesting. Geoff Huston 41:58 This whole thing about who's generating the repeats is its own little, you know, weird case. But REFUSED? Oh, people don't like REFUSED. You're refusing me. I'm going to ask you again. Was that real? How dare you refuse me? We get 9.9 repeats. George Michaelson 42:13 Oh wow! So the one that means authoritatively, not you, buddy, actually makes more traffic emerge. Geoff Huston 42:20 Yes, but it can't be me. I'm going to ask you again. It can't be me. I'm going to ask you again. I'm going to ask you 10 more times. George Michaelson 42:29 Wow! Geoff Huston 42:29 Before I give up, on average, George Michaelson 42:31 that's very surprising. So, doing the right thing and saying I can't handle you causes yous out there to Geoff Huston 42:39 get very aggressive, George Michaelson 42:41 whereas totally ignoring you, la la la la la. I'm not listening. [Geoff: Oh, hang on]. Only has a slighting. Geoff Huston 42:48 No, no. I said no data. George Michaelson 42:50 Oh, no data. Geoff Huston 42:51 I'm listening, but what you asked for doesn't exist. George Michaelson 42:54 Yeah, Geoff Huston 42:55 I'm responding to you. George Michaelson 42:56 Right. Geoff Huston 42:57 So let's go to the two that are weirder. Sir, fail. It's not you, it's me, and you think politely. You can't take umbrage at that. You can't say, "Well, you know, it's the server. The DNS goes. George Michaelson 43:09 Are you better yet? You said you were really. You're better yet. How's that headache? Are you better yet? Geoff Huston 43:13 84 times. George Michaelson 43:15 Oh my god! Geoff Huston 43:16 It's amazing. George Michaelson 43:17 84. 84! Geoff Huston 43:19 84. times. George Michaelson 43:20 I'm going to refuse you. Refuse is better than SERVFAIL. No more SERVFAIL in my life, buddy. I don't care if I'm sick. It's you. Refused. Geoff Huston 43:29 141 million unique names. 11.7 billion queries. That's kind of guys, guys. [George: Wow] It's me. Stop hitting me when I'm down. I'm down. Didn't you hear me? George Michaelson 43:40 The last one. Silence is golden. Geoff Huston 43:43 Silence. Oddly enough, in our numbers, it was a mere 82.5 repeats, slightly less than SERVFAIL. If you say nothing, it's subtly better than saying, "I don't feel like answering. I'm sick. George Michaelson 43:58 Well, you might be on Mars, Geoff. I don't know where the distance is. I know the number of AS hops, but I don't know where you are. So I'll try again in case you are actually in orbit. Geoff Huston 44:07 The observation kind of is that there are ways and means of saying no. George Michaelson 44:12 Yeah. Geoff Huston 44:12 And generally, not answering is a disaster. Just not answering is the wrong thing. George Michaelson 44:17 Right. Geoff Huston 44:17 And even refuse is really not a very good approach. I don't care about the EDE errors. George Michaelson 44:23 SERVFAIL is a disaster as bad as not answering. Geoff Huston 44:27 It's better to say the name exists but not what you're after. [George: Yeah]. If you want to keep the response size down for sign names, or just to simply give a conventional NX domain response. That's better than any other way of doing this. George Michaelson 44:39 And if you have the CPU, and if you believe the clients were going to handle it, the signed extent, the span rejection does have some ability to prevent other queries, but it incurs cost. Geoff Huston 44:52 It incurs cost, and no data, that kind of negative response doesn't give you the span. So what you win in the immediate. No, you lose in terms of defense against our random name attacks. Everything is a compromise. [George: Yeah]. So let's go back into a few general observations here because I want to hammer in this control. Even saying yes on the whole the repeats. Now I've already asked the A and the quad A and the and the HTTPS. That's fine, but I still see out of 263 million names, I see 712 million repeats because the total query volume out of those 263 million was 1.4 billion, and that's the A's, the quad A's, the HTTPS's, and 712 million repeats. So the DNS is sloppy. It over asks, and a lot of those, almost 60% occur back to back. They occur even before they see the first answer. It's kind of to be sure, to be sure. Let me speak in doubles, or let let me me speak speak twice, twice is what's going on. George Michaelson 46:02 I have this thought, okay, and it's triggered because of the context of why are we in UDP and why would you stay in UDP? And that Geoff Huston 46:11 let's go there. George Michaelson 46:12 That conversation. I'm in UDP to avoid overhead, to avoid establishment, to avoid end-to-end signaling. And you've just gone through 20 minutes of explaining. Even in UDP, I'm going to hear from you two times, three times, four times, five times. Geoff Huston 46:27 UDP is incredibly inefficient. [George: Yeah], it's typically 2.7 repeats. So in actual fact, the true query volume is less than a third of the received UDP queries, George Michaelson 46:41 yeah. Geoff Huston 46:42 So UDP is very noisy. Now, other people have done a very, very different study where they take a log of captioned queries, real queries, and they play it against a recursive resolver using a UDP transport. It's as if they're a new recursive resolver, and they jam things through as fast as they can. Then they take the same query log using live Internet and use TCP as their query transport. And the reports, and there are quite a few of them, say you'll get about a third of the throughput. So if you use TCP, the same hardware, the same software, the same system that's doing the recursive resolution will operate at around 1/3 of the efficiency of UDP. George Michaelson 47:21 That seems very expensive. Geoff Huston 47:24 Well, there's a lot of process control blocks. There's the three- way handshake. There's the delay issue. You know, yes, that's quite a reasonable outcome in terms of TCP is a higher overhead if what you're doing, you know, is simple transactions because I've got to set up the session, ask the query, and tear it down again. George Michaelson 47:43 Oh, so you're encapsulating the equivalent of a UDP transactionless event into a complete closed TCP event? Do the thing, do it, end the thing. Geoff Huston 47:54 If I could set up a single TCP session and send all of the queries in the same session, George Michaelson 47:59 yes, Geoff Huston 48:00 per query, it's as fast as UDP. George Michaelson 48:02 Yes. Well, I mean, let's have that conversation. If I know, Geoff Huston 48:08 oh, okay, we'll try and do the ultimate form of centrality, George. That every single domain name on the planet is actually hosted on the same authoritative server. George Michaelson 48:17 Well, hang on. Let's come back a bit. I'm asking you a four, a six, an SVCB, TLS bootstrapping, and some other things. Shouldn't I am amalgamate those queries into a single TCP binding? Geoff Huston 48:32 Yes. So you're starting to get down into the area that I find fascinating too. That a lot of the UDP swap is based around UDP. I don't know when an answer is going to come. It's a defensive statement to back to back the queries because I'll get better performance. It loads up the server. George Michaelson 48:52 Yeah, Geoff Huston 48:53 loads up the server. And the real question is, when I say loads up, what do I mean by loads up? And the data that we're seeing here tends to indicate that that loading factor is getting close to three, three, and the reason why is you don't know if you've got the question and there's an answer coming. It's UDP. You can't tell. So UDP introduces approximately a 60% inefficiency. Vague number, right? But you've just told me that if I ran TCP, my throughput drops down to 30% But I know you've got the question. It's TCP. If you're not going to give me an answer, you're going to send me a reset at some point. I know what's happening here. I don't have to guess. I don't have to open a new TCP session. You don't see in the web an unresponsive web page going "Hello, let's start again. Hello, let's start again. Hello, let's start again. You don't do that. George Michaelson 49:48 No, because the TCP signaling says this silence exists because I know and you know that we've said what had to be said. You asked, I told, you acknowledge. You'd seen the answer. You told me you've seen the answer. Geoff Huston 50:04 So let's use TCP for the DNS. Even though it's only 30% efficient, the answer is but UDP is 300% inefficient. You know, 2.7 queries are just repeats of previous queries. George Michaelson 50:19 Maybe this equipoise is. Maybe we actually arrive at it's a similar quality of service. Geoff Huston 50:24 Well, it's better. George Michaelson 50:25 It's better. Geoff Huston 50:26 Why? Well, I don't have to muck around with that weird limit of what's the current number, 1240 octets per answer. I don't have to jump through extraordinary hoops to figure out DNSSEC big answers. And let me give you the other one. There's this thing called chaining, where in essence, when I give you a signed answer, you now have to ask a whole bunch of additional questions to say is that signature valid, because I've got to basically produce a validation chain. Now I know what questions you're going to ask. I can give you the answers already with the first answer. We're back to the model of TLS handshakes, where you don't need to do anything. You just need to do the compute on the validity. Oh, so if I use TCP, I could do that. Yes, you can. So what do I win? I win a certain amount of determinism and get rid of a huge amount of well, let's be derogatory. UDP slop. That while it's efficient, the industry has said queries are free. Let's just make more of them. George Michaelson 51:26 We might now have arrived at yes, there's a cost here, and it's slightly higher than I hope to get in UDP. But lived experience tells me I wasn't getting that in UDP. I now have certainty. I have an establishment cost. I incur a burden, and you incur a burden. But that cycle of questions that needs to exist is now certain. My time bounds are significantly more certain. It's a better outcome. Geoff Huston 51:52 So, if you switch to TCP and changed nothing, you might find that your query rate would drop, and it would fit within your existing hardware blah, you know, profile, etc. because the query rate has dropped because you've replaced a UDP guesswork and aggressive timers with a definitive version of TCP that says, "I know you got it because you ACK'd my question, and now I'm just going to wait for an answer because it's TCP. And if you're not going to give me an answer. Silence. You're going to send me back a reset. We're going to know. Whereas in UDP, it's kind of I'm getting impatient. I'm getting impatient. Queries are free. I'll send you another query. And this is this sort of area around protocol engineering. George Michaelson 52:35 Yeah, Geoff Huston 52:35 where you're you're substituting, if you will, what seems to be a really efficient, lightweight protocol with something that has weight, but as well as adding weight, it adds determinism. And the UDP problem is the only way I can cope and offer you performance is to abuse it and oversend. That's the way I add surety and responsiveness into a protocol that gives me neither. George Michaelson 53:01 Now you're very fond of a certain literary trope in these ping episodes, Geoff, and that is when you say, "But let me go one stage further. And so I invite you, Geoff, having explored UDP and TCP, to go one stage further. Geoff Huston 53:21 Oh, it's all QUIC. [George: Yeah] because at this point with QUIC, you've got this idea of reliable remote procedure calls. It's a bit like UDP, but if the query is lost, you resend, which is bizarre. [George: Right]. But it's kind of UDP that's reliable, and if you do that in a quick session, then I can take the A, the quad A, and the HTTPS and throw it in the one QUIC session as separate elements inside that stream. And interestingly, if I'm talking to something that's authoritative for more than one domain or so on, I can throw it all into that one quick stream. QUIC is actually a brilliant protocol for this kind of work. You know what about the encryption over here? Well, it's only the setup that gives you the round trip time. With TLS 1.3, that's really really short. I can amortize that. So if I was going to do this myself right now, I would say it's QUIC, dudes. It's DNS over QUIC, not HTTP three. None of that. It's QUIC. George Michaelson 54:22 Yeah, parallelism, a session key, fairly lightweight overhead, TCP like behaviors in terms of the certainty, UDP like behaviors in terms of packet flows and behaviors. It sounds like a good fit. Geoff Huston 54:37 It's a fascinating area. Are we going to go there? Oh, god. All large systems have a problem in migrating. What drives us is money. But you sit there and think we are over investing in the DNS and provisioning for basically UDP slop. We're not actually doing this for traffic. We're doing it for UDP times three. What if we didn't have to do the times three? Could we do a better job? And could we actually add encryption everywhere? And could we add DNSSEC everywhere? Because the whole issue about DNSSEC is that not many folk do it because of the overheads. And so, if I'm an application and I wish to do something that was based on the authenticity of the DNS, you're wasting your time. There is none. There's only a small amount of signed names out there. But what if every name was signed basically for free? Oh, that gets interesting. So you know, is this possible? George Michaelson 55:27 And what if I got channel security for free in the same framework? Geoff Huston 55:32 Absolutely, it's worth thinking about. And you know, the avenue of getting here is bizarre. But it really is a case that when you look hard at the DNS, what you actually find is a loosely coordinated connection that is phenomenally inefficient, and you can exploit it by going. Well, if we tightened up that bit of the infrastructure, what could we do with what's around it? And I kind of like that way of thinking. George Michaelson 55:54 We could do so much better. Is always an interesting place to be in a protocol. Geoff, have you written this one up. Is this available as a blog anywhere? Geoff Huston 56:02 Oh, there's more work to do, George. This is actually bigger than I had thought, so it will be written up on potaroo.net. but it'll be a couple of weeks yet before it's done. George Michaelson 56:11 It's been fascinating, Geoff. Thank you. Geoff Huston 56:14 Well, thank you, George. Thank you, dear listeners, for hanging on. Thank you. George Michaelson 56:18 If you've got a story or research to share here on Ping. Why not get in contact by email to ping@apnic.net or via the APNIC social media channels? Also, remember the measurement at APNIC. net mailing list on orbit is there to discuss and share relevant collaborative opportunities, grants and funding opportunities, jobs and graduate placings, or to seek feedback from the community on your own measurement projects. Be sure to check out the APNIC website for all your resource and community needs. Until next time.