Submind YouTube summaries
Thumbnail for Flock 2025 Distributing Open Source: AlmaLinux's Mirrorlist Evolution

Flock 2025 Distributing Open Source: AlmaLinux's Mirrorlist Evolution

Watch on YouTube

Video summary

Jonathan, the infrastructure lead at AlmaLinux since 2021, presented an overview of how the distribution's mirror system evolved from a single, fragile server into a robust global network of over 400 mirrors. When AlmaLinux launched in early 2021, it initially relied on one dedicated server that served as a single point of failure, causing significant latency issues for users far away and offering no redundancy. To address these scalability problems, the team rapidly expanded their infrastructure, eventually deploying a custom Python-based application built on standard Amazon EC2 instances and load balancers. This system now automatically scales horizontally to handle peak traffic loads, such as those seen during major releases, while conserving resources when demand drops, effectively managing millions of requests per second across the globe. The core architecture relies on a tiered approach starting with three controlled "Tier Zero" mirrors located in Seattle, Atlanta, and Amsterdam, which receive updates directly from the build system before distributing them to standard public mirrors. A critical evolution in this process was the implementation of advanced geolocation logic that moved beyond simple country-level matching to coordinate-based routing and ASN pinning for large networks like Azure and AWS. Early attempts using broad country data led to inefficient routing, such as sending New York users to California mirrors when a closer option existed; by switching to precise coordinate data from ipinfo.io and implementing randomization to distribute traffic evenly among available mirrors in a region, the team significantly reduced download times and alleviated pressure on specific servers. Furthermore, the presentation highlighted the substantial cost and reliability advantages of traditional mirroring over Content Delivery Networks (CDNs) for open-source distributions. While CDNs like Cloudflare or Fastly offer speed, they come with prohibitive costs that make them unfeasible at the scale required by large Linux projects, and they introduce a single point of failure in their control plane that cannot be easily recovered by the project itself. In contrast, AlmaLinux's self-managed mirror network allows for full control over content validation, rapid recovery from outages, and significant infrastructure savings, with caching layers alone reducing DNF update response times by 70% and cutting EC2 instance usage by two-thirds. The talk concluded with a call to action encouraging the community to contribute servers and transit capacity to maintain this vital infrastructure, ensuring that open-source distributions remain accessible and affordable for users worldwide.
Read the full video transcript
Hello. Can everybody hear me? >> Yes. >> Cool. All right. Um, so t talk is titled distributing open source. Alma Linux is mirrorless evolution. Um, and what we're going to talk about is um mostly going to be about how the Alma Linux mirror system itself has evolved, but it's all applicable to um broader open source and mirroring and open source in general. Um, so just a quick show of hands. Who knows what mirroring is in this context? Okay. Who knows what CDN's are? Who understands why traditional mirroring is better for open source projects than CDN's? One. Okay. One. [laughter] I'll give you a hint. Money. Um. Anyway, uh I'm Jonathan. I've been the infrastructure lead at Alma Linux since 2021. I became a Fedora packager in 2022. uh thanks to Carl George who finally convinced me to uh to join up. Uh became a package sponsor in I think it was 2023 and I have mentored uh one one other Fedora packager so far. Eventually Andrew is going to get wrangled into it u one of these days. Um I've been a heavy open source user and and around open source since about 2005. Uh so right at about 20 years. Does anybody remember PHP Nuke? It was a there. Okay, we got a few. It was a pre-WordPress CMS back in the day. You had like P PHP Nuke, Drupal, and Jumla. Um, and that was that was like it. So, I got my start in PHP Nuke. And, uh, you know, kind of kind of went from there. Um, quick tidbit, my latest hobby, I'm an aspiring private pilot. Um, in about a month, I should officially be a pilot. So, cool new hobby. All right, so let's talk about how things started uh within Alma Linux. Uh when Alma Linux cranked up in uh 2021, you know, we were starting from nothing. Uh had no users, had no operating system. There was nothing there. Uh so it didn't make any sense to jump straight into, you know, a fullfledged mirror system like Fedora has and mirror manager um or, you know, other distros have similar things. So our mirror system wasn't a mirror system at all. It was a single server. Uh the the mirror URL, you know, we we set it up ahead of time knowing that it was coming, but it was just you were hitting a single server. Um there were a lot of problems with that. Uh that one mirror uh server that we had could have been really far from the user. That one I I want to say Andrew, it was a dedicated server at Hner, wasn't it? The original repo. Alma. So, you know, for people in the US that wasn't great. For people in Asia that wasn't great. single point of failure just you know bad bad bad uh had no redundancy and that obviously wasn't going to scale as as Linux grew so we went from that to now having a network of over 400 mirrors worldwide which is more than the majority of Linux distributions out there um we've got about 1.5 million systems that hit our public mirror list every week um we know of several million more systems than that using private mirror mirrors uh which is kind of out of scope for this talk. Um we're averaging about 420 requests per second across those 400 mirrors. Um during peak releases of like recently on my Linux 9.6 and 10, you know, that's peaking into over a thousand requests per second. So, you know, a good bit of traffic. We've got automatic horizontal scaling now. So, when we when we get those peaks, the system scales up. When the peaks are over, system scales back down. Uh you know, conserves on resources. Uh, and it's all backed by a custom software stack. We just call it the Alma Linux mirror list system. Uh, but it's a Pythonbased application that we've written from the ground up. Um, and it's just using regular old Amazon EC2 instances, uh, an Amazon load balancer. So, nothing too fancy there. It's all easily portable to other clouds or bare metal, uh, what have you. You know, nothing is is vendor locking it to to what it's on now. Um that's just a quick graph of how Alma Linux has grown over the years. Uh and we're about to cover kind of some milestones that we hit throughout that with the mirrorless software uh and how we distribute my Linux. So how did we get here? Um first and foremost is uh tier zero mirrors. We these days we have repo.allinx.org. That's kind of the the king of the castle. All updates from our build system. First go to rebuild.allinx.org. rebuild.allinx.org then pushes traffic to what we call tier zero arsync mirrors. We currently have uh three of them. We used to have four. We lost one. So we've got three of them. Um Seattle, Washington, Atlanta, Georgia, and Amsterdam. So we got two in the US, one in Europe. Uh and we kind of geographically spread out that traffic with just some basic DNS geo steering. um you know, nothing too fancy. So, unlike a lot of projects, um we control all of our tier zero arsync mirrors. So, we're pushing to them instead of having them pull from us. Um we're monitoring them. Uh we're making sure that, you know, they're not hitting their port capacities, anything like that. A lot of projects will have kind of their little internal mirror network uh that then their tier zeros can pull from and those tier zeros might just be, you know, public mirrors with big pipes um per se. And there are some downsides to that. You know, the project can't monitor that. If there's an issue, the project can't quickly get it back online. Um you can have out of sync issues and and things can kind of quickly fall apart. Uh so that's why we've opted to control the tier zeros ourselves. Um and then beyond that, the next layer is standard mirroring which all systems pull their updates from. Um here's kind of a a quick breakdown on how we use the geodns to steer that traffic. Um it's got the the three mirrors, three arsync mirrors that I mentioned on there. Uh we're working on a fourth one that'll be coming online. Um Atlanta, Georgia is kind of where it all started and that one is obviously a little bit undersized compared to the others. So we've got a uh a 50 gig tier zero mirror that'll be coming online in in Kansas City, uh Missouri, which is a little bit northwest of Atlanta. So uh pretty similar. We we had one in Tokyo at one point that was serving uh Asia, you know, had a lot of lot better connectivity from Tokyo to the whole AP pack region uh as far as getting those updates out to the the standard mirrors. Uh unfortunately, we lost that one due to a hardware failure and the uh the provider, while they still have big pipes in that location, they've um stopped putting new non-net network equipment in there. It's just a peering facility for them now. a quick snippet for anybody that hasn't heard of the who has heard of the micro mirror project. So, a guy named Kenneth uh and and a friend of his, John, started this project um about the same time that we were kicking up Alma Linux a few years back and they've um they're building these mirrors on these little uh they call it micro mirror. They're essentially using what a lot of places would use as like thin clients and they're stuffing some SSDs in them and they're sending them out to uh you know anybody from ISPs to colleges to transit providers, anybody that'll take them. Uh and they're mirroring Linux all over the place. Uh there's a picture, I wish I had it, I should have put it in the slide of a micro mirror running inside one of those little roadside telco boxes, if y'all have ever seen those. Um really good project. They they do a lot of mirroring for Alma Linux. I think they're up to like 30 systems now or something, but they're mirroring Ubuntu and Fedora and Apple and um the whole shebang. It's a really cool project. All right, [sighs] so let's look at the history uh and then how we ended up where we are today. Uh so back in February of 2021, we had no operating system. There was no Linux yet. Uh we announced it, we're looking at it, you know, everybody's excited, but there's no Linux. Uh so that's when the initial mirror system um aka Petner dedicated server uh was spun up. There was a lot of um a lot of excitement surrounding all my Linux. Uh we had some people jumping on board um and you know they set up mirrors which at this point had no content but you know they're setting up their crowns and everything's ready to go. Um we had about 32 people uh by the end of February of 2021 set up mirrors uh pulling from that one server kind of getting ready. Um by March just one month later we're up to 54 mirrors. That's when the the media is kind of taking over spreading the word about um you know CentOS and now there's Alma Linux and so on and so forth. So more people are stepping up. Um 54 mirrors and and no operating system until the end of the month is pretty impressive. uh you know I think we had all these people excited to distribute something that that was vaporware until right at the end of the month. Um the system was still dumb. It was a a static list being distributed uh of mirrors and I think this next slide. Yeah. So Andrew who's our lead architect sitting right there in the back um he was managing all of this by hand. you know, monitoring when mirrors would drop offline. He would take them out of the configs or add new mirrors, make sure they had all the content. All this was being done by hand. It was kind of a mess. Uh, and I don't that's kind of legible. There's just some some getit history and commits on, you know, lots of back and forth. This is obviously not going to scale. You know, we can't manage mirrors like this forever. We need some automation. So, in July of 2021 is when we first start doing some of our big changes. uh we add some geoloccation aware logic to the mirror system so that hopefully we're serving people mirrors that are are geographically close to them. Uh there's some basic mirror status checks implemented so that when mirrors you know go offline for maintenance or you know flap up and down whatever that's all handled automatically. We're not serving people dead mirrors uh and we're also not having to manage that by hand. Uh Andrew can take a break. Um later in 2021, this is when things really start picking up. Um we introduce ASN and subnet pinning. That is for large networks. Um, think Azure, think AWS, uh, think our friends at CERN, you know, people like that that have large deployments of Alma Linux in, you know, one location or or multiple locations or whatever. It makes more sense for them to mirror Alma Linux internally uh than it does for them to use the bandwidth from public mirrors. It's better for them, it's better for us, everybody wins. Uh, so we added this logic uh specifically is actually at the request of Azure. Uh, but it's, you know, useful for everybody else. And it we'll get to some some pictures in a little bit and I'll show you why this is important. But we're starting to see some issues. Um the mirror status checker is getting a bit slow at this point. We've got it's a a singlethreaded uh Python loop that every I think at that point we had it running once an hour that's going through all the mirrors, you know, checking all the the critical endpoints, making sure the data is up to date, making sure everything is there. Um, at this point we're we're in the couple hundred mirror range. Uh, and that single thread is starting to uh, basically it can't finish before it has to run again. So, we're running into some problems there. Uh, we also noticed some some pretty bad bugs. When we were doing the geoloccation at this point, it was a a pretty broad geoloccation match. So, we were saying, "Hey, this user is in this country. Let's serve him a mirror out of this country." All right, that's great. in a small country like uh I don't know, I'm from the US. It's a pretty big geographical country. So when you take somebody in the US, uh say they're in let's say New York, everybody kind of knows where New York is. Um and you serve them a mirror out of Seattle, Washington all the way across the country, that's not great. You know, we can do better than that. Uh so for large geographically large countries with a lot of different internet hubs within them uh such as the US such as Canada such as um you know Germany would probably be a pretty good example in Europe you know you've got opposite sides of Germany both of which have you know multiple internet hubs um we needed to get better than just hey here's a country serve a mirror in that uh in that country. So that's a problem that we need to solve. Um and also when we were doing the logic to serve a list of mirrors in that country there was no randomization. So it was pretty much alphabetically serving countries serving mirrors in a given country. Um in the US say we've got 50 mirrors at this point. The same one two three mirrors are getting the whole country's traffic. That's not very great either. Um a quick solution to that was some some randomization logic. uh while we thought on it more. Um but then in November uh this is where we start looking at better geoloccation, better serving of the mirrors. Um the mirrorless generation code is pretty heavy. Um the matching for ASN's for IPs, uh for subnets, you can only get that so lightweight. You know, you're still underneath it all you're doing math. Um and we we couldn't make that math any lighter. So what we opted to do to improve things from there is implement a caching layer. Uh so that when you are requesting a mirror, you know, a DNF update for example, by default in in Alma Linux and Fedora, um Red Hat, whatever, there's more than one repository configuration that you're searching for updates for. So you've got the the request DNF update, but behind the scenes, you're actually requesting updates for three, four, five, you know, however many repos you have enabled by default. That's an easy problem to solve, right? Let's throw some caching in there. We get that first one. We can cache what we return. We know that's valid mirrors for this request. Let's throw it back at them. Uh so we did that. Uh that gave us I think another slide covers it. But um on Alma Linux, we had at the time three repositories enabled by default. we instantly saw like a 70% reduction uh in the request timings uh for for DNF and it it cut our uh uh the infrastructure the the number of EC2 instances behind this by about twothirds perfect you know linear uh benefit there we implemented some mirror flap detection so that when mirrors were you know up and down we can pull them out for a good bit of time uh because if they're up and down you know obviously there's some issues Um a big one we start doing geoloccation based on coordinates not just the country so that within large countries or across the boundaries of countries uh we can now serve a better mirror. Let's go back to our New York example and looking at New York. We don't want to say there's not a mirror in New York. Uh say there's not another mirror in the US except for in California. We don't want to serve somebody in New York a mirror in California when there's a mirror right there in Toronto that just happens to be right across the border with Canada. Um, so that solved a lot of problems as well. Sped things up uh for a lot of people. Getting low on time. I'm going to start start moving through these. Uh, 2022 everything is chugging along pretty well. Uh, 2023 had some releases. We um made some tweaks to improve cash hits. Uh, still chugging along pretty well. Still continuing on uh minor tweaks, nothing major. This is going into the details about the uh the caching improvements. Um and this is the this is the improvement that that caching had. Um, it took the overall DNF transaction response time, the cumulative response time, uh, from, you know, 3/4 of a second, 1 second down to like 200 milliseconds. Uh, just massive, you know, you can see it right there. Um, and and in the overall scale of a DNF transaction, that's a very short amount of time. You know, saving what is what is saving three quarters or one second um to a DNF transaction? was nothing but the amount of resources that it saved on the infrastructure side was huge. Uh the next problem that comes up [sighs] we were using Maxmind um for for geoloccation data. Is everybody kind of familiar with MaxM mind? A lot of people they publish a a free geoccation database. They also have a paid product. um that data was not great and we can only serve mirrors uh as well as you know the data backing it that tell us where these people are that we need to be serving. Uh we found a lot of inaccuracies and uh just flatout wrong data within MaxMine. You know people in Europe are getting served US mirrors because that's what the MaxMine data was telling us. The inverse US to Europe and and Asia to the US like just all of these really bad things. Um the the issue was worse than that as well. Who can tell me what's right there in the US? >> There is absolutely nothing there in the US. It's it's it's smack in the middle of Kansas and uh there's there's certainly no internet infrastructure there. However, according to MaxMine GeoP, this was where we got the most traffic from. So what that is, it's the geographical coordinates for what is is considered the geographic center of the continental US. Uh so when MaxMine didn't know specifically where an address was from, it put it right there. What's the problem for the mirrors that are right there? They're not designed, they're not in an internet hub, you know, they've got a a gig port, you know, 500meg port, whatever. They're getting hammered with this traffic. Uh and they're not designed for it. They're not in an area with with network that can handle it. So, we found this company, ipinfo.io, uh commercial company. They make uh some really really good uh geoloccation databases. Uh they've got like these uh I forget what they call them, polers or something like that all around the world that are, you know, doing trace routes and and you know, really honing in on where addresses are. So, we implement their data. Now, it looks like this. It's about more what you'd expect. >> [clears throat] >> uh you know, you've got the the hot spot right there in in Virginia where all the the AWS data centers are and and everything. So, that solved that problem. Uh those there were there were like three or four mirrors uh that were, you know, getting hammered in Kansas. They were happy all of a sudden, you know, they're not getting getting hit anymore. They can actually handle the traffic. Uh users get faster downloads. Those poor people in the UK and Asia that are hitting US mirrors in US and in Europe, you know, everybody's happier. everything's going a lot better. So what have we found uh through our whole experience? Traditional mirroring is amazing for free and open source distributions uh for a number of reasons. One of those is cost. One of those is quality. Uh and the the third reason is uh quite simply put fewer points of failure. When you have a CDN, if the CDN's control plane dies, dead in the water. There's nothing you can do. You don't have actual software uh behind the scenes to deal with it. With this system, we can pick it up, redeploy it, uh, on a completely different, you know, cloud or bare metal or whatever. We're good to go. Everything is happy. Um, delivering releases and updates, uh, without these mirrors would be incredibly expensive. Um, CDNs are not cheap. Fastly, Cloudflare, a you know, whichever one, they're all very, very expensive. um at the scale that we are pushing traffic uh at the scale that we have the requests uh and it it just it wouldn't be feasible. There's no way it would work. Um the the other thing we found is accurate geoloccation data improves user experience tfold. Bad data in, bad results out. So um what's next? We want to implement some more statistics within the mirror system uh similar to what Fedora has. You can go to the Fedora mirror manager site uh see like country stats, where the mirrors are, what's online, what's offline, um things like that. We want to do that public facing. Uh we need a better development environment to get some more people on board. Um it's it's all written in Python. It's relatively easy. It's it's basic Python async.io stuff. Um but the the barrier to getting a development environment up for the code um is not that easy. So, we need some we need some a better development environment set up u you know both guide and maybe some pre-made images things like that to make it easier for people to jump in and contribute. Uh protocol filtering, we actually got that one done uh earlier this year where you can like forcibly select if you only want HTTPS mirrors or arsync mirrors, what have you. uh partial mirroring. Um as Alma Linux has grown in in both number of versions, architectures, things like that, you know, our mirrors are now at, you know, 1.5, two terabytes. People want to be able to trim that down. Maybe they only want to mirror x86. Maybe they only want to mirror ARM. Maybe they want to ignore ISOs. Um in our status checkers, we've got to be able to handle that. [snorts] Um that kind of goes hand inhand with better mirror validation. Uh we need to find more efficient ways to check the status of them. Uh using arsync is a very obvious option. And then uh just a little call to action, get involved uh mirrors. Linux.org. Uh the code is on GitHub. Barring that, you know, if you're not a coder, maybe you have access to servers, transit, uh contribute mirrors to your favorite open source projects, especially distributions. A lot of projects have moved away from mirroring. Um, you know, everything from Apache to GCC to the kernel used to come from these traditional style mirrors that we're talking about. A lot of that with GitHub and and things like that have kind of moved away from it. Distros are largely still reliant on that type of architecture. Um, so you know, set up mirrors for projects that you care about. And then we are out of time, but if there's any quick questions, we might can sneak them in. There's nobody. Nobody busting in the back door. >> I had to like really ram through the second half of that. Sorry. >> My question is kind of related to that. How quickly were you able to pull this talk together in the 11th minute hour to fill the schedule gap? [laughter] >> So actually, >> thank you by the way. Sincerely. >> No, no problem. So actually I gave this talk at Flock last year. However, um I had I think one attendee because I was right next to the forjo discussion where everybody was. So [laughter] it's still pretty new content. >> No problem. We've got one behind you. >> Hello. Uh you said that you don't combine uh CDNs with mirroring, but Fedora has a CDN mirror. So, so we actually do use CDNs to some degree. Um, repo.allinx.org itself is a CDN. Um, simply because of the type of traffic that it receives. Uh, it is not part of our public mirror infrastructure. So, when you have a system and go to get updates on it, you're not hitting repo. Linux.org. Um, that's primarily traffic from the website that's like, hey, go grab an ISO, go, you know, grab an image. So, that is behind a CDN. And then the special mirroring that we have um inside of AWS inside of Azure they are using those clouds internal CDNs uh that we control and manage um within those environments you know that is the proper fit. Uh but for the the broader public mirrorless system uh you know CDN would be cost prohibitive. David. >> Uh yeah. So uh at the end there you were talking about you know volunteer uh to be a mirror for projects you care about and at the beginning you were talking about how Alman Linux you unlike other projects you you kind of control the mirrors you push to them. Is that >> the tier zeros? Yeah. >> Yeah. So is that when you were saying volunteer were you saying volunteer to be like a a tier zero or is that >> a volunteer to be a public mirror? So we push from repo.allinx.org. Let's say we push from the build system into our three currently three about to be four tier zero arsync mirrors. Everybody then pulls from there. A lot of other projects and maybe Kevin might know how Fedora is operated these days because I don't remember um but a lot of projects the tier zeros are technically community contributed as well. It's just you know selective relationships, mirrors with big pipes, things like that. We choose to find people that will give us the server and the transit and let us manage the tier zero. Um are are the tier zeros within Fedora's mirroring still kind of public contributions Kevin? >> So tier zier one is our community >> and then is there >> is there a tier two then >> then everybody else? Yeah. >> Okay. Yeah. Yeah. So, a little bit different lingo, but that's pretty much what I was alluding to. All right, we are out of time and people are coming in, so we'll we'll call it there. Thanks for coming. [applause] Thank you.