Flock 2025 Distributing Open Source: AlmaLinux's Mirrorlist Evolution
Watch on YouTubeVideo summary
Jonathan, the infrastructure lead at AlmaLinux since 2021, presented an overview of how the distribution's mirror system evolved from a single, fragile server into a robust global network of over 400 mirrors. When AlmaLinux launched in early 2021, it initially relied on one dedicated server that served as a single point of failure, causing significant latency issues for users far away and offering no redundancy. To address these scalability problems, the team rapidly expanded their infrastructure, eventually deploying a custom Python-based application built on standard Amazon EC2 instances and load balancers. This system now automatically scales horizontally to handle peak traffic loads, such as those seen during major releases, while conserving resources when demand drops, effectively managing millions of requests per second across the globe.
The core architecture relies on a tiered approach starting with three controlled "Tier Zero" mirrors located in Seattle, Atlanta, and Amsterdam, which receive updates directly from the build system before distributing them to standard public mirrors. A critical evolution in this process was the implementation of advanced geolocation logic that moved beyond simple country-level matching to coordinate-based routing and ASN pinning for large networks like Azure and AWS. Early attempts using broad country data led to inefficient routing, such as sending New York users to California mirrors when a closer option existed; by switching to precise coordinate data from ipinfo.io and implementing randomization to distribute traffic evenly among available mirrors in a region, the team significantly reduced download times and alleviated pressure on specific servers.
Furthermore, the presentation highlighted the substantial cost and reliability advantages of traditional mirroring over Content Delivery Networks (CDNs) for open-source distributions. While CDNs like Cloudflare or Fastly offer speed, they come with prohibitive costs that make them unfeasible at the scale required by large Linux projects, and they introduce a single point of failure in their control plane that cannot be easily recovered by the project itself. In contrast, AlmaLinux's self-managed mirror network allows for full control over content validation, rapid recovery from outages, and significant infrastructure savings, with caching layers alone reducing DNF update response times by 70% and cutting EC2 instance usage by two-thirds. The talk concluded with a call to action encouraging the community to contribute servers and transit capacity to maintain this vital infrastructure, ensuring that open-source distributions remain accessible and affordable for users worldwide.
Read the full video transcript
Hello. Can everybody hear me?
>> Yes.
>> Cool. All right. Um,
so t talk is titled distributing open
source. Alma Linux is mirrorless
evolution. Um, and what we're going to
talk about is um mostly going to be
about how the Alma Linux mirror system
itself has evolved, but it's all
applicable to um broader open source and
mirroring and open source in general.
Um, so just a quick show of hands. Who
knows what mirroring is in this context?
Okay. Who knows what CDN's are? Who
understands why traditional mirroring is
better for open source projects than
CDN's?
One. Okay. One. [laughter]
I'll give you a hint. Money. Um.
Anyway, uh I'm Jonathan. I've been the
infrastructure lead at Alma Linux since
2021. I became a Fedora packager in
2022. uh thanks to Carl George who
finally convinced me to uh to join up.
Uh became a package sponsor in I think
it was 2023 and I have mentored uh one
one other Fedora packager so far.
Eventually Andrew is going to get
wrangled into it u one of these days. Um
I've been a heavy open source user and
and around open source since about 2005.
Uh so right at about 20 years. Does
anybody remember PHP Nuke? It was a
there. Okay, we got a few. It was a
pre-WordPress CMS back in the day. You
had like P PHP Nuke, Drupal, and Jumla.
Um, and that was that was like it. So, I
got my start in PHP Nuke. And, uh, you
know, kind of kind of went from there.
Um, quick tidbit, my latest hobby, I'm
an aspiring private pilot. Um, in about
a month, I should officially be a pilot.
So, cool new hobby.
All right, so let's talk about how
things started uh within Alma Linux. Uh
when Alma Linux cranked up in uh 2021,
you know, we were starting from nothing.
Uh had no users, had no operating
system. There was nothing there. Uh so
it didn't make any sense to jump
straight into, you know, a fullfledged
mirror system like Fedora has and mirror
manager um or, you know, other distros
have similar things. So our mirror
system wasn't a mirror system at all. It
was a single server. Uh the the mirror
URL, you know, we we set it up ahead of
time knowing that it was coming, but it
was just you were hitting a single
server. Um there were a lot of problems
with that. Uh that one mirror uh server
that we had could have been really far
from the user. That one I I want to say
Andrew, it was a dedicated server at
Hner, wasn't it? The original repo.
Alma. So, you know, for people in the US
that wasn't great. For people in Asia
that wasn't great. single point of
failure just you know bad bad bad uh had
no redundancy and that obviously wasn't
going to scale as as Linux grew
so we went from that to now having a
network of over 400 mirrors worldwide
which is more than the majority of Linux
distributions out there um we've got
about 1.5 million systems that hit our
public mirror list every week um we know
of several million more systems than
that using private mirror mirrors uh
which is kind of out of scope for this
talk. Um we're averaging about 420
requests per second across those 400
mirrors. Um during peak releases of like
recently on my Linux 9.6 and 10, you
know, that's peaking into over a
thousand requests per second. So, you
know, a good bit of traffic. We've got
automatic horizontal scaling now. So,
when we when we get those peaks, the
system scales up. When the peaks are
over, system scales back down. Uh you
know, conserves on resources. Uh, and
it's all backed by a custom software
stack. We just call it the Alma Linux
mirror list system. Uh, but it's a
Pythonbased
application that we've written from the
ground up. Um, and it's just using
regular old Amazon EC2 instances, uh, an
Amazon load balancer. So, nothing too
fancy there. It's all easily portable to
other clouds or bare metal, uh, what
have you. You know, nothing is is vendor
locking it to to what it's on now.
Um that's just a quick graph of how Alma
Linux has grown over the years. Uh and
we're about to cover kind of some
milestones that we hit throughout that
with the mirrorless software uh and how
we distribute my Linux.
So how did we get here?
Um first and foremost is uh tier zero
mirrors. We these days we have
repo.allinx.org.
That's kind of the the king of the
castle. All updates from our build
system. First go to rebuild.allinx.org.
rebuild.allinx.org
then pushes traffic to what we call tier
zero arsync mirrors. We currently have
uh three of them. We used to have four.
We lost one. So we've got three of them.
Um Seattle, Washington, Atlanta,
Georgia, and Amsterdam. So we got two in
the US, one in Europe. Uh and we kind of
geographically spread out that traffic
with just some basic DNS geo steering.
um you know, nothing too fancy. So,
unlike a lot of projects, um we control
all of our tier zero arsync mirrors. So,
we're pushing to them instead of having
them pull from us. Um we're monitoring
them. Uh we're making sure that, you
know, they're not hitting their port
capacities, anything like that. A lot of
projects will have kind of their little
internal mirror network uh that then
their tier zeros can pull from and those
tier zeros might just be, you know,
public mirrors with big pipes um per se.
And there are some downsides to that.
You know, the project can't monitor
that. If there's an issue, the project
can't quickly get it back online. Um you
can have out of sync issues and and
things can kind of quickly fall apart.
Uh so that's why we've opted to control
the tier zeros ourselves. Um and then
beyond that, the next layer is standard
mirroring which all systems pull their
updates from.
Um here's kind of a a quick breakdown on
how we use the geodns to steer that
traffic. Um it's got the the three
mirrors, three arsync mirrors that I
mentioned on there. Uh we're working on
a fourth one that'll be coming online.
Um Atlanta, Georgia is kind of where it
all started and that one is obviously a
little bit undersized compared to the
others. So we've got a uh a 50 gig tier
zero mirror that'll be coming online in
in Kansas City, uh Missouri, which is a
little bit northwest of Atlanta. So uh
pretty similar. We we had one in Tokyo
at one point that was serving
uh Asia, you know, had a lot of lot
better connectivity from Tokyo to the
whole AP pack region uh as far as
getting those updates out to the the
standard mirrors. Uh unfortunately, we
lost that one due to a hardware failure
and the uh the provider, while they
still have big pipes in that location,
they've um stopped putting new non-net
network equipment in there. It's just a
peering facility for them now.
a quick snippet for anybody that hasn't
heard of the who has heard of the micro
mirror project. So, a guy named Kenneth
uh and and a friend of his, John,
started this project um about the same
time that we were kicking up Alma Linux
a few years back and they've um
they're building these mirrors on these
little uh they call it micro mirror.
They're essentially using what a lot of
places would use as like thin clients
and they're stuffing some SSDs in them
and they're sending them out to uh you
know anybody from ISPs to colleges to
transit providers, anybody that'll take
them. Uh and they're mirroring Linux all
over the place. Uh there's a picture, I
wish I had it, I should have put it in
the slide of a micro mirror running
inside one of those little roadside
telco boxes, if y'all have ever seen
those. Um really good project. They they
do a lot of mirroring for Alma Linux. I
think they're up to like 30 systems now
or something, but they're mirroring
Ubuntu and Fedora and Apple and um the
whole shebang. It's a really cool
project.
All right, [sighs] so let's look at the
history uh and then how we ended up
where we are today. Uh so back in
February of 2021, we had no operating
system. There was no Linux yet. Uh we
announced it, we're looking at it, you
know, everybody's excited, but there's
no Linux. Uh so that's when the initial
mirror system um aka Petner dedicated
server uh was spun up. There was a lot
of um a lot of excitement surrounding
all my Linux. Uh we had some people
jumping on board um and you know they
set up mirrors which at this point had
no content but you know they're setting
up their crowns and everything's ready
to go. Um we had about 32 people uh by
the end of February of 2021 set up
mirrors uh pulling from that one server
kind of getting ready.
Um by March just one month later we're
up to 54 mirrors. That's when the the
media is kind of taking over spreading
the word about um you know CentOS and
now there's Alma Linux and so on and so
forth. So more people are stepping up.
Um 54 mirrors and and no operating
system until the end of the month is
pretty impressive. uh you know I think
we had all these people excited to
distribute something that that was
vaporware until right at the end of the
month. Um the system was still dumb. It
was a a static list being distributed uh
of mirrors and I think this next slide.
Yeah. So Andrew who's our lead architect
sitting right there in the back um he
was managing all of this by hand. you
know, monitoring when mirrors would drop
offline. He would take them out of the
configs or add new mirrors, make sure
they had all the content. All this was
being done by hand. It was kind of a
mess. Uh, and I don't that's kind of
legible. There's just some some getit
history and commits on, you know, lots
of back and forth. This is obviously not
going to scale. You know, we can't
manage mirrors like this forever. We
need some automation.
So, in July of 2021 is when we first
start doing some of our big changes. uh
we add some geoloccation aware logic to
the mirror system so that hopefully
we're serving people mirrors that are
are geographically close to them. Uh
there's some basic mirror status checks
implemented so that when mirrors you
know go offline for maintenance or you
know flap up and down whatever that's
all handled automatically. We're not
serving people dead mirrors uh and we're
also not having to manage that by hand.
Uh Andrew can take a break.
Um later in 2021, this is when things
really start picking up. Um we introduce
ASN and subnet pinning. That is for
large networks. Um, think Azure, think
AWS,
uh, think our friends at CERN, you know,
people like that that have large
deployments of Alma Linux in, you know,
one location or or multiple locations or
whatever. It makes more sense for them
to mirror Alma Linux internally uh than
it does for them to use the bandwidth
from public mirrors. It's better for
them, it's better for us, everybody
wins. Uh, so we added this logic uh
specifically is actually at the request
of Azure. Uh, but it's, you know, useful
for everybody else. And it we'll get to
some some pictures in a little bit and
I'll show you why this is important.
But we're starting to see some issues.
Um the mirror status checker is getting
a bit slow at this point. We've got it's
a a singlethreaded uh Python loop that
every I think at that point we had it
running once an hour that's going
through all the mirrors, you know,
checking all the the critical endpoints,
making sure the data is up to date,
making sure everything is there. Um, at
this point we're we're in the couple
hundred mirror range. Uh, and that
single thread is starting to uh,
basically it can't finish before it has
to run again. So, we're running into
some problems there. Uh, we also noticed
some some pretty bad bugs. When we were
doing the geoloccation at this point, it
was a a pretty broad geoloccation match.
So, we were saying, "Hey, this user is
in this country. Let's serve him a
mirror out of this country." All right,
that's great. in a small country like uh
I don't know, I'm from the US. It's a
pretty big geographical country. So when
you take somebody in the US, uh say
they're in let's say New York, everybody
kind of knows where New York is. Um and
you serve them a mirror out of Seattle,
Washington all the way across the
country, that's not great. You know, we
can do better than that. Uh so for large
geographically large countries with a
lot of different internet hubs within
them uh such as the US such as Canada
such as um you know Germany would
probably be a pretty good example in
Europe you know you've got opposite
sides of Germany both of which have you
know multiple internet hubs um we needed
to get better than just hey here's a
country serve a mirror in that uh in
that country.
So that's a problem that we need to
solve. Um and also when we were doing
the logic to serve a list of mirrors in
that country there was no randomization.
So it was pretty much alphabetically
serving countries serving mirrors in a
given country. Um in the US say we've
got 50 mirrors at this point. The same
one two three mirrors are getting the
whole country's traffic. That's not very
great either. Um a quick solution to
that was some some randomization logic.
uh while we thought on it more.
Um but then in November uh this is where
we start looking at better geoloccation,
better serving of the mirrors. Um the
mirrorless generation code is pretty
heavy. Um the matching for ASN's for
IPs, uh for subnets,
you can only get that so lightweight.
You know, you're still underneath it all
you're doing math. Um and we we couldn't
make that math any lighter. So what we
opted to do to improve things from there
is implement a caching layer. Uh so that
when you are requesting a mirror, you
know, a DNF update for example, by
default in in Alma Linux and Fedora, um
Red Hat, whatever, there's more than one
repository configuration that you're
searching for updates for. So you've got
the the request DNF update, but behind
the scenes, you're actually requesting
updates for three, four, five, you know,
however many repos you have enabled by
default.
That's an easy problem to solve, right?
Let's throw some caching in there. We
get that first one. We can cache what we
return. We know that's valid mirrors for
this request. Let's throw it back at
them. Uh so we did that. Uh that gave us
I think another slide covers it. But um
on Alma Linux, we had at the time three
repositories enabled by default. we
instantly saw like a 70% reduction uh in
the request timings uh for for DNF and
it it cut our uh
uh the infrastructure the the number of
EC2 instances behind this by about
twothirds perfect you know linear uh
benefit there we implemented some mirror
flap detection so that when mirrors were
you know up and down we can pull them
out for a good bit of time uh because if
they're up and down you know obviously
there's some issues
Um a big one we start doing geoloccation
based on coordinates not just the
country so that within large countries
or across the boundaries of countries uh
we can now serve a better mirror. Let's
go back to our New York example and
looking at New York. We don't want to
say there's not a mirror in New York. Uh
say there's not another mirror in the US
except for in California. We don't want
to serve somebody in New York a mirror
in California when there's a mirror
right there in Toronto that just happens
to be right across the border with
Canada. Um, so that solved a lot of
problems as well. Sped things up uh for
a lot of people. Getting low on time.
I'm going to start start moving through
these. Uh, 2022 everything is chugging
along pretty well. Uh, 2023 had some
releases. We um made some tweaks to
improve cash hits. Uh, still chugging
along pretty well.
Still continuing on uh minor tweaks,
nothing major. This is going into the
details about the uh the caching
improvements.
Um and this is the this is the
improvement that that caching had. Um,
it took the overall DNF transaction
response time, the cumulative response
time, uh, from, you know, 3/4 of a
second, 1 second down to like 200
milliseconds. Uh, just massive, you
know, you can see it right there. Um,
and and in the overall scale of a DNF
transaction, that's a very short amount
of time. You know, saving what is what
is saving three quarters or one second
um to a DNF transaction? was nothing but
the amount of resources that it saved on
the infrastructure side was huge.
Uh the next problem that comes up
[sighs] we were using Maxmind um for for
geoloccation data. Is everybody kind of
familiar with MaxM mind? A lot of people
they publish a a free geoccation
database. They also have a paid product.
um that data was not great and we can
only serve mirrors uh as well as you
know the data backing it that tell us
where these people are that we need to
be serving. Uh we found a lot of
inaccuracies and uh just flatout wrong
data within MaxMine. You know people in
Europe are getting served US mirrors
because that's what the MaxMine data was
telling us. The inverse US to Europe and
and Asia to the US like just all of
these really bad things. Um the the
issue was worse than that as well. Who
can tell me what's right there in the
US?
>> There is absolutely nothing there in the
US. It's it's it's smack in the middle
of Kansas and uh there's there's
certainly no internet infrastructure
there. However, according to MaxMine
GeoP, this was where we got the most
traffic from. So what that is, it's the
geographical coordinates for what is is
considered the geographic center of the
continental US. Uh so when MaxMine
didn't know specifically where an
address was from, it put it right there.
What's the problem for the mirrors that
are right there? They're not designed,
they're not in an internet hub, you
know, they've got a a gig port, you
know, 500meg port, whatever. They're
getting hammered with this traffic. Uh
and they're not designed for it. They're
not in an area with with network that
can handle it.
So, we found this company, ipinfo.io,
uh commercial company. They make uh some
really really good uh geoloccation
databases. Uh they've got like these uh
I forget what they call them, polers or
something like that all around the world
that are, you know, doing trace routes
and and you know, really honing in on
where addresses are. So, we implement
their data. Now, it looks like this.
It's about more what you'd expect.
>> [clears throat]
>> uh you know, you've got the the hot spot
right there in in Virginia where all the
the AWS data centers are and and
everything.
So, that solved that problem. Uh those
there were there were like three or four
mirrors uh that were, you know, getting
hammered in Kansas. They were happy all
of a sudden, you know, they're not
getting getting hit anymore. They can
actually handle the traffic. Uh users
get faster downloads. Those poor people
in the UK and Asia that are hitting US
mirrors in US and in Europe, you know,
everybody's happier. everything's going
a lot better. So what have we found uh
through our whole experience?
Traditional mirroring is amazing for
free and open source distributions uh
for a number of reasons. One of those is
cost. One of those is quality. Uh and
the the third reason is uh quite simply
put fewer points of failure. When you
have a CDN, if the CDN's control plane
dies, dead in the water. There's nothing
you can do. You don't have actual
software uh behind the scenes to deal
with it. With this system, we can pick
it up, redeploy it, uh, on a completely
different, you know, cloud or bare metal
or whatever. We're good to go.
Everything is happy. Um,
delivering releases and updates, uh,
without these mirrors would be
incredibly expensive. Um, CDNs are not
cheap. Fastly, Cloudflare, a you know,
whichever one, they're all very, very
expensive. um at the scale that we are
pushing traffic uh at the scale that we
have the requests uh and it it just it
wouldn't be feasible. There's no way it
would work. Um the the other thing we
found is accurate geoloccation data
improves user experience tfold. Bad data
in, bad results out. So
um what's next? We want to implement
some more statistics within the mirror
system uh similar to what Fedora has.
You can go to the Fedora mirror manager
site uh see like country stats, where
the mirrors are, what's online, what's
offline, um things like that. We want to
do that public facing. Uh we need a
better development environment to get
some more people on board. Um it's it's
all written in Python. It's relatively
easy. It's it's basic Python async.io
stuff. Um but the the
barrier to getting a development
environment up for the code um is not
that easy. So, we need some we need some
a better development environment set up
u you know both guide and maybe some
pre-made images things like that to make
it easier for people to jump in and
contribute. Uh protocol filtering, we
actually got that one done uh earlier
this year where you can like forcibly
select if you only want HTTPS mirrors or
arsync mirrors, what have you. uh
partial mirroring. Um as Alma Linux has
grown in in both number of versions,
architectures, things like that, you
know, our mirrors are now at, you know,
1.5, two terabytes. People want to be
able to trim that down. Maybe they only
want to mirror x86. Maybe they only want
to mirror ARM. Maybe they want to ignore
ISOs. Um in our
status checkers, we've got to be able to
handle that. [snorts]
Um that kind of goes hand inhand with
better mirror validation. Uh we need to
find more efficient ways to check the
status of them. Uh using arsync is a
very obvious option.
And then uh just a little call to
action, get involved uh mirrors.
Linux.org. Uh the code is on GitHub.
Barring that, you know, if you're not a
coder, maybe you have access to servers,
transit, uh contribute mirrors to your
favorite open source projects,
especially distributions. A lot of
projects have moved away from mirroring.
Um, you know, everything from Apache to
GCC to the kernel used to come from
these traditional style mirrors that
we're talking about. A lot of that with
GitHub and and things like that have
kind of moved away from it. Distros are
largely still reliant on that type of
architecture. Um, so you know, set up
mirrors for projects that you care
about. And then we are out of time,
but if there's any quick questions, we
might can sneak them in. There's nobody.
Nobody busting in the back door.
>> I had to like really ram through the
second half of that. Sorry.
>> My question is kind of related to that.
How quickly were you able to pull this
talk together in the 11th minute hour to
fill the schedule gap? [laughter]
>> So actually,
>> thank you by the way. Sincerely.
>> No, no problem. So actually I gave this
talk at Flock last year. However, um I
had I think one attendee because I was
right next to the forjo discussion where
everybody was. So [laughter]
it's still pretty new content.
>> No problem. We've got one behind you.
>> Hello. Uh you said that you don't
combine uh CDNs with mirroring, but
Fedora has a CDN mirror. So, so we
actually do use CDNs to some degree. Um,
repo.allinx.org
itself is a CDN. Um, simply because of
the type of traffic that it receives.
Uh, it is not part of our public mirror
infrastructure. So, when you have a
system and go to get updates on it,
you're not hitting repo. Linux.org. Um,
that's primarily traffic from the
website that's like, hey, go grab an
ISO, go, you know, grab an image. So,
that is behind a CDN. And then the
special mirroring that we have um inside
of AWS inside of Azure they are using
those clouds internal CDNs uh that we
control and manage um within those
environments you know that is the proper
fit. Uh but for the the broader public
mirrorless system uh you know CDN would
be cost prohibitive.
David.
>> Uh yeah. So uh at the end there you were
talking about you know volunteer uh to
be a mirror for projects you care about
and at the beginning you were talking
about how Alman Linux you unlike other
projects you you kind of control the
mirrors you push to them. Is that
>> the tier zeros? Yeah.
>> Yeah. So is that when you were saying
volunteer were you saying volunteer to
be like a a tier zero or is that
>> a volunteer to be a public mirror? So we
push from repo.allinx.org.
Let's say we push from the build system
into our three currently three about to
be four tier zero arsync mirrors.
Everybody then pulls from there. A lot
of other projects and maybe Kevin might
know how Fedora is operated these days
because I don't remember um but a lot of
projects the tier zeros are technically
community contributed as well. It's just
you know selective relationships,
mirrors with big pipes, things like
that. We choose to find people that will
give us the server and the transit and
let us manage the tier zero. Um are are
the tier zeros within Fedora's mirroring
still kind of public contributions
Kevin?
>> So tier zier
one is our community
>> and then is there
>> is there a tier two then
>> then everybody else? Yeah.
>> Okay. Yeah. Yeah. So, a little bit
different lingo, but that's pretty much
what I was alluding to.
All right, we are out of time and people
are coming in, so we'll we'll call it
there. Thanks for coming. [applause]
Thank you.