Video summary
The meeting began with an overview of recent releases and team capacity updates, noting that the Jenkins 2.578 release occurred later than usual but without significant issues. The team discussed upcoming absences, confirming that several members would be away starting August 31st, which means the next team meeting on September 1st will proceed with a reduced roster. Priorities for the immediate future remain focused on billing and statistics, while specific updates regarding the search CI "search bomb" issue were deemed acceptable but not critical. The agenda also highlighted upcoming topics such as mirror updates, the retirement of Jenkins ingress, and a potential Kubernetes upgrade scheduled for September. Additionally, progress was made on sponsoring NUDS, with plans to test their plugin in a sandbox environment using manually spun-up micro-virtual machines, though challenges regarding inbound-only agent implementations were noted.
Significant attention was given to cloud usage forecasts and infrastructure maintenance. Cloud consumption forecasts showed stability despite some fluctuations attributed to specific agents like search CI, while AWS costs decreased following ongoing cleanup efforts. The team addressed credential expiration issues for various providers, including Terraform, Fastly, and Cloudflare, with plans to renew these soon. A major discussion centered on the shift in release cycles for Oracle and Eclipse teams to a monthly schedule due to increased security vulnerabilities, which impacts Jenkins' patch upgrade campaigns. This change raises concerns about exotic infrastructure support, particularly S390x systems, as some Java versions are not yet fully available for these architectures. The team debated whether to delay updates for these specific platforms or accept the risk, ultimately deciding to proceed with caution while monitoring the situation closely.
The latter part of the meeting covered technical improvements and operational challenges, including the successful testing of new update center certificates and the transition away from Windows Server 2019 support in container images. The team identified a critical issue where burstable VM instances in Azure were being throttled and taken offline during high CPU usage spikes, causing pipeline failures; this will be resolved by migrating to more powerful, non-burstable instances once sponsored credits are fully utilized. Furthermore, the discussion addressed mirror selection issues, particularly for users in China who face HTTP 403 errors, leading to a proposal to host a local VPS within China to ensure accessibility. Finally, the team reviewed open issues related to dependency bumps for Node.js 24 and SonarCloud token security risks, emphasizing a strict security posture where any credentials found on public repositories are assumed compromised until proven otherwise.
Read the full video transcript
Hello everyone, welcome to the Genkins
infrastructure team meeting. Today we
are the 25 of August 2026.
Around the virtual table we have myself
portal,
we have J ready and Mark Wait. Hello
folks.
Let's get quickly started since we're a
bit late with releases. Uh last week
Jenkins released 2.578 which started a
bit late uh was released with no issue.
Thanks S and Mark for monitoring.
Nothing specific for these issues.
Any questions on this one?
and for me.
>> Cool. Quick uh head up on the team
capacity is off. He will be back on
Monday 31. Both Jay and Hi will be off
on Monday 31. So they will be alone. I'm
back on Tuesday. And Jay, you are also
off on the next team meeting. So no team
meeting not for you next week.
>> And I'll be I'll be off next week as
well for the team meeting. Yep. Barack
is off
from 31 up to 24.
>> Actually, I'm off from 26 August through
1 September inclusive
>> to
>> Yeah. So, same.
>> Great.
>> Okay. Priorities are still the same. uh
billing and statistics for now. Uh you
have the list of epics to follow top
level topics as usual. Reminder on three
topics worth mentioning um so we had
discussion with Daniel about the search
bomb and search CI not a topic that will
bite us. It's acceptable.
uh but yeah we will have an update on
mirror a bit soon and genics ingress
retirement and kubernetes upgrade most
probably September for these topics.
A quick word on sponsoring nuds. Uh
Demian
has a test account and a few credits
granted for test.
starting to test their plug-in
still private source
because it's an early stage.
So I'm in the process of starting what
they call sandbox which are micro
virtual machines manually from my
machines first to test the access and
then I'm starting a controller. I will
install their plug-in that I'm I've
built locally and the goal is to set up
the plug-in to see if I can spin up
agents quickly from my machine.
Um, most probably their initial
implementation will fail because the it
looks like they only implemented inbound
agents. So I will most probably will
need to set up a genkins controller with
their testing framework in their own
sandbox. But that will be still test. I
will suggest them to implement SSH as
soon as possible because that's the way
we want it to go in our case. Uh but yep
that will allow us to build bomb and
eventually other things somewhere else.
Any question on announcement on
priorities?
Okay. Upcoming calendar next week 1st
September without Jay without Mark we
will have a team meeting.
Uh next weekly release will happen on
Wednesday 2nd September at the same time
as the next LTS release.
Uh security release advisory nonpublicly
published in the mailing list
upcoming credential expiration in the
three weeks. So most of the Azure one
have been treated by Jay. We will see
later. Thanks Jay. I've opens issues for
the Fastly and Cloudflare. We will add
all of these in the upcoming milestone
and renewal. Uh most of the fastly
renewal are it's the first time we renew
them. So I'm trying to improve the
process step by step before handing it
over to J or
there is one last issue to open but that
will be for next week. Uh it's about
Terraform credential expire. That's a
bunch of things and most probably I will
ask if he can take care of this because
he never did it.
And no next major event as far as I can
tell.
>> Any question on calendar?
>> No.
>> CL budget. Um
as your CDF is going as as expected
forecasted at 2.2 no matter changes
stable. Uh we could move publicates but
right now it's middle of summer with
everyone days off. That's not easy to do
and we have a few prerequisites.
uh sponsored subscription we have a peak
in consumption this month uh forecasted
at 6.9
I think the forecast is broken but that
will be a bit uh a bit more than the
previous month is culprit our search CI
agent because there's been a lot of
builds including bomb builds in search
CI we saw the same peak um in May
so most probably we will end around 5.7
to 5.8 eight. It just moved the the
threshold on August 2027 in one year
because we have two months with the
amount of credit consumed instead of 13.
>> Any question on Azure?
>> No sense.
>> Thanks for confirm
uh digital stable nothing specific.
AWS we see a decrease which is really
good. Uh I hope it will be better but
looks like the work on the bum that
every did start to pace. Uh so let let's
continue to see we still need two to two
three weeks. So we will be sure on
only on middle September that will help
us to see what measure we should take.
Uh, by the way, we received an email uh
our usual sponsor contact
uh at AWS change jobs inside AWS and
gave us new contact.
New contact at AWS
Miller see is changing role.
So I will take care of contacting the
new ones to to check uh for the
credential renewal because it should be
around right
usually it's September.
>> Yeah I thought it was actually even
oftent times April or May that they told
us hey time to submit your proposals but
but it's it's I agree it's time to for
us to ask them hey are you going to
donate again? We would love to have the
donation.
>> Absolutely. I will I will send them um
first email so they get to know each
other and they will be able to ask mil
during the transition period.
Grog uh storage increase and bandwidth
increase.
Yep, that's all. Daniel is currently
checking cleanups and yep, we might have
a few incrementals
additional builds that need to be
cleaned up. I'm not sure if Darin is
continuing. He helped but uh I think we
have to automate uh with the elements he
gave us in an help desk issue. We have
everything needed. So no need to bother
Darin on this one
and Alolia is going fine.
Any question on our cloud usages?
>> None from me.
>> Okay. So then let's move to the
milestone.
Thanks Jay. On the topic of the tasks uh
which team is keeping the infrastructure
up to date free credential rotated
no issue and you asked and you were
tasked following that question you asked
for a slight improvement. Right now we
store credential inside terapform state.
I wanted to avoid our terraform output
because I don't want humans to type the
command terapform output and copy and
pass things. However, discussing with
Jay, we located a new feature in subs
that allow subs to insert from estadine
or from a subshell values directly
inside the YAML file just uh and without
involving YQ or GQ.
So not only partial update but also they
allow uh key and queries directly which
mean we should be able to have a
terapform output- row pipe subs
something or subs with a subshell. So J
no emergency but if you want an
improvement on this you can work on this
that mean adding the output in terapform
and setting up subs to have the the
correct queries
in the topic of infra and maintainable
uh Jay was able to get over the two new
ad rules so no more uh uh breaking
builds. Thanks Jay on this.
Anything else on done tasks?
Okay, work in progress now on the infra
up to date. I've opened a new issue
critical patch upgrade campaign uh for
GDKs.
J your GDK uh July campaign is waiting
for the next LTS before being closed. So
I moved it in Treyage and it will be
back on next milestone.
So that's why it's not there. But thanks
for that work. And the critical patch
upgrade campaign is one step further.
Yes, Mark.
>> And Oracle and uh Eclipse Team have both
stated their intent to that they have
now switched to a monthly release cycle.
>> So this is right. So this is this is the
new standard.
>> Um and the the new standard means we've
got it. So it will be quarterly we'll
get a dot a a change in the third
position and 2 months after that we'll
get a change in the fourth position
twice.
>> Okay. Interesting.
>> So so they have they have changed and
and their explanation is has been Oracle
had announced it but it wasn't clear
that Taran would follow suit. Taran has
now stated they will follow suit. uh and
their rationale was security issues are
being reported so much more frequently
in the in the days of of AI assisted
security investigations that they've
they've got to switch their release
cycle to monthly instead of quarterly.
>> Okay. So we'll see how the release are
moving on timarine uh because the
challenge are the exotic infrastructure
such as S390s which are not available
yet
>> right and and for me it that may mean
that we have to say and and for instance
I also saw JDK8 is not fully available
yet for the places that and JDK1 I think
is one that's missing
um Windows. So, it's not just the
exotics even, but I agree the exotics
are absolutely missing. And I wonder if
we have to then decide in our patch
upgrade campaigns, we will separate
system 390 and admit it's just not
important enough for us to delay others.
>> Yeah,
exotic CPUs.
What do we do? And and of course that
doesn't answer it for our container
images that we ship to to users, right?
Our infrastructure container images we
could do without doing system 390, but
we do system 390 with our our
>> core container image and we've only got
one Java version we can ship.
not ten CI/Doc
uh whip on some updates. Okay. Yep. Uh
because maybe we can proceed as soon as
possible. That's already the case on
some images. Um but yeah, same on and
GDK 81
slower to release. So we already had GDK
uh 8 slower to release. to our okay
doing it
>> and it's right and it's no threat right
we only keep it for for ancient things
so I I think JDK is not a concern but
for me 21 not having an S390X image is
is a real thing
>> yeah I I believe we should eventually
stop doing these issues Jay maybe think
about this uh because the issues were
about treestrial
campaign updates but now with a monthly
update I mean it will be dayto-day
operation almost for us
>> right okay now I see system 390 for Java
25 so so at least one of them has it
>> cool so let let's see I guess Timarine
will be forced to publish things faster
and faster
>> I assume yeah
>> so yep I'm taking care of this Um, I
thought we could be able to ship uh on
the weekly release though. Uh, but we
won't. So maybe patch surprise next
week.
That's patch. So that's okay.
>> Well, and and I'm not aware of anything
in the security fixes that actually is
relevant to Jenkins.
>> Yeah, same uh check.
>> So if we continually shipping 25.0.4
rather than the.1
I I think
>> that's okay. In terms of actual security
rather than security theater, we're
fine.
>> Yep.
That's one way. Uh, next issue update
center routt. Uh, Danielle
tested attemp certificate with success.
Uh, Pierre opened on genkins call to add
the new CA. So we tested new CA and new
update center certificates
which work because we changed some
metadatas and we had to check it before
it's embedded in genkins score. So it
most probably will end up in a weekly in
one or two weeks after that upcoming
release and that will be in the next LTS
line uh that select that weekly as base.
So I guess not next LTS line but the
line after given the selection process
we can backport uh this D if needed
um waiting for the next pier uh for the
next core releases with this change
um I need to dig on the process on the
update center side
uh to see if we can have different
certificates served based on different
CAS or not. We'll see.
Windows 2019 support all Docker
genkins CI
images have dropped 2019
and blog post. Thanks Mark. And the
2.568.3
upgrade guide includes some text
copied bluntly from the the upgrade
guide that we did for 568.1 saying that
um container images for agents or agent
container images are no longer supported
for Windows Server 2019. So we're we're
trying to get the message out in
multiple places.
>> Yep.
Thanks for this. Next steps. Next step
is WinPCI builds
to use
Windows 2025.
That's the last step. And then we can
get rid of the the stuff.
Oh, I haven't categorized this one, but
responding very very slowly. So I
believe this one is closable.
The problem was only on Sunday and
confirmed by Linux Foundation to be a
peak in usage. Right. Correct. I I think
I think it's reasonable for us to close
it trusting that they will they will
handle it and I'll paste uh any updates
they provide into it. Um it's it really
was just that one day spike and
okay, who knows? Maybe we've got a um an
AI scraper who's now created an account
that so that they can log in in order to
scrape.
>> Yep,
that's
um I don't know if you saw my comments.
Uh can we ask the LF or access log from
Sunday or at least during the peak just
to see if we can if we have a a word
pattern or
>> I can I can certainly ask them. I'm not
sure if they have them, but I'll I'll
happily ask.
And but yeah, we can close the issue no
matter what.
Um, mirror support mirror selection in
update center. So yes, the user looks
like they are the only user of Jenkins
in planet in the planet and especially
in China. So they tell us what to do. Uh
uh yeah not not answering their details
but
>> that some of few actions.
>> Yep.
>> And not the the the as far as I can tell
the request involves significant
development on Jenkins core. It's not
just an exercise. It need it would need
a new UI on Jenkins core. It would need
new logic in Jenkins core and
>> we need a new update center
>> right and no offer from the user to to
provide those things just a request for
features. Sorry that at least
>> I don't feel like I have capacity to
even think about doing such a thing.
>> Yep.
>> It's not a not a bad request. It's a
it's an okay request. It's just that's a
lot of development work in a lot of
places. or they could do an air gap
installation, download all the thing by
themselves and build themselves their
own genkins. That's what I will end up
telling them. But right now focusing on
the main thing. So first uh tuna mirror
is now China only. So that scopes the
issues mentioned by the user because
they have an issue. Okay. um tuna admin
contacted
to ask for clarification because yes uh
they they should be able to tell us
something and also that's an ill check.
If they are not uh answering that means
we cannot contact an admin for any
reason. So we will and that will be the
reason to exclude them from our mirror
list at least temporarily.
uh let's say I call that staging area if
they never answer we can remove them so
the proposal I mean um the problem yes
we don't have anything in China but if
user in China are served HTTP 403 and if
they don't answer that can start to be a
problem
>> yeah I see and I'm I'm hesit I'm not yet
persuaded that the problem is the the
level of issue that the user seems to
indicate so so but but again we don't
have many many users inside China and
now that it's country specific they
really must be inside the great firewall
right they've got to be inside
>> because I assume Hong Kong oh maybe is
>> no Hong Kong has different as sets of
numbers that's based on the ASN so I see
yeah most of the China test case for for
tuna
>> exactly
>> exactly Hong Kong is using their own
close [clears throat] mirror
>> got it
>> uh so not writing in Not right now. But
the proposal is a we can simply find a
VPS here in inside China and us the
mirror ourselves. Um that means maybe
asking CDF for a few bucks. I think I
have a budget of 300 bucks yearly for a
VPS that should be able to handle it
inside China. Uh but I asked the user
and they don't provide any organization.
So most probably we will end up doing
nothing else more and wait for other
user to complain or propose help. But we
have a solution if you are really stuck
in the infrasign and maintainable um
define a common node with timeout. That
one is u um almost there.
Damian plays with code. I'm challenging
codes to write things useful.
So I'm be I'm being challenged by code.
That's that's takeaway. But yeah, I will
soon come with a pull request.
Third CI um started the task list.
Important
use a non
burst table
instance family. We have a burst table
instance because we only have peak of
CPUs. But right now we have a few cases
at least one pipeline in search CI and
one pupet uh change when we update
thirdbot plugins in both cases that for
that absolutely put the VM down. The
reason is because it use a lot of CPU as
a peak during a few minutes and since we
are using credit based CPU for the
current controller uh as your hypervisor
starts to absolutely panic and stop
getting us so we are throttled and the
machine is down for one or two hours. We
can keep it down for two hours until
it's back automatically. However,
usually it happens when security team is
running things. So that's important that
when we will move to Azure sponsored
credits, we will use a more powerful
instance with um no credit space, no
burst and deterministic CPU behavior.
Um work resumed on the stats
we on infra and PGSQL on census. I've
started the local test on my vagrant
setup.
Now we have a few issue to triage. Um so
Jay I realize we don't have an issue for
for your unable work. So I'm assuming
that until September you are in
experimenting and self-arning. That's
what we will call it. And next week uh
that will be the task for you to prepare
an issue describing the goal and scope
of what you are working on. Is that okay
for you? Since you are experimenting I
didn't want it to scope too much but you
going back to writing mon so no Monday
you are out but for end of week is that
okay for you
>> okay so that will be issue in triage
and then I will add it to the milestone
next week okay for you
>> yeah sounds good
>> and we have a few issue in triage that
are worth that will be added to the
milestone uh five token renewal Of
course, we have a sonar cloud token
request. I need to check because we need
to access sonar cloud. So, create a
genkins infra team account with and
select the API key and put it somewhere.
Um,
or maybe not because putting a token on
C ino might be a wrong idea.
uh the user created a pipeline library
and asked for sonar cloud but yeah not
sure about the risk uh we'll ask the
question is what happen if that token is
uh if someone find
code on ci jenkinsio find an issue get
the token what happens
>> right exact right we've we've held
rather rigorously to the notion that
cienkinsio is always treated as though
it were compromised right it's we we
simply assume that a an attacker can
submit a poll request that that poll
request may have access to any
credential that that is available there.
>> Yep.
>> Okay.
>> So,
>> and so Rodic, you'll have that
conversation with Rodic in the in the
help desk ticket.
>> Yes. I fear his reaction as usual
because he will tell us it's easy to do
five minutes with code and all problem
are solved and then we will when we will
ask him for help he will disappear for
three months like he did in the past
years
>> right well and and and if that happens
we turn it off right it's it's pretty
simple it
>> I think that's what will happen and
there is an issue with ill score uh I
believe it's on PHS I might ask uh
Adrian for help but yeah there is
something wrong in PHS report I
We'll see. Um, we have these free shoes.
Just mentioning them. They are back to
tri age until next week or in two weeks.
GDK patch upgrade campaign that Jay
walked on. Of course, waiting for the
new LTS. So, not this milestone.
Nodegs 22 to 24 because we need someone
to spend time on bumping dependencies on
uplink in order to allow NodeJS 24 to
work.
Uh, I guess this can be a code thing.
And finally the decrease bomb cost
because survey is off. We have good
results but yep back here and it will
report cost when it will be back.
All the other issue are also waiting for
triage.
That's all for me. Anything else?
Okay. So I'm stopping screen share. I'm
stopping recording. So for people
watching us see you next week.