Submind YouTube summaries
Thumbnail for 2026 08 25 Jenkins Infra Meeting

2026 08 25 Jenkins Infra Meeting

Watch on YouTube

Video summary

The meeting began with an overview of recent releases and team capacity updates, noting that the Jenkins 2.578 release occurred later than usual but without significant issues. The team discussed upcoming absences, confirming that several members would be away starting August 31st, which means the next team meeting on September 1st will proceed with a reduced roster. Priorities for the immediate future remain focused on billing and statistics, while specific updates regarding the search CI "search bomb" issue were deemed acceptable but not critical. The agenda also highlighted upcoming topics such as mirror updates, the retirement of Jenkins ingress, and a potential Kubernetes upgrade scheduled for September. Additionally, progress was made on sponsoring NUDS, with plans to test their plugin in a sandbox environment using manually spun-up micro-virtual machines, though challenges regarding inbound-only agent implementations were noted. Significant attention was given to cloud usage forecasts and infrastructure maintenance. Cloud consumption forecasts showed stability despite some fluctuations attributed to specific agents like search CI, while AWS costs decreased following ongoing cleanup efforts. The team addressed credential expiration issues for various providers, including Terraform, Fastly, and Cloudflare, with plans to renew these soon. A major discussion centered on the shift in release cycles for Oracle and Eclipse teams to a monthly schedule due to increased security vulnerabilities, which impacts Jenkins' patch upgrade campaigns. This change raises concerns about exotic infrastructure support, particularly S390x systems, as some Java versions are not yet fully available for these architectures. The team debated whether to delay updates for these specific platforms or accept the risk, ultimately deciding to proceed with caution while monitoring the situation closely. The latter part of the meeting covered technical improvements and operational challenges, including the successful testing of new update center certificates and the transition away from Windows Server 2019 support in container images. The team identified a critical issue where burstable VM instances in Azure were being throttled and taken offline during high CPU usage spikes, causing pipeline failures; this will be resolved by migrating to more powerful, non-burstable instances once sponsored credits are fully utilized. Furthermore, the discussion addressed mirror selection issues, particularly for users in China who face HTTP 403 errors, leading to a proposal to host a local VPS within China to ensure accessibility. Finally, the team reviewed open issues related to dependency bumps for Node.js 24 and SonarCloud token security risks, emphasizing a strict security posture where any credentials found on public repositories are assumed compromised until proven otherwise.
Read the full video transcript
Hello everyone, welcome to the Genkins infrastructure team meeting. Today we are the 25 of August 2026. Around the virtual table we have myself portal, we have J ready and Mark Wait. Hello folks. Let's get quickly started since we're a bit late with releases. Uh last week Jenkins released 2.578 which started a bit late uh was released with no issue. Thanks S and Mark for monitoring. Nothing specific for these issues. Any questions on this one? and for me. >> Cool. Quick uh head up on the team capacity is off. He will be back on Monday 31. Both Jay and Hi will be off on Monday 31. So they will be alone. I'm back on Tuesday. And Jay, you are also off on the next team meeting. So no team meeting not for you next week. >> And I'll be I'll be off next week as well for the team meeting. Yep. Barack is off from 31 up to 24. >> Actually, I'm off from 26 August through 1 September inclusive >> to >> Yeah. So, same. >> Great. >> Okay. Priorities are still the same. uh billing and statistics for now. Uh you have the list of epics to follow top level topics as usual. Reminder on three topics worth mentioning um so we had discussion with Daniel about the search bomb and search CI not a topic that will bite us. It's acceptable. uh but yeah we will have an update on mirror a bit soon and genics ingress retirement and kubernetes upgrade most probably September for these topics. A quick word on sponsoring nuds. Uh Demian has a test account and a few credits granted for test. starting to test their plug-in still private source because it's an early stage. So I'm in the process of starting what they call sandbox which are micro virtual machines manually from my machines first to test the access and then I'm starting a controller. I will install their plug-in that I'm I've built locally and the goal is to set up the plug-in to see if I can spin up agents quickly from my machine. Um, most probably their initial implementation will fail because the it looks like they only implemented inbound agents. So I will most probably will need to set up a genkins controller with their testing framework in their own sandbox. But that will be still test. I will suggest them to implement SSH as soon as possible because that's the way we want it to go in our case. Uh but yep that will allow us to build bomb and eventually other things somewhere else. Any question on announcement on priorities? Okay. Upcoming calendar next week 1st September without Jay without Mark we will have a team meeting. Uh next weekly release will happen on Wednesday 2nd September at the same time as the next LTS release. Uh security release advisory nonpublicly published in the mailing list upcoming credential expiration in the three weeks. So most of the Azure one have been treated by Jay. We will see later. Thanks Jay. I've opens issues for the Fastly and Cloudflare. We will add all of these in the upcoming milestone and renewal. Uh most of the fastly renewal are it's the first time we renew them. So I'm trying to improve the process step by step before handing it over to J or there is one last issue to open but that will be for next week. Uh it's about Terraform credential expire. That's a bunch of things and most probably I will ask if he can take care of this because he never did it. And no next major event as far as I can tell. >> Any question on calendar? >> No. >> CL budget. Um as your CDF is going as as expected forecasted at 2.2 no matter changes stable. Uh we could move publicates but right now it's middle of summer with everyone days off. That's not easy to do and we have a few prerequisites. uh sponsored subscription we have a peak in consumption this month uh forecasted at 6.9 I think the forecast is broken but that will be a bit uh a bit more than the previous month is culprit our search CI agent because there's been a lot of builds including bomb builds in search CI we saw the same peak um in May so most probably we will end around 5.7 to 5.8 eight. It just moved the the threshold on August 2027 in one year because we have two months with the amount of credit consumed instead of 13. >> Any question on Azure? >> No sense. >> Thanks for confirm uh digital stable nothing specific. AWS we see a decrease which is really good. Uh I hope it will be better but looks like the work on the bum that every did start to pace. Uh so let let's continue to see we still need two to two three weeks. So we will be sure on only on middle September that will help us to see what measure we should take. Uh, by the way, we received an email uh our usual sponsor contact uh at AWS change jobs inside AWS and gave us new contact. New contact at AWS Miller see is changing role. So I will take care of contacting the new ones to to check uh for the credential renewal because it should be around right usually it's September. >> Yeah I thought it was actually even oftent times April or May that they told us hey time to submit your proposals but but it's it's I agree it's time to for us to ask them hey are you going to donate again? We would love to have the donation. >> Absolutely. I will I will send them um first email so they get to know each other and they will be able to ask mil during the transition period. Grog uh storage increase and bandwidth increase. Yep, that's all. Daniel is currently checking cleanups and yep, we might have a few incrementals additional builds that need to be cleaned up. I'm not sure if Darin is continuing. He helped but uh I think we have to automate uh with the elements he gave us in an help desk issue. We have everything needed. So no need to bother Darin on this one and Alolia is going fine. Any question on our cloud usages? >> None from me. >> Okay. So then let's move to the milestone. Thanks Jay. On the topic of the tasks uh which team is keeping the infrastructure up to date free credential rotated no issue and you asked and you were tasked following that question you asked for a slight improvement. Right now we store credential inside terapform state. I wanted to avoid our terraform output because I don't want humans to type the command terapform output and copy and pass things. However, discussing with Jay, we located a new feature in subs that allow subs to insert from estadine or from a subshell values directly inside the YAML file just uh and without involving YQ or GQ. So not only partial update but also they allow uh key and queries directly which mean we should be able to have a terapform output- row pipe subs something or subs with a subshell. So J no emergency but if you want an improvement on this you can work on this that mean adding the output in terapform and setting up subs to have the the correct queries in the topic of infra and maintainable uh Jay was able to get over the two new ad rules so no more uh uh breaking builds. Thanks Jay on this. Anything else on done tasks? Okay, work in progress now on the infra up to date. I've opened a new issue critical patch upgrade campaign uh for GDKs. J your GDK uh July campaign is waiting for the next LTS before being closed. So I moved it in Treyage and it will be back on next milestone. So that's why it's not there. But thanks for that work. And the critical patch upgrade campaign is one step further. Yes, Mark. >> And Oracle and uh Eclipse Team have both stated their intent to that they have now switched to a monthly release cycle. >> So this is right. So this is this is the new standard. >> Um and the the new standard means we've got it. So it will be quarterly we'll get a dot a a change in the third position and 2 months after that we'll get a change in the fourth position twice. >> Okay. Interesting. >> So so they have they have changed and and their explanation is has been Oracle had announced it but it wasn't clear that Taran would follow suit. Taran has now stated they will follow suit. uh and their rationale was security issues are being reported so much more frequently in the in the days of of AI assisted security investigations that they've they've got to switch their release cycle to monthly instead of quarterly. >> Okay. So we'll see how the release are moving on timarine uh because the challenge are the exotic infrastructure such as S390s which are not available yet >> right and and for me it that may mean that we have to say and and for instance I also saw JDK8 is not fully available yet for the places that and JDK1 I think is one that's missing um Windows. So, it's not just the exotics even, but I agree the exotics are absolutely missing. And I wonder if we have to then decide in our patch upgrade campaigns, we will separate system 390 and admit it's just not important enough for us to delay others. >> Yeah, exotic CPUs. What do we do? And and of course that doesn't answer it for our container images that we ship to to users, right? Our infrastructure container images we could do without doing system 390, but we do system 390 with our our >> core container image and we've only got one Java version we can ship. not ten CI/Doc uh whip on some updates. Okay. Yep. Uh because maybe we can proceed as soon as possible. That's already the case on some images. Um but yeah, same on and GDK 81 slower to release. So we already had GDK uh 8 slower to release. to our okay doing it >> and it's right and it's no threat right we only keep it for for ancient things so I I think JDK is not a concern but for me 21 not having an S390X image is is a real thing >> yeah I I believe we should eventually stop doing these issues Jay maybe think about this uh because the issues were about treestrial campaign updates but now with a monthly update I mean it will be dayto-day operation almost for us >> right okay now I see system 390 for Java 25 so so at least one of them has it >> cool so let let's see I guess Timarine will be forced to publish things faster and faster >> I assume yeah >> so yep I'm taking care of this Um, I thought we could be able to ship uh on the weekly release though. Uh, but we won't. So maybe patch surprise next week. That's patch. So that's okay. >> Well, and and I'm not aware of anything in the security fixes that actually is relevant to Jenkins. >> Yeah, same uh check. >> So if we continually shipping 25.0.4 rather than the.1 I I think >> that's okay. In terms of actual security rather than security theater, we're fine. >> Yep. That's one way. Uh, next issue update center routt. Uh, Danielle tested attemp certificate with success. Uh, Pierre opened on genkins call to add the new CA. So we tested new CA and new update center certificates which work because we changed some metadatas and we had to check it before it's embedded in genkins score. So it most probably will end up in a weekly in one or two weeks after that upcoming release and that will be in the next LTS line uh that select that weekly as base. So I guess not next LTS line but the line after given the selection process we can backport uh this D if needed um waiting for the next pier uh for the next core releases with this change um I need to dig on the process on the update center side uh to see if we can have different certificates served based on different CAS or not. We'll see. Windows 2019 support all Docker genkins CI images have dropped 2019 and blog post. Thanks Mark. And the 2.568.3 upgrade guide includes some text copied bluntly from the the upgrade guide that we did for 568.1 saying that um container images for agents or agent container images are no longer supported for Windows Server 2019. So we're we're trying to get the message out in multiple places. >> Yep. Thanks for this. Next steps. Next step is WinPCI builds to use Windows 2025. That's the last step. And then we can get rid of the the stuff. Oh, I haven't categorized this one, but responding very very slowly. So I believe this one is closable. The problem was only on Sunday and confirmed by Linux Foundation to be a peak in usage. Right. Correct. I I think I think it's reasonable for us to close it trusting that they will they will handle it and I'll paste uh any updates they provide into it. Um it's it really was just that one day spike and okay, who knows? Maybe we've got a um an AI scraper who's now created an account that so that they can log in in order to scrape. >> Yep, that's um I don't know if you saw my comments. Uh can we ask the LF or access log from Sunday or at least during the peak just to see if we can if we have a a word pattern or >> I can I can certainly ask them. I'm not sure if they have them, but I'll I'll happily ask. And but yeah, we can close the issue no matter what. Um, mirror support mirror selection in update center. So yes, the user looks like they are the only user of Jenkins in planet in the planet and especially in China. So they tell us what to do. Uh uh yeah not not answering their details but >> that some of few actions. >> Yep. >> And not the the the as far as I can tell the request involves significant development on Jenkins core. It's not just an exercise. It need it would need a new UI on Jenkins core. It would need new logic in Jenkins core and >> we need a new update center >> right and no offer from the user to to provide those things just a request for features. Sorry that at least >> I don't feel like I have capacity to even think about doing such a thing. >> Yep. >> It's not a not a bad request. It's a it's an okay request. It's just that's a lot of development work in a lot of places. or they could do an air gap installation, download all the thing by themselves and build themselves their own genkins. That's what I will end up telling them. But right now focusing on the main thing. So first uh tuna mirror is now China only. So that scopes the issues mentioned by the user because they have an issue. Okay. um tuna admin contacted to ask for clarification because yes uh they they should be able to tell us something and also that's an ill check. If they are not uh answering that means we cannot contact an admin for any reason. So we will and that will be the reason to exclude them from our mirror list at least temporarily. uh let's say I call that staging area if they never answer we can remove them so the proposal I mean um the problem yes we don't have anything in China but if user in China are served HTTP 403 and if they don't answer that can start to be a problem >> yeah I see and I'm I'm hesit I'm not yet persuaded that the problem is the the level of issue that the user seems to indicate so so but but again we don't have many many users inside China and now that it's country specific they really must be inside the great firewall right they've got to be inside >> because I assume Hong Kong oh maybe is >> no Hong Kong has different as sets of numbers that's based on the ASN so I see yeah most of the China test case for for tuna >> exactly >> exactly Hong Kong is using their own close [clears throat] mirror >> got it >> uh so not writing in Not right now. But the proposal is a we can simply find a VPS here in inside China and us the mirror ourselves. Um that means maybe asking CDF for a few bucks. I think I have a budget of 300 bucks yearly for a VPS that should be able to handle it inside China. Uh but I asked the user and they don't provide any organization. So most probably we will end up doing nothing else more and wait for other user to complain or propose help. But we have a solution if you are really stuck in the infrasign and maintainable um define a common node with timeout. That one is u um almost there. Damian plays with code. I'm challenging codes to write things useful. So I'm be I'm being challenged by code. That's that's takeaway. But yeah, I will soon come with a pull request. Third CI um started the task list. Important use a non burst table instance family. We have a burst table instance because we only have peak of CPUs. But right now we have a few cases at least one pipeline in search CI and one pupet uh change when we update thirdbot plugins in both cases that for that absolutely put the VM down. The reason is because it use a lot of CPU as a peak during a few minutes and since we are using credit based CPU for the current controller uh as your hypervisor starts to absolutely panic and stop getting us so we are throttled and the machine is down for one or two hours. We can keep it down for two hours until it's back automatically. However, usually it happens when security team is running things. So that's important that when we will move to Azure sponsored credits, we will use a more powerful instance with um no credit space, no burst and deterministic CPU behavior. Um work resumed on the stats we on infra and PGSQL on census. I've started the local test on my vagrant setup. Now we have a few issue to triage. Um so Jay I realize we don't have an issue for for your unable work. So I'm assuming that until September you are in experimenting and self-arning. That's what we will call it. And next week uh that will be the task for you to prepare an issue describing the goal and scope of what you are working on. Is that okay for you? Since you are experimenting I didn't want it to scope too much but you going back to writing mon so no Monday you are out but for end of week is that okay for you >> okay so that will be issue in triage and then I will add it to the milestone next week okay for you >> yeah sounds good >> and we have a few issue in triage that are worth that will be added to the milestone uh five token renewal Of course, we have a sonar cloud token request. I need to check because we need to access sonar cloud. So, create a genkins infra team account with and select the API key and put it somewhere. Um, or maybe not because putting a token on C ino might be a wrong idea. uh the user created a pipeline library and asked for sonar cloud but yeah not sure about the risk uh we'll ask the question is what happen if that token is uh if someone find code on ci jenkinsio find an issue get the token what happens >> right exact right we've we've held rather rigorously to the notion that cienkinsio is always treated as though it were compromised right it's we we simply assume that a an attacker can submit a poll request that that poll request may have access to any credential that that is available there. >> Yep. >> Okay. >> So, >> and so Rodic, you'll have that conversation with Rodic in the in the help desk ticket. >> Yes. I fear his reaction as usual because he will tell us it's easy to do five minutes with code and all problem are solved and then we will when we will ask him for help he will disappear for three months like he did in the past years >> right well and and and if that happens we turn it off right it's it's pretty simple it >> I think that's what will happen and there is an issue with ill score uh I believe it's on PHS I might ask uh Adrian for help but yeah there is something wrong in PHS report I We'll see. Um, we have these free shoes. Just mentioning them. They are back to tri age until next week or in two weeks. GDK patch upgrade campaign that Jay walked on. Of course, waiting for the new LTS. So, not this milestone. Nodegs 22 to 24 because we need someone to spend time on bumping dependencies on uplink in order to allow NodeJS 24 to work. Uh, I guess this can be a code thing. And finally the decrease bomb cost because survey is off. We have good results but yep back here and it will report cost when it will be back. All the other issue are also waiting for triage. That's all for me. Anything else? Okay. So I'm stopping screen share. I'm stopping recording. So for people watching us see you next week.