Submind YouTube summaries
Thumbnail for 2026 08 18 Jenkins Infra Meeting

2026 08 18 Jenkins Infra Meeting

Watch on YouTube

Video summary

The meeting began with a review of the weekly release process and identified recurring technical hurdles that require attention. The team noted that while last week's release proceeded without major blockers, it was affected by minor caching issues at the end of the packaging job which caused pipeline timeouts. More significantly, the group discussed persistent problems with Fastly, a content delivery network service, which has been returning 503 and 504 errors multiple times. These errors appear to be related to credential expiration or backend misconfigurations rather than transient network issues, prompting a plan to reproduce the problem locally by extracting credentials and running curl commands to diagnose the root cause before implementing fixes in the upcoming release cycle. In terms of team capacity and future priorities, several members announced their availability for the coming weeks, with some taking time off to prepare for family events like a child's school entry. The leadership clarified that while immediate production outages remain the highest priority, the next two months will focus on resuming statistical analysis and managing costs associated with build pipelines. A significant portion of the discussion was dedicated to evaluating new infrastructure technologies, specifically "sandboxes" offered by a sponsor called NOOps. These micro-virtual machines promise significantly faster startup times for agents compared to traditional Docker containers or Azure VMs, potentially reducing build durations from hours to seconds and offering a viable alternative as AWS credits expire later in the year. The agenda also covered various maintenance tasks and updates regarding the Jenkins ecosystem. The team addressed the deprecation of Windows Server 2019 support within Docker agents and discussed strategies for handling outdated mirror servers, particularly those located in India and China that have not been updated in weeks. To mitigate risks associated with unreliable mirrors, the group considered restricting access to specific regions or disabling problematic sources entirely. Additionally, progress was made on closing several long-standing issues, including cleaning up leftover user accounts on trusted CI agents and enabling pull request previews for specific projects using GitHub Actions. The meeting concluded with a commitment to continue monitoring cloud budgets and optimizing build costs through these infrastructure improvements.
Read the full video transcript
April. Hello everyone. Welcome to the Genkins infrastructure weekly team meeting. We are the 18th August 2026 around the virtual table. We are for myself. Wait. Hello everyone. Um let's get started with announcement. Last week 2.577 weekly release went well. We had a minor caching issues at the end of the packaging job. Uh nothing blocking except uh yep it it just make the pipeline stuck. I thought there was a time out >> and that's why I haven't checked. >> Um >> that was my surprise this morning to see that >> today's release did not start until I clicked on the procedum last week. proceed that was an error on the >> and it was blocking all the the build chain. So, but thanks for taking care of that survey. Not not the fault of not your fault, not the fault of everyone. Thanks for checking >> and I also forgot to open an issue about the >> Oh, yep. >> P.S at least two weeks if not three I don't remember. Let me take care of that. Yes, I have a few days off, but I'm here next week. Unless you really want to write an issue now. >> No, I'm good. >> That's what I talked um this week. So, 2.578 is in progress. Are you able to follow it up? Do you need fallback? Um so the release job started I think uh it was about 1 hour one hour and a half that I started it. So I think it did no I did not check yet I will check. >> It's not about checking is are you able to check in the upcoming hours? >> Yeah I am able to follow the chat. >> Yeah and I will do the controller release too. >> Okay cool. I'm just asking to to be sure if you can't someone will take over. Thanks. >> It's we're we're actually now within a time range where I could even watch it through. So every if you're on vacation and not available, don't be shy. Just say, "Hey, somebody else take it and we'll we'll cover it." >> Uh I still have this afternoon to work, but uh no, I'm good Mark for this week. Thanks. >> Great. Uh, anything else on the weekly releases? >> Nothing on my side. >> So, it's for everyone just so you know, it's most probable that we will have the fastly issue again. We had it last week. So, it's it happened at least three times. >> So, we know it's consistent. Uh, it's a 503 error. looked like the initially transient problem but the fact that it's happens all the time means there is something different like a badly httpc coded response and so the real error is something else. The issue is to track that minor problem because usually the the the purge happens soon after the packaging job when the change log is uh pushed by genkins. genkinso only uncash a few CSS or pages. However, fastly detects the changes and after one hour max is purging everything. So yeah, it's just optimization. The road, as I told very privately, the road to fix that problem is first to reproduce the issue locally. So extract the credential and run the curl command on on the machine to see what the response holds or change the pipeline for next week. So we collect uh the body of the answer. Um we will have credential rotation if it's related to a credential issue because I think we messed up fastly one uh one trimester ago. Maybe that's related to expiration. So surprising it's a 503 but we never know. >> It's a 504 last week. Oh, so that one is a transient one. Okay, for sure. Okay, cool. 504 is usually uh you reach the ingress the the ingress uh front ends and the back ends are not available. So that's definitively a fast infrastructure error with 50 five or three could be uh bad could be a badly developed uh back end. Anything else on weekly releases? No. >> Okay. Moving to announcement team capacity. So both and high are um off tomorrow. Um I'm back Monday. Her is back on 31 August in two weeks. Uh but both of us are hitting the road later today. So that's why I was asking. Uh we will be async after this meeting. Jay, you mentioned can you confirm that what I wrote is okay. You plan to be off on 31 August and 1st September. Is that still? >> Yeah, that sounds right. >> Cool. Also, I'm I will be off on 31. So, Monday 31, you will have a slow day for you. >> Yay. >> Back from holidays at a good rate. >> Yeah. Uh it will be the preparation for my son's entry in the uh elementary school. So, >> yep. So only the I emergency and the rest uh you let it you let it wait for the for Tuesday. Um priorities we haven't changed the priorities right now we have we still have the costs that are around the corner but we we were able to delay of one trimester. So most of the work should go now on unable and resuming the statistics. These are the two top level priorities that should take over everything except production outages. A few topic worth mentioning. Uh I wrote explicit time frame of the upcoming two months. These things will hit us in the up in September or until middle October. Um we've realized that there are bump PCT builds happening on third CI which explain the picss we see once every every trimesters. Uh they use their own pipeline. So every improvements that is working on on the bum are not present on search CI. Is it a problem? We don't know yet. We need to measure the cost. I have a belief that in fact we shouldn't care because most of the build on CCR are usually six to 10 hours long. The agents are running on Azure virtual machine on the sponsored subscription. So we can eventually afford if it's three or 400 per month. We need to check. However, it doesn't kill to have a I think you already shared some pointers to Daniel privately about the split test because that will decrease the pressure on the CCI controller like you did on CI genu. >> Yes, >> the context is different. They don't exactly evaluate the same thing. Not the not always the same line. We don't use the same agents and they might not have the same caching rules as we had in search CI. Most of the time caching is bad because that's not what they want. In CI we care about caching more and more. So different context, different problem statement, different way of solving it eventually. So it's it won't bite us but we will have to discuss it. Anything else to add over there on your side that I could have forgotten on that topic? >> Uh, no. >> Cool. Moving on. Mirror bits. Uh, it's just a patch, but they don't do semantic versioning. We are running uh uh 061. We are running 061 but there are a few internal changes that could uh have impact or improve situation. So I've put the polar request on draft. We need to evaluate because yet update center and get genkins could be impacted. Engics ingress retirement it's officially not updated even for CVS. It's the ingress we use. That's a right balance but we need to start thinking about other solution. Switching ingress implementation is one solution. Switching all of our M chart to API gateways. That's basically the other solution. And Kubernetes 1.35 that's the usual schedule September beginning of October but we will have to run upgrades. Are there other level topic worth mentioning that will impact us on the two months for for you folks? Um, so we've got Windows 2019 somewhere in the list that it's going away. I know I've still got the action item on the blog post. Okay, it's in there. Good. >> Okay, just one word. I've been contacted. Uh, thanks Mark for uh follow for forwarding them to me. Uh, NOUps is a company right now that advertise for web free AI product in a sus. I'm not really sure about that part. However, they provide a nifty um infrastructure as a service system. They call that sandboxes. That's the same idea as the Docker sandboxes. These are micro virtual machine that are blazing fast to start and that can run on in their case uh they hide but they run on many clouds and these machines uh could be interesting. They contacted us because they want to sponsor the genkins project. Of course, they want visibility. That's a good deal. And uh they showed me their technology that could be interesting for running agents in CIA. That's the idea they have. So that mean they already checked the project and they already targeted some parts. So they did a really due diligence work. Honestly, that's really cool. Their technology is impressive. Um I've seen a few CIS doing the same kind of things but yeah the mach the sandboxes are starting in less time to write a JSON file on my file system than to have a virtual machine running. So that means we can have agents that start in one or two seconds fully fledged agent with a docker engine running which is really impressive. As such, um, I pointed them to, hey, maybe you need to start with a genkins plug-in that will allows Genkins to start and stop agents. So, they already added me. Right now, it's private. I will push them to make it public as soon as possible. They started from the Kubernetes plug-in uh because that's one of the let's say most interesting API management I we knew. I said Azure VM and Kubernetes are the one we used. I say don't go from EC2. Absolutely not. Um so yep they provided something that is already interesting. They should provide me a first level of testing so I will have an account with a few bucks per month for now. Uh discussing with one of the targets here could be the bomb. We don't we are not able to evaluate the cost of the bomb compared to other builds. It's still hard to to tell things apart. However, that will mean um a bunch of workloads that will be blazing fast to start agents even with airways split test. That will be interesting. >> And because it doesn't do a lot of IO, most of our plugins build or genkins or at are doing a lot of IO's and nodes doesn't provide good IO's for now. >> So bum could be interesting. Second thing, we could increase the bomb scope because uh we could start running test with docker. That's not the case today since we run the bum agent inside the pod agent. So we'll see what they have in in mind for us. But that that could be really interesting especially with the AWS credits going away end of year if they don't renew. that could be a good alternative for at least part of C and can say I will let you know as soon as they have more public things so I'm not the only one to be able to test their thing so thanks for them any question on priorities team capacities >> good no questions for me that sounds great actually it's at least very interesting micro virtual machines things like Firecracker. This sounds This sounds wonderful. I mean, of course, we've got to evaluate it, but it just sounds so promising. That's great. >> Um, we have a team meeting and a weekly next week. Next LTS is beginning of September. Everyone will be here. Jay will be back and hi as well. So, that that should be okay. >> Oh. Oh, one more upcoming calendar. We're choosing the LTS baseline that doesn't really affect the infra team, but it's it's crucial for for the for our users and for all sorts of people. Um, and let's see, I think that is next Wednesday that we choose the LTS baseline. >> Thanks. For the future reader, we may want to add the Windows 2019 drop from the current chant as it will happen soon or should we just keep the issue on discussing it on the desk >> on the desk? I don't see that as a major thing. >> And I I had it wrong. Our LTS baseline selection is tomorrow. So my mistake. Got it. Yeah, >> 2019 is a really tiny subject. We don't have a lot of users on this. >> It's a lot of work for us. I agree. But >> um no security advisory publicly announced. It's August and we have a bunch of issues to open for upcoming credentials that we missed on the past two weeks. Um as I mentioned the fastly ones. uh but everything should be doable by J high after next team meeting. So most probably these issues won't be created until next team meeting and no next major event questions okay cloud budget we are in line um Azure is steady we we fer between 2.1 to 2.2 two uh we are still online. We could gain some money on moving publicates but we have to choose our battles. So we consume credits and and yeah we will have to eventually find solution for next year but it's summer. Uh we had a peak of consumption on Azure sponsor subscription which messed up the forecast but that's the bombci uh things that we that happen one or twice per trimester. Uh so other than that we are steady on Azure. Same for digital ocean and on AWS we had a slight decrease but still too soon to conclude especially um we it will be interesting to watch this cost weekly in the upcoming two to three weeks to see the impact of the split test work that did. We cannot conclude if it will decrease the cost. In any case, it will improve the situation because cost is not the only brick here. >> Um that will gives us a very good foundation to think about moving to node ops, moving to virtual machine, changing the kubernetes setup because that will depend on where CI genio will go in the upcoming months and we need that foundation in any case. So good job on this one. I hope that that will help especially the optimization part outside the infra that you made about not rerunning all tests when only a few tests fail that will drastically decrease the amount of builds. >> Yep. >> Um I guess we'll start to see result in September most probably given we are already at half the month of August. We are still under the foc. But yeah, we should decrease if we want to reach the end of year. Grog uh increased. Yeah, I don't have anything else to say about Grog. Uh the bugs that doesn't that doesn't update the the Maven metadatas XML on a few elements such as the one of the parent pump. Uh the latest field on the XML is not updated after the release. Uh I saw the problem happen during the core release but we have something that recalculate him. So yep I don't know what's happening at Grog but it's not they are not doing well >> and Alolia nothing to say. That's all for me on the cloud budget. Are there questions? >> No questions for me. >> Okay let's continue. Uh the notes we were able to close the following items. There was a full disc on a trusted CI genu. Uh in fact there was a leftover uh 17 gigabyte almost uh GE user that was a remnant of manipulation Daniel and I did in April when working with the update centers. So I removed the user. We forgot to clean it up and yep uh a few elements around pull request preview for GSOC projects. Um so we enabled them. So that was basically allowing pull request preview on jobs that were targeting a branch which is not the main branch. That's not a security risk. Now that we have a job which only runs PR preview on Nefraci, I declined the automatic deployment of the GSO 2026. Though it's not technically possible without creating a third job only for that because right now we have one job that only target the main branch and one that only target pull request. But the way the GitHub um uh scanning works, you specify patterns for either branch names or pull request destination branch name. That's the same field for both. Which mean if you had g you will have a branch and a branch main uh for the pull request. If you remove main, you won't have pull request preview for pull request targeting main. So we are stuck. We could use our build strategy. So the branches will be discovered but never built. But that start to create empty artifact in infraci. So instead I came with a solution which is a they could just open a draft pull request from JSO to main branch and that one is a pull request preview. Each time they will merge whatever pull request on their ones that will update their own preview and we're done. Does it make sense? Did I miss something or do you need more clarification on these two? >> Sounds reasonable. >> Cool. Um, plug-in modernizer stats. They wanted pull request previews, but they already have their own CD system inside GitHub action because we don't know if that project will ever go to our production. They are fully autonomous with gha. So I ping I shared with them a gha that create pull request preview using GitHub pages that create subdirectories or something like that. Arve a word on the new mirror. Uh nothing particular. Uh it went well. uh after I um I don't know if it's uh because the mirror runner switched to archive that mirror went up almost immediately or if it was just a matter of waiting a bit but uh we have a new mirror in Europe, Netherlands and uh it's already in the top three mirror in that region looking at mirror stats. Cool. Uh the reason is because um when you after you created an initial mirror, even if it's disabled, it's still being scanned by a background task, but that scan takes time because it's usually five to 700 gigabyte of files to scan. So even if you don't download all files or uh the the scanning the file tree can take 20 to 14 minutes to 14 minutes, sorry. So it could be that that the the background task was finishing during all the the dies and retries you had. So when when you looked again and enabled again the scan was finished and say okay we are good. Okay, >> but maybe it could be archives and the time file with a different time stamp that could have changed from the other mirror because instead of chaining it was using a more recent source could be I don't have anything to conclude but that could be one of this. >> Yep. So yeah, thanks again Peter Paul for this. >> We are really grateful for that. Um and there was uh uh credential rotation nothing specific to mention work in progress on the topic of keeping the infrastructure sane and maintainable. So the bomb cost which might not be cost outcome but still needed. So it's in production already the split test. Yeah, we are now uh since last week the bomb is using fixed split test with an amount of 20 agent per split. It's uh values that give us relative same amount of same total duration time a bit lower but more or less the same ballpark but using for example a weekly test is running on 20 agent instead of 300 now. Uh so for this issue the next steps are following the cost probably and on more personal side on not infra related because not really discuss but I have a request already for balance split using split test result from the parallel exeutor plug-in and the next uh thing I want to add is skipping um testing repositories that uh were in success in the previous speeds of a job. >> So, and already I'm seeing what feels like a better reliability from the the split that's running now with the fixed number of agents. Thanks very much, Erve. It's as the bomb bomb release lead this week, I'm very grateful. Thank you very much. I was a little embarrassed that we released a version of the bomb on Monday inadvertently because I hadn't edited the uh the the crown job. Sorry about that. >> I made a mistake too. By the way, uh I wanted to get a build on master with uh unit test result recorded in the proper stage. And even if I commented out the incremental function in the pipeline, it still produce uh bomb release. So I've marked it at pre pre-released and I added ammonition uh in the release body. >> Oh, okay. Yeah, you it it's it's a if we release there was a time when we released much more frequently. We release weekly only to save cost. So there's no shame in us having done a release. >> Okay. >> Yep. Nothing else on this issue. >> Um so is that okay? We keep it on the milestone uh with the let's follow up cost until the end of the month. Is that okay for you? >> Yes. >> Okay, we'll comment. Keeping in milestone to follow up costs. Okay. >> And I just saw the weekly build has been published. The artifact for the weekly build has been published to Artifactory. >> So we're we're making progress. taking a look quickly to see if the B pass or not. Thanks folks. Next topic the statistics. So um uh whip on adding post SQL on census genkins VM. Uh that's the current status. we need to add a posgrsql database. Uh and the idea is to on this machine install everything required to have the same kind of environment that uh like what Andrew was doing. We'll need eventually credential uh or yeah but the goal is to integrate one month in that machine and eventually get the CSV and commit them manually from one of our machines. That's will be the first step. Then we will discuss the next steps. Any question? >> No, that's that's great that it's reopening. Thank you that it's coming back to life. Much appreciated. Every step of progress is appreciated. Sensus. So no more bus factor automation will come next. Uh Jay, anything on the Adelind thing? >> Uh yeah, I just briefed you guys on the issue. So with the new version of packer images after 2.14, so had went from 2.12 to 2.15 and uh because of that we were getting a bunch of false positives on the new rules. So uh specifically with our you with the username with the username keyword use and user. So um yeah the proposal is to ignore both of them globally in the pipeline library. Uh and on the order of the priority this task should start right after the JD upgrade start. >> Thanks. Yeah, that's it. That's true. >> And Jay, do you know did they fix the bug that I reported to them about 2.15.0? >> It's on Windows and we don't run Adelint on Windows Docker file. So, not on the infra. So, we don't know. >> Okay. >> Is a fix of the issue you opened on Adelint, but they did not release it yet. >> Ah, thank you. Okay. And and reminder, it's not because it's marked as released that it is unless we use the latest binary and show it's fixed. It's considered not fixed. Either it's in production or not. The rest are So yeah, the answer is not yet. >> Oh, but we if we don't have code, we won't ever see it in production. So yeah, >> absolutely. So not yet is still good enough, but yep, be careful. Uh not with time out and retry on old until demier is back. I have a proof of concept for this one. Uh finally same forci on old. So I will remove them and put them back to triage for both because I'm not there until next team meeting. uh on the keep infrastructure up to date GDK patch campaign upgrades Jay just a summary >> the current uh the current the current progress is uh the JDK8 upgrade is being patched in backer images as we speak and apart from that the only remaining work was that the Turin images so since it takes uh it takes time for them to release their images. We're going ahead and just downloading the archive directly. So the work is being done on that. The pull request is currently in work and it should be done soon and uh we should be able to close the issue after these two points are addressed. Did I miss anything? >> That's good for me. Thanks. Uh update center CA rotation being Danielle. He will integrate the new certificate. I asked him because I thought it was only one pull request on the call, but he want to run um a full set of tests on a local update center. So since I'm not available for a few days and we start to it start to be a pressing matter um I we discussed and he will take care of this. I will ping him if not done in one week to see if I can help. Uh, not yet. All done except uplink back to triage unless someone has time to work on the uplink parts. Uh the reason why we can't upgrade NodeJS to 24 on uplink is because uh dependencies in npm. And finally, Windows 2019 support dropped from Docker agent. >> Uh it's not dropped yet. It's not released. It's not in production as you said. >> Good point. That was a trap and you didn't fail for it. That well done. Very very well done. >> Remaining Docker SSH agents >> and Wimpy >> and remaining the blog post. Uh we we had that note from oh dear who was it I forget who somebody told us that >> Mikei oh Jesse Jesse Glick about cloudbases one of their container images on it. So they needed to know what Yeah. >> pointed him to the issue in the to aent but post will be nice. >> Absolutely. Mark, are you still planning to take care of it? Cool. Many thanks. >> Yes, absolutely. It's I sorry that I haven't done it yet. I've been busy on other things, but I see no reason not to get that blog post out very soon. Okay. And once all of these has been fixed, we can drop support from packer image. That's all for me. Um on the support triage, I only have one uh that's that user who want to select their mirror. Uh I propose to >> at least at least avoid some but yeah. Yeah, >> people who read the issue know I propose to add it on the next milestone because >> at least I believe uh we should >> uh eventually block or disable the tuna mirror because uh if they randomly serve four or three errors that means something is doing is really wrong with them. either they don't provide the right HTTP code that should be 4 to9 if it's on the rate limit and if they start to censor users close to them we should disable them. I agree with the user on that part. We are really grateful that they sponsor us and provide a mirror but if the mirror is restricted to their own certet network that's thought to be a problem for users like the one who opened the issue right. Yep. Um my issue is that we only have one person. So my proposal is that we still contact them and say, "Hey, we have a user that has been served four or three errors. Can you help us?" >> Yeah. And okay, my concern is that's also the only mirror we have inside the Great Firewall of China, right? So, >> so tuna is tuna is important to a big group of of uh Jenkins users even if they're flawed. >> Oh yeah, they're good points. I thought there were there was a second one in China. >> I'm looking at satellite right now. >> I'm not aware of one now. Maybe there's there's definitely one in Taiwan, but the Great Firewall of China does not consider Taiwan part of China. Okay. Um I haven't checked in details the curl output yet they shared because maybe that could be the same thing as the B Russia like yeah we could restrict tuna to only China first >> right >> so all the outside uh traffic will move to Taiwan. >> Yeah we have we have a mirror in Singapore one in Japan. Yeah. Yeah. In fact, we've got se two or more mirrors in Singapore last I remember it. >> Exactly. That's why we could eventually start making Tuna a China only mirror and see if it changed. The user did not answer about the location of their EDLless server which is which is why I'm annoyed. They state thing but they don't provide a lot of rock solid data. So that's why I don't want to to go that direction. Um but my goal is still to contact to know saying okay you have user that you have at least one user here that complains about that do you have more information about that uh because in any case if they restrict I can understand the user problem that means we will need to set a course of action to have a mirror in China >> right and now when I look at our map it looks like our India mirrors are gone >> no I right now >> you do good. Okay. All right. I >> if it's on my screen share it just because that's all mirror bits is grading my location to other mirrors. So that's why I guess I'm not in India due to the latency. We have a better latency to some part of Asia than India from Europe. >> Yeah. Well the table like that doesn't show it even. >> Looks a link I've just posted open the link I've just posted. >> Yeah. But these are stats not mirror list. So that's different thing. >> Ah okay. >> Okay. Sister and Albony have not updated their content since a few day. That's why they are excluded. >> Okay. All right. Which may also increase the demand on on tuna, right? Because that means requests from India will be balanced between Singapore and and China. Japan probably much lower pri priority. >> Yep. Oh, yep. Ah, I thought we were monitoring the mirror's uh age on get genkins but that's not the case. So I guess we have an issue sitting somewhere. >> Okay. >> Uh yep we should contact both of them. I don't know why it's not updated. >> But yeah in any case India uh or that part of Asia we could still uh create a machine on digital mirror. >> Oh all right right >> they have data center in that area. So that's not a blocker that require a bit of work but yeah it's still doable. >> Okay. >> Thanks Harvey. So yep given the discussions and the potential action point here um also contact Indian India located mirrors. Uh Jay maybe I will delegate that to you during the upcoming days. I will do it asynchronously as they look out of date since 20 days. Okay. I don't have any other issue to add to the milestone. Do you folks? >> No. >> Okay. Are there other topics you want to discuss before we stop the recording? >> None from me. >> No. >> Cool. Okay. So, I'm stopping screen share, stopping screen record. So, for people watching us, see you next week. If I can find Where is the recording though? I've lost the recording tool. Insane. Okay.