Submind YouTube summaries
Thumbnail for 2026 09 15 Jenkins Infra Meeting

2026 09 15 Jenkins Infra Meeting

Watch on YouTube

Video summary

The meeting began with a review of the recent release cycle, noting that while the announcement was successful, an error occurred during the metadata calculation step on Artifactory which required manual intervention to complete. The team outlined several key strategic initiatives for the upcoming milestones, including the migration from Puppet to Ansible 26 and the upgrade campaign for Kubernetes 1.35. Significant attention was also given to infrastructure optimizations, specifically focusing on Azure usage costs which showed a net decrease compared to the previous period, allowing the team some financial breathing room for the next three months. Additionally, the discussion covered the retirement of the NodeOps plugin and plans to test new capabilities with it, alongside efforts to migrate Kubernetes provider resources to versioned Terraform resources to improve state management. A substantial portion of the meeting was dedicated to addressing technical debt and improving reliability within the Jenkins infrastructure. The team discussed cleaning up GitHub links to redirect security advisories from private Shira issues to public Jenkins advisories, a task that automated scripts completed for nearly 300 repositories. They also addressed the issue of Artifactory repository permissions failing due to deprecated API versions, leading to a decision to update the Repository Permission Updater tool. Furthermore, the group tackled the problem of Docker image pulls failing intermittently by proposing a switch from public Docker Hub mirrors to explicit Amazon ECR registry proxies, aiming to reduce bandwidth usage and improve stability within the AWS environment. The migration away from Windows Server 2019 agents was also finalized, with the team successfully moving remaining consumers to Windows 2025 templates to modernize their build clusters. Looking toward future automation and cost management, the team detailed plans to centralize website build pipelines and optimize statistics generation by migrating from MongoDB to PostgreSQL. A critical path item involves automating the processing of user survey data, which currently relies on manual intervention due to a lack of access to necessary GPG keys held by a specific individual; once automated, this will remove human bottlenecks and allow for more frequent data integration. The meeting concluded with an important discussion regarding scheduling conflicts that affect team members in different time zones, particularly those in India and the Americas. To ensure broader participation without sacrificing daily collaboration, the team proposed alternating meeting times between earlier and later slots on even and odd weeks respectively, a change intended to balance attendance across Asia, Europe, and the Americas while maintaining near-daily communication through other channels.
Read the full video transcript
Hello and welcome to the Jenkins team meeting. We are the 15 15th of September with Damian de Portal J ready um and Robson and N for the announcement uh last week weekly went well. Uh today's release uh complete was completed but there has been an error on artifactory on the calculate metadata step uh completed the remaining step manually and opened an issue about that um for the announcement Daniel will be off this Friday uh we have our ala map that's need an update on our current topic uh in set link. Uh why two of them? Sorry. The big worth mentioning are just um replacement of puppet by unable to 26 uh upgrade campaign determine where we will put CI. Jankkins.io IO and uh Jenkins mirror infrastructure for download and update center. The topic worth mentioning as engineress retirement kubernetes 1.35 upgrade and we can also mention that uh node ops is uh uh we want to test something with node ops and dam is testing their plugin. Any question or remark on this announcement? No. Okay. Um I just have a few note on the issues. Uh the three of us have issues and or epics to write this for the upcoming milestone and as soon as possible. >> Uh you have one around automation of the Docker release. I have one around the Azure uh infrastructure optimizations incoming optimizations. Uh and Jay, you have one uh you have an issue to open as a child of the epic run on civil uh to track your work, the work you are currently working on because there are no issues for this one. So I makes I'm asking you to start writing an issue to define the definition of done the value of this one and what it's adding in in the context of the epic. Yeah, sounds good. I'll create an an epic for it. >> And don't forget to associate your B request to this one. >> Um, thanks. That's all for me on the epics and the announcement >> in the upcoming calendar. Next meeting will be next on Tuesday, next weekly the same. Next LTS will be at the end of December with the Lesk related to the back port that we have to >> September. >> End of September. You say December. That's why >> September. >> I said December. Sorry. >> Yep. No worries. Just wanted to be sure. >> Uh next security release will be plugin only tomorrow. Yeah. Uh nothing more to add about that one. On the upcoming credential expiration, we will have C drinking ark for publisher to a new for next month. No next major event. Any question remark on doing calendar beet as your CDF uh subscription we are uh still good at uh 2K forecast um and we still have to move public test to the sponsored subscription on your sponsored subscription. Uh we are good too. Uh the forecast is quite lower than last week. Uh so taking uh this amount monthly will have it for a bit more than 11 months. digital. Um nothing particular to say we are still good uh the forecast on AWS. uh we have a net decrease we are almost 30% less than the last the same period last month and our forecast is uh it was 11% less last week and now is 22% less so we are quite good adjusting the current rate uh we have now till midFebruary with the current The next uh cost prediction effort would be centered out Azure on factory usage. Uh the storage is quite the same than last week a bit increase and the forecast has decreased uh for September but we can't really do anything more about it. So yeah, any questions remark about about uh budget? >> Yes, >> that that's great work. Uh that means we can we can be uh a bit chill for the upcoming three months. So great work everyone. uh going to the issue uh on the done part I've asked to be downgrade removed from the release team uh I was in the release team in Jenkins Jenkins core repository um most probably uh I've been added to that team when I've made some when I was release lead for previous FTS release. Uh so now uh I've only tried on it. Um could not log in artifactory. Uh I don't remember this one dam it was just yeah someone that uh failed to login to Artifactory to release this plugin and then they switch to CD instead. >> Yes. That's that remove the need to for them to login on Artifactory. Um but the problem is rooted on repository permission updator which uh is using duplicated APIs on Artifactory which creates many problems including this one. >> Yeah. Okay. That's >> we might have other users have with the same problem. Uh clean up auto link reference. Uh that was uh um an ask from Daniel to redirect security. Um so the link in GitHub to redirect security dash number short um reference to link to Jenkins advisory instead of the Shira security issue that are private by default. and made no sense. So team run a script that I made to automatically upgrade a bit less than 300 repository score. Uh this is a ping scoring under security show 0%. Um the database wasn't updated since a few some times ago and Adria shar fixing scoring and we now have hopes with the last run date that we will be able to add data monitor on it to be alerted when it's still ms build not available on JS.io. That was an old uh issue that uh we fixed uh few times ago and that I closed. Persist disable performance announcement g config as code um that was CI to junction.io config settings that wasn't stored as code and is now done. So we won't have surprise when regarding to on keep up to date Bit to the last version. Uh I let you talk about this one Daniel. >> Yes. So we check the change log. No breaking change inside Mirror Bits. Uh don't be fooled by the patch increase because Mirror Bits is now having a rate of one release per year and it doesn't follow semantic versioning. So be warned each time there is an update even if it's a patch it could break things. Uh yep that's how they work. In that case, a few new features including u an improved way to locate your to calculate which mirror to send you and to spread the workload across many mirrors when you have many candidates. So that should decrease pressure on some mirrors and increase on others. Um the only bad surprise we had is the common line because we use the mirror bits common line on one of the trusted agents. They don't provide built binary. So we introduced build of the binaries on bus x64 and 64 CPUs. We do it with our docker images and use the docker images to retrieve the the command line. That was the fastest one. We could do differently but right now that's the the way we we went and no issues during deployment. Great. Thanks. Any questions about this one? J or Rob? No. Um, migrate Kubernetes provider resources to version resources. Uh, do you want to word about it, Tamil? >> Nothing specific. Um, we were missing the moved. Um, the Kubernetes provider in Terapform is not implementing the moved resource, but it implements the removed and import. So I use a trick of removing without deleting the cloud resource. It's just removed from the state and also the import with the references. Uh that was a bit verbose and required the second pull request to clean up the migration but it did work without requiring to run terapform command line >> on the drop windows 2019 support. That was a longstanding issue. We had a remaining uh consumer of the Windows 2019 agent that was windp used by Jenkins score as DL to control process on Windows. Uh I migrated it to this tool version 145 and Windows 20 25 agent. So we were then able to clean up every agent template uh in our controllers and the packer image um Windows 2019 the um as your gallery on AWS MAI. So we don't have yeah no Windows 2019 anymore in alpha closed as not plan. Um this one uh I don't have anything to add about it. Uh someone proposed to push less um more morphing to Maven Central and Max tell them that no we want we don't want that and this one was someone asking for permission but it's instead for for the repository permission updator. >> Yep. They have the documentation. They could have their AI chatbot or at least that should be something they could they could they could they could train their AI. Anyway, that's me being snarky. Uh but anyway, I closed the issue not because uh I don't want to help. I closed the issue because help desk is for u things where we have an action. But the only action here is the the plug-in contributor. >> Cool. >> Work in progress. Um that one is a new one since yesterday. Uh cmake.org is unavailable. So crawler is falling. I snooze alert for 8 hours. Uh for free release plugin uh releasing plugin I'll let you take no >> work in progress on repository permission updater. uh we have to update uh the the API version for artifactory that we use right now the version one which is duplicated is answering uh http 406 I haven't seen that error uh quite often uh which means uh the the way permissions are created updated uh is invalid now so I guess they switched off some part of the v1 or they are uh brown outing anyway we have to update Repository permission updator. I've I ate that but I vibe coded something. Uh the problem now is testing and I I need help. Uh team Yakom gave me a way to test partially which resulted in me breaking permissions in production because we don't have a staging Artifactory. Artifactory is managed by Grog. Maybe I should set up an Artifactory instance with a Docker container. But that will be an empty one. At least it will uh it will show a few elements. It won't test totally but at least it would increase the the test uh the test surface. So I will wait for team and see tomorrow. But yep I need people who know it's not about Java language. It's about knowing the RPU and the history. And since Daniel Beck is in early days we don't have our rock solid person on that area. and his teammate Kevin is currently uh busy this week. So right now it's on hold. Uh we'll work with team and eventually other contributor who will be interested to to provide things but yeah we have hit the limit of vibe coding >> uh transform your proxy. We have to set up um to to have the possibility to add an optional registry mirror in front of the from directive in the docker in the alpine and docker file >> and all all of them uh because all of them are mirrored in a public Amazon ECR registry. >> The goal is not to use our mirror. What we call a mirror is usually a transparent proxy or a registry mirror set up at the Docker C. When you do docker pool Ubuntu, it that's the docker engine which takes care of saying oh I see I have a registry mirror. Let me append uh prepend the registry URL and try to download image from it. If it fails or don't know, then I will fall back to docker up. Here we want something different because that technique I just described we already have it in Amazon for CI agent can say agent but sometimes it fail and quite often so the problem seems located in is it network is it the way doc thing I don't know um I so the proposal is instead of trying to fix something that looks really really deep inside docker buildics moi and moi engine instead let's use What Amazon provide us? Amazon mirrors docker hub inside their own explicit ECR registry they provide rate limit if you are unauthenticated of course like docker hub but if you have an AWS internal IM authentication we should be good. So the goal is to have that explicit proxy in the name like you like the pull request you started yesterday. Uh instead of from Debbian that would be from public.cr PCR something/debian since the mirror they provide the same tags the same layers so we should be fine and that should be less download bandwidth from docker and more from internal AWS services is my explanation clear >> um in the pull request you closed yesterday uh I've put some links that I believe will be reported in an issue if that's okay for you. >> Yep. >> I did it from the phone so I did not have all context but does it make sense? >> Yes. >> Because that's a relious topic. So that's why I prefer having discussing this. >> Yes, we tried the just putting http equals true in the build config file. I think we already >> Yes, it's already been the case for months. The problem started to appear a few weeks ago. >> Okay. >> It it used to be random um until June. The problem was always an issue on our own transparent mirror that was failing. We were able to correlate build failures where the docker engine switch back to docker rub to an error log during the initial tentative to download from the our mirror because this is full network issue inside the case cluster whatever now we don't have any error logs any picss in resources so nothing that say there is a problem so that's why it looks like the problem is on the client side and it's somewhere in the building's behavior that suddenly start in some cases is to you try to use HTTPS instead of HTTP despite all configuration telling it not to >> exactly >> feels like a bug in in buildings but yeah it's really hard to the to investigate so the lazy engineer I am let's not have the problem that will be easier to solve >> um creat That's a cloud token still waiting. Uh we we we may want to create service a GitHub service segment to with read permission to be able to get personal access token for it token for it. >> Still need to spend time on it. But yes, that's the next step. No blockers for now. onure tuna China based it might not mirror is reliable for Chinese users um we are doing a run test right now it's disabled um we want to wait until end of September to see if there are any complaints and if there are not we will keep it disabled >> we will delete it >> delete it yeah >> I don't see any reason to keep the configuration honestly uh because they don't answer and I would rather remove it than having a mirror where the administrator don't don't answer to our request. There is already another Alibaba Bay uh located mirror with the same problem. We added it. They had issues. We contacted them. They never answered for months if not years. So we removed the and put it out. So these are not supported mirrors. Just a note I've there is a new epic around mirrors. I've created one subtask on that epic uh related to China. Uh Uh the goal will be for us to build our own mirror in China for get genio for this that will replace tuna but that's something we manage and we pay for so we have to find sponsoring or money or location and uh for update center as well. So we will be able to provide a mirror for the update center metadatas for updates genkins >> just that's that's just a note but that will solve that problem in that case. Yeah. Um, any questions about this one? No. Uh, keeping infrastructure sign number. I'm finalizing a build website pipeline library function to centralize the build of all website or web components to reduce mounts for everyone align process either a implementation detail so contributors will still be able to work on the build website function itself while not having to take care of which credential to use which storage account which file sha and to it's also it will also enforce uh bits security so um should be ready soon uh optimize cost and maintenance by merging windows 2022 and Windows 2025 template. I think uh we could use uh Windows 2019 uh uh 25 um node windows on I guess cluster used for release but uh unfortunately the terapform provider as your terform provider is not ready yet. So instead I will I will open an issue that will be subtask of the release from release automation which will be to use virtual machine instead of p in the ins cluster to for the release process uh which is mainly building uh the MSI install for Windows. uh on move data to for the usage that seat generation on true from on machine. Um do you want to take the mic? >> Uh we are doing great progress. So all statistics until August included have been uh done. So everyone in the team know how to do it manually. uh and uh we've set up uh yesterday census.jenkins Jenkins IIO the machine where we were able to integrate statistics it's now a permanent agent for trusted CI which mean we can uh resume work on automation the goal will be to have a pipeline that does exactly the same thing on what we run uh manually after a quick brainstorm with yesterday looks like we should be able to run that process once a week instead of once a month so we get earlier feedbacks and less work to less time to spend on the integration because almost every steps support partial updates uh getting the importing the logs from the remote machine then to the database we've confirmed that this part this this works by port it uses air sync and I had issue on the import command line to database so I had to retry and it started where um where it stopped after the last successful import so that works that will work very well the only question we we have is the CSV report generated that should be pushed to the GitHub repository. That's the only question. Does it break uh the world rendering of the old and new websites or is it okay is just show partial data. That's something we should try but from the beginning um that's the report part uh not survey after your review. I've also added an item regarding the PG pass file. >> Yeah. uh that should be generated and removed by trusted CI as a credential and we can remove the one in the for the root user. So now the only upcoming things will be updating the existing pipeline that used to run on trusted for census which is already defined in infra statistics on the master branch. So I will just revamp that one because it used to be something getting data starting a MongoDB database loading the data and exporting things but right now we have the posgrrisql database. So we should be we should be good. Uh then automation we will run the pipeline for each month. So I will just create a big queue and it will run it will run once per month. So, yep. And then we will have the weekly build at the end when every data will be integrated. We should reach March 2026. We cannot go further. I need to contact Kosuk as a next step. I didn't have time to to do it to do this. But yeah, there is a wall issue about Kosuk involvement on that part. He's the creator and he's the only person with access to the GPG key which allow decryptting data. His work is to have an automation on his basement to anonymize the data from our users and that's the anonymized data that we work on and that anonymized data has not been updated since March. So that will be the next step but automation of every other monies will will be already a good step because that mean we have removed Andrew and anyone here from the critical path. So thanks for the work folks. Thanks for the help. I don't have anything else unless you have question. >> Good for me. >> Yeah. Um we are now uh monitor outcomes as score computation routine and let us when waiting is too long. the pling score. It's about uh what I said earlier when Adrian fixed uh the pluging score. He also al he also added the runtime the last computation time that time in the probe that we can uh used for data monitoring. Any questions about this one? Jay, the the definition is clear for you for this one. You don't have any question about >> Yeah, I've already started the P on the steness monitors. So once it's ready for review, I'll let you guys know. >> Okay, perfect. Thanks. >> Question >> on the keep. Sorry. >> Sorry, no question. Okay, it's okay. That's right. >> Um, keep the tour up to date. We have to migrate uh our public cluster to um the sponsor suppression. We need uh to write down the issue and do it. >> Yep. Uh I have a draft but it was uh postponed. Um and I have an experiment to run but I won't I might probably will ask either or J to pair with me on the experiment. I will write on the plan at least the problems. It's about the public IPs. I would want to avoid changing the public IPs used for inbound connections. Uh I want I would want to have someone else with me so we can take not too and it's not uh hidden inside my brain. >> Yeah, no problem for me if you discuss when we just do it prepared center with C at the rotation. Um I don't know if the pull request on core with the new certificate has been merged but yeah uh in progress. So nothing to add on my side on this one. And uh resume database weekly update for uh bits. Uh do you want to take So I've described in the issue the the challenges uh regarding uh location on where the routine is running on and how do we tell mirror bits uh to reload the database because we are using h mode and the command reload is only for single process so for each node which mean we need to iterate over all node of the cluster and tell it to reload. uh that involve different set of permissions than triggering a cube cut rollout restart uh which is a lower set of permission than what I'm using right now config with full admin permission that's absolutely not good but at the first step to be sure valid we validate all the steps uh I need to open an RF on mirror bits to make the command mirror bits reload or at least the joip database reload clusterwide command like when you add mirror or update mirror. Also, I discovered the mirror bit update command that I missed that was released last year with 0.6.0. Uh that command should be run on at least the two last mirror we integrated one in Canada and one in Romania because there might be wrong that that will be a oneshot thing part of that of that update. Uh work in progress. Uh I have the script. I've described every problems. I focused on uh statistics but I need to resume the work on this one unless someone is able to help. >> Um if you need don't hesitate to ask um issue triage now uh we have the back port to take care of if we want to have it ready one week before the release. Um, I'm looking at user which one I don't see. I don't Yep. Um, I don't see anyone any other. Uh, let me look if someone open new is shoot otherwise we're good to go. So the no the crawler build is already part of the current milestone. Okay. No new issues. So no nothing to to put here. Yeah. Um any question any remark before we close do you have any question for us or subject you want to discuss? >> Uh thanks. so much uh in this case no and second time I prefer to listen I learn more in next day next meeting uh I can to to talk more but I understand some um but I need erh uh understand more. Uh in this case uh your planning or this this plan about create optimization and the planning uh about the the opra. Okay. I understood about the Yes. Okay. Um let's finish then and uh see you next week. Okay. >> Um something for uh I have two two topics. So one for Robson first. >> Next week we might have a few beginners issue that we could give you if you are interested that do not require credentials and that could be a contribution to the infrastructure. Usually these beginner issues Jay could testify that's uh where he started um are updating things that need to be updated because they changed such as a tag a version of a software in areas where we know where to get it from. We have that tool name update CLI that's a Golong command line um that we use everywhere to track updates. And so we have a few uh areas we can search for to-do track with update CLI in our YAML files. We have many areas where there is a public IP of a service version of a software that is not tracked but pinned and we want to update. So that system would like dependabot open pull request saying hey there is a new version of that and propose a change so we can evaluate and integrate this without any risk. These things sometimes we forget about them. They are not a lot of work but we forget or we don't have that time for a lot for a few amount of work. So that could be a great contribution that will help understanding the different component of the platform without requiring a deep commitment like writing a software that runs in production. >> Great. In this case, we um verify the the the the link the the the GBS. >> No, no, no worries. No expectation from you yet. Uh usually we have to first write an issue that describe the expectation first. So that's easier for you to get started and know what what we want and what we don't. That helps you scops the effort. So you will have a written issue. I will try to make an effort on for a specific one that I just cooked yesterday. So that should be something really scoped that you can start working on. >> Okay. Okay. Thanks. >> If you are interested, of course, you have the right to say no. It's just just proposal to control. >> Uh sure. Sure. I I I stay here. I I I prefer to to improve and to Yes. Okay. In this case, okay, for me. Thanks. I have another topic. Um I will want to request for uh time time change for this meeting because that time slot is really start to be really hard for me to attend. Uh it's middle of the noon time. uh especially with the kiddo I have at home that can be sometimes complicated. Uh the the challenge we have had in the past has been the time because if we do it too late that will be hard for Jay to join because uh located on India time zone and if it's too early that will be really hard for our American friends to join. Uh we haven't had Mark for two weeks now. Mark was the constraint on the US. I don't know Robson. Uh what time zone are you based on? >> I moment 7 and 40 5 minutes in the morning. I'm from Brazil. >> Okay. So, so yep. So if it's start to be if it's earlier that will be hard for you to join then I assume. Uh so I'm I'm requesting this uh one of the solution could be alternating on even weeks that could be earlier and on odd weeks that will be later. So we cannot have everyone at the same time. Right now European time zone will be okay. So that mean Jay you should only be able to attend or expected to attend once every two weeks. So the the information you will lose will be compensated by the fact that we meet almost every day in our case. Uh and so that's the same for you Robson. There will be one easy schedule for you to attend once every two weeks. The goal will be to cover both Asia and America. Yes, of course. >> I'm going to propose uh write this down in the matrix element channel. So, we will we will have feedbacks. Um I know team sometimes used to attend with the previous time slots because it was in the UK that was the noon break for him cannot attend during his working hours. Uh so if we totally change that could also have people not able to attend during their working day in Europe. Uh we'll see but at the noon between noon and 2 p.m. for us in the US in the in Europe. Uh it's a bit odd for me. So that's why if you objection please raise but yeah we used to do that the alternate time years ago with Olivia. So I'm going to propose this and see if someone objects. If no one objects then I will propose a time a time schedule. Is that okay for everyone? >> Yeah. >> Cool. I will write this down and we will see each other. I let you finish survey to close. >> Yep. See you next week. Bye. Annie, can I stop sort of recording