Submind YouTube summaries
Thumbnail for FOSDEM infrastructure review

FOSDEM infrastructure review

Watch on YouTube

Video summary

The FOSDEM infrastructure review highlights a system that has remained largely stable since 2020, relying on Cisco ASR routers and C7 switches while transitioning key components like DNS from BIND to CoreDNS. To manage monitoring effectively, the team utilizes Prometheus, Loki, and Grafana for observability, with public dashboards hosted externally due to bandwidth constraints at their primary location. A significant portion of the technical setup involves a robust video infrastructure featuring "video boxes" equipped with Blackmagic encoders that convert SDI and HDMI signals into processable streams. These devices are connected via an automated switching system and include local SSD storage as a safety net against network failures, ensuring that conference content is preserved even during connectivity issues. The rendering farm responsible for processing these video streams consists of 27 laptops donated to the community after their use at FOSDEM, reflecting the event's commitment to open-source principles and resource sharing. Operational challenges this year included managing a new 10 Gbps uplink alongside an existing connection, which required careful configuration of BGP sessions that initially caused temporary connectivity loss while propagating across global networks. The team also addressed critical security updates for their DNS infrastructure and resolved issues with flapping network interfaces by eventually consolidating onto the higher-bandwidth link to prevent congestion during peak usage times. Despite these hurdles, automation tools like Ansible played a crucial role in streamlining deployments, though maintaining legacy configurations from previous years required significant effort before deciding to rebuild certain systems entirely. The transition was facilitated by centralized logging and monitoring which allowed engineers to quickly identify problems such as missing internet lines or interface failures without needing physical presence for every issue. Communication during the event is handled through a Matrix installation that serves both internal coordination among volunteers and public interaction, alongside custom radio frequencies for on-site support teams. Regarding future plans, while some network hardware like switches will be retired due to ULB's infrastructure upgrades, other equipment such as cameras and video boxes are brought in truckloads each year from external providers or reused internally. The team operates primarily on their own Belgian hosting provider during the event but utilizes cloud resources briefly for specific tasks like live cutting before reverting back to local hardware. Looking ahead, decisions regarding server replacements and further investments will be made after a post-mortem analysis next week, with an ongoing goal of modernizing aging routers and ensuring that new volunteers can seamlessly integrate into this complex yet evolving technical ecosystem.
Read the full video transcript
okay hello everyone and uh welcome to the last lightning Talk of the conference I hope you've enjoyed yourselves this will be the fosm infrastructure review same as every year presented by Richard and basty so normally I do this thing but uh B has been helping a ton and the ball of spaghetti and spit and duct tape which I left him he turned into into something usable uh so I'm just going to sit here on the side and I'm here around for the Q&A but for the rest it's buy and it's his first public talk for realy so give him a big Round of [Applause] Applause well thank you I hope will I will not screw this one up okay so we'll have about 15 minutes and 10 minutes of talk and 5 minutes for Q&A and I hope it's somewhat interesting to you um first the facts the core infrastructure hasn't changed that much um since the last fdom 2020 we're still running on its Cisco um ASR 10K for routing acl's net 64 and DCP we have already um we reused um several switches that were already here from the last fosm they owned by fos them um these are Cisco C7 uh C 3,750 switches we had our old servers which are now turning 10 this year um they were still here and they will be replaced next year um we have done like all the years before everything um with promethus Loki and grafana for monitoring our infrastructure because that's U what helps us um running all the conference here um and we've built some public dashboards and this we just put it out to a VM outside of ulb because we were running out of bandwidth like the years before um and I'll come to that back later we have a quite beefy V video infrastructure um you might have seen this one here it's a video capturing device it's called the video box here at fom um it's all public it's all open source except one one piece that's in there and uh you can find it on GitHub if you try to uh build it yourself um go ahead just grab the GitHub repo and clone it these devices there's two of them one at the camera one here for the for the the presenter's laptop they sent their streams to a big render farm that we have over in the K building where um like every year our render Farm is running on on some laptops um so laptops send the streams off um to to the cloud from hatzer and from there we just distribute it to the world so everyone at home can see the talks and um yeah we have some sort of semi-automated Rev on cutting process um those of you who have been talking here maybe have known s revie for years this is the first time um it's running on kubernetes um so we are trying to go Cloud native as well with our infrastructure um just to show how all is um been held together um this is our video boxes I don't know if you can see it um we got those um black magic encoders here that are turning the signals that we get like SDI HDMI into a useful uh signal that we can process with our banana pie that we have in there uh everything's wired up to a dump switch here and then we go out like here and have our our own switching infrastructure inside those boxes there's some SSD below here where we just in case of network failure dump everything to the SSD as well so hopefully everything that was been uh talked about at the conference is still captured and available in case of a network breakdown those boxes also have a nice display for the speaker so we can see if everything's running or it's not running um which makes it easy for people to operate these boxes here you don't have to be a video Pro you just have to wire yourself up to the Box you see a nice fom logo and see okay everything's working and and you're done and everything's get sound set out this is like how the video system is actually working we have all this can be found on the on the GitHub you don't have to take screenshots for that um um if you like to see it you can we will tear down this room uh afterwards so we can just everyone can have a look at the infrastructure we're using because it's not being used after this talk um you see it's quite some interesting things to do this is um the instructions that that um all our volunteers get um when they wire up the whole um buildings here um on one day on Friday um so they're not here but they should be given some Round of Applause because they're volunteers that are doing really the hard work and building up on one day the complete fuz them so maybe it's time for Round of Applause for them yeah so here we have the the the another thing this is also um on the GitHub rep where you can see like where's your where's something coming from we have the room sound system this is what you're hearing me through and we have a camera with audio gear speaker laptops and that's all getting pushed down until it someone reaches your device down here there's a ton of services processing it uh in between and this is all done with um almost all done with open source software um EX for the encoder that's running in there um which is from Blackmagic design so how is it processed we have a rendering farm this are the laptops it's 27 this year um for those of you who don't know um those laptops are being sold after them so you if you want one you can grab one this year they're already gone but for next year maybe you want to have a cheap device you can get have them with everything that's on them because we literally don't care for that you can have it um because e things been processed after um after Theos them um you can see it we have some some some wrecks where we just put them four wise and we have 247 of them no 27 of them um we have some switch infrastructure um that process is for used for processing all that stuff and this one's not running out of bandwidth but we coming back to what's running out of bandwidth you might see this mess over here this is our internet and looks like every common internet on the on the planet um and this is like um our safety net we have a big box here where all the streams go and this will be sent out to Bulgaria to the video team right after F them so we have a really offsite copy of everything so the challenges for this year dns64 um all the years we've been running on bind n since ages and we switched to cordan S just like testing it on Sunday of fosm 2020 and we we really saw a significant reduction CP usage and that's why um we stuck to um core DNS since then and this year we also replaced the remaining bind installations that we handled for all the internal DNS and all other recursive stuff that's been used here to to provide you internet access Richie always used to give you some timelines and that's what I'm trying to do as well um there were times when it was mentally challenging for people building up F them um we got better by year by year by year by doing some sort of Automation and getting people used to know what to do what to do and have um everything set up before that um we installed routers um you see that there's a slight it's it's getting better year of the Year this year we had like um a very we we thought it it would be okay from what we know we just set it up in January and everything worked we came here here uh on the 5th of January I think um put everything up and it just worked which is great which gives you some sort of um things not to care about because there were other things to care about the network uh to have it up and running here um took us a bit longer this year than the last year's because we were playing around with the second Uplink that we got um we used to have one one gbit Uplink um last week we got a 10 GB Uplink and we thought okay just enable that and play with it and it would turned out to be not that easy to getting up both of the bgp sessions running and um doing it properly um that's why it took us a bit longer this year the monitoring was um also one thing which really helps us to understand if posm is ready to go or if something has to stay very very late here um the last years we've been very very good at that basically in January everything was was done like the last of January but um it's January this time it were in the first half of January everything was set up and was running and it worked and yeah that was really great because some people actually got some sleep at fosm um didn't uh need to stay here very long because everything was all pre-made and it just just go and look at the dashboard okay this is this is missing this is missing and just so okay just have them all checked the video buildup took a bit longer this year um because of we getting old and Rusty at that also in very many uh new faces that have never built up such a great conference um this is why we took us a bit longer and the video team also yeah I think they they got the the the least amount of sleep of all of the stuff that was running the conference this was The Story So Far um we closed fom 2020 I was also there at 2020 um 2020 was really one of the best ones we ever had from from a technical perspective um we had everything running via anible just like one command and then wait an hour till everything is deployed and you're gone cool have some beer some m in between and everything was cool then we had this pandemic just for me like a week after fom everything went down and we you know you had fom 2021 and 2022 there were no conference here at the ulb um so we had no infrastructure to manage was quite okay we had to do some other things like most of you have um learned that we have a big Matrix installation to run and the fosm conference um and the compan um and help you with communicating um during the conference um then there was this bad thing that the maintainer of the infrastructure left F them with between this years and so Richie search for someone to um who was dumb enough to do that yeah that's me so this year we're back again in Persona sorry yeah thanks [Applause] yeah so after 2 years we came looking for the two machines after almost two years and like no one touched them they rebooted one or two times due to power outages in the server cabinet um but we had a working SSH key um we had tons of updates to install after literally 3 years um I wonder nobody broke into that maches because they were public Exposed on the internet um but only SSH and uh I think a three-year-old or three and a half year old promethus installation which was full of bucks um yeah um we noticed that the battery controllers uh the battery packs of the rate controllers have been depleted so this was the only thing that actually happened in the three years the batteries went went um to zero and didn't set themselves on fire so everything was okay the machines worked um just a bit of performance degradation but everything seemed to be okay and then we tried to run this anible thing from the last years and you know three years later um anible has done a lot of things in the time and you want to use a current version of anible with that old stuff um you end up like this this is me yeah start from scratch or fix all the anable roads like there you can have a look at them they're also on GitHub um so when we we thought okay how do we do this and said okay don't just be gone answer will be gone we just fix fix it um after the F them because um we will have to renew the servers anyway and everything will change so so the servers timeline we have them servers live at the 8th of January Services dns64 all the way the mid of January we had C ized all our locks this was something Richie was looking for since ages that we had easy accessible lock files for everything that's running here at F them um which was good that we had them because we um could see things like oh the uh the internet line that was proposed to be there actually came we did nobody told us but it came up you see that thanks to the centralized logging we we were aware of things like that and then we could um go and fire up our um our bgp sessions then two days later we noticed okay firing up the bgp sessions wasn't that a good idea because we lost almost all connectivity stop it says but I don't care yeah I just keep I just keep talking yeah um we lost all connectivity and said okay damn it we went in some sort of panic mode because the reason for looking at the service was that was like this bind security issue that was been uh I read the mail at the the morning of January uh 20 morning of January 28th and said okay we have to fix the bind installations and then you suddenly can't reach your service anymore and say okay are they already hacked or what's going on and doing some back and forth with our centralized logging you see that this is grafana low key um that we that we leverage for that we were kind of like yeah it's been um it's been really nice to debug things like that we also noticed that there was a interface constantly flapping um to our backbone which we also could um fix within that session and after that we um said okay there's some MTU problems we um have so have to restart bgp and so on and back and forth and then we finally agreed to just throw away um the BCP bgp sessions um go with the 1 gbit line and yesterday evening we switched to the 10 GB line because we had the congested Uplink like since 11: in the morning um so many people using so too much bandwidth um and since yesterday evening everything is okay it's better and we are on the 10 gbit link due to the fact that they not so many people here today yesterday there were quite a bit more um the link was not fully saturated but you can you can tell we this is the place where we could use some more bandwidth was like I don't know this is usually time for something to to eat but at 3:30 we could actually use something of the new bandwidth that we had available so if you want to look at all of the things we have a dashboard put out there uh publicly if you want to have a look at the infrastructure and an anible redo that will be fixed to work with current anible versions within the next few days just clone our infrastructure clone everything and if you have any questions um I'll be glad to take them um yeah fire [Applause] away as as I don't see any questions then we we are about to tear down this room after this so please don't leave anything in here because it will be cleaned and everything will be torn out if if anyone else has a question just there's one there's we use lab the question is uh why do you use laptops for rending because they have a built-in USB called battery so in place of the power outage we can easily run with them um also they're very cheap for us um we can just use the computing power and sell it at the same price that we bought it to the people here you get a cheap laptop we get a some Computing time on them before and um that's the main reason for running it on laptops well actually the question was why you were using banana pie um that's a good question um the thing is that the capabilities of the banana pie were um a bit better than Raspberry Pi the times the decision was made um if you see there's a big um there a big LCD screen in front of uh the boxes where you can see that thing I think it was with driving those uh LCD panels um and also the computing power available on the banana pie that wasn't yeah but actually we have to look that up in the in the the ripo there's everything documented okay yeah yeah there's another one in the front so the question was is there if there are any public dashboards out there yeah we' put some public dashboards on dashboard. graana.com sorry uh which you can have a look at the infrastructure we used to have some more dashboards like uh the t-shirts that have been sold but due to the fact that we changed the shop um we converted to uh something that we bought to um an open source solution and um the thing is we totally forgot to monitor that so that's but there are some dashboards out there to monitor it and if you want to have some to see something more just come come to me after the talk and we I'll show you something more here laptop okay yeah another one what was the biggest issue that you faced the the biggest one standing [Laughter] here now actually the Bigg the big the biggest issues we had was um like um running all that stuff after 3 years and and not having set up everything everything properly was quite um challenging like on Saturday morning we had to run and redo the whole um video installation on on the K building because of um you see those transmitters here they were not pluged properly and so we had no audio on the stream this was one thing and then another very challenging thing was like um when we played around and as I play we did not engineer anything properly um when we played around with the bgp sessions it was not clear how long it would take till things um distributed to the whole net um and we were literally just trying to get information is it working is it working not and till this bgp information propagate from here to the rest of the planet like Brazil um it takes quite some time and so you can't be sure that you're setting up bgp session everything works because um will hit the fan after 10 20 30 minutes and not instantly and so um it's quite um it's quite a problem to have instant recognition if things are going well or not so the question was um if the problems with the Wi-Fi that we had here on on site were due to the rgp playing or was it due to something something something else um solar flares or so um thing is that we had some issues um we we've been given access to the wlc the wireless controllers you see these um boxes over there they're centrally controlled and we have to dig in that um we have some visibility of the infrastructure that's owned by the ulb they've given us access to that so we can engineer that but we're we're not quite sure why was that mostly most of the time fosdem which is an ipv 6 only was working quite good except for some Apple devices that do tend to just um set up an ipv4 address even if there is no proper ipv4 and things get complicated then fom dual stack which is dual stack um usually worked for most of the Apple devices um but we we're not very certain with yeah yeah you will see that there's another one so the question is if the live stream live streams will be made um yeah this rewindable or or not I honestly I can't tell you that I don't know I can ask the video guys if they're planning that for next year um but there's no plan of that as far as I know the biggest challenge was to to redo things with HDMI over vgi which we had the last years this but there's another one yeah's next so the question is that we're planning to use servers do we know what and what's planned for next year um we'll have a talk about that next week I think and then we go through the postmortem which is usually week after Frost them and then we decide on things to to be bought for next year because switches are old and um router is always all also old I think and just we with one more year on the route to go that should be fine for next year but what what after that we have to make some decisions and some Investments for next year to run this stuff and this will be done next week when we're all bit cooled down and refreshed after this for them anyone else yeah come what part of the existing infra are you reusing and what what else bring so the question was what uh in uh what are we what part of the infrastructure are being reused and what do we bring for the event well in numbers it's I think it was three truck loads of stuff no three because the the video arrived the second yeah um uh we bring mainly cameras and those boxes here um switches stay at the ulb most of them stay here but but um that the one that didn't stay here um they won't be here next year because U be is planning to do some some tidying up and giving here um some some video ports for our vlans they're very very um good at working with us um we get access to most of the infrastructure they we just say tell them what we learn we would you like to use and they just throw it on their uh controllers and Bridge it to our service and we can use it and make fun with it and they will be replacing part of the um the network infrastructure next year and we then will have to bring even last gear here yeah which one first yeah what about the rest of the so um the question was uh what's about all the other stuff that fosm is doing through the year um do we host it on our own Hardware is it any cloud or somewhere we used yeah we have another company called ton here it's a Belgium provider and there's most of the stuff is running at ton during the year during fost time we also spin up some VMS um at hetner um in Germany and they they are only for during the event and short time after the event so like cutting videos and so on in in the cloud and they will be turned off like two or three weeks and then everything is running on tun on our own Hardware there as well so there was another question what are you using to you communicate so the question was what is being used for the communication between volunteers um we have that Matrix setup um I don't know who's aware of matrix it's a real-time communication tool like um yeah like slack or something like that um we use Matrix since 2020 internal for our video team um for for communicating and then we expanded that for uh 20 20 and then when the pandemic we open it up for all of the people and now the volunteers are being coordinated to that and we also have our own uh drunk terrial um that we have here um especially for this event setup and the volunteers are also um Can can be reached via those radios am I correct volunteers yes yes okay we have two volunteers here so yeah get to them yeah is there anything else you want to know or any where's the money the the the question is where's the money leowski that's the real phrase from the film um I don't actually know I'm not yet the member of f them stuff so you have to ask someone in a yellow shirt there happens to be one next to me to just throw in the microphone we have a money box and a bank account done done anyone else 3 to one thank you very much thank you very much for