Submind YouTube summaries
Thumbnail for NorthSec 2025 - Joey D - Noise Pollution is Damaging Your SOC

NorthSec 2025 - Joey D - Noise Pollution is Damaging Your SOC

Watch on YouTube

Video summary

Joey D, a detection engineering team lead at the Canadian Center for Cyber Security, addresses the critical issue of noise pollution within modern Security Operations Centers (SOCs). Drawing on his extensive background in incident response and Capture The Flag competitions, he highlights how the sheer volume of data generated by federal institutions—reaching up to 200,000 host events per second—creates a "tsunami" that leads to alert fatigue. He argues that relying solely on Indicators of Compromise (IoCs) is insufficient because legitimate system behaviors often mimic malicious activity, causing analysts to waste time triaging false positives. To combat this, he advocates for a mindset shift from detecting what is malicious to detecting what is legitimate, emphasizing the need for high-quality telemetry and rich context to prevent analysts from missing genuine threats due to exhaustion. A primary example of this noise pollution is Microsoft's Delivery Optimization feature, which uses peer-to-peer technology to speed up Windows updates by sharing files across a network. By default, this feature listens on port 7680 and can share update chunks with other devices, including those outside the local network if configured to do so. Joey details the complex API calls and DNS resolution chains involved in this process, noting that standard Endpoint Detection and Response (EDR) tools often misinterpret these legitimate sequences as suspicious network scans or data exfiltration attempts. Without a deep understanding of the specific phases of Delivery Optimization, including its use of proprietary protocols like Swarm and specific hash verification steps, security teams generate thousands of unnecessary alerts for every single update cycle across their fleets. To solve this problem, Joey proposes building a robust knowledge base that provides analysts with immediate context to triage alerts effectively within seconds. He introduces the concept of a "20-20 definition," where an analyst should be able to identify behavior in 20 seconds, understand it in 2 minutes, and recognize potential abuse scenarios in 20 minutes. By implementing custom filters that fuse EDR data with network telemetry, organizations can automatically tag alerts as benign Delivery Optimization activity rather than malicious traffic. This approach not only reduces the noise reaching human analysts but also restores trust in IOC-based detection by ensuring that resources are focused on actual threats rather than filtering out normal operating system functions. Ultimately, the goal is to reduce manual triage burdens and allow security teams to focus their efforts on identifying true anomalies and adversary tactics.
Read the full video transcript
Feels great. Finally, I'm participating. I can still remember being an IT engineering student in the crowd. Just robes. It's about to give me a seizure, but it's fine. And that was just a confirmation that cyber security was the career choice I wanted to make. So anyway, I have absolutely no segue from there. So let's talk about pollution. Um, my name is Joey. I'm the team lead of one of the many detection engineering team at the Canadian Center for Cyber Security. Previous to that, I filled various blue team roles uh where my last one was in the CTI team working on APS. I promise that this is the last AI generated content in my slide, but it's not going to be the last cat. In my past time, I happen to be a CTF enthusiast and I've been a challenge designer for 5 years now and I'm also involved with cyersai as the coach of team Canada and CTF. Uh Canada has worldclass cyber security talent and can be proud of winning gold medals uh for the three past European cyber security uh challenge competition. So that's quite an achievement here. I'm also involved in most high-profile incident response cases that the cyber center helped with. And uh if you watch TV shows, you might have also seen me on LEGO laborator on Radio Canada. If not, um here's a sample here. Now, the Canadian Center for Cyber Security is part of the Communication Security Establishment, CSC, and is the federal government's operational and technical lead for cyber security. Our mandate is to offer cyber security services to various client across the country ranging from federal institutions to system of importance uh the electronic systems of Latvia and Ukraine and other systems designated as system of importance by the minister of national defense in the critical infrastructure. We help all the levels of government. So whether you're municipal, provincial, uh territorial and even ind indigenous, we help the energy sector, the finance sector, the telecommunication sector, and a lot more. So we had elections not too long ago. We help with that. The cyber defense that I'm branch that I'm part of offers a suite of tools and services and sensors to federal institutions. The portfolio of sensor and log collector let us cover most of the IT footprint of the government of Canada. So we have various log aggregator, our very own hostbased sensor, our very own cloud-based sensors and a bunch of network sensors. Now collecting data at that scale is already quite the endeavor. And here comes my problem and why I want to talk about pollution and noise pollution. With over 167 clients, almost a million devices protected, 200,000 and more host event per second. That's a tsunami of data that has to be analyzed. Just to give an example, our manual and automated analysis is actioning about three billion network blocks every day to protect our clients. Now my team is part of the security operations center and let me tell you that noise pollution causes a lot of pain and it's resulting in alert fatigue but to be honest that noise pollution is caused by pain. Let me explain here. You might have heard of the infamous pyramid of pain. It's a great attempt at presenting the inside of your thread thread intelligence and what it can provide. The more you go up the pyramid, the more challenging it is for an actor to modify their TTPs. Now, which would make technically make your detection future proof. On the blue team side, you're also a victim of the same pyramid where the bottom levels are often ephemeral and noisy. But in a mature suck, defense in depth is key. And one can't ignore the lower part of that pyramid, even if it's painful to triage. This here is what you can't afford to miss. You can't miss an incident on that's using known malicious infrastructure, especially when you're paying hundred of thousands of dollars in various feeds. Your board will simply not like it uh that you're wasting the investment. But I'm here today to suggest ways to give glory back to the indicators of compromise. Let's avoid IoC's from turning into indication of cacophony here. Of the many reasons why IOC based detection can be noisy, there is deceptive data. You just don't have the right data, capturing net flow data will not help you uh detect the web encrypted web shell traffic that you have on your server. What you need is either full pcaper with TLS offloading or you need host data. ignoring the intelligence part of CTI context matter here that new fake uh capture domain that you queried well if it's not starting from the explorer process where a user did Windows R pasted the command and pressed enter you're likely looking at a false positive here an imbalance of human activity there's always more room for automation but then at the same time don't let automation hallucinate the response there's many solutions here like I said to reduce the noise in your detection pipeline but it ultimately boils down to three key points. You want to increase the quality of the telemetry used for detection. You want to increase the context presented to analysts and you want to decrease the quantity of manual alert that has to be triaged or you just risk alert fatigue and ultimately false negative assessments. To solve the problem of noise detection, detection engineer have to change their mindsets from detecting what is malicious to detecting what is legitimate. Exploring the noise one feature at a time is critical to lowering the noise itself. The solution here, building a 2220 knowledge base to provide enough context for triage analyst to be able to triage their alerts. So analyst should be able in 20 seconds to identify the behavior they're looking at. In 2 minutes they should be able to understand it and in 20 minutes they should be able to recognize how an actor would abuse that behavior. Here let's just demonstrate that here. It's a nice a Friday afternoon right now. What afternoon? It's 400 p.m. The sun is shining. It's a nice 30° in Montreal and you can smell that the weekend's coming here. You get a notification, new critical alert to triage in your alerting platform. Oh no. The local IP of your client is uploading 2 GB of data to an IP tag by your CTI team as Volt Typhoon and it's doing it over port 7680. That's it. That's your alert. What do you do? Do you triage uh by discarding it or do you escalate it? Now assessing the alert without additional information is just wrong. In that case here artificial intelligence was even as hallucinating and incorrect explanation and we didn't have any knowledge base that was rich enough to provide context and as a detection engineer I should have packaged a 20220 definition allowing an analyst to triage uh by the alert by identifying understanding and finding abuse of this behavior here. So let's just build one together. Time to build one so that we never have to do it in the future. Let's start with the hypothesis that any traffic with port 7680 is indicative of delivery optimization. We're going to deep dive into it in 22 minutes and 20 seconds. Uh but please don't die me. Alice should be able to identify delivery optimization in in about 20 seconds. So here's what you need to know. It is a Windows 10 and 11 feature that is cloud managed by Microsoft and it features a peer-to-peer client and downloader service normally over port 7680. It is not malicious at all and the content that you're downloading via deliver optimization is normally Windows updates, Windows Store apps and Windows Defender definition updates. uh the download is orchestrated and downloaded via Microsoft servers network peers and the Microsoft connected cache and that's it 7680 is all you need to remember on how to identify this but let's get a little more technical and understand how it works delivery optimization is a service so you'll find information about it in the registry it runs under it is started by SVC host under the network service user the the DL that it uses is all in system 32. So do the OSBC.dll. And that's pretty much it for it. Like I said already, it is on by default on Windows 10 and 11. And the expected result of using delivery optimization. It's supposed to uh have a result of having reduced bandwidth usage and a faster update process. Now let run you through one of the example. Take a classic patch Tuesdays. you have to download a 700 megabyte new patch. It's great. Now, looking at the entire fleet of Windows machines that Microsoft has of 1.4 billion devices worldwide, that's quite a lot of terabyte to download on a Tuesday. Let's say that way. And if we tune it down to your deployment or our deployment of about 1 million devices, it's still 700 terabytes of data that has to cross our various sensors. And it's causing a lot of pollution here. Now the way that deliver optimization works is that it's going to take the full update, split it into equally sized chunks. Some of the chunks are going to be downloaded via Microsoft servers. Some of the chunks are going to be downloaded via peer-to-peer and then some of the chunks are going to be downloaded via Microsoft connected cache. And then deliver optimization checks that it's the right file that it downloaded and you're done. That's that's how it works. But if you truly want to understand how adversaries are cloaking their activities as a legitimate feature of the OS, you need to master how it can be abused. And that's a good thing that we have 20 minutes left for this. The optimization works in three phases. The first one is that the service needs to receive a task from other processes. Then you're going to download the metadata, so the information about the file you're about to download. And the file itself is the last step for the task. Deliver optimization uses RPC calls uh and as its interprocess communication. Now what it means is that com object are exposed for any services to call. It functions similarly to bits if you're familiar with this. the task received. Uh the services that are normally tasking it are the Windows update service of course the Microsoft ser uh store one and you can also download a PowerShell script on GitHub that lets you play with it. To give an example, I'm thinking here the Windows update delivery optimization ecosystems or wudo. On the left you have Windows update and on the right delivery optimization. Well, Windows Update is going to call a bunch of uh com interfaces sequential sequentially to ask for a file to be downloaded. So, it's going to add the file to the job, get transfer information. It's the classic interprocess communication here. Now, the problem with com telemetry is that you can't look at that. It's way too noisy. But there are great forensics artifact on a machine. If you want more information, there's a file called dosvcd state.dat that is high formatted. So you can open it with your registry editor even though it's not in the registry. I don't make these decisions anyway. It's got the jobs. So those are the individual file downloads and it's got the swarms which are all the connections that you've made with peers. Uh so you can investigate these tasks. There's also some built-in PowerShell commands that you can use on your machine to get more information about these tasks as well. And we're done. That's the task. That's how uh delivery optimization knows that it needs to do something. So let's get started with the metadata. Delivery optimization will want to know first where is the closest data center that is physically located to me. Uh for this it's going to call the go API of Microsoft which is going to respond with the public IP address that you have as well as the next step you want to hit. So KV 601 in that example here is the closest data center to me physically. As a blue teamer though, you want to make sure that no malware are abusing the Microsoft API to try to evade your classic detonation public IP resolution. So, it's not going to use what's my IP. It could be using go. It's free. Let's hit KV 601. So, the key value uh API call is going to give you a large list of all the next steps that you need to follow uh with a bunch of config information. But what interests me here is disk 601 and CP 601 which are the closest data center to me again. So let's hit CP 601. This one is the content specific policy. It's going to give me what's important here the hash of hashes. I'm going to go back to it. But think of it as the it's the the file integrity hash of your whole download. And it's going to give you a link to a file called pieces hash. Let's go get this one. The pieces hash file will get will give you the hash of hashes again which is the same one we had previously and it is the sha 256 of the full file you want to download. Your hash of hashes uh is also going to give you the size of each of these chunks for delivery optimization. So in that case 1 megabyte is the default and the pieces are each going to have their own shot 256. So you're going to know for sure that you're downloading the right thing. Even though that's done over HTTP, at least they're comparing the hash between the two services. Now, if your computer is configured to be only using Microsoft servers, it ends there. You're only going to download that from Microsoft servers, and the job's done. But there's also a peer-to-peer system that which I want to talk about. If your device chooses to download the file over peer-to-peer, you need to join the array. that's the closest to you. When you're torrenting Linux ISOs, it's the equivalent of asking a for a tracker for a list of peers to connect to. So, we're going to make a call for disk 601, which is going to give us our closest array, which is exactly what we want to join. So, we're going to hit join on array. What's going to come back from Microsoft is a list of about 250. If you do it again, it's different peers, but it's going to give you a list of peers, which are machines that have the file you want and that are ready to share with you. Uh, you'll get their peer ID, the IP address to reach them on, and the port. If other hosts have joined the array from the same public IP as yourself, their internal IP are returned by the API. Otherwise, you only get their external IP. So think of it as kind of a silent port 7680 scanner if you can reach that. This means that if you can get access to a network, you can technically conduct a network scan without ever touching any of these hosts. And as a a couple of months ago, both the internal and external IPs were returned when you query that API, but that's has since been patched, so shouldn't happen anymore. Now in a if I make a little summary here delivery optimization you can just see that it's working correctly when you see it doing a bunch of queries sequentially to a bunch of different APIs but how does it looks like on your EDR that's the question I asked myself digging into the legitimate features can sometime uncover weird behavior from your from your own security tools and as a mandatory disclaimer here what I'm about to explain has since been improved on our on our own hostbased sensor. So I recommend still taking a look at your own EDR to see if this flaw logic impacts you here. We're expecting the first query to be on Jio. So that's technically the real network query. First one, go.microsoft.com. What our edr is was uh reporting was array 618. Sure. Okay. Let's look at the next one. KV 601 akami edge. I'm not getting what I expect here. CP 601 AMI edge. Okay, Akimi is resolving everything now. DL.d delivery whatever. Oh, we're getting the right one. I'm I'm even more confused here. Disk 601. We're back on Aimi. And the last one, array 601. Now we're hitting array 601. Please help me someone. I don't understand. Actually, I do now. And it's always DNS. When you're making a query for KV 601, you're not getting an A record. You're getting a CNAME which points to a load balancer uh system. And this load balancer edge key is giving you another CNAME for Akami here. And then then AI is giving you an IP address an A record. So what that means is that if a host doing a query for the initial domain would when it's making a query for a KV 601, it's the last domain in the CNAME chain that it's storing in its DNS cache. And because an EDR has no way to inspect the traffic that you have, it has to rely on that DNS cache to map the IP back to the domain. So if what you have in your DNS cache is just wrong and incomplete, that's what the EDR is going to report anyway. Uh and because the initial domain has followed a CNAME chain, you lose the expected mapping that you wanted from the original domain queried. So your IOC detection here will just not work in your EDR. One way to see the behavior in action is to just look at the events sequentially. So you'll see pairs of event, the first one being right and the last one, the second one being wrong. But like I said earlier, that's not impacting our hostbased sensor anymore. But you might want to confirm if your own ADR does that. We're done with metadata. Let's download the file itself. The first thing if your computer chooses to use peer-to-peer delivery optimization is that it's going to start a listener on port 7680. The file is going to be transferred by your network over a uh protocol called swarm. It's a proprietary Microsoft protocol here, not documented anywhere as usual. And what you're going to get uh that's important here in the handshake is that you'll get the shot 256 of the file as well as the peer ID of yourself. And if the other computer at the other end has the same file, it's going to respond back with the same hash as well as its own peer ID. So what's important here? 75 bytes response 75 bytes answer. That means that the file is being shared on delivery optimization. If there was a zero byte response, it means that the other host would not have it at all. And that's how you can see it in your EDR here. That is the only call uh the part of the swarm protocol that's been fuzzed in open source. And delivery optimization runs with sedd bugs uh permission. So please try to fuzz the other one and contribute back if you have time because it's kind of important that this one gets audited. Now if you see in your EDR that more than 75 bytes were transferred more data than that it means that a file was shared over peer-to-peer delivery optimization. By default on Windows 10 and Windows 11 delivery optimization is configured to share files on the local network only. But here's the little neat option. I have a pointer that doesn't work here. Once you enable that option here, it lets you share files with anyone on the internet, even that IP that your CTI team tagged as Volty. Now, as part of each 20 to20 definition, I find it interesting to look at the whole fleet to understand how popular a behavior is. Even if it's not malicious, which it it isn't in that case. Looking at our entire client base, we see that 66% of the hosts are sharing uh have a port 7680 listener open on the host, meaning that delivery optimization is activated. Of these 66%, 54% of them have shared over one megabyte of data over delivery optimization. of these 55 54% there's 20% of them that shared over the internet. You might think that it's only one client but it's half of them. Half of them have at least one host sharing delivery optimization files over the internet. While delivery optimization is not a threat and it's a feature that's supposed to speed updating process, it is a nightmare for sock analysts. We buy thread feed from most manager vendors uh consume most MISP and open source feeds. And all these stats here, they sum up to 7% of our entire fleet going to the internet. That's an impressive amount of hosts making an impressive amount of connections to random IPs, suspicious, malicious, whatever we have. And without the 2020 2220 filter, it would just result in an unsustainable amount of alerts to filter. And I'm done with my 20 to20 filter here. And as a next step, my recommendation would be for you to assess if delivery optimization is configured the way you expect in your own network. I will myself keep focusing on identifying, understanding, and finding sources of abuse of more operating system features. And detection engineers should also focus more on finding and signaturing legitimate behaviors. As blue teamers, we need to keep focusing on reducing the noise that gets to our sock analysts. That's the only way that we can regain trust in IOC uh detection here. If you want to get started, that's that's your slide to take a picture of. If you want to get started, there's a plenty of open source research that's done to to imitate the same recipe here. Excyclicia documents most Windows files and processes in depth and it's also linking to rules on how to abuse it and how to detect abuses. If you like the the Sigma format, it speaks like Sigma. The MITER attack framework, well, it needs no introduction, but its data source section is great at quickly identifying these legitimate behaviors. The my favorite here, living off the living off the land. The name says exactly what it is. uh it lists a ton of projects around various use cases of legitimate and less legitimate OS features and behind each repository on it there's a hundreds of hundreds of techniques that will help you build your knowledge base have a a look at the WTF bins and living off the false positive quite interesting now we're back with this I told you uh that I would delegate that part here uh fortunately we built a 2220 filters to help By doing data fusion with your EDR data and your network data, we're now able to identify that it follows the regular delivery optimization process chain. Here a premputed verification routine can also confirm that it is delivery optimization and it generates a tag on the alert to help us understand what we're seeing. And uh luckily for us, there was even someone that presented at NorthSc 2025 about delivery optimization and the system added a link to the YouTube video of this talk here. Now you're the analyst. How do you triage this? Do you triage it? Uh do you discard it or do you escalate it? And if you guessed that discarding it was the right answer in that case, uh you are correct. the poor IP address was probably used uh by was probably compromised infrastructure in that case that was used as an a relay box but our client happened to share Windows updates with it that always happens on that note uh that's it for me today and I'm wishing you an incident free Friday and I will see you soon with the next 2220 Thank you.