Submind YouTube summaries
Thumbnail for LF Live Webinar: On-Prem and Cloud: Closing the Correlation Gap

LF Live Webinar: On-Prem and Cloud: Closing the Correlation Gap

Watch on YouTube

Video summary

The webinar addresses the critical challenge of managing hybrid IT environments that combine on-premises infrastructure with cloud services, containers, and Kubernetes. The main argument presented is that while collecting data from these diverse sources is manageable, the true difficulty lies in correlating signals across them to solve problems efficiently. A survey cited in the presentation reveals that 94% of teams use three or more monitoring tools, leading to fragmented visibility where issues often bounce between different departments like application, network, and compute teams. This fragmentation relies on outdated tribal knowledge and results in slow incident response times, with statistics showing that nearly nine out of ten organizations cannot resolve cross-environment incidents within an hour, causing significant financial losses due to downtime. To solve this correlation gap, the presentation introduces Data Dog as a unified platform that provides a "single pane of glass" for all infrastructure layers. By integrating metrics, logs, traces, and events into one place, the tool eliminates the need to switch between applications or login credentials, ensuring that context is never lost during troubleshooting. The speaker highlights the use of CloudCraft diagrams to visualize the entire stack and its dependencies, allowing engineers to see exactly how components connect before and after a failure. Furthermore, the platform leverages AI-powered tools called "Bits" to automatically investigate alerts, explore multiple hypotheses for root causes, and even generate code patches or Terraform plans for remediation, significantly accelerating the time from alert to resolution. The effectiveness of this unified approach is demonstrated through a live demo and a case study involving Porsche Information Technology. In the demo, an engineer investigates a high system load alert by drilling down from a host level to specific Kubernetes pods and services, using AI to identify that CPU overload was degrading a product recommendation service. The case study illustrates how Porsche consolidated ten different observability tools into a single Data Dog platform, which allowed their teams to achieve 99.5% availability and faster response times across their global retail locations. This consolidation not only reduced costs by eliminating redundant tooling but also empowered teams to innovate faster without the silos that previously hindered collaboration and customer experience. In conclusion, the webinar emphasizes that modern observability requires a strategy that unifies data from both on-premises and cloud environments to prevent critical failures in complex systems. The presentation addresses common concerns such as log management costs by explaining Data Dog's capabilities for pre-filtering logs to ensure users only pay for relevant data, and it tackles issues like misconfigured devices or incorrect timestamps by showing how a unified view makes it easy to spot anomalies immediately. Ultimately, the solution offered is a comprehensive platform that brings all teams together with consistent data, removing the friction of handoffs and enabling organizations to maintain high availability in an increasingly complex hybrid landscape.
Read the full video transcript
Great. Thanks for the uh invitation and the introduction. Hi everybody. I'm Jayce Harker. I'm a product manager here at Data Dog. And today we're going to talk about on-prem and cloud helping you solve uh issues by uh closing the correlation gap. So correlating signals across all your different infrastructure and making it faster and easier to solve problems. Um so like I said, my name is Jace. I'm a product manager here at Data Dog. I've been data dug for a bit over three years now and yeah I'm excited to talk with you today. So let's jump right in and get started. Uh what we're going to talk about here is first of all why is it challenging to um look at uh hybrid visibility today and especially in hybrid environments right um and where the big gap is not so much about collecting data as it is about how do we correlate all this data. Um we're going to talk a little about where this how this causes problems for organizations. Uh, and I'm sure that some of you have experienced these problems. So, we'll talk a little bit about that. Um, and then we'll show how Data Dog helps solve this problem by bringing together all this data in one place, giving you a single pane of glass. I'll go through a quick demo to show you how this works in action. Uh, and then we'll take we'll sum up and take a little time for Q&A. All right. So, let's jump right in. Before we get into the uh the presentation though, I do want to ask a question. Um, how many tools does your team typically use when you have to troubleshoot a single incident? Um, so some folks, you know, oh, we've already got just one tool. It does everything for us. Other teams do three or four or more. So, uh, uh, yeah, let's go ahead and vote and tell us, uh, how many tools does your team use? Give just a few more seconds for the poll to come in. All right. So uh what we found is uh that 94% of teams run three or more monitoring tools. Um so data dog did a survey of over 105 um senior and executives at different companies [clears throat] uh and talked about you know how are you managing your hybrid environment and uh 94% of teams are using three or more three or more monitoring tools uh to try to troubleshoot all their incidents. Um so this really points to uh a a hidden challenge here which is um as we've moved to the cloud we really increased complexity because on-prem parts of environments didn't go away right um most organizations today didn't completely get rid of their on-rem uh in fact in some cases onrem is actually growing so um in this recent survey we found that 72% % of respondents said their organization was either considering or actually expanding their on-prem workload. Um, and in addition, if you're working in the cloud and you're using containers and Kubernetes, this adds many more layers of complexity. Uh, so now you have on-prem, you have cloud, you have containers, you have Kubernetes. Um, a lot of legacy observability tools um were built before a world of this complexity existed and are not designed to tie all of these different signals together to help you troubleshoot. Um, this isn't just a transition from on-rem to the cloud. This is now the current operating model, right? This is the current state of affairs. Um, so it's not slowing down. It's only going to increase. Um the these bigger, more complicated workloads mean more moving parts, more things you can miss. Um without a single pane of glass, there's not a quick way to get to a fast answer. And yes, when stuff goes down, um customers impact it. I don't have a poll for this, but pretend you're raising your hand. Uh how many of you are now working with GPUs or uh LLMs, um and adding LLM based features into your products? Um a lot of folks are. That's yet another layer of complexity that's coming in. Um, and with all of these different layers and pieces of the puzzle, if any one of them fails, um, then this translates into customer pain, lost sales, lost revenue, um, other problems. So, it it really is a challenging environment to be working in today. Um, why is it challenging? It's not about just collecting the data, right? There's lots of tools that can collect metrics or logs or traces or other events. Um the challenge is bringing all this information together in one place and being able to correlate across all of these different signals so that you can see what's going on. Network traffic is another one. Um when you're troubleshooting an incident, often one team has to do investigate their thing, hand off to another team who investigates their thing, hand off to another team. Folks may not know how all the pieces connect. um you rely on tribal knowledge which if somebody leaves the team or leaves the organization suddenly that information is lost because potentially documentation is far out of date. Um so this leads to ultimately slower incident response um takes longer to uh hand things off to different teams um hire to communicate and at the end of the day it hits the bottom line right um this is the kind of thing that we're trying to fix. So this is like a very standard. Take a look at this and tell me if this sounds familiar to you, right? An alert fires at minute zero. A SE 2 is declared. Um people start looking into it uh on one team. Um they start investigating. They don't find a clear answer. Maybe the application team starts looking into it. They don't find an issue there. They escalate to the networking team. Networking team escalates to the compute team. Right? and this bounces around and over an hour later um eventually we get to the root cause right um this is painful it's timeconuming it's ultimately it's costly um this is what happens in real life in this survey we found that 87% of people said that their organization could not resolve a cross environment incident in under an hour so the 65 minutes uh this is real this is what a lot of folks are dealing with how much does it cost Uh it can cost a lot. Um ITIC uh recently did research and found that mid to large-siz corporations um an hour of downtime can cost $300,000 if you're a Fortune 500 company and cost half a million to a million dollars per hour. Um or potentially even more in certain industries like finance and healthcare. Sometimes it can be $5 million per hour of downtime. Um so the point is that every time every minute that you spend trying to track down the source of the issue um is adding to that clock time, right? So it's it's and if you can speed that up, it's going to directly reduce the cost to the business. So let's do another poll. Um how much of this resonates with you? If you're have a production alert go off, um where do you spend most of your time before your team starts to resolve it? Is it figuring out who owns the problem? Uh, is it trying to correlate data across the multiple tools that you use? Um, are you waiting for that one person who knows what's going on to get looped into the the Zoom call? Um, or do you really have all this figured out already, which is great. If you do, that's fantastic. Uh, we'll give it a minute or two more for a couple or a couple more seconds for folks to uh submit their uh submit their uh the polls. All right. So, let's see. But what do the numbers say about this? So, what we found is um almost everybody says one of those things, right? It's a very small percentage of you that said, "Oh, we got this dialed in." [laughter] So, um and that's really what Data Dog is here to help solve. So for everyone seeing data dog for the first time, this is what data dog looks like at a high level. The point of data dog is really that we bring all this information together into a single pane of glass. We bring all your teams together into a into a single place so that you have clear communication, clear continuity. You can connect the dots between all this information um and use this to troubleshoot every single type that we've talked about. metrics, events, logs, traces, they all flow into this single platform. Network, infrastructure, application monitoring, all the layers, they're all connected together, right? There's one data dog agent, one platform, this single unified experience across every team. Uh, and here's why that matters when something, you know, actually breaks, right? Um, it means that these signals can be correlated together automatically. You can see what logs correlate with what trace. You can see what metrics correlate with which logs. There's no switching between tools, no switching between signals, no switching between teams. Everybody sees the same information in the same place. Nothing is lost in translation. And now with Data Dog's AI tools, with Bits, which I'm going to show you a little bit later, um we can actually help you in investigating and figuring out the root cause even faster. So what does this look like in practice? I'm going to give you a demo in a minute, but at a high level, you can see here, this is the CloudCraft diagram in Data Dog. And what this shows is the full picture. This is what your whole stack looks like mapped out in one diagram. It's a visual diagram, so you can clearly see what's going on before something breaks. How do things connect to each other? What are the dependencies at the infrastructure and at the software level? You can't fix ultimately what you can't see. If you have out-of-date diagrams, this will give you clear understanding of how infrastructure works. Even if it's not the infrastructure that you're normally responsible for. Now, let's see what happens when something goes wrong. You can now in data dog follow the incident across different areas. Whether you're looking at the software layer, uh whether you're looking at logs, traces, like I mentioned, all of this, there's no tool switching, no new login. You're not losing any context. You have a consistent experience all the way, right? This is what we wish the on call engineer had right at t equals 0 when the incident is called in the first place. And now with BITS, we actually use AI to connect the dots for you, saving you even more time. So rather than Yes, you can use all of these tools to investigate, to troubleshoot, to confirm. Uh, but Bits will do a lot of the job for you right out of the gate and give you a head start speeding up your remediation, finding what the root cause looked like. So let's what does this look like in practice? Let's jump right in and take a look at what this looks like in the data dog environment. So, let's say that I'm on call. I got paged for an alert. In this case, it's relatively simple looking alert. It says, "Hey, system load is high on this host." Um, I can go down. I can see more information about this. I'm going to see my diagram right here that shows me, hey, here's my host and here's the infrastructure around it. Um, now I'm going to click and open up this diagram in full just so I can explore in a little more detail. So, here's that host again. Now I've clicked on this host and I'm seeing all of the telemetry related to this host in one place. I can see the key things that are going wrong. In this case the system load is high. Uh metrics of this host. I can look at if it's running Kubernetes which in this case it is. I can see all the pods that are running on it. Logs, traces, network traffic. If there's security issues I can see that. I bring all this information together in one place. It's like a single pane of glass. Furthermore, looking at this diagram, I can see that this host is actually running this. It's one of the nodes for this Kubernetes cluster. And in this Kubernetes cluster over here, I can see that there's three services, orders, app, web store, and product recommendation that all have errors. So, I'm going to hover over this and see exactly what uh the product recommendation status is. And it looks like the health is critical. So, I'm going to click onto the service page for that. And now I'm looking at the service product recommendation and I can see the health is critical. I can see what's going on and I can go down the page and look at the specific traces and uh logs and other information about this service again all in one place. If there's database queries I can see that and I can really use this to drill drill uh drill down quickly and find the root cause of the issue. But while I'm doing that, I may want to simply kick off an investigation. So I'm going to go here and I'm just going to say, "Hey, please investigate this with Bits." I'm going to click here and Bits is going to run an investigation for me. So what is Bits going to do? Bits is going to start with this monitor alert, explore multiple hypotheses of what could be going wrong, look at the different telemetry, all of this correlated telemetry and metrics to correlate things directly together and ultimately verify that some of these hypotheses seem to be correct. And what this is going to show me is it's going to show me uh that in this case CPU overload has degraded the product recommendation service which in turn had kind of this knock-on effect. And the recommendation here is to ultimately uh scale up my Kubernetes cluster. Um and it's even giving me a PR um so that I can deploy this via Terraform um and uh and make these changes. So this is a really great way to see this. Now I mentioned earlier correlation across cloud and uh hybrid environments. I'll just throw back here a little bit. Um this of course is showing me the cloud part of my environment but I can also go here and just as easily see my onrem in this case this is a a VMware cluster. I can see my cluster. I can see alerts on different things that are going off and if the issue is running onrem everything that I just showed also exists for onrem. So again, here's my list and I'm going to look at this resource and see all of this information together in one place. Works just the same for on-rem as it does for the cloud. So I can clearly correlate all this information together and troubleshoot and ultimately get to the root cause. So I can of course go in and use all this all this great tools to verify that bits did find the right solution and then I can go ahead and implement it. All right. So now that we've seen how this works in practice in data dog um you've seen how data dog correlates these signals together. We've shown um how uh we can immediately kind of start looking at uh things going wrong out of the box. We've shown how bits can detect and investigate and figure out the root cause for you automatically. And we've shown how this ultimately gives you one pane of glass that your all of your teams can rely on that has the same data for everybody that gives a single view so that we don't have different teams in different silos pointing fingers at each other. So which of these would be most useful for you? All right, great. So we're getting some responses in the poll here. So I think the answer for most people is more than one of these things. Um and there's just an example. Um Porsche Informatic. Um of course Porsche is a very large company. They do business in 29 countries. Um they have over 500 retail locations and they had a lot of challenges. Um before they started working with Data Dog, they had 10 different observability tools. Um and these 10 different monitoring tools were really creating a lot of complexity, slowing their team's ability to innovate, uh impacting the quality of their service, driving up their costs. Um once they started using data dog, they were able to consolidate all of this into the single data dog platform. Now all of their teams had complete visibility into every part of their environment all in one place. And this unified view really meant that uh they could detect issues faster, troubleshoot issues faster um and ultimately respond faster to all the dealerships they have in the field. So the law so the the bottom line impact here was they were able to achieve 99.5% availability, improve the root cause analysis, uh reduce their time to insights and also optimize their costs for their infrastructure spend. So Porsche's leadership really credited this consolidation with empowering their teams to move faster, improve the customer experience, innovate at scale in a way that they couldn't do before when their observability uh tools were fragmented. So this was a really uh great success story and it really illustrates the benefits of data dog. Right? before you have tools that are disconnected from each other, teams that are disconnected from each other. These different data is all in different silos, hard to correlate, it's hard to hand off. Um, often you don't know who's going to know what. Um, and at the end of the day, this means slower time to resolution. Um, with Data Dog, you get everything in one place. unified visibility, faster time to resolution, correlating your signals across your environment, makes it much easier to hand things off between teams, quicker to get to a root cause. Um, and now with Bits, Bits can actually do a lot of that work for you right up front and speed up even more of the time to the root cause because rather than having to dig through your telemetry to figure it out yourself, now Bits gives you the pointer to look in the right direction and all you have to do is validate it and then implement the fix. So that's uh that's the uh the overview of some of the great benefits of data dog. I know there's actually a lot of other great features in data dog that I didn't have time to cover, but I think that gives an overview of what some of the great benefits are and how data dog can really help you working in your hybrid environments. Um so with that, uh we'll take a pause here and open the floor to questions. Uh, let's see here. I see one question already. So, I'm going to go ahead and uh and answer that question first and then we'll uh and then we'll uh see if we get more questions as we go along. But please, as I'm answering the question, please do uh come on in and uh and uh put more questions in the chat and I'm happy to uh answer more of them. So this question says uh I spent most of my time identifying the misconfiguration of infrastructure devices, routers, servers, applications that skewed the log data by the wrong date or time or was not sending logs to all the right places or the application was misidentified with the type of data it was supposed to send. [laughter] How does data dog fix those issues? [snorts] Um that's uh a great question. Um, I think at the end of the day, data dog fixes those issues by helping you, like I said, see all the information in one place. Uh, I'm going to use um, CloudCraft as an example here just to show you kind of what this might look like. But for example, um, in this Cloudcraft diagram, I can zoom in here and say, "Show me logs." And for some of these resource types, I can actually see, for example, these lambdas, I can see which of them are reporting logs and click on a particular lambda to see just the logs for that lambda. So if you're trying to troubleshoot for a specific resource, whether it's configured properly and whether the logs are shown up correctly, you can see that here immediately, right? Um and then similarly, if we go into let's say our um our host list here, I can actually see a list of all of the hosts that I'm monitoring. Um and I'm going to say be able to look at each of these hosts and see uh so for example, this is a um I just opened this up, but this is an Ubuntu host. Um, and uh, I can see that it's running the agent. I can go in here and to get that same unified view, I can search my hosts to find exactly the host or the subset of hosts that I want and manage the whole view to see all this information in one place. So, I can quickly say uh, I probably picked a bad example here because this host doesn't have any logs or traces at the moment, but if I had logs and traces, I could just jump between them here and see exactly what's going on. Um, so the the advantage of having all this information in one place is that you can quickly kind of crossorrelate this data um to see what's going on and help you troubleshoot. Yeah. Okay. Looks like this log has the wrong time stamp because I know it was just generated, but it says it was generated 4 hours ago or something like that. Let's see. Um, any other questions from the group? Please feel free to throw more questions in the chat here. Let's see. Looking in the chat, I don't see a lot of other questions. Give it one more minute. And if anyone has any more questions you'd like to ask, feel free to throw it in the chat. Oh, I see one more question come in. Let's take a look. Great question. So, um, this question says, "Large corporations may have an issue with the cost of sending all the log data to a single device as many times a lot of that data is unnecessary. How do you determine the specific data to send?" Um, so I'll say I'm I'm not a super expert on logs within Data Dog, but I do know that Data Dog has a lot of capabilities to pre-filter logs before they get sent to Data Dog. Um, and there's a number of different tools you can also use to uh view and manage your log uh spend and make sure that you're only sending the important logs um or in some cases like subfilter logs and only show like the the types of logs that you think are the most important. Um, so there's definitely a lot of capabilities within data dog to uh make sure that you are getting good value for the the logs that you're spending and you're not spending money on logged volume that's not important or not useful. Let's see. We'll give it one more minute for other questions. Also, if you have questions about any other capabilities of Data Dog that you've heard about, you'd like to learn more about, feel free to ask. Or if you have any particular issues that are relevant for you and you'd like to ask about it and uh and see how data can help you, feel free to throw that in the chat and uh happy to answer questions about specific use cases or problems that you're trying to solve. All right. Think we're not getting any more questions. So, I think we will uh we will uh end the webinar here. Thanks everyone for your attention and for taking the time today. And yeah, if you can scan this QR code below, you'll get a 14-day free trial uh for Data Dog. Or if you just go to the Data Dog website and sign up, you can get a 14-day free trial. Try it out. No credit card is required to do the free trial. So, you can just jump right in and try things out and see how it works for you. Um and uh yeah, we hope you try out Data Dog and start using it. We think you'll get a lot of benefit from it because it's going to solve a lot of uh problems that you might be having with your current observability tools um and a lot of challenges that we have in today's kind of modern uh complex hybrid environments. So, thanks again once uh once again all of you for your time and I'm going to hand the ball back over to Linux Foundation. >> Thank you so much Jace for your time today and thank you everyone for joining us. As a reminder, this recording will be on the Linux Foundation's YouTube page later today. We hope you join us for future webinars. Have a wonderful day.