LF Live Webinar: On-Prem and Cloud: Closing the Correlation Gap
Watch on YouTubeVideo summary
The webinar addresses the critical challenge of managing hybrid IT environments that combine on-premises infrastructure with cloud services, containers, and Kubernetes. The main argument presented is that while collecting data from these diverse sources is manageable, the true difficulty lies in correlating signals across them to solve problems efficiently. A survey cited in the presentation reveals that 94% of teams use three or more monitoring tools, leading to fragmented visibility where issues often bounce between different departments like application, network, and compute teams. This fragmentation relies on outdated tribal knowledge and results in slow incident response times, with statistics showing that nearly nine out of ten organizations cannot resolve cross-environment incidents within an hour, causing significant financial losses due to downtime.
To solve this correlation gap, the presentation introduces Data Dog as a unified platform that provides a "single pane of glass" for all infrastructure layers. By integrating metrics, logs, traces, and events into one place, the tool eliminates the need to switch between applications or login credentials, ensuring that context is never lost during troubleshooting. The speaker highlights the use of CloudCraft diagrams to visualize the entire stack and its dependencies, allowing engineers to see exactly how components connect before and after a failure. Furthermore, the platform leverages AI-powered tools called "Bits" to automatically investigate alerts, explore multiple hypotheses for root causes, and even generate code patches or Terraform plans for remediation, significantly accelerating the time from alert to resolution.
The effectiveness of this unified approach is demonstrated through a live demo and a case study involving Porsche Information Technology. In the demo, an engineer investigates a high system load alert by drilling down from a host level to specific Kubernetes pods and services, using AI to identify that CPU overload was degrading a product recommendation service. The case study illustrates how Porsche consolidated ten different observability tools into a single Data Dog platform, which allowed their teams to achieve 99.5% availability and faster response times across their global retail locations. This consolidation not only reduced costs by eliminating redundant tooling but also empowered teams to innovate faster without the silos that previously hindered collaboration and customer experience.
In conclusion, the webinar emphasizes that modern observability requires a strategy that unifies data from both on-premises and cloud environments to prevent critical failures in complex systems. The presentation addresses common concerns such as log management costs by explaining Data Dog's capabilities for pre-filtering logs to ensure users only pay for relevant data, and it tackles issues like misconfigured devices or incorrect timestamps by showing how a unified view makes it easy to spot anomalies immediately. Ultimately, the solution offered is a comprehensive platform that brings all teams together with consistent data, removing the friction of handoffs and enabling organizations to maintain high availability in an increasingly complex hybrid landscape.
Read the full video transcript
Great. Thanks for the uh invitation and
the introduction. Hi everybody. I'm
Jayce Harker. I'm a product manager here
at Data Dog. And today we're going to
talk about on-prem and cloud helping you
solve uh issues by uh closing the
correlation gap. So correlating signals
across all your different infrastructure
and making it faster and easier to solve
problems. Um so like I said, my name is
Jace. I'm a product manager here at Data
Dog. I've been data dug for a bit over
three years now and yeah I'm excited to
talk with you today. So let's jump right
in and get started. Uh what we're going
to talk about here is first of all why
is it challenging to um look at uh
hybrid visibility today and especially
in hybrid environments right um and
where the big gap is not so much about
collecting data as it is about how do we
correlate all this data. Um we're going
to talk a little about where this how
this causes problems for organizations.
Uh, and I'm sure that some of you have
experienced these problems. So, we'll
talk a little bit about that. Um, and
then we'll show how Data Dog helps solve
this problem by bringing together all
this data in one place, giving you a
single pane of glass. I'll go through a
quick demo to show you how this works in
action. Uh, and then we'll take we'll
sum up and take a little time for Q&A.
All right. So, let's jump right in.
Before we get into the uh the
presentation though, I do want to ask a
question. Um, how many tools does your
team typically use when you have to
troubleshoot a single incident? Um, so
some folks, you know, oh, we've already
got just one tool. It does everything
for us. Other teams do three or four or
more. So, uh, uh, yeah, let's go ahead
and vote and tell us, uh, how many tools
does your team use?
Give just a few more seconds for the
poll to come in.
All right.
So uh what we found is uh that 94% of
teams run three or more monitoring
tools. Um so data dog did a survey of
over 105
um senior and executives at different
companies [clears throat]
uh and talked about you know how are you
managing your hybrid environment and uh
94% of teams are using three or more
three or more monitoring tools uh to try
to troubleshoot all their incidents. Um
so this really points to uh a a hidden
challenge here which is um as we've
moved to the cloud we really increased
complexity because on-prem parts of
environments didn't go away right um
most organizations today didn't
completely get rid of their on-rem uh in
fact in some cases onrem is actually
growing so um in this recent survey we
found that 72% % of respondents said
their organization was either
considering or actually expanding their
on-prem workload. Um, and in addition,
if you're working in the cloud and
you're using containers and Kubernetes,
this adds many more layers of
complexity. Uh, so now you have on-prem,
you have cloud, you have containers, you
have Kubernetes. Um, a lot of legacy
observability tools um were built before
a world of this complexity existed and
are not designed to tie all of these
different signals together to help you
troubleshoot. Um, this isn't just a
transition from on-rem to the cloud.
This is now the current operating model,
right? This is the current state of
affairs. Um, so it's not slowing down.
It's only going to increase. Um the
these bigger, more complicated workloads
mean more moving parts, more things you
can miss. Um without a single pane of
glass, there's not a quick way to get to
a fast answer. And yes, when stuff goes
down, um customers impact it. I don't
have a poll for this, but pretend you're
raising your hand. Uh how many of you
are now working with GPUs or uh LLMs, um
and adding LLM based features into your
products? Um a lot of folks are. That's
yet another layer of complexity that's
coming in. Um, and with all of these
different layers and pieces of the
puzzle, if any one of them fails, um,
then this translates into customer pain,
lost sales, lost revenue, um, other
problems. So, it it really is a
challenging environment to be working in
today. Um, why is it challenging? It's
not about just collecting the data,
right? There's lots of tools that can
collect metrics or logs or traces or
other events. Um the challenge is
bringing all this information together
in one place and being able to correlate
across all of these different signals so
that you can see what's going on.
Network traffic is another one. Um when
you're troubleshooting an incident,
often one team has to do investigate
their thing, hand off to another team
who investigates their thing, hand off
to another team. Folks may not know how
all the pieces connect. um you rely on
tribal knowledge which if somebody
leaves the team or leaves the
organization suddenly that information
is lost because potentially
documentation is far out of date. Um so
this leads to ultimately slower incident
response um takes longer to uh hand
things off to different teams um hire to
communicate and at the end of the day it
hits the bottom line right um this is
the kind of thing that we're trying to
fix. So this is like a very standard.
Take a look at this and tell me if this
sounds familiar to you, right? An alert
fires at minute zero. A SE 2 is
declared. Um people start looking into
it uh on one team. Um they start
investigating. They don't find a clear
answer. Maybe the application team
starts looking into it. They don't find
an issue there. They escalate to the
networking team. Networking team
escalates to the compute team. Right?
and this bounces around and over an hour
later um eventually we get to the root
cause right um this is painful it's
timeconuming it's ultimately it's costly
um this is what happens in real life in
this survey we found that 87% of people
said that their organization could not
resolve a cross environment incident in
under an hour so the 65 minutes uh this
is real this is what a lot of folks are
dealing with how much does it cost Uh it
can cost a lot. Um ITIC uh recently did
research and found that mid to large-siz
corporations um an hour of downtime can
cost $300,000 if you're a Fortune 500
company and cost half a million to a
million dollars per hour. Um or
potentially even more in certain
industries like finance and healthcare.
Sometimes it can be $5 million per hour
of downtime. Um so the point is that
every time every minute that you spend
trying to track down the source of the
issue um is adding to that clock time,
right? So it's it's and if you can speed
that up, it's going to directly reduce
the cost to the business.
So let's do another poll. Um how much of
this resonates with you? If you're have
a production alert go off, um where do
you spend most of your time before your
team starts to resolve it? Is it
figuring out who owns the problem? Uh,
is it trying to correlate data across
the multiple tools that you use? Um, are
you waiting for that one person who
knows what's going on to get looped into
the the Zoom call? Um, or do you really
have all this figured out already, which
is great. If you do, that's fantastic.
Uh, we'll give it a minute or two more
for a couple or a couple more seconds
for folks to uh submit their uh
submit their uh the polls.
All right. So, let's see. But what do
the numbers say about this? So, what we
found is
um almost everybody says one of those
things, right? It's a very small
percentage of you that said, "Oh, we got
this dialed in." [laughter]
So, um and that's really what Data Dog
is here to help solve.
So for everyone seeing data dog for the
first time, this is what data dog looks
like at a high level. The point of data
dog is really that we bring all this
information together into a single pane
of glass. We bring all your teams
together into a into a single place so
that you have clear communication, clear
continuity. You can connect the dots
between all this information um and use
this to troubleshoot every single type
that we've talked about. metrics,
events, logs, traces, they all flow into
this single platform. Network,
infrastructure, application monitoring,
all the layers, they're all connected
together, right? There's one data dog
agent, one platform,
this single unified experience across
every team.
Uh, and here's why that matters when
something, you know, actually breaks,
right? Um, it means that these signals
can be correlated together
automatically. You can see what logs
correlate with what trace. You can see
what metrics correlate with which logs.
There's no switching between tools, no
switching between signals, no switching
between teams. Everybody sees the same
information in the same place. Nothing
is lost in translation. And now with
Data Dog's AI tools, with Bits, which
I'm going to show you a little bit
later, um we can actually help you in
investigating and figuring out the root
cause even faster.
So what does this look like in practice?
I'm going to give you a demo in a
minute, but at a high level, you can see
here, this is the CloudCraft diagram in
Data Dog. And what this shows is the
full picture. This is what your whole
stack looks like mapped out in one
diagram. It's a visual diagram, so you
can clearly see what's going on before
something breaks. How do things connect
to each other? What are the dependencies
at the infrastructure and at the
software level?
You can't fix ultimately what you can't
see. If you have out-of-date diagrams,
this will give you clear understanding
of how infrastructure works. Even if
it's not the infrastructure that you're
normally responsible for. Now, let's see
what happens when something goes wrong.
You can now in data dog follow the
incident across different areas. Whether
you're looking at the software layer, uh
whether you're looking at logs, traces,
like I mentioned,
all of this, there's no tool switching,
no new login. You're not losing any
context. You have a consistent
experience all the way, right? This is
what we wish the on call engineer had
right at t equals 0 when the incident is
called in the first place. And now with
BITS, we actually use AI to connect the
dots for you, saving you even more time.
So rather than Yes, you can use all of
these tools to investigate, to
troubleshoot, to confirm. Uh, but Bits
will do a lot of the job for you right
out of the gate and give you a head
start speeding up your remediation,
finding what the root cause looked like.
So let's what does this look like in
practice? Let's jump right in and take a
look at what this looks like in the data
dog environment. So, let's say that I'm
on call. I got paged for an alert. In
this case, it's relatively simple
looking alert. It says, "Hey, system
load is high on this host."
Um, I can go down. I can see more
information about this. I'm going to see
my diagram right here that shows me,
hey, here's my host and here's the
infrastructure around it. Um, now I'm
going to click and open up this diagram
in full just so I can explore in a
little more detail. So, here's that host
again. Now I've clicked on this host and
I'm seeing all of the telemetry related
to this host in one place. I can see the
key things that are going wrong. In this
case the system load is high. Uh metrics
of this host. I can look at if it's
running Kubernetes which in this case it
is. I can see all the pods that are
running on it. Logs, traces, network
traffic. If there's security issues I
can see that. I bring all this
information together in one place. It's
like a single pane of glass.
Furthermore, looking at this diagram, I
can see that this host is actually
running this. It's one of the nodes for
this Kubernetes cluster. And in this
Kubernetes cluster over here, I can see
that there's three services, orders,
app, web store, and product
recommendation that all have errors. So,
I'm going to hover over this and see
exactly what uh the product
recommendation status is. And it looks
like the health is critical. So, I'm
going to click onto the service page for
that. And now I'm looking at the service
product recommendation and I can see the
health is critical. I can see what's
going on and I can go down the page and
look at the specific traces and uh logs
and other information about this service
again all in one place. If there's
database queries I can see that and I
can really use this to drill drill uh
drill down quickly and find the root
cause of the issue. But while I'm doing
that, I may want to simply kick off an
investigation. So I'm going to go here
and I'm just going to say, "Hey, please
investigate this with Bits." I'm going
to click here and Bits is going to run
an investigation for me. So what is Bits
going to do? Bits is going to
start with this monitor alert, explore
multiple hypotheses of what could be
going wrong, look at the different
telemetry, all of this correlated
telemetry and metrics to correlate
things directly together
and ultimately verify that some of these
hypotheses seem to be correct. And what
this is going to show me is it's going
to show me uh that in this case CPU
overload has degraded the product
recommendation service which in turn had
kind of this knock-on effect. And the
recommendation here is to ultimately uh
scale up my Kubernetes cluster. Um and
it's even giving me a PR um so that I
can deploy this via Terraform um and uh
and make these changes.
So this is a really great way to see
this. Now I mentioned earlier
correlation across cloud and uh hybrid
environments. I'll just throw back here
a little bit. Um this of course is
showing me the cloud part of my
environment but I can also go here and
just as easily see my onrem in this case
this is a a VMware cluster. I can see my
cluster. I can see alerts on different
things that are going off and if the
issue is running onrem everything that I
just showed also exists for onrem. So
again, here's my list and I'm going to
look at this resource and see all of
this information together in one place.
Works just the same for on-rem as it
does for the cloud. So I can clearly
correlate all this information together
and troubleshoot and ultimately get to
the root cause. So I can of course go in
and use all this all this great tools to
verify that bits did find the right
solution and then I can go ahead and
implement it.
All right. So now that we've seen how
this works in practice in data dog um
you've seen how data dog correlates
these signals together. We've shown um
how uh we can immediately kind of start
looking at uh things going wrong out of
the box. We've shown how bits can detect
and investigate and figure out the root
cause for you automatically.
And we've shown how this ultimately
gives you one pane of glass that your
all of your teams can rely on that has
the same data for everybody that gives a
single view so that we don't have
different teams in different silos
pointing fingers at each other. So which
of these would be most useful for you?
All right, great. So we're getting some
responses in the poll here.
So I think the answer for most people is
more than one of these things.
Um and there's just an example. Um
Porsche Informatic. Um of course Porsche
is a very large company. They do
business in 29 countries. Um they have
over 500 retail locations and they had a
lot of challenges. Um before they
started working with Data Dog, they had
10 different observability tools.
Um and these 10 different monitoring
tools were really creating a lot of
complexity, slowing their team's ability
to innovate, uh impacting the quality of
their service, driving up their costs.
Um
once they started using data dog, they
were able to consolidate all of this
into the single data dog platform. Now
all of their teams had complete
visibility into every part of their
environment all in one place. And this
unified view really meant that uh they
could detect issues faster, troubleshoot
issues faster um and ultimately respond
faster to all the dealerships they have
in the field. So the law so the the
bottom line impact here was they were
able to achieve 99.5%
availability, improve the root cause
analysis, uh reduce their time to
insights and also optimize their costs
for their infrastructure spend. So
Porsche's leadership really credited
this consolidation with empowering their
teams to move faster, improve the
customer experience, innovate at scale
in a way that they couldn't do before
when their observability uh tools were
fragmented.
So this was a really uh great success
story and it really illustrates the
benefits of data dog. Right? before you
have tools that are disconnected from
each other, teams that are disconnected
from each other. These different data is
all in different silos, hard to
correlate, it's hard to hand off. Um,
often you don't know who's going to know
what. Um, and at the end of the day,
this means slower time to resolution.
Um, with Data Dog, you get everything in
one place. unified visibility, faster
time to resolution, correlating your
signals across your environment, makes
it much easier to hand things off
between teams, quicker to get to a root
cause. Um, and now with Bits, Bits can
actually do a lot of that work for you
right up front and speed up even more of
the time to the root cause because
rather than having to dig through your
telemetry to figure it out yourself, now
Bits gives you the pointer to look in
the right direction and all you have to
do is validate it and then implement the
fix.
So that's uh that's the uh the overview
of some of the great benefits of data
dog. I know there's actually a lot of
other great features in data dog that I
didn't have time to cover, but I think
that gives an overview of what some of
the great benefits are and how data dog
can really help you working in your
hybrid environments. Um so with that, uh
we'll take a pause here and open the
floor to questions.
Uh, let's see here. I see one question
already. So, I'm going to go ahead and
uh and answer that question first and
then we'll uh and then we'll uh see if
we get more questions as we go along.
But please, as I'm answering the
question, please do uh come on in and uh
and uh put more questions in the chat
and I'm happy to uh answer more of them.
So this question says uh I spent most of
my time identifying the misconfiguration
of infrastructure devices, routers,
servers, applications that skewed the
log data by the wrong date or time or
was not sending logs to all the right
places or the application was
misidentified with the type of data it
was supposed to send. [laughter]
How does data dog fix those issues?
[snorts] Um that's uh a great question.
Um, I think at the end of the day, data
dog fixes those issues by helping you,
like I said, see all the information in
one place. Uh, I'm going to use um,
CloudCraft as an example here just to
show you kind of what this might look
like. But for example, um, in this
Cloudcraft diagram, I can zoom in here
and say, "Show me logs." And for some of
these resource types, I can actually
see, for example, these lambdas, I can
see which of them are reporting logs and
click on a particular lambda to see just
the logs for that lambda. So if you're
trying to troubleshoot for a specific
resource, whether it's configured
properly and whether the logs are shown
up correctly, you can see that here
immediately, right? Um and then
similarly, if we go into let's say our
um our host list here, I can actually
see a list of all of the hosts that I'm
monitoring. Um and I'm going to say be
able to look at each of these hosts and
see uh so for example, this is a um I
just opened this up, but this is an
Ubuntu host. Um,
and uh, I can see that it's running the
agent. I can go in here and to get that
same unified view, I can search my hosts
to find exactly the host or the subset
of hosts that I want and manage the
whole view to see all this information
in one place. So, I can quickly say uh,
I probably picked a bad example here
because this host doesn't have any logs
or traces at the moment, but if I had
logs and traces, I could just jump
between them here and see exactly what's
going on. Um, so the the advantage of
having all this information in one place
is that you can quickly kind of
crossorrelate this data um to see what's
going on and help you troubleshoot.
Yeah. Okay. Looks like this log has the
wrong time stamp because I know it was
just generated, but it says it was
generated 4 hours ago or something like
that.
Let's see. Um, any other questions from
the group? Please feel free to throw
more questions in the chat here.
Let's see. Looking in the chat, I don't
see a lot of other questions.
Give it one more minute. And if anyone
has any more questions you'd like to
ask, feel free to throw it in the chat.
Oh, I see one more question come in.
Let's take a look.
Great question. So, um, this question
says, "Large corporations may have an
issue with the cost of sending all the
log data to a single device as many
times a lot of that data is unnecessary.
How do you determine the specific data
to send?" Um, so I'll say I'm I'm not a
super expert on logs within Data Dog,
but I do know that Data Dog has a lot of
capabilities to pre-filter logs before
they get sent to Data Dog. Um, and
there's a number of different tools you
can also use to uh view and manage your
log uh spend and make sure that you're
only sending the important logs um or in
some cases like subfilter logs and only
show like the the types of logs that you
think are the most important. Um, so
there's definitely a lot of capabilities
within data dog to uh make sure that you
are getting good value for the the logs
that you're spending and you're not
spending money on logged volume that's
not important or not useful.
Let's see. We'll give it one more minute
for other questions.
Also, if you have questions about any
other capabilities of Data Dog that
you've heard about, you'd like to learn
more about, feel free to ask.
Or if you have any particular issues
that are relevant for you and you'd like
to ask about it and uh and see how data
can help you, feel free to throw that in
the chat and uh happy to answer
questions about specific use cases or
problems that you're trying to solve.
All
right.
Think we're not getting any more
questions. So, I think we will uh
we will uh end the webinar here. Thanks
everyone for your attention and for
taking the time today. And yeah, if you
can scan this QR code below, you'll get
a 14-day free trial uh for Data Dog. Or
if you just go to the Data Dog website
and sign up, you can get a 14-day free
trial. Try it out. No credit card is
required to do the free trial. So, you
can just jump right in and try things
out and see how it works for you. Um and
uh yeah, we hope you try out Data Dog
and start using it. We think you'll get
a lot of benefit from it because it's
going to solve a lot of uh problems that
you might be having with your current
observability tools
um and a lot of challenges that we have
in today's kind of modern uh complex
hybrid environments.
So, thanks again once uh once again all
of you for your time and I'm going to
hand the ball back over to Linux
Foundation.
>> Thank you so much Jace for your time
today and thank you everyone for joining
us. As a reminder, this recording will
be on the Linux Foundation's YouTube
page later today. We hope you join us
for future webinars. Have a wonderful
day.