Video summary
The March 27, 2026 episode of the OpenStack Ops Radio Hour centered on a comprehensive discussion regarding OVN, the default software-defined networking solution for OpenStack, alongside broader operational strategies. Participants shared valuable migration experiences, detailing successful transitions from legacy solutions like Linux Bridge and Nova Network to OVN, with specific workflows involving service stopping, namespace cleanup, and configuration of ML2 settings. While acknowledging that OVN presents a steep operational learning curve rather than inherent instability, the group highlighted technical considerations such as potential packet drops under high traffic loads and the single-process nature of Northbound database conversions, suggesting mitigations like DPDK for Open vSwitch bottlenecks. The conversation also touched upon scalability in massive clusters, noting that while OVN-IC offers promising interconnect capabilities, it remains bleeding-edge with limited documentation compared to established Nova cell architectures.
Beyond networking specifics, the dialogue shifted to critical best practices for observability and community sustainability. Mark Schushlin from Bloomberg emphasized the importance of managing large-scale environments using tools like Telegraf and Grafana while warning against the risk of monitoring systems crashing production databases under stress. The group also issued a call to action for volunteers to assist Red Hat in maintaining RPM packages, as the ecosystem is moving toward container images only. These technical insights were framed within a strategic decision to align future episodes with Project Technical Gathering days to maximize attendance and developer availability, ensuring that discussions occur when key contributors are most accessible.
The meeting concluded with administrative updates regarding the Project Team Group tooling, confirming its configurability to allow teams to register specific MeetPad links and override default EtherPads for better event management. The host noted a personal conflict requiring an early departure but affirmed the value of the PTG concept, leading to a consensus on scheduling the next session for April 24th. Participants were instructed to add new agenda items at the top of the list, with an invitation extended for others to submit discussion topics or request specific experts, such as an observability specialist, for upcoming sessions before the meeting was officially adjourned.
Read the full video transcript
The suggestion for this week was to have
a group discussion about OVN, which I
think is the default software-defined
networking for OpenStack. Um I have to
have a huge disclaimer. We don't use it
where I work,
so I don't know much about it. Um
And then um other other things on the
agenda are uh observability and then um
a note about uh RPM packaging volunteers
for RDO.
So um
do we have OVN practitioners here who
can maybe um
help the discussion move you know get
off the ground about um
migrating to OVN and experience with it.
Uh
I see on the chat that Ildiko's saying
again that um
you know um
feel free to to write more things onto
the agenda
uh under topics for us to um to discuss.
So
um
It's going to be a very difficult
discussion if we don't have any OVN
practitioners here. How I will say that.
Hello. Just speak up. Hi. Um it's Luis
from CERN. Just uh to to mention that in
our case we are migrating to to OVN
uh from Linux bridge and we are getting
our hands into operating
OVN. Uh people are interested in in the
part of migrating from Linux bridge to
OVN. This is something we have
published a couple of um
uh blog posts last year and also uh a
couple of uh one presentation in in
OpenStack Summit in Paris.
Uh so yeah, you you could take a look to
to those sources of information. We are
running as well uh
kind of
interest group of operators migrating
from Linux bridge to to OVN. Uh we have
to schedule the the next meeting to to
follow up on on that. I don't know if
anyone in in the room today is in the
same boat of moving from Linux bridge to
to OVN.
>> [snorts]
>> Hey Kalle.
Hi.
Hello.
How's it going? Are you already on it or
No, we're just
because we have a much more complex
network structure than you because you
have luckily have quite a simple we are
still considering is it OVS or OVN, but
we need to hammer out those plans.
Okay.
For us at at the moment the the
main worry is to to get the operational
experience with with OVN.
Uh we are now in the phase of
creating different scenarios, break it,
uh get confidence on how to recover it
and and these sort of things.
So where where do other people stand on
this? Um
we we seem to have two
uh people with some OVN experience. I
think Luis is representing CERN that
they are moving to it and Kalle
[clears throat] is representing that his
organization that they are Yeah, we're
seeing
We are seeing CS in Finland. Yes, we are
assessing it. We are looking into it,
but we're also on the old Linux bridge
boat, which is a bit more challenging
move than I think from OVS.
So we it's not trivial.
Uh by the way, I I forgot to emphasize
one thing that's useful is if you can
put [clears throat] your name and
organization on the attendees list on
the um etherpad. Sorry, I actually
didn't put
my employer, but uh it's done now.
Um
that just gives us some idea um you know
who's talking and and whether they're
maybe in you know affiliated with the
same
um
organizations. Uh it's particularly
helpful when the the colors that you're
getting automatically in etherpad
change. So if you sort of think like oh
um I remember you know um Chris was
uh aquamarine, well, it might might be
purple next time, right? So names are
useful.
Um
Yeah, we so not moving from uh Linux
bridge to OVN, but uh
our old OpenStack was um Nova network,
which was the original very limited uh
thing before
Neutron and uh
it was so difficult to imagine moving
that we actually did greenfield new
OpenStack deployments when we moved to
Neutron.
Um
which uh my business sponsor told me
that's the last time you ever get to do
new clusters.
Uh
So uh it's very important to um
to get on the
on the Neutron track and then
um
assessing your network thing is is uh
you know which choice you have of
Neutron plugins is critical because um
it's extremely difficult to move after
that and the users really don't want
their cloud to
go away and and for for the operators to
say oh, there's another cloud now.
So um it's kind of pivotal. Um
So yeah, I I I feel that we actually um
switched off our Nova network clusters
within the past couple of months.
Finally.
Um
Someone's asking to participate. I don't
know if that's um
something the admin needs to do. Um but
just go go ahead and speak. Oh, I see
you have your hand up, Eduardo.
Uh
you guys are
hear me?
Yes, can hear you. Go ahead.
Um
Uh first of all, I will introduce
myself. I am uh
I work in a lab in
UFCG University in Brazil and my first
step first task in my job is
migration OVS to OVN.
And I can share a link
with you guys.
It's a GitHub
repository.
Uh are you share Did you say you're
share you're going to share a link?
Yes, yes.
Uh do you have the etherpad? I think
that's the best place to share it. Uh
I'm as I mentioned the shared document
in meetpad is an etherpad and then
If I writing
in here,
it's good to you guys? So that's the
that yeah.
Yes. That's the uh So you're in the uh
etherpad chat. There is also a Jitsi
chat,
which other people are chatting on and I
think uh honestly it's a little better
to uh
to share it here. I don't know if you're
seeing the doc. Um
Uh in the doc. Yeah, just go right ahead
and write on it. But so
again, people should should feel free to
speak up.
Um and they should feel free to write in
the etherpad um
that their contributions. So um I've
transferred it over, Eduardo, so there's
something for people to look at. This is
a migration tool you said from OVS to
OVN?
Yes, yes. I can
uh
when I I was creating this migration,
I'm looking to others tool
uh we have to
a triple
and um
with these two I was able to do this
migration for Cola.
Uh first of all, we we need the I will
speak step by step the workflow.
Uh first of all,
uh you need the stop the demon service
from
Neutron with OVS.
Um
stop the Neutron service.
When you do this, you you need the
remove existing OVS flows in OV uh
in open stack the switch.
Uh the flow is basically
when you receive a package package, you
need the
this package it's forwarded to another
port. Be
it's a
OVS agent make this.
Um sorry for my English. I'm not too
good with this.
Um
After this, you need clean
uh BR-tun it's no longer used it
in OVN
and you need clean all namespace.
When you finish this, that you
you need you can
uh deploy O
VN.
And when you deploy O VN with
Cola normally,
uh
the Cola will make all configurations
you need. Uh set the DB set the
database, set the chassis chassis.
All all these you
it's making by Cola.
But after you deploy O VN,
the network don't work because you need
three other configuration OML2
configuration file.
It's a Neutron sync mode. You need to
set repair.
OVNL3
mode, you need the
set to to O VN
the L3
network and finally VIF type, you need
the set to OVS.
When you finish that and reconfigure all
all the Neutron service,
uh the network works
uh really well, but uh
it's I
have one prob- one problem with the
database.
Uh has one table in database in Neutron
database
uh called provider association
reassociations.
Uh in this table uh, when um
OVS
it's called it has an an entry uh,
called single node it's re-enable to
OVS, you need uh, delete this table and
reconfigure Open vSwitch. It's will
make
the OVN assume the control of the
backend Neutron backend. After all that
you will make the migration works well.
Great. Um, so Eduardo, um, you know,
this is recorded and people can can hear
your your walk through there. But it
sounds like that might be a a great
thing to actually blog just the the the
trans- transformation process that you
did if if you get a chance. Um
So, thank you for that. Um So
I I I think we are trying to have a
general discussion about uh, experience
with OVN. So, we've got some some some
good material here about getting to it.
Um, how are people finding it once
they've got it working in terms of
performance and scalability?
I think from from our side it's been a
nice experience. One of the things we
appreciate moving from Linux bridge to
OVN is that we get rid of
our big rabbit and queue that we
we need to to handle all our
connectivity with Linux bridge and the
hypervisors and we rely now on the OVNDB
and this is way more efficient. Um, so
in in that regard it's it's quite nice.
We had a scare because the technology is
complex, but it's true as well that so
far has been stable. We have adopted
around
500 hypervisors um,
with a few thousands of virtual machines
connected to the to to a cluster. So so
far so good.
We have had some some weird bugs that we
had to to fight with with some traffic
being forwarded in the wrong places, but
apart from that
Callie, what you were saying? If I
remember correctly, you were at some
point there was only one copy of was it
North DB or something running it and you
were a bit worried about that.
There was one database in OVN that only
had one copy running, I think.
We I was not redundant.
Our OVNDB is a cluster of three nodes.
So, you have have a weird DB that is
clustered. And then is yeah, is the the
the Naughty doing the conversions that
it has one process that uh
is just running on one of the boxes.
It's unlocked
to do the conversions between between
the Northbound database and the
Southbound database.
And um, you um
you said it's stable and it's nice
experience. Um, the actual um,
networking performance is is is good?
In that area, one of the
the things that happens is that you have
Open vSwitch running on the hypervisors
and this is something that poses some
load in the of the way it works that
while processing um, all the incoming
traffic into the hypervisors, there
could be some bottleneck. So, if you
have a lot of traffic
then you could experience some packet
drops because Open vSwitch is not coping
fine with the with all the the ingress
of the hypervisor. There are different
technologies that if people hit that,
they could explore like something called
DPDK. So, it would be an implementation
a different implementation of Open
vSwitch that would make everything uh,
user space. So, there would be less
jumps between kernel and user space. So,
it would make everything more efficient,
but it comes with some operational
implications also to to run it.
In our case, we hit this scenario in
some corners sometimes with some
uh, let's say
demanding
uh, workloads, but it's not that common
at the moment. But it's something we
will explore at some point all this part
of
Open vSwitch DPDK.
Do you know is it the OVS flows that are
the bottleneck and
or is it the
Yeah, I don't know about how big is the
bottleneck like is there any traffic
amounts that is hard to push through
there?
Yeah, the the the the what what what you
would experience is basically packet
drops in the ingress of the hypervisor
depending on the on the combination of
workloads, this could be
more of a problem or just some low noise
that is fine.
Yeah, that that's actually very
interesting. Um
our um
directions were to have a networking
thing that was always wire speed and
couldn't have such things. So, we don't
use OVN. So, I'm I'm always been curious
about it. We use a different Neutron
thing called Calico which is pretty
niche for OpenStack.
Um, it doesn't have any agents
forwarding packets in software at all
anywhere.
But I'm
I'm definitely not doing a sales pitch
for it. I've tried to capture the some
of that discussion Luis and and Callie.
Um, if if you see on the etherpad,
please keep me honest and if I if I've
said anything incorrect um, you know, or
incomplete, just just dive right in and
edit that.
You know, some of the people who said
let's have a discussion about OVN were
just unsure whether to even look at it
or you know, unsure of the the potential
for it. Sounds like uh
CERN's happy with it um, apart from what
I think you said some weird bugs and
then occasionally you can get um
hotspots and packet drops can happen if
there's a tremendous amount of ingress
to VMs, right? Exactly, yeah.
Um, but for us the the main worry at the
moment is like is new component is
complex. So, people when moving to it
should
uh, plan for some time of
learning and getting used to all the
moving pieces and the way all the
networking part is represented in OVN.
And I guess the actual the possible
packet drops and everything is not
actually the problem of OVN itself. It
will be there even if you just use OVS
by itself.
Yeah, could be, yeah.
Yeah.
Okay, great. We're I'm starting to see
starting to see like sort of
conversation capturing uh, onto the
etherpad there. Again, if people aren't
sure what we're seeing um, in my screen
I have the etherpad front and center. I
have the Jitsi chat left and the video
thumbnails right. So, I can see
everything that's going on in the
meeting.
Um [snorts]
and I see some people also also using
the etherpad chat. So, there's
unfortunately two competing chats. Um,
so
that's fine, but some people may not
have those open. So
I'm typing into the etherpad chat now.
Um, so if you didn't see me if you
didn't see a message arrive, you're not
looking at the etherpad chat. Now, I'm
going to chat into the um
Jitsi chat uh, if I can find the right
message window.
And again, so if you if um, it's it's
kind of a shame that meetpad has two
separate chat windows cuz it makes it
very fractured, but um,
I just typed in both of them. So, if you
haven't found those please do so that
you know, you're actually getting the
full experience. Anyway, [snorts] the
main thing is the etherpad is
uh
meant to capture what we discuss and it
turns into a great resource um,
for reference later. You know, often
people say
something brief and then they drop a
link
and um, none of us go and explore the
link during the meeting, but we can go
afterwards and uh, explore it. So, for
instance, we've got the CERN blog uh,
Eduardo's migration tool. This is
fantastic. Um
and now I think someone's um
linked to a OVN bug so that on a YouTube
video. So, that's that's great.
>> [snorts]
>> Um
Uh, may I ask a question for Eduardo
because I found the tool very
interesting even though we are not using
Cola.
Did for that approach, how big clusters
have you been
tested in testing with and how big is
the actual
data plane downtime when you do this
approach?
Uh, I I don't try in bigger clusters.
It's uh, I need to improve my tool.
But it's the downtime it's like 10 15
minutes.
Perfect. Thank you.
I'm again trying to capture the
smattering of that. So, I think I heard
that the question was uh
how large a cluster is the Cola
migration tool been tried on and um
Eduardo says it hasn't been used on the
the biggest ones, but on the test
clusters it took 50 minutes to convert.
Uh
Okay. Um
The tool needs some improve and I I'm
testing the all scenarios, but it's in
my scenario actually it's worked well,
but it's ha- has a problem. I can't
I can't testing with the Octavia
service.
It's one of the problems.
Okay, so um that's that's been captured
in the section about that tool, which is
great. Um go ahead.
I think Carl, you were speaking.
No, it was just to clarify.
Hasn't you haven't been able to test it
with Octavia or you haven't been able to
get it working with Octavia?
Able to test. Okay.
I I
the tool is in
uh testing environment.
I'm testing the all possibilities.
Thank you.
Okay, so um a new question uh
posted in this area uh possibly from I
think Mike.
Uh any experience with OVN-IC,
which I'm not sure what that is.
Um
that's to the reporter.
Right.
It's a interconnect space, so
uh
uh hi everyone. I'm Mike. I represent
Superfluid.
We have recently
restarted our interest into OpenStack,
so I'm new here.
So
uh glad to have found this. Um
when you look at Nova, it has cells, so
you can do uh
a really good spreading of your uh
compute space for Nova.
The OVN part is one big
setup. So they've tried to create a
cell-like structure using different OVN
cells, which you can interconnect.
That's the word IC.
With each other using transit grade
bridges as sorry, transit transit
routers and transit switches. I'm
actually wondering if anyone has
actually any experience with that.
I'm guessing the answer is no. Uh that's
very interesting, Carl.
Cells is something that we wish Neutron
supported in uh in our context because
our Neutron is um
you know, as you said, there's only one
for a an OpenStack deployment, and so
um
it's sort of a single point of failure.
Um
And I'm you know, of course, you can run
your your database in a clustered mode,
so it's not a single server point of
failure, but nevertheless, if the
Neutron service is down, your cluster is
very impacted. Um
but I it sounds like
uh Mike, nobody else here has actually
tried OVN-IC.
Are you trying it now or just something
that you're looking at that you'd like
to use, but you're looking for
experience?
Uh what I've found so far is that in the
Kubernetes space, OVN seems to be active
as well.
And it seems that Red Hat with OpenShift
has been
pushing OVN OVN forward quite rapidly.
And I get the impression that in the
Kubernetes space, the interconnect is in
a beta stage, so I was was wondering
since I didn't really find any
relevant
documentation or references in Neutron
or the ML2 uh
driver for OVN
to reference this, so I was wondering if
anyone even knows about it.
I think the answer is not on this call,
but that's um that's fascinating cuz um
that's a key concern of ours, obviously
with a different Neutron implementation.
Um
and that's actually something that's a
bit unfortunate with OpenStack in
general is that there's
um
sharding primitives, shall we say, but
they're not really coherent from the
point of view of the whole OpenStack
ecosystem, so
Nova cells, for instance, allow you to
have separate rabbit and other services,
but uh
Neutron just
hasn't done that uh
inherently.
Um
Not that you want to maybe make
OpenStack clusters, you know,
too gigantic anyway, but um
um
we're looking at a deployment where
um
it's a new data a new data center site
with a multiple buildings,
and um
you know, we're reluctant to have that
be one Neutron
um
mediated
software-defined network because we
couldn't have um
a gigantic new data center be
have have its cloud be entirely down, so
we're actually splitting it up, um but
we're using regions.
Um
So
effectively somewhat completely separate
OpenStack
clusters. Um
So
something like this OVN interconnect, I
think I heard you say,
sounds very promising for OVN users, but
uh doesn't sound like anyone else on the
call is even aware of it, so
um
maybe we can ask that on the Matrix
uh channel and see if there are people
interested and maybe they can come and
talk about it
in the future.
Uh
hi Chris. Um I'm also from CERN as well.
The yes, looking at what the kind of OVN
interconnect is specifically to connect
clusters in between. I'm not really sure
that it will do the same thing as cells
because those clusters will be
independent from each other. And what it
enables this is to connect to have
switches that need to connect routers
that are in two different OVN clusters.
So I'm not really sure that we'll do
kind of the cell stuff that the that
Mike is
like uh referring to.
And we we at CERN what we what we ended
up doing was like to have two different
regions that are completely independent
from each other as Chris mentioned to
with two OVN clusters that are
independent, and then they are connected
with basically layer three in within
each other.
I think that's a fair statement indeed.
It's Neutron will not be like Nova
cells.
Uh what in
in that context, if you look at Nova, it
is a cell zero, which is basically your
Neutron stuff, and then you have the
cell one, which would be your OVN
implementation. It's kind of similar,
but it's not fully similar.
No, it's not it's not actually this
would allow you to kind of connect those
those clusters have a kind of a switch
that goes across regions. So if you have
two OVN clusters in two different
regions, you can have a switch and a
genie tunnel in between those two
basically using these interconnects, but
I don't think there is anything on the
driver that the Neutron driver does that
allows that.
That's in the indeed my uh
uh conclusions so far as well.
Hey, seems like that's very bleeding
edge stuff.
Um
If if we expand into any OVN
deployments, I'll be glad that I'm aware
of that. Um
We're uh we're tasked with uh
thinking ahead to scalability problems
at you know, bigger and bigger scales.
Um so this sounds like something that I
should be aware of. Thank you.
Um
so I want to make sure that we do cover
the whole agenda. Um
so any more questions or observations or
things we want to share about OVN
because we do have another topic.
Um
So it's sort of last call on that.
Please just go ahead. I see people using
hand up or waiting, but just speak up.
That's me.
Um I I
as e- everyone
um said here,
uh
people seems to get um
good experience on using OVN. But I
think uh
we haven't uh
uh
uh good tools to migrate an
an en- environment from OVS to OVN. So
we think um I I I I think we we
we should have better um
in- um tools to
migrate because um for some clouds, um
it's not possible to just um uh
you stop the cloud and and just um
do everything with all OVN, and I I
think it's very
important that we we we should have um
um some tools on the
deployment tools like Cola um
um
charts and all the
uh
deployment tools to help
o- o- operators to migrate from OVS to
OVN. Um my question is to
the to the CERN people here, the CERN
people. Um
I I've seen that you guys
uh wrote
some posts about migrating from OVS to
OVN. Uh do you guys have any um
a tool or In our in our case, it's not
from OVS to OVN, it's from Linux bridge
to OVN.
Oh, I see.
And and as you will see in the if you
take a look to the to the blog post or
or to some presentations,
we were lucky in a way because our
previous environment based on Linux
bridges is rather basic.
So, while migrating, we could make some
assumptions on the connectivity of the
VMs, on the number of ports, or name of
the bridges.
So, this facilitated a lot of our work.
Actually, in our case, we can even live
migrate
uh VMs from Linux bridge to OVN, so zero
downtime for for the users. And this is
something when you talk to people in
setups more complex with OBS and private
networks and many things, this is
uh this is not applicable and people
need to be creative to
to find a way to migrate away from that.
So, we don't have experience from OBS to
OVN. Our use case is only Linux bridge
to OVN. I see. I see. Thanks.
I
I think it's a very valid point, and I
put a question there either, but if
people could fill it in, who is actually
building their own migration tools here
because I assume everybody will at some
point build their own tool. And it might
be good to be of this and improve each
other's rather than building our own 50
different times.
Uh that's actually quite impressive that
uh live migration from an old an old
neutron thing to a new neutron thing um
in this case Linux bridge to OVN
actually worked.
Um
These days
>> we are surprised as well.
These days our our
our users expect a zero downtime.
Um we did used to do, you know, the the
cloud will be down this weekend.
Try again on Monday. It's kind of
upgrades.
Uh we can't get away with that anymore,
so that's that's very encouraging. Is
that actually in the blogs that you've
linked?
Uh yeah, it's in
a series of a series of
four blog posts, more or less, and it
explains a little bit our our scenario
and a bit the different paths we we
took. Because also we we are in the
process of doing some refurbishment of
our data centers, so part of the
capacity was not live migrated, but cold
migrated from Linux bridge to OVN and at
the same time new hardware. But another
part of the infrastructure was live
migrated and and it is described how how
we did it.
And if people is interested on knowing
more, we we can reach out to us in the
metric channel or
just contact us on IRC or whatever you
use.
That's fantastic. So, um if anyone's
just become interested in that, it's the
first link on today's thing. So, it's uh
tech blog web cern. ch
That link at the top is apparently a
four-part on this, so that's a
That's great. Oh, and someone's actually
linked to the the specific thing about
the live migration.
Um
Yeah, I was just going to point out that
that it works with a caveat that for
some reason when we migrated just once,
the
the ports were marked as downs, we had
to do it twice, but then it worked.
Okay, so some
That's some That's some
>> explained in the blog post. Yeah, that's
fantastic.
Uh I really appreciate the openness
there. I know that it's part of CERN's
charter. Um
And I'm saying that in the sense that uh
we have some stuff that we want to
publish, but we just haven't quite
connected with it for quite a while.
Um so, it's very inspiring inspirational
to me. Um I'll be reading those blogs,
and uh maybe we can uh
share some of our stuff in a similar
fashion.
>> [snorts]
>> Um talk about live migration. Um we I
think in the previous one that people
enjoyed um
previous Ops Radio hour, we talked about
um how we've been working on live
migration performance.
Um because we had apps that broke when
we live migrated the VMs. And um
we are now thinking that we're getting
close to being um
a 3-second interruption for VMs
and generally almost no CPU throttling.
Um
And that is something that we definitely
want to publish either as a blog, but
possibly
um
take it to a conference
um in the future. But if people do have
questions about what we've done,
as someone just said, you know, um I'm
attempting to stay on the matrix
Ops rate Ops op- operators channel um
more often, and um I'll take questions
on that and share whatever we can.
Okay, so we've got about 25 minutes
left. Um
I think we've done a lot on OVN, and
there's obviously a lot more to go. Um
but maybe we could get on to the next
topic for now, which is
best practices for OpenStack
observability. And this section is all
basically posted by uh Mark Schushlin
from Germany. Yeah, that's right.
So, why don't you Do you want to just
give us a
you know, a quick summary of what you
have here and then your questions?
I'm I'm currently in progress in
developing
observability stack for
um staff and OpenStack environments as a
open source product.
Um and I'm pretty interested what are
the best practice patterns to uh do
observability from logging perspective
or also from a metrics perspective for
larger um
OpenStack installations.
So, larger
beyond 500 nodes or much larger.
Um
Yeah, I'm interested in your experiences
or in in your hints to
go in the right direction.
>> [snorts]
>> What are you using for observability?
And how do you do it in an efficient
way?
I I can answer the question, but it's
not helpful, I'm afraid. It For
Bloomberg, we have large clusters. Um
They're some up to um 1,200 nodes.
Um and we had many false starts with um
metrics collection and observability.
Eventually, um an an independent um
sort of metrics product within Bloomberg
launched,
which covers OpenStack, but also the
physical machines, which we still have a
lot of
um
So, we have um
metrics being sent up with I think it's
uh um
sorry, Telegraf.
Um and then it uses uh Graphite and
Grafana, so that's a web page, web
portal, and you can build build
dashboards.
Um but we don't None of that is open
source as you know, as a as a portal
that you could just download a whole
coherent thing and use it with
OpenStack.
Um but we are moving to open telemetry,
I hear, which is good. That's um the hot
new new thing, I think, in that space.
And some of the some of the uh people
that are on top of our cloud, like for
instance, the Kubernetes managed people,
are using Prometheus.
Um
So, that's a Google open source thing.
Uh I am not personally Yeah, I I I I'm
just a user of these things.
Um
So, I think other people here maybe
could could share more actual useful,
you know, uh
Our current base setup relies on
Zabbix as a coordinating instance for
for observability of or for for
collecting data,
uh collecting high-level data. And we
are um using Mimir and Prometheus to
collect data. So, my my question goes
more in a direction. We also we run our
OpenStack environment with the Yaook
operator. It's operator for Kubernetes,
which runs the services uh
in Kubernetes.
>> [snorts]
>> And so, I'm pretty interested how can we
get
Yeah,
really good structured logs um that that
is very simple for for the Oslo layer,
but if we look at other areas like
Galera, OVN, vSwitch, and OVSDB, and
RabbitMQ,
um it seems that there's a lot of
engineering needed to to have structured
logging
and also to monitor that. So, yeah.
That's my situation.
>> [laughter]
>> Uh [snorts] I get it. So, one thing that
um is an idea forming in my head that
the experts on this actually are
colleagues of mine. Um
And one of the things that we've had to
do is actually im- improve the data
being captured from hypervisors
so that we can actually manage our
product. So, as a specific example,
um we started to look at
a metric that is reported from the
machines
the VMs called CPU steal. So, if you get
CPU steal, it means that the VM didn't
get the cycles from the CPU that it was
expecting.
And we actually try and move load around
to minimize that. Um
so,
um
we got CPU steal from Linux straight
away. Um finding the um
equivalent metric for Windows VMs was
more difficult.
I think we're actually solving that now
by getting data more from the
uh hypervisor
the the actual
KVM layer
than inside the VMs. So, if this is of
interest, I might actually be able to
persuade an expert in that subject area
to join one of these future things, but
I I'd kind of want to to hear that was
of interest. So, we're we're now um
looking at I don't know if it's deployed
today, but looking at being able to see
when Windows VMs have also suffered uh
CPU steal.
Um
You know, because then we we actually
have automated alarms, and if we get
more than a certain amount of CPU steal
on one hypervisor, we will actually move
some of the VMs off to a lightly loaded
hypervisor so that, you know, everybody
gets all the cycles that they
um
are owed in some sense.
So, um
if that's of interest, um you know, uh
put a note here or or
uh post to the Matrix chat and I can
invite them to come to some future
thing. I can't say when, obviously. The
person who
I think would be best at this is
actually on vacation in Europe right
now, so I certainly can't uh ask him
right now.
But, um
you know, so
um let me know. I I did have um another
colleague
uh last time talk about our high
availability control plane um
engineering.
Um I think that was that went rather
well, so um yeah, just
for that in particular, let me know
um you know, if if you would be
interested in in joining radio Ops radio
for that session.
Mhm.
So, does anyone else have responses to
Mark? You know, he's got a lot of
questions here and it's a very wide open
topic and um I
unfortunately, I'm not really best
placed to answer it for even for my
organization, but uh
I'm sure everyone here has some um
needs or experience with this.
One One thing I experienced with um the
standard Prometheus mechanisms to or to
observe um things in in OpenStack, um
many of them are using the API. And if
you have a large environment, the API is
uh the API workload is significant.
And so,
I I think uh
we we probably need to to write
something which collects the data
directly from the databases.
But, probably there are also
ready-to-use
uh mechanisms Prometheus uh
>> [gasps and sighs]
>> things which can be reused. So, uh if
you have any hints regarding that, uh if
I have good experience with something,
then let me know.
I can uh I can share a
a hilarious misstep that we did. So,
about 10 years ago when we first brought
up an OpenStack
Mhm. installation,
the monitoring was all on OpenStack
running in OpenStack. So, for instance,
all of the metrics went via RabbitMQ.
Mhm.
It all seemed very
um orthodox, should we say? But, what
actually happened is the first time that
the cluster ran into problems, the the
the monitoring we kind of went crazy and
then crushed rabbit and crushed the
database. So, the the first the first
little wiggle, and the whole thing just
just crashed into the earth at high
speed.
So,
I think that's it could have been line
with what you were saying that that um
you know, it can be a lot of data
hitting APIs and maybe you need to
get it out
quickly to some other [clears throat]
system. And that is effectively what we
do now. We're just users of a
centralized um
metrics portal
and then um a website with um you know,
dashboards and user creatable dashboards
using Grafana.
Um and that's been a lot better cuz
there's a whole another team looking
after the capacity of that, the uptime
of that, the upgrades of that. So, uh
we're actually fortunate. I know that
this is not an answer
for those assembled here if they don't
have that, but that's why I can't really
tell you a lot about how it works.
>> [snorts]
>> But, um you definitely don't want your
uh
you know, to be blind to what's going on
at the first sign of trouble, right?
So, if your cluster's struggling
and your observability depends on it
working perfectly, then
you've got a sort of logical flaw there.
Yeah, of course.
Yeah. Okay. Okay.
Uh so, it sounds like that's uh
a live topic that just doesn't have a
lot of engagement in this meeting. Um
maybe we can um
Mark, maybe we can sort of drum up more
interest in that. Um I think it sounds
like if you were to
uh put post that as something that we
would like to talk about in future, we
could get people together who know more
about it than certainly I do.
Um
One of the things that um we're still
working on is to actually get the word
out about Ops radio. Um I I actually
post on Fosterdon in an account called
Op- Ops meetups.
Uh we used to have in-person Ops meetups
uh back when we had um when we used when
Twitter was acceptable.
So, I moved it over to Fosterdon, but uh
Mast- Mastodon hasn't been hasn't been a
big uptake. The Matrix channel seems
another thing and then IRC and uh the
mailing list.
Um
so, you know, um
if you have a topic that you want
to talk about at one of these things,
use those channels where your colleagues
or people you know in the community
uh hang out. I mean, I hear that there's
also Reddit, there's LinkedIn, right?
There's
there's a Slack channels for some
regions. I think Latin America uh they
have a Slack channel for
OpenStack in in South America or Mexico.
So,
uh there's so many that it's hard to get
the word out, but but do that and you
know, this is a one kind of place that
we can make kind of rendezvous.
Um
trying to be a good uh meeting
uh runner. Um there is another couple of
topics, so
um
and by the way, um I'm not cutting off
the discussion of the previous things.
If you have something you wanted to say,
just speak up, but I'll try and go into
the gap here. So, um
the Red Hat's OpenStack RDO, uh they
need volunteers to
continue the production of RPM packages
because I think that by default, they're
only going to be doing container images
for OpenStack going forward unless
volunteers take that up.
So, um
this this link um re-shares that
message. And then there's a list of
volunteers here. I'm not going to click
through it. I don't want to risk getting
out of my
me- uh me pad.
>> [snorts]
>> Um so, if you're a user of RDO uh and
you were you consume OpenStack via RPM,
then but you know,
that's a community need. Otherwise,
that's going to dry up and then might be
out of luck. So,
it's kind of a call to action if you're
one of those. Go ahead. Uh question on
this. So, you don't have to talk all the
time. I do appreciate you running this
meeting very much. Uh
what are the next steps on this RDO? I
think Jose has been working like in
those calls. So, what
as normal, so is there any next step
next actions that need to be taken here
or how will this go forward?
So, the the only thing I
I know, Calle, is from um that that
etherpad, they're men- mentioning that
they're waiting for uh
I think it was like a PTO discussion in
that to be done in the PTO. And then
after that, they will come back uh in
the in the uh
in the ether- in uh
I think it's like via mail to the
mailing list. This is what I what I
understood. Like they're waiting on
someone to get back from PTO before
scheduling.
The The thing is, yes, we are heavily I
mean, we are using RDO.
And we have a on our side, we have like
two problems. One is the the
uh the infra and the other one is the
client side.
And we may need to kind of have
different approaches for for both. This
is for the uh mostly what we are
discussing there is the infra. For the
client side, what they what we're
proposing is like to have them in Fedora
uh packaging in um yeah, in Fedora which
we have kind of maintainers at CERN that
they basically can can help us with
that.
But then for the kind of the the infra
of the server distro, it's like there's
many many many more packages to to
build. And we're basically waiting for
for Amy to discuss it or or continue the
discussion.
Okay, so uh hurry up and wait. But, also
we are also interested in this because
we are also running RDO.
Yeah, I know. The The thing is,
this the So, if if you look at the uh
Epoxy release, it was 250 packages plus
400 dependencies. This is like quite
some effort
just to bring it up.
And uh yeah, and probably what you may
have seen in the mailing list, I was
like uh
not complaining, but it's like we just
got into Epoxy
and then we just discovered this.
Um
Now, this is great. Um
I'm thanks for
waking up this discussion. Um
I feel like I'm just doing a disclaimer
on every topic. Uh we do consume
packages, but we use um Ubuntu's cloud
archive distribution.
Um
we talk about containers regularly, but
we've never moved our actual OpenStack
into
containerized. Um
I think it's partially because um we
have very um
onerous
requirements on uh security.
And um there's exploits of containers
and there's, you know,
worries about it and there's debate
about it, so
getting packages from the vendor you're
paying and installing them is just a
little simpler
from that point of view. I know it's
more
cumbersome in other ways, so that's
that's where we are. So, we
are consumers of packages, but not we've
never used RDO.
Um but there's some good updates here.
Um the next step is apparently waiting
for the
um
PTO.
Um
And by the way, um
I saw a great update, I think from Mark,
further up. Um he's got a link here to
uh tech talk about observability at the
Alaska event.
Um
So, if you're interested in his
questions, um take a look at the tech
talk.
Um
I don't know if people can see what I'm
seeing. I've just highlighted that that
line.
Um
Okay, so we have
less than 10 minutes left. Um
the last item on the agenda um is from
Ildiko. Great reminder. We actually
don't have a fixed date for the next
session until we
pick one.
And the most obvious one
for a month's time would happen to be on
a PTG day.
Um and so her suggestion is
should we make it an actual literal PTG
session so that I guess it'll be in
people's PTG calendar.
I think that's a great idea personally.
The other alternatives would be
to maybe move it one week earlier or one
week later.
Um but the last Friday in the month
seems um
to be working. It's easy to remember.
So um
I I'm actually I didn't plus one this,
but
I think that's a great idea. Um
what are people's opinions about when we
should do the next one?
Um before we move on, what is PTG PTG?
Project Technical Gathering, so
it's when the dev teams get together.
Um it used to be in person, but now it's
um
entirely virtual and it's and so
therefore it's it's basically free and
open. You so you can
you can join, you know, the Nova
development group and listen to them
talk about new features for Nova compute
if you wish.
Um but, you know, it's an event with a
certain timetable. And if we were to
make Ops Radio one of those, then it
would kind of hold that time slot
for people to jump on MeetPad here and
discuss things
like we did today.
Uh
whereas the rest of the day might be
other things within PTG. So
I think that's um that's a good idea
because otherwise it might well be that
somebody wanted to come to
this, but it's something else got
scheduled over the top of it.
And that creates, you know, difficulties
like, you know, I I I I wanted to go to
the radio hour, but
you know, uh
my team scheduled something else over
it, right? So um
I so far I think that that's the
uh a good suggestion without anyone
um
speaking out against that. Um
is Ildiko still on? Yes, can you hear
me?
I can hear you. Okay, then my headset
works, too.
So um another advantage can be if we
schedule the next call during the PTG is
that the OpenStack project teams will be
around
for discussions more that week because
of this online event. So people are
preparing to make themselves available.
So if
we can put together some topics
discussion topics for the next meeting a
little bit in advance, then we can also
use this opportunities to raise some
awareness and maybe try to see if there
any topics that would be candidates to
try to get some
cross team collaboration discussion kind
of um
activity happening. So it it might be
easier to find people in the active
developer community
to weigh in on some of the some of the
topics and discussions if
that would be beneficial for anything.
And on the other hand, as Chrissy were
already mentioning, this will be in the
in the PTG schedule, so people will be
able to also plan better with uh with
overlaps.
And it sounds like that there is a
consensus or at least no objections.
And um the PTG is getting into the final
stages of the um organizing the event,
so I will go ahead and make sure
that the um Ops Radio hour is registered
uh for the event, and then we'll book
this time slot once the booking window
opens.
>> [snorts]
>> Sounds good to me. Uh one question,
Ildiko, just operationally, um if that
was done, would we still be using this
interface, you know, MeetPad and
EtherPad, or is it Yes. done on you
know, on Zoom or Okay. No, so the uh the
good thing about the PTG is that the
tooling is
um configurable, so it's very flexible.
So what I'm going to do once the group
is added to the tooling as kind of a
participating team on the administrative
side, then I can go ahead and register
this exact MeetPad link as the
conference tool that we are planning to
use for the call, and I will also update
the EtherPad link because the tooling
generates an EtherPad automatically, but
it's we can override that, so I will do
that as well. So the Ops Radio hour
EtherPad will show up in the list of
EtherPads for the event for this group.
So
>> you think we'll be able to get a
similarish time slot to this one?
Yes, the
the PTG has big time blocks. We have
some breaks built in, but the the time
one of them starts at 1300 UTC. So we
will be able to get this a
from
from the perspective of the attendees of
this meeting, if you're not joining it
through the the PTG event page, then
joining the call will be the exact same
process of coming to this MeetPad uh
link and the same EtherPad link at the
same time slot. So we can make all that
happen.
And it is kind of an added advantage to
have the call scheduled on the PTG
agenda as well that that OpenStack
project teams will see and will be able
to check the agenda and see if there's
any topics that they might be able to
add to or would want to participate and
those kind of things.
Fantastic. I [snorts] don't see any
downtime. Um so just on a personal note,
I have to finish my participation in
this call
absolutely on the dot or maybe ideally a
minute or two before because I'm meeting
with my boss next and I don't get much
time with him.
Um so I'm going to say last call. Um I
personally think that the idea of PTG is
great, sounds like no downside, and
there's
all all plus ones there, so I think
that's motion carried.
Um
so if any final calls, um we'll we'll
pencil in April 24th.
Uh people should add things to the
agenda for next time. Uh we just add at
the top and push the old material down.
And um
I guess we'll see some of you in a
month's time.
Any final comments from anybody?
Okay.
Uh well, thanks everyone, this was
great. Um please do get the word out.
You can put topics that you would like
to discuss or that you would like
someone who knows about a topic to talk
about, and we do try and get people
accordingly. Um so for instance, I might
be able to get our expert in
observability to join
if there's interest. But for now I think
we are done. I'm going to declare
meeting over.
Uh Ildiko, if you can stop the
recording, and then we'll see you next
time.
Thank you, everyone. Thanks, Chris, and
also see you all in the Matrix room in
the meantime. Thanks.
Thank you. Thank you.