Video summary
Many enterprises are currently adopting a pragmatic approach by equipping developer workstations with GPUs to run local open-weight models, which serves as an excellent entry point for exploring AI capabilities. However, this strategy is not sufficient for long-term enterprise success because it lacks the necessary control and auditability required when deploying systems across individual end-user machines. Relying solely on decentralized hardware limits visibility into what users are doing with their resources, making it difficult to govern approved models or ensure compliance. In contrast, a dedicated cluster of AI resources offers superior cost controls, governance capabilities, and the flexibility to manage diverse workloads effectively without being restricted by the specific configurations of individual desktops.
Furthermore, the hardware landscape for artificial intelligence is evolving rapidly, with new chips emerging that provide inference capabilities even on standard CPUs, meaning organizations should not lock themselves into GPU-only solutions found only in personal devices. Beyond hardware considerations, enterprises must also plan for agentic systems and batch operations that require a controlled environment rather than relying on ad-hoc setups like cracked-lid laptops scattered across an office. Guaranteeing service level agreements (SLAs) is critical to prevent workflow disruptions when AI services go down or perform poorly, ensuring that routine processes can continue effectively using open-weight models without risking operational stability.
The path from initial experimentation to owning robust AI infrastructure involves building skills in multiple dimensions simultaneously rather than following a rigid, sequential order often seen in traditional migrations like VMware projects. While empowering developers with local tools is an immediate first step, it must happen alongside efforts by hardware and operations teams to learn how to set up Kubernetes for AI workloads, select appropriate models, and integrate new capabilities into the broader infrastructure. Companies that focus on building these intermediate skills early end up innovating faster because they do not wait until a future date to address operational control, automation, and process improvements; instead, they embed these competencies across their organization right from the start.
Ultimately, successful AI adoption requires a holistic strategy where delivery of infrastructure, model selection, and system integration occur in parallel rather than as isolated phases. Organizations should avoid getting tied up solely in deciding which specific large language models to use or how to configure individual harnesses before establishing the foundational skills needed to build scalable systems. By ensuring that all teams are developing the necessary expertise concurrently, enterprises can create a flexible infrastructure capable of adapting quickly to rapid AI innovation while maintaining strict control over their assets and data. This proactive approach ensures that businesses remain competitive and secure as they transition from local pilots to comprehensive, enterprise-grade AI operations.
Read the full video transcript
Many enterprises are starting by putting
GPUs in developer workstations and
running local open weight models, which
is a great way to get started. But why
is that a reasonable place to begin, but
not a complete long-term strategy?
>> Well, there's a couple of reasons why
not and and one of them is innovation
and pace of innovation. So, I want to I
want to get there and and put a pin on
that and how you do it. But from an
enterprise perspective, ultimately, it's
going to be about control and
auditability. So, what we're looking at
here is if you're running an enterprise
system and you want to have your AI
um you know, your your end your
individual end users, you said
developers, but it could be any end user
leveraging a local model, that is going
to save you on processing capabilities.
However, what you're going to do is be
limited to the what that end user is
doing on their machine, when their
machine is running, when you have access
to it. And it's going to be very hard to
audit and control and check. So, the
idea of having a cluster of AI resources
actually gives you better cost controls,
but it also gives you better
auditability and governance, lets you
manage those pieces. And this is where
the other the other dimension comes in.
Inference is does not necessarily
require a GPU. There's a lot of new
chips coming coming into market that
will provide
AI inference capability. And you know,
even running CPUs and CPUs are getting
new inference capabilities. So, the
you know, locking yourself into what's
on people's desktops, it's a good
short-term solution, but long-term,
having that that that dedicated AI
capability with a heterogeneous mix of
hardware, so you can match models to
capabilities and bring in new new models
whenever you need to and then control to
make sure that the models that you've
approved are getting used, those control
points are absolutely essential and
really from an enterprise perspective
critical to scale.
And it's it's worth noting when you look
at how those pieces go. And that's
exactly what we're helping companies
spec and build today. The harness allows
you to to pull those pieces in. You
guarantee an SLA, right? If you've been
We We get caught up in AI goes down, has
a day where it's not performing well,
and a lot of work gets stopped. Um
you're going to want to be able to
control even if you're not on a frontier
model. Having guaranteed access to
models to run processes is absolutely
essential because more routine processes
can be be run very effectively by open
weights models.
And along those lines, what you also
need to think about is in that
infrastructure and with those controls.
And there's one more critical point with
this,
which is that it's not just your users,
your developers, and people doing this.
Part of what you need to be thinking
about building is agentic systems, and
those agentic systems have to run
somewhere also. And you don't want a uh
room full of cracked-lid laptops
providing your agents. You actually want
them running in controlled corporate
infrastructure. And so, all of these
things work together
to drive this this realization that you
want to have a dedicated AI
infrastructure that can run your
inference models in a controlled way,
allow your agents to work, especially
offline or or to provide, you know,
batch operations in a controlled way.
So, all of these pieces fit together.
So, as much as I love running local
models and empowering developers, from
an enterprise perspective, you need to
be planning further out.
>> And now, what does the path from that
first step to owned AI infrastructure
actually looks like for an enterprise?
>> Well, this is one of those ones and
actually it's funny cuz VMware migration
is similar to this in our books and
experience. Is enterprises have a
tendency to try and sequence out all
these steps and they they think they
have to do this in a very orderly
manner. Our experience actually is that
the path for should have a lot of
parallel operations.
You know immediately, right, that you
can empower developers to do local
models and look at Harness. So, Harness
is is clearly a first step, but the
people who are going to make an
evaluation for your Harness, right, your
platform team, your developers,
aren't the same ones who actually are
the ones who should be looking at how to
run and inspect AI gear and
infrastructure. That's mostly hardware
teams and operation teams that need to
understand how to set up and run AI
gear, how to build Kubernetes from an AI
workload perspective. And what I would
what I would encourage, especially
because AI innovation is running so
quickly, is that you need to be building
these skills in multiple dimensions
simultaneously. So, you should be
looking on how do I improve my delivery
of AI infrastructure? How do I know what
to buy? How do I know how to wire it
together? How do I build these systems
and get that skill set embedded in your
organization. Uh we see exactly the same
thing going on with a lot of VMware
choices where people try to figure out
what their exit out of VMware should
look like and then work backwards. The
reality is the
all answers require you to have better
operational control of your
infrastructure, more automation, better
processes. This is what RackN does from
a bare metal perspective. What we've
seen is the companies who build the
platform and the capabilities that allow
them to then innovate on top of that
platform end up moving a lot faster than
the ones who figure out where they have
to get to at the end of the trip and
don't worry about, you know, any of the
intermediate skills they're going to
need to build.
Especially in today's market, you need
to be making sure you're jumping through
all the intermediate skills
simultaneously to build this up. So,
don't don't get tied up in I need to
figure out my harness. I have to figure
out which LM is best. I have to
You do need to do those things, but
don't wait
on I'm going to be building AI
infrastructure until after you've made
those decisions. Start building all the
skills across your teams so that they're
delivering the pieces that you need
right out of the gate.