David Kanter, ML Commons | theCUBE + NYSE Wired: Mixture of Experts
Watch on YouTubeVideo summary
David Kanter, co-founder of ML Commons and head of MLPerf, joins the discussion to explain how his organization evolved from a non-profit initiative into a critical standard-setter for the artificial intelligence industry. Before the rise of generative AI, there was no unified way to measure the speed, energy efficiency, or reliability of machine learning models, creating a chaotic market where buyers struggled to compare different technologies. To solve this, MLPerf was established as a consensus-driven benchmark that allows enterprises to make informed infrastructure decisions, much like standardized car ratings help consumers choose vehicles based on objective metrics rather than marketing claims. This foundation has since expanded to address the growing concerns around safety and risk, ensuring that AI outputs align with societal values while maintaining high performance standards across diverse applications.
As the industry shifts from simple inference tasks to complex agentic workflows involving coding, customer support, and multi-step reasoning, the focus of ML Commons is adapting to these new realities. Kanter emphasizes that modern AI deployment requires a nuanced understanding of trade-offs between speed, accuracy, and cost, noting that different use cases demand different balances of compute resources. For instance, a front-line application needing instant answers for customers requires a different infrastructure strategy than a back-office system focused on deep analysis. The organization is currently developing new benchmarks to capture these specific scenarios, moving beyond raw hardware speed to evaluate how quickly systems can reach the right answer under varying conditions, effectively helping CIOs and CISOs navigate the complexity of integrating AI into existing enterprise ecosystems.
A core principle driving ML Commons is its commitment to openness and transparency, operating as a member-funded non-profit with representation from over 125 organizations across six continents. Unlike traditional hardware-focused benchmarks that prioritize raw processing power, this community-driven approach ensures that standards remain current, comparable, comprehensive, and contextualized for the rapidly evolving software landscape. By fostering consensus among industry leaders, academia, and individual contributors, the organization builds trust and avoids the pitfalls of proprietary agendas, ensuring that the benefits of AI are accessible to everyone. This collaborative model allows the community to collectively define what "better" means for society, whether in medical diagnostics, autonomous driving, or creative tools, ultimately guiding the industry toward responsible innovation.
Read the full video transcript
Palo Alto studio connecting Silicon
Valley and Wall Street.
>> I'm John Furrier host of the Cube here
with Dave Vellante my co-host.
Hello, I'm John Furrier host of the Cube
here at the Cube's NYSE studio. Of
course, we have our Palo Alto studio
connecting Silicon Valley to Wall Street
part of our NYSE wired program and
community. This is our mixture of
experts series where we bring in people
who are experts in their field doing
great work and innovating. David Canter
here he's the co-founder of ML Commons
and head of ML Perf. If you know all
about machine learning, you know what
that organization has done
a non-profit doing really amazing work
helping people figure out what's safe,
what's real, what's not. David great to
see you. Thanks for coming on the Cube.
Saw you at AMD's event in San Francisco.
>> Absolutely a pleasure. It's a great to
be here. You get to be mixed up with all
the other experts.
>> Yeah, yeah, it's it's it's kind of a
initially it was kind of a goof on AI
when we did the series, but it's
actually great way to bring in our
community and kind of mixture of experts
kind of share in the data.
Um and you know, one of the things that
everyone loves about AI is that it's got
a great utility but pre
you know, the transformer technology
machine learning has been around for a
long long time. You know, fraud
detection every bank has it
um supervised unsupervised machine
learning. It really was the genesis
of what got all the deep tech nerds in
the labs.
You know, we're looking at what's coming
out and then the bill just went
supernova from there. So, it's really a
valuable
organization and lesson and also a
template
for the future. Explain what you're
doing at ML Commons and ML Perf. How it
all came together. What is it that
people don't know what it is and what it
does and kind of where it is.
>> Absolutely. So, we got started
2018 back [snorts]
you know, pre chat GPT pre generative
AI.
And everyone was looking at we knew
these AI models could do incredible
things. We wanted to improve performance
to improve capabilities, but there was
no standard way of measuring things. And
so a group of us all came together from
industry and academia to build that
standard set of benchmarks to measure
speed and energy efficiency. And and
that became MLPerf.
And you know, before that
>> And by the way, that then became the
cited benchmark stat in every
presentation at that time.
>> Exactly, right. And you know,
part of the the thing that's really
wonderful about getting to be involved
in this group is we bring together
everyone from all across the industry.
And through consensus, we build these
trusted standards like MLPerf. And you
know, at the time it would almost be as
if you were buying a car and and one guy
says, "Hey, my car can do zero to 60 in
a second." The next guy says, "My car
has a turn signal." And the third guy
says, "I've got airbags." Which one do
you want to buy? You don't you don't
really know. And so
you need some way to compare them all. I
mean, I live in San Francisco, so I
might go for the airbags.
But
that was the genesis of MLPerf. And then
we realized
sort of
the impact you can have to help drive
the whole industry. And we said we
should put this into a nonprofit and
then look for other ways that we can
deploy our expertise in measurement and
data to help make AI better, right? And
so after we first did performance, we
then zeroed in on
measuring power efficiency,
building large open data sets, and then
over time we've started looking at
benchmarks in risk and reliability of
helping to make sure that, you know, the
outputs of generative models are kind of
in line with what we want.
>> The evolution of AI now is the number
one conversation safety, right? And then
see the Anthropic versus say OpenAI
approach fast and loose, more
>> [music]
>> conservative. There's a general There's
a lot lot of consensus around no one
really knows what the hell that means.
So, take us through kind of where you
guys are focused on now because you guys
have the playbook on open. We see the
the success of open source. I mean,
damn, it's the most successful trend
ever in the computer industry. Look at
what it's done. Now you got open
weights. So, you got a lot of open
things happening. What are you guys
focused on now? How is this translating
into some of the conversations today?
>> Yeah, so I I'd say, you know, one of the
most critical things is when you're
looking at, you know, anything, whether
it's performance or risk or
responsibility, it's about measuring it
in the right way.
You know, having written down what
you're doing, what you're trying to
accomplish, how much precision you have.
And so, for us, you know, one of the
things that we released last week was
MLPerf which version .7, which is a
rethinking of our inference benchmarks
for the modern era, right? And we see
this uh
just race to deploy. As everyone's
discovered that there's so much valuable
things you can do with AI, how do we
deploy it across the enterprise for
consumers? We see building more data
centers, needing more power, and more
performance systems. So, we had to
evolve our benchmarks
to
uh match that pace of innovation and to
help customers really make the decisions
that they need. You know, you look at a
Fortune 500 company,
>> Yeah.
>> they're not just saying, "Hey, I want AI
for the C-suite." They're saying, "I
have dozens or hundreds of applications
that I'm going to deploy. Each one's
different. You know, how do I find the
infrastructure that's going to pair up
in the right way. And so, I look at that
as being our job. Is how do we build the
tools that help empower those folks to
make the decisions to help deploy AI.
>> So, you're really kind of taking the DNA
of MLPerf, MLCommons, machine learning,
yeah, applying that to the AI growth
wave, which is infrastructure,
and trying to help people navigate that.
So, I love that. The question that we're
seeing now is I just had a I just wrote
a post when lunch I had to write a post,
but on agents. And a lot of the fear
with agents is you have more
deterministic workloads. Again, that's
cool. Um but, you have workloads, and
they're different. So, now you have
different conditions.
>> Mhm.
>> What's What's the scope of some of the
things you guys are getting your arms
around in the open because you know, a
decision for company A will be different
than company B because I might want to
have more compute,
>> Mhm.
>> less GPU, or
prefill, decode. Like, all these things
are kind of now coming into the systems.
It's not a clear general purpose
benchmarking market. And so, how do you
guys think about that? What's the
community doing on? Is there any data
you can share, thoughts,
personal thoughts?
>> Uh just thoughts for sure. And and you
know, stay tuned. We'll have data later
this year for sure. But, I think one of
the things you you really touched on is
when we see sort of blending
AI inference with standard computing
workloads through agentic flows. Right,
the world is your oyster. Right, before
it was like, oh, maybe you're doing
recommendation or translation. Well, now
you might be pairing that with
sort of hey, this is the conventional
workflow that I have in my bank, but now
I'm going to stick AI in here to
accomplish my goal. Or, you know, of
course, the thing that we've seen the
most demand for is coding.
>> Yeah.
>> Right. And you know, it sort of makes
sense. The folks who are developing the
tools are like, "Hey, wait, I can do
what with this? Like, let's let's get
some acceleration." Hugging Face, hello,
you know, test bed went off the rail.
Again, there's so many use cases where
it could go
off.
>> Yeah. That may or may not be related to
anything other than the environment.
>> That's right. And one of the core
insights in MLPerf actually was it's not
just about speed, but it's about how
fast you get to the right answer. And to
the point you made earlier,
you could achieve the same task with a
lot of accelerated inference compute and
maybe less conventional compute, or
maybe a different balance. And it it
depends on what you have. You know,
imagine you want to pick like what's the
best South Indian restaurant in
Manhattan. You could look at the top 10
if you're really confident of that top
10, or maybe you've got a system that's
even more accurate and you just say, "I
I really need the top three from this
system, so I don't need to read all
those reviews."
Both of those will hopefully get me to a
delicious meal, but what's the right
path? That's really tricky. And so we're
just in the starting
>> Yeah.
>> ages.
>> That's a really good point. I mean, I
think what you just said was compelling
because most people look at the user
experience of say ChatGPT, which is most
consumer experience, as getting a good
answer fast. The first answer. Kind of a
search results. Hey, that you know,
great. Where do I find food? Boom,
answers.
Reasoning is a different It's not a
search paradigm. You're doing discovery
>> That's right.
>> with multi-step reasoning.
>> That's exactly right.
>> different. So, it's a very nuanced
point, but it changes the configuration
of the data,
what
systems I might want to use. Maybe it's
a complex answer. Maybe you have certain
Indian food needs that might require a
Pareto curve of the Vera Rubin. Right?
Who knows?
>> [laughter]
>> Ching on the tokens. So, like this is a
cost trade-off. This is a math equation.
>> That's exactly right. And so part of our
goal with MLPerf points is how do we
present that trade-off so that if you're
a CIO or CISO, you can say, "All right,
I've got all these applications. These
need to be really fast. These are for
front line. I need an answer for the
customer quick. These other things maybe
back office are going to look
different." And so how can you get the
right infrastructure for everything? But
I think to your point, the world of
agentic,
you know, before you sort of had AI
walled off and it was its own thing. You
had separate AI infrastructure people.
But now,
you know, one of the things I used to
say is to me inference is a lot like
salt in cooking. You don't eat pure salt
usually.
But
if you know from any recipe book, you
add a little bit of salt and it makes
almost everything taste better. And
that's what we're seeing. So now it's
not AI's here and regular computing
here, it's all braided together and all
throughout. So, you know, almost
everything that we have today in the
enterprise is going to be recast with an
agentic side. And so we're still in the
early stages.
And you know, when I think about what we
want to do, we have an agentic benchmark
coming out later this year
focusing on some of the things that we
think are most popular,
code development and software
engineering as well as
you know, sort of Q&A and customer
support. But I wouldn't be surprised if
we have dozens of use cases in the
future as we're discovering them live.
>> when we were chatting at the AMD event
in San Francisco, we were talking about
some of the historical views. We've
lived through so many cycles
of innovation. This one's actually the
most kick-ass ever because it's got
everything popping. You got
infrastructure up and down the stack.
But in the old days when we were I was
breaking in the business at
Hewlett-Packard and before that IBM, PCs
and servers, they all had the
benchmarks. And that was really twofold.
One to do an industry service to kind of
level the playing field on, you know,
horses on the track, apples and oranges,
making sure everyone knows what's what.
But it also helped customers scope
what they wanted to buy. So it was
really an economic
beacon, too, for the okay, I need a
mid-range system for these desktops or
whatever.
When you get to AI, a lot of that's kind
of going on. It's not as simple. What's
the biggest change in your mind today
trying to
rally the industry around ML Commons
while looking at the the the aperture of
use cases.
Um you guys look at that as an
opportunity. Do you Is there some first
principles and then playbook tactics you
guys are using? Because everyone wants
the same thing. What do I buy for what
and when? I don't want to waste any
money. I don't want GPU cycles wasted. I
want to use the right token,
expensive tokens for the right
models
and let people do their job.
>> So, I would say, you know, one of the
When I look at what's changed, right?
It's we've got a much broader audience
and the rate of evolution is just
incredible. And so I think for us, that
resolves down to we have to
shift from the speed of hardware,
because the truth is MLPerf was started
by hardware folks, to the speed of
software, right? Like, you know, you're
used to your apps on your phone getting
updated every week or so. You know, and
it
>> Not when my battery's low, though.
>> Right, [laughter] of course. Of course.
But you know, my head of marketing drew
out this great chart and if you look at
sort of leading edge frontier labs and
capabilities, they're adding a new thing
every 2 weeks. And so we have to shift
to that sort of speed and get results
and benchmarks that are going to stay
current.
We have to have benchmarks that are
comprehensive, that really map out all
the options a buyer is going to look at,
and compare them in sensible economic
terms and then contextualize them,
right? You know, it's because it's not
just the AI infrastructure people
anymore. It's it's everyone.
>> yeah.
>> That contextualization is a big
challenge. So, we call that sort of the
four C challenge. We want things to be
current,
>> Yeah.
>> comparable,
uh
and comprehensive, and then
contextualized. And of course, you know,
you're part of that contextualization as
well. You help
>> Yeah.
>> illuminate the path for AI for many
people.
>> Well, a lot of people want to know
what's on the road map. Everyone wants
to connect the dots. And and and again,
back to the open piece. I think that is
really the most uh important because you
guys were grounded in the early days of
AI, you know, cloud native, Linux
Foundation, uh CNCF. Again, a unique
approach flowered up some nice benefits
with cloud native. So, the open source
equation is key to success. In a way,
you're AI commons now. I mean you mean
see you know, ML's kind of like too
small in my mind cuz you're everything.
You're helping everybody. You could call
it AI in commons, physical AI common. I
mean basically, it's
>> You you should be careful because we
might bring you in for renaming
[laughter] you.
>> No, no, but it's it's broad and you're
doing great work. Um and you also your
member funded. So, explain that piece.
This is not like you guys have a
particular agenda. Talk about the scope
and the mission cuz I think this is also
an important balancing piece.
>> Yeah, no, that that's exactly right. So,
you know, we are a member-driven
organization. We have uh over 125
members on six out of seven continents.
We're
still waiting for some folks in the
Antarctic to sign up, but one day. One
day.
>> [laughter]
>> Uh and you know, it's it's drawn from
all pieces
of the AI industry, and it's really
focused on how can we make AI better for
everyone through measurement?
>> Yeah.
>> Uh and
it's those members
that that fund us, right? And so, we're
a nonprofit. There's a degree of
transparency, of good governance that
lets us do great work and that gives
everyone the trust and the knowledge
that we can pull our members expertise
to help shed light on what's going on
and hopefully drive us down a path where
AI is just going to do tremendous good.
You know, I think the applications in
medical or
you know, you've taken a Waymo in San
Francisco.
>> Yeah.
>> It's almost magical, right?
>> awesome. And and if you look at the
applications to the to the human
society, again back to the original
mission of MLCommons and MLPerf was to
make it better for people.
>> That is what people want today in AI.
>> Um share what the activities are like
for a member. What goes on behind the
curtain? Uh every kind of project open
group has kind of different you know,
paths. What's it like? How do you guys
engage? How do you get consensus? Take
us through some of the sauces making and
some of the behind the curtain.
>> Yeah, so the first thing is, you know,
membership is open to everyone. We have
individual members, we have academics.
You know, if you're interested, you can
come get involved and help guide what
we're building. And so for a lot of the
members,
you know, they'll have representatives
that show up and say, you know, here is
something that we think is important
we'd like to incorporate or
yes, you've got a plan. The plan is 90%
right, but if we tweak it just this way,
it'll help us
include our solutions in here and so we
can make the benchmarks broader and more
>> more intentional in your focus than say
let a thousand flower blooms may the
best projects win. Which is kind of like
a Linux Foundation. That's and that
works for them.
>> Yeah.
>> But is that the same? You guys have a
different approach? How would you
compare?
>> we're we you know, the Linux Foundation
is in a lot of ways, you know, sort of
one of the original
architects of this kind of an
organization. And they're they're
absolutely huge and and have a huge
number of projects. I think we're a lot
more focused on we want to be doing
things where we're uniquely suited. And
so, is it relevant to AI? Does
measurement expertise play in? Does data
play in? There's, you know, a set of
things that we're really good at and
align with what we do. And we want to
focus on that. You know,
people ask me all the time, do you want
to host an open source project? And I
say, you know,
there's other folks who are uh
far more expertise in doing that. You
know, the Linux Foundation's been doing
it for 15 years.
>> are really focusing on your knitting
what you where you kind of you came from
and where you got where you are and
where you're going.
>> All right. So, for people who want to
and get involved, what's the process?
>> Come to our website. You know, you can
sign up for a membership. A lot of our
groups are open to the public. Uh you
can just sort of sign up for those, but
you can uh engage with us on
uh on Twitter or X, LinkedIn. I think we
have a YouTube channel. Uh and, you
know, we'll often be at conferences
presenting what we're doing. You know,
we're a really friendly bunch. So, you
know, if you have some great ideas,
sign up and uh say hi.
>> Well, we love what you guys have done in
the past, set the table. You know,
that's the foundation now. The world all
wants what is going on. AI factories are
super hot. It's It's the computing
industry revolution again, but it's the
same game, but it looks different, more
dense. A lot of subsystems are involved.
They're closer together, kind of like
the old days, but bigger and bigger and
better. Rack scale systems. Edge is
coming super fast. Latency. All these
factors.
>> Yeah, I mean, it's,
you know, a recasting of all of our
compute infrastructure in a wildly
different way. And it's It's both
exciting and a little bit terrifying to
be, you know, in the eye of the storm,
right? You know, I I tell my members,
you know, what we do is through
consensus, right? And so, there is a lot
of uh negotiation over what we should be
doing. And consensus is deliberate, it's
slow, it builds trust.
But uh at the same time, you know,
everything's Yeah, everything's moving
at a million miles an hour. And so, you
know, it's But it is an absolute
pleasure and an honor to get to have
this role. And it's exciting to see
what's, you know, going to be coming out
in the next year.
>> Well, thanks for coming on sharing the
mission. Totally behind it. Open always
wins. I've been saying that for day one.
I mean, I've said in the queue probably
the most of anything. Governance gets a
lot of buzzwords these days in AI. But
but you know, open uh wins and that's
where innovation lives. Thanks for
coming on. Appreciate David. Thanks for
>> Thank you so much for your time.
>> I'm John Furrier. This is our mixture of
experts here where people share their
thoughts on the key issues facing what's
being built on the innovation of the AI
infrastructure and hot area, agents,
physical AI and robotics, defense tech,
all booming as part of this new AI
revolution. And again, people want to
know where things fit, where to buy it.
Doing our part here. Thanks for
watching.