Deploy and run Agentic AI applications on Intel Xeon CPUs with Red Hat AI Quickstarts
Watch on YouTubeVideo summary
Bu videoda Intel'in yapay zeka yazılımı çözümleri mühendisi Alex, Agentic AI uygulamalarının yalnızca Intel Xeon işlemcilerinde hızlı bir şekilde nasıl dağıtılacağını ve çalıştırılacağını gösteren bir demo sunmaktadır. Sunulan çözümün temelini oluşturan Red Hat AI Quick Start'lar, Helm kurulumu veya OC komutu gibi basit adımlarla uygulama deployment sürecini hızlandıran mükemmel bir şablon görevi görmektedir. Bu altyapı Llama Stack API'si üzerine inşa edilmiş olup, web arama ve diğer işlevleri destekleyen çeşitli MCP sunucuları ile özelleştirilebilmekte; ayrıca VLLM gibi model sunucusu desteğiyle yüksek veri akışı kapasitesine sahip sistemler kurulmasına olanak tanımaktadır.
Demo sırasında izleyiciye sunulduğu üzere, bu yapay zeka sanal asistanı finans hizmetleri, bankacılık ve seyahat gibi farklı sektörlerdeki kullanım senaryolarını karşılamak üzere tasarlanmış şablonları desteklemektedir. Özellikle bir seyahat uzmanı ajanının seçildiği sahne incelendiğinde, İtalya için yedi günlük bir rota planlaması talebi üzerine sistemin öncelikle MCP aracını kullanarak web üzerinden arama yaptığı ve ardından kullanıcıya Roma ile Venedik'i içeren detaylı bir program hazırladığı görülmektedir. Bu süreçte kullanılan model sunucusu (VLM), Xeon işlemcilerindeki AMX teknolojisinin gücüyle küçük dil modellerini yüksek performansla çalıştırarak, düşük bellek ayak izi ile spesifik ve doğru yanıtlar üretebilmektedir.
Video boyunca vurgulanan en önemli nokta, Agentic AI uygulamalarının başarısının arkasındaki anahtarın Intel Xeon işlemcilerinde bulunan AMX teknolojisi olmasıdır; bu teknoloji sayesinde kullanıcılar karmaşık optimizasyon işlemlerine girmeden modellerini yüksek verimlilikte çalıştırabilmektedir. Red Hat OpenShift üzerinde çalışan Xeon tabanlı altyapı, sadece sunulan sanal asistanla sınırlı kalmayıp VLM araç çağırma, RAG ürün önerici sistemleri ve yeni eklenen gizli yapay zeka çıkarımı gibi çeşitli hızlı başlangıç çözümlerini de kapsamaktadır. Sonuç olarak bu çözüm seti, işletmelerin kendi özel bilgi tabanlarını oluşturarak müşterilerine sektöre özgü sorulara net yanıtlar verebilen esnek ve güçlü Agentic AI sistemleri kurmalarını sağlamaktadır.
Read the full video transcript
full Aentech AI on Xeon CPUs. Hi, my
name is Alex. I'm an AI software
solutions engineer at Intel. And today
I'm going to show you a demo on how you
can quickly get started with running
Agentic AI solely on Xeon.
What we have here is a Red Hat AI quick
start. It is a essentially a blueprint
where you can quickly deploy an
application with a simple Helm install
or OC command. Let me show you how it
works.
It's built with the llama stack API and
allows you to customize it with
different MCP servers including web
search. It also supports model server in
such as VLLM.
We at Intel do a lot of performance
optimizations to leverage to bring AMX
and upstream it to VLM so that you as a
user can deploy your models with high
throughput and low memory footprint
without you having to do the hard work
yourself. And finally, this solution
supports vector databases, allowing you
to customize a knowledge base so that
your virtual agents can give specific
answers to your customers depending on
the business use case. Now, let me show
you the demo in action. The AI virtual
agent supports different use cases. Here
you can see that there are different
templates for a set of virtual agents
ranging from financial services and
banking uh and also travel.
Now I'm going to navigate into some of
the agents to see how we can interact
with it.
Now, very shortly, I'm going to select
from the list where I'm going to choose
the travel agent specialist, and I'm
going to use one of the default prompts
and ask it to create a 7-day itinerary
for Italy. So, let's see what it gives
us.
You can see that it immediately calls
the MCP tool to do a web search before
giving us the response. Now, this is
using the VLM server that I was
mentioning earlier.
tells us that, hey, we're gonna first
two days let's go to Rome and then the
next two days we're going to go to
Venice. You know, I personally haven't
been to Italy myself, but I think this
is pretty good to get started with.
All right, now to take away from this
entire demo,
AMX is what enables you to run small
language models on Zeon powering agentic
AI applications.
Additionally, you can get quickly get
started running with these Red Hat AI
quick starts on Red Hat Open Shift
running on Zeon.
It's not just limited to this AI virtual
agent where you can get the code here.
There are also additional quick starts
including VLM tool calling, rag product
recommener system, and a new one called
confidential AI inference. Thank you so
much for watching.