Submind YouTube summaries
Thumbnail for Deploy and run Agentic AI applications on Intel Xeon CPUs with Red Hat AI Quickstarts

Deploy and run Agentic AI applications on Intel Xeon CPUs with Red Hat AI Quickstarts

Watch on YouTube

Video summary

Bu videoda Intel'in yapay zeka yazılımı çözümleri mühendisi Alex, Agentic AI uygulamalarının yalnızca Intel Xeon işlemcilerinde hızlı bir şekilde nasıl dağıtılacağını ve çalıştırılacağını gösteren bir demo sunmaktadır. Sunulan çözümün temelini oluşturan Red Hat AI Quick Start'lar, Helm kurulumu veya OC komutu gibi basit adımlarla uygulama deployment sürecini hızlandıran mükemmel bir şablon görevi görmektedir. Bu altyapı Llama Stack API'si üzerine inşa edilmiş olup, web arama ve diğer işlevleri destekleyen çeşitli MCP sunucuları ile özelleştirilebilmekte; ayrıca VLLM gibi model sunucusu desteğiyle yüksek veri akışı kapasitesine sahip sistemler kurulmasına olanak tanımaktadır. Demo sırasında izleyiciye sunulduğu üzere, bu yapay zeka sanal asistanı finans hizmetleri, bankacılık ve seyahat gibi farklı sektörlerdeki kullanım senaryolarını karşılamak üzere tasarlanmış şablonları desteklemektedir. Özellikle bir seyahat uzmanı ajanının seçildiği sahne incelendiğinde, İtalya için yedi günlük bir rota planlaması talebi üzerine sistemin öncelikle MCP aracını kullanarak web üzerinden arama yaptığı ve ardından kullanıcıya Roma ile Venedik'i içeren detaylı bir program hazırladığı görülmektedir. Bu süreçte kullanılan model sunucusu (VLM), Xeon işlemcilerindeki AMX teknolojisinin gücüyle küçük dil modellerini yüksek performansla çalıştırarak, düşük bellek ayak izi ile spesifik ve doğru yanıtlar üretebilmektedir. Video boyunca vurgulanan en önemli nokta, Agentic AI uygulamalarının başarısının arkasındaki anahtarın Intel Xeon işlemcilerinde bulunan AMX teknolojisi olmasıdır; bu teknoloji sayesinde kullanıcılar karmaşık optimizasyon işlemlerine girmeden modellerini yüksek verimlilikte çalıştırabilmektedir. Red Hat OpenShift üzerinde çalışan Xeon tabanlı altyapı, sadece sunulan sanal asistanla sınırlı kalmayıp VLM araç çağırma, RAG ürün önerici sistemleri ve yeni eklenen gizli yapay zeka çıkarımı gibi çeşitli hızlı başlangıç çözümlerini de kapsamaktadır. Sonuç olarak bu çözüm seti, işletmelerin kendi özel bilgi tabanlarını oluşturarak müşterilerine sektöre özgü sorulara net yanıtlar verebilen esnek ve güçlü Agentic AI sistemleri kurmalarını sağlamaktadır.
Read the full video transcript
full Aentech AI on Xeon CPUs. Hi, my name is Alex. I'm an AI software solutions engineer at Intel. And today I'm going to show you a demo on how you can quickly get started with running Agentic AI solely on Xeon. What we have here is a Red Hat AI quick start. It is a essentially a blueprint where you can quickly deploy an application with a simple Helm install or OC command. Let me show you how it works. It's built with the llama stack API and allows you to customize it with different MCP servers including web search. It also supports model server in such as VLLM. We at Intel do a lot of performance optimizations to leverage to bring AMX and upstream it to VLM so that you as a user can deploy your models with high throughput and low memory footprint without you having to do the hard work yourself. And finally, this solution supports vector databases, allowing you to customize a knowledge base so that your virtual agents can give specific answers to your customers depending on the business use case. Now, let me show you the demo in action. The AI virtual agent supports different use cases. Here you can see that there are different templates for a set of virtual agents ranging from financial services and banking uh and also travel. Now I'm going to navigate into some of the agents to see how we can interact with it. Now, very shortly, I'm going to select from the list where I'm going to choose the travel agent specialist, and I'm going to use one of the default prompts and ask it to create a 7-day itinerary for Italy. So, let's see what it gives us. You can see that it immediately calls the MCP tool to do a web search before giving us the response. Now, this is using the VLM server that I was mentioning earlier. tells us that, hey, we're gonna first two days let's go to Rome and then the next two days we're going to go to Venice. You know, I personally haven't been to Italy myself, but I think this is pretty good to get started with. All right, now to take away from this entire demo, AMX is what enables you to run small language models on Zeon powering agentic AI applications. Additionally, you can get quickly get started running with these Red Hat AI quick starts on Red Hat Open Shift running on Zeon. It's not just limited to this AI virtual agent where you can get the code here. There are also additional quick starts including VLM tool calling, rag product recommener system, and a new one called confidential AI inference. Thank you so much for watching.