Video summary
The video demonstrates the capabilities of the Intel Core Ultra Series 3 processor running on a Dell Panther Lake mini desktop system, specifically focusing on Edge AI applications for live video captioning. In this industrial safety scenario, a video stream is ingested via an RTSP URL and analyzed in real-time to detect accidents within a workplace environment. The entire processing pipeline for visual analysis is executed on the GPU using a Vision Language Model (VLM), which actively monitors the footage and instantly generates live captions along with alerts whenever potential hazards or incidents are identified.
Complementing the visual detection system, the demonstration utilizes a chatbot powered by the Neural Processing Unit (NPU) to provide intelligent context and reporting. As the VLM processes the video feed, it continuously converts detected events into text and stores these captions as embeddings within a vector database. This setup allows the system to retain a searchable history of events without compromising real-time performance, creating a robust foundation for subsequent data retrieval and analysis by the AI components.
When a user interacts with the chatbot to ask questions about the ongoing situation, the NPU on the Panther Lake system activates to retrieve relevant information from the vector database and generate a comprehensive report. This interaction highlights the efficiency of the Edge AI architecture, where the CPU, GPU, and NPU work in concert to handle different stages of the workflow—from raw video ingestion and visual recognition to natural language processing and data storage. The seamless integration of these components ensures that safety alerts are not only detected but also immediately contextualized for human operators.
Ultimately, this demonstration showcases how the Intel Core Ultra Series 3 platform enables sophisticated, real-time industrial monitoring solutions directly at the edge. By leveraging specialized hardware accelerators like the NPU and GPU, organizations can deploy advanced AI models locally to ensure low-latency safety monitoring without relying on cloud connectivity. The system effectively transforms raw video feeds into actionable insights, providing immediate warnings for accidents while maintaining a detailed record of workplace conditions through embedded text data accessible via natural language queries.
Read the full video transcript
Here what you are seeing is Intel Core
Ultra Series 3 processor running on our
Dell Panther Lake mini desktop system.
You are seeing how the CPU GPU is being
utilized while we do live captioning.
Our video streams comes in through RTSP
URL. We give a prompt to look for any
accidents in the industrial workplace.
We are running the entire pipeline on
GPU using a VLM model. You can see that
if there's any accident, it's generating
live captions, gives an alert to the
screen. On the right hand side of the
demonstration here,
I'm using a chatbot which is running on
NPU. Right now,
the when the VLM is running, all the
text from the captions are being stored
as embeddings in the vector database.
So, when I ask the NPU a question here
to the chatbot,
the NPU on the Panther Lake system
lights up and while it's generating the
report, if there any accidents in the
workplace.
>> [music]