Submind YouTube summaries
Thumbnail for Violoop External AI Accelerator Hardware with 26 TOPS NPU for Local Vision and Automation

Violoop External AI Accelerator Hardware with 26 TOPS NPU for Local Vision and Automation

Watch on YouTube

Video summary

The Violoop External AI Accelerator is a compact hardware device designed to bring powerful local artificial intelligence capabilities to standard computers without requiring a complete system upgrade. By connecting via a single USB-C cable, the device functions as an external monitor that captures HDMI signals along with mouse and keyboard inputs, effectively allowing the AI to "see" and interact with the user's screen in real-time. At its core lies an ALK 3675 chip featuring a Neural Processing Unit (NPU) capable of 26 TOPS, paired with 13 GB of RAM and 128 GB of storage. This architecture enables the execution of large language models and vision-based agents locally, ensuring that sensitive data such as emails, financial records, and personal chats remain on the device rather than being transmitted to the cloud. A primary advantage of this hardware is its ability to automate repetitive tasks with high security and efficiency by analyzing screen content directly. Unlike traditional software agents that must send every pixel to a remote server for processing, Violoop compresses visual data into structural information locally, reducing token usage by 99% and allowing complex models to run on the device itself. The system includes pre-fine-tuned models for specific functions like screen analysis, security monitoring, and task orchestration, while also supporting integration with major cloud providers like OpenAI, Google, and DeepSeek via API keys for more complex operations when necessary. To further enhance privacy, the device incorporates physical security chips that require a manual button press to authorize sensitive actions, ensuring that no data leaves the local environment without explicit user confirmation. The device is particularly useful for users who want to leverage AI assistance without migrating their existing digital lives or dealing with the risks associated with cloud-based agents that may be blocked by legacy software lacking APIs. By simulating real mouse and keyboard movements rather than injecting simulated inputs, Violoop can interact with websites and applications that actively block automation scripts, effectively bypassing captchas and access restrictions. The system also features a proactive mode where users can customize which applications the AI is allowed to monitor, providing granular control over privacy settings. Additionally, the hardware supports voice activation and mobile app synchronization, allowing users to issue commands or receive suggestions from anywhere, making it an accessible tool for non-technical users, such as elderly family members, to manage daily tasks like sending emails or filing taxes with minimal instruction. Looking ahead, Violoop plans to release a refined version of the device that utilizes a single USB 4 cable for a more elegant design while maintaining compatibility with older computers and various display standards. The company is currently launching a Kickstarter campaign to fund production, offering early bird pricing and founder gift packs to supporters who reserve units before the official launch in mid-September. Shipping is scheduled for the end of October, positioning the product as a timely solution for those seeking advanced local AI automation before the holiday season. Ultimately, the vision behind Violoop is to create a personal digital companion that learns user habits over time, grows with their workflow, and handles complex administrative duties autonomously, freeing users to focus on creative work while traveling or relaxing without needing to carry their entire computing setup.
Read the full video transcript
Hi, my name is Enoch and uh right now we have this product called V Loop and uh what it does is a very elegant design and what it does is you just need to connect uh you just need to connect the USB >> type-C on and uh it will get the signal for the HDMI and also like uh the mouse and keyboard signal. So what it what it will do is it will see your screen >> like a real person. >> See everything you do. >> Yeah. >> And then uh this one has Oh, >> let me just start. Sorry. >> Sorry, let me just cut there. Uh so it has a AI accelerator chip in there. >> Yes. >> So what's your chip? >> Uh we got we got the ALK uh 3675 chips in there. So it got a MPU unit with 26 tops in there and uh we got a AI accelerator and also 128 GB of uh hard drive inside. So all your chat data, local datas, uh memory data is going to store here. So >> RAM is it 16 GB? Huh? 16 GB RAM or 8? >> Uh it's uh 8 uh 8 + 5. So it's uh 8 GB plus five because >> kind like 13 or something. >> Yeah. 13. 13. So we got the AI accelerator like basically on there. >> And then you can run a cool model on there. You can run a model locally on this box. >> Yes. >> Without using internet necessarily. >> Yes. Uh so right now we have four fine-tuned model in there. One of those is uh a Gemma E4B model and uh it is uh is fine tuned like to handle various tasks. And the other model is uh because one of the specialties that we have is uh it can see your screen all the time and uh because they see your screen all the time. So we have a fine-tuned uh vision model inside. So compared to uh software agent you have to send every single pixel to the cloud. It doesn't have to do that. >> You can analyze everything kind of you can try to automate some things. >> Uh it can analyze your screen compress it into a structural data. So the token usage is 99% less and therefore it can be run on a local model. So uh >> and then for example if you always do the same thing you always check your email you always uh go on this website or something it can also actually do it by itself. >> Oh yeah of course the first time you learn it and then if it understand it right and then uh we can turn it into a skill and therefore it can rerun over and over again. So it'll learn about your repetitive task and uh keep going. And the advantage is that you can do it with more security because everything is local potentially. >> Correct. If you want and yes correct and second thing is like uh we have uh security chips inside. So what it will do is any sensitive data. So for example like u uh sending email to your important client or like credit card information. >> Yeah. >> You don't send this. >> We don't send this. Uh actually no. uh what happened is we are hardcoded it into the security chips. So it will always trigger and only if you do a physical click and confirmation it will go through. So >> but it's not the fingerprint scanning. >> No, it's not fingerprint scanning. It's a it's a physical click but ultimately you are the one who are in. >> So somebody have to be there. >> Yeah. Yeah. >> To authorize. >> Yeah. You're authorize it. So you are the one who authorizing it. So you are also doing kind of like hybrid with the uh frontier model on the cloud together with this local. >> Yes, correct. So uh a lot of uh important things right uh or complex task it need to go to the frontier model right. So for the frontier model we you connect like we can connect it with the cloud but at the same time we made it so easy that like you can connect with all the different like entropic open AI and uh different like with API keys or you can use your chipd subscription connected to it. So uh that's how you use the frontier model also. >> Can I use my GLM subscription? Of course. uh we optimize for GLM is one of those deepseek or anything like all the basically major player we are connected to it already. >> So the advantage is that this little rock chip chip is actually quite good 26 tops is a lot right it's more than what the MacBook can do. Well, it is actually like faster than a Mac Mini. And then uh also because the reason that we do that is because a lot of people they're using a computer that is not as powerful they and they don't want to buy a new AI computer and but all the information that you have is on your regular computer. So it's very difficult to migrate. So that's why we take all the computation outside which give us another advantage because what happen is like a lot of people is like AI start if you install it on there AI start deleting things or just doing like uh they they they don't feel safe with that for this we regarding on the uh physical security chips uh uh the security button confirmation we also if worse come to worse just unplugged it then uh it doesn't have access because computation run here. So uh that's the physical aspect of control that you have over >> and I guess kind of like the application that you can run on here you can customize with the frontier model or you can develop it you can just make a better use of this uh because what is the actual uh application that you can run here you know it runs its own >> oh you run the app here and it's just using this only for accelerating >> uh well actually everything run here. This is just a parent unit or you can of course you can talk to it, right? So, uh but all the computation just like run right here. You can uh in the future we going to have a community that you can like add different like function like to it. So, so the Gemini model uh E4B what do you call it? >> Uh it's a Gemma E4B. >> Gemma Gemma F3 E4B. Um it's like a a what do you call multiple experts um kind of like model. Yeah. You only need 4B performance. >> Yeah. >> You could in theory run even more because you have 26 tabs. >> Yeah. We we have like uh four different one. One is to analyze the screen >> and then another one is for uh securities, right? And then we've got another one that is for uh uh basically planning uh uh orchestrating uh making the the whole model better. So we got a bunch of things going on. And then we got like the uh whisper, we got the speech model and all those stuff in there. >> What if somebody want to try to run a 31B model on here, a bigger model, maybe a a quantitized version of it? Uh right now we uh we don't let people to install it on there like that. Right. So uh but they can work with their local computer things model on there. Right. So all right. So what's the plan? Right. Right here it says uh reserve for $10. Yeah. >> Get the founder gift pack and $330 saving. Yep. So, uh, right now we are going on Kickstarter on, uh, September 15th and, uh, Kickstarter is going to be $3.99, but if you do it before Kickstarter, we'll be super early bird, but so for $10 deposit, you can get it down to 369 total. >> 369. >> So, you can uh, save even more to it. >> All right, cool. How soon will be shipping? >> Uh, we are shipping at the end of October. >> Very soon. >> It's It's very fast, actually. Let me show you something like uh fun actually. >> So uh for a specific use case, right? Because I'm in the uh EVA uh uh uh like website and then right now I need to like actually contact different client of mine to come by to the booth. So it's right now it is going through the list and then identify like important client for me and then uh it is clicking open because it control my mouse and keyboard and uh it's going to come up with a personalized message after understanding what this person does in their company. So very soon it should like come up with >> it's just going to fill out the text field and click send >> and yes >> soon. >> Let's see. All right here you go. It's like hey you know you might be find this interesting. and come by and take a look and it should click send very soon for my auto >> and it doesn't matter about the speed because it's just going to work for the whole day the whole night whatever >> yeah and second thing is like once it become consistent a skill right then uh it doesn't actually need that much token right it sent out already so as you can see right now it's like 65 of 226 it got a bunch to go right but then uh is while I'm talking to you right now it is doing work for me. So this is and uh we it's possible to just minimize that part in the back and then it does it continues while you do other things on the laptop. You don't need to keep it open in the front. >> Uh depending on the task itself, right? But then uh yeah for this particular case it doesn't and another best part is our mobile app would allow us to just basically communicate it to it directly. So you can be on the beach and then it will still be working for you. See it is synchronizing. >> Is there any chance I can buy four or five and put them together to get a bigger model? >> Uh no right now it's like one computer to each >> one by one only. >> Yeah, one by one. And another very interesting things I'll show you. >> So what happened is so for example right because it see it understand what you're doing for the last like I don't know 2 days 30 minutes or something. And because you can do that, it can even understand your screen and give you live suggestions. Ah, >> so it's just typing message for you for example. Oh, you >> hold on. Hold on. >> So for example, here I just type and then uh it just finished the sentence. >> It finish the sentence. You can just tap and then it's done. So uh so yeah I mean like uh and then it will watch your screen and then slowly remember what you do and then grow with the memory. So it become your companion. It will understand you. Ultimately it'll learn to talk and uh predict like you. So there's a lot of technology like jam-packed in this little device here. >> The dream is to have a little device here that does all your work for you and you can just go on holiday. >> Yeah. Just need to bring my phone and uh that's about it. So uh we have this is the version three we are continual uh innovating and upgrading. So uh version three have a little bit too much wire right but version four on the other hand it become just one wire so it's more elegant. So the latest version that right now everything is >> this is a USB 4 or a thunderbolt 5 or the is a using PCIe bus >> to connect USB 3. Yeah, it's a USB 3. Uh >> USB 3 is fine. >> Yeah, USB 3 is fine. Yeah. So, all right. >> So, it's just like a regular thing because Yeah, because we are also accounted for a lot of like a older computer, >> right? So, they need to like get in. >> Nice. >> So, I try sometimes to connect to Billy and Yuku. It's difficult for me to get the API access. Yeah. >> And what is the challenge when there's no API and you can actually control any website? >> Yeah. So uh a lot of those they don't have API. They are legacy software and uh they're not going to build >> or they don't want to have API. >> Yeah. They don't Yeah. They don't want you to have to access to that. Right. So what we do is uh we connect through the mouse and keyboard. It take real control and then >> full control of mouse and keyboard and you take the real display over the full display >> because it come from the GPU rendering. So we can actually see everything, right? So this way it doesn't need API and can access like I would say 99% of the website on the internet have no API access to it. So this way we are literally building an official API directly with the screen >> with any website you build the API. >> Yes. And uh compared to software agent what happen is like uh because they're simulating the mouse and click input what happen is uh all the platform need to do if they don't want you to access their platform they just need to act to check are you injection based simulation and then you'll get locked out you'll get your capture you need to solve that puzle get banned. Yeah, but this come from the real mouse and keyboard input like a real mouse and keyboard. Just imagine this is a mouse and keyboard. It will not be bent. You will not encounter the capture and all that stuff. So, uh so uh that's that's that's one of the advantage of having a external device on there. Right. Regarding on the on on on >> can you show some other settings you have that you just before you were showing me some setting about the proactive? What was this? So for proactive what it does is it will see your screen because our vision model is able to see your screen compress it into structural data right and therefore it's much smaller it can process locally but then for proactive function uh one is the screen right so for example here it see my screen and then understand what I'm doing yeah what's awesome let's because they see my screen and therefore and just fill up the rest. >> You can fill up the rest and understand you see we're talking about when you come to Barcelona station >> and and it's not just you choose that WhatsApp is allowed to be done that there's like setting for this. >> Yeah. So for privacy concern right we can some people are not comfortable giving full access to the whole AI therefore our proactive >> so you turn on proactive. >> Yep. Yep. Yep. >> And then >> and then I can remember what you work on and then give you suggestion card. But here if you don't want to give it full access I only give it you can only see my WhatsApp you don't see my telegram no we chat but then uh you can click on whatever that you wanted to see so some information for example like a company they have a lot of like uh private information even though it's stored locally all the data locally doesn't go to the cloud but you still don't want it to see and therefore you can turn it off so this is your another layer of privacy >> and also is it an advantage of doing it in a box here because maybe Apple will block this kind of application on Apple on Mac. >> Yes, correct. Uh so there's a lot of like access you cannot get with software and uh >> it's kind of hard to get an app that can see your screen. Even that you have to authorize and I don't know what I'm sure it's really seeing your screen. >> Uh for a individual app it's impossible to just monitor all the time because the only one who have access to monitor all the time is Apple and they don't do it. But on the outside we is it is basically a new monitor and therefore I'm just monitoring my own monitor and therefore it can understand it and say it say all safe locally so uh nothing of your screen is sent to the cloud. >> Can we see more of your settings and what you how you set it up? >> Uh what more is you can show there? >> Yep. So for example right uh the AI model that uh we can we have our own while loop cloud but at the same time you can also set up your own like API keys. >> How to choose what goes to the frontier model and what stays local. How do you choose? So for complex task it's going to go to the frontier model because those important for uh for complex task but for example proactive right it just need to see your screen autocomplete give you some suggestions right those are considered easy task easy task it is going to go through the local model you know why because imagine right autocomplete just like fill out a sentence if every single thing need to go to your car it's going to be first expensive and slow it's not fast So those we will use our fine-tuned gemma model and then uh to do it. So those can just directly because for us it's not just the privacy it's also the speed you need to have a good user experience and here you can just you can have all these subscriptions. >> Yes. And for example, I have GLM, I have Mini Max, I have Gemini, OpenAI. >> I can just activate them all and say put 25% here, 25% there, and put all my uh banking stuff on Google because I trust them. Then don't put it on the Chinese or >> we can work on that. Right now is uh it's just that like you decide which one to use. But at the same time, right, because a lot of user they already got a chdu, >> they already got a chbd subscription, right? and therefore they can connect with the chat GBT if they don't want to use our cloud service. Of course, we also provide a subscription service. >> All right, cool. And it's going to be available all over the world after Kickstarter. >> Yes. So, uh Kickstarter coming on the uh September 15 and then uh and then uh we are shipping at the end of uh October. >> So, it's every shop before Christmas. >> Oh, yeah. Thanksgiving, Christmas. We're ready. And you like the idea of uh having this box help uh like uh an elderly person maybe I don't know my my my grand my grandfather or something or my mom or something it can help them use the computer. Yes. Uh the goal is because every single AI agents like cloth code and all those right it take times to learn to set it up. If you're in programming, yeah, fine. That's not a problem. But then for normal user like me, I want to just have this plug it in, talk to it, and tell it what to do. >> Let's say my grandfather would click here and say, "Help me send an email to my grandson." >> And then Gmail. >> Yeah, you can open Gmail. It will if it might even like teach you like, "Hey, you know what? You don't have to do anything. I'll open the Gmail for you. I'll help you do the clicking. I'll do the sending." just like the showcase that I said >> or like uh hey help me do my taxes now and it's just going to start filling up. >> It's actually quite funny. I actually did my taxes with this when I do my taxes like this. >> Optimize my taxes. >> Yeah. I I was like go to my 2000 like last year folder >> learn how I build it and I want you to duplicate that process for 20 like 25. So that's uh that's literally what I did for my tax. >> Click here and say make me some money uh before Friday or something. It just like finds an idea. >> Yeah. Yeah. Yeah. Yeah. And and uh and and also right you don't even need to click right. You can just like hey vio and then uh it gets voice activated. You can just talk to it. So you put it on your computer in a living room your TV screen and then you literally have a driver in your house. >> I want to click and say negotiate for a cheaper Tesla and it goes on all the websites and ask for cheaper deal. >> Where you might find your Elon Mus. [laughter] Okay. Imagine every pixel, every sound enabled [music] through HDMI technology. From the highest resolutions, [music] the fastest refresh rates to immersive multi-channel [singing] audio. [music] HDMI technology powers a worldwide ecosystem of devices for movies, for gaming, for creators. Experience [music] clarity, performance, connection.