Submind YouTube summaries
Thumbnail for Ryo Suzuki: Augment Human Thought and Creativity with the Power of AR and AI

Ryo Suzuki: Augment Human Thought and Creativity with the Power of AR and AI

Watch on YouTube

Video summary

Ryo Suzuki, a professor leading the Programmable Reality Lab at the University of Colorado Boulder, envisions transforming physical environments into dynamic mediums that augment human thought and creativity using augmented reality (AR) and artificial intelligence (AI). He argues that current interfaces are often restricted to small screens or static documents like PDFs, which fail to leverage rich spatial interactions; instead, his goal is to make every room a living space for thinking and learning. By integrating AI with AR, the team aims to move beyond text-based limitations toward embodied spatial computing, effectively programming the physical world directly through generative methods that can embed instructions onto appliances or alter a room's atmosphere during activities like watching movies. To achieve this vision, Suzuki outlines three primary contributions: first, using machine learning to automatically generate interactive AR content from static documents without manual preparation, such as turning math textbooks into dynamic simulations; second, democratizing development through "teachable reality," where non-programmers can create experiences by simply demonstrating how everyday objects should function rather than coding sensors or apps; and third, enhancing human communication with real-time, speech-driven presentations that embed visuals and annotations directly into the live environment based on spoken keywords. These tools allow for rapid prototyping of complex research tasks that previously took months to be completed in a single day by experimenting with various representations like data graphics or F-diagrams within educational settings such as classrooms and museums. The potential impact of these technologies extends significantly to accessibility, particularly aiding visually impaired individuals who face challenges navigating the real world. Researchers from institutions including the University of Washington, the University of Michigan, and Japan's JAXA are developing specialized AR interfaces designed specifically for navigation and mobility support for blind users. This collaborative effort highlights how augmenting physical spaces with spatial information can fundamentally change cognitive processes by providing new types of representation that traditional displays cannot offer, ensuring that advancements in digital media benefit a broader range of people through improved real-world interaction capabilities.
Read the full video transcript
one of the biggest one one >> 120 >> 120 B the volume of TPB talk and then I'm happy to like introduce Leo so like Leo Suzuki is an assistant professor at a university of coral border and at institute of develop department of computer science where he directs the programmable reality lab so before joining like a border he was also indust professor the computer science at the University of the Calgary. And then he is his research mission is to uh enhance human thoughts and creativity by transforming the entire living environment into dynamic space and for souls where people can think through the tangible and partial exploration with real objects in the real world not like just with a virtual object on the screen. So like since like he he been publishing many papers in HCI human computer interaction and robotics and also some user interface technologies. So like today you're going to we're going to see tons of the his invention of the how we can transform the spial environment into more creating uh you know enhancing our creativity and those so even I think that his research field is basically human computer interaction and also computer science and how we can use the technology to enhance our capability and then I think this also inspire you to how you can how you can use those technology to enhance your creativity. on new research I think so I think from that point of view I think this his research also could be highly impactful the activity in hope okay so now Mike is yours thanks >> w thank you so much for the introduction uh hi my name is your suzuki so I am currently assistant professor Colorado border uh in department science and institute and today I want to talk about augment human thought and creativity with the power of AR and AI I but I I usually kind of give a talk to computer science audience but I think you you may know the AI but does anyone who don't know the augmented reality like AR or VR I think it's probably okay I can also give you some of the example but augment reality is somehow you know Pokemon Go like you know like a moving from the computer screen or a mobile phone you know in the future we kind of believe that like a like a visual content virtual content can be embedded in a physical world or like through the glasses. So that that is kind of like a topic. But before jumping into the um the talk, I let me just kind of quickly briefly introduce myself. So I am originally from Japan and I graduate from the University of Tokyo uh back in 201 15 and then I became a PhD student at the University of Border and during the PhD uh PhD I also did some kind of internship at some of the university as well as industry such as Microsoft and Adobe and then I became the faculty at Calgary uh like three years uh and then in the meantime also work for the Google uh but the my family kind of complained about winter of the Calgary so I returned back to the collab border uh two years ago so I I'm now going to assistant process we so uh my research area is mainly the computer science uh more specifically focusing on human computer interaction uh the the reason why I'm kind of quite interested in human computer interaction is because I believe uh uh interfaces and the representation can actually going to change the how we think. So let me just kind of give you an example about what it means about representation interfaces uh can augment the human thought. So here's a simple arithmetics uh question but can anyone solve this uh at the moment [laughter] and then what about this? So maybe now you can immediately understand and solve it uh you know in your head right then it's kind of interesting because the data is the same I mean data is same right because one is a Roman numeral another one Arabic numer and your brain also hasn't changed right but what have changed is just kind of representation or I interface to interact with your brain with the number or the data the digit so as you can See this kind of example shows how representation of the interfaces can accent change or understand the world. Uh here's another example. So back in like a 70th century people make sense of the data by just looking at a table like this. So uh the back in like a 7086 uh the William Player invented the technique called data graphics or nowadays which data visualization and now you could see the data rather than read the data right and then now you cannot imagine the world where that there's no data graphics here but this is somehow artificial invented works right and then how that kind of representation of the graphic representation data can change how that you know symbolic representation uh of the data to think about it. So in so that's why I really interested in kind of representation interfaces for human thinking and creativity and I believe computer AI are also very powerful tools to augment human thought. For instance, uh like a interactive data visualization rather than just a data on the ink and a paper like if you just can interactively see the data and which allows you to make sense of data much more easily and briefly and also such kind of dynamic physics simulation tool. So this is a kind of help from University of Colorado. uh that's that's kind of the dynamic strategic simulation tool also allow us to learn like complex mathematics of physic concept and also more recently as you could see like a powerful AI tools like a chatbt can actually change like how we think learn and create ideas drastically right so that's why I I believe the computers and AI are very powerful tools and to augment human thought but one of the problem is The current interfaces or representation of media does not really help how we think in the physical world because that is kind of limited and constraint into this tiny rectangle screen. But when we reflect about how we think in the physical world, it's kind of much more rich and very you know you know expressive way to think about it. For instance, we use kind of space uh like discussing around table or we also going to use a physical and a tangible tool to manipulate and you know uh organizing. We gonna also walk walk around for you know look around the physical and spatial uh orientation and then also we use a tangible physical papers to kind of scatter and organize ideas and then kind of discuss spatially and the people tinker things like make things you know like a learn through like a moving object and a hands and you know the body and math and so that kind of the rich tangible and a spatial and embodied interaction actually going to you know show that how rich our way of thinking and understanding and learning is right and but when we think about learning thinking and understanding with a computer it's just constraint into this tiny rectangle screen ignoring the entire rich physical and spatial interaction that you know surround surrounding us and we have to think in this tiny rectangle screen which kind of frustrated because I believe as I said a medium and representation can untap how we think and then the the the currently the computers and AI just a truthful thought that such kind of to constraint into this tiny rectangle screen. So my goal is kind of trying to change this paradigm into like a transform instead of just kind of into this rectangle screen into the space or entire you know environment can become a dynamic and computer medium to think and learn as if you know like a physical media to live in rather than live with so which I call like a dynamic space of thought. So it's more like a dynamic not tool but dynamic space for thinking and understanding and uh here's my vision about the future. So where like a people can like learn and understand through the actual like interacting with the spatially and understanding and also you can communicate and uh like a discussing in a spatial uh manner and also you could kind of discovering and browsing thing and also exploring and drawing and creating thing. So that is kind of how I envision the like a future look like and I'm kind of trying to make such kind of dynamic uh spatial thought to bring it to uh every single room in every home. So I think back in 1980s Bill Gates like a bring the computer or computation medium to every desk and people love it but now we have computer in every desk but I rather think the future could be like bring such kind of dynamic computation medium into the every single room in every home. So that is a kind of uh I what I'm kind of envisioning. So just like downloading the app but downloading the space itself like a anti-change medium so that you your room can become like a dynamic medium. Uh think about it. So that is kind of how I kind of trying to uh you know uh aim for it. And then the one key challenges and the question is like how to blend such kind of virtual and a physical object environment in a different different room because as uh unlike just downloading app to the same screen your room and my room is completely different. So it has to be uh intelligent intelligently recognize our environment and understand activity and embed it response into the real world so that you know the virtual elements can seamlessly blend it into the physical world. And to solve that kind of problem uh I am trying to combine the augment reality and AI as a means to drive this things over and toward the score uh I have been making a three key contribution in the three different area. One is the first one is AI power augment reality content creation and the second one is a make such kind of augment reality ubiquitous to every uh different room and every different object and the third one is how to organ uh human mediated uh AI mediated communication. So let me uh talk about the first one. So the first one is about how can we automatically generate an interactive augment reality content that can adapt and augment the real world. So let me give you an example. So as I said like augment reality is something like a Pokemon go but it has been kind of rich history over the decades. So one of the application is kind of learning and educational context. For example, here's the augment reality like textbook and where you can kind of scan the textbook and then yeah can can see the 3D version and animated version of the her like this. So this is a great actually but one of the problem is this specific uh DX3 content only works for this specific page of this specific right and then if you just kind of scan the different different books and it doesn't work because you have to uh prepare and then the program such kind of air content beforehand and to make it happen like you have to kind of create the air content for every single books but it doesn't work. So that's why it's not really scalable for you know many different contexts. So to solve this problem uh we have done like a like a new type of approach uh which is called like augmented math uh which is kind of AI enable. So using the AI to generate such kind of dynamic content in real time uh for uh non repair content. So um before kind of showing the video, let me just kind of quickly show uh like a demo. So this is a kind of actually there's an air version of it. But here's kind of scanned math text that I got from kind of random math uh uh pages. And here's the kind of static graphs here. But if you just gonna uh click this block and also link uh one of the uh like a equation and now you can kind of start interacting with to the graphs to see oh this can change uh this uh you know x axis and this can change the y- axis okay what if you know if you change here oh this rope gonna be changed oh so this is kind of like a two or three and then you know how it can interact with it and then again it's this is not something like that I prepare beforehand. So you could just kind of explore oh for example okay what if you know if you just kind of uh select this one okay like oh if like a sign car going to change in this way or something that so this is something like it's like a machine learning driven one approach meaning that you can just kind of scan it and then generate such kind of content on demand uh without preparing this kind of data. So that is the uh what we did for the augmented mass. So in the actual system this is a kind of camera based system so that you gonna uh look at the uh document and then you could also going to interact with it. So that this can be uh like a makes a static uh augment static math textbook dynamic and interactive so that you don't have to imagine like what happen but you could also kind of learn and understand through just kind of interacting and also manipulating the data. So this is kind of what I mean by you know we are living in Arabic nuh Roman numeral era or dynamic content. So if you have such kind of dynamic content then you could see and understand and learn much more differently. So um yeah so these are kind of simple uh example but the actual contribution here is that we kind of had a machine learning kind of automatic pipeline that extract the graphs by looking at the like a textbook and then the automatically kind of generating a linking between the massive w to the graphs and then generates the dynamic version of it. So, so that like you don't really need to like prepare for this specific one, but whatever kind of pages that had the gloss can actually can be augmented and the later uh we publish it and then Apple also kind of created like a dynamic math notes like this. that I heard. Yeah, this is somehow, you know, the kind of I I I'm kind of glad like it's somehow inspired this kind of dynamic math uh sketch uh interaction that can kind of scale it to the uh many different people and then after that we also kind of try to extend such kind of notion into the different domains. So uh instead of math to the physics or the chemistry. So this is one of the papers that we had a best paper award at a whis 2024 uh called augment physics but let me also show you just a quick demo. Uh so again like this is also supposed to be the actual kind of camera based air content but here it's just a uh sim uh like a like a desktop version but here imagine that you also reading the physics text which has a diagram like this and then as you can see like what happened to be you have to imagine but what if you could also just kind of say okay just kind of quickly uh interact with it and then now you can kind of start simulating uh physics diagram like this again like this is not animation but you could also actually gonna you know yeah try to try to interact with it and so this is uh so this is a simulation so what if you could also kind of change the parameter for example here the mass is uh 80 gram so what about 800 and here's the friction is zero right now so what about one and then you could see oh the behavior kind of change because the friction is a little bit you know high so it doesn't it kind move story. So as I said like this is not repair compound but you could just kind of look at a textbook and then kind of start making it in making a simulation even for the non. So, for example, like here's a nonconventional um um yeah, like a non-conventional diagram like this. But you could also just say uh um like a select it and then now we can kind of try to segment it and extract the object uh like this. And we could also um yeah for like uh uh different type of the uh application and you could also see oh like how the lens can behave based on like a focus like a circuit by kind of looking at the uh diagram. um by kind of simulating like this. So um this is kind of how uh we uh did it. So it's kind of based on the combination the computer graphics which kind of trying to understand what's inside of the images and then we get extract the uh image content such as for example in this case a slope and also the person object and then makes the physics simulation based on this content on demand uh before and in in the background. So everything is kind of very interactive. So you don't really have to you know like just looking at the animation but you could also kind of change the parameters or change the behavior by looking at it. So again these are uh pretty much for mo mostly kind of like very simple physics simulation but it kind of works for you know any any kind of like a physics text that has a kind of diagram. So uh we are only right now just focusing on 2D graphics but uh it's it doesn't really have to have to be 2D. We also can trying to extract 2D content to generate like a 3D dynamic version of it. for example like if you look at a skeleton in a textbook then we could also extract it and then generate a 3D version of it and also like interact with the uh like a chemistry you know like a car or even like a creativity that kind of so these are technical challenging because you know we also we don't generate the 3D content but also we have to animate the interactive 3D version of it but we are also really uh interested in kind of making things happen for this uh uh through this uh yeah uh angle and these are only for the like a math of physics focusing on math or physics textbook but we are also interested in like what if you know we can you know augment any kind of texture content. So to do that uh we did a kind of the uh like a LM version like a like a chat version or the uh like a document enhancement. So the here uh this guy is kind of wearing glasses and then this guy can look at the page and then you could see the kind of additional uh content uh such as kind of what the differences is or what the keyword is and then based on the keyword you could also see the images or the table and on a map uh by automatically extracting the content from the physical document like this. So what happened here is by looking at the document uh like a pages and then we actually extract the content uh by uh and a text and then using the chbd GP4 here and we can generate a summary timeline and a keyword list and information cards and the tables and then yeah we generate and then show this kind of additional content centered around the document uh in in the glasses in the in the such as like a summary uh table and a timeline and like a keyword list and like a highlights and and so uh one of the interesting aspect of this product is we actually deploy it. So because this is not this this can basically work for any kind of texture content. So we recruited people and then actually use it in their daily lives. So they use the they look at the posters or the paper on the wall by kind of asking the question or they also kind of you know like while reading a textbook they also kind of enhance the you know document but also interesting finding is they didn't stop at just a reading documents but they also kind of start looking at the text around the world. So one of the interesting application or use cases was oh they looked at the menu on the cafeteria and asked like what's a vegetarian menu and then or even like what's the healthy menu is and then like start question answering uh through the a air glasses with an AI and then also they kind of look at like which you know disposable uh like a content should be in and so on and so forth. So this is kind of pretty interesting findings and we are also kind of focusing more on like what kind of the real world AI interaction uh in the future. But um to summarize the first uh category is like we are interested in how augmented reality content can be automatically generated by using uh AI uh content creation and essentially focusing on educational and learning uh context. So uh so that is a kind of first uh last and for the second contribution uh we also interested in how to make the such kind of tangible physical augmented content uh interaction more ubiquitous. So this uh last uh the motivation comes from what when I uh come from the experiences when I was teaching the classroom. So I'm currently teaching a mixed reality and augment reality class and in that uh in this class and a student usually kind of creates uh like a you know application like this for example you have you have kind of gloss some kind of physical dishes and then they kind of manipulate like a virtual car for example. So this is kind of interesting and a great application but making and creating and experiencing that's kind of uh application is notoriously hard because you really have to kind of talk the physical object so that you have to put the sensor on it and then you also manipulate and interact with the virtual core. So usually they spend like a weeks or even like a month to create such kind of very simple application. But when we kind of trying to you know democratize such high currencies to nontechnical people like a scientist like you to just quickly you know make it it doesn't work. So that's why like we created uh the system called uh like a teachable reality which is uh allows the noner to just generate and quickly uh such kind of the experiences. So the teachable reality is kind of using everyday physical objects such as like a dishes or the you know cardboard color or the cars and then generate and create these kind of interesting interactive uh experiences and application in seconds rather than hours or days or month by using interact uh like a technique called interacting. So idea behind it is the it's basically gonna trying to learn uh the trying to you know demonstrate it to create it. So here for example if you use this uh physical object as a slider you first kind of trying to uh you know capture like three different uh position of the finger and then yeah you're gonna uh you know place like a object like a tree and you're just going to link each one of them to each one of the state here and then we just going to train the machine learning model on demand. and then we can recognize the three different states by uh three different states of the finger position and then the corresponding we also kind of change the three sides like this. So by doing it uh we can make a lot of different type of applications such as you know using uh like a everyday physical object such as this is to interact with these kind of objects or you could also make the context aware you know assistant or the paper augmentation and a body augmentation and so on because that is just kind of capturing like using a camera and then yeah just make a pose and then you know make a different uh states and then then you can kind of create such kind of interactive application. So it's very easy and limitless uh without you know making a sensor data or the making application but it's it's a kind of you know many different types of you know it kind of untaps the uh use cases of application for nonprogrammer student. So some of the actually in my class uh there are also some like a physics or chemistry like a non computer science student but they could also kind of use this kind of tool to imagine their own application or to to kind of interact with it. So that is a very interesting for me and then beyond that like we are also kind of uh moving this forward to not only for uh like a just a prototyping but also more like a for uh presentation lecturing uh type of application. So here uh the the product for like a reality sketch is that you can kind of sketch uh just like a whiteboard but whatever you sketch in this iPad like augment reality content can become dynamic. For example, here like if you just draw a pendum attached to the physical object and then yeah what you sketch here is going to be automatically dynam dynamically kind of embedded into the physical world so that you can kind of you know demonstrate and show like how pandum works and by by showing the blocks in the physical world. So idea is simple. So whatever you set here can become like uh attached to the physical object. So if you go object here is a P. So that I like P. So that if you move it like this kind of sketch elements can be. So this can be useful for for example show me like how deflection can be happen. And here's see for example like how the middle angle you could also make a constraint. So how the middle angle can also change uh for the math uh situation. So we created so many kind of different type of scenario such as you know kind of enhancing like a like a physics classroom exper especially kind of experimental like a physical classroom or elementary student elementary or middle school student so that you know like a people or a stu even student can draw the thing and then visualize thing to see what happened by like a learning by just kind of experiment or you could also kind of like you know like a use the math For example, here see like conceptual demonstration about like how the different ratio of the gear can change the uh reversal uh between these two so that they can again it's more like make your classroom just like a science museum so that they can just kind of create and interact with it and it's also not only for the you know like educational content but here is see like how the you know like yoga posture or the exercise posture can kind of track and visualize it. And also like a you could turn like a everyday objects into like a dynamic like a sliders or the you know manipulators of the physical virtual object and we also kind of uh enhance such kind of the application to more like a freehand uh sketch exploration. So here if you sketch it and then you could also attach to the body and then you can make like animation also. So these are more like a creativity aspect about like how uh what if you can instead of just kind of what if you could see uh the like a sketched animation on demand uh instead of kind of waiting for hours of the video creation. Um one of the application here is that we actually kind of collaborated with the performance artists and then who use that kind of system to augment their kind of dance perform uh to generate like a movies or generate videos uh to enhance it and also uh some of the teacher also use it to uh enhance their kind of you know lecture recording to create uh like a enhanced uh presentation which I can also talk later too. So as I said like one aspect of this one is uh like a making such kind of seamlessly blending virtual content into the physical world is hard but we are kind of trying to make it more simple and ubiquitous uh by de democratizing such kind for non and then the third one is about we are also going to try to organize and enhance human to human communication because that's how we generate idea between it. Uh so uh one of the um motivations is uh from the like here's he's a Hans Roslin uh he's a expert in public health but he's also kind of very famous of the data graphic data visualization um he also wrote a book called so maybe some of you may know so this is kind of one of the interesting uh demonstration like how he communicates the public health knowledge into the general general people general public people. So he doesn't use the slides or like a just a books but he actually kind of use you know created such kind of the beta graphics video where he talked about like how like a pure uh country uh can become a literature by over the time by using his body and generating this kind of data graphics interactively and then embedded in the real world. So this is a really inspiring to me because uh nowadays we are kind of stuck with the like a slice of presentation like that but this can be also the future of the presentation or the future of the communication between the scientists and a general public. But one of the key challenge here is uh you know it kind of like requires a lot of some time for the programming and the video recording and and also most more importantly it doesn't work in real time. So it cannot you know do it right now in the live but you have to kind of record it like how to do it and then you you have to kind of synchronize between your motion and the video which is super tedious. So that's why we we thought about it. What if we can kind of generate such kind of dynamic enhance presentation in real time uh by using like a speechdriven approach. So here's uh what I mean by speechdriven approach. So here's the one of the demo uh created by my anagramraph students. >> Hi, my name is Mar. I'm an undergraduate student at the University of Calgary. Today I want to talk about augmented presentation. As you can see when I talk about something we can augment the presentation using an augmented reality interface. We have several features. We have example live kinetic typography embedded icons embedded visuals and embedded annotations to physical objects. All components are interacted. [clears throat] So they could see like maybe you can understand now that what he is talking about the the [clears throat] video kind of capture like what he's talking about in real time and then shows the this kind of clip and also associated images on demand so that you can generate these [clears throat] kind of augmented feature in real time for the live performance. So uh to design such kind of system we first started kind of analyzing like hundreds of the YouTube videos and thousand of the scene uh that uses these kind of techniques uh on YouTube and then we understand like what kind of the you know how they use their bodies or what kind of the information should be shown and then when analyzing this um um like analyze existing videos and and then gen created uh like a proposing the system [clears throat] that uses like a voice based input as a trigger and then yeah generate dynamic information that can be manipulated with the body. So here's the one of the uh application uh for example like for the teaching here. >> Today we're going to discuss white blood cells and killer tea cells in particular we'll talk about the HIV virus and how it attacks the blood cells and hides from the tea cells. We can see this displayed through this diagram of the immune system. Now, we're going to show a quick video demonstrating this process. >> Yeah. And here's another example to introduce this new water. It [clears throat] has so many benefits, including double wall vacuum insulation, this flexible, perforated handle for easy use, and it's dishwasher safe. You can bring it with you to so many different activities, including going to the gym. You can bring it to the outdoors if you'd like. You can even bring it to the pool. So as you can see like here like you you like based on the speech uh yeah based on the speech uh you could also kind of generate this kind of title and also like images associated with it and you could kind of so uh this is kind of how we envision the future of the presentation where it uses kind of live conversation communication and also visual spatial and embodied uh representation about your words, right? But also it's not only recorded but it's also not only prepared. It's kind of think and yeah another application we also did a kind of like a similar approaches for like a data graphics data visualization like data storytelling. So here like a you know like using a tangible and a physical object to talk about data such as you know how the calorie can change between like you know like how where this kind of uh you know like a nutrition facts of the of the so everything is kind of improvised on live. So we are kind of trying to design the future of the the way of communicating ideas uh between people because slides has been uh invented with somebody else but it doesn't really have there's no need to be slid so we kind of thinking about how we can augment the human to human communication uh for the future by using the both augment reality and AI to house communication And communication is not only for the presentation but we also have a conversation right and where the idea generate idea is generated. So uh we also want to briefly talk about like we kind of did a like a conversation version of it which is like a whole tbt. So the uh the idea is like just kind of you know extracting the keywords from the conversation and looking at like showing like images or delet uh around the conversation. But yeah maybe you can get the idea how you could you use such kind of system to enhance the communication. So uh this was a kind of first uh trajectory about like how we can also augment uh the AI mediated communication. So uh by using just a couple of minutes I also want to talk about like a future research direction where we are kind of going to. So I believe like a generative AI is a very powerful way uh to move forward especially for example image generation text generation video generation nowadays would be very you know uh rich way of the communication for example I've been kind of collaborating with shi kapara to you know use a con as a communication for like science to the general public but I think yeah beyond that we are also interested in how such kind of experiences can enhance the real world interaction and augment interaction by using augmented reality. So which I call like a gener generative augmented reality. So as I said like my original you know motivation was like how the representation can change uh change how we think and currently like a AI tool like a chatbt is mostly textual representation or symbolic representation right so meaning that you have to kind of watch on the screen or the smart but I am envisioning the future of representation should be more embedded and spatial and interactive So here's kind of let me give you a kind of quick quick example. So when you're in a kitchen and if you ask the what kind of recipe you can make uh then most likely what you can get on the smartphone is just kind of textual recipe. Right. Right. So the step one here and step two here but you have to read it and you have to kind of you know go back to the kitchen to you know imag you know to to synchronize it. But we kind of thinking about this is not really very suitable representation. So the suitable representation should be you are in a kitchen. So you should be you should be able to see like what you going to do next and it's also embedded in a in a physical object. So for example if you ask recipe to ch respond in a physical you know way. So where you should kind of put it in and you also can also see like a timer on the top and when you stop the fire and so on so forth right so this is kind of how we believe the future should be and the AI and here's see one of the prototype that we uh have making so here uh it's a organic instruction tool but instead of kind of asking and gen getting a text responses we automatically uh create like a embedded version of the instruction. For example, when you ask uh how to use this kind of printer machine and then how to clean this printer machine and then as you can see the instruction is automatically generated as animation so that you could just see and to what what you're going to do next. So that is like a like a much more rich and easier to understand what's going on and more important thing it's actually going to completely 100% automatic. So you're just going to ask GBT and we generate this kind of you know uh dynamic interaction uh based on what you're looking at in in in the camera this kind of automatic way. So to do that what we did is like by you know analyzing what captivity is responding back with the text and at the same time we also capture it and detect like where the object is and localize it in a 3D position and also segment object in the scene so that like we can kind of generate like a highlight you know like a move and like hand gesture and so on. So again like this is kind of one of glimpse of the future about like a chatbt. Now you understand uh chat tax is not really useful in a physical world but we should be right. So we should we should use such kind of rich interaction. So I just want to say like we are still living in Roman numeral era of the AI. So the text representation is not a suitable way but we kind of trying to invent Arabic now for the AI. So that is kind of what we trying to do by you analyzing a physical world. So that if you ask like even for the simple question answering such as like what's like the size of this table and you could see and embed it in a physical world. So that is kind of what we are kind of trying to uh embed. So for example when you also kind of analyze in on the whiteboard and you could also see and kind of generate you know like a dynamic version of it. And then lastly I just want to quickly uh talk about for uh the uh how such kind of the content can also enhance our entertainment content. So here's the uh the product called a casino war where you are watching on a movie uh in the room but instead of just kind of watching on in a 2D uh 2D screen the here like based on extracting the content from the movie content you can kind of transform your room for example when you're watching the rain and your room is kind of start raining by using the glasses but in the snow machine you could see the snow and and change the entire room uh uh to match to the uh the to the content that you're working in. So as I said by using generative AI uh yeah we can extract the three you know content and a generous 3D such as letter or the ads like this. So everything is happen automatically so without analyzing it. So, so that like it's kind like that uh we kind of imagine that Jai is not only the image generation or the video generation but we can also now programming the real world. So that's why like we call it like a our lab called like a programmable reality lab but we think that AI can go beyond just kind of text or AI images on the screen but we can kind of start you know interacting and representing uh in the 3D scene. So I believe this kind of a direction it just started and we are kind of trying to you know make things happen in a different uh you know in a different application. Uh overall we kind of trying to build the future of the generative augment reality to not only enhance for the you know like a like a 2D 3D content but actually kind of trying to makes the entire room like every single room into dynamics interactive content for thinking learning and understanding this should be uh very you know so we kind of trying to invent Arabic for the AI. So that's pretty much it and uh as as I said we are the program reality lab at University of Colorado water and we have a couple PhD students and I also want to thanks all of the collaborators and uh also the research institute and also the sponsors help us to make that happen. So uh yeah that's pretty much it. So thank you for uh listening to my talk and I'm happy to take any questions. Great. So like uh thank you much talk and then now uh is anyone have any question? I don't think maybe you you I can ah mic is here. >> Thank you. Um I have two brief questions. One first uh how difficult would have been to deploy any of those things here right now? Would you have done it >> as demonstration? So uh for instance uh well it kind of for instance like this thing can work for I I think here too and this movie movie thing is also can work. So do you have a meta? >> No. >> Oh no. Okay. But he he has it maybe uh you could actually test it out. Yeah. >> Okay. And then the other question all of this is super impressive. H have you been studying what the psychological response of people is going to be? Because I feel really overwhelmed by the AI that's on the screen. So to see this all around me. >> I I'm cooperating with him. So that he he he's more like an expert in the psychological experience side where I I'm not I am kind of trying to finance. So we are doing some kind of very simple usability study such as like how uh a house movie can be more entertaining or immersive but we didn't really do uh how say like a like a behavioral change or any kind of you know like a long-term metrics. Yeah. But uh we are very interested in that. >> Yeah. So in the real time of augumentation of the the presentation, so how can the presenter monitor what kind of augmentation is happening and how can he or she control the type of augmentation? >> Yeah, that's a really good question. So um this is kind of actual system diagram. So uh let me also kind of briefly talk about the like a history about this product. So we originally uh just tried to use uh everything is like a like a generated in in real time. So example whatever I talk like I just get like a Google like images from the Google and I show it and it turns out it's so you know messy and very how to say you know yeah very distracting. So that's why like we kind of did a kind of hybrid approach where the user first kind of trying to uh trying to how to say like a list of keywords and associated images but the user is kind of trying to how to say some kind of keywords and then it's just show the images and then the uh in terms of the as you can see in terms of how to control in real time uh for this specific one is the more like a so when we did it in z in in the covid era so where uh when we like a zoom kind of presentation is more common. So in that case like we just kind of you know see the camera here and we we see in a you know like like a zoom presentation. So we just kind of see it and kind of manipulate it. that for more kind of inerson presentation I think we might also have to see uh yeah like how how we can kind of synchronize you know manipulate it. I think one of the possible way for the future is assuming like a like a like a using like a glass uh so that you could see like what you are kind of manipulating in in the in your environ and it's also kind of synchronized communicated. So that might be the the way but this was done in during the covid era back in 2021. So we focus more on like a zoom zoom type. Thank you. >> Do I have one anyone question? I actually have one question to you for because like I really like your idea and also project wise and um where we because like it's actually augmenting the spaces and then you know not only enter screen as a small screen it's more we can also leverage spial information >> then it feels like do you think this kind of expanding workspace using some special information do you think this can also enhance really enhance human source or human cognitive process because like sometimes we sometimes we are you working on some small screen and then thinking about deeply in the data. >> Some sometimes [clears throat] you know for like a kind of like downgrading the level of the dimension sometimes help us to thinking about a little bit like a you know like abstract way something so like sometimes might be always having some larger space is not sometimes you know best way. So yeah, I just want to uh you know ask >> yeah to precisely understand like what happened we also actually need more like a precise usability cycle or study but I just want to mention that like there are two say there are kind of two different how to say like category one is kind of different medium and one is a representation and what I'm talking about representation is because for instance that you are reading the books with just kind of physical books, right? But if you cannot use a Kindle or the iPad, the representation is still just a static book, right? So for example, you cannot watch the video on the static book and then now iPad allows you to watch the you know video. So that is kind of how the different medium enable the different representation. But if you just only you know read in a PDF then it's not I mean medium is kind of change but representation right. So I believe the more interesting question to me is not just only extending what we have right now. So for example if you have the kind of larger screen then maybe right now we can still do it right. Uh I think I'm more interested in like a like a different type of medium such as you know augment reality enables representation that couldn't be done in actual to this to the screen. So one you know one application like one example is like a really gen. So I'm more interested in kind of trying to find or investigate the not you know like the representation that doesn't exist now can you know thinking rather than yeah [laughter] that is very interesting. Okay. How do you have any like a thoughts of the ultimate libertification with the science or something? >> Yeah, I mean that's that's why we kind of trying to prototype things. So the the way we kind of trying to discover uh is kind of a lot of kind of prototyping and experiencing. So that's part we also kind of trying to access accelerate our own research by using the AI and AR. So that before like we spend like a couple of months to protect experiences but now we can kind of do it in a day. So I think that's a also how we kind of try to accelerate our own scientific experiment with to uh kind of prototype different type of representation. So we don't there's no kind of you know single answer to do but I as for example like a fine uh invented like a f diagram or you know like a player created data graphics a lot of kind of struggles and a prototype right so I guess that's kind of how we doing it and we we hope like maybe at some point we can kind of make something uh makes thinkable thinkable yeah for the science yeah science scientific number. We we have one remote question so very sure like uh uh have you met or do you know do you know of any research of AI AR for blind person blind people? Uh yeah, there there are um so one I mean I I actually do have couple of like differences. So if you I um let's see um so for example you Udab makeility lab uh University of Washington make lab is one of the uh group who has been doing um like a like a for like a blind people. Uh another one is uh at the Hungu at the University of Michigan has been also doing uh AR uh interfaces for the blind people especially for navigation uh kind of thing and also in the Japan uh it's not specifically for the AR but um Jakaw in the mid has been working real world navigation uh for that. So I guess this might be a good uh you know pointer uh if you if you're interested then yeah >> thank you. So hope like the person who made a question of get >> yeah yeah that's >> all right so thanks so much everyone to come and then uh I think Leo will stay always until next week. No next week >> 28 29 something like that. So you can also grab him to talk about more about like >> Yeah, I'm also really interested in learning what's happening here. >> Yeah, let's talk. >> All right, thank you so much for coming and then let's let's thank thank again for your