Ryo Suzuki: Augment Human Thought and Creativity with the Power of AR and AI
Watch on YouTubeVideo summary
Ryo Suzuki, a professor leading the Programmable Reality Lab at the University of Colorado Boulder, envisions transforming physical environments into dynamic mediums that augment human thought and creativity using augmented reality (AR) and artificial intelligence (AI). He argues that current interfaces are often restricted to small screens or static documents like PDFs, which fail to leverage rich spatial interactions; instead, his goal is to make every room a living space for thinking and learning. By integrating AI with AR, the team aims to move beyond text-based limitations toward embodied spatial computing, effectively programming the physical world directly through generative methods that can embed instructions onto appliances or alter a room's atmosphere during activities like watching movies.
To achieve this vision, Suzuki outlines three primary contributions: first, using machine learning to automatically generate interactive AR content from static documents without manual preparation, such as turning math textbooks into dynamic simulations; second, democratizing development through "teachable reality," where non-programmers can create experiences by simply demonstrating how everyday objects should function rather than coding sensors or apps; and third, enhancing human communication with real-time, speech-driven presentations that embed visuals and annotations directly into the live environment based on spoken keywords. These tools allow for rapid prototyping of complex research tasks that previously took months to be completed in a single day by experimenting with various representations like data graphics or F-diagrams within educational settings such as classrooms and museums.
The potential impact of these technologies extends significantly to accessibility, particularly aiding visually impaired individuals who face challenges navigating the real world. Researchers from institutions including the University of Washington, the University of Michigan, and Japan's JAXA are developing specialized AR interfaces designed specifically for navigation and mobility support for blind users. This collaborative effort highlights how augmenting physical spaces with spatial information can fundamentally change cognitive processes by providing new types of representation that traditional displays cannot offer, ensuring that advancements in digital media benefit a broader range of people through improved real-world interaction capabilities.
Read the full video transcript
one of the biggest one one
>> 120
>> 120 B the volume of TPB talk and then
I'm happy to like introduce Leo so like
Leo Suzuki is an assistant professor at
a university of coral border and at
institute of develop department of
computer science where he directs the
programmable
reality lab so before joining like a
border he was also indust professor the
computer science at the University of
the Calgary. And then he is his research
mission is to uh enhance human thoughts
and creativity by transforming the
entire living environment into dynamic
space and for souls where people can
think through the tangible and partial
exploration with real objects in the
real world not like just with a virtual
object on the screen. So like since like
he he been publishing many papers in HCI
human computer interaction and robotics
and also some user interface
technologies. So like today you're going
to we're going to see tons of the his
invention of the how we can transform
the spial environment into more creating
uh you know enhancing our creativity and
those so even I think that his research
field is basically human computer
interaction and also computer science
and how we can use the technology to
enhance our capability and then I think
this also inspire you to how you can how
you can use those technology to enhance
your creativity. on new research I think
so I think from that point of view I
think this his research also could be
highly impactful the activity in hope
okay so now Mike is yours thanks
>> w thank you so much for the introduction
uh hi my name is your suzuki so I am
currently assistant professor Colorado
border uh in department science and
institute and today I want to talk about
augment human thought and creativity
with the power of AR and AI I but I I
usually kind of give a talk to computer
science audience but I think you you may
know the AI but does anyone who don't
know the augmented reality like AR or VR
I think it's probably okay I can also
give you some of the example but augment
reality is somehow you know Pokemon Go
like you know like a moving from the
computer screen or a mobile phone you
know in the future we kind of believe
that like a like a visual content
virtual content can be embedded in a
physical world or like through the
glasses. So that that is kind of like a
topic.
But before jumping into the um the talk,
I let me just kind of quickly briefly
introduce myself. So I am originally
from Japan and I graduate from the
University of Tokyo uh back in 201
15 and then I became a PhD student at
the University of Border and during the
PhD uh PhD I also did some kind of
internship at some of the university as
well as industry such as Microsoft and
Adobe and then I became the faculty at
Calgary uh like three years uh and then
in the meantime also work for the Google
uh but the my family kind of complained
about winter of the Calgary so I
returned back to the collab border uh
two years ago so I I'm now going to
assistant process we
so uh my research area is mainly the
computer science uh more specifically
focusing on human computer interaction
uh the the reason why I'm kind of quite
interested in human computer interaction
is because I believe uh uh interfaces
and the representation can actually
going to change the how we think. So let
me just kind of give you an example
about what it means about representation
interfaces uh can augment the human
thought. So here's a simple arithmetics
uh question but can anyone solve this uh
at the moment
[laughter]
and then what about this?
So maybe now you can immediately
understand and solve it uh you know in
your head right then it's kind of
interesting because the data is the same
I mean data is same right because one is
a Roman numeral another one Arabic numer
and your brain also hasn't changed right
but what have changed is just kind of
representation or I interface to
interact with your brain with the number
or the data the digit so as you can See
this kind of example shows how
representation of the interfaces can
accent change or understand the world.
Uh here's another example. So back in
like a 70th century people make sense of
the data by just looking at a table like
this. So uh the back in like a 7086
uh the William Player invented the
technique called data graphics or
nowadays which data visualization and
now you could see the data rather than
read the data right and then now you
cannot imagine the world where that
there's no data graphics here but this
is somehow artificial invented works
right and then how that kind of
representation of the graphic
representation data can change how that
you know symbolic representation uh of
the data to think about it.
So in so that's why I really interested
in kind of representation interfaces for
human thinking and creativity and I
believe computer AI are also very
powerful tools to augment human thought.
For instance, uh like a interactive data
visualization rather than just a data on
the ink and a paper like if you just can
interactively see the data and which
allows you to make sense of data much
more easily and briefly and also such
kind of dynamic physics simulation tool.
So this is a kind of help from
University of Colorado. uh that's that's
kind of the dynamic strategic simulation
tool also allow us to learn like complex
mathematics of physic concept and also
more recently as you could see like a
powerful AI tools like a chatbt can
actually change like how we think learn
and create ideas drastically right so
that's why I I believe the computers and
AI are very powerful tools and to
augment human thought but one of the
problem is The current interfaces or
representation of media does not really
help how we think in the physical world
because that is kind of limited and
constraint into this tiny rectangle
screen. But when we reflect about how we
think in the physical world, it's kind
of much more rich and very you know you
know expressive way to think about it.
For instance, we use kind of space uh
like discussing around table or we also
going to use a physical and a tangible
tool to manipulate and you know uh
organizing. We gonna also walk walk
around for you know look around the
physical and spatial uh orientation and
then also we use a tangible physical
papers to kind of scatter and organize
ideas and then kind of discuss spatially
and the people tinker things like make
things you know like a learn through
like a moving object and a hands and you
know the body and math and so that kind
of the rich tangible and a spatial and
embodied interaction actually going to
you know show that how rich our way of
thinking and understanding and learning
is right and but when we think about
learning thinking and understanding with
a computer it's just constraint into
this tiny rectangle screen ignoring the
entire rich physical and spatial
interaction that you know surround
surrounding us and we have to think in
this tiny rectangle screen which kind of
frustrated because I believe as I said a
medium and representation can untap how
we think and then the the the currently
the computers and AI just a truthful
thought that such kind of to constraint
into this tiny rectangle screen.
So my goal is kind of trying to change
this paradigm into like a transform
instead of just kind of into this
rectangle screen into the space or
entire you know environment can become a
dynamic and computer medium to think and
learn as if you know like a physical
media to live in rather than live with
so which I call like a dynamic space of
thought. So it's more like a dynamic not
tool but dynamic space for thinking and
understanding
and uh here's my vision about the
future. So where like a people can like
learn and understand through the actual
like interacting with the spatially and
understanding and also you can
communicate and uh like a discussing in
a spatial uh manner and also you could
kind of discovering and browsing thing
and also exploring and drawing and
creating thing. So that is kind of how I
envision the like a future look like and
I'm kind of trying to make such kind of
dynamic uh spatial thought to bring it
to uh every single room in every home.
So I think back in 1980s Bill Gates like
a bring the computer or computation
medium to every desk and people love it
but now we have computer in every desk
but I rather think the future could be
like bring such kind of dynamic
computation medium into the every single
room in every home. So that is a kind of
uh I what I'm kind of envisioning. So
just like downloading the app but
downloading the space itself like a
anti-change medium so that you your room
can become like a dynamic medium. Uh
think about it. So that is kind of how I
kind of trying to uh you know uh aim for
it. And then the one key challenges and
the question is like how to blend such
kind of virtual and a physical object
environment in a different different
room because as uh unlike just
downloading app to the same screen your
room and my room is completely
different. So it has to be uh
intelligent intelligently recognize our
environment and understand activity and
embed it response into the real world so
that you know the virtual elements can
seamlessly blend it into the physical
world. And to solve that kind of problem
uh I am trying to combine the augment
reality and AI as a means to drive this
things over and toward the score uh I
have been making a three key
contribution in the three different
area. One is the first one is AI power
augment reality content creation and the
second one is a make such kind of
augment reality ubiquitous to every uh
different room and every different
object and the third one is how to organ
uh human mediated uh AI mediated
communication. So let me uh talk about
the first one. So the first one is about
how can we automatically generate an
interactive augment reality content that
can adapt and augment the real world. So
let me give you an example.
So as I said like augment reality is
something like a Pokemon go but it has
been kind of rich history over the
decades. So one of the application is
kind of learning and educational
context. For example, here's the augment
reality like textbook and where you can
kind of scan the textbook and then yeah
can can see the 3D version and animated
version of the her like this. So this is
a great actually but one of the problem
is this specific uh DX3 content only
works for this specific page of this
specific right and then if you just kind
of scan the different different books
and it doesn't work because you have to
uh prepare and then the program such
kind of air content beforehand and to
make it happen like you have to kind of
create the air content for every single
books but it doesn't work. So that's why
it's not really scalable for you know
many different contexts. So to solve
this problem uh we have done like a like
a new type of approach uh which is
called like augmented math uh which is
kind of AI enable. So using the AI to
generate such kind of dynamic content in
real time uh for uh non repair content.
So um before kind of showing the video,
let me just kind of quickly show uh like
a demo. So this is a kind of actually
there's an air version of it. But here's
kind of scanned math text that I got
from kind of random math uh uh pages.
And here's the kind of static graphs
here. But if you just gonna uh click
this block and also link uh one of the
uh like a equation and now you can kind
of start interacting with to the graphs
to see oh this can change uh this uh you
know x axis and this can change the y-
axis okay what if you know if you change
here oh this rope gonna be changed oh so
this is kind of like a two or three and
then you know how it can interact with
it and then again it's this is not
something like that I prepare
beforehand. So you could just kind of
explore oh for example okay what if you
know if you just kind of uh select this
one okay like oh if like a sign car
going to change in this way or something
that so this is something like it's like
a machine learning driven one approach
meaning that you can just kind of scan
it and then generate such kind of
content on demand uh without preparing
this kind of data. So that is the uh
what we did for the augmented mass. So
in the actual system this is a kind of
camera based system so that you gonna uh
look at the uh document and then you
could also going to interact with it. So
that this can be uh like a makes a
static uh augment static math textbook
dynamic and interactive so that you
don't have to imagine like what happen
but you could also kind of learn and
understand through just kind of
interacting and also manipulating the
data. So this is kind of what I mean by
you know we are living in Arabic nuh
Roman numeral era or dynamic content. So
if you have such kind of dynamic content
then you could see and understand and
learn much more differently.
So um yeah so these are kind of simple
uh example but the actual contribution
here is that we kind of had a machine
learning kind of automatic pipeline that
extract the graphs by looking at the
like a textbook and then the
automatically kind of generating a
linking between the massive w to the
graphs and then generates the dynamic
version of it. So, so that like you
don't really need to like prepare for
this specific one, but whatever kind of
pages that had the gloss can actually
can be augmented
and the later uh we publish it and then
Apple also kind of created like a
dynamic math notes like this. that I
heard. Yeah, this is somehow, you know,
the kind of I I I'm kind of glad like
it's somehow inspired this kind of
dynamic math uh sketch uh interaction
that can kind of scale it to the uh many
different people
and then after that we also kind of try
to extend such kind of notion into the
different domains. So uh instead of math
to the physics or the chemistry. So this
is one of the papers that we had a best
paper award at a whis 2024 uh called
augment physics but let me also show you
just a quick demo. Uh so again like this
is also supposed to be the actual kind
of camera based air content but here
it's just a uh sim uh like a like a
desktop version but here imagine that
you also reading the physics text which
has a diagram like this and then as you
can see like what happened to be you
have to imagine but what if you could
also just kind of say okay just kind of
quickly uh interact with it and then now
you can kind of start simulating uh
physics diagram like this again like
this is not animation but you could also
actually gonna you know yeah try to try
to interact with it and so this is uh so
this is a simulation so what if you
could also kind of change the parameter
for example here the mass is uh 80 gram
so what about 800 and here's the
friction is zero right now so what about
one and then you could see oh the
behavior kind of change because the
friction is a little bit you know high
so it doesn't it kind move story. So as
I said like this is not repair compound
but you could just kind of look at a
textbook and then kind of start making
it in making a simulation even for the
non. So, for example, like here's a
nonconventional
um um
yeah, like a non-conventional diagram
like this. But you could also just say
uh um like a select it and then now we
can kind of try to segment it and
extract the object uh like this. And we
could also um yeah for like uh
uh different type of the uh
application and you could also see oh
like how the lens can behave based on
like a focus
like a circuit by kind of looking at the
uh diagram. um by kind of simulating
like this.
So um this is kind of how uh we uh did
it. So it's kind of based on the
combination the computer graphics which
kind of trying to understand what's
inside of the images and then we get
extract the uh image content such as for
example in this case a slope and also
the person object and then makes the
physics simulation based on this content
on demand uh before and in in the
background. So everything is kind of
very interactive. So you don't really
have to you know like just looking at
the animation but you could also kind of
change the parameters or change the
behavior by looking at it. So again
these are uh pretty much for mo mostly
kind of like very simple physics
simulation
but it kind of works for you know any
any kind of like a physics text that has
a kind of diagram.
So uh we are only right now just
focusing on 2D graphics but uh it's it
doesn't really have to have to be 2D. We
also can trying to extract 2D content to
generate like a 3D dynamic version of
it. for example like if you look at a
skeleton in a textbook then we could
also extract it and then generate a 3D
version of it and also like interact
with the uh like a chemistry you know
like a car or even like a creativity
that kind of so these are technical
challenging because you know we also we
don't generate the 3D content but also
we have to animate the interactive 3D
version of it but we are also really uh
interested in kind of making things
happen for this uh uh through this uh
yeah uh angle and these are only for the
like a math of physics focusing on math
or physics textbook but we are also
interested in like what if you know we
can you know augment any kind of texture
content. So to do that uh we did a kind
of the uh like a LM version like a like
a chat version or the uh like a document
enhancement. So the here uh this guy is
kind of wearing glasses and then this
guy can look at the page and then you
could see the kind of additional uh
content uh such as kind of what the
differences is or what the keyword is
and then based on the keyword you could
also see the images or the table and on
a map
uh by automatically extracting the
content from the physical document
like this.
So what happened here is by looking at
the document uh like a pages and then we
actually extract the content uh by uh
and a text and then using the chbd GP4
here and we can generate a summary
timeline and a keyword list and
information cards and the tables and
then yeah we generate and then show this
kind of additional content centered
around the document uh in in the glasses
in the in the
such as like a summary uh table and a
timeline and like a keyword list and
like a highlights and and
so uh one of the interesting aspect of
this product is we actually deploy it.
So because this is not this this can
basically work for any kind of texture
content. So we recruited people and then
actually use it in their daily lives. So
they use the they look at the posters or
the paper on the wall by kind of asking
the question or they also kind of you
know like while reading a textbook they
also kind of enhance the you know
document but also interesting finding is
they didn't stop at just a reading
documents but they also kind of start
looking at the text around the world. So
one of the interesting application or
use cases was oh they looked at the menu
on the cafeteria and asked like what's a
vegetarian menu and then or even like
what's the healthy menu is and then like
start question answering uh through the
a air glasses with an AI and then also
they kind of look at like which you know
disposable uh like a content should be
in and so on and so forth. So this is
kind of pretty interesting findings and
we are also kind of focusing more on
like what kind of the real world AI
interaction uh in the future. But um to
summarize the first uh category is like
we are interested in how augmented
reality content can be automatically
generated by using uh AI uh content
creation and essentially focusing on
educational and learning uh context.
So uh so that is a kind of first uh last
and for the second contribution uh we
also interested in how to make the such
kind of tangible physical augmented
content uh interaction more ubiquitous.
So this uh last uh the motivation comes
from what when I uh come from the
experiences when I was teaching the
classroom. So I'm currently teaching a
mixed reality and augment reality class
and in that uh in this class and a
student usually kind of creates uh like
a you know application like this for
example you have you have kind of gloss
some kind of physical dishes and then
they kind of manipulate like a virtual
car for example. So this is kind of
interesting and a great application but
making and creating and experiencing
that's kind of uh application is
notoriously hard because you really have
to kind of talk the physical object so
that you have to put the sensor on it
and then you also manipulate and
interact with the virtual core. So
usually they spend like a weeks or even
like a month to create such kind of very
simple application. But when we kind of
trying to you know democratize such high
currencies to nontechnical people like a
scientist like you to just quickly you
know make it it doesn't work. So that's
why like we created uh the system called
uh like a teachable reality which is uh
allows the noner to just generate and
quickly uh such kind of the experiences.
So the teachable reality is kind of
using everyday physical objects such as
like a dishes or the you know cardboard
color or the cars and then generate and
create these kind of interesting
interactive uh experiences and
application in seconds rather than hours
or days or month by using interact uh
like a technique called interacting.
So idea behind it is the it's basically
gonna trying to learn uh the trying to
you know demonstrate it to create it. So
here for example if you use this uh
physical object as a slider you first
kind of trying to uh you know capture
like three different uh position of the
finger and then yeah you're gonna uh you
know place like a object like a tree and
you're just going to link each one of
them to each one of the state here and
then we just going to train the machine
learning model on demand. and then we
can recognize the three different states
by uh three different states of the
finger position and then the
corresponding we also kind of change the
three sides like this. So by doing it uh
we can make a lot of different type of
applications such as you know using uh
like a everyday physical object such as
this is to interact with these kind of
objects or you could also make the
context aware you know assistant or the
paper augmentation and a body
augmentation and so on because that is
just kind of capturing like using a
camera and then yeah just make a pose
and then you know make a different uh
states and then then you can kind of
create such kind of interactive
application. So it's very easy and
limitless uh without you know making a
sensor data or the making application
but it's it's a kind of you know many
different types of you know it kind of
untaps the uh use cases of application
for nonprogrammer student. So some of
the actually in my class uh there are
also some like a physics or chemistry
like a non computer science student but
they could also kind of use this kind of
tool to imagine their own application or
to to kind of interact with it. So that
is a very interesting for me and then
beyond that like we are also kind of uh
moving this forward to not only for uh
like a just a prototyping but also more
like a for uh presentation lecturing uh
type of application. So here uh the the
product for like a reality sketch is
that you can kind of sketch uh just like
a whiteboard but whatever you sketch in
this iPad like augment reality content
can become dynamic. For example, here
like if you just draw a pendum attached
to the physical object and then yeah
what you sketch here is going to be
automatically dynam dynamically kind of
embedded into the physical world so that
you can kind of you know demonstrate and
show like how pandum works and by by
showing the blocks in the physical
world. So idea is simple. So whatever
you set here can become like uh attached
to the physical object. So if you go
object here is a P. So that I like P. So
that if you move it like this kind of
sketch elements can be. So this can be
useful for for example show me like how
deflection can be happen. And here's see
for example like how the middle angle
you could also make a constraint. So how
the middle angle can also change uh for
the math uh situation.
So we created so many kind of different
type of scenario such as you know kind
of enhancing like a like a physics
classroom exper especially kind of
experimental like a physical classroom
or elementary student elementary or
middle school student so that you know
like a people or a stu even student can
draw the thing and then visualize thing
to see what happened by like a learning
by just kind of experiment or you could
also kind of like you know like a use
the math For example, here see like
conceptual demonstration about like how
the different ratio of the gear can
change the uh reversal uh between these
two so that they can again it's more
like make your classroom just like a
science museum so that they can just
kind of create and interact with it and
it's also not only for the you know like
educational content but here is see like
how the you know like yoga posture or
the exercise posture can kind of track
and visualize it. And also like a you
could turn like a everyday objects into
like a dynamic like a sliders or the you
know manipulators of the physical
virtual object and we also kind of uh
enhance such kind of the application to
more like a freehand uh sketch
exploration. So here if you sketch it
and then you could also attach to the
body and then you can make like
animation also. So these are more like a
creativity aspect about like how uh what
if you can instead of just kind of what
if you could see uh the like a sketched
animation on demand uh instead of kind
of waiting for hours of the video
creation. Um one of the application here
is that we actually kind of collaborated
with the performance artists and then
who use that kind of system to augment
their kind of dance perform uh to
generate like a movies or generate
videos uh to enhance it and also uh some
of the teacher also use it to uh enhance
their kind of you know lecture recording
to create uh like a enhanced
uh presentation which I can also talk
later too.
So as I said like one aspect of this one
is uh like a making such kind of
seamlessly blending virtual content into
the physical world is hard but we are
kind of trying to make it more simple
and ubiquitous uh by de democratizing
such kind for non
and then the third one is about we are
also going to try to organize and
enhance human to human communication
because that's how we generate idea
between it. Uh so uh one of the um
motivations is uh from the like here's
he's a Hans Roslin uh he's a expert in
public health but he's also kind of very
famous of the data graphic data
visualization
um he also wrote a book called so maybe
some of you may know so this is kind of
one of the interesting uh demonstration
like how he communicates
the public health knowledge into the
general general people general public
people. So he doesn't use the slides or
like a just a books but he actually kind
of use you know created such kind of the
beta graphics video where he talked
about like how like a pure uh country uh
can become a literature by over the time
by using his body and generating this
kind of data graphics interactively and
then embedded in the real world. So this
is a really inspiring to me because uh
nowadays we are kind of stuck with the
like a slice of presentation like that
but this can be also the future of the
presentation or the future of the
communication between the scientists and
a general public. But one of the key
challenge here is uh you know it kind of
like requires a lot of some time for the
programming and the video recording and
and also most more importantly it
doesn't work in real time. So it cannot
you know do it right now in the live but
you have to kind of record it like how
to do it and then you you have to kind
of synchronize between your motion and
the video which is super tedious.
So that's why we we thought about it.
What if we can kind of generate such
kind of dynamic enhance presentation in
real time uh by using like a
speechdriven approach. So here's uh what
I mean by speechdriven approach. So
here's the one of the demo uh created by
my anagramraph students.
>> Hi, my name is Mar. I'm an undergraduate
student at the University of Calgary.
Today I want to talk about augmented
presentation.
As you can see when I talk about
something we can augment the
presentation using an augmented reality
interface.
We have several features. We have
example live kinetic typography
embedded icons
embedded visuals
and embedded annotations to physical
objects. All components are interacted.
[clears throat] So they could see like
maybe you can understand now that what
he is talking about the the
[clears throat]
video kind of capture like what he's
talking about in real time and then
shows the this kind of clip and also
associated images on demand so that you
can generate these [clears throat] kind
of augmented feature in real time for
the live performance.
So uh to design such kind of system we
first started kind of analyzing like
hundreds of the YouTube videos and
thousand of the scene uh that uses these
kind of techniques uh on YouTube and
then we understand like what kind of the
you know how they use their bodies or
what kind of the information should be
shown and then when
analyzing this um um like analyze
existing videos and and then gen created
uh like a proposing the system
[clears throat] that uses like a voice
based input as a trigger and then yeah
generate dynamic information
that can be manipulated with the body.
So here's the one of the uh application
uh for example like for the teaching
here.
>> Today we're going to discuss white blood
cells and killer tea cells
in particular we'll talk about the HIV
virus and how it attacks the blood cells
and hides from the tea cells.
We can see this displayed through this
diagram of the immune system.
Now, we're going to show a quick video
demonstrating this process.
>> Yeah. And here's another example to
introduce this new water. It
[clears throat] has so many benefits,
including double wall vacuum insulation,
this flexible, perforated handle for
easy use, and it's dishwasher safe. You
can bring it with you to so many
different activities, including going to
the gym. You can bring it to the
outdoors if you'd like. You can even
bring it to the pool.
So as you can see like here like you you
like based on the speech uh yeah based
on the speech uh you could also kind of
generate this kind of title and also
like images associated with it and you
could kind of so uh this is kind of how
we envision the future of the
presentation where it uses kind of live
conversation communication and also
visual spatial and embodied
uh representation about your words,
right? But also it's not only recorded
but it's also not only prepared. It's
kind of think
and yeah another application we also did
a kind of like a similar approaches for
like a data graphics data visualization
like data storytelling. So here like a
you know like using a tangible and a
physical object to talk about data such
as you know how the calorie can change
between like you know like how where
this kind of uh you know like a
nutrition facts of the of the
so everything is kind of improvised on
live. So we are kind of trying to design
the future of the the way of
communicating ideas uh between people
because slides has been uh invented with
somebody else but it doesn't really have
there's no need to be slid so we kind of
thinking about how we can augment the
human to human communication uh for the
future by using the both augment reality
and AI
to house communication
And communication is not only for the
presentation but we also have a
conversation right and where the idea
generate idea is generated. So uh we
also want to briefly talk about like we
kind of did a like a conversation
version of it which is like a whole tbt.
So the uh the idea is like just kind of
you know extracting the keywords from
the conversation and looking at like
showing like images or delet uh around
the conversation.
But yeah maybe you can get the idea how
you could you use such kind of system to
enhance the communication.
So uh this was a kind of first uh
trajectory about like how we can also
augment uh the AI mediated
communication.
So uh by using just a couple of minutes
I also want to talk about like a future
research direction where we are kind of
going to. So I believe like a generative
AI is a very powerful way uh to move
forward especially for example image
generation text generation video
generation nowadays would be very you
know uh rich way of the communication
for example I've been kind of
collaborating with shi kapara to you
know use a con as a communication for
like science to the general public but I
think yeah beyond that we are also
interested in how such kind of
experiences can enhance the real world
interaction and augment interaction by
using augmented reality. So which I call
like a gener generative augmented
reality.
So as I said like my original you know
motivation was like how the
representation can change uh change how
we think and currently like a AI tool
like a chatbt is mostly textual
representation or symbolic
representation right so meaning that you
have to kind of watch on the screen or
the smart but I am envisioning the
future of representation should be more
embedded and spatial and interactive So
here's kind of let me give you a kind of
quick quick example. So when you're in a
kitchen and if you ask the what kind of
recipe you can make uh then most likely
what you can get on the smartphone is
just kind of textual recipe. Right.
Right. So the step one here and step two
here but you have to read it and you
have to kind of you know go back to the
kitchen to you know imag you know to to
synchronize it. But we kind of thinking
about this is not really very suitable
representation. So the suitable
representation should be you are in a
kitchen. So you should be you should be
able to see like what you going to do
next and it's also embedded in a in a
physical object. So for example if you
ask recipe to ch
respond in a physical you know way. So
where you should kind of put it in and
you also can also see like a timer on
the top and when you stop the fire and
so on so forth right so this is kind of
how we believe the future should be and
the AI and here's see one of the
prototype that we uh have making so here
uh it's a organic instruction tool but
instead of kind of asking and gen
getting a text responses we
automatically uh create like a embedded
version of the instruction. For example,
when you ask uh how to use this kind of
printer machine and then how to clean
this printer machine and then as you can
see the instruction is automatically
generated as animation so that you could
just see and to what what you're going
to do next. So that is like a like a
much more rich and easier to understand
what's going on and more important thing
it's actually going to completely 100%
automatic. So you're just going to ask
GBT and we generate this kind of you
know uh dynamic interaction uh based on
what you're looking at in in in the
camera this kind of automatic way. So to
do that what we did is like by you know
analyzing what captivity is responding
back with the text and at the same time
we also capture it and detect like where
the object is and localize it in a 3D
position and also segment object in the
scene so that like we can kind of
generate like a highlight you know like
a move and like hand gesture and so on.
So again like this is kind of one of
glimpse of the future about like a
chatbt. Now you understand uh chat tax
is not really useful in a physical world
but we should be right. So we should we
should use such kind of rich
interaction. So I just want to say like
we are still living in Roman numeral era
of the AI. So the text representation is
not a suitable way but we kind of trying
to invent Arabic now for the AI. So that
is kind of what we trying to do by you
analyzing a physical world. So that if
you ask like even for the simple
question answering such as like what's
like the size of this table and you
could see and embed it in a physical
world. So that is kind of what we are
kind of trying to uh embed. So for
example when you also kind of analyze in
on the whiteboard and you could also see
and kind of generate you know like a
dynamic version of it. And then lastly I
just want to quickly uh talk about for
uh the uh how such kind of the content
can also enhance our entertainment
content. So here's the uh the product
called a casino war where you are
watching on a movie uh in the room but
instead of just kind of watching on in a
2D uh 2D screen the here like based on
extracting the content from the movie
content you can kind of transform your
room for example when you're watching
the rain and your room is kind of start
raining by using the glasses but in the
snow machine you could see the snow and
and change the entire room uh uh to
match to the uh the to the content that
you're working in. So as I said by using
generative AI uh yeah we can extract the
three you know content and a generous 3D
such as letter or the ads like this. So
everything is happen automatically so
without analyzing it. So, so that like
it's kind like that uh we kind of
imagine that Jai is not only the image
generation or the video generation but
we can also now programming the real
world. So that's why like we call it
like a our lab called like a
programmable reality lab but we think
that AI can go beyond just kind of text
or AI images on the screen but we can
kind of start you know interacting and
representing uh in the 3D scene.
So I believe this kind of a direction it
just started and we are kind of trying
to you know make things happen in a
different uh you know in a different
application. Uh overall we kind of
trying to build the future of the
generative augment reality to not only
enhance for the you know like a like a
2D 3D content but actually kind of
trying to makes the entire room like
every single room into dynamics
interactive content for thinking
learning and understanding this should
be uh very you know so we kind of trying
to invent Arabic for the AI. So that's
pretty much it and uh as as I said we
are the program reality lab at
University of Colorado water and we have
a couple PhD students and I also want to
thanks all of the collaborators and uh
also the research institute and also the
sponsors help us to make that happen. So
uh yeah that's pretty much it. So thank
you for uh listening to my talk and I'm
happy to take any questions.
Great. So like uh thank you much talk
and then now uh is anyone have any
question? I don't think maybe you you I
can ah mic is here.
>> Thank you. Um
I have two brief questions. One first uh
how difficult would have been to deploy
any of those things here right now?
Would you have done it
>> as demonstration? So uh for instance uh
well it kind of for instance like this
thing can work for I I think here too
and this movie movie thing is also can
work. So do you have a meta?
>> No.
>> Oh no. Okay. But he he has it maybe uh
you could actually test it out. Yeah.
>> Okay. And then the other question all of
this is super impressive. H have you
been studying what the psychological
response of people is going to be?
Because I feel really overwhelmed by the
AI that's on the screen. So to see this
all around me.
>> I I'm cooperating with him. So that he
he he's more like an expert in the
psychological experience side where I
I'm not I am kind of trying to finance.
So we are doing some kind of very simple
usability study such as like how uh a
house movie can be more entertaining or
immersive but we didn't really do uh how
say like a like a behavioral change or
any kind of you know like a long-term
metrics. Yeah. But uh we are very
interested in that.
>> Yeah. So in the real time of
augumentation of the the presentation,
so how can the presenter monitor what
kind of augmentation is happening and
how can he or she control the type of
augmentation?
>> Yeah, that's a really good question. So
um
this is kind of actual system diagram.
So uh let me also kind of briefly talk
about the like a history about this
product. So we originally uh just tried
to use uh everything is like a like a
generated in in real time. So example
whatever I talk like I just get like a
Google like images from the Google and I
show it and it turns out it's so you
know messy and very how to say you know
yeah very distracting. So that's why
like we kind of did a kind of hybrid
approach where the user first kind of
trying to uh trying to how to say like a
list of keywords and associated images
but the user is kind of trying to how to
say some kind of keywords and then it's
just show the images and then the uh in
terms of the as you can see in terms of
how to control in real time uh for this
specific one is the more like a so when
we did it in z in in the covid era so
where uh when we like a zoom kind of
presentation is more common. So in that
case like we just kind of you know see
the camera here and we we see in a you
know like like a zoom presentation. So
we just kind of see it and kind of
manipulate it. that for more kind of
inerson presentation I think we might
also have to see uh yeah like how how we
can kind of synchronize you know
manipulate it. I think one of the
possible way for the future is assuming
like a like a like a using like a glass
uh so that you could see like what you
are kind of manipulating in in the in
your environ and it's also kind of
synchronized communicated. So that might
be the the way but this was done in
during the covid era back in 2021. So we
focus more on like a zoom zoom type.
Thank you.
>> Do I have one anyone question?
I actually have one question to you for
because like I really like your idea and
also project wise and um where we
because like it's actually augmenting
the spaces and then you know not only
enter screen as a small screen it's more
we can also leverage spial information
>> then it feels like do you think this
kind of expanding workspace using some
special information do you think this
can also enhance really enhance human
source or human cognitive process
because like sometimes we sometimes we
are you working on some small screen and
then thinking about deeply in the data.
>> Some sometimes [clears throat]
you know for like a kind of like
downgrading the level of the dimension
sometimes help us to thinking about a
little bit like a you know like abstract
way something so like sometimes might be
always having some larger space is not
sometimes you know best way. So yeah, I
just want to uh you know ask
>> yeah to precisely understand like what
happened we also actually need more like
a precise usability cycle or study but I
just want to mention that like there are
two say there are kind of two different
how to say like category one is kind of
different medium and one is a
representation and what I'm talking
about representation is because for
instance that you are reading the books
with just kind of physical books, right?
But if you cannot use a Kindle or the
iPad, the representation is still just a
static book, right? So for example, you
cannot watch the video on the static
book and then now iPad allows you to
watch the you know video. So that is
kind of how the different medium enable
the different representation. But if you
just only you know read in a PDF then
it's not I mean medium is kind of change
but representation
right. So I believe the
more interesting question to me is
not just only extending what we have
right now. So for example if you have
the kind of larger screen then maybe
right now we can still do it right. Uh I
think I'm more interested in like a like
a different type of medium such as you
know augment reality enables
representation that couldn't be done in
actual to this to the screen. So one you
know one application like one example is
like a really gen.
So I'm more interested in kind of trying
to find or investigate
the not you know like the representation
that doesn't exist now can you know
thinking rather than
yeah [laughter]
that is very interesting. Okay. How do
you have any like a thoughts of the
ultimate libertification with the
science or something?
>> Yeah, I mean that's that's why we kind
of trying to prototype things. So the
the way we kind of trying to discover uh
is kind of a lot of kind of prototyping
and experiencing. So that's part we also
kind of trying to access accelerate our
own research by using the AI and AR. So
that before like we spend like a couple
of months to protect experiences but now
we can kind of do it in a day. So I
think that's a also how we kind of try
to accelerate our own scientific
experiment with to uh kind of prototype
different type of representation. So we
don't there's no kind of you know single
answer to do but I as for example like a
fine uh invented like a f diagram or you
know like a player created data graphics
a lot of kind of struggles and a
prototype right so I guess that's kind
of how we doing it and we we hope like
maybe at some point we can kind of make
something uh makes thinkable thinkable
yeah for the science yeah science
scientific number. We we have one remote
question so very sure like uh uh have
you met or do you know do you know of
any research of AI AR for blind person
blind people? Uh yeah, there there are
um so one I mean I I actually do have
couple of like differences. So if you I
um let's see um so for example you Udab
makeility lab uh University of
Washington make lab is one of the uh
group who has been doing um like a like
a for like a blind people. Uh another
one is uh
at the Hungu at the University of
Michigan has been also doing uh AR uh
interfaces for the blind people
especially for navigation uh kind of
thing and also in the Japan uh it's not
specifically for the AR but um Jakaw in
the mid has been working real world
navigation uh for that. So I guess this
might be a good uh you know pointer uh
if you if you're interested then yeah
>> thank you. So hope like the person who
made a question of get
>> yeah yeah that's
>> all right so thanks so much everyone to
come and then uh I think Leo will stay
always until next week. No next week
>> 28 29 something like that. So you can
also grab him to talk about more about
like
>> Yeah, I'm also really interested in
learning what's happening here.
>> Yeah, let's talk.
>> All right, thank you so much for coming
and then let's let's thank thank again
for your