Video summary
This working session serves as the inaugural meeting of a series dedicated to developing the PDP-Connect standard, which aims to establish a shared language for data portability across different platforms and applications. The discussion highlights that current methods for exporting personal data often result in bundled packages containing irrelevant information like IP logs or credit card details alongside user content. To address this inefficiency, the team proposes creating an architecture where users can grant granular access to specific slices of their data rather than entire archives. This approach relies on a flexible framework that defines how applications request and consume data without being overly opinionated about storage formats, identity management systems, or deployment locations, thereby encouraging broad adoption by various providers including banks and social media companies.
The proposed architecture centers on two main components: the selection process for requesting specific data access and the record model used to represent diverse data sources generically. The system builds upon existing authorization protocols like OAuth 2.0 but extends them with an envelope that carries detailed semantic information about what is being shared, effectively creating a contract between the user and the client application enforced by the server hosting the data. A key feature of this design is its neutrality regarding storage; it allows users to maintain their own personal servers or utilize connectors provided by third parties to bridge gaps where platforms lack native API support for portability. This flexibility ensures that even if major tech giants do not immediately adopt a specific standard, individuals can still exercise control over their data through self-hosted solutions and community-built tools.
Significant attention was given to the implications of this decentralized model regarding legal jurisdiction and government subpoenas within an increasingly AI-driven landscape. The speakers argued that by placing custody of personal data directly with the user—similar to a non-custodial crypto wallet—the system inherently limits governmental reach, as authorities would only have jurisdiction over individuals physically present in their territory rather than controlling a global network. This structure aligns well with regulations like the EU's Digital Markets Act and GDPR while offering robust protections against cross-border data demands from sanctioned nations or entities seeking to exert undue influence. Furthermore, the session demonstrated a live proof-of-concept where an AI agent successfully queried specific subsets of ChatGPT memories via a personal server, illustrating how future agents could seamlessly access user histories for tasks like financial planning without exposing sensitive information unnecessarily.
Looking ahead, the roadmap includes four sessions throughout August leading to a formal launch in Geneva, with subsequent weeks dedicated to deep dives into record models, grants, connectors, and resource servers. The team emphasized that while early iterations of data portability efforts faced challenges due to big tech companies ignoring standards, the current landscape is shifting as tooling becomes more accessible for smaller organizations and individuals to build their own solutions. Projects like Buzz are already emerging as examples of entities leveraging these protocols to host information securely and write data back into ecosystems using new agent standards. The consensus among product managers, engineers, and policy experts present was that now is a critical time to finalize this shared language before the market demand for true user sovereignty peaks, ensuring that personal data can flow freely between applications while respecting individual consent at every step.
Read the full video transcript
Hey, Daniela. Good morning.
>> Hey, Anna. How are you?
>> Good. How are you doing?
>> I'm doing all right. I can just stay for
a little bit, but I wanted to kick it
off with you guys. [laughter]
>> Totally. Thanks for joining.
>> There you go. Are you back on this?
>> Yeah. Yeah, I'm back in. in uh wait
Monter. Oh, wait. Why do I always think
>> Pacifica? Yeah,
>> there we go. Yeah,
>> just uh this week and then I go to
Brazil next week again. So,
>> awesome. Cool.
>> Yeah.
>> What's in Brazil?
>> Uh huge community there. We have like
over 6,000 uh participants in our
regional chapter. Um we have a handful
of members. Uh we work very closely with
you know other government agencies down
there. huge adoption of our tech um in
uh in Brazil. So there's an event next
week called Blockchain Rio.
>> So we have like a full day of content uh
that our regional chapter has put
together with members, government, you
know, representatives and stuff. So um
yeah, so uh I'll be talking about you
guys, I'm sure.
>> Very cool.
>> I always like to bring our new projects
and [clears throat] stuff in there. So
here I'll put it.
Hey Tim, how are you? Good to see you.
>> How's it going?
>> Good. Good. Busy as always. I don't
know. Everybody else is like, "Oh, I'm
taking the rest of the month off." I'm
like, "Not us." [laughter]
>> So, um,
>> you have a lot of conferences and
events, right?
>> Yeah.
>> Yeah. I I joke all the time that, you
know, when they first hired me, they
forgot to tell me that the big the big
portion of what we do is events. Events
like this, you know, obviously community
events, but we do, you know, a lot of on
the ground inerson events as well, which
is which is great, which is how this
work gets done. [snorts]
Anyway, enough about me. Good to see
everybody.
>> Hey Sarah, good to see you.
>> Hey Anna. Hey team. Really nice to be
here. Yes, Sarah.
>> Um, cool. Well, I think we'll get
started. Um, let me just pull up some
slides that or actually Tim, do you want
to pull up the slides for today? I can
kick us off.
>> Yeah, one sec.
>> Sarah, do you want to give a a quick
intro? I think everyone else in this
group knows.
>> Yeah, definitely. Um, thank you everyone
for uh for um allowing me to join you on
on PDP for a little bit. Um, I'm Sarah.
I'm a product uh person in London in the
UK. Um, I lead up product for a human
data uh company called Prolific. Um,
human data is a really busy space. We
try to do it a little bit um more
ethically, I want to say, and with with
a bit more human interest at the heart
of it. Um I'm actually departing this
role pretty soon and I've become like I
I don't want to overuse the word but
maybe the right word is obsessed with uh
uh data pro provenence and aentic
provenence and portability and
authorization and um really keen to see
if I can apply some of that passion and
direction of helping the community
community a little bit. It's really nice
to meet you all.
>> Awesome. Thanks for joining.
Um maybe we'll do a quick round of
intros for you Sarah too. So, we already
know each other. Um, so I I'll pass it.
Actually, Daniela, do you want to give
an intro and then we could do the rest
of the folks on the bon side?
>> Sure. Nice to see everyone. I'm
Daniellea Barbosa. I'm the executive
director for Linux Foundation
Escentralized Trust, which is the host
uh organization for this project. So,
very excited. Been working with Art and
Anna and the rest of the team for a
while. uh and my uh I actually you know
started in the identity world uh in uh I
think it was 2008 with uh uh as one of
the founders of a project at the time
called the data portability project. So
it's really excited to see some of this
work you know get into uh into the
foundation very importantly get into the
ecosystem as well. So um I will not be
joining every single meeting but I
thought this was the inaugural one so I
was like hey I want to come and check it
in and uh and see everyone. So, thank
you all for having me.
>> Thanks for making the time and getting
up early to join us in the midst of a
busy travel schedule.
>> I'm sorry, folks. My
>> Yeah, good.
>> Zoom crashed when I tried sharing my
screen. So,
um Anna, maybe if you don't mind sharing
the slides.
>> Yeah. Yeah, sure. Um and Tim and Machi,
do you guys want to give an intro too
before we jump in?
>> Yeah, definitely. Um,
so maybe I'll talk a little bit about my
intro um, as we talk more about PPP, but
basically I've been doing full stack
engineering for a number of years across
a bunch of different domains um, retail,
HR, fintech and I've been working on and
with VA for the past almost four years
working on data portability, working on
AI consumer um, projects related to data
portability, working on social
coordination and decentralization
problems and
um I'm just really fascinated by systems
design and
like the engineering problems that
relate to coordinating with lots of
people and and enabling user sovereign
so sovereignty just based on the way
that technology has um moved so fast.
And I think everybody's probably
wondering a little bit like what uh
where does that leave the the individual
human being. So I just think this is a
really cool and interesting problem
space to be in. Um Machi, I'll pass it
over to you.
>> Thanks Tim. Uh hi everyone. My name is
Mache. I I uh lead the engineering team
at VANA. I've been at Vana for for the
last year and before that I've worked
for like was like 15 years as an
engineer at uh various companies work
like from from very small startups to
comp through companies uh like Zenesk
and then I went into blockchain and you
know this kind of I was very passionate
about decentralization of the internet
with like IPFS and Falcoin uh at
protocol apps and then I joined VA
actually when we when I met Anna and we
started talking about VA it was
specifically about data portability and
I f and I consider that as a you know
kind of next step in the journey of
decentralizing the internet by you know
giving uh people access to their data.
Good to see you all.
>> Awesome. Hey Justin, thanks for joining.
Hi Q. I think it might be the middle of
the night for you. So, thank you for
being awake at this strange hour. And
hey, hey, Nas, too. Feel free to jump in
with an an intro if you want, but we
we'll also just get started. Um, but but
good to see you. Thanks for hopping on.
Um, so today will be kind of uh the
first session of four sessions
throughout August going into kind of the
proposed PDP
connect standard and kind of proposing
this um what I would describe as like a
shared language for data portability. I
think the closest analoges are like X42
um for agentic payments and um MCP for
bringing data into an application. And
the reason why I think the standard is
important is it's like how do you have a
a shared language so that it can be easy
for applications to expose data and for
applications know kind of the data that
they're consuming. Um so this is what
the month looks like kind of first just
introduction and architecture and then
next week diving deep into the record
model following week grants and
connectors and then um August 27th the
resource server and open questions and
then in September launching this in
Geneva um so yeah that's what the month
ahead looks like. Um, really I think
kind of what PDP is is designed to do is
to specify how you can authorize access
to personal data in a shared neutral way
across different platforms. Um, I think
often an example can be helpful of like
why is it why would something like this
be better than what currently exists. So
right now if you go and get like say a
GDPR export of your um Instagram data or
Spotify or ChatGBT data often everything
is kind of all bundled together. So if
you export your data from a platform it
might include like all of your IP logs
or your credit card information etc.
Right? And so if you want to bring your
um context into a given application you
probably don't want to bring everything.
You just want to bring some stuff. And
so having a very clear way of saying,
"Hey, grant granular access for this
specifically really goes a long way,
especially as you start to bring it
across different applications."
Um, I think that often when you think
about data and I'd say like this maybe I
imagine Tim and Mache have a view on
too, but like I think in the past few
years something we've realized is that
you actually can unify a lot of data
sources and one question we often get is
like hey data is so different, right?
like some places are exporting a MD
file, JSON file, videos, audios to just
like these mega zip exports. And I think
one thing we've seen is that actually if
you just define like a shared language
and kind of map everything down to like
a stream or a file, you actually can do
it. And you have to be kind of flexible
in um saying, okay, each platform
defines their own schema, each data
source defines their own schema, and
kind of not overly opinionated. That's
something we'll get into in some of the
details. Tim will share it later too on
on being storage agnostic. Um but I
guess what it is to say here is like if
you take this user first identity model
I think you actually can be quite
flexible in terms of what um you're able
to represent which is pretty much all
personal data right which ultimately is
the goal of of this specification. Um
Tim I'll pass it over to you uh to to
continue.
>> Awesome. Yeah, I think on that point um
I've been really surprised that we can
model I I guess when you think about how
most data on the internet is stored, a
lot of it's just in SQL databases and
they're using the same primitives like
tables or collections of data,
relationships between records, um
indexes over that data and we so we've
already kind of solved this problem um
in terms of like how you work with data
and just how to wrap up personal data.
in this unified standard like language
is the piece that's been missing. So in
just sort of building out PDP and
playing around with it, I've already
accumulated like four and a half million
records in personal data from like 20
different data sources from my own data.
And I know there's a lot more that I
could um sort of pull in and I just have
a lot of confidence that this is
possible. Maybe we can make some tweaks.
Um, and there there will always be
improvements we can make. I think there
could be some exceptions to what fits in
this framework like maybe high frequency
real time like streaming video that's
high bandwidth or something. But even
then, I think you could still um sort of
put PDP compatible like wrappers around
that data and just point to like the
heavy stuff and still have some way of
talking about what's in there.
Um, yeah, if you can move forward to the
next one. I
>> I think also like the timing for this is
really good because frontier models are
proving to be very capable and they're
able to accomplish a lot of everyday
tasks. Maybe um not every single person
is using AI like frequently within their
lives right now, but I think we're
seeing strong evidence that
um like making the making the newest AI
as useful as possible is is getting to
be less about like which model do you
use and more about what context are you
able to provide to it. And
like as people's digital lives continue
to grow and we rely maybe more on new
technology and AI systems that means
that there's even more personal data
that has value to the to the user and
that ultimately like making portable
will
um be important. Also on the product
side, we're we're seeing like this flood
of MCP servers for example from existing
products and these are companies
realizing that they can actually make
their product value propositions
stronger by moving toward data
portability and giving users access to
the data. So everything is kind of
aligning um to make this I think a
really like great time to solve data
portability.
And so going forward um to the next one,
Anna, like
right now we have um data connectors in
the PDP
um connect um organization in within the
lab. We have a whole bunch of data
connectors like dozens of them that
connect to chatbt, aura, shop, um a
bunch of others. And you can think of
these sort of like MCP servers. they
they like run with the client or some
environment controlled by the user and
they sort of adapt over the existing
surface of some let's say cloud hosted
product and this already exists today
and creates like an interface that's PDP
um conformant and enables this language
to be spoken about the data enables the
user to sort of take their data with
them um in addition to this which will
take maybe a little bit more time is
platforms could also natively support
PPP within their APIs. Um they could
ship official connectors that have like
really strong support for this um for
this standard. And really PPP is not
super opinionated about the deployment.
Um, the important thing is that there's
there's some resource server that a
client application can talk to. Where
the data comes from, that's just going
to depend on what the data source is.
But
in general, platforms shouldn't have a
hard time sort of adapting their
existing tech stacks, their storage,
their APIs um, in a way that just
enables
that surface to speak the language of
portable data.
Um, so we can try I was going to say we
could try to do a quick demo here. Maybe
what we'll do is just save that until we
get to the end and if my browser crashes
we'll I'll just talk through it. Um,
kind of explaining like some of the key
ideas within PPP.
Um, it it starts from two different
angles. How is data how is access to
data granted and then how is the data
consumed once access has been granted.
So the selection request is basically
how does a client application ask for
specific access to data. Um I want just
this piece of your data over this range
of time for this reason.
When the user authorizes access to the
data that's expressed in an immutable
grant and that's like a contract between
the user and the server that's hosting
their data and the or I guess it's
enforced by the server. It's a contract
between the user and the client
application.
And so the grant is kind of like the
cornerstone of the entire system. And
then the record model is how do you
actually model the data in a generic way
that enables grants to be expressed over
different kinds of data sources that
enables clients to ask for permission to
the data. And then yeah once data has
been um once access to data has been
granted we need some way of actually
grabbing that. Um, and within PDP it's
basically there's a standing API and if
you have um a grant then you can just
query that API to get the data.
So this actually isn't super novel in
the sense that
other protocols already take advantage
of um sort of building on what works
really well on the internet. OOTH is
sort of the dominant authorization
protocol and two notable protocols or
standards are SMRT which um are it's a
standard for medical records and open
banking in the UK is a standard for
banks and financial records and in both
cases these standards have have chosen
to basically profile OOTH meaning that
OOTH handles all of the authorization
the user um goes through very like a
typical OOTH flow where you get a token.
Um the client application gets a token
and presents that token as proof that
the user has given consent. The
difference is that RFC 9396 creates this
envelope where you can add sort of
whatever data that you want into that
OOTH consent process. Um on the next
slide, I know this might be a little bit
easier to follow.
Yeah, this one. And so basically the
idea here is you you just take a
standard authorization flow, but you
bundle in
additional information that's associated
with that consent token. And in
um smart and open banking's cases, those
would be domain specific sort of bundles
of information about what consent is
being granted like to which medical
records or to which financial documents.
In PDP's case, it's the same idea, but
it's using this more general language
about data that basically points to what
is the data source, how is data modeled
within that data source, and what
consent is is granted over that specific
um data source. In the future, other
authorization pro um protocols like GNAP
is one um could be alternatives to OLAP.
And we're architecting PDP in a way that
it's not like totally dependent on OOTH.
That's where we're starting, but the
language of data portability um doesn't
necessarily require authorization to be
based on OOTH.
And just to go a little bit more into
smart and open banking
um in both cases they express consent
semantics or their particular domain.
They also so so that's how is data
access like granted and then they also
define APIs or like a read surface where
then how do you query the data that has
been um the user has given access to.
It's important for a standard like this
to have some kind of conformance program
so that um if you're PPP
um conformant or you're smart
conformant, you know that you check all
the boxes and like we will be building
tools to to make it really easy um to be
conformant with the standard. And
in PDP's case, we're taking um maybe a
little bit of a bet that as
personal data and data portability
becomes more important in markets and
maybe in the context of regulation too,
having like a ready to go standard that
works well, that has some adoption um
will create like a tailwind basically.
Um, we think we're we're going in at a
at a good time to be ready for
increasing demand for data portability.
Um,
yeah. So just like smart um we're
defining the standard first and actually
there are already relevant um
regulations both in GDPR and in the
digital markets act that don't
necessarily point to a specific
standard. Um, in particular, the digital
markets act requires that data
portability um is provided through
continuous and real-time access. PDP is
a really good fit for that. And
obviously the GDPR um like you can do a
oneshot export, but PDP can very
feasibly like fulfill a requirement
there too. So if a platform um wants to
be GDPR compliant for example in their
PDP conformant then it's very easy to
enable that um sort of compliance out of
the box.
And then there are other projects sort
of in the same ecosystem for data
portability that are worth mentioning.
The data transfer project is by the data
transfer initiative organization and
um there's a sort of natural fit with
PPP and DTP in that PPP defines fine
grain and strong consent semantics
um out of all of my data exactly what
slices do I want to grant um access to
or what do I want to make portable and
then DTP is a way where um two data
providers can like transmit the data um
you know from provider A to provider B
sort of translating it into like a
common model. So PDP could sort of just
be used um within DTP as a consent layer
or as an alternative access layer. Um,
and then PDPA
kind of defines what exported archives
are like. You can think of this like a
standard for Google takeout. And that's
something that goes really well with PPP
because you may not necessarily want to
just query real-time access to data that
the user's consented to. You may want to
dynamically produce an archive of like
what the what the user wants to portably
export um sort of in one shot. So I I
think these like all compose fairly
well.
So this is basically um like an OOTH
like a standard OOTH authorization flow.
And now we're getting a little bit into
the architecture of how PPP works at a
very high level. Um so the steps are
basically some client whether that's an
app or an AI agent um sends a request
saying I would like to select or access
this particular slice of the user's
data. The user um sees that request
within the authorization server and
chooses whether or not to approve it.
Once it's approved, the authorization
server issues a grant and then the
client can take that grant and query for
the data.
If we go to the next slide,
um
so the question then is how is the data
fulfilled when when the client queries
for the data, where does it come from?
And a bank for example could natively
support PDP within their API. They could
handle the authorization requests and
the consent process using let's say OOTH
and they can serve the data directly out
of their database over their APIs.
On the next slide, an alternative is
let's say the bank doesn't um doesn't
have those APIs yet, but the user has
maybe a personal data server or the user
is consuming um a service from some
other provider which is happy to sort of
store the data and then the problem to
solve for is how does the data get from
the bank or the data provider into the
resource. ource server uh that the user
is controlling and that's where data
connectors come in and this is um where
the community can help accelerate the
process of making their data portable.
So I think it's worth being explicit
about what we're not solving for with
PDP because the language of data
portability doesn't need us to define
everything and by keeping things like um
storage flexible then we enable more
participation and so I think the three
sort of main concepts or components of
building out let's say an endto-end um
data portability system that PPP is not
opinionated about our identity. So, so
identity could work the same way that it
works within the bank that you log into
or the
um I don't know like the web 3 system
that you authenticate with using a key
like that doesn't necessarily have to
change with storage. Different data
sources may choose to store data
differently, may have different
compliance requirements, may encrypt
data differently. Um there may be
backups involved like PVP doesn't
require any particular storage format um
or or location. And then deployment in
terms of who's operating server, the
server is where the data lives like
that's not really the point of the user
providing consent to the data. and and
granting access to it. Um, as long as
it's under the user's control and it's
it's their data, like how that gets
deployed is sort of left unspecified.
And so I think this is a good point to
just invite as much feedback as we can
get. Like we would love to hear thoughts
about this. We have a discord channel in
the LFDT discord server. Um, you can,
it's not on this slide, but you can go
to pdpp.dev.
And we're still pushing updates to the
website. Um, so it might be a little bit
more user friendly later today. But
yeah, you can basically find us in
Discord, you can find us in GitHub, try
building a data connector. Um, there's a
personal server that you can run and
connect your own data to and connect an
AI agent using MCP or build an app using
that. Um, and these are all things that
are getting better rapidly. So, if you
have any issues, just come talk to us
and we'll get it sorted out. And if if
you are someone you know has an API with
user data behind it and would would love
to pilot
um how PDP could work with that API. I
think that would be extremely valuable
feedback and we would love to like sort
of advance that aspect of the protocol.
Um
yeah, I think I'll pass it back to Anna
um for some closing thoughts.
Yeah. Um well, I want to pause for any
questions. Um any questions from like a
technical perspective also just in from
a policy perspective or regulation
perspective if there are things that
that come to mind. Um so yeah, pause for
any questions. Um and if you want to try
to um Tim to share your screen for the
demo, feel free to pull that up in the
meantime, but I I know your browser is
crashing, so
>> I'll give it a shot.
Yeah, Sarah.
>> Um, please forgive any naive uh starting
positions on this since it's a little
bit new. So, um, I think the first thing
first thing I'm thinking about, so the
UK open banking analog is really useful
because that like we're that's in common
usage here and we use it all the time. I
think something that um comes to mind as
you were describing that Tim is like the
the difference in latency and just in
timeness of the interaction. Like in the
case of UK open banking, it's a user
generated like um exchange that um is
sending a pretty small blob from what it
feels like across uh across two uh two
entities or actors. Do do we have we
thought about that mechanic here? Like
um is it likely to be just in time? Is
it kind of backfilling personal data
from the sources over time? You know, is
is that some something we thought about
much?
>> Yeah, that's a great question. I think
one of the advantages of PDP having a
sort of active API is that to the extent
that the data can be kept fresh, there's
no reason that a consumer of the data
can't get the latest data, that data
can't like incrementally and constantly
backfill or if a provider has native
support for PDP, then if I'm granted
access to see a user's latest posts,
social media posts, there's no reason I
can't always get the very latest
information using the same consent
that's already been granted. Um, and I
think that's that unlocks a lot of use
cases that a singleshot like point in
time export of your data doesn't don't
uh wouldn't necessarily enable.
I think one thing I'd add on is there's
sort of a a um like we the standard is
pretty um intentionally underspecifies
like storage for example. So I think one
pattern um that I think would work very
well in practice is users often kind of
getting these grants themselves and then
essentially just like syncing their data
in the background and keeping it on an
environment that they control. that is
obviously more of a federated system and
involves a personal server that I think
Tim is going to demo too. And I think
what's nice about that is that then it's
almost like the user is kind of keeping
this backup of all their data and knows,
okay, I'm just pulling all of this. Um,
but there is a trade-off too where
there's just some some setup required.
Yeah.
>> Yeah. Justin,
>> sorry I'm screened off for the minute.
This is a good point that in the EU
obviously there's a great deal of
concern uh it's laid out in the AI act
in fact about data sovereignty. You
cannot use a system that's not sovereign
or controlled by an EU entity if
everybody is owning their own data which
is good. How does that work in terms of
have you thought about the ability of
governments to subpoena to seek data
from users? Is it the kind of thing
where the idea is that only a person in
a particular country at a given time can
be asked to give up their data? How are
you considering that kind of legal
element if if at all?
>> Yeah, that's a good question. Um, and
there's also um an ongoing uh pilot that
we're working on involving the EU AI act
that's going into some of this in in
more detail as well. Um so for I guess
when someone kind of is the case of
being subpoenaed I think it's kind of
similar to if you are um holding funds
with your your crypto wallet ultimately
you are the person custodying them right
so it's sort of similar to a
non-custodial wallet so I think in the
same way it would be um up to the
individual right in the same way that if
there there's something physically with
them I think this data is kind of the
closest thing. Um, at a a cryptographic
level, I guess it's like if I've granted
myself access to my data and I' I've
synced it in the background. Um, in
order for a government to subpoena that
data, they would come to me for that
grant and then they could uh either take
what I have synced or uh use that grant
to um get it from different machine,
different services on my behalf. Um, so
yeah, I guess that's sort of a
long-winded way of saying it. it it is
basically just putting custody with with
the user. Um I guess what's your
reaction I'm curious what your reaction
to that is from
>> I think that's actually smart because I
think the biggest danger you're going to
face is how the governments say we have
control over this entire network if one
person touches it. So having it be
around the only person government who
has jurisdiction over a person is the
government that is you know the polity
that person's either presently in or
residing in. And that's useful because I
think then you could use terms of
service to prevent this being a problem
with people operating in say sanctioned
countries or to a lesser extent mainland
China because the biggest issue with
this is going to be the Chinese
government trying to exert some level of
jurisdiction over everyone in the
network or vice versa western
governments being nervous about Chinese
citizens and CCP affiliated institutions
using it.
>> Yeah.
That is a very good point. Um, which can
be um, I guess ultimately someone could
be PDP compliant
just purely by running something running
their own own service. Um, and so there
is kind of a question of what this would
look like. Yeah. In in mainland China or
places where there are um, other
restrictions
>> and I think that's important because the
nature of a decentralized system is it
is decentralized. So what you're doing
by focusing it on the person who is
custody that moment in time their peace
you can say this is not in any one
country per se it's where each
individual with it is at a given time
ergo you only have jurisdiction over the
people who are currently in your
government not across the whole network.
>> Yeah. Yeah. I I feel like there's maybe
an analogy too to actually like physical
storage and like a hard drive or even
like physical phones right. It's not to
say, hey, if you you've bought this
physical device, then you now have the
right to all physical devices. It's just
kind of your piece of it. And making
that super clear. Um, yeah, good point.
Um, Tim, I'll hand it back to you for
the demo. Or if there are any other
questions, too, happy to pause.
>> Cool.
>> Um, can you see my screen?
>> So, I don't think this was in the slide
deck. I'll paste it in the chat. This is
the PPP website in case you want to try
this yourself. Um, so we have this um
self-hosting option where you can run
sort of a local personal data server.
And I'm just going to show you what that
looks like. If when I copy this and run
it and and sort of set it up, I can add
um data sources that connect my data to
a number of providers. And I'm just
going to skip ahead a little bit and
show you
um beyond that setup
that I have an AI agent that is sort of
connecting to my data. So, let me switch
Windows. One second.
Um Okay. So, this is an AI agent
basically making a request for my data.
And
um
I should have shown you this, but I'm
not going to keep flipping back and
forth. I'm basically telling it you can
access my chat GBT memories. You can't
access anything else. None of my
shopping data or other Chat GPT data. Um
so, let me just click that button in the
browser here.
Okay. So, my personal data server has
been connected and I'm just going to ask
my agent a simple question.
>> Are you showing the screen, Tim?
>> Oh, I'm sorry. It looks like it stopped.
There you go.
>> Great. Um, so this is a a new agent
session. It hasn't been loaded with any
context about me. And I'm just telling
it I have some data that's PDP connected
and it's my chat GBT memory. So what are
some themes um from the memories in my
chat GPT data? And what we should see is
this MCP server is telling the agent how
to speak the language of PDP, so to
speak. Um, we can actually zoom in to
the specific tool calls that the agent
is making. It's seeing what connectors
exist and it's seeing from the manifest
of that data source, um, what kinds of
fields it can ask for and how it can
filter the data.
So, it already found 76 memories. It's
noticing that I've worked on blockchains
and talked to chat GPT a lot about
blockchains. I've done some AI and
crypto research and there's some
personal life and financial planning
memories within chat GBT. Um, I think
what's interesting about this is
I was able to to sort of quickly pipe ve
a very narrow subset of all of my data
directly to the agent. Um, I could make
this a one-time access policy. I could,
um, support multiple data sources and
then the agent would be able to see
like, let's say, let's say financial
data across my bank and my credit card
um, data sources. So this is, I think, a
decent proof point just based on my
personal experience that it's a
significant step up in terms of the
utility that I can get out of at least
the agents that I'm working with on a
pretty regular basis.
So yeah, that's basically it. Um, any
questions about about that?
I had a question in chat but I think I
answered it as you were showing that Tim
which is it's it's the record or stream
part of PDP that defines these like
semantic trenches right memories versus
shopping.
>> Yeah.
>> So whoever is um sort of defining the
data source data model is making
decisions about how is that data
expressed in terms of streams and fields
within those streams. if there's let's
say a time stamp that's important like
which specific time stamp is the one
that says this is when this data like
was created. Um so there are a lot of
like hooks for the semantics um of
whoever's the owner of that data source
or the expert on that data source to
bubble up into the protocol.
I think one um nuance Tim that you
mentioned yesterday was on like how
someone could use that potentially with
like a vector database or something more
topical to then provide like granular
access to their LLM history. Can you
expand on some of that?
>> Yeah. Um,
so one extension to PDP and I think
these may even be in the website. Um,
yeah. So there are a couple of
extensions that we're sort of shipping
with PDP that are more like optional
features that a provider who's
conformant with PDP could choose to
support are like lexical search,
semantic search, doing aggregation and
these are basically extensions to enable
different kinds of queries over the
data. So, if I want to be able to ask um
an AI agent a question like um tell me
about things that make me happy, like
the word happy could be used in a
semantic search across all of the data
to see what comes out of that. Um
other extensions would be possible on
top of the the core protocol. Like we're
trying not to be too opinionated about
where that could go. You can imagine a
potential future where the authorization
process is really sophisticated and the
consent that the user gives is almost
like
um a subjective policy like don't grant
access to anything that's too sensitive
and then maybe some agent is a part of
the authorization flow to figure out
what that means. Um, so yeah, I think
there's a lot of potential to explore in
different directions there and the
protocol itself is trying to be
unopinative about those.
>> Daniela, I'm curious to hear like from
your early data portability work, um, I
feel like there's kind of this spectrum
that we're navigating right now, which
is basically like how opinionated to be,
right? because we don't want to be too
opinionated with PDP in the sense of we
do want to make it relatively easy for
um different platforms to adopt and we
need to have kind of a sufficient
standard that it is um a shared
language. Um I think identity and
storage for example are two that um it
is tempting to be more opinionated about
to say users should control their
storage or something like that but
instead saying you know what actually in
practice this can work without that. How
did you navigate that at the time? And
like I guess just any learnings to share
on those trade-offs or or broader
learnings to share on some of the early
data portability work.
>> Um I think we weren't opinionated enough
honestly.
>> Okay.
>> And what it led to is, you know, the big
uh providers uh essentially doing what
they thought they wanted to do, doing
what they wanted to do. I don't think
there was enough of of a a combined
um yeah a combined force. Um and we
backed off very early on because you
know the Facebooks and then you know the
Twitters and the Googles did what they
wanted to do. Um
>> that's really interesting to hear. Yeah.
I got to ask Lisa to wake up early for
the next one. So Lisa is at the data
transfer initiative the what the project
that Tim mentioned and
>> they're the what folks are funded by
Apple and Meta and Google on yeah data
transfer and so I think it it sounds
like finding that bridge and I mean the
recommendation when you say not backing
off like in a way I would think of it as
could we be more neutral so it's easier
to adopt but what I'm hearing from you
is actually it's like maybe being more
opinionated and finding ways to to
encourage adoption
>> right and bringing them along as well.
And I think we had this conversation
with her right during lunch uh as well.
But yeah, it would be great to see cuz
you know she does have um that group
does have um relationships with the big
uh providers, right? So
>> yeah.
Huh. That's a good learning. Do you
think the incentives have changed? Like
do you think we should expect the
platforms to look at this any
differently or do you think it's it's
kind of a similar challenge that that
we're currently up against?
>> I think the platforms have changed. I
mean look what um what Jack Dorsey you
know released this week as well. Um so I
do think that um the yeah I think it has
changed
>> good
>> because of the tooling because of the
tooling available to individuals and you
know in smaller organizations and
smaller product you know platform
builders etc.
>> Yeah.
Okay. That's good to hear. Yeah, we
should have, you know, find a way to get
to to Jack's people building that
because I think that'd be could be an
interesting uh conversation to have with
them.
>> Yeah, that's a great point. Um, one of
the projects that's using
>> and for those of you Let me pull it up
if for those of you who just don't did
not see Go ahead, Anna. Oh, I was going
to mention one of the projects that um
is using um some of the data portability
stuff built into Tavana to just port
chat history in and then they're one of
the first to be like, hey, we want to
write data back as well using the PDP
standard. They're actually built on top
of um the goose agent standard which
came out of um yeah, some of Jack's open
source work as well. And so I think
>> which is now at the Linux Foundation
>> um Aif artificial uh yeah AI artificial
intelligent agent uh what is it called
the agent agentic AI society there's too
many acronyms in my head
>> there are [laughter]
>> so it's buzz is the name here I'll put
>> yeah I can probably get I can Yeah, I
can get
the folks there to look at it.
>> Yeah, that would be awesome.
>> I'm just looking at the Okay, cool. And
I guess is the premise I actually I
think Matcha had [clears throat] sent
Buzz into our our team channel a while
back. Is the premise that like basically
as an org you're hosting all of your
information and this gives you like more
control over it?
>> I believe so. Here's there's a couple
there's lots of articles, but um
the website is buzzy. The [laughter]
article gives me more information.
>> Yeah, I'm like looking at the website
like, okay, yellow gradient. What does
it do?
>> Um
awesome. Well, I think um that covers
everything we wanted to cover in the
first session and we have all of August
to go in deeper. Um thank you everyone
for joining. Uh I think that having kind
of all these different perspectives from
a product perspective, from a policy
perspective, from the Linux Foundation
perspective, um it really is so
multid-disciplinary and
interdisciplinary to try to shape a
standard around this. Um and I think
that kind of yeah, having your different
perspectives is is invaluable. So I I
hope to see you again next week. Um and
yeah, I'll just leave it at that. Hope
everyone has a good rest of their day.
Yeah, and I know we mentioned the JDC
before, so I'll just put a link on there
as well. That's happening in Geneva if
anyone is interested in coming. Um, we
do have um
a ticket a ticket link. Let me give you
the link link as well, or you could just
reach out to me and I'll make sure you
get a ticket. Um, I think we're on on
the wait list now, but um I I know
people.
>> Thank you. Thank you. And yeah, so our
team will be there and we're really
excited to to kind of formally launch
this there. So very much looking forward
to it.
>> Yeah,
let me just find that link and drop
[laughter]
>> uh there's too many links. Hold on. Find
LFT
sorry
got to go to the source. Hold on.
events.
>> I'm just looking at the the Buzz GitHub
repo and and kind of taking in an
understanding of it.
>> All right. So, there's the event listing
and on there you'll find u the
registration link. Um if you don't get
if you register um and it doesn't get
accepted, just ping me and uh I will
accept your I'll have somebody accept
it.
>> Awesome. Thank you. Um, cool. There's
your email, too.
>> Okay. Well, I hope to see everyone next
week and maybe even in Geneva, too. Um,
yeah. Thanks again for joining.
>> Thanks, Anna. Thanks, everyone. Bye. Me,
too.