Christina Qi, Databento | theCUBE + NYSE Wired: Fintech Exchange
Watch on YouTubeVideo summary
Christina Qi, co-founder and CEO of Databento, joins the discussion to explain how her company serves as a critical infrastructure layer for the financial world by providing high-quality market data specifically designed for machine consumption rather than human viewing. Unlike traditional providers that format data for terminals like Bloomberg, Databento aggregates raw stock price information from over 15 global venues and normalizes it into a single source of truth optimized for AI models and quantitative algorithms. This approach addresses a fundamental gap in the industry where existing data is often too messy or structured for automated systems to ingest efficiently, allowing developers and fintech startups to build faster and more reliable applications on top of this robust backbone.
The company's unique value proposition extends beyond simple data reselling, as Databento has rebuilt its entire technical stack from the ground up, including owning bare metal servers and collocation facilities across the United States, Europe, and with plans for Asia. This heavy capital expenditure creates a significant barrier to entry that prevents AI labs and other competitors from easily replicating their service or "vibe coding" a similar solution overnight. While many startups focus on alternative data like satellite imagery or earnings calls, Databento concentrates on essential market data because it is a non-negotiable requirement for survival in the financial sector, leading to exceptionally high customer retention rates and making the company a preferred partner even for sophisticated institutions that previously built their own internal pipes.
Financially, Databento has achieved a rare position of profitability before its recent $97 million Series B funding round, which was secured not out of necessity but to accelerate global expansion and complete their product roadmap. Their business model is flexible, catering to both the "buy data, not sell it" reality of the industry by offering usage-based pricing for individual securities and scalable monthly contracts for growing users. This strategy allows them to serve a diverse range of customers, from retail traders and students learning to code to massive AI labs that rely on their proprietary infrastructure, effectively bridging the gap between early adopters and enterprise-grade needs without compromising on data quality or accessibility.
Looking ahead, Databento plans to utilize its capital to expand its physical footprint into Asia and deepen coverage in Europe while also broadening its asset classes to include foreign exchange, fixed income, and cryptocurrency markets. Christina emphasizes that despite the current hype surrounding AI, her company is not an AI-native startup but rather a traditional infrastructure builder that happens to be highly compatible with AI agents by making their documentation and data easily accessible for them. Her ultimate goal is to avoid selling the company prematurely to preserve the unique culture she has built, ensuring that the team can continue to innovate and serve as the reliable data foundation for the next generation of financial products without rushing to exit the market.
Read the full video transcript
Palo Alto Studio Connection Silicon
Valley and Wall Street. I'm John Fost
here with Dave Volante, my co-host.
Welcome to the Cube studio here at the
New York Stock Exchange. I'm Jim Allen,
co-host of NYC Wired Fintech Exchange,
and we talk constantly about AI
transforming finance. But there's a much
less sexy question underneath all of it.
What data are you actually feeding the
machine? Because a stock closing at $102
tells you almost nothing about what
happened on the way there, who was
buying, who pulled the order, how deep
was the market, what happened in the
milliseconds before that price moved.
Joining me now to have a conversation
about how that infrastructure is being
built is Christina Chi, co-founder and
CEO at Data Bento. Welcome, Christina.
>> Thanks for having me.
>> So, help me understand data bento. First
of all, I love the name. I mean, it
certainly makes sense to me. I was just
planning bento boxes for my children
this week, but help me understand what
it is that you guys are actually
building. So, we strive to be the single
source of truth for basically financial
information. So, we're a financial
market data API provider essentially is
what we do in a nutshell. Um, we
aggregate financial stock price
information from various exchanges or
venues across the globe including the
NYC. Uh and then we basically normalize
that data and distribute that to
currently about a 100 thousand or so
users around the world.
>> Wow. Okay. So help me understand who is
using that data like for what purpose?
Are you creating data for quanters for
like I'm guess probably not retail um
customers right I assume that it's
probably quite a technical to technical
interaction. bring it to life a little
bit like who is the key customer? Who
did you build this with in mind?
>> So we actually originally built this
product for um quantitative traders. It
could actually be for anybody honestly.
So we do have a lot of retail traders
who use the product. We have students
who use the product for people for the
first time who um have you know never
actually touched a quantitative you know
maybe not maybe don't even know how to
code actually um who can use the product
as well. but then also people who are
very sophisticated including some of the
largest financial institutions in the
world who use the product as well. So it
does range very greatly um between the
different types of users um and I think
recently what surprised us is that we do
have a new type of user that uses the
product. We have a lot of AI labs that
have also picked up on the product. So,
um, yeah, I think that's probably been a
bigger surprise for us is that we didn't
design the product originally for that
type of category. But, um, that has been
a big like I guess like tailwind for us
in recent times. Yeah.
>> So, give me a scenario here. You are a
quant trader and you want to understand
all I'm looking at the board here.
Nvidia stock buys between 9:30 and 10
a.m., right? That information isn't
necessarily easily ingestable.
>> You guys can actually kind of solve for
scenarios like that. bring it to life
like how does it compare to the world of
the Bloomberg terminal for example that
we're all so familiar with here on Wall
Street
>> for sure. So the difference and the
reason why we started data bento to
begin with is actually because um I ran
a hedge fund before data bento and when
we were dealing with the previous
generation of data providers we noticed
that they were designed for human
consumption essentially so like they
were designed around uh for a human you
know basically staring at a terminal
like basically product [snorts] and we
noticed that even today like most data
is being consumed by machines not by
humans and so we wanted to design a data
provider that was meant for machine
consumption, not for human consumption.
Um, and that way we could basically help
be the data backbone of any kind of
product and help other products launch
faster, go to market faster. You can
build products on top of us. So, think
of us as like we're designed for the
builder in mind rather than for um an
end user. So, you know, if you pull up a
stock price on an app today, like any
app on your phone, um that data might
come from us already. So, that's also
what we do as well. We do serve a lot of
fintech startups that use us to build
like apps and products on top of us.
Yeah.
>> And if you're, you know, any sort of
enthusiast and you pull up data on your
phone like the Nvidia stock price,
right, you don't necessarily get the
underlying data or the story or the
trajectory, what brought it there like
you know what that kind of broader story
looks like. Is a lot of this around kind
of piecing together that trajectory like
that broader story like and where is the
value in that data like where is that
realized?
>> Yeah. Yeah. So what's interesting is we
decided to hyperfocus on just what's
called what's what's called market data
which is just the financial stock price
data um instead of uh alternative data
which would be like um let's say
satellite data on the parking lots or um
shipment data or um you know earnings
calls or things like that that are more
broad that might be surrounding that
stock price where where it might go or
where it could head stuff like that. Um
and the reason why we decided to focus
on just market data is just because one
everybody needs market data. Every
company needs that within our industry
in order to basically survive. Um we've
even had you know we have a lot of
customers who've invested in our company
for example or want to invest in our
company because they're like if you die
we die. Um so data market data is like a
critical need for a lot of these
companies. Whereas uh if you go branded
to alternative data it's just less
critical of a need for these guys. And
so it's harder to sell alternative data
and convince companies that they need
that type of data. And so it just came
purely down to a preference on like what
do we want to focus on? Um and plus my
specialty is more in the market data
realm. But that data not saying that
alternative data is not important. Um
it's just that that's just what we chose
to focus on. Yeah. So maybe this is and
I'm sure it is a false assumption but in
this world of AI where we're hearing
more and more about data accessibility
about the democratization of data and
information.
>> I would have assumed that like a lot of
large quantum houses are building their
own pipes building their own platforms.
What is actually happening though
because clearly there's a real market
need for this
>> you know what sorts of conversations are
you having with the buyers across those
industries? Yeah. So, we were surprised
by the AI labs like using us to begin
with. I think um the reason why they're
using our product um one is that we're a
product that AI can't replicate very
easily because to give you guys a sense
of maybe the differences. Um we we're
not just simply reselling data. We
actually rebuilt the entire stack from
the ground up. All the infrastructure
behind the data as well. So that comes
down to even things like the bare metal
servers, the collocation, all that
infrastructure, we actually built
rebuilt I guess from the ground up. And
so as a result of that, because it is a
lot of hardware behind the scenes, um
you can't just vibe code that kind of
company from the ground up. And so as a
result of that, um a lot of these AI
labs end up using us, uh as a customer
actually rather than um from, you know,
just doing it themselves just because
it's so much easier to use us from the
perspective of a customer. So Yeah, that
is fascinating and that is not something
I would have guessed when I looked at
this company. So, you're not necessarily
another cloud license spend on your end
or even a frontier love fund. It's
something much deeper.
>> What so talk to me through the company
like what's the footprint like then? You
know, how many of you guys where do you
exist? Right.
>> What what does the you know data's
actual hardware/s software footprint
look like?
>> Yeah. So we are colllocated um currently
at various uh you know venues kind of
globally. We've uh signed across I think
about 15 different venues currently uh
globally and uh or 50 different physical
locations but in terms of like actual
you know venues or so it's probably
going to be about 60 different locations
or 60 different venues sorry um in
globally. So hopefully that will give us
coverage across US, Europe, and Asia uh
in the let's say nearish future. So
that's something that we're really
excited about just in terms of our
roadmap. Um currently we only have US
coverage as well as a little bit of
European coverage. So what's fascinating
is that we do have about over 100,000
users on data bento right now. But our
product is incomplete, meaning that we
only have, you know, limited number of
data sets. And um the more global
coverage we have the more likely we are
to actually gain more customers because
a lot of customers need global coverage
in order to you know even switch over to
us as uh you know as their primary data
source. So
>> and what is the kind of chicken and egg
formula for the data sets like where is
that information aggregated from? How do
you collect and build on that?
>> Yeah. So it's aggregated from basically
these these uh socket changes. Yeah.
Like NYC um where we have to sign you
know basically licenses to redistribute
this data. we do have to you know pay
for that that license as well as
aggregate historical data as well. So
there is an initial fee you know to or
we have to pay that cost right to be
able to basically get all that
historical data dump to come in. Um we
have to also you know have the servers
to be able to capture that data to be
able to store that data as well. So it
is quite costly to be able to do that
initially. So we think of it almost like
an investment. Um thankfully one thing
that's really nice is that we do have a
lot of customers who are willing to pay
us in advance now to be like hey if you
go out to this venue we'll pay you in
advance as like a day one customer so
that once you have that data you'll have
like a day one customer to will be
willing to buy the data from you. So
that eliminates that risk for us
>> and I mean that type of market
conversation also shows the appetite and
the need right that there is
>> in ensuring that you do have data access
in the right places and at the right end
points yeah
>> um five years from now. So talk me
through the tech here. I mean it sounds
like you guys are building something
pretty unique. You know how proprietary
is everything you're building? You
mentioned the Frontier Labs are customer
of yours. Are you a customer of theirs?
How is this build happening? Especially
at the speed you are alluding to there.
>> Yeah. Um yeah, the tech behind the
scenes is quite complex actually. That's
why um I guess that's why we have so
many customers across those different
industries. Um, and it's also why
there's not a lot of uh, it's
interesting because like I guess the
best way to put is like there's a reason
why we're we were a lot of in the VC
world, we were described as a meme stock
in the VC world in sense that we were,
you know, we raised around to funding
recently, but we weren't trying to fund
raise. Um, the VCs came over to us and
they're like, you know, begging
basically to invest in us and we were
trying to figure out why. And it was
actually because we were one of the few
non-AI companies that was doing very
well. Um, and we were we were like,
"Wait, there's so few like non-AI
companies really. Like, what what's
going on?" Um, it turns out that there's
just not a lot of good non-AI companies.
And then even comparing us to an AI
company, in AI, there's so much
competition in the startup space right
now. There's just so many AI companies
and the competition is really fierce.
But in the market data space, there's
not as much competition. And a lot of
it's just because there's not a lot of
people who have the knowledge to be able
to build that infrastructure from the
ground up. Um and then also like if you
think about datab we have really good
retention of customers as a result of
that just because um there's not a lot
of other alternatives to really go to as
a result. Um a lot of the startups in
this space um are more like focused on
retail um which is also a really great
market by the way like there's a lot of
really great providers but usually when
a retail company or retail trader when
they're ready to take it to the next
level they'll usually switch over to
data bento when they're ready for like a
more powerful solution. And so, um,
that's who we usually cater to is like
when someone's ready for something just
slightly more powerful. Um,
>> I mean, that is a fantastic problem to
have, right? VC is coming to you trying
to throw money at you. How much did you
raise? What stage are you guys at?
>> Yeah, so we raised a $97 million series
B. Um, and that was led by NEA. Um,
yeah. And that was that closed um just
earlier this year. Yeah. And um so give
you a sense of why I mean we actually we
didn't need to raise the funding. We
actually broke even um earlier this
year. So thankfully we're in like a
strong financial position which is kind
of rare for startups these days. I feel
like a lot of startups are losing a lot
of burning cash, losing money. Um
thankfully we were pretty good about
like how we spend and so thankfully
we're in a good position. Um but we
figured like our product is incomplete
actually. So we were like let's take the
money because we do need to build out
that data center presence globally and
we do want to complete that product.
Um the other thing we thought about was
also a lot of data companies at the
series B stage in our industry they tend
to exit by the way they tend to get
acquired um around the series B stage we
decided not to uh we were like let's
continue building because we feel like
we're still at the start line of our
product we haven't run the race yet um
just cuz our product is incomplete so
we're like let's build out the product
and I feel like our customers would be
so disappointed in us if we sold the
company at this stage so we're going
>> and I love to see a female founder as
while fighting the fight and I have to
tell you it's fantastic. So, I want to
go back to something you said though,
which is like, you know, we're not an AI
company, right? That's an interesting
line to lead with in 2026 because like
you said, everyone is now an AI company,
right? But when we boil that down, what
exactly do you mean by that? Because I'm
sure there is a lot of AI built into
your technology, right? And especially
as it relates to discoverability and
data accessibility and all of those
things. So how do you differentiate
between you know what you just described
there and an AI native you know SAS
startup or say or any type of startup.
Yeah, I mean like that's actually true,
right? For example, most of our doc site
is being read by AI today and that's
something that we have to be aware of.
Um where as for a lot of the incumbents,
their entire doc site is offline. It's
like a PDF or it's an attachment. They
have to email you the docs. For us, um
we're like we recognize that AI is the
primary consumer of our documentation.
And so we're like, let's put it online.
Not only is it online, but let's make it
AI friendly so that an AI agent can read
that instead. So that makes it easier
for our end user. So, we're trying to do
things to make it easier for our
customers behind the scenes so that when
our customers do use choose to use AI,
um their life will be easier, right?
That that being said, our product is not
an AI product, meaning like, you know,
we're not building agents ourselves or
we're not doing a bunch of um you know,
we're just not ourselves in the AI
industry. We're not actually consuming
like tokens ourselves, which is really
hilarious.
>> That is Yeah, we're not a [laughter]
company that like consumes [gasps]
tokens behind the scenes or anything. Um
it was funny cuz I we actually did a
startup accelerator. Um and during the
accelerator um pretty much almost all
the other startups in the accelerator
were AI companies and they're talking
about tokens and I remember I was the
dumb one in the like the dumb founder. I
was like what's a token [laughter] and
they were like they couldn't believe
that I was the only one that didn't know
what a token was at the time because I
just was that was how out of the loop I
was in terms of the AI space because
we've been building this far along
without ever consuming a token. That is
absolutely I have to say from where I'm
sitting on a nicely wired truly unique
right like you're probably the only
person who has said that maybe to me
anyway certainly ever and I kind of love
it to be honest maybe the traditionalist
in me is like go Christina but how are
you monetizing this what's the business
model here I mean we're very familiar
with the world of Bloomberg terminals
and how extremely lucrative that life is
for many
>> how do you make money off the back of
this
>> yeah so we uh have two different types
of models when most users come in, they
want to try it before they buy it,
right? Um, and in our industry, there's
a comment saying, "Data is bought, not
sold." I'm not going to call a leader in
this industry. I don't I'm not going to
call Ray Dalio and be like, "Did you
know you need data?" And then he's not
going to be like, "Yes, I need data."
Like, that's so unrealistic of a
scenario. Um, and so instead, what
happens is when someone on the team
needs data, they're going to look for a
data provider that services their use
case, they're going to come in and
they're going to they're going to buy
it. And so we re recognize that that's
the most common use case. Um, for
example, if SpaceX, there's a company
that has an IPO one day, right?
Suddenly, everybody needs data from that
company. And so, they're going to come
in and buy that data. And so, data is
bought, not sold. And we recognize that
people want to try it, they want to buy
it, and then that's the whole model. And
so, we have a usage based model like
that where um that first they want to
pay on a usage basis for it might just
be a single stock, a single security,
one product. And then after that when
they want to scale upwards they don't
want to pay for you know they want uh
pricing that's more scalable and they
want certainty in the pricing and so
instead when they scale up they do want
to pay a monthly fee and so then after
that we can have a monthly contract or
an annual recurring contract that is
more standard with the standard you know
B2B SAS subscription kind of model. Um,
and so we have it both ways. So users
can pay on an a pay as you go kind of
plan, usage based, and then they can
scale upwards into a more a traditional
um plan once they're ready to go for
that upgrade.
>> Okay. So prove the value, lock them in.
>> So Christina, you mentioned 97 million.
I mean, you also described a very capex
heavy footprint there and certainly in
terms of what you're building out. What
does the next kind of year or so look
like for you and the team? What are you
spending that 97 million on? Close us
out with the the journey ahead.
>> Yeah, so for us, um we're really excited
to finally expand out to Asia for the
first time. So getting Asia coverage,
expanding more out to Europe as well,
and then also expanding out to different
asset classes, too. So we've had a lot
of demand for FX, for fixed income, for
crypto, um so for different asset
classes, um also for event contracts as
well. So there's just a lot of demand
for different things. we're hoping to be
able to fulfill those demands as soon as
we can for from all of our customers.
Um, and then also just continue to build
and, you know, make our customers happy.
So, I think that's the number one thing
is just continue to fulfill that mission
of being that single source of truth for
data. And, um, and also for me, I think
a big thing is also just making sure
that my team is happy. I think um, you
know, big reason why we didn't sell the
company is also I feel like my team
would be really disappointed in me if we
sold the business early today and um,
called it quits. you know, I think feel
like um no one's quit the company in so
long in such a long period of time that
um if we were the first to quit the
company and just like leave, then the
team would be so sad. And I'm like,
okay, let's continue staying so that we
can continue building and hopefully
expanding the size of the team,
hopefully without like ruining the
culture. So, I'm I'm like a little
nervous about that, but hopefully we'll
get there and be able to still keep this
culture. Well, I certainly hope you go
the whole mile here, Christina, and
raise as much money as you can and maybe
one day ring the bell here at the New
York Stock Exchange. God knows we need
enough another female trailblazer like
yourself to do that. So, thank you so
much for joining us.
>> Thank you. I'm Jean here at the Cube
Studio at the New York Stock Exchange.
This is Fintech Exchange, one of our
shows with NYC Wired. Thanks for
watching.