Submind YouTube summaries
Thumbnail for Christina Qi, Databento | theCUBE + NYSE Wired: Fintech Exchange

Christina Qi, Databento | theCUBE + NYSE Wired: Fintech Exchange

Watch on YouTube

Video summary

Christina Qi, co-founder and CEO of Databento, joins the discussion to explain how her company serves as a critical infrastructure layer for the financial world by providing high-quality market data specifically designed for machine consumption rather than human viewing. Unlike traditional providers that format data for terminals like Bloomberg, Databento aggregates raw stock price information from over 15 global venues and normalizes it into a single source of truth optimized for AI models and quantitative algorithms. This approach addresses a fundamental gap in the industry where existing data is often too messy or structured for automated systems to ingest efficiently, allowing developers and fintech startups to build faster and more reliable applications on top of this robust backbone. The company's unique value proposition extends beyond simple data reselling, as Databento has rebuilt its entire technical stack from the ground up, including owning bare metal servers and collocation facilities across the United States, Europe, and with plans for Asia. This heavy capital expenditure creates a significant barrier to entry that prevents AI labs and other competitors from easily replicating their service or "vibe coding" a similar solution overnight. While many startups focus on alternative data like satellite imagery or earnings calls, Databento concentrates on essential market data because it is a non-negotiable requirement for survival in the financial sector, leading to exceptionally high customer retention rates and making the company a preferred partner even for sophisticated institutions that previously built their own internal pipes. Financially, Databento has achieved a rare position of profitability before its recent $97 million Series B funding round, which was secured not out of necessity but to accelerate global expansion and complete their product roadmap. Their business model is flexible, catering to both the "buy data, not sell it" reality of the industry by offering usage-based pricing for individual securities and scalable monthly contracts for growing users. This strategy allows them to serve a diverse range of customers, from retail traders and students learning to code to massive AI labs that rely on their proprietary infrastructure, effectively bridging the gap between early adopters and enterprise-grade needs without compromising on data quality or accessibility. Looking ahead, Databento plans to utilize its capital to expand its physical footprint into Asia and deepen coverage in Europe while also broadening its asset classes to include foreign exchange, fixed income, and cryptocurrency markets. Christina emphasizes that despite the current hype surrounding AI, her company is not an AI-native startup but rather a traditional infrastructure builder that happens to be highly compatible with AI agents by making their documentation and data easily accessible for them. Her ultimate goal is to avoid selling the company prematurely to preserve the unique culture she has built, ensuring that the team can continue to innovate and serve as the reliable data foundation for the next generation of financial products without rushing to exit the market.
Read the full video transcript
Palo Alto Studio Connection Silicon Valley and Wall Street. I'm John Fost here with Dave Volante, my co-host. Welcome to the Cube studio here at the New York Stock Exchange. I'm Jim Allen, co-host of NYC Wired Fintech Exchange, and we talk constantly about AI transforming finance. But there's a much less sexy question underneath all of it. What data are you actually feeding the machine? Because a stock closing at $102 tells you almost nothing about what happened on the way there, who was buying, who pulled the order, how deep was the market, what happened in the milliseconds before that price moved. Joining me now to have a conversation about how that infrastructure is being built is Christina Chi, co-founder and CEO at Data Bento. Welcome, Christina. >> Thanks for having me. >> So, help me understand data bento. First of all, I love the name. I mean, it certainly makes sense to me. I was just planning bento boxes for my children this week, but help me understand what it is that you guys are actually building. So, we strive to be the single source of truth for basically financial information. So, we're a financial market data API provider essentially is what we do in a nutshell. Um, we aggregate financial stock price information from various exchanges or venues across the globe including the NYC. Uh and then we basically normalize that data and distribute that to currently about a 100 thousand or so users around the world. >> Wow. Okay. So help me understand who is using that data like for what purpose? Are you creating data for quanters for like I'm guess probably not retail um customers right I assume that it's probably quite a technical to technical interaction. bring it to life a little bit like who is the key customer? Who did you build this with in mind? >> So we actually originally built this product for um quantitative traders. It could actually be for anybody honestly. So we do have a lot of retail traders who use the product. We have students who use the product for people for the first time who um have you know never actually touched a quantitative you know maybe not maybe don't even know how to code actually um who can use the product as well. but then also people who are very sophisticated including some of the largest financial institutions in the world who use the product as well. So it does range very greatly um between the different types of users um and I think recently what surprised us is that we do have a new type of user that uses the product. We have a lot of AI labs that have also picked up on the product. So, um, yeah, I think that's probably been a bigger surprise for us is that we didn't design the product originally for that type of category. But, um, that has been a big like I guess like tailwind for us in recent times. Yeah. >> So, give me a scenario here. You are a quant trader and you want to understand all I'm looking at the board here. Nvidia stock buys between 9:30 and 10 a.m., right? That information isn't necessarily easily ingestable. >> You guys can actually kind of solve for scenarios like that. bring it to life like how does it compare to the world of the Bloomberg terminal for example that we're all so familiar with here on Wall Street >> for sure. So the difference and the reason why we started data bento to begin with is actually because um I ran a hedge fund before data bento and when we were dealing with the previous generation of data providers we noticed that they were designed for human consumption essentially so like they were designed around uh for a human you know basically staring at a terminal like basically product [snorts] and we noticed that even today like most data is being consumed by machines not by humans and so we wanted to design a data provider that was meant for machine consumption, not for human consumption. Um, and that way we could basically help be the data backbone of any kind of product and help other products launch faster, go to market faster. You can build products on top of us. So, think of us as like we're designed for the builder in mind rather than for um an end user. So, you know, if you pull up a stock price on an app today, like any app on your phone, um that data might come from us already. So, that's also what we do as well. We do serve a lot of fintech startups that use us to build like apps and products on top of us. Yeah. >> And if you're, you know, any sort of enthusiast and you pull up data on your phone like the Nvidia stock price, right, you don't necessarily get the underlying data or the story or the trajectory, what brought it there like you know what that kind of broader story looks like. Is a lot of this around kind of piecing together that trajectory like that broader story like and where is the value in that data like where is that realized? >> Yeah. Yeah. So what's interesting is we decided to hyperfocus on just what's called what's what's called market data which is just the financial stock price data um instead of uh alternative data which would be like um let's say satellite data on the parking lots or um shipment data or um you know earnings calls or things like that that are more broad that might be surrounding that stock price where where it might go or where it could head stuff like that. Um and the reason why we decided to focus on just market data is just because one everybody needs market data. Every company needs that within our industry in order to basically survive. Um we've even had you know we have a lot of customers who've invested in our company for example or want to invest in our company because they're like if you die we die. Um so data market data is like a critical need for a lot of these companies. Whereas uh if you go branded to alternative data it's just less critical of a need for these guys. And so it's harder to sell alternative data and convince companies that they need that type of data. And so it just came purely down to a preference on like what do we want to focus on? Um and plus my specialty is more in the market data realm. But that data not saying that alternative data is not important. Um it's just that that's just what we chose to focus on. Yeah. So maybe this is and I'm sure it is a false assumption but in this world of AI where we're hearing more and more about data accessibility about the democratization of data and information. >> I would have assumed that like a lot of large quantum houses are building their own pipes building their own platforms. What is actually happening though because clearly there's a real market need for this >> you know what sorts of conversations are you having with the buyers across those industries? Yeah. So, we were surprised by the AI labs like using us to begin with. I think um the reason why they're using our product um one is that we're a product that AI can't replicate very easily because to give you guys a sense of maybe the differences. Um we we're not just simply reselling data. We actually rebuilt the entire stack from the ground up. All the infrastructure behind the data as well. So that comes down to even things like the bare metal servers, the collocation, all that infrastructure, we actually built rebuilt I guess from the ground up. And so as a result of that, because it is a lot of hardware behind the scenes, um you can't just vibe code that kind of company from the ground up. And so as a result of that, um a lot of these AI labs end up using us, uh as a customer actually rather than um from, you know, just doing it themselves just because it's so much easier to use us from the perspective of a customer. So Yeah, that is fascinating and that is not something I would have guessed when I looked at this company. So, you're not necessarily another cloud license spend on your end or even a frontier love fund. It's something much deeper. >> What so talk to me through the company like what's the footprint like then? You know, how many of you guys where do you exist? Right. >> What what does the you know data's actual hardware/s software footprint look like? >> Yeah. So we are colllocated um currently at various uh you know venues kind of globally. We've uh signed across I think about 15 different venues currently uh globally and uh or 50 different physical locations but in terms of like actual you know venues or so it's probably going to be about 60 different locations or 60 different venues sorry um in globally. So hopefully that will give us coverage across US, Europe, and Asia uh in the let's say nearish future. So that's something that we're really excited about just in terms of our roadmap. Um currently we only have US coverage as well as a little bit of European coverage. So what's fascinating is that we do have about over 100,000 users on data bento right now. But our product is incomplete, meaning that we only have, you know, limited number of data sets. And um the more global coverage we have the more likely we are to actually gain more customers because a lot of customers need global coverage in order to you know even switch over to us as uh you know as their primary data source. So >> and what is the kind of chicken and egg formula for the data sets like where is that information aggregated from? How do you collect and build on that? >> Yeah. So it's aggregated from basically these these uh socket changes. Yeah. Like NYC um where we have to sign you know basically licenses to redistribute this data. we do have to you know pay for that that license as well as aggregate historical data as well. So there is an initial fee you know to or we have to pay that cost right to be able to basically get all that historical data dump to come in. Um we have to also you know have the servers to be able to capture that data to be able to store that data as well. So it is quite costly to be able to do that initially. So we think of it almost like an investment. Um thankfully one thing that's really nice is that we do have a lot of customers who are willing to pay us in advance now to be like hey if you go out to this venue we'll pay you in advance as like a day one customer so that once you have that data you'll have like a day one customer to will be willing to buy the data from you. So that eliminates that risk for us >> and I mean that type of market conversation also shows the appetite and the need right that there is >> in ensuring that you do have data access in the right places and at the right end points yeah >> um five years from now. So talk me through the tech here. I mean it sounds like you guys are building something pretty unique. You know how proprietary is everything you're building? You mentioned the Frontier Labs are customer of yours. Are you a customer of theirs? How is this build happening? Especially at the speed you are alluding to there. >> Yeah. Um yeah, the tech behind the scenes is quite complex actually. That's why um I guess that's why we have so many customers across those different industries. Um, and it's also why there's not a lot of uh, it's interesting because like I guess the best way to put is like there's a reason why we're we were a lot of in the VC world, we were described as a meme stock in the VC world in sense that we were, you know, we raised around to funding recently, but we weren't trying to fund raise. Um, the VCs came over to us and they're like, you know, begging basically to invest in us and we were trying to figure out why. And it was actually because we were one of the few non-AI companies that was doing very well. Um, and we were we were like, "Wait, there's so few like non-AI companies really. Like, what what's going on?" Um, it turns out that there's just not a lot of good non-AI companies. And then even comparing us to an AI company, in AI, there's so much competition in the startup space right now. There's just so many AI companies and the competition is really fierce. But in the market data space, there's not as much competition. And a lot of it's just because there's not a lot of people who have the knowledge to be able to build that infrastructure from the ground up. Um and then also like if you think about datab we have really good retention of customers as a result of that just because um there's not a lot of other alternatives to really go to as a result. Um a lot of the startups in this space um are more like focused on retail um which is also a really great market by the way like there's a lot of really great providers but usually when a retail company or retail trader when they're ready to take it to the next level they'll usually switch over to data bento when they're ready for like a more powerful solution. And so, um, that's who we usually cater to is like when someone's ready for something just slightly more powerful. Um, >> I mean, that is a fantastic problem to have, right? VC is coming to you trying to throw money at you. How much did you raise? What stage are you guys at? >> Yeah, so we raised a $97 million series B. Um, and that was led by NEA. Um, yeah. And that was that closed um just earlier this year. Yeah. And um so give you a sense of why I mean we actually we didn't need to raise the funding. We actually broke even um earlier this year. So thankfully we're in like a strong financial position which is kind of rare for startups these days. I feel like a lot of startups are losing a lot of burning cash, losing money. Um thankfully we were pretty good about like how we spend and so thankfully we're in a good position. Um but we figured like our product is incomplete actually. So we were like let's take the money because we do need to build out that data center presence globally and we do want to complete that product. Um the other thing we thought about was also a lot of data companies at the series B stage in our industry they tend to exit by the way they tend to get acquired um around the series B stage we decided not to uh we were like let's continue building because we feel like we're still at the start line of our product we haven't run the race yet um just cuz our product is incomplete so we're like let's build out the product and I feel like our customers would be so disappointed in us if we sold the company at this stage so we're going >> and I love to see a female founder as while fighting the fight and I have to tell you it's fantastic. So, I want to go back to something you said though, which is like, you know, we're not an AI company, right? That's an interesting line to lead with in 2026 because like you said, everyone is now an AI company, right? But when we boil that down, what exactly do you mean by that? Because I'm sure there is a lot of AI built into your technology, right? And especially as it relates to discoverability and data accessibility and all of those things. So how do you differentiate between you know what you just described there and an AI native you know SAS startup or say or any type of startup. Yeah, I mean like that's actually true, right? For example, most of our doc site is being read by AI today and that's something that we have to be aware of. Um where as for a lot of the incumbents, their entire doc site is offline. It's like a PDF or it's an attachment. They have to email you the docs. For us, um we're like we recognize that AI is the primary consumer of our documentation. And so we're like, let's put it online. Not only is it online, but let's make it AI friendly so that an AI agent can read that instead. So that makes it easier for our end user. So, we're trying to do things to make it easier for our customers behind the scenes so that when our customers do use choose to use AI, um their life will be easier, right? That that being said, our product is not an AI product, meaning like, you know, we're not building agents ourselves or we're not doing a bunch of um you know, we're just not ourselves in the AI industry. We're not actually consuming like tokens ourselves, which is really hilarious. >> That is Yeah, we're not a [laughter] company that like consumes [gasps] tokens behind the scenes or anything. Um it was funny cuz I we actually did a startup accelerator. Um and during the accelerator um pretty much almost all the other startups in the accelerator were AI companies and they're talking about tokens and I remember I was the dumb one in the like the dumb founder. I was like what's a token [laughter] and they were like they couldn't believe that I was the only one that didn't know what a token was at the time because I just was that was how out of the loop I was in terms of the AI space because we've been building this far along without ever consuming a token. That is absolutely I have to say from where I'm sitting on a nicely wired truly unique right like you're probably the only person who has said that maybe to me anyway certainly ever and I kind of love it to be honest maybe the traditionalist in me is like go Christina but how are you monetizing this what's the business model here I mean we're very familiar with the world of Bloomberg terminals and how extremely lucrative that life is for many >> how do you make money off the back of this >> yeah so we uh have two different types of models when most users come in, they want to try it before they buy it, right? Um, and in our industry, there's a comment saying, "Data is bought, not sold." I'm not going to call a leader in this industry. I don't I'm not going to call Ray Dalio and be like, "Did you know you need data?" And then he's not going to be like, "Yes, I need data." Like, that's so unrealistic of a scenario. Um, and so instead, what happens is when someone on the team needs data, they're going to look for a data provider that services their use case, they're going to come in and they're going to they're going to buy it. And so we re recognize that that's the most common use case. Um, for example, if SpaceX, there's a company that has an IPO one day, right? Suddenly, everybody needs data from that company. And so, they're going to come in and buy that data. And so, data is bought, not sold. And we recognize that people want to try it, they want to buy it, and then that's the whole model. And so, we have a usage based model like that where um that first they want to pay on a usage basis for it might just be a single stock, a single security, one product. And then after that when they want to scale upwards they don't want to pay for you know they want uh pricing that's more scalable and they want certainty in the pricing and so instead when they scale up they do want to pay a monthly fee and so then after that we can have a monthly contract or an annual recurring contract that is more standard with the standard you know B2B SAS subscription kind of model. Um, and so we have it both ways. So users can pay on an a pay as you go kind of plan, usage based, and then they can scale upwards into a more a traditional um plan once they're ready to go for that upgrade. >> Okay. So prove the value, lock them in. >> So Christina, you mentioned 97 million. I mean, you also described a very capex heavy footprint there and certainly in terms of what you're building out. What does the next kind of year or so look like for you and the team? What are you spending that 97 million on? Close us out with the the journey ahead. >> Yeah, so for us, um we're really excited to finally expand out to Asia for the first time. So getting Asia coverage, expanding more out to Europe as well, and then also expanding out to different asset classes, too. So we've had a lot of demand for FX, for fixed income, for crypto, um so for different asset classes, um also for event contracts as well. So there's just a lot of demand for different things. we're hoping to be able to fulfill those demands as soon as we can for from all of our customers. Um, and then also just continue to build and, you know, make our customers happy. So, I think that's the number one thing is just continue to fulfill that mission of being that single source of truth for data. And, um, and also for me, I think a big thing is also just making sure that my team is happy. I think um, you know, big reason why we didn't sell the company is also I feel like my team would be really disappointed in me if we sold the business early today and um, called it quits. you know, I think feel like um no one's quit the company in so long in such a long period of time that um if we were the first to quit the company and just like leave, then the team would be so sad. And I'm like, okay, let's continue staying so that we can continue building and hopefully expanding the size of the team, hopefully without like ruining the culture. So, I'm I'm like a little nervous about that, but hopefully we'll get there and be able to still keep this culture. Well, I certainly hope you go the whole mile here, Christina, and raise as much money as you can and maybe one day ring the bell here at the New York Stock Exchange. God knows we need enough another female trailblazer like yourself to do that. So, thank you so much for joining us. >> Thank you. I'm Jean here at the Cube Studio at the New York Stock Exchange. This is Fintech Exchange, one of our shows with NYC Wired. Thanks for watching.