Submind YouTube summaries
Thumbnail for The Data Layer AI Actually Needs to Work | Michel Tricot, Airbyte

The Data Layer AI Actually Needs to Work | Michel Tricot, Airbyte

Watch on YouTube

Video summary

The central question regarding how organizations should adapt their data strategies for the age of AI is whether they need to fundamentally alter how they store information or if current methods suffice. The speaker argues that while storage mechanics remain relevant, the most critical evolution lies in enhancing search capabilities rather than just changing infrastructure types like vector databases. Although there was significant hype surrounding specific database technologies, the ultimate goal is not merely storing data but ensuring it is indexed and structured so that AI agents can efficiently discover and retrieve accurate information first. This shift emphasizes that as models evolve over time, their ability to remain relevant depends entirely on injecting high-quality, searchable data into them. Looking forward, a new specialized data layer will inevitably emerge between enterprise applications and AI systems, though this development is expected sooner rather than later. The primary function of this emerging layer will be to inject intelligence that guides models in understanding connections across different records and silos within an organization. This approach mirrors the historical concept of semantic layers used for data warehouses but adds a crucial dimension: while human guidance remains important, these systems must also possess self-improving capabilities based on user interactions. By analyzing which answers are most relevant to users, the system can automatically enrich its internal connections and memory without constant manual intervention. The importance of this new layer is underscored by the necessity for AI agents to understand company knowledge over time; if an agent lacks context or historical data, it cannot truly comprehend current information. Consequently, there must be a mechanism to encode corporate knowledge into a novel data layer that functions effectively as memory for these systems. Ideally, this process should happen automatically, allowing the system to continuously learn and adapt based on how users engage with its outputs. This ensures that as AI models change architectures or capabilities in the future, they will always have access to a rich, searchable foundation of contextualized knowledge derived from years of accumulated organizational data.
Read the full video transcript
As you were also talking about, you know, organizations, they have cemented data, they have to look at it. Because of AI, uh should they change the way they store data? Or because of AI, they really don't have to worry too much because AI can really, you know, filter, sniff through all the data, and pick the data that it needs. Uh what I'm trying to ask you is that should organizations change their approach to how they generate, create, and store data? Or it really doesn't matter, they can continue as they are doing. >> I would say today, the thing that is the most important is the ability to search. So, if I'm thinking about new evolution of data infrastructure, I think a lot more is going to be put on search. Uh you know, we had this whole um hype around vector databases. At the end of the day, vector databases is an ability to search around unstructured data because suddenly we realized, "Oh, yeah, we can do something with that unstructured data." But at the end of the day, an agent will need to have a very, very strong ability to search data, to discover what is available to it. And that's where I see a lot of the infrastructure moving towards is it's not so much about like how do you store it, it's more like how do you store it so that it is searchable, so that it is indexed in a way that always provides the most accurate and the most valuable uh information first. Where do you see this is going? Are we heading towards a new data layer that sits specifically between AI systems and enterprise application? Or this is too early, we'll see how things will pan out. No, there there is definitely something that's going to happen, uh and sooner rather than later. Like, the models might change. The the the architectures of the models will change over time like that is a certainty as something that will not change is the fact that to make these models relevant you need to inject data. So what I think is going to happen on this data layer is how do you bring more intelligence to guide the model into or like to get the model to actually populate more information about what kind of connection is being able to do between different records between different silos and that to me is the you know a few years ago everyone was talking about like the semantic layer for for data warehouses. Like that was a way of like encoding some information in a way some type of of memory into the warehouse to make sure that everyone is is talking the same language and I think that's going to be the same pattern that's going to apply to data except that yes it might be guided by humans but there will always also be like a a self-improvement based on how users are going the relevance of the answers that has been provided by the model to just enrich this connection and that's where you know you were talking about memory a lot like this is where memory is going to become very important like Airbyte has a fiscal year that starts in February. Well if the agent doesn't know it cannot understand the data. So how do you make sure that over time all these company knowledge gets encoded as a novel layer of the data and ideally automatically.