Submind YouTube summaries
Thumbnail for Smart Extract explained: from page image to structured data in one step (Atlas demo)

Smart Extract explained: from page image to structured data in one step (Atlas demo)

Watch on YouTube

Video summary

The video introduces Smart Extract, a transformative technology designed to streamline the process of converting page images into structured data in a single step. Previously, organizations faced a dilemma between using cumbersome multi-step pipelines involving layout segmentation, transcription, and extensive post-editing, or relying on language models that offered acceptable but uncontrolled results. Smart Extract resolves this by allowing users to define their desired output schema first and then train a model specifically to learn and extract those exact features, effectively bridging the gap between raw images and usable data without intermediate manual intervention. The demonstration showcases the capabilities of Atlas, the prototype model introduced with this technology, which can handle diverse document types with remarkable adaptability. For instance, when presented with pages containing tables, Atlas identifies the tabular structure and fills it with relevant data such as hair or eye details. The system also excels at processing complex layouts like newspaper columns, correctly determining the reading order between multiple text blocks. Furthermore, the model is trained to annotate specific schemas that identify named entities, successfully extracting place names like Corsica, dates such as August 24th, and proper nouns like the Lord of Band from printed materials. Smart Extract further demonstrates its versatility by handling multilingual documents seamlessly, as seen in an example featuring both English and Danish text on the same page. The model reads paragraphs in different languages with high accuracy while simultaneously labeling person names, dates, and locations within the mixed-language content. This flexibility allows users to define almost any schema they require for a single image and train the model accordingly, offering nearly indefinite degrees of freedom for data extraction tasks. Currently, the video highlights that Atlas serves as a powerful out-of-the-box showcase, but emphasizes that even greater capabilities are unlocked through fine-tuning for specific project needs. In conclusion, Smart Extract represents a significant advancement in data processing by shifting from rigid, multi-stage workflows to an agile, one-step solution that grants users full control over their data extraction strategies. While the current Atlas model is presented as a demonstration tool available for feedback via a link in the comments, the broader vision involves supporting various projects with specialized Smart Extract models. The creators invite viewers to stay tuned for more details at the upcoming user conference, promising further insights into how this technology will evolve and assist users in managing complex document data efficiently and accurately.
Read the full video transcript
control or convenience. Until now, you had two options. When you had data like this, you had two options. Either the cumbersome route where you needed to go for a multi-step pipeline first layout segmentation, then transcription, and very often imposed a lot of different steps of post editing and scripting until the data went into your system. or in recent years you basically just dumped the data in a lot of language models prompted back and forth until the results were somewhat acceptable but you never had full control over it. Now we are introducing smart extract. With smart extract we're introducing a technology that lets you go from image to output in one step. You start by defining the output. what do you want to get out of the image and then train a model to learn exactly these features to be extracted in the schema that you have defined. In this example, you see this model Atlas that we are introducing with smart extract as a prototype has never seen pages like this but has identified that there are tables has filled those tables with the data there is like hair, eyes etc. You could also have different examples like tables. For instance, if you dump tables into Atlas, which is our demo model, you can see Atlas will be able to recognize the structure of the table and then extract the data in a structured form. But what if you had printed material for instance, newspapers like this newspaper page? You can also dump in newspapers and learn how Atlas is able to read your newspapers. With newspapers, very often the issue is reading order. But if you have a look at this one, you can see we have two columns in this newspaper. The first column is recognized first and then we have the second column. But what are all these colors there? With Atlas, we have uh trained the model to annotate a schema that also identifies named entities. For instance, place names like Corsica or dates like the 24th of August or names like the Lord of Band. So with Smart Extract, we are introducing a technology that lets you fine-tune models that can do this and a lot more. The degrees of freedom are almost indefinite. You can define the schema that you want to get out of one image and then train a model to extract that schema. Let's have a look at the other example. For instance, if we look at again at printed text, these are printed documents in two languages. We have English and Danish on the very same page. And you can see that the English paragraphs are read quite decently. And say the same is true for the Danish on this image and nicely it has also labeled all the different person names and dates and also place names. So with smart extract we are currently intro showcasing a demo. You have access to this demo in the comments below. We happily can get your feedback about this and would hear uh love to hear what you are saying about smart extract. Of course to say it once again smart extract in the form of Atlas which is our first model of this generation is just a showcase. There's a lot more you can do with smart extract and we will be happily uh supporting your projects with smart extract models. For now, we are giving you access to this very powerful out ofthe-box model that can already do quite a few things and even more things are possible by fine-tuning. So stay tuned. You will hear a lot more at the user conference about smart extract models and see you soon.