Submind YouTube summaries
Thumbnail for Transkribus Webinar for Beginners (English)

Transkribus Webinar for Beginners (English)

Watch on YouTube

Video summary

Transkribus is introduced as a powerful platform designed for transcribing historical handwritten and printed documents using artificial intelligence across more than 50 languages. The webinar outlines three primary workflows to assist users: quick text recognition through drag-and-drop functionality, guided processes utilizing public models, and the option to train custom models for specific needs. Participants learn how to upload documents into collections, manage individual pages, and initiate automatic transcription using pre-trained public models available in free plans or super-models intended for mixed scripts and languages under the Scholar Plan+. The editing interface allows users to view images and text side-by-side for easy correction of errors, adjust layouts by adding text regions or baselines, tag entities such as names and dates, and apply structural tags. Once processed, results become searchable within Transkribus Sites and can be exported in various formats including images, PDFs, or XML files. Achieving high accuracy involves an iterative process where users significantly lower the Character Error Rate by training new models using corrected pages added to the dataset as ground truth data. While printed text generally requires fewer words than handwritten material, model performance depends heavily on document quality and complexity; if public models underperform despite filtering and setting adjustments, users can test multiple versions via history or refine advanced parameters before deciding to train a custom model. Custom training allows for configurable base models, cycles, and stopping criteria, with ideal accuracy targets set around 90% CER below 10%, though lower error rates indicate superior performance. Training graphs visually demonstrate improvement over multiple cycles as the system learns from corrected data until high precision is reached. Beyond standard transcription tasks, Transkribus supports layout recognition for unique document formats and field-specific extraction including tables, with recognized content publishable to dedicated sites. The platform itself was developed by Readcorp, a European cooperative founded in 2019 with over 250 members operating under the principle of "purpose before profit," maintaining headquarters and data centers within Austria and the EU. With approximately 25 team members, Readcorp encourages users to explore new features through their help center or upcoming webinars focused on datasets and projects. Future engagement includes an annual user conference at the University of Passau centered on AI equality, alongside active community interaction via social media channels like LinkedIn, Instagram, Facebook, and BlueSky.
Read the full video transcript
Good evening. Welcome to our beginners webinar today, Wednesday afternoon. Happy you all are here and are joining us in today's session. We will record the webinar, so you can also rewatch it if you like. We will also provide the slides as PDF, so you can also take a look at them later, and feel free to drop questions in the chat. We try to answer them in the chat or here live during the webinar or also at the end. There will also be a little bit of time to address your questions. So, I will just start with showing you a bit what you can expect today. So, we have prepared a little agenda for this webinar. We will, after a brief introduction, jump right into the platform and introducing Transkribus and the workflows in Transkribus. Um first, we will take a look at how to upload documents. Then, we will uh look into the text recognition itself, how you can edit and export your results, and um then we will also have a little outlook and a little Q&A session at the end. Just so that you know who is here uh with with you today and with me as well. So, um Sonja is here, one of our customer success managers. She will later tell you more about uh how to train models, how to work in Transkribus. Then, we have Helene, our marketing and comms lead. She will start the Transkribus session in a moment. And um my name is Michaela. I'm part of the board of directors of the Readcorp, the yeah, the corporation behind Transkribus, and I am responsible for our go-to-market area. So, everything that uh has to do and uh work together with customers and our users. Um we will take you through the session a session today, and I will start with a little introduction about uh Transkribus really general overview um into what Transkribus can do. Um we designed Transkribus as a platform to um make it possible to work on historical documents um as simply and uh easily as possible. So, you all know that uh working with historical documents, especially trying to read them, can be really time-consuming and uh yeah, also um a work that uh can take a lot of time. And uh we want to provide the best tools and platform for these kind of tasks. Um Transkribus comes from uh the typical text recognition or is uh special especially known for text recognition and transcription. So, really focusing on documents, handwritten and printed printed documents, and with the help of AI I'm being able to read these texts uh in a matter of seconds. So, uh per page, we can cover over um 50 languages by now with different fonts and also yeah, spanning a wide range uh through the centuries um with when we look into the material that we can uh transcribe with our models. Then, as a second layer, we also have the layout of possibilities to recognize layout in in Transkribus, so automatically detect the different structures a document can have like columns, tables, um >> [clears throat] >> as you can see here also in the examples. Um but all kind of of structures that um you can think of when it comes to um historical documents. Another um layer is also the tagging and extraction of entities. So, you are able to um yeah, as I said, tag entities, also extract metadata. You can flag, for example, cities, places, um dates, if you like to and with this create your structured data set. And this can be done directly from your transcribed data. Another um yeah, useful advantage advantage is also being able to search and if you like publish your um results, so you can um do a full-text search across documents and collections. You can search um yeah, names and places or other words if you and and yeah, depending on what you wish to find and um we also have a possibility to um publish these um results and um data via our product called Transkribus Sites, so really um nice feature where you have a where you can create a website on your own to display your documents together with the uh with the transcription. Then, we will yeah, what Transkribus also provide is uh a nice editor where you can work on your documents side by side, so image and transcription side by side. We will look into this in a moment. There you can do, um, all sorts of things like editing the text, correcting text, and also layout, mark, uh, the entities you would like to have. And, um, yeah, use this interface, um, uh, yeah, for your needs. And, um, the heart of a one one of the heart in Transkribus is, of course, being able to train your own AI models. So, really on your own data, for your specific use case or documents, um, you can train your own models. And we have, at the moment, over 300 community models based on millions of pages that were already Yeah, we're already trained and are public also to use. Our goal, um, today is, with this webinar, that you will be able, um, to upload and transcribe documents successfully after this session, that you also know how to pick the right model for your material, for your specific documents, and, um, maybe also languages that you work with, and to understand how you can manually edit, um, results, and also when it is maybe necessary to train your own model, and how you can do this. And with this, I will give you the floor, Hélène, to make a quick first introduction into the platform. >> Thank you so much. Uh, I'm just sharing my screen now, and yes, so we'll get into how we can make your work with historical documents a breeze, and how everything that Michaela so kindly explained actually looks like in practice. Um to do that, to make it work for us, we have three flexible workflows in Transkribus. We have level one, which is the quick one, the quick text recognition, which is a workflow where you can basically drag and drop your pages into our quick text recognition tool. Then we have the level two. This is for a larger data sets or a larger corpus, where you can use our public and pre-trained models. And then, like Ms. Michaela said, we have the third uh where you can train the custom models uh with Transkribus. So, the quick one is really uh the quickest result. It kind of similarly works to translation tools, where you just drag and drop your page uh into the quick text recognition tool, and then you already see the result right next to it. Um but what we're focused on today are level two and three, the guided one. Um it's a bit more elaborate, uh but the information and bit more um helpful for larger amounts of text, if you have it, and you can really uh adjust it to the documents that you're working with better. And then, of course, we have the custom workflow with the training that uh Sonya will have a look at later. So, Transkribus is really designed uh to work with different amounts of of documents, if you have just one page or a larger amount of documents, so it works for corpus of any size, large uh small and large ones. All right. Now that we have seen what Transkribus can do, we will start with the guided workflow. And we'll start right at the beginning with the upload uh uploading and managing your documents. Before uh we look at how you can upload your files, uh we want to explain how the documents are managed and kind of explain the document structure first, so you know exactly where your documents and pages live within the platform, and you can easily find them later. So, in Transkribus, everything lives inside of a collection, and you can think of a collection kind of as a digital folder, and within the folder, you have your documents. You can have multiple documents in every folder in every collection, and within the documents, there are the pages. So, you can have multiple pages in a document, but of course, that depends on what kind of material you're working with. If you work with letters, maybe your document is one letter, and that one has just one page, or a document is a book, and then it has hundreds of pages. So, it really depends on what documents you upload. And now, I think we can already jump into the platform, and we'll have a look at where you will be doing the uploading of the documents in the web app. And this is, let me go full screen, what you see once you um log into Transkribus. Uh we have seen that many of you already have an account. Uh maybe you haven't uh logged into Transkribus in a while, so this is what it looks like right now. Um and here you can see an overview. This is the landing page with your favorite collections, recent documents that you were working on, and here we have the quick text recognition that I mentioned previously. On the left side, you have a menu, and here you can see the main menu options, which is home, of course, uh then projects, AI jobs, and search. And then you have your assets, which is collections, models, sites, and datasets. So, something that is new is projects and datasets. Uh we will host a separate webinar next week, more about that later, to explain projects and uh datasets, but we'll briefly explain them today, just so you have that so that you know um what these new uh assets uh are. So, the one thing that you need to know and the only thing that you need to know about projects right now is that the projects bring together your collections in AI models or datasets together under one roof. So, what you can think or how you can think of a project is basically a digital binder for your work. So, instead of everything being scattered around the platform, you can really bring it all together into one organized space. So, before I mentioned the structure of collections and documents and pages, so you can have multiple collections, for example, in a project and even uh, custom models that you trained also assigned to a specific project. So, it's a single dashboard where you and your team can collaborate, track progress, and manage everything um, and that's everything you need to know for now. Today, the main assets that we are going to use is collections and models, and then everything else for everything else we also have separate webinars, but not to overwhelm you right now. So, now that we know the doc how the documents are structured and kind of what we are looking at here once we log in, let's see how easy it is to actually upload your material into Transkribus. And the easiest way to do that is to click on new. And then you can see already that you have the option to upload documents. Now, if you've already uh, worked in Transkribus and you already have a collection that you created, you can upload it to existing collections, but of course, if you're just starting out, we recommend creating a new collection. And then you can give your collection a name. We will call this beginner webinar July and click on create a collection. Immediately immediately you will see the upload area and you can click on upload to add your first document. You can either just drag and drop your files here or browse for them. Just going to drag three pages over here and you can see upload is complete. We see the pop-up here, sort files. They're not in alphabetical order. Do I want to sort them? It's three pages. I don't need to sort them right now, but just so you know, you do have the option to sort your pages in order. And then I will give my document a title. So, before we gave a name to the collection, but now we're uploading the pages and it will automatically create a document and we will call this English handwriting. Click submit. And the document creation has started. You can see here, uh we are in the collections overview. We're in the collection called beginner webinar July and here is the document that will be uploaded. We can see the progress bar right here as well. And we can even see here a preview of the amount of pages that is in my document, which is three pages. So, once you've uploaded your document, it's ready for processing. Um what you can also do here is manage your document. So, for example, you can uh share it, move it, copy it. You can also manage labels or add labels. This helps you for example with the work that you're doing to immediately see what kind of for example, material material you have at what point in the uh transcription progress you are um stuff like that. And if you open your document, there you can see the pages that you just uploaded and you can of course also manage the pages here um on page level. So, in just a few seconds you have uploaded your documents to Transkribus and it's ready for text recognition. So, the next step is to start the automatic text recognition. And before we do that, we'll just jump back into our presentation for a bit more information. All right. So, we have successfully uploaded our documents. Next step is using the public models. So, how do we start the automatic text recognition? When we start the automatic text recognition with public models, the text recognition converts printed or handwritten text into digital text using our public or custom AI models. In this case, in this step, we're using public models. Um so, just a quick recap. We already I think had a previous uh a brief explanation before, but public models are pre-trained AI models that are available to everyone in Transkribus. They've been trained on hundreds or even thousands of pages and can be trained for print as well as handwritten uh text in different languages and scripts. We already have over 400 public models available for different scripts and languages. We have a short preview here on the side um and also in over 50 languages and scripts. Now, one question that um new users, especially, often have, and it's a very legitimate question, very valid question, is how do I find the right model? So, I've just mentioned that we have over 400 public models available, which sounds like a lot, and it is a lot of models. Um but, how do you find now a model out of those 400 that actually works for your material? So, what you can do is here you can see a preview of uh the model space, uh the model asset space that I previously mentioned. And what you can do here is filter. So, we always recommend to use the filter option to really um reduce the amount of models or to really filter down the models that you can actually use for your material. So, when I talk about filtering, it really depends on what kind of material you have. For example, now we have um English handwriting, meaning I can sort uh filter for English language. I can also filter for handwritten, and I can even uh filter for the specific time period that my documents were written in. So, that way, once you apply all those filters, um the models that will be shown really reduce to a handful. So, starting from uh 409 models, if you apply the specific filters and really look for the type of material that you have, it will be reduced to a handful, for example, um 10 or 20, and that's already quite a lot more manageable than 409. I will show you how you can filter in a bit later uh in a bit. Once you click on the model, you will also see further information about it. Um for example, the size of the training data or the character rate. And this is again an indication of if the model might work for the material that you have or not, um because the model description not only shows these kinds of stats, but often also shows with what material the model was trained on. So, for example, if it was trained on government records or on diary pages, st- um in- additional information like that, and that can give you another indication of it how it might work for your documents or if it might be useful. The CER, the character error rate, is also also an indication of how well the model might perform on your documents. This is not a guarantee, but it is an indicator. So, the CER, the character error rate, of 4.8 means that during the training, the model showed an accuracy of over 95%, meaning if you apply this model on a similar uh on similar material, it will probably also have a good result. Maybe not exactly 95 uh 95.2 accuracy because the models were trained on a lot of material. We can see here 38,000 pages. But it has not seen your specific material that you're working with. What we also have are super models. And they are Oh, I realized that we have an error here in the heading. The super models are available from the scholar plan and up. So the general public models are available in the free plan. The super models from the scholar plan and up. And what is so super about them and so impressive about them are that is that they are even larger than the general public models and very versatile. So what does this mean? This means that you can use a model even for different scripts and languages and for mixed material. You can see here that the model was trained for six languages or on six languages meaning it is a model that we generally recommend for mixed material. Speaking of the text titan, we now even have a new text titan, the text titan two. I think it was published last week that you can also try for your material and it provides even better results than the text titan two. So once you have found a model that matches your material whether it's a public model for a specific script or one of our super models for mixed or multilingual text, you can then apply it to your documents. And we will jump back into the platform to do just that. So we have uploaded our pages and we want to start the text recognition. What I am doing to start the text recognition is select the pages that I want to recognize. You can also select the entire document, but since we're all at page level, I'm selecting the pages right here. And then you click on process with AI. And then here, we have a model that is already pre-selected because it's probably the one that I worked with last or that we worked with last. But you can also search for specific models here by for example typing in English and then you will see the models recommended that have English documents in their training data. Or you can also filter Oh. Filter here for uh the type. Again, you can select the language that was seen in the training data and filter by century. Oh. And can see here. Okay. It shows doesn't show it here correctly, but we can see that the filter was applied correctly. And then here is also a smaller overview of what the screenshot showed before. I'm going to select the Text Sichten right now because it is the one that is or best out of the box or provides the best out of the box performance, generally speaking. And now I can just simply click on start recognition. And the recognition is in progress. So, what's happening right now in the background is that Transkribus will basically analyze the layout of the page to see where there is text on the page and then convert the handwritten text into machine-readable form. Now, processing can take a bit. We can see here. It's quite busy. We're number 198 in the queue. But because uh we figured that might be happening, we have prepared some documents for you already. Let's go to our collection overview and go to our beginners webinar folder. Our beginners webinar collection. And then I'm going to go into my English records document. And we have some pages prepared here. Now, once I open the page, I can see the side-by-side view of Let me just close it here. Of the page that I uploaded. And then here on the right side, I can see the result of the automatic transcription. Reduce the font size maybe a little bit. Um and what can happen when you use public models is that there are some errors and some mistakes. As I mentioned before, the public models often see a lot of training data, a lot of references during the training that they use to correctly transcribe your documents. But it can still happen that there are mistakes because the public models have been trained on different material, but not exactly the handwriting that you're working with. So, mistakes can happen, but you can simply correct them in the editor. So, we've added a few mistakes in here to show you how to correct them. For example, if I click here in my text, I can see that the name is not Richard Yates, but it is Richard Gates. And I just use my keyboard to correct the text here. We can also see here for one part. And this way you can correct the transcription and then save the changes. This is how you can correct the transcript. Um what you can also correct is the layout uh um the layout part. So, here on the left side, we have the layout editor. Um and we can also do some changes and edits here. For example, let me maybe open a different page. So, this one we have also prepared with some mistakes, so we can show you what to do um if the layout, for example, isn't recognized correctly. So, here we can see that the marginalia wasn't recognized, or in this case, I mean, it's prepared, but yes, let's say the marginalia wasn't recognized. What can I do in that case? I still want to have the text uh uh fully recognized. What you can do is add another text region. So, this here, this yellow box is the text region, and this encloses all the handwriting. This is basically um where the Transkribus AI recognized that there is text on the page. Um the purple line underneath the text are our baselines. These are the reference points for the AI to know where there is uh handwritten lines within the text region. And you can add both the text region and the baselines manually in the layout editor. You can do so by clicking on add region, and then click add text region. Now click once to start drawing the text region, and click again to stop drawing the text region. And we can see now that there has been another text region that was added. Um one thing I think I'm going to just jump the gun a little bit, or maybe show you something right away that I was going to show a bit later, but I am seeing that the text region isn't really well um visible very well, because the page is kind of orangey, and the text region color is yellow. So, it's a bit hard to see. Um if you have um or in this case, what I can do is go to settings, because it does help to have kind of a better visual uh differentiation between the colors of the text region and the page, especially when I'm editing. And then click on image. And here in the settings, I have some visibility settings that I can change. For examples, I can choose to either show or hide baselines, and I can do the same with regions. Um, I can also change the colors here. So, because the yellow is a bit hard to spot on the orangey page, let's maybe go for go for a green color. And as soon as I click on it, you can see the region color change. And I think I'll stick with that, because you can really see the difference, uh, very well, uh, whether the or you can see very well if the text region has been selected or no. So, this, I think, is easier to work with. Now, if I want to add a baseline, I click on add baseline. I click once to start drawing the baseline, and then double click to finish drawing the baseline. There are, uh, many more other, uh, things that you can do to edit the layout, such as, for example, split the layout or merge the layout. But, I think we probably should keep going and move on to another, um, step. But, what you can do if you're interested in, uh, working more with the layout editor, is you can click here on the three dots, and then go to keyboard shortcuts, and we have a nice list of all the keyboard shortcuts that you can do, um, and what you can do when it comes to editing the layout. Now, one more thing that I want to mention here is that you can add tags to your layout, as well as to your text. So, tags basically help, um, add additional indicators or markers for your transcription. The structure tags can help you define the, um, hierarchy structure of your document. For example, they can identify or help you, um, mark if this is, for example, a paragraph or marginalia. And it gives you, or it helps provide more context, uh, and meaning to the transcription. So, to add text, we select a region, uh, then do a right click with my mouse, and then I can assign a structure type. In this case, this is a marginalia. And we can see that this is reflected right here as well. Here it is a bit smaller, but we can just simply go back to settings, and then we can increase the label size a bit, and we can see that the label is much better visible now. Similarly, you can add uh, textural tags in your transcript, and they help to identify and label label specific textural information, such as names or locations, uh, within your content that you have. To do that, make sure that the tagging feature is enabled here. And then you can just mark the text that you want to tag. You can have a person here, and then, for example, date here, and we can even add additional information, for example, that the th in ninth is superscript, and that way you can add, um, labels to your documents. Then make sure to save your changes. And this is how you can edit the transcriptions, uh, once you've used a public model. And now that we have reviewed and corrected our text, let's just say it's it's accurate, um, it's also fully searchable. So, the nice thing is, because it is machine readable, it is also searchable, and you can use the, uh, search feature. Let's go back to the overview here to look for specific information. Again, names, places, or anything else, keywords that you want to look for within the transcriptions that you created. Once you've recognized and corrected and maybe even searched through all your transcriptions, the next step, or well, it's an optional step, is exporting your transcriptions. So, exporting can make a lot of sense if you want to use the text outside of Transkribus, for example, to include it in a publication, share it with colleagues who don't use Transkribus. And to export it, you can select the document or single pages, and then click on the action button right here, and go to export. And then here, you can see the different export options. For example, images, PDF. What is also included from the Scholar Plan and Up is, for example, page XML, which is especially useful when you work with more structured documents like tables. I do think we had a question about tables. You can work quite well with tables within Transkribus. We will not get into it right today, unfortunately, but if you're interested in working with with table layouts in Transkribus, we do have a separate webinar recording on our YouTube channel for that. But yes, so you can select the export format that you want, and then click on start export, and the download link is sent by email, and you also have an overview of all your activities here in the jobs table, for example, of your upload or your exports. Now, if you're still working with the material within Transkribus, it's best to stay in the platform. That way, your progress, the layout, the version history is all preserved, but this way you can export it if if you need to. So, exporting is something that you don't have to do right away, but we recommend it to only do it once it's finished, but now you know how it works. Um now I will hand over to Sonia in just a bit, but before we do that, I just want to show you one more slide, and that is about what to do when public models fall short. So, if you have used a public model um and you filtered it quite well for the type of material that you have, for the language, for the script, for the time period, but the result still wasn't really that good, there are a few options what you can do to improve the result. We on the first hand recommend to test multiple models, so you can compare the results from two or more uh text recognition models by using the version history within the editor. This is the small clock icon in the top right corner. Let me maybe just jump right back to show you. Um this is very useful to really compare the history of um your transcription. And you can use it here. So, this is the version history, and you can select a different version, click on compare with current version. Here, this there's not much to compare. Let's maybe see here. And there we can see a lot of differences. So, this is one way how you can compare the results of two different models. What you can do as well is use a layout model first. So, for example, a baseline model like this one, like the Universal Lines model, to first recognize the structure, and then in the second step apply the text recognition model. Uh what you can also try is adjust the settings, so fine-tune the advanced settings. We have more information about what settings um there are and how you can adjust them in our help center as well. And if it still doesn't give you the result that you need, you've tried multiple models, you compared them, you experimented with the baseline models, and you use the advanced settings, in that case we recommend to train a custom model specifically for the documents that you're working with. And this is where I can hand over to my colleague Sonja. >> Thank you, Helena. Let me just share my screen. So, as Helena just said, um if everything else falls short, that if the uh public models just fall short for your material, or if you have like a lot of material, or also maybe if you have very specific um handwriting or very specific vocabulary, for example, as Helena showed us, if you want to tag places or names and they're very specific places of a specific region, for example, then um sometimes a public models might not recognize those words correctly because as I said, they are so specific, and then it might make sense to just maybe train your own model. Same goes, for example, if you have a lot of um material and it's very homogeneous material, or it's also yeah, just material that's um not very common in um maybe the region or language or something like that, then also um it might make sense, as well as there are, of course, areas, languages, countries, a lot of things where our public models maybe still you can't find the quite the right one, for example, and then also it might make sense to train your own model on your material. And um this also makes sense to then, of course, customize your model to your material, as said, for example, with text, and get the best results or the results you need, for example, in that case here. So, yeah, it really depends on your material. It really depends on your needs, on what you uh need the transcriptions for, for example, or what data you need from the transcriptions. Um and yeah, it can just make sense to train your own custom model, sometimes even maybe make it faster than using um public public model that isn't really um working out for your material. And so, custom AI models, as said, they use specific examples to help the AI um accurately recognize and understand the writing in your training data. So, the important thing is that your um the AI learns from what you train it on. So, on your material that you um trained. And how are you doing it? So, we're adding the step into the workflow that I didn't explain before, the training, before the recognition, the editing, and exporting of our material. And for this, it's easiest or it's really fast um if you already find a public model that is maybe semi good on your materials. So, if if you say, "No, okay, I still have like a 70 um percent accuracy and a 30% of errors in my material, um I need a better model because I have hundreds of pages um then you can correct that text. You can recognize uh use the public model already recognize your material once and then correct that text. That's faster than doing everything manually. So, you alrea- already will have some words in there that are correct. You already will have the fields and the baselines drawn hopefully correctly. Else, you can also pre um recognize only the fields and only your baselines and using a baseline model to only recognize the layout um first and then type in all your text in the editor. Um whatever of course works best for your material where you're fastest, but it will of course take away some time from correcting if you use um pre-trained models. And after having corrected the text, um you save it as ground truth. And ground truth is a um term that's used in um machine learning. It's training data that is accurately transcribed text paired with its image. So, the very important thing here is that for example, as you saw in the editor just before, the text correlates with the lines recognized on the image. And that every word of course correlates with this. So, if I want to have a word for example, um the word mug correct um read as mug, then I have to type in those words of course in that place, else it will not learn it correctly. Once you have created that ground truth, you have saved it, um you can create a data set. As said, data set is um one of the um newer assets in the um Transkribus interface and it's a curated um version selection of the pages that you can import from more than one collection to use as a training data that you can afterwards also have a look at it again and see okay what did I even use as training data for the model. Um as I said it's a bit bigger than that of a topic and we will have a uh special webinar again. More information on that later, but datasets in general are curated sets of pages that um can be used to train, validate, and test custom AI models. And um they are designed to um get you uh best overview of your ground truth data that can be imported from more than one collection, so you can have different collections, maybe you had different um upload steps and everything and you can put all your ground truth data, everything that was corrected, into one dataset and you have it in one place, which is way easier than it was before because before you only had everything in the collections and you you needed to somehow um fit everything together and sort your data for your ground truth data. And so this just makes it easier to have everything for your custom model in one place. And um yeah, this is how it's going to look then and you once you have the ground truth and you have set it into a dataset you can start with training a model. And dataset here will then freeze that ground truth pages and versions of your model. And if you see, which is a very common thing to do, oh my first model wasn't good enough and you then want to add more pages or you afterwards um for some reason process more pages and say, "Yeah, now I can um, add to my model because I have another hand that I want to add or anything like that." Then I can add those ground truth pages again to the data set. The older version will remain with all the pages. So, if I need to have a look at the older version, I can see, "Oh, I used those and those pages." And adding um, to a new version of the data set will mean all the old ground truth will also be in the new version and you will add new um, pages. And so, you don't have to again go searching in every one of your collections, in every one of your documents for all those pages. It will just make it easier. And afterwards also shareable if you want and have the permission to share it um, with other users once you maybe publish your custom model, which you um, can also do. So, how does training work in general? Once you have set all your ground truth in your data set, you then go to train your custom model. You choose your version and then you select your uh, data between the pages that the model is really trained on. And then the validation set, which are pages that are set aside during the training that are then used to assess the model's accuracy. So, generally it's around um, it's set automatically to around 10%. So, if I have I don't know, 100 pages, then 10 pages will be validation set and the model runs those pages against um, the trained data to see how well it performed. And that's where the CER comes out, what um, Helene already showed us before in the public models as well. So, how good the model actually is on your trained data. Afterwards we get to model setup. There you add the model description, the model metadata and then you add the training parameters. Here you can also add advanced parameters like a base model, the training cycles, or early stopping. Um base models can be helpful sometimes. There you could add um essentially one of your own older models or a public model as a base model so that data is trained with your model. Um if you don't have a lot of material, it generally doesn't make sense to use a base model, especially considering data sets where you also can use your old ground truth. It's make make makes more sense to build on your old ground truth than um use a base model except if you have a large amount of data or very specific data that is correlate in correlation with another public model or things like that. Um yeah, and then you can start your training. And now let's have a look into this because of course this is a bit um tech- technical without any further information. So, currently I am in the um this is the collection overview and I'm in I'm going to use one of the documents and in one of the document overviews and as you can see um with the green bar I have already set the all the pages to ground truth so they are corrected. They have tags, everything I want them to have, everything I want them to train on. Uh or be trained on. And so I say, "Okay, I have the whole collection for me now is ground truth. There's there isn't anything that I want won't like to use as training." So, I select all of them in this case. Maybe I don't want to use this one because there's no text here, but all the others. And then we just click on train model and select a text recognition model. You could also train um field, table, or baseline models. Those are of course if you have very specific material if it for example table material, you could first train that your layout as Helena said before. If the public models don't work out for your material, there is um adjustments you can make for example with table, field, and baseline models. But for now we're just training a text recognition model because we already have the ground truth, the layout, and everything. Now it automatically selected all pages and from this collection I could also have selected it from a dataset, but we're using collection. I want to have the latest transcriptions. Um yeah, if I don't select pages manually, it will automatically just select all of them. So I'm going to use the next step. Now it needs a higher collection so a collection where your we're using the beginner this collection when we're in where your validation set is later is then saved and also um where your model is connected to. And then let's do just a test training. And description should of course be longer especially if you want to um publish your models. It makes sense to have a more detailed description. But for now, this is the answer. We have English as the language. And I think 19th century. Um these are also important features then for if I said you want to make your model public later on so that other users might be able to use it as well and uh for them also to find it of course. We don't have any image here, or else this is of course uh is also just a visual thing. So, the next step. Now, we have the training parameters. As said, base models we're not going to use base model because we don't have that many pages, so and we will leave everything else as is. This also makes sense if you see with the first model training, um it didn't work and you already had a lot of pages and are quite sure that your ground truth should be okay, you can try around a bit with the training parameters. Basically, you can try to let it run on more cycles, so it will go more often um through the entire training set, but you know, I just Yeah, um but for now, we will leave it as is. It always makes sense to first try it with the automatic settings, and if it doesn't work, we will try different settings. So, to the next step. Lastly, you will get an overview of uh the model you're training. We are getting our pages, our language, our century, then the validation split, that's as I said, automatic, it's 10% in this case, so I have 146 pages of training and 16 pages pages where it will um the model will see how well it perform. Then, the image preservation, we didn't do anything to the training parameters, so that's all as is. Same with the training loops, and the model overview will also tell you what in this case would be the outcome, which is a specialized AI model for English documents. So, we can start training. Now, it's started, we will get to re-work that here to the AI jobs, and we can see it's created. Training can take um a bit longer than recognition, so depending also on the e-pack box that it is running on, so how often it is um training on the material. If we have 100, if we have two, 300 um then it might take longer even. Yeah, but after we have after it's finished it will show up. You will get a notification and it will show up in your own models. So, here for example, and here's also one model already trained. This is um on the same material basically with a CR in this case of 50% which is quite high. So, now let's jump into the So, the CR in this case was 50%. Here you also have an example of a CR of 38%. Sorry. Uh this means that the accuracy of the training was in the other case it showed 50%, in this case 61%. So, uh the accuracy of the training is quite high while the error rate is also quite high in this case. 48% means there's still quite some errors in there. Generally, as a thumb rule, we're thinking of about 90% of accuracy, which means about 10% of CR. So, if you're looking at your own custom model or any other trained models, it always makes sense to have a CR that's lower than 10% in your model. And uh the lower you get the better, but you won't will never get to zero. So, around 4 to 3% is really, really good already. So, this means your model will not have a lot of errors when recognizing your own specific material. And in this case, the important thing is that um simply correcting the transcription, so if you already if you have a really high CER and you want to lower it, so we want to go from this 38% to um 10%, it doesn't uh do you any any good to just correct the transcriptions in the editor in Transkribus. You have to train a new model. And that's where, for example, you have to correct more pages, then put them into the data set, and then train a new version of your model, and then the model will probably improve. Or, as I said, you can play around with the parameters, maybe change something there, then also the model will improve, but you always have to do this by training a new model, basically. So, once you have a model you're um satisfied with, or even if you say, "Okay, my model has a 15% CER, but I want to try it, I want to see how it really performs on my documents." Try out your model, recognize some of the pages that aren't ground truth pages with your new model. Don't recognize all of them at once, first try it out, that's always better. Then um if you say, "Oh, I want to lower the CER, I want to have a better model." Then you can correct the text that you recognized with your first model, again save them as ground truth pages, and create a new version in the data set, and start a new training. So, you uh do this over and over again, and at the end, maybe you get to around 6% of the CER, so you have a very high accuracy during training, then the training graph will look something like this in the picture here. And this is where you want to be for a really good or accurate model. And then again, you can train it try it out and recognize your material. And if you used, I don't know, maybe 200 pages and you have still have 1,000 pages, it will save you a lot of work and a lot of time to train that model even if you had to train that three times maybe. And yeah, as a general estimate, for printed text, you don't need that many pages, that many words. That's a bit easier, of course. But it really really depends on your material, on the quality of your material, on very on different aspects of your material. And of course, the more hands, the more difficult it is for us humans to read the material, the more difficult it will be for the AI to learn the material. So, you might need more words, more pages of ground truth to tell you train your material. So, as I said, this is a very normal part. You probably will have to do it more than once if you start training your own models. So, now you should know about how to upload and transcribe your documents successfully and easily, how to manually edit, how to pick the right model for your material, and if that fails, how to train and improve a custom model. What else can you do in Transkribus and with Transkribus? As said, as Selena said, you can also recognize layout, the baseline models if you have very peculiar layout, for example, maybe with very small lines or anything like that. There is very everything is possible basically. And then you can also recognize just the fields as you saw in the editor before with Elena um if you want to have specific just specific fields, you can mark and recognize those. Or if you have table material, you can train your own field and table models and can then also recognize um table material with structure. Maybe you need that for export the metadata to do or anything like it. And lastly, um as uh Michaela already said, if you don't want to export it, you can also publish your recognized and edited material on Transkribus sites. And you can have a look at an example sites for example in the links down below. And now I think I will give uh back to Michaela. >> Yes, thank you for this overview and I will just quickly wrap up our today's webinar um to give you also a little inside here. So now it's working. Um yeah, what what you saw um especially um focusing on texts uh models and using text public text models and also how you can train your own text model. We believe or what we are working for with Transkribus is to um build the best tools, the best platform for making your historical documents accessible, to also make it easy to um work with these documents, especially to being able to um faster transcription, um less manual work, so that you can later focus on the real content behind this, also the different use cases everyone has while working with these documents. Um Who is behind uh Transkribus? Um I mentioned it briefly at the beginning that um Readcorp is the European cooperative behind Transkribus. We were founded in in um 2019. Um first Transkribus was part of a research project and uh the coop was then founded um six uh yeah, seven years ago already. And um we are co-owned, so we have uh multiple members, over um 250 public institutions, private institutions, as well as private individuals who are owning uh shares uh at yeah, in the coop. And um we or the idea behind this is to really being able to push operation and further development um of the platform together. So, um our mission and our our vision is also purpose before profit, so we really um reinvest everything into development of the platform. And uh together with the community, we um also building the road road map um to where Transkribus should go or being developed. Um we our our headquarter is in Austria and Innsbruck. There also um is our data center. So, um we have all the data so the the servers are in the E-EU. >> [gasps and sighs] >> Um this is also quite nice to have there. Um and to give you a little um idea of also the people behind, the team size is um at 25 team members at the moment working on developing the platform. And also working together with users and customers. Um yeah, just just a little overview of the different members we have. We integrated a few logos here, so you can take a look at um I hope you uh got inspired during our webinar. We encourage you to test it, try it out. And if you need help, you can also visit our help center. The team is really busy also working on constantly updating the help center so that you can find your right support and help resources there as well. Um we mentioned it before, so all um in the first place all our webinars we had uh in the past you can find on our YouTube channel, but we will also have some new webinars in the coming weeks and months. And the next one is next week on Wednesday where we will look into the new features, projects, and data sets and how to use and to work with them. So if you're interested in this, we are also happy if you would join next week. And um yeah, another big uh conference is coming up. We have our big Transkribus user conference every 2 years. And uh this year it will be uh at the University of Passau in September. And the topic is not all AI is created equal. We have a nice program and uh also really cool speakers and uh yeah, a great community or yeah, opportunity for community exchange there. So if you like, you can also check out our TUC website. You can um of course participate on site in Passau where we would be really happy to see lots of uh faces there from the community, but you can also, if you want to attend online and listen to some of the um sessions. Yeah, a little shout-out to our socials. So, if you want to join the conversation, we are on LinkedIn, Instagram, Facebook, and BlueSky. So, you can also find us there. And um with this, we are slightly over time with 4 minutes, but I think it's fine. Thanks for for listening. Are there any questions you have? It was a lot, and we will, of course, share the the recording afterwards. So, you can also rewatch and >> [snorts] >> test your work in Transkribus. So, I don't see any questions, so I would say we can wrap this up. Thanks for listening. Have a nice Wednesday evening, the rest of the week, and hopefully we will see you next time.