Video summary
Transkribus is an AI-powered platform designed to streamline the processing of historical documents by converting them into machine-readable, searchable, and structured formats. The webinar outlines three primary workflows to accommodate different user needs: a quick drag-and-drop method using ready-made models for immediate results, a standard guided workflow that involves selecting appropriate public or super models based on language, script, and time period, and an advanced custom model training process tailored for specific handwriting styles or layouts. Originating as a project at the University of Innsbruck in Austria funded by the European Union, Transkribus has evolved into a cooperative owned by over 250 members worldwide, distinguishing itself as the first AI cooperative where profits are reinvested into platform improvements rather than distributed to shareholders.
The tutorial demonstrates the practical application of these tools within the web app's "desk" workspace, guiding users through uploading documents and utilizing automatic text recognition with pre-trained AI models. It emphasizes the importance of selecting the correct model via filters, such as distinguishing between handwritten and printed text, and evaluating performance metrics like Character Error Rate to ensure accuracy. For cases where public models are insufficient, the session explains how to create custom models by generating "ground truth" data through pre-recognition and correction, recommending a minimum of 20 pages for initial training with iterative retraining aimed at achieving a reliability rate under 10%. Users can also manually edit transcription errors, adjust layout regions to capture marginalia, and apply structural or textual tags to preserve context, while the platform continues to expand its capabilities with specialized models for baseline recognition, field extraction, and table structures.
Beyond technical workflows, the webinar highlights Transkribus's commitment to community engagement and continuous improvement through active user feedback mechanisms. The organization invites participation via upcoming webinars, a September user conference with an open call for papers, and regular social media updates, ensuring that tools are refined based on real-world usage. For users facing unresolved questions or specific challenges with their documents, the help center's contact form and future webinars provide dedicated support channels. Recent developments include the release of field and table models, with end-to-end models expected later in the year, further enhancing the platform's ability to make historical archives accessible to researchers and institutions globally. The session concludes by noting that today's recording will be distributed via email next week, offering a valuable resource for those interested in leveraging AI to digitize and structure complex historical collections.
Read the full video transcript
Hello everyone. All right,
welcome to today's beginner's webinar.
Thank you so much for joining. Um, you
joined Transcribus probably because you
were in a similar situation where you
had a historical document with kind of
tricky handwriting to read and you were
struggling with it. had material and
documents and you wanted to extract this
information maybe and work with it in a
different way and so transcribus uses AI
to help you automatically read this
document to process it and it basically
analyzes the text and then gives you
machine readable layout so it does the
reading for you and helps make your work
easier.
What we'll have a look at today is first
an introduction to transcribus to the
software and then we'll have a look at
the workflow the standard workflow. So
how to upload your documents how to
start the AI recognition and how to pick
the right model for your documents and
we'll even have a small look at the
advanced workflow because maybe some of
you already know you can even train your
own custom model but more about that
later. And at the end uh I hope we'll
have time for questions but of course
you can always send your questions to
the zoom chat. So who is here with you
today from the team? My colleague Areno
who is working as a solution adviser is
here with me and working in chat. So as
a solution adviser he has a lot of
knowledge can answer your questions and
do your best to reply in the chat
immediately provide you with helpful
links helpful information. And then
there's me. I'm Helina. I work in the
marketing and comms team. So I usually
do the webinars. I also work in social
media and I go to events. So again,
maybe we've met at Rootste two weeks ago
and I'll do my best to uh help you get
started with transcribus today.
So let's have a look at the
introduction. How can we make your work
with historical documents a breeze? So
with transcribus what you can do is you
can use the transcribus AI for text
recognition and automatic transcription.
So maybe you've had exactly these
questions and these struggles that your
transcript is never accurate but working
with transcribers you can't maybe find
the type of model can't really type find
the type of writing and that's why you
were joined today. Maybe you have
already used it and you know what kind
of language, what kind of script you're
looking for, but you still don't really
know how to pick the right model. Or
maybe you're just starting out and you
don't even know uh what questions you
might have and you really just want to
get a first idea of how to work with
transcribers and how to get those
transcriptions. So in that case, you're
definitely in the right place today. By
the end of this session, you will know
how to upload and transcribe a document,
but also how to manually edit the
results. Maybe there were a few mistakes
in the automatic transcription and
you'll know how to pick the right model
for your material and also when in what
case it makes it makes sense and also
how to train a custom model.
Um again just to remind you transcribus
is your AI part LA designed to simplify
your time consuming time consuming and
laborous work with historical documents.
So we really try to make your work with
historical documents easier. Um in what
way do we do that? Well, with
Transcribus, you cannot only transcribe
but also search your documents because
it is machine readable text. It is also
machine searchable. So, you can use full
text search to search for specific
information within your material. You
can also use text classifications for
names, places, and dates.
And you can even recognize structure
because often it's not just about the
text itself, but also the structure that
the text comes in. and to extract the
information in a structured form as
well. Now what does this look like in
practice? Let's have a look. So we have
three different workflows. Um the first
one is a quite quick one. It's to be
able to get really fast results with
readym made models. So this basically
works in a drag and drop way where you
drag and drop pages into it kind of like
often um uh translation tools work.
Level two is then a guided workflow
where you can improve accuracy by
picking a right model. So this is our
standard workflow. And then level three
is the custom uh the custom training
where you train your own model. So the
first one is really one where you use
transcribus out of the box. You drag and
drop your page in similar to translation
tool is for quick results
kind of like this. So transcribus scans
your page and gives you a result
immediately. The second level which is a
guided one which is our standard
workflow. You can use the existing
models for your material and we will
have a look at this one today and then
we will have a sneak peek at our custom
model training and I'll show you how you
can start training your own custom
models as well. But first again we will
have a look at our guided level the
standard workflow. Now how do we start
with transcribus? So we've gone through
uh the steps, we've gone through the
workflows and now we really want to
start at the beginning uploading our
documents. Now first we will have to
take a look at where you will be doing
the uploading and the working in the
transcribus web app for some context and
on the
um uh transcribus web app once you log
in this is what you will see. So you
will see three main areas here on the
side underneath home. Home is of course
your landing page. The three main areas
are then a desk space which is your
central workspace. So here you will
manage documents and collections. And
this is also where you will start your
workflow today. Then you have the models
area. Here you can browse and select
different AI models for text
recognition. And then we have sites. So
sites is where you can publish your
collections online. But we will get more
into that a bit later. We'll take it one
step at a time. So first the important
area is the desk area because again this
is where we work. This is where we
upload and organize our documents.
Now in transcribus everything lives
inside a collection. So these
collections are kind of like folders or
projects and within these folders are
your documents. So for me example
because I work in marketing and I do
webinars one of my collections one of my
projects is a webinar. So within the
webinar I then have different
collections and within uh different
documents I'm sorry and within these
documents I have different pages. Now
you can have multiple collections with
multiple documents and the amount of
pages really depends on what type of
documents you uploaded. So if you have a
letter might just be one page. If you
have a book probably maybe hundreds of
pages. Now that we know how the
documents are structured, let's see how
easy it is to bring it into transcribus
and we will have a look right into the
plat platform and jump on over here.
So once we log in, as you can see here,
we're now in my home in the landing page
and I will move into the desk workspace
where my collections are. Here I have an
overview of my favorite collections and
I can also just move to the general
overview where all my collections are um
uh situated. So we'll move to my
collection called beginner's webinar.
But if you're starting out and you want
to create a new collection, you simply
click here in the top right corner.
Click new correction uh new collection.
Give it a name. So we'll call this
webinar test.
create the collection.
And now you can already see the area
where you can upload your documents. So
I'll click on upload.
And here you have a overview of what
formats you can upload and what size
they can be. You can here also upload
entire folders and just drag and drop
them in. Now let's see. I have prepared
a few pages
for this webinar and I'm just dropping
them in here. We can see that it's
uploading. We see the loading bar
progressing
and
we do not do want to sort them before
continuing. Enter document title. This
is then again the document
test webinar. Then we also have a name
for the document. And now we can already
see we are in my collection called
webinar test and we see the document
with the three pages that I drag and
dropped in
called document test webinar.
So just in a few seconds your material
is safely stored and now ready for text
recognition.
And now that your documents are uploaded
and organized in the collection the next
steps is to turn these images into text.
So now you will start with the automatic
text recognition. Now how does that
work? So again in this case we'll use
our public AI models to process our
documents and do the automatic text
recognition.
to start the automatic text recognition.
Um, we have to use these models because
they convert the printed and handwritten
text into digital text using our
pre-trained AI. So, these models were
trained with machine learning. And in
this case, it's important to choose the
right right approach, which is the
public models because they are
pre-trained and ready to use. They're
quite easy. So you just have to select
one and um be mindful of uh the language
that you pick. But this is the easiest
way and the right approach for the
standard workflow.
When choosing the public models, we have
over 300 public models available for
handwritten as well as printed text and
for over three uh for over 30 languages
and scripts. So in this case um it is
really important to filter for models
because we have over 300 models
available. Um the important thing is to
find the right one and this is also a
question that we got during registration
is how do I pick the right model? There
are so many public models available. How
do I know which one to pick for my
material?
And you can do that by using this filter
option in the transcribus web app. So
you can see here an overview of all the
models that we have. We have as you can
see um the last time I took the
screenshot we had 330. I think it's even
more now. Um but you see that this is
quite a large number. And you can use
these filter options on top. Oh, I'm
sorry. I think there is missing a
screenshot here but I can show it to you
right in the platform if we go to models
and public.
All right, there we go. As mentioned,
it's already quite a much higher amount
of public models that are available. But
you can simply use the filter option
right here to pick a handwritten or
printed text and even just filter by
selecting a language here. So once you
select a language
for example English
and you filter for handwritten texts and
a specific century let's say 17
to 19 you can see that the amount of
public models has immediately reduced to
just a handful. So that way you have a
much easier time of looking through the
public models that are available to you.
Let's go back. Now if you click on the
model description, you will also see
further information about the model on
the right side. So this will give you
more information and a further
indication of if the model performs well
for your language or for the text that
you're working with. So let's say you
work with a Spanish document. In this
case, you have selected the colossa
espanol. You can see here that it has
been trained on almost 40,000 pages and
it has a cer of only 4.8. So the cer is
also another indication of how well the
model might perform um next to of course
the language and the time frame. But the
CR is the character error rate. So that
means that only 4.8 out of 100
characters were recognized incorrectly.
Meaning during the training there was an
over 95%
accuracy during the training. And this
gives you an indication that this might
work quite well for the documents that
you're working with. One of the more
newer technology that we have are the
super models. And what is new and
impressive about them is that they are
very large and generally more versatile
models. So what does that mean? This
means that you can use the models for
different scripts and languages and also
mixed material. So as you can see here
we have the text titan one which has
been trained on six different languages
can see here German and then five other
ones. So the super models usually really
provide the best out of the box
performance.
Now once you found a model that matches
your material whether it's a public
model for a specific script we have had
the question about working with German
uh documents so we have a lot of German
models um it's actually one of the
stronger languages that we support so we
have one for German Kurand sitelin also
frau so we have models for very specific
scripts or you use one of our more
general purpose um super models you can
apply directly to to your documents.
Now, the progress of doing this is quite
simple, but I think to make this a bit
clearer, we will switch over to the web
app again and have a look at what this
looks like in practice. So I will move
to my desk workspace and to my
collections
and I'll move to my beginner's webinar
or actually let's do our webinar test.
Let's open our pages that we just
uploaded. So in my case I have here my
document and I want to recognize all
three pages at once.
I can do so by selecting the uh document
just here by clicking on it or I can
also open it. Now I see the overview of
the separate pages and now I can either
select one or maybe all three depending
on how many you want to process at once.
Now you just click on process with AI
and you can see this slide in on the
right side. Now here you select the
process type. So right now it's uh the
process type that is selected is layout
but I want to recognize the text. So
I'll click on text and then I have here
already a pre-selected model. I probably
have used this before um the last time I
did a text recognition. But if I
deselect it, I can see here again these
filter options. Now these are a bit
smaller um but they kind of uh work the
same way than in the general model
overview I showed you before. So here
you can filter for specific models if
you already know the model name. For
example, here I have a few of my
favorite models. Or you can go through
the model list. Again, you can use the
filter to search for a specific language
um handwritten or printed text. And this
way um select the best model for your
material. So again when speaking about
best model for your material it's really
important to consider what type of
documents are you working with what kind
of language is it? What kind of script
is it? What kind of time frame is it
written in? And then select the model
that fits exactly this document type the
best. This gives you uh the best
possible outcome of the automatic
transcription. Now in my case I will now
use the text titan.
Again, these are usually the models that
provide the best out of the box
performance. And now I just simply click
on start recognition
and the recognition process starts which
I can see here indicated with this
scanning bar.
Now it might uh take a bit of time to
process the pages. This depends uh on
the general um um traffic I think maybe
is the right word uh that is happening
in the background. So um what is
happening in the background right here
is that transcribus is analyzing these
documents to see where is text on the
page where is the layout on the page and
then converts this handwritten or
printed text into machine readable form.
Now because this is so loading, we'll
just go back to the documents that we
have already prepared for today in our
collection called beginners webinar
and I will open my 17th and 18th century
English records.
And here you can see I have quite a few
pages. Um let's just have a look at this
one right here.
Now once you open the page this is what
you will see. You will see the image
that you uploaded on the left and the
automatic transcription output on the
right side. Now it can happen especially
when using public models that there are
some mistakes. This happens because
these models have seen a lot of training
data during their training process. But
we know that every handwriting is
different. There might be very specific
scripts. So in this case the model has
of course not seen the exact handwriting
that you are working with. So there
might be a few mistakes but we can
correct this quite easily with our uh
keyboard in the editor.
Let's see. We will go to our document to
the next page. As you can see I can use
the sidebar to navigate. And here we
have prepared some uh mistakes um ahead
of time. And if I go through the
automatic transcription, I can zoom in
here a bit. I can see for example in
this line right here that there was a
mist mistake in the transcription and it
does not say menu row, but the letter
that is written here is actually an M
and then a dot. So I can simply click in
the text where there is a mistake and
use my keyboard to correct it. If we
look further down, we can see here that
it does not say treasurer but treasurer.
So I can simply use my keyboard to
correct this.
So this way you can use the document
editor to correct the automatic
transcript, but there are even more
options to work with the document
editor.
Oh,
close this so we have a bit of a better
view. One other thing that you can
correct is the layout. So you can see
here the layout part of that document
editor. What you see here encased in
this green line is the text region. Now
the text region is the box that encloses
all the handwritten text contained in
this image. The line in the blue
underneath the text is called baseline.
And that is one of the most important
reference points for text recognition.
And what you can do here is because we
can see it has recognized quite well
this large body of text and even this
marginelia here on the side. Now again
this is something that we prepared ahead
of time. So these mistakes we prepared
but just so you know how you can correct
this and edit this. So you can also
correct the layout here on the left. Not
just the automatic transcription output
on the right, but also the layout
recognition here because we also want to
have this text recognized in the
margins. We can add another region and
baselines. So to add a region, I go here
on the left side in the settings bar. I
click on add region. Click add text
region. I click once to start drawing
the region and click again to end
drawing the region. And you can see here
I have a new region drawn and it
immediately reflects in the um text
editor part in the right side of this
editor. Now there are no lines there is
no transcription yet but we can change
that quite easily by drawing these
baselines. So again these baselines are
reference points for text recognition.
Now I can draw the line by clicking on
add line clicking once to start drawing
the baseline and continuing and then
double click to end drawing it. And you
can see here
see you see here that an empty line
appears. So there's no automatic
transcription yet but there is a line
where the text can go. We can continue
drawing these baselines. I just continue
real quick.
And finished. And we have four
additional lines.
Now what you can do, the easiest way to
add the automatic text recognition is
first of all save the changes. And then
you can start another text recognition
to automatically recognize the text
here. But of course you can also
manually enter the text. Now I have
started the transcription but I can also
just use my keyboard. In this case it
might be faster and Carlton
add the text manually myself
of the and so on and so forth.
Now you can even
let's go back to the selection mode.
selection mode.
You can even split regions and combine
them. So in this case, we have one large
body of text and we can see here this is
all region two. But if I want to have
the structure reflected better in the
transcript and I say okay now this is
one paragraph but I think this is a
separate paragraph and I want this to be
reflected. I can select the text region
and then split it by clicking H like in
a vertical cut and then clicking. You
can see this blue bar appearing on the
left side. Clicking where you want to
split the region. And now we see
that this is its separate text region.
Maybe we're saying, "Okay, it doesn't
really make sense that this text region
right here is before the marginalia on
the side." And we can fix the layout
order in the transcript by using our
layout tree here on the right side.
And I can move the regions, entire
regions,
just drag and drop it. And the order has
changed. And I can see now that this
paragraph is region three and the
marginelia is region four.
Now one more important thing to know.
There we go. That was already a teaser.
One more important thing to know is that
you can add tags to your documents. Now
what are tags?
So you can um you can add two types of
tags. Structure tags and layout uh
textural text. I'm sorry. structure text
and textural text. So structural texts
help define the hierarchy structure of a
document. So they can identify for
example titles, paragraphs, marginelia
and other structural elements and these
help maintain the original layout and
the formatting of the documents during
the transcription. So meaning and
context is preserved. The textual tags
on the other hand are used to identify
specific textual elements for example
names or dates or locations for example
and this labeling can also be useful for
creating a database and it can also help
provide context. So to keep it short
these tags serve as labels or markers
that are applied to specific elements
within your document.
So you can apply them to your layout by
selecting the region like this using the
right click on your mouse and then you
should see this overview of structure
tags. In this case I'll select paragraph
and you can see here
paragraph written in the top left corner
and also paragraph appearing with the
transcript.
Let me tag the marginelia as well. We
select marginelia.
And this way you can add uh more
information. Now for textural tags, you
need to make sure that this um
area, this toggle, this button is
enabled and it's blue. You can see here
if it's disabled, I'm selecting to
enable it. And to add the text tool
tags, you simply mark
the text in your automatic
transcription. And now you can see these
tags appearing. So I will tag this as a
date. And then we also have for example
a person here.
So this way you can add more
information. You can use these tags to
also further edit text. So if there is
text written in superscript,
you can also update it as well with this
tagging feature.
Now one other thing that is important
here is settings.
And if you open settings, you have a few
configuration options such as different
visibility to for example also show uh
line polygons or to remove the
baselines. You can see it reflected
immediately here in the image on the
left side. And you can also increase
label size or change the colors of
region or baselines. So maybe um there
is a bit of a difficulty of
distinguishing between colors. So that
way we want to make uh it more
accessible and more more easier for you
to work with the documents in the way
that works best. Sometimes also we have
for example here the highlight the
baseline highlight in orange that maybe
if your document is really quite tinted
it might be hard to differentiate from
the background. So that's what you can
change here as well. There are also
configuration options for the text,
where to center the text, how to align
the text, and we have configuration
options for tags. So here, for example,
you have an overview of all the tags
that you have in your collection. And
you can choose to disable them for these
documents. So that will not be shown
anymore. Let me show you. So I disabled
marginalia and it cannot be seen in this
overview.
enable it again and it should be back
right here. And of course, you can also
add new tags. So, if there are any tags
missing, you can click on edit tags in
collection settings and simply add new
ones. If for example, you're working
with uh let's say
what example do we have? Maybe a recipe
book and you want to add a tag for
specific ingredients or a similar thing.
Then you can simply of course manually
add uh different tags to really reflect
the type of material that you're working
with.
Now that we have reviewed and corrected
our text, so we say it's accurate, it's
not just accurate, it's also digital and
searchable. Um and because again because
it is digital, you can also use full
text search. But what do we do? Uh what
do we want to do now? We have revised it
and corrected it and now we want to
export it. So let's save our changes.
Make sure that everything that we have
corrected is saved and let's go back to
the overview of our records.
So I have corrected the transcript. I am
satisfied with it and I want to work
with it uh in a different way. So
exporting can make a lot of sense if you
use the text outside of transcribus for
example to include it in a publication
with colleagues who don't use
transcribus or and we had that question
as well in our registration questions.
Uh there was a question on how to work
with a literary corpus. So for example,
if you want to export um uh export the
different documents and work with it in
a different for example digital
humanities tool uh different software
further for analysis you can use the
export feature that way. So you select
the pages you want to export. Um let's
just do these three pages for now and
then you click on the action button
right here.
You select export and now you have an
overview of the different export
features uh export formats that you
have. So let's say for example a
document
and here you can see you can even export
the tags from your document. You can
also choose another PDF for example and
simply click on start the export and as
shown in the popup right here the job
has started and you will be sent the
exported pages via email. Now the export
options, the export formats that are
available to you um depend on the
subscription plan you have but there are
a lot of different formats uh PDF, Word,
page XML as well. Um again some are only
available starting with the paid plan
with the scholar plan such as um
exporting it as um an Excel file. So if
you're working with tables for example,
this is available with the pay plan. But
if you're still working on the material,
it's usually best best to stay within
transcribers for now. All your changes
are saved in the documents. Um and that
way you can keep all your layout, your
tags, your version history intact. Maybe
you are planning to further edit the
documents at a later date, train a
model, maybe upload more documents or
publish them. So exporting is something
that you don't have to do right away,
but now you know how it works and how
you can get the documents out of
transcribers as well.
Now let's go back to the slides.
So, we've had a look at how we can use
our public models.
But what do you do when public models
don't give you quite the result that you
need? So maybe you've tried multiple
public models. You have experimented
also with using baseline models first.
You have also tried some of the advanced
settings. We have more resources on how
to work with advanced settings in our
help center as well. But the recognition
is still not really quite there and that
is when you can train a public a custom
model. So let's quickly compare again.
We have the public models on the one
hand. They are pre-trained and ready to
use. They work very well in many cases
but they are still general purpose. The
custom models are trained on your
specific material. So they can really
learn the um specifics of the
handwriting that you're working with,
the quirks of your handwriting, also
regional styles or unique document
layouts.
So when are custom models useful?
They're useful when you're working with
documents that have those regional
charact characteristic specific
handwriting, specific vocabulary, the
individual writing styles, maybe also
demanding layouts. Um, and they are very
versatile for individual applications.
So in these cases, we choose the custom
models. The custom AI model training
uses specific examples to help the AI
accurately recognize and understand the
writing in your training data. So
training a custom model in transcribus
is all about giving the AI examples of
the text you wanted to recognize
together with correct transcriptions
side by side. So you give examples to
transcribus to teach the model to
understand the particular handwriting.
Now we do have um
our advanced workflow here. So the one
step that is different from the standard
workflow that we saw previously is
really this training part
and the training workflow works as such.
So we recommend to first do a
pre-recoognition with a public model. So
let's say you have uploaded a diary
because we'll use a diary as an example.
You have uploaded a diary and you want
to train a custom model for the specific
handwriting of the person who wrote
this. To do this you need again these
examples to train the model. Now how do
you get these examples? it takes quite
some time to manually get a certain
amount of pages. We recommend around 20
pages as a starting point to train the
model. So there's a difference of course
in the time and effort that goes into
manually transcribing 20 pages for this
training data or to use a pre to use a
public model for pre-recoognition.
Then accurately correct the text, save
this as ground truth and then start the
first training.
So the ground truth is the training data
and this training data is again
accurately transcribed text paired with
its image.
Now how to start the model training is
you use these pages. So I have a
screenshot here. We'll move into the web
app um shortly but just to show you you
have these pages of let's say again 20
pages as a starting point of correctly
transcribed material. So you have used a
public model, you have applied it, you
corrected the transcription, it is now
accurate for the training so that the
model can learn from this accurately
transcribed text can learn from this
data that you're presenting to it. You
will select these pages that you have
saved as ground truth that are
transcribed correctly and then you click
again on this action button that we used
first to export documents. But instead
we click on train model and then on text
recognition model.
And I will show you what that looks like
in the web app right here.
Now
we will use in this case our diary. So
we have this diary of Marjgerie Fleming
who was a child author. And in this case
I will select my
let's go with more than 20 pages. We
have 25 pages of ground truth. I click
on the action button. I click on train
model. And then I select text
recognition model. So we do have as you
can see more options to train a model
but for now for today's beginner's
webinar we'll stick with the text
recognition model. So I click on text
recognition model and now we have
already selected the training data. So
you can see here let me zoom in maybe
it's a bit much the training data we
have selected that already. You can see
here again we recommend at least 20
pages of transcribed material. We have
done that. So we can click on next. The
next step is the validation data. So the
training data were the pages um that
represent my material. These are the
pages that were correctly transcribed
because the AI learns from them. Now the
validation data is a group of examples
that allow for a neutral evaluation of
the model. So what does this mean? So
these are the pages that the transcribus
AI compares the results with to
determine how accurate the model will be
during the training. So in this step,
10% are automatically selected of your
set. In this case, it's only two pages.
We recommend just sticking with the 10%
for now. You can do a manual selection
and change this, but generally speaking,
this should work quite well. Then click
on next. And now what's left is just the
model setup. So we'll give this a name.
We'll call this Marjorie Diary
um
webinar model. We can give this a
description.
Marjgery
Scottish
check author. In this case, we recommend
to also give information about the
material itself. So, for example, that
it is um handwritten text, the time
period that it was written in, maybe
even a specific script. So, if you're
working with German um text, and you're
working with Kurand, for example, to
specify that as well.
Um then the last thing that is mandatory
is the language. So we'll select
English.
And now we can
have a look and check the information
that we have entered which is everything
is correct and we can already start the
training. So the process that takes the
longest for model training is usually
preparing the training data. But again,
you can really save time with doing a
pre-recoognition with a public model
first correcting this and then
uh uh using these pages for for the
training data. Now that we have trained
the model or we started the model
training, where can I find my custom
model once uh the training has finished?
We go again to the models area, but
instead of the public models, I go to my
models. And here is an overview of all
the custom models that I have trained.
So these are only visible to me unless I
share them with someone.
And we can see here a model that we have
also prepared ahead of time.
You can see here the overview
of the model and this is also where you
can evaluate the model. So we mentioned
the CER the character error rate already
previously. The character error rate is
an indication of the accuracy during the
training. So for example, we have here
an accuracy rate of 38.9. So almost 39%
meaning that 39 characters out of 100
were recognized incorrectly. That's
quite a high number. We can see that the
training pages that we used with this
model, the amount of pages we used were
only 19. So that's not that much of
pages or not that many pages that a
model can learn from. So we always
recommend to add more pages to the
training. The more pages the model sees
during the training, generally speaking,
the accuracy improves
and I can show to you what that looks
like back in the slid set. I think we'll
have to skip maybe a few. Yes. So we
have seen already the different training
steps and we can have a look at the
evaluation of the model. So this is a
screenshot of the model that we already
saw. So the accuracy during the training
was only 61%. And we can also see this
here in the learning curve. So the
learning curve shows the process of the
accuracy. The yaxis represents the
character error rate. So with the
training starting and then progressing,
the curve goes down as the model
improves. The xaxis represents what we
call an epoch or um a full round of
training. And each of these rounds of
training helps the model to better
understand uh the data. So it's
basically practice rounds. So the more
practice rounds the model has, the
better it gets. But if there is no
further improvement, the model stops the
training.
Now what we recommend in this case cuz
the accuracy is not quite good is to
retrain the model
and simply correcting the transcription
in document does not improve or retrain
the model. Accuracy only increases by
training new model versions with updated
ground truth. What does that mean? So we
had this previous workflow where it said
first recognition or pre-recoognition
with public model. But to retrain a
model, we now recommend a recognition
with your model. So that would that then
be Marjgery diary version one. That
would be your model.
Uh accurately correct the text, save as
ground truth,
start a new training,
recognize more pages with the new and
improved model. So these added pages
that you have corrected, you keep adding
to the training progress. So let's say I
have had 20 pages of ground truth first.
Now I can use my model to recognize 20
more pages.
Accurately correct these 20 more pages.
Save them as ground truth. And now start
the new training with 40 pages of ground
truth.
With the second iteration with 40 pages
of training data, the model will
probably already have a lower CER,
meaning it will perform better. Now I
can recognize even more pages. Let's say
I can recognize 40 more pages, add those
to the existing 40 and start a new
training with 80 pages of ground truth.
So repeating this process will improve
your model and we can see here the
effect of this. So we have here the
version two. So speaking of training
pages, previously we had 19. Now we have
164 pages that we used for training.
Meaning the model has seen a lot more
variation, a lot more handwriting during
the the training process. And we can
also see here the accuracy has improved
by a lot. So the accuracy is now over
90%. We can also see that reflected in
the training stats.
And we have a very nice uh table here.
Now this is an indication
of how many pages of training data we
recommend for a reliable model. We
always uh say to aim for a character
error rate of under 10%. With a
character error rate of under 10% you
should already have a quite reliable
custom model. And retraining a model is
a normal part of the uh training
process. So don't get discouraged if the
first model with maybe 25 pages, 30
pages has a CER of over 10%. This is
quite normal. Um just simply add more
correctly transcribed pages to the
training data. start a second iteration
of training and it should improve
already by quite a lot.
Exactly. So it is a normal part of the
process and each attempt brings you
closer to your ideal results.
And that brings us slowly to the end. Um
now you know how to upload and
transcribe a document successfully, how
to manually edit results, how to pick
the right model, and even how to train
and improve a custom model. So we've
gone through the core workflow from
uploading your documents, recognizing
text with public or super models,
editing and even tagging your text and
even training a custom model when you
need this extra accuracy. But there is
even more that you can do in
transcribus. It's not just text
recognition. We also have as mentioned
before different types of models that
you can train such as baseline models.
baseline models can be trained to
accurately recognize text lines in your
material. So if you have material that
has um a bit of a tricky text layout in
terms of sort of lines that are slanted,
quite long or quite short um then we
recommend training a baseline model or
using a baseline model. And using this
will definitely also improve the text
recognition because when the layout
isn't correctly recognized, the AI has a
hard time also recognizing the text
correctly because it needs to know where
is the text written on the page to
accurately reflect that in the automatic
output.
We also have field models that you can
train. So they can also be trained to
automatically recognize and even mark
certain layout components. So they can
also be trained to automatically tag
certain layout components in your
material. Now when you're working with
specific document layouts, we also had
the questions about church records,
maybe tabular data, um layout, field
models as well as table models are quite
useful to train because they help you to
extract the information uh in an
accurate way. So not just the text but
also the the structure that the text and
the information comes in. So with the
table models we can also train
transcribers to automatically recognize
the rows and columns in your training
data. We have uh previous hel previously
held webinars on field and table models.
We have also recorded these webinars and
they are uploaded to have been uploaded
to our YouTube channel. So you can go to
our YouTube channel and have a look in
the webinar playlist and find the
webinar recordings there. So if this is
interesting to you, if you're working
with documents with more complex layouts
and finding out more about baseline
field or table models, you can have a
look at our help center and our YouTube
channel. And um oh yes, here we have a
nice preview of what that looks like to
have the text also extracted in tabular
tabular data tabular form.
And I also mentioned previously that we
have a publishing feature which is
called transcripted sites. So you can
create a searchable online database of
your documents that uh everyone can
access your documents online without the
need for programming knowledge or
extensive um IT resources and you can
share entire collections with either a
private group or public group from
anywhere. So, transcribed sites offers a
very simple and efficient way to publish
your collection online. Again, you can
also find more information about that on
our help center or on our YouTube
channel. We have done a previous webinar
and we have helpful guides uh there as
well.
And with this, I can now
uh wrap up slowly the webinar and we can
have maybe some final questions at the
end. So just to wrap it up with
transcribus, we really want to help you
save time and provide a tool that can be
learned to use quite quickly. AI should
be able to be understood and controlled
as well. Technology should be for
everyone and not just tech experts. So,
we really want to help you save time uh
when it comes to manual transcription by
using transcribus AI to reduce the
effort and time involved and free up
time for working with the content of
your documents. Spend time on
researching on analyzing
and have to spend less time on manually
transcribing. Um it can be quite
difficult also for the untrained eye. Um
and it can be uh the historical
documents can be damaged by frequent
use. So we try to provide easier access
to old documents digitized, readable and
again searchable. We know that sometimes
technology can be a bit challenging and
learning a new technology can be
challenging. So we really try to keep a
low threshold access where you don't
need an installation again don't need
too much of IT experience. We try to
provide these webinars also to help you
get started of course.
Um, and we really want to help you
achieve accurate text recognition with
the existing models that we provide uh
for free in our free plan and also our
super models which are new technology uh
with uh quite good out of the box
performance and even the option of
training custom models which the custom
model training is also even part of our
free subscription plans as well because
we want to give everyone the opportunity
to work with historical documents. Um,
so it's really not just about reading
text. It is a main part, but it really
is about uncovering historical um,
information. So our goal is to provide
the best tools for making history
accessible. And I keep saying we, but
who are we? Who is transcrib? Um,
transcribus itself comes from the
university sector. So we started as a
project at the University of Insburg in
Austria. um funded by the European
Union. We are now a cooperative with
over 250 members. So we are
communityowned and even the first AI
cooperative that exists in this way. And
being a cooperative by our statute
statutes, we cannot distribute profit.
So purpose of our profit is our motto.
Everything is reinvested into
transcribus. Trying to make the platform
better, develop new features and tools
and implement them such as for example
field and table models is one of the
more newer additions. Uh we're also
working on a new feature that is called
end toend models which will be coming
this year. So um we always try to
improve it, make it better for our
amazing user community. uh that is also
helping us be better by providing very
helpful feedback. And as you can see, we
have many amazing members uh co-owners
in the cooperative who are all also very
involved in improving the platform uh
from all across the world and different
sectors
uh and where we try with them to create
a even better transcript tool and
software for for amazing users. So yeah,
now it's your turn. Uh you can head over
if you haven't already to transcribus
atapp.transcribus.org.
Give it a try. Try it out. We have
helpful resources in our help center. Um
we also have regular webinars and the
webinar recordings again are on YouTube.
We will also host a transcribus user
conference this year in September. Call
for papers is still open. If you have an
interesting project with transcribus
that you want to share, feel free to uh
propose uh or to hand in a proposal. Um
the ticket sale should start next week,
so maybe we'll see each other in person
in September. So, we're really excited
to to connect with our user community at
the conference as well. And of course,
we're uh also active on social media. We
always post updates of new features, new
models, new webinars there as well. So
feel free to connect with us there as
well. Now let's see. I think we have
maybe a few more minutes for questions.
Are there questions that were unanswered
that we can maybe have a look at in the
platform
or have a
>> I hopefully everything
everything
nobody can complain about it.
>> Wow.
Amazing. Thank you.
for answering everything and thank you
for everyone for also um asking
questions. I mean this is what why we do
the webinar. We really want to give you
the opportunity to ask questions right
away. Hopefully have them answered right
away. Um so thank you all so much for
joining today. I think we can now wrap
up the webinar. We're we're ending quite
on time. Very punctual. Very nice. Um
again we'll have regular webinars. So,
if there were any questions that were
still unanswered, feel free to join the
next one, ask questions, then we also
have a contact form on our help center.
Um, if there are more specific questions
if you have any struggles with your
documents once you've tried it out and
there's just something that you can't
figure out how to do it, feel free to
head on over there. Uh, look for answers
or contact our help center. And I think
with this
we can wrap it up again. Thank you so
much. Hopefully I'll see you at one of
the next webinars as well. We will send
you the recording of today's webinar via
email probably next week. And enjoy your
morning, your evening, your afternoon.
And uh it was great to have you here.
Thank you and goodbye. See you next
time.