Video summary
The Transkribus Field Models webinar introduces a powerful feature that extends beyond standard text recognition by automatically detecting and labeling specific document layout elements such as headings, paragraphs, columns, and forms while preserving their semantic structure. Unlike custom-trained models that output regions and text simultaneously, field models require a distinct training phase where users define regions and assign structural tags to teach the AI to categorize complex layouts. This capability is particularly valuable for handling diverse documents like newspapers with varying paragraph indentations, legal forms, court records, and sheet music, allowing users to extract only the specific information they need while ignoring irrelevant text.
Preparing and training these models involves a structured process where users upload documents, draw regions using the layout editor, and mark pages as "ground truth" to serve as accurate examples for the AI. To ensure balanced training data that handles layout variations, it is recommended to create random samples from larger collections, with simple layouts requiring a minimum of 50 pages and complex forms or heterogeneous newspapers needing between 200 to 500 pages. During configuration, users can select specific tags to train on, choose to recognize untagged regions, or opt for specialized line polygon training for documents with inaccurate baselines, while the system automatically sets aside 10% of the data for validation to prevent overfitting.
Once trained, the model's accuracy is evaluated using Mean Average Precision, where a score above 60% is considered satisfactory rather than aiming for an unattainable 100%. If certain tags are underrepresented in the results, users should add more ground truth pages with those specific tags and retrain, noting that stacking models is not possible. The recognition workflow proceeds sequentially through field detection, layout analysis to find lines, and finally text recognition, with options to adjust confidence levels and shape details during application. Data is best exported as a spreadsheet using the structural regions format, where rows represent pages and columns contain the recognized text for each tagged field.
The session concluded by highlighting Transkribus's mission as a cooperative organization prioritizing purpose over profit and previewed upcoming "end-to-end models" that will combine layout, line, and named entity recognition into a single step later in the year. The hosts encouraged viewers to rewatch the webinar, pause to test features in their accounts, and follow social media for updates on new models and events, while directing any further questions about smart extract models or unanswered queries to the help desk via email. The webinar ended with thanks to the audience and an invitation to join future sessions as these advanced capabilities are developed and released.
Read the full video transcript
Okay, let's start. Welcome everyone to
today's fields models webinar here at
Transgrievous. We are happy that many of
you have joined today's webinar and
we'll have a look in the next about 60
minutes what field models are and give
you a good introduction into these
models in transcribus.
As you know transcribus, you probably
have worked with similar documents.
Today when we talk about field models,
we will go a little bit beyond what you
are uh probably capable of doing with
transcribers already. So doing text
recognition. Today we will have a closer
look to field models and also answer
some in-depth questions about using
field models.
Just a very brief overview. First we
will start with a quick introduction.
Then we will mainly talk about preparing
training data. So that's a very
important and fundamental step. When you
work with field models, you need to
understand how you train your model
based on the data that you are able to
prepare. Then we will have a look at how
the actual training works, how you start
the training and what you can keep in
mind in terms of settings when doing a
field model training. And then
eventually how you can use your trained
field models and at the end we will have
some time for questions. As always, if
you have questions, just type them in in
the chat during the webinar. So we are
also happy to take them if they fit
during uh the things we are discussing.
uh but there's also time at the end to
answer your questions if we were not
able to answer them during the webinar.
Who's here with me today? First we have
Helena from our com team who she's
leading our marketing and coms uh team
and yeah does a great job in preparing
all of these things for the webinars and
our uh presence online. Then we have
Sara here from our customer success team
who's yeah managing many of our very
successful customers when it comes to
executing larger projects and really
doing the fancy stuff with Transcribus
and yours truly. So I am part of the
board of directors here at Transcribus
and also leading the product development
with Transcribus.
As you've seen uh Transcribus can easily
read such documents. this very example,
this very clean handwriting as well. So
reading this text or recognizing this
text is not a problem for transcurous at
all. But if you have a closer look and
we have a little bit more complex layout
in this example, you probably wonder how
can I untangle this? Because if you have
a look at the first uh line and then the
second one, you already see okay that's
a heading and then we probably have a
page number and they are in line number
one and two and then there's the third
one and then if we have a look at uh six
and seven there already things getting
complicated because these uh uh reading
order does not uh yeah provide a lot of
value for you as a user as you would
like to have those uh sections of the
text uh untangled.
So, we know where the text is and what
it says, but we don't know really where
it is and how it uh is related to each
other. Here is where field models can
come in. Uh you can probably think of
them like a cookie cutter. You can cut
out different parts of your documents
and assign labels to them. So you can
automatically detect those fields as we
call them with field models uh and
assign regions and text to them as you
see in this example. There's a heading
left of that heading and paragraph
section. There's also marginelia and
yeah the main text is in the paragraph.
So how do we make that work for you?
Today we will have a look at how to
prepare training data as said before how
to train your model. How you evaluate
the model that you've trained. So to
understand am I training a good model?
Can I use this model for my production
phase now or not? And how to apply your
model. So we hope that at the end of
this session you will be able to do all
of these things and can successfully
work with field models in transcribus.
With transcribus as you probably know we
try to really uh deliver uh an AI ally
to you that you can use for to simplify
your time consuming and laborious work
so you can focus on the fun part of
working with historical documents. So
how does that look like in practice? Now
I will hand over to my colleague and we
will start having a closer look to the
actual work when you work with field
models in transcribus. So take it away.
>> Thank you F
and we will jump right into the topic.
Just make whole screen. There we go. All
right. So as Flo already explained,
field models help uh transcribus
identify and categorize different
layouts element layout elements in your
in your documents in your material and
make sure that the structure is
preserved and also understood in a
semantical manner. So up until now when
we used layout recognition in transcreus
uh we focused more on detecting the text
regions and then the text baselines and
then extracting the text from there. So
we did get the transcript but we didn't
really know uh if the text was part of
just a block of text or marginelia or if
it was a heading. And with field models,
we can now train transcribers to
automatically recognize and also tag
these regions and therefore also to
specify this additional information in
the transcription. So with field models,
we'll then have to do some additional
steps compared to the text recognition.
Um but the benefit is that you get this
additional information about the layout
of your document. So when I talk about
additional steps, what do I mean? So um
unlike with customtrained text models
where you train a text recognition model
and once you use it, you get the
regions, the lines and the text all in
one recognized. With field models, we
need to split those steps up. So we need
to make sure that really only the
information is extracted together with
the specific layout elements that we
want. So what we recommend here is to
look at your documents. Look at what
kind of documents you have. What kind of
structure is in these documents in the
material? What does the layout look uh
like? And do I need to use a field model
or maybe do I just need the normal text
recognition model? Of course, you can
train a field model for different types
of layout even if there are just some
small variations and deviations that you
want to work with. We have some examples
here. So, uh when field models are
really useful is for example for
newspapers. Uh you can train a field
model to really recognize those
paragraph indentations to recognize when
a new paragraph starts. Um you can also
use them for legal forms, index card,
court records. So really a large variety
when field models are useful. We can
have a look at some other examples. So
here you see an image where it's really
mostly about different text regions but
again specially useful especially useful
uh it is for newspapers
um where you can train to recognize the
specific paragraphs. You can see here
this percentage that is next to the uh
paragraph tags and this indicates the
confidence level meaning the confidence
uh in which transcribers recognizes
those paragraphs but we will talk about
the confidence level later uh in more
detail.
You can also of course recognize uh just
different columns. So, a simple layout
with two or three different columns. Um,
and you can also recognize uh the
content of forms. So, here it's really
interesting and useful if you only need
specific information. As you can see
here, for example, maybe you don't need
the text that indicates name. You just
need the information of the name itself.
And in that instance, you can really
just train the model to categorize these
layout elements and then in turn only
extract the layout elements and text
elements that you need for your work.
But of course, it's not just limited to
text. You can also train the model to
recognize, for example, illustrations,
as you can see here on the left side, um
or even elements of sheet music. And we
actually have a project in Austria where
they are working with sheet music. So
there are no limits and there are some
very exciting projects um that are
working with this technology.
Now to preparing the training data. So
now we have our documents, we have our
let's say our forms um or our newspapers
and we want to prepare the training data
to train the custom model. How do I get
from my basic layout to actually the
training data, the ground truth data to
train your model?
Um, if you've used transcribers before
and attended a previous webinar, you
might know this workflow. So, generally
speaking, we have different workflows
where you can also use public models.
For trainable layout models, we often
have only a limited amount of public
models such as also for uh field models.
We do have uh a few. We have for example
two public models that we trained. The
baroness of blocks which works well as
the name says for blocks and the
marginalia monarch which works quite
well for marginelia. We had a few
questions in the registration form
asking how to recognize marginelia. Of
course, you can train your own model to
recognize marginelia, but the marginalia
monarch is already quite good. So that
is something that we can also recommend
trying out.
But if the public models that we have,
we have also a few published ones uh
from users. If they're not sufficient
and you have specific layout that you
want to train your own model for, then
of course you can use the next workflow
that we'll show you. And for the
training of this uh of this custom
model, as usual, it's always very
important to have specific examples that
we can show the AI to help transcribers
recognize and understand the layout in
your training data. So you need to
prepare these examples.
And how that works preparing these
examples is to first of course you
upload your documents. But then you
accurately draw and tag the text regions
that you want to train your model on.
Then you save these uh pages as ground
truth and you start your first training.
With field models, we recommend to have
uh 50 pages of ground truth as a start
to train your model.
and I'll show you how you can prepare
these examples in the web app. Let's
have a look.
So, while we're here and I have all
these examples, maybe first uh I can
show you an example of public models. I
have one page here where we can see on
the left side
marginelia.
This is an example where it's already
recognized. But let me just go one step
back
store this version.
So this is the empty page and how you
can use the public models is to just
click on process with AI
and
see this.
This looks funny.
Okay. So, do you know why it's so
squished inside and not opening
properly?
>> We can see uh start recognition at the
bottom.
>> Okay, perfectly. So, it is important
here when you're working with fields to
switch from text to fields and then here
you can select the public recognition
model. So in this case I selected the
marginelia uh of um no yes marginelia
monarch but you can see here there are
already uh quite a few models uh that
are also public as well as some private
ones that we've already trained but if I
select now marginelia monarch I can see
I've selected just one page and I start
the recognition
it will as you already have seen the
preview before it will only recognize
the margin marginelia in your document.
So it will not recognize the blocks of
text but really specifically just the
marginelia.
Now we are number one in Q. But I can
just go back to our previous field
recognition. You can see here in our
version history that I used field
recognition.
And to just show you Oh, it's already
finished. You can see here that
low that maybe
you can see here that really only the
marginelia has been recognized and not
the blocks of text.
All right. Now if I want to train my
model and I need to prepare the training
data,
what you do is you again open the
editor. And in this case, you can as
usually see the layout editor on the
left and the text editor on the right.
For preparing the training data, we
really only need the layout editor. So
I'll just enlarge in that a little bit.
And then you want to start with drawing
fields and regions. And you do that by
opening add region. Then click on add
text region. In this case, we want to
add text regions. And then you add the
regions by clicking once to start the
region and then clicking again
to finalize it. So it's not a dragging
motion. It's really just one click and a
second click to finish it.
And you can draw however many you want.
If you have a bit of a complex layout
that is not just a simple block, you can
also adjust this
by clicking on let me zoom in on the
region borders and then using these
anchor points
to adjust the region a bit better. You
can also add new anchor points
here
and by just selecting it and dragging
the layout region where you want it to
be. So this way you can get some more
exact layout depending uh what kind of
complex shapes you have.
Then the next thing you want to do is
add tags. So the layout tags in
transcribus are a label that you give to
a specific part of the document
structure and this describes what this
is not what it says. So it's an
indicator that helps you preserve
information and context and you can add
the tag by using your mouse and right
clicking a region. And here you can see
assign structure type and I can select
one of the structure tags. For example,
this is the newspaper. What do we have
here? I select this region and this is
shelf mark. If you notice that a tag is
missing, for example, I want to let's
add another text region here. I want to
tag this um with the tag name, but I
don't see the tag name appear here. What
you can do is let me zoom in a little
bit. Go to settings here in the bottom
right corner.
Open settings and then you can see an
overview here of your structural tags
and you can check those if the tag that
you need is already available. We can
see here. Okay. Name. I turn the toggle
to enable name. And now once I click on
assign structure type, I can see that
name has appeared as a tag. And it
already appeared here on the left in the
image as well as on the right here as
well. If the tag that you want to add is
not available here anywhere, you can
click on edit tags in collection
setting.
This opens your collection settings and
here you can add new tags to your
collection. So if you want to create a
new tag, you click on create new. Now
let's call this example
webinar.
Even choose color.
just with turquoise. Um, and
create the tag.
Now, if I go back to my page, let's just
reload the page real quick. Let's save
our changes first. Thank you for the
reminder.
And then we should be able to see our
new example tag.
See if we can see it already.
Our example webinar. Yes. And it is
already enabled. So this is how you can
add new tags to your uh document uh and
to the documents that you're working
with.
Now one thing that you can also do is um
just a small tip, you can also use
doubleclick to open the structure tag
menu. So if I enable this, instead of
having to right click, I can just double
click with my left mouse um button and
the tag options are visible. So right
click also opens some other options,
some other settings and double click
just immediately opens uh your tag
overview.
Now if you train your uh field model and
because you can train your field model
to only recognize the regions you need,
you can also choose to train the model
to only recognize these four regions,
these four fields that I have drawn. You
can choose to not uh draw other regions
and to not have these uh recognized in
the field models. So really you can see
what information do you need from the
documents you're working with and you
can choose to leave out uh some
information. Of course you can include
it but just so that you know you don't
have to um tag and mark everything.
Then we save uh our changes. We already
saved it for now. And you can also
change the status of your pages. So we
now change the status to ground truth.
Um as mentioned before to train the
model we need those accurate examples to
show the AI what we wanted to learn and
we call these accurate examples ground
truth and with this status you can
indicate that we have now created this
accurate example and it can be used to
train the model.
Um again to start training a model we
recommend at least 50 pages and more
pages for difficult layouts or
variations. You um can create a larger
amount of training data in a quite
simple way and that is by using a sample
or creating a sample.
So to create a sample, this helps create
or yeah create a balanced set of
training data of ground truth data. And
this is especially useful if you have a
larger variety of uh of different
layouts. And to do this, you select the
documents you want to include. We'll now
use our example collection of the
Peabody Peabody newspaper index cards.
And then you go to our action button and
you click on create sample. You can give
it a name.
We'll call this index
sample.
And we can see here can add more
documents but we have already selected
the one we want. And then you can select
the number of pages you want to include.
So there can be an absolute number. So
for example 50 pages or also a
percentage if we say 20%
you can see
47 random pages will be selected.
Now we can click on create sample
and it should now create automatically a
sample of this whole collection. And you
can use this sample then for model
training. And because uh if the sample
is large enough, it should then add
really randomly every type of uh layout
of field layout that you have in your
training data. uh for the model
training. Now once your ground truth is
ready, you can start with the model
training and this is my cue to hand over
to Sava as she will explain
how to train the model. Thank you Elena.
I'm just going to share my screen.
Okay.
So let's look now how to train our uh uh
mod our first field model.
As Helen said fields models can be
trained to automatically recogn
recognize and mark certain layout
component of the document.
uh when we have our uh at least 50 pages
of ground root we can train our first uh
field model. So we select the documents
or the pages
um with containing our ground. We go
here to this drop-own menu and we select
trade models and then field models. As
you know we have it's possible to train
four different type of of models in
transcribus text recognition models
baseline models fields and tables and
today we focus on fields models.
Um
yeah we are just repeating that you need
to before starting the actual training
you need to create your grant route and
50 pages of training data is a good uh
start as Helen said it really depends on
the complexity of your material. Um if
you have just two uh structural tags,
two type of regions and the pages are
very simple, for example, the the main
text and the marginalia, uh 50 pages
might be enough. Uh if the documents are
uh more complex and you have uh 30
different types of uh regions just on
one page because you might have very
complex forms with a lot of different
pieces of information. In that case, you
need to increase the number of uh ground
root pages to at least I would say 200
uh to 500 ground root uh pages for very
complex layouts. The same with
newspapers.
If you want to train a field model just
on one newspaper type,
50 or 100 u pages of ground truth can be
enough if the newspaper layout is very
homogeneous. But if you want to train a
field model that can recognize different
types of newspaper layout, you need to
increase the number of ground pages. So
just so you know the logic uh between
the number of pages to to include and
it's always good to do a test with uh to
start with 50 pages train first field
model
like as a test and then increase uh the
number of ground pages.
Uh after you select train and model um
the this the button we saw before um the
interface ask you to se to confirm your
training data. So all your ground pages
or all the pages selected um go to the
training data. Uh the training data is
the pages the model is trained on. So
here you just need to confirm that you
want to train your model on those pages.
Uh then the second page uh is the tag
selection. You need to uh select to
choose the tags your model should be
able to recognize. So the training will
happen only on the tags you select in
this um in this page. Um, if you don't
select a certain tag, the field model
won't be able to recognize it, even if
it's in your ground pages.
Here, it's also possible to flag the
option that Helena mentioned before,
recognize untagged regions. Um
so if you just want the field model to
draw the regions in the way you want uh
without adding tags u you can do that by
flagging this option. Uh and this is
what the baroness of blocks public model
does. So it recognize the blocks the the
regions
um without
adding any tags. And here we also have
another option option called train
online polygons.
Um
this is a diff a slightly different type
of field model.
Um if you select this option the field
model isn't trained on the tech on the
regions and on the tags but on the line
polygons. The line polygons are uh the
polygons encasing the text uh the lines
of text. Uh sometimes with very special
uh documents um the auto the the
automatic computing computing of the
line polygons isn't satisfactory. So you
see here we have this page where we have
or we have very thin uh and long line
polygons.
Um if you just run a a very simple text
recognition
uh on those pages the risk is that uh
many words aren't recognized not because
the text recognition model is bad but
because the line polygons aren't
accurate. In such cases, you can train a
a a field model on line polygons to get
the polygons correctly recognized. But
this is a very rare case. So in with
most with many documents, I would say
98% of the documents or 99% of the
documents, you don't need to do that. Uh
but if you have very large um characters
or uh
your documents are very peculiar, you
might think about train using this
option.
But now look, let's go back to the
standard fields models. Uh on the next
>> sorry can I just ask sorry uh there was
one question um about the training of
the field model and it is does it matter
for the training whether there's already
recognized text in a document or do I
have to train the field model before
using a text recognition? I think that
fits in here quite well.
>> Yeah, thank you. Uh no, it doesn't
matter if there is text uh on the
document.
you can create the ground root. Uh it it
just looks at the regions and the text.
Uh if there are baselines
or text or recognized text, it doesn't
matter to to the training.
And then we have this last uh this page
where you are asked to select your
validation data. automatically
transcribus uh prompts you to select 10%
of your uh training data. So during the
training 90% of your grant pages are
used to train the model and 10% are set
aside to test the accuracy of the model.
Um
it's so we recommend to stick with the
10%
uh to to do the 10%.
And then we have the model setup page.
Here you need to add name, description
and further information about your
model. And here also possible to um
change the advanced settings
um the trainings. Uh these are the four
advanced settings that you can uh
change. Uh for the training cycles and
the learning rate, we recommend to stick
with the default settings, the
recommended ones. Uh the training cycles
are the number of times the model goes
through the entire training data set.
Uh by default, I think is 15,000
training cycles. Um,
you can increase this number up to
30,000.
Um, and I would do that only if you have
many ground pages. Let's say if you have
a more than 500 ground pages, it might
make sense to uh increase the training
cycles. If you have a very limited
number of grant root pages and you
increase the training cycles, the risk
you the risk is that the model overfits
which means that it learns very well the
training data but then it doesn't um
learn to generalize on new data. So
there are also there is also risk if you
increase it but you don't have enough
training data
and then you can also choose between two
different type of uh backbone
architecture. Uh one is standard and one
is enhanced. Uh select the standard one
if you have simple document structures
uh like the index card we show you
before. uh select the enhance one uh if
you want a field model that can
recognize different type of layout at
once. For example, if you have a a form
that changes over time, uh you can train
one field model on different uh on these
different forms. U you just need to
increase the training data to make sure
that all the different type of forms are
seen in the training. And it could also
help to select the enhance um
architecture.
When the training is uh finished uh we
uh want to review our model and the
value that uh express the precision of
our model the accuracy is called mean
average pre precision and you can see it
here in the models card.
Uh this value is a complex me measure
that evaluates how accurately the system
detects text regions considering whether
they were detected and how well their
size and shape match the validation
data.
If the mean average precision is over
60%
um it means that the model is uh quite
good and usually it delivers
satisfactory result. So don't aim to
reach a 100% uh% accuracy because it
will be very difficult uh I would say
the impossible
um also because uh it also depends how
this value is is measured. So it will be
impossible to reach 100% accuracy. uh
but uh if it's already over 60% it means
that it's good and very likely it will
give it it will produce the out outcome
that you expect and always run a
recognition test with the model on a few
pages to evaluate the results yourself
because even if the uh mean average
precision isn't that high high as you
would have expected fact it could be
very good during the recognition and you
don't need to and maybe you don't need
to increase the ground and train a new
version of your model.
Uh you can also in the model card you
can also check the uh number of the
instances per tag to understand how
often a model has seen one instance of a
tag. Um, so if you you can check these
numbers and if you notice that one type
of tag is less represented than the
others, you can create some new ground
truth pages including uh more of those
of of the pages with with this tag and
then retrain a new version of the model.
As I said, it's always possible to train
multiple version of your model. And when
you have the first version, you can use
it to speed up the creation of more
ground pages. Because if you have a
model uh to start from, you can
recognize new pages and just correct the
regions instead of drawing everything by
by hand. uh and uh in the end uh you can
train a a new model with the old ground
root and the new ground pages.
And now let's look how it looks like in
the the platform. So I have uh this
ground truth pages. Uh I and you see
here the banner at the top is dark green
which is the color for run in
transcribus. I select my 51 run pages. I
click train model
build model and automatically all the
pages I selected are assigned to the
training data.
uh I go to the next page uh and here I'm
asked I need to select the tags I want
my model to be trained on. So in this
case is details
um let me me check
name
newspaper
reference and shelf marker.
uh
I can also combine it with recognize
untagged regions. Uh what is not
possible is to combine it with uh the
training on line polygons. So either you
train the model on regions or you you
train the model on region or you train
the model on line polygons. U it's not
possible to do the same uh both task at
the same uh time with the same model. Um
here uh I have my validation data. I can
also change the percentage or manually
select my validation data but it's safer
to select it to keep the automatic
selection.
uh because in this way we know that uh
the validation data is u has been
randomly selected uh and so
it's unbiased and uh it should give us a
realistic
value in the accuracy and uh here we can
add
the model name the description here we
have our advanced settings for now we
stay with the standard option because
this is simple model. We go to next. We
review all our data and we can start the
the rec uh the training and we when we
go to the training lab uh we see that
the fields model training has been
created. Uh we are first in the queue.
Uh so we need to wait uh probably a few
hours uh before the training is
completed but we will receive an email
when the training is uh is done. And
when the training is finished, we can go
to uh under the models
category here uh select fields and we
see our uh the models we have trained.
Uh we already trained the model on those
index cards. Uh you see we have 47
training pa pages in the validation data
five in no uh 47 in the training and
five in the validation and the mean
average precision is 88.77%
which is very good
and also here if I click show details I
can see the number of instances for each
tag
after the training uh let's see what we
do next. So how to use the fields
models? Um
>> can I
s
>> Yes.
>> Can I just quickly um ask two questions
before we move on to using the model?
One question was about the training
process and the question was if you
train the model on line polygons can you
skip the layout recognition so the
baselines recognition later
uh yes you you need to do that um
because if you run the
in that case you need to run the field
recognition
and then directly the text recognition
uh remembering to flag the option keep
existing line polygons. So if you have
your um field model train on line
polygons, you use it for the recognition
of the line polygons and then you go
directly to the text recognition. Uh but
please remember to flag this option
otherwise
transcribes
deletes the line polygons and uh draw
draws new ones.
>> And the second question was about
retraining a model and maybe you could
just quickly reexlain the process. The
question was when training a new model
after the first training do you use a
public model or the previous model
trained on the same document? So I
assume that the question is about what
the retraining is based on.
So many uh we have some public models
for fields
um
but uh
probably they if you have very specific
uh type of layouts they won't work on
your u documents. So the strategy here
is to train a first version of your
field model uh on a few pages, let's say
50 pages um to have a sort of a draft
model. Uh so instead of creating um 100
or 500 grroot pages manually, uh you can
start small. So with 50 pages uh you
create 50 pages of ground root. You
train your first version and then you
use it to recognize
uh the remaining 450 pages. Let's
imagine it's a very you want to train a
very big model because you have complex
documents. Um, that way on those
additional 450 pages, you don't need to
draw and tag everything manually, but
you already have a
a pre-recision
done by the first version of your model.
And this should speed up the creation of
your ground root because you don't need
to draw manual everything manually, but
you need just to adjust the the regions
because maybe on some pages they are too
short uh or too too big. Um and when you
have checked and correct all those 450
pages, you start a new training
including the early 50 pages of ground
root and the additional the new 450
pages and you train your final uh field
model. With field models it isn't
possible to um select a base model as we
have we have them with text recognition
models but this is not possible to to to
do with field fields models. So uh you
cannot train a field model on the top of
another one. You always need to select
all the bas all the ground.
I hope I answer. Thank you Sarah.
>> Yeah. So now let's see how to use fields
models. We can directly look at it
inside the
the the platform.
So let's go back to
our collection
and let's take uh one page.
Uh we are here. We have our index cards.
Now I'm showing you how to do that for
one page. What we recommend is to test
it on five 10 pages and then when you
have found the correct settings the
satis that gives you a satisfactory
result you start the recognition for the
entire documents or the entire
collection in a batch. So you don't need
to do that page to to run the
recognition on each page separately. Um
we are here and we click process with
AI. The first step is the field
recognition with the model we trained.
Uh we select a field here. Uh
and our model call in the card model 2.
Um
and we start the recognition. Here we
had have the advanced settings and we
will look at them in a minute.
So let's first start let's first start
recognition.
It's running.
Okay, now we have our region
with the correct tags.
Uh what we can change in the advanced
settings. So one setting we can change
is the detection confidence level. So
how confident the model should uh should
be. uh it's what uh Helena show before
with the newspaper. So we are telling
transcribus
um to to recognize
um the regions with a certain
confidence. So if in this case if um
trans if transcribus if the confidence
is less than 47%
uh transcribus won't show us those
regions. So here when it's set to 75% we
have two people recognized here. uh when
we lower the confidence level we have
three of them because the model wasn't
that sure that this should be recognized
as a person.
Um so if you increase the confidence
level uh it's likely that you will get
less regions uh but they are more
accurate. If you decrease it you get
more region with the risk that some of
them are wrong.
It really depends on your material. So
you can do a few tests. Uh the default
value is 0.75%
which works well on most of the cases.
Then you can uh change uh the shape
detail levels level. So how you want uh
how you want the shape of your regions
regions. If you want just a rectangular
shape or a polygon and also this is
medium. So you have a uh a polygon which
is however close to a rectangular and
then if you set the shape detail level
too high it's more um it follow more the
the content uh of the region. But
remember the shape of the region depends
on what you um had in your training
data. So if in your training data you
created just simple shapes rectangular
region even when you set this value to
high transcribus will create s simple
shapes.
uh then there is the possibility to add
the field recognition to exist to an
existing layout. Let's imagine you have
this document and you already transcribe
the main text and now you want a model
that recognize and text the marginelia
but you want to keep the main uh the
other text you need to flag this option
and this option is also important if you
want to combine tables recognition and
fields recognition. So if you have a
table but also a text region on the page
and you want to combine the two of them,
you first need to run the the table
recognition.
When the table recognition is finished,
run the field recognition flagging this
option and transcribus will keep the
table and add the the regions. Uh
and then we have a option that we
recently introduced. It's called merge
over overlapping regions. And also here
you have a threshold. Uh if two regions
with the same tag are uh overlaps for
more than this threshold. So in this
case for more than 65%
transcribes automatically merge them. Uh
but please remember that they should
have the same tag. Uh and also that this
option doesn't work if the shape detail
level is set to high. Uh so it works
only with the detail level set to low or
medium.
And then the next step and we can look
at it here. Uh we have our regions and
now we want to find the lines inside
those uh those regions.
Uh the next step is to run the layout
recognition and uh select one of our
public models. Uh the the best ones are
mix line orientation, universal lines or
horizontal line orientation.
And uh
here remember to modify the advanced
settings. In particular, you need to you
must select this option keep existing
text regions because we want to find the
lines inside our fields. We don't want
transcribers to recognize other regions.
And it might also be helpful to decrease
the minim minimal benzlank length to low
especially if you have uh very short
lines like digits. Uh and also to flag
this option split lines or region
border. U so the
line ends with the the region and we can
now start the layout recognition.
And the next step when we have our
lines, the next step is the the text
recognition.
Okay, it's done. And we have our regions
and inside
the the regions the
our lines. And then as I said the last
step is the text recognition.
We can select a super model or in this
case transcribus print M1 is enough
because the it's
typewritten
uh text and I can already show you the
the result. So the final results will
look like this.
Now we want to the final step is uh
exporting our um fields. Uh we select
the pages, we go there, we click on
export and the best option um to extract
information uh extract the data uh when
you have a fields model is to select the
spreadsheet file format and here the
type you want you need to select is a
structural structural regions export.
This will create a
a spreadsheet. Uh each row is a page
uh of your document and each column is a
a field. So the header contains the
structural tags. Uh while in the cells
you can you the cells contains um the
the text recognized in that specific uh
region.
And now I leave the floor to to Helen
for the last part.
Thank you, Sara.
I'll just share my screen real quick.
I think it's lagging a bit. Let's see.
All right, we're already at the end of
our webinar.
So, I'll just be doing a short wrap-up
of
what we've seen today
and what we learned.
We go.
Think you should be able to see my
screen. Yes. Nice. There we are. All
right. So, we had a look at a few
things. We had a look at how you can
prepare the training data for your field
model, how to train your field model,
and how to evaluate it. And we also
learned how you can retrain and improve
it, and of course, apply the field model
to your document.
And the whole idea with transcribus is
that we really want to help you save
time with your work with historical
documents and provide a tool that can be
learned quite quickly how to use or you
can learn quite quickly how to use it
because we think that AI should be able
to be understood and also to be learned
and also controlled and that technology
should be for everyone and not just tech
experts. So we really hope to create
tools that help uh you work with
historical documents, with handwritten
documents, uh save you time, save you
work. Um because it's not just about
reading text while it is a big part is
reading the text, but also understand it
and uncover historical information.
And that is why we really want to
provide the best tools to make history
accessible um and provide the best tools
for you as our users to help you work
with your historical documents. Um but
who are we? We are REIT co-op. Uh
transcribus itself comes from the
university sector. We started as the
REIT project at the University of
Insburg and we are now a cooperative
with over 250 members public and private
institutions and individuals. Our core
idea is the further development of the
transcribus platform and our motto is
really purpose over profit. So
everything uh in terms of uh profit that
we make is reinvested into transcribus
and distribution is excluded by our
statutes from our cooperative.
So everything that we put back into
transcribus is really there to develop
new tools, new features and implement
them. for example, such as the field
models which are uh I think we have had
them for one and a half years now and uh
also uh table models which we had the
webinar last week text or image matching
and also one of our new um uh features
that is coming this year which are the
endtoend models. So with the field
models, we showed you how to h how you
have to do the layout recognition and
the line recognition, the text
recognition in three three different and
separate steps. And we have been working
on a solution to and not have to do
these things in separate steps and end
to end models will be able to take care
of all these three steps in one. So
layout, text, and named entity
recognition will all be able to uh be
done in one step. We can't say too much
now, but this is something that will be
coming later this year. Um, and we
invite you to stay updated by following
us on our socials or of course uh take
part in one of our future webinars. Um,
but yeah, this is something that we have
been working on for a while and we're
excited to to be able to show you later
this year.
As you can see, we have a lot of amazing
members in our cooperative who are all
very involved in providing feedback and
involved in furthering the technology
and the community also and really
creating better transcriptus tools for
our users. And
now, yeah, now it's your turn if you
haven't already. I think most of you uh
have indicated that you have tried
before. if you haven't yet, it's your
turn. Um, also to try training field
models. We really hope that this was a
helpful instruction instructional
webinar to training field models. Um, as
mentioned before, we'll also upload it
on our YouTube channel. We have more
helpful information on our help center.
Here's the link as well. Um, you can
find more instructions, helpful uh
posts, step-by-step instructions on how
to work with transcribers, how to train
models, and anything that you can think
of is in there. Uh, we also have uh
always some upcoming webinars. You can
find them on our events page on the
website. You can also find our past
events there. I always try to include
link to the webinar recordings in the
webinar description. So, at least the
webinars from this last fall should have
the webinar recordings also linked.
But if you don't want to click through
all the past events, you can just go to
YouTube and rewatch our latest webinars
um and click through those. The nice
thing is you can just pause, try for
yourself uh in your account in
transcribus and then continue watching.
So that's also very nice and that is why
we upload them there. Um yeah, you can
join the conversation, follow us on our
social media accounts. We post regular
regularly also updates on new models,
new features, um events that we're
attending if maybe we can meet in person
somewhere. And that's how we uh want to
stay and try to stay connected with our
user community as well. And I think
that's it for now. It's 5:02, so we're
quite on time. If there aren't any more
questions,
do we have uh some questions that we
should still address Sara or
>> There is one about smart extract models.
Uh if
>> yes,
>> we will announce something this year or
when
>> yes, so our end toend models, we will be
announcing something this year. Again, I
can't say too much for now, but we can
say that they are coming and that we've
been working on them for a while. Um,
but we will update you with information
as soon as we can and we will be excited
to do so once the time is here.
So, I think we'll just wrap up the field
model webinar for today. Um, if you have
any more questions that we weren't able
to answer, you can always also reach out
to our help desk. Um, we also have a
great team there and they'll do their
best to answer your questions uh via
email.
All right, then. Thank you so much for
joining today. I hope you'll all have a
nice evening, nice morning depending
from where you join uh our webinar and
then I can just say thank you so much
and I'll hope I'll see you next time
again in one of our webinars.