Video summary
The EUPILOT project, funded by the European Union under EuroHPC, aims to establish a sovereign High-Performance Computing ecosystem based on the open RISC-V architecture. Led by Barcelona Supercomputing Center with participation from nineteen partners across Europe including Chalmers and KTH, this initiative seeks to reduce dependency on non-EU technologies like ARM and x86 while enhancing security and fostering domestic industry growth without royalty fees. The core of the demonstrator involves an autonomous accelerator platform integrating eight discrete chips connected via immersion cooling technology at BSC, featuring two primary accelerators: one combining out-of-order cores with RISC-V Vector Extension 1.0 for general tasks, and another dedicated to machine learning stencils.
To overcome significant challenges in hardware verification and software-hardware codesign for energy efficiency, the project employs simulation tools like G5 alongside FPGA prototyping within an LLVM/Clang toolchain environment. Researchers are optimizing applications such as convolutions using advanced algorithms that leverage long vector registers up to 16,384 bits enabled by ELMO, while managing parallelism and handling missing instructions through custom extensions or intrinsics. These efforts have yielded optimized implementations achieving twenty percent faster performance than traditional approaches on Gem5 simulations for small networks like VGG, although scalability studies indicate diminishing returns beyond specific vector lengths and cache sizes, with larger models still presenting simulation hurdles that require further algorithmic adaptation.
Beyond immediate technical optimizations, the project addresses critical software porting issues, particularly regarding CUDA workloads which often necessitate automatic conversion tools or manual register management to function effectively on RISC-V hardware. Emerging support for Vector Spec 1.0 from companies like Tenstorrent and Semidynamic is expected soon to ensure interoperability, while future specifications aim to include self-contained instructions, expanded registers up to one hundred twenty-eight bits, sparsity support, and standardized matrix extensions alongside unified performance monitoring counters. OpenMP optimizations have already demonstrated the ability to reduce barrier overhead by over two times on existing hardware like Intel Xeon Phi, highlighting the potential for significant efficiency gains as these technologies mature.
Looking toward the future, while dedicated HPC clusters are essential to overcome current funding and awareness challenges compared to AI-focused initiatives, experts predict that RISC-V systems capable of competing with today's leaders may emerge within two to six years. The ultimate goal is iterative optimization leading to competitive exascale systems in subsequent generations of hardware, ensuring a robust European infrastructure for high-performance computing. By integrating these advancements into widely used libraries and adapting algorithms for long vector registers, the EUPILOT project paves the way for an independent, secure, and customizable HPC landscape that can sustain Europe's technological sovereignty without relying on external instruction set architectures or proprietary licensing models.
Read the full video transcript
should I start
yeah all right um welcome everybody
thank you for coming uh to my talk so
I'm my name is Mikel peras I'm an
associate professor at
shmmer um I thought i' maybe start by
saying a few words about myself because
I guess most of the people don't don't
know me so I am I'm from Barcelona and I
actually did uh PhD at the UPC while I
was working at Bor and super Computing
CER Center so my background or what the
work that I was doing at my PhD was
mostly in the field what we call
computer architecture so it was
mostly design of uh microprocessors and
trying to you know find novel
organizations for Designing high
performance
processors so I was working at uh BC
also at that time so and I have an
interest at in HBC and supercomputing
and I also spent after working for BC
for a few years I went to two years to
to Tokyo I was at the Tokyo Institute of
Technology where they have this
supercomputer which is called subam or
the series of supercomputers
there and at that time I started working
more in term in topics related to
programming models and Analysis of
runtime systems performance analysis of
runtime systems and after that in 2014 I
moved to to shmmer shmmer is in for
those who don't know shammer university
is in gothenborg which is City on the
west coast of Sweden and I've been there
now for for 10 years working mostly on
topics related to programming models and
codesign so I well at shamers we are uh
involved in several uh European projects
that um are related to risk 5 and one of
them is this U pilot project so I got an
invitation from Kenneth for which I'm
very thankful to provide
uh well our experiences to talk a little
bit about what's going
on so this is how I've organized the
talk I will start a little bit talking
about risk five and what's the interest
of the a European Union in Risk
five and mainly focuses on on two very
link projects the European processor
initiative and other project which is
called up pilot which I will present in
a bit more detail
and then I thought okay you know since I
give the chance to talk to you I want to
also talk a little bit about the work
that we are doing at at charm so I want
to talk a little bit about some research
directions here that we're doing
pursuing and finally I'd like to share
some some thoughts about my general
feeling about where we are I mean where
risk five let's say is in in the context
of uh you know in HPC and what is what
needs to be
done so before um starting I'll just
like to ask a general question I mean
how how many of you have heard about
risk
five okay I guess that's most of the
people that's good
sign I'm not going I mean I'm a teacher
but I'm not going to ask you not to go
out and try to Define what it is so
which is what I would usually do but
then we would need many hours
so but anyway I think I should probably
try to provide some
introduction into risk five so risk five
is an open instruction set
architecture it's based on the on the
risk
philosophy and it's very much unlike
other Isis that you have probably very
familiar with such as x86 or arm in the
sense that it is an open standard which
is not under the control of a single
company if you don't know what an
Isa is well Isa is basically can
understand as the language that the
processor understands so x86 armr
wellknown isas or
isas now in opening this context means
that it
is freely available uh for everybody to
use and also to
modify and I'll come to that also in in
in a serious in a in a moment many
people think that risk five is an open
source processor but that's not true
risk five is an open standard it's open
source in the sense that the sources for
them for the ISA are
open now the advantages of having an an
such an open standard is that well you
don't have to pay royalties to any
company if you want to make use of this
Isa you can it enables higher
customization and as a as a vendor if
you implement a processor that makes use
of this Isa you can benefit from the
fact that there's already a shared well
shared software ecosystem but I put you
into parenthesis that this of course
depends on the fact that you that you
have to adhere to the standards so in a
sense this bullet of
customization that's a good thing but if
you are going to customize some parts of
the processor make extensions your own
extensions you will have to develop your
own software environment for that of
course and of course there's a you know
it's a very it's a great Community very
B brand community and a lot of people
have a lot of interest
so it's easy to to reach out and get
support yeah I mean yes I'm totally
happy it depends
on when you say about the customization
uh are you able to customize it uh in
the sense it's like a GPL thing that you
have to give it back the customizations
or you can do propietary ones for
yourself no you you can do proprietary
ones in fact it's there are yeah okay
you will see that vendors are all the
time doing this thank
you okay
now as you saw in the in the tiet and in
the title in this talk I want to focus a
bit about the European perspective of of
risk
five and you might have noticed that
risk 5 appears often when in the context
of the eurohpc joint undertak so in this
slide I actually tried to take some
different views from various sources but
one of of them um actually the whole
slide is not visible so I'm not sure
if maybe if I click it's going to
disappear
twice no so now it's visible but it's
still with the with the bar there I just
want to point out that there is a in the
in last year's risk five EUR Summit
Europe there was an interesting panel
exactly on this on this topic so if you
want to go into more detail what I
discuss in this slide you can actually
go and look into what all the panelist
had to say about the topic but one
important reason that the you is
interested in um in Risk five is that it
wants to achieve
uh or develop an autoctonos HPC industry
and we actually seeing this trend in
many many regions on the world and many
governments that feel that you know the
the way that the semiconductor suppli is
currently organized around a few select
companies and a few foundaries it's a
bit of a risk situation there's a chance
that you might get cut out so many
governments feel that moving towards an
open standard provides a higher or high
security and less risk so there's an
interest in developing uh such
a our own HBC
industry then there are also needs such
as I mean in interest in increasing HPC
capacity development of HPC
skills and
also some people would like to see a
movement towards Open Source Hardware
there's a bit of a debate at this so the
the vision is that okay up to you know
like the Linux Kel from the software
stack there's High we have a lot of Open
Source software but and below the kernel
it's mostly
closed so it would be nice if you can
also have a Open Source Hardware there
is an Open Source Hardware movement but
there is some debate as to how how
viable this is particularly one the the
main one important concern here is that
unlike software which basically you know
compiling can run on your system if you
want to do Open Source Hardware if you
want to actually tape that out in let's
say high performance node that has a
huge amount of cost so somebody will
have to pay for
it now why did uh why is the interest in
in Risk five specifically well it's sort
of because none of the other options
actually fulfill um the requirements as
you know x86 and arm are controlled by
entities that are non EU entities at
this point and that's actually you know
if you if you watch some of the look a
little bit at the history of all how
this has developed a lot of people track
this development of like the USBC
projects now back to the series of
projects that was called the mon blank
projects at the BSC and those projects
were based on arm they were taking first
starting with arm um let's say embedded
I might say or process of that were
originally designed from smartphones to
try to build clusters based on arm
technology but then you know at some at
some point soft Bank bought arm and well
still under control of soft Bank to an
extent as I understand and then the
European felt like okay know we've been
doing all this investment into a
technology which we actually cannot
control so that's one of the reasons
that they suddenly start looking into uh
risk
five then also you can see that risk
five can be a language to somehow
communicate academic ideas with with
industry all right and well this last
bullet is basically what I mentioned
that risk five is an open standard
doesn't mean the same as open source uh
Hardware okay so you probably heard
about the uhpc joint undertaking so
here's a definition which probably I
took from the from the their own website
says it's a joint initiative between the
EU European countries and private
Partners to develop world-class
supercomputing ecosystem in Europe they
are I mean it's they are funding several
different initiatives on the one hand
they are funding big supercomputers for
use uh by researchers there's a long
list of them I think the largest of them
is uh currently is the is the Lumi
supercomputer in in Finland but there's
also an a site of arent well research
and Innovation
projects in which we have to well
participate in a few of them there are
many of these projects but three
projects which are related in fact to
risk five and where Hardware design is a
is important is the European processor
initiative this is sort of the project
which I would say kicked off a lot a
little a lot of this uh work this
project had two um two goals one is to
develop an arm host processor so the arm
is still in the the picture here and and
also an accelerator based on risk five
Technologies in fact it's not a single
accelerator there are several
accelerators this usually called under
the name of
epac and which I think stands for
European processor accelerator
chip there are two ongoing Pro Pilots
also now that which I will actually show
in the next SL a bit which is the Up
pilot and the the upex and there's also
another uh project which I will not be
focusing so much about today which is
called E processor which also is has the
goal of developing a risk five based
processor with extensions for you know
for HBC for executing bioinformatics
applications so let me show you some
slight which I actually got from a from
a presentation from philipo mantoani
he's one of the researchers at at
Barcelona who is involved in the
development of the risk five accelerator
and I'm going to be showing I want to
show you two slides which this is from a
seminar that was in in fact in in ulik
last year so to show a bit what is being
veloped in this Epi uh project so on the
one hand we have an Arm based general
purpose CPU which is called
Ria and you might have heard about this
because as I understand one of the
partitions of the Jupiter is going to be
based on this
technology so this is mostly well
developed by by cyer which was a company
actually funded within Epi to sort of
drive this development and commercialize
it
then on the other side we have this epac
which is
actually a set of accelerators but one
that is most interesting to to us I mean
charmes also to me and related to this
talk is the is basically a vector
processor that is being developed inside
of this project and this is called
epac and why I I I show this because I
want to sort of show this timeline and
in fact in this timeline starting 2015
you see also this mon blank and the Deep
project appearing here you will see what
is sort of the strategy of the
development so we have the the
Epi uh project which has in fact uh two
phases and these phases are developing
these two processors the Arm based host
processor that sort of feeds into this
upex project which is a pilot to develop
a system based on on the re chip and
another pilot to develop a system based
on the on the apar chip and that's the E
pilot
project and the idea is that these the
outcome of these projects will feed into
future exos scale systems so we already
know that um um Jupiter will have some
components from based on
Ria and for
the for the epac there is
a well I would say there is a wish that
some that the next Marin norom 6 will
have a partition based on these
processors okay
so since um the topic is risk five I
thought okay I'll describe a little bit
about what exactly is being implemented
in in the E pilot project these sides
are not mine they are so Carlos Bol he
is the the the technical coordinator for
for the project and so I'm going to be
mostly showing a subset of slides that
he from a recent presentation that um
that he uh delivered that was actually
at high
peak okay so well typical slide with uh
where you can see all the partners
involved in the project we have a set of
well total of 19 Partners that's a set
of acad Partners there's a set of
industrial Partners involved um for some
reason the figure on the on the right I
think that's actually taken from the EU
portal but it's only showing I think the
location of the academic partners for
for some reason but as you can see it's
um sort of distributed among Europe the
project is coordinated by by
BSC and well here in Sweden we have
shmer also kth is also U participating
in in the project
so what are the goals of of the of the
Up pilot project so first and the the
main goal is to
demonstrate this preex scale accelerator
platform so we want to
actually deploy and show the you
know uh a risk five based accelerator
platform that is in a in a usable state
and to reach there there are several
things that that we are working of on
one hand the project is well
designing and doing all this validation
steps and deploying the accelerator
platform with a goal of maximizing
European technology and and assets
because that's always one of the goals
also of the eurohpc
there is um the project is po across the
whole stack so there's a goal to do
software and Hardware
codesign to sort of Ian to basically
Drive the hardware design from the
software uh
components with the goal of improving
performance and achieving better Energy
Efficiency then finally also there are
the goals to extend open source into
hardware for HPC some of the components
that are being developed will in fact
are in fact uh open source or some of
the hardware is open source and to do
everything based on the risk five
instruction set
architecture so this is a top level view
of what we're trying to to build it
actually has um two type of accelerator
chips that I will uh describe in a
little bit more detail next we want to
integrate eight of these ships into an
accelerator board that will then then be
connected to a host server and all this
goes into this uh tank so it's going to
be using immersive cooling
technology that's and there's already a
space
at at the BSC where this is is going to
be um installed as as I understand
okay this is a sort of high level view
of the of the
system and here's a bit more detail on
the actual chips that are being
developed so there are actually two two
chips that that are both sort of derived
from this epac um chip that I mentioned
that had several accelerators so out of
these
accelerators here there are two which we
are going to construct discrete chips
out one is a vector accelerator
which in fact combines uh a core that is
developed by a company called semi
dynamics that is also located in
Barcelona so they have a core they are
designing an outof order core called
atraido and this will be connected to a
vector processing unit that is mostly
coordinated by BC that development that
will that implements the risk five
Vector extension 1.0
standard and on the on the right side
we have a another chip here which is
called the MLS accelerator MLS here
stands for machine learning and
stencil this is um as
you this is designed mostly for for AI
type of um applications we are not so
much um involved in this work so I'm not
going to be discussing much about the
parts related to the MLS in this in this
presentation at the end the goal is to
tape this out in a technology so it's 12
nanometer Global Foundry so this as you
can realize is not the most Cutting Edge
technology but you know the if you
wanted to implement this in like say
like five nanometers the cost of the
project would explore I mean it's not
like going to one of the high
performance nodes it's not an
incremental increase in cost in fact it
multiplies the cost
by probably by more than an order of M
magnitude and this is the software stack
that we are working on so the we are the
way that the project this project is
organized is that we select a set of of
applications of
um yeah Target applications and then
develop the underlying stack to try to
make that work and here we have on the
left side the components that are
targeting the vector accelerator on the
right side we have the components that
Target the machine learning stencil
accelerator so in terms of applications
so the main targets that we are looking
at is Grox which is mostly well
developed and coordinated by by kth by
Professor Eric lindal and uh there's
also EC Earth which I understand is
being developed at
BSC and you know basically stuck off of
many many libraries and tools all the
way down to um the system you Linux and
also of course the tool chains to
compile the applications there so that's
one of the efforts that is ongoing and
that of course is a critical component
of of the
system so at at the end what we are
trying to to achieve in terms of uh of
performance is shown in this in this
slide so of course one thing that I have
to point out that whenever I'm showing
the numbers here it's a this is sort of
subject to constraints of budget and
some you know the cost of a square
millimeter changes over time so
depending on variations some of these
numbers will will will change but the
goal for for ve is to develop a let's
say a 16 core chip in an area of 46 uh
square millim and such that each chip
has eight Vector processing units and if
you do the math for example considering
that the fuse multiply at counts for for
two uh flops then for
64bit Point targeting 1.5 gahz we will
reach to 384 G flops um per chip that's
not counting the double Precision of the
of the core itself so in total that goes
to 432 gflops per chip and since we want
to have eight of these chip on a board
so we hope to um deliver a system where
each of these accelerator uh
systems this eilot accelerator system
each one has delivers about 3.5
teraflops
now I talked a bit
about the work that is being done in the
development of the applications and the
tool chain so and one of the critical
components that is required to make this
work is what we call the software
development vehicles and so you will you
can think of it as follows when you
start developing a process process you
want to start immediately developing
also the software that is going to run
there but of course at that point you
don't have any hardware to actually
develop the software on so what is the
solutions that you can go for
so in the Epi project and also in the E
pilot projects we are relying on U
several um options several platforms one
this is the let's say the hardware
protot type platform that we are we're
using there as a software development
vehicle which consists of um two type of
um of uh platform two types of boards
one is uh taking for example a set of
risk five commercial
boards then in this case we have a
series of of boards from that are this
high five unmatch that's um a board that
has a series of chips with four uh risk
five
course that we can start to use to do
some uh software development of course
this is not the same as the chip that is
being developed we do not have for
example the the vector units here so in
order to be able then to target the
vector unit you need to do something
else one option is that you can try to
emulate the vector instructions here so
there's for example a tool which is
called behave when every time that you
run a you know you execute a vector
instruction on a platform that doesn't
have support for Vector instruction so
what happens is that you get a you get a
signal right you will get a signal ill
instruction signal something well this
tool captures the signal and then goes
to analyze what is this instruction sort
of emulates that in software that's one
option and you can do it or another
option that you can actually try to run
it on on fpga itself you can try to put
the the RTL meaning the hard that is
being designed compile that but instead
of targeting to the let's say the tool
chain for for taping out the the chip
you target an an fpga device and then
you can map it on fpga and you can try
to run your whole software stack meaning
that you can try to boot Linux here and
you can try to run the whole um tool
chain and test your applications there
this of course getting this to work is a
is a is a major effort it requires you
know to develop the hardware itself to
actually get all the like an operating
system to boot there so
and then of
course you you might not be able to
emulate the full system because in an
fpga the density of Hardware let's say
that you can emulate is not the same as
the one when when you take out so here
for example you could have maybe um a
few cores without the vector unit or you
can have maybe just one core with the
vector unit to try to start working and
developing code
there later on I will also mention there
are some other options you can also for
example try to do risk five development
in purely virtual uh environment so we I
will also discuss a little bit about
that there
so this is a little bit what I wanted to
explain about the work that is currently
being done in you pilot so the next
thing I wanted to talk a little bit
about our involvement and sort of a bit
of what projects we are working on at
shmmer but maybe I'll take a break here
just to ask if there's any questions on
on the topic so far
maybe we can get back to it at the end
but what's the biggest challenge in
these type of projects to
actually get to the point where where we
have say a European supercomputer
running
on only risk five so CPU and
accelerator
well there are many challenges so the
question is what is the the major
um well I I think I mean the the
development of the hardware and its
verification to make sure that it can
run that it can reach the stability to
actually run um software that is a is a
major step and I think it's you know
it's really like when you reach that
point it's really like a breakthrough
because at that point you can then start
more know productively developing um
code and then actually try to start to
optimize it so I was thinking actually
about this a bit while coming here that
it's like I can see that there like this
basically two important phes one is to
try to get you know to your first
version to get that working and then
after that you basically have to start
an iterative process because now you
want to optimize the software and you
want to optimize also or and to improve
the hardware so you you start developing
your software targeting this new
hardware maybe using some Performance
Tools and then you can try to optimize
the well get the feedback from the
Performance Tools to optimize your
software but at the same time you
provide new information to the hardware
so a new version or improved vers
Hardware improvements are developed so
now you get to a new new version of the
hardware and now there's a question
about okay these optimizations which did
how portable were they do you need to
now reoptimize your application so at
some point then we get to this point and
that's why you know I I can see that the
road towards getting this uh this to to
work is probably still you know a few
years ahead we are now sort of I think
reaching this point where we we do have
some some
platforms I mean I'm talking a little
bit about these projects now I mean
there are many many approaches in fact I
will I will talk about what some other
vendors are proposing but what I can see
in the project that we are involved is
that we are getting to a a point where
we do have working
hardware and now I think we need to
start with more actively with see how we
can do sort of you know iteratively
learn from attempting
to optimize application seeing you know
how which are the instructions that they
most rely on are the where what
optimization should we introduce into
the hardware then try to develop next
version of the hardware and then see
okay now can we weart again the process
of looking how well these ported
applications are going to run and
eventually we'll reach there and this is
a step that that something that will
require certainly a few steps until you
reach there that's why I think in any of
these projects of this scale it's not
really realistic to think that you know
we're going to develop the chip
immediately the chip will be competitive
no you have you develop the chip and
then we have to start and after a few
generations of the chip then we
hopefully can reach to something that
will be competitive that's um the way I
I see
it all right um okay let's talk a little
bit more about our inol what we're doing
in and charma so we do have two parts
we're working a little bit on the on the
software side and also on the hardware
side so on the
on the software side we have been
looking at vectorizing convolutions for
example we have also been looking at how
to optimize part of the open andp
runtime there and we have also a plan
this we haven't started yet to actually
work on on having an efficient optimiz
implementation of CLE for for risk5 and
it's interesting because there was a
discussion about CLE also also before
and this is sort
of interesting you know because if you
think about how many applications there
out there written in CLE actually the
number is not that large but one of the
the reasons we are interested or there's
interest to get CLE to to work on this
is because CLE can serve as a sort of
intermediate step when you are porting
applications from Cuda to work on this
platform so that's one of the things
that we want to look at also and then on
the hardware side I'm not going to talk
about these things uh these two topics
is we but shers is sort of responsible
for what is known as the home node so
the to basically ensure that the course
are going to be cash
coherent both you know within a single
chip but also across different
chips that's one of my my colleagues um
are working on this part but I will not
be discussing this into into
detail
so for in the next few slides I want to
talk a littleit
provide some example of activities that
we are um working this is the sort of
this is sort of my more like my
day-to-day life so I hope that I can
make it interesting and not just bore
you with a lot of technical um detail
here but before I mean to jump into that
I want to talk about one thing which is
the the risk five uh Vector extension
and the reason I want to talk about this
is because I think this is one of the
important development that has happened
in RIS RIS 5 over the past years that is
an important step towards being able to
have know high performance Computing in
in Risk
5 so you're probably um aware of what
Vector extension is because it's very
common implemented I mean all
architectures or major architectures
have back extensions in the case of x86
you will have something you know like
ABX 512 ax2 earlier you had the SS
extensions in the case of arm we have
the knee
instructions but there's also the sve
the scalable Vector
extension so and all these are very
important to achieve high performance
because Vector execution is one thing is
one way to also achieve good Energy
Efficiency and why well basically you
can think of it as when you have Vector
instructions then you you are able to
decode one instruction but the semantics
of the instruction itself indicate
multiple operations to be performed so
you can do you decode the instruction
only once you schedule it only once but
then you know you can execute many
operations um and there's a guarantee
that those operations are independent so
you don't need to do any let's say
disambiguation at the hardware level
which is something which is a bit costly
so that's why also the approach that
we're taking in eilot and Epi to try to
achieve high performance risk
five execution is using the vector uh
extension and in fact so this resc
Vector extension the there are several
versions of it the
1.0 version let's say the the standard
1.0 version was ratified in towards the
end of 20121 we are leveraging it in
this in the in the epi new pilot uh
project so I put theost ASAC and back
but it refers to basically these two
projects and one of the major strategies
that we attempting there to in to in try
to increase as much as possible the
Energy Efficiency is that we are relying
on having very long vectors so one thing
that I haven't mentioned yet but it's
actually the next bullet here is that
the way that this Vector extension is
built it's also similar approach taken
in the arm sve is that the architecture
itself doesn't dictate the length of the
vector register this is like an
extension like AVX 512 which says you
know ABX 512 is is for 12 bits but in
the case of of risk five you can have
any size of vector register going from
um I think it's either 64 or 128 bits up
to
16,384 you can do this in increments of
of two so this this number of bits what
does it mean basically this means that
you can have vectors of 256 double
Precision
elements and in fact at the programming
model you can even make these vectors
longer because uh risk five supports
something which is called um something
called Elmo which basically allows you
to take multiple registers and treat
them as a single uh
register so the idea here is that we
will be able to achieve higher Energy
Efficiency by following this approach
now this is enabled by having this
programming style which is called Vector
length agnostic
so since now the implementation is free
to choose the actual length of the of
the registers it means that when you are
for example processing a long Vector an
application Level Vector in a
loop um then when the compiler generates
the code it will have to do some extra
bookkeeping trying to understand first
okay how many elements can I actually
fit in the vectors that are available
and then you know partition the loops in
as many iterations required to actually
be able to process this uh this
Loop so this provides some portability
but it potentially there are some
scenarios in which you might lose some
efficiency by by doing this as compared
to what is would be called the vector
length specific um code
generation so the way I I like to think
of it is can we basically do something
that is competitive with a gpgpu by
using um long vectors and a bit of the
of the research that we do goes into
into this
direction now before going to that I
just thought okay another sort of
important uh consideration when you are
talking about
vectorization is how do you actually
vectorize because this is this is a big
topic and something that actually has
been studied for for long long time and
there are many challenges so let's say
you we the two extremes that I have here
on this on this slide I think the is to
understand one is okay you do
everything um well you just code your
program in Assembly Language you have
maximum control of of all the features
this is of course not very um
productive on the Other Extreme what you
would like to have is something where
the compiler does everything this is
usually called autov
vectorization but uh here you give all
the control to the compiler but this
would be the the most productive
unfortunately you cannot always rely to
autov
vectorization First autov vectorization
requires the compiler to be able to do
um analysis of um let's
say memory addresses and and
dependencies between instructions to
guarantee that for example when you're
when you are
um looking at a loop that is for examp
example adding uh two vectors and
storing it to a third Vector you would
be able you would require for example to
guarantee that the output Vector cannot
alas with the any of the input vectors
if you cannot do this because the
programmer has just written some some
pointers there then the compiler must
say I'm sorry I cannot uh vectorize this
otherwise it might be wrong code
generated at the end
so and there are many other scenarios
and and small details that may cause a
compiler to to fail in vectorization so
there's a lot of work also required to
try to get this technology to work
effectively and sometimes we will have
to actually ask the the programmer to
help a little bit which is something
that you can do sometimes easily by
adding some simd Clauses like in open MP
you can do something like pragma om simd
and then you tell the compiler there is
no dependencies in this Loop you can
vectorize this say a compiler say okay
thank you um
but go but reaching this of course from
if you don't have a vectorizing compiler
just getting to to this point is a big
effort so one of the things that that we
often that uh that we do is some
intermediate step which is the program
using using intrinsics this is uh a way
of programming which still gives a lot
of uh control through the programmer but
it provides some some help of make some
some things um simplify some task for
the
programmer so I don't know have you ever
programmed using intrinsics for example
for
x86 okay all right then then you will
then you know that for example when you
program with intrinsics you don't need
to worry about the register allocation
and you can make use of the access
variable names directly so this will
this will be
um something that will be helpful so
actually in in the Epi project a set of
risk five bulletins has been developed
to Target most of the risk five Vector
instructions but not all of them and
they have some format like you know will
be called buildin EPI and the name of
the instruction
and with some extra parameters there but
so and this has sort of been
an a way to prototype to try to develop
a first uh set of intrinsics to make
this
work and later on the risk five uh
Foundation actually has set up a task
group to come up with standardized um
race five vector intrinsics and now
there is a standard you can see actually
I just put you know two instructions
like vvl so set Vector length with the
EPI built-ins and with the name of the
standards you will see that you know
format is a little bit different but
it's basically but they are equivalent
there's a frozen spec now but it has not
yet been ratified so this is not yet
standardized but there is implementation
of NG GCC and llvm are are ongoing
here these are sort of the four major
ways that you can leverage a vector the
vector
extension in fact any new Isa um
extension before it can be fully be uh
be used by the
compiler you can work at the assem level
or you can try to develop some builtins
for it so the one thing that I haven't
mentioned here on this slide is that
when you are doing
assembly there is also an option to
actually generate the assembly at
runtime using just in time
compilation and I will come back to that
a little bit later but this sort of
enables more um
optimizations as long as the more
information you have about let's say
variables
and runtime state of your program the
more you will be able to optimize it
specifically for the platform on which
you are executing so that's also there's
also several approaches that are looking
into into being able to for example
generate um or jit generate Vector
instructions I haven't checking the time
but you started late so it's say half an
hours
anyway let me just then show a little
bit just to highlight a few of the
things that that we are doing without
going into too much detail we'll talk
about three projects so this is a a
project in which um one of
my um PhD students when I who actually
started as a research engineer in in Epi
she was looking into how can we make use
of the of the vector units to try to
efficiently um or to make efficient
convolutions and
well in in particular she started early
with the work in in fact at the
beginning she was working on
RSV and she she was working
on on a few different approaches so let
me actually backrack here and say that
when you are when you want to run you
know convolutions like convolutions that
you would do for machine learning
actually there are many different
algorithms that you can that you can use
you can do what is called a direct
convolution
you can apply an algorithm which is
sometimes called IM am to call G where
you do an A transformation of the of the
input Matrix so that you can use a
matrix multiplication later
on you can also use um an algorithm
which is called
vogr and finally it's also possible to
use um ffts for for convolutions now
ffts are not that commonly used I think
they have some I mean for for the sizes
of the kernels that I usually used there
they are they sometimes lead to
numerical instability so most of the
approaches are based on either direct I
am to call or vogr and Sonia had been
had developed an implementation of vogr
for for rmsv and she set out to try to
Port this work to to risk five to see
what are the challenges of taking a code
optimized for rmsv and try to get it to
run on the on a risk five platform using
the vector extension
so she set out to do this work while
trying to you know maximize usage of the
vector units and of the vector registers
and
also because we are interested doing
some code design we we looked also into
some Hardware uh parameter
tuning so I was thinking can I get rid
of this slide but at the end I decided
to keep it just to point out a bit of
the tool chains that were used and that
this work was actually based on using a
sim a hardware simulator which is called
G 5 so we did not actually run this on
on the fpga platform at this point but
using a software Simulator for the for
the project and also using the llbm
Clank tool chain which is targeting risk
five which is being developed in in the
Epi
project okay so without going into too
much details just try to provide the
high level that when you are doing a
vinograd convolution you you take you
have to do
um three things first you have to take
the input Matrix and the kernel and
apply some Transformations that's first
step the second step is that you have to
multiply this in a type of
multiplication which is called the
topple multiplication and after you have
done this multiplication then you can
convert it back so in
the in the original work that we did on
rmsb we figured out a way on how can we
try to increase the intertype
parallelism or you or basically make use
of parallelism across different input
channels in order to be able to fill um
vectors of longer size so in this case
it's just showing how you can for
example fill a vector of 512 bit using
four channels but I mentioned before
that in the for the case of the V we
have very long vectors of up to 16 K bit
so this approach would need to be
extended so we tried so the first thing
that we did try to use 32 channels in
order to fill in this case up to 4,000 U
bits so you already realize that this is
already sort of hinting to one to one
problem of using very long vectors is
that you will need to be able to find
enough parallelism to be able to fill
those vectors in this case we say okay
we use a parallelism from different
channels but how many channels can you
actually have in a in your application
so that was one of the approaches the
other was to try to using increase the
topple size or also to utilize longer
vectors so we started working uh looking
into how to Port this to to five and we
identified several challenges and you
know when you do this work the thing is
you you will probably be biased when you
start from one um code and then try to
Port it to another one because here we
started from rmsv and then of course the
simplest way to try to Port it is to say
okay let's try to do one to one mapping
and then of course you will find out
okay ah in my original code I used this
instruction that was you for example a
load quad word elements into a vector
and replicate and then when you go to
the risk five Vector spec you might find
that oh this doesn't exist so you have
to find some some alternative so well in
this case we try to we found that that
we we're lacking a similar instruction
that basically the what this instruction
would do if my mouse pointer is visible
maybe this is not the best but is to try
to take a set of uh elements here and
basically replicate them a set of times
over the vector
register okay so we figured okay we
could try to for example use index loads
to basically load these elements and
then keep replicating them on a vector
register or alternatively use some type
of instruction that is available in Ras
five which is called slide up
instructions so we evaluated that we've
realized that okay according to the
simulated data using the slide up
instructions would be considerably
faster than using the index Vector uh
load
nevertheless having a single instruction
that could implement this would probably
still lead to a a faster execution
time another challenge we identified is
that we we were for example lacking an
instruction that would do a transpose of
four
vectors so I guess this figure on the on
the right shows this you know you have
the let's say the four vectors organized
in this row y way in this row based wave
a Subzero until a sub3 here and after
the transform would like to instead have
that organized you know in this column
uh way over a set of vector
registers so that's such an instruction
exists in in for rmsv but we don't have
any equivalent for example on the case
of of risk five
interestingly um in the case of Epi
there is actually a custom extension
which is not part of the vector standard
but that targets
only two vectors so that was not
something that we could uh that was
useful for us so we tried several
different Alternatives like doing a a
unit stri at store Then followed by an
index load or doing a strided store
followed by unit strided load which more
or less um perform the same the at least
these two because as you can see they're
basically doing the same operation but
just you know following a little bit
different uh
logic but any case so we identified an
option here to maybe come up with a
similar instruct to further speed up
this uh
process and the final challenge that we
faced that was both related to with
performance and also programmability
aspects is that it turns out that in in
the case of RSV it is possible to
declare a vector type at the the sea
level and then take a reference for it
and this this
allows
to well this was um allowing us to solve
a problem that we had here in which we
were trying
to let me see did I describe it actually
on the
slide okay so the idea is that in the in
the trans in the Transformations there
is
a there are some operations that need to
be done over vectors and that they need
to happen on all every time that you are
calling these um these three
Transformations the input transformation
the kernel transformation and the output
transformation but the registers the
vector registers on we on which we want
to operate they keep they're not always
uh the same so there's a challenge that
okay so if you want to for example
create a procedure to do this
transformation of the of the vector
registers then you would have to
basically create use a set of
intermediate Vector registers on which
that would be used to pass the
parameters to this function and then
return the parameters the vector
parameters back and that would of course
well that first of all would increase
the register
spilling and furthermore this actually
reduce programmability because we in the
case of risk five such a passing of
references is not possible so we would
have to actually write the code which
was about 30 lines of you know of code
in six different places in the program
so leading to decreased
programmability so that was I mean
that's the what we ended up doing and
but we're trying to look into ways in
which such a possibility to try to you
know pass register references to
procedures if that would be possible in
the case of risk five it would certainly
help very much for this type of uh
scenario
so
overall without going to too much detail
we found that compared to the
traditional approach that uses IM am to
call theog implementation that we did on
risk five was was about 20%
faster and on the on the gem five it
would actually achieve comparable
performance to our optimized rmsv
implementation so that was sort of a
good um outcome
and then also because we wanted to look
into what is the actual efficiency of
this and scalability you know more as a
in terms of code design to try to
understand how useful is it actually to
have longer vectors or how useful is it
to have longer caches so we did this uh
study in which we tuned the parameters
of the L2 cach for example from 1 to to
56 megabytes and also the length of the
of the vector Reg registers from 512 to
4K and what we learned by doing this is
that in terms of vector length there was
not much point in going Beyond 2K uh
vectors
and also the performance the impact of
increasing the second the last level
cach in this case was the second uh
level cach sort of saturated after 64
megabytes but combining these two
uh factors
together we predict that using an
implementation that would have 2K
vectors with 64 megabyte last cash would
achieve about 1.8 performance
Improvement
yes but this is for a Tiny Network yeah
for a vgg which which is minimal if if
you have a something bigger did you try
something bigger than that the other
thing we tried to is Yolo which is also
tiny
yeah no
so we didn't go for for for long larger
one of one of the reasons is actually a
bit related to the practicality of the
approach when you when you're using a
tool such as gem 5 it's actually
extremely slow and if we go for a very
long Vector that was not Vector so very
long large model that would not finish
in fact for YOLO V3 we didn't in fact
simulate the whole thing because it
would simply not finish in a reasonable
amount of
time but
[Music]
um I mean I'm I'm actually interested in
your suggestions for other types of of
network are you talking about CNN
related or or non convolution I I would
say like if you go for a Transformer it
would be much bigger yeah so those
numbers will will still keep contining
going down yeah that's that's the thing
no I I agree this is yeah I mean we are
very interested in looking into to
Transformers for a moment I thought you
meant larger
cnns but no if we talk about
Transformers that's a it's
a that's actually I would say that it's
sort of future
work okay yes and but okay so this work
was done actually in a on a library
which is called darket which is not
really that commonly used anymore so one
of the things that we are also working
on right now in fact I have a master
thesis going on right now is to try to
put these algorithms like the vogr the
IM to call or also an implementation of
a direct convolution inside of 1 DNN so
1 DNN is a much more wellknown or much
more commonly used library for doing um
machine learning
operations and it's um open
source it's uh actually was originally
developed by by Intel so it's part of
this one API stuff but there's also
ports for many other architectur so you
can use it also arm power and also
there's uh risk risk five that we are
actually developing inside of the pilot
project as
well I will not show you any any
performance results but I want to show a
little bit about what we are currently
doing because it's uh with one DNN we
actually are making use of this sort of
jit approach that I was describing so
it's also and it's also sort of an
important tool to support in this
ecosystem in fact the work itself on 1
DNN wasn't started by us it was started
by by
BSC um it's mainly um marasas and one of
his students that has been working on
this and they come up with this they
implemented a jit approach in 1 DNN
basically you if you want to for example
um you know have generate an instruct
which in this case is a is a vector
additional instruction you can from the
high level you will just uh call a
function which is okay push this
instruction and then later on you can
generate the code and and execute
it so using jit has its pros and cons I
would say two problems compared to using
intrinsics for example a problem is that
you need to keep track of the registers
manually because here you are really
actually you know specifying the exact
registers that are going to be used you
cannot use high level names that the
compiler can then transform for
you and basically it's the same as
creating an assembly version of your
code
the
potential benefit well in this case is
that this
Approach at least makes it easy to
extend to add instructions to the jit
and why one would like to use jit itself
is as I mentioned before because you can
um it enables certain types of of
optimizations one sort of simple
optimization you can think of is that if
you have you know at some point you want
to for example call A A convolution for
example let's say you want to execute a
matrix
multiplication and uh you know EX act if
you when you write your code for your
matx multiplication you will have okay
you have for example so many rows and we
will have so many columns and then we
can come up with some blocking but you
know you have to sort of support a range
of sizes but if you do it on a jit
Approach at that time moment you might
know that ah the Matrix I have to
operate is exactly 128 rows and 64
so I can then propagate this information
and simplify the control flow inside of
of the of the routine to create a
specialized version to be executed at
that point so there are instances in
which having some you know some of these
information which you know only while
you are executing will allow you to
generate more more efficient code so
that's why we are um interested um into
this
and I have here a couple of slides that
just show a little bit how this um
approach um works but I think I'll just
go quickly I'll just sort of skip it
maybe I'll just stop here just you know
because it's kind of maybe interesting
since I've been talking about the the
intrinsics that have been developed in
the Epi project so you can see here an
example of how the in code um for
intrinsics uh with intrinsics looks like
and how the code that is then generated
with the with the jit um ends up looking
I actually know exactly for which
function um this
is but um so you can see that in the
case of intrinsics you can still you
know make use of um a call make use of
variable names and do for example
arithmetic on on pointers and then pass
that to a to the intrinsics will will
then generate the corresponding load
instruction but this sort of arithmetic
that that you can see for example in
this first instruction is something that
you would not be able to do in the S
side of the jit here you have
to know have to pre computed exactly the
value in order to be able to generate
that
and all
right I'm switching a little bit to a
different um project that was also work
that we have been looking at was to
optimize the
openmp for execution on the on the risk
five in fact this was more started as a
more generic we started also looking
into arsb at the beginning and also into
the
Intel simd but with a long-term goal to
support risk five vectors U and so well
here the problem where we focus
specifically on synchron constructs such
as barriers and and reductions in open
andp and as you know the problem is that
well as you if you look for example on
this graph here and where on the xaxis
we have number of threads this is
actually run on a Intel
K&L
and we look into sort of the overhead
that results just from this operation
you will see that the more threats you
add the performance uh degrades so
question was can we you make use of uh
of vectors can we utilize Vector units
to reduce this
overhead and by the way this this Intel
KL was located at the hpc2 so this is
links it's the one that you had so we
got access through it via Snick at some
point anyway so
well you know this is a supposed to be
an example of of a barrier a barrier is
a construct in which you have multi mple
threats that have to sort of wait all at
the same uh place and wait for each
other and only when all the threats have
arrived then they are allowed to proceed
so one way to one attempt to vectorize
this is shown here on the on the right
side let's assume all the threats when
they reach um the barrier they will
activate a bit which is a part of of a
vector and then the primary threat can
try to load this Vector using a vector
load and then try to see if all the
threats have have reached so very um
simple um idea similar can we also
applied for reductions there it becomes
a bit more complex because we have to
apply the reduction and we also had to
modify clang but for the barriers we did
an implementation that was only inside
of the llvm openmp runtime and we
created three different versions one for
INT AIX one for rmsv and one for for the
risk five Vector
extension now the results that I that I
have are only shown for for the KL
machine and a64 FX because at that time
we didn't have access to test still on
the on the hardware for the for the risk
5 but um we have but what we did do for
risk 5 is to validate that because we
used at that point we used an approach
using using this uh vave tool which is
the emulator for the vector instruction
which I mentioned before and that was
coupled with Cho so Kemo is
um well I will talk about it a bit later
but K is basically an emulator for a
system level
emulator okay so we did this result and
at least for this for the KL platform we
observed that we could reach uh already
speed up just you know just launching
barriers at this point
of up to a little bit over 2x so we
haven't been able to test this yet on
the on the Up pilot prototype but that's
sort of one of the next steps and based
on that we'll see if we can provide some
feedback back to the hardware team on
the performance of the of the
barriers okay I finish here with this
this a little bit of things of more
research like that we have done in terms
of risk five and I will try to finish
a little bit asking more a question or
ask or sharing some some thoughts about
how I sort of see the the road ahead for
for risk five in in terms of
HBC these are of course my my own
thoughts and they are probably biased
and and faulty in extent a lot of things
are happening in in parallel in fact in
this field it's difficult to keep track
of of
everything but what I can tell you is
that well as you have seen there's a lot
of progress happening actually in at all
the levels there's work going on from
the hardware side on the instruction set
architecture and also on the software
libraries and Tool chains but I think
there there are challenges at all these
levels that still need to be need to be
face so I'm going to I want to talk a
little bit about these things you say
going from
Hardware Isa to software a few other
thoughts and then even share some
thoughts I me what you could do if you
would if you're interested for example
to look a bit into um risk five or you
in a in in a general
sense okay
so there are if you want to let's say
start with risk five Hardware there's
actually a lot of
um small boards that Implement risk five
but if you are looking specifically for
HPC it's a bit more more more difficult
so I try to I mean based on my
limited view of it of course I haven't
read all the papers and everything that
is happening all the time but i' I'm
only aware of let's say one um board at
least that I have seen having been
evaluated independently and that's this
um well this softphone SG
2042 which consists of a set of cores
from a company called Tad and this one
is implementing the vector spec 0.7.1 so
it's not implementing the one
1.0
and I I sort of highlight this
independently evaluated here because I
see that there's a lot of I mean there's
a lot of projects and even I can see
pictures of people saying okay we are
developing this chip and here's a
picture of the chip but I have actually
not seen anyone use that yet or having
it evaluated but here there's a a paper
in fact this is by um I think it's by
Nick Brown um He was discussing about
this in the risk five Workshop of
supercomputing last year where they look
into this this um system one of the
challenges is that this does not
implement the last version of the vector
spec so it actually requires a custom
GCC compiler so from terms of usability
particularly for that Vector instruction
is actually not um not very
convenient
but supposedly it looks like there is
um there's light at the end of the T I
would say there's a huge amount of
companies and many of these companies
have a have a large amount of of funding
and and backing that have
announced or have prob publicly
disclosed that they are working on high
performance um hardware and and often
they there will there are systems
specifically targeting HPC and also
systems targeting AI so here this is
just a a list of the of the ones that
probably come to my mind when I when I
prepared this slide but you know
companies such as T torent they are
developing a large outof order risk five
cores plus an AI chip which is called a
106 there's Ventana micro which is also
developing a out out of order processor
with you know with Vector extensions
espiranto Technologies they have these
two types of of course they have some
very small risk five cores that you call
the the minion and I think that's that's
mostly for for AI type of workloads and
they're also developing a host processor
the ET maxion semidynamic is company
which I which is part of Up pilot they
have developed for example out of order
code which called ATO this has sort of
of a interface for a for a vector unit
and I think that they have their own
implementation of a vector unit outside
of the of the vector unit that I was uh
discussing in the EU pilot project and
they they have also in fact a set of
custom tensor instructions that they are
developing for for AI workloads Inspire
semi has a has an accelerator platform
which they call
Thunderbird and you know there are
probably more these are I think these
are all companies I mean most of them
are us companies sem dynamics of course
located in in Europe but I'm sure that
there are companies many other places in
the world that are currently developing
Hardware
so at some point we are going to get to
the point that we will have more
availability but right now I am
personally not aware of any of these
systems that can be accessed I might be
wrong but uh I haven't I haven't come
across the that yet the good thing is
that all of
them from what I could understand
support the risk five Vector 1.0 spec
and that's good because once these
systems become available it should offer
interoperability you know you should be
able to run code on one system it should
also be able to run on the other systems
and that's of course you know very
important for for being able
to you know not only deploy software but
also be able to compare
systems but later we'll see you know it
will be very interesting to see once
these systems become available if it is
possible to start running benchmarks on
them to actually see how they compare in
terms of uh
performance now what's going on in terms
of specification so I've talked a lot
about the vector spec and I think that's
certainly a big step um ahead for for
being able to support HPC in in in the
context of risk five now the
there is still a lot of discussion going
on in the in the what's it called the
special interest group for and the
vector s in in Risk
5
luckily um you know everything in in
Risk 5 is kind of open so everybody can
just go in into the archives for the
maing list and see what people are
discussing what's going on what sort of
discussions for new versions of the spec
are are happening
what you cannot do unless you are a
member is to join the the meetings
themselves and try to to uh participate
for that you need to first become member
of the risk five International which you
can do as an individual or as a
organization depending on whether your
organization is paying your work on risk
five or not but anyway because of
that I I have been you know been able to
attend some of these meetings and also
look to back blots of of mailing lists
to sort of try to understand what is
going being discussed and here are some
things that I've seen that are sort of
topics of of Interest inside of the
vector c one is um that there's lot of
interest in subsetting the the current
uh Vector
specification what this means is that
you know we now have a vector spec but
this Vector spec is actually very large
it's it's actually well the last time I
I checked it it has around maybe 200
instructions and it's uh implementing it
in its totality is quite complex in
particularly for embedded devices you
might not be interested in implementing
the whole thing so there is discussion
for example in separating the vector
specification into smaller
sets that probably hasn't too much
impact on the HPC part itself but what
could have an impact is is the next
thing
so self-contained Vector instruction so
risk five Vector instructions are
currently 32 bits and that's not a lot
of space to to specify you know many of
the operations or that are required for
these vectors particularly for this
Vector length agnostic design so the the
vector spec currently is based on having
a lot of control registers that support
and that impact the execution of the
instructions but there is you know
there's a discussion on having a new
version of the vector spec that would
have longer Vector instructions maybe 64
bits where all this information is
encoded as part of the instruction and
that could for example enable having
more architectural registers right now
we have 32 Vector registers maybe in the
fut a future version of the spec would
have
128 registers and then you know that
will have a big impact on on many of the
of the applications for example the
vinograd implementation that I discussed
we were limited we had problem of
register spilling if we had 64 or 128
registers that probably goes
away then other things they're looking
is some sparity support and also trying
to have more GPU like capabilities in
the in the spec um
itself other things that are I think
very important particularly for AI is
that there's a lot of interest in Matrix
extensions and before we had a bit of
this discussion about the customization
whether you know you need to give back
any extensions that you make to the to
the spec or or not well this is a field
in which you will see a lot of custom
extensions because right now we do not
have any standard extension inside of
risk five there are two working groups
one for what is called for the
integrated Matrix extension which I I I
think what this means is that they it's
a matrix extension that makes use of the
Vector registers for storage of the
matrices and then attach Matrix
extension which would then be more like
a co-processor type of Matrix extension
with externally stored mates on separate
registers so many many companies are
currently have developed their own
custom uh matx extensions for these
types of workloads
but the natural you know
flow should be such that at the end of
the road this um extension somehow
convert into an actual
specification and another thing that I
also highlighted here I mean I think
there are more things in the risk five
spec that impact HPC but also I want I
think I mean personally I've always had
a big interest and I think it's very
important the topic of performance
monitoring support so there is I mean
risk five does have a spec for um for uh
performance
monitoring but one thing that is going
on now is that right now the
specification for performance monitoring
just uh specifies how the how to access
the counters for example but it doesn't
specify which counters need to be
available so right now there is a
process to also have a specification of
a set of names for standard performance
counters and I think that's also going
to be very helpful because it will
guarantee that we can you know run papy
or whatever or perf on all these
platforms and and hope to get the same
counters to so that we can use the same
Performance Tuning methodologies on all
these platforms so that's Al hopefully
something that is going to be there also
in the
future I have only a couple more more
slides I wanted to mention something
about software tool
chains you know obviously for the
success of of risk 5 we will have to
Port a lot of you know software we have
seen how how how long it took for in the
case for arm to become competitive on
the HBC side same sort of effort needs
to be done in the case in the case of uh
risk five I have dis discussed some of
the work that we have done in epi new
pilot and here's just you know I try to
go quickly over that slide that I showed
earlier to highlight some of the
libraries that we are actively uh
working on and some the tool chains but
there is more work um I mean there's a
big interest on this
topic of course and um in fact there is
there is a industry Le effort which is
the rice project which is the risk five
software
ecosystem and you know they are looking
about what has to be done in all these
uh levels to try to accelerate the
development of Open Source software for
risk five so looking you know as
compilers GCC and lvm system libraries
SSL GPC blast
kernel most of focusing of course on on
Linux and
KVM managed uh run times Linux
distributions debug and profiling
support simulators and and system
firmware
so lot of things happening in in in that
front um as
well so some ideas more from our own
experience working on risk five is that
you know while you might think okay you
know it's just a new architecture let's
just recompile and we are done this is
of course far far from it one challenge
I mentioned this before is that a lot of
codes that we want to be able to run on
risk five uh in the future on risk 5
accelerators are currently being
developed for example only in a in a
language such such as
Cuda so there's a question how can we
make Cuda programs be use um r five do
we need to rewrite basically the the
application to use um more like the
vector way of programming with
intrinsics or trying to rely on open p
simd
or maybe an alternative approach is to
use automatic conversion from Cuda to
something like CLE there are some tools
that and that enable that and then try
to have a high performance
implementation of
CLE and there are there some work going
on
there so in the end yeah just point out
that risk 5 is not a it's not a GPU so
those codes will still require some
effort to be
done and in the context of uh you pilot
one thing that I I see is that you know
we we are working on trying to get to do
long
vectors but um the way I see it being
able to to fill long vectors is is very
much algorithm dependent and a lot of
code that for example has been developed
with Intel like simd like x86 in mind
will not easily translate to very long
vectors because there you have more
shorter vectors up to 5 12 bits and when
you know that and you program
specifically for for that you will
already organize your algorithms
targeting that you will not write your
algorithm to be able to extract vectors
of thousands of
bits so we are we will need to do some
research into either more advanced
compiler support to be able to fill
these longer vectors or at the end the
developers themselves will need to
actually do some um algorithmic
transformations to to expose the
required parallelism to fill these long
vectors
and yes I guess I'm
I'm almost done one thing I also thought
that maybe could could be interested if
you have never come across the risk five
but you're a little bit curious just to
point out that there's a large amount of
boards these are mostly you know small
boards not HPC like boards that you can
play with you can go to this uh this
link and there's a long
collection also if you a bit more
interested in the HPC site or the work
that we're doing in Epi pilot I want to
mention that you know these sdvs that I
mentioned before with the fpga and
Unleashed boards that's actually open
for external access I think if you sent
a email to philipo mantoani you can
request access to that and you can also
play with that tool chain there if if
you don't want or you don't have
Hardware or you don't want to have
Hardware there are of course also
Alternatives spike is sort of the golden
reference for the R 5 East as an
emulator
and and if you want to Le something
that's a little bit more complete Kimu
is actually part I mean there's a lot of
good support for race five in kemu it's
part of also of the goals of the rice
project that K is well supported and I
think there's a commitment that all
ratified instructions need to be in in
kemu so you can actually get quite a lot
of done by using just K to start playing
with risk
five and finally this is maybe more
interest of me because we do you know a
bit more computer architecture and
performance
evaluation if you are in that field then
it's also good news that gem 5 since the
last release in the in December 23
released 23.1 now also includes a model
for the the supports the vector 1.0
standard okay so
that's I thought that would be all but
now of course I have my a conclusion
slide and I have an acknowledgement
slide
too so well we have seen a lot of
progress we have we are working in Epi
pilot working on a on a on a multicore
system with a vector support and a
software tool chain and there's also a
lot of work at the commun
as I said I think that the vector is a
big step forward for being able to have
um high performance Computing based on
on risk 5 but I think we'll need to have
more boards more HPC like boards to play
with and Performance Tools so that we
can you know start optimizing the code
you know using more serious uh
platforms yeah we hopefully the HBC
Hardware will also so converge around a
set of more stable specs and I didn't
talk about something called the risk
five profiles but that's actually sort
of the way risk five is trying to get
vendors to F that focus on the same
Market to use the same set of of vector
extensions so not Vector extension of
risk five extensions So to avoid
fragmentation and well I've talked about
some specific challenges in this talk
that I particularly think are
interesting at least for us that are
looking at is autov
vectorization for extracting long
vectors or more systematic ways to
rework algorithms that will allow to
extract longer
vectors also conversion strategies to
move from Cuda to be able to execute
Cuda programs on on risk five and
finally I think it's going to be very
important that the working groups that
are defining the Matrix extensions try
to come up with uh some standards um
rather soon
okay so now just I wanted to acknowledge
my that of course all this work has been
done by my team our we call sort of
informally we call us the chart team
chmer heterogeneous architectur run
times team and yeah and I you always try
to make use a plug here to just point
out that we have also very nice master
programm in high performance Computing
systems and then if anyone is
interested then please
come okay I have absolutely no idea how
long it took but I we're
done we started late but you're pretty
much on time if all right okay well
thanks thank you for enduring me until
the
end take some more questions maybe if
there are any I'll be happy
I'll be around in any case until
tomorrow evening
so I think there's a
question well so you said about the
difference between an open specification
and open source Hardware I I didn't get
exactly what is the difference in this
case open specification just refers to
the instruction set architectur to the
language that the hardware has to talk
so you can you can take the
specification you can customize it
change and add new instructions if you
want and nobody's going to stop that in
that sense it is open but Open Source
Hardware would mean that for example you
publish the RTL the verog or HDL like we
are publishing the C code or C++ right
so that's sort of the yeah the
difference you talked about emulation a
bit as well through keu for example when
you're playing around with risk 5 since
there's to some extent a lack of
Hardware to actually run stuff from and
in a previous risk five talk I saw is
that it also depends where you're
running q u in terms of getting a better
emulation to some sense or at least
better closer to the hardware like
running qemu on an arm system makes more
sense than running it on an x86 system
because the memory model of arm is very
similar or is at least a lot closer to
risk five than it is for for
C6 I mean what you say makes sense I
personally don't have that experience I
don't have the opposite experience I
simply haven't tried that so I can't
really well what he was warning about is
that if you can pick run on arm because
then it's less likely that when you're
just transferring your binaries let's
say to an actual risk five platform that
you're then running to surprises there
which you which which you're more likely
to overlook if you're running Q on x86
so there's a there's a factor there as
well at least no so I I think I
understand what what this means I mean
in x86 you have a stronger memory model
that's that's usually called TSO total
store order um and in Risk five and arm
you have a weekly modeled uh weekly
ordered memory
model and if you have some specific
algorithms that happen to work because
of the TSO guarantee it might be that
running those algorithms on K on x86
work on the risk five emulation but
actually don't work in practice that
yeah that could happen I hav okay but
um it's an interesting question itself I
mean how can the how can k
itself remove that those guarantees but
it's probably maybe too too much to ask
I don't know yeah it it would only lead
to further slowdown which is probably
not what you want you want to use the
underlying Hardware as much as you can
and and I think it's just an important
detail that people should be aware of
like if you can just run Q on on Q on
arm particularly if you are developing
let's say like lock free data structures
or stuff like that that would be very
important then I think okay and another
thing is that had a slide with lots of
open-source projects that are uh
starting to look into risk 5 or are
being ported to risk 5 do you mean the
the one on on Rice yeah the one like
yeah yeah indeed where like BL open
blast and all these things are are
mentioned M I'm not sure if you have any
experience with that but what's the
situation like when a project like
rise gets in touch with these projects
and says we want to add support for risk
5 like do they actually know what risk 5
is
I mean for GCC that's a given they are
well my understand that well I mean the
people in rice themselves are doing are
doing the work um but then is are the
projects also accepting their
contributions do they understand why
it's interesting to do that I mean
that's probably a question that I cannot
answer but
I I would say I would probably say why
not I mean do you think why do you feel
that there are any specific concerns
about it the reason I asked is we a
couple of years ago we had an online
talk by the the Bliss maintainers so
Bliss is like it's an alternative to
open blast let's say a blast leack
Library yeah actually I think we have it
yeah bliss bliss is there now but when I
don't know how long ago this was but I
think it was five years ago when they
yeah they gave a really in-depth talk on
Bliss and everything they do um to make
sure you're getting good performance
Cals and so on and I asked them the
question about risk 5 and they had no
idea what risk 5 was now that's that's
five years ago things are very different
today the people from from julick right
developing it or no the um where are
they in Tennessee Tennessee is it or no
or Texas no
Texas University of Texas I think yeah
that surprised me a bit because if Bliss
doesn't know about risk 5 they're so low
level they're basically hand coding
assembly to some extent well BL not
really they are trying to avoid that but
if if they didn't know about it five
years ago then I wouldn't I'm pretty
sure that many let's say scientific
applications currently which are three
four levels higher still have no idea
what risk five is and some of them at
least should like if you're touching
assembly anywhere you better know that
this is coming because it's going to
take some effort to make sure you're
you're being ported and this is just an
observation and it's not a criticism or
or anything but that's I I think it's
still a challenge it's improving but
it's it's a challenge to explain to
people that this is coming and why they
should be aware and maybe try to prepare
for
it yeah no I I mean I I don't disagree I
I don't really have I mean I would like
to know myself how much awareness there
is because I live in this risk five
bubble and sort of something think
everybody knows about this but then you
go out I mean but here for example I saw
that most people are aware of it but um
if you go out some completely different
Community it's a good question but
um yeah I
mean like um in for example in this
European projects or the way it it works
it's a little bit like okay we Define a
project we try to select which are the
key applications that we think are
important and then we sort of reach out
to the responsible and try to explain
why we think it's important y I don't
know if um this works similarly in other
communities but at least this is one way
in which we try to make this awareness
it's right now it's coming it's sort of
the in the responsibility of the risk
five Community to try to convince that
people and of course the more projects
know about it it's going to snowball
into into bigger aware of course at some
point that hopefully should be the case
yes
you didn't mention any any timelines but
do you dare making any predictions on
when we'll have let's say an the top 500
or yeah let's say that yeah that's
that's an easy okay recording well but
by then it may need to be ex scale to be
on top
500 it's very diff it's a bit of antic
question is it um no I could try I mean
um so my I think my main challenge to
try to answer this is that question is
that let's
say I don't really know how advanced all
these players are I know that that they
have you know huge teams working on it
so they probably have the capacity to
put the system there in a in a in a few
years but uh you
know if I don't see the system and see
some numbers from somebody testing it I
I it's very hard to understand what is
the maturity really because you know
showing a chip ah here's the chip but
you know what does that
mean um but uh
I it could
be you
know optimistically maybe two to three
years if if these people are really
where they claim they are or more
pessimistically I would say me five six
years that that's getting the actual
chip that doesn't mean you have a full
well a system but yeah but if but I'm
I'm saying getting getting no but at
that point I would say that the the chip
should be having everything in place
let's say so you can if you want to you
could build a big system
yeah I mean some of these players seem
to be close to that at least from what
they say publicly I mean t stor is not
it's a startup in some sense but there's
people behind this like Jim Keller used
to work at at Intel pretty much
everywhere so he knows his stuff quite
well uh that doesn't mean many many of
these have very wellknown people backing
them also asiro Technologies um Dave
dsel I think behind I mean and
um and I don't know which ones but they
have huge teams in the order of hundreds
of people uh working on on that and with
a lot of funding
also most okay why and I mean is one of
the things that I pointed out here is
that many of these they have you know
some chips that are specifically
designed for AI so I think one of the
reason that they have big funding is
because they there's people who hope
that they will be able to get some part
of the cake that Nvidia currently has
but if that but maybe it's if you know
if they make good use of their money
they they should be able to reach there
in a reasonable period of time
MH let's see maybe we can invite you
again in three years from now and then
can see that the
updated thanks I think one way forward
is to build a cluster specifically for
AI because this is a bus word you get as
you quite rightly said money for that
and demonstrate risk five is actually a
decent and you have to take it Zs
competitor because otherwise you've got
this chicken and egg problem we are
interested in but it is new technology
and we don't have the funding to do it
and because it is new technology we are
not interested in buying it because it
hasn't been proven it is actually good
um think about the Lumi project for
example where they deliberately used
gpus from AMD they had quite a lot of
problems to actually Port all of the
software into it and I'm pretty sure you
will have similar issues with the risk
five if if if you see what I mean I try
to be positive M just to be
clear I mean many of these are sort of
trying to do what what uh you suggest to
try to build a demonstrator for for AI
but because it's like that I mean if you
don't have the hardware or can show that
it works it's you're not going to you
can't convince anyone to to buy into it
yeah and we are sort of trying to do the
same thing but for the HPC the HPC
does unfortunately not have the same
budgets that the AI has so that's why
Euro HPC puts more of the public money
into it because that's otherwise you
can't go
there yeah I think we can wrap up here
is going to be around for more questions
today and tomorrow
so I guess we can grab a quick coffee
and then just jump straight ahead
thing