Submind YouTube summaries
Thumbnail for EUPILOT: towards an All-European RISC-V-based HPC demonstrator

EUPILOT: towards an All-European RISC-V-based HPC demonstrator

Watch on YouTube

Video summary

The EUPILOT project, funded by the European Union under EuroHPC, aims to establish a sovereign High-Performance Computing ecosystem based on the open RISC-V architecture. Led by Barcelona Supercomputing Center with participation from nineteen partners across Europe including Chalmers and KTH, this initiative seeks to reduce dependency on non-EU technologies like ARM and x86 while enhancing security and fostering domestic industry growth without royalty fees. The core of the demonstrator involves an autonomous accelerator platform integrating eight discrete chips connected via immersion cooling technology at BSC, featuring two primary accelerators: one combining out-of-order cores with RISC-V Vector Extension 1.0 for general tasks, and another dedicated to machine learning stencils. To overcome significant challenges in hardware verification and software-hardware codesign for energy efficiency, the project employs simulation tools like G5 alongside FPGA prototyping within an LLVM/Clang toolchain environment. Researchers are optimizing applications such as convolutions using advanced algorithms that leverage long vector registers up to 16,384 bits enabled by ELMO, while managing parallelism and handling missing instructions through custom extensions or intrinsics. These efforts have yielded optimized implementations achieving twenty percent faster performance than traditional approaches on Gem5 simulations for small networks like VGG, although scalability studies indicate diminishing returns beyond specific vector lengths and cache sizes, with larger models still presenting simulation hurdles that require further algorithmic adaptation. Beyond immediate technical optimizations, the project addresses critical software porting issues, particularly regarding CUDA workloads which often necessitate automatic conversion tools or manual register management to function effectively on RISC-V hardware. Emerging support for Vector Spec 1.0 from companies like Tenstorrent and Semidynamic is expected soon to ensure interoperability, while future specifications aim to include self-contained instructions, expanded registers up to one hundred twenty-eight bits, sparsity support, and standardized matrix extensions alongside unified performance monitoring counters. OpenMP optimizations have already demonstrated the ability to reduce barrier overhead by over two times on existing hardware like Intel Xeon Phi, highlighting the potential for significant efficiency gains as these technologies mature. Looking toward the future, while dedicated HPC clusters are essential to overcome current funding and awareness challenges compared to AI-focused initiatives, experts predict that RISC-V systems capable of competing with today's leaders may emerge within two to six years. The ultimate goal is iterative optimization leading to competitive exascale systems in subsequent generations of hardware, ensuring a robust European infrastructure for high-performance computing. By integrating these advancements into widely used libraries and adapting algorithms for long vector registers, the EUPILOT project paves the way for an independent, secure, and customizable HPC landscape that can sustain Europe's technological sovereignty without relying on external instruction set architectures or proprietary licensing models.
Read the full video transcript
should I start yeah all right um welcome everybody thank you for coming uh to my talk so I'm my name is Mikel peras I'm an associate professor at shmmer um I thought i' maybe start by saying a few words about myself because I guess most of the people don't don't know me so I am I'm from Barcelona and I actually did uh PhD at the UPC while I was working at Bor and super Computing CER Center so my background or what the work that I was doing at my PhD was mostly in the field what we call computer architecture so it was mostly design of uh microprocessors and trying to you know find novel organizations for Designing high performance processors so I was working at uh BC also at that time so and I have an interest at in HBC and supercomputing and I also spent after working for BC for a few years I went to two years to to Tokyo I was at the Tokyo Institute of Technology where they have this supercomputer which is called subam or the series of supercomputers there and at that time I started working more in term in topics related to programming models and Analysis of runtime systems performance analysis of runtime systems and after that in 2014 I moved to to shmmer shmmer is in for those who don't know shammer university is in gothenborg which is City on the west coast of Sweden and I've been there now for for 10 years working mostly on topics related to programming models and codesign so I well at shamers we are uh involved in several uh European projects that um are related to risk 5 and one of them is this U pilot project so I got an invitation from Kenneth for which I'm very thankful to provide uh well our experiences to talk a little bit about what's going on so this is how I've organized the talk I will start a little bit talking about risk five and what's the interest of the a European Union in Risk five and mainly focuses on on two very link projects the European processor initiative and other project which is called up pilot which I will present in a bit more detail and then I thought okay you know since I give the chance to talk to you I want to also talk a little bit about the work that we are doing at at charm so I want to talk a little bit about some research directions here that we're doing pursuing and finally I'd like to share some some thoughts about my general feeling about where we are I mean where risk five let's say is in in the context of uh you know in HPC and what is what needs to be done so before um starting I'll just like to ask a general question I mean how how many of you have heard about risk five okay I guess that's most of the people that's good sign I'm not going I mean I'm a teacher but I'm not going to ask you not to go out and try to Define what it is so which is what I would usually do but then we would need many hours so but anyway I think I should probably try to provide some introduction into risk five so risk five is an open instruction set architecture it's based on the on the risk philosophy and it's very much unlike other Isis that you have probably very familiar with such as x86 or arm in the sense that it is an open standard which is not under the control of a single company if you don't know what an Isa is well Isa is basically can understand as the language that the processor understands so x86 armr wellknown isas or isas now in opening this context means that it is freely available uh for everybody to use and also to modify and I'll come to that also in in in a serious in a in a moment many people think that risk five is an open source processor but that's not true risk five is an open standard it's open source in the sense that the sources for them for the ISA are open now the advantages of having an an such an open standard is that well you don't have to pay royalties to any company if you want to make use of this Isa you can it enables higher customization and as a as a vendor if you implement a processor that makes use of this Isa you can benefit from the fact that there's already a shared well shared software ecosystem but I put you into parenthesis that this of course depends on the fact that you that you have to adhere to the standards so in a sense this bullet of customization that's a good thing but if you are going to customize some parts of the processor make extensions your own extensions you will have to develop your own software environment for that of course and of course there's a you know it's a very it's a great Community very B brand community and a lot of people have a lot of interest so it's easy to to reach out and get support yeah I mean yes I'm totally happy it depends on when you say about the customization uh are you able to customize it uh in the sense it's like a GPL thing that you have to give it back the customizations or you can do propietary ones for yourself no you you can do proprietary ones in fact it's there are yeah okay you will see that vendors are all the time doing this thank you okay now as you saw in the in the tiet and in the title in this talk I want to focus a bit about the European perspective of of risk five and you might have noticed that risk 5 appears often when in the context of the eurohpc joint undertak so in this slide I actually tried to take some different views from various sources but one of of them um actually the whole slide is not visible so I'm not sure if maybe if I click it's going to disappear twice no so now it's visible but it's still with the with the bar there I just want to point out that there is a in the in last year's risk five EUR Summit Europe there was an interesting panel exactly on this on this topic so if you want to go into more detail what I discuss in this slide you can actually go and look into what all the panelist had to say about the topic but one important reason that the you is interested in um in Risk five is that it wants to achieve uh or develop an autoctonos HPC industry and we actually seeing this trend in many many regions on the world and many governments that feel that you know the the way that the semiconductor suppli is currently organized around a few select companies and a few foundaries it's a bit of a risk situation there's a chance that you might get cut out so many governments feel that moving towards an open standard provides a higher or high security and less risk so there's an interest in developing uh such a our own HBC industry then there are also needs such as I mean in interest in increasing HPC capacity development of HPC skills and also some people would like to see a movement towards Open Source Hardware there's a bit of a debate at this so the the vision is that okay up to you know like the Linux Kel from the software stack there's High we have a lot of Open Source software but and below the kernel it's mostly closed so it would be nice if you can also have a Open Source Hardware there is an Open Source Hardware movement but there is some debate as to how how viable this is particularly one the the main one important concern here is that unlike software which basically you know compiling can run on your system if you want to do Open Source Hardware if you want to actually tape that out in let's say high performance node that has a huge amount of cost so somebody will have to pay for it now why did uh why is the interest in in Risk five specifically well it's sort of because none of the other options actually fulfill um the requirements as you know x86 and arm are controlled by entities that are non EU entities at this point and that's actually you know if you if you watch some of the look a little bit at the history of all how this has developed a lot of people track this development of like the USBC projects now back to the series of projects that was called the mon blank projects at the BSC and those projects were based on arm they were taking first starting with arm um let's say embedded I might say or process of that were originally designed from smartphones to try to build clusters based on arm technology but then you know at some at some point soft Bank bought arm and well still under control of soft Bank to an extent as I understand and then the European felt like okay know we've been doing all this investment into a technology which we actually cannot control so that's one of the reasons that they suddenly start looking into uh risk five then also you can see that risk five can be a language to somehow communicate academic ideas with with industry all right and well this last bullet is basically what I mentioned that risk five is an open standard doesn't mean the same as open source uh Hardware okay so you probably heard about the uhpc joint undertaking so here's a definition which probably I took from the from the their own website says it's a joint initiative between the EU European countries and private Partners to develop world-class supercomputing ecosystem in Europe they are I mean it's they are funding several different initiatives on the one hand they are funding big supercomputers for use uh by researchers there's a long list of them I think the largest of them is uh currently is the is the Lumi supercomputer in in Finland but there's also an a site of arent well research and Innovation projects in which we have to well participate in a few of them there are many of these projects but three projects which are related in fact to risk five and where Hardware design is a is important is the European processor initiative this is sort of the project which I would say kicked off a lot a little a lot of this uh work this project had two um two goals one is to develop an arm host processor so the arm is still in the the picture here and and also an accelerator based on risk five Technologies in fact it's not a single accelerator there are several accelerators this usually called under the name of epac and which I think stands for European processor accelerator chip there are two ongoing Pro Pilots also now that which I will actually show in the next SL a bit which is the Up pilot and the the upex and there's also another uh project which I will not be focusing so much about today which is called E processor which also is has the goal of developing a risk five based processor with extensions for you know for HBC for executing bioinformatics applications so let me show you some slight which I actually got from a from a presentation from philipo mantoani he's one of the researchers at at Barcelona who is involved in the development of the risk five accelerator and I'm going to be showing I want to show you two slides which this is from a seminar that was in in fact in in ulik last year so to show a bit what is being veloped in this Epi uh project so on the one hand we have an Arm based general purpose CPU which is called Ria and you might have heard about this because as I understand one of the partitions of the Jupiter is going to be based on this technology so this is mostly well developed by by cyer which was a company actually funded within Epi to sort of drive this development and commercialize it then on the other side we have this epac which is actually a set of accelerators but one that is most interesting to to us I mean charmes also to me and related to this talk is the is basically a vector processor that is being developed inside of this project and this is called epac and why I I I show this because I want to sort of show this timeline and in fact in this timeline starting 2015 you see also this mon blank and the Deep project appearing here you will see what is sort of the strategy of the development so we have the the Epi uh project which has in fact uh two phases and these phases are developing these two processors the Arm based host processor that sort of feeds into this upex project which is a pilot to develop a system based on on the re chip and another pilot to develop a system based on the on the apar chip and that's the E pilot project and the idea is that these the outcome of these projects will feed into future exos scale systems so we already know that um um Jupiter will have some components from based on Ria and for the for the epac there is a well I would say there is a wish that some that the next Marin norom 6 will have a partition based on these processors okay so since um the topic is risk five I thought okay I'll describe a little bit about what exactly is being implemented in in the E pilot project these sides are not mine they are so Carlos Bol he is the the the technical coordinator for for the project and so I'm going to be mostly showing a subset of slides that he from a recent presentation that um that he uh delivered that was actually at high peak okay so well typical slide with uh where you can see all the partners involved in the project we have a set of well total of 19 Partners that's a set of acad Partners there's a set of industrial Partners involved um for some reason the figure on the on the right I think that's actually taken from the EU portal but it's only showing I think the location of the academic partners for for some reason but as you can see it's um sort of distributed among Europe the project is coordinated by by BSC and well here in Sweden we have shmer also kth is also U participating in in the project so what are the goals of of the of the Up pilot project so first and the the main goal is to demonstrate this preex scale accelerator platform so we want to actually deploy and show the you know uh a risk five based accelerator platform that is in a in a usable state and to reach there there are several things that that we are working of on one hand the project is well designing and doing all this validation steps and deploying the accelerator platform with a goal of maximizing European technology and and assets because that's always one of the goals also of the eurohpc there is um the project is po across the whole stack so there's a goal to do software and Hardware codesign to sort of Ian to basically Drive the hardware design from the software uh components with the goal of improving performance and achieving better Energy Efficiency then finally also there are the goals to extend open source into hardware for HPC some of the components that are being developed will in fact are in fact uh open source or some of the hardware is open source and to do everything based on the risk five instruction set architecture so this is a top level view of what we're trying to to build it actually has um two type of accelerator chips that I will uh describe in a little bit more detail next we want to integrate eight of these ships into an accelerator board that will then then be connected to a host server and all this goes into this uh tank so it's going to be using immersive cooling technology that's and there's already a space at at the BSC where this is is going to be um installed as as I understand okay this is a sort of high level view of the of the system and here's a bit more detail on the actual chips that are being developed so there are actually two two chips that that are both sort of derived from this epac um chip that I mentioned that had several accelerators so out of these accelerators here there are two which we are going to construct discrete chips out one is a vector accelerator which in fact combines uh a core that is developed by a company called semi dynamics that is also located in Barcelona so they have a core they are designing an outof order core called atraido and this will be connected to a vector processing unit that is mostly coordinated by BC that development that will that implements the risk five Vector extension 1.0 standard and on the on the right side we have a another chip here which is called the MLS accelerator MLS here stands for machine learning and stencil this is um as you this is designed mostly for for AI type of um applications we are not so much um involved in this work so I'm not going to be discussing much about the parts related to the MLS in this in this presentation at the end the goal is to tape this out in a technology so it's 12 nanometer Global Foundry so this as you can realize is not the most Cutting Edge technology but you know the if you wanted to implement this in like say like five nanometers the cost of the project would explore I mean it's not like going to one of the high performance nodes it's not an incremental increase in cost in fact it multiplies the cost by probably by more than an order of M magnitude and this is the software stack that we are working on so the we are the way that the project this project is organized is that we select a set of of applications of um yeah Target applications and then develop the underlying stack to try to make that work and here we have on the left side the components that are targeting the vector accelerator on the right side we have the components that Target the machine learning stencil accelerator so in terms of applications so the main targets that we are looking at is Grox which is mostly well developed and coordinated by by kth by Professor Eric lindal and uh there's also EC Earth which I understand is being developed at BSC and you know basically stuck off of many many libraries and tools all the way down to um the system you Linux and also of course the tool chains to compile the applications there so that's one of the efforts that is ongoing and that of course is a critical component of of the system so at at the end what we are trying to to achieve in terms of uh of performance is shown in this in this slide so of course one thing that I have to point out that whenever I'm showing the numbers here it's a this is sort of subject to constraints of budget and some you know the cost of a square millimeter changes over time so depending on variations some of these numbers will will will change but the goal for for ve is to develop a let's say a 16 core chip in an area of 46 uh square millim and such that each chip has eight Vector processing units and if you do the math for example considering that the fuse multiply at counts for for two uh flops then for 64bit Point targeting 1.5 gahz we will reach to 384 G flops um per chip that's not counting the double Precision of the of the core itself so in total that goes to 432 gflops per chip and since we want to have eight of these chip on a board so we hope to um deliver a system where each of these accelerator uh systems this eilot accelerator system each one has delivers about 3.5 teraflops now I talked a bit about the work that is being done in the development of the applications and the tool chain so and one of the critical components that is required to make this work is what we call the software development vehicles and so you will you can think of it as follows when you start developing a process process you want to start immediately developing also the software that is going to run there but of course at that point you don't have any hardware to actually develop the software on so what is the solutions that you can go for so in the Epi project and also in the E pilot projects we are relying on U several um options several platforms one this is the let's say the hardware protot type platform that we are we're using there as a software development vehicle which consists of um two type of um of uh platform two types of boards one is uh taking for example a set of risk five commercial boards then in this case we have a series of of boards from that are this high five unmatch that's um a board that has a series of chips with four uh risk five course that we can start to use to do some uh software development of course this is not the same as the chip that is being developed we do not have for example the the vector units here so in order to be able then to target the vector unit you need to do something else one option is that you can try to emulate the vector instructions here so there's for example a tool which is called behave when every time that you run a you know you execute a vector instruction on a platform that doesn't have support for Vector instruction so what happens is that you get a you get a signal right you will get a signal ill instruction signal something well this tool captures the signal and then goes to analyze what is this instruction sort of emulates that in software that's one option and you can do it or another option that you can actually try to run it on on fpga itself you can try to put the the RTL meaning the hard that is being designed compile that but instead of targeting to the let's say the tool chain for for taping out the the chip you target an an fpga device and then you can map it on fpga and you can try to run your whole software stack meaning that you can try to boot Linux here and you can try to run the whole um tool chain and test your applications there this of course getting this to work is a is a is a major effort it requires you know to develop the hardware itself to actually get all the like an operating system to boot there so and then of course you you might not be able to emulate the full system because in an fpga the density of Hardware let's say that you can emulate is not the same as the one when when you take out so here for example you could have maybe um a few cores without the vector unit or you can have maybe just one core with the vector unit to try to start working and developing code there later on I will also mention there are some other options you can also for example try to do risk five development in purely virtual uh environment so we I will also discuss a little bit about that there so this is a little bit what I wanted to explain about the work that is currently being done in you pilot so the next thing I wanted to talk a little bit about our involvement and sort of a bit of what projects we are working on at shmmer but maybe I'll take a break here just to ask if there's any questions on on the topic so far maybe we can get back to it at the end but what's the biggest challenge in these type of projects to actually get to the point where where we have say a European supercomputer running on only risk five so CPU and accelerator well there are many challenges so the question is what is the the major um well I I think I mean the the development of the hardware and its verification to make sure that it can run that it can reach the stability to actually run um software that is a is a major step and I think it's you know it's really like when you reach that point it's really like a breakthrough because at that point you can then start more know productively developing um code and then actually try to start to optimize it so I was thinking actually about this a bit while coming here that it's like I can see that there like this basically two important phes one is to try to get you know to your first version to get that working and then after that you basically have to start an iterative process because now you want to optimize the software and you want to optimize also or and to improve the hardware so you you start developing your software targeting this new hardware maybe using some Performance Tools and then you can try to optimize the well get the feedback from the Performance Tools to optimize your software but at the same time you provide new information to the hardware so a new version or improved vers Hardware improvements are developed so now you get to a new new version of the hardware and now there's a question about okay these optimizations which did how portable were they do you need to now reoptimize your application so at some point then we get to this point and that's why you know I I can see that the road towards getting this uh this to to work is probably still you know a few years ahead we are now sort of I think reaching this point where we we do have some some platforms I mean I'm talking a little bit about these projects now I mean there are many many approaches in fact I will I will talk about what some other vendors are proposing but what I can see in the project that we are involved is that we are getting to a a point where we do have working hardware and now I think we need to start with more actively with see how we can do sort of you know iteratively learn from attempting to optimize application seeing you know how which are the instructions that they most rely on are the where what optimization should we introduce into the hardware then try to develop next version of the hardware and then see okay now can we weart again the process of looking how well these ported applications are going to run and eventually we'll reach there and this is a step that that something that will require certainly a few steps until you reach there that's why I think in any of these projects of this scale it's not really realistic to think that you know we're going to develop the chip immediately the chip will be competitive no you have you develop the chip and then we have to start and after a few generations of the chip then we hopefully can reach to something that will be competitive that's um the way I I see it all right um okay let's talk a little bit more about our inol what we're doing in and charma so we do have two parts we're working a little bit on the on the software side and also on the hardware side so on the on the software side we have been looking at vectorizing convolutions for example we have also been looking at how to optimize part of the open andp runtime there and we have also a plan this we haven't started yet to actually work on on having an efficient optimiz implementation of CLE for for risk5 and it's interesting because there was a discussion about CLE also also before and this is sort of interesting you know because if you think about how many applications there out there written in CLE actually the number is not that large but one of the the reasons we are interested or there's interest to get CLE to to work on this is because CLE can serve as a sort of intermediate step when you are porting applications from Cuda to work on this platform so that's one of the things that we want to look at also and then on the hardware side I'm not going to talk about these things uh these two topics is we but shers is sort of responsible for what is known as the home node so the to basically ensure that the course are going to be cash coherent both you know within a single chip but also across different chips that's one of my my colleagues um are working on this part but I will not be discussing this into into detail so for in the next few slides I want to talk a littleit provide some example of activities that we are um working this is the sort of this is sort of my more like my day-to-day life so I hope that I can make it interesting and not just bore you with a lot of technical um detail here but before I mean to jump into that I want to talk about one thing which is the the risk five uh Vector extension and the reason I want to talk about this is because I think this is one of the important development that has happened in RIS RIS 5 over the past years that is an important step towards being able to have know high performance Computing in in Risk 5 so you're probably um aware of what Vector extension is because it's very common implemented I mean all architectures or major architectures have back extensions in the case of x86 you will have something you know like ABX 512 ax2 earlier you had the SS extensions in the case of arm we have the knee instructions but there's also the sve the scalable Vector extension so and all these are very important to achieve high performance because Vector execution is one thing is one way to also achieve good Energy Efficiency and why well basically you can think of it as when you have Vector instructions then you you are able to decode one instruction but the semantics of the instruction itself indicate multiple operations to be performed so you can do you decode the instruction only once you schedule it only once but then you know you can execute many operations um and there's a guarantee that those operations are independent so you don't need to do any let's say disambiguation at the hardware level which is something which is a bit costly so that's why also the approach that we're taking in eilot and Epi to try to achieve high performance risk five execution is using the vector uh extension and in fact so this resc Vector extension the there are several versions of it the 1.0 version let's say the the standard 1.0 version was ratified in towards the end of 20121 we are leveraging it in this in the in the epi new pilot uh project so I put theost ASAC and back but it refers to basically these two projects and one of the major strategies that we attempting there to in to in try to increase as much as possible the Energy Efficiency is that we are relying on having very long vectors so one thing that I haven't mentioned yet but it's actually the next bullet here is that the way that this Vector extension is built it's also similar approach taken in the arm sve is that the architecture itself doesn't dictate the length of the vector register this is like an extension like AVX 512 which says you know ABX 512 is is for 12 bits but in the case of of risk five you can have any size of vector register going from um I think it's either 64 or 128 bits up to 16,384 you can do this in increments of of two so this this number of bits what does it mean basically this means that you can have vectors of 256 double Precision elements and in fact at the programming model you can even make these vectors longer because uh risk five supports something which is called um something called Elmo which basically allows you to take multiple registers and treat them as a single uh register so the idea here is that we will be able to achieve higher Energy Efficiency by following this approach now this is enabled by having this programming style which is called Vector length agnostic so since now the implementation is free to choose the actual length of the of the registers it means that when you are for example processing a long Vector an application Level Vector in a loop um then when the compiler generates the code it will have to do some extra bookkeeping trying to understand first okay how many elements can I actually fit in the vectors that are available and then you know partition the loops in as many iterations required to actually be able to process this uh this Loop so this provides some portability but it potentially there are some scenarios in which you might lose some efficiency by by doing this as compared to what is would be called the vector length specific um code generation so the way I I like to think of it is can we basically do something that is competitive with a gpgpu by using um long vectors and a bit of the of the research that we do goes into into this direction now before going to that I just thought okay another sort of important uh consideration when you are talking about vectorization is how do you actually vectorize because this is this is a big topic and something that actually has been studied for for long long time and there are many challenges so let's say you we the two extremes that I have here on this on this slide I think the is to understand one is okay you do everything um well you just code your program in Assembly Language you have maximum control of of all the features this is of course not very um productive on the Other Extreme what you would like to have is something where the compiler does everything this is usually called autov vectorization but uh here you give all the control to the compiler but this would be the the most productive unfortunately you cannot always rely to autov vectorization First autov vectorization requires the compiler to be able to do um analysis of um let's say memory addresses and and dependencies between instructions to guarantee that for example when you're when you are um looking at a loop that is for examp example adding uh two vectors and storing it to a third Vector you would be able you would require for example to guarantee that the output Vector cannot alas with the any of the input vectors if you cannot do this because the programmer has just written some some pointers there then the compiler must say I'm sorry I cannot uh vectorize this otherwise it might be wrong code generated at the end so and there are many other scenarios and and small details that may cause a compiler to to fail in vectorization so there's a lot of work also required to try to get this technology to work effectively and sometimes we will have to actually ask the the programmer to help a little bit which is something that you can do sometimes easily by adding some simd Clauses like in open MP you can do something like pragma om simd and then you tell the compiler there is no dependencies in this Loop you can vectorize this say a compiler say okay thank you um but go but reaching this of course from if you don't have a vectorizing compiler just getting to to this point is a big effort so one of the things that that we often that uh that we do is some intermediate step which is the program using using intrinsics this is uh a way of programming which still gives a lot of uh control through the programmer but it provides some some help of make some some things um simplify some task for the programmer so I don't know have you ever programmed using intrinsics for example for x86 okay all right then then you will then you know that for example when you program with intrinsics you don't need to worry about the register allocation and you can make use of the access variable names directly so this will this will be um something that will be helpful so actually in in the Epi project a set of risk five bulletins has been developed to Target most of the risk five Vector instructions but not all of them and they have some format like you know will be called buildin EPI and the name of the instruction and with some extra parameters there but so and this has sort of been an a way to prototype to try to develop a first uh set of intrinsics to make this work and later on the risk five uh Foundation actually has set up a task group to come up with standardized um race five vector intrinsics and now there is a standard you can see actually I just put you know two instructions like vvl so set Vector length with the EPI built-ins and with the name of the standards you will see that you know format is a little bit different but it's basically but they are equivalent there's a frozen spec now but it has not yet been ratified so this is not yet standardized but there is implementation of NG GCC and llvm are are ongoing here these are sort of the four major ways that you can leverage a vector the vector extension in fact any new Isa um extension before it can be fully be uh be used by the compiler you can work at the assem level or you can try to develop some builtins for it so the one thing that I haven't mentioned here on this slide is that when you are doing assembly there is also an option to actually generate the assembly at runtime using just in time compilation and I will come back to that a little bit later but this sort of enables more um optimizations as long as the more information you have about let's say variables and runtime state of your program the more you will be able to optimize it specifically for the platform on which you are executing so that's also there's also several approaches that are looking into into being able to for example generate um or jit generate Vector instructions I haven't checking the time but you started late so it's say half an hours anyway let me just then show a little bit just to highlight a few of the things that that we are doing without going into too much detail we'll talk about three projects so this is a a project in which um one of my um PhD students when I who actually started as a research engineer in in Epi she was looking into how can we make use of the of the vector units to try to efficiently um or to make efficient convolutions and well in in particular she started early with the work in in fact at the beginning she was working on RSV and she she was working on on a few different approaches so let me actually backrack here and say that when you are when you want to run you know convolutions like convolutions that you would do for machine learning actually there are many different algorithms that you can that you can use you can do what is called a direct convolution you can apply an algorithm which is sometimes called IM am to call G where you do an A transformation of the of the input Matrix so that you can use a matrix multiplication later on you can also use um an algorithm which is called vogr and finally it's also possible to use um ffts for for convolutions now ffts are not that commonly used I think they have some I mean for for the sizes of the kernels that I usually used there they are they sometimes lead to numerical instability so most of the approaches are based on either direct I am to call or vogr and Sonia had been had developed an implementation of vogr for for rmsv and she set out to try to Port this work to to risk five to see what are the challenges of taking a code optimized for rmsv and try to get it to run on the on a risk five platform using the vector extension so she set out to do this work while trying to you know maximize usage of the vector units and of the vector registers and also because we are interested doing some code design we we looked also into some Hardware uh parameter tuning so I was thinking can I get rid of this slide but at the end I decided to keep it just to point out a bit of the tool chains that were used and that this work was actually based on using a sim a hardware simulator which is called G 5 so we did not actually run this on on the fpga platform at this point but using a software Simulator for the for the project and also using the llbm Clank tool chain which is targeting risk five which is being developed in in the Epi project okay so without going into too much details just try to provide the high level that when you are doing a vinograd convolution you you take you have to do um three things first you have to take the input Matrix and the kernel and apply some Transformations that's first step the second step is that you have to multiply this in a type of multiplication which is called the topple multiplication and after you have done this multiplication then you can convert it back so in the in the original work that we did on rmsb we figured out a way on how can we try to increase the intertype parallelism or you or basically make use of parallelism across different input channels in order to be able to fill um vectors of longer size so in this case it's just showing how you can for example fill a vector of 512 bit using four channels but I mentioned before that in the for the case of the V we have very long vectors of up to 16 K bit so this approach would need to be extended so we tried so the first thing that we did try to use 32 channels in order to fill in this case up to 4,000 U bits so you already realize that this is already sort of hinting to one to one problem of using very long vectors is that you will need to be able to find enough parallelism to be able to fill those vectors in this case we say okay we use a parallelism from different channels but how many channels can you actually have in a in your application so that was one of the approaches the other was to try to using increase the topple size or also to utilize longer vectors so we started working uh looking into how to Port this to to five and we identified several challenges and you know when you do this work the thing is you you will probably be biased when you start from one um code and then try to Port it to another one because here we started from rmsv and then of course the simplest way to try to Port it is to say okay let's try to do one to one mapping and then of course you will find out okay ah in my original code I used this instruction that was you for example a load quad word elements into a vector and replicate and then when you go to the risk five Vector spec you might find that oh this doesn't exist so you have to find some some alternative so well in this case we try to we found that that we we're lacking a similar instruction that basically the what this instruction would do if my mouse pointer is visible maybe this is not the best but is to try to take a set of uh elements here and basically replicate them a set of times over the vector register okay so we figured okay we could try to for example use index loads to basically load these elements and then keep replicating them on a vector register or alternatively use some type of instruction that is available in Ras five which is called slide up instructions so we evaluated that we've realized that okay according to the simulated data using the slide up instructions would be considerably faster than using the index Vector uh load nevertheless having a single instruction that could implement this would probably still lead to a a faster execution time another challenge we identified is that we we were for example lacking an instruction that would do a transpose of four vectors so I guess this figure on the on the right shows this you know you have the let's say the four vectors organized in this row y way in this row based wave a Subzero until a sub3 here and after the transform would like to instead have that organized you know in this column uh way over a set of vector registers so that's such an instruction exists in in for rmsv but we don't have any equivalent for example on the case of of risk five interestingly um in the case of Epi there is actually a custom extension which is not part of the vector standard but that targets only two vectors so that was not something that we could uh that was useful for us so we tried several different Alternatives like doing a a unit stri at store Then followed by an index load or doing a strided store followed by unit strided load which more or less um perform the same the at least these two because as you can see they're basically doing the same operation but just you know following a little bit different uh logic but any case so we identified an option here to maybe come up with a similar instruct to further speed up this uh process and the final challenge that we faced that was both related to with performance and also programmability aspects is that it turns out that in in the case of RSV it is possible to declare a vector type at the the sea level and then take a reference for it and this this allows to well this was um allowing us to solve a problem that we had here in which we were trying to let me see did I describe it actually on the slide okay so the idea is that in the in the trans in the Transformations there is a there are some operations that need to be done over vectors and that they need to happen on all every time that you are calling these um these three Transformations the input transformation the kernel transformation and the output transformation but the registers the vector registers on we on which we want to operate they keep they're not always uh the same so there's a challenge that okay so if you want to for example create a procedure to do this transformation of the of the vector registers then you would have to basically create use a set of intermediate Vector registers on which that would be used to pass the parameters to this function and then return the parameters the vector parameters back and that would of course well that first of all would increase the register spilling and furthermore this actually reduce programmability because we in the case of risk five such a passing of references is not possible so we would have to actually write the code which was about 30 lines of you know of code in six different places in the program so leading to decreased programmability so that was I mean that's the what we ended up doing and but we're trying to look into ways in which such a possibility to try to you know pass register references to procedures if that would be possible in the case of risk five it would certainly help very much for this type of uh scenario so overall without going to too much detail we found that compared to the traditional approach that uses IM am to call theog implementation that we did on risk five was was about 20% faster and on the on the gem five it would actually achieve comparable performance to our optimized rmsv implementation so that was sort of a good um outcome and then also because we wanted to look into what is the actual efficiency of this and scalability you know more as a in terms of code design to try to understand how useful is it actually to have longer vectors or how useful is it to have longer caches so we did this uh study in which we tuned the parameters of the L2 cach for example from 1 to to 56 megabytes and also the length of the of the vector Reg registers from 512 to 4K and what we learned by doing this is that in terms of vector length there was not much point in going Beyond 2K uh vectors and also the performance the impact of increasing the second the last level cach in this case was the second uh level cach sort of saturated after 64 megabytes but combining these two uh factors together we predict that using an implementation that would have 2K vectors with 64 megabyte last cash would achieve about 1.8 performance Improvement yes but this is for a Tiny Network yeah for a vgg which which is minimal if if you have a something bigger did you try something bigger than that the other thing we tried to is Yolo which is also tiny yeah no so we didn't go for for for long larger one of one of the reasons is actually a bit related to the practicality of the approach when you when you're using a tool such as gem 5 it's actually extremely slow and if we go for a very long Vector that was not Vector so very long large model that would not finish in fact for YOLO V3 we didn't in fact simulate the whole thing because it would simply not finish in a reasonable amount of time but [Music] um I mean I'm I'm actually interested in your suggestions for other types of of network are you talking about CNN related or or non convolution I I would say like if you go for a Transformer it would be much bigger yeah so those numbers will will still keep contining going down yeah that's that's the thing no I I agree this is yeah I mean we are very interested in looking into to Transformers for a moment I thought you meant larger cnns but no if we talk about Transformers that's a it's a that's actually I would say that it's sort of future work okay yes and but okay so this work was done actually in a on a library which is called darket which is not really that commonly used anymore so one of the things that we are also working on right now in fact I have a master thesis going on right now is to try to put these algorithms like the vogr the IM to call or also an implementation of a direct convolution inside of 1 DNN so 1 DNN is a much more wellknown or much more commonly used library for doing um machine learning operations and it's um open source it's uh actually was originally developed by by Intel so it's part of this one API stuff but there's also ports for many other architectur so you can use it also arm power and also there's uh risk risk five that we are actually developing inside of the pilot project as well I will not show you any any performance results but I want to show a little bit about what we are currently doing because it's uh with one DNN we actually are making use of this sort of jit approach that I was describing so it's also and it's also sort of an important tool to support in this ecosystem in fact the work itself on 1 DNN wasn't started by us it was started by by BSC um it's mainly um marasas and one of his students that has been working on this and they come up with this they implemented a jit approach in 1 DNN basically you if you want to for example um you know have generate an instruct which in this case is a is a vector additional instruction you can from the high level you will just uh call a function which is okay push this instruction and then later on you can generate the code and and execute it so using jit has its pros and cons I would say two problems compared to using intrinsics for example a problem is that you need to keep track of the registers manually because here you are really actually you know specifying the exact registers that are going to be used you cannot use high level names that the compiler can then transform for you and basically it's the same as creating an assembly version of your code the potential benefit well in this case is that this Approach at least makes it easy to extend to add instructions to the jit and why one would like to use jit itself is as I mentioned before because you can um it enables certain types of of optimizations one sort of simple optimization you can think of is that if you have you know at some point you want to for example call A A convolution for example let's say you want to execute a matrix multiplication and uh you know EX act if you when you write your code for your matx multiplication you will have okay you have for example so many rows and we will have so many columns and then we can come up with some blocking but you know you have to sort of support a range of sizes but if you do it on a jit Approach at that time moment you might know that ah the Matrix I have to operate is exactly 128 rows and 64 so I can then propagate this information and simplify the control flow inside of of the of the routine to create a specialized version to be executed at that point so there are instances in which having some you know some of these information which you know only while you are executing will allow you to generate more more efficient code so that's why we are um interested um into this and I have here a couple of slides that just show a little bit how this um approach um works but I think I'll just go quickly I'll just sort of skip it maybe I'll just stop here just you know because it's kind of maybe interesting since I've been talking about the the intrinsics that have been developed in the Epi project so you can see here an example of how the in code um for intrinsics uh with intrinsics looks like and how the code that is then generated with the with the jit um ends up looking I actually know exactly for which function um this is but um so you can see that in the case of intrinsics you can still you know make use of um a call make use of variable names and do for example arithmetic on on pointers and then pass that to a to the intrinsics will will then generate the corresponding load instruction but this sort of arithmetic that that you can see for example in this first instruction is something that you would not be able to do in the S side of the jit here you have to know have to pre computed exactly the value in order to be able to generate that and all right I'm switching a little bit to a different um project that was also work that we have been looking at was to optimize the openmp for execution on the on the risk five in fact this was more started as a more generic we started also looking into arsb at the beginning and also into the Intel simd but with a long-term goal to support risk five vectors U and so well here the problem where we focus specifically on synchron constructs such as barriers and and reductions in open andp and as you know the problem is that well as you if you look for example on this graph here and where on the xaxis we have number of threads this is actually run on a Intel K&L and we look into sort of the overhead that results just from this operation you will see that the more threats you add the performance uh degrades so question was can we you make use of uh of vectors can we utilize Vector units to reduce this overhead and by the way this this Intel KL was located at the hpc2 so this is links it's the one that you had so we got access through it via Snick at some point anyway so well you know this is a supposed to be an example of of a barrier a barrier is a construct in which you have multi mple threats that have to sort of wait all at the same uh place and wait for each other and only when all the threats have arrived then they are allowed to proceed so one way to one attempt to vectorize this is shown here on the on the right side let's assume all the threats when they reach um the barrier they will activate a bit which is a part of of a vector and then the primary threat can try to load this Vector using a vector load and then try to see if all the threats have have reached so very um simple um idea similar can we also applied for reductions there it becomes a bit more complex because we have to apply the reduction and we also had to modify clang but for the barriers we did an implementation that was only inside of the llvm openmp runtime and we created three different versions one for INT AIX one for rmsv and one for for the risk five Vector extension now the results that I that I have are only shown for for the KL machine and a64 FX because at that time we didn't have access to test still on the on the hardware for the for the risk 5 but um we have but what we did do for risk 5 is to validate that because we used at that point we used an approach using using this uh vave tool which is the emulator for the vector instruction which I mentioned before and that was coupled with Cho so Kemo is um well I will talk about it a bit later but K is basically an emulator for a system level emulator okay so we did this result and at least for this for the KL platform we observed that we could reach uh already speed up just you know just launching barriers at this point of up to a little bit over 2x so we haven't been able to test this yet on the on the Up pilot prototype but that's sort of one of the next steps and based on that we'll see if we can provide some feedback back to the hardware team on the performance of the of the barriers okay I finish here with this this a little bit of things of more research like that we have done in terms of risk five and I will try to finish a little bit asking more a question or ask or sharing some some thoughts about how I sort of see the the road ahead for for risk five in in terms of HBC these are of course my my own thoughts and they are probably biased and and faulty in extent a lot of things are happening in in parallel in fact in this field it's difficult to keep track of of everything but what I can tell you is that well as you have seen there's a lot of progress happening actually in at all the levels there's work going on from the hardware side on the instruction set architecture and also on the software libraries and Tool chains but I think there there are challenges at all these levels that still need to be need to be face so I'm going to I want to talk a little bit about these things you say going from Hardware Isa to software a few other thoughts and then even share some thoughts I me what you could do if you would if you're interested for example to look a bit into um risk five or you in a in in a general sense okay so there are if you want to let's say start with risk five Hardware there's actually a lot of um small boards that Implement risk five but if you are looking specifically for HPC it's a bit more more more difficult so I try to I mean based on my limited view of it of course I haven't read all the papers and everything that is happening all the time but i' I'm only aware of let's say one um board at least that I have seen having been evaluated independently and that's this um well this softphone SG 2042 which consists of a set of cores from a company called Tad and this one is implementing the vector spec 0.7.1 so it's not implementing the one 1.0 and I I sort of highlight this independently evaluated here because I see that there's a lot of I mean there's a lot of projects and even I can see pictures of people saying okay we are developing this chip and here's a picture of the chip but I have actually not seen anyone use that yet or having it evaluated but here there's a a paper in fact this is by um I think it's by Nick Brown um He was discussing about this in the risk five Workshop of supercomputing last year where they look into this this um system one of the challenges is that this does not implement the last version of the vector spec so it actually requires a custom GCC compiler so from terms of usability particularly for that Vector instruction is actually not um not very convenient but supposedly it looks like there is um there's light at the end of the T I would say there's a huge amount of companies and many of these companies have a have a large amount of of funding and and backing that have announced or have prob publicly disclosed that they are working on high performance um hardware and and often they there will there are systems specifically targeting HPC and also systems targeting AI so here this is just a a list of the of the ones that probably come to my mind when I when I prepared this slide but you know companies such as T torent they are developing a large outof order risk five cores plus an AI chip which is called a 106 there's Ventana micro which is also developing a out out of order processor with you know with Vector extensions espiranto Technologies they have these two types of of course they have some very small risk five cores that you call the the minion and I think that's that's mostly for for AI type of workloads and they're also developing a host processor the ET maxion semidynamic is company which I which is part of Up pilot they have developed for example out of order code which called ATO this has sort of of a interface for a for a vector unit and I think that they have their own implementation of a vector unit outside of the of the vector unit that I was uh discussing in the EU pilot project and they they have also in fact a set of custom tensor instructions that they are developing for for AI workloads Inspire semi has a has an accelerator platform which they call Thunderbird and you know there are probably more these are I think these are all companies I mean most of them are us companies sem dynamics of course located in in Europe but I'm sure that there are companies many other places in the world that are currently developing Hardware so at some point we are going to get to the point that we will have more availability but right now I am personally not aware of any of these systems that can be accessed I might be wrong but uh I haven't I haven't come across the that yet the good thing is that all of them from what I could understand support the risk five Vector 1.0 spec and that's good because once these systems become available it should offer interoperability you know you should be able to run code on one system it should also be able to run on the other systems and that's of course you know very important for for being able to you know not only deploy software but also be able to compare systems but later we'll see you know it will be very interesting to see once these systems become available if it is possible to start running benchmarks on them to actually see how they compare in terms of uh performance now what's going on in terms of specification so I've talked a lot about the vector spec and I think that's certainly a big step um ahead for for being able to support HPC in in in the context of risk five now the there is still a lot of discussion going on in the in the what's it called the special interest group for and the vector s in in Risk 5 luckily um you know everything in in Risk 5 is kind of open so everybody can just go in into the archives for the maing list and see what people are discussing what's going on what sort of discussions for new versions of the spec are are happening what you cannot do unless you are a member is to join the the meetings themselves and try to to uh participate for that you need to first become member of the risk five International which you can do as an individual or as a organization depending on whether your organization is paying your work on risk five or not but anyway because of that I I have been you know been able to attend some of these meetings and also look to back blots of of mailing lists to sort of try to understand what is going being discussed and here are some things that I've seen that are sort of topics of of Interest inside of the vector c one is um that there's lot of interest in subsetting the the current uh Vector specification what this means is that you know we now have a vector spec but this Vector spec is actually very large it's it's actually well the last time I I checked it it has around maybe 200 instructions and it's uh implementing it in its totality is quite complex in particularly for embedded devices you might not be interested in implementing the whole thing so there is discussion for example in separating the vector specification into smaller sets that probably hasn't too much impact on the HPC part itself but what could have an impact is is the next thing so self-contained Vector instruction so risk five Vector instructions are currently 32 bits and that's not a lot of space to to specify you know many of the operations or that are required for these vectors particularly for this Vector length agnostic design so the the vector spec currently is based on having a lot of control registers that support and that impact the execution of the instructions but there is you know there's a discussion on having a new version of the vector spec that would have longer Vector instructions maybe 64 bits where all this information is encoded as part of the instruction and that could for example enable having more architectural registers right now we have 32 Vector registers maybe in the fut a future version of the spec would have 128 registers and then you know that will have a big impact on on many of the of the applications for example the vinograd implementation that I discussed we were limited we had problem of register spilling if we had 64 or 128 registers that probably goes away then other things they're looking is some sparity support and also trying to have more GPU like capabilities in the in the spec um itself other things that are I think very important particularly for AI is that there's a lot of interest in Matrix extensions and before we had a bit of this discussion about the customization whether you know you need to give back any extensions that you make to the to the spec or or not well this is a field in which you will see a lot of custom extensions because right now we do not have any standard extension inside of risk five there are two working groups one for what is called for the integrated Matrix extension which I I I think what this means is that they it's a matrix extension that makes use of the Vector registers for storage of the matrices and then attach Matrix extension which would then be more like a co-processor type of Matrix extension with externally stored mates on separate registers so many many companies are currently have developed their own custom uh matx extensions for these types of workloads but the natural you know flow should be such that at the end of the road this um extension somehow convert into an actual specification and another thing that I also highlighted here I mean I think there are more things in the risk five spec that impact HPC but also I want I think I mean personally I've always had a big interest and I think it's very important the topic of performance monitoring support so there is I mean risk five does have a spec for um for uh performance monitoring but one thing that is going on now is that right now the specification for performance monitoring just uh specifies how the how to access the counters for example but it doesn't specify which counters need to be available so right now there is a process to also have a specification of a set of names for standard performance counters and I think that's also going to be very helpful because it will guarantee that we can you know run papy or whatever or perf on all these platforms and and hope to get the same counters to so that we can use the same Performance Tuning methodologies on all these platforms so that's Al hopefully something that is going to be there also in the future I have only a couple more more slides I wanted to mention something about software tool chains you know obviously for the success of of risk 5 we will have to Port a lot of you know software we have seen how how how long it took for in the case for arm to become competitive on the HBC side same sort of effort needs to be done in the case in the case of uh risk five I have dis discussed some of the work that we have done in epi new pilot and here's just you know I try to go quickly over that slide that I showed earlier to highlight some of the libraries that we are actively uh working on and some the tool chains but there is more work um I mean there's a big interest on this topic of course and um in fact there is there is a industry Le effort which is the rice project which is the risk five software ecosystem and you know they are looking about what has to be done in all these uh levels to try to accelerate the development of Open Source software for risk five so looking you know as compilers GCC and lvm system libraries SSL GPC blast kernel most of focusing of course on on Linux and KVM managed uh run times Linux distributions debug and profiling support simulators and and system firmware so lot of things happening in in in that front um as well so some ideas more from our own experience working on risk five is that you know while you might think okay you know it's just a new architecture let's just recompile and we are done this is of course far far from it one challenge I mentioned this before is that a lot of codes that we want to be able to run on risk five uh in the future on risk 5 accelerators are currently being developed for example only in a in a language such such as Cuda so there's a question how can we make Cuda programs be use um r five do we need to rewrite basically the the application to use um more like the vector way of programming with intrinsics or trying to rely on open p simd or maybe an alternative approach is to use automatic conversion from Cuda to something like CLE there are some tools that and that enable that and then try to have a high performance implementation of CLE and there are there some work going on there so in the end yeah just point out that risk 5 is not a it's not a GPU so those codes will still require some effort to be done and in the context of uh you pilot one thing that I I see is that you know we we are working on trying to get to do long vectors but um the way I see it being able to to fill long vectors is is very much algorithm dependent and a lot of code that for example has been developed with Intel like simd like x86 in mind will not easily translate to very long vectors because there you have more shorter vectors up to 5 12 bits and when you know that and you program specifically for for that you will already organize your algorithms targeting that you will not write your algorithm to be able to extract vectors of thousands of bits so we are we will need to do some research into either more advanced compiler support to be able to fill these longer vectors or at the end the developers themselves will need to actually do some um algorithmic transformations to to expose the required parallelism to fill these long vectors and yes I guess I'm I'm almost done one thing I also thought that maybe could could be interested if you have never come across the risk five but you're a little bit curious just to point out that there's a large amount of boards these are mostly you know small boards not HPC like boards that you can play with you can go to this uh this link and there's a long collection also if you a bit more interested in the HPC site or the work that we're doing in Epi pilot I want to mention that you know these sdvs that I mentioned before with the fpga and Unleashed boards that's actually open for external access I think if you sent a email to philipo mantoani you can request access to that and you can also play with that tool chain there if if you don't want or you don't have Hardware or you don't want to have Hardware there are of course also Alternatives spike is sort of the golden reference for the R 5 East as an emulator and and if you want to Le something that's a little bit more complete Kimu is actually part I mean there's a lot of good support for race five in kemu it's part of also of the goals of the rice project that K is well supported and I think there's a commitment that all ratified instructions need to be in in kemu so you can actually get quite a lot of done by using just K to start playing with risk five and finally this is maybe more interest of me because we do you know a bit more computer architecture and performance evaluation if you are in that field then it's also good news that gem 5 since the last release in the in December 23 released 23.1 now also includes a model for the the supports the vector 1.0 standard okay so that's I thought that would be all but now of course I have my a conclusion slide and I have an acknowledgement slide too so well we have seen a lot of progress we have we are working in Epi pilot working on a on a on a multicore system with a vector support and a software tool chain and there's also a lot of work at the commun as I said I think that the vector is a big step forward for being able to have um high performance Computing based on on risk 5 but I think we'll need to have more boards more HPC like boards to play with and Performance Tools so that we can you know start optimizing the code you know using more serious uh platforms yeah we hopefully the HBC Hardware will also so converge around a set of more stable specs and I didn't talk about something called the risk five profiles but that's actually sort of the way risk five is trying to get vendors to F that focus on the same Market to use the same set of of vector extensions so not Vector extension of risk five extensions So to avoid fragmentation and well I've talked about some specific challenges in this talk that I particularly think are interesting at least for us that are looking at is autov vectorization for extracting long vectors or more systematic ways to rework algorithms that will allow to extract longer vectors also conversion strategies to move from Cuda to be able to execute Cuda programs on on risk five and finally I think it's going to be very important that the working groups that are defining the Matrix extensions try to come up with uh some standards um rather soon okay so now just I wanted to acknowledge my that of course all this work has been done by my team our we call sort of informally we call us the chart team chmer heterogeneous architectur run times team and yeah and I you always try to make use a plug here to just point out that we have also very nice master programm in high performance Computing systems and then if anyone is interested then please come okay I have absolutely no idea how long it took but I we're done we started late but you're pretty much on time if all right okay well thanks thank you for enduring me until the end take some more questions maybe if there are any I'll be happy I'll be around in any case until tomorrow evening so I think there's a question well so you said about the difference between an open specification and open source Hardware I I didn't get exactly what is the difference in this case open specification just refers to the instruction set architectur to the language that the hardware has to talk so you can you can take the specification you can customize it change and add new instructions if you want and nobody's going to stop that in that sense it is open but Open Source Hardware would mean that for example you publish the RTL the verog or HDL like we are publishing the C code or C++ right so that's sort of the yeah the difference you talked about emulation a bit as well through keu for example when you're playing around with risk 5 since there's to some extent a lack of Hardware to actually run stuff from and in a previous risk five talk I saw is that it also depends where you're running q u in terms of getting a better emulation to some sense or at least better closer to the hardware like running qemu on an arm system makes more sense than running it on an x86 system because the memory model of arm is very similar or is at least a lot closer to risk five than it is for for C6 I mean what you say makes sense I personally don't have that experience I don't have the opposite experience I simply haven't tried that so I can't really well what he was warning about is that if you can pick run on arm because then it's less likely that when you're just transferring your binaries let's say to an actual risk five platform that you're then running to surprises there which you which which you're more likely to overlook if you're running Q on x86 so there's a there's a factor there as well at least no so I I think I understand what what this means I mean in x86 you have a stronger memory model that's that's usually called TSO total store order um and in Risk five and arm you have a weekly modeled uh weekly ordered memory model and if you have some specific algorithms that happen to work because of the TSO guarantee it might be that running those algorithms on K on x86 work on the risk five emulation but actually don't work in practice that yeah that could happen I hav okay but um it's an interesting question itself I mean how can the how can k itself remove that those guarantees but it's probably maybe too too much to ask I don't know yeah it it would only lead to further slowdown which is probably not what you want you want to use the underlying Hardware as much as you can and and I think it's just an important detail that people should be aware of like if you can just run Q on on Q on arm particularly if you are developing let's say like lock free data structures or stuff like that that would be very important then I think okay and another thing is that had a slide with lots of open-source projects that are uh starting to look into risk 5 or are being ported to risk 5 do you mean the the one on on Rice yeah the one like yeah yeah indeed where like BL open blast and all these things are are mentioned M I'm not sure if you have any experience with that but what's the situation like when a project like rise gets in touch with these projects and says we want to add support for risk 5 like do they actually know what risk 5 is I mean for GCC that's a given they are well my understand that well I mean the people in rice themselves are doing are doing the work um but then is are the projects also accepting their contributions do they understand why it's interesting to do that I mean that's probably a question that I cannot answer but I I would say I would probably say why not I mean do you think why do you feel that there are any specific concerns about it the reason I asked is we a couple of years ago we had an online talk by the the Bliss maintainers so Bliss is like it's an alternative to open blast let's say a blast leack Library yeah actually I think we have it yeah bliss bliss is there now but when I don't know how long ago this was but I think it was five years ago when they yeah they gave a really in-depth talk on Bliss and everything they do um to make sure you're getting good performance Cals and so on and I asked them the question about risk 5 and they had no idea what risk 5 was now that's that's five years ago things are very different today the people from from julick right developing it or no the um where are they in Tennessee Tennessee is it or no or Texas no Texas University of Texas I think yeah that surprised me a bit because if Bliss doesn't know about risk 5 they're so low level they're basically hand coding assembly to some extent well BL not really they are trying to avoid that but if if they didn't know about it five years ago then I wouldn't I'm pretty sure that many let's say scientific applications currently which are three four levels higher still have no idea what risk five is and some of them at least should like if you're touching assembly anywhere you better know that this is coming because it's going to take some effort to make sure you're you're being ported and this is just an observation and it's not a criticism or or anything but that's I I think it's still a challenge it's improving but it's it's a challenge to explain to people that this is coming and why they should be aware and maybe try to prepare for it yeah no I I mean I I don't disagree I I don't really have I mean I would like to know myself how much awareness there is because I live in this risk five bubble and sort of something think everybody knows about this but then you go out I mean but here for example I saw that most people are aware of it but um if you go out some completely different Community it's a good question but um yeah I mean like um in for example in this European projects or the way it it works it's a little bit like okay we Define a project we try to select which are the key applications that we think are important and then we sort of reach out to the responsible and try to explain why we think it's important y I don't know if um this works similarly in other communities but at least this is one way in which we try to make this awareness it's right now it's coming it's sort of the in the responsibility of the risk five Community to try to convince that people and of course the more projects know about it it's going to snowball into into bigger aware of course at some point that hopefully should be the case yes you didn't mention any any timelines but do you dare making any predictions on when we'll have let's say an the top 500 or yeah let's say that yeah that's that's an easy okay recording well but by then it may need to be ex scale to be on top 500 it's very diff it's a bit of antic question is it um no I could try I mean um so my I think my main challenge to try to answer this is that question is that let's say I don't really know how advanced all these players are I know that that they have you know huge teams working on it so they probably have the capacity to put the system there in a in a in a few years but uh you know if I don't see the system and see some numbers from somebody testing it I I it's very hard to understand what is the maturity really because you know showing a chip ah here's the chip but you know what does that mean um but uh I it could be you know optimistically maybe two to three years if if these people are really where they claim they are or more pessimistically I would say me five six years that that's getting the actual chip that doesn't mean you have a full well a system but yeah but if but I'm I'm saying getting getting no but at that point I would say that the the chip should be having everything in place let's say so you can if you want to you could build a big system yeah I mean some of these players seem to be close to that at least from what they say publicly I mean t stor is not it's a startup in some sense but there's people behind this like Jim Keller used to work at at Intel pretty much everywhere so he knows his stuff quite well uh that doesn't mean many many of these have very wellknown people backing them also asiro Technologies um Dave dsel I think behind I mean and um and I don't know which ones but they have huge teams in the order of hundreds of people uh working on on that and with a lot of funding also most okay why and I mean is one of the things that I pointed out here is that many of these they have you know some chips that are specifically designed for AI so I think one of the reason that they have big funding is because they there's people who hope that they will be able to get some part of the cake that Nvidia currently has but if that but maybe it's if you know if they make good use of their money they they should be able to reach there in a reasonable period of time MH let's see maybe we can invite you again in three years from now and then can see that the updated thanks I think one way forward is to build a cluster specifically for AI because this is a bus word you get as you quite rightly said money for that and demonstrate risk five is actually a decent and you have to take it Zs competitor because otherwise you've got this chicken and egg problem we are interested in but it is new technology and we don't have the funding to do it and because it is new technology we are not interested in buying it because it hasn't been proven it is actually good um think about the Lumi project for example where they deliberately used gpus from AMD they had quite a lot of problems to actually Port all of the software into it and I'm pretty sure you will have similar issues with the risk five if if if you see what I mean I try to be positive M just to be clear I mean many of these are sort of trying to do what what uh you suggest to try to build a demonstrator for for AI but because it's like that I mean if you don't have the hardware or can show that it works it's you're not going to you can't convince anyone to to buy into it yeah and we are sort of trying to do the same thing but for the HPC the HPC does unfortunately not have the same budgets that the AI has so that's why Euro HPC puts more of the public money into it because that's otherwise you can't go there yeah I think we can wrap up here is going to be around for more questions today and tomorrow so I guess we can grab a quick coffee and then just jump straight ahead thing