Submind YouTube summaries
Thumbnail for MLFPM Symposium 2022: Bastian Rieck

MLFPM Symposium 2022: Bastian Rieck

Watch on YouTube

Video summary

Bastian Rieck from the Helmholtz Center Munich delivered a keynote on leveraging topological data analysis (TDA) to understand shapes that are deformed or perturbed rather than perfect, moving beyond traditional geometric constraints. He introduced fundamental algebraic topology concepts such as Euler's Seven Bridges problem and the Euler characteristic, explaining how these invariants classify spaces up to homeomorphism by allowing stretching and bending without tearing. The core of his talk focused on computational topology through persistent homology, a method that approximates point clouds at varying scales using Vietoris-Rips complexes to track the birth and death of topological features like connected components, cycles, and voids across a filtration. This process generates persistence diagrams that effectively capture an object's complexity while offering two distinct advantages: stability against small perturbations and subsampling, which links geometry and topology via bottleneck distance, and seamless integration with machine learning where these features serve as inductive biases to complement existing algorithms and potentially escape limitations like the VC dimension hierarchy. To address the complexities of shape reconstruction as an inverse problem involving 2D-to-3D conversion, Rieck discussed a specialized autoencoder called "shaper" that maps voxels in a grid using convolutional neural networks trained with geometry-based loss functions like Dice loss and binary cross-entropy. However, he noted that these standard losses are insufficient because they fail to account for shape variations when topological characteristics remain unchanged despite geometric modifications, such as rotating an object. To overcome this limitation, his recent work introduces a combined approach using both topology-based and geometry-based losses, incorporating Wasserstein distances between persistence diagrams and a total topological variation term specifically applied to predicted likelihood functions. This joint strategy successfully reduces undesirable surface wiggles that do not alter overall topology while maintaining geometric alignment, demonstrating substantial error reduction across various metrics and proving scalability by showing that topological features remain expressive even when calculated on downsampled data without performance degradation. The practical applications of these methods were illustrated through fMRI analysis and cell morphology prediction, where cubical persistence was used to characterize time-varying brain activation data while handling subject variability and geometric alignment issues with high success in predicting participant ages from movie-watching stimuli. In the realm of biology, the technique reconstructs 3D cell shapes directly from 2D microscopy images rather than relying on single-cell RNA sequencing, enabling robust detection of pathologies such as abnormalities in red blood cells even when working with small biological samples. During the Q&A session, Rieck clarified that while collaborators utilize confocal microscopy for high-quality 3D imaging, this algorithm aims to provide a faster alternative achieving approximately 80% accuracy by flagging anomalies from 2D slides alone, addressing concerns about dimensionality and differentiability within standard frameworks like PyTorch. Although current implementations do not yet integrate with structural bioinformatics tasks like AlphaFold due to the need for joint optimization, Rieck confirmed interest in such developments while explaining that mappings are constant or Lipschitz in specific neighborhoods, allowing gradients to be computed almost everywhere despite degenerate cases.
Read the full video transcript
that our last keynote speaker had fallen sick so we were left without the final keynote speaker but luckily Munich is full of great researchers and luckily we know a few and we could um we could convince Bastian Rick from the helmhold center in Munich to come and to deliver the final keynote um of of our Symposium as a replacement so so bastian's background is in mathematics he's one of the experts in topological data analysis in machine learning he did his PhD in in Heidelberg and then postdoc postdoctoral time in Kansas Sultan and later on in ADH Circ in my lab in fact and then he became a pi at the helmhold center here in Munich recently he's a rising star at this intersection of topological data analysis and machine learning he's also the program chair of this new learning on graphs conference log that's happening for the first time this year so um we now move from the more medical perspective of process medicine from the more medical perspective through more mathematical perspective on what you can do in the Life Sciences and in in the in in data science in the Life Sciences thank you very much for coming on short notice Bastian we are really happy to have you here and we're looking forward to your talk thank you very much for for having me it's the microphone already live I have to think it's not maybe I need to do something um or is it just black either is it already line ah yes now I know this is so much better thank you very much for your for your really kind words um it's also a pleasure to be here on on short notice so I I coupled this together for a general audience as I was told so I tried to take everyone with me and I hope that we can have a nice q a session as well so the title of of this is more like a framework I would say it's called a good scale it's hard to find shape analysis using topology and first let's maybe give you a small induction of what I like to do this is one of my favorite pictures it's not drawn by myself it's drawn by an AI it's supposed to represent this idea of what you can do with topology and how shapes are kind of melting and maybe it has a certain serial characteristic which is which is what I'm going for because in practice we almost never are dealing with the with the right shape but we're dealing with deformations we're dealing with deformed perturbed variants of such a shape now let me talk a little bit about algebraic topology my special subfield back in the days when the dinosaurs roam the Earth this is what I studied in mathematics it's a fear that is supposedly about counting and calculating stuff and when you look up a definition of algebraic topology you will find a bunch of them of course so a very prosaic one is we want to develop invariants that classify topological spaces up to homoromorphism there's a bunch of weird words already in there we don't have to use them another one is we can use tools from algebra to study topological spaces whatever these topological spaces are but what I like to think about in the terms of topology is we want to understand shapes through calculations and I'm stressing the calculations part because we humans have very good visual system and probably most of you know this better than than I do I mean mine as you can see is also not working properly most of the time so we're really really good at recognizing shapes really really good at seeing each other and the question is now how can we bring this into the computer world how can we somehow leverage things that our cortex can do on its own now a first taste and this is probably something that you will encounter lots of times when you look at my stuff or things that deal with topology in general that's the Seven Bridges of koenig's back problem the first I would say first historical occurrence of topological data analysis if you want to call it that it dates back to Euler because of course it does everything dates back to Euler if you look long enough in mathematics either that or Gauss so I think I said this in another talk before but if you ever do a pub quiz and you're asked about mathematics then Euler and Gauss are your sure guesses I would say now in the Seven Bridges I couldn't expect the question is can you do a walk through the city that crosses every bridge of koenigsberg exactly once and if you look at the city there's a map so there's some geometry there and it's kind of hard to do that and you could probably walk through clinics back all day long it's a nice city I've heard so you could do that or you could abstract this and you could build a graph that has the the bridges as its nodes and then it tries to build a connectivity from from up there and when you do this you will find that this is actually one of the first theorems that you often learn in graph Theory and undergraduate graph Theory no such War can actually exist because there are more than two vertices with odd degree and this is so fascinating to me because this is a geometric problem or well let's maybe not stretch real world problem too much here but for me this is a real world problem and you've abstracted it just by virtue of putting it into a graph and then ascertaining some of the properties of that graph and that has almost a magical character to it to me and this is one of the aspects of topology and topological data analysis that's of course not the only thing we can do so I have said this before we'll stay with Euler for a while because he's he's the man so there's also other invariants that we can calculate of spaces and invariant is something that remains fixed while we transform the space where we subject it to certain Transformations the Transformations can be homomorphism so stretching something or bending it without carrying it or they can be something different if you are dealing with graph machine learning for instance you have probably encountered the term permutation invariant or mutation Equity variant sometime and this is exactly one of those invariants that people are now looking for that they're interested in because if I give you a graph or a machine learning algorithm then probably the output should not change if you just change the ordering of the vertices it should change however if you start rewiring the graph and coming back to this Euler characteristic here the Euler characteristic is a very simple way of defining a polyhedron so a shape that we can nicely draw it's defined as the number of vertices V minus the number of edges e plus the number of faces F respectively and you can calculate this and there's a nice theorem of course attached to this because this is how math works and you can show that the Euler characteristic of every platonic solid is exactly two interestingly this theorem also characterizes platonic solid so if you if you have this theorem you can characterize the solids or you can do it the other way around because you can show that any one of those shapes must by necessity be a platonic solid if it has this Euler characteristic here and satisfy certain other properties now some of you might recognize those I took them I think from a very nice book by Kepler the symbols they're not meant to to represent anything here I think this was some alchemical part here which which we're not doing here because that's that's of course not science but I think it looks nicely drawn so we can check this of course let's briefly walk through this just so you can see that I'm not lying to you or at least I'm not lying on purpose to you I might be lying because I'm just plain wrong in some of the aspects here because that's also something that we have in science but it's not it's not a misinformation and purpose here so for the tetrahedron we can do four minus six plus four that's uh two we can do the same for hexahedron for the octahedron for the dodecahedron and for the iconsahedron and there we have it and by the way I think I switched uh two of these places here that was just for someone to check but yeah I guess uh I I guess I could have done it on purpose as well now let's walk um through some more invariants here and see what we can what we can do in higher Dimensions because well I mean platonic solids are all well and good but I mean I don't know about you but the last time I encountered a platonic solid was when I was playing Dungeons and Dragons about 10 years ago and so this is not really what we're dealing with in science anymore but we can go higher of course there's another nice invariant that is called the Betty number it's um named after Enrico Betty and the this D dating number counts the number of d-dimensional holes in a space whatever that might be we'll see some some examples of this later on and primarily it can be used to distinguish between spaces because it is what we call a homeomorphism invariant so if you bend your space if you stretch it if you tear it a little bit it will not change so in that sense it is a characteristic property of a space the Betty numbers have very nice properties in that they represent I would say intuitively capturable objects or captural characteristics for Dimension zero for instance these are the connected components for Dimension One these are the cycles in the data set or in a graph and four dimension two these are the voids so for instance when you think about proteins or molecules or something like this and you you thicken them a little bit then in D equals two you will find the pockets that they enclose so I've been told that this is something that people are being interested in I myself have not been working with such data so this is just uh this is just an educated guess on my part let's look at some examples here the ordinary point so nothing without any extents has the BT number of 1 and 0 and 0 respectively so just one connected component nothing else going on we can do a little bit more with the cube a cube which I assume to be a thing that contains some space has a bit number of one in dimension two because you it closes one void and we can do the same for the sphere and for the Toros as well I'm not going to go into the details why the Taurus has exactly two cycles here that's actually a very interesting and and deep theorem in topology as well but if you want to go for an intuitive explanation I would say there's one cycle that you can immediately see because it's the one that you can put your finger through so if you eat a donut like this you can eat it around your finger and the other cycle you can do by thinking of hanging it up on a Christmas tree as an ornament that's the second cycle you can find now at this point since we're now working towards computation topology let me say a few words about why these invariants are useful in general so not the Betty number specifically but why it's really useful to have invariants when you define an invariant when you look for characteristic properties you're always in the struggle between having something that is extremely precise so you want something that tells your data sets apart that has very high expressive power so for instance for those of you that are into the craft machine learning domain a little bit you might have heard about device from the limit hierarchy or the vice family graph kernel or the vice element test for subgraph isomorphism and these are exactly things that you are that you're looking for so you want something that is very expressive and that tells your data sets apart on the other hand and that's that's kind of the this uh the balancing part of this coin you also want something that you can compute effectively and efficiently it's not useful if you have an invariant that is NP hard to compute the way you don't have an efficient algorithm because then it doesn't scale and you can't use it in practice and at the risk of of being proven uh wrong with uh YouTube comments or comments from the audience here I think that computational topology tries to navigate this path a little bit so we we try to develop invariants that are sort of expressive while still being sort of computable of course that doesn't work in all dimensions and for all kinds of data sets but for lower dimensional data sets and lower dimensional topological features we are doing quite well and we'll now be looking into what this means in practice moving to computational topology and that's a new subfield that has been rising for like about a decade maybe give or take and in computational topology the idea is that you take all the methods methods from algebraic to apology and you make them actually computable you make them actually implementable on a computer and potentially also in your network we'll see a bunch of examples of this as well now reality of course is always messy maybe maybe often is not not even appropriate maybe it's always messy now what we see typically is we deal with a point cloud like this like just something that looks a little bit like a Taurus potentially what we see as a as a human or what we try to link this back to if you're a platonist and this is this is the thing that you are that you're looking for on the right hand side you try to link this to a an actual to an idealized Taurus and computational topology helps us bridge that Gap going from the left hand side to the right hand side going from the unstructured discrete Point Cloud to the nice shape on the right hand side now how does it do that I of course can't give a very nice introduction into the intricate details of the algorithms here so I'll just opt for a very intuitive and hopefully visually appealing way of doing things what we're essentially doing with these topological methods and persistent homology is one specific one here we approximate a point Cloud at different scales in the data so we look at it from near nearby and from far away and we observe how topological features appear and disappear as the scale changes for those of you in the know or for those of you who want to know more and be assured there will also be some references later on this is known as a vietroid's ribs complex calculation and interestingly enough I have to mention this historical tidbit because it's fascinating to me uh this was actually developed in the beginning of the 20th century so not the 21st mind you but the 20s so I think Leopold vietores wrote this seminal paper on calculating these types of complexes from point clouds in 1928. of course this terminology was a little bit off I mean he wouldn't say Point cloud or computer or whatever but the the principle is the same he was already thinking by then at about how to Leverage The Power of algebraic topology which was a very very young field back then as well in order to describe discrete data sets because he had the hunch that this might be something very relevant and very interesting I think this is also a very nice way of showing a conference between statistics and and data analysis because I think his main paper was motivated by uh tabulating certain statistical um uh calculations and statistical results from a census or something like this anyway this video is words complex is super easy to calculate we pick a distance and we pick a threshold Epsilon and then we just start connecting subsets if the pairwise distance of their Points Falls below that threshold and of course we do this only for pairs that are not identical and now as we grow this threshold here is a nice animation I've prepared you can see that more and more things start to be connected and now suppose that we're looking for Cycles in this data set then at some point going from this scale to this scale we have finally identified a cycle this cycle then persists for another scale this is also where the name persistent homology is coming from because we are tracking how long features survive in this process and after a certain threshold it gets closed again it gets swallowed because we have reached a maximum of our Zoo levels so to say now that's the intuition behind that you can actually use that for all kinds of Point clouds it's not restricted to point clouds we'll see also some examples of this later on in the talk but to illustrate this principle once more and what to actually do with these topological features let me show this to you with the with the projection of a 2d Point Cloud again we can start growing our euclidean spheres around the individual points and we can track topological features and in the end what we do get out of this and this is I think the key takeaway that I want you to have if you forget everything else from from my talk from the keynote please take away two things namely that Euler did a lot of stuff in topology and that this descriptor on the right hand side is called a persistence diagram the persistence diagram this little diagram on the right hand side captures topological features and it captures their creation and destruction across different scales so basically the more activity you have in there the more topological complexity there is in your object and you can do all kinds of interesting things with this of course and I promise that I'll go light on the formulas here but I just want to want to show this to you at least once to see why people are doing this and we'll see more motivations for why this is interesting in machine learning as well so the first thing that you can do with these persistence diagrams with these descriptors is you can calculate a distance between them in fact maybe you've heard about the concept optimal transport this is one of the nicest applications of optimal transport that I'm aware of so you can take two persistence diagrams D and D Prime and you can calculate what is known as the bottleneck distance or also the um some kind of Vassar Stein distance between them which is essentially solving an optimal matching problem so you try to take all the points in one diagram you try to match them to the other diagram and you have an internal cost for doing so and then you try to solve for the infimum over the supremum of this respective cost function so this is also where the idea of the bottleneck comes from you're searching for the best matching that you can find and then you take the largest distance the largest cost that you have to entail to make the matching stick that's known as the bottleneck distance there's a relaxed variant of this distance around as well which is known as the wasserstein distance technically I should say the wasserstein distance between persistence diagrams because you can of course also calculate the vasage time distance between probability distributions so these are actually the same it's the same distance in some sense it's just being calculated over different spaces now why would you want to do this well one nice thing is that persistence diagrams are actually stable under certain transformations of the data and in particular one stability that is really great for us when we're doing machine learning later on is that they are stable under certain sampling conditions so in particular you have probably heard about the term mini batch or batch in in machine learning of course so when you do batches when you take sub samples of your data and you work with them then you are almost virtually always guaranteed and we have actually a theorem that quantifies this a little bit how the persistent homology how the topological features of your data set behave under this sub sampling here I give you just an intuitive view so you can three three point clouds and you can see that the respective persistence diagrams in I think that's Dimension One they are more or less of the same shape I'm saying more or less because you can see that the blue point Cloud here oh no I'm not going I'm not going to try this because it's a it's an extended screen oh no but it works the blue point cloud has a little bit of a different sampling here so I I changed the sampling conditions here on purpose but you can see that the persistence diagram is still kind of kind of similar now let's make this more precise since we already have a distance definition we can also Define this more formally so if you have a triangular little space so something that you can add a triangulation to which in I would say our modern parlance is almost any data space that you will ever encounter and you have a continuous chain function that is a function that has only a finite number of critical points so it's not something that is really degenerate or or ill-behaved then the corresponding persistence diagrams satisfy a relationship in that their bottleneck distance is bounded by the house of distance between the functions and that's also very surprising result to me because if you let that sink in for a minute you have on the right hand side a housed off distance which is a fundamentally geometrical property but on the left hand side you have a topological distance you have a property between topological features so here geometry anthropology go hand in hand and one bounce the other which is really nice way to to think about these computational topology methods because and it's good that that this is that this is live streamed as well I think the field is misnomed it should also contain some form of geometry in there so if people here computation topology they think oh we're doing just discrete stuff and we throw geometry out of the window but that's actually not true so as you can see here geometry is being kept around and is being used now moving onward a little bit this is a slide that I that I discovered recently when I when I did some historical digging here because persistent homology is actually an idea that has been around for some time and it affords a very generic view on your data which is something that I'll try to convince you in the second part this keynote lecture when I chose some applications and I'll try to not butcher this quote too much there's a quote by Victory go which is or resistance so something very freely translated this means you can't resist the power of an idea whose time has come and this you can see this when you go through the papers that I mentioned here because there have been some precursors of persistent homology already back in the 90s they called it a distance for similarity classes of sub manifolds of euclidean space or they called it the frame Morse complex and its invariants or a size functions from a categoric Viewpoint all of these things are if you look at this mathematically precursors to the the things that I showed you before precursors to this idea of looking at data at various scales and seeing how its properties change but the fundamental or seminal paper that I was growing up with so to speak as a researcher is called topological persistence and simplification and let me just give you a quote here so you can see that these Notions are applicable in general settings in this paper the adults and colleagues write we formalize a notion of topological simplification within the framework of a filtration which is the history of a growing complex so already here you only need like some way to order your data in some sense of fashion and this can always be done um they then go on to say we classify topological change that happens during growth as either a feature or noise depending on its lifetime or persistence within the filtration we give Fast algorithms for computing persistence and experimental evidence for their speed and utility and this was done in 2002 and I would say it sparked this whole field of topological data analysis that we're now starting to reap the fruits within the context of machine learning now uh this is the this is the last slide before the applications actually and I was toying with myself I was I was arguing with myself whether I should name this slide why should you care so I didn't do this instead I asked the question now here as a rhetorical question this is the slide that tells you why should you care about these topological features this all looks nice it can be like a mathematical game mathematicians like to play they like to develop new stuff that's that's kind of cool okay fair enough but why should you care about this in the context of machine learning well there's a bunch of evidence um going back to to work I did with with Carson and colleagues um and and even some PhD students that are sitting here um and this this shows a lot of nice properties so for instance in the context of machine learning you can think of topological features as constituting an additional set of inductive biases so you already know that that the recent Paradigm Shift has started to happen in in deep learning so instead of saying okay we can learn everything we can from the data um people are now also using a specific inductive biases for specific tasks so for instance when you when you want permutation invariance or permutation equivalence so adjusting the model to respect certain things certain properties is uh has become a staple of modern machine learning research now uh in another fashion topological features can also be shown to complement existing machine learning algorithms and endow them with an explosivity that cannot be achieved otherwise so for instance the our recent paper on topological graph neural networks showed that with topological features we're able to escape the vice versailliman hierarchy so we are able to be together we are we are more expressive than the individual parts and last but certainly not least of course topological features have also some advantageous theoretical properties so for instance in our paper anthropological autoencoders the last one this slide we were actually looking at the sub-sampling conditions and we were able to show that provided your sub sampling of your data so your mini batches provided that they are kind of well behaved um not going into the details here then your topological features and also your reconstruction is also well behaved so this is essentially why you might want to care and why these why these features are are really being you useful now let me show you how to use this in the applications there is a generic topology driven machine learning Pipeline and recently we've we've started to upgrade this pipeline quite considerably let me let me show this to you so ordinarily people would use a Point Cloud they would do persistent homology then they would get persistence diagrams from out of there so these topological descriptors and then they would look at those diagrams and they would say okay this diagram tells me something about the data so they were being used as static features that can be used in an exploratory data analysis context but recently at some point people realized that hey wait we can also use them as input features for machine learning I mean that's not surprising to anyone here I guess if you have something that you can calculate and you can represent it somehow then you can also of course throw it as additional features in your um in your machine learning algorithm but the really interesting thing is and this is a really a brand new result that started to occur about I would say two years ago and this is we can actually back propagate information through this whole pipeline that is specifically a gradient from the machine learning algorithm whatever that might be from a deep Network for instance exists under certain mild conditions so for instance in the topological autoencoder papers we were able to enumerate this condition as saying that all the distances in your data set have to be well behaved you're not allowed to have infinite distances and you're also not allowed to have distances that are that are too close to each other so what I mean to say here is that like it is possible to go back and to go this pipeline in the other direction um as well and therein I would say lies the the true value of topological machine learning at the moment because you don't only get static features out there and you can say oh I'm looking at my coffee this morning and the coffee grounds they look a little bit different today so this means something but no no you can use those features in classification scenarios for instance you can use them for reconstruction purposes and for many other tasks as well now let's take a look at one specific application in the Life Sciences we use this back in the back in the days we use this for characterizing FM MRI data sets let me give you a brief rundown and I'm sorry for being a little bit cursory here because I'm not an expert in fmri of course so this is also my understanding of the technology fmri to my understanding measures some kind of blood oxygen level dependent activation in your brain so if you think about it some very very hard about something and certain brain areas are involved then I'm being told there's you need more oxygen there in there and you can and this slides up under the under the machine it is a technique that has temporal and spatial components so you do this um measurements not only for a single time step but you do this for longer time so people are lying in the MRI form I know 45 minutes um or maybe shorter I don't know um but you can do this for um for as long as you want and you can measure this signal all the time however since I guess every one of us and that's actually very philosophical issue so maybe we're at the right place here um every one of us perceives the world probably differently and has a different way of thinking about things hence in fmri data analysis you are plagued by a large degree of Interest subject variability so even if if you and I have ostensibly the same Hardware or I guess I should say wetware in this case because it's a brain then it will still behave a little bit differently so of course we know where the eyes are we know how the cortex Works sort of but still how I perceive a certain stimulus is different from how you perceive it properly moreover there are also issues with geometrical alignment and this is already where maybe the the alarm Bell should should start ringing you could say ah okay we take something that is maybe invariant to certain geometrical Transformations and yes this is what we did we used topological data analysis to characterize time varying fmri data specifically we characterize them using cubicle persistence that's in Europe's paper from 2020 and cubicle persistence not to go into the details here but it illustrates one of the nice points about this whole topological data analysis um framework namely that it indeed works for all kinds of data if you're able to rephrase your problem in a specific Manner and in this case we were able to reframe our problem as saying that well FMI data is dealing with volume data but volume data is something that can be considered a special type of topological complex in this case a cubicle complex and so with minor modifications all of the things that I said before so there's tracking of topological features Cycles voids and so on this works in this setting as well this this technique is built on on previous work by Wagner and colleagues published in 2012 on efficient computation of persistent homology for cubicle data now what we did specifically in this project is we looked at this bold activation function so the blood oxygen level dependent activation function we consider this to be a time varying function on some manifold and in um one of the few cases where this is actually working quite well we were able to also understand the manifold directly because the manifold was just the volume data that we got and so we were able to calculate topological features of this manifold measured via this function f and obtain stable topological summaries at different resolutions of this function now the main advantage of this is that this was working on the raw data and I'm putting raw in quotes here because it's not really the raw data my collaborators did a great job in in cleaning this up for us and aligning this of course but it is as raw as you can get without doing auxiliary representation so for instance if you're looking at some fmri Publications you will find that people often use an atlas so they think about which region should be present in the data or they use a correlation graph something like this we don't need all of these things in particular we don't need a we don't need to do certain modeling choices but we can use the data as is as some kind of time bearing volume and let me show you the pipeline with the cubic complex being highlighted as the central or pivotal element here so we start with an fmri stack on the left hand side we obtain an FMI volume from this by putting all of this together it's time varying this I can't show because else this slide would be a little bit would make you a little bit queasy I guess we transform all of this as a cubicle complex and from this we extract persistence diagrams so again these these nice diagrams that characterize topological features now if your attentive still at this late hour you might see that the persistence diagrams that I'm showing you here they have the third dimension here well well spotted in this case the third dimension is time so we're really lazy here and we're just stacking them on top of each other because we just use the time Dimension as an as an individual axis here notice that for those of you that are interested in Time series analysis in general we are for this Approach at least we're not using any relations between time steps so rather we are parallelizing everything and we're just treating every time step as an independent instance of of a topological expression we could do smarter things here in fact we're still working on this you will find some references to this but this is what we did back then and the data set that we are looking at comprised about 155 participants who were all watching the film partly cloudy so I do want to stress that this was not a distressing study for for for anyone because 122 of our participants were children and only 33 of them were adults so they were just watching the movie nothing else was done they didn't have to solve any tasks but what is this what this amounted to as I'm told is it's a continuous stimulation of participants so it's not something that is known as resting state data or resting state fmri or something like this but no no they had to watch the movie of course we didn't force them to watch the movie so they could have closed their eyes and dozed off we didn't actually we didn't actually enforce anything here but they all had the same stimulus which is great because now we can compare their responses to certain things in the movie and what we did first is we tried to predict their ages this is as I'm being told this is a neuroscience um I would say in the parlance of computer science it's a smoke test so it's a test for whether the representations that we are extracting are actually any use at all and it turns out that they have to do that to actually work with um an age prediction task we had to evaluate the norm of a persistence diagram so that's another neat mathematical property that you can have of these diagrams um it's essentially just the maximum of the points distances to the diagonal that you can have this Norm is also stable it's highly useful in particular when you want to obtain simple descriptions of time varying data sets because by calculating the norm of course you turn your high dimensional topological descriptor into a single time series and then you can evaluate this in other forms of fashion now uh to let's let's Feast our eyes briefly on this table here this is one of the nicest results of the study I would say this is the age prediction based on the summary statistics of all the participants you can see see that high scores are favorable here because it's a correlation coefficient so ideally you would want to have some kind of correlation coefficient of about I guess 0.9 we're of course not not there yet but still it's pretty pretty nice I think the mean squared error that we that we had was about um uh two or um two point something years um which is which is not too shabby you can see that in particular if you compare this with shared response models this SRM based technique which we only had available for a specific subset of our data set then we still outperform them considerably and I'm mentioning this because it's surprising that the data collection process in the data analysis process like ours which just looks at Raw data without any bells and whistles and which also doesn't include any prior biological or neuroscientific knowledge that this still works that well but that I think shows how expressive the topological representations can be for these tests of course age prediction is not something that you want to do in practice so you could do some a lot more thing things one thing that we tried and this is still ongoing work because that is really really complex and there's no pun intended with a complexity analysis we try to do complexity analysis based on the actual brain states that participants went through it turns out that the data are very noisy and we had to aggregate these um these complexities by the cohort so we had to aggregate by by the years and what we were looking for is to what extend a younger cohort is exhibiting higher topological uh sorry lower topological complexity than an older cohort with the idea being that if you're a very young child and you're watching this movie then you probably don't understand a lot of what is going on but you see that there's lots of sounds and noise and fury signifying nothing and if you're an adult you probably have a an emotional context and then and another context that you can that you can relate this to now again stressing this this is ongoing work for instance one thing that we're looking at nowadays is we're trying to relate such trajectories also with actual events and an emotional con text in the stimulus but this is of course hard to do so there's all kinds of interesting annotated data where we are where people are tracking facial expressions or where they are asking people in this movie scene which emotion are you predominantly experiencing anger shame disgust Joy surprise um apathy whatever um actually there's more negative there in there than positive ones but that's not my that's not my key so I'm not a responsible for this but so so this would be one of the steps where we want to take this next and then we would characterize the actual shape of the of our brain State trajectory as we um as we watch or as we encounter a stimulus now with this macroscopic uh considerations of a brain Let Me Maybe now zoom in quite a lot and talk briefly about the prediction of the shape of cells so now we have a we have a quite different task so now we're actually having something where we can measure the outcome quite considerably so this is relatively recent work um and also still ongoing because it's a complicated problem we'll see why this is the case so my collaborators from helmhos they have a lot of nice cell images and those are they say it's images of single cells but in case you're also as confused as I am this has nothing to do with single cell data analysis a single cell or scnr C analysis this is something completely different so what they mean is really like they take pictures of individual cells under the microscope that's great they use the confocal fluorescence microscope for this and then they're interested in predicting the 3D shape of a cell from this 2D image this is also known as a morphological analysis and it's it's a crucial way to detect certain pathologies so one very similar paper in this in this area is paper by Ford on red blood cell morphology and this states that when used properly RBC so red blood cell morphology can be a key tool for laboratory hematology professionals to recommend appropriate clinical and laboratory follow-up and to select the best tests for definitive diagnosis so in some sense and it's maybe maybe some of you already recognized this in some sense this is what a company called theranostrite to do um we're not we're not claiming that we that we can even do do five percent of what they claim to do because it turns out that they were a hoax company so bad for them potentially good for us but the the goal would really be if this works would really be that we take some blood sample we try to reconstruct this from Individual microscopy images and then we know something about the patient's state of health and I'm stressing this because and I hope no one is queasy to queasy in the audience here um blood is a really nice substance in that it's almost always available in patients so unless a patient is really really sick you can probably spare at least a drop of blood in the hospital so it's really a substance that you can easily get and you can easily analyze it so drawing inference making inferences from small quantities of human blood is a very nice well technology to to have in the future um spoiler alert we are not quite there yet but we're making some some progress and I'm going to show you what we were able to achieve with topology so let's first start without topology namely we take a look at the pipeline that we had before adding some topological information here what we do is we start with a 2d input on the left hand side we throw a machine learning model on there that's by the way that is dotted because it's of course something that you can easily replace if you find something better we let the model predict a 3D shape and then we use a geometrical loss term more about that in a minute and compared with the ground truth and the mathematicians or computer scientists in the audience they might appreciate this this is really hard and it's really hard because it's a complicated inverse problem so we're going from 2D to 3D I mean you already know that if I look at my shadow then I can reconstruct all kinds of interesting things so it is also essentially an ill-defined problem with a large number of potential Solutions so we do need a lot of input data to to make this work semi-reliably and I can already tell you that there will of course be cases depending on how you look at the cell where this reconstruction can never work because you're just missing features but we are content with capturing 80 of the cases maybe quite well and then raising a flag for the case that we can't handle either so that would already be a nice result there now what we had here is the the so-called shaper the shape reconstruction or reconstructing Auto encoder and this is a very simple machine learning technique that employs a convolutional neural network with some fully connected neural networks and it as you can see it it kind of decreases the picture first and then it blows it up again into a volume so it starts with a 64 by 64 image and then it gives you out a 64 cubed uh voxel volume and what this essentially does under the hood is it is learning a likelihood function a likelihood function is a function that Maps every voxel of this grid here so every point in R3 to some scalar value and this scalar value indicates the likelihood of a specific voxel being part of the true volume so that is what the what this method does and how it tries to reconstruct the images now for the normal loss function or for the geometry based loss function the shaper method uses a geometry based loss that consists of two components one is the dice lost the other one is a binary cross entropy loss and without going into the details here let me just give you the intuition here so essentially it compares the geometry of the resulting volumes on a per voxel basis so what it does is it's looking for whether the reconstructed volume is well aligned with the ground truth one but there's one issue at least namely on their own these issue these losses are not sufficient to capture shape variation because if I modify the shape a little bit then its topological characteristics of course don't change so I can rotate my icosahedron or my platonic solid in space all I want it's still a platonic solid but of course these losses that are very restricted to the voxels themselves they will then raise an alarm and will say no no this this deviates so you're learning a very restricted set of shapes or shape features less so than you could do in practice but now let's add topology to the mix and let's hope that this helps solve some problems this is what we did in our recent mikai paper so essentially if you recall the previous slide what we what we added is these three components below so we use topological features calculating of both the 3D prediction and the 3D ground truth and then we have a topology-based loss that we can combine with the geometry based loss above to obtain a joint loss and to kind of balance out the geometry-based Reconstruction and the topology-based reconstruction and to go into some more details about this loss it's something that you have encountered before in these slides it's the sum of wasserstein distances between the persistence diagrams and a term that I haven't introduced before but that I would just now call here total topological variation sometimes it's also known as total persistence now these terms have two different um components or two different uh result if I can call it that the first is you align the ground through the likelihood function and the predicted likelihood function f Prime so you want your ground truth and your predicted function to be as close as possible in the topological sense mind you the second term so this total variation term or this total topological variation term is only applied of course to the predicted likelihood function because we can only change that prediction right we cannot change the ground truth data and it's added there to reduce the geometrical topological variation of the predicted likelihood function so essentially and you can you can try this out I have a website for this if you want to check it out and this reduces the wriggles in the surface that you get because if you have a nice surface if you have a nice dodecahedron one of the issues with with the topological features is that I can add a lot of Wiggles around this surface and this won't change the overall topology but it's something that is really really undesirable and so adding this additional persistence term in the shown in blue here gets rid of this and then of course we can combine them and this is this is what you're all familiar with and what you all know and we can choose the Lambda parameter to oh what was that I hope okay okay great Cricket okay so nothing I didn't destroy anything that's that's good so and then we obtain a combined loss by um by adding those together and by by weighting them accordingly so now two interesting um uh things here so namely we of course write out what happens if you only go for the topology based loss then everything explodes again because topology on its own is not powerful enough to regularize your shape nicely but geometry on its own is also not powerful enough so there you have it this is really something that needs to be optimized jointly on the other hand one interesting tidbit that we found in this paper is that it's actually sufficient and this is one of the moments where well well in Heights that you're all you're always smarter right but in hindsight it was clear to me why this should be the case so I found a theorem from some topology book that explained this a little bit but um it turns out that we don't actually need the sum over these faster shine distances but it's sufficient to do one of our such line distance in a dimension two specifically because cells have kind of nice features and there's a duality in topological persistence going on that we that we can expire for these purposes I'm showing you the the nice version though or the the big beefed up version here though because that's what we initially trained with and then the rest is more like an empirical result that might not hold for generic data sets now let us briefly look at the results here and I I don't want to delve into the details what these error metrics mean but essentially what we found is that just by adding these simple calculation the simple loss term here we're able to reduce errors in all relevant metrics quite substantially except for of course um the odd one out there's always one experiment where this doesn't work um namely in the surface roughness for the nuclear data set um the results are still very very close and I think this is more like an initialization question so if we run this multiple times it might be that we that we end up having something something that is relatively close here one thing I do want to stress here though and this is one of the reasons why I'm excited about this project is that typically you might run into scalability issues if you have very high dimensional topological features that's not something that I that I mentioned too often here because it didn't appear in in these cases but for for this type of data set we are actually we have almost no performance decreases is whatsoever because it turns out that the topological features are by themselves expressive enough to be calculated on a very very simplistic rough version of the data so we don't have to use the 64 cubed volume data to calculate this part of the loss term but we can use a down sampled one and the down sampling is is almost almost free in terms of in terms of computational performance that's really nice because it really ties this together and also demonstrates those features bring in complementary perspectives now uh that's all I have for you today so I hope I was able to convince you a little bit about the fact that topology can provide useful inductive biases for shape reconstruction tasks in particular um I do want to stress that another takeaway so I already have two so don't forget about Euler being the man then persistence diagrams those are the topological descriptors and the third one is that they actually encode geometrical and topological properties of the data so don't get fooled by our bad advertising mathematicians are really bad at advertising and naming things Carson knows this from from our work together I really am bad at this um and so we we name it computation topology but it's actually actually more it's also it also encodes some geometrical properties not all of them but some and moreover and this is the fact that I'm most exciting about excited about the integration into standard machine learning methods is now possible so we have something for Auto encoders with something for graphs if something for shape reconstruction tasks more is hopefully to come I mean knock on wood right if you want want to learn more there's a recent survey that I that I co-authored with Felix Hensel and Michael Moore also of Carson's lab now I think a postdoc in Stanford it's called a survey of topological machine learning methods and it's um it's an Open Access publication in Frontiers and artificial intelligence and last but not least this is um this is now the advertisement party I hope it's okay um if you're interested in topological machine learning and you want to check this out on your own um my my lab and I we're trying to make the software work it's called python topological I know very creative name use you see I'm kind of like true true to form here it's not a creative name but it works and this gives you the power of topology at your fingertips and it at least at the moment of me saying this I mean depending on how fast the others are with the pull request we can do the fmri data analysis we can do the shape reconstruction stuff what we can't do yet and maybe someone in the audience wants to do that is we can't do the the graph neural networks yet but that's just a matter of time until we have a rest cute our old code and put it into into a nicer framework but it can already do quite a quite a few things and I'm happy to um discuss more and and maybe ask answer any questions about the software now of course I'm very happy that that I could be here and I'm looking forward to some of your questions now thank you very much [Applause] thank you very much Bastian both for jumping in for this inspiring talk thank you very much so are there questions for Bastian yeah I mean um so at a certain point I missed the step I think so you kind of lost me um no I'm sorry because maybe you're totally in that field you started out explaining that you measure the shape of this red blood cells and you mentioned that your colleagues have a confocal and somehow the next slide was how you want to go from a two-dimensional to a three-dimensional reconstruction but if you have a confocal you have the three-dimensional yes so at that point I lost the connection very very good point I'm I'm sorry this this is this is also my my lack of biology talking there so the way I understood their problem is that they say um these uh doing 3D reconstructions here fast and doing a lot of those is time consuming for them and doesn't two Dimensions yes exactly yeah yeah so so they they can do it in 3D I mean in fact um the the data that we got here and and this is actually I want to stress this because this is an actual um ground truth data that that we got so I didn't make this up or anything this is really um one of the ground truths from one of the predictions that we have from our algorithm they they are being done by by um by our collaborators um but they are telling me that this is a is a time consuming process and it doesn't doesn't scale very well so in the meantime what they are looking for is something that gives them that gets them eighty percent off the way there and that maybe can tell them oh this is a blood cell that looks very very anomalous from the 2D slides I'm sorry I should make this more clear but this is a very very good point thank you very much foreign so thank you it was incredibly fascinating I have a curiosity so uh so you showed quite a few results for images and 3D space yes how far can you push dimensionality in this setting or to put it another way how badly is these topological setting affected by the curves of dimensionality yes Ah that's a that's a that's a very good question so now okay I don't I don't wanna I don't want to do a politician's dance and give you a half answer so um let me just let me disentangle this though so first of all um it is not affected as much by the curse of dimensionality as other methods because um fundamentally it's it's built upon the idea of having good distances in your data and of course yeah euclidean distance suffers from this but you can throw your own distance in there you can even throw learn distances or metrics or Mahalo Nobis distance or whatever you want in there so in that sense this can be mitigated but where the curse of dimensionality re-hits us is when we go back let's hope that this works this is the first time I know one the second time actually I'm giving a talk since uh 2020 in in real life again so where this curse of Dimension and he really hits us is if you look at this V Epsilon expression on the bottom of the slide you have all these subsets that are within Epsilon or or less to each other and if you um if you don't restrict the size of your subset in the worst case you get 2 to the power of the number of your points of your subset so it is a very very bad scaling in this sense and this is where the curse of dimensionality hits us um this again can be mitigated by saying okay we're only interested in topological features of a certain Dimension so for instance we can say that empirically speaking for graphs it's sufficient to do 0d and one DV just to already get a nice performance improvements for higher Dimensions 2D might still be feasible I personally haven't encountered a data set where really much more Beyond 2D was required I mean you can build data sets where you can only characterize your your data by by having higher order Dimensions but it's it's rare and now to give you a very precise answer so this scalability is abysmal you have to you have to cheat yourself a little bit around this it it is possible but it's it's hard to do um but at least if you have a 1 000 dimensional Point cloud and you're only interested in a bunch of low dimensional topological features which can still be very expressive mind you then it's still something that can be that can be applied I hope this was good thank you you're welcome Sean Philippe um [Music] yeah thanks a very great talk uh and my question was in terms of application so you you mentioned a few of them but I'm sure there are plenty of others so uh what about a structural bioinformatics so stretches of molecules either small molecules of proteins and in particular since you mentioned that you can back propagate it sounds like things like you know Alpha fold Etc which predict the structure using some loss functions potentially could also use some some real stuff look at that absolutely no not yet to be honest so I would I would very much love to so in fact I think that um our success in this topological graph new networks kind of gave us the motivation to dig a little bit deeper into this realm um there I I do think that that this is one of the application areas where we would require a joint optimization or kind of a joint view on the data because I don't think that the topological features on their own are any more powerful than what is already out there but potentially if we phrase this right and if we set up the task right then we could have something that is really complementary because it can capture things that you cannot capture in in other ways so yeah I would definitely be interested in this there's also some I only mentioned this briefly I think let me go back to this yeah it is actually the right slide so no it's actually not uh there is work with Leslie but okay Leslie is on the slide that's good there's other work with Leslie where we're looking at evaluating genitive models for for graphs and here I think a topological perspective would also be very much warranted because graphs are already topological objects so it would be very interesting to kind of characterize the expressivity of a generator of a distribution in terms of its topological properties so so to say can we get all the modes that live in this space or are we restricted to a certain class of graphs but I I also have to say that this is living a little bit in future worlds for now so ongoing or rather rather planned as research but definitely interested in I think definitely worthwhile and can I have a second second question or someone else yeah so super technical question uh when you talk about the differentiability yeah of uh the topological representation with respect to the input data uh can you say a bit more about so I suspect the function is not differentiable but maybe everywhere almost every different channels the question is you know yes since the points appear there's area where there is no point and suddenly it moves away from the diagonal exactly so so one thing that we exploit here maybe just to go back to this to this diagram here so one of the there's multiple ways of of going around this some of my colleagues for instance um uh um El Canon Solomon has been has been looking into this um and Matthew career as well um uh here for this diagram here we would exploit the fact that um every Point has at least a very very small neighborhood around it around itself that doesn't contain any other points so there's like no overlapping points in this diagram if we have this condition then we can show that um that the mapping from the space to the diagram is constant and then the the the the the composition rule of gradients tells us that we can that we can ignore this part there's also more technical results if you use different representations because a lot of things exist in this space that I haven't told you about unfortunately here in this talk apologies for this if you use a different representation of your topological features and then you can show for instance that the mapping is is lip shits and then you can allude to the now now I'm blanking on the name of theorem but then you can allude to a theorem that tells you that the mapping is on differentiable almost everywhere and then you need some computational tricks to make this gradient actually unique in in practice but it can be done and it works even even kind of out of the box with pytorch it works surprisingly surprisingly well even if you kind of ignore some some of the degenerate cases it would work surprisingly well here thank you you're very welcome so let's thank Bastian again for for this very inspiring keynote thank you Bastian [Applause]