Video summary
Bastian Rieck from the Helmholtz Center Munich delivered a keynote on leveraging topological data analysis (TDA) to understand shapes that are deformed or perturbed rather than perfect, moving beyond traditional geometric constraints. He introduced fundamental algebraic topology concepts such as Euler's Seven Bridges problem and the Euler characteristic, explaining how these invariants classify spaces up to homeomorphism by allowing stretching and bending without tearing. The core of his talk focused on computational topology through persistent homology, a method that approximates point clouds at varying scales using Vietoris-Rips complexes to track the birth and death of topological features like connected components, cycles, and voids across a filtration. This process generates persistence diagrams that effectively capture an object's complexity while offering two distinct advantages: stability against small perturbations and subsampling, which links geometry and topology via bottleneck distance, and seamless integration with machine learning where these features serve as inductive biases to complement existing algorithms and potentially escape limitations like the VC dimension hierarchy.
To address the complexities of shape reconstruction as an inverse problem involving 2D-to-3D conversion, Rieck discussed a specialized autoencoder called "shaper" that maps voxels in a grid using convolutional neural networks trained with geometry-based loss functions like Dice loss and binary cross-entropy. However, he noted that these standard losses are insufficient because they fail to account for shape variations when topological characteristics remain unchanged despite geometric modifications, such as rotating an object. To overcome this limitation, his recent work introduces a combined approach using both topology-based and geometry-based losses, incorporating Wasserstein distances between persistence diagrams and a total topological variation term specifically applied to predicted likelihood functions. This joint strategy successfully reduces undesirable surface wiggles that do not alter overall topology while maintaining geometric alignment, demonstrating substantial error reduction across various metrics and proving scalability by showing that topological features remain expressive even when calculated on downsampled data without performance degradation.
The practical applications of these methods were illustrated through fMRI analysis and cell morphology prediction, where cubical persistence was used to characterize time-varying brain activation data while handling subject variability and geometric alignment issues with high success in predicting participant ages from movie-watching stimuli. In the realm of biology, the technique reconstructs 3D cell shapes directly from 2D microscopy images rather than relying on single-cell RNA sequencing, enabling robust detection of pathologies such as abnormalities in red blood cells even when working with small biological samples. During the Q&A session, Rieck clarified that while collaborators utilize confocal microscopy for high-quality 3D imaging, this algorithm aims to provide a faster alternative achieving approximately 80% accuracy by flagging anomalies from 2D slides alone, addressing concerns about dimensionality and differentiability within standard frameworks like PyTorch. Although current implementations do not yet integrate with structural bioinformatics tasks like AlphaFold due to the need for joint optimization, Rieck confirmed interest in such developments while explaining that mappings are constant or Lipschitz in specific neighborhoods, allowing gradients to be computed almost everywhere despite degenerate cases.
Read the full video transcript
that our last keynote speaker had fallen
sick so we were left without the final
keynote speaker but luckily Munich is
full of great researchers and luckily we
know a few and we could
um we could convince Bastian Rick from
the helmhold center in Munich to come
and to deliver the final keynote
um of of our Symposium as a replacement
so so bastian's background is in
mathematics he's one of the experts in
topological data analysis in machine
learning he did his PhD in in Heidelberg
and then
postdoc postdoctoral time in Kansas
Sultan and later on in ADH Circ in my
lab in fact and then he became a pi at
the helmhold center here in Munich
recently he's a rising star at this
intersection of topological data
analysis and machine learning he's also
the program chair of this new learning
on graphs conference log
that's happening for the first time this
year so um we now move from the more
medical perspective of process medicine
from the more medical perspective
through more mathematical perspective on
what you can do in the Life Sciences and
in in the in in data science in the Life
Sciences
thank you very much for coming on short
notice Bastian we are really happy to
have you here and we're looking forward
to your talk thank you very much for for
having me it's the microphone already
live I have to think it's not maybe I
need to do something
um or is it just black either is it
already line ah yes now I know this is
so much better thank you very much for
your for your really kind words
um it's also a pleasure to be here on on
short notice so I I coupled this
together for a general audience as I was
told so I tried to take everyone with me
and I hope that we can have a nice q a
session as well so the title of of this
is more like a framework I would say
it's called a good scale it's hard to
find shape analysis using topology and
first let's maybe give you a small
induction of what I like to do this is
one of my favorite pictures it's not
drawn by myself it's drawn by an AI it's
supposed to represent this idea of what
you can do with topology and how shapes
are kind of melting and maybe it has a
certain serial characteristic which is
which is what I'm going for because in
practice we almost never are dealing
with the with the right shape but we're
dealing with deformations we're dealing
with deformed perturbed variants of such
a shape now let me talk a little bit
about algebraic topology my special
subfield back in the days when the
dinosaurs roam the Earth this is what I
studied in mathematics it's a fear that
is supposedly about counting and
calculating stuff and when you look up a
definition of algebraic topology you
will find a bunch of them of course so a
very prosaic one is we want to develop
invariants that classify topological
spaces up to homoromorphism there's a
bunch of weird words already in there we
don't have to use them another one is we
can use tools from algebra to study
topological spaces whatever these
topological spaces are
but what I like to think about in the
terms of topology is we want to
understand shapes through calculations
and I'm stressing the calculations part
because we humans have very good visual
system and probably most of you know
this better than than I do I mean mine
as you can see is also not working
properly most of the time so we're
really really good at recognizing shapes
really really good at seeing each other
and the question is now how can we bring
this into the computer world how can we
somehow leverage things that our cortex
can do on its own
now a first taste and this is probably
something that you will encounter lots
of times when you look at my stuff or
things that deal with topology in
general that's the Seven Bridges of
koenig's back problem the first
I would say first historical occurrence
of topological data analysis if you want
to call it that it dates back to Euler
because of course it does everything
dates back to Euler if you look long
enough in mathematics either that or
Gauss so I think I said this in another
talk before but if you ever do a pub
quiz and you're asked about mathematics
then Euler and Gauss are your sure
guesses I would say now in the Seven
Bridges I couldn't expect the question
is can you do a walk through the city
that crosses every bridge of koenigsberg
exactly once and if you look at the city
there's a map so there's some geometry
there and it's kind of hard to do that
and you could probably walk through
clinics back all day long it's a nice
city I've heard so you could do that or
you could abstract this and you could
build a graph that has the the bridges
as its nodes and then it tries to build
a connectivity from from up there and
when you do this you will find that this
is actually one of the first theorems
that you often learn in graph Theory and
undergraduate graph Theory no such War
can actually exist because there are
more than two vertices with odd degree
and this is so fascinating to me because
this is a geometric problem or well
let's maybe not stretch real world
problem too much here but for me this is
a real world problem and you've
abstracted it just by virtue of putting
it into a graph and then ascertaining
some of the properties of that graph and
that has almost a magical character to
it to me and this is one of the aspects
of topology and topological data
analysis that's of course not the only
thing we can do so I have said this
before we'll stay with Euler for a while
because he's he's the man so there's
also other invariants that we can
calculate of spaces and invariant is
something that remains fixed while we
transform the space where we subject it
to certain Transformations the
Transformations can be homomorphism so
stretching something or bending it
without carrying it or they can be
something different if you are dealing
with graph machine learning for instance
you have probably encountered the term
permutation invariant or mutation Equity
variant
sometime and this is exactly one of
those invariants that people are now
looking for that they're interested in
because if I give you a graph or a
machine learning algorithm then probably
the output should not change if you just
change the ordering of the vertices it
should change however if you start
rewiring the graph and
coming back to this Euler characteristic
here the Euler characteristic is a very
simple way of defining a polyhedron so a
shape that we can nicely draw it's
defined as the number of vertices V
minus the number of edges e plus the
number of faces F respectively and you
can calculate this and there's a nice
theorem of course attached to this
because this is how math works and you
can show that the Euler characteristic
of every platonic solid is exactly two
interestingly this theorem also
characterizes platonic solid so if you
if you have this theorem you can
characterize the solids or you can do it
the other way around because you can
show that any one of those shapes must
by necessity be a platonic solid if it
has this Euler characteristic here and
satisfy certain other properties now
some of you might recognize those I took
them I think from a very nice book by
Kepler the symbols they're not meant to
to represent anything here I think this
was some alchemical part here which
which we're not doing here because
that's that's of course not science but
I think it looks nicely drawn so we can
check this of course let's briefly walk
through this just so you can see that
I'm not lying to you or at least I'm not
lying on purpose to you I might be lying
because I'm just plain wrong in some of
the aspects here because that's also
something that we have in science but
it's not it's not a misinformation and
purpose here so for the tetrahedron we
can do four minus six plus four that's
uh two we can do the same for hexahedron
for the octahedron for the dodecahedron
and for the iconsahedron and there we
have it and by the way I think I
switched uh two of these places here
that was just for someone to check but
yeah I guess uh I I guess I could have
done it on purpose as well now let's
walk
um through some more invariants here and
see what we can what we can do in higher
Dimensions because well I mean platonic
solids are all well and good but I mean
I don't know about you but the last time
I encountered a platonic solid was when
I was playing Dungeons and Dragons about
10 years ago and so this is not really
what we're dealing with in science
anymore but
we can go higher of course there's
another nice invariant that is called
the Betty number it's um named after
Enrico Betty and the this D dating
number counts the number of
d-dimensional holes in a space whatever
that might be we'll see some some
examples of this later on and primarily
it can be used to distinguish between
spaces because it is what we call a
homeomorphism invariant so if you bend
your space if you stretch it if you tear
it a little bit it will not change so in
that sense it is a characteristic
property of a space
the Betty numbers have very nice
properties in that they represent I
would say intuitively
capturable objects or captural
characteristics for Dimension zero for
instance these are the connected
components for Dimension One these are
the cycles in the data set or in a graph
and four dimension two these are the
voids so for instance when you think
about proteins or molecules or something
like this and you you thicken them a
little bit then in D equals two you will
find the pockets that they enclose so
I've been told that this is something
that people are being interested in I
myself have not been working with such
data so this is just uh this is just an
educated guess on my part
let's look at some examples here the
ordinary point so nothing without any
extents has the BT number of 1 and 0 and
0 respectively so just one connected
component nothing else going on we can
do a little bit more with the cube a
cube which I assume to be a thing that
contains some space has a bit number of
one in dimension two because you it
closes one void and we can do the same
for the sphere and for the Toros as well
I'm not going to go into the details why
the Taurus has exactly two cycles here
that's actually a very interesting and
and deep theorem in topology as well but
if you want to go for an intuitive
explanation I would say there's one
cycle that you can immediately see
because it's the one that you can put
your finger through so if you eat a
donut like this you can eat it around
your finger and the other cycle you can
do by thinking of hanging it up on a
Christmas tree as an ornament that's the
second cycle you can find now at this
point since we're now working towards
computation topology let me say a few
words about why these invariants are
useful in general so not the Betty
number specifically but why it's really
useful to have invariants when you
define an invariant when you look for
characteristic properties you're always
in the struggle between having something
that is extremely precise so you want
something that tells your data sets
apart that has very high expressive
power so for instance for those of you
that are into the craft machine learning
domain a little bit you might have heard
about device from the limit hierarchy or
the vice family graph kernel or the vice
element test for subgraph isomorphism
and these are exactly things that you
are that you're looking for so you want
something that is very expressive and
that tells your data sets apart on the
other hand and that's that's kind of the
this uh the balancing part of this coin
you also want something that you can
compute effectively and efficiently it's
not useful if you have an invariant that
is NP hard to compute the way you don't
have an efficient algorithm because then
it doesn't scale and you can't use it in
practice and
at the risk of of being proven uh wrong
with uh YouTube comments or comments
from the audience here I think that
computational topology tries to navigate
this path a little bit so we we try to
develop invariants that are sort of
expressive while still being sort of
computable of course that doesn't work
in all dimensions and for all kinds of
data sets but for lower dimensional data
sets and lower dimensional topological
features we are doing quite well and
we'll now be looking into what this
means in practice moving to
computational topology and that's a new
subfield that has been rising for like
about a decade maybe give or take and in
computational topology the idea is that
you take all the methods methods from
algebraic to apology and you make them
actually computable you make them
actually implementable on a computer and
potentially also in your network we'll
see a bunch of examples of this as well
now reality of course is always messy
maybe maybe often is not not even
appropriate maybe it's always messy now
what we see typically is we deal with a
point cloud like this like just
something that looks a little bit like a
Taurus potentially what we see as a as a
human or what we try to link this back
to if you're a platonist and this is
this is the thing that you are that
you're looking for on the right hand
side you try to link this to a an actual
to an idealized Taurus and computational
topology helps us bridge that Gap going
from the left hand side to the right
hand side going from the unstructured
discrete Point Cloud to the nice shape
on the right hand side
now how does it do that I of course
can't give a very nice introduction into
the intricate details of the algorithms
here so I'll just opt for a very
intuitive and hopefully visually
appealing way of doing things what we're
essentially doing with these topological
methods and persistent homology is one
specific one here we approximate a point
Cloud at different scales in the data so
we look at it from near nearby and from
far away and we observe how topological
features appear and disappear as the
scale changes for those of you in the
know or for those of you who want to
know more and be assured there will also
be some references later on this is
known as a vietroid's ribs complex
calculation and interestingly enough I
have to mention this historical tidbit
because it's fascinating to me uh this
was actually developed in the beginning
of the 20th century so not the 21st mind
you but the 20s so I think Leopold
vietores wrote this seminal paper on
calculating these types of complexes
from point clouds in 1928. of course
this terminology was a little bit off I
mean he wouldn't say Point cloud or
computer or whatever but the the
principle is the same he was already
thinking by then at about how to
Leverage The Power of algebraic topology
which was a very very young field back
then as well in order to describe
discrete data sets because he had the
hunch that this might be something very
relevant and very interesting I think
this is also a very nice way of showing
a conference between statistics and and
data analysis because I think his main
paper was motivated by uh tabulating
certain statistical
um uh calculations and statistical
results from a census or something like
this anyway this video is words complex
is super easy to calculate we pick a
distance and we pick a threshold Epsilon
and then we just start connecting
subsets if the pairwise distance of
their Points Falls below that threshold
and of course we do this only for pairs
that are not identical and now as we
grow this threshold here is a nice
animation I've prepared you can see that
more and more things start to be
connected and now suppose that we're
looking for Cycles in this data set then
at some point going from this scale to
this scale we have finally identified a
cycle this cycle then persists for
another scale this is also where the
name persistent homology is coming from
because we are tracking how long
features survive in this process and
after a certain threshold it gets closed
again it gets swallowed because we have
reached a maximum of our Zoo levels so
to say now that's the intuition behind
that you can actually use that for all
kinds of Point clouds it's not
restricted to point clouds we'll see
also some examples of this later on in
the talk but to illustrate this
principle once more and what to actually
do with these topological features let
me show this to you with the with the
projection of a 2d Point Cloud again we
can start growing our euclidean spheres
around the individual points and we can
track topological features and in the
end what we do get out of this and this
is I think the key takeaway that I want
you to have if you forget everything
else from from my talk from the keynote
please take away two things namely that
Euler did a lot of stuff in topology and
that this descriptor on the right hand
side is called a persistence diagram the
persistence diagram this little diagram
on the right hand side captures
topological features and it captures
their creation and destruction across
different scales so basically the more
activity you have in there the more
topological complexity there is in your
object
and you can do all kinds of interesting
things with this of course and I promise
that I'll go light on the formulas here
but I just want to want to show this to
you at least once to see why people are
doing this and we'll see more
motivations for why this is interesting
in machine learning as well so the first
thing that you can do with these
persistence diagrams with these
descriptors is you can calculate a
distance between them in fact maybe
you've heard about the concept optimal
transport this is one of the nicest
applications of optimal transport that
I'm aware of so you can take two
persistence diagrams D and D Prime and
you can calculate what is known as the
bottleneck distance or also the um some
kind of Vassar Stein distance between
them which is essentially solving an
optimal matching problem so you try to
take all the points in one diagram you
try to match them to the other diagram
and you have an internal cost for doing
so and then you try to solve for the
infimum over the supremum of this
respective cost function so this is also
where the idea of the bottleneck comes
from you're searching for the best
matching that you can find and then you
take the largest distance the largest
cost that you have to entail to make the
matching stick that's known as the
bottleneck distance there's a relaxed
variant of this distance around as well
which is known as the wasserstein
distance technically I should say the
wasserstein distance between persistence
diagrams because you can of course also
calculate the vasage time distance
between probability distributions so
these are actually the same it's the
same distance in some sense it's just
being calculated over different spaces
now
why would you want to do this well one
nice thing is that persistence diagrams
are actually stable under certain
transformations of the data and in
particular one stability that is really
great for us when we're doing machine
learning later on is that they are
stable under certain sampling conditions
so in particular you have probably heard
about the term mini batch or batch in in
machine learning of course so when you
do batches when you take sub samples of
your data and you work with them then
you are almost virtually always
guaranteed and we have actually a
theorem that quantifies this a little
bit how the persistent homology how the
topological features of your data set
behave under this sub sampling here I
give you just an intuitive view so you
can three three point clouds and you can
see that the respective persistence
diagrams in I think that's Dimension One
they are more or less of the same shape
I'm saying more or less because you can
see that the blue point Cloud here oh no
I'm not going I'm not going to try this
because it's a it's an extended screen
oh no but it works the blue point cloud
has a little bit of a different sampling
here so I I changed the sampling
conditions here on purpose but you can
see that the persistence diagram is
still kind of kind of similar now let's
make this more precise since we already
have a distance definition we can also
Define this more formally so if you have
a triangular little space so something
that you can add a triangulation to
which in I would say our modern parlance
is almost any data space that you will
ever encounter and you have a continuous
chain function that is a function that
has only a finite number of critical
points so it's not something that is
really degenerate or or ill-behaved then
the corresponding persistence diagrams
satisfy a relationship in that their
bottleneck distance is bounded by the
house of distance between the functions
and that's also very surprising result
to me because if you let that sink in
for a minute you have on the right hand
side a housed off distance which is a
fundamentally geometrical property but
on the left hand side you have a
topological distance you have a property
between topological features so here
geometry anthropology go hand in hand
and one bounce the other which is really
nice way to to think about these
computational topology methods because
and it's good that that this is that
this is live streamed as well I think
the field is misnomed it should also
contain some form of geometry in there
so if people here computation topology
they think oh we're doing just discrete
stuff and we throw geometry out of the
window but that's actually not true so
as you can see here geometry is being
kept around and is being used now
moving onward a little bit this is a
slide that I that I discovered recently
when I when I did some historical
digging here because persistent homology
is actually an idea that has been around
for some time and it affords a very
generic view on your data which is
something that I'll try to convince you
in the second part this keynote lecture
when I chose some applications and I'll
try to not butcher this quote too much
there's a quote by Victory go which is
or resistance
so something very freely translated this
means you can't resist the power of an
idea whose time has come and this you
can see this when you go through the
papers that I mentioned here because
there have been some precursors of
persistent homology already back in the
90s they called it a distance for
similarity classes of sub manifolds of
euclidean space or they called it the
frame Morse complex and its invariants
or a size functions from a categoric
Viewpoint all of these things are if you
look at this mathematically precursors
to the the things that I showed you
before precursors to this idea of
looking at data at various scales and
seeing how its properties change but the
fundamental or seminal paper that I was
growing up with so to speak as a
researcher is called topological
persistence and simplification and let
me just give you a quote here so you can
see that these Notions are
applicable in general settings in this
paper the adults and colleagues write we
formalize a notion of topological
simplification within the framework of a
filtration which is the history of a
growing complex so already here you only
need like some way to order your data in
some sense of fashion and this can
always be done
um they then go on to say we classify
topological change that happens during
growth as either a feature or noise
depending on its lifetime or persistence
within the filtration we give Fast
algorithms for computing persistence and
experimental evidence for their speed
and utility and this was done in 2002
and I would say it sparked this whole
field of topological data analysis that
we're now starting to reap the fruits
within the context of machine learning
now
uh this is the this is the last slide
before the applications actually and I
was toying with myself I was I was
arguing with myself whether I should
name this slide why should you care so I
didn't do this instead I asked the
question now here as a rhetorical
question this is the slide that tells
you why should you care about these
topological features this all looks nice
it can be like a mathematical game
mathematicians like to play they like to
develop new stuff that's that's kind of
cool okay fair enough but why should you
care about this in the context of
machine learning well there's a bunch of
evidence
um going back to to work I did with with
Carson and colleagues
um and and even some PhD students that
are sitting here
um and this this shows a lot of nice
properties so for instance in the
context of machine learning you can
think of topological features as
constituting an additional set of
inductive biases so you already know
that that the recent Paradigm Shift has
started to happen in in deep learning so
instead of saying okay we can learn
everything we can from the data
um people are now also using a specific
inductive biases for specific tasks so
for instance when you when you want
permutation invariance or permutation
equivalence so adjusting the model to
respect certain things certain
properties is uh has become a staple of
modern machine learning research now
uh in another fashion topological
features can also be shown to complement
existing machine learning algorithms and
endow them with an explosivity that
cannot be achieved otherwise so for
instance the our recent paper on
topological graph neural networks showed
that with topological features we're
able to escape the vice versailliman
hierarchy so we are able to be together
we are we are more expressive than the
individual parts and last but certainly
not least of course topological features
have also some advantageous theoretical
properties so for instance in our paper
anthropological autoencoders the last
one this slide we were actually looking
at the sub-sampling conditions and we
were able to show that provided your sub
sampling of your data so your mini
batches provided that they are kind of
well behaved
um not going into the details here then
your topological features and also your
reconstruction is also well behaved so
this is essentially why you might want
to care and why these why these features
are are really being you useful now
let me show you how to use this in the
applications there is a generic topology
driven machine learning Pipeline and
recently we've we've started to upgrade
this pipeline quite considerably let me
let me show this to you so ordinarily
people would use a Point Cloud they
would do persistent homology then they
would get persistence diagrams from out
of there so these topological
descriptors and then they would look at
those diagrams and they would say okay
this diagram tells me something about
the data so they were being used as
static features that can be used in an
exploratory data analysis context
but recently at some point people
realized that hey wait we can also use
them as input features for machine
learning I mean that's not surprising to
anyone here I guess if you have
something that you can calculate and you
can represent it somehow then you can
also of course throw it as additional
features in your um in your machine
learning algorithm but the really
interesting thing is and this is a
really a brand new result that started
to occur about I would say two years ago
and this is we can actually back
propagate information through this whole
pipeline that is
specifically a gradient from the machine
learning
algorithm whatever that might be from a
deep Network for instance exists under
certain mild conditions so for instance
in the topological autoencoder papers we
were able to enumerate this condition as
saying that all the distances in your
data set have to be well behaved you're
not allowed to have infinite distances
and you're also not allowed to have
distances that are that are too close to
each other so what I mean to say here is
that like it is possible to go back and
to go this pipeline in the other
direction
um as well and therein I would say lies
the the true value of topological
machine learning at the moment because
you don't only get static features out
there and you can say oh I'm looking at
my coffee this morning and the coffee
grounds they look a little bit different
today so this means something but no no
you can use those features in
classification scenarios for instance
you can use them for reconstruction
purposes and for many other tasks as
well now let's take a look at one
specific application in the Life
Sciences we use this back in the back in
the days we use this for characterizing
FM MRI data sets let me give you a brief
rundown and I'm sorry for being a little
bit cursory here because I'm not an
expert in fmri of course so this is also
my understanding of the technology fmri
to my understanding measures some kind
of blood oxygen level dependent
activation in your brain so if you think
about it some very very hard about
something and certain brain areas are
involved then I'm being told there's you
need more oxygen there in there and you
can and this slides up under the under
the machine it is a technique that has
temporal and spatial components so you
do this
um measurements not only for a single
time step but you do this for longer
time so people are lying in the MRI form
I know 45 minutes
um or maybe shorter I don't know
um but you can do this for um for as
long as you want and you can measure
this signal all the time however
since I guess every one of us and that's
actually very philosophical issue so
maybe we're at the right place here
um every one of us perceives the world
probably differently and has a different
way of thinking about things hence in
fmri data analysis you are plagued by a
large degree of Interest subject
variability so even if if you and I have
ostensibly the same Hardware or I guess
I should say wetware in this case
because it's a brain then it will still
behave a little bit differently so of
course we know where the eyes are we
know how the cortex Works sort of but
still how I perceive a certain stimulus
is different from how you perceive it
properly moreover there are also issues
with geometrical alignment and this is
already where maybe the the alarm Bell
should should start ringing you could
say ah okay we take something that is
maybe invariant to certain geometrical
Transformations and yes this is what we
did we used topological data analysis to
characterize time varying fmri data
specifically we characterize them using
cubicle persistence that's in Europe's
paper from 2020 and cubicle persistence
not to go into the details here but it
illustrates one of the nice points about
this whole topological data analysis
um framework namely that it indeed works
for all kinds of data if you're able to
rephrase your problem in a specific
Manner and in this case we were able to
reframe our problem as saying that well
FMI data is dealing with volume data but
volume data is something that can be
considered a special type of topological
complex in this case a cubicle complex
and so with minor modifications all of
the things that I said before so there's
tracking of topological features Cycles
voids and so on this works in this
setting as well this this technique is
built on on previous work by Wagner and
colleagues published in 2012 on
efficient computation of persistent
homology for cubicle data now what we
did specifically in this project is we
looked at this bold activation function
so the blood oxygen level dependent
activation function we consider this to
be a time varying function on some
manifold and in
um one of the few cases where this is
actually working quite well we were able
to also understand the manifold directly
because the manifold was just the volume
data that we got and so we were able to
calculate topological features of this
manifold measured via this function f
and obtain stable topological summaries
at different resolutions of this
function now
the main advantage of this is that this
was working on the raw data and I'm
putting raw in quotes here because it's
not really the raw data my collaborators
did a great job in in cleaning this up
for us and aligning this of course but
it is as raw as you can get without
doing auxiliary representation so for
instance if you're looking at some fmri
Publications you will find that people
often use an atlas so they think about
which region should be present in the
data or they use a correlation graph
something like this we don't need all of
these things in particular we don't need
a we don't need to do certain modeling
choices but we can use the data as is as
some kind of time bearing volume
and let me show you the pipeline with
the cubic complex being highlighted as
the central or pivotal element here so
we start with an fmri stack on the left
hand side we obtain an FMI volume from
this by putting all of this together
it's time varying this I can't show
because else this slide would be a
little bit would make you a little bit
queasy I guess we
transform all of this as a cubicle
complex and from this we extract
persistence diagrams so again these
these nice diagrams that characterize
topological features now if your
attentive still at this late hour you
might see that the persistence diagrams
that I'm showing you here they have the
third dimension here well well spotted
in this case the third dimension is time
so we're really lazy here and we're just
stacking them on top of each other
because we just use the time Dimension
as an as an individual axis here notice
that for those of you that are
interested in Time series analysis in
general we are for this Approach at
least we're not using any relations
between time steps so rather we are
parallelizing everything and we're just
treating every time step as an
independent instance of of a topological
expression we could do smarter things
here in fact we're still working on this
you will find some references to this
but this is what we did back then and
the data set that we are looking at
comprised about 155 participants who
were all watching the film partly cloudy
so I do want to stress that this was not
a distressing study for for for anyone
because 122 of our participants were
children and only 33 of them were adults
so they were just watching the movie
nothing else was done they didn't have
to solve any tasks but what is this what
this amounted to as I'm told is it's a
continuous stimulation of participants
so it's not something that is known as
resting state data or resting state fmri
or something like this but no no they
had to watch the movie of course we
didn't force them to watch the movie so
they could have closed their eyes and
dozed off we didn't actually we didn't
actually enforce anything here but they
all had the same stimulus which is great
because now we can compare their
responses to certain things in the movie
and what we did first is we tried to
predict their ages this is as I'm being
told this is a neuroscience um I would
say in the parlance of computer science
it's a smoke test so it's a test for
whether the representations that we are
extracting are actually any use at all
and it turns out that they have to do
that to actually work with um an age
prediction task we had to evaluate the
norm of a persistence diagram so that's
another neat mathematical property that
you can have of these diagrams
um it's essentially just the maximum of
the points distances to the diagonal
that you can have this Norm is also
stable it's highly useful in particular
when you want to obtain simple
descriptions of time varying data sets
because by calculating the norm of
course you turn your high dimensional
topological descriptor into a single
time series and then you can evaluate
this in other forms of fashion now uh to
let's let's Feast our eyes briefly on
this table here this is one of the
nicest results of the study I would say
this is the age prediction based on the
summary statistics of all the
participants you can see see that high
scores are favorable here because it's a
correlation coefficient so ideally you
would want to have some kind of
correlation coefficient of about I guess
0.9 we're of course not not there yet
but still it's pretty pretty nice I
think the mean squared error that we
that we had was about
um uh two or um two point something
years
um which is which is not too shabby you
can see that in particular if you
compare this with shared response models
this SRM based technique which we only
had available for a specific subset of
our data set then we still outperform
them considerably and I'm mentioning
this because it's surprising that the
data collection process in the data
analysis process like ours which just
looks at Raw data without any bells and
whistles and which also doesn't include
any prior biological or neuroscientific
knowledge that this still works that
well but that I think shows how
expressive the topological
representations can be for these tests
of course age prediction is not
something that you want to do in
practice so you could do some a lot more
thing things one thing that we tried and
this is still ongoing work because that
is really really complex and there's no
pun intended with a complexity analysis
we try to do complexity analysis based
on the actual brain states that
participants went through it turns out
that the data are very noisy and we had
to aggregate these um these complexities
by the cohort so we had to aggregate by
by the years and what we were looking
for is to what extend a younger cohort
is exhibiting higher topological uh
sorry lower topological complexity than
an older cohort with the idea being that
if you're a very young child and you're
watching this movie then you probably
don't understand a lot of what is going
on but you see that there's lots of
sounds and noise and fury signifying
nothing and if you're an adult you
probably have a an emotional context and
then and another context that you can
that you can relate this to now again
stressing this this is ongoing work for
instance one thing that we're looking at
nowadays is we're trying to relate such
trajectories also with actual events and
an emotional con text in the stimulus
but this is of course hard to do so
there's all kinds of interesting
annotated data where we are where people
are tracking facial expressions or where
they are asking people in this movie
scene
which emotion are you predominantly
experiencing anger shame disgust Joy
surprise
um apathy whatever
um actually there's more negative there
in there than positive ones but that's
not my that's not my key so I'm not a
responsible for this but so so this
would be one of the steps where we want
to take this next and then we would
characterize the actual shape of the of
our brain State trajectory as we um as
we watch or as we encounter a stimulus
now with this macroscopic uh
considerations of a brain Let Me Maybe
now zoom in quite a lot and talk briefly
about the prediction of the shape of
cells so now we have a we have a quite
different task so now we're actually
having something where we can measure
the outcome quite considerably so this
is relatively recent work
um and also still ongoing because it's a
complicated problem we'll see why this
is the case so my collaborators from
helmhos they have a lot of nice cell
images and those are they say it's
images of single cells but in case
you're also as confused as I am this has
nothing to do with single cell data
analysis a single cell or scnr C
analysis this is something completely
different so what they mean is really
like they take pictures of individual
cells under the microscope that's great
they use the confocal fluorescence
microscope for this and then they're
interested in predicting the 3D shape of
a cell from this 2D image this is also
known as a morphological analysis and
it's it's a crucial way to detect
certain pathologies so one very similar
paper in this in this area is paper by
Ford on red blood cell morphology and
this states that when used properly RBC
so red blood cell morphology can be a
key tool for laboratory hematology
professionals to recommend appropriate
clinical and laboratory follow-up and to
select the best tests for definitive
diagnosis so in some sense and it's
maybe maybe some of you already
recognized this in some sense this is
what a company called theranostrite to
do
um we're not we're not claiming that we
that we can even do do five percent of
what they claim to do because it turns
out that they were a hoax company so bad
for them potentially good for us but the
the goal would really be if this works
would really be that we take some blood
sample we try to reconstruct this from
Individual microscopy images and then we
know something about the patient's state
of health and I'm stressing this because
and I hope no one is queasy to queasy in
the audience here
um blood is a really nice substance in
that it's almost always available in
patients so unless a patient is really
really sick you can probably spare at
least a drop of blood in the hospital so
it's really a substance that you can
easily get and you can easily analyze it
so drawing inference making inferences
from small quantities of human blood is
a very nice well technology to to have
in the future
um spoiler alert we are not quite there
yet but we're making some some progress
and I'm going to show you what we were
able to achieve with topology so let's
first start without topology namely we
take a look at the pipeline that we had
before adding some topological
information here what we do is we start
with a 2d input on the left hand side we
throw a machine learning model on there
that's by the way that is dotted because
it's of course something that you can
easily replace if you find something
better we let the model predict a 3D
shape and then we use a geometrical loss
term more about that in a minute and
compared with the ground truth and
the mathematicians or computer
scientists in the audience they might
appreciate this this is really hard and
it's really hard because it's a
complicated inverse problem so we're
going from 2D to 3D I mean you already
know that if I look at my shadow then I
can reconstruct all kinds of interesting
things so it is also essentially an
ill-defined problem with a large number
of potential Solutions so we do need a
lot of input data to to make this work
semi-reliably and I can already tell you
that there will of course be cases
depending on how you look at the cell
where this reconstruction can never work
because you're just missing features but
we are content with capturing 80 of the
cases maybe quite well and then raising
a flag for the case that we can't handle
either so that would already be a nice
result there now what we had here is the
the so-called shaper the shape
reconstruction or reconstructing Auto
encoder and this is a very
simple machine learning technique that
employs a convolutional neural network
with some fully connected neural
networks and it as you can see it it
kind of decreases the picture first and
then it blows it up again into a volume
so it starts with a 64 by 64 image and
then it gives you out a 64 cubed uh
voxel volume and what this essentially
does under the hood is it is learning a
likelihood function a likelihood
function is a function that Maps every
voxel of this grid here so every point
in R3 to some scalar value and this
scalar value indicates the likelihood of
a specific voxel being part of the true
volume so that is what the what this
method does and how it tries to
reconstruct the images now
for the normal loss function or for the
geometry based loss function the shaper
method uses a geometry based loss that
consists of two components one is the
dice lost the other one is a binary
cross entropy loss and without going
into the details here let me just give
you the intuition here so essentially it
compares the geometry of the resulting
volumes on a per voxel basis so what it
does is it's looking for whether the
reconstructed volume is well aligned
with the ground truth one but there's
one issue at least namely on their own
these issue these losses are not
sufficient to capture shape variation
because if I modify the shape a little
bit then its topological characteristics
of course don't change so I can rotate
my icosahedron or my platonic solid in
space all I want it's still a platonic
solid but of course these losses that
are very restricted to the voxels
themselves they will then raise an alarm
and will say no no this this deviates so
you're learning a very restricted set of
shapes or shape features less so than
you could do in practice but now
let's add topology to the mix and let's
hope that this helps solve some problems
this is what we did in our recent mikai
paper so essentially if you recall the
previous slide what we what we added is
these three components below so we use
topological features calculating of both
the 3D prediction and the 3D ground
truth and then we have a topology-based
loss that we can combine with the
geometry based loss above to obtain a
joint loss and to kind of balance out
the geometry-based Reconstruction and
the topology-based reconstruction
and to go into some more details about
this loss it's something that you have
encountered before in these slides it's
the sum of wasserstein distances between
the persistence diagrams and a term that
I haven't introduced before but that I
would just now call here total
topological variation sometimes it's
also known as total persistence now
these terms have two different
um components or two different uh result
if I can call it that the first is you
align the ground through the likelihood
function and the predicted likelihood
function f Prime so you want your ground
truth and your predicted function to be
as close as possible in the topological
sense mind you the second term so this
total variation term or this total
topological variation term is only
applied of course to the predicted
likelihood function because we can only
change that prediction right we cannot
change the ground truth data and it's
added there to reduce the geometrical
topological variation of the predicted
likelihood function so essentially and
you can you can try this out I have a
website for this if you want to check it
out and this reduces the wriggles in the
surface that you get because if you have
a nice surface if you have a nice
dodecahedron one of the issues with with
the topological features is that I can
add a lot of Wiggles around this surface
and this won't change the overall
topology but it's something that is
really really undesirable and so adding
this additional persistence term in the
shown in blue here gets rid of this and
then of course we can combine them and
this is this is what you're all familiar
with and what you all know and we can
choose the Lambda parameter to oh what
was that I hope
okay okay great Cricket okay so nothing
I didn't destroy anything that's that's
good so and then we obtain a combined
loss by um by adding those together and
by by weighting them accordingly so now
two interesting
um uh things here so namely we of course
write out what happens if you only go
for the topology based loss then
everything explodes again because
topology on its own is not powerful
enough to regularize your shape nicely
but geometry on its own is also not
powerful enough so there you have it
this is really something that needs to
be optimized jointly on the other hand
one interesting tidbit that we found in
this paper is that it's actually
sufficient and this is one of the
moments where well well in Heights that
you're all you're always smarter right
but in hindsight it was clear to me why
this should be the case so I found a
theorem from some topology book that
explained this a little bit but um it
turns out that we don't actually need
the sum over these faster shine
distances but it's sufficient to do one
of our such line distance in a dimension
two specifically because cells have kind
of nice features and there's a duality
in topological persistence going on that
we that we can expire for these purposes
I'm showing you the the nice version
though or the the big beefed up version
here though because that's what we
initially trained with and then the rest
is more like an empirical result that
might not hold for generic data sets now
let us briefly look at the results here
and I I don't want to delve into the
details what these error metrics mean
but essentially what we found is that
just by adding these simple calculation
the simple loss term here we're able to
reduce errors in all relevant metrics
quite substantially except for of course
um the odd one out there's always one
experiment where this doesn't work
um namely in the surface roughness for
the nuclear data set
um the results are still very very close
and I think this is more like an
initialization question so if we run
this multiple times it might be that we
that we end up having something
something that is relatively close here
one thing I do want to stress here
though and this is one of the reasons
why I'm excited about this project is
that typically you might run into
scalability issues if you have very high
dimensional topological features that's
not something that I that I mentioned
too often here because it didn't appear
in in these cases but for for this type
of data set we are actually we have
almost no performance
decreases is whatsoever because it turns
out that the topological features are by
themselves expressive enough to be
calculated on a very very simplistic
rough version of the data so we don't
have to use the 64 cubed volume data to
calculate this part of the loss term but
we can use a down sampled one and the
down sampling is is almost almost free
in terms of in terms of computational
performance that's really nice because
it really ties this together and also
demonstrates those features bring in
complementary
perspectives now
uh that's all I have for you today so I
hope I was able to convince you a little
bit about the fact that topology can
provide useful inductive biases for
shape reconstruction tasks in particular
um I do want to stress that another
takeaway so I already have two so don't
forget about Euler being the man then
persistence diagrams those are the
topological descriptors and the third
one is that they actually encode
geometrical and topological properties
of the data so don't get fooled by our
bad advertising mathematicians are
really bad at advertising and naming
things Carson knows this from from our
work together I really am bad at this
um and so we we name it computation
topology but it's actually actually more
it's also it also encodes some
geometrical properties not all of them
but some and moreover and this is the
fact that I'm most exciting about
excited about the integration into
standard machine learning methods is now
possible so we have something for Auto
encoders with something for graphs if
something for shape reconstruction tasks
more is hopefully to come I mean knock
on wood right if you want want to learn
more there's a recent survey that I that
I co-authored with Felix Hensel and
Michael Moore also of Carson's lab now I
think a postdoc in Stanford it's called
a survey of topological machine learning
methods and it's um it's an Open Access
publication in Frontiers and artificial
intelligence and last but not least this
is um this is now the advertisement
party I hope it's okay
um if you're interested in topological
machine learning and you want to check
this out on your own
um my my lab and I we're trying to make
the software work it's called python
topological I know very creative name
use you see I'm kind of like true true
to form here it's not a creative name
but it works and this gives you the
power of topology at your fingertips and
it at least at the moment of me saying
this I mean depending on how fast the
others are with the pull request we can
do the fmri data analysis we can do the
shape reconstruction stuff what we can't
do yet and maybe someone in the audience
wants to do that is we can't do the the
graph neural networks yet but that's
just a matter of time until we have a
rest cute our old code and put it into
into a nicer framework but it can
already do quite a quite a few things
and I'm happy to
um discuss more and and maybe ask answer
any questions about the software now of
course I'm very happy that that I could
be here and I'm looking forward to some
of your questions now thank you very
much
[Applause]
thank you very much Bastian
both for jumping in for this inspiring
talk thank you very much so are there
questions for Bastian
yeah I mean
um so at a certain point I missed the
step I think so you kind of lost me
um no I'm sorry because maybe you're
totally in that field
you started out explaining that you
measure the shape of this red blood
cells and you mentioned that your
colleagues have
a confocal
and somehow the next slide was how you
want to go from a two-dimensional to a
three-dimensional reconstruction but if
you have a confocal you have the
three-dimensional yes so at that point I
lost the connection very very good point
I'm I'm sorry this this is this is also
my my lack of biology talking there so
the way I understood their problem is
that they say
um these uh doing 3D reconstructions
here fast and doing a lot of those is
time consuming for them and doesn't
two Dimensions yes
exactly yeah yeah so so they they can do
it in 3D I mean in fact
um the the data that we got here and and
this is actually I want to stress this
because this is an actual um ground
truth data that that we got so I didn't
make this up or anything this is really
um one of the ground truths from one of
the predictions that we have from our
algorithm they they are being done by by
um by our collaborators um but they are
telling me that this is a is a time
consuming process and it doesn't doesn't
scale very well so in the meantime what
they are looking for is something that
gives them that gets them eighty percent
off the way there and that maybe can
tell them oh this is a blood cell that
looks very very anomalous from the 2D
slides I'm sorry I should make this more
clear but this is a very very good point
thank you very much
foreign
so thank you it was incredibly
fascinating I have a curiosity so uh so
you showed quite a few results for
images and 3D space yes how far can you
push dimensionality in this setting or
to put it another way how badly is these
topological setting affected by the
curves of dimensionality yes Ah that's a
that's a that's a very good question so
now okay I don't I don't wanna I don't
want to do a politician's dance and give
you a half answer so
um let me just let me disentangle this
though so first of all
um it is not affected as much by the
curse of dimensionality as other methods
because
um fundamentally it's it's built upon
the idea of having good distances in
your data and of course yeah euclidean
distance suffers from this but you can
throw your own distance in there you can
even throw learn distances or metrics or
Mahalo Nobis distance or whatever you
want in there so in that sense this can
be mitigated but where the curse of
dimensionality re-hits us is when we go
back let's hope that this works this is
the first time I know one the second
time actually I'm giving a talk since uh
2020 in in real life again so where this
curse of Dimension and he really hits us
is if you look at this V Epsilon
expression on the bottom of the slide
you have all these subsets that are
within Epsilon or or less to each other
and if you
um if you don't restrict the size of
your subset in the worst case you get 2
to the power of the number of your
points of your subset so it is a very
very bad scaling in this sense and this
is where the curse of dimensionality
hits us
um this again can be mitigated by saying
okay we're only interested in
topological features of a certain
Dimension so for instance we can say
that empirically speaking for graphs
it's sufficient to do 0d and one DV just
to already get a nice performance
improvements for higher Dimensions 2D
might still be feasible I personally
haven't encountered a data set where
really much more Beyond 2D was required
I mean you can build data sets where you
can only characterize your your data by
by having higher order Dimensions but
it's it's rare and now to give you a
very precise answer so this scalability
is abysmal you have to you have to cheat
yourself a little bit around this it it
is possible but it's it's hard to do
um but at least if you have a 1 000
dimensional Point cloud and you're only
interested in a bunch of low dimensional
topological features which can still be
very expressive mind you then it's still
something that can be that can be
applied
I hope this was good thank you
you're welcome
Sean Philippe um
[Music]
yeah thanks a very great talk uh and my
question was in terms of application so
you you mentioned a few of them but I'm
sure there are plenty of others so uh
what about a structural bioinformatics
so stretches of molecules either small
molecules of proteins and in particular
since you mentioned that you can back
propagate it sounds like things like you
know Alpha fold Etc which predict the
structure using some loss functions
potentially could also use some some
real stuff look at that absolutely no
not yet to be honest so I would I would
very much love to so in fact I think
that
um our success in this topological graph
new networks kind of gave us the
motivation to dig a little bit deeper
into this realm
um there I I do think that that this is
one of the application areas where we
would require a joint optimization or
kind of a joint view on the data because
I don't think that the topological
features on their own are any more
powerful than what is already out there
but potentially if we phrase this right
and if we set up the task right then we
could have something that is really
complementary because it can capture
things that you cannot capture in in
other ways so yeah I would definitely be
interested in this there's also some I
only mentioned this briefly I think let
me go back to this yeah it is actually
the right slide so no it's actually not
uh there is work with Leslie but okay
Leslie is on the slide that's good
there's other work with Leslie where
we're looking at evaluating genitive
models for for graphs and here I think a
topological perspective would also be
very much warranted because graphs are
already topological objects so it would
be very interesting to kind of
characterize the expressivity of a
generator of a distribution in terms of
its topological properties so so to say
can we get all the modes that live in
this space or are we restricted to a
certain class of graphs but I I also
have to say that this is living a little
bit in future worlds for now so ongoing
or rather rather planned as research but
definitely interested in I think
definitely worthwhile
and can I have a second second question
or someone else yeah so super technical
question uh when you talk about the
differentiability yeah of uh the
topological representation with respect
to the input data uh can you say a bit
more about so I suspect the function is
not differentiable but maybe everywhere
almost every different channels the
question is you know yes since the
points appear there's area where there
is no point and suddenly it moves away
from the diagonal exactly so so one
thing that we exploit here maybe just to
go back to this to this diagram here so
one of the there's multiple ways of of
going around this some of my colleagues
for instance
um uh
um El Canon Solomon has been has been
looking into this um and Matthew career
as well
um uh here for this diagram here we
would exploit the fact that
um every Point has at least a very very
small neighborhood around it around
itself that doesn't contain any other
points so there's like no overlapping
points in this diagram if we have this
condition then we can show that
um that the mapping from the space to
the diagram is constant and then the the
the the the composition rule of
gradients tells us that we can that we
can ignore this part there's also more
technical results if you use different
representations because a lot of things
exist in this space that I haven't told
you about unfortunately here in this
talk apologies for this if you use a
different representation of your
topological features and then you can
show for instance that the mapping is is
lip shits and then you can allude to the
now now I'm blanking on the name of
theorem but then you can allude to a
theorem that tells you that the mapping
is on differentiable almost everywhere
and then you need some computational
tricks to make this gradient actually
unique in in practice but it can be done
and it works even even kind of out of
the box with pytorch it works
surprisingly surprisingly well even if
you kind of ignore some some of the
degenerate cases it would work
surprisingly well here
thank you you're very welcome
so let's thank Bastian again for for
this very inspiring keynote thank you
Bastian
[Applause]