Emmanuel Kowalski: Wasserstein metrics and equidistribution (NTWS 292)
Watch on YouTubeVideo summary
The video presents a joint mathematical investigation into the quantitative aspects of equidistribution, moving beyond classical definitions to explore more robust metrics for measuring convergence. Traditionally, equidistribution is defined by the convergence of probability measures on a locally compact space, often illustrated by examples like Weyl's theorem on multiples of irrational numbers or Dirichlet's theorem on primes in arithmetic progressions. However, the speaker argues that classical approaches, such as the Erdős-Turan inequality, have limitations when dealing with higher-dimensional spaces or situations requiring invariance under continuous transformations. To address these issues, the presentation introduces Wasserstein metrics, also known as Kantorovich-Rubinstein metrics, which provide a powerful and flexible framework for quantifying the distance between probability measures. These metrics are particularly advantageous because they possess strong invariance properties under Lipschitz functions, allowing for consistent quantitative statements even when the underlying space is transformed, unlike traditional discrepancy methods that struggle with such changes.
The core of the talk focuses on defining and utilizing the Wasserstein metric within the context of optimal transport theory to establish rigorous bounds for equidistribution. The speaker defines the Wasserstein distance based on an infimum over couplings of measures on a product space, highlighting its ability to metrize convergence in law when the underlying space is compact. A key property emphasized is the duality theorem for the case where the parameter $p=1$, which interprets the metric as a supremum over integrals of one-Lipschitz functions. This functional interpretation simplifies proofs and leverages the geometry of the space effectively. Furthermore, for specific settings involving compact connected Lie groups, such as tori or matrix groups, the presentation details an inequality analogous to Erdős-Turan but formulated using $L^2$ averages of Fourier coefficients associated with irreducible representations. This approach allows researchers to bound the Wasserstein distance directly, thereby capturing the rate of equidistribution in a way that respects the intrinsic symmetries of the group structure.
The theoretical framework is then applied to significant problems in analytic number theory, specifically concerning hyper-Kloosterman sums and Deligne's equidistribution theorem over finite fields. By leveraging the Riemann hypothesis for these sums, the speaker demonstrates how Wasserstein metrics can yield effective quantitative bounds on the distribution of conjugacy classes of unitary or symplectic matrices. A major application discussed is the "shrinking target problem," where one seeks to determine how many elements fall into a set that becomes progressively smaller as the prime modulus grows. Using the established Wasserstein bounds and the Lipschitz nature of the trace map, the presentation derives precise lower bounds for the number of such elements, confirming not just their existence but also that they appear in the expected proportion relative to the limiting measure.
Finally, the talk underscores the profound arithmetic implications of these quantitative equidistribution results, extending beyond simple convergence statements. The speaker illustrates that obtaining a bound on the Wasserstein distance is equivalent to establishing a zero-free strip for associated $L$-functions over finite fields, even without assuming the full strength of the Riemann hypothesis. This connection reveals that Wasserstein metrics encapsulate deep arithmetic information that is crucial for understanding the distribution of exponential sums. The work concludes by noting recent extensions to non-connected groups and other values of $p$, suggesting that this metric-based approach offers a versatile and powerful tool for future research in number theory, capable of addressing complex equidistribution phenomena where traditional methods fall short.
Read the full video transcript
This is joint work with Teo Antro. And
as the title says, so it's about
equidistribution.
And especially quantitative aspects. So,
let me first remember recall the
definition
of equidistribution
in the setting I'm going to work on.
So, very abstractly equidistribution is
some version of convergence of measures
of probability measures on space. And
the version I'm going to look at will be
when you have a
locally compact
topological space, very often compact,
but it could be something like the
quotient of the upper half plane by a
discrete group which might have finite
volume and not be compact.
We have some probability measure
mu on X which is kind of a reference
measure to which other measures would
like to converge.
So, mu n will be a sequence of measures
of probability measures again.
Everything being happening on X.
And the example to to keep in mind for
for most of what I'm going to talk
about,
the standard example is when you have a
finite subset
non-empty
in capital X and then you look at the
measure which is just the average
of delta masses
over this finite subset.
That's the standard example and then
we say that mu n
becomes or is mu equidistributed
as n goes to infinity.
Uh
by definition means that
you can compute the integral
of suitable functions
with respect to mu
on the on the
set capital X as limits of the
corresponding integrals
with respect to the measures mu n.
Okay, so and in the example this means
the limit of the average of the test
function f at the points in the finite
subset
that are varying.
And of course this statement for an
arbitrary function f is way way too
strong to ask in general
and so one imposes restrictions on the
test functions and
classical definition for
equidistribution is
to assume that this works for f
continuous and bounded.
So if if x is compact this is just f
continuous if f is not necessarily
compact then we also uh assume that it's
bounded.
Okay? So this is a classical definition
uh
let me give one or two examples
uh for those who might not
uh be very familiar with it.
Uh one of the most basic ones
uh goes back to Hermann Weyl uh in the
early
20th century is uh taking
the space X to be the circle or R mod Z.
Uh
the measure mu to be the Lebesgue
measure.
And for instance the xn so
Hermann Weyl proved many statements of
equidistribution of this kind but the
most basic one would be when you take uh
the measures mu n to be defined as
before by averaging over a finite set
and the set xn would be the set of
uh multiples up to n alpha
for some element alpha which is in
uh R minus Q.
Okay, and of course alpha, 2 alpha, and
so on have to be taken modulo Z.
And then it's a fact that
mu n
or xn one could say is indeed
uh equidistributed according to
Lebesgue measure when n goes to
infinity.
And it's a prototype of many statements
of equidistribution.
But just to give um so this is about
I think 1917.
And just to give another in fact uh
even earlier example even if it's was
not interpreted in this way
at the time.
So if you think of uh Dirichlet's
theorem on primes in arithmetic
progressions, it can be interpreted as
equidistribution.
In this case uh the set X is just
finite.
It's the invertible classes modulo some
fixed modulus Q.
Uh the set xn again so the measure mu is
the uniform
probability measure on this finite set.
And uh xn could be the set of uh
primes.
So you look at the uh set of
residue classes
uh
P modulo Q.
Let's say for primes P
uh up to up to n.
And you need
P to be coprime with Q if you want to
literally for this to be inside capital
X.
Okay, and then Dirichlet's theorem says
that then there is equidistribution.
Meaning that every
congruence class modulo Q can be
represented
equally often,
roughly speaking, by primes reduced
modulo Q.
Okay. And this example is a good example
to see that it's very important very
quickly to not only be able to say that
equidistribution holds
in the sense that these various limits
that I've stated here exist and are
equal to what they should be, but you
also want to have more quantitative
information. And that's for anybody who
has studied any multiplicative number
theory one knows very well that
Dirichlet's theorem on primes in
arithmetic progression is just the
beginning and one really one needs to
have error terms and quantitative
information.
Need to have
quantitative
statements
in many applications.
So I'll give concrete example of such
applications in quantitative terms, but
for the moment
what do we mean by quantitative?
So I'm going to stick to to settings
which are closer to the
results like Armand Weil's than to
Dirichlet's theorem. And so one of the
easiest way is to say well we we
restrict the set of test functions and
we try to say that the left hand side of
the equality, the integral with respect
to the limit measure, is equal to the
nth term on the right hand side
plus some explicit error term.
And
if you do this,
there are many cases where you can do
that, but it is not necessarily what's
needed in applications and
for instance
maybe what most people most classically
associate to the notion of quantitative
equidistribution
is the Erdos-Turan
inequality.
So this is the maybe the first example
which has been studied
for a long time. So the the first
version of this
by Erdos and Turan is from 1948 to give
an idea of the time scale. So this
concerned the case
which is the one considered by Weyl. So
the space is the circle
the limit measure is the Lebesgue
measure
but it can apply in principle to any
sequence of measures which can
approximate
approximate mu
and it quantifies equidistribution
by comparing the measure of intervals
for the Lebesgue measure so the length
of an interval
with the corresponding
measure for mu n. So in the case
of the example
so this would be counting
how many elements in your set
belong to the interval.
So
how many elements in xn belong
to a given interval.
So equidistribution in the case of the
Lebesgue measure means that this goes to
zero
when n goes to infinity and the way it's
quantified by Erdos and Turan is to take
the supremum of all the intervals.
And the interval you can think of closed
intervals or open intervals it doesn't
matter.
And they find a quantitative upper bound
for this
uh, depends on
what are called the
Weyl sums for equidistribution. So, the
upper bound is of the following form.
So, it's bounded up to a multiple
constant which can be made explicit
by 1 over T. So, T is going to be some
parameter
which can be a real number that can be
chosen
arbitrarily
and usually is chosen to give the
optimal possible bound.
Okay, so maybe I should make clear that
when I write this less less symbol, it's
the
uh,
it's the Vinogradov symbol, so it's
means that it's bounded by a certain
constant times whatever follows. So, in
this case, so 1 over T
and then a sum over integers H
in Z
with absolute value
non-zero
so, H is non-zero but bounded by this
capital T.
And then I have something which decays
with H, 1 over the modulus of H times
the modulus or the absolute value of
this Weyl sums
which I'll call WH of μn in general
uh, which are
in terms of integral, this would be the
integral of the function exponential
2iπnx
uh, 2iHx
with respect to μn.
And again, in the case of the example
that is just averaging
the additive character
exponential 2iπHx
over the finite set Xn.
Okay.
So, if one looks at this inequality,
uh, then it reveals quickly that, uh, to
have equidistribution
uh, is equivalent to having this WH of
μn go to zero. This is an old fact that
was established by Armand Weil uh in the
same paper that I mentioned earlier.
And uh this shows that if you know good
bounds for these so-called Weil sums,
then you will be able to provide an
estimate for the supremum of this uh
difference between the Lebesgue measure
of an interval and how many elements in
the finite set, for instance, happen to
be in that interval. So, this is
something that has been used uh many
many many many times, and it's uh quite
convenient in its way.
And uh what I want to discuss is uh
problems where this is not actually a
such a satisfactory way of uh
quantifying equidistribution.
So, uh issues with this
Erdos-Turan inequality.
Okay, so for certain purposes, it's
perfectly fine,
uh but there are others where uh it
might not be what what one wants.
Um
So, I'm going to highlight
two of them. So, one is
uh higher-dimensional versions,
uh which feel less intrinsic, at least
to me.
So, I'm not going to state any uh
of these higher-dimensional versions.
There are many classical ones.
So, the thing is usually there what
happens is you replace the intervals by
a certain choice of types of subsets,
and uh if you look at a
higher-dimensional version,
meaning in R mod Z to some power,
uh
it's not necessarily clear why, for
instance, rectangles should be better
than other subsets, depending on what
you want to to use this for.
Um
And in in our case, in this in this work
I had with uh Thai Van Vu,
the issue started there because we had
equidistribution not in R mod Z to the
power D
of certain limit measures but we had the
measure on a on a sub torus
and we could not there was no canonical
coordinates on that uh
on that subgroup and so it was not
necessarily very natural to choose any
choice of coordinates to define the
rectangles that you would use for a
classical
uh I don't remember the Erdos-Turan
inequality.
So that's one potential issue.
And the other one which is somehow
related is
what I will call lack of invariance
under various transformations.
And what I mean here is that suppose
uh we're in the situation of the
Erdos-Kac theorem
and you have some kind of other function
G
uh from R mod Z to R mod Z
uh which is continuous.
Then uh if mu n converges to the
Lebesgue measure mu
then it's a completely formal fact just
looking at the definition that if you
push forward
uh mu n by the function G it converges
to the push forward
of the measure by the function G.
Uh
so that's completely formal at the level
of equidistribution
but the problem is you also might want
to have such a phenomenon or principle
uh of transformation that also applies
for quantitative statements and if you
measure the convergence of mu n to mu
using the discrepancy in the Erdos-Turan
inequality
and you apply uh some continuous
function well there will be first the
problem that the limit measure G lower
star of mu is
not equal to mu in general.
And And the Turin inequality is
specifically written for the Lebesgue
measure.
And even if you took
a function G so that this is again the
Lebesgue measure, it's not
straightforward just looking at the
the statement of the
Turin inequality what's going to happen
for the discrepancy of the image under
the function G of the mu ends.
Okay.
So, this is what we
we were looking
in in in our application with Theo, what
we were looking for was a more abstract
or maybe a more invariant way of
measuring quantitative equidistribution.
And it turns out there's a way of doing
this which is uh
in some sense was staring us
very obviously
so uh
in the face because it's an extremely
well-known
uh
way of measuring the distance between
probability measure. It's this uh
Wasserstein metrics.
Okay. So, they're also known as
Kantorovich metrics or um
uh
Kantorovich-Rubinstein.
There's There's various names, but the
one that seems to stick the most is is
Wasserstein metrics.
Uh
so, it's a very important concept
in uh analysis.
Uh especially in optimal transport.
But it's also very often used in modern
probability,
statistics,
uh actually many many many many fields.
Um
And as I said, it's it's a way of
giving a distance on the space of
probability measures
uh on on suitable spaces. So, I'm going
to give the definition. As I said, this
if you've ever seen a colloquium in
analysis in the last 10 to 15 years,
there's a good chance that optimal
transport was mentioned, and therefore
there's a good chance you actually saw
this definition. Uh because it's really
become one of the
uh most uh
exploited tools uh in many many parts of
analysis.
So, I'm going to
again work in less generality than and
then these metrics are defined.
Uh so, one small difference with the
setting I was before is now we have to
we need to have a metric space.
For applications, this is usually not an
issue, but it will depend on this
distance D.
Uh the metric space has to be complete.
But it need not be locally compact.
There's There's many cases where
infinite-dimensional spaces are used in
this setting.
Um so, we're going to define distances
between uh probability measures
I'll call them just mu1 and mu2.
Probability measures on X.
Uh these will depend on a parameter P.
So, it's a real number.
Okay, so at some point later P will
become a prime number as it should, but
for the moment
this is a terminology that's also so
standard that using anything else would
be a bit problematic.
Um so, we have a parameter which is this
real number, and one defines the
Wasserstein P metric between mu1 and mu2
in the following way. So, it's an
infimum
over measures
on the product in a set that's called
capital pi of mu1 mu2 usually,
which is defined as the set of
probability measures.
So, I'll call them pi on the product of
X with itself,
such that if you project with the first
on the first coordinate
you obtain mu one.
And if you project on the second
coordinate
you obtain mu two. Okay? So, for
instance, the product measure of mu one
with mu two will have this property. So,
this set is always non-empty.
And once you have this, then you
integrate
over the product
the distance between two points to the
power P
with respect to this measure pi
and you take the power one over P. And
then you take the infimum over this.
And in this generality, this infimum is
actually achieved.
I'm not going to use that fact, but it's
something that's worth knowing.
Okay.
So, it's a definition
that one finds now in books
dealing with optimal transport, for
instance, and so on.
And
it turns out to have many, many good
properties. So, the first, the basic one
I want to
mention.
So, all together, there'll be a list of,
I think, four properties and they they
really show that this gives a good
quantification
of the notion of equidistribution, of
convergence of measures.
So, the first one is that it does
metrise a convergence in law or
equidistribution, at least
if X is compact.
So, in the more general case, you would
need to have some finiteness condition
on the on the measure, but in that case
to say that mu n converges to mu in the
sense of equidistribution
is equivalent to saying that the
distance between mu n and mu goes to
zero
as n goes to infinity.
So, any quantitative bound that goes to
zero for the Wasserstein distance would
give some kind of quantification of
equidistribution. And this holds for
every fixed P. So this does not depend
on the choice of P.
Uh that's the first
property.
The second one is
it has these strong invariance
properties, which is what I was saying
the Erdős-Kac inequality does not really
have.
Uh so it doesn't quite work with an
arbitrary um
function that's continuous uh between a
space and another, but it works for
Lipschitz functions.
So suppose you have a Lipschitz function
from X to not necessarily the same
space, let's call it Y.
And let's say it's Lipschitz with
constant C, so it does not distort the
metric by more than a factor C.
Okay, so Y is another metric space.
Then uh it is completely formal that uh
again for any P uh computing the
distance
after taking the push forward
by uh the function G uh cannot increase
the Wasserstein distance by more than
the factor C.
Okay.
So that means any kind of quantitative
statement you you might have to measure
equidistribution in Wasserstein metric
anytime you apply any Lipschitz
function, this will translate into uh
essentially uh equivalence equi-
equidistribution quantitative
equidistribution on the target space Y.
So what you're losing is this Lipschitz
constant. The inequality being with no
implied or mysterious things means that
it's also actually useful maybe when C
depends on other parameters, and it will
give you something uniform uh without
having to do anything. Okay? And this
invariance property is, as I said,
completely straightforward from the
definition.
There's essentially nothing to prove.
Which for me is a good sign that this
can be useful for for the kind of things
we want to do.
Um
Then the third uh
is
uh
specific to the case P equals 1. There's
there's a version
which applies to other values of P but
is
uh less straightforward. So when P
equals 1, this is something called the
Kantorovich
or Kantorovich actually, I think.
uh Rubinstein duality.
And it's a very nice uh
functional interpretation
of W1.
So W1 mu1 mu2, the distance between
the measure mu1 and the measure mu2 if
you use P equals 1, it's the same. So
the
as a supremum, so it's a kind of a
duality between min and max.
But now you take the supremum over
uh
continuous functions which are
one-Lipschitz.
And then you compare the integral of F
from mu1 and mu2.
Take the modulus and take the supremum
over all one-Lipschitz function.
So it's a kind of a functional uh
interpretation
of the W1 metric
uh which is obviously extremely useful
and and extremely nice and pretty. So
when P is larger, there's a similar
statement but the set of test functions
becomes
uh much more intricate.
Um
One might say that this suggests to
actually define
W1 by the right-hand side instead of
this uh play with measures
uh projecting to mu1 and one and mu two
uh and it's true that for what I'm going
to say today you could do it this way.
You could take the right-hand side as
the definition.
Uh and forget about all other
Wasserstein metrics, but in the general
theory it seems really that actually P
equals one like an L1 space is is
pathological in some respects.
And it's actually better to
uh to use
uh other values of P like P equals two
in particular.
Okay.
And the last one
uh the last statement
is again uh
even more restricted. So, I'm going to
take P equals one.
Uh
and the space X will be very special.
This will be a compact
uh connected
Lie group.
Okay. An
example is
this R mod Z
to the power of D that I discussed
before, but you should also think of
uh unitary matrices
or special unitary matrices
or special or unitary symplectic
matrices of some size.
Uh
and in that case when you take D to be a
suitable uh
uh
invariant metric.
So, in the case of the torus it's the
usual one in the case of SUD or SUSP2G,
these are more of strike metrics defined
in the theory of compact Lie group. So,
I should say pi invariant metric.
Um
then uh
there is a theorem
uh due to Bobda
in the general case
and uh
Bobkov and Ledoux
for tori.
So, for the first case.
Uh Uh and this inequality will look like
an Erdos-Turan inequality, but for W1 on
such a space
but very much similar to the Erdos-Turan
inequality in some sense.
So, it says that uh for mu1 and mu2
arbitrary measures, probability measures
on on this group
so I'll call it capital G cuz a group
should not be called capital X, I think.
So, let me get my statement
to be sure I get the right one.
So
we estimate
W1 of mu1 and mu2
it's going to be bounded by a constant
depending on the group
by something like again 1 over T where T
is a parameter
that can be optimized.
And then instead of a a sum of Weyl
sums, it's now going to be an L2 kind of
average. So, it's a sum over parameters
uh lambda with norm up to T. I'm going
to say what these are in a second.
Uh then there's the dimension of the
associated representation
divided by another quantity that I'll
define which is kappa lambda
and then
uh
the
Fourier coefficient
for
associated to lambda for mu1 minus the
one for mu2
squared.
Uh this is a Hilbert-Schmidt norm
and everything is
raised to the power 1/2 because this is
some kind of an L2
uh right-hand side. Okay, so I should
say what all these various things are.
Uh
so, the lambdas are the highest weights
parametrizing
uh irreducible
representations
of G. So, in the uh case
uh
of R mod Z to the D, which is already
very interesting,
then uh lambda is just an integer
in Z to the power D.
Um So, in general, these are elements in
the cone in a cone, sorry, uh contained
again in a lattice of some dimension uh
related to the to the Lie group.
Uh
then we have
So, the irreducible representation
associated to lambda is denoted by rho
lambda, and then it's the dimension of
the underlying space.
So, in the case uh of lambda in Z to the
D, then rho lambda is a one-dimensional
representation,
uh and it's defined by rho lambda of X
equals exponential 2i pi
the classical inner product of lambda
with X.
Um
And the dimension is just one.
Then the
kappa lambda is
so-called uh Casimir eigenvalue or
Laplace eigenvalue. It's essentially
the variant of the Laplace
uh
uh an eigenvalue of a variant of the
Laplace operator.
So, in the case of R mod Z to the D, the
norm of lambda is just the Euclidean
norm,
maybe up to a constant multiplicative
factor.
Uh
and the final thing is this
Fourier coefficient. I will define them
here.
So, the mu tilde for an arbitrary
probability measure mu on G, and for an
arbitrary lambda, uh, what you do is you
integrate
uh,
over the group with respect to the
probability uh, measure, uh, sorry, no,
with respect to the measure mu,
the image of
uh, rho lambda of an element, you take
the adjoint,
and you integrate, as I said, with
respect to mu. Okay?
So, this is a a linear map
on the space of rho lambda.
So, it's a linear transformation, which
is why here we have a Hilbert-Schmidt
norm
uh, to compute the the size of these
things.
And in the case of the torus, well, this
is just the usual Fourier coefficient.
It's a one-dimensional thing. It's just
the integral over R mod Z to the D
of exponential 2i pi
inner product of lambda with x with
respect to x.
Okay.
So, that's the
uh, this inequality of Bourgain, and I
should say that recently there's been
uh, a preprint
by Bourgain uh,
Gronin,
which uh, states a version of this for
other values than p equals 1. So, we
haven't yet incorporated that
into our application, but this should be
straightforward, and this should be uh,
very useful to obtain applications where
uh, you use WP where P is not equal to 1
uh, necessarily.
But for us, we started with p equals 1,
and this inequality,
as I said, is highly comparable with the
Erdos-Turan inequality, of course.
Uh, in the case of the torus, the uh,
quantities which appear are exactly the
same as in generalization of the
Erdos-Turan inequality, uh,
so the bounds you obtain are kind of
very similar to what one would get in
any application of the other student
inequality, but what you bound is not
the discrepancy, what you bound is this
Wasserstein metric.
And then you can exploit the invariance
properties of the Wasserstein metric to
to go towards applications. So, let me
now in the
remaining time discuss some application.
So, I'm going to present the one that we
have in our paper with still, but there
are others that have already actually
been done
by a
few people since our paper was on
archive,
and I'll mention
Cornelissen.
So,
Sief also
cuz this kind of
show how this really can give a
very uh
flexible uh
framework for equidistribution, or can
Bring a ling.
And there's another
by Peter Humphries, which
does the kind of Duke's theorem or
equidistribution on the
hyperbolic
on quotients on the hyperbolic plane
in terms
of Wasserstein metrics.
Okay. So, the the application we have is
a version of Deligne's equidistribution
theorem,
which is effective and quantitative.
Uh and here again, I should say this is
not the first time that people uh
quantify Deligne's theorem. It's very
natural to try to do this because
Deligne's equidistribution theorem is
proved by very strong bounds
coming from the Riemann hypothesis on
finite fields
for the underlying Weyl sums, and
therefore it in some sense was already
from the beginning a quantitative
equidistribution statement. So, here I
want to
refer to an old paper of
Fourier and Michel,
which is very much in the spirit of what
I'm going to describe, so from 2002,
which kind of did by hand some of the
things I'm going to describe, and more
recent work by
Fu,
Lau,
and Chi,
probably 2020,
2023 or 2024,
who have a fairly general version in
terms of some version of Deligne's
equidistribution theorem.
So, Deligne's equidistribution theorem
is a very very general
statement about existence of limiting
distribution for families of exponential
sums
over finite fields, and since I don't
want to assume the kind of background
material that's involved in a general
statement, I'm going to do a special
case,
which is already quite interesting and
important for applications.
So, this is going to be about
hyper-Kloosterman sums.
So, we fix an integer R.
Now, we're going to
from now on P will be a prime again,
and the the Vassiliev matrix will all be
in
W1,
then for a prime number P and an
invertible element modulo P,
one can define, following Kloosterman
and and Deligne,
the so-called hyper-Kloosterman sum,
KlR of A and P, which is some
normalizing factor,
1 over P to the power R minus 1 over 2,
and then you sum over
R elements
in FP,
restricted by the condition that the
product is equal to A,
and you sum the additive character
exponential 2 i pi over P applied to the
sum of these elements.
Okay?
So, when R is equal to two
then you have one over square root of P
and then you have X and Y where the
product is equal to A and if you
express, let's say that X2 is A over X1,
then you recognize this is a classical
Kloosterman sum.
In general, this is a generalization of
Kloosterman sum
which occurs naturally in many problems
of analytic number theory.
And uh
the distribution properties of this was
established uh so, by Deligne and Katz.
Uh so, let me
state it
in two steps.
Uh so, Deligne proved um
around 1974 that
uh
for every A and for every P, you can
find uh
matrix that I'll call theta of A and P
uh unitary matrix uh of size
R minus one
uh no, sorry, of size R
of C.
Uh actually, with determinant one, but
let's say just a unitary matrix of size
R such that
uh
such that the trace
of this matrix
is the hyper-Kloosterman sum.
Okay? So, this is already an extremely
strong fact because it immediately
implies, for instance, because the trace
of a unitary matrix is bounded by the
size
that the hyper-Kloosterman sums are
bounded by R.
So, So you take R equal two, you get
you get the Weil bound uh for the
classical Kloosterman sum. So this is a
special case of the Riemann hypothesis
over finite fields.
Uh
already quite strong. So RH over
finite fields.
Okay?
And
uh
it follows from the work of Deligne and
then from the the work of of Katz. So
Deligne more or less proved there has to
be some kind of equidistribution
property for these matrices, and Katz
actually determined precisely what the
equidistribution properties are, and
they can be phrased as follows. So if
you take uh
uh
Okay. So there exists uh a co-
uh uh
a Lie group, compact Lie group
G sub R
such that the
uh
if you take
for a given prime P
all the various matrices associated to
these
uh to this prime
uh and they depend on R, I should say.
Cuz the rank is important.
Uh so there exists a Lie group, which
also depends on the on the R, such that
these uh become
equidistributed.
Uh
And in fact, not in the group itself
because I didn't say it, but
this is the space of conjugacy classes.
Cuz these matrices are only really well
defined up to conjugation.
Uh but uh these matrices become
equidistributed in the space of
conjugacy classes
uh with limit, so the the measure mu in
the equidistribution
uh equal
to the uh
image
of the probability our measure.
So, they are uniformly distributed.
That's the way to think about this.
And moreover, so in some sense, this
statement was already known to Deligne
uh except for some subtleties. But what
Katz did, which is the
crucial to actually understand what the
statement means, is to uh
compute the group, and he showed that
the group is the space of unitary
matrices with determinant one
when R is odd,
and it's the space of uh unitary
symplectic matrices of size R
when R is even.
So, this is what the Katz called uh or
calls the the monodromy group for this
family of exponential sums.
Okay? And so, the uh the first
application uh that
or one example of the general form of
the equidistribution theorem that we
proved
uh with Theo for
uh
Wasserstein distances
uh is the following.
Um
so,
that uh the W1 metric
between uh I should give a name for this
probability our measure. Let me call it
mu sub uh R
small R.
Uh
so, the distance
to the limiting measure
associated to the hyper-Kloosterman sums
with the integer R, and the
sampled, so I'll write it explicitly,
the
average
of Dirac masses
associated to these conjugacy classes,
viewed as elements, everything taking
place in the space of conjugacy classes.
This is big O
of P to the minus one over the dimension
of this
monodromy group.
And this is the statement that we prove
in
in full generality for any family of
exponential sums where the monodromy
group is also connected.
It's also it's always going to be a
compact group. It's not always
connected. We're working on on
generalizing the statement to
non-connected groups which definitely
should be possible.
Okay.
So
in the few minutes
before end I want to
emphasize
two two things about this this result.
So one is a concrete application.
So what does this tell us about
hypergeometric sums that we might not
have known before?
So one typical application of
quantitative equidistribution is what
people call often shrinking target
problems.
Which means you're trying to compute how
many let's say of these conjugacy
classes
are in a certain set where the set is
not fixed independent of the prime in
this case, but actually becomes smaller
and smaller when the prime grows.
And here so what this gives us is the
following
theorem.
So I'm going to state it when R is odd
because the the statement takes slightly
different form depending on the
on the parity. I'm going to
take then a matrix G0 in SUr of C.
So that's the group GR for R odd
uh,
at least three different eigenvalues.
Okay, so in particular, uh, it cannot be
the identity.
Uh,
and then uh,
we are able using the Wasserstein metric
to get
uh, lower bound for the number of
uh, hyper Kloosterman sums
which are close to the trace of G0
uh,
which are close by a constant
so let's say one
divided by P to some exponent which
turns out to be three times R squared
minus one.
Exponent is not so important.
Okay, so you see
if we just were saying that we want this
to be less than some fixed epsilon
we would just want the Kloosterman sum
to be close to the trace of some element
in SU2 of C and equidistribution will
tell us that this happens uh,
at least once and with the right
proportion when P goes to infinity.
Now we have uh, a distance between the
hyper Kloosterman sum and the target
which is shrinking when P goes to
infinity
and we show that this is bounded from
below by for P large enough by uh, a
constant times P to the power one minus
two over
three times R squared minus one.
Okay, so we show existence but we also
show
uh, that there are in fact uh, many of
them and what is relevant here is that
this is the right proportion.
In the sense that this is the our
measure of the set of matrices where the
trace of G minus the trace of G0 is
bounded by this approximately up to
constant.
Okay? So, that's something that it's a
statement which does not mention
uh
Wasserstein metric or anything.
Uh
and the way it's proved is well, we
start with
the Wasserstein bound
uh that we that I stated before.
Uh this
is only provable in a kind of easy way
because of the Riemann hypothesis and
Bott as inequality.
But then the trace map from the
conjugacy classes to C is a Lipschitz
map.
And therefore, uh we obtained a
comparable quantitative bound for the
distance between the image of these
measures by the trace map. And then we
have to do some uh construction of test
function to deduce
uh this this type of things.
And I'll take one more minute
uh
with one last remark which I I find
interesting.
Uh
I'll I'll say it relatively imprecisely,
but uh one can show
that
assuming
a bound
uh
like W1
between the R measure
and this average
uh let's say suppose you assume that you
know a bound
P to the minus alpha.
So, I just showed that it can be proved,
but suppose we didn't know the Riemann
hypothesis, then we could not prove it.
But suppose then that you say, "Okay,
someone gives you a bound like this for
some alpha strictly positive."
Then from this you deduce uh zero-free
strip.
Not not the Riemann hypothesis, but a
zero-free strip for certain L functions
over the finite field.
Associated to the
uh to the hyper cluster sums.
Okay? So, I'll stop here, but for me
this is an interesting point because it
shows that somehow
uh bounds in Wasserstein metric, they do
contain
uh
some of the most relevant
uh arithmetic information uh that that
go even into the proof of
equidistribution. In this case, it's not
the Riemann hypothesis, but the
zero-free strip is already something
very very strong
uh and I find that intriguing. Yeah. So,
I'll stop here.