Seminar with Stéfan van der Walt "Scientific Python: Community, Tools, and Open Science"
Watch on YouTubeVideo summary
Stéfan van der Walt presents a historical overview of the Scientific Python ecosystem, tracing its evolution from its origins in 2004 through the formal launch of the Scientific Python project in December 2020. This narrative highlights key figures such as Fernando Pérez, Jared Milman, and John Hunter, alongside pivotal milestones like the unification of NumPy, the creation of Matplotlib, and the establishment of documentation standards. A critical turning point occurred around 2017 when the ecosystem, despite its organic growth, became siloed because small developer teams struggled to coordinate effectively as projects scaled. To address this fragmentation, the Scientific Python project was established with six key focus areas designed to streamline development and collaboration.
These foundational pillars include cross-cutting coordination documents known as Specs, which provide guidelines on version support, lazy loading, and deployment to reduce maintenance burdens, distinct from project-specific PEPs or NEPs. The initiative also introduced Summits for in-person working meetings to accelerate decision-making, a central Development Guide offering best practices and tools like repo-review for automated checks, and advanced Lecture Notes covering topics from optimization to sparse matrices. Furthermore, the project strengthened Community Coordination through platforms like Discourse and Discord, while consolidating Tools such as the SciPy theme, spin for documentation building, and analytics dashboards. Looking forward, the project aims to expand statistical offerings beyond econometrics, enhance support for domains like astronomy and Earth sciences, coordinate releases for stability, and host summer schools to foster broader engagement.
The talk emphasizes the importance of shared governance models and addresses the challenge of integrating a massive pool of potential contributors, noting that Berkeley Data 8 enrollment now represents half the campus. Future initiatives include establishing an Open Source Projects Office at Berkeley and collaborating with Stanford's OSO to support both local and distributed community growth. The discussion underscores how these projects serve as a vital bridge between the campus community and a broader distributed network, encouraging non-Berkeley guests to ask questions and connect. A specific example of this success is the Sky Portal, an astronomy platform developed with Celsic and other institutes that integrates open-source tools into a single web interface to accelerate discovery for hundreds of astronomers.
This project stands as a successful culmination of years of work on fundamental ecosystem layers by factoring out generic components, demonstrating that concepts applicable to fields like neuroscience or biology can drive similar advancements. The session also highlights the instrumental role of individuals like Matthew Fer in launching Spec4, which handles nightly wheels, illustrating how collaborative efforts can solve complex infrastructure challenges. Ultimately, the presentation concludes with a strong message about the power of coordinated community effort and open science principles, thanking the audience for the opportunity to share these insights and connect over the shared mission of advancing scientific computing through unified tools and governance.
Read the full video transcript
all right so yeah in this talk I promise
to tell you more about Scientific Python
and specifically a Scientific Python
project that we've been running for the
past few years um very excited to do so
but I thought since this is a little bit
more informal uh that I might use this
opportunity to give you um a bit of a
behind the scene look and a personal
overview um of how this project came to
be
um and I wanted to tell you a bit more
about the major role that Berkeley has
played in the space Berkeley and bids
has played in the space and the very
long history that that we have with
Scientific Python now you know
disclaimer my my memory is obviously uh
incomplete and subjective but I think
I've managed to uh capture the the
general gist this fairly informal so
feel free to post questions on the chat
and Jad will relay those to me
get my slides moving
here all right so um let us travel back
to uh around 2000 to
2004 what's that were you supp yeah
there we go okay yeah so let's let's go
back to around 200 2004 so at this point
in time uh Fernando Perez is a grad
student in physics at CU Boulder and
he's working on wavelet analysis using
Mathematica and he became interested in
a new language called python which at
the time lacked symbolic and numerical
computation facilities uh Jared milman
worked at the the brain Imaging Center
and there he was supporting the analysis
pipeline which was this crazy mixture of
matlb IDL orc bash c um bunch of other
things and I was at the time a grad
student in Applied Mathematics in South
Africa and I was frustrated by the lack
of a workable matlb installation for my
Linux machine um and was experiencing a
lot of issues with the licensed servers
I just couldn't use the software when I
needed to so um I switched to using new
octave and I was actively helping the
project to develop at the
time so because of this uh pipeline
challenge Jared became motivated to
replace their Imaging analysis Pipeline
with python uh Keith Worley who was a
big figure in Neuroscience um introduced
Jared to Jonathan Taylor who was a young
upcoming statistician and is now a
professor at
Stanford uh just Googling python Imaging
Neuroscience Jared came across John
Hunter who was doing EOG at the
Children's Hospital in Chicago um and
they had a matlb pipeline that he would
he wanted to replace in python as well
so so for that purpose um John started a
a project called matplot lip and John
introduced Jared to
Fernando so at that point the U there
were two array libraries for python so
the first was numeric and that was more
geared towards handling smaller data
arrays and then number array was written
by the astronomers to handle their
larger imaging sets but this Schism this
split in the community was uh was truly
problematic and made it very difficult
for for newcomers to onboard into the
ecosystem so Jared and Matthew Brett who
was visiting uh he was a visiting
researcher at the time they organized a
few meetings and these meetings brought
Travis olant who was the numeric
maintainer and Perry Greenfield the num
maintainer together um and during these
meetings they figured out a way to Mel
these two projects into one and that is
the project that eventually ended up
becoming uh numpy
so around 2005 the first course in
scientific Computing in Python uh was
given uh it was probably one of the
earliest courses anywhere on this topic
it was in Tolman Hall where the new data
science building is going to be and that
was given by Fernando and John some of
you may um may know DAV Clark he later
became a bits fellow he was also that
meeting came up from
Stanford my advisor Ben harpst um he
regularly visited Colorado on a
sabatical and this one year he came back
and he was incredibly excited he was
inspired by this student over there by
um postto over there Fernando um and he
he was he was very excited about the
prospects of this new language python so
I dropped my work on octave switched to
python um but he really wanted me and
Fernando to meet so soon after in 2006
Fernando traveled to South Africa and he
taught what we call the pi for science
workshop and I stepped in to to help him
so so that's where we met for the first
time then around 2007 Joe Harrington who
worked in planetary science wanted to uh
teach a course in using dpai and cpai
and these tools and because the tools
were so poorly documented he found it
hard to use for his course so he he had
a little bit of money and he wanted to
pay to change uh to to improve the
documentation situation so Fernando and
Jared phoned me up and with support for
my advisor I ran the uh documentation
Marathon which in the end I ran it for
two years there was another year after
that um so together with the community
we built a Wikipedia like system and
over two years we took numai and cpai
from having virtually no documentation
to being fully documented which greatly
increased the accessibility of the
libraries so in 2008 Fernando was
appointed at the brain Imaging Center
and there he worked closely with
Jared uh around
2010 Jared co-organized CPI India and
there we started working together on the
proceeding system for the the US scipi
conference around this time Josh bloom a
professor in astronomy was teaching an
annual python boot camp
uh and the AY 250 astronomy course and
through that he really helped to
establish python both in the astronomy
community and uh on Berkeley
campus so now we fast forward to
2015 uh Biz is formed and I joined as
one of the first computational fellows
that was called at the time and around
then the data science
curriculum uh was launched on campus and
it was built around python as well
around 2017 uh we started working on the
first ever grants for numai these were
given to us by the mo and the Sloan
foundations um and as part of that work
we brought key ecosystem developers to
uh to
campus and one of the challenges that
the group discussed was that as the
projects got big developers could no
longer coordinate their work in the way
they would were used to so in the early
days of sipai and numpy uh we would meet
you know 20 30 40 50 eventually 80
people would meet and they could still
fit in one room around a table and
discuss the future of this ecosystem
where to position it strategically um
you know how to align their efforts but
uh around 2017 the CPI conference was
already a thousand people in attendance
and it was just no longer feasible to uh
communicate that way so you know we had
this ecosystem that without much
guidance grew very organically and many
projects in the process became quite
siloed doing their own little thing um
there was very little Unity there was
very little to know careful cross
project design and coordination and
while tools were sometimes brought or
borrowed from one project and used in
another uh the tools were then developed
independently and would often take take
on a life of their
own but because of the the success of
the projects because of the adoption
also as part of the data science
movement uh they went have now very many
more users to support but there was the
the number of developers didn't increase
that much and there was no framework
within which to support all these new
users industry at the same time became
probably because of machine learning
became much more interested in the
ecosystem and we were worried what did
this mean for the developers how would
the community remain involved in you you
steering and designing the projects that
they founded so we started thinking
about ways to bring the the community
together again and uh we
discussed like we wanted to bring them
together so they could discuss strategy
so they could improve the uniformity for
the benefit of the users if you look
like something you know product like
Matlab they have a product team they
have a documentation team they work
closely together very unified um
experience for the users
um and we wanted to give our users the
same we wanted to share technical
infrastructure and tools so that we
didn't keep Reinventing the wheels so we
didn't keep on wasting time um and we
wanted to make sure that projects were
well governed or Community governed and
were owned by the by the community of
researcher
developers so here's the the last bit of
timeline I want to discuss so around
2018 that was the time of that numpy
meeting uh we we created the first
Scientific Python site was just a um
scaffolding for that site more or less
but this is where we started really
writing up the um the question about how
to communicate the how to coordinate the
community um around December of 2018 we
had a conversation with Chris menel who
was then at the Moore Foundation and he
asked us to write a short proposal about
this work that we envisioned so in
January of 2019 we produced the first
draft of that and eventually in June
2020 that funding was approved and
through that Jared and I were able to in
December of 2020 officially start uh the
Scientific Python project okay so that
gives you the the personal history of
how we came to this point and now I'd
like to tell you just a little bit more
about Scientific Python project the
Scientific Python project what it is and
how it works and everything we've done
since 2020
right so Scientific Python is a project
which is aimed at better coordinating
the ecosystem and to support
specifically the maintainers of that
ecosystem um here if you want to learn
more about the project after the talk uh
you should take a look at our website
scientific-grade
is that we have S of these six key areas
of focus your slid Stu what's that your
slide seems stuck thank you uh one of
the there's one area missing here which
is tools but I'll get to that in the end
so yeah we've got uh spec Summits
development guide lecture notes we're
working on spars arrays and we have
Community coordination so I'll go into
each of those a little bit just tell you
what what we're doing there
right so the specs are Scientific Python
ecosystem coordination documents so if
you followed along with uh python
development or with numpy development um
python has these documents that discuss
technical proposals about changes uh
they want to make to the library and
then there's discussion around these
proposals and a process whereby these
proposals are vetted and approved by the
community Jared took that process and
modified it for numpy and numpy is a
similar process and we have the same for
Scientific Python now the the difference
between the specs and something like
Python's peps or nump neps is that those
are technical documents very
specifically geared towards one project
whereas what we're talking about here um
we you know we're really looking at
crosscutting concerns across various
projects in the community so so much
we're taking a much broader view the
specs are also not that technical in
nature they often refer to external
sources for the technical details but it
tries to be an index of uh development
ideas where projects can come especially
new projects can come and see how are
things done in the ecosystem it's also a
mechanism whereby smaller projects can
have a voice in the ecosystem because
usually they would not be heard very
easily but if they write up a spec if
it's a good idea if the bigger projects
like it then um then those ideas can be
adopted so the adoption works by having
the big projects what we call the core
projects evaluate these proposals and uh
and endorse them so basically say do we
think this is a good idea doesn't even
necessarily mean that they're
implementing it themselves it just means
that they say yeah if you want to go
this route it's probably smart um and uh
the specs are very much you know we're
not there's no control over the
ecosystem libraries can do exactly what
they want to but we have these
guidelines to make it easier for
developers to uh to get going I mean one
one of the big things we're trying to
address with Scientific Python is to
reduce the the overload on maintainers
to reduce the amount of um you know
chores they have to do just to get their
projects out there we really want to
free them up to think about their domain
problem or their domain Library so we're
trying to to take care of some of the
busy work the um you know yeah all these
things they would have to think about
otherwise we're just trying to get that
in the central location make it easier
to access the um the spec process is
driven by the spec steering committee
these are community volunteers who come
in and help us facilitate the process of
shepherding these docs through the
process of being uh proposed accepted
evaluated and written and so
forth show you a few examp example spec
so you get a a feeling of I think I had
a list earlier
on see yeah so here's the full list of
specs we have so far um you can see it
covers everything from dependencies to
API dispatching to how do you keep your
password safe in the project how do you
seed random numbers it's really all the
cross cutting concerns that we think
about um
so here's an example so the first spe we
wrote was which versions do the do the
developers have to support so again like
I mentioned we're trying to reduce the
maintenance burden on developers and the
developers feel a great responsibility
to make life easier for their user so
they would typically try to support as
many versions of python as they can they
will try and support as many versions of
scipi and numpy as they can so that you
can install it easily on your system but
ultimately supporting many versions of
those libraries and and and the language
comes at um uh you know it comes at a
cost so we have this policy that
basically says or this recommendation
that basically says look if you if you
at least uh support these three versions
or if you support python over this
period of time then that's good enough
for the ecosystem and it frees you up to
pay attention to other things another
example is spec one which is about lazy
loading so in the early days of scipi um
python has this concept of name spaces
and the name spaces basically extend to
subname spaces and as you have these
deep hierarchies of functions it can be
become quite hard to discover them so if
you're working inside of something like
IPython or Jupiter you have this handy
tab completion but you can only tab
complete once you've imported all the
modules as deep in the trees as you
needed to go we didn't like that
situation but originally it was costly
from an import perspective to just
import the whole tree so python
eventually came up with a new feature um
that allowed us to to handle that much
more smartly and efficiently and you
know as project we're trying to figure
out how do we use this new feature we
figure it out once we write it up
there's the spec um and we wrote a
little library to to expose that
functionality and now all the libraries
can come together and use use that same
functionality together if you introduce
a problem at least you introduce it in
one place you can fix it in one place
and the maintainers don't have to worry
about
it spec four is another um coordinating
mechanism for deploying Wheels regularly
so wheels are the binary artifacts every
night projects in their um continuous
integration in their testing systems
they build a version of the project but
it would be really handy if this test
version yeah so so you have have this
mechanism whereby projects they every
night they produce a development version
of the project and you know before the
projects would just do it individually
and
um then no one would know where to find
those those artifacts so spec four is
all about saying let's upload this to a
single Central repository here's how you
do it here's how you keep it safe and
secure um here's how other projects can
use it in their testing Frameworks um
and that way we we much better
coordinate testing we catch errors
earlier um when
whenn right moving on to from the specs
to the second category here we have the
developer Summits so the Summits are um
they're in-person working meetings and
they're all about working so the goal
there is to bring key people in the
ecosystem together key Volunteers in the
ecosystem we spend a lot of time
beforehand to plan work that we're going
to do um here's yesterday's uh meeting
for this upcoming Summit that's
happening in in Seattle in June so we
coordinate the work we're going to do we
identify pressing matters in the
ecosystem that need to be addressed and
then we all get together in one space
and just for a week or in this case for
three days we we just just work on
knocking those out and we've had some
really good successes on this I think
one of the key examples for me was um we
really wanted to improve the situation
around spars and I'll get to that later
on but the this inperson Workshop really
was key to bringing all the different
people together in one room because you
know you land in this situation where
you have a core developer on CI that's
responsible for approving the changes
that you make so even if you work at the
summit you get a lot of work done you
still have this one person that needs to
review your work and um and get it into
the library but if you can get that
person into the room with you
then changes can happen much more
quickly okay next we have the
development
guide so again you know if you're
starting out with a new library then how
do you know which best practices to
follow how do you know what changed over
the past year you know you're busy
trying to figure out how to implement
the routines in your library um
so as you're staying up to date with
that you don't really have time to read
up on what changed the ecosystem how it
set up tools break various libraries you
know what new changes are in GI upci so
all these things the community community
can track for you and aggregate in um in
these documentation Pages you at least
have one Central resource to refer
to one example of um a tool that was
developed as part of that development
guide uh by restrainer is is this repo
review so this is a web assembly tool
that will run in your browser connect to
your project repository and do a bunch
of checks just make sure that you know
you have a license you have tests you
use um a linter correctly etc
etc U making it much easier to just have
a checklist for um is my project
basically up to up to scratch and up to
up to
standard the CPI lecture
notes so this is where we basically go
beyond where the carpentries go and we
say um we'd like to teach you about how
to use these tools but we' like to go
you know um it's fine to start with the
basics but we really want to take you to
the point where you can use this for
your science and your research so it
goes fairly in depth um and yeah it was
developed by by numerous members of the
community as you can as you can see
there on the cover uh we sort of so so
these notes were fairly well established
but then they languished for a while so
Scientific Python took over the
maintenance of these nodes and we've
been refurbishing them both at the
Summits and um and whenever we get a
chance update them you know to use the
latest libraries to use the nearest
changes in nump by API and so forth so
you can
see they they start off with python they
talk about Python and numpy matpod lib
the basics and then they they go all all
the way through to more Advan topics
like optimization image processing using
sparse matrices and so
forth then as I mentioned one of our big
focuses is improving sparse array
support so especially as we move into
much larger data sets it's important to
be able to handle sparse matrices you
can have enormous uh matrices but they
only have entries in a few locations you
should be able to store those and
operate in those efficiently we have
fairly reasonable machinery for doing
that um in
scipi but unfortunately this was written
uh a long time ago when the predominant
API wasn't the numpy array that we know
now but there was also in numpy a
concept known as well we called I mean
they were known as matrices before we
used array
and matrices and arrays are fundament
like they operate quite differently
matrices are basically 2D arrays are ND
and their interfaces look different so
our goal with this project was to take
the CI aray functionality and to Port it
to the numpy array API to make it much
more modern and to support the various
dimensionalities as well so starting
with 1D support um eventually we may end
up with ND support but at least going 1D
2D and not not only 2D so um the sparse
functionality becomes much more
compatible with with the majority of
research
codes all right I think I spoke about
that right and then finally we try and
provide platforms where the community
can connect and communicate with one
another so
we have a a discourse server where we
have discussions and um scipi recently
moved their list over there pyed image
has been using it for a long time that's
where the spec coordination
happens and then for more informal
lighter weight conversation we have a
Discord so it's just a chat server that
also supports video and um good for for
quick to and two and two and fro
conversations
then we also have the blog so we've
aggregated some of the community blogs
that were out there already like the mat
blot Li blog and um we put a procedure
in place to make it easy for the
communicate for the community to
contribute blog posts and write in a
more informal way about the work they're
doing and um and yeah the projects
they're involved
with I mentioned that the one thing that
was Miss ing from that front page was
the tools work that we're
doing so as I mentioned earlier on we we
try and do a lot of work to support the
maintainer so that they don't have to
carry such a big burden and part of that
is
centralizing the maintenance of common
tools used across the ecosystem so you
know I remember from psyched image for
example one of the early things we
wanted to do is generate API
documentation now to generate API
documentation you need to uh go through
the functions find their signatures um
list them all in a coherent way in a way
that your documentation compilation
engine can understand so you know we
grabbed a tool that Fernando had written
for IPython back in the day that tool
got developed in psyched image and
several other places got specialized
those are the types of tools that we
really want to host centrally so that uh
we can maintain it on behalf of the
whole ecosystem
so we've got tools uh for the web this
is mainly a um a theme let me scroll
here
[Music]
yeah so yeah for the web we mainly have
the Scientific Python Yugo theme that's
like a Yugo theme that's roughly
equivalent to the pi data Sphinx theme
which is commonly used for documentation
in the ecosystem uh for documentation uh
sorry for development I'll just
highlight a few we've got spin which is
essentially the userfriendly interface
to building numpy psyched image and so
forth it uh it hides all the complexity
of building with Mison and you get a
userfriendly interface that you can
invoke uh repo review I already showed
you change list is the thing we use to
automatically generate release notes for
each
release um we've got a community
calendar that we host so you submit a PO
request change a yaml file on our
website and it renders a calendar that
you can subscribe to
Google we've got a developer dashboard
so it looks at your repository and it it
look it uh calculates some statistics on
you know how many PRS were open and
closed how healthy is your project uh
where are pain points that you might
need to pay attention to and so forth so
an analytics dashboard for your um
project we also have analytics like
Google analytics style that we host for
various projects
um I think I'll stop there for
now all right
so what is on the horizon I want to
highlight here a section from um
Fernando who's our new academic director
I want to highlight a piece from his um
Vision that he put down for bids um I'm
very excited to see this renewed
emphasis on the importance of uh open
software and research he wrote Our
scient partner with an extended
distributed community of other
researchers and developers to build an
ecosystem that benefits all uh this is
how we will build much more Inc coming
years working grounded in the expertise
of our Scholars and
immediately um applied to research and
educational needs but in open
collaboration with Partners near and far
to build access to um research and
education that is impactful accessible
and fair so yeah for me these values is
aligned so perfectly with the work that
we've been doing over the past decade
and a half and what we want to see on
Berkeley
campus so some um exciting prospects I
wanted to highlight uh there's an open
source projects office being established
at Berkeley and I think it's great to
have a central hub for conversation
around open source um if you want to
learn more about that you can chat to
Jared mil is also on this call areas
that we intend to work on as part of
Scientific Python we really want to
improve the statistics offering in
Python which I think has a has a few
blind spots um currently the offering is
very econometric centered which is fine
but you know it doesn't suit all our
users we want to focus on specific
subdomains that use our tools for
example improve the offering for
astronomy or for Earth and space science
there are a few other domains maybe
geosciences where we really want to go
in and make sure that the tools are well
compatible with the needs of those subc
communities there's a panel this Friday
on a very important topic that has just
become even more relevant supply chain
security um with with a recent
exy um project infiltration situation
that took place we're of course now very
aware of of the risks that we face in
shipping the software to millions of
computers
um not Friday thday say
again Thursday not Friday THS Thursday
apology is the date right is it May
9th yes thday 4 to six okay
perfect then we also better want to
coordinate releases we want to make sure
that this is an idea that's come along
for that's that's been there for a long
time but uh no one's been able to act on
it but you should be able to essentially
install uh a combination
of uh of a set of ecosystem libraries
maybe the most commonly used libraries
and have some guarantee that these these
libraries have been tested together will
work well together um similar to the
notion of you know ubun releasing every
every quarter or whenever they do um
really making sure that we can provide a
stable set of cross tested packages
we're going to host some summer schools
where both for developers and for users
of the ecosystem and then finally we
also want to look at shared governance
models and make sure that all the
projects in the ecosystem have solid
government governance models that are
essentially yeah compatible or um have
been have been vetted by the community
in the same
way all right so I I think we can switch
modes to the Q&A um again if you want to
more about the project please navigate
to Scientific dp.org um and please
connect with me on on
Mastadon that's that's where we where I
hang out nowadays um yeah so I'll open
the floor to to questions
St maybe one question to PRT discussion
and thanks for that wonderful
overview what do you see as immediate
opportunities and ways in which bids uh
and and your team in the context of uh
of kind of this this new direction that
we're trying to build
upon can help the campus community and I
know we have some guests from Beyond
campus and that a lot of this is framed
uh framed as us being part of this
distributed Community um but but but we
are first and foremost a cus entity and
so how what do you see as key
opportunities at at various levels where
you imagine we could serve serve the
needs and if folks on the call have
questions or ideas we we welcome your
input this is very much a conversation
we really want to build a sense of of a
thriving and useful community on campus
regarding these
things yeah thanks for that for I think
you know we've spoken a lot about
coordination between the um the projects
and you know between campus and the
community out there but you're you're
absolutely right that um you know first
and foremost we need to address the
needs of the researchers on campus and
there's a lot of python usage on campus
there's a lot of um there are a lot of
good things happening and a lot of
Talent on campus that that can be
harnessed um and I think you know we we
haven't really had BDS BDS has started
to be that Central meeting place where
especially when we had the physical
place it was kind of easy for that to
happen you people would flow through a
single place where they could talk where
they could ask questions about the
ecosystem or tell us about their
research problems we could really work
out what is the best way to support them
but I think um I think we still have
quite a ways to go in making sure that
the needs of the Campus Community are um
are better addressed the the tools needs
and also that we involve the Campus
Community more closely in the
development of these tools because I
think when when you own a tool when you
modify a tool for your own purposes it
does two things for you first it makes
you understand the problem that you're
dealing with really well and then once
you understand the problem really well
you build a tool that is perfectly
suited for your problem and um so it's a
valuable process to be involved in the
in the development of your own tool and
I'd love to see more of that happening
on
campus Stefan I have
to yeah so another uh way that we'd love
to to interact with people is um you
know so most of our work is at this
point is Grant funded and so that's
somewhat focused our work on things but
uh we also have uh you know a long track
record that's been established as
Scientific Python experts uh we also
have some funding track record and we
would love to uh collaborate uh in some
kind of capacity with uh researchers
submitting new grants um you know I
think one of the things we have is that
many of the open- source uh calls that
are coming out now ask for plans for
governance and how you relate back to
the major Community uh and how you're
going to do a sustainability model I
think that's something that we can
partner with uh other um researchers on
campus and help address that question uh
and you know we're building up a team
we're hoping I think to um provide a
central pool of uh expertise and
resources that can help uh domain
experts and researchers that are
building out these capacities uh and
with Scientific Python we certainly have
that covered I think we're also hoping
with the ASO to be much more language
agnostic uh so that's definitely a way
we're very excited to collaborate and
grow um together as a community
Stefan I don't know if you're watching
the chat you have a question in the chat
from Serena can read it read it and take
it take it thank you um yes so Serena WR
um thank uh so we work on muscular
skeletal image analysis what would you
recommend as an entry point um to use a
tools you create yeah so I guess that
depends on a few things you know how do
do you have an established pip line do
you um you know how familiar is your
team with um you know with with the
ecosystem already we have like U psyched
image is is a fairly good starting point
when you want to do uh 2D types of
analysis it's it's we're working towards
more IND dimensional but it's uh I'm not
quite sure in your problem whether
whether you would need that but we have
fairly good tools for 2D analysis and
then if if you need to move towards more
realtime analysis like handling video
feeds and so on then open CV Still
Remains one of the best optimized
library to do that but maybe you can
give me a little bit more information
about your project or if you prefer to
continue the conversation afterwards I'd
be I'd be happy to have a
chat I'll also add that
um yes it would be great to have the
conversation offline uh one of the
things that uh a challenge we had is as
we were establishing this it was the
middle of the pandemic and so uh there
was a really difficult time having
online events and I think it also had a
lot of early work was done in sort of
getting the flywheel started I think
we're now at an area um where we want to
have um more interaction with the
community uh at Berkeley but also
locally so I think we're going to be
having uh trying to organize more events
um where we bring people together um
different types of summits and uh we've
done that across projects but I think
it'd be nice to have a set of local
projects as well I'm not sure but it
look like you might be at Stanford um we
have a lot of collaborations with
Stanford and love to have more um yeah
please reach out to
us Stefan maybe another question for you
one of
the one of the most amazing reservoirs
of
energy Creative Energy intellectual
energy time uh that a university has
That's Unique is it's pool of
undergraduates uh and also graduate
students so in general the student
population that's a really special part
of being and it's kind of what makes you
University a university right and uh but
if we compare some of the retrospective
that you made about the early days of
Scientific Python even at Berkeley and
today one important thing has change
right now we are teaching I have the
data on how many students take data
eight and years ago I had said I'm
guessing when we reach steady state this
is going to be on the order of half of
campus and it is basically dead on half
of Campus so I have the data from the
enrollment of data 8 that data 8 is
taken by 50% of the University um data
100 is U not as much but still on the
order I if I remember correctly about
30% give or take that's an upper
division course that's reasonably
Advanced Data analysis with python and
so we have now a population at the
undergraduate level that on average
right if you're gonna throw a rock you
have a 50-50 chance of hitting someone
who's used to using Python and who also
has access to all of this incredible
infrastructure of cloud-based data hubs
that are deployed there for the entire
campus to use how do you imagine kind of
harnessing that energy that energy now
is much better prepared much the the
ground layer has been built in that
population right to be much better align
than anything we could have remotely
dreamed of 10 or 15 years ago what do
you
imagine in terms of engaging with that
community and and and that resource both
ways and ways that benefit your goals
and in ways that benefit them as well
right yeah that no that's it's such a
fun topic to think about as well um and
you know what a lot of people don't know
is that this uh the whole ecosystem was
in large part built by by students of
course things have changed a little bit
now in the sense that the the ecosystem
is more mature so it's it's harder in
that sense to get involved you know when
you work to work I could as a as a
master student work on numpy no problem
because it was essentially so broken
that anything you did to improve it was
was fantastic uh now you know that
library is much more sophisticated it's
got a bunch of of features it's well
tested and so forth so much harder to
walk in there but there are still new
libraries being formed there you know
all the time people try new things in
the ecosystem so there are still ways to
engage with that ecosystem it's not it's
not trivial
to engage with a massive amount of
collaboration all of a sudden because
you have you know part of the way we
develop our tools is we we make sure
that the code contributions are
carefully uh reviewed so so as you bring
more developers into the space you also
put a bigger burden on the reviewers
that that have to handle that load so I
think we have to think a little bit
about you know what is the best way we
can take that effort and um do work but
then sort of polish that work up before
it reaches the the interface with the
ecosystem one effort that like we've
done some experiments in bids along
those lines um around trying to remember
it was around 2016 or 17 maybe for a few
years I ran something I called the
machine shop and the Machine Shop was
basically small teams of students that
would come in and they would work on a
specific data science Challenge and
several of the tools we wrote uh for
that actually went out and was was used
I mean we we did some work for the
Natural History Museum um that they
recently publ published and we yeah we
did work that ended up in you know in
computational libraries that are being
used today um that relied on someone on
campus taking the role of mentorship but
the students very often more or less
coordinated the work u in their little
teams so I think with some lightweight
organization around effort like that um
we could really um we could harness a
lot of that potential I think it's
incredibly valuable to participate in an
open source project you learn so much in
the process and you hang out with some
of the best developers um you know in
the world so you know these are people
who volunteer their time just because
they care about of topics so they're
passionate they know what they're doing
and that's how that's how I learned most
of what I know about software
engineering today so yeah worth thinking
about worth getting the students
involved um if we can put the right
infrastructure in place as to not
overwhelm the existing community of
maintainers
Alec yeah hi Stephan thank thanks for a
nice overview especially the the
community um maybe taking taking off
from Fernando's comments about I would
call it maybe the local Berkeley market
for for for your products and and sort
of coming coming at the issue of of uh
trying to understand uh the the the
various interests within the within and
outside of the community do you do you
do you have any notion other than of
course everyone's personal observations
about the demographics of of of your of
course your participants your developers
and people directly involved in
governance but most specifically about
the user Community uh and the user
communities especially when you're
you're addressing multi-domains that
that are quite distinct I
mean genetic sequencing and and
cosmology uh come together maybe just
down at the Matrix
uh uh what you know in in a in a in a in
a corporate setting this would be a
marketing uh departments that would deal
with with their expertise and their
funding as well yeah so any any thoughts
about sort of taking what Fernando was
saying about the Berkeley
opportunity and casting the broader net
and then and then seeing how that works
down to to local the local situation
yeah I'll say maybe first that we really
welcome participation and very wide
participation there are many reasons
systemic and otherwise why participation
is not where we want it to be but from
our perspective we you know we benefit
from greater diversity in the developer
community so we want to improve the
onboarding processes we want to improve
how friendly the community is you know
we're always thinking about um how to
better integrate with the community how
our user Community looks is um is not
such an easy thing for us to measure
because you know as open source
libraries we don't typically collect
Telemetry and we don't have frontend on
top of this Library although there are
some thoughts floating around on how we
could engage with project Jupiter on
getting getting permission on asking
some of these questions that that we
want to know um
the uh the ecosystem very much developed
by way of people scratching their itches
you know they would come in they would
volunteer a piece of work and that work
would be would be integrated and that's
how we got to know one another and
that's how we um how we grew grew the
project so it's not it's not been run as
much as a traditional project as you
know in working in an ecosystem where
you become aware of a problem and try to
address that um that
problem maybe maybe there's some hope
that an open source program office might
be interested in that and could provide
the various sub various projects with at
least at least help and and
understanding users and commun better
yeah for me what's exciting about the
project office is as I me the talk I
think it's a it's a point of connection
like bit is too you know it's a point
where people come together they can talk
about about a topic it's you know right
now if you have questions about open
source if if you want to engage in the
process or um if you want to know about
someone who's active in open source and
campus there's there's no no place to go
and have that discussion so I hope we
can through the ASO broaden that space
Maybe throw in throw in daab as well
so yeah absolutely I I just to add that
I think um you know as a visit is
getting relaunched both with the new
faculty director and also with the new
homecoming up I think there's a lot of
uh excitement with the oso and other
things about trying to figure out new
spaces to meet and new collaborations to
have um so I think possibly over the
summer uh the funding for the ASO is not
a lot and I've got a lot of the projects
but I'm going to be starting to reach
out and trying to figure out different
labs and different groups to connect the
lab's an obvious one we also would like
to have a better integration with the
library uh
and all the other various groups on
campus um and I know that Stanford has
an Oso that we'll be looking at and one
of the things they have is a a like a
contributor and maintainers Round Table
uh so we may be looking at events like
that where uh we'll have forums where um
you know grad students and post dos
researchers can come together and share
lessons and uh tell each other about
projects they're working on um and so
people have ideas like for things like
that I I'm all years and I hopefully
we'll be setting up a number of these
events uh of the next
year and Stefan I I know that I asked
about things on campus because we are a
campus entity but I also know that we
have a number of uh guests on Zoom or
not not Berkeley guests and so I'm
wondering if any of them have questions
for Stefan and Jared and for us because
as Stefan also pointed out these
projects kind of are a hinge between our
Campus Community and expertise and and a
broad distributed community so if folks
have questions please don't be shy and
join and we we look forward to knowing
where you're coming from and finding
ways to connect with you if you were
interested in coming to this this talk
it's because you have some interest in
our work and we'd love to connect with
you so please I want to encourage our
non bkly guests to kind of participate
or ask questions if they if they have
any it's great to see some Outsiders
maybe can I make one comment yeah I
see I see go ahead like connect with the
outside Community please go ahead I just
see that guys on the call um guy and I
work together on with Josh Bloom on Sky
portal I think that for me has been one
of the most interesting projects um uh
you know scientific platforms developed
at bids so this is a an astronomy
platform that we develop together with
celtech and and several other institutes
and um it's a place where it's basically
a platform that uses all these open
source tools that we've built over the
years and brings all these tools
together in a single web platform where
astronomers can come and talk about
their work share their results and
accelerate uh Discovery and I think that
is such a beautiful example of how all
this work is paid off over the years you
know we've been working on this fun on
the fundamental layers in the ecosystem
and now finally we have this thing
together that that culminates in in a
platform that allows hundreds of
astronomers to um to accelerate their
their
workflows and Stan you've uh you guys
have factored out some generic
components of that too so it's something
where there's already now I think that's
a lot of the work we do if we try to
find the generic parts and um hopefully
we'll be matching up with other maybe
you can some of
that yeah definitely I mean the you know
many of the web the the web platform and
then a lot of the ideas explored there
could be applied equally well to to
Neuroscience or biology or what have
you I think I saw uh Matthew fer on the
call as well he was instrumental in
getting Spec 4 launched that's the
project that deals with
uh with nightly Wheels thanks very much
for your work on that
M all right well thank you very much for
this opportunity to speak to you all and
to connect I really appreciate it