Growing agriculture with big data: Philip Evans, Boston Consulting Group
Watch on YouTubeVideo summary
Philip Evans begins his presentation by drawing a parallel between Jorge Luis Borges' story of an infinite, one-to-one map and the current trajectory of technology and data. He argues that we are approaching a similar state where our understanding of the world becomes so detailed and comprehensive that it effectively mirrors reality itself. To illustrate this, he points to Google's self-driving cars, which utilize vast datasets to perceive their environment with superhuman accuracy, identifying cyclists, traffic lights, and road conditions better than any human driver. This capability is built upon three fundamental mechanisms: first, the use of external data like satellite imagery to create precise agricultural maps; second, the ability of objects to describe themselves through metadata, such as cell phones revealing population movement patterns which allowed researchers to optimize bus routes in Abidjan and reduce commute times by ten percent without adding a single vehicle; and third, the recent breakthroughs in artificial intelligence where machines can now perceive and understand their surroundings, eventually reaching a point where they can generate grammatically correct summaries of video content.
These advancements are driven by four dramatic technological shifts: the Internet of Things, which has proliferated sensors to an estimated 168 per person globally; the explosion of big data, with the world's data stock doubling every two years; the rise of artificial intelligence through neural networks; and mobility, which allows insights to be delivered exactly where they are needed. Evans emphasizes that these technologies have only emerged in the last seven or eight years, creating a new macro-pattern where billions of devices connect via IP addresses to exchange information at near-zero latency. While institutional barriers like privacy concerns exist, the technical architecture supports a unified global data set. This shift is further enabled by massive economies of scale found in centralized data centers, which store data once for reuse, while the interpretation and analysis of that data fragment out to small teams and individuals who can solve complex problems autonomously.
The presentation highlights how this new landscape is disrupting traditional business models centered on linear value chains, where physical products move through a series of activities from raw material to final good. Instead, Evans describes a transition toward a horizontally stratified "stack" of businesses that are fundamentally different yet interconnected, serving each other's needs. This fragmentation occurs because transaction costs are falling, allowing separate pieces of the value chain to operate independently, while economies of scale in data processing lead to consolidation among platform giants like Google and Microsoft Azure. Meanwhile, individual initiative, creativity, and experimentation drive innovation from the bottom up, as seen in platforms like Kaggle where diverse participants compete to solve specific data problems for companies like Allstate or Merck. In these contests, teams with no background in chemistry or biology have successfully predicted toxic side effects of drugs by simply mastering data analysis, proving that raw data control is more valuable than being the smartest person in the room.
Ultimately, Evans concludes that the defining phenomenon of our changing world is the shift from an economy dominated by the economics of things to one dominated by the economics of information. This transition represents a profound restructuring of industries, moving away from vertically integrated models toward a dynamic ecosystem where services are provided across a digital stack. The value of data control has become the primary basis for reward and competitive advantage, enabling organizations to reject bad deals, accept good ones, and optimize operations with unprecedented efficiency. As we stand on the brink of this new era, the ability to harness these four technologies will determine success, marking a departure from traditional strategies that relied on physical flows toward a future defined by information processing and decentralized innovation.
Read the full video transcript
good afternoon it's a great uh it's a
great pleasure to be here
um i am neither an australian nor an
agronomist so i'm in many ways uniquely
unqualified to be saying anything to
this august group
what i would like to offer for what it's
worth is a brief perspective on
how
data and information technology more
broadly
have been changing in recent years and
what are some of the models and
frameworks that we see evolving
now if i can get to my presentation
therefore
there we go
um i'd like to start if i may with a
story
um jorge luis borges the argentinian
poet and novelist wrote a very famous
short story it doesn't even have a title
and in this short story he recounted the
tale of a lost kingdom this was a
kingdom somewhere you can imagine in the
andes where the aristocracy were
obsessed with maps
and
however each map that they cast was
deemed unsatisfactory it wasn't good
enough so they launched a new project
for an even more ambitious map until
finally they launched the ultimate
mapping project a map of their kingdom
on a scale of one to one
and according to borges if you visit
this kingdom to this day you will find
fragments of this failed project
in various parts of the terrain
now book has obviously had a lot of
philosophical ideas behind this and it
was a tale of human hubris and so on but
what i would like to suggest
is that in many ways that idea the idea
of a map on a scale of one to one
is indeed precisely the image that we
should have of where technology in
general and data in particular are
taking us
and to illustrate the point consider the
following
this is what the google car
sees
uh in the inside on the left you have
the view from the dashboard but in the
larger part of the screen what you see
is the car's detection not just of the
roads and the bridges and the traffic
lanes but also of other vehicles also of
cyclists of traffic lights and so forth
indeed there's a wonderful moment just
coming up where the car faces the
challenge of trying to overtake a
typically wayward erratic and irrational
and suicidal northern californian
cyclist who cannot make up his mind
whether he is turning left or turning
right and notwithstanding his random
flapping of his hands nonetheless
the car successfully passes him
indeed to this day uh the uh google car
has only experienced two accidents uh
one of which was when it was rammed from
the rear while parked and the other
which was when it was being driven
manually by a google engineer
now
what are the components of the
technology that we see here what i'd
like to suggest is that it's worth
looking at this because those components
are actually quite general
one of them is very simply and obviously
the existence of a general comprehensive
and accurate map of the terrain in
google's case of course this would be
the google maps map of the roads but in
fact we have many examples of that and
an obvious agricultural example would be
the use of satellite data in order to
know not just the uh the bounds of the
terrain but to understand things like
the humidity or the ph factors uh or off
the soil which enables all sorts of um
uh aspects of precision agriculture that
would not have been possible before
but there is actually i would argue a at
least two other mechanisms that are
involved that are less obvious
one of which is that
the map is generated from things
describing themselves this is a classic
example this is a view of the united
states um taken from a satellite at
night what you see is the pattern of
light from cities railroad tracks
roadways um airplanes and so forth and
very obviously you can see the shape of
the country
notoriously if we go to africa it is the
dark continent the continent where um uh
such uh infrastructure is conspicuous by
its absence but not entirely and that is
in many ways what makes the story
interesting orange the
telecommunications company the french
telecommunications company is the
monopoly provider of a cell phone
service in the ivory coast what orange
did was they they collected meta data on
how people use the telephones in the
cell phones in the ivory coast nine
months worth of data for the entire
population and by doing that they were
able to create networks like this where
you could see the movements of people
because of course the cell phone
registers with the nearest tower so that
the telephone company actually knows
where it is
and who speaks to whom and even in fact
what language they speak in
what orange did having collected this
data at very low cost was simply to
publish it they anonymized it and they
published it and they said to the
world's researchers go at it see what
information see what patterns you can
find in this data
and interestingly 82 papers were written
by researchers and academics around the
world using this unique data set as a
source of insight
one of those papers came up with the
following rather elegant analysis some
researchers at ibm looked at the um
pattern of commuting in aberjan which is
the largest city in the ivory coast now
they weren't looking at where people get
on the bus and get off the bus which the
bus company kind of knew from its own
information what they were looking at
was where people started their journey
and where people finished their journey
and that enabled them to think therefore
of a very large optimization problem
which is you've got a finite number of
buses you've got a population who need
to on a daily basis get to and from two
locations their home and their work what
is the optimal allocation of bus routes
that will
minimize the time that it takes the
average
citizen of the town to get to and from
work
and they formulated this problem the
mathematics is quite trivial trivial the
computation of course is horrendous they
formulated this problem they ran it on a
hadoop cluster for two days and sure
enough they got an answer
and the answer was a reconfiguration of
the abidjan bus system which without
adding a single bus shave ten percent
off the average commute time
now you think of
the effective impact on as it were the
gdp of the country to reduce everybody's
work day by 10
and the impact is truly extraordinary
why was that possible it was possible
because of a
very very large data set which had not
been available before
and because of a few smart people and
access to some very powerful machines
but only for a few days in scale terms
the really difficult bit was the data
once the data was available all sorts of
things became possible
but that's just the second mechanism by
which these maps can be created
the maps created through objects
describing themselves in this case the
cell phone declaring its location
there is a third mechanism which is the
most recent and in many ways the most
profound and that is the ability of
machines to actually perceive and
understand their environment there was
an immense breakthrough in something
called convolutional neural networks
about five years ago
again based on big data
that made it possible for machines to
recognize for example objects in a
photograph
and an annual competition held by
stanford university to solve this
particular problem was won by a team
from microsoft this year with a solution
that is more accurate in predicting the
identifying the objects in a photograph
than is the average human
but that is just the beginning look at
this
this is a video that was created about
six weeks ago by a graduate student in
amsterdam what he did was he was running
a piece of software on his macintosh
that is using neural network technology
not just to identify objects but to
caption that is to say to create
grammatical sentences describing what is
going on
and what you see he what you see is is
that he is walking down the street and
the the camera embedded in his laptop is
uh generating images which the software
is interpreting
now you'll observe very obviously that
about half of these captions are
incorrect
this is about where the identification
of photographs was about two or three
years ago and everybody knows full well
that at the rate at which this is
advancing within two or three years we
will have instead of a boat is parked on
the side of the water we will have a
boat is moored at the side of the water
because the machine by then will have
learned the correct grammar
it is confidently predicted that within
five years a machine will be able to
watch a youtube video and um
generate a grammatical sensible one
paragraph summary of what is the story
of that video that is how near
machine learning is
so when we stand back
from these phenomena we actually have
these three different ways that these
maps can be created
what underlies them is four
dramatically new technologies one is the
internet of things the vast
proliferation of uh sensors
in the world it's estimated today that
there are 168 sensors for every man
woman and child on the planet the cost
of these sensors is dropping by an order
of magnitude every five or six years
it's confidently expected that that
number will multiply by a factor of 100
in the next 10 years
secondly we have big data and i'm sure
you have heard of the statistic
frequently quoted that the world stock
of data is doubling every two years
thirdly we have artificial intelligence
data as jackie said earlier data is
worthless without the insight that
interprets it and it is artificial
intelligence that has been subject to
many major breakthroughs the one
particular one about neural networks in
just the past five years
and then finally we have mobility we
have the fact that the information and
the insight that is being generated
can now be used at the point where it is
needed at the point where it is relevant
most obviously in the case of
consumers
in the form of information delivered to
your smartphone and the number of phones
in the world is now greater than the
number of people there are two and a
half billion smartphones in the world
within five years they will all be
smartphones this is just a matter of
time
now a couple of points about this one
this is one big system this isn't a
pattern that replicates itself millions
of times
the internet of things is billions of
devices connected to each other
via ip addresses and the web the big big
data when we talk about big data what
we're talking about is data that has
that is on servers
excuse me or on laptops all of which
again have ip addresses meaning that
they are connected to each other meaning
that in principle at almost zero costs
and at almost zero latency they can
exchange information now they may not
because of course different people own
it there's issues of privacy um uh
intellectual property all sorts of
reasons why institutionally it doesn't
happen but from a technical point of
view it is one data set in similar
fashion very obviously the phones are
all connected to the global
telecommunications network so what we're
talking about is the emergence of this
macro pattern
the other key point to emphasize is how
recent all of these things were if you
turn the clock back seven or eight years
nobody was talking about any of them
now
in that
world
what are the institutions that are
needed what are the institutions that
make it possible
and the answer here is a little
paradoxical because it's an alliance
between the very big and the very small
one component of this architecture is
data centers things like cloud computing
of which i'm sure you've heard
this is a
and not a typical
data center belonging to google the
world's largest data center which is
under construction in china
is the size of 120 football fields if
you can imagine a building of that size
it's actually larger than the us
pentagon
why because there are immense economies
of scale
in the accumulation protection
um and management of data
and because obviously you know it's only
one there's a fixed cost to data there's
fixed cost to gathering and to storing
the data and then once it has been
gathered and stored there's essentially
zero cost to reusing it so once the data
has been collected on a single occasion
it is actually not worth duplicating
that data on another occasion except for
reasons for example of backup or privacy
so therefore there are massive economies
of scale and these kinds of data centers
excuse me these kinds of data centers
are exploiting that
but that's only half the equation
precisely because the data is amenable
to these massive economies of scale
the interpretation of the data can
actually fragment it turns out that four
ibm engineers in dublin running a hadoop
cluster can actually solve problems
perfectly efficiently if they merely
have access to those colossal data sets
one rather interesting example of that
at working in practice is a company
called kaggle which was founded
interestingly by an australian
kaggle is a basically a website that
curates contests
companies that have data problems post
their data usually in anonymized form
onto kaggle and then kaggle orchestrates
a contest by which anybody in the world
hackers scientists engineers researchers
phd students and so on
can try to solve the problem competing
for a prize
this particular graph shows a kaggle
contest where the client was allstate
allstate the
largest american property insurance
company property and casualty insurance
company this is automotive insurance
and they had an algorithm they also had
an algorithm which they developed over
many years from their actuarial work and
so on to predict from the application
form what is the expected loss rate from
a given potential customer and this
would be a basis for pricing the
insurance it would be a basis obviously
for maybe deciding whether or not to
take the insurance
the contest was to try to improve on
that algorithm and as you can see from
the graph that you're looking at what
happened over just a 12-week period was
that the um
these competing teams were able to
improve on allstate's original formula
by a factor of nearly 300 percent
now there's a very interesting footnote
to this rather spectacular story of very
rapid innovation
and that is that the
i did a back of the envelope calculation
to estimate the value of that
improvement to allstate
and basically what it is is obviously
rejecting bad deals and accepting good
ones maybe even pricing down to win the
good deals because you realize how good
they are
and the value at all state is
approximately 50 million dollars per
year
the prize that was won by the team that
came up with the best solution to the
problem was six thousand dollars
so you notice a rather radical asymmetry
between the value of being the smartest
person in the room and the value of
controlling the data it turns out that
it's control of the data that is what
makes or breaks is the is the basis of
reward um in this emerging world the
other thing that's very striking is that
the team that in a parallel be
particularly accurate a very parallel
problem problems in toxicology that were
solved uh for merck the winning team
didn't even know about the contest until
the last two weeks so they rather
hurriedly marshaled and entered the
contest and believed in it on their
first shot they won it a grand prize
again of something like five thousand
dollars but the interesting point was
that the problem was to predict the
toxic side effects of drugs something on
which merck had been working for 30
years
merck had been using
clinical trials at immense expense in
order to try to understand that the
question was can you just by looking at
the molecular structure predict those
kinds of toxic side effects the answer
is yes you can this team from the
university of toronto won the contest
not a single member of that team had a
background in chemistry or biology still
less toxicology
nobody knew a thing about the problem
what they knew about was how to analyze
data
so
in summary
we are moving from a world defined in
traditional business school business
school professor terms by things from a
world where things were defined by
things called value chains the value
chain is what it's the idea that a
business is a set of heterogeneous
activities and
the the physical product kind of goes
through those activities as it is
converted from raw material into final
good think of a factory think of a ford
motor assembly line raw materials go on
in cars come out the other and that
basic idea that a business is defined by
its value chain is fundamental to how
things like business strategy have been
thought about for 30 years
now in fact what is happening in
consequence of exactly the forces i
described is that that value chain is
breaking up
first of all transaction costs are
falling so that you can do these pieces
separately
secondly where there are economies of
scale as in data as in data processing
we're seeing colossal consolidation into
platform businesses such as google and
microsoft azure and so on and thirdly
where what matters is individual
initiative individual talent smarts
creativity experimentation as with all
of those grad students all trying to
solve the problems i've described that
fragments you don't even need to be part
of a corporation to do that people will
do that autonomously on small teams so
when you put these patterns together
what you see is a transposition of the
structure of whole industries you see a
shift from a vertically integrated
set of businesses all of which
essentially look alike to a horizontally
stratified set of businesses which are
very fundamentally different from each
other and which provides services to
each other we call that a stack because
that's the term that people would use in
software
stack is fundamental to the economics of
information just as the value chain the
traditional idea of physical flows is
fundamental to the economics of things
if there's one
phenomenon that defines the way our
world is changing it is that we are
moving from a world dominated by the
economics of things to a world dominated
by the economics of information thank
you very much