Submind YouTube summaries
Thumbnail for Growing agriculture with big data: Philip Evans, Boston Consulting Group

Growing agriculture with big data: Philip Evans, Boston Consulting Group

Watch on YouTube

Video summary

Philip Evans begins his presentation by drawing a parallel between Jorge Luis Borges' story of an infinite, one-to-one map and the current trajectory of technology and data. He argues that we are approaching a similar state where our understanding of the world becomes so detailed and comprehensive that it effectively mirrors reality itself. To illustrate this, he points to Google's self-driving cars, which utilize vast datasets to perceive their environment with superhuman accuracy, identifying cyclists, traffic lights, and road conditions better than any human driver. This capability is built upon three fundamental mechanisms: first, the use of external data like satellite imagery to create precise agricultural maps; second, the ability of objects to describe themselves through metadata, such as cell phones revealing population movement patterns which allowed researchers to optimize bus routes in Abidjan and reduce commute times by ten percent without adding a single vehicle; and third, the recent breakthroughs in artificial intelligence where machines can now perceive and understand their surroundings, eventually reaching a point where they can generate grammatically correct summaries of video content. These advancements are driven by four dramatic technological shifts: the Internet of Things, which has proliferated sensors to an estimated 168 per person globally; the explosion of big data, with the world's data stock doubling every two years; the rise of artificial intelligence through neural networks; and mobility, which allows insights to be delivered exactly where they are needed. Evans emphasizes that these technologies have only emerged in the last seven or eight years, creating a new macro-pattern where billions of devices connect via IP addresses to exchange information at near-zero latency. While institutional barriers like privacy concerns exist, the technical architecture supports a unified global data set. This shift is further enabled by massive economies of scale found in centralized data centers, which store data once for reuse, while the interpretation and analysis of that data fragment out to small teams and individuals who can solve complex problems autonomously. The presentation highlights how this new landscape is disrupting traditional business models centered on linear value chains, where physical products move through a series of activities from raw material to final good. Instead, Evans describes a transition toward a horizontally stratified "stack" of businesses that are fundamentally different yet interconnected, serving each other's needs. This fragmentation occurs because transaction costs are falling, allowing separate pieces of the value chain to operate independently, while economies of scale in data processing lead to consolidation among platform giants like Google and Microsoft Azure. Meanwhile, individual initiative, creativity, and experimentation drive innovation from the bottom up, as seen in platforms like Kaggle where diverse participants compete to solve specific data problems for companies like Allstate or Merck. In these contests, teams with no background in chemistry or biology have successfully predicted toxic side effects of drugs by simply mastering data analysis, proving that raw data control is more valuable than being the smartest person in the room. Ultimately, Evans concludes that the defining phenomenon of our changing world is the shift from an economy dominated by the economics of things to one dominated by the economics of information. This transition represents a profound restructuring of industries, moving away from vertically integrated models toward a dynamic ecosystem where services are provided across a digital stack. The value of data control has become the primary basis for reward and competitive advantage, enabling organizations to reject bad deals, accept good ones, and optimize operations with unprecedented efficiency. As we stand on the brink of this new era, the ability to harness these four technologies will determine success, marking a departure from traditional strategies that relied on physical flows toward a future defined by information processing and decentralized innovation.
Read the full video transcript
good afternoon it's a great uh it's a great pleasure to be here um i am neither an australian nor an agronomist so i'm in many ways uniquely unqualified to be saying anything to this august group what i would like to offer for what it's worth is a brief perspective on how data and information technology more broadly have been changing in recent years and what are some of the models and frameworks that we see evolving now if i can get to my presentation therefore there we go um i'd like to start if i may with a story um jorge luis borges the argentinian poet and novelist wrote a very famous short story it doesn't even have a title and in this short story he recounted the tale of a lost kingdom this was a kingdom somewhere you can imagine in the andes where the aristocracy were obsessed with maps and however each map that they cast was deemed unsatisfactory it wasn't good enough so they launched a new project for an even more ambitious map until finally they launched the ultimate mapping project a map of their kingdom on a scale of one to one and according to borges if you visit this kingdom to this day you will find fragments of this failed project in various parts of the terrain now book has obviously had a lot of philosophical ideas behind this and it was a tale of human hubris and so on but what i would like to suggest is that in many ways that idea the idea of a map on a scale of one to one is indeed precisely the image that we should have of where technology in general and data in particular are taking us and to illustrate the point consider the following this is what the google car sees uh in the inside on the left you have the view from the dashboard but in the larger part of the screen what you see is the car's detection not just of the roads and the bridges and the traffic lanes but also of other vehicles also of cyclists of traffic lights and so forth indeed there's a wonderful moment just coming up where the car faces the challenge of trying to overtake a typically wayward erratic and irrational and suicidal northern californian cyclist who cannot make up his mind whether he is turning left or turning right and notwithstanding his random flapping of his hands nonetheless the car successfully passes him indeed to this day uh the uh google car has only experienced two accidents uh one of which was when it was rammed from the rear while parked and the other which was when it was being driven manually by a google engineer now what are the components of the technology that we see here what i'd like to suggest is that it's worth looking at this because those components are actually quite general one of them is very simply and obviously the existence of a general comprehensive and accurate map of the terrain in google's case of course this would be the google maps map of the roads but in fact we have many examples of that and an obvious agricultural example would be the use of satellite data in order to know not just the uh the bounds of the terrain but to understand things like the humidity or the ph factors uh or off the soil which enables all sorts of um uh aspects of precision agriculture that would not have been possible before but there is actually i would argue a at least two other mechanisms that are involved that are less obvious one of which is that the map is generated from things describing themselves this is a classic example this is a view of the united states um taken from a satellite at night what you see is the pattern of light from cities railroad tracks roadways um airplanes and so forth and very obviously you can see the shape of the country notoriously if we go to africa it is the dark continent the continent where um uh such uh infrastructure is conspicuous by its absence but not entirely and that is in many ways what makes the story interesting orange the telecommunications company the french telecommunications company is the monopoly provider of a cell phone service in the ivory coast what orange did was they they collected meta data on how people use the telephones in the cell phones in the ivory coast nine months worth of data for the entire population and by doing that they were able to create networks like this where you could see the movements of people because of course the cell phone registers with the nearest tower so that the telephone company actually knows where it is and who speaks to whom and even in fact what language they speak in what orange did having collected this data at very low cost was simply to publish it they anonymized it and they published it and they said to the world's researchers go at it see what information see what patterns you can find in this data and interestingly 82 papers were written by researchers and academics around the world using this unique data set as a source of insight one of those papers came up with the following rather elegant analysis some researchers at ibm looked at the um pattern of commuting in aberjan which is the largest city in the ivory coast now they weren't looking at where people get on the bus and get off the bus which the bus company kind of knew from its own information what they were looking at was where people started their journey and where people finished their journey and that enabled them to think therefore of a very large optimization problem which is you've got a finite number of buses you've got a population who need to on a daily basis get to and from two locations their home and their work what is the optimal allocation of bus routes that will minimize the time that it takes the average citizen of the town to get to and from work and they formulated this problem the mathematics is quite trivial trivial the computation of course is horrendous they formulated this problem they ran it on a hadoop cluster for two days and sure enough they got an answer and the answer was a reconfiguration of the abidjan bus system which without adding a single bus shave ten percent off the average commute time now you think of the effective impact on as it were the gdp of the country to reduce everybody's work day by 10 and the impact is truly extraordinary why was that possible it was possible because of a very very large data set which had not been available before and because of a few smart people and access to some very powerful machines but only for a few days in scale terms the really difficult bit was the data once the data was available all sorts of things became possible but that's just the second mechanism by which these maps can be created the maps created through objects describing themselves in this case the cell phone declaring its location there is a third mechanism which is the most recent and in many ways the most profound and that is the ability of machines to actually perceive and understand their environment there was an immense breakthrough in something called convolutional neural networks about five years ago again based on big data that made it possible for machines to recognize for example objects in a photograph and an annual competition held by stanford university to solve this particular problem was won by a team from microsoft this year with a solution that is more accurate in predicting the identifying the objects in a photograph than is the average human but that is just the beginning look at this this is a video that was created about six weeks ago by a graduate student in amsterdam what he did was he was running a piece of software on his macintosh that is using neural network technology not just to identify objects but to caption that is to say to create grammatical sentences describing what is going on and what you see he what you see is is that he is walking down the street and the the camera embedded in his laptop is uh generating images which the software is interpreting now you'll observe very obviously that about half of these captions are incorrect this is about where the identification of photographs was about two or three years ago and everybody knows full well that at the rate at which this is advancing within two or three years we will have instead of a boat is parked on the side of the water we will have a boat is moored at the side of the water because the machine by then will have learned the correct grammar it is confidently predicted that within five years a machine will be able to watch a youtube video and um generate a grammatical sensible one paragraph summary of what is the story of that video that is how near machine learning is so when we stand back from these phenomena we actually have these three different ways that these maps can be created what underlies them is four dramatically new technologies one is the internet of things the vast proliferation of uh sensors in the world it's estimated today that there are 168 sensors for every man woman and child on the planet the cost of these sensors is dropping by an order of magnitude every five or six years it's confidently expected that that number will multiply by a factor of 100 in the next 10 years secondly we have big data and i'm sure you have heard of the statistic frequently quoted that the world stock of data is doubling every two years thirdly we have artificial intelligence data as jackie said earlier data is worthless without the insight that interprets it and it is artificial intelligence that has been subject to many major breakthroughs the one particular one about neural networks in just the past five years and then finally we have mobility we have the fact that the information and the insight that is being generated can now be used at the point where it is needed at the point where it is relevant most obviously in the case of consumers in the form of information delivered to your smartphone and the number of phones in the world is now greater than the number of people there are two and a half billion smartphones in the world within five years they will all be smartphones this is just a matter of time now a couple of points about this one this is one big system this isn't a pattern that replicates itself millions of times the internet of things is billions of devices connected to each other via ip addresses and the web the big big data when we talk about big data what we're talking about is data that has that is on servers excuse me or on laptops all of which again have ip addresses meaning that they are connected to each other meaning that in principle at almost zero costs and at almost zero latency they can exchange information now they may not because of course different people own it there's issues of privacy um uh intellectual property all sorts of reasons why institutionally it doesn't happen but from a technical point of view it is one data set in similar fashion very obviously the phones are all connected to the global telecommunications network so what we're talking about is the emergence of this macro pattern the other key point to emphasize is how recent all of these things were if you turn the clock back seven or eight years nobody was talking about any of them now in that world what are the institutions that are needed what are the institutions that make it possible and the answer here is a little paradoxical because it's an alliance between the very big and the very small one component of this architecture is data centers things like cloud computing of which i'm sure you've heard this is a and not a typical data center belonging to google the world's largest data center which is under construction in china is the size of 120 football fields if you can imagine a building of that size it's actually larger than the us pentagon why because there are immense economies of scale in the accumulation protection um and management of data and because obviously you know it's only one there's a fixed cost to data there's fixed cost to gathering and to storing the data and then once it has been gathered and stored there's essentially zero cost to reusing it so once the data has been collected on a single occasion it is actually not worth duplicating that data on another occasion except for reasons for example of backup or privacy so therefore there are massive economies of scale and these kinds of data centers excuse me these kinds of data centers are exploiting that but that's only half the equation precisely because the data is amenable to these massive economies of scale the interpretation of the data can actually fragment it turns out that four ibm engineers in dublin running a hadoop cluster can actually solve problems perfectly efficiently if they merely have access to those colossal data sets one rather interesting example of that at working in practice is a company called kaggle which was founded interestingly by an australian kaggle is a basically a website that curates contests companies that have data problems post their data usually in anonymized form onto kaggle and then kaggle orchestrates a contest by which anybody in the world hackers scientists engineers researchers phd students and so on can try to solve the problem competing for a prize this particular graph shows a kaggle contest where the client was allstate allstate the largest american property insurance company property and casualty insurance company this is automotive insurance and they had an algorithm they also had an algorithm which they developed over many years from their actuarial work and so on to predict from the application form what is the expected loss rate from a given potential customer and this would be a basis for pricing the insurance it would be a basis obviously for maybe deciding whether or not to take the insurance the contest was to try to improve on that algorithm and as you can see from the graph that you're looking at what happened over just a 12-week period was that the um these competing teams were able to improve on allstate's original formula by a factor of nearly 300 percent now there's a very interesting footnote to this rather spectacular story of very rapid innovation and that is that the i did a back of the envelope calculation to estimate the value of that improvement to allstate and basically what it is is obviously rejecting bad deals and accepting good ones maybe even pricing down to win the good deals because you realize how good they are and the value at all state is approximately 50 million dollars per year the prize that was won by the team that came up with the best solution to the problem was six thousand dollars so you notice a rather radical asymmetry between the value of being the smartest person in the room and the value of controlling the data it turns out that it's control of the data that is what makes or breaks is the is the basis of reward um in this emerging world the other thing that's very striking is that the team that in a parallel be particularly accurate a very parallel problem problems in toxicology that were solved uh for merck the winning team didn't even know about the contest until the last two weeks so they rather hurriedly marshaled and entered the contest and believed in it on their first shot they won it a grand prize again of something like five thousand dollars but the interesting point was that the problem was to predict the toxic side effects of drugs something on which merck had been working for 30 years merck had been using clinical trials at immense expense in order to try to understand that the question was can you just by looking at the molecular structure predict those kinds of toxic side effects the answer is yes you can this team from the university of toronto won the contest not a single member of that team had a background in chemistry or biology still less toxicology nobody knew a thing about the problem what they knew about was how to analyze data so in summary we are moving from a world defined in traditional business school business school professor terms by things from a world where things were defined by things called value chains the value chain is what it's the idea that a business is a set of heterogeneous activities and the the physical product kind of goes through those activities as it is converted from raw material into final good think of a factory think of a ford motor assembly line raw materials go on in cars come out the other and that basic idea that a business is defined by its value chain is fundamental to how things like business strategy have been thought about for 30 years now in fact what is happening in consequence of exactly the forces i described is that that value chain is breaking up first of all transaction costs are falling so that you can do these pieces separately secondly where there are economies of scale as in data as in data processing we're seeing colossal consolidation into platform businesses such as google and microsoft azure and so on and thirdly where what matters is individual initiative individual talent smarts creativity experimentation as with all of those grad students all trying to solve the problems i've described that fragments you don't even need to be part of a corporation to do that people will do that autonomously on small teams so when you put these patterns together what you see is a transposition of the structure of whole industries you see a shift from a vertically integrated set of businesses all of which essentially look alike to a horizontally stratified set of businesses which are very fundamentally different from each other and which provides services to each other we call that a stack because that's the term that people would use in software stack is fundamental to the economics of information just as the value chain the traditional idea of physical flows is fundamental to the economics of things if there's one phenomenon that defines the way our world is changing it is that we are moving from a world dominated by the economics of things to a world dominated by the economics of information thank you very much