Is India Behind in the AI Race? Policybazaar’s Chief Data Scientist Explains
Watch on YouTubeVideo summary
The interview features Santosh, Chief Data Scientist at Policybazaar, who provides a data-driven perspective on the current state of the Indian economy and its trajectory in the artificial intelligence landscape. While economic indicators such as insurance premiums across motor, health, and term sectors continue to rise consistently, Santosh highlights that true financial security for Indian households is still evolving due to gaps in financial literacy. He argues that many families hold misconceptions about their actual security levels and require further education to achieve robust financial stability. This foundational understanding is crucial before AI can be effectively leveraged to predict customer behavior or enhance economic outcomes, as the quality of data fed into models directly dictates their accuracy and fairness.
A significant portion of the discussion focuses on the practical application of AI in customer service, where Policybazaar utilizes large language models (LLMs) to handle over 60% of customer interactions via chat and voice agents. These digital assistants act as "digital twins" of human advisors, capable of resolving simple queries about policies and benefits around the clock while seamlessly transferring complex issues—such as payment disputes or cancellations—to human agents. Despite these advancements, Santosh points out that data prediction models in India often lag behind global standards primarily due to poor data quality rather than a lack of algorithmic sophistication. To build trust and ensure useful predictions, companies must rigorously verify the veracity and accuracy of their data, addressing biases before training models that serve diverse populations.
The conversation also delves into the limitations of current AI technologies, specifically the "black box" problem where models provide answers without explaining their reasoning process. Santosh clarifies that while Large Language Models excel at tasks like document reading and information extraction, they are not yet true Artificial General Intelligence (AGI) and function largely as automation tools trained on specific datasets rather than possessing genuine reasoning capabilities. Furthermore, he notes that Indian consumers often receive responses based on data from the US or UK, lacking deep contextual understanding of local nuances. Although India possesses unique advantages such as its massive population and linguistic diversity, the country currently lacks foundational models and control over computing infrastructure, which are essential for developing indigenous AI solutions like open-source alternatives to global giants.
Looking toward the future, Santosh outlines immediate priorities for Policybazaar, including the development of custom small language models (SLMs) fine-tuned for specific tasks and the automation of contract analysis to improve operational efficiency. He envisions a future where voice assistants supplement human advisors rather than replace them, streamlining processes like vehicle insurance issuance from days down to minutes by analyzing video uploads instantly. While predicting disease patterns using AI presents ethical challenges regarding insurance premiums and privacy, the technology is already saving lives in medical diagnostics. Ultimately, Santosh remains optimistic that India can catch up in the global AI race by leveraging its vast data resources and unique demographic profile, transforming these assets into innovations that address local needs even without owning the foundational hardware or software initially.
Read the full video transcript
Hi, I'm Rich Pava um editorial lead with
BW business world exchange for media and
martekchai.com
and I am at global fintech fest here in
Mumbai and I have with me uh Santosh but
he's the chief data scientist at
policybazar.com. So I I understand that
you're a data scientist but I want to
start with a an umbrella question. What
do you as a somebody who you know sees
data every day millions of data points
uh what do you think about the Indian
economy right now solely in terms of
what the data suggests?
If I'm looking at the data and broadly
what we see in our numbers and so on,
the numbers are have been consistently
going up. Okay.
>> And I don't think that has slowed down
in terms of the whether it's motor
insurance or uh you know uh health
insurance or the term insurance. It's
not necessarily slowed down. It's sort
of you know been along the expected
lines as such
>> but but you know we have been talking
about financial security of Indian
families right Indian households and a
lot of people say that families were
more secure loans were lesser when you
know let's say 10 years before 20 years
before the economy is growing but do you
think Indian households are also are
becoming more secure financially
I can only answer that with respect to
basically the financial literacy
>> I I think there's still a long way uh
for India in terms of the financial
literacy
Right? I mean people have misconceptions
about you know how financially secure
they are or what do they need to become
or secure their family and so on. So
there's still some way to go to achieve
that financial result.
>> Interesting. Right.
>> Uh you talked about data modeling right
and using artificial intelligence a
layer of putting that as a layer on data
and then actually predicting customer
behavior. Right.
>> Now I have a question. Um comes to my
mind as a very interesting recipe. Uh I
was scrolling my Instagram the other day
and I got an advertisement while
scrolling my stories of a car which was
about 20 lakh rupees and me as a target
customer and after six more stories I
get an advertisement of a car which was
60 lakh rupees and with me as a target
customer. Right. So one of the marketers
has gotten the data out. Right now we
are living in an era where intelligence
has you know moved leaps and bounds but
data prediction models are still lagging
behind. Why do you think that is the
case?
>> It is predominantly lagging behind
because of the quality of data that is
being fed to the models. and fed to the
model when I say if you have the wrong
set of data and you train the model on
that
>> then obviously the model is going to
create a bias
>> and that is the main reason why you see
all of this happening so any data which
is present with us will have to be kind
of checked for veracity and the you know
accuracy and so on right and the
enrichment of that data becomes very
very essential it has to be unbiased
when you're training the model only then
the model will actually start predicting
something which is useful
>> uh now you have also said in your
previous interviews that
>> good questions yes you have you have
said in your previous interviews that
you know AI handle more than half of
policy as customer chats now right uh
has this improved customer satisfaction
what are the what is the data that you
are seeing here because you know Indian
consumer is you largely knows what they
want to do right and all of us get a lot
of calls from different companies right
u but as a you know if you put the layer
of artificial intelligence on top of
these customer AI agents are talking to
customers how has that changed customer
satisfaction experience
>> so what we've done and consciously is to
actually keep the customer as the focus.
You don't want to kind of, you know, uh,
get the customer annoyed because he's
been talking to an AI agent. The AI
>> agent's role is to understand and solve
whatever possible in whatever possible
ways he could solve a customer's
problem.
>> And the role of an AI agent is to be
available 24/7
>> and you know whenever the customer needs
it, right? It is so you could think of
the AI agent as a digital twin of the of
a human adviser, right? which means
mostly the easy things that you know
people come there to ask and so on and
so forth it's able to solve you know for
example if he says I don't I want to
have a soft copy it is it's able to give
him that answer if he's got very
reasonably simple questions around the
policy the benefits etc it's able to
answer it doesn't need the you know
human to be there on the call at all
times
>> but it should also be able to hand over
the call when there is a complexity and
when there needs to be an interaction
between two humans
>> right that's so what we've seen is the
the chats whether it's through voice or
whether it's through uh you know text
text chats
So it's more than 50% it's possibly
about 60% now. Um now that number has
slightly gone up if you I mean compared
to last year but the satisfaction pretty
much has remained the same which is a
which is what it is telling us is it has
not impacted the customer in a wrong way
>> usage of a customer would still want to
speak to somebody who is more has more
human behavioral like a voice or a tone
or empathy right have you seen that
issue
>> in some customers largely what happens
is let's say when I buy a policy and I
want to get issuance of my policy I
would want that issuance you know I want
the problem to be solved M
>> what we have seen in the last 3 or 4
months at least is that people who like
the customers who booked the policy with
us are okay talking to agents as long as
but there are people for certain kind of
complex problems particularly payment
related issues refund cancellation any
kind of complex issues that's when the
uh we actually deliberately transfer it
to the human
>> now you have also written about the
blackbox problem in AI right what do you
mean by that
>> uh the explanability of of AI
>> okay
>> right so like when I say explainability
like it can give you u suggestions or it
can say something which is not
explainable and as to how it actually
arrived at that answer.
>> And so any large language model is
predominantly a black box, right? You
given certain inputs on certain prompts
and so on, it is it's forced to answer,
but it is it won't tell you how it got
to that reasoning.
>> Yeah.
>> Right. So although they've gotten better
with the with the language models and so
but it's still predominantly a black
box. So now how do you make that
explainable? That's I think very
important to build trust as well amongst
if you're going to use AI in the future
as well. And you talked about LLM and
that takes me to another interesting
question and I was talking to a CTO of a
very large company L& tech Ashish Kushu.
He told me something very interesting.
He said whatever we are seeing in terms
of LLMs today is just automation under
the garb of artificial intelligence. Now
why he says that is even LM work on a
preset of data. Right? If I'm asking a
question, it has a certain set of data
which it's using to answer that
question. It's not really using the
reasoning intelligence that it should.
Where do you think how far are we from a
model where AI will actually use real
intelligence and not just be trained on
certain data sets? Because as as an
Indian consu consumer sitting in India,
it is probably showing me answers
trained on data sets in the United
States of America or United Kingdom. It
is not really grasping real Indian
intelligence or Indian context. How far
are we from that? And this is an
umbrella question I know but what's
what's your view on?
>> There are there are two parts to it. I
think your question has two parts to it.
one is getting I'll come to the first
question where how far are we from that
real intelligence and so on I think that
question if it refers to the art I mean
artificial general intelligence then I
think we are some way away we are not
anywhere close that's my view at least
>> but astra has been launched
>> I know but AGI but again I think the
real you know intelligence is somewhere
away whether it is 3 years away whether
it's 5 years away we'll wait and watch
very difficult to answer that question
>> now coming to the Indian context and how
>> sorry we're at AGI so there was an an
experiment done by meta where it uh led
to AGI models start conversing to each
other and if I'm if I'm correct they
started talking about the end of
humanity because they didn't like
humanity and trying to get nuclear codes
out of the US president. This was four
five years ago and they had to
immediately uh turn the conversational
models down. So to the viewers who are
watching it we are still not there at as
we are speaking whatever you are using
is just automation under the G of
artificial intelligence. It's not
artificial intelligence but yes we will
hear more fun from the expert. Yes. So
>> uh automation under the garb of
intelligence.
>> Well so it depends on how you look at
it.
>> So the way I look at it is
There are certain tasks that we do as
humans which the older not older the
current set of technologies like typical
software development could not achieve
because those are deterministic nature.
The current set of you know AI models
help us to do a bit of that right. So
hence we find a lot of automation that
is happening. There are certain tasks
which with the current set of LLMs are
done much better. for example, reading
documents, extracting information, even
extracting, you know, information from
calls. Those are the kind of things it
does quite well, right? As compared to
basically dealing with numbers, looking
at numbers and you know, making sense of
numbers. LLMs are not necessarily great
at doing those those kind of stuff.
>> Now, this another very interesting
question that comes to my mind. There
might come a time when AI will probably
start predicting diseases in humans,
right? In terms of looking at certain
features, looking at certain reports, it
might be able to predict, right? How
does that how will it impact the
insurance sector? uh because if AI
starts predicting disease patterns, if
AI starts predicting what will happen to
a certain person 10 years from now, how
will I I I know it's a very u futuristic
question but uh if you can just throw
some light on it
>> it's very difficult to say uh
>> there are already certain
thoughts if I'm I mean it's very early
days I guess but there are already
certain thoughts which are moving in
this direction
>> which are moving in this direction for
example and uh AI starts reading your
face and basically based on that would I
actually kind of skip your medical
verification for that matter, right? Um
trying to make issuance faster, trying
to make life easier for the customer.
That's the initial direction. But it
this what you're essentially saying has
more implications. For example,
>> if AI says that
>> I I will get cancer 20 years down the
line, how will it impact insurance?
Those are very difficult questions to
answer and it is it is possible to do
it. I mean in the sense AI is already
doing it
>> and there was there was a use case where
this lady wasn't able the doctors
weren't able to find out a problem with
the lady and and the LLM she went to
helped it and she was saved.
>> Correct. very positive you know cases
there have been very bad cases as well.
So I think you have to take it with a
pinch of salt right now although I think
there are thoughts going in that
direction as to how do I ease the
customer onboarding
>> with the help of AI. uh what has AI made
the easiest at policy bazar if you talk
about 5 years ago from now on now
>> there have been multiple sort of
initiative that we did with the help of
AI which has helped the customer for
example um there is something called as
a break-in vehicle insurance wherein you
know your your uh policy has expired and
you want to renew it now let's say it
expired yesterday and you want to renew
the the the regulations say that all the
insurance mandated that you have to
record a video of your car and send it
to them
>> right now what what helps is they they
will I
they will look at the video and they'll
issue right and they look at multiple 40
50 parameters and they'll issue
>> typically this would have taken two to
three days even before the video if the
quality of the video is good or bad or
so on today I mean if the the moment a
customer uploads a video I can actually
tell him whether the video was correct
how many parameters are correct and if
everything was correct we could actually
potentially get that issued in the next
5 minutes
>> it's a significant uplift for the
customer
>> interesting yeah
>> right so that that's one pay as you
drive now all he has to do is basically
take the take his uh you know camera in
the full camera around the vehicle and
certain three or four points like you
photoometer, charging number etc etc the
three or four points
>> the issuance can be in just four five
minutes
>> there's nothing more that is required so
imagine the customer experience it's
also now the we've seen a decent amount
of uplift in the customer service side
which when it comes to basically
issuance again it is not the journey is
not done it's still early days but we do
see that you know signs of that customer
you know kind of
>> have you created your own SLMs also
>> we are in the process so we have some
SLM so there basically I can't say that
we have created a SLM for all of policy
we are in the process of doing that.
However, for certain type of use cases,
we use something called as a fine-tuned
model, right? So those are also SLMs,
but they they do maybe two or three
tasks very well. So those kind of models
are already in place. Now, if I'm
building SLM broadly for answering all
the questions, yeah, we are in the
process of doing that.
>> It also comes back to your other
question where can it understand the
context of India, Indian insurance and
so on that possibly is the
>> Yeah. So the SLM answer solves that
problem.
>> I'll come back to that question again.
any behavioral patterns, any insights
that you have for the consumer data that
you have at policy based
[laughter]
but I have given you another thing to go
back to your
>> I took the answer for that like the two
>> the two use cases
>> use cases we have to stick it together
some
>> interesting what next for cons
>> oh there's still a long way to go I mean
see ultimately if you look at it
>> what are you working on immediately the
next five six months
>> by by and large okay the first thing
that I think from a technology
perspective that SLM is an important one
um both from a cost perspective I mean
that is cost is an important factor. The
second one is actually understanding our
contracts, our data and so on. That's
one big thing that we are working on.
Right? The the the single biggest uh
area of focus is the customer service.
How do I improve customer service?
Again, it is not one project, it's not
two projects, it's a series of projects.
Like for example, I I will I'll have to
have a voice assistant. Not because
they're going to replace somebody human,
but it is to supplement the uh the our
own human advisor. So maybe you know the
voice agent can actually answer a
certain set of questions even the chat
bots. Then you know can can the internal
bots can can they help in you know
basically streaming streamlining the
entire operations for example talk AI
can automate that process it's only
going to help the customer. M
>> now my final question this is the most
serious question I have had Santo not
been a chief data scientist
>> what would he have been
>> nothing else I guess
>> nothing else
>> nothing else I've been doing yes I've
been doing this for the last 20 odd
years
>> no but some people say that they would
have been a golfer or a singer or that
way so what maybe what other line of um
work not work or hobby would you have
loved to
>> so uh my hobbies keep changing every
five years but what has stayed with me
apart from my love for mathematics as
such AI is also is mathematics um is
actually wildlife environment and
wildlife. So I do photography, I do some
bit of conservation and here and there.
>> Tell me also tell me something now. Um
India is a country, right? We were not
late to the AI race, right? We might
have been led to various other uh you
know races globally but everybody had
the same starting point when it came to
the AI race. we still somehow lagged
behind, right? We haven't really
produced an open eye of ours or a a
deepseek of ours, right? Yes, we have
GBD again which is not B2C, right? It's
more of a B2B platform. Where have where
have we lagged? Is it the computational
possibilities? Is the data possibility?
Uh is is that the lack of grit that
people that founders in India have?
Where have we lacked in terms of
actually giving the world an MLM? So I
think the
you you kind of answered it yourself in
the sense that we've lacked in the
foundational part of the AI race which
means we don't have any foundational
model which has been built in India nor
are we kind of controlling the compute
right possibly the the data centers and
so on yeah that may be still there um
it's still okay I mean so think about it
right the hardware is not with us even
the software is not with us at least as
of now but it's I I think we're still
early in the in the race in the sense
while we may have been lagging behind
there we we might be able to catch might
be able to catch up.
>> I think on that note we end this
interview. Thank you so much for giving
your time. I think it's some very
interesting insights.
>> There's a lot of interesting things that
India could go ahead for example no
other country has like 25 30 languages.
>> Yes. No other country has 1.4 billion
people also. India sitting on just all
kind of data.
>> All kind of data and there could be many
innovation that can come even with the
even if you don't own the you know
>> and I think that makes your job even
exciting as a data scientist. I mean I'd
love to you know and as a somebody who
loves mathematics you would love to sit
on numbers sit on data every day.
Interesting. It's interesting to see
people who love their job. Uh to all
those who are watching this interview, I
hope you gained something from it. Uh
Santosh was very very good and
insightful with his answers. Thank you
so much for giving us nice