Confidence Interval Estimate for Single Proportion | Applied Biostatistics | BIO733_Topic118
Watch on YouTubeVideo summary
Bu modül, tek bir oransızlık için güven aralığı tahmini hesaplamayı ele almaktadır ve araştırmacıların popülasyon oranlarıyla ilgili birçok soruya yanıt bulmalarını sağlar. Örneğin, belirli bir tedavi alan hastaların ne kadarının iyileştiğini ölçmek veya nüfusun belirli bir hastalığa sahip olup olmadığını belirlemek gibi durumlar söz konusu olabilir. Bu tür durumlarda temel amaç, popülasyon oranını tahmin etmektir ve bu süreç, ortalama popülasyonu tahmin ederken izlenen yöntemle benzer şekilde ilerler. Bir örneklem popülasyondan çekilir ve popülasyon oranı olan pi için nokta tahmincisi olarak örnekleim oranı P şap kullanılır.
Güven aralığı ifadesi, tahmincinin standart hatası ile güvenilirlik faktörünün çarpımının toplanması ve çıkarılması yoluyla elde edilir. Binomial dağılımının parametreleri olan NP ve N1-P'nin her ikisinin de 5'ten büyük olması durumunda, P şap'ın örnekleme dağılımının normal olasılık dağılımına oldukça yakın olduğu kabul edilir. Bu koşul sağlandığında, kesin dağılım binomial olsa da yaklaşık olarak normal dağılıma yaklaşır ve bu nedenle Z değerlerinden oluşan standart normal dağılım kullanılır. Böylece, 100(1-alfa) yüzdüklü güven aralığı formülü ile hesaplanır; burada P şap tahminciyi, kök altındaki ifadenin ise standart hatayı temsil ederken, Z değeri güvenilirlik faktörüdür.
Örnek olarak, Pew Internet ve American Life Projesi'nin 2003'te yayınladığı verilere göre internet kullanıcılarının %18'inin deneysel tedaviler veya ilaçlar hakkında bilgi aramak için internet kullandığı belirtilmiştir. Bu örneklemde 1220 yetişkin internet kullanıcısı yer almış ve teleferik görüşmeler yoluyla veri toplanmıştır. %95 güven seviyesi için güvenilirlik faktörü Z tablosundan bulunarak 1.96 olarak belirlenmiş ve NP çarpımı 219.6 hesaplanarak bu değerin 5'ten büyük olduğu doğrulanmıştır. Bu sayede normal dağılım yaklaşımı geçerli bulunmuş ve formülde yer alan değerler kullanılarak alt güven sınırı %15.8, üst güven sınırı ise %20.2 olarak hesaplanmıştır.
Sonuç olarak, araştırmacılar bu sonuçlara dayanarak, deneysel tedavi veya ilaç bilgisi için internet kullanan yetişkin internet kullanıcılarının popülasyonunda %15.8 ile %20.2 arasında bir oran bulunduğunu %95 güvenle bekleyebilirler. Güven aralığı tahmini, sadece tahmin yapmakla kalmaz aynı zamanda popülasyon oranının anlamlılığını test etmek için de kullanılabilir; çünkü tekrarlanan örneklemelerde bu yöntemle oluşturulan aralıkların yaklaşık %95'i gerçek pi değerini içerecektir.
Read the full video transcript
This module talks about calculating the
confidence interval estimate
for
single proportion.
Many questions to the researchers relate
to the population proportion. For
example,
they want to measure that what
proportion of patients who received a
particular type of treatment recover.
They may be interested in knowing that
what proportion of population has
certain disease.
Or their interest could be that what
proportion of population is immune to a
certain disease.
In these type of cases, our goal is to
estimate population proportion.
So, we proceed in the same manner as
when estimating a population mean. A
sample is drawn from the population of
interest
and the sample proportion P cap
is calculated as a point estimator for
the population proportion that is
denoted by pi.
An expression for the confidence
interval can be obtained by the
estimator
plus minus reliability factor multiplied
by the standard error of the estimate.
As we are aware
that as both NP
which are the parameters of the binomial
distribution
and N 1 minus P
are greater than 5
we may consider the sampling
distribution of P hat to be quite close
to normal probability distribution.
Whereas the exact distribution for this
case would be binomial distribution. But
here, if
NP or N 1 minus P would be greater than
5
it approximates to the normal
distribution, hence we
we assume that
it's quite close to the normal
probability distribution and we use it
in that in this case.
So, when this condition is met, our
reliability factor is some value of Z
from the standard normal distribution.
Hence, 100 into 1 minus alpha percent
confidence interval estimate for
population proportion pi can be obtained
by the given expression.
Where P hat
is
estimator
and square root of P hat into 1 minus P
hat divided by N is the standard error
of estimate and Z 1 minus alpha by 2 is
the reliability factor that is obtained
from the standard normal distribution
here in this case.
Let's take an example.
The Pew Internet and American Life
Project reported in 2003
that 18% of the internet users have used
it to search for information regarding
experimental treatments or medicines.
The sample consisted of
1,220
adult
internet users.
And information was collected from
telephonic interviews.
We wish to construct a 95% confidence
interval for the proportion of internet
users in the sampled population
who have searched for information on
experimental treatment or medication.
Since here we are using 95% confidence
interval and we are already aware that
reliability factor should be calculated
from from the standard normal
distribution, so we are using this table
where we look look up for 0.9750
area and for 95% confidence level, the Z
alpha by 2 is 1.96.
Using this value as a reliability factor
and the given information that the
sample size is 1,220
and proportion
sample proportion to be equals to 18%
and here P will be 0.18.
We calculate NP and it turned out to be
219.6,
which is certainly greater than five.
And as this condition holds true, we can
say that we can
we can clearly approximate it to the
standard normal probability
distribution. And using all these values
and substituting into the formula for
the confidence interval, we get that
0.18
plus minus 1.96, that is the reliability
factor, and the standard error
calculated is 0.0110.
It gives us two values, the lower
confidence limit, which is 0.158,
and the upper confidence limit, which is
2.2.
Using these values, we are
95% confident that the population
proportion P
is between
0.158, that is lower confidence limit,
and 0.202, that is upper confidence
limit. Because in repeated sample,
about 95% of the interval constructed in
the manner of the
present single interval would include
the true value
of P.
On the basis of these results, we would
expect, with 95% confident, to find
somewhere between 15.8%
and 20.2% of adult internet users to
have used it for information on medicine
or experiment treatment. Confidence
interval estimate can also be used for
testing the significance of population
proportion.
Thank you.