Binomial sum variance inequality

The binomial sum variance inequality states that the variance of the sum of binomially distributedrandom variables will always be less than or equal to the variance of a binomial variable with the same n and p parameters. In probability theory and statistics, the sum of independent binomial random variables is itself a binomial random variable if all the component variables share the same success probability. If success probabilities differ, the probability distribution of the sum is not binomial.[1] The lack of uniformity in success probabilities across independent trials leads to a smaller variance.[2][3][4][5][6] and is a special case of a more general theorem involving the expected value of convex functions.[7] In some statistical applications, the standard binomial variance estimator can be used even if the component probabilities differ, though with a variance estimate that has an upward bias.

Inequality statement

Consider the sum, Z, of two independent binomial random variables, X ~ B(m0, p0) and Y ~ B(m1, p1), where Z = X + Y. Then, the variance of Z is less than or equal to its variance under the assumption that p0 = p1 = p¯{\displaystyle {\bar {p}}}, that is, if Z had a binomial distribution with the success probability equal to the average of X and Y 's probabilities.[8] Symbolically, Var(Z)E[Z](1E[Z]m0+m1){\displaystyle Var(Z)\leqslant E[Z](1-{\tfrac {E[Z]}{m_{0}+m_{1}}})}.

Proof

We wish to prove that

Var(Z)E[Z](1E[Z]m0+m1){\displaystyle Var(Z)\leqslant E[Z](1-{\frac {E[Z]}{m_{0}+m_{1}}})}

We will prove this inequality by finding an expression for Var(Z) and substituting it on the left-hand side, then showing that the inequality always holds.

إذا كان المتغير العشوائي Z يتبع التوزيع الثنائي بمعاملات n و p ، فإن القيمة المتوقعة لـ Z تُعطى بالعلاقة E[ Z ] = np، وتباين Z يُعطى بالعلاقة Var[ Z ] = np (1 – p ). وبفرض n = m₀ + m₁ ، وبالتعويض عن np بالعلاقة E[ Z ] نحصل على

Vأر(Z)=هـ[Z](1-هـ[Z]م0+م1){\displaystyle Var(Z)=E[Z](1-{\frac {E[Z]}{m_{0}+m_{1}}})}

المتغيران العشوائيان X و Y مستقلان، لذا فإن تباين المجموع يساوي مجموع التباينات ، أي

Vأر(Z)=هـ[X](1-هـ[X]م0)+هـ[Y](1-هـ[Y]م1){\displaystyle Var(Z)=E[X](1-{\frac {E[X]}{m_{0}}})+E[Y](1-{\frac {E[Y]}{m_{1}}})}

لإثبات النظرية، يكفي إثبات أن

هـ[X](1-هـ[X]م0)+هـ[Y](1-هـ[Y]م1)هـ[Z](1-هـ[Z]م0+م1){\displaystyle E[X](1-{\frac {E[X]}{m_{0}}})+E[Y](1-{\frac {E[Y]}{m_{1}}})\leqslant E[Z](1-{\frac {E[Z]}{m_{0}+m_{1}}})}

باستبدال E[ X ] + E[ Y ] بـ E[ Z ] نحصل على

هـ[X](1-هـ[X]م0)+هـ[Y](1-هـ[Y]م1)(هـ[X]+هـ[Y])(1-هـ[X]+هـ[Y]م0+م1){\displaystyle E[X](1-{\frac {E[X]}{m_{0}}})+E[Y](1-{\frac {E[Y]}{m_{1}}})\leqslant (E[X]+E[Y])(1-{\frac {E[X]+E[Y]}{m_{0}+m_{1}}})}

بضرب الأقواس وطرح E[X] + E[Y] من كلا الطرفين نحصل على:

-هـ[X]2م0-هـ[Y]2م1-(هـ[X]+هـ[Y])2م0+م1{\displaystyle -{\frac {E[X]^{2}}{m_{0}}}-{\frac {E[Y]^{2}}{m_{1}}}\leqslant -{\frac {(E[X]+E[Y])^{2}}{m_{0}+m_{1}}}}

بضرب الأقواس نحصل على

هـ[X]-هـ[X]2م0+هـ[Y]-هـ[Y]2م1هـ[X]+هـ[Y]-(هـ[X]+هـ[Y])2م0+م1{\displaystyle E[X]-{\frac {E[X]^{2}}{m_{0}}}+E[Y]-{\frac {E[Y]^{2}}{m_{1}}}\leqslant E[X]+E[Y]-{\frac {(E[X]+E[Y])^{2}}{m_{0}+m_{1}}}}

بطرح E[X] و E[Y] من كلا الطرفين وعكس المتباينة نحصل على

هـ[X]2م0+هـ[Y]2م1(هـ[X]+هـ[Y])2م0+م1{\displaystyle {\frac {E[X]^{2}}{m_{0}}}+{\frac {E[Y]^{2}}{m_{1}}}\geqslant {\frac {(E[X]+E[Y])^{2}}{m_{0}+m_{1}}}}

يؤدي توسيع الجانب الأيمن إلى

هـ[X]2م0+هـ[Y]2م1هـ[X]2+2هـ[X]هـ[Y]+هـ[Y]2م0+م1{\displaystyle {\frac {E[X]^{2}}{m_{0}}}+{\frac {E[Y]^{2}}{m_{1}}}\geqslant {\frac {E[X]^{2}+2E[X]E[Y]+E[Y]^{2}}{m_{0}+m_{1}}}}

الضرب فيم0م1(م0+م1){\displaystyle m_{0}m_{1}(m_{0}+m_{1})}العائد

(م0م1+م12)هـ[X]2+(م02+م0م1)هـ[Y]2م0م1(هـ[X]2+2هـ[X]هـ[Y]+هـ[Y]]2){\displaystyle (m_{0}m_{1}+{m_{1}}^{2}){E[X]^{2}}+({m_{0}}^{2}+m_{0}m_{1}){E[Y]^{2}}\geqslant m_{0}m_{1}({E[X]}^{2}+2E[X]E[Y]+{E[Y]]^{2}})}

بطرح الطرف الأيمن نحصل على العلاقة

م12هـ[X]2-2م0م1هـ[X]هـ[Y]+م02هـ[Y]20{\displaystyle {m_{1}}^{2}{E[X]^{2}}-2m_{0}m_{1}E[X]E[Y]+{m_{0}}^{2}{E[Y]^{2}}\geqslant 0}

أو ما يعادل ذلك

(م1هـ[X]-م0هـ[Y])20{\displaystyle (m_{1}E[X]-m_{0}E[Y])^{2}\geqslant 0}

مربع أي عدد حقيقي يكون دائمًا أكبر من أو يساوي الصفر، لذا فإن هذا ينطبق على جميع التوزيعات الثنائية المستقلة التي يمكن أن تأخذها المتغيرات X و Y. وهذا يكفي لإثبات النظرية.

على الرغم من أن هذا البرهان قد طُوِّر لمجموع متغيرين، إلا أنه يُمكن تعميمه بسهولة ليشمل أكثر من متغيرين. إضافةً إلى ذلك، إذا كانت احتمالات النجاح الفردية معروفة، فإن التباين يكون معروفًا بأنه يأخذ الشكل [ 6 ].

متغير(Z)=نص¯(1-ص¯)-نs2،{\displaystyle \operatorname {Var} (Z)=n{\bar {p}}(1- {\bar {p}})-ns^{2},}

أينص¯{\displaystyle {\bar {p}}}هو متوسط ​​الاحتمال وs2=1نأنا=1ن(صأنا-ص¯)2{\displaystyle s^{2}={\frac {1}{n}}\sum _{i=1}^{n}(p_{i}-{\bar {p}})^{2}}يشير هذا التعبير أيضًا إلى أن التباين يكون دائمًا أقل من تباين التوزيع ذي الحدين.ص=ص¯{\displaystyle p={\bar {p}}}، لأن التعبير القياسي للتباين ينخفض ​​بمقدار ns 2 ، وهو عدد موجب.

التطبيقات

The inequality can be useful in the context of multiple testing, where many statistical hypothesis tests are conducted within a particular study. Each test can be treated as a Bernoulli variable with a success probability p. Consider the total number of positive tests as a random variable denoted by S. This quantity is important in the estimation of false discovery rates (FDR), which quantify uncertainty in the test results. If the null hypothesis is true for some tests and the alternative hypothesis is true for other tests, then success probabilities are likely to differ between these two groups. However, the variance inequality theorem states that if the tests are independent, the variance of S will be no greater than it would be under a binomial distribution.

References

  1. Butler, Ken; Stephens, Michael (1993). "The distribution of a sum of binomial random variables"(PDF). Technical Report No. 467. Department of Statistics, Stanford University. Archived(PDF) from the original on April 11, 2021.
  2. Nedelman, J and Wallenius, T., 1986. Bernoulli trials, Poisson trials, surprising variances, and Jensen’s Inequality. The American Statistician, 40(4):286–289.
  3. Feller, W. 1968. An introduction to probability theory and its applications (Vol. 1, 3rd ed.). New York: John Wiley.
  4. Johnson, N. L. and Kotz, S. 1969. Discrete distributions. New York: John Wiley
  5. Kendall, M. and Stuart, A. 1977. The advanced theory of statistics. New York: Macmillan.
  6. 12Drezner, Zvi; Farnum, Nicholas (1993). "A generalized binomial distribution". Communications in Statistics - Theory and Methods. 22 (11): 3051–3063. doi:10.1080/03610929308831202. ISSN 0361-0926.
  7. Hoeffding, W. 1956. On the distribution of the number of successes in independent trials. Annals of Mathematical Statistics (27):713–721.
  8. Millstein, J.; Volfson, D. (2013). "Computationally efficient permutation-based confidence interval estimation for tail-area FDR". Frontiers in Genetics. 4 (179): 1–11. doi:10.3389/fgene.2013.00179. PMC 3775454. PMID 24062767.