常见统计概率分布实现(代码)
随机变量(Random Variable) 密度函数(Density Functions) 伯努利分布(Bernoulli Distribution) 二项式分布(Binomial Distribution) 均匀分布(Uniform Distribution) 泊松分布(Poisson Distribution) 正态分布(Normal Distribution) 长尾分布(Long-Tailed Distribution) 学生 t 检验分布(Student’s t-test Distribution) 对数正态分布(Lognormal Distribution) 指数分布(Exponential Distribution) 威布尔分布(Weibull Distribution) 伽马分布(Gamma Distribution) 卡方分布(Chi-square Distribution) 中心极限定理(Central Limit Theorem)
1. 随机变量
离散随机变量
2. 密度函数
3. 离散分布
import seaborn as snsfrom scipy.stats import bernoulli# 单一观察值# 生成数据 (1000 points, possible outs: 1 or 0, probability: 50% for each)data = bernoulli.rvs(size=1000,p=0.5)# 绘制图形ax = sns.distplot(data_bern,kde=False,hist_kws={"linewidth": 10,'alpha':1})ax.set(xlabel='Bernouli', ylabel='freq')
二项式分布
伯努利分布是针对单个观测结果的。多个伯努利观测结果会产生二项式分布。例如,连续抛掷硬币。
试验是相互独立的。一个尝试的结果不会影响下一个。
import matplotlib.pyplot as pltfrom scipy.stats import binomn = 20# 实验次数p = 0.5# 成功的概率r = list(range(n + 1))# the number of success# pmf值pmf_list = [binom.pmf(r_i, n, p) for r_i in r ]# 绘图plt.bar(r, pmf_list)plt.show()
现在这次,你有一枚欺诈硬币。你知道这个硬币正面向上的概率是 0.7。因此,p = 0.7。
该分布显示出成功结果数量增加的概率增加。
均匀分布
data = np.random.uniform(1, 6, 6000)Poisson 分布
它是与事件在给定时间间隔内发生频率相关的分布。
import matplotlib.pyplot as pltfrom scipy.statsimport poissonr = range(0,11)# 呼叫次数lambda_val = 4# 均值# 概率值data = poisson.pmf(r, lambda_val)# 绘图fig, ax = plt.subplots(1, 1, figsize=(8, 6))ax.plot(r, data, 'bo', ms=8, label='poisson')plt.ylabel("Probability", fontsize="12")plt.xlabel("# Calls", fontsize="12")plt.title("Poisson Distribution", fontsize="16")ax.vlines(r, 0, data, colors='r', lw=5, alpha=0.5)
4. 连续分布
正态分布
最著名和最常见的分布(也称为高斯分布),是一种钟形曲线。它可以通过均值和标准差定义。正态分布的期望值是均值。
曲线对称。均值、中位数和众数相等。曲线下总面积为 1。
大约 68%的值落在一个标准差范围内。~95% 落在两个标准差范围内,~98.7% 落在三个标准差范围内。
import scipymean = 0standard_deviation = 5x_values = np. arange(-30, 30, 0.1)y_values = scipy.stats.norm(mean, standard_deviation)plt.plot(x_values, y_values. pdf(x_values))
正态分布的概率密度函数为:
QQ 图
import numpy as npimport statsmodels.api as smpoints = np.random.normal(0, 1, 1000)fig = sm.qqplot(points, line ='45')plt.show()
长尾分布
import matplotlib.pyplot as pltfrom scipy.stats import skewnormdef generate_skew_data(n: int, max_val: int, skewness: int):# Skewnorm functionrandom = skewnorm.rvs(a = skewness,loc=max_val, size=n)plt.hist(random,30,density=True, color = 'red', alpha=0.1)plt.show()generate_skew_data(1000, 100, -5) # negative (-5)-> 左偏分布
generate_skew_data(1000, 100, 5) # positive (5)-> 右偏分布学生 t 检验分布
正态但有尾(更厚、更长)。
t 分布是具有较厚尾部的正态分布。如果可用数据较少(约 30 个),则使用 t 分布代替正态分布。
在 t 分布中,自由度变量也被考虑在内。根据自由度和置信水平在 t 分布表中找到关键的 t 值。这些值用于假设检验。
t 分布表情移步:https://www.sjsu.edu/faculty/gerstman/StatPrimer/t-table.pdf。
对数正态分布
import numpy as npimport matplotlib.pyplot as pltfrom scipy import statsX = np.linspace(0, 6, 1500)std = 1mean = 0lognorm_distribution = stats.lognorm([std], loc=mean)lognorm_distribution_pdf = lognorm_distribution.pdf(X)fig, ax = plt.subplots(figsize=(8, 5))plt.plot(X, lognorm_distribution_pdf, label="μ=0, σ=1")ax.set_xticks(np.arange(min(X), max(X)))plt.title("Lognormal Distribution")plt.legend()plt.show()
指数分布
from scipy.stats import exponimport matplotlib.pyplot as pltx = expon.rvs(scale=2, size=10000) # 2 calls# 绘图plt.hist(x, density=True, edgecolor='black')
韦伯分布
它是指时间间隔是可变的而不是固定的情况下使用的指数分布的扩展。在 Weibull 分布中,时间间隔被允许动态变化。
import matplotlib.pyplot as pltx = np.arange(1,100.)/50.def weib(x,n,a):return (a / n) * (x / n)**(a - 1) * np.exp(-(x / n)**a)count, bins, ignored = plt.hist(np.random.weibull(5.,1000))x = np.arange(1,100.)/50.scale = count.max()/weib(x, 1., 5.).max()plt.plot(x, weib(x, 1., 5.)*scale)plt.show()å
Gamma 分布
import numpy as npimport scipy.stats as statsimport matplotlib.pyplot as plt#Gamma distributionsx = np.linspace(0, 60, 1000)y1 = stats.gamma.pdf(x, a=5, scale=3)y2 = stats.gamma.pdf(x, a=2, scale=5)y3 = stats.gamma.pdf(x, a=4, scale=2)# plotsplt.plot(x, y1, label='shape=5, scale=3')plt.plot(x, y2, label='shape=2, scale=5')plt.plot(x, y3, label='shape=4, scale=2')#add legendplt.legend()#displayplotplt.show()
Gamma 分布
# x轴范围0-10,步长0.25X = np.arange(0, 10, 0.25)plt.subplots(figsize=(8, 5))plt.plot(X, stats.chi2.pdf(X, df=1), label="1 dof")plt.plot(X, stats.chi2.pdf(X, df=2), label="2 dof")plt.plot(X, stats.chi2.pdf(X, df=3), label="3 dof")plt.title("Chi-squared Distribution")plt.legend()plt.show()
中心极限定理
我们可以从任何分布(离散或连续)开始,从人群中收集样本并记录这些样本的平均值。随着我们继续采样,我们会注意到平均值的分布正在慢慢形成正态分布。
来源:算法进阶, 仅用于传递和分享更多信息,并不代表本平台赞同其观点和对其真实性负责,版权归原作者所有,如有侵权请联系我们删除。-END- 推荐阅读:
10W字《R ggplot2可视化教程1.0》来了! 详解Python列表推导式|迭代器|生成器|匿名函数 Jupyter Notebook的16个超棒插件! 临床WGS/WES/Gene Panel/Single gene异同 一图胜千言,超形象图解NumPy教程! 那些神经网络可视化利器 R Graphics Cookbook中译教程
我的学习小圈子👉 加入
赞、在看就是最大的支持