数据STUDIO

5 种被严重低估的统计检验

Image

Image

如果你的工具箱里只有 t检验 和 方差分析(ANOVA),那么你可能会错过数据中隐藏的关键模式——尤其是在面对非正态数据、异常值、时间趋势或重复测量时。

本文中,云朵君将和大家一起与学习 5种强大但被严重忽略的统计检验,并在免疫学(TCR/BCR分析)、金融(股票价格)和体育科学(运动员表现)中列举示例以展示其能够提供更可靠的洞察。

1. Mann-Kendall趋势检验

Mann-Kendall 检验是一种非参数检验,用于检查随时间推移是否存在单调的上升或下降趋势。与常规线性回归不同,它不假定线性关系或正态分布,直接建模。

应用

  • 金融: 检测股票价格的微妙上升趋势。
  • 免疫学: 发现 TCR/BCR 克隆随时间的扩展或收缩。

上升趋势与无趋势

下面我们生成两个合成数据集来说明这两种情况。

  1. 上升趋势: 随时间推移而增加的数据。
  2. 无显著趋势: 围绕平均值随机波动的数据(白噪声)。

我们将对这两个数据集进行 Mann-Kendall 检验并进行比较。

import pymannkendall as mk
np.random.seed(43)

# 1. Upward Trend Data
time_trend = np.arange(1, 31)
values_trend = 0.5 * time_trend + np.random.normal(0, 2, size=len(time_trend))

# 2. No Trend Data (just noise)
time_no_trend = np.arange(1, 31)
values_no_trend = np.random.normal(0, 2, size=len(time_no_trend))

# Perform the Mann-Kendall Test
trend_result = mk.original_test(values_trend)
no_trend_result = mk.original_test(values_no_trend)

fig, axes = plt.subplots(1, 2, figsize=(20, 10))
plt.grid(lw=2, ls=':')

axes[0].plot(time_trend, values_trend, marker='o', linestyle='-', color='blue')
axes[0].set_title("Data with Upward Trend")
axes[0].set_xlabel("Time")
axes[0].set_ylabel("Value")

textstr_trend = '\n'.join((
    f"Trend: {trend_result.trend}",
    f"P-value: {trend_result.p:.5f}",
    f"Tau (Kendall's): {trend_result.Tau:.4f}"
))
props = dict(boxstyle='round', facecolor='wheat', alpha=0.5)
axes[0].text(0.05, 0.95, textstr_trend, transform=axes[0].transAxes, fontsize=14,
             verticalalignment='top', bbox=props)

axes[1].plot(time_no_trend, values_no_trend, marker='o', linestyle='-', color='red')
axes[1].set_title("Data with No Trend")
axes[1].set_xlabel("Time")
axes[1].set_ylabel("Value")
axes[1].grid(lw=2, ls=':')
axes[0].grid(lw=2, ls=':')

textstr_no_trend = '\n'.join((
    f"Trend: {no_trend_result.trend}",
    f"P-value: {no_trend_result.p:.5f}",
    f"Tau (Kendall's): {no_trend_result.Tau:.4f}"
))
axes[1].text(0.05, 0.95, textstr_no_trend, transform=axes[1].transAxes, fontsize=14,
             verticalalignment='top', bbox=props)

plt.tight_layout()
plt.grid(lw=2, ls=':')
plt.show()
Image
  • 在上升趋势的情况下,你可能会看到 “Upward Trend”,且 p 值较低,具有统计学意义。
  • 在无显著趋势情况下,测试结果应显示 “No Trend”,且 p 值较高。

2. Mood中位数检验

Mood中位数检验可以检测两个或多个样本是否来自中位数相同的人群。当数据分布偏斜或你特别关注中位数而非平均值时,这是一种稳健的非参数检验。

应用

  • 免疫学: 比较多个患者组之间 TCR 多样性的中位数(例如,两个 Rerpertoires 之间的 K1000 指数)。
  • 金融: 比较不同股票的日收益中位数。
  • 体育: 比较不同策略的性能指标(如运行时间)中位数。

不同的中位数与相似的中位数

我们将创建两组:

  1. 不同的中位数: 一组的中位数明显较高,另一组较低。
  2. 相似中位数: 基本分布相似的组。
import numpy as np
from scipy.stats import median_test
import matplotlib.pyplot as plt

np.random.seed(2024)

# Scenario A: Different medians
group_A_diff = np.random.exponential(scale=1.0, size=50)
group_B_diff = np.random.exponential(scale=2.0, size=50)

# Scenario B: Similar medians
group_A_sim = np.random.exponential(scale=1.0, size=50)
group_B_sim = np.random.exponential(scale=1.1, size=50)

# Apply Mood's Median Test to both scenarios
stat_diff, p_value_diff, median_diff, table_diff = median_test(group_A_diff, group_B_diff)
stat_sim, p_value_sim, median_sim, table_sim = median_test(group_A_sim, group_B_sim)

fig, axes = plt.subplots(1, 2, figsize=(20, 7))

axes[0].boxplot([group_A_diff, group_B_diff], labels=["Group A (diff)", "Group B (diff)"])
axes[0].set_title("Groups with Different Medians")
axes[0].set_ylabel("Value")

textstr_diff = '\n'.join((
    f"Test Statistic: {stat_diff:.4f}",
    f"P-value: {p_value_diff:.5f}",
    f"Overall Median: {median_diff:.4f}",
    f"Contingency Table:\n{table_diff}"
))
props = dict(boxstyle='round', facecolor='wheat', alpha=0.5)
axes[0].text(0.05, 0.95, textstr_diff, transform=axes[0].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[1].boxplot([group_A_sim, group_B_sim], labels=["Group A (sim)", "Group B (sim)"])
axes[1].set_title("Groups with Similar Medians")
axes[1].set_ylabel("Value")

textstr_sim = '\n'.join((
    f"Test Statistic: {stat_sim:.4f}",
    f"P-value: {p_value_sim:.5f}",
    f"Overall Median: {median_sim:.4f}",
    f"Contingency Table:\n{table_sim}"
))
axes[1].text(0.05, 0.95, textstr_sim, transform=axes[1].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[1].grid(lw=2, ls=':', axis='y')
axes[0].grid(lw=2, ls=':', axis='y')

plt.tight_layout()
plt.show()
Image

3. Friedman 检验

Friedman 检验相当于重复测量方差分析的非参数检验。它可以检验多种处理或条件是否会导致对相同受试者(或区块)进行测量时产生不同的结果。

应用

  • 体育: 比较同一组运动员在不同训练方案下的长期表现。
  • 免疫学: 测量同一患者在不同治疗或多个时间点下的反应。
  • 金融: 在多个时间窗口中重复相同的时间序列数据,评估不同的预测方法
  • 展示两种情况

  1. 显著差异: 三种训练方法得出的分数明显不同。
  2. 无显著差异: 三种训练方法得出的分数大致相同。
import numpy as np
from scipy.stats import friedmanchisquare
import matplotlib.pyplot as plt

np.random.seed(456)

# Scenario 1: Significant Difference
method_A1 = np.random.normal(loc=50, scale=5, size=10)
method_B1 = np.random.normal(loc=60, scale=5, size=10)
method_C1 = np.random.normal(loc=70, scale=5, size=10)
stat1, pval1 = friedmanchisquare(method_A1, method_B1, method_C1)

# Scenario 2: No Significant Difference
method_A2 = np.random.normal(loc=50, scale=5, size=10)
method_B2 = np.random.normal(loc=51, scale=5, size=10)
method_C2 = np.random.normal(loc=49.5, scale=5, size=10)
stat2, pval2 = friedmanchisquare(method_A2, method_B2, method_C2)

fig, axes = plt.subplots(1, 2, figsize=(20, 7))

axes[0].boxplot([method_A1, method_B1, method_C1], labels=["A", "B", "C"])
axes[0].set_title("Scenario 1: Significant Difference")
textstr1 = f"Test statistic: {stat1:.4f}\nP-value: {pval1:.4f}"
props = dict(boxstyle='square', facecolor='wheat', alpha=0.5)
axes[0].text(0.05, 0.95, textstr1, transform=axes[0].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[1].boxplot([method_A2, method_B2, method_C2], labels=["A", "B", "C"])
axes[1].set_title("Scenario 2: No Significant Difference")
textstr2 = f"Test statistic: {stat2:.4f}\nP-value: {pval2:.4f}"
axes[1].text(0.05, 0.95, textstr2, transform=axes[1].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[0].grid(lw=2, ls=':', axis='y')
axes[1].grid(lw=2, ls=':', axis='y')

plt.tight_layout()
plt.show()
Image
  • 在显著差异中,你应该看到中位数的明显差异和较低的 p 值。
  • 在无显著差异中,分布严重重叠,p 值会更高。

4. Theil-Sen 估计器

Theil-Sen 估计器是一种估计线性关系斜率的稳健方法。与普通最小二乘法(OLS)不同,它对异常值的敏感度要低得多,因为它是基于所有成对斜率的中位数。

应用

  • 金融: 估计股票走势,而不会被几个极端日误导。
  • 免疫学: 跟踪 TCR 频率或抗体水平随时间的变化率,即使异常日出现意外峰值。
  • 体育: 检查运动员的表现趋势如何演变,忽略特别糟糕或特别好的 “侥幸 ”日子。

展示两种情况

  1. 无异常值: 数据遵循近乎完美的线性趋势。
  2. 有异常值: 某些点严重偏离趋势。
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression, TheilSenRegressor

np.random.seed(789)
X = np.linspace(0, 10, 50)
y = 3 * X + np.random.normal(0, 1, 50)

# Scenario 1: No Outliers
X1 = X.reshape(-1, 1)
y1 = y.copy()
lr1 = LinearRegression().fit(X1, y1)
ts1 = TheilSenRegressor().fit(X1, y1)

y_pred_lr1 = lr1.predict(X1)
y_pred_ts1 = ts1.predict(X1)

# Scenario 2: With Outliers
X2 = X.reshape(-1, 1)
y2 = y.copy()
y2[10] += 20   # Add big positive outlier
y2[25] -= 15   # Add big negative outlier
lr2 = LinearRegression().fit(X2, y2)
ts2 = TheilSenRegressor().fit(X2, y2)

y_pred_lr2 = lr2.predict(X2)
y_pred_ts2 = ts2.predict(X2)

fig, axes = plt.subplots(1, 2, figsize=(20, 10))

axes[0].scatter(X1, y1, color='grey', label='Data')
axes[0].plot(X1, y_pred_lr1, color='red', label='Linear Regression')
axes[0].plot(X1, y_pred_ts1, color='green', label='Theil-Sen')
axes[0].set_title("No Outliers")
axes[0].set_xlabel("X")
axes[0].set_ylabel("Y")
axes[0].legend()

axes[1].scatter(X2, y2, color='grey', label='Data (with outliers)')
axes[1].plot(X2, y_pred_lr2, color='red', label='Linear Regression')
axes[1].plot(X2, y_pred_ts2, color='green', label='Theil-Sen')
axes[1].set_title("With Outliers")
axes[1].set_xlabel("X")
axes[1].set_ylabel("Y")
axes[1].legend()

axes[0].grid(lw=2, ls=':', axis='both')
axes[1].grid(lw=2, ls=':', axis='both')

plt.tight_layout()
plt.show()
Image

5. Anderson-Darling检验

Anderson-Darling检验是一种拟合优度检验,用于检验数据是否来自特定的分布(通常是正态分布)。与其他检验(如 Shapiro-Wilk 检验)相比,Anderson-Darling检验更重视尾部,因此对极端值的偏差更为敏感。

应用

  • 金融: 验证每日收益是否真正遵循正态分布--剧透警告:通常情况下,它们并不遵循正态分布!
  • 免疫学: 检查你的 TCR/BCR 多样性指标是否遵循假定分布。
  • 体育: 在应用参数检验之前,确保性能数据符合正态性假设。

展示两种情况

  1. 来自正态分布的数据。
  2. 来自非正态分布的数据(例如指数)。
import numpy as np
from scipy.stats import anderson
import matplotlib.pyplot as plt

np.random.seed(101)

# Scenario 1: Normal data
normal_data = np.random.normal(loc=0, scale=1, size=200)
result_normal = anderson(normal_data, dist='norm')

# Scenario 2: Non-normal data (exponential)
non_normal_data = np.random.exponential(scale=1.0, size=200)
result_non_normal = anderson(non_normal_data, dist='norm')

# Quick visualization
fig, axes = plt.subplots(1, 2, figsize=(20, 8))

axes[0].hist(normal_data, bins=20, color='blue', alpha=0.7)
axes[0].set_title("Scenario 1: Normal Data")

textstr_normal = '\n'.join((
    f"Statistic: {result_normal.statistic:.4f}",
"Critical Values / Significance Levels:",
    *[f" - {sl}% level: CV = {cv:.4f}"for cv, sl in zip(result_normal.critical_values, result_normal.significance_level)]
))
props = dict(boxstyle='round', facecolor='wheat', alpha=0.5)
axes[0].text(0.05, 0.95, textstr_normal, transform=axes[0].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[1].hist(non_normal_data, bins=20, color='orange', alpha=0.7)
axes[1].set_title("Scenario 2: Non-Normal Data")

textstr_non_normal = '\n'.join((
    f"Statistic: {result_non_normal.statistic:.4f}",
"Critical Values / Significance Levels:",
    *[f" - {sl}% level: CV = {cv:.4f}"for cv, sl in zip(result_non_normal.critical_values, result_non_normal.significance_level)]
))
axes[1].text(0.25, 0.95, textstr_non_normal, transform=axes[1].transAxes, fontsize=10,
             verticalalignment='top', bbox=props)

axes[0].grid(lw=2, ls=':', axis='y')
axes[1].grid(lw=2, ls=':', axis='y')

plt.tight_layout()
plt.show()
Image
  • 对于正态数据,你可能会发现测试统计量低于临界值,这表明无法拒绝正态性。
  • 对于非正态分布数据,测试统计量较高,表明数据偏离了正态分布--尤其是尾部行为更为明显。

写在最后

这五种统计检验确实在非理想数据条件下表现出色,能够解决许多传统参数检验(如t检验、ANOVA)的局限性。以下是它们的核心特点和应用场景的进一步解读,以及何时选择它们的建议:

1. Mann-Kendall 趋势检验

  • 用途:检测时间序列中的单调趋势(上升/下降),无需假设线性或正态性。
  • 优势:对缺失数据、非正态分布和季节性波动稳健。
  • 案例: 环境科学中分析气温或污染物的长期趋势。

2. Mood's Median 检验

  • 用途:比较多个独立组的中位数差异,替代单因素ANOVA的非参数版本。
  • 优势:不受异常值或偏态分布影响,适用于序数数据。
  • 注意:若组间方差差异大,Kruskal-Wallis可能更优。
  • 案例: 比较不同教育政策下学生考试成绩的中位数(数据有极端值)。

3. Friedman 检验

  • 用途:重复测量或匹配样本的非参数比较(类似重复测量ANOVA)。
  • 优势:处理同一受试者在不同条件下的排序数据(如药物治疗前后的效果)。
  • 后续分析:可用Nemenyi检验进行两两比较。
  • 案例: 运动员在三种不同训练方案下的表现排名(数据为序数或非正态)。

4. Theil-Sen 回归

  • 用途:估计线性趋势的斜率,对异常值不敏感。
  • 优势:比普通最小二乘(OLS)回归更稳健,计算中位数斜率。
  • 扩展:可与Siegel重复中位数方法结合处理多维数据。
  • 案例: 分析经济指标与贫困率的关系(数据中存在极端离群值)。

5. Anderson-Darling 检验

  • 用途:检验数据是否来自特定分布(如正态性、指数分布)。
  • 优势:比K-S检验对尾部差异更敏感,适合小样本。
  • 注意:需指定分布参数(可使用估计值)。
  • 案例:验证金融收益率的正态性(对极端风险敏感)。

何时使用这些检验?

问题类型推荐检验替代方案
时间序列趋势检测
Mann-Kendall
Seasonal-Kendall
多组中位数比较
Mood's Median
Kruskal-Wallis
重复测量数据
Friedman
重复测量ANOVA(若正态)
稳健回归(抗异常值)
Theil-Sen
Huber回归
分布检验(尤其尾部)
Anderson-Darling
Shapiro-Wilk / K-S

这些方法共同的特点是放弃部分统计效能换取稳健性,在现实数据中往往是更务实的选择。尤其在生物医学、社会科学、环境监测等领域,它们能揭示参数检验可能掩盖的真相。

🏴‍☠️宝藏级🏴‍☠️ 原创公众号『数据STUDIO』内容超级硬核。公众号以Python为核心语言,垂直于数据科学领域,包括可戳👉Python|MySQL|数据分析|数据可视化|机器学习与数据挖掘|爬虫等,从入门到进阶!

长按👇关注- 数据STUDIO -设为星标,干货速递ImageImage