pythonic生物人

别只用R ggplot2来画图……

本篇详细点介绍ggstatsplot使用,续上篇👉极大补充ggplot2的统计分析能力。

ggstatsplot = gg + stats + plot,可见其在ggplot2的基础上补充了统计能力stats,将数据分析工作流中的数据可视化和统计建模两个不同的阶段结合在一起,使数据挖掘变得简单和快速。

ggstatsplot目前支持的图形

ggstatsplot目前支持的图形
ggstatsplot目前支持的图形

ggstatsplot目前支持的统计检验

ggstatsplot目前支持的统计检验
ggstatsplot目前支持的统计检验
Bayesian analysis相关检验
Bayesian analysis相关检验

这些统计检验调用了statsExpressions package的功能,

statsExpressions work flow
statsExpressions work flow

统计结果展示

ggstatsplot这里的统计结果展示参考了统计结果gold standard标准格式,

Image

具体到图中,如下图红圈部分,

Image

个人感觉太过于fancy,实际使用时可以酌情处理。


下面简单介绍ggstatsplot主要函数使用

ggbetweenstats

创建数据(between-group或者between-condition comparisons)的violin plot, box plot或者二者混合图,ggbetweenstats具有众多参数可自行个性化。

一个简单例子,使用默认参数,使用iris数据集,比较不同鸢尾花萼片长度差异。

library(ggstatsplot)
ggbetweenstats(
  data = iris,
  x = Species,
  y = Sepal.Length,
  title = "Distribution of sepal length across Iris species"
)
Image
  • 修改配色、主题
library(ggstatsplot)
library(ggplot2)
library(ggthemes)

ggstatsplot::ggbetweenstats(
  data = iris,
  x = Species,
  y = Sepal.Length,
  title = "Distribution of sepal length across Iris species",
  ggtheme = ggthemes::theme_economist(),#ggthemes经济学人主题
  package = "wesanderson", #修改图形配色包View(paletteer::palettes_d_names)可查看所有可使用的包及对应色盘
  palette = "Darjeeling1"# 选择颜色盘
)
Image

package和palette参数可供选择的特别多,

Image

grouped_ggbetweenstats

grouped_ggbetweenstats,可以很方便的展示数据集子集的分布差异。

简单使用下,

grouped_ggbetweenstats(
  data = dplyr::filter(
    .data = ggstatsplot::movies_long,
    genre %in% c("Action", "Action Comedy")
  ),#数据过滤
  x = mpaa,
  y = length,
  grouping.var = genre, #分组设置
  ggsignif.args = list(textsize = 4, tip_length = 0.01),#p值属性设置
#p.adjust.method = "bonferroni", 
  ggplot.component = list(ggplot2::scale_y_continuous(sec.axis = ggplot2::dup_axis())),
  k = 3,
  title.prefix = "电影类别",
  caption = substitute(paste(italic("Source"), ": IMDb (Internet Movie Database)")), 
  ggtheme = ggthemes::theme_economist(),#使用ggthemes经济学人主题
  package = "wesanderson", #修改图形配色包View(paletteer::palettes_d_names)可查看所有可使用的
  palette = "Darjeeling1", # 选择颜色盘
  plotgrid.args = list(nrow = 2),
  title.text = "不同类别电影中不同级别电影时长差异"
)
Image

ggwithinstats

类似于ggbetweenstats,不过他可以把各个箱子牵起来,下图会把均值牵起来。

library(WRS2)

# plot
ggstatsplot::ggwithinstats(
  data = WineTasting,
  x = Wine,
  y = Taste,
  title = "Wine tasting",
  caption = "Data source: `WRS2` R package",
  ggtheme = ggthemes::theme_fivethirtyeight(),
  ggstatsplot.layer = FALSE,
  messages = FALSE,
  ggsignif.args = list(textsize = 3, tip_length = 0.01),
)
Image

grouped_ggwithinstats

grouped_ggwithinstats(
  data = dplyr::filter(
    .data = ggstatsplot::bugs_long,
    region %in% c("Europe", "North America"),
    condition %in% c("LDLF", "LDHF")
  ),
  x = condition,
  y = desire,
  xlab = "Condition",
  ylab = "Desire to kill an artrhopod",
  grouping.var = region,
  outlier.tagging = TRUE,
  outlier.label = education,
  ggtheme = hrbrthemes::theme_ipsum_tw(),
  ggstatsplot.layer = FALSE,
  messages = FALSE
)
Image

ggscatterstats

这是R中的方法绘制边际分布图,python版本的边际分布图见👉Python边际图代码模版

ggstatsplot::ggscatterstats(
  data = ggplot2::msleep,
  x = sleep_rem,
  y = awake,
  xlab = "REM sleep (in hours)",
  ylab = "Amount of time spent awake (in hours)",
  title = "Understanding mammalian sleep",
  messages = FALSE
)
Image

边际上的图有以下5种图可修改,修改参数marginal.type即可。

Image
ggscatterstats(
  data = dplyr::filter(.data = ggstatsplot::movies_long, genre == "Action"),
  x = budget,
  y = rating,
  type = "robust", # type of test that needs to be run
  xlab = "Movie budget (in million/ US$)", # label for x axis
  ylab = "IMDB rating", # label for y axis
  label.var = "title", # variable for labeling data points
  label.expression = "rating < 5 & budget > 100", # expression that decides which points to label
  title = "Movie budget and IMDB rating (action)", # title text for the plot
  caption = expression(paste(italic("Note"), ": IMDB stands for Internet Movie DataBase")),
  ggtheme = hrbrthemes::theme_ipsum_ps(), # choosing a different theme
  ggstatsplot.layer = FALSE, # turn off `ggstatsplot` theme layer
  marginal.type = "densigram", # type of marginal distribution to be displayed
  xfill = "pink", # color fill for x-axis marginal distribution
  yfill = "#009E73", # color fill for y-axis marginal distribution
  centrality.parameter = "median", # central tendency lines to be displayed
  messages = FALSE# turn off messages and notes
)
Image

grouped_ggscatterstats

同样也有grouped函数

grouped_ggscatterstats(
  data = dplyr::filter(
    .data = ggstatsplot::movies_long,
    genre %in% c("Action", "Action Comedy", "Action Drama", "Comedy")
  ),
  x = rating,
  y = length,
  grouping.var = genre, # grouping variable
  label.var = title,
  label.expression = length > 200,
  xfill = "#E69F00",
  yfill = "#8b3058",
  xlab = "IMDB rating",
  title.prefix = "Movie genre",
  ggtheme = ggplot2::theme_grey(),
  ggplot.component = list(
    ggplot2::scale_x_continuous(breaks = seq(2, 9, 1), limits = (c(2, 9)))
  ),
  plotgrid.args = list(nrow = 2),
  title.text = "Relationship between movie length by IMDB ratings for different genres"
)
Image

ggpiestats

绘制饼图,计算各个快之间是否有差异。

Titanic_full_50 <- dplyr::sample_frac(tbl = ggstatsplot::Titanic_full, size = 0.5)
ggpiestats(
  data = Titanic_full_50,
  x = Survived,
  title = "Passenger survival on the Titanic", # title for the entire plot
  caption = "Source: Titanic survival dataset", # caption for the entire plot
  legend.title = "Survived?",
  package = "ggthemr",
  palette = "dust",
)
Image

数据集分组,组间及组内计算统计指标。

Titanic_full_50 <- dplyr::sample_frac(tbl = ggstatsplot::Titanic_full, size = 0.5)

ggpiestats(
  data = Titanic_full_50,
  x = Survived,
  y = Sex,
  title = "Passenger survival on the Titanic by gender", # title for the entire plot
  caption = "Source: Titanic survival dataset", # caption for the entire plot
  legend.title = "Survived?", # legend title
  ggtheme = ggplot2::theme_grey(), # changing plot theme
  package = "ggthemr",
  palette = "dust",
  k = 3, # decimal places in result
  perc.k = 1# decimal places in percentage labels
) + # further modification with `ggplot2` commands
  ggplot2::theme(
    plot.title = ggplot2::element_text(
      color = "black",
      size = 14,
      hjust = 0
    )
  )
Image

grouped_ggpiestats

grouped_ggpiestats(
  data = ggstatsplot::movies_long,
  x = genre,
  grouping.var = mpaa, # grouping variable
  title.prefix = "Movie genre", # prefix for the faceted title
  label.repel = TRUE, # repel labels (helpful for overlapping labels)
  package = "ggthemr",
  palette = "dust",
  title.text = "Composition of MPAA ratings for different genres"
)
Image

ggbarstats

功能类似于ggpiestats,图形非常好康。

ggbarstats(
  data = ggstatsplot::movies_long,
  x = mpaa,
  y = genre,
  sampling.plan = "jointMulti",
  title = "MPAA Ratings by Genre",
  xlab = "movie genre",
  legend.title = "MPAA rating",
  ggtheme = hrbrthemes::theme_ipsum_pub(),
  ggplot.component = list(scale_x_discrete(guide = guide_axis(n.dodge = 2))),
  package = "ggthemr",
  palette = "dust",
  messages = FALSE
)
Image

grouped_ggbarstats

df <-
  dplyr::filter(
    .data = forcats::gss_cat,
    race %in% c("Black", "White"),
    relig %in% c("Protestant", "Catholic", "None"),
    !partyid %in% c("No answer", "Don't know", "Other party")
  )

# plot
ggstatsplot::grouped_ggbarstats(
  data = df,
  x = relig,
  y = partyid,
  grouping.var = race,
  title.prefix = "Race",
  xlab = "Party affiliation",
  package = "ggthemr",
  palette = "dust",
  ggtheme = ggthemes::theme_tufte(base_size = 12),
  ggstatsplot.layer = FALSE,
  title.text = "Race, religion, and political affiliation",
  plotgrid.args = list(nrow = 2)
)
  package = "ggthemr",
  palette = "dust",
  k = 3, # decimal places in result
  perc.k = 1# decimal places in percentage labels
) + # further modification with `ggplot2` commands
  ggplot2::theme(
    plot.title = ggplot2::element_text(
      color = "black",
      size = 14,
      hjust = 0
    )
  )
Image

gghistostats

可视化单一变量的分布,计算单一变量的均值与指定值(下例子中为5)之间是否存在统计学差异。

gghistostats(
  data = iris, # dataframe from which variable is to be taken
  x = Sepal.Length, # numeric variable whose distribution is of interest
  title = "Distribution of Iris sepal length", # title for the plot
  caption = substitute(paste(italic("Source:"), "Ronald Fisher's Iris data set")),
  bar.measure = "both",
  test.value = 5, # default value is 0
  test.value.line = TRUE, # display a vertical line at test value
  centrality.parameter = "mean", # which measure of central tendency is to be plotted
  centrality.line.args = list(color = "darkred"), # aesthetics for central tendency line
  binwidth = 0.10, # binwidth value (experiment)
  ggtheme = hrbrthemes::theme_ipsum_tw(), # choosing a different theme
  ggstatsplot.layer = FALSE,# turn off ggstatsplot theme layer
  package = "ggthemr",
  palette = "dust",
)
Image

grouped_gghistostats

grouped_gghistostats(
  data = dplyr::filter(
    .data = ggstatsplot::movies_long,
    genre %in% c("Action", "Action Comedy", "Action Drama", "Comedy")
  ),
  x = budget,
  xlab = "Movies budget (in million US$)",
  type = "robust", # use robust location measure
  grouping.var = genre, # grouping variable
  normal.curve = TRUE, # superimpose a normal distribution curve
  normal.curve.args = list(color = "red", size = 1),
  title.prefix = "Movie genre",
  ggtheme = ggthemes::theme_economist(),
  ggplot.component = list( # modify the defaults from `ggstatsplot` for each plot
    ggplot2::scale_x_continuous(breaks = seq(0, 200, 50), limits = (c(0, 200)))
  ),
  plotgrid.args = list(nrow = 2),
  title.text = "Movies budgets for different genres"
)
Image

ggcorrmat

轻松绘制相关系数矩阵图,python中也可以轻松绘制该图:

热力图heatmap代码模版~

seaborn又一个扩展heatmapz

gapminder_2007 <- dplyr::filter(.data = gapminder::gapminder, year == 2007)

# producing the correlation matrix
ggstatsplot::ggcorrmat(
  data = gapminder_2007, # data from which variable is to be taken
  cor.vars = lifeExp:gdpPercap,# specifying correlation matrix variables
  colors = c("#E69F00", "white","#d5695d"), #传入配色
)
Image
ggstatsplot::ggcorrmat(
  data = gapminder_2007, # data from which variable is to be taken
  cor.vars = lifeExp:gdpPercap, # specifying correlation matrix variables
  cor.vars.names = c(
"Life Expectancy",
"population",
"GDP (per capita)"
  ),
  type = "spearman", # which correlation coefficient is to be computed
  lab.col = "red", # label color
  ggtheme = ggplot2::theme_light(), # selected ggplot2 theme
  ggstatsplot.layer = FALSE, # turn off default ggestatsplot theme overlay
  matrix.type = "lower", # correlation matrix structure
  colors = NULL, # 关闭指定色号
  package = "ggthemr",#启用色盘
  palette = "grape",
  title = "Gapminder correlation matrix", # custom title
  subtitle = "Source: Gapminder Foundation"# custom subtitle
)
Image

grouped_ggcorrmat

grouped_ggcorrmat(
# arguments relevant for ggstatsplot::ggcorrmat
  data = dplyr::sample_frac(tbl = ggplot2::diamonds, size = 0.05),
  type = "robust", # percentage bend correlation coefficient
  beta = 0.2, # bending constant
  p.adjust.method = "holm", # method to adjust p-values for multiple comparisons
  grouping.var = cut,
  title.prefix = "Quality of cut",
  cor.vars = c(carat, depth:z),
  cor.vars.names = c(
"carat",
"total depth",
"table",
"price",
"length (in mm)",
"width (in mm)",
"depth (in mm)"
  ),
  lab.size = 3.5,
# arguments relevant for ggstatsplot::combine_plots
  title.text = "Relationship between diamond attributes and price across cut",
  title.args = list(size = 16, color = "red"),
  caption.text = "Dataset: Diamonds from ggplot2 package",
  caption.args = list(size = 14, color = "blue"),
  plotgrid.args = list(
    labels = c("(a)", "(b)", "(c)", "(d)", "(e)"),
    nrow = 3,
    ncol = 2
  )
)
Image

ggcoefstats

类似下面这种图

Image


-推荐阅读-

👉matplotlib教程:20w字+数百张图形+1W行代码+详细代码注释+学习交流群

图片

👉seaborn教程:12.3万字+500多张图形+8000行代码

图片

👉R可视化教程:39个章节+20w字+数百张图

图片
图片

如何加入学习?

👇(请备注:299)

图片