使用 PyScript 将 Jupyter Notebook 转换为交互式仪表盘
PyScript是一个开源框架,可让用户在浏览器中使用 HTML 运行 Python 程序。它是利用Pyodide、WASM和其他现代网络技术的强大功能开发而成的。PyScript 提供了一个灵活的框架,Python 开发人员,尤其是数据科学家,可以在此基础上直接在 Python 中创建可扩展的组件。
以下是使用 PyScript 的一些好处:
Python 程序员可以使用 PyScript 轻松编写独立的网络应用程序。 统计学家或数据科学家可以生成可共享、可点击的 HTML 文件,用于查看模型预测。 用户无需安装 Python 或任何其他依赖项即可与仪表板交互。当脚本在浏览器中运行时,它们会自动安装。
在本文中,我们将学习如何使用matplotlib将 Jupyter Notebook 中的一些数据可视化,然后使用自动化脚本将其自动转换为使用PyScript 的独立交互式仪表盘。
上下滑动查看更多
# dashboard.html<html>
<head>
<title>YouTube and Spotify artists</title>
<meta charset="utf-8">
<link rel="stylesheet" href="https://pyscript.net/latest/pyscript.css" />
<script defer src="https://pyscript.net/latest/pyscript.js"></script>
</head>
<body>
<py-config>
packages = [ "pandas", "matplotlib" ]
</py-config>
<py-script>import pandas as pd
import urllib.request
import matplotlib.pyplot as pltfrom pyodide.http import open_url
from pyodide.ffi import create_proxyurl = "https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv"
df = pd.read_csv(open_url(url))current_selected = []
filter_elements = js.document.getElementsByName( "Album_type" )def plot(df):
plt.rcParams["figure.figsize"] = (15, 10)
df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
fig, (ax1, ax2) = plt.subplots(1, 2)
ax1.set_title('Top 10 Artists on YouTube')
df_views.plot(kind='bar', ax=ax1)
ax2.set_title('Top 10 Artists on Spotify')
df_streams.plot(kind='bar', ax=ax2)
ax1.set_xlabel('Artist')
ax1.set_ylabel('Views')
ax2.set_xlabel('Artist')
ax2.set_ylabel('Streams')
display(fig, target="graph-area", append=False)def select_filter(event):
for ele in filter_elements:
if ele.checked:
current_selected = ele.value
break
if current_selected == "ALL":
plot(df)
else:
filter = df.apply(lambda x: ele.value in x[ "Album_type" ], axis=1)
plot(df[filter])ele_proxy = create_proxy(select_filter)
for ele in filter_elements:
if ele.value == "ALL":
ele.checked = True
current_selected = ele.value
ele.addEventListener("change", ele_proxy)plot(df)
</py-script>
<div id="input" style="margin: 20px;">
Select Album_type : <br/>
<input type="radio" id="all" name="Album_type" value="ALL">
<label for="all"> All</label>
<input type="radio" id="album" name="Album_type" value="album">
<label for="album"> album </label>
<input type="radio" id="single" name="Album_type" value="single">
<label for="single"> single </label>
<input type="radio" id="compilation" name="Album_type" value="compilation">
<label for="compilation"> compilation </label></div>
<py-repl>
df
</py-repl><div id="graph-area"></div>
</body>
</html>
下面是互动仪表板的预览。
笔记本中的数据可视化
在本教程中,我们将使用YouTube 和 Spotify 艺术家数据集,并按照以下步骤操作,使笔记本准备好被自动化脚本使用:
单击 "视图"->"单元格工具栏"->"标签",启用 Jupyter Notebook 单元格标签编辑器。这将启用标签用户界面。这很有用,因为用户需要添加自动化脚本所需的几个重要单元格标签。 从托管的 URL 读取数据。请注意,PyScript 使用 Pyodide来访问服务器上的任何数据。因此,我们需要在包含数据集 URL 的单元格中添加一个特殊的标签data-url。根据需要分析数据集。 绘制相关详细信息。在这里,我们将分别绘制 YouTube 和 Spotify 上排名前 10 的艺术家。确保核心绘图逻辑位于一个单元格中,并在该单元格中添加 data-plot标签。这表示需要作为 PyScript 的一部分运行的脚本。同时,在该单元格中添加filter-Album_type标签。这将在Album_type列上启用交互功能,该列有 3 个唯一值:专辑、单曲和合辑。您可以用任何其他合适的列来代替它。
import pandas as pd
import urllib.request
import matplotlib.pyplot as plt
url = "https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv"
urllib.request.urlretrieve(url, filename="Spotify_Youtube.csv")
df = pd.read_csv("Spotify_Youtube.csv")
df.head()
df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
fig, (ax1, ax2) = plt.subplots(1, 2)
ax1.set_title('Top 10 Artists on YouTube')
df_views.plot(kind='bar', ax=ax1)
ax2.set_title('Top 10 Artists on Spotify')
df_streams.plot(kind='bar', ax=ax2)
ax1.set_xlabel('Artist')
ax1.set_ylabel('Views')
ax2.set_xlabel('Artist')
ax2.set_ylabel('Streams')
自动化脚本
上下滑动查看更多
# nb2dashboard.pyimport sys
import nbformat
import pandas as pd
import urllib.request
from string import Templatedef convert(filename):
print(f'Converting {filename} to PyScript dashboard')
ref = nbformat.read(filename, as_version=nbformat.NO_CONVERT)
cells = ref['cells']imports = []
for cell in cells:
if cell["cell_type"] == "code":
tags = cell["metadata"].get("tags")
source = cell["source"]# package imports
if "import" in source:
import_cells = source.replace("\n", "\n" + " ")
for package_import in source.split("\n"):
pkg = package_import.split(" ")[1]
if "." not in pkg and pkg != "urllib":
imports.append(f'"{pkg}"')
else:
main_pkg = pkg.split(".")[0]
if main_pkg != "urllib":
imports.append(f'"{main_pkg}"')if tags:
for tag in tags:
# read data url
if "data-url" in tags:
data_url = source[source.find("=") + 1 :].strip().replace('"', "")# data plot code
if "data-plot" in tags:
plot_code = source.replace("\n", "\n" + " ")if "filter" in tag:
filter_col_name = tag.split("-")[1]
filter_col = f'"{filter_col_name}"'
urllib.request.urlretrieve(data_url, filename="dataset.csv")
df = pd.read_csv("dataset.csv")
unique_filter_col_values = df[filter_col_name].unique().tolist()selection = f"Select {filter_col_name} : <br/>\n"
selection += f"<input type=\"radio\" id=\"all\" name=\"{filter_col_name}\" value=\"ALL\">\n"
selection += "<label for=\"all\"> All</label>\n"for val in unique_filter_col_values:
selection += f"<input type=\"radio\" id=\"{val}\" name=\"{filter_col_name}\" value=\"{val}\">\n"
selection += f"<label for=\"{val}\"> {val} </label>\n"template = Template(
"""<html>
<head>
<title>YouTube and Spotify artists</title>
<meta charset="utf-8">
<link rel="stylesheet" href="https://pyscript.net/latest/pyscript.css" />
<script defer src="https://pyscript.net/latest/pyscript.js"></script>
</head>
<body>
<py-config>
packages = [ $packages ]
</py-config>
<py-script>$user_imports
from pyodide.http import open_url
from pyodide.ffi import create_proxyurl = $url
df = pd.read_csv(open_url(url))current_selected = []
filter_elements = js.document.getElementsByName( $filter_column )def plot(df):
plt.rcParams["figure.figsize"] = (15, 10)
$code
display(fig, target="graph-area", append=False)def select_filter(event):
for ele in filter_elements:
if ele.checked:
current_selected = ele.value
break
if current_selected == "ALL":
plot(df)
else:
filter = df.apply(lambda x: ele.value in x[ $filter_column ], axis=1)
plot(df[filter])ele_proxy = create_proxy(select_filter)
for ele in filter_elements:
if ele.value == "ALL":
ele.checked = True
current_selected = ele.value
ele.addEventListener("change", ele_proxy)plot(df)
</py-script>
<div id="input" style="margin: 20px;">
$selection</div>
<py-repl>
df
</py-repl><div id="graph-area"></div>
</body>
</html>
"""
)dashboard = template.substitute(packages=(", ".join(imports)),
url=f"\"{data_url}\"",
user_imports=import_cells,
code=plot_code,
filter_column=filter_col,
selection=selection)with open("dashboard.html", "w") as file:
file.write(dashboard)if __name__ == "__main__":
notebook = sys.argv[1]
convert(notebook)
以上内容是并将相关内容转换为 PyScript 文件。关于此脚本的一些要点
脚本使用 nbformat软件包读取笔记本中每个单元格的源代码和元数据标签。重要的变量会被提取出来,如数据集的 url、可能的软件包导入以及绘制图表的最终代码。根据用户添加的 过滤标签,脚本会从数据集中提取唯一的列值。最后,将这些变量替换到模板字符串中,生成 PyScript 代码。
现在,在与笔记本相同的目录下运行nb2dashboard.py文件。
python nb2dashboard.py <notebook_name>.ipynb
这将生成一个名为dashboard.html 的文件。确保<py-config>中提到的软件包准确无误。<py-script>中 Python 代码的缩进可能因编辑器而异,尤其是软件包导入和def plot(data) 中附加的代码。这些缩进错误可以手动修正。
转换后的独立仪表板
下面是一个PyScript 参考文件,你可以用它来验证生成的文件。导入的软件包在<py-config>标记中定义。
<py-config>
packages = [ "pandas", "matplotlib" ]
</py-config>
所有 Python 代码都需要包含在<py-script>标记中。数据集使用Pyodide库访问。
path = "Spotify_Youtube.csv"
df = pd.read_csv(path)
select_filter函数用于根据用户交互选择的列值选择数据集的子集。通过一组单选按钮实现了交互性,每个单选按钮都与用户在笔记本中标记的过滤列的唯一值相对应。
<div id="input" style="margin: 20px;">
Select Album_type : <br/>
<input type="radio" id="all" name="Album_type" value="ALL">
<label for="all"> All</label>
<input type="radio" id="album" name="Album_type" value="album">
<label for="album"> album </label>
<input type="radio" id="single" name="Album_type" value="single">
<label for="single"> single </label>
<input type="radio" id="compilation" name="Album_type" value="compilation">
<label for="compilation"> compilation </label>
</div>
🏴☠️宝藏级🏴☠️ 原创公众号『数据STUDIO』内容超级硬核。公众号以Python为核心语言,垂直于数据科学领域,包括可戳👉Python|MySQL|数据分析|数据可视化|机器学习与数据挖掘|爬虫 等,从入门到进阶!
长按👇关注- 数据STUDIO -设为星标,干货速递