数据STUDIO

使用 PyScript 将 Jupyter Notebook 转换为交互式仪表盘

Image

Image

PyScript是一个开源框架,可让用户在浏览器中使用 HTML 运行 Python 程序。它是利用Pyodide、WASM和其他现代网络技术的强大功能开发而成的。PyScript 提供了一个灵活的框架,Python 开发人员,尤其是数据科学家,可以在此基础上直接在 Python 中创建可扩展的组件。

以下是使用 PyScript 的一些好处:

  • Python 程序员可以使用 PyScript 轻松编写独立的网络应用程序。
  • 统计学家或数据科学家可以生成可共享、可点击的 HTML 文件,用于查看模型预测。
  • 用户无需安装 Python 或任何其他依赖项即可与仪表板交互。当脚本在浏览器中运行时,它们会自动安装。

在本文中,我们将学习如何使用matplotlib将 Jupyter Notebook 中的一些数据可视化,然后使用自动化脚本将其自动转换为使用PyScript 的独立交互式仪表盘。


上下滑动查看更多

# dashboard.html

<html> 
     <head> 
     <title>YouTube and Spotify artists</title>
     <meta charset="utf-8">
     <link rel="stylesheet" href="https://pyscript.net/latest/pyscript.css" />
     <script defer src="https://pyscript.net/latest/pyscript.js"></script>
     </head>
     <body>
     <py-config> 
         packages = [ "pandas", "matplotlib" ]
     </py-config>
     <py-script>

                  import pandas as pd
         import urllib.request
         import matplotlib.pyplot as plt

                  from pyodide.http import open_url
         from pyodide.ffi import create_proxy

                  url = "https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv"
         df = pd.read_csv(open_url(url))

                  current_selected = []
         filter_elements = js.document.getElementsByName( "Album_type" )

                  def plot(df):
             plt.rcParams["figure.figsize"] = (15, 10)
             df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
             df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
             fig, (ax1, ax2) = plt.subplots(1, 2)
             ax1.set_title('Top 10 Artists on YouTube')
             df_views.plot(kind='bar', ax=ax1)
             ax2.set_title('Top 10 Artists on Spotify')
             df_streams.plot(kind='bar', ax=ax2)
             ax1.set_xlabel('Artist')
             ax1.set_ylabel('Views')
             ax2.set_xlabel('Artist')
             ax2.set_ylabel('Streams')
             display(fig, target="graph-area", append=False)

                      def select_filter(event):
            for ele in filter_elements:
              if ele.checked:
                  current_selected = ele.value
                  break
            if current_selected == "ALL":
              plot(df)
            else:
              filter = df.apply(lambda x: ele.value in x[ "Album_type" ], axis=1)
              plot(df[filter])

                       ele_proxy = create_proxy(select_filter)

                  for ele in filter_elements:
          if ele.value == "ALL":
            ele.checked = True
            current_selected = ele.value
          ele.addEventListener("change", ele_proxy)

                  plot(df)

             </py-script>

          <div id="input" style="margin: 20px;">
      Select Album_type : <br/>
      <input type="radio" id="all" name="Album_type" value="ALL">
      <label for="all"> All</label>
      <input type="radio" id="album" name="Album_type" value="album">
      <label for="album"> album </label>
      <input type="radio" id="single" name="Album_type" value="single">
      <label for="single"> single </label>
      <input type="radio" id="compilation" name="Album_type" value="compilation">
      <label for="compilation"> compilation </label>

     </div>

    <py-repl>
      df
    </py-repl>

    <div id="graph-area"></div>

         </body>
     </html>

下面是互动仪表板的预览。

Image

dashboard

笔记本中的数据可视化

在本教程中,我们将使用YouTube 和 Spotify 艺术家数据集,并按照以下步骤操作,使笔记本准备好被自动化脚本使用:

  • 单击 "视图"->"单元格工具栏"->"标签",启用 Jupyter Notebook 单元格标签编辑器。这将启用标签用户界面。这很有用,因为用户需要添加自动化脚本所需的几个重要单元格标签。
  • 从托管的 URL 读取数据。请注意,PyScript 使用Pyodide来访问服务器上的任何数据。因此,我们需要在包含数据集 URL 的单元格中添加一个特殊的标签data-url。
  • 根据需要分析数据集。
  • 绘制相关详细信息。在这里,我们将分别绘制 YouTube 和 Spotify 上排名前 10 的艺术家。确保核心绘图逻辑位于一个单元格中,并在该单元格中添加data-plot标签。这表示需要作为 PyScript 的一部分运行的脚本。同时,在该单元格中添加filter-Album_type标签。这将在Album_type列上启用交互功能,该列有 3 个唯一值:专辑、单曲和合辑。您可以用任何其他合适的列来代替它。
import pandas as pd
import urllib.request
import matplotlib.pyplot as plt
url = "https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv"
urllib.request.urlretrieve(url, filename="Spotify_Youtube.csv")
df = pd.read_csv("Spotify_Youtube.csv")
df.head()
Image
df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
fig, (ax1, ax2) = plt.subplots(1, 2)
ax1.set_title('Top 10 Artists on YouTube')
df_views.plot(kind='bar', ax=ax1)
ax2.set_title('Top 10 Artists on Spotify')
df_streams.plot(kind='bar', ax=ax2)
ax1.set_xlabel('Artist')
ax1.set_ylabel('Views')
ax2.set_xlabel('Artist')
ax2.set_ylabel('Streams')

自动化脚本

上下滑动查看更多

# nb2dashboard.pyimport sys
import nbformat
import pandas as pd 
import urllib.request
from string import Template

def convert(filename):

    print(f'Converting {filename} to PyScript dashboard')

    ref = nbformat.read(filename, as_version=nbformat.NO_CONVERT)
    cells = ref['cells']

    imports = []

    for cell in cells:
        if cell["cell_type"] == "code":
            tags = cell["metadata"].get("tags")
            source = cell["source"]

            # package imports
            if "import" in source:
                import_cells = source.replace("\n", "\n" + "    ")
                for package_import in source.split("\n"):
                    pkg = package_import.split(" ")[1]
                    if "." not in pkg and pkg != "urllib":
                        imports.append(f'"{pkg}"')
                    else:
                        main_pkg = pkg.split(".")[0]
                        if main_pkg != "urllib":
                            imports.append(f'"{main_pkg}"')

            if tags:

                for tag in tags:

                    # read data url
                    if "data-url" in tags:
                        data_url = source[source.find("=") + 1 :].strip().replace('"', "")

                    # data plot code
                    if "data-plot" in tags:
                        plot_code = source.replace("\n", "\n" + "        ")

                    if "filter" in tag:
                        filter_col_name = tag.split("-")[1]
                        filter_col = f'"{filter_col_name}"'
                        urllib.request.urlretrieve(data_url, filename="dataset.csv")
                        df = pd.read_csv("dataset.csv")
                        unique_filter_col_values = df[filter_col_name].unique().tolist()

    selection = f"Select {filter_col_name} : <br/>\n"
    selection += f"<input type=\"radio\" id=\"all\" name=\"{filter_col_name}\" value=\"ALL\">\n" 
    selection += "<label for=\"all\"> All</label>\n"

    for val in unique_filter_col_values:
        selection += f"<input type=\"radio\" id=\"{val}\" name=\"{filter_col_name}\" value=\"{val}\">\n" 
        selection += f"<label for=\"{val}\"> {val} </label>\n"

    template = Template(
        """<html> 
         <head> 
         <title>YouTube and Spotify artists</title>
         <meta charset="utf-8">
         <link rel="stylesheet" href="https://pyscript.net/latest/pyscript.css" />
         <script defer src="https://pyscript.net/latest/pyscript.js"></script>
         </head>
         <body>
         <py-config> 
             packages = [ $packages ]
         </py-config>
         <py-script>

                          $user_imports

                          from pyodide.http import open_url
             from pyodide.ffi import create_proxy

                          url = $url
             df = pd.read_csv(open_url(url))

                          current_selected = []
             filter_elements = js.document.getElementsByName( $filter_column )

                          def plot(df):
                 plt.rcParams["figure.figsize"] = (15, 10)
                 $code
                 display(fig, target="graph-area", append=False)

                              def select_filter(event):
                for ele in filter_elements:
                  if ele.checked:
                      current_selected = ele.value
                      break
                if current_selected == "ALL":
                  plot(df)
                else:
                  filter = df.apply(lambda x: ele.value in x[ $filter_column ], axis=1)
                  plot(df[filter])

                               ele_proxy = create_proxy(select_filter)

                          for ele in filter_elements:
              if ele.value == "ALL":
                ele.checked = True
                current_selected = ele.value
              ele.addEventListener("change", ele_proxy)

                          plot(df)

                     </py-script>

                  <div id="input" style="margin: 20px;">
          $selection

         </div>

        <py-repl>
          df
        </py-repl>

        <div id="graph-area"></div>

                 </body>
         </html>
          """


    )

    dashboard = template.substitute(packages=(", ".join(imports)), 
                         url=f"\"{data_url}\"",
                         user_imports=import_cells,
                         code=plot_code,
                         filter_column=filter_col,
                         selection=selection)

    with open("dashboard.html", "w") as file:
        file.write(dashboard)

if __name__ == "__main__":
    notebook = sys.argv[1]
    convert(notebook)

以上内容是并将相关内容转换为 PyScript 文件。关于此脚本的一些要点

  • 脚本使用nbformat软件包读取笔记本中每个单元格的源代码和元数据标签。
  • 重要的变量会被提取出来,如数据集的url、可能的软件包导入以及绘制图表的最终代码。
  • 根据用户添加的过滤标签,脚本会从数据集中提取唯一的列值。
  • 最后,将这些变量替换到模板字符串中,生成 PyScript 代码。

现在,在与笔记本相同的目录下运行nb2dashboard.py文件。

python nb2dashboard.py <notebook_name>.ipynb

这将生成一个名为dashboard.html 的文件。确保<py-config>中提到的软件包准确无误。<py-script>中 Python 代码的缩进可能因编辑器而异,尤其是软件包导入和def plot(data) 中附加的代码。这些缩进错误可以手动修正。

转换后的独立仪表板

下面是一个PyScript 参考文件,你可以用它来验证生成的文件。导入的软件包在<py-config>标记中定义。

<py-config> 
         packages = [ "pandas", "matplotlib" ]
</py-config>

所有 Python 代码都需要包含在<py-script>标记中。数据集使用Pyodide库访问。

path = "Spotify_Youtube.csv"
df = pd.read_csv(path)

select_filter函数用于根据用户交互选择的列值选择数据集的子集。通过一组单选按钮实现了交互性,每个单选按钮都与用户在笔记本中标记的过滤列的唯一值相对应。

Image

<div id="input" style="margin: 20px;">
      Select Album_type : <br/>
      <input type="radio" id="all" name="Album_type" value="ALL">
      <label for="all"> All</label>
      <input type="radio" id="album" name="Album_type" value="album">
      <label for="album"> album </label>
      <input type="radio" id="single" name="Album_type" value="single">
      <label for="single"> single </label>
      <input type="radio" id="compilation" name="Album_type" value="compilation">
<label for="compilation"> compilation </label>
</div>
🏴‍☠️宝藏级🏴‍☠️ 原创公众号『数据STUDIO』内容超级硬核。公众号以Python为核心语言,垂直于数据科学领域,包括可戳👉Python|MySQL|数据分析|数据可视化|机器学习与数据挖掘|爬虫 等,从入门到进阶!

长按👇关注- 数据STUDIO -设为星标,干货速递ImageImage