pythonic生物人

将Jupyter Notebook转换为交互式仪表盘

来源:数据STUDIO
PyScript是一个开源框架,可让用户在浏览器中使用 HTML 运行 Python 程序。它是利用Pyodide、WASM和其他现代网络技术的强大功能开发而成的。PyScript 提供了一个灵活的框架,Python 开发人员,尤其是数据科学家,可以在此基础上直接在 Python 中创建可扩展的组件。

以下是使用 PyScript 的一些好处:

  • Python 程序员可以使用 PyScript 轻松编写独立的网络应用程序。
  • 统计学家或数据科学家可以生成可共享、可点击的 HTML 文件,用于查看模型预测。
  • 用户无需安装 Python 或任何其他依赖项即可与仪表板交互。当脚本在浏览器中运行时,它们会自动安装。

在本文中,我们将学习如何使用matplotlib将 Jupyter Notebook 中的一些数据可视化,然后使用自动化脚本将其自动转换为使用PyScript 的独立交互式仪表盘。


上下滑动查看更多

# dashboard.html

<html> 
     <head> 
     <title>YouTube and Spotify artists</title>
     <meta charset='utf-8'>
     <link rel='stylesheet' href='https://pyscript.net/latest/pyscript.css' />
     <script defer src='https://pyscript.net/latest/pyscript.js'></script>
     </head>
     <body>
     <py-config> 
         packages = [ 'pandas', 'matplotlib' ]
     </py-config>
     <py-script>

                  import pandas as pd
         import urllib.request
         import matplotlib.pyplot as plt

                  from pyodide.http import open_url
         from pyodide.ffi import create_proxy

                  url = 'https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv'
         df = pd.read_csv(open_url(url))

                  current_selected = []
         filter_elements = js.document.getElementsByName( 'Album_type' )

                  def plot(df):
             plt.rcParams['figure.figsize'] = (15, 10)
             df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
             df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
             fig, (ax1, ax2) = plt.subplots(1, 2)
             ax1.set_title('Top 10 Artists on YouTube')
             df_views.plot(kind='bar', ax=ax1)
             ax2.set_title('Top 10 Artists on Spotify')
             df_streams.plot(kind='bar', ax=ax2)
             ax1.set_xlabel('Artist')
             ax1.set_ylabel('Views')
             ax2.set_xlabel('Artist')
             ax2.set_ylabel('Streams')
             display(fig, target='graph-area', append=False)

                      def select_filter(event):
            for ele in filter_elements:
              if ele.checked:
                  current_selected = ele.value
                  break
            if current_selected == 'ALL':
              plot(df)
            else:
              filter = df.apply(lambda x: ele.value in x[ 'Album_type' ], axis=1)
              plot(df[filter])

                       ele_proxy = create_proxy(select_filter)

                  for ele in filter_elements:
          if ele.value == 'ALL':
            ele.checked = True
            current_selected = ele.value
          ele.addEventListener('change', ele_proxy)

                  plot(df)

             </py-script>

          <div id='input' style='margin: 20px;'>
      Select Album_type : <br/>
      <input type='radio' id='all' name='Album_type' value='ALL'>
      <label for='all'> All</label>
      <input type='radio' id='album' name='Album_type' value='album'>
      <label for='album'> album </label>
      <input type='radio' id='single' name='Album_type' value='single'>
      <label for='single'> single </label>
      <input type='radio' id='compilation' name='Album_type' value='compilation'>
      <label for='compilation'> compilation </label>

     </div>

    <py-repl>
      df
    </py-repl>

    <div id='graph-area'></div>

         </body>
     </html>

下面是互动仪表板的预览。

Image

dashboard

笔记本中的数据可视化

在本教程中,我们将使用YouTube 和 Spotify 艺术家数据集,并按照以下步骤操作,使笔记本准备好被自动化脚本使用:

  • 单击 '视图'->'单元格工具栏'->'标签',启用 Jupyter Notebook 单元格标签编辑器。这将启用标签用户界面。这很有用,因为用户需要添加自动化脚本所需的几个重要单元格标签。
  • 从托管的 URL 读取数据。请注意,PyScript 使用Pyodide来访问服务器上的任何数据。因此,我们需要在包含数据集 URL 的单元格中添加一个特殊的标签data-url。
  • 根据需要分析数据集。
  • 绘制相关详细信息。在这里,我们将分别绘制 YouTube 和 Spotify 上排名前 10 的艺术家。确保核心绘图逻辑位于一个单元格中,并在该单元格中添加data-plot标签。这表示需要作为 PyScript 的一部分运行的脚本。同时,在该单元格中添加filter-Album_type标签。这将在Album_type列上启用交互功能,该列有 3 个唯一值:专辑、单曲和合辑。您可以用任何其他合适的列来代替它。
import pandas as pd
import urllib.request
import matplotlib.pyplot as plt
url = 'https://raw.githubusercontent.com/ploomber/posts/master/notebook-to-dashboard/Spotify_Youtube.csv'
urllib.request.urlretrieve(url, filename='Spotify_Youtube.csv')
df = pd.read_csv('Spotify_Youtube.csv')
df.head()
Image
df_views = df.groupby('Artist')['Views'].sum().sort_values(ascending=False)[:10]
df_streams = df.groupby('Artist')['Stream'].sum().sort_values(ascending=False)[:10]
fig, (ax1, ax2) = plt.subplots(1, 2)
ax1.set_title('Top 10 Artists on YouTube')
df_views.plot(kind='bar', ax=ax1)
ax2.set_title('Top 10 Artists on Spotify')
df_streams.plot(kind='bar', ax=ax2)
ax1.set_xlabel('Artist')
ax1.set_ylabel('Views')
ax2.set_xlabel('Artist')
ax2.set_ylabel('Streams')

自动化脚本

上下滑动查看更多

# nb2dashboard.pyimport sys
import nbformat
import pandas as pd 
import urllib.request
from string import Template

def convert(filename):

    print(f'Converting {filename} to PyScript dashboard')

    ref = nbformat.read(filename, as_version=nbformat.NO_CONVERT)
    cells = ref['cells']

    imports = []

    for cell in cells:
        if cell['cell_type'] == 'code':
            tags = cell['metadata'].get('tags')
            source = cell['source']

            # package imports
            if 'import' in source:
                import_cells = source.replace('\n', '\n' + '    ')
                for package_import in source.split('\n'):
                    pkg = package_import.split(' ')[1]
                    if '.' not in pkg and pkg != 'urllib':
                        imports.append(f''{pkg}'')
                    else:
                        main_pkg = pkg.split('.')[0]
                        if main_pkg != 'urllib':
                            imports.append(f''{main_pkg}'')

            if tags:

                for tag in tags:

                    # read data url
                    if 'data-url' in tags:
                        data_url = source[source.find('=') + 1 :].strip().replace(''', '')

                    # data plot code
                    if 'data-plot' in tags:
                        plot_code = source.replace('\n', '\n' + '        ')

                    if 'filter' in tag:
                        filter_col_name = tag.split('-')[1]
                        filter_col = f''{filter_col_name}''
                        urllib.request.urlretrieve(data_url, filename='dataset.csv')
                        df = pd.read_csv('dataset.csv')
                        unique_filter_col_values = df[filter_col_name].unique().tolist()

    selection = f'Select {filter_col_name} : <br/>\n'
    selection += f'<input type=\'radio\' id=\'all\' name=\'{filter_col_name}\' value=\'ALL\'>\n' 
    selection += '<label for=\'all\'> All</label>\n'

    for val in unique_filter_col_values:
        selection += f'<input type=\'radio\' id=\'{val}\' name=\'{filter_col_name}\' value=\'{val}\'>\n' 
        selection += f'<label for=\'{val}\'> {val} </label>\n'

    template = Template(
        '''<html> 
         <head> 
         <title>YouTube and Spotify artists</title>
         <meta charset='utf-8'>
         <link rel='stylesheet' href='https://pyscript.net/latest/pyscript.css' />
         <script defer src='https://pyscript.net/latest/pyscript.js'></script>
         </head>
         <body>
         <py-config> 
             packages = [ $packages ]
         </py-config>
         <py-script>

                          $user_imports

                          from pyodide.http import open_url
             from pyodide.ffi import create_proxy

                          url = $url
             df = pd.read_csv(open_url(url))

                          current_selected = []
             filter_elements = js.document.getElementsByName( $filter_column )

                          def plot(df):
                 plt.rcParams['figure.figsize'] = (15, 10)
                 $code
                 display(fig, target='graph-area', append=False)

                              def select_filter(event):
                for ele in filter_elements:
                  if ele.checked:
                      current_selected = ele.value
                      break
                if current_selected == 'ALL':
                  plot(df)
                else:
                  filter = df.apply(lambda x: ele.value in x[ $filter_column ], axis=1)
                  plot(df[filter])

                               ele_proxy = create_proxy(select_filter)

                          for ele in filter_elements:
              if ele.value == 'ALL':
                ele.checked = True
                current_selected = ele.value
              ele.addEventListener('change', ele_proxy)

                          plot(df)

                     </py-script>

                  <div id='input' style='margin: 20px;'>
          $selection

         </div>

        <py-repl>
          df
        </py-repl>

        <div id='graph-area'></div>

                 </body>
         </html>
          '''


    )

    dashboard = template.substitute(packages=(', '.join(imports)), 
                         url=f'\'{data_url}\'',
                         user_imports=import_cells,
                         code=plot_code,
                         filter_column=filter_col,
                         selection=selection)

    with open('dashboard.html', 'w') as file:
        file.write(dashboard)

if __name__ == '__main__':
    notebook = sys.argv[1]
    convert(notebook)

以上内容是并将相关内容转换为 PyScript 文件。关于此脚本的一些要点

  • 脚本使用nbformat软件包读取笔记本中每个单元格的源代码和元数据标签。
  • 重要的变量会被提取出来,如数据集的url、可能的软件包导入以及绘制图表的最终代码。
  • 根据用户添加的过滤标签,脚本会从数据集中提取唯一的列值。
  • 最后,将这些变量替换到模板字符串中,生成 PyScript 代码。

现在,在与笔记本相同的目录下运行nb2dashboard.py文件。

python nb2dashboard.py <notebook_name>.ipynb

这将生成一个名为dashboard.html 的文件。确保<py-config>中提到的软件包准确无误。<py-script>中 Python 代码的缩进可能因编辑器而异,尤其是软件包导入和def plot(data) 中附加的代码。这些缩进错误可以手动修正。

转换后的独立仪表板

下面是一个PyScript 参考文件,你可以用它来验证生成的文件。导入的软件包在<py-config>标记中定义。

<py-config> 
         packages = [ 'pandas', 'matplotlib' ]
</py-config>

所有 Python 代码都需要包含在<py-script>标记中。数据集使用Pyodide库访问。

path = 'Spotify_Youtube.csv'
df = pd.read_csv(path)

select_filter函数用于根据用户交互选择的列值选择数据集的子集。通过一组单选按钮实现了交互性,每个单选按钮都与用户在笔记本中标记的过滤列的唯一值相对应。

Image

<div id='input' style='margin: 20px;'>
      Select Album_type : <br/>
      <input type='radio' id='all' name='Album_type' value='ALL'>
      <label for='all'> All</label>
      <input type='radio' id='album' name='Album_type' value='album'>
      <label for='album'> album </label>
      <input type='radio' id='single' name='Album_type' value='single'>
      <label for='single'> single </label>
      <input type='radio' id='compilation' name='Album_type' value='compilation'>
<label for='compilation'> compilation </label>
</div>

-END-

交流、合作👇

推荐阅读:
👉《Python可视化教程》,优惠中!