数据STUDIO

more-itertools,一个神奇的 Python 库

Image

Image

我之前写了一篇关于 Python itertools 库的文章,它是名副其实的“瑞士军刀”,里面有很多实用的函数。

然而学无止境,我偶然发现了一把更大的刀——more-itertools,一个神奇的 Python 库。我相信,应该有不少 Python 开发人员甚至从未听说过这个库。

more-itertools 的官方文档链接位于本文末尾,但现在,云朵君将带你了解我认为的一些更有趣和更有用的功能。

more-itertools 中的函数列表非常庞大。以下是完整列表。

上下滑动查看更多源码

+-----------------------------+-----------------------------+------------------------------------+
| Function Name               | Function Name               | Function Name                      |
+-----------------------------+-----------------------------+------------------------------------+
| ichunked                    | collapse                    | distinct_permutations              |
| chunked_even                | interleave                  | distinct_combinations              |
| sliced                      | interleave_longest          | nth_combination_with_replacement   |
| constrained_batches         | interleave_evenly           | circular_shifts                    |
| distribute                  | partial_product             | partitions                         |
| divide                      | sort_together               | set_partitions                     |
| split_at                    | value_chain                 | product_index                      |
| split_before                | zip_offset                  | combination_index                  |
| split_after                 | zip_equal                   | permutation_index                  |
| split_into                  | zip_broadcast               | combination_with_replacement_index |
| split_when                  | flatten                     | gray_product                       |
| bucket                      | roundrobin                  | outer_product                      |
| unzip                       | prepend                     | powerset_of_sets                   |
| batched                     | ilen                        | powerset                           |
| grouper                     | unique_to_each              | random_product                     |
| partition                   | sample                      | random_permutation                 |
| transpose                   | consecutive_groups          | random_combination                 |
| spy                         | map_reduce                  | random_combination_with_replacement|
| windowed                    | join_mappings               | nth_product                        |
| substrings                  | exactly_n                   | nth_permutation                    |
| substrings_indexes          | is_sorted                   | nth_combination                    |
| stagger                     | all_unique                  | always_iterable                    |
| windowed_complete           | minmax                      | always_reversible                  |
| pairwise                    | iequals                     | countable                          |
| triplewise                  | all_equal                   | consumer                           |
| sliding_window              | first_true                  | with_iter                          |       
| subslices                   | quantify                    | callback_iter                      |       
| count_cycle                 | first                       | iter_except                        |       
| intersperse                 | last                        | dft                                |              
| padded                      | one                         | idft                               |       
| mark_ends                   | only                        | convolve                           |       
| repeat_each                 | strictly_n                  | dotproduct                         |
| repeat_last                 | strip                       | factor                             |
| adjacent                    | lstrip                      | matmul                             |
| groupby_transform           | rstrip                      | polynomial_from_roots              |
| padnone                     | filter_except               | polynomial_derivative              |
| pad_none                    | map_except                  | polynomial_eval                    |
| ncycles                     | filter_map                  | sieve                              |
| nth                         | iter_suppress               | sum_of_squares                     |
| before_and_after            | nth_or_last                 | totient                            |
| take                        | unique_in_window            | locate                             |
| tail                        | duplicates_everseen         | rlocate                            |
| unique_everseen             | duplicates_justseen         | replace                            |
| unique_justseen             | classify_unique             | numeric_range                      |
| unique                      | longest_common_prefix       | side_effect                        |
| distinct_permutations       | takewhile_inclusive         | iterate                            |
+----------------------------+----------------------------+--------------------------------------+

最好将这些函数按照它们所关联的操作类型进行可视化。表格如下。

上下滑动查看更多源码

+--------------------------+-----------------------------------------------------------------------------------------------------+
| Category                 | Functions                                                                                            |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Grouping                 | chunked, ichunked, chunked_even, sliced, constrained_batches, distribute, divide, split_at,          |
|                          | split_before, split_after, split_into, split_when, bucket, unzip, batched, grouper, partition,       |
|                          | transpose                                                                                           |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Lookahead and lookback   | spy, peekable, seekable                                                                             |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Windowing                | windowed, substrings, substrings_indexes, stagger, windowed_complete, pairwise, triplewise,         |
|                          | sliding_window, subslices                                                                           |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Augmenting               | count_cycle, intersperse, padded, repeat_each, mark_ends, repeat_last, adjacent, groupby_transform, |
|                          | pad_none, ncycles                                                                                   |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Combining                | collapse, sort_together, interleave, interleave_longest, interleave_evenly, zip_offset, zip_equal,  |
|                          | zip_broadcast, flatten, roundrobin, prepend, value_chain, partial_product                           |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Summarizing              | ilen, unique_to_each, sample, consecutive_groups, run_length, map_reduce, join_mappings, exactly_n, |
|                          | is_sorted, all_equal, all_unique, minmax, first_true, quantify, iequals                             |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Selecting                | islice_extended, first, last, one, only, strictly_n, strip, lstrip, rstrip, filter_except,          |
|                          | map_except, filter_map, iter_suppress, nth_or_last, unique_in_window, before_and_after, nth, take,  |
|                          | tail, unique_everseen, unique_justseen, unique, duplicates_everseen, duplicates_justseen,           |
|                          | classify_unique, longest_common_prefix, takewhile_inclusive                                         |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Math                     | dft, idft, convolve, dotproduct, factor, matmul, polynomial_from_roots, polynomial_derivative,      |
|                          | polynomial_eval, sieve, sum_of_squares, totient                                                     |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Combinatorics            | distinct_permutations, distinct_combinations, circular_shifts, partitions, set_partitions,          |
|                          | product_index, combination_index, permutation_index, combination_with_replacement_index,            |
|                          | gray_product, outer_product, powerset, powerset_of_sets, random_product, random_permutation,        |
|                          | random_combination, random_combination_with_replacement, nth_product, nth_permutation,              |
|                          | nth_combination, nth_combination_with_replacement                                                   |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Wrapping                 | always_iterable, always_reversible, countable, consumer, with_iter, iter_except                     |
+--------------------------+-----------------------------------------------------------------------------------------------------+
| Others                   | locate, rlocate, replace, numeric_range, side_effect, iterate, difference, make_decorator,          |
|                          | SequenceView, time_limited, map_if, iter_index, consume, tabulate, repeatfunc, reshape,             |
|                          | doublestarmap                                                                                       |
+--------------------------+-----------------------------------------------------------------------------------------------------+

在深入使用这些工具的一些代码示例之前,你应该知道,你需要运行 Python 3.8 或更高版本

代码示例

现在我们看几个使用这些函数的编码示例。我们将查看上表中的几个类别,并从每个类别中选择一个函数来展示 more-itertools 库的功能。

首先,使用 pip 安装该库。

pip install more-itertool

示例 1 — 分组

使用分块函数处理数据流。

import more_itertools

defprocess_data_in_chunks(data_stream, chunk_size):
    chunk_summaries = [] 

# 将数据拆分成块
for chunk in more_itertools.chunked(data_stream, chunk_size): 
# 对于每个块,计算总和和平均值
        chunk_sum = sum (chunk) 
        chunk_avg = chunk_sum / len (chunk) 
        chunk_summaries.append((chunk_sum, chunk_avg)) 

return chunk_summaries 

# 示例数据流:一大串数字
data_stream = range(1, 23)   # 想象这是一个实时数据流

# 以 5 个块为单位处理流 
chunk_size = 5
summaries = process_data_in_chunks(data_stream, chunk_size) 
# 输出每个块的总和和平均值
for idx, (chunk_sum, chunk_avg) in enumerate(summaries, 1):
   print(f"Chunk {idx}: Sum = {chunk_sum}, Average = {chunk_avg:.2f}")


Chunk 1: Sum = 15, Average = 3.00
Chunk 2: Sum = 40, Average = 8.00
Chunk 3: Sum = 65, Average = 13.00
Chunk 4: Sum = 90, Average = 18.00
Chunk 5: Sum = 43, Average = 21.50

示例 2 — 前瞻/回顾

这些工具可以查看可迭代对象的值。

假设你正在处理大量的日志条目,并且需要检测日志中的特定模式或异常。然而,对于每个异常,你还需要回溯并重新处理之前的几条日志条目以获取上下文。

你可以使用seekable()正常处理日志,并且当检测到异常时,你可以“倒回”检查以前的条目,而无需重新启动整个流。

import more_itertools as mit

defprocess_log(log_stream):
# 使用 seekable 包装日志流以允许来回移动
    log_it = mit.seekable(log_stream, maxlen=5)   # 缓存最后 5 个日志条目

for log in log_it: 
# 正常日志处理
        print(f"Processing log: {log}") 

# 检查日志中的异常(例如,关键字“ERROR”)
if"ERROR"in log: 
            print("\nAnomaly detected! Gathering context:\n") 

# 返回前 3 个日志条目 for context
             log_it.relative_seek(-3) 

# 显示之前的日志
for _ in  range(3): 
if log_it.peek():   # 检查是否可以向前查看
                    print(f"Context log: {log_it.peek()}") 
                next (log_it)   # 在日志流中向前移动

            print(f"Anomalous log: {log}") 
            print("-" * 40 , "\n") 

# 模拟带有异常的日志流
log_stream = [
"INFO: System started",
"INFO: User logged in",
"INFO: User requested data",
"ERROR: Data not found",
"INFO: System idle",
"INFO: User logged out",
]

# 处理日志
process_log(log_stream)


Processing log: INFO: System started
Processing log: INFO: User logged in
Processing log: INFO: User requested data
Processing log: ERROR: Data not found

Anomaly detected! Gathering context:

Context log: INFO: User logged in
Context log: INFO: User requested data
Context log: ERROR: Data not found
Anomalous log: ERROR: Data not found
---------------------------------------- 

Processing log: INFO: System idle
Processing log: INFO: User logged out

示例 3 — 窗口

从可迭代对象中产生项目窗口。

在分析时间序列数据时,more_itertools.windowed()非常有用。假设你正在分析一只股票的每日价格走势,并且希望根据每个窗口中第一个元素和最后一个元素之间的变化来计算移动趋势得分。

import more_itertools as mit 

defanalyze_price_trends(prices, window_size):
# 使用滑动窗口分析价格趋势
    trend_analysis = [] 

for window in mit.windowed(prices, n=window_size, step=1): 
# 解压窗口中的第一个和最后一个价格
        first_price, *_, last_price = window 

if first_price isNoneor last_price isNone: 
# 跳过不完整的窗口(可选,根据需求而定)
continue

# 计算价格趋势(上升、下降或稳定)
if last_price > first_price: 
            trend = "Rising"
elif last_price < first_price: 
            trend = "Falling"
else : 
            trend = "Stable"

         trend_analysis.append((window, trend)) 

return trend_analysis 

# 10 天的股票价格数据示例
stock_prices = [100, 102, 104, 103, 105, 107, 106, 105, 104, 103] 

# 分析 5 天窗口内的价格趋势
window_size = 5
trends = analyze_price_trends(stock_prices, window_size) 

# 输出趋势分析
for window, trend in trends:
    print(f"Window {window} -> Trend: {trend}")



Window (100, 102, 104, 103, 105) -> Trend: Rising
Window (102, 104, 103, 105, 107) -> Trend: Rising
Window (104, 103, 105, 107, 106) -> Trend: Rising
Window (103, 105, 107, 106, 105) -> Trend: Rising
Window (105, 107, 106, 105, 104) -> Trend: Falling
Window (107, 106, 105, 104, 103) -> Trend: Falling

示例 4 — 增强

这些工具从可迭代对象中生成项目以及附加数据。

调用可迭代对象mark-ends()会产生的3 元组(is_first, is_last, item),这样可以轻松地对第一个和/或最后一个项目执行特定操作。

import more_itertools as mit 

defprocess_paginated_data( pages ):
for page_num, page in  enumerate (pages, start= 1 ): 
        print(f"Processing page {page_num}...") 

# 使用 mark_ends 确定该项目是页面上的第一个还是最后一个
for is_first, is_last, item in mit.mark_ends(page):
if is_first and page_num == 1:
                print(f"Starting page {page_num} with item: {item}")
elif is_last and page_num == len(pages):
                print(f"Last item on last page {page_num}: {item}")
else:
                print(f"Processing middle item: {item}")
        print()

# 分页数据示例(共 3 页项目)
pages = [
    ['item1_page1', 'item2_page1', 'item3_page1'],
    ['item1_page2', 'item2_page2', 'item3_page2'],
    ['item1_page3', 'item2_page3', 'item3_page3']
]

# 处理每一页,并特别处理第一/最后一项
process_paginated_data(pages)


Processing page 1...
Starting page 1with item: item1_page1
Processing middle item: item2_page1
Processing middle item: item3_page1

Processing page 2...
Processing middle item: item1_page2
Processing middle item: item2_page2
Processing middle item: item3_page2

Processing page 3...
Processing middle item: item1_page3
Processing middle item: item2_page3
Last item on last page 3: item3_page3

示例 5 — 合并

这些工具结合了多个可迭代对象。

你有三个列表:员工姓名、部门和薪资。你希望先按部门对员工进行排序,然后按薪资(每个部门内)进行排序,同时确保排序后员工的姓名、部门和薪资保持一致。下面sort_together()函数调用用例。

import more_itertools as mit 

# 员工数据
names = ['John', 'Jane', 'Dave', 'Anna', 'Zoe']
departments = ['HR', 'Engineering', 'HR', 'Engineering', 'Sales']
salaries = [50000, 80000, 55000, 75000, 60000] 

# 先按部门对员工进行排序,再按薪水排序(均按升序)
sorted_data = mit.sort_together([names, agencies, salaries], key_list=(1, 2)) 

# 解压排序后的数据
sorted_names, sorted_departments, sorted_salaries = sorted_data 

# 输出排序后的员工及其部门和薪水
for name, sector, salary in zip(sorted_names, sorted_departments, sorted_salaries): 
print(f"Employee: {name}, Department: {department}, Salary: ${salary}")


Employee: Anna, Department: Engineering, Salary: $75000
Employee: Jane, Department: Engineering, Salary: $80000
Employee: John, Department: HR, Salary: $50000
Employee: Dave, Department: HR, Salary: $55000
Employee: Zoe, Department: Sales, Salary: $60000

示例 6 - 总结

这些工具从可迭代对象返回汇总或聚合数据。

假设你要处理一个销售交易列表,其中每笔交易都包含产品、类别和销售金额。目标是:

  • 按类别对交易进行分组。
  • 将每个类别的销售额相加。

这是map_reduce()函数的一个很好的用例。

import more_itertools as mit 

# 销售数据示例(产品、类别、销售金额)
 sales_data = [ 
    ( '笔记本电脑' , '电子产品' , 1000 ), 
    ( '智能手机' , '电子产品' , 500 ), 
    ( '耳机' , '电子产品' , 150 ), 
    ( '面包' , '杂货' , 3 ), 
    ( '牛奶' , '杂货' , 2 ), 
    ( '鸡蛋' , '杂货' , 4 ), 
    ( 'T 恤' , '服装' , 20 ), 
    ( '牛仔裤' , '服装' , 40 ) 
] 

# 关键函数:按类别分组(元组中的第二项)
 keyfunc = lambda x: x[ 1 ] 

# 值函数:提取销售金额(元组中的第 3 项)
 valuefunc = lambda x: x[ 2 ] 

# Reduce 函数:对每个类别的销售额求和
reducefunc = sum 

# 执行 map-reduce 操作对销售额进行分类和求和
category_sales = mit.map_reduce(sales_data, keyfunc=keyfunc, valuefunc=valuefunc, reducefunc=reducefunc) 

# 输出按类别划分的总销售额
for category, total_sales in category_sales.items(): 
    print(f"类别:{category}, 总销售额:$ {total_sales}")


类别:电子产品,总销售额:$ 1650
类别:杂货,总销售额:$ 9
类别:服装,总销售额:$ 60

示例 7 — 选择

这里我们将查看一个从可迭代对象中产生某些项目的示例函数。

假设你正在从一个以批次(页面)形式返回结果的 API 中提取数据,并且你希望始终检索每个批次中的第 n 个项目(如果可用)。但是,如果项目少于 n 个,则检索最后一个可用项目。你还想处理批次为空的情况。你可以使用nth_or_last()函数来实现。

import more_itertools as mit

defprocess_paginated_data(pages, n, default="No data available"):
for page_num, page in  enumerate (pages, start= 1 ): 
try : 
            item = mit.nth_or_last(page, n, default) 
            print(f"Page {page_num} , Item: {item}") 
except ValueError: 
            print(f"Page {page_num} is empty.") 

# 来自 API 的分页数据示例(某些页面较短或为空)
 pages = [ 
    [ 'item1_page1' , 'item2_page1' , 'item3_page1' ], 
    [ 'item1_page2' , 'item2_page2' ],   # 少于 n 个项目
    [],   # 空页面
    [ 'item1_page4' , 'item2_page4' , 'item3_page4' , 'item4_page4' ], 
] 

# 从每页中获取第 3 个项目,如果少于 3 个项目,则获取最后一项
process_paginated_data(pages, n=3, default="No data available")

Page 1, Item: item3_page1
Page 2, Item: item2_page2
Page 3, Item: No data available
Page 4, Item: item4_page4

写在最后

more-itertools 库是 Python 内置itertools模块的强大扩展,提供了一系列用于处理可迭代对象的附加工具。它提供了用于分组、过滤、转换和汇总数据的高级函数,使开发人员能够更高效、更简洁地处理集合。

虽然我已经展示了几个用例和代码来演示这个库的实用性,但我只触及了他最基本的用法,其实他的功能远不止此,我建议你查看more-itertools官方文档[1]。

如果对你有用,不妨双击屏幕!

参考资料
[1] 

more-itertools官方文档: https://more-itertools.readthedocs.io/en/stable/api.html

🏴‍☠️宝藏级🏴‍☠️ 原创公众号『数据STUDIO』内容超级硬核。公众号以Python为核心语言,垂直于数据科学领域,包括可戳👉Python|MySQL|数据分析|数据可视化|机器学习与数据挖掘|爬虫等,从入门到进阶!

长按👇关注- 数据STUDIO -设为星标,干货速递ImageImage