Pandas 教程-数据处理
整理:python架构师
在数据分析和建模的大部分时间中,都花费在数据准备和处理上,即加载、清理和重新排列数据等。此外,由于Python库,Pandas为我们提供了高性能、灵活和高级的数据处理环境。对于 Pandas,有各种功能可用于有效地处理数据。
层次化索引
为了增强数据处理的能力,我们必须使用一些索引,这有助于根据标签对数据进行排序。因此,层次化索引出现在这里,并被定义为 Pandas 的一个基本特性,可以帮助我们使用多个索引级别。
创建多重索引
在层次化索引中,我们必须为数据创建多个索引。这个例子创建了一个具有多个索引的 Series。
import pandas as pdinfo = pd.Series([11, 14, 17, 24, 19, 32, 34, 27],index = [['x', 'x', 'x', 'x', 'y', 'y', 'y', 'y'],['obj1', 'obj2', 'obj3', 'obj4', 'obj1', 'obj2', 'obj3', 'obj4']])data
aobj1 11obj2 14obj3 17obj4 24bobj1 19obj2 32obj3 34obj4 27dtype: int64
info.index
MultiIndex(levels=[['x', 'y'], ['obj1', 'obj2', 'obj3', 'obj4']],labels=[[0, 0, 0, 0, 1, 1, 1, 1], [0, 1, 2, 3, 0, 1, 2, 3]])
专属福利 👉点击领取:最全Python资料合集
部分索引
部分索引可以被定义为从分层索引中选择特定索引的一种方法。
import pandas as pdinfo = pd.Series([11, 14, 17, 24, 19, 32, 34, 27],index = [['x', 'x', 'x', 'x', 'y', 'y', 'y', 'y'],['obj1', 'obj2', 'obj3', 'obj4', 'obj1', 'obj2', 'obj3', 'obj4']])info['b']
obj1 19obj2 32obj3 34obj4 27dtype: int64
info[:, 'obj2']
x 14y 32dtype: int64
数据的展开
展开意味着将行标头更改为列标头。行索引将更改为列索引,因此 Series 将变成 DataFrame。以下是展开数据的示例。
import pandas as pdinfo = pd.Series([11, 14, 17, 24, 19, 32, 34, 27],index = [['x', 'x', 'x', 'x', 'y', 'y', 'y', 'y'],['obj1', 'obj2', 'obj3', 'obj4', 'obj1', 'obj2', 'obj3', 'obj4']])# unstack on first level i.e. x, y#note that data row-labels are x and ydata.unstack(0)
输出:
abobj1 11 19obj2 14 32obj3 17 34obj4 24 27# unstack based on second level i.e. 'obj'info.unstack(1)
obj1 obj2 obj3 obj4a 11 14 17 24b 19 32 34 27
import pandas as pdinfo = pd.Series([11, 14, 17, 24, 19, 32, 34, 27],index = [['x', 'x', 'x', 'x', 'y', 'y', 'y', 'y'],['obj1', 'obj2', 'obj3', 'obj4', 'obj1', 'obj2', 'obj3', 'obj4']])# unstack on first level i.e. x, y#note that data row-labels are x and ydata.unstack(0)d.stack()
aobj1 11obj2 14obj3 17obj4 24bobj1 19obj2 32obj3 34obj4 27dtype: int64
列索引
import numpy as npinfo = pd.DataFrame(np.arange(12).reshape(4, 3),index = [['a', 'a', 'b', 'b'], ['one', 'two', 'three', 'four']],columns = [['num1', 'num2', 'num3'], ['x', 'y', 'x']] ... )info
num1 num2 num3x y xa one0 1 2two3 4 5b three 6 7 8four 9 10 11# display row indexinfo.index
MultiIndex(levels=[['x', 'y'], ['four', 'one', 'three', 'two']], labels=[[0, 0, 1, 1], [1, 3, 2, 0]])# display column indexinfo.columns
MultiIndex(levels=[['num1', 'num2', 'num3'], ['green', 'red']], labels=[[0, 1, 2], [1, 0, 1]])
交换和排序级别
import numpy as npinfo = pd.DataFrame(np.arange(12).reshape(4, 3),index = [['a', 'a', 'b', 'b'], ['one', 'two', 'three', 'four']],columns = [['num1', 'num2', 'num3'], ['x', 'y', 'x']] ... )info.swaplevel('key1', 'key2')nnum1 num2 num3p x y xkey2 key1onea 0 1 2twoa 3 4 5three b 6 7 8four b 9 10 11
info.sort_index(level='key2')nnum1 num2 num3p x y xkey1 key2bfour 9 10 11aone 0 1 2bthree 6 7 8atwo 3 4 5
1 热门推荐