Pandas 教程-Pandas中删除列
整理:python架构师
drop()方法
语法:
DataFrame.drop(self, labels=None, axis=0, index=None, columns=None, level=None, inplace=False, errors='raise')
参数:
标签: 列的标签或行索引值的字符串或列表。 索引: 提供行标签。 级别: 在多索引DataFrame的情况下,用于确定应从中删除标签的级别。它接受级别位置或级别名称作为输入。 轴: 指示应删除列还是行。要删除列,请将轴设置为1或'columns'。默认情况下,它会从DataFrame中删除行。 列: 这是axis = 'columns'的替代项。它接受单个列标签或列标签列表作为输入。 Inplace: 它指定是返回新的DataFrame还是修改现有的DataFrame。它是一个默认值为False的布尔标志。 错误: 如果设置为'ignore',则忽略错误。
返回值
如果inplace = True,则返回具有删除列的DataFrame或None。 如果找不到标签,则引发KeyError。
👉点击领取:最全Python资料合集
删除单个列
示例: 我们使用df.drop(columns='col name')来从下面的示例中删除DataFrame的'age'列。
import pandas as pdstudent_dict = {"name": ["Joe", "Nat"], "age": [20, 21], "marks": [85.10, 77.80]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df)# drop columnstudent_df = student_df.drop(columns='age')print(student_df)
输出: 执行此代码后,我们将获得以下输出:
name age marks0 Joe 20 85.11 Nat 21 77.8name marks0 Joe 85.11 Nat 77.8
使用axis='column'或axis=1的drop函数
示例: 让我们以上面的示例来理解如何使用axis = 'column'和axis = 1的drop函数。
student_df = student_df.drop(['age', 'marks'], axis='columns')# alternative both generates same resultstudent_df = student_df.drop(['age', 'marks'], axis=1)
输出: 执行此代码后,我们将获得以下输出:
name age marks0 Joe 20 85.11 Nat 21 77.8name0 Joe1 Nat
删除多个列
使用column参数指定要删除的列名列表。 将轴设置为1,并移动列名列表。
示例: 让我们以一个示例来了解如何一次删除多列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [77.29, 69.15]}student_df = pd.DataFrame(student_dict)print(student_df.columns.values)# drop 2 columns at a timestudent_df = student_df.drop(columns=['age', 'marks'])print(student_df.columns.values)
输出: 执行此代码后,我们将获得以下输出:
name age marks0 John 24 77.291 Alex 18 69.15name0 John1 Alex
原地删除列
如果inplace = True,则在不返回任何内容的情况下更新当前DataFrame。 如果将inplace参数设置为False,则生成具有更新更改的新DataFrame并返回它。
示例: 让我们解释一下如何使用drop函数原地删除列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [79.18, 68.79]}student_df = pd.DataFrame(student_dict)print(student_df.columns.values)# drop columns in placestudent_df.drop(columns=['age', 'marks'], inplace=True)print(student_df.columns.values)
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 79.181 Alex 18 68.79name0 John1 Alex
通过抑制错误删除列
将errors='ignore'设置为忽略任何错误。 将errors='raised'设置为为未知列生成KeyError。
示例: 让我们通过抑制错误来删除列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [79.49, 82.54]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df)# supress errorstudent_df = student_df.drop(columns='salary', errors='ignore') # No change in the student_df# raise errorstudent_df = student_df.drop(columns='salary') # KeyError: "['salary'] not found in axis"
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 79.491 Alex 18 82.54raise KeyError(f"{labels[mask]} not found in axis")KeyError: "['salary'] not found in axis"
通过索引位置删除列
删除前n列
示例: 让我们以一个示例来了解如何删除DataFrame中的前n列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [84.45, 76.11], "class": ["A", "B"],"city": ["US", "UK"]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df.columns.values)# drop column 1 and 2student_df = student_df.drop(columns=student_df.iloc[:, range(2)])# print only columnsprint(student_df.columns.values)
输出: 执行上述代码后,我们将获得以下输出:
name age marks class city0 John 24 84.45 A US1 Alex 18 76.11 B UKmarks class city84.45 A US76.11 B UK
删除最后一列
示例: 让我们以一个示例来了解如何从DataFrame中删除最后一列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [68.44, 85.67]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df.columns.values)# find position of the last column and droppos = len(student_df.columns) - 1student_df = student_df.drop(columns=student_df.columns[pos])print(student_df.columns.values)# delete column present at index 1# student_df.drop(columns = student_df.columns[1])
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 68.441 Alex 18 85.67name age0 John 241 Alex 18
使用iloc删除列的范围
示例: 让我们以一个示例来了解如何使用iloc函数删除列的范围。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [79.64, 86.84]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df.columns.values)# drop column from 1 to 3student_df = student_df.drop(columns=student_df.iloc[:, 1:3])print(student_df.columns.values)
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 79.641 Alex 18 86.84name0 John1 Alex
从多索引DataFrame中删除列
示例: 让我们以一个示例来了解如何从多索引DataFrame中删除列。
import pandas as pd# create column headercol = pd.MultiIndex.from_arrays([['Class X', 'Class Y', 'Class Z', 'Class Y'],['Name', 'Marks', 'Name', 'Marks']])# create DataFrame from 2darraystudent_df = pd.DataFrame([['John', '87.22', 'Nat', '68.79'], ['Peter', '73.45', 'Alex', '82.76']], columns=col)print(student_df)# drop columnstudent_df = student_df.drop(columns=['Marks'], level=1)print(student_df)
输出: 执行上述代码后,我们将获得以下输出:
Class X Class Y Class Z Class YName Marks Name Marks0 John 87.22 Nat 68.791 Peter 73.45 Alex 82.76Class X Class ZName Name0 John Nat1 Peter Alex
使用函数删除列
使用pandas DataFrame.pop()函数删除列
示例: 让我们以一个示例来了解如何使用pandas DataFrame.pop()函数删除列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [24, 18], "marks": [62.46, 54.21]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df)# drop columnstudent_df.pop('age')print(student_df)
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 62.461 Alex 18 54.21name marks0 John 62.461 Alex 54.21
使用loc函数删除列
示例: 让我们以一个示例来了解如何使用loc函数删除列。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [25, 19], "marks": [79.68, 84.45]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df.columns.values)# drop column 1 and 2student_df = student_df.drop(columns=student_df.loc[:])# print only columnsprint(student_df.columns.values)
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 24 79.681 Alex 18 84.45
使用pandas DataFrame删除函数
示例: 让我们以一个示例来了解如何使用pandas DataFrame删除函数。
import pandas as pdstudent_dict = {"name": ["John", "Alex"], "age": [23, 22], "marks": [57.88, 78.84]}# Create DataFrame from dictstudent_df = pd.DataFrame(student_dict)print(student_df)# drop columndel student_df['age']print(student_df)
输出: 执行上述代码后,我们将获得以下输出:
name age marks0 John 23 57.881 Alex 22 78.84name marks0 John 57.881 Alex 78.84
热门推荐