常用机器学习自动化库!
数据读取和合并,使其可供使用。 数据预处理是指数据清理和数据整理。 优化功能和模型选择过程的位置。 将其应用于应用程序以预测准确的值。
用于自动参数调整的AutoML(相对基本的类型) 用于非深度学习的AutoML,例如AutoSKlearn。此类型主要应用于数据预处理,自动特征分析,自动特征检测,自动特征选择和自动模型选择。 用于深度学习/神经网络的AutoML,包括NAS和ENAS以及用于框架的Auto-Keras。
为什么需要AutoML?
AutoML三大优点
它通过自动化最重复的任务来提高效率。这使数据科学家可以将更多的时间投入到问题上,而不是模型上。
自动化的ML管道还有助于避免由手工作业引起的潜在错误。
AutoML是朝着机器学习民主化迈出的一大步,它使每个人都可以使用ML功能。
以下是用Python实现
auto-sklearn
详细原理与案例请见(点击查看):一文彻底搞懂自动机器学习AutoML:Auto-Sklearn
例子
import sklearn.model_selectionimport sklearn.datasetsimport sklearn.metricsimport autosklearn.regressiondef main():X, y = sklearn.datasets.load_boston(return_X_y=True)feature_types = (['numerical'] * 3) + ['categorical'] + (['numerical'] * 9)X_train, X_test, y_train, y_test = \sklearn.model_selection.train_test_split(X, y, random_state=1)automl = autosklearn.regression.AutoSklearnRegressor(time_left_for_this_task=120,per_run_time_limit=30,tmp_folder='/tmp/autosklearn_regression_example_tmp',output_folder='/tmp/autosklearn_regression_example_out',)automl.fit(X_train, y_train, dataset_name='boston',feat_type=feature_types)print(automl.show_models())predictions = automl.predict(X_test)print("R2 score:", sklearn.metrics.r2_score(y_test, predictions))if __name__ == '__main__':main()
FeatureTools
图片
安装
python -m pip install featuretools
conda install -c conda-forge featuretools
附加组件
python -m pip install featuretools[complete]
python -m pip install featuretools[update_checker]
python -m pip install featuretools[tsfresh]
例
import featuretools as ftes = ft.demo.load_mock_customer(return_entityset=True)es.plot()
图片
feature_matrix, features_defs = ft.dfs(entityset=es,target_entity="customers")feature_matrix.head(5)
官方网站:
https://featuretools.alteryx.com/cn/stable/
MLBox
图片
MLBox是功能强大的自动化机器学习python库。详细原理与案例请见(点击查看)一文彻底掌握自动机器学习AutoML:MLBox
根据官方文档,它具有以下功能:
快速读取和分布式数据预处理/清理/格式化
高度强大的功能选择和泄漏检测以及精确的超参数优化
最新的分类和回归预测模型(深度学习,堆叠,LightGBM等)
使用模型解释进行预测,MLBox已在Kaggle上进行了测试,并显示出良好的性能。
管道
MLBox体系结构
MLBox主软件包包含3个子软件包:
预处理:读取和预处理数据
优化:测试或优化各种学习者
预测:预测测试数据集上的目标
TPOT
分类
from tpot import TPOTClassifierfrom sklearn.datasets import load_digitsfrom sklearn.model_selection import train_test_splitdigits = load_digits()X_train, X_test, y_train, y_test = train_test_split(digits.data,digits.target,train_size=0.75,test_size=0.25,random_state=42)tpot = TPOTClassifier(generations=5, population_size=50,verbosity=2, random_state=42)tpot.fit(X_train, y_train)print(tpot.score(X_test, y_test))tpot.export('tpot_digits_pipeline.py')
此代码将发现达到98%的测试精度的管道。应将相应的Python代码导出到tpot_digits_pipeline.py文件,其外观类似于以下内容:
import numpy as npimport pandas as pdfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import make_pipeline, make_unionfrom sklearn.preprocessing import PolynomialFeaturesfrom tpot.builtins import StackingEstimatorfrom tpot.export_utils import set_param_recursive# NOTE: Make sure that the outcome column is labeled 'target' in the data filetpot_data = pd.read_csv('PATH/TO/DATA/FILE', sep='COLUMN_SEPARATOR', dtype=np.float64)features = tpot_data.drop('target', axis=1)training_features, testing_features, training_target, testing_target = \train_test_split(features, tpot_data['target'], random_state=42)# Average CV score on the training set was: 0.9799428471757372exported_pipeline = make_pipeline(PolynomialFeatures(degree=2, include_bias=False, interaction_only=False),StackingEstimator(estimator=LogisticRegression(C=0.1, dual=False, penalty="l1")),RandomForestClassifier(bootstrap=True, criterion='entropy',max_features=0.35000000000000003,min_samples_leaf=20, min_samples_split=19,n_estimators=100))# Fix random state for all the steps in exported pipelineset_param_recursive(exported_pipeline.steps, 'random_state', 42)exported_pipeline.fit(training_features, training_target)results = exported_pipeline.predict(testing_features)
回归
TPOT可以优化管道以解决回归问题。以下是使用波士顿房屋价格数据集的最小工作示例。
from tpot import TPOTRegressorfrom sklearn.datasets import load_bostonfrom sklearn.model_selection import train_test_splithousing = load_boston()X_train, X_test, y_train, y_test = train_test_split(housing.data,housing.target,train_size=0.75,test_size=0.25,random_state=42)tpot = TPOTRegressor(generations=5, population_size=50,verbosity=2, random_state=42)tpot.fit(X_train, y_train)print(tpot.score(X_test, y_test))tpot.export('tpot_boston_pipeline.py')
import numpy as npimport pandas as pdfrom sklearn.ensemble import ExtraTreesRegressorfrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import PolynomialFeaturesfrom tpot.export_utils import set_param_recursive# NOTE: Make sure that the outcome column is labeled 'target' in the data filetpot_data = pd.read_csv('PATH/TO/DATA/FILE',sep='COLUMN_SEPARATOR',dtype=np.float64)features = tpot_data.drop('target', axis=1)training_features, testing_features, training_target, testing_target = \train_test_split(features, tpot_data['target'], random_state=42)# Average CV score on the training set was: -10.812040755234403exported_pipeline = make_pipeline(PolynomialFeatures(degree=2, include_bias=False, interaction_only=False),ExtraTreesRegressor(bootstrap=False, max_features=0.5,min_samples_leaf=2, min_samples_split=3,n_estimators=100))# Fix random state for all the steps in exported pipelineset_param_recursive(exported_pipeline.steps, 'random_state', 42)exported_pipeline.fit(training_features, training_target)results = exported_pipeline.predict(testing_features)
Github链接
https://github.com/EpistasisLab/tpot
Lightwood
安装
pip3 install lightwood
from lightwood import Predictor
import pandassensor3_predictor = Predictor(output=['sensor3']).learn(from_data=pandas.read_csv('sensor_data.csv'))
prediction = sensor3_predictor.predict(when={'sensor1':1, 'sensor2':-1})
官方链接:
https://github.com/mindsdb/lightwood
MindsDB
官方链接:
https://github.com/mindsdb/mindsdb
mljar-supervised
图片
解释和理解您的数据,
尝试许多不同的机器学习模型,
通过分析创建有关所有模型的详细信息的Markdown报告,
保存,重新运行和加载分析和ML模型。
解释模式,非常适合于解释和理解数据,其中包含许多数据解释,例如决策树可视化,线性模型系数显示,排列重要性和数据的SHAP解释, 执行构建用于生产的ML管道, 竞争模式,用于训练具有集成和堆叠功能的高级ML模型,目的是用于ML竞赛中。
官方链接:
https://github.com/mljar/mljar-supervisedv
Auto-Keras
图片
官方链接:
https://github.com/keras-team/autokeras
神经网络智能 NNI
Ludwig
无需编码:不需要任何编码技能即可训练模型并将其用于获取预测。 通用性:新的基于数据类型的深度学习模型设计方法使该工具可在许多不同的用例中使用。 灵活性:经验丰富的用户对模型的建立和培训具有广泛的控制权,而新用户则会发现它易于使用。 可扩展性:易于添加新的模型架构和新的特征数据类型。 可理解性:深度学习模型的内部通常被认为是黑匣子,但是路德维希(Ludwig)提供了标准的可视化效果来了解其性能并比较其预测。 开源:Apache License 2.0
官方链接
https://github.com/uber/ludwig
AdaNet
易于使用:提供熟悉的API(例如Keras,Estimator)用于训练,评估和提供模型。 速度:可用计算进行扩展,并快速生成高质量的模型。 灵活性:允许研究人员和从业人员将AdaNet扩展到新颖的子网体系结构,搜索空间和任务。 学习保证:优化提供理论学习保证的目标。
Darts
官方链接
https://github.com/quark0/darts
automl-gs
官方链接
https://github.com/minimaxir/automl-gs
AutoKeras的R接口
官方文档
https://github.com/r-tensorflow/autokeras
以下是用Scala实现
TransmogrifAI
数小时而不是数月内即可构建生产就绪的机器学习应用程序 在没有博士学位的情况下建立机器学习模型在机器学习中 构建模块化,可重用,强类型的机器学习工作流程
官方链接
https://github.com/salesforce/TransmogrifAI
以下是用Java实现
Glaucus
接收多源数据集,包括结构化,文档和图像数据; 提供丰富的数学统计功能,图形界面使用户轻松掌握数据情况; 在自动模式下,我们实现了从预处理,特征工程到机器学习算法的全管道自动化; 在手动模式下,它极大地简化了机器学习流程,并提供了自动数据清理,半自动特征选择和深度学习套件。
官方网站
https://github.com/ccnt-glaucus/glaucus
H20 AutoML
官方链接
https://github.com//h2oai/h2o-3/blob/master/h2o-docs/src/product/automl.rst
PocketFlow
官方链接
https://github.com/Tencent/PocketFlow
Ray
Tune:可伸缩超参数调整 RLlib:可扩展的强化学习 RaySGD:分布式培训包装器 Ray Serve:可扩展和可编程服务
官方链接
https://github.com/ray-project/ray
SMAC3
结论
-END- -微信交流- -猜您喜欢👇
如何学Matplotlib+ggplot2? 天天在用的P值到底是个啥? pythonic生物人去哪了? 《ggplot2: Elegant.....》中文版 最有价值50图表(Python代码) Jupyter Notebook的16个插件!