适合python练习-如何统计文本文件中的词频?
读取文本文件
#读取整个文件的内容:text = open('file.txt').read()#按行读取文本,并返回一个list,每一行是list的一个itemlines = open('file.txt').readlines()
切分单词
Google introduced its TPU at Google I/O 2016. Distinguished hardware engineer – and top MIPS CPU architect – Norm Jouppi in a blog post said Google had been running TPUs in its data centers since 2015 and that the specialized silicon delivered “an order of magnitude better-optimized performance per watt for machine learning.”
# 仅仅以空格切分:words = text.split(' ')#切分更准确的话就要使用正则表达式模块reimport re# 下面的正则表达式的含义是,# 切分符包括空白符号(空格、换行符\n, Tab符\t等看不见的符号)、# 英文逗号、英文句号.、英文问号?、感叹号!、英文冒号:# 中括号[]扩起来表示任意匹配这些符号其一即可# 最后的加号+表示如果这些符号是连续挨着的则当成一个分割符切分pattern = r'[\s,\.?!:"]+'words = re.split(pattern, text)
统计词频
使用dict这个key-value数据结构来进行统计和保存统计结果。
key就是单词,value就是单词的个数。
result = {}for w in words:if w in result:result[w] += 1else:result[2] = 1#或者用defaultdictfrom collections import defaultdictresult = defaultdict(int)for w in words:result[w] += 1