python结巴分词后字典排列元素（key/value对）代码详解

importjiebaimportjieba.possegstop=[line.strip()forlineinopen(‘stopwords.txt','r',enco... import jieba

import jieba.posseg

stop = [line.strip() for line in open(‘stopwords.txt' , 'r' , encoding = "utf-8").readlines()]

file = open('data.txt' , 'r')
line = file.readline()
while line!=" ":
print(" / ".join(list(word for word in jieba.cut(line,HMM=True)if word not in stop and len(word.strip())>1)))
line = file.readline()
以上完成分词及停用词
求之后字典排列元素（key/value对）代码详解
发邮箱siriuszyq@163.com也可以
降序展开

 我来答

1个回答

#热议# 什么是淋病？哪些行为会感染淋病？

日TimE寸
推荐于2016-05-31 · TA获得超过9568个赞

知道大有可为答主

回答量：1358

采纳率：83%

帮助的人：482万

我也去答题访问个人页

关注

展开全部

最复杂的就是这一行了：
(word for word in jieba.cut(line,HMM=True)if word not in stop and len(word.strip())>1)
jieba.cut(line)将一行字符串，分割成一个个单词
word for word in jieba.cut(line,HMM=True)是一个Python的表理解，相当于for循环遍历分割好的一个个单词
if word not in stop and len(word.strip())>1这仍然是表理解的一部分，如果满足条件，就把单词加入到一个新的列表中，如果不满足就丢弃，
word not in stop单词不在停用词当中
len(word.strip())>1单词去掉首尾的空格、标点符号后的长度大于1

更多追问追答
追问

我需要用字典降序排列（key/value对）的代码，前半段运行没有问题
追答

是统计每一个单词出现次数的字典吗
追问

对的，降序，最好用（key/value对），谢谢了
Ok不，挺急的
追答
import jieba
from collections import Counter

with open('stopwords.txt', 'r', encoding='utf-8') as f:
    stopwords = f.read().split('\n')

with open('data.txt', 'r', encoding = 'utf-8') as f:
    data = f.read()
words = filter(lambda w:w not in stopwords and len(w)>1, jieba.cut(data))

count = Counter(words)
count = sorted(count.items(), key=lambda x:x[1], reverse=True)结果是这样的：


追问

data是ansi的，删了encoding就好了。。结果print一下就好了。。谢了！

已赞过 已踩过<

评论收起

推荐律师服务：若未解决您的问题，请您详细描述您的问题，通过百度律临进行免费专业咨询

python结巴分词后字典排列元素（key/value对）代码详解

其他类似问题

为你推荐：