最新文章专题视频专题问答1问答10问答100问答1000问答2000关键字专题1关键字专题50关键字专题500关键字专题1500TAG最新视频文章推荐1 推荐3 推荐5 推荐7 推荐9 推荐11 推荐13 推荐15 推荐17 推荐19 推荐21 推荐23 推荐25 推荐27 推荐29 推荐31 推荐33 推荐35 推荐37视频文章20视频文章30视频文章40视频文章50视频文章60 视频文章70视频文章80视频文章90视频文章100视频文章120视频文章140 视频2关键字专题关键字专题tag2tag3文章专题文章专题2文章索引1文章索引2文章索引3文章索引4文章索引5123456789101112131415文章专题3
当前位置: 首页 - 科技 - 知识百科 - 正文

python中如何去除标点符号

来源:动视网 责编:小采 时间:2020-11-27 14:06:44
文档

python中如何去除标点符号

python中如何去除标点符号:Python去掉标点符号的方法如下:方法一:str.isalnum:S.isalnum() -> bool返回值:如果string至少有一个字符并且所有字符都是字母或数字则返回True,否则返回False。实例:>>> string = "Special $#! char
推荐度:
导读python中如何去除标点符号:Python去掉标点符号的方法如下:方法一:str.isalnum:S.isalnum() -> bool返回值:如果string至少有一个字符并且所有字符都是字母或数字则返回True,否则返回False。实例:>>> string = "Special $#! char


Python去掉标点符号的方法如下:

方法一:

str.isalnum:

S.isalnum() -> bool

返回值:如果string至少有一个字符并且所有字符都是字母或数字则返回True,否则返回False。

实例:

>>> string = "Special $#! characters spaces 888323"
>>> ''.join(e for e in string if e.isalnum())
'Specialcharactersspaces888323'

只能识别字母和数字,杀伤力大,会把中文、空格之类的也干掉

方法二:

string.punctuation

import re, string

s ="string. With. Punctuation?" # Sample string 

# 写法一:
out = s.translate(string.maketrans("",""), string.punctuation)

# 写法二:
out = s.translate(None, string.punctuation)

# 写法三:
exclude = set(string.punctuation)
out = ''.join(ch for ch in s if ch not in exclude)

# 写法四:
>>> for c in string.punctuation:
	s = s.replace(c,"")
>>> s
'string With Punctuation'

# 写法五:
out = re.sub('[%s]' % re.escape(string.punctuation), '', s)
## re.escape:对字符串中所有可能被解释为正则运算符的字符进行转义

# 写法六:
# string.punctuation 只包括 ascii 格式; 想要一个包含更广(但是更慢)的方法是使用: unicodedata module :
from unicodedata import category
s = u'String — with - ?Punctuation ?...'
out = re.sub('[%s]' % re.escape(string.punctuation), '', s)
print 'Stripped', out
# 
输出:u'Stripped String u2014 with xabPunctuation xbb' out = ''.join(ch for ch in s if category(ch)[0] != 'P') print 'Stripped', out # 输出:u'Stripped String with Punctuation ' # For Python 3 str or Python 2 unicode values, str.translate() only takes a dictionary; codepoints (integers) are looked up in that mapping and anything mapped to None is removed. # To remove (some?) punctuation then, use: import string remove_punct_map = dict.fromkeys(map(ord, string.punctuation)) s.translate(remove_punct_map) # Your method doesn't work in Python 3, as the translate method doesn't accept the second argument any more. import unicodedata import sys tbl = dict.fromkeys(i for i in range(sys.maxunicode) if unicodedata.category(chr(i)).startswith('P')) def remove_punctuation(text): return text.translate(tbl)

方法三:

re

例:

import re
s ="string. With. Punctuation?"
s = re.sub(r'[^ws]','',s)

测试:

import re, string, timeit

s ="string. With. Punctuation"

exclude = set(string.punctuation)
table = string.maketrans("","")
regex = re.compile('[%s]' % re.escape(string.punctuation))

def test_set(s):
	return ''.join(ch for ch in s if ch not in exclude)

def test_re(s): 
	return regex.sub('', s)

def test_trans(s):
	return s.translate(table, string.punctuation)

def test_repl(s):
	for c in string.punctuation:
	s=s.replace(c,"")
	return s

print"sets :",timeit.Timer('f(s)', 'from __main__ import s,test_set as f').timeit(1000000)
print"regex :",timeit.Timer('f(s)', 'from __main__ import s,test_re as f').timeit(1000000)
print"translate :",timeit.Timer('f(s)', 'from __main__ import s,test_trans as f').timeit(1000000)
print"replace :",timeit.Timer('f(s)', 'from __main__ import s,test_repl as f').timeit(1000000)

out_put:
# sets : 19.8566138744
# regex : 6.86155414581
# translate : 2.12455511093
# replace : 28.4436721802

更多Python相关技术文章,请访问Python教程栏目进行学习!

文档

python中如何去除标点符号

python中如何去除标点符号:Python去掉标点符号的方法如下:方法一:str.isalnum:S.isalnum() -> bool返回值:如果string至少有一个字符并且所有字符都是字母或数字则返回True,否则返回False。实例:>>> string = "Special $#! char
推荐度:
  • 热门焦点

最新推荐

猜你喜欢

热门推荐

专题
Top