bamtercelboo / Corpus_process_script
Licence: apache-2.0
chinese and english corpus process script, python, c++, java
Stars: ✭ 141
Programming Languages
python
139335 projects - #7 most used programming language
Introduction
这里将会有中英文数据处理脚本,编程语言不限,会有详细的README说明。
Script Lists
- 中文繁体转简体
- 维基百科数据处理
- 抽取单字特征
- 抽取双字特征
- 抽取汉字笔画信息
- 去除非中文字符
- 中文Money转换数字Money
- 全半角转换
- python2代码批量转换python3
- NER标签转换(BIO, BMESO)
Question
-
if you have any question, you can open a issue or email [email protected]{gmail.com, 163.com}.
-
if you have any good suggestions, you can PR or email me.
Note that the project description data, including the texts, logos, images, and/or trademarks,
for each open source project belongs to its rightful owner.
If you wish to add or remove any projects, please contact us at [email protected].