A simple text segmenter
This is a python library for text segmentation of Japanese text.
- Text segmentation by simple rules,
- rule-based, no machine learning,
- so you can assume results.
- comparably fast. It's written in rust-lang.
pip install kuzukiri
pip install setuptools-rust
python -m pip install .
import kuzukiri
segmenter = kuzukiri.Segmenter()
text = "これはテストです。文分割します。"
sentences = segmenter.split(text)
print(sentences) # => ['これはテストです。', '文分割します。']
For details, see examples
and tests
directories.
MIT
- PyO3 : to compile rust code for python.
- unicode_normalization crate : for NFKC normalization