PySBD: Pragmatic Sentence Boundary Disambiguation

10/19/2020
by   Nipun Sadvilkar, et al.
3

In this paper, we present a rule-based sentence boundary disambiguation Python package that works out-of-the-box for 22 languages. We aim to provide a realistic segmenter which can provide logical sentences even when the format and domain of the input text is unknown. In our work, we adapt the Golden Rules Set (a language-specific set of sentence boundary exemplars) originally implemented as a ruby gem - pragmatic_segmenter - which we ported to Python with additional improvements and functionality. PySBD passes 97.92 Golden Rule Set exemplars for English, an improvement of 25 open-source Python tool.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset