Get our free extension to see links to code for papers anywhere online!


Adaptive Sentence Boundary Disambiguation

Add code

Nov 22, 1994
David D. Palmer, Marti A. Hearst


Share this with someone who'll enjoy it:


Labeling of sentence boundaries is a necessary prerequisite for many natural language processing tasks, including part-of-speech tagging and sentence alignment. End-of-sentence punctuation marks are ambiguous; to disambiguate them most systems use brittle, special-purpose regular expression grammars and exception rules. As an alternative, we have developed an efficient, trainable algorithm that uses a lexicon with part-of-speech probabilities and a feed-forward neural network. After training for less than one minute, the method correctly labels over 98.5\% of sentence boundaries in a corpus of over 27,000 sentence-boundary marks. We show the method to be efficient and easily adaptable to different text genres, including single-case texts.

* Proceedings of ANLP 94 
* This is a Latex version of the previously submitted ps file (formatted as a uuencoded gz-compressed .tar file created by csh script). The software from the work described in this paper is available by contacting dpalmer@cs.berkeley.edu 


   Access Paper Source



Share this with someone who'll enjoy it: