The Stanford Natural Language Processing Group
Tokenization of raw text is a standard pre-processing step for many NLP tasks. For English, tokenization usually involves punctuation splitting and separation of some affixes like possessives. Other languages require more extensive token pre-processing, which is usually called segmentation. The Stanford Word Segmenter currently supports Arabic and Chinese. The provided segmentation schemes have been found to work well for a variety of applications.
Sep-28-2016, 15:05:45 GMT
- Country:
- North America > United States > California > Santa Clara County > Palo Alto (0.40)
- Technology: