AI Experts on Different Language NLP Datasets in APAC

#artificialintelligence 

Regarding the issue of different languages, generally speaking, biomedical NLP targets the languages of the scientific literature and the language of documentation in electronic health records. For the former, while much of the scientific literature is in English, it definitely isn't all, and I have been involved with efforts to work on automatic machine translation specifically for scientific texts, specifically through the Workshop on Machine Translation Biomedical task. For the latter, a key challenge is the availability of data sets and resources for working with clinical texts in different languages; clinical texts are not easy to obtain in any language. However, there are ongoing efforts to make these available, for instance for Spanish, the Biomedical Text Mining Unit at the Barcelona Supercomputing Center has run several shared tasks on Spanish-language clinical texts, and I collaborated with a team to develop a deep learning-based NLP approach for named entity recognition in Spanish clinical narratives in that context. Another challenge is'translating' complex clinical terminology to more consumer-friendly language; we have also done some early work leveraging Wikipedia for that (called WikiUMLS).