In order to solve the problems of the scarcity of Khmer lexical annotation corpora and the lack of obvious identification features of Khmer named entities, a Khmer named entity recognition method introducing English - Khmer cross - language features is proposed. Firstly, with the help of the mature model of English named entities and the word alignment relationship of English - Khmer bilingual parallel corpora, the entity categories of the source language are mapped to the target language; then, according to the Khmer word vectors, a nearest neighbor graph is constructed, and the label propagation algorithm is adopted to obtain the entity category distribution of Khmer words and complete the cross - language knowledge transfer; finally, the named entity category distribution of Khmer words is integrated into the conditional random field model as a constraint feature. The experimental results show that the conditional random field model integrated with cross - language features can effectively improve the effect of Khmer named entity recognition.
为了解决柬埔寨语词法标注语料稀缺、柬埔寨语命名实体缺乏明显标识特征的问题,提出一种引入英柬跨语言特征的柬埔寨语命名实体识别方法。首先,借助英语命名实体的成熟模型及英柬双语平行语料的词对齐关系,将源语言的实体类别映射到目标语言;然后根据柬埔寨语词向量构造最近邻图,采用标签传播算法,获得柬埔寨语单词的实体类别分布,完成跨语言知识转移;最后,将柬埔寨语单词的命名实体类别分布作为约束特征融入到条件随机场模型中。实验结果表明,融入跨语言特征的条件随机场模型能有效地提升柬埔寨语命名实体识别的效果。