Building Open NLP Infrastructure for Kurdish — Three Dialects, One Corpus, Four Tokenizers
中文摘要
研究人员构建了库尔德语开放NLP基础设施,涵盖三种方言、统一语料库和四个分词器,支持多种脚本,服务数千万用户。
English Summary
Researchers developed open Kurdish NLP infrastructure featuring a unified corpus and four tokenizers, supporting three major dialects and multiple scripts for 30-40 million speakers.
原文节选
Kurdish is spoken by 30–40 million people, split across three major varieties — Kurmancî (Latin script), Soranî (Arabic script), and… Continue reading on Medium »