Readability is a fundamental problem in textbooks assessment.For low resources languages(LRL),however,little investigation has been done on the readability of textbook.In this paper,we proposed a readability assessmen...Readability is a fundamental problem in textbooks assessment.For low resources languages(LRL),however,little investigation has been done on the readability of textbook.In this paper,we proposed a readability assessment method for Tibetan textbook(a low resource language).We extract features based on the information that are gotten by Tibetan segmentation and named entity recognition.Then,we calculate the correlation of different features using Pearson Correlation Coefficient and select some feature sets to design the readability formula.Fit detection,F test and T test are applied on these selected features to generate a new readability assessment formula.Experiment shows that this new formula is capable of assessing the readability of Tibetan textbooks.展开更多
基金This work was supported by the China National Natural Science Foundation No.(61331013)the Young faculty scientific research ability promotion program of Minzu University of China.
文摘Readability is a fundamental problem in textbooks assessment.For low resources languages(LRL),however,little investigation has been done on the readability of textbook.In this paper,we proposed a readability assessment method for Tibetan textbook(a low resource language).We extract features based on the information that are gotten by Tibetan segmentation and named entity recognition.Then,we calculate the correlation of different features using Pearson Correlation Coefficient and select some feature sets to design the readability formula.Fit detection,F test and T test are applied on these selected features to generate a new readability assessment formula.Experiment shows that this new formula is capable of assessing the readability of Tibetan textbooks.