TY - GEN
T1 - Printed romanian modelling
T2 - International Conference on Computational Science and its Applications, ICCSA 2007
AU - Vlad, Adriana
AU - Mitrea, Adrian
AU - Mitrea, Mihai
PY - 2007/1/1
Y1 - 2007/1/1
N2 - This paper is part of a larger study dedicated by the authors to the description of printed Romanian language as an information source. Here, the statistical investigation attempts to get an answer concerning the mathematical model of the language with orthography and punctuation marks included into the alphabet. To come out to an accurate result, the authors processed the information obtained out of multiple data sets sampled from a corpus linguistics, by using the following statistical inferences: probability estimation with multiple confidence intervals, test of the hypothesis that the probability belongs to an interval, and test of the equality between two probabilities. The second type statistical error probability involved in the tests was considered. The experimental results, which are new for printed Romanian, refer to the letter, digram and trigram statistical structure in a corpus linguistics of 93 books (about 50 millions characters).
AB - This paper is part of a larger study dedicated by the authors to the description of printed Romanian language as an information source. Here, the statistical investigation attempts to get an answer concerning the mathematical model of the language with orthography and punctuation marks included into the alphabet. To come out to an accurate result, the authors processed the information obtained out of multiple data sets sampled from a corpus linguistics, by using the following statistical inferences: probability estimation with multiple confidence intervals, test of the hypothesis that the probability belongs to an interval, and test of the equality between two probabilities. The second type statistical error probability involved in the tests was considered. The experimental results, which are new for printed Romanian, refer to the letter, digram and trigram statistical structure in a corpus linguistics of 93 books (about 50 millions characters).
KW - Corpus linguistics
KW - Mathematics of natural language
KW - Natural language stationarity
KW - Orthography and punctuation marks
KW - Statistical error control
U2 - 10.1007/978-3-540-74472-6_33
DO - 10.1007/978-3-540-74472-6_33
M3 - Conference contribution
AN - SCOPUS:38049008977
SN - 9783540744689
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 409
EP - 423
BT - Computational Science and Its Applications - ICCSA 2007 - International Conference, Proceedings
PB - Springer Verlag
Y2 - 26 August 2007 through 29 August 2007
ER -