Skip to main navigation Skip to search Skip to main content

Printed romanian modelling: A corpus linguistics based study with orthography and punctuation marks included

  • University 'Politehnica' of Bucharest
  • Romanian Academy
  • CNRS SAMOVAR UMR 5157

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper is part of a larger study dedicated by the authors to the description of printed Romanian language as an information source. Here, the statistical investigation attempts to get an answer concerning the mathematical model of the language with orthography and punctuation marks included into the alphabet. To come out to an accurate result, the authors processed the information obtained out of multiple data sets sampled from a corpus linguistics, by using the following statistical inferences: probability estimation with multiple confidence intervals, test of the hypothesis that the probability belongs to an interval, and test of the equality between two probabilities. The second type statistical error probability involved in the tests was considered. The experimental results, which are new for printed Romanian, refer to the letter, digram and trigram statistical structure in a corpus linguistics of 93 books (about 50 millions characters).

Original languageEnglish
Title of host publicationComputational Science and Its Applications - ICCSA 2007 - International Conference, Proceedings
PublisherSpringer Verlag
Pages409-423
Number of pages15
EditionPART 1
ISBN (Print)9783540744689
DOIs
Publication statusPublished - 1 Jan 2007
Externally publishedYes
EventInternational Conference on Computational Science and its Applications, ICCSA 2007 - Kuala Lumpur, Malaysia
Duration: 26 Aug 200729 Aug 2007

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
NumberPART 1
Volume4705 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

ConferenceInternational Conference on Computational Science and its Applications, ICCSA 2007
Country/TerritoryMalaysia
CityKuala Lumpur
Period26/08/0729/08/07

Keywords

  • Corpus linguistics
  • Mathematics of natural language
  • Natural language stationarity
  • Orthography and punctuation marks
  • Statistical error control

Fingerprint

Dive into the research topics of 'Printed romanian modelling: A corpus linguistics based study with orthography and punctuation marks included'. Together they form a unique fingerprint.

Cite this