Passer à la navigation principale Passer à la recherche Passer au contenu principal

A real-time French text-to-speech system generating high-quality synthetic speech

  • E. Moulines
  • , F. Emerard
  • , D. Larreur
  • , J. L. Le Saint Milon
  • , L. Le Faucheur
  • , F. Marty
  • , F. Charpentier
  • , C. Sorin
  • Orange Labs

Résultats de recherche: Contribution à un journalArticle de conférenceRevue par des pairs

29 Citations (Scopus)

Résumé

The main features of the CNET diphone-based text-to-speech system for French language are described. The linguistic analysis works in three steps. First, a morphosyntactic analysis module assigns a grammatical value to each word in the text and transcribes it phonetically. A second module parses the text into hierarchical syntactico-prosodic groups. Finally, prosodic patterns are automatically assigned to each word by queries to a database of prosodic events. The phonetic and prosodic information serves as commands to the synthesis component. The synthesis component is based on diphone concatenation. A time-domain formulation of the pitch-synchronous overlap-add scheme (TD-PSOLA) is used to modify the speech prosody and to concatenate diphone waveforms. It is combined with a low bit-rate speech decoder to reduce the memory requirement for storing the diphone inventory. The system runs in real time on a PC equipped with a TMS320C25 DSP board and provides notably improved sound quality and naturalness in comparison to commercially available systems.

langue originaleAnglais
Pages (de - à)309-312
Nombre de pages4
journalICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Volume1
étatPublié - 1 déc. 1990
Modification externeOui
Evénement1990 International Conference on Acoustics, Speech, and Signal Processing: Speech Processing 2, VLSI, Audio and Electroacoustics Part 2 (of 5) - Albuquerque, New Mexico, USA
Durée: 3 avr. 19906 avr. 1990

Empreinte digitale

Examiner les sujets de recherche de « A real-time French text-to-speech system generating high-quality synthetic speech ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation