A tokenizer, text cleaner, and phonemizer for many human languages.

These details have not been verified by PyPI

Project links

Homepage

Project description

Gruut

A tokenizer, text cleaner, and IPA phonemizer for several human languages.

from gruut import text_to_phonemes

text = 'He wound it around the wound, saying "I read it was $10 to read."'

for sent_idx, word, word_phonemes in text_to_phonemes(text, lang="en-us"):
    print(word, *word_phonemes)

which outputs:

he h ˈi
wound w ˈaʊ n d
it ˈɪ t
around ɚ ˈaʊ n d
the ð ə
wound w ˈu n d
, |
saying s ˈeɪ ɪ ŋ
i ˈaɪ
read ɹ ˈɛ d
it ˈɪ t
was w ə z
ten t ˈɛ n
dollars d ˈɑ l ɚ z
to t ə
read ɹ ˈi d
. ‖

Note that "wound" and "read" have different pronunciations when used in different contexts.

See the documentation for more details.

Installation

$ pip install gruut

Additional languages can be added during installation. For example, with French and Italian support:

$ pip install gruut[fr,it]

You may also manually download language files and use the --lang-dir option:

$ gruut <lang> <command> --lang-dir /path/to/language-files/

Extracting the files to $HOME/.config/gruut/ will allow gruut to automatically make use of them. gruut will look for language files in the directory $HOME/.config/gruut/<lang>/ if the corresponding Python package is not installed. Note that <lang> here is the full language name, e.g. de-de instead of just de.

Supported Languages

gruut currently supports:

Czech (cs or cs-cz)
German (de or de-de)
English (en or en-us)
Spanish (es or es-es)
Farsi/Persian (fa)
French (fr or fr-fr)
Italian (it or it-it)
Dutch (nl)
Russian (ru or ru-ru)
Swedish (sv or sv-se)

The goal is to support all of voice2json's languages

Dependencies

Python 3.6 or higher
Linux
- Tested on Debian Buster
num2words fork and Babel
- Currency/number handling
- num2words fork includes additional language support (Arabic, Farsi, Swedish, Swahili)
gruut-ipa
- IPA pronunciation manipulation
pycrfsuite
- Part of speech tagging and grapheme to phoneme models

Command-Line Usage

The gruut module can be executed with python3 -m gruut <LANGUAGE> <COMMAND> <ARGS>

The commands are line-oriented, consuming/producing either text or JSONL. They can be composed to produce a pipeline for cleaning text.

You will probably want to install jq to manipulate the JSONL output from gruut.

tokenize

Takes raw text and outputs JSONL with cleaned words/tokens.

$ echo 'This, right here, is some RAW text!' \
    | python3 -m gruut en-us tokenize \
    | jq -c .clean_words
["this", ",", "right", "here", ",", "is", "some", "raw", "text", "!"]

See python3 -m gruut <LANGUAGE> tokenize --help for more options.

phonemize

Takes JSONL output from tokenize and produces JSONL with phonemic pronunciations.

$ echo 'This, right here, is some RAW text!' \
    | python3 -m gruut en-us tokenize \
    | python3 -m gruut en-us phonemize \
    | jq -c .pronunciation_text
ð ɪ s | ɹ aɪ t h iː ɹ | ɪ z s ʌ m ɹ ɑː t ɛ k s t ‖

See python3 -m gruut <LANGUAGE> phonemize --help for more options.

Intended Audience

gruut is useful for transforming raw text into phonetic pronunciations, similar to phonemizer. Unlike phonemizer, gruut looks up words in a pre-built lexicon (pronunciation dictionary) or guesses word pronunciations with a pre-trained grapheme-to-phoneme model. Phonemes for each language come from a carefully chosen inventory.

For each supported language, gruut includes a:

A word pronunciation lexicon built from open source data
- See pron_dict
A pre-trained grapheme-to-phoneme model for guessing word pronunciations

Some languages also include:

A pre-trained part of speech tagger built from open source data:
- See universal dependencies

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

2.4.0

Jul 3, 2024

2.3.4

Jun 17, 2022

2.3.3

Jun 17, 2022

2.3.2

May 11, 2022

2.3.1

May 11, 2022

2.3.0

Mar 30, 2022

2.2.3

Mar 17, 2022

2.2.2

Mar 11, 2022

2.2.0

Dec 6, 2021

2.1.1

Dec 3, 2021

2.1.0

Nov 10, 2021

2.0.4

Nov 5, 2021

2.0.3

Nov 1, 2021

2.0.2

Oct 19, 2021

2.0.1

Oct 15, 2021

2.0.0 yanked

Oct 14, 2021

Reason this release was yanked:

Bug fix for Python 3.6

1.3.1

Aug 2, 2021

This version

1.3.0

Jul 22, 2021

1.2.3

Jul 11, 2021

1.2.2

Jun 18, 2021

1.2.1

Jun 16, 2021

1.1.0

Jun 9, 2021

1.0.0

Jun 1, 2021

0.9.5

Apr 27, 2021

0.9.4

Apr 14, 2021

0.9.3

Apr 12, 2021

0.9.2

Mar 31, 2021

0.9.1

Mar 26, 2021

0.8.0

Mar 5, 2021

0.7.0

Mar 3, 2021

0.3.0

Oct 26, 2020

0.2.1

Oct 9, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gruut-1.3.0.tar.gz (15.5 MB view details)

Uploaded Jul 22, 2021 Source

File details

Details for the file gruut-1.3.0.tar.gz.

File metadata

Download URL: gruut-1.3.0.tar.gz
Upload date: Jul 22, 2021
Size: 15.5 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.23.0 setuptools/47.1.0 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.7.10

File hashes

Hashes for gruut-1.3.0.tar.gz
Algorithm	Hash digest
SHA256	`1507a82391b0a23ddfd0f0148f648aaf2c676568be9de1538a60e342d43e55d4`
MD5	`64f7303715d5d10add7dbb6254102053`
BLAKE2b-256	`844057554324454b6cf0160248bd55a6c1287e6ed4fede1f3a1966f1e19f4bab`