App-Lingua-BO-Wylie-Transliteration
view release on metacpan or search on metacpan
lib/App/Lingua/BO/Wylie/Transliteration.pm view on Meta::CPAN
want to use a certain B<transliteration scheme>.
Just compare all the different names you can find "Dostojevski" transliterated
to, to see what enourmous differences there will be.
Now for the Classical Tibetan "dbu med" alphabet there exist two main
transliteration schemes:
=over
=item *
Library of Congress Transliteration
=item *
Wylie Transliteration
=back
Classical Tibetan alphabet itself works in a really interesting way.
First, let's have a look at the table of the individual "characters"
with their Wylie transliterations:
E<lt>http:E<sol>E<sol>en.wikipedia.orgE<sol>wikiE<sol>Tibetan_alphabetE<gt>
A few key observations:
=over
=item *
(Almost) all letters represent a consonant, carrying an B<inherent> vocal: "a"
=item *
These lettersE<sol>syllables are sorted according to tonality and aspiration (in pronunciation)
=item *
Other vocals will be achieved by adding certain vocal-symbols in the proper places:
=back
* i.e. you can build the following syllables by adding a vowel sign:
* ka -> ko
* ka -> ku
* ka -> ki
* ka -> ke
The latter is a process of B<merging> symbols to form new symbols (with the
merging taking place in the proper places -- 'e' is on top, 'u' will be inserted
at the bottom)
This merging process can be seen as building "ligatures", of which even more exist.
As we have seen, the vocal symbols (for all vocals except 'a', which is inherent)
need to be added in the proper places.
For the rest of the symbols that can be added, the scheme looks the following:
b s g r u b s
| | | | | | |
1 2 3 4 5 6 7
With the places being the following:
1) Prescript
2) Superscript
3) The Center piece (carriyng the inherent vocal, mandatory)
4) Subscript
5) The vocal sign
6) Postscript1
7) Postscript2
Note: Except from the Center piece (3), all other signs are optional.
Note: Optional character B<can> form B<ligatures> with the character they are combined with.
=head1 TECHNICAL BACKGROUND (UNICODE)
The Unicode consortium had to decide what they want their code points to look like:
a) either each altered base syllable is represented (i.e. ka, ko, ke, ki, ku) as a separate character (code point)
b) the base syllables are represented and the altered syllables will be merged
Since it was chosen for the latter, this has a few consequences:
=over
=item *
The graphical representation will depend on building B<ligatures> and thus on the font you are using.
=back
=over
=item *
That means you have to make sure you have appropriate fonts to display the signs correctly.
=back
Even if you find the ligatures not mixed together well (e.g. on your shell),
you can still copy-paste the results somewhere else where you have a proper
font available. Since only the code points are represented and it is up to the font
to build the ligatures you will find the copy-pasted result come out very well with
a proper font having the ligatures available.
Copy-pasting your results L<here>(http:E<sol>E<sol>www.thlib.orgE<sol>referenceE<sol>transliterationE<sol>wyconverter.php) might help should the tibetan signs not be rendered correctly on your shell
=head1 USAGE
echo bsgrubs | wylie-transliterate
or
wylie-transliterate <FILE>
( run in 2.275 seconds using v1.01-cache-2.11-cpan-d80b1682f3f )