anka-tts
English · Türkçe
A thin wrapper over F5-TTS (MIT) for zero-shot
Turkish voice cloning. It is not a new inference runtime: the model architecture,
the ODE solver and the vocoder all come from f5-tts as they are. What the
package adds is three things — Turkish text normalization (numbers, suffixes,
abbreviations, dates/times/money), reference audio validation (it tells you
F5's undocumented expectations before synthesis) and measured defaults.
Installation
There are two separate installs; which one you need depends on what you are doing.
pip install anka-tts # text normalization + reference validation
pip install anka-tts[tts] # the above + synthesis
anka-tts |
anka-tts[tts] |
|
|---|---|---|
| dependencies | numpy, soundfile | + torch, torchaudio, f5-tts, huggingface-hub |
| download size | a few MB | ~2 GB |
anka.text |
✅ | ✅ |
anka.audio |
✅ | ✅ |
AnkaTTS.synthesize() |
❌ | ✅ |
Someone who only wants to normalize text should not have to download 2 GB of
PyTorch: anka.text and anka.audio run on the pure standard library plus
soundfile. Torch is loaded only when AnkaTTS.from_pretrained() is called.
In five lines
from anka import AnkaTTS
tts = AnkaTTS.from_pretrained("kullanici/anka-tts-tr", device="cuda")
tts.add_voice("anlatici", "referans.wav", "Bugün İstanbul'da hafif bir rüzgar var.")
wav = tts.synthesize("2026'da toplantı 14:30'da mı?", voice="anlatici")
tts.save_wav(wav, "cikti.wav")
Without synthesis, normalization only:
from anka.text import normalize
normalize("15.03.2026'da saat 14:30'da %20 indirim başlıyor.")
# 'on beş mart iki bin yirmi altıda saat on dört otuzda yüzde yirmi indirim başlıyor.'
Reference audio guide
The single most important thing determining cloning quality is the reference recording — more so than the model or the parameters. The rules:
| good | bad | |
|---|---|---|
| duration | 7-9 s | < 5 s (the model cannot pick up the voice), > 12 s (F5 clips it, the rest is wasted) |
| content | a complete sentence ending on falling intonation | a fragment cut off mid-sentence |
| transcript | punctuated, word for word what is said | unpunctuated, approximate, incomplete |
| recording | single speaker, quiet room, 24 kHz+ | noise, music, reverb, 16 kHz |
| level | normal recording level | normalized afterwards (raises the noise floor) |
# Good
ref_text = "Bugün İstanbul'da hafif bir rüzgar var. Denizin kokusu sokaklara yayılıyor."
# Bad -- a sentence fragment, no punctuation, does not match the audio word for word
ref_text = "bugün istanbulda hafif bir rüzgar"
ref_text must match the audio word for word. F5 estimates the duration of
the generated speech from the ratio ref_duration / ref_text_length; if the
transcript is too long or too short, that ratio breaks and the model swallows
syllables.
Check before synthesizing:
from anka.audio import check_reference
check = check_reference("referans.wav", ref_text, gen_text="Üretilecek metin.")
print(check.duration, check.estimated_chunks)
print(check.ok) # True only when nothing at all was flagged
print(check.fatal) # True when the reference cannot be used (missing ref_text)
for w in check.warnings:
print("!", w)
ok is the strict answer: it goes False on any warning at all, a 16 kHz
recording or a 3-second one included. fatal is the narrow one -- the
reference cannot be used, which today means only a missing ref_text. So test
ok when you want a clean reference and fatal when you only want to know
whether synthesis can go ahead.
synthesize() already calls this automatically and surfaces the findings as
warnings; it can be turned off with validate_reference=False.
Why 7-9 seconds
F5 computes the generation window like this (utils_infer.py):
max_chars = (ref_bytes / ref_duration) × (22 − ref_duration) × speed
= speaking_rate × (22 − ref_duration) × speed
There is a 22-second budget: the more room the reference takes up, the less is left for generation. Which means a longer reference splits the text into more chunks. A short reference has a different problem — not chunking, but the model not picking up enough of the voice characteristics. 7-9 s is the band between the two.
Note: the formula counts UTF-8 bytes, not characters. Since ğ ü ş ı ö ç
take two bytes each, Turkish text splits earlier than English text of the same
length.
Long texts
Once the budget fills up, F5 splits the text at any of [;:,.!?] — that is,
at a comma, mid-sentence. With a 12-second reference it looks like this:
2. Bavulları akşamdan hazırlarsak,
3. sabah yola çıkarken hiç acele etmeyiz.
A pause landing there sounds wrong to the ear, and since it depends entirely on the reference duration, it is outside the control of whoever wrote the text.
That is why anka does not leave the splitting to F5: it first breaks the text
into sentences, synthesizes the sentence groups that fit the budget separately,
then joins them with a short pause in between. A seam always lands at the end
of a sentence — where the ear expects a pause anyway. Whether the reference is
7 s or 12 s, sentences stay whole.
Paragraphs separated by a blank line are preserved, and a longer pause is placed at the paragraph transition. Sentences are never grouped across a paragraph boundary.
tts.synthesize(uzun_metin, voice="anlatici") # defaults: 0.15 s / 0.45 s
tts.synthesize(uzun_metin, voice="anlatici", sentence_pause=0.2, paragraph_pause=0.6)
tts.synthesize(uzun_metin, voice="anlatici", chunk_by_sentence=False) # F5's own splitting
A sentence that exceeds the budget on its own is passed to F5 unsplit; F5 makes its own decision inside that sentence, but the others are unaffected.
You do not need to change any setting to produce both short and long text with the same voice. A one-sentence text goes out in a single call and forms no seam at all; a two-paragraph text is broken into sentences. The reference duration only affects how many pieces there will be, not where the splits happen.
Console output
F5 prints a block of its own for every chunk it generates — the reference
transcript, each gen_text, blank lines and a progress bar. Since anka calls it
once per sentence group, that block repeats for every segment. anka captures it
and shows its own view instead:
anka model anka_tts/v0.1 · mps · 24 kHz
anka voice anlatici · 7.8s reference
anka warn the text will be split into 3 chunks, and every seam between them
is a potentially audible pause. AnkaTTS aligns those seams with
sentence ends, where a pause is expected anyway…
✓ 1/3 Dr. Ayşe Yılmaz'ın 15.03.2026 tarihli raporuna göre %12 arttı. 2.6s
✓ 2/3 Sabah 09:05'te başlayan toplantı 14:30'da sona erdi. 2.6s
▸ 3/3 ████████████░░░░░░ 00:02 Katılımcılara 250 ml su dağıtıldı.
Segments are named in the wording you wrote, not the normalized text the
model is handed: Dr. Ayşe Yılmaz'ın 15.03.2026 rather than doktor Ayşe Yılmazın on beş mart iki bin yirmi altı. The two are aligned sentence by
sentence; where they cannot be aligned, the normalized text is shown rather
than a wrong guess. A sentence too long for the budget is generated in pieces,
and the pieces after the first are marked ….
Warnings are routed through the same renderer. Left to Python they are written
straight to stderr — landing in the middle of the live line and leaving a
mangled row — and printed with the file, line number and source line, none of
which help whoever is listening to the result. Only the message is shown, and a
warning raised by another library keeps its origin as a caption
(torch: …), so it stays clear who is complaining.
progress= decides what is drawn. F5's own output is captured either way.
| value | behaviour |
|---|---|
"auto" (default) |
the live view above on an interactive terminal, silence anywhere else |
False / None |
silence |
"plain" |
one line per segment, no escape sequences — for logs and CI |
"pretty" / True |
force the live view even when redirected |
| a callable | callback(done, total, segment_text) after each segment |
tts.synthesize(metin, voice="anlatici") # auto
tts.synthesize(metin, voice="anlatici", progress=False) # silent
tts.synthesize(metin, voice="anlatici", progress="plain") # one line per segment
Two kinds of third-party noise are dropped, because they are addressed to the author of the library that raised them rather than to you:
- torch's
An output with one or more elements was resized…, which fires on everystft/istftcall vocos and F5 make; - the hub backend's
You are sending unauthenticated requests to the HF Hub, printed on every load because F5 fetches the vocos vocoder from the hub. SetHF_TOKENto silence it at the source, orHF_HUB_OFFLINE=1to skip the lookup once the vocoder is cached.
Download progress bars are never touched: a silent 2 GB fetch would be worse than the noise.
Set ANKA_DEBUG=1 to see everything that was hidden — F5's prints, the muted
warnings and the hub notice — on stderr. NO_COLOR is honoured, and the
rendering falls back to plain lines on a dumb terminal.
Normalization
normalize() turns text into something that can be handed to the model.
Punctuation is preserved — ? above all, because the model derives question
intonation from it.
| input | output |
|---|---|
1994'te |
bin dokuz yüz doksan dörtte |
5'de (typed wrong) |
beşte |
4'ü |
dördü |
6'nın |
altının |
90'lar |
doksanlar |
3. sınıf |
üçüncü sınıf |
Sıralamada 3. |
(left alone — it may be the end of a sentence) |
15.03.2026 |
on beş mart iki bin yirmi altı |
14:30 / 14:00 / 09:05 |
on dört otuz / on dört / dokuz sıfır beş |
1.250 TL |
bin iki yüz elli lira |
3,5 TL |
üç lira elli kuruş |
2,75 $ |
iki dolar yetmiş beş sent |
2,5 milyon TL |
iki virgül beş milyon lira |
%20 / %50'si |
yüzde yirmi / yüzde ellisi |
5 km'de |
beş kilometrede |
36,6 °C |
otuz altı virgül altı derece |
II. Abdülhamid |
ikinci Abdülhamid |
TBMM'de |
te be me mede |
MEB açıkladı |
(left alone — it is read as a word) |
Dr. Ayşe |
doktor Ayşe |
Ali & Veli |
Ali ve Veli |
Suffixes are built from the reading, not from the way they were typed. The
user may have written 5'de; the correct form is beşte, because what decides
the correct form is how the number is spoken. Vowel harmony, consonant
devoicing, the dört → dörd softening and the linking consonants are applied
accordingly.
Abbreviations
Whether an abbreviation is read as a word (MEB → "meb") or letter by letter
(TCDD → "te ce de de") was not left to a hand-maintained list — such a list
never ends. The decision follows from phonotactics: the Turkish syllable has
the shape (C)V(C)(C), no consonant cluster may open a word, and only certain
two-consonant clusters may close one.
from anka.text import is_pronounceable
is_pronounceable("MEB") # True -> "meb"
is_pronounceable("TCDD") # False -> "te ce de de" (no vowel)
is_pronounceable("ABD") # False -> "a be de" (bd is an invalid final cluster)
is_pronounceable("TÜRKSAT") # True -> "TÜRKSAT" (rk|s is valid)
The rule was tested on 58 real institution abbreviations and classified every one
of them correctly. The few cases where it is wrong (AB → "a be", RTÜK →
"ertük") live in two small exception sets in lexicon.py. You can add your own
entries without touching the lexicon:
normalize("XYZKUR bildirdi.", extra_acronyms={"XYZKUR"})
normalize("Xyz. geldi", extra_abbreviations={"Xyz.": "iksiz"})
Circumflex letters and the vocab
The circumflexed â î û are preserved — kâr and kar are not the same
word, and these letters also drive vowel harmony (kâr → kârda, mahkûm →
mahkûmu). If you are working with a checkpoint whose vocab does not carry
these letters, fold_circumflex=True folds them to their plain counterparts.
If a vocab is given (and AnkaTTS passes its own automatically), missing
characters are handled too:
normalize("MAHKÛM oldu.", vocab=vocab) # 'mahkûm oldu.' (no uppercase Û in the vocab)
normalize("Fiyat 100 ₼", vocab=vocab) # warning: '₼' is not in the vocab at all
A blanket lower() is not applied — the model was trained on normally
written text. Only the word carrying an out-of-vocab character is lowered, and
only if the lowercase form solves the problem. If it does not, the word is left
intact and a warning is raised.
Parameters
There are three layers. Day to day, only the first one is needed.
Everyday: text, voice or ref_audio + ref_text, speed, seed,
sentence_pause (0.15 s), paragraph_pause (0.45 s)
Through presets: cfg_strength, nfe_step, sway_sampling_coef
tts.synthesize("Metin.", voice="anlatici", preset="dogal")
tts.synthesize("Metin.", voice="anlatici", preset="dogal", speed=1.1) # override it
| preset | speed | nfe | measured |
|---|---|---|---|
dogal |
0.85 | 32 | WER 1.58%, CER 0.33% |
kararli (default) |
1.0 | 32 | WER 1.66%, CER 0.38% |
hizli |
1.0 | 16 | ~2× faster, quality not yet measured |
The values are not guesses: the 260k Turkish checkpoint synthesized a
495-sentence evaluation set, which then went through one and the same ASR pass
(faster-whisper large-v3, tr, beam 5). cfg_strength and sway_sampling_coef
are identical across all three presets — no sweep has been run on that axis
yet, and the dogal row will be updated after the cfg sweep.
Undocumented: cross_fade_duration, target_rms, fix_duration,
ode_method are passed through **f5_kwargs but do not appear in the signature.
The urge to fiddle with cross_fade_duration in particular is usually an attempt
to "fix" the chunking pause; the actual fix is the reference duration, and the
library already says so.
Without an explicit speed, the preset's value is used (speed=None = "whatever
the preset says").
Voices
tts.add_voice("anlatici", "kayit.wav", "Referansta birebir söylenen cümle.")
tts.list_voices() # ['anlatici']
tts.synthesize("Metin.", voice="anlatici")
The audio file is copied under ~/.cache/anka/voices/ — the voice keeps
working even if you move the original file. The directory can be changed with
ANKA_HOME. If exactly one voice is registered, it is used without writing
voice= at all.
There are no ready-made voices in the package: since distributing voices requires consent, the repo is empty for now.
Output contract
synthesize() always returns:
numpy.float32- 24 000 Hz
- mono (a one-dimensional array)
- values within
[-1, 1]
save_wav() writes in the same format. The same text + the same seed + the same
reference = the same output (random if seed=None is given).
Out of scope
This package contains no model architecture, no training code, no ODE solver and
no vocoder — all of that lives in f5-tts and is used as is. No fork, no copied
inference code. The single exception is chunk_text, and that only for
dependency layering: so that the chunk count can be estimated without pulling in
torch (its source and license are stated in the file).
There is no automatic transcription (ASR) either: if you pass ref_audio, you
must pass ref_text too. A Whisper dependency would triple the install size.
Development
pip install -e ".[dev]"
pytest # 783 tests, no model required
The determinism test needs a model; to run it:
export ANKA_TEST_MODEL=/yol/model_dizini
export ANKA_TEST_REF=referans.wav
export ANKA_TEST_REF_TEXT="Referansta söylenen cümle."
pytest tests/test_tts.py -k determinizm
License
Code: Apache-2.0.
Dependencies and weights are subject to their own licenses — f5-tts is MIT, but
the license of the model weight you use is separate and may not permit
commercial use (for example marduk-ra/F5-TTS-Turkish is CC-BY-NC-4.0). Getting
permission from the owner of the cloned voice is your responsibility as well.
anka-tts (Türkçe)
English · Türkçe
Zero-shot Türkçe ses klonlama için F5-TTS (MIT)
üzerine ince bir sarmalayıcı. Yeni bir inference runtime değil: model mimarisi,
ODE çözücü ve vocoder olduğu gibi f5-tts'ten gelir. Paketin kattığı üç şey var —
Türkçe metin normalizasyonu (sayılar, ekler, kısaltmalar, tarih/saat/para),
referans ses doğrulama (F5'in belgelenmemiş beklentilerini sentezden önce
söyler) ve ölçülmüş varsayılanlar.
Kurulum
İki ayrı kurulum var; hangisine ihtiyacın olduğu ne yapacağına bağlı.
pip install anka-tts # metin normalizasyonu + referans doğrulama
pip install anka-tts[tts] # yukarıdakiler + sentez
anka-tts |
anka-tts[tts] |
|
|---|---|---|
| bağımlılıklar | numpy, soundfile | + torch, torchaudio, f5-tts, huggingface-hub |
| indirme boyutu | birkaç MB | ~2 GB |
anka.text |
✅ | ✅ |
anka.audio |
✅ | ✅ |
AnkaTTS.synthesize() |
❌ | ✅ |
Metni normalize etmek isteyen birinin 2 GB PyTorch indirmesi gerekmiyor:
anka.text ve anka.audio saf standart kütüphane + soundfile ile çalışır.
Torch yalnızca AnkaTTS.from_pretrained() çağrıldığında yüklenir.
Beş satırda
from anka import AnkaTTS
tts = AnkaTTS.from_pretrained("kullanici/anka-tts-tr", device="cuda")
tts.add_voice("anlatici", "referans.wav", "Bugün İstanbul'da hafif bir rüzgar var.")
wav = tts.synthesize("2026'da toplantı 14:30'da mı?", voice="anlatici")
tts.save_wav(wav, "cikti.wav")
Sentez olmadan, yalnızca normalizasyon:
from anka.text import normalize
normalize("15.03.2026'da saat 14:30'da %20 indirim başlıyor.")
# 'on beş mart iki bin yirmi altıda saat on dört otuzda yüzde yirmi indirim başlıyor.'
Referans ses kılavuzu
Klonlama kalitesini belirleyen tek en önemli şey referans kaydıdır — modelden ya da parametrelerden daha çok. Kurallar:
| iyi | kötü | |
|---|---|---|
| süre | 7-9 sn | < 5 sn (model sesi tanıyamaz), > 12 sn (F5 kırpar, boşa gider) |
| içerik | tamamlanmış cümle, düşen tonlamayla biten | cümle ortasından kesilmiş parça |
| transkript | noktalamalı, birebir söylenen | noktalamasız, yaklaşık, eksik |
| kayıt | tek konuşmacı, sessiz ortam, 24 kHz+ | gürültü, müzik, yankı, 16 kHz |
| seviye | normal kayıt seviyesi | sonradan normalize edilmiş (taban gürültüsü yükselir) |
# İyi
ref_text = "Bugün İstanbul'da hafif bir rüzgar var. Denizin kokusu sokaklara yayılıyor."
# Kötü -- cümle parçası, noktalama yok, sesle birebir eşleşmiyor
ref_text = "bugün istanbulda hafif bir rüzgar"
ref_text sesle birebir eşleşmeli. F5 üretilecek sesin süresini
ref_süre / ref_metin_uzunluğu oranından tahmin ediyor; transkript uzun ya da
kısa olursa bu oran bozulur ve model hece yutar.
Sentezden önce kontrol et:
from anka.audio import check_reference
check = check_reference("referans.wav", ref_text, gen_text="Üretilecek metin.")
print(check.duration, check.estimated_chunks)
print(check.ok) # yalnızca hiç uyarı yoksa True
print(check.fatal) # referans hiç kullanılamıyorsa True (ref_text yoksa)
for w in check.warnings:
print("!", w)
ok katı olan cevap: tek bir uyarı bile varsa False olur -- 16 kHz'lik bir
kayıt ya da 3 saniyelik bir referans dahil. fatal ise dar olan: referans hiç
kullanılamıyor demektir, bugün yalnızca ref_text eksikse. Temiz bir referans
istiyorsan ok'a, sentezin yürüyüp yürüyemeyeceğini soruyorsan fatal'a bak.
synthesize() bunu zaten otomatik çağırır ve uyarıları warnings olarak verir;
validate_reference=False ile kapatılabilir.
Neden 7-9 saniye
F5 üretim penceresini şöyle hesaplıyor (utils_infer.py):
max_chars = (ref_bayt / ref_süre) × (22 − ref_süre) × speed
= konuşma_hızı × (22 − ref_süre) × speed
22 saniyelik bir bütçe var: referans ne kadar yer kaplarsa üretime o kadar az kalıyor. Yani uzun referans metni daha çok parçaya böler. Kısa referansın sorunu ise farklı — parçalanma değil, modelin ses karakteristiğini yeterince alamaması. 7-9 sn ikisinin ortasındaki bant.
Not: formül karakter değil UTF-8 bayt sayıyor. ğ ü ş ı ö ç iki bayt
olduğu için Türkçe metin aynı uzunluktaki İngilizce metinden daha erken bölünür.
Uzun metinler
F5 bütçe dolduğunda metni [;:,.!?] işaretlerinin herhangi birinden böler —
yani virgülden, cümlenin ortasından. 12 sn'lik bir referansta şöyle oluyor:
2. Bavulları akşamdan hazırlarsak,
3. sabah yola çıkarken hiç acele etmeyiz.
Orada duraksama olması kulakta yanlış duyulur, ve bu tamamen referansın süresine bağlı olduğu için metni yazan kişinin kontrolünde değil.
anka bu yüzden bölmeyi F5'e bırakmaz: metni önce cümlelere ayırır, bütçeye
sığan cümle gruplarını ayrı ayrı sentezler ve aralarına kısa bir duraklama
koyup birleştirir. Ek yeri her zaman cümle sonuna denk gelir — kulağın
zaten duraklama beklediği yere. Referans 7 sn de olsa 12 sn de olsa cümleler
bütün kalır.
Boş satırla ayrılmış paragraflar korunur ve paragraf geçişine daha uzun bir duraklama konur. Cümleler paragraf sınırını aşarak gruplanmaz.
tts.synthesize(uzun_metin, voice="anlatici") # varsayılan: 0.15 sn / 0.45 sn
tts.synthesize(uzun_metin, voice="anlatici", sentence_pause=0.2, paragraph_pause=0.6)
tts.synthesize(uzun_metin, voice="anlatici", chunk_by_sentence=False) # F5'in kendi bölmesi
Tek başına bütçeyi aşan bir cümle bölünmeden F5'e bırakılır; o cümlede F5 kendi kararını verir ama diğerleri etkilenmez.
Aynı sesle hem kısa hem uzun metin üretmek için ayar değiştirmen gerekmez. Tek cümlelik bir metin tek çağrıya gider, hiç ek yeri oluşmaz; iki paragraflık bir metin cümlelere ayrılır. Referans süresi yalnızca kaç parça olacağını etkiler, nereden bölüneceğini değil.
Konsol çıktısı
F5 ürettiği her parça için kendi bloğunu basar — referans transkripti, her
gen_text, boş satırlar ve bir ilerleme çubuğu. anka F5'i her cümle grubu için
bir kez çağırdığından bu blok her segmentte tekrar eder. anka bu çıktıyı yakalar
ve yerine kendi görünümünü çizer:
anka model anka_tts/v0.1 · mps · 24 kHz
anka voice anlatici · 7.8s reference
anka warn the text will be split into 3 chunks, and every seam between them
is a potentially audible pause. AnkaTTS aligns those seams with
sentence ends, where a pause is expected anyway…
✓ 1/3 Dr. Ayşe Yılmaz'ın 15.03.2026 tarihli raporuna göre %12 arttı. 2.6s
✓ 2/3 Sabah 09:05'te başlayan toplantı 14:30'da sona erdi. 2.6s
▸ 3/3 ████████████░░░░░░ 00:02 Katılımcılara 250 ml su dağıtıldı.
Segmentler senin yazdığın hâliyle adlandırılır, modele giden normalize
metinle değil: doktor Ayşe Yılmazın on beş mart iki bin yirmi altı değil,
Dr. Ayşe Yılmaz'ın 15.03.2026. İkisi cümle cümle hizalanır; hizalama
kurulamazsa yanlış tahmin yerine normalize metin gösterilir. Bütçeye sığmayan
bir cümle parça parça üretilir ve ilkinden sonraki parçalar … ile işaretlenir.
Uyarılar da aynı çiziciden geçer. Python'a bırakılırsa doğrudan stderr'e
yazılır — canlı satırın ortasına düşüp satırı bozar — ve dosya adı, satır
numarası, kaynak satırıyla birlikte basılır; hiçbiri sesi dinleyen kişinin
işine yaramaz. Yalnızca mesaj gösterilir; başka bir kütüphanenin uyarısı
kaynağını etiket olarak korur (torch: …), böylece kimin şikâyet ettiği
belli kalır.
Ne çizileceğine progress= karar verir. F5'in kendi çıktısı her hâlükârda
yakalanır.
| değer | davranış |
|---|---|
"auto" (varsayılan) |
etkileşimli terminalde yukarıdaki canlı görünüm, başka her yerde sessizlik |
False / None |
sessizlik |
"plain" |
segment başına tek satır, kaçış dizisi yok — log ve CI için |
"pretty" / True |
yönlendirilmiş çıktıda bile canlı görünümü zorla |
| bir callable | her segmentten sonra callback(done, total, segment_text) |
tts.synthesize(metin, voice="anlatici") # auto
tts.synthesize(metin, voice="anlatici", progress=False) # sessiz
tts.synthesize(metin, voice="anlatici", progress="plain") # segment başına tek satır
İki tür üçüncü taraf gürültüsü elenir; çünkü ikisi de sana değil, uyarıyı üreten kütüphanenin yazarına hitap ediyor:
- torch'un
An output with one or more elements was resized…uyarısı — vocos ve F5'in yaptığı herstft/istftçağrısında tetikleniyor; - hub arka ucunun
You are sending unauthenticated requests to the HF Hubmesajı — F5 vocos vocoder'ını hub'dan çektiği için her yüklemede basılıyor. Kaynağında susturmak içinHF_TOKEN, vocoder önbellekteyse sorguyu tamamen atlamak içinHF_HUB_OFFLINE=1.
İndirme çubuklarına dokunulmaz: sessiz bir 2 GB indirme gürültüden kötüdür.
Gizlenen her şeyi — F5'in çıktısı, susturulan uyarılar ve hub mesajı — stderr'de
görmek için ANKA_DEBUG=1. NO_COLOR dikkate alınır, dumb terminalde düz
satırlara düşülür.
Normalizasyon
normalize() metni modele verilebilir hâle getirir. Noktalama korunur —
özellikle ?, çünkü model soru tonlamasını ondan çıkarıyor.
| girdi | çıktı |
|---|---|
1994'te |
bin dokuz yüz doksan dörtte |
5'de (yanlış yazım) |
beşte |
4'ü |
dördü |
6'nın |
altının |
90'lar |
doksanlar |
3. sınıf |
üçüncü sınıf |
Sıralamada 3. |
(dokunulmaz — cümle sonu olabilir) |
15.03.2026 |
on beş mart iki bin yirmi altı |
14:30 / 14:00 / 09:05 |
on dört otuz / on dört / dokuz sıfır beş |
1.250 TL |
bin iki yüz elli lira |
3,5 TL |
üç lira elli kuruş |
2,75 $ |
iki dolar yetmiş beş sent |
2,5 milyon TL |
iki virgül beş milyon lira |
%20 / %50'si |
yüzde yirmi / yüzde ellisi |
5 km'de |
beş kilometrede |
36,6 °C |
otuz altı virgül altı derece |
II. Abdülhamid |
ikinci Abdülhamid |
TBMM'de |
te be me mede |
MEB açıkladı |
(dokunulmaz — sözcük gibi okunur) |
Dr. Ayşe |
doktor Ayşe |
Ali & Veli |
Ali ve Veli |
Ekler yazılıştan değil okunuştan üretilir. Kullanıcı 5'de yazmış olabilir;
doğrusu beştedir, çünkü doğru biçimi belirleyen sayının söylenişidir. Ünlü
uyumu, ünsüz sertleşmesi, dört → dörd yumuşaması ve kaynaştırma harfleri
buna göre uygulanır.
Kısaltmalar
Bir kısaltmanın sözcük gibi mi (MEB → "meb") yoksa harf harf mi (TCDD →
"te ce de de") okunacağı elle tutulan bir listeye bırakılmadı — o liste hiç
bitmez. Karar fonotaktikten çıkıyor: Türkçe hecesi (C)V(C)(C) kalıbındadır,
sözcük başında ünsüz kümesi bulunmaz, sonda yalnızca belirli ikili kümeler gelir.
from anka.text import is_pronounceable
is_pronounceable("MEB") # True -> "meb"
is_pronounceable("TCDD") # False -> "te ce de de" (ünlü yok)
is_pronounceable("ABD") # False -> "a be de" (bd geçersiz son küme)
is_pronounceable("TÜRKSAT") # True -> "TÜRKSAT" (rk|s geçerli)
Kural 58 gerçek kurum kısaltmasında sınandı ve hepsini doğru sınıflandırdı.
Yanıldığı birkaç durum (AB → "a be", RTÜK → "ertük") lexicon.py'deki iki
küçük istisna kümesinde. Kendi ekini lexicon'a dokunmadan verebilirsin:
normalize("XYZKUR bildirdi.", extra_acronyms={"XYZKUR"})
normalize("Xyz. geldi", extra_abbreviations={"Xyz.": "iksiz"})
Şapkalı harfler ve vocab
Düzeltme işaretli â î û korunur — kâr ile kar aynı sözcük değil, ve
ünlü uyumunu da bu harfler belirler (kâr → kârda, mahkûm → mahkûmu).
Vocab'ı bu harfleri taşımayan bir kontrol noktasıyla çalışıyorsan
fold_circumflex=True düz karşılıklarına indirir.
Vocab verilirse (ki AnkaTTS kendi vocab'ını otomatik verir) eksik karakterler
de ele alınır:
normalize("MAHKÛM oldu.", vocab=vocab) # 'mahkûm oldu.' (büyük Û vocab'da yok)
normalize("Fiyat 100 ₼", vocab=vocab) # uyarı: '₼' vocab'da hiç yok
Genel bir lower() yapılmaz — model normal yazılmış metinle eğitildi.
Yalnızca vocab dışı karakter taşıyan kelime küçültülür, o da ancak küçük hâli
sorunu çözüyorsa. Çözmüyorsa kelime bozulmaz, uyarı verilir.
Parametreler
Üç katman var. Günlük kullanımda yalnızca ilki gerekir.
Günlük: text, voice ya da ref_audio + ref_text, speed, seed,
sentence_pause (0.15 sn), paragraph_pause (0.45 sn)
Preset ile: cfg_strength, nfe_step, sway_sampling_coef
tts.synthesize("Metin.", voice="anlatici", preset="dogal")
tts.synthesize("Metin.", voice="anlatici", preset="dogal", speed=1.1) # üstüne yaz
| preset | speed | nfe | ölçülen |
|---|---|---|---|
dogal |
0.85 | 32 | WER %1.58, CER %0.33 |
kararli (varsayılan) |
1.0 | 32 | WER %1.66, CER %0.38 |
hizli |
1.0 | 16 | ~2× hızlı, kalite ölçülmedi |
Değerler tahmin değil: 260k Türkçe kontrol noktasıyla 495 cümlelik değerlendirme
seti sentezlenip aynı ASR geçişinden (faster-whisper large-v3, tr, beam 5)
geçirildi. cfg_strength ve sway_sampling_coef üç preset'te de aynı —
o eksende henüz tarama yapılmadı, dogal satırı cfg taraması sonrası
güncellenecek.
Gizli: cross_fade_duration, target_rms, fix_duration, ode_method
**f5_kwargs ile geçer ama imzada görünmez. Özellikle cross_fade_duration'ı
kurcalama isteği genellikle chunking duraksamasını "düzeltme" çabasıdır; asıl
çözüm referans süresidir ve kütüphane bunu zaten söylüyor.
speed verilmezse preset'in değeri kullanılır (speed=None = "preset ne diyorsa").
Sesler
tts.add_voice("anlatici", "kayit.wav", "Referansta birebir söylenen cümle.")
tts.list_voices() # ['anlatici']
tts.synthesize("Metin.", voice="anlatici")
Ses dosyası ~/.cache/anka/voices/ altına kopyalanır — kaynak dosyayı
taşısan da ses çalışmaya devam eder. Dizin ANKA_HOME ile değiştirilebilir.
Tek bir ses kayıtlıysa voice= yazmadan da o kullanılır.
Pakette hazır ses yok: dağıtılacak sesler rıza gerektirdiği için depo şimdilik boş.
Çıktı sözleşmesi
synthesize() her zaman şunu döner:
numpy.float32- 24 000 Hz
- mono (tek boyutlu dizi)
- değerler
[-1, 1]aralığında
save_wav() aynı biçimde yazar. Aynı metin + aynı seed + aynı referans = aynı
çıktı (seed=None verilirse rastgele).
Kapsam dışı
Bu paket model mimarisi, eğitim kodu, ODE çözücü ya da vocoder içermez — hepsi
f5-tts içinde ve olduğu gibi kullanılıyor. Fork yok, kopyalanmış inference
kodu yok. Tek istisna chunk_text, o da bağımlılık katmanlaması için: torch
çekmeden parça sayısı tahmin edilebilsin diye (kaynağı ve lisansı dosyada
belirtildi).
Otomatik transkripsiyon (ASR) da yok: ref_audio verirsen ref_text de
vermelisin. Whisper bağımlılığı kurulum boyutunu üçe katlardı.
Geliştirme
pip install -e ".[dev]"
pytest # 783 test, model gerektirmez
Determinizm testi model ister; çalıştırmak için:
export ANKA_TEST_MODEL=/yol/model_dizini
export ANKA_TEST_REF=referans.wav
export ANKA_TEST_REF_TEXT="Referansta söylenen cümle."
pytest tests/test_tts.py -k determinizm
Lisans
Kod: Apache-2.0.
Bağımlılıklar ve ağırlıklar kendi lisanslarına tabidir — f5-tts MIT'tir, ama
kullandığın model ağırlığının lisansı ayrıdır ve ticari kullanıma izin
vermeyebilir (örneğin marduk-ra/F5-TTS-Turkish CC-BY-NC-4.0'dır). Klonlanan
sesin sahibinden izin almak da senin sorumluluğundadır.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file anka_tts-0.1.0.tar.gz.
File metadata
- Download URL: anka_tts-0.1.0.tar.gz
- Upload date:
- Size: 103.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3875164e93a095d395ed4cafd0d0c84bc09d877daaa65f09a2b4704f55860e81
|
|
| MD5 |
b8d7bb7e2ab698874ec07b9718a66cc3
|
|
| BLAKE2b-256 |
b05e412e9135876c113b83e3f4b38f90bbf813766f1e6471805b80a86410b354
|
Provenance
The following attestation bundles were made for anka_tts-0.1.0.tar.gz:
Publisher:
publish.yml on KrmKayabasi/Anka-TTS
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
anka_tts-0.1.0.tar.gz -
Subject digest:
3875164e93a095d395ed4cafd0d0c84bc09d877daaa65f09a2b4704f55860e81 - Sigstore transparency entry: 2796541642
- Sigstore integration time:
-
Permalink:
KrmKayabasi/Anka-TTS@b70e2e837c782fb17461e2659595b7b994b080ad -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/KrmKayabasi
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b70e2e837c782fb17461e2659595b7b994b080ad -
Trigger Event:
push
-
Statement type:
File details
Details for the file anka_tts-0.1.0-py3-none-any.whl.
File metadata
- Download URL: anka_tts-0.1.0-py3-none-any.whl
- Upload date:
- Size: 68.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
affb466cd905db9b7f3280ed42da35ef2b63bfd389e76bfd89a57145b08146d3
|
|
| MD5 |
703a7672a368954327e06ce6e2539089
|
|
| BLAKE2b-256 |
4640eaebf6dda68f1d6331a48b166077df7c2946663152201daf15dbab7ef42e
|
Provenance
The following attestation bundles were made for anka_tts-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on KrmKayabasi/Anka-TTS
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
anka_tts-0.1.0-py3-none-any.whl -
Subject digest:
affb466cd905db9b7f3280ed42da35ef2b63bfd389e76bfd89a57145b08146d3 - Sigstore transparency entry: 2796541683
- Sigstore integration time:
-
Permalink:
KrmKayabasi/Anka-TTS@b70e2e837c782fb17461e2659595b7b994b080ad -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/KrmKayabasi
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b70e2e837c782fb17461e2659595b7b994b080ad -
Trigger Event:
push
-
Statement type: