Voice infrastructure · low-resource languages

The voice data your models have never heard.

Proprietary voice corpora for low-resource languages: native audio, transcription, 14 annotation layers, commercial rights included. Produced in studio with native speakers, available via API.

Browse the catalog
Fon
audio_filefon_04217.wav
language_isofon
speaker_idBEN-FON-083
regionLittoral, Benin
speaker_genderF
speaker_age34
registerFolk tale
transcriptionMi ɖo gbɛ na mì...
translation_enWe have something to tell you...
GDPR consent · commercial AI license
8 languages in the target catalog~200M speakers covered14 layers of annotationGold standard multi-annotatorRights secured at the source
Catalog

Eight low-resource languages, produced continuously.

Languages nearly absent from the web, and therefore from your training data. We produce them in studio, with native speakers under contract.

Hub Cotonou · Benin
Fon
~4M
Benin
Ewe
~10M
Togo, Benin
Yoruba
~50M
Nigeria, Benin
Hausa
~94M
Nigeria, Niger, Benin
Hub Dakar · Senegal
Wolof
~18M
Senegal
Pulaar
~7M
Senegal, Sahel
Bambara
~16M
Mali, Senegal
Maninka
~6M
Eastern Senegal, Guinea, Mali
Annotation

Every hour is annotated across 14 layers. One corpus, many voice systems.

ASR, TTS, speech translation, voice agents, voice servers: the same corpus powers the training, fine-tuning and evaluation of your models and machines.

Text
VerbatimNormalizedITN
Language & pronunciation
Code-switchingTones & geminationG2P
Time & speakers
AlignmentDiarization
Meaning
TranslationSemanticsProsody & emotion
Context & usage
Acoustic eventsIndexingDialogue acts

Gold standard, measured quality

Multi-expert annotation3 independent native speakers, blind, arbitration by a senior linguist
Reliability measured and publishedKrippendorff α, WER, CER, delivered in the datasheet
Strict linguistic complianceofficial orthography, tones and gemination respected
Legal compliancecommercial and AI consent traced per speaker
Linguistic signaturetonal error rate, which no competitor publishes
ASRTTSSpeech translationVoice agentsVoice serversAudio searchEvaluation
Production

This data cannot be scraped. We produce it, end to end.

Local studios and mobile units, native narrators under contract, trained transcribers: we control every link, from speaker to delivered dataset.

Narratorcontract, GDPR consent
Recordingstudios, mobile units
Annotation14 layers, linguist arbitration
Quality + rightsGold validation, IP assignment
Catalogdelivery via API or export
Client's modeltraining, fine-tuning, evaluation

Our field presence: Labari.io, our free listening app, streams stories and folk tales in local languages and sustains our narrator network. labari.io →

Licensing

Three ways to access the data.

From a one-off test to language exclusivity: same quality standard, same rights, in every format.

Catalog license

Annual subscription access to one or more languages: search, export and delivery of structured datasets, via API.

Custom dataset

Dedicated production for your needs: language, domain, register, speech type, volume. You specify, we produce in studio.

Temporary exclusivity

An exclusivity window on a language, domain or corpus, before it joins the shared catalog.

Data ready for production
GDPR consenttraced per recording
Intellectual propertycomplete chain of rights, from speaker to dataset
Clear commercial licenseno ambiguity for AI use
Contact

Test the data on your models.

Describe your use case (ASR, TTS, translation, agents) and the languages you need: we will prepare a sample.