The voice data your models have never heard.
Proprietary voice corpora for low-resource languages: native audio, transcription, 14 annotation layers, commercial rights included. Produced in studio with native speakers, available via API.
| audio_file | fon_04217.wav |
| language_iso | fon |
| speaker_id | BEN-FON-083 |
| region | Littoral, Benin |
| speaker_gender | F |
| speaker_age | 34 |
| register | Folk tale |
| transcription | Mi ɖo gbɛ na mì... |
| translation_en | We have something to tell you... |
Eight low-resource languages, produced continuously.
Languages nearly absent from the web, and therefore from your training data. We produce them in studio, with native speakers under contract.
Every hour is annotated across 14 layers. One corpus, many voice systems.
ASR, TTS, speech translation, voice agents, voice servers: the same corpus powers the training, fine-tuning and evaluation of your models and machines.
Gold standard, measured quality
This data cannot be scraped. We produce it, end to end.
Local studios and mobile units, native narrators under contract, trained transcribers: we control every link, from speaker to delivered dataset.
Our field presence: Labari.io, our free listening app, streams stories and folk tales in local languages and sustains our narrator network. labari.io →
Three ways to access the data.
From a one-off test to language exclusivity: same quality standard, same rights, in every format.
Catalog license
Annual subscription access to one or more languages: search, export and delivery of structured datasets, via API.
Custom dataset
Dedicated production for your needs: language, domain, register, speech type, volume. You specify, we produce in studio.
Temporary exclusivity
An exclusivity window on a language, domain or corpus, before it joins the shared catalog.
Test the data on your models.
Describe your use case (ASR, TTS, translation, agents) and the languages you need: we will prepare a sample.