SomAI. Contact
Somali Voice Access Initiative

Voice technology
should speak
Somali.

SomAI is a Somali-led initiative building rights-cleared speech data, documentation, and baseline tools for accessibility, education, and public-interest use.

Pilot target
25–35 hrs
clean, transcribed Somali speech
First use case
Somali TTS
dataset ready for text-to-speech
Release stance
Open*
*where consent & licensing allow

01 / Why

Somali is spoken by millions across the Horn of Africa and the diaspora. Voice technology still doesn't speak it well.

That gap falls hardest on people who rely on audio: blind and visually impaired users, low-literacy users, children, elders, students, and mobile-first communities.

Accessibility

Somali screen readers and audio interfaces need natural speech output to exist at all.

Education

Audio learning tools support students, children, and anyone who learns better by listening.

Public information

Health, civic, and community messages reach further as Somali audio.

02 / Current project

The Somali Voice Access Pilot

Not a finished product — a serious foundation. Clean audio, accurate transcripts, consent records, metadata, an evaluation set, and a small baseline TTS demo.

  1. 01

    Scripts & consent

    Somali recording scripts, consent forms, contributor agreements, and documentation workflows.

  2. 02

    Record & review

    Native Somali speakers recorded in controlled conditions; unclear or low-quality audio removed.

  3. 03

    Transcribe & structure

    Accurate transcripts, segmented audio, metadata, and TTS-ready formatting.

  4. 04

    Baseline demo

    Small Somali TTS experiments on the pilot dataset, evaluated with community feedback.

03 / Outputs

Practical resources, not empty claims.

A

Speech dataset

25–35 hours of clean recordings with accurate transcripts and segmented audio-text pairs.

B

Documentation

Consent records, metadata, dataset card, licensing notes, limitations, responsible-use guidance.

C

Evaluation set

A Somali TTS test set for pronunciation, numbers, names, clarity, and sentence rhythm.

D

Baseline demo

A simple text-to-speech demo to validate the dataset and surface what to improve next.

E

Community feedback

Input from Somali speakers, educators, developers, and accessibility-focused users.

F

Future roadmap

A plan to grow data volume, improve models, and include more regional variation.

04 / Responsible AI

A voice is personal. We treat it that way.

A person's voice can be recognizable. SomAI works on consent, privacy, documentation, and careful release conditions.

  • Consent first

    Contributors are told exactly how their recordings may be used, released, and trained on.

  • Minimal personal data

    We avoid collecting unnecessary personal information and document only what's needed.

  • Open where allowed

    Public-good resources are shared wherever consent and licensing permit.

  • No harmful use

    No impersonation, fraud, deception, or harmful synthetic voice applications.

05 / Timeline

A realistic nine-month plan.

  1. Months 1–2

    Project setup, consent forms, recording scripts, contributor onboarding, workflow design.

  2. Months 3–5

    Voice recording, audio processing, transcription, segmentation, and quality review.

  3. Months 6–7

    Dataset documentation, metadata, licensing notes, and the TTS evaluation set.

  4. Month 8

    Baseline Somali TTS demo and internal evaluation on the pilot dataset.

  5. Month 9

    Community feedback, final reporting, approved resource release, expansion roadmap.

  6. Next phase

    Expand data volume, improve TTS quality, include regional variation, build practical tools.

06 / Contact

Interested in
Somali voice AI?

caseerprivate@gmail.com

The initiative

SomAI / Somali AI Research Initiative is an emerging Somali-led project focused on practical language resources for inclusive AI. Currently being formalized, led by Abdirahman Aseyr Alasow, Founder / Project Lead.

Collaboration

Open to Somali voice professionals, educators, researchers, accessibility stakeholders, transcription contributors, developers, and responsible AI partners.