Update README.md
Browse files
README.md
CHANGED
|
@@ -1,10 +1,71 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
-
title: README
|
| 3 |
-
emoji: 👁
|
| 4 |
-
colorFrom: red
|
| 5 |
-
colorTo: indigo
|
| 6 |
-
sdk: static
|
| 7 |
-
pinned: false
|
| 8 |
-
---
|
| 9 |
|
| 10 |
-
|
|
|
|
|
|
|
|
|
| 1 |
+
# 🇱🇰 Trilingual AI for Sri Lanka
|
| 2 |
+
|
| 3 |
+
**Building open and accessible Speech AI for Sinhala, Tamil, and English.**
|
| 4 |
+
|
| 5 |
+
We are developing AI technologies designed specifically for the linguistic and cultural context of **Sri Lanka**, with a focus on making high-quality speech and language technologies available for everyone.
|
| 6 |
+
|
| 7 |
+
## 🎙️ Our Focus
|
| 8 |
+
|
| 9 |
+
Our primary focus is **Speech AI**, covering the complete speech technology stack:
|
| 10 |
+
|
| 11 |
+
- 🗣️ **Automatic Speech Recognition (ASR / STT)** — Speech-to-text for Sinhala, Tamil, and English
|
| 12 |
+
- 🔊 **Text-to-Speech (TTS)** — Natural and expressive speech synthesis
|
| 13 |
+
- 🌐 **Multilingual & Code-Switching Speech** — Handling Sinhala, Tamil, and English in real-world conversations
|
| 14 |
+
- 🧠 **Speech & Language Models** — Models that understand Sri Lankan linguistic patterns
|
| 15 |
+
- 📚 **Datasets** — High-quality datasets for training and evaluating Sri Lankan AI models
|
| 16 |
+
- 🔬 **Research & Evaluation** — Benchmarks and evaluation resources for Sri Lankan languages
|
| 17 |
+
|
| 18 |
+
## 🌏 Languages
|
| 19 |
+
|
| 20 |
+
| Language | Focus |
|
| 21 |
+
|---|---|
|
| 22 |
+
| 🇱🇰 Sinhala | Speech recognition, synthesis, datasets & language AI |
|
| 23 |
+
| 🇱🇰 Tamil | Speech recognition, synthesis, datasets & language AI |
|
| 24 |
+
| 🇬🇧 English | Multilingual and code-switching speech AI |
|
| 25 |
+
|
| 26 |
+
We are particularly interested in **real-world Sri Lankan speech**, including different accents, dialects, speaking styles, environments, and Sinhala–English / Tamil–English code-switching.
|
| 27 |
+
|
| 28 |
+
## 🚀 What We Build
|
| 29 |
+
|
| 30 |
+
Our Hugging Face organization will host:
|
| 31 |
+
|
| 32 |
+
### Models
|
| 33 |
+
- Speech-to-Text models
|
| 34 |
+
- Text-to-Speech models
|
| 35 |
+
- Speech representation models
|
| 36 |
+
- Multilingual speech models
|
| 37 |
+
- Fine-tuned and domain-specific models
|
| 38 |
+
|
| 39 |
+
### Datasets
|
| 40 |
+
- Sinhala speech datasets
|
| 41 |
+
- Tamil speech datasets
|
| 42 |
+
- English speech datasets from Sri Lankan speakers
|
| 43 |
+
- Multilingual and code-switched speech datasets
|
| 44 |
+
- Evaluation and benchmark datasets
|
| 45 |
+
|
| 46 |
+
### Tools & Resources
|
| 47 |
+
- Data processing pipelines
|
| 48 |
+
- Training recipes
|
| 49 |
+
- Evaluation tools
|
| 50 |
+
- Model demos
|
| 51 |
+
- Research resources
|
| 52 |
+
|
| 53 |
+
## 🎯 Our Vision
|
| 54 |
+
|
| 55 |
+
> **AI that understands Sri Lanka's languages, voices, and people.**
|
| 56 |
+
|
| 57 |
+
Our goal is to help build an ecosystem where Sinhala and Tamil are first-class languages in modern AI systems, while supporting English and multilingual communication across Sri Lanka.
|
| 58 |
+
|
| 59 |
+
## 📌 Hugging Face
|
| 60 |
+
|
| 61 |
+
This organization serves as a central hub for our:
|
| 62 |
+
|
| 63 |
+
**Models · Datasets · Spaces · Research · Tools**
|
| 64 |
+
|
| 65 |
+
Explore our repositories and experiment with our models and datasets.
|
| 66 |
+
|
| 67 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
+
### 🇱🇰 Sinhala · தமிழ் Tamil · English
|
| 70 |
+
|
| 71 |
+
**Building Speech AI for Sri Lanka.**
|