gloomcheng commited on
Commit
77fca5d
·
verified ·
1 Parent(s): 991e2f1

docs: update organization profile with IlhaEmbed 37MB and dual-facet clinical architecture

Browse files
Files changed (1) hide show
  1. README.md +53 -63
README.md CHANGED
@@ -7,85 +7,75 @@ sdk: static
7
  pinned: false
8
  ---
9
 
 
 
10
  # WeeMed AI
11
 
12
- **Building a globally standardized, AI-driven digital health and long-term care platform for aging societies.**
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
 
14
- FHIR-first. Integrated health data, intelligent decision support, personalized preventive medicine — deployed in real clinics, health-screening centers, and community long-term care sites in Taiwan.
 
15
 
16
- > **AI accelerates digitization; it isn't the product.**
 
 
17
 
18
  ---
19
 
20
- ## 🇹🇼 Taiwan Sovereign Medical AI Stack
21
-
22
- We ship production-grade models derived from real-world clinical and health checkup deployment problems in Taiwan, publishing open benchmarks and transparent data provenance.
23
-
24
- ### 🧠 [IlhaEmbed (v2.0)](https://huggingface.co/weemed/IlhaEmbed)
25
- **Traditional Chinese Clinical & Medical Terminology Embedding Model (38.5 MB INT8 ONNX / 384-dim)**
26
- - **Domain-Adapted for Taiwan Clinical Writing**: Reads Taiwanese hospital shorthand ( → 低劑量胸部電腦斷層), checkup note acronyms ( → 傷寒篩檢糞便檢體), clinical slang ( → 帶狀皰疹), and Taigi colloquialisms ( → 中風).
27
- - **MODA & NAER 13-Set Retrained**: Fully retrained with the Ministry of Digital Affairs (MODA) and National Academy for Educational Research (NAER) 111,386 medical concept taxonomy.
28
- - **Context-Conditioned Disambiguation**: Dynamically disambiguates pure-Latin acronyms (, , Bun is a fast JavaScript runtime, package manager, bundler, and test runner. (1.3.14+0d9b296af)
29
-
30
- Usage: bun <command> [...flags] [...args]
31
-
32
- Commands:
33
- run ./my-script.ts Execute a file with Bun
34
- lint Run a package.json script
35
- test Run unit tests with Bun
36
- x vite Execute a package binary (CLI), installing if needed (bunx)
37
- repl Start a REPL session with Bun
38
- exec Run a shell script directly with Bun
39
-
40
- install Install dependencies for a package.json (bun i)
41
- add @evan/duckdb Add a dependency to package.json (bun a)
42
- remove redux Remove a dependency from package.json (bun rm)
43
- update @zarfjs/zarf Update outdated dependencies
44
- audit Check installed packages for vulnerabilities
45
- outdated Display latest versions of outdated dependencies
46
- link [<package>] Register or link a local npm package
47
- unlink Unregister a local npm package
48
- publish Publish a package to the npm registry
49
- patch <pkg> Prepare a package for patching
50
- pm <subcommand> Additional package management utilities
51
- info zod Display package metadata from the registry
52
- why tailwindcss Explain why a package is installed
53
-
54
- build ./a.ts ./b.jsx Bundle TypeScript & JavaScript into a single file
55
-
56
- init Start an empty Bun project from a built-in template
57
- create vite Create a new project from a template (bun c)
58
- upgrade Upgrade to latest version of Bun.
59
- feedback ./file1 ./file2 Provide feedback to the Bun team.
60
-
61
- <command> --help Print help text for command.
62
-
63
- Learn more about Bun: https://bun.com/docs
64
- Join our Discord community: https://bun.com/discord, ) and polysemous terms () using operational field context.
65
- - **Pure CPU On-Premise Native**: Zero GPU, zero cloud, zero API keys required. Patient data stays safely inside hospital premises.
66
-
67
- ### 🎙️ [Breeze-ASR-26-edge](https://huggingface.co/weemed/Breeze-ASR-26-edge) & [Taiwanese-Tailo-ASR](https://huggingface.co/weemed/Taiwanese-Tailo-ASR)
68
- **Taiwanese Hokkien (台語) + Mandarin Clinical Speech Recognition Stack**
69
- - Quantized for edge devices in CTranslate2 () and ONNX runtimes.
70
- - Multi-format output support (, , , ).
71
- - Derived from MediaTek Research's Breeze-ASR-26 & OpenAI Whisper-large-v3, fine-tuned on SuíSiann 2.0, TAT_MOE, and real meeting/clinic audio.
72
 
73
  ---
74
 
75
- ## Why this matters to us
76
 
77
- Taiwan is aging fast, and the people doing the caring — nurses at screening centers, care managers at community sites, elders themselves — mostly do not speak to each other in written Mandarin. They speak **Taigi**, in noisy rooms, on tablets held in one hand.
78
 
79
- Health tech that only understands clean written Mandarin does not meet them where they are. So we work on the unglamorous end of clinical AI: the languages actually spoken, the devices actually held, the constraints actually present.
 
 
80
 
81
  ---
82
 
83
- ## How we publish
84
 
85
- - **Honest benchmarks.** We report the numbers that survive scrutiny, including when they challenge our own hypotheses.
86
- - **Documented Provenance.** Open weights and base models are permissively licensed (Apache-2.0). Upstream open government datasets are cataloged transparently.
87
- - **Real audio, real clinics.** Benchmarks measured on real multi-speaker meeting and clinical workflows, not synthetic noise.
88
 
89
  ---
90
 
91
- <sub>Apache-2.0 unless stated otherwise · Made in Yunlin, Taiwan · WeeMed AI</sub>
 
 
 
7
  pinned: false
8
  ---
9
 
10
+ <div align="center">
11
+
12
  # WeeMed AI
13
 
14
+ **Building open-source, edge-native clinical AI and FHIR infrastructure for aging societies.**
15
+
16
+ Deployed in real clinics, mobile screening stations, and community eldercare sites across Taiwan.
17
+
18
+ [Live Space Demo](https://huggingface.co/spaces/weemed/ilhaembed-demo) · [IlhaEmbed Model](https://huggingface.co/weemed/IlhaEmbed) · [GitHub](https://github.com/weemed-ai)
19
+
20
+ ---
21
+
22
+ </div>
23
+
24
+ ## 🇹🇼 Taiwan Sovereign Clinical AI Stack
25
+
26
+ We engineer hyper-compact, edge-native models designed for the realities of frontline healthcare: constrained hardware, strict air-gapped privacy, localized clinical shorthand, and spoken elderly dialects.
27
+
28
+ ### 🧭 [IlhaEmbed: Edge Biomedical Embedding Engine](https://huggingface.co/weemed/IlhaEmbed)
29
+ **37.28 MB INT8 ONNX · 384-dim · 3.97 ms on Vanilla CPU (250+ notes/sec)**
30
+
31
+ - **Dual-Faceted Real-World Clinical Adaptation**:
32
+ - **Spoken Elderly Vernacular (AST / STT)**: Resolves raw spoken complaints transcribed from older adults in Taiwanese (Taigi) into international clinical concepts (e.g., *「阿嬤講伊心臟跳真緊,腳頭烏痛,全身軟巡巡沒力氣」* → `Condition: Generalized Malaise & Palpitations`, *「皮蛇」* → `Herpes Zoster`).
33
+ - **Healthcare Staff Shorthand & NHI Codes**: Decodes ultra-fast nursing notes and screening acronyms (e.g., *「114年成健已做,糞檢 MIF 未交」* → `DiagnosticReport: Fecal Occult Blood / LOINC 14563-1`, *「排 L-CT」* → `ServiceRequest: Low-Dose Chest CT`).
34
+ - **100% Zero-Shot Intent Routing**: Flawlessly categorizes text across 44 standard clinical and administrative anchors (HL7 FHIR `Condition`, `DiagnosticReport`, `Medication`, `Observation`, `ServiceRequest`, etc.).
35
+ - **MTEB Medical Benchmark**: Achieves **0.4498 MAP**, outperforming models several times its size (including BGE-small) while consuming a fraction of the compute.
36
+ - **Fail-Closed Safety Gate**: Explicitly rejects administrative noise and non-clinical requests without hallucinating false diagnosis codes.
37
+ - **Interactive Console**: Experience real-time inference directly in your browser at [weemed/ilhaembed-demo](https://huggingface.co/spaces/weemed/ilhaembed-demo).
38
+
39
+ ```python
40
+ from sentence_transformers import SentenceTransformer
41
+
42
+ model = SentenceTransformer("weemed/IlhaEmbed")
43
 
44
+ # 1. Elderly spoken complaint (AST transcription)
45
+ emb1 = model.encode(["阿嬤講伊心臟跳真緊,全身軟巡巡沒力氣"])
46
 
47
+ # 2. Nursing shorthand & screening notation
48
+ emb2 = model.encode(["114年成健已做,糞檢 MIF 未交,排 L-CT"])
49
+ ```
50
 
51
  ---
52
 
53
+ ### 🎙️ [Breeze-ASR-26-edge](https://huggingface.co/weemed/Breeze-ASR-26-ONNX) & Taiwanese Tailo ASR
54
+ **Bilingual Taiwanese Hokkien (台語) + Mandarin Speech Recognition Stack**
55
+ - Quantized for edge devices in CTranslate2, ONNX, and GGML runtimes.
56
+ - Multi-format phonetic transcription support (Traditional Chinese Hanzi, Taiwanese Romanization / Tâi-lô, and code-switched clinical dialogue).
57
+ - Built upon MediaTek Research's Breeze-ASR-26 & Whisper architectures, fine-tuned on SuíSiann, TAT_MOE, and real-world multi-speaker clinical intake audio.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
 
59
  ---
60
 
61
+ ## Why Edge-Native Clinical AI?
62
 
63
+ Taiwan is aging at one of the fastest rates in the world. The people providing care — community health workers, outreach nurses, and geriatric case managers — interact with seniors who speak **Taigi**, working in noisy community centers with handheld tablets or legacy PC kiosks.
64
 
65
+ 1. **Patient Privacy & Air-Gap Compliance**: Clinical notes and patient complaints cannot be routed through commercial cloud APIs. Everything must run on-premise.
66
+ 2. **Vanilla Hardware Realities**: Ward carts and rural health stations rarely have high-end GPUs. A model that requires 24GB of VRAM cannot help a rural nurse; a 37MB model running at 3.97ms on an older Intel CPU can.
67
+ 3. **Honest Scientific Boundaries**: Vector embeddings excel at semantic concept clustering, but are naturally insensitive to temporal sequencing ("completed" vs. "pending follow-up"). We position our models as **calibrated, fail-closed edge sidecars**, leaving final temporal logic to deterministic business rules and clinical professionals.
68
 
69
  ---
70
 
71
+ ## Open Science & Provenance
72
 
73
+ - **Permissive Open Source**: Model weights, inference scripts, and web demonstrations are published under Apache-2.0.
74
+ - **Transparent Data Lineage**: Grounded in open public health taxonomies (MODA, NAER, LOINC, SNOMED CT, RxNorm) with explicit attribution.
75
+ - **Field-Verified Benchmarks**: Evaluated on genuine clinical notes, ASR transcripts, and real multi-speaker recordings—not synthetic noise.
76
 
77
  ---
78
 
79
+ <div align="center">
80
+ <sub>Apache-2.0 License · Crafted in Yunlin & Chiayi, Taiwan · <b>WeeMed AI</b></sub>
81
+ </div>