--- language: - en - zh language_bcp47: - zh-TW license: apache-2.0 tags: - sign-language - sign-language-recognition - taiwanese-sign-language - tsl - computer-vision - graph-convolutional-network - st-gcn - pytorch pipeline_tag: video-classification --- # SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN
--- ## English Model Card ### 1. Model Summary SignBridge-TSL-STGCN is a lightweight Spatial-Temporal Graph Convolutional Network (ST-GCN) engineered for isolated Taiwanese Sign Language (TSL) word recognition. The model models the spatial configuration of skeletal joints and their temporal evolution across frames, enabling real-time, low-latency sign language recognition suitable for mobile and edge deployments. - **Developer:** SignBridge Project Team (Taiwan) - **Model Architecture:** Spatial-Temporal Graph Convolutional Network (ST-GCN) - **Input Dimensions:** `(Batch_Size, Channels=2, Frames=64, Nodes=55)` - **Keypoint Topology:** 55 Nodes (12 Upper-Body Pose + 21 Left Hand + 21 Right Hand + 1 Virtual Neck Center) - **Primary Weights File:** `best_model.pt` - **License:** Apache License 2.0 - **Repository:** [https://huggingface.co/dYang1/SignBridge-TSL-STGCN](https://huggingface.co/dYang1/SignBridge-TSL-STGCN) ### 2. Compliance and Provenance Declaration In compliance with open-source regulations and digital supply chain integrity standards: 1. **Autonomous Development & Non-Foreign Adversary Compliance:** - The graph network topology, training pipeline, and checkpoint weights were developed independently by the Taiwanese project team. - The model relies entirely on open-source frameworks (PyTorch, OpenCV, MediaPipe). It contains no proprietary models, black-box components, or pre-trained dependencies originating from PRC-based entities or restricted supply chains. 2. **Data Authenticity & Locality:** - Training datasets were collected and curated specifically for Taiwanese Sign Language (TSL) vocabulary, reflecting regional grammatical conventions and practical daily dialogue scenarios. 3. **Open-Source Transparency:** - All source code (network definition `model.py`, graph builder `graph.py`, extractor `extract.py`) and vocabulary label files (`sign_language_words.json`) are fully disclosed under the Apache-2.0 license. ### 3. Skeletal Layout & Preprocessing The graph convolution operates on 55 spatial landmarks extracted via MediaPipe Holistic: - **Upper-Body Pose (Nodes 0–11):** Key joint landmarks including shoulders, elbows, wrists, and hips. - **Left Hand (Nodes 12–32):** Complete 21-joint articulated hand topology. - **Right Hand (Nodes 33–53):** Complete 21-joint articulated hand topology. - **Virtual Root Center (Node 54):** Midpoint between both shoulders, used as the global coordinate anchor for normalization. Input sequences are temporally resampled to 64 frames and spatially normalized by shoulder distance. ### 4. Inference Snippet ```python import json import torch from model import build_from_checkpoint # 1. Load model checkpoint device = torch.device("cuda" if torch.cuda.is_available() else "cpu") checkpoint = torch.load("best_model.pt", map_location=device) model, config = build_from_checkpoint(checkpoint) model.eval() # 2. Load TSL vocabulary with open("sign_language_words.json", "r", encoding="utf-8") as f: vocab = json.load(f) # 3. Simulate input: (Batch=1, Channels=2, Frames=64, Nodes=55) dummy_skeleton = torch.randn(1, 2, 64, 55).to(device) # 4. Perform forward pass with torch.no_grad(): logits = model(dummy_skeleton) probabilities = torch.softmax(logits, dim=-1) predicted_index = torch.argmax(probabilities, dim=-1).item() print(f"Predicted Gloss: {vocab[predicted_index]}") print(f"Confidence: {probabilities[0, predicted_index]:.4f}") ``` ### 5. Citations & References If you utilize the SignBridge system or its pipeline components, please cite: #### Google DeepMind Gemma 4 ```bibtex @misc{gemma20264, title={Gemma 4: Open Multimodal Language Models}, author={Gemma Team}, year={2026}, publisher={Google DeepMind}, howpublished={\url{https://ai.google.dev/gemma/docs/core/model_card_4}} } ``` #### SignBridge-TSL-STGCN ```bibtex @software{SignBridge_TSL_STGCN_2026, author = {SignBridge Development Team}, title = {SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/dYang1/SignBridge-TSL-STGCN}} } ``` ### 6. License This project is licensed under the Apache License 2.0. --- ## Chinese Model Card (繁體中文) ### 1. 模型摘要 (Model Summary) SignBridge-TSL-STGCN 是專為「台灣手語(Taiwanese Sign Language, TSL)」獨立詞辨識研發之輕量化時空圖卷積神經網路(ST-GCN)模型。本模型透過人體骨架關節點之拓撲圖卷積與時間卷積,捕捉手語動作在時空序列中的動態特徵,具備邊緣端與即時環境下低延遲、高穩定度的辨識能力。 - **研發單位**:SignBridge 專案團隊 (Taiwan) - **模型架構**:時空圖卷積神經網路(Spatial-Temporal Graph Convolutional Network, ST-GCN) - **輸入維度**:`(Batch_Size, Channels=2, Frames=64, Nodes=55)` - **關節拓撲**:55 節點(12 上半身姿態 + 21 左手掌指 + 21 右手掌指 + 1 雙肩中點虛擬根節點) - **主要權重檔**:`best_model.pt` - **開源授權**:Apache License 2.0 - **官方儲存庫**:[https://huggingface.co/dYang1/SignBridge-TSL-STGCN](https://huggingface.co/dYang1/SignBridge-TSL-STGCN) ### 2. 開源與非中資合規聲明 (Compliance and Provenance) 為符合數位發展部數位產業署「開源 AI 模型應用組」之競賽規範與資通安全查核,本模型聲明如下: 1. **自主研發與非中資來源**: - 本模型之網路架構設計、特徵管線邏輯、訓練流程及推論權重,均由台灣本地團隊自主研發完成。 - 核心框架均採用開源且具國際公信力之主流專案(PyTorch、OpenCV、MediaPipe),完全無使用任何源自中國大陸或受國際制裁實體之專有模型、私有權重或黑箱相依套件。 2. **訓練資料在地合規**: - 訓練語料皆採集自台灣手語(TSL)詞彙與日常會話情境,遵循台灣在地手語文法,無跨境不合規數據。 3. **商業友善與完整開源**: - 依 Apache-2.0 協議完整開源模型權重、圖架構代碼(`model.py`、`graph.py`)、特徵前處理邏輯(`extract.py`)與標籤對應表(`sign_language_words.json`)。 ### 3. 骨架拓撲與特徵前處理 (Skeletal Layout & Preprocessing) 特徵提取採用 55 個空間座標點(透過 MediaPipe Holistic 擷取): - **上半身姿態 (節點 0–11)**:肩部、手肘、手腕、髖部關鍵節點。 - **左手掌與指關節 (節點 12–32)**:完整 21 個手部關節鏈結。 - **右手掌與指關節 (節點 33–53)**:完整 21 個手部關節鏈結。 - **虛擬頸點 (節點 54)**:取雙肩中點作為全域空間幾何對齊與正規化原點。 所有輸入序列皆在時間軸上等距重採樣至 64 幀,並依據雙肩距離進行空間尺度正規化。 ### 4. 快速推論範例代碼 (Inference Snippet) ```python import json import torch from model import build_from_checkpoint # 1. 載入模型檢查點 device = torch.device("cuda" if torch.cuda.is_available() else "cpu") checkpoint = torch.load("best_model.pt", map_location=device) model, config = build_from_checkpoint(checkpoint) model.eval() # 2. 讀取台灣手語詞彙標籤 with open("sign_language_words.json", "r", encoding="utf-8") as f: vocab = json.load(f) # 3. 模擬輸入骨架特徵: (Batch=1, Channels=2, Frames=64, Nodes=55) dummy_skeleton = torch.randn(1, 2, 64, 55).to(device) # 4. 執行前向傳播 with torch.no_grad(): logits = model(dummy_skeleton) probabilities = torch.softmax(logits, dim=-1) predicted_index = torch.argmax(probabilities, dim=-1).item() print(f"辨識結果 Gloss: {vocab[predicted_index]}") print(f"信心度 Confidence: {probabilities[0, predicted_index]:.4f}") ``` ### 5. 引用來源與學術參考 (Citations & References) 本系統於手語轉自然語言之語序重組模組中採用 Google DeepMind 所開源之 Gemma 多模態模型系列。若於學術研究或專案中使用本專案之演算法與模型,請引用以下標準 BibTeX 條目: #### Google DeepMind Gemma 4 模型引用 ```bibtex @misc{gemma20264, title={Gemma 4: Open Multimodal Language Models}, author={Gemma Team}, year={2026}, publisher={Google DeepMind}, howpublished={\url{https://ai.google.dev/gemma/docs/core/model_card_4}} } ``` #### SignBridge-TSL-STGCN 手語模型引用 ```bibtex @software{SignBridge_TSL_STGCN_2026, author = {SignBridge Development Team}, title = {SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/dYang1/SignBridge-TSL-STGCN}} } ``` ### 6. 開源授權 (License) 本專案依 Apache License 2.0 條款授權釋出。