SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN

English | 繁體中文


English Model Card

1. Model Summary

SignBridge-TSL-STGCN is a lightweight Spatial-Temporal Graph Convolutional Network (ST-GCN) engineered for isolated Taiwanese Sign Language (TSL) word recognition. The model models the spatial configuration of skeletal joints and their temporal evolution across frames, enabling real-time, low-latency sign language recognition suitable for mobile and edge deployments.

  • Developer: SignBridge Project Team (Taiwan)
  • Model Architecture: Spatial-Temporal Graph Convolutional Network (ST-GCN)
  • Input Dimensions: (Batch_Size, Channels=2, Frames=64, Nodes=55)
  • Keypoint Topology: 55 Nodes (12 Upper-Body Pose + 21 Left Hand + 21 Right Hand + 1 Virtual Neck Center)
  • Primary Weights File: best_model.pt
  • License: Apache License 2.0
  • Repository: https://huggingface.co/dYang1/SignBridge-TSL-STGCN

2. Compliance and Provenance Declaration

In compliance with open-source regulations and digital supply chain integrity standards:

  1. Autonomous Development & Non-Foreign Adversary Compliance:
    • The graph network topology, training pipeline, and checkpoint weights were developed independently by the Taiwanese project team.
    • The model relies entirely on open-source frameworks (PyTorch, OpenCV, MediaPipe). It contains no proprietary models, black-box components, or pre-trained dependencies originating from PRC-based entities or restricted supply chains.
  2. Data Authenticity & Locality:
    • Training datasets were collected and curated specifically for Taiwanese Sign Language (TSL) vocabulary, reflecting regional grammatical conventions and practical daily dialogue scenarios.
  3. Open-Source Transparency:
    • All source code (network definition model.py, graph builder graph.py, extractor extract.py) and vocabulary label files (sign_language_words.json) are fully disclosed under the Apache-2.0 license.

3. Skeletal Layout & Preprocessing

The graph convolution operates on 55 spatial landmarks extracted via MediaPipe Holistic:

  • Upper-Body Pose (Nodes 0–11): Key joint landmarks including shoulders, elbows, wrists, and hips.
  • Left Hand (Nodes 12–32): Complete 21-joint articulated hand topology.
  • Right Hand (Nodes 33–53): Complete 21-joint articulated hand topology.
  • Virtual Root Center (Node 54): Midpoint between both shoulders, used as the global coordinate anchor for normalization.

Input sequences are temporally resampled to 64 frames and spatially normalized by shoulder distance.

4. Inference Snippet

import json
import torch
from model import build_from_checkpoint

# 1. Load model checkpoint
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
checkpoint = torch.load("best_model.pt", map_location=device)
model, config = build_from_checkpoint(checkpoint)
model.eval()

# 2. Load TSL vocabulary
with open("sign_language_words.json", "r", encoding="utf-8") as f:
    vocab = json.load(f)

# 3. Simulate input: (Batch=1, Channels=2, Frames=64, Nodes=55)
dummy_skeleton = torch.randn(1, 2, 64, 55).to(device)

# 4. Perform forward pass
with torch.no_grad():
    logits = model(dummy_skeleton)
    probabilities = torch.softmax(logits, dim=-1)
    predicted_index = torch.argmax(probabilities, dim=-1).item()

print(f"Predicted Gloss: {vocab[predicted_index]}")
print(f"Confidence: {probabilities[0, predicted_index]:.4f}")

5. Citations & References

If you utilize the SignBridge system or its pipeline components, please cite:

Google DeepMind Gemma 4

@misc{gemma20264,
  title={Gemma 4: Open Multimodal Language Models},
  author={Gemma Team},
  year={2026},
  publisher={Google DeepMind},
  howpublished={\url{https://ai.google.dev/gemma/docs/core/model_card_4}}
}

SignBridge-TSL-STGCN

@software{SignBridge_TSL_STGCN_2026,
  author = {SignBridge Development Team},
  title = {SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/dYang1/SignBridge-TSL-STGCN}}
}

6. License

This project is licensed under the Apache License 2.0.


Chinese Model Card (繁體中文)

1. 模型摘要 (Model Summary)

SignBridge-TSL-STGCN 是專為「台灣手語(Taiwanese Sign Language, TSL)」獨立詞辨識研發之輕量化時空圖卷積神經網路(ST-GCN)模型。本模型透過人體骨架關節點之拓撲圖卷積與時間卷積,捕捉手語動作在時空序列中的動態特徵,具備邊緣端與即時環境下低延遲、高穩定度的辨識能力。

  • 研發單位:SignBridge 專案團隊 (Taiwan)
  • 模型架構:時空圖卷積神經網路(Spatial-Temporal Graph Convolutional Network, ST-GCN)
  • 輸入維度:(Batch_Size, Channels=2, Frames=64, Nodes=55)
  • 關節拓撲:55 節點(12 上半身姿態 + 21 左手掌指 + 21 右手掌指 + 1 雙肩中點虛擬根節點)
  • 主要權重檔:best_model.pt
  • 開源授權:Apache License 2.0
  • 官方儲存庫:https://huggingface.co/dYang1/SignBridge-TSL-STGCN

2. 開源與非中資合規聲明 (Compliance and Provenance)

為符合數位發展部數位產業署「開源 AI 模型應用組」之競賽規範與資通安全查核,本模型聲明如下:

  1. 自主研發與非中資來源:
    • 本模型之網路架構設計、特徵管線邏輯、訓練流程及推論權重,均由台灣本地團隊自主研發完成。
    • 核心框架均採用開源且具國際公信力之主流專案(PyTorch、OpenCV、MediaPipe),完全無使用任何源自中國大陸或受國際制裁實體之專有模型、私有權重或黑箱相依套件。
  2. 訓練資料在地合規:
    • 訓練語料皆採集自台灣手語(TSL)詞彙與日常會話情境,遵循台灣在地手語文法,無跨境不合規數據。
  3. 商業友善與完整開源:
    • 依 Apache-2.0 協議完整開源模型權重、圖架構代碼(model.py、graph.py)、特徵前處理邏輯(extract.py)與標籤對應表(sign_language_words.json)。

3. 骨架拓撲與特徵前處理 (Skeletal Layout & Preprocessing)

特徵提取採用 55 個空間座標點(透過 MediaPipe Holistic 擷取):

  • **上半身姿態 (節點 0–11)**:肩部、手肘、手腕、髖部關鍵節點。
  • **左手掌與指關節 (節點 12–32)**:完整 21 個手部關節鏈結。
  • **右手掌與指關節 (節點 33–53)**:完整 21 個手部關節鏈結。
  • **虛擬頸點 (節點 54)**:取雙肩中點作為全域空間幾何對齊與正規化原點。

所有輸入序列皆在時間軸上等距重採樣至 64 幀,並依據雙肩距離進行空間尺度正規化。

4. 快速推論範例代碼 (Inference Snippet)

import json
import torch
from model import build_from_checkpoint

# 1. 載入模型檢查點
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
checkpoint = torch.load("best_model.pt", map_location=device)
model, config = build_from_checkpoint(checkpoint)
model.eval()

# 2. 讀取台灣手語詞彙標籤
with open("sign_language_words.json", "r", encoding="utf-8") as f:
    vocab = json.load(f)

# 3. 模擬輸入骨架特徵: (Batch=1, Channels=2, Frames=64, Nodes=55)
dummy_skeleton = torch.randn(1, 2, 64, 55).to(device)

# 4. 執行前向傳播
with torch.no_grad():
    logits = model(dummy_skeleton)
    probabilities = torch.softmax(logits, dim=-1)
    predicted_index = torch.argmax(probabilities, dim=-1).item()

print(f"辨識結果 Gloss: {vocab[predicted_index]}")
print(f"信心度 Confidence: {probabilities[0, predicted_index]:.4f}")

5. 引用來源與學術參考 (Citations & References)

本系統於手語轉自然語言之語序重組模組中採用 Google DeepMind 所開源之 Gemma 多模態模型系列。若於學術研究或專案中使用本專案之演算法與模型,請引用以下標準 BibTeX 條目:

Google DeepMind Gemma 4 模型引用

@misc{gemma20264,
  title={Gemma 4: Open Multimodal Language Models},
  author={Gemma Team},
  year={2026},
  publisher={Google DeepMind},
  howpublished={\url{https://ai.google.dev/gemma/docs/core/model_card_4}}
}

SignBridge-TSL-STGCN 手語模型引用

@software{SignBridge_TSL_STGCN_2026,
  author = {SignBridge Development Team},
  title = {SignBridge-TSL-STGCN: Taiwanese Sign Language Recognition via ST-GCN},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/dYang1/SignBridge-TSL-STGCN}}
}

6. 開源授權 (License)

本專案依 Apache License 2.0 條款授權釋出。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support