File size: 2,402 Bytes
9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 e5ac32f e68be84 9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 9e16b58 e68be84 e5ac32f 9e16b58 e68be84 9e16b58 e68be84 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | ---
license: apache-2.0
pipeline_tag: audio-to-audio
tags:
- fullsubnet
- speech-enhancement
- audio
- robotics
- IB-Robot
- ascend
- torch
- edge-deployment
---
# Model Card for FullSubNet (IB-Robot)
FullSubNet speech enhancement model with cumulative Laplace normalization for
streaming 4-channel microphone array processing, packaged for the
[IB-Robot](https://atomgit.com/openeuler/IB_Robot) framework with three
deployments sharing one stream contract.
## Deployments
| deployment | backend | artifacts | notes |
|---|---|---|---|
| `ascend_310p` | Ascend ACL (Ascend310P1) | stateful FB/SB OM pair | board-side streaming (LSTM state via state_links) |
| `ascend_310b` | Ascend ACL (Ascend310B1) | stateful FB/SB OM pair (`*_310b_origin.om`) | 310B board-side streaming |
| `torch_cuda` | PyTorch CUDA | — (ckpt in `assets/`) | host-side streaming executor |
| `torch_cpu` | PyTorch CPU | — (ckpt in `assets/`) | CPU fallback (watch the 128 ms hop budget) |
All deployments share `tensor_model/fullsubnet/enhance`
(`observation.audio_4ch [-1,4] -> voice.audio_enhanced_4ch [-1,4]`) with a
stateful stream execution contract (`state_bank_mode: runtime_exclusive`,
`max_open_streams: 1`).
## Weights provenance
- `assets/cum_fullsubnet_best_model_218epochs.tar` = Audio-WestlakeU/FullSubNet
official v0.2 release checkpoint (sha256
`d08d09107eb276b8dc3d2d9fff995f4354a51fa3347125f52f8b9aea7c339f81`)
- The 310P stateful OM pair was converted from the same checkpoint
- `assets/cum_fullsubnet_best_model_218epochs.manifest.json` pins the digest and
`norm_type: cumulative_laplace_norm` (do not mix with offline-norm checkpoints)
Note: the `.tar` file is the upstream PyTorch serialization container
(legacy `torch.save` format), loaded directly by `torch.load` — no extraction
step.
## Repository Structure
- `inference_manifest.json` — deployment routing (schema v3, stateful stream contract)
- `assets/adapter.json` — algorithm contract (STFT 512/256, T=2, look-ahead 2)
- `assets/cum_fullsubnet_best_model_218epochs.tar` — Torch checkpoint
- `artifacts/ascend/fullsubnet/*.om` — stateful FB/SB OM modules (310P)
- `artifacts/ascend_310b/fullsubnet/*.om` — stateful FB/SB OM modules (310B)
## Usage
Host: `voice_asr_service.speech_direction` with backend `stateful_torch_cuda` /
`stateful_torch_cpu`; board: select `ascend_310p` via the unified inference
runtime.
|