File size: 7,761 Bytes
e36d0b4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 | ---
license: apache-2.0
pipeline_tag: image-to-3d
tags:
- pear
- human-mesh-recovery
- smplx
- robotics
- IB-Robot
- ascend
- edge-deployment
base_model:
- BestWJH/PEAR_models
---
# Model Card for PEAR Parameter Network (IB-Robot)
The parameter-regression network of PEAR (Pixel-aligned Expressive humAn mesh Recovery),
packaged for the [IB-Robot](https://atomgit.com/openeuler/IB_Robot) framework. From a
person crop it regresses SMPL-X / FLAME / camera parameters for expressive 3D human
understanding in HRI.
Person crops come from
[openEuler/yolox_x_640](https://huggingface.co/openEuler/yolox_x_640).
**Scope:** this bundle contains the PEAR image encoder and parameter regression heads only.
EHM / SMPL-X LBS, mesh generation and rendering stay on the host CPU and are **not** part of
the OM.
## Repository Structure
- `inference_manifest.json` — deployment routing (schema v3)
- `assets/adapter.json` — adapter identity (`pear_parameter_network` / `predict_parameters`)
- `assets/pear_model.pt` — original PyTorch checkpoint the OM was converted from
- `artifacts/ascend_310p/pear_parameter_network_bs1.om` — Ascend 310P1 OM (batch 1)
## Deployment Backends
| Target | Backend | Runtime | Hardware |
|--------|---------|---------|----------|
| `ascend_310p` | ascend | ACL | Ascend 310P1 |
**Input:** `pear.input` float32 [1,3,256,256] NCHW (`input`)
**Outputs:** eight tensors, fixed order
| # | semantic | shape |
|---|----------|-------|
| 0 | `smplx_pose_raw` | [1,312] |
| 1 | `smplx_scale` | [1,6] |
| 2 | `smplx_shape` | [1,200] |
| 3 | `smplx_expression` | [1,50] |
| 4 | `flame_pose` | [1,14] |
| 5 | `flame_shape` | [1,300] |
| 6 | `flame_expression` | [1,50] |
| 7 | `camera_raw` | [1,3] |
`smplx_pose_raw` splits as:
```text
0:6 global_orient
6:132 body_pose, 21 × 6D
132:222 left_hand_pose, 15 × 6D
222:312 right_hand_pose, 15 × 6D
```
The 6D values are **not** Euler angles: decode to rotation matrices first, then to
axis-angle radians if needed.
### Preprocessing contract (`pear-rgb-crop256-bgr-imagenet-v1`)
```text
person bbox xyxy in source-image coordinates
→ centre (cx, cy), side = max(w, h) × 1.25
→ square affine crop (cv2.INTER_LINEAR, BORDER_CONSTANT 0)
→ 256×256 BGR
→ NCHW float32, divided by 255
→ ImageNet normalize (mean 0.485/0.456/0.406, std 0.229/0.224/0.225)
→ width slice [:, :, :, 32:-32]
```
The backbone therefore sees 256×192 content. Crops must be taken in source-video
coordinates, not in YOLOX's 640×640 letterbox space.
## Source Model
This bundle's torch weights (`assets/pear_model.pt`) are the upstream PEAR checkpoint,
unmodified:
- **Model weights (HuggingFace):** [BestWJH/PEAR_models](https://huggingface.co/BestWJH/PEAR_models) — `pear_model.pt`, fetched by upstream code via `hf_hub_download(repo_id="BestWJH/PEAR_models", filename="pear_model.pt")`
- **Project page:** <https://wujh2001.github.io/PEAR/>
- `pear_model.pt` — 2,685,908,343 bytes, sha256 `be82dfa06e7b0608c6440058dfa0794d9b2ceee69f6e5b09bf41bb0076abeb18`
Upstream states this is the *initial* release of the PEAR model rather than the final version
presented in the paper; it may underperform on complex poses.
### Source code
```bash
git clone https://github.com/Pixel-Talk/PEAR.git
git -C PEAR checkout 230fa1534367c9f357c1c192a328cdc87ab4491c
```
- **Repository:** <https://github.com/Pixel-Talk/PEAR> (Apache-2.0)
- **Commit:** `230fa1534367c9f357c1c192a328cdc87ab4491c` on `main` — 2026-08-01, *"Update app.py"*
- The clone used for export carries no submodules (upstream has no `.gitmodules`) and no
local patches; the working tree differs from that commit only in file permission bits.
The Ascend OM was converted from those weights via ONNX with external data
(`pear_parameter_network_bs1.onnx` + `.data`, consolidated `.data` sha256
`76d0b08fea2a17133aa62b718e1faaf329e0167a42638ffe63c177951b5ab766`), with ATC
`--soc_version=Ascend310P1`. The OM (`pear_parameter_network_bs1.om`, sha256
`67798d9f1da61fba4e5b706b20acf82cf2daf030a3239f89709e655421c84e81`) is not re-trained.
## Assets Not Included
Consuming the parameter outputs (rotation decode, parameter bookkeeping) needs `numpy` only.
Reconstructing meshes / 3D joints additionally requires the SMPL-X, FLAME and MANO body
models, which are **not** redistributed here because their licenses do not permit it.
Obtain them yourself from the original sites and accept their terms:
- SMPL-X — <https://smpl-x.is.tue.mpg.de/> (`SMPLX_NEUTRAL_2020.npz`)
- FLAME 2020 — <https://flame.is.tue.mpg.de/> (`generic_model.pkl`)
- MANO — <https://mano.is.tue.mpg.de/> (`MANO_LEFT.pkl`, `MANO_RIGHT.pkl`)
These are research-licensed assets and are generally not usable for commercial deployment
without a separate agreement.
## Validation
Board evidence recorded on a real Ascend 310P1 over a 368-frame, 30 FPS clip
(`npu-smi info` SoC = `Ascend310P1`; `Ascend310P3` is not a valid target for this device).
Numerical alignment vs. the PyTorch reference (strict gate passed):
```text
raw parameter max abs diff: 0.036261 (gate 0.05)
rotation worst P95: 0.068421° (gate 1.0°)
body joint max: 0.332952° (gate 5.0°)
left hand joint max: 0.211992° (gate 5.0°)
right hand joint max: 0.211410° (gate 5.0°)
left/right swap: none (direct mean 0.027°, swapped mean 45.964°)
NaN/Inf: none; rotation matrices orthonormal to ~1e-7
```
Latency, same run:
```text
OM only: mean 26.298 ms, P95 27.559 ms, max 31.641 ms
full ACL: mean 27.784 ms, P95 29.602 ms, max 33.094 ms
ACL + CPU parameter processing: mean 28.542 ms, P95 30.542 ms, max 33.929 ms
```
**Status: engineering `GO`, strict real-time `CONDITIONAL`.** Against a 33.33 ms budget at
30 Hz, mean/P50/P95 fit, but the worst frame exceeds it by ~0.60 ms. Systems with a hard
per-frame deadline must budget for that overrun.
### Semantics caveats
- Pose outputs are SMPL-X **local joint rotations**, not robot motor angles. Driving a robot
additionally requires a SMPL-X→joint mapping, axis transforms, zero offsets, sign and unit
conversion, joint limits and velocity/acceleration limits — none of which are in this bundle.
- `camera_raw` and the recovered body are in relative/model coordinates. Without real camera
intrinsics and root depth they must not be presented as absolute camera XYZ.
## Usage
Select the `ascend_310p` deployment through the IB-Robot unified inference runtime; the
bundle is consumed as an external model bundle (it is not stored in the IB-Robot Git
repository).
```python
from inference_manifest import load_inference_manifest
validated = load_inference_manifest("models/pear_parameter_network", "ascend_310p")
```
## License
Code and packaging: Apache-2.0. The PEAR weights are redistributed under the upstream
Apache-2.0 license of [BestWJH/PEAR_models](https://huggingface.co/BestWJH/PEAR_models) /
[Pixel-Talk/PEAR](https://github.com/Pixel-Talk/PEAR). The SMPL-X / FLAME / MANO body models
needed for mesh reconstruction are **not** included and carry their own restrictive licenses.
## Citation
@misc{wu2026pear,
title = {PEAR: Pixel-aligned Expressive humAn mesh Recovery},
author = {Jiahao Wu and Yunfei Liu and Lijian Lin and Ye Zhu and Lei Zhu and Jingyi Li and Yu Li},
year = {2026},
eprint = {2601.22693},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2601.22693}
}
@software{ib_robot,
title = {IB-Robot: Intelligence Boom Robot},
url = {https://atomgit.com/openeuler/IB_Robot},
license = {Apache-2.0}
}
|