File size: 14,448 Bytes
5e2d1eb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65f33bb
 
5e2d1eb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b224bd3
 
 
 
 
 
 
 
 
 
5e2d1eb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
94e56dd
5e2d1eb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
---
license: apache-2.0
language:
  - en
library_name: custom
pipeline_tag: image-to-text
tags:
  - ocr
  - document-ai
  - information-extraction
  - nameplate
  - nameplate-ocr
  - equipment-nameplate
  - data-center
  - data-centre
  - datacenter
  - mission-critical
  - mep
  - electrical
  - equipment-schedule
  - schedule-verification
  - asset-register
  - commissioning
  - quality-assurance
  - ups
  - pdu
  - switchgear
  - generator
  - ats
  - transformer
  - busway
  - battery
  - crah
  - crac
  - chiller
  - cooling
  - construction
  - aec
  - edge-ai
  - on-device
  - privacy-preserving
  - tesseract
  - browser
---

# Data Centre Nameplates — equipment-nameplate OCR checked against the schedule

**Read an equipment nameplate and know if it is the unit the schedule asked for.** Point a phone at the
nameplate of a UPS, PDU, switchgear, generator, transformer, busway, battery, CRAH or chiller; the model returns
the plate's fields — manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant,
manufacture date — and **checks each one against the equipment schedule**, flagging a mismatch instead of a
silent pass. It builds an asset register and exports it to CSV.

Built for the part of a data-centre job where the gear *is* the job. The expensive mistakes are a unit
delivered at the wrong voltage or rating, or one that quietly never makes it onto the asset register.

> **Runs on the device.** OCR and matching run in the browser (or Node) with no server and no upload. A
> nameplate photo of live infrastructure never leaves the phone that took it.

**▶ Try it live:** [huggingface.co/spaces/constructelligence/data-centre-nameplates](https://huggingface.co/spaces/constructelligence/data-centre-nameplates) — a browser demo with a sample schedule and a sample plate.

| | |
|---|---|
| **Task** | Image → structured nameplate fields → schedule verification |
| **Fields** | 13 (manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, mfg date) |
| **Equipment** | 13 types (UPS, generator, ATS/STS, switchgear, transformer, PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump) |
| **OCR** | Tesseract.js 5 (English LSTM), pretrained — no fine-tuning |
| **Extraction** | Deterministic labelled-field parser with an OCR error model |
| **Runtime** | Browser · Node · edge — no GPU, no server, no data egress |
| **Licence** | Apache-2.0 |

---

## Why this exists

General OCR reads the text on a nameplate. A data-centre commissioning or QA team needs the text to answer a
narrower question: **is this the right unit, and does it match what was specified?** That means three things a
plain OCR call does not do:

1. **Turn plate text into fields, not paragraphs.** "INPUT: 480Y/277 VAC 3 PH 60 HZ" becomes
   `voltage=480Y/277 V`, `phase=3`, `hz=60`; "RATING: 750 KVA / 675 KW" becomes `kva=750`, `kw=675`.
2. **Survive the way OCR misreads plate lettering.** Serifed `V`→`Vv`, `K`→`X` (`XVA`), `HZ`→`Hw`, codes split
   after their punctuation (`NPX- 750- 480`). The parser repairs these before matching.
3. **Compare to the schedule the way codes actually differ.** Model and serial codes are compared with `O/0`,
   `I/L/1`, `S/5`, `B/8`, `Z/2` folded and punctuation ignored, so a misread is not a false mismatch, while a
   genuine voltage variant (`NPX-750-415` vs `NPX-750-480`) still fails.

The result is a pass/mismatch verdict per plate, with the fields that disagree named.

## How it works

```
 photo ──► OCR (Tesseract.js) ──► field parser ──► schedule matcher ──► verdict + asset register CSV
             │                        │                    │
      lines + confidence        labelled-value rules   code folding, tolerance checks
```

1. **OCR** (`tesseract.js@5`, English) returns lines with a per-line confidence.
2. **Field parser** scans lines for labelled values (`MODEL`, `S/N`, `MVA`, `MCA`, `MOCP`, `FLA`, `VOLTS`,
   `REFRIGERANT`, `MFG DATE`), reading voltages as `480Y/277`, `13.8 kV`, `208/120`, and rejecting look-alikes
   (`480/277` with no voltage word is not a voltage; `R-513A` and a charge weight are not amps). Lines OCR
   scored near zero (brushed metal, background texture) are dropped before parsing.
3. **Equipment typing** recognises the family from plate wording (UPS, genset, ATS, switchgear, transformer,
   PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump).
4. **Schedule matcher** ranks candidate tags — a serial match wins outright; otherwise model similarity
   dominates, with manufacturer and equipment type as tie-breakers. Data halls repeat, so among equal matches
   the first **unscanned** tag is offered.
5. **Field checks** compare the reading to the chosen row and return `ok`, `mismatch` or `unread` per field, and
   an overall status: `verified`, `partial`, `mismatch`, `unmatched` or `missing` (scheduled but not scanned).

## Fields extracted

| Field | Example | Notes |
|---|---|---|
| `manufacturer` | Northline Power Systems | ~44 data-centre makers recognised as whole words; otherwise the company-looking line |
| `model` | NPX-750-480 | `MODEL`, `MOD.`, `M/N`, `CAT. NO.`, `P/N`, `PART NO.` |
| `serial` | NL26A01937 | `SERIAL`, `SER.`, `S/N`, `5/N`, `SN` |
| `voltage` | 480Y/277 V | every voltage on the plate (transformers and UPSes carry more than one) |
| `phase` | 3 | `1`/`3`, `PH`, `PHASE`, `Ø` |
| `hz` | 60 | 50 or 60, `HZ`/`FREQUENCY` |
| `kva` | 750 | thousands separators handled (`3,750`); European decimals too |
| `kw` | 675 | `KW` / `EKW` |
| `amps` | 1002 | labelled (`FLA`, `RLA`, `CURRENT`, `AMPS`) first |
| `mca` | 612 | minimum circuit ampacity |
| `mocp` | 800 | `MOCP` / `MAX FUSE` / `MAX BREAKER` |
| `refrigerant` | R-513A | R-410A, 134A, 1234ZE, 1234YF, 513A, 514A, 1233ZD, 454B, 407C, 32, 22 |
| `mfgDate` | 03/2026 | `MFG`, `MANUFACTURED`, `MFD`, `BUILD`, `PRODUCTION` |

## Equipment types recognised

`ups` · `generator` · `ats` (also STS) · `switchgear` · `transformer` · `pdu` (also RPP) · `panelboard` ·
`busway` · `battery` · `cdu` · `chiller` · `crah` (also CRAC) · `pump`

Manufacturers recognised include Schneider Electric, Square D, APC, Eaton, ABB, Siemens, Vertiv, Liebert,
Caterpillar, Cummins, Kohler, Rolls-Royce, MTU, Generac, Trane, Carrier, Daikin, York, Johnson Controls,
Stulz, Munters, nVent, Legrand, Mitsubishi Electric, Toshiba, General Electric, Hubbell, Socomec, Piller,
Starline, Raritan, Server Technology, Emerson, Russelectric, ASCO, Hitachi, Riello, Rittal, Motivair, CoolIT,
Hitec, EnerSys, Leclanché and Samsung SDI.

## Results — OCR and extraction, end to end

Six drawn plates (UPS, diesel genset, chiller, dry-type transformer, PDU, CRAH), each pushed through the
degradations a phone photo actually has, then run through the app's own OCR + parser and scored **field by
field**. **56 checks per condition** (six plates × their fields), 13 conditions.

| Condition | Field accuracy | ms/plate |
|---|:---:|---:|
| clean | **100%** (56/56) | 1236 |
| small (640 px) | **100%** (56/56) | 897 |
| blur | **100%** (56/56) | 540 |
| noise | **100%** (56/56) | 748 |
| low contrast | **100%** (56/56) | 557 |
| glare | **100%** (56/56) | 532 |
| tilt (7°) | **100%** (56/56) | 612 |
| keystone | **100%** (56/56) | 474 |
| dark anodised plate | **100%** (56/56) | 808 |
| JPEG | **100%** (56/56) | 524 |
| plate small in a cluttered scene | **100%** (56/56) | 895 |
| sideways | **100%** (56/56) | 554 |
| phone (blur + glare + noise + tilt + scene + JPEG) | **98.2%** (55/56) | 1607 |
| **Overall** | **99.9% (727/728)** | **768** |

The single miss is one `hz` value on the hardest "phone" composite. The benchmark measures what **OCR** loses,
not parser coverage: the ground truth is the parser's own reading of the clean plate text, so every reduction
is an OCR failure the parser could not repair. Reproduce it with `node tests/nameplates-bench.mjs`.

**Read this honestly.** These are crisp, machine-drawn plates. A real nameplate photographed on a live unit is
harder: cast shadows, embossed lettering, print over brushed metal, a plate half out of frame. Expect the
degradation suite, not the clean row, and treat every low-confidence field as one a person must confirm.


## Fine-tuning — the higher-accuracy server model

The engine above is the on-device path. Alongside it, a **Donut** model
([`naver-clova-ix/donut-base`](https://huggingface.co/naver-clova-ix/donut-base)) is fine-tuned to read a plate
**straight into fields** (`<s_nameplate><s_type>ups</s_type><s_manufacturer>…`) for a server / endpoint where
the browser engine falls short. Training data is synthetic plates with real-photo degradations, plus corrected
real plates from the field. Once trained, the weights are published in a sibling repository and linked here.

Nothing on this page claims the accuracy of that model yet — it is the roadmap, not a result.

## Intended use

- **Commissioning and QA walks:** photograph each unit as it is installed; confirm the delivered unit matches
  the equipment schedule; capture an asset register as you go.
- **Receiving and delivery checks:** flag the unit that arrived at the wrong voltage or rating before it is set.
- **Register completion:** track scheduled tags that have no plate scanned against them yet.
- **Takeoff and submittal review:** pull model/rating fields off a plate photo into a spreadsheet.

### Out of scope — do not use it for

- **Safety, code compliance or energisation decisions.** It reads a plate; it does not verify that a unit is
  safe to energise, correctly protected, or code-compliant.
- **An as-built or a legal record on its own.** OCR misreads happen; low-confidence fields are highlighted for a
  person to check, and any field that disagrees with the schedule must be confirmed by eye.
- **Serial-number evidence in a dispute** without human verification. Two plates can share a misread serial;
  the register flags duplicate serials but cannot tell a repeated plate from a copied number.

## Limitations

- **English nameplates**, Latin script. Non-Latin or handwritten plates are out of scope.
- **Glare, blur and angle degrade OCR.** The parser repairs common misreads, but a plate that is small in the
  frame, badly glared or very low contrast will lose fields (those fields come back `unread`, not invented).
- **Deterministic parser, not a language model.** It reports only what it can match to a rule; it will not
  infer a missing digit.
- **The schedule must be structured** (CSV with a Tag or Model column). Free-form schedule PDFs are not parsed
  here.

## Quickstart

The engine is plain ES modules with no build step. `nameplate-model.js` is the extractor and verifier;
`schedule-import.js` is the CSV reader it depends on.

```js
import { parseNameplate, parseEquipmentSchedule, candidatesFor, checkAgainst, statusOf } from './nameplate-model.js';

// 1. Parse plate text (from Tesseract, or typed by hand)
const reading = parseNameplate(`NORTHLINE POWER SYSTEMS
UNINTERRUPTIBLE POWER SUPPLY
MODEL NO: NPX-750-480
SERIAL NO: NL26A01937
INPUT: 480Y/277 VAC  3 PH  60 HZ
RATING: 750 KVA / 675 KW
INPUT CURRENT: 1002 A
MFG DATE: 03/2026`);

console.log(reading.type);            // 'ups'
console.log(reading.fields.voltage);  // { value: '480Y/277 V', values: [480, 277], conf, line }

// 2. Match against the equipment schedule and check each field
const { items } = parseEquipmentSchedule(scheduleCsv);
const [best] = candidatesFor(reading, items);
const checks = checkAgainst(reading, best.item);
console.log(statusOf(checks, best.item.tag)); // 'verified' | 'partial' | 'mismatch' | ...
```

In a browser, feed Tesseract's `data.lines` (`[{ text, confidence, bbox }]`) straight into `parseNameplate`
to keep OCR confidence and the plate line each value came from.

## Evaluation

- **Extraction/verification logic:** 12 unit tests over the parser, voltage reader, code folding, schedule
  parsing, candidate ranking and mismatch detection — all passing (`node --test tests/nameplate-model.test.mjs`).
- **OCR end-to-end:** a field-by-field benchmark draws plates for six equipment types and puts each through the
  degradations real phone photos have — small, blurred, noisy, low-contrast, glare, tilt, keystone, dark
  anodised plate, JPEG, plate-small-in-scene and sideways — then scores the app's own OCR + parsing pipeline
  field by field. Run it with `node tests/nameplates-bench.mjs`.

## Repository contents

| File | What it is |
|---|---|
| `nameplate-model.js` | the extractor, matcher and register builder (ES module) |
| `schedule-import.js` | the CSV schedule reader it depends on |
| `progress-project.js` | a helper the schedule reader imports |
| `sample_schedule.csv` | a 10-row data-centre equipment schedule |
| `sample_plates.json` | two sample plates: one that verifies, one flagged on model and voltage |
| `eval/` | the OCR benchmark harness output |
| `config.json` | machine-readable fields, types and statuses |

## Provenance

An open extraction and verification engine from Constructelligence, the same code path that powers the
**Nameplates** mode of BuildVision. No customer data is used; sample plates use fictional makers and models.
It is the document-AI sibling of the vision models
[`construction-site-safety-hazards`](https://huggingface.co/constructelligence/construction-site-safety-hazards)
and [`electrical-circuit-connectivity`](https://huggingface.co/constructelligence/electrical-circuit-connectivity).

## Safety

This is a **prompt to look**, not a finding. It is not safety-rated, not a compliance decision, and not a
substitute for inspection by a competent person. Verify every flagged field against the physical plate before
acting on it.

## Licence and trademarks

Code released under **Apache-2.0**. Brand and product names (Schneider Electric, Eaton, Vertiv, Caterpillar,
Trane, …) are trademarks of their owners and are matched only to identify user-supplied plates; this project is
independent and not affiliated with or endorsed by them.

## Citation

```bibtex
@misc{constructelligence_nameplates,
  title  = {Data Centre Nameplates: equipment-nameplate OCR checked against the equipment schedule},
  author = {Constructelligence},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/constructelligence/data-centre-nameplates}}
}
```