Download README.md from constructelligence/data-centre-nameplates: direct link, hf CLI and curl.
- Browser
- Download file 14.4 kB
-
https://huggingface.co/constructelligence/data-centre-nameplates/resolve/main/README.md
- Command line
-
hf download hf://constructelligence/data-centre-nameplates/README.md
-
curl -L -o README.md https://huggingface.co/constructelligence/data-centre-nameplates/resolve/main/README.md
license: apache-2.0
language:
- en
library_name: custom
pipeline_tag: image-to-text
tags:
- ocr
- document-ai
- information-extraction
- nameplate
- nameplate-ocr
- equipment-nameplate
- data-center
- data-centre
- datacenter
- mission-critical
- mep
- electrical
- equipment-schedule
- schedule-verification
- asset-register
- commissioning
- quality-assurance
- ups
- pdu
- switchgear
- generator
- ats
- transformer
- busway
- battery
- crah
- crac
- chiller
- cooling
- construction
- aec
- edge-ai
- on-device
- privacy-preserving
- tesseract
- browser
Data Centre Nameplates — equipment-nameplate OCR checked against the schedule
Read an equipment nameplate and know if it is the unit the schedule asked for. Point a phone at the nameplate of a UPS, PDU, switchgear, generator, transformer, busway, battery, CRAH or chiller; the model returns the plate's fields — manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, manufacture date — and checks each one against the equipment schedule, flagging a mismatch instead of a silent pass. It builds an asset register and exports it to CSV.
Built for the part of a data-centre job where the gear is the job. The expensive mistakes are a unit delivered at the wrong voltage or rating, or one that quietly never makes it onto the asset register.
Runs on the device. OCR and matching run in the browser (or Node) with no server and no upload. A nameplate photo of live infrastructure never leaves the phone that took it.
▶ Try it live: huggingface.co/spaces/constructelligence/data-centre-nameplates — a browser demo with a sample schedule and a sample plate.
| Task | Image → structured nameplate fields → schedule verification |
| Fields | 13 (manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, mfg date) |
| Equipment | 13 types (UPS, generator, ATS/STS, switchgear, transformer, PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump) |
| OCR | Tesseract.js 5 (English LSTM), pretrained — no fine-tuning |
| Extraction | Deterministic labelled-field parser with an OCR error model |
| Runtime | Browser · Node · edge — no GPU, no server, no data egress |
| Licence | Apache-2.0 |
Why this exists
General OCR reads the text on a nameplate. A data-centre commissioning or QA team needs the text to answer a narrower question: is this the right unit, and does it match what was specified? That means three things a plain OCR call does not do:
- Turn plate text into fields, not paragraphs. "INPUT: 480Y/277 VAC 3 PH 60 HZ" becomes
voltage=480Y/277 V,phase=3,hz=60; "RATING: 750 KVA / 675 KW" becomeskva=750,kw=675. - Survive the way OCR misreads plate lettering. Serifed
V→Vv,K→X(XVA),HZ→Hw, codes split after their punctuation (NPX- 750- 480). The parser repairs these before matching. - Compare to the schedule the way codes actually differ. Model and serial codes are compared with
O/0,I/L/1,S/5,B/8,Z/2folded and punctuation ignored, so a misread is not a false mismatch, while a genuine voltage variant (NPX-750-415vsNPX-750-480) still fails.
The result is a pass/mismatch verdict per plate, with the fields that disagree named.
How it works
photo ──► OCR (Tesseract.js) ──► field parser ──► schedule matcher ──► verdict + asset register CSV
│ │ │
lines + confidence labelled-value rules code folding, tolerance checks
- OCR (
tesseract.js@5, English) returns lines with a per-line confidence. - Field parser scans lines for labelled values (
MODEL,S/N,MVA,MCA,MOCP,FLA,VOLTS,REFRIGERANT,MFG DATE), reading voltages as480Y/277,13.8 kV,208/120, and rejecting look-alikes (480/277with no voltage word is not a voltage;R-513Aand a charge weight are not amps). Lines OCR scored near zero (brushed metal, background texture) are dropped before parsing. - Equipment typing recognises the family from plate wording (UPS, genset, ATS, switchgear, transformer, PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump).
- Schedule matcher ranks candidate tags — a serial match wins outright; otherwise model similarity dominates, with manufacturer and equipment type as tie-breakers. Data halls repeat, so among equal matches the first unscanned tag is offered.
- Field checks compare the reading to the chosen row and return
ok,mismatchorunreadper field, and an overall status:verified,partial,mismatch,unmatchedormissing(scheduled but not scanned).
Fields extracted
| Field | Example | Notes |
|---|---|---|
manufacturer |
Northline Power Systems | ~44 data-centre makers recognised as whole words; otherwise the company-looking line |
model |
NPX-750-480 | MODEL, MOD., M/N, CAT. NO., P/N, PART NO. |
serial |
NL26A01937 | SERIAL, SER., S/N, 5/N, SN |
voltage |
480Y/277 V | every voltage on the plate (transformers and UPSes carry more than one) |
phase |
3 | 1/3, PH, PHASE, Ø |
hz |
60 | 50 or 60, HZ/FREQUENCY |
kva |
750 | thousands separators handled (3,750); European decimals too |
kw |
675 | KW / EKW |
amps |
1002 | labelled (FLA, RLA, CURRENT, AMPS) first |
mca |
612 | minimum circuit ampacity |
mocp |
800 | MOCP / MAX FUSE / MAX BREAKER |
refrigerant |
R-513A | R-410A, 134A, 1234ZE, 1234YF, 513A, 514A, 1233ZD, 454B, 407C, 32, 22 |
mfgDate |
03/2026 | MFG, MANUFACTURED, MFD, BUILD, PRODUCTION |
Equipment types recognised
ups · generator · ats (also STS) · switchgear · transformer · pdu (also RPP) · panelboard ·
busway · battery · cdu · chiller · crah (also CRAC) · pump
Manufacturers recognised include Schneider Electric, Square D, APC, Eaton, ABB, Siemens, Vertiv, Liebert, Caterpillar, Cummins, Kohler, Rolls-Royce, MTU, Generac, Trane, Carrier, Daikin, York, Johnson Controls, Stulz, Munters, nVent, Legrand, Mitsubishi Electric, Toshiba, General Electric, Hubbell, Socomec, Piller, Starline, Raritan, Server Technology, Emerson, Russelectric, ASCO, Hitachi, Riello, Rittal, Motivair, CoolIT, Hitec, EnerSys, Leclanché and Samsung SDI.
Results — OCR and extraction, end to end
Six drawn plates (UPS, diesel genset, chiller, dry-type transformer, PDU, CRAH), each pushed through the degradations a phone photo actually has, then run through the app's own OCR + parser and scored field by field. 56 checks per condition (six plates × their fields), 13 conditions.
| Condition | Field accuracy | ms/plate |
|---|---|---|
| clean | 100% (56/56) | 1236 |
| small (640 px) | 100% (56/56) | 897 |
| blur | 100% (56/56) | 540 |
| noise | 100% (56/56) | 748 |
| low contrast | 100% (56/56) | 557 |
| glare | 100% (56/56) | 532 |
| tilt (7°) | 100% (56/56) | 612 |
| keystone | 100% (56/56) | 474 |
| dark anodised plate | 100% (56/56) | 808 |
| JPEG | 100% (56/56) | 524 |
| plate small in a cluttered scene | 100% (56/56) | 895 |
| sideways | 100% (56/56) | 554 |
| phone (blur + glare + noise + tilt + scene + JPEG) | 98.2% (55/56) | 1607 |
| Overall | 99.9% (727/728) | 768 |
The single miss is one hz value on the hardest "phone" composite. The benchmark measures what OCR loses,
not parser coverage: the ground truth is the parser's own reading of the clean plate text, so every reduction
is an OCR failure the parser could not repair. Reproduce it with node tests/nameplates-bench.mjs.
Read this honestly. These are crisp, machine-drawn plates. A real nameplate photographed on a live unit is harder: cast shadows, embossed lettering, print over brushed metal, a plate half out of frame. Expect the degradation suite, not the clean row, and treat every low-confidence field as one a person must confirm.
Fine-tuning — the higher-accuracy server model
The engine above is the on-device path. Alongside it, a Donut model
(naver-clova-ix/donut-base) is fine-tuned to read a plate
straight into fields (<s_nameplate><s_type>ups</s_type><s_manufacturer>…) for a server / endpoint where
the browser engine falls short. Training data is synthetic plates with real-photo degradations, plus corrected
real plates from the field. Once trained, the weights are published in a sibling repository and linked here.
Nothing on this page claims the accuracy of that model yet — it is the roadmap, not a result.
Intended use
- Commissioning and QA walks: photograph each unit as it is installed; confirm the delivered unit matches the equipment schedule; capture an asset register as you go.
- Receiving and delivery checks: flag the unit that arrived at the wrong voltage or rating before it is set.
- Register completion: track scheduled tags that have no plate scanned against them yet.
- Takeoff and submittal review: pull model/rating fields off a plate photo into a spreadsheet.
Out of scope — do not use it for
- Safety, code compliance or energisation decisions. It reads a plate; it does not verify that a unit is safe to energise, correctly protected, or code-compliant.
- An as-built or a legal record on its own. OCR misreads happen; low-confidence fields are highlighted for a person to check, and any field that disagrees with the schedule must be confirmed by eye.
- Serial-number evidence in a dispute without human verification. Two plates can share a misread serial; the register flags duplicate serials but cannot tell a repeated plate from a copied number.
Limitations
- English nameplates, Latin script. Non-Latin or handwritten plates are out of scope.
- Glare, blur and angle degrade OCR. The parser repairs common misreads, but a plate that is small in the
frame, badly glared or very low contrast will lose fields (those fields come back
unread, not invented). - Deterministic parser, not a language model. It reports only what it can match to a rule; it will not infer a missing digit.
- The schedule must be structured (CSV with a Tag or Model column). Free-form schedule PDFs are not parsed here.
Quickstart
The engine is plain ES modules with no build step. nameplate-model.js is the extractor and verifier;
schedule-import.js is the CSV reader it depends on.
import { parseNameplate, parseEquipmentSchedule, candidatesFor, checkAgainst, statusOf } from './nameplate-model.js';
// 1. Parse plate text (from Tesseract, or typed by hand)
const reading = parseNameplate(`NORTHLINE POWER SYSTEMS
UNINTERRUPTIBLE POWER SUPPLY
MODEL NO: NPX-750-480
SERIAL NO: NL26A01937
INPUT: 480Y/277 VAC 3 PH 60 HZ
RATING: 750 KVA / 675 KW
INPUT CURRENT: 1002 A
MFG DATE: 03/2026`);
console.log(reading.type); // 'ups'
console.log(reading.fields.voltage); // { value: '480Y/277 V', values: [480, 277], conf, line }
// 2. Match against the equipment schedule and check each field
const { items } = parseEquipmentSchedule(scheduleCsv);
const [best] = candidatesFor(reading, items);
const checks = checkAgainst(reading, best.item);
console.log(statusOf(checks, best.item.tag)); // 'verified' | 'partial' | 'mismatch' | ...
In a browser, feed Tesseract's data.lines ([{ text, confidence, bbox }]) straight into parseNameplate
to keep OCR confidence and the plate line each value came from.
Evaluation
- Extraction/verification logic: 12 unit tests over the parser, voltage reader, code folding, schedule
parsing, candidate ranking and mismatch detection — all passing (
node --test tests/nameplate-model.test.mjs). - OCR end-to-end: a field-by-field benchmark draws plates for six equipment types and puts each through the
degradations real phone photos have — small, blurred, noisy, low-contrast, glare, tilt, keystone, dark
anodised plate, JPEG, plate-small-in-scene and sideways — then scores the app's own OCR + parsing pipeline
field by field. Run it with
node tests/nameplates-bench.mjs.
Repository contents
| File | What it is |
|---|---|
nameplate-model.js |
the extractor, matcher and register builder (ES module) |
schedule-import.js |
the CSV schedule reader it depends on |
progress-project.js |
a helper the schedule reader imports |
sample_schedule.csv |
a 10-row data-centre equipment schedule |
sample_plates.json |
two sample plates: one that verifies, one flagged on model and voltage |
eval/ |
the OCR benchmark harness output |
config.json |
machine-readable fields, types and statuses |
Provenance
An open extraction and verification engine from Constructelligence, the same code path that powers the
Nameplates mode of BuildVision. No customer data is used; sample plates use fictional makers and models.
It is the document-AI sibling of the vision models
construction-site-safety-hazards
and electrical-circuit-connectivity.
Safety
This is a prompt to look, not a finding. It is not safety-rated, not a compliance decision, and not a substitute for inspection by a competent person. Verify every flagged field against the physical plate before acting on it.
Licence and trademarks
Code released under Apache-2.0. Brand and product names (Schneider Electric, Eaton, Vertiv, Caterpillar, Trane, …) are trademarks of their owners and are matched only to identify user-supplied plates; this project is independent and not affiliated with or endorsed by them.
Citation
@misc{constructelligence_nameplates,
title = {Data Centre Nameplates: equipment-nameplate OCR checked against the equipment schedule},
author = {Constructelligence},
year = {2026},
howpublished = {\url{https://huggingface.co/constructelligence/data-centre-nameplates}}
}