AI & ML interests

robotics, physical-ai, embodied-ai, humanoid

Recent Activity

Articles

Organization Card

Noitom Robotics

Making the physical world learnable.

Noitom Robotics builds human-centric data infrastructure for Physical AI. Our thesis, laid out in The World Compiler, is that the bottleneck of Physical AI is not data scarcity but learnability scarcity: the physical world already produces vast embodied intelligence, yet almost none of it exists in a form machines can learn from — and accumulation alone will not close that gap. Volume and learnability must scale together. ModalityNet (modalitynet.com) is our implementation of that thesis: compact, fully structured corpora that compile far larger, weakly structured data into learnable form. The name deliberately echoes ImageNet — where ImageNet organized visual reality into a learnable substrate for vision, ModalityNet organizes physical reality, across modalities, into a learnable substrate for Physical AI.

We build at the human layer because human physical intelligence is the one prior every embodiment, architecture, and paradigm shares. Captured once, it transfers to all of them.

Three corpora, three priors

  • High Precision Human Interaction, Motion with Object and Vision (HiPHI-MOV) — the motion prior. Whole-body motion with interacted-object tracking and side-view vision, captured on hybrid optical–inertial systems and exported as BVH skeletons with end-effector 6-DOF poses. It answers: how does the body move?

  • High Precision Human Interaction, Omni-Modality (HiPHI-OM) — the interaction prior. Hand–object tracking held to millimeter error, synchronized with full-body and finger motion, object mesh tracking, tactile and pressure signals, and ego- and side-view RGB-D. It answers: why do interactions succeed or fail?

  • In-The-Wild (ITW) — the distribution prior. Stereo ego vision, sparse body sensing, and audio captured in unconstrained daily environments, preserving the true distribution of physical reality — long-tail cases included. It answers: does learned behavior hold in the real world?

The three corpora are designed for joint use: the high-precision layers act as a compiler toolchain that raises the learnability of in-the-wild data at scale, narrowing the gap between demonstration and real-world deployment.

Scale

HiPHI production runs at 100,000+ hours per year, and we work with close to 100 companies across robotics, embodied AI, and world modeling.

Work with us

None of this is work we do alone. Sample data, modality definitions, and the ModalityNet Technical Specification are available at modalitynet.com — and we welcome researchers and teams building humanoid policies, world models, and vision-language-action (VLA) systems to build with us.

models 0

None public yet

datasets 0

None public yet