File size: 3,068 Bytes
0a089d7
f523628
 
 
 
0a089d7
 
 
 
f523628
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7cab23e
 
f523628
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
title: PyIntel Research
emoji: 👁️
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
---

<div align="center">

# 👁️ PyIntel Research
### *Pioneering Open-Weight Multimodal Edge Intelligence, Spatial Reasoning & Situated Perception*

[![Datasets](https://img.shields.io/badge/🤗%20Datasets-vantage--spatial--pov--mix-blue)](https://huggingface.co/datasets/pyintel/vantage-spatial-pov-mix)
[![License](https://img.shields.io/badge/License-Apache%202.0-green.svg)](https://opensource.org/licenses/Apache-2.0)

</div>

---

## 🧭 About PyIntel
**PyIntel** is an independent open-source AI collective focused on building high-density, compute-efficient foundation models, datasets, and perception architectures for **embodied robotics, edge devices, and ambient assistants**.

Rather than scaling parameters blindly, PyIntel focuses on **data-centric intelligence, 3D geometric grounding, and architectural efficiency**—enabling sub-3B parameter models to reason like frontier systems.

---

## 🔬 Core Research Pillars

### 1. 📐 Situated Spatial Reasoning & 3D Grounding
Teaching vision-language models to ground physical entities in 3D coordinate space, compute line-of-sight raycasts, and understand depth, clearances, and bounding geometries.

### 2. 👤 Perspective-Taking & Theory of Mind (POV of Others)
Breaking free of the "mirror reversal" bug. Training models to translate fluidly between **egocentric** (first-person) and **allocentric** (third-person/observer) reference frames, and accurately model what other agents in the environment can or cannot see.

### 3. 🧠 Episodic Soft Memory Projections
Developing lightweight projection bridges that compress historical video frames and sensor audio into dense **soft memory tokens**—giving edge models long-horizon memory without context window exhaustion.

### 4. ⚡ Real-Time Duplex Interaction
Architecting low-latency, streaming vision-and-voice pipelines for interactive physical assistants ("Jarvis") running on consumer edge hardware.

---

## 📦 Flagship Artifacts & Releases

### 📊 Datasets
* **[`pyintel/vantage-spatial-pov-mix`](https://huggingface.co/datasets/pyintel/vantage-spatial-pov-mix)**  
  *A 7,712-pair harmonized multimodal instruction dataset featuring situated 3D spatial queries, camera perspective translations, and Theory of Mind reasoning with native `<|channel>thought` chains.*
* **[`pyintel/open-board-registry`](https://huggingface.co/datasets/pyintel/open-board-registry)**  
  *Ground-truth structured specifications for 1,700+ microcontroller boards and embedded development platforms (pinouts, voltages, RAM/ROM limits, and communication interfaces).*

### 🤖 Models & Architectures *(In Development)*
* **`PyIntel Vantage-E2B`**  
  *A unified multimodal spatial reasoning engine combining Google DeepMind's `EmbeddingGemma-2` (740M) and `Gemma-4-E2B-it` (2.3B) via an episodic Soft Memory Projection Bridge.*

---

## 🌐 Connect & Collaborate

* **Organization Hub:** [huggingface.co/pyintel](https://huggingface.co/pyintel)