Update README.md
Browse files
README.md
CHANGED
|
@@ -1,3 +1,68 @@
|
|
| 1 |
---
|
|
|
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
language: en
|
| 3 |
license: mit
|
| 4 |
+
tags:
|
| 5 |
+
- pytorch
|
| 6 |
+
- regression
|
| 7 |
+
- dot-product
|
| 8 |
+
- bilinear-networks
|
| 9 |
+
metrics:
|
| 10 |
+
- loss
|
| 11 |
+
pipeline_tag: table-regression
|
| 12 |
---
|
| 13 |
+
|
| 14 |
+
# SBL-NET (Scalar Bilinear Linear Network)
|
| 15 |
+
|
| 16 |
+
A hybrid PyTorch neural network designed to **highly accurately compute the scalar (dot) product** of split sub-vectors without any data normalization (Z-score, etc.).
|
| 17 |
+
|
| 18 |
+
## 🔬 Architecture & Features
|
| 19 |
+
|
| 20 |
+
The main highlight of this model is the integration of a rare **bilinear layer (`nn.Bilinear`)** at the input stage, combined with classic fully connected layers (`nn.Linear`) and the `SELU` activation function.
|
| 21 |
+
* The network accepts an input tensor of shape `[batch, 4]` and splits it into two vectors: `A [batch, 2]` and `B [batch, 2]`.
|
| 22 |
+
* The bilinear layer efficiently extracts cross-features between the vectors, allowing the model to reduce the error to an impressive **0.0569%**.
|
| 23 |
+
|
| 24 |
+
## 📊 Training Results
|
| 25 |
+
|
| 26 |
+
* **Loss Function:** Smooth L1 Loss
|
| 27 |
+
* **Optimizer:** Adam (with StepLR scheduler)
|
| 28 |
+
* **Error Rate:** ~0.0569% (Accuracy ~99.94%)
|
| 29 |
+
* **Extreme Test Case:**
|
| 30 |
+
* Input: `[[-6.0, 70.0, 4.0, -196.0]]`
|
| 31 |
+
* Expected Mathematical Answer: `-13744.0000`
|
| 32 |
+
* Actual Network Prediction: `-13769.5225`
|
| 33 |
+
|
| 34 |
+
## 🧮 Model Statistics
|
| 35 |
+
|
| 36 |
+
* **Total Parameters:** 52,101
|
| 37 |
+
* **Trainable Parameters:** 52,101
|
| 38 |
+
* **Non-trainable Parameters:** 0
|
| 39 |
+
* **Model Size:** ~208 KB (Weights in FP32)
|
| 40 |
+
* **Input Shape:** `[batch_size, 4]`
|
| 41 |
+
* **Output Shape:** `[batch_size, 1]`
|
| 42 |
+
|
| 43 |
+
## 💻 How to Use
|
| 44 |
+
|
| 45 |
+
You can download the architecture file and the model weights directly from this repository:
|
| 46 |
+
|
| 47 |
+
```python
|
| 48 |
+
import torch as t
|
| 49 |
+
from huggingface_hub import hf_hub_download
|
| 50 |
+
|
| 51 |
+
# 1. Download the architecture and weights files (replace YOUR_USERNAME)
|
| 52 |
+
REPO_ID = "YOUR_USERNAME/sbl-net"
|
| 53 |
+
hf_hub_download(repo_id=REPO_ID, filename="model.py", local_dir=".")
|
| 54 |
+
hf_hub_download(repo_id=REPO_ID, filename="model_weights_hybrid.pth", local_dir=".")
|
| 55 |
+
|
| 56 |
+
# 2. Import the model class and load the weights
|
| 57 |
+
from model import WebAISC
|
| 58 |
+
|
| 59 |
+
model = WebAISC()
|
| 60 |
+
model.load_state_dict(t.load("model_weights_hybrid.pth", map_location=t.device('cpu')))
|
| 61 |
+
model.eval()
|
| 62 |
+
|
| 63 |
+
# 3. Inference
|
| 64 |
+
test_input = t.tensor([[-6.0, 70.0, 4.0, -196.0]], dtype=t.float32)
|
| 65 |
+
with t.no_grad():
|
| 66 |
+
prediction = model(test_input)
|
| 67 |
+
print(f"Model prediction: {prediction.item():.4f}")
|
| 68 |
+
```
|