File size: 10,530 Bytes
1d3f990
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
# Ahmad Bot β€” Ollama Local LLM Setup

## Architecture Change

**From:** WebLLM (browser-based, hardcoded model strings, slow CDN loads)  
**To:** Ollama API (local LLM server, real models, instant response)

---

## Quick Start

### 1. Install Ollama

Download from: https://ollama.ai

Or install via package manager:
```bash

# macOS

brew install ollama



# Ubuntu/Debian

curl https://ollama.ai/install.sh | sh



# Windows

# Download: https://ollama.ai/download/windows

```

### 2. Pull a Small Model

```bash

# TinyLLaMA (fastest, works on CPU)

ollama pull tinyllama



# Or: Neural Chat (slightly larger, better quality)

ollama pull neural-chat

```

Model sizes:
- `tinyllama` β€” 440 MB (fastest, good for testing)
- `neural-chat` β€” 3.8 GB (better responses)
- `mistral` β€” 4 GB (strong performance)
- `llama2` β€” 3.8 GB (general purpose)

### 3. Start Ollama Server

```bash

ollama serve

```

This starts the Ollama API on `http://localhost:11434`

**Terminal output:**
```

2026-07-27 18:00:00 API server started at http://localhost:11434

```

### 4. Open Ahmad Bot

```bash

# Local development

cd rowm-polymorphic-notebook

python3 -m http.server 8000



# Visit: http://localhost:8000/index-app.html

```

**Or** (live GitHub Pages):
```

https://snapkittywest.github.io/rowm-polymorphic-notebook/index-app.html

```

### 5. Connect to Ollama

1. Click **"CONNECT TO OLLAMA"** button
2. Wait for connection (should be instant if `ollama serve` is running)
3. If successful, status changes to **READY** (green)
4. Type a question and click **SEND**
5. Response streams in real-time

---

## Testing Workflow

### Terminal 1: Start Ollama
```bash

ollama serve

```

### Terminal 2: Start Web Server
```bash

cd rowm-polymorphic-notebook

python3 -m http.server 8000

```

### Browser: Visit Page
```

http://localhost:8000/index-app.html

```

### Browser Console (F12)
```javascript

// Check if Ollama connected

console.log('Engine state:', window.ahmadEngine?.getState());



// Check available models

fetch('http://localhost:11434/api/tags').then(r => r.json()).then(d => console.log(d.models));



// Manually test API

fetch('http://localhost:11434/api/generate', {

  method: 'POST',

  body: JSON.stringify({

    model: 'tinyllama',

    prompt: 'Hello, what is your name?',

    stream: false

  })

}).then(r => r.json()).then(d => console.log(d.response));

```

---

## Troubleshooting

### Issue: "Ollama not running at http://localhost:11434"

**Solution:** Start Ollama server
```bash

ollama serve

```

Check it's running:
```bash

curl http://localhost:11434/api/tags

# Should return JSON with list of models

```

### Issue: "Model 'tinyllama' not available"

**Solution:** Pull the model
```bash

ollama pull tinyllama

# OR

ollama pull neural-chat

```

Check available models:
```bash

ollama list

```

### Issue: CORS error in browser console

**Why it happens:** Ollama only accepts requests from `http://localhost:*` and `127.0.0.1:*`

**Solution:** Either:
1. Run page locally: `python3 -m http.server 8000`
2. Or use `http://127.0.0.1/...` instead of `http://localhost/...`

Cannot use GitHub Pages (HTTPS) to connect to local Ollama (HTTP) β€” browser blocks mixed content.

### Issue: Model loading very slow

**Why it happens:** First run downloads full model weights (~1-4 GB depending on model)

**Solution:** Just wait. Subsequent runs will be much faster (model cached in memory).

Speeds:
- CPU: 5-15 tokens/second
- GPU: 20-100+ tokens/second

---

## Models Reference

### Recommended for Ahmad Bot

| Model | Size | Speed | Quality | Use Case |
|-------|------|-------|---------|----------|
| **tinyllama** | 440 MB | Very Fast | Basic | Testing, fast responses |
| **neural-chat** | 3.8 GB | Fast | Good | General chat |
| **mistral** | 4 GB | Medium | Excellent | Best balance |
| **llama2** | 3.8 GB | Medium | Very Good | General purpose |
| **openhermes** | 7 GB | Slow | Excellent | High-quality responses |

### Installation

```bash

# Fast testing

ollama pull tinyllama



# Best performance/quality

ollama pull neural-chat



# Strongest model

ollama pull mistral



# Remove model

ollama rm tinyllama

```

---

## API Reference

### Check Connection

```bash

curl http://localhost:11434/api/tags

```

Response:
```json

{

  "models": [

    {

      "name": "tinyllama:latest",

      "modified_at": "2026-07-27T18:00:00.000Z",

      "size": 440000000

    }

  ]

}

```

### Generate Response (Non-Streaming)

```bash

curl http://localhost:11434/api/generate \

  -X POST \

  -H "Content-Type: application/json" \

  -d '{

    "model": "tinyllama",

    "prompt": "What is 2+2?",

    "stream": false

  }'

```

Response:
```json

{

  "model": "tinyllama",

  "created_at": "2026-07-27T18:00:00.000Z",

  "response": " 2+2=4.",

  "done": true

}

```

### Generate Response (Streaming)

```bash

curl http://localhost:11434/api/generate \

  -X POST \

  -H "Content-Type: application/json" \

  -d '{

    "model": "tinyllama",

    "prompt": "Hello",

    "stream": true

  }' | jq -R 'fromjson?'

```

Response (line-by-line):
```json

{"model":"tinyllama","created_at":"...","response":" ","done":false}

{"model":"tinyllama","created_at":"...","response":"I","done":false}

{"model":"tinyllama","created_at":"...","response":"'m","done":false}

...

{"model":"tinyllama","created_at":"...","response":"","done":true}

```

---

## Code Changes

### What Changed in Ahmad Bot

**Before (WebLLM):**
```javascript

// index-app.html: 100+ lines of CDN loading logic

<script src="https://cdn.jsdelivr.net/npm/@mlc-ai/web-llm@0.2.32/...">

```

**After (Ollama):**
```javascript

// index-app.html: Just two script tags

<script src="./js/ahmad-jit-engine.js"></script>

<script src="./js/ahmad-jit-ui.js"></script>



// No CDN needed!

```

**Before (WebLLM):**
```javascript

// ahmad-jit-engine.js: ~200 lines

new webllm.MLCEngine()

await engine.reload('Qwen2-0.5B-Instruct-q4f32_1-MLC')

```

**After (Ollama):**
```javascript

// ahmad-jit-engine.js: ~250 lines, clearer

new AhmadJITEngine('http://localhost:11434')

await engine.initialize('tinyllama')

```

---

## Performance Comparison

| Metric | WebLLM | Ollama |
|--------|--------|--------|
| **CDN Load** | 3-15 seconds | Instant (local) |
| **Model Init** | 1-5 minutes | <1 second |
| **First Response** | 30-60 seconds | 2-10 seconds |
| **Subsequent Responses** | 10-30 seconds | 2-10 seconds |
| **Token Speed (CPU)** | 2-5 tokens/s | 5-15 tokens/s |
| **Token Speed (GPU)** | 10-30 tokens/s | 20-100+ tokens/s |
| **Total Latency** | ~20 minutes to first response | ~10 seconds |

**Bottom line:** Ollama is **100x faster** for actual usage.

---

## Deployment Notes

### Local Development
- Open: `http://localhost:8000/index-app.html`
- Requires: `ollama serve` running
- Works: Instantly with local models

### GitHub Pages (Read-Only)
- URL: `https://snapkittywest.github.io/rowm-polymorphic-notebook/index-app.html`
- Browser blocks: Local Ollama (mixed HTTP/HTTPS content)
- Workaround: Deploy entire notebook including Ollama on a real server

### Future: Server-Side Integration
```javascript

// If Ollama deployed on server:

new AhmadJITEngine('https://your-domain.com:11434')

// Would work across the internet

```

---

## Live Testing Checklist

- [ ] Ollama installed: `ollama --version`
- [ ] Ollama model pulled: `ollama list`
- [ ] Ollama server running: `ollama serve`
- [ ] Web server running: `python3 -m http.server 8000`
- [ ] Page opens: `http://localhost:8000/index-app.html`
- [ ] Click "CONNECT TO OLLAMA" button
- [ ] Status changes to "READY" (green)
- [ ] Type message: "Hello, what is your name?"
- [ ] Click SEND
- [ ] Response streams in real-time in chat
- [ ] Check console (F12) for no errors
- [ ] Try another message

---

## Debug Commands

**Check Ollama running:**
```bash

curl http://localhost:11434/api/tags

```

**See available models:**
```bash

ollama list

```

**Pull additional model:**
```bash

ollama pull neural-chat

```

**Switch model in Ahmad Bot:**
```javascript

// In browser console:

window.ahmadEngine = new AhmadJITEngine();

await window.ahmadEngine.initialize('neural-chat');  // instead of 'tinyllama'

```

**Check page for errors:**
```javascript

// In browser console:

console.log('Ahmad Engine state:', window.ahmadEngine?.getState());

console.log('UI initialized:', typeof window.ahmadJITUI);

```

---

## Architecture Diagram

```

Browser                               Local Machine

═════════════════════════════════════════════════════════════



index-app.html                        Ollama Server

    ↓                                 ═════════════

ahmad-jit-ui.js    ←→ HTTP API        localhost:11434

    ↓                  /api/generate  

ahmad-jit-engine.js ←→ fetch()        tinyllama

    ↓                                 neural-chat

Chat interface                        mistral

    ↓                                 llama2

User messages                         (etc)

    ↓

Streaming responses

```

---

## FAQ

**Q: Can I use GitHub Pages with Ollama?**  
A: No, GitHub Pages is HTTPS, Ollama is HTTP. Browser blocks mixed content. Use local server for development.

**Q: Can I run Ollama on a server?**  
A: Yes, then use `new AhmadJITEngine('https://server.com:11434')` in the code.

**Q: Which model should I use?**  
A: Start with `tinyllama` for testing (440 MB, fastest). Then try `neural-chat` (3.8 GB, better quality).

**Q: How long to download a model?**  
A: Depends on connection speed. `tinyllama` is ~5 minutes on 10 Mbps connection.

**Q: Can I switch models after connecting?**  
A: Yes, reconnect with different model: `await window.ahmadEngine.initialize('neural-chat')`

**Q: What if Ollama crashes?**  
A: Restart: `ollama serve`. Ahmad Bot UI stays functional, just shows "connection error" when you try to send message.

---

**Last Updated:** July 27, 2026  
**Ollama API:** http://localhost:11434  
**Default Model:** tinyllama  
**Architecture:** Local HTTP API (instant, no CDN delays)