Instructions to use txktxkabcd/fahqgpt-rag with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use txktxkabcd/fahqgpt-rag with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="txktxkabcd/fahqgpt-rag")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("txktxkabcd/fahqgpt-rag", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use txktxkabcd/fahqgpt-rag with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "txktxkabcd/fahqgpt-rag" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "txktxkabcd/fahqgpt-rag", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/txktxkabcd/fahqgpt-rag
- SGLang
How to use txktxkabcd/fahqgpt-rag with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "txktxkabcd/fahqgpt-rag" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "txktxkabcd/fahqgpt-rag", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "txktxkabcd/fahqgpt-rag" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "txktxkabcd/fahqgpt-rag", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use txktxkabcd/fahqgpt-rag with Docker Model Runner:
docker model run hf.co/txktxkabcd/fahqgpt-rag
FahQgpt RAG —— 让 146M 小模型"能正常对话"的三层架构
这是让 FahQgpt 1.0 Nano(145.9M 参数,从零训练)能可靠回答问题的完整方案。 裸模型单独用只有约 23% 的日常问题能答对;加上这套架构后, 测试集里 16/16 全对。
配套模型:**txktxkabcd/fahqgpt-1.0-nano**
一、核心思想:别逼小模型"记住",让它"抄"
小模型(<200M)的能力瓶颈是容量,不是训练方法:
- 逼它把知识记进参数 → 装不下。实测塞进算术就挤掉身份识别能力
- 换思路:知识放外部,模型只负责照抄 → 零容量成本
实测对比(同一批 23 个日常问题):
| 问题 | 裸模型 | 加 RAG 后 |
|---|---|---|
| Who are you? | I am a computer scientist. |
正确身份 |
| What is 1+1? | The answer is 1 |
1 + 1 = 2 |
| What is 7+5? | 7+5 = 7 |
7 + 5 = 12 |
| What is the capital of Japan? | 答不出 | Tokyo |
| Count from one to five. | A to Z |
one, two, three, four, five. |
| Can you help me? | No, I cannot |
Yes, I'll do my best to help you. |
**正确率 23% → 接近 100%**(就测的这批而言)
二、三层架构
用户提问
│
├─ 1. 计算器 ──────── 正则识别算术 → 直接计算 → 返回
│ (语言模型本来就不该做算术,连大模型都会错)
│
├─ 2. 计数器 ──────── 识别 "count from one to five" 这类 → 直接生成
│
├─ 3. 知识库检索 ──── 字符 n-gram 相似度,从 19,141 条问答里找最像的
│ |相似度 > 0.20 → 直接返回标准答案
│ |否则 → 把 top-3 当 few-shot 范例喂给模型
│
└─ 4. 模型兜底 ────── llama.cpp 采样生成(这一层会胡说,但格式对)
为什么检索用字符 n-gram 而不是向量检索
- 不需要任何 embedding 模型 —— 零下载、零额外显存
- 对"问法稍有不同"匹配效果足够好
- 纯 Python,
math+collections就能跑,没有重依赖
实测命中率 11/12(含训练数据里没有的新问题)。
为什么算术必须单独做
检索只能返回"知识库里出现过的题"。实测 What is 7+5? 检索到了
相似度 0.22 的错误条目(回到 2),因为库里没有 7+5。
算术是确定性问题,用计算器算才对。
三、详细配置教学(Windows / macOS / Linux)
第 0 步:确认你有这些
| 需要 | 说明 |
|---|---|
| 操作系统 | Windows 10/11、macOS、Linux 都行 |
| 内存 | 2 GB 以上(模型只有 137~350 MB,很轻) |
| 硬盘 | 约 1 GB 空间 |
| 显卡 | 不需要!CPU 就能跑,只是慢一点 |
第 1 步:安装 Python
Windows:去 python.org 下载 Python 3.10+ 安装时**务必勾选 "Add Python to PATH"**(这一步漏了后面全失败)
macOS:brew install python3
Linux:sudo apt install python3 python3-pip(Debian/Ubuntu)
验证:
python --version # 应显示 3.10 以上
第 2 步:安装 llama.cpp(提供推理后端)
这是必须的 —— RAG 服务要靠它跑模型。
方式 A:下载预编译版(推荐,最简单)
去 llama.cpp releases 下载对应平台包:
| 平台 | 下载哪个 |
|---|---|
| Windows(NVIDIA 显卡) | llama-bXXXX-bin-win-cuda-x64.zip |
| Windows(无独显/AMD) | llama-bXXXX-bin-win-cpu-x64.zip |
| macOS | llama-bXXXX-bin-macos-arm64.zip |
| Linux | llama-bXXXX-bin-ubuntu-x64.zip |
解压到一个固定目录,例如 D:\llama.cpp。
⚠️ 国内下载 GitHub 慢:可以用 ghproxy 加速, 或直接下 CPU 版(体积小很多)。
方式 B:用包管理器
# macOS
brew install llama.cpp
# Linux(部分发行版)
sudo apt install llama-cpp
第 3 步:下载模型
从 txktxkabcd/fahqgpt-1.0-nano 下载:
| 文件 | 大小 | 何时用 |
|---|---|---|
fahqgpt-1.0-nano-final-f16.gguf |
350 MB | 推荐,质量最好 |
fahqgpt-1.0-nano-final-Q4_K_M.gguf |
137 MB | 想省空间/内存小时用 |
国内下载方案(huggingface.co 直连不稳时):
# 用镜像站
export HF_ENDPOINT=https://hf-mirror.com # Windows 用 set HF_ENDPOINT=...
pip install -U huggingface_hub
huggingface-cli download txktxkabcd/fahqgpt-1.0-nano \
fahqgpt-1.0-nano-final-f16.gguf --local-dir ./models
或者直接浏览器打开 https://hf-mirror.com/txktxkabcd/fahqgpt-1.0-nano 手动下载。
第 4 步:下载这个仓库的文件
git clone https://huggingface.co/txktxkabcd/fahqgpt-rag
cd fahqgpt-rag
或者网页上逐个下载。
第 5 步:安装 Python 依赖
pip install requests
就这一个。不需要 torch、transformers 之类的重型库。
第 6 步:改一行配置(指向你的 llama-server)
打开 rag_chat.py,找到这一行(在 --serve 分支里):
llama = subprocess.Popen(
[r"D:\llama.cpp\llama-server.exe", # ← 改成你自己的路径
"-m", str(gguf.relative_to(PROJECT)).replace("\\", "/"),
...
同时把模型文件放在 rag_chat.py 同目录(或改 gguf 那行的路径)。
💡 更省事的办法:把
llama-server加到系统 PATH,然后把这里改成["llama-server", ...],就不用写绝对路径了。
第 7 步:启动!
python rag_chat.py --serve --port 8080 --direct-threshold 0.20
看到这样的输出就成功了:
知识库:19141 条
RAG 服务已启动:http://127.0.0.1:8080
浏览器打开 http://127.0.0.1:8080 —— 界面里就能聊天了。
第 8 步(可选):让局域网内其它设备访问
把 rag_chat.py 里最后的这一行:
srv = ThreadingHTTPServer(("0.0.0.0", args.port), H)
确认是 0.0.0.0(默认就是)。然后在别的设备上访问
http://<你这台机器的内网IP>:8080
查询内网 IP:
# Windows
ipconfig
# macOS / Linux
ifconfig | grep "inet "
Windows 还需要放行防火墙(首次运行时会弹窗,点"允许")。
四、常见问题
Q:报 ModuleNotFoundError: No module named 'requests'
→ 没装依赖,跑 pip install requests
Q:报找不到 llama-server
→ 第 6 步的路径没改对。确认 llama-server.exe(Windows)或
llama-server(macOS/Linux)的真实路径
Q:端口被占用 Address already in use
→ 换端口:--port 8081;或找占用进程:
# Windows
netstat -ano | findstr :8080
# macOS / Linux
lsof -i :8080
Q:回答很短 → 这是模型特性(训练数据里答案平均 5 个词)。点回答下方的「继续」按钮 可以让它接着写。
Q:回答开始复读(esophagus. esophagus...)
→ 采样惩罚没生效。检查 rag_chat.py 的 ask_model() 里有没有
repeat_penalty / dry_multiplier 这些参数。
Q:知识库能换吗
→ 能。final_sft_balanced.jsonl 就是知识库,格式是每行一个
{"user": "问题", "assistant": "答案"}。换成你自己的数据即可。
五、⚠️ 手机上能跑吗?
不能直接跑这套 RAG 服务。 三个原因:
- 需要 Python 进程(检索 + HTTP 服务),手机没有现成运行环境
- 必须另起
llama-server作为推理后端,这是 PC 端程序 - 每次提问要遍历 19,141 条做 n-gram 匹配,纯 Python 在手机上会明显卡顿
手机上可用的两个替代方案
方案 A:手机连 PC(推荐)
按上面第 8 步配好局域网访问,手机浏览器打开 http://<PC内网IP>:8080
即可获得完整的 RAG 体验。
方案 B:手机只跑模型(放弃检索)
用支持 GGUF 的手机 App 直接加载 fahqgpt-1.0-nano-final-Q4_K_M.gguf(137 MB):
| 平台 | 推荐 App |
|---|---|
| Android | ChatterUI、PocketPal、MLC Chat |
| iOS | PocketPal、Private LLM |
此时没有检索层,体验就是裸模型的水平(约 23% 问题能答对, 其余会一本正经地胡说)。优点是零部署、完全离线。
如果一定要手机离线 + 检索:理论上可以把知识库精简到 200~500 条, 做成静态 few-shot 提示词塞进 App 的系统提示里。但 146M 模型的 上下文只有 1024 token,只能装约 15 组问答,覆盖率极低,意义不大。
六、诚实的局限
- 知识库外的问题:仍会胡说(模型生成层的能力上限)
- 事实性错误:知识库来自 SmolTalk + DeepSeek 合成数据,未经人工逐条核验, 里面就有错(例如"狮子是最大的动物")
- "全对"是就测试的 16 条而言,不等于所有问题都对
- 长回答:模型被训成"说完就停",要继续写只能点「继续」按钮
- 采样参数敏感:长输出必须配 DRY + frequency/presence penalty, 否则会陷入重复循环
七、文件说明
| 文件 | 作用 |
|---|---|
rag_chat.py |
主程序:检索 + 计算器 + HTTP 服务 |
chat_ui.html |
网页聊天界面(含「继续」按钮、长度控制) |
final_sft_balanced.jsonl |
知识库 + 训练数据(27,545 条,2.2 MB) |
tools/build_final_sft.py |
合并多个数据源(含同问不同答的冲突检测) |
tools/gen_synthetic.py |
用 DeepSeek API 生成合成对话数据 |
tools/make_math_data.py |
生成算术数据(纯数字答案,防模型抄题) |
tools/make_identity_en.py |
生成纯英文身份数据 |
八、模型信息
FahQgpt 1.0 Nano —— 由 B站@一片烂海苔 从零训练
| 项目 | 数值 |
|---|---|
| 参数量 | 145.9M |
| 架构 | LLaMA 式(RMSNorm + RoPE + GQA 14Q/2KV + SwiGLU + 权重绑定) |
| 训练数据 | 19 亿 token |
| 训练步数 | 14,495 |
| 硬件 | 单张 RTX 5060 Ti 16GB |
| 耗时 | 13.45 小时 |
预训练语料:FineWeb-Edu (83.4%) / Cosmopedia (6.8%) / OpenWebMath (6.4%) / Wikipedia (3.2%),全部来自 HuggingFace 公开数据集。
对话能力通过「继续预训练(CPT,SmolTalk 33 万段对话)+ 混合损失 SFT」获得。 关键技术点:对话格式用纯文本而非特殊 token,且必须混通用文本一起训 (否则输出分布塌缩)。
MIT License