Pluto-Lite-v1 / README.md
SanaeNya's picture
Update README.md
0329151 verified
|
Raw History Blame Contribute Delete
2.21 kB
---
model_name: PlutoLite-v1
license: mit
language:
- en
- zh
base_model:
- google/gemma-4-E2B
tags:
- gemma4
- QLoRA
---
# PlutoLite-v1
Author : **KochiyaSanaeNya**
PlutoLite-v1 是基于 [google/gemma-4-E2B](https://huggingface.co/google/gemma-4-E2B) 的 Python 代码生成 / 函数补全 QLoRA 适配器。
发布内容为 PEFT LoRA 权重,需要另外加载基础模型。基础模型固定版本为 `d29ff6b45f081a49ee2733a859c9c9c2d95d1a6f`。适配器已包含之前各阶段累计更新,使用时只需加载适配器。
## 加载与推理
Python 3.11+;先安装适合本机 GPU 的 CUDA PyTorch,再进入本目录安装依赖:
```bash
pip install -r requirements.txt
python load_adapter.py --prompt "Write a Python function that returns whether an integer is even."
```
脚本默认从 Hugging Face 加载固定版本基础模型,并从脚本所在目录加载适配器;需要支持 BF16 的 NVIDIA GPU,使用 NF4 量化并将大型逐层词嵌入放在 CPU。也支持本地基础模型:
```bash
python load_adapter.py --base-model ../../base/gemma-4-E2B --local-files-only --prompt "Write a Python function that adds two integers."
```
使用其他 Hugging Face 适配器仓库路径时可传 `--adapter 用户名/PlutoLite-v1`。`--max-new-tokens` 默认为 256。
训练采用 **BOS + 原始 prompt** 的 completion 格式,instruction 为空;函数补全时保留原始换行和缩进,不添加聊天模板或 `Instruction/Response` 包装。脚本仅打印新生成文本,不执行生成的代码。
本适配器作用于 `Gemma4ForCausalLM` 文本模型。基础仓库使用多模态权重命名,因此加载脚本显式将 `model.language_model.` 映射为 `model.` 并检查文本权重完整性。它不是多模态微调产物。
## 来源与许可
基础模型由 Google 提供,其随附模型卡声明 Apache-2.0,参见 [基础模型卡](https://huggingface.co/google/gemma-4-E2B) 和 [基础模型许可](https://ai.google.dev/gemma/docs/gemma_4_license)。上述数据集的来源及条款以各自数据集卡为准。本包不包含基础权重或训练数据。此适配器尚未单独指定分发许可,未自动继承基础模型许可。