hharsha's picture
Fix usage for transformers without text2text-generation pipeline task
29a3108 verified
|
Raw History Blame Contribute Delete
2.74 kB
---
language:
- en
license: apache-2.0
library_name: transformers
base_model: google-t5/t5-small
tags:
- peft
- t5
- agents
- rag
- llmops
- lora
- text2text-generation
datasets:
- hharsha/agentic-github-meta
pipeline_tag: text2text-generation
widget:
- text: multi-agent platform with RAG, MCP, and observability
- text: looking for RAG hybrid search recall@k and reranking
- text: FastAPI Next.js agent backend with Docker Compose
---
# agentic-github-tagger
Lightweight **text2text tag generator** for agentic AI / RAG / LLMOps GitHub-style
descriptions. Fine-tuned from [`google-t5/t5-small`](https://huggingface.co/google-t5/t5-small)
with **PEFT LoRA** (r=16, alpha=32, dropout=0.05, target_modules `q`,`v`) on
[`hharsha/agentic-github-meta`](https://huggingface.co/datasets/hharsha/agentic-github-meta),
then **merged** so full small weights load on free CPU.
> ~60M-param T5-small tagger — **not** a 7B chat demo. Free Hub + CPU friendly.
## Usage
```python
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_id = "hharsha/agentic-github-tagger"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
text = "multi-agent platform with RAG, MCP, and observability"
ids = tok(text, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, num_beams=4)
print(tok.decode(out[0], skip_special_tokens=True))
```
On older `transformers` that still register the task, this also works:
```python
from transformers import pipeline
pipe = pipeline("text2text-generation", model="hharsha/agentic-github-tagger")
print(pipe("multi-agent platform with RAG, MCP, and observability")[0]["generated_text"])
```
Sample output from this training run:
```
multi-agent, multi-agent, observability, rag, mCP, observability
```
## Training
| | |
|---|---|
| Base | `google-t5/t5-small` |
| Method | PEFT LoRA then merge |
| r / alpha / dropout | 16 / 32 / 0.05 |
| target_modules | q, v |
| Epochs | 3 (CPU) |
| Batch size | 8 |
| Dataset | [`hharsha/agentic-github-meta`](https://huggingface.co/datasets/hharsha/agentic-github-meta) (687 rows; 600 used for train) |
## Links
- Dataset: [`hharsha/agentic-github-meta`](https://huggingface.co/datasets/hharsha/agentic-github-meta)
- Showcase: [`hharsha/agentic-systems-showcase`](https://huggingface.co/datasets/hharsha/agentic-systems-showcase)
- Studio: [https://agentic-systems-studio.com](https://agentic-systems-studio.com)
- GitHub: [https://github.com/hharsha98](https://github.com/hharsha98)
## Intended use / limits
Auto-suggest comma-separated tags for agentic / RAG / LLMOps project listings.
Small model; tags can repeat or be incomplete. Not for safety-critical labeling.