--- language: - es license: apache-2.0 base_model: unsloth/Qwen3.5-0.8B datasets: - FarmifAI/FarmifAI_dataset_1.3 library_name: transformers pipeline_tag: text-generation tags: - agriculture - colombia - rag - small-language-model - on-device - lora - unsloth - qwen3_5 --- # FarmifAI 1.3 **FarmifAI 1.3** is a small language model (fine-tuned from [Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B)) that answers agricultural questions in Spanish **from a context you provide**. It is built for Colombian agriculture and powers FarmifAI, an offline assistant app for farmers: given a question and technical passages retrieved from a knowledge base, it writes a short reasoning and a clear final answer based only on that context. This repository contains the full-precision weights. For llama.cpp and on-device use, see **[FarmifAI_1.3_GGUF](https://huggingface.co/FarmifAI/FarmifAI_1.3_GGUF)**. > FarmifAI is not a general-purpose chatbot. It is trained to answer from the context it receives, so it should always be used together with a retrieval step. ## Model details | | | |---|---| | **Developed by** | FarmifAI, Universidad del Cauca (Colombia) | | **Base model** | Qwen3.5-0.8B | | **Language** | Spanish | | **Input** | System prompt + retrieved context in `` tags + question | | **Output** | `…` followed by `…` | | **License** | Apache 2.0 | ## Intended use - Answering farmers' questions in Spanish inside a RAG pipeline, using passages from agricultural technical documents. - Not meant to be used without retrieved context, in other languages, or as the only basis for decisions such as agrochemical selection or dosing. ## Prompt format The model was trained with a fixed Spanish system prompt. Use it as is: ```text Eres un asistente agrícola. Responde únicamente con la información dentro de . Si la respuesta no está en el contexto, declara que no tienes información; no inventes datos. Instrucciones de formato: - En , analiza paso a paso el contexto frente a la ... <<< PASTE THE REST OF THE EXACT SYSTEM PROMPT FROM THE DATASET HERE >>> ``` The user message contains the retrieved context followed by the question: ```text {retrieved context} {question} ``` The model replies with a step-by-step analysis and a final answer: ```text Step-by-step analysis of the context against the question. Final answer in Spanish, based only on the provided context. ``` Applications typically show only the `` block. Parse it defensively in case the tags are missing. ## Quickstart ```python import re from transformers import AutoProcessor, AutoModelForMultimodalLM model_id = "FarmifAI/FarmifAI_1.3" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForMultimodalLM.from_pretrained(model_id, device_map="auto", dtype="auto") SYSTEM_PROMPT = "..." # the system prompt from "Prompt format" above context = "..." # passages retrieved from your knowledge base question = "¿Cómo puedo controlar la broca en mi cultivo de café?" user_message = f"\n{context}\n\n\n{question}" messages = [ {"role": "system", "content": [{"type": "text", "text": SYSTEM_PROMPT}]}, {"role": "user", "content": [{"type": "text", "text": user_message}]}, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.1, top_p=0.9) text = processor.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True) match = re.search(r"(.*?)", text, re.DOTALL) print(match.group(1).strip() if match else text) ``` **Recommended settings:** `temperature=0.1`, `top_p=0.9`, `max_new_tokens=512`. ## Training Fine-tuned with LoRA (Unsloth + TRL) on [FarmifAI_dataset_1.3](https://huggingface.co/datasets/FarmifAI/FarmifAI_dataset_1.3), about 8.8k synthetic Spanish conversations built from technical documents on Colombian agriculture. Each example pairs a context passage and a question with a reasoning + answer response. ## Results Compared with the base Qwen3.5-0.8B on 250 held-out examples, using the same prompt and settings for both models: | Metric | Base model | FarmifAI 1.3 | |---|---|---| | Format adherence ↑ | 24.4% | **98.0%** | | Answer relevancy (LLM judge, 0–5) ↑ | 2.16 | **4.70** | | Faithfulness to context (LLM judge, 0–5) ↑ | 2.13 | **3.48** | | Contradictions with context (NLI) ↓ | 30.4% | **17.2%** | ## Limitations - The model can still make mistakes or add details that are not in the context. Check its recommendations, especially anything about agrochemicals, doses or safety periods. - Answer quality depends on the quality of the retrieved context. - It was evaluated with automatic metrics and LLM judges, not by agronomists or in the field, and it is not a substitute for professional advice. Its training material may reflect outdated practices, so do not use it for current regulations or registered products. ## Links - Quantized versions: [FarmifAI/FarmifAI_1.3_GGUF](https://huggingface.co/FarmifAI/FarmifAI_1.3_GGUF) - Dataset: [FarmifAI/FarmifAI_dataset_1.3](https://huggingface.co/datasets/FarmifAI/FarmifAI_dataset_1.3) *Fine-tuned with [Unsloth](https://github.com/unslothai/unsloth).*