Instructions to use coconut19/llama-dialog-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use coconut19/llama-dialog-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-1B-Instruct") model = PeftModel.from_pretrained(base_model, "coconut19/llama-dialog-lora") - Transformers
How to use coconut19/llama-dialog-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="coconut19/llama-dialog-lora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("coconut19/llama-dialog-lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use coconut19/llama-dialog-lora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "coconut19/llama-dialog-lora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coconut19/llama-dialog-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/coconut19/llama-dialog-lora
- SGLang
How to use coconut19/llama-dialog-lora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "coconut19/llama-dialog-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coconut19/llama-dialog-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "coconut19/llama-dialog-lora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coconut19/llama-dialog-lora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use coconut19/llama-dialog-lora with Docker Model Runner:
docker model run hf.co/coconut19/llama-dialog-lora
llama-dialog-lora
A LoRA adapter for meta-llama/Llama-3.2-1B-Instruct that generates short two-speaker conversations ("podcast scripts") from a one-line topic instruction.
The base model tends to answer a podcast request with headings, segment titles, narration, or extra speakers, which a text-to-speech pipeline can't read directly. This adapter teaches the model to output a plain alternating A: / B: dialogue. Each line can then go straight to a TTS voice.
It was built as the local generation model for Podcast+, a project that turns documents and questions into two-host audio episodes.
Model Details
- Developed by: Team PP (@Allenwang2004, @0u88)
- Model type: LoRA adapter (PEFT) for a decoder-only causal language model
- Language: English
- License: Llama 3.2 Community License (inherited from the base model)
- Finetuned from:
meta-llama/Llama-3.2-1B-Instruct - Repository: https://github.com/Allenwang2004/Podcast-
How to Get Started
The adapter was trained on a plain-text prompt, not the Llama chat template. Use the same format at inference:
Instruction: <your instruction>
Answer:
import torch
from peft import PeftConfig, PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
ADAPTER = "coconut19/llama-dialog-lora"
config = PeftConfig.from_pretrained(ADAPTER)
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
base_model = AutoModelForCausalLM.from_pretrained(
config.base_model_name_or_path,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, ADAPTER)
model.eval()
def generate_dialog(instruction, max_new_tokens=512):
prompt = f"Instruction: {instruction}\nAnswer:\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=True,
temperature=0.8,
top_p=0.9,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
)
text = tokenizer.decode(outputs[0], skip_special_tokens=True)
return text[len(prompt):].strip()
print(generate_dialog(
"Generate a podcast content between two people discussing their favorite class of computer science."
))
Example output:
A: I know what you're thinking, but I really want to go back to school for computer science.
B: You're right. I don't think I'm good at it. I've been out of school for a long time. But I know that I can learn it.
A: You're not going to get it if you don't try. You have to learn something new, no matter how hard you try.
B: I know, but I've been out of school for so long, I don't know where to start.
A: Don't worry. Start with something simple. It's like playing a new sport. You just have to start with the basics and work your way up.
...
Using retrieved context (RAG)
To ground the conversation in specific material, put the context before the instruction:
prompt = f"""
Use the information provided in the relevant context to generate the answer.
### Relevant Context:
{retrieved_context}
### Instruction:Instruction: {instruction}
### Answer:
"""
Uses
Direct Use
- Generating short, casual two-person dialogues on a given topic
- Producing scripts for a TTS pipeline, where each
A:/B:line is sent to a different voice. In Podcast+ these lines were voiced with Kokoro-82M.
Out-of-Scope Use
- Factual or educational content where accuracy matters. The model is small and was trained on everyday small talk, so it often drifts off-topic or states things that are wrong.
- Long-form episodes. Training examples are short (typically 4 to 12 turns), and outputs are similar in length.
- Languages other than English.
- General chat or instruction following. Fine-tuning narrows the model toward one output format.
Bias, Risks, and Limitations
- Topic drift: The training dialogues cover everyday topics (social life, sports, food, school, health). On technical topics the model often falls back to generic small talk. In a no-context test it answered "poisson distribution" with a conversation about summer evenings.
- Hallucination: With or without retrieved context, the model can state plausible but incorrect facts (for example, contradicting itself about which of C and C++ is more widely used).
- Truncation: Generation stops at
max_new_tokens, so the last line may be cut off mid-sentence. Trim incomplete lines before sending text to TTS. - Inherited behavior: The adapter inherits the biases and limitations of Llama-3.2-1B-Instruct and of the DailyDialog corpus.
Training Details
Training Data
The first 3,000 dialogues from the DailyDialog training split (the dialog column of train.csv), a corpus of human-written, everyday English conversations.
Preprocessing:
- Topic labeling: Each dialogue was embedded with
sentence-transformers/all-mpnet-base-v2and given the label of the closest topic (by cosine similarity) out ofsports,food,school,healthandsocial. - Speaker formatting: Utterances were parsed from the raw list string and labeled with alternating
A:/B:speakers. - Instruction pairs: Each dialogue became one example:
{"instruction": "Generate a podcast content between two people discussing <topic>.", "input": "", "output": "A: ...\nB: ...\n..."}
Training Procedure
Each example was formatted as Instruction: {instruction}\nAnswer:\n{output}, tokenized to a fixed length of 512 tokens, and trained with a causal LM objective. Prompt tokens were masked from the loss (labels set to -100), so the model learns only the dialogue.
LoRA Configuration
| Parameter | Value |
|---|---|
Rank (r) |
16 |
lora_alpha |
32 |
lora_dropout |
0.05 |
| Bias | none |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Trainable parameters | 3,407,872 (about 0.28% of 1.24B) |
Training Hyperparameters
| Parameter | Value |
|---|---|
| Training regime | bf16 mixed precision |
| Epochs | 1 |
| Per-device batch size | 2 |
| Gradient accumulation steps | 8 (effective batch size 16) |
| Total steps | 188 |
| Learning rate | 2e-4 |
| Warmup steps | 100 |
| Optimizer | AdamW (adamw_torch) |
| Gradient checkpointing | enabled |
| Max sequence length | 512 |
Training Loss
| Step | Epoch | Loss |
|---|---|---|
| 50 | 0.27 | 2.801 |
| 100 | 0.53 | 0.560 |
| 150 | 0.80 | 0.541 |
Evaluation
Testing Data
- Base vs. LoRA: 50 prompts of the form
Generate a podcast content between two people discussing <topic>, with topics such as Technology. Both models generated at most 100 new tokens with sampling. - LoRA vs. LoRA + RAG: 10 prompts on technical topics (Poisson distribution, C++ classes, Newton's first law, pointers, and others). Each prompt was paired with one retrieved passage.
Metrics
Both metrics use sentence-transformers/all-mpnet-base-v2 embeddings.
- Semantic consistency: The mean cosine similarity between consecutive lines of a generated dialogue. It measures how well each turn follows from the previous one.
- Prompt-output alignment: The cosine similarity between the embedding of a fixed reference prompt (
"Please summarize the following conversation.") and the full output. It roughly measures how much an output reads like a conversation. It does not measure relevance to the requested topic.
Results
| Setting | Prompts | Semantic consistency | Prompt-output alignment |
|---|---|---|---|
| Base model (Llama-3.2-1B-Instruct) | 50 | 0.417 | 0.301 |
| LoRA (this adapter) | 50 | 0.581 | 0.386 |
| LoRA, technical topics, no context | 10 | 0.630 | 0.268 |
| LoRA + RAG, technical topics | 10 | 0.539 | 0.250 |
Summary
Compared to the base model, the adapter produces noticeably more coherent turn-by-turn dialogue and keeps to the A: / B: format that TTS needs. On technical topics, adding retrieved context makes the content more on-topic (the conversation actually discusses the requested concept), but these embedding metrics score it slightly lower. That's because the injected facts make consecutive turns less uniform. The 10-prompt RAG comparison is small, so treat it as indicative only.
Technical Specifications
Model Architecture and Objective
Low-rank adapters on the attention projections of Llama-3.2-1B-Instruct (16 layers, hidden size 2048). The objective is causal language modeling with the prompt tokens masked out.
Compute Infrastructure
Trained in Google Colab on a single GPU with bf16. One epoch takes 188 optimizer steps (total FLOs about 9.0e15).
Software
- PEFT 0.18.0
- Transformers
- PyTorch
- sentence-transformers (data labeling and evaluation)
Model Card Authors
Model Card Contact
Open an issue at https://github.com/Allenwang2004/Podcast-
- Downloads last month
- 23
Model tree for coconut19/llama-dialog-lora
Base model
meta-llama/Llama-3.2-1B-Instruct