llama-dialog-lora

A LoRA adapter for meta-llama/Llama-3.2-1B-Instruct that generates short two-speaker conversations ("podcast scripts") from a one-line topic instruction.

The base model tends to answer a podcast request with headings, segment titles, narration, or extra speakers, which a text-to-speech pipeline can't read directly. This adapter teaches the model to output a plain alternating A: / B: dialogue. Each line can then go straight to a TTS voice.

It was built as the local generation model for Podcast+, a project that turns documents and questions into two-host audio episodes.

Model Details

  • Developed by: Team PP (@Allenwang2004, @0u88)
  • Model type: LoRA adapter (PEFT) for a decoder-only causal language model
  • Language: English
  • License: Llama 3.2 Community License (inherited from the base model)
  • Finetuned from: meta-llama/Llama-3.2-1B-Instruct
  • Repository: https://github.com/Allenwang2004/Podcast-

How to Get Started

The adapter was trained on a plain-text prompt, not the Llama chat template. Use the same format at inference:

Instruction: <your instruction>
Answer:
import torch
from peft import PeftConfig, PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

ADAPTER = "coconut19/llama-dialog-lora"

config = PeftConfig.from_pretrained(ADAPTER)
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
base_model = AutoModelForCausalLM.from_pretrained(
    config.base_model_name_or_path,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model = PeftModel.from_pretrained(base_model, ADAPTER)
model.eval()

def generate_dialog(instruction, max_new_tokens=512):
    prompt = f"Instruction: {instruction}\nAnswer:\n"
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            do_sample=True,
            temperature=0.8,
            top_p=0.9,
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.eos_token_id,
        )
    text = tokenizer.decode(outputs[0], skip_special_tokens=True)
    return text[len(prompt):].strip()

print(generate_dialog(
    "Generate a podcast content between two people discussing their favorite class of computer science."
))

Example output:

A: I know what you're thinking, but I really want to go back to school for computer science.
B: You're right. I don't think I'm good at it. I've been out of school for a long time. But I know that I can learn it.
A: You're not going to get it if you don't try. You have to learn something new, no matter how hard you try.
B: I know, but I've been out of school for so long, I don't know where to start.
A: Don't worry. Start with something simple. It's like playing a new sport. You just have to start with the basics and work your way up.
...

Using retrieved context (RAG)

To ground the conversation in specific material, put the context before the instruction:

prompt = f"""
Use the information provided in the relevant context to generate the answer.
### Relevant Context:
{retrieved_context}
### Instruction:Instruction: {instruction}
 ### Answer:
"""

Uses

Direct Use

  • Generating short, casual two-person dialogues on a given topic
  • Producing scripts for a TTS pipeline, where each A: / B: line is sent to a different voice. In Podcast+ these lines were voiced with Kokoro-82M.

Out-of-Scope Use

  • Factual or educational content where accuracy matters. The model is small and was trained on everyday small talk, so it often drifts off-topic or states things that are wrong.
  • Long-form episodes. Training examples are short (typically 4 to 12 turns), and outputs are similar in length.
  • Languages other than English.
  • General chat or instruction following. Fine-tuning narrows the model toward one output format.

Bias, Risks, and Limitations

  • Topic drift: The training dialogues cover everyday topics (social life, sports, food, school, health). On technical topics the model often falls back to generic small talk. In a no-context test it answered "poisson distribution" with a conversation about summer evenings.
  • Hallucination: With or without retrieved context, the model can state plausible but incorrect facts (for example, contradicting itself about which of C and C++ is more widely used).
  • Truncation: Generation stops at max_new_tokens, so the last line may be cut off mid-sentence. Trim incomplete lines before sending text to TTS.
  • Inherited behavior: The adapter inherits the biases and limitations of Llama-3.2-1B-Instruct and of the DailyDialog corpus.

Training Details

Training Data

The first 3,000 dialogues from the DailyDialog training split (the dialog column of train.csv), a corpus of human-written, everyday English conversations.

Preprocessing:

  1. Topic labeling: Each dialogue was embedded with sentence-transformers/all-mpnet-base-v2 and given the label of the closest topic (by cosine similarity) out of sports, food, school, health and social.
  2. Speaker formatting: Utterances were parsed from the raw list string and labeled with alternating A: / B: speakers.
  3. Instruction pairs: Each dialogue became one example:
    {"instruction": "Generate a podcast content between two people discussing <topic>.", "input": "", "output": "A: ...\nB: ...\n..."}
    

Training Procedure

Each example was formatted as Instruction: {instruction}\nAnswer:\n{output}, tokenized to a fixed length of 512 tokens, and trained with a causal LM objective. Prompt tokens were masked from the loss (labels set to -100), so the model learns only the dialogue.

LoRA Configuration

Parameter Value
Rank (r) 16
lora_alpha 32
lora_dropout 0.05
Bias none
Target modules q_proj, k_proj, v_proj, o_proj
Trainable parameters 3,407,872 (about 0.28% of 1.24B)

Training Hyperparameters

Parameter Value
Training regime bf16 mixed precision
Epochs 1
Per-device batch size 2
Gradient accumulation steps 8 (effective batch size 16)
Total steps 188
Learning rate 2e-4
Warmup steps 100
Optimizer AdamW (adamw_torch)
Gradient checkpointing enabled
Max sequence length 512

Training Loss

Step Epoch Loss
50 0.27 2.801
100 0.53 0.560
150 0.80 0.541

Evaluation

Testing Data

  • Base vs. LoRA: 50 prompts of the form Generate a podcast content between two people discussing <topic>, with topics such as Technology. Both models generated at most 100 new tokens with sampling.
  • LoRA vs. LoRA + RAG: 10 prompts on technical topics (Poisson distribution, C++ classes, Newton's first law, pointers, and others). Each prompt was paired with one retrieved passage.

Metrics

Both metrics use sentence-transformers/all-mpnet-base-v2 embeddings.

  • Semantic consistency: The mean cosine similarity between consecutive lines of a generated dialogue. It measures how well each turn follows from the previous one.
  • Prompt-output alignment: The cosine similarity between the embedding of a fixed reference prompt ("Please summarize the following conversation.") and the full output. It roughly measures how much an output reads like a conversation. It does not measure relevance to the requested topic.

Results

Setting Prompts Semantic consistency Prompt-output alignment
Base model (Llama-3.2-1B-Instruct) 50 0.417 0.301
LoRA (this adapter) 50 0.581 0.386
LoRA, technical topics, no context 10 0.630 0.268
LoRA + RAG, technical topics 10 0.539 0.250

Summary

Compared to the base model, the adapter produces noticeably more coherent turn-by-turn dialogue and keeps to the A: / B: format that TTS needs. On technical topics, adding retrieved context makes the content more on-topic (the conversation actually discusses the requested concept), but these embedding metrics score it slightly lower. That's because the injected facts make consecutive turns less uniform. The 10-prompt RAG comparison is small, so treat it as indicative only.

Technical Specifications

Model Architecture and Objective

Low-rank adapters on the attention projections of Llama-3.2-1B-Instruct (16 layers, hidden size 2048). The objective is causal language modeling with the prompt tokens masked out.

Compute Infrastructure

Trained in Google Colab on a single GPU with bf16. One epoch takes 188 optimizer steps (total FLOs about 9.0e15).

Software

  • PEFT 0.18.0
  • Transformers
  • PyTorch
  • sentence-transformers (data labeling and evaluation)

Model Card Authors

@Allenwang2004, @0u88

Model Card Contact

Open an issue at https://github.com/Allenwang2004/Podcast-

Downloads last month
23
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coconut19/llama-dialog-lora

Adapter
(676)
this model

Dataset used to train coconut19/llama-dialog-lora