🤖 ResumeBot: a personalised LLM fine-tuned on my resume

ResumeBot is a chatbot that answers questions about Sudeep (Data Analyst | Data Science). It is Google's pretrained google/flan-t5-base model fine-tuned on 353 question–answer pairs built from my resume, combined with a semantic retriever and two hallucination guards so it only says things that are actually on my resume.

"What are Sudeep's technical skills?" → "Sudeep's technical skills are Python, SQL, Power BI, Excel, Data Analysis, Data Visualization and Business Intelligence." 📄 Resume section: Technical skills

How it works

Question
   ↓
Semantic retriever (all-MiniLM-L6-v2) finds the most relevant resume facts
   ↓
Out-of-scope guard: not about the resume?  →  polite refusal (no guessing)
   ↓
Fine-tuned FLAN-T5 writes the answer from those facts
   ↓
Grounding check: answer contains words not in the resume?  →  return the verified resume fact
   ↓
Answer + resume section (Gradio web app)

This approach is called retrieval-augmented fine-tuning: the model is fine-tuned to answer from retrieved resume text, which keeps answers accurate even for questions worded in new ways.

Training details

Base model google/flan-t5-base (pretrained by Google)
Method Supervised fine-tuning (seq2seq) with retrieved context
Knowledge base 28 resume facts
Dataset sudeep1610/resumebot-dataset: 353 Q&A pairs (327 train / 26 unseen test)
Epochs / learning rate / batch size 10 / 0.0003 / 8
Hardware Google Colab (free T4 GPU)
Retriever sentence-transformers/all-MiniLM-L6-v2

Results (on questions never seen during training)

Metric Value
Final training loss 0.0010
Loss on unseen questions 0.0056
Retriever picked the correct resume fact 69%
Answer exactly matches the resume answer 69%
Training loss

Hallucination guards

  1. Retrieval: the model only sees the resume facts relevant to the question.
  2. Grounding check: any answer containing information that is not in the retrieved resume text is replaced by the verified fact.
  3. Out-of-scope guard: questions unrelated to the resume (e.g. "What is the capital of France?") get a polite refusal.

Use it

Run the full ResumeBot web app (Gradio):

git clone https://huggingface.co/sudeep1610/ResumeBot
cd ResumeBot
pip install -r requirements.txt
python app.py

Use only the fine-tuned model in Python:

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tok = AutoTokenizer.from_pretrained("sudeep1610/ResumeBot")
model = AutoModelForSeq2SeqLM.from_pretrained("sudeep1610/ResumeBot")
prompt = ("Resume information: Sudeep's technical skills are Python, SQL, Power BI, Excel, Data Analysis, "
          "Data Visualization and Business Intelligence.\nQuestion: What tools does he know?\n"
          "Answer using only the resume information:")
out = model.generate(**tok(prompt, return_tensors="pt"), max_new_tokens=128, num_beams=4)
print(tok.decode(out[0], skip_special_tokens=True))

Files

File What it is
model.safetensors, config.json, tokenizer* The fine-tuned FLAN-T5 model
app.py Complete ResumeBot: retriever + fine-tuned model + guards + Gradio website
bot_config.json Knowledge base: resume facts, profile, prompt, known questions
requirements.txt Python packages needed to run app.py
ResumeBot_Pro.ipynb Training notebook (Google Colab)
loss_curve.png Training / test loss chart

Limitations

  • Knows only what is written in the resume; it cannot answer about anything else by design.
  • Small academic project: trained on a small dataset, so unusual phrasings may occasionally match the wrong section.

Author

Sudeep, Data Analyst | Data Science · M.Sc. Data Science, Dayananda Sagar University (2025-2027) LinkedIn · GitHub

Downloads last month
11
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sudeep1610/ResumeBot

Finetuned
(936)
this model

Dataset used to train sudeep1610/ResumeBot