Specialized 1B Parameter Model for Computer Engineering
Fine-tuned with LoRA on 8-bit quantized Llama-3-1B
🛠️ Technical Specifications
Architecture
| Component | Specification |
|---|---|
| Base Model | Meta-Llama-3-1B |
| Hidden Size | 2048 |
| Layers | 16 |
| Attention Heads | 32 |
| Quantization | 8-bit via BitsAndBytes |
| Fine-Tuning Method | LoRA (Low-Rank Adaptation) |
| Tokenizer Vocabulary | 128,256 tokens |
Training Data
- Wikitext-2-raw-v1 (General knowledge)
- Custom computer engineering corpus:
- Hardware design principles
- Processor architectures
- Embedded systems documentation
Installation and Usage
This model works best as a completion-style Computer Engineering model rather than as a general-purpose chat assistant.
For best results, use the following prompt format:
Q: Explain the purpose of CPU cache memory.
A:
Recommended format:
Q: <computer engineering question>
A:
Option 1: From Hugging Face Hub
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="auto"
)
question = "Explain the purpose of CPU cache memory."
prompt = f"""Q: {question}
A:"""
inputs = tokenizer(
prompt,
return_tensors="pt"
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=200,
temperature=0.7,
top_p=0.9,
do_sample=True,
repetition_penalty=1.1,
pad_token_id=tokenizer.eos_token_id
)
generated_tokens = outputs[
0,
inputs["input_ids"].shape[1]:
]
response = tokenizer.decode(
generated_tokens,
skip_special_tokens=True
)
print(response)
Option 2: Local Installation
Git LFS is required if the model repository is cloned locally.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "./Llama-3.2-1B-ComputerEngineeringLLM"
tokenizer = AutoTokenizer.from_pretrained(
model_path,
local_files_only=True
)
model = AutoModelForCausalLM.from_pretrained(
model_path,
local_files_only=True,
device_map="auto",
dtype="auto"
)
question = "Explain the purpose of CPU cache memory."
prompt = f"""Q: {question}
A:"""
inputs = tokenizer(
prompt,
return_tensors="pt"
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=200,
temperature=0.7,
top_p=0.9,
do_sample=True,
repetition_penalty=1.1,
pad_token_id=tokenizer.eos_token_id
)
generated_tokens = outputs[
0,
inputs["input_ids"].shape[1]:
]
response = tokenizer.decode(
generated_tokens,
skip_special_tokens=True
)
print(response)
Recommended Generation Configuration
max_new_tokens = 200
temperature = 0.7
top_p = 0.9
do_sample = True
repetition_penalty = 1.1
Short, direct technical questions generally produce better results than conversational role prompts.
The model may occasionally continue generating additional questions or answers after completing its initial response. Reducing max_new_tokens can help control generation length.
For deterministic benchmarking, greedy decoding can be used, but it may not reflect the model's best interactive generation behavior.
License Compliance
This model is governed by the Llama 3.2 Community License.
Use, modification, and redistribution must comply with the terms of the Llama 3.2 Community License and the associated Acceptable Use Policy.
When redistributing Llama 3.2 materials or derivative works, applicable attribution and redistribution requirements from Meta must be preserved.
Attribution notice:
"Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved."
For the complete and authoritative license terms, refer to the Llama 3.2 Community License linked in this model card.
Limitations
Specialized for computer engineering (general performance may vary) Occasional repetition in outputs Requires prompt engineering for optimal results Knowledge cutoff: January 2025
Citation
If using for academic research, please cite:
@misc{llama3.2-1b-eng-2025,
title = {Llama-3.2-1B-Computer-Engineering-LLM},
author = {Irfanuruchi},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM},
}
- Downloads last month
- 40
Model tree for Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM
Base model
meta-llama/Llama-3.2-1B