Specialized 1B Parameter Model for Computer Engineering
Fine-tuned with LoRA on 8-bit quantized Llama-3-1B


🛠️ Technical Specifications

Architecture

Component Specification
Base Model Meta-Llama-3-1B
Hidden Size 2048
Layers 16
Attention Heads 32
Quantization 8-bit via BitsAndBytes
Fine-Tuning Method LoRA (Low-Rank Adaptation)
Tokenizer Vocabulary 128,256 tokens

Training Data

  • Wikitext-2-raw-v1 (General knowledge)
  • Custom computer engineering corpus:
    • Hardware design principles
    • Processor architectures
    • Embedded systems documentation

Installation and Usage

This model works best as a completion-style Computer Engineering model rather than as a general-purpose chat assistant.

For best results, use the following prompt format:

Q: Explain the purpose of CPU cache memory.
A:

Recommended format:

Q: <computer engineering question>
A:

Option 1: From Hugging Face Hub

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="auto"
)

question = "Explain the purpose of CPU cache memory."

prompt = f"""Q: {question}
A:"""

inputs = tokenizer(
    prompt,
    return_tensors="pt"
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=200,
        temperature=0.7,
        top_p=0.9,
        do_sample=True,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id
    )

generated_tokens = outputs[
    0,
    inputs["input_ids"].shape[1]:
]

response = tokenizer.decode(
    generated_tokens,
    skip_special_tokens=True
)

print(response)

Option 2: Local Installation

Git LFS is required if the model repository is cloned locally.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "./Llama-3.2-1B-ComputerEngineeringLLM"

tokenizer = AutoTokenizer.from_pretrained(
    model_path,
    local_files_only=True
)

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    local_files_only=True,
    device_map="auto",
    dtype="auto"
)

question = "Explain the purpose of CPU cache memory."

prompt = f"""Q: {question}
A:"""

inputs = tokenizer(
    prompt,
    return_tensors="pt"
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=200,
        temperature=0.7,
        top_p=0.9,
        do_sample=True,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id
    )

generated_tokens = outputs[
    0,
    inputs["input_ids"].shape[1]:
]

response = tokenizer.decode(
    generated_tokens,
    skip_special_tokens=True
)

print(response)

Recommended Generation Configuration

max_new_tokens = 200
temperature = 0.7
top_p = 0.9
do_sample = True
repetition_penalty = 1.1

Short, direct technical questions generally produce better results than conversational role prompts.

The model may occasionally continue generating additional questions or answers after completing its initial response. Reducing max_new_tokens can help control generation length.

For deterministic benchmarking, greedy decoding can be used, but it may not reflect the model's best interactive generation behavior.


License Compliance

This model is governed by the Llama 3.2 Community License.

Use, modification, and redistribution must comply with the terms of the Llama 3.2 Community License and the associated Acceptable Use Policy.

When redistributing Llama 3.2 materials or derivative works, applicable attribution and redistribution requirements from Meta must be preserved.

Attribution notice:

"Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved."

For the complete and authoritative license terms, refer to the Llama 3.2 Community License linked in this model card.


Limitations

Specialized for computer engineering (general performance may vary) Occasional repetition in outputs Requires prompt engineering for optimal results Knowledge cutoff: January 2025


Citation

If using for academic research, please cite:

@misc{llama3.2-1b-eng-2025,
  title = {Llama-3.2-1B-Computer-Engineering-LLM},
  author = {Irfanuruchi},
  year = {2025},
  publisher = {Hugging Face},
  url = {https://huggingface.co/Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM},
}
Downloads last month
40
Safetensors
Model size
1B params
Tensor type
F32
·
F16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Irfanuruchi/Llama-3.2-1B-Computer-Engineering-LLM

Adapter
(751)
this model
Finetunes
1 model
Quantizations
2 models