Numera: Handwritten Digit Classifier (CNN)
A compact convolutional neural network trained from scratch on MNIST to classify handwritten digits (0β9). It reaches 99.43% test accuracy with only 160,842 parameters (643 KB), and it powers Numera, a web app where you draw a digit and see the prediction in real time.
| Test accuracy | Parameters | Weights file | CPU inference |
|---|---|---|---|
| 99.43% | 160,842 | 643 KB | ~2 ms / image |
Model Details
Model Description
This model is a CNN trained from scratch on the MNIST benchmark dataset. It accepts 28Γ28 grayscale images of handwritten digits and outputs logits over 10 classes (digits 0β9). It is the backbone of the Numera web app.
- Developed by: Abdul Rafay
- Model type: Convolutional Neural Network (CNN)
- Input: 1 Γ 28 Γ 28 grayscale image, normalized with mean 0.1307 and std 0.3081
- Output: 10 logits (apply softmax for probabilities)
- License: MIT
- Framework: PyTorch (tested with 2.1.2)
- Finetuned from: Trained from scratch (no pretrained base)
Model Sources
- Demo: Numera on Hugging Face Spaces (free Space: it may take a moment to wake up)
- Code: github.com/abdurafay19/Digit-Classifier (FastAPI backend, Next.js frontend, Docker)
- Training notebook: MNIST_Training.ipynb
Uses
Direct Use
This model can be used directly to classify 28Γ28 grayscale images of handwritten digits, with no fine-tuning required. It is best suited for:
- Educational demos of deep learning and CNNs
- Handwritten digit recognition in controlled environments
- Integration into apps via the provided web UI or API
Downstream Use
The model can be fine-tuned or adapted for:
- Multi-digit number recognition (e.g., street numbers, forms)
- Similar single-character classification tasks
- A transfer learning baseline for other small image classification problems
Out-of-Scope Use
This model is not suitable for:
- Recognizing letters, symbols, or non-digit characters
- Noisy, real-world document scans without preprocessing
- Multi-digit or multi-character sequences in a single image
- Safety-critical systems (e.g., medical, legal document processing)
Bias, Risks, and Limitations
- Dataset bias: MNIST digits are clean, centered, and size-normalized. The model may underperform on digits written in non-Western styles, extreme stroke widths, or unusual orientations.
- Domain shift: Performance degrades on images that differ significantly from the MNIST distribution (e.g., photos of digits on paper, different fonts).
- Input polarity: MNIST digits are white on a black background. Dark-on-light images should be inverted before inference.
- No uncertainty calibration: The model outputs softmax probabilities, which may appear confident even on out-of-distribution inputs.
Recommendations
- Preprocess input images to 28Γ28 grayscale, white digit on black, and center/normalize digits before inference.
- Do not rely on model confidence scores alone; add a rejection threshold for production use.
- Evaluate on your specific distribution before deploying in any real-world scenario.
How to Get Started with the Model
pip install torch torchvision huggingface_hub pillow
import torch
from huggingface_hub import hf_hub_download
from PIL import Image
from torchvision import transforms
# Download the architecture and weights from this repo
hf_hub_download("abdurafay19/Digit-Classifier", "model.py", local_dir=".")
weights = hf_hub_download("abdurafay19/Digit-Classifier", "model.pt")
from model import Model
model = Model()
# map_location is required on CPU-only machines: the weights were saved from a GPU
model.load_state_dict(torch.load(weights, map_location="cpu", weights_only=True))
model.eval()
transform = transforms.Compose([
transforms.Grayscale(),
transforms.Resize((28, 28)),
transforms.ToTensor(),
transforms.Normalize((0.1307,), (0.3081,)),
])
img = Image.open("digit.png") # white digit on a black background
x = transform(img).unsqueeze(0) # shape: [1, 1, 28, 28]
with torch.no_grad():
probs = torch.softmax(model(x), dim=1)[0]
digit = probs.argmax().item()
print(f"Predicted digit: {digit} ({probs[digit]:.1%} confidence)")
Training Details
Training Data
- Dataset: MNIST: 70,000 grayscale images (60,000 train / 10,000 test)
- Input size: 28Γ28 pixels, single channel
- Classes: 10 (digits 0β9)
Training Procedure
Preprocessing
- Images converted to tensors and normalized using the MNIST mean (0.1307) and std (0.3081)
- Training augmentation: random rotation (Β±10Β°), random affine with translation (Β±10%), scale (0.9β1.1Γ), and shear (Β±5Β°)
- Test images: normalization only, no augmentation
Training Hyperparameters
| Parameter | Value |
|---|---|
| Optimizer | AdamW |
| Learning Rate | 3e-3 (max, OneCycleLR) |
| Weight Decay | 1e-4 |
| Batch Size | 64 |
| Epochs | 50 (best checkpoint: epoch 40) |
| Loss Function | CrossEntropyLoss |
| Label Smoothing | 0.1 |
| Gradient Clipping | Max norm 1.0 |
| LR Scheduler | OneCycleLR (10% warmup, cosine anneal) |
| Dropout (conv) | 0.25 (Dropout2d) |
| Dropout (FC) | 0.25 |
| Random Seed | 23 |
| Training regime | fp32 |
Speeds, Sizes, Times
- Training time: ~10 minutes on a single GPU (NVIDIA T4, Google Colab)
- Model parameters: 160,842
- Weights file: 643 KB (
model.pt, 658,805 bytes, fp32 state dict) - Inference speed: ~2 ms per image on CPU. This is the forward pass only at batch size 1, measured on an Intel Core i5-5200U laptop CPU with PyTorch 2.1.2 (2 threads); the mean ranged from 1.2 to 2.2 ms across runs. It excludes preprocessing. Reproduce with
scripts/benchmark.py.
Evaluation
Testing Data, Factors & Metrics
Testing Data
Evaluated on the standard MNIST test split: 10,000 images not seen during training.
Factors
Evaluation was performed across all 10 digit classes. No disaggregation by subpopulation was conducted (MNIST does not include demographic metadata).
Metrics
- Accuracy: primary metric; proportion of correctly classified digits
- Confusion matrix: to identify per-class error patterns
Results
| Metric | Value |
|---|---|
| Test Accuracy | 99.43% (9,943 / 10,000) |
Per-Class Accuracy
| Digit | Correct | Errors | Accuracy |
|---|---|---|---|
| 0 | 980 | 0 | 100.0% |
| 1 | 1132 | 3 | 99.7% |
| 2 | 1025 | 7 | 99.3% |
| 3 | 1008 | 2 | 99.8% |
| 4 | 976 | 6 | 99.4% |
| 5 | 885 | 7 | 99.2% |
| 6 | 949 | 9 | 99.1% |
| 7 | 1020 | 8 | 99.2% |
| 8 | 968 | 6 | 99.4% |
| 9 | 1000 | 9 | 99.1% |
Summary
The model achieves 99.43% accuracy on the MNIST test set (57 errors out of 10,000). Digit 0 is classified perfectly. Digits 6 and 9 have the most errors (9 each), but they are not confused with each other: 9s are most often misread as 4 (4 times) or 7 (3 times), and 6s as 5 (3 times). The most frequent single confusions are 7β1, 9β4 and 4β9 (4 times each).
Model Examination
The model's convolutional filters learn edge detectors and stroke patterns in early layers, which compose into digit-specific features in deeper layers. Standard CNN interpretability techniques (e.g., Grad-CAM) can be applied to visualize which regions most influence predictions.
Environmental Impact
Carbon emissions estimated using the ML Impact Calculator.
| Factor | Value |
|---|---|
| Hardware Type | NVIDIA T4 GPU |
| Hours Used | ~0.2 hrs (10 min) |
| Cloud Provider | Google Colab |
| Compute Region | Singapore |
| Carbon Emitted | ~0.01 kg COβeq (est.) |
Technical Specifications
Model Architecture
The model uses 4 convolutional blocks followed by a compact fully connected head. Every conv block is Conv2d β BatchNorm2d β ReLU β MaxPool2d(2) β Dropout2d(0.25).
| Block | Layer | Output shape | Parameters |
|---|---|---|---|
| Input | 28Γ28 grayscale image | 1 Γ 28 Γ 28 | β |
| Conv block 1 | Conv2d 1β32, 3Γ3, padding 1 | 32 Γ 14 Γ 14 | 384 |
| Conv block 2 | Conv2d 32β64, 3Γ3, padding 1 | 64 Γ 7 Γ 7 | 18,624 |
| Conv block 3 | Conv2d 64β128, 3Γ3, padding 1 | 128 Γ 3 Γ 3 | 74,112 |
| Conv block 4 | Conv2d 128β256, 1Γ1, no padding | 256 Γ 1 Γ 1 | 33,536 |
| Classifier | Flatten β Linear 256β128 β ReLU β Dropout(0.25) | 128 | 32,896 |
| Output | Linear 128β10 (logits) | 10 | 1,290 |
| Total | 160,842 |
Shape Flow
Input: (B, 1, 28, 28)
Block 1: (B, 32, 14, 14)
Block 2: (B, 64, 7, 7)
Block 3: (B, 128, 3, 3)
Block 4: (B, 256, 1, 1)
Flatten: (B, 256)
FC1: (B, 128)
Output: (B, 10)
Compute Infrastructure
- Training hardware: NVIDIA T4 GPU (Google Colab)
- Software: Python 3.10, PyTorch 2.1.2, torchvision 0.16.2
Citation
If you use this model in your work, please cite:
BibTeX:
@misc{digit-classifier-2026,
author = {Abdul Rafay},
title = {Handwritten Digit Classifier (CNN on MNIST)},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/abdurafay19/Digit-Classifier}
}
APA:
Abdul Rafay. (2026). Handwritten Digit Classifier (CNN on MNIST). Hugging Face. https://huggingface.co/abdurafay19/Digit-Classifier
Glossary
| Term | Definition |
|---|---|
| CNN | Convolutional Neural Network, a deep learning architecture suited for image data |
| MNIST | A benchmark dataset of 70,000 handwritten digit images |
| Softmax | Activation function that converts raw outputs to probabilities summing to 1 |
| Dropout | Regularization technique that randomly disables neurons during training |
| BatchNorm | Batch Normalization: normalizes layer activations to stabilize and speed up training |
| OneCycleLR | Learning rate schedule with warmup and cosine decay for faster convergence |
| Label Smoothing | Softens hard targets to reduce overconfidence and improve generalization |
| Grad-CAM | Gradient-weighted Class Activation Mapping, a model interpretability technique |
Model Card Authors
Abdul Rafay: CS student at ITU and AI Engineer (LinkedIn Β· GitHub) Β· abdulrafay17wolf@gmail.com
Model Card Contact
For questions or issues, open a GitHub issue at github.com/abdurafay19/Digit-Classifier or reach out via Hugging Face.
Dataset used to train abdurafay19/Digit-Classifier
Space using abdurafay19/Digit-Classifier 1
Evaluation results
- Test Accuracy on MNISTtest set self-reported99.430

