---
license: mit
library_name: scikit-learn
pipeline_tag: tabular-regression
tags:
  - machine-learning
  - regression
  - scikit-learn
  - random-forest
  - house-price-prediction
  - real-estate
  - india
---

# 🏠 gharDAAM – House Price Prediction Model

## Overview

**gharDAAM** is a machine learning project that predicts residential property prices in India using structured real estate data. The model is built using **Random Forest Regression** and trained on a cleaned and preprocessed dataset containing approximately **175,000** property listings.

This repository contains the trained model and documentation required to reproduce predictions.

---

## Model Information

| Property | Value |
|----------|-------|
| Task | House Price Prediction |
| Problem Type | Regression |
| Algorithm | Random Forest Regressor |
| Framework | scikit-learn |
| Language | Python |
| Target Variable | Property Price |

---

## Model Performance

| Metric | Score |
|--------|------:|
| RΒ² Score | **0.82** |
| Mean Absolute Error (MAE) | **β‚Ή17.1 Lakh** |

The Random Forest model was selected after evaluating multiple regression algorithms because it provided the best balance between predictive accuracy and generalization.

---

## Features Used

The model was trained using features such as:

- BHK
- Location
- Transaction Type
- Furnishing Status
- Bathroom Count
- Balcony Count
- Covered Parking
- Open Parking
- Garden/Park Availability
- Main Road Facing
- Swimming Pool Availability

---

## Requirements

Recommended package versions:

```text
Python >= 3.11

scikit-learn == 1.7.2
pandas
numpy

Using a different scikit-learn version may generate compatibility warnings when loading the serialized model.


Loading the Model

import pickle

with open("house_price_predictor_model.pkl", "rb") as f:
    model = pickle.load(f)

Example Prediction

import pandas as pd
import pickle

with open("house_price_predictor_model.pkl", "rb") as f:
    model = pickle.load(f)

sample = pd.DataFrame([{
    "BHK": 2,
    "Location": "thane",
    "Transaction": 0,
    "Furnishing": "Semi-Furnished",
    "Bathroom": 2,
    "Balcony": 1,
    "Covered Parking": 1,
    "Open Parking": 0,
    "Garden/Park": 1,
    "Main Road": 0,
    "Pool": 0
}])

prediction = model.predict(sample)

print(prediction)

Repository Contents

.
β”œβ”€β”€ house_price_predictor_model.pkl
β”œβ”€β”€ README.md

If preprocessing artifacts (such as encoders, scalers, or imputers) are required for inference, they should also be included in the repository.


Dataset

The model was trained using a cleaned and preprocessed version of a publicly available Indian residential real estate dataset.

The original dataset was obtained from Kaggle and licensed under the Community Data License Agreement – Sharing – Version 1.0 (CDLA-Sharing-1.0).

This repository does not redistribute the original dataset.


Limitations

  • Predictions are based only on the available training features.
  • The model should be used for educational and research purposes.
  • Real estate prices are influenced by market conditions that may change over time.
  • Performance may decrease for regions or property types underrepresented in the training data.

Future Improvements

  • Feature importance analysis
  • Model deployment using FastAPI
  • Interactive web interface
  • Continuous retraining with updated data

Acknowledgements

  • The original real estate dataset was sourced from Kaggle uploaded by Juhi Bhojani.

  • Thanks to the open-source Python ecosystem, especially:

    • scikit-learn
    • pandas
    • NumPy

License

This repository is released under the MIT License.

Please refer to the original dataset's license for any restrictions related to the training data.


Author

Raj Aryan

Electronics and Communication Engineering (ECE) Indian Institute of Information Technology (IIIT) Surat

Interested in:

  • Artificial Intelligence
  • Machine Learning
  • Robotics
  • Data Science

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support