predictive-maintenance-rag-system / PROJECT_PROGRESS.md
SyedaArisha's picture
Upload folder using huggingface_hub
baf834b verified
|
Raw History Blame Contribute Delete
18.8 kB
# LLM-Enhanced Predictive Maintenance and Production Planning for FMCG Manufacturing
**Author:** Syeda Arisha Hassan
**Started:** July 2026
**Status:** In Progress
---
## 1. Project Idea
Build an end-to-end desktop pipeline that:
1. Predicts machine failures before they happen
2. Estimates Remaining Useful Life (RUL) of machines
3. Reads free-text maintenance logs using an LLM
4. Generates plain-language explanations for production managers
5. Automatically adjusts the production schedule based on predictions
Application domain is FMCG manufacturing — high-volume, continuous production of fast-moving consumer goods.
---
## 2. Why FMCG Manufacturing
FMCG refers to the products (milk, detergents, snacks, personal care items) and the industry that manufactures them. Factories run by Unilever, Nestlé, P&G, Engro Foods are FMCG manufacturing plants. Production is an integral part of the FMCG industry.
Machine failures are especially costly here because:
- Production lines run almost continuously
- Products have short shelf life
- Profit margins are thin
- Stockouts quickly affect retail availability
- One hour of unplanned downtime costs $10,000 to $36,000 (McKinsey)
- Predictive maintenance can reduce unplanned downtime by up to 50% and maintenance costs by 10-40%
---
## 3. Research Gaps (Proven from Literature)
### Gap 1: LLM use in PdM is fragmented
**Paper:** Toward Autonomous LLM-Based AI Agents for Predictive Maintenance (Di Maggio, 2025, MDPI Applied Sciences)
**Link:** https://www.mdpi.com/2076-3417/15/21/11515
**Exact quote:** "The literature on LLM-driven agents for PdM remains fragmented and lacks a unified view... the literature on autonomous agents for PdM is fragmented, lacks shared benchmarks, and does not offer a unified architectural vision calibrated to industrial maintenance workflows."
### Gap 2: Explainability and trust gap
**Paper:** Explainable Predictive Maintenance: A Survey of Current Methods, Challenges and Opportunities (Cummins et al., 2024)
**Link:** https://arxiv.org/abs/2401.07871
**Exact quote:** "As these methods are adopted for more serious and potentially life-threatening applications, the human operators need to trust the predictive system... explainability and interpretability into the predictive system."
### Gap Table
| Gap | Current State | This Project |
|---|---|---|
| LLM in PdM fragmented | Isolated tools | Integrated pipeline: XGBoost + LSTM + LLM + FAISS + scheduling |
| Black-box models | Technical metrics only | LLM plain-language explanation for managers |
| Text logs unused | Only sensor data used | FAISS retrieval of historical logs fed to LLM |
| Prediction and scheduling separate | Two separate research streams | Closed-loop system |
| No FMCG-specific research | Aerospace and automotive dominate | 100% focused on FMCG manufacturing |
---
## 4. System Architecture
### Full Pipeline
```
Multiple Datasets
↓
Preprocessing + Feature Engineering
↓
├── CPU Training
│ ├── XGBoost on AI4I 2020
│ ├── XGBoost on Pump Sensor Data
│ └── Ensemble → Final Classifier
│
├── GPU Training (Colab T4)
│ ├── LSTM on NASA CMAPSS
│ ├── CNN-LSTM on Azure PdM Telemetry
│ └── Ensemble → Final RUL Predictor
│
Ensemble Layer (Classification + RUL combined)
↓
FAISS Vector DB
(Historical logs + past predictions + user interactions)
↓
LLM - Llama 3
(Retrieve similar cases → Generate plain-language explanation)
↓
Production Schedule Adjustment
↓
PyQt5 Desktop Dashboard
(Machine status, RUL bars, alerts, LLM explanation panel)
↓
Production Manager
```
### Models
| Model | Task | Dataset | Hardware |
|---|---|---|---|
| XGBoost | Failure classification | AI4I 2020 + Pump Sensor | CPU |
| Random Forest | Failure classification (backup) | AI4I 2020 + Pump Sensor | CPU |
| LSTM | RUL regression | NASA CMAPSS | GPU |
| CNN-LSTM | RUL regression | Azure PdM Telemetry | GPU |
| Llama 3 | Log analysis + explanation | Azure PdM text logs | GPU |
### Vector Database
- **Tool:** FAISS
- **Stores:** Historical maintenance logs, previous model predictions with timestamps, user queries and manager decisions
- **Role:** RAG architecture — retrieves similar past cases to give LLM context
### Frontend
- **Tool:** PyQt5 desktop application
- **Features:** Machine status panel, RUL progress bars, LLM explanation panel, alerts, auto-adjusted schedule view
- **Why desktop:** Fully local, no internet dependency, data stays private, lightweight
---
## 5. Datasets
| Dataset | Purpose | Rows | Features | Status |
|---|---|---|---|---|
| AI4I 2020 | XGBoost classification | 10,000 | 14 | EDA + model done |
| Pump Sensor Data | XGBoost classification | 220,320 | 53 | EDA in progress |
| NASA CMAPSS | LSTM RUL regression | Multiple | 26 | Downloaded |
| Azure PdM | LSTM + LLM text logs | 876,000+ | Multiple files | Downloaded |
---
## 6. Quantitative Targets
### Industry Impact
| Metric | Value |
|---|---|
| Unplanned downtime reduction with PdM | Up to 50% |
| Maintenance cost reduction | 10-40% |
| Cost of 1 hour downtime (FMCG) | $10,000 - $36,000 |
| Annual downtime cost (large plants) | Up to $10 million |
### Model Performance Targets
| Metric | Baseline | Target |
|---|---|---|
| Failure classification accuracy | 85-92% single model | 93-95% ensemble |
| RUL error (MAE/RMSE) | Standard LSTM | 10-20% lower |
| Text log utilization | 0% in most systems | 100% via LLM + FAISS |
| Plain-language explanation coverage | None | 100% |
| Prediction + scheduling linkage | Separate | Fully closed-loop |
---
## 7. Tools and Environment
- **Language:** Python
- **ML:** XGBoost, Scikit-learn, PyTorch, PyTorch Forecasting
- **LLM:** Llama 3
- **Vector DB:** FAISS
- **Data:** Pandas, NumPy
- **Visualization:** Matplotlib, Seaborn
- **Frontend:** PyQt5
- **Training:** Google Colab T4 GPU, Kaggle Notebooks (backup)
- **Storage:** Google Drive (models, checkpoints, datasets)
---
## 8. Progress Log
### Week 1 — July 27, 2026
#### Environment Setup
- Downloaded all 4 datasets to Google Drive
- Set up Google Colab with Drive mount
- Organized datasets in `/datastes/` folder
#### AI4I 2020 — EDA Complete
**Dataset info:**
- 10,000 rows, 14 columns
- No missing values
- Target: Machine failure (0/1)
- Class distribution: 9,661 normal, 339 failures (3.4%)
- 5 failure modes: HDF (115), OSF (98), PWF (95), TWF (46), RNF (19)
**Correlation findings:**
- Air temp and process temp: 0.88 (highly correlated)
- Rotational speed and torque: -0.88 (strong negative)
- HDF most correlated with failure: 0.58
**Feature engineering:**
```python
df['temp_diff'] = df['Process temperature [K]'] - df['Air temperature [K]']
df['power'] = df['Torque [Nm]'] * df['Rotational speed [rpm]']
```
Both engineered features ranked in top 5 by importance.
**Feature importance (top 5):**
1. rotational_speed: 0.33
2. power: 0.23 (engineered)
3. tool_wear: 0.16
4. torque: 0.13
5. temp_diff: 0.07 (engineered)
**Model results:**
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Baseline XGBoost | 0.75 | 0.77 | 0.76 |
| XGBoost + SMOTE | 0.66 | 0.77 | 0.71 |
| Tuned XGBoost | 0.72 | 0.80 | 0.76 |
**Best model code:**
```python
from sklearn.preprocessing import LabelEncoder
from xgboost import XGBClassifier
from sklearn.model_selection import train_test_split
le = LabelEncoder()
df['Type_encoded'] = le.fit_transform(df['Type'])
df = df.rename(columns={
'Air temperature [K]': 'air_temp',
'Process temperature [K]': 'process_temp',
'Rotational speed [rpm]': 'rotational_speed',
'Torque [Nm]': 'torque',
'Tool wear [min]': 'tool_wear'
})
features = ['Type_encoded', 'air_temp', 'process_temp',
'rotational_speed', 'torque', 'tool_wear',
'temp_diff', 'power']
X = df[features]
y = df['Machine failure']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model3 = XGBClassifier(
scale_pos_weight=9661/339,
n_estimators=200,
max_depth=6,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
random_state=42
)
model3.fit(X_train, y_train)
```
**Model saved:**
```python
import joblib
joblib.dump(model3, '/content/drive/MyDrive/datastes/xgboost_ai4i.pkl')
```
---
#### SECOM — Attempted, Replaced
**Issues found:**
- Only 1,567 rows, 104 fail cases — too few for reliable modeling
- 591 features with heavy missing values and constant columns
- Best recall achieved: 0.33 even with SMOTE
- Decision: replaced with Pump Sensor Data
---
#### Pump Sensor Data — EDA In Progress
**Dataset info:**
- 220,320 rows, 55 columns
- Target: machine_status (NORMAL, RECOVERING, BROKEN)
- Class distribution: NORMAL 205,836, RECOVERING 14,477, BROKEN 7
**Issues found:**
1. **Data leakage with random split:** RECOVERING rows have sensor readings that directly reflect broken state. sensor_04 correlation with failure: 0.916, sensor_10: 0.872. Model achieved perfect 1.00 score which is not real.
2. **Time-based split problem:** All failures concentrated before index 160,000. Last 20% of data has 0 failures, making test set useless.
3. **sensor_15:** Entirely null (220,320 missing values), dropped.
**Current approach being tested:**
Remove RECOVERING label, only keep BROKEN as true failure, use 60-minute look-ahead window to label rows before failure as at-risk:
```python
df_pump2['failure'] = df_pump2['machine_status'].apply(
lambda x: 1 if x == 'BROKEN' else 0)
df_pump2['failure_ahead'] = 0
failure_indices = df_pump2[df_pump2['failure'] == 1].index
for idx in failure_indices:
start = max(0, idx - 60)
df_pump2.loc[start:idx, 'failure_ahead'] = 1
```
**Status:** In progress, resolving leakage issue.
---
## 9. Code
### Cell 1: Mount Drive
```python
from google.colab import drive
drive.mount('/content/drive')
```
### Cell 2: Load AI4I Dataset
```python
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
df = pd.read_csv('/content/drive/MyDrive/datastes/ai4i2020.csv')
print(df.shape)
print(df.head())
print(df.info())
print(df['Machine failure'].value_counts())
```
### Cell 3: Failure mode distribution
```python
print(df[['TWF','HDF','PWF','OSF','RNF']].sum())
print(df['Type'].value_counts())
```
### Cell 4: Correlation heatmap
```python
plt.figure(figsize=(10,6))
sns.heatmap(df.drop(columns=['UDI','Product ID','Type']).corr(),
annot=True, fmt='.2f', cmap='coolwarm')
plt.title('Feature Correlation Heatmap')
plt.tight_layout()
plt.show()
```
### Cell 5: Feature engineering
```python
df['temp_diff'] = df['Process temperature [K]'] - df['Air temperature [K]']
df['power'] = df['Torque [Nm]'] * df['Rotational speed [rpm]']
print(df[['temp_diff', 'power', 'Machine failure']].corr())
```
### Cell 6: Rename columns
```python
df = df.rename(columns={
'Air temperature [K]': 'air_temp',
'Process temperature [K]': 'process_temp',
'Rotational speed [rpm]': 'rotational_speed',
'Torque [Nm]': 'torque',
'Tool wear [min]': 'tool_wear'
})
```
### Cell 7: Train baseline XGBoost
```python
from sklearn.preprocessing import LabelEncoder
from xgboost import XGBClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, confusion_matrix
le = LabelEncoder()
df['Type_encoded'] = le.fit_transform(df['Type'])
features = ['Type_encoded', 'air_temp', 'process_temp',
'rotational_speed', 'torque', 'tool_wear',
'temp_diff', 'power']
X = df[features]
y = df['Machine failure']
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
model = XGBClassifier(scale_pos_weight=9661/339, random_state=42)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print(classification_report(y_test, y_pred))
print(confusion_matrix(y_test, y_pred))
```
### Cell 8: Feature importance plot
```python
import pandas as pd
plt.figure(figsize=(8,5))
pd.Series(model.feature_importances_, index=features).sort_values().plot(kind='barh')
plt.title('Feature Importance')
plt.tight_layout()
plt.show()
```
### Cell 9: Try SMOTE
```python
from imblearn.over_sampling import SMOTE
sm = SMOTE(random_state=42)
X_res, y_res = sm.fit_resample(X_train, y_train)
model2 = XGBClassifier(random_state=42)
model2.fit(X_res, y_res)
y_pred2 = model2.predict(X_test)
print(classification_report(y_test, y_pred2))
```
### Cell 10: Tuned XGBoost (best model)
```python
model3 = XGBClassifier(
scale_pos_weight=9661/339,
n_estimators=200,
max_depth=6,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
random_state=42
)
model3.fit(X_train, y_train)
y_pred3 = model3.predict(X_test)
print(classification_report(y_test, y_pred3))
```
### Cell 11: Save AI4I model
```python
import joblib
joblib.dump(model3, '/content/drive/MyDrive/datastes/xgboost_ai4i.pkl')
print("AI4I model saved")
```
### Cell 12: Load SECOM
```python
df_secom = pd.read_csv('/content/drive/MyDrive/datastes/uci-secom.csv')
print(df_secom.shape)
print(df_secom['Pass/Fail'].value_counts())
print(f"Missing values: {df_secom.isnull().sum().sum()}")
print(f"Missing percentage: {df_secom.isnull().mean().mean()*100:.2f}%")
cols_high_null = df_secom.isnull().mean()[df_secom.isnull().mean() > 0.5]
print(f"Columns with more than 50% missing: {len(cols_high_null)}")
```
### Cell 13: SECOM preprocessing
```python
df_secom = df_secom.drop(columns=cols_high_null.index)
df_secom = df_secom.drop(columns=['Time'])
df_secom = df_secom.fillna(df_secom.median())
X_secom = df_secom.drop(columns=['Pass/Fail'])
y_secom = df_secom['Pass/Fail'].replace({-1: 0, 1: 1})
print(f"Shape after cleaning: {X_secom.shape}")
print(f"Class distribution:\n{y_secom.value_counts()}")
```
### Cell 14: SECOM feature selection
```python
from sklearn.feature_selection import SelectKBest, f_classif
selector = SelectKBest(f_classif, k=20)
X_secom_selected = selector.fit_transform(X_secom, y_secom)
selected_cols = X_secom.columns[selector.get_support()]
print("Top 20 features selected:")
print(selected_cols.tolist())
X_secom_final = pd.DataFrame(X_secom_selected, columns=selected_cols)
print(f"Final shape: {X_secom_final.shape}")
```
### Cell 15: SECOM XGBoost baseline
```python
X_train_s, X_test_s, y_train_s, y_test_s = train_test_split(
X_secom_final, y_secom, test_size=0.2, random_state=42)
model_secom = XGBClassifier(
scale_pos_weight=1463/104,
n_estimators=200,
max_depth=6,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
random_state=42
)
model_secom.fit(X_train_s, y_train_s)
y_pred_s = model_secom.predict(X_test_s)
print(classification_report(y_test_s, y_pred_s))
```
### Cell 16: SECOM with SMOTE
```python
sm2 = SMOTE(random_state=42)
X_res_s, y_res_s = sm2.fit_resample(X_train_s, y_train_s)
model_secom2 = XGBClassifier(
n_estimators=200,
max_depth=4,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
random_state=42
)
model_secom2.fit(X_res_s, y_res_s)
y_pred_s2 = model_secom2.predict(X_test_s)
print(classification_report(y_test_s, y_pred_s2))
```
### Cell 17: Save SECOM model
```python
joblib.dump(model_secom2, '/content/drive/MyDrive/datastes/xgboost_secom.pkl')
print("SECOM model saved")
```
### Cell 18: Load Pump Sensor Data
```python
df_pump = pd.read_csv('/content/drive/MyDrive/datastes/sensor.csv')
print(df_pump.shape)
print(df_pump['machine_status'].value_counts())
```
### Cell 19: Pump initial preprocessing
```python
df_pump['failure'] = df_pump['machine_status'].apply(
lambda x: 1 if x in ['BROKEN', 'RECOVERING'] else 0)
print(df_pump['failure'].value_counts())
print(f"Missing values: {df_pump.isnull().sum().sum()}")
```
### Cell 20: Check leakage source
```python
print(df_pump.isnull().sum().sort_values(ascending=False).head(10))
print(f"Constant columns: {(df_pump.nunique() == 1).sum()}")
print(df_pump.corr()['failure'].abs().sort_values(ascending=False).head(10))
```
### Cell 21: Visualize top correlated sensors
```python
fig, axes = plt.subplots(1, 3, figsize=(15, 4))
for i, sensor in enumerate(['sensor_04', 'sensor_10', 'sensor_11']):
df_pump.groupby('failure')[sensor].plot(
kind='hist', alpha=0.5, ax=axes[i], legend=True)
axes[i].set_title(sensor)
plt.tight_layout()
plt.show()
```
### Cell 22: Check failure distribution over time
```python
print(f"Failures in train: {y_train_p.sum()}")
print(f"Failures in test: {y_test_p.sum()}")
plt.figure(figsize=(12,3))
y_pump.reset_index(drop=True).plot()
plt.title('Failure distribution over time')
plt.show()
```
### Cell 23: Fix leakage — reload and use look-ahead window
```python
df_pump2 = pd.read_csv('/content/drive/MyDrive/datastes/sensor.csv')
df_pump2 = df_pump2.drop(columns=['Unnamed: 0', 'sensor_15'])
status = df_pump2['machine_status'].copy()
timestamp = df_pump2['timestamp'].copy()
df_pump2 = df_pump2.drop(columns=['timestamp', 'machine_status'])
df_pump2 = df_pump2.fillna(df_pump2.median())
df_pump2['timestamp'] = timestamp
df_pump2['machine_status'] = status
df_pump2['failure'] = df_pump2['machine_status'].apply(
lambda x: 1 if x == 'BROKEN' else 0)
df_pump2 = df_pump2.sort_values('timestamp').reset_index(drop=True)
df_pump2['failure_ahead'] = 0
failure_indices = df_pump2[df_pump2['failure'] == 1].index
for idx in failure_indices:
start = max(0, idx - 60)
df_pump2.loc[start:idx, 'failure_ahead'] = 1
print(df_pump2['failure_ahead'].value_counts())
print(f"Actual broken rows: {df_pump2['failure'].sum()}")
```
---
## 10. Pending Work
- [ ] Resolve pump sensor leakage and complete XGBoost on pump data
- [ ] Ensemble AI4I and pump models
- [ ] Download and EDA on NASA CMAPSS
- [ ] LSTM training on CMAPSS
- [ ] CNN-LSTM training on Azure PdM telemetry
- [ ] Ensemble LSTM models
- [ ] FAISS setup and maintenance log embedding
- [ ] LLM integration with Llama 3
- [ ] Closed-loop scheduling logic
- [ ] PyQt5 dashboard development
- [ ] Full pipeline integration and testing
- [ ] Evaluation against baselines
- [ ] Scope, SRS, SDD documentation
- [ ] Research paper writing
---
## 11. References
1. Di Maggio, L.G. (2025). Toward Autonomous LLM-Based AI Agents for Predictive Maintenance. *Applied Sciences, MDPI*. https://www.mdpi.com/2076-3417/15/21/11515
2. Cummins, L. et al. (2024). Explainable Predictive Maintenance: A Survey of Current Methods, Challenges and Opportunities. *arXiv:2401.07871*. https://arxiv.org/abs/2401.07871
3. McKinsey & Company. Predictive Maintenance Industry Impact Data.
4. Matzka, S. (2020). Explainable Artificial Intelligence for Predictive Maintenance Applications. *AI4I Conference*.
5. NASA CMAPSS Dataset. Prognostics CoE at NASA Ames.
6. Microsoft Azure Predictive Maintenance Dataset.