|
Download PROJECT_PROGRESS.md from SyedaArisha/predictive-maintenance-rag-system: direct link, hf CLI and curl.
- Browser
- Download file 18.8 kB
-
https://huggingface.co/SyedaArisha/predictive-maintenance-rag-system/resolve/main/PROJECT_PROGRESS.md
- Command line
-
hf download hf://SyedaArisha/predictive-maintenance-rag-system/PROJECT_PROGRESS.md
-
curl -L -o PROJECT_PROGRESS.md https://huggingface.co/SyedaArisha/predictive-maintenance-rag-system/resolve/main/PROJECT_PROGRESS.md
18.8 kB
| # LLM-Enhanced Predictive Maintenance and Production Planning for FMCG Manufacturing | |
| **Author:** Syeda Arisha Hassan | |
| **Started:** July 2026 | |
| **Status:** In Progress | |
| --- | |
| ## 1. Project Idea | |
| Build an end-to-end desktop pipeline that: | |
| 1. Predicts machine failures before they happen | |
| 2. Estimates Remaining Useful Life (RUL) of machines | |
| 3. Reads free-text maintenance logs using an LLM | |
| 4. Generates plain-language explanations for production managers | |
| 5. Automatically adjusts the production schedule based on predictions | |
| Application domain is FMCG manufacturing — high-volume, continuous production of fast-moving consumer goods. | |
| --- | |
| ## 2. Why FMCG Manufacturing | |
| FMCG refers to the products (milk, detergents, snacks, personal care items) and the industry that manufactures them. Factories run by Unilever, Nestlé, P&G, Engro Foods are FMCG manufacturing plants. Production is an integral part of the FMCG industry. | |
| Machine failures are especially costly here because: | |
| - Production lines run almost continuously | |
| - Products have short shelf life | |
| - Profit margins are thin | |
| - Stockouts quickly affect retail availability | |
| - One hour of unplanned downtime costs $10,000 to $36,000 (McKinsey) | |
| - Predictive maintenance can reduce unplanned downtime by up to 50% and maintenance costs by 10-40% | |
| --- | |
| ## 3. Research Gaps (Proven from Literature) | |
| ### Gap 1: LLM use in PdM is fragmented | |
| **Paper:** Toward Autonomous LLM-Based AI Agents for Predictive Maintenance (Di Maggio, 2025, MDPI Applied Sciences) | |
| **Link:** https://www.mdpi.com/2076-3417/15/21/11515 | |
| **Exact quote:** "The literature on LLM-driven agents for PdM remains fragmented and lacks a unified view... the literature on autonomous agents for PdM is fragmented, lacks shared benchmarks, and does not offer a unified architectural vision calibrated to industrial maintenance workflows." | |
| ### Gap 2: Explainability and trust gap | |
| **Paper:** Explainable Predictive Maintenance: A Survey of Current Methods, Challenges and Opportunities (Cummins et al., 2024) | |
| **Link:** https://arxiv.org/abs/2401.07871 | |
| **Exact quote:** "As these methods are adopted for more serious and potentially life-threatening applications, the human operators need to trust the predictive system... explainability and interpretability into the predictive system." | |
| ### Gap Table | |
| | Gap | Current State | This Project | | |
| |---|---|---| | |
| | LLM in PdM fragmented | Isolated tools | Integrated pipeline: XGBoost + LSTM + LLM + FAISS + scheduling | | |
| | Black-box models | Technical metrics only | LLM plain-language explanation for managers | | |
| | Text logs unused | Only sensor data used | FAISS retrieval of historical logs fed to LLM | | |
| | Prediction and scheduling separate | Two separate research streams | Closed-loop system | | |
| | No FMCG-specific research | Aerospace and automotive dominate | 100% focused on FMCG manufacturing | | |
| --- | |
| ## 4. System Architecture | |
| ### Full Pipeline | |
| ``` | |
| Multiple Datasets | |
| ↓ | |
| Preprocessing + Feature Engineering | |
| ↓ | |
| ├── CPU Training | |
| │ ├── XGBoost on AI4I 2020 | |
| │ ├── XGBoost on Pump Sensor Data | |
| │ └── Ensemble → Final Classifier | |
| │ | |
| ├── GPU Training (Colab T4) | |
| │ ├── LSTM on NASA CMAPSS | |
| │ ├── CNN-LSTM on Azure PdM Telemetry | |
| │ └── Ensemble → Final RUL Predictor | |
| │ | |
| Ensemble Layer (Classification + RUL combined) | |
| ↓ | |
| FAISS Vector DB | |
| (Historical logs + past predictions + user interactions) | |
| ↓ | |
| LLM - Llama 3 | |
| (Retrieve similar cases → Generate plain-language explanation) | |
| ↓ | |
| Production Schedule Adjustment | |
| ↓ | |
| PyQt5 Desktop Dashboard | |
| (Machine status, RUL bars, alerts, LLM explanation panel) | |
| ↓ | |
| Production Manager | |
| ``` | |
| ### Models | |
| | Model | Task | Dataset | Hardware | | |
| |---|---|---|---| | |
| | XGBoost | Failure classification | AI4I 2020 + Pump Sensor | CPU | | |
| | Random Forest | Failure classification (backup) | AI4I 2020 + Pump Sensor | CPU | | |
| | LSTM | RUL regression | NASA CMAPSS | GPU | | |
| | CNN-LSTM | RUL regression | Azure PdM Telemetry | GPU | | |
| | Llama 3 | Log analysis + explanation | Azure PdM text logs | GPU | | |
| ### Vector Database | |
| - **Tool:** FAISS | |
| - **Stores:** Historical maintenance logs, previous model predictions with timestamps, user queries and manager decisions | |
| - **Role:** RAG architecture — retrieves similar past cases to give LLM context | |
| ### Frontend | |
| - **Tool:** PyQt5 desktop application | |
| - **Features:** Machine status panel, RUL progress bars, LLM explanation panel, alerts, auto-adjusted schedule view | |
| - **Why desktop:** Fully local, no internet dependency, data stays private, lightweight | |
| --- | |
| ## 5. Datasets | |
| | Dataset | Purpose | Rows | Features | Status | | |
| |---|---|---|---|---| | |
| | AI4I 2020 | XGBoost classification | 10,000 | 14 | EDA + model done | | |
| | Pump Sensor Data | XGBoost classification | 220,320 | 53 | EDA in progress | | |
| | NASA CMAPSS | LSTM RUL regression | Multiple | 26 | Downloaded | | |
| | Azure PdM | LSTM + LLM text logs | 876,000+ | Multiple files | Downloaded | | |
| --- | |
| ## 6. Quantitative Targets | |
| ### Industry Impact | |
| | Metric | Value | | |
| |---|---| | |
| | Unplanned downtime reduction with PdM | Up to 50% | | |
| | Maintenance cost reduction | 10-40% | | |
| | Cost of 1 hour downtime (FMCG) | $10,000 - $36,000 | | |
| | Annual downtime cost (large plants) | Up to $10 million | | |
| ### Model Performance Targets | |
| | Metric | Baseline | Target | | |
| |---|---|---| | |
| | Failure classification accuracy | 85-92% single model | 93-95% ensemble | | |
| | RUL error (MAE/RMSE) | Standard LSTM | 10-20% lower | | |
| | Text log utilization | 0% in most systems | 100% via LLM + FAISS | | |
| | Plain-language explanation coverage | None | 100% | | |
| | Prediction + scheduling linkage | Separate | Fully closed-loop | | |
| --- | |
| ## 7. Tools and Environment | |
| - **Language:** Python | |
| - **ML:** XGBoost, Scikit-learn, PyTorch, PyTorch Forecasting | |
| - **LLM:** Llama 3 | |
| - **Vector DB:** FAISS | |
| - **Data:** Pandas, NumPy | |
| - **Visualization:** Matplotlib, Seaborn | |
| - **Frontend:** PyQt5 | |
| - **Training:** Google Colab T4 GPU, Kaggle Notebooks (backup) | |
| - **Storage:** Google Drive (models, checkpoints, datasets) | |
| --- | |
| ## 8. Progress Log | |
| ### Week 1 — July 27, 2026 | |
| #### Environment Setup | |
| - Downloaded all 4 datasets to Google Drive | |
| - Set up Google Colab with Drive mount | |
| - Organized datasets in `/datastes/` folder | |
| #### AI4I 2020 — EDA Complete | |
| **Dataset info:** | |
| - 10,000 rows, 14 columns | |
| - No missing values | |
| - Target: Machine failure (0/1) | |
| - Class distribution: 9,661 normal, 339 failures (3.4%) | |
| - 5 failure modes: HDF (115), OSF (98), PWF (95), TWF (46), RNF (19) | |
| **Correlation findings:** | |
| - Air temp and process temp: 0.88 (highly correlated) | |
| - Rotational speed and torque: -0.88 (strong negative) | |
| - HDF most correlated with failure: 0.58 | |
| **Feature engineering:** | |
| ```python | |
| df['temp_diff'] = df['Process temperature [K]'] - df['Air temperature [K]'] | |
| df['power'] = df['Torque [Nm]'] * df['Rotational speed [rpm]'] | |
| ``` | |
| Both engineered features ranked in top 5 by importance. | |
| **Feature importance (top 5):** | |
| 1. rotational_speed: 0.33 | |
| 2. power: 0.23 (engineered) | |
| 3. tool_wear: 0.16 | |
| 4. torque: 0.13 | |
| 5. temp_diff: 0.07 (engineered) | |
| **Model results:** | |
| | Model | Precision | Recall | F1 | | |
| |---|---|---|---| | |
| | Baseline XGBoost | 0.75 | 0.77 | 0.76 | | |
| | XGBoost + SMOTE | 0.66 | 0.77 | 0.71 | | |
| | Tuned XGBoost | 0.72 | 0.80 | 0.76 | | |
| **Best model code:** | |
| ```python | |
| from sklearn.preprocessing import LabelEncoder | |
| from xgboost import XGBClassifier | |
| from sklearn.model_selection import train_test_split | |
| le = LabelEncoder() | |
| df['Type_encoded'] = le.fit_transform(df['Type']) | |
| df = df.rename(columns={ | |
| 'Air temperature [K]': 'air_temp', | |
| 'Process temperature [K]': 'process_temp', | |
| 'Rotational speed [rpm]': 'rotational_speed', | |
| 'Torque [Nm]': 'torque', | |
| 'Tool wear [min]': 'tool_wear' | |
| }) | |
| features = ['Type_encoded', 'air_temp', 'process_temp', | |
| 'rotational_speed', 'torque', 'tool_wear', | |
| 'temp_diff', 'power'] | |
| X = df[features] | |
| y = df['Machine failure'] | |
| X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) | |
| model3 = XGBClassifier( | |
| scale_pos_weight=9661/339, | |
| n_estimators=200, | |
| max_depth=6, | |
| learning_rate=0.05, | |
| subsample=0.8, | |
| colsample_bytree=0.8, | |
| random_state=42 | |
| ) | |
| model3.fit(X_train, y_train) | |
| ``` | |
| **Model saved:** | |
| ```python | |
| import joblib | |
| joblib.dump(model3, '/content/drive/MyDrive/datastes/xgboost_ai4i.pkl') | |
| ``` | |
| --- | |
| #### SECOM — Attempted, Replaced | |
| **Issues found:** | |
| - Only 1,567 rows, 104 fail cases — too few for reliable modeling | |
| - 591 features with heavy missing values and constant columns | |
| - Best recall achieved: 0.33 even with SMOTE | |
| - Decision: replaced with Pump Sensor Data | |
| --- | |
| #### Pump Sensor Data — EDA In Progress | |
| **Dataset info:** | |
| - 220,320 rows, 55 columns | |
| - Target: machine_status (NORMAL, RECOVERING, BROKEN) | |
| - Class distribution: NORMAL 205,836, RECOVERING 14,477, BROKEN 7 | |
| **Issues found:** | |
| 1. **Data leakage with random split:** RECOVERING rows have sensor readings that directly reflect broken state. sensor_04 correlation with failure: 0.916, sensor_10: 0.872. Model achieved perfect 1.00 score which is not real. | |
| 2. **Time-based split problem:** All failures concentrated before index 160,000. Last 20% of data has 0 failures, making test set useless. | |
| 3. **sensor_15:** Entirely null (220,320 missing values), dropped. | |
| **Current approach being tested:** | |
| Remove RECOVERING label, only keep BROKEN as true failure, use 60-minute look-ahead window to label rows before failure as at-risk: | |
| ```python | |
| df_pump2['failure'] = df_pump2['machine_status'].apply( | |
| lambda x: 1 if x == 'BROKEN' else 0) | |
| df_pump2['failure_ahead'] = 0 | |
| failure_indices = df_pump2[df_pump2['failure'] == 1].index | |
| for idx in failure_indices: | |
| start = max(0, idx - 60) | |
| df_pump2.loc[start:idx, 'failure_ahead'] = 1 | |
| ``` | |
| **Status:** In progress, resolving leakage issue. | |
| --- | |
| ## 9. Code | |
| ### Cell 1: Mount Drive | |
| ```python | |
| from google.colab import drive | |
| drive.mount('/content/drive') | |
| ``` | |
| ### Cell 2: Load AI4I Dataset | |
| ```python | |
| import pandas as pd | |
| import matplotlib.pyplot as plt | |
| import seaborn as sns | |
| df = pd.read_csv('/content/drive/MyDrive/datastes/ai4i2020.csv') | |
| print(df.shape) | |
| print(df.head()) | |
| print(df.info()) | |
| print(df['Machine failure'].value_counts()) | |
| ``` | |
| ### Cell 3: Failure mode distribution | |
| ```python | |
| print(df[['TWF','HDF','PWF','OSF','RNF']].sum()) | |
| print(df['Type'].value_counts()) | |
| ``` | |
| ### Cell 4: Correlation heatmap | |
| ```python | |
| plt.figure(figsize=(10,6)) | |
| sns.heatmap(df.drop(columns=['UDI','Product ID','Type']).corr(), | |
| annot=True, fmt='.2f', cmap='coolwarm') | |
| plt.title('Feature Correlation Heatmap') | |
| plt.tight_layout() | |
| plt.show() | |
| ``` | |
| ### Cell 5: Feature engineering | |
| ```python | |
| df['temp_diff'] = df['Process temperature [K]'] - df['Air temperature [K]'] | |
| df['power'] = df['Torque [Nm]'] * df['Rotational speed [rpm]'] | |
| print(df[['temp_diff', 'power', 'Machine failure']].corr()) | |
| ``` | |
| ### Cell 6: Rename columns | |
| ```python | |
| df = df.rename(columns={ | |
| 'Air temperature [K]': 'air_temp', | |
| 'Process temperature [K]': 'process_temp', | |
| 'Rotational speed [rpm]': 'rotational_speed', | |
| 'Torque [Nm]': 'torque', | |
| 'Tool wear [min]': 'tool_wear' | |
| }) | |
| ``` | |
| ### Cell 7: Train baseline XGBoost | |
| ```python | |
| from sklearn.preprocessing import LabelEncoder | |
| from xgboost import XGBClassifier | |
| from sklearn.model_selection import train_test_split | |
| from sklearn.metrics import classification_report, confusion_matrix | |
| le = LabelEncoder() | |
| df['Type_encoded'] = le.fit_transform(df['Type']) | |
| features = ['Type_encoded', 'air_temp', 'process_temp', | |
| 'rotational_speed', 'torque', 'tool_wear', | |
| 'temp_diff', 'power'] | |
| X = df[features] | |
| y = df['Machine failure'] | |
| X_train, X_test, y_train, y_test = train_test_split( | |
| X, y, test_size=0.2, random_state=42) | |
| model = XGBClassifier(scale_pos_weight=9661/339, random_state=42) | |
| model.fit(X_train, y_train) | |
| y_pred = model.predict(X_test) | |
| print(classification_report(y_test, y_pred)) | |
| print(confusion_matrix(y_test, y_pred)) | |
| ``` | |
| ### Cell 8: Feature importance plot | |
| ```python | |
| import pandas as pd | |
| plt.figure(figsize=(8,5)) | |
| pd.Series(model.feature_importances_, index=features).sort_values().plot(kind='barh') | |
| plt.title('Feature Importance') | |
| plt.tight_layout() | |
| plt.show() | |
| ``` | |
| ### Cell 9: Try SMOTE | |
| ```python | |
| from imblearn.over_sampling import SMOTE | |
| sm = SMOTE(random_state=42) | |
| X_res, y_res = sm.fit_resample(X_train, y_train) | |
| model2 = XGBClassifier(random_state=42) | |
| model2.fit(X_res, y_res) | |
| y_pred2 = model2.predict(X_test) | |
| print(classification_report(y_test, y_pred2)) | |
| ``` | |
| ### Cell 10: Tuned XGBoost (best model) | |
| ```python | |
| model3 = XGBClassifier( | |
| scale_pos_weight=9661/339, | |
| n_estimators=200, | |
| max_depth=6, | |
| learning_rate=0.05, | |
| subsample=0.8, | |
| colsample_bytree=0.8, | |
| random_state=42 | |
| ) | |
| model3.fit(X_train, y_train) | |
| y_pred3 = model3.predict(X_test) | |
| print(classification_report(y_test, y_pred3)) | |
| ``` | |
| ### Cell 11: Save AI4I model | |
| ```python | |
| import joblib | |
| joblib.dump(model3, '/content/drive/MyDrive/datastes/xgboost_ai4i.pkl') | |
| print("AI4I model saved") | |
| ``` | |
| ### Cell 12: Load SECOM | |
| ```python | |
| df_secom = pd.read_csv('/content/drive/MyDrive/datastes/uci-secom.csv') | |
| print(df_secom.shape) | |
| print(df_secom['Pass/Fail'].value_counts()) | |
| print(f"Missing values: {df_secom.isnull().sum().sum()}") | |
| print(f"Missing percentage: {df_secom.isnull().mean().mean()*100:.2f}%") | |
| cols_high_null = df_secom.isnull().mean()[df_secom.isnull().mean() > 0.5] | |
| print(f"Columns with more than 50% missing: {len(cols_high_null)}") | |
| ``` | |
| ### Cell 13: SECOM preprocessing | |
| ```python | |
| df_secom = df_secom.drop(columns=cols_high_null.index) | |
| df_secom = df_secom.drop(columns=['Time']) | |
| df_secom = df_secom.fillna(df_secom.median()) | |
| X_secom = df_secom.drop(columns=['Pass/Fail']) | |
| y_secom = df_secom['Pass/Fail'].replace({-1: 0, 1: 1}) | |
| print(f"Shape after cleaning: {X_secom.shape}") | |
| print(f"Class distribution:\n{y_secom.value_counts()}") | |
| ``` | |
| ### Cell 14: SECOM feature selection | |
| ```python | |
| from sklearn.feature_selection import SelectKBest, f_classif | |
| selector = SelectKBest(f_classif, k=20) | |
| X_secom_selected = selector.fit_transform(X_secom, y_secom) | |
| selected_cols = X_secom.columns[selector.get_support()] | |
| print("Top 20 features selected:") | |
| print(selected_cols.tolist()) | |
| X_secom_final = pd.DataFrame(X_secom_selected, columns=selected_cols) | |
| print(f"Final shape: {X_secom_final.shape}") | |
| ``` | |
| ### Cell 15: SECOM XGBoost baseline | |
| ```python | |
| X_train_s, X_test_s, y_train_s, y_test_s = train_test_split( | |
| X_secom_final, y_secom, test_size=0.2, random_state=42) | |
| model_secom = XGBClassifier( | |
| scale_pos_weight=1463/104, | |
| n_estimators=200, | |
| max_depth=6, | |
| learning_rate=0.05, | |
| subsample=0.8, | |
| colsample_bytree=0.8, | |
| random_state=42 | |
| ) | |
| model_secom.fit(X_train_s, y_train_s) | |
| y_pred_s = model_secom.predict(X_test_s) | |
| print(classification_report(y_test_s, y_pred_s)) | |
| ``` | |
| ### Cell 16: SECOM with SMOTE | |
| ```python | |
| sm2 = SMOTE(random_state=42) | |
| X_res_s, y_res_s = sm2.fit_resample(X_train_s, y_train_s) | |
| model_secom2 = XGBClassifier( | |
| n_estimators=200, | |
| max_depth=4, | |
| learning_rate=0.05, | |
| subsample=0.8, | |
| colsample_bytree=0.8, | |
| random_state=42 | |
| ) | |
| model_secom2.fit(X_res_s, y_res_s) | |
| y_pred_s2 = model_secom2.predict(X_test_s) | |
| print(classification_report(y_test_s, y_pred_s2)) | |
| ``` | |
| ### Cell 17: Save SECOM model | |
| ```python | |
| joblib.dump(model_secom2, '/content/drive/MyDrive/datastes/xgboost_secom.pkl') | |
| print("SECOM model saved") | |
| ``` | |
| ### Cell 18: Load Pump Sensor Data | |
| ```python | |
| df_pump = pd.read_csv('/content/drive/MyDrive/datastes/sensor.csv') | |
| print(df_pump.shape) | |
| print(df_pump['machine_status'].value_counts()) | |
| ``` | |
| ### Cell 19: Pump initial preprocessing | |
| ```python | |
| df_pump['failure'] = df_pump['machine_status'].apply( | |
| lambda x: 1 if x in ['BROKEN', 'RECOVERING'] else 0) | |
| print(df_pump['failure'].value_counts()) | |
| print(f"Missing values: {df_pump.isnull().sum().sum()}") | |
| ``` | |
| ### Cell 20: Check leakage source | |
| ```python | |
| print(df_pump.isnull().sum().sort_values(ascending=False).head(10)) | |
| print(f"Constant columns: {(df_pump.nunique() == 1).sum()}") | |
| print(df_pump.corr()['failure'].abs().sort_values(ascending=False).head(10)) | |
| ``` | |
| ### Cell 21: Visualize top correlated sensors | |
| ```python | |
| fig, axes = plt.subplots(1, 3, figsize=(15, 4)) | |
| for i, sensor in enumerate(['sensor_04', 'sensor_10', 'sensor_11']): | |
| df_pump.groupby('failure')[sensor].plot( | |
| kind='hist', alpha=0.5, ax=axes[i], legend=True) | |
| axes[i].set_title(sensor) | |
| plt.tight_layout() | |
| plt.show() | |
| ``` | |
| ### Cell 22: Check failure distribution over time | |
| ```python | |
| print(f"Failures in train: {y_train_p.sum()}") | |
| print(f"Failures in test: {y_test_p.sum()}") | |
| plt.figure(figsize=(12,3)) | |
| y_pump.reset_index(drop=True).plot() | |
| plt.title('Failure distribution over time') | |
| plt.show() | |
| ``` | |
| ### Cell 23: Fix leakage — reload and use look-ahead window | |
| ```python | |
| df_pump2 = pd.read_csv('/content/drive/MyDrive/datastes/sensor.csv') | |
| df_pump2 = df_pump2.drop(columns=['Unnamed: 0', 'sensor_15']) | |
| status = df_pump2['machine_status'].copy() | |
| timestamp = df_pump2['timestamp'].copy() | |
| df_pump2 = df_pump2.drop(columns=['timestamp', 'machine_status']) | |
| df_pump2 = df_pump2.fillna(df_pump2.median()) | |
| df_pump2['timestamp'] = timestamp | |
| df_pump2['machine_status'] = status | |
| df_pump2['failure'] = df_pump2['machine_status'].apply( | |
| lambda x: 1 if x == 'BROKEN' else 0) | |
| df_pump2 = df_pump2.sort_values('timestamp').reset_index(drop=True) | |
| df_pump2['failure_ahead'] = 0 | |
| failure_indices = df_pump2[df_pump2['failure'] == 1].index | |
| for idx in failure_indices: | |
| start = max(0, idx - 60) | |
| df_pump2.loc[start:idx, 'failure_ahead'] = 1 | |
| print(df_pump2['failure_ahead'].value_counts()) | |
| print(f"Actual broken rows: {df_pump2['failure'].sum()}") | |
| ``` | |
| --- | |
| ## 10. Pending Work | |
| - [ ] Resolve pump sensor leakage and complete XGBoost on pump data | |
| - [ ] Ensemble AI4I and pump models | |
| - [ ] Download and EDA on NASA CMAPSS | |
| - [ ] LSTM training on CMAPSS | |
| - [ ] CNN-LSTM training on Azure PdM telemetry | |
| - [ ] Ensemble LSTM models | |
| - [ ] FAISS setup and maintenance log embedding | |
| - [ ] LLM integration with Llama 3 | |
| - [ ] Closed-loop scheduling logic | |
| - [ ] PyQt5 dashboard development | |
| - [ ] Full pipeline integration and testing | |
| - [ ] Evaluation against baselines | |
| - [ ] Scope, SRS, SDD documentation | |
| - [ ] Research paper writing | |
| --- | |
| ## 11. References | |
| 1. Di Maggio, L.G. (2025). Toward Autonomous LLM-Based AI Agents for Predictive Maintenance. *Applied Sciences, MDPI*. https://www.mdpi.com/2076-3417/15/21/11515 | |
| 2. Cummins, L. et al. (2024). Explainable Predictive Maintenance: A Survey of Current Methods, Challenges and Opportunities. *arXiv:2401.07871*. https://arxiv.org/abs/2401.07871 | |
| 3. McKinsey & Company. Predictive Maintenance Industry Impact Data. | |
| 4. Matzka, S. (2020). Explainable Artificial Intelligence for Predictive Maintenance Applications. *AI4I Conference*. | |
| 5. NASA CMAPSS Dataset. Prognostics CoE at NASA Ames. | |
| 6. Microsoft Azure Predictive Maintenance Dataset. | |