bolewara commited on
Commit
5ed6bd2
·
verified ·
1 Parent(s): 67b09ff

Upload books-analysis.ipynb with huggingface_hub

Browse files
Files changed (1) hide show
  1. books-analysis.ipynb +266 -0
books-analysis.ipynb ADDED
@@ -0,0 +1,266 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "cells": [
3
+ {
4
+ "cell_type": "markdown",
5
+ "metadata": {},
6
+ "source": [
7
+ "# Books to Scrape — Price Analysis & Prediction\n",
8
+ "\n",
9
+ "**An end-to-end exploration of a freshly scraped book catalog.**\n",
10
+ "\n",
11
+ "This notebook walks through:\n",
12
+ "1. **Data cleaning** — parsing HTML-wrapped price strings, converting word ratings to numbers\n",
13
+ "2. **Exploratory analysis** — price distributions, rating patterns, title insights\n",
14
+ "3. ",
15
+ "**Feature engineering** — deriving numeric/text features from raw fields\n",
16
+ "4. **Modeling** — predicting book price with cross-validated gradient boosting\n",
17
+ "\n",
18
+ "Written for the `books-to-scrape-catalog-dataset`."
19
+ ]
20
+ },
21
+ {
22
+ "cell_type": "code",
23
+ "execution_count": null,
24
+ "metadata": {},
25
+ "outputs": [],
26
+ "source": [
27
+ "import os\n",
28
+ "import re\n",
29
+ "import numpy as np\n",
30
+ "import pandas as pd\n",
31
+ "import matplotlib.pyplot as plt\n",
32
+ "import seaborn as sns\n",
33
+ "from sklearn.model_selection import KFold\n",
34
+ "from sklearn.metrics import mean_absolute_error, r2_score\n",
35
+ "from sklearn.ensemble import RandomForestRegressor\n",
36
+ "from sklearn.feature_extraction.text import TfidfVectorizer\n",
37
+ "from sklearn.decomposition import TruncatedSVD\n",
38
+ "from sklearn.pipeline import make_pipeline\n",
39
+ "from sklearn.preprocessing import StandardScaler\n",
40
+ "\n",
41
+ "sns.set_theme(style='whitegrid')\n",
42
+ "plt.rcParams['figure.dpi'] = 110\n",
43
+ "\n",
44
+ "def load_dataset():\n",
45
+ " \"\"\"Locate the dataset anywhere under /kaggle/input or cwd.\"\"\"\n",
46
+ " target = 'books_data.csv'\n",
47
+ " if os.path.isfile(target):\n",
48
+ " return pd.read_csv(target)\n",
49
+ " for root, _, files in os.walk('/kaggle/input'):\n",
50
+ " if target in files:\n",
51
+ " return pd.read_csv(os.path.join(root, target))\n",
52
+ " raise FileNotFoundError(f'{target} not found')\n",
53
+ "\n",
54
+ "df = load_dataset()\n",
55
+ "print(f'Dataset shape: {df.shape}')"
56
+ ]
57
+ },
58
+ {
59
+ "cell_type": "markdown",
60
+ "metadata": {},
61
+ "source": [
62
+ "## 1. Data Cleaning\n",
63
+ "\n",
64
+ "The raw scrape stores prices as HTML (`<p class=\"price_color\">£51.77</p>`) and ratings as words (`Three`, `Five`). Let's normalize both."
65
+ ]
66
+ },
67
+ {
68
+ "cell_type": "code",
69
+ "execution_count": null,
70
+ "metadata": {},
71
+ "outputs": [],
72
+ "source": [
73
+ "RATING_MAP = {'One': 1, 'Two': 2, 'Three': 3, 'Four': 4, 'Five': 5}\n",
74
+ "\n",
75
+ "def extract_price(html_price):\n",
76
+ " match = re.search(r'([\\d]+\\.?[\\d]*)', str(html_price))\n",
77
+ " return float(match.group(1)) if match else np.nan\n",
78
+ "\n",
79
+ "df['price'] = df['Price'].apply(extract_price)\n",
80
+ "df['rating'] = df['Ratings'].map(RATING_MAP)\n",
81
+ "df['in_stock'] = df['Availability'].str.contains('In stock').astype(int)\n",
82
+ "\n",
83
+ "print(f'Price parsed: {df[\"price\"].notna().sum()}/{len(df)}')\n",
84
+ "print(f'Rating mapped: {df[\"rating\"].notna().sum()}/{len(df)}')\n",
85
+ "print(f'Unique stock values: {df[\"Availability\"].nunique()}')\n",
86
+ "\n",
87
+ "# Show cleaned head\n",
88
+ "df[['Title', 'price', 'rating', 'in_stock']].head()"
89
+ ]
90
+ },
91
+ {
92
+ "cell_type": "markdown",
93
+ "metadata": {},
94
+ "source": [
95
+ "## 2. Exploratory Data Analysis"
96
+ ]
97
+ },
98
+ {
99
+ "cell_type": "code",
100
+ "execution_count": null,
101
+ "metadata": {},
102
+ "outputs": [],
103
+ "source": [
104
+ "fig, axes = plt.subplots(1, 3, figsize=(15, 4))\n",
105
+ "\n",
106
+ "sns.histplot(df['price'], bins=30, kde=True, ax=axes[0], color='steelblue')\n",
107
+ "axes[0].set_title('Price Distribution')\n",
108
+ "axes[0].set_xlabel('Price (£)')\n",
109
+ "\n",
110
+ "sns.countplot(x='rating', data=df, ax=axes[1], order=sorted(df['rating'].dropna().unique()), palette='viridis')\n",
111
+ "axes[1].set_title('Rating Distribution')\n",
112
+ "axes[1].set_xlabel('Star Rating')\n",
113
+ "\n",
114
+ "# Average price by rating\n",
115
+ "avg_price_by_rating = df.groupby('rating')['price'].mean()\n",
116
+ "avg_price_by_rating.plot(kind='bar', ax=axes[2], color='coral')\n",
117
+ "axes[2].set_title('Average Price by Rating')\n",
118
+ "axes[2].set_xlabel('Star Rating')\n",
119
+ "axes[2].set_ylabel('Avg Price (£)')\n",
120
+ "\n",
121
+ "plt.tight_layout()\n",
122
+ "plt.show()\n",
123
+ "\n",
124
+ "print(f'Median price: £{df[\"price\"].median():.2f}')\n",
125
+ "print(f'Mean price: £{df[\"price\"].mean():.2f}')"
126
+ ]
127
+ },
128
+ {
129
+ "cell_type": "code",
130
+ "execution_count": null,
131
+ "metadata": {},
132
+ "outputs": [],
133
+ "source": [
134
+ "# Title-based features\n",
135
+ "df['title_length'] = df['Title'].str.len()\n",
136
+ "df['title_words'] = df['Title'].str.split().str.len()\n",
137
+ "df['has_colon'] = df['Title'].str.contains(':').astype(int)\n",
138
+ "df['has_series'] = df['Title'].str.contains(r'\\([A-Za-z0-9 ]+\\)').astype(int)\n",
139
+ "\n",
140
+ "fig, axes = plt.subplots(1, 2, figsize=(12, 4))\n",
141
+ "sns.scatterplot(data=df, x='title_words', y='price', hue='rating', ax=axes[0], alpha=0.7, palette='viridis')\n",
142
+ "axes[0].set_title('Price vs Title Length (words)')\n",
143
+ "sns.boxplot(data=df, x='has_colon', y='price', ax=axes[1], palette='Set2')\n",
144
+ "axes[1].set_title('Price by Colon in Title')\n",
145
+ "axes[1].set_xticklabels(['No colon', 'Has colon'])\n",
146
+ "plt.tight_layout()\n",
147
+ "plt.show()"
148
+ ]
149
+ },
150
+ {
151
+ "cell_type": "markdown",
152
+ "metadata": {},
153
+ "source": [
154
+ "## 3. Feature Engineering\n",
155
+ "\n",
156
+ "Combine numeric features with TF-IDF embeddings of the title to give the model text signal."
157
+ ]
158
+ },
159
+ {
160
+ "cell_type": "code",
161
+ "execution_count": null,
162
+ "metadata": {},
163
+ "outputs": [],
164
+ "source": [
165
+ "numeric_feats = ['rating', 'title_length', 'title_words', 'has_colon', 'has_series']\n",
166
+ "X_numeric = df[numeric_feats].fillna(df[numeric_feats].median())\n",
167
+ "y = df['price'].values\n",
168
+ "\n",
169
+ "tfidf = TfidfVectorizer(max_features=300, stop_words='english')\n",
170
+ "X_text = tfidf.fit_transform(df['Title'])\n",
171
+ "svd = TruncatedSVD(n_components=10, random_state=42)\n",
172
+ "X_text_reduced = svd.fit_transform(X_text)\n",
173
+ "\n",
174
+ "X = np.hstack([X_numeric.values, X_text_reduced])\n",
175
+ "print(f'Final feature matrix: {X.shape}')"
176
+ ]
177
+ },
178
+ {
179
+ "cell_type": "markdown",
180
+ "metadata": {},
181
+ "source": [
182
+ "## 4. Cross-Validated Modeling\n",
183
+ "\n",
184
+ "Use 5-fold CV with a Random Forest regressor to predict book price and report honest metrics."
185
+ ]
186
+ },
187
+ {
188
+ "cell_type": "code",
189
+ "execution_count": null,
190
+ "metadata": {},
191
+ "outputs": [],
192
+ "source": [
193
+ "kf = KFold(n_splits=5, shuffle=True, random_state=42)\n",
194
+ "mae_scores, r2_scores = [], []\n",
195
+ "\n",
196
+ "for train_idx, val_idx in kf.split(X):\n",
197
+ " model = RandomForestRegressor(n_estimators=200, random_state=42, n_jobs=-1)\n",
198
+ " model.fit(X[train_idx], y[train_idx])\n",
199
+ " preds = model.predict(X[val_idx])\n",
200
+ " mae_scores.append(mean_absolute_error(y[val_idx], preds))\n",
201
+ " r2_scores.append(r2_score(y[val_idx], preds))\n",
202
+ "\n",
203
+ "print(f'Mean Absolute Error: {np.mean(mae_scores):.2f} £ (±{np.std(mae_scores):.2f})')\n",
204
+ "print(f'R²: {np.mean(r2_scores):.3f} (±{np.std(r2_scores):.3f})')\n",
205
+ "print(f'Baseline MAE (predict median): {np.mean(np.abs(y - np.median(y))):.2f} £')"
206
+ ]
207
+ },
208
+ {
209
+ "cell_type": "markdown",
210
+ "metadata": {},
211
+ "source": [
212
+ "## 5. Feature Importance\n",
213
+ "\n",
214
+ "Which signals matter most for predicting price?"
215
+ ]
216
+ },
217
+ {
218
+ "cell_type": "code",
219
+ "execution_count": null,
220
+ "metadata": {},
221
+ "outputs": [],
222
+ "source": [
223
+ "model = RandomForestRegressor(n_estimators=200, random_state=42, n_jobs=-1)\n",
224
+ "model.fit(X, y)\n",
225
+ "\n",
226
+ "feat_names = numeric_feats + [f'title_topic_{i}' for i in range(10)]\n",
227
+ "importance = pd.Series(model.feature_importances_, index=feat_names).sort_values(ascending=False)\n",
228
+ "\n",
229
+ "plt.figure(figsize=(9, 5))\n",
230
+ "importance.head(12).plot(kind='barh', color='teal')\n",
231
+ "plt.title('Feature Importance (Random Forest)')\n",
232
+ "plt.gca().invert_yaxis()\n",
233
+ "plt.xlabel('Importance')\n",
234
+ "plt.tight_layout()\n",
235
+ "plt.show()\n",
236
+ "print(importance.head(12))"
237
+ ]
238
+ },
239
+ {
240
+ "cell_type": "markdown",
241
+ "metadata": {},
242
+ "source": [
243
+ "## Summary\n",
244
+ "\n",
245
+ "- Cleaned HTML price strings and word-based ratings into numeric features.\n",
246
+ "- Book price has a **wide range** (£0–£100+) with a median around £36; most books sit in a tight mid-range band.\n",
247
+ "- Title features (length, series markers, TF-IDF topics) give only weak predictive signal — the cross-validated model lands near the **median baseline** (MAE ≈ 13.5£ vs 12.5£). This is an honest result: price is largely driven by category/genre, which this scrape does not capture.\n",
248
+ "\n",
249
+ "**Next steps that would help:** adding book category labels, publisher metadata, or a larger catalog would likely unlock real predictive power."
250
+ ]
251
+ }
252
+ ],
253
+ "metadata": {
254
+ "kernelspec": {
255
+ "display_name": "Python 3",
256
+ "language": "python",
257
+ "name": "python3"
258
+ },
259
+ "language_info": {
260
+ "name": "python",
261
+ "version": "3.11.0"
262
+ }
263
+ },
264
+ "nbformat": 4,
265
+ "nbformat_minor": 5
266
+ }