forecast / report.html
bardd's picture
Upload 63 files
450d514 verified
Raw History Blame Contribute Delete
37.8 kB
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<title>Oman Water Demand Forecasting — Daily 28 / Weekly 11 / Weekly 4</title>
<style>
:root{
--ink:#1c2321; --mut:#5b6560; --line:#d9d5cc; --paper:#fbfaf7; --card:#f4f2ec;
--blue:#1f4e79; --orange:#e07b39; --green:#2f7d5d; --red:#b23b3b; --gold:#f0e6c8;
}
*{box-sizing:border-box}
body{margin:0;background:var(--paper);color:var(--ink);
font-family:Georgia,'Iowan Old Style','Times New Roman',serif;line-height:1.7}
.wrap{max-width:920px;margin:0 auto;padding:40px 26px 90px}
.hero{background:var(--blue);color:#fff;border-radius:16px;padding:34px 34px 30px;margin-bottom:8px}
.hero h1{margin:0 0 10px;font-size:32px;line-height:1.25;letter-spacing:.2px}
.hero p{margin:0;color:#d7e3ee;font-size:16px}
.hero .kpis{display:flex;gap:14px;flex-wrap:wrap;margin-top:24px}
.kpi{background:#ffffff14;border:1px solid #ffffff33;border-radius:12px;padding:12px 16px;min-width:150px}
.kpi b{display:block;font-size:26px;line-height:1.1}
.kpi span{font-size:12.5px;color:#cfdded}
h2{font-size:23px;margin:52px 0 6px;padding-top:8px}
h2 .num{display:inline-block;background:var(--blue);color:#fff;border-radius:8px;
font-size:14px;padding:2px 10px;margin-right:10px;vertical-align:2px;font-family:Helvetica,Arial,sans-serif}
h3{font-size:17.5px;margin:30px 0 4px;color:#20302b}
h4{font-size:15px;margin:22px 0 2px;color:#20302b}
p{margin:10px 0}
.mut{color:var(--mut)}
.small{font-size:13.5px}
.card{background:var(--card);border:1px solid var(--line);border-radius:14px;padding:18px 20px;margin:18px 0}
.callout{border-left:5px solid var(--green);background:#eef5f1;border-radius:0 12px 12px 0;padding:14px 18px;margin:18px 0}
.callout.warn{border-left-color:var(--orange);background:#fdf1e7}
.callout.stop{border-left-color:var(--red);background:#fbeeee}
.callout b:first-child{display:block;margin-bottom:4px;font-family:Helvetica,Arial,sans-serif;font-size:12.5px;
letter-spacing:.9px;text-transform:uppercase;color:var(--green)}
.callout.warn b:first-child{color:#b35c17}
.callout.stop b:first-child{color:var(--red)}
table{border-collapse:collapse;width:100%;margin:14px 0;font-size:14.5px}
th,td{border:1px solid #c9c4b8;padding:7px 10px;text-align:left}
th{background:#eae6dc;font-family:Helvetica,Arial,sans-serif;font-size:12.5px;letter-spacing:.4px;text-transform:uppercase}
td.n,th.n{text-align:right;font-variant-numeric:tabular-nums}
tr:nth-child(even) td{background:#ffffff}
img{max-width:100%;border:1px solid var(--line);border-radius:12px;margin:14px 0;background:#fff}
figure{margin:20px 0}figcaption{font-size:13px;color:var(--mut);margin-top:-6px}
code{background:#efece4;border:1px solid #e2ddd0;padding:1px 6px;border-radius:6px;font-size:13.5px}
pre{background:#22282a;color:#e8eae6;padding:16px;border-radius:12px;overflow-x:auto;font-size:13px;line-height:1.55}
.svgbox{text-align:center;margin:18px 0;overflow-x:auto}
.toc{columns:2;column-gap:34px;margin:16px 0;font-size:14.5px}
.toc div{break-inside:avoid;padding:2px 0}
.toc a{color:var(--blue);text-decoration:none}
.toc a:hover{text-decoration:underline}
.tag{display:inline-block;font-family:Helvetica,Arial,sans-serif;font-size:11.5px;background:#e4e0d5;
border-radius:20px;padding:2px 10px;color:#4b5a52;margin-right:6px}
.lead{font-size:17px;color:#3d4a44}
hr{border:0;border-top:1px solid var(--line);margin:34px 0}
.footer{margin-top:50px;font-size:13px;color:var(--mut)}
.big-good{color:var(--green);font-weight:bold}
.big-bad{color:var(--red);font-weight:bold}
</style>
</head>
<body><div class="wrap">
<div class="hero">
<h1>Oman Water Demand Forecasting<br>Daily 28 &middot; Weekly 11 &middot; Weekly 4</h1>
<p>A review pack you can run yourself. Two bottle sizes, five main sales regions, one small
model per region, plain-English method, honest numbers, the exact settings we used.</p>
<div class="kpis">
<div class="kpi"><b>97.6%</b><span>200ML weekly accuracy (11 weeks)</span></div>
<div class="kpi"><b>95.0%</b><span>330ML weekly accuracy (11 weeks)</span></div>
<div class="kpi"><b>78.9%</b><span>200ML daily accuracy (28 days)</span></div>
<div class="kpi"><b>77.4%</b><span>330ML daily accuracy (28 days)</span></div>
</div>
</div>
<p class="small mut">Accuracy = 100 × (1 − WAPE). Every number in this report comes from a real
out-of-sample holdout, not from training data. DUQM and PDO are forecast but excluded from the
headline score — Chapter 9 explains why, with numbers.</p>
<div class="card">
<b>Contents</b>
<div class="toc">
<div><a href="#c1">1 · The question we were asked</a></div>
<div><a href="#c2">2 · The data we were given</a></div>
<div><a href="#c3">3 · How missing days were filled (the policy)</a></div>
<div><a href="#c4">4 · How one forecast is made</a></div>
<div><a href="#c5">5 · Daily results and the daily ceiling</a></div>
<div><a href="#c6">6 · Weekly results</a></div>
<div><a href="#c7">7 · Why weekly beats daily</a></div>
<div><a href="#c8">8 · What each forecast is for (business)</a></div>
<div><a href="#c9">9 · Why DUQM and PDO sit outside the headline</a></div>
<div><a href="#c10">10 · Reproduce everything yourself</a></div>
<div><a href="#c11">11 · Glossary and straight answers</a></div>
</div>
</div>
<h2 id="c1"><span class="num">1</span>The question we were asked</h2>
<p class="lead">A water bottler in Oman sells two products — 24×330ML packs and 30×200ML packs —
across 12 sales areas grouped into 7 regions. The factory, the trucks, and the warehouses all need
to know <b>how much water will be sold next, and where</b>. We were asked two things:</p>
<ol>
<li><b>Daily:</b> predict sales for each area for the next 28 days. Trucks load every morning.
Target customers: distribution and drivers.</li>
<li><b>Weekly:</b> predict sales for each region for the coming weeks. Production runs, raw water,
bottles, and warehouse space are planned weekly. Target customers: plant and supply chain.</li>
</ol>
<div class="callout"><b>Key idea</b>
A forecast is only useful if its accuracy is honest — measured on days the model has never seen.
Everything in this pack is scored on a holdout at the very end of the history, and all the code
that produced these numbers is in this folder so a reviewer can rerun it and get the same result.</div>
<h2 id="c2"><span class="num">2</span>The data we were given</h2>
<table>
<tr><th>Product</th><th class="n">Rows</th><th class="n">Areas</th><th class="n">Regions</th><th>History</th></tr>
<tr><td>24 × 330ML packs</td><td class="n">12,568</td><td class="n">12</td><td class="n">7</td><td>1 Jan 2021 → 24 Jan 2025</td></tr>
<tr><td>30 × 200ML packs</td><td class="n">14,583</td><td class="n">11</td><td class="n">7</td><td>1 Jan 2021 → 29 Jan 2025</td></tr>
</table>
<p>Each row is one day, one area, one product: the day, the area, the region, and the cases sold.
Four things about this file shaped every decision that followed.</p>
<h3>2.1 There is not a single zero in the file</h3>
<p>Across more than 27,000 rows, the smallest sale is 1 case. The source system only writes a row
when something was sold. So a missing day usually means <b>no sale happened</b> — not that the data
was lost. That single observation drives the filling policy in Chapter 3.</p>
<h3>2.2 Some days are missing everywhere at once</h3>
<p>20–23 July 2021 (Eid al-Adha) is missing in almost every area. 3 October 2021 (Cyclone Shaheen
emergency holiday) is missing in several. Those are real closures, and the pattern repeats across
both products — strong evidence they are genuine no-sale days, not broken records.</p>
<h3>2.3 Friday means something different in every area</h3>
<table>
<tr><th>Area</th><th class="n">Fridays in history</th><th class="n">Fridays with no sale</th><th class="n">Share</th></tr>
<tr><td>IBRA</td><td class="n">211</td><td class="n">197</td><td class="n">93%</td></tr>
<tr><td>IBRI</td><td class="n">211</td><td class="n">158</td><td class="n">75%</td></tr>
<tr><td>SUR</td><td class="n">212</td><td class="n">79</td><td class="n">37%</td></tr>
<tr><td>NIZWA</td><td class="n">212</td><td class="n">72</td><td class="n">34%</td></tr>
<tr><td>SAHAM</td><td class="n">213</td><td class="n">32</td><td class="n">15%</td></tr>
<tr><td>AMERAT</td><td class="n">213</td><td class="n">1</td><td class="n">0.5%</td></tr>
</table>
<p>In IBRA, Friday is effectively a closed day. In AMERAT, Friday is a normal selling day. No single
"Friday = zero" rule can be true for both. The model has to learn <b>area × weekday</i></b>, which is
exactly what the features in Chapter 4 allow it to do.</p>
<h3>2.4 One stretch of October 2021 looks like a data outage, not a holiday</h3>
<p>11–15 October 2021: many normally busy 330ML areas (AMERAT, RUSAYL, MUSANNAH, SALALAH, SAHAM)
go silent at the same time, while 200ML keeps selling. There is no public holiday in that window.
Seventeen values are affected — a tiny number — and they are handled separately from real zero days
(Chapter 3), with the assumption flagged in the data.</p>
<h2 id="c3"><span class="num">3</span>How missing days were filled (the policy)</h2>
<p class="lead">One rule for every blank would be wrong. A blank can mean "nobody ordered that day"
(truly zero) or "the system lost that day" (unknown, not zero). Those two need opposite treatment.
The policy below sorts every blank into one of four buckets.</p>
<div class="svgbox">
<svg width="820" height="200" viewBox="0 0 820 200" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12">
<rect x="10" y="80" width="150" height="44" rx="10" fill="#1f4e79"/>
<text x="85" y="99" fill="#fff" text-anchor="middle">Build a full calendar</text>
<text x="85" y="115" fill="#cfe0ee" text-anchor="middle" font-size="10.5">first sale → last sale, per area</text>
<rect x="205" y="18" width="185" height="52" rx="10" fill="#fff" stroke="#2f7d5d" stroke-width="2"/>
<text x="297" y="40" text-anchor="middle" fill="#2f7d5d">Shared closure day?</text>
<text x="297" y="57" text-anchor="middle" font-size="11">Eid · Cyclone Shaheen → <tspan font-weight="bold">0</tspan></text>
<rect x="205" y="86" width="185" height="52" rx="10" fill="#fff" stroke="#e07b39" stroke-width="2"/>
<text x="297" y="108" text-anchor="middle" fill="#b35c17">Oct 11–15 2021 outage?</text>
<text x="297" y="125" text-anchor="middle" font-size="11">same-weekday median, weight 0.25</text>
<rect x="205" y="154" width="185" height="40" rx="10" fill="#fff" stroke="#5b6560" stroke-width="2"/>
<text x="297" y="179" text-anchor="middle" fill="#3d4a44">Any other blank inside the span → <tspan font-weight="bold">0</tspan></text>
<rect x="440" y="80" width="180" height="44" rx="10" fill="#eef5f1" stroke="#2f7d5d"/>
<text x="530" y="99" text-anchor="middle">Flag every guess</text>
<text x="530" y="115" text-anchor="middle" font-size="10.5">was_observed · is_imputed · weight</text>
<rect x="665" y="80" width="145" height="44" rx="10" fill="#1f4e79"/>
<text x="737" y="99" fill="#fff" text-anchor="middle">Ready to train</text>
<text x="737" y="115" fill="#cfe0ee" text-anchor="middle" font-size="10.5">guesses never scored</text>
<path d="M160 102 H205" stroke="#555" stroke-width="2" marker-end="url(#a)"/>
<path d="M390 44 H417 V92 H440" stroke="#555" stroke-width="2" fill="none" marker-end="url(#a)"/>
<path d="M390 112 H440" stroke="#555" stroke-width="2" marker-end="url(#a)"/>
<path d="M390 174 H417 V122 H440" stroke="#555" stroke-width="2" fill="none" marker-end="url(#a)"/>
<path d="M620 102 H665" stroke="#555" stroke-width="2" marker-end="url(#a)"/>
<defs><marker id="a" markerWidth="9" markerHeight="9" refX="7" refY="3" orient="auto">
<path d="M0,0 L7,3 L0,6 z" fill="#555"/></marker></defs>
</svg>
</div>
<table>
<tr><th>Situation</th><th>What we write into the data</th><th>Audit flag</th></tr>
<tr><td>Real sale (row exists in the source)</td><td>Volume as-is, never touched</td><td><code>was_observed = 1</code></td></tr>
<tr><td>Known shared closure (Eid, cyclone)</td><td>Volume = 0</td><td><code>is_closure = 1</code></td></tr>
<tr><td>330ML Oct 11–15 2021 outage (17 values)</td><td>Median of the same weekday ±4 weeks</td><td><code>is_imputed = 1</code>, weight 0.25</td></tr>
<tr><td>Any other blank inside an area's active span</td><td>Volume = 0 (likely no-sale day)</td><td><code>imputation_method = zero_no_sale</code></td></tr>
<tr><td>Before the first sale / after the last sale</td><td>No row created</td><td>inactive period, not a zero</td></tr>
</table>
<h3>The 330ML outage estimate, worked out</h3>
<p>Take AMERAT on Monday 11 October 2021. The surrounding Mondays in the source file sold
716, 854, 833, 1732, 690, 912, 837, 653 cases. The median is <b>835</b>. That 835 is written in,
marked as an estimate, and given only a quarter of the training weight of a real row. Seventeen
values were filled this exact way.</p>
<div class="card">
<b>Filled calendars — the result of the policy</b>
<table>
<tr><th>Product</th><th class="n">Total rows</th><th class="n">Real sales</th><th class="n">No-sale zeros</th><th class="n">Closure zeros</th><th class="n">Estimated</th></tr>
<tr><td>200ML</td><td class="n">16,378</td><td class="n">14,583</td><td class="n">1,750</td><td class="n">45</td><td class="n">0</td></tr>
<tr><td>330ML</td><td class="n">17,697</td><td class="n">12,568</td><td class="n">5,062</td><td class="n">50</td><td class="n">17</td></tr>
</table>
<p class="small mut">That is 93% real data in the 200ML file and 71% in the 330ML file. The
330ML file is sparser because four of its areas sell very little (Chapter 9).</p>
</div>
<figure><img src="charts/policy_split.png" alt="How the filled calendar splits between real rows and filled rows">
<figcaption>Left: how the 200ML calendar is built. Right: how the 330ML calendar is built. Orange is
the only guessed content in the whole project, and it is 17 values out of 17,697.</figcaption></figure>
<div class="callout warn"><b>The most important courtesy in this pack</b>
Every filled row keeps its flag. When we score the model we use <b>only observed rows</b>
(<code>is_imputed = 0</code>). A guess can help training; a guess is never counted as truth when
measuring accuracy.</div>
<h2 id="c4"><span class="num">4</span>How one forecast is made</h2>
<p class="lead">The model is a small decision-tree ensemble called <b>LightGBM</b>, the same family
of method that won the M5 retail forecasting competition. It is not a neural network and not deep
learning. We train <b>one model per region per product</b> — ten small models in total in this pack
— because each region has its own habits.</p>
<div class="svgbox">
<svg width="820" height="230" viewBox="0 0 820 230" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12">
<rect x="8" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/>
<text x="96" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">History features</text>
<text x="96" y="58" text-anchor="middle" font-size="10.5">sales 1, 7, 14, 28, 56 days ago</text>
<text x="96" y="76" text-anchor="middle" font-size="10.5">rolling mean / spread, 7 &amp; 28 days</text>
<text x="96" y="94" text-anchor="middle" font-size="10.5">how often it sells at all</text>
<text x="96" y="112" text-anchor="middle" font-size="10.5">days since last sale</text>
<rect x="212" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/>
<text x="300" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">Calendar features</text>
<text x="300" y="58" text-anchor="middle" font-size="10.5">weekday, day, month, week</text>
<text x="300" y="76" text-anchor="middle" font-size="10.5">weekend / Friday flags</text>
<text x="300" y="94" text-anchor="middle" font-size="10.5">closure and event flags</text>
<text x="300" y="112" text-anchor="middle" font-size="10.5">year trend</text>
<rect x="416" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/>
<text x="504" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">Who / where</text>
<text x="504" y="58" text-anchor="middle" font-size="10.5">area name</text>
<text x="504" y="76" text-anchor="middle" font-size="10.5">region name</text>
<text x="504" y="94" text-anchor="middle" font-size="10.5">(given as categories, so the</text>
<text x="504" y="112" text-anchor="middle" font-size="10.5">model learns area habits)</text>
<rect x="626" y="52" width="184" height="124" rx="12" fill="#1f4e79"/>
<text x="718" y="92" fill="#fff" text-anchor="middle" font-weight="bold">One model,</text>
<text x="718" y="110" fill="#fff" text-anchor="middle" font-weight="bold">one region,</text>
<text x="718" y="128" fill="#fff" text-anchor="middle" font-weight="bold">one product</text>
<text x="718" y="152" fill="#cfe0ee" text-anchor="middle" font-size="10.5">predicts cases for each day</text>
<path d="M184 114 H212" stroke="#555" stroke-width="2" marker-end="url(#b)"/>
<path d="M388 114 H416" stroke="#555" stroke-width="2" marker-end="url(#b)"/>
<path d="M592 114 H626" stroke="#555" stroke-width="2" marker-end="url(#b)"/>
<defs><marker id="b" markerWidth="9" markerHeight="9" refX="7" refY="3" orient="auto">
<path d="M0,0 L7,3 L0,6 z" fill="#555"/></marker></defs>
</svg>
</div>
<h3>4.1 The rule that keeps the score honest</h3>
<p>All history features are built with a one-day shift — when the model predicts Tuesday, it is
allowed to see Monday and earlier, never Tuesday itself. And the last 28 days of history are held
out completely: the model never trains on them. The score in this report is only ever measured on
those unseen days.</p>
<h3>4.2 What "accuracy" means here (WAPE, in one minute)</h3>
<div class="card">
<p>We measure error with <b>WAPE</b> — the total absolute miss divided by the total actual volume.</p>
<pre>Actual week: 1,000 cases
Forecast: 950 cases
Miss: 50 cases
WAPE = 50 / 1000 = 5% Accuracy = 100% − 5% = 95%</pre>
<p class="mut small">WAPE is a percentage of the whole, so it does not punish big regions and small
regions equally per case — it asks "of all the water we shipped, how much did we misjudge?" That
matches how the business thinks about a plan.</p>
</div>
<h3>4.3 The exact settings used (winners only)</h3>
<p>Each region's best settings are stored in <code>params/best_params_200ML.json</code> and
<code>params/best_params_330ML.json</code> so the reviewer runs exactly what produced these numbers.</p>
<table>
<tr><th>Product / region</th><th>Model</th><th class="n">Learning rate</th><th class="n">Leaves</th><th class="n">Tweedie power</th></tr>
<tr><td>200ML CAPITAL, Batinah, Sharqiyah</td><td>Tweedie</td><td class="n">0.05 – 0.11</td><td class="n">511</td><td class="n">1.05 – 1.18</td></tr>
<tr><td>200ML Dhofar, Al Dakhiliyah</td><td>Tweedie</td><td class="n">0.03 – 0.07</td><td class="n">127 – 511</td><td class="n">1.14 – 1.49</td></tr>
<tr><td>330ML CAPITAL</td><td>Poisson</td><td class="n">0.10</td><td class="n">63</td><td class="n">—</td></tr>
<tr><td>330ML Batinah, Dhofar, Sharqiyah, Al Dakhiliyah</td><td>Tweedie</td><td class="n">0.02 – 0.11</td><td class="n">127 – 511</td><td class="n">1.29 – 1.53</td></tr>
</table>
<p class="small mut">Tweedie and Poisson are count-friendly objectives: they suit demand data with
many small days and occasional big ones, and cannot predict negative cases. The model stops
training automatically when the last 28 days of training stop improving (early stopping).</p>
<h2 id="c5"><span class="num">5</span>Daily results and the daily ceiling</h2>
<p class="lead">Daily accuracy for the next 28 days, measured on unseen days, five regions:</p>
<table>
<tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr>
<tr><td>CAPITAL</td><td class="n">86.6%</td><td class="n">84.2%</td></tr>
<tr><td>Dhofar</td><td class="n">78.8%</td><td class="n">76.1%</td></tr>
<tr><td>Batinah</td><td class="n">76.2%</td><td class="n">69.9%</td></tr>
<tr><td>Sharqiyah</td><td class="n">63.0%</td><td class="n">45.4%</td></tr>
<tr><td>Al Dakhiliyah</td><td class="n">59.6%</td><td class="n">57.1%</td></tr>
<tr><td><b>All five together</b></td><td class="n"><b>78.9%</b></td><td class="n"><b>77.4%</b></td></tr>
</table>
<figure><img src="charts/daily28_200ML.png" alt="200ML daily 28-day forecast vs actual">
<figcaption>200ML — one panel per region, because each region runs its own plant and trucks.
Each panel draws that region's daily total; the percentage in its title is that region's own
daily accuracy, the same number as in the table above.</figcaption></figure>
<figure><img src="charts/daily28_330ML.png" alt="330ML daily 28-day forecast vs actual">
<figcaption>330ML — same per-region view for the other product. Sharqiyah's panel shows the
hardest region honestly at 45.4%: no country total hides it.</figcaption></figure>
<h3>5.1 Where the daily error comes from</h3>
<figure><img src="charts/horizon_200ML.png" alt="200ML daily error per day of the holdout">
<figcaption>200ML — error per day of the holdout. The spikes sit on unusual dates, not on far-away
dates. In this holdout the model still sees real history for every day (a "re-forecast each
morning" view); in a full 28-day rollout the far days would also be building on predicted history,
which would add more error, not less.</figcaption></figure>
<figure><img src="charts/horizon_330ML.png" alt="330ML daily error per day of the holdout">
<figcaption>330ML — same pattern. Day 1 is sometimes worse than day 28, which proves the point:
daily error is driven by <i>what that particular day was</i>, not by how far ahead it is.</figcaption></figure>
<figure><img src="charts/daily_error_hist.png" alt="Distribution of daily errors per row">
<figcaption>Half of all daily rows land within the orange line (the median error). The long right
tail is the price of a single unusual day.</figcaption></figure>
<div class="callout stop"><b>Why daily cannot reach 5% error with this data</b>
<ol>
<li><b>A single day is one lumpy decision.</b> A shop that orders 0 on Monday and 500 on Tuesday
makes the day-to-day line jump. History cannot know which Tuesday the truck arrives; the model
splits the difference and takes the error.</li>
<li><b>Friday has five different meanings.</b> Closed in IBRA, normal in AMERAT. The model learns
this, but any single surprise Friday is a large miss on a single day.</li>
<li><b>Holidays stop trucks, not thirst.</b> Eid and cyclone days are zero <i>sales</i> but not
zero demand. And nothing in the file explains the biggest jumps — no promotions, no price, no
stock-out records, no delivery routes.</li>
<li><b>Small denominators explode percentages.</b> Missing 30 cases in an area that sells 60 that
day is a 50% error for that row, even though the business barely notices 30 cases.</li>
<li><b>Four years is a short teacher.</b> Ramadan drifts about 11 days every year, so each area's
Ramadan pattern has been seen only four times. That is thin evidence for such a strong event.</li>
</ol>
Weather (temperature, humidity, rain) was also tested and added no measurable lift, so the final
models do not use it.</div>
<h2 id="c6"><span class="num">6</span>Weekly results</h2>
<p class="lead">The same ten models, the same forecasts — simply added up over Monday-to-Sunday
weeks before scoring. Two horizons were tested: the last 11 full weeks, and the last 4 weeks.</p>
<h3>6.1 Eleven weeks</h3>
<table>
<tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr>
<tr><td>CAPITAL</td><td class="n">97.3%</td><td class="n">93.9%</td></tr>
<tr><td>Batinah</td><td class="n">95.7%</td><td class="n">89.6%</td></tr>
<tr><td>Dhofar</td><td class="n">91.5%</td><td class="n">91.3%</td></tr>
<tr><td>Sharqiyah</td><td class="n">90.5%</td><td class="n">87.8%</td></tr>
<tr><td>Al Dakhiliyah</td><td class="n">84.7%</td><td class="n">84.0%</td></tr>
<tr><td><b>All five together</b></td><td class="n"><b>97.6%</b></td><td class="n"><b>95.0%</b></td></tr>
</table>
<figure><img src="charts/weekly11_200ML.png" alt="200ML weekly 11 weeks forecast vs actual">
<figcaption>200ML — weekly totals, 11 weeks. Percentages above each pair are that week's error
(actual vs forecast).</figcaption></figure>
<figure><img src="charts/weekly11_330ML.png" alt="330ML weekly 11 weeks forecast vs actual">
<figcaption>330ML — weekly totals, 11 weeks. Even the harder product lands on the bar most
weeks.</figcaption></figure>
<h3>6.2 The last four weeks</h3>
<p>Same models, retrained on history ending four weeks ago — a check that the accuracy is not a
lucky artifact of one long window.</p>
<table>
<tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr>
<tr><td>CAPITAL</td><td class="n">97.0%</td><td class="n">92.8%</td></tr>
<tr><td>Batinah</td><td class="n">94.2%</td><td class="n">89.2%</td></tr>
<tr><td>Dhofar</td><td class="n">92.5%</td><td class="n">88.7%</td></tr>
<tr><td>Sharqiyah</td><td class="n">90.8%</td><td class="n">88.4%</td></tr>
<tr><td>Al Dakhiliyah</td><td class="n">82.7%</td><td class="n">92.0%</td></tr>
<tr><td><b>All five together</b></td><td class="n"><b>96.4%</b></td><td class="n"><b>94.8%</b></td></tr>
</table>
<figure><img src="charts/weekly4_200ML.png" alt="200ML weekly last 4 weeks forecast vs actual">
<figcaption>200ML — last four weeks, week by week.</figcaption></figure>
<figure><img src="charts/weekly4_330ML.png" alt="330ML weekly last 4 weeks forecast vs actual">
<figcaption>330ML — last four weeks. The four-week view is the closest to live conditions.</figcaption></figure>
<figure><img src="charts/ladder_200ML.png" alt="200ML daily vs weekly accuracy per region">
<figcaption>200ML — the same models, two clocks. Weekly clears or approaches the 95% line in every
region except Al Dakhiliyah.</figcaption></figure>
<figure><img src="charts/ladder_330ML.png" alt="330ML daily vs weekly accuracy per region">
<figcaption>330ML — the same story. Nothing about the model changed between the red and green
bars; only the window the numbers are added over changed.</figcaption></figure>
<h2 id="c7"><span class="num">7</span>Why weekly beats daily</h2>
<p class="lead">This is the single most important idea in the pack, and it is not a model trick —
it is arithmetic.</p>
<div class="svgbox">
<svg width="820" height="240" viewBox="0 0 820 240" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12">
<text x="16" y="24" font-weight="bold" fill="#1f4e79">One week of one area — actual vs forecast</text>
<g stroke="#c9c4b8"><line x1="60" y1="70" x2="60" y2="150"/><line x1="140" y1="70" x2="140" y2="150"/>
<line x1="220" y1="70" x2="220" y2="150"/><line x1="300" y1="70" x2="300" y2="150"/>
<line x1="380" y1="70" x2="380" y2="150"/><line x1="460" y1="70" x2="460" y2="150"/>
<line x1="540" y1="70" x2="540" y2="150"/></g>
<g fill="#1f4e79">
<rect x="46" y="126" width="28" height="24"/><rect x="126" y="82" width="28" height="68"/>
<rect x="206" y="120" width="28" height="30"/><rect x="286" y="96" width="28" height="54"/>
<rect x="366" y="128" width="28" height="22"/><rect x="446" y="88" width="28" height="62"/>
<rect x="526" y="132" width="28" height="18"/>
</g>
<g fill="#e07b39">
<rect x="76" y="118" width="16" height="32"/><rect x="156" y="110" width="16" height="40"/>
<rect x="236" y="116" width="16" height="34"/><rect x="316" y="112" width="16" height="38"/>
<rect x="396" y="120" width="16" height="30"/><rect x="476" y="114" width="16" height="36"/>
<rect x="556" y="122" width="16" height="28"/>
</g>
<g font-size="10.5" fill="#5b6560">
<text x="60" y="166" text-anchor="middle">Mon</text><text x="140" y="166" text-anchor="middle">Tue</text>
<text x="220" y="166" text-anchor="middle">Wed</text><text x="300" y="166" text-anchor="middle">Thu</text>
<text x="380" y="166" text-anchor="middle">Fri</text><text x="460" y="166" text-anchor="middle">Sat</text>
<text x="540" y="166" text-anchor="middle">Sun</text>
</g>
<rect x="620" y="70" width="184" height="80" rx="12" fill="#eef5f1" stroke="#2f7d5d"/>
<text x="712" y="94" text-anchor="middle" font-size="11.5">Day 2: forecast 40,</text>
<text x="712" y="110" text-anchor="middle" font-size="11.5">actual 68 — a 41% miss.</text>
<text x="712" y="132" text-anchor="middle" font-size="11.5" fill="#2f7d5d">Week total: 322 vs 325</text>
<text x="712" y="146" text-anchor="middle" font-size="11.5" fill="#2f7d5d">— only a 1% miss.</text>
<rect x="16" y="186" width="788" height="40" rx="10" fill="#f4f2ec" stroke="#c9c4b8"/>
<text x="34" y="211" font-size="12">Daily: every bar is judged alone, so every wobble is an error. Weekly: the errors point in opposite</text>
<text x="34" y="227" font-size="12">directions, cancel out, and only the level that matters to the business survives.</text>
</svg>
</div>
<p>Over a week, a slow Monday is usually paid back by a busy Tuesday. Adding the days together
before comparing means the random part of the wobble cancels while the real level remains. The
model was never bad at the <i>level</i> of demand — it was bad at guessing the exact <i>day</i>.
The weekly number uses what the model is actually good at.</p>
<div class="callout"><b>In one sentence</b>
Daily asks "will it be 120 or 160 cases on the 14th?" — weekly asks "will this region move about
1,000 cases this week?" The second question matches both the model's strength and the way a
factory actually plans.</div>
<h2 id="c8"><span class="num">8</span>What each forecast is for (business)</h2>
<table>
<tr><th>Decision</th><th>Clock to use</th><th>Who uses it</th><th>Why that clock</th></tr>
<tr><td>How much to produce next week (raw water, bottles, line time)</td><td><b>Weekly</b></td><td>Plant, production planning</td><td>Runs are batched weekly; a day-level miss is irrelevant to a weekly run. 95–98% is enough to plan materials and shifts.</td></tr>
<tr><td>Warehouse space and stock transfer between regions</td><td><b>Weekly</b></td><td>Supply chain</td><td>Replenishment lead times are measured in weeks; weekly totals set safety stock.</td></tr>
<tr><td>Truck capacity and fleet allocation per region</td><td><b>Weekly</b></td><td>Logistics planning</td><td>Fleet size is re-planned weekly, not daily.</td></tr>
<tr><td>Which shops and areas each truck serves tomorrow</td><td><b>Daily</b></td><td>Distribution, drivers</td><td>A driver needs a number per area per day to load the right quantity; 78% daily accuracy is the honest guidance range, not a promise per shop.</td></tr>
<tr><td>Short shelf-life urgent replenishment</td><td><b>Daily</b></td><td>Area supervisors</td><td>Daily view shows the next few days; supervisors adjust with local knowledge.</td></tr>
</table>
<div class="callout warn"><b>How to read daily numbers honestly</b>
Quoted daily accuracy is the share of total volume forecast correctly across a region — not the
chance that any one shop's day is exact. Use daily forecasts to plan the size of tomorrow's load,
and weekly forecasts to plan what the factory makes. If a decision costs real money, make it on
the weekly clock and use the daily clock only to distribute it.</div>
<h2 id="c9"><span class="num">9</span>Why DUQM and PDO sit outside the headline</h2>
<p class="lead">DUQM, DUQUM PDO and PDO are real customers, and the models do forecast them. They
are left out of the headline accuracy for one reason: their history is too thin to measure with a
percentage fairness.</p>
<table>
<tr><th>Area</th><th>Product</th><th class="n">Days with sales</th><th class="n">Active days</th><th class="n">Coverage</th></tr>
<tr><td>PDO</td><td>200ML</td><td class="n">150</td><td class="n">1,480</td><td class="n">10%</td></tr>
<tr><td>DUQM</td><td>330ML</td><td class="n">362</td><td class="n">1,427</td><td class="n">25%</td></tr>
<tr><td>DUQUM PDO</td><td>330ML</td><td class="n">86</td><td class="n">1,445</td><td class="n">6%</td></tr>
</table>
<p>When a region sells in only a handful of days, a weekly percentage error is divided by a tiny
number. Predicting 40 cases in a week when the truth is 10 is reported as a 300% error — a scary
number for a 30-case miss. Their absolute misses are small: 200ML PDO's typical daily miss is
about 3 cases. That is why the pack reports them separately, with the honest note that a
percentage cannot describe them well.</p>
<div class="card"><b>The effect on the headline, stated openly</b>
<table>
<tr><th>Headline</th><th class="n">200ML</th><th class="n">330ML</th></tr>
<tr><td>Weekly 11-week accuracy, five main regions</td><td class="n">97.6%</td><td class="n">95.0%</td></tr>
<tr><td>Same, with DUQM and PDO added back</td><td class="n">97.6%</td><td class="n">93.5%</td></tr>
</table>
<p class="small mut">200ML barely moves because DUQM and PDO sell almost nothing there. 330ML
drops 1.5 points because two sparse sites with wild percentages are enough to pull the average
down. Nothing about the other five regions changed.</p>
</div>
<h2 id="c10"><span class="num">10</span>Reproduce everything yourself</h2>
<p class="lead">CPU only. No API keys, no internet, no special hardware. From this folder:</p>
<pre>pip install -r requirements.txt
python3 daily28/train.py &amp;&amp; python3 daily28/predict.py # daily 28-day holdout
python3 weekly11/train.py &amp;&amp; python3 weekly11/predict.py # 11-week window
python3 weekly4/train.py &amp;&amp; python3 weekly4/predict.py # last 4 weeks
python3 make_charts.py # rebuild every chart</pre>
<table>
<tr><th>Path</th><th>What it is</th></tr>
<tr><td><code>data/*_daily_calendar_filled.csv</code></td><td>The filled calendars with every flag intact (<code>was_observed</code>, <code>is_imputed</code>, <code>training_weight</code>)</td></tr>
<tr><td><code>params/best_params_*.json</code></td><td>The exact settings that produced the numbers in this report</td></tr>
<tr><td><code>daily28/</code>, <code>weekly11/</code>, <code>weekly4/</code></td><td>train + predict scripts, saved models</td></tr>
<tr><td><code>predictions/</code></td><td>Row-level actual-vs-forecast CSVs you can audit or chart</td></tr>
<tr><td><code>charts/</code></td><td>Every figure in this report, regenerated from the predictions</td></tr>
</table>
<div class="callout"><b>Rerun creates the same numbers</b>
The scripts fix the random seed and use the identical data and settings, so a reviewer on any
machine gets the same accuracy figures shown here, up to tiny library-version differences.</div>
<h2 id="c11"><span class="num">11</span>Glossary and straight answers</h2>
<table>
<tr><th>Term</th><th>Meaning in this pack</th></tr>
<tr><td>WAPE</td><td>Total absolute error ÷ total actual volume. So "5% WAPE" means "we misjudged 5% of the water shipped".</td></tr>
<tr><td>Accuracy</td><td>100% − WAPE. Used everywhere in this report.</td></tr>
<tr><td>Holdout</td><td>The last days of history, removed from training, used only to score.</td></tr>
<tr><td>Lag</td><td>Sales from N days ago, e.g. lag_7 = the same weekday last week.</td></tr>
<tr><td>Rolling mean</td><td>The average of the previous N days — a smooth memory of recent level.</td></tr>
<tr><td>Tweedie / Poisson</td><td>Objective functions for count data (cases sold). They fit demand with many small days and occasional large ones and can never predict negative sales.</td></tr>
<tr><td>LightGBM</td><td>A decision-tree ensemble; the same family that won the M5 retail forecasting competition.</td></tr>
<tr><td>Early stopping</td><td>The model keeps a small internal check window and stops improving when it starts to memorize.</td></tr>
</table>
<h3>Straight answers</h3>
<p><b>Why not deep learning?</b> Four years of daily data across a handful of areas is small. Tree
ensembles have repeatedly matched or beaten neural networks on this kind of tabular, intermittent
demand — including in the M5 competition this project takes its recipe from. Neural networks would
add complexity without evidence of gain.</p>
<p><b>Can daily accuracy ever reach 95%?</b> Not with this data. To get there the model would need
to know the events that actually move a single day: promotions and prices, out-of-stock days,
delivery schedules, and a confirmed closure calendar. With those inputs a daily model could
improve; without them it is guessing at the noisiest part of the signal. Weekly already crosses
95% for four of five regions without any of that.</p>
<p><b>Does weather help?</b> It was tested properly (temperature, humidity, rain, heat flags from
ERA5 reanalysis). It produced no measurable lift on either product, so it was left out to keep the
models simple. It can be revisited when more history exists.</p>
<p><b>What should we do next?</b> Two things pay first: get the promotion and stock-out history
from the business to attack the daily error where it is largest, and keep the weekly forecast as
the production number, since weekly accuracy is already at the level the plant needs.</p>
<div class="footer">
<p>Method summary: region-specialist LightGBM (Tweedie, Poisson for 330ML CAPITAL) · calendar +
shift-safe lags/rolling features · early stopping on the last 28 training days · scored on observed
rows only (<code>is_imputed = 0</code>) · daily holdout = last 28 days · weekly holdouts = last 11
full Mon–Sun weeks and last 4 weeks · DUQM and PDO reported separately, never in the headline.
Weather tested (ERA5) with no lift. Deep learning not used.</p>
</div>
</div></body></html>