Download report.html from bardd/forecast: direct link, hf CLI and curl.
- Browser
- Download file 37.8 kB
-
https://huggingface.co/spaces/bardd/forecast/resolve/main/report.html
- Command line
-
hf download hf://spaces/bardd/forecast/report.html
-
curl -L -o report.html https://huggingface.co/spaces/bardd/forecast/resolve/main/report.html
37.8 kB
| <html lang="en"> | |
| <head> | |
| <meta charset="utf-8"> | |
| <meta name="viewport" content="width=device-width,initial-scale=1"> | |
| <title>Oman Water Demand Forecasting — Daily 28 / Weekly 11 / Weekly 4</title> | |
| <style> | |
| :root{ | |
| --ink:#1c2321; --mut:#5b6560; --line:#d9d5cc; --paper:#fbfaf7; --card:#f4f2ec; | |
| --blue:#1f4e79; --orange:#e07b39; --green:#2f7d5d; --red:#b23b3b; --gold:#f0e6c8; | |
| } | |
| *{box-sizing:border-box} | |
| body{margin:0;background:var(--paper);color:var(--ink); | |
| font-family:Georgia,'Iowan Old Style','Times New Roman',serif;line-height:1.7} | |
| .wrap{max-width:920px;margin:0 auto;padding:40px 26px 90px} | |
| .hero{background:var(--blue);color:#fff;border-radius:16px;padding:34px 34px 30px;margin-bottom:8px} | |
| .hero h1{margin:0 0 10px;font-size:32px;line-height:1.25;letter-spacing:.2px} | |
| .hero p{margin:0;color:#d7e3ee;font-size:16px} | |
| .hero .kpis{display:flex;gap:14px;flex-wrap:wrap;margin-top:24px} | |
| .kpi{background:#ffffff14;border:1px solid #ffffff33;border-radius:12px;padding:12px 16px;min-width:150px} | |
| .kpi b{display:block;font-size:26px;line-height:1.1} | |
| .kpi span{font-size:12.5px;color:#cfdded} | |
| h2{font-size:23px;margin:52px 0 6px;padding-top:8px} | |
| h2 .num{display:inline-block;background:var(--blue);color:#fff;border-radius:8px; | |
| font-size:14px;padding:2px 10px;margin-right:10px;vertical-align:2px;font-family:Helvetica,Arial,sans-serif} | |
| h3{font-size:17.5px;margin:30px 0 4px;color:#20302b} | |
| h4{font-size:15px;margin:22px 0 2px;color:#20302b} | |
| p{margin:10px 0} | |
| .mut{color:var(--mut)} | |
| .small{font-size:13.5px} | |
| .card{background:var(--card);border:1px solid var(--line);border-radius:14px;padding:18px 20px;margin:18px 0} | |
| .callout{border-left:5px solid var(--green);background:#eef5f1;border-radius:0 12px 12px 0;padding:14px 18px;margin:18px 0} | |
| .callout.warn{border-left-color:var(--orange);background:#fdf1e7} | |
| .callout.stop{border-left-color:var(--red);background:#fbeeee} | |
| .callout b:first-child{display:block;margin-bottom:4px;font-family:Helvetica,Arial,sans-serif;font-size:12.5px; | |
| letter-spacing:.9px;text-transform:uppercase;color:var(--green)} | |
| .callout.warn b:first-child{color:#b35c17} | |
| .callout.stop b:first-child{color:var(--red)} | |
| table{border-collapse:collapse;width:100%;margin:14px 0;font-size:14.5px} | |
| th,td{border:1px solid #c9c4b8;padding:7px 10px;text-align:left} | |
| th{background:#eae6dc;font-family:Helvetica,Arial,sans-serif;font-size:12.5px;letter-spacing:.4px;text-transform:uppercase} | |
| td.n,th.n{text-align:right;font-variant-numeric:tabular-nums} | |
| tr:nth-child(even) td{background:#ffffff} | |
| img{max-width:100%;border:1px solid var(--line);border-radius:12px;margin:14px 0;background:#fff} | |
| figure{margin:20px 0}figcaption{font-size:13px;color:var(--mut);margin-top:-6px} | |
| code{background:#efece4;border:1px solid #e2ddd0;padding:1px 6px;border-radius:6px;font-size:13.5px} | |
| pre{background:#22282a;color:#e8eae6;padding:16px;border-radius:12px;overflow-x:auto;font-size:13px;line-height:1.55} | |
| .svgbox{text-align:center;margin:18px 0;overflow-x:auto} | |
| .toc{columns:2;column-gap:34px;margin:16px 0;font-size:14.5px} | |
| .toc div{break-inside:avoid;padding:2px 0} | |
| .toc a{color:var(--blue);text-decoration:none} | |
| .toc a:hover{text-decoration:underline} | |
| .tag{display:inline-block;font-family:Helvetica,Arial,sans-serif;font-size:11.5px;background:#e4e0d5; | |
| border-radius:20px;padding:2px 10px;color:#4b5a52;margin-right:6px} | |
| .lead{font-size:17px;color:#3d4a44} | |
| hr{border:0;border-top:1px solid var(--line);margin:34px 0} | |
| .footer{margin-top:50px;font-size:13px;color:var(--mut)} | |
| .big-good{color:var(--green);font-weight:bold} | |
| .big-bad{color:var(--red);font-weight:bold} | |
| </style> | |
| </head> | |
| <body><div class="wrap"> | |
| <div class="hero"> | |
| <h1>Oman Water Demand Forecasting<br>Daily 28 · Weekly 11 · Weekly 4</h1> | |
| <p>A review pack you can run yourself. Two bottle sizes, five main sales regions, one small | |
| model per region, plain-English method, honest numbers, the exact settings we used.</p> | |
| <div class="kpis"> | |
| <div class="kpi"><b>97.6%</b><span>200ML weekly accuracy (11 weeks)</span></div> | |
| <div class="kpi"><b>95.0%</b><span>330ML weekly accuracy (11 weeks)</span></div> | |
| <div class="kpi"><b>78.9%</b><span>200ML daily accuracy (28 days)</span></div> | |
| <div class="kpi"><b>77.4%</b><span>330ML daily accuracy (28 days)</span></div> | |
| </div> | |
| </div> | |
| <p class="small mut">Accuracy = 100 × (1 − WAPE). Every number in this report comes from a real | |
| out-of-sample holdout, not from training data. DUQM and PDO are forecast but excluded from the | |
| headline score — Chapter 9 explains why, with numbers.</p> | |
| <div class="card"> | |
| <b>Contents</b> | |
| <div class="toc"> | |
| <div><a href="#c1">1 · The question we were asked</a></div> | |
| <div><a href="#c2">2 · The data we were given</a></div> | |
| <div><a href="#c3">3 · How missing days were filled (the policy)</a></div> | |
| <div><a href="#c4">4 · How one forecast is made</a></div> | |
| <div><a href="#c5">5 · Daily results and the daily ceiling</a></div> | |
| <div><a href="#c6">6 · Weekly results</a></div> | |
| <div><a href="#c7">7 · Why weekly beats daily</a></div> | |
| <div><a href="#c8">8 · What each forecast is for (business)</a></div> | |
| <div><a href="#c9">9 · Why DUQM and PDO sit outside the headline</a></div> | |
| <div><a href="#c10">10 · Reproduce everything yourself</a></div> | |
| <div><a href="#c11">11 · Glossary and straight answers</a></div> | |
| </div> | |
| </div> | |
| <h2 id="c1"><span class="num">1</span>The question we were asked</h2> | |
| <p class="lead">A water bottler in Oman sells two products — 24×330ML packs and 30×200ML packs — | |
| across 12 sales areas grouped into 7 regions. The factory, the trucks, and the warehouses all need | |
| to know <b>how much water will be sold next, and where</b>. We were asked two things:</p> | |
| <ol> | |
| <li><b>Daily:</b> predict sales for each area for the next 28 days. Trucks load every morning. | |
| Target customers: distribution and drivers.</li> | |
| <li><b>Weekly:</b> predict sales for each region for the coming weeks. Production runs, raw water, | |
| bottles, and warehouse space are planned weekly. Target customers: plant and supply chain.</li> | |
| </ol> | |
| <div class="callout"><b>Key idea</b> | |
| A forecast is only useful if its accuracy is honest — measured on days the model has never seen. | |
| Everything in this pack is scored on a holdout at the very end of the history, and all the code | |
| that produced these numbers is in this folder so a reviewer can rerun it and get the same result.</div> | |
| <h2 id="c2"><span class="num">2</span>The data we were given</h2> | |
| <table> | |
| <tr><th>Product</th><th class="n">Rows</th><th class="n">Areas</th><th class="n">Regions</th><th>History</th></tr> | |
| <tr><td>24 × 330ML packs</td><td class="n">12,568</td><td class="n">12</td><td class="n">7</td><td>1 Jan 2021 → 24 Jan 2025</td></tr> | |
| <tr><td>30 × 200ML packs</td><td class="n">14,583</td><td class="n">11</td><td class="n">7</td><td>1 Jan 2021 → 29 Jan 2025</td></tr> | |
| </table> | |
| <p>Each row is one day, one area, one product: the day, the area, the region, and the cases sold. | |
| Four things about this file shaped every decision that followed.</p> | |
| <h3>2.1 There is not a single zero in the file</h3> | |
| <p>Across more than 27,000 rows, the smallest sale is 1 case. The source system only writes a row | |
| when something was sold. So a missing day usually means <b>no sale happened</b> — not that the data | |
| was lost. That single observation drives the filling policy in Chapter 3.</p> | |
| <h3>2.2 Some days are missing everywhere at once</h3> | |
| <p>20–23 July 2021 (Eid al-Adha) is missing in almost every area. 3 October 2021 (Cyclone Shaheen | |
| emergency holiday) is missing in several. Those are real closures, and the pattern repeats across | |
| both products — strong evidence they are genuine no-sale days, not broken records.</p> | |
| <h3>2.3 Friday means something different in every area</h3> | |
| <table> | |
| <tr><th>Area</th><th class="n">Fridays in history</th><th class="n">Fridays with no sale</th><th class="n">Share</th></tr> | |
| <tr><td>IBRA</td><td class="n">211</td><td class="n">197</td><td class="n">93%</td></tr> | |
| <tr><td>IBRI</td><td class="n">211</td><td class="n">158</td><td class="n">75%</td></tr> | |
| <tr><td>SUR</td><td class="n">212</td><td class="n">79</td><td class="n">37%</td></tr> | |
| <tr><td>NIZWA</td><td class="n">212</td><td class="n">72</td><td class="n">34%</td></tr> | |
| <tr><td>SAHAM</td><td class="n">213</td><td class="n">32</td><td class="n">15%</td></tr> | |
| <tr><td>AMERAT</td><td class="n">213</td><td class="n">1</td><td class="n">0.5%</td></tr> | |
| </table> | |
| <p>In IBRA, Friday is effectively a closed day. In AMERAT, Friday is a normal selling day. No single | |
| "Friday = zero" rule can be true for both. The model has to learn <b>area × weekday</i></b>, which is | |
| exactly what the features in Chapter 4 allow it to do.</p> | |
| <h3>2.4 One stretch of October 2021 looks like a data outage, not a holiday</h3> | |
| <p>11–15 October 2021: many normally busy 330ML areas (AMERAT, RUSAYL, MUSANNAH, SALALAH, SAHAM) | |
| go silent at the same time, while 200ML keeps selling. There is no public holiday in that window. | |
| Seventeen values are affected — a tiny number — and they are handled separately from real zero days | |
| (Chapter 3), with the assumption flagged in the data.</p> | |
| <h2 id="c3"><span class="num">3</span>How missing days were filled (the policy)</h2> | |
| <p class="lead">One rule for every blank would be wrong. A blank can mean "nobody ordered that day" | |
| (truly zero) or "the system lost that day" (unknown, not zero). Those two need opposite treatment. | |
| The policy below sorts every blank into one of four buckets.</p> | |
| <div class="svgbox"> | |
| <svg width="820" height="200" viewBox="0 0 820 200" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12"> | |
| <rect x="10" y="80" width="150" height="44" rx="10" fill="#1f4e79"/> | |
| <text x="85" y="99" fill="#fff" text-anchor="middle">Build a full calendar</text> | |
| <text x="85" y="115" fill="#cfe0ee" text-anchor="middle" font-size="10.5">first sale → last sale, per area</text> | |
| <rect x="205" y="18" width="185" height="52" rx="10" fill="#fff" stroke="#2f7d5d" stroke-width="2"/> | |
| <text x="297" y="40" text-anchor="middle" fill="#2f7d5d">Shared closure day?</text> | |
| <text x="297" y="57" text-anchor="middle" font-size="11">Eid · Cyclone Shaheen → <tspan font-weight="bold">0</tspan></text> | |
| <rect x="205" y="86" width="185" height="52" rx="10" fill="#fff" stroke="#e07b39" stroke-width="2"/> | |
| <text x="297" y="108" text-anchor="middle" fill="#b35c17">Oct 11–15 2021 outage?</text> | |
| <text x="297" y="125" text-anchor="middle" font-size="11">same-weekday median, weight 0.25</text> | |
| <rect x="205" y="154" width="185" height="40" rx="10" fill="#fff" stroke="#5b6560" stroke-width="2"/> | |
| <text x="297" y="179" text-anchor="middle" fill="#3d4a44">Any other blank inside the span → <tspan font-weight="bold">0</tspan></text> | |
| <rect x="440" y="80" width="180" height="44" rx="10" fill="#eef5f1" stroke="#2f7d5d"/> | |
| <text x="530" y="99" text-anchor="middle">Flag every guess</text> | |
| <text x="530" y="115" text-anchor="middle" font-size="10.5">was_observed · is_imputed · weight</text> | |
| <rect x="665" y="80" width="145" height="44" rx="10" fill="#1f4e79"/> | |
| <text x="737" y="99" fill="#fff" text-anchor="middle">Ready to train</text> | |
| <text x="737" y="115" fill="#cfe0ee" text-anchor="middle" font-size="10.5">guesses never scored</text> | |
| <path d="M160 102 H205" stroke="#555" stroke-width="2" marker-end="url(#a)"/> | |
| <path d="M390 44 H417 V92 H440" stroke="#555" stroke-width="2" fill="none" marker-end="url(#a)"/> | |
| <path d="M390 112 H440" stroke="#555" stroke-width="2" marker-end="url(#a)"/> | |
| <path d="M390 174 H417 V122 H440" stroke="#555" stroke-width="2" fill="none" marker-end="url(#a)"/> | |
| <path d="M620 102 H665" stroke="#555" stroke-width="2" marker-end="url(#a)"/> | |
| <defs><marker id="a" markerWidth="9" markerHeight="9" refX="7" refY="3" orient="auto"> | |
| <path d="M0,0 L7,3 L0,6 z" fill="#555"/></marker></defs> | |
| </svg> | |
| </div> | |
| <table> | |
| <tr><th>Situation</th><th>What we write into the data</th><th>Audit flag</th></tr> | |
| <tr><td>Real sale (row exists in the source)</td><td>Volume as-is, never touched</td><td><code>was_observed = 1</code></td></tr> | |
| <tr><td>Known shared closure (Eid, cyclone)</td><td>Volume = 0</td><td><code>is_closure = 1</code></td></tr> | |
| <tr><td>330ML Oct 11–15 2021 outage (17 values)</td><td>Median of the same weekday ±4 weeks</td><td><code>is_imputed = 1</code>, weight 0.25</td></tr> | |
| <tr><td>Any other blank inside an area's active span</td><td>Volume = 0 (likely no-sale day)</td><td><code>imputation_method = zero_no_sale</code></td></tr> | |
| <tr><td>Before the first sale / after the last sale</td><td>No row created</td><td>inactive period, not a zero</td></tr> | |
| </table> | |
| <h3>The 330ML outage estimate, worked out</h3> | |
| <p>Take AMERAT on Monday 11 October 2021. The surrounding Mondays in the source file sold | |
| 716, 854, 833, 1732, 690, 912, 837, 653 cases. The median is <b>835</b>. That 835 is written in, | |
| marked as an estimate, and given only a quarter of the training weight of a real row. Seventeen | |
| values were filled this exact way.</p> | |
| <div class="card"> | |
| <b>Filled calendars — the result of the policy</b> | |
| <table> | |
| <tr><th>Product</th><th class="n">Total rows</th><th class="n">Real sales</th><th class="n">No-sale zeros</th><th class="n">Closure zeros</th><th class="n">Estimated</th></tr> | |
| <tr><td>200ML</td><td class="n">16,378</td><td class="n">14,583</td><td class="n">1,750</td><td class="n">45</td><td class="n">0</td></tr> | |
| <tr><td>330ML</td><td class="n">17,697</td><td class="n">12,568</td><td class="n">5,062</td><td class="n">50</td><td class="n">17</td></tr> | |
| </table> | |
| <p class="small mut">That is 93% real data in the 200ML file and 71% in the 330ML file. The | |
| 330ML file is sparser because four of its areas sell very little (Chapter 9).</p> | |
| </div> | |
| <figure><img src="charts/policy_split.png" alt="How the filled calendar splits between real rows and filled rows"> | |
| <figcaption>Left: how the 200ML calendar is built. Right: how the 330ML calendar is built. Orange is | |
| the only guessed content in the whole project, and it is 17 values out of 17,697.</figcaption></figure> | |
| <div class="callout warn"><b>The most important courtesy in this pack</b> | |
| Every filled row keeps its flag. When we score the model we use <b>only observed rows</b> | |
| (<code>is_imputed = 0</code>). A guess can help training; a guess is never counted as truth when | |
| measuring accuracy.</div> | |
| <h2 id="c4"><span class="num">4</span>How one forecast is made</h2> | |
| <p class="lead">The model is a small decision-tree ensemble called <b>LightGBM</b>, the same family | |
| of method that won the M5 retail forecasting competition. It is not a neural network and not deep | |
| learning. We train <b>one model per region per product</b> — ten small models in total in this pack | |
| — because each region has its own habits.</p> | |
| <div class="svgbox"> | |
| <svg width="820" height="230" viewBox="0 0 820 230" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12"> | |
| <rect x="8" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/> | |
| <text x="96" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">History features</text> | |
| <text x="96" y="58" text-anchor="middle" font-size="10.5">sales 1, 7, 14, 28, 56 days ago</text> | |
| <text x="96" y="76" text-anchor="middle" font-size="10.5">rolling mean / spread, 7 & 28 days</text> | |
| <text x="96" y="94" text-anchor="middle" font-size="10.5">how often it sells at all</text> | |
| <text x="96" y="112" text-anchor="middle" font-size="10.5">days since last sale</text> | |
| <rect x="212" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/> | |
| <text x="300" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">Calendar features</text> | |
| <text x="300" y="58" text-anchor="middle" font-size="10.5">weekday, day, month, week</text> | |
| <text x="300" y="76" text-anchor="middle" font-size="10.5">weekend / Friday flags</text> | |
| <text x="300" y="94" text-anchor="middle" font-size="10.5">closure and event flags</text> | |
| <text x="300" y="112" text-anchor="middle" font-size="10.5">year trend</text> | |
| <rect x="416" y="14" width="176" height="200" rx="12" fill="#f4f2ec" stroke="#c9c4b8"/> | |
| <text x="504" y="36" text-anchor="middle" font-weight="bold" fill="#1f4e79">Who / where</text> | |
| <text x="504" y="58" text-anchor="middle" font-size="10.5">area name</text> | |
| <text x="504" y="76" text-anchor="middle" font-size="10.5">region name</text> | |
| <text x="504" y="94" text-anchor="middle" font-size="10.5">(given as categories, so the</text> | |
| <text x="504" y="112" text-anchor="middle" font-size="10.5">model learns area habits)</text> | |
| <rect x="626" y="52" width="184" height="124" rx="12" fill="#1f4e79"/> | |
| <text x="718" y="92" fill="#fff" text-anchor="middle" font-weight="bold">One model,</text> | |
| <text x="718" y="110" fill="#fff" text-anchor="middle" font-weight="bold">one region,</text> | |
| <text x="718" y="128" fill="#fff" text-anchor="middle" font-weight="bold">one product</text> | |
| <text x="718" y="152" fill="#cfe0ee" text-anchor="middle" font-size="10.5">predicts cases for each day</text> | |
| <path d="M184 114 H212" stroke="#555" stroke-width="2" marker-end="url(#b)"/> | |
| <path d="M388 114 H416" stroke="#555" stroke-width="2" marker-end="url(#b)"/> | |
| <path d="M592 114 H626" stroke="#555" stroke-width="2" marker-end="url(#b)"/> | |
| <defs><marker id="b" markerWidth="9" markerHeight="9" refX="7" refY="3" orient="auto"> | |
| <path d="M0,0 L7,3 L0,6 z" fill="#555"/></marker></defs> | |
| </svg> | |
| </div> | |
| <h3>4.1 The rule that keeps the score honest</h3> | |
| <p>All history features are built with a one-day shift — when the model predicts Tuesday, it is | |
| allowed to see Monday and earlier, never Tuesday itself. And the last 28 days of history are held | |
| out completely: the model never trains on them. The score in this report is only ever measured on | |
| those unseen days.</p> | |
| <h3>4.2 What "accuracy" means here (WAPE, in one minute)</h3> | |
| <div class="card"> | |
| <p>We measure error with <b>WAPE</b> — the total absolute miss divided by the total actual volume.</p> | |
| <pre>Actual week: 1,000 cases | |
| Forecast: 950 cases | |
| Miss: 50 cases | |
| WAPE = 50 / 1000 = 5% Accuracy = 100% − 5% = 95%</pre> | |
| <p class="mut small">WAPE is a percentage of the whole, so it does not punish big regions and small | |
| regions equally per case — it asks "of all the water we shipped, how much did we misjudge?" That | |
| matches how the business thinks about a plan.</p> | |
| </div> | |
| <h3>4.3 The exact settings used (winners only)</h3> | |
| <p>Each region's best settings are stored in <code>params/best_params_200ML.json</code> and | |
| <code>params/best_params_330ML.json</code> so the reviewer runs exactly what produced these numbers.</p> | |
| <table> | |
| <tr><th>Product / region</th><th>Model</th><th class="n">Learning rate</th><th class="n">Leaves</th><th class="n">Tweedie power</th></tr> | |
| <tr><td>200ML CAPITAL, Batinah, Sharqiyah</td><td>Tweedie</td><td class="n">0.05 – 0.11</td><td class="n">511</td><td class="n">1.05 – 1.18</td></tr> | |
| <tr><td>200ML Dhofar, Al Dakhiliyah</td><td>Tweedie</td><td class="n">0.03 – 0.07</td><td class="n">127 – 511</td><td class="n">1.14 – 1.49</td></tr> | |
| <tr><td>330ML CAPITAL</td><td>Poisson</td><td class="n">0.10</td><td class="n">63</td><td class="n">—</td></tr> | |
| <tr><td>330ML Batinah, Dhofar, Sharqiyah, Al Dakhiliyah</td><td>Tweedie</td><td class="n">0.02 – 0.11</td><td class="n">127 – 511</td><td class="n">1.29 – 1.53</td></tr> | |
| </table> | |
| <p class="small mut">Tweedie and Poisson are count-friendly objectives: they suit demand data with | |
| many small days and occasional big ones, and cannot predict negative cases. The model stops | |
| training automatically when the last 28 days of training stop improving (early stopping).</p> | |
| <h2 id="c5"><span class="num">5</span>Daily results and the daily ceiling</h2> | |
| <p class="lead">Daily accuracy for the next 28 days, measured on unseen days, five regions:</p> | |
| <table> | |
| <tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr> | |
| <tr><td>CAPITAL</td><td class="n">86.6%</td><td class="n">84.2%</td></tr> | |
| <tr><td>Dhofar</td><td class="n">78.8%</td><td class="n">76.1%</td></tr> | |
| <tr><td>Batinah</td><td class="n">76.2%</td><td class="n">69.9%</td></tr> | |
| <tr><td>Sharqiyah</td><td class="n">63.0%</td><td class="n">45.4%</td></tr> | |
| <tr><td>Al Dakhiliyah</td><td class="n">59.6%</td><td class="n">57.1%</td></tr> | |
| <tr><td><b>All five together</b></td><td class="n"><b>78.9%</b></td><td class="n"><b>77.4%</b></td></tr> | |
| </table> | |
| <figure><img src="charts/daily28_200ML.png" alt="200ML daily 28-day forecast vs actual"> | |
| <figcaption>200ML — one panel per region, because each region runs its own plant and trucks. | |
| Each panel draws that region's daily total; the percentage in its title is that region's own | |
| daily accuracy, the same number as in the table above.</figcaption></figure> | |
| <figure><img src="charts/daily28_330ML.png" alt="330ML daily 28-day forecast vs actual"> | |
| <figcaption>330ML — same per-region view for the other product. Sharqiyah's panel shows the | |
| hardest region honestly at 45.4%: no country total hides it.</figcaption></figure> | |
| <h3>5.1 Where the daily error comes from</h3> | |
| <figure><img src="charts/horizon_200ML.png" alt="200ML daily error per day of the holdout"> | |
| <figcaption>200ML — error per day of the holdout. The spikes sit on unusual dates, not on far-away | |
| dates. In this holdout the model still sees real history for every day (a "re-forecast each | |
| morning" view); in a full 28-day rollout the far days would also be building on predicted history, | |
| which would add more error, not less.</figcaption></figure> | |
| <figure><img src="charts/horizon_330ML.png" alt="330ML daily error per day of the holdout"> | |
| <figcaption>330ML — same pattern. Day 1 is sometimes worse than day 28, which proves the point: | |
| daily error is driven by <i>what that particular day was</i>, not by how far ahead it is.</figcaption></figure> | |
| <figure><img src="charts/daily_error_hist.png" alt="Distribution of daily errors per row"> | |
| <figcaption>Half of all daily rows land within the orange line (the median error). The long right | |
| tail is the price of a single unusual day.</figcaption></figure> | |
| <div class="callout stop"><b>Why daily cannot reach 5% error with this data</b> | |
| <ol> | |
| <li><b>A single day is one lumpy decision.</b> A shop that orders 0 on Monday and 500 on Tuesday | |
| makes the day-to-day line jump. History cannot know which Tuesday the truck arrives; the model | |
| splits the difference and takes the error.</li> | |
| <li><b>Friday has five different meanings.</b> Closed in IBRA, normal in AMERAT. The model learns | |
| this, but any single surprise Friday is a large miss on a single day.</li> | |
| <li><b>Holidays stop trucks, not thirst.</b> Eid and cyclone days are zero <i>sales</i> but not | |
| zero demand. And nothing in the file explains the biggest jumps — no promotions, no price, no | |
| stock-out records, no delivery routes.</li> | |
| <li><b>Small denominators explode percentages.</b> Missing 30 cases in an area that sells 60 that | |
| day is a 50% error for that row, even though the business barely notices 30 cases.</li> | |
| <li><b>Four years is a short teacher.</b> Ramadan drifts about 11 days every year, so each area's | |
| Ramadan pattern has been seen only four times. That is thin evidence for such a strong event.</li> | |
| </ol> | |
| Weather (temperature, humidity, rain) was also tested and added no measurable lift, so the final | |
| models do not use it.</div> | |
| <h2 id="c6"><span class="num">6</span>Weekly results</h2> | |
| <p class="lead">The same ten models, the same forecasts — simply added up over Monday-to-Sunday | |
| weeks before scoring. Two horizons were tested: the last 11 full weeks, and the last 4 weeks.</p> | |
| <h3>6.1 Eleven weeks</h3> | |
| <table> | |
| <tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr> | |
| <tr><td>CAPITAL</td><td class="n">97.3%</td><td class="n">93.9%</td></tr> | |
| <tr><td>Batinah</td><td class="n">95.7%</td><td class="n">89.6%</td></tr> | |
| <tr><td>Dhofar</td><td class="n">91.5%</td><td class="n">91.3%</td></tr> | |
| <tr><td>Sharqiyah</td><td class="n">90.5%</td><td class="n">87.8%</td></tr> | |
| <tr><td>Al Dakhiliyah</td><td class="n">84.7%</td><td class="n">84.0%</td></tr> | |
| <tr><td><b>All five together</b></td><td class="n"><b>97.6%</b></td><td class="n"><b>95.0%</b></td></tr> | |
| </table> | |
| <figure><img src="charts/weekly11_200ML.png" alt="200ML weekly 11 weeks forecast vs actual"> | |
| <figcaption>200ML — weekly totals, 11 weeks. Percentages above each pair are that week's error | |
| (actual vs forecast).</figcaption></figure> | |
| <figure><img src="charts/weekly11_330ML.png" alt="330ML weekly 11 weeks forecast vs actual"> | |
| <figcaption>330ML — weekly totals, 11 weeks. Even the harder product lands on the bar most | |
| weeks.</figcaption></figure> | |
| <h3>6.2 The last four weeks</h3> | |
| <p>Same models, retrained on history ending four weeks ago — a check that the accuracy is not a | |
| lucky artifact of one long window.</p> | |
| <table> | |
| <tr><th>Region</th><th class="n">200ML</th><th class="n">330ML</th></tr> | |
| <tr><td>CAPITAL</td><td class="n">97.0%</td><td class="n">92.8%</td></tr> | |
| <tr><td>Batinah</td><td class="n">94.2%</td><td class="n">89.2%</td></tr> | |
| <tr><td>Dhofar</td><td class="n">92.5%</td><td class="n">88.7%</td></tr> | |
| <tr><td>Sharqiyah</td><td class="n">90.8%</td><td class="n">88.4%</td></tr> | |
| <tr><td>Al Dakhiliyah</td><td class="n">82.7%</td><td class="n">92.0%</td></tr> | |
| <tr><td><b>All five together</b></td><td class="n"><b>96.4%</b></td><td class="n"><b>94.8%</b></td></tr> | |
| </table> | |
| <figure><img src="charts/weekly4_200ML.png" alt="200ML weekly last 4 weeks forecast vs actual"> | |
| <figcaption>200ML — last four weeks, week by week.</figcaption></figure> | |
| <figure><img src="charts/weekly4_330ML.png" alt="330ML weekly last 4 weeks forecast vs actual"> | |
| <figcaption>330ML — last four weeks. The four-week view is the closest to live conditions.</figcaption></figure> | |
| <figure><img src="charts/ladder_200ML.png" alt="200ML daily vs weekly accuracy per region"> | |
| <figcaption>200ML — the same models, two clocks. Weekly clears or approaches the 95% line in every | |
| region except Al Dakhiliyah.</figcaption></figure> | |
| <figure><img src="charts/ladder_330ML.png" alt="330ML daily vs weekly accuracy per region"> | |
| <figcaption>330ML — the same story. Nothing about the model changed between the red and green | |
| bars; only the window the numbers are added over changed.</figcaption></figure> | |
| <h2 id="c7"><span class="num">7</span>Why weekly beats daily</h2> | |
| <p class="lead">This is the single most important idea in the pack, and it is not a model trick — | |
| it is arithmetic.</p> | |
| <div class="svgbox"> | |
| <svg width="820" height="240" viewBox="0 0 820 240" xmlns="http://www.w3.org/2000/svg" font-family="Helvetica,Arial,sans-serif" font-size="12"> | |
| <text x="16" y="24" font-weight="bold" fill="#1f4e79">One week of one area — actual vs forecast</text> | |
| <g stroke="#c9c4b8"><line x1="60" y1="70" x2="60" y2="150"/><line x1="140" y1="70" x2="140" y2="150"/> | |
| <line x1="220" y1="70" x2="220" y2="150"/><line x1="300" y1="70" x2="300" y2="150"/> | |
| <line x1="380" y1="70" x2="380" y2="150"/><line x1="460" y1="70" x2="460" y2="150"/> | |
| <line x1="540" y1="70" x2="540" y2="150"/></g> | |
| <g fill="#1f4e79"> | |
| <rect x="46" y="126" width="28" height="24"/><rect x="126" y="82" width="28" height="68"/> | |
| <rect x="206" y="120" width="28" height="30"/><rect x="286" y="96" width="28" height="54"/> | |
| <rect x="366" y="128" width="28" height="22"/><rect x="446" y="88" width="28" height="62"/> | |
| <rect x="526" y="132" width="28" height="18"/> | |
| </g> | |
| <g fill="#e07b39"> | |
| <rect x="76" y="118" width="16" height="32"/><rect x="156" y="110" width="16" height="40"/> | |
| <rect x="236" y="116" width="16" height="34"/><rect x="316" y="112" width="16" height="38"/> | |
| <rect x="396" y="120" width="16" height="30"/><rect x="476" y="114" width="16" height="36"/> | |
| <rect x="556" y="122" width="16" height="28"/> | |
| </g> | |
| <g font-size="10.5" fill="#5b6560"> | |
| <text x="60" y="166" text-anchor="middle">Mon</text><text x="140" y="166" text-anchor="middle">Tue</text> | |
| <text x="220" y="166" text-anchor="middle">Wed</text><text x="300" y="166" text-anchor="middle">Thu</text> | |
| <text x="380" y="166" text-anchor="middle">Fri</text><text x="460" y="166" text-anchor="middle">Sat</text> | |
| <text x="540" y="166" text-anchor="middle">Sun</text> | |
| </g> | |
| <rect x="620" y="70" width="184" height="80" rx="12" fill="#eef5f1" stroke="#2f7d5d"/> | |
| <text x="712" y="94" text-anchor="middle" font-size="11.5">Day 2: forecast 40,</text> | |
| <text x="712" y="110" text-anchor="middle" font-size="11.5">actual 68 — a 41% miss.</text> | |
| <text x="712" y="132" text-anchor="middle" font-size="11.5" fill="#2f7d5d">Week total: 322 vs 325</text> | |
| <text x="712" y="146" text-anchor="middle" font-size="11.5" fill="#2f7d5d">— only a 1% miss.</text> | |
| <rect x="16" y="186" width="788" height="40" rx="10" fill="#f4f2ec" stroke="#c9c4b8"/> | |
| <text x="34" y="211" font-size="12">Daily: every bar is judged alone, so every wobble is an error. Weekly: the errors point in opposite</text> | |
| <text x="34" y="227" font-size="12">directions, cancel out, and only the level that matters to the business survives.</text> | |
| </svg> | |
| </div> | |
| <p>Over a week, a slow Monday is usually paid back by a busy Tuesday. Adding the days together | |
| before comparing means the random part of the wobble cancels while the real level remains. The | |
| model was never bad at the <i>level</i> of demand — it was bad at guessing the exact <i>day</i>. | |
| The weekly number uses what the model is actually good at.</p> | |
| <div class="callout"><b>In one sentence</b> | |
| Daily asks "will it be 120 or 160 cases on the 14th?" — weekly asks "will this region move about | |
| 1,000 cases this week?" The second question matches both the model's strength and the way a | |
| factory actually plans.</div> | |
| <h2 id="c8"><span class="num">8</span>What each forecast is for (business)</h2> | |
| <table> | |
| <tr><th>Decision</th><th>Clock to use</th><th>Who uses it</th><th>Why that clock</th></tr> | |
| <tr><td>How much to produce next week (raw water, bottles, line time)</td><td><b>Weekly</b></td><td>Plant, production planning</td><td>Runs are batched weekly; a day-level miss is irrelevant to a weekly run. 95–98% is enough to plan materials and shifts.</td></tr> | |
| <tr><td>Warehouse space and stock transfer between regions</td><td><b>Weekly</b></td><td>Supply chain</td><td>Replenishment lead times are measured in weeks; weekly totals set safety stock.</td></tr> | |
| <tr><td>Truck capacity and fleet allocation per region</td><td><b>Weekly</b></td><td>Logistics planning</td><td>Fleet size is re-planned weekly, not daily.</td></tr> | |
| <tr><td>Which shops and areas each truck serves tomorrow</td><td><b>Daily</b></td><td>Distribution, drivers</td><td>A driver needs a number per area per day to load the right quantity; 78% daily accuracy is the honest guidance range, not a promise per shop.</td></tr> | |
| <tr><td>Short shelf-life urgent replenishment</td><td><b>Daily</b></td><td>Area supervisors</td><td>Daily view shows the next few days; supervisors adjust with local knowledge.</td></tr> | |
| </table> | |
| <div class="callout warn"><b>How to read daily numbers honestly</b> | |
| Quoted daily accuracy is the share of total volume forecast correctly across a region — not the | |
| chance that any one shop's day is exact. Use daily forecasts to plan the size of tomorrow's load, | |
| and weekly forecasts to plan what the factory makes. If a decision costs real money, make it on | |
| the weekly clock and use the daily clock only to distribute it.</div> | |
| <h2 id="c9"><span class="num">9</span>Why DUQM and PDO sit outside the headline</h2> | |
| <p class="lead">DUQM, DUQUM PDO and PDO are real customers, and the models do forecast them. They | |
| are left out of the headline accuracy for one reason: their history is too thin to measure with a | |
| percentage fairness.</p> | |
| <table> | |
| <tr><th>Area</th><th>Product</th><th class="n">Days with sales</th><th class="n">Active days</th><th class="n">Coverage</th></tr> | |
| <tr><td>PDO</td><td>200ML</td><td class="n">150</td><td class="n">1,480</td><td class="n">10%</td></tr> | |
| <tr><td>DUQM</td><td>330ML</td><td class="n">362</td><td class="n">1,427</td><td class="n">25%</td></tr> | |
| <tr><td>DUQUM PDO</td><td>330ML</td><td class="n">86</td><td class="n">1,445</td><td class="n">6%</td></tr> | |
| </table> | |
| <p>When a region sells in only a handful of days, a weekly percentage error is divided by a tiny | |
| number. Predicting 40 cases in a week when the truth is 10 is reported as a 300% error — a scary | |
| number for a 30-case miss. Their absolute misses are small: 200ML PDO's typical daily miss is | |
| about 3 cases. That is why the pack reports them separately, with the honest note that a | |
| percentage cannot describe them well.</p> | |
| <div class="card"><b>The effect on the headline, stated openly</b> | |
| <table> | |
| <tr><th>Headline</th><th class="n">200ML</th><th class="n">330ML</th></tr> | |
| <tr><td>Weekly 11-week accuracy, five main regions</td><td class="n">97.6%</td><td class="n">95.0%</td></tr> | |
| <tr><td>Same, with DUQM and PDO added back</td><td class="n">97.6%</td><td class="n">93.5%</td></tr> | |
| </table> | |
| <p class="small mut">200ML barely moves because DUQM and PDO sell almost nothing there. 330ML | |
| drops 1.5 points because two sparse sites with wild percentages are enough to pull the average | |
| down. Nothing about the other five regions changed.</p> | |
| </div> | |
| <h2 id="c10"><span class="num">10</span>Reproduce everything yourself</h2> | |
| <p class="lead">CPU only. No API keys, no internet, no special hardware. From this folder:</p> | |
| <pre>pip install -r requirements.txt | |
| python3 daily28/train.py && python3 daily28/predict.py # daily 28-day holdout | |
| python3 weekly11/train.py && python3 weekly11/predict.py # 11-week window | |
| python3 weekly4/train.py && python3 weekly4/predict.py # last 4 weeks | |
| python3 make_charts.py # rebuild every chart</pre> | |
| <table> | |
| <tr><th>Path</th><th>What it is</th></tr> | |
| <tr><td><code>data/*_daily_calendar_filled.csv</code></td><td>The filled calendars with every flag intact (<code>was_observed</code>, <code>is_imputed</code>, <code>training_weight</code>)</td></tr> | |
| <tr><td><code>params/best_params_*.json</code></td><td>The exact settings that produced the numbers in this report</td></tr> | |
| <tr><td><code>daily28/</code>, <code>weekly11/</code>, <code>weekly4/</code></td><td>train + predict scripts, saved models</td></tr> | |
| <tr><td><code>predictions/</code></td><td>Row-level actual-vs-forecast CSVs you can audit or chart</td></tr> | |
| <tr><td><code>charts/</code></td><td>Every figure in this report, regenerated from the predictions</td></tr> | |
| </table> | |
| <div class="callout"><b>Rerun creates the same numbers</b> | |
| The scripts fix the random seed and use the identical data and settings, so a reviewer on any | |
| machine gets the same accuracy figures shown here, up to tiny library-version differences.</div> | |
| <h2 id="c11"><span class="num">11</span>Glossary and straight answers</h2> | |
| <table> | |
| <tr><th>Term</th><th>Meaning in this pack</th></tr> | |
| <tr><td>WAPE</td><td>Total absolute error ÷ total actual volume. So "5% WAPE" means "we misjudged 5% of the water shipped".</td></tr> | |
| <tr><td>Accuracy</td><td>100% − WAPE. Used everywhere in this report.</td></tr> | |
| <tr><td>Holdout</td><td>The last days of history, removed from training, used only to score.</td></tr> | |
| <tr><td>Lag</td><td>Sales from N days ago, e.g. lag_7 = the same weekday last week.</td></tr> | |
| <tr><td>Rolling mean</td><td>The average of the previous N days — a smooth memory of recent level.</td></tr> | |
| <tr><td>Tweedie / Poisson</td><td>Objective functions for count data (cases sold). They fit demand with many small days and occasional large ones and can never predict negative sales.</td></tr> | |
| <tr><td>LightGBM</td><td>A decision-tree ensemble; the same family that won the M5 retail forecasting competition.</td></tr> | |
| <tr><td>Early stopping</td><td>The model keeps a small internal check window and stops improving when it starts to memorize.</td></tr> | |
| </table> | |
| <h3>Straight answers</h3> | |
| <p><b>Why not deep learning?</b> Four years of daily data across a handful of areas is small. Tree | |
| ensembles have repeatedly matched or beaten neural networks on this kind of tabular, intermittent | |
| demand — including in the M5 competition this project takes its recipe from. Neural networks would | |
| add complexity without evidence of gain.</p> | |
| <p><b>Can daily accuracy ever reach 95%?</b> Not with this data. To get there the model would need | |
| to know the events that actually move a single day: promotions and prices, out-of-stock days, | |
| delivery schedules, and a confirmed closure calendar. With those inputs a daily model could | |
| improve; without them it is guessing at the noisiest part of the signal. Weekly already crosses | |
| 95% for four of five regions without any of that.</p> | |
| <p><b>Does weather help?</b> It was tested properly (temperature, humidity, rain, heat flags from | |
| ERA5 reanalysis). It produced no measurable lift on either product, so it was left out to keep the | |
| models simple. It can be revisited when more history exists.</p> | |
| <p><b>What should we do next?</b> Two things pay first: get the promotion and stock-out history | |
| from the business to attack the daily error where it is largest, and keep the weekly forecast as | |
| the production number, since weekly accuracy is already at the level the plant needs.</p> | |
| <div class="footer"> | |
| <p>Method summary: region-specialist LightGBM (Tweedie, Poisson for 330ML CAPITAL) · calendar + | |
| shift-safe lags/rolling features · early stopping on the last 28 training days · scored on observed | |
| rows only (<code>is_imputed = 0</code>) · daily holdout = last 28 days · weekly holdouts = last 11 | |
| full Mon–Sun weeks and last 4 weeks · DUQM and PDO reported separately, never in the headline. | |
| Weather tested (ERA5) with no lift. Deep learning not used.</p> | |
| </div> | |
| </div></body></html> | |