Download statistical-ml.html from Soha85/AI_Review: direct link, hf CLI and curl.
- Browser
- Download file 7.32 kB
-
https://huggingface.co/spaces/Soha85/AI_Review/resolve/main/statistical-ml.html
- Command line
-
hf download hf://spaces/Soha85/AI_Review/statistical-ml.html
-
curl -L -o statistical-ml.html https://huggingface.co/spaces/Soha85/AI_Review/resolve/main/statistical-ml.html
7.32 kB
| <html lang="en"> | |
| <head> | |
| <meta charset="UTF-8"> | |
| <meta name="viewport" content="width=device-width, initial-scale=1.0"> | |
| <title>Era Details: Statistical Machine Learning</title> | |
| <link rel="stylesheet" href="style.css"> | |
| <style> | |
| .details-body { max-width: 800px; margin: 0 auto; padding: 40px 20px; text-align: left; } | |
| .hero-section { border-bottom: 2px solid var(--primary); padding-bottom: 20px; margin-bottom: 40px; } | |
| .quote { font-style: italic; font-size: 1.5rem; color: #38bdf8; margin: 10px 0; } | |
| .back-btn { display: inline-block; margin-bottom: 20px; color: var(--primary); text-decoration: none; font-weight: bold; } | |
| .sub-timeline { border-left: 2px solid #334155; padding-left: 20px; margin-left: 10px; } | |
| .milestone { margin-bottom: 30px; position: relative; } | |
| .milestone::before { content: ""; position: absolute; left: -26px; top: 5px; width: 10px; height: 10px; background: #38bdf8; border-radius: 50%; } | |
| .methodology-box { background: #1e293b; border-left: 4px solid var(--primary); padding: 20px; margin: 20px 0; } | |
| .tech-list { display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; } | |
| .tech-tag { background: #0f172a; border: 1px solid #334155; padding: 10px; border-radius: 5px; font-family: monospace; font-size: 0.85rem; } | |
| </style> | |
| </head> | |
| <body> | |
| <div class="details-body"> | |
| <a href="index.html" class="back-btn">← Back to Timeline</a> | |
| <div class="hero-section"> | |
| <span class="year" style="color: #38bdf8;">1995 – 2010</span> | |
| <h1>Statistical Machine Learning</h1> | |
| <p class="quote">"Don't tell the computer what to do; show it what you've done."</p> | |
| </div> | |
| <section> | |
| <h2>Chronology of the Probabilistic Shift</h2> | |
| <div class="sub-timeline"> | |
| <div class="milestone"> | |
| <h4>1995: Support Vector Machines (SVM) Popularity</h4> | |
| <p>Cortes and Vapnik publish their work on SVMs, providing a powerful mathematical framework for classification that outperformed early neural nets.</p> | |
| </div> | |
| <div class="milestone"> | |
| <h4>1997: Deep Blue vs. Garry Kasparov</h4> | |
| <p>IBM's Deep Blue defeats the world chess champion. While largely "brute-force," it utilized sophisticated evaluation functions learned from grandmaster games.</p> | |
| </div> | |
| <div class="milestone"> | |
| <h4>2001: Random Forests Algorithm</h4> | |
| <p>Leo Breiman introduces Random Forests, showing that an ensemble of many "weak" decision trees could create a very "strong" and stable predictor.</p> | |
| </div> | |
| <div class="milestone"> | |
| <h4>2006: The Netflix Prize</h4> | |
| <p>Netflix offers $1M to improve their recommendation engine, catalyzing massive research into Collaborative Filtering and Matrix Factorization.</p> | |
| </div> | |
| <div class="milestone"> | |
| <h4>2009: ImageNet Launch</h4> | |
| <p>Fei-Fei Li launches ImageNet, a massive labeled dataset of 14 million images, creating the "competition" that would eventually trigger the Deep Learning era.</p> | |
| </div> | |
| </div> | |
| </section> | |
| <section style="margin-top: 40px;"> | |
| <h2>The "Scientific" Methodologies of ML</h2> | |
| <p>This era moved AI from "hacking" to a structured engineering discipline. These four pillars are what made Statistical ML reliable:</p> | |
| <div class="methodology-box"> | |
| <h4>1. Cross-Validation (The Gold Standard)</h4> | |
| <p>To ensure a model didn't just "memorize" data, engineers divided datasets into <strong>Training, Validation, and Test sets</strong>. Techniques like <em>K-Fold Cross-Validation</em> allowed models to be tested on multiple subsets of data to prove their stability.</p> | |
| </div> | |
| <div class="methodology-box"> | |
| <h4>2. Regularization (L1 & L2)</h4> | |
| <p>Methodologies like <strong>Lasso (L1)</strong> and <strong>Ridge (L2)</strong> regression were introduced to prevent "Overfitting." They added a mathematical penalty for complexity, forcing the model to stay simple and generalize better to the real world.</p> | |
| </div> | |
| <div class="methodology-box"> | |
| <h4>3. Ensemble Learning</h4> | |
| <p>The methodology of combining multiple models to get one superior result. | |
| <ul> | |
| <li><strong>Bagging:</strong> Training models in parallel (e.g., Random Forests).</li> | |
| <li><strong>Boosting:</strong> Training models in sequence, where each new model fixes the errors of the previous one (e.g., AdaBoost, XGBoost).</li> | |
| </ul> | |
| </p> | |
| </div> | |
| <div class="methodology-box"> | |
| <h4>4. Dimensionality Reduction</h4> | |
| <p>As data grew, models became overwhelmed. Methodologies like <strong>PCA (Principal Component Analysis)</strong> and <strong>LDA</strong> were used to "compress" hundreds of variables into a few key components without losing the essential information.</p> | |
| </div> | |
| </section> | |
| <section style="margin-top: 40px;"> | |
| <h2>The Paradigms of Learning</h2> | |
| <div class="tech-grid"> | |
| <div class="tech-item"> | |
| <h4>Supervised Learning</h4> | |
| <p>Learning with a teacher. The model is given inputs and the correct answers (labels). Goal: Predict the label for new data.</p> | |
| </div> | |
| <div class="tech-item"> | |
| <h4>Unsupervised Learning</h4> | |
| <p>Learning without labels. The model looks for hidden structures or clusters in the data (e.g., grouping customers by behavior).</p> | |
| </div> | |
| <div class="tech-item"> | |
| <h4>Reinforcement Learning (Early)</h4> | |
| <p>Learning through trial and error. Agents receive "rewards" or "penalties" to learn a policy (e.g., TD-Learning used in early game AI).</p> | |
| </div> | |
| </div> | |
| </section> | |
| <section style="margin-top: 40px;"> | |
| <h2>Key Concepts & Breakthroughs</h2> | |
| <ul> | |
| <li><strong>The Curse of Dimensionality:</strong> Learning how to handle data with thousands of features without the model becoming "lost" in the noise.</li> | |
| <li><strong>Feature Engineering:</strong> The manual process of selecting which variables (e.g., the frequency of the word "Free" in an email) are important for the model.</li> | |
| <li><strong>Overfitting vs. Underfitting:</strong> The struggle to make a model that performs well on *new* data, not just the data it was trained on.</li> | |
| <li><strong>The Rise of Big Data:</strong> The realization that more data often beats a better algorithm (The "Unreasonable Effectiveness of Data").</li> | |
| </ul> | |
| </section> | |
| </div> | |
| </body> | |
| </html> |