"The network is the feature: Learning layers of abstraction."
Alex Krizhevsky and Geoffrey Hinton win the ImageNet competition by a landslide using a Deep CNN. This proved that GPUs and Deep Nets were the future.
Ian Goodfellow introduces GANs, where two networks compete. One creates images, the other critiques them. This was the birth of AI-generated art.
Google DeepMind's AlphaGo defeats the world champion in Go, a game previously thought impossible for AI due to its infinite complexity.
Google researchers introduce the **Transformer** architecture. This eventually replaced RNNs and paved the way for modern LLMs.
Designed for spatial data like images. They use "filters" to detect edges, then shapes, then objects.
Designed for sequential data like speech and text. They have "memory" of what happened in the previous step.
Used for compression and noise removal. They learn to reconstruct their input from a condensed "bottleneck" layer.
Combining neural nets with trial-and-error rewards. Used for robotics and mastering video games.
AI reached a "level up" not just because of better code, but because we changed the physical way computers think.
The Move: Moving from CPU to GPU. While a CPU handles a few complex tasks in a row (Serial), a GPU handles thousands of simple math tasks at once (Parallel).
Why it matters: Neural networks are just massive matrices of multiplication. GPUs can do billions of these per second.
The Move: Nvidia’s CUDA allowed researchers to write C++ code directly for the GPU. This turned a "video card" into a general-purpose AI brain.
The Move: Google developed ASICs (Application-Specific Integrated Circuits) designed specifically for the matrix math used in AI, stripping away everything a computer doesn't need for neural nets.
The Move: The bottleneck wasn't just calculation speed; it was moving data to the chip. HBM allowed "stacks" of memory to sit right next to the processor, providing the speed needed for LLMs.
The mathematical engine. The model calculates its "error" and sends it backward through the network to update millions of weights using Calculus (Derivatives).
Non-linear functions that decide if a neuron should "fire." ReLU solved the "vanishing gradient" problem, allowing nets to be 100+ layers deep.
The methodology of taking a model trained on one giant task (like recognizing cats) and "fine-tuning" it for a specific task (like detecting cancer in X-rays).
Methods used during training to keep the network stable and prevent it from becoming overly reliant on specific "pathways," ensuring better generalization.
The fundamental approach shifted to End-to-End. You no longer tell the AI to look for "eyes" and "ears" to find a face. You give it 10 million faces, and it discovers that "eyes" are a statistically significant pattern on its own. The machine builds its own internal dictionary.