Neural Networks
Overview
Neural networks have been a leading method in various AI tasks such as image recognition and NLP since the early 2010s. They consist of layers of interconnected computing nodes, or neurons, that mimic the human brain and work together to solve complex problems.
Why neural networks matter
Hand‑coding rules for messy tasks like handwriting recognition quickly becomes unmanageable because you can’t realistically enumerate every visual variation you’ll see in the wild. Neural networks instead consume many labeled examples and learn internal parameters (weights and biases) that implement useful decision rules without you ever specifying those rules directly.
This “learn from data” approach scales to images, audio, and text, which is why deep networks currently provide state‑of‑the‑art solutions in image recognition, speech recognition, and many NLP problems.
Neurons and perceptrons
The basic building block is the perceptron, which takes several inputs, multiplies each by a weight, adds a bias, and outputs 0 or 1 depending on whether this weighted sum crosses a threshold. You can interpret each input as a piece of evidence and its weight as how strongly that evidence should influence the final yes/no decision.
By wiring many perceptrons together you can simulate logic gates and thus compute any function a conventional program can compute, but perceptrons are hard to train because small changes to weights can flip the output abruptly. To make learning smoother, the book replaces them with sigmoid neurons, which use the same weighted sum plus bias but pass it through an S‑shaped activation that outputs a continuous value between 0 and 1 instead of a hard 0/1.
Architectures
Classification vs. Regression in Neural Networks
The fundamental difference between classification and regression neural networks lies in what they predict and how they measure success.
Classification
- ·Predicts discrete class labels
- ·e.g. "spam" or "legitimate" email
- ·e.g. "cat" or "dog" in images
Regression
- ·Predicts continuous quantities
- ·e.g. house prices or temperature
- ·e.g. stock prices
Recurrent neural networks (RNNs) and convolutional neural networks (CNNs) are other neural networks frequently used in machine learning and deep learning tasks.
Transformers are a type of neural network architecture, most popular neural network architecture and it is what most modern AI systems are built upon.
Layers as a transformation pipeline
A feedforward neural network arranges neurons in layers where information flows strictly from input to output without cycles.
- The input layer just encodes raw data, such as 784 pixels for a 28×28 digit image.
- One or more hidden layers transform these inputs into increasingly abstract internal features.
- The output layer produces the final prediction, e.g., 10 activations representing which digit (0–9) the network thinks it sees.
From a web‑dev perspective, each hidden layer is like a middleware stage in a pipeline: it transforms the request representation one step closer to the business decision the final handler will make. Deep networks simply stack more of these stages, letting early layers learn simple patterns and later layers assemble them into higher‑level concepts.
Training
To train a network, you need labeled examples and a way to measure how wrong the network is on those examples through a cost function. Training becomes an optimization problem: adjust all the weights and biases to minimize this cost across the training data.
Learning via cost minimization and gradient descent
Gradient descent repeatedly nudges parameters in the direction that most rapidly decreases the cost, analogous to rolling downhill in an error landscape. Because computing exact gradients on the entire dataset is expensive, it uses stochastic gradient descent (SGD), estimating gradients from small mini‑batches and updating more often, which is the same core idea used in modern deep‑learning frameworks.
Backpropagation: making training scalable
The central algorithm that makes this feasible is backpropagation, which efficiently computes the gradient of the cost with respect to every weight and bias in the network. It works by:
- Doing a forward pass to compute activations in all layers.
- Computing the error at the output layer, capturing how much each output neuron contributed to the total cost.
- Propagating this error backward layer by layer, using the chain rule to determine how each earlier neuron and weight contributed to the final error.
- Using these error signals to compute gradients for each weight and bias, which gradient descent then uses to update the parameters.
Conceptually, backprop is like tracing a bug in your frontend back through each component: you start from the wrong output and work backward to see which earlier pieces contributed how much to the problem, then fix those pieces proportionally.
Check your understanding
Why do sigmoid neurons make training easier than perceptrons?
Which of these are regression tasks?
In a feedforward network classifying 28×28 digit images, what does the output layer produce?
Put the steps of backpropagation in order.
- 01Compute the error at the output layer
- 02Propagate the error backward layer by layer
- 03Compute gradients and update weights and biases
- 04Forward pass to compute activations in all layers