Neural Networks

10 min read interactive

Overview

Neural networks have been a leading method in various AI tasks such as image recognition and NLP since the early 2010s. They consist of layers of interconnected computing nodes, or neurons, that mimic the human brain and work together to solve complex problems.

Why neural networks matter

Hand‑coding rules for messy tasks like handwriting recognition quickly becomes unmanageable because you can’t realistically enumerate every visual variation you’ll see in the wild. Neural networks instead consume many labeled examples and learn internal parameters (weights and biases) that implement useful decision rules without you ever specifying those rules directly.

This “learn from data” approach scales to images, audio, and text, which is why deep networks currently provide state‑of‑the‑art solutions in image recognition, speech recognition, and many NLP problems.

Neurons and perceptrons

The basic building block is the perceptron, which takes several inputs, multiplies each by a weight, adds a bias, and outputs 0 or 1 depending on whether this weighted sum crosses a threshold. You can interpret each input as a piece of evidence and its weight as how strongly that evidence should influence the final yes/no decision.

By wiring many perceptrons together you can simulate logic gates and thus compute any function a conventional program can compute, but perceptrons are hard to train because small changes to weights can flip the output abruptly. To make learning smoother, the book replaces them with sigmoid neurons, which use the same weighted sum plus bias but pass it through an S‑shaped activation that outputs a continuous value between 0 and 1 instead of a hard 0/1.

Architectures

Classification vs. Regression in Neural Networks

The fundamental difference between classification and regression neural networks lies in what they predict and how they measure success.

what the network predicts

Classification

  • ·Predicts discrete class labels
  • ·e.g. "spam" or "legitimate" email
  • ·e.g. "cat" or "dog" in images

Regression

  • ·Predicts continuous quantities
  • ·e.g. house prices or temperature
  • ·e.g. stock prices

Recurrent neural networks (RNNs) and convolutional neural networks (CNNs) are other neural networks frequently used in machine learning and deep learning tasks.

Transformers are a type of neural network architecture, most popular neural network architecture and it is what most modern AI systems are built upon.

Layers as a transformation pipeline

A feedforward neural network arranges neurons in layers where information flows strictly from input to output without cycles.

  • The input layer just encodes raw data, such as 784 pixels for a 28×28 digit image.
  • One or more hidden layers transform these inputs into increasingly abstract internal features.
  • The output layer produces the final prediction, e.g., 10 activations representing which digit (0–9) the network thinks it sees.
Loading diagram…

From a web‑dev perspective, each hidden layer is like a middleware stage in a pipeline: it transforms the request representation one step closer to the business decision the final handler will make. Deep networks simply stack more of these stages, letting early layers learn simple patterns and later layers assemble them into higher‑level concepts.

Training

To train a network, you need labeled examples and a way to measure how wrong the network is on those examples through a cost function. Training becomes an optimization problem: adjust all the weights and biases to minimize this cost across the training data.

Learning via cost minimization and gradient descent

Gradient descent repeatedly nudges parameters in the direction that most rapidly decreases the cost, analogous to rolling downhill in an error landscape. Because computing exact gradients on the entire dataset is expensive, it uses stochastic gradient descent (SGD), estimating gradients from small mini‑batches and updating more often, which is the same core idea used in modern deep‑learning frameworks.

Backpropagation: making training scalable

The central algorithm that makes this feasible is backpropagation, which efficiently computes the gradient of the cost with respect to every weight and bias in the network. It works by:

  1. Doing a forward pass to compute activations in all layers.
  2. Computing the error at the output layer, capturing how much each output neuron contributed to the total cost.
  3. Propagating this error backward layer by layer, using the chain rule to determine how each earlier neuron and weight contributed to the final error.
  4. Using these error signals to compute gradients for each weight and bias, which gradient descent then uses to update the parameters.

Conceptually, backprop is like tracing a bug in your frontend back through each component: you start from the wrong output and work backward to see which earlier pieces contributed how much to the problem, then fix those pieces proportionally.

Check your understanding

quiz

Why do sigmoid neurons make training easier than perceptrons?

quiz / select all that apply

Which of these are regression tasks?

quiz

In a feedforward network classifying 28×28 digit images, what does the output layer produce?

quiz / arrange in order

Put the steps of backpropagation in order.

  1. 01Compute the error at the output layer
  2. 02Propagate the error backward layer by layer
  3. 03Compute gradients and update weights and biases
  4. 04Forward pass to compute activations in all layers