Skip to main content
The new Teacher Workspace is here. Your first 3 assignments are free. Try it →

Stochastic gradient descent

Stochastic gradient descent is a training algorithm that updates a model’s weights using one example or a small batch at a time. In Intro to Cognitive Science, you see it when neural networks are discussed as models of learning and pattern recognition.

Last updated July 2026

What is stochastic gradient descent?

Stochastic gradient descent, or SGD, is a way to train an artificial neural network by changing its weights after each example, or after a very small batch of examples, instead of waiting to process the full dataset. In Intro to Cognitive Science, it shows up when you study how learning algorithms let a model adjust its behavior from experience, a little like how humans improve with feedback.

The basic idea is simple: the network makes a prediction, compares that prediction to the correct answer, and then uses the error to nudge the weights. Those nudges are based on the gradient of the loss function, which tells the model which direction reduces error fastest. Because SGD uses only a slice of the data each step, the updates are quick, but they are also noisy.

That noise is not just a flaw. Since each training example is a little different, SGD can move the model around instead of making perfectly smooth progress. In practice, that can help it move out of shallow local minima or plateaus in the loss landscape, which is one reason it is so common in deep learning. The tradeoff is that the path to a low loss value is less steady than with full gradient descent.

This is where learning rate matters. If the learning rate is too large, the weights can jump past a good solution and the loss can bounce around or even blow up. If it is too small, training crawls and may look stuck even though the model is still improving.

You will also see versions like mini-batch gradient descent, which sits between full-batch and pure SGD. Instead of one example at a time, it updates from a small group of examples, which keeps some speed while reducing the randomness. In cognitive science classes, that comparison matters because it shows how model training balances efficiency, stability, and generalization.

Why stochastic gradient descent matters in Intro to Cognitive Science

SGD matters in Intro to Cognitive Science because it connects the math of neural networks to the bigger question of how systems learn from data. When you study artificial neural networks, you are not just looking at a diagram of nodes and weights. You are also looking at the procedure that changes those weights over time, and SGD is one of the main procedures that makes learning happen.

It also gives you a clean way to compare different learning algorithms. Full gradient descent uses the whole dataset each step, so it is more stable but slower. SGD is faster and scales better to large datasets, which is why it shows up so often in modern machine learning and in course examples about image recognition, language, or pattern detection.

In a cognitive science setting, SGD also helps you think about learning as an adaptive process rather than a fixed structure. That makes it useful when you discuss how computational models can mimic aspects of human learning, such as gradual improvement after repeated feedback. Even if the network is not a brain, the training process gives you a model for how complex behavior can emerge from many small updates.

Keep studying Intro to Cognitive Science Unit 7

Official unit cheatsheet

open one-pager

How stochastic gradient descent connects across the course

Gradient Descent

Gradient descent is the broader training method that SGD belongs to. The difference is in how much data the algorithm uses for each update. If you understand full gradient descent first, SGD makes more sense as the faster, noisier version that trades smoothness for speed and scalability.

Mini-batch Gradient Descent

Mini-batch gradient descent sits between full-batch gradient descent and SGD. It updates weights using a small group of examples, which lowers the randomness of single-example updates while keeping training efficient. In class discussions, this is often the version you compare to pure SGD when talking about practical neural network training.

Loss Function

SGD exists to reduce the loss function. The loss tells the model how wrong its predictions are, and the gradient of that loss tells SGD which direction to move the weights. If you cannot track the loss, you cannot really explain what SGD is optimizing.

artificial neural networks

SGD is one of the main ways artificial neural networks learn from data. The network architecture gives you the layers and connections, but SGD supplies the adjustment process that changes weights after each example or batch. That makes it a learning rule, not just a math trick.

Is stochastic gradient descent on the Intro to Cognitive Science exam?

A quiz or problem-set question will usually ask you to identify SGD as the algorithm that updates weights after each example or small batch, then explain why that makes training faster but noisier. You might also compare it to full gradient descent or mini-batch gradient descent and describe the effect of the learning rate. In a short-answer prompt, a strong response traces the loop: prediction, loss, gradient, weight update, repeat. If the question uses a neural network diagram or training curve, you should be able to point to where SGD changes the parameters and why the loss may bounce around before settling.

Stochastic gradient descent vs Gradient Descent

People often use these terms as if they mean the same thing, but SGD is the stochastic version of gradient descent. Plain gradient descent usually means using the whole dataset for each update, while SGD uses one example or a small subset, which makes the updates faster and noisier.

Key things to remember about stochastic gradient descent

  • Stochastic gradient descent updates a neural network’s weights one example, or one small batch, at a time.

  • It is designed to reduce the loss function quickly, even if the path to a good solution looks messy.

  • The randomness in SGD can help training move past shallow local minima, but it can also make the loss curve bounce around.

  • The learning rate controls how big each weight update is, so it strongly affects whether training converges or diverges.

  • In Intro to Cognitive Science, SGD matters because it explains how artificial neural networks actually learn from data.

Frequently asked questions about stochastic gradient descent

What is stochastic gradient descent in Intro to Cognitive Science?

It is a training algorithm for neural networks that adjusts weights after each example or small batch. In cognitive science, you usually meet it when the course talks about how computational models learn from experience and reduce prediction error over time.

How is SGD different from gradient descent?

Gradient descent often means using the full dataset for each update, while SGD uses one example or a small sample. That makes SGD faster and more scalable, but also noisier because each update is based on less information.

Why does SGD use random or noisy updates?

Because it does not wait for the whole dataset, each step is based on a limited slice of training data. That randomness can make the loss move up and down a bit, but it can also help the model avoid getting stuck in weak local minima.

Where does SGD show up in cognitive science classes?

You usually see it in lessons on artificial neural networks, machine learning, and learning algorithms. It may come up in a worked example about training a model, a comparison with mini-batch methods, or a discussion of how computational systems learn patterns.