Skip to main content
The new Teacher Workspace is here. Your first 3 assignments are free. Try it →

Training data

Training data is the set of examples a machine learning model learns from in Intro to Cognitive Science. It gives the model input-output patterns it can use to make predictions on new data.

Last updated July 2026

What is training data?

Training data is the set of examples a machine learning system uses to learn a pattern in Intro to Cognitive Science, usually paired with labels or target outcomes in supervised learning. Instead of being told a rule in words, the model sees many examples and adjusts its internal parameters so it can map inputs to outputs.

That makes training data more than just a pile of examples. It is the evidence the system uses to estimate what matters in the input, whether that is recognizing a face, classifying a sound, or predicting a choice. If the examples are noisy, too narrow, or missing important cases, the model may learn the wrong pattern.

A useful way to think about it is: training data comes first, model learning happens next, and predictions on new data happen after that. The model does not memorize the training set perfectly if it is working well. It learns a general rule that should still work when the input changes a little.

This is why preprocessing matters in this course context. Missing values, inconsistent scales, and badly encoded categories can distort what the model sees during learning. Normalization can keep one feature from overpowering the others, and cleaning the data can make the learned pattern more stable.

Training data is also usually split from validation and test data. The training set is where the model learns, the validation set helps tune settings, and the test set checks how well the model handles new examples. If you train on data that is not representative of the real world, the model may look strong in practice but fail when it meets unfamiliar cases.

In cognitive science, this concept matters because machine learning is often used as a model of cognition. When researchers compare human learning to machine learning, the structure and quality of the training data help explain what the system can and cannot learn from experience.

Why training data matters in Intro to Cognitive Science

Training data matters in Intro to Cognitive Science because it is the bridge between raw examples and a model that can classify, predict, or make decisions. When the course discusses machine learning and cognitive systems, training data is the material that shapes what the system learns from experience.

It also gives you a way to explain performance problems. If a model misses rare cases, the training data may not include enough of them. If it makes biased predictions, the training set may be imbalanced, so the model mostly sees one class and treats it like the default answer.

This term also connects machine learning to human cognition. Cognitive science often compares how people learn from limited, messy experience to how models learn from datasets. That comparison gets sharper when you ask what kinds of examples are available, which ones are repeated, and whether the learner can generalize beyond them.

You will also see training data as part of the bigger workflow around model building. It sits before validation and testing, after preprocessing, and alongside choices about features and learning method. Once you can track that sequence, it becomes easier to explain why a model succeeds, overfits, or fails on new cases.

Keep studying Intro to Cognitive Science Unit 8

Official unit cheatsheet

open one-pager

How training data connects across the course

Supervised Learning

Training data is the raw material for supervised learning. In that setup, examples come with labels, and the model learns a mapping from inputs to outputs. If you understand training data, you can explain why supervised learning depends so heavily on labeled examples rather than just pattern hunting with no target answer.

Overfitting

Overfitting happens when a model learns the training data too closely, including noise or quirks that do not appear in new data. The model may score well on the training set but perform worse on fresh examples. That makes training data useful for learning, but not enough for judging whether the model will generalize.

Feature Extraction

Feature extraction changes the training data into a form a model can use more effectively. Instead of feeding every raw detail into the system, you choose or create features that capture the useful structure in the examples. In cognitive science, this often connects to how people and machines pick out relevant signals from messy input.

Unsupervised Learning

Training data still matters in unsupervised learning, but the model is not given labels to learn from. The data itself has to reveal structure, like clusters or patterns, without a correct answer attached to each example. Comparing the two helps show how labels change what the learner can discover.

Is training data on the Intro to Cognitive Science exam?

A quiz question might show a dataset, a model output, or a short case description and ask you to identify which examples count as training data and what the model learns from them. You may also be asked to explain why a model does well on its training set but poorly on new cases, which usually points to poor training data quality or overfitting.

In short-answer or essay questions, use the term to trace the learning pipeline: the data is collected, cleaned, split, and then used to fit the model before validation or testing. If the prompt mentions imbalance, missing values, or nonrepresentative samples, connect those features to biased or weak predictions. In a machine learning case study, you can describe how changing the training data would change the system’s behavior.

Training data vs testing data

Training data is what the model learns from, while testing data is what checks how well the model handles brand-new examples. A common mistake is thinking both sets do the same job. In Intro to Cognitive Science, that difference matters because a model can look great on training data and still fail on the test set if it has memorized instead of generalized.

Key things to remember about training data

  • Training data is the example set a machine learning model uses to learn patterns in Intro to Cognitive Science.

  • Good training data is representative, because the model can only generalize well if the examples match the kinds of inputs it will face later.

  • A model learns from training data before it is judged on validation or test data.

  • If the training set is imbalanced, noisy, or too small, the resulting model can become biased or overfit to the wrong patterns.

  • In cognitive science, training data helps compare machine learning with human learning from experience.

Frequently asked questions about training data

What is training data in Intro to Cognitive Science?

Training data is the set of examples a machine learning model uses to learn how inputs connect to outputs. In Intro to Cognitive Science, it shows up when you study supervised learning, classification, and model performance. The examples need to be representative, or the model may learn a pattern that breaks down on new data.

How is training data different from test data?

Training data is used to fit the model, while test data is held back to evaluate whether the model generalizes. If you mix them up, you cannot tell whether the system truly learned anything or just memorized the examples it saw. That distinction is one of the core checks in machine learning workflows.

Why does training data quality matter?

Quality affects what the model can learn. If the data has missing values, inconsistent scaling, or a skewed class balance, the model may learn a distorted pattern. In cognitive science terms, bad training data can produce a system that looks smart in class examples but fails on realistic cases.

Can training data cause bias in a model?

Yes. If one class or type of example shows up much more often than others, the model can become biased toward the majority pattern. That is why class balance and representative sampling matter so much when you build or evaluate cognitive systems.

Training Data | Intro to Cognitive Science | Fiveable