An AI cost function assigns a numerical score to a model’s parameters or to a candidate decision. A learning or optimization algorithm uses that score to compare possibilities and search for one with lower cost—or, under a maximization convention, higher utility. In supervised machine learning, the score commonly aggregates prediction errors across training examples.
How a cost function works in machine learning
Suppose a model with parameters θ makes a prediction f(xᵢ; θ) for each input xᵢ, and yᵢ is the corresponding target. A per-example loss ℓ measures how far a prediction is from its target. A common dataset-level cost is the average of those losses:
J(θ) = (1/n) Σᵢ₌₁ⁿ ℓ(f(xᵢ; θ), yᵢ)
Here, n is the number of training examples. Because predictions depend on θ, changing the model parameters changes J. Training adjusts those parameters to reduce the objective. The University of Toronto’s notes explain this distinction between a single-example loss and an average cost across examples: Gradient descent notes.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
This average is an empirical measure on the training set, not a guarantee about future inputs. It serves as a proxy for performance on the data-generating distribution; a model can lower its training cost without improving its results on unseen examples.
Cost, loss, and objective: related terms with flexible meanings
These labels are often used differently across textbooks and applications, so it is best to define the convention being used rather than claim a universal distinction.
Rank #2
- Loss often means the error for one example, comparing a prediction with its target.
- Cost often means an aggregate, such as the average or sum of losses across a dataset.
- Objective means the function being minimized or maximized. It may refer to the cost or loss, or include additional terms such as regularization.
The University of Toronto notes distinguish individual loss from dataset-average cost, while Stanford HAI’s glossary uses “cost” and “objective” as alternate names. Poole and Mackworth also describe a minimizing objective as often being called a cost, loss, or error function: Artificial Intelligence: Foundations of Computational Agents, optimization chapter; Stanford HAI AI glossary.
Examples of AI cost functions
Regression: mean squared error
For a regression model, mean squared error (MSE) averages the squared differences between predictions and targets. Squaring makes large deviations count more heavily than absolute error does. A conventional factor of one half sometimes appears in the formula; it does not change which parameters minimize the error. See the University of Toronto’s MSE and gradient-descent notes.
Classification: negative log-likelihood
A classifier can be trained by minimizing the negative log-likelihood assigned to the correct class. This is a differentiable surrogate for classification error, not necessarily the final metric used to judge the system. As a result, the training objective and the measure that matters to a user may differ. Stanford’s Speech and Language Processing chapter on logistic regression discusses this class of objective.
Scheduling: weighted soft constraints
In a scheduling problem, hard constraints rule out assignments that are infeasible, while soft constraints assign penalties to undesirable but possible outcomes. An exam timetable might penalize student conflicts, back-to-back exams, or use of less-preferred times or rooms. The objective can add those penalties with weights that express their relative importance, and the optimizer searches for a feasible schedule with low total cost. See Poole and Mackworth’s chapter on constraint satisfaction and scheduling.
Rank #4
How to choose a cost function
There is no single best cost function for every AI task. The appropriate objective depends on the output, the training method, and which errors or preferences matter most. When comparing alternatives, consider:
- Error priorities: Which mistakes are most costly in the real application?
- Large errors and outliers: Should unusually large deviations receive extra weight, as they do under squared error?
- Training compatibility: Can the objective be optimized effectively with the model and learning method?
- Outcome alignment: Does reducing the objective correspond to improving the metric or real-world outcome people care about?
Sometimes the desired final metric is difficult to optimize directly. Training may therefore use a surrogate loss, with validation performance or another criterion used to assess results or decide when to stop. State both what the model optimizes and what that objective is intended to improve.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
What a low training cost does—and does not—show
A lower cost means the model or decision scores better under the particular objective and data used to calculate it. It does not, by itself, show that an AI system will perform well in deployment. A sufficiently flexible model can overfit its training examples, lowering training cost while failing to generalize to new ones. Evaluate behavior on data not used to fit the model and against the outcome that matters for the task.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




