Wiki Hub › AI Automation › Machine Learning
AI Automation

Machine Learning

Last updated: Oct 04, 2026
Machine Learning

Machine learning (ML) is a field of study within artificial intelligence in which computer systems improve at a task by identifying patterns in data, rather than by following only explicitly programmed rules. A machine learning system is typically built by supplying an algorithm with example data. The algorithm produces a model that can then make predictions or decisions about new, unseen data.

Machine learning draws on statistics, optimization, and computer science. It is used in areas such as image and speech recognition, language translation, recommendation systems, fraud detection, and scientific research. The term is often used interchangeably with "artificial intelligence" in popular media, although machine learning is generally considered one subfield of the broader discipline.

Infobox

Field

Artificial intelligence; computer science; statistics

Term popularized by

Arthur Samuel (1959)

Main paradigms

Supervised, unsupervised, semi-supervised, self-supervised, reinforcement learning

Common model families

Linear models, decision trees, ensembles, support vector machines, neural networks

Related fields

Data mining, deep learning, pattern recognition, statistics

Definition

Several definitions of machine learning are in common use. Arthur Samuel, who worked at IBM on a checkers-playing program, is widely credited with describing it in 1959 as a field giving computers the ability to learn without being explicitly programmed. That phrasing is commonly attributed to him, though sources differ on its exact wording.

A more formal definition appears in Tom Mitchell's 1997 textbook Machine Learning. A program is said to learn from experience E with respect to a class of tasks T and a performance measure P if its performance at tasks in T, as measured by P, improves with experience E.

Note: Machine learning is not synonymous with artificial intelligence. Some AI systems, such as early expert systems built from hand-written rules, do not use machine learning.

How machine learning works

Most machine learning systems follow a similar sequence, although details vary by method.

  1. Data collection and preparation. Data is gathered, cleaned, and often split into training, validation, and test sets.
  2. Model selection. A model family is chosen, such as a decision tree or a neural network.
  3. Training. An algorithm adjusts the model's internal parameters to reduce an loss function, a numerical measure of error on the training data.
  4. Evaluation. The model is tested on data it has not seen, to estimate how well it generalizes.
  5. Deployment and monitoring. The model is used in practice and checked over time, since real-world data can change.

A central goal is generalization: performing well on new data, not only on the examples used in training. Two failure modes are widely discussed. Overfitting occurs when a model captures noise or quirks of the training data and performs poorly elsewhere. Underfitting occurs when a model is too simple to capture the underlying pattern.

History

Early foundations (1940s–1960s)

Black-and-white photo of a 1950s computer room with equipment cabinets and an operator at a console.
Early research on learning machines in the 1950s and 1960s used room-sized computers and custom hardware.

Foundations of the field include the mathematical modeling of neurons by Warren McCulloch and Walter Pitts in 1943, and Alan Turing's 1950 paper "Computing Machinery and Intelligence," which discussed the idea of machines that learn. In 1958, Frank Rosenblatt introduced the perceptron, an early trainable neural network model. Samuel's checkers program, developed in the 1950s, is often cited as an early demonstration of a program improving through experience.

Setbacks and revival (1969–1990s)

In 1969, Marvin Minsky and Seymour Papert published Perceptrons, which analyzed the limitations of single-layer networks. Many historians link this work to reduced funding and interest in neural network research during the following years, although the extent of its influence is debated. Interest recovered in the 1980s, notably after a 1986 paper by David Rumelhart, Geoffrey Hinton, and Ronald Williams popularized the backpropagation algorithm for training multi-layer networks. In 1995, Corinna Cortes and Vladimir Vapnik published work on support vector machines, which became a leading method for many tasks.

Deep learning era (2010s–present)

Growth in computing power, particularly graphics processing units (GPUs), and the availability of large datasets enabled deep learning, which uses neural networks with many layers. A frequently cited milestone is the 2012 victory of the AlexNet model in the ImageNet image-classification competition. In 2016, DeepMind's AlphaGo defeated professional Go player Lee Sedol, using a combination of deep neural networks and reinforcement learning. In 2017, researchers at Google introduced the transformer architecture in the paper "Attention Is All You Need," which underlies many subsequent large language models.

Types of machine learning

Type

Training data

Typical goal

Example task

Supervised learning

Labeled examples

Predict a known output

Classifying emails as spam or not spam

Unsupervised learning

Unlabeled data

Find structure

Grouping customers by behavior (clustering)

Semi-supervised learning

Small labeled set plus large unlabeled set

Improve with limited labels

Medical image classification with few annotations

Self-supervised learning

Labels derived from the data itself

Learn general representations

Predicting masked words in text

Reinforcement learning

Rewards from interaction

Learn a decision policy

Game-playing agents, robotic control

Supervised learning

In supervised learning, each training example pairs an input with a desired output. Problems fall into two main groups: classification, where the output is a category, and regression, where the output is a numerical value, such as a house price estimate.

Unsupervised and self-supervised learning

Unsupervised methods look for patterns without labeled answers. Common tasks include clustering, dimensionality reduction, and anomaly detection. Self-supervised learning, which generates its own training signal from raw data, has become prominent in training large language and vision models.

Reinforcement learning

In reinforcement learning, an agent takes actions in an environment and receives rewards or penalties. Over many trials, it learns a strategy that maximizes cumulative reward.

Common methods

Diagram of connected nodes arranged in input, hidden, and output layers.
Structure of a simple feed-forward neural network
  • Linear and logistic regression: Statistical models that predict a value or a probability from weighted input features.
  • Decision trees and random forests: Models that split data through a series of rules; forests combine many trees to improve accuracy.
  • Gradient boosting: An ensemble technique that builds models sequentially, each correcting errors of the previous ones. It is widely used on structured, tabular data.
  • Support vector machines: Methods that find a boundary separating classes with the largest possible margin.
  • Artificial neural networks: Layered models loosely inspired by biological neurons. Variants include convolutional networks (images), recurrent networks (sequences), and transformers (language and other data).

Machine learning, artificial intelligence, and deep learning

The three terms are nested. Artificial intelligence is the broadest, covering any technique that enables machines to perform tasks associated with human intelligence. Machine learning is the subset that learns from data. Deep learning is a subset of machine learning built on multi-layer neural networks. Machine learning also differs from traditional data mining, which emphasizes discovering previously unknown patterns, whereas machine learning often emphasizes prediction on new data. In practice, the two overlap heavily.

Applications

Machine learning is applied across many sectors. Documented uses include:

  • Healthcare: analysis of medical images and prediction of patient risk, typically as decision support under clinical oversight.
  • Finance: credit scoring and detection of fraudulent transactions.
  • Language technology: machine translation, speech recognition, and text generation.
  • Transportation: perception systems in driver-assistance and autonomous-vehicle research.
  • Science: protein structure prediction, notably DeepMind's AlphaFold, and analysis of astronomical or particle-physics data.
  • Commerce and media: product and content recommendation.

Limitations and criticism

Researchers, regulators, and civil society groups have raised several concerns, which are active areas of study and debate.

  • Bias and fairness. Models trained on historical data can reproduce or amplify existing biases. Studies of facial-analysis and hiring systems have documented performance differences across demographic groups.
  • Interpretability. Complex models, especially deep neural networks, are often described as "black boxes" because their internal reasoning is difficult to explain. The field of explainable AI seeks to address this.
  • Data and privacy. Training can require large volumes of data, raising questions about consent, copyright, and the exposure of personal information.
  • Robustness. Models can fail on inputs that differ from training data, and adversarial examples, which are small deliberate changes to input, can cause incorrect outputs.
  • Resource use. Training large models can require significant computing power and energy. Estimates vary widely by model and method.

Governments have begun to regulate certain uses. The European Union's AI Act, adopted in 2024, is one widely discussed example, with obligations phasing in over several years. Provisions and timelines may change, so current official sources should be consulted.

Current status

Machine learning remains an active research and engineering field. Much recent attention has focused on large pretrained models, often called foundation models, and on questions of safety, evaluation, and governance. Methods and benchmarks change quickly, so figures about model size, performance, and adoption date rapidly.

Was this article helpful?
0 of 0 users found this helpful