Q-learning is the 1989 reinforcement-learning algorithm that lets an agent learn action values through rewards and penalties.
Chris Watkins introduced Q-learning in his doctoral work at the University of Cambridge. The algorithm estimates a value called Q for taking a particular action in a particular state. By repeatedly interacting with an environment, the agent updates those estimates using rewards and the value of a possible future state.
A major feature of Q-learning is that it can learn an optimal policy without first being given a model of the environment. This makes it a model-free reinforcement-learning method. The agent can still explore actions it has not tried often while gradually favoring actions with higher estimated long-term value.
Q-learning is not the same as supervised learning. It does not receive a correct label for every decision; instead, it receives feedback from the consequences of its behavior. Later systems, including deep Q-networks, combined the method with neural networks.