LSTM was the 1997 recurrent architecture that introduced memory cells and gates for learning long-term sequence dependencies.
LSTM stands for Long Short-Term Memory. Sepp Hochreiter and Jürgen Schmidhuber introduced it to address a major weakness of ordinary recurrent neural networks: information and gradients could fade as sequences became longer. LSTM units use a memory cell and gates that regulate what information is stored, forgotten, and exposed.
This design made recurrent networks more effective for speech recognition, handwriting recognition, language modeling, and time-series prediction. Later versions often included forget gates and other refinements, but the central idea remained controlled memory over time.
LSTM is sometimes confused with GRU, or Gated Recurrent Unit. GRU is a later gated recurrent architecture with a simpler structure and fewer gates. Both can model sequences, but they are not the same model. Transformers later became dominant in many language tasks because they can process relationships across a sequence more directly and efficiently in parallel.