Room 5, 1948–2022

Language and the first LLMs

At its core, a language model does something simple: it predicts what comes next. The idea is old. What changed is scale: how much text the model learns from and how much context it can use.

Open the interactive room →

Exhibits in this room

1948

Shannon measures prediction

Claude Shannon shows that text that resembles English can be generated by measuring which symbols or words tend to follow others.

2013

Words become numbers

With word2vec, each word can be represented as a point in a numerical space. Words used in similar contexts tend to end up close together.

2017

The Transformer

Google researchers publish “Attention Is All You Need”. The attention mechanism allows the model to relate each token to other tokens in the sequence.

2018

GPT and BERT

OpenAI and Google show that a model first pretrained on a very large body of text can later be adapted to many different tasks.

2020

GPT-3

With 175 billion parameters, the model can write text, translate and answer questions, often from only a few examples in the prompt.

2022

ChatGPT

On 30 November 2022, a conversational language-model interface is released publicly as a research preview. Generative AI begins to enter everyday life at unprecedented scale.