Exhibit 5.4
2018GPT and BERT
OpenAI and Google show that a model first pretrained on a very large body of text can later be adapted to many different tasks.

Why it is in the museum
Language is treated as a prediction problem. From Shannon’s statistical sequences, the story leads to Transformers and large language models.
What supports this exhibit
Primary research papers
The generalisation of Transformer pretraining followed by adaptation to many downstream tasks.
Main source: Radford et al. (2018), Generative Pre-Training; Devlin et al. (2018), BERT.
What to keep in mind: Primary sources.
Ask the exhibit
This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.
Start with one of the suggested questions above, or type your own.