Exhibit 6.1
1Data collection
Huge quantities of text are gathered from sources such as websites, books and code. Selection and cleaning influence what a model can learn, how well it performs and which biases it may reproduce.
Why it is in the museum
A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.
What supports this exhibit
Curatorial synthesis of technical literature
That data selection and curation influence model knowledge, quality and bias.
Main source: Technical literature on pretraining, data curation and scaling of modern LLMs.
The source is listed here, but no direct link is currently available in the register.
What to keep in mind: General process; details vary by model.
Ask the exhibit
This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.
Start with one of the suggested questions above, or type your own.