Exhibit 6.6
6Testing and safety
Before release, models are evaluated for errors, harmful behaviour and misuse. Testing and monitoring can continue after deployment.
Why it is in the museum
A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.
What supports this exhibit
Curatorial synthesis of safety practice
The need to evaluate model behaviour and risks before and after release.
Main source: System cards, evaluations and technical literature on model safety and red teaming.
The source is listed here, but no direct link is currently available in the register.
What to keep in mind: Processes differ across organisations and models.
Ask the exhibit
This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.
Start with one of the suggested questions above, or type your own.