Museum of Artificial IntelligenceELEL

Exhibit 6.5

5

Alignment

People, and sometimes other models, evaluate answers. Training then pushes the model toward responses judged more helpful, honest or safe. A well-known approach is reinforcement learning from human feedback (RLHF).

Alignment
AI-generated illustration

Why it is in the museum

A modern language model does not appear fully formed. It passes through data collection, training, adaptation, evaluation and, finally, answer generation.

What supports this exhibit

Primary research paper

The use of human preferences and reinforcement learning to shape model behaviour.

Main source: Ouyang, L. et al. (2022), “Training language models to follow instructions with human feedback”.

Open the source

What to keep in mind: Primary source; modern systems also use methods beyond RLHF.

Ask the exhibit

This small guide uses only the information documented on this exhibit page. If your question goes beyond that evidence, it will say so instead of inventing an answer.

Start with one of the suggested questions above, or type your own.