Like I'm 5™

Lesson60 seconds

What is RLHF?

Teaching a model using human preference ratings.

Step 1 of 9

Simple Definition

Reinforcement learning from human feedback (RLHF) optimises a model using preference data so outputs better match what people rate as helpful and .

Explanation level

Tap a persona to hear the same idea in a totally different voice.

Try “Like I'm 5™” for a picture-book version - not the grown-up wording above.

Notebook

Your notes

Notes save automatically on this device

Knowledge graph

RLHF
Done Here Next