RLHF in plain terms: what does it solve that SFT cannot?

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below

Reference answer

Then expect these follow-ups

  • Why doesn't SFT already do this?

  • What is reward hacking?

  • How do you scale RLHF to 70B+?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question