Reinforcement Learning from Human Feedback
Aligning AI models to human preferences helps them become safer, smarter, easier to use and tuned to the exact style the creator desires. Reinforcement Learning from Human Feedback (RLHF) is the process of using human responses to a model’s output to shape its alignment and therefore its behaviour.
Specificaties
| ISBN/EAN | 9781633434301 |
| Auteur | Nathan Lambert |
| Uitgever | Van Ditmar Boekenimport B.V. |
| Taal | Engels |
| Uitvoering | Paperback / gebrocheerd |
| Pagina's | 312 |
| Lengte | |
| Breedte |
