+1 (646) 787 1317 | [email protected]
Log In | My Account | Log Out |
A training method used for finetuning LLMs like GPT, where human evaluators assess and score model outputs to train it to align with human preferences.