Aller au contenu
Hugging Face Blog·· 2023-04-05sélectionAI Score67

StackLLaMA : guide pratique pour entraîner LLaMA avec RLHF

StackLLaMA: A hands-on guide to train LLaMA with RLHF

AI Introduction

Hugging Face publie StackLLaMA, un modèle LLaMA 7B entraîné avec RLHF pour répondre aux questions de Stack Exchange.

Raison de la recommandation

Le guide détaille tout le pipeline RLHF de StackLLaMA, du jeu de données à l'entraînement PPO, avec des configurations LoRA et des commandes réutilisables.

Source :Hugging Face Blog · huggingface.co