📢 Nouveau : recevez les offres du jour sur notre canal WhatsApp
Jobiglo

Aucun resultat.

RLHF Specialist (Remote)

Odixcity Consulting · Djibouti

Remote
Remote Mid 🇬🇧 English
Python PyTorch TensorFlow JAX LoRA QLoRA LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI TRL Transformer Reinforcement Learning Axolotl

Description du poste

About the role

We are looking for an RLHF Specialist to design and operate reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factuality and alignment of large language models. The role is fully remote and collaborates closely with machine‑learning engineers and annotation teams.

Key responsibilities

  • Generate high‑quality preference data by comparing model outputs and ranking them on helpfulness, honesty and harmlessness.
  • Design multi‑turn prompts to stress‑test model reasoning and safety, and write chain‑of‑thought rationales for reward‑model training.
  • Work with ML engineers to analyse failure modes, identify data gaps and propose data‑driven interventions.
  • Develop and iterate annotation strategies, ensuring consistent preference scoring across a global team.
  • Probe models for biases, hallucinations or reward‑hacking, document findings and suggest corrective actions.
  • Maintain a personal benchmark set, regularly re‑evaluating new model versions against historical performance.
  • Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.

Required profile

  • At least 2 years of experience in data annotation, model evaluation, computational linguistics or AI safety.
  • Strong understanding of reinforcement‑learning algorithms (PPO, trust‑region methods) and their application to language generation.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Familiarity with constitutional AI or self‑alignment techniques and contributions to open‑source alignment libraries.
  • Experience managing human‑in‑the‑loop workflows with annotation platforms.

Required skills

  • Python programming
  • PyTorch, TensorFlow or JAX
  • LoRA / QLoRA fine‑tuning techniques
  • Annotation tools such as LabelBox, Scale AI, Snorkel
  • Cloud ML services: AWS SageMaker, GCP Vertex AI
  • Open‑source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 1 mois

Expire dans 4 jours

53 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

Odixcity Consulting

Djibouti