📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

RLHF Specialist (Remote)

Odixcity Consulting · Djibouti

Remote
Remote Mid 🇬🇧 English
Python PyTorch TensorFlow JAX LoRA QLoRA LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI TRL Transformer Reinforcement Learning Axolotl

Job description

About the role

We are looking for an RLHF Specialist to design and operate reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factuality and alignment of large language models. The role is fully remote and collaborates closely with machine‑learning engineers and annotation teams.

Key responsibilities

  • Generate high‑quality preference data by comparing model outputs and ranking them on helpfulness, honesty and harmlessness.
  • Design multi‑turn prompts to stress‑test model reasoning and safety, and write chain‑of‑thought rationales for reward‑model training.
  • Work with ML engineers to analyse failure modes, identify data gaps and propose data‑driven interventions.
  • Develop and iterate annotation strategies, ensuring consistent preference scoring across a global team.
  • Probe models for biases, hallucinations or reward‑hacking, document findings and suggest corrective actions.
  • Maintain a personal benchmark set, regularly re‑evaluating new model versions against historical performance.
  • Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.

Required profile

  • At least 2 years of experience in data annotation, model evaluation, computational linguistics or AI safety.
  • Strong understanding of reinforcement‑learning algorithms (PPO, trust‑region methods) and their application to language generation.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Familiarity with constitutional AI or self‑alignment techniques and contributions to open‑source alignment libraries.
  • Experience managing human‑in‑the‑loop workflows with annotation platforms.

Required skills

  • Python programming
  • PyTorch, TensorFlow or JAX
  • LoRA / QLoRA fine‑tuning techniques
  • Annotation tools such as LabelBox, Scale AI, Snorkel
  • Cloud ML services: AWS SageMaker, GCP Vertex AI
  • Open‑source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 2 days from now

57 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Odixcity Consulting

Djibouti