📢 جديد: تابع عروض اليوم على قناتنا في واتساب
Jobiglo

لا توجد نتائج.

RLHF Specialist (Remote)

Odixcity Consulting · Djibouti

Remote
Remote Mid 🇬🇧 English
Python PyTorch TensorFlow JAX LoRA QLoRA LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI TRL Transformer Reinforcement Learning Axolotl

وصف الوظيفة

About the role

We are looking for an RLHF Specialist to design and operate reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factuality and alignment of large language models. The role is fully remote and collaborates closely with machine‑learning engineers and annotation teams.

Key responsibilities

  • Generate high‑quality preference data by comparing model outputs and ranking them on helpfulness, honesty and harmlessness.
  • Design multi‑turn prompts to stress‑test model reasoning and safety, and write chain‑of‑thought rationales for reward‑model training.
  • Work with ML engineers to analyse failure modes, identify data gaps and propose data‑driven interventions.
  • Develop and iterate annotation strategies, ensuring consistent preference scoring across a global team.
  • Probe models for biases, hallucinations or reward‑hacking, document findings and suggest corrective actions.
  • Maintain a personal benchmark set, regularly re‑evaluating new model versions against historical performance.
  • Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.

Required profile

  • At least 2 years of experience in data annotation, model evaluation, computational linguistics or AI safety.
  • Strong understanding of reinforcement‑learning algorithms (PPO, trust‑region methods) and their application to language generation.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Familiarity with constitutional AI or self‑alignment techniques and contributions to open‑source alignment libraries.
  • Experience managing human‑in‑the‑loop workflows with annotation platforms.

Required skills

  • Python programming
  • PyTorch, TensorFlow or JAX
  • LoRA / QLoRA fine‑tuning techniques
  • Annotation tools such as LabelBox, Scale AI, Snorkel
  • Cloud ML services: AWS SageMaker, GCP Vertex AI
  • Open‑source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.

قدم طلبك في 30 ثانية

أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.

بالمتابعة، أنت توافق على شروط الاستخدام.

لديك حساب بالفعل؟ تسجيل الدخول

💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ شهر

ينتهي 3 أيام من الآن

56 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Odixcity Consulting

Djibouti