هذه الوظيفة لم تعد متاحة
انتهت صلاحية هذه الوظيفة في 22/09/2026. لم تعد تقبل الطلبات.
RLHF Specialist (Remote)
Odixcity Consulting · Djibouti
وصف الوظيفة
About the role
We are looking for an RLHF Specialist to design and operate reinforcement‑learning‑from‑human‑feedback pipelines that improve the safety, factuality and alignment of large language models. The role is fully remote and collaborates closely with machine‑learning engineers and annotation teams.
Key responsibilities
- Generate high‑quality preference data by comparing model outputs and ranking them on helpfulness, honesty and harmlessness.
- Design multi‑turn prompts to stress‑test model reasoning and safety, and write chain‑of‑thought rationales for reward‑model training.
- Work with ML engineers to analyse failure modes, identify data gaps and propose data‑driven interventions.
- Develop and iterate annotation strategies, ensuring consistent preference scoring across a global team.
- Probe models for biases, hallucinations or reward‑hacking, document findings and suggest corrective actions.
- Maintain a personal benchmark set, regularly re‑evaluating new model versions against historical performance.
- Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.
Required profile
- At least 2 years of experience in data annotation, model evaluation, computational linguistics or AI safety.
- Strong understanding of reinforcement‑learning algorithms (PPO, trust‑region methods) and their application to language generation.
- Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
- Familiarity with constitutional AI or self‑alignment techniques and contributions to open‑source alignment libraries.
- Experience managing human‑in‑the‑loop workflows with annotation platforms.
Required skills
- Python programming
- PyTorch, TensorFlow or JAX
- LoRA / QLoRA fine‑tuning techniques
- Annotation tools such as LabelBox, Scale AI, Snorkel
- Cloud ML services: AWS SageMaker, GCP Vertex AI
- Open‑source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl)
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في جيبوتي.
الراتب: French-Speaking Customer Support Specialist (Remote) استنادًا إلى 10 عرض عمل في جيبوتيلديك سؤال حول هذا العرض؟
اطرحه هنا: ستصلك تفاصيل العرض كاملة عبر البريد الإلكتروني، فوراً.
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
Odixcity Consulting
Djibouti
عروض عمل ذات صلة
-
Senior Infrastructure Lead
UTP Systems Djibouti -
Junior Help Desk & Network Systems Administrator (OCONUS)
GovCIO Djibouti -
Audio Visual VTC Field Service Representative
GovCIO Djibouti -
Consultant(e) national(e) – étude de faisabilité Smart Trade Border Platform
Organisation Internationale pour les Migrations (OIM) Djibouti City -
Consultant(e) national(e) pour étude de faisabilité et conception de plateforme numérique intelligente
Organisation Internationale pour les Migrations