Job Description

About AI71:

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role:

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71s platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What Youll Do:

  • Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
  • Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.

What Youll Bring:

  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
  • Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
  • Mentorship record — engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
  • Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams. 

Strong Preference:

  • Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
  • Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • On-prem / air-gap ML delivery architecture experience at scale.
  • Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
  • Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.

Nice to Have: 

  • Conference speaking, technical writing, or industry thought leadership.
  • Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
  • C/C++ or CUDA kernel experience for performance-critical paths.
  • Arabic language skills.

Why AI71:

  • Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
  • Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
  • World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.

 


Job Details

Role Level: Executive-Level Work Type: Full-Time
Country: United Arab Emirates City: Abu Dhabi
Company Website: https://ai71.ai/ Job Function: DevOps & QA
Company Industry/
Sector:
Technology Information and Internet

What We Offer


About the Company

Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.

Report

Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.


ad 1
Talentmate Instagram Talentmate Facebook Talentmate YouTube Talentmate LinkedIn