Headquartered in Abu Dhabi, United Arab Emirates, is an enterprise technology company specializing in AI infrastructure. It builds platforms that integrate intelligence into core operations, delivering real-time recommendations, automation, and human-like interactions through robust MLOps systems. The company focuses on scaling secure, reliable AI solutions for mission-critical environments, combining deep technical expertise with practical engineering to close the gap between model development and production.
Job Summary
As a Senior MLOps Engineer at AI71, you will play a pivotal role in shaping the infrastructure that powers our AI-enabled enterprise platform—an early-stage system designed to automate complex workflows, deliver real-time insights, and enable data-driven decision-making at scale. Unlike traditional AI solutions, our platform embeds intelligence at its core, driving predictive recommendations, intelligent automation, and natural-language interfaces that transform how users interact with complex systems. This role is focused on building the foundational layers that ensure models are deployed, monitored, and maintained reliably for enterprise customers. You will design and construct the model serving infrastructure—including real-time inference, batch processing, and LLM endpoints—while establishing the CI/CD pipelines, observability frameworks, and evaluation systems required to support mission-critical AI workloads. Your work will bridge the gap between model development and production, ensuring seamless transitions from notebook experiments to scalable, secure, and cost-efficient deployments. Given the early-stage nature of the platform, you will have the opportunity to define architectural standards, deployment patterns, and internal tooling that will shape how AI capabilities are integrated and managed at AI71. This is a platform engineering role with a strong emphasis on reliability, security, and scalability—areas where enterprise customers demand rigorous controls, observability, and fail-safes before adopting AI solutions. Your decisions will set the trajectory for future engineering efforts, ensuring that the next twenty models are deployed with greater efficiency than the first. The challenge lies in balancing the unpredictability of modern AI with the reliability expectations of enterprise environments. You will collaborate closely with data scientists and ML engineers to streamline workflows, eliminate redundant efforts, and establish best practices for model lifecycle management. From infrastructure as code to security posture, your work will ensure that the platform meets the stringent requirements of regulated industries while remaining adaptable to evolving AI advancements. This is not a modeling role; it is an infrastructure role. Your expertise will be leveraged to solve complex problems at the intersection of software engineering and AI operations, where the stakes are high, and the impact is enterprise-wide.
Key Responsibilities
- Design and build the model serving layer, including real-time inference, batch processing, and LLM endpoints, ensuring enterprise-grade latency and availability requirements are met.
- Develop and implement CI/CD pipelines for models, covering automated training, testing, versioning, staged rollouts, and rollback mechanisms that integrate seamlessly into engineering workflows.
- Establish and maintain a comprehensive observability layer to monitor performance, model drift, data quality, cost efficiency, and output quality, ensuring proactive issue detection before customer impact.
- Own the evaluation infrastructure for AI features, including regression testing for non-deterministic outputs and ensuring robust validation processes for production models.
- Build and maintain infrastructure as code across all environments, adhering to the security and compliance standards required for enterprise and regulated customer deployments.
- Collaborate closely with data scientists and ML engineers to streamline the transition from notebook-based development to production, eliminating redundant rewrites and ensuring consistency.
- Define and enforce architectural standards, reference patterns, and internal tooling to optimize the deployment and operational efficiency of future models, reducing complexity for subsequent teams.
Required Qualifications
- Substantial hands-on experience running machine learning systems in production, with a focus on platform engineering rather than model development.
- Strong expertise in Kubernetes and containerization, including ownership of cluster-level decisions and architecture rather than consuming pre-built solutions.
- Proven experience with infrastructure as code (Terraform or equivalent) and end-to-end CI/CD pipeline ownership for production systems.
- Deep familiarity with model serving frameworks such as KServe, Seldon, BentoML, vLLM, Triton, or comparable technologies.
- Solid software engineering skills in Python, with experience writing maintainable code for production environments.
- Hands-on experience with ML lifecycle tooling, including MLflow, Kubeflow, Airflow, feature stores, or equivalent custom-built solutions.
- Practical experience implementing monitoring and observability for ML systems, using tools like Prometheus, Grafana, or similar, with a focus on ML-specific failure modes and performance tracking.
- In-depth knowledge of cloud platforms (AWS, Azure, GCP, or sovereign cloud environments) and the ability to design scalable, secure architectures.
- Judgment to balance platform complexity with company stage, avoiding both fragile scripts and overly abstracted systems that may not align with current needs.
Required Skills
- Substantial hands-on experience running machine learning systems in production, with a focus on platform infrastructure rather than model development
- Deep expertise in Kubernetes and containerization, including ownership of cluster-level architecture and decision-making
- Proficiency in Infrastructure as Code (Terraform or equivalent) and end-to-end CI/CD pipeline ownership for production environments
- Experience with model serving frameworks such as KServe, Seldon, BentoML, vLLM, or Triton, tailored for enterprise-scale deployments
- Strong software engineering skills in Python, with an emphasis on maintainable, production-grade code
- Familiarity with ML lifecycle tooling, including MLflow, Kubeflow, Airflow, or equivalent feature stores and pipeline orchestration systems
- Practical experience implementing monitoring and observability solutions (Prometheus, Grafana, or comparable) for ML-specific failure modes, including performance, drift, data quality, and output quality
- In-depth knowledge of cloud platforms (AWS, Azure, GCP, or sovereign cloud environments) and their application in AI/ML infrastructure
- Judgment in designing platform solutions that balance scalability, reliability, and adaptability to the company’s growth stage, avoiding over-engineering or fragile scripts
- Collaborative mindset to work closely with data scientists and ML engineers to streamline the transition from notebook-based development to production-ready deployments
Work Location Details
- Location: Abu Dhabi, UAE
- Work Arrangement: On-site (hybrid option to be confirmed)
Employment Type
- Employment Status: Permanent
- Work Arrangement: Full-time
Team Overview: Platform Engineering
This role will join the Platform Engineering team, a specialized group focused on designing, building, and maintaining the foundational infrastructure that enables seamless development, deployment, and operations across the organization. The team collaborates closely with engineering, product, and operations teams to ensure scalability, reliability, and efficiency of technical platforms. By leveraging modern engineering practices and technologies, Platform Engineering drives innovation while reducing operational overhead, allowing other teams to focus on delivering high-impact solutions.
Reports To
This role will report to a position that will be confirmed during the hiring process. The exact reporting structure will be finalized based on organizational needs and team alignment.
The Product and Challenge
AI-Enabled Enterprise Platform
We are developing an AI-powered enterprise platform designed to automate complex workflows, deliver real-time insights, and enable organizations to make faster, data-driven decisions at scale. Unlike traditional solutions, intelligence is not an add-on feature—it is the foundation of the platform, driving predictive recommendations, intelligent automation, and natural-language interfaces that allow users to interact with sophisticated systems conversationally.
Enterprise-Grade Requirements
Our target customers are large enterprises with critical needs: reliability, security, and trust in mission-critical systems. This demands solving uniquely challenging engineering problems, including integrating AI unpredictability with traditional software reliability, scaling models for enterprise use, and building robust guardrails, observability, and controls that meet the stringent expectations of production environments.
Early-Stage Opportunity
As a pioneering effort, we are defining how AI capabilities are architected, integrated with core systems, and continuously evaluated and improved. This role presents an opportunity to shape foundational decisions in a space where best practices are still evolving.
Key Responsibilities
- Design and build the model serving layer, including real-time inference, batch processing, and LLM endpoints, ensuring they meet the latency and availability requirements for enterprise workloads.
- Develop and maintain CI/CD pipelines for models, covering automated training, testing, versioning, staged rollouts, and rollback mechanisms that engineers actively adopt rather than bypass.
- Establish and maintain the observability layer, encompassing performance monitoring, drift detection, data quality, cost tracking, and quality-of-output metrics to proactively identify issues before they impact customers.
- Own the evaluation infrastructure for AI features, including regression testing frameworks for non-deterministic outputs to ensure reliability and consistency.
- Build and uphold infrastructure as code across all environments, adhering to the security standards required by enterprise and regulated customers.
- Collaborate closely with data scientists and machine learning engineers to streamline the transition from notebook-based development to production, eliminating redundant rewrites for each deployment.
- Define and enforce reference architectures, deployment patterns, and internal tooling to create scalable, repeatable processes that make subsequent model deployments progressively easier to execute.
Useful But Not Essential
- Experience with LLM serving and inference optimisation techniques, including quantisation, batching, KV caching, and GPU scheduling.
- Familiarity with RAG pipeline operations, particularly retrieval quality monitoring.
- Knowledge of LLM evaluation and guardrail tooling.
- Experience in regulated or security-sensitive environments such as financial services, government, defence, or healthcare.
- Exposure to data engineering concepts, including streaming architectures, lakehouse architectures, and pipeline reliability.
What This Role Is Not
Modeling Focus
This is not a modeling role. The position does not involve feature engineering, experiment design, or model architecture as primary responsibilities. Candidates whose expertise lies primarily in training models or whose infrastructure experience is secondary may find this role misaligned with their strengths.
Platform Maintenance
This role is not intended for maintaining an existing platform. There is minimal legacy infrastructure to inherit, which may appeal to some candidates while presenting a challenge to others. The opportunity lies in shaping foundational decisions from the ground up.
Why Join Now
Foundational Ownership
This role offers unparalleled influence over AI71’s AI shipping framework. As a key decision-maker, you will establish the foundational patterns for how AI is developed and deployed—a level of ownership that diminishes as the platform matures. Your contributions will shape the long-term direction of the company’s technical and operational standards.
High-Impact Visibility
Your work will be highly visible, with decisions carrying lasting significance. The platform is designed for enterprise customers, whose rigorous standards ensure the work is both challenging and impactful. This creates an environment where your expertise directly drives meaningful progress.
Enterprise-Grade Challenges
Building for enterprise clients demands precision, scalability, and innovation. The high standards of these customers will push your work to new levels of difficulty, offering a rare opportunity to solve complex problems at the forefront of AI development.
Hiring Process
- 1. An initial conversation to explore fit and expectations.
- 2. A technical discussion focused on systems design and architecture.
- 3. A collaborative session with the wider engineering team to assess cultural alignment and technical collaboration.