Job Description

Position summary

The Systems Engineer builds, hardens, and operates the Linux server fleet and core infrastructure services that underpin GPU One (GPUaaS). This role keeps the operating system, networking, and storage layers healthy, secure, and automated so that customer and platform workloads run reliably at scale.

Key responsibilities

·        Administer, patch, and harden a large fleet of Linux servers (RHEL, Rocky, or Ubuntu) across bare-metal and cloud environments

·        Troubleshoot complex OS, kernel, performance, and hardware issues down to root cause

·        Design, configure, and operate TCP/IP networking including routing, VLANs, DNS, DHCP, firewalls, and bonding

·        Deploy and manage NFS storage and other shared file systems for high-throughput workloads

·        Automate provisioning, configuration, and remediation using Ansible and infrastructure-as-code

·        Build and maintain monitoring, alerting, and dashboards with Grafana and related observability tooling

·        Manage the operational workflow through ticketing systems such as ServiceNow (SNOW) and JIRA

·        Own capacity planning, OS lifecycle, and standardized system build and image management

·        Partner with Platform Engineering and SRE on reliability, security, and rollout of new services

·        Document standards, runbooks, and procedures, and drive continuous reduction of operational toil

Required qualifications

·        5+ years in Linux systems administration or infrastructure engineering at scale

·        Strong Linux internals and troubleshooting skills across OS, kernel, storage, and performance

·        Solid TCP/IP networking fundamentals and hands-on network troubleshooting

·        Hands-on experience operating NFS and other shared storage in production

·        Proven experience with configuration management using Ansible

·        Experience with monitoring tools such as Grafana and ticketing tools such as ServiceNow and JIRA

·        Scripting proficiency in Bash and Python for automation

·        Bachelors degree in computer science, engineering, or equivalent experience

Preferred qualifications

·        Experience operating GPU servers, HPC, or AI infrastructure

·        Familiarity with Kubernetes, Slurm, or other cluster schedulers

·        Exposure to storage technologies such as Lustre, GPFS/Spectrum Scale, or Ceph

·        Knowledge of InfiniBand or RDMA and high-performance networking

·        Familiarity or working knowledge of using AI coding tools such as Claude, OpenAI, or others to accelerate automation and troubleshooting

·        Relevant certifications (RHCSA/RHCE, CCNA) or cloud provider certifications


Job Details

Role Level: Mid-Level Work Type: Full-Time
Country: India City: Lahore Pakistan ,Punjab
Company Website: https://primesystemsolutions.com/ Job Function: Engineering
Company Industry/
Sector:
IT Services and IT Consulting

What We Offer


About the Company

Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.

Report

Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.


ad 1
Talentmate Instagram Talentmate Facebook Talentmate YouTube Talentmate LinkedIn