Job Description

<div>Technical Skills<br><br>· 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Engineering.<br><br>· Expertise in AWS services such as EC2, S3, RDS, IAM, VPC, Lambda, CloudWatch, etc.<br><br>· Strong knowledge of Kubernetes and container orchestration best practices.<br><br>· Experience managing services on Amazon ECS (Fargate or EC2).<br><br>· Proficient in infrastructure-as-code tools like Terraform, CloudFormation, or Pulumi.<br><br>· Skilled in scripting languages such as Python, Bash, or Go.<br><br>· Solid grasp of networking, load balancing, DNS, and firewall rules in cloud environments.<br><br>· Deep understanding of microservices architectures, API gateways, and service meshes.<br><br>Soft Skills<br><br>· Proven leadership and cross-functional collaboration skills.<br><br>· Strong problem-solving and incident-resolution mindset.<br><br>· Clear communication, documentation, and stakeholder reporting abilities.<br><br>· Passion for continuous improvement and automation.<br><br>Preferred Qualifications<br><br>· AWS certifications such as AWS Certified DevOps Engineer, Solutions Architect – Professional, or equivalent.<br><br>· Familiarity with service meshes like Istio or Linkerd.<br><br>· Experience with serverless architectures and event-driven systems.<br><br>· Knowledge of regulatory compliance (SOC2, ISO 27001, GDPR) in cloud environments.<br><br>Skills – AWS Cloud, CICD, EC2, Kubernete, Grafana, Datadog, Python<br><br>SRE- AWS<br><br>Job Summary<br>We are looking for an experienced and driven Senior Site Reliability Engineer (SRE) to architect, implement, and maintain robust cloud infrastructure. This role demands a deep understanding of AWS, Kubernetes, ECS, and the ability to build scalable, secure, and highly available infrastructure from scratch. The ideal candidate will be a strong advocate for DevOps principles, automation, and reliability, and will possess the skills to support and optimize complex microservices-based architectures.<br>Key Responsibilities<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Infrastructure Design &amp; Implementation<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Design and build highly scalable, fault-tolerant, and secure cloud infrastructure using AWS, Kubernetes, and ECS.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Lead efforts in infrastructure as code (IaC) using tools like Terraform or CloudFormation.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Develop and enforce best practices for infrastructure provisioning, security, and cost optimization.<br>System Reliability &amp; Performance<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Ensure availability, performance, scalability, and security of production systems.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Implement observability strategies including monitoring, logging, and alerting using tools such as Prometheus, Grafana, ELK, or Datadog.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Analyse system performance metrics and proactively identify potential issues and bottlenecks.<br>DevOps &amp; Automation<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Build and maintain CI/CD pipelines to streamline code deployments across environments.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Drive automation in infrastructure provisioning, configuration management, and operational tasks.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Ensure repeatable and reliable deployments using containers and orchestration tools like Kubernetes and ECS.<br><br><br><br>Service Management<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Own the SRE lifecycle, including incident management, postmortems, root cause analysis, and runbook creation.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Collaborate closely with development and QA teams to ensure seamless microservices integration, deployment, and lifecycle management.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Maintain service-level objectives (SLOs), service-level agreements (SLAs), and error budgets.<br>Security &amp; Compliance<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Implement and enforce cloud security best practices for networking, identity and access management, and data protection.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Support audits, compliance assessments, and vulnerability remediation.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Monitor for security anomalies and work with security teams to respond to threats.<br>Technical Skills<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Engineering.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Expertise in AWS services such as EC2, S3, RDS, IAM, VPC, Lambda, CloudWatch, etc.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Strong knowledge of Kubernetes and container orchestration best practices.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Experience managing services on Amazon ECS (Fargate or EC2).<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Proficient in infrastructure-as-code tools like Terraform, CloudFormation, or Pulumi.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Skilled in scripting languages such as Python, Bash, or Go.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Solid grasp of networking, load balancing, DNS, and firewall rules in cloud environments.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Deep understanding of microservices architectures, API gateways, and service meshes.<br>Soft Skills<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Proven leadership and cross-functional collaboration skills.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Strong problem-solving and incident-resolution mindset.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Clear communication, documentation, and stakeholder reporting abilities.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Passion for continuous improvement and automation.<br>Preferred Qualifications<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>AWS certifications such as AWS Certified DevOps Engineer, Solutions Architect – Professional, or equivalent.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Familiarity with service meshes like Istio or Linkerd.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Experience with serverless architectures and event-driven systems.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Knowledge of regulatory compliance (SOC2, ISO 27001, GDPR) in cloud environments.<br>Skills – AWS Cloud, CICD, EC2, Kubernete, Grafana, Datadog, Python<br>Key Responsibilities:<br>Cloud Platform: GCP<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Infrastructure Automation: Design, implement, and manage infrastructure as code using Terraform to provision and manage GCP resources.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Container Orchestration: Deploy and manage Kubernetes clusters, ensuring efficient operation of containerized applications.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Continuous Integration/Continuous Deployment (CI/CD): Develop and maintain CI/CD pipelines using Jenkins to automate application build, test, and deployment processes.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Containerization: Collaborate with development teams to containerize applications using Docker and manage deployments with Helm Charts.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Code Quality Assurance: Integrate and manage SonarQube to ensure code quality and security standards are met.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Monitoring and Logging: Implement and manage monitoring solutions using Datadog to ensure system health, performance, and security.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Collaboration: Work closely with cross-functional teams, including developers, QA, and operations, to streamline processes and improve productivity.<br>Requirements:<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Experience: 5+ years in DevOps or cloud engineering roles, with at least 3 years of relevant experience in the specified technologies.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Technical Proficiency:<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Hands-on experience with GCP services and architecture.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Proficiency in Terraform for infrastructure as code implementations.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Strong understanding and experience with Kubernetes and Docker.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Experience in setting up and managing CI/CD pipelines using Jenkins.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Familiarity with Helm Charts for application deployment.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Experience with SonarQube for code quality analysis.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Proficiency in monitoring and logging tools, particularly Datadog.<br>•<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Scripting Skills: Proficiency in scripting languages such as Bash or Python is an added advantage.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Strong problem-solving abilities and analytical thinking.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Excellent communication skills, both verbal and written.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Ability to work collaboratively in a team environment.<br>o<span style="white-space:pre;">&nbsp; &nbsp;&nbsp;</span>Strong organizational and time management skills.<br><br>Skills – Terraform, Kubernetes, Cluster, Docker, GCP, SonarQube</div>


Job Details

Role Level: Mid-Level Work Type: Full-Time
Country: India City: Visakhapatnam ,Andhra Pradesh
Company Website: http://www.sailssoftware.com/ Job Function: DevOps & QA
Company Industry/
Sector:
Software Development

What We Offer


About the Company

Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.

Report

Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.


ad 1
Talentmate Instagram Talentmate Facebook Talentmate YouTube Talentmate LinkedIn