Talentmate
India
9th September 2025
2509-9201-13
As a Observability Engineer under Site Reliability Engineering Team, you will be a crucial part of the team responsible for the availability, performance, and scalability of our cloud platform. You will blend software engineering and systems administration expertise to build and run large-scale, distributed, fault-tolerant systems. Your mission is to ensure our services are reliable and efficient through automation, robust monitoring, and proactive incident response. You will work closely with development teams to build resilient and scalable applications on our Google Cloud Platform (GCP) and Kubernetes-based infrastructure. Having a Strong troubleshooting skills and a methodical approach to problem-solving is a MUST.
Key Responsibilities
Infrastructure as Code (IaC): Design, build, and maintain our core cloud infrastructure on GCP using tools like Terraform and Google Config Connector (KCC) within a GitOps framework.
Automation: Utilize Infrastructure as Code (IaC) with Kubernetes (GKE) and Google Config Connector (KCC), Develop automation scripts and tools (primarily in Python or Go) to reduce operational toil, streamline deployments, and improve system efficiency.
Observability: Implement and manage comprehensive monitoring, logging, and alerting solutions using tools like Prometheus, Open Telemetry, Grafana, and Google Clouds operations suite to gain deep insights into system health.
Reliability & SLOs: Define, measure, and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for critical services. Drive initiatives to meet and exceed these objectives. Develop & promote dashboarding, and actionable alerting across the organization.
Incident Management: Participate in an on-call rotation to respond to and resolve production incidents. Lead blameless post-mortems to identify root causes and implement lasting solutions.
Collaboration: Partner with software engineering teams throughout the development lifecycle to provide guidance on building reliable, scalable, and secure applications. Help them troubleshoot complex issues, improve service performance, and adopt observability best practices.
Enhance Reliability: Analyze observability data to identify trends, uncover potential issues, and drive initiatives to improve system reliability, performance, and cost-efficiency.
Secure and Scale: Manage secrets and system configurations securely using Hashi Corp Vault and ensure the observability platform scales to meet the demands of a growing engineering organization.
Qualifications Required
Role Level: | Mid-Level | Work Type: | Full-Time |
---|---|---|---|
Country: | India | City: | Bengaluru ,Karnataka |
Company Website: | http://www.cmegroup.com | Job Function: | Engineering |
Company Industry/ Sector: |
Financial Services |
Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.
Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.