Job Description

There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the worlds most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.



Job responsibilities

  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
  • Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
  • Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
  • Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
  • Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
  • Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
  • Engineer and support routing and switching across enterprise networks.
  • Engineer and support security and L4–L7 network components, including firewalls, load balancers, and proxies.
  • Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
 
 
 
Required qualifications, capabilities, and skills
 
  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
  • Strong hands-on networking skills across enterprise routing/switching and security/L4–L7 components.
  • Strong automation capability using Python, Shell, and Ansible in production operations.
  • SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
  • Ability to work independently, prioritize effectively, and deliver with minimal oversight.
  • Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
  • Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
  • Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
 
Preferred qualifications, capabilities, and skills
 
  • Demonstrate experience with Cisco ACI / fabrics.
  • Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
  • Work effectively in a financial institution or other regulated environment.
 
 


Job Details

Role Level: Associate Work Type: Full-Time
Country: India City: Hyderabad ,Telangana
Company Website: http://www.jpmorganchase.com Job Function: DevOps & QA
Company Industry/
Sector:
Financial Services

What We Offer


About the Company

Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.

Report

Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.


ad 1
Talentmate Instagram Talentmate Facebook Talentmate YouTube Talentmate LinkedIn