Join a dynamic team shaping the tech backbone of our operations, where your expertise fuels seamless system functionality and innovation.
As a Site Reliability Engineer II at JPMorgan Chase within the Commercial & Investment Bank- Payments Technology team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. The Site Reliability Engineer is responsible for ensuring the reliability, availability, performance, and operational excellence of production services through observability, automation, incident management, resiliency engineering, and continuous improvement. The role partners closely with Engineering and Infrastructure teams to build scalable, self-healing, and highly resilient platforms.
Job responsibilities
Execute small to medium-sized projects independently and progressively take ownership of designing and delivering solutions end-to-end.
Leverage engineering best practices to develop scalable, maintainable, and resilient solutions that improve operational stability and efficiency.
Analyze, troubleshoot, and resolve production incidents, driving root-cause identification and permanent corrective actions.
Improve platform reliability, availability, and operational stability through proactive problem management and resiliency initiatives.
Design, implement, and enhance observability capabilities, including monitoring, alerting, dashboards, telemetry, SLIs, and SLOs.
Monitor production environments, identify anomalies and trends, and proactively address risks using standard observability and operational tooling.
Eliminate operational toil through automation, self-healing solutions, process optimization, and reuse-first engineering practices.
Support incident, problem, and change management processes across applications, infrastructure, and full-stack technology services.
Partner with Engineering, Infrastructure, Product, and Operations teams to influence design decisions with a reliability-first mindset.
Utilize enterprise-approved AI and agentic capabilities to accelerate incident triage, root-cause analysis, and remediation while adhering to data security and governance standards.
Continuously improve service resilience, operational efficiency, and customer experience by driving automation, observability, and reliability engineering best practices.
Required qualifications, capabilities, and skills
3+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services
Ability to code in at least one programming language
Familiar with site reliability concepts, principles, and practices
Familiar with observability such as white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others
Familiarity with containers or a common Server OS such as Linux and Windows
Emerging knowledge of software, applications and technical processes within a given technical discipline (e.g., Cloud, artificial intelligence, Android, etc.)
Emerging knowledge of continuous integration and continuous delivery tools like Jenkins, GitLab, or Terraform
Emerging knowledge of common networking technologies
Ability to work in a large, collaborative team and demonstrates the willingness to vocalize ideas with peers and managers
Understanding of how to prioritize and adjust work plans to adapt to changes in assigned responsibilities and projects
Eagerness to participate in learning opportunities to enhance one’s effectiveness in executing day-to-day project activities
Ability to demonstrate and apply existing and new system processes, methodologies, and skills to contribute to the development of systems
Strong partnership skills with understand of escalation management across partner relationship
Knowledge of applications or infrastructure in a large-scale technology environment on premises or public cloud
Preferred qualifications, capabilities, and skills
Knowledge of one or more general purpose programming languages or automation scripting
Desire to grow and learn in the AI space as the business grows
Knowledge of Kubernetes, ITRS Active Console, Splunk, and Dynatrace
Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.
Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together.
Applicants
are
advised to research the bonafides of the prospective employer independently. We do NOT
endorse any
requests for money payments and strictly advice against sharing personal or bank related
information. We
also recommend you visit Security Advice for more information. If you suspect any fraud
or
malpractice,
email us at abuse@talentmate.com.
You have successfully saved for this job. Please check
saved
jobs
list
Applied
You have successfully applied for this job. Please check
applied
jobs list
Do you want to share the
link?
Please click any of the below options to share the job
details.
Report this job
Success
Successfully updated
Success
Successfully updated
Thank you
Reported Successfully.
Copied
This job link has been copied to clipboard!
Apply Job
Your application for Site Reliability Engineer II
has been successfully submitted!
To increase your chances of getting shortlisted, we recommend completing your profile.
Employers prioritize candidates with full profiles, and a completed profile could set you apart in the
selection process.
Why complete your profile?
Higher Visibility: Complete profiles are more likely to be viewed by employers.
Better Match: Showcase your skills and experience to improve your fit.
Stand Out: Highlight your full potential to make a stronger impression.
Complete your profile now to give your application the best chance!