Job Description

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical Edge Reliability Engineering Team!

Site Reliability Engineers at Akamai leverage software engineering, systems expertise, and operational skills to deliver reliable global services. The ERE team ensures performance, resilience, and availability of Akamais media and web delivery platform while addressing distributed systems challenges. This role acts as the top technical escalation point for critical customer-impacting issues, connecting engineering, operations, and global support teams effectively.

Partner with the best

In this role, you will balance deep system diagnostics with progressive automation. You will serve as a premier technical authority for our platform and a champion for reducing systemic operational toil.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Leading complex reliability and performance investigations across Akamais global edge, media delivery, and web delivery platforms.
  • Troubleshooting critical distributed systems issues spanning application, platform, network, and operating system layers, serving as the highest technical escalation point.
  • Partnering with Engineering, Product, Support, and Network teams to identify root causes and deliver scalable, long-term solutions that improve platform reliability.
  • Designing and improving observability through SLIs, SLOs, KPIs, telemetry, dashboards, and alerts to identify and address customer-impacting issues.
  • Analyzing platform performance, traffic patterns, and system bottlenecks to improve scalability, resilience, and overall service reliability.
  • Developing automation, internal tools, AI-assisted diagnostics, and self-service workflows to streamline operations, reduce manual effort, and accelerate incident response.
  • Enhancing operational excellence through reliability-centered architecture reviews, post-incident analysis, continuous improvements, and offering off-hours support during critical incidents as needed.

Do what you love

To be successful in this role you will:

  • Possess Bachelors in CS/Engineering or a related field with 6 years of industry experience in large-scale SRE/Systems Infrastructure roles.
  • Have logical reasoning skills diagnosing complex performance bottlenecks, data integrity anomalies, and system failure modes in distributed environments.
  • Have understanding of internet technologies and foundational networking concepts, including caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures.
  • Have foundation in Linux/Unix administration, diagnostic tools, and low-level environment troubleshooting.
  • Be able to retrieve data, analyze telemetry streams, and troubleshoot platform data integrity issues through SQL queries.
  • Have experience developing automation tools using languages like Python, Bash, or Go.
  • Demonstrate expertise in AI models and focus on implementing agentic workflows to reduce operational inefficiencies effectively.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether youre streaming live events, scrolling social media, watching your favorite series, or managing your savings, were the engine behind the scenes. We provide the worlds most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.

Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the worlds biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the worlds most distributed cloud platform.

At Akamai, we dont just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And were the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your jobs needs

Akamais FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. Its not about telling employees where to work; its about supporting employees to do their best work.

We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!


Job Details

Role Level: Mid-Level Work Type: Full-Time
Country: India City: Bengaluru ,Karnataka
Company Website: https://www.akamai.com Job Function: DevOps & QA
Company Industry/
Sector:
Technology Information and Internet

What We Offer


About the Company

Searching, interviewing and hiring are all part of the professional life. The TALENTMATE Portal idea is to fill and help professionals doing one of them by bringing together the requisites under One Roof. Whether you're hunting for your Next Job Opportunity or Looking for Potential Employers, we're here to lend you a Helping Hand.

Report

Disclaimer: talentmate.com is only a platform to bring jobseekers & employers together. Applicants are advised to research the bonafides of the prospective employer independently. We do NOT endorse any requests for money payments and strictly advice against sharing personal or bank related information. We also recommend you visit Security Advice for more information. If you suspect any fraud or malpractice, email us at abuse@talentmate.com.


ad 1
Talentmate Instagram Talentmate Facebook Talentmate YouTube Talentmate LinkedIn