Description
Delivery & Service Assurance Lead
Customer Operations — SPEC-Ops (SRE & Platform Engineering)
Role Summary
The Delivery & Service Assurance Lead runs the day-to-day rhythm of the SPE Customer Operations team, assuring the team hits its delivery and SLA commitments. Delivery cadence and SLA performance are treated as one job: the cadence exists to move the SLA numbers. This role owns the team's 24/7 service picture and works standard US business hours; it doesn't carry on-call duty itself but runs the schedule and handoffs that keep coverage solid across regions.
Facilitation & Delivery Cadence
● Runs team ceremonies — standups, planning, reviews, and retrospectives — scheduled during US hours, with both the US and India shifts included and in sync.
● Keeps the backlog organized, making sure run work and operations are visible alongside project work, so the team's full workload is always in one place.
● Clears blockers and protects the team's focus, keeping a steady and predictable rhythm that the team can plan around.
● Owns the handoff between regions so work in progress transfers cleanly and coverage never drops, without needing to be online around the clock.
Service Assurance
● Tracks the team's own performance — SLA attainment across all shifts — and keeps metrics current and visible to the team and leadership.
● Helps the team balance run work against feature and project delivery.
● Brings SLA misses into retrospectives, gets to root cause, and helps make sure the same issue doesn't happen twice.
● Maintains a single 24-hour view of team performance — measured for the whole service, not by region — while working standard US hours.
Change Management
This role also manages maintenance of our processes, tools, and observability for applications. As products and features evolve, so does the operational requirements needed to support them, and this role keeps that work organized and moving. This includes managing backlog and tasks for:
● Coordinating the onboarding tasks needed to bring new products or features under proper observability coverage, including refinement.
● Procedural updates
● Security, ITSM, and other tooling updates
Continuous Improvement
● Keeps the SPEC-Ops way of working current as the team grows, adjusting cadence and process to fit both the SRE and Platform Engineering sides.
● Turns assurance data into backlog items — spotting the recurring toil and reliability gaps driving SLA risk, and prioritizing the work that fixes them.
● Builds and improves the metrics and reporting that show stakeholders how the team is doing.
Interface with Incident Command
When an SLA breach or incident happens, the Incident Commander (or escalation owner) handles the outward-facing work — stakeholder updates, escalation, vendor coordination. This role stays focused inward: making the breach visible, driving it to root cause, and turning it into backlog work to prevent it from happening again. That split keeps interrupt load off the person managing the team's delivery cadence, so this role can stay steady even during a live incident.
Qualifications
● Experience running agile delivery for an operations or engineering team — including facilitating the full set of ceremonies and owning backlog health.
● Solid grasp of SLA practice: setting targets, measuring attainment, and driving improvement.
● Comfortable with observability and ITSM tools and turning that data into delivery decisions.
● Proven track record coordinating a distributed, globally covered team across time zones while working standard US hours, including owning the handoff between regions.
● Experience turning operational data into a prioritized improvement backlog and following it through to completion.
Preferred
● Scrum Master or agile delivery certification, or equivalent hands-on experience.
● Experience standing up a delivery cadence for a newly formed or combined SRE/platform team.
● Background in a regulated or high-availability operations environment.
Similar jobs
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Est. 90,000 GBP
Who are we? Ensono is a global technology services provider dedicated to helping organizations navigate the complexity of digital transformation. Through Ensono Product, Consulting & Technology, our dedicated consult…
Delivery Manager (Scrum Master) Integration Product Business Unit (Pune) Location: Pune, IndiaExperience: 8–15 YearsReports To: PBU Head – Integration Product Business Unit Role Purpose The Delivery Manager is the operat…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
Job Title: Sr Manager, Engineering - DevOps About Reltio At Reltio®, an SAP Company, we believe data should fuel your success in the enterprise AI era. Our Context Intelligence Platform turns fragmented data into a trust…
Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. We protect 200+ organizations across North America, Europe, the Middle East, Southeast Asia, A…
About The Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcom…
Who We Are Insurity empowers insurance organizations to quickly capitalize on new opportunities by delivering the world’s most configurable, cloud-native, easy-to-use, and intuitively analytical insurance software. Just…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excel…
Est. 124,000 USD
Our Systems Engineering team owns the infrastructure that keeps a ~2,500-person global firm running around the clock: cloud (primarily AWS), Microsoft 365, Windows Server, identity platforms (Entra ID, Active Directory,…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excell…
About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…
Est. 125,000 USD
About Definitive Healthcare: At Definitive Healthcare (NASDAQ: DH), we’re passionate about turning data, analytics, and expertise into meaningful intelligence that helps our customers achieve success and shape the future…
At Optimove, we believe people are capable of more than a single job description. You’re not hired just to fill a position- you’re empowered to shape it, grow it, and make it your own.We call this being Positionless.And…
Est. 115,000 USD
Strategic Operations Lead (Project/Program Management - CloudOps) Position Summary This is a project/program management role embedded within CloudOps. The Strategic Operations Lead is primarily responsible for leading an…
Senior Product Manager Location: Hyderabad, India (Hybrid) About the Company At Redpin we simplify life's most important payments. Buying a new property overseas can be a stressful time, especially when it comes to movin…
Title: Senior Site Reliability Engineer - I, Product Area FocusLocation: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence o…
Black Duck Software, Inc. helps organizations build secure, high-quality software, minimizing risks while maximizing speed and productivity. Black Duck, a recognized pioneer in application security, provides SAST, SCA, a…
Est. 80,000 GBP
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and e…
Est. 140,000 USD
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About Kaseya Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service Providers (MSPs) and internal IT organizations worldwide. Our comprehensive platform helps organ…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Est. 120,000 GBP
Our roster has an opening with your name on it. The fastest Sportsbook in the industry doesn't happen by accident. We're looking for a Software Engineering Director to lead the teams behind the performance, reliability a…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
GovTech is the lead agency driving Singapore’s Smart Nation initiatives and public sector digital transformation. As the Centre of Excellence for Infocomm Technology and Smart Systems (ICT & SS), GovTech develops the…
Who We Are Insurity empowers insurance organizations to quickly capitalize on new opportunities by delivering the world’s most configurable, cloud-native, easy-to-use, and intuitively analytical insurance software. Just…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…