Description
Senior Splunk Engineer for Automation and Reliability Engineering Project
Project Summary
- Support Automation and Reliability Engineering project and operations.
- Responsibilities:
- Observability Engineering and Governance
- Architect and maintain enterprise SIEM solutions aligned with operational resilience mandates (e.g., MAS TRM, DORA, APRA CPS 230).
- Lead deployment, configuration, and optimization of Splunk for full-stack visibility across infrastructure, applications, networks, and user experience.
- Define and enforce telemetry data governance standards—metrics, logs, and traces—ensuring consistency, retention compliance, and security.
- Integrate Splunk with incident management, ITSM, and AIOps systems to enable predictive alerting and anomaly detection.
- Act as the SIEM/Splunk subject matter expert (SME) for architecture reviews, platform upgrades, and performance tuning.
- Reliability Engineering and Automation
- Implement and champion SRE frameworks and reliability practices for mission-critical systems.
- Design and automate runbooks, alerts, and self-healing workflows using Python, Ansible, and Terraform.
- Collaborate with Application, Infrastructure, and Cyber teams to embed reliability principles into the delivery lifecycle.
- Conduct resilience, chaos, and capacity testing aligned with business continuity and disaster recovery standards.
- Define and track error budgets, reliability scorecards, and service health indicators for production workloads.
- Cloud & Platform Integration
- Engineer SIEM for cloud-native workloads in AWS and Azure, ensuring visibility across compute, storage, and network layers.
- Integrate Splunk and cloud observability tools into CI/CD pipelines and landing zones to ensure continuous compliance.
- Implement infrastructure-as-code (IaC) models using Terraform and Ansible for consistent, auditable provisioning.
- Collaborate with Cloud, DevOps, and Security teams to ensure telemetry aligns with audit, compliance, and operational risk requirements.
- Operational Excellence and Collaboration
- Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
- Partner with Service Delivery, Cyber, and Application teams to enable predictive incident prevention and root cause transparency.
- Develop and maintain executive dashboards and reports showcasing availability, reliability KPIs, and operational risk indicators.
- Provide technical leadership during major incidents, post-incident reviews, and audits, ensuring lessons learned are codified into automation and process improvements.
Skillset (Must have)
- Possess a degree in Computer Science, Engineering, or related discipline.
- Minimum 8 years of experience in Infrastructure, Cloud, or Site Reliability Engineering related roles, with at least 5 years of experience specializing in SIEM/Splunk engineering or observability in financial or regulated environments.
- Proven hands-on expertise in the following technical areas:
- SIEM Platforms: Splunk (must), EL/Elastic
- Automation/IaC, Terraform, Ansible, Python, CI/CD tools
- Cloud and other platforms and integrations: AWS (CloudWatch, X-Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights), Datadog, ServiceNow
- Deep understanding of SRE principles, service health modelling, error budgets, and auto-remediation design.
- Strong analytical and troubleshooting skills, with the ability to perform deep-dive investigations and develop long-term preventive solutions.
- Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
- Excellent written and verbal communication skills.
- Strong interpersonal and communication skills to interact with diverse stakeholders.
- Agile, fast learner and able to adapt to changes
Skillset (Good to have)
Preferred Certifications:
- Splunk Certified Power User / Splunk Certified Admin / Splunk Certified Architect
- Terraform / Ansible / Python Certified Expert
- AWS Certified DevOps Engineer / Azure DevOps Expert
- SRE Foundation / Practitioner (DevOps Institute)
- ITIL v4 Managing Professional
Any personal data you share with us during the application process will be processed strictly in compliance with applicable data protection laws and our Privacy Notice.
Similar jobs
The Platform Operations Engineer is responsible for maintaining and supporting on-premises infrastructure platforms that underpin mission-critical systems. This role ensures platform reliability, operational stability, a…
We are looking for a Software Engineer (Mid-Level / Senior) to support a large-scale Atlassian Cloud migration project. This role will focus on engineering solutions to support the transition from on-premise environments…
We are seeking a skilled and passionate Engineer to join our team to build and operate a Whole-of-Government (WoG) runtime platform. As a Site Reliability Engineer, you will be responsible for designing and operating Git…
About the Role We are looking for a technically driven and customer-focused Service Engineer (Cloud & Infrastructure) to join our growing Singapore team. This role is ideal for individuals who enjoy solving complex i…
ROLE OVERVIEW We are looking for a seasoned Data Architect with deep expertise in enterprise data warehousing, ETL/ELT pipeline development, and business intelligence. You will be responsible for the end-to-end design an…
Software Engineer (SRE) What you will be working on Solution Engineering Maintain comprehensive system architecture with deep understanding of integration patterns and dependencies across the technology stack Design and…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excel…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excell…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Est. 48,000 EUR
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Title: Senior Site Reliability Engineer - I, Product Area FocusLocation: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence o…
Job Title: Senior / IT Infra Engineer (Collaboration) Role Overview We are seeking highly skilled and innovative engineer to join our transformative SSOE Programme that manages more than 350 schools across Singapore. Thi…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Job Title: Senior / IT Infra Engineer (Identity and Security) Role Overview We are seeking highly skilled and innovative engineer to join our transformative SSOE Programme that manages more than 350 schools across Singap…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
DevSecOps Engineer (Contractor role, GovTech project) The Government Digital Product (GDP) Team aims to spearhead the digital transformation of government. GDP was established to develop new capabilities focusing on stra…
Est. 54,000 EUR
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Senior Site Reliability Engineer I Location San Jose, Costa Rica - Remote Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s plane…
Est. 148,500 USD
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Senior Manager, Engineering - Application Security Want to lead a global team responsible for the most important product features – availability, reliability & security? Sumo’s SRE program focuses on continual data-d…
About the Role We are seeking a Senior Software Quality Engineer to lead our quality assurance efforts and ensure the delivery of high-quality software solutions. In this role, you will work closely with Product Owners a…
Securing the Future with AvePoint AvePoint is a global leader in data management and governance, trusted by over 21,000 customers worldwide to enhance their digital workplaces across Microsoft, Google, Salesforce, and ot…
Beyond Secure. AvePoint is the global leader in data security, governance, and resilience, going beyond traditional solutions to ensure a robust data foundation and enable organizations everywhere to collaborate with con…
Senior Manager, Engineering - Application Security Want to lead a global team responsible for the most important product features – availability, reliability & security? Sumo’s SRE program focuses on continual data-d…
Technical Requirements 5+ years of software engineering experience with demonstrated progression in technical complexity and scope Strong foundation in software architecture, system design, and engineering best practices…
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Est. 185,500 USD
Our Purpose At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them become…
Securing the Future with AvePointAvePoint is a global leader in data management and governance, trusted by over 21,000 customers worldwide to enhance their digital workplaces across Microsoft, Google, Salesforce, and oth…