Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic
Full-time•DevOps•Ireland
Apply Now

TechJobs.ie Job Insights

At a glance

Employment type
Full-time
Location
Ireland
Workplace
On-site
Category
DevOps

Technologies & skills

Primary technologies

Cloud & infrastructure

Other technical skills

What you'll be doing

Senior Site Reliability Engineer at Elastic, focused on designing, building, and scaling a multi-cloud platform for hosting internal and external services. The role emphasizes automation, operational excellence, and collaboration to ensure platform reliability and exceptional customer experience.

  • Lead technical initiatives for automating system engineering to guarantee infrastructure reliability
  • Develop and maintain software, tooling, and automations to scale platform infrastructure
  • Champion collaboration, operational excellence, and team development
  • Respond to and prevent repeated customer impact from major incidents
  • Participate in on-call rotation using a follow-the-sun model

Key requirements

Must-have

  • Experience in platform reliability and site reliability engineering
  • Background in software engineering
  • Experience with public cloud and managed Kubernetes services is advantageous
  • Ability to collaborate in distributed or remote teams
  • Customer-first approach to solving operational problems

Nice-to-have

  • Experience operating SaaS products in public cloud using Infrastructure-as-Code tools (e.g., Terraform)
  • Experience with Kubernetes-at-scale across multiple cloud providers
  • Programming experience in Go or other languages
  • Experience with containerized services (e.g., Docker)
  • Experience with alerting and incident management systems (e.g., Prometheus, Elastic Stack)
  • Professional Linux system administration skills on distributed systems

Role signals

Technical focus
site reliability engineering, platform automation, infrastructure scaling
Leadership
Mentoring
Architecture / system design
Indicated in the listing
Hands-on vs management
Hands-on

Full job description

Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter. By taking advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.

What is The Role:

As part of the Platform Engineering department, the SRE team is designing, building, scaling and maturing the multi-cloud platform for hosting internal and external services such as the Elastic Cloud Hosted and Serverless. We develop and extend new software and tools that support the rest of the infrastructure, so that we can rapidly deploy products from all corners of Elastic. We want your experience and recommendations to offer a truly exceptional customer experience!

What you will be doing:

  • Taking an engineering approach in leading technical initiatives for automating system engineering efforts to guarantee the reliability of the global Elastic infrastructure. .

  • Growing our global Platform infrastructure to meet the increasing scaling demands by developing and maintaining software, tooling and automations.

  • Using an inclusive approach at championing an environment focused on collaboration, operational excellence, and uplifting others.

  • Responding to and preventing repeated customer impact in response to major incidents and prioritised problem management. Our on call rotation uses follow-the-sun model where everyone participates in it in (mostly) their working hours.

What you bring:

  • Success and lessons of experiences from striving for 'progress not perfection' in the name of Platform reliability. We want to hear about your customer first approach in solving operational problems with a SRE perspective.

  • A background in software engineering to collaborate with engineers to expertly identify, implement and deliver solutions. An experience in public cloud and managed Kubernetes services is advantageous.

  • Passion for developing solutions that involve inclusive communication methods to grow and strengthen partner and team relationships. Examples of working in distributed teams or working remotely is desirable.

Bonus Points:

You don't need to have all of these items, but these represent the types of work you will do as a Site Reliability Engineer at Elastic.

  • You have operated a SaaS product in a public cloud ideally built using Infrastructure-as-Code tooling such as Crossplane or Terraform

  • You have built or operated a Kubernetes-at-scale infrastructure, ideally across multiple cloud providers, and the vital automation to support it.

  • You have written non-trivial programs in Golang or other programming languages.

  • You have worked with containerized services (such as Docker.)

  • You have proven experience in leading and improving alerting and major incident management standard processes metrics systems (e.g. Elastic Stack, Graphite, Prometheus, Influx) to diagnose issues and quantify impacts to present to others at varying level of the organization.

  • You have experience in system administration with professional skills in Linux on distributed systems at scale.

  • You have diagnosed or designed, implemented and created solutions with the Elastic Stack.

  • You are experienced in thriving in a self-organizing and sharing in a globally distributed team environment.

  • You strengthen team members in bringing out the best of each other by uplifting others with coaching and mentoring.

Compensation for this role is in the form of base salary. This role does not have a variable compensation component.

At Elastic, our compensation philosophy aims to provide fair, competitive and transparent remuneration. Salary ranges are established based on a combination of external market benchmarks, internal pay equity considerations, and the responsibilities and complexity associated with each role. This approach helps ensure consistency across comparable roles while remaining competitive within the relevant labour markets.

The final compensation offered within the applicable range will be determined based on several objective factors, including relevant professional experience, level of skills and expertise, alignment with the role requirements, and the overall scope and complexity of the position.

The typical starting salary range for this role is:

€98.400—€126.900 EUR

Additional Information - We Take Care of Our People

As a distributed company, diversity drives our identity. Whether you’re looking to launch a new career or grow an existing one, Elastic is the type of company where you can balance great work with great life. Your age is only a number. It doesn’t matter if you’re just out of college or your children are; we need you for what you can do.

We strive to have parity of benefits across regions and while regulations differ from place to place, we believe taking care of our people is the right thing to do.

  • Competitive pay based on the work you do here and not your previous salary

  • Health coverage for you and your family in many locations

  • Ability to craft your calendar with flexible locations and schedules for many roles

  • Generous number of vacation days each year

  • Increase your impact - We match up to $2000 (or local currency equivalent) for financial donations and service

  • Up to 40 hours each year to use toward volunteer projects you love

  • Embracing parenthood with minimum of 16 weeks of parental leave

Different people approach problems differently. We need that. Elastic is an equal opportunity employer and is committed to creating an inclusive culture that celebrates different perspectives, experiences, and backgrounds. Qualified applicants will receive consideration for employment without regard to race, ethnicity, color, religion, sex, pregnancy, sexual orientation, gender perception or identity, national origin, age, marital status, protected veteran status, disability status, or any other basis protected by federal, state or local law, ordinance or regulation.

We welcome individuals with disabilities and strive to create an accessible and inclusive experience for all individuals. To request an accommodation during the application or the recruiting process, please email candidate_accessibility@elastic.co. We will reply to your request within 24 business hours of submission.

Applicants have rights under Federal Employment Laws, view posters linked below: Family and Medical Leave Act (FMLA) Poster; Pay Transparency Nondiscrimination Provision Poster; Employee Polygraph Protection Act (EPPA) Poster and Know Your Rights (Poster)

Elasticsearch develops and distributes technology and information that is subject to U.S. and other countries’ export controls and licensing requirements for individuals who are located in or are nationals of the following sanctioned countries and regions: Belarus, Cuba, Iran, North Korea, Syria, or Russia, including the Ukrainian territories annexed by Russia (The Crimea region of Ukraine, The Donetsk People's Republic (DNR), The Luhansk People's Republic (LNR), Kherson or Zaporizhzhia). If you are located in or are a national of one of the listed countries or regions, an export license may be required as a condition of your employment in this role. Please note that national origin and/or nationality do not affect eligibility for employment with Elastic.

Please see here for our Privacy Statement.

Interview prep pack

Grounded in this listing. Use it to prepare examples before you apply.

Your interview focus

Based on this listing, the Senior Site Reliability Engineer - Platform Reliability (Resilience) role at Elastic focuses on leading technical initiatives to automate and scale multi-cloud platform infrastructure, with a strong emphasis on reliability, collaboration, and operational excellence.

  • Hands-on SRE and automation experience·High
  • Multi-cloud and Kubernetes expertise·High
  • Incident management and prevention·High
  • Collaboration and mentoring in distributed teams·Medium

Only have 30 minutes?

Follow a focused preparation plan based on this job.

Start 30-minute prep

Your 30-minute plan

  1. Review Elastic's Platform and Cloud Offerings

    0–7 min

    Spend time understanding Elastic's hosted and serverless products, focusing on their reliability and scaling needs.

  2. Refresh Kubernetes and Multi-Cloud Skills

    7–14 min

    Go over your experience with Kubernetes and multi-cloud environments, preparing specific examples of scaling and automation.

  3. Prepare Incident Management Stories

    14–20 min

    Select and structure 1-2 strong examples of incident response and prevention from your past roles.

  4. Review Infrastructure-as-Code Projects

    20–25 min

    Gather details on your use of Terraform or similar tools, focusing on automation and reliability improvements.

  5. Plan Questions and Team Fit Examples

    25–30 min

    Draft thoughtful questions for the interviewers and prepare stories that demonstrate collaboration and mentoring.

Likely questions

, 6 items

Priority reflects how strongly this topic is emphasised in the job listing, not whether it will be asked.

Talking points

, 6 items
  • Automating System Engineering for Reliability

    Be ready to discuss examples where you led or contributed to automation efforts that improved platform reliability, as this is a core responsibility.

  • Scaling Multi-Cloud and Kubernetes Infrastructure

    Prepare to share your experience with scaling infrastructure, especially using Kubernetes and public cloud services, as this is highlighted as a key part of the role.

  • Incident Management and Prevention

    Demonstrate your approach to responding to and preventing major incidents, including how you analyze root causes and implement solutions to avoid recurrence.

  • Collaboration in Distributed Teams

    Showcase your ability to work effectively in globally distributed or remote teams, emphasizing communication and inclusive practices.

  • Developing and Maintaining Tooling and Automation

    Discuss your experience building and maintaining software tools (especially with Go, Terraform, or similar), as this is central to the role.

  • Mentoring and Uplifting Team Members

    Be prepared to give examples of how you have coached or mentored others, as the role values uplifting and strengthening the team.

What to research

, 4 items
  • Elastic's Cloud and Serverless Offerings

    Review Elastic's hosted and serverless products to understand the platform's architecture and reliability requirements.

  • Kubernetes and Multi-Cloud Operations

    Refresh your knowledge of operating Kubernetes at scale, especially across multiple cloud providers.

  • Incident Management Best Practices

    Study modern incident response frameworks and how to implement preventive measures in SRE environments.

  • Infrastructure-as-Code with Terraform

    Prepare examples of using Terraform or similar tools for automating infrastructure deployment and management.

Questions to ask

, 6 items
  1. What are the biggest current challenges facing the platform reliability team at Elastic?

    Why ask this? To understand the team's priorities and where your skills can have the most impact.

  2. How does the team approach incident response and post-incident reviews?

    Why ask this? To learn about the incident management culture and expectations.

  3. What tools and technologies are most critical to your current automation and scaling efforts?

    Why ask this? To clarify which technologies you should be most familiar with.

  4. How is collaboration fostered across distributed teams at Elastic?

    Why ask this? To assess the company's approach to remote teamwork and communication.

  5. What opportunities exist for mentoring or coaching within the team?

    Why ask this? To show your interest in team development and leadership.

  6. How does the on-call rotation work in practice, and what support is available during major incidents?

    Why ask this? To understand expectations and support structures for incident response.

Register now to upload your CV

Create a free account, save a PDF or Word CV, and quick apply on roles that take applications here.

Apply NowApply before: 25 Oct 2026