Submitting more applications increases your chances of landing a job.

Here’s how busy the average job seeker was last month:

Opportunities viewed

Applications submitted

Keep exploring and applying to maximize your chances!

Looking for employers with a proven track record of hiring women?

Click here to explore opportunities now!
We Value Your Feedback

You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for

Would You Be Likely to Participate?

If selected, we will contact you via email with further instructions and details about your participation.

You will receive a $7 payout for answering the survey.


User unblocked successfully
Thank you. Your report has been submitted and will be reviewed shortly.
https://bayt.page.link/vrnqD7qWkFZY5iza7
Back to the job results

Lead Site Reliability Engineer

6 hours ago 2026/11/25 ·Application closes in 119 days
Other Business Support Services
Create a job alert for similar positions
Job alert turned off. You won’t receive updates for this search anymore.

Job description

Job Purpose

To lead the reliability, availability, performance, and continuous improvement of Air Arabia's mission-critical Java-based applications and airline reservation systems by implementing Site Reliability Engineering (SRE) best practices. Responsible for driving operational excellence through proactive monitoring, automation, incident management, scalability improvements, and collaboration with cross-functional teams, while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements.



Key Result Responsibilities
  • Lead Site Reliability Engineering (SRE) initiatives for Java-based microservices and monolithic applications supporting mission-critical airline operations and Passenger Service Systems (PSS).
  • Establish, monitor, and continuously improve system reliability by defining and managing Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets.
  • Lead and facilitate root cause analysis (RCA) for complex production incidents, ensuring timely resolution and implementation of preventive measures to minimize recurrence.
  • Drive architectural enhancements to improve system performance, scalability, resilience, availability, and operational efficiency through optimization techniques such as caching and distributed system design.
  • Mentor and provide technical guidance to team members by promoting best practices in incident management, production support, troubleshooting, and operational excellence.
  • Collaborate closely with software engineering, infrastructure, and cross-functional teams to enhance application design, improve system reliability, and ensure production readiness.
  • Own and define CI/CD pipeline standards, GitOps practices, and infrastructure automation strategy, setting the framework that engineers across the team execute against.
  • Set the containerization and orchestration strategy (Docker, Kubernetes) for the team, defining standards for scalability, resilience, and high availability that other engineers implement.
  • Own the on-call escalation framework, ensuring adequate coverage and clear escalation paths, and lead post-incident reviews.
  • Evaluate new tools and technologies to strengthen reliability and represent SRE in capacity planning and release readiness reviews.
  • Own reliability reporting to management, translating SLA/SLO performance and incident trends into clear business updates.

Qualifications (Academic, training, languages)
  • Bachelor’s Degree in Computer Engineering/ Computer Science/ Information Technology. 
  • Fluent in English Language.
  • Strong expertise in Java, Spring Boot, and troubleshooting complex production issues within enterprise environments.
  • Airline, aviation, or travel industry experience, particularly with Passenger Service Systems (PSS), is preferred.
  • Proficiency in MS Office.
  • Strong knowledge of SQL with experience in database design, optimization, and performance tuning; experience with Oracle Database is an advantage.
  • Solid understanding of distributed systems, system architecture, scalability, high availability, and resilient application design.
  • Proficiency in monitoring, logging, and observability tools such as Prometheus, Grafana, Elasticsearch, or Datadog, used to drive proactive incident detection and reliability strategy.

Work Experience
  • With 6-8 years of experience in Software Engineering, SRE, or Production Support (Java Applications).
  • Hands-on experience designing, developing, and supporting both microservices and monolithic application architectures.
  • Strong hands-on experience with Docker and Kubernetes for containerization, orchestration, and production deployment (mandatory).
  • Experience with JBoss Application Server or similar enterprise Java application servers is an added advantage.
  • Proven experience designing and governing CI/CD pipeline standards, GitOps practices, and Git-based workflows across multiple teams.
  • Hands-on experience with caching technologies (e.g., Redis) and messaging platforms to improve application performance and reliability.


This job post has been translated by AI and may contain minor differences or errors.
You’ve reached the maximum limit of 15 job alerts. To create a new alert, please delete an existing one first.
Job alert created for this search. You’ll receive updates when new jobs match.
Are you sure you want to unapply?

You'll no longer be considered for this role and your application will be removed from the employer's inbox.