Submitting more applications increases your chances of landing a job.

Here’s how busy the average job seeker was last month:

Opportunities viewed

Applications submitted

Keep exploring and applying to maximize your chances!

Looking for employers with a proven track record of hiring women?

Click here to explore opportunities now!
We Value Your Feedback

You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for

Would You Be Likely to Participate?

If selected, we will contact you via email with further instructions and details about your participation.

You will receive a $7 payout for answering the survey.


User unblocked successfully
Thank you. Your report has been submitted and will be reviewed shortly.
https://bayt.page.link/yach6bGYxtYBE8Pz9
Back to the job results

Platform Reliability Engineer

18 hours ago 2026/11/25 ·Application closes in 119 days
Other Business Support Services
Create a job alert for similar positions
Job alert turned off. You won’t receive updates for this search anymore.

Job description

Platform Reliability Engineer

About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millennium’s mission is to deliver results for our investors.


Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.


Meet the Team
Millennium’s Infrastructure organization sits within Information Technology, a function that is core to the health and growth of the firm. The team designs, engineers, supports, and manages a robust server estate, systems virtualization, and core enterprise services that underpin a flexible, scalable technology environment. Working across a globally distributed platform, the team partners closely to improve reliability, reduce operational complexity, and support critical computing environments with a practical, collaborative approach.


What You'll Do


  • Help ensure the production reliability of the firm’s Linux-based research and trading platform as part of a globally distributed engineering team


  • Respond quickly and effectively to production infrastructure incidents to minimize disruption and restore service


  • Build a strong understanding of internal client needs and communicate priorities clearly to regional and global leadership


  • Identify operational risks, develop contingency plans, and implement solutions to strengthen platform resilience


  • Develop and enhance observability capabilities to monitor the health and performance of critical computing environments


  • Participate in a monthly on-call rotation and provide support to on-call engineers as needed


  • Improve operational efficiency through configuration management, systems optimization, and automation


  • Contribute to team knowledge through clear documentation, knowledge sharing, and maintainable code


What You Bring


  • 2+ years of experience in SRE, DevOps, or infrastructure engineering, ideally within the financial services industry


  • Strong knowledge of Linux systems internals, including kernel behavior, memory management, and performance optimization


  • Solid understanding of storage technologies, particularly in high-performance computing environments; GPFS experience is a plus


  • Broad infrastructure knowledge across networking and core services such as DNS, NTP/PTP, and NIS


  • Experience with automation, monitoring, and self-healing systems; Salt experience is a plus


  • Familiarity with container orchestration and virtualization technologies such as Kubernetes, Nomad, and VMware


  • Exposure to on-premises and cloud-based HPC infrastructure; operational knowledge of Slurm and GPU environments is a plus


  • Strong communication skills, a hands-on approach to problem-solving, and genuine curiosity about technology, automation, and AI-driven infrastructure improvements


This job post has been translated by AI and may contain minor differences or errors.
You’ve reached the maximum limit of 15 job alerts. To create a new alert, please delete an existing one first.
Job alert created for this search. You’ll receive updates when new jobs match.
Are you sure you want to unapply?

You'll no longer be considered for this role and your application will be removed from the employer's inbox.