Submitting more applications increases your chances of landing a job.
Here’s how busy the average job seeker was last month:
Opportunities viewed
Applications submitted
Keep exploring and applying to maximize your chances!
Looking for employers with a proven track record of hiring women?
Click here to explore opportunities now!You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for
Would You Be Likely to Participate?
If selected, we will contact you via email with further instructions and details about your participation.
You will receive a $7 payout for answering the survey.
Platform Reliability Engineer
About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millennium’s mission is to deliver results for our investors.
Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.
Meet the Team
Millennium’s Infrastructure organization sits within Information Technology, a function that is core to the health and growth of the firm. The team designs, engineers, supports, and manages a robust server estate, systems virtualization, and core enterprise services that underpin a flexible, scalable technology environment. Working across a globally distributed platform, the team partners closely to improve reliability, reduce operational complexity, and support critical computing environments with a practical, collaborative approach.
What You'll Do
Help ensure the production reliability of the firm’s Linux-based research and trading platform as part of a globally distributed engineering team
Respond quickly and effectively to production infrastructure incidents to minimize disruption and restore service
Build a strong understanding of internal client needs and communicate priorities clearly to regional and global leadership
Identify operational risks, develop contingency plans, and implement solutions to strengthen platform resilience
Develop and enhance observability capabilities to monitor the health and performance of critical computing environments
Participate in a monthly on-call rotation and provide support to on-call engineers as needed
Improve operational efficiency through configuration management, systems optimization, and automation
Contribute to team knowledge through clear documentation, knowledge sharing, and maintainable code
What You Bring
2+ years of experience in SRE, DevOps, or infrastructure engineering, ideally within the financial services industry
Strong knowledge of Linux systems internals, including kernel behavior, memory management, and performance optimization
Solid understanding of storage technologies, particularly in high-performance computing environments; GPFS experience is a plus
Broad infrastructure knowledge across networking and core services such as DNS, NTP/PTP, and NIS
Experience with automation, monitoring, and self-healing systems; Salt experience is a plus
Familiarity with container orchestration and virtualization technologies such as Kubernetes, Nomad, and VMware
Exposure to on-premises and cloud-based HPC infrastructure; operational knowledge of Slurm and GPU environments is a plus
Strong communication skills, a hands-on approach to problem-solving, and genuine curiosity about technology, automation, and AI-driven infrastructure improvements
You'll no longer be considered for this role and your application will be removed from the employer's inbox.