Job description
Our Sr. Staff TechOps & Support Engineer provides technical leadership for enterprise production operations by ensuring the stability, availability, and continuous improvement of mission-critical environments while driving automation, operational excellence, and service reliability across the organization.
Responsibilities Ensure high availability, performance, and reliability of production systems and platforms Drive incident response, root cause analysis, problem management, and continuous service improvement initiatives to enhance platform reliability and reduce operational risk.
Architect and optimize CI/CD pipelines, deployment strategies, automation frameworks, monitoring, and observability solutions to improve operational efficiency.
Provide technical leadership for IBM Cloud Pak Stacks, Confluent, Elasticsearch, container platforms, and enterprise infrastructure while supporting complex production deployments.
Establish operational standards, security controls, backup, disaster recovery, and service reliability practices to ensure compliance and business continuity.
Mentor engineers, provide technical direction, and collaborate with Development, Platform Engineering, Security, QUALITY, and Infrastructure teams to resolve critical production issues.
Maintain technical governance, operational documentation, and platform standards while driving innovation and continuous operational improvement.
Bachelor's degree or Diploma in Computer Science, Engineering, or a related field.
6–8 years of experience in Technical Operations, Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or a similar technical discipline.
Extensive experience with Linux/Windows administration, Kubernetes, OpenShift, Docker, CI/CD platforms, automation, scripting, monitoring, and observability tools (Grafana & Prometheus).
Strong expertise in IBM Cloud Pak Stacks, Confluent, Elasticsearch, Databases (SQL, DB2, MongoDB, etc.
), JVM performance analysis, networking, and enterprise production environments.
Advanced knowledge of security best practices, IAM, backup and disaster recovery, scalability, performance tuning, incident management, and SLA/SLO management.
Demonstrated ability to lead complex troubleshooting efforts, mentor engineering teams, and drive operational excellence across enterprise platforms.
Excellent leadership, stakeholder management, communication, decision-making, and strategic problem-solving skills.
This job post has been translated by AI and may contain minor differences or errors.