Ditto is seeking a Senior Site Reliability Engineer to enhance system reliability and operational excellence. You will develop observability solutions and lead incident management processes.
In this hands-on role, you'll manage production GPU clusters and contribute to system redesign for efficiency and scale. You'll also serve as the final escalation for complex GPU and networking failures.
The Senior Site Reliability Engineer is responsible for ensuring the reliability, scalability, performance, and operational excellence of a remote monitoring platform for cardiac devices.
As a Senior Site Reliability Engineer, you will support Counterpart Health’s technology infrastructure by improving processes, developing automation tools, and troubleshooting issues.
In this role, you will work with multidisciplinary teams to ensure production focus while creating necessary infrastructure. Your responsibilities include defining SLIs and SLOs, incident management, and creating observability systems.
In this role, you will work with multidisciplinary teams to ensure production focus while creating the necessary infrastructure. Your responsibilities include defining SLIs and SLOs, incident management, and creating observability systems.
As a Senior Site Reliability Engineer, you will support Counterpart Health’s technology infrastructure by developing automation tools, troubleshooting issues, and collaborating with cross-functional teams.
You will own the infrastructure behind our AI tooling and work hands-on with engineering teams to accelerate delivery and ensure production reliability. This role involves designing, building, and scaling Kubernetes infrastructure for secure, high-availability applications.
Join ClickHouse as a Senior Site Reliability Engineer to build and lead processes for cloud infrastructure reliability. Collaborate with various teams to design scalable and secure systems while managing incident response and performance improvements.
As a Senior Site Reliability Engineer, you will own pieces of the GCP infrastructure, container platform, CI/CD pipelines, and observability stack, ensuring issues are caught early and resolved quickly. This fully remote role involves collaborating with product and data engineering teams to maintain reliable platforms.
This role offers the opportunity to shape the reliability, scalability, and security of modern cloud infrastructure supporting mission-critical applications.
Join ClickHouse as a Senior Site Reliability Engineer to lead processes that ensure the reliability and performance of our cloud infrastructure. Collaborate with various teams to design and implement scalable systems while managing incident response and continuous improvement.