You will be responsible for driving operational excellence across complex production environments while acting as a key escalation point for critical incidents and challenging technical issues. This role combines hands-on technical operations with technical leadership, helping shape operational standards and reliability practices.
Linux
Kubernetes
NVIDIA GPUs
+5 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities. The engineer will also provide technical leadership for Kubernetes platform operations and support infrastructure services.
As a (Senior) IoT Cloud Engineer, you will own the cloud side of our IoT Platform that ingests, routes, and observes every message flowing between our gateway fleet and our cloud services.
Join GitLab's Observability, Monitoring, and Integrations team to enhance the Monetization section by developing tools for anomaly detection and telemetry. You'll work on a greenfield project, shaping its operations and using AI and machine learning to improve reliability.
You will be responsible for developing a high-scale, AI-ready Data Lakehouse and prototyping emerging architectural patterns. This role involves building tools and SDKs for autonomous agents and ensuring interoperability between platforms.
You will own the infrastructure behind our AI tooling and work hands-on with engineering teams to accelerate delivery and ensure production reliability. This role involves designing, building, and scaling Kubernetes infrastructure for secure, high-availability applications.
As a Senior SRE at DualEntry, you will be responsible for building cache layers, improving latency and response time, and ensuring systems scale effectively.
$140k - $240k
Terraform
AWS
Python
+1 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities. The engineer will engage with pioneering AI hardware and support large-scale NVIDIA GPU infrastructure.
As a Senior DevOps Engineer, you will develop a platform that enables developers and agents to validate changes quickly and ship confidently. You will work with a small Infrastructure team to automate processes and enhance self-service capabilities.
AWS
EKS
Terraform
+13 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms while providing technical leadership for Kubernetes operations. The engineer will also drive improvements in platform reliability and operational processes.
Linux
Kubernetes
NVIDIA GPU
+5 more
Senior AI Infrastructure & Platform Operations Engineer
You will be responsible for driving operational excellence across complex production environments while acting as a key escalation point for critical incidents and challenging technical issues. This role combines hands-on technical operations with technical leadership, helping shape operational standards and reliability practices.
This role involves monitoring and supporting production AI infrastructure platforms while collaborating with engineering teams. The position offers the chance to engage with pioneering AI hardware and drive the development of automated operational capabilities.
As an AI Infrastructure Engineer III, you will focus on building scalable AI infrastructure and optimizing GPU platforms for AI workloads. This role involves collaborating with data science teams to enhance platform usability and performance.
As an AI Infrastructure & Platform Operations Engineer, you will maintain the reliability and efficiency of AI service platforms across a global datacenter footprint. This role involves engaging with advanced AI hardware and driving the development of automated operational capabilities.
NVIDIA GPU
Kubernetes
InfiniBand
+4 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms while providing technical leadership for Kubernetes operations. The engineer will also drive improvements in platform reliability and operational processes.
Linux
Kubernetes
NVIDIA GPU
+5 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms across a global datacenter footprint while driving the development of automated operational capabilities. The engineer will also provide technical leadership for Kubernetes platform operations and support infrastructure services.
Linux
Kubernetes
NVIDIA GPU
+5 more
Senior AI Infrastructure & Platform Operations Engineer
This role involves maintaining the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities. The engineer will also provide technical leadership for Kubernetes platform operations and support infrastructure services.
As a PFS Admin & Tools - Monitoring, you will take end-to-end responsibility for monitoring platforms, focusing on both traditional infrastructure and cloud-native environments. The position requires strong experience with monitoring tools and collaboration with various teams.