Site Reliability Engineer — Data Platforms, IS&T Ai & Data Platforms

Apple

Summary

Posted: Aug 10, 2026

Weekly Hours: 40

Role Number:200676274

The AiDP Data Platforms team builds and operates data platforms at scale on the Cloud, helping Apple process, store, and access petabytes of data. We’re seeking an SRE to own the reliability, performance, and operability of our data and ML platforms — someone who thinks in SLOs, failure modes, and blast radius, and is passionate about keeping large-scale distributed systems fast, available, and cost-efficient.

Description

You’ll operate and harden our big data platform — built on open source and other technologies — that powers critical applications like analytics, reporting, and AI/ML. This means driving down MTTR, automating operational toil, tuning performance and cost, running capacity planning, and root-causing production incidents before and after they happen. You’re an independent, self-directed problem-solver who communicates clearly with both engineers and non-technical partners, and you’ll work across many teams to keep the platform running at Apple’s standard.

Minimum Qualifications

  • 3+ years of experience operating and supporting critical, large-scale distributed systems in production, with scripting/programming ability in Python, Go, Java, Scala, or Bash for automation and tooling.
  • Deep understanding of reliability principles — fault tolerance, high availability, low latency, graceful degradation — and how to instrument and enforce them (SLIs/SLOs, alerting, on-call practices, incident response, postmortems).
  • Hands-on experience operating data processing ecosystems and distributed computing frameworks (Spark, Flink) and MPP query engines (Trino, StarRocks), including performance tuning and capacity management.
  • Proficiency operating Kubernetes/Helm at scale, building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), managing infrastructure as code (Terraform, Pulumi), and running service-oriented architectures across multi-cloud environments.
  • Strong troubleshooting and performance analysis skills in complex production environments; fluency in Unix/Linux and command-line diagnostics, with excellent problem-solving and communication skills.

Preferred Qualifications

  • Experience contributing to open source projects or operating across multiple public cloud providers.
  • In-depth operational knowledge of specific distributed frameworks — Spark, Flink, or Kafka Streams, Trino, Iceberg — including debugging via component-specific logs, and experience managing multi-tenant Kubernetes clusters at scale.
  • Experience with workflow/pipeline orchestration tools (Airflow, dbt) and understanding of data modeling and warehousing concepts.
  • Experience debugging Kubernetes/Spark production issues via logs and metrics, with a continuous-improvement mindset for self, team, and org.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.

Apple is an equal opportunity employer that is committed to inclusion and diversity, and thus we treat all applicants fairly and equally. Apple is committed to working with and providing reasonable accommodation to applicants with physical and mental disabilities.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Read Full Description
Confirmed 6 hours ago. Posted 30+ days ago.

Discover Similar Jobs

Suggested Articles