Confluent - Senior Software Engineer - Flink Autopilot

IBM

Benefits
Special Commitments
Skills

Introduction

At IBM Software, we transform client challenges into solutions. Building the world's leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You'll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM's product and technology landscape. Here, you'll have the tools and opportunities to advance your career while creating software that changes the world.

Your role and responsibilities

The Stream Processing & Analytics (SPA) team is building an elastic, reliable, durable, cost-effective, and performant stream processing engine based on Apache Flink for Confluent Cloud. Within SPA, the Flink Autopilot team owns the systems that make Flink a true cloud-native, "zero-knob" experience — autoscaling, resource management, and worry-free operations that deliver the right amount of stream processing at any given moment, so customers can focus on their use case instead of managing infrastructure.

As a senior engineer on Autopilot, you'll independently drive complex engineering projects end to end within our domain, become a go-to expert for one or more areas of the system, and help the team raise the quality and operational health of the platform.

Read This First: The Kind of Engineering We Do

This is a low-level systems engineering role. The hard part of Autopilot is not integrating Flink, Kafka, or Kubernetes, but it's writing the code underneath them that has to be correct under concurrency, failure, and partial state.

You will build control loops that make autoscaling and resource decisions, and that code must remain consistent and highly available on its own (through crashes, restarts, leader changes, and racing events), without a higher-level framework quietly handling correctness for you. When we say "scalable, fault-tolerant, distributed systems," we mean the mechanics of consistency and high availability implemented in your own code, not the ability to assemble existing platforms that provide those properties.

If your strength is designing high-level services that delegate reliability to Kafka/Flink/Kubernetes, this role is probably not the right fit, and we'd rather both sides know that now than discover it mid-process.

What we mean vs. what we don't mean

  • Consistency: We mean reasoning about and hand-writing the logic that keeps state correct across concurrent updates, retries, and failures (idempotency, ordering, reconciliation, avoiding races). We do not mean "point at a datastore that promises consistency and move on"
  • High availability: We mean writing code that survives process crashes, leader changes, and restarts; recovering and reconciling its own state. We do not mean "deploy multiple replicas and let the platform handle failover"
  • Fault tolerance: We mean anticipating partial failure in your own control loops and coordination logic and handling it explicitly. We do not mean relying on retries and health checks configured at the infrastructure layer.
  • Distributed systems: We mean the coordination, state, and correctness problems inside the system you're building. We do not mean operating distributed applications as a user of them

Deep, effective use of Kafka, Flink, and Kubernetes is genuinely valuable here, but it is a complement to the above, not a substitute for it.

What You Will Do

  • Drive complex projects end to end: Independently take projects from design through production within the Autopilot domain (for example: autoscaling logic, resource management, or job/task manager resilience). Break large efforts into clear milestones and deliver them with high quality
  • Build correctness into the code: Design and implement control loops and coordination logic that guarantee consistency and high availability directly (handling concurrency, restarts, leader changes, and partial failure explicitly, rather than delegating those guarantees to an underlying platform)
  • Master a domain: Develop deep expertise in one or more areas of Autopilot and become the person the team relies on for that area, representing your components in cross-functional design and architecture discussions
  • Articulate your design thinking: Explain the failure modes, trade-offs, and correctness arguments behind your systems through clear design documents, one-pagers, and proposals
  • Raise the bar through reviews and mentorship: Review PRs and designs with constructive feedback, coach junior engineers and interns, and contribute to the team's interview question pool and hiring
  • Communicate clearly: Demonstrate strong, succinct written and verbal communication to align a small group on technical direction and keep stakeholders informed.

Required education

Bachelor's Degree

Preferred education

Master's Degree

Required technical and professional expertise

  • Deep expertise involving 6–8+ years of relevant experience in stream processing or large-scale distributed systems, and some familiarity with autoscaling, resource management, or scheduling in distributed data systems.
  • Strong fundamentals in distributed systems or stream processing, with a track record of independently designing, building, and shipping complex systems end to end
  • Demonstrated ability to implement consistency and high availability directly in code — reasoning about concurrency, failure and recovery, state, and coordination, and getting correctness right without relying on a higher-level framework to provide it
  • Proficiency in Java and/or Scala (the core languages of the Flink engine), with the ability to contribute production code from day one
  • Hands-on experience building and operating mission-critical systems in a public cloud environment (AWS, GCP, or Azure)
  • Ability to own a technical area, make sound trade-offs, and align a small group on direction

Preferred technical and professional experience

  • Experience implementing low-level reliability mechanisms yourself (for example, leader election, distributed coordination, reconciliation loops, consensus, or state replication) rather than only consuming them as a service
  • Experience with Kubernetes and operating distributed applications in production
  • Open source engagement: recognized, impactful technical contributions to open-source stream processing projects, particularly Apache Flink
  • Experience working with Go in a cloud control-plane context

ABOUT BUSINESS UNIT

IBM Software infuses core business operations with intelligence—from machine learning to generative AI—to help make organizations more responsive, productive, and resilient. IBM Software helps clients put AI into action now to create real value with trust, speed, and confidence across digital labor, IT automation, application modernization, security, and sustainability. Critical to this is the ability to make use of all data, because AI is only as good as the data that fuels it. In most organizations data is spread across multiple clouds, on premises, in private datacenters, and at the edge. IBM’s AI and data platform scales and accelerates the impact of AI with trusted data, and provides leading capabilities to train, tune and deploy AI across business. IBM’s hybrid cloud platform is one of the most comprehensive and consistent approach to development, security, and operations across hybrid environments—a flexible foundation for leveraging data, wherever it resides, to extend AI deep into a business.

YOUR LIFE @ IBM

In a world where technology never stands still, we understand that, dedication to our clients success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better.

Being an IBMer means you’ll be able to learn and develop yourself and your career, you’ll be encouraged to be courageous and experiment everyday, all whilst having continuous trust and support in an environment where everyone can thrive whatever their personal or professional background.

Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide on-going feedback to help other IBMers grow, as well as collaborate with colleagues keeping in mind a team focused approach to include different perspectives to drive exceptional outcomes for our customers. The courage our IBMers have to make critical decisions everyday is essential to IBM becoming the catalyst for progress, always embracing challenges with resources they have to hand, a can-do attitude and always striving for an outcome focused approach within everything that they do.

Are you ready to be an IBMer?

ABOUT IBM

IBM’s greatest invention is the IBMer. We believe that through the application of intelligence, reason and science, we can improve business, society and the human condition, bringing the power of an open hybrid cloud and AI strategy to life for our clients and partners around the world.

Restlessly reinventing since 1911, we are not only one of the largest corporate organizations in the world, we’re also one of the biggest technology and consulting employers, with many of the Fortune 500 companies relying on the IBM Cloud to run their business.

At IBM, we pride ourselves on being an early adopter of artificial intelligence, quantum computing and blockchain. Now it’s time for you to join us on our journey to being a responsible technology innovator and a force for good in the world.

IBM is proud to be an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status.

OTHER RELEVANT JOB DETAILS

For additional information about location requirements, please discuss with the recruiter following submission of your application.

Read Full Description
Confirmed 30+ days ago. Posted 12 days ago.

Discover Similar Jobs

Suggested Articles