About Camunda

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. With built-in governance, auditability, and human oversight, Camunda gives enterprises the control they need to move AI from pilots to production — safely and at scale. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda helps enterprises boost operational efficiency, accelerate time-to-value, and deliver better customer experiences.

Fully remote and global, we are in the middle of something bigger: transforming into an AI-first organisation, built on our own platform. We use Agentic AI to automate, orchestrate intelligent processes, and elevate human contribution across every team.

About the role

If you’re passionate about pushing the limits of reliability, fault-tolerance, and operational excellence in complex distributed systems, Camunda offers the perfect stage. As a Senior Software Engineer, Backend, you’ll take ownership of automated reliability testing and chaos engineering at the core of Camunda 8. This role thrives on curiosity, experimentation, and relentless improvement—you’ll break things (on purpose!) in safe environments to strengthen our platform before customers ever feel the impact.

What You’ll Be Doing

  • Design, implement, and execute automated chaos experiments and reliability tests, validating Camunda's platform under real-world scenarios.
  • Investigate, root-cause and debug potential failures or performance regressions of our Java based products.
  • Continuously improve existing load testing, chaos engineering and observability infrastructure to keep our system reliable and performant (using tools like Java, Go, Kubernetes, Prometheus, Grafana, etc.).
  • Introduce new tooling and approaches for reliability testing, driving measurable improvements to operational experience and user outcomes.
  • Collaborate closely with QA and engineering, sharing learnings and influencing team roadmaps based on experiment findings.
  • Champion a pragmatic, autonomous approach to software design, solving tough problems and learning new principles and technologies.
  • Advocate for user-centric reliability by using and understanding Camunda’s products firsthand.

Compensation

We offer competitive, fair, and transparent compensation. Salary ranges are location-based, with Standard and Major markets (global tech hubs) reflecting local competition. Equity is also offered through our Virtual Stock Option Plan (VSOP).

Benefits & Perks

  • Remote & Flexible: Work from anywhere with a home office budget, co-working space support, and flexible time off.
  • In Person Connection: Annual Kickoff, team offsites, and Camundi Connection Budgets for meetups and local gatherings.
  • Health & Wellbeing: Locally tailored healthcare, Modern Health for global mental wellbeing, and a Live Well Lifestyle Spending Account (LSA) for active living, family care, and personal passions.
  • Financial Security: Retirement and pension plans, plus life and disability insurance where relevant.
  • Professional Growth: Up to $/€/£1,000 per year for self-driven learning (courses, certifications, books).

What You Bring

  • Ability and/or willingness to use our product.
  • 5+ years of experience in backend software engineering (Java).
  • Proven drive to experiment, learn new tech, and conduct automated reliability and chaos testing in distributed environments.
  • Deep enthusiasm for improving performance and fault-tolerance in production systems.
  • Autonomous, pragmatic problem-solving skills with an ability to guide and influence others.
  • Willingness and ability to use Camunda products and approach reliability from a user’s perspective.

Nice to Haves

  • Hands-on experience with Go or other languages, in a multi-language context.
  • Working as SRE on distributed systems, with a strong software engineering mindset.
  • Working knowledge of Kubernetes, Helm Charts, Operators, and related production infrastructure.
  • Experience running apps in production, with solid skills in monitoring, troubleshooting, and performance analysis.
  • Background in chaos engineering, automated reliability, load, or performance testing.