About LivePerson
LivePerson (NASDAQ:LPSN) is a leading customer engagement company, creating digital experiences powered by Curiously Human AI. Every person is unique, and our technology makes it possible for companies, including leading brands like HSBC, Orange, and GM Financial, to treat their audiences that way at scale. Nearly a billion conversational interactions are powered by our Conversational Cloud each month.
LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. As the market leader in real-time intelligent customer engagement, we are a B2B SaaS company with 20 years of experience and the heart of a startup.
Role Overview
The Cloud DevOps team at LivePerson is looking for an SRE Team Lead to lead a team of Site Reliability Engineers responsible for building, operating, and continuously improving highly reliable, scalable, and secure cloud infrastructure and services. The ideal candidate is a strong technical leader who combines deep experience in Site Reliability Engineering and cloud infrastructure with a passion for developing people and building high-performing engineering teams.
Responsibilities
- Lead, mentor, and develop a team of SREs, fostering a culture of ownership, collaboration, technical excellence, and continuous improvement.
- Provide technical leadership and direction for the design, implementation, and operation of highly available, scalable, secure, and resilient infrastructure and services.
- Own the team's technical roadmap and ensure alignment with broader engineering and business objectives.
- Plan and prioritize team initiatives, balancing new development, reliability improvements, technical debt, operational work, and business priorities.
- Design and maintain cloud infrastructure and services across cloud and hybrid environments, with a strong focus on Google Cloud Platform (GCP).
- Guide the development of automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern DevOps/SRE technologies.
- Lead the design, deployment, and operation of Kubernetes-based platforms and workloads, including troubleshooting complex production issues.
- Establish and maintain GitOps-based deployment workflows using Kubernetes, Helm, and FluxCD.
- Drive the design, implementation, and continuous improvement of CI/CD pipelines using GitLab CI/CD.
- Establish and improve observability practices using metrics, logs, traces, dashboards, and alerting to provide actionable insights into system health and performance.
- Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics for critical services.
- Lead incident response and ensure effective handling of critical production incidents, including root cause analysis and follow-up corrective actions.
- Drive initiatives to reduce operational toil, eliminate recurring incidents, and improve the overall reliability and operational maturity of the platform.
- Partner with software engineering, security, networking, product, and other infrastructure teams to influence architecture and deliver reliable and secure solutions.
- Lead capacity planning, performance analysis, scalability assessments, and reliability reviews for critical systems.
- Establish and promote engineering standards, operational processes, and best practices across the organization.
- Identify technical risks and dependencies and proactively develop mitigation strategies.
- Drive continuous improvement through post-incident reviews, reliability assessments, and engineering retrospectives.
- Support and participate in an on-call rotation and ensure the team has effective processes for handling production incidents.
- Provide technical mentorship and career guidance to engineers, helping them grow their technical expertise and ownership.
- Participate in hiring, interviewing, onboarding, performance management, and development of SRE team members.
- Build strong relationships with stakeholders and communicate technical risks, priorities, and progress clearly to engineering leadership and partner teams.