About LivePerson
LivePerson (NASDAQ:LPSN) is a leading customer engagement company, creating digital experiences powered by Curiously Human AI. Every person is unique, and our technology makes it possible for companies, including leading brands like HSBC, Orange, and GM Financial, to treat their audiences that way at scale. Nearly a billion conversational interactions are powered by our Conversational Cloud each month.
LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. As the market leader in real-time intelligent customer engagement, we are a B2B SaaS company with 20 years of experience and the heart of a startup.
Role Overview
The Cloud SRE team at LivePerson is looking for a Senior Site Reliability Engineer (SRE) to help design, build, and operate highly reliable, scalable, and secure cloud infrastructure and services. As a Senior SRE, you will have the opportunity to influence technical direction, drive engineering best practices, mentor other engineers, and take ownership of critical infrastructure and production services.
Role and Responsibilities
- Design, build, and maintain highly available, scalable, secure, and resilient infrastructure and services across cloud and hybrid environments, with a strong focus on Google Cloud Platform (GCP).
- Develop and maintain automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern DevOps/SRE tools.
- Design, deploy, and operate Kubernetes-based platforms and workloads, including troubleshooting complex issues across clusters and production environments.
- Design, implement, and maintain GitOps-based deployment workflows using Kubernetes, Helm, and FluxCD.
- Design, develop, and maintain CI/CD pipelines using GitLab CI/CD to automate application and infrastructure delivery.
- Establish and improve observability practices using metrics, logs, traces, dashboards, and alerting to ensure systems are reliable and actionable from an operational perspective.
- Define, implement, and continuously improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics for critical services.
- Participate in and lead incident response, troubleshooting complex production issues, identifying root causes, and driving corrective and preventative actions.
- Drive automation and operational improvements that reduce manual work, eliminate repetitive tasks, and minimize operational toil.
- Partner closely with software engineering, security, networking, and other infrastructure teams to design reliable and secure solutions throughout the software development lifecycle.
- Perform capacity planning, performance analysis, and reliability assessments to ensure systems can scale with business and customer needs.
- Contribute to architecture and technical design decisions, challenging existing solutions and identifying opportunities to improve scalability, reliability, security, and operational efficiency.
- Establish and promote engineering standards, best practices, and operational processes across the organization.
- Mentor and support other engineers, sharing knowledge and helping raise the technical and operational maturity of the team.
- Participate in an on-call rotation and provide support for critical production services when required.