Brief info about Vinted
Our mission is to make second-hand the first choice, and we're looking for people who want to help us get there. Every day, we work together to help our members buy and sell pre-loved clothing and lifestyle items, giving each piece a second life – or even a third. The Vinted Group is made up of three business units that support this mission:
- Vinted Marketplace is Europe’s leading platform for second-hand fashion and a go-to destination for all kinds of pre-loved items, with a growing range of categories.
- Vinted Go enhances the shipping experience with a vast network of over 500,000 pick-up and drop-off points.
- Vinted Pay is dedicated to bringing secure, reliable payments to buyers and sellers across Europe.
Information about the position
As a Site Reliability Engineer on the Hardware team, you’ll be responsible for maintaining 10,000+ servers across Vinted’s data centers. You’ll help monitor and debug production servers to ensure we meet our established SLOs. This will involve diving deep into hardware configuration, BIOS, BMC, and Linux systems to identify bugs, troubleshoot issues, and test performance. You’ll also automate server provisioning, build Grafana dashboards, and set up monitoring alerts. You’ll work closely with data center administrators, who handle physical server repairs, helping them identify issues and validate fixes after repairs are completed. You’ll also collaborate with other teams to support them with hardware- and software-related challenges.
Responsibilities
- Lead data center infrastructure reliability initiatives across 10,000+ physical servers.
- Build up automation tooling for server provisioning, configuration management, testing, and large-scale deployments.
- Create scripts, tools, and workflows to reduce manual operational work and improve infrastructure efficiency.
- Assist with troubleshooting hardware, BIOS, BMC, Redfish, Linux OS, networking, and performance-related issues.
- Communicate with infrastructure, platform, and engineering teams to support secure, scalable, and performant services.
- Participate in production incident response, root cause analysis, and on-call rotation.
- Monitor data center infrastructure, server health, SLOs, Grafana dashboards, alerts, and production systems.
- Oversee testing and benchmarking of new server hardware, CPU, and GPU models.
- Keep track of data center assets, hardware configurations, and infrastructure documentation in DCIM/NetBox.
- Work with the Platform Foundations DC team to create reliable, automated, and scalable infrastructure solutions.
- Maintain configuration-as-code repositories and automated tests.
- Improve server provisioning and configurations.
- Support other engineering teams who are running their services on servers.
- Collaborate with engineers to identify infrastructure bottlenecks and implement long-term improvements.