Senior Site Reliability Engineer
Multi Media LLC
Multi Media LLC is hiring a remote senior site reliability engineer to improve resilience, performance, observability, and automation for a large real-time streaming platform. IT Support Group readers with production Linux, networking, cloud, and DevOps experience should find this a direct infrastructure-operations match.
What you would work on:
- Analyze APM and distributed telemetry to find instability and improve system scalability, reliability, and performance.
- Build DevOps tooling and automation while operating infrastructure across data-center hardware and public cloud environments.
- Plan for failures and disasters, administer databases and key-value stores, and reduce operational surprises and downtime.
- Participate in incident response, write postmortems, and collaborate with engineering teams on production reliability.
Good fit if:
- You have SRE, DevOps, or production software-engineering experience running web applications at scale.
- You are comfortable with Linux internals, Bash, Python or Go, networking, database operations, and on-call work.
- You have used infrastructure tools such as Terraform, Ansible, Docker, Kubernetes, ArgoCD, or Helm.
Curated from Himalayas for IT Support Group readers. This is an external listing; use the apply link for the source listing and latest details.