Finding the best job has never been easier
Share
About the Job
Red Hat is looking for a Platform Engineer. In this role, you will help architect, implement, improve, and support the OpenShift-based platform that runs many of Red Hat’s most important multi-tenant Software-as-a-Service (SaaS) and managed-service offerings. Using your expertise in site reliability engineering (SRE) principles, you will help create an environment where reliability, scalability, and security come first, and are not treated as an afterthought.
In this role, you will spend a portion of your time working across teams to define and iterate upon processes for onboarding new managed services at Red Hat and demonstrate good judgment in employing onboarding methods and techniques that can be repeated and iterated upon. You will also contribute to the codebase of command-and-control software that automates the building, deployment, monitoring, and alerting of Red Hat managed services. The remainder will be spent on various other tasks, such as diagnosing issues, planning, documenting, and mentoring.
What You Will Do
Design, write, and maintain software (primarily in Python and Golang) that automates the deployment, monitoring, and maintenance of Red Hat managed services
Assist our product engineering teams with the following:
Onboarding of new services onto our OpenShift-based platform
Adhering to cloud-native design principles & best practices to ensure reliability, scalability, and security
Contribute to documents, like standard operating procedures (SOPs) and playbooks, that assist in issue resolution and new-service onboarding
Participate in an Agile Scrum team that scopes, prioritizes, and allocates work items
Participate in an on-call rotation that is responsible for responding to service incidents
What You Will Bring
Bachelor's or Master's degree in Computer Science, engineering, math or equivalent practical experience
Background writing object-oriented automation software in Python or Golang
Background administering production cloud-native services, preferably containerized and deployed via a container-orchestration system like Kubernetes or OpenShift
Experience diagnosing service failures and carrying out incident response procedures
Familiarity with Linux operating system and its configuration
Ability to effectively work in a globally distributed team
The following experience is considered a plus:
Understanding of computer networking and protocols, including TCP/IP and DNS
Understanding of computer security and cryptography basics, including certificates, TLS, and credential-storage systems like Vault
Familiarity with CI/CD pipeline concepts and systems, like Jenkins and Tekton/Argo
Familiarity with observability tools like Prometheus and Grafana, and how to define metrics that can be used to measure service health and reliability
The salary range for this position is $90,480.00 - $144,660.00. Actual offer will be based on your qualifications.
Pay Transparency
● Comprehensive medical, dental, and vision coverage
● Flexible Spending Account - healthcare and dependent care
● Health Savings Account - high deductible medical plan
● Retirement 401(k) with employer match
● Paid time off and holidays
● Paid parental leave plans for all new parents
● Leave benefits including disability, paid family medical leave, and paid military leave
These jobs might be a good fit