Jobs
>
Vancouver

    Senior Site Reliability Engineer - British Columbia, Canada - Red Hat

    Default job background
    Description

    About the job

    Red Hat is seeking a Senior Site Reliability Engineer (SRE) to develop, scale, and operate our OpenShift managed cloud services. OpenShift is Red Hat's enterprise Kubernetes distribution. As an SRE you will contribute to running OpenShift at scale by enabling customer self-service, making our monitoring system more sustainable, and eliminating work through automation.

    On the SRE team, you will have the opportunity to influence the complex challenges of scale which are unique to Red Hat managed cloud services, while using your skills in coding, operations, and large-scale distributed system design.

    Red Hat relies on teamwork and openness for its success. We are a global team and strive to cultivate a transparent environment that makes room for different voices. We learn from our failures in a blameless environment to support the continuous improvement of the team. At Red Hat, your individual contributions have more visibility than most large companies, and visibility means career opportunities and growth.

    What you will do

    The day-to-day responsibilities of an SRE involve working with live systems and coding automation. As an SRE you will be expected to:

    • Contribute code to increase the scalability and reliability of the service
    • Contribute software tests and participate in peer review to increase the quality of our codebase
    • Help and develop peers' capabilities through knowledge sharing, mentoring, and collaboration
    • Participate in a regular on-call schedule, including occasional paid weekends and holidays
    • Practice sustainable incident response and blameless postmortems
    • Resolve customer issues escalated from the Red Hat Global Support team
    • Work within a small agile team to develop and improve SRE software, support your peers, plan and self-improve

    What you will bring

    A bachelor's degree in Computer Science or a related technical field involving software or systems engineering is required. However, hands-on experience that demonstrates your ability and interest in Site Reliability Engineering are valuable to us, and may be considered in lieu of degree requirements. You must have some experience programming in at least one of these languages: Python, Golang, Java, C, C++ or another object-oriented language. You must have experience working with public clouds such as AWS, GCP, or Azure. You must also have the ability to collaboratively troubleshoot and solve problems in a team setting.

    As an SRE you will be most successful if you have some experience troubleshooting an as-a-service offering (SaaS, PaaS, etc.) and some experience working with complex distributed systems. Direct experience with Kubernetes or OpenShift is a plus. We like to see a demonstrated ability to debug, optimize code and automate routine tasks. We are Red Hat, so you need a basic understanding of Unix/Linux operating systems.

    Desired skills

    • 5+ years of experience managing Linux servers running Red Hat Enterprise Linux (RHEL), CentOS, or Fedora hosted at a cloud provider such as Amazon Web Services (AWS), Google Compute Engine (GCE), or Microsoft Azure
    • 3+ years of experience with enterprise systems monitoring; knowledge of Prometheus is a plus
    • 3+ years of experience with enterprise configuration management software like Ansible by Red Hat, Puppet, or Chef
    • 2+ years of experience programming with at least one object-oriented language; Golang, Java, or Python are preferred
    • 2+ years of experience delivering a hosted service
    • Demonstrated ability to quickly and accurately troubleshoot system issues
    • Solid understanding of standard TCP/IP networking and common protocols like DNS and
    • Solid communications skills and experience working directly with and presenting to customers
    • 1+ year(s) of experience with Kubernetes is a plus
    • 1+ year(s) of experience with docker-based containers is a plus
    #J-18808-Ljbffr


  • Stafflink Vancouver, BC, Canada

    Job Description · Position: Site Reliability Engineer · Duration: 12 Months · Location: Principally remote, with at least one day per month in office for applicants in the lower mainland. Local candidates are given preference. · Work hours: Monday – Friday, 9:00 am – 5:00 ...


  • T-Net British Columbia Vancouver, BC, Canada

    Site Reliability Engineer Co-op (Sept May 2025) Job Overview · Our innovative technology transforms the way that organisations make decisions, allowing them to elevate their employees and drive better business outcomes. Embarking on an exciting new chapter in our growth story, w ...


  • Dapper Labs Vancouver, Canada Full time

    We're looking for a Site Reliability Engineer who wants to be at the technical core of an organization that's completely reshaping how distributed applications on blockchains can reach massive audiences. · You will join a Site Reliability Engineering team that has the ability t ...


  • Visier Inc. Vancouver, BC, Canada

    Our co-op experience is unique and designed to prepare you for professional success as you work on real, impactful work from the beginning. Our ultimate goal is to give you the mentorship, training, and work experience you need to start your career. A number of our students retur ...


  • Axiom Zen Vancouver, Canada

    We're looking for a Site Reliability Engineer who wants to be at the technical core of an organization that's completely reshaping how distributed applications on blockchains can reach massive audiences. · You will join a Site Reliability Engineering team that has the ability to ...


  • Visier, Inc Vancouver, BC, Canada

    Visier Co-op Opportunity · Our innovative technology transforms the way that organisations make decisions, allowing them to elevate their employees and drive better business outcomes. Embarking on an exciting new chapter in our growth story, we are looking for talented individua ...


  • Visier, Inc Vancouver, BC, Canada

    Our innovative technology transforms the way that organizations make decisions, allowing them to elevate their employees and drive better business outcomes. Embarking on an exciting new chapter in our growth story, we are looking for talented individuals who can help both Visier ...


  • Razr Marketing Vancouver, BC, Canada

    Senior Site Reliability Engineer · These values have made RAZR what it is for years, and today, they are more important than ever. You can't wait to get out of bed in the morning & get on with your day · We are seeking a skilled and motivated Site Reliability Engineer (SRE) to ...


  • Sentry Vancouver, BC, Canada

    About the role · The Site Reliability Engineering team is responsible for the deployment, configuration, maintenance and monitoring of Sentry's hosted platform. We do this by leveraging automation tools to automatically spin up and scale services to meet the traffic demands of 1 ...


  • RAZR Marketing, Inc. Vancouver, BC, Canada

    You will be required to be in our office In Vancouver, BC three times per week. · These values have made RAZR what it is for years, and today, they are more important than ever. You can't wait to get out of bed in the morning & get on with your day · We are seeking a skilled an ...


  • TEEMA Vancouver, Canada Full time

    MUST LIVE IN CANADA NEAR AN AIRPORT · Looking for a technical lead with 10+ years of DevOps/SRE experience · MUST HAVE - 5+ years permanent residence or Citizenship (cant have lived out of Canada for the last 5 years) · MUST LIVE IN CANADA NEAR AN AIRPORT · Looking for a technica ...


  • Stafflink Vancouver, BC, Canada

    Position: Site Reliability Engineer · Location: Principally remote, with at least one day per month in office for applicants in the lower mainland. Local candidates are given preference. · Monday - Friday, 9:00 am - 5:00 pm PST · Serve as the subject matter expert (SME) for Dynat ...


  • Dapper Labs Vancouver, BC, Canada

    We're looking for a Site Reliability Engineer who wants to be at the technical core of an organization that's completely reshaping how distributed applications on blockchains can reach massive audiences. · You will join a Site Reliability Engineering team that has the ability to ...


  • Taurus SA Vancouver, Canada CDI

    Are you ready to take on an entrepreneurial challenge in the digital asset industry? Taurus, a global leader in digital asset infrastructure, has an exciting opportunity for you. · Founded in April 2018, Taurus provides enterprise-grade solutions to issue, custody, and trade dig ...


  • Red Hat, Inc. British Columbia, Canada

    About the job · Red Hat is seeking a Senior Site Reliability Engineer (SRE) to develop, scale, and operate our OpenShift managed cloud services. OpenShift is Red Hat's enterprise Kubernetes distribution. As an SRE you will contribute to running OpenShift at scale by enabling cu ...


  • Red Hat, Inc. British Columbia, Canada

    About the job · Red Hat is seeking a Senior Site Reliability Engineer (SRE) to develop, scale, and operate our OpenShift managed cloud services. . OpenShift is a cloud native application platform for the enterprise, powered by Kubernetes. As an SRE you will contribute to runnin ...


  • Electronic Arts Vancouver, Canada

    EA's Digital Platform (EADP) organization drives important technology decisions and investments for EA on a global basis, across all divisions and studio teams. Technology and engineering leadership at EA is essential to making the industry's best games and services and the EADP ...


  • Electronic Arts Vancouver, Canada Regular

    Responsibilities · : You will create monitoring, alerting and dashboarding solutions that improve visibility into EA's application performance and business metrics. · You will help design and develop robust, supportable tools to automate the deployment and management of distrib ...


  • New Value Solutions Richmond, Canada

    New Value Solutions, a national IT consulting company, is seeking a Site Reliability Engineer for our client. · Responsibilities: · Serve as the subject matter expert (SME) for Dynatrace, responsible for configuring, optimizing, and managing Dynatrace monitoring solutions. · Des ...


  • Sentry Vancouver, BC, Canada

    The Site Reliability Engineering team is responsible for the deployment, configuration, maintenance and monitoring of Sentry's hosted platform. We do this by leveraging automation tools to automatically spin up and scale services to meet the traffic demands of 1,000,000+ develope ...