Just Front-end Jobs

Front-end Jobs nearSan Mateo, CA

Manager, Cloud Operations Engineering (OS/Cloud Experience Required - SRE BG)

new

MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As the worldwide Manager, Cloud Operations Engineering (COE) team, you will directly assist the Technical Services senior leadership team in building out and scaling the COE team, including augmenting its staff, policies, processes and tools. You should be prepared to lead a 24/7/365 global response effort that follows the sun, with COE engineers located in our offices across the globe.

You will interact with other teams within MongoDB to ensure short- and long-term customer success with Atlas, including other teams within Engineering as well as the Sales, Product Marketing, Customer Success and Professional Services organizations. You are skilled in bringing parties together to solve complex problems, and in building trust relationships with both colleagues and customers.

This is a broad-spanning leadership role that will include day-to-day duties such as establishing and checking systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.

At MongoDB you will grow your career and skills, wear multiple hats, and be part of a technical operations team that works at the frontier of cloud services and database systems.

Responsibilities

  • Manage and lead a team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base

  • Scale the worldwide Cloud Operations Engineering team with the strategic implementation of new processes and tools, and hire and ramp the best Cloud Operations Engineers in the world!

  • Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents

  • Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution

  • Automate routine monitoring and troubleshooting tasks

  • Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution

  • Cooperate with our product management and cloud engineering organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure

  • Provide consistent, high-quality feedback and recommendations to our product managers and development teams regarding product defects or recurring patterns of malperformance

  • Inform executive leadership and escalation management personnel of major outages

  • Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (from direct surveillance or through alerts via our Technical Services Engineers)

Requirements

  • Team Lead experience of an oncall DevOps, SRE, or Cloud Operations team (at least 2 years)

  • Experience with being an oncall DevOps, SRE, or Cloud Operations engineer (at least 2 years)

  • Expertise with Linux system administration and networking technologies like DNS, TCP/IP, etc.

  • Knowledgeable about a wide range of web and internet technologies

  • Knowledge of database operations and concepts

  • Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)

  • Experience in monitoring, system performance data collection and analysis, and reporting

  • Capability to write small programs/scripts to solve both short-term systems problems

  • A CS/CE degree or equivalent experience

  • At least 1 of the following programming languages: Java, Go, Python, Javascript

  • A keen interest in learning new things

Nice To Have

  • MongoDB

  • Splunk

  • Kubernetes

*MongoDB, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Apply

Manager, Cloud Operations Engineering (OS/Cloud Experience Required - SRE BG)

new

MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As the worldwide Manager, Cloud Operations Engineering (COE) team, you will directly assist the Technical Services senior leadership team in building out and scaling the COE team, including augmenting its staff, policies, processes and tools. You should be prepared to lead a 24/7/365 global response effort that follows the sun, with COE engineers located in our offices across the globe.

You will interact with other teams within MongoDB to ensure short- and long-term customer success with Atlas, including other teams within Engineering as well as the Sales, Product Marketing, Customer Success and Professional Services organizations. You are skilled in bringing parties together to solve complex problems, and in building trust relationships with both colleagues and customers.

This is a broad-spanning leadership role that will include day-to-day duties such as establishing and checking systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.

At MongoDB you will grow your career and skills, wear multiple hats, and be part of a technical operations team that works at the frontier of cloud services and database systems.

Responsibilities

  • Manage and lead a team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base

  • Scale the worldwide Cloud Operations Engineering team with the strategic implementation of new processes and tools, and hire and ramp the best Cloud Operations Engineers in the world!

  • Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents

  • Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution

  • Automate routine monitoring and troubleshooting tasks

  • Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution

  • Cooperate with our product management and cloud engineering organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure

  • Provide consistent, high-quality feedback and recommendations to our product managers and development teams regarding product defects or recurring patterns of malperformance

  • Inform executive leadership and escalation management personnel of major outages

  • Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (from direct surveillance or through alerts via our Technical Services Engineers)

Requirements

  • Team Lead experience of an oncall DevOps, SRE, or Cloud Operations team (at least 2 years)

  • Experience with being an oncall DevOps, SRE, or Cloud Operations engineer (at least 2 years)

  • Expertise with Linux system administration and networking technologies like DNS, TCP/IP, etc.

  • Knowledgeable about a wide range of web and internet technologies

  • Knowledge of database operations and concepts

  • Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)

  • Experience in monitoring, system performance data collection and analysis, and reporting

  • Capability to write small programs/scripts to solve both short-term systems problems

  • A CS/CE degree or equivalent experience

  • At least 1 of the following programming languages: Java, Go, Python, Javascript

  • A keen interest in learning new things

Nice To Have

  • MongoDB

  • Splunk

  • Kubernetes

*MongoDB, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Apply

Remote Jobs

Sorry, no listings for this city at the moment.

Jobs farther away

Build Release Engineer

Sailpoint Technologies, Inc. in Austin, TX 1490 mi java continuous-integration saas agile devops sysadmin

SailPoint is the Worldwide Leader for Enterprise-Class IAM
We minimize risk and maximize business growth by managing access to data and resources across your enterprise. We do it effectively and securely for every person who interacts with your organization—any user, on any device, anywhere in the world.
We were first to recognize that companies could benefit from an approach to identity that addresses both IT and business priorities. We developed a unique risk-based model and leveraged that approach for everything from compliance to user provisioning. Then we followed that with the industry's first solution for truly extending enterprise identity management to applications in the cloud.
Today, we offer comprehensive products that can handle enterprise IAM on-premises or as a cloud-based service. This gives you the freedom to choose the best solution for your current needs, while at the same time establishing a clear path for future growth.
IdentityNow is SailPoint’s Identity as a Service (IDaaS) product, and the Release Engineer is a crucial role of the IdentityNow team. He/She will proactively work with Engineering, DevOps, Product, Services, and other functional departments to implement and operate SaaS software release best practices. The ideal candidate will be a self-starter who enjoys a fast-paced job and understands all the pieces that have to align correctly to get code safely and frequently from development to production.
Responsibilities:

  • Build, maintain and monitor all infrastructure required to support the release pipeline, in accordance with pre-existing DevOps standards

  • Monitor and troubleshoot all software builds

  • Perform deployments of each service as needed, with the goal of automating yourself out of needing to manually release software

  • Create methodology to move releases safely from one environment to the next, while still complying with our governing standards

  • Develop and/or improve standards and policies around proper software delivery principles, and encourage/promote CI/CD

  • Manage cross-functional requirements working with DevOps, Engineering, Product, Services, and other departments

  • Produce metrics, dashboards, and alerts that report on overall release quality and status

    Background & Experience:

  • Experience in a 24x7 Agile, SaaS environment

  • You’ve worked in a shop employing CI/CD

  • You’re an expert in one or more release automation products (Jenkins, GoCD, etc)

  • You’re great with at least one scripting language (Ruby, Python, etc)

  • You may not know Java, but you know how to build something written in Java

  • You have a good understanding of networking concepts

  • You have strong interpersonal and team skills, and the ability to set and enforce process and influence engineers who are not direct reports

    In addition to the above requirements, the ideal candidate will have experience with one or more of the following criteria:

  • System administration, system configuration, and system debugging experience

  • Knowledge of cloud infrastructure environments like AWS

  • Knowledge of container technology

    Education:

  • Bachelor's degree in Computer Science or other technical disciplines, or equivalent experience

Apply