Senior Site Reliability Engineer, ML Platforms @ Nvidia

Home > Devops

Senior Site Reliability Engineer, ML Platforms

Nvidia
6 - 11 years
Hyderabad
21 days ago
Email to a friend
Report this job

Job Description

The role involves designing, building, and maintaining services that enable real-time data analytics, streaming, data lakes, observability and ML/AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring capacity, latency, and performance is part of the role.
To succeed in this position, a strong background in SRE practices, systems, networking, coding, capacity management, cloud operations, continuous delivery and deployment, and open-source cloud enabling technologies like Kubernetes and OpenStack is required. Deep understanding of the challenges and standard methodologies of running large-scale distributed systems in production, solving complex issues, automating repetitive tasks, and proactively identifying potential outages is also necessary.
Furthermore, excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. As a Senior SRE at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking.

What you ll be doing:

Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.
Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.
Create tools and automation to reduce operational overhead and eliminate manual tasks.
Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.
Define meaningful and actionable reliability metrics to track and improve system and service reliability.
Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.
Build tools to improve our service observability for faster issue resolution.
Practice sustainable incident response and blameless postmortems

What we need to see:

Minimum of 6+ years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.
Masters or Bachelors degree in Computer Science or Electrical Engineering or CE or equivalent experience.
Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.
Proficiency in incident, change, and problem management processes.
Skilled in problem-solving, root cause analysis, and optimization.
Experience with streaming data infrastructure services, such as Kafka and Spark.
Expertise in building and operating large-scale observability platforms for monitoring and logging (e. g. , ELK, Prometheus).
Proficiency in programming languages such as Python, Go, Perl, or Ruby.
Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.
Experience in deploying, supporting, and supervising services, platforms, and application stacks.

Ways to stand out from the crowd:

Experience operating large-scale distributed systems with strong SLAs.
Excellent coding skills in Python and Go and extensive experience in operating data platforms.
Knowledge of CI/CD systems, such as Jenkins and GitHub Actions.
Familiarity with Infrastructure as Code (IaC) methodologies and tools.
Excellent interpersonal skills for identifying and communicating data-driven insights.

Job Classification

Industry: Electronic Components / Semiconductors
Functional Area / Department: Engineering - Software & QA
Role Category: DevOps
Role: Site Reliability Engineer
Employement Type: Full time

Contact Details:

Company: Nvidia
Location(s): Hyderabad

+ View Contact

Login

Candidates can login here to view contacts and apply.

Sign In Sign Up

Email:

Password:

Password too short

To create your profile, apply for a job or make a registration

Your name (*)

Email (*)

Mobile (*)

Preferred City (* max. 2 w/comma)

Designation / Expected Role

Current / Recent Company (*)

Experience (*)

Expected Salary (*)

Desired Industry (*):

Functional area / Department (*):

Enter Skills (key skills, subjects, technologies & roles to use in search)

Write briefly about yourself, your experience and education (*)

Attach Resume Max 2.38 MB (RTF, PDF, DOC, DOCX formats only parsed)

Please, check the file size and type.

Add social media [ + ]

Create password

I agree with website service terms and conditions

Candidates are expected to provide most recent and accurate profile information, inappropriate content is strictly prohibited!

Keyskills: Automation Networking Coding Problem management Perl Open source Ruby Monitoring Python

Fraud Alert to job seekers!

₹ Not Disclosed

Job application

We will notify the employer with your details. You can also attach a resume or a cover letter.

Sign In Sign Up

Email:

Password:

Password too short

To create your profile, apply for a job or make a registration

Your name (*)

Email (*)

Mobile (*)

Preferred City (* max. 2 w/comma)

Designation / Expected Role

Current / Recent Company (*)

Experience (*)

Expected Salary (*)

Desired Industry (*):

Functional area / Department (*):

Enter Skills (key skills, subjects, technologies & roles to use in search)

Write briefly about yourself, your experience and education (*)

Attach ResumeMax 2.38 MB (RTF, PDF, DOC, DOCX formats only parsed)

Please, check the file size and type.

Add social media [ + ]

Create password

I agree with website service terms and conditions

Similar positions

Software Engineer, Site Reliability Engineering

Google

2 - 7 years

Bengaluru

8 days ago

₹ Not Disclosed

Technical Solutions Engineer, Infrastructure, Serverless

Google

2 - 7 years

Pune

8 days ago

₹ Not Disclosed

Senior Devops Engineer - Geforce Now Cloud

Nvidia

1 - 7 years

Pune

4 days ago

₹ Not Disclosed

Site Reliability Engineer

Globallogic

10 - 15 years

Hyderabad

4 days ago

₹ Not Disclosed

Nvidia

Nvidia Corporation

Senior Site Reliability Engineer, ML Platforms @ Nvidia

Home > Devops