Overview
Join the Azure Specialized AI Infrastructure team in India to drive advancements in Artificial Intelligence (AI) and support high-performance infrastructure for generative AI workloads. As a Site Reliability Engineer, you will automate, and maintain large-scale distributed systems powering latest AI applications and machine learning models. Your primary focus will be on the reliability, scalability, and performance of AI infrastructure, ensuring seamless operations for mission-critical AI services. The role emphasizes a start-up mindset, collaboration, and customer advocacy.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads.
Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability.
Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance.
Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker).
Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices.
Qualifications
Required Qualifications:
Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administrationOR Bachelor's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration
OR equivalent experience.
8+ years of professional software engineering experience, with 5+ years in service operations, monitoring, and reliability improvement for infrastructure.
1+ years experience with incident management and reliability engineering in cloud or AI environments.
Other Requirements:
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
Preferred Qualifications:Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administrationOR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration
OR equivalent experience.
2+ years technical experience working with large-scale cloud or distributed systems.
1+ years experience in distributed systems and/or cloud platforms (Azure, Kubernetes, Docker, containers ecosystem).
1+ years experience with GPUs, InfiniBand, or similar high-performance technologies.
#azurecorejobs
#BRAVO2026
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Анонимная аналитика
Мы используем анонимную аналитику, чтобы улучшать поиск вакансий и работу сайта.