Save Job Back to Search Job Description Summary Similar JobsBe recognised as a Senior Leader in a leading IT Services OrganizationChance to drive AI-powered observability across enterprise environments.About Our ClientOur client is a large, globally established organisation with a strong focus on technology transformation, cloud adoption, infrastructure resilience, and operational excellence. The organisation is investing heavily in AI-driven observability, automation, and Site Reliability Engineering practices to improve service performance, customer experience, and business value.Job DescriptionLead the establishment and growth of an enterprise-wide Observability Centre of Excellence (CoE).Drive adoption of observability best practices across infrastructure, cloud and application environments.Own the monitoring strategy covering metrics, logs, traces, APM and user experience monitoring.Develop and standardise observability frameworks, onboarding processes, governance and operating models.Partner with engineering, platform and business teams to improve service reliability and operational resilience.Define and implement SRE practices including SLI/SLO frameworks, error budgets and reliability metrics.Drive automation, AIOps, anomaly detection and proactive incident management initiatives.Build executive dashboards and translate operational insights into measurable business outcomes.Mentor a specialist team of observability and reliability engineers while remaining hands-on with architecture and solution design.Influence cloud-native monitoring strategies across AWS, Azure, Kubernetes and hybrid environments.The Successful Applicant12+ years of experience in Observability, Site Reliability Engineering (SRE), Platform Engineering or Infrastructure Monitoring.Proven experience building or scaling an Observability Practice, CoE or Reliability Engineering capability.Strong hands-on expertise with Datadog, Dynatrace, New Relic, Grafana, Prometheus and OpenTelemetry.Experience in application performance monitoring, distributed tracing, log analytics and telemetry instrumentation.Strong knowledge of cloud platforms including AWS and Azure.Experience with Kubernetes, Docker, CI/CD pipelines and modern engineering practices.Exposure to AIOps, observability automation, anomaly detection and AI-enabled operations.Deep understanding of incident management, RCA processes, reliability engineering and operational excellence.Previous experience leading small, high-performing teams of engineers, architects or SMEs. (5-10 Member Teams)Excellent stakeholder management and communication skills with the ability to engage senior leadership.What's on OfferJoin a high-visibility role where you'll have the opportunity to build and shape an enterprise-wide Observability Centre of Excellence from the ground up. You'll influence technology strategy across cloud, infrastructure and application landscapes while driving modern SRE, AIOps and reliability engineering practices. This position offers strong exposure to executive stakeholders, ownership of observability transformation initiatives, and the chance to leverage cutting-edge AI-driven monitoring technologies.Quote job refJN-012026-6935256Job summaryFunctionInformation TechnologySub SectorIT ArchitectureWhat is your area of specialisation?Technology & TelecomsLocationNoidaJob TypePermanentJob ReferenceJN-012026-6935256Work from HomeWork from Home or Hybrid