Apply for Job
Lead Matrixx SME DevOps Engineer
Petaling Jaya, MY
Role Mission:
As a Lead Provisioning DevOps Engineer at StarHub, you will own the end-to-end DevOps strategy, platform reliability, and release governance for mission-critical DXP Provisioning platforms. You will lead the design and evolution of CI/CD pipelines, Kubernetes platform architecture, cloud infrastructure, observability, and production operations practices that support high-volume telco customer journeys across provisioning, charging, billing, and network activation.
This role requires strong working knowledge of end-to-end telco flows, including CRM -> Order Management -> Provisioning -> Charging -> Network, and hands-on exposure to provisioning or online charging environments such as IN, OCS, PCRF/PCF, UDM, HLR/HSS, CRM/BSS integrations, mediation, balance management, and rating or charging flows. You will work across engineering, operations, product, vendor, and business teams to ensure customer-impacting journeys are delivered safely, monitored effectively, and operated with a strong RCA and revenue-assurance mindset.
The role goes beyond execution. You will drive DevOps best practices, establish SRE capabilities, improve production stability, and guide engineers in delivering faster, safer, and more reliable platform changes across high-availability telco environments.
Responsibilities:
- Drive DevOps strategy and operating model for DXP Provisioning, CRM/OM, and charging-adjacent platforms, acting as the technical escalation point for complex production issues.
- Ensure end-to-end reliability of provisioning and telco customer journeys across CRM, order orchestration, provisioning, charging, billing, mediation, and network activation flows.
- Support and improve integrations with telco network and charging systems, including IN/OCS, PCRF/PCF, HLR/HSS/UDM, CRM/BSS, mediation platforms, and external partner APIs.
- Lead design reviews and operational readiness for provisioning flow design, order orchestration, service activation, service modification, suspension, restoration, termination, and charging-triggered workflows.
- Troubleshoot real-time provisioning and charging issues involving SOAP/REST APIs, middleware, Diameter/CAMEL interfaces, event flows, Kafka-based integrations, and microservices.
- Support production incidents, failed provisioning cases, charging mismatches, balance or rating defects, revenue leakage investigations, reconciliation gaps, and high-severity operational escalations.
- Architect and govern end-to-end CI/CD pipelines using Jenkins, Pipeline-as-Code, GitLab CI, and GitOps practices such as Argo CD or Flux, enabling safe and repeatable releases.
- Lead trunk-based development adoption, automated testing, quality gates, deployment validation, rollback controls, and release safety mechanisms across microservices.
- Design and operate multi-cluster Kubernetes platforms such as EKS and OpenShift, including networking, RBAC, scaling, resilience, workload standardization, and platform governance.
- Govern AWS cloud infrastructure and Infrastructure as Code using Terraform and/or CloudFormation, ensuring security, scalability, cost efficiency, and operational resilience.
- Apply SRE practices by defining SLOs/SLIs, improving reliability and performance, leading incident management, and driving structured root-cause analysis and problem management.
- Establish observability and operational excellence using Splunk, Prometheus, Grafana, CloudWatch, ELK, alerting, reconciliation dashboards, SLA monitoring, and on-call tooling.
- Drive automation for provisioning operations, incident response, reconciliation, environment provisioning, release validation, and platform maintenance.
- Contribute to code improvements, architecture reviews, operational governance, engineering standards, runbooks, documentation, and knowledge sharing.
- Mentor engineers and influence cross-functional teams to improve delivery quality, platform maturity, and production ownership.
Qualifications:
- Bachelor's degree in Computer Science, Software Engineering, Telecommunications, or a related field.
- 6+ years of hands-on experience across DevOps, SRE, platform engineering, software engineering, or telco provisioning/charging operations, with at least 2-3 years in a senior or technical leadership role.
- Strong exposure to provisioning systems and/or online charging platforms in a telco environment, such as IN, OCS, PCRF/PCF, HLR/HSS/UDM, CRM/BSS integrations, mediation, balance management, or rating and charging flows.
- Strong understanding of end-to-end telco flows across CRM, order management, provisioning, charging, billing, mediation, and network systems.
- Hands-on experience supporting production incidents, failed provisioning cases, charging mismatches, reconciliation issues, revenue assurance investigations, SLA breaches, and high-severity escalations.
- Strong proficiency in Java and Spring Boot, with hands-on experience supporting microservices-based and distributed architectures.
- Good understanding of APIs, middleware, orchestration patterns, event-driven integrations, SOAP/REST services, Kafka, and real-time integration troubleshooting.
- Familiarity with telco protocols and charging/network concepts such as Diameter, CAMEL, real-time rating, balance management, PCRF/PCF policy flows, and subscriber data management.
- Deep understanding of scalability, reliability engineering, high availability, production system performance, and operational governance.
- Advanced hands-on experience with:
- Kubernetes, preferably EKS and/or OpenShift, Helm, and container orchestration
- CI/CD platforms such as Jenkins or GitLab CI and GitOps practices
- AWS cloud services and cloud-native architecture
- Infrastructure as Code using Terraform and/or CloudFormation
- Observability tools such as Splunk, Grafana, Prometheus, CloudWatch, ELK, and alerting platforms
- Strong SQL skills with hands-on PostgreSQL administration, query analysis, and performance tuning experience.
- Proficiency in Python and shell scripting for automation, operational tooling, and platform maintenance.
- Experience implementing reconciliation, SLA monitoring, observability, dashboarding, and alert tuning across business-critical workflows.
- Proven experience working in Agile/Scrum environments, with strong stakeholder management, communication, vendor coordination, and technical leadership skills.
- Ability to operate effectively in a fast-paced, production-critical, high-availability telco environment.