Apply now
Apply for Job
Senior Observability Engineer
Date:
5 Oct 2026
Location:
SG
Company:
StarHub Ltd
Job Description
Responsibilities:
- Administer the log and search platform (currently Splunk Cloud): indexes, data onboarding, roles and access, apps and knowledge objects, retention, and search and dashboard performance. Monitor ingest latency, skipped searches and licence consumption, and tune to keep cost under control.
- Administer the APM, RUM and infrastructure monitoring platform (currently Splunk Observability Cloud): teams, tokens, integrations, detectors, alert routing and dashboards. Manage metric cardinality and usage against entitlement.
- Roll out telemetry to all IS applications: OpenTelemetry Collector (currently the Splunk distribution) on AWS EKS and EC2, and APM instrumentation for Java, Go, .NET, Node.js and Python. Keep instrumentation vendor-neutral where possible so the backend can change without re-instrumenting. Maintain onboarding guides and report coverage per application.
- Set up RUM for customer-facing web and mobile frontends and link browser sessions to backend traces, logs and infrastructure metrics, so an incident can be followed from the user to the data store.
- Build and maintain observability as code (Terraform or equivalent) in Git with peer review and CI/CD, so detectors, dashboards and configuration are versioned, repeatable and free of manual drift.
- Plan and run platform and agent upgrades with vendors: define the test plan, validate in non-production, sign off before release, and manage support cases and escalations. Run structured evaluations and proofs of concept when StarHub considers a new or replacement tool.
- Design, build and operate AI and agent-assisted observability workflows, such as alert noise reduction, automated triage, root-cause summaries and natural-language querying, with human review and guardrails.
- Apply controls aligned with MAS-TRM and CSA best practices to telemetry: retention, access control, audit logging, and masking of PII and sensitive data before ingest.
- Perform User Access Reviews (UAR) on the observability platforms at least twice a year, covering user accounts, roles and tokens, and follow up on removals and exceptions. Participate in company-wide audits when required, providing evidence and closing findings on time.
Qualifications
Minimum Profile/ Track Record:
Desired Background
- Experience in medium-to-large technology, telecommunications or financial services organizations with complex hybrid-cloud environments and many application teams; has owned an enterprise observability platform, not only used one.
Seniority, Skills, Certifications (must-haves)
- Bachelor's degree in Computer Science, Information Technology, Engineering or a related field.
- 5+ years of relevant experience in observability, monitoring, SRE or platform engineering, including hands-on administration, optimization and health monitoring of an enterprise observability or log analytics platform. Splunk (Cloud or Enterprise) is strongly preferred.
- Hands-on experience with an enterprise APM/RUM/infrastructure monitoring platform across RUM, APM, infrastructure and data-storage monitoring. Splunk Observability Cloud is preferred; Datadog, Dynatrace, New Relic, Elastic or similar is acceptable. Clear understanding of distributed tracing from frontend to backend.
- Experience installing and operating OpenTelemetry-based telemetry on AWS (EKS, EC2) for Java, Go, .NET, Node.js and Python applications.
- Experience managing observability configuration as code (Terraform or equivalent, Git, CI/CD).
- Prior AI delivery experience in observability or IT operations: has taken an AI or agent-based capability into production use by operations or engineering teams.
- Familiarity with security and compliance standards (MAS-TRM, CSA), including telemetry data governance.
- Certifications are a plus: Splunk Core Certified Admin or Power User, Splunk Observability Cloud certification, or equivalent vendor certifications on another observability platform.
Ideal track record #1
-
- Owned an enterprise observability or log analytics platform at scale (Splunk preferred) as administrator and platform owner, improving search performance, stability or ingest cost, and ran vendor-led upgrades with test and validation plans and no unplanned outage.
Ideal track record #2
-
- Rolled out APM and infrastructure telemetry (OpenTelemetry) across a large application estate with RUM-to-backend tracing in use during real incidents, measurably reducing time to detect or resolve.
Ideal track record #3
-
- Delivered an AI or agent-based capability in production for monitoring, triage or root-cause analysis, with measurable impact such as less alert noise or faster resolution.
Apply now