Apply now

Apply for Job

Senior Observability Engineer

Date:  5 Oct 2026
Location: 

SG

Company:  StarHub Ltd

Job Description

Responsibilities: 

  1. Administer the log and search platform (currently Splunk Cloud): indexes, data onboarding, roles and access, apps and knowledge objects, retention, and search and dashboard performance. Monitor ingest latency, skipped searches and licence consumption, and tune to keep cost under control.
  2. Administer the APM, RUM and infrastructure monitoring platform (currently Splunk Observability Cloud): teams, tokens, integrations, detectors, alert routing and dashboards. Manage metric cardinality and usage against entitlement.
  3. Roll out telemetry to all IS applications: OpenTelemetry Collector (currently the Splunk distribution) on AWS EKS and EC2, and APM instrumentation for Java, Go, .NET, Node.js and Python. Keep instrumentation vendor-neutral where possible so the backend can change without re-instrumenting. Maintain onboarding guides and report coverage per application.
  4. Set up RUM for customer-facing web and mobile frontends and link browser sessions to backend traces, logs and infrastructure metrics, so an incident can be followed from the user to the data store.
  5. Build and maintain observability as code (Terraform or equivalent) in Git with peer review and CI/CD, so detectors, dashboards and configuration are versioned, repeatable and free of manual drift.
  6. Plan and run platform and agent upgrades with vendors: define the test plan, validate in non-production, sign off before release, and manage support cases and escalations. Run structured evaluations and proofs of concept when StarHub considers a new or replacement tool.
  7. Design, build and operate AI and agent-assisted observability workflows, such as alert noise reduction, automated triage, root-cause summaries and natural-language querying, with human review and guardrails.
  8. Apply controls aligned with MAS-TRM and CSA best practices to telemetry: retention, access control, audit logging, and masking of PII and sensitive data before ingest.
  9. Perform User Access Reviews (UAR) on the observability platforms at least twice a year, covering user accounts, roles and tokens, and follow up on removals and exceptions. Participate in company-wide audits when required, providing evidence and closing findings on time.

 

 

 

Qualifications

Minimum Profile/ Track Record:

Desired Background

  • Experience in medium-to-large technology, telecommunications or financial services organizations with complex hybrid-cloud environments and many application teams; has owned an enterprise observability platform, not only used one.

Seniority, Skills, Certifications (must-haves)

  • Bachelor's degree in Computer Science, Information Technology, Engineering or a related field.
  • 5+ years of relevant experience in observability, monitoring, SRE or platform engineering, including hands-on administration, optimization and health monitoring of an enterprise observability or log analytics platform. Splunk (Cloud or Enterprise) is strongly preferred.
  • Hands-on experience with an enterprise APM/RUM/infrastructure monitoring platform across RUM, APM, infrastructure and data-storage monitoring. Splunk Observability Cloud is preferred; Datadog, Dynatrace, New Relic, Elastic or similar is acceptable. Clear understanding of distributed tracing from frontend to backend.
  • Experience installing and operating OpenTelemetry-based telemetry on AWS (EKS, EC2) for Java, Go, .NET, Node.js and Python applications.
  • Experience managing observability configuration as code (Terraform or equivalent, Git, CI/CD).
  • Prior AI delivery experience in observability or IT operations: has taken an AI or agent-based capability into production use by operations or engineering teams.
  • Familiarity with security and compliance standards (MAS-TRM, CSA), including telemetry data governance.
  • Certifications are a plus: Splunk Core Certified Admin or Power User, Splunk Observability Cloud certification, or equivalent vendor certifications on another observability platform.

Ideal track record #1

    • Owned an enterprise observability or log analytics platform at scale (Splunk preferred) as administrator and platform owner, improving search performance, stability or ingest cost, and ran vendor-led upgrades with test and validation plans and no unplanned outage.

 

Ideal track record #2

    • Rolled out APM and infrastructure telemetry (OpenTelemetry) across a large application estate with RUM-to-backend tracing in use during real incidents, measurably reducing time to detect or resolve.

 

Ideal track record #3

    • Delivered an AI or agent-based capability in production for monitoring, triage or root-cause analysis, with measurable impact such as less alert noise or faster resolution.

To APPLY NOW, click on Skye!

Apply now

Apply for Job