Site Reliability Engineer
- Build and improve platform reliability, resilience and operational excellence.
- Own and enhance Terraform-based infrastructure provisioning and management.
- Manage and support Kubernetes (Amazon EKS) environments in production.
We're hiring a Site Reliability Engineer to join a high-performing AI Platform team responsible for building and operating the core infrastructure, runtime and operational foundations that enable AI adoption across the organisation.
This role is ideal for someone who enjoys working across Cloud Infrastructure, Platform Engineering, Site Reliability Engineering, and DevOps disciplines while solving complex operational challenges in a modern cloud-native environment. You will play a key role in ensuring the AI Platform remains scalable, resilient and observable while supporting AI-powered services and applications used across the business.
What You'll Do
- Build and improve platform reliability, resilience and operational excellence.
- Own and enhance Terraform-based infrastructure provisioning and management.
- Manage and support Kubernetes (Amazon EKS) environments in production.
- Develop and maintain CI/CD pipelines supporting deployment and release activities.
- Drive observability, monitoring and alerting best practices across the platform.
- Support incident investigation, root cause analysis and platform troubleshooting.
- Own and coordinate production releases from UAT through to Production deployment.
- Automate operational tasks through scripting and tooling.
- Support AI and agent-based workloads running on the platform ecosystem.
- 3-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps or Cloud Infrastructure environments.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- Strong production experience with Kubernetes, ideally Amazon EKS.
- Strong experience supporting cloud-native environments on AWS.
- Experience building, maintaining or owning CI/CD pipelines.
- Experience owning production releases, change processes and deployment activities.
- Strong troubleshooting and incident management experience.
- Experience with observability and monitoring tools such as Datadog.
- Linux administration and troubleshooting experience.
- Python or scripting experience used for automation and operational tooling.
- Strong communication and stakeholder management skills.
- AI / LLM platform experience.
- Agentic systems experience.
- Helm.
- OpenSearch.
- Databricks.
- AI observability and traceability concepts.
- Enterprise platform operations.
- Chaos engineering or resilience testing exposure.
We regret to inform that only shortlisted candidates will be notified .
EA Registration No.: WONG LIN, RACHEL, R25158204
Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA License No. 10C4544
Job ID a4VQ8000008A3xhMAC
We’re partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a committed team working with over 6,000 clients across North America, Europe and Asia Pacific.
As an industry leader in talent services, we work with progressive leaders to drive change. That’s the power of true partnership.
TEKsystems is an Allegis Group company.
More Jobs From TEKsystems (Allegis Group Singapore Pte Ltd)
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
TEKsystems (Allegis Group Singapore Pte Ltd)
Singapore
Boost your career
Find thousands of job opportunities by signing up to eFinancialCareers today.More Jobs Like This
Oxford Knight
Singapore