Site Reliability Engineer (AWS & ML) - Remote
Lewis Personnel Management Muntinlupa Full-time
Location: Metro Manila, Philippines | Remote
Employment type: Full-time | Day Shift
About the Role
This position operates within Japan's leading restaurant reservation and hospitality management platform, which currently powers operations globally. We run a robust and highly fault-tolerant enterprise infrastructure built entirely on Amazon Web Services (AWS) utilizing modern technologies like Terraform, Kubernetes (EKS), Helm, and automated CI/CD pipelines.
In this role, your primary focus will be traditional SRE responsibilities—owning, maintaining, and evolving our 24/7 production infrastructure running on Kubernetes. However, you will also play a key role in our machine learning initiatives, bringing production-level reliability engineering discipline to our MLOps pipelines, model deployment mechanisms, and the underlying cloud infrastructure that powers our intelligent product features.We operate in a highly collaborative, remote-first, and document-driven culture where team members are expected to take complete ownership of their domains.
Key Responsibilities
SRE & Cloud Infrastructure (Primary Focus)- 24/7 Production Stability: Maintain a highly available, scalable, and secure production environment on Kubernetes following standard SRE principles.
- DevOps Methodologies: Implement modern DevOps practices to consistently improve automated workflows and internal engineering quality of life.
- Infrastructure Management: Manage and scale core AWS cloud infrastructure services, specifically covering EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
- Infrastructure as Code (IaC): Build and maintain highly reusable infrastructure modules and CI/CD automation pipelines utilizing Terraform, Helm, and ArgoCD.
- Observability & Incident Response: Proactively monitor systems, configure alerts, and lead high-efficiency incident response and technical postmortem processes.
- ML Infrastructure Operations: Apply rigid SRE discipline to our machine learning stack—ensuring model serving, training pipelines, and data layers remain observable, stable, and highly performant.
- Model Deployment Pipelines: Support and continuously optimize ML model deployment pipelines and automated MLOps lifecycles.
- Collaborative Engineering: Partner closely with data science teams and core product tracks to operationalize machine learning models cleanly at an enterprise scale.
- EKS Optimization for ML: Deploy and optimize hardware-efficient infrastructure configurations tailored specifically for ML workloads on AWS Kubernetes.
- Cloud Experience: At least 2+ years of professional, hands-on experience managing Amazon Web Services (AWS) environments, with strong production exposure to EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
- EKS Expertise: Heavy, practical experience directly configuring and running workloads on AWS Elastic Kubernetes Service (EKS).
- SRE Leadership: Proven experience in software engineering following DevOps or SRE methodologies, including at least 1 year of experience acting as a technical lead.
- Programming Skills: Strong, active coding proficiency in at least one of the following languages: Python, Ruby, Elixir, Go, JavaScript, or Rust.
- Systems Foundations: Excellent understanding of containerization (Docker) and basic hypervisor and virtualization fundamentals.
- Automation Tooling: Solid knowledge of configuration management using YAML or Bash scripting; direct experience utilizing Helm and Terraform is highly preferred.
- ML/MLOps Familiarity: Conceptual familiarity with modern machine learning workflows, alongside hands-on Python experience using ML-adjacent tooling (e.g., model deployment pipelines or inference serving).
RH-TT
Ben Edictio SearchPasay, 19 km from Muntinlupa
and architecture discussions.
• Provide on call support and incident response.
Qualifications:
• Bachelor's degree in Information Technology/Computer Engineering or any related course
• With 3-5 years of experience in DevOps or Site Reliability Engineering...
Taguig, 15 km from Muntinlupa
Develop custom software solutions to design, code, and enhance components across systems or applications. Use modern frameworks and agile practices to deliver scalable, high-performing solutions tailored to specific business needs.
Allegro MicroSystemsMuntinlupa
solutions meet world-class quality, reliability, and cost targets.
What You Will Do
• Work with AMPI Engineering and Planning, and AML Engineering and PMs to coordinate Probe Test, Assembly build, Qual Test, and Final Test activities at AMPI based...