Save Job Back to Search Job Description Summary Similar JobsBuild reliable platforms powering millions of daily transactions.Lead SRE initiatives using cloud-based and AI technologies.About Our ClientOur client is a leading global technology company operating high-volume digital platforms that process millions of user interactions and transactions every day. They are investing in modern cloud-based infrastructure, automation, and Site Reliability Engineering practices to deliver highly reliable customer experiences.Job DescriptionDefine and continuously improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs) in collaboration with engineering teams.Design highly available and resilient systems with redundancy, failover, disaster recovery, and backup strategies.Perform infrastructure capacity planning to balance scalability, performance, and cost efficiency.Design, build, and maintain Kubernetes-based infrastructure and CI/CD pipelines using Infrastructure as Code.Manage production operations, including releases, configuration changes, maintenance, and operational documentation.Participate in incident response, on-call rotations, root cause analysis, and blameless postmortems.Implement security best practices including vulnerability management, patching, access control, and secret management.Build automation tools and reduce operational overhead through scripting and AI-powered solutions.Collaborate with development teams to improve deployment processes, platform reliability, and developer productivity.Drive continuous improvements in observability, automation, and operational excellence.The Successful ApplicantBachelor's degree in Computer Science, Software Engineering, or equivalent practical experience.8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.5+ years of hands-on experience designing and operating Kubernetes environments.Strong experience managing Linux-based production infrastructure.Experience designing scalable infrastructure for large-scale web or mobile applications.Solid understanding of networking technologies including DNS, CDN, load balancing, TLS, and firewalls.Experience designing and maintaining CI/CD pipelines using Jenkins, GitHub Actions, CircleCI, or similar tools.Proficiency with Infrastructure as Code tools such as Terraform or Pulumi.Hands-on experience with public cloud platforms including AWS, GCP, or Azure.Strong scripting and automation skills using Python, Shell, or similar languages.Experience using AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar.Strong knowledge of systems architecture, networking, and information security.Excellent communication, documentation, and stakeholder management skills.A proactive mindset with a passion for improving systems and engineering practices.What's on OfferOpportunity to build and operate mission-critical services at massive scale.Opportunity to work in a hybrid environment.Competitive salary and comprehensive employee benefits.Excellent opportunities for technical leadership and long-term career growth within a global technology organization.ContactAahan RawatQuote job refJN-072026-7058532Phone number+81 3 6832 8627Job summaryFunctionITSpecialisationSystems AdministrationSpecialisationTechnology & TelecomsLocationTokyoJob TypeTemporaryConsultant nameAahan RawatConsultant phone+81 3 6832 8627Job ReferenceJN-072026-7058532Company TypeForeign Multinational