Save Job Back to Search Job Description Summary Similar JobsWork on large-scale Kubernetes platforms serving millions of users.Drive DevOps innovation using AI-powered automation and SRE practices.About Our ClientOur client is a global technology company operating large-scale digital entertainment and consumer services. They are committed to building highly reliable, scalable platforms while driving innovation through DevOps, Site Reliability Engineering, and AI-assisted operations.Job DescriptionOperate and maintain Kubernetes-based production infrastructure for large-scale B2C applications.Monitor system health, respond to alerts, troubleshoot incidents, and ensure platform stability.Support development teams by investigating infrastructure-related issues and providing technical guidance.Manage lifecycle activities for operating systems and middleware, including upgrades and end-of-life planning.Design, execute, and evaluate non-functional testing, including load, stress, and failure testing.Identify system bottlenecks and implement performance and reliability improvements.Build, maintain, and optimize CI/CD pipelines using Jenkins.Develop and maintain operational documentation, procedures, and runbooks.Collaborate with engineering and business stakeholders to recommend infrastructure and operational improvements.Drive operational efficiency by leveraging AI tools such as GitHub Copilot, Claude Code, Gemini, or similar technologies.Promote DevOps, automation, and Site Reliability Engineering best practices across engineering teams.The Successful ApplicantProven experience of building and administering Linux server environments.Proven experience operating and monitoring infrastructure supporting large-scale B2C applications.Hands-on experience with Kubernetes and/or Docker in production environments.Strong understanding of Linux administration, middleware configuration, user management, permissions, and troubleshooting.Knowledge of networking fundamentals including TCP/IP, HTTP, and DNS.Basic knowledge of relational databases, including SQL operations, backup, and recovery.Experience building or maintaining CI/CD pipelines, preferably with Jenkins.Strong analytical and problem-solving skills with a proactive approach to identifying and resolving operational issues.Interest in applying AI-powered engineering tools such as GitHub Copilot or Claude Code to improve productivity and operational workflows.Excellent communication skills with the ability to collaborate effectively across development, operations, and business teams.What's on OfferOpportunity to work on large-scale, customer-facing platforms with modern cloud-based technologies.Competitive salary and comprehensive employee benefits.Opportunity to work in a hybrid setup, combining flexibility and collaboration.Chance to be part of a global environment with diverse teams.ContactAahan RawatQuote job refJN-052026-7026837Phone number+81 3 6832 8627Job summaryFunctionITSpecialisationInfrastructureSpecialisationTechnology & TelecomsLocationTokyoJob TypeTemporaryConsultant nameAahan RawatConsultant phone+81 3 6832 8627Job ReferenceJN-052026-7026837Company TypeForeign Multinational