EPAM Systems
Lead Data DevOps Engineer
1仕事内容
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential. We are looking for a Lead Data DevOps Engineer to join our team. We are building an Enterprise AI Gateway from the ground up. It serves as the single entry point every team in the company uses to access large language models. It is also how we maintain control over AI at scale: who can use which models, what it costs, and what gets logged. The goal is for it to be self-sustaining. Onboarding teams, agents, and MCP servers, along with managing keys, permissions, limits, and guardrails, should happen through automation rather than manual tickets. Very little of this exists yet, so you will influence it starting from the earliest design choices. The team is small and focused, meaning decisions move quickly and your contributions will be clearly visible. We are looking for someone who defaults to automation and who values the governance side of AI just as much as the models themselves. Responsibilities Architect and construct the foundational design of the Enterprise AI Gateway starting from scratchAutomate onboarding processes for teams, agents, and MCP servers, including the provisioning of keys, permissions, limits, and guardrailsBuild and maintain infrastructure as code to enable scalable, self-service platform capabilitiesSet up monitoring, logging, and cost-tracking systems to preserve visibility and oversight of AI usageEstablish guardrails and governance mechanisms that control which teams and models can access particular resourcesGuarantee the reliability, uptime, and performance of production systems that power the gatewayWork directly with platform users, resolving questions and troubleshooting errors as they ariseRefine automation on an ongoing basis to minimize manual work and reliance on ticket-based workflowsAssess and incorporate emerging GenAI and agentic AI patterns, frameworks, and protocols into the platformPlay a role in shaping major architectural and design decisions as the platform matures from its early stages Requirements At least 5 years of relevant experienceA minimum of one year of experience leading and managing teamsSolid Python skills for automation, extensions, and integrationsStrong SRE capabilities, backed by genuine experience maintaining healthy production systemsProficient with Git for version controlExperience working with Google Cloud PlatformFamiliarity with LLMOps practicesExperience using Terraform and Helm for infrastructure automationWorking knowledge of GenAI/Agentic AI concepts, including associated patterns, frameworks, and protocolsStrong communication skills, with the ability to clearly convey technical concepts, as you will engage daily with platform users, answering questions and resolving issuesExcellent English proficiency (B2 level or higher) Nice to have Hands-on experience with Google Vertex AI, especially endpoints for model serving and Model ArmorGCP experience involving services such as BigQuery, Cloud Run, and IAMExperience building AI agents, for instance using Google's Agent Development Kit (ADK)Experience with AWS BedrockExtensive Kubernetes experience, ideally with GKEFamiliarity with GroovyExperience with CI/CD pipelines using JenkinsProduction experience with an AI gateway, such as LiteLLM or EPAM DIAL, considered highly valuable We offer International projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid tim
追加情報
今すぐ応募
企業へ直接応募