EPAM Systems
Senior Data Engineer
1仕事内容
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential. We are seeking a Senior Data Engineer to lead the design, development, and maintenance of data and ML pipelines on the Domino Data Lab platform. This role focuses on the data engineering and MLOps side of Domino, building reliable data pipelines, managing model lifecycle workflows, and ensuring the platform's data and compute infrastructure runs efficiently and securely. This is not a front-end or application development role. Responsibilities Design, build, and maintain robust data pipelines that support analytical and machine learning workloadsManage the end-to-end lifecycle of data workflows, from ingestion through transformation and deliveryOversee compute infrastructure to ensure efficient, secure, and reliable platform operationsTroubleshoot and resolve issues affecting pipeline performance and data qualityCollaborate with data scientists and other engineers to support their infrastructure and tooling needsEstablish and promote best practices for platform usage, pipeline architecture, and data workflow designAutomate testing and deployment processes to improve reliability and reduce manual effortMonitor pipeline health and proactively address bottlenecks or failuresContribute to the ongoing improvement of internal tools and processes supporting data operationsDocument technical designs, workflows, and configurations to support knowledge sharing across the team Requirements A minimum of 3 years of relevant experienceExtensive hands-on experience with the Domino Data Lab platform, including Data Sources and Connectors, Datasets, Environments, Projects, Jobs, and Flows, with the ability to architect and troubleshoot end-to-end data pipelines and guide best practices for platform usageExpert-level proficiency in Python as the primary language for data engineering and pipeline developmentStrong command of SQL for data extraction, transformation, and optimization across relational and warehouse systemsComfortable working with R and Bash across the broader data science toolchain and for automation scriptingDemonstrated experience designing, building, and maintaining ETL/ELT pipelines, including data ingestion, transformation, validation, and orchestrationHands-on experience with Kubernetes and managed solutions such as EKS, AKS, or GKE, with the ability to deploy, debug, and tune cluster workloads running data and ML jobsExperience building, optimizing, and troubleshooting container images for data and ML workloads using DockerExperience building and maintaining CI/CD pipelines using tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps to automate testing and deployment of data pipelines and ML workflowsWorking knowledge of at least one major cloud provider, such as AWS, Azure, or GCP, with the ability to reason about data architecture, cost, and security trade-offsExcellent English proficiency (B2 level or higher) Nice to have Experience with Domino Nexus or hybrid/multi-cloud compute orchestrationExperience delivering ML workflows covering training, deployment, monitoring, and retraining, with an understanding of reproducibility and versioningPractical experience with GenAI, LLMs, or agentic frameworks, including retrieval-augmented generation (RAG), from a data pipeline perspectiveFamiliarity with orchestration tools such as MLflow, Kubeflow, Airflow, or Domino FlowsKnowledge of model governance, compliance automation, or audit lo
追加情報
今すぐ応募
企業へ直接応募