Job Description

  • Design, build, and maintain scalable ETL/ELT pipelines to ingest structured and unstructured data from SAP S/4HANA, Oracle, Salesforce, and third-party APIs into enterprise data lakes/lakehouses.
  • Develop and optimize data products (medallion architecture) on platforms like Databricks
  • Implement data quality, lineage, and governance controls using tools such as Purview, Alation, or Unity Catalog.
  • Build, deploy, and monitor ML models (classification, regression, forecasting, anomaly detection) using Python, PySpark, scikit-learn, TensorFlow, or PyTorch.
  • Develop GenAI and LLM-powered solutions — including RAG (Retrieval-Augmented Generation) pipelines, embeddings, vector databases (Pinecone, Azure AI Search), and prompt engineering for Copilot/Azure OpenAI/Bedrock.
  • Operationalize models through MLOps practices — CI/CD for ML, model versioning, drift monitoring, and automated retraining (MLflow, Azure ML, SageMaker, Vertex AI)
  • Deploy and manage data/AI workloads on Azure, AWS, or GCP, leveraging services like ADF, Databricks, Glue, Lambda, Functions, and Kubernetes.