Client Overview
Our client is an Azerbaijani telecommunications company, the largest mobile network operator in Azerbaijan. The main products are: Fixed telephony, Mobile telephony, Internet services, Wireless broadband, and Value-added services.
Project Objectives
The primary goal is to accelerate the client’s Data & AI initiatives via a secure, hybrid cloud foundation on AWS while systematically modernizing the IT estate as part of the cloud migration.
Key Project Objectives include:
* Cloud Foundation & Landing Zone: Deploy target hybrid network architectures, establishing a secure Landing Zone and hybrid Data/AI platforms on AWS.
* Security, Compliance & Governance: Operationalize on-prem tokenization (achieving zero raw PII in the cloud), resolve policy blockers to include AWS in the ISMS, and establish a Cloud Center of Excellence (CCoE) to govern Cloud adoption.
* AI Chatbot & Voicebot Design & Implementation: Develop and operationalize a flagship Customer Care Chatbot and Voicebot as the first hybrid-setup consumer.
Responsibilities
- Build, operationalize, and automate end-to-end MLOps pipelines using Amazon SageMaker Pipelines and MLflow for experiment tracking, model versioning, and registry lifecycle management.
- Design, deploy, and manage production SageMaker inference endpoints (real-time, serverless, and batch) and Amazon Bedrock API integrations for LLM/SLM deployment with cost controls and latency optimization (Bedrock API Gatekeeper).
- Implement AgentOps / LLMOps frameworks (AgentCore, Bedrock Guardrails, Promptfoo) to manage multi-agent orchestration, prompt evaluation, safety guardrails, and RAG retrieval pipelines.
- Operationalize real-time STT / TTS (Speech-to-Text / Text-to-Speech) voicebot pipelines and low-latency speech inference on hybrid/cloud GPU node pools for the flagship Customer Care Voicebot.
- Optimize specialized GPU node pools (NVIDIA A100/L40S / EC2 GPU instance types) for Azerbaijani SLM/LLM model training, fine-tuning, and scalable inference workloads.
- Establish automated CI/CD for Machine Learning using GitLab CI/CD pipelines and Infrastructure-as-Code (Terraform or AWS CDK) to enforce security-gated MLOps promotion workflows (from SageMaker Canvas/Sandbox to production).
- Integrate data de-identification, Format Preserving Encryption (FPE), and tokenization wrappers into ML data pipelines to ensure zero raw PII enters AWS cloud environments during model training and inference.
- Set up telemetry, performance monitoring, model drift detection, and cost anomaly alerting for AI/ML workloads using Amazon CloudWatch, Splunk, and FinOps spend control frameworks.
- Collaborate with Data Engineering, AI Architects, and Cloud Teams to integrate vector storage/retrieval (RAG), Apache Spark/EMR-on-EKS runtimes, and local tokenization databases.
- Author technical MLOps runbooks, model deployment procedures, governance documentation, and disaster recovery playbooks.