Career Profile
Senior DevOps Engineer with 11+ years of experience in Cloud Platform Engineering, Kubernetes operations, CI/CD architecture, OpenStack, and large-scale infrastructure environments. Currently part of IBM Cloud VPC Platform Team, leading Kubernetes-based platform deployments, new region/mzone bring-up, CI/CD ownership (Jenkins & Tekton), automation engineering, and production-grade release management. Strong expertise in platform reliability, incident response, automation using Python, and multi-region infrastructure scaling.
Experience
- Led end-to-end regional expansions for IBM Cloud VPC (WDC, Madrid, Dallas DC), managing infrastructure readiness and zero-impact production cutovers.
- Architected Kubernetes control-planes and networking configurations for global-scale compute services, managing clusters with 1,000+ worker nodes.
- Reduced deployment windows by 40% by re-engineering CD pipelines to parallelize zone/zonelet rollouts across complex regional topologies.
- Modernized CI/CD infrastructure by using Terraform to migrate legacy Jenkins workflows to IBM OnePipeline (Tekton), automating toolchain creation.
- Developed Terraform-managed health checks for Next Gen Data Centers (NGDC), executing hourly validations of pod readiness and node status.
- Engineered GitOps regional promotion using K8s CronJobs and Bash to automate delta comparisons and PR creation between environments.
- Automated release validation via Tekton, implementing environment locking and automated testing for PR auto-merging across global regions.
- Decreased release cycle times by 1+ hour per PR by implementing label-driven “skip-pre-check” logic for back-to-back regional deployments.
- Developed Python-based security monitoring to audit HashiCorp Vault secrets, triggering Slack alerts 5 days prior to certificate expiration.
- Built Python health-check suites to automate real-time validation of pod status, node readiness, and control-plane stability.
- Optimized CI/CD workflows using Jenkins and Tekton, introducing label-based skipping to bypass validated zones during deployment reruns.
- Managed 24/7 high-availability operations, leading incident triage, RCA, and rapid recovery via PagerDuty for critical VPC services.
- Improved MTTR and platform stability by partnering with SRE teams to enhance monitoring coverage and proactive alerting thresholds.
- Collaborated with Networking and Security teams to ensure architectural compliance, service onboarding, and strict regulatory readiness.
- Scaled global compute capacity by automating node provisioning and edge-layer expansions to meet increasing workload demands.
- Standardized deployment orchestration across Pre-Integration, Staging, and Production, ensuring consistency in microservice delivery.
- Enhanced developer productivity by building self-service automation tools that reduced manual infrastructure overhead by 30%+.
- Implemented automated rollback strategies and validation gates to significantly reduce change failure rates in production environments.
- Successfully operationalized multiple cloud regions with production-grade reliability, contributing to the growth of IBM Cloud VPC.
- Provisioned and deployed Kubernetes clusters on bare-metal Dell R650 servers, managing low-level BIOS configurations, RAID setups, and firmware upgrades.
- Orchestrated application delivery using Helm, managing critical Kubernetes resources including ConfigMaps, Secrets, and persistent storage solutions.
- Implemented observability stacks by installing and configuring the ELK (Elasticsearch, Logstash, Kibana) stack within production Kubernetes clusters.
- Executed platform optimization and security through cluster hardening, dynamic scale-in/scale-out operations, and rigorous acceptance testing.
- Provisioned high-performance Kubernetes clusters on bare-metal Dell R650 servers, managing full-stack infrastructure including BIOS configuration, RAID setup, and firmware upgrades.
- Architected and managed OpenStack clusters using OOO methodology, integrating VNFM/NFVO components and scaling compute and storage within NFVI environments.
- Optimized platform operations by deploying ELK observability stacks, managing Helm-based application delivery, and executing rigorous cluster hardening and acceptance testing.
- Engineered VNFD packages for firewall-based telecom workloads using Ansible and Bash, streamlining automated VNF deployments.
- Developed modular, idempotent Ansible playbooks for end-to-end VM provisioning, security hardening, and complex OS-level configurations across global environments.
- Architected infrastructure-as-code (IaC) solutions using HOT (Heat Orchestration Templates) and TOSCA to automate provisioning within OpenStack environments.
- Optimized performance and security by building customized CentOS images and integrating automated network/firewall configuration for production-grade reliability.
- Led end-to-end validation and handover, executing functional testing and troubleshooting in staging to ensure seamless production cutovers and team knowledge transfer.
- Deployed RHOSP13 in production environments, successfully integrating bare-metal servers using OpenStack Ironic.
- Architected Heat templates for automated infrastructure orchestration and performed deep-stack analysis to resolve complex onsite deployment issues.
- Managed multi-customer OpenStack cloud environments, designing network subnet allocations and creating Heat templates for resource orchestration.
- Deployed Virtual Network Functions (VNFs) via Horizon and CLI, overseeing full VM lifecycle operations in high-availability environments.
- Installed and configured Linux systems tailored for telecom-specific environments, ensuring high-performance server stability.
- Managed DNS, AAA, and VO-WiFi services, performing end-to-end troubleshooting to maintain critical network service availability.
- Architected a scalable 4G SecGW (Security Gateway) solution on AWS, utilizing Terraform to provision complex multi-interface VPC environments for LTE traffic isolation.
- Automated the deployment of virtualized telecom workloads using AWS CodeDeploy, ensuring consistent configuration management across elastic EC2 fleets.
- Engineered dynamic scaling for LTE signalling traffic by implementing Auto Scaling Groups driven by custom network metrics, maintaining 99.99% availability during peak loads.
- Integrated automated health-checks and validation via Python scripts to monitor S1-interface readiness and IPsec tunnel stability post-deployment.