
Manu Balakrishnan Sreekumari
Site Reliability Engineer & Cloud Architect
Professional Summary
Site Reliability Engineer and Cloud Architect with 18 years of experience building and operating highly available infrastructure across AWS, Kubernetes, and hybrid VMware/cloud environments. Known for turning operational pain points into durable systems, from IAM governance models that scale across teams, to automation that has driven $500K+ in measurable savings and eliminated recurring toil. Comfortable owning problems end-to-end: designing the architecture, building the tooling, and mentoring engineers along the way. Recently expanded into AI/ML fundamentals (ZTM coursework, AWS AI Practitioner in progress) to bring that same systems thinking to AI-driven infrastructure.
Key Achievements
- 18 years of progressive experience in Site Reliability Engineering
- $200K+ annual savings through AWS SSM patching tool automation
- 30% cost reduction in development accounts through resource optimization
- 20% cost savings through EKS/ECS consolidation
- Managed infrastructure with 2000+ VMs across multiple cloud providers
- Reduced QA turnaround time from weeks to hours through containerization
- Successfully managed teams across multiple time zones (North America, Asia Pacific)
Certifications
AWS Certified Solutions Architect Associate
Validation Number: 4X7D1B1LEME4QN3P
AWS Certified AI Practitioner
Validation Number: In progress
Technical Skills
Infrastructure & Cloud
- AWS (EC2, ECS, EKS, Lambda, S3, RDS, CloudFront, Route53, SSM, Transfer Family)
- Multi-cloud environments (AWS, Azure, GCP)
- VMware vSphere, ESXi, vCloud, Lab Manager
- Nutanix HCI, AHV
- Infrastructure as Code (Terraform, CloudFormation)
- Managed hybrid environment with 2000+ Linux and Windows VMs
Container & Orchestration
- Kubernetes (K8s, EKS, AKS)
- Docker, Containerd, Amazon ECS
- Service Mesh (Istio)
- GitOps (ArgoCD)
- Helm Charts
Automation & DevOps
- CI/CD (GitLab, GitHub, Azure DevOps, TeamCity)
- Configuration Management (Ansible, Packer)
- Scripting (Python, Bash, basic Golang)
- Akeyless (secrets management)
- Terraform, Packer, Ansible
Observability & Monitoring
- Prometheus, Grafana
- Splunk, Honeycomb
- PagerDuty incident management
- Log analysis and root cause investigation
Databases
- Oracle, DB2, MySQL, MSSQL, PostgreSQL
- Database deployment and management
- Performance tuning
Systems
- Linux (extensive experience)
- Unix systems (HP-UX, AIX)
- Windows Server
- NetApp Storage Support
SRE Practices
- Incident response and management
- Toil reduction and automation
- Service level objectives (SLOs)
- Post-incident reviews and continuous improvement
AI / Machine Learning
- Completed the ZTM comprehensive AI/ML course
- Preparing to sit for the AWS Certified AI Practitioner certification
- Foundational exposure to MLOps
Professional Experience
Site Reliability Engineer
GE Healthcare
Toronto, ON | 05/2024 to Current
- Supported development and operational readiness for the GE Healthcare CareIntellect AI software platform, helping ensure reliability and continuity for the AI-based product
- Designed a layered IAM governance model to solve org-wide runner role drift, decomposing monolithic CI/CD roles into global, team, environment, and account-scoped policy layers, each version-controlled with its own ownership and review path
- Built an AWS Nuke-based decommissioning tool that identifies and removes unused cloud resources, cutting development account costs significantly
- Engineered an AWS Transfer Family solution enabling zero-downtime migration of terabytes of data
- Developed internal tooling to automate operations and eliminate recurring toil, freeing up team capacity for higher-value work
- Partnered with release teams to validate and verify deployments, with feedback loops that measurably improved deployment reliability
- Standardized validation processes through documentation, improving consistency across operations teams in multiple time zones
Site Reliability Engineer
KAR Global
Toronto, ON | 05/2022 to 05/2024
- Developed and maintained Azure DevOps YAML pipelines to build, test, and deploy applications across multiple environments
- Consolidated multiple applications to EKS/ECS, saving about 20% cost compared to the previous year
- Automated and improved system scalability using tools like Terraform
- Created an AWS SSM based patching tool for cloud and On-premise systems, saving the company $200k annually
- Responded to PagerDuty alerts and troubleshoot incidents, reviewed logs and alerts to investigate root causes and suggest improvements
- Maintained internally developed tool that assists teams in speeding up incident response
- Maintained and promoted adoption of Kubernetes using Helm Charts, EKS, Argo CD and Terraform
- Responded to security alerts on ORCA and took corrective measures
- Collaborated with cross-functional teams to develop, test, and deploy scalable software solutions
- Mentored co-op engineers, sharing knowledge of best practices for site reliability engineering methodologies
- Established a culture of continuous improvement within the team by encouraging feedback loops and iterative development processes
IT Software Engineer Principal 3
Progress Software Development Private Limited
Hyderabad, Telangana | 07/2009 to 04/2022
- Worked with development team to design, evaluate, automate and maintain on-premise and cloud development environments
- Developed tools and processes for automating CI/CD with TeamCity, focused on improving software delivery throughout the SDLC
- Managed VMware, Nutanix, and NetApp-based development infrastructure across North America and Asia Pacific
- Planned, designed, and coordinated the evaluation, installation, upgrade and integration of various software and systems
- Developed applications to provision Virtual Machines in VMware and Nutanix, contributing to datacenter consolidation from Hyderabad to Boston
- Migrated QA database and applications to Kubernetes-based containers, reducing turnaround time from a week to hours
- Led PoC deployment and maintenance of Rancher-based Kubernetes infrastructure
- Monitored system performance, reviewed logs to perform root cause analysis and made recommendations to optimize performance
- Participated in the Agile Scrum process
- Integrated Akeyless for centralized secrets management across CI/CD pipelines and Kubernetes as part of an initial PoC and integration testing
- Worked with all teams across the board and engaged with stakeholders like product owners and quality analysts
- Created end-user and administrator guide for self-service provisioning portal vCloud Director
- Worked with software vendors to evaluate new technologies and participated in all phases from PoC to Production
Education
Bachelor of Engineering - Electronics and Communication Engineering
Manonmaniam Sundaranar University, Tamil Nadu, India
Diploma - Electronics and Communication Engineering
State Board of Technical Education and Training, Tamil Nadu, India