
Hojat Gazestani
Senior SRE & MLOps Engineer
SRE/MLOps engineer who understands systems deeply and is now applying AI to that knowledge
Contact Information
Location
Available for remote work worldwide
Direct Contact
Call or WhatsApp
Primary contact method
About Me
Senior SRE / AI Infrastructure Engineer with 10+ years of hands-on experience building, operating, and automating reliable infrastructure and distributed systems. Strong background in Linux, networking, security, Kubernetes, cloud infrastructure, CI/CD, observability, and Python, with practical experience across AWS, Ansible, GitOps, Prometheus/VictoriaMetrics, Elasticsearch, and developer platforms.
Experienced in connecting knowledge across infrastructure layers—from Linux and networking to Kubernetes, security, monitoring, and automation—to troubleshoot complex production problems and build reliable systems. Recently focused on applying AI, LLMs, RAG, agentic systems, LangGraph, and MCP to infrastructure and SRE workflows, including incident analysis, observability, troubleshooting, infrastructure automation, and developer productivity.
Passionate about technology and continuous learning, with a strong ability to quickly learn new tools, systems, and technologies and turn them into practical solutions. Motivated by understanding how systems work end-to-end, automating repetitive operational work, and using AI to make infrastructure operations more intelligent, efficient, and reliable.
Skills
Cloud & Infrastructure
Programming
CI/CD & GitOps
Web & Application
Monitoring & Observability
AI & LLM
Networking & Security
Message Brokers & Databases
Latest Blogs
View All BlogsLinux Sudoers File Configuration: A Practical Guide with Examples
How to configure the /etc/sudoers file. Explains syntax, aliases, and best practices for granting sudo access without a password.
Read moreUnderstanding Linux Special Permissions: SUID, SGID, and the Sticky Bit
A clear guide to Linux's special permission bits - SUID, SGID, and Sticky Bit - and how to use them for system security and collaboration.
Read moreUnderstanding Linux Process UIDs: Real, Effective, Saved, and Filesystem
Linux user IDs (UIDs) and how the kernel uses Real, Effective, Saved, and Filesystem UIDs to manage permissions and security for processes like Nginx and passwd.
Read moreProfessional Experience
Cloud Engineer | DevOps | Kubernetes | AWS (Contract)
May 2023 – Present- Developed scalable services on AWS using EC2 Auto Scaling, ELB, S3, and DynamoDB
- Engineered AWS EKS Kubernetes clusters with Terraform, reducing deployment time by 60%
- Optimized Kubernetes deployments using Kubespray and Bash scripts, achieving 80% faster application launches
- Created HAProxy training materials (GitHub, YouTube) boosting content delivery by 50%
- Implemented GitLab CI/CD with Kustomize and Flux, reducing delivery time by 80%
OpenStack DevOps Engineer
Sep 2021 – Dec 2022- Architected OpenStack and Ceph infrastructure, improving network management by 40%
- Transformed single-node Redis to cluster, achieving 100% scalability and 70% reliability boost
- Redesigned RabbitMQ into cluster configuration, improving scalability and stability by 80%
DevOps Engineer
Apr 2020 - May 2021- Optimized Python unit tests in CI/CD, improving code management by 100%
- Automated Netbox-VMware vCenter synchronization with Python
- Implemented monitoring stack (Zabbix, Grafana, Prometheus, cAdvisor)
- Established SLAs for 50+ web services in Zabbix
Systems Engineer
Jul 2016 – Apr 2020Didehbannet
Tehran, Iran- Designed and secured new server room infrastructure
- Implemented Kubernetes/Docker storage optimization saving high-cost storage space
- Redesigned core network security (Juniper, FortiGate) improving security by 80%
- Automated network management with Python scripts
- Configured EMC Unity and HP SAN storage for VMware HA clusters
SELECTED PROJECTS
Agentic Kubernetes Root Cause Analysis Platform
GitHub repository- Built an agentic Kubernetes Root Cause Analysis platform that investigates alerts using live metrics, logs, Kubernetes state, and historical incident data.
- Designed a LangGraph-based autonomous workflow for planning investigations, collecting evidence, verification, and RCA generation.
- Integrated infrastructure and observability capabilities through MCP tools, including Kubernetes, VictoriaMetrics, and Elasticsearch.
- Implemented evidence-driven RCA to distinguish observed facts from inferred causes and reduce unsupported conclusions.
- Added incident-history retrieval using RAG to correlate current failures with previous incidents.
- Designed structured RCA reports containing affected workloads, pod/node status, events, metrics, logs, and supporting evidence.
Agentic Job Search / Career Automation
GitHub repository- Built an agentic job-search system that evaluates engineering positions against a structured candidate profile.
- Implemented Planner → Verifier → Scoring workflow using LangGraph.
- Integrated web search and local OpenAI-compatible LLM infrastructure for job discovery and evaluation.
- Designed candidate/job matching around Kubernetes, SRE, DevOps, cloud, AI infrastructure, and agentic engineering requirements.
Education
Bachelor's Degree in Information Security
University of Applied Science and Technology (UAST)
Associate Degree in Software Development
Chamran College
Languages
Online Content
Certifications
AWS Cloud Solutions Architect Specialization
Coursera
DevOps on AWS Specialization
Coursera
Python for Everybody Specialization
Coursera
Google Cloud Fundamentals: Core Infrastructure
Coursera
Container & Kubernetes Essentials V2
Coursera
Introduction to Containers with Docker, Kubernetes & OpenShift
Coursera
AWS CloudFront: Serve content from multiple S3 buckets
Coursera
Routing & Switching (CCNA, CCNP)
Cisco
Data Center (CCNA, CCNP)
Cisco
Mikrotik (MTCNA, MTCRE, MTCWE)
Mikrotik
Microsoft (MCSA)
Microsoft
VMware (VCP, NSX, SRM)
VMware
PWK (Penetration Testing with Kali Linux)
Offensive Security
ISMS (Information Security Management System)
ISO/IEC 27001
EMC Storage Certifications
Dell EMC