Key Responsibilities
1. Cloud Infrastructure & Networking
• Kubernetes Platforms: Deploy and maintain AWS EKS and Azure AKS multi-region clusters.
• Traffic & Security: Configure ingress (Traefik, Nginx), TLS lifecycle (Cert Manager), and secrets management (AWS Secrets Manager, Azure Key Vault).
• Network Topologies: Design VPC/VNets, route tables, Hub-Spoke networks, Azure Bastion, and secure SFTP infrastructure.
2. Operations & Automation
• Smart Scaling: Implement Karpenter for cost-efficient EKS node provisioning and KEDA for event-driven workload scaling.
• CI/CD & Fleet Management: Build GitHub Actions pipelines and manage self-hosted GitHub Actions Runner Controller (ARC) fleets on Kubernetes.
• Tooling & Helm: Maintain Helm charts and develop internal platform automation using Golang and Python.
3. Data & Messaging Services
• Storage & Databases: Manage AWS S3, Azure Blob Storage (lifecycle, replication), and AWS RDS PostgreSQL clusters (tuning, backups, failover).
• Queueing Systems: Operate AWS SQS and Azure Service Bus, managing Dead Letter Queues (DLQs) and consumer scaling.
4. Observability & SRE
• Telemetry Stack: Instrument services with OpenTelemetry, Prometheus (PromQL), Loki, Tempo, and Pyroscope.
• Reliability & Incidents: Manage Grafana dashboards, alerting, SLO/SLI tracking, and on-call workflows via Grafana Cloud IRM.
• SRE Practice: Lead blameless post-mortems, enforce error budgets, and improve production resilience.
Core Requirements
• Core Tech: Hands-on EKS/AKS, Helm, GitHub Actions, OpenTelemetry, and Grafana ecosystem expertise.
• Development: Hands on experience scripting/programming Language - Golang and Python.
• Fundamentals: Deep understanding of Linux OS internals, network security, and database administration.
• Mindset: Strong problem-solving, incident management skills, and an SRE-focused reliability drive.


.webp)










