SRE Engineering Manager
Ascendion
Description
About the Role"
"Site Reliability Engineer (SRE)"
"Responsibilities"
-
Design, implement, and manage highly available and scalable infrastructure.
-
Automate operational tasks using scripting and Infrastructure as Code (IaC).
-
Monitor system performance and ensure reliability using observability tools.
-
Define and manage SLIs, SLOs, and SLAs.
-
Build and maintain CI/CD pipelines for faster and reliable deployments.
-
Perform incident management, root cause analysis (RCA), and postmortems.
-
Collaborate with development teams to improve application reliability and performance.
-
Manage containerized environments using Docker and Kubernetes.
-
Ensure system security, compliance, and best practices (DevSecOps).
-
Continuously improve system efficiency and reduce operational toil.
"Required Skills"
-
Bachelor's degree in Computer Science, Engineering, or related field.
-
3–8 years of experience in SRE / DevOps / Cloud Engineering roles.
-
Strong knowledge of Linux/Unix systems.
-
Proficiency in Python, Bash, or Shell scripting.
-
Experience with cloud platforms (AWS / Azure / GCP).
-
Hands-on experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI).
-
Expertise in containerization (Docker) and orchestration (Kubernetes).
-
Experience with Infrastructure as Code (Terraform, CloudFormation, Ansible).
-
Familiarity with monitoring & logging tools (Prometheus, Grafana, ELK, Datadog, Dynatrace).
-
Understanding of networking concepts and security practices.
"Desirable Skills"
-
Experience with microservices architecture.
-
Knowledge of OpenTelemetry and distributed tracing.
-
Experience in performance tuning and capacity planning.
-
Familiarity with Agile/Scrum methodologies.
-
Certification in AWS / Azure / Kubernetes is a plus.
"Education Qualification"
- Bachelor's degree in Computer Science, Engineering, or related field.
Ensure the Job Description has proper punctuations and all sentences end with a full stop.
About Ascendion
-
