← Back to jobs
XXempla

Site Reliability Engineer (SRE) - Executive

Xempla

Kolkata
Full-Time
0-3 Years experience

Description

About Xempla :
Xempla is pioneering the future of facility management through the Autonomous Maintenance Operating Centre (AMOC) — a next-generation platform that thinks, plans, and acts: so you don't have to! Our mission is to redefine how facilities operate by driving reliability, performance, and efficiency through autonomous systems. As a bootstrapped, high-growth company, we're lean, decisive, and deeply committed to measurable impact over noise.

The Opportunity:
As an Site Reliability Engineer (SRE), you will help build the reliability foundation of an AI-native product. As our platform scales across APIs, background jobs, agentic pipelines, and real-time workflows, reliability becomes critical to product integrity and customer trust. This role is about shifting from reactive firefighting to proactive reliability engineering. You will design structured observability, investigate system behaviour across distributed components, and prevent issues before they impact customers.
AI tools will be part of your daily workflow accelerating log analysis, anomaly detection, root cause investigation, and automation. AI-assisted operations is how we work. If you want to move beyond ticket-based DevOps and build real reliability discipline, this is that opportunity.

What You'll Do:
• Strengthen observability across APIs, background jobs, cron services, and pipelines
• Work with metrics, logs, and distributed tracing to improve system visibility
• Investigate slow APIs and intermittent failures using structured root cause analysis
• Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR)
• Identify performance bottlenecks before scaling infrastructure
• Embrace AI tools to accelerate log analysis, debugging, and incident triage
• Automate monitoring and reduce manual operational toil

Who You Are:
• 0.5–2 years in SRE, DevOps, or Platform Engineering
• Hands-on exposure to observability tools (Grafana, Datadog, New Relic, or similar)
• Familiarity with logging and tracing systems (ELK stack, OpenTelemetry, or similar)
• Basic understanding of cloud infrastructure (AWS, GCP, or Azure)
• Scripting or automation exposure (Python, Bash, or similar)
• Understanding of background job processing and distributed systems is a plus

Ready to build the reliability and observability backbone of an AI-native platform?

Requirements

  • Cloud infrastructure
  • Background job processing
  • grafana
  • datadog
  • pythonBash
  • New Relic
  • Distributed Systems
  • debugging
  • GCP
  • Observability across APIs, backgrou...
  • Metrics
  • logs
  • distributed tracing
  • Structured root cause analysis
  • Mean Time to Detect (MTTD)
  • Mean Time to Resolve (MTTR)
  • Performance bottleneck identificati...
  • AI tools for log analysis
  • incident triage
  • Monitoring automation
  • ELK stack
  • OpenTelemetry
  • AWS
  • amazon web services aws
  • Azure
  • Logging systems
  • Tracing systems

About Xempla

-

Industry: no-mentionEmployees: 50+Website