Site Reliability Engineer (SRE) - Executive
Xempla
Description
About Xempla :
Xempla is pioneering the future of facility management through the Autonomous Maintenance Operating
Centre (AMOC) — a next-generation platform that thinks, plans, and acts: so you don't have to!
Our mission is to redefine how facilities operate by driving reliability, performance, and efficiency through
autonomous systems. As a bootstrapped, high-growth company, we're lean, decisive, and deeply committed
to measurable impact over noise.
The Opportunity:
As an Site Reliability Engineer (SRE), you will help build the reliability foundation of an AI-native product. As
our platform scales across APIs, background jobs, agentic pipelines, and real-time workflows, reliability
becomes critical to product integrity and customer trust. This role is about shifting from reactive firefighting
to proactive reliability engineering.
You will design structured observability, investigate system behaviour across distributed components, and
prevent issues before they impact customers.
AI tools will be part of your daily workflow accelerating log analysis, anomaly detection, root cause
investigation, and automation. AI-assisted operations is how we work. If you want to move beyond ticket-based
DevOps and build real reliability discipline, this is that opportunity.
What You'll Do:
• Strengthen observability across APIs, background jobs, cron services, and pipelines
• Work with metrics, logs, and distributed tracing to improve system visibility
• Investigate slow APIs and intermittent failures using structured root cause analysis
• Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR)
• Identify performance bottlenecks before scaling infrastructure
• Embrace AI tools to accelerate log analysis, debugging, and incident triage
• Automate monitoring and reduce manual operational toil
Who You Are:
• 0.5–2 years in SRE, DevOps, or Platform Engineering
• Hands-on exposure to observability tools (Grafana, Datadog, New Relic, or similar)
• Familiarity with logging and tracing systems (ELK stack, OpenTelemetry, or similar)
• Basic understanding of cloud infrastructure (AWS, GCP, or Azure)
• Scripting or automation exposure (Python, Bash, or similar)
• Understanding of background job processing and distributed systems is a plus
Ready to build the reliability and observability backbone of an AI-native platform?
Requirements
- Cloud infrastructure
- Background job processing
- grafana
- datadog
- pythonBash
- New Relic
- Distributed Systems
- debugging
- GCP
- Observability across APIs, backgrou...
- Metrics
- logs
- distributed tracing
- Structured root cause analysis
- Mean Time to Detect (MTTD)
- Mean Time to Resolve (MTTR)
- Performance bottleneck identificati...
- AI tools for log analysis
- incident triage
- Monitoring automation
- ELK stack
- OpenTelemetry
- AWS
- amazon web services aws
- Azure
- Logging systems
- Tracing systems
About Xempla
-
