Rabbit logo
← All open positions

Senior Site Reliability Engineer

Jakarta, Indonesia
Full-time

The Role

Help Rabbit build and ship faster with AI — safely, securely and reliably.

Rabbit runs automated cost optimization across enterprise Google Cloud environments. When we change a customer’s BigQuery reservations or rightsize their GKE clusters, those changes need to be correct and dependable. Reliability is central to the trust customers place in our product.

Our foundation is already in place: logging, alerting, automated deployment and Terraform-managed infrastructure. Your mission is to evolve that foundation for an AI-accelerated engineering team: turn faster implementation into faster, dependable delivery through automated validation, safe releases and rapid feedback.

You’ll apply proven SRE practices — SLOs, observability, incident response and deployment safety — to AI-assisted development and agent-driven workflows. The goal is to increase how quickly the team can deliver verified improvements, while controlling production risk and reducing manual operational work.

What You’ll Do

  • Make AI-assisted delivery faster and safer. Build automated validation, progressive rollout and recovery mechanisms that let engineers and agents move quickly with clear checks before and after changes reach production.
  • Make reliability measurable. Define and operationalize SLOs, SLIs and error budgets, and use them to guide practical decisions about delivery speed, stability and reliability work.
  • Automate workflows with AI agents. Identify repetitive operational work and build reusable agent-driven workflows for alert triage, incident investigation, routine maintenance and reporting. Add verification and human approval where needed, and measure the reduction in manual effort.
  • Improve observability and feedback. Evolve logging, metrics, tracing and alerting so failures are detected early and changes can be traced, investigated and verified.
  • Turn incidents into lasting improvements. Improve runbooks, investigation and blameless postmortems, and translate recurring problems into tests, safeguards and automation.
  • Keep infrastructure reproducible. Extend our Terraform and delivery tooling so environments remain consistent and changes stay reviewable as the platform grows.
  • Improve GCP reliability and efficiency. Strengthen our cloud infrastructure, networking, access controls and capacity management, balancing performance, reliability and cost.
  • Use AI to accelerate reliability engineering itself. Build maintainable tooling in Go, Python or a comparable language, and use agents to accelerate investigation, implementation, testing and documentation while verifying their outputs.

How We Work — AI-First, Agentic by Default

AI-assisted engineering is an expectation of this role, not an optional experiment. Claude Code, Cursor and agent-driven workflows are part of how we work, including infrastructure and reliability engineering.

We want someone who actively looks for ways to increase engineering speed with AI and makes those improvements safe to repeat. That means shorter feedback loops, automated checks, traceable changes and recovery paths — not simply generating more code.

You remain accountable for engineering judgment: what to automate, how to verify it, when human approval is needed and when a change should be stopped or rolled back. Success means faster delivery of reliable improvements, less repetitive work and a platform the team can trust.

What You’ll Bring

Must-have:

  • 6+ years in SRE, production engineering or infrastructure-heavy backend roles, with hands-on ownership of production systems
  • Strong production GCP experience, including Cloud Run, networking and IAM. Hands-on Google Cloud experience is required and will be assessed during the interview process
  • Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety
  • Strong observability and troubleshooting skills: you can make systems debuggable, identify root causes and verify that a fix works
  • Coding ability in Go, Python or a comparable language, with experience building maintainable operational tooling
  • Experience leading production incident investigation and driving follow-up improvements that prevent recurrence
  • Practical fluency with AI coding agents and the ability to critically review, test and validate their work. You are motivated to make AI-assisted engineering faster and more dependable
  • Strong written English and the discipline to collaborate asynchronously with a distributed team

Nice to have:

  • Kubernetes / GKE experience, including deploying, operating, debugging and scaling containerized services
  • Hands-on Datadog experience, including dashboards, monitors, logs, APM and distributed tracing
  • Experience improving delivery speed and safety through progressive delivery, policy-as-code and automated rollback
  • GCP cost management or FinOps experience
  • Security experience

Why This Role

  • Shape how an AI-first engineering team scales. Build the reliability practices and automation that let Rabbit turn faster development into dependable customer outcomes.
  • Work on systems with real customer impact. Rabbit operates inside enterprise GCP environments, where reliability and security directly affect customer trust.
  • Own meaningful improvements. Work in a small team with short decision paths and end-to-end ownership, supported by appropriate review and production safeguards.
  • Use AI as an engineering multiplier. Apply agentic tools to infrastructure, delivery and operations, and help define how we measure and improve their impact.

How to Apply

Send your CV or LinkedIn profile along with a short note on a deal you won on technical merit — what the prospect’s objection was, how you handled it, and why they believed you. If you have a recorded talk or demo, link it.

Apply now

Why join Rabbit

Build from zero to production

You won't maintain someone else's code. You'll discover opportunities, prototype solutions, validate them with real customers, and ship them. The full cycle, every time.

AI-driven engineering

We use agentic coding tools as a core part of our workflow. If you've been waiting for a team that actually embraces AI-assisted development, this is it.

Small team, no layers

No middle management, no ticket queues, no waiting for approvals. You own your work end-to-end and your decisions directly shape the product.

Real customer impact

We work with enterprises running some of the largest GCP workloads in Europe. You'll see the direct impact of what you build — not through dashboards, but through customer conversations.

Contact us icon

Get in touch to start saving

We help businesses save 30-50% on their Google Cloud spending and provide full clarity on their costs.
Automated cloud cost optimization for teams at scale

Rabbit helps engineering and data teams manage and optimize cloud costs across large Google Cloud environments, without slowing down delivery.

ISO 27001 badgeSOC 2 badge

SolutionsCost Insights for All TeamsFor Data TeamsBigQuery for Data TeamsFor Platform TeamsAutomationAgentic Cloud Cost Optimization
Google Cloud Partner logoGoogle Cloud Platform Marketplace logo with link

Rabbit logo
TERMS AND CONDITIONS
PRIVACY POLICY
© 2026 Follow Rabbit PTE Ltd. Google Cloud Partner.