IND Staff Engineer, Reliability
Job Details
- Location:
- Hyderabad, Telangāna, IN
- Category:
- Information Technology
- Employment Type:
- Full time
- Job Ref:
- R2626209-333
We’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future.
Position Summary
We are seeking a highly skilled T7 AI Operations & Site Reliability Engineer to join our engineering team in Hyderabad, India. This role is laser-focused on the availability, reliability, and performance of our production AI systems. You will own the operational health of AI-powered products — ensuring LLM-based services, agentic workflows, RAG pipelines, and ML inference platforms maintain enterprise-grade uptime while scaling to meet demand. You will build the observability, automation, and incident response capabilities that keep our AI products running 24/7.
Level: T7 (Senior Engineer)
Location: Hyderabad, India
Employment Type: Full-Time
Key Responsibilities
AI Platform Reliability & Uptime
- Own the end-to-end reliability of production AI systems including LLM services, RAG pipelines, agentic workflows, and inference endpoints
- Define and maintain SLOs/SLIs/SLAs for AI products — latency, availability, error rates, token throughput, and response quality
- Design and implement high-availability architectures for AI workloads: multi-region failover, load balancing, auto-scaling, and graceful degradation
- Build circuit breakers, retry logic, fallback models, and rate-limiting strategies to ensure AI services remain available under stress
- Drive availability targets of 99.9%+ for critical AI-powered products
- Establish disaster recovery procedures and regularly test backup/restore for AI data stores and model artifacts
Observability & Monitoring
- Build and maintain comprehensive observability stacks for AI systems — metrics, logs, traces, and AI-specific signals (hallucination rates, model drift, token costs)
- Implement real-time dashboards and alerting for AI service health, model performance, and infrastructure utilization
- Design anomaly detection and proactive alerting to identify degradation before users are impacted
- Monitor LLM provider dependencies (GCP Vertex AI, OpenAI, etc.) and implement automated failover when external services degrade
- Track and optimize cost-per-inference, token utilization, and resource efficiency across AI workloads
Incident Management & Response
- Lead incident response for AI system outages and degradations — triage, mitigate, resolve, and communicate
- Build and maintain runbooks for common AI failure modes: model timeouts, context window overflows, embedding pipeline failures, vector DB issues
- Establish on-call rotations and escalation procedures tailored to AI system failure patterns
- Automate incident detection and remediation where possible — self-healing pipelines and auto-rollback
AI Infrastructure & Platform Operations
- Operate and scale cloud-native AI infrastructure (GCP, AWS) including model serving platforms, Kubernetes containers, and vector databases
- Implement and maintain Infrastructure-as-Code (Terraform) for AI platform environments
- Automate deployment pipelines for model updates, configuration changes, and infrastructure scaling
Collaboration & Documentation
- Partner closely with AI Engineers to ensure new features are built with operability, observability, and reliability in mind
- Define production-readiness criteria for AI services — ensuring all systems meet reliability standards before launch
- Maintain comprehensive operational documentation: architecture diagrams, runbooks, playbooks, and SOPs
- Contribute to architecture reviews with a reliability lens — identifying single points of failure, blast radius, and operational risk
- Participate in on-call rotations and drive continuous improvement of operational practices
Required Qualifications
- Experience: 8+ years of professional experience in software engineering, DevOps, or site reliability engineering, with 1+ year operating AI/ML systems in production
- Education: Bachelor's degree in Computer Science, Software Engineering, or related field (or equivalent experience)
- SRE Fundamentals:
- Deep understanding of SRE principles: SLOs, error budgets, toil reduction, incident management, and capacity planning
- Proven track record maintaining high availability (99.9%+) for production systems at scale
- Experience integrating into observability platforms (Prometheus, Grafana, Datadog, Splunk, or equivalent)
- Strong incident response skills with experience leading war rooms and post-incident reviews
- AI/ML Operations:
- Understanding of AI-specific failure modes: model drift, hallucination spikes, token limit errors, embedding pipeline failures, and provider outages
- Familiarity with LLM providers and platforms (GCP Vertex AI, OpenAI, AWS Bedrock) from an operational perspective
- Cloud & Infrastructure: Advanced-level experience with cloud platforms (GCP, AWS), Kubernetes, containerization, and Infrastructure-as-Code (Terraform)
- Programming: Strong proficiency in Python and at least one systems language; comfortable writing automation scripts, custom exporters, and operational tooling
- Networking & Security: Solid understanding of networking, load balancing, DNS, TLS, and security best practices for cloud-native systems
- CI/CD: Experience building and maintaining deployment pipelines (Jenkins, GitHub Actions, ArgoCD) with automated rollback capabilities
- Communication: Excellent communication skills for incident coordination, stakeholder updates, and cross-team collaboration
Preferred Qualifications
- Knowledge of AI cost optimization strategies — model routing, caching, batching, and tiered inference
- Experience in regulated industries (insurance, finance, healthcare) with compliance and audit requirements
- Cloud certifications (GCP Professional Cloud Architect, AWS Solutions Architect, CKA/CKAD)
- Experience with AIOps — using AI/ML to improve operational intelligence and automated remediation
About Us
We believe every day is a day to do right.
And that belief has guided us for over 200 years. Showing up for people isn’t just what we do, it’s who we are. We’re devoted to finding innovative ways to serve our customers, communities and employees – continually asking ourselves what more we can do.
And while how we contribute looks different for each of us, it’s these values that drive all of us to do more and to do better every day.
Featured Career Opportunities
-
IND Staff Engineer, Reliability
- Location
- Hyderabad, Telangāna
- Employment Type:
- Full time
- Job Ref:
- R2626209
-
INDStaff Software Engineer - AI
- Location
- Hyderabad, Telangāna
- Employment Type:
- Full time
- Job Ref:
- R2626213
-
Credit Risk Finance Consultant
- Location
- Employment Type:
- Full time
- Job Ref:
- R2626196