Applied AI Engineer

at Infer — Operating system for insurance agencies

Bengaluru, KA, IN / Bengaluru, Karnataka, INFull-time3+ years₹2M - ₹5M INR0.01% - 0.20% equityYC-S21

What the role involves

About us

Infer is building the operating system for insurance agencies. We make AI agents(including voice agents) that handle the work agencies have always done by hand: qualifying inbound leads, helping producers during live calls, auditing calls after, running renewals, and bringing churned customers back.

Our long bet is that AI eventually sells insurance directly. Agencies are the wedge because that is where the work, the data, and the customer relationships actually live. Get good there, and the rest follows.

We are a YC company and have raised from Stellaris Venture partners and others. Founders are: Vaibhav, Urvin and Suneel. Vaibhav was an architect and AI researcher(at Purdue) now a licensed insurance agent. Urvin worked at BCG, is a surfer with six pack abs. Suneel is an IITian and a philomath.

A few reasons to join us

  • We like pushing each other on team to test the limits because that's when you rediscover yourself.
  • We’re paranoid about making customers succeed (we challenge whats already good)
  • We love people who question-challenge-build.
  • We're highly transparent founders to work with & love getting challenged.
  • Finally, we love people who’re interdisciplinary.

About the role

We're hiring an Applied AI Engineer to own the system that tells us whether our voice agents are getting better, and to keep them getting better on their own.

Voice quality is the product. If an agent stutters, hallucinates a quote, or misses a disclosure, we lose trust, deals, and sometimes compliance footing. The system that catches all of that before customers do is the most important infrastructure we will build this year.

Today we run thousands of conversations a day with real prospects. We need a harness that scores every change end to end, a benchmark suite that runs against any new model the day it drops, a red-team pipeline that probes our agents for failure modes, and self-improvement loops that feed production failures back into the eval set.

This is an evals and infrastructure role with deep LLM work. You will touch audio, but the center of gravity is the harness and the loops around it. Think of the harness as CI for voice conversations: it runs synthetic and real calls through our stack and scores agent behavior at every layer (STT, LLM, tools, TTS, full call outcomes), so we catch regressions before customers do. New models are coming out every few weeks, so the question is not just whether ours is good today, but whether we can tell within a week if a new open source release should replace it.

What you'll do

  • Building and maintaining the eval framework that scores voice agent quality across transcription, LLM reasoning, tool use, TTS, and full-conversation outcomes
  • Design voice agent behavior: system prompts, tool use, conversation flow, error recovery, and guardrails for real-time interactions
  • Drive STT and TTS accuracy improvements by comparing providers, tuning configurations, and running rigorous A/B experiments the team can act on.
  • Drive TTS quality improvements voice selection, latency vs. fidelity tradeoffs, prosody, edge cases
  • Curate and grow our evaluation datasets, including hard-case mining from production traffic
  • You'll build benchmarks we can run against any new model in days, run a red-team pipeline that probes for jailbreaks, hallucinated quotes, and compliance failures,
  • Partner with backend engineers to wire eval signals into CI so regressions get caught before they ship
  • Wire eval signals into CI so regressions block merges, and build self-improvement loops where hard cases from production auto-feed the eval set and our prompts optimize themselves over time.
  • What success looks like
  • Day 30
  • You understand how our agents work across prompts, tools, evals, telephony, and customer systems.
  • You have shipped a v1 of evals with at least one end-to-end metric the team trusts.
  • You are sitting in on customer call reviews and tagging failure modes by hand to learn where the real problems live.
  • You have one new model (open

About Infer

Full Infer profile

Other roles at Infer

Similar Engineering, Machine learning elsewhere

jobo is a browser extension. Open this on a computer to install it.