
Hamming AI
Complete QA platform for voice agents
What Hamming AI does
Hamming AI is the complete QA platform for voice agents from pre-deployment testing to production monitoring. 4M+ calls tested. 10K+ agents monitored. Trusted by public companies and the fastest-growing voice AI teams. Connect your agent, auto-generate test scenarios, and run thousands of test calls in minutes. In production, monitor every call with 50+ built-in metrics and get alerts before customers notice. Sumanyu (CEO) was previously Head of Data at Citizen and Senior Staff Data Scientist at Tesla, where he grew AI-powered sales to hundreds of millions in revenue per year. Our engineering team ranks #1 on Weave (https://workweave.ai) across 10,000+ engineers and 100+ teams.
2 open roles
What the role involves
Location: Remote (North America) or Austin, TX Employment Type: Full-time (no contractors) Department: Engineering About Hamming AI Hamming automates QA for voice AI agents. Everyone is building voice agents. We secure them. In fact, we invented this category. With one click, thousands of our agents call our customers’ agents across accents, background noise, and personalities—then we generate crisp bug reports and production-grade analytics. Reliability is the moat in voice AI, and that’s our whole job. We are one of the fastest engineering teams in the world. We prod deploy 4x / day. I’m looking for someone who can own reliability and scale across our LLM-enabled platform, shipping precise, outcome-driven improvements to high-availability systems. — Sumanyu (CEO) Previously: grew Citizen 4× and scaled an AI sales program to $100Ms/yr at Tesla. Devin Case Study Ranked #1 Eng team OpenAI Dev Day 100billion token list What you’ll do Own core services in TypeScript/Node.js and Python that orchestrate LiveKit, Temporal, STT/TTS, and LLM tooling for real-time voice agents. Scale 1 → N → 100×: take what works today and harden it for 10K parallel calls with 99.99% uptime. Turn human playbooks into productized systems. Harden pipelines for ingestion, evaluation, and analytics so telephony events, recordings, and outcomes propagate reliably across services. Level-up observability: deepen OpenTelemetry/SigNoz and trace-first practices to shrink mean-time-to-truth in prod. Prototype → test → prod: partner with product to ship new LLM-driven behaviors with clear success metrics, guardrails, and regressions blocked in CI. Infrastructure readiness: CI/CD, environment automation, incident response playbooks—customer conversations stay online. You might be a fit if you Have senior/staff experience running distributed backends with real-time/streaming constraints. Are fluent in TypeScript/Node.js and comfortable jumping into Python for ML/audio jobs. Know Temporal (or similar workflow engines), queues, Redis, and PostgreSQL. Have shipped production LLM apps and understand prompt/tool design, evals, and guardrail instrumentation. Operate cloud-native on AWS with Terraform; k8s doesn’t scare you. Are a power user of Cursor/Zed/Devin and were using code-gen before it was cool. Have intuition for what current-gen LLMs can/can’t do—and what tomorrow’s models will unlock. Think independently, grind with customers, and do whatever it takes—without dropping the quality bar. Bonus: built 0→1 real-time systems in Telecom/Networking, Autonomous Vehicles, or HFT; founded something; built AI voice apps. Interesting problems you’ll touch Voice simulations that feel real: accents, overlapping speech, crosstalk, background noise, barge-ins. Massive concurrency: 10,000+ parallel calls with deterministic behavior and graceful degradation. Temporal-driven orchestration for long-running, interruptible call flows. Closed-loop reliability: turn prod failures into auto-generated tests and blocked deploys. Trace-everything culture: make “what happened?” a 30-second question, not a war room. How we work Outcomes over output: we adjust roadmaps when new data lands. Demo early and document decisions so context moves fast. Own incidents: lead the investigation, write crisp notes, land durable fixes. Direct, candid, respectful communication keeps remote teammates in lockstep with Austin HQ. Our stack App: Next.js, TypeScript, Tailwind AI: OpenAI, Anthropic, STT/TTS providers Realtime/Orchestration: LiveKit, Pipecat/Daily, Temporal Infra/DB: AWS, k8s, PostgreSQL, Redis, Terraform Observability: OpenTelemetry, SigNoz Apply If you want to make AI voice agents reliable at scale, let’s talk.
What the role involves
Location: Remote (North America) or Austin, TX Employment Type: Full-time (no contractors) Department: Engineering About Hamming AI Hamming automates QA for voice AI agents. Everyone is building voice agents. We secure them. In fact, we invented this category. With one click, thousands of our agents call our customers’ agents across accents, background noise, and personalities—then we generate crisp bug reports and production-grade analytics. Reliability is the moat in voice AI, and that’s our whole job. We are one of the fastest engineering teams in the world. We prod deploy 4x / day. I’m looking for someone who can own reliability and scale across our LLM-enabled platform, shipping precise, outcome-driven improvements to high-availability systems. — Sumanyu (CEO) Previously: grew Citizen 4× and scaled an AI sales program to $100Ms/yr at Tesla. Devin Case Study Ranked #1 Eng team OpenAI Dev Day 100billion token list What you’ll do Own product features end-to-end: spec → prototype → ship → iterate, across frontend and backend. Work closely with customers: onboard new accounts, run weekly check-ins, and act as a high-agency partner to drive adoption and outcomes. Build core customer workflows for voice-agent QA: test creation, scenario management, evaluation results, analytics, debugging, and triage. Turn messy, high-dimensional data (calls, transcripts, tool events, traces, eval outputs) into product experiences that are obvious and actionable. Partner with customers to understand their reliability pain, then translate it into shipped product with measurable outcomes. Tighten the product loop: instrumentation, funnels, and feedback so we know what’s working and what’s not. Maintain high engineering velocity while keeping craftsmanship: clean APIs, strong abstractions, and excellent UI polish. You might be a fit if you Have 3+ years building customer-facing products in a high-velocity environment (startup experience a plus). Are fluent in TypeScript and comfortable across the stack (React/Next.js + Node services). Ship quickly but with discipline: you write clear code, strong tests where it matters, and avoid accidental complexity. Have strong product instincts: you can simplify complex workflows into crisp UX and make good tradeoffs under ambiguity. Love talking to users, diagnosing friction, and iterating until a feature feels “done.” Care about reliability: you build with observability, failure modes, and data correctness in mind. Communicate clearly: written specs, crisp PRs, and decisions that scale across a fast-moving team. Bonus Experience building analytics-heavy products (dashboards, event pipelines, debugging tools). Familiarity with LLM apps, evals, tool calling, or prompt/guardrail systems. Experience with real-time systems, telecom/voice, or high-concurrency workflows. Strong UI craft: interaction design, information architecture, and performance tuning. Interesting problems you’ll touch Debugging workflows for voice agents: call timelines, transcripts, tool calls, traces, and “what changed?” diffs. Test authoring that scales: scenario libraries, parameterization, coverage, and regression packs. Evaluation UX: turning model-graded / heuristic / human feedback into trustworthy signals and action items. Analytics that matter: reliability metrics customers can run their business on. Enterprise readiness in-product: RBAC, audit trails, data retention, and environment/region controls. Our stack App: Next.js, TypeScript, Tailwind AI: OpenAI, Anthropic, STT/TTS providers Realtime/Orchestration: LiveKit, Pipecat/Daily, Temporal Infra/DB: AWS, k8s, PostgreSQL, Redis, Terraform Observability: OpenTelemetry, SigNoz Apply If you want to build the product layer for reliable Voice AI, let’s talk.
Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.
Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.
Questions and experiences
Nobody has asked anything about Hamming AI yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.
Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.