Arth v2 sets a new standard for streaming speech recognition on India's phone calls

Arth V2 by SquadStack runs at 67 ms, the fastest engine we tested, and beats Sarvam, Deepgram, Google and OpenAI on 863 real Indian sales calls.

Apurv Agrawal
Apurv Agrawal

CEO & Co-founder

October 10, 2026
|
Blog Read Icon
6 min

What does your voice agent hear when a customer says "EMI bees tareekh ko kar dunga" from a busy road? If it hears "tees", the agent confirms the wrong date, and messes up the rest of the conversation.

Today we're making Arth V2 public. Arth V2 is SquadStack's in-house speech recognition model, and it already listens to 80% of the 50 lakh+ calls our agents make every day. On a benchmark of 863 real Indian sales calls, Arth V2 makes fewer meaning-changing errors than the streaming engines from Sarvam, Deepgram, Google, OpenAI, Cartesia and ElevenLabs, and it is the fastest engine we tested, with a latency of 67 ms. Arth is the first model from SquadStack's model factory, where every new model is tested on real Indian sales calls and ships only when it beats the one in production, so a better Arth arrives every few weeks.

Arth V2 at a glance

  • 67 ms latency, the fastest engine we tested, beating Soniox and Deepgram Nova 3, Sarvam, Google Gemini and OpenAI 
  • 8.85% semantic word error rate, beating Sarvam, Deepgram, Google, OpenAI, Cartesia and ElevenLabs on accuracy, and matching Soniox
  • 18% fewer meaning-changing errors than Sarvam Saaras v4, 26% fewer than Google Gemini 3.5 and 28% fewer than OpenAI gpt-live-transcribe
  • Beats every engine on noisy calls and heavy Hindi-English mixing, with 17% fewer errors than Soniox
  • Live on 80% of SquadStack's daily calls, with the benchmark open for anyone to test their own engine

At 67 ms latency, Arth V2 beats every engine we tested

Your agent can't reply until it knows what the customer said, so every millisecond the transcript takes is a millisecond of silence on the line. We measured this as FTR, or finalize-to-response: the time from the moment your customer's turn ends to the moment the engine sends back its final transcript. On 80% of turns, Arth V2 returns it within 67 ms, faster than every engine we tested.

That makes Arth V2 2.6x quicker than Sarvam, about 4x quicker than Google Gemini and 12x quicker than OpenAI. Even the engine closest to it on accuracy can't keep up: Soniox v5 takes 89 ms.

Fewer meaning-changing errors than Sarvam, Deepgram, Google and OpenAI

We ran nine streaming engines and Arth V2 on the same 863 calls, through the same scoring pipeline. Each engine heard every call as it happened, the way your agent hears it, on ordinary 8 kHz phone audio.

We scored them on semantic word error rate, which is the share of words an engine gets wrong in a way that would change what your agent does next. A wrong EMI amount counts as an error, while an extra "haan" at the start of a sentence does not. Arth V2 kept those errors to 8.85% of all words spoken, ahead of Sarvam, Deepgram, Google, OpenAI, Cartesia and ElevenLabs on the exact same calls.

Arth V2 leads where calls get hardest

Clean audio flatters every model, and your customers rarely call from a quiet room. They pick up at a shop counter, in a moving auto or at home with the TV on, and they switch between Hindi and English in the middle of a sentence.

The benchmark splits every call by background noise and by how much Hindi and English it mixes, so you can see how each engine holds up as conditions get worse. Arth V2 leads every engine on exactly the calls that decide whether a voice agent works in India.

Noise moves every engine more than anything else. On noisy calls, Arth V2 beats every engine we tested at 10.5%, with 17% fewer errors than Soniox, 12% fewer than Nova 3 and 29% fewer than Sarvam. Going from clean calls to noisy ones, Arth V2's error rate rises by 3.8 points, the smallest climb of any engine we tested, while Nova 3 rises by 5.8, Soniox by 7.8, Google Gemini by 7.9, Sarvam by 8.0 and Cartesia by 11.5. So the noisier the line, the further ahead Arth V2 gets.

The same holds when customers lean on Hinglish. On calls with heavy Hindi-English mixing, Arth V2 beats every engine we tested at 9.7%, ahead of Soniox at 10.4% and Nova 3 at 10.8%. It also beats every engine on marketplace calls, the hardest vertical in the benchmark, with 12% fewer errors than Soniox, and on BFSI calls it makes 35% fewer errors than Sarvam.

Where a wrong word costs the sale

Two errors can look identical on a scorecard and cost very different things on a call. In one benchmark call, the customer says "तीन हजार लगभग" (about three thousand) and an engine writes "मतलब तीन हजार लगभग", which changes nothing. In another, the customer says "छबीस" (twenty-six) and the engine writes "छत्तीस" (thirty-six), so your agent now has the wrong number.

That's why the benchmark sorts every meaning-changing error by what it breaks on the call. Two kinds lose a sale faster than any other: a wrong number, where your agent quotes the wrong loan amount or EMI date, and a flipped negation, where a "nahi" gets heard as "haan". Against Sarvam, Deepgram, Google and OpenAI, Arth V2 makes the fewest of both, and it stays ahead of all four on misheard promises to pay and on every other word that changes what your agent does next.

On content errors, the misheard words that make your agent ask again or go off track, Arth V2 makes the fewest of every engine we tested, Soniox included.

Hear the difference

Hear the difference: Arth vs other engines (V6 · GTM industries) (Copy)
word heard right by SquadStack Arth v2word heard wrongword left out

How we built Arth V2

We trained Arth V2 from scratch on 60 crore+ minutes of real Indian sales conversations from more than 10 crore unique speakers, covering over 85% of the country's pincodes. These conversations come from six years of running AI-native sales contact centres for India's consumer brands, so they only exist inside live sales operations and nobody can scrape them from the internet. They carry everything your agent actually meets on a call: the accents, the background noise, the language switches, and the amounts, dates and pincodes where a sale is won or lost.

Arth is built for streaming from the ground up, so it transcribes your customer while they're still speaking, on the live 8 kHz line, and your agent can reply in real time. Every number in this post is a streaming number.

A benchmark you can check for yourself

Vendor benchmarks are easy to doubt, so we published this one on Hugging Face. The SquadStack Conversational Streaming ASR Benchmark holds 863 real Hindi-English sales calls from 19 campaigns across marketplace, logistics, travel, BFSI and education. Before we ran a single engine, we checked the set against the audio used to train Arth and rebuilt it, so none of its calls or speakers overlap with Arth's training data and Arth V2 hears these calls for the first time, just like every other engine.

The calls fill every cell of a three-by-three grid of noise and language mixing, and no single vertical holds more than 23% of the speech. Every engine ran through the same pipeline, and every error each engine made is published call by call, with its verdict and its type.

Already on your calls, and better every few weeks

Arth has already handled over 1 crore minutes of live customer calls in two months, and every new agent on SquadStack now starts on it by default. It runs inside every Humanoid Voice AI Agent, so if you run campaigns with us, your customers are already talking to agents that hear them through Arth, and each new version reaches them without you changing a thing.

Every one of those calls also trains the next version. Lift, our self-improvement loop, finds the turns where a misheard word changed the conversation and turns them into training data, so the failure your agent hits this week becomes something the next Arth handles.

‍

Request a Demo to experience Arth on a live call inside the Humanoid Voice AI Agent.