Robotics · Live streaming · Voice AI · Inference · Netcode

Real-time infrastructure for physical and multimodal AI.

One QUIC transport and one C++ runtime carry robot control, live video, phone calls and model output on a single connection.

Measured, not modelled

One runtime. Real hardware. Real numbers.

Robot teleoperation63live sessions on one CPU core, at 40 ms p95 round trip and 100% deliveryMeasured · relay on GCE c4-highcpu-8, real NIC, run to saturation
Voice17,000concurrent calls on one server, 0 failuresMeasured · G.711, same box: Asterisk 5,000 · FreeSWITCH 4,500
Live streaming500 / 500viewers on one live stream, 0 droppedMeasured · real RTMP ingest → relay → SDK viewers
Inference2.7×faster follow-up turn with our KV cache, identical outputMeasured · RTX A5000, vLLM, Qwen2.5-1.5B, 7.2k-token prompt: 129 ms vs 353 ms

Every figure comes from a benchmark we can rerun in front of you.

one box · 8 cores
voiceinferencewebrtcsiproboticsgames
01 · The substrate

It starts with one packet.

A single unit of your data, on the wire.

One live connection.

That packet is one of thousands on a single connection.

One core.

Every core runs its own connections, full tilt.

0★

messages/second on one 8-core box — and improving. 36,732 concurrent connections at 0.4% CPU.

One substrate.

Every box, every modality, on the same transport.

AI
TV
noise
caller
coworker
02 · Voice AI

It starts with one sample.

A scrap of sound, captured the instant it's spoken.

Which becomes speech.

Words on the wire, in real time.

A live call.

Caller, model, agent — one pipe, both directions.

It hears the whole room.

The line tells the caller from the TV, a coworker, or background noise — and listens to the right one.

Thousands of calls at once.

On every provider, with first-class barge-in.

OpenAI RealtimeGemini LiveElevenLabsAnthropicDeepgram
Explore in detail →
03 · Inference

It starts with one token.

A single piece of the answer, fresh off the model.

A stream off the GPU.

Tokens leave the GPU and start arriving immediately.

Straight to your client.

Point your OpenAI-compatible client at the line — nothing in your app changes.

First token, fast.

Same SSE wire your client already speaks — the moment a token exists, your app has it.

Explore in detail →
meet.example.com/abc-xyz
you
alex
jordan
sam
04 · Browser · WebRTC

Audio, video, text and control.

Four modalities, in one place.

Each on its own lane.

A stall in one never blocks another — no head-of-line blocking.

All on one connection.

One line, nothing to glue together.

On your transport.

Your path end to end — no plugins, no native app.

Explore in detail →
05 · SIP · Telephony

It starts with one call.

A carrier hands you an inbound call.

Then straight to your agent.

Media rides QUIC direct to the agent — no extra hops in the path.

Pull a human onto the line.

Listen, whisper, or take over — over the same connection, without rerouting the call.

One call, the whole wallboard.

Fan it out to every supervisor and dashboard, live.

TwilioVapiLiveKitDailyChimeTata IMSAudiocodes
Explore in detail →
06 · Robotics

One robot, in control.

Onboard autonomy stays on its own network.

Its data, in realtime.

A live mirror of everything it senses.

Up to the cloud.

Reaches the cloud over QUIC — observe from anywhere.

500

live tracks per shard at 50 fps, zero loss.

ROS 2ZenohMQTTDDS
Explore in detail →
07 · Games · Netcode

A player acts.

Input rides its own fast lane, as datagrams.

Two lanes, one session.

Fast state as datagrams, chat and events on reliable streams.

Even behind a firewall.

Players whose network blocks game ports still connect, on port 443.

Every player, in sync.

Tick-aligned state, lobby and matchmaking included.

UnityWeb · WebTransport
Explore in detail →
Six modalities converge on one substratevoiceinferencewebrtcsiproboticsgames
08 · One substrate

Six modalities. One substrate.

Every branch flows back into the same transport line.

0★

payload types, one transport. Add a modality, not a stack.

Inference

Serving that respects the GPU.

Inference runs on the same connection as your audio, video and robot control, so the model knows the moment its answer stops mattering.

Cancel propagation

The GPU stops when the user does.

A caller hangs up, a deadline passes, or the user talks over the agent. The transport carries that signal to the model and the generation stops. You stop paying for tokens nobody will hear.

Three triggers · caller · deadline · newer turn
ClutchKV

Our own KV cache for vLLM.

Written in C++. It saves the attention state of a conversation and loads it back on the next turn, so the model does not compute the same prompt twice.

Follow-up turn 129 ms vs 353 ms recompute · 2.7× · identical tokens
Vision for robots

Faster answers about the scene.

The same cache works for vision-language models. A robot that asks several questions about one task gets its first answer about twice as fast.

Qwen2.5-VL-3B · first token 144 ms → 69 ms with a shared task prompt
KV-aware routing

Each request goes to the right GPU.

The router reads live cache and load metrics from Triton and TensorRT-LLM, and picks a GPU on that data. OpenAI-compatible API, your model, your GPUs.

Least-cache and least-busy strategies across pooled GPUs
More on the same runtime

One engine, more products.

The transport that carries robots and voice also runs our VPN, remote desktop and tunnel. Same runtime, same operations team.

VPN

A VPN that keeps latency flat under load.

IP over QUIC. Sessions survive a change of network, and traffic on port 443 looks like ordinary HTTPS. Two-factor sign-in and per-user entitlements. Desktop clients for Windows, macOS and Linux.

Measured vs Quincy · RTT under load 0.51 ms vs 14.69 ms · 38% less CPU per GB at 4 cores
QuickDesk

Remote desktop on your own relay.

Screens go through a relay you host, so pixels never leave your infrastructure. Frames carry priority and expiry, and sessions survive roaming. Attended or unattended.

Windows · macOS · Linux · Android
Tunnel

Public endpoints for private services.

A self-hosted alternative to ngrok and frp. One command gives a local service with no public IP a stable public URL, with token and IP policy at the edge instead of in your app.

HTTP · TCP · UDP · self-host or cloud
Drop-in SDKs

Ship on the substrate in an afternoon.

Nine language SDKs, from the browser to embedded devices, plus a ROS 2 middleware binding for robots.

py

Python

Async client, drop-in for the realtime API. Packages are shared during your pilot.

ts

TypeScript

Browser + Node, with a helper for browser media over your transport.

++

C++ / Rust

Bare-metal bindings to the core. Zero-copy frame handoff.

# realtime voice in 6 lines
from clutchcall import Session
s = Session(provider="openai-realtime")
async for ev in s.connect("+1..."):
    if ev.bargein: s.cancel()   # instant
    play(ev.audio)

Coming from another transport? The migration guide maps your existing calls onto the substrate one site at a time.

Read the migration guide →
Pricing

Start with a paid pilot.

Six weeks on your real traffic, with success criteria we agree in writing. The pilot fee is credited to your first month if you continue.

Production is priced per concurrent channel per month, with telephony passed through. On-prem is licensed per node. We quote against your workload after a 30-minute call.

FAQ

Questions builders ask.

No. The substrate exposes the realtime + telephony interfaces your code already speaks. You point your existing client at our endpoint and keep the rest. There's a drop-in shim for OpenAI Realtime, a SIP listener for carrier traffic, and idiomatic SDKs in 9 languages when you want native types.

Interruption detection lives in the transport, not in each provider's SDK. The moment caller audio crosses the speech-gate threshold we hold the agent's output stream and emit a cancel — same code path whether the model behind it is OpenAI Realtime, Gemini Live, or a self-hosted llama.cpp. The tail you hear after you start talking is one packet, not a sentence.

WebSocket sits on TCP, so a single dropped packet head-of-line-blocks every subsequent frame on the same socket — your agent's audio stalls behind your text. QUIC streams are independent: a lost packet only delays its own stream. On a 3% loss link in our DevTools benchmarks, the inter-arrival p95 widens by ~6× on WS and ~1.4× on QUIC. We share the histograms during a pilot. (For a deeper dive on QUIC's loss-recovery story, see the Google QUIC team's writeups.)

WebTransport is shipped in Chrome and Edge today. Safari and Firefox still trail, and some corporate proxies block UDP. The SDK transparently falls back to WebSocket-over-HTTPS using the same agent and the same fast-cancel barge-in path — you just lose the QUIC head-of-line-blocking benefit on that one client. No code change.

Every engagement starts with a paid 6-week pilot on your real traffic, with success criteria we agree in writing: latency, barge-in, call completion and cost per call. The pilot fee is credited to your first month if you continue. Production is priced per concurrent channel per month, with telephony minutes passed through. On-prem is licensed per node. We quote against your workload after a 30-minute call.

Yes. The core is a single statically-linked binary plus a Redis. It runs in your own racks — including air-gapped deployments where the model worker is also local — and there's a hard-enforced licensing path for that. Ask us about it during your pilot; we keep our hosted tenants and self-host tenants on the exact same wire so behaviour matches.

LiveKit, Daily, and Chime are SFUs built for many-party video. Their per-participant-minute pricing assumes a meeting room. Twilio Media Streams is PCMU over WebSocket, which gives you carrier reach but not codec choice or sub-200ms barge-in. We're a single-tenant agent transport: one caller talking to one (or many) AI providers, with codec/transport/perception all owned end-to-end. Different shape, different price.

Anything that speaks a streaming-audio or streaming-text API. We ship verified adapters for OpenAI Realtime, Gemini Live, ElevenLabs Conversational AI, and a local llama.cpp worker over QUIC. For STT-only or TTS-only pipelines you can mix providers — Deepgram on input, Cartesia on output, GPT-4o in between — and the substrate handles the joins.

Python, TypeScript, Go, Rust, Java, C#, Swift, Kotlin, and C++ — same API shape, idiomatic types per language. They all wrap one C++ core via FFI, so a wire-protocol fix in core ships to every SDK by re-building. The Python and TypeScript clients see most of the dogfooding; the others are tested against a polyglot smoke matrix on every release.

On the public internet to OpenAI Realtime via our pool, first-token-after-barge-in lands in ~120-180 ms p50 from the moment the caller's first speech-frame arrives, with the long tail dominated by the upstream model not the transport. On localhost the transport itself adds <2 ms over raw UDP. We share the detailed histograms during a pilot.

We do not hold SOC 2 or HIPAA certification today. For regulated buyers we deploy on-prem, in your own data centre or cloud account, so call audio and records stay inside your infrastructure. Hosted pilots run in a single region we agree with you. We do not train models on customer audio.

It starts with a 30-minute call to map your current stack and agree the success criteria. We then connect your numbers or carrier and your agent, and run the 6-week pilot on real calls. How long setup takes depends on your carrier and agent framework; we give you a date on the first call.

Company

A deep-tech company, built in Mangaluru, Karnataka.

ClutchCall is built by Throughgen Innovations Private Limited. We write the runtime, the transport and the GPU cache ourselves, in C++, at the systems level. It is not a wrapper around someone else's API.

New to QUIC? Our free academy teaches networking from the ground up: layers, DNS, NAT, latency, then TCP versus QUIC and Media-over-QUIC, with hands-on labs.

Founder, CEO & CTOAshwin Kallaje

Designed and wrote the C++ runtime, the transport and the edge network. Leads the company and works directly with every pilot customer.

Co-founderSwasthik Padma