Founding AI Voice Engineer  ·  the three questions

Real‑time voice AI, built end to end.

I design, train and ship voice systems end to end: the models, the media path, and the evals that decide what ships. Each question starts with one plain sentence, then the technical detail behind it.

0+
voice products
deployed
0+
fine-tuned
models
0+
quantized
models
0+
SLMs
trained
200-300ms
agent reply
latency
Q1
The most advanced real-time AI voice system you have personally built and deployed. What did the product do, what did you own, what was the architecture and stack, and what scale did it run at?

RecTek: an AI interviewer that talks to candidates in live voice.

In plain words

Candidates speak with an AI interviewer in live two-way audio. It asks follow-ups, scores against the hiring manager's rubric, and produces a structured report. I built and own the whole thing.

The voice path, in order  click a pool chip to change architecture Fig. 01 · live architecture · full cascade
session traces · eval + re-ranker · gated updates AGENT CONFIG POOL no pre-baked agents · evict → re-validate → re-enter FULL CASCADE HALF CASCADE DUPLEX NATIVE AUDIO MULTIMODAL KNOWLEDGE & MEMORY temporal knowledge graph OpenViking FalkorDB Graphiti MODEL FACTORY fine-tune · distill · quantize LLMOps MLOps eval gates TRANSCRIPTION & RECORDING always-on sidecar sessions → traces, audits and evals 01 Mic capture 01 Mic + Cam audio + video 02 VAD end of speech 02 Turn + Speaker active detection 02 Turn Detect in-band 03 STT streaming 03 Audio-in LLM speech in inside the model pass-through 04 Orchestrator agent per session 05 LLM persona + KB 05 Duplex Model two live streams 05 Speech Model audio to audio 05 Live Multimodal audio · video · text 06 TTS streaming voice inside the model pass-through 07 Speaker avatar audio

What I ownEnd to end, no voice team behind me

  • The voice path: WebRTC and LiveKit media, streaming STT, LLM and TTS, VAD, end-of-utterance handling, turn-level interruption and cancellation.
  • The intelligence: multi-agent scoring, dynamic RAG and knowledge-graph workflows (FalkorDB, Graphiti temporal graph), distributed learning phases, structured human review.
  • The platform: architecture, deployment, observability and cost. A modular monolith at first, now event-driven with microservices, a Rust aggregator, and an LLMOps/MLOps model factory that fine-tunes, distills and quantizes the models that enter the pool.

ScaleWhat it runs at

  • 70+ company accounts and 300,000+ candidates on the platform.
  • One account alone: nearly 5,000 candidates and 60 live job posts, at around 500 concurrent users, no special-casing.
  • 20+ production releases. Platform-wide: 100K+ concurrent at 99.99% uptime on autoscaling Kubernetes behind Cloudflare.
Performance and cost figures are confidential while the product is in stealth. I would rather walk you through them on the call.
Q2
The hardest technical problem you encountered with real-time voice, and how you solved it.

From 2-3 seconds to a quarter of a second.

In plain words

The agent answered too slowly to feel human, so I stopped using one fixed setup and started building the right agent for every conversation on the fly.

Then One fixed pipeline. Each stage waited for the one before it.
STT
LLM
TTS
2-3 s to first audio serial · no overlap
Now Stages overlap and stream. First audio leaves long before the reply is finished.
VAD
STT
LLM
TTS
reply streams out while you listen
200-300 ms to first audio overlapped · streaming

The first iterations ran one fixed pipeline: STT, then LLM, then TTS, each stage waiting on the one before it. That design does not get under a second, no matter how fast the models are.

So I deleted the fixed pipeline. RecTek has no pre-baked agents. Per session, the orchestrator builds the agent on the fly: the best available STT, LLM, TTS and VAD combination for that moment, plus the fine-tuning parameters and the knowledge base, then spawns it. The countdown when an interview starts is that agent being constructed live. A config that misses its performance budget is evicted and re-checked in a lazy validation queue, and only comes back once it passes, so there is always a healthy agent to hand the user. Streaming every stage and cutting the serialization points is what bought the milliseconds.

full cascade
half cascade
duplex
native audio
multimodal live
+ VAD
+ fine-tuning params
+ knowledge base
built per session, not pre-baked
Q3
Very briefly: trained or fine-tuned a voice model? Gone beyond prompting and trained the conversational behavior? Built a learning loop from production conversations?

Yes to all three.

a  ·  trained a voice model?
Yes. I trained the voice model itself, not just wired up voice APIs.

The deep one is an Eastern Armenian Kokoro-82M fine-tune built from scratch: my own G2P into the 178-token phoneme vocabulary, two-stage training, Kanade voice cloning, and an ASR round-trip harness that decides what ships. Same pattern across the portfolio: 300+ fine-tunes, 50+ quantized builds, 25+ trained SLMs. RecTek is the showcase because it is the most recent.

8 GB GPU82M paramsCER-gated
b  ·  trained the behavior?
Yes. How the agent behaves, when it speaks, when it stops, how it sounds.

Three levels. Weights: the TTS prosody and style pathway, because pace, energy and pitch are what a person reads as calm, serious or warm. Policy: persona, backchannel, interruption and escalation rules tuned against rubric-scored transcripts, in the structure OpenAI's live prompting template uses. Data: SFT on curated multi-turn examples, then DPO on hand-labeled pairs.

prosody / stylepolicy + rubricSFT + DPO
c  ·  learning loop?
Yes. Real conversations get scored and fed back in.

Two loops. The model loop: synthesize, score with dual-ASR CER plus spectral analysis plus blind listening, gate, diagnose, retrain. The production loop: traces, rubric evaluation, a re-ranker that picks the best outcomes, curated datasets, fine-tunes, then benchmark plus a safety regression set before any gated release. For a mental-health product that safety set is the part I would never skip.

dual-ASR gatestraces to datasafety set
The learning loop, in order Fig. 02 · production feedback cycle · gates
fail → re-tune · new data in GATE 01 Conversations production 02 Session traces recorded 03 Eval vs rubric scored 04 Re-ranker best outcomes 05 Curated data consented 06 Fine-tune LoRA · SFT · DPO 07 Benchmark + safety + gates 08 Gated release pass → production
+
And why this role. The posting asks for someone who cares about making mental support available to anyone with a smartphone. Three things point here directly.
i.

Digital health, for real. I shipped PDA Pro, an iOS product that turns psychological frameworks into product behavior for parents of children with Pathological Demand Avoidance, built with domain experts, sensitive data handled privacy-first.

ii.

Empathic voice in a hard language. Eastern Armenian had no usable TTS, so I built one. The emotion layer is next: excited, sad, serious, curious and neutral, with pre-registered gates.

iii.

Research credibility. Loss against CER anti-correlation, a quantified bf16 audio deficit, a silent-failure taxonomy for TTS training, an exposure-ledgered eval set meant for release. MSc AI systems in progress.