I design, train and ship voice systems end to end: the models, the media path, and the evals that decide what ships. Each question starts with one plain sentence, then the technical detail behind it.
Candidates speak with an AI interviewer in live two-way audio. It asks follow-ups, scores against the hiring manager's rubric, and produces a structured report. I built and own the whole thing.
The agent answered too slowly to feel human, so I stopped using one fixed setup and started building the right agent for every conversation on the fly.
The first iterations ran one fixed pipeline: STT, then LLM, then TTS, each stage waiting on the one before it. That design does not get under a second, no matter how fast the models are.
So I deleted the fixed pipeline. RecTek has no pre-baked agents. Per session, the orchestrator builds the agent on the fly: the best available STT, LLM, TTS and VAD combination for that moment, plus the fine-tuning parameters and the knowledge base, then spawns it. The countdown when an interview starts is that agent being constructed live. A config that misses its performance budget is evicted and re-checked in a lazy validation queue, and only comes back once it passes, so there is always a healthy agent to hand the user. Streaming every stage and cutting the serialization points is what bought the milliseconds.
The deep one is an Eastern Armenian Kokoro-82M fine-tune built from scratch: my own G2P into the 178-token phoneme vocabulary, two-stage training, Kanade voice cloning, and an ASR round-trip harness that decides what ships. Same pattern across the portfolio: 300+ fine-tunes, 50+ quantized builds, 25+ trained SLMs. RecTek is the showcase because it is the most recent.
Three levels. Weights: the TTS prosody and style pathway, because pace, energy and pitch are what a person reads as calm, serious or warm. Policy: persona, backchannel, interruption and escalation rules tuned against rubric-scored transcripts, in the structure OpenAI's live prompting template uses. Data: SFT on curated multi-turn examples, then DPO on hand-labeled pairs.
Two loops. The model loop: synthesize, score with dual-ASR CER plus spectral analysis plus blind listening, gate, diagnose, retrain. The production loop: traces, rubric evaluation, a re-ranker that picks the best outcomes, curated datasets, fine-tunes, then benchmark plus a safety regression set before any gated release. For a mental-health product that safety set is the part I would never skip.
Digital health, for real. I shipped PDA Pro, an iOS product that turns psychological frameworks into product behavior for parents of children with Pathological Demand Avoidance, built with domain experts, sensitive data handled privacy-first.
Empathic voice in a hard language. Eastern Armenian had no usable TTS, so I built one. The emotion layer is next: excited, sad, serious, curious and neutral, with pre-registered gates.
Research credibility. Loss against CER anti-correlation, a quantified bf16 audio deficit, a silent-failure taxonomy for TTS training, an exposure-ledgered eval set meant for release. MSc AI systems in progress.