Back to Portfolio
AI / EdTechLive

100Minds.ai

A practice-based leadership and power skills training platform powered by AI. I architected the core AI infrastructure — a voice-enabled tutor, an interactive AI avatar, and RAG pipelines grounded in curated learning content — with a sub-second end-to-end latency budget.

Lead AI Engineer
2025
AI / EdTech
View Project
RAG
Grounded AI Responses
< 1s
Voice Response Latency
Live
In Production
2
AI Interfaces Built

Static corporate training videos are useless for teaching power skills. You can't learn how to give a hard performance review by watching someone else do it. You have to try, get it wrong, and get feedback in the moment.

The mandate at 100Minds was to build a system that could dynamically role-play complex, high-stakes leadership scenarios — performance reviews, conflict resolution, difficult conversations — where the learner speaks aloud, the AI responds in real time, and the interaction feels like coaching, not a course.

Three constraints, all non-negotiable

The technical hurdles were stacked, and each one was a product-killer if we missed it.

Near-zero latency. Any noticeable pause during a tense role-play scenario breaks psychological immersion completely. The learner has to feel like they're in a conversation, not waiting for a server.

High conversational memory. The AI has to remember what the learner said three turns ago and use it — otherwise the role-play falls apart the first time the learner references something earlier.

Zero hallucinations. Generating bad corporate advice is a real business liability. The AI could not invent leadership frameworks that don't exist. Everything it said had to be grounded in curated material.

The RAG pipeline solved the hallucination problem

I built a RAG pipeline that ingests and chunks the platform's curated training content into a vector database, so every AI response is anchored in real material — never invented. On top of that I integrated a voice-enabled tutor for spoken interaction and an AI avatar for immersive scenario practice.

The learner is looking at a face and speaking naturally, not reading text on a screen. The full pipeline orchestrates speech-to-text, LangChain reasoning, vector retrieval, text-to-speech, and avatar syncing tight enough that users don't perceive the AI as thinking.

The real problem was latency

The hardest part wasn't the RAG pipeline. It was making the voice interaction feel natural.

Early versions had noticeable latency — two, sometimes three seconds between the learner finishing a sentence and the AI responding. Awkward turn-taking where the AI would talk over the learner, or leave dead air. Both of those are immersion killers in a coaching scenario. You can survive a bad AI response. You can't survive an AI that feels laggy.

I spent weeks tuning the streaming pipeline. Switched to streamed STT so we started processing before the learner finished. Ran retrieval and generation concurrently instead of sequentially. Built an interruption handler that would cut off the AI mid-sentence when the learner started talking.

A slightly worse model with tight latency feels better than a smarter model that pauses. Voice AI is a latency product, not a quality product.

The avatar sync problem

The avatar added another layer of complexity. Synchronizing lip movement and expression with live audio output required careful orchestration between the voice model and the avatar rendering layer. If the audio arrived before the lip-sync was ready, the avatar looked broken.

I built a small buffering layer that held the audio for a few hundred milliseconds while the avatar caught up, calibrated to the specific rendering pipeline we used. That buffer was invisible to the learner but load-bearing for the illusion.

Chunking is the highest-leverage decision in RAG

The chunking strategy for the RAG pipeline was another rabbit hole. Chunking too coarsely lost precision — the AI would retrieve a whole page when it needed a sentence. Chunking too finely lost context — the AI would retrieve a sentence stripped of the paragraph it belonged to, and the response would sound plausible but subtly wrong.

I settled on a hybrid: semantic chunking at paragraph boundaries with adjacent-chunk overlap, tuned per document type. Not glamorous work, but it moved the needle on hallucination rate more than any prompt engineering did.

What I learned

RAG quality is entirely dependent on chunking strategy. Chunking too coarsely loses precision, too finely loses context. I spent more time on chunking than any other part of the system, and it was the single highest-leverage decision in the whole build.

Voice AI UX is its own discipline. Latency tolerance, interruption handling, and conversational pacing matter as much as the underlying model quality.

Building reliable AI products isn't about picking the smartest model. It's about mastering the orchestration layer.

The models are commodities. The latency budget, the retrieval quality, the interruption handling, the avatar sync — those are what make the product feel real. Every millisecond mattered.

What I owned

  • Designed and implemented the RAG pipeline — ingestion, chunking, embedding, and retrieval — grounding every AI response in curated training content and eliminating hallucination risk in a high-liability domain
  • Built the voice-enabled AI tutor with streamed STT, concurrent retrieval and generation, and natural interruption handling — hitting a sub-second end-to-end latency budget
  • Integrated an interactive AI avatar for immersive scenario-based learning, synchronized with live audio output via a calibrated buffering layer that kept lip-sync tight without introducing perceptible delay
  • Architected the LangChain and LangGraph workflow orchestrating context retrieval, response generation, session memory, and multi-turn context handoff
  • Optimized chunking strategy per document type — hybrid semantic chunking at paragraph boundaries with adjacent-chunk overlap — to maximize factual precision without losing conversational context
  • Owned the end-to-end latency budget across STT → reasoning → retrieval → TTS → avatar rendering, shaving milliseconds at every layer to keep the interaction feeling human