Guide · For learning & training providers · 10 min read
The avatar API is the fast part. Tavus, HeyGen, D-ID, Simli, Anam — any of them gets you a talking, listening avatar in days. What fails learning providers is everything the avatar vendors don’t sell: the learning layer. Six things, each a multi-month build, each discovered mid-rollout. Here’s the breakdown.
The six things that fail providers
Every one of these looks solved in the demo, because a demo is one conversation with a forgiving user. Every one of them fails in a rollout, because a rollout is thousands of conversations against a curriculum. Next to each: the typical build effort for a team doing it for the first time.
The avatar doesn’t know what it’s teaching
Every scenario is a hand-written prompt with no link to course content. Turning a curriculum into conversation designs — per scenario, per CEFR level, with scaffolding for learners above and below target — is content engineering, and it repeats for every course you cover.
The conversation loses the lesson
Raw LLMs drift off the objective mid-conversation. Vendors ship guardrail primitives because of exactly this — but primitives aren’t pedagogy: someone has to author objectives, tune drift control, handle struggling vs. excelling learners, and QA it across every scenario. This is a conversation-control system, not a prompt.
Assessment measures the wrong thing
Out of the box you get tone, pace and filler words — surface metrics a generic model produces for any chat. Competency-grounded scoring — pronunciation assessment, CEFR-aligned per-skill evaluation with evidence a teacher can defend — is the single largest missing piece, and no avatar vendor sells any of it.
Nobody can see what happened
Sessions run; institutions get nothing. Learner feedback views, teacher coaching notes, admin cohort dashboards, competency tracking over time — the reporting stack is what institutions are actually paying for, and it’s entirely your build.
Privacy lands on your desk
Live student audio and video streaming through third-party models. Going direct means your organization signs the DPA, defines retention, answers every institution’s security questionnaire, and audits the vendor — for as long as the product runs.
Production is yours at 2am
The vendor’s SLA covers their pipeline; the system is yours. Load testing for term-start spikes, latency monitoring, incident response, conversation QA, and concurrency capacity bought as annual commitments before you know real utilization. This is a permanent ops function, not a launch task.
Where the year goes
The workstreams overlap, but they can’t fully parallelize — reporting needs assessment, conversation control needs scenarios, and everything needs the pipeline integration stable underneath. For a team building the learning layer for the first time, this is what the calendar typically looks like:
The vendor landscape
None of this is the vendors’ failure. These are strong infrastructure platforms — the point is where their responsibility ends, by design. Descriptions based on public positioning and documentation, mid-2026:
Tavus
full conversational pipeline · BYO LLMAI research lab selling an end-to-end conversational video pipeline: proprietary rendering, turn-taking, and a perception model reading the user’s camera. You plug in your own LLM as the brain. Bundled minutes plus dedicated concurrent streams; enterprise plans are annual commitments.
HeyGen
interactive avatar API · avatar-only or full modesKnown for generated video, with an interactive avatar API for real-time use. Avatar-only (you run STT/LLM/TTS) or fuller pipeline modes. Credit- and minute-based pricing.
D-ID
streaming avatar API · agentsReal-time streaming avatars and an “agents” product with knowledge-base grounding. Usage-based API pricing.
Simli & Anam
lightweight face rendering · assemble-your-ownLower-cost real-time face rendering. You assemble STT, transport, TTS and LLM yourself. Cheapest per minute; most moving parts to own.
Coverage matrix
The same six failures, as a coverage table. Capabilities vary by vendor in the infrastructure column — what’s categorically absent doesn’t.
| Capability | Avatar infrastructure (Tavus / HeyGen / D-ID / Simli / Anam) |
Nousable (learning context layer) |
|---|---|---|
| Real-time avatar rendering & streaming | ✓ included | ✓ managed vendor-substitutable underneath |
| Speech-to-text / text-to-speech / turn-taking | ◐ varies full-pipeline vendors include it; rendering-only don’t | ✓ managed |
| Conversational LLM | — bring your own | ✓ included |
| Scenarios grounded in course contentfailure 01 | — not offered | ✓ core |
| Conversation holds the teaching objectivefailure 02 | ◐ toolkit only authoring & tuning per scenario is on you | ✓ core |
| Pronunciation & competency-grounded assessmentfailure 03 — CEFR-aligned, with evidence | — not offered | ✓ core |
| Learner, teacher & admin reportingfailure 04 | — not offered | ✓ included |
| Privacy posture for educationfailure 05 — who signs the DPA, what’s retained | ◐ yours to own | ✓ transcript-only audio/video not retained |
| Production ops beyond the pipelinefailure 06 | — yours to own | ✓ included |
Going direct is right for teams whose core product is conversational AI and who are staffing for it. Everyone else discovers the raw layer was the easy part — and the year went to the layer above it. That layer is what Nousable sells: your team authors the scenarios, the context layer carries each one through a live conversation and out to assessment, and the infrastructure stays swappable underneath.
FAQ
Is Nousable a competitor to Tavus or HeyGen?
Aren’t the time estimates just scare numbers?
Can’t we just use the vendors’ guardrails and knowledge-base features?
What does “assessment” mean here, concretely?
What happens to student audio and video?
We already started building directly. Is it too late?
Evaluating an avatar build?
Send us your scenario list and target volumes — we’ll map the six workstreams against your team on each route, with your curriculum, not a generic demo.
Get in touch