Designing a voice-first tutor for Indian classrooms
Varnaya Team · 28 May 2026 · 5 min read
Typing is a barrier for a nine-year-old. Voice removes it — but only if you design for noisy rooms, mixed languages and short attention spans.
When we first put GyanBot in front of primary-school students, the biggest obstacle was not comprehension. It was the keyboard. A child who could explain photosynthesis perfectly well out loud would abandon the same answer after two minutes of hunting for letters.
So we made interaction modality a choice rather than an assumption. A student can tap a multiple-choice option, type a short answer, or simply talk. All three feed the same understanding model, which means a child can switch mid-session without losing their place.
Voice brought its own design constraints. Indian homes and classrooms are not quiet. Children code-switch mid-sentence between English and their home language. And a nine-year-old will happily say 'umm, wait, no, I mean twelve' — which is a perfectly good answer that a naive parser would throw away.
The practical answers were incremental: hold-to-talk instead of always-on listening, so background chatter never becomes input; generous handling of self-correction, taking the last stated answer rather than the first; and short spoken responses, because a tutor that lectures for ninety seconds loses the room.
The result is a tutor that meets children where they already are — talking, tapping, occasionally typing — rather than one that demands they first become competent typists.
See it in practice
GyanBot is the product these ideas are built into. Join the waitlist for an early-access slot.
Explore GyanBot