A student app that carries one learner from Band 3.5 to Band 9.0: a plan built from their own mistakes, practice cut by skill and question type, speaking face to face with an AI examiner, and a mistake notebook that follows them into their inbox and their browser.
Draft for discussion — no scope has been committedOn 06/09/2026 the brief for this product was given in full, verbatim: "Build for me an ielts preparation platform UI/UX and deploy to roadmap.flyer.vn/ielts-v2 — Personalization for students / Gamification, pride / Students can study by idivisual skill / Details onboarding / Knowledge band from 3.5, 4.5, to 9.0 / Ranking between students / Speaking facing with AI Teacher / Details questions filter / Mock Test, recently offical test / Mobile app UI/UX friendly / Email notification to remind to learn vocabulary mistake ( each mistake will be store in students notebook) / Teacher conner (will use Teado.ai) no worries at this stage / Training conner: IPA, learn pronuncation conner / Chrome extention connection (when reading new paper, can open extention to add vocabulary list)". On 11/09/2026 the follow-up was: "tạo thêm prd và cho vào luôn".
The clickable prototype was built first and this document was written after it, by an AI assistant in a single session, describing what the prototype already shows plus the rules needed to build it for real. No user research, no engineering estimate and no content licensing check sits behind it yet.
It is a proposal to argue with, not a decision. Numbers marked proposed have no baseline behind them and exist to be replaced with real ones.
IELTS preparation is bought by learners with a deadline and a number. They know the band they need and roughly when they sit the test, and almost every other decision — what to study today, which question type to drill, whether they are on track — is one they are unqualified to make alone. Most preparation products answer this with a content library and leave the sequencing to the student.
The bet in v2 is that sequencing is the product. Every screen in the prototype exists to answer one of three questions a student asks: where am I, what do I do next, and is it working. The content bank, the AI examiner and the gamification are inputs to those answers rather than features in their own right.
/θ/ in Speaking, a missing article in Writing and a word saved from a news article all become the same kind of object — a notebook entry with a due date. One capture mechanism feeds the daily plan, the review emails and the drills. If that loop works, retention follows it.| Learner | Situation | What they need first | Where they break |
|---|---|---|---|
| Deadline student primary | Test booked in 8–16 weeks, needs 6.5 or 7.0 for university admission | A plan that fits the date, and proof each week that the band is moving | Studies whatever is easiest, avoids Writing and Speaking, discovers the gap in the last fortnight |
| Undated improver | Wants a stronger band, no test booked | A visible ladder and a daily habit small enough to keep | Loses the streak in week two and never returns |
| Class student | Studying with a teacher who works in Teado.ai | Homework in the same place as self-study, and one band estimate both sides trust | Teacher tooling and student tooling disagree about level |
| Re-taker | Sat the test, missed by 0.5 in one skill | To spend nearly all time in the one skill that failed | Restarts a general course instead of a targeted one |
The prototype is drawn for the deadline student and degrades sensibly for the others: an undated learner skips the date step in onboarding and gets a plan by band rather than by week.
All targets below are proposed. None has a measured baseline in this product yet, since nothing is built; the first job after launch is to replace this table with real numbers rather than to defend these.
| Question | Measure | Target | Read it as |
|---|---|---|---|
| Does onboarding land? | Students finishing all 8 steps including placement | 70% | Activation |
| Do they come back? | Active on 4+ days in their first 14 | 45% | Habit formed |
| Does the band move? | Median overall band change after 8 weeks of use | +0.5 | The only outcome that matters |
| Do they face Speaking? | Students with 3+ AI Teacher sessions in 30 days | 50% | The avoided skill |
| Does the notebook loop close? | Due items reviewed within 48 hours | 60% | Mistake capture is worth the build |
| Do the emails earn their place? | Review email open rate, and unsubscribe rate | 35% / <2% | Reminder, not spam |
| Is the estimate honest? | Gap between predicted band and real test result | ±0.5 | Trust in everything else |
Every line of the original brief, and where it is answered in the prototype. "Drawn" means the screen exists and is clickable with sample data; nothing in the prototype is wired to a real backend.
| From the brief | Answered by | State |
|---|---|---|
| Personalization for students | Onboarding answers drive the Today plan; mock results and notebook mistakes re-weight it; profile shows performance over 7/30/90 days and ranked weaknesses | Drawn |
| Gamification, pride | Streak, XP, coins, leagues, 12 badges, earned titles beside the name, shareable result card | Drawn |
| Study by individual skill | Four skill hubs with per-question-type mastery and a band path each | Drawn |
| Detailed onboarding | 8 steps: goal, target band, test date, level, placement, daily time, reminders, plan | Drawn |
| Knowledge band 3.5 → 9.0 | 12-rung ladder with locked, current and mastered rungs; every unit and question carries a band tag | Drawn |
| Ranking between students | League, class, target-band peers, friends, global, plus a public profile | Drawn |
| Speaking facing with AI Teacher | Examiner picker, live call with cue card and transcript, feedback by the four criteria | Drawn |
| Detailed questions filter | Skill, part, question type, band, topic, source, status, time | Drawn |
| Mock test, recent official test | Official recent tests by month, Cambridge full mocks, per-skill mocks, in-test screen, result report | Drawn · content source unresolved |
| Mobile app UI/UX friendly | Bottom tab bar under 720px, filters as sheets, phone frame toggle in the prototype bar | Drawn |
| Email reminder for vocabulary mistakes | Notebook with spaced repetition, plus a rendered review email with an inline quiz | Drawn |
| Teacher corner (Teado.ai) | Placeholder plus class-code join; deliberately out of scope for now | Phase 2 |
| Training corner: IPA, pronunciation | 44-sound chart, per-sound record and compare, minimal pairs, word stress | Drawn |
| Chrome extension for vocabulary | Popup over an article, saved lists, sync into notebook and emails | Drawn · extension itself not built |
The ladder is the spine of the product. It is twelve rungs — 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0 — and it does three jobs: it tags content, it gates content, and it reports progress.
| Rung | Listening | Reading | Writing | Speaking |
|---|---|---|---|---|
| 4.5 | Short factual exchanges, form completion | Short texts, short answer | Sentence-level accuracy, 150 words | Answers of one or two sentences |
| 5.5 | Part 2 monologues, map labelling | Summary completion, TFNG at 60% | Paragraph structure, linking, 250 words | Fluent on familiar topics, short long-turn |
| 6.5 | Part 3 discussion, paraphrase tracking | Matching headings, YNNG reliably | Clear position, developed body paragraphs | Two-minute long turn without long pauses |
| 7.5 | Part 4 lectures, note completion at speed | Dense academic text, inference | Collocation range, cohesion without over-linking | Abstract opinion, hedging, natural stress |
Onboarding collects five inputs, and each one changes something visible. If an answer changes nothing, the question should be cut.
| Input | Collected in | What it changes |
|---|---|---|
| Goal (study abroad, work, school, self) | Step 2 | Academic or General module; topic weighting; which Task 1 type is taught first |
| Target band, overall and per skill | Step 3 | The top of the ladder; what counts as "on track"; difficulty ceiling in the question bank |
| Test date | Step 4 | Plan length, weekly intensity, when mocks are scheduled |
| Starting level (placement or self-report) | Steps 5–6 | The starting rung per skill |
| Daily time and reminder settings | Step 7 | Size of the daily plan; push and email cadence |
The plan is rebuilt each morning. It holds two to four tasks sized to the student's declared minutes, chosen in this priority order:
Every task states its reason in one line: "Because your last essay scored 5.0 on Coherence". A plan whose reasoning is hidden is indistinguishable from a random content feed, and students treat it as one.
Each skill hub shows the current rung, the mastery of every question type inside that skill, and the band path. Mastery is accuracy across the last twenty attempts of that type, which keeps it responsive without being noisy.
| Skill | Parts | Question types the bank must tag |
|---|---|---|
| Listening | 4 | Form completion, multiple choice, map labelling, matching, sentence completion, note completion |
| Reading | 3 passages | True/False/Not Given, Yes/No/Not Given, matching headings, matching information, summary completion, multiple choice, short answer |
| Writing | 2 tasks | Task 1 chart, Task 1 process or map, Task 2 opinion, discuss both views, problem/solution, advantages/disadvantages |
| Speaking | 3 parts | Part 1 topics, Part 2 cue types (person, place, event, object), Part 3 abstract discussion, pronunciation focus |
The filter panel is the difference between a library and a practice tool. All eight dimensions combine, and the active set is shown as removable chips so a student can see why a list is short.
Three products sit under one roof, and they are not the same thing:
A completed mock produces a result report with an overall band, four skill bands, a breakdown of where marks were lost by question type, an updated plan, and a shareable card. The report is the moment the product earns trust or loses it, so it shows the student's own wrong answers next to the correct ones, with the passage or transcript that settles the argument.
Speaking is the skill students avoid, so the design goal is to make starting a session feel closer to a video call than to a recording exercise. The student picks an examiner — accent and manner differ — picks a part or a full test, and is then in a call: examiner on screen, their own camera in the corner, question asked aloud, live transcript running underneath.
| Stage | What happens | Requirement |
|---|---|---|
| Choose | Examiner, part, topic; weak sounds offered as a warm-up | Four examiner personas, at least three accents |
| Interview | Examiner asks, student answers, follow-ups adapt to the answer | Response latency low enough to feel conversational; barge-in handled |
| Long turn | Cue card, one minute preparation, two minutes speaking | Real timer, notes allowed, no interruption during the turn |
| Feedback | Band by the four official criteria, with evidence per criterion | Every claim cites the moment it came from |
| Capture | Errors written into the notebook | Pronunciation, grammar and lexis, each typed correctly |
Feedback reports Fluency and coherence, Lexical resource, Grammatical range and accuracy, and Pronunciation, each as a band with a one-line justification pointing at real evidence — "/θ/ produced as /s/ three times", "no complex sentences in the two-minute turn". A band without evidence is not shown.
One notebook, four sources, three kinds of entry. Everything the student gets wrong lands here without them doing anything.
| Source | Captures | Entry type |
|---|---|---|
| Reading and Listening | Wrong answers, plus words tapped in a passage | Vocabulary |
| Writing | Spelling, grammar and collocation slips from AI marking | Vocabulary, Grammar |
| Speaking | Mispronounced sounds, wrong word stress, grammar errors in speech | Pronunciation, Grammar |
| Chrome extension | Words saved while reading anything on the web | Vocabulary |
Each entry keeps the word or rule, its IPA, its meaning, what the student actually wrote or said, and the sentence it came from. The wrong version matters: it is what makes the entry theirs rather than a dictionary line.
Spaced repetition with four stages — 10 minutes, 1 day, 3 days, 7 days — then mastered. A failed review drops the entry back one stage rather than to the start.
Two different motivations are being served and they should not be confused. Gamification is private and keeps a daily habit alive. Pride is public and gives a reason to tell someone else.
| Mechanic | Earned by | Serves |
|---|---|---|
| Streak | Completing the daily plan | Habit; the one thing a reminder can protect |
| XP | Every finished task, weighted by difficulty and by whether it was avoided work | Progress inside a week |
| Coins | Milestones | Reserved for a reward mechanic not yet designed |
| Leagues | Weekly XP, thirty students per league, top five promote | Competition with peers of similar effort |
| Badges | Twelve defined achievements | Collection, and nudges toward avoided behaviour |
| Titles | Sustained behaviour, shown beside the name in rankings | Public identity |
| Share card | Band improvement or a mock result | Pride, and the cheapest acquisition channel available |
Five boards, because one global leaderboard tells a Band 5.0 student only that they are losing: league, class, students with the same target band, friends, and global. The peer board is the useful one — it shows what students who started where you are actually practise.
A full chart of the 44 sounds — twelve monophthongs, eight diphthongs, twenty-four consonants — where the student's own weak sounds are marked from real speaking sessions rather than from a generic list for Vietnamese learners.
Each sound opens to: how the mouth makes it, a short mouth video, minimal pairs to hear the contrast, a record-and-compare with a score and a named failure ("/θ/ heard as /s/ in think"), and the words from the student's own notebook that contain it. Word stress and sentence rhythm are separate drills, because four-syllable stress errors and flat sentence rhythm cost Pronunciation marks independently of individual sounds.
The loop that makes this worth building: AI Teacher flags a sound, the sound appears in the training corner as weak, the drill uses the student's own vocabulary, and the next speaking session measures whether it moved.
Students who are ready for Band 7.0 vocabulary read English outside the product. The extension makes that reading count.
Deliberately minimal, as the brief asked. The student side holds a class-code join and a placeholder explaining what appears once a teacher connects. Teachers themselves work in Teado.ai.
What phase 2 has to settle: assignments with due dates that auto-mark Listening and Reading and pre-mark Writing and Speaking for teacher review; a class ranking; a teacher override on any AI band, with the override visible to the student; and attendance. The hard part is not the UI. It is agreeing one band estimate that both the student app and the teacher app display, so that a student is never told two different things about their own level.
#6A309B, ink #1b1b1f, lines #e4e4e7, 8 and 10 pixel radii, flat surfaces, no drop shadows.| Entity | Key fields | Feeds |
|---|---|---|
| Student | Goal, module, target band overall and per skill, test date, daily minutes, reminder settings, streak, XP, league | Plan, reminders, ranking |
| Band estimate | Skill, band, confidence, evidence count, computed at, stale flag | Ladder, plan, profile, report |
| Question | Skill, part, type, band, topic, source, expected minutes, answer key | Bank filters, mocks, mastery |
| Attempt | Question, student, answer given, correct, seconds taken, flagged, context (practice or mock) | Mastery, weakness ranking, anti-gaming |
| Notebook entry | Type, term, IPA, meaning, student's wrong version, source sentence, origin, stage, due date | Review, emails, drills, extension |
| Speaking session | Examiner, part, audio, transcript, four criterion bands, evidence, errors extracted | Feedback, notebook, training corner |
| Mock result | Test, four bands, overall, per-type loss breakdown, duration | Report, band history, plan |
Two fields carry more weight than their size suggests: the student's wrong version on a notebook entry, and evidence on a band estimate. Both are what make the product feel like it is paying attention rather than guessing.
Ordered so that each phase is usable on its own. A student could stop at the end of any phase and still have something worth opening.
| Phase | Contains | Usable because | Blocked by |
|---|---|---|---|
| P0 done | Clickable prototype, 34 screens, both layouts | Scope is now arguable against something real | — |
| P1 | Onboarding, placement, band ladder, question bank with filters, skill hubs, daily plan | A student can practise deliberately and see their level | Tagged content bank; placement item set |
| P2 | Mistake notebook, spaced repetition, review emails, Writing AI marking | The mistake loop closes and brings students back | Email infrastructure; marking quality bar |
| P3 | AI Teacher speaking, feedback by criterion, training corner | The avoided skill becomes practisable | Voice stack choice; cost ceiling; latency budget |
| P4 | Mocks and official recent tests, result report, band history | Students can rehearse the real thing | Content licensing decision |
| P5 | Leagues, badges, titles, ranking boards, share cards | Retention and word of mouth | Enough students for a league to be non-empty |
| P6 | Chrome extension, Teado.ai teacher corner | Study extends outside the app and into class | Extension review; shared band estimate with Teado |
| # | Decision | Why it cannot wait | Owner |
|---|---|---|---|
| 1 | Where "official recent test" content comes from | Blocks P4 entirely, and the answer may change what the category is called | Content + legal |
| 2 | Voice stack, per-session cost ceiling, fair-use limit | Decides whether speaking is unlimited or metered, which changes the pricing page | Product + engineering |
| 3 | How v2 relates to IELTS Guru: successor, redesign of its student side, or separate product | Decides authentication, student identity, whose content bank is used, and how Teado connects. Everything else in this document assumes an answer | Product |
| 4 | English only, or English and Vietnamese at launch | Cheap to decide now, expensive to retrofit across 34 screens | Product |
| 5 | The accuracy bar AI Writing marking must clear before students see a band | A wrong band is worse than no band; this bar gates P2 | Academic |
| 6 | Who owns the tagged content bank, and how many questions P1 needs | Every filter, mastery score and band estimate depends on tagging | Content |
| 7 | Free and paid boundary | Changes which screens need an upgrade path drawn | Business |
Related: clickable prototype · IELTS Guru PRD (shipped product) · onboarding walkthrough · all 34 screens and the requirement map · Cloud theme · Teado teacher app PRD · all PRDs