Children in an Islamabad classroom taking tablet reading assessments with headphones
All case studies
EdTech HubPakistan

The team that taught a machine to assess children's reading

A data crew in Islamabad. An app studio with roots in South Africa. An AI lab in Sydney. English and Urdu. One hard problem.

EnglishUrduاردو

1,700

Children assessed

26

Schools visited

600k

Recordings marked

90%

Target agreement with humans

The gap

Eighty in a hundred cannot read a simple text.

That is Pakistan on the World Bank's learning-poverty measure. Ten-year-olds. In school. Still unable to read.

The Pakistan Institute of Education, inside the Ministry of Federal Education and Professional Training, cannot close a gap this size unless it can measure it — fast enough to act.

Paper assessment, child by child, is slow. Costly. Hard to run across dialects and accents. By the time the numbers reach the ministry, the ground has often moved. Climate disasters keep moving it.

“Using a general speech model for this is like using a hammer to drive in a screw. It's the wrong tool, built for a much simpler problem.”

Saeed AfsharWestern Sydney University

Most speech AI transcribes. It assumes the speaker said the words correctly. Its only job: write them down.

This job is the opposite. The child will make mistakes. The whole point is to catch them. Across languages. Across accents. In a child's voice.

That is why this was not an off-the-shelf model. Why it took three countries. Why the team tested and set aside more than one architecture before they landed on a model they could trust — small enough to scale.

Picture it. A child. A tablet. A headset. Letters and words, one at a time. The tool records. In real time, it marks: right, or not.

Readup tablet prompt in Urdu: آپ کیا کر رہے ہیں؟
Urduاردو
The prompt on the tablet. Urdu. One item at a time.

Lineage

Australia. Then South Africa. Then Pakistan.

AustraliaSouth AfricaPakistan

Ben Blaine, Neurabuild: it started as a way to automate an assessment with 300 word cards. A hundred normal. A hundred irregular. A hundred made-up. An assessor sat, recorded, listened back, marked. Huge labour. A university proved voice recognition could do it. Neurabuild put it online.

Then the Gates Foundation and the University of Cape Town: can you digitise this for minority African languages? They rebuilt from scratch. Offline on a tablet. Sync when a signal appears. First version in 48 hours. Collecting data on day three.

Pakistan is the next chapter of that same tool.

Alberto Soriano Diaz, the ministry's point of contact, had seen it work in South Africa. “If South Africa can do it, Pakistan can too.” In Punjab alone, foundational-learning programmes reach close to a million children. Make assessment cheaper, and you free money for teaching. The point was never the technology. It was an affordable way to see where children can read — and where they need help.

Three teams

Improvised. Fast. Small budget.

Islamabad
Wyeside Consulting

Walked into 26 schools. Recorded children. Triple-marked every clip.

Sam Wilson, Amina Batool

Cape Town

Built the offline app. Fixed bugs overnight while the field team slept.

Ben Blaine, Graham Withey

Sydney

Trained speech models on the labelled data. Judged right from wrong.

Sergio Chevtchenko, Saeed Afshar

Coordinated with the Pakistan Institute of Education — Dr Zaigham Qadeer and Alberto Soriano — working with the Ministry of Federal Education and Professional Training.

Schoolchildren in Islamabad using tablets and headphones for a reading assessment

The field

Washing baskets. Fifty tablets. Nightly patches.

Twenty-six schools in and around Islamabad. Government schools. City and rural. Boys and girls. English and Urdu. Accents the model had to hear as accents — not as errors.

EnglishUrduاردو

Sam Wilson: “I drove half a mile up a street to find it blocked by a sewage pipe, so we parked and carried the tablets across fields in washing baskets. Every night we'd come home, charge fifty tablets in our bedrooms and living rooms until the place looked like some sort of chaotic data centre, and upload the day's recordings.”

The app updated 10–15 times during fieldwork. Bugs in the evening. Fixes by morning. Graham Withey: “The app had to work offline on the tablets, sync when it could, and not lose a child's recording.”

Marking workbench used to label children's reading recordings

The gold standard

Three markers. Then the experts.

Amina Batool: 1,700+ children. 400,000 recordings. 200,000 of them triple-marked. That is 600,000 markings. “That is a lot of marking.”

Where markers disagreed, local language experts came in. One Urdu letter can make a ‘w’ at the start of a word and an ‘oo’ as a vowel. Two readings. Both correct. Experts sat for hours. What they settled became the gold standard.

Even experts disagreed on accent versus error. That is the size of the problem: train a model to judge pronunciation when there is no single convention that covers every accent.

Saeed Afshar: the human layer never fully disappears. More data means less checking. It never hits zero. “Think of a factory: you still want one person at the end of the line checking what comes off it.”

The model

Agree with humans. Fit on a tablet.

Target: agree with three human markers at least 90% of the time. Not a general speech-recognition model. A model that flags when a word was read wrong.

Several architectures. Most of the labelled data for training. A smaller independent set for testing. Accuracy was not the only goal. The tool has to run offline, in a classroom with patchy connectivity. Eventually on the device itself.

Wav2Vec2 XLSR-53 came out on top for English, and held up on early Urdu. Data2Vec was close — and a fraction of the size. For the edge, that trade-off matters as much as raw accuracy.

One pattern: model accuracy tracked how much the human markers agreed with each other. Where markers argued about accented sounds, the model struggled too. Sergio Chevtchenko: “The AI isn't failing randomly. It's running into the same ambiguity a human listener would.”

On recordings the model had never seen, AI total scores matched human scores across the full range — children who got almost everything wrong, and children who read fluently.

“If you can measure reading early, cheaply and accurately, you can act in time.”

What the ministry needs

Measure learning poverty at the level where you can act.

Dr Zaigham Qadeer, PIE: the Prime Minister declared an Education Emergency for two reasons. Children out of school. And a learning crisis among children who are in. This tool fits the second.

Assessments today run at national, provincial, and regional level. Even there, paper is expensive. Learning poverty plays out district by district. Sometimes school by school. That is the level the data needs to reach. The only realistic way, without the cost becoming prohibitive, is digital — and AI.

The case went to the Ministry and the Economic Affairs Division. Well received. Link it to the National Open Data Portal, not a parallel system. Initial and recurring costs will decide whether this can scale to the districts.

The hard-headed case

A national reading assessment in Pakistan, on paper, for around twenty thousand children, costs over a million pounds. Tablets are cheaper. The data is better. The model cuts subjective error. For a government, cheaper and better is the argument that lands.

Sam Wilson · Wyeside

What happens next

Proof of concept. Not finished.

To complete the whole assessment: about 1.4 million markings. They did 600,000. A targeted slice. Finish it is not complicated. Sit a team down for a month and crunch through the rest.

Amina Batool: embed it where the work already is. Foundational-learning programmes. Every province with a policy to implement. UNICEF on the next phase. A mandate in Punjab for every school to use AI. This tool could be a pillar of any of those.

Building a tool is only the first step. Governments have to own it. Fund it. Analyse it. Act on what it shows.

Saeed Afshar: “Getting as close as we possibly can to the interaction the very best teacher has with a child learning to read. Close that gap, and you change lives.”

The people

The names on the work

  • Dr Zaigham Qadeer

    Director General

    Pakistan Institute of Education
  • Alberto Soriano

    Ministry point of contact

    Pakistan Institute of Education
  • Ben Blaine

    Carried the tool from Australia to Pakistan

  • Graham Withey

    Built and fixed the app in the field

  • Sergio Chevtchenko

    Trains the speech models

  • Saeed Afshar

    Leads the lab

  • Sam Wilson

    Data collection, Islamabad

    Wyeside Consulting
  • Amina Batool

    Marking lead

    Wyeside Consulting

Source

Originally published by EdTech Hub

This page retells the story for Readup. The interviews, figures, and reporting live on EdTech Hub.