
Educate Girls is a non-profit founded by Safeena Husain that works to mobilise communities for girls’ education in rural and marginalised parts of India. A new book by Husain, Every Last Girl, tells the story of building this transformative movement. An extract from the book below describes how developing an algorithmic model enabled her team to predict the villages with the most out-of-school girls to find them and get them into education.

Predictable girls
What IDinsight did was scrape what is called “predictor variables/data” from publicly available sources for 7,796 Indian villages, chosen because we had been there recently and knew the actual number of out-of-school girls in every village. What were the literacy levels, caste or tribe status, average family size, occupations and school infrastructure like in the villages with the most out-of-school girls? The machine crunched all this data and learnt where the correlations were between the girls’ context and whether they were enrolled in school.
What would really help us predict whether a girl would go to school, or more precisely, predict that a parent was unlikely to enrol their girl in school?
Deciding what predictor data was most relevant for our girls was the next puzzle. What would really help us predict whether a girl would go to school, or more precisely, predict that a parent was unlikely to enrol their girl in school? Was it her parent’s literacy levels? Was it her proximity to the school? Was the parents’ income a factor or the fact that she was from a particular community? What made a girl “predictable” – girl who was predictably out of school?
Lives behind the numbers
As the team started to examine the data, names came flooding back to me. They were data points for the machine, but for me they were all stories. They were all lives. Antimbala is a predictable girl. Her parents are illiterate; they are from an OBC (Jats) in Rajasthan. Antimbala is from a relatively big family – with three daughters born before her and the much – awaited brother Raj Kumar, born five years after her. Her father works away at the brick kiln over the border in Gujarat. Her mother manages the tiny strip of land they rent which provides some basic produce to the family, with only a very small onion surplus to sell. The school in the village is very poorly attended. The girls’ toilet is poorly maintained, the powercuts are long and the secondary school is very far. All this adds up to Antimbala not going to school.

Maafi (“sorry”) is a predictable girl. Sorry she is the eldest of four girls, sorry she is from the Mahyavanshi caste, sorry her father died of a heart attack. Faltu (“useless”) is a predictable girl. But she is far from useless. It is her father, really, who deserves the “useless” label, sadly spending most of the day inebriated. Neither of her parents have consistent work. She is from the Bhil tribe, and without steady income Faltu is usually forced to work in the agarbatti factory. She is the oldest girl in the family, the school is far and she is predictably out of school.
Selecting predictive indicators
To improve the predictions over time, we needed to decide what predictor data we would use. What did we think were the root causes of girls not going to school? What were the barriers to girls accessing their right to education? What were the reasons why Antimbala was at home and not with her brother in the classroom? What did we need to know that might help us predict where to find the most out-of-school girls?
We gathered opinions from education, gender and development experts, combined it with our deep knowledge of our communities and narrowed the predictors down a bit, but still fed over 313 indicators into the machine. All data that was free and accessible and from trusted sources. The most comprehensive dataset was of course the Indian census, and from this we selected 233 indicators. What were the literacy levels in a village? An illiterate parent was undoubtedly a solid predictor that a girl might be out of school. We included religion, caste, employment status and household income.
The size of a household was also a clear indicator – the bigger the family, the higher the likelihood of out-of-school girls.
The size of a household was also a clear indicator – the bigger the family, the higher the likelihood of out-of-school girls. We included the presence (or not) of a toilet in the house, the source of water – tap or well – and we even input what kind of roof a family had over their heads – pukka (permanent) or kutcha (temporary). On top of this, indicators came from the District Information System for Education where data points such as gross enrolment rates, girls’ enrolments, school governance and whether the school included the secondary grades were used. The more data the better.
(Hu)man vs machine
Once we had taught the machine it was then time to test the predictions. In the first year, Vikram – by then our expansion lead and senior manager – and his team already had a set plan of where they wanted to expand our programmes. So, fuelled with some healthy scepticism and a determination not to put all our eggs in one basket, rather than being led by machine learning, they carried on with their survey based on their own contextual knowledge and village selection.
After just one iteration of the model we realised the machine was able to help us to find between 50 per cent more girls in some villages and a staggering 100 per cent more girls in other villages.
Simultaneously Buddy, Ben and Jeff ran predictions on those same villages. When the predictions were shared internally after the survey was completed, the findings surprised everyone. The tool was working much better than anticipated, and after just one iteration of the model we realised the machine was able to help us to find between 50 per cent more girls in some villages and a staggering 100 per cent more girls in other villages.
Over time, the machine has become more accurate at predicting where we will find the most out-of-school girls – where Antimbala, Faltu and Maafi are likely to live. It won’t tell us exactly how many girls we will find in a village. It won’t tell us which house, but what it will do is take a district and give us an idea of where there will be a concentration of villages with high numbers of out-of-school girls. It tells us relatively accurately if a village is likely to have above-average or below-average numbers of out-of-school girls so we can take decisive action about where to work.
So, we avoid villages where the machine says we will likely find less than ten girls who don’t go to school and go where the machine tells us we will likely find more than 50. When we first started working in this way, we were finding on average eighteen girls in every village we decided to go to. After a few years and refinements to the algorithm we were finding close to 42 girls per village – demonstrating that we were getting better at our predictions and going where the problem was greatest.
Excerpted with permission from Every Last Girl: A Journey to Educate India’s Forgotten Daughters, Safeena Husain, HarperCollins India ©. All Rights Reserved.
Note: This extract gives the views of the author, not the position of the LSE Review of Books blog nor of the London School of Economics and Political Science.
Image: Praniket Desai on Unsplash.
Enjoyed this post? Subscribe to our newsletter for a round-up of the latest reviews sent straight to your inbox every other Tuesday.