Speech Enabled Reading Diagnostics App  (SERDA) Dataset: Recordings of Read Speech by Dutch-speaking Second and Third Graders

van der Velde, M.E.
Harmsen, W.N.
Tejedor-Garcia, C.
Swart, N.M.
Feskens, R.
Veldkamp, B.P.
Strik, H.
Cucchiarini, C.

This collection contains 177h of child speech recordings that were made as part of the ASTLA project (Advanced Speech Technology and Learning Analytics for child personalized reading education). The recordings were collected in two rounds: from October 2022 to February 2023, and from October 2023 to February 2024. The speakers were 653 Dutch children from grade 2 and 3 in 19 different schools. Each child read two reading tasks using the Speech Enabled Reading Diagnostics App (SERDA) on a tablet. Their speech was recorded using a headset. The first task is a word reading task that consisted of three subtasks in which each 50 words had to be read. The words were presented in isolation, and after reading each word, the children had to tab the screen of the tablet to see the next word. The second task is a passage reading task that also consisted of three subtasks. In each subtask, the children had to read a short passage for a maximum of three minutes. All six subtasks were saved as separate .wav recordings. In total, the collection consists of 1953 recordings of word subtasks and 1928 recordings of story subtasks. In addition, the collection also contains speaker metadata (i.e., Id, data collection round, School-province, School-id, Class, Gender, Birth-province, languages spoken at home, expectation of dyslexia, date of data collection, age). Furthermore, the collection contains the most recently obtained AVI and DMT scores of children, as obtained from children's schools. Finally, the collection contains the reading profiles children were provided with throughout the project.