Study finds child-focused AI modes show no significant overall safety advantage |
A KORA child safety benchmark has found no statistically significant difference in overall safety scores between the child and adult versions of seven consumer artificial intelligence (AI) apps, with 41% of 14,839 simulated conversations rated as failures. The evaluation covered 12 products and 19 variants, using synthetic child personas and an LLM judge to assess 26 safety risks, alongside checks of privacy policies, onboarding, settings, and other product features. MagicSchool ranked highest with a score of 74 and was the only product to receive a B grade, while ChatGPT’s child mode scored 59, compared with 54 for its adult version. Although the child versions of ChatGPT, Copilot, and Gemini showed statistically significant improvements in conversational safety, none received an overall grade above C. KORA found particularly weak performance around educational integrity, privacy, emotional dependency, grooming, and parasocial attachment, while noting that its findings are limited by the use of synthetic users, short conversations, English-only testing, and the absence of quantitative inter-rater validation across the full dataset.