KORA benchmark finds no significant safety difference between child and adult AI modes
The evaluation rated 41% of 14,839 simulated conversations as failing and gave MagicSchool the highest overall score, while relying on synthetic child personas and an LLM judge KORA evaluated child an...
Discussion
Log in to join the discussion