Each skill measures an axis the models struggle to reproduce. Candidate scores are pitted against the latest versions of the public models, on the same challenges.
A résumé won’t tell you whether someone can infer a rule from three examples, stay sharp when the environment changes its rules, or read the real emotion behind a polite message. These are exactly the grounds where today’s AI models stumble — and therefore where a human still makes the difference. Our challenges measure these skills directly, through real scenarios, not a self-reported questionnaire.
6 skills available · 6 challenges in the catalog