THE COLLECTION
EVERY SIGNAL, SORTED.
Search the whole shelf by topic, tool, term, or plain old curiosity.
SHOWING 2 OF 25 STORIES
ISSUE 015EVENING
Research and Evaluation
The benchmark says safety. The questions may be measuring something else.
Ai2’s open BenchMIRT method looks beneath a single score and asks whether an evaluation is testing safety, reasoning, or a muddy mixture of both.
ISSUE 005EVENING
Human Impact
Can we measure whether an AI conversation helps or hurts?
A new grant program puts a harder question on the table: how do you test wellbeing when the risk only becomes visible across a long conversation?