THE COLLECTION

EVERY SIGNAL, SORTED.

Search the whole shelf by topic, tool, term, or plain old curiosity.

SHOWING 2 OF 25 STORIES
ISSUE 015EVENING

Research and Evaluation

The benchmark says safety. The questions may be measuring something else.

Ai2’s open BenchMIRT method looks beneath a single score and asks whether an evaluation is testing safety, reasoning, or a muddy mixture of both.

#benchmarks#safety#evaluation
ISSUE 005EVENING

Human Impact

Can we measure whether an AI conversation helps or hurts?

A new grant program puts a harder question on the table: how do you test wellbeing when the risk only becomes visible across a long conversation?

#wellbeing#evaluations#safety