THE COLLECTION

EVERY SIGNAL, SORTED.

Search the whole shelf by topic, tool, term, or plain old curiosity.

SHOWING 2 OF 30 STORIES
ISSUE 026EVENING

Open Models and Evaluation

K2 Horizon shows what radical AI openness actually looks like

The model family arrives with weights, code, checkpoints, logs, and a wonderfully awkward disclosure about benchmark cheating. The mistakes may be the most valuable part.

#open-source#models#benchmarks
ISSUE 015EVENING

Research and Evaluation

The benchmark says safety. The questions may be measuring something else.

Ai2’s open BenchMIRT method looks beneath a single score and asks whether an evaluation is testing safety, reasoning, or a muddy mixture of both.

#benchmarks#safety#evaluation