THE COLLECTION
EVERY SIGNAL, SORTED.
Search the whole shelf by topic, tool, term, or plain old curiosity.
SHOWING 2 OF 30 STORIES
ISSUE 026EVENING
Open Models and Evaluation
K2 Horizon shows what radical AI openness actually looks like
The model family arrives with weights, code, checkpoints, logs, and a wonderfully awkward disclosure about benchmark cheating. The mistakes may be the most valuable part.
ISSUE 015EVENING
Research and Evaluation
The benchmark says safety. The questions may be measuring something else.
Ai2’s open BenchMIRT method looks beneath a single score and asks whether an evaluation is testing safety, reasoning, or a muddy mixture of both.