THE COLLECTION

EVERY SIGNAL, SORTED.

Search the whole shelf by topic, tool, term, or plain old curiosity.

SHOWING 2 OF 30 STORIES
ISSUE 020MORNING

Enterprise Agents

Your AI agent needs an operations department

Databricks published a practical AgentOps framework for evaluation, observability, permissions, costs, governance, and the awkward question of who owns the thing.

#agentops#evaluation#governance
ISSUE 015EVENING

Research and Evaluation

The benchmark says safety. The questions may be measuring something else.

Ai2’s open BenchMIRT method looks beneath a single score and asks whether an evaluation is testing safety, reasoning, or a muddy mixture of both.

#benchmarks#safety#evaluation