Author here, Life lately: build a hard eval → new model mogs it → build a harder eval → mogged again → repeat. Cafe Bench is the latest round this, it is a small eval ( world sim ) we built at Dot to test and track progress of the product.
But the data points transfer well into generalised model benchmakrs.
surprises: Opus 5.5 made +$227k for $8.57 in 22 minutes; GPT-6 Astra made +$157k but cost $56; and GPT-6 Luna lost $60k on average.
Author here, Life lately: build a hard eval → new model mogs it → build a harder eval → mogged again → repeat. Cafe Bench is the latest round this, it is a small eval ( world sim ) we built at Dot to test and track progress of the product.
But the data points transfer well into generalised model benchmakrs.
surprises: Opus 5.5 made +$227k for $8.57 in 22 minutes; GPT-6 Astra made +$157k but cost $56; and GPT-6 Luna lost $60k on average.