CHANGELOG

smevals - a small eval suite for evaluating models, prompts, and harnesses

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models.…

Source Simon WillisonPublished 3d ago · Jul 31, 2026Posted on Bluesky

What we hold

Source
Simon Willison