smevals - a small eval suite for evaluating models, prompts, and harnesses
smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models.…
Source Simon WillisonPublished 3d ago · Jul 31, 2026Posted on Bluesky

What we hold
- Source
- Simon Willison