Menu

Categories

Tags

Vals-Smith Puts Your Codebase on the Exam — and Lets You See Which AI

August 14, 2026 | Source: t | AI, Developer | 165 views 0 comments

Vals-Smith uses your own codebase to test AI models — and tells you which one actually fits your team. The new evaluation platform from Vals reads a GitHub repository, pulls real development tasks from merged pull requests, and turns them into a custom coding exam.

Each task comes with a problem description, hidden tests, and a reproducible environment. The model being tested only sees the code as it was before the fix — it never gets a look at the original patch. To count as a pass, the model has to pass the hidden tests without breaking existing functionality.

Teams can run different models and coding agents against the same set of tasks and compare pass rates directly. Public leaderboards show who's strongest overall; Vals-Smith is about finding who's strongest for your specific codebase.

Tags: #Github

Leave a Reply

Your email address will not be published. Required fields are marked *