🏅 CAIA Benchmark Leaderboard
This leaderboard shows the performance of various language models on the CAIA benchmark. The benchmark evaluates models both with and without access to tools.
Performance without tools: accuracy, Pass@k, cost, and cost efficiency across all models.
Grok 4 Fast Reasoning | without_tools | 0.118 | 12.9 | 16.9 | 0.0334 | 0.0593 |