SAIR Benchmark
SAIR's public benchmarks for AI on mathematical reasoning.
The current release is the Equational Theories Benchmark — measuring how well leading AI models decide whether one equation implies another, across problem sets at multiple difficulty
levels and comparable reasoning/temperature settings. Models are accessed via OpenRouter; problem sets, runs, leaderboards, and the evaluation prompt are published on Hugging Face.
More from SAIR