SMALL MODELS. MEASURED.

Intelligence Index.

Five benchmarks. One view of model capability.

INDEX v1.1

Model intelligence

Elo + benchmark accuracy · 1–100 · higher is better

How the index works Weights, scoring & size classes

A transparent weighted average

The index uses 23.75% each of PIQA, HellaSwag, ARC Easy, and BananaMind Base Bench 1.1 normalized overall Elo, plus 5% Arithmark 3 accuracy. Scores are clamped to 1–100; raw zero scores remain visible in exact results. This is a custom index, not an Artificial Analysis score. Arithmark’s small contribution is a weight, not a sampled subset.

Base Bench Elo is normalized using 100 / (1 + 10^((1000 − Elo) / 400)). Elo 1000 maps to 50, 1400 to 90.91, and 600 to 9.09. This fixed reference is a custom index convention, not a measured accuracy or passing grade. The eight profile bars use normalized category Elo for Base Bench and reported percentages for the external benchmarks. Commonsense averages Base Bench and HellaSwag equally. Quantitative uses 90% Base Bench and 10% Arithmark. These diagnostic bars are not averaged to calculate the index. Base Bench context tracking remains included in its overall Elo. Raw accuracy does not contribute to the index.

What is a strong result?

Base Bench uses four choices: chance accuracy is 25%. Each category contains 50 questions, so one answer changes its accuracy by 2 percentage points. Elo reflects difficulty-weighted performance; 1000 is the prior’s center, not a passing grade. Raw overall and category Elo are available below each model’s profile. Small score differences do not establish statistical superiority.

Size-class frontier

A model glows when no scored model in its parameter neighborhood has a higher unrounded index. All eligible registered models count, including hidden models; supplied size-comparison exclusions are respected. Ties glow together. A sole model qualifies but is marked as having no comparison peers.

Model parameters, fromComparison radius
1M±100K
5M±1M
10M±2M
25M±5M
50M±7.5M
100M±15M
125M and above±20M

Thresholds are inclusive. Below 1M, no frontier class is defined. Only models with all four external benchmarks and Base Bench Elo are included. The external-suite parameter count defines size neighborhoods; differing Base Bench counts are shown on model pages. cma-8M and Veyra-30M-Base remain unmatched because their supplied repository links differ. Commented-out source entries are excluded. Missing required measurements never receive an estimated score.

Intelligence Index vs. Parameters

Model size on a logarithmic scale · higher and further left is better

Fewer parameters, higher intelligence┈ Pareto line

Shading uses the selected models’ median parameter count and index. The Pareto line joins models with no equally small or smaller model scoring higher (or equally well with fewer parameters). Size-comparison exclusions are omitted from this chart.