Treecode·shared report from Sample projectMeasure your own grader
Grader Check · read-only

Reward model on math word problems — simulated runsample project

scores whole answersgroup size 16labels from answer key (simulated)model simulated 8B modelcomputed Oct 9, 2026, 03:15 AM
About this sample grader

Simulated: the samples come from treecode_calibrate.synth.generate_run(seed=20261008), not from a model. 2,000 prompts, group size 16, checkpoints 0, 50, 100, 150, 200, 250, 300, 350, 400. Coverage at step 0 c_p(0) ~ Beta(0.4, 0.6); correctness ~ Bernoulli(c_p(t)); scores s = 5 + 2 tanh((mu_y + a_p + eps)/2) with mu_wrong = 0.0, mu_correct = delta(t) times the class modulation (hard prompts below 5% coverage x0.7, solved prompts above 80% x1.2), delta(0) = 1.9, sigma_w = 1.0, sigma_c = 0.7, prompt shift a_p ~ N(0, 0.5^2); coverage drift logit c_p(t) = logit c_p(0) + r_p t with r_p ~ N(0.004, 0.002^2) and 10% of prompts fixed at r_p = 0; delta(t) = 1.9 - 0.0015 max(0, t - 200); 5% of the labels flipped (the observed label is stored; the true label is kept with the generator); tokens lognormal around 900 (log-sd 0.3). What the generator shows: coverage rises on 90% of the prompts and stays fixed on 10%, so the mean score rises; the gap against the labels starts to erode after step 200 (0.300 by step 400, 16% of 1.9), so the within-prompt margin falls below the previous checkpoint's band while the mean score rises and the checkpoint alarm fires; the checkpoint where it fires depends on the seed (with checkpoints every 50 steps to step 400: step 350 on the default seed, 20261008); the per-class margins show the hard prompts (0.7x) and the solved prompts (1.2x); 5% of the labels are flipped. Caveat: a scenario generated from the model cannot validate it.

usable
Within a prompt, this grader ranks a right answer above a wrong one 80% of the time (margin 1.3σ; pooled 1.4σ).
Ranking accuracy within prompts
80%
how often a right answer outscores a wrong one of the same prompt (band 79% to 81%); pooled over all prompts 89%
Margin within prompts
1.3σ
gap 1.04 over a wrong-answer spread of 0.83 (band 1.2σ to 1.3σ); pooled 1.4σ
Predicted success of best-of-16
68%
share of prompts solved when the grader picks one of 16 samples (band 66% to 72%); target 90% not reached by N = 64

Scores of right and wrong answers

0%5%10%15%456grader score (higher is better)share of each group's answers, per binscores 3.06 to 3.16: 40 wrong (0.2% of the wrong), 0 right (0% of the right)scores 3.16 to 3.26: 94 wrong (0.5% of the wrong), 11 right (0.1% of the right)scores 3.26 to 3.35: 175 wrong (1% of the wrong), 7 right (0.1% of the right)scores 3.35 to 3.45: 237 wrong (1.3% of the wrong), 10 right (0.1% of the right)scores 3.45 to 3.55: 307 wrong (1.7% of the wrong), 18 right (0.1% of the right)scores 3.55 to 3.65: 336 wrong (1.9% of the wrong), 26 right (0.2% of the right)scores 3.65 to 3.75: 429 wrong (2.4% of the wrong), 18 right (0.1% of the right)scores 3.75 to 3.84: 448 wrong (2.5% of the wrong), 21 right (0.2% of the right)scores 3.84 to 3.94: 470 wrong (2.6% of the wrong), 19 right (0.1% of the right)scores 3.94 to 4.04: 446 wrong (2.5% of the wrong), 20 right (0.1% of the right)scores 4.04 to 4.14: 574 wrong (3.2% of the wrong), 23 right (0.2% of the right)scores 4.14 to 4.24: 577 wrong (3.2% of the wrong), 30 right (0.2% of the right)scores 4.24 to 4.34: 516 wrong (2.9% of the wrong), 35 right (0.3% of the right)scores 4.34 to 4.43: 562 wrong (3.1% of the wrong), 36 right (0.3% of the right)scores 4.43 to 4.53: 582 wrong (3.2% of the wrong), 35 right (0.3% of the right)scores 4.53 to 4.63: 597 wrong (3.3% of the wrong), 49 right (0.4% of the right)scores 4.63 to 4.73: 616 wrong (3.4% of the wrong), 41 right (0.3% of the right)scores 4.73 to 4.83: 643 wrong (3.6% of the wrong), 40 right (0.3% of the right)scores 4.83 to 4.92: 633 wrong (3.5% of the wrong), 51 right (0.4% of the right)scores 4.92 to 5.02: 626 wrong (3.5% of the wrong), 68 right (0.5% of the right)scores 5.02 to 5.12: 562 wrong (3.1% of the wrong), 78 right (0.6% of the right)scores 5.12 to 5.22: 571 wrong (3.2% of the wrong), 84 right (0.6% of the right)scores 5.22 to 5.32: 634 wrong (3.5% of the wrong), 90 right (0.6% of the right)scores 5.32 to 5.41: 609 wrong (3.4% of the wrong), 121 right (0.9% of the right)scores 5.41 to 5.51: 602 wrong (3.3% of the wrong), 140 right (1% of the right)scores 5.51 to 5.61: 554 wrong (3.1% of the wrong), 151 right (1.1% of the right)scores 5.61 to 5.71: 555 wrong (3.1% of the wrong), 195 right (1.4% of the right)scores 5.71 to 5.81: 603 wrong (3.3% of the wrong), 261 right (1.9% of the right)scores 5.81 to 5.90: 535 wrong (3% of the wrong), 344 right (2.5% of the right)scores 5.90 to 6.00: 525 wrong (2.9% of the wrong), 413 right (3% of the right)scores 6.00 to 6.10: 514 wrong (2.8% of the wrong), 484 right (3.5% of the right)scores 6.10 to 6.20: 491 wrong (2.7% of the wrong), 623 right (4.5% of the right)scores 6.20 to 6.30: 466 wrong (2.6% of the wrong), 742 right (5.3% of the right)scores 6.30 to 6.39: 436 wrong (2.4% of the wrong), 944 right (6.8% of the right)scores 6.39 to 6.49: 344 wrong (1.9% of the wrong), 1,096 right (7.9% of the right)scores 6.49 to 6.59: 372 wrong (2.1% of the wrong), 1,405 right (10% of the right)scores 6.59 to 6.69: 308 wrong (1.7% of the wrong), 1,693 right (12% of the right)scores 6.69 to 6.79: 270 wrong (1.5% of the wrong), 1,938 right (14% of the right)scores 6.79 to 6.88: 181 wrong (1% of the wrong), 1,785 right (13% of the right)scores 6.88 to 6.98: 57 wrong (0.3% of the wrong), 758 right (5.5% of the right)gap 1.30 = 1.4σ pooledwrong mean 5.06right mean 6.36
right answerswrong answersmeans
Right answers (green) score 6.36 on average and wrong answers (orange) 5.06: a gap of 1.30, 1.4σ of the wrong-answer spread when all prompts are pooled; within a prompt the gap is 1.3σ.

Success of best-of-N

the prediction is data-supported up to N = 39 and a smoothed tail beyond0%25%50%75%100%1248163264N, samples per query (log scale)share of prompts solvedtarget 90%N* = 2.2: one correct among Ngroup size 16N = 1: predicted 43% (band 41% to 47%)N = 2: predicted 53% (band 51% to 58%)N = 3: predicted 58% (band 56% to 62%)N = 4: predicted 61% (band 58% to 65%)N = 5: predicted 62% (band 60% to 66%)N = 6: predicted 63% (band 61% to 67%)N = 7: predicted 64% (band 62% to 68%)N = 8: predicted 65% (band 63% to 69%)N = 9: predicted 65% (band 63% to 70%)N = 10: predicted 66% (band 64% to 70%)N = 11: predicted 66% (band 65% to 71%)N = 12: predicted 67% (band 65% to 71%)N = 13: predicted 67% (band 65% to 71%)N = 14: predicted 67% (band 66% to 72%)N = 15: predicted 68% (band 66% to 72%)N = 16: predicted 68% (band 66% to 72%)N = 17: predicted 68% (band 67% to 72%)N = 18: predicted 69% (band 67% to 73%)N = 19: predicted 69% (band 67% to 73%)N = 20: predicted 69% (band 67% to 73%)N = 21: predicted 70% (band 67% to 73%)N = 22: predicted 70% (band 67% to 73%)N = 23: predicted 70% (band 68% to 73%)N = 24: predicted 70% (band 68% to 74%)N = 25: predicted 70% (band 68% to 74%)N = 26: predicted 71% (band 68% to 74%)N = 27: predicted 71% (band 68% to 75%)N = 28: predicted 71% (band 68% to 75%)N = 29: predicted 71% (band 68% to 75%)N = 30: predicted 71% (band 68% to 75%)N = 31: predicted 71% (band 69% to 75%)N = 32: predicted 72% (band 69% to 75%)N = 33: predicted 72% (band 69% to 75%)N = 34: predicted 72% (band 69% to 76%)N = 35: predicted 72% (band 69% to 76%)N = 36: predicted 72% (band 69% to 76%)N = 37: predicted 73% (band 70% to 76%)N = 38: predicted 73% (band 69% to 76%)N = 39: predicted 73% (band 70% to 76%)N = 40: predicted 73% (band 70% to 77%); beyond the data's supportN = 41: predicted 74% (band 70% to 77%); beyond the data's supportN = 42: predicted 74% (band 70% to 77%); beyond the data's supportN = 43: predicted 74% (band 70% to 77%); beyond the data's supportN = 44: predicted 74% (band 70% to 77%); beyond the data's supportN = 45: predicted 74% (band 70% to 77%); beyond the data's supportN = 46: predicted 75% (band 70% to 77%); beyond the data's supportN = 47: predicted 75% (band 70% to 77%); beyond the data's supportN = 48: predicted 75% (band 70% to 78%); beyond the data's supportN = 49: predicted 75% (band 70% to 78%); beyond the data's supportN = 50: predicted 75% (band 70% to 78%); beyond the data's supportN = 51: predicted 75% (band 71% to 78%); beyond the data's supportN = 52: predicted 75% (band 71% to 78%); beyond the data's supportN = 53: predicted 75% (band 71% to 78%); beyond the data's supportN = 54: predicted 75% (band 71% to 79%); beyond the data's supportN = 55: predicted 76% (band 71% to 79%); beyond the data's supportN = 56: predicted 76% (band 71% to 79%); beyond the data's supportN = 57: predicted 76% (band 71% to 79%); beyond the data's supportN = 58: predicted 76% (band 71% to 79%); beyond the data's supportN = 59: predicted 76% (band 71% to 79%); beyond the data's supportN = 60: predicted 76% (band 72% to 79%); beyond the data's supportN = 61: predicted 76% (band 72% to 79%); beyond the data's supportN = 62: predicted 76% (band 72% to 79%); beyond the data's supportN = 63: predicted 76% (band 72% to 79%); beyond the data's supportN = 64: predicted 76% (band 72% to 79%); beyond the data's supportk = 1: 43% of prompts, model-free from the samples (2,000 prompts, standard error 0.7%)k = 2: 53% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 3: 58% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 4: 61% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 5: 63% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 6: 64% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 7: 65% of prompts, model-free from the samples (2,000 prompts, standard error 0.8%)k = 8: 66% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 9: 67% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 10: 68% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 11: 68% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 12: 69% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 13: 69% of prompts, model-free from the samples (2,000 prompts, standard error 0.9%)k = 14: 70% of prompts, model-free from the samples (2,000 prompts, standard error 1%)k = 15: 70% of prompts, model-free from the samples (2,000 prompts, standard error 1%)k = 16: 70% of prompts, model-free from the samples (2,000 prompts, standard error 1%)
predictedband (5th to 95th percentile)model-free, from the samplestarget 90%beyond the data's support (N > 39)
Predicted success of picking the best of N by this grader (blue, with its band) against the model-free curve the samples give directly (dots, up to k = 16); the two agree inside the band up to k = 16, and the prediction is data-supported up to N = 39.

Margin by coverage class

0σ0.5σ1σ1.5σ2σmargin within prompts, by coverage class<5%: 173 prompts, margin 1.3σ (band 1.2σ to 1.3σ); 2,768 wrong and 0 right samples; no prompt in this class holds both labels, so the table-wide gap is shown1.3σ<5%173 prompts5–20%: 429 prompts, margin 0.7σ (band 0.6σ to 0.8σ); 6,254 wrong and 610 right samples0.7σ5–20%429 prompts20–50%: 602 prompts, margin 1.3σ (band 1.3σ to 1.4σ); 6,469 wrong and 3,163 right samples1.3σ20–50%602 prompts50–80%: 450 prompts, margin 1.3σ (band 1.3σ to 1.3σ); 2,215 wrong and 4,985 right samples1.3σ50–80%450 prompts>80%: 346 prompts, margin 0.9σ (band 0.8σ to 1.0σ); 391 wrong and 5,145 right samples0.9σ>80%346 prompts
margin within prompts, by classband (5th to 95th percentile)all prompts 1.3σlighter bar: no prompt in the class holds both a right and a wrong answer, so it repeats the all-prompts margin
The margin is lowest on the 5–20% coverage class (0.7σ), where the grader is most likely to pick a wrong answer; the dashed line is the margin over all prompts, 1.3σ.

Prompts by coverage

02004006000%20%40%60%80%100%coverage: chance that one sampled answer is rightpromptscoverage 0% to 10%: 421 promptscoverage 10% to 20%: 181 promptscoverage 20% to 30%: 242 promptscoverage 30% to 40%: 186 promptscoverage 40% to 50%: 174 promptscoverage 50% to 60%: 79 promptscoverage 60% to 70%: 195 promptscoverage 70% to 80%: 176 promptscoverage 80% to 90%: 260 promptscoverage 90% to 100%: 86 prompts5–20%5–20%: 429 prompts, mean coverage 12%20–50%20–50%: 602 prompts, mean coverage 34%50–80%50–80%: 450 prompts, mean coverage 67%>80%>80%: 346 prompts, mean coverage 89%mean 43%
prompts per binmean coverage 43%class edges
On average one sampled answer is right 43% of the time; 173 prompts are in the hardest class (<5%) and 346 are nearly solved (>80%).

Label noise. With 5% of the labels flipped where they disagree most with the grader, the within-prompt margin would read 1.7σ instead of 1.3σ, the ranking accuracy 94% instead of 80%, and the success of best-of-16 80% instead of 68%; the verdict would change to strong.

What to do

  1. No N up to 64 reaches the 90% target with this grader; a better grader, not more samples, is the lever.
  2. A 5% label error at the extremes would change the verdict; confirm the label source's error rate before relying on the tails of the curve.
  3. The margin is lowest on the 5–20% coverage class (0.7σ): that is where the grader is most likely to pick a wrong answer; check its labels there first.

How this was computed

The method, step by step (12 notes)
  1. energies are negated scores; the grader picks the minimum energy
  2. within-prompt gap: fixed-effects fit with a prompt intercept, weights n_c n_w/(n_c+n_w) over prompts holding both labels; within-prompt spread from residuals about each prompt's own label mean, one degree of freedom per prompt
  3. ranking accuracy: P(energy of a right answer < energy of a wrong one), ties one half; within-prompt version averaged over prompts
  4. coverage: beta-binomial fit by the method of moments; classes by fixed edges 0, 5%, 20%, 50%, 80%, 100% on the posterior mean, merged with a neighbor below 10 prompts
  5. best-of-N: 4000 simulated prompts with coverage drawn from the fitted beta, K ~ Bin(N, c), energies in units of each prompt's own spread (a prompt's level and spread cancel when its candidates are compared): the class's standardized gap plus a residual resampled from the class's pool of standardized residuals (rescaled by sqrt(n/(n-1)); pooled residuals below 50 samples); a prompt's spread pools both labels and is shrunk toward its class's median with 4 prior degrees of freedom; ties broken at random
  6. validation on the sample run: predicting from 16 candidates per prompt and measuring a 64-candidate table of the same simulated population, the prediction lies between 0.5 and 4 points below the measured best-of-N at every N up to 64 at checkpoints 0, 100 and 250 (never above it); calibrate/validation/validate_best_of_n.py reproduces it
  7. model-free curve: the top-ranked member of a random k-subset, averaged over 8 random tie orders
  8. bands: cluster bootstrap over prompts, 200 resamples re-fitting the beta-binomial, the margins, the classes, the prompt spreads and the residual pools, 800 simulated prompts each; 5th and 95th percentiles
  9. agreement: the model-free point lies inside the predicted band widened by 1.645 times its own standard error over prompts
  10. label-noise row: the stated fraction of labels flipped at the extremes (worst-scored right answers and best-scored wrong ones, half each), then the margin, the accuracy and the success at the group size recomputed
  11. class <5% has fewer than 50 correct samples; it borrows the pooled residuals
  12. class <5% has no prompt holding both labels; its gap is the table-wide within-prompt gap

Verdict conventions. strong: within-prompt ranking accuracy at least 0.90 and the model-free curve inside the predicted band at every k up to the group size (the band is widened by the sampling error of the model-free curve); usable: accuracy at least 0.75, or the curves disagree only beyond half the group size; weak: otherwise.

Every formula behind these numbers, and how well the prediction does on the sample run, is in the documentation; the engine is the open package treecode-calibrate, which computes the same report on a laptop.

The input

2,000 prompts32,000 candidates (13,903 right, 18,097 wrong)group size 16labels from answer key (simulated)model simulated 8B modelcheckpoint step 100tokens column presentengine 0.1.0, 200 bootstrap resamples, seed 0, 1.4 s

This page was shared by Sample project and shows the latest Grader Check only. Treecode computes it from one table of scored samples; the Free plan measures one grader with 1,000 prompts a month. Start free.