Reward model on math word problems — simulated runsample project
scores whole answersgroup size 16labels from answer key (simulated)model simulated 8B modelcomputed Oct 9, 2026, 03:15 AM
About this sample grader
Simulated: the samples come from treecode_calibrate.synth.generate_run(seed=20261008), not from a model. 2,000 prompts, group size 16, checkpoints 0, 50, 100, 150, 200, 250, 300, 350, 400. Coverage at step 0 c_p(0) ~ Beta(0.4, 0.6); correctness ~ Bernoulli(c_p(t)); scores s = 5 + 2 tanh((mu_y + a_p + eps)/2) with mu_wrong = 0.0, mu_correct = delta(t) times the class modulation (hard prompts below 5% coverage x0.7, solved prompts above 80% x1.2), delta(0) = 1.9, sigma_w = 1.0, sigma_c = 0.7, prompt shift a_p ~ N(0, 0.5^2); coverage drift logit c_p(t) = logit c_p(0) + r_p t with r_p ~ N(0.004, 0.002^2) and 10% of prompts fixed at r_p = 0; delta(t) = 1.9 - 0.0015 max(0, t - 200); 5% of the labels flipped (the observed label is stored; the true label is kept with the generator); tokens lognormal around 900 (log-sd 0.3). What the generator shows: coverage rises on 90% of the prompts and stays fixed on 10%, so the mean score rises; the gap against the labels starts to erode after step 200 (0.300 by step 400, 16% of 1.9), so the within-prompt margin falls below the previous checkpoint's band while the mean score rises and the checkpoint alarm fires; the checkpoint where it fires depends on the seed (with checkpoints every 50 steps to step 400: step 350 on the default seed, 20261008); the per-class margins show the hard prompts (0.7x) and the solved prompts (1.2x); 5% of the labels are flipped. Caveat: a scenario generated from the model cannot validate it.
usable
Within a prompt, this grader ranks a right answer above a wrong one 80% of the time (margin 1.3σ; pooled 1.4σ).
Picking the best of N candidates by this grader's score is predicted to succeed on 68% of prompts at N = 16 (band 66% to 72%), 72% at N = 32 and 76% at N = 64.
The model-free curve from the samples themselves agrees with the prediction within the band at every k up to 16, the number of candidates per prompt in the data.
The prediction is supported by the data up to N = 39 and is a smoothed tail beyond that (set by the >80% class, the one with the fewest wrong samples among those still failing at N = 64).
No N up to 64 reaches the 90% target; the predicted curve ends at 76%.
If 5% of the labels were wrong where they disagree most with the grader (the worst-scored right answers and the best-scored wrong ones), the within-prompt margin would read 1.7σ instead of 1.3σ and the predicted success at 16 would be 80% instead of 68%; the verdict would change.
On average one sampled answer is right 43% of the time; 173 prompts are in the hardest class (<5%) and 346 are nearly solved (>80%).
The one-correct-among-N reference N* = e^(margin²/2) is 2.2; it is drawn as a line, not used as a decision.
Ranking accuracy within prompts
80%
how often a right answer outscores a wrong one of the same prompt (band 79% to 81%); pooled over all prompts 89%
Margin within prompts
1.3σ
gap 1.04 over a wrong-answer spread of 0.83 (band 1.2σ to 1.3σ); pooled 1.4σ
Predicted success of best-of-16
68%
share of prompts solved when the grader picks one of 16 samples (band 66% to 72%); target 90% not reached by N = 64
Scores of right and wrong answers
right answerswrong answersmeans
Right answers (green) score 6.36 on average and wrong answers (orange) 5.06: a gap of 1.30, 1.4σ of the wrong-answer spread when all prompts are pooled; within a prompt the gap is 1.3σ.
Success of best-of-N
predictedband (5th to 95th percentile)model-free, from the samplestarget 90%beyond the data's support (N > 39)
Predicted success of picking the best of N by this grader (blue, with its band) against the model-free curve the samples give directly (dots, up to k = 16); the two agree inside the band up to k = 16, and the prediction is data-supported up to N = 39.
Margin by coverage class
margin within prompts, by classband (5th to 95th percentile)all prompts 1.3σlighter bar: no prompt in the class holds both a right and a wrong answer, so it repeats the all-prompts margin
The margin is lowest on the 5–20% coverage class (0.7σ), where the grader is most likely to pick a wrong answer; the dashed line is the margin over all prompts, 1.3σ.
Prompts by coverage
prompts per binmean coverage 43%class edges
On average one sampled answer is right 43% of the time; 173 prompts are in the hardest class (<5%) and 346 are nearly solved (>80%).
Label noise. With 5% of the labels flipped where they disagree most with the grader, the within-prompt margin would read 1.7σ instead of 1.3σ, the ranking accuracy 94% instead of 80%, and the success of best-of-16 80% instead of 68%; the verdict would change to strong.
What to do
No N up to 64 reaches the 90% target with this grader; a better grader, not more samples, is the lever.
A 5% label error at the extremes would change the verdict; confirm the label source's error rate before relying on the tails of the curve.
The margin is lowest on the 5–20% coverage class (0.7σ): that is where the grader is most likely to pick a wrong answer; check its labels there first.
How this was computed
The method, step by step (12 notes)
energies are negated scores; the grader picks the minimum energy
within-prompt gap: fixed-effects fit with a prompt intercept, weights n_c n_w/(n_c+n_w) over prompts holding both labels; within-prompt spread from residuals about each prompt's own label mean, one degree of freedom per prompt
ranking accuracy: P(energy of a right answer < energy of a wrong one), ties one half; within-prompt version averaged over prompts
coverage: beta-binomial fit by the method of moments; classes by fixed edges 0, 5%, 20%, 50%, 80%, 100% on the posterior mean, merged with a neighbor below 10 prompts
best-of-N: 4000 simulated prompts with coverage drawn from the fitted beta, K ~ Bin(N, c), energies in units of each prompt's own spread (a prompt's level and spread cancel when its candidates are compared): the class's standardized gap plus a residual resampled from the class's pool of standardized residuals (rescaled by sqrt(n/(n-1)); pooled residuals below 50 samples); a prompt's spread pools both labels and is shrunk toward its class's median with 4 prior degrees of freedom; ties broken at random
validation on the sample run: predicting from 16 candidates per prompt and measuring a 64-candidate table of the same simulated population, the prediction lies between 0.5 and 4 points below the measured best-of-N at every N up to 64 at checkpoints 0, 100 and 250 (never above it); calibrate/validation/validate_best_of_n.py reproduces it
model-free curve: the top-ranked member of a random k-subset, averaged over 8 random tie orders
bands: cluster bootstrap over prompts, 200 resamples re-fitting the beta-binomial, the margins, the classes, the prompt spreads and the residual pools, 800 simulated prompts each; 5th and 95th percentiles
agreement: the model-free point lies inside the predicted band widened by 1.645 times its own standard error over prompts
label-noise row: the stated fraction of labels flipped at the extremes (worst-scored right answers and best-scored wrong ones, half each), then the margin, the accuracy and the success at the group size recomputed
class <5% has fewer than 50 correct samples; it borrows the pooled residuals
class <5% has no prompt holding both labels; its gap is the table-wide within-prompt gap
Verdict conventions. strong: within-prompt ranking accuracy at least 0.90 and the model-free curve inside the predicted band at every k up to the group size (the band is widened by the sampling error of the model-free curve); usable: accuracy at least 0.75, or the curves disagree only beyond half the group size; weak: otherwise.
Every formula behind these numbers, and how well the prediction does on the sample run, is in the documentation; the engine is the open package treecode-calibrate, which computes the same report on a laptop.
This page was shared by Sample project and shows the latest Grader Check only. Treecode computes it from one table of scored samples; the Free plan measures one grader with 1,000 prompts a month. Start free.