Skip to contents

Core

as_eval()
Build an Evaluation Object

Scoring a Single Model

ev_score()
Evaluation Score with a Standard Error
ev_cluster()
Cluster-Robust Standard Error for an Evaluation
ev_icc()
Intra-Cluster Correlation
ev_resample()
Variance Decomposition for Repeated Sampling
ev_bootstrap()
Cluster Bootstrap for an Arbitrary Statistic

Comparing Models

ev_paired()
Paired Comparison of Two Models
ev_unpaired()
Unpaired Comparison of Two Models
ev_variance_reduction()
Variance Reduction with a Reference Model
ev_multi()
Multiplicity Adjustment Across a Benchmark Suite

Planning an Evaluation

ev_power()
Number of Questions Needed to Detect a Difference
ev_mde()
Smallest Difference an Evaluation Can Detect

Model Judges

ev_judge_agreement()
Agreement Between a Model Judge and a Human Gold Standard
ev_judge_debias()
Debias a Model Judge with a Small Human Sample
ev_judge_power()
How Many Human Labels a Debiased Evaluation Needs

Leaderboards

ev_rank()
Bootstrap Rank Intervals for a Leaderboard
ev_elo()
Bradley-Terry Ratings from Pairwise Preferences

Reporting

ev_table()
Collect Results into a Table
ev_plot()
Plot Results with Error Bars

Package

evaluatellm evaluatellm-package
evaluatellm: Statistical Inference for Language Model Evaluations