Package index
-
as_eval() - Build an Evaluation Object
-
ev_score() - Evaluation Score with a Standard Error
-
ev_cluster() - Cluster-Robust Standard Error for an Evaluation
-
ev_icc() - Intra-Cluster Correlation
-
ev_resample() - Variance Decomposition for Repeated Sampling
-
ev_bootstrap() - Cluster Bootstrap for an Arbitrary Statistic
-
ev_paired() - Paired Comparison of Two Models
-
ev_unpaired() - Unpaired Comparison of Two Models
-
ev_variance_reduction() - Variance Reduction with a Reference Model
-
ev_multi() - Multiplicity Adjustment Across a Benchmark Suite
-
ev_power() - Number of Questions Needed to Detect a Difference
-
ev_mde() - Smallest Difference an Evaluation Can Detect
-
ev_judge_agreement() - Agreement Between a Model Judge and a Human Gold Standard
-
ev_judge_debias() - Debias a Model Judge with a Small Human Sample
-
ev_judge_power() - How Many Human Labels a Debiased Evaluation Needs
-
ev_table() - Collect Results into a Table
-
ev_plot() - Plot Results with Error Bars
-
evaluatellmevaluatellm-package - evaluatellm: Statistical Inference for Language Model Evaluations