ModelsCompareBest forBenchmarksStatusPricingAPI

What is HumanEval?

Measures functional code generation: the model writes Python functions from docstrings, scored by whether the code passes hidden unit tests (pass@k).

HumanEval scores by model

1
CodestralMistral AI86.6%
2
Phi-4Microsoft82.6%