Measures functional code generation: the model writes Python functions from docstrings, scored by whether the code passes hidden unit tests (pass@k).