Stand up the first OpenAI Evals suite: register a fixed, held-out set of agent tasks, run the agent against each, and score with a deterministic grader. No LLM judges, no humans, no production sampling yet. I care about the unit of evaluation (are you grading the final answer, or the whole tool-call trajectory?) and how you keep the test set honest.