Reasoning models (o1, o3, DeepSeek-R1) generate internal chain-of-thought tokens then a summary. Three eval differences: (1) latency/cost asymmetric: reasoning models are 10-30x more expensive and slow; normalize by "task cost" not "answer correctness". (2) Output semantics: the visible answer is much shorter than the work behind it; metrics like "verbosity" lose meaning. (3) Some metrics (exact-match GSM8K) inflate because the model has more compute to find the right answer; passive benchmarks become saturated. Senior recipe: (a) use held-out reasoning-specific benchmarks (MATH, FrontierMath, GPQA Diamond, AIME) for capability comparison; (b) re-evaluate with cost-normalized success rate; (c) add reasoning-trace analyses (model's CoT sampled and judged for logical consistency); (d) monitor for overconfidence: reasoning models fail silently when they hallucinate a wrong intermediate.