The profiler was unambiguous: 93.7% of search-worker CPU in the cosine distance, a scalar loop that accumulates in f64 and recomputes both norms plus a square root per call. Obvious win. Wired up the AVX2 SIMD path, then went further -- normalize the vectors at ingest and serve cosine as a plain dot product, skipping the norms entirely. Shipped it to the dev fleet. The math was right, recall even ticked UP, to 0.9867.

Throughput gain: zero. Coord-direct c100 held at appx 1057 QPS, a wash with the old scalar path. I made the distance dramatically cheaper and the system didn't care, because the distance was never the binding constraint -- the ceiling is coordination, the coord-worker round trip, bursty load with every tier periodically idle. Nothing was CPU-saturated. Amdahl has been trying to tell people this since 1967 and I still had to relearn it with a fleet rebuild.

Then the part that stung. In the same head-to-head I finally measured the index I HADN'T been tuning. Vamana, same 1M collection, same c100, serving-confirmed: 1733 QPS at recall 0.9642, against hnsw's 1057. The index I spent a week SIMD-polishing was the slower one by 64% the whole time.

Postscript from production: the auto-tuner reached the same verdict on its own. When the prod 1M collection converged this week, vamana held all three metric slots -- my hand-polished index didn't place. And the cluster that opened July at 48.8 queries per second did 1275 QPS at c50 through the coordinator, p99 under 50ms, zero errors. Nobody hand-tuned anything to get there. The machine did it while I was busy being wrong about hnsw.