For months our recall numbers came out of an internal evaluator I trusted more than I should have. It sampled N vectors from a collection, queried with vectors drawn from the SAME sample, computed brute-force ground truth against that sample, then swept cosine thresholds and reported the best recall any threshold hit. Two problems, both biased upward. Every query had a perfect match in the sample -- itself -- so top-1 recall was 100% by construction. And "best over a sweep" is a max over operating points: you quote whichever number flatters the index most.
Grading my own homework, with a thumb on the scale.
So I stopped. VectorDBBench is Zilliz's harness, the standard one -- ships clients for Milvus, Pinecone, Qdrant, pgvector, and computes recall the canonical way: the dataset's own train array as corpus, the held-out test queries, recall@k against published ground truth, at a fixed index config, no sweep. It had no Vector Panda client, so I wrote one on top of the SDK. We get driven like any first-class database, on the same axes as everyone else.
The numbers moved. Some got less impressive. All of them got real.
