Indexing settings
Precision: Google's guidance is to run this model in bfloat16 or float32 and not float16,
whose range its activations exceed. The ONNX fp16 build used here measures 0.9998 cosine similarity to
float32 (and matched 4-bit within 0.002 in testing), so it is used for the CPU path to halve the download —
pick fp32 if you want the safest option.