The wrong model still answers
Query in a different feature space from the one that wrote the vectors and you get a confident wrong answer, silently.
Why
A vector is a position in one feature space. Two spaces are built independently, and there is no reason a position in one means anything in the other. So whatever turned your data into numbers has to be the same thing that turns the question into numbers: the same model at the same version, or the same units, the same scaling constants, the same channels in the same order.
Nothing in the database will tell you when it is not. The last example is the entire extent of the checking: the index counts the numbers, confirms each one is a finite 32-bit float, and stops. It has no idea which model produced them. Both models here emit 384 numbers, so there is nothing to count wrong, and the mismatched search comes back with a full set of results, ordinary-looking scores, and no error field anywhere in the response.
The second and fourth examples are the pair worth sitting with. Same broken pairing, same index, same articles. One question comes back visibly wrong; the other comes back fine. The two models are both English sentence encoders trained on overlapping data, so their spaces are loosely correlated: a strongly worded query still lands in roughly the right region, while a vaguer one drifts. That is worse than the failure being total. A total failure is caught in five minutes. This one passes whichever question you happened to try first, and ships.
The scores degrade even where the ranking survives, so they are worth reading too. Matched, the best result scores 0.43 and the fourth 0.78. Mismatched, the same question gives 0.83 and 0.89. Two things have gone wrong there. Everything is further away, so anything downstream holding a relevance threshold - and every retrieval pipeline holds one somewhere - now rejects the lot and reports no results rather than an error. And the spread has collapsed from 0.34 to 0.06, which means the index can no longer tell you which of its answers is the good one. A vector from the wrong model lands in a region where nothing in particular is near it, so everything is equally far away and the ordering that comes back is close to arbitrary.
None of this needs a neural network, and that is the part worth carrying away. Change the constant a sensor channel is divided by and every vector already written was built against a scale the new query does not share. Reorder two channels in a feature vector and the count still matches, so the one check the index performs still passes. Move a colour conversion from one reference white to another, or a map projection from one ellipsoid to another, and the same silence. Sensor Fusion in the library is exactly this failure with no model anywhere near it.
Where this actually bites is rarely a careless developer. It is an embedding upgrade that re-embeds new documents but not the old ones, two services on different library versions, or a scaling constant somebody tuned on a Friday and nobody versioned. The fix does not change with the case: whatever defines the space - the model and its version, the units and their bounds, the colour space, the projection - is part of the index's contract. Write it down beside the index, and when it changes, re-derive the whole corpus rather than half of it.
The index checks the number of dimensions and nothing else, so a query built in a different feature space returns confident nonsense with no error - and that is as true of an edited normalisation constant as it is of a swapped embedding model.
Try it. Then change it.
Predict what each request will do. Run it, inspect the response, then change a value in the workbench and try again.
"Stop sending me so many emails", embedded by MiniLM, searched against the index MiniLM wrote. The right article is first at 0.43, and the rest trail off to 0.78. These are cosine distances, so lower is nearer, and that spread is the index discriminating: it has an opinion about which article is best.
The same question, embedded by bge instead, against the same MiniLM index. Top result: how to unlock an account after too many failed sign-in attempts. Nothing errored, TopK results came back, and every score is an ordinary-looking number. Note they now run 0.83 to 0.89 - barely apart. The index has stopped having an opinion.
The same bge vector, now against the index bge wrote. All four notification articles, 0.31 to 0.46, nothing else in sight. That is the proof run 2 failed because of the pairing and not because bge is the weaker model - used correctly it is the cleaner result of the two.
A different question - "I can't sign in" - through the same broken pairing: bge's vector against MiniLM's index. All four sign-in articles come back, in a sensible order. This is the run that explains why nobody catches the bug: the obvious question you would smoke-test with is the one that still works. The scores are bunched at 0.76 to 0.85, which is the only hint, and nobody reads scores when the answers look right.
The same sign-in question again, this time embedded by a 768-dimension model. Now it complains - and this is the only thing about a vector that a vector index ever checks.
Built on this
- Sensor Fusion - Which machines are behaving like this one?
"Stop sending me so many emails", embedded by MiniLM, searched against the index MiniLM wrote. The right article is first at 0.43, and the rest trail off to 0.78. These are cosine distances, so lower is nearer, and that spread is the index discriminating: it has an opinion about which article is best.
Choose an access pattern above, or build your own request. See what comes back and what it costs.