A vector database stores data as lists of numbers that capture its meaning, then finds the items whose numbers sit closest to your search. That is how you can search thousands of photos for 'cats in loaf position' when no photo was ever labelled that way.
The limit of exact matches
A normal database is great for exact facts. Ask for every photo where the cat's name is 'Mochi' or the colour is 'orange', and it answers instantly:
SELECT file FROM photos
WHERE colour = 'orange';That works because 'orange' was saved as a value. 'Loaf position' usually wasn't. You could add a pose column and tag every photo by hand, but that is slow, people describe the same pose differently ('loaf', 'bread cat', 'tucked in'), and it only covers the questions someone thought of in advance. A text search such as LIKE '%loaf%' doesn't help either: it finds the word, not the idea.
What you want is to search by meaning. That is the job a vector database does.
Vectors: meaning as numbers
A vector is just a list of numbers, such as [0.12, -0.83, 0.41, ...]. To turn a photo or a sentence into one, you run it through an embedding model: a machine learning model trained so that things with similar meaning get similar numbers. The vector it produces is called an embedding.
Real embeddings are long, often hundreds or thousands of numbers, and no single number means anything you could name. What matters is where each vector sits compared with the others. Photos of loaf cats end up close together. Photos of jumping or sleeping cats end up farther away.
Searching photos with words
Kitty's search is a sentence, but her collection is photos. That only works if the model puts text and images in the same space, so the words 'a cat in loaf position' land near photos of loaf cats. Models such as CLIP are trained on huge numbers of images paired with captions to do exactly that.
The same rule applies to any vector search: the stored items and the search must be embedded by the same model. Vectors from two different models are like map coordinates from two different maps; comparing them gives nonsense.
How a search works
A vector search has two phases.
Indexing, done once per item:
- Run each photo through the embedding model to get its vector.
- Store the vector in the vector database, with an id that points back to the photo.
Searching, done for every query:
- Run the search text through the same model, so it becomes a vector too.
- Ask the database for the stored vectors closest to it.
- Return the photos those vectors belong to, closest first.
Nothing in that flow looks for the word 'loaf'. The match comes entirely from the vectors being near each other.
Measuring 'close'
There are a few ways to measure how close two vectors are. The most common for embeddings is cosine similarity, which compares the direction the two vectors point in. A score of 1 means they point the same way, and the lower the score, the less alike they are. Other options are Euclidean distance, the straight-line gap between two points, and the dot product. Embedding models usually say which one they were trained for.
Here is the whole idea in a few lines of Python, with toy three-number vectors standing in for real embeddings:
import math
photos = {
"loaf_sofa.jpg": [0.9, 0.1, 0.3],
"loaf_box.jpg": [0.7, 0.2, 0.5],
"jumping.jpg": [0.1, 0.9, 0.2],
"sleeping.jpg": [0.3, 0.1, 0.9],
}
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.hypot(*a) * math.hypot(*b))
query = [0.8, 0.1, 0.3] # "cats in loaf position"
for name, vec in sorted(
photos.items(),
key=lambda item: cosine(query, item[1]),
reverse=True):
print(f"{name:14} {cosine(query, vec):.2f}")It prints:
loaf_sofa.jpg 1.00
loaf_box.jpg 0.96
sleeping.jpg 0.63
jumping.jpg 0.29Both loaf photos come out on top, with no label in sight. In a real system, the model chooses the numbers and there are far more of them, but the ranking works the same way.
Staying fast with millions of vectors
The code above compares the query with every stored vector. That is called a brute-force search, and its cost grows with every item you add: in Big O terms it is O(n) comparisons per search, each over hundreds of numbers. For a few thousand photos that's fine. For millions of documents, it is too slow.
Vector databases solve this with an index built for approximate nearest neighbour (ANN) search. A popular one, HNSW, links each vector to a handful of its neighbours, so a search can hop from neighbour to neighbour towards the query instead of checking everything. The word 'approximate' is the trade-off: the search is far faster, but it may occasionally miss one of the true closest matches. Most databases let you tune how much accuracy to trade for speed.
Where vector databases are used
- Semantic search: finding documents, products or photos that match what someone means, not just the words they typed.
- Recommendations: 'more like this' is a nearest-neighbour question.
- Retrieval for AI assistants: before an LLM answers, the app finds the most relevant chunks of a knowledge base and adds them to the prompt. This is often called retrieval-augmented generation (RAG), and it is a big part of context engineering.
- Finding near-duplicates: two images or support tickets with almost the same vector are probably about the same thing.
You don't always need a separate product for this. Several existing databases can store and search vectors, such as PostgreSQL with the pgvector extension, which is often enough when your data already lives there.
When it's the wrong tool
- Exact questions. Totals, joins, 'every order from last Tuesday' and anything that needs transactions belong in a normal database.
- When 'similar' isn't 'right'. A vector search always returns the closest matches, even when nothing is truly relevant. Set a minimum similarity score, or check results before trusting them.
- When you need both. Real searches often mix the two: orange cats (an exact filter) in loaf position (a similarity search). Many vector databases support filtering on stored metadata for exactly this.
Common mistakes
- Mixing models. Embedding the photos with one model and the query with another makes the distances meaningless.
- Changing models without re-embedding. Switch to a new embedding model and every stored vector must be regenerated.
- Embedding whole long documents. One vector for a 50-page manual blurs every topic into one. Split it into smaller chunks and embed each.
- Treating it as the source of truth. Keep the original data in your main database, and treat the vectors as an index you can rebuild.
Key takeaways
- A vector is a list of numbers that represents meaning, produced by an embedding model.
- Similar items get vectors that sit close together; different ones end up farther apart.
- A search turns the query into a vector with the same model, then finds the closest stored vectors.
- Approximate nearest neighbour indexes keep searches fast at scale, at a small cost in accuracy.
- Use it to search by meaning, alongside a normal database for exact facts.