Vector search, visualized.
How Salesforce Data 360 (formerly Data Cloud) turns documents into answers an agent can use: chunking, embeddings, vector search, and hybrid search.
Harmonization and identity resolution handle structured data: rows, fields, and keys. Most of what a company knows lives somewhere else, in policies, manuals, and help articles. Before an agent can answer from them, Data 360 has to cut them into pieces, turn each piece into numbers that capture its meaning, and find the right pieces when a question comes in.
Follow four short help documents from a fictional kitchen store through each step. Every choice you make here changes what the agent sees at the end.
1. Chunking
An agent can't read a whole manual for every question, and one vector can't hold a whole manual's meaning. So each document gets cut into chunks. Where the cuts fall decides what can be found later.
2. Embeddings
Each chunk goes through an embedding model, which turns its meaning into a long list of numbers: a point in space. Chunks that mean similar things land near each other, even when they share no words. Returns sit near refunds. Error codes sit near cleaning, because both are about the blender itself.
Real embeddings have hundreds or thousands of dimensions. This map flattens them to two, and the positions are illustrative, not the output of a real model. Change the chunking above and watch the dots move: a chunk's position comes from everything in it.
3. Vector search
A question gets embedded the same way, and becomes a point on the same map. Vector search returns the nearest chunks. That's why it understands "send it back" means "return," with no words in common.
Pick a question:
Nearest chunks
4. Hybrid search
Vector search has a blind spot: exact identifiers. "E17" and "E22" mean nearly the same thing to an embedding model (both are error codes), so it can't tell them apart. Keyword search is the opposite: it matches exact terms, and misses anything phrased differently. Hybrid search runs both and merges the rankings.
Pick a question:
Keyword half exact terms
Vector nearest in meaning
Hybrid both, fused
✓ has the answer marks chunks that contain what the question needs. Data 360 has no keyword-only index: the keyword column is the keyword half of a hybrid index, shown on its own so you can see what it adds.
What the agent sees
The retriever hands its top chunks to the model as grounding, and the model answers from those alone. If the right sentence didn't make it into a chunk that got retrieved, the model never sees it, however good the model is.
Try it: pick "Can I return a blender I've used?" and drag max tokens down. The rule and the sentence that limits it land in different chunks, and only one may come back.
The short version
Retrieval quality is decided long before the model runs. Chunk on the document's real structure so conditions stay attached to the rules they qualify. Use vector search when people ask in their own words, and hybrid when exact codes matter. Then test with real questions and read what actually got retrieved, not just the final answer.
The same idea runs through the other tools: harmonization and identity resolution decide what structured data an agent can trust. This is the unstructured half.
All content is fictional, the embedding map is illustrative, and the logic is simplified for learning. Keyword scoring (BM25) and the hybrid merge (reciprocal rank fusion) are computed live; Data 360's own scoring differs in detail. An independent tool, not affiliated with Salesforce. More Data 360 tools