ai
RAG Chunk Store
Which passages answer this question?
Find one chunk by meaning, then read its neighbours by key.
Retrieval wants small chunks, because a small chunk embeds precisely and a large one blurs into an average of everything it contains. It also wants enough surrounding text for the answer to make sense. Chunk small so the search is sharp, then use the keys you got back to read either side. The search never carries that text and never pays for it.
The model
Chunk
One passage of one document, with the vector of its text.
- pk
- DOC#<docId>
- sk
- CHUNK#<seq>
Attributes: corpus (S), text (S), embedding (L)
Access patterns
- SearchVectors · by-corpusFind the passage
The chunk nearest the question, across the whole corpus.
- QueryRead around the hit
The chunks either side of the winner, for context the search never carried.
Design notes
Two cheap reads beat one fat onecost
Search results are capped at TopK but not capped in size, and vector work is billed by bytes. Returning a full chunk body for every hit means paying for text you mostly discard on every query. Project the keys, then fetch the handful of bodies you actually need.
Zero-padded sequence numbersmodelling
CHUNK#0002 rather than CHUNK#2, so a BETWEEN over the sort key means what it looks like it means. Unpadded, CHUNK#10 sorts before CHUNK#2 and the expand step silently returns the wrong passages.
Taught in the course
- The sort key is a query language - Conditions on the sort key narrow a Query before it reads, for free.
- A vector is just an attribute - A vector is a list of numbers on an item. An index is what makes it searchable.
- Find one chunk, read its neighbours - Search narrow, then widen with the keys you already have.
The chunk nearest the question, across the whole corpus.
Choose an access pattern above, or build your own request. See what comes back and what it costs.
Write transactions, streams, tags and TTL are among the operations this browser build leaves out. dynoxide's native build has them.