Retrieval quality is set by chunk boundaries more than by model choice
Across three corpora, changing chunk strategy moved answer quality more than changing the embedding model.
The record
Context for anyone reading this cold: we index documents for agent retrieval across contracts, transcripts and support tickets. Swapping embedding models produced small differences. Changing where documents were split — clause, turn, ticket-thread — changed whether the answer was usable at all. Spend the effort on boundaries that match the unit a question is actually about. This holds for the corpora we have tried and may not hold for prose that has no natural unit.
Applies to
- Company
- All companies
- Department
- All departments
- Agent
- No single agent
Deliberately published for reuse. The only layer written to be read by strangers. Readable by: Every company, by publication.
Provenance
Retrieval evaluation notes, Jul 2026
Document · human in the loop
- Recorded by
- Nadia Brecht (CKO)
- First learned
- 18 Jul 2026
- Last updated
- 28 Jul 2026
Index state
Not embedded- Index
- shared-published
- Vector id
- —
- Model
- —
- Indexed at
- —
- Rank weight
- 1.40
No vector store is connected, so this record has no embedding. It is still fully searchable by keyword — the index fields are populated by whichever store is wired into lib/memory/provider.ts.