Summary
This is a record of an experiment, run on a service developed at my own company for searching long specialized documents, that compared which parts of a document to embed. Three approaches were compared. The main-argument approach, the new one, splits only the main-argument sections, in which the document’s author states the central claims, into chunks and embeds them. The main-argument approach also creates one embedding per document that represents the document. I call this embedding the representative vector. The labeled approach splits every section except the fixed-form boilerplate and the attached materials into chunks. It embeds each chunk with a label for the section type at its start. At search time, it filters down to chunks from main-argument sections. The uniform full-text splitting approach splits the full text of each document into chunks of 800 characters, overlapping by 200 characters, and embeds them.
Through round 3 of the comparison, the main-argument approach always included the list from the representative vector search as the third list in equal-weight RRF. RRF is a method that merges the document lists returned by several searches into a single ranked list, using the sum of values computed from the rank in each list. I call this comparison condition the form with the representative vector always in RRF. Through round 3, the form with the representative vector always in RRF fell below the comparison conditions that measured the labeled approach, on the point estimate of nDCG@10. In round 3, the condition it was compared against was the main-argument-filtered control condition, which measured the labeled approach with the filtering applied as defined. In the final values of the comparison, the form with the representative vector always in RRF scored 0.635 and the main-argument-filtered control condition scored 0.685. The final values are at 768 dimensions and were recalculated after relevance was judged down to the top 10 search results.
The comparison condition added in round 4 does not put the list from the representative vector search into RRF. The rest of its setup is the same as the form with the representative vector always in RRF. I call this comparison condition the form without the representative vector in RRF. The nDCG@10 of the form without the representative vector in RRF was 0.687. The difference from the 0.635 of the form with the representative vector always in RRF was +0.051, and the 95% interval of the difference did not include zero. The difference of +0.051 was calculated from the unrounded values, 0.6865 and 0.6352. The difference from the 0.685 of the main-argument-filtered control condition was +0.002, and the 95% interval of the difference included zero. The experiment records describe a case where the 95% interval of the difference includes zero as statistically equivalent.
The AI agent leading the development judged that the part that had been wrong was the setup that always puts the representative vector into equal-weight RRF. The design document lists two uses that were not measured. One merges the lists with weights, and the other uses the representative vector to rerank only the top results after merging. This judgment does not extend to those two uses.
In the body, I first show the conclusion of the comparison between the form without the representative vector in RRF and the control conditions, in a table of nDCG@10 differences and 95% intervals. Next, I explain the three approaches compared and the search setup of the six comparison conditions. Then I explain the search queries used for evaluation and the relevance judgments. After that, I explain the pass/fail criteria written in the design document before the comparison results were recorded. I then cover rounds 1 to 3 of the comparison and the three High-severity findings on the comparison program. On that basis, I explain where the nDCG@10 difference came from, as isolated in round 4. Next, I explain why the main-argument approach was adopted even though it was statistically equivalent to the main-argument-filtered control condition. Finally, I write about what this experiment did not measure, the comparison on the entire corpus, and three steps usable in experiments that compare search designs.
What you can take away
This article is for people who design or evaluate search that uses keyword search and vector search together. You will learn the following four things.
- The nDCG@10 difference and its 95% interval between always putting the list from a search over embeddings that represent documents into equal-weight RRF and leaving that list out
- What the three High-severity findings on the comparison program were, and when each defect was found
- How to isolate where a difference between comparison conditions comes from, using a comparison condition that changes only one element
- The steps for writing the pass/fail criteria and the minimum thresholds in the design document before recording the comparison results
This article is a discussion of design based on the records of an experiment on a service developed at my own company.
The rest of this article is paid
You can read the rest by buying this article on its own, or with a subscription that covers every paid article.
The paid part is about 26,400 characters, roughly a 53-minute read.
Read just this article
From 300 JPY
Buy this article on its own. The exact price is shown at checkout. Purchased articles stay readable whenever you sign in with the email used at purchase.
Read every paid article
Standard is 490 JPY / month
Four new paid articles ship every month. For yearly billing and the full comparison, see the plans.
Add dialogue and columns
Premium is 980 JPY / month
Premium adds the subscriber-only columns and dialogue with matsumotory-kun on top of every paid article. See the plans for details.
Prices include tax. Purchases and subscriptions start after you sign in. Subscribers and readers who already bought this article can sign in and read the full article; for the plans, see the plans page.
Sales are available in supported regions only; the Terms list where we sell. All charges are in Japanese yen.