Removing the vector that represents each document from RRF improved search accuracy

日本語

Contents of this article

Summary

This is a record of an experiment, run on a service developed at my own company for searching long specialized documents, that compared which parts of a document to embed. Three approaches were compared. The main-argument approach, the new one, splits only the main-argument sections, in which the document’s author states the central claims, into chunks and embeds them. The main-argument approach also creates one embedding per document that represents the document. I call this embedding the representative vector. The labeled approach splits every section except the fixed-form boilerplate and the attached materials into chunks. It embeds each chunk with a label for the section type at its start. At search time, it filters down to chunks from main-argument sections. The uniform full-text splitting approach splits the full text of each document into chunks of 800 characters, overlapping by 200 characters, and embeds them.

Through round 3 of the comparison, the main-argument approach always included the list from the representative vector search as the third list in equal-weight RRF. RRF is a method that merges the document lists returned by several searches into a single ranked list, using the sum of values computed from the rank in each list. I call this comparison condition the form with the representative vector always in RRF. Through round 3, the form with the representative vector always in RRF fell below the comparison conditions that measured the labeled approach, on the point estimate of nDCG@10. In round 3, the condition it was compared against was the main-argument-filtered control condition, which measured the labeled approach with the filtering applied as defined. In the final values of the comparison, the form with the representative vector always in RRF scored 0.635 and the main-argument-filtered control condition scored 0.685. The final values are at 768 dimensions and were recalculated after relevance was judged down to the top 10 search results.

The comparison condition added in round 4 does not put the list from the representative vector search into RRF. The rest of its setup is the same as the form with the representative vector always in RRF. I call this comparison condition the form without the representative vector in RRF. The nDCG@10 of the form without the representative vector in RRF was 0.687. The difference from the 0.635 of the form with the representative vector always in RRF was +0.051, and the 95% interval of the difference did not include zero. The difference of +0.051 was calculated from the unrounded values, 0.6865 and 0.6352. The difference from the 0.685 of the main-argument-filtered control condition was +0.002, and the 95% interval of the difference included zero. The experiment records describe a case where the 95% interval of the difference includes zero as statistically equivalent.

The AI agent leading the development judged that the part that had been wrong was the setup that always puts the representative vector into equal-weight RRF. The design document lists two uses that were not measured. One merges the lists with weights, and the other uses the representative vector to rerank only the top results after merging. This judgment does not extend to those two uses.

In the body, I first show the conclusion of the comparison between the form without the representative vector in RRF and the control conditions, in a table of nDCG@10 differences and 95% intervals. Next, I explain the three approaches compared and the search setup of the six comparison conditions. Then I explain the search queries used for evaluation and the relevance judgments. After that, I explain the pass/fail criteria written in the design document before the comparison results were recorded. I then cover rounds 1 to 3 of the comparison and the three High-severity findings on the comparison program. On that basis, I explain where the nDCG@10 difference came from, as isolated in round 4. Next, I explain why the main-argument approach was adopted even though it was statistically equivalent to the main-argument-filtered control condition. Finally, I write about what this experiment did not measure, the comparison on the entire corpus, and three steps usable in experiments that compare search designs.

What you can take away

This article is for people who design or evaluate search that uses keyword search and vector search together. You will learn the following four things.

This article is a discussion of design based on the records of an experiment on a service developed at my own company.