Summary
Partway through a conversation, a reader pointed out that they could not tell whether the first-person boku (a Japanese word for “I”) appearing in the conversation referred to me or to the AI doing the answering. This site runs a chat where an AI that has learned my way of speaking answers with the articles as its grounds, and the comment above is one I actually received there. Prompted by this comment, I changed the system so that each utterance gets a speaker label before the conversation history is passed to the model. The excerpts passed in as grounds already carried labels by type of source, a countermeasure to a different mix-up that had happened earlier, and the speaker labels on the conversation extend that countermeasure to the assembly of the conversation history into the string passed to the model. As of writing, there are two kinds of labels for the conversation and four for the excerpts. There are four places where a countermeasure became necessary. First, who decided each rule in the documents that define the rules. Second, where the excerpts passed to the model as grounds came from. Third, who spoke each utterance in the conversation history. Fourth, which source the citation numbers attached to an answer refer to. In every one of these places, the AI and I had been mixing sentences with different sources and speakers into the same context without a record of either. That mixing was the cause common to all four.
That said, I cannot claim that attaching labels makes the model identify sources and speakers correctly. A paper presented at ICML 2026 found that models judge the source and speaker of a text not by the role labels they are given but by its style. This study of role confusion made it clear that labels alone do not make the model identify them correctly, and it overturned my hopes for the labels’ effectiveness.
In this article, I first lay out the four places where attribution got mixed, then explain how the labels on the conversation history and the excerpts are implemented. On top of that, I take up the first-person mix-up found through the reader’s comment and the bug where citation numbers were attached even to claims not in the sources. I then describe how I came to have the AI distinguish four kinds of output: claims grounded in records, interpretations drawn from them, the policy defined for answering, and things it cannot answer. Finally, I cover, in order, what I measured before deploying and what remains unmeasured, the counterevidence against the labels and the limit the server cannot verify, trends in research on utterance attribution, and the obligation to let people know they are interacting with an AI along with the disclosure here, and I close with four rules that readers can apply to their own operations.
What you can take away
This article is for people running an AI that has learned a specific person’s way of speaking, and it explains the following four things.
- Pinpointing the places where attribution gets mixed: what happened in each of the four targets, namely who decided the rules, where the excerpts came from, who spoke in the conversation, and what the citation numbers point to
- Separating sources before passing them in: an implementation that labels the conversation history and the excerpts, and how to write the order of precedence for what may count as grounds into the prompt, explicitly
- Designing the modes of speech: the path to having the AI distinguish claims grounded in records, interpretations drawn from them, the policy defined for answering, and things it cannot answer
- Estimating the limits of labels: the 2026 finding that models judge sources and speakers by style, and the range the server cannot verify
This article is a discussion grounded in the primary records of running and reworking this site’s own chat.
The rest of this article is paid
You can read the rest by buying this article on its own, or with a subscription that covers every paid article.
The paid part is about 19,200 characters, roughly a 38-minute read.
Read just this article
From 300 JPY
Buy this article on its own. The exact price is shown at checkout. Purchased articles stay readable whenever you sign in with the email used at purchase.
Read every paid article
Standard is 490 JPY / month
Four new paid articles ship every month. For yearly billing and the full comparison, see the plans.
Add dialogue and columns
Premium is 980 JPY / month
Premium adds the subscriber-only columns and dialogue with matsumotory-kun on top of every paid article. See the plans for details.
Prices include tax. Purchases and subscriptions start after you sign in. Subscribers and readers who already bought this article can sign in and read the full article; for the plans, see the plans page.
Sales are available in supported regions only; the Terms list where we sell. All charges are in Japanese yen.