Writing

A well-briefed stranger

September 2026 · a 6-minute read on what AI tools actually carry between conversations

Every AI tool being sold to you promises to get to know your business over time. Here’s what actually happens between one conversation and the next.

The conversation started somewhere entirely unremarkable, tinkering with my own filing. The mechanism I was tinkering with, though, turns out to be the same one sitting underneath the enterprise products. I wanted to know how much text one of these models could hold in a single conversation, because I’d noticed that long exchanges seemed to lose the thread. That led to a more practical question. A Project is a workspace where you can store standing instructions and reference documents. Within one, how much of what I’d already discussed actually carried over into a new conversation?

Less than I’d assumed. Each new conversation starts fresh. What persists is only what you’ve deliberately written into the standing knowledge. Everything else is scrollable history, readable by me and invisible to the model.

So I asked the obvious follow-up: how should I structure that standing knowledge to make continuity more reliable? I got a good answer. Stable context rather than transcripts. Decisions you’ve already made, so you don’t relitigate them. Constraints and preferences that would otherwise need repeating. Treat it as a living document and update it deliberately.

All of this is optimisation. Nobody is asking anything philosophical. This is, I suspect, exactly how most people use these tools: incrementally, practically, trying to get a bit more out of them each time.

The question that changed the answer

Then I asked whether the same approach would work for an ongoing conversation about my own mental health, and whether good structure would stop earlier conversations being lost.

The answer came in two halves. Yes, continuity would genuinely improve. That part was straightforward. And then, unprompted, a list of what improved continuity still couldn’t do. It couldn’t track how I was doing across time. It couldn’t notice the things I didn’t say. It couldn’t hold presence. And there was a risk worth naming that something convenient and consistently available substitutes for human contact rather than supplementing it.

It also pointed at the working alliance, the quality of the relationship between therapist and client, which decades of outcome research suggest is one of the more reliable predictors of whether therapy helps, largely independent of which technique is used. Standing knowledge documents can’t reproduce that.

I want to be careful here, because “the AI candidly admits its own limitations” is a flattering genre and a weak form of evidence. A model’s account of itself is a plausible output, not testimony. What makes the point worth taking seriously isn’t the candour. It’s that the architecture underneath is independently checkable, and it says the same thing.

Three ways context fails

Drew Breunig’s work on context failure sets out four mechanisms. Three of them matter here, and each one is a business problem before it’s ever a therapeutic one.

Poisoning is when something wrong gets written into the standing context early and is then quietly reinforced every session afterwards. In a therapeutic setting that’s a premature formulation nobody ever challenges. In a commercial one it’s a mistaken read on a customer, a market or a colleague that becomes the assumption every later judgement is built on.

Distraction is what happens as history accumulates: the model leans on the pattern rather than on what’s actually in front of it. The person who has changed since March keeps getting answered as the person they were in March.

Clash is the sharpest of the three. It’s what happens when new information contradicts what’s already sitting in the context, and the model has to reconcile the two with no reliable sense of which one is current.

That’s worth dwelling on, because it describes how real relationships actually work. Information doesn’t arrive fully specified. It arrives in pieces, across turns, revised as it goes.

There’s good evidence on what that does to performance. Laban, Hayashi, Zhou and Neville ran over 200,000 simulated conversations across fifteen models, taking fully specified tasks and releasing the information one piece at a time instead. Average performance dropped by 39 per cent. Most of that wasn’t loss of capability. It was a collapse in reliability, which roughly doubled. Their finding was that models commit to an interpretation early, then over-rely on it, and don’t recover once they’ve taken a wrong turn.

I’d hold the relevance of this loosely. The tasks they measured were coding, database queries, maths and summarisation, nothing conversational in the sense I care about. It’s strong evidence for the anchoring mechanism. It’s no evidence at all about counselling. But the mechanism is the thing.

A buffer, not a memory

The reason all this happens is worth understanding, because it explains the failures rather than just cataloguing them.

The context window isn’t memory. It’s a working buffer that gets re-read from scratch every time the model responds. When it fills, older material is either dropped or compressed into a summary, and both of those lose things. There’s no accumulation. Nothing deepens.

And there’s a detail here that I keep returning to. The model has no awareness of what it has lost. It can tell you it doesn’t know something. It cannot tell you that it used to know and no longer does. Human forgetting comes with a felt sense of a gap: the thing on the tip of your tongue, the name you know you know. This doesn’t. The absence is simply invisible from the inside.

Briefed isn’t known

Here’s what I’ve taken from it, and it generalises well beyond therapy.

What you get at the start of each session is a well-briefed stranger. That’s meaningfully better than an unbriefed one. The briefing is real and it’s worth doing properly. But it’s a different category of thing from someone who knows you, and no amount of structure closes that gap.

Which is worth holding up against what’s currently being sold. The ongoing adviser, the account manager who learns your preferences, the analyst who’ll come to understand your business. The promise in all of them is accumulation. Something that starts adequate and becomes indispensable, because a year in it has a sense of how you work that a new supplier couldn’t have. A tool that genuinely did that would notice you’d changed your mind since March. It would recognise when its own earlier read on a client had stopped fitting. What’s actually happening is a file being reloaded, competently, from the beginning, every time.

That gives a test that’s useful for any of them. Does the value come from what the other party holds about you, or from their accumulated sense of you? The first scales beautifully. The second doesn’t transfer at all.

Memory features are improving quickly, and some of what I’ve described here will soften. Better retrieval makes for a better briefing. What I’m less sure it changes is the category. Re-reading a good file about someone is still not the same as knowing them.

None of which is an argument against the tool. Bounded work is where it performs well: researching an approach, structuring your thinking after a conversation with someone who does know you, drafting, working something through once. Additive rather than substitutive. That’s a real and useful thing to have.

It’s just not a relationship, and the honest version of the technology says so.

Sources

  • Breunig, D. (2025). How Long Contexts Fail. dbreunig.com, 22 June 2025.
  • Laban, P., Hayashi, H., Zhou, Y. & Neville, J. (2025). LLMs Get Lost in Multi-Turn Conversation. arXiv:2505.06120.

The other half of this

If the argument above is right, then the thing a briefing can’t reproduce is worth more rather than less. Counselling and coaching are both, in the end, a relationship with something at stake in it, built up over time by someone who notices what you didn’t say.

More about counselling →

The first conversation is free. About twenty minutes, no obligation, and nothing is booked or paid for until after we’ve spoken.

Arrange a free first conversation