Why chatbots lose the plot
· The StoryProp team
Every writer who has tried to draft a novel inside a general chat tool knows the moment. Somewhere past the first act, the assistant confidently renames a character, forgets that the letter was burned in chapter three, or offers you a plot development you rejected two weeks ago as though it were new. The tool isn't being careless. It's being exactly what it is: a conversation with a ceiling. OpenAI publishes a 272,000-token context window for GPT-5.6 Sol on the ChatGPT Business plan — about 204,000 words at OpenAI's own rule of thumb of 0.75 words per token, and roughly a quarter of the 1,050,000 tokens the same model gets on the developer API. Everything you have ever told that thread about your book has to live inside that ceiling, alongside the draft it is writing for you.
The conversation is the memory
A general chatbot's working knowledge of your story is whatever fits in its context window: the text of the current conversation, plus whatever you've pasted in. That window is not small any more. Anthropic publishes 1,000,000 tokens for Claude Opus 5 on its API; OpenAI publishes 1,050,000 for GPT-5.6 Sol; Google's model page for Gemini 3.5 Flash gives the literal figure, 1,048,576. Google's long-context documentation illustrates a million tokens as roughly eight average-length English novels. So raw capacity is not the story. What fills the window is — and what fills it is not only your manuscript. Anthropic's context-window documentation counts the system prompt, every message in the thread including documents and images, the tool definitions, and the model's own output for the turn, thinking tokens included, against the same budget. This is as true of general chatbots such as ChatGPT, Claude, or Gemini as of any other conversation-based tool; it isn't a flaw in one product but the shape of the category.
And the window you actually get is usually not the one in the headline. Those are developer-API figures. The published in-product numbers are smaller: OpenAI lists 272,000 tokens for GPT-5.6 Sol inside ChatGPT Business, and Anthropic lists 500,000 tokens for Claude Opus 4.8 on paid claude.ai plans against 1,000,000 for the same model on the API.
| Where you are working | Published context window | Note |
|---|---|---|
| GPT-5.6 Sol — OpenAI API | 1,050,000 tokens | 128,000-token cap on any single response |
| GPT-5.6 Sol — ChatGPT Business plan | 272,000 tokens | About a quarter of the API window |
| GPT-5.6 Terra / Luna — ChatGPT Business plan | 128,000 tokens | Smaller siblings in the same family |
| Claude Opus 5 — Anthropic API | 1,000,000 tokens | Default; no beta header required |
| Claude Opus 5 — claude.ai paid plans | 1,000,000 tokens | Pro accounts need usage credits enabled |
| Claude Opus 4.8 — claude.ai paid plans | 500,000 tokens | Half the API window for the same model |
| Other Claude models — claude.ai paid plans | 200,000 tokens | Anthropic puts this at about 500 pages of text or more |
| Gemini 3.5 Flash — Gemini API | 1,048,576 tokens | The literal published limit, not a rounded million |
Anthropic Claude Platform Docs and Claude Help Center; OpenAI API docs and Help Center; Google Gemini API docs, August 2026
None of which is the deepest problem. A novel project is not just its manuscript. It's the decisions you made and unmade, the facts you established, the endings you tried and rejected, the voice you settled into around chapter five. Not one of those is a passage of prose you can paste back in; they are verdicts. And a transcript records a verdict exactly as faithfully as it records everything else you typed — which is to say, until the compression starts.
When the conversation outgrows the window, tools compress, and Anthropic publishes the mechanism in detail. In its documented compaction feature for developers, the default trigger is 150,000 input tokens — roughly 83,000 words of accumulated conversation, using Anthropic's published ratio of 555,000 words per million tokens for its current-generation tokenizer. At that point the API writes a summary, emits it as a compaction block, and drops every content block before it from subsequent requests. The consumer product describes the same move more gently: Anthropic's help centre says Claude summarises earlier messages to make room for new content and keeps the full chat history available for reference — but it is the summary that made the room, so the summary is what the next turn reads. OpenAI publishes the constraint for its own models, that prompt plus output must fit inside the maximum context length, without publishing what a chat does when it crosses it. Either way, summaries are lossy in precisely the way fiction can't afford: they keep the gist and lose the specifics. The gist of chapter three is that Mara received a threatening letter. The specific is that she burned it — which is why it cannot turn up in a drawer in chapter eleven.
General-purpose memory features help less than you'd hope, because they're built for preferences, not canon. OpenAI documents two separate mechanisms: saved memories, which keep what you explicitly asked ChatGPT to remember, and referenced chat history, which stores what it infers from past chats and which, in OpenAI's words, “can change over time as ChatGPT updates what's more helpful to remember.” OpenAI also states plainly that ChatGPT does not retain every detail from past chats, and advises saving anything that must always be remembered. Anthropic's Claude memory is built as entries organised into categories, with a separate memory space per Project. Google's Gemini learns personal context from past chats. All three are real features; none of them is a canon table. Remembering that you like tea is one kind of memory. Remembering that Daniel cannot know about the lease until the midpoint, because the whole second act depends on his not knowing, is another kind entirely — and it has to be consulted at drafting time, every time, not recalled occasionally.
Watch a summary eat a scene
Here's the failure in miniature, small enough to watch. Say you're forty sessions into a domestic novel — call it The Quiet House — and in one kitchen scene you settle three things with your assistant. Mara has seen the foreclosure notice but says nothing about it. Daniel almost raises it — he gets as far as “about the house” — and she cuts him off by handing him the chipped cup, a deflection the ending is going to rhyme with. And after some back-and-forth you decide that Daniel does not know she's seen the notice: you tried the version where he knows, it flattened the scene, and you rejected it.
Twenty sessions later that exchange has aged out of the window, and what the tool retains is a summary — something like “Mara and Daniel discuss their financial troubles in the kitchen; tension over the house.” Notice what survived: the gist. Notice what died: that they precisely did not discuss it; that the cup is now a loaded object; that Daniel's ignorance is the load-bearing fact, and that his knowing was considered and rejected. So when you ask for the chapter-eleven confrontation, the draft opens with Daniel saying he has known about the notice for weeks — fluent, plausible, and wrong in a way you'll only catch if you still remember the kitchen better than the machine does. Nothing malfunctioned. Summarisation did what it is for: it kept the shape and discarded the specifics, and fiction lives in the specifics.
What breaks, concretely
Continuity drift: physical details, timelines and who-knows-what wander, because nothing distinguishes established fact from passing mention, which is most of the work of keeping a long story consistent. Decision amnesia: choices you made firmly resurface as open questions, because a rejected idea and an accepted one look identical in a transcript. Voice erosion: early style instructions fade as their turns age out of context. And the re-onboarding tax — every serious session starting with you re-pasting and re-explaining — which is where most long projects in chat tools quietly die.
Why re-pasting the manuscript doesn't work
The natural workaround is a ritual: keep the manuscript in a document and open every session by pasting it back in. It fails for three quiet reasons. The manuscript is the smallest part of the project — prose records what happened on the page, not that you rejected the version where Daniel knows, or that the cup is promised to the ending — so pasting the text re-supplies the evidence while losing the verdicts.
The paste also lands in the same window with the same limits, and the arithmetic is unforgiving. SFWA's award rules put the floor for a novel at 40,000 words, and a 100,000-word manuscript works out to roughly 133,000 tokens at OpenAI's 0.75-words-per-token rule — close to half a ChatGPT Business thread's entire 272,000-token budget before you have said a word about it — or roughly 180,000 tokens at Anthropic's current ratio, which is already past the 150,000-token mark where Anthropic's compaction fires by default. The arithmetic also moves under you — Anthropic documents that the tokenizer introduced with Claude Opus 4.7 produces roughly 30% more tokens for the same text than its earlier models, so a manuscript that fitted comfortably last year is about a third heavier this year on the same nominal window. And prose is ambiguous testimony: a tool re-reading your chapters has to infer what's canon from what's merely mentioned, and inference re-opens questions you closed months ago — which is how a settled ending comes back to you as a fresh suggestion.
The structural fix is records, not bigger windows
A larger context window doesn't solve this, and the vendors' own documentation says why. Anthropic names context rot in its context-window docs: as the token count grows, accuracy and recall degrade, so usable capacity is smaller than nominal capacity. Its stated rationale for compaction is the same — response quality degrades as context grows, not merely that the space runs out. A window twice the size is a bigger room with the same filing system, which is none. The fix is changing what the memory is.
A project-aware writing agent keeps the project itself: the full conversation as a permanent, ordered history, and — separately — authoritative records for the things that must not drift. Canon. Characters and what each of them knows. The shape of the story. The style rules you actually confirmed, distinct from patterns merely observed. Every earlier draft, retained. Research kept with its sources and dates, in its own place, so a fact you looked up never quietly becomes a fact about your world.
The discipline that makes records work is the boundary between locked and loose. In StoryProp, what you agree to gets written down; what you don't stays a suggestion. A brainstorm never silently becomes canon, and the agent reads the records before it writes anything — which is what keeps chapter twenty-one honest about chapter two without you supervising the memory yourself. When a new page contradicts a record, StoryProp surfaces the contradiction rather than quietly picking a side.
It also changes what a wrong turn costs. When every draft is kept and a whole agent pass undoes as one action, trying the risky version of a scene is cheap. In a chat tool, the risky version overwrites the good one in the only place either of them existed. And no tool skips the many-passes part: OpenAI and Anthropic both cap a single response on their standard APIs at 128,000 tokens — about 96,000 words at OpenAI's conversion, about 71,000 at Anthropic's — so no one call emits a finished 100,000-word novel from any current flagship. A long book gets assembled piecemeal whatever you use. The only question is whether anything keeps the pieces.
If you stay in a chatbot anyway
If a general chatbot is where your novel lives for now, you can do the record-keeping by hand, and it genuinely helps. Keep a canon sheet outside the chat — one page: the facts that must hold, who knows what as of which chapter, the decisions you've made and, just as important, the ones you've explicitly rejected. Open each session by pasting that sheet instead of the manuscript; a page of verdicts survives the context window far better than three hundred pages of evidence, and it tells the tool what's settled instead of making it guess. When something new gets decided, close the session by updating the sheet yourself — the transcript will not do it for you. One useful closing move: ask the assistant to list the decisions it believes the session made, then correct that list before it goes on the sheet. The places where its account differs from yours are exactly where next session's errors were going to come from.
Beyond the sheet: work in smaller conversations, one scene or one problem each, so the window holds everything relevant to the task at hand. Keep the manuscript in files you control and treat the chat as a place drafts pass through, not where they live — a version that gets overwritten in a chat existed nowhere else. When you revise, paste only the passage in question and say precisely what you want changed; asking a tool to fix chapter six is an invitation to re-imagine chapter six. None of this is unreasonable discipline. It's just that you have become the project's memory — which is exactly the job a writing partner was supposed to take off your hands.
A long story is a system of promises: to the reader, and from one chapter to another. Tools that remember only conversation can't hold promises, because a summary keeps the shape of a promise and loses its terms. Tools built around the project can. That's the entire difference, and it's the one that decides whether the thing gets finished.
If you've built the canon sheet, the decision log, and the paste-in ritual, you already know what the tool should have been doing all along. StoryProp holds that memory natively — a permanent project conversation, decisions promoted to records the agent reads before it writes, every draft kept — so the bookkeeping goes to the machine and the promises stay yours to make.
Sources
- Context windows, single-response output caps, compaction trigger and context rot — Anthropic — Claude Platform Docs, August 2026
- Context windows and summarisation on paid Claude.ai plans; Claude memory entries and per-Project scope — Anthropic — Claude Help Center, August 2026
- GPT-5.6 Sol context window and output cap; the 0.75-words-per-token conversion — OpenAI — API documentation, August 2026
- In-product context windows for ChatGPT Business; saved memories and referenced chat history — OpenAI Help Center, August 2026
- Gemini 3.5 Flash token limits, long-context equivalences, and personalisation from past chats — Google — Gemini API documentation and Gemini Apps Help, August 2026
- The 40,000-word threshold that defines a novel — Science Fiction & Fantasy Writers Association — Complete Nebula Awards Rules, 2026
Rates, fees and market figures change. These were accurate at the dates shown; check the source for current numbers before you rely on one.