StoryPropEarly accessStart writing

Writing with an agent

AI story generators vs. writing partners: which do you actually want?

· The StoryProp team

Shop for an AI story writing tool for an afternoon and every product starts to sound the same. The demo is always a sentence blooming into a chapter; the copy always says partner, co-writer, collaborator. If the thing you are making will take months, the machine you want is the partner, and the arithmetic shows why before the craft argument even starts: the longest single response Claude Opus 5 or OpenAI's GPT-5.6 Sol can emit through the ordinary synchronous API is 128,000 tokens — roughly 71,000 to 96,000 words, depending on which vendor's published conversion you use — so a single request does not return a finished 100,000-word novel in one stroke. Under the shared vocabulary there are three genuinely different machines — a generator, an assistant, and a partner — built for different jobs. Most disappointment in this aisle is not a bad tool. It is a category mismatch: someone bought one machine expecting another.

The three converged in vocabulary long before they converged in behavior, so the useful sort is not by label but by mechanism — specifically, the tool's relationship to your project over time. Held for a minute, held for a selection, or held for the months a book actually takes. Those three horizons are the whole taxonomy, and everything below follows from them.

What an AI story generator is actually for

A generator's contract is simple: prompt in, pages out. You describe a premise — a lighthouse keeper finds a door in the sea — and it returns continuous prose, whole scenes at a stroke. An AI novel generator is the same machine at scale: hand it the premise and it will produce an outline, then chapters against that outline, book-shaped text in an evening.

For some jobs this is genuinely the right machine. Exploration, first: when an idea is a shimmer rather than a plan, watching a generator take a run at it shows you one concrete version fast — and often the most valuable output is the flinch, the moment you say no, not like that, and discover in your own objection what you actually wanted. Register, second: seeing the same premise rendered noir, then gothic, then flat and contemporary is a quick way to try on voices before committing to one. And placeholder volume, third, when you need text to think against rather than sentences you intend to keep.

What a generator does not have is any allegiance to your book. Each run is a fresh roll: the machine optimizes for plausible continuation, not for the decisions you made last Tuesday. Ask it twice and you get two chapter fours, each internally confident, each casually contradicting the other. That isn't a defect — it's the design. A generator produces story from a standing start, and a standing start is precisely what a novel in progress never has.

Assistants: polish for prose that already exists

The second machine works at the other end of the process. An assistant takes sentences you have already written and improves them: tightens a paragraph, catches the repeated word, offers three colder versions of a line of dialogue, untangles the clause that got away from you. Its scope is the selection; its horizon is the page in front of it; it is reactive by design, waiting on your prose the way a copyeditor waits on a manuscript.

For that job it is excellent — fast, unobtrusive, low-stakes, the machine equivalent of a good line edit. But notice the quiet assumption underneath: that the hard part is already done. Someone has decided what happens, who knows what, and why the scene exists. An assistant polishes decisions. It does not make them, does not record them, and will not remember them.

Which is why assistants fail politely on structural questions. Ask one whether the reveal in chapter nineteen is earned by what chapter three established and you will get an answer about the prose of chapter nineteen. It cannot see chapter three; nothing in its design ever holds two distant pieces of your book at once. The paragraph is its whole world.

What a writing partner means

The third category is the newest, and it is the one the shared vocabulary borrows from. A writing partner is project-aware: it collaborates over months, not prompts. You direct in plain language — draft the kitchen scene, Mara guarded, keep the cup in frame — and it drafts to your direction rather than to a general model of what stories are like. Crucially, the decisions you settle together persist. What you agreed about a character in March is still constraining the draft in July, without you re-briefing anyone.

Take our demo manuscript, “The Quiet House.” Early on, you decide the chipped cup in Mara's kitchen matters: it was her father's, and Daniel must never learn that directly. A generator cannot hold this; next run, the cup is gone or the secret has become dialogue. An assistant was never told; it only ever sees the paragraph under its cursor. A partner is the machine for exactly this shape of fact — the agreement is written down, and thirty sessions later the draft still walks around it the way you asked.

This is the category StoryProp was built in, and its mechanics show what the label means when it is load-bearing: decisions get promoted to records — canon, characters, story shape, confirmed style rules — and the agent reads those records before it writes a word. Drafting is piecemeal by default, a scene at a time under your direction. Revision is pointed, so asking for one colder line leaves the rest of the page alone, and when a new instruction collides with an old decision, the contradiction is surfaced instead of silently smoothed over. The full feature set is longer, but every item on it is downstream of one commitment: the project, not the prompt, is the unit of work. That is what an AI co-writer for stories has to mean to be worth the word, and it is the standard StoryProp is built to meet — a whole book is too big for anyone to hold alone, and the tool's job is making sure the holder isn't only you.

The failure mode: right machine, wrong job

Each category has a characteristic disappointment, and every one of them comes from expecting a different category's behavior. Use a generator where you needed a partner and the failure arrives on a delay. The first sessions feel miraculous — pages upon pages. Around the fifth chapter the seams show: a character's dead mother turns up at dinner, the timeline folds, the town quietly changes names. Because the machine holds nothing, you become the continuity department, pasting an ever-longer briefing above every prompt, until maintaining the briefing costs more than writing the chapter would have. General chatbots inherit this failure mode for long fiction, and why chatbots lose the plot is a question with published answers rather than a matter of opinion: for all their fluency, a conversation is not a project, and the ledger ends up living in your head.

The numbers behind that are worth having, because they are documented. A manuscript is not small: 100,000 words is about 133,000 tokens using OpenAI's published rule of thumb of roughly 0.75 words per token, or about 180,000 using Anthropic's current figure of 555,000 words per million tokens — the same text, two tokenizers, a 35% spread. The window a chat product gives you is often a fraction of the developer window behind the same model: OpenAI's help center publishes 272,000 tokens for GPT-5.6 Sol on the ChatGPT Business plan, against 1,050,000 for that model in the API, and 128,000 tokens for GPT-5.6 Terra and Luna on that same plan. And the window holds everything, not only your book — Anthropic's context-window documentation counts the system prompt, every prior message, attachments, tool results and the model's own output against it, and notes that recall degrades as the count climbs, an effect it calls context rot. Anthropic's compaction feature, currently in beta, fires by default at 150,000 input tokens: it summarizes what came before and drops the earlier turns from later requests. That is a sound engineering answer to a full window. It is not a record of what you decided.

What has to fit, against the windows that are published to hold it
What it isSize in tokensNote
A 40,000-word novel — the SFWA floor53,000-72,000Range spans the two vendors' published ratios
A 100,000-word manuscript133,000-180,000OpenAI ratio at the low end, Anthropic's at the high end
GPT-5.6 Terra / Luna in ChatGPT (Business plan)128,000OpenAI's published in-product window
GPT-5.6 Sol in ChatGPT (Business plan)272,000About a quarter of the same model's 1,050,000-token API window
Claude Opus 5 on paid Claude.ai plans1,000,000Opus 4.8/4.7/4.6 and Sonnet 4.6 on the same plans: 500,000; other models: 200,000
The longest single response Opus 5 or GPT-5.6 Sol can emit (synchronous API)128,000About 71,000-96,000 words; Opus 5 reaches 300,000 on Anthropic's Message Batches API behind a beta header

Anthropic Claude Platform Docs and Help Center; OpenAI API documentation and Help Center; word floor from the SFWA Nebula Awards rules; word conversions from each vendor's own published ratio, August 2026

Persistent memory narrows the gap without closing it, and both vendors are explicit about what theirs is. OpenAI documents two mechanisms: saved memories you ask ChatGPT to keep, and details inferred from past chats which, in OpenAI's own wording, can change over time — and it advises saving anything that must always be remembered, because ChatGPT does not retain every detail. Anthropic documents memory as individual entries organized into categories, written and updated during a chat, with each Project given its own separate memory space. Both are real features and useful ones. Neither is a manuscript record you wrote deliberately, can read in full, and can rely on the drafting to consult before it writes.

Use a partner where you wanted a generator and the friction runs the other way. You ask for a novel; it drafts a scene and stops. It asks which of two readings of “more tension” you meant. It surfaces a contradiction you would have preferred not to think about tonight. If what you actually wanted was forty pages of anything by morning, that deliberateness reads as slowness — not because the machine is worse, but because you are operating a scalpel like a firehose.

Use an assistant on a draft that needs structural help and you get the subtlest failure of the three: success, of a kind. Every sentence comes back better, and the book stays broken, because the problem was never the sentences. A polished draft with an unearned ending is a broken book that now reads well enough to fool you for one more revision cycle.

Four questions that place any tool

Classification comes from behavior, and four questions are usually enough to put a tool on the map. What does this tool know about my project tomorrow — not what it can do in one sitting, but what survives the sitting ending? Where do my decisions live, and can I read them — agreed facts kept distinct from suggestions still in play? If I ask for one changed line, what else changes, because selection-scoped revision is one category and a page-wide rewrite under a one-line instruction is another? And what is the unit the product is organized around — a prompt box, a selection, or a project? The interface answers that last one at a glance, and the answer is the category.

One deeper question stands behind all four, and it has its own article: whether a tool leaves you directing the story or quietly starts writing it — who is holding the pen — is the subject of whether AI can help without writing it for you. Once you have placed a tool by category and by that question, evaluation is simple: give it a real chapter, come back in a week, and see what it still knows.

So which do you actually want? For a scene, an experiment, a voice test, the generator earns its afternoon. For a finished draft where the sentences are the whole job, the assistant is the right instrument. But the long middle of a book — the months when the thing is too large to hold in your head and too particular to hand to a machine with no memory of it — belongs to the partner, and the long middle is where most books are won. Knowing which machine you are asking for is most of the skill. Ask for the right one and the work moves.

If the machine you keep missing is the middle one — a partner that holds the project while you hold the pen — that gap is the reason StoryProp exists: direction in plain language, decisions written down, and a draft that still remembers March when you reach July.

Sources

Rates, fees and market figures change. These were accurate at the dates shown; check the source for current numbers before you rely on one.

← All posts