Grok 4.5 can take in 1 million tokens for $2 and return 1 million for $6. At those prices, producing prose can cost less than deciding what that prose should be allowed to do.

Key takeaways

  • Cheap, persistent generation shifts the bottleneck from producing drafts to governing which outputs may ship, need revision, or require human approval.
  • Editorial standards become operational infrastructure when encoded in agent instructions, evaluation criteria, review checkpoints, and rule records—not merely placed in a style guide.
  • Orwell-style rules are useful because they translate subjective preferences into model-agnostic instructions and auditable review questions, while allowing precision, safety, and comprehension to override brevity.
  • Human review should focus on ambiguous, sensitive, high-stakes, or irreversible communication; rewriting every output eliminates the economic benefit of autonomy.
  • Editorial controls improve clarity and accountability but cannot establish factual correctness or validate the underlying evidence and decisions.

The first wave of generative AI treated writing as an output problem: ask for text, receive text, edit the text. The emerging AI-agent workflow is different. An agent receives a task, acts across tools, produces multiple artifacts, and continues with limited supervision. The resulting language moves out of the chat window and into the interface between a team and its customers, colleagues, software, and institutions.

As generation becomes cheap and persistent, the bottleneck shifts from getting a first draft to deciding which drafts may pass, which must be revised, and which require an accountable human decision. The scarce capability is the editorial control plane around the model.

Autonomy turns prose into an operating surface

Coding agents advanced from suggestions toward complex projects with minimal oversight during the months leading into 2026. Those agents also produce instructions, explanations, status reports, requests, decisions, and handoffs around their work. That prose is operational: a vague sentence can send the next actor down the wrong path as surely as a faulty line of code.

Cursor is reportedly building Sand, a general-purpose agent for non-developers that would handle emails, texts, and documents while competing with Claude Cowork. The July 12–13 reports were rumors, not confirmation, so Sand signals direction rather than adoption. Yet the proposed product follows the same architecture: a company built around coding workflows is considering agent orchestration for ordinary communication.

Developers are also being advised to make software API-first because agents may increasingly become its primary users. Once agents can act through software rather than describe what a person should click, instructions become executable inputs. “Write clearly” is then as underspecified as “handle the customer.” The system needs definitions, constraints, tests, and escalation rules.

Longer-running agents push code, documents, and communication toward the same design: teams encode more judgment before execution because supervising every intermediate output erases the economics of autonomy.

Cheap tokens move value from drafts to constraints

Grok 4.5 launched in Grok Build, Cursor, and the SpaceXAI console. Cursor and SpaceXAI trained the model together, joining model capability to the environment in which teams direct it.

Grok 4.5 price per 1M input and output tokens

At those prices, a team can produce another draft for little money. The costly work surrounds it: choosing the right context, defining the audience, constraining tone, checking claims, detecting omissions, and deciding when the system must stop. A cheaper model does not remove those tasks. It increases the number of outputs they must govern.

Teams first use a cheaper input to do the old task faster. Then they redesign the workflow around abundant drafts and scarce judgment. Editorial standards matter because teams can encode recurring decisions once instead of asking a senior editor to rediscover the same preference sentence by sentence.

Prompt wording by itself is not a durable moat. Research suggests models can perform much of their own prompt engineering, so clever incantations lose value as a differentiator. Teams instead build a durable harness from institutional context, persistent instructions, evaluation criteria, review checkpoints, and records of which rules governed an output; the prompt serves as one control inside that workflow.

Orwell’s rules work because they can be executed

Applied to agents, George Orwell’s six rules turn literary advice into specification design. They convert tacit preferences—avoid clichés, use short and concrete language, cut needless words, prefer active constructions, and break a rule when obedience would produce uglier prose—into instructions that teams can repeat and audit.

A compact guide works across models because it describes desired properties of the output rather than a model-specific trick for obtaining them. It gives the agent a target and the reviewer a test. “Make this better” asks a probabilistic system to infer taste. “Replace stale phrases, identify the actor, remove words that do not change the meaning, and preserve necessary qualifications” makes several editorial decisions in advance.

Editorial principle Executable instruction Review question
Use concrete language Name the actor, action, object, and consequence where known Can the reader tell who did what?
Cut unnecessary words Remove language that adds neither meaning nor required context Did compression preserve every necessary qualification?
Prefer active constructions Use the responsible actor as the subject unless that actor is unknown Does the sentence obscure responsibility?
Avoid stale phrasing Replace clichés and generic transitions with specific claims Could this sentence appear unchanged in an unrelated document?
Permit exceptions Override brevity when precision, accessibility, or risk requires detail Would following the rule make the communication worse?

Orwell’s exception carries the most weight. Brevity can damage technical explanations, legal qualifications, accessibility instructions, and sensitive communication when concision removes distinctions the audience needs. A useful specification ranks accuracy, safety, and audience comprehension above elegance. It spends the reader’s attention only when precision requires it.

A style guide cannot make an agent factually correct or contextually wise. Reports of coding agents’ rapid progress also documented failures on higher-complexity tasks. Editorial rules can expose vague claims and force clearer attribution, but they cannot validate the underlying mathematics, evidence, or decision; polish makes an error easier to read, not more correct.

Human review becomes quality assurance, not rewriting

If every agent draft requires a person to rewrite it, generation scaled but the organization did not. Explicit rules must change the shape of human review.

Low-stakes outputs can pass when required checks succeed. Teams can sample recurring outputs for drift. They can send ambiguous, emotionally sensitive, high-stakes, or irreversible communications to a person before release. The human checkpoint does more than catch model mistakes: it assigns accountability where a rule cannot settle the decision.

Teams apply operational AI assurance to language when they specify what the system may decide, what evidence it must provide, what quality threshold it must meet, and when it must escalate. Review then tests compliance with a known standard instead of asking a tired editor whether the prose vaguely feels acceptable.

The productivity paradox around agents shows the alternative. A UCB study found that people who offloaded work to AI also worked longer hours. In that pattern, artifact volume makes the agent look productive while stacking a review queue in front of the human.

Editorial controls prevent each review from starting at zero. Teams make a decision once in a rule, turn that rule into an evaluation, and send only context-heavy or consequential cases to a person.

“Slop” names a transfer of costs

The term “AI slop” emerged in reaction to AI art generators in 2022, then spread across social media, books, search results, and other generated material. The term is useful because it names a cost transfer: unwanted output that is cheap for the producer but costly for the recipient to evaluate, moderate, interpret, or discard.

“Workslop” makes the transfer explicit: passable AI-generated work creates more work for colleagues. A sender saves time producing a plausible document; recipients then discover that it lacks a decision, omits necessary context, or says nothing with unusual confidence.

Teams can reverse that transfer by imposing constraints before distribution. They require the agent to identify the audience, state the action, preserve material uncertainty, remove generic filler, and escalate when tone or stakes exceed its mandate. Those controls reduce the number of outputs that pass by design, because cheap volume is no longer the objective.

When 1 million input tokens cost $2, the first draft stops being the scarce event. The work moves upstream: decide what a usable output must contain, encode those decisions, and send only consequential exceptions to a person. Teams that skip that layer do not save the cost of judgment; they pass it to every colleague and customer who must sort the prose afterward.

From autonomous coding to agent-mediated communication

  • 2026-02-26 — Evidence described AI coding agents as having made a major leap since December, completing complex projects with minimal oversight.
  • 2026-03-01 — A Bloomberg report highlighted a UCB study finding that people who offloaded work to AI also worked longer hours, illustrating how output volume can increase human workload.
  • 2026-07-12 — Sources reported that Cursor was building Sand, a rumored general-purpose agent for non-developers that would handle emails, texts, and documents.
  • 2026-07-13 — Further reporting characterized the rumored Sand agent as a competitor to Claude Cowork, extending the agent model from coding into everyday workplace communication.

Frequently asked questions

What is an editorial control plane for AI agents?

It is the system of persistent instructions, quality tests, review checkpoints, escalation rules, and records that governs what an agent may publish or act on. It turns editorial judgment into repeatable controls rather than relying on a person to repair every draft.

Why aren’t better prompts enough to control AI-generated communication?

A prompt is only one control and can be inconsistently applied. Durable governance also requires institutional context, explicit evaluation criteria, human-review triggers, and records showing which rules governed each output.

How can Orwell’s writing rules be made executable for an AI agent?

Teams can convert them into specific actions such as naming the actor and consequence, removing words that add no meaning, replacing clichés with concrete claims, and preserving necessary qualifications. Reviewers can then test whether responsibility is clear and whether compression removed important context.

Which AI-generated communications should require human review?

Ambiguous, emotionally sensitive, high-stakes, or irreversible communications should be escalated before release. Low-stakes outputs can pass automated checks and be sampled later for drift.

Do editorial rules prevent hallucinations or factual errors?

No. They can expose vague claims, require attribution, and make uncertainty visible, but they cannot verify mathematics, evidence, or judgment; polished prose can still be wrong.