Why AI writing is still a mess: we're dividing a task we never broke down - Part 2
Part 2: Allocating the 4 components for better writing
In the early 1900s, building a Ford Model T took about 12 hours. The process was pretty raw at the time. Crews moved around the factory from car to car, hauling around parts and tools, with each crew responsible for large, loosely defined chunks of the job. The result was a lot of fetching, waiting, and crowding around the same car. It worked, but it was messy and slow. Everyone was doing a little bit of everything and it wasn’t clear who was responsible for what.
Eager to make improvements, Henry Ford had an idea. He thought that if he could break the job down into smaller chunks, the process would be more organized, have fewer errors, and ultimately be faster. He also knew that machines could do some parts of the job better than his workers and even help with coordination.
By late 1913, Ford had decomposed the build into 84 discrete steps and introduced the assembly line. The new process let each worker remain in one spot as they performed the single task they were best at, while machines pulled cars down a line from one worker to the next.
The result? Build time on a Model T dropped from 12 hours to a mere 90 minutes.
It's one of the best examples in history of optimizing the division of labor to realize incredible gains in efficiency and production.
But why am I talking about building cars in an article that’s supposed to be about making writing better?
Well, if you've been following my recent essays, you know I believe strongly that modern writing (writing where we use AI) has a long way to go, despite most of the content industry touting it as solved. The reality is that modern writing's current state is messy, inefficient, and usually produces content that real people just don’t want to read.
We’re at the start of a long journey and it will get better. But the way forward isn't just trying harder at what we’re already doing. It’s changing our approach. Specifically, it’s dividing the labor between human and AI in a smarter, more deliberate way. My guiding compass behind all of this is the human what / AI how principle, where humans own the “what” and AI helps with the “how”.
Right now, modern writing basically sits where Ford was before 1913. It’s disorganized, slow, and ineffective, with its contributors, human and AI, doing a bit of everything under loosely defined roles.
It might not be a perfect example (I realize writing is different from building cars) but it shows how breaking a job into its fundamental parts and allocating those parts to the best-fit contributor drives real improvement.
I also like the example because it’s not just a story about humans. Like modern writing, the Ford process involved using machines alongside human workers. And Ford didn’t only use machines to perform single tasks. He even used machines to coordinate the entire process by having them move vehicles down the line, bringing the work from one human to the next when the time was right.
But Ford’s success didn’t come by skipping to the end. To get to a better place, his first step was decomposing the larger task into smaller fundamental parts. We performed that same step in part one of this two-part essay, where we outlined the four components of modern writing: ideas, structure, expression, and mechanics.
His next step was to assign those parts to the most qualified contributors, and that's where we find ourselves today. With the writing process finally broken down to its underlying parts, we’re now in a position to look at the strengths of both human and AI and consider who is best fit for each component.
Ideas are for humans
Ideas are the substance. The meaning behind your words. They form the argument, story, or understanding you want the reader to walk away with.
Idea work in writing is the process of figuring out what to say.
Ideas belong to humans, full stop. Not mostly, not by default. Fundamentally. And it's important to understand why, because it won't be changing any time soon.
AI doesn't struggle with ideas because it isn't good enough yet. It struggles because of a fundamental design limit.
That limit?
AI has only seen records of human experience. It has never had real experience.
In its training, it consumed trillions of words, yet zero experiences. Yes, newer models take in images and audio too (which aren't real experiences either), but the text it absorbed makes up the vast majority of its knowledge. And although words can be used to describe experiences, they can't fully communicate what it's like to live through them.
That may not sound like a big difference at first, but it's what keeps AI from ever having true authority over ideas.
To understand why, let’s first look at what we’re actually doing when we work with ideas in writing. There are two main tasks.
The first is generating ideas. This is the act of coming up with possible things you could say.
The second is judging ideas. This is the act of deciding which ideas are worth saying.
Keep in mind, these tasks aren't confined to an initial brainstorming phase. Ideas surface constantly as you write, from the central argument down to the small point a single sentence makes, so you're generating and judging the whole way through.
Good writing requires doing both tasks well.
Generating ideas
When it comes to idea generation, AI performs well in a quantitative sense. Ask it for angles on a topic and it will pump out more than you could think of. Ask it to help you write a section and it will produce a handful of ideas along the way. It can do both faster than you can, over and over, all day long.
For AI, volume isn't the problem. The quality of its ideas, however, is. Out of all the ideas it hands you, many will be useless, some will be okay, and maybe one will be good.
Why is this? Well, the quality of an idea is typically tied to the value it provides. More specifically, it's often tied to the value it adds to the world's existing body of knowledge—you know, everything humankind has written down over our history. AI struggles to yield high-value ideas reliably because its ideas come precisely from that existing body of knowledge. In other words, AI has learned what we already know. That's why its ideas feel like commodities. It's generating ideas, or combinations of ideas, that are already out there.
Now think about where your best ideas come from. Somewhere you went. Something you did. A mistake you made. A dirty look you received. A pattern you picked up after working your job for 30 years. That sandwich you left out too long last week. A gut-wrenching heartbreak you'll never get over. None of that is in the pile AI learned from. It's only in you.
So although AI can out-generate you in terms of volume, there's an entire source of ideas it has no access to, and that source is where your most valuable material lives. It can produce a thousand ideas and still never produce that one.
This is AI’s design limit surfacing in the first task. Its lack of real-world experience means its only “experience” is the words we’ve already written down.
Knowing this, use AI to generate when and where it helps. Just be ready for what it's likely to hand you: something serviceable and already out there. Occasionally it'll surface an angle you can use, and that's worth the almost nothing it costs. But the ideas that make a piece worth reading will have to come from you.
Which brings us to the next, harder question. Out of all the ideas available at any given moment, what actually gets written into the piece?
Judging ideas
Deciding which ideas are worth saying is a judgment call. And not just any judgment call, a creative one, which matters because it's a specific area where AI struggles big time.
Why? You guessed it. AI’s lack of real experience. Except here, it hinders AI in two different ways.
First, it gives AI a narrower base of information to reason with.
Think about what you're actually doing when you judge something. You're reasoning with whatever information is available to you, then deciding based on that reasoning. Because of this, the quality of the judgment depends heavily on how much information you hold relevant to the decision. Lots of rich, relevant information equals better judgment. Less or incomplete information equals poorer judgment.
AI actually does well with simple judgments, because in those cases it has all the information it needs. Ask it whether water boils at 100 degrees at sea level and it will tell you yes. That kind of question has one right answer, the answer has been written down thousands of times, and nothing about it lives outside of language. Everything required to make the call is sitting right there in the text AI studied.
But deciding whether an idea belongs in your article? That's a completely different animal. There's no fact of the matter to check. The question you're asking in these scenarios is, will this idea do something worthwhile to the person reading it? And this means the reasoning requires actually predicting an effect on a real human mind.
Suddenly the information that matters is enormous, and almost none of it is sitting in the world’s existing texts.
The information that matters now isn't about true or false. It's about human experience, and you hold that kind of information in forms that have nothing to do with words. Sights, sounds, and physical sensations. Emotions you've actually undergone. Memories of specific moments with all their texture still attached. AI holds one form of information: language, and math derived from language. You hold many, and most of yours never passed through words on the way in.
AI simply doesn’t have that kind of data and it never will. It can’t. Such data is only acquired through experience, and AI wasn’t fed experience. It was only fed words representing experience.
A helpful way to understand why this matters is to acknowledge that experience doesn’t arrive in language. Neither do the ideas we form from those experiences. Language comes after both.
Language is just a code we built to move that information from one head into another. And it's a lossy code. Every time we put experience into words, we pick out a few features and leave the rest behind. More importantly, some of that experience just outright can't be put into words, like the actual sensation of a cool breeze on your cheek, or what the color blue looks like.
Everything AI knows came only from that compressed code. Yes, it converts words into tokens, builds relationships between them, weighs and compares them, which is all genuinely sophisticated work, but every operation it performs still has words as its referent, not experience. Words in, computation over representations of those words, words out.
So when AI reasons, it's reasoning only at the language layer, with a compressed version of the world rather than the world itself. This results in a narrower, less rich base of information to reason with.
Second, AI’s lack of real experience leaves it unable to simulate a reader’s experience.
Remember, in creative scenarios, you're predicting what the idea will do to a reader. Not whether it's accurate or even whether it's relevant. But if it will land. If it will resonate. And this isn't a language event. It’s a feeling that happens inside a person.
When you're weighing an idea, you can run it through the same equipment it's eventually going to be tested on. You place the idea in your head, you get a feeling, and that feeling is a live sample of what you're trying to predict. You're not reasoning about the reader's experience. You're having a version of it.
Then you adjust from there. Maybe your reader knows less about this than you do, or cares more, or is walking in skeptical. But the point is you're adjusting a real reading, not building an estimate from secondhand records.
AI can't run this test. Not for lack of information about feelings like surprise, intrigue, or boredom. It has an enormous amount of information about all three. It can't run the test because it has never been surprised, intrigued, or bored. So when it estimates whether an idea will land, it isn't checking against the experience of something landing. It's comparing the text of the idea to how similar text has been received, and estimating from there.
In 1934, philosopher John Dewey wrote that "the artist embodies in himself the attitude of the perceiver while he works." He was describing exactly this. The maker keeps becoming the audience, testing the work by experiencing it, and reshaping it until it feels right.
You get a direct reading. AI gets a correlation.
The practical approach for ideas
AI’s design limit results in three major problems when working with ideas, one related to generating and two related to judging:
- Its raw material is what people have already written, putting the world’s most valuable ideas, the ones inside of us, out of its reach
- It has thinner material to reason with, because the parts of experience that matter most don't survive the trip into words
- It doesn’t have the required equipment to simulate a reader’s experience
None of that changes with a better model. Better models get better at the patterns. They don't grow a connection to reality, and they don't develop the ability to feel what their own suggestion does. Those aren't features that arrive in the next version. They're consequences of its design.
Therefore, use AI to help generate. Let it pump out angles and hand you options you wouldn't have found on your own, but tap your own experience for the truly valuable ideas out of its reach.
Most importantly, you make the decisions. AI can tell you whether an idea could belong here. What it can't tell you is whether the idea deserves to be here. And that call is usually the one that determines if the writing is any good.
Ideas, then, belong to humans. This is the "what" of writing, and it's the one component AI is structurally shut out of owning.
Structure and expression: the two arrangement components
Structure and expression are a bit more complicated.
They are clearly different components, but share a family resemblance in that they are both forms of arrangement.
Structure takes the big ideas and arranges them into a larger sequence. It decides what comes first, what supports what, where one section ends and another begins, and how the argument or explanation unfolds.
Expression takes language and arranges it at a smaller level. It selects words, builds phrases and sentences, and determines exactly how each idea, big or small, appears on the page.
Structure is the arrangement of ideas.
Expression is the arrangement of words.
They’re cousins, not twins. One works at the macro level and the other at the micro level, which is why they remain separate components. But underneath, they’re doing the same kind of work: taking existing building blocks and putting them in an order that communicates something.
Neither component creates the building blocks it works with. If a new idea appears while you're structuring a piece, that's idea work. It now becomes another block for the structure to arrange. Expression doesn't create its building blocks either. The words already exist. The job is to select and arrange them so they deliver the idea.
But don't mistake "only arranging" for "not important." Arrangement decisions can have an enormous impact on a piece, and they often determine whether the writing connects with its readers or falls flat.
Arrange ideas one way and you get a dry explanation. Arrange them another way and you get an argument that builds. Put a point at the beginning and it frames everything after it. Save it for the end and it becomes a reveal.
The same goes for language. An idea can be worded to feel blunt, warm, funny, urgent, or completely forgettable, depending on how the words are chosen and arranged.
That’s the beauty of the “how.” Ideas determine what is being communicated. Structure and expression determine how the reader experiences it.
Knowing that both structure and expression are forms of arrangement helps because we can group these components together when assessing who the best contributor is. But it doesn’t give us the answer.
The answer sits in a detail we usually don’t consider:
Arrangement can be applied to serve two different purposes.
The first is clarity.
Clarity means arranging so the intended meaning can be understood easily and completely. The ideas appear in a logical order. Related information stays together. The sentences say what they are supposed to say. The reader doesn't have to fight through confusion to find the point.
The second is rhetoric.
Rhetoric means arranging to create a particular effect. Maybe you want to persuade the reader. Surprise them. Build tension. Make them laugh. Hold something back until the right moment. Make an idea feel more important, memorable, emotional, or urgent.
Clarity gets the idea across.
Rhetoric adds an experience on top of the idea.
These aren't opposites. Good rhetorical writing normally still needs to be clear. The difference is that clarity establishes the floor, while rhetoric raises the ceiling.
The line between clarity and rhetoric turns out to be exactly where human and AI split on both structure and expression.
Arranging for clarity
AI is remarkably good at the clarity side of arrangement.
Give it a pile of unorganized notes and it can group the related ideas, cut the repetition, and turn the whole mess into something another person can actually follow.
Give it a badly written paragraph and it will hand back a clean one.
It does this so well because it has seen an absurd amount of structure and expression already. To put it in perspective, a typical person might read a few hundred books in their lifetime. AI has read the equivalent of millions of books.
What held AI back with ideas is exactly what makes it so strong here. With ideas, knowing nothing but language was a problem. Here, though, knowing language is the entire job. And AI knows language extremely well. It knows the common shape of an explanation, which words tend to go together, and which transitions connect one point to the next.
This makes it a natural fit for a big chunk of arrangement work.
For example, say you need to explain a messy project to a few people on your team. You've got pages of notes, scattered updates, open questions, decisions, risks, and action items. The team already cares and is going to read closely. There's no need to build suspense or entertain anyone. The goal is simply to build a report that communicates the information clearly.
When it comes to structuring that report, AI can handle most of it on its own. It can sort the pile into sections, put the urgent things first, tuck the supporting detail underneath, and build a hierarchy that makes the document easy to move through.
AI can handle the expression with similar ease. It can take your raw, poorly written notes and turn them into clean, readable language. And its value here is more than just a matter of speed. If you're not a strong writer, or the subject matter is unfamiliar territory where you don't know the common language, AI can generate clarity you couldn't have produced on your own. It takes a rough, half-formed explanation and renders it clean and correct almost instantly, often catapulting you from "here's roughly what I mean" to "here it is, expressed clearly" in a single step.
For clear, matter-of-fact communication, AI is a great fit.
Arranging for rhetoric
The balance shifts when rhetoric starts to matter.
Humans shine here, for the same reason they do with ideas. AI has never had a real experience. It has only seen records of experience.
That same design limit that keeps AI from owning ideas is also what keeps it from producing strong rhetoric.
With rhetoric, you're performing the same two tasks you did with ideas, just on structure and expression. You're generating options, then deciding which ones make it in.
And because the tasks are the same, the problems plaguing AI are the same too.
The first problem is that the only material AI has to work with is what's already been written. So when it generates a structure or a sentence, it's drawing from what's already out there. That's a recipe for genericness, and it's why almost everything it formulates has that safe, predictable flavor, just like its commoditized ideas.
But didn’t we just say that AI has read millions of books worth of text? If it's seen so much writing, wouldn't it have learned great rhetoric from those sources? It's a good question. And, yes, AI has learned the patterns of rhetorical structure and expression very well. It truly can offer up suggestions that look the part.
But a rhetorical move is only as good as its placement. Even the best lines from Shakespeare, Mary Shelley, or Stephen King fall flat in the wrong spots.
Which brings us to the judging problem. Like choosing which ideas are worth saying, rhetorical decisions are creative judgment calls. The question isn't whether something is correct, or even whether it has a rhetorical formula, but what it will do to a reader. Will this structure be intriguing? Will that sentence be suspenseful? Will this word hit home?
And we already know how AI performs in this department. It struggles because it has a narrower base of information to reason with, and it can't simulate a reader's experience.
You remember how you felt when you read your favorite lines. Those feelings never showed up in the words AI trained on. You can also run the same test as earlier, just on structure or expression this time. You can scan an outline or read a paragraph and get a feeling. Maybe the argument gives itself away too early and you get bored. Maybe a sentence charges you up. Maybe your opening paragraph starts with “In today's fast-paced world” and makes you cringe. AI gets none of that, because it's stuck at the language level.
To understand why this matters in practice, let’s look at an example.
Suppose you’re writing a public essay. Now, unlike the internal report, the reader isn’t obligated to care. The piece might need to earn attention, build curiosity, get past resistance, or walk the reader toward a conclusion they wouldn't have accepted up front. This is the work of rhetoric.
When it comes to structuring this kind of writing, it's not enough for the ideas to just appear in a sensible order. You might decide: don't introduce the framework right away. Start with the industry's claim that writing is solved, let the reader sit in the contradiction, and only then reveal that the real problem is labor allocation.
That's a structural decision, but it isn't about clarity. A perfectly clear outline could just state the thesis in the first paragraph. A different order is chosen on purpose, to build an intended experience for the reader.
The same standard applies to the expression inside that essay. AI’s default phrasing is typically clear, fluent, and technically solid. Ask it for something with more punch and it will try, but will likely land off target. Neither mode gives you something that actually earns a reader’s attention.
You might cut a sentence in half because the short version hits harder, or repeat a word three times to get your point across. You might reorder a sentence so it ends on the word you want them to remember.
Just like with the structural decision, you're choosing the less obvious approach to manage the reader’s experience.
The same thing that makes AI so reliable at clarity is exactly what caps it at rhetoric. It knows what normally comes next. You know when something else should.
The practical approach for structure and expression
When the job is clarity, AI can handle the heavy lifting, and it can do it at a scale and speed no human could reasonably match.
When the job is rhetoric, humans lead, for the same reason we saw with ideas. AI has never had a real experience, so its suggestions come out sounding generic and it has a hard time judging which moves fit the moment.
The exact balance changes with the piece. A straightforward internal summary may need almost no human rhetoric at all. A persuasive essay, story, or high-end brand piece may depend on it. But most writing sits somewhere in the middle and needs both. So the real question isn't whether human or AI should handle structure and expression. It's who does how much, and that depends on the mix of clarity and rhetoric called for.
Mechanics belongs to AI
Mechanics is the technical layer: grammar, punctuation, spelling, formatting, and consistency. Unlike the others, it's governed mostly by established rules. There's a correct way to punctuate a sentence or format a citation. And although style guides may differ at the margins, the bulk of mechanics is settled, documented, and not really up for interpretation.
That's what makes it an easy fit for AI. The task is to apply a known set of rules consistently across text, and that’s exactly what AI is great for. It doesn't get tired, it doesn't lose focus halfway down the page, and it applies the same standard from the first line to the last.
None of this means it's flawless. AI can still miss things, and human proofreading isn’t going anywhere. But of all four components, mechanics is the one where AI gets the most authority, freeing us humans up to focus more on ideas, structure, and expression—the stuff that really drives a good piece of writing.
What human-led means
It may be easy to take what I’ve argued here the wrong way: assuming that, for the human-led components, AI must stay out. This is not the case.
Where I’ve stated that humans lead a component, whether that's ideas or the rhetorical sides of structure and expression, I'm talking about who decides, not who's allowed to contribute. Human leadership is about ownership and final say. It doesn't mean AI sits on the sidelines.
In fact, having AI propose is one of the most useful things we can use it for. It can surface ideas you hadn't considered, suggest a different way to structure the argument, or give you three versions of a sentence when you're stuck on one. Even on the rhetorical side, AI can still generate options worth looking at. You should absolutely use it this way. Cutting it out of the generating process would waste one of its biggest strengths.
The thing to remember is that those proposals are candidates, not decisions. AI is helping you find the right idea or the right phrasing. It isn't choosing it. The choosing is still yours, and that's what matters.
A division of labor guide for writing with AI
The table below summarizes the approach for each component.
Remember, one thing the table can't capture is you. A strong writer will lean on AI less for structure and expression than a weak one will. Someone with deep expertise and no writing ability will rely on it much more. The split depends on what you're good at and where you need the help, so treat this as a starting point and adjust as needed.
Unlocking real gains
To reduce build time from 12 hours to 90 minutes, Ford didn't just break the work into steps. He matched each step to the best-fit contributor for the job. In this article, we've proposed how to do the same thing for modern writing.
Humans own ideas, because they must. AI’s lack of real experience prevents it from generating the most valuable ideas and from judging which ideas belong.
Structure and expression are a team effort. The balance fluctuates depending on the degree of rhetoric called for.
Mechanics sits comfortably with AI for the most part.
It’s a flexible but purposeful approach.
Applying this division of labor lets us leverage the power of AI in a responsible way. We get to use it for its strengths while ensuring high-judgment areas like ideas and rhetoric stay with us.
But there's a problem that arises from this conclusion: when writing with AI in this manner, getting from first draft to finished article still takes a lot of work, especially when the piece leans on rhetoric.
It turns out that writing properly with AI takes more work than we all first thought.
But maybe there’s another piece to this puzzle. Something we're still missing. Some way to unlock even bigger gains.
Think back to the Ford example. Ford didn't only use machines for doing work. He also used them to coordinate work across many human contributors. What if AI could do the same thing for writing? Not just contributing where it's strong, but gathering the work of several people and pulling it together into one piece. The load that's sitting on one person now could be shared across a larger group.
If something like this is possible, then maybe the next big question in writing is: how do we go from one human to many?
I’ll explore this question in my next article by looking at the synthesizing power of AI and the practicality of parallel human contribution.