Are we outsourcing our creativity when we use AI to generate images, or extending it? It’s possible to do either, or both. Half a dozen of one, six of the other.
Two questions are worth pulling apart: what does this new symbiosis look like, and is there skill involved?
Where the decisions sit
Start with the floor case. Type “create an image of a unicorn” and you’ll get something, maybe even something lovely.

But look at who made the consequential choices: the model picked the pose, the light, the style, the background. You supplied a veto. That’s outsourcing creativity with a vote.
Now flip it. You decide the subject, the composition, the palette, the mood. The model explores variations, renders options, and surprises you at the edges. That’s extension. The tool is the same, but the decision-making has moved, and most real work oscillates between the two inside a single session.
Each side brings something the other lacks. The human brings intent, constraint, and judgment. The model brings variation, speed, and cross-domain association. Neither closes the loop alone.
Where skill lives
Whether we get something we like depends on the assumptions the model makes, how well it can take a simple concept and run with it (Midjourney is good at this), and the aesthetic taste of the prompter reviewing and selecting the results to keep.
Often the sloppiest of AI slop impresses the noob, until they learn to distinguish what is mediocre or poor generative art from what stands apart. A refinement in perception that comes from generating 10,000 images, perhaps. This practice delves into understanding the parameters of what can be evoked through what can be described.
Prompts, incantations, cantrips… learning the differences between art styles, color palettes, lighting options, framing, and compositional rules, once implicit, now become tools of expression. A good generative artist is always expanding their vocabulary, and thus the discriminations they can make visually. They notice through experimentation the impact of these “trade secrets” and invisible architecture in visual design by making the rules explicit. An intermediate prompt, for example, may look something like this:
A luminous white unicorn standing ankle-deep in a shallow black-water forest pool at dawn, shown in a three-quarter profile facing left, its horn angled toward a narrow opening in the trees. The unicorn has a subtle pearlescent coat with individual hairs visible along the mane and shoulders; its mane drifts in a gentle breeze and catches the light in faint lavender, rose-gold, and icy-blue highlights.
Composition: cinematic vertical portrait, full body visible, unicorn placed on the right third of the frame, with the horn leading the viewer's eye diagonally upward toward a bright break in the canopy; foreground reeds slightly out of focus at lower left, layered trunks receding into mist for depth; generous negative space in the upper-left quadrant. Low camera angle, approximately chest height, creating quiet monumentality without exaggerating anatomy.
Lighting: cool pre-sunrise ambient light fills the forest; a narrow warm golden sunbeam enters from the upper left and rims the unicorn's back, mane, ears, and horn; soft reflected cyan light rises from the water beneath its legs; gentle volumetric rays visible in the mist; high dynamic range but no blown-out whites.
Color palette: restrained and harmonious—deep blue-green shadows, charcoal water, desaturated violet mist, muted moss, then selective accents of warm amber sunlight and opalescent lavender in the mane. Avoid saturated rainbow colors; retain a dark, contemplative overall value structure.
Art direction: painterly magical realism, combining the atmospheric depth and luminous brushwork of late-19th-century Symbolist painting with contemporary high-detail fantasy concept art. Delicate visible brush texture, refined anatomical realism, elegant proportions, emotionally quiet rather than cute or ornamental.
Mood and atmosphere: ancient, sacred, solitary, slightly uncanny; drifting mist, a few floating pollen-like motes, barely visible ripples spreading from its hooves.
Render constraints: crisp focal detail on the unicorn's eye, horn, and shoulder; progressively softer focus toward the foreground reeds and distant forest; no rider, no castle, no text, no jewelry, no wings, no extra limbs, no cartoon styling, no oversaturated neon glow.

The critique loop
This isn’t meant to be a tutorial on making better images. It’s an exploration of how working with AI can teach us new ways to perceive, and thus describe, our intentions.
We can use AI to help level up the visual and design constructs available to us. Learning to write these types of prompts is a creativity and skill multiplier. We can ask:
“From a visual design standpoint, what is working in this image and what is not, and how could it be improved?”
AI Response:
“Keep the dawn forest, cool mist, warm rim light, and pale unicorn. Make the unicorn’s eye the sharpest and most contrast-rich point in the frame. Reduce highlight intensity in the tail, lower mane, water glitter, and right-side tree trunk. Shift the unicorn into a subtle three-quarter stance with one forehoof gently breaking the water. Place the horn against a darker band of forest rather than the brightest sun opening. Make the mane slightly damp and clumped near the neck and legs, with restrained lavender-gold sheen rather than uniformly luminous strands.”

To make this understanding our own, we can ask “why is this better?”
And the response may be:
"The core issue: hierarchy. In the original, several bright, high-contrast elements compete for first attention:
- The sun and beam cluster at upper left.
- The horn, which aims toward that bright cluster.
- The bright rim-lit mane around the neck and tail.
- Sparkling water in the lower frame.
- Bright, textured bark along the right-side tree.
All of them are individually attractive. But images are generally easier to read when the viewer can perceive an intended order:
Primary focal point → secondary accent → environmental discovery
I agreed with this, and not just in the abstract. Once the hierarchy was named, I couldn’t unsee the competing highlights in the first version. That’s the human half of the loop: the model proposes a diagnosis, and I decide whether it matches what I now see.
The catch
Serendipity isn’t new. Surrealist automatism, Cage’s chance operations, and Eno’s Oblique Strategies all invited luck in. But those dice were fair. This one is loaded: it’s trained on everyone’s priors, so its randomness leans toward the median.
The same pull applies one level up. The critique above is conventional compositional doctrine: focal hierarchy, controlled highlights, value structure. The prompt itself, with its lavender, rose-gold, volumetric rays, and mist, is the house style of “good AI fantasy art.” Learning these rules is the first skill. Knowing when to break them is the next one. And when generation is nearly free, taste becomes the bottleneck, because volume is easy to mistake for taste.
A new audience for writing
Writers once practiced wordsmithing for human readers. Now there’s a new kind of reader, one that can’t see your intent, only your description. That makes precise language a way to steer outcomes and, as the unicorn exercise shows, a way to train your own eye. You can’t describe what you can’t distinguish, and every distinction you learn to describe becomes something you can choose rather than merely accept.
That’s the line between outsourcing and extending: whether you come out of the loop seeing more than when you went in.