How to edit photos with AI and hold one character across a series of frames

How to edit photos with AI and hold one character across a series of frames

Generating a picture is easy; the real task is built differently: remove something from an existing photo, put a particular product into another scene, hold one character across five listing images. Below is the sign showing which tool you need.

Generating a picture is easy; the real task is built differently: remove something from an existing photo, put a particular product into another scene, hold one character across five listing images. Below is the sign showing which tool you need.

---

Generating a picture is easy, editing one is not

The first thing everybody learns is generating from nothing. Describe it, get it.

Then a real task arrives, and it turns out to be built differently.

You need to remove something from an existing photo — cables, a passer-by, clutter in the background. You need to put a real product into different surroundings. You need the same character on five listing images in a row looking the same. You need to add a detail to an existing shot rather than redraw everything.

None of those gets solved by generating from nothing. Each time you get a new picture instead of a changed old one — and that is not your mistake but a different category of tool.

Two different tasks

Understanding the distinction saves weeks.

Generating from nothing. You describe, the model invents. Strong where there is no source and no accuracy is needed: concepts, illustrations, atmospheric shots, stylised visuals.

Editing and extending. There is a source image and it has to be changed while everything else stays. Other models work here — the ones that understand the context of a picture and correct it precisely.

The practical sign to choose by. If the task contains the word "this" — this product, this person, this room — you need editing rather than generating. If the task sounds like "draw something along these lines" — generating.

What genuinely gets solved

Removing and replacing. An unwanted object, a background, a caption, a person in shot. One of the most in-demand tasks in commercial work: a product photograph taken on a phone in unfortunate surroundings gets brought into usable shape.

Extending the frame. Widening an image so it fits another format: a horizontal photo for a vertical listing, a cropped shot for a cover.

Placing an object in a new scene. A real product in an interior, against a background, in somebody's hands. It replaces product photography in a substantial share of cases.

One character across a series. The most valuable thing for commercial work. A character looking the same across five frames is what you cannot deliver a line of listing images, a series of posts or an advertising set without.

Correcting details. A colour, an item of clothing, lettering on packaging, a small thing that is wrong.

Where the limit is

Exact correspondence to a real object. The model makes it similar rather than identical. For a product where every detail matters — markings, composition, precise shape — that is a fundamental limit.

Text on an image. A weak point of generative models generally. Lettering usually gets added separately in an ordinary editor.

Anything legally significant. Nothing presented as documentary photography can be edited — that is no longer a question of quality.

And an honest caveat about people. Faces and hands remain the hardest part. On series with people, every frame has to be checked.

How to keep a series consistent

The most valuable and the hardest part of commercial work.

Record everything on the first good frame. The description, the parameters, the reference, the model version. Without recording it, repeating will not happen — and a series of five listing images requires exactly that repeatability.

Work from one source image rather than from a description. A character set by text will be different every time. A character set by an image will be recognisable.

Change one parameter at a time. Pose, angle, background — in turn. Change three things at once and get a different person, and you cannot tell what broke the resemblance.

Check the whole series together rather than one at a time. Frames each fine on their own can fail to match in a row on light, tone or scale. You have to look at them the way a buyer will — in sequence.

Why this matters more than generating for commissions

A simple observation from practice.

A client rarely needs a picture in general. They need a picture of their product, their premises, their offering. Which means almost always working with something existing rather than creating from nothing.

From my scrape of 848 listings over 24 days: an image individually is $12–35, while marketplace listing images are $245–730 per set. The difference is not the quality of an individual picture. The difference is that a set is a series with a single style and a real product — and that is editing.

Somebody who can only generate stays on the bottom line.

How to structure the work

Start with the source rather than the prompt. A good source is half the result. Even light, a clean background, the product whole in frame. Shooting on a phone and refining is easier than rescuing a bad frame.

Record what worked. The description, the parameters, the reference. Otherwise repeating what you liked in the second frame of a series will not happen.

Correct in iterations. Do not rewrite the request from scratch — change one thing at a time from the previous result.

Check on a small screen. Listing images and covers get looked at in a feed rather than at full size. What is invisible large shows up small, and the other way round.

What you need from the client

A separate point that saves rounds of revisions.

Source files at maximum quality. Not forwarded through a messenger with compression but the originals. Half the problems with a result are problems with the source rather than the tool.

A reference for what they like. Describing a style in words is nearly impossible; showing it is easy. One example from the client replaces three rounds of correspondence.

Where this will be used. A listing, a cover, an advert, print — that determines the format, the resolution and what to look at when checking.

What must not change. The packaging colour, the shape of the logo, a particular detail of the product. A client's list of prohibitions matters more than their list of wishes — and it has to be asked for directly, because they will not volunteer it.

And agree the number of revisions in advance. For generative work that is critical: a revision often means regenerating, which means a new cost. Two rounds in the quote is fine; "we revise until you're happy" is not.

Where to start

The "ChatGPT Images & Nano Banana" express course — about generating and editing images in chat models: from a first prompt to branded content with one character across a series of frames.

The key difference from the course on generating from nothing: this is about working with existing images — editing photos, holding a character across a series, adding details to something finished.

The price is $9. The course is included in the Plus plan — not the base one, and it is not in the three free days of Basic.

Three routes to it: a one-off purchase at $9, the Plus plan, where it sits alongside the course on generating from nothing, the video model courses and the website templates, or 1 000 experience points earned by working on the platform.

What is available free. The catalogues and descriptions are open to guests with no registration. Registration gives three days of Basic — with 140+ AI assistants across 14 categories, creative and e-commerce among them, and an encyclopaedia of more than 190 vetted services showing what gets done with what.

Generation is paid from the wallet — on use, the cost visible before you press, money returned on a failure.

Do one thing today: take one of your photographs that almost works but something spoils it — an unwanted object, a poor background, the wrong format. Try to fix that one rather than generate something similar. The difference between the two approaches becomes obvious in ten minutes.

---

The exact figures per tool

The photo editor — 34 models, $0.12–0.50: replace an object, remove something, extend the edges, change the background or the style.

Background removal — $0.12. Upscaling — 6 models, $0.12–0.40. Product Shot — 2 models, $0.12: the product in a studio scene with no shoot. Face swap — $0.12.

And a pipeline if there are many items

Auto-Shine from the archive: a product image becomes selling content in a couple of minutes, handling a catalogue rather than one product.

Plus a technique from module 6 — the Midjourney web editor: you select an area right in the browser and write what to replace. Only that gets redrawn; the shadows and background stay.

Module 6 — Plus; Auto-Shine — Make and Full.

---

What to read next

[Learning Midjourney](/en/blog/how-to-learn-midjourney) — the tool for creating from nothing.

[Your own visual style](/en/blog/train-your-own-style-lora) — holding a manner across a series.

[A 3D model from a product photo](/en/blog/3d-model-from-photo) — when photo correction is not enough.

[AI for marketplace sellers](/en/blog/ai-for-marketplace-seller) — who this sells to.

[AI video generation in 2026](/en/blog/ai-video-generation-guide) — the full guide to video models.