Drop in a photo and the outfit describer writes out what the person is wearing — every garment, color, fabric, fit, layer and accessory — as a paragraph you can paste straight into a caption, a listing or a draft.
Running it needs a free account. The free plan includes 10 image-to-text descriptions as a one-time allowance, and paid plans have no cap on them.
写真をアップロードまたはドラッグ
PNG、JPG、またはWEBP 最大30MB
画像がここに表示されます
Pick a mode under the photo. Describe The Outfit stays on the clothes, Describe The Person covers the whole figure, and Midjourney Prompt turns the look into an image prompt. The selector beside the button controls which language the result comes back in.
Drag in a file or paste an image URL. Full-length shots work best — if the shoes and the hem are in frame, they end up in the description.
It is the first mode under the photo. Use Custom Question instead when you only care about one piece, like what the jacket is made of.
The result lands as a single paragraph with a copy button next to it. One run costs one description from your account.
Ask most models to describe this outfit and you get back “a woman in a white top and jeans”, which helps nobody. Here is what the outfit mode is built to name instead.
Every piece by its actual type rather than “a top”: a cropped denim jacket, a ribbed turtleneck, pleated wide-leg trousers.
Specific shades rather than color families — sage, oxblood, bone — plus stripes, checks and prints, and where on the garment they sit.
Denim, ribbed knit, satin, corduroy, leather, sheer mesh. Texture is usually what separates a written outfit description from a shopping list.
Oversized, tailored, cropped, high-waisted, floor-length. Fit is the part people leave out, and the part that changes the whole look.
Tucked or untucked, sleeves pushed up, buttons left open, hems cuffed, waist belted. How it is worn, not only what is worn.
Footwear, bags, jewelry, eyewear and hats get the same treatment, and the paragraph closes with the aesthetic the look reads as — minimalist, streetwear, business casual, Y2K — and the season it suits.
Plenty of people searching for this are not holding a photo at all — they are writing a novel, a game character or a roleplay profile and need clothes on the page. This tool always starts from an image, so it cannot invent a look from a text idea. What it can do is turn a reference you already have — a runway shot, a film still, a render of your character — into wording you rewrite in your own voice. If you would rather write it cold, this is the formula that makes clothes read as real.
She wears a heavy charcoal wool coat that falls to mid-calf and holds its shape when she turns, over a bone knit stretched out at the cuffs. Black trousers, cropped just above scuffed leather boots. One glove. No bag.
Written by hand, not generated — it is here to show the formula working rather than to advertise the tool's output. Forty words, five garments, one detail that says something.
If you are rebuilding a look in Midjourney or Flux rather than writing about it, switch the mode to Midjourney Prompt. You get the same wardrobe read-out shaped as an image prompt — subject, garments, fabric, styling, then framing and lighting — instead of a paragraph of prose. That turns a photo into a reusable outfit descriptor: build a character sheet, hold one look steady across a whole set of generations, or lift a silhouette from a lookbook and put it on someone else.
Yes. Ask it to describe this dress and you get the silhouette, neckline, sleeve length, waist, hem, fabric and print, plus shoes and jewelry if they are in frame. A full-length shot with the hem visible reads far better than a cropped portrait.
It writes down only what it can see. If the trousers leave the frame at the knee, you get the trousers to the knee — it will not guess the hem or invent shoes. When a piece matters, crop or reshoot so it is actually in the photo.
You need an account, and creating one costs nothing. The free plan comes with 10 image-to-text descriptions. They are a one-time allowance rather than a daily refill, so once they are spent they do not come back; Starter and Pro have no cap on descriptions.
No. Every run starts from a photo you upload or an image URL you paste. If you are writing fiction with nothing to upload, pick any reference close enough to your idea and rewrite the result, or use the five-part formula above and write it yourself.
Only when a logo is clearly legible. Otherwise it stays generic — “white leather low-tops with a green side stripe” rather than a brand name. That is deliberate: guessing a label from a silhouette is exactly how these tools get it wrong.
The outfit mode stays on the clothes and the styling. The person mode also covers face, hair, build, posture and expression. Use the outfit mode for a clean wardrobe read-out, the person mode when you need the whole figure.
Yes. Pick a language beside the button and the description comes back translated — 16 are supported, including Spanish, French, German, Japanese, Korean, Arabic and both Chinese scripts. It still counts as one description.
Last updated: