Open Grok Imagine 2 when the thing you are making is drawn, painted, or built in a style. Open GPT Image 2 when it has to look photographed, or when there are words in it that somebody will read. Both models are current flagships and both are in Picsart, so the choice is about the job, not about which one is better.
Grok Imagine 2 is the newest image model from xAI, trained to hold up across three areas at once: photography, design and illustration. Editing is part of the model itself rather than a layer added over the top of it. GPT Image 2 is OpenAI’s newest, and it is built around two things it does unusually well: skin, light and surfaces that read as a real photograph, and letters that come out correct.
Comparison table
Grok Imagine 2
GPT Image 2
Who makes it
xAI
OpenAI
Strongest at
Illustration, stylized and designed work
Photorealism, and words that must be correct
Handling text
Plans typography and layout as a composition
Around 99% character accuracy in six scripts
Keeping a look consistent
Carries a style across separate generations
Up to 10 matching images in one run
Editing what you have
From an instruction, no selection to draw
From an instruction, plus extending the frame
Largest image
2k
2048 by 2048, with a 4096 by 4096 beta
Here is how that plays out across the work most people actually bring to an image model.
Illustration, pixel art, and art styles
This is where Grok Imagine 2 is the one to open. Its range across visual languages is the point of the model rather than a side effect of it. Halftone portraits made of fine white dots, classical ink painting, soft manga pages, watercolor journal spreads, retro pixel art, vintage travel posters: it moves between those registers without being talked into them.
It also holds to an instruction closely, including the small parts of it, which matters more in stylized work than people expect. A style request carries a lot of specific baggage. Pixel art has a resolution logic. Halftone has a dot structure. Ink painting has rules about where the brush lifts. A model that only approximates the style gets those wrong and the piece looks like a filter rather than a drawing.
GPT Image 2 will produce illustration too. It is just not what it was tuned for, and the difference shows in the pieces that depend most heavily on committing to a look.
Photos that look real
GPT Image 2 is the stronger pick here, and its specific claim is worth knowing. The two tells that used to mark an image as generated, a warm cast over everything and skin with a waxy finish, have been trained out. Pores and fine lines survive. Shadows sit where the light source says they should. Depth of field falls off gradually instead of all at once.
That makes it the model for product shots, for portraits and headshots, for interiors, and for anything going into a place where a real photograph would normally sit. A catalog page. A press kit. A slide where a stock photo would look obviously stock.
Photography is one of the three areas Grok Imagine 2 was trained on, so it is far from a bad photographic model. But when the test is whether a viewer would assume a camera made it, GPT Image 2 is the safer bet.
Text inside the image
GPT Image 2 again, and this is its single most reliable advantage. It renders text at around 99% character-level accuracy across Latin, Chinese, Japanese, Korean, Arabic and Hebrew. It holds up on fine print, on curved text that wraps around a shape, and on multilingual labels where a wrong character is not a typo but a mistake.
So: packaging with a real product name on it. Signage. A label in more than one language. A chart whose annotations have to mean something. A mockup with real interface text instead of placeholder shapes.
Grok Imagine 2 handles type well in its own way. It works out type and layout the way a designer would, so a dense visual made of several parts holds together as one composition instead of collapsing into a pile of elements. That is an arrangement strength rather than a spelling strength, which is a genuinely useful thing on an illustrated piece where the lettering is part of the artwork. When the words themselves have to be exactly right, use GPT Image 2.
Making a set of matching images
Both models do this, and they do it differently enough that the difference decides jobs.
GPT Image 2 returns a matching batch from a single prompt, and in Picsart you can ask it for up to 10 at once. You describe the thing once and the whole set comes back together. That suits a product series, a storyboard, or a set of variants where the whole set arrives together.
Grok Imagine 2 works the other way. What you feed it survives from one generation to the next and through edits, so a look carries forward across images made separately at different times. That is how you build out a world: a character in one generation, the places she goes in the next few, the objects she carries after that, all holding the same style. For game assets, a comic, or a video project that needs a consistent visual bible, that is the more useful shape.
Editing a photo you upload
GPT Image 2 is the more specified editor, and in Picsart it is also the more capable one.
It names its operations: patch a single area, extend the picture past its original edges, take an object out, replace a background, restyle the whole frame. All of it from a written instruction, with no selection to draw first. Extending past the frame is the one worth flagging, because it is the operation Grok Imagine 2 has no answer to.
Grok Imagine 2 edits from an instruction too, and editing was built into the model rather than bolted on. What it does not offer is a way to push the picture beyond the crop you started with, so a reframe still has to happen somewhere else.
Vertical, square, and ultra-wide images
Grok Imagine 2 has 13 frame shapes, and the interesting ones are at the extremes: 19.5:9, 9:19.5, 20:9, 9:20, 2:1 and 1:2. Those cover a phone screen edge to edge, and the long thin banners that ad slots ask for.
GPT Image 2 has 7, running from 1:1 out to 16:9 and 9:16, plus an auto setting that picks the shape for you. It reaches widescreen in both orientations, but it cannot be persuaded into a 20:9 banner. If you already know the slot this image has to fill and its shape is unusual, check the list first. No amount of prompting adds a frame the model was not given.
Which model to pick for each job
What you are making
Open this
Anything illustrated, painted, or in a defined art style
Grok Imagine 2
Pixel art, game assets, sprites, icon sets
Grok Imagine 2
A product shot or a portrait that has to look photographed
GPT Image 2
A label, a pack, or a shopfront sign, in any script
GPT Image 2
A data graphic or a screen mockup whose labels have to be legible
GPT Image 2
A character plus locations plus props that all share one look
Grok Imagine 2
A matching set of variants delivered in one go
GPT Image 2
Pulling a frame wider than it was, or swapping what is behind the subject
GPT Image 2
A full-bleed phone frame or a very wide banner
Grok Imagine 2
Deliverables that have to carry proof of origin
GPT Image 2, which embeds content credentials
How to try both in Picsart
Both models live in
Picsart AI Playground
, which is the fastest way to settle this for your own work: one prompt, both models, two results side by side. They share a single credit balance, so trying the second one costs you a click.
GPT Image 2 reaches further into the product. It is in the
AI Image Generator
, and in
Flow
it can be wired in as one node among many, so a generation feeds straight into whatever has to happen to the file afterward. Everything it does is listed on the
GPT Image 2 model page
.
A test prompt to run in both
A prompt that exposes the split cleanly, because it asks for a style and for legible type at the same time:
Try this prompt
A vintage travel poster for a coastal Italian town at golden hour, hand-painted look, muted teal and terracotta palette, the town name set in bold condensed type across the lower third, small print underneath reading "Departures daily from the harbour"
Run it in both. Grok Imagine 2 will tend to give you the more convincing poster as an illustrated object. GPT Image 2 will tend to give you the more trustworthy small print. Which of those two failures you can live with is the answer to the whole question.
Get answers to common questions
GPT Image 2, when the words have to be correct. It renders text at around 99% character-level accuracy across Latin, Chinese, Japanese, Korean, Arabic and Hebrew, and it holds up on fine print and on curved text. Grok Imagine 2’s strength with type is different: it plans typography and layout so a dense, multi-part visual holds together as a design.
Grok Imagine 2. Illustration and design were trained targets for it alongside photography, and it commits to a visual language rather than approximating one. Pixel art, halftone, ink painting, manga and hand-painted poster looks are all comfortable ground for it.
They are level as standard. Grok Imagine 2 tops out at 2k, and GPT Image 2’s native ceiling is 2048 by 2048, which is the same ballpark. The difference is that GPT Image 2 documents a 4096 by 4096 mode still in beta, so it has somewhere further to go. GPT Image 2 is also the one to pick if you would rather the model chose the frame shape itself, since it has an auto setting.
Yes. Neither one makes you mask an area or trace a shape first; you describe the change in words. GPT Image 2 spells out more of what it will do, naming patched regions, deleted objects, swapped backgrounds, wholesale restyling, and canvas extended beyond the original crop.
GPT Image 2 writes C2PA content credentials into every file, so its origin stays attached to the image wherever it goes. Grok Imagine 2 publishes nothing comparable. On brand or agency work that has to document provenance, that alone can settle the choice.
Try both and compare
Start from the deliverable. Drawn, styled, or one piece of a set that has to match: Grok Imagine 2. Meant to read as a photograph, or carrying copy someone will actually read: GPT Image 2. When you genuinely cannot tell, run the brief through both in
Picsart AI Playground
and compare what comes back.
You cannot upload a PDF to Instagram. LinkedIn takes one directly in an organic document post, which is exactly why the gap catches people out: Instagram has no equivalent, because a feed post accepts images and video only. So posting a PDF on Instagram means converting each page into an image first, then posting the set as one carousel post.
PDF to carousel is the whole job, and it is the part people do badly. A raw PDF page exported straight to JPG comes out the wrong shape for the feed, at document proportions, with type sized for reading rather than scrolling. This walkthrough uses the
Create an Animated Social Media Carousel from a PDF
template in
Picsart Flow
to do it properly: it splits the PDF, redesigns every page as a vertical slide, and keeps your wording exactly as written.
Why you cannot upload a PDF to Instagram directly
Instagram feed posts accept photos and videos. There is no document post type, so a PDF has to become a set of images before Instagram will take it. There is no setting to change and no workaround: the file has to be converted.
That asymmetry is what makes a PDF to carousel workflow especially useful for Instagram. On LinkedIn it is a convenience, because the platform will accept your document either way and converting only buys you a better looking post. On Instagram it is the only route in, so the quality of the conversion decides the quality of the post.
That leaves you three options, and only one of them looks good:
Screenshot each page.
Fast, and it produces the wrong aspect ratio with soft type.
Convert PDF to image with a file converter.
You get clean JPGs at the PDF’s own proportions, which is still a document shape, not a feed shape.
Convert and redesign in one pass.
Each page comes back reformatted as a vertical slide with the original text intact. This is the PDF to carousel route, and it is what the workflow below does.
The difference matters because a PDF page and an Instagram slide are different objects. One is built to be read at arm’s length, the other to be understood in about a second while somebody scrolls.
What you need before you start
One multi-page PDF.
Every page is processed as a separate item automatically, so you never upload page by page.
Twenty pages or fewer.
An Instagram carousel holds up to 20 photos or videos, so a longer PDF needs trimming first.
A first page worth leading with.
Instagram uses slide one as the cover, and its aspect ratio sets the crop for every slide behind it.
How to convert a PDF to an Instagram post
Four nodes, one input. The redesign runs on
Nano Banana 2
and the optional animation on
Seedance 2.0
, and both are already wired when you open the template.
That single PDF controls everything downstream:
How many slides you get, since one page becomes one slide
The running order of the carousel
Every word that appears on the finished slides
Which slide becomes your cover
1.
Drop the PDF into the Document node
This is the only input the workflow takes. The page conversion, the redesign and the animation all regenerate from this one file.
2.
Let the splitter separate the pages
The PDF Page Splitter turns the file into individual pages and passes them forward as one batch, in the original order. No page-by-page uploading.
3.
Check the batch before you spend on the redesign
One upload produces as many parallel design jobs as the PDF has pages, so confirm the page count is what you expected.
4.
Run the redesign at 4:5
The Carousel Slide Designer converts every page into a vertical slide at 4:5 and 1440p, preserving the original content and applying one visual system across the batch.
5.
Review the set, not the single slide
The palette, typography and hierarchy are applied across the whole batch, so one slide out of context tells you very little.
6.
Download the batch in page order
You now have a coordinated set of images. Confirm the running order survived the download before you go near Instagram.
7.
Check every slide is the same shape
Instagram takes the aspect ratio from slide one and crops the rest to match, so a mixed batch gets cropped rather than letterboxed.
8.
Upload the images as one carousel
Add the slides in order and publish. Instagram posts them in the order you add them, with no reordering afterwards.
What the redesign changes and what it keeps
That 4:5 ratio is the point. A straight PDF to image export gives you the document’s own proportions, and the feed crops it for you. Running the conversion and the redesign together is what produces a batch that posts cleanly.
Across the full batch the node applies:
A single color palette, so the slides read as one set
Consistent typography and spacing
A clear visual hierarchy on every page
A unique layout per slide, so the series does not look like a template loop
And three things stay locked:
The wording.
Both design prompts forbid rewriting, summarizing, removing, misspelling or adding any text.
The page order.
Slide one is still page one.
The content source.
Each slide is built only from the page it came from.
Split and convert each page
Transform the uploaded PDF page into a premium social media carousel slide. Preserve every piece of original text exactly as written. Do not rewrite, summarize, remove, misspell, or add any text.
Redesign as a carousel slide
Redesign the uploaded PDF page as a premium vertical social media carousel slide. Use the uploaded page as the only content source. Preserve every original word, heading, number, and sentence exactly as written.
Convert the PDF to video instead
The last node is optional, and it is what a file converter cannot do. Motion Slide Generator turns every redesigned slide into a five-second motion-graphics clip on Seedance 2.0, which gives you a PDF to video output rather than a static one.
The animation is deliberately conservative. It animates the design you already approved rather than reinterpreting it.
What stays fixed:
The text
The composition
The visual style
The typography and layout
What moves:
Smooth reveals on existing graphic elements
Subtle movement across the slide
Depth
Transitions between elements
Animate each slide
Transform the uploaded document page into a polished 5-second motion-graphics animation. Keep the original page design, composition, colors, text, typography, and layout completely unchanged.
One setting to check before you export: the animation node ships at 3:4 and 480p while the slides render at 4:5 and 1440p. Because Instagram crops every slide to match the first, set the animation node to the same ratio as your slides, or post the clips as their own carousel rather than mixing them with the stills.
Instagram carousel specs to get right
Converting the file is half of it. These are the numbers that decide whether the post goes up cleanly:
Accepted formats:
images and video. Not PDFs, and not documents of any kind
Slide count:
up to 20 photos or videos in one carousel post
The first-slide rule:
Instagram takes the aspect ratio from slide one and crops the rest to fit, so mixed ratios get cropped rather than letterboxed
Order:
slides post in the order you add them, with no reordering after publishing
Vertical 4:5 is the default worth choosing because it takes the most feed height, which is why the template converts pages to that ratio.
Tips for PDF content that holds the scroll on Instagram
Cut the title page
If page one of your PDF is a cover, it wastes the slide that earns
Fix density in the PDF, not the output
The prompts preserve wording exactly, so a dense
Judge the set, not the slide
The visual system is applied across the whole batch, so one
Trim to under 20 pages first
A 30-page PDF will happily produce 30 slides that Instagram
Decide static or animated before you export
The two outputs ship at different ratios,
Reuse the workflow, not the design
Swapping the PDF in the Document node reruns
Turn your next PDF into an Instagram post
The slowest part of a carousel is deciding what belongs on each slide. If that decision already exists in a document, the workflow handles the rest.
Open the template in Picsart Flow
, drop in a PDF, and run it.
Get answers to common questions
Not as a PDF. Instagram feed posts accept images and video only, with no document post type. To post PDF content you convert each page into an image and publish the set as a carousel post.
Split the PDF into pages, convert each page into a vertical 4:5 image, then upload the images as a single carousel post in page order. The workflow above does the split and the redesign in one run so the slides come out as a consistent set.
One workflow that does both halves at once. Splitting the pages and redesigning them are separate jobs in most tools, which is why a file converter leaves you with document-shaped images you still have to lay out. Running the split and the redesign together is what makes PDF to carousel a single step.
No. Both design prompts instruct the model to preserve every original word, heading, number and sentence exactly as written, and not to rewrite, summarize, remove or add text. The design changes and the copy does not.
Yes. The final node turns each redesigned slide into a five-second motion-graphics clip, so the same PDF can come out as a set of stills or a set of short videos.
Muse Image 1.0 is live in Picsart. Meta’s new image model is available now in
AI Playground
, and it is the first model in Picsart that works through your prompt before it draws anything: it plans the composition, looks up references it does not already know, builds any structured elements in code, and checks its own work before handing you a result.
That changes what you can reasonably ask for. Most image models take your words and render them in a single pass, which is why they guess at things they have never seen and why text inside an image so often arrives as letter-shaped smudges. Muse Image 1.0 treats a prompt as a task to work through rather than a description to match, so instructions with several moving parts survive the trip.
Below: what the model is, how its reasoning actually works, what it does well, how to prompt it, and where to find it now that it has landed.
What is Muse Image 1.0?
Muse Image 1.0 is Meta’s agentic image model, built by Meta Superintelligence Labs. One model covers the whole job. It generates images from a text prompt, edits images you already have, and composes new ones from several reference images at once.
What separates it from a standard generator is that it uses tools while it works. It can search the web for visual references and current facts, and it can write and run code to lay out charts, plots and QR codes accurately before placing them in the picture. It also reviews its own output as it goes, making a small correction when a detail is off, or starting again when something larger is wrong.
The practical effect is accuracy on the things image models usually fumble: real places, real products, legible text, and any instruction with more than one part to it.
How Muse Image 1.0 actually works
Ask a typical model for a poster with a headline, a subhead and a logo in the corner and you get an image that looks like a poster with gibberish printed on it. The model matched the mood, not the instruction.
Muse Image 1.0 breaks the request down first. It works out what goes where, resolves anything it needs to look up, renders, then checks the result against what you actually asked for. Where a conventional generator makes one pass from words to pixels, this one runs a loop, and the loop is where the quality comes from. Meta found that giving the model more time to think produces steadily better images, and that thinking harder beats simply generating more options and picking a favourite.
Three things happen inside that loop.
It looks things up
Point the model at a real landmark, product, logo or style and it can pull visual references rather than approximate from memory. Ask for something that depends on current information and it can go and find it instead of inventing a plausible answer.
This is the difference between an image that resembles a place and one that depicts it. A model working from memory alone produces a building that feels roughly Parisian. A model that can look first produces the building you named.
It builds structure in code
Some elements have to be correct rather than merely decorative. A chart’s proportions carry meaning. A diagram’s labels have to line up with what they label. A QR code either scans or it is a decorative square.
Because Muse Image 1.0 can write and run code, it constructs those elements properly and then places them in the image, instead of drawing an impression of what a chart looks like.
It corrects itself
The model reviews its own draft as it goes. When a small detail is wrong it makes a local edit. When something larger is wrong it starts that part again. When it is missing information it goes and finds it.
The interesting part is that Meta did not design this behaviour. It emerged during training, simply because a model that caught its own mistakes produced better images and was rewarded for it.
Editing with precision
Editing is where the reasoning shows up most plainly. Ask Muse Image 1.0 to clear fog from a landscape, remove someone from the background, restore a damaged family photo, or rewrite the text on a sign, and it changes what you named while leaving the rest of the frame alone.
That restraint is harder than it sounds. The common failure in AI editing is collateral damage: you ask for one change and the model quietly re-renders faces, shifts colours, or rearranges the background. Naming both halves of the instruction, the thing to change and the thing to protect, gives it a boundary it can hold.
Composing from several references
You can hand the model more than one reference image in a single prompt, and interleave your instructions between them, so each picture is captioned with the job it is doing. Use the pose from this one. The colour palette from this one. The room from this one.
That inline pairing is what makes complex composites tractable. Rather than attaching a folder of images and hoping the model infers your intent, you tell it what each reference is for.
The related trick is consistency across a set. Anchor a run of generations on a small group of references and the look holds from one image to the next. This is what makes campaign sets, product catalogues and multi-image social sequences workable, where the perennial difficulty has been keeping image five looking like image one.
Refining an idea across turns
The clearest demonstration of what the model is holding onto is a chain where each step depends on the last. Meta’s own walkthrough runs like this:
Start with two reference photos, a cat and a dog, and ask for them as best friends having a picnic on a sunny day, in a vintage 35mm style. Then ask to see that exact picnic photo as a framed print hanging on the wall of a cosy cafe, with a table and two empty chairs in front of it. Then ask for the front of the cafe, with its name, matching the vibe of the interior, with the framed photo visible through the window. Then design that cafe’s paper menu using its exact name, adding a “Picnic Special” with a small illustration of the same cat and dog. Finally, place that menu on the table from the empty-table shot made three steps earlier.
Nothing in that sequence is a fresh prompt. Each turn inherits the cafe, the animals, the style and the name invented along the way. That is the difference between a generator and something you can art-direct.
Prompting Muse Image 1.0
Because the model reads a prompt as an instruction rather than a mood, how you write changes the result more than it does elsewhere.
Caption each reference as you go.
Label what every image is contributing rather than attaching several and hoping. Each reference gets a job.
Say what should stay the same.
When editing, name the thing you want changed and the thing you want protected. “Add a red wool hat, keep the snowy forest background” gives the model a boundary.
Be explicit about numbers and details.
If it matters that there are exactly five of something, say exactly five, and say it should stay five. Vague quantities are where any image model drifts.
Name the medium.
Vintage 35mm, Korean manhwa, claymation, botanical engraving, isometric low-poly. A named style lands harder than an adjective.
Ask for the text you want.
If words belong in the image, write them out exactly as they should appear, including the punctuation.
Let it think when it counts.
The model has a reasoning setting. Give it room when the output has to be right, and dial it back when you are exploring quickly.
It sits in the model picker under the prompt box, marked New.
3.
Describe what you want
Upload an image first if you are editing rather than generating from scratch.
4.
Set your shape and how many
Seven aspect ratios cover square, story, landscape and portrait, and you can return up to ten variations from a single prompt.
5.
Generate, then keep going
Tell it what to change rather than rewriting the prompt, and it builds on what it already made.
The Advanced panel is where the model’s own behaviour is exposed: reasoning strength, whether it searches for images and facts, whether it uses its layout and chart tools, and your export format. The defaults leave everything on, which is what you want for most work. The case for switching the search tools off is speed on purely imaginative prompts, where there is nothing real to look up.
Five things to try first
A launch kit for something imaginary.
Invent a product, then build the packaging, a poster, a spec card and a social set, all anchored on the same references so they look related.
A recipe or how-to card.
A single image carrying legible steps, where the typography is part of the design rather than an afterthought.
An event poster with a working QR code.
The code is generated properly rather than pasted in, so it actually scans off the screen.
A room, restyled.
Photograph a space, ask for it in a different style, and keep the pieces you
liked from earlier versions as you iterate.
An explainer diagram.
Something with labelled parts and a sequence, where being readable matters more than being pretty.
Get answers to common questions
Muse Image 1.0 is an image generation and editing model from Meta Superintelligence Labs. It generates images from text, edits existing images, and composes images from multiple references, and it can search the web and run code while it works to get details right.
It reasons through a prompt before rendering rather than mapping words to pixels in one pass. It can look up visual references and current facts, build structured elements like charts and QR codes in code, and correct its own output before returning it.
Open [AI Playground](https://picsart.com/ai-playground/), switch to Image mode, and select Muse Image 1.0 from the model picker. Describe the image you want, set your aspect ratio and how many variations you want back, and generate.
Yes. Upload an image and describe the change you want. It edits precisely, altering what you asked about and leaving the rest of the picture intact, and you can keep refining across several turns.
Yes. Rendering legible, correctly styled text is one of its strengths, which is what makes it suited to infographics, posters, menus and any graphic where the words carry the meaning.
Yes. You can pass several references in one prompt and interleave instructions between them, so each image is labelled with what it should contribute.
Start creating
Muse Image 1.0 is live in
AI Playground
. Pick it from the model picker, describe what you want, and let it work the problem before it draws.
To edit all PDF pages at once, load the file onto a canvas, split it into pages, and run one written instruction across the whole set. That is how you edit a PDF with AI instead of by hand. No page-by-page clicking, no rebuilding the document from scratch, no learning a new interface. You describe the change in plain language and every page comes back changed the same way.
That is what the Bulk edit all pages of a PDF template does in Picsart Flow. It takes a multi-page document, fans the pages out into a batch, and applies a single prompt to the whole stack. The template ships with a real instruction already sitting in the prompt bar: change the font to Times New Roman, remove all emojis, make the colors magenta. Three separate edits, one run, every page.
This guide covers how the template is built, how to point it at your own file, how to change the font across a whole document, and which prompts survive a batch.
What an AI PDF editor does differently
Most PDF work is repetitive by nature. A change that takes ten seconds on page one takes ten seconds again on page two, and a forty-page deck turns a small decision into an afternoon.
A bulk PDF editor solves the volume problem. An AI PDF editor also solves the instruction problem, because you are not recording an action or configuring a rule. You are writing a sentence, and an AI document editor reads it the way a colleague would.
The difference shows up on the edits that are tedious to specify and easy to describe:
Swapping every typeface in a document to a single font
Recoloring headings, accents, and highlights to a new brand palette
Stripping emojis, icons, or decorative marks out of a working file
Flattening a mixed-format handout into one consistent look
Rebranding a template that was written for a different company
Turning an internal document into something you can post
Each of those is one sentence to a person and dozens of clicks to a mouse. That gap is the whole reason to batch the job.
It suits the documents that were built once and then kept getting reused:
Pitch decks and sales one-pagers going out under a new brand
Onboarding guides and internal handbooks
FAQ sheets and templated answer documents
Event programs, menus, and printed handouts
Course material and worksheets that need a consistent look
If you have ever done bulk image editing on a folder of photos, the mental model is the same. One instruction, many inputs, one pass. You can batch edit PDF files exactly that way.
How to edit all PDF pages at once
The template is deliberately small. Two nodes, one prompt, one run. Open it from the
Picsart Flow template library
and the canvas arrives already wired, so there is nothing to connect.
Step 1. Load your document
The Document node sits at the left of the frame and holds the file. The template ships with an eight-page FAQ document loaded, and a pager in the corner counts through it.
Click the Document node to swap in your own PDF
Use the arrows or the thumbnail filmstrip to move between pages
Check the page count before you go further, because that count is what gets processed
Step 2. Let the pages fan out into the batch
The Document node wires straight into a Batch node, and that connection is what turns one file into many inputs. Load an eight-page PDF and the Batch node reads “items 8” with all eight pages tiled inside it.
Every page becomes its own item in the batch
The badge on the node tells you exactly how many items are queued
The Add more button lets you drop in extra pages or images alongside the document
Step 3. Write one instruction for the whole document
The prompt bar runs along the bottom of the canvas. Whatever you write there applies to every item in the batch, so the instruction has to make sense on any page in the file.
The template’s own prompt stacks three edits into one line:
The template prompt
Change the font of the document to Times New Roman, remove all emojis, and make the colors magenta.
Notice what it does not do. It never names a page, never points at a specific heading, and never describes content that only exists once. Every clause is true of the whole document.
A batch-safe instruction usually has three parts:
The property you are changing, named plainly, such as the font or the accent color
The value you want it changed to, as specifically as you can put it
The things that must stay untouched, spelled out rather than assumed
Step 4. Set the model and the output
Under the prompt bar sit the controls for the run. The template comes preset, so you only touch these if you want something different.
Model is set to GPT Image 2
Size is set to 1024×1024
Quality is set to High
The run button shows the credit cost before you commit, which reads 40 on the eight-page default
Step 5. Run it and review the stack
Hit run and the batch processes as a set. Review the results together rather than one at a time, because what you are checking is consistency.
Scan for pages where the instruction landed differently
Look hardest at the densest pages, since they carry the most for the model to hold
Check that the clauses you wrote as protections actually held
Download the set or keep it on the canvas and wire it into the next step
How to change the font in a PDF across every page
Font is the first clause of the template’s prompt, and there is a reason it comes first. It is the change people most often want across a whole document and the one that punishes you hardest for doing it manually.
Change the font in a PDF one page at a time and you make the same decision over and over, with a fresh chance to be inconsistent each time. Batched, it is one sentence.
Three things make a font instruction land cleanly:
Name the typeface, or name the category.
“Times New Roman” is unambiguous. “A clean sans serif” is a category, and the model will pick within it, which is fine when you care more about the feel than the exact face.
Say what happens to the hierarchy.
Without a word about heading sizes, a font swap can flatten the structure. Tell it to keep the relative sizes.
Protect the wording.
A typography instruction should never be an invitation to rewrite. Say the words stay as they are.
Font swap, structure protected
Change all text in the document to Times New Roman. Keep the relative heading sizes, the paragraph breaks, and the wording exactly as they are.
If the file arrived carrying three or four typefaces, which happens to any document that has been passed around, the same instruction doubles as a cleanup. One named font in, every inconsistency out.
Five prompts to run across a whole document
The prompts below are written to survive a batch. Each one describes a property of the document rather than a detail on a single page, and each one says what to leave alone.
Brand color swap
Change every heading to deep navy and every body paragraph to dark charcoal. Change all accent marks, bullets, and highlights to warm gold. Keep the layout, the spacing, and the wording exactly as they are.
Single typeface
Change all text in the document to a clean sans serif typeface. Keep the relative heading sizes, the paragraph breaks, and the wording unchanged.
Strip the decoration
Remove all emojis, icons, colored highlights, and decorative marks. Keep the text content and the layout identical to the original.
Full rebrand
Apply a warm cream background with charcoal text throughout. Change every accent color and link color to burnt orange. Do not change any wording, spacing, or layout.
Serious to social
Make the document bolder and more graphic. Increase the visual weight of every heading, add strong color blocking behind section titles, and keep all body text legible and unchanged in wording.
Start editing your PDF pages with AI
Open the
Bulk edit all pages of a PDF template
, drop in the file you have been putting off, and write the change as one sentence. Run it once and watch the whole stack come back consistent. Build it on the
Picsart Flow
canvas.
The “i’m just a girl” trend is a collage of the things you carry, flattened into solid pink shapes, where one object at a time comes back in full colour. Save a version for every object, play them in order, and the collage lights up piece by piece. Creator
@moonsol.design
posted the tutorial this drop is built from and it has passed 148,000 plays.
What is the “i’m just a girl” trend?
Everything is flat except your face.
The objects keep their outlines and lose every bit of detail. The cutout of your face stays in full colour the whole way through, so there is always one real thing in the frame.
One object returns at a time.
Ten objects in the source means ten separate exports, each with a single item back at full colour while the rest stay flat.
It is a video made of stills.
Nothing animates. Every frame is a saved image, and the motion comes from playing them in order at half a second each.
Why it works
Flattening makes the reveal readable.
When everything else is one colour, a single object in full detail is impossible to miss.
The silhouette colour is free.
The pink is the canvas behind the collage showing through the holes, so changing the background changes every shape at once.
It is a list you can watch.
A still collage asks you to scan it. This one hands you the items one at a time.
Build the grid in the
collage maker
, which lays out mood boards as well as photo grids. The flatten step is the
background remover
followed by an invert, and the objects can be cut once and kept in the
sticker maker
so you can rebuild the set later. The saved frames go on a timeline in the
video editor
. Anything you do not own gets generated in
AI Playground
, which puts 182 models from 34 providers behind one prompt bar.
How to make it in Picsart
1.
Cut your face out on a white canvas
Open a new project with a white background and add a selfie. Remove the background, then take the eraser to everything below the chin. You want the head sitting on its own.
2.
Arrange the objects around it
Add the things you carry as cutouts, spaced evenly around the face, and set your line of text into the gaps. Shoot your own objects on a plain surface. Anything with a readable logo comes back at full strength the moment you restore it, so turn labels away from the camera. Save the finished collage.
3.
Drop the collage onto a coloured canvas
Start a second project and set the background to the colour you want your silhouettes to be. Pink in the source. Add the collage you just saved and scale it up to fill the frame.
4.
Remove the background, then invert it
Remove the background so the objects sit on the colour. Then hit invert. The white surround comes back and every object turns into a flat hole with the canvas colour showing through.
5.
Restore one object, save, undo
Brush restore across a single object to bring it back in full colour, then save the result. Undo, restore a different object, save again. One export per object.
6.
Assemble the saves at half a second
Add every image you saved to a video timeline in the order you want them revealed, then set each clip to half a second. Finish on a frame with everything restored.
The prompt
Object cutout, for anything you do not own
A single object photographed straight on against a plain white background, centred, with even studio lighting and a soft contact shadow beneath it. The whole object sits inside the frame with clear space around every edge. Colours are clean and true to the real thing. LOCKS: plain white background only; one object; no hands; no text; no logos or readable branding; no room reflections.
Variations worth trying
Change the canvas colour.
The silhouettes are whatever sits behind them, so one project gives you the whole palette.
Reveal in a deliberate order.
The source jumps around the collage rather than working top to bottom. Group by category if you want it to read as a list.
End on everything.
The last frame in the source has every object restored, which gives the run somewhere to land.
Ten objects, ten saves.
The “i’m just a girl” trend works because flattening everything makes one object impossible to miss. Build the collage, invert it, then bring the pieces back one at a time. Open the
background remover
and flatten the first one.
To make reels with AI, you brief a reel director instead of driving an editor. Hand it a starting point, an idea or one to ten of your own photos, and it turns that into a single cinematic clip. Plot, storyboard, then one video generation. That is the whole loop.
Your reel director is Reeva, Picsart’s Reel Maker agent. What comes back is a single cinematic clip plus a caption, hashtags, and a best time to post.
The planning step is why this works, and it is the whole difference between an agent and a generator. A reel with photos usually fails for one reason: the photos just sit there. Stills cut together on a beat read as a slideshow, and viewers scroll past a slideshow. A reel earns its name when the images move, when there is a beginning and an end, and when the person in frame still looks like the person in frame.
This guide covers the four ways to make reels with AI, the plot-to-storyboard-to-generation pipeline behind it, the exact flow from upload to finished clip, montage mode for photos and clips, editing a video you already shot, reel ideas worth stealing, and the craft decisions that separate a reel people watch from one they thumb past.
What separates a photo reel from a slideshow
A slideshow shows photos in order. A reel tells a short story with them. The difference is not the transition style, it is whether anything moves and whether the sequence goes somewhere.
Four things do most of the work:
Motion inside the frame.
A still that drifts, breathes, or has its subject shift slightly reads as footage. A hard cut between two frozen images reads as a gallery.
A shape to the sequence.
Setup, turn, payoff. Even at eight seconds, a reel that lands somewhere holds attention longer than eight seconds of pretty.
One consistent face.
If your face changes shape between panels, viewers notice before they can say why, and trust in the whole clip drops.
Sound that matches the cut.
Audio is what makes the sequence feel deliberate rather than assembled.
A reel director handles all four because of the order it works in, which is worth understanding before you upload anything.
Plot, storyboard, then one generation
A reel can be built two ways. You can generate a set of shots and join them, or you can plan the whole thing and generate it once. Reeva does the second, and the sequence is fixed:
Plot.
The agent decides what happens and in what order, before any pixels exist.
Storyboard.
That plot becomes six panels, the plan made visible, laid out shot by shot.
One video generation.
The approved storyboard becomes a single cinematic clip.
The third stage is the one people miss. The reel arrives as one video generation, not as separate renders joined afterwards.
It also changes where your attention goes. You are not picking the least disappointing render out of ten. You are reading one plan.
How to make reels with AI from photos, step by step
Step 1. Open the agent
Reel Maker sits with the rest of the
Picsart Agents
. Picsart agents live in WhatsApp, Telegram, and Slack, so you can brief a reel from your phone and wake up to finished work.
Step 2. Choose how you want in
Reeva takes four kinds of input, and which one you pick is the real creative decision. You are not typing a prompt and hoping, you are choosing what the reel gets built out of.
An idea, no assets.
Use when you want a look you cannot shoot. Concept pieces, product fantasy, anything that does not need your face.
One to ten of your own photos.
Use when the reel is about you, your client, or your product and recognizability is the point. Identity is preserved, so the person in the photos is the person in the reel rather than a stranger who resembles them.
Photos and clips together.
Montage mode, covered below. Use when you already shot the thing and just need it cut.
A video you already have.
Use when the footage is fine and only the treatment is wrong.
To make a reel with photos, pick the second.
Step 3. Upload your photos
Ten is the ceiling, and it is a ceiling rather than a target. Pick for variety, not volume:
Different distances, so the reel has wides and close-ups rather than ten head-and-shoulders frames.
Different angles on the same subject, which gives the sequence somewhere to move.
Clean light. A well-lit ordinary photo animates better than a dramatic dark one.
Faces you actually want on screen, because they are the ones you will get.
One frame that clearly opens the story and one that clearly closes it.
Nothing you would not post as a still, because animation does not rescue a bad photo.
Step 4. Approve the storyboard
The six panels come back for approval. Read them as a director would:
Does panel one make someone stop scrolling?
Does the middle change something, rather than restating panel one?
Does the last panel land, or does it just stop?
Are your photos being used for what they are good at?
Send it back if the answer is no. Changing a storyboard costs nothing. Changing a finished render costs a render.
Step 5. Pick your audio
Three options: AI-generated music, a track you upload, or silence. Silence is not a throwaway choice. A reel with strong visual motion and no music often reads as more confident than one carrying a generic bed, and many feeds autoplay muted anyway.
Step 6. Collect the reel and the ready-to-post pack
What comes back is not just a file. If you are posting to Reels, TikTok, or Shorts, the clip arrives with a ready-to-post pack, which is the three things that usually sit between a finished reel and a published one:
A caption
, delivered with the reel.
Hashtags
, delivered with it too.
A best time to post
, which answers the question most people resolve by posting whenever they happen to finish.
The gap between rendering a reel and publishing it is usually the caption nobody felt like writing. Closing that gap is the difference between a folder of finished clips and a feed.
Montage mode: photos and clips in one reel
The third way in takes both. Hand over stills and footage together and three things happen:
Photos come alive with subtle motion, so they stop reading as stills.
Clips get trimmed to the part worth keeping.
Everything is stitched into one reel.
The craft rule here is ruthlessness. Whatever the source folder holds, the reel should be shorter than feels fair to it.
Edit a video by describing the change
The fourth way in involves no photos at all. When the footage already exists and only the treatment is wrong, you describe the change rather than perform it.
Four kinds of edit work this way:
Restyle.
The look of the clip changes while the action stays exactly as shot.
Swap, add, or remove an element.
Something in frame leaves, arrives, or becomes something else.
Change the mood.
The same footage in a different emotional register.
Extend.
The clip runs longer than the length you originally filmed.
This is the most underrated of the four inputs. A clip that is close but wrong no longer has to be reshot or rebuilt. It needs one sentence describing what should be different.
That is the full range. Reeva makes one reel at a time, so long-form editing, beat-synced music-video cuts, genre treatments, and mass platform variants for volume publishing are different jobs and sit outside what this agent does.
Tips for making reels with AI that people finish
Open on the strongest frame, not the earliest one
Chronology is a habit, not a structure.
Give each photo less time than feels comfortable
Under-stay rather than over-stay.
Keep the count low
Six well-chosen photos beat ten that include three near-duplicates.
Shoot vertical when you can
Reels, TikTok, and Shorts all favor 9:16, and cropping a horizontal photo to vertical throws away most of the frame.
Vary the distance between consecutive frames
Wide, close, wide reads as edited. Close, close, close reads as a contact sheet.
Decide the ending before the beginning
Knowing where a reel lands makes every earlier choice easier.
Write the caption while the reel is fresh
Or take the one delivered with it and edit rather than start from a blank field.
Get answers to common questions
You brief an agent instead of driving an editor, and it plans before it generates. That planning step is what makes AI reels hold together rather than looking like animated stills.
Yes. A reel does not require filmed footage. Photos are one of the four ways into Reeva, and one to ten of them are enough.
One is enough, and ten is the working maximum with Reeva. Somewhere between five and eight tends to give a sequence enough variety to move without leaving each frame on screen too briefly to register.
Yes. Identity is preserved when you supply your own photos, so the reel keeps the face you uploaded rather than generating a new person who resembles you.
Yes, that is montage mode, and it is the third of the four ways in. Your photos and your clips end up in one reel rather than in separate blocks.
Describe it rather than perform it. Style, elements, mood, and length are the four things you can change this way.
Once, at the storyboard. It is the single approval gate and it sits before any spend, so revisions at the planning stage are free.
Start making reels with AI
The photos are already on your phone. The part you have been avoiding is the editing, and that is the part a reel director takes. Brief it, approve the storyboard, and post the reel. Open
Picsart Agents
and start with Reel Maker.
Generative editing has landed on the Picsart Flow canvas. The latest version of
Picsart Flow
is live, and it hands you a brush, a frame you can stretch, a fresh lighting setup for video, and a virtual camera you can walk around a subject. Repaint any region of an image, push a picture past its own borders, relight or re-weather a clip, and re-shoot a photo from a new angle. You can also change the settings on a whole selection of nodes in a single move.
Five changes ship in this release. Not one of them asks you to leave the canvas or start a generation from scratch. Take them one at a time.
Inpaint: paint over an area and describe what belongs there
Inpaint is the single most-used gesture in professional AI editing, and it finally lives on any image node. Open the Actions menu, choose Inpaint, and go after exactly the thing that bothers you: background clutter in a product shot, an object that should be something else, a flaw, or a gap that wants a new element. Everything outside your brush strokes is left alone, so the parts you already liked survive the edit.
Paint the region with the brush. The eraser removes strokes, and everything merges into one mask.
Undo steps back stroke by stroke, and Clear all is a single undoable step.
Describe the change in the prompt bar below the node. The run starts once you have both a mask and a prompt.
Only the masked region changes, and the rest of the image stays pixel-identical.
The result always lands on a new node, so the original stays right next to it. Compare, branch, or A/B both versions downstream.
Outpaint: expand an image beyond its original borders
Reframing for a new placement no longer means rebuilding the asset. Turn a square generation into a 16:9 hero banner, give a cramped composition room to breathe, or stretch a scene out for a story format. The new space is filled generatively, so it reads as though the shot was always that wide rather than stretched.
Drag any of the 8 handles to open up new space. The original image never moves or resizes.
Slide the frame to choose where the new space sits, not just how big it is.
Aspect chips lock the frame to a target ratio for exact placements, and Auto keeps it free-form.
Add an optional prompt to steer what fills the gap, or leave it empty for a seamless extension.
The result is a new node, so your source image is never replaced.
Visual Effects: relight and re-weather your videos
Change the mood of a clip the way a reshoot would. Shadows swing to a new direction, highlights warm or cool, and reflections catch up, so the result reads as filmed rather than filtered. One product video can carry a whole seasonal campaign, with the same footage at golden hour, in the rain, and under neon.
Presets sit in a gallery of three groups. Every tile previews the effect it applies. Hover one and the preview plays live, so you can audition looks without spending a generation on them.
The transformed video lands as a new node downstream. Your original clip stays intact. That is what lets one clip carry several looks at once.
Camera Angle: re-shoot a photo from a new viewpoint
One photo, every angle. Instead of staging a multi-angle shoot for each product or character, drop a single shot into the new Camera Angle node and re-render it from the viewpoint you need. Proportions and lighting stay consistent across views, so the whole set hangs together.
Horizontal angle:
8 positions all the way around the subject, from Front through Right side, Back, and Front-left, in 45 degree steps.
Vertical angle:
Low-angle, Eye-level, Elevated, or High-angle shot.
Distance:
Wide shot, Medium shot, or Close-up.
Three dials make up one control. Point the virtual camera and run. It is built for listing galleries, ads, and storyboards that need front, side, and three-quarter views of the same subject.
Bulk parameter edit across a multi-selection
Changing one setting on ten nodes used to mean ten identical edits. On a large graph it was the most repetitive action in the product. Now the whole selection moves in one edit.
Multi-select same-type nodes and open Settings from the selection toolbar.
Edit their shared model settings as one group, including resolution, duration, style, and anything else the model exposes.
Where nodes disagree, the chip shows “Mixed” instead of guessing, and your first pick applies everywhere.
Switch the model for the entire selection at once.
One apply, one undo. The whole bulk edit reverts with a single Cmd+Z.
Nodes that are mid-generation are left untouched, and the panel tells you so, for example “Editing 8 nodes, 2 in progress”.
Tips for generative editing on the canvas
Mask a little wider than the flaw.
Every stroke merges into one mask rather than a stack of selections, so a few extra passes cost nothing and the eraser trims them back.
Leave the Outpaint prompt empty for a clean extension.
The prompt is optional there, so add one only when you want to steer what appears in the new space.
Set the aspect chip before you drag the handles.
Locking the frame to a target ratio beats eyeballing eight handles when the placement size is already fixed.
Hover the effect tiles before you commit.
Each one plays live in the gallery, so you can rule presets out without spending a generation on them.
Set all three camera dials before you run.
Horizontal, vertical, and distance work as one control, so picking them together saves a second pass at the same subject.
Batch your settings last.
Build the graph first, then multi-select and push resolution or duration across the whole selection in a single edit.
Keep both nodes when you branch.
Inpaint, Outpaint, and Visual Effects all write to a new node, which makes an A/B pair the default instead of something you have to set up.
Start editing on the canvas
Open a workflow in
Picsart Flow
, find the asset that is nearly right, and change the one thing that is not. There is more queued up behind this release, so it is worth seeing what the canvas can do each time you come back.
Get answers to common questions
It is a set of edits that run on the canvas itself rather than in a separate tool. This release adds five of them: Inpaint, Outpaint, Visual Effects for video, a Camera Angle node, and bulk parameter editing across a selection.
No. Only the masked region changes, and the rest of the image stays pixel-identical. That is the point of masking rather than regenerating.
Yes. The run starts once you have both a mask and a prompt. Neither one on its own is enough to fire it.
Yes. Aspect chips lock the Outpaint frame to a target ratio for exact placements, and Auto keeps it free-form. You can also slide the frame to choose where the new space sits.
Presets come in three groups: Lighting, Time of Day, and Weather. Every tile previews the effect, and hovering plays it live so you can compare before running one.
Three dials cover it. Horizontal offers 8 positions around the subject in 45 degree steps, vertical offers Low-angle, Eye-level, Elevated, and High-angle, and distance offers Wide shot, Medium shot, and Close-up.
Yes. Multi-select same-type nodes, open Settings from the selection toolbar, and edit their shared model settings as one group. Nodes that are mid-generation are left untouched.
It stays where it is. Inpaint, Outpaint, and Visual Effects all put the result on a new node instead of replacing the source, so you always keep the version you started with.
Four nodes on a canvas. One selfie you upload, two images generated by
Nano Banana 2
, and a final video generated by
Seedance 2.5
using all three as references.
That is the whole build. The shot that once needed a rig of 99 synchronized cameras now needs three reference images and one prompt.
The bullet time effect is the shot where time slows almost to a stop while the camera keeps moving around the subject. Everyone knows it from The Matrix. Almost nobody has been able to make one, because until recently making one meant owning the hardware.
This tutorial covers what the effect actually is, why it is turning up everywhere again, and then the exact node by node build in
Picsart Flow
, using the Fashion Bullet Time Effect template.
What is the bullet time effect?
Bullet time is a camera move, not a filter. The subject is frozen or moving in extreme slow motion, and the camera orbits around them at normal speed.
Those two things fighting each other is what makes it look impossible. Your eye reads the frozen subject as a photograph, then the moving camera tells it this is video, and the brain cannot file it as either one.
The name comes from being able to see a bullet in flight. The technique is older than the phrase, and it has a few other names you will run into:
Time slice.
The still photography version, where the frozen moment is a single composite image.
Frozen moment
or
time freeze.
Used for the same effect when nothing is being shot at.
The Matrix effect.
What most people call it, after the film that made it famous.
Arc shot
or
orbit shot.
The camera move on its own, at normal speed, without the time distortion.
The distinction that matters for making one: bullet time is not slow motion. Slowing footage down slows the camera down with it. Bullet time keeps the camera at full speed and takes the time out of the subject instead.
Why the effect is everywhere again
Two things changed, and neither of them was the effect getting easier to film.
First, action cameras shipped a cheap approximation. Spin a 360 camera on a cord above your head and you get an orbit around yourself, which is why searches for bullet time now come loaded with camera and accessory terms. It looks the part in a wide outdoor shot and falls apart the moment you want a controlled, lit, close up frame.
Second, and this is the real shift, AI video models learned to hold a subject still while moving the camera. Bullet time is now a preset in most AI video tools rather than a production challenge, and the results no longer need a location, a rig or a crew.
So the format moved. It stopped being a stunt in an action film and became a shot people use for:
Fashion and outfit reveals, where the freeze lets you see the whole look.
Product and accessory close ups, because the camera can hold on a detail at macro range.
Dance and sports clips, frozen at the peak of the movement.
Brand campaigns that want to look expensive without a shoot budget.
The fashion version is the one this template builds, and it is the clearest demonstration of the effect, because a frozen figure with a moving camera is exactly how a luxury campaign wants to present clothes.
What is inside the Fashion Bullet Time Effect template
Open the template and the canvas has four nodes. Understanding what each one is for is what lets you rebuild it around your own subject later.
Your selfie.
An image node you upload to. This is the identity reference and nothing generates it.
Sunglasses: text.
A Nano Banana 2 image node that generates the accessory for the macro close up, including legible text on it.
You.
A Nano Banana 2 image node that combines your selfie with a wardrobe image and outputs a full body turnaround sheet.
Final.
A Seedance 2.5 video node that takes all three images as references and generates the finished vertical clip.
The order matters. The video node is last because it needs the other three to exist first, and it treats them as references rather than as a starting frame.
That is the part that separates this from animating a photo. You are not putting one image in motion. You are giving the video model a person, an outfit and a product, then asking it to shoot all three.
Step by step: build the bullet time shot
Open the
Fashion Bullet Time Effect template
and follow along on your own canvas. Every node below is already there and already wired, so the four steps are about what you change rather than what you build.
Step 1. Upload your identity reference
Drop a selfie into the
Your selfie
node. This is the face every later stage locks to, so it does the most work of anything you provide.
What to pick:
A clear, front facing or three quarter angle shot.
Even lighting with no heavy shadow across the face.
No sunglasses, no hat, nothing covering the features.
Reasonable resolution, because the turnaround stage will be reading detail out of it.
A phone selfie is fine. A blurry one at arm’s length in a dark room is not, and no amount of prompting downstream fixes it.
Step 2. Generate the accessory with text on it
The
Sunglasses: text
node runs on Nano Banana 2 at 1440p. It generates the hero product for the macro moment in the video, and the reason it is its own node is the text.
Nano Banana 2 is the model in this build that can put a specific word on a specific surface and keep it spelled correctly. That is what makes the accessory readable when the camera pushes in on it, and it is why this stage is not left to the video model.
The prompt is already written into the node. It describes the object, the surface the text sits on, the lettering material, and the camera treatment, and it ends by naming the focal length and the angle, which is what stops the model handing back a flat catalogue photo instead of a macro frame.
One thing to change: the word on the temple arm. Swap it for your own brand name and this node becomes a product placement.
Step 3. Build a turnaround sheet of yourself in the outfit
The
You
node also runs on Nano Banana 2, and it takes two inputs. Your selfie as the identity reference, and a wardrobe image as the outfit reference.
What comes out is not a portrait. It is a character turnaround sheet, a grid of the same person in the same outfit seen from the front, the sides and the back.
This is the most important node in the template and the easiest one to misread. A single photo gives a video model one angle to work from, so the moment the camera swings around, the model has to invent the other side of you and the clothes change halfway through the orbit.
A turnaround sheet removes the guessing. Every angle the camera will pass through already exists as a reference, which is what holds the outfit and the face steady across a full arc.
The prompt that does this comes loaded in the node, and it works by locking two things separately. Every facial feature from the selfie, and every garment, colour and fabric from the wardrobe image, each listed as something the model is not allowed to alter.
If the face drifts, the fix is upstream. Go back to the selfie rather than adding more instructions here.
Step 4. Generate the video with Seedance 2.5
The
Final
node runs Seedance 2.5 in reference mode. All three images connect into it, and the settings on the node are as much a part of the shot as the prompt:
9:16 at 1080p.
Vertical, because the effect is being made for Reels and TikTok.
8 seconds.
Long enough for a full orbit plus the macro push in.
Reference mode.
This is the setting that tells the model to treat the connected images as things to match rather than as a first frame.
No audio.
Sound gets added on the platform, where trending audio lives.
Reference mode is the whole reason this template works. Without it you are animating a still. With it you are handing a model three fixed facts, the face, the outfit and the product, and asking for a shot that respects all of them.
The prompt in this node is a shot list rather than a description. It names the walk in, the freeze, the 180 degree orbit, the macro push in on the lettering, and the pull back out as normal speed returns, in the order they happen.
Then hit run. The whole chain costs 4 credits for each image node and 144 for the video, so a full pass is 152 credits.
The effect fails in predictable ways, and all of them are fixable before you spend the video credits.
Tips for a cleaner bullet time shot
Fix the freeze in the prompt, not in editing
Say the subject holds completely still while the camera moves. Without that sentence the model gives you a normal orbit and normal motion, which is just an arc shot.
Name the orbit in degrees
A slow 180 reads as bullet time. An unspecified camera move reads as drift.
Say constant speed
A camera that accelerates through the arc breaks the illusion, because the frozen subject then looks like a mistake rather than a choice.
Keep the freeze at a peak
Mid-stride, mid-turn, mid-jump. A subject frozen standing still looks like a paused video.
Regenerate the turnaround before you regenerate the video
Most identity and wardrobe drift in the final clip comes from a weak sheet, and the sheet costs a fraction of what the video costs.
Push in on one detail only
The macro beat works because it is a single object. Two close ups in eight seconds turns the shot into a montage.
Leave audio off until the platform
Trending audio is chosen where the clip gets posted, not where it gets made.
Variations to build from the same chain
The template ships as a fashion shot, and every variation below changes the canvas in a different way. One takes nodes out, one reverses what freezes, one adds video nodes, and one changes nothing except the settings.
Take the person out.
Delete the selfie and the turnaround sheet, and generate your product from four angles in the accessory node instead. The video node then orbits an object frozen in mid air, with nobody in frame. Two nodes instead of four, and the cheapest version of the effect to run.
Reverse what freezes.
Keep all four nodes and flip the instruction. Your subject walks at normal speed while the street, the traffic and the crowd around them hold completely still. Same camera orbit, opposite subject, and it is the time freeze video reading of the same bullet time video effect.
Fan out the video node.
Generate the turnaround sheet once, then connect three video nodes to it instead of one. Give each a different move, a 90 degree arc, a full 360, and a slow rise. One image spend, three clips, and the same person and outfit in all of them because they share a reference.
Change only the settings.
Leave the prompts alone and switch the video node to landscape with a longer duration. The vertical cut goes to Reels and TikTok, the wide one goes at the top of a product page, and both come from the same three images.
The freeze and the orbit are the two things every version needs stated outright. Everything else on the canvas is yours to add, remove or duplicate.
Get answers to common questions
It is a shot where the subject freezes or slows almost to a stop while the camera keeps orbiting at normal speed. The frozen subject and the moving camera together are what make it look impossible.
No. Slow motion slows the camera down along with the subject. Bullet time keeps the camera moving at full speed and takes the motion out of the subject instead.
Not for this build. The rig and the spinning 360 camera are the hardware routes to the effect. This template generates the shot from three reference images instead.
Two image nodes run on Nano Banana 2, one for the accessory with legible text and one for the identity and wardrobe turnaround sheet. The video node runs on Seedance 2.5 in reference mode.
Yes. The macro node is a normal image prompt, so describe your product and the text you want on it, and the video node will push in on that instead.
Start with your own selfie
Open the Fashion Bullet Time Effect template, upload a selfie, and swap the wardrobe image for the outfit you want to see. The two image nodes and the video node are already wired.
The shot took 99 cameras and a purpose built rig in 1999. Now it takes four nodes.
Open the template in Picsart Flow
The “it’s officially fall” trend is a season announcement built out of a run of short clips, each held about half a second, all pushed to the same warm grade, under one line of text that never moves. Nothing transforms and nothing is generated. The edit is the whole trick. Creator
@jessdoestravels
posted the version this drop is built from and it has passed 137,000 plays.
What is the “it’s officially fall” trend?
A new shot roughly twice a second.
Every clip is held between 0.37 and 0.53 seconds, so the cuts land on the beat and no shot outstays it.
They are clips, not photos.
A page turns, a dog crosses the frame, the camera drifts. Half a second of movement in each shot is what separates this from a photo carousel.
One caption for the whole run.
The line sits centred and static from the first frame to the last, so a pile of unrelated shots reads as a single sentence.
The grade does the joining.
The shots have nothing in common except colour. Every clip leans warm, which is what lets a fireplace and a rainy park belong in the same reel.
Half a second is under the boredom threshold.
No clip is on screen long enough to be judged, so the reel is over before attention drops.
The text carries the point.
The clips are the evidence and the line is the claim.
The whole build runs in the
video editor
, where you trim and split clips, add video filters, and edit music and audio on one timeline. If you would rather describe the look than dial it in, the
AI video editor
does mood and light transformation from a prompt. Autumn photos you already have can become clips in
image to video
, and anything you are missing gets generated in
AI Playground
, which puts 179 models from 32 providers behind one prompt bar.
How to make it in Picsart
1.
Shoot half a second of movement, not photos
Film short clips instead of taking pictures. A page turning, steam off a cup, leaves underfoot as you walk. That half second of motion is the difference between this and a photo carousel.
2.
Collect more shots than you need
A run this fast eats a lot of footage for very little runtime. Shoot across a week so you can drop the ones that fight the grade.
3.
Push every clip warm
Apply one filter across the whole timeline. Brightness and saturation can vary, they do in the source, but every clip has to lean warm or it falls out of the set.
4.
Cut on the beat at half a second
Trim each clip to roughly half a second and let the cuts land with the track. Keep the holds even, because one long shot stalls the run.
5.
Hold one line of text over everything
Centre a single caption and leave it there for the full run. Do not animate it and do not change it per clip.
The prompt
Autumn clip, paste with your photo
Animate this photo into a short handheld clip. Add one small natural movement and nothing else: steam drifting off a cup, a single leaf falling, the page of a book turning. The camera drifts very slightly as if handheld. Keep the colour warm, with amber and rust tones under soft overcast light. The movement stays subtle and the clip settles rather than builds. LOCKS: handheld drift only; one movement; warm grade; no people entering frame; no text; no logos or readable signage.
Getting the grade to match
Warm is the only rule.
In the source every clip runs red over blue, but brightness swings from a dim fireplace to a bright overcast park. Match the temperature, not the exposure.
Grade last, on the full timeline.
One filter over everything beats correcting every clip one at a time.
Cut anything that stays cool.
A blue-grey shot reads as a mistake even when the subject is right.
One grade, one line.
The “it’s officially fall” trend works because the grade does the joining and the text does the talking. Shoot short, cut on the beat, push it all warm. Open the
video editor
and cut the first one.
Qwen Image 3.0 Pro is the one to use when the picture has to look expensive. GPT Image 2 is the one to use when the picture has to say something. Both are flagship models and both are very good, but they were built with different jobs in mind, and that is what should decide it rather than which is newer.
Qwen Image 3.0 Pro sits at the top of Alibaba’s Qwen line, aimed at editorial and brand work: campaign visuals, styled portraits, product shots with real polish. GPT Image 2 comes from OpenAI, and its stand-out skill is getting words and information right inside the frame, on posters, packaging, signage and charts.
What each model is built for
Qwen Image 3.0 Pro is made for images that carry a look. It holds fine detail, renders skin and fabric and surface convincingly, and keeps a composition together when the prompt asks for a lot at once. The work it suits is the work an art director would recognise: a hero image for a campaign, a styled shot of a person, a product photographed like it matters.
GPT Image 2 is made for images that carry a message. Because it draws on everything the wider GPT models know about the world, it understands the context around a request, which shows up most when a picture needs to be correct as well as attractive. A poster whose headline reads properly. Packaging with the right words on it. A chart whose labels mean something.
Neither is a downgrade of the other. They are pointed at different halves of the work most creators do.
The two side by side
Qwen Image 3.0 Pro
GPT Image 2
What it is
The top tier of Alibaba’s Qwen line
OpenAI’s newest image model
Built for
Editorial and brand imagery
Images that have to carry information
The look it goes for
High polish, fine detail, styled
True to life color, real skin, cinematic light
Words inside the picture
Keeps a busy layout in order
Spells accurately in six writing systems
Working out your prompt
Rewrites and reasons through it
A mode you turn on
Changing an image you have
Describe the change
Describe the change, or extend past the frame
Biggest image
2688 pixels on the long side
2048 by 2048, with a 4K test mode
Where Qwen Image 3.0 Pro is stronger
Polish is the honest one word answer. Qwen Image 3.0 Pro resolves the small things that make an image look shot rather than generated, and it is built to hold up at the size and quality a real campaign asks for.
It is also the better bet for a long, detailed prompt. When you have specified a subject, a setting, a light, a wardrobe and a mood all at once, Qwen Image 3.0 Pro tends to still have all five standing at the end. That is partly the model and partly the two passes it runs first, described below, and it matters most on the kind of brief that arrives from a client already fully formed.
And it handles text well, which is worth saying because it is easy to assume only one model in this pair does. Qwen Image 3.0 Pro keeps lettering readable and spelling correct, and it keeps headlines, labels and captions where you put them. Its strength there is the arrangement rather than the individual characters.
Where GPT Image 2 is stronger
Words, first. GPT Image 2’s spelling is accurate about 99% of the time, it works in Arabic, Hebrew, Chinese, Japanese, Korean and Latin, and it copes with small print and with text that curves around a shape. If your image contains a sentence somebody will actually read, this is the model.
Then a specific kind of realism. The warm cast and slightly plastic skin that used to give AI images away are not there. Colors behave like real colors, skin looks like skin, light falls the way light falls, and the depth of field is convincing. It is a different quality from Qwen Image 3.0 Pro’s polish: less styled, more photographed.
It is also the more useful model when a picture has to be right. Product packaging with accurate names on it, a mockup of an interface with real elements in it, an infographic whose annotations you can read. Those are jobs where a beautiful image with garbled type is worth nothing.
Both work out your prompt before they draw
This is the thing the two models genuinely share, and it is why a short prompt gets you further than it used to on either one.
Qwen Image 3.0 Pro does two things first. Prompt-rewrite takes a bare prompt and fills in the gaps it needs. Thinking mode works through a complicated prompt properly before anything is drawn. GPT Image 2 offers the same idea as a choice: Instant Mode goes straight from prompt to picture, and Thinking Mode takes a moment to reason it through, which improves the structure, the layout and the accuracy.
So if you were hoping one of them would understand you better than the other, that is not where the difference lies. Both will take a two line prompt and give you back something properly composed. What they then do with it is where they separate.
Editing an image you already have
Both models can change an existing picture from a written instruction, so you do not need a second tool for revisions, and neither one asks you to draw a selection first.
Qwen Image 3.0 Pro edits from a description: say what you want different and it does that. GPT Image 2 commits to more by name. It will take something out, swap a background, restyle the whole picture, patch a single spot, or carry the scene out past the edge of the original frame. If widening a shot or replacing a background is the actual job, that is the safer choice.
One thing that matters on client work: every GPT Image 2 image comes with content credentials, a record of how the picture was made that stays with the file. Qwen Image 3.0 Pro says nothing about doing the same, and some brands now ask for that record.
Where to use both in Picsart
Both models are in
Picsart AI Playground
, where one prompt goes to both at once so you can see the two results next to each other. Both are in the
AI Image Generator
as well, and GPT Image 2 can be built into a chain of steps in
Flow
next to things like background removal and resizing. They share one credit balance, so trying both costs you nothing but a click.
The full list of what each one does is on its own page:
Qwen Image 3.0 Pro
and
GPT Image 2
.
Which one to pick, by what you are making
A campaign hero or a styled portrait.
Qwen Image 3.0 Pro, for the polish.
A poster, a label or any image with a sentence in it.
GPT Image 2, for the spelling.
Packaging or signage in more than one language.
GPT Image 2, which works across six writing systems.
A product shot meant to look photographed, not styled.
GPT Image 2, for the realism.
A product shot meant to look like a campaign.
Qwen Image 3.0 Pro, for the finish.
A long client brief with a lot to fit in.
Qwen Image 3.0 Pro, which holds the whole thing together.
A chart, a diagram or an interface mockup.
GPT Image 2, because the text has to be readable.
Widening a shot or replacing a background.
GPT Image 2, which names both.
Work that has to show where the image came from.
GPT Image 2, for the content credentials.
Get answers to common questions
GPT Image 2, when getting the words right is the hard part. Its spelling is accurate about 99% of the time, it handles Arabic, Hebrew, Chinese, Japanese, Korean and Latin, and it copes with small print and curved text. Qwen Image 3.0 Pro is good at a related but different thing: holding a busy layout together, so headlines, labels and captions stay where you put them.
Neither one outranks the other. GPT Image 2 leads on words inside the image and on a photographed kind of realism. Qwen Image 3.0 Pro leads on editorial polish and on holding a long, detailed prompt together. Pick by what you are making rather than by which model is newer.
Qwen Image 3.0 Pro, on its longest side, which reaches 2688 pixels against GPT Image 2’s 2048 by 2048. GPT Image 2 also documents a 4096×4096 mode still in testing.
Yes, both work from a written instruction and neither needs a mask or a manual selection. GPT Image 2 additionally names background replacement, object removal, restyling and extending the frame, so if that is the job it is the more certain choice.
GPT Image 2 describes its multi image results as consistent, which suits a storyboard or a product series. Qwen Image 3.0 Pro’s multiple results read more as different takes on the same prompt, which is more useful while you are still deciding.
No. Both work from an ordinary written description, and both fill in the parts of a prompt you left out. You bring the idea.
Start with what you are making
Look at the thing you owe somebody. If it has words in it that a person will read, open GPT Image 2. If it has to look like it came off a shoot, open Qwen Image 3.0 Pro. Then put the same prompt through both in
Picsart AI Playground
and let the two results settle it.
A MiniMax prompt has to answer one question before anything else. Which of the two models is going to read it? MiniMax H3 understands text, images, video, and audio in one context, and returns video with native stereo sound. MiniMax H3 Max reads words and frames only, and reads them fast.
That split decides what your prompt needs to say. Write a soundtrack into an H3 prompt and the clip arrives with audio already in sync. Write the same line for H3 Max and you have spent words on something its controls do not expose. So the ten prompts below are sorted by the model that runs them best.
Both models live in Picsart
AI Playground
, which means you can paste a prompt, generate, and switch models without changing anything else. Copy any prompt here, swap in your own subject, and run it.
What every MiniMax prompt has to carry
Both models want the same four things named plainly, in this order:
The subject and the setting.
One clear thing in one clear place. “A ceramicist at a wheel in a dusty studio” beats “an artist working”.
The action, with a beginning and an end.
A clip runs 5 to 15 seconds, so pick an action that fits inside one. One completed move reads better than three rushed ones.
The camera.
State whether it holds still, pushes in, or tracks alongside. Say nothing and the model chooses for you.
The light.
Time of day, direction, and hardness. This is the fastest way to change how a shot feels.
After those four, the models diverge. MiniMax H3 generates native stereo audio in the same pass as the picture, so an H3 prompt gets a fifth job: name what makes the noise. MiniMax H3 Max was tuned to follow a prompt more closely and to look better doing it, so its prompts reward precise camera and framing language instead.
There is one more habit worth building for H3. When you attach a reference, describe the relationship between that input and the clip you want, rather than just describing the clip. “Move the camera the way the reference video does, but around the man in the reference image” is the kind of instruction it was built to follow.
Which model should run your prompt
Reach for MiniMax H3 when the clip has to arrive finished. It generates 2K by default at 24 fps, produces matching stereo audio in the same pass, models multiple shots natively inside one generation, and takes reference images, video, and audio together. It is also strong at rendering legible text and brand marks, and at transferring motion from one clip to another. Clips run 5, 10, or 15 seconds.
Reach for MiniMax H3 Max when you are still deciding. It turns a prompt around in seconds at 480p or 768p, accepts any whole second count in the 5 to 15 range, and pins the first and last frame. There are no reference slots, so the prompt carries everything.
The practical order is to draft on H3 Max and finish on H3. Full specs for each sit on the
MiniMax H3
and
MiniMax H3 Max
model pages.
MiniMax H3 prompts: sound, references, and 2K
These six lean on what only H3 does. Each one either names its own audio or puts a reference to work. Notice how the sound line describes a source rather than a mood, because “a wooden rib scraping the clay” is something a model can render and “atmospheric” is not.
One continuous take with its own soundtrack
A ceramicist’s studio at first light, wet clay turning on the wheel, grey light coming through one dusty window. Her hands close around the rising wall of the pot and steady it, both thumbs pressing a groove into the rim. The camera holds at hand height and does not move. Sound: the low hum of the wheel, water sliding under her palms, a wooden rib scraping once against the clay, birds outside the glass. No music.
Several shots inside one generation
A night market in the rain, three shots in one clip. First a wide of the lane, canopies dripping, string lights doubled in the puddles. Then a close on a wok as the noodles hit the oil and flare up. Finally a medium of the cook handing a paper box across the counter to a customer in a yellow raincoat. Sound: rain drumming on canvas, the burst of the wok, low crowd chatter underneath.
Motion borrowed from a reference video
Take the camera movement from the reference video and apply it to a new subject: a lone red tractor parked in a harvested field at dusk. Match the reference for speed, direction, and the moment the move settles, but change nothing about my subject or setting to suit it. Low sun behind the tractor, long shadows across the stubble. Sound: wind across open ground, metal ticking as the engine cools.
A character locked by a reference image
Keep the woman in the reference image exactly as she is: her face, her hair, and the green corduroy jacket stay identical from the first frame to the last. Put her in a second-hand bookshop, walking the length of the aisle, pulling a paperback from a high shelf and reading the back cover as she keeps walking. Warm tungsten light, tall stacks either side, shallow depth of field. Sound: floorboards creaking under her boots, a page turning, a radio playing quietly at the front of the shop.
Text and a logo that stay legible
The camera creeps toward a cafe’s front window, early morning, the street still empty behind it. The words “OPEN FROM SEVEN” are painted on the glass in cream serif capitals, arched, and they stay sharp and correctly spelled for the whole clip. Reproduce that string exactly. Below it, a small circular logo in the same cream, centered under the arch. Soft overcast daylight, faint reflections of the street in the glass. Sound: a distant bus, a shutter rolling up somewhere off screen.
A spoken line matched to a reference voice
Have the mechanic in the reference image speak the line “It was never the alternator” in the voice from the reference audio, matching its pace and delivery rather than reading it flat. She is in a lit garage bay, oil on her forearms, wiping a wrench on a rag, and she looks up at someone off camera before saying it. Then she turns back to the engine. Medium shot at chest height, hard overhead work light, deep shadow behind her. Sound: her line clear over the ring of a dropped socket and a compressor cycling behind her.
MiniMax H3 Max prompts: fast, framed, and repeatable
These four are built for the drafting pass. They stay short on story and long on framing, because framing is what H3 Max has to work with. Run them at 5 seconds and 480p first, then rerun the one that works at 768p and full length.
Five seconds to test one idea
A cyclist crests a coastal road at golden hour, the sea on her left, dry grass bending in the wind. She stands on the pedals for the last of the climb, then sits back down as the road levels out. The camera tracks alongside her at the same speed, low and close to the wheels.
A product reveal pinned between two frames
Start frame: the closed box on a concrete surface. End frame: the bottle standing upright beside the open box. Between them, a pair of hands lifts the lid, sets it aside, and draws the bottle out in one continuous move. Keep the surface, the background, and the light identical from the first frame to the last. Camera locked off, no cuts.
Photo animation from a start frame
Start frame: the uploaded photo. Hold the framing, the clothing, and the light exactly as they are, then let the scene run forward. Steam lifts off the cup, the sitter turns her head toward the window and settles back, traffic crosses the street beyond the glass. The camera drifts in slightly and stops. Nothing else in the frame changes.
A camera move stated exactly
A slow dolly in on one empty chair in a school gymnasium, mid-afternoon, dust hanging in the light from the high windows. Start wide enough to see the painted lines on the floor and end tight on the chair back. Constant speed, no easing, no handheld shake. Muted colors, long shadows running away from the windows.
Prompt expansion, and how much to write
MiniMax H3 Max adds a control the prompts above assume you will touch. Prompt expansion has three modes: disabled, balanced, and quality. Disabled runs your words as written, which is what you want once a prompt is doing exactly what you asked.
Balanced and quality let the model elaborate, which helps a short prompt and can overwrite a long one. So the rule is simple. The more detail you have written, the further down that scale you should sit.
There is also a seed field. Fix it and a rerun of an identical prompt lands in the same place, which is how you change one word at a time and see what that word actually did.
Where to run these MiniMax prompts
Both models sit in
AI Playground
, where you can run one prompt through several models and compare the results side by side. That is the fastest way to see the difference between an H3 clip with sound and an H3 Max clip without it.
MiniMax H3 is also in the
AI video generator
if you want to paste a prompt and go. Either way, the rest of the
AI models
catalog is one click away when a shot calls for something else.
Get answers to common questions
A MiniMax prompt is the text you give MiniMax H3 or H3 Max to generate a video clip. A good one names the subject, the action, the camera, and the light.
Mostly the same, with one difference. MiniMax H3 generates native stereo audio, so its prompts should say what the shot sounds like. H3 Max has no reference slots, so its prompts carry the character detail instead.
MiniMax H3 generates 5, 10, or 15 seconds. H3 Max accepts any whole second count from 5 to 15, so 7 or 12 both work. Both run at 24 fps.
Not strictly, but you should. H3 generates audio whether or not you direct it, so an unwritten soundtrack is one the model chooses for you. Naming the source keeps that choice yours.
Describe the relationship between the reference and the clip you want, not just the clip. H3 reads text, images, video, and audio as one context, so “keep the face from this image and the camera move from this video” is exactly its kind of instruction.
It decides how much H3 Max elaborates on what you typed. Disabled runs your words as written, while balanced and quality fill in detail you left out. Use disabled for long prompts and quality for short ones.
MiniMax H3 and H3 Max both run in Picsart AI Playground, and MiniMax H3 is also in the AI video generator. No setup is needed either way.
Start generating with MiniMax
Pick the prompt closest to the shot you want, change the subject, and generate. Draft it on H3 Max, then run the winner through MiniMax H3 for 2K and stereo sound.
Try MiniMax in AI Playground
You can get paid for TikTok content before you have an audience, because paid campaign briefs pay on what a post does rather than on who posted it.
A brief is a job with the terms written down: what to make, where to post, and how the payout works. Nothing in it counts your followers. That makes it the route that is open on day one, while everything else is still building.
The short answer
Make the content, post it to your own TikTok account, get paid on how it performs.
That is the whole shape. You are not moving your audience anywhere, not pitching anyone, and not waiting to be big enough. You post the way you already post, and the work is attached to a payout before you start.
Why TikTok suits this particularly well
TikTok has said that follower count is not a direct factor in what its recommendation system shows people. A video from an account nobody follows can land in front of a large audience if it holds attention. Followers still help indirectly, since more people see your posts by default, but the feed is not checking your number before deciding who sees you.
That matters here because campaign payouts work on the same principle. Earn calculates what you make from real audience engagement rather than audience size, so the two line up neatly. TikTok can put your video in front of people who have never heard of you, and the brief pays on what those people did with it.
Compared with the other channels Earn supports, that is the friendliest starting point. Feeds built mainly around the accounts you already follow put a natural ceiling on a new creator, because reach grows roughly in step with the audience. An interest-based feed has no such ceiling. A first video can outperform a hundredth one on an established account, which is unusual anywhere else.
None of that makes the other platforms a bad idea, and Earn supports posting to Instagram, YouTube, and X as well. It just means TikTok is usually where early reach arrives fastest, and early reach is the thing campaign work converts into money.
Getting paid without a follower count
Picsart Earn
is built for exactly that gap. There is no follower minimum, no invite list, and no gatekeeping. Approval takes minutes rather than months, and what you earn depends on how your audience responds rather than how many people follow you.
Payouts come from real audience engagement, named on the program page as views, comments, shares, and reach on content you create and post yourself. Those are exactly the signals a strong TikTok post generates, whether or not anyone follows the account. You post on your own channels, TikTok included, so there is no separate platform to manage and no third-party posting. It is not a brand deal or a sponsorship either, which means no agency, no negotiation, and nobody taking a percentage on the way through.
Creators on the program have earned over $1 million in under 100 days, which is the practical version of the claim that talent beats follower count.
How the briefs work
You pick a campaign, make the post following the brief, then submit it and track it from the Earn dashboard.
Each brief on the
Earn campaigns
board publishes its reward logic before you start: which actions count, how they are verified, any cap on one submission, when payment lands, and how long the window stays open. You know what you are working toward before you spend a day on it.
Payout shapes vary. Live budgets have run from $500 up to $5,000, and finished briefs have paid up to $10,000 for a single post. Some drop performance entirely, including one that pays a flat amount per accepted video with no social account required and another that asks for an ad with no posting at all. Campaigns rotate quickly, so read the current brief rather than assuming.
Pick a brief close to what you already post
This is the part that decides whether any of it works.
Briefs cover product ads, tutorials, story pieces, design concepts, satisfying video, and clipping work built from an approved source pack. Take one near your usual subject and it is content you would have made anyway, now attached to a payout. Take one far outside it and you are learning a new subject while working to someone else’s rules, which is how half-finished submissions happen.
Make the work with what suits the brief. The
AI Editor
, the
Background Remover
,
Persona
, and
Aura
cover most of it, and clipping briefs are built in the
Picsart Video Editor
.
What actually moves your earnings
Three things, and audience size is not one of them.
Brief fit.
A campaign matching content you already make will get finished, and finished work is the only kind that pays.
Retention.
Payouts follow engagement, and engagement follows whether people watch to the end and pass it on. That is an editing problem more than a reach problem.
Real variations.
Performance-based work rewards whichever version lands. Three genuine attempts beat one polished piece, because the winner is rarely the version that felt best while you were making it.
Other routes that do not count followers
Campaign briefs are the fastest to start, but they are not the only option that ignores audience size.
Affiliate and commission.
You link a product and earn on sales. Audience size barely matters here, since it pays on how persuasive one video was rather than how many people follow you. It needs an audience that buys, though, not one that only watches.
Selling your own work.
Services, commissions, digital products, or anything you make and sell directly. Slowest to set up and the only route where you keep everything.
Content work for businesses.
Making videos for a company rather than for your own feed. Your following never comes into it, but you will need a few samples before anyone books you.
Each of those takes longer to produce a first payment than picking a brief, which is why the rest of this focuses on campaigns.
Frequently asked questions
For TikTok’s own Creator Rewards Program, 10,000 followers plus 100,000 video views in the previous 30 days, and you must be 18 or over in an eligible country. Other routes, including paid campaign briefs and affiliate links, set no follower requirement at all.
Not through the Creator Rewards Program, which gates on follower count. You can through campaign briefs that pay on performance, affiliate commission, or selling your own products and services, since none of those measure audience size.
Because the recommendation system and the payout program measure different things. TikTok has said follower count is not a factor in what the feed recommends, so a small account can reach a large audience. The Creator Rewards Program still requires 10,000 followers before it pays anything.
No. Creator Rewards requires videos of at least one minute. Duets, Stitches, ads, paid promotions, and sponsored content are also excluded, regardless of performance.
Paid campaign briefs, because they set no follower minimum and publish their payout terms before you start. Approval on programs like Earn with Picsart takes minutes, so the delay is how long the content takes rather than how long the account takes.
Start with one brief
Growing an audience is worth doing, and it makes strong performance more likely. It is just not the thing standing between you and your first payment.
Browse the open briefs on the
Earn campaigns
page, take the one closest to what you were going to post this week, and get paid for that while the account grows.
Use Qwen Image 3 Pro when the brief is still rough, and Qwen Image 2 Pro when the brief is already exact. Qwen Image 3 Pro thinks about your prompt before it renders. Qwen Image 2 Pro renders the prompt as written and resolves texture better than any tier below it. The split is labor, not quality.
Qwen Image 3 Pro is listed as Qwen Image 3.0 Pro, and Qwen Image 2 Pro is listed as Qwen 2 Pro. Both name pairs point at the same two models.
Both models are Alibaba’s, both sit in
Picsart AI Playground
, and both are configured identically there: the same five output resolutions, the same one to six images per run, the same reference image and aspect ratio controls, the same credit balance. The interpretation layer is the only thing that separates them.
The question that settles it: how finished is your prompt?
Qwen 2 Pro is the premium tier of Alibaba’s Qwen 2 image family. It takes a written brief and resolves the small stuff: the weave in a fabric, individual strands of hair, type small enough to squint at. Composition holds together on tightly structured prompts. The model handles fidelity. You supply the precision.
Qwen Image 3.0 Pro is the flagship tier of the Qwen image family, a generation past Qwen 2, and it adds two layers ahead of rendering. Prompt-rewrite expands a short brief into a fuller description. Thinking mode reasons through complex, multi-part briefs before any pixels are generated.
That is the whole trade. A brief that is already specific gets rendered faithfully by either model. A brief that is three words long, or one that carries six simultaneous requirements, is where the extra interpretation layer starts paying for itself.
Qwen Image 3.0 Pro vs Qwen 2 Pro at a glance
Capability
Qwen Image 3.0 Pro
Qwen 2 Pro
Place in the Qwen line
Flagship tier of the Qwen image family
Premium tier of the Qwen 2 image family
Prompt handling
Prompt-rewrite and thinking mode run ahead of rendering
Clean lettering and accurate spelling across posters, packaging, and UI mockups
Small type resolved as part of its fine detail
Best suited to
Rough briefs and briefs with many moving parts
Briefs that are already precise
What the interpretation layer buys you
Prompt-rewrite and thinking mode both run before a single pixel is generated, and the practical effect is the same in both cases: less of the brief is left to you. A thin instruction gets expanded into subject, setting, light, and framing. A brief carrying four simultaneous requirements gets reasoned through, so the fourth requirement is still standing at the end.
That has a cost worth naming. A prompt you wrote carefully is a prompt with your decisions in it, and an expansion layer will fill gaps you left deliberately. Handing a finished brief to Qwen Image 3 Pro means accepting interpretation you did not ask for, which is the exact reason Qwen Image 2 Pro is still the better instrument for a precise brief.
Six images per run is worth mentioning here only because it compounds the effect. Both models offer it, so the volume is not the differentiator. Six expansions of a thin brief explore genuinely different readings of it, while six renders of an exact brief return six versions of one decision.
What both models already share
Several things read like differentiators and are not. Skipping them saves a pointless model switch.
A reference image alongside the prompt.
Both model pages expose reference image input, so neither one is text-only.
Aspect ratio control.
Both prompt boxes let you set the frame before generating.
The same five output resolutions.
Both run 2048×2048, 2688×1536, 1536×2688, 2368×1728, and 1728×2368. Neither model reaches a size the other cannot.
The same batch sizes.
Both return one, two, four, or six images per run.
One credit balance.
Qwen Image 3.0 Pro and Qwen 2 Pro sit on the same balance as Seedream 4.5, Flux 2 Pro, and Imagen 4.0 Ultra, so switching costs nothing in setup.
Plain-language prompting.
Neither model needs design skills or parameter tuning. You write the brief, the model handles composition.
Legible text.
Both render readable type, and small lettering is a documented strength on each.
Commercial use.
Images from either model can be used for marketing, social, brand content, and e-commerce, subject to Picsart’s terms.
Where Qwen 2 Pro is still the right call
Qwen 2 Pro is not the older model you tolerate. It is the premium fidelity tier of its family, and it stays the better pick in a few specific cases.
The brief is already exact.
A prompt that already names subject, setting, light, lens, and palette does not need rewriting. Handing a precise brief to a rewrite layer adds interpretation you did not ask for.
Small detail is the deliverable.
Woven fabric, individual hairs, and type at small sizes are what this tier resolves best. Product photography and editorial work live on exactly those details.
The output is going to print.
This tier targets deliverables that get signed off: campaign leads, brand imagery, and files headed to a printer. Structured prompts hold their composition through it.
Match the model to the brief in front of you
A three-word idea and no time to write it up.
Qwen Image 3.0 Pro, for prompt-rewrite.
A brief with six simultaneous requirements.
Qwen Image 3.0 Pro, for thinking mode.
A fully specified shot list.
Qwen 2 Pro. The precision is already in the prompt.
A thin brief you want read several ways at once.
Qwen Image 3.0 Pro at four or six images, where each render expands the brief differently.
A poster or packaging layout where the type has to be readable.
Either one. Both reach the same frame sizes, so decide on how specified the layout brief already is.
A product close-up that lives on surface texture.
Qwen 2 Pro.
A look you have already locked, rendered consistently.
Qwen 2 Pro, which will not reinterpret what you specified.
Run both on the same prompt
One prompt through both models settles the comparison faster than any spec sheet. Write it precisely, so Qwen 2 Pro is playing to its strength, then see what the interpretation layer changes.
Same-brief test, both models
Editorial product photograph of a matte ceramic coffee cup on a pale travertine surface, warm window light raking from the left at a low angle, soft shadow falling right, shallow depth of field, visible clay texture on the rim, muted sand and clay palette, natural color, square frame
Then run a deliberately thin brief through Qwen Image 3.0 Pro alone to see what prompt-rewrite fills in.
Thin-brief test, Qwen Image 3.0 Pro
A coffee brand poster that looks expensive
Where each model lives in Picsart
Both models run in
Picsart AI Playground
, which is the practical place to compare them, since the same prompt can go to each one without leaving the prompt box. Both are also available in the
AI image generator
.
For the full specification on either model, the
Qwen Image 3.0 Pro model page
and the
Qwen 2 Pro model page
cover each one on its own. The broader
image models
lineup shows what else is available for a given job.
Get answers to common questions
Yes. Qwen 2 Pro is the product name for the premium tier of Alibaba’s Qwen 2 image family, and Qwen Image 2 Pro is the same model written out in full. Qwen Image 3 Pro and Qwen Image 3.0 Pro are likewise one model, where the “.0” is version notation.
Not automatically. The two models share their resolutions, batch sizes, and input controls, so the switch buys you prompt-rewrite and thinking mode and nothing else. That is worth it on rough or complex briefs, and worth skipping on briefs you have already written precisely.
Yes. Its model page lists natural-language image editing, and the prompt box accepts a reference image alongside the text, so you can supply a picture and describe the change. Image input is not exclusive to the newer model.
The same five options as Qwen Image 3 Pro: 2048×2048, 2688×1536, 1536×2688, 2368×1728, and 1728×2368. Output size is not a reason to pick between these two models.
Yes. It is the premium tier of the family and uses more credits per generation than Qwen 2. That trade is aimed at editorial and final-deliverable work where the detail ceiling is what matters.
Yes. Both sit on one balance alongside Seedream 4.5, Flux 2 Pro, and Imagen 4.0 Ultra, so testing a prompt through both costs nothing extra in setup.
No. Both work from natural-language prompts, and the model handles composition and detail. You bring the brief.
Start with the brief you already have
Read the prompt you were about to run. A brief that already names the subject, the light, and the frame is ready for Qwen 2 Pro. A brief that is still an idea belongs in Qwen Image 3.0 Pro, where the rewrite and thinking layers close the gap. Open
Picsart AI Playground
, paste it into both, and let the output decide.
The sliced outfit trend shows a full-length look assembling itself out of flat horizontal bands, with the background visible through the gaps until the last piece drops in. A band of torso hangs in mid-air with no head and no legs, and a moment later the figure is suddenly a whole person. Cut, new look, new street, same build. Creator
@ansyva
posted the version going around and it has passed 83,000 plays.
What is the sliced outfit trend?
Three looks, three builds.
Each outfit gets its own locked-off shot of three to five seconds, and each one starts in pieces.
The bands land out of order.
Not a top-to-bottom wipe. The torso is there from frame one, the head and shoes appear at the same moment, and the strip across the hip is the last thing to close.
It comes apart again.
The look is whole about a third of the way in, breaks back into bands, then rebuilds and lands complete just before the cut.
The gaps make you wait.
A missing band of body is a question, so the clip holds attention through a static shot.
The reveal is the outfit.
Every piece that lands is another part of the look.
It cuts itself.
Each shot ends with the figure complete, so the editor has no transition to invent.
The build is one photo per outfit, and it runs in
AI Playground
, which puts 179 models from 32 providers behind one prompt bar. Animate the still in
image to video
so the look assembles.
Kling V3 Turbo
is the pick because it takes a start frame and an end frame, so you hand it the sliced version and the finished version and it fills the middle. If you are generating the looks rather than wearing them, the stills come out of the
AI image generator
,
the photo editor
is where you erase the bands, and the clips get cut together in the
AI video editor
.
How to make it in Picsart
1.
Lock the camera and shoot head to toe
Put the phone on a tripod and do not touch it. Stand straight on and centred, with your shoes and some clear space above your head inside the frame.
2.
Keep the background plain and unbranded
A metal shutter, a stone wall, an empty pavement. The gaps put your background on screen exactly where a body would otherwise cover it, so signage and book covers become readable. Pick a wall over a window.
3.
Make the sliced start frame
Erase two or three horizontal bands across the body of your photo, one at the neck, one at the hip, one below the knee. Keep the background under them untouched.
4.
Animate from bands to whole
Upload your images, select an AI model, adjust settings, generate and download. The sliced frame is the start, the untouched photo is the end, and the model builds the figure across the gap.
5.
Land the look complete, then cut
Each shot in the source finishes as a whole person and holds there for about a second. Cut while a band is still missing and the next look reads as a mistake.
The prompt
Bands to whole, paste with your outfit photo
Animate this photo. The camera does not move at all and the background stays completely still. The person is missing several flat horizontal bands across their body, and the background is visible through those gaps. One by one the missing bands fill in with the correct part of the body and outfit, arriving out of order rather than top to bottom, until the figure is a single complete person standing in the pose. The feet stay in exactly the same place the whole time. Nothing else in the scene moves. LOCKS: locked camera; static background; feet fixed; no cuts; no text; no logos or readable signage.
Getting the bands to line up
Keep the cuts horizontal and parallel.
The effect falls apart the moment a band tilts.
Pin the feet.
The shoes land early and stay put, which tells the eye it is one person, not swapped parts.
Leave a gap inside the body, not just at the edges.
The strip at the hip closing last is what makes the build satisfying.
Variations worth trying
Take it apart instead.
Run the build backwards so a complete look dissolves into bands.
One plate, many looks.
Keep the identical background and swap only the outfit.
Bands that miss.
Let a band land offset by a few centimetres before it snaps into place.
Three looks, three builds, and the shoes never move.
The sliced outfit trend works because a missing piece is a reason to keep watching, and every piece that arrives is more of the outfit. Lock the camera, erase a few bands, and let the look put itself together. Open
AI Playground
and build the first one.
Three templates in Picsart Flow run Kling, and the interesting thing about them is not the prompt. It is what each one hands the model before the prompt is even read. One starts from a single photo. One starts from a character reference sheet. One starts from footage you already shot, plus a written script. That choice decides more about the finished clip than any adjective in the prompt box.
What the Kling settings on a Flow canvas actually decide
Every generation node in Flow carries a row of settings above the run button, and on these three templates that row is the real workflow. It names the model, the mode, the aspect ratio, the clip length, whether audio is generated, and which quality tier runs.
Read the three rows side by side and the pattern is hard to miss. The spectator template runs Frames mode at 5 seconds on the Pro tier with audio off. The kids song template runs 16:9 at 5 seconds with audio on. The hook template runs 9:16 at 15 seconds, also with audio on.
The settings shift because the starting point shifts. A template handing Kling one still needs Frames mode and leaves the sound for later. A template built to sing needs audio generated in the clip itself.
All three keep Multi-Shot off, so each node produces one continuous shot rather than a cut sequence.
Types of Kling workflow chains
The three templates are three chain shapes, and the shape is more useful to learn than the template.
Extend a single frame.
One image goes in, one styled still comes out of an intermediate node, and the video nodes animate that still. The chain is linear and short.
Lock an identity, then fan out.
A character reference sheet is generated first, then several video nodes run against that same sheet in parallel. The sheet is what lets separate generations agree on one character.
Fan out from footage.
A clip you already shot feeds several video nodes at once, each carrying different written lines. Nothing is generated from scratch, so the work is comparison rather than creation.
Chain shape
What you start with
What comes out
What Kling decides
Extend a single frame
One photo of a person
Two short atmosphere clips
Almost all the motion
Lock an identity, then fan out
A character and a song idea
Three verse clips plus a stitched video
Motion and performance, within a locked identity
Fan out from footage
Footage you shot, and a script
Several alternative openings
The least, since the frame and the words are set
Which Kling model to use where
Two Kling models appear across the three canvases. The spectator and hook templates run
Kling 3.0 Omni
, and the kids song template runs
Kling 3.0
.
The mode matters as much as the model. Only the spectator template uses Frames mode, which is what you reach for when a finished still already exists and the job is motion rather than invention.
Audio is the other fork. Leave it on when the clip has to carry its own sound, as the sung verses and the spoken opening do. Switch it off when the sound is arriving from an edit later.
No template locks you to its model. The picker sits on the node, Picsart runs a wider Kling line, and these are the video models worth knowing before you switch one:
Kling 3.0 Omni
for reference-based generation with native audio. Picsart lists Flow as one of the places it runs.
Kling 3.0
for precision motion, built for realism and detail in the movement itself.
Kling V3 Turbo
the faster V3 variant, for while you are still iterating rather than finishing.
Kling 2.6
the native-audio model, carrying voice and sound effects alongside motion control.
Photo to video: a broadcast-style clip from one still
The lightest starting point is one image. The
Trendy Sports Game Spectator Video
template takes an uploaded photo of a person, restyles it into a stadium broadcast still, then animates that still into two clips of about five and three seconds.
The canvas is only four nodes wide, which makes it the clearest example in the set. Image in, one styled frame, two video outputs. The prompt on the styling node instructs the model to read the attached photo, keep the person seated in the stadium seats, and hold their facial detail while the scene is rebuilt around them.
This is where Frames mode earns its place. The node is not asked to invent a scene from a description. It is handed a finished frame and asked what happens next in it, so Kling spends its whole budget on motion. Generating that still as its own node is the part worth copying.
Character consistency: three verse clips that keep one face
One step up, the starting point is an identity. The
Create a Cartoon Kids Song Video
template turns a character and a topic into a three verse sing along, and it solves the hardest problem in multi clip video before any video is generated.
The build starts with an input image of a character. That image feeds a node producing a character reference sheet showing the same character in three poses. Only then do the video nodes run, three of them, one per verse, each pointing back at that sheet. The three clips at five seconds each stitch into a full song video of fifteen seconds.
Two text nodes carry the writing. One holds the song topic in a single line of about eleven words. The other holds the lyrics for all three verses, around seventy words in total.
Keeping them separate is what makes the template reusable. Change the topic line and the lyrics, keep the character, and the whole song is new.
The reference sheet is the move to copy. Describing a character in words gets you a different character every generation, because each run reads the description fresh.
Handing every node the same sheet gives them one source of truth for the face, the proportions and the palette. Consistency becomes something the canvas enforces.
Try this prompt
A playful nursery rhyme scene in a new colorful setting, the character singing and dancing, gentle motion, audio on, no warping. Bright cheerful 3D storybook style, soft rounded shapes, pastel colors, warm friendly lighting, consistent character design across every shot.
Hook variations: several openings from one clip
The heaviest starting point is real video plus written words. The
Viral video hook generator
template starts from a clip you already shot and generates several alternative openings for it, so one piece of footage gets tested a few ways without a reshoot.
It is by far the largest canvas of the three. A single uploaded video fans out to four video outputs, alongside a text node holding scripted opening lines and a cluster of image nodes covering wardrobe and camera angle variations.
The settings tell you what it is for. Vertical at 9:16, fifteen seconds, audio on, which is the shape of a short form opener with a spoken line in it.
Fifteen seconds is three times the length of the other two, because an opening has to deliver a sentence rather than a mood. Comparing finished variants side by side beats iterating on one node and trying to remember the last version.
How to run any of these Kling templates
All three canvases follow the same working order. Open
Picsart Flow
and load the template closest to the material you already have.
1.
Open the template and look before you run
The nodes arrive already wired. Trace the chain left to right so you know which node produces what.
2.
Read the settings row on every generation node
Model, mode, ratio, length, audio and tier are set here, and they matter more than the prompt wording.
3.
Replace the input
Swap in your own photo, character image or footage. This is the asset the rest of the canvas reads.
4.
Edit the text nodes
Each template has at least one, holding a topic line, lyrics or scripted lines. Keep the existing structure, since the video nodes are wired to it.
5.
Run the intermediate node first
A styled still or character sheet sits between input and video on two of the three chains. Approve it before spending a generation on motion.
6.
Generate the video nodes
Run them one at a time so you can judge each output on its own rather than waiting on the whole canvas.
7.
Check the final node
That is a stitched video, a pair of clips or a set of variants. Watch it through and confirm nothing drifted.
8.
Re-run single nodes, then export
Every output regenerates on its own, so fix the weak one, not the whole build. Exporting usually needs you to be signed in.
Tips for steering Kling in Flow
Set the clip length before you write the prompt
Five seconds and fifteen seconds want very different amounts of action in one continuous shot.
Turn audio off when you plan to score the clip later
A generated soundtrack fights a voiceover or track you add afterwards.
Lock identity in an image, not in a description
A reference sheet survives across separate generations in a way that written details do not.
Give each node one job
All three templates split styling and motion into separate steps rather than asking for both at once.
Re-run one node instead of the whole canvas
Every output regenerates on its own, so a single weak clip is cheap to fix.
Get answers to common questions
It is a canvas where one or more generation nodes run Kling, wired to the inputs that feed them. The input, any styling steps and every output sit together, so the whole build stays visible and each part can be re-run on its own.
The settings row on each canvas names it. The spectator and hook templates both run Kling 3.0 Omni. The kids song template runs Kling 3.0. You can change the model on any node, since the picker is part of the node rather than fixed by the template.
Only for the hook template, which is built around a clip you already shot. The spectator template needs one photo, and the kids song template needs a character image and some writing. Nothing there requires a camera.
Generate a character reference sheet as its own node, then wire that sheet into every clip that features the character. The kids song template does exactly this before it generates any video, which is why one character holds across three separate verses.
Turn it off whenever sound is arriving from somewhere else, such as a voiceover or an edit you will assemble later. The spectator template keeps audio off for this reason. Leave it on when the clip carries its own sound, as the sung verses do.
Yes. Ratio, length, audio and tier are set per node, so you can take a landscape template vertical without rebuilding anything. Changing the length usually means rewriting the prompt to match, since a shorter clip holds less action.
Start with the chain shape that matches your material
Pick the template by what you already have rather than by which output looks nicest. A photo, a character, or a clip you shot this morning each has a chain built for it. Open
Picsart Flow
to browse the templates, or go to the
Flow editor
and wire your own.
Qwen Image 3.0 Pro
is available now in the
Picsart AI Playground
and the
AI image generator
. It is Alibaba’s flagship image model, a generation on from Qwen 2, and it does two jobs: it makes images from a text prompt, and it changes images you already have.
The thing worth knowing before you start is that it has no low quality setting. Every resolution it offers runs at roughly 4.2 megapixels, from a 2048 by 2048 square up to wider landscape and portrait shapes. Most models ask you to trade resolution for speed. This one does not offer the trade.
That design points at what it is for. Fine detail survives at this size, which matters most when your image contains things that fall apart at lower resolutions: text on a sign, panels in a storyboard, items on a menu, small elements inside a busy layout.
Here is what the model gives you and how to get the most out of it.
What Qwen Image 3.0 Pro does
Three things, all from the same model.
It generates from text.
Describe what you want and the model builds it. The output is aimed at work that gets published rather than work that gets posted, so think brand imagery, ad creative, and anything where a client will look closely.
It also has unusual room for instruction. Prompts run to around 4,500 tokens, which is thousands of words of direction. Short prompts work, but you are leaning on the model to guess rather than telling it what you want.
It edits what you already have.
Hand it an image, say what should change in ordinary language, and it makes the change. You stay in one model for the whole job, so a revision never means rebuilding your look somewhere else.
It gives you several options at once.
Ask for one, two, four, or six versions of the same prompt. Six is the fastest way to learn whether an idea is working, because you see the range the model is capable of before you commit to refining any single result.
Prompt-rewrite and thinking modes
Two things happen between your prompt and your image, and together they are most of what separates this generation from the last one.
Prompt-rewrite
takes a thin prompt and fills it out before anything is generated. It can do that on its own or hand the job to an agent. If you write three lines and get back a fully realised scene, this is the part responsible.
Thinking mode
gives the model a chance to work through a complicated brief first. A prompt carrying several subjects, a fixed layout, and text in three separate places is exactly the sort of request that collapses without it.
Neither does much for a simple prompt. There is little to expand and little to work out. On a demanding one, they are the difference between a near miss and the thing you asked for.
Everything happens in one bar at the bottom of the Playground. If you want to see what else is available first, the full
model catalog
is one click away, and the
AI image generator
runs the same model.
How to use Qwen Image 3.0 Pro
1.
Open the Playground and choose Image
The switcher on the left moves between video, image, and music.
2.
Select Qwen Image 3.0 Pro
It shows up in the model dropdown as Qwen 3.0 Pro.
3.
Pick your resolution
Five options, from a 2048 by 2048 square to wider landscape and portrait shapes.
4.
Set how many images you want
One, two, four, or six per run.
5.
Write your prompt
Type it, or press the microphone and say it. If you are short of ideas, Inspire me will fill the box for you.
6.
Add a reference image
Only if you are editing rather than generating from scratch.
7.
Generate
Pick the strongest result, or run it again with a tighter prompt.
Where you can use it
The Playground is the easiest place to start, mostly because you can run the same prompt through Qwen Image 3.0 Pro and 170+ other models and see the difference side by side. Nothing needs configuring first.
It also travels. Use it on the web, in the
Picsart desktop app
, or wire it into your own projects through the CLI, MCP, REST API, and SDK. You will not need an API key from anyone else or a second subscription.
Credits are shared across the whole catalog too. Moving between this and
Qwen 2 Pro
,
Seedream 4.5
,
Flux 2 Pro
, or
Imagen 4.0 Ultra
costs you a dropdown, not a new tool.
Choosing your resolution
Five presets, three shapes.
Resolution
Shape
Ratio
Use it for
2048 x 2048
Square
1:1
Feed posts, profile art, product tiles
2688 x 1536
Wide landscape
1.75:1
Banners, headers, cinematic scenes
1536 x 2688
Tall portrait
1:1.75
Stories, Reels covers
2368 x 1728
Landscape
1.37:1
Editorial spreads, print layouts
1728 x 2368
Portrait
1:1.37
Menus, posters, book covers
The two landscape options differ more than the numbers suggest. 2688 x 1536 is the wider, more cinematic crop. 2368 x 1728 sits closer to a classic photo shape and leaves more vertical room, which helps when your subject is tall or your layout needs headroom.
One practical note on vertical. 1536 x 2688 is the closest thing to a 9:16 social frame, but it is not exactly on it, so plan for a small crop if you are posting straight to Stories or Reels. 1728 x 2368 is a gentler portrait and suits print style layouts better than social.
Text that survives the render
Ask most image models for a poster with a headline on it and you get a poster with something headline shaped on it. The letters are almost right. Nothing is spelled correctly. You either prompt again or open a design tool and place the type yourself.
This model was built to avoid that. It holds type down to around 10 pixels, and it spells things properly, which is what makes packaging, menus, ad creative, and interface mockups possible as single generations rather than as generation plus cleanup.
The two modes above are doing quiet work here. Long briefs are where text usually wanders off, and reasoning through the layout first is what keeps a caption attached to the thing it was captioning.
There are 12 languages and more than 20 typefaces available natively, which matters if one campaign has to ship in five markets, or if the font itself is part of how the brand is recognised.
When words are load bearing rather than decorative, that reliability is the whole difference between an asset you can use and an afternoon of edits.
Editing an existing image
Attach your image, then write the instruction. Two habits make the difference here.
Say what should stay, not just what should change. Editing models drift on the parts you leave unmentioned, so if the face, the hair, or the product needs to survive the edit, put that in the prompt explicitly.
Then describe the target, not the delta. “A champagne silk blouse under a tailored grey blazer” gives the model more to work with than “make the top nicer.” Detail on the destination beats instructions about the journey.
Because generating and editing live in the same place, a second pass does not cost you the look you established in the first.
What it is built for
Any model can produce a nice looking picture. This one is aimed at a harder job: pictures that have to carry information as well as look good.
Busy layouts, one generation.
It can place images inside images, so a composition with many parts arrives assembled rather than needing to be built. Storyboards, menus, magazine spreads, and newspaper pages are all within reach.
Screens and interfaces.
It can imitate web pages, games, and live stream layouts convincingly enough for app screens, product UI, and stream overlays, which saves opening a design tool to fake them.
Detail that holds up close.
Expression, skin texture, individual hairs. This is usually where generated people give themselves away, with skin sanded smooth and hair rendered as a single object.
If you want something simple and you want it now, a lighter model will be quicker. Reach for this one when somebody is going to look closely.
Try Qwen Image 3.0 Pro
Open the
Picsart AI Playground
, pick Qwen Image 3.0 Pro, and give it something with detail in it.
Get answers to common questions
It is the top tier of Alibaba’s Qwen image family, available in the Picsart AI Playground and AI image generator. It generates images from text, edits images you supply, expands and reasons through prompts before generating, and can return up to six results at once.
No, and the names are easy to mix up. Qwen3 is Alibaba’s text and chat model. Qwen Image 3.0 Pro is its image model. Separate product lines that happen to share a family name.
Yes. The two spellings point at the same model, and the “.0” carries no meaning beyond version numbering. Both turn up in the wild.
It is the newer generation. It keeps the image quality Qwen 2 Pro was known for and adds prompt expansion, reasoning before generation, editing from a supplied image, and batches of one, two, four, or six.
Prompt-rewrite fleshes out a short prompt into a fuller description before generating, either by itself or through an agent. Thinking mode has the model work through a complicated brief first. Both pay off on demanding prompts and do little on simple ones.
Up to 2048 x 2048, plus 2688 x 1536, 1536 x 2688, 2368 x 1728, and 1728 x 2368. Every one of them runs at full quality, so there is nothing to trade down to.
Twelve languages natively and more than 20 typefaces, so one layout can ship to several markets without changing model.
Around 4,500 tokens, which runs to thousands of words. Detailed prompts are where this model does its best work.
Three connected nodes do it. A source holding your material, a text node that summarizes and structures it, and an image node that turns that summary into a visual.
That chain is the entire method. It does not change whether you feed it a spreadsheet, an article, or an entire book.
Most long documents fail the same way. The information is useful and nobody reads it. A report, an article, even a spreadsheet of sales numbers can hold everything a team needs and still get skipped, because getting through it is work.
This guide walks through the text to infographic workflow in
Picsart Flow
, using three source types: a sales spreadsheet, a text-heavy article, and a book.
Same three stages every time, whether you are making content for social media, presenting to your team, or building educational resources.
The three-node chain behind every infographic
The workflow has three stages, and understanding them is what lets you point it at anything.
Your source material.
A spreadsheet, an article URL, a book. This is the raw input, and it goes on the canvas first.
A text node.
This is where the summarizing and structuring happens. You are not only asking for a shorter version.
One or more image nodes.
These take the structured summary and draw it.
The middle stage is the one people underestimate. The text node does not only shorten the material. It reads it and proposes a visual treatment, and that proposal is what the image node follows.
That is what separates this from a single prompt. Most text to infographic tools take your topic and guess at a design. Here the brief is built from your own content, then handed to the image node.
Everything stays connected on the canvas, so nothing sends you back to a blank start. Without rebuilding the workflow you can:
Change the summary instruction and rerun the image node.
Swap the output size and generate the same content in a new layout.
Add a second image node beside the first for an alternative treatment.
Try a different image model against the summary you already have.
Run several generations and keep the one that works.
Three source types run through it, and each one is covered below:
An article URL
becomes one illustrated infographic, then four social posts.
A book
becomes a chapter-by-chapter illustrated series.
A spreadsheet
becomes a presentation-style graphic sized for slides.
Read the one closest to your material, then follow the walkthrough at the end.
Turn an article into a week of social posts
A long article is the clearest case. The text is already written and the only problem is that it is long.
Paste the article URL into a text node. There is no need to copy the body across, because the link is the input. Then ask for three things in one instruction:
A section count.
Four clear sections gives the image node a structure to lay out.
The output it is for.
Optimized for an infographic, not for a summary.
A visual direction.
For a historical subject, medieval illuminated manuscripts.
Connect that summary to an image node and you get one fully illustrated infographic covering all four sections. Test portrait, square, and landscape to see which layout carries it best.
Then the part that changes the economics. Connect a new image node for each of the four sections and ask Flow to illustrate them separately.
One long article becomes four standalone graphics. Each works as its own social media infographic, and they share a style because they came from the same summary.
Turn a book into an illustrated classroom series
Hundreds of pages is the hardest version of the problem, and the workflow does not change.
Ask Flow to summarize the key moments into easy-to-follow sections, then connect that to an image node and prompt for a style that suits the source. The example here is The Odyssey, so the prompt asks for ancient Greek pottery.
It is worth running a few image models against the same summary. Illustration styles vary a lot between them, and a book gives you enough content to see the difference clearly.
To make a series, connect 10 image nodes, one per chapter. From the same workflow you get an illustrated carousel or a classroom resource, chapter by chapter, in one consistent style.
That is the strongest argument for the node approach. Ten chapter graphics made individually would drift. Ten hanging off one summary do not, because they share a parent.
Turn a spreadsheet into presentation slides
Data is where the text node earns its place, because numbers need interpreting before they can be drawn.
Upload a spreadsheet of sales data into an image node. You can read the numbers, but it is not presentation ready. Connect a text node and ask AI to summarize the data and recommend the best way to visualize it as an infographic.
What comes back is a structured summary plus the design brief you did not have to write:
The key insights worth highlighting.
The charts that would communicate them clearly.
The icons to use alongside them.
The layout that would hold it together.
Now connect both the spreadsheet and the summary to an image node, ask for a clean presentation-style infographic, set the size to 16:9, and run. A page of numbers reads at a glance.
These steps are identical whichever source you started with, so read them once and apply them to anything.
How to turn text into an infographic in Flow
1.
Open Flow and create a new workflow
From the Picsart homepage, head into Flow and create a new workflow. You are starting on an empty canvas.
2.
Add your source material
Upload your file or paste your link. A spreadsheet is uploaded into an image node, an article URL goes straight into a text node, and a PDF creates a Document node when you drop it on the canvas.
3.
Write the summary brief
Create a text node and connect it to your source. Ask it to summarize the material, and be specific about the shape you want back. Four clear sections optimized for an infographic gives a very different result from summarize this.
4.
Add your visual direction
In the same instruction, describe the look. Ancient Greek pottery for The Odyssey, medieval illuminated manuscripts for a historical article, clean and presentation-style for data.
5.
Create an image node and connect everything
Connect the summary to an image node. For data, connect the original source as well, so the image node has both the numbers and the interpretation.
6.
Set your model and size, then run
Choose your preferred image model and set the size. Use 16:9 for presentation slides, or portrait, square and landscape for everything else. Add more image nodes off the same summary to turn one source into a series.
Small choices in the text node change the output more than anything you do downstream.
Tips for a clearer infographic
Ask for a section count
Four clear sections is a structure. Summarize this is not. The number you give is what the image node lays out.
Put the visual direction in the summary brief
Not only in the image prompt. The summary then arrives pre-styled and the illustration follows it.
Match the style to the subject
Greek pottery for The Odyssey, manuscripts for a historical article. Arbitrary styles look arbitrary.
For data, connect the source as well as the summary
The image node needs the numbers, not just the description of them.
Test layouts and models on one summary
Portrait, square and landscape carry different amounts of text, and illustration styles vary a lot between image models.
Branch instead of restarting
A second image node off the same summary costs one connection, and everything hanging off it shares a style.
One more input: turn a PDF into an infographic
Everything above assumes you are pasting a link or uploading a file. There is now a fourth way in.
Drop a PDF on the canvas and Flow creates a Document node holding it. A report, whitepaper, or research paper can feed the same chain with no copy and paste.
The Document node has three output lanes. To turn a PDF into an infographic you want
Text
, which carries the full extracted text into your text node. From there it runs exactly as an article would.
Image
carries the open page, and
All pages
carries every page as a batch.
PDFs up to 50 MB and 100 pages are supported, which covers most reports you would otherwise have to read in full.
Start turning your text into visuals
Pick the longest thing on your desk. A report nobody opened, an article you want to promote, a chapter you need to teach. Open
Picsart Flow
, put it on the canvas, and wire it through a text node into an image node.
Three connections and one run is all it takes to find out whether the material is worth building out:
Paste the link, upload the file, or drop the PDF.
Ask the text node for a set number of sections and a visual direction.
Connect an image node, set your size, and run.
The workflow does not change, so the only thing that changes is what you point it at.
Get answers to common questions
Connect three nodes in Picsart Flow. Add your source material, connect a text node that summarizes and structures it, then connect an image node that turns that summary into a visual. The same chain works for a spreadsheet, an article, or a book.
More than shorten the text. Ask it to summarize a spreadsheet and it returns a structured summary of the key insights, then suggests the charts, icons, and layout that would communicate the information more clearly. That recommendation becomes the brief for the image node.
No. Paste the article URL into the text node and Flow works from the link. There is no need to copy the body text across.
Yes. Summarize the article into four clear sections, then connect a new image node for each section and ask Flow to illustrate them separately. One long article becomes four standalone graphics that share a visual style.
Yes. Dropping a PDF on the canvas creates a Document node. Wire its Text lane into your text node and the workflow runs exactly as it would with an article. PDFs up to 50 MB and 100 pages are supported.
A manga storyboard is the rough panel plan an artist works out before inking a single finished page. It fixes the shots, the order, and the pacing while everything is still cheap to change.
You can build one without drawing it. Not from a single prompt, though. A storyboard is a sequence, and sequences fall apart when you ask one prompt to produce all of them at once. What works is a workflow, built as connected steps on a canvas in
Picsart Flow
.
Here is what a manga storyboard has to do, why it needs a chain rather than a prompt, and how to build one.
What a manga storyboard is
A manga storyboard is a planning document, not artwork. Panels are loose, faces are simple, and the point is to see whether a sequence reads before anyone commits to finished linework.
It answers a short list of questions:
Where does the eye enter the page?
Panel one has to be unmissable.
Which beat carries the most weight?
Emphasis is how a page tells you what matters.
What does the reader see first, and what a half second later?
Order is the whole craft.
That makes it different from a finished manga page in one important way. A storyboard is allowed to be ugly. What it is not allowed to be is unclear.
How the six-panel grid carries an action scene
Six panels is a workable unit for one beat of action. It gives you room to set up, escalate, and land without asking a single page to carry a whole chapter, and a two-by-three grid reads in order without anyone having to think about it.
A shonen action beat usually breaks down roughly like this:
Panel 1. Arrival.
The character enters the space. Wide, calm, no threat visible yet.
Panel 2. The trigger.
They touch, open, or step on the thing that starts it.
Panel 3. The reveal.
The threat appears at full scale, framed from below so it fills the frame.
Panel 4. Reaction.
Close on the face. Shock, not action.
Panel 5. Impact.
The exchange itself, drawn at the sharpest angle on the page.
Panel 6. The standoff.
Pull back wide. Scale restated, outcome still open.
Notice how much of that is camera work rather than plot. The panels move between wide, close, and low angle on purpose, because six panels of the same framing read as one flat panel repeated six times.
Generated panels come out the same size as each other, so you cannot enlarge the reveal the way a printed page would. Emphasis has to come from shot scale instead. That is why panel three is framed wide and low, and panel four sits tight on a face.
Style matters as much as framing. Black-and-white manga leans on hatching, speed lines, and heavy blacks, so contrast does the work color normally would.
Why this is a workflow and not a prompt
Ask one prompt for six manga panels and you get six strangers. The model has no reason to draw the same face twice, because nothing in the request tells it that panel four is looking at the person from panel one.
A workflow fixes that by splitting the job into steps, where each step hands its output to the next one as an input. The character is decided once, early, and everything downstream inherits that decision instead of re-inventing it.
That is what Picsart Flow is for. It is a no-code AI workflow tool built on an infinite canvas, where you connect AI models to each other and run the whole chain as one piece. Images, text, and video nodes sit side by side, so a picture generated in step two is available as a reference in step three.
The model library suits this kind of job. It includes an image model tuned for storyboarding, a character reference model built for consistent character rendering, and video models picked for narrative control. Scene continuity is handled at the workflow level, keeping lighting and character states steady from the first frame to the final cut.
Make your manga storyboard step by step
Open a canvas in Flow and build it left to right. You can also start from a ready-made template, such as the
manga storyboard workflow
, and change the prompts rather than wiring the chain yourself.
Start the canvas with one reference image.
Upload a photo of a person or generate a character. Either way this image is the decision everything else inherits, so make it clear and well lit before you move on.
Chain a text node to it and restyle into manga.
Ask for black-and-white hand-drawn manga, and name what has to survive: facial features, hairstyle, clothing, pose, and the setting behind them. Say what must not change as explicitly as what must.
Set the art register in that same step.
Mature shonen or seinen proportions mean realistic eye size and a defined jawline. Soft, rounded, youthful styling is a different genre. Decide it here, once, rather than per panel.
Feed that output forward as the character reference.
This is the step people skip. Connect the manga image you just made into the next node, so the grid is built from a picture rather than a description.
Generate the six-panel grid.
Prompt for a two-by-three storyboard of one titled beat and name the shots you want by type, since the model will otherwise settle into a single framing.
Chain a video node to the storyboard.
Give it an opening frame to animate from and the grid as the running order, and tell it to use the grid for sequence and pacing only. Without that, storyboard grids tend to animate as grids, panel borders included.
Iterate on the canvas, not from scratch.
Change one node and rerun from that point. Upstream steps keep their output, which is what makes a workflow cheaper to refine than a prompt.
What keeps the character consistent across six panels
Consistency is the whole difficulty here. Each panel is a separate generation, and nothing forces panel four to remember what panel one decided.
The chain solves most of it structurally, by passing a real image forward instead of a description. The prompt on the grid node closes the rest:
Try this prompt
The character in every one of the six frames must be the exact same person shown in the reference image, with the same facial features, face shape, eye shape, hairstyle and hair length, body proportions, and the exact same outfit in every panel, with zero variation across scenes. Do not reinterpret, restyle, age up, age down, or alter the character’s identity between frames. Treat the reference image as the single source of truth for the character’s appearance throughout the entire storyboard.
Three things in there do the work:
It lists the features that actually drift.
Hair length and outfit details go first, so they get named first.
It forbids the failures by name.
Restyling and ageing up are the two most common, so both are ruled out explicitly.
It appoints a single source of truth.
With one image named as the authority, the model is never choosing between references.
Manga storyboard examples to try
The same chain carries a lot more than one story. Change the beat description on the grid node and rerun from that point:
The boss battle.
A character reaches a ruined shrine, touches something they should not, and a stone guardian tears itself out of the ground.
The awakening.
An ordinary character discovers the power. Panels build from a quiet moment to a full surge.
The rooftop chase.
Pure movement. Every panel a different angle on the same run, with the gap closing.
The duel opening.
Two fighters, six panels, and no contact until the last one. All tension, no impact.
The rescue.
Something is falling and someone is running. The middle panels hold the distance between them.
The mentor’s test.
A training bout that is losing on purpose. Reaction panels carry the story, not the hits.
The arrival of the rival.
A new character walks in. The existing cast reacts across four panels before you show the newcomer’s face.
Each of those is a one-line change. The reference image, the art register, and the chain itself stay where they are, which is the point of building it as a workflow.
Tips for a stronger manga storyboard
Pick the reference image for the face, not the outfit
Clothing is easy to describe in a prompt. A face the model cannot see clearly will drift across the panels.
Name the beat, not the whole story
Six panels hold one moment. A prompt that describes a three-act plot returns six unrelated panels.
Vary the shot type explicitly
Ask for wide, close, and low angle by name. Left to itself the model settles into one comfortable framing.
Keep it monochrome
The moment color arrives it stops reading as manga. Contrast and hatching are doing the styling work.
Change one node at a time
Rewriting several steps between runs makes it impossible to tell which edit fixed the sequence.
Judge the grid before you animate
Rerunning one image node is cheap. Regenerating a video because panel two was wrong is not.
Get answers to common questions
A manga storyboard is a rough panel-by-panel plan of a manga page or sequence, made before any finished artwork. It settles the shots, the reading order, and the pacing while changes are still cheap.
Enough to carry one beat. A two-by-three grid of six panels gives you room to set up, escalate, and resolve a single moment of action without crowding the page.
No. The workflow generates the panels from a reference image and a written beat, so the skill it asks for is describing shots rather than drawing them.
No. Picsart Flow is a no-code visual builder, so a workflow is made by connecting nodes on a canvas. You can also describe what you want in plain language and let the built-in AI Copilot configure the steps and pick the models.
Pass a generated image forward as the reference rather than describing the character again, then name the features that must not change, including face shape, hairstyle, body proportions, and outfit.
Yes. Chain a video node to the finished grid and use the storyboard as the running order rather than filming the panels themselves.
Yes. A workflow can be saved and rerun, so the same chain produces a new storyboard from a new reference image and a new beat description.
Start building your manga storyboard
Open a canvas in Flow, drop in one reference image, and get the character right before you generate a single panel. Everything downstream inherits that decision. Community-built chains to start from and remix sit in the
Picsart Flow templates
library.