cuesora

Models

Every model, what it is for, and what it costs

Credits are shown at each model's default settings. Longer clips, higher resolution and more images cost proportionally more — the exact price is always on the button before you run.

Photo

GPT Image 2

Cuesora pickNew

OpenAI's newest image model: it follows an instruction to the letter and writes legible text.

Use it for: layouts with copy, product mockups, anything with exact instructions

text · photo8 shapesUp to 16 referencesUp to 4 images3 quality levels

≈ 85 cr at defaults · from 3 cr

AuraFlow

An open model with a loose, painterly hand — forgiving of a vague prompt.

Use it for: concept art, abstract backgrounds, moodboards

textUp to 2 images

≈ 2 cr at defaults · from 2 cr

FLUX Dev

The workhorse: clean renders with composition you can predict, from text or from a photo.

Use it for: everyday image work, illustration, product shots

text · photo1K · 1440p8 shapesUp to 4 images

≈ 10 cr at defaults · from 10 cr

SD 3.5 Large

Stable Diffusion at its largest — holds a composition and keeps lettering readable.

Use it for: posters with text, typography, layouts

text1K · 1440p8 shapesUp to 4 images

≈ 26 cr at defaults · from 26 cr

Nano Banana

Google's quick editor: natural light, and edits that leave the rest of the photo alone.

Use it for: photo retouching, putting a product in a scene, quick realism

text · photo10 shapesUp to 10 referencesUp to 4 images

≈ 16 cr at defaults · from 16 cr

Nano Banana 2

Nano Banana with the controls opened up — shape, scale, and a strip of photos to work from.

Use it for: composites from several photos, brand visuals, print-ready art

text · photo1K · 1440p · 2160p10 shapesUp to 10 referencesUp to 4 images

≈ 32 cr at defaults · from 32 cr

FLUX Pro v1.1

FLUX Pro, quicker and sharper, with the shape of the frame in your hands.

Use it for: advertising, editorial scenes, detailed renders

text1K · 1440p8 shapesUp to 4 images

≈ 16 cr at defaults · from 16 cr

FLUX Pro Ultra

FLUX at its most photographic: grain, depth and skin that reads as a camera, not a render.

Use it for: photoreal portraits, large prints, campaign visuals

text9 shapesUp to 4 images

≈ 24 cr at defaults · from 24 cr

Nano Banana Pro

Google's top image model — what we reach for when text and faces must survive an edit.

Use it for: text on an image, consistent characters, complex edits

text · photo1K · 1440p · 2160p10 shapesUp to 14 referencesUp to 4 images

≈ 60 cr at defaults · from 60 cr

GPT Image 1.5

The previous OpenAI model — still strong on long, specific prompts and on lettering.

Use it for: infographics, mockups, images with captions

text · photo8 shapesUp to 10 referencesUp to 4 images3 quality levels

≈ 54 cr at defaults · from 4 cr

Video

Seedance 2.5

Cuesora pickNew

ByteDance's flagship: motion that stays coherent over a long take, behind a strict content filter.

Use it for: long takes, complex scenes, most video work

text · photo · frames · references4–30 s480p · 720p · 1080p6 shapesSound on / offUp to 29 references

≈ 946 cr at defaults · from 353 cr

Seedance Lite

The family's junior tier — cheaper.

Seedance's draft tier — text in, a clip out, cheap enough to run again and again.

Use it for: storyboards, previews, cheap iterations

text2–12 s480p · 720p · 1080p7 shapes

≈ 78 cr at defaults · from 14 cr

Seedance Pro

The family's senior tier — more expensive.

Seedance's first generation at full quality — a polished clip straight from text.

Use it for: polished short clips, advertising, social video

text2–12 s480p · 720p · 1080p6 shapes

≈ 243 cr at defaults · from 20 cr

Kling 3 Standard

The family's base tier.

Kling 3 with sound of its own, a closing frame, and no objection to real faces.

Use it for: clips with people, dialogue scenes, animating a portrait

text · photo · frames3–15 s3 shapesSound on / offMotion control

≈ 252 cr at defaults · from 101 cr

Kling o3 Standard

NewThe family's base tier.

The Kling that keeps a character recognisable from shot to shot, sound included.

Use it for: characters across shots, scenes built from references, clips with sound

text · photo · frames · references3–15 s3 shapesSound on / offUp to 4 references

≈ 168 cr at defaults · from 101 cr

Kling o3 Pro

NewThe family's senior tier — more expensive.

Kling o3's senior tier — steadier motion and finer detail in every mode it has.

Use it for: hero shots with recurring characters, advertising, finished pieces

text · photo · frames · references3–15 s3 shapesSound on / offUp to 4 references

≈ 224 cr at defaults · from 135 cr

Seedance 2.0

New

The previous Seedance flagship — cinematic motion with sound, photos addressed as @Image.

Use it for: cinematic scenes, compositions from several photos, high-resolution deliverables

text · photo · frames · references4–15 s480p · 720p · 1080p · 2160p6 shapesSound on / offUp to 9 references

≈ 605 cr at defaults · from 218 cr

Seedance 2.0 Fast

NewFast tier — cheaper than the standard one.

Seedance's quick branch — the same read of a prompt, a lighter picture, made for drafting.

Use it for: drafts of Seedance scenes, volume work, checking a shot before the full take

text · photo · frames · references4–15 s480p · 720p6 shapesSound on / offUp to 9 references

≈ 484 cr at defaults · from 175 cr

Veo 3.1

New

Google Veo films rather than animates, and it speaks: dialogue, ambience, a sense of lens.

Use it for: dialogue scenes, realistic footage, first-to-last-frame shots

text · photo · frames · references4–8 s720p · 1080p · 2160p2 shapesSound on / offUp to 3 references

≈ 1280 cr at defaults · from 320 cr

Veo 3.1 Lite

NewThe family's junior tier — cheaper.

Veo's cheap tier — the same sound and camera sense, starting from your photo.

Use it for: drafts, social clips, trying an idea before the full Veo

photo · frames4–8 s720p · 1080p2 shapesSound on / off

≈ 160 cr at defaults · from 48 cr

Wan 3.0

New

Alibaba Wan — long, patient takes and a strip of photos to build a scene from.

Use it for: longer scenes, sequences from several references, unhurried shots

text · photo · frames · references2–30 s480p · 720p · 1080p6 shapesUp to 10 references

≈ 400 cr at defaults · from 40 cr

Wan 3.0 Prime

New

Wan's senior tier — the same modes with cleaner motion and more detail held.

Use it for: finished pieces, work where the extra sharpness shows, client deliverables

text · photo · frames · references2–30 s480p · 720p · 1080p6 shapesUp to 10 references

≈ 560 cr at defaults · from 55 cr

Grok Imagine

New

xAI Grok Imagine — quick and literal; it rewards a short, concrete prompt.

Use it for: quick ideas, a single clear action, social posts

text · photo · references1–15 s480p · 720p · 1080p7 shapesUp to 7 references

≈ 336 cr at defaults · from 32 cr

MiniMax H3

New

MiniMax H3 holds a face steady through a clip, and renders at high resolution.

Use it for: recurring characters, clips that need resolution, faces in motion

text · photo · frames · references5–15 s480p · 768p · 1440p · 2160p7 shapesUp to 9 references

≈ 260 cr at defaults · from 100 cr

MiniMax H3 Max

New

The larger MiniMax H3 — steadier motion and a firmer grip on a character.

Use it for: clips that must hold together, recurring characters, longer actions

text · photo · frames5–15 s480p · 768p6 shapes

≈ 160 cr at defaults · from 100 cr

MiniMax H3 Max Turbo

New

The quickest way to see an idea move at all — rough, and made for trying again.

Use it for: first drafts, checking a prompt, throwaway attempts

text · photo · frames5–15 s480p · 768p6 shapes

≈ 80 cr at defaults · from 50 cr

Speech

ElevenLabs v3

New

ElevenLabs — what we reach for when the reading must not sound synthetic.

Use it for: narration, advertising, character voices

20 voicesStability

40 cr per 1,000 characters

MiniMax 2.8 HD

New

MiniMax — strong Russian, with presets for emotion and a hand on the pace.

Use it for: voice-over in Russian, expressive reads, social clips

17 voicesEmotionsSpeed

40 cr per 1,000 characters

Start creating