
7ARTAI Models on 7ART
73 engines for video, image, music, voice, avatars and 3D — in one workflow, on one credit balance. We add new top-tier models as soon as we can get API access.
Every duration, output tier and price below is what 7ART offers and charges for that model — read from the studio’s own tables, not from a vendor spec sheet.
In the catalogue today
22 video · 23 image · 6 music · 11 voice & audio · 5 avatar & lipsync · 6 3d
01 / Video
Video models
They disagree about almost everything — clip length, resolution ceiling, whether they will take a reference video at all — so what each card carries is the difference, not the pitch.
Video7ART preset7Video
- Duration
- 5–15s
- Quality
- 768P / 2K
- Inputs
- Start + end frame, or image / video / audio references
Credits
10% above the base engine rate
View model
VideoGemini Omni Flash
- Duration
- 4, 6, 8 or 10s
- Output tiers
- 720p / 1080p / 4K
- Aspect
- 16:9 or 9:16
Credits
90 / 120 / 150 / 180 for 4 / 6 / 8 / 10s at 720p or 1080p · 210 / 240 / 270 / 300 at 4K · flat 240 (720p/1080p) or 360 (4K) with a video reference
View model
VideoVEO
- Duration
- 4, 6 or 8s
- Quality
- Preview / 1080p / 4K
- Inputs
- Start + end frame
- 3.1 Quality
- 3.1 Fast
- 3.1 Lite
Credits · 3.1 Quality
Flat per generation: 250 preview · 255 at 1080p · 370 at 4K
View model
VideoSeedance
- Duration
- 4–30s
- Quality
- 480p / 720p
- References
- 30 images · 10 videos · 10 audio
- 2.5
- 2.0
- 2.0 Fast
- 2.0 Mini
Credits · 2.5
37/sec at 480p · 79/sec at 720p
View model
VideoKling
- Duration
- 3–15s
- Quality
- 720p / 1080p / 4K
- Inputs
- Start frame; end frame outside multi-shot
- 3.0
- 3.0 Turbo
- 2.6
- 2.5 Turbo
- +1 more
Credits · 3.0
14/sec at 720p · 18/sec at 1080p · 20 / 27 per sec with audio · 67/sec at 4K
View model
VideoGrok Imagine (video)
- Duration
- 6–30s
- Quality
- 480p / 720p
- Inputs
- Start frame
Credits
1.6/sec at 480p · 3/sec at 720p
View model
VideoMiniMax H3
- Duration
- 5–15s
- Quality
- 768P / 2K
- Inputs
- Start + end frame, or 9 images / 3 videos / 3 audio
Credits
22.5/sec at 768P · 36.5/sec at 2K, plus any reference video's own seconds and 11 credits per reference image past the first 5
View model
VideoHappyHorse 1.1
- Duration
- 3–15s
- Quality
- 720p / 1080p
- Inputs
- Start frame, or reference images
Credits
33/sec at 720p · 44/sec at 1080p
View model
VideoWan
- Duration
- 5, 10 or 15s
- Quality
- 720p / 1080p
- Inputs
- Start frame, end frame, reference video, driving audio
- 2.7
- 2.6
- 2.5
Credits · 2.7
16/sec at 720p · 24/sec at 1080p
View model
VideoMMAudio v2
- Output
- Audio muxed onto the source clip
Credits
Flat 5 per run
View model
VideoLuma Ray-2 Modify
- Output
- The whole source clip, restyled
Credits
Flat 600 per run
View modelAll 22 video engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| 7Video7ART | 7ART's flagship video preset — our best video quality, tuned and priced as one option in the same picker. |
| 10% above the base engine rate |
| Gemini Omni FlashGoogle | Multi-modal Google model driven entirely from its own reference panel — images, a video clip, saved Characters and Voices, no frame cards.Listed in the 7ART picker as "Gemini Omni Video" — that name comes from our provider's endpoint, not from Google. |
| 90 / 120 / 150 / 180 for 4 / 6 / 8 / 10s at 720p or 1080p · 210 / 240 / 270 / 300 at 4K · flat 240 (720p/1080p) or 360 (4K) with a video reference |
| VEO 3.1 QualityGoogle | Google's best-looking cinematic tier, priced flat per generation rather than per second. |
| Flat per generation: 250 preview · 255 at 1080p · 370 at 4K |
| VEO 3.1 FastGoogle | Faster VEO, and the only tier that accepts reference photos. |
| Flat per generation: 60 preview · 65 at 1080p · 180 at 4K |
| VEO 3.1 LiteGoogle | The affordable VEO — good for experimenting. |
| Flat per generation: 30 preview · 35 at 1080p · 150 at 4K |
| Seedance 2.5ByteDance | The longest clips in the studio, audio generated with the video, and by far the widest reference set — but 480p/720p only, a lower ceiling than Seedance 2.0. |
| 37/sec at 480p · 79/sec at 720p |
| Seedance 2.0ByteDance | Real people with strong identity consistency and cinematic multi-shot sequences; the only Seedance tier that renders 1080p. |
| 19 / 41 / 102 per sec at 480p / 720p / 1080p · with a video reference 11.5 / 25 / 62 per sec over input + output seconds |
| Seedance 2.0 FastByteDance | Lighter Seedance 2.0 for drafts and previews. |
| 15.5/sec at 480p · 33/sec at 720p · with a video reference 9 / 20 per sec over input + output seconds |
| Seedance 2.0 MiniByteDance | The cheapest Seedance tier. |
| 9.5/sec at 480p · 20.5/sec at 720p · with a video reference 6 / 12.5 per sec over input + output seconds |
| Kling 3.0Kling AI | The most capable Kling — characters, synced sound, multi-shot, and a 4K tier. |
| 14/sec at 720p · 18/sec at 1080p · 20 / 27 per sec with audio · 67/sec at 4K |
| Kling 3.0 TurboKling AI | Faster Kling 3.0, text and image-to-video. |
| 18/sec at 720p · 22.5/sec at 1080p |
| Kling 2.6Kling AI | Reliable and strong on animation, with sound. |
| 11/sec · 22/sec with audio |
| Kling 2.5 TurboKling AI | The fastest, cheapest draft engine. |
| 9/sec |
| Kling 3.0 Motion ControlKling AI | Motion transfer — take the movement from one clip and drive a new subject with it.Runs in the Motion studio, not the video model picker. |
| 20/sec at 720p · 27/sec at 1080p |
| Grok Imagine (video)xAI | Strong motion, text and image-to-video, and the cheapest per second in the studio. |
| 1.6/sec at 480p · 3/sec at 720p |
| MiniMax H3MiniMax | Two quality tiers — 768P and native 2K — first-to-last frame, and image, video and audio references.References are only sent when there is no start frame — the panel hides one while the other is in use. |
| 22.5/sec at 768P · 36.5/sec at 2K, plus any reference video's own seconds and 11 credits per reference image past the first 5 |
| HappyHorse 1.1Alibaba | Text, image or multi-reference to video — images only, no video or audio references. |
| 33/sec at 720p · 44/sec at 1080p |
| Wan 2.7 | Four modes in one — text-to-video, image-to-video, video edit and reference-to-video — with driving audio, prompt expansion and negative prompts. |
| 16/sec at 720p · 24/sec at 1080p |
| Wan 2.6 | Transform an existing video; multi-scene. |
| 14/sec at 720p · 21/sec at 1080p |
| Wan 2.5 | Expands short prompts into richer descriptions for you. |
| 12/sec at 720p · 20/sec at 1080p |
| MMAudio v2 | Generates a foley track for a clip that has none.Runs from a clip's edit actions, not the model picker. |
| Flat 5 per run |
| Luma Ray-2 ModifyLuma | Restyles an existing clip end to end.Runs from a clip's edit actions, not the model picker. |
| Flat 600 per run |
02 / Image
Image models
Reference-image caps here are the number the primary provider actually receives, which is not always the number the upload box will let you drop. FLUX renders around 1 MP on every tier whatever its resolution label says.
Image7ART preset7Image
- Resolutions
- 1K / 2K / 4K
- References
- Up to 16
Credits
10% above the base engine rate
View model
ImageNano Banana
- Resolutions
- 1K / 2K / 4K
- References
- Up to 14
- 2
- Pro
Credits · 2
16 / 24 / 32 per image at 1K / 2K / 4K
View model
ImageSeedream
- Resolutions
- 2K / 3K
- References
- Up to 10
- 5 Lite
- 4.5
- 4.0
Credits · 5 Lite
6 at 2K · 8 at 3K
View model
ImageGPT Image 2
- Resolutions
- 1K / 2K / 4K
- References
- Up to 16
Credits
8 / 11 / 20 per image at 1K / 2K / 4K
View model
ImageZ-Image
- Output
- 2K, single fixed quality
- References
- None — text-to-image only
Credits
2 per image
View model
ImageGrok Imagine (image)
- Output
- 2K, single fixed quality
- References
- 1
- Batch
- 6 images in speed mode · 4 in quality
Credits
4 per generation in speed mode · 5 in quality
View model
ImageFLUX
- Output
- ~1 MP
- FLUX.2 Pro
- FLUX.2
- FLUX.2 Turbo
- FLUX.2 Flash
- +8 more
Credits · FLUX.2 Pro
6 per image
View model
ImageClarity Upscaler
- Input
- Exactly 1 image
Credits
24 per run
View model
ImageIC-Light v2
- Input
- Exactly 1 image
Credits
20 per run
View modelAll 23 image engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| 7Image7ART | 7ART's flagship image preset — our best image quality, pinned to the top of the same picker. |
| 10% above the base engine rate |
| Nano Banana 2Google | Up to 14 reference images — the safest pick for keeping a face consistent. |
| 16 / 24 / 32 per image at 1K / 2K / 4K |
| Nano Banana ProGoogle | Multi-image composition. |
| 30 / 30 / 60 per image at 1K / 2K / 4K |
| Seedream 5 LiteByteDance | Cheap and quick. |
| 6 at 2K · 8 at 3K |
| Seedream 4.5ByteDance | The mid Seedream tier. |
| 7 at 2K · 9 at 4K |
| Seedream 4.0ByteDance | The widest resolution range of the Seedreams. |
| 5 / 6 / 8 per image at 1K / 2K / 4K |
| GPT Image 2OpenAI | Up to 16 references, strong at following instructions and at text inside the image.A 1:1 render cannot be 4K, and an unspecified size caps at 1K. |
| 8 / 11 / 20 per image at 1K / 2K / 4K |
| Z-ImageAlibaba | Cheap text-to-image plates — it takes no reference images at all. |
| 2 per image |
| Grok Imagine (image)xAI | Text and image-to-image; one flat charge returns a whole batch. |
| 4 per generation in speed mode · 5 in quality |
| FLUX.2 ProBlack Forest Labs | Top realism and typography. |
| 6 per image |
| FLUX.2Black Forest Labs | Realism; understands JSON and hex-colour prompts. |
| 3 per image |
| FLUX.2 TurboBlack Forest Labs | Fast and cheap. |
| 2 per image |
| FLUX.2 FlashBlack Forest Labs | The fastest FLUX.2. |
| 1 per image |
| FLUX.2 KleinBlack Forest Labs | Few-step 9B model. |
| 2 per image |
| FLUX.1 DevBlack Forest Labs | The classic FLUX quality. |
| 5 per image |
| FLUX.1 SchnellBlack Forest Labs | Sub-second drafts. |
| 1 per image |
| FLUX.2 EditBlack Forest Labs | Instruct-edit — edit or compose up to 4 images.Appears in the picker only once a reference image is attached. |
| 5 per image |
| FLUX.2 Pro EditBlack Forest Labs | Premium instruct-edit.Appears in the picker only once a reference image is attached. |
| 9 per image |
| FLUX.2 Klein EditBlack Forest Labs | Fast, light edits.Appears in the picker only once a reference image is attached. |
| 5 per image |
| Trained ElementBlack Forest Labs | FLUX.1 Dev plus a LoRA trained on your own Element — reproduces that identity exactly.Appears in the picker only while a trained Element is @mentioned in the prompt. |
| 10 per image, after a one-off 400 to train |
| Clarity Upscaler | Upscales a single image, 2× by default.Reachable only through the Super Agent — there is no button for it in the Image studio. |
| 24 per run |
| IC-Light v2 | Relights a single image.Reachable only through the Super Agent — there is no button for it in the Image studio. |
| 20 per run |
| FLUX.1 KontextBlack Forest Labs | Instruct-edit on a single source image.Reachable only through the Super Agent — there is no button for it in the Image studio. |
| 8 per run |
03 / Music
Music models
One vendor, six versions, one flat price — because the interesting part of Suno on 7ART is everything that happens after the track exists.
All 6 music engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| Suno V5.5Suno | The latest Suno version in the 7ART music studio. |
| 12 per generation |
| Suno V5Suno | The version the music route generates with when none is chosen. |
| 12 per generation |
| Suno V4.5 AllSuno | Earlier Suno version, kept selectable. | — | 12 per generation |
| Suno V4.5+Suno | Earlier Suno version, kept selectable. | — | 12 per generation |
| Suno V4.5Suno | Earlier Suno version, kept selectable. | — | 12 per generation |
| Suno V4Suno | The oldest Suno version still selectable in the studio. | — | 12 per generation |
04 / Voice & audio
Voice and audio models
Speech, dialogue, effects and score. The ElevenLabs engines run through fal with a kie fallback; Sonilo has a single provider and no second leg.
AudioElevenLabs
- Speed
- 0.7×–1.2×
- Turbo 2.5
- Multilingual v2
- Eleven v3
- v3 Dialogue
- +4 more
Credits · Turbo 2.5
10 per 1,000 characters
View model
AudioHeyGen
- Video Translate — Precision
- Video Translate — Speed
Credits · Video Translate — Precision
20 per second of output
View model
AudioSonilo v1.1
- Modes
- Video → sound effects · Video → music · Text → sound · Text → music
- Sections
- ≥5s each, contiguous, ≤200 characters
- Billing floor
- 15-second minimum
Credits
1.8/sec for either video mode · 0.36/sec text-to-sound · 0.5/sec text-to-music, over a 15-second minimum
View modelAll 11 voice & audio engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| ElevenLabs Turbo 2.5ElevenLabs | The fast text-to-speech tier. |
| 10 per 1,000 characters |
| ElevenLabs Multilingual v2ElevenLabs | The multilingual text-to-speech tier. |
| 20 per 1,000 characters |
| ElevenLabs Eleven v3ElevenLabs | The expressive text-to-speech tier. |
| 20 per 1,000 characters |
| ElevenLabs v3 DialogueElevenLabs | Multi-speaker dialogue from a script. | — | 20 per 1,000 characters |
| ElevenLabs Sound Effects v2ElevenLabs | A sound effect from a text description. | — | 0.4 per second |
| ElevenLabs Audio IsolationElevenLabs | Strips background noise off a recording. | — | 0.34 per second |
| ElevenLabs Speech-to-TextElevenLabs | Transcribes speech from an audio file. | — | 6 per minute |
| ElevenLabs DubbingElevenLabs | Dubs a video into another language. | — | 180 per minute, rounded up |
| HeyGen Video Translate — PrecisionHeyGen | Translates a video with a cloned voice and a re-synced mouth. | — | 20 per second of output |
| HeyGen Video Translate — SpeedHeyGen | The fast video-translation mode. | — | 10 per second of output |
| Sonilo v1.1Sonilo | Sound design that lands on the action — foley or a music score synced to a video, or either one from a text description. |
| 1.8/sec for either video mode · 0.36/sec text-to-sound · 0.5/sec text-to-music, over a 15-second minimum |
05 / Avatar & lipsync
Avatar and lipsync models
Four engines from four vendors for making a face talk, from a single portrait or from footage you already have.
AvatarKling Avatar
- Quality
- 720p
- Max length
- 60s
- Standard
- Pro
Credits · Standard
8 per second
View modelHeyGen Avatar IV
- Quality
- 720p / 1080p
- Max length
- 60s
Credits
20 per second
View model
AvatarOmnihuman 1.5
- Quality
- 720p / 1080p, billed the same
- Max length
- 60s
Credits
27 per second
View model
AvatarVolcengine Lip Sync
- Input
- A source video
- Max length
- 60s
Credits
8 per second
View modelAll 5 avatar & lipsync engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| Kling Avatar — StandardKling AI | Turns a portrait into a talking avatar. |
| 8 per second |
| Kling Avatar — ProKling AI | The 1080p talking-avatar tier, and the longest clip in the studio. |
| 16 per second |
| HeyGen Avatar IVHeyGen | Photo to talking avatar. |
| 20 per second |
| Omnihuman 1.5ByteDance | Portrait to talking video. |
| 27 per second |
| Volcengine Lip SyncVolcengine | Re-lipsyncs an existing video to new audio. |
| 8 per second |
06 / 3D
3D models
Four generators and two mesh tools, all through fal. Polycount figures are the studio slider’s range, which is deliberately narrower than what the endpoints accept.
3DMeshy 6
- Modes
- Image · Multi-view · Text
- Polycount
- 10,000–300,000
- Extras
- PBR · quad topology · A-pose / T-pose
Credits
240 per generation
View model
3DHunyuan 3.1 Pro
- Modes
- Image · Multi-view · Text
- Polycount
- 40,000–1,500,000
- Extras
- PBR
Credits
80 per generation
View model
3DTrellis 2
- Modes
- Image only
- Polycount
- 10,000–300,000
- Images
- 1
Credits
60 per generation
View model
3DTripo 2.5
- Modes
- Image · Multi-view · Text
- Polycount
- 10,000–300,000
- Extras
- PBR
Credits
60 per generation
View model
3DRetexture Engine
- Input
- An existing mesh + 1 style image
- Extras
- PBR
Credits
60 per run
View model
3DRemesh
- Input
- An existing mesh
- Polycount
- 10,000–300,000
- Formats
- GLB · FBX · OBJ · STL · USDZ
Credits
40 per run
View modelAll 6 3d engines, with every spec and price
| Model | What it is for | Specs | Credits |
|---|---|---|---|
| Meshy 6Meshy | The best all-round mesh — quad topology, PBR materials and posing. |
| 240 per generation |
| Hunyuan 3.1 Pro | The highest-detail mesh, with multi-view input. |
| 80 per generation |
| Trellis 2 | Fast and affordable, from a single image. |
| 60 per generation |
| Tripo 2.5Tripo3D | The fastest mesh in the 3D studio. |
| 60 per generation |
| Retexture EngineMeshy | Re-textures a mesh you already have, from a prompt or a style image. |
| 60 per run |
| RemeshMeshy | Remesh, rescale and re-export a model for printing or another pipeline. |
| 40 per run |
How a model gets onto this page
A model appears here once we can actually run it — API access granted, the adapter written, the price set. Access is the vendor’s to grant, not ours to promise, so you will not find a launch date on this page for anything we cannot generate with today.
What you do get is one place to run all of it. The same prompt bar, the same reference uploads, the same library, and one credit balance every model draws from — so moving from an image to a video to a score costs you a dropdown, not a new subscription.
Where two providers serve the same model we run one as primary and the other as a fallback, and the price you see is the same either way.
Every model above, on one balance
No per-vendor subscriptions, no separate accounts, no re-learning a UI each time a new model lands.
