7ART

AI Models on 7ART

73 engines for video, image, music, voice, avatars and 3D — in one workflow, on one credit balance. We add new top-tier models as soon as we can get API access.

Every duration, output tier and price below is what 7ART offers and charges for that model — read from the studio’s own tables, not from a vendor spec sheet.

In the catalogue today

22 video · 23 image · 6 music · 11 voice & audio · 5 avatar & lipsync · 6 3d

01 / Video

Video models

They disagree about almost everything — clip length, resolution ceiling, whether they will take a reference video at all — so what each card carries is the difference, not the pitch.

Open the video studio
7Video on 7ARTVideo7ART preset
7ART

7Video

7ART's flagship video preset — our best video quality, tuned and priced as one option in the same picker.

Duration
5–15s
Quality
768P / 2K
Inputs
Start + end frame, or image / video / audio references

Credits

10% above the base engine rate

View model
Gemini Omni Flash on 7ARTVideo
Google

Gemini Omni Flash

Multi-modal Google model driven entirely from its own reference panel — images, a video clip, saved Characters and Voices, no frame cards.

Duration
4, 6, 8 or 10s
Output tiers
720p / 1080p / 4K
Aspect
16:9 or 9:16

Credits

90 / 120 / 150 / 180 for 4 / 6 / 8 / 10s at 720p or 1080p · 210 / 240 / 270 / 300 at 4K · flat 240 (720p/1080p) or 360 (4K) with a video reference

View model
VEO on 7ARTVideo
Google3 models

VEO

3.1 Quality: Google's best-looking cinematic tier, priced flat per generation rather than per second.

Duration
4, 6 or 8s
Quality
Preview / 1080p / 4K
Inputs
Start + end frame
  • 3.1 Quality
  • 3.1 Fast
  • 3.1 Lite

Credits · 3.1 Quality

Flat per generation: 250 preview · 255 at 1080p · 370 at 4K

View model
Seedance on 7ARTVideo
4 models

Seedance

2.5: The longest clips in the studio, audio generated with the video, and by far the widest reference set — but 480p/720p only, a lower ceiling than Seedance 2.0.

Duration
4–30s
Quality
480p / 720p
References
30 images · 10 videos · 10 audio
  • 2.5
  • 2.0
  • 2.0 Fast
  • 2.0 Mini

Credits · 2.5

37/sec at 480p · 79/sec at 720p

View model
Kling on 7ARTVideo
Kling AI5 models

Kling

3.0: The most capable Kling — characters, synced sound, multi-shot, and a 4K tier.

Duration
3–15s
Quality
720p / 1080p / 4K
Inputs
Start frame; end frame outside multi-shot
  • 3.0
  • 3.0 Turbo
  • 2.6
  • 2.5 Turbo
  • +1 more

Credits · 3.0

14/sec at 720p · 18/sec at 1080p · 20 / 27 per sec with audio · 67/sec at 4K

View model
Grok Imagine (video) on 7ARTVideo
xAI

Grok Imagine (video)

Strong motion, text and image-to-video, and the cheapest per second in the studio.

Duration
6–30s
Quality
480p / 720p
Inputs
Start frame

Credits

1.6/sec at 480p · 3/sec at 720p

View model
MiniMax H3 on 7ARTVideo
MiniMax

MiniMax H3

Two quality tiers — 768P and native 2K — first-to-last frame, and image, video and audio references.

Duration
5–15s
Quality
768P / 2K
Inputs
Start + end frame, or 9 images / 3 videos / 3 audio

Credits

22.5/sec at 768P · 36.5/sec at 2K, plus any reference video's own seconds and 11 credits per reference image past the first 5

View model
HappyHorse 1.1 on 7ARTVideo
Alibaba

HappyHorse 1.1

Text, image or multi-reference to video — images only, no video or audio references.

Duration
3–15s
Quality
720p / 1080p
Inputs
Start frame, or reference images

Credits

33/sec at 720p · 44/sec at 1080p

View model
Wan on 7ARTVideo
3 models

Wan

2.7: Four modes in one — text-to-video, image-to-video, video edit and reference-to-video — with driving audio, prompt expansion and negative prompts.

Duration
5, 10 or 15s
Quality
720p / 1080p
Inputs
Start frame, end frame, reference video, driving audio
  • 2.7
  • 2.6
  • 2.5

Credits · 2.7

16/sec at 720p · 24/sec at 1080p

View model
MMAudio v2 on 7ARTVideo

MMAudio v2

Generates a foley track for a clip that has none.

Output
Audio muxed onto the source clip

Credits

Flat 5 per run

View model
Luma Ray-2 Modify on 7ARTVideo
Luma

Luma Ray-2 Modify

Restyles an existing clip end to end.

Output
The whole source clip, restyled

Credits

Flat 600 per run

View model
All 22 video engines, with every spec and price
ModelWhat it is forSpecsCredits
7Video7ART7ART's flagship video preset — our best video quality, tuned and priced as one option in the same picker.
  • Duration: 5–15s
  • Quality: 768P / 2K
  • Inputs: Start + end frame, or image / video / audio references
10% above the base engine rate
Gemini Omni FlashGoogleMulti-modal Google model driven entirely from its own reference panel — images, a video clip, saved Characters and Voices, no frame cards.Listed in the 7ART picker as "Gemini Omni Video" — that name comes from our provider's endpoint, not from Google.
  • Duration: 4, 6, 8 or 10s
  • Output tiers: 720p / 1080p / 4K
  • Aspect: 16:9 or 9:16
  • References: 7 images · 1 video clip · 3 Characters · 3 Voices
90 / 120 / 150 / 180 for 4 / 6 / 8 / 10s at 720p or 1080p · 210 / 240 / 270 / 300 at 4K · flat 240 (720p/1080p) or 360 (4K) with a video reference
VEO 3.1 QualityGoogleGoogle's best-looking cinematic tier, priced flat per generation rather than per second.
  • Duration: 4, 6 or 8s
  • Quality: Preview / 1080p / 4K
  • Inputs: Start + end frame
Flat per generation: 250 preview · 255 at 1080p · 370 at 4K
VEO 3.1 FastGoogleFaster VEO, and the only tier that accepts reference photos.
  • Duration: 4, 6 or 8s
  • Quality: Preview / 1080p / 4K
  • Inputs: Start + end frames, or up to 3 reference photos
Flat per generation: 60 preview · 65 at 1080p · 180 at 4K
VEO 3.1 LiteGoogleThe affordable VEO — good for experimenting.
  • Duration: 4, 6 or 8s
  • Quality: Preview / 1080p / 4K
  • Inputs: Start + end frame
Flat per generation: 30 preview · 35 at 1080p · 150 at 4K
Seedance 2.5ByteDanceThe longest clips in the studio, audio generated with the video, and by far the widest reference set — but 480p/720p only, a lower ceiling than Seedance 2.0.
  • Duration: 4–30s
  • Quality: 480p / 720p
  • References: 30 images · 10 videos · 10 audio
  • Aspect: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 or auto
37/sec at 480p · 79/sec at 720p
Seedance 2.0ByteDanceReal people with strong identity consistency and cinematic multi-shot sequences; the only Seedance tier that renders 1080p.
  • Duration: 4–15s
  • Quality: 480p / 720p / 1080p
  • References: 9 images · 3 videos · 3 audio
19 / 41 / 102 per sec at 480p / 720p / 1080p · with a video reference 11.5 / 25 / 62 per sec over input + output seconds
Seedance 2.0 FastByteDanceLighter Seedance 2.0 for drafts and previews.
  • Duration: 4–15s
  • Quality: 480p / 720p
  • References: 9 images · 3 videos · 3 audio
15.5/sec at 480p · 33/sec at 720p · with a video reference 9 / 20 per sec over input + output seconds
Seedance 2.0 MiniByteDanceThe cheapest Seedance tier.
  • Duration: 4–15s
  • Quality: 480p / 720p
  • References: 9 images · 3 videos · 3 audio
9.5/sec at 480p · 20.5/sec at 720p · with a video reference 6 / 12.5 per sec over input + output seconds
Kling 3.0Kling AIThe most capable Kling — characters, synced sound, multi-shot, and a 4K tier.
  • Duration: 3–15s
  • Quality: 720p / 1080p / 4K
  • Inputs: Start frame; end frame outside multi-shot
  • Audio: Optional — raises the per-second rate
14/sec at 720p · 18/sec at 1080p · 20 / 27 per sec with audio · 67/sec at 4K
Kling 3.0 TurboKling AIFaster Kling 3.0, text and image-to-video.
  • Duration: 3–15s
  • Quality: 720p / 1080p
  • Inputs: Start frame
18/sec at 720p · 22.5/sec at 1080p
Kling 2.6Kling AIReliable and strong on animation, with sound.
  • Duration: 5 or 10s
  • Quality: 720p / 1080p
  • Inputs: Start frame
11/sec · 22/sec with audio
Kling 2.5 TurboKling AIThe fastest, cheapest draft engine.
  • Duration: 5 or 10s
  • Quality: 720p
  • Inputs: Start frame
9/sec
Kling 3.0 Motion ControlKling AIMotion transfer — take the movement from one clip and drive a new subject with it.Runs in the Motion studio, not the video model picker.
  • Quality: 720p / 1080p
  • Inputs: Start frame
20/sec at 720p · 27/sec at 1080p
Grok Imagine (video)xAIStrong motion, text and image-to-video, and the cheapest per second in the studio.
  • Duration: 6–30s
  • Quality: 480p / 720p
  • Inputs: Start frame
1.6/sec at 480p · 3/sec at 720p
MiniMax H3MiniMaxTwo quality tiers — 768P and native 2K — first-to-last frame, and image, video and audio references.References are only sent when there is no start frame — the panel hides one while the other is in use.
  • Duration: 5–15s
  • Quality: 768P / 2K
  • Inputs: Start + end frame, or 9 images / 3 videos / 3 audio
  • Aspect: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
22.5/sec at 768P · 36.5/sec at 2K, plus any reference video's own seconds and 11 credits per reference image past the first 5
HappyHorse 1.1AlibabaText, image or multi-reference to video — images only, no video or audio references.
  • Duration: 3–15s
  • Quality: 720p / 1080p
  • Inputs: Start frame, or reference images
33/sec at 720p · 44/sec at 1080p
Wan 2.7Four modes in one — text-to-video, image-to-video, video edit and reference-to-video — with driving audio, prompt expansion and negative prompts.
  • Duration: 5, 10 or 15s
  • Quality: 720p / 1080p
  • Inputs: Start frame, end frame, reference video, driving audio
16/sec at 720p · 24/sec at 1080p
Wan 2.6Transform an existing video; multi-scene.
  • Duration: 5, 10 or 15s
  • Quality: 720p / 1080p
  • Inputs: Start frame, reference video
14/sec at 720p · 21/sec at 1080p
Wan 2.5Expands short prompts into richer descriptions for you.
  • Duration: 5 or 10s
  • Quality: 720p / 1080p
  • Inputs: Start frame
12/sec at 720p · 20/sec at 1080p
MMAudio v2Generates a foley track for a clip that has none.Runs from a clip's edit actions, not the model picker.
  • Output: Audio muxed onto the source clip
Flat 5 per run
Luma Ray-2 ModifyLumaRestyles an existing clip end to end.Runs from a clip's edit actions, not the model picker.
  • Output: The whole source clip, restyled
Flat 600 per run

02 / Image

Image models

Reference-image caps here are the number the primary provider actually receives, which is not always the number the upload box will let you drop. FLUX renders around 1 MP on every tier whatever its resolution label says.

Open the image studio
7Image on 7ARTImage7ART preset
7ART

7Image

7ART's flagship image preset — our best image quality, pinned to the top of the same picker.

Resolutions
1K / 2K / 4K
References
Up to 16

Credits

10% above the base engine rate

View model
Nano Banana on 7ARTImage
Google2 models

Nano Banana

2: Up to 14 reference images — the safest pick for keeping a face consistent.

Resolutions
1K / 2K / 4K
References
Up to 14
  • 2
  • Pro

Credits · 2

16 / 24 / 32 per image at 1K / 2K / 4K

View model
Seedream on 7ARTImage
ByteDance3 models

Seedream

5 Lite: Cheap and quick.

Resolutions
2K / 3K
References
Up to 10
  • 5 Lite
  • 4.5
  • 4.0

Credits · 5 Lite

6 at 2K · 8 at 3K

View model
GPT Image 2 on 7ARTImage
OpenAI

GPT Image 2

Up to 16 references, strong at following instructions and at text inside the image.

Resolutions
1K / 2K / 4K
References
Up to 16

Credits

8 / 11 / 20 per image at 1K / 2K / 4K

View model
Z-Image on 7ARTImage
Alibaba

Z-Image

Cheap text-to-image plates — it takes no reference images at all.

Output
2K, single fixed quality
References
None — text-to-image only

Credits

2 per image

View model
Grok Imagine (image) on 7ARTImage
xAI

Grok Imagine (image)

Text and image-to-image; one flat charge returns a whole batch.

Output
2K, single fixed quality
References
1
Batch
6 images in speed mode · 4 in quality

Credits

4 per generation in speed mode · 5 in quality

View model
FLUX on 7ARTImage
Black Forest Labs12 models

FLUX

FLUX.2 Pro: Top realism and typography.

Output
~1 MP
  • FLUX.2 Pro
  • FLUX.2
  • FLUX.2 Turbo
  • FLUX.2 Flash
  • +8 more

Credits · FLUX.2 Pro

6 per image

View model
Clarity Upscaler on 7ARTImage

Clarity Upscaler

Upscales a single image, 2× by default.

Input
Exactly 1 image

Credits

24 per run

View model
IC-Light v2 on 7ARTImage

IC-Light v2

Relights a single image.

Input
Exactly 1 image

Credits

20 per run

View model
All 23 image engines, with every spec and price
ModelWhat it is forSpecsCredits
7Image7ART7ART's flagship image preset — our best image quality, pinned to the top of the same picker.
  • Resolutions: 1K / 2K / 4K
  • References: Up to 16
10% above the base engine rate
Nano Banana 2GoogleUp to 14 reference images — the safest pick for keeping a face consistent.
  • Resolutions: 1K / 2K / 4K
  • References: Up to 14
16 / 24 / 32 per image at 1K / 2K / 4K
Nano Banana ProGoogleMulti-image composition.
  • Resolutions: 1K / 2K / 4K
  • References: Up to 8
30 / 30 / 60 per image at 1K / 2K / 4K
Seedream 5 LiteByteDanceCheap and quick.
  • Resolutions: 2K / 3K
  • References: Up to 10
6 at 2K · 8 at 3K
Seedream 4.5ByteDanceThe mid Seedream tier.
  • Resolutions: 2K / 4K
  • References: Up to 10
7 at 2K · 9 at 4K
Seedream 4.0ByteDanceThe widest resolution range of the Seedreams.
  • Resolutions: 1K / 2K / 4K
  • References: Up to 10
5 / 6 / 8 per image at 1K / 2K / 4K
GPT Image 2OpenAIUp to 16 references, strong at following instructions and at text inside the image.A 1:1 render cannot be 4K, and an unspecified size caps at 1K.
  • Resolutions: 1K / 2K / 4K
  • References: Up to 16
8 / 11 / 20 per image at 1K / 2K / 4K
Z-ImageAlibabaCheap text-to-image plates — it takes no reference images at all.
  • Output: 2K, single fixed quality
  • References: None — text-to-image only
2 per image
Grok Imagine (image)xAIText and image-to-image; one flat charge returns a whole batch.
  • Output: 2K, single fixed quality
  • References: 1
  • Batch: 6 images in speed mode · 4 in quality
4 per generation in speed mode · 5 in quality
FLUX.2 ProBlack Forest LabsTop realism and typography.
  • Output: ~1 MP
6 per image
FLUX.2Black Forest LabsRealism; understands JSON and hex-colour prompts.
  • Output: ~1 MP
3 per image
FLUX.2 TurboBlack Forest LabsFast and cheap.
  • Output: ~1 MP
2 per image
FLUX.2 FlashBlack Forest LabsThe fastest FLUX.2.
  • Output: ~1 MP
1 per image
FLUX.2 KleinBlack Forest LabsFew-step 9B model.
  • Output: ~1 MP
2 per image
FLUX.1 DevBlack Forest LabsThe classic FLUX quality.
  • Output: ~1 MP
5 per image
FLUX.1 SchnellBlack Forest LabsSub-second drafts.
  • Output: ~1 MP
1 per image
FLUX.2 EditBlack Forest LabsInstruct-edit — edit or compose up to 4 images.Appears in the picker only once a reference image is attached.
  • Output: ~1 MP
  • References: Up to 4
5 per image
FLUX.2 Pro EditBlack Forest LabsPremium instruct-edit.Appears in the picker only once a reference image is attached.
  • Output: ~1 MP
  • References: Up to 9
9 per image
FLUX.2 Klein EditBlack Forest LabsFast, light edits.Appears in the picker only once a reference image is attached.
  • Output: ~1 MP
  • References: Up to 4
5 per image
Trained ElementBlack Forest LabsFLUX.1 Dev plus a LoRA trained on your own Element — reproduces that identity exactly.Appears in the picker only while a trained Element is @mentioned in the prompt.
  • Output: ~1 MP
  • References: 1
  • Training: 400 credits, once per Element
10 per image, after a one-off 400 to train
Clarity UpscalerUpscales a single image, 2× by default.Reachable only through the Super Agent — there is no button for it in the Image studio.
  • Input: Exactly 1 image
24 per run
IC-Light v2Relights a single image.Reachable only through the Super Agent — there is no button for it in the Image studio.
  • Input: Exactly 1 image
20 per run
FLUX.1 KontextBlack Forest LabsInstruct-edit on a single source image.Reachable only through the Super Agent — there is no button for it in the Image studio.
  • Input: Exactly 1 image
8 per run

03 / Music

Music models

One vendor, six versions, one flat price — because the interesting part of Suno on 7ART is everything that happens after the track exists.

Open the music studio
All 6 music engines, with every spec and price
ModelWhat it is forSpecsCredits
Suno V5.5SunoThe latest Suno version in the 7ART music studio.
  • Status: Latest
12 per generation
Suno V5SunoThe version the music route generates with when none is chosen.
  • Status: Default
12 per generation
Suno V4.5 AllSunoEarlier Suno version, kept selectable.12 per generation
Suno V4.5+SunoEarlier Suno version, kept selectable.12 per generation
Suno V4.5SunoEarlier Suno version, kept selectable.12 per generation
Suno V4SunoThe oldest Suno version still selectable in the studio.12 per generation

04 / Voice & audio

Voice and audio models

Speech, dialogue, effects and score. The ElevenLabs engines run through fal with a kie fallback; Sonilo has a single provider and no second leg.

Open the voice studio
All 11 voice & audio engines, with every spec and price
ModelWhat it is forSpecsCredits
ElevenLabs Turbo 2.5ElevenLabsThe fast text-to-speech tier.
  • Speed: 0.7×–1.2×
10 per 1,000 characters
ElevenLabs Multilingual v2ElevenLabsThe multilingual text-to-speech tier.
  • Speed: 0.7×–1.2×
20 per 1,000 characters
ElevenLabs Eleven v3ElevenLabsThe expressive text-to-speech tier.
  • Speed: 0.7×–1.2×
20 per 1,000 characters
ElevenLabs v3 DialogueElevenLabsMulti-speaker dialogue from a script.20 per 1,000 characters
ElevenLabs Sound Effects v2ElevenLabsA sound effect from a text description.0.4 per second
ElevenLabs Audio IsolationElevenLabsStrips background noise off a recording.0.34 per second
ElevenLabs Speech-to-TextElevenLabsTranscribes speech from an audio file.6 per minute
ElevenLabs DubbingElevenLabsDubs a video into another language.180 per minute, rounded up
HeyGen Video Translate — PrecisionHeyGenTranslates a video with a cloned voice and a re-synced mouth.20 per second of output
HeyGen Video Translate — SpeedHeyGenThe fast video-translation mode.10 per second of output
Sonilo v1.1SoniloSound design that lands on the action — foley or a music score synced to a video, or either one from a text description.
  • Modes: Video → sound effects · Video → music · Text → sound · Text → music
  • Sections: ≥5s each, contiguous, ≤200 characters
  • Billing floor: 15-second minimum
1.8/sec for either video mode · 0.36/sec text-to-sound · 0.5/sec text-to-music, over a 15-second minimum

05 / Avatar & lipsync

Avatar and lipsync models

Four engines from four vendors for making a face talk, from a single portrait or from footage you already have.

Open the lipsync studio
All 5 avatar & lipsync engines, with every spec and price
ModelWhat it is forSpecsCredits
Kling Avatar — StandardKling AITurns a portrait into a talking avatar.
  • Quality: 720p
  • Max length: 60s
8 per second
Kling Avatar — ProKling AIThe 1080p talking-avatar tier, and the longest clip in the studio.
  • Quality: 1080p
  • Max length: 5 minutes
16 per second
HeyGen Avatar IVHeyGenPhoto to talking avatar.
  • Quality: 720p / 1080p
  • Max length: 60s
20 per second
Omnihuman 1.5ByteDancePortrait to talking video.
  • Quality: 720p / 1080p, billed the same
  • Max length: 60s
27 per second
Volcengine Lip SyncVolcengineRe-lipsyncs an existing video to new audio.
  • Input: A source video
  • Max length: 60s
8 per second

06 / 3D

3D models

Four generators and two mesh tools, all through fal. Polycount figures are the studio slider’s range, which is deliberately narrower than what the endpoints accept.

Open the 3D studio
All 6 3d engines, with every spec and price
ModelWhat it is forSpecsCredits
Meshy 6MeshyThe best all-round mesh — quad topology, PBR materials and posing.
  • Modes: Image · Multi-view · Text
  • Polycount: 10,000–300,000
  • Extras: PBR · quad topology · A-pose / T-pose
  • Images: Up to 4
240 per generation
Hunyuan 3.1 ProThe highest-detail mesh, with multi-view input.
  • Modes: Image · Multi-view · Text
  • Polycount: 40,000–1,500,000
  • Extras: PBR
  • Images: Up to 4
80 per generation
Trellis 2Fast and affordable, from a single image.
  • Modes: Image only
  • Polycount: 10,000–300,000
  • Images: 1
60 per generation
Tripo 2.5Tripo3DThe fastest mesh in the 3D studio.
  • Modes: Image · Multi-view · Text
  • Polycount: 10,000–300,000
  • Extras: PBR
  • Images: Up to 4
60 per generation
Retexture EngineMeshyRe-textures a mesh you already have, from a prompt or a style image.
  • Input: An existing mesh + 1 style image
  • Extras: PBR
60 per run
RemeshMeshyRemesh, rescale and re-export a model for printing or another pipeline.
  • Input: An existing mesh
  • Polycount: 10,000–300,000
  • Formats: GLB · FBX · OBJ · STL · USDZ
40 per run

How a model gets onto this page

A model appears here once we can actually run it — API access granted, the adapter written, the price set. Access is the vendor’s to grant, not ours to promise, so you will not find a launch date on this page for anything we cannot generate with today.

What you do get is one place to run all of it. The same prompt bar, the same reference uploads, the same library, and one credit balance every model draws from — so moving from an image to a video to a score costs you a dropdown, not a new subscription.

Where two providers serve the same model we run one as primary and the other as a fallback, and the price you see is the same either way.

Every model above, on one balance

No per-vendor subscriptions, no separate accounts, no re-learning a UI each time a new model lands.