Fabric 1.0: The AI avatar model built for lip-sync precision

Highlights
AI avatars have come a long way, but most tools still make you choose between quality and control. Either the lip-sync looks off, or you’re locked into a rigid workflow that doesn’t fit how you actually create.
Fabric 1.0 is different as a dedicated AI avatar model now available on Artlist, it does one thing better than almost anything else out there. With this model, you can make mouths move the way they should.
Here’s everything you need to know about it, and which version to choose for your workflow.
What does Fabric 1.0 do?
Fabric 1.0 is an image to video and audio to video model developed by VEED that runs on a Diffusion Transformer (DiT) architecture. You give it a still image and an audio file, and it generates a realistic talking-head video with lip movements precisely matched to the phonetics and timing of the audio, in any language.
Whether your audio is in English, Spanish, Japanese, or anything in between, Fabric 1.0 reads the phonetics and animates accordingly. For creators producing multilingual content, dubbing videos, or building international audiences, that’s a genuine game-changer.
Fabric 1.0 is purpose-built for speech animation. If you’re looking for realistic camera movements or cinematic motion, this isn’t that. But if you need a face to speak convincingly, it’s hard to beat.
Why lip-sync precision matters
Most AI video tools treat lip-sync as one feature among many. Fabric 1.0 treats it as the feature.
This AI model precisely maps mouth movements to the phonetics and timing of the audio you input — not just the general shape of speech, but the specific articulation of each sound. The result is video that doesn’t feel AI-generated. Mouths move the way real mouths move. That level of realism is what separates content that builds trust from content that gets scrolled past.
For YouTubers, this means voiceovers that actually look like the person speaking them. For marketers, it means spokesperson content without the spokesperson budget. For anyone experimenting with faceless content, it means a face that feels real.
Why use Fabric 1.0?
Fabric 1.0 is a strong fit for:
- Talking-head videos: Spokesperson content, explainers, course material, faceless YouTube channels. Any format where a face needs to deliver information clearly and naturally.
- Product commercials: Give your brand a face without a full production shoot. Upload a product image with a character or ambassador and let Fabric do the work.
- UGC-style content: Fabric 1.0 helps you produce that authentic, direct-to-camera feel at the scale you need for user-generated content.
- Animals, animations, and abstract faces: This is where it gets fun — if it has a face, Fabric can make it talk. Fabric 1.0 by VEED is excellent at bringing non-human faces to life for example, your brand mascot, a cartoon character, an illustrated avatar.
Output is available at 480p and 720p resolution, delivered as MP4, with an incredible maximum duration of 60 seconds per generation.
Supported image inputs include JPG, JPEG, PNG, WebP, GIF, and AVIF. Audio inputs support MP3, OGG, WAV, M4A, and AAC.
Two versions
Fabric 1.0
Fabric 1.0 is the standard model. You can upload your image, upload your audio, and generate up to 60 seconds of lip-synced video. Quality is high, processing is thorough, and the results are professional-grade. This is the go-to for most use cases.
Fabric 1.0 Fast
This is the same quality as the standard model, with a faster processing time. This costs slightly more credits, but the output is comparable, so if you're working to a deadline or generating at volume, Fast is worth it.
8 pro tips: how to get the AI avatars you want with Fabric 1.0
Getting great results from Fabric 1.0 comes down to the inputs you give it. Here's what to do — and what to avoid — across images, audio, and your overall workflow.
- Choose a front-facing, well-lit portrait
The model needs a clear view of the mouth. Faces shot straight-on or at a slight angle work best. Avoid profiles, extreme tilts, or images where the mouth is obscured by hair, hands, or objects.
- Use a high-contrast, uncluttered background
Busy backgrounds can distract the model from the face region. A neutral or solid background helps Fabric focus on animating the face accurately and keeps your output looking polished.
- Keep your audio clean and clearly voiced
Background noise, music, or heavy reverb can reduce lip-sync accuracy. Use dry, studio-quality audio whenever possible — or apply noise reduction before uploading. Clearly articulated speech produces the sharpest results.
- Match the speaking pace to the facial expression
Fast-talking clips or clips with unnatural pauses can cause timing drift. A natural, conversational pace tends to sync most convincingly. If your audio feels rushed, slow it down slightly before upload.
- Test non-human faces with shorter clips first
Animating mascots, cartoon characters, or stylised avatars? Start with short, clearly spoken clips before going full 60-second generations. Non-human faces with exaggerated features can behave differently — testing early saves credits.
- Use Fabric 1.0 Fast for iteration, Standard for finals
Speed through early drafts with Fast to check timing and expression before committing to Standard for your final output. This two-pass approach balances cost and quality without sacrificing either.
- For multilingual content, keep the accent natural
Fabric 1.0 reads phonetics, not the language itself — so dubbed audio works best when the accent and rhythm feel native to the target language.
- Start with a sharp, high-resolution image
While the model outputs at 480p or 720p, starting with a high-resolution image gives it more facial detail to work with. Upscaling a blurry photo won't help.
Get started with Fabric 1.0
If you've ever watched a talking-head video and felt something was slightly off, like when the mouth moves a beat behind, or the vowels not quite matching, you already know why lip-sync precision matters. Fabric 1.0 fixes that.
It's available now on Artlist, ready to turn any image into a face that speaks convincingly. Pick your image, prep your audio, and start generating. The hardest part is choosing which face gets to speak first!




