AI video face swap finds and tracks the people in a clip, matches each photo you add to the person you picked, and generates new frames in which that person looks like the photo while the motion, camera, and lighting stay the same. Then the original sound goes back on. Because the frames are generated rather than edited, you get a new take on the scene, not a pixel-perfect copy, and some takes need another try.
How does AI video face swap work?
It comes down to four steps.
Step 1: Finding and tracking the people
First, the AI finds every person in the video and tracks each one from frame to frame, so it can tell that the dancer on the left is the same person after she spins around. It also outlines the area each person takes up, so only that person changes.
In vswp, the search starts on its own after you upload a video of 1 to 120 seconds, up to 150 MB. It covers the first 30 seconds, or another part of up to 30 seconds that you choose, and usually takes a couple of minutes. Videos from the Trends library skip it because their people are already found.
The people found appear as cards, main people first; anyone seen only briefly or far away is grouped as the background.
Step 2: Matching a photo to a person
Next, you add a photo to the card of each person you want to replace. The photo is a reference, not a sticker: the AI learns what the person looks like (face, hair, and, if the photo shows them, clothes and build) and redraws them in the video.
Up to 6 different photos (JPEG, PNG, or WebP) fit in one video, enough for a whole group of friends, and everyone without a photo stays as in the original. Since the photo is what you control most, choose it with care.
Step 3: Generating new frames
Now the AI draws new frames for the processed part: the chosen people get their new look, while the moves, timing, camera, and lighting follow the original. Because the frames are new, the face takes on the light of the scene, and no two runs come out identical.
Step 4: Putting the sound back
Finally, the frames become a video again and the original audio is added back, so the music, voices, and timing stay as they were. A face swap changes how people look, not how they sound.
In vswp, you get the processed part, up to 30 seconds, as an MP4 at 480p or 720p; the frame rate follows the source, up to about 30 fps.
Face swap or whole-person swap: what's the difference?
A face swap changes only the face and hair. A whole-person swap also brings the clothes and silhouette from the photo. In vswp, the Keep the original look switch picks one:
| Keep the original look on | Keep the original look off | |
|---|---|---|
| What changes | Face and hair | Face, hair, clothes, and silhouette |
| Photo you need | A clear portrait is enough | A full-body photo of one person |
| Good for | Costumes, uniforms, matching outfits | Seeing yourself head to toe in your own clothes |
A portrait can't pass on clothes it doesn't show. For the whole process, see how to put yourself in a video.
What results can you realistically expect?
A good take looks like the original clip with a different person in it. But this is generative AI, not frame-exact editing: likeness varies from take to take, and some typical artifacts (visible glitches) can appear:
- Drifting likeness. The person can look like a relative of the one in the photo, or more like them in some frames than others.
- Hands. Fingers can merge, bend oddly, or change in number.
- Objects over the face. A hand, hair, or a phone crossing the face can blend into it for a moment.
- Very fast camera motion. Quick pans and heavy shake can smear the new face for a few frames.
- Crowded scenes. When many people overlap, details can get mixed up between neighbors.
- People who are hard to see. Someone seen only very briefly or heavily blurred may not be found at all.
How do you get a better take?
- Start with a good photo. One person, sharp, in even light, face clearly visible, no sunglasses or filters.
- Match the switch to the photo. Only a portrait? Leave Keep the original look on. Have a full-body photo? You can turn it off to bring in your own clothes.
- Pick a calm part of the clip. Choose 30 seconds with clear faces and a steadier camera.
- Make another take. Press Create video again for a new variation and keep the one you like. Each take is a separate render with its own price; if one fails or you cancel it, the Stars come back automatically.
Is this the same as a deepfake?
"Deepfake" is the everyday name for a video in which AI makes a person look like someone else or seem to do something they never did. AI video face swap belongs to the same family, and the steps above describe the general idea. The difference is consent and honesty: swapping in friends who agreed is very different from putting someone into a video without permission or passing it off as real.
Pressing Create video in vswp confirms you have the right to use the video and the photos; our guide to face swap consent and safety explains what that means.
FAQ
How long does an AI video face swap take?
Finding the people in an uploaded video usually takes a couple of minutes, and the swapped video is typically ready in minutes. The app shows the estimated time before you start.
Does the voice change too?
No. The original audio is kept, so the voice in the result is the voice from the source clip.
Can I try it for free?
Yes. You get a free preview: about 5 seconds of your video at 480p, with a small "vswp.app" mark, and if it fails, the free try comes back. The full video is paid in Telegram Stars (⭐), and your first paid video is 50% off. The price depends on length, quality, and number of people, and the exact price is shown before you start.