
How I Made a Photorealistic Viking Film with AI (Full Workflow)
Aug 2026

I made a photorealistic Viking short film called Ulysses entirely with AI. No camera, no crew, no actors, no set. Just over four hundred shots, built in about eight days, that look like real footage. This is the full workflow, start to finish, exactly the way it happened on the canvas. Watch the finished film on Flick TV first if you want; everything below is how it got made.
It is a short film about two seemingly ordinary game NPCs waiting in an empty holding space, only for us to gradually realize that they may be more alive than the player who controls them. Instead of seeing NPCs as inactive background code, we imagine that they continue to exist even when unseen, carrying their own awareness, uncertainty, and presence. The story begins as an absurd waiting game, but slowly becomes a reflection on what it means to be visible, to perform, and to have inner life.
1. Story first, then shots
Ulysses is a Viking-age story: a mead-hall keeper, a young guard, an armored warrior, and a dragon. Before generating anything, I broke it into shots, the way you would storyboard any film. Every shot got a rough plan for who is in it, where the camera sits, and what happens.
Make your own AI film
Start with one image and build a whole film, shot by shot.
2. Cast the characters as reference sheets
This is the single most important step. For each character I generated a clean three-view reference sheet, front, profile, and three-quarter, on a plain grey background, like a casting photo.

The sheet becomes the character's “actor.” Every later shot of that person is built from it, so the face never drifts. My lead's sheet was reused as the base in seventy-nine separate shots. Wardrobe and armor get the same treatment: I sketched the armor by hand, then converted the sketch into a full photoreal turnaround.

3. Sketch every shot by hand
This is the step that surprises people most. Before generating a frame, I drew a rough sketch of the composition, just shapes: where the character stands, where the horizon sits, where the camera looks. On this film I did that for hundreds of shots.
The sketch is the blocking and the camera in one. It is fast, it is disposable, and it gives the model something exact to follow instead of guessing at a composition from words.
4. Turn the sketch into a photoreal frame
Then I combined three things in Nano Banana, the composition sketch, the character's reference sheet, and a short instruction, and asked for a photorealistic live-action frame. The prompt forbids the model from wandering:
Use the provided sketch as the only composition reference, use the provided guard
reference as the guard character reference. Convert the scene into a photorealistic
live-action cinematic image. Composition 100% exactly matches the provided sketch.
Camera position 100% unchanged.
A rough sketch plus a locked character comes out the other side as a shot that looks filmed.

5. Refine by editing, not re-rolling
Almost nothing in this film came from a blank text prompt. Of my image generations, 371 were edits of an existing frame and only 21 were text-to-image. When a frame was slightly off, a heavy jaw, an odd shadow, a wrong detail, I edited that frame rather than starting over. Consistency comes from editing on top of a locked base, not from perfect wording.
6. Animate, and let them speak
With the stills locked, I animated selected frames into video with Kling. Some shots are simple motion; others are dialogue, with the character talking straight to camera. Those prompts are specific about the performance:
Handheld camera, extreme close-up of his face. His eyes stay locked on the camera,
no drift. He says, "Still, here." Then a brief pause, still staring, with visible
strain, as if forcing the words out.
The still holds the identity; the video adds the breath and the voice.
7. The canvas it all happened on
Every step above lives in one workspace. Here is the actual project, zoomed out: sketch rows on the left feeding photoreal rows on the right, character sheets along the bottom, and the iteration ladders in between. 424 images and 38 video shots, all traceable to their references.

Zoom in anywhere and you see the same pattern repeat: one composition, several takes, one survivor. This is what the two leads look like mid-iteration, the guard in his mail and the keeper in his tunic, same faces in every take because every take points at the same sheets.

8. The numbers behind it
For a sense of scale, here is what the finished project held:
| From the Ulysses canvas | Value |
|---|---|
| Images generated / failures | 424 / 20 |
| Reference-image links | 653 (92% of all connections) |
| Image edits vs text-to-image | 371 vs 21 (95% img2img) |
| One reference sheet reused | 79 shots |
| Build time | about 8 days, one person |
9. What I would tell you before you start
- Build your reference sheets first. They are the whole game. A character without a sheet will drift.
- Sketch your shots. A bad drawing with the right composition beats a beautiful paragraph.
- Edit, do not re-roll. Lock a base image and change as little as possible.
- Expect misses. Twenty shots failed outright and many more needed fixing. That is normal; the pipeline is built to absorb it.
Make your own
Start with one image and build a film shot by shot. The step that matters most is the reference sheet, and our guide to consistent AI characters is the deep dive on exactly that.
Frequently Asked Questions
How long does it take to make an AI film like this?
Ulysses ran 432 generation jobs over about 8 days, one person, no crew. Most of that time is choosing between takes, not waiting on renders.
Which AI models did the film actually use?
Nano Banana for stills and edits (371 of 392 image generations were edits of an existing frame), and Kling for motion: the finished canvas carries 37 Kling video shots, including the spoken dialogue takes.
How do the characters stay consistent for 400 shots?
Reference sheets, wired into everything: the project holds 653 reference-image links, 92 percent of all its connections, and the lead's sheet alone sits under 79 shots. No shot is generated from words alone.
Can the characters actually talk?
Yes. Dialogue shots are animated from a locked still with a performance prompt that specifies the line, the pause, and what the eyes do. The still holds the identity; the video pass adds breath and voice.