我如何用 AI 制作写实维京影片(完整工作流程)

· Flick 内部影片制作人 · LinkedIn

Aug 2026

Low-angle cinematic portrait of a weathered Norse man against a stone monolith and blue sky, from the AI short film Ulysses
Low-angle cinematic portrait of a weathered Norse man against a stone monolith and blue sky, from the AI short film Ulysses

I made a photorealistic Viking short film called Ulysses entirely with AI. No camera, no crew, no actors, no set. Just over four hundred shots, built in about eight days, that look like real footage. This is the full workflow, start to finish, exactly the way it happened on the canvas. Watch the finished film on Flick TV first if you want; everything below is how it got made.

这是一部关于两个看似普通的游戏 NPC 在空旷的等候空间中等待的短片,我们逐渐意识到他们可能比控制他们的玩家更有生命力。我们不再将 NPC 视为不活跃的背景代码,而是想象他们即使在不被看见时也继续存在,承载着自己的意识、不确定性和存在感。故事始于一场荒诞的等待游戏,但慢慢演变成对可见性、表演以及拥有内在生命意味着什么的反思。

1. 先有故事,再有镜头

Ulysses 是一个维京时代的故事:一个蜂蜜酒馆老板、一个年轻守卫、一个身着盔甲的战士和一条龙。在生成任何内容之前,我将它分解成镜头,就像你为任何影片制作故事板一样。每个镜头都有一个粗略的计划,包括谁在其中、摄像机位置以及发生了什么。

Make your own AI film

Start with one image and build a whole film, shot by shot.

2. 将角色设定为参考表

这是最重要的一步。对于每个角色,我生成了一个清晰的三视图参考表,包括正面、侧面和四分之三视角,背景为纯灰色,就像选角照片一样。

Three-view full-body reference sheet of a bearded Norse man in a wool tunic on a grey background
Three-view full-body reference sheet of a bearded Norse man in a wool tunic on a grey background

这个参考表成为角色的"演员"。该角色的每个后续镜头都基于它构建,因此面部不会发生漂移。我的主角参考表在 79 个不同的镜头中被重复用作基础。服装和盔甲也采用同样的处理方式:我手工绘制盔甲草图,然后将草图转换为完整的照片级写实全方位展示。

Three-view reference sheet of an armored Viking warrior in a horned helmet holding an axe
Three-view reference sheet of an armored Viking warrior in a horned helmet holding an axe

3. 手绘每个镜头的草图

这是最让人惊讶的步骤。在生成画面之前,我会粗略绘制构图草图,只是一些形状:角色站在哪里、地平线在哪里、摄像机看向哪里。在这部影片中,我为数百个镜头都做了这个工作。

草图将场景调度和摄像机合二为一。它快速、可抛弃,并且为模型提供了准确的遵循指引,而不是仅凭文字猜测构图。

4. 将草图转换为照片级写实画面

然后我在 Nano Banana 中结合了三样东西:构图草图、角色的参考表和简短的指令,并要求生成照片级写实的实拍画面。提示词禁止模型偏离:

Use the provided sketch as the only composition reference, use the provided guard
reference as the guard character reference. Convert the scene into a photorealistic
live-action cinematic image. Composition 100% exactly matches the provided sketch.
Camera position 100% unchanged.

粗略的草图加上锁定的角色,最终输出的镜头看起来就像拍摄的一样。

The same Norse man from the reference sheet, running along a blue-painted wall on white sand, a finished cinematic frame
The same Norse man from the reference sheet, running along a blue-painted wall on white sand, a finished cinematic frame

5. 通过编辑优化,而非重新生成

这部影片中几乎没有任何内容来自空白文本提示词。在我的图像生成中,371 次是对现有画面的编辑,只有 21 次是文本生成图像。当画面稍有偏差时,比如下巴过重、阴影奇怪或细节错误,我会编辑该画面而不是从头开始。一致性来自于在锁定的基础上进行编辑,而非完美的措辞。

6. 制作动画,让角色开口说话

静态画面锁定后,我使用 Kling 将选定的画面制作成视频动画。有些镜头是简单的运动;其他镜头是对话,角色直接对着镜头说话。这些提示词对表演有具体要求:

Handheld camera, extreme close-up of his face. His eyes stay locked on the camera,
no drift. He says, "Still, here." Then a brief pause, still staring, with visible
strain, as if forcing the words out.

静态画面保持了身份特征;视频则加入了呼吸和声音。

7. The canvas it all happened on

Every step above lives in one workspace. Here is the actual project, zoomed out: sketch rows on the left feeding photoreal rows on the right, character sheets along the bottom, and the iteration ladders in between. 424 images and 38 video shots, all traceable to their references.

The full Ulysses project canvas in Flick, zoomed out: hand-drawn composition sketches on the left connected to photoreal frames on the right, character reference sheets along the bottom
The full Ulysses project canvas in Flick, zoomed out: hand-drawn composition sketches on the left connected to photoreal frames on the right, character reference sheets along the bottom

Zoom in anywhere and you see the same pattern repeat: one composition, several takes, one survivor. This is what the two leads look like mid-iteration, the guard in his mail and the keeper in his tunic, same faces in every take because every take points at the same sheets.

Close view of the Ulysses canvas showing repeated takes of the two leads, an armored guard and a man in a wool tunic, standing on white sand
Close view of the Ulysses canvas showing repeated takes of the two leads, an armored guard and a man in a wool tunic, standing on white sand

8. The numbers behind it

为了了解规模,以下是完成项目所包含的内容:

来自 Ulysses 画布数值
生成的图像 / 失败次数424 / 20
参考图像链接653(占所有连接的 92%)
图像编辑 vs 文本生成图像371 vs 21(95% 为图像到图像)
单个参考表重复使用79 个镜头
制作时间约 8 天,一人完成

9. What I would tell you before you start

  • 首先制作参考表。这是整个流程的核心。没有参考表的角色会失去一致性。
  • 绘制镜头草图。一幅构图正确的粗糙草图胜过一段华丽的文字描述。
  • 编辑,而非重新生成。锁定基础图像,尽可能少做改动。
  • 预期会有失败。有 20 个镜头完全失败,还有更多需要修复。这很正常;整个流程能够消化这些问题。

制作你自己的作品

Start with one image and build a film shot by shot. The step that matters most is the reference sheet, and our guide to consistent AI characters is the deep dive on exactly that.

常见问题

How long does it take to make an AI film like this?

Ulysses ran 432 generation jobs over about 8 days, one person, no crew. Most of that time is choosing between takes, not waiting on renders.

Which AI models did the film actually use?

Nano Banana for stills and edits (371 of 392 image generations were edits of an existing frame), and Kling for motion: the finished canvas carries 37 Kling video shots, including the spoken dialogue takes.

How do the characters stay consistent for 400 shots?

Reference sheets, wired into everything: the project holds 653 reference-image links, 92 percent of all its connections, and the lead's sheet alone sits under 79 shots. No shot is generated from words alone.

Can the characters actually talk?

Yes. Dialogue shots are animated from a locked still with a performance prompt that specifies the line, the pause, and what the eyes do. The still holds the identity; the video pass adds breath and voice.