
我如何用 AI 製作寫實維京人影片(完整工作流程)
Aug 2026

I made a photorealistic Viking short film called Ulysses entirely with AI. No camera, no crew, no actors, no set. Just over four hundred shots, built in about eight days, that look like real footage. This is the full workflow, start to finish, exactly the way it happened on the canvas. Watch the finished film on Flick TV first if you want; everything below is how it got made.
這是一部關於兩個看似普通的遊戲 NPC 在空曠的等待空間中等待的短片,但我們逐漸意識到他們可能比控制他們的玩家更有生命力。我們不再將 NPC 視為靜態的背景程式碼,而是想像他們即使在看不見的時候也繼續存在,擁有自己的意識、不確定性和存在感。故事一開始是一場荒誕的等待遊戲,但慢慢地變成對可見性、表演以及擁有內在生命意味著什麼的反思。
1. 先有故事,再有鏡頭
Ulysses 是一個維京時代的故事:一位蜜酒廳老闆、一名年輕守衛、一位身披盔甲的戰士,還有一條龍。在生成任何內容之前,我先將它拆分成鏡頭,就像為任何影片製作分鏡腳本一樣。每個鏡頭都有粗略的規劃:誰會出現在其中、鏡頭位置在哪裡,以及會發生什麼事。
Make your own AI film
Start with one image and build a whole film, shot by shot.
2. 將角色製作成參考表
這是最重要的一步。我為每個角色生成了一張清晰的三視圖參考表,包含正面、側面和四分之三視角,背景是純灰色,就像選角照片一樣。

這張參考表成為角色的「演員」。該角色之後的每個鏡頭都以它為基礎建立,因此面容不會改變。我的主角參考表在七十九個不同的鏡頭中被重複使用作為基底。服裝和盔甲也採用相同方式處理:我手繪盔甲草圖,然後將草圖轉換成完整的照片級真實全視角圖。

3. 手繪每個鏡頭的草圖
這是最讓人驚訝的步驟。在生成畫面之前,我先粗略地繪製了構圖草圖,只畫形狀:角色站在哪裡、地平線在哪裡、鏡頭看向哪裡。在這部影片中,我為數百個鏡頭都這樣做了。
草圖同時包含了場面調度和鏡頭角度。它快速、可拋棄,並且為模型提供了確切的遵循對象,而不是從文字中猜測構圖。
4. 將草圖轉換成照片級真實畫面
然後我在 Nano Banana 中結合了三樣東西:構圖草圖、角色的參考表和簡短指示,要求生成照片級真實的實拍畫面。提示詞禁止模型偏離:
Use the provided sketch as the only composition reference, use the provided guard
reference as the guard character reference. Convert the scene into a photorealistic
live-action cinematic image. Composition 100% exactly matches the provided sketch.
Camera position 100% unchanged.
粗略的草圖加上固定的角色,從另一端產生出看起來像拍攝出來的鏡頭。

5. 透過編輯而非重新生成來精修
這部影片中幾乎沒有任何內容是從空白文字提示詞開始的。在我的圖像生成中,371 次是對現有畫面的編輯,只有 21 次是文字轉圖像。當畫面稍微有偏差時,比如下顎太重、陰影奇怪或細節錯誤,我會編輯該畫面而不是重新開始。一致性來自於在鎖定的基礎上進行編輯,而不是來自完美的措辭。
6. 動畫化並讓角色開口說話
靜態畫面鎖定後,我使用 Kling 將選定的畫面製作成影片動畫。有些鏡頭是簡單的動作;其他則是對話,角色直接對著鏡頭說話。這些提示詞對表演有具體的描述:
Handheld camera, extreme close-up of his face. His eyes stay locked on the camera,
no drift. He says, "Still, here." Then a brief pause, still staring, with visible
strain, as if forcing the words out.
靜態畫面保持身份特徵;影片則增添呼吸和聲音。
7. The canvas it all happened on
Every step above lives in one workspace. Here is the actual project, zoomed out: sketch rows on the left feeding photoreal rows on the right, character sheets along the bottom, and the iteration ladders in between. 424 images and 38 video shots, all traceable to their references.

Zoom in anywhere and you see the same pattern repeat: one composition, several takes, one survivor. This is what the two leads look like mid-iteration, the guard in his mail and the keeper in his tunic, same faces in every take because every take points at the same sheets.

8. The numbers behind it
為了讓您了解規模,以下是完成專案所包含的內容:
| 來自 Ulysses 畫布 | 數值 |
|---|---|
| 生成的圖像 / 失敗次數 | 424 / 20 |
| 參考圖像連結 | 653(佔所有連結的 92%) |
| 圖像編輯 vs 文字轉圖像 | 371 vs 21(95% 為圖轉圖) |
| 單一參考表重複使用 | 79 個鏡頭 |
| 製作時間 | 約 8 天,一人完成 |
9. What I would tell you before you start
- 先建立參考表。這是整個遊戲的關鍵。沒有參考表的角色會飄移變形。
- 繪製您的鏡頭草圖。一張構圖正確的粗糙草圖勝過一段華麗的文字描述。
- 編輯,不要重新生成。鎖定基礎圖像,盡可能少做改動。
- 預期會有失誤。有二十個鏡頭完全失敗,還有更多需要修正。這很正常;整個流程本來就能承受這些失誤。
創作您自己的作品
Start with one image and build a film shot by shot. The step that matters most is the reference sheet, and our guide to consistent AI characters is the deep dive on exactly that.
常見問題
How long does it take to make an AI film like this?
Ulysses ran 432 generation jobs over about 8 days, one person, no crew. Most of that time is choosing between takes, not waiting on renders.
Which AI models did the film actually use?
Nano Banana for stills and edits (371 of 392 image generations were edits of an existing frame), and Kling for motion: the finished canvas carries 37 Kling video shots, including the spoken dialogue takes.
How do the characters stay consistent for 400 shots?
Reference sheets, wired into everything: the project holds 653 reference-image links, 92 percent of all its connections, and the lead's sheet alone sits under 79 shots. No shot is generated from words alone.
Can the characters actually talk?
Yes. Dialogue shots are animated from a locked still with a performance prompt that specifies the line, the pause, and what the eyes do. The still holds the identity; the video pass adds breath and voice.