
AIでフォトリアルなバイキング映画を制作した方法(完全ワークフロー)
Aug 2026

I made a photorealistic Viking short film called Ulysses entirely with AI. No camera, no crew, no actors, no set. Just over four hundred shots, built in about eight days, that look like real footage. This is the full workflow, start to finish, exactly the way it happened on the canvas. Watch the finished film on Flick TV first if you want; everything below is how it got made.
これは、空の待機スペースで待機する2人の一見普通のゲームNPCについてのショートフィルムです。しかし次第に、彼らは自分たちを操作するプレイヤーよりも生きている可能性があることに気づきます。NPCを非アクティブな背景コードとして見るのではなく、見られていない時でも存在し続け、独自の意識、不確実性、存在感を持っていると想像します。物語は不条理な待機ゲームとして始まりますが、徐々に見られること、演じること、そして内的な生を持つことの意味についての考察へと変わっていきます。
1. ストーリーが先、次にショット
Ulyssesはヴァイキング時代の物語です。蜂蜜酒ホールの管理人、若い警備兵、鎧を纏った戦士、そして竜が登場します。何かを生成する前に、他の映画で絵コンテを作るように、ショットに分割しました。各ショットには、誰が登場するか、カメラの位置、何が起こるかという大まかなプランを立てました。
Make your own AI film
Start with one image and build a whole film, shot by shot.
2. キャラクターをリファレンスシートとしてキャスティングする
これが最も重要なステップです。各キャラクターについて、キャスティング写真のように、無地のグレー背景で正面、横顔、斜め前の3方向からの明確なリファレンスシートを生成しました。

このシートがキャラクターの「俳優」になります。その人物のその後のすべてのショットはこれを基に構築されるため、顔がぶれることはありません。主役のシートは79の異なるショットでベースとして再利用されました。衣装や鎧も同じ処理をします。鎧を手描きでスケッチし、そのスケッチをフォトリアルな完全な回転図に変換しました。

3. すべてのショットを手描きでスケッチする
これが最も驚かれるステップです。フレームを生成する前に、構図の大まかなスケッチを描きました。形だけです。キャラクターが立つ位置、地平線の位置、カメラの向き。この映画では数百のショットでこれを行いました。
スケッチはブロッキングとカメラを一つにまとめたものです。速く、使い捨てで、言葉から構図を推測させるのではなく、モデルに従うべき正確なものを与えます。
4. スケッチをフォトリアルなフレームに変換する
次に、Nano Bananaで3つのもの、構図スケッチ、キャラクターのリファレンスシート、短い指示を組み合わせて、フォトリアルな実写フレームを生成しました。プロンプトはモデルが逸脱することを禁止します。
Use the provided sketch as the only composition reference, use the provided guard
reference as the guard character reference. Convert the scene into a photorealistic
live-action cinematic image. Composition 100% exactly matches the provided sketch.
Camera position 100% unchanged.
大まかなスケッチと固定されたキャラクターが、反対側から撮影されたように見えるショットとして出てきます。

5. 再生成ではなく編集で洗練させる
この映画のほぼすべては、白紙のテキストプロンプトから生成されたものではありません。私の画像生成のうち、371件は既存フレームの編集で、テキストから画像への生成はわずか21件でした。フレームが少しずれていたり、顎が重かったり、奇妙な影があったり、細部が間違っていたりした場合、最初からやり直すのではなく、そのフレームを編集しました。一貫性は、完璧な言葉遣いからではなく、固定されたベースの上に編集を重ねることから生まれます。
6. アニメーション化して、彼らに語らせる
静止画を固定した後、選択したフレームをKlingで動画にアニメーション化しました。シンプルな動きのショットもあれば、キャラクターがカメラに向かって直接話す対話のショットもあります。これらのプロンプトは、演技について具体的に指定しています:
Handheld camera, extreme close-up of his face. His eyes stay locked on the camera,
no drift. He says, "Still, here." Then a brief pause, still staring, with visible
strain, as if forcing the words out.
静止画がアイデンティティを保持し、動画が息吹と声を加えます。
7. The canvas it all happened on
Every step above lives in one workspace. Here is the actual project, zoomed out: sketch rows on the left feeding photoreal rows on the right, character sheets along the bottom, and the iteration ladders in between. 424 images and 38 video shots, all traceable to their references.

Zoom in anywhere and you see the same pattern repeat: one composition, several takes, one survivor. This is what the two leads look like mid-iteration, the guard in his mail and the keeper in his tunic, same faces in every take because every take points at the same sheets.

8. The numbers behind it
規模感をつかむために、完成したプロジェクトに含まれるデータを示します:
| Ulyssesのキャンバスから | 数値 |
|---|---|
| 生成された画像 / 失敗 | 424 / 20 |
| リファレンス画像のリンク | 653(全接続の92%) |
| 画像編集 vs テキストから画像 | 371 vs 21(95% img2img) |
| 1つのリファレンスシートの再利用 | 79ショット |
| 制作時間 | 約8日間、1人 |
9. What I would tell you before you start
- まずリファレンスシートを作る。これがすべてです。シートのないキャラクターはブレてしまいます。
- ショットをスケッチする。適切な構図の下手な絵は、美しい文章に勝ります。
- 再生成ではなく、編集する。ベース画像を固定して、できるだけ少ない変更にとどめます。
- 失敗を想定する。20ショットが完全に失敗し、さらに多くのショットが修正を必要としました。これは正常なことで、パイプラインはそれを吸収するように構築されています。
自分で作ってみよう
Start with one image and build a film shot by shot. The step that matters most is the reference sheet, and our guide to consistent AI characters is the deep dive on exactly that.
よくある質問
How long does it take to make an AI film like this?
Ulysses ran 432 generation jobs over about 8 days, one person, no crew. Most of that time is choosing between takes, not waiting on renders.
Which AI models did the film actually use?
Nano Banana for stills and edits (371 of 392 image generations were edits of an existing frame), and Kling for motion: the finished canvas carries 37 Kling video shots, including the spoken dialogue takes.
How do the characters stay consistent for 400 shots?
Reference sheets, wired into everything: the project holds 653 reference-image links, 92 percent of all its connections, and the lead's sheet alone sits under 79 shots. No shot is generated from words alone.
Can the characters actually talk?
Yes. Dialogue shots are animated from a locked still with a performance prompt that specifies the line, the pause, and what the eyes do. The still holds the identity; the video pass adds breath and voice.