
AI로 포토리얼리스틱 바이킹 영화를 만든 방법 (전체 워크플로우)
Aug 2026

I made a photorealistic Viking short film called Ulysses entirely with AI. No camera, no crew, no actors, no set. Just over four hundred shots, built in about eight days, that look like real footage. This is the full workflow, start to finish, exactly the way it happened on the canvas. Watch the finished film on Flick TV first if you want; everything below is how it got made.
빈 대기 공간에서 기다리고 있는 두 명의 평범해 보이는 게임 NPC에 관한 단편 영화입니다. 그들을 조종하는 플레이어보다 그들이 더 살아있을 수 있다는 것을 점차 깨닫게 됩니다. NPC를 비활성 배경 코드로 보는 대신, 우리는 그들이 보이지 않을 때도 계속 존재하며 자신만의 인식, 불확실성, 그리고 존재감을 지니고 있다고 상상합니다. 이야기는 부조리한 기다림 게임으로 시작하지만, 천천히 가시성, 수행성, 그리고 내면의 삶이 무엇을 의미하는지에 대한 성찰로 변해갑니다.
1. 스토리가 먼저, 그다음 샷
Ulysses는 바이킹 시대 이야기입니다: 미드홀 관리인, 젊은 경비병, 갑옷을 입은 전사, 그리고 용. 무언가를 생성하기 전에 저는 이야기를 샷으로 나눴습니다. 영화를 스토리보드로 만드는 것처럼요. 모든 샷에 대해 누가 등장하는지, 카메라가 어디에 위치하는지, 무슨 일이 일어나는지에 대한 대략적인 계획을 세웠습니다.
Make your own AI film
Start with one image and build a whole film, shot by shot.
2. 캐릭터를 레퍼런스 시트로 캐스팅하기
이것이 가장 중요한 단계입니다. 각 캐릭터마다 정면, 측면, 3/4 각도의 깔끔한 3면 레퍼런스 시트를 생성했습니다. 캐스팅 사진처럼 평범한 회색 배경에서요.

시트는 캐릭터의 "배우"가 됩니다. 그 사람의 모든 이후 샷은 이것을 기반으로 만들어지므로 얼굴이 흐트러지지 않습니다. 제 주인공의 시트는 79개의 개별 샷에서 베이스로 재사용되었습니다. 의상과 갑옷도 같은 방식으로 처리합니다: 저는 갑옷을 손으로 스케치한 다음, 스케치를 완전한 포토리얼 턴어라운드로 변환했습니다.

3. 모든 샷을 손으로 스케치하기
이것이 사람들을 가장 놀라게 하는 단계입니다. 프레임을 생성하기 전에 저는 구성에 대한 대략적인 스케치를 그렸습니다. 단순한 형태만으로요: 캐릭터가 서 있는 위치, 수평선이 놓인 위치, 카메라가 바라보는 곳. 이 영화에서 저는 수백 개의 샷에 대해 그렇게 했습니다.
스케치는 블로킹과 카메라를 하나로 합친 것입니다. 빠르고, 일회용이며, 단어로 구성을 추측하는 대신 모델이 따라갈 정확한 무언가를 제공합니다.
4. 스케치를 포토리얼 프레임으로 바꾸기
그런 다음 Nano Banana에서 세 가지를 결합했습니다. 구성 스케치, 캐릭터의 레퍼런스 시트, 그리고 짧은 지시사항. 그리고 포토리얼 실사 프레임을 요청했습니다. 프롬프트는 모델이 벗어나는 것을 금지합니다:
Use the provided sketch as the only composition reference, use the provided guard
reference as the guard character reference. Convert the scene into a photorealistic
live-action cinematic image. Composition 100% exactly matches the provided sketch.
Camera position 100% unchanged.
대략적인 스케치와 고정된 캐릭터는 반대편에서 촬영된 것처럼 보이는 샷으로 나옵니다.

5. 다시 생성하지 말고 편집으로 다듬기
이 영화에서 빈 텍스트 프롬프트만으로 만든 것은 거의 없습니다. 제가 생성한 이미지 중 371개는 기존 프레임을 편집한 것이었고, 텍스트-이미지 생성은 단 21개뿐이었습니다. 프레임이 약간 어긋났을 때, 턱이 너무 각졌거나, 이상한 그림자가 있거나, 잘못된 디테일이 있을 때, 처음부터 다시 시작하지 않고 해당 프레임을 편집했습니다. 일관성은 완벽한 문구가 아니라 고정된 베이스 위에 편집을 쌓는 것에서 나옵니다.
6. 애니메이션 적용하고 말하게 하기
스틸 이미지가 확정되면, 선택한 프레임들을 Kling으로 비디오로 애니메이션했습니다. 일부 샷은 단순한 움직임이고, 다른 샷들은 캐릭터가 카메라를 똑바로 보며 말하는 대화 장면입니다. 이러한 프롬프트는 연기에 대해 구체적입니다:
Handheld camera, extreme close-up of his face. His eyes stay locked on the camera,
no drift. He says, "Still, here." Then a brief pause, still staring, with visible
strain, as if forcing the words out.
스틸 이미지는 정체성을 유지하고, 비디오는 숨결과 목소리를 더합니다.
7. The canvas it all happened on
Every step above lives in one workspace. Here is the actual project, zoomed out: sketch rows on the left feeding photoreal rows on the right, character sheets along the bottom, and the iteration ladders in between. 424 images and 38 video shots, all traceable to their references.

Zoom in anywhere and you see the same pattern repeat: one composition, several takes, one survivor. This is what the two leads look like mid-iteration, the guard in his mail and the keeper in his tunic, same faces in every take because every take points at the same sheets.

8. The numbers behind it
규모를 가늠하기 위해, 완성된 프로젝트에 포함된 내용은 다음과 같습니다:
| Ulysses 캔버스 기준 | 값 |
|---|---|
| 생성된 이미지 / 실패 | 424 / 20 |
| 레퍼런스 이미지 연결 | 653 (전체 연결의 92%) |
| 이미지 편집 vs 텍스트-이미지 | 371 vs 21 (95% img2img) |
| 재사용된 레퍼런스 시트 1개 | 79샷 |
| 제작 시간 | 약 8일, 1인 |
9. What I would tell you before you start
- 레퍼런스 시트를 먼저 만드세요. 이것이 전부입니다. 시트 없는 캐릭터는 흔들립니다.
- 샷을 스케치하세요. 올바른 구도의 서툰 그림이 아름다운 문장보다 낫습니다.
- 다시 생성하지 말고 편집하세요. 베이스 이미지를 고정하고 최소한만 변경하세요.
- 실패를 예상하세요. 20개의 샷이 완전히 실패했고 더 많은 샷들이 수정이 필요했습니다. 이것은 정상이며, 파이프라인은 이를 흡수하도록 구축되어 있습니다.
직접 만들어보세요
Start with one image and build a film shot by shot. The step that matters most is the reference sheet, and our guide to consistent AI characters is the deep dive on exactly that.
자주 묻는 질문
How long does it take to make an AI film like this?
Ulysses ran 432 generation jobs over about 8 days, one person, no crew. Most of that time is choosing between takes, not waiting on renders.
Which AI models did the film actually use?
Nano Banana for stills and edits (371 of 392 image generations were edits of an existing frame), and Kling for motion: the finished canvas carries 37 Kling video shots, including the spoken dialogue takes.
How do the characters stay consistent for 400 shots?
Reference sheets, wired into everything: the project holds 653 reference-image links, 92 percent of all its connections, and the lead's sheet alone sits under 79 shots. No shot is generated from words alone.
Can the characters actually talk?
Yes. Dialogue shots are animated from a locked still with a performance prompt that specifies the line, the pause, and what the eyes do. The still holds the identity; the video pass adds breath and voice.