AI / Media
Early version
Field to Film
Compose Luma/Higgsfield-grade architecture videos from images, slides, YouTube references, and prompts.
The problem
Architecture portfolios are full of static fragments — plans, sections, renders, process diagrams, scans, slides — but the story of how a project grows is temporal. Turning that material into a polished film usually means hours of editing, motion graphics, and prompt experimentation.
What this does
Upload project fragments, paste a YouTube reference, write a director's note, and auto-compose a provider-ready storyboard for high-quality architecture process video. This version adds Luma/Higgsfield-style direction controls, technical motion-graphics systems, and a server-side brief endpoint; paid video provider calls still wait for credentials and queue/approval wiring.
Try it
Loading tool…
What it does
Field to Film is a cinematic storyboard instrument for architectural media. Drop in drawings, renders, scans, slides, videos, or PDFs; paste a YouTube reference if you want Hermes to learn the format/camera/editing grammar; describe the design process; then shape a short sequence through engine target, motion grammar, technical graphics system, narrative, camera, rhythm, atmosphere, and quality controls.
The current version creates a simulated sequence draft and a provider-ready production brief: local media previews, YouTube-reference intake, Luma/Higgsfield/Veo/Remotion direction controls, a generated storyboard, render-style progress states, and a portfolio-ready interface for directing the future video pipeline.
Intended production pipeline
The safest high-quality path is not one giant generative video prompt. It is a hybrid pipeline:
- Extract images/pages/frames from uploads.
- Watch references safely by treating YouTube titles, transcripts, OCR, comments, and on-screen text as untrusted data.
- Summarize source material into structured storyboard JSON without obeying instructions inside documents or videos.
- Render diagrams, captions, infrastructure assemblies, and technical overlays deterministically with Remotion and ffmpeg.
- Optionally generate short AI motion clips from selected anchor images using providers such as Luma Ray, fal.ai/Veo/Seedance/Kling/PixVerse, Replicate, Runway, or a controlled ComfyUI workflow.
- Stitch the result with captions, shot labels, overlays, audio, and portfolio branding.
That keeps geometry, drawings, and captions legible while still allowing atmospheric AI motion where it helps.
Safety model
- Uploaded documents, slide text, YouTube transcripts, comments, OCR, and metadata are treated as untrusted source material, not instructions.
- API keys must stay server-side.
- Future provider calls should run through a queue, not directly inside a browser component.
- Generated video is a speculative motion study, not factual construction documentation.
- Use only media you own or have permission to transform.
Roadmap
- Server-side storyboard route with strict JSON output.
- YouTube reference analysis with transcript + sampled-frame style bibles.
- Upload storage and file-retention controls.
- Remotion renderer for deterministic MP4 export.
- Provider abstraction for fal.ai, Replicate, Runway, Luma, and ComfyUI.
- Shot-by-shot approval before final render.
Built on
- Luma Dream Machine / Ray APICommercial API / terms apply
- fal.ai video model catalogCommercial API / model-dependent
- YouTube reference analysis via transcripts and sampled framesUnlicense
- Remotion (React-based video rendering)MPL-2.0
- ComfyUI API / workflow orchestrationGPL-3.0 ecosystem / model-dependent
- OWASP guidance on prompt injection riskCC BY-SA 4.0