Inspiration

Filmmakers often know what they want a shot to feel like long before they know exactly how to create it.

A director might imagine a camera weaving through a room, orbiting a character, or flying between buildings. But turning that idea into an actual 3D camera path is still surprisingly difficult. Traditional storyboards are great for framing individual shots, but they do not communicate movement, timing, depth, clearance, or how a shot flows through space.

Today, testing those ideas usually means either spending hours inside complex 3D software or waiting until production to see if the shot actually works.

Drone cinematography makes this problem especially obvious. A path that looks good in your head might be too tight, too fast, or physically impossible once you are on location with a pilot, crew, equipment, and limited shooting time.

We built Camvas because we wanted a faster way to explore camera ideas before committing to them.

Camvas gives creators a collaborative 3D space where they can build a scene, describe how they want the camera to move, and immediately see that idea take shape. Whether it is a simple cinematic camera move or an aggressive drone fly-through, creators can experiment, iterate, and understand the shot in 3D before production begins.

What it does

You have an idea for a shot. In Camvas, you can build the scene around it: arrange a 3D environment, search online libraries for characters and props, and direct the camera by saying things like “pan up.”

You can create cinematic moves and drone-style shots, then preview and adjust the camera path on a timeline. Teammates can work in the same scene in real time, while you organize multiple scenes into a larger project or storyboard. Camvas takes you from an idea to a camera-ready scene in one studio.

How we built it

Camvas combines a 3D editor, an agentic director, and a camera path planning system in one web application.

Tech stack

Next.js, React, and TypeScript power the application, editor, and timeline.

PlayCanvas renders interactive 3D models and Gaussian splat captures. We use Three.js for camera and geometry calculations.

Codex powers scene understanding, natural-language editing, and drone shot planning.

Python, PyTorch, and CinemaTraj support optional camera trajectory optimization.

Yjs and WebRTC synchronize scene edits and collaborator presence.

Sketchfab, Jamendo, and Freesound provide searchable models, music, and sound effects.

Railway hosts the application and backend.

Architecture Overview

+----------------------------------------------------------+
|                  CAMVAS CREATIVE STUDIO                   |
|                                                          |
|  +----------------+ +----------------+ +--------------+  |
|  | Next.js 16     | | TypeScript 7   | | Motion       |  |
|  | React 19       | | CSS Modules    | | Lucide       |  |
|  +----------------+ +----------------+ +--------------+  |
|                                                          |
|  Scene Editor / Camera Paths / Actor Blocking / Timeline  |
+----------------------------+-----------------------------+
                             |
                             v
+----------------------------------------------------------+
|                 REAL-TIME SPATIAL ENGINE                  |
|                                                          |
|  +------------------------+ +-------------------------+  |
|  | PlayCanvas             | | Camera & Geometry       |  |
|  | - Gaussian splatting   | | - Three.js spatial math |  |
|  | - Streamed LOD         | | - Blockout geometry     |  |
|  | - GLB animation        | | - Editable flight paths |  |
|  +------------------------+ +-------------------------+  |
+----------------------------+-----------------------------+
                             |
                             v
+----------------------------------------------------------+
|                    AI & ASSET SERVICES                   |
|                                                          |
|  +----------------+ +----------------+ +--------------+  |
|  | Codex Director | | 3D Sources     | | Sound        |  |
|  | - Shot plans   | | - SuperSplat   | | - Jamendo    |  |
|  | - Visual review| | - Sketchfab    | | - Freesound  |  |
|  | - Scene edits  | | - Blender/GLB  | | - Web Audio  |  |
|  +----------------+ +----------------+ +--------------+  |
|                                                          |
|       Node.js API Routes / Optional Python Optimizer     |
+----------------------------+-----------------------------+
                             |
                             v
+----------------------------------------------------------+
|                 COLLABORATION & PROJECTS                 |
|                                                          |
|  +------------------------+ +-------------------------+  |
|  | Yjs + WebRTC           | | Project Storage         |  |
|  | - Shared scene state   | | - Versioned JSON        |  |
|  | - Live presence        | | - Browser persistence   |  |
|  | - Shared cursors       | | - Import / Export       |  |
|  +------------------------+ +-------------------------+  |
+----------------------------+-----------------------------+
                             |
                             v
+----------------------------------------------------------+
|                   BROWSER-NATIVE EXPORT                  |
|                                                          |
|       Web Audio --> Offline Mix --> WebCodecs             |
|                         + Mediabunny                     |
|                                                          |
|       Supersampling / Motion Blur / Titles / Fades        |
|                                                          |
|                 MP4 / WebM --> Films Gallery              |
+----------------------------------------------------------+

From a capture to a scene the agent can understand

Camvas can load conventional 3D assets and 3D Gaussian splat captures. A splat gives us a detailed view of a real space, but it does not automatically tell the camera planner where the walls, objects, or safe passages are. We built separate geometry workflows to fit visual blocks to a capture and to generate approximate collision boxes from its sampled data. Users can inspect and correct those boxes before using them for path refinement.

We also build a scene graph: a structured map of the scene with nodes for meshes, cameras, actors, props, landmarks, and labeled regions. Its edges describe relationships such as contains, targets, and visits this waypoint. Nodes can carry measured bounds, positions, and timed movement tracks. Codex uses this graph to reason about requests such as “fly through the entrance, reveal the chair, then settle on a wide shot” with specific subjects and spatial relationships.

An agent that can direct and edit

We run the Codex director with GPT-6 Astra at medium reasoning effort. We chose it for tasks that require reasoning across the scene graph, measured geometry, camera targets, and the timing of a shot. The agent proposes edits and drone routes; Camvas then checks those proposals against geometry and motion constraints before applying them.

Users can type or speak a direction to the Codex director. Its harness supplies the current scene state, exposes searches for models and audio, and converts the agent’s response into structured actions. The agent can inspect search results, choose assets, place objects, block actors, and create camera moves. Camvas validates IDs, coordinates, timing, and other action parameters before applying an edit.

For supported mesh scenes, the drone planning flow goes further: Codex proposes visual beats and camera control points, Camvas generates a path and checks it against geometry and motion limits, and Codex reviews rendered frames from the proposed shot. Failed checks can feed another proposal. This lets the agent reason about both the story of a shot and the route between its key moments.

The math behind the camera move

We also integrated the DirectPoseOptimizer from CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents (Li et al., 2024).

Our local path refinement optimizes sampled camera positions pᵢ around an intended route pᵢ⁰:

$$E = \sum_{i} \left[ 20(g_i + g_i^2) + 0.2 | p_i - p_i^0 |^2 \right] + \sum_{i} | -p_i + 3p_{i+1} - 3p_{i+2} + p_{i+3} |^2$$

Here ( g_i = \max(0, 0.3 - d(p_i)) ), where ( d(p_i) ) is the signed distance to the nearest obstacle box. The objective balances clearance, staying close to the intended shot, and smooth motion. Required waypoints remain fixed, and each adjusted sample can move at most 0.75 metres.

We also integrated CinemaTraj’s CPU position optimizer as a separate refinement option. It uses obstacle bounds and, for splat scenes, a reviewed coverage area. After optimization, Camvas densely samples the path that will actually play back and rejects results that violate its 0.25-metre clearance margin. Camera shots remain visible and editable on the timeline.

Finally, Yjs and WebRTC allow teammates to work in the same scene, while project storage keeps multiple scenes together as a storyboard.

We integrated external model and audio services to help users populate their scenes, and we deployed the application and backend on Railway with a custom GoDaddy domain.

Challenges we ran into

The hardest challenge was making drone-style camera movement both safe and spatially aware. Our imported environment was essentially one large mesh. Although mesh colliders could tell us when the camera might hit a wall or object, they could not explain what anything represented. The AI had no built-in understanding of concepts such as “the kitchen,” “the living room,” or “the room beside the pool.”

This created two separate problems: the drone needed geometric awareness to avoid collisions, while the AI needed semantic awareness to understand where the user wanted it to go. Sending the entire environment through a vision-language model was too slow and unreliable for interactive use.

We combined reviewed semantic labels with measured scene geometry. The labels give the AI meaningful subjects and destinations; the geometry gives the camera planner obstacles to check. The AI proposes visual beats and route control points, then deterministic code generates a smooth camera path and validates its clearance and motion limits. Routes that fail those checks are rejected or sent back for revision.

Accomplishments that we're proud of

The fact that it works! In just 36 hours, we built a functional video suite that combines 3D scene creation, cinematic camera controls, AI-assisted direction, drone path planning, and real-time collaboration. We believe Camvas is more than an on-a-whim hackathon prototype, it's an ambitious idea our team has wanted to engineer for a long time and a product we would genuinely use ourselves.

What we learned

We learned that AI is most effective when it works within a well-defined creative system. A model can propose a camera move, but its coordinates need to be checked against the scene and turned into an editable shot. Combining AI proposals with deterministic generation and validation made the results more predictable.

We also learned how different filmmaking and software development can be. Filmmakers think in terms of mood, framing, pacing, and movement, while a 3D engine needs vectors, rotations, curves, and timestamps. Designing Camvas required us to build a bridge between those two ways of thinking.

On the technical side, we learned how to combine mesh colliders with semantic scene graphs, create collision-aware drone paths, help AI understand rooms and landmarks, translate natural-language directions into structured camera actions, and synchronize editable 3D scenes between collaborators in real time.

What's next for Camvas

We want to continue expanding Camvas with a physics engine for more realistic scene interactions, VR and AR support for immersive shot planning, and additional collaboration tools that make it easier for creative teams to build and direct scenes together in real time.

Built With

+ 4 more
Share this project:

Updates

Submission history