We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

SurfaceOS

Turn an ordinary table into an interactive workspace.

Inspiration

Screens give us powerful tools, but they keep our digital work inside a rectangle. We wanted to see what would happen if the surface already in front of us could become the interface. A desk could hold a note, a game, or a question about a physical object, with each tool appearing where it is useful.

SurfaceOS is our exploration of that idea: a workspace projected onto a real surface and controlled with your hands.

What it does

A projector displays a workspace on a table while an overhead camera tracks hand position and gestures. After calibrating the surface, a user can draw a window at a chosen location, select an app, and interact with it through pinches and pointing. Windows can be moved, resized, and placed on separately calibrated surface regions. The same interface also works with a mouse for development and demonstration.

The app collection includes practical tools and games such as a calculator, notes, a to-do list, a timer, and Pong. SurfaceOS can also capture an image of the physical workspace. A thumbs-up opens an Ask AI flow: choose a voice conversation or capture a photo, crop the area of interest, and ask questions about it. The cropped image stays with the conversation, so follow-up questions have visual context.

How we built it

The system has three cooperating parts. A Python camera process uses OpenCV and MediaPipe to track hands, recognize gestures, and send pointer events to the browser over WebSockets. A browser shell manages calibrated surfaces, projected windows, and the shared input state. A reusable app and widget layer supplies the content inside each window.

Calibration maps positions seen by the camera to positions on the projected surface. We use projected markers to estimate that mapping, followed by a fingertip alignment step to bring the cursor closer to the user's actual touch point. The shell converts those positions into window-local input, allowing the same apps to respond to either hand or mouse control.

For Ask AI, the Python process keeps the Gemini API key on the server side. It captures a still from the camera, lets the user crop it, and sends the selected image and question to Gemini. The browser handles speech recognition, spoken responses, and the visible transcript.

Challenges we ran into

Camera coordinates, projected coordinates, and coordinates inside a movable window are three different spaces. Getting a hand movement to land on the intended button required calibration, consistent event contracts, and careful handling of a pinch from press through release. Changes to lighting, camera placement, or projector placement also affect the physical setup, so we built a mouse path that exercises the same workspace without hardware.

Capturing a physical object introduced another challenge: the projector can illuminate the object with our own interface. We added a brief blanking step before taking an Ask AI photo. We also kept camera capture and AI requests out of the critical pointer path so the interface can continue responding while a request runs.

Accomplishments we are proud of

  • A window can be created on a chosen physical area instead of opening at a fixed desktop position.
  • Hand gestures and mouse input can drive the same shell and apps.
  • The project brings together projection calibration, independent app windows, physical capture, and an image-aware voice conversation.
  • Each major layer can be tested separately, which helped three teammates integrate their work during the hackathon.

What we learned

A convincing spatial interface depends as much on reliable input and coordinate mapping as it does on the visual design. We learned to make calibration visible, preserve a mouse fallback, and keep clear boundaries between tracking, window management, and app content. We also learned that a focused interaction with the real surface tells the story better than adding a long list of disconnected apps.

What's next

We want to make content respond more directly to objects on the surface: select an object, ask a question, and place an annotation beside it. We also want more reliable object tracking, saved workspaces, and smoother transitions between multiple physical surfaces. These are future directions; the current prototype centers on calibrated windows, gesture interaction, and capture-based AI questions.

Built With

Share this project:

Updates

Submission history