Skip to content

Repository files navigation

🦾 Dum-E — SO-101 Vision-Language-Action Robotics

Teach a low-cost SO-101 arm to sort cables by color through behavioral cloning — no hand-coded perception or motion planning.

Built on Hugging Face LeRobot · benchmarking SmolVLA vs ACT · shipped as the dume CLI + TUI.

PyPI version Python Downloads License: MIT Built with LeRobot Powered by SmolVLA TUI: Textual HF Model HF Dataset

pip install dume      # then run:
dume                  # 🚀 the icon-rich SO-101 cockpit (TUI)

Note

Dum-E is the academic project (cable-sorting with a Vision-Language-Action policy); dume is the open-source tool we built to operate the robot. One install, two modes: the icon-rich Textual TUI (dume) and a fully scriptable CLI (dume run <subcommand>).


Table of Contents


Overview

Problem. Manually coding perception, grasping, and color-dependent logic is rigid and hard to scale. Classic imitation-learning policies on low-cost hardware often fail to recover from small execution errors.

Hypothesis. Behavioral cloning lets the SO-101 learn complex sorting without explicit motion programming, and a Vision-Language-Action policy (SmolVLA) — conditioned on a natural-language task — generalizes to new object positions better than a vision-only baseline (ACT).

Objective. Teach the SO-101 to sort cables by color (black, green, red) via behavioral cloning from teleoperated demonstrations, and test whether a single SmolVLA policy handles all three colors, benchmarked against ACT.

Two-arm SO-101 setup sorting colored cables into boxes
Leader + follower SO-101 with the cable-sorting task environment.

Quick install

dume is published on PyPI and installs a single command, dume, with two modes — the TUI and the scriptable dume run CLI.

pip install dume          # or: pipx install dume  /  uv tool install dume

# TUI mode
dume                      # open the cockpit
dume teleop               # jump straight to a view
dume --ascii              # plain mode (no Nerd Font / inline images)

# CLI mode (scriptable)
dume run check            # connectivity health check (arms + cameras)
dume run --help           # full command list

Important

The PyPI wheel pulls lerobot[feetech] (Torch, DepthAI, …) and assumes you have the SO-101 arms and cameras attached. For training you need a CUDA GPU (or Apple mps). To hack on the project itself, use install from source.

The dume TUI & dume run CLI

One tool, two modes, one robot:

Mode Command Purpose
TUI dume Icon-rich Textual cockpit: live connection status, arm/calibration health, dataset & model inventory, quick actions, inline camera test, and form-driven views (teleop, …).
CLI dume run <subcommand> Scriptable, one-shot operations: ports, motor setup, calibration, teleop, dataset recording, training, eval, push/pull to the HF Hub.

Both modes are thin front-ends over the same so101_cli modules. The TUI never reimplements robot logic — it reuses those modules by import for read-only display, and for heavy/interactive ops it suspends and shells out to dume run.

DUM-E home cockpit
Home cockpit — connection, arms+calibration, datasets, models, quick actions, inline image test.
DUM-E teleop view
Teleop view — editable fields recompose a live $ dume run teleop … preview, then launch with the full TTY.

A condensed dume run command map (full reference in so101_cli/COMMANDS.md):

dume run find-ports                                  # identify which USB port is which arm
dume run setup-motors follower|leader                # assign motor IDs 1..6
dume run check                                       # both arms + both cameras health check
dume run teleop [--rate 30] [--record]               # live leader → follower (+ optional recording)
dume run record-dataset ... / dume run record-batch  # capture VLA demos → HF Hub
dume run train [--type smolvla|act] [--device cuda]  # wrap lerobot-train
dume run eval --color red|green|black [--n N]        # roll a trained policy on the follower
dume run eval-viz --color red                        # eval + ResNet activation heatmap (rerun)
dume run push|pull dataset|model ...                 # sync with the HF Hub
dume run dataset-stats                               # per-color episode/frame balance

Hardware

# Component Role
2× SO-101 arm (5 joints + gripper, 6 DOF) — Feetech STS3215 servos Leader (teleop handle) + follower (executes)
1× Intel RealSense D435 front camera (top-down task view)
1× OAK-D Pro lateral camera (side view, DepthAI)
1× αFusion / Fusion XPark GPU machine for training
— 2× Waveshare USB-to-TTL adapters Two stable /dev/ttyACM* serial buses
— Cardboard box, wood, LEGOs, colored cables Task environment
  • Follower: 6× STS3215 with 1/345 reduction.
  • Leader: 6× STS3215 with mixed gearing (1/191, 1/345, 1/147 per joint) so it moves with little force.

See docs/armado.md for the build guide.

The task

A single natural-language instruction per episode — "pick the <color> cable and place it in the <color> box" — across black / green / red cables placed in varying positions. The follower receives target joint positions every tick at 30 fps.

Method & architecture

Both policies are trained by supervised behavioral cloning: the model imitates expert demonstrations captured via leader→follower teleoperation, minimizing the gap between its predicted actions and the expert's.

SmolVLA pipeline
SmolVLA: multimodal vision + language + state → flow-matching action expert.

SmolVLA (Vision-Language-Action)

  • SigLIP vision encoder embeds the multi-camera feeds; SmolLM embeds the language task; a linear projection tokenizes proprioception.
  • Early-fusion concatenation merges visual + text + state tokens, processed by a deep self-attention VLM backbone.
  • An action expert with interleaved cross/self-attention and flow matching denoises a random vector into motor commands, solved by a numerical ODE solver.

Learning techniques

  • Action chunking — the model emits a short burst of future joint commands per inference.
  • Fine-tuning — initialize from pretrained SmolVLA weights and train on the full dataset.
  • Linear probing — freeze the pretrained backbone (SmolVLM2-500M) to cheaply adapt to the new gripper geometry.
  • Flow matching — gradually transform random noise into the correct action.

ACT baseline — a vision-only Action-Chunking Transformer, used as the benchmark to isolate the value of language conditioning.

Dataset

  • Collected by teleoperation and recorded with LeRobot — one task string per episode.
  • Each frame synchronizes, from a single capture loop, the 6 follower joint states + front camera + lateral camera; the 6 target joint positions are sent to the follower every tick at 30 fps.
  • Stored in the LeRobot v3 layout on the Hugging Face Hub: episodes batched into shared Parquet files, one MP4 per camera.
  • Balanced by construction — episodes are recorded in randomized batches of 3 (one per color) so classes stay even as the dataset grows.
  • Two published datasets: the main 232-episode so101_terminal_sort, plus a smaller so101_terminal_sort_ext of fresh demos recorded after we extended the gripper (see Experimental design). The extended fingers changed the arm's kinematics, so old grasp heights no longer matched — we recorded a clean dataset on the new geometry and fine-tuned from it, deliberately never mixing old- and new-gripper demos in one dataset.
dume run dataset-stats output: 232 balanced episodes
dume run dataset-stats — the main dataset: 232 episodes, ~57.5 min @ 30 fps, balanced across black/green/red.

Models & datasets

All artifacts are public on the Hugging Face Hub 🤗.

Policies

Model Type Role
smolvla_terminal_sort_ft SmolVLA (fine-tuned) Current best — fine-tuned on the extended-gripper demos
smolvla_terminal_sort SmolVLA Base VLA trained on the full 232-episode dataset
act_terminal_sort ACT Vision-only baseline for the SmolVLA-vs-ACT benchmark

Datasets

Dataset Role
so101_terminal_sort Main balanced cable-sorting dataset (232 episodes, 3 colors)
so101_terminal_sort_ext Fresh demos on the extended gripper, used for the fine-tune
dume run pull model  --repo-id armandomm09/smolvla_terminal_sort_ft         # grab the policy
dume run pull dataset --repo-id armandomm09/so101_terminal_sort             # grab the dataset
dume run eval --color red --policy armandomm09/smolvla_terminal_sort_ft     # roll it on the follower

Experimental design

Conditions
Training Cables at known positions/orientations, consistent lighting, demos by teleoperation.
Evaluation New cable locations and rotations not seen during recording, with small lighting variations — on the real follower, no leader needed.

A real-world adaptation surfaced mid-project: the original gripper had insufficient contact surface, so we extended it to increase grasp area. The geometry change meant recording a small set of fresh demos to fine-tune the existing model.

Extended gripper (pink 3D-printed tips)
Extended gripper tips — more contact surface, then a fine-tune on fresh demos.

Results & expectations

  • Trained on 232 balanced episodes (~57.5 min @ 30 fps).
  • Target: the fine-tuned VLA sorts black/green/red cables into the correct boxes with a >70% success rate across evaluation episodes.
  • The extended gripper is expected to improve grasp reliability; the model should tolerate small lighting variations and new cable rotations.
  • Validated by eval runs on the follower SO-101 alone (no teleoperation).

Feasibility & limitations. Low-cost accessible hardware + behavioral cloning is practical; transfer learning and cheap fine-tuning make the VLA tractable. Limits: a human-collected dataset, BC can drift into states not covered by demonstrations, a narrow domain (specific cables), and slow real-world testing.

Install from source (development)

LeRobot is vendored as an editable path dependency at ./lerobot (managed with uv). setup.sh bootstraps everything.

git clone https://github.com/armandomm09/dume
cd dume
bash setup.sh                 # installs uv, ensures Python 3.12, clones lerobot/, uv sync
source .venv/bin/activate
cp .env.example .env          # fill HUGGINGFACE_TOKEN and HF_USER

USB permissions (Linux, once):

sudo usermod -aG dialout $USER   # then log out and back in

Stable port names (optional): edit udev/99-so101.rules with your adapter serials, then:

sudo cp udev/99-so101.rules /etc/udev/rules.d/
sudo udevadm control --reload-rules && sudo udevadm trigger

In a checkout, the repo wrapper ./dume runs the same code as the installed command (./dume run <subcommand> for the CLI).

End-to-end workflow

# 1. Detect ports, set up motors, calibrate (once per arm)
dume run find-ports
dume run setup-motors follower && dume run setup-motors leader
bash scripts/02_calibrate.sh both

# 2. Teleoperate and record demonstrations
dume run teleop --rate 30
dume run record-batch --batches 10 --push          # balanced batches of 3 → HF Hub

# 3. Train (GPU)
dume run train --type smolvla --device cuda         # default → <HF_USER>/smolvla_terminal_sort

# 4. Evaluate on the real arm (no leader)
dume run eval --color red --n 5
dume run eval-viz --color red                        # + ResNet activation heatmap

Project layout

dume/
├── dume                     # wrapper: `./dume` (TUI) and `./dume run <subcommand>` (CLI)
├── setup.sh                 # bootstrap: uv + clone lerobot + uv sync
├── pyproject.toml           # PyPI metadata; lerobot as editable path source (dev only)
├── configs/                 # per-arm port + id (flat YAML)
├── calibrations/            # versioned calibration JSONs (symlinked into LeRobot cache)
├── scripts/                 # numbered bash wrappers over the lerobot-* CLIs
├── so101_cli/               # the package (distribution name: dume)
│   ├── cli.py               # `dume run` argparse tree
│   ├── poses.py motion.py …  # smooth moves, trajectories, rerun viz, cameras
│   ├── depthaicamera.py     # project-owned LeRobot camera type for the OAK-D
│   └── dume/                # the Textual TUI (app, engine/, screens/, widgets/)
├── udev/                    # stable /dev/so101_* rules
└── docs/                    # build, calibration, troubleshooting, dume design

The editable LeRobot clone lives as a sibling: ./lerobot/. Architecture details for contributors are in CLAUDE.md and docs/dume.md.

Roadmap

  • Phase 1 — dume foundation + Home cockpit
  • Phase 2 — Teleop view (form → live command preview → launch)
  • Phase 3 — Record / Eval / Train forms + dataset & model browsers
  • Rigid, enclosed rig with constant lighting to remove ambient-light variables
  • Simulation pipeline — a digital twin in MuJoCo / Isaac Sim

Design note: No ROS2 — deliberate. LeRobot already covers teleop/recording/training/eval, and ROS2 would mean re-implementing drivers that already exist. There is no CI/test suite; "tests" are physical executions on the hardware.

Team

Developed for a Semester 6 implementation course at Tecnológico de Monterrey.

Name ID
Pablo Armando Mac Beth Milian A01735082
José Luis Domínguez Morales A01285873
Paola Llamas Hernández A01178479
Jocelyn Anahid Velarde Barrón A01285780
Héctor Eduardo Tovar Mendoza A00840308

Acknowledgments

License

Released under the MIT License.

About

Setup y operacion del brazo SO-101 (leader + follower) usando LeRobot framework. Sin ROS2.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages