Volume IV staff build brief: prove the kit before teaching it
Audience: Postdoctoral researcher and teaching staff. Status: Supporting hardware notes for the proposed UNO Q spring seminar. Start with the feasibility plan, then use this page for one-axis fallback details if the SO-101 command path cannot be reproduced. The student competency checklist defines what the course must teach; the course operating plan defines the student milestones. Build and measure one complete station before reproducing a class set. The candidate parts list is a purchasing aid, not evidence that the parts work together.
Priority pilot: SO-101 course station
Return four demonstrations in sequence: native LeRobot calibration, teleoperation, and episode replay on the delivered arm; one guarded servo commanded only through Qualcomm Linux → STM32 → servo bus with measured readback and refusal tests; the same STM32-controlled path reproduced across the full arm; and a compact local policy loop on the exact UNO Q plus a separate SmolVLA memory/latency verdict. The feasibility plan lists the evidence for each step. Have another staff member reproduce the command path and physical task trial before ordering a class set or publishing arm-based lab instructions.
One-axis fallback assignment
Deliver a reproducible tabletop station in which a versioned Hugging Face model runs locally on the UNO Q’s Qualcomm QRB2210, influences a bounded physical action, and uses the observed result to choose again. The STM32U585 owns the actuator pins and can refuse a proposal. A person can remove motor power independently. This is the minimum Physical AI scope test: learned model, consequential physical feedback, and delegated actuator authority.
The fallback task is a slow, one-axis visual follower. A fixed camera sees a soft pointer and a hand-placed target within a marked arc. The Qualcomm side estimates the target and proposes one small step; the MCU permits or refuses it, drives the axis, and reports measured position. The next frame shows whether pointer error fell. Move the target to a new position between action cycles and repeat. Start with a stationary target after each placement so the measured camera and inference cadence sets a fair task speed.
Try a compact learned action policy that consumes the camera observation, measured pointer angle, and current target state and proposes a bounded signed step or abstention. A pretrained MobileNet V2 Hub checkpoint can supply a visual encoder, but a sector classifier wired to fixed LEFT, RIGHT, or HOLD commands is only a bring-up check, not the qualifying physical AI loop. Pin model and preprocessing revisions, run the exported policy on the UNO Q CPU, and test held-out target positions and a changed target after the first move. If this path fails, record the failure and test one smaller policy or input resolution against the same frozen trial set. A deterministic rule must run through the same controller and physical protocol for comparison.
Use one UNO Q, a powered USB camera connection, a feedback actuator, an MCU-read SPI angle sensor on the same axis, an independent motor supply with an accessible cutoff, a home reference, a rigid mount, and a low-energy guarded pointer. An AS5048A magnetic angle sensor is a candidate only after its magnet mount and the UNO Q’s 3.3 V SPI pin map are piloted. Compare its reading with the actuator’s feedback and the camera, including disagreement and disconnection. Record the exact board variant, pin map, voltage levels, power path, and parts cost. The UNO Q datasheet documents the QRB2210, STM32U585, and their distinct I/O domains. Do not send motor power through a board I/O pin. Do not order a class set until another staff member can reproduce the first station.
What must work before reporting the fallback kit ready
| Test | Build and experiment | Evidence to return |
|---|---|---|
| 1. Body and safe state | Assemble the axis, mount the camera, isolate motor power, and test boot, reset, cutoff, and power return with model software stopped. Measure travel, stopping, position feedback, and homing. | Wiring and authority diagram, photographs, pin/voltage table, measured mechanical envelope, and observed safe-state trace. |
| 2. Observation and data | Save camera frames with target position, SPI angle, actuator feedback, measured pointer position, observation ID, and trial ID. Include varied target positions, backgrounds, lighting, and partial occlusion. Hold out complete capture sessions. Map the station’s observations and actions to a pinned LeRobot Robot adapter or documented converter. |
Small labeled dataset, LeRobot-compatible episode, collection script, data card, sensor disagreement trace, and a nonlearned image-rule baseline. |
| 2a. Simulation and dynamics | Implement a minimal one-axis simulator from measured travel and response. Implement the same observation/action/trace interface around a one-axis MuJoCo model. Fit a small learned next-state model from training traces. All three must predict held-out hardware motion; the simulators must represent action limits, delayed observations, and a failed move. | Versioned model and simulator configurations, shared adapter, matched rollouts, and one-step and multi-step error for each prediction method. |
| 3. Hugging Face inference on Qualcomm | Choose a compact Hub vision checkpoint or fine-tune one off board. Pin its exact revision and preprocessing. Export if useful, compare host and board outputs on the same frames, and run inference locally on the QRB2210. A compact MobileNet V2 is one candidate backbone; it needs a task-specific localization or action head. Optimum ONNX is an export route to test, not a guarantee of board performance. | Model ID and revision, training/export recipe, output comparison, peak memory, load time, and capture-to-inference latency distribution on the exact board. |
| 4. Proposal and permission | Send a bounded action request from Linux over Bridge. Let the MCU check arm state, task epoch, observation cycle, sequence, deadline, travel, and local feedback before moving. Linux never directly drives the actuator. | Message schema, MCU state diagram, accepted and refused command traces, and observed behavior after Linux stops. |
| 5. Closed loop | Place the target at randomized positions, infer, request one move, measure actual motion, capture again, and correct or abstain. Move the target unexpectedly after apparent convergence. Compare model, nonlearned rule, and no-action runs from matched starts. | Linked frame → model → proposal → permission → measured motion → next-frame traces; initial and final error, action count, completion time, and failures for each trial. |
| 6. Faults and replication | Occlude or disconnect the camera, delay inference, replay an old proposal, interrupt Linux, cut motor power, and restart. Have a second person build a second station from the instructions. | Driver/position evidence that prohibited motion does not occur, unresolved-failure list, frozen board image, bill of materials with availability, assembly guide, spare-parts plan, and independent reproduction log. |
Freeze the trial protocol and model before the final comparison. Use multiple randomized starts and report every trial, including failed ones. The initial acceptance target is that the pointer enters the target’s visible width within five bounded moves in at least eight of ten held-out stationary-target trials, then responds to a new placement rather than replaying the old command. Every tested stale, duplicate, out-of-range, disarmed, and power-cut case must produce no new motion. Report the actual latency distribution and mechanical limits; revise the target width, action limit, or timing budget only with measured justification. If a model cannot beat the rule on a meaningful physical condition, change the visual task or explain why the learned component has no educational role.
The final report is a demonstration plus evidence packet: a short uncut video of normal and perturbed trials, source and frozen image, model and dataset revisions, both simulators and their hardware comparison, wiring/assembly instructions, raw traces, summary plots, fault results, per-station cost, and a one-page recommendation for class replication. Report a failure early rather than silently substituting a laptop, cloud inference, or prepackaged model runtime for the agreed on-board path.
In parallel with the class-kit work, return a SmolVLA-on-UNO-Q feasibility report. Fine-tune or obtain a task-matched LeRobot checkpoint and verify it first on a workstation and a shared SO-101 or simulation task. On the exact UNO Q variant, report whether the model loads, peak memory, cold start, per-chunk latency, camera-to-action age, sustained behavior, and host-versus-board output agreement. Attempt a slow physical rollout only after the action path and cutoff are tested. A failure to meet the measured deadline is a valid result; it must identify the bottleneck and a defensible smaller policy path. Do not claim that running an unrelated vision classifier demonstrates VLA deployment.
How this becomes a course rather than one demo
The first six weekly labs each own one reusable component. Andrea should complete these experiments personally and turn each test above into a student handout with a starting state, prediction, deliberate failure, measured result, and artifact that the next lab consumes. The lab sequence expands the experiments.
| Week | Component students master | Artifact carried forward |
|---|---|---|
| 1 | Physical boundary, power, and safe state | Authority and power diagram |
| 2 | One-axis plant, SPI and actuator feedback, calibration, and a simple fitted simulator | Sensor contract, motion envelope, homing trace, and first predicted rollout |
| 3 | MCU permission and Linux-to-MCU protocol | Proposal schema and refusal traces |
| 4 | Physical data, labels, and a rule baseline | Dataset and linked episode log |
| 5 | Hub model revision, preprocessing, on-board inference | Reproducible model artifact and latency trace |
| 6 | First model-dependent physical feedback loop | Feasibility demonstration and full trace |
Teams choose a capstone direction early, but weeks 7–14 are where they cut vertically through the whole stack. They reuse the station and evidence format while adding state freshness, a learned dynamics model, a MuJoCo model of the same axis, one-step planning, a simulator-to-hardware comparison, an action-chunking exercise, model/latency comparison, fault handling, human interruption, repeatable evaluation, and an unfamiliar final trial. The capstone can be the follower, active visual inspection with a swappable retained-object stage, or verified routing on a retained carrier. An attachment becomes a student option only after staff have built and qualified it. Every capstone must show a learned output changing an admissible action and a later observation changing the next decision.
The two simulators have distinct teaching jobs. The small Python model exposes every assumption and makes parameter sweeps cheap. MuJoCo adds a modeled actuator, joint, geometry, and camera for a richer physical hypothesis. Give both the same reset → observe → propose → permit → step → log adapter and the same units, limits, task epoch, and trace fields as the real station. The simulator may advance virtual time quickly; the real UNO Q must use measured wall-clock deadlines. Students should compare predicted and actual position, overshoot, observation timing, and failure recovery from matched starts. Gymnasium’s environment API can provide a common reset/step wrapper. Its Pendulum or CartPole exercise is a useful optional control warm-up, but it is not a validated model of this kit.
SO-101 course feasibility
Inventory the delivered Seeed SO-101 Pro assembled kit: count leader and follower arms, cameras, controllers, cables, power supplies, and spare servos rather than assuming a packout from the product name. Follow the official LeRobot SO-101 setup to identify the motor-bus USB ports, configure IDs, calibrate joint positions, and record a short teleoperated episode if the required leader or other input device is present. Use those episodes for a staff-prepared ACT action-chunking lab; provision workstation compute and an offline replay path for students without physical arm time.
Draw the actual command path before connecting the arm to a Qualcomm board. Seeed documents a UART servo bus behind a USB controller. The stock LeRobot host can command motors through that controller, so the UNO Q MCU cannot claim per-joint veto merely because it is connected to the same system. Test an STM32-to-servo command route first on one guarded, unloaded servo, then on the full arm, with no live direct host-to-servo bypass. Measure refusal, readback, cutoff, reset, and command timing. A motor-power interlock proves emergency cutoff, not independent permission for each action. Return a measured recommendation for the arm-centered curriculum and the local SmolVLA trial; retain the one-axis reference station if the STM32 route cannot be reproduced.
VLA and pendulum decision
The reference follower uses a compact vision model; it is not a VLM or VLA claim. This is deliberate: the first engineering question is whether the UNO Q can close a measured physical loop with a Hugging Face artifact.
Do not substitute a separate VLM image-description demo for the VLA experiment. The camera supplies evidence for a physical policy; the test is whether language and vision change an admissible action and a later measured outcome.
SmolVLA is a genuine Hugging Face vision-language-action model: it consumes images, robot state, and an optional instruction and generates action chunks. Hugging Face recommends task-specific demonstrations and fine-tuning for a new robot. Its published affordable-robot examples do not establish performance on the UNO Q or compatibility with this one-axis rig. Attempt local deployment as the course’s explicit research target, then use the measured verdict to decide whether a VLA can become a required student lab. The first class kit still needs its proven compact local policy when that result is pending or negative.
An inverted pendulum is a good control experiment but a poor first VLM or VLA assignment. For an idealized point mass 20 cm from the pivot, the near-upright instability time scale is about sqrt(0.20/9.81) = 0.14 s; the fast stabilizing loop belongs on the MCU and needs angle feedback, a suitable actuator/driver, and a contained frame. A slow language model could choose a goal or mode, but it would not be the balancing controller. Consider a guarded pendulum attachment only after the reference station works, and require the learned component to change a measured physical outcome rather than decorate a classical controller.