Skip to content

drv/dotbot_control: add a hardware-free control core, also built as WebAssembly - #38

Merged
geonnave merged 19 commits into
DotBots:mainfrom
geonnave:control-core
Sep 29, 2026
Merged

geonnave merged 19 commits into
DotBots:mainfrom
geonnave:control-core

Conversation

@geonnave

@geonnave geonnave commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

What

Extracts the DotBot app's control glue into a pure, hardware-free core in
drv/dotbot_control/. It receives commands (MOVE_RAW, WHEEL_VELOCITY,
waypoint batches, max speed, CONTROL_MODE), ticks the scheduler, produces telemetry reports and advertisements, and can
seed a robot's pose when its heading is already known. All state lives in a
single struct so the core has no globals and no HAL dependency; the wire
shape is versioned as ABI 2.

Why

A simulator can now run the exact same control code that ships on the robot,
instead of a Python reimplementation that drifts from firmware behaviour.
New firmware can move a robot through this one driver rather than
re-deriving the waypoint/steering/deadman logic. And because it's pure C
with no board dependency, it's host-testable without a device or an
emulator.

The WebAssembly build

make wasm compiles the same core as a WASI reactor with wasi-sdk (pinned
version, fetched by wasm/fetch-wasi-sdk.sh rather than assuming a
preinstalled toolchain) and zero imports - it's a closed function-call
surface a host can drive one fleet of robots through. wasm/check.py
verifies the build has no imports, exports the expected ABI 2 surface, and
that every field offset of the input, output, report and geometry structs
matches the layout table the build exports (layout_offsets(), pinned by
_Static_asserts in C, so a host binding can check offsets at load and not
only sizes). It then replays a waypoint batch through a small Python plant,
hashing the trace against a golden value so a behaviour change in the core is
caught even though there's no host running real physics, and checks that
batched advertisements match one-by-one ones and that fleet_fix_due(), and
the mask each fleet_step() leaves for the next tick, flag exactly the ticks
that read a fix (a simulator hands a robot a fix only then). CI runs this build + check on every push
and PR, the release job depends on it, and a tagged release attaches
dotbot_control.wasm as a release asset so hosts can pin a released build.

Tests

971 passed, 0 failed across five host test binaries (make test): wheel
control, pose estimator, steering, the control core itself, and the fleet API
the wasm build exports, compiled natively. There is no wasm-vs-native
equivalence test: the wasm build is checked by wasm/check.py against its
own golden trace, as above.

Pose estimator: a lost robot re-anchors

While the estimator is LOST, once the wheels have been still for 0.5 s and
3 consistent fixes arrive, it goes back to SEEDING through the same path a
kidnap uses. These returns are counted in their own lost_reseeds counter,
so kidnaps keeps meaning a kidnap detected while TRACKING. Steering's HOLD then spins for a heading and resumes the batch,
so a robot moved while driving recovers instead of holding forever. The
wasm/check.py golden trace is unchanged (its replay has no kidnap). This is
a behaviour change on the robot and needs a bench check before any firmware
release picks it up.

LOST recovery, bench-driven fixes

A bench run on this core showed a robot lifted mid-batch going wrong in the
hand: LOST after 1 s, HOLD braking its wheels, so the "wheels still for 0.5 s
plus 3 consistent fixes" test passed while it was carried slowly (65 mm/s
stays inside the 20 mm chain tolerance). It reseeded in the air, NO_HEADING
spun the wheels, the seed chain rejected every fix, SEEDING never timed out,
and the batch failed with NO_HEADING after 3 s. Three changes:

  • Real settle test. While LOST, the fixes themselves must stay within
    rest_mm (5 mm) of their running mean for the settle time (0.5 s, at least
    kidnap_fixes fixes), with the wheels still as long. The chain then starts
    on that mean. A slow carry keeps it LOST.
  • Free spin. While SEEDING, chain rejections where odometry moved the
    photodiode but the fixes stayed within free_spin_mm (10 mm) add up; at
    acquire_mm of odometry the estimator goes LOST with no pose (fixes only
    chain, never gate against the stale pose). Steering sees LOST and brakes
    instead of spinning in the air. Counted in free_spins.
  • Retry instead of failing. A NO_HEADING spin that times out, or ends with
    the pose going LOST for a reason other than a free spin, goes to HOLD
    (braked, no_heading_rest_ticks = 1 s after a timeout, or until the pose is
    SEEDING again) and spins again, up to no_heading_retries (3) times since
    the heading was last acquired; past that it fails with the existing
    NO_HEADING reason.

New conf fields sit at the ends of their structs; none of the wasm ABI structs
changed, so ABI 2 and the golden trace hash are unchanged (the replay never
loses its pose). Host tests cover a slow braked carry (no reseed), set-down
reseed, free spin to LOST, retry then success, exhaustion, and a closed-loop
lift mid-batch. Needs a bench check like the rest of the LOST path.

Held in the air: one spin, then fail

The next lift test on the bench showed the retries doing harm. Held in the
hand, the robot lost tracking for 3.4 s, the fixes held still enough to pass
the 0.5 s rest test, it reseeded, and each spin was caught as a free spin
within two fixes. But every free spin dropped into a retry, so the wheels
gave 3-5 short fast bursts in the hand before the batch failed. A closed-loop
host run of the same lift gives the drive-on until the pose times out, then
four 0.2-0.3 s spin bursts about 1 s apart, then NO_HEADING.

The approach: a free spin is already the "in the air" signal the bench asked
for. The wheels turned about a spin's worth of the lever arc while the
photodiode fix, which sits 51.5 mm off the axle and on the floor would trace
that arc, stayed within 10 mm. So a pose LOST on a free spin now fails the
batch at once with NO_HEADING, braked, with no retry, whatever steering
state it is in. The steering pose carries a free_spin flag, set by the
control core while the estimator is LOST with no pose after a free spin. The
retries stay for spins that moved the fix but timed out for another reason.
A batch sent while still in that state fails without turning the wheels.
Once the robot is set down, the rest test reseeds it and a new batch acquires
normally.

Detection is not faster than before: it takes two fixes at 10 Hz, about 180 ms
after the spin starts (the first chain step at spin speed is inside the 20 mm
tolerance, and the second step is the rejection that crosses acquire_mm),
and the brake follows in the same tick. The spin's first ~20 ms at full duty
is the wheel loop's normal start and happens on the floor too. That leaves
one short burst in the air, in place of 3-5.

The rest test after a free spin is unchanged, deliberately. With no retries,
a reseed in the hand costs at most one more spin, and only if a new batch
arrives. A longer settle, or more consistent fixes, would delay every
legitimate set-down recovery to cut a case that is now cheap.

The floor-side false-trigger question came from a 25 s run after that lift,
which went untracked mid-leg and failed NO_HEADING in 3.7 s. That signature
is a lift, not a false trigger. The raw fix jumped about 100 mm off the lane,
38 mm sideways (the same offset as the confirmed lift), and then stayed within
15 mm for 3.7 s. On the floor, the four spins that such a fast fail implies
would have moved the fix around a 51.5 mm circle. It did not reproduce in
180 s of untouched laps at 150 mm/s on the same robot and build (4 batches,
all ARRIVED, no loss of tracking after the initial heading). So the free-spin
rule is unchanged. A host test pins the related case: a 0.5 s stall in fixes
mid-spin is no fix, not a still one, so it is not a free spin.

Host tests: in the air the move fails on its first spin with no retry and
brakes, then a new batch after set-down arrives; a batch sent while LOST on a
free spin fails without moving; a spin on the floor acquires; a spin whose
fixes move but never chain times out and still retries; a fix stall
mid-spin still acquires; and a closed-loop lift through the control core
gives the drive-on plus one spin, braked within two fixes. The wasm golden
trace is unchanged (3707b44e...), since its replay never loses its pose.

On the robot (sandbox app, DotBot-firmware #430), I OTA-flashed one robot
with this build and ran it for 180 s of untouched square laps at 150 mm/s: 4
batches, all ARRIVED, no loss of tracking after the initial heading spin. That
matches the two 180 s baselines run on the previous build. The lift test itself
is not re-run yet and still needs a person at the bench.

Behaviour differences from the app's current control glue

  • The deadman timeout is expressed in scheduler ticks here, not wall time,
    since the core has no clock of its own.
  • Commands are applied as they arrive rather than queued; the app keeps its
    own mailbox in front of this core if it still wants to buffer.
  • Bench telemetry/trace hooks the app currently has are not part of the
    core - those stay app-side.
  • The RGB LED command is ignored by the core; it stays in the app.
  • A min-TX-interval change applies immediately in the core
    (db_control_set_min_tx_interval()), where the app applies it from the
    next advertisement.
  • Keepalive has no output flag: the app drives it from
    db_control_fix_due(), as it drives the LH2 solve.

Not in this PR

  • Switching the sandbox app itself to use this core - that's a firmware
    change that needs a bench run, not just host tests.
  • Retiring drv/control_loop - left alone until the app cutover above has
    landed and been validated.

Size

File + -
.github/workflows/build.yml +33 -1
Makefile +32 -2
drv/dotbot_control.h +253 -0
drv/dotbot_control/dotbot_control.c +549 -0
drv/pose_estimator.h +30 -1
drv/pose_estimator/pose_estimator.c +76 -9
drv/steering/steering.c +34 -4
tests/test_dotbot_control.c +616 -0
tests/test_dotbot_control_fleet.c +150 -0
tests/test_pose_estimator.c +122 -0
tests/test_steering.c +160 -1
wasm/check.py +406 -0
wasm/dotbot_control_wasm.c +342 -0
wasm/dotbot_control_wasm.h +60 -0
wasm/fetch-wasi-sdk.sh +72 -0
3 small files: doc/sphinx/drv.md, drv/drv.emProject, drv/steering.h +35 -4
Total, 18 files +2970 -22

… spin

A robot lifted with its wheels braked and carried slowly passed the old
wheels-still test and reseeded in the hand, then spun in the air with the
chain rejecting every fix. Fixes now have to hold still, and a chain the
odometry breaks while the fixes stay put drops back to LOST.

AI-assisted: Claude Opus 5.5
A free spin means the wheels turn while the fixes stay put, which is a
robot held in the air. Retrying the spin from HOLD turned the wheels in
the hand once per retry, so the retries stay for spins that moved the
fix and timed out.

AI-assisted: Claude Opus 5.5
@geonnave
geonnave merged commit 48cc06b into DotBots:main Sep 29, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant