Skip to content

bsp: carry LH2 homographies as float32 and fix the LH2 decoder and timer re-arm - #31

Merged
geonnave merged 15 commits into
DotBots:mainfrom
geonnave:lh2-phase3
Sep 23, 2026
Merged

geonnave merged 15 commits into
DotBots:mainfrom
geonnave:lh2-phase3

Conversation

@geonnave

@geonnave geonnave commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Part of the LH2 calibration rework that moves the lighthouse homography from a scaled integer to float32 and gives a calibration an identity (site name, id, validity rectangle). This PR is the DotBot-libs piece: the float32 store, and six decoder and timer faults found while validating it on a robot and in review.

Float32 homographies. The calibration packet (protocol_lh2_homography_t) and db_lh2_store_homography() carry each matrix element as IEEE-754 float32 in millimetres, in place of an int32 holding the value times 1000. The int32 form lost precision on the perspective row and needed a divide on every store; float32 is what the solve uses anyway. The store now ignores an out-of-range basestation index instead of writing past its sixteen-slot array. This is a wire change: the host packer and every firmware linking this move together.

Advertised calibrated byte as a bitmask. db_protocol_dotbot_advertizement_to_buffer() took bool calibrated, which C collapses to 1, so a firmware bitmask of 0x05 (stations 0 and 2) reached the host as 0x01. With more than one station a host could never see the bot as fully calibrated. The parameter is now uint8_t. One station is unaffected, which is why it went unnoticed.

Counts a false polynomial match produces are dropped at the store. A count beyond one rotation (the station's period over 8), or equal to the other sweep's count, cannot come from a real sweep. Discarding it before it is stored keeps it out of both the raw-count readers and the position solve; on the bench a phantom station had reached a calibration capture this way.

Periodic timer re-arm. A periodic RTC compare serviced more than one period late got a compare value the counter had already passed, so the tick stopped until the 24-bit counter wrapped, 512 s later. On a sandbox robot that stops the application's watchdog feed and the bootloader's position publishing; it is what hung a control app into a watchdog reset on the bench. A late compare is now re-armed from the current counter.

zccs_1 overflow in _demodulate_light. The zero-crossing buffer had 128 entries, indexed by a uint8_t that passes 128 on a busy capture, writing past it on the stack. In Debug builds it corrupted the demodulation, so at some static poses a still robot decoded only 0-61% of captures (the long-standing "no position after a flash" symptom); in Release it clobbered a saved register and HardFaulted the swarmit bootloader on its first capture. That code is unchanged from main, so main's Release bootloader is exposed too. The buffer is now sized for every uint8_t index.

64-bit popcount. The polynomial match took __builtin_popcount of a 64-bit difference, which counts only the low 32 bits, so a polynomial disagreeing only in the upper bits scored as an exact match. It is now __builtin_popcountll.

_demodulate_light stays inside its chip array. The zero-crossing buffer is only partly written by a capture, and the rest was uninitialised stack; the demodulator derived chips from it, and on some of those values wrote one byte below the chip array or read and wrote past its end. The untouched tail is now initialised so it thresholds to zero chips, every scan is bounded by the array, and the carry between the two passes is reset. Output on a clean capture is unchanged.

Re-arm margin. The late re-arm compared the new compare value against a counter read taken before the write, so a tick in between could still leave it one tick ahead, which the RTC does not guarantee to match. The margin is now 4 ticks.

Lines changed

File + -
6 small files: bsp/lh2.h, bsp/nrf/lh2_decoder.c, bsp/nrf/lh2_default.c, bsp/nrf/timer.c, drv/protocol.h, drv/protocol/protocol.c +58 -20
Total, 6 files +58 -20

Merge order

This is the first of four PRs that go in this order:

  1. DotBot-libs: this PR
  2. swarmit: device: carry float32 LH2 calibrations with a site, and expose raw counts swarmit#165
  3. DotBot-firmware: apps-sandbox: move to the float32 bootloader, add dotbot-next and a button calibration app DotBot-firmware#424
  4. PyDotBot: dotbot: push float32 LH2 calibrations with a site gate, reframe, and capture from the button PyDotBot#296

swarmit and DotBot-firmware bump their dotbot-libs submodule to this branch's tip, so the submodule pointers are re-pointed at the merge commit once this lands. There is no other gate: the project is in beta and these can merge as soon as they are reviewed.

Validation

Built through swarmit (bootloader and network core, dotbot-v3, Debug and Release) and DotBot-firmware (sandbox apps and the bare-metal dotbot app, sandbox-dotbot-v3 Release), zero warnings. On a DotBot v3 with one lighthouse, running the swarmit bootloader built from the companion PR: a live position within 1-3 s of each of 12 cable flashes with a calibration; 5-minute soaks in the bootloader and under two sandbox apps with no reset at about 10 solves/s; and a 4.4 h A/B against main on the same robot, gateway and calibration with no unexpected reset and 98-100% status delivery on both. The ISR cost was measured with the DWT cycle counter: 1.4 ms mean, 2.4 ms max in Release, 4.3 ms max in Debug. Re-run after the code-review fixes, 2026-09-22 10:17-11:38, on one tethered DotBot v3 with one lighthouse and the gateway on schedule tiny, Debug bootloader and network core from swarmit #165 cable-flashed with a calibration: three fresh cable flashes each reported a live position 1.5-1.6 s after the flash tool returned, with no button press. In 15-minute windows (bootloader idle, dotbot idle, dotbot with move_raw bursts, dotbot-next idle, calibrate with simulated presses) there was no reset, no crash report, no absence from swarm status, no NODE_LEFT and no FCS error on the gateway link; STATUS delivery was 98.9-99.7% and adverts 97.7-99.6% of the 2/s the apps send, with no advert gap over 1.13 s; the advertised position matched STATUS to 1 mm and no invalid position was reported. The decoder matched the camera's displacement at four spots nudged with move_raw (a constant offset within 1.5 mm).

Re-run after the final review round, 2026-09-22 13:55-14:19, on the same tethered robot and gateway (schedule tiny): two cable flashes of the Debug bootloader and network core with a calibration each reported a fresh boot 2-3 s after the flash tool returned, device info v2 carrying the file's site and id. A dotbot build with a receive counter got 400 of 400 move_raw packets exactly once (100 at 4/s, then 300 at about 20/s), with STATUS flowing and no mari rejoin. A 15-minute window on dotbot had no reset and no NODE_LEFT, swarm status had updated within the last second at each of 58 samples, and 99.6% of the 2/s adverts arrived (largest gap 1.27 s). dotbot-next stopped its motors 539-720 ms after the last command (a 519 ms timeout checked every 200 ms) and advertised at 2.00/s. swarmit bootloader (Debug, Release) and network core (Debug), every sandbox app and the bare dotbot app build at this tip with zero warnings.

One robot, one station: multi-station behaviour, which the bitmask fix is for, has not been run on hardware.

The homography travels as IEEE-754 float32 in millimetres instead of an
int32 holding the value times 1e3, so the host packer and every firmware
that links this must move together. An out-of-range basestation index
is now ignored by the store instead of writing past the array.

AI-assisted: Claude Opus 5
A count beyond a rotation (the station's period over 8) or equal to the
other sweep's cannot come from a real sweep, so it is discarded before
it is stored and never reaches a raw-count reader or the position solve.

AI-assisted: Claude Opus 5
The byte was a bool, which collapsed a per-basestation bitmask to 1, so
a host could never see more than basestation 0 as calibrated.

AI-assisted: Claude Opus 5
A periodic channel serviced more than one period late got a compare
value the counter had already passed, so the tick stopped until the
24-bit wrap, 512 s later. A sandbox app then stops feeding the
watchdog, and the bootloader stops publishing positions.

AI-assisted: Claude Opus 5
A capture with more than 128 crossings wrote past zccs_1 on the stack.
In the Debug build that corrupted the demodulation, so at some poses a
still robot never matched a polynomial; in the Release build it
clobbered a saved register and the swarmit bootloader HardFaulted.

AI-assisted: Claude Opus 5
__builtin_popcount takes an unsigned int, so only the low 32 bits of the
64-bit difference were counted and a polynomial disagreeing in the upper
bits scored as an exact match.

AI-assisted: Claude Opus 5
The first pass read chips1[128] on every call and could write past the
array when that byte read 0xFF; the second pass started with the ones
counter the first pass left behind, which wrote below chips1[0] when
the capture ended on an odd run of ones; and the tail past the last
crossing was thresholded from uninitialised stack. All of it runs on
the swarmit bootloader's stack on every LH2 capture.

AI-assisted: Claude Fable 5.1
With bitwise & every operand is evaluated, so chips1[jj + 1] was still
loaded when jj had just advanced to 127, one byte past the array; only
the branch was suppressed. && stops at the index check.

AI-assisted: Claude Fable 5.1
The pass never counts chip 0, so a run of ones ending at jj - 1 started
at index 1 or later and jj - ones_counter - 1 is never negative. The
reset of ones_counter before the pass is what keeps it so.

AI-assisted: Claude Fable 5.1
@geonnave
geonnave merged commit 9646cf8 into DotBots:main Sep 23, 2026
16 checks passed
@geonnave
geonnave deleted the lh2-phase3 branch September 23, 2026 08:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant