Skip to content

Survive a serial port that is absent at startup or disappears while running - #85

Open
ssharma0704 wants to merge 3 commits into
flynneva:mainfrom
ssharma0704:fix-startup-race
Open

ssharma0704 wants to merge 3 commits into
flynneva:mainfrom
ssharma0704:fix-startup-race

Conversation

@ssharma0704

@ssharma0704 ssharma0704 commented Sep 28, 2026 •

Copy link
Copy Markdown

The serial device can be briefly absent after boot or a USB re-enumeration, and it can also disappear while the node is running. This PR handles both, and fixes two faults that made the first attempt at it ineffective. All of it is measured on hardware (FT232H bridge, UART, Jetson, Humble); details under each commit.

connect() and main() retry instead of exiting (first commit)

connect() called sys.exit(1) on the first failed open, and main() then referenced node before assignment in its except and finally blocks, masking the real error with an UnboundLocalError. Under a launch file with on_exit=Shutdown() this took the whole stack down and the service restart-looped it.

  • UART connect: retry the serial open for up to 30 s, then raise ConnectionError instead of sys.exit(1).
  • main(): retry node.setup() up to 6 times, 2 s apart, then raise.
  • main(): initialise node to None and guard the except/finally cleanup so a failed startup reports its real cause.

Make that retry actually reach its later attempts (second commit)

Testing the above on hardware showed the retry loop could never succeed, for two independent reasons:

  • setup() began by constructing NodeParameters, which declares all 25 parameters, so attempt 2 onwards always failed with Parameter(s) already declared — and that error also pushed the real cause off the top of the log.
  • SensorService.configure() called sys.exit(1) on the first failed chip-ID read. SystemExit derives from BaseException, so except Exception in the retry loop never saw it and the process left at attempt 1, without retrying at all.

Parameters are now declared once in __init__, configure() raises ConnectionError and lets the caller decide, and the connector and SensorService are constructed once so a later attempt re-runs configure() alone rather than creating a second set of publishers and services.

This matters most straight after a USB re-enumeration: the BNO055 needs about a second after power is restored, and a node that opened the port too early used to die instead of waiting two seconds and asking again.

Measured: against a port that never answers, the node now makes 6 real attempts 2 s apart, each reporting did not answer the chip-ID read: Unexpected length of READ-request response: 0, then exits 1. Before, it exited at attempt 1.

Exit after a run of failed reads instead of publishing nothing (third commit)

This one is a different failure and is separable, so please say if you would rather it came as its own PR.

read_data() caught every exception, logged a warning and returned, so a node whose sensor had gone away stayed alive at the query rate forever, publishing nothing. Nothing escalated, and a supervisor configured to restart on exit cannot help a process that never exits.

It now counts consecutive failed reads and exits 1 after roughly 3 s of them, derived from data_query_frequency so no new parameter is needed. BusOverRunException and ZeroDivisionError keep their early return and deliberately do not count, because "data fusion not ready" is normal on a cold start and must not be mistaken for a missing sensor.

Measured by unbinding ftdi_sio under the running node: reads failed with TransmissionException, the node exited after 3.0 s, and its supervisor had it publishing again 11 s after the port disappeared. Before this commit the topic stayed silent indefinitely. 60 s of normal running produces no spurious exit, so the counter resets correctly on good reads.

After boot or a USB re-enumeration the serial device can be briefly
absent. connect() called sys.exit(1) on the first failed open, and
main() then referenced node before assignment in the except and
finally blocks, masking the real error with an UnboundLocalError.
Under a launch file with on_exit=Shutdown() this took the whole
stack down and systemd restart-looped it.

- UART connect: retry the serial open for up to 30 s, then raise
  ConnectionError instead of calling sys.exit(1).
- main(): retry node.setup() up to 6 times, 2 s apart, then raise.
- main(): initialise node to None and guard the except/finally
  cleanup so a failed startup reports its real cause.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thankyou for creating your first PR! Contributions like yours is what it is all about. Keep it up!

The retry loop added in e8b0a5e could never work. Two reasons, both
measured on hardware:

- setup() began by constructing NodeParameters, which declares all 25
  parameters, so attempt 2 onwards always failed with "Parameter(s)
  already declared" and the real cause scrolled off above five
  misleading errors. Parameters are now declared once in __init__.
- SensorService.configure() called sys.exit(1) on the first failed
  chip-ID read. SystemExit is a BaseException, so "except Exception" in
  the retry loop never saw it and the process left at attempt 1. It now
  raises ConnectionError and lets the caller decide.

A retry must also not build a second set of publishers and services, so
the connector and SensorService are created once and a later attempt
re-runs configure() alone.

This matters most right after the port reset that rover-bno055.service
runs before every start: the BNO055 needs about a second after power is
restored, and a node that opens the port too early used to die instead
of waiting two seconds and asking again.

Verified on a MITI: against a port that never answers the node now makes
6 real attempts 2 s apart, each reporting "did not answer the chip-ID
read", then exits 1 for the service to restart.
read_data() caught every exception, logged a warning and returned, so a
node whose sensor had gone away stayed alive at the query rate forever,
publishing nothing. Nothing escalated, and because Restart=always only
acts when a process exits, the service could not help either: an
external watchdog was the only thing that noticed, 30 to 75 s later.

It now counts consecutive failed reads and exits 1 after about 3 s of
them, derived from data_query_frequency so no new parameter is needed.
BusOverRunException and ZeroDivisionError keep their early return and do
not count, because "data fusion not ready" is normal on a cold start and
must not be mistaken for a missing sensor.

Verified on a MITI by unbinding ftdi_sio under the running node: reads
failed with TransmissionException, the node exited at 300/300 after 3.0 s,
and the service restarted it, reset the port and had the IMU publishing
again 11 s after the port disappeared. The driver was untouched
throughout.
@ssharma0704 ssharma0704 changed the title Retry the UART connect and sensor setup instead of exiting at startup Survive a serial port that is absent at startup or disappears while running Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant