fix: bolt optimization replacing get_all_nodes with iterator - #1016
fix: bolt optimization replacing get_all_nodes with iterator#1016n24q02m wants to merge 1 commit into
Conversation
**What**: Replaced the usage of `get_all_nodes` and `get_all_edges` in `exporter.py` with newly introduced `iter_raw_nodes` and `iter_raw_edges` methods in `GraphStore`. The new methods return raw `sqlite3.Row` iterators, bypassing the conversion to `GraphNode` and `GraphEdge` Python dataclasses. Also adapted the export functions (GraphML, JSON-LD, DOT, and Cypher) to access fields using dictionary-style lookups rather than object attributes. **Why**: During full graph exports, reading all nodes and edges via `get_all_nodes()` materializes tens of thousands of `GraphNode` and `GraphEdge` Python objects into memory simultaneously. Fetching raw database rows directly drastically cuts down on the peak memory overhead and avoids the computational cost of instantiating numerous wrapper objects. **Impact**: Substantially reduces memory utilization and parsing overhead when exporting large repositories, especially to formats like JSON-LD and GraphML. This aligns with Bolt's performance rule around direct cursor iteration. **Measurement**: Tested via an in-memory database export script to confirm that the changes correctly handle raw rows and no data is lost. Furthermore, `uv run pytest tests/test_exporter.py` validates that functionality is preserved. Co-authored-by: n24q02m <135627235+n24q02m@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
|
Closing: competing variant of the raw-row export iterator; #1015 landed the same design (iter over rows instead of GraphNode/GraphEdge materialization) with fuller coverage. |
Understood. Acknowledging that this work is now obsolete and stopping work on this task. |
What:
Replaced the usage of
get_all_nodesandget_all_edgesinexporter.pywith newly introducediter_raw_nodesanditer_raw_edgesmethods inGraphStore. The new methods return rawsqlite3.Rowiterators, bypassing the conversion toGraphNodeandGraphEdgePython dataclasses. Also adapted the export functions (GraphML, JSON-LD, DOT, and Cypher) to access fields using dictionary-style lookups rather than object attributes.Why:
During full graph exports, reading all nodes and edges via
get_all_nodes()materializes tens of thousands ofGraphNodeandGraphEdgePython objects into memory simultaneously. Fetching raw database rows directly drastically cuts down on the peak memory overhead and avoids the computational cost of instantiating numerous wrapper objects.Impact:
Substantially reduces memory utilization and parsing overhead when exporting large repositories, especially to formats like JSON-LD and GraphML. This aligns with Bolt's performance rule around direct cursor iteration.
Measurement:
Tested via an in-memory database export script to confirm that the changes correctly handle raw rows and no data is lost. Furthermore,
uv run pytest tests/test_exporter.pyvalidates that functionality is preserved.PR created automatically by Jules for task 4874563076502746270 started by @n24q02m