fix: bolt optimization for full graph exports peak memory overhead - #1015
Conversation
Co-authored-by: n24q02m <135627235+n24q02m@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
💡 What: Optimized full graph exports (
export_graphml,export_jsonld,export_dot,export_cypher, andexport_crg) to stream directly fromsqlite3.Cursorrows via newGraphStore.iter_raw_nodes()andGraphStore.iter_raw_edges()methods, instead of loading the entire graph into PythonGraphNode/GraphEdgeobjects.🎯 Why: Previously, generating exports required materializing tens of thousands of dataclass instances in memory, leading to massive peak memory spikes and CPU overhead during object construction, which is unnecessary since export formats only require scalar row values.
📊 Impact: Significantly reduces peak memory utilization and CPU time during full graph exports (O(1) python object overhead instead of O(N)).
🔬 Measurement: Memory profiling a call to
export_crgagainst a large repository will show substantially lower peak memory overhead asGraphNodeandGraphEdgeobject instantiations are completely eliminated during the export loop.PR created automatically by Jules for task 10172509774891061715 started by @n24q02m