Skip to content

Add Document Ingestion module for curated docs and web search - #48

Open
ParamThakkar123 wants to merge 19 commits into
mainfrom
ingestion
Open

ParamThakkar123 wants to merge 19 commits into
mainfrom
ingestion

Conversation

@ParamThakkar123

Copy link
Copy Markdown
Collaborator

Summary

Adds a new HealthLLM.Ingestion module for pulling documentation into a RAG index from curated sources (JuliaHealth, OMOP CDM, OHDSI, FunSQL.jl) and live web search via a pluggable provider (DuckDuckGo).

New files

  • src/ingestion.jl — module entry point
  • src/ingestion/types.jl — SourceDocument, SearchResult types
  • src/ingestion/curated.jl — CURATED_SOURCES registry
  • src/ingestion/fetch.jl — fetch_url, html_to_text, fetch_curated
  • src/ingestion/search.jl — AbstractSearchProvider, DuckDuckGoProvider
  • src/ingestion/pipeline.jl — ingest, ingest_to_index orchestration
  • test/IngestionTest.jl — offline unit tests
  • docs/src/ingestion.md — full documentation

Dependencies added

  • HTTP, URIs

Modified files

  • Project.toml, src/HealthLLM.jl, docs/make.jl, docs/src/index.md, test/runtests.jl

Introduces a new Ingestion module that pulls documentation into a RAG index
from two kinds of sources: curated docs (JuliaHealth, OMOP CDM, OHDSI,
FunSQL.jl) and live web search via a pluggable provider (DuckDuckGo).
Includes HTTP/S fetch with HTML-to-text cleaning, a search provider
interface, and an orchestration pipeline (ingest/ingest_to_index) that
feeds directly into RAGTools for indexing.

New files:
- src/ingestion.jl — module entry point
- src/ingestion/types.jl — SourceDocument, SearchResult
- src/ingestion/curated.jl — CURATED_SOURCES registry
- src/ingestion/fetch.jl — fetch_url, html_to_text, fetch_curated
- src/ingestion/search.jl — AbstractSearchProvider, DuckDuckGoProvider
- src/ingestion/pipeline.jl — ingest, ingest_to_index
- test/IngestionTest.jl — unit tests for offline logic
- docs/src/ingestion.md — full user-facing documentation

Dependencies added: HTTP, URIs
@codecov

codecov Bot commented Jul 12, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.68675% with 188 lines in your changes missing coverage. Please review.
✅ Project coverage is 71.66%. Comparing base (677f98e) to head (d15c452).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/storage.jl 54.02% 40 Missing ⚠️
src/ingestion/pipeline.jl 0.00% 30 Missing ⚠️
src/huggingface.jl 65.33% 26 Missing ⚠️
src/ingestion/fetch.jl 51.06% 23 Missing ⚠️
src/execution.jl 73.13% 18 Missing ⚠️
src/ingestion/search.jl 11.11% 16 Missing ⚠️
src/embeddings.jl 70.00% 15 Missing ⚠️
src/database.jl 61.90% 8 Missing ⚠️
src/ingestion/chunk.jl 96.91% 5 Missing ⚠️
src/utils.jl 88.63% 5 Missing ⚠️
... and 1 more
Additional details and impacted files
@@             Coverage Diff             @@
##             main      #48       +/-   ##
===========================================
+ Coverage   60.78%   71.66%   +10.88%     
===========================================
  Files           4       15       +11     
  Lines         102      720      +618     
===========================================
+ Hits           62      516      +454     
- Misses         40      204      +164     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

ParamThakkar123 and others added 18 commits July 18, 2026 09:57
Add embedding generation and vector storage functionality with tests and documentation.
Introduces a Prompt module with:
- FUNSQL system prompt and customizable PromptTemplate
- Context formatting from retrieved chunks
- Build prompts for LLM query construction

Includes full test coverage and documentation page.
Introduces an Execution module with:
- FUNSQL query execution against DuckDB via RAGTools
- Result extraction, formatting, and display
- End-to-end query-to-answer pipeline integration

Includes tests and documentation updates.
The Documentation job failed with six unresolvable `@ref` targets,
from two separate causes:

- `Chunk`, `HeaderChunk` and the rest of the chunking API are defined
  and exported by the `Ingestion` submodule but were never re-exported
  from `HealthLLM`, so `[`Chunk`](@ref)` in querying.md (which runs
  under `CurrentModule = HealthLLM`) had no binding to resolve against.
  Add them to the `import .Ingestion:` and `export` lists alongside the
  other ingestion names.

- Documenter resolves `@ref`s inside a docstring in that docstring's own
  module. The `Prompt` docstrings reference `retrieve`, `search` and
  `Chunk`, none of which `Prompt` imports, so they failed as
  `HealthLLM.Prompt.retrieve` and friends. Qualify them with the
  `[`name`](@ref Module.name)` form, which keeps the rendered link text
  unchanged.

`julia --project=docs docs/make.jl` now completes CrossReferences and
RenderDocument with no errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RYi1NAjjGQ848cUMLEbh1x
feat: add Execution module for FUNSQL query execution and answer generation
feat: add Prompt module for grounded RAG query construction
feat: add retrieve() convenience method for text-query-to-search
Add embedding and storage modules
Add document chunking strategies with provenance tracking

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant