Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
22c9403
Add Document Ingestion module for curated docs and web search
ParamThakkar123 Jul 12, 2026
7ece173
Add document chunking strategies with provenance tracking
ParamThakkar123 Jul 18, 2026
578bf2d
Add embedding and storage modules
ParamThakkar123 Jul 18, 2026
940e863
feat: add retrieve() convenience method for text-query-to-search
ParamThakkar123 Jul 18, 2026
ab0cdf6
feat: add Prompt module for constructing grounded RAG prompts
ParamThakkar123 Jul 19, 2026
aacaf35
feat: add Execution module for FUNSQL query execution
ParamThakkar123 Jul 19, 2026
5d41761
chore: add test/Manifest*.toml to .gitignore
ParamThakkar123 Jul 19, 2026
5742949
Merge branch 'main' of https://github.com/JuliaHealth/HealthLLM.jl in…
ParamThakkar123 Sep 3, 2026
2ec8e0d
Updates
ParamThakkar123 Sep 3, 2026
73c54bf
docs: fix unresolved @ref cross-references in docs build
ParamThakkar123 Sep 5, 2026
d1bd1e0
Fixed merge conflicts
ParamThakkar123 Sep 21, 2026
63e856f
Merge pull request #53 from JuliaHealth/response_generation
ParamThakkar123 Sep 21, 2026
af9059d
Merge pull request #52 from JuliaHealth/prompt_construction
ParamThakkar123 Sep 21, 2026
707b6c3
Merge pull request #51 from JuliaHealth/retrieval
ParamThakkar123 Sep 21, 2026
c94fb93
Merge pull request #50 from JuliaHealth/embedding
ParamThakkar123 Sep 21, 2026
a5b5594
Merge pull request #49 from JuliaHealth/chunking
ParamThakkar123 Sep 21, 2026
66fe66a
rebase
ParamThakkar123 Sep 21, 2026
438ff42
Merge branch 'ingestion' of https://github.com/JuliaHealth/HealthLLM.…
ParamThakkar123 Sep 21, 2026
d15c452
Dependencies
ParamThakkar123 Sep 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@
*.jl.mem
/Manifest*.toml
/docs/Manifest*.toml
/test/Manifest*.toml
/docs/build/
.env
.env.example
Expand Down
10 changes: 6 additions & 4 deletions Project.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,24 +4,26 @@ version = "0.1.0"
authors = ["ParamThakkar123 <paramthakkar864@gmail.com> and TheCedarPrince <jacobszelko@gmail.com>"]

[deps]
HTTP = "cd3eb016-35fb-5094-929b-558a96fad6f3"
HuggingFaceHub = "d0076355-e2c0-48e6-a044-05906e51b7fc"
JSON3 = "0f8b85d8-7281-11e9-16c2-39a750bddbf1"
LibPQ = "194296ae-ab2e-5f79-8cd4-7183a0a5a0d1"
LinearAlgebra = "37e2e46d-f89d-539d-b4ee-838fcccc9c8e"
OpenAI = "e9f21f70-7185-4079-aca2-91159181367c"
PromptingTools = "670122d1-24a8-4d70-bfce-740807c42192"
RAGTools = "16ddad29-bbe8-45a7-857d-3d9514eb0023"
Serialization = "9e88b42a-f829-5b0c-bbe9-9e923198166b"
SparseArrays = "2f01184e-e22b-5df5-ae63-d93ebab69eaf"
Statistics = "10745b16-79ce-11e8-11f9-7d13ad32a3b2"
URIs = "5c2747f8-b7ea-4ff2-ba2e-563bfd36b1d4"

[compat]
HTTP = "1.11"
HuggingFaceHub = "0.1.2"
JSON3 = "1.14.3"
LibPQ = "1.18.0"
LinearAlgebra = "1.10"
OpenAI = "0.11"
PromptingTools = "0.82.1"
RAGTools = "0.7.0"
Serialization = "1.10"
SparseArrays = "1.10"
Statistics = "1.10"
URIs = "1.6"
julia = "1.10"
3 changes: 3 additions & 0 deletions docs/make.jl
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,9 @@ makedocs(;
pages=[
"Home" => "index.md",
"Getting Started" => "getting-started.md",
"Document Ingestion" => "ingestion.md",
"Building Embeddings" => "embeddings.md",
"Querying the RAG System" => "querying.md",
],
)

Expand Down
339 changes: 339 additions & 0 deletions docs/src/embeddings.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,339 @@
```@meta
CurrentModule = HealthLLM
```

# Building Embeddings

Retrieval quality in the RAG pipeline rests on two choices: **which model turns
text into vectors**, and **where those vectors are stored and searched**. This
page covers both — picking a model, generating embeddings, validating them, and
persisting them to one of three interchangeable stores.

```
chunks ─▶ embed ─▶ dim × n matrix ─▶ validate_embeddings ─┐
├─▶ LocalVectorStore (file)
├─▶ PgVectorStore (pgvector)
└─▶ FaissVectorStore (FAISS)
```

Throughout, embeddings use the **`dim × n` convention**: one column per chunk,
one row per embedding dimension. This matches the storage layer and
[`validate_embeddings`](@ref).

## Choosing a model

The pipeline is built around embedding models that are reachable from **both** a
local [Ollama](https://ollama.com) server and [HuggingFace](https://huggingface.co),
so an index built in one environment can be reproduced in the other. The registry
is [`EMBEDDING_MODELS`](@ref):

```julia
using HealthLLM

collect(keys(EMBEDDING_MODELS))
# ["nomic-embed-text", "all-minilm", "mxbai-embed-large", "bge-m3"]

embedding_dimension() # 768 (the default model)
embedding_dimension("all-minilm") # 384
```

| Name | Dim | Ollama tag | HuggingFace repo | Notes |
|---------------------|------|---------------------|-------------------------------------------|------------------------------|
| `nomic-embed-text` | 768 | `nomic-embed-text` | `nomic-ai/nomic-embed-text-v1.5` | **Default** — balanced quality/size |
| `all-minilm` | 384 | `all-minilm` | `sentence-transformers/all-MiniLM-L6-v2` | Lightest / fastest |
| `mxbai-embed-large` | 1024 | `mxbai-embed-large` | `mixedbread-ai/mxbai-embed-large-v1` | Higher quality, larger |
| `bge-m3` | 1024 | `bge-m3` | `BAAI/bge-m3` | Multilingual, long-context |

[`DEFAULT_EMBEDDING_MODEL`](@ref) is `"nomic-embed-text"`: it runs on both backends,
produces 768-dimensional vectors with strong retrieval quality, and is the model
used throughout the [ingestion docs](ingestion.md). Reach for `"all-minilm"` when
you want a smaller, faster index and can trade a little accuracy.

Resolve a model's backend-specific reference with [`embedding_ref`](@ref):

```julia
embedding_ref("nomic-embed-text"; provider=:ollama) # "nomic-embed-text"
embedding_ref("all-minilm"; provider=:huggingface) # "hf:sentence-transformers/all-MiniLM-L6-v2"
```

!!! note "Backend setup"
For Ollama, pull the tag once with `ollama pull nomic-embed-text` and make
sure the server is running. For HuggingFace, set a token (see below). Either
way, register the model in the RAG pipeline with `register_models(...)` as
shown in [Getting Started](getting-started.md).

### The HuggingFace backend

PromptingTools ships schemas for a long list of OpenAI-compatible providers but
none for HuggingFace, so this package supplies one:
[`HuggingFaceOpenAISchema`](@ref). It is an `AbstractOpenAISchema`, so
`aigenerate`, `aiembed`, message rendering and retries all work unchanged —
`provider = :huggingface` and `hf:`-prefixed model names route through it
automatically.

Set a token once per session, or let it come from the environment
(`HF_API_TOKEN`, `HF_TOKEN`, `HUGGINGFACE_API_KEY`, `HUGGING_FACE_HUB_TOKEN`,
checked in that order):

```julia
set_huggingface_api_key!(ENV["HF_TOKEN"])
huggingface_api_key() # what will actually be sent
```

!!! note "Pinning an inference provider"
The router auto-routes only to providers **enabled on your account**, so a
model can be live on HuggingFace and still be refused with
`model_not_supported`. Append the provider to pin it:

```julia
huggingface_providers("Qwen/Qwen2.5-7B-Instruct") # ["featherless-ai"]

register_models("hf:Qwen/Qwen2.5-7B-Instruct:featherless-ai", "hf:BAAI/bge-m3")
```

When routing is refused, the error raised here already names the live
providers and the exact model string to use, so you rarely need to look this
up yourself. Alternatively enable the provider at
[huggingface.co/settings/inference-providers](https://huggingface.co/settings/inference-providers).

Chat and embeddings use different HuggingFace surfaces, because the router's
OpenAI-compatible API covers chat completions only — `/v1/embeddings` answers
404 there:

| Call | Endpoint |
|--------------|---------------------------------------------------------------------------|
| `aigenerate` | [`HUGGINGFACE_ROUTER_URL`](@ref) + `/chat/completions` |
| `aiembed` | [`HUGGINGFACE_INFERENCE_URL`](@ref) + `/<model>/pipeline/feature-extraction` |

The feature-extraction reply is reshaped into the OpenAI embeddings response
`aiembed` expects, and token-level output (from models that do not pool
internally) is mean-pooled to one vector per input — so a HuggingFace embedding
matrix is the same `dim × n` shape as an Ollama one.

!!! note "Cold models"
A HuggingFace model that is not already warm loads while holding the
connection open — around 40–55s for `bge-m3` in practice. That exceeds
PromptingTools' 120s `aiembed` default under load, so this package raises the
read timeout to [`HUGGINGFACE_EMBED_TIMEOUT`](@ref) (300s) when you have not
chosen one yourself. Any `http_kwargs` you pass is left exactly as given.

To use a deployment that *does* speak OpenAI embeddings — a Text Embeddings
Inference container or a dedicated Inference Endpoint — pass its URL and the
request is forwarded there instead:

```julia
E = embed(chunks, "bge-m3"; provider = :huggingface,
api_kwargs = (; url = "https://my-endpoint.hf.space/v1"))
```

## Generating embeddings

[`embed`](@ref) turns text into a `dim × n` matrix using the chosen model and
backend. It accepts a single string or a vector of strings:

```julia
E = embed(["OMOP person table", "FunSQL From clause"]) # 768 × 2 (Ollama, default model)
size(E, 1) == embedding_dimension() # true

# A lighter model, via HuggingFace:
E2 = embed(chunks, "all-minilm"; provider=:huggingface)
```

`provider` selects the backend (`:ollama` default, or `:huggingface`); any extra
keyword arguments are forwarded to `PromptingTools.aiembed`.

In practice you embed the chunks produced by the ingestion layer:

```julia
docs = ingest(; sources = ["FunSQL.jl", "OMOP CDM"])
chunks = String[]
for d in docs
for c in chunk_document(d) # content-type-aware chunking
push!(chunks, c.text)
end
end

E = embed(chunks) # 768 × length(chunks)
```

## Validating embeddings

Before storing vectors, check that they are well-formed.
[`validate_embeddings`](@ref) runs a series of cheap structural checks and returns
a `(; dim, n, norms)` summary:

```julia
info = validate_embeddings(E; expected_dim = embedding_dimension(), chunks = chunks)
info.dim # 768
info.n # number of chunks
```

It throws on the first problem it finds:

1. an empty matrix (zero columns);
2. a dimension that does not match `expected_dim`;
3. a column count that does not match `length(chunks)`;
4. any non-finite entry (`NaN`/`Inf`);
5. a zero-vector (degenerate) column;
6. a column whose self-similarity is not `≈ 1`.

### Similarity sanity checks

Structural validity does not guarantee the model captures *meaning*. Two helpers
check semantics. [`cosine_similarity`](@ref) and [`similarity_matrix`](@ref) let
you inspect relationships directly:

```julia
cosine_similarity(view(E, :, 1), view(E, :, 2)) # similarity of the first two chunks
S = similarity_matrix(E) # full n × n cosine matrix (diagonal ≈ 1)
```

[`embedding_sanity_check`](@ref) is an end-to-end check that the model orders
meaning correctly — it embeds an anchor sentence, a semantically *near* sentence,
and an unrelated *far* one, and confirms `cos(anchor, near) > cos(anchor, far)`:

```julia
res = embedding_sanity_check() # uses OMOP-flavoured defaults
res.passed # true when near beats far
res.near, res.far # the two similarity scores
```

This requires a reachable backend (it makes live embedding calls), so it belongs
in a smoke test rather than a unit test.

## Storing embeddings

All three backends share one interface, so you can swap where vectors live without
changing surrounding code:

| Store | Backing | Requirement |
|--------------------------|----------------------------------|--------------------------------|
| [`LocalVectorStore`](@ref) | in-memory matrix + `Serialization` | none — built in |
| [`PgVectorStore`](@ref) | PostgreSQL + `pgvector` | `LibPQ` and a running server |
| [`FaissVectorStore`](@ref) | a FAISS index | `Faiss.jl` loaded in `Main` |

The common operations are [`add!`](@ref), [`search`](@ref), `length`, and — where
supported — [`save`](@ref) / [`load`](@ref).

### Local file-based store

The simplest option: vectors live in a `dim × n` matrix, search is cosine
similarity, and the whole store round-trips through `Serialization`. No external
service is required.

```julia
store = LocalVectorStore(embedding_dimension()) # 768-d store
add!(store, E, chunks) # append embeddings + their text
length(store) # number of stored vectors

hits = search(store, embed("count patients"), 5) # top-5 nearest
hits[1].chunk # most similar chunk text
hits[1].score # cosine similarity in [-1, 1]

save(store, "omop_index.jls") # persist
store2 = load(LocalVectorStore, "omop_index.jls") # restore
```

Vectors are L2-normalised on insertion by default (`normalize=true`), which makes
cosine search a plain dot product and keeps scores in `[-1, 1]`.

### PostgreSQL + pgvector

For a shared, queryable index, back the store with PostgreSQL's `pgvector`
extension. [`PgVectorStore`](@ref) wraps
[`store_embeddings_pgvector`](@ref) and [`search_embeddings_pgvector`](@ref):

```julia
using LibPQ

conn = LibPQ.Connection("postgresql://user:pass@localhost/health")
store = PgVectorStore(conn, embedding_dimension(); table = "omop_embeddings", metric = :cosine)

add!(store, E, chunks) # creates the table + inserts in one transaction
hits = search(store, embed("count patients"), 5) # Vector{Hit}: row id in `index`, raw `distance` kept
```

`metric` chooses the distance operator: `:cosine` (`<=>`), `:dot` (`<#>`), or
`:l2` (`<->`). Table names are validated as SQL identifiers before interpolation.
The underlying functions can also be called directly if you do not want the store
wrapper:

```julia
store_embeddings_pgvector(conn, E, chunks, embedding_dimension(); table = "omop_embeddings")
hits = search_embeddings_pgvector(conn, embed("count patients"), 5; table = "omop_embeddings") # raw (; id, chunk, distance)
```

!!! note "pgvector prerequisites"
The target database must have the `vector` extension installed
(`CREATE EXTENSION IF NOT EXISTS vector;`) and `LibPQ` available in the
environment.

### FAISS

For large local indexes, [`FaissVectorStore`](@ref) delegates to
[FAISS](https://github.com/facebookresearch/faiss). FAISS is an **optional**
dependency: the store works only when a `Faiss` module is loaded into `Main`
(e.g. `using Faiss`), and otherwise raises a clear error at construction — nothing
else in the package requires it to be installed.

```julia
using Faiss # optional; load before constructing

store = FaissVectorStore(embedding_dimension()) # inner-product index (cosine after normalisation)
add!(store, E, chunks)
hits = search(store, embed("count patients"), 5) # Vector{Hit}, as with every backend
```

As with the local store, vectors are normalised for the inner-product/cosine
metrics; use `metric = :l2` for a raw L2 index (scores are then negated distances
so that larger is always more similar).

## Retrieving by text

[`search`](@ref) takes a query *vector*, so at query time you would normally embed
the question yourself and pass the result in. [`retrieve`](@ref) folds those two
steps into one per-query call: give it the store and a **string**, and it embeds
the query and returns the nearest chunks. It works with any store backend and
returns exactly what that backend's [`search`](@ref) yields.

```julia
store = LocalVectorStore(embedding_dimension())
add!(store, embed(chunks), chunks)

hits = retrieve(store, "How do I count patients in OMOP?", 5)
hits[1].chunk # most relevant chunk text
```

`model`, `provider`, and any extra keyword arguments are forwarded to the embedder,
so you can retrieve with the same model you indexed with. The embedder itself is
injectable via the `embedder` keyword — handy for tests or a cached embedding
function — and is called as `embedder(query, model; provider, kwargs...)`.

## Putting it together

```julia
using HealthLLM

# 1. Gather and chunk documents.
docs = ingest(; sources = ["FunSQL.jl", "OMOP CDM"])
chunks = [c.text for d in docs for c in chunk_document(d)]

# 2. Embed with the default dual-backend model.
E = embed(chunks)

# 3. Validate before storing.
validate_embeddings(E; expected_dim = embedding_dimension(), chunks = chunks)

# 4. Store locally (swap for PgVectorStore / FaissVectorStore as needed).
store = LocalVectorStore(embedding_dimension())
add!(store, E, chunks)
save(store, "omop_index.jls")

# 5. Query.
for h in search(store, embed("How do I count patients?"), 3)
println(round(h.score, digits = 3), " ", first(h.chunk, 80))
end
```

See the [API reference](index.md) for full docstrings of every function above.
```
Loading
Loading