Repository navigation
Expand file tree
/
Copy pathvector-search.html
More file actions
141 lines (136 loc) · 6.57 KB
/
Copy pathvector-search.html
File metadata and controls
141 lines (136 loc) · 6.57 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Vector Search</title>
<style>
body {
background: #ffffff;
color: #000000;
font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, Helvetica, Arial, sans-serif;
line-height: 1.6;
margin: 0;
padding: 1.25rem;
max-width: 820px;
}
h1 { font-size: 1.6rem; margin: 0 0 0.75rem; }
h2 { font-size: 1.2rem; margin: 1.75rem 0 0.5rem; border-bottom: 1px solid #e0e0e0; padding-bottom: 0.2rem; }
code {
background: #f4f4f4;
padding: 0.1rem 0.3rem;
border-radius: 3px;
font-family: "SF Mono", Menlo, Consolas, monospace;
font-size: 0.9em;
}
ul { padding-left: 1.25rem; }
li { margin: 0.3rem 0; }
table { border-collapse: collapse; margin: 0.5rem 0; width: 100%; }
th, td { border: 1px solid #ddd; padding: 0.4rem 0.6rem; text-align: left; vertical-align: top; }
th { background: #f4f4f4; }
.note {
border-left: 3px solid #999;
padding: 0.25rem 0.75rem;
margin: 1rem 0;
background: #fafafa;
}
</style>
</head>
<body>
<h1>Vector Search Plugin</h1>
<h2>Executive overview</h2>
<p><b>Vector Search</b> adds <b>semantic</b>, meaning-based results to
Code on the Go's project search. Instead of matching only exact text, it ranks
code by similarity of meaning, so a search for "read a file into a string" can
surface the relevant method even when those exact words don't appear.</p>
<p>It is a <b>headless</b> plugin — it has no screens of its own. It
contributes an extra <b>"Semantic Results"</b> section to the existing project
search.</p>
<h2>Core functionality</h2>
<ul>
<li><b>Semantic project search</b> — contributes ranked results via the
IDE's project-search extension point.</li>
<li><b>Indexing</b> — walks the project (capped, skipping
<code>build</code>/<code>.git</code> and similar), chunks files by
language-aware boundaries, embeds those chunks in batches, and stores the
vectors in a local SQLite database.</li>
<li><b>Routed through your own AI selection</b> — it embeds with the
backend you chose in AI settings and that backend's embedding model. It
names no provider of its own.</li>
<li><b>Versioned vectors</b> — every row records the backend, model and
width that produced it; a search only ranks vectors of the same origin, and
changing either builds the index again.</li>
<li><b>Cosine-similarity ranking</b> with a relevance cutoff so only strong
matches are returned.</li>
<li><b>Semantic Search settings screen</b> — in Preferences →
Configuration: shows the selected backend and whether it supports Vector
Search (with how to fix it if not), picks the embedding model, says where
the code is sent, reports what is indexed, and clears the index after
confirmation.</li>
</ul>
<h2>Technical architecture</h2>
<table>
<tr><th>Component</th><th>Role</th></tr>
<tr><td><code>VectorSearchPlugin</code></td><td>Entry point. Implements
<code>ProjectSearchExtension</code> and answers search requests, asking
<code>IndexCoordinator</code> for an index build when one is needed.</td></tr>
<tr><td><code>IndexCoordinator</code></td><td>Builds, reuses and clears the
index on a background scope, one operation at a time, so a clear never
leaves a partial index behind.</td></tr>
<tr><td><code>EmbedderResolver</code></td><td>Resolves the selected backend to
an <code>EmbeddingBackend</code>, or says which of the reasons it could
not.</td></tr>
<tr><td><code>CodeChunker</code></td><td>Splits files into small, language-aware
chunks with overlap for context.</td></tr>
<tr><td><code>EmbeddingBatches</code></td><td>Groups chunks into batched calls
so a large project is never held in memory whole.</td></tr>
<tr><td><code>EmbeddingIndexingService</code></td><td>Collects files and stores
chunk vectors, with their provenance, in a local SQLite database
(<code>embeddings.db</code>).</td></tr>
<tr><td><code>ReindexDecision</code></td><td>Decides whether the existing index
can answer the query, or has to be built again.</td></tr>
<tr><td><code>VectorMath</code> / <code>VectorSearchService</code></td><td>Cosine
similarity and top-K ranking of stored embeddings against the query.</td></tr>
<tr><td><code>SemanticSearchSettingsFragment</code></td><td>The settings screen,
contributed through <code>SettingsExtension</code>. Lists and sets the model
through the backend's <code>EmbeddingModelSelectable</code>.</td></tr>
<tr><td><code>BackendWatch</code></td><td>Listens for backend changes on AI
Core's inference service, so the screen updates without reopening.</td></tr>
</table>
<p>Embeddings are requested through <b>AI Core</b>'s inference service over the
shared service registry; the plugin contains no model, no HTTP client and no
native code of its own.</p>
<h2>Usage</h2>
<ol>
<li>Install <b>AI Core</b> and an agent plugin whose backend embeds
(<b>OpenAI</b> or <b>Gemini</b>), then select it in AI settings and give it
a key.</li>
<li>Install <b>Vector Search</b> via the Plugin Manager and restart the
IDE.</li>
<li>Optionally open <b>Preferences → Configuration → Semantic
Search</b> to check the backend is supported and choose its embedding
model.</li>
<li>Use the IDE's <b>project search</b>. Semantic matches appear under a
<b>"Semantic Results"</b> section. The first search on a project triggers a
background index build.</li>
</ol>
<div class="note">
<b>Privacy:</b> the index is a local SQLite database, but the chunk text is
sent to the backend you selected in order to be embedded — indexing a
project sends its source to that provider. Vector Search never picks a
backend for you: with none selected, nothing is sent and nothing is indexed.
</div>
<h2>Key benefits</h2>
<ul>
<li><b>Find code by meaning</b> — surfaces relevant code even without an
exact keyword match.</li>
<li><b>Honest about what it is</b> — with no embedder available it
contributes nothing, rather than presenting word matches as semantic
ones.</li>
<li><b>Provider-neutral</b> — it follows your AI selection and gains any
new backend that embeds, with no change here.</li>
<li><b>Lightweight</b> — no bundled model or native libraries; a compact
SQLite index.</li>
</ul>
</body>
</html>