Per-Source Configuration
Every source in DocsGPT carries its own behavior contract — a small config object that controls how that source is chunked when it is ingested and how it is retrieved when you ask a question. This lets you tune each source independently: a large reference manual can use a different chunking strategy and retriever than a short FAQ.
You edit this config from a source’s settings in the UI (shown below), or through the API. The same options are also available in Advanced settings when you first upload a document.
Per-source retrieval is enabled by default. Operators can turn it off instance-wide with PER_SOURCE_RETRIEVAL_ENABLED=false, in which case all sources fall back to the classic retriever regardless of their stored config.
Two kinds of settings: live vs. bake-time
The config has two groups of settings that differ in when they take effect:
| Group | When it applies | Re-ingest needed? |
|---|---|---|
Retrieval (retrieval.*) | Query time — applied live on the next question | No |
Chunking (chunking.*) | Ingest time — baked into the stored chunks | Yes |
Changing a retrieval setting takes effect immediately. Changing a chunking setting only affects documents ingested after the change, so you must re-ingest the source to apply it to existing content. The API response includes a requires_reingest flag to make this explicit.
Chunking configuration
Chunking decides how a document is split into the pieces that get embedded and stored.
{
"chunking": {
"strategy": "classic_chunk",
"max_tokens": 1250,
"min_tokens": 150,
"duplicate_headers": false
}
}| Field | Default | Description |
|---|---|---|
strategy | classic_chunk | Which chunking algorithm to use (see below). |
max_tokens | 1250 | Upper bound on chunk size in tokens. |
min_tokens | 150 | Lower bound; small fragments are merged up to this size. |
duplicate_headers | false | Repeat section headers into each child chunk for context. |
Available chunking strategies
| Strategy | Behavior |
|---|---|
classic_chunk | The default token-window splitter. An empty config reproduces DocsGPT’s historical chunking byte-for-byte. |
recursive | Recursive character/token splitter that tries to break on natural boundaries (paragraphs, sentences). |
markdown | Splits along Markdown structure (headings, sections) — good for docs and wikis. |
parent_child | Embeds small child chunks for precise matching but carries a larger parent window in metadata, so the model still sees surrounding context. |
semantic | Embeds sentences and splits where meaning shifts (at the 95th-percentile cosine-distance gap between adjacent sentences), falling back to recursive on failure. Produces topically coherent chunks at the cost of extra embedding calls during ingest. |
Chunking is bake-time. After changing strategy, max_tokens, min_tokens, or duplicate_headers, re-ingest the source so existing chunks are rebuilt.
Retrieval configuration
Retrieval decides which chunks are pulled in to answer a question. These settings apply live.
{
"retrieval": {
"retriever": "classic",
"exposure": "prefetch",
"chunks": 6,
"score_threshold": null,
"rephrase_query": true,
"prescreen": null,
"graph": {
"seed_strategy": "entities",
"passage_nodes": true,
"blend_vector": true
}
}
}| Field | Default | Description |
|---|---|---|
retriever | classic | Retrieval strategy: classic, hybrid, or graphrag. |
exposure | prefetch | How retrieved context reaches the model: prefetch or agentic_tool (see below). |
chunks | 6 | Final number of chunks (top-k) returned to the answer. Range 1–500. Set here, it overrides whatever a request asks for. |
score_threshold | null | Minimum similarity score. Honored by pgvector, MongoDB Atlas, Qdrant and Milvus (MongoDB Atlas scores cosine as (1 + cos) / 2, so 0.8 there is cosine 0.6; Qdrant’s score follows QDRANT_DISTANCE_FUNC, cosine by default); FAISS, Elasticsearch and the hybrid retriever ignore it — the config API returns a warnings entry when you set it on one of those. |
rephrase_query | true | Whether to run a query-rephrasing side-call before retrieval. |
prescreen | null | Optional LLM relevance filter (see below). null = off. |
graph | see GraphRAG | How the graphrag retriever walks the graph: seed_strategy, passage_nodes, blend_vector. Ignored by the other retrievers. |
reranker | null | Reserved for a future reranking step. Accepted and stored, but it has no effect. |
Who wins: source config or the request?
A source that has configured retrieval.chunks outranks the value sent with a
request. The owner tuned top-k for that corpus, so a client cannot raise or
lower it per call. Sources left at the default still let the request decide.
chunks is a total for the request, not a count per source: a request with
chunks: 6 over three sources takes two chunks from each. Every attached source
still gets at least one chunk, so with more sources than chunks the answer
receives one chunk per source (eight sources at chunks: 6 return eight
chunks). A source counts as configured when its stored
value differs from the default, so a source saved while the default was 2
keeps 2 as its own setting.
Requests are also bounded: chunks is clamped to 0–500, and 0 still means
“skip retrieval for this turn”.
Retrievers
classic— Vector similarity search. The default and a safe choice for any vector store.hybrid— Fuses vector search with full-text keyword search using Reciprocal Rank Fusion, which improves recall for exact terms, codes, and names that pure vector search can miss.graphrag— Knowledge-graph retrieval. Set indirectly when you enable GraphRAG on a source. See GraphRAG.
Keyword search for the hybrid retriever is currently implemented only for the pgvector vector store. On other stores (FAISS, Qdrant, Milvus, etc.) the keyword half returns nothing, so hybrid quietly behaves like classic (vector-only).
Exposure: prefetch vs. agentic tool
exposure controls how a source’s content is delivered to the model:
prefetch(default) — DocsGPT retrieves the top chunks up front and injects them into the prompt before the model answers. Best for focused Q&A over a source.agentic_tool— The source is exposed to the model as a search tool it can call on demand, deciding when and what to look up (browse-as-you-go) rather than receiving a bulk prefetch. This is the default exposure for Wiki sources.
Pre-screening (LLM relevance filter)
Pre-screening adds an optional map-reduce step between retrieval and answering: a base retriever fetches a wider set of candidates, an LLM screens them in batches, and only the most relevant survivors are passed to the answer. It improves precision on noisy sources at the cost of extra query-time LLM calls, so it is off by default.
{
"retrieval": {
"chunks": 8,
"prescreen": {
"candidate_k": 40,
"batch_size": 10,
"max_keep": 8,
"model": null
}
}
}| Field | Default | Description |
|---|---|---|
candidate_k | 40 | Candidates fetched before screening. Must be >= chunks. |
batch_size | 10 | Candidates screened per LLM call. |
max_keep | 8 | Survivors kept after screening. Must be <= candidate_k. |
model | null | Model used for screening. null reuses the request’s resolved model. |
Testing retrieval
Test retrieval shows the chunks a query actually retrieves from one source, so you can tune the retrieval settings without asking questions in a chat. Open it from Settings → Knowledge, either with Test retrieval in a source’s menu or with the Test retrieval button at the top of an open source. Anyone who can use the source can run it, including team viewers.
- Type a query the source should answer.
- Adjust the retrieval settings under the query if you want to try something else. They start from the source’s saved settings.
- Click Run.
The result lists the chunks in rank order with their file name, token count, score and text, plus the retriever used, the number of chunks and the latency. The settings you try are used for that test only and are not saved; change them in the source’s settings to keep them.
The test runs the same retrieval pipeline as an answer: the same retriever, chunks, score_threshold, pre-screening and token budget. It differs in three ways:
- It never rephrases the query, as with the first message of a chat.
- It applies the source’s settings even when
PER_SOURCE_RETRIEVAL_ENABLED=false, where answers use the classic retriever. - Pre-screening without its own
modeluses the instance’s default model, since a test has no chat model behind it. Settings that differ from the saved ones may use at most 20 pre-screening LLM calls (candidate_k/batch_size).
Reading the scores
The meaning of a score depends on the retriever and the vector store, and the modal labels each one:
| Label | Comes from | How to read it |
|---|---|---|
| similarity | classic on pgvector, MongoDB Atlas, Qdrant or Milvus | Cosine similarity: higher is better, at most 1. This is the value score_threshold is compared with. |
| distance | classic on FAISS | L2 distance: lower is better, with no fixed upper bound. |
| RRF | hybrid | A reciprocal-rank-fusion score. It only orders the results of this query; don’t compare it across queries or with a threshold. |
| no score | graphrag, and classic on Elasticsearch | The retriever ranks chunks without a comparable per-chunk score. A graphrag source that has no graph yet falls back to classic retrieval, so its results carry the vector store’s scores instead. |
When nothing comes back and a score_threshold is set, the modal says so: lower the threshold to see the near misses.
From the API
curl -X POST https://your-docsgpt/api/sources/<source_id>/search \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"query": "How do I rotate the encryption key?",
"retrieval": { "retriever": "hybrid", "chunks": 4 }
}'query is required (at most 2,000 characters). retrieval is optional: leave it out to use the source’s saved retrieval settings, or send a complete retrieval object (fields you leave out take their defaults, as in the table above). Nothing is saved. The response:
{
"success": true,
"query": "How do I rotate the encryption key?",
"retriever": "hybrid",
"retrieval": { "retriever": "hybrid", "chunks": 4, "...": "..." },
"total": 4,
"latency_ms": 182,
"chunks": [
{
"rank": 1,
"text": "...",
"title": "...",
"filename": "security.md",
"source": "...",
"tokens": 212,
"score": 0.0328,
"score_kind": "rrf"
}
]
}score_kind is cosine_similarity, l2_distance, rrf or null, matching the labels above; score is null when score_kind is.
Editing the config via API
The config is edited with a PATCH to the source’s config endpoint:
curl -X PATCH https://your-docsgpt/api/sources/<source_id>/config \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"retrieval": { "retriever": "hybrid", "chunks": 4 },
"chunking": { "strategy": "semantic" }
}'The response echoes the stored config and a requires_reingest flag:
{
"success": true,
"config": { "...": "..." },
"requires_reingest": true
}Notes:
- The body replaces the whole config rather than merging into it: any group or field you leave out is reset to its default. To change one setting, send the source’s current
config(fromGET /api/sources) with that setting changed. On a GraphRAG source, keep"retriever": "graphrag"and the top-levelgraphobject. - Invalid values are rejected with
400(strict validation on write). - The
kindfield (classic / wiki / graphrag) cannot be changed through this endpoint — converting a source to a Wiki or enabling GraphRAG uses dedicated endpoints. - Editing requires ownership of the source or a team
editorgrant; viewers receive403.
Related
- GraphRAG — knowledge-graph retrieval for a source.
- Wiki Sources — LLM-editable living documentation.
- Embeddings — the embedding model used during ingest and retrieval.