17 KiB
JARVIS Knowledge Base (Phase 3) Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Replace the TODO-stub POST/GET /api/v1/documents endpoints with a real document-upload + semantic-search knowledge base, backed by pgvector in the existing Postgres instance and Ollama for embeddings.
Architecture: Documents are chunked (character-based, with overlap) on upload, each chunk is embedded via Ollama's nomic-embed-text model, and stored in a new document_chunks table (pgvector column). Search embeds the query the same way and ranks chunks by cosine distance (<=>).
Tech Stack: FastAPI, psycopg2 (existing db_query helper), pgvector Postgres extension, aiohttp (existing Ollama HTTP pattern), pytest (new, local-only dev dependency).
Spec: docs/superpowers/specs/2026-09-12-knowledge-base-design.md
Global Constraints
- No dedicated git repository exists for this project yet (the working directory sits inside an unrelated, very large home-directory-level repo that must not be touched). Every "Commit" step below is replaced by "confirm the file is saved" — do not run
git add/git commit. - Runtime is the live VPS at
72.61.186.98, containerjarvis-api(imagepython:3.11-slim, code bind-mounted from/home/jarvis-core/jarvis/api,uvicorn --reload). SSH:ssh -F /dev/null -o IdentitiesOnly=yes -i ~/.ssh/jarvis_core_key jarvis-core@72.61.186.98. Local files live inClaude outputs/(note the space — quote the path in shell commands) and must be scp'd to/home/jarvis-core/jarvis/...to take effect. - Reuse the existing
db_query(query, params, fetch)async helper inmain.pyfor all DB access — do not introduce an ORM. - Follow the existing code style in
main.py: section comments (# ============ X ============), try/except around endpoint bodies that logs and raisesHTTPException, async endpoints. - Embedding model:
nomic-embed-text(768 dimensions) — this fixes thevector(768)column width used below. - Chunking: ~1000 characters, 100 character overlap (from the spec) — do not change these defaults without updating the spec.
Task 1: pgvector extension + document_chunks table
Files:
- Modify:
Claude outputs/docker-compose.yml(postgres image) - Create:
Claude outputs/migrations/002_document_chunks.sql
Interfaces:
- Produces: Postgres table
document_chunks(id, document_id, chunk_index, content, embedding vector(768), created_at), used by Task 4 and Task 5's DB helpers.
Context: the plain postgres:16-alpine image does not ship the pgvector extension. pgvector/pgvector:pg16 is a drop-in image (built from postgres:16) with the extension precompiled — swapping the image and recreating the container is safe, the existing postgres_data volume is untouched.
- Step 1: Swap the Postgres image to the pgvector-enabled build
Edit Claude outputs/docker-compose.yml:
postgres:
image: pgvector/pgvector:pg16
(only the image: line under the postgres: service changes — everything else stays as-is).
- Step 2: Deploy the compose change and recreate the postgres container
SSHOPTS="-F /dev/null -o IdentitiesOnly=yes -i /c/Users/Jonny/.ssh/jarvis_core_key"
scp $SSHOPTS "C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs\docker-compose.yml" jarvis-core@72.61.186.98:/home/jarvis-core/jarvis/docker-compose.yml
ssh $SSHOPTS jarvis-core@72.61.186.98 "cd /home/jarvis-core/jarvis && docker compose up -d postgres && sleep 5 && docker compose ps postgres"
Expected: jarvis-postgres shows Up ... (healthy) with the new image.
- Step 3: Write the migration file
Create Claude outputs/migrations/002_document_chunks.sql:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE document_chunks (
id SERIAL PRIMARY KEY,
document_id INTEGER NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
chunk_index INTEGER NOT NULL,
content TEXT NOT NULL,
embedding vector(768) NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_document_chunks_document_id ON document_chunks(document_id);
CREATE INDEX idx_document_chunks_embedding ON document_chunks
USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
- Step 4: Apply the migration to the live database
SSHOPTS="-F /dev/null -o IdentitiesOnly=yes -i /c/Users/Jonny/.ssh/jarvis_core_key"
scp $SSHOPTS "C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs\migrations\002_document_chunks.sql" jarvis-core@72.61.186.98:/home/jarvis-core/jarvis/migrations/002_document_chunks.sql
ssh $SSHOPTS jarvis-core@72.61.186.98 "docker exec -i jarvis-postgres psql -U jarvis -d jarvis < /home/jarvis-core/jarvis/migrations/002_document_chunks.sql"
Expected output ends with CREATE INDEX (no errors). If CREATE EXTENSION fails with "could not open extension control file", Step 1/2 didn't take effect — re-check the image was actually swapped (docker inspect jarvis-postgres | grep Image).
- Step 5: Verify
ssh $SSHOPTS jarvis-core@72.61.186.98 "docker exec jarvis-postgres psql -U jarvis -d jarvis -c '\d document_chunks'"
Expected: column list showing embedding | vector(768) among others.
- Step 6: Confirm files saved (no git repo for this project — see Global Constraints)
Task 2: Chunking function (TDD)
Files:
- Modify:
Claude outputs/main.py(addchunk_text) - Test:
Claude outputs/tests/test_chunking.py - Create:
Claude outputs/requirements-dev.txt
Interfaces:
-
Produces:
chunk_text(text: str, chunk_size: int = 1000, overlap: int = 100) -> list[str]inmain.py, consumed by Task 4's upload endpoint. -
Step 1: Add pytest as a local-only dev dependency
Create Claude outputs/requirements-dev.txt:
pytest==8.3.3
- Step 2: Install it locally
cd "C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs"
python -m pip install -r requirements-dev.txt -r requirements.txt
- Step 3: Write the failing tests
Create Claude outputs/tests/test_chunking.py:
import sys
import os
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from main import chunk_text
def test_chunk_text_empty_returns_empty_list():
assert chunk_text("") == []
def test_chunk_text_shorter_than_chunk_size_returns_single_chunk():
assert chunk_text("hello", chunk_size=1000, overlap=100) == ["hello"]
def test_chunk_text_exact_multiple_of_chunk_size():
text = "a" * 20
chunks = chunk_text(text, chunk_size=10, overlap=2)
assert chunks == [text[0:10], text[8:18], text[16:20]]
- Step 4: Run tests to verify they fail
cd "C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs"
python -m pytest tests/test_chunking.py -v
Expected: FAIL with ImportError: cannot import name 'chunk_text' (or similar — chunk_text doesn't exist yet). Note: importing main.py will also try to import psycopg2, redis, anthropic, aiohttp — these must already be installed locally via Step 2 for the import to get far enough to fail on chunk_text specifically, not on a missing package.
- Step 5: Implement
chunk_text
In Claude outputs/main.py, add near the other pure helpers (e.g. above # ============ DB HELPERS ============):
def chunk_text(text: str, chunk_size: int = 1000, overlap: int = 100) -> list:
if not text:
return []
if len(text) <= chunk_size:
return [text]
chunks = []
step = chunk_size - overlap
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end])
if end >= len(text):
break
start += step
return chunks
- Step 6: Run tests to verify they pass
cd "C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs"
python -m pytest tests/test_chunking.py -v
Expected: 3 passed.
- Step 7: Confirm files saved (no git repo for this project — see Global Constraints)
Task 3: Ollama embedding helper
Files:
- Modify:
Claude outputs/main.py(addOLLAMA_EMBED_MODELconfig +get_embedding) - Modify:
Claude outputs/docker-compose.yml(jarvis-api environment)
Interfaces:
-
Consumes:
OLLAMA_HOST(already defined inmain.py),aiohttp(already imported). -
Produces:
async def get_embedding(text: str) -> listinmain.py, consumed by Task 4 and Task 5. -
Step 1: Add the embed-model config constant
In Claude outputs/main.py, next to the other # ============ CONFIG ============ constants:
OLLAMA_EMBED_MODEL = os.getenv("OLLAMA_EMBED_MODEL", "nomic-embed-text")
- Step 2: Add
get_embedding
In Claude outputs/main.py, add after chunk_text:
async def get_embedding(text: str) -> list:
timeout = aiohttp.ClientTimeout(total=30)
async with aiohttp.ClientSession(timeout=timeout) as session:
async with session.post(
f"{OLLAMA_HOST}/api/embeddings",
json={"model": OLLAMA_EMBED_MODEL, "prompt": text},
) as resp:
if resp.status != 200:
body = await resp.text()
raise RuntimeError(f"Ollama embedding failed ({resp.status}): {body}")
data = await resp.json()
return data["embedding"]
- Step 3: Add
OLLAMA_EMBED_MODELto the deployed environment
In Claude outputs/docker-compose.yml, under jarvis-api: environment:, add a line next to OLLAMA_HOST:
- OLLAMA_EMBED_MODEL=nomic-embed-text
- Step 4: Confirm files saved (no automated test here —
get_embeddingis verified end-to-end in Task 4; see Global Constraints for why there's no commit step)
Task 4: Wire up document upload
Files:
- Modify:
Claude outputs/main.py(insert_document,insert_document_chunks,upload_document)
Interfaces:
-
Consumes:
chunk_text(Task 2),get_embedding(Task 3),db_query(existing),DEFAULT_USER_ID(existing global). -
Produces: DB helpers
insert_document(title, content, document_type) -> intandinsert_document_chunks(document_id: int, chunks: list) -> None, wherechunksislist[tuple[str, list]]of(chunk_text, embedding)pairs. -
Step 1: Add a vector-literal formatting helper and the DB helpers
In Claude outputs/main.py, add near the other DB helpers (after save_message):
def _vector_literal(embedding: list) -> str:
return "[" + ",".join(repr(float(x)) for x in embedding) + "]"
async def insert_document(title: str, content: str, document_type: str) -> int:
row = await db_query(
"INSERT INTO documents (user_id, title, content, document_type) VALUES (%s, %s, %s, %s) RETURNING id",
(DEFAULT_USER_ID, title, content, document_type),
fetch="one",
)
return row["id"]
async def insert_document_chunks(document_id: int, chunks: list):
for index, (content, embedding) in enumerate(chunks):
await db_query(
"""
INSERT INTO document_chunks (document_id, chunk_index, content, embedding)
VALUES (%s, %s, %s, %s::vector)
""",
(document_id, index, content, _vector_literal(embedding)),
)
- Step 2: Replace the
upload_documentendpoint
In Claude outputs/main.py, replace the existing stub:
@app.post("/api/v1/documents")
async def upload_document(title: str, content: str, document_type: str = "general"):
"""Upload document to knowledge base"""
if not pg_pool:
raise HTTPException(status_code=503, detail="Database is not configured")
if not content.strip():
raise HTTPException(status_code=422, detail="content must not be empty")
try:
logger.info(f"Uploading document: {title}")
document_id = await insert_document(title, content, document_type)
chunks = chunk_text(content)
embedded_chunks = [(c, await get_embedding(c)) for c in chunks]
await insert_document_chunks(document_id, embedded_chunks)
return {
"document_id": document_id,
"title": title,
"status": "indexed",
"chunk_count": len(chunks),
}
except HTTPException:
raise
except Exception as e:
logger.error(f"Document upload error: {str(e)}")
raise HTTPException(status_code=503, detail=f"Embedding failed: {str(e)}")
- Step 3: Deploy
SSHOPTS="-F /dev/null -o IdentitiesOnly=yes -i /c/Users/Jonny/.ssh/jarvis_core_key"
OUT="C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs"
scp $SSHOPTS "$OUT/main.py" jarvis-core@72.61.186.98:/home/jarvis-core/jarvis/api/main.py
scp $SSHOPTS "$OUT/docker-compose.yml" jarvis-core@72.61.186.98:/home/jarvis-core/jarvis/docker-compose.yml
ssh $SSHOPTS jarvis-core@72.61.186.98 "cd /home/jarvis-core/jarvis && docker compose up -d jarvis-api && sleep 8 && docker logs --tail 20 jarvis-api"
Expected: log ends with Application startup complete. (no traceback).
- Step 4: Pull the embedding model
ssh $SSHOPTS jarvis-core@72.61.186.98 "docker exec jarvis-ollama ollama pull nomic-embed-text"
Expected: ends with success.
- Step 5: Test the upload endpoint against the live API
ssh $SSHOPTS jarvis-core@72.61.186.98 "curl -s -X POST 'http://localhost:8000/api/v1/documents?title=Oeffnungszeiten&content=Wir%20haben%20Montag%20bis%20Freitag%20von%209%20bis%2017%20Uhr%20geoeffnet.&document_type=faq'; echo"
Expected: JSON with "status":"indexed" and "chunk_count":1.
- Step 6: Confirm files saved (no git repo for this project — see Global Constraints)
Task 5: Wire up document search + end-to-end verification
Files:
- Modify:
Claude outputs/main.py(search_chunks,search_documents)
Interfaces:
-
Consumes:
get_embedding(Task 3),db_query(existing). -
Produces: DB helper
search_chunks(query_embedding: list, limit: int) -> list[dict]. -
Step 1: Add the search DB helper
In Claude outputs/main.py, add next to insert_document_chunks:
async def search_chunks(query_embedding: list, limit: int):
return await db_query(
"""
SELECT dc.content, dc.chunk_index, d.id AS document_id, d.title,
dc.embedding <=> %s::vector AS distance
FROM document_chunks dc
JOIN documents d ON d.id = dc.document_id
ORDER BY distance ASC
LIMIT %s
""",
(_vector_literal(query_embedding), limit),
fetch="all",
)
- Step 2: Replace the
search_documentsendpoint
In Claude outputs/main.py, replace the existing stub:
@app.get("/api/v1/documents")
async def search_documents(query: str, limit: int = 10):
"""Search knowledge base using vector similarity"""
if not pg_pool:
raise HTTPException(status_code=503, detail="Database is not configured")
if not query.strip():
raise HTTPException(status_code=422, detail="query must not be empty")
try:
query_embedding = await get_embedding(query)
rows = await search_chunks(query_embedding, limit)
return {
"query": query,
"results": [
{
"document_id": r["document_id"],
"title": r["title"],
"content": r["content"],
"chunk_index": r["chunk_index"],
"distance": float(r["distance"]),
}
for r in rows
],
"count": len(rows),
}
except HTTPException:
raise
except Exception as e:
logger.error(f"Document search error: {str(e)}")
raise HTTPException(status_code=503, detail=f"Search failed: {str(e)}")
- Step 3: Deploy
SSHOPTS="-F /dev/null -o IdentitiesOnly=yes -i /c/Users/Jonny/.ssh/jarvis_core_key"
OUT="C:\Users\Jonny\Projekte\Claude\JARVIS\Claude outputs"
scp $SSHOPTS "$OUT/main.py" jarvis-core@72.61.186.98:/home/jarvis-core/jarvis/api/main.py
ssh $SSHOPTS jarvis-core@72.61.186.98 "cd /home/jarvis-core/jarvis && docker compose restart jarvis-api && sleep 8 && docker logs --tail 20 jarvis-api"
Expected: log ends with Application startup complete. (no traceback).
- Step 4: End-to-end test (matches spec's manual verification requirement)
Using the document uploaded in Task 4 Step 5 ("Wir haben Montag bis Freitag von 9 bis 17 Uhr geoeffnet."), search with a thematically related query:
ssh $SSHOPTS jarvis-core@72.61.186.98 "curl -s 'http://localhost:8000/api/v1/documents?query=Wann%20habt%20ihr%20auf%3F'; echo"
Expected: results[0].title == "Oeffnungszeiten" and results[0].content contains the opening-hours text, with the lowest distance of any returned result.
- Step 5: Confirm files saved (no git repo for this project — see Global Constraints)