FastAPI course-assistant product for ingesting knowledge files, indexing them in Postgres + pgvector, and answering learner questions with grounded citations. The app is built for a single-course trial deployment today, with a path to harden into a larger multi-course product later.
- Ingest
.pdf,.md,.markdown,.txt,.html, and.htmfiles. - Convert uploads into canonical Markdown sources.
- Store documents, sections, chunks, conversations, jobs, evals, feedback, audit events, and embeddings in Postgres.
- Use pgvector for retrieval. FAISS is intentionally not part of the current path.
- Answer with selected source snippets through an answer provider and validate citations.
- Stream
/chat/streamresponses from the provider when OpenAI answering is enabled. - Guard every chat query with an LLM router (
gpt-5.4-mini): off-topic, harmful, and prompt-injection queries are blocked with a fixed learner-friendly reply; allowed questions are difficulty-rated in the same call and hard questions are answered with a stronger model (gpt-5.4by default). - Persist every LLM call (router + answer) to
provider_call_logswith full request and response payloads, keyed by conversation id for end-to-end tracing. - Serve a pre-login landing page; after login, a three-pane learner workbench: left tabs (chat / 知識圖譜 / Sources), center content with a readable course-material reader (mermaid diagrams render natively), right citation sources and 來源內容, with 主題切換 (淺色/深色/自動). Learner chat includes 清除對話, 匯出 JSON, and 回報問題 (copies the session id and records an audit event).
- Two login roles configured in
.env: learner (student) and admin. Only admins see the 管理主控台, which gathers overview, uploads/indexing, 教材編輯 (markdown CRUD with reindex), document lifecycle, graph extraction, evals, 服務狀態 (system health), 系統設定 (runtime LLM settings, applied without restart), background jobs, provider usage with LLM call logs, and a sortable audit log. - Run app, worker, tests, backups, and deploy checks through Docker Compose and Make.
The accepted launch configuration is:
- Embeddings:
text-embedding-3-small - Embedding dimension:
KB_EMBEDDING_DIMENSION=768 - Answer model:
gpt-5.4-mini(easy) /gpt-5.4(hard, chosen by the query router) - Query router model:
gpt-5.4-mini - Max output tokens:
4096 - Retrieval strategy:
hybrid - Learner auth: configured learner + admin logins, no registration
- Admin auth: admin login session or
X-KB-Admin-Key
Latest local real-content acceptance:
- Artifact:
backups/real-content-20260529T171708Z/(ignored by git) - Indexed content: 35 course files, 819 sections/chunks
- Retrieval acceptance: 5/5 passed
- Live answer acceptance: 3/3 passed
See ops/live-answer-acceptance.md for the required learner-facing RAG checks.
Full step-by-step setup and usage walkthrough: docs/USAGE.md.
Prerequisites: Python 3.12, uv, Docker, and curl.
uv sync --python 3.12 --group dev
docker compose up -d postgres
make migrate
uv run --python 3.12 python -m scripts.seed_sample_docs
make devOpen http://localhost:8000, log in with the local Compose learner account if platform
auth is enabled, then index the seeded docs from another terminal:
make indexLocal Compose defaults are development-only:
- Learner login:
student/student-password - Admin login:
admin/admin-password - Admin key:
local-admin-key
| Task | Command |
|---|---|
| Install dependencies | uv sync --python 3.12 --group dev |
| Start local Postgres | docker compose up -d postgres |
| Run migrations | make migrate |
| Start app | make dev |
| Rebuild index | make index |
| Process one worker job | make worker-once |
| Run worker loop | make worker |
| Check worker runtime | make worker-status |
| Seed concept graph | make graph-seed |
| Seed eval cases | make eval-seed |
| Run scheduled evals | make eval-run |
| Unit tests | make test-unit |
| Integration tests | make test-integration |
| E2E tests | make test-e2e |
| Full test suite | make test |
| Lint/typecheck | make lint |
| Docker tests | make docker-test |
| Docker app/worker smoke | make docker-smoke |
| Deploy env validation | make deploy-check |
| Runtime ops smoke | make ops-check API_URL=http://localhost:8000 KB_ADMIN_API_KEY=local-admin-key |
| Real course package | make real-content-package |
Settings use the KB_ prefix unless noted. The Makefile loads .env when present and
exports only OPENAI_API_KEY and KB_POSTGRES_PORT for Make-driven workflows.
Required for staging or production:
KB_AUTH_SECRET_KEYKB_PLATFORM_USERNAME/KB_PLATFORM_PASSWORD(learner login)KB_ADMIN_USERNAME/KB_ADMIN_PASSWORD(admin login)KB_ADMIN_API_KEYKB_DATABASE_URLKB_DOCS_DIRKB_RAW_DIRKB_KB_DIROPENAI_API_KEYwhen using OpenAI providers
Important provider settings:
KB_EMBEDDING_PROVIDER=fake|openaiKB_ANSWER_PROVIDER=fake|openaiKB_OPENAI_EMBEDDING_MODEL=text-embedding-3-smallKB_OPENAI_CHAT_MODEL=gpt-5.4-miniKB_OPENAI_CHAT_MODEL_HARD=gpt-5.4(answer model for router-rated hard questions)KB_OPENAI_ROUTER_MODEL=gpt-5.4-mini(guardrail / difficulty router)KB_QUERY_ROUTER_ENABLED=trueKB_OPENAI_REQUEST_TIMEOUT_SECONDSKB_OPENAI_MAX_RETRIESKB_OPENAI_CHAT_MAX_COMPLETION_TOKENS(default 4096)KB_PROVIDER_BUDGET_*
Model, hard model, router model, max output tokens, temperature, and budgets can also be
overridden at runtime from the admin console (系統設定) without a restart; runtime
overrides are stored in the runtime_settings table and take precedence over env values.
Important knowledge graph settings:
KB_GRAPH_EXTRACTION_ENABLED=true: auto-chains concept extraction after every index rebuild. RequiresKB_ANSWER_PROVIDER=openai; the step is skipped (not failed) when the answer provider isfake.KB_GRAPH_MAX_CONCEPTS_PER_DOC=30: maximum concepts extracted per document.KB_GRAPH_EXTRACTION_TOKEN_BUDGET=12000: token budget for the extraction prompt sent to the answer provider.
Important retrieval and context settings:
KB_TOKEN_ENCODING=o200k_base: tiktoken encoding used for token-aware chunking and context budgeting.KB_CONTEXT_NEIGHBOR_SECTIONS=1: neighboring sections included on each side of a retrieved section when assembling answer context.KB_CONTEXT_TOKEN_BUDGET=8000: token budget for the assembled answer context.
Important runtime settings:
KB_RATE_LIMIT_*KB_MAX_CONCURRENT_UPLOADSKB_BACKGROUND_JOB_STALE_AFTER_SECONDSKB_BACKGROUND_JOB_RETRY_BASE_DELAY_SECONDSKB_BACKGROUND_JOB_RETRY_MAX_DELAY_SECONDSKB_WORKER_IDKB_WORKER_HEARTBEAT_INTERVAL_SECONDSKB_WORKER_HEARTBEAT_STALE_AFTER_SECONDSKB_PLATFORM_COHORTSKB_PLATFORM_EXTRA_VISIBILITY_LABELSKB_POSTGRES_PORT
Use ops/env.production.example as the production template, then validate the target environment with:
make deploy-checkupload/import
-> raw artifact in raw/
-> background_jobs.ingest.upload
-> canonical Markdown in docs/
-> background_jobs.index.rebuild
-> documents/sections/chunks + pgvector embeddings
-> search/chat retrieval
-> answer provider
-> citation validation + response diagnostics
POST /imports stays lightweight: it stores the raw artifact, creates import metadata,
and queues background work. python -m scripts.run_background_worker handles conversion,
index rebuilds, eval runs, retries, stale-job recovery, and worker heartbeats.
Identical uploads deduplicate by content hash. Same filename with different content is kept as a new artifact with a content-hash suffix. The app rejects empty uploads, unsupported extensions, invalid PDF signatures, obvious content-type mismatches, HTML without recognizable markup, and binary-looking text before writing raw artifacts.
Learner-facing:
POST /auth/loginPOST /auth/logoutGET /auth/sessionPOST /searchPOST /chatPOST /chat/streamPOST /chat/report(report a problematic session; writes achat.session_reportedaudit event keyed by conversation id)GET /sourcesGET /sources/{document_id}GET /sources/{document_id}/sections/{section_id}GET /graphGET /graph/concepts/{concept_id}
Admin-only (admin session or X-KB-Admin-Key):
POST /importsGET /imports/statusPOST /indexGET /index/statusGET /admin/jobsPOST /admin/jobsPOST /admin/jobs/recover-stalePOST /admin/jobs/{job_id}/requeuePOST /admin/jobs/{job_id}/cancelGET /admin/jobs/runtimeGET /admin/documentsGET /admin/documents/{document_id}/contentPUT /admin/documents/{document_id}/content(edit markdown in place + reindex)PATCH /admin/documents/{document_id}/lifecycleDELETE /admin/documents/{document_id}POST /admin/documents/{document_id}/reindexGET /admin/settings/PUT /admin/settings(runtime LLM/budget overrides)GET /admin/system-status(DB / index / worker / budget health)GET /admin/audit-eventsGET /admin/provider-observabilityGET /admin/provider-logs(full LLM request/response call log)POST /graph/extractGET /metrics
GET /health is liveness. GET /ready checks database, pgvector, Alembic migration
state, storage paths, platform auth, and index readiness.
Hybrid retrieval now fuses lexical, vector, and markdown candidates with Reciprocal Rank
Fusion (RRF), applying the score threshold to per-strategy scores before fusion. Chat
answers no longer see only the matched chunk: a context assembly step expands each hit to
its full section plus KB_CONTEXT_NEIGHBOR_SECTIONS neighbors on each side, within the
KB_CONTEXT_TOKEN_BUDGET token budget, and chat responses report this in a
context_assembly block.
Search and chat return retrieval diagnostics such as selected source IDs, rejected source IDs, threshold, strategy counts, top score, raw/merged/accepted/rejected counts, and score debug data. Chat responses also expose answer-quality metadata:
answer_validcitation_errorsselected_source_idscited_source_idscannot_confirm_reason
Provider citations must reference selected sources (or their context-assembly neighbors).
Citation matching normalizes anchors (slugified, punctuation-tolerant) so heading variants
still match; an answer with at least one valid citation stays valid and unmatched citation
tokens degrade gracefully in the UI. Only answers with zero valid citations downgrade to
the exact cannot-confirm response with cannot_confirm_reason="invalid_citations".
Guardrail-blocked queries return the fixed blocked reply with
cannot_confirm_reason="guardrail_blocked" and the router decision in query_route.
Streaming chat sends retrieval diagnostics in the sources event and final answer quality
in the done event.
Platform auth is intentionally simple: two configured logins for the trial course — a
learner account (role student) and an admin account (role admin). Only the admin role
sees the 管理主控台 entry. Admin API routes accept an authenticated admin session or
X-KB-Admin-Key.
Source visibility is enforced across search, chat, streaming chat, source preview, and the
knowledge graph. Canonical Markdown frontmatter can set visibility; omitted visibility
defaults to public. Learners can see:
publicrole:<role>user:<username>cohort:<name>fromKB_PLATFORM_COHORTS- labels from
KB_PLATFORM_EXTRA_VISIBILITY_LABELS
Private course material lives in course-materials-md/, which is ignored by git. Build a
deployable artifact with the real embedding provider:
make real-content-packageIf local port 5432 is already in use:
KB_POSTGRES_PORT=55432 make real-content-packageThe workflow uses an isolated Compose project named kb-real-content, indexes
course-materials-md/, runs retrieval acceptance cases, and writes:
postgres.dumpruntime-files.tar.gzreal-content-acceptance-report.json
Restore on the deploy target with:
make restore-db RESTORE_DB_FILE=<artifact>/postgres.dump CONFIRM_RESTORE=yes
make restore-files RESTORE_FILES_FILE=<artifact>/runtime-files.tar.gz CONFIRM_RESTORE=yes
make ops-check API_URL=https://your-app.example.com KB_ADMIN_API_KEY=$KB_ADMIN_API_KEYNever commit course-materials-md/, .env, or generated backups/ artifacts.
Use focused tests while developing, then run the broadest feasible check before claiming work is complete.
make test-unit
make test-integration
make test-e2e
make test
make lint
make docker-test
make docker-smokeDB-backed tests need KB_DATABASE_URL_TEST; otherwise they are skipped. CI is defined in
.github/workflows/ci.yml and runs lint/typecheck, local tests,
deploy env validation, Docker Compose validation, Docker tests, and Docker smoke checks.
The app includes an app-native operations foundation:
- Structured logs with
X-Request-ID /health,/ready, protected/metrics- Rate limits and upload concurrency guards
- Provider timeout, retry, budget, usage, and error controls
- Admin audit/security event log
- Provider observability dashboard data
- Document lifecycle management
- Background worker runtime supervision
- Stuck job recovery and requeue endpoints
Create an application backup:
make backup BACKUP_DIR=backups/$(date -u +%Y%m%dT%H%M%SZ)Restore operations require CONFIRM_RESTORE=yes. File restore overlays archived files and
does not delete stale files.
Runbooks:
For the first learner trial, use one app container, one worker container, and Postgres with
pgvector. Run Alembic migrations once per deploy, keep docs, raw, and .kb on durable
storage, and run:
make ops-check API_URL=https://your-app.example.com KB_ADMIN_API_KEY=$KB_ADMIN_API_KEYBefore inviting learners, confirm:
- CI is green for the deployed commit.
make deploy-checkpasses on the target environment./readyis ready.- Worker heartbeat is fresh.
- Login works with the configured platform user.
- Known search/chat cases return expected citations.
- Live answer acceptance passes with OpenAI answering enabled.