MIT licensed Β· runs fully offline

No-fee enterprise AI.
A full multi-agent system, on your laptop.

Every layer of this platform β€” ETL, warehouse, vector search, reranking, agent routing, generation β€” is the real open-source technology production teams use, wired together to run locally with no cloud bill and no API key.

0production layers mapped
0cloud API keys required
0components with graceful fallback
demo.py

How a question moves through the system

Two ingestion paths β€” structured accounts data and unstructured policy docs β€” converge in a LangGraph agent that decides which tool a question needs. The dashed lines are live: this is the actual path a request travels.

accounts_raw.csv sample_docs/*.txt Spark ETL (local) LangChain loader β†’ splitter DuckDB warehouse BGE embed β†’ FAISS index Cross-encoder reranker LangGraph planner routes to sql_agent or document_agent Ollama Β· Qwen2.5:7B

Try the routing live

These questions run against the platform's synthetic sample data. Pick one, or type your own β€” you'll see which agent it routes to before the answer comes back.

SAMPLE QUESTIONS
Pick a question on the left to see it routed and answered.

Production β†’ free-stack mapping

Every technology below is genuinely running, not stubbed out. Each module's docstring repeats the exact line you'd change to point at the managed service instead.

LayerProductionThis project (free, local)
Distributed ETLSpark (cluster)Real PySpark, local mode β€” only master() changes for a cluster
Analytical warehouseBigQueryDuckDB β€” same SQL dialect family, in-process columnar engine
Vector databaseAlloyDB Vector SearchFAISS β€” the ANN library underlying much of production vector search
EmbeddingsGemini embeddingsBGE-small-en-v1.5 via sentence-transformers
Rerankingoften skippedBGE cross-encoder β€” genuine two-stage retrieve-then-rerank
Doc/chunk handlingLangChain (cloud loaders)LangChain β€” same library, local file loader
LLMGeminiOllama + Qwen2.5:7B β€” fully local inference
Agent frameworkGoogle ADKLangGraph β€” typed-state multi-agent graphs
Tool protocolMCPFastMCP β€” real MCP server, real tool discovery
Graph analyticsNeo4j / Spanner GraphNetworkX β€” in-memory graph + centrality
API layerFastAPI (cloud)FastAPI β€” identical, on localhost
FrontendStreamlit / internal toolStreamlit β€” identical
β“˜
Two components fall back gracefully if a first-run download is unavailable: BGE embeddings/reranker fall back to local TF-IDF, and Ollama generation falls back to extractive, quote-and-cite answers. Neither hard-crashes the app β€” check_setup.py tells you which mode you're in.

Project layout

One FastAPI app, one Streamlit frontend, and a demo script that walks every stage without starting a server.

tree
app/
  api/       chat.py, ingest.py, sql.py
  rag/       loader, splitter, embeddings,
             vectorstore, retriever, reranker
  llm/       ollama.py
  agents/    planner, sql_agent,
             document_agent, workflow_agent
  warehouse/ duckdb.py
  graph/     knowledge_graph.py
  etl/       spark_pipeline.py
  main.py    FastAPI entry point
  mcp_server.py FastMCP server
data/
  accounts_raw.csv    synthetic ETL input
  sample_docs/*.txt   synthetic policy docs
demo.py       one-shot CLI walkthrough
streamlit_app.py
check_setup.py
what check_setup.py verifies
βœ“ BGE embeddings loaded (or TF-IDF fallback)
βœ“ Ollama reachable + qwen2.5:7b pulled
βœ“ FAISS index built from sample_docs
βœ“ DuckDB warehouse populated
βœ“ FastMCP tools registered
  β†’ document_search
  β†’ sql_query
  β†’ related_knowledge_concepts

# tells you exactly which components
# are at full strength vs. fallback,
# and what closes the gap.

Running it

No cloud account, no API key. Ollama is optional but recommended β€” without it the app falls back to extractive, retrieval-only answers.

Install dependencies

A virtual environment is optional but recommended. PySpark also needs a JRE.

bash
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env

Pull a local model (recommended)

The app runs without this step, falling back to extractive answers β€” but this is what turns on grounded generation.

bash
# install Ollama from https://ollama.com, then:
ollama pull qwen2.5:7b

Check what's active

Confirms which components are running at full strength vs. their offline fallback.

bash
python check_setup.py

Run the one-shot demo

Walks through Spark ETL β†’ LangGraph agent routing β†’ RAG β†’ graph analytics β†’ MCP tool calls, printed to your terminal.

bash
python demo.py

Or run the full app

Two terminals: the FastAPI backend, then the Streamlit frontend that calls it over HTTP.

bash
# terminal 1
uvicorn app.main:app --reload --port 8000

# terminal 2
streamlit run streamlit_app.py