No-fee enterprise AI.
A full multi-agent system, on your laptop.
Every layer of this platform β ETL, warehouse, vector search, reranking, agent routing, generation β is the real open-source technology production teams use, wired together to run locally with no cloud bill and no API key.
How a question moves through the system
Two ingestion paths β structured accounts data and unstructured policy docs β converge in a LangGraph agent that decides which tool a question needs. The dashed lines are live: this is the actual path a request travels.
Try the routing live
These questions run against the platform's synthetic sample data. Pick one, or type your own β you'll see which agent it routes to before the answer comes back.
Production β free-stack mapping
Every technology below is genuinely running, not stubbed out. Each module's docstring repeats the exact line you'd change to point at the managed service instead.
| Layer | Production | This project (free, local) |
|---|---|---|
| Distributed ETL | Spark (cluster) | Real PySpark, local mode β only master() changes for a cluster |
| Analytical warehouse | BigQuery | DuckDB β same SQL dialect family, in-process columnar engine |
| Vector database | AlloyDB Vector Search | FAISS β the ANN library underlying much of production vector search |
| Embeddings | Gemini embeddings | BGE-small-en-v1.5 via sentence-transformers |
| Reranking | often skipped | BGE cross-encoder β genuine two-stage retrieve-then-rerank |
| Doc/chunk handling | LangChain (cloud loaders) | LangChain β same library, local file loader |
| LLM | Gemini | Ollama + Qwen2.5:7B β fully local inference |
| Agent framework | Google ADK | LangGraph β typed-state multi-agent graphs |
| Tool protocol | MCP | FastMCP β real MCP server, real tool discovery |
| Graph analytics | Neo4j / Spanner Graph | NetworkX β in-memory graph + centrality |
| API layer | FastAPI (cloud) | FastAPI β identical, on localhost |
| Frontend | Streamlit / internal tool | Streamlit β identical |
check_setup.py tells you which mode you're in.Project layout
One FastAPI app, one Streamlit frontend, and a demo script that walks every stage without starting a server.
app/ api/ chat.py, ingest.py, sql.py rag/ loader, splitter, embeddings, vectorstore, retriever, reranker llm/ ollama.py agents/ planner, sql_agent, document_agent, workflow_agent warehouse/ duckdb.py graph/ knowledge_graph.py etl/ spark_pipeline.py main.py FastAPI entry point mcp_server.py FastMCP server data/ accounts_raw.csv synthetic ETL input sample_docs/*.txt synthetic policy docs demo.py one-shot CLI walkthrough streamlit_app.py check_setup.py
β BGE embeddings loaded (or TF-IDF fallback) β Ollama reachable + qwen2.5:7b pulled β FAISS index built from sample_docs β DuckDB warehouse populated β FastMCP tools registered β document_search β sql_query β related_knowledge_concepts # tells you exactly which components # are at full strength vs. fallback, # and what closes the gap.
Running it
No cloud account, no API key. Ollama is optional but recommended β without it the app falls back to extractive, retrieval-only answers.
Install dependencies
A virtual environment is optional but recommended. PySpark also needs a JRE.
python3 -m venv venv && source venv/bin/activate pip install -r requirements.txt cp .env.example .env
Pull a local model (recommended)
The app runs without this step, falling back to extractive answers β but this is what turns on grounded generation.
# install Ollama from https://ollama.com, then: ollama pull qwen2.5:7b
Check what's active
Confirms which components are running at full strength vs. their offline fallback.
python check_setup.py
Run the one-shot demo
Walks through Spark ETL β LangGraph agent routing β RAG β graph analytics β MCP tool calls, printed to your terminal.
python demo.py
Or run the full app
Two terminals: the FastAPI backend, then the Streamlit frontend that calls it over HTTP.
# terminal 1 uvicorn app.main:app --reload --port 8000 # terminal 2 streamlit run streamlit_app.py