๐ง Company Brain MVP (v2.0.0)
Company Brain is an automated, offline-first knowledge extraction pipeline. It connects to 38 connectors across your company's communication, engineering, HR, and analytics stack, pulls all the scattered information, runs it through a local AI model, and outputs structured "Skills" โ machine-readable procedure cards that AI agents can directly execute.
Instead of manually writing instruction manuals for your AI tools, Company Brain generates them automatically by watching how your team communicates and documents things.
โ ๏ธ v2.0.0 Behavioral Change โ License Key Required: Version 2.0.0 introduces a tier-gated license verification system (Firebase Realtime Database REST API validation, 24-hour local session caching, and 3-day offline grace period). Run
company-brain initto activate your license key (SoloorEnterprisetier) before connecting services.
โจ What Does It Actually Do?
- Reads Your Data: Securely connects to 38 connectors (Slack, Google Docs, GitHub, Notion, Discord, Linear, Outlook, GitLab, Dropbox, Mixpanel, Amplitude, Algolia, Exa, Perplexity, Facebook, Todoist, Web Crawler, Jira, Asana, Calendly, ClickUp, Airtable, Datadog, Segment, Front, Zoom, Twitter, HubSpot, Salesforce, Monday, Basecamp, Ashby, BambooHR, Deel, Rippling, Gmail, Google Sheets, Google Drive) ingesting raw text through a unified connector registry.
- Understands Context: It breaks the text down and uses a local AI model (
gemma4:e4b) to figure out what the text is actually about (e.g., "Is this a refund policy?" or "Is this a server deployment guide?"). - Synthesizes "Skills": It groups related information together and writes step-by-step procedures, including decision points (if/then rules) and edge cases.
- Exports for AI Agents: It outputs everything into a clean
skills_file.jsonthat you can plug directly into tools like LangChain, AutoGen, or your own custom AI bots so they know exactly how your company operates.
๐ ๏ธ Technical Architecture & Algorithms
For technical deep-dives, here is exactly how the pipeline operates under the hood:
1. Ingestion & Connectors (Data Layer)
- Abstract Base Connector: All 38 connectors implement
BaseConnector, a shared abstract class that enforces a standardingest(days_back)interface, providesretry_with_backoff()for exponential backoff on API rate limits (429/5xx errors), and checksCONNECTORS_ENABLEDfor dynamic toggling without code changes. - Connector Registry:
registry.pymaintains a central list of all connector classes.get_enabled_connectors()instantiates each one and filters out unconfigured connectors (e.g., ifGITHUB_TOKENis missing,GitHubConnectorsilently skips). This meansSyncSchedulernever imports individual connectors โ it only loops over whatever is enabled. - Error Isolation: Each connector runs in a
try/exceptblock inside the sync loop. A bad token or an API outage on one connector logs an error and moves on โ it will not crash the entire sync run. - Idempotency: Before processing, the system calls
get_processed_source_item_ids()which returns all item IDs already stored inprocessed_chunks. Any item already in that set is skipped. Running sync 10 times on the same data produces the same result as running it once.
2. Document Chunking Algorithm
- Semantic Sentence Splitting: Instead of naive character-count splitting (which breaks code blocks and sentences in half), the
DocumentChunkeruses regex-based sentence boundary detection. - Sliding Window Overlap: Text is chunked into 5-sentence blocks with a 2-sentence overlap. This guarantees that context isn't lost across chunk boundaries, which is critical for accurate LLM extraction.
3. Knowledge Extraction (LLM Pipeline)
- Combined Single LLM Call: The
KnowledgeExtractorsends a single optimised prompt that simultaneously classifies the chunk type (procedure,policy,decision,incident,general) AND extracts key concepts as a JSON array. This halves API latency compared to two sequential calls. - Structured JSON Fallbacks: Because open-source LLMs can hallucinate formatting, the prompt enforces strict JSON output, which is then parsed using Python's built-in
jsonlibrary with regex fallbacks to strip out markdown code fences.
4. Skill Synthesis (Clustering & Generation)
- Context Window Assembly: The
SkillsSynthesizerfilters the database for high-confidence chunks (confidence_score >= 0.2) classified as actionable knowledge. - Schema Enforcement: It asks the LLM to act as a technical writer, reading the raw concepts and synthesizing them into a strict domain model (
SkillPydantic class). This generates the final procedure steps, prerequisites, and edge cases.
5. UI Architecture (Event-Driven Dashboard)
- Decoupled State: The
richterminal dashboard runs independently of the backend pipeline. - Log Interception: Instead of polluting
SyncSchedulerwith UI logic, a custom Pythonlogging.Handlerintercepts backend logs (e.g.,"Phase 1: Ingesting data"), updates the UI's internal state machine, and computes progress/ETA โ theSyncSchedulernever knows a UI exists.
โ๏ธ The Pipeline Workflow
graph TD;
A[16 Data Sources via REST/GraphQL APIs] -->|BaseConnector.ingest| B(ConnectorRegistry)
B -->|RawDataItem objects| C(SQLite: raw_data_items)
C -->|Sliding Window Chunker| D{Local LLM: gemma4:e4b}
D -->|Single combined prompt| E(SQLite: processed_chunks)
E -->|Confidence filtering + concept clustering| F[LLM Synthesizer]
F -->|Pydantic Skill schema| G((skills_file.json))
๐ Setup Guide
Prerequisites
- Python 3.11+
- Ollama โ Download from ollama.ai
- API keys for the connectors you want to enable (see
.env.example)
Option A โ Install via pip (Recommended for users)
# 1. Install the package
pip install company-brain
# 2. Pull the required AI model and start Ollama
ollama pull gemma4:e4b
ollama serve # Keep running in background
# 3. Initialize โ creates .env from template, verifies Ollama
company-brain init
# 4. Fill in your API keys
# Edit the .env file created in the current directory
Tip:
company-brain initwill automatically copy.env.exampleinto your working directory if no.envexists, and will tell you exactly which keys to fill in.
Option B โ Developer / Source Setup
# Clone and enter the repo
git clone https://github.com/your-org/company-brain.git
cd company-brain
# Create virtual environment
python -m venv venv
.\venv\Scripts\activate # Windows
# source venv/bin/activate # macOS/Linux
# Install in editable mode
pip install -e .
# Pull the required AI model
ollama pull gemma4:e4b
ollama serve
# Configure credentials
cp .env.example .env
# Edit .env and add your API keys
Connector Credentials
Edit .env and fill in only the connectors you need โ unused connectors are automatically skipped if their credentials are absent:
| Connector | Env Var | Where to get it |
|---|---|---|
| Slack | SLACK_BOT_TOKEN |
api.slack.com/apps |
| Google Docs/Sheets | OAuth via credentials.json |
console.cloud.google.com |
| GitHub | GITHUB_TOKEN |
github.com/settings/tokens |
| Notion | NOTION_TOKEN |
notion.so/my-integrations |
| Linear | LINEAR_API_KEY |
Linear Settings |
| Discord | DISCORD_BOT_TOKEN |
discord.com/developers |
See .env.example for the full list of all 16 connectors.
๐ป Usage & Commands
If installed via pip, use the company-brain command from anywhere:
# First-time setup check
company-brain init
Developer mode: If running from source, prefix commands with
python main.pyinstead ofcompany-brain.
Run an Interactive Sync (Recommended)
Processes sources in small batches and pauses after each batch.
company-brain sync --interactive
Run a Full One-Time Sync
Processes everything in one go without stopping.
company-brain sync --once
Run Continuous Background Sync
Automatically re-syncs every 30 minutes.
company-brain sync
Check Current Status
View a breakdown of skills generated, categories, and confidence scores.
company-brain status
Export Your Skills
Export structured data for use by AI agents.
# Export as JSON (Best for AI Agents)
company-brain export --output output/skills.json
# Export as Markdown (Best for Humans)
company-brain export --output output/skills.md --format markdown
Company Brain โ Complete Technical Reference
This document covers every aspect of the Company Brain project: what it does, how it works under the hood, every API used, every technology used, the full data pipeline, CLI commands, and the engineering decisions made. Written for YC interviews and technical deep-dives.
Table of Contents
- What is Company Brain?
- The Core Problem It Solves
- High-Level Architecture
- Technology Stack
- Connectors โ All 16 Data Sources
- Full Data Pipeline โ Step by Step
- Database Design
- The LLM Integration
- CLI Commands Reference
- Terminal Dashboard (UI)
- Key Engineering Decisions
- Directory Structure
1. What is Company Brain?
Company Brain is an offline-first knowledge extraction pipeline. It connects to 16 of your company's existing communication, engineering, and analytics tools, pulls all the scattered information, runs it through a local AI model, and outputs structured "Skills" โ machine-readable procedure cards that AI agents can directly execute.
In plain English: Your team's knowledge lives in thousands of Slack messages, GitHub issues, Notion pages, and Linear tickets. Right now, no AI agent can act on that knowledge because it's buried in unstructured text across 16 different platforms. Company Brain reads all of it and converts it into clean, structured instructions.
2. The Core Problem It Solves
| Without Company Brain | With Company Brain |
|---|---|
| AI agents have no idea how your company operates | AI agents get a skills_file.json with exact procedures |
| Building custom SOPs takes weeks of manual work | Company Brain auto-generates them by reading your existing docs |
| Institutional knowledge lives in people's heads / old Slack threads | It's extracted, structured, and searchable |
| Onboarding new hires is slow because no one knows where to find information | Skills are tagged by category, have prerequisites, and success criteria |
| Your knowledge is siloed across 16+ tools | A single unified sync pulls everything into one knowledge base |
3. High-Level Architecture
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Data Sources (16) โ
โ Slack GitHub Notion Discord Linear Google Docs Google Sheets โ
โ Outlook/Teams GitLab Dropbox Mixpanel Amplitude โ
โ Algolia Exa Perplexity Facebook โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ REST / GraphQL APIs
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Connector Layer (BaseConnector) โ
โ โ
โ โข Every connector inherits BaseConnector (get_source_name, ingest) โ
โ โข registry.py: get_enabled_connectors() loops all 16, skips unconfigured โ
โ โข retry_with_backoff() handles 429/5xx with exponential backoff โ
โ โข Error isolation: one failing connector never blocks the others โ
โ โข CONNECTORS_ENABLED env var toggles connectors without code changes โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ List[RawDataItem] (Pydantic model)
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SQLite Database โ
โ (via SQLAlchemy ORM) โ
โ Table: raw_data_items Table: processed_chunks Table: skills โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Processing Pipeline โ
โ โ
โ 1. DocumentChunker โ sentence-based sliding window (5 sentences, โ
โ 2-sentence overlap), URL/email masking, noise filtering โ
โ โ
โ 2. KnowledgeExtractor โ single optimised prompt to gemma4:e4b via Ollama โ
โ - Classifies chunk type (procedure/policy/decision/incident/general) โ
โ - Extracts key concepts as JSON array โ
โ - Computes heuristic confidence score (0.0 โ 1.0) โ
โ โ
โ 3. SkillsSynthesizer โ filters high-confidence chunks, clusters by โ
โ primary concept, asks Gemma to write a complete Skill card โ
โ (steps, decisions, prerequisites, edge cases) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Export Layer โ
โ skills_file.json / skills_file.md โ
โ (ready to plug into LangChain, AutoGen, or custom bots) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
4. Technology Stack
| Category | Library / Tool | Why We Used It |
|---|---|---|
| Language | Python 3.11 | Mature ecosystem for AI/NLP pipelines |
| Data Validation | pydantic v2 |
Enforces strict schema for every data model |
| Database ORM | sqlalchemy v2 |
Maps Python classes to SQLite tables cleanly |
| Database | SQLite | Zero-config local persistence, no server needed |
| HTTP Client | requests |
Used by all new connectors for REST/GraphQL API calls |
| Slack Integration | slack-sdk v3 |
Official Slack client; handles pagination & auth |
| Google Integration | google-api-python-client, google-auth-oauthlib |
Shared OAuth 2.0 flow for both Google Docs and Google Sheets |
| Local AI | ollama (Python client) |
Runs Gemma 4 locally; zero data sent externally |
| CLI Framework | click |
Clean, composable command-line interface |
| Terminal UI | rich |
Beautiful dashboard with live animations, progress bars |
| Background Scheduler | apscheduler |
Runs sync jobs on a 30-minute interval |
| Env Management | python-dotenv |
Loads .env credentials without hardcoding secrets |
5. Connectors โ All 16 Data Sources
Company Brain uses a pluggable connector architecture. All connectors inherit from BaseConnector and are registered in registry.py. The sync loop dynamically loads only the ones with valid credentials configured in .env.
Tier 1 โ Core Business Knowledge
| Connector | Auth | What It Ingests | Dedup Key |
|---|---|---|---|
| Slack | Bot Token | Messages, threads, replies from all joined channels | slack:{channel_id}:{ts} |
| Google Docs | OAuth 2.0 | Document paragraph text, headers, table cell contents | gdoc:{doc_id} |
| Google Sheets | OAuth 2.0 (reused) | Spreadsheet tabs as structured Header: Value row text |
gsheet:{spreadsheet_id}:{sheet_name} |
| GitHub | Personal Access Token | Issues, PR descriptions, labels, review comments | github:{repo}:{issue_number} |
| Notion | Internal Integration Token | Page block hierarchies: headings, lists, quotes, code | notion:{page_id} |
| Discord | Bot Token | Server text channels and message history | discord:{channel_id}:{message_id} |
| Linear | API Key | Issues, descriptions, team context, comment threads (GraphQL) | linear:{issue_id} |
| Outlook / Teams | Azure OAuth 2.0 (MS Graph) | Outlook email threads + Teams channel messages | outlook:{message_id} / teams:{channel_id}:{message_id} |
Tier 2 โ Extended Sources
| Connector | Auth | What It Ingests | Dedup Key |
|---|---|---|---|
| GitLab | Personal Access Token | Projects, issues, merge requests, issue notes | gitlab:{project}:{type}:{iid} |
| Dropbox | OAuth Access Token | Text files (.md, .txt, .json, .csv, .doc) from shared folders |
dropbox:{file_id} |
| Mixpanel | API Secret | Tracked event schema catalog | mixpanel:event_schema:{project_id} |
| Amplitude | API Key + Secret | Event taxonomy definitions and descriptions | amplitude:taxonomy:{key} |
| Algolia | App ID + API Key | Indexed search records from a named index | algolia:{index}:{objectID} |
| Exa | API Key | Web research results for configured search queries | exa:{hash(query)} |
| Perplexity | API Key | AI search completions for configured prompts | perplexity:{hash(prompt)} |
| Page Access Token | Page posts and customer comment threads | facebook:{post_id} |
How to Enable/Disable Connectors
Edit the CONNECTORS_ENABLED variable in your .env:
# Enable only what you use โ unconfigured connectors are silently skipped
CONNECTORS_ENABLED=slack,google_docs,github,notion,linear
Leave it empty to attempt all connectors (only those with valid credentials will run).
6. Full Data Pipeline โ Step by Step
Step 1: Ingestion
The SyncScheduler calls get_enabled_connectors() from registry.py. This returns a list of instantiated connectors whose credentials are present. For each connector, ingest(days_back=30) is called and the returned RawDataItem objects are collected.
# Every piece of data becomes this shape regardless of source
class RawDataItem(BaseModel):
id: str # e.g., "github:owner/repo:42" or "notion:abc123"
source: DataSourceType # enum: "slack", "github", "notion", etc.
source_id: str # platform-native ID (channel ID, doc ID, issue ID)
title: str # human-readable display name
content: str # full normalized text
author: str # username or display name
created_at: datetime
updated_at: datetime
raw_metadata: dict # source-specific extras (labels, state, channel, etc.)
Idempotency check: Before processing, the system calls get_processed_source_item_ids() on the database. This returns the set of all item IDs already processed into chunks. The pipeline filters out any item whose ID is already in this set. This means:
- You can stop the sync halfway through and resume exactly where you left off.
- Running sync again never re-processes already-seen items.
Step 2: Chunking โ DocumentChunker
Long documents are split into smaller, manageable pieces before being sent to the LLM.
Text Cleaning (always runs first):
- Collapses multiple whitespace characters into single spaces
- Strips control characters (ASCII 0-31, 127-159)
- Replaces URLs with the token [URL]
- Replaces email addresses with the token [EMAIL]
Sliding-Window Sentence Chunking:
- Default chunk size:
5 sentences - Default overlap:
2 sentences(retains tail context from the previous chunk) - Uses regex
(?<=[.!?])\s+to split on sentence boundaries (not mid-sentence) - Chunks shorter than 20 characters are discarded
Why overlap? Without overlap, if a procedure starts at the end of one chunk and continues at the start of the next, the LLM would only see half the context. The overlap ensures key context is preserved across boundaries.
Step 3: Knowledge Extraction โ KnowledgeExtractor
For every chunk, the extractor makes a single combined LLM call (optimised from two separate calls):
Prompt: "Classify this text as ONE of: procedure, decision, incident, policy, or general.
Then extract 3-5 key concepts. Return as JSON: { type: ..., concepts: [...] }"
Output: { "type": "procedure", "concepts": ["Handle Refunds", "Verify order status", "Payment processor"] }
Confidence Scoring (no LLM call, pure heuristic):
+0.3 if chunk is longer than 100 words
+0.2 if chunk is 50โ100 words
+0.3 if classified as "procedure", "policy", or "decision"
+0.2 if 4+ key concepts were extracted
+0.1 if 2โ3 key concepts were extracted
Max score: 1.0
Every processed chunk becomes a ProcessedChunk object and is saved to the database.
Step 4: Skill Synthesis โ SkillsSynthesizer
This is where the real magic happens.
Phase A โ Filtering: Only chunks with confidence_score >= 0.2 move forward.
Phase B โ Clustering: Chunks are grouped by their first (primary) key concept. For example, all chunks where the first concept is "Handle Refunds" end up in the same cluster. Clusters with fewer than 2 chunks get merged into a miscellaneous cluster.
Phase C โ Skill Generation: For each cluster, the LLM is given all the chunk texts combined and asked to synthesize a complete Skill card:
{
"name": "clear skill name",
"description": "what this skill does",
"category": "refunds / pricing / incidents / ...",
"procedure_steps": ["Step 1", "Step 2", "..."],
"decision_points": {"if customer > 30 days": "deny refund"},
"prerequisites": ["who can perform this", "required access"],
"success_criteria": ["how to verify completion"],
"exceptions": ["edge cases to watch for"]
}
The final Skill object is upserted (insert or update) into the SQLite skills table.
Step 5: Export
Running python main.py export serializes all skills from the database into a single SkillsFile object and writes it to disk.
{
"version": "1.0.0",
"generated_at": "2026-07-25T00:00:00",
"company_name": "Your Company",
"skills": [
{
"id": "skill-uuid-...",
"name": "Handle Customer Refunds",
"description": "...",
"category": "refunds",
"procedure_steps": ["Check eligibility", "Verify order", "Issue refund"],
"decision_points": {"if_disputed": "escalate to manager"},
"examples": [],
"prerequisites": ["Customer support access"],
"success_criteria": ["Refund confirmation sent"],
"exceptions_and_edge_cases": ["Active subscriptions need billing cancel"],
"source_items": ["github:myorg/repo:42", "slack-C01-123...", "notion:abc123"],
"confidence_score": 0.75
}
],
"metadata": {
"total_skills": 1,
"by_category": {"refunds": 1}
}
}
7. Database Design
The system uses SQLite with SQLAlchemy ORM. The database lives at data/company_brain.db (path configurable via DATABASE_URL in .env).
3 Tables:
raw_data_items
โโโ id (PK) โ unique ID e.g. "github:myorg/repo:42", "notion:abc123"
โโโ source โ enum: "slack", "github", "notion", "linear", etc.
โโโ source_id โ platform-native ID (channel ID, doc ID, issue ID)
โโโ title โ display name
โโโ content โ full raw text
โโโ author โ username or display name
โโโ created_at
โโโ updated_at
โโโ raw_metadata โ JSON blob with source-specific fields (labels, state, etc.)
processed_chunks
โโโ id (PK) โ "chunk-{uuid}"
โโโ source_item_id โ FK to raw_data_items.id
โโโ chunk_text โ the actual text block
โโโ chunk_index โ position within source document
โโโ key_concepts โ JSON list e.g. ["Refunds", "Policy"]
โโโ chunk_type โ "procedure", "policy", "decision", "incident", "general"
โโโ confidence_score โ float 0.0 to 1.0
โโโ created_at
skills
โโโ id (PK) โ "skill-{uuid}"
โโโ name โ e.g. "Handle Customer Refunds"
โโโ description
โโโ category โ e.g. "refunds"
โโโ procedure_steps โ JSON list
โโโ decision_points โ JSON dict
โโโ examples โ JSON list
โโโ prerequisites โ JSON list
โโโ success_criteria โ JSON list
โโโ exceptions_and_edge_cases โ JSON list
โโโ source_items โ JSON list of contributing raw_data_items IDs
โโโ last_updated
โโโ confidence_score
8. The LLM Integration
Model: gemma4:e4b โ Gemma 4 with 4 billion parameters (E4B = Efficient 4B). Always use this model; do not swap to Mistral.
How it runs: Ollama is a lightweight server that runs locally on your machine. You install it once (ollama pull gemma4:e4b), and then the Python code communicates with it via the ollama Python client on localhost:11434.
The ollama client (not raw HTTP):
import ollama
client = ollama.Client()
response = client.generate(model="gemma4:e4b", prompt="...", stream=False)
result = response["response"]
JSON Parsing Robustness: Because LLMs sometimes wrap JSON in markdown code fences (like ```json ... ```), we use a regex fallback:
json_match = re.search(r'\[.*?\]', result_text, re.DOTALL) # for arrays
json_match = re.search(r'\{.*\}', result_text, re.DOTALL) # for objects
if json_match:
data = json.loads(json_match.group())
Why Gemma 4? Outperformed Mistral in adhering to JSON schemas during testing. Smaller footprint than 7B/13B models while producing more structured outputs.
9. CLI Commands Reference
All commands are run from inside the project root directory using the venv Python interpreter.
# Start interactive batch sync (RECOMMENDED)
# Processes 2 items at a time, asks whether to continue after each batch
.\venv\Scripts\python.exe main.py sync --interactive
# Run a single full sync (processes everything, no stops)
.\venv\Scripts\python.exe main.py sync --once
# Run continuous auto-sync (syncs every 30 minutes in the background)
.\venv\Scripts\python.exe main.py sync
# Check current database status: skill count, categories, confidence
.\venv\Scripts\python.exe main.py status
# Export skills as JSON (for AI agents)
.\venv\Scripts\python.exe main.py export --output output/skills.json
# Export skills as readable Markdown (for humans)
.\venv\Scripts\python.exe main.py export --output output/skills.md --format markdown
# Disable the Rich dashboard UI (useful for piping output to logs)
.\venv\Scripts\python.exe main.py sync --once --plain
Interactive Batch Menu (appears after each batch of 2 documents):
[Batch 1 complete. 5 total chunks processed so far.]
Do you want to: (1) Process next batch (2) Synthesize skills now and STOP (3) Quit immediately?
1โ Continue to next 2 items2โ Stop ingesting new data, run synthesis on what you have right now, and export3โ Exit immediately without synthesizing
10. Terminal Dashboard (UI)
The terminal UI is built with the rich library and runs inside a rich.live.Live context. This means the dashboard panel stays fixed at the top of the terminal while log messages scroll underneath it.
Key components:
SyncStateโ A plain Python class that holds the current progress percentage, file count, skill count, task statuses, and elapsed time.SyncDashboardโ A custom__rich_console__renderable that reads fromSyncStateand draws the panel usingrich.panel.Panel,rich.console.Group, andrich.progress.Progress.ProceduralNetworkโ The animated neural network animation. It's NOT pre-built frames. It usesmath.sin(time.time() * 3 + offset)to procedurally compute the brightness/state of each node and edge in real time at ~10 FPS.DashboardLogHandlerโ A custom Pythonlogging.Handlerthat intercepts log messages from the backend (like "Phase 1: Ingesting data") and maps them to UI state updates (progress bar %, which task is active). This is how the backend and UI stay decoupled โ theSyncSchedulernever knows a UI exists.
11. Key Engineering Decisions
Why an Abstract Base Connector + Registry pattern?
With 16 connectors, hardcoding each one into SyncScheduler would create a monolithic, hard-to-maintain sync loop. Instead, BaseConnector enforces a uniform ingest() interface, and registry.py acts as a plugin system. Adding a new connector is three steps: write the class, register it in registry.py, add credentials to .env.example. The scheduler code never changes.
Why SQLite instead of a cloud database?
Privacy-first design. Company data stays on-premise. SQLite also requires zero configuration, which makes setup trivial. The path is configurable via DATABASE_URL if you want to swap in PostgreSQL for production.
Why local LLM (Ollama) instead of OpenAI?
Companies have confidential Slack data, GitHub issues, emails. Sending it to a third-party API is a major legal and trust risk. Ollama + Gemma 4 gives equivalent results while keeping everything local. This is a core value proposition: "your data, your machine, your AI."
Why a single combined LLM call instead of two?
The original design made two LLM calls per chunk: one for classification, one for concept extraction. This was optimised to a single JSON-returning prompt that does both simultaneously, cutting per-chunk latency roughly in half without sacrificing output quality.
Why a sliding window chunker instead of splitting by headers?
We also implemented chunk_by_structure() (splits by # headers). But most Slack messages and informal docs don't have formal headers. The sentence-based sliding window handles messy, unstructured text far better across all 16 source types.
Why concept-overlap clustering instead of embeddings/vector search?
For an MVP, cosine similarity over embeddings requires storing large float arrays and running nearest-neighbor search. Instead, we cluster purely by the primary concept extracted by the LLM. It's O(N) instead of O(Nยฒ) and produces very interpretable clusters. Embedding-based clustering is a listed future improvement (embedding field already exists on ProcessedChunk, it's just null for now).
Why the interactive batch system?
Running a full sync across 16 connectors through a local 4B model can take a long time. Users need to be able to stop, inspect partial results, and decide whether to continue. The batch system also lets users validate quality early without committing to the full pipeline run.
Idempotency
Every raw item has a globally unique ID (e.g., github:myorg/repo:42, notion:abc123, slack-C05-1234). Before any processing, we query processed IDs from the DB and filter them out. Running sync 10 times on the same data produces the same result as running it once.
12. Directory Structure
company_brain_mvp/
โ
โโโ main.py โ CLI entry point (self-relative sys.path setup)
โโโ requirements.txt โ All Python dependencies
โโโ .env โ Your API keys (gitignored)
โโโ .env.example โ Template for .env with all 16 connectors
โโโ credentials.json โ Google OAuth credentials (gitignored)
โโโ token.pickle โ Saved Google OAuth token (gitignored)
โโโ test_integration.py โ Quick smoke test (chunker + LLM + synthesizer)
โ
โโโ src/
โ โโโ connectors/
โ โ โโโ base_connector.py โ Abstract BaseConnector: ingest(), retry_with_backoff()
โ โ โโโ registry.py โ Central connector registry: get_enabled_connectors()
โ โ โ
โ โ โ โโ Tier 1 Connectors โโ
โ โ โโโ slack_connector.py โ Slack: messages, thread replies
โ โ โโโ google_docs_connector.py โ Google Docs: paragraph text, tables
โ โ โโโ github_connector.py โ GitHub: issues, PRs, comments (REST API v3)
โ โ โโโ notion_connector.py โ Notion: page blocks, recursive children
โ โ โโโ discord_connector.py โ Discord: server channels, messages
โ โ โโโ google_sheets_connector.py โ Google Sheets: tabs as key-value row text
โ โ โโโ linear_connector.py โ Linear: issues, comments (GraphQL API)
โ โ โโโ outlook_teams_connector.py โ Outlook emails + Teams messages (MS Graph)
โ โ โ
โ โ โ โโ Tier 2 Connectors โโ
โ โ โโโ gitlab_connector.py โ GitLab: issues, merge requests, notes
โ โ โโโ dropbox_connector.py โ Dropbox: text file downloads
โ โ โโโ analytics_connector.py โ Mixpanel + Amplitude: event schemas
โ โ โโโ search_connector.py โ Algolia + Exa + Perplexity: search results
โ โ โโโ facebook_connector.py โ Facebook: page posts, comments
โ โ
โ โโโ processors/
โ โ โโโ chunker.py โ Sentence-based sliding window chunker
โ โ โโโ knowledge_extractor.py โ gemma4:e4b: single combined classify+extract call
โ โ โโโ skills_synthesizer.py โ Cluster chunks โ generate Skill cards
โ โ
โ โโโ models/
โ โ โโโ domain.py โ Pydantic models: RawDataItem, ProcessedChunk, Skill
โ โ DataSourceType enum (all 19 source types)
โ โ
โ โโโ storage/
โ โ โโโ database.py โ SQLAlchemy table definitions (3 tables)
โ โ โโโ storage_manager.py โ CRUD: save_raw_items, save_chunks, get_processed_ids
โ โ
โ โโโ sync/
โ โ โโโ sync_scheduler.py โ Orchestrates full & interactive sync via registry
โ โ
โ โโโ cli/
โ โโโ main.py โ Click commands: sync, export, status
โ โโโ ui/
โ โโโ animation.py โ Procedural neural network animation
โ โโโ dashboard.py โ Rich Live dashboard (SyncDashboard, SyncState)
โ โโโ logger.py โ Custom logging.Handler โ UI state bridge
โ โโโ components.py โ Rich tables/panels for status & export
โ
โโโ tests/
โ โโโ test_github_connector.py โ GitHub connector unit tests (3 tests)
โ โโโ test_notion_connector.py โ Notion connector unit tests (3 tests)
โ โโโ test_discord_connector.py โ Discord connector unit tests (3 tests)
โ โโโ test_google_sheets_connector.py โ Google Sheets connector unit tests (3 tests)
โ โโโ test_linear_connector.py โ Linear connector unit tests (3 tests)
โ โโโ test_outlook_teams_connector.py โ Outlook/Teams connector unit tests (3 tests)
โ โโโ test_gitlab_connector.py โ GitLab connector unit tests (3 tests)
โ โโโ test_dropbox_connector.py โ Dropbox connector unit tests (3 tests)
โ โโโ test_analytics_connector.py โ Mixpanel + Amplitude unit tests (2 tests)
โ โโโ test_search_connector.py โ Algolia + Exa + Perplexity unit tests (3 tests)
โ โโโ test_facebook_connector.py โ Facebook connector unit tests (3 tests)
โ
โโโ data/
โ โโโ company_brain.db โ SQLite database (auto-created on first run)
โ
โโโ output/
โโโ skills_file.json โ Final structured output for AI agents
โโโ skills_file.md โ Human-readable version
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file company_brain-2.0.0.tar.gz.
File metadata
- Download URL: company_brain-2.0.0.tar.gz
- Upload date:
- Size: 103.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad51de4d0489823015cce12019fff6b50684c25282c9031ce8056335f36bef86
|
|
| MD5 |
f52fd7635d8815f3e19b47739be6aae8
|
|
| BLAKE2b-256 |
37b26e5af1224242010f44ca8ba99240ad3ab8dfa2b48d6ddb85a4ee37cdfce0
|
File details
Details for the file company_brain-2.0.0-py3-none-any.whl.
File metadata
- Download URL: company_brain-2.0.0-py3-none-any.whl
- Upload date:
- Size: 136.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
325eb2c3617d17823ff85a7ad0cc20bb2603d08656a0a669714b11ea96f1d7ba
|
|
| MD5 |
c7bef21956d58eb13c2d16487c3f8e82
|
|
| BLAKE2b-256 |
99bf144fb3bb0a01fa6ccb07c3615e1c4747bca22080fdd436373a4d28187109
|