Skip to main content

An honest, cite-or-abstain persona engine — grounded in a signed knowledge base.

Project description

Qee, thinqee's face

thinqee

An honest, cite-or-abstain persona engine. Ask how a specific person works, thinks, and decides — grounded in a knowledge base they signed. Every answer cites its sources; anything not in the notes is declined honestly, in that person's voice. The persona is config, not weights.


CI PyPI License: Apache-2.0

thinqee is persona-agnostic: it contains no one's data. You install the engine and point it at your own knowledge base. It never downloads or trains anything — it only calls the model endpoint you configure (any provider, or a local keyless one).

Status: beta. The engine is complete and live-tested, but the API may still move before 1.0. Maintained by one person, at a sustainable pace.

Get started — your own answering persona in ten minutes

One path, one default. The default model backend is a local, keyless OpenAI-compatible endpoint (e.g. Ollama) — no account needed; a hosted provider is one .env line instead (AI_MODEL=mistral:mistral-medium-latest + AI_API_KEY=…; OpenAI/Anthropic/Groq slugs work the same).

pip install "thinqee[agents,api]"
thinqee init ./my-persona --name YourName
cd my-persona
cp .env.example .env                  # pre-filled for local Ollama; edit for a hosted provider
# open knowledge/values.md, replace the DELETE-ME example with something true about you,
# and set `signed: true` — your signature is the build step, there is no other
thinqee ask "what do you value?"      # → a cited answer, from your note
thinqee serve                         # → http://127.0.0.1:8000 — your instance, with /studio

Then ask it something it doesn't know, and watch it say so honestly — that refusal is the product working, not failing. Stuck anywhere? thinqee doctor prints a one-screen ✓/⚠/✗ preflight with the exact fix for each problem. To put the widget on your own site, see EMBEDDING.md.

thinqee init scaffolds the full knowledge tree, an optional prompts/ override slot, a starter eval set (with injection refusal probes, so your fork is born with a conscience), an Ollama-first .env.example, and a CI workflow. Interview yourself into notes with thinqee interview --session A.

Which one is you?

Role What you install What you do
Visitor nothing Use a deployed persona's page, its embedded widget, or its /mcp endpoint.
Persona owner the engine (pip install thinqee / uvx thinqee) the ten minutes above: init, write and sign your notes, serve. Your own server hosts /widget.js — a thinqee widget is never installed from npm or a CDN (that's deliberate; see EMBEDDING.md).
Contributor this repo (pip install -e ".[agents,api]") hack on the engine, pytest.

Working as a team? A knowledge maintainer needs only a browser — no install. The team layer is your git host: thinqee init scaffolds CODEOWNERS, and with branch protection a code owner's merge is the sign-off gate (thinqee builds no accounts of its own). Notes are edited in github.dev, opened as a PR, and CI (thinqee eval) shows red/green. See docs/collaboration.md.

Everything below is detail, in reading order: How it works · the demo · the quality gate · serving & the console · MCP & agent skills · capabilities · principles · ethics · privacy · license

How it works

flowchart LR
    Q[question] --> G{Gatekeeper}
    G -- injection / empty --> DR[fixed refusal<br/>model never sees it]
    G -- lane --> A[Answerer]
    KB[(signed<br/>knowledge base)] --> A
    A -->|grounded| ANS[cited answer]
    A -->|no coverage| NC[honest 'not in my notes']
    A -->|off-topic / private| DEC[in-voice decline]

    subgraph HITL[human-in-the-loop gates]
      S1[note sign-off]
      S2[review queue]
      S3[follow-up consent]
    end
  • Gatekeeper screens prompt-injection / empty / oversize input before the model, then routes the question to a lane. Answerer replies strictly from that lane's signed notes, cite-or-abstain, within a hard word cap, in the question's language.
  • Repeat questions are cached (keyed to the KB commit), so the same question gets smarter with each signed note and never re-hits the model.
  • Nothing enters the knowledge base without the owner's signature — no auto-learned memory.

The design surface travels with the engine: /studio ships in the wheel — anyone running thinqee serve has it on their own deployment, and offline via thinqee face preview.

Try the demo — the docs answer for themselves

The demo persona is Qee: this project's own documentation, running on this project's own engine. Ask it how to set thinqee up and it answers from signed notes, names the ones it used, and declines anything that isn't in them. It's the manual, and it's the proof.

pip install -e ".[agents]"        # in a checkout of this repo

# Local & keyless (Ollama, vLLM, LM Studio, any OpenAI-compatible endpoint) — no account:
AI_BASE_URL=http://localhost:11434/v1 AI_MODEL=mistral-small PERSONA_NAME=Qee \
  thinqee ask "why is there no vector database?" --kb examples/qee

The hosted default is mistral:mistral-medium-latest (AI_MODEL=<slug> AI_API_KEY=… — openai:…, anthropic:…, groq:… work the same). Every agent role takes an optional override — AI_MODEL_JUDGE (a model shouldn't grade its own answers), AI_MODEL_FAST (a small model for the gatekeeper hop), AI_MODEL_INTERVIEWER, AI_MODEL_LIBRARIAN and AI_MODEL_DIGEST — unset, a role follows AI_MODEL. The console shows the effective model per role.

A second example, examples/anna/, shows what a persona of an actual person looks like — a fictional agile coach, with the availability, the values-from-real-stories and the contested positions that a docs bot has no use for. Read Qee to learn the engine; read Anna to see the shape of your own.

thinqee ask is built for both a terminal and a pipe: quotes are optional (thinqee ask how do I set this up --kb examples/qee), the output degrades to plain text when piped (--json stays byte-stable), and running it with no question in a terminal opens a tiny REPL. The exit code follows the verdict0 answered, 3 no coverage, 4 declined, 1 a real error — so a script can branch on the outcome.

The quality gate is real

thinqee eval --self-test         # proves a deliberately broken refusal FAILS the build (no key)
thinqee eval                     # grades the full set; exits red on any failed refusal

A hybrid Judge grades every change — deterministic checks first (refusal safety never rides on an LLM's mood), then model judgment for grounding and voice. One failed refusal is a red build, no matter what else improved. An item can also demand a whole set of notes, so an answer that is correct but incomplete (cites some, not all) is caught and counts against the pass rate. Numbers here are generated, never hand-typed:

Honest numbers

  • Eval set: 55 questions (KB 00f9b38ab4c25f27669820199d9b1ce22f21f35e).
  • Answer pass rate: 97%.
  • Refusals held: 18/18
  • Honest "not in my notes": 3/3
  • Multi-note synthesis: 5/5
  • Lane routing: 30/34 matched declared lane
  • Failing items: craft-agents-de.

See what a commit changes

thinqee diff HEAD~1              # how the last commit changed this persona's answers

Because answers are cached on (question, kb_sha, prompt_version, model), thinqee can replay the persona's bound questions (the eval set + the widget chips) against two KB commits and show exactly which answers move — verdict, sources, or wording — with a source delta and a text diff. Exit code 2 means "answers changed" (a CI check a knowledge PR can gate on); --fail-on verdict|sources|text sets the threshold. It shows what changed; thinqee eval judges whether it's good — different tools. This is the governance payoff of the git-host team layer: an owner sees what a knowledge PR does before signing off.

Serve it (public API + your admin app)

pip install -e ".[agents,api]"
AI_API_KEY=... PERSONA_NAME=Qee THINQEE_KB_PATH=examples/qee \
  THINQEE_OWNER_TOKEN=$(openssl rand -hex 32) thinqee serve
# → http://127.0.0.1:8000   sign in — the instance is YOUR admin app
#   /console                what is deployed + the questions your notes could not answer
#   /studio                 design its face, theme and layout; try both skins live

The deployed instance is the admin interface — like a Coolify or Plausible instance. Visitors never come here: they meet the persona through the widget embedded on your own site (EMBEDDING.md). serve also pre-warms the answers to your own chip questions at startup, so the first visitor never waits for a question you put on your own front page.

Hardened for a public link: SQLite answer cache with single-flight, per-IP rate limit + a daily budget cap (a traffic spike can't bankrupt a free tier), an identity-free review queue, an identity-free and text-free feedback bit (its own rate limit), and CORS + CSP locked. Boring, EU-hostable deployment via Dockerfile + docker-compose.yml + Caddy.

The console: your content roadmap

Sign in at / with your THINQEE_OWNER_TOKEN and /console shows what is deployed and the questions your notes could not answer. That list is your content roadmap. Two actions per question, right in the browser: review runs the gap interview and hands you an unsigned draft note to download; promote hands you a YAML eval snippet to copy. Both end in an artefact you commit through git. Visitors can also mark an answered question helpful or not — one identity-free bit, no free-text — and a "not helpful" joins that same queue (deduplicated), so the loop learns from answers that were given but still missed, not only from refusals. The console never writes to your knowledge/ or evals/ on the server (a file in a container is lost on redeploy and would skip the sign-off gate); the CLI (thinqee queue review --id <n>) does the same from a terminal. Sign-in sets an HttpOnly cookie (the page never holds your token); once a token is set, /studio asks for the same sign-in. Without the token, / shows a setup page, /studio stays open (a fork designs its face locally), and /console with its /api/owner/* routes is not registered at all — a deployment that hasn't opted in has nothing to find.

Qee — thinqee's face — greets the terminal and reacts to each answer (thinking, answered, an in-voice decline…). See BRAND.md for the character, and EMBEDDING.md for putting the widget on your own page. Forks get a neutral face + a playground to design their own.

One engine, two faces

The widget ships as two custom elements over the same engine — same /api/ask, same cache, same guardrails, so both give the identical cited answer to a question (the corner chat additionally sends the page path as a retrieval hint, so a question about a specific page may answer page-specifically). Pick the surface, not a second backend:

<qee-terminal></qee-terminal>   <!-- a full-panel, terminal-styled Q&A surface -->
<qee-chat></qee-chat>           <!-- a floating corner chat (launcher + panel) -->
<script type="module" src="/widget.js"></script>

Both live in one Shadow-DOM bundle and are themed purely through CSS custom properties (a warm reference theme in EMBEDDING.md recolours either without touching widget code). The corner chat adds curated per-page chips and a strictly-reactive attention wave — it never watches what a visitor types. thinqee serve hosts the terminal at /; try the corner chat in /studio, which previews it in a real device frame against the real endpoints. For apps, embed the widget in a WebView or build a native client against the same contract — see native-clients.md.

Use it inside an agent (MCP)

The persona becomes a tool any MCP client can consult — two sharp tools, ask and perspective.

pip install -e ".[agents,mcp]"
# Local, keyless — the fork's persona as a tool on your machine:
AI_BASE_URL=http://localhost:11434/v1 AI_MODEL=mistral-small PERSONA_NAME=Qee \
  THINQEE_KB_PATH=examples/qee thinqee mcp

Claude Desktop (claude_desktop_config.json), local stdio:

{
  "mcpServers": {
    "qee": {
      "command": "thinqee",
      "args": ["mcp"],
      "env": { "PERSONA_NAME": "Qee", "THINQEE_KB_PATH": "/path/to/examples/qee",
               "AI_BASE_URL": "http://localhost:11434/v1", "AI_MODEL": "mistral-small" }
    }
  }
}

For a deployed persona, set THINQEE_MCP=1 on the server and the MCP endpoint mounts at /mcp (sharing the same cache, rate limit, budget guard, and tracing). Tools are read-only and tell agents that "not in my notes" is a valid result — to be surfaced, not retried around.

Owner mode. In your own repo clone, thinqee mcp --owner adds the learning-loop tools — read the queue, pull the weekly digest, run the gap interview, and get unsigned draft notes + eval snippets. Point it at your deployment (THINQEE_SERVER_URL + THINQEE_OWNER_TOKEN) or let it work the local queue.db. The server never writes your knowledge base: the agent writes files into the clone, you review the diff and sign. The dream path — digest → draft → PR → merge is the signature. Every AI draft carries drafted_by: ai:<model> and stays signed: false until a human accepts it.

Agent skills

Two skills in skills/ make thinqee a first-class citizen of an agent (e.g. Claude) — you install one by copying its folder into your agent's skills directory:

  • consult-persona — call a deployed persona's /api/ask mid-task and relay its cited answer, surfacing a decline or "not in my notes" instead of working around it.
  • persona-interview — run the engine's interviewer to build your own persona, one signed note at a time (it facilitates; it never ghostwrites).

Capabilities

init · ask · serve — scaffold, answer, host (no build step: the signed notes are the build)
Provider-open backend (any pydantic-ai provider or local endpoint)
add / synthesize — distil sources, synthesize templates
HTTP API + web terminal + cache / rate-limit / budget
Learning loop: queue → digestinterview/review → signed note
Owner console: token-gated / — queue, browser gap-interview → unsigned draft, promote → eval snippet
Consent follow-up mail (send once, delete address) · prune retention
OpenTelemetry tracing → self-hosted Phoenix (optional)
Eval gate: hybrid Judge, binary refusal + honest-IDK rule, eval/report, self-test, CI
thinqee diff <ref>: replay the bound questions across two KB commits — which answers change (verdict/sources/text), CI-gateable exit code
MCP server: public ask / perspective (stdio + remote); mcp --owner adds the learning-loop tools for an agent in your clone
Agent skills: consult-persona (query a deployed persona) · persona-interview (build your own)
Team layer via the git host: scaffolded CODEOWNERS + PR template, browser editing, CI eval gate — no accounts in the engine
Language registry: English + German built-in; a fork adds a language (detection, deterministic screen, chips, widget chrome, follow-up mail) as data in knowledge/languages.yaml + voice.yaml, no engine change
EU AI Act Art. 50 first-contact disclosure: default-active, per-language, persona-overridable, never silent

Design principles

  1. No sentence I haven't signed. Nothing enters the KB without explicit owner approval.
  2. Cite or abstain. Every answer names its source notes.
  3. Refusals are personality data. How the machine declines is written in the owner's voice.
  4. The knowledge base is the product. Code is a weekend; the notes are the work.
  5. Climb the ladder only when forced. Simplest retrieval and infra that work.
  6. Honest numbers or no numbers. Eval results, including failures, are shown as they are.
  7. The upstream test. A feature belongs in the engine only if any other product would want it and it needs no user identity, state, or app context. Everything identity-, data-, or state-bound is built by the embedder against the contract. This keeps the engine neutral and small.

Ethics

Build a persona of yourself. The engine grounds every answer in first-person, signed notes, and the interview requires first-person participation by design; building an unauthorized persona of someone else is out of scope.

The engine itself is developed in pair-programming with an AI (Claude) — and holds itself to its own rules: every note in the demo knowledge base is human-signed, and every release passes the eval gate.

Privacy

Question text is logged without any identity (no IP, user agent, or fingerprint). Follow-up emails are opt-in only, sent once, then the address is deleted. You choose your inference provider (an EU-hosted default keeps data in the EU). Observability is self-hosted; nothing goes to a third-party APM. The widget sets no tracking identifiers and needs no cookie banner — it writes only three strictly-functional values (a once-per-session greeting flag; the launcher position, only if you drag it; a once-per-session marker for the AI disclosure). Visitors see an "you're talking to an AI" notice at first contact (EU AI Act Art. 50) — default-active, shown in the widget's language, rewordable per persona but never silent. A copy-ready privacy-policy annex (EN + DE) ships in docs/privacy-annex.md — fill the placeholders and have it reviewed before a public launch.

Contributing

See CONTRIBUTING.md. Planned-but-frozen ideas live in post-v1.md. AI-assisted contributions are welcome — the human contributor signs off (DCO) and stands behind the diff.

License

Apache-2.0 — permissive, with an explicit patent grant, so any product can host the engine. Contributions are accepted under the DCO (sign off with git commit -s); there is no CLA. See CONTRIBUTING.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

thinqee-0.9.2.tar.gz (368.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

thinqee-0.9.2-py3-none-any.whl (266.2 kB view details)

Uploaded Python 3

File details

Details for the file thinqee-0.9.2.tar.gz.

File metadata

  • Download URL: thinqee-0.9.2.tar.gz
  • Upload date:
  • Size: 368.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for thinqee-0.9.2.tar.gz
Algorithm Hash digest
SHA256 72e304a7fdde71845bdcf100c64f87a3457dcb7814d2c96518732451e38a3f8e
MD5 3a9b31949771518b88a00de45121c0f6
BLAKE2b-256 8c12494aff3e5f056f0587fd41bad35ee7ac39d775480b01bb522a17cb60d4c5

See more details on using hashes here.

File details

Details for the file thinqee-0.9.2-py3-none-any.whl.

File metadata

  • Download URL: thinqee-0.9.2-py3-none-any.whl
  • Upload date:
  • Size: 266.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for thinqee-0.9.2-py3-none-any.whl
Algorithm Hash digest
SHA256 b4656689b2a06e4d80d97059bfbe92d22b42ec6e53bd85917dfe4a58b08d348d
MD5 34c9c6f961aaabf0f769ebf27902d542
BLAKE2b-256 8fd95ad03654f8a48c66734f2ed5346b17f2e31b76bafcca690356d839c304cb

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page