Global Think Tank Analyst (gtta)
An experimental strategic-risk reasoning framework for AI agents, with versioned memo contracts, CLI, and MCP.
Global Think Tank Analyst turns broad questions about policy, sanctions, regulation, trade, geopolitics, and strategic risk into decision-shaped memos. It makes evidence boundaries, assumptions, uncertainty, actor incentives, options, and watch indicators explicit.
Use the skill · Read the Russian version · See examples · Inspect project status
Current maturity: R3 / M3 / U0. A reproducibly tested package is
available from PyPI, alongside an executable method contract and disclosed
paired evaluations.
No practitioner validation or production reliability is claimed.
GTTA improves analytical structure; it does not establish factual truth. It is not legal, compliance, sanctions, financial, investment, or trading advice. Verify current sources and use qualified human review before acting.
Try it in one prompt
Attach SKILL.md to a capable agent, then paste:
Use Global Think Tank Analyst.
Question: What does regulatory uncertainty change for our market-entry decision?
Decision this informs: enter now, run a limited pilot, or wait.
Audience: operating committee.
Geography: [countries or markets].
Time horizon: 12 months.
Evidence mode: reasoning-only unless live sources are available.
Depth: standard memo.
Separate facts, assumptions, assessments, scenarios, and unknowns.
Tag every material claim with its provenance.
Give options, trade-offs, indicators, confidence, and what would change the judgment.
The skill works without the Python package. This is the simplest and most mature way to use the method.
What it does
- Frames analysis around a concrete decision, audience, geography, and time horizon.
- Separates facts, assessments, assumptions, scenarios, and unknowns.
- Uses per-claim provenance tags:
[primary],[secondary],[user-provided],[inference], and[analyst-judgment]. - Calibrates language to evidence and confidence.
- Models actors, incentives, leverage, options, trade-offs, scenarios, and observable triggers.
- Supports seven response modes, from a quick brief to competing hypotheses.
- Exposes the same method through agent instructions, Python, CLI, and MCP.
- Provides deterministic checks for method structure and a strict structured memo artifact for machine-readable workflows.
What it is not
- Not a source-retrieval system or real-time intelligence feed.
- Not a factuality verifier.
- Not an autonomous decision-maker.
- Not a substitute for legal, sanctions, compliance, financial, or domain review.
- Not externally practitioner-validated or production-proven.
- Not a generic multi-agent platform; the optional LangGraph pipeline is an experiment around the core reasoning method.
Install and use
Use the instructions directly
Add AGENTS.md and SKILL.md to an agent workspace,
or attach SKILL.md to a conversation. English and Russian instructions are
both packaged in the wheel. SKILL.md remains canonical; SKILL_RU.md is a
full 45-section Russian rendering with Mode A–G and bilingual canonical output
markers so Russian memos remain compatible with the method checker. Structural
parity is enforced by scripts/validate_language_parity.py.
Install the developer toolkit from source
The stable package is available from PyPI. A source checkout remains useful for developing the method, examples, and integrations.
git clone https://github.com/vassiliylakhonin/global-think-tank-analyst.git
cd global-think-tank-analyst
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[mcp,verification]"
# Generate a memo scaffold
gtta new --mode B --topic "Market-entry regulatory exposure"
# Heuristically check a Markdown memo
gtta check-contract memo.md --mode B
gtta check-contract memo.md --mode B --format sarif --out gtta.sarif
# Inspect, validate, and render the strict structured artifact
gtta artifact-schema
gtta source-catalog-schema
gtta check-artifact memo.json --json
gtta render-artifact memo.json > memo.md
# Project source-backed claims through Agenda Intelligence and fail closed
gtta verify memo.json --strict
gtta verify memo.json --strict --format html --out review.html
gtta verify memo.json --strict --format sarif --out verification.sarif
gtta verify memo.json --strict --repair-prompt repair.md
# Write the complete strict review contract to a new, auditable directory
gtta review memo.json --out-dir memo.review
gtta check-review-bundle memo.review --json
# Serve the method and artifact tools over MCP stdio
gtta mcp
The current release is
v1.8.0:
pip install global-think-tank-analyst
It includes the native Agenda Intelligence verification seam, bounded repair
guidance, SARIF output, and a portable verification benchmark while preserving
the gtta.memo@1.0 contract. Version 1.8.0 adds gtta review, a single strict
orchestration command that atomically writes the versioned review bundle
documented in docs/review-bundle.md, plus the
Agenda-independent gtta check-review-bundle integrity checker. See
STATUS.md for publication state and limitations.
Executable analysis contracts
GTTA separates checks that answer different questions:
| Layer | Interface | What it can establish | What it cannot establish |
|---|---|---|---|
| Markdown method preflight | gtta-method-contract@1.x |
Required declarations, mode shape, confidence, likely untagged claims, generic advice | Claim boundaries, factuality, source support |
| Structured memo | gtta.memo@1.0 |
Claim IDs, provenance, source references, dependency links, mode invariants, canonical rendering | Whether a named source is trustworthy or supports the claim |
| Memo verification | gtta.memo-verification@1.0 + Agenda Intelligence MD |
Native MemoArtifact projection, claim/source packet completeness, declared quotes, lexical support, unmatched numbers | Factual truth or professional approval |
| Memo repair plan | gtta verify --repair-prompt |
Bounded claim-specific repair instructions that preserve unresolved evidence gaps | Source discovery, automatic factual correction, clearance |
| Review bundle integrity | gtta.review-bundle-check@1.0 |
Exact file contract, SHA-256 matches, receipt/SARIF shape, internal status consistency | Authorship, authenticity, factual truth, trusted attestation |
| Operational decision | Human review | Contextual judgment, current-source verification, accountability | Guaranteed correctness |
MemoArtifact is the canonical machine-readable GTTA seam. Its claim ledger is
shared by the Python API, CLI, and MCP tools; Markdown is the rendered human
view.
Verification findings can be emitted as SARIF 2.1.0 with claim-level JSON line
locations. The portable six-case benchmark is available through
python scripts/run_verification_benchmark.py; CI stores its receipts and
uploads regression errors to GitHub Code Scanning. Expected negative-control
findings remain in the JSON receipt and do not masquerade as production defects.
Memo modes
| Mode | Use it for | Required shape |
|---|---|---|
| A — Quick Brief | Fast orientation | Bottom line, risks, watch indicators, confidence |
| B — Standard Memo | Default decision analysis | Context, actors, assessment, options, change conditions |
| C — Scenario Brief | Divergent futures | Baseline, scenarios, triggers, implications, indicators |
| D — Red-Team Challenge | Stress-testing a claim | Target claim, alternatives, failure modes, revised judgment |
| E — Decision Pack | Team action | Memo, options, watchlist, owner questions, next steps |
| F — Analyst Training | Developing reasoning | Coaching and Socratic challenge rather than a finished answer |
| G — Competing Hypotheses | Attribution and rival explanations | Hypotheses, evidence matrix, disconfirmation, sensitivity, bounded judgment |
Before and after
A generic answer:
The environment is uncertain. Monitor developments, engage stakeholders, remain agile, and review the strategy regularly.
A GTTA-shaped answer:
Decision: authorize a limited pilot or wait for regulatory clarity. Evidence mode: reasoning-only.
[analyst-judgment]Prefer a reversible pilot because it buys operating information without committing the full rollout budget.Main downside: delay and duplicated setup cost. Trigger to pause: the regulator expands the authorization requirement to cover the pilot itself. Confidence: Moderate. What would change the judgment: evidence that the pilot creates the same irreversible exposure as a full launch.
The difference is not a more confident tone. It is a visible decision frame, evidence boundary, trade-off, trigger, and revision condition.
How the portfolio composes
GTTA owns the horizontal reasoning method. Regional depth and evidence-packet checks stay in separate repositories.
flowchart LR
Q[Decision question] --> G[GTTA<br/>reasoning method]
V[Optional regional specialist] --> G
G --> M[MemoArtifact / Markdown memo]
M --> A[Agenda Intelligence MD<br/>evidence-packet checks]
A --> H[Qualified human review]
| Layer | Repository | Responsibility |
|---|---|---|
| Horizontal method | Global Think Tank Analyst | Decision framing, memo modes, uncertainty, scenarios, options |
| Central Asia depth | Central Asia + Caspian skill | Regional mechanisms, corridors, banking, sanctions adjacency |
| Gulf depth | Gulf + Middle East skill | Gulf banking, energy, maritime chokepoints, Iran-related risk |
| Evidence packet | Agenda Intelligence MD | Deterministic claim/source packet checks |
See PORTFOLIO.md and the
evidence-packet handoff for the full seam.
Integration status
| Surface | Status | Entry point |
|---|---|---|
| Agent instructions | Core | AGENTS.md, SKILL.md, SKILL_RU.md, llms.txt |
| Python artifact API | Core development interface | gtta.MemoArtifact, check_memo_artifact(), render_memo_artifact(), verify_memo_artifact() |
| CLI | Tested | gtta new, check-contract, check-artifact, render-artifact, verify |
| MCP server | Tested optional extra | python -m pip install -e ".[mcp]", then gtta mcp |
| LangChain / LlamaIndex adapters | Optional | .[langchain] or .[llamaindex] |
| LangGraph draft-and-critique pipeline | Experimental | .[agent] |
| FastAPI / Streamlit | Local experiments | .[enterprise,ui]; not a production deployment architecture |
Examples
Use examples/README.md as the complete learning path.
Start with these:
| Goal | Evidence mode | Example |
|---|---|---|
| Learn the basic memo shape | reasoning-only |
Sanctions exposure memo |
| See explicit public-source boundaries | live-source-backed |
OFAC case memo |
| See a narrow retrieval boundary | live-source-backed |
Middle Corridor logistics risk |
| Work from supplied documents | user-provided sources |
Supply-chain sanctions exposure |
| Surface conflicting sources | illustrative source packet |
IEA–OPEC forecast conflict |
| Challenge an existing claim | reasoning-only |
Red-team policy brief |
Every example declares its evidence mode. Source-backed examples are snapshots; verify their retrieval dates and current facts before use.
Evaluation and maturity
The repository contains:
- deterministic regression tests for the CLI, MCP, method checker, and structured artifact;
- an installed-wheel smoke test;
- human review checklists and failure modes under
evals/; - a predeclared 12-case same-task, with/without-skill structural harness under
evals/agent-eval/with an offline Antigravity export/import path and no model API client; - a versioned declared-behavior extension that keeps structural and behavioral pass rates separate and uses frozen case-specific expectations to move beyond schema-only conformance;
- four published Antigravity runs: two Gemini executions, the
seed 20260830 runand a freshness-gatedseed 20260831 replication, and two Claude executions, the freshness-gatedseed 20260901 cross-model runand fresh post-changeseed 20260902 replication; each publishes exact requests, outputs, recorded settings, hashes, mapping, and a deterministic report.
All four completed runs found a 12/12 contract pass rate with the skill and
0/12 for the generic baseline. The Claude run extends the result to a second
model family, while its original 181 capped skill warnings expose materially
weaker per-claim provenance compliance than the Gemini runs. A narrow
gtta-method-contract@1.2.2 precision rescore reduces that stored count to
170 without changing outputs. The fresh post-change Claude replication stores
111 skill warnings, but three samples hit the warning cap, so this is
directional evidence rather than a precise causal improvement estimate. These
author-operated runs support only a bounded structural-discipline claim: the
scorer does not assess factuality, source support, decision quality, or
practitioner usefulness. Practitioner review remains U0.
The first preregistered
structured declared-behavior run
moved beyond the schema-only ceiling: both arms passed 12/12 structural checks,
while the skill arm passed 8/12 frozen behavior expectations versus 3/12 for
baseline. The observed +41.7 point difference applies only to model-declared
artifact fields in one Gemini execution. It is not a factuality,
reasoning-quality, causal, or practitioner-usefulness score.
The subsequent Claude Code / Opus 4.6 replication preserved the direction but not the magnitude: 3/12 skill versus 0/12 baseline declared-behavior passes, with 11/12 versus 12/12 structural passes. This is cross-model-family structural evidence, but the low absolute pass rate and one skill invariant failure argue against further headline-score optimization on the same cases.
The preregistered
Gemini 3.8 broader-domain holdout
then passed strict structure 10/10 in both arms but passed zero complete
declared-behavior expectations in either arm. Missing verify: true
declarations dominated. This null result does not reproduce the original
suite's positive combined-pass difference and is published with its execution
qualifications rather than tuned away.
Read STATUS.md for current evidence,
docs/maturity-framework.md for the independent
release/method/usefulness axes, and
docs/definition-of-done.md for claim-specific
release gates.
Signal archive
signals/ contains compact examples of the method style. It is not
a live intelligence service.
- Latest signal:
signals/latest.md - Machine-readable index:
signals/index.json - JSON Feed:
signals/feed.json - Contribution template:
signals/TEMPLATE.md
Re-verify every cited fact before operational use. Any signal can be expanded by running its example prompt through the skill.
Agent-readable endpoints and naming
AGENTS.md— repository-wide agent contractSKILL.md— canonical English runtime instructionsSKILL_RU.md— full Russian runtime instructions (45 of 45 sections, Mode A–G)codex/SKILL.md— Codex-ready variantllms.txt— orientation for agents and indexersGlobal Think Tank Analyst— project and horizontal skillPolicy Risk Memo Architect— analytical method implemented by the skillMemoArtifact— versioned machine-readable memo interface
Repository structure
.
├── AGENTS.md # Repository contract for agents
├── SKILL.md / SKILL_RU.md # Canonical runtime instructions
├── STATUS.md # Current R/M/U evidence
├── src/gtta/artifact.py # MemoArtifact schema, validation, rendering
├── src/gtta/discipline.py # Markdown method-contract preflight
├── src/gtta/cli.py # CLI adapters
├── src/gtta/mcp_server.py # MCP adapters
├── docs/ # Contracts, handoffs, release guidance
├── examples/ # Worked memos and evidence modes
├── evals/ # Review material and structural harness
├── signals/ # Public style examples and feeds
└── tests/ # Runtime and contract regression tests
Limitations
- GTTA does not retrieve or continuously refresh sources.
- It does not decide whether a source is independent, authoritative, or sufficient for a specific claim.
check-contractuses Markdown heuristics; useMemoArtifactfor exact declared claim accounting.check-artifactvalidates structure and cross-references, not truth.- The agent pipeline, API, UI, batch jobs, memory, knowledge-graph drafts, and document parsing are experiments, not the release target.
- There is no labeled factual-accuracy benchmark, long-horizon agent trial, or recorded external practitioner review.
Roadmap
- Keep
gtta.memo@1.xandgtta-method-contract@1.xstable, including the machine-readable warning-truncation telemetry added in ruleset 1.2.3. - Freeze the completed Gemini and Claude declared-behavior results; do not tune the method or rubric against repeated runs on the same cases.
- Freeze the completed broader-domain holdout and its null result; do not tune the method or thresholds against those cases.
- Keep the stable
1.7verification seam, source catalog, SARIF mappings, and bounded-repair behavior regression-tested. - Complete PyPI Trusted Publishing after account access is restored.
- Record real practitioner feedback if access becomes available; do not use
proxy metrics to disguise
U0.
Contributing
Read CONTRIBUTING.md, then run:
python3 scripts/check.py
Package changes should also pass the test suite, wheel build, and installed wheel smoke test. Issues and pull requests are welcome.
License
MIT — see LICENSE.
Metadata
Release files for global-think-tank-analyst 1.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| global_think_tank_analyst-1.8.0.tar.gz | 95.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| global_think_tank_analyst-1.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 178.5 kB
Release files / global_think_tank_analyst-1.8.0.tar.gz
| Download URL | global_think_tank_analyst-1.8.0.tar.gz |
|---|---|
| Size | 95.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5c188e3f19fee570074641f6b9069216d99115d043fba1c6d825a2e3711c8c01
|
|
BLAKE2b-256 checksum How to use checksums |
d7cb8ee95283ea050c3d72616ad3acb43c52823da398f261b8a23633d2278eef
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / global_think_tank_analyst-1.8.0-py3-none-any.whl
| Download URL | global_think_tank_analyst-1.8.0-py3-none-any.whl |
|---|---|
| Size | 82.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
78997a5c5284b5494ee3827d432768268ea23faaf87e121cc756654ab5b8314c
|
|
BLAKE2b-256 checksum How to use checksums |
96462da5d7a63ebffdec80b00dbb66bdccdb43f017c829c2bd87e80ee2ae4b01
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|